AI/ ai · healthcare · llms · research

AI Model Compresses Diagnosis Codes to Catch More Conditions

A new training method rewards AI models for catching every likely diagnosis in a patient's record, not just the easiest one.

A new AI framework rewards language models for catching every likely diagnosis in a patient's chart, not just the one that's easiest to guess.

Researchers built a system called CARing that reworks how large language models predict a patient's next diagnoses from electronic health records. Standard reinforcement learning rewards a model once it gets the right answer in a reasoning chain, but patients often have several valid diagnoses at once, so a model can rack up points by repeatedly nailing the same easy one while ignoring the rest. CARing compresses ICD diagnosis codes into compact "semantic ID" tokens so the model can reason about diseases more directly, then adds a coverage reward that only pays off when the model spreads its guesses across different correct diagnoses. Tested on the MIMIC-III and MIMIC-IV hospital record datasets, it beat other EHR-trained models on weighted F1 and recall, hitting a top-30 recall of 46.04% and 46.52% in its reasoning mode.

Most clinical AI benchmarks reward a single correct answer, which is a poor match for real medicine, where patients routinely carry multiple active conditions. A model trained that way can look accurate on paper while quietly missing comorbidities that matter for treatment. Fixing the reward function, rather than just scaling up the model, is a cheaper and more targeted way to close that gap.

Even so, a top-30 recall near 46% means the model is still wrong more often than right about which conditions actually apply - a reminder that diagnosis prediction remains a research exercise, not something ready for a doctor's desk.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →