Longer context windows don't automatically make AI better at tracking a patient's medical history.
Researchers evaluated four open-weight language models on MedLoCoMo, a benchmark for longitudinal clinical reasoning, testing five different ways of feeding those models a patient's history. One method just stuffs the full record into the context window; others select recent events, group history into episodes, extract semantic facts, or blend several of those approaches into a hybrid. The episodic and hybrid strategies produced the most accurate answers overall, and held up best when the evidence a question depended on was buried far back in the record. The strategy that leaned on recent events degraded the fastest as relevant information got older.
That's a problem for the industry's current pitch that bigger context windows solve long-document reasoning: more history available doesn't mean more history gets used correctly, and in a clinical setting, missed evidence is a patient-safety issue, not just a UX quirk. Just as concerning, the study found that models which answered supported questions well didn't reliably recognize when a question had no real answer in the chart, meaning confident-sounding wrong answers remain a live risk even in well-performing systems.
Context length is a spec sheet number; knowing where to look in a patient's chart is a different skill, and this paper suggests most models still haven't learned it.