AI/ ai · healthcare · llms · research

Feeding AI More Patient History Doesn't Make It Smarter

A new benchmark study finds that how clinical AI models select and organize a patient's medical history matters more than how much they can read at once.

Longer context windows don't automatically make AI better at tracking a patient's medical history.

Researchers evaluated four open-weight language models on MedLoCoMo, a benchmark for longitudinal clinical reasoning, testing five different ways of feeding those models a patient's history. One method just stuffs the full record into the context window; others select recent events, group history into episodes, extract semantic facts, or blend several of those approaches into a hybrid. The episodic and hybrid strategies produced the most accurate answers overall, and held up best when the evidence a question depended on was buried far back in the record. The strategy that leaned on recent events degraded the fastest as relevant information got older.

That's a problem for the industry's current pitch that bigger context windows solve long-document reasoning: more history available doesn't mean more history gets used correctly, and in a clinical setting, missed evidence is a patient-safety issue, not just a UX quirk. Just as concerning, the study found that models which answered supported questions well didn't reliably recognize when a question had no real answer in the chart, meaning confident-sounding wrong answers remain a live risk even in well-performing systems.

Context length is a spec sheet number; knowing where to look in a patient's chart is a different skill, and this paper suggests most models still haven't learned it.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →