AI/ ai · healthcare · llm-reliability · clinical-ai

Study Finds AI Models Struggle to Update Patient Risk Estimates

New research shows large language models often double down on prior judgments as ICU patient data changes, and prompting doesn't fix it.

A new study finds that AI models asked to track a patient's condition over time often get worse, not better, at predicting what happens next.

Researchers tested several large language models on real intensive-care patient histories pulled from electronic health records, asking each model to update its risk estimate as new evidence came in. Telling a model its own prior assessment before showing it new data made predictions less accurate more often than it helped, a pattern that held across two separate outcomes. The models also reacted harder to bad news than to equally strong good news about a patient's breathing, and shifting a model's starting assumption about risk from 10 percent to 90 percent swung its final estimate by 26.2 percentage points, even when the actual evidence given was identical. Rewording the prompts did not fix any of this; only a filtering method the researchers built, called Evidence-Validated Longitudinal Update, produced fewer but more trustworthy revisions.

Hospitals are already testing LLMs to help flag patients who are getting worse, and clinical judgment is supposed to update cleanly as new vitals and labs arrive. This research suggests the models carry their own stubborn priors into that process, the same anchoring bias that worries us in overworked residents, except harder to audit from the outside. That is a narrower, more specific problem than generic AI errors, and one that deployed systems will need to test for directly rather than assume away.

A chatbot that gets more confident the longer it is wrong is not a new assistant for the ICU. It is a familiar failure mode wearing a lab coat.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →