A team of researchers built a system to catch the data errors that make AI health predictions look better than they actually are.
EHR2Trace converts electronic health records from different hospital systems into a standardized, traceable format for training so-called patient world models, AI systems meant to predict how a patient's condition will change over time. The tool separates when a clinical event happened from when that information became available to clinicians, and it distinguishes a medication being ordered from it actually being dispensed and administered, details that often get flattened together in raw records. Tested across three clinical datasets covering 846.4 million events, the system passed every applicable validation check except one unit-consistency issue on the MIMIC-IV dataset, and it caught all 28 faults the researchers deliberately planted to test it.
The researchers ran a controlled experiment showing why the distinction matters: when later diagnoses were incorrectly assigned back to a patient's admission time, measured model performance looked artificially inflated. A model trained on that distorted timeline then lost accuracy when tested on histories that respected what information was actually available at each moment. That's a direct problem for any hospital tool meant to flag a deteriorating patient in real time, since deployment only ever has access to the past, not the future the training data accidentally leaked in.
It's a reminder that in clinical AI, the plumbing, whether a model is quietly cheating by peeking at outcomes it shouldn't know yet, matters as much as the model architecture built on top of it.