AI/ ai · healthcare · benchmarks · agent memory

New Benchmark Finds AI Health Agents Choke on Too Much Memory

A new benchmark shows that medical AI agents actually get worse at answering clinical questions the more conversation history they accumulate.

Researchers have built a benchmark that measures something most AI health apps never get tested on: whether they can actually remember you correctly over months of use.

The new benchmark, called MedMemoryBench, synthesizes long-horizon medical conversations from clinically grounded, synthetic patient profiles using a human-agent collaborative pipeline. The resulting dataset covers roughly 2,000 sessions and 16,000 interaction turns, all expertly validated. Rather than testing memory in one static snapshot, the benchmark evaluates agents while their memory is still accumulating, mirroring how a real production system builds up a patient's history over time. The researchers say the work was driven by the demands of an unnamed health management agent already serving tens of millions of active users.

The headline finding is what the researchers call memory saturation: the more clinical history an agent accumulates, the worse it gets at retrieving the right facts and reasoning about them. That is a bad look for mainstream AI architectures, which the benchmark shows struggle most with complex medical reasoning and lose accuracy once the conversation gets noisy.

Most memory benchmarks to date have been built around casual chat, not clinical stakes. This one is a reminder that an AI health agent getting your medical history wrong isn't a quirky bug - it's the kind of mistake that belongs in an incident report.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →