AI/ ai · llm-agents · research · benchmarking

Giving AI Agents Memory Costs Cache Space, Not Accuracy

A new study finds AI agent memory adds cache overhead but no measurable accuracy gain across eight benchmarks.

A new arXiv study finds that giving AI agents persistent memory barely changes their accuracy - but it does cost extra cache memory.

Researchers tested a three-tier multi-agent architecture that splits long-context inference across cooperating agents. Splitting the work this way cut peak memory use dramatically: 14.3 MiB per query, versus 35.5 MiB and 35.3 MiB for single-pass and retrieval-augmented baselines. The same architecture also included a persistent memory tier that stores and recalls past reasoning traces, the kind of feature agent papers often credit with accuracy gains. Across eight controlled benchmark pairs with 100 trials each, that memory tier added about 0.368 MiB of cache and produced no statistically detectable accuracy improvement.

Decomposing inference across agents is a real and substantial efficiency win here. But the common claim that agent memory improves accuracy didn't hold up under controlled testing. The authors argue the null result is structural, not a fluke: standard benchmarks score each question independently, so stored memory has nothing useful to retrieve unless traces are reset between conditions - and getting that reset right took four separate corrections to their own measurement setup.

That's a useful caution for anyone citing an ablation table: an uncontrolled benchmark can manufacture a benefit that was never actually there.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →