AI/ ai-agents · benchmarks · llm-research · multi-agent-systems

Study Tests AI Agents on Facts That Change Mid Task

A new benchmark shows a lean agent system beats a call-heavy rival at tracking which version of a fact is true when the underlying document keeps changing.

Researchers built a test that asks AI agents a trickier question than usual: not just "what's the answer," but "what was true at the moment this question was asked."

The setup takes six standard benchmarks, including MMLU, MATH, and HumanEval, and turns them into 31,119 episodes where facts get revised mid-stream by an event feed. An agent has to find the document version that was valid at query time, then answer based on that version, rather than just reusing whatever it last computed. The researchers' own system, called RIAG, splits that work into two parts: a deterministic step that figures out which version is current, and a separate reasoning step that solves the task. It hit 54.24 percent joint accuracy using about 0.62 model calls per query. The strongest comparison method managed only 32.22 percent, and needed 18 calls per query to get there.

That gap matters beyond the leaderboard. Most agent benchmarks assume the facts hold still while the model thinks. Real pipelines, like news databases, medical records, or codebases, don't work that way, so an agent that can't tell a stale cached answer from a current one will confidently serve the wrong one. RIAG's trick is architectural, not just a bigger model: cache by document identity, try the cheap path first, only call in audit and repair when something looks off.

The efficiency claim is the real headline here, not the accuracy number. Going from 18 calls to under one per query is a 29x cut in compute for a system that's also more accurate, which is the kind of tradeoff that actually ships in production rather than staying in a paper.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →