A new scoping review says much of the evidence behind AI agent upgrades doesn't actually prove the upgrade helped.
Researchers combed through 348 studies of language-model agents, the systems that plan, act, and recover from errors across multiple steps, looking at a common design move: swapping one component, such as an error detector or a stopping rule, for another. Of 222 studies that measured how well a new component performed on its own, only 142 also checked whether the overall task got done better, and 49 relied on proxy measures instead of real task outcomes. Three specific cases make the point. One monitoring tool is credited with a package-level completion gain that can't be clearly traced to better detection. A selection method's improvement was validated only against an offline proxy, not a real run. A termination method cuts premature, unsupported stops and matches prior completion rates, but doesn't beat them. The review's core finding: a local metric and a task metric improving at the same time is not proof that the local fix caused the task-level gain.
That gap matters because this is exactly how agent tooling gets marketed. A new component ships, a benchmark number ticks up, and the explanation offered is better decision quality, when the real cause could just as easily be a different execution path that happened to dodge a known failure mode. Buyers and researchers evaluating agent frameworks have little way to tell the two apart without matched, controlled comparisons, which this review says are mostly missing.
So next time a vendor claims its new agent component makes smarter calls, ask for the head-to-head evidence, not just a chart that points up.