AI/ ai-benchmarks · ai-agents · arxiv-research · ai-safety

Paper Finds AI Agent Benchmarks Score Broken Tools as Passing

A new arXiv paper finds agent benchmarks often grade tools by what they claim to do, not what actually happens, letting broken tools score as fixed.

Benchmarks that certify AI agents may be certifying agents that never actually did anything.

A paper posted to arXiv on September 30, 2026, titled "Do Agent Benchmarks Do What They Say? An Executable-Contract Audit of Tool-Using Agent Environments" (arXiv:2609.37315), audited 34 mutating tools across four popular agent-testing benchmarks. The researchers treated each tool's documented interface as a contract, then checked whether the tool's underlying code did what it advertised, instead of trusting the tool's own report of success. They confirmed seven tool defects and one evaluator flaw at specific, pinned versions of the benchmark code. When they injected known defects to test their own checker, it never raised a false alarm across 25 flags, but missed most real problems - in 29 of 33 scored misses, the evaluator's test coverage technically touched the defect without any check catching it.

This isn't about buggy tools alone - it's about scores nobody can actually trust. In the AgentDojo benchmark, at least 5 of the paper's 25 examined mutating tools diverged from their advertised behavior, and in tau2-bench, no scored run in 1,120 test paths ever reached a known defect the paths were built to isolate - the evaluator rewarded an agent for refueling a suspended phone line and penalized the correctly repaired tool. In the clinical-records benchmark, a tool tells the agent every write succeeded under an undisclosed no-write design, and the grader counts that message as proof rather than checking whether any record actually changed.

Benchmarks are the report card the whole agentic-AI pitch leans on; if the grading script cannot tell a real write from a tool that silently does nothing, every leaderboard score built on it is decoration, not data.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →