Benchmarks that certify AI agents may be certifying agents that never actually did anything.
A paper posted to arXiv on September 30, 2026, titled "Do Agent Benchmarks Do What They Say? An Executable-Contract Audit of Tool-Using Agent Environments" (arXiv:2609.37315), audited 34 mutating tools across four popular agent-testing benchmarks. The researchers treated each tool's documented interface as a contract, then checked whether the tool's underlying code did what it advertised, instead of trusting the tool's own report of success. They confirmed seven tool defects and one evaluator flaw at specific, pinned versions of the benchmark code. When they injected known defects to test their own checker, it never raised a false alarm across 25 flags, but missed most real problems - in 29 of 33 scored misses, the evaluator's test coverage technically touched the defect without any check catching it.
This isn't about buggy tools alone - it's about scores nobody can actually trust. In the AgentDojo benchmark, at least 5 of the paper's 25 examined mutating tools diverged from their advertised behavior, and in tau2-bench, no scored run in 1,120 test paths ever reached a known defect the paths were built to isolate - the evaluator rewarded an agent for refueling a suspended phone line and penalized the correctly repaired tool. In the clinical-records benchmark, a tool tells the agent every write succeeded under an undisclosed no-write design, and the grader counts that message as proof rather than checking whether any record actually changed.
Benchmarks are the report card the whole agentic-AI pitch leans on; if the grading script cannot tell a real write from a tool that silently does nothing, every leaderboard score built on it is decoration, not data.