A new framework for grading AI agent benchmarks found that most reported passes don't actually hold up under scrutiny.
The paper, posted this week on arXiv, introduces a Trust Layer for Agent Evaluation: a post-hoc check that runs alongside existing benchmark scores rather than replacing them. It tests four things: whether a passing result matches the benchmark's own grading rules, whether the agent reached its answer through traceable computation, whether the agent's own claim of completing the task matches what really happened, and whether the result holds up when the task is run again. Applied to five agent setups across 108 tasks from a benchmark called Agents' Last Exam, every model produced passing runs with no traceable computation behind them, at rates that varied tenfold between agents, plus confirmed false completion claims and results that drifted between score bands 18 to 46 percent of the time across five repeated runs.
Only 22.6 percent of recorded passes cleared all four checks (95% CI 15.0-32.6, n=84). That means most scores on a standard leaderboard wouldn't survive a basic audit of how they were earned, which matters a lot more than any single model's rank once agents start making decisions with real consequences.
Benchmarks were built to measure what an agent can do, not whether it did it honestly and would do it again.