BenchGuard uses AI to check the homework of AI benchmark makers - and it's finding real mistakes.
Researchers built BenchGuard, a system that puts frontier language models to work auditing agent benchmarks instead of just taking them, cross-checking specifications, tasks, and reference solutions, and optionally using agent solutions or execution traces as extra evidence. Run on ScienceAgentBench, it surfaced 12 issues the benchmark's own authors confirmed were real, including fatal errors that made some tasks impossible to solve. On a 50-task subset of BIXBench called Verified-50, BenchGuard's findings matched 83.3% of the issues human experts had already flagged. A full audit of that 50-task bioinformatics set cost under $15, and a preliminary run on ProgramBench suggests the approach works across different benchmark formats too.
That 83.3% match rate is less a victory lap than a sanity check: it shows an AI auditor can reliably reproduce painstaking human review for pocket change. The 12 confirmed bugs in ScienceAgentBench matter more directly, since some agents graded as failing on that benchmark may simply have been asked to solve an unsolvable task. For labs burning compute chasing leaderboard rankings, that is an expensive kind of noise worth catching before it skews results.
It is a tidy bit of recursion: AI models now checking the tests used to grade AI models, for less than the price of a nice dinner.