AI/ ai · benchmarks · llm-agents · ai-research

AI Auditor Catches Bugs in Two Science Benchmarks

BenchGuard, an LLM-based auditor, found 12 author-confirmed flaws in ScienceAgentBench and matched 83.3% of known BIXBench issues for under $15 a run.

BenchGuard uses AI to check the homework of AI benchmark makers - and it's finding real mistakes.

Researchers built BenchGuard, a system that puts frontier language models to work auditing agent benchmarks instead of just taking them, cross-checking specifications, tasks, and reference solutions, and optionally using agent solutions or execution traces as extra evidence. Run on ScienceAgentBench, it surfaced 12 issues the benchmark's own authors confirmed were real, including fatal errors that made some tasks impossible to solve. On a 50-task subset of BIXBench called Verified-50, BenchGuard's findings matched 83.3% of the issues human experts had already flagged. A full audit of that 50-task bioinformatics set cost under $15, and a preliminary run on ProgramBench suggests the approach works across different benchmark formats too.

That 83.3% match rate is less a victory lap than a sanity check: it shows an AI auditor can reliably reproduce painstaking human review for pocket change. The 12 confirmed bugs in ScienceAgentBench matter more directly, since some agents graded as failing on that benchmark may simply have been asked to solve an unsolvable task. For labs burning compute chasing leaderboard rankings, that is an expensive kind of noise worth catching before it skews results.

It is a tidy bit of recursion: AI models now checking the tests used to grade AI models, for less than the price of a nice dinner.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →