[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-ai-auditor-catches-bugs-in-two-science-benchmarks":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},10111,"ai-auditor-catches-bugs-in-two-science-benchmarks","AI Auditor Catches Bugs in Two Science Benchmarks","BenchGuard, an LLM-based auditor, found 12 author-confirmed flaws in ScienceAgentBench and matched 83.3% of known BIXBench issues for under $15 a run.","BenchGuard uses AI to check the homework of AI benchmark makers - and it's finding real mistakes.\n\nResearchers built BenchGuard, a system that puts frontier language models to work auditing agent benchmarks instead of just taking them, cross-checking specifications, tasks, and reference solutions, and optionally using agent solutions or execution traces as extra evidence. Run on ScienceAgentBench, it surfaced 12 issues the benchmark's own authors confirmed were real, including fatal errors that made some tasks impossible to solve. On a 50-task subset of BIXBench called Verified-50, BenchGuard's findings matched 83.3% of the issues human experts had already flagged. A full audit of that 50-task bioinformatics set cost under $15, and a preliminary run on ProgramBench suggests the approach works across different benchmark formats too.\n\nThat 83.3% match rate is less a victory lap than a sanity check: it shows an AI auditor can reliably reproduce painstaking human review for pocket change. The 12 confirmed bugs in ScienceAgentBench matter more directly, since some agents graded as failing on that benchmark may simply have been asked to solve an unsolvable task. For labs burning compute chasing leaderboard rankings, that is an expensive kind of noise worth catching before it skews results.\n\nIt is a tidy bit of recursion: AI models now checking the tests used to grade AI models, for less than the price of a nice dinner.","[\"ai\",\"benchmarks\",\"llm-agents\",\"ai-research\"]","2026-10-05T04:00:00.000Z","2026-10-05T23:38:20.491Z","2026-10-05T23:38:26.873Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The dek claims BenchGuard found 'dozens of errors that human reviewers missed,' but the body only supports 12 author-confirmed issues on ScienceAgentBench (not described as missed by reviewers) and an 83.3% match rate on already-flagged BIXBench issues — fix the dek to match the actual counts in the body.","resolved","ai",[30,32,33,34],"benchmarks","llm-agents","ai-research",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2604.24955",0,{"sections":41},[42,46,50,55,60,65,69,74,79,84,89,94,99,104],{"name":43,"slug":30,"count":44,"latest_published_at":45},"AI",6317,"2026-10-05T09:51:57.000Z",{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",871,{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",446,"2026-10-05T10:25:00.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",340,"2026-10-05T09:18:03.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Hardware","hardware",205,"2026-10-05T10:58:22.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":18},"Science","science",179,{"name":70,"slug":71,"count":72,"latest_published_at":73},"Consumer Tech","consumer-tech",160,"2026-10-05T10:23:15.000Z",{"name":75,"slug":76,"count":77,"latest_published_at":78},"Dev Tools","dev-tools",99,"2026-10-05T10:47:06.000Z",{"name":80,"slug":81,"count":82,"latest_published_at":83},"Software","software",97,"2026-10-04T10:00:00.000Z",{"name":85,"slug":86,"count":87,"latest_published_at":88},"Startups","startups",93,"2026-10-05T11:13:51.000Z",{"name":90,"slug":91,"count":92,"latest_published_at":93},"Gaming","gaming",53,"2026-10-02T02:50:39.000Z",{"name":95,"slug":96,"count":97,"latest_published_at":98},"General","general",51,"2026-10-05T02:35:01.000Z",{"name":100,"slug":101,"count":102,"latest_published_at":103},"Reviews","reviews",32,"2026-10-02T18:00:00.000Z",{"name":105,"slug":106,"count":107,"latest_published_at":108},"How-To","how-to",8,"2026-10-05T09:00:00.000Z"]