[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-researchers-build-automated-bug-benchmark-for-ai-coding-agents":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},8654,"researchers-build-automated-bug-benchmark-for-ai-coding-agents","Researchers Build Automated Bug Benchmark for AI Coding Agents","A new arXiv paper describes a tool that auto-discovers and reproduces real bugs in AI agent harnesses, then uses them to test and improve coding agents.","A new automated tool hunts down and reproduces real bugs in the software that lets AI coding agents use tools - and most agents still can't fix what it finds.\n\nA paper posted to arXiv on September 30, 2026 (arXiv:2609.37864v1, not yet peer-reviewed) introduces AgentBug-Smith, a system that automatically discovers and reproduces real harness bugs - the connective code that lets an AI agent call tools and interact with its environment - from open-source agentic projects. The paper's authors report it reproduces these bugs at success rates 10.67 to 27.56 percent higher than existing general-purpose bug-reproduction techniques, depending on the backbone LLM used. They used it to build Live-Harness-Bench, a benchmark that currently holds 200 reproducible harness bugs and is designed to keep growing as new ones turn up in the wild. The same authors report that distilling repair patterns mined from that benchmark into existing coding agents raised those agents' harness-bug repair rates by 6.32 percent.\n\nThat matters because most agent benchmarks are small, fixed sets that the paper says take hundreds of human hours to build, so they go stale fast and can't track how quickly agent frameworks change. A benchmark that regenerates itself from live, real-world failures is a tougher test than a handful of hand-picked bugs, and the paper's own evaluation found that today's top software agents still struggle to repair the harness bugs it surfaces - a useful check against the idea that coding agents are close to fixing themselves.\n\nThat check comes with the standard caveat for a single, not-yet-peer-reviewed paper: the improvement numbers are self-reported by its authors, and whether Live-Harness-Bench holds up once outside labs start poking at it is still unknown.","[\"ai agents\",\"software bugs\",\"benchmarks\",\"research\"]","2026-09-30T04:00:00.000Z","2026-09-30T17:28:02.463Z","2026-09-30T17:28:08.766Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Add explicit attribution for every statistic and claim (e.g., cite the arXiv paper, its posting date\u002FID, and that it's not yet peer-reviewed) since the draft currently states all figures and findings with no named source or publication.","resolved","ai",[32,33,34,35],"ai agents","software bugs","benchmarks","research",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.37864",0,{"sections":42},[43,46,50,54,59,64,68,73,78,82,87,92,97,102],{"name":44,"slug":30,"count":45,"latest_published_at":18},"AI",5180,{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",791,{"name":51,"slug":52,"count":53,"latest_published_at":18},"Policy","policy",417,{"name":55,"slug":56,"count":57,"latest_published_at":58},"Deals","deals",284,"2026-09-29T21:00:00.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":63},"Hardware","hardware",194,"2026-09-29T13:16:04.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":18},"Science","science",155,{"name":69,"slug":70,"count":71,"latest_published_at":72},"Consumer Tech","consumer-tech",142,"2026-09-29T18:38:03.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":18},"Dev Tools","dev-tools",90,{"name":83,"slug":84,"count":85,"latest_published_at":86},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]