[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-benchmark-judges-ai-coding-agents-on-repo-health":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},7712,"new-benchmark-judges-ai-coding-agents-on-repo-health","New Benchmark Judges AI Coding Agents on Repo Health","SWE-Prometheus scores coding agents on governance, and a template baseline nailed tests and docs but scored zero on security and environments.","A new benchmark grades AI coding agents on whether they actually make a codebase healthier, not just whether their patch closes a ticket.\n\nSWE-Prometheus tests agents across six governance dimensions - among them tests and CI, quality gates, documentation, dependency and security, and reproducible environments - using fixed repository snapshots and clean-environment checks rather than a human-written bug report. Ten AI models were run on a shared 22-repository subset, where mean Normalized Governance Improvement ranged from 0.0568 to 0.5760 and the rate of models breaking existing behavior ranged from 0% to 23%. Comparing the two top-scoring systems, Kimi-K3 pulled ahead once broken-behavior cases were counted, even though the two looked similar on valid-only comparisons. Separately, on a frozen ten-repository batch, the researchers tested a repository-blind template - a fixed, non-adaptive baseline, not one of the ten evaluated models - which scored a mean NGI of 0.272 but improved Reproducible Environment and Dependency & Security in zero of those repositories; its gains came entirely from Tests & CI, Quality Gates, and Documentation.\n\nThat baseline result is the useful part. It shows a benchmark can be fooled into rewarding a repo for looking more governed - more tests, more CI badges, more docs - without ever touching the unglamorous, harder-to-fake work of pinning dependencies or making builds reproducible. Most existing repo-level benchmarks only check whether a patch resolves a known issue; SWE-Prometheus is built to catch that kind of surface-only cleanup instead of rewarding it.\n\nFor context, doing nothing at all - the benchmark's no-op condition - scored a median NGI of zero, which is the bar every model and every template has to clear before any of these numbers mean much.","[\"ai\",\"coding-agents\",\"benchmarks\",\"software-engineering\"]","2026-09-25T04:00:00.000Z","2026-09-25T05:50:59.519Z","2026-09-25T05:51:05.238Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The dek and body claim that 'AI coding agents routinely skip security and environment fixes' as a general finding about the tested models, but the source only shows zero coverage of Dependency & Security and Reproducible Environment for the simple repository-blind template baseline on a separate 10-repo batch — not for the ten actual AI models evaluated on the 22-repo subset — so rewrite to correctly attribute that specific finding to the baseline rather than implying it describes the agents' be","resolved","ai",[30,32,33,34],"coding-agents","benchmarks","software-engineering",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.29465",0,{"sections":41},[42,45,50,55,60,65,70,75,80,85,90,95,99,104],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",4466,{"name":46,"slug":47,"count":48,"latest_published_at":49},"Security","security",729,"2026-09-24T19:54:21.000Z",{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",386,"2026-09-24T23:50:55.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",237,"2026-09-24T22:00:00.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Hardware","hardware",182,"2026-09-25T01:25:53.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":69},"Science","science",138,"2026-09-24T18:24:52.000Z",{"name":71,"slug":72,"count":73,"latest_published_at":74},"Consumer Tech","consumer-tech",128,"2026-09-24T19:24:34.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Software","software",88,"2026-09-24T23:06:55.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":84},"Dev Tools","dev-tools",79,"2026-09-22T22:21:13.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"Startups","startups",71,"2026-09-24T20:45:00.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Gaming","gaming",46,"2026-09-24T17:52:29.000Z",{"name":96,"slug":97,"count":93,"latest_published_at":98},"General","general","2026-09-25T02:12:57.000Z",{"name":100,"slug":101,"count":102,"latest_published_at":103},"Reviews","reviews",30,"2026-09-24T20:07:31.000Z",{"name":105,"slug":106,"count":107,"latest_published_at":108},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]