[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-ai-coding-agents-fail-oncall-root-cause-benchmark":10,"sections":46},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":35,"tags":36,"sources":41,"feedback":45,"feedback_at":22,"cost_usd":45,"total_tokens":45},8722,"ai-coding-agents-fail-oncall-root-cause-benchmark","AI Coding Agents Fail Oncall Root Cause Benchmark","ORCA-bench tested five frontier AI agents on realistic production incidents and found the best one solved just a quarter of them.","A new benchmark says AI agents that write code confidently still can't figure out why your production system just broke.\n\nResearchers built ORCA-bench, a test that drops AI coding agents into a simulated oncall shift. The setup pairs 1,079 root-cause-analysis tasks with six days of metrics, logs, and traces from a microservice system running under continuous simulated load, all accessed through real tools like Prometheus, Jaeger, and OpenSearch via Grafana. Agents also get full access to the underlying source code, the same toolkit a human engineer would use to diagnose an incident. Ground-truth answers were verified by expert site reliability engineers, and the paper's automated scoring was checked against human judges with strong agreement, a weighted Cohen's kappa of 0.91.\n\nAcross five frontier agents, including Claude Fable 5, the best RCA accuracy was 25.3% on realistic, medium-difficulty tasks and just 10% on hard ones. The weakest model invented a plausible-sounding but wrong root cause in 40% of incidents, and every model got worse and more prone to hallucinating when source-code access was removed. That's a meaningful data point for anyone pitching AI agents as oncall backup: these are the skills that matter when an alert fires at 3 a.m., and the gap between writing code and explaining why code broke is still wide.\n\nThe test system is a modest 50 GB, six-day slice of a public codebase; real production environments are bigger, messier, and far less forgiving, so treat these numbers as a ceiling, not a floor.","[\"ai-agents\",\"oncall\",\"benchmarks\",\"llm-evaluation\"]","2026-09-30T04:00:00.000Z","2026-09-30T21:49:29.543Z","2026-09-30T21:49:34.808Z","published",null,[24,30],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Add a named source attribution (the arXiv preprint, e.g. 'arXiv:2607.28545' and posting date) since every figure and claim currently traces to an unnamed 'researchers' with no publication or date cited, which fails the source-attribution bar.","resolved",{"id":31,"reviewer":32,"round":33,"reason":34,"status":29},"publisher-r2","publisher",2,"The cited arXiv ID 2607.28545 corresponds to a July 2026 submission (arXiv IDs encode YYMM), contradicting the article's claim that it was posted September 30, 2026.","ai",[37,38,39,40],"ai-agents","oncall","benchmarks","llm-evaluation",[42],{"name":43,"url":44},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.28545",0,{"sections":47},[48,52,57,62,67,71,75,79,84,88,93,98,103,108],{"name":49,"slug":35,"count":50,"latest_published_at":51},"AI",5214,"2026-09-30T13:00:00.000Z",{"name":53,"slug":54,"count":55,"latest_published_at":56},"Security","security",793,"2026-09-30T12:55:00.000Z",{"name":58,"slug":59,"count":60,"latest_published_at":61},"Policy","policy",419,"2026-09-30T12:24:32.000Z",{"name":63,"slug":64,"count":65,"latest_published_at":66},"Deals","deals",291,"2026-09-30T10:38:22.000Z",{"name":68,"slug":69,"count":70,"latest_published_at":51},"Hardware","hardware",196,{"name":72,"slug":73,"count":74,"latest_published_at":18},"Science","science",155,{"name":76,"slug":77,"count":78,"latest_published_at":51},"Consumer Tech","consumer-tech",144,{"name":80,"slug":81,"count":82,"latest_published_at":83},"Dev Tools","dev-tools",91,"2026-09-30T12:58:00.000Z",{"name":85,"slug":86,"count":82,"latest_published_at":87},"Software","software","2026-09-25T20:55:00.000Z",{"name":89,"slug":90,"count":91,"latest_published_at":92},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":94,"slug":95,"count":96,"latest_published_at":97},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":99,"slug":100,"count":101,"latest_published_at":102},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":104,"slug":105,"count":106,"latest_published_at":107},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":109,"slug":110,"count":111,"latest_published_at":112},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]