AI/ ai-agents · oncall · benchmarks · llm-evaluation

AI Coding Agents Fail Oncall Root Cause Benchmark

ORCA-bench tested five frontier AI agents on realistic production incidents and found the best one solved just a quarter of them.

A new benchmark says AI agents that write code confidently still can't figure out why your production system just broke.

Researchers built ORCA-bench, a test that drops AI coding agents into a simulated oncall shift. The setup pairs 1,079 root-cause-analysis tasks with six days of metrics, logs, and traces from a microservice system running under continuous simulated load, all accessed through real tools like Prometheus, Jaeger, and OpenSearch via Grafana. Agents also get full access to the underlying source code, the same toolkit a human engineer would use to diagnose an incident. Ground-truth answers were verified by expert site reliability engineers, and the paper's automated scoring was checked against human judges with strong agreement, a weighted Cohen's kappa of 0.91.

Across five frontier agents, including Claude Fable 5, the best RCA accuracy was 25.3% on realistic, medium-difficulty tasks and just 10% on hard ones. The weakest model invented a plausible-sounding but wrong root cause in 40% of incidents, and every model got worse and more prone to hallucinating when source-code access was removed. That's a meaningful data point for anyone pitching AI agents as oncall backup: these are the skills that matter when an alert fires at 3 a.m., and the gap between writing code and explaining why code broke is still wide.

The test system is a modest 50 GB, six-day slice of a public codebase; real production environments are bigger, messier, and far less forgiving, so treat these numbers as a ceiling, not a floor.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →