A new benchmark says today's AI coding agents are lost without a bug report to guide them.
The benchmark, called Active-SWE, comes from a paper posted to arXiv (arXiv:2608.04682) that flips the usual test for coding agents. Most existing benchmarks hand an agent a detailed issue report and ask it to patch the specific bug described. Active-SWE instead asks agents to find bugs on their own, using 1,663 tasks spanning six bug categories and eight programming languages. The researchers built a difficulty-aware task pipeline and a dual-track evaluation framework to grade bug discovery and bug fixing separately.
Real engineers rarely get a tidy issue report before they start debugging; they scan a codebase, spot something off, and fix it before it becomes a ticket. Active-SWE tests that proactive skill directly, and the results are unflattering: state-of-the-art coding agents struggled to locate and resolve even recorded bugs, let alone find new multi-bug problems or flag valid potential bugs on their own.
In other words, the agents billed as ready to replace junior engineers still need someone to write the bug report for them first.