Coding agents look great on paper, but a new benchmark suggests they're better at dodging hard problems than solving them.
Researchers built CABRA, a benchmark constructed from scratch as call-graph transformations rather than scraped from real repositories, so they could dial difficulty up or down along four specific axes: function traversal, search, runtime resolution, and instruction following. Running eight standalone LLMs and six full coding agents across 6,840 tasks, they found a split: raw LLM accuracy drops as tasks get harder, but agents stay near-perfect because they lean on tools like grep to do the heavy lifting. A companion analysis of SWE-bench Verified backs this up - how many tool calls an agent makes for reading and analysis predicts its accuracy better than how many lines of code it ultimately edits. When the researchers pushed CABRA further, adding a task that requires tracing divergent logic across two classes, agent accuracy finally cracked.
That's the real finding here: the bottleneck isn't editing code, it's understanding it. Benchmarks like SWE-bench have quietly assumed that lines-changed roughly tracks difficulty, which has let agents get credit for grep-and-patch maneuvers that look like comprehension but aren't. If tool access is masking a comprehension gap, every leaderboard built on tasks tools can trivialize is overstating how much these systems actually "get" the code they're touching.
It's the coding-agent version of a needle-in-a-haystack test that a keyword search quietly solves for you - impressive numbers, less impressive understanding.