Coding agents like Claude Code burn real money on work they have already done.
Researchers analyzed 1,200 coding-agent trajectories from Claude Code and Mini-SWE-Agent across four configurations on the SWE-bench Verified benchmark, then tested fixes on more than 10,000 additional trajectories from held-out SWE-bench Verified and Pro tasks. They identified three recurring cost-inefficient behaviors: agents re-running searches that had already surfaced the needed files (subsumed retrieval), generating near-duplicate scripts to solve the same sub-problem, and re-running full test suites when a narrower check would do. These habits show up in 79 to 98 percent of tasks and account for as much as 22.75 percent of a task's total cost.
The researchers tested three fixes: structure-aware retrieval tools, skills the agent writes for itself from past runs, and skills written by human developers. Structure-aware retrieval sometimes backfired, pushing costs up by as much as 28.14 percent by changing how agents delegated work. Developer-written skills won clearly, cutting cost by up to 41.73 percent, roughly double what agent-generated skills achieved, because humans wrote general-purpose guidance instead of narrow, trace-specific notes.
It is a useful reminder that the fix for an AI agent's bad habits is still, for now, a person telling it what to stop doing.