Researchers just measured how AI coding agents actually spend their token budgets, and the answer is: unpredictably.
The study analyzed trajectories from eight frontier language models tackling SWE-bench Verified, a standard benchmark of real-world coding tasks. It found that agentic coding tasks consume roughly 1000x more tokens than ordinary code chat or reasoning, with most of that cost coming from input tokens rather than the answers models generate. Running the identical task twice can produce token totals that differ by up to 30x, and spending more doesn't reliably buy better results; accuracy tends to peak at moderate cost and flatten out after that. Efficiency also varies sharply by model: Kimi-K2 and Claude-Sonnet-4.5 burned through more than 1.5 million additional tokens on average compared to GPT-5 on the same tasks.
For teams budgeting for AI coding assistants, that's the real finding: there is currently no reliable way to estimate what a task will cost before running it. Human-rated task difficulty barely correlates with actual token spend, and the models themselves are bad at guessing their own usage, with correlations as weak as 0.39 and a consistent tendency to underestimate.
Model makers love to tout accuracy benchmarks. This one is a reminder that the meter is running, and nobody, including the AI, can tell you how fast.