AI/ ai · llm-reasoning · benchmarks · token-budgets

Research Shows AI Token Budget Cutoffs Flip Results at Scale

A new study finds advisory stopping beats strict cutoffs at small budgets but loses on accuracy at 32k tokens despite using fewer tokens overall.

Cutting off an AI mid-calculation helps or hurts depending entirely on how much budget you give it.

Researchers replayed 19,200 traces covering 120 AIME, BrUMO and HMMT competition problems, run through two configurations of one model with a fixed 16-attempt cap. They compared two ways of handling a token budget that runs out mid-derivation: stopping immediately (strict) or letting the current attempt finish (advisory). At a 4,000-token cap, advisory's accuracy gains came mostly from converting an abstention into a correct answer, since strict stopping leaves an unfinished prefix the selector can't use. But the advantage doesn't hold as budgets grow: in the high-budget condition, advisory's accuracy came in 0.42 percentage points below strict running at a 32,000-token cap, even though advisory used only 59% as many tokens on average.

That reversal matters because AI vendors routinely sell 'let it think longer' as an unambiguous upgrade. This study suggests the right stopping rule flips depending on budget size, and that measuring accuracy against realized cost rather than the stated cap can make a worse method look better on paper. The researchers also found that giving a selector more candidate answers to pick from doesn't reliably improve accuracy, undercutting a second common assumption in budget-based benchmarking.

In other words: before trusting any 'we let the model think more' accuracy chart, ask what cap, cost and selector it's hiding.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →