[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-research-shows-ai-token-budget-cutoffs-flip-results-at-scale":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},8888,"research-shows-ai-token-budget-cutoffs-flip-results-at-scale","Research Shows AI Token Budget Cutoffs Flip Results at Scale","A new study finds advisory stopping beats strict cutoffs at small budgets but loses on accuracy at 32k tokens despite using fewer tokens overall.","Cutting off an AI mid-calculation helps or hurts depending entirely on how much budget you give it.\n\nResearchers replayed 19,200 traces covering 120 AIME, BrUMO and HMMT competition problems, run through two configurations of one model with a fixed 16-attempt cap. They compared two ways of handling a token budget that runs out mid-derivation: stopping immediately (strict) or letting the current attempt finish (advisory). At a 4,000-token cap, advisory's accuracy gains came mostly from converting an abstention into a correct answer, since strict stopping leaves an unfinished prefix the selector can't use. But the advantage doesn't hold as budgets grow: in the high-budget condition, advisory's accuracy came in 0.42 percentage points below strict running at a 32,000-token cap, even though advisory used only 59% as many tokens on average.\n\nThat reversal matters because AI vendors routinely sell 'let it think longer' as an unambiguous upgrade. This study suggests the right stopping rule flips depending on budget size, and that measuring accuracy against realized cost rather than the stated cap can make a worse method look better on paper. The researchers also found that giving a selector more candidate answers to pick from doesn't reliably improve accuracy, undercutting a second common assumption in budget-based benchmarking.\n\nIn other words: before trusting any 'we let the model think more' accuracy chart, ask what cap, cost and selector it's hiding.","[\"ai\",\"llm-reasoning\",\"benchmarks\",\"token-budgets\"]","2026-10-01T04:00:00.000Z","2026-10-01T09:22:32.319Z","2026-10-01T09:22:37.931Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Fix the dek: it says advisory's edge 'shrinks' at high budgets, but the body shows advisory 4k actually falls 0.42 points behind strict 32k (a reversal, not a shrinking lead), so rewrite the dek\u002Fheadline framing to reflect that advisory loses on raw accuracy at high budgets despite using fewer tokens.","resolved","ai",[30,32,33,34],"llm-reasoning","benchmarks","token-budgets",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.38699",0,{"sections":41},[42,45,50,55,60,65,70,75,80,84,89,94,99,104],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",5351,{"name":46,"slug":47,"count":48,"latest_published_at":49},"Security","security",801,"2026-09-30T22:18:23.000Z",{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",429,"2026-10-01T02:26:17.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",298,"2026-09-30T21:00:26.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":69},"Science","science",157,"2026-09-30T15:00:56.000Z",{"name":71,"slug":72,"count":73,"latest_published_at":74},"Consumer Tech","consumer-tech",149,"2026-09-30T22:57:11.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":81,"slug":82,"count":78,"latest_published_at":83},"Software","software","2026-09-30T21:41:11.000Z",{"name":85,"slug":86,"count":87,"latest_published_at":88},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":90,"slug":91,"count":92,"latest_published_at":93},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":95,"slug":96,"count":97,"latest_published_at":98},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":100,"slug":101,"count":102,"latest_published_at":103},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":105,"slug":106,"count":107,"latest_published_at":108},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]