AI models that reason step-by-step often keep reasoning long after they have no shot at the right answer, burning compute on questions they were always going to miss.
That is the starting point of a preprint titled 'Nice Fold or Hero Call: Learning Budget-Efficient Thinking under Policy-Dependent Solvability' (arXiv:2605.11625), which has not yet been peer-reviewed. The authors propose Budget-Efficient Thinking, or BET, a two-stage training method that pairs a behavioral cold-start phase with a reinforcement-learning technique called GRPO, tuned with a reward that treats extra reasoning tokens as an investment rather than a free resource. The system learns three moves: answer easy questions fast, fold early on ones it cannot solve, and spend extra tokens only on hard-but-winnable problems - what the paper calls a nice fold versus a hero call. Tested across seven benchmarks and three base models, BET cut reasoning-token usage by 54 percent while nudging accuracy up by as much as 3.2 percent, with similar gains carrying over zero-shot to science QA and logic tasks the models were not trained on.
The real story here is not the token savings, it is the reframing. Most efficiency methods treat a model's pass rate as a difficulty score and squeeze every question by the same amount. BET instead tries to estimate whether more thinking would actually help, which matters as reasoning models make inference cost, not training cost, the expense companies actually pay for at scale.
Fifty-four percent is a big number for a paper that has not yet cleared peer review, and the source material offers no head-to-head test against the budget controls that production reasoning systems already ship. Poker metaphors aside, teaching a model to recognize a lost cause is still just abstention with better branding.