A new self-supervised training method cuts AI reasoning length by up to 25% without ever telling the model to be shorter.
Researchers fine-tuned reasoning models to predict their own confidence at points along their reasoning traces, using just 600 training problems. The training loss never mentions length, efficiency, or stopping - only confidence. At inference time, the models run standard generation with no early-stopping trick bolted on. Tested on Gemma, Qwen, Nemotron, and GPT-OSS models across math, science, and coding benchmarks, the fine-tuned versions matched baseline accuracy while generating up to 25% fewer tokens, on par with methods built explicitly to shorten reasoning.
That matters because inference cost is the real tax on reasoning models: every extra token a model rambles through costs money and time, and most fixes so far have meant either penalizing length during training or bolting on early-stopping logic at inference. This result suggests efficiency can show up as a side effect of teaching a model to track its own certainty, without engineers having to explicitly optimize for brevity at all.
Still, this is one early paper trained on 600 problems, not a peer-reviewed, large-scale result - and getting a model to accurately judge its own confidence is a notoriously shaky foundation to build on.