A new training trick stops AI reasoning models from padding out answers to questions they already know how to solve.
The issue starts with reinforcement learning post-training, the fine-tuning stage where models practice solving problems and get rewarded for correct answers. Longer responses are often treated as a proxy for better reasoning on hard problems, but researchers found that same length creep leaks into easy problems too, with no accuracy benefit. They call this the length-scaling tax, or LST: extra words on already-solved queries that don't make the answer any better. Their fix, called Length Self-Distillation, routes easy prompts through a distillation step instead of standard reinforcement learning, while hard, unsolved prompts still get the normal training process. The twist is that the "teacher" guiding the distillation isn't a separate model at all, it's just a running average of the model's own earlier weights.
This matters because verbosity has a real price tag. Every extra sentence a model generates costs more compute, more latency, and more money at scale, and that cost compounds across millions of queries. A method that trims padding without needing extra data or an outside model is cheap to bolt onto training pipelines that already exist, which is probably why the gains are notable: the tax dropped from 19.0% to -3.7% on single-turn reasoning tasks, and from 31.4% to 13.7% on multi-turn agentic tasks, with accuracy holding steady or improving.
Still, this is a patch on a training method, not a rethink of why models learned to pad in the first place, since the underlying incentive, rewarding length as a stand-in for effort, hasn't gone away.