AI/ ai · model-distillation · reasoning-models · small-language-models

Small Models Learn to Solve But Not When to Stop

A new study distilling Qwen3-8B into smaller models finds they learn to solve problems but not to reliably recognize when they've solved them.

A new arXiv paper says shrinking AI reasoning models breaks something more important than raw problem-solving: knowing when to shut up.

Researchers distilled Qwen3-8B, a larger reasoning model, into three smaller student models with 4 billion, 1.7 billion and 0.6 billion parameters, using on-policy distillation, where a student learns from a stronger teacher's feedback on its own outputs. They tested each student in "thinking mode" (extended reasoning before an answer) and "non-thinking mode" (answering directly). Solving ability improved at every size but always stayed capped below what the teacher could reach, and the smaller the student, the bigger that gap. The ability to stop reasoning at the right moment did not transfer nearly as well: in thinking mode, students increasingly cut their reasoning short during training, and smaller students lost more of that skill than larger ones.

The researchers found the teacher only signals a stop where the student already stops on its own, so distillation cannot teach new stopping behavior, only reinforce existing stops that happen to land on a correct answer. That is a problem for weak students, who have few correct stops to reinforce in the first place. The smallest models often compute the right answer internally but never commit to it, either failing to mark it as final or writing past it.

That looks like a structural limit, not a training bug, which is a less comfortable story than the usual "just distill harder" pitch for squeezing reasoning into small, on-device models.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →