Researchers have found a smarter way to let AI models tutor themselves through math problems, without the self-correction backfiring.
The technique, called STEPS (Selective On-Policy Self-Distillation), has a model act as its own teacher, using privileged context only available during training to coach its own reasoning. Earlier versions of this self-distillation approach applied that coaching to every token in a response, which researchers found can over-constrain the model and bake in biases from information it will not have at inference time. STEPS instead targets only the critical spans in a model's practice answers, applying one kind of correction to spans that are on track and another to spans heading toward an error, then phases out entirely in favor of GRPO (Group Relative Policy Optimization), the standard reward-based reasoning method, after a short window.
Across four models from three model families, STEPS beat plain GRPO training on accuracy, including a 2.76 percentage point gain on Qwen3-8B across four math benchmarks plus the GPQA-Diamond science test, for roughly 2.7% more training compute. Even when the model graded itself instead of using a stronger outside annotator, it still gained 1.90 points, meaning the technique does not depend on access to a better teacher model.
Self-distillation has been sold as a way to squeeze more reasoning out of models without bigger datasets or more parameters. STEPS's actual contribution is narrower and more useful: knowing when to stop hand-holding.