AI/ ai · machine-learning · self-distillation · reasoning-models

AI Researchers Fix Hidden Confidence Bug in Self-Teaching Models

A new method called E2-OPSD fixes a confidence-drift bug in self-distillation training, boosting reasoning benchmarks without added compute.

A popular trick for teaching AI models to reason has a confidence problem, and researchers just patched it.

Self-distillation lets one model train itself: it plays teacher when given the answer and student when it only sees the question, skipping the need for a second model entirely. Researchers found that during this process the student's uncertainty keeps climbing past the teacher's and never settles back down, a glitch they call entropy overshoot. The cause is twofold: the teacher's confidence is tied to clues specific to the final answer rather than reusable reasoning steps, and the standard training math keeps spreading the student's predictions without reeling them back in. Their fix, called E2-OPSD, swaps the real answer for a similar already-solved problem during teaching and adjusts each correction based on the gap between student and teacher confidence.

That matters because self-distillation is attractive precisely because it is cheap, and most fixes in this space bolt on extra models or compute to patch problems like this. E2-OPSD does not: same training setup, no new forward passes, but up to 4.3 points better on math reasoning benchmarks and up to 5.5 points better on out-of-domain tests.

It is a narrow, plumbing-level result, one technique fixing one failure mode. But plumbing fixes like this tend to get quietly copied into every lab's training recipe once they are this cheap to adopt.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →