Researchers have pinned down a mathematical explanation for why a popular AI training shortcut sometimes backfires.
The technique is on-policy distillation, where a smaller "student" model learns by generating its own responses and getting corrected by a teacher model, rather than just copying the teacher's answers directly. A new paper examines what happens when a student learns from multiple teachers in sequence, and finds the method comes down to a choice between two statistical approaches. One, called forward KL divergence, essentially averages the teachers' opinions. The other, reverse KL divergence, gives more weight to a confident teacher's preferences. The researchers built algorithms for both and proved bounds on how fast they converge.
This matters because on-policy distillation has been gaining traction as a way to shrink AI models without the usual problem of "catastrophic forgetting," where fine-tuning on new data wipes out old capabilities. But the paper shows the reverse-KL approach, which is the one that better preserves a confident teacher's judgment, is also more easily thrown off by a teacher that is uncertain about the correct answer. Worse, its token-by-token structure can let the student lock onto wrong answers early in long outputs and compound the error over the rest of the response.
In other words, the same math that makes on-policy distillation good at avoiding forgetting is also what makes it fragile under noisy supervision. That's a useful caution for anyone treating distillation as a free lunch for building cheaper models.