AI/ distillation · llm training · machine learning · research

New Research Says More Supervision Can Hurt AI Distillation

A new study finds that adding more training supervision during AI model distillation can lower accuracy instead of improving it.

A new study finds that giving an AI model more training signal during distillation can make it worse, not better.

On-policy distillation lets a smaller 'student' model learn by generating its own answers and getting graded by a larger 'teacher' model. When the two models use different tokenizers - the systems that break text into chunks - researchers have assumed you need to align as much of the shared vocabulary as possible to teach well. A new arXiv paper tested that assumption across three teacher-student pairs on math reasoning and code generation tasks. Restricting the training signal to a compact set of 16 well-matched vocabulary positions matched the accuracy of full-coverage alignment and beat other cross-tokenizer methods; adding extra supervision to patch over mismatched vocabulary spots lowered accuracy instead of raising it.

That is a notable result in a field that usually treats more data and more coverage as a free lunch. The researchers traced the drop to gradient conflict: the patched-in supervision pointed in directions that fought the reliable signal, and that conflict grew stronger as training went on. For anyone building smaller, cheaper models by distilling from bigger ones, it is a concrete argument against bolting on extra loss terms just because the coverage gap is there to fill.

Sometimes the fix for a mismatch is not to cover it - it is to ignore it.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →