AI/ ai · model-distillation · llm-training · arxiv

Researchers Find a Blind Spot in How AI Models Learn From Teachers

A new arXiv preprint (2609.33455) names a distillation flaw called PISA and proposes Trajectory Dropout, a cheap fix that lifts math benchmark results.

A new arXiv preprint says a popular way of teaching cheap AI models to reason like expensive ones has a hidden blind spot - and a simple fix for it.

The paper, arXiv:2609.33455 ("What Shared Prefixes Hide: Trajectory Dropout for On-Policy Distillation"), studies on-policy distillation, a method where a student model trains on text it generated itself while a stronger teacher model grades every token along the way. The authors find that when the student and teacher are working from the same reasoning prefix, the teacher's corrective signal often gets diluted - a problem they name Prefix-Induced Supervision Attenuation, or PISA. It shows up two ways: an overconfident student can shrug off a disagreeing teacher, and tokens that depend on earlier reasoning steps get correction signals no stronger than those for trivial next-word guesses. Their fix, Trajectory Dropout, randomly deletes chunks of the student's own reasoning trajectory during training while the teacher still scores the full, undropped version, forcing sharper feedback exactly where it was going missing.

On-policy distillation is already the budget option for getting smaller models to reason well without full reinforcement learning, and this paper argues a chunk of that budget was being wasted on weak gradients nobody noticed. The authors report the technique improves average performance across six math reasoning benchmarks and two out-of-domain ones, and that it bolts onto existing OPD setups with what they describe as negligible extra compute. That's a meaningful claim in a training pipeline where labs are increasingly measured on results per dollar, not just raw benchmark wins.

A caveat worth flagging: this is a v2 replacement preprint posted today, not a peer-reviewed paper, and the circulated abstract doesn't name authors, an institution, or publish the actual score deltas behind "improves average performance" - so the size of the gain is still an open question until a fuller version or published table shows up. Distillation papers have promised free lunches before; the mechanism here is plausible, but the magnitude is the part to watch.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →