A new training method gives AI models a way to learn from paths they never actually took.
Researchers behind a technique called TISD (Trajectory-Intervention Self-Distillation) tackled a specific flaw in on-policy self-distillation (OPSD), a popular way of training AI models by having a stronger "teacher" model provide dense feedback on a student's own generated outputs. The problem: when the teacher disagrees with the student mid-generation and prefers a different next step, OPSD can flag that disagreement but can't teach the student what would happen if it actually took that alternate path - unless the student happens to stumble onto it itself. TISD fixes this by forcing the student down the teacher's preferred branch at the disagreement point, letting the student finish generating from there, then having the teacher grade the entire resulting trajectory. Tested on coding models and science-domain benchmarks, the method beat a prior technique called SDPO by 1.2 percentage points on coding accuracy and by up to 0.8 points on science tasks.
This is a plumbing fix for a specific data-collection bottleneck in AI training, not a new architecture or a bigger model. It matters because dense-feedback training methods like self-distillation are becoming a standard alternative to reinforcement learning for teaching models to reason step by step, and studies like this quietly determine whether that approach scales efficiently or keeps hitting the same blind spots. The gains are incremental, but they attack a structural limitation - the training method itself refusing to explore paths the student hasn't already sampled - that no amount of extra data would otherwise solve.
A single-digit percentage-point bump over one baseline, across a couple of benchmark families, is not evidence this fixes reasoning training generally - it is evidence the fix is worth testing at bigger scale before anyone calls it a breakthrough.