A new AI training method lets language models catch their own bad reasoning mid-stream and rewrite it before it derails the final answer.
Researchers describe the technique, called Segment-wise On-Policy Distillation (Seg-OPD), in a paper posted to arXiv. The standard approach, on-policy distillation, has a larger teacher model grade every token a smaller student model generates while reasoning. The catch: if the student's chain-of-thought goes wrong early, token-by-token teacher feedback keeps scoring everything built on that bad start, reinforcing the mistake instead of correcting it. Seg-OPD instead flags uncertain segments in the student's reasoning, has the teacher redraft just those segments, and trains the student to prefer the teacher's redraft over its own version, on top of the usual token-level supervision.
That targets a real gap in how reasoning models get trained: once a chain-of-thought answer starts badly, most methods have no clean way to intervene mid-stream, only to keep grading the wreckage. On math and competitive programming benchmarks, Seg-OPD-trained models beat existing distillation baselines by an average of 5.22 percent and were more successful at revising their own flawed steps.
It is a narrow, technical fix, not a new model or product launch, but if this kind of mid-reasoning revision training generalizes beyond the benchmarks tested here, it points toward models that correct themselves as they think rather than just double-checking their work afterward.