A new training method teaches AI reasoning models to distinguish a fixable misstep from a dead end - and to stop punishing the ones that were never really wrong.
The technique, called counterfactual recoverability, targets on-policy distillation, where a smaller "student" model learns by having a stronger "teacher" model correct its reasoning as it goes. Instead of treating every deviation from the teacher's path as an error to erase, researchers replay each error state down two paths: let the teacher finish the thought, or roll the student back and let it retry. Whichever path succeeds determines the label - recoverable, irreversible-but-avoidable, or ambiguous - and that label decides whether training keeps the trajectory, discards it, or handles it the conventional way. A recoverability score built from these branch tests predicted fixability with an AUC of 1.000 on AIME math problems, versus 0.392 for simply measuring how far the student had strayed.
That distinction changed outcomes. Training guided by recoverability hit 0.578 success on held-out AIME2025 problems, against 0.517 for the strongest baseline. Average scores across repeated attempts (average@32) rose too: AIME2024-2025 climbed from 0.2656 to 0.3125, about 4.7 percentage points, and GPQA-Diamond went from 0.2702 to 0.3070, about 3.7 points.
The real finding here isn't the score bump - it's confirmation that most distillation setups have been throwing away salvageable reasoning. Punishing any drift from a teacher model, regardless of whether the student could have talked its way back to the right answer, wastes training signal on turns that were never fatal.
Still, a few points on two benchmarks is not a breakthrough, and an AUC of 1.000 on branch diagnostics is the kind of suspiciously clean number that belongs in a follow-up study before it belongs in a pitch deck.