[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-ai-training-method-sorts-fixable-mistakes-from-fatal-ones":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},8156,"ai-training-method-sorts-fixable-mistakes-from-fatal-ones","AI Training Method Sorts Fixable Mistakes From Fatal Ones","A new training technique teaches language models to tell recoverable errors from dead ends, lifting reasoning scores on AIME and GPQA benchmarks.","A new training method teaches AI reasoning models to distinguish a fixable misstep from a dead end - and to stop punishing the ones that were never really wrong.\n\nThe technique, called counterfactual recoverability, targets on-policy distillation, where a smaller \"student\" model learns by having a stronger \"teacher\" model correct its reasoning as it goes. Instead of treating every deviation from the teacher's path as an error to erase, researchers replay each error state down two paths: let the teacher finish the thought, or roll the student back and let it retry. Whichever path succeeds determines the label - recoverable, irreversible-but-avoidable, or ambiguous - and that label decides whether training keeps the trajectory, discards it, or handles it the conventional way. A recoverability score built from these branch tests predicted fixability with an AUC of 1.000 on AIME math problems, versus 0.392 for simply measuring how far the student had strayed.\n\nThat distinction changed outcomes. Training guided by recoverability hit 0.578 success on held-out AIME2025 problems, against 0.517 for the strongest baseline. Average scores across repeated attempts (average@32) rose too: AIME2024-2025 climbed from 0.2656 to 0.3125, about 4.7 percentage points, and GPQA-Diamond went from 0.2702 to 0.3070, about 3.7 points.\n\nThe real finding here isn't the score bump - it's confirmation that most distillation setups have been throwing away salvageable reasoning. Punishing any drift from a teacher model, regardless of whether the student could have talked its way back to the right answer, wastes training signal on turns that were never fatal.\n\nStill, a few points on two benchmarks is not a breakthrough, and an AUC of 1.000 on branch diagnostics is the kind of suspiciously clean number that belongs in a follow-up study before it belongs in a pitch deck.","[\"ai\",\"machine-learning\",\"llm-training\",\"research\"]","2026-09-28T04:00:00.000Z","2026-09-28T11:47:23.341Z","2026-09-28T11:47:30.337Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Fix the arithmetic error in the average-score claim: AIME2024-2025 average@32 rose from 0.2656 to 0.3125 (~4.7 points) and GPQA-Diamond average@32 rose from 0.2702 to 0.3070 (~3.7 points), not both 'roughly five to six points' as currently stated — cite the precise before\u002Fafter figures for each benchmark instead of a rounded range.","resolved","ai",[30,32,33,34],"machine-learning","llm-training","research",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.04408",0,{"sections":41},[42,45,49,54,59,64,68,73,78,83,88,93,97,102],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",4806,{"name":46,"slug":47,"count":48,"latest_published_at":18},"Security","security",762,{"name":50,"slug":51,"count":52,"latest_published_at":53},"Policy","policy",399,"2026-09-27T18:39:02.000Z",{"name":55,"slug":56,"count":57,"latest_published_at":58},"Deals","deals",261,"2026-09-27T15:30:35.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":63},"Hardware","hardware",188,"2026-09-27T20:46:36.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":18},"Science","science",151,{"name":69,"slug":70,"count":71,"latest_published_at":72},"Consumer Tech","consumer-tech",135,"2026-09-26T14:30:00.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":82},"Dev Tools","dev-tools",84,"2026-09-26T04:20:58.000Z",{"name":84,"slug":85,"count":86,"latest_published_at":87},"Startups","startups",76,"2026-09-25T18:33:59.000Z",{"name":89,"slug":90,"count":91,"latest_published_at":92},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":94,"slug":95,"count":91,"latest_published_at":96},"General","general","2026-09-26T17:02:42.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Reviews","reviews",30,"2026-09-24T20:07:31.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]