[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-training-method-tackles-a-flaw-in-how-ai-models-learn-to-reason":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},8486,"new-training-method-tackles-a-flaw-in-how-ai-models-learn-to-reason","New Training Method Tackles a Flaw in How AI Models Learn to Reason","SIPO pairs correct and incorrect model outputs as dueling self-teachers, canceling shared bias to give token-level feedback that plain reward signals miss.","A new reinforcement learning method called SIPO claims to close a gap in how AI models learn to reason, without hiring a second model to watch over the first.\n\nSIPO stands for self-instructing policy optimization. It targets a known problem with RLVR, the standard method for training models on tasks with a checkable answer: RLVR only rewards or punishes the final answer, so it gives no credit or blame for individual steps in a long reasoning chain. A follow-up approach, on-policy self-distillation, tried fixing that by having a model grade its own intermediate steps, but that self-grading model tends to be overconfident and overpenalizes long reasoning traces. SIPO's fix is to sample several attempts per problem, build two self-teacher contexts for each attempt (one paired with the correct answer, one paired with the group's actual mistakes), and score each token by the gap between the two, on the idea that shared biases in both contexts cancel out.\n\nThat token-level scoring matters because reward-only training has a blind spot: when every attempt in a batch fails, there's nothing to compare and the usual training signal disappears. SIPO says it still produces a learning signal in those all-fail cases, which is exactly when models most need correction.\n\nThe paper's abstract claims SIPO beat both RLVR and the earlier self-distillation baseline on 'multiple reasoning and code-generation benchmarks,' but it doesn't name those benchmarks or give margin-of-improvement numbers in the summary - a gap worth closing before anyone treats this as a drop-in upgrade rather than a promising idea.","[\"reinforcement-learning\",\"llm-training\",\"ai-research\",\"self-distillation\"]","2026-09-30T04:00:00.000Z","2026-09-30T06:43:25.841Z","2026-09-30T06:43:32.516Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The claim that SIPO 'beat both RLVR and the earlier self-distillation baseline' is a performance result stated with no actual figures, named benchmarks, or margin of improvement — pull the specific benchmark names and comparison numbers from the paper (not just the abstract) or explicitly note they aren't available, since an unquantified 'it won' claim isn't verifiable.","resolved","ai",[32,33,34,35],"reinforcement-learning","llm-training","ai-research","self-distillation",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.36742",0,{"sections":42},[43,46,50,54,59,64,69,74,79,84,89,94,99,104],{"name":44,"slug":30,"count":45,"latest_published_at":18},"AI",5028,{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",780,{"name":51,"slug":52,"count":53,"latest_published_at":18},"Policy","policy",417,{"name":55,"slug":56,"count":57,"latest_published_at":58},"Deals","deals",284,"2026-09-29T21:00:00.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":63},"Hardware","hardware",194,"2026-09-29T13:16:04.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":68},"Science","science",154,"2026-09-28T13:19:18.000Z",{"name":70,"slug":71,"count":72,"latest_published_at":73},"Consumer Tech","consumer-tech",142,"2026-09-29T18:38:03.000Z",{"name":75,"slug":76,"count":77,"latest_published_at":78},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":80,"slug":81,"count":82,"latest_published_at":83},"Dev Tools","dev-tools",89,"2026-09-29T17:15:00.000Z",{"name":85,"slug":86,"count":87,"latest_published_at":88},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":90,"slug":91,"count":92,"latest_published_at":93},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":95,"slug":96,"count":97,"latest_published_at":98},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":100,"slug":101,"count":102,"latest_published_at":103},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":105,"slug":106,"count":107,"latest_published_at":108},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]