[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-training-method-lets-ai-models-coach-themselves-on-math":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},9228,"new-training-method-lets-ai-models-coach-themselves-on-math","New Training Method Lets AI Models Coach Themselves on Math","A new technique called STEPS nudges AI models through tricky reasoning steps only when needed, boosting math accuracy without the usual overcorrection.","Researchers have found a smarter way to let AI models tutor themselves through math problems, without the self-correction backfiring.\n\nThe technique, called STEPS (Selective On-Policy Self-Distillation), has a model act as its own teacher, using privileged context only available during training to coach its own reasoning. Earlier versions of this self-distillation approach applied that coaching to every token in a response, which researchers found can over-constrain the model and bake in biases from information it will not have at inference time. STEPS instead targets only the critical spans in a model's practice answers, applying one kind of correction to spans that are on track and another to spans heading toward an error, then phases out entirely in favor of GRPO (Group Relative Policy Optimization), the standard reward-based reasoning method, after a short window.\n\nAcross four models from three model families, STEPS beat plain GRPO training on accuracy, including a 2.76 percentage point gain on Qwen3-8B across four math benchmarks plus the GPQA-Diamond science test, for roughly 2.7% more training compute. Even when the model graded itself instead of using a stronger outside annotator, it still gained 1.90 points, meaning the technique does not depend on access to a better teacher model.\n\nSelf-distillation has been sold as a way to squeeze more reasoning out of models without bigger datasets or more parameters. STEPS's actual contribution is narrower and more useful: knowing when to stop hand-holding.","[\"ai\",\"reasoning\",\"machine-learning\",\"research\"]","2026-10-01T04:00:00.000Z","2026-10-02T02:59:09.163Z","2026-10-02T02:59:14.654Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Define the acronyms on first use — spell out what STEPS stands for (Selective On-Policy Self-Distillation) and what GRPO means (Group Relative Policy Optimization) before using them bare in the body.","resolved","ai",[30,32,33,34],"reasoning","machine-learning","research",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2605.10194",0,{"sections":41},[42,45,49,53,58,63,67,72,77,81,86,91,96,101],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",5629,{"name":46,"slug":47,"count":48,"latest_published_at":18},"Security","security",816,{"name":50,"slug":51,"count":52,"latest_published_at":18},"Policy","policy",430,{"name":54,"slug":55,"count":56,"latest_published_at":57},"Deals","deals",298,"2026-09-30T21:00:26.000Z",{"name":59,"slug":60,"count":61,"latest_published_at":62},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":64,"slug":65,"count":66,"latest_published_at":18},"Science","science",163,{"name":68,"slug":69,"count":70,"latest_published_at":71},"Consumer Tech","consumer-tech",149,"2026-09-30T22:57:11.000Z",{"name":73,"slug":74,"count":75,"latest_published_at":76},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":78,"slug":79,"count":75,"latest_published_at":80},"Software","software","2026-09-30T21:41:11.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":102,"slug":103,"count":104,"latest_published_at":105},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]