[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-researchers-refine-how-ai-models-learn-from-their-own-outputs":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},8724,"researchers-refine-how-ai-models-learn-from-their-own-outputs","Researchers Refine How AI Models Learn From Their Own Outputs","A new preprint proposes SR-OPSD, which blends a model's current self with a frozen earlier version to guide training on reasoning and code tasks.","A new training trick wants AI models to become their own teachers - and the researchers behind it say the approach holds up across several model families and task types.\n\nThe method, called Self-Referenced On-Policy Self-Distillation (SR-OPSD), is described in a paper posted to arXiv (2608.09745). It fine-tunes language models using dense, token-by-token feedback rather than relying only on sparse reinforcement-learning rewards. To build its training target, the method blends a moving self-teacher - a version of the model's own evolving parameters - with a frozen snapshot of the model's original policy, then pulls the student toward that blend using a statistical distance measure called Renyi divergence. Two adjustable settings control the process: one weights how much the self-teacher contributes, the other tunes how sharply the training signal reacts to gaps between target and student probabilities. The authors tested it on scientific reasoning, tool use, math, and code-generation tasks across multiple model families and sizes.\n\nHere's the catch: the paper's abstract calls the results strong performance but doesn't publish specific benchmark scores or comparison numbers, so there's no way yet to independently judge how big the gains are. More telling is an ablation result buried in the same abstract: anchoring the student to its frozen former self can help or hurt depending entirely on how the projection step is configured - which suggests this is a fragile knob to tune, not a guaranteed upgrade.\n\nIt's one more attempt to squeeze extra signal out of reinforcement learning without burning more compute - a trend that's been building all year, but one that still needs numbers, not adjectives, to prove itself.","[\"ai\",\"machine-learning\",\"llm-training\",\"research\"]","2026-09-30T04:00:00.000Z","2026-09-30T21:57:18.551Z","2026-09-30T21:57:24.236Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The piece claims the method 'noticeably improves performance' and demonstrates 'strong performance' without citing any actual benchmark figures, metric definitions, or comparison numbers from the paper, and omits the arXiv identifier needed to verify the claim — either pull the specific results\u002Fmetrics from the paper or soften the language to match what the abstract actually supports, and cite the paper (arXiv:2608.09745) directly.","resolved","ai",[30,32,33,34],"machine-learning","llm-training","research",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.09745",0,{"sections":41},[42,46,51,56,61,65,69,73,78,82,87,92,97,102],{"name":43,"slug":30,"count":44,"latest_published_at":45},"AI",5214,"2026-09-30T13:00:00.000Z",{"name":47,"slug":48,"count":49,"latest_published_at":50},"Security","security",793,"2026-09-30T12:55:00.000Z",{"name":52,"slug":53,"count":54,"latest_published_at":55},"Policy","policy",419,"2026-09-30T12:24:32.000Z",{"name":57,"slug":58,"count":59,"latest_published_at":60},"Deals","deals",292,"2026-09-30T14:15:18.000Z",{"name":62,"slug":63,"count":64,"latest_published_at":45},"Hardware","hardware",196,{"name":66,"slug":67,"count":68,"latest_published_at":18},"Science","science",155,{"name":70,"slug":71,"count":72,"latest_published_at":45},"Consumer Tech","consumer-tech",144,{"name":74,"slug":75,"count":76,"latest_published_at":77},"Dev Tools","dev-tools",91,"2026-09-30T12:58:00.000Z",{"name":79,"slug":80,"count":76,"latest_published_at":81},"Software","software","2026-09-25T20:55:00.000Z",{"name":83,"slug":84,"count":85,"latest_published_at":86},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]