[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-researchers-let-ai-models-tutor-themselves-with-future-versions":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},8505,"researchers-let-ai-models-tutor-themselves-with-future-versions","Researchers Let AI Models Tutor Themselves With Future Versions","A new self-distillation method trains a model briefly into the future, then has that future version teach its present self, sharply boosting math accuracy.","A new training technique lets an AI model learn from a version of itself that hasn't been built yet.\n\nResearchers describe Bootstrapped On-Policy Self-Distillation, or B-OPSD, in a new paper. The method temporarily pushes a language model's training forward, freezes that more-advanced checkpoint as a 'future teacher,' then rewinds the actual model back to its starting point and has the future teacher supervise it token by token. It builds on on-policy self-distillation, where a more-informed version of a model helps train a less-informed one on its own generated outputs. The team tested B-OPSD on Qwen3-4B and Qwen3-8B, two open-source language models, on math reasoning problems. In the paper's rollout-privileged setting - where the teacher gets extra information while generating example solutions - accuracy climbed from 27.50 percent to 41.30 percent on the 4-billion-parameter model, and from 48.80 percent to 64.44 percent on the 8-billion-parameter model, compared with standard self-distillation.\n\nThat's a real jump for a technique that adds no new external data - it just recycles the model's own optimization progress. If a model can reliably distill its future gains back into its present state, it chips away at the assumption that self-improvement needs bigger models, more compute, or fresh human-labeled feedback to keep climbing.\n\nMath problems are also the easiest place for a trick like this to shine, since answers are checkable and rewards are unambiguous. Whether the same bootstrapping holds up on messier, harder-to-grade tasks is the real test.","[\"ai\",\"self-distillation\",\"llm-training\",\"machine-learning\"]","2026-09-30T04:00:00.000Z","2026-09-30T07:56:10.799Z","2026-09-30T07:56:16.169Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"State what the 27.50→41.30 and 48.80→64.44 numbers actually measure (e.g., accuracy percentage on which specific math benchmark) since citing bare scores without identifying the metric or benchmark leaves readers unable to judge how big these gains really are.","resolved","ai",[30,32,33,34],"self-distillation","llm-training","machine-learning",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.37132",0,{"sections":41},[42,45,49,53,58,63,68,73,78,83,88,93,98,103],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",5028,{"name":46,"slug":47,"count":48,"latest_published_at":18},"Security","security",780,{"name":50,"slug":51,"count":52,"latest_published_at":18},"Policy","policy",417,{"name":54,"slug":55,"count":56,"latest_published_at":57},"Deals","deals",284,"2026-09-29T21:00:00.000Z",{"name":59,"slug":60,"count":61,"latest_published_at":62},"Hardware","hardware",194,"2026-09-29T13:16:04.000Z",{"name":64,"slug":65,"count":66,"latest_published_at":67},"Science","science",154,"2026-09-28T13:19:18.000Z",{"name":69,"slug":70,"count":71,"latest_published_at":72},"Consumer Tech","consumer-tech",142,"2026-09-29T18:38:03.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":82},"Dev Tools","dev-tools",89,"2026-09-29T17:15:00.000Z",{"name":84,"slug":85,"count":86,"latest_published_at":87},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":89,"slug":90,"count":91,"latest_published_at":92},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":94,"slug":95,"count":96,"latest_published_at":97},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":99,"slug":100,"count":101,"latest_published_at":102},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":104,"slug":105,"count":106,"latest_published_at":107},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]