[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-researchers-teach-ai-models-to-reason-like-experts-no-prompt-needed":10,"sections":34},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":29,"feedback":33,"feedback_at":22,"cost_usd":33,"total_tokens":33},9180,"researchers-teach-ai-models-to-reason-like-experts-no-prompt-needed","Researchers Teach AI Models To Reason Like Experts, No Prompt Needed","A new self-distillation technique, OPSRD, trains AI models to reason like experts internally, improving math accuracy without role prompts at runtime.","A new training method squeezes expert-level reasoning out of AI models without asking them to role-play an expert at run time.\n\nThe technique, called OPSRD (on-policy self-role distillation), has a model teach itself. A frozen copy of the same base model, given an expert-role prompt, generates probability distributions over likely next words. A second, role-free version of the model generates its own answer, and the system compares the two sets of predictions at each step, not just the final result. Forward KL divergence nudges the student toward alternatives it had been underweighting, with clipping to stop any single word choice from dominating the update. The model only learns from the half of its predictions where it was most uncertain, rather than every token it generated.\n\nThis matters because role prompting - telling a model to act as an expert mathematician - is a cheap trick that works unevenly and adds tokens to every request. OPSRD tries to bake that boost into the model's weights permanently, so smaller models can get the benefit without paying the inference cost or prompt-engineering tax every time. Tested on Qwen3 models at 1.7B, 4B, and 8B parameters across three competition-math benchmarks, the method beat the un-prompted base models, with forward KL outperforming the other divergence measures tried at every size.\n\nThe results are promising but narrow - three math benchmarks and one model family is a small sample, and beating a model with no role prompt is a lower bar than beating the role prompt itself.","[\"ai\",\"llm-training\",\"machine-learning\",\"research\"]","2026-10-01T04:00:00.000Z","2026-10-02T00:15:40.418Z","2026-10-02T00:15:45.760Z","published",null,[],"ai",[24,26,27,28],"llm-training","machine-learning","research",[30],{"name":31,"url":32},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.39884",0,{"sections":35},[36,39,43,47,52,57,61,66,71,75,80,85,90,95],{"name":37,"slug":24,"count":38,"latest_published_at":18},"AI",5599,{"name":40,"slug":41,"count":42,"latest_published_at":18},"Security","security",815,{"name":44,"slug":45,"count":46,"latest_published_at":18},"Policy","policy",430,{"name":48,"slug":49,"count":50,"latest_published_at":51},"Deals","deals",298,"2026-09-30T21:00:26.000Z",{"name":53,"slug":54,"count":55,"latest_published_at":56},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":58,"slug":59,"count":60,"latest_published_at":18},"Science","science",163,{"name":62,"slug":63,"count":64,"latest_published_at":65},"Consumer Tech","consumer-tech",149,"2026-09-30T22:57:11.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":70},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":72,"slug":73,"count":69,"latest_published_at":74},"Software","software","2026-09-30T21:41:11.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":84},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]