[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-method-lets-ai-models-pick-which-teacher-trains-each-token":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},6445,"new-method-lets-ai-models-pick-which-teacher-trains-each-token","New Method Lets AI Models Pick Which Teacher Trains Each Token","The paper 'Who Teaches Which Token?' (arXiv:2609.15404) argues AI experts should grade individual tokens, not entire answers, to train smarter models.","A new training method teaches AI models token by token, not answer by answer.\n\nThe paper \"Who Teaches Which Token? Verifier-Gated Multi-Expert On-Policy Distillation for Scientific Reasoning\" (arXiv:2609.15404) proposes VG-OPD, a way to fold several specialist \"teacher\" models into one student model. Standard multi-teacher distillation hands each training prompt to a single domain expert and treats every token in that expert's response as equally important. The paper's authors found that assumption doesn't hold: a teacher's useful signal is concentrated in a handful of tokens, not spread evenly across a response. VG-OPD instead runs a verifier that checks each teacher's answer against specific criteria, then only lets that teacher's guidance shape the exact tokens where it demonstrably helps.\n\nTested on 4B and 8B parameter student models across seven scientific reasoning benchmarks, VG-OPD ranked first on five of them, with the largest gains on knowledge-intensive science questions. The paper's own ablations show those gains come from placing supervision on the right tokens, not from adding more teachers or more distillation loss. Misplacing the same supervision budget, the authors report, was the single most damaging change they tested, worse than skipping distillation altogether.\n\nThat is a useful data point for anyone building many-teachers-one-student training pipelines. More experts and more distillation are not automatically better, and sloppy, unverified distillation can drag a model below the performance of plain reinforcement learning.","[\"ai\",\"machine-learning\",\"distillation\",\"research\"]","2026-09-16T04:00:00.000Z","2026-09-17T16:47:14.268Z","2026-09-17T16:47:26.819Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Attribute the technique to its actual source (cite the arXiv paper, e.g. 'Who Teaches Which Token?', arXiv:2609.15404) instead of the unnamed 'Researchers describe...' framing, so the core claims aren't sourced to an anonymous group.","resolved","ai",[30,32,33,34],"machine-learning","distillation","research",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.15404",0,{"sections":41},[42,46,51,56,61,65,69,74,79,83,88,93,98,103],{"name":43,"slug":30,"count":44,"latest_published_at":45},"AI",3852,"2026-09-17T08:27:09.000Z",{"name":47,"slug":48,"count":49,"latest_published_at":50},"Security","security",648,"2026-09-17T04:00:00.000Z",{"name":52,"slug":53,"count":54,"latest_published_at":55},"Policy","policy",338,"2026-09-11T04:00:00.000Z",{"name":57,"slug":58,"count":59,"latest_published_at":60},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":62,"slug":63,"count":64,"latest_published_at":50},"Hardware","hardware",154,{"name":66,"slug":67,"count":68,"latest_published_at":50},"Science","science",114,{"name":70,"slug":71,"count":72,"latest_published_at":73},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":75,"slug":76,"count":77,"latest_published_at":78},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":80,"slug":81,"count":82,"latest_published_at":50},"Dev Tools","dev-tools",73,{"name":84,"slug":85,"count":86,"latest_published_at":87},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":89,"slug":90,"count":91,"latest_published_at":92},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":94,"slug":95,"count":96,"latest_published_at":97},"General","general",41,"2026-09-08T01:57:23.000Z",{"name":99,"slug":100,"count":101,"latest_published_at":102},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":104,"slug":105,"count":106,"latest_published_at":107},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]