[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-routing-method-beats-full-parameter-rl-training":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},9118,"new-routing-method-beats-full-parameter-rl-training","New Routing Method Beats Full-Parameter RL Training","A tiny routing module outperforms full-parameter reinforcement learning on math reasoning tasks while training just 0.466 percent of the model's parameters.","A new technique trains less than half a percent of a language model's parameters and still beats training all of them on math reasoning.\n\nResearchers describe T-Router, a parameter-efficient reinforcement learning method that adds a compact routing layer to a pretrained 8.95-billion-parameter model instead of retraining its weights. The router keeps a compressed bank of earlier computations and uses a small recurrent controller to decide which past results a given layer should reuse, and how heavily to weight them. Only this interface, 41.73 million parameters, or 0.466% of the backbone, gets trained using correctness-based rewards, while the backbone itself stays frozen. After reinforcement learning on the GSM8K math dataset, T-Router scored 83.64 on a composite math benchmark, versus 73.79 for a full-parameter reinforcement learning baseline called GRPO run on the same backbone, and it also beat LoRA, a popular lightweight fine-tuning technique, which scored 77.28 at a similar parameter budget. On a separate AIME benchmark, accuracy rose from 48.33% with the baseline to 60.56%.\n\nThe result matters because it undercuts a core assumption in reasoning-model training: that sharpening a model's reasoning requires updating most or all of its weights. If a sliver of well-targeted parameters can outperform full-parameter reinforcement learning, that is a real cost and infrastructure win for anyone training reasoning models on a budget, not just an academic curiosity.\n\nWorth noting: these gains are measured on math benchmarks like GSM8K and AIME, not the messier reasoning tasks most products actually need, and the comparisons come from the paper's own evaluation rounds rather than independent replication.","[\"reinforcement-learning\",\"parameter-efficient-ai\",\"math-reasoning\",\"ai-research\"]","2026-10-01T04:00:00.000Z","2026-10-01T20:57:07.102Z","2026-10-01T20:57:10.053Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The headline ('Full Model Retraining') and dek ('full fine-tuning') mischaracterize the actual baseline described in the body, which is a full-parameter reinforcement learning run (GRPO) on an already-pretrained model, not full fine-tuning or retraining from scratch — rewrite the headline and dek to say the method beats full-parameter RL training, matching the body's own framing.","resolved","ai",[32,33,34,35],"reinforcement-learning","parameter-efficient-ai","math-reasoning","ai-research",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.39109",0,{"sections":42},[43,46,50,54,59,64,68,73,78,82,87,92,97,102],{"name":44,"slug":30,"count":45,"latest_published_at":18},"AI",5570,{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",815,{"name":51,"slug":52,"count":53,"latest_published_at":18},"Policy","policy",430,{"name":55,"slug":56,"count":57,"latest_published_at":58},"Deals","deals",298,"2026-09-30T21:00:26.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":63},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":18},"Science","science",163,{"name":69,"slug":70,"count":71,"latest_published_at":72},"Consumer Tech","consumer-tech",149,"2026-09-30T22:57:11.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":79,"slug":80,"count":76,"latest_published_at":81},"Software","software","2026-09-30T21:41:11.000Z",{"name":83,"slug":84,"count":85,"latest_published_at":86},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]