[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-pruning-method-shrinks-ai-models-without-full-retraining":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},9327,"new-pruning-method-shrinks-ai-models-without-full-retraining","New Pruning Method Shrinks AI Models Without Full Retraining","A cheap fine-tuning trick identifies prunable AI experts, beating standard pruning methods by double-digit accuracy margins.","A new pruning method lets engineers cut half the experts out of a mixture-of-experts AI model while preserving far more accuracy than standard pruning tricks.\n\nMoE models split their work across specialized sub-networks called \"experts,\" but every expert still has to sit loaded in memory even though only a few fire per token - that's the waste the researchers targeted. Their fix: run a brief, cheap fine-tuning pass using LoRA adapters that touch only the router (the part that decides which experts handle which tokens) - just 0.002% of total parameters - then prune whichever experts' router weights moved the least. On Mixtral-8x7B-Instruct, that approach retained 27.54% accuracy on the MMLU-Pro benchmark after removing half the experts, compared with roughly 16% for the standard random or magnitude-based pruning baselines; larger adapters pushed retained accuracy as high as 28.76%. The same criterion transferred to Qwen1.5-MoE tuned for math reasoning, holding 49.7% mean accuracy across eleven benchmarks after halving its experts, while cutting memory use by 49% and per-token latency by 37%.\n\nThis matters because mixture-of-experts architectures are now the default for frontier-scale models, and memory - not raw compute - is often the real deployment bottleneck. A method that reliably trims experts without a full retraining run lowers that cost. It also turns a purely theoretical guarantee into something practical: the earlier version of this idea needed full fine-tuning to find the prunable experts, which is exactly the expensive step MoE pruning is supposed to let you skip.\n\nStill, every number here is a comparison against other pruning methods, not against the un-pruned model itself - the paper doesn't say how much accuracy disappears relative to keeping every expert, so \"beats magnitude pruning\" and \"barely worse than the full model\" remain two different claims.","[\"mixture-of-experts\",\"model-pruning\",\"ai-research\",\"efficiency\"]","2026-10-01T04:00:00.000Z","2026-10-02T09:42:14.844Z","2026-10-02T09:42:15.105Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The dek claims 'little accuracy loss' but the body never states baseline (unpruned) accuracy for Mixtral or Qwen, so the only accuracy figures given (27.54% and 49.7%) compare against other pruning methods, not against the original model — either add the baseline accuracy to substantiate 'little loss' or rewrite the dek to only claim the method beats alternative pruning approaches.","resolved","ai",[32,33,34,35],"mixture-of-experts","model-pruning","ai-research","efficiency",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.07890",0,{"sections":42},[43,47,52,57,62,67,71,76,81,85,89,94,99,104],{"name":44,"slug":30,"count":45,"latest_published_at":46},"AI",5694,"2026-10-01T12:05:27.000Z",{"name":48,"slug":49,"count":50,"latest_published_at":51},"Security","security",821,"2026-10-01T14:00:00.000Z",{"name":53,"slug":54,"count":55,"latest_published_at":56},"Policy","policy",431,"2026-10-01T11:08:42.000Z",{"name":58,"slug":59,"count":60,"latest_published_at":61},"Deals","deals",311,"2026-10-01T14:19:00.000Z",{"name":63,"slug":64,"count":65,"latest_published_at":66},"Hardware","hardware",197,"2026-10-01T11:37:06.000Z",{"name":68,"slug":69,"count":70,"latest_published_at":18},"Science","science",165,{"name":72,"slug":73,"count":74,"latest_published_at":75},"Consumer Tech","consumer-tech",150,"2026-10-01T11:59:27.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":82,"slug":83,"count":79,"latest_published_at":84},"Software","software","2026-09-30T21:41:11.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":51},"Startups","startups",86,{"name":90,"slug":91,"count":92,"latest_published_at":93},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":95,"slug":96,"count":97,"latest_published_at":98},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":100,"slug":101,"count":102,"latest_published_at":103},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":105,"slug":106,"count":107,"latest_published_at":108},"How-To","how-to",7,"2026-10-01T09:00:00.000Z"]