A new pruning method lets engineers cut half the experts out of a mixture-of-experts AI model while preserving far more accuracy than standard pruning tricks.
MoE models split their work across specialized sub-networks called "experts," but every expert still has to sit loaded in memory even though only a few fire per token - that's the waste the researchers targeted. Their fix: run a brief, cheap fine-tuning pass using LoRA adapters that touch only the router (the part that decides which experts handle which tokens) - just 0.002% of total parameters - then prune whichever experts' router weights moved the least. On Mixtral-8x7B-Instruct, that approach retained 27.54% accuracy on the MMLU-Pro benchmark after removing half the experts, compared with roughly 16% for the standard random or magnitude-based pruning baselines; larger adapters pushed retained accuracy as high as 28.76%. The same criterion transferred to Qwen1.5-MoE tuned for math reasoning, holding 49.7% mean accuracy across eleven benchmarks after halving its experts, while cutting memory use by 49% and per-token latency by 37%.
This matters because mixture-of-experts architectures are now the default for frontier-scale models, and memory - not raw compute - is often the real deployment bottleneck. A method that reliably trims experts without a full retraining run lowers that cost. It also turns a purely theoretical guarantee into something practical: the earlier version of this idea needed full fine-tuning to find the prunable experts, which is exactly the expensive step MoE pruning is supposed to let you skip.
Still, every number here is a comparison against other pruning methods, not against the un-pruned model itself - the paper doesn't say how much accuracy disappears relative to keeping every expert, so "beats magnitude pruning" and "barely worse than the full model" remain two different claims.