AI/ mixture-of-experts · llm-training · ai-research · model-architecture

New Method Fixes Mixture-of-Experts Upcycling Collapse

A new training trick lets smaller mixture-of-experts models match bigger rivals by forcing tokens to pull from different domain specialists.

A new technique called DivMoE fixes a routing problem that has made fine-grained mixture-of-experts upcycling nearly useless.

Researchers describe DivMoE, a method for converting dense language models into sparse mixture-of-experts models without training from scratch. The team found that when fine-grained experts are split from a single source model, the router collapses and accuracy drops close to random guessing - on Qwen3-1.7B, one existing method scored just 23.2% average accuracy across 15 benchmarks, barely above training from scratch at 22.2%. DivMoE instead builds experts from models that have already gone through domain-specific continual pre-training, then forces each token to draw from experts in different domain groups rather than letting the router pick favorites. Across two base models and 15 benchmarks, DivMoE hit 55.6% average accuracy versus 51.6% for the best existing upcycling baseline, and beat the dense base model on every single benchmark.

That gap matters because upcycling exists to make scaling cheaper, not to produce a model that performs worse than the one you started with. A fine-grained DivMoE model with 12 billion parameters matched the 16 billion parameter Moonlight-MoE at 64.5% average accuracy after fine-tuning, suggesting the routing fix - not just extra experts - was the missing piece for getting fine-grained MoE to pay off.

Mixture-of-experts has been the industry's favorite way to grow models without growing compute bills; this paper is a reminder that the cheap route to get there has had a hidden tax most labs were paying without noticing.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →