A new technique called DivMoE fixes a routing problem that has made fine-grained mixture-of-experts upcycling nearly useless.
Researchers describe DivMoE, a method for converting dense language models into sparse mixture-of-experts models without training from scratch. The team found that when fine-grained experts are split from a single source model, the router collapses and accuracy drops close to random guessing - on Qwen3-1.7B, one existing method scored just 23.2% average accuracy across 15 benchmarks, barely above training from scratch at 22.2%. DivMoE instead builds experts from models that have already gone through domain-specific continual pre-training, then forces each token to draw from experts in different domain groups rather than letting the router pick favorites. Across two base models and 15 benchmarks, DivMoE hit 55.6% average accuracy versus 51.6% for the best existing upcycling baseline, and beat the dense base model on every single benchmark.
That gap matters because upcycling exists to make scaling cheaper, not to produce a model that performs worse than the one you started with. A fine-grained DivMoE model with 12 billion parameters matched the 16 billion parameter Moonlight-MoE at 64.5% average accuracy after fine-tuning, suggesting the routing fix - not just extra experts - was the missing piece for getting fine-grained MoE to pay off.
Mixture-of-experts has been the industry's favorite way to grow models without growing compute bills; this paper is a reminder that the cheap route to get there has had a hidden tax most labs were paying without noticing.