A new technique called MoRE lets AI models pack in dramatically more "experts" without paying the usual computational tax.
Mixture-of-experts models split their work across many specialized sub-networks, or "experts," using a router to pick which ones process each token, and the industry trend has been toward more, smaller experts. The paper argues the standard router does not scale with that trend, since its cost grows with both hidden dimension and expert count, becoming the bottleneck once expert counts get large. MoRE compresses the router's weight matrix to a much lower rank, cutting that cost and, per the authors' math, permitting roughly h/r times more experts at the same compute budget - paired with a custom Triton kernel so the savings hold at inference, not just on paper. Tested on a synthetic phonebook-memorization task and on knowledge-intensive Q&A benchmarks after pretraining, MoRE reportedly improves results on both while matching baseline reasoning performance, though the abstract does not publish the specific benchmark scores or comparison figures behind that claim.
Router efficiency is a quietly important problem: as labs push toward hundreds of tiny experts per layer, a routing step whose cost scales linearly with expert count turns into real overhead at inference time. A cheaper router is the unglamorous kind of fix that can matter more for production costs than another benchmark headline, since it changes what's economically deployable rather than just what's trainable in a lab.
Code is on GitHub, but without independently verified benchmark numbers, this reads as promising engineering rather than a proven performance leap.