Mixture-of-experts models are getting more capable without proportionally more compute, but only if the router assigns tokens to the right experts. A new paper proposes teaching that router to care about being wrong.
Today's sparse MoE routers score how well each expert matches a token, then send the token to the top few matches. Those affinity scores are trained through the general language-modeling loss and kept balanced by load-based regularizers, but nothing ties them directly to which tokens the model is actually getting wrong. Researchers introduce two supervision methods that close that gap: one adds a small head that predicts each expert's likely error and uses it to dampen the routing score before top-K selection, the other tunes the router's own affinities directly against the model's objective, both using either Itakura-Saito divergence or an exponential negative log-likelihood loss. Tested across two sparse MoE backbones on four multiple-choice QA benchmarks, both methods hold the same sparse compute budget and expert-combination policy as the baseline. On the Granite backbone specifically, the team reports roughly a 2.3 percentage point average accuracy gain over a parameter-matched baseline router, and the stronger-supervision variant pushes the ARC-Challenge gain to 2.94 points, a result the plain baseline routing method did not reach.
The interesting part isn't the score bump. It's that routing, the part of an MoE model that decides where compute goes, has mostly been optimized as a side effect of the main training loss. Giving it its own error signal, without adding inference cost or touching the sparsity budget, is a cheap lever as more labs default to MoE architectures to scale models affordably.
Two to three points on multiple-choice benchmarks, on one named backbone, is a modest result to hang a routing rethink on. Whether it holds on larger production MoEs or open-ended generation is a question this paper doesn't answer.