AI/ ai · interpretability · mixture-of-experts

A New Way to Peek Inside Mixture of Experts AI Models

RouterInterp shows AI mixture-of-experts models split tasks into tiny overlapping skills rather than tidy subject specialties.

A new interpretability method just found that AI experts don't specialize the way everyone assumed.

Researchers built a tool called RouterInterp to figure out why mixture-of-experts models send certain tokens to certain expert networks. The popular theory held that each expert handles one coherent topic, like code or biology. Testing that idea against gpt-oss-20b, the team found the opposite: experts specialize in scattered, fine-grained features that don't map to any single domain. They call this the Superposed Specialisation Hypothesis, and RouterInterp uses sparse autoencoders to pull out the specific features driving each routing decision, then turns them into plain-language explanations.

Mixture-of-experts architecture is now a standard way to scale AI models without scaling compute costs on every token. But nobody could cleanly explain what the experts were actually doing, which made auditing and debugging these models harder than it should be. RouterInterp reportedly beats older token-statistics methods at predicting routing decisions by about 65 percent, giving researchers a more reliable lens on a part of these systems that has mostly been a black box.

Mixture-of-experts models have become a default way to scale AI without a matching default way to explain them; this work is a step toward closing that gap, not the last word on it.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →