Phones that train AI models together without sharing personal data have had a specialization problem, and a new paper offers a fix.
On-device AI models that use a Mixture-of-Experts design, the architecture behind a lot of today's efficient large language models, run into trouble under federated learning: when multiple phones train the same model without pooling their data, their internal experts and routing preferences drift apart instead of converging. Researchers describe a fix called FedAlign-MoE in a new paper. The method aligns how different phones route requests to experts and checks whether same-numbered experts across devices are actually doing the same job before merging them. In testing on non-IID, meaning unevenly distributed, data, the researchers report faster convergence and higher accuracy than existing approaches, with lighter computation and less network traffic.
This matters because Mixture-of-Experts is the main way companies are scaling up on-device AI without blowing through a phone's compute budget. Federated learning was already the privacy-friendly answer to training on personal data without centralizing it; bolting MoE onto that approach introduces exactly the coordination problem this paper tries to solve.
For now, it is one paper's benchmark numbers, not a product. The real test is whether it holds up on actual phones instead of simulated splits of a dataset.