AI/ federated-learning · optimizers · machine-learning · ai-research

Researchers Tune Federated Learning by Mixing Preconditioners

A new federated learning method mixes shared and local optimizer preconditioners, trimming bias and lifting language model accuracy by up to 19 points.

A new federated learning method fixes a subtle bug: local optimizers quietly drift apart even when every client starts from the same model.

Researchers propose FedMIX-P, which blends a shared global preconditioner with each client's own local one at every training step, not just when a round starts. The problem it targets: in federated training, clients such as phones or hospitals train adaptive optimizers like SOAP, Sophia, or Muon on their own data, and those optimizers build up local statistics about gradient geometry. That local bookkeeping can drift apart across clients even when everyone shares identical model weights, biasing the combined update after averaging. The authors prove the mixing reduces that mismatch and that the method still converges on nonconvex problems, without requiring local optimizers to agree with each other first.

This matters because federated learning only works if local training doesn't quietly pull the shared model in different directions. That's a plumbing problem, not a glamorous one, but plumbing is most of what breaks in federated systems once they leave the lab.

In experiments using SOAP, Sophia, and Muon across vision and language tasks, FedMIX-P beat plain local optimizers, with 60 million to 350 million parameter language models seeing accuracy gains up to 19.47 percentage points and lower validation loss - a sign that most federated learning progress right now is about taming the optimizer, not building a bigger model.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →