AI/ federated-learning · state-space-models · mamba2 · ai-research

Study Shows How Mamba2 Models Behave in Federated Learning

New math explains why Mamba2 state space models wobble under federated training, and tests nine algorithms across six text domains to prove it.

A new paper works out, mathematically, how Mamba2-style models behave when you train them across scattered devices instead of one big machine.

Researchers derived math describing how selective state space models, an alternative to transformers that process sequences step by step with a running internal state, behave during federated learning, where multiple devices train a shared model without pooling their data. They built formulas predicting how stable training stays under two standard federated methods, FedAvg and FedProx, based on the model's internal stability and how much each device's data differs from the rest. They checked the single-layer version of this math against a controlled teacher-student setup and found it held up. Then they ran nine federated learning algorithms on Mamba2 language models across six different text domains, using the theory to explain the gaps between algorithms.

Most federated learning research assumes the model architecture barely matters, that whatever recipe works for transformers will transfer to anything else. This paper argues the opposite for state space models: their recurrent, input-dependent dynamics change how federated training behaves, not just how fast.

It's a narrow, technical result, but it matters if state space models keep eating into transformers' turf. Teams building on-device or privacy-preserving models will eventually need SSM-specific playbooks, not borrowed transformer assumptions.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →