AI/ mixture-of-experts · distributed-training · gpu-memory · ai-research

New Method Doubles Training Speed for Mixture-of-Experts Models

A ring-based routing scheme called RelayMoE cuts memory use in distributed MoE training, letting labs train longer sequences without buying more GPUs.

A new training method called RelayMoE cuts the memory overhead of training huge AI models built from many specialized sub-networks.

Researchers describe RelayMoE, which changes how Mixture-of-Experts (MoE) models (AI systems split into many specialized expert sub-networks, where only some handle each input) move data between GPUs during training. The standard approach bundles all routed data into one giant buffer before sending it across the network, which eats memory. RelayMoE instead passes expert weights or tokens around a ring of GPUs, computing as data arrives and discarding intermediate results once used. Tested on production-scale models from 30 billion to 57 billion parameters, it delivered a 2x speedup in isolated layer tests and up to 2.02x faster full-model training within the same GPU memory budget, while letting researchers train sequences up to 2.85 times longer.

Memory, not raw compute, is usually the ceiling on how big a model a given GPU cluster can train, and MoE architectures (the structure behind many current large language models) make that worse because routing data to the right experts multiplies memory use. A technique that frees up memory without new hardware effectively stretches an existing GPU budget further, which matters more as labs chase longer context windows and bigger expert counts rather than just bigger parameter totals.

The gains are measured against Megatron-LM, the industry's reference training framework, not against whatever custom pipeline a given lab already runs in production, so real-world improvements will likely vary.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →