AI/ ai · diffusion-models · mixture-of-experts · model-compression

Researchers Compress MoE Diffusion Models Without Losing Accuracy

A new technique cuts storage and compute costs for mixture-of-experts diffusion language models by up to 7x while keeping accuracy above 96%.

A new compression method lets mixture-of-experts diffusion language models run up to 7 times faster without gutting their accuracy.

Researchers built ITC-MoE, a framework that compresses the expert sub-networks inside mixture-of-experts diffusion language models - AI systems that generate text by filling in tokens in parallel rather than one at a time, using a crowd of specialized expert networks and routing each token to only a few of them. The team found compression needs vary a lot: some parts of the expert weight matrices matter more than others, and some tokens are far more sensitive to compression than "cold" tokens that already stick to a narrow set of experts. ITC-MoE uses gradient and activation signals to decide how aggressively to shrink each weight matrix, then gives sensitive tokens a small low-rank patch to recover lost precision while cold tokens get routed through a smaller expert shortlist. On a 30-billion-parameter model (SDAR-30B-A3B-Chat-b32), the method cut parameters by 30% while holding MultiArith accuracy at 96.33% and delivering up to a 7.22x end-to-end speedup.

Diffusion language models are the leading alternative to the autoregressive, one-token-at-a-time transformers behind most chatbots, promising faster generation through parallelism - but mixture-of-experts versions carry enormous parameter counts that make them expensive to store and run. A compression method that targets token-level and weight-level redundancy, instead of pruning everything equally, could make these models practical to deploy outside research labs and narrow the cost gap with smaller dense models.

Most prior MoE compression work focused on standard transformers, not diffusion models. Whether ITC-MoE's token-aware tricks hold up beyond the single 30B model tested here is the question the paper leaves open.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →