AI/ ai · mixture-of-experts · model-scaling · llm-research

A New Recipe Lets Looped AI Models Scale Past Two Loops

Researchers diagnosed why reusing the same AI model layers repeatedly used to stop helping after two loops, then fixed the two root causes.

A new training recipe lets AI language models loop through the same block of layers up to a dozen times instead of just twice, without the usual drop in quality.

The team, publishing as LOOM, studied looped mixture-of-experts models: setups that reuse the same layers multiple times and split work among specialized sub-networks called experts, rather than adding new parameters to get deeper. They found two reasons looping usually stalls after two passes: internal signal variance snowballs with each extra loop and destabilizes the model, and the router kept picking the same experts every time, so extra loops added computation without adding anything new. LOOM fixes both by capping the variance growth, re-feeding the original input at every loop, assigning a different router to each loop, and carrying earlier outputs forward through an added 'looping residual' connection. Tested on models from 100 million to 1.7 billion parameters, it scaled cleanly to 9-12 loops, with the 1.7B model, trained on 60 billion tokens, cutting its perplexity score from 9.62 to 7.77 and lifting average zero-shot accuracy from 42.4% to 47.7%.

This matters because looping is one of the few ways to make a model smarter without buying more GPUs or adding parameters, and until now that trick topped out fast. If the fix holds at the scale labs actually care about, compute-constrained teams get another lever besides 'bigger' for squeezing out gains, though LOOM's biggest test here, 1.7 billion parameters, is still tiny by frontier standards.

It's a reminder that scaling laws aren't one law: parameters, training data, and now loop count each hit their own wall that someone eventually has to engineer around.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →