Researchers have found a much cheaper way to train looped language models, AI systems that reuse the same layers over and over to simulate extra depth.
The team's recipe cuts the token budget needed to train a looped model from scratch from the 7.7 trillion tokens used in the prior Ouro project to just 310 billion, a roughly 25-fold reduction. They get there with standard pretraining followed by targeted mid-training, a learning-rate warmup, and tighter exit-gate regularization, the mechanism that decides when to stop looping. In controlled tests, their 1.4-billion-parameter looped model beat a same-size, same-data dense model on all 12 benchmarks tried, including gains of 14 points on GSM8K, 10 points on MATH, and 22 points on DROP. The researchers also built a lightweight way to convert existing dense models into looped ones, using a single learned mixing parameter and a smoothed exit loss, and showed it improved Qwen3-1.7B on every benchmark tested.
Looping has been pitched as a way to get more reasoning power out of fewer parameters, but the training cost made that trade-off mostly theoretical outside a frontier lab's budget. This recipe makes recurrence testable, and potentially shippable, for teams working with far smaller compute, and it isolates the gains from recurrence itself rather than differences in data or training schedule. The retrofit method matters just as much in practice: it means labs do not need to start from scratch to see whether looping helps their existing checkpoints.
At matched inference compute, the 1.4B looped model approached the performance of a 3.9B dense model while using little more than a third of the parameters, a reminder that scaling up is not the only lever left to pull.