AI/ ai · transformers · llm efficiency · research

Looped AI Models Get Smarter by Sharing Memory, Not Growing It

A new looped transformer design lets later passes reuse the first pass's cache instead of building their own, cutting memory use without hurting accuracy.

A new transformer design manages to use less memory and still produce better predictions than the model it's built from.

Looped Transformers save on parameters by running the same set of layers over a token multiple times instead of stacking new ones. The catch is that each pass through the loop writes its own key-value cache, so memory use climbs with every extra recursion even though the parameter count stays flat. Researchers built a variant called the Looped Prediction Transformer that lets only the first pass write a cache; every later pass reads that shared cache and keeps just a small window of its own. Tested at 150 million to 1 billion parameters with five recursions, the hybrid version cut validation perplexity on the FineWeb-Edu dataset by 1.12 to 1.82 compared to a standard transformer of the same size, while using 76-79% less context memory.

This matters because memory, not parameter count, is turning into the real limit on how much extra compute a model can spend thinking per token. Most existing tricks for shrinking a growing key-value cache trade away some accuracy to get there. This one does the opposite, and the researchers trace the gain to the shared cache acting as a kind of gradient shortcut back to the first pass.

The tests stop at 1 billion parameters, far short of the models running today's chatbots, so the open question is whether this frontier holds once someone tries it at a scale that actually matters.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →