AI/ scaling laws · attention mechanism · llm research · arxiv

New Theory Links AI Scaling Laws To Attention Nonlinearity

A new arXiv paper argues nonlinear attention, not data structure, explains why deeper language models keep improving.

A new theoretical paper argues that the reason bigger language models keep getting better might come down to one mechanism: nonlinearity in attention.

Researchers have long tried to explain why model performance improves in a predictable pattern as you add more layers. One existing theory pins this on linear-attention models, which can't selectively focus on relevant tokens. Those models learn according to the overall statistical structure of the data - strong patterns get learned first, weaker ones later, however deep the model gets. The new paper shows that once attention is nonlinear, as it is in real large language models, that constraint disappears. Nonlinear attention lets a model focus on relevant tokens at every layer, so strong and weak patterns get learned in parallel rather than in sequence - and across every data spectrum the researchers tested, this produced a consistent inverse relationship between loss and depth.

That reframes depth scaling as a property of how attention works, not just how much structure happens to be sitting in the training data. It also means the smooth loss curves researchers lean on to plan training runs may owe more to attention's nonlinearity than to dataset shape. The paper ties this to the central limit theorem: shared errors across layers set a floor on how low loss can go, while small layer-to-layer differences keep adding up to produce further gains.

It's a theory paper, and "all tested data spectra" is doing a lot of load-bearing work - plenty of tidy scaling explanations have come apart once checked against real frontier-scale training runs.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →