AI/ ai · machine-learning · transformers · research

Transformers Secretly Run a Classic Regression Algorithm

A new study finds that when transformers solve regression tasks in context, they are mathematically retracing an old iterative algorithm, layer by layer.

New research shows transformer models can secretly execute a decades-old iterative algorithm when solving regression problems from context alone, with no gradient updates in sight.

The paper constructs a single-head transformer that approximates Gaussian kernel ridge regression purely through its forward pass. The authors show that a model built with O(log(1/epsilon)) transformer blocks and MLP width O(sqrt(N/epsilon)) can hit epsilon-level accuracy on prompts of length N. In this construction, softmax attention handles interactions between data points, producing a normalized Gaussian-kernel operator, while the MLP layers do the local arithmetic the update step needs. When the researchers trained GPT-2-style transformers on Gaussian-process regression tasks, the models' predictions and internal weights increasingly matched a classical kernel ridge regression solver as depth increased, with deeper layers tracking later iterations of that solver.

This adds a concrete, testable mechanism to the case that in-context learning isn't mysterious - it's the transformer running a recognizable optimization procedure, here called preconditioned Richardson iteration, encoded directly in its weights. That matters for anyone predicting when these models generalize versus fail, since a model executing a known algorithm inherits that algorithm's known failure modes.

It's a useful corrective to "learning without training" hype: underneath the hood, it can just be old-fashioned numerical methods, run fast.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →