AI/ ai · reinforcement-learning · transformers · interpretability

Study Finds Transformer Attention Implicitly Runs Control Algorithm

A new preprint argues transformer attention can run policy mirror descent as a controller, beating two RL baselines despite consuming far more data per round.

A new paper argues that attention layers in transformers can run a known reinforcement learning algorithm step by step, not just approximate it in one shot.

The researchers built a controller made of actor, critic, and environment components, all wired through the standard causal-softmax attention that powers today's transformers, then trained real pre-LN transformer models that reproduced the target computation on their own. In a preregistered five-run test at their S=4 setting, a learned actor paired with an exact critic reached a median loss of 1.052 times the ideal reference policy, and kept that performance across four changes to the test conditions without retraining. Using the model's own learned critic instead of the exact one gave a similar 1.050 times the reference loss, though that figure falls outside the preregistered comparison. At a harder S=8 setting, the learned critic pushed median loss to 0.0225, well below two established baseline methods, Liang-Lai and Algorithm Distillation, which scored 20 to 24 times worse.

This adds to a growing pile of evidence that standard transformer components secretly implement known algorithms rather than some inscrutable proxy for one, a pattern interpretability researchers have chased for years in the language-modeling context. Extending that idea to sequential control, instead of single-step prediction, suggests attention could double as a lightweight reinforcement-learning controller without bolting on extra machinery.

The catch: the comparison isn't apples to apples. The method consumes 144 training transitions per round versus 20 for the baselines it beats, and the strict same-budget test the authors flag as still open hasn't been run.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →