AI/ ai · linear-attention · transformers · machine-learning

Researchers Find a Cheaper Way to Match Transformer Attention

SwiLA is a new attention layer that keeps linear attention's constant memory footprint while closing much of its quality gap with standard attention.

A new attention layer keeps its memory footprint flat as context grows, and it is starting to match, and in some cases beat, the standard attention transformers rely on.

Researchers have introduced Switching Linear Attention (SwiLA), a sequence layer built to close the long-standing gap between linear attention and standard softmax attention. Softmax attention is expressive but needs a key-value cache that grows with every token, which gets expensive at long context lengths. Linear attention avoids that by keeping a fixed-size recurrent state, but it has historically sacrificed modeling quality to do it. SwiLA's fix lets each output dimension dynamically pick among several linear attention components at test time, derived from a framework that treats the state update as online expectation-maximization over a mixture of linear regressions.

That matters because the core trade-off in sequence modeling, expressivity versus memory, has mostly been an either-or choice. Efficient architectures like linear attention and state-space models have chased lower memory costs for years without fully closing the quality gap to standard transformers. SwiLA's results, which include beating softmax attention on some associative recall and in-context learning benchmarks, suggest that gap is more about architecture design than an inherent ceiling.

Benchmarks are not production systems. The real test for any linear-attention variant is whether it holds up at the scale and hardware constraints of the models people actually ship, something this paper does not claim to answer yet.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →