A new compression technique squeezes AI model memory down to a sixth of its original size without tanking accuracy.
Researchers posted a paper to arXiv on September 30, 2026 describing STEPQuant, a post-training quantization method for the recurrent states used in linear-attention models built on the Delta-rule architecture. Unlike standard transformers, these models swap growing key-value caches for a fixed-size memory that still eats server RAM at scale. The paper's authors found errors don't hit evenly: mistakes in long-lived memory rows compound over many decoding steps, and different key rows sway outputs more than others. STEPQuant allocates precision accordingly, spending more bits on the values that matter and fewer on the ones that don't.
Tested on Qwen3.8-27B and Kimi-Linear-48B-A3B-Instruct, the researchers report their 6-bit configuration matches FP32 accuracy while compressing recurrent states more than 5x, and cuts total serving memory by up to 68.7% once built into the SGLang inference engine. That's a real dent in the hardware bill for anyone running linear-attention models at scale, where memory capacity, not raw compute, often caps how many users a server can serve at once.
Linear attention has been pitched for years as the transformer's leaner cousin; this paper is a reminder that "efficient" claims still need the unglamorous work of figuring out which bits you can actually afford to throw away.