A new quantization method compresses the memory that powers fast, linear-attention AI models down to 8 bits without losing accuracy.
The method, called LeapQuant, targets a bottleneck in linear-attention architectures like Gated DeltaNet and Kimi Delta Attention, which hybrid LLMs increasingly use to handle long context cheaply by compressing conversation history into a fixed-size recurrent state. Constantly reading and updating that state at full precision is expensive, and naive quantization degrades quality because errors pile up and a few outlier values dominate. LeapQuant quantizes the state only once per window of tokens instead of after every update, and keeps the state's largest outliers in high precision so rounding errors don't compound. Tested training-free across the Qwen, Kimi, and GLM model families, it matched FP32 baseline accuracy while cutting memory and compute costs.
The kernel-level math sped up 2.05x to 3.70x, but once you account for everything else involved in running a model, the real-world gain on Nvidia's B200, RTX PRO 6000, and RTX 5090 GPUs settled at 1.47x end-to-end. That gap between headline kernel numbers and actual inference speed is exactly the kind of detail that gets flattened in press coverage, and it's a useful reminder that a faster core operation doesn't automatically mean a faster model.
Efficient long-context inference is turning into as big a battleground as raw model quality, and shaving memory costs without retraining is the unglamorous work that actually ships.