A new compression scheme lets AI models hang onto long conversations using a fraction of the memory they normally need.
Researchers behind a paper called ResidualKV built a system that splits an AI model's key-value cache - the running memory it uses to track everything said so far - into two parts: a small set of important reference tokens and compressed "residual" codes for everything else. Instead of permanently deleting older tokens to save space, as many existing methods do, ResidualKV keeps a lightweight record of them and reconstructs the full detail only when the model's attention mechanism actually needs it. Tested across Llama, Qwen, LLaVA-OV, and Qwen3-VL models, the method held performance close to systems using the full, uncompressed cache while using just 13-16% of the storage and 30% of the attention computation on the LongBench benchmark. On multimodal tasks the savings were steeper still: 8-10% storage and 10% computation, plus decoding speedups of up to 1.5x with cache quantization and 3.4x without it. Code is posted on GitHub.
Long-context inference is one of the biggest cost centers in running AI models - every extra token of conversation history eats memory and slows attention calculations, which is why many chat products quietly cap or summarize history behind the scenes. A method that keeps near-full accuracy while cutting storage by 85% or more could let products hold onto much longer histories without a matching jump in GPU memory and latency costs.
These are benchmark numbers, not results from a production chatbot under real traffic, so how much of that efficiency survives outside LongBench and similar test suites is still an open question.