AI/ ai · llm · kv-cache · inference

New Technique Slashes AI Chat Memory Use by Up to 90 Percent

A new compression method keeps AI models fast and accurate while storing only a fraction of the data needed to remember long conversations.

A new compression scheme lets AI models hang onto long conversations using a fraction of the memory they normally need.

Researchers behind a paper called ResidualKV built a system that splits an AI model's key-value cache - the running memory it uses to track everything said so far - into two parts: a small set of important reference tokens and compressed "residual" codes for everything else. Instead of permanently deleting older tokens to save space, as many existing methods do, ResidualKV keeps a lightweight record of them and reconstructs the full detail only when the model's attention mechanism actually needs it. Tested across Llama, Qwen, LLaVA-OV, and Qwen3-VL models, the method held performance close to systems using the full, uncompressed cache while using just 13-16% of the storage and 30% of the attention computation on the LongBench benchmark. On multimodal tasks the savings were steeper still: 8-10% storage and 10% computation, plus decoding speedups of up to 1.5x with cache quantization and 3.4x without it. Code is posted on GitHub.

Long-context inference is one of the biggest cost centers in running AI models - every extra token of conversation history eats memory and slows attention calculations, which is why many chat products quietly cap or summarize history behind the scenes. A method that keeps near-full accuracy while cutting storage by 85% or more could let products hold onto much longer histories without a matching jump in GPU memory and latency costs.

These are benchmark numbers, not results from a production chatbot under real traffic, so how much of that efficiency survives outside LongBench and similar test suites is still an open question.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →