AI/ ai · llm · research · efficiency

New Technique Cuts AI Model Memory Use by 32x

A new arXiv preprint shows a learned selector can shrink LLM KV caches by up to 32x without hurting accuracy, letting models run cheaper at long context.

A new method can shrink the memory used by long-context AI models by up to 32 times without hurting accuracy.

The approach, called KV-Kaizen and detailed in an arXiv preprint (arXiv:2609.37988, "KV-Kaizen: Learning Context-Adaptive Cache Compression Choices"), targets the key-value cache that large language models use to track context. That cache can grow bigger than the model's own weights as context length increases, and since decoding is memory-bound, a bloated cache slows down every token the model generates. Instead of the common fix - evicting less-relevant tokens, which risks losing something the model needs later - KV-Kaizen trains a selector that decides, layer by layer, how to compress the cache before generation begins, using three levers: sharing cache across layers, storing it at lower precision, or truncating its low-rank representation. On a 14-billion-parameter model, that got the decode-time cache down 32x with no accuracy loss, and on models 7 billion parameters and up, a 4x cache reduction cost nothing in accuracy.

That matters because cache size, not raw compute, is often what caps how fast and how cheaply a model can process long documents, chat histories, or codebases. The paper also reports that a compressed model outperforms a smaller uncompressed model using the same cache budget, an argument for training big models first and compressing afterward rather than shrinking them from the start.

It is one preprint, not a shipped product, but if the gains hold outside the paper's own benchmarks, this is a cheaper route to long-context performance than simply buying more memory.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →