AI/ ai · llm inference · kv cache · model efficiency

New technique shrinks AI memory footprint without deleting history

Researchers describe a compression method that shrinks the memory long-reasoning AI models consume by compacting old reasoning steps instead of deleting them.

A new paper proposes a way to shrink the memory bill of long AI reasoning chains without throwing anything away.

When a model works through a chain-of-thought answer, it builds a cache of key-value states for every token it generates, and that cache grows with every step. Most existing fixes handle this by evicting old tokens from the cache, which frees memory but risks deleting states the model might need to revisit later in a long reasoning chain. A paper posted October 5, 2026 on arXiv proposes iS-KV, which instead keeps recent tokens stored in full and folds older ones into a compact, bounded-size mathematical representation built incrementally. The method also re-syncs older compressed data whenever that representation updates, which the authors say is necessary because naively updating it for new tokens while leaving old coordinates alone causes the stored history to drift out of accuracy.

On DeepSeek-R1-Distill-Llama-8B, iS-KV held accuracy at 82.6% against 83.6% uncompressed, while shrinking the persistent cache more than fourfold. On Qwen3-8B, it reached 89.2% accuracy at a 5.64-fold compression ratio, and the authors report it beat token-eviction baselines under matched memory budgets.

Memory, not raw compute, is increasingly the ceiling on how long a model can "think" before cost or latency make it impractical, so shaving a small accuracy loss off a multi-fold memory cut is a real lever on what reasoning chains are affordable to run. That near-lossless result also suggests there's headroom to push the compression ratio further before it meaningfully hurts output quality.

Still, this is a preprint tested on two 8-billion-parameter models, not a shipped feature, so whether the gains hold at frontier scale is an open question.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →