AI/ ai · llms · kv-cache · research

Researchers Cut AI Chatbot Memory Use in Half Without Deleting Data

AttSVD compresses the memory AI models use to track long conversations by up to 50 percent, without discarding any tokens.

A new compression method lets AI models remember long conversations using up to half the memory, without throwing anything away.

Researchers behind a method called AttSVD target the key-value cache, the memory system transformer-based AI models use to track every earlier word as they generate new text. That cache grows with every token added to a conversation, and it becomes the single biggest memory cost once conversations get long. Most existing fixes solve this by deleting less-important tokens and hoping nothing valuable gets lost. AttSVD instead keeps every token, but compresses how each one is stored, using a per-prompt mathematical shortcut called SVD that is tuned to whatever the model's attention mechanism is actually reading. Tested on an agentic benchmark and the LongBench suite, it reportedly matched the performance of an uncompressed cache while using as little as half the memory.

Long-context AI, the kind that can hold an entire codebase or book in its working memory, is expensive partly because of this cache problem. Token-eviction methods are a one-way bet: once a token is thrown out, it is gone, and the model never finds out if it mattered later. A compression method that keeps everything but shrinks it is a genuinely different trade-off, and tuning that compression per prompt rather than applying one fixed rule is the detail worth watching.

Whether this beats eviction-based methods in real deployments, not just in papers, is the open question - shrinking the cache per token is useful, but the underlying architecture still grows that cache with every token added, no matter how cheaply each one is stored.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →