Compressing an AI model's memory to handle long documents turns out to create predictable holes in what it can remember.
A new paper on arXiv examines chunked KV-cache compression, a technique that shrinks the memory large language models use during long-context inference by squeezing blocks of tokens into fewer cache entries at a fixed interval. The researchers found a token's "phase" - where it sits relative to those compression boundaries - determines how easy it is to retrieve later. In large open-weight models using this compression, retrieval accuracy swung by as much as 40 percentage points depending on phase alone. To confirm the effect wasn't a fluke of any one architecture, the team pretrained their own transformers from scratch across several compression designs, reproduced the same pattern, and used causal interventions to show that different attention components specialize in retrieving information from different phases.
That specialization is the real finding here: it is not noise, it is a structural habit the model learns during training. A single average benchmark score can hide this entirely, since good performance at some phases cancels out failure at others. For anyone deploying long-context models in production, that means a model could pass a standard retrieval benchmark while still reliably dropping information that happens to land in the wrong spot in a document.
It is a reminder that "long-context support" is a marketing phrase doing a lot of work. The actual guarantee underneath it depends on exactly how a model manages its cache, and this paper suggests the honest answer is: unevenly.