Researchers have built a way to trim AI memory caches by 90 percent while barely denting accuracy.
The technique, called CORE, targets the key-value cache that models use to remember earlier parts of a long conversation or document. Instead of just ranking which cached entries to keep and discarding the rest, CORE calculates the exact error that eviction introduces and redistributes that lost information into what's retained. It bundles two signals, query usefulness and a coverage measure borrowed from linear algebra, into one lightweight system that handles both decisions at inference time. Tested across three model backbones, it beat the strongest version of the RULER benchmark by up to 3.78 points at 90 percent compression, with further tests on LongBench and repeated-eviction scenarios showing it holds up over repeated use.
This is the unglamorous but real bottleneck behind long-context AI: every extra token a model remembers costs memory and bandwidth, and most existing compression methods just guess at what to throw away without accounting for what that guess costs in output quality. CORE's contribution is treating eviction as a measurable error to minimize rather than a ranking problem to approximate. If the approach generalizes, it's the kind of efficiency gain that lowers serving costs for anything leaning on long context, chatbots, coding agents, document analysis, without reaching for a bigger model.
The results so far come from benchmark suites, not live production traffic, so how well the savings translate to the sprawling, unpredictable contexts of real agent workloads is still an open question.