AI chatbots that cache old conversations to save memory might be leaking more than anyone realized.
A new arXiv paper tests a basic assumption behind long-running AI agents: that once an old observation is dropped from a chatbot's cache to save space, the information is actually gone. Researchers removed a single earlier observation from an agent's history while keeping everything else the same, then checked whether later answers still depended on it. They overwhelmingly did. The researchers call this "semantic materialization" - a later cached event can act as a standalone readout of a fact whose original source was deleted. Worse, it can be done deliberately: a carefully worded prompt that never states a value outright raised the odds of recovering that deleted value from 6% to 51% on the Qwen3-8B model.
This matters because cache-trimming, or "eviction," is the standard way AI systems manage memory over long conversations. The assumption has always been that trimming old data is safe as long as answer quality holds up. This research says that assumption is wrong - a system can look fine on accuracy while still quietly retaining information it was supposed to forget.
That has real teeth for privacy. If an agent is told to "forget" a piece of sensitive data, deleting the original message may not be enough - traces can persist in later cached computations. It is the AI equivalent of wiping a hard drive but leaving the file's fingerprints in some other file's metadata. Anyone building compliance claims around chat memory deletion should read this before promising customers anything gets truly erased.