AI/ ai · llm memory · privacy · research

New Research Splits AI Memory Control Into Two Steps

A new study shows splitting what an AI agent recalls from how it says it cuts leakage and errors, but recall quality remains unsolved.

AI agents that remember you are more useful and more dangerous. A new arXiv paper tackles that tension by splitting memory control into two separate jobs: deciding what a system is allowed to recall, and deciding how it phrases what it recalls.

Researchers tested two inference-time methods, called factor-compiled and permission-semantic admission, across four model backbones and an external benchmark of four 300-sample tasks. Compared to dumping stored memories verbatim into the prompt, both methods cut judge-assessed failure rates by 6.7 and 8.8 percentage points, with cross-domain leakage (the agent spilling details from one context into an unrelated one) dropping by as much as 29.5 percentage points. Tightening what gets admitted cut cross-domain failures by another 17.5 percentage points. Changing only how approved memories were phrased, by contrast, made no statistically significant difference.

That is the real finding here: the fix for leaky, sycophantic AI memory is not better wording, it is better gatekeeping. Every assistant now adding "remembers you" features is making exactly this admission call, mostly without naming it as one.

The catch is that tighter filters did not come free. Both methods increased personalization failures, and the more selective one missed its own preregistered bar for preserving useful memory. Translation: nobody has cracked how to keep an AI's memory both safe and genuinely personal yet.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →