AI chatbots that claim to remember you often remember the wrong things, because they forget why a fact was true in the first place.
Researchers describe a framework called Gated Memory that adds two checkpoints before a conversational AI writes anything to its long-term memory store. The first checkpoint, an admission gate, checks a candidate fact against the full conversation before any extraction happens. The second stage, enrichment, tags each fact with its scope, source (stated vs inferred), and resolved time and place, and blocks any fact that invents an entity not present in the conversation. On the LoCoMo-10 benchmark, which includes emotionally dense conversations, the system scored 2.6% higher in relative LLM-judge accuracy than a baseline that shared the same retrieval and generation setup.
Most memory work in this space has focused on what happens after a fact is stored: retrieval, deduplication, pruning. This paper argues the real damage happens earlier, at the moment a sentence gets flattened into a subject-relation-object triple and loses the context that would have told the system this is temporary or this is who I am. That is a sharper diagnosis than most memory papers offer, even if the benchmark gain is modest.
A 2.6% bump on one benchmark is not proof this scales to the messy, contradictory things real users tell their chatbots over months. But it is a rare case of an AI memory paper fixing a cause instead of patching a symptom.