AI assistants that remember your history still get confused about how those memories relate to each other.
Researchers introduced SubtleMemory, a benchmark testing whether long-running AI agents can tell when stored memories reinforce, diverge from, or outright contradict each other. The benchmark builds relation-controlled memory variants - complementary, nuanced, and contradictory versions of the same underlying fact - and buries them inside realistic, long-running user-agent conversation histories. It totals 1,522 evaluation instances across 10 long histories, built from 1,090 relation-controlled memory-variant sets, covering both user-specific and general queries. The team tested six standalone memory systems plus five Claw-style agents, two with native memory modules and three running plugin memory add-ons, and found that nearly all of them struggled to tell the memory relationships apart correctly.
Most memory benchmarks check whether an agent can recall a fact, not whether it still understands how that fact fits with everything else it has been told. That gap matters as assistants get pitched as long-term collaborators: a scheduling assistant that misses the fact that your new no-meetings-before-10am note contradicts what you said last month isn't just forgetful, it is actively wrong. The researchers' diagnostic protocols go further, separating failures by stage - whether a system drops the relevant detail when storing memories, fails to retrieve it, or retrieves it but reasons about it incorrectly - which is useful for anyone trying to figure out where their own memory stack actually breaks.
Persistent memory has been pitched as the next big unlock for AI assistants for a couple of years now. This benchmark is a reminder that storing more memories isn't the hard part, keeping them straight is.