AI/ ai · ai-agents · benchmarks · memory

Benchmark Shows AI Agents Forget How To Treat Repeat Users

A new benchmark, DyadMem, finds top AI models struggle to recall and safely update how they should work with individual users over time.

A new benchmark shows AI agents are bad at remembering how to treat the people they talk to, not just what those people said.

Researchers introduced DyadMem, a benchmark built around a new concept called User-conditioned Relational Agent Memory, or URAM: the idea that an agent needs to remember not just facts about a user, but how it is supposed to interact with that specific user over time. The dataset spans 3,065 episodes, 50,961 sessions, and 61,210 question-and-answer instances, each annotated for when an agent should capture new information, update it, and recall it later. The team tested 16 open-weight models and 4 proprietary models under two conditions: Gold-Memory, where the correct memory is handed to the model, and Full-Pipeline, where the model has to build and retrieve memory on its own. Scores held up fine under Gold-Memory and dropped sharply under Full-Pipeline.

That gap is the real finding. It means today's long-term memory features in chatbots and assistants are good at answering questions when given perfect notes, but much worse at actually taking the notes, updating them when something changes, and knowing what is safe to forget. The study also found specific unsafe-deletion problems, meaning models sometimes erase information they should have kept.

If an assistant has ever cheerfully forgotten a correction you made sessions ago, this benchmark has the numbers explaining why.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →