A new benchmark suggests AI companion apps barely need the long-term memory they are built around.
Researchers released RealCompanion, a benchmark built from ten real relationships between people and an AI companion - 27,218 messages spanning up to 120 days. The dataset includes the raw conversations plus four derived files: a profile, a persona, a chat ground truth, and a question set, each one citing the exact messages behind it. Every answer label comes with a reasoning trace that was checked stage by stage against the source conversation, an unusual amount of rigor for a field that typically fabricates its test subjects.
The results undercut the pitch behind "memory" as a product feature. A simple recency window - no memory retrieval at all - found the correct message for 95.9% of questions, and at the rate memory-dependent questions occur naturally, 96% of the benefit credited to pulling up old messages actually came from messages that did not require memory. None of the memory-need detectors the researchers tested could reliably spot, in real conversations, when older context actually mattered; tests built from authored questions over the same histories leaked the cue in ways real usage does not.
The study also found three different agent systems reconstructed a user's persona to the same accuracy despite a 31-fold spread in cost - its own quiet argument against paying more for "deeper" memory.