A new AI memory system remembers conversations in more accurate-sounding detail, but its own creators say that doesn't always mean better answers.
A paper posted to arXiv (2609.30797, September 28, 2026) introduces HasMem, a memory architecture for AI chat agents that compresses long conversation histories without retraining the underlying language model. Instead of storing full transcripts, HasMem starts from fixed reference summaries, then uses a controller to resize them and a component called the Writer to re-encode entries as they shrink or grow. On a 535-question recall test built from the Multi-Session Chat dataset, the main version scored a lexical overlap measure (F1) of 95.3, a 4.4-point gain over the baseline, while using just 93.6% of the space the reference summaries took up. Against a simpler rule-based compression method, six HasMem configurations also beat exact-match accuracy by 8.0 to 23.6 percentage points.
But the researchers' own numbers include a caveat worth flagging: across both of the paper's full-scale evaluations, including a second 500-question benchmark called LongMemEval-S, gains in lexical overlap came paired with lower exact-match accuracy. In plain terms, the system got better at recalling text that sounds like the original conversation, without a guaranteed matching improvement in getting the actual answer right. For teams building agents meant to hold onto weeks of chat history, that's the gap between sounding consistent and being correct.
It's an early-stage research result, not a shipping product, and until exact-match accuracy catches up to the overlap scores, treat "sounds right" and "is right" as two different claims.