Most of the hype around AI agents with "better memory" turns out to be about how much history they can see, not how smart their retrieval is.
Researchers built a streaming-recall benchmark that separately tests retention and selection, running 300 seeded episodes per condition. When they held access to history fixed, giving an agent query-aware selection, meaning it chooses what to surface based on the question asked, improved recall of required facts by 15.5 percentage points. But in a mixed comparison that let both retention and access vary at once, the reported advantage ballooned to 68.7 points, and 53.2 of those points came purely from extra access to history rather than smarter selection. Under a fixed memory budget, query-aware, dense, and even oracle selection methods all hit the same retention ceiling, and every one of 319 observed failures traced back to facts being evicted from memory, not to bad ranking. Recall fell to zero once a fact receded far enough into the past. The same pattern held on the SQuAD benchmark, where dense retrieval beat simple keyword matching but retention stayed the real bottleneck.
That reframes a chunk of existing agent-memory research. Papers comparing methods with different history budgets are mostly measuring whose memory is bigger, not whose retrieval algorithm is sharper. For anyone building long-running agents with limited context windows, this suggests bounded retention policy, what gets thrown away and when, deserves as much engineering attention as the retrieval model sitting on top of it.
In other words, "our agent remembers more" is often a context-window claim wearing an algorithm's clothes.