Turns out all that effort building fancy graph databases for AI agent memory might not be worth it: in a new benchmark, plain keyword search beats most graph-based retrieval methods anyway.
Researchers built a controlled evaluation framework to compare how AI agents retrieve memories from past conversations, testing several retrieval methods against a shared set of 5W-style conversational histories. Localized graph configurations traversed a common base graph directly, while AdaptiveGraph added chronological edges and a diffusion-based ranking technique called Personalized PageRank on top of that same graph. The team also ran BM25, a classic lexical search algorithm, over the same extracted notes, plus OpenClaw, a system that searches raw conversation text instead of a processed graph. On the LongMemEval-S benchmark, AdaptiveGraph led the graph-based methods with a 0.844 MRR score, a measure of how high the right memory ranks in the results, but BM25 (0.867) and OpenClaw (0.880) both beat it; on a second benchmark, ATANT Core, the localized graph configurations actually outperformed AdaptiveGraph's diffusion approach, while BM25 again took the lead in the hardest stress-round tests.
The deeper finding isn't that graphs are inherently worse, it's that retrieval architecture can't be judged apart from how memories get extracted and tagged in the first place. The researchers traced many of the graph method's missed retrievals back to extraction steps that simply failed to tag relevant details in the notes. That's a more useful lesson than simply labeling graphs good or bad: a sophisticated index built on sloppy inputs will lose to a dumb index built on clean ones.
For teams racing to bolt graph databases onto every AI memory system, that's an inconvenient reminder: the decades-old keyword index still shows up and wins.