A new research paper tackles a boring but costly problem: AI chat assistants waste huge amounts of computation hunting for the right memory to bring up.
Researchers built a system called Madeleine that learns to associate conversational cues with the memories they should trigger, without calling a large language model during actual use. Today's long-term memory systems ask an LLM to reason about which memory is relevant every time you chat, burning anywhere from hundreds to over a thousand model calls per memory bank and thousands of tokens per query. Madeleine instead trains a smaller query encoder offline, using an LLM to simulate fictional life histories and the memory cues that pop up within them. Once trained, it runs with zero LLM calls at query time and plugs into existing vector-based memory systems by swapping out just that encoder.
On the LoCoMo-Plus benchmark, Madeleine paired with an existing memory system called HyperMem scored 66.6, the best result reported under the test's official protocol. Used by itself, it nearly matched HyperMem's own released score, 52.4 versus 52.9, while using about one twenty-first of the context tokens and no LLM calls at all, and it boosted a separate system called T-Mem by 26.2 points. That efficiency gain matters for anyone running a chatbot with memory at scale, since those LLM calls were previously a direct hit to latency and cost.
The trick is training the system to guess what you will need before you ask for it, which works neatly right up until real human memory turns out messier than a simulated one.