AI agents that lean on stored memories to guide new tasks may finally get a smarter way to judge which memories are worth pulling up.
UpliftMem is a new training method that scores memory sets by their uplift - the boost in task success versus the same agent working with no memory at all. Measuring uplift directly is expensive, since testing every candidate memory set means running extra full task attempts, so the researchers use a statistical method called expected value of sample information to decide which comparisons are worth that cost during training. A single scoring model learns from those targeted comparisons and then picks memory sets at run time without needing any extra test rollouts. Across three benchmarks - ALFWorld household tasks, WebShop online shopping, and BigCodeBench coding problems - it beat other memory-retrieval baselines on success rate.
Most agent memory systems retrieve whatever past example looks most similar to the current task, which isn't the same as useful and can actively steer an agent wrong. This work reframes the real bottleneck as training efficiency - figuring out which comparisons to run - rather than just building a bigger memory store, a distinction that matters once agent rollouts get expensive at scale.
If that efficiency holds beyond these three benchmarks, it points toward agent memory that gets pickier, not just bigger, as teams start relying on it for real work.