Long-running AI agents are supposed to remember everything that matters, but a new study finds they systematically forget the rare stuff.
Researchers audited how language agents actually use external memory under repeated retrieval and found a consistent pattern: memory clusters around a small core of common situations while rare ones pile up in a long tail where prediction errors build. Random-walk agents produced a milder, log-normal-like version of this skew; agents driven by semantic LLM policies produced a sharper truncated-power-law pattern. Based on that finding, the team built Core-Tail World Model (CTWM), a rank-based controller that spends prompt budget using a single tunable exponent while keeping a summarized version of the tail instead of dropping it. Across three testbeds, CTWM kept full state and transition coverage, cut prompt tokens by 5.9% and tail prediction error by 13.6% on Synthetic Graph World, saved tokens consistently on ALFWorld, and trimmed tokens by 24.48% on LongMemEval without losing accuracy.
This matters because most agent memory systems get graded on task success or raw token cost, metrics that can hide exactly this kind of quiet degradation on uncommon cases. For teams running coding assistants or support agents on tight context budgets, that blind spot is where expensive mistakes hide - the edge case nobody tested for. CTWM's single exponent also gives engineers an actual dial to turn, rather than just shrinking the context window and hoping.
Worth noting: two of the three testbeds are synthetic or scripted environments, and LongMemEval is still a benchmark, not a production workload, so the real test is whether this pattern and fix hold up once agents meet genuinely messy, open-ended memory.