AI agents that browse the web and complete multi-step tasks have a forgetting problem, and a new training method called AMBER is built to fix it by refusing to let them delete anything.
Long tasks generate too much history to keep in an AI model's context window, so researchers have tried shrinking that history down. One popular method lets an agent learn to overwrite its own memory, keeping only a fixed-size summary as it works. Researchers from the paper behind AMBER found a flaw in that approach: agents trained this way tend to delete the exact things they need later, including corrective feedback the environment gave them after a mistake. AMBER instead uses an append-only memory, so an agent writes notes as it goes but can never erase them, and it is trained end-to-end with reinforcement learning rather than hand-curated examples. On the WebArena Lite benchmark, AMBER beat overwrite-based memory by 4.09 percentage points on average success and raised the share of tasks solved consistently across five repeated runs by 4.8 points.
This matters because "just summarize the history" has become the default fix for AI agents running out of context, and this paper is evidence that default is quietly sabotaging reliability. An agent that forgets why its last attempt failed is an agent doomed to repeat it, which is exactly the kind of silent failure mode that is easy to miss in a demo but costly in production.
The tradeoff is one every memory system eventually hits: append-only avoids deletion mistakes but has to manage growing token costs instead, so the real test is whether that balance holds up on tasks longer than anything in WebArena Lite.