A new reinforcement learning agent has learned a skill most memory systems lack: knowing when to forget on purpose.
Researchers behind ALER (Adaptive Learnable Experience Rewriting) built an agent that pairs a standard LSTM with a slot-based memory module. Instead of just storing and retrieving facts, ALER can overwrite a single memory slot using a Gumbel-Softmax write mechanism, then blend that update with its recurrent state through a learned gate before acting. The team tested it on new maze environments called Rune-Mazes, where symbols can invert, cancel, reset, or repeat a hidden cue the agent has to track. Against seven baseline methods, ALER hit at least 0.82 success across sixteen Endless T-Maze setups, at least 0.99 on five Rune T-Maze tasks, and beat PPO-LSTM in eight of ten pixel-based memory configurations.
Most memory benchmarks for RL only test retention: can the agent hold onto information until it's needed. That's a narrower problem than what real environments demand, where a later observation can flip the meaning of something learned earlier. ALER's contribution is less the architecture itself than the reframing: treating memory updates as rewriting and fusion operations, not just storage, and building tasks that actually require more memory states to solve.
The mazes are still toy problems, and a slot-memory LSTM is a modest architecture next to the attention-heavy memory tricks powering today's large models. Whether this rewriting idea scales past gridworlds is the open question the paper doesn't answer.