AI/ reinforcement-learning · test-time-training · memory · decision-transformer

New Model Stretches AI Memory 20x Beyond Its Training Window

Decision Titan combines a Decision Transformer with test-time training to show offline RL agents can recall events far outside their context window.

A new reinforcement learning model can remember events that happened 20 times further back than its own context window should allow.

Researchers built Decision Titan by bolting Test-Time Training (TTT) layers onto a Decision Transformer, an architecture used for offline RL. TTT stores episodic memory inside a neural network's parameters, updating them with gradient descent during both training and testing. The team tested the model in X-Maze, an extended version of the T-Maze benchmark built specifically to probe long-range memory. They found Decision Titan could recall information from 20 times beyond its context window and still perform reasonably well on sequences 1.7 times longer than anything it trained on.

That's a meaningful result because memory has been the stubborn problem in sequential decision-making. RNNs lose track of old information as gradients vanish, and Transformers get expensive fast because attention costs scale quadratically with sequence length. TTT offers a third path: stash what matters in the weights themselves rather than a fixed-size hidden state or an ever-growing attention window.

The catch, and the researchers say so themselves, is that how well this works depends heavily on details most papers wave past: the time embeddings used and how the relevant information gets encoded in the first place. Get those wrong and the long-term memory advantage evaporates.

X-Maze is a toy problem, not a warehouse robot or a game-playing agent juggling a day's worth of state. Whether this scales past maze corridors is the next experiment, not this one.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →