AI/ ai · llm-memory · reinforcement-learning · research

New RL Technique Helps AI Models Recall Old Conversations

A reinforcement-learning method called StateTree trains smaller models to track facts, timestamps, and changed preferences across many chat sessions.

A new training method teaches AI assistants to actually keep track of what you told them weeks ago.

Researchers describe StateTree, a reinforcement-learning method for improving how language models handle long, multi-session conversations. The problem: relevant facts get scattered across sessions, preferences change over time, and training on very long contexts is expensive and starved of good data. StateTree gets around that by building an artificial task - key-value records are hidden across sessions in a binary-tree structure, and the model must traverse from root to leaf, cross-referencing timestamps to pick the right branch and find a hidden question among decoy answers. Trained with curriculum reinforcement learning on contexts of just 10,000 tokens, the resulting models generalized to 128,000-token conversations, with a 14B version reaching 59.00% accuracy on the LongMemEval benchmark - beating QwenLong-L1-32B's 45.20% despite having less than half the parameters.

Beating a model more than twice its size on a memory benchmark is a real efficiency win, not just a bigger number. It suggests the bottleneck in long-term dialogue reasoning was never purely model scale - it was the absence of training tasks that force genuine cross-session reasoning instead of surface pattern matching. For anyone building a personalized assistant without frontier-lab compute budgets, that's the more useful lesson than the leaderboard score.

Still, LongMemEval is a benchmark, not a living relationship - remembering a quiz question across sessions is a long way from remembering that a user switched jobs and stopped caring about the old one.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →