AI/ ai · llm-research · interpretability · arxiv

Researchers Pinpoint Why AI Models Cling to Old Answers

New research shows LLMs sometimes answer with outdated facts because attention skews toward old values, and a no-training fix corrects most of these errors.

Large language models can forget that you corrected them, even when the right answer is sitting right there in the conversation.

Researchers built a benchmark called Controlled In-Context Memory (CICM) to test whether models track updated facts, preferences, or goals during a conversation instead of defaulting to earlier ones. They found that even frontier reasoning models sometimes answer with a stale value instead of the current one, a failure the team calls "stale binding." Using probes on open-source models, they showed the correct answer was often still recoverable internally even when the model's spoken output was wrong, meaning the information wasn't lost, just not selected. Component-level tests on Qwen and Pythia traced the problem to "attention drift," where several outdated values that look similar to the model can collectively out-compete the single current value for attention.

That reframes a familiar AI annoyance, an agent citing your old preference or an outdated fact, as a selection problem rather than a memory problem, which changes how it gets fixed. Using a one-layer transformer, the researchers mathematically modeled how old values gang up on attention, then built an intervention that redirects attention toward the current value without any retraining, correcting most old-value errors while preserving nearly all answers that were already right.

It's a tidy fix in a controlled benchmark, but real conversations rarely hand a model one clean, well-defined "current value" to redirect attention toward.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →