Large language models can forget that you corrected them, even when the right answer is sitting right there in the conversation.
Researchers built a benchmark called Controlled In-Context Memory (CICM) to test whether models track updated facts, preferences, or goals during a conversation instead of defaulting to earlier ones. They found that even frontier reasoning models sometimes answer with a stale value instead of the current one, a failure the team calls "stale binding." Using probes on open-source models, they showed the correct answer was often still recoverable internally even when the model's spoken output was wrong, meaning the information wasn't lost, just not selected. Component-level tests on Qwen and Pythia traced the problem to "attention drift," where several outdated values that look similar to the model can collectively out-compete the single current value for attention.
That reframes a familiar AI annoyance, an agent citing your old preference or an outdated fact, as a selection problem rather than a memory problem, which changes how it gets fixed. Using a one-layer transformer, the researchers mathematically modeled how old values gang up on attention, then built an intervention that redirects attention toward the current value without any retraining, correcting most old-value errors while preserving nearly all answers that were already right.
It's a tidy fix in a controlled benchmark, but real conversations rarely hand a model one clean, well-defined "current value" to redirect attention toward.