AI agents built on memory can convince themselves of things that never happened.
A new paper proposes FAME, a training-free framework for catching false memory in autonomous agents: beliefs that drift from spurious correlations, shifting environments, or conflicting knowledge rather than from the facts at hand. Instead of grading an agent's final answer, FAME tracks how its internal beliefs move under counterfactual scenarios, reading concept drift straight out of the agent's hidden states before it ever generates a response. That sidesteps the need for reward models or sampling multiple answers to guess at what the agent actually believes. Tested across math reasoning (GSM-Symbolic), code generation (GitChameleon), and complex reasoning (BigBench-Hard), FAME hit AUROC scores of 76.2% to 96.7%, beating the best existing baseline by 3.4 to 23.3 percentage points.
The bigger point: answer checking alone keeps missing this failure mode, according to the researchers. That matters because agent memory is becoming the default way to stretch a model past its training data, through retrieval, long context, and multi-step tool use, and each of those mechanisms can quietly encode a wrong belief that still produces a confident, plausible answer. A detector built on belief states rather than outputs is a more honest proxy for what's actually happening inside the system.
Still, detecting a false memory and fixing one are different problems. FAME flags the drift; it doesn't tell the agent how to correct it.