AI/ ai · llm research · ai memory · machine learning

AI Models Can Know a Fact and Still Fail to Retrieve It

New research shows language models can memorize a fact without learning to retrieve it from every phrasing, and the gap traces to a specific internal state.

A new study finds that teaching a language model a fact and teaching it to answer questions about that fact are two different skills.

Researchers trained models in two stages to pull these apart. In stage one, one model saw facts phrased five different ways, such as "The capital of X is Y" and "The capital of X:", while a control model saw the same facts only as plain statements. In stage two, both models learned a new batch of facts, but this time only as statements. Despite getting identical training on the new facts, the model that had earlier practiced varied phrasings could answer questions about them in multiple formats, while the statement-only model mostly drew a blank unless asked in the exact format it had trained on.

The researchers traced this to what they call the "context state," the hidden state right before a model produces its answer. Models exposed to varied phrasings built more consistent context states across different question forms, and directly manipulating that state could switch retrieval of an already-learned fact on or off. The effect held across different kinds of facts, like capitals and currencies, suggesting it is a general wiring issue rather than a quirk of any one topic.

It is a reminder that quizzing a model one way and getting silence does not mean the fact was never learned - it might just be stuck behind the wrong question.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →