AI/ ai research · llm evaluation · model interpretability · open-source ai

Study Finds LLMs Know More Than They Can Actually Say

A new framework traces AI answers back to training data, showing models often fail to retrieve facts they know and can't tell real knowledge from guesses.

A new diagnostic tool checks whether an AI model's correct answers are real knowledge or lucky guesses.

Researchers built a framework called LUMOS and ran it on OLMo 2, an open-source model whose full training data is public, letting them trace specific facts from the training corpus straight through to what the model says out loud. They found models store rare facts internally with 84 percent separability, meaning the information is clearly there, but only surface the correct answer 54 percent of the time when asked directly. That retrieval gap shrinks as models get larger. When models were asked to judge their own answers, they hit 83 percent accuracy on facts drawn from their training data, but dropped to 49 percent, essentially a coin flip, on facts outside it, and asking for chain-of-thought reasoning only made the wrong answers sound more confident.

This matters because most knowledge benchmarks grade a model purely on its output, treating a wrong answer as proof of ignorance. LUMOS shows that framing is often false: the model may know the fact and simply fail to retrieve it, a distinction that changes how much weight anyone should put on a single wrong answer during an eval or a product demo. It also undercuts a popular safety assumption, since self-reflection and chain-of-thought prompting are routinely pitched as ways to catch a model's own mistakes, yet here they just amplified false confidence instead of fixing it.

The catch is that this kind of audit only works because OLMo 2 publishes its training data; try the same trace on GPT or Claude and there is nothing to point at.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →