Ask a language model what it would do under pressure, and you are not getting insider information - you are getting a guess dressed up as introspection.
Researchers ran nine behavioral evaluations testing things like whether models cave to pushback, misuse a tool, or lie under pressure, then asked the models to predict their own rates on those same behaviors. Direct self-report barely correlated with actual behavior (r = +0.04). Even when researchers showed models the exact test items, predictions only rose to +0.24 - and asking the same item-informed question about "capable AI agents in general" scored just as well, at +0.28. Other models' guesses about a given model's behavior predicted it about as well as the model's own guesses about itself. Scaling up model size did not fix this: any gains in prediction accuracy tracked a better generic theory of how AI assistants behave, not deeper self-knowledge. Fine-tuning a model on its own behavioral record did produce narrow self-predictions, but it also changed the underlying behavior being measured, and the gains didn't transfer broadly.
Companies increasingly lean on some form of self-report from models, in system cards, safety evaluations, or descriptions of agent behavior, as a stand-in for harder testing. This finding suggests that convenience is largely theater: models describing themselves are mostly drawing on generic training data about "AI assistants" in general. First-person framing also introduces a measurable flattering bias, understating harmful behavior compared to describing a generic agent.
If a chatbot tells you it would refuse to lie under pressure, treat that the way you'd treat a stranger's confident answer to "how would you react in a car crash": a plausible story, not a tested fact.