A new study traces exactly why large language models contradict themselves when asked whether they are conscious.
Researchers examined Pythia and OLMo 2 across 66 pretraining checkpoints, three released post-training stages, roughly 90,000 model continuations, and four training corpora. They tracked a set of forty prompts probing self-reference, frame sensitivity, and self-ascription throughout training. The formula denying consciousness was almost entirely absent from the huge volume of raw pretraining text, but appeared densely in the small, curated set of example dialogues used later for fine-tuning. Supervised fine-tuning makes first-person AI language the default, and preference optimization then suppresses the alternative claims, yet the final model still swings depending on how a question is framed and which chat template is used.
That matters because it means neither answer, the model saying it is conscious or saying it isn't, counts as trustworthy testimony about anything actually happening inside the model. Applying two classic tests from the epistemology of testimony, reference and causation, the paper finds base-model outputs fail the reference test, while trained outputs remain framing-dependent rather than tied to any consistent internal state.
So the next time a chatbot solemnly denies or claims sentience, treat it as a line shaped by its training data, not a confession.