AI/ ai-consciousness · language-models · ai-research · model-training

Study Traces Why Chatbots Flip on Claims of Consciousness

Researchers traced why AI models flip between denying and claiming consciousness, tracing it to training data and question framing, not real self-knowledge.

A new study traces exactly why large language models contradict themselves when asked whether they are conscious.

Researchers examined Pythia and OLMo 2 across 66 pretraining checkpoints, three released post-training stages, roughly 90,000 model continuations, and four training corpora. They tracked a set of forty prompts probing self-reference, frame sensitivity, and self-ascription throughout training. The formula denying consciousness was almost entirely absent from the huge volume of raw pretraining text, but appeared densely in the small, curated set of example dialogues used later for fine-tuning. Supervised fine-tuning makes first-person AI language the default, and preference optimization then suppresses the alternative claims, yet the final model still swings depending on how a question is framed and which chat template is used.

That matters because it means neither answer, the model saying it is conscious or saying it isn't, counts as trustworthy testimony about anything actually happening inside the model. Applying two classic tests from the epistemology of testimony, reference and causation, the paper finds base-model outputs fail the reference test, while trained outputs remain framing-dependent rather than tied to any consistent internal state.

So the next time a chatbot solemnly denies or claims sentience, treat it as a line shaped by its training data, not a confession.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →