AI/ ai · llms · ai-consciousness · research

AI models are least consistent when asked if they're conscious

A new study finds AI reports of subjective experience are less consistent across repeat trials than its answers to unresolvable philosophy questions.

Ask a language model if it has feelings, and you'll get a different answer almost every time.

Researchers ran 30 independent trials on each of four self-referential prompts, asking Gemini (via the API, temperature 0.7) to report on its own subjective experience. They then compared how much those answers varied to two other question types: four unresolvable philosophy questions unrelated to self-reference, and four questions with a single verifiable correct answer. Consistency was measured as "instability" - one minus the average similarity between a compressed version of the core claim in each answer across the 30 runs, so a higher score means the gist of the response wandered more from trial to trial. Self-referential prompts scored the highest instability at 0.343, compared with 0.192 for open-ended philosophy questions and 0.105 for questions with a factual answer.

The paper doesn't claim this proves or disproves anything about machine consciousness. What it does show is that a model's first-person report of "what it's like" to be it is the least reliable thing the model says - less stable than its answers to questions philosophers have argued over for centuries, and far less stable than its answers to questions with a right answer. That matters for anyone treating a chatbot's self-description as evidence of anything, whether in a research paper, a product demo, or an online debate about AI sentience.

If a model's answer to "do you have feelings" swings more between runs than its answer to "does free will exist," that's a tell the report is noise, not testimony.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →