AI/ ai-ethics · llm-research · ai-alignment · arxiv

Study Finds LLMs Contradict Their Own Moral Rules

A new arXiv study finds GPT, Mistral, and Llama models contradict their own stated ethical principles up to 78% of the time when the same scenario is reframed.

A new study finds that large language models routinely contradict their own moral judgments when the same ethical scenario is simply reworded.

Researchers built sets of scenarios that were logically identical but framed differently across three schools of ethics: deontology, utilitarianism, and virtue ethics. They ran these through GPT, Mistral, and Llama models, then converted each response into structured logical statements to check for contradictions against the model's own prior answers. Contradiction rates hit as high as 78% within a single school of ethical thought. The models weren't just disagreeing across different frameworks, which might be expected. They were flip-flopping within one.

This matters because these models are already being used to draft content moderation calls, summarize policy debates, and advise on morally loaded questions, all places where consistency isn't a nice-to-have. A model that condemns an action in one phrasing and condones it in another isn't reasoning from principles at all. It's pattern-matching to surface features of the prompt, then dressing the output up in the language of ethics.

The researchers frame this as a warning shot for AI alignment work broadly: you can't align a system's values to yours if the system can't even align its values to itself. That's a more fundamental problem than the usual alignment conversation about whose values get encoded. Before anyone argues about which ethics to train in, this suggests checking whether the model can hold any ethical position steady at all.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →