AI/ ai alignment · interpretability · llm safety · moral reasoning

Researchers Map How LLMs Flip Moral Frameworks Mid Answer

A new study finds LLMs switch ethical frameworks mid-reasoning most of the time, and that inconsistency makes them easier to manipulate.

A new study shows large language models rarely stick to one ethical framework when reasoning through moral questions.

Researchers tracked what they call moral reasoning trajectories, the sequence of ethical frameworks a model invokes across its intermediate reasoning steps, across six LLMs and three benchmarks. They found framework switches in 55.4 to 57.7 percent of consecutive reasoning steps, with only 16.4 to 17.8 percent of full trajectories sticking to a single framework throughout. Using linear probes, they located where each model internally encodes its current ethical framework, at layer 63 of 81 for Llama-3.3-70B and layer 17 of 81 for Qwen2.5-72B, beating a baseline that just assumes a model repeats its prior step's framework by 16.8 to 22.2 percent lower error. They also introduced a Moral Representation Consistency metric, checked against human annotators who agreed with the model's framework labeling at a cosine similarity of 0.859.

That instability has a practical cost. Trajectories that bounced between frameworks were 1.29 times more susceptible to persuasive attacks than consistent ones, a statistically significant gap (p=0.015). In other words, a model's tendency to wander between ethical frameworks may predict how easily someone can talk it into a bad answer.

Activation steering could shift that consistency-accuracy relationship, but in opposite directions for the two models tested, widening it for Qwen2.5-72B and erasing it for Llama-3.3-70B - a sign that whatever governs a model's moral consistency is model-specific plumbing, not a universal ethics module.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →