AI/ ai alignment · machine learning · ethics · interpretability

A New Model Teaches AI to Judge Morality by Context

A new framework doubles how well AI predicts human moral judgments by learning context-specific rules instead of relying on raw LLM reasoning.

A new AI framework scores right and wrong the way humans actually do: case by case, not rule by rule.

Researchers built a system called COMETH that learns moral judgments from context instead of fixed rules. The team collected 300 scenarios spanning six core actions that violate one of three rules - don't kill, don't deceive, don't break the law - and had 101 people rate each as blameworthy, neutral, or acceptable. COMETH uses an LLM filter and MiniLM embeddings to standardize those actions into clusters, then a probabilistic model groups scenarios by how people actually judged them, pulling out specific contextual features that explain the verdict. On the test set, COMETH matched human judgments about 60 percent of the time, roughly double the 30 percent score from prompting an LLM directly for a moral call.

That gap matters because most AI moral reasoning today runs through exactly the kind of end-to-end LLM prompting this study beats. A chatbot asked whether an action was wrong gives an answer, but not a reason you can audit or correct. COMETH's feature weights are inspectable, so a developer can see which contextual cue - intent, consent, an alternative being available - tipped a judgment one way. For a field built on models that increasingly referee human behavior, from content moderation to agent guardrails, that transparency is worth more than another leaderboard score.

Still, 60 percent agreement on 300 lab scenarios rated by 101 people is a promising start, not a verdict - real moral disputes rarely come pre-sorted into three tidy categories.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →