AI/ ai · llm benchmarks · ai research · social reasoning

LLMs Partly Mirror Human Blame Judgments, Study Finds

A new benchmark probes how 32 language models assign blame and responsibility, finding bigger models align more with human judgment but still fall short.

Large language models get partial credit on an old human skill: figuring out who's to blame.

A team of researchers built the first benchmark testing how LLMs assign responsibility and blame, grounding the work in Attribution Theory, the decades-old social psychology framework for how people judge the causes of others' behavior. Their dataset, SAB-Bench, combines a Vignette subset drawn from classic attribution-theory scenarios with a Reality subset built from real-world social narratives, totaling 7,639 responsibility and blame judgment questions. They ran 32 representative LLMs and five non-LLM baselines through the benchmark, then used a probing method to examine how five key attribution dimensions show up in the models' hidden-state representations. The goal: see not just what models conclude, but whether their internal reasoning tracks the same factors humans weigh.

The results land in a familiar middle ground for AI evaluation: measurable but incomplete agreement with human judgments, and that agreement improves as models get bigger. More notably, some attribution dimensions are decodable from specific spots in a model's hidden states, and their influence on the final judgment lines up with what Attribution Theory predicts for humans. That's a meaningful data point for anyone deploying LLMs in moderation, hiring, or legal-adjacent tools, where deciding who caused what is the whole job.

Attribution Theory has been mapping how people assign credit and blame since the 1950s. Now there's a public benchmark - code and data posted openly - for checking whether the systems reasoning alongside us are playing by anything close to the same rules.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →