Large language models get partial credit on an old human skill: figuring out who's to blame.
A team of researchers built the first benchmark testing how LLMs assign responsibility and blame, grounding the work in Attribution Theory, the decades-old social psychology framework for how people judge the causes of others' behavior. Their dataset, SAB-Bench, combines a Vignette subset drawn from classic attribution-theory scenarios with a Reality subset built from real-world social narratives, totaling 7,639 responsibility and blame judgment questions. They ran 32 representative LLMs and five non-LLM baselines through the benchmark, then used a probing method to examine how five key attribution dimensions show up in the models' hidden-state representations. The goal: see not just what models conclude, but whether their internal reasoning tracks the same factors humans weigh.
The results land in a familiar middle ground for AI evaluation: measurable but incomplete agreement with human judgments, and that agreement improves as models get bigger. More notably, some attribution dimensions are decodable from specific spots in a model's hidden states, and their influence on the final judgment lines up with what Attribution Theory predicts for humans. That's a meaningful data point for anyone deploying LLMs in moderation, hiring, or legal-adjacent tools, where deciding who caused what is the whole job.
Attribution Theory has been mapping how people assign credit and blame since the 1950s. Now there's a public benchmark - code and data posted openly - for checking whether the systems reasoning alongside us are playing by anything close to the same rules.