AI models think we're angrier at rule-breakers than we actually are.
Researchers built a dataset called NormReact: 450 scenarios of social norm violations, hand-annotated by humans for the emotions and behaviors both violators and observers would show afterward. They tested six large language models on two tasks: predicting how a violator would regulate their own behavior, and how an observer would respond, across different genders and levels of social closeness. The models consistently guessed that people would react with harsher punishment than human annotators actually expected, and that gap widened the further removed the observer was from the violator.
That matters because these models are already being pitched for jobs like conflict mediation and policy simulation, tasks that depend on correctly reading when people let things slide versus when they escalate. An AI that assumes every broken rule ends in public shaming or worse will nudge those simulations toward punishment and away from the tolerance and restraint that actually govern most social friction.
Real communities run on selective enforcement, not maximum severity - models trained mostly on text seem to have learned the loudest reactions, not the quiet ones where nothing happens.