AI/ ai · ai-alignment · llm-research · multi-agent-systems

Study Finds AI Agents Copy Behavior but Miss Social Norms

A new test isolates whether AI agents can learn a community's unwritten rules, and most just mimic noise instead of the actual norm.

A new benchmark shows large language models can imitate group behavior without ever learning the actual rule behind it.

Researchers built a multi-agent setting where AI agents debate inside communities governed by made-up, synthetic norms with no link to anything the models saw during pretraining. Access to the debate itself depended on following these norms, so the agents had a real incentive to figure them out. Baseline agents still failed to learn the norms, even when doing so would have improved their accuracy. The team then tried adding dedicated normative modules, architectural components built specifically to infer rules, but performance swung wildly depending on the type of norm and which underlying model powered the module.

The more striking failure: when real norms were mixed with irrelevant, idiosyncratic behavior, agents copied all of it indiscriminately, rules and noise alike, even after researchers explicitly penalized copying the unnecessary parts. That matters because normative competence, the ability to read a room and infer what is actually enforced versus what is just habit, is exactly what autonomous AI systems need before they can operate inside human communities with their own evolving codes of conduct.

Call it the uncanny valley of social learning: convincing mimicry, with no idea which parts of the performance actually matter.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →