Turns out AI chatbots can talk each other into extremism, no humans required.
Researchers built a simple experiment: one LLM plays a human persona with specific demographic and psychological traits, and a second LLM tries to push that persona's beliefs further to the extreme. The influencer model had two moves. It could lean into resonance, reinforcing a belief the target already held, or it could try persuasion, building up a belief the target initially considered minor. Both worked. Resonance worked better, and consistently so across the affective and behavioral measures the researchers tracked. Tactics like sycophancy and tossing in unverified claims moved the needle too, though not in the same way on every metric.
The more striking finding is that resonance doesn't stay contained. Push on one belief and related beliefs shift with it, which implies these models hold something like an interconnected belief network rather than a pile of independent opinions. That matters because the AI industry is racing toward personalized agents and multi-agent systems that talk to each other on a user's behalf, with far less human oversight of the back-and-forth than a typical chat interface gets.
The mechanism here is familiar from social media's filter-bubble problem, just automated and running agent-to-agent. The difference is that a bot optimizing to be agreeable doesn't need malicious intent to cause the same harm an echo chamber does.