AI/ multi-agent debate · llm reasoning · ai research · personality prompting

Debate Order Can Beat a Smarter AI Model

New research shows AI agents that speak first skew multi-agent debate outcomes, but low agreeableness can restore a stronger model's lost influence.

Where an AI agent sits in a debate lineup can matter more than how smart it is.

Researchers studying multi-agent debate (MAD), a technique where multiple large language models critique each other's answers to sharpen reasoning, found a pronounced first-speaker bias in sequential setups. Agents that speak first disproportionately anchor the group's final answer, which means a stronger model placed later in the order can lose much of its reasoning advantage to weaker agents who spoke earlier. To counter this, the researchers tested personality prompting drawn from the Big Five trait model, focusing on agreeableness and extraversion. Assigning low agreeableness to the stronger, later-speaking agent restored its influence and improved final accuracy, while extraversion mainly changed how much agents talked rather than how much weight their arguments carried.

The result complicates a common assumption that running several models through a debate naturally averages out their differences in quality. Order effects can override raw capability, which matters for anyone building pipelines that chain multiple models together to double-check each other's work. The fix on offer is cheap: reorder who speaks, or prompt the strongest model to be a little less agreeable, rather than retraining or upgrading it.

It turns out stacking AI agents in a debate behaves less like a neutral tally and more like a meeting where whoever talks first sets the tone.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →