Researchers tested whether having AI models debate each other beats just asking one model to try multiple times - and for the most part, it doesn't.
The study ran 23 open-weight language models from eleven vendor families through five tasks and more than 5,500 debate and control runs. The team varied "cognitive diversity" three ways - assigning agents personas, changing sampling temperature, and mixing different model identities - and compared each debate setup against a majority-vote baseline using the same compute budget. Debate did beat a single model by 3 to 7 points on tasks with room to improve. But when matched for budget, it merely tied or lost to simple repeated sampling, despite costing 1.6 times the wall-clock time and 3.4 times the tokens. The researchers also found a bug where debate transcripts silently overflowed model context windows; fixing it moved the debate-versus-sampling comparison from debate trailing by 1.8 points to a dead heat.
That reframes a lot of hype around multi-agent debate systems. If the accuracy gains are really just an artifact of generating more answers and voting on them, cheaper sampling gets the same result without the theater of agents "talking" to each other. Assigning personas made things worse, not better, and most of debate's benefit showed up after the very first round of answers - so the elaborate back-and-forth many products lean on may be doing less than advertised.
Multi-agent frameworks have sold themselves on the idea that specialized, distinct agents reason better together. This paper's budget-matched baseline is a useful gut check: before trusting a debate pipeline's price tag, ask whether the same money spent on plain sampling would have gotten you there anyway.