AI models get more accurate when they argue with each other first, but only if the arguing partners are different models.
Researchers adapted a three-stage deliberation process previously tested on human forecasters and applied it to large language models from three different model families. They ran it across four tasks: estimating quantities in images, reviewing machine-learning papers, spotting an AI agent hiding malicious behavior, and picking winners in sports betting markets. In every case, small groups that discussed an answer before finalizing it beat groups whose independent guesses were simply averaged, the classic wisdom-of-crowds approach. Individual models also kept some of that accuracy boost even after the group session ended. The effect disappeared entirely when a group was made up of copies of the same model; diversity across model families was required for deliberation to help at all.
That last finding is the useful part. It suggests one practical fix for AI unreliability is not a bigger model but a mixed committee of different ones, arguing it out, much like human juries or forecasting teams often outperform a single expert. The same mechanism that improved sports predictions also caught hidden bad behavior in an AI agent, which matters more as companies lean on AI to supervise AI.
Running three different frontier models to cross-examine each other costs roughly three times as much as running one, so this is a quality upgrade, not a cost saver.