A new arXiv study finds that bigger AI models don't automatically win arguments with smaller ones.
Researchers tested seven open-weight language models on three language-understanding tasks, pairing them up to see what happens when one model disagrees with another. They measured persuasion as the probability that a model abandons its original answer after hearing a peer's dissenting explanation. The shifts were large: when models disagreed, the "listener" often dropped its first judgment entirely. Neither a model's own consistency nor its size reliably predicted this. Some models that were almost perfectly steady on their own turned out to be among the easiest to talk out of a correct answer, while small models proved just as capable of persuading, and resisting persuasion, as much larger ones.
That has real implications for multi-agent AI systems, which increasingly mix models from different labs and size classes on the assumption that scale confers authority. The study found the size of a judgment shift depends more on how persuadable the listener is than on how convincing the speaker is, so a small model running as a dissenting agent can overturn a much larger model's verdict, depending on which two models are paired.
In other words, you can't predict how a committee of AI agents will behave by looking at each member's spec sheet. The outcome hinges on the specific combination deployed, not which model has the most parameters.