Throwing more AI agents at a problem does not reliably make it smarter. Whether that works depends entirely on what kind of problem it is.
Researchers ran teams of up to 30 copies of the same large language model, across 13 open-weight models, sorting tasks using an old framework for grading human group work. On tasks where the team only needs one member to be right (the paper calls these disjunctive tasks), adding agents raises the odds that someone nails it by 5 to 20 percentage points. Simple majority voting barely captures that gain, though: the study found the vote matches the model's own most likely single answer to within half a point on average. Letting agents revise their answers after seeing what peers said does help, but one peer to argue with produces almost as much improvement as 29. On Fermi-style estimation tasks, where averaging independent guesses should work well, scaling barely moved the needle, because about 87% of the error comes from a bias the model itself carries into every one of its copies.
That's a real problem for anyone assuming agent swarms are a cheap reliability upgrade. The fix is not more agents, it's matching the combination method (revision, voting, or averaging) to the task's structure, and even mixing different model families only helped on the estimation problems, not the multiple-choice-style ones.
So before a product roadmap adds "multi-agent" as a feature and multiplies its compute bill by 10, it is worth asking whether the problem at hand rewards more voices, or just repeats the same blind spot 30 times over.