Throwing more AI agents at a task does not reliably make the answer smarter - it depends heavily on the task.
A new study tested eight ways of orchestrating teams of small language models, each with 7 to 9 billion parameters, scaling from three to thirty model calls per task. Across benchmarks, the payoff split sharply: on arithmetic word problems like GSM8K and GSMHard, accuracy rose by as much as 17 points as teams grew. On ARC, GPQA, and MMLU, the gain topped out at four points, regardless of architecture. One setup, Proposer-Critic, where one model drafts and another critiques, captured most of the arithmetic gains and beat every other architecture at the largest team size - but it ranked among the weakest performers elsewhere. No single architecture won across every task type.
The researchers explain the pattern with a formula that splits any accuracy change into two parts: how many correct answers a team's proposals cover, and how well a downstream step converts that coverage into a final answer. Arithmetic problems apparently leave room for more right answers to surface as teams scale, and a critic can exploit that room. Multiple-choice tests either run out of that room quickly or resist a critic's attempts to sort through it, and on code generation the conversion step barely helps at all. At equal call budgets, token cost between architectures still varied by a factor of 2.1.
That's a useful corrective for the "just add more agents" instinct spreading through agentic AI product design. More calls buy more candidate answers, not more intelligence - and which architecture cashes that opportunity in depends entirely on what you're asking it to do.