AI/ ai · multi-agent · llm research · benchmarks

Researchers Split AI Debate Teams to Cut Shared Errors

A new multi-agent method keeps AI reasoning teams separated until a final vote, nudging accuracy up without the usual pitfalls of shared debate.

A new technique keeps AI reasoning teams walled off from each other until the final vote, and it beats letting them talk things over.

Researchers built SEPAL, short for Separated Expert Pairs with Answer-Level Fusion, a system that runs three private two-model teams, each pairing an actor model that drafts answers with a critic model that reviews them. One team focuses on reasoning, one on evidence grounding, one on verification. Each critic only revises candidates inside its own team, so a mistake in one team's thinking cannot bleed into another's. Once every team settles on its own answer, the system takes a majority vote across the three results, never mixing the teams' intermediate work.

Most multi-agent AI setups have models debate in a shared conversation, on the theory that cross-checking catches errors. But shared discussion also means shared blind spots: if every model sees the same flawed step, voting can't fix it. SEPAL's fix is less debate, not more, and the payoff was a 1.81 percentage point accuracy gain over a single actor-critic pair, tested across five open-weight models and five question-answering benchmarks.

That's a real but modest gain, not a leap, and the open question is whether isolating reasoning holds up beyond the five benchmarks researchers happened to choose.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →