A new paper proposes a way to stop AI committees from confusing confidence with correctness.
Researchers introduce a method called Bayesian Dialectical Argumentation, or BDA, for aggregating answers from a "council" of multiple large language models that deliberate on a question together. Instead of just tallying votes or raw confidence scores, BDA treats each model's moves in the debate - proposing an answer, challenging another model's answer, or conceding - as evidence about that model's reliability, the same logic used in classical annotator-reliability models from crowdsourcing research. It then weights each model's input by its inferred reliability, producing a posterior probability meant to track actual odds of being correct rather than how decisive the group sounds. The method needs no extra LLM calls, and across binary and multi-class benchmarks the researchers report it beats other zero-cost aggregation methods on calibration while holding up better when some models behave as persistent adversaries.
Multi-model setups increasingly ship a confidence number next to an answer, but that number usually measures decisiveness, not correctness - a gap that matters once these councils feed into automated decisions nobody double-checks. BDA's other trick is handling persistently unreliable models by inverting their signal instead of simply being outvoted by them, a different failure mode than the standard majority-vote setup most councils still use.
Call it a reminder that polling several chatbots and averaging their answers isn't a free calibration fix - a lesson the multi-agent debate literature has had to relearn since the first chain-of-thought voting papers.