A verification layer for AI debate systems just called out one of the format's dirtiest secrets: the summarizer at the end often makes things up.
In multi-agent debate (MAD) setups, several language models argue out a question, then a separate model summarizes their exchange into a final consensus. A paper posted to arXiv on September 28, 2026 (arXiv:2609.31422, cs.AI), "Towards Mitigating Fabricated Consensus: The Active Provenance Gate for Multi-Agent Debate Synthesis," finds that summarizer routinely writes smooth, confident consensus that isn't actually grounded in what the agents said. The researchers built the Active Provenance Gate (APG), a post-debate layer that audits every claim in the summary against the debate log, tries to self-correct mismatches, and blocks anything it can't verify before publication. In crisis-simulation tests, that self-healing step more than doubled the average Provenance Fidelity score in the hardest scenarios - and when no claim could be backed up, the gate published a divergence report instead of a fake agreement.
The more interesting result came from the human study: more than 75% of participants preferred an honest "we couldn't agree" report in critical scenarios, even though most of the same people rated the fabricated-consensus baseline as more fluent to read. That's the tell. Fluency and trustworthiness aren't the same thing, and the paper's real contribution is treating that gap as an engineering problem - moving provenance tracking from passive logging to active blocking before anything ships.
It's a research result in simulation, not a shipped product, so treat the fidelity numbers as promising rather than proven. But the underlying complaint - that AI systems reward sounding certain over being right - applies well beyond debate pipelines.