A new study shows AI medical answer systems can be unanimously wrong, and the usual way of trusting consensus cannot tell the difference.
Researchers built ProbeGuard, a certified abstention framework for medical question-answering systems that vote across multiple AI samples before settling on an answer. Instead of only checking whether the samples agree, ProbeGuard tracks how that agreement formed: whether disagreement was resolved by evidence, how long minority answers held out, and whether retrieving new information changes anything. For unanimous votes, it checks whether the stated reasoning behind each sample actually coheres, then runs an active probe that feeds in counter-evidence to see if the agreement survives. A calibration step then turns those checks into a statistical bound on how often the system is allowed to be wrong.
On the MedQA benchmark, 13.4% of unanimous votes were wrong, and every agreement-based signal tested missed all of them. ProbeGuard's process signals pushed the ability to separate correct from incorrect consensus from random guessing to 0.696 AUROC, and the certified version answered six in ten unanimous questions at an observed error rate of 9.0%, improving to roughly nine in ten as more calibration data accumulated.
Multi-round AI systems have leaned on consensus as a stand-in for safety. This paper is a reminder that a room full of confident models repeating the same mistake is still a room full of wrong answers.