AI/ llm-voting · test-time-scaling · ensemble-inference · ai-research

Why Your LLM Ensemble Might Bury the Right Answer

A new analysis shows that finding the correct answer in LLM voting ensembles is not the same as that answer actually winning the vote.

Sampling more answers from a large language model does not guarantee the right one wins the vote.

Researchers studying LLM ensemble voting - the practice of asking a model the same question many times and picking the most common answer - mapped out exactly when a correct answer, once it shows up among the samples, still has enough remaining calls left to actually win. They derived a precise threshold for that "recoverability" window and showed it only shrinks as more responses come in. They also found that lumping all wrong answers into one bucket does not improve accuracy, while simply reordering the input text can shift which answer ends up winning by double digits. From the threshold, they built a stopping rule that locks in the final answer early once no amount of further sampling could change it, cutting 28-30% of calls in a 16-call budget without altering a single output.

Test-time scaling - burning extra inference calls and voting on the results - is one of the standard ways labs squeeze more accuracy out of existing models without retraining them, and it is not free. This paper quantifies how much of that extra spend is wasted once a correct answer has already lost its chance to win, and converts that waste into calls you can skip. For anyone running multi-sample pipelines at scale, that is a near-free cost cut.

It will not make a mediocre model smarter. It just stops you from paying to find an answer you were never going to use.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →