A new technique lets language models reach confident answers using far fewer guesses.
Self-consistency is the standard trick for squeezing reliable answers out of an LLM: run the same prompt many times, generate several reasoning paths, and take whichever final answer shows up most often. A new paper argues that approach throws away information, because each reasoning trajectory doesn't just hand back one answer. It comes from a probability distribution over possible answers, visible in the model's log-probabilities. The researchers built an algorithm called ASC-D that reads that distribution after every trajectory and stops sampling as soon as it's statistically confident it has identified the true most-likely answer.
Fewer trajectories mean fewer, cheaper model calls, which matters to anyone running reasoning pipelines at scale rather than toy demos. Tested on the MMLU-Redux benchmark, ASC-D needed 46.4 to 95.6 percent fewer trajectories than existing adaptive self-consistency baselines, while still posting the highest rate of correctly certified answers within a fixed compute budget across three open-source models.
It's a reminder that the costly part of asking an AI the same question five times and voting was never the five times. It was ignoring the math the model had already done.