A new paper explains why asking a chatbot the same question dozens of times eventually stops paying off.
Self-consistency is a common trick for boosting large language model accuracy: have the model reason through a problem several times, then take the most common answer as a majority vote. Researchers noticed these gains plateau as you add more attempts, but nobody had a way to predict when. This paper borrows "diversity combining" math from wireless communications - the same framework engineers use to combine signals from multiple noisy antennas - and treats each reasoning attempt as a noisy observation. The correlation between those attempts caps how much a majority vote can ever help, no matter how many you sample. Across 5 models and 12 benchmarks, the authors found that simply varying the prompt template for each attempt cuts that correlation in 55 of 57 test cases, with the biggest payoff on open-ended question answering, and they built an Adaptive-K rule that samples just four attempts to predict the ideal number to use, keeping 96-103% of the accuracy you'd get from sampling 32.
Self-consistency sampling isn't free - every extra reasoning attempt is another full model call, which means real compute cost. This gives developers a formula for knowing when to stop paying for near-zero marginal accuracy, plus a cheap trick, varying the prompt template, that recovers some of the lost gains without extra sampling. That's unglamorous work, but it's the kind that actually shows up on a production compute bill.
It's also a useful reminder that a chunk of what gets called "reasoning" in LLMs is really correlated guesses voting on themselves, bounded by the same math that governs how radios cope with static.