A new paper finally explains the math behind why giving AI models more thinking time sometimes pays off and sometimes backfires.
Researchers built an analytically solvable model of an LLM-as-a-judge setup, using Bayesian linear regression with a reward-weighted sampler, to study inference-time scaling - the practice of generating many candidate answers and picking the best one via a judge model instead of training a bigger one. They found that when the judge's reward signal is close to the true quality signal, pulling more samples steadily lowers error, decaying at a predictable 1/k-squared rate. But when the reward is badly misspecified, there's a finite number of samples beyond which more sampling makes results worse, and for any fixed sample count there's an optimal sampling temperature. They confirmed the pattern experimentally using an LLM as judge, and found the benefit of extra inference compute shrinks as tasks get harder.
That matters because companies are pouring inference budget into best-of-k sampling and judge-based reranking on the assumption that more samples always means better answers. This gives engineers an actual formula for when that bet pays off, and a warning for when it quietly stops working - which, inconveniently, is exactly on the hardest problems where teams lean hardest on extra compute.
So the industry's favorite inference-time scaling story comes with an asterisk, and the asterisk gets bigger exactly where it matters most.