A new algorithm can crown the best large language model in a lineup without running every model through exhaustive, expensive tests.
The paper treats LLM selection as a multi-armed bandit problem, but with a twist. Instead of scoring each model alone, the system runs head-to-head comparisons between pairs of model outputs, a setup the researchers call dueling feedback. It also weighs the fact that querying different LLMs costs different amounts of money. The authors assume a single best model, called a Condorcet winner, exists in any given lineup, a claim they say holds up across several real-world datasets. Their algorithm, based on a Track-and-Stop approach, is proven to converge on the cheapest possible testing strategy as the required confidence level approaches certainty.
For any team paying to compare a growing menu of paid APIs, this matters because it replaces guesswork with a formal stopping rule. Query just enough to be confident, then stop, rather than overspending on exhaustive tests or underspending and picking wrong. That is a real cost problem now that GPT, Claude, and Gemini-class models all charge by the token and differ wildly in price.
It is still a math paper, not a product. Turning an asymptotically optimal guarantee into something a product manager can act on requires someone to build the tool, and to handle the messier reality where there might not be one clear winner at all.