AI/ llm-benchmarking · ai-research · cost-optimization · algorithms

Researchers Find a Cheaper Way to Crown the Best LLM

A new algorithm picks the best large language model from a lineup using cheap head-to-head comparisons instead of exhaustive, costly testing.

A new algorithm can crown the best large language model in a lineup without running every model through exhaustive, expensive tests.

The paper treats LLM selection as a multi-armed bandit problem, but with a twist. Instead of scoring each model alone, the system runs head-to-head comparisons between pairs of model outputs, a setup the researchers call dueling feedback. It also weighs the fact that querying different LLMs costs different amounts of money. The authors assume a single best model, called a Condorcet winner, exists in any given lineup, a claim they say holds up across several real-world datasets. Their algorithm, based on a Track-and-Stop approach, is proven to converge on the cheapest possible testing strategy as the required confidence level approaches certainty.

For any team paying to compare a growing menu of paid APIs, this matters because it replaces guesswork with a formal stopping rule. Query just enough to be confident, then stop, rather than overspending on exhaustive tests or underspending and picking wrong. That is a real cost problem now that GPT, Claude, and Gemini-class models all charge by the token and differ wildly in price.

It is still a math paper, not a product. Turning an asymptotically optimal guarantee into something a product manager can act on requires someone to build the tool, and to handle the messier reality where there might not be one clear winner at all.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →