AI/ ai · ai-agents · benchmarks · evaluation

Study Finds AI Agent Leaderboards Reliable for Products, Not Models

A new study finds that popular agent benchmarks reliably rank specific chatbot-tool setups but not the underlying models, and adding more tasks barely helps.

Agent leaderboards are reliable at ranking specific AI products, but much shakier at ranking the underlying models that power them, according to a new study.

Researchers built a Bayesian variance-decomposition framework and ran it against 22 benchmarks from the Holistic Agent Leaderboard and the Harbor Index. They found that rankings of fixed model-scaffold systems - a given model paired with a specific set of tools, prompts, and instructions, often called a scaffold - were highly reliable, scoring between 0.935 and 0.994. Rankings meant to compare the underlying models on their own, independent of scaffold, scored far lower, between 0.148 and 0.841. Swapping the scaffold could change which model appears to win, and adding more tasks to a benchmark raised model-ranking reliability by at most 0.097 when limited scaffold coverage was the real bottleneck.

That gap matters because companies cite these leaderboards to justify which model they deploy or crown as the best one, but a leaderboard built to rank products isn't the same instrument as one that ranks models. The researchers also found that pooling several diverse benchmarks, rather than adding tasks to just one, raised cross-task reliability from 0.44 to 0.75 at the same evaluation budget, while cutting projected cost by as much as 83 percent.

So the next time a model tops an agent leaderboard, the useful question isn't which model won - it's how much of that win came from the scaffolding doing the work.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →