AI/ ai · benchmarks · enterprise-ai · evaluation

New Protocol Measures AI Systems, Not Model Names

A new protocol called IB2 grades AI systems by serving route, not model name, after finding identical weights pass or fail depending on setup.

A new benchmarking protocol argues the AI industry has been grading the wrong thing: the model's name, not the actual system a business runs.

The paper's authors audited 18 existing benchmarks and found every one scores the advertised model identifier, ignoring that real performance depends on the serving route, precision, output contract, and harness wrapped around those weights. Their proposed protocol, called IB2, has three parts: a preflight check that verifies a serving route can even execute the evaluation before scoring starts, a scoring rule that keeps failures in the tally instead of discarding them, and an adjudication process kept blind to which system produced which result. They tested it across eleven systems using 128 locked tasks and 987 assertions covering document, spreadsheet, chart, tool, and database work. The task set itself stays sealed; the protocol is what's being released, not the benchmark.

The results undercut a basic assumption behind AI leaderboards: that a model's name is a stable stand-in for what it can do. Two runs on identical weights failed the binding check in different ways, while switching serving configuration alone moved one system's declared score from 77.38 to 82.54. For a company picking a vendor off a leaderboard number, that swing is the difference between a system that works and one that quietly doesn't.

Worth some skepticism, though: the underlying task corpus is sealed and self-reported, so outside teams can't reproduce the eleven-system comparison. The numbers illustrate the method, not an independently verified ranking.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →