AI/ ai-leaderboards · benchmarking · chatbot-arena · ai-safety

A Statistical Fix for Rigged AI Leaderboards

A new certificate math can bolt onto any AI leaderboard, flagging exactly when enough votes have been rigged to make its rankings untrustworthy.

A new statistical certificate can tell you exactly how many votes would need to be faked before an AI leaderboard's ranking should be trusted.

Researchers propose a "certified corruption budget" that is recalculated after every new vote or benchmark record and published alongside each head-to-head ranking claim. The guarantee: with high confidence, either the ranking is correct, or more votes than the published budget number were corrupted. Unlike older methods that assume good-faith voting or only bound manipulation one step at a time, this one holds even against an attacker who watches every published certificate and times a burst of fake votes to maximize damage. It also handles labs that quietly test many private model variants and publish only the best one, charging just a small, log-scaled penalty for that selection.

The timing matters. Chatbot Arena and similar leaderboards now shape which models get press coverage, enterprise deals, and funding, which makes them worth gaming. Replaying the method on 1.8 million real Chatbot Arena votes, the researchers found that a few hundred rigged votes were enough to make standard statistical confidence intervals certify a false ranking, while their method correctly flagged the risk and stayed valid throughout. Clearly separated models, they found, can absorb roughly 2,000 forged votes before the certificate breaks.

It is not a fix for vote rigging, just an honest price tag for it. That is still more than most leaderboards currently offer: a number, updated live, telling you how much to trust what you are looking at.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →