A new paper argues that AI benchmarks are not neutral measuring sticks - they are tools that concentrate power in a handful of well-funded labs.
The paper draws on philosopher Iris Marion Young's theories of oppression and structural injustice to analyze benchmarking culture in AI research. It argues that leaderboards reward state-of-the-art performance with prestige, citations, trust, and institutional influence, and that as the cost of building competitive systems rises, those rewards increasingly flow to industry-backed labs. The authors map current benchmarking practices onto four of Young's "faces of oppression," arguing the harms appear even when no individual lab or researcher does anything wrong. They frame the problem as structural injustice: normalized, individually defensible practices and network effects that quietly narrow which research directions get pursued.
If the analysis holds, benchmark culture isn't just a metrics problem - it's a gatekeeping mechanism that decides whose research counts. That matters most for smaller labs and academic groups that can't afford the compute-heavy runs needed to top a leaderboard, even when their underlying ideas are sound.
Benchmarks became AI's de facto currency because they are easy to cite and easy to game. A paper questioning that currency's fairness is unlikely to change what conference reviewers and funders reward anytime soon.