Most AI benchmarks measure one thing: can the model beat a human working alone. A new position paper says that's the wrong yardstick, and it's quietly steering the whole field.
The paper argues that today's dominant evaluation paradigm chases superhuman, autonomous performance, which implicitly treats replacing humans as the goal. Instead, the authors say, the AI community should evaluate human-AI teams together, measuring how well a person paired with a model performs versus either working solo. That shift, they argue, would push developers to build systems that complement human judgment rather than just outrun it on a leaderboard.
This matters because benchmarks are not neutral. They function as targets, and labs build toward what gets scored. If the industry's scorecards only reward solo-machine performance, the paper's logic goes, that's what labs will keep optimizing for, even when the resulting tools are worse collaborators in practice. A benchmark that scored teamwork instead could change which capabilities get prioritized, like knowing when to defer to a human or how to explain a decision, not just raw accuracy.
It's a fair critique of an incentive structure, though the harder problem is designing a team-performance benchmark that's as easy to run and compare as a leaderboard score, and this paper is a proposal, not a working replacement.