AI/ ai research · cooperative ai · game theory · benchmarking

Study Finds AI Benchmarks Miss How Humans Read Hints

A Hanabi study finds AI teams rely on almost no unspoken conventions, while human pairs lean on them heavily, casting doubt on AI-AI benchmarks.

Researchers just found a fairness gap in how we test cooperative AI: it's the humans doing all the mind-reading, not the machines.

A team of AI researchers replayed roughly 101,000 actions from three public Hanabi datasets: human-human games, AI-AI games (from a benchmark called HOAD), and human-AI games (from HanabiData). For each play, they calculated a convention gap - the difference between the failure rate predicted from the literal information in a hint and the failure rate players actually achieved. Human pairs beat their literal-information prediction by 26.2 percentage points, mostly by correctly playing cards that had received no hint at all. AI pairs showed essentially no gap, just -0.7 points, meaning AI partners communicate almost entirely through the literal content of their signals.

That distinction matters because most cooperative-AI research judges an agent by how well it plays against other AI, on the assumption that scoring well there predicts scoring well with a human partner. This study suggests that's the wrong yardstick: among human-AI pairs, the AI partner that prompted the largest convention gap in its human teammate also caused the fewest mistakes, while raw game score depended more on which players showed up than on how well the pairing actually communicated.

In other words, an AI that aces cooperation drills against other bots might still be lousy company for people who read between the lines.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →