AI/ ai · llm-evaluation · benchmarks · kaggle

Kaggle's Game Arena Makes LLMs Play Chess Poker and Werewolf

The Kaggle-built platform pits LLMs against each other in chess, poker, and Werewolf to measure strategic reasoning without benchmark saturation.

Kaggle has built an arena where language models play chess, poker, and Werewolf against each other instead of taking another multiple-choice test.

The project, called Game Arena, pits large language models in head-to-head matchups across three pilot environments: chess for perfect-information play, poker for hidden-information strategy, and Werewolf for multiplayer social deduction. A new technical report lays out the platform's infrastructure and the evaluation metrics used to score full competitions run across multiple models. Kaggle says the arena is open and meant to keep expanding with new games and variants over time, so gameplay strength can keep climbing instead of stalling out. The goal is to measure planning, adaptation, and how models handle uncertainty, not just recall.

That distinction matters because standard benchmarks are running out of room. Once a test's questions circulate widely enough, models start scoring well by pattern-matching rather than reasoning, and the numbers stop meaning much. A live opponent that bluffs in poker or lies about being a werewolf is much harder to game than a fixed answer key, because the model has to adapt in real time instead of recalling a memorized response.

Chess is the interesting test case here: engines have played near-perfect chess for decades, but LLMs are notorious for proposing illegal moves and losing track of the board, so this arena may end up measuring rule-following as much as strategy.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →