A new AI benchmark shows only one agent setup can out-learn human-written game bots, and only for half the games.
Researchers built AAArena, a testbed of 12 adversarial games stocked with 1,920 archived human-written programs, and ran it like a real competitive ladder. Agents had to read each game's rules, pick opponents, study match replays, and rewrite their own strategy code within a fixed budget of matches and evaluation time, all without changing the underlying model's weights. Among the model and harness configurations tested, only Opus 5.5 running inside Claude Code won gold medals, taking first place in 6 of the 12 ladders. Every other configuration evaluated won zero golds, and no one, including that top setup, topped the remaining 6 ladders.
That split is the real result here, not the win count. It separates the question of whether a model can learn a strategy from whether it can write working code, and the answer is: sometimes, and only for the easier half of the field. The researchers also found performance fell as a game's rules got more complex, and that agents learned faster with richer feedback and when they could choose easier opponents to face first.
A 6-of-12 record against programs humans coded by hand is a real result, but it's a long way from the clean sweep a passing glance at the leaderboard might suggest.