A new benchmark says AI coding agents can write code, but building an actual playable game is a different story.
Researchers built SWE-Game, a set of 247 tasks drawn from 41 executable Godot games across 13 gameplay categories in 2D and 3D. The tasks span five types: building a game from a short brief, following a full design document, completing a skeleton, fixing 83 planted bugs, and porting a game from Godot to Unity. To check whether an agent's code actually works, the benchmark runs the games, replays certified reference inputs, and reviews agent-recorded demos of their own features, with separate vision-language rubrics scoring the game's visual presentation. Across six models tested, Opus5 topped every one of the five task categories.
On the three tasks that build something from scratch, even the best score stayed under 60 out of 100, and turning a short brief into a working game topped out at 50.38. Researchers traced most failures to mundane causes: skipped requirements and broken gameplay logic, not an inability to produce code that runs at all. The benchmark's own validation is telling too: actually running the games and checking engine state caught human-flagged bugs with 92.59 percent accuracy, versus 78.41 percent for a video-watching AI judge, a reminder that watching gameplay footage is a weaker substitute for playing the game.
Coding benchmarks love to reward passing unit tests. This one grades on whether the thing is actually fun and functional, a tougher bar, and most models still clear it poorly.