Coding agents can build a game. Building the game you actually asked for is a different story.
Researchers have released A2Z GameSpec-Bench, a benchmark of 100 long-form Game Design Documents meant to test whether AI coding agents can faithfully implement detailed specs rather than just producing something plausible-looking. Each GDD gets converted into a dependency-aware contract listing rules, constraints, and prerequisite relationships that must all hold together. The benchmark then checks agent output two ways: inspecting the source code directly, and running agent-generated test policies through scenario replays and adaptive playtesting. Across evaluations, current agents struggled to satisfy interdependent requirements consistently across both code and actual play.
This matters because most coding-agent benchmarks test short, self-contained prompts, which is not how real software gets specified. Game design documents are a useful stress test precisely because they are long, cross-referential, and demand that visual rendering, game logic, and player interaction all agree with each other. If agents can't hold that many interlocking constraints at once in a game, the same failure mode likely shows up in any sufficiently detailed app spec, from checkout flows to onboarding logic.
The one encouraging number: requirement-specific feedback improved spec fidelity by 10.9% over two rounds of self-revision, compared to agents just trying to fix their own mistakes blind. That is a small win dressed up as a benchmark paper, and it is also the real finding. Generic self-correction has limits; agents do better when someone tells them exactly which rule they broke. Anyone pitching "fully autonomous app generation" should read the fine print before the demo video.