A new benchmark just caught AI coding agents failing at something simpler than it sounds: building a whole project from nothing.
Zero2Repo hands an agent a product requirements document, an interface contract spelling out how the code should connect to the rest of the system, and an empty folder. The agent has to produce a working repository in its target language: Python, TypeScript, Go, or C++ in this release. Every task comes from a real, version-pinned open-source project, converted into a spec, a reproducible environment, and a set of hidden tests that only reveal themselves after the agent submits. Those tests are checked two ways: a reference implementation has to pass them, and broken implementations have to fail them, so the grading itself gets audited before any agent touches it.
On 11 tasks built from repositories that frontier models have almost certainly already seen during training, the best agent still finished only 10. The failures weren't half-built projects; failing submissions passed 90 to 99 percent of the hidden tests. Most misses traced back to one overlooked detail or a rule mentioned once in the spec, not a missing feature.
That's a useful distinction. It suggests the bottleneck for these agents isn't writing code, it's reading carefully, and a benchmark that can tell the difference is worth more than a leaderboard score, even though these agents had the easiest version of this test since the source material was likely already in their training data.