The best AI coding agents still fail roughly half of basic Roblox Studio tasks on their first attempt, according to a new benchmark.
Researchers built OpenGameEval, a framework that runs language models as coding agents inside live Roblox Studio sessions and grades them with automated checks on both the edited scene and a simulated play session. They tested 13 frontier models against 84 human-curated tasks, giving each model 16 attempts per task. The top-performing model solved 51.7% of tasks on a single try, but that dropped to 39.4% when it had to succeed five times out of five. Six tasks went unsolved by every model tested.
The benchmark's useful trick is splitting "look around" tools from "make a change" tools, so researchers can measure exploration separately from final success. That split shows models that inspected every object a reference solution touched passed up to 13.4 percentage points more often than models that skipped the look-around step entirely. The failures, in other words, often start before any code gets written.
Most agentic coding leaderboards only care whether a model reaches the right answer, however it gets there. OpenGameEval's exploration-tracking approach is a reminder that knowing where to look may matter as much as knowing what to type.