AI/ ai · benchmarks · roblox · agentic-coding

OpenGameEval tests AI coding agents inside Roblox Studio

A new benchmark finds top AI coding agents solve barely half of Roblox Studio tasks on a first try, and exploration habits predict success.

The best AI coding agents still fail roughly half of basic Roblox Studio tasks on their first attempt, according to a new benchmark.

Researchers built OpenGameEval, a framework that runs language models as coding agents inside live Roblox Studio sessions and grades them with automated checks on both the edited scene and a simulated play session. They tested 13 frontier models against 84 human-curated tasks, giving each model 16 attempts per task. The top-performing model solved 51.7% of tasks on a single try, but that dropped to 39.4% when it had to succeed five times out of five. Six tasks went unsolved by every model tested.

The benchmark's useful trick is splitting "look around" tools from "make a change" tools, so researchers can measure exploration separately from final success. That split shows models that inspected every object a reference solution touched passed up to 13.4 percentage points more often than models that skipped the look-around step entirely. The failures, in other words, often start before any code gets written.

Most agentic coding leaderboards only care whether a model reaches the right answer, however it gets there. OpenGameEval's exploration-tracking approach is a reminder that knowing where to look may matter as much as knowing what to type.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →