A new benchmark finds that AI-written no-code bug fixes, the kind that just tell you to flip a setting instead of patching code, often don't actually fix anything.
Researchers built an automated pipeline that checks whether a proposed no-code fix (changing a setting, upgrading to a version where the bug is already fixed, or adjusting a workflow) actually resolves the reported bug, rather than relying on a developer to verify it by hand. They ran 322 no-code fixes generated by 12 different LLM configurations through a real browser: an executor agent carried out each fix's instructions, and a separate checker confirmed whether the bug persisted. Three executors took turns running the tests, OpenCUA-72B, Claude Sonnet 5, and Meta's Muse Glimmer. Resolution rates across all 322 fixes ranged from 14.6% to 49.7% depending on which executor ran them, and the strongest single combination, Claude Opus 4.6's fixes carried out by Claude Sonnet 5, cleared 74.1%.
The more telling number isn't the top score, it's the variance. Swapping only the executor shifted a configuration's resolution rate by 38.8% on average, and the three executors agreed on the same verdict for under half the fixes. That means a fix's reported success rate today says as much about which agent tested it as about which model wrote it, a shaky basis for routing automated fixes straight to users.
Even the executors' calls only matched human judgment 66.1% to 88.1% of the time, a reminder that automated verification is a useful sanity check right now, not yet a replacement referee.