AI/ ai agents · benchmarking · computer-use ai · open-weight models

New Benchmark Measures Speed of AI Agents That Use Computers

cua-speedrun standardizes testing of computer-use AI agents and finds speed does not scale predictably with model choice or hardware upgrades.

A new benchmark says the AI agents that click and type their way through software are getting good - just not fast.

Researchers built cua-speedrun, a standardized testbed for computer-use agents, the systems that complete tasks by operating a graphical interface like a human would. Earlier benchmarks already showed these agents beating people on complex, multi-step tasks, but testing environments varied so much between labs that speed and cost numbers weren't comparable. cua-speedrun fixes that by running every agent through the same virtual machine setup, execution pipeline, and interface across four existing benchmarks. That let the team isolate how reasoning effort, the software harness wrapping each model, and network latency actually affect how fast and cheaply a task gets done.

The headline result is that no model family wins on performance, speed, and cost all at once, and every open-weight model tested trails the frontier on at least one of those measures. Two findings run against intuition: cranking up a model's reasoning effort sometimes finishes tasks faster, not slower, and speeding up the environment's input-output can actually slow the agent down overall.

That's a reminder that for agents acting on screens, raw model capability is only half the story - the surrounding plumbing matters just as much.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →