AI/ ai-agents · benchmarks · browser-agents · computer-use

New Benchmark Shows AI Agents Fail at Finishing the Job

A new browser-agent benchmark finds top AI systems complete fewer than 3% of complex, artifact-producing tasks end to end.

A new benchmark finds that today's AI web agents can search just fine, but almost all of them botch the final product.

The benchmark, called KNOWS, tests browser-based agents on open-ended tasks that require more than fetching a fact: agents must retrieve information across multi-step workflows and synthesize it into a finished artifact, such as a document, presentation, or spreadsheet. The researchers built a task design rubric to ensure each test reflects real assistant work, and paired every task with an evaluator that combines deterministic checks with LLM judgment. They ran frontier computer-use agents and browser-based harnesses against it. The best performer scored decently on partial-success metrics but fully succeeded on fewer than 3% of the complex, long-horizon tasks.

The gap comes down to visual work. Agents that completed more than half of a task's steps still produced unusable artifacts because they stumbled on visual and spatial steps, like formatting a slide or laying out a spreadsheet. That is the part of "assistant" work that current agent benchmarks mostly skip, and it is the part that actually determines whether a human would accept the output.

So the pitch of an AI agent quietly finishing your deck while you grab coffee is, per this data, still mostly aspirational.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →