A new benchmark finds that today's AI web agents can search just fine, but almost all of them botch the final product.
The benchmark, called KNOWS, tests browser-based agents on open-ended tasks that require more than fetching a fact: agents must retrieve information across multi-step workflows and synthesize it into a finished artifact, such as a document, presentation, or spreadsheet. The researchers built a task design rubric to ensure each test reflects real assistant work, and paired every task with an evaluator that combines deterministic checks with LLM judgment. They ran frontier computer-use agents and browser-based harnesses against it. The best performer scored decently on partial-success metrics but fully succeeded on fewer than 3% of the complex, long-horizon tasks.
The gap comes down to visual work. Agents that completed more than half of a task's steps still produced unusable artifacts because they stumbled on visual and spatial steps, like formatting a slide or laying out a spreadsheet. That is the part of "assistant" work that current agent benchmarks mostly skip, and it is the part that actually determines whether a human would accept the output.
So the pitch of an AI agent quietly finishing your deck while you grab coffee is, per this data, still mostly aspirational.