A web-browsing AI agent jumped from last place to second, and the fix was plumbing, not reasoning.
WebFovea, a vision-based web agent, placed second in the WebRetriever Challenge 2026 with a score of 57.0 out of 100 on Protocol III of the WebRetriever benchmark. The task: start at an entry URL, operate a live website's own interface, and return a verifiable answer. The team used the same underlying multimodal model across all four of its submissions, then went hunting for failures in the code connecting model to browser. They found a coordinate-space bug that sent every click to three-quarters of its intended position, silent failures on native dropdowns, iframes, and text boxes, and leftover chat-template tokens that contaminated 4.9% of task episodes.
That harness work, not a model swap, pushed the team's official score from 31.0 to 57.0 across submissions. It is a useful corrective for anyone crediting benchmark jumps to a smarter model: a capable LLM is necessary here but clearly not sufficient, and most of the visible failures traced back to parsing, execution, and reporting bugs rather than bad reasoning.
Going from 31.0 to 57.0 is a real gain, about 1.8x, not the doubling it might sound like, and it still leaves WebFovea well short of a perfect run on live, unpredictable websites.