AI/ ai-agents · computer-use-agents · training-data · benchmarks

Synthetic Desktop Data Makes Computer-Use Agents Click Better

A new synthetic desktop training environment called DeskForge sharply improved an AI agent's click accuracy and its ability to finish real multi-step tasks.

A research team has built a synthetic desktop factory to train AI agents on exactly where to click.

The project, called DeskForge, runs real applications inside a controllable environment that varies window layout, screen resolution, and app content to generate dense, labeled snapshots of what is on screen. Each snapshot fuses a screenshot, the accessibility tree, and window geometry, then records the outcome of every action taken. That pipeline produced DeskForge-1M, a dataset of 1.2 million annotated desktop observations covering 159.7 million individual on-screen elements. The team fine-tuned four vision-language models on 200,000 examples from that data; one of them, Qwen3.5-4B, gained 11.51 percentage points on the ScreenSpot-Pro benchmark and 10.11 points on OSWorld-G.

The more interesting number is what happens to actual task completion. Using the exact same planner, the fine-tuned Qwen3.5-4B solved 50 of 119 WebArena-Infinity tasks, up from 31, and 15 of 100 OpenApps tasks, up from 3. Nothing about the agent's reasoning changed, only its eyes - a sign that grounding, not planning, has been the quieter bottleneck for computer-use agents that still fumble basic clicks outside curated demos.

Still, 15 out of 100 is a pass rate no product could ship on. Better synthetic training data narrows the gap; it does not close it.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →