A research team has built a synthetic desktop factory to train AI agents on exactly where to click.
The project, called DeskForge, runs real applications inside a controllable environment that varies window layout, screen resolution, and app content to generate dense, labeled snapshots of what is on screen. Each snapshot fuses a screenshot, the accessibility tree, and window geometry, then records the outcome of every action taken. That pipeline produced DeskForge-1M, a dataset of 1.2 million annotated desktop observations covering 159.7 million individual on-screen elements. The team fine-tuned four vision-language models on 200,000 examples from that data; one of them, Qwen3.5-4B, gained 11.51 percentage points on the ScreenSpot-Pro benchmark and 10.11 points on OSWorld-G.
The more interesting number is what happens to actual task completion. Using the exact same planner, the fine-tuned Qwen3.5-4B solved 50 of 119 WebArena-Infinity tasks, up from 31, and 15 of 100 OpenApps tasks, up from 3. Nothing about the agent's reasoning changed, only its eyes - a sign that grounding, not planning, has been the quieter bottleneck for computer-use agents that still fumble basic clicks outside curated demos.
Still, 15 out of 100 is a pass rate no product could ship on. Better synthetic training data narrows the gap; it does not close it.