GUI-HARVEST fixes AI agents that click, type, and navigate software, not by retraining the AI, but by rewriting the code that watches and controls it.
Researchers describe a system that targets the harness, the surrounding code that captures screenshots, executes clicks and keystrokes, and judges whether a task succeeded. Instead of fine-tuning the underlying vision-language model, GUI-HARVEST compares before-and-after screenshots to work out what an agent's action actually did, runs the same task repeatedly to isolate which steps cause failures, and turns recurring failure patterns into specific code edits that get tested before they're kept. On the OSWorld-Verified benchmark, this lifted Qwen3-VL-32B-Instruct's score by 12.33 points. A harness built this way also improved GPT-5's score by 13.87 percentage points on a different benchmark, WindowsAgentArena, with no further tuning done for that model.
That portability is the interesting part. Most efforts to make GUI agents more reliable chase bigger or better-trained models. This one leaves the model alone and fixes the scaffolding instead, which is cheaper and, per the results, transfers across backbones the researchers never touched directly.
The paper also reports beating two baseline approaches, Self-Harness and Meta-Harness, on the same setup, evidence the GUI-specific diagnosis is doing real work rather than repeating a generic self-improvement trick. Still, these are benchmark numbers from the team that built the tool, on test suites designed for this kind of agent, not a product running in the wild.