AI/ ai agents · automation · benchmarks · research

Reusable Code Policies Cut Computer-Use Agent Costs 217x

A new agent method turns recurring computer tasks into reusable code, cutting costs up to 217 times while boosting reliability over standard agents.

AI agents that automate computer tasks usually replan from scratch every time, even on jobs they have done a hundred times before.

Researchers describe what they call neuro-symbolic computer use: instead of an agent re-deriving every click for a recurring workflow, a learned policy handles it. The policy locks in the parts of a task that stay the same across runs, like ordering, variables, loops, and branches, as executable code, and hands off only the parts that depend on what is on screen, like locating a button or checking whether a step worked, to neural models. The system builds these policies through an iterative loop: it runs the policy, uses judges to diagnose where it failed, and has a coding model rewrite the broken section using an agent's own attempt to recover from that failure point. A verifier then double-checks each state-changing action before it runs live.

The team tested this on two agent benchmarks, OSWorld-Verified and ScienceBoard, using a metric called Pass^3, which only counts a task as solved if the agent completes it successfully in three straight attempts rather than getting lucky once. The learned policies posted the best Pass^3 scores in all four test settings, beating the base agent by 3.6 to 15.8 points, while cutting cost per run by 15 to 217 times and latency by 3.4 to 5.1 times. Policies trained only on synthetic task variations also transferred to the original, unseen tasks, outperforming a prior system called AutoRPA by 8.6 to 17.5 points.

That reliability-per-dollar tradeoff matters more than another leaderboard win: most computer-use demos chase flashy one-off tasks, but repetitive workflows are what businesses actually want automated. Whether policies built this way handle truly novel tasks, rather than variations on ones they have already seen, remains the open question here.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →