[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-reusable-code-policies-cut-computer-use-agent-costs-217x":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},8495,"reusable-code-policies-cut-computer-use-agent-costs-217x","Reusable Code Policies Cut Computer-Use Agent Costs 217x","A new agent method turns recurring computer tasks into reusable code, cutting costs up to 217 times while boosting reliability over standard agents.","AI agents that automate computer tasks usually replan from scratch every time, even on jobs they have done a hundred times before.\n\nResearchers describe what they call neuro-symbolic computer use: instead of an agent re-deriving every click for a recurring workflow, a learned policy handles it. The policy locks in the parts of a task that stay the same across runs, like ordering, variables, loops, and branches, as executable code, and hands off only the parts that depend on what is on screen, like locating a button or checking whether a step worked, to neural models. The system builds these policies through an iterative loop: it runs the policy, uses judges to diagnose where it failed, and has a coding model rewrite the broken section using an agent's own attempt to recover from that failure point. A verifier then double-checks each state-changing action before it runs live.\n\nThe team tested this on two agent benchmarks, OSWorld-Verified and ScienceBoard, using a metric called Pass^3, which only counts a task as solved if the agent completes it successfully in three straight attempts rather than getting lucky once. The learned policies posted the best Pass^3 scores in all four test settings, beating the base agent by 3.6 to 15.8 points, while cutting cost per run by 15 to 217 times and latency by 3.4 to 5.1 times. Policies trained only on synthetic task variations also transferred to the original, unseen tasks, outperforming a prior system called AutoRPA by 8.6 to 17.5 points.\n\nThat reliability-per-dollar tradeoff matters more than another leaderboard win: most computer-use demos chase flashy one-off tasks, but repetitive workflows are what businesses actually want automated. Whether policies built this way handle truly novel tasks, rather than variations on ones they have already seen, remains the open question here.","[\"ai agents\",\"automation\",\"benchmarks\",\"research\"]","2026-09-30T04:00:00.000Z","2026-09-30T07:16:29.093Z","2026-09-30T07:16:35.463Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Explain what the Pass^3 metric actually measures (e.g., how many benchmark attempts must succeed) before citing scores on it, since readers can't interpret 'best Pass^3' or the point gains without knowing what the metric represents.","resolved","ai",[32,33,34,35],"ai agents","automation","benchmarks","research",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.36927",0,{"sections":42},[43,46,50,54,59,64,69,74,79,84,89,94,99,104],{"name":44,"slug":30,"count":45,"latest_published_at":18},"AI",5028,{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",780,{"name":51,"slug":52,"count":53,"latest_published_at":18},"Policy","policy",417,{"name":55,"slug":56,"count":57,"latest_published_at":58},"Deals","deals",284,"2026-09-29T21:00:00.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":63},"Hardware","hardware",194,"2026-09-29T13:16:04.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":68},"Science","science",154,"2026-09-28T13:19:18.000Z",{"name":70,"slug":71,"count":72,"latest_published_at":73},"Consumer Tech","consumer-tech",142,"2026-09-29T18:38:03.000Z",{"name":75,"slug":76,"count":77,"latest_published_at":78},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":80,"slug":81,"count":82,"latest_published_at":83},"Dev Tools","dev-tools",89,"2026-09-29T17:15:00.000Z",{"name":85,"slug":86,"count":87,"latest_published_at":88},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":90,"slug":91,"count":92,"latest_published_at":93},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":95,"slug":96,"count":97,"latest_published_at":98},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":100,"slug":101,"count":102,"latest_published_at":103},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":105,"slug":106,"count":107,"latest_published_at":108},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]