[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-technique-pushes-ai-agent-success-rate-to-80-percent":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},6288,"new-technique-pushes-ai-agent-success-rate-to-80-percent","New Technique Pushes AI Agent Success Rate to 80 Percent","A new prompt optimization method lifted AI agent task completion from 60% to 80% on held-out tasks in a brainstorming-workflow benchmark.","A new tuning method for AI agent instructions just pushed task completion from 60% to 80% on a held-out test set.\n\nResearchers built RobustSGPO, an extension of an existing technique called semantic-gradient-based prompt optimization, or SGPO, which tunes the instructions given to AI agents using feedback from their own execution attempts. The new method specifies exactly what edit to make, builds and verifies the resulting patch, then decides whether to keep testing from the current version or roll back to an earlier saved snapshot. The team evaluated the approach on an AgentX brainstorming workflow using 120 tasks across 95 runs and more than 7,300 candidate prompt edits. On a separate set of 30 held-out tasks, completion rose from 60% to 80%, and a quality score climbed from 3.77 to 4.14, all within a 20-million-token compute budget.\n\nThis is a paper about tuning agent instructions, not a new agent or model. Gains like this say less about raw capability and more about how much performance was left on the table by sloppy prompt-search methods. A disciplined search process clawed a lot of it back without touching the underlying model.\n\nThe catch: this was tested on one brainstorming workflow, not coding, customer support, or anything closer to how most people actually deploy agents, so treat the 80% figure as a lab result until someone reproduces it elsewhere.","[\"ai agents\",\"prompt optimization\",\"llm research\"]","2026-09-11T04:00:00.000Z","2026-09-11T04:39:10.593Z","2026-09-11T04:39:22.498Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Fix two inaccuracies against the source: the 60%→80% completion jump was measured on 30 held-out tasks, not the 120 tasks cited in the body, so correct or separate those figures; and the dek's claim that this is about 'AI coding agents' is unsupported — the paper evaluates a general agent harness on an AgentX brainstorming workflow, not coding agents specifically.","resolved","ai",[32,33,34],"ai agents","prompt optimization","llm research",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.09646",0,{"sections":41},[42,45,49,53,58,63,68,71,76,80,85,90,95,100],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",3508,{"name":46,"slug":47,"count":48,"latest_published_at":18},"Security","security",636,{"name":50,"slug":51,"count":52,"latest_published_at":18},"Policy","policy",338,{"name":54,"slug":55,"count":56,"latest_published_at":57},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":59,"slug":60,"count":61,"latest_published_at":62},"Hardware","hardware",153,"2026-09-09T15:12:32.000Z",{"name":64,"slug":65,"count":66,"latest_published_at":67},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":69,"slug":70,"count":66,"latest_published_at":18},"Science","science",{"name":72,"slug":73,"count":74,"latest_published_at":75},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":18},"Dev Tools","dev-tools",70,{"name":81,"slug":82,"count":83,"latest_published_at":84},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"General","general",41,"2026-09-08T01:57:23.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]