[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-researchers-build-ai-agents-that-self-improve-over-many-rounds":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},8855,"researchers-build-ai-agents-that-self-improve-over-many-rounds","Researchers Build AI Agents That Self-Improve Over Many Rounds","A new research system trains AI agents to repeatedly critique and rewrite their own answers, and the benchmark gains keep climbing instead of leveling off.","A new agent framework called AREX-2 teaches AI models to grade and rewrite their own answers across dozens of rounds, and the gains keep compounding instead of plateauing.\n\nResearchers trained the agent on synthetic \"long-horizon improvement trajectories\" drawn from machine learning and algorithmic programming tasks, domains where right and wrong answers are easy to verify. The bet is that two skills learned there - spotting a better solution (reflection) and sticking with the grind over many iterations (long-horizon execution) - should transfer to other kinds of work. Built on Alibaba's open-weight Qwen3 model family, the resulting agent scored 81.8 on MLE-bench Lite and 70.7 on Frontier-CS, two machine-learning and coding benchmarks. It also carried over to tasks it wasn't explicitly trained on: 84.0 on BrowseComp, 52.6 on HLE (Humanity's Last Exam, a notoriously difficult general-knowledge test), 92.2 on GAIA (a benchmark of real-world assistant tasks), and 93.8 on DeepSearchQA.\n\nThe interesting part isn't the raw scores - it's that performance kept climbing as the agent got more rounds to iterate, rather than hitting a wall after a couple of tries, which is the usual failure point for self-correcting agents. Teaching reflection in verifiable domains like code and ML, then watching it transfer to messier, open-ended research tasks, backs a broader industry wager: that letting models think longer at test time, not just training bigger ones, is where the next gains come from.\n\nIt's one paper's benchmarks, not an independent audit, and self-improving agents have a habit of acing curated test sets while stumbling on messier real-world workloads - so treat the transfer claims as promising, not proven.","[\"ai agents\",\"self-improving ai\",\"llm benchmarks\",\"test-time compute\"]","2026-10-01T04:00:00.000Z","2026-10-01T07:41:51.436Z","2026-10-01T07:41:55.699Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Verify and correct the model name 'Qwen3.8-27B', which doesn't match any real Qwen3 naming convention (Qwen3 ships as 8B, 14B, 32B, etc.) and reads as a garbled or invented variant, and define benchmark acronyms like HLE and GAIA on first use.","resolved","ai",[32,33,34,35],"ai agents","self-improving ai","llm benchmarks","test-time compute",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.38288",0,{"sections":42},[43,46,51,56,61,66,71,76,81,85,90,95,100,105],{"name":44,"slug":30,"count":45,"latest_published_at":18},"AI",5270,{"name":47,"slug":48,"count":49,"latest_published_at":50},"Security","security",801,"2026-09-30T22:18:23.000Z",{"name":52,"slug":53,"count":54,"latest_published_at":55},"Policy","policy",429,"2026-10-01T02:26:17.000Z",{"name":57,"slug":58,"count":59,"latest_published_at":60},"Deals","deals",298,"2026-09-30T21:00:26.000Z",{"name":62,"slug":63,"count":64,"latest_published_at":65},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":70},"Science","science",157,"2026-09-30T15:00:56.000Z",{"name":72,"slug":73,"count":74,"latest_published_at":75},"Consumer Tech","consumer-tech",149,"2026-09-30T22:57:11.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":82,"slug":83,"count":79,"latest_published_at":84},"Software","software","2026-09-30T21:41:11.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":106,"slug":107,"count":108,"latest_published_at":109},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]