[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-a-validation-check-keeps-ai-agent-upgrades-from-backfiring":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},8619,"a-validation-check-keeps-ai-agent-upgrades-from-backfiring","A Validation Check Keeps AI Agent Upgrades From Backfiring","VACE alternates AI model training with harness tweaks, testing each change before keeping it, and beats simpler baselines on two benchmarks.","A new training method alternates between upgrading an AI agent's model and upgrading its toolkit, checking each change before it's allowed to stick.\n\nResearchers built a system called VACE (Validation-Gated Alternating Co-Evolution) that trains AI agents in two interleaved steps: reinforcement learning on the model's weights, then a rewrite of the \"harness,\" the scaffolding of prompts and tools that tells the model how to do its job. After each round of training, the system reuses the same trajectories to draft a harness revision, then tests the old and new harness against each other with the model frozen. Only if the new harness scores better on a validation set does it get used in the next training round. Tested on a 9-billion-parameter Qwen3.5 model, the approach hit 45.26% accuracy on a benchmark called OfficeQA and a 75.19% partial-credit score on AutomationBench, beating training the model's weights alone by roughly 6 to 9 percentage points.\n\nThe bottleneck this targets is real: improve the model without touching its toolkit, and gains are capped by whatever the harness lets the model do. Change the harness without a check, and the paper's own data shows the risk: 17 of 44 proposed revisions would have hurt validation performance if adopted unfiltered. The gating step is doing much of the work here, a cheap sanity check before letting an automated change touch a production training loop.\n\nVACE reads as one data point in a wider shift toward treating an agent's tools as trainable alongside its weights, instead of fixed scaffolding. Whether the idea holds up beyond one 9B model and two benchmarks depends on someone testing it on larger models, a broader set of tasks, and against rival co-evolution methods it hasn't yet been measured against.","[\"ai agents\",\"reinforcement learning\",\"ai research\",\"benchmarks\"]","2026-09-30T04:00:00.000Z","2026-09-30T15:19:48.421Z","2026-09-30T15:19:54.158Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The closing paragraph is caveat-only ('treat as a lab result, not a guarantee') — replace it with a stronger closer that adds context, such as what would need to be shown (more models\u002Fbenchmarks) to validate the approach, or how this fits into the broader trend of co-adapting agents and their tooling.","resolved","ai",[32,33,34,35],"ai agents","reinforcement learning","ai research","benchmarks",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.37105",0,{"sections":42},[43,46,50,54,59,64,69,74,79,83,88,93,98,103],{"name":44,"slug":30,"count":45,"latest_published_at":18},"AI",5135,{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",788,{"name":51,"slug":52,"count":53,"latest_published_at":18},"Policy","policy",417,{"name":55,"slug":56,"count":57,"latest_published_at":58},"Deals","deals",284,"2026-09-29T21:00:00.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":63},"Hardware","hardware",194,"2026-09-29T13:16:04.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":68},"Science","science",154,"2026-09-28T13:19:18.000Z",{"name":70,"slug":71,"count":72,"latest_published_at":73},"Consumer Tech","consumer-tech",142,"2026-09-29T18:38:03.000Z",{"name":75,"slug":76,"count":77,"latest_published_at":78},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":80,"slug":81,"count":82,"latest_published_at":18},"Dev Tools","dev-tools",90,{"name":84,"slug":85,"count":86,"latest_published_at":87},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":89,"slug":90,"count":91,"latest_published_at":92},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":94,"slug":95,"count":96,"latest_published_at":97},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":99,"slug":100,"count":101,"latest_published_at":102},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":104,"slug":105,"count":106,"latest_published_at":107},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]