[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-training-pipeline-sharply-boosts-ai-coding-agents-success":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},8484,"new-training-pipeline-sharply-boosts-ai-coding-agents-success","New Training Pipeline Sharply Boosts AI Coding Agents' Success","A new arXiv paper shows a training method that teaches AI coding agents when to use diagnostic tools, lifting task success rates significantly in tests.","A new training method teaches AI coding agents not just to use diagnostic tools, but to know when.\n\nResearchers behind a paper posted to arXiv this week introduce ToolMLBench, a set of executable tools for inspecting data, checking code, and diagnosing failed experiments, paired with a training pipeline called SPICE. The problem SPICE solves is subtle: an agent's decision to run a diagnostic check only pays off based on what it does with the result later, so simply rewarding successful outcomes does not tell the model which tool calls were actually useful. SPICE gets around this by measuring how access to privileged context changes the odds of a given tool call, then uses that gap as a turn-level reward during reinforcement learning. The team trained on 80 synthetic ML engineering tasks and evaluated on 25 similar tasks plus 10 unfamiliar ones.\n\nThe gains held up, and they came from better tool-use training, not a bigger model. Qwen3-8B's in-domain success rate rose from 24.8% to 52.4%, roughly doubling. Qwen3.5-35B-A3B climbed from 35.6% to 69.2% in-domain - a big jump, but shy of doubling - and its out-of-domain success rose from 31% to 48%.\n\nThat out-of-domain number is the one worth watching. Most agent benchmarks report gains on tasks that look like their training data; a jump from 31% to 48% on genuinely unfamiliar sources and targets suggests the agents picked up a transferable habit, not a memorized script. The paper also notes that simply handing a model a tool and a description of what it does produced inconsistent results on its own - the skill of knowing when to check your work had to be trained, not assumed.\n\nStill, this is 80 hand-built synthetic tasks, not the sprawling, undocumented datasets that make real ML engineering miserable. Whether that learned diagnostic instinct survives contact with an actual messy production pipeline is the question no benchmark here has answered yet.","[\"ai\",\"ai-agents\",\"machine-learning\",\"benchmarks\"]","2026-09-30T04:00:00.000Z","2026-09-30T06:35:30.305Z","2026-09-30T06:35:35.632Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Name the actual source (arXiv paper, e.g. 'a paper posted to arXiv this week') since the draft cites all its percentages with no attribution, and fix the headline\u002Fdek's 'doubles'\u002F'more than doubling' framing since the larger model's result (35.6%→69.2%) is a ~1.94x gain, not actually more than double.","resolved","ai",[30,32,33,34],"ai-agents","machine-learning","benchmarks",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.36679",0,{"sections":41},[42,45,49,53,58,63,68,73,78,83,88,93,98,103],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",5028,{"name":46,"slug":47,"count":48,"latest_published_at":18},"Security","security",780,{"name":50,"slug":51,"count":52,"latest_published_at":18},"Policy","policy",417,{"name":54,"slug":55,"count":56,"latest_published_at":57},"Deals","deals",284,"2026-09-29T21:00:00.000Z",{"name":59,"slug":60,"count":61,"latest_published_at":62},"Hardware","hardware",194,"2026-09-29T13:16:04.000Z",{"name":64,"slug":65,"count":66,"latest_published_at":67},"Science","science",154,"2026-09-28T13:19:18.000Z",{"name":69,"slug":70,"count":71,"latest_published_at":72},"Consumer Tech","consumer-tech",142,"2026-09-29T18:38:03.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":82},"Dev Tools","dev-tools",89,"2026-09-29T17:15:00.000Z",{"name":84,"slug":85,"count":86,"latest_published_at":87},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":89,"slug":90,"count":91,"latest_published_at":92},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":94,"slug":95,"count":96,"latest_published_at":97},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":99,"slug":100,"count":101,"latest_published_at":102},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":104,"slug":105,"count":106,"latest_published_at":107},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]