[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-method-teaches-ai-to-fix-its-own-broken-tool-calls":10,"sections":46},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":35,"tags":36,"sources":41,"feedback":45,"feedback_at":22,"cost_usd":45,"total_tokens":45},9458,"new-method-teaches-ai-to-fix-its-own-broken-tool-calls","New Method Teaches AI to Fix Its Own Broken Tool Calls","ContractRL repairs broken AI tool calls field by field instead of regenerating them, cutting tokens while beating supervised baselines on accuracy.","A new training method lets AI models repair their own malformed tool calls field by field, instead of regenerating the whole object from scratch.\n\nResearchers describe a reinforcement-learning system called ContractRL that treats a broken JSON tool call like a chart to correct, not a draft to rewrite. When a tool call fails a schema or validation check, the model doesn't start over. It receives the candidate, the specific verifier error, a pointer to the bad field, and a record of what it has already tried. A rules-based filter blocks edits that violate the contract before a deterministic validator checks the fix. Across 192 test cases and five seeds, ContractRL reached 93.62% semantic success using an average of 34.4 generated tokens, compared with 90.76% and 44.9 tokens for a supervised patch-only baseline, and 91.48% and 137.2 tokens for full object regeneration.\n\nThat gap matters more than the headline number suggests. Schema violations are one of the most common ways AI agent pipelines quietly break, and the standard fix today is to resend the whole object and hope. A cheaper, field-level, auditable repair step is unglamorous plumbing, but it is the kind of plumbing that decides whether an agent system actually holds up running thousands of tool calls a day.\n\nOne figure needs a footnote the paper itself doesn't provide. A separate ablation on the reinforcement-learning step alone reports supervised ContractRL climbing from 91.86% to 93.75% success, a number close to but not the same as the 93.62% headline result, apparently drawn from a different seed count and evaluation slice. The authors don't reconcile the two, so until this gets peer-reviewed or replicated, treat the accuracy gains as promising rather than settled.","[\"reinforcement-learning\",\"ai-agents\",\"tool-calling\",\"research\"]","2026-10-02T04:00:00.000Z","2026-10-02T20:04:48.020Z","2026-10-02T20:04:52.956Z","published",null,[24,30],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The closing paragraph's '~four percentage points, tight CI' claim is misattributed to the 192-case\u002Ffive-seed benchmark (which actually shows a 2.86-point gap, 93.6% vs 90.8%, no CI given) when that CI figure (+3.96, 95% CI [1.37,6.62]) actually comes from a separate three-seed paired evaluation — fix the attribution so the stat isn't presented with conflicting values from two different experiments.","resolved",{"id":31,"reviewer":32,"round":33,"reason":34,"status":29},"publisher-r2","publisher",2,"The numbers for the RL step are internally inconsistent — ContractRL's headline result is 93.62% (vs. 90.76% supervised baseline), but a later unlabeled sentence claims RL fine-tuning pushed success from 91.86% to 93.75%, figures that don't match either number above and aren't reconciled or explained as a different test.","ai",[37,38,39,40],"reinforcement-learning","ai-agents","tool-calling","research",[42],{"name":43,"url":44},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.00328",0,{"sections":47},[48,51,55,60,65,70,75,80,85,90,95,100,105,110],{"name":49,"slug":35,"count":50,"latest_published_at":18},"AI",5765,{"name":52,"slug":53,"count":54,"latest_published_at":18},"Security","security",831,{"name":56,"slug":57,"count":58,"latest_published_at":59},"Policy","policy",437,"2026-10-01T18:10:00.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Deals","deals",317,"2026-10-01T22:00:00.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":69},"Hardware","hardware",198,"2026-10-01T17:38:48.000Z",{"name":71,"slug":72,"count":73,"latest_published_at":74},"Science","science",168,"2026-10-01T18:35:55.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Consumer Tech","consumer-tech",155,"2026-10-01T19:54:10.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":84},"Dev Tools","dev-tools",96,"2026-10-01T16:57:03.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"Software","software",93,"2026-09-30T21:41:11.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Startups","startups",90,"2026-10-01T21:55:22.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"Gaming","gaming",53,"2026-10-02T02:50:39.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":106,"slug":107,"count":108,"latest_published_at":109},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":111,"slug":112,"count":113,"latest_published_at":114},"How-To","how-to",7,"2026-10-01T09:00:00.000Z"]