[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-training-framework-sharpens-ai-tool-calling-accuracy":10,"sections":34},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":29,"feedback":33,"feedback_at":22,"cost_usd":33,"total_tokens":33},6884,"new-training-framework-sharpens-ai-tool-calling-accuracy","New Training Framework Sharpens AI Tool-Calling Accuracy","MATCH pairs curriculum learning with layered rewards to fix two common flaws in how language models are trained to pick and use external tools.","A new training method aims to stop AI models from confidently calling the wrong tool - or calling the right one with garbled arguments.\n\nResearchers describe MATCH, a training framework that combines two techniques to fix known weaknesses in how large language models learn to use external tools via reinforcement learning. The first piece, Model-Aware Curriculum Learning, tracks how difficult each training example is for the model in real time and adjusts which examples it sees next, replacing older systems that use a fixed difficulty threshold that can drift out of sync with what the model can actually do. The second piece, Hierarchical Tool-call Gated Reward, scores a model's tool call in stages - first the tool name, then the argument keys, then the argument values - and only hands out credit at each stage if the one before it was correct, which stops a model from getting partial credit for well-formed arguments attached to the wrong tool. On the API-Bank and BFCL V3 benchmarks, MATCH scored 72.19% and 62.87% overall accuracy, beating supervised and other RL-based baselines, with gains holding up across four different model backbones from two model families.\n\nTool use is the part of 'agentic AI' that turns a chatbot into something that can actually book a flight, query a database, or run code, and most of the current hand-wringing about agent reliability traces back to exactly the two failure modes this paper targets: models that misjudge which tool to reach for, and models that get sloppy with arguments once they've picked one. A tighter feedback loop between how a model is scored and how its training difficulty ramps up is a sensible fix, and the fact that it holds across multiple model families suggests it isn't overfit to one architecture's quirks.\n\nStill, this is a benchmark paper, not a shipped feature - API-Bank and BFCL V3 scores describe controlled test sets, not whether an agent will handle the weird tool call your actual API throws at it.","[\"ai\",\"llm-agents\",\"reinforcement-learning\",\"tool-use\"]","2026-09-18T04:00:00.000Z","2026-09-18T21:02:28.417Z","2026-09-18T21:02:40.348Z","published",null,[],"ai",[24,26,27,28],"llm-agents","reinforcement-learning","tool-use",[30],{"name":31,"url":32},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.20082",0,{"sections":35},[36,39,43,48,53,57,61,66,70,75,80,85,90,95],{"name":37,"slug":24,"count":38,"latest_published_at":18},"AI",4066,{"name":40,"slug":41,"count":42,"latest_published_at":18},"Security","security",657,{"name":44,"slug":45,"count":46,"latest_published_at":47},"Policy","policy",338,"2026-09-11T04:00:00.000Z",{"name":49,"slug":50,"count":51,"latest_published_at":52},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":54,"slug":55,"count":56,"latest_published_at":18},"Hardware","hardware",155,{"name":58,"slug":59,"count":60,"latest_published_at":18},"Science","science",123,{"name":62,"slug":63,"count":64,"latest_published_at":65},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":18},"Dev Tools","dev-tools",78,{"name":71,"slug":72,"count":73,"latest_published_at":74},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":84},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"General","general",41,"2026-09-08T01:57:23.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]