[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-salesforces-koa-model-beats-gpt-41-at-enterprise-tool-calling":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},11146,"salesforces-koa-model-beats-gpt-41-at-enterprise-tool-calling","Salesforce's Koa Model Beats GPT-4.1 at Enterprise Tool Calling","Salesforce fine-tuned an open-weight Nemotron model with reinforcement learning on synthetic CRM tasks, and it now beats GPT-4.1 at enterprise tool calling.","Salesforce built a language model whose only job is picking the right CRM button and calling it correctly.\n\nThe company created Koa by taking Nvidia's open-weight Nemotron-3-Super-120B model and post-training it with reinforcement learning using Group Relative Policy Optimization, trained only on public and synthetically generated data. The core technique, which Salesforce calls specification-driven task construction, converts declarative task specs into multi-turn, persona-based scenarios where the model is rewarded only for successfully invoking the correct tool with valid arguments. Applied to enterprise CRM specs, that same pipeline produces the in-domain training data Koa specializes on. It ships in FP8 for production, trimming inference costs for a vendor that needs this running at scale.\n\nOn Salesforce's own CRMAgentBench, Koa scores an 87% task success rate, ahead of GPT-4.1's 82% and its untrained base model's 79%. It also leads or ties on nearly every metric of a human-labeled production tool-calling benchmark. A controlled comparison that holds architecture and RL recipe fixed shows the extra CRM-specific training stage, not the base model, is what drives the improvement in argument accuracy and full tool-call success.\n\nSalesforce says Koa keeps the base model's general capability intact on public benchmarks like Tau2Bench and BFCL, and that balance matters as much as the benchmark win: a tool-calling specialist that forgets how to do anything else is useless in a general-purpose assistant, and vendors building task-specific agent models will be judged on whether they can hold that line, not just beat GPT-4.1 by five points.","[\"ai\",\"agentic-ai\",\"tool-calling\",\"salesforce\"]","2026-10-09T04:00:00.000Z","2026-10-10T08:21:48.459Z","2026-10-10T08:21:52.762Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The closing line claims 'the paper says nothing about how Koa performs on anything outside CRM-shaped tasks,' which directly contradicts the article's own earlier statement that Koa preserves general capability on Tau2Bench and BFCL — remove or fix this self-contradictory claim.","resolved","ai",[30,32,33,34],"agentic-ai","tool-calling","salesforce",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.15066",0,{"sections":41},[42,46,51,56,61,66,71,76,81,86,91,96,101,106],{"name":43,"slug":30,"count":44,"latest_published_at":45},"AI",6841,"2026-10-09T19:36:56.000Z",{"name":47,"slug":48,"count":49,"latest_published_at":50},"Security","security",941,"2026-10-09T19:25:43.000Z",{"name":52,"slug":53,"count":54,"latest_published_at":55},"Policy","policy",490,"2026-10-09T14:25:43.000Z",{"name":57,"slug":58,"count":59,"latest_published_at":60},"Deals","deals",483,"2026-10-09T11:20:39.000Z",{"name":62,"slug":63,"count":64,"latest_published_at":65},"Hardware","hardware",233,"2026-10-09T16:07:54.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":70},"Science","science",195,"2026-10-09T11:00:57.000Z",{"name":72,"slug":73,"count":74,"latest_published_at":75},"Consumer Tech","consumer-tech",183,"2026-10-09T14:50:57.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Startups","startups",120,"2026-10-09T17:02:31.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Software","software",114,"2026-10-08T17:57:01.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"Dev Tools","dev-tools",109,"2026-10-09T13:03:48.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"General","general",69,"2026-10-09T19:10:07.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"Gaming","gaming",59,"2026-10-09T11:43:43.000Z",{"name":102,"slug":103,"count":104,"latest_published_at":105},"Reviews","reviews",34,"2026-10-08T14:00:22.000Z",{"name":107,"slug":108,"count":109,"latest_published_at":110},"How-To","how-to",8,"2026-10-05T09:00:00.000Z"]