AI/ ai · agentic-ai · tool-calling · salesforce

Salesforce's Koa Model Beats GPT-4.1 at Enterprise Tool Calling

Salesforce fine-tuned an open-weight Nemotron model with reinforcement learning on synthetic CRM tasks, and it now beats GPT-4.1 at enterprise tool calling.

Salesforce built a language model whose only job is picking the right CRM button and calling it correctly.

The company created Koa by taking Nvidia's open-weight Nemotron-3-Super-120B model and post-training it with reinforcement learning using Group Relative Policy Optimization, trained only on public and synthetically generated data. The core technique, which Salesforce calls specification-driven task construction, converts declarative task specs into multi-turn, persona-based scenarios where the model is rewarded only for successfully invoking the correct tool with valid arguments. Applied to enterprise CRM specs, that same pipeline produces the in-domain training data Koa specializes on. It ships in FP8 for production, trimming inference costs for a vendor that needs this running at scale.

On Salesforce's own CRMAgentBench, Koa scores an 87% task success rate, ahead of GPT-4.1's 82% and its untrained base model's 79%. It also leads or ties on nearly every metric of a human-labeled production tool-calling benchmark. A controlled comparison that holds architecture and RL recipe fixed shows the extra CRM-specific training stage, not the base model, is what drives the improvement in argument accuracy and full tool-call success.

Salesforce says Koa keeps the base model's general capability intact on public benchmarks like Tau2Bench and BFCL, and that balance matters as much as the benchmark win: a tool-calling specialist that forgets how to do anything else is useless in a general-purpose assistant, and vendors building task-specific agent models will be judged on whether they can hold that line, not just beat GPT-4.1 by five points.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →