A new paper proposes using small language models to double check whether an AI agent's tool calls actually make sense for the task at hand, not just whether they are technically permitted.
Researchers built a dataset of multi-tool tasks that require pulling from several Model Context Protocol (MCP) servers, the standard that lets AI agents reach external tools and data sources. They trained small language models to act as a relevance classifier, checking each tool call an agent makes against the task it was assigned and flagging ones that do not fit. The team tested three training approaches - prompt optimization, supervised fine-tuning, and reinforcement learning via a method called GRPO - to get the models better at catching mismatches. The point is a lightweight checker that can run on-premises and fast enough to review every call an agent makes, not just a sample.
Most agent security today only asks whether an agent is allowed to call a given tool, which says nothing about whether that call actually serves the task. An agent, or someone deliberately steering one, could stay fully within its permissions while still wandering off-task, pulling data it does not need or chaining calls toward a different goal. As companies connect more agents to more tools through protocols like MCP, that gap between permitted and appropriate looks like the more interesting attack surface.
It is a sensible idea, but it mostly pushes the trust problem down a level: now you have to trust a small model's read on intent, and nothing here shows yet how that judgment holds up once it leaves a curated research dataset.