AI/ ai-agents · llm-security · prompt-injection · tool-use

New Defense Aims to Stop LLM Agents From Leaking Data via Tools

A new system called PACE checks every tool call an AI agent makes against its actual permissions, not just its intentions, before letting it run.

Researchers have built a gatekeeper for AI agents that checks what a tool call actually does, not just whether it looked safe going in.

The system, called Provenance-Aware Capability Enforcement (PACE), targets a specific problem with LLM agents that can browse the web, call APIs, or use plug-in "skills." Those agents pull in tool metadata, web pages, and stored memory that can be poisoned to hijack later actions. Researchers found that screening an artifact once, before an agent uses it, doesn't work: a safe input and a malicious one can produce identical-looking evidence at admission time, so a single check can't tell them apart. PACE instead intercepts every tool call right before execution, confining which influence paths are allowed and verifying the call's effects against the authority granted in the original request. Tested across eight benchmarks and three model families, the approach produced the lowest attack success rate in 62 of 79 test columns, with native task performance dropping at most three points compared to an undefended agent.

This matters because most agent-security work still treats prompt injection as an input-filtering problem: scan the text, flag the bad stuff, move on. PACE's bet is that filtering text is the wrong layer entirely, since a cleverly worded attack and an innocent request can look identical on paper. Checking the actual tool call against what the agent was authorized to do is a narrower, more mechanical guarantee, and that narrowness is exactly why it held up against adaptive attacks that found zero successful exploits in 30 attempts.

It's not a universal fix. The approach only governs what happens at the tool-call boundary, and the researchers are upfront that this is "the last boundary a deployment can still act on" -- not the first line of defense. Expect this kind of runtime authority-checking to show up in agent frameworks well before anyone solves the upstream poisoning problem it's working around.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →