Security/ ai safety · llm agents · security research · benchmarks

New Monitor Cuts Multi-Turn AI Agent Attack Success From 84% to 25%

DART, a new runtime monitor, tracks shifts in an AI agent's internal representations to catch multi-step attacks that single-action checks miss.

A new monitoring layer can tell when an AI agent's plans are drifting toward something dangerous, even if each individual step looked fine.

Researchers describe DART, a runtime defense detailed in an October 2 arXiv paper, that watches how an AI agent's internal representations shift as a conversation unfolds. Rather than judging single actions or states in isolation, it tracks the accumulated pattern across turns and flags the specific context that triggered a shift, then intervenes with a targeted reminder instead of just shutting the agent down. Tested across six models on two multi-turn benchmarks, DART cut attack success on MT-AgentRisk from 84% to 25%, with a 12% false-alarm rate and an 8% hit to benign non-refusal. On the tougher ASEval benchmark, it brought attack success down from 97% to 52% with no cost to benign requests. A key technical fix (denoising the safety signal so ordinary conversational drift does not get mistaken for an attack) nearly tripled detection on ASEval, from a 7%-40% range up to 60%-85%.

This matters because agentic AI systems are increasingly built to chain tool calls and browse, book, and purchase on a user's behalf, and the attacks that worry security researchers rarely look harmful at any single step. DART also beat ToolShield, previously the best-performing multi-turn defense, on the MT-AgentRisk benchmark across all six tested models, and it runs cheap: 0.14 to 0.56 seconds of overhead per step, with no extra model required.

Even with that win, a 52% attack success rate on ASEval is a reminder that this is a research benchmark result, not a solved problem. Half the attacks still got through on the harder test.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →