A new monitoring layer can tell when an AI agent's plans are drifting toward something dangerous, even if each individual step looked fine.
Researchers describe DART, a runtime defense detailed in an October 2 arXiv paper, that watches how an AI agent's internal representations shift as a conversation unfolds. Rather than judging single actions or states in isolation, it tracks the accumulated pattern across turns and flags the specific context that triggered a shift, then intervenes with a targeted reminder instead of just shutting the agent down. Tested across six models on two multi-turn benchmarks, DART cut attack success on MT-AgentRisk from 84% to 25%, with a 12% false-alarm rate and an 8% hit to benign non-refusal. On the tougher ASEval benchmark, it brought attack success down from 97% to 52% with no cost to benign requests. A key technical fix (denoising the safety signal so ordinary conversational drift does not get mistaken for an attack) nearly tripled detection on ASEval, from a 7%-40% range up to 60%-85%.
This matters because agentic AI systems are increasingly built to chain tool calls and browse, book, and purchase on a user's behalf, and the attacks that worry security researchers rarely look harmful at any single step. DART also beat ToolShield, previously the best-performing multi-turn defense, on the MT-AgentRisk benchmark across all six tested models, and it runs cheap: 0.14 to 0.56 seconds of overhead per step, with no extra model required.
Even with that win, a 52% attack success rate on ASEval is a reminder that this is a research benchmark result, not a solved problem. Half the attacks still got through on the harder test.