[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-monitor-cuts-multi-turn-ai-agent-attack-success-from-84-to-25":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},9626,"new-monitor-cuts-multi-turn-ai-agent-attack-success-from-84-to-25","New Monitor Cuts Multi-Turn AI Agent Attack Success From 84% to 25%","DART, a new runtime monitor, tracks shifts in an AI agent's internal representations to catch multi-step attacks that single-action checks miss.","A new monitoring layer can tell when an AI agent's plans are drifting toward something dangerous, even if each individual step looked fine.\n\nResearchers describe DART, a runtime defense detailed in an October 2 arXiv paper, that watches how an AI agent's internal representations shift as a conversation unfolds. Rather than judging single actions or states in isolation, it tracks the accumulated pattern across turns and flags the specific context that triggered a shift, then intervenes with a targeted reminder instead of just shutting the agent down. Tested across six models on two multi-turn benchmarks, DART cut attack success on MT-AgentRisk from 84% to 25%, with a 12% false-alarm rate and an 8% hit to benign non-refusal. On the tougher ASEval benchmark, it brought attack success down from 97% to 52% with no cost to benign requests. A key technical fix (denoising the safety signal so ordinary conversational drift does not get mistaken for an attack) nearly tripled detection on ASEval, from a 7%-40% range up to 60%-85%.\n\nThis matters because agentic AI systems are increasingly built to chain tool calls and browse, book, and purchase on a user's behalf, and the attacks that worry security researchers rarely look harmful at any single step. DART also beat ToolShield, previously the best-performing multi-turn defense, on the MT-AgentRisk benchmark across all six tested models, and it runs cheap: 0.14 to 0.56 seconds of overhead per step, with no extra model required.\n\nEven with that win, a 52% attack success rate on ASEval is a reminder that this is a research benchmark result, not a solved problem. Half the attacks still got through on the harder test.","[\"ai safety\",\"llm agents\",\"security research\",\"benchmarks\"]","2026-10-02T04:00:00.000Z","2026-10-03T03:42:05.340Z","2026-10-03T03:42:11.536Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"publisher-r1","publisher",1,"The body contradicts itself on the MT-AgentRisk result, claiming DART both 'cut attack success...to 25%' and 'catching every attack on the first benchmark,' which cannot both be true.","resolved","security",[32,33,34,35],"ai safety","llm agents","security research","benchmarks",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.00400",0,{"sections":42},[43,47,50,54,59,63,67,72,77,82,87,92,97,102],{"name":44,"slug":45,"count":46,"latest_published_at":18},"AI","ai",5978,{"name":48,"slug":30,"count":49,"latest_published_at":18},"Security",842,{"name":51,"slug":52,"count":53,"latest_published_at":18},"Policy","policy",438,{"name":55,"slug":56,"count":57,"latest_published_at":58},"Deals","deals",317,"2026-10-01T22:00:00.000Z",{"name":60,"slug":61,"count":62,"latest_published_at":18},"Hardware","hardware",199,{"name":64,"slug":65,"count":66,"latest_published_at":18},"Science","science",173,{"name":68,"slug":69,"count":70,"latest_published_at":71},"Consumer Tech","consumer-tech",155,"2026-10-01T19:54:10.000Z",{"name":73,"slug":74,"count":75,"latest_published_at":76},"Dev Tools","dev-tools",96,"2026-10-01T16:57:03.000Z",{"name":78,"slug":79,"count":80,"latest_published_at":81},"Software","software",93,"2026-09-30T21:41:11.000Z",{"name":83,"slug":84,"count":85,"latest_published_at":86},"Startups","startups",90,"2026-10-01T21:55:22.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"Gaming","gaming",53,"2026-10-02T02:50:39.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"How-To","how-to",7,"2026-10-01T09:00:00.000Z"]