[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-filter-cuts-ai-agent-hijack-attacks-from-85-to-2":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},4955,"new-filter-cuts-ai-agent-hijack-attacks-from-85-to-2","New Filter Cuts AI Agent Hijack Attacks From 85% to 2%","PIPES, a new provenance-checking filter for AI agent tool data, cut attack success from 84.7% to 2.3% without hurting normal performance.","A new filter promises to stop rogue data from hijacking AI agents - and slashed attack success rates by roughly 37x in lab tests.\n\nResearchers describe PIPES (Provenance-Informed, Prior-Enforced Screening) in a paper posted to arXiv on August 14, 2026. The system checks every piece of data a tool-using AI agent pulls in - search results, API responses, file contents - against two things: whether the format matches what that field is supposed to contain, and whether the source is trusted enough to make the claim it's making. Content that breaks either rule gets flagged for removal, a warning, a block, or escalation to a human. Tested against adaptive, PAIR-style jailbreak attacks across six benchmark splits (three from VitaBench, three from AgentDyn), PIPES cut average attack success from 84.7 percent down to 2.3 percent, while performance on legitimate tasks barely moved - 92.5 percent with the defense versus 90.6 percent without.\n\nThe interesting part isn't the numbers, it's the target. So-called state-corruption attacks don't trick a model with clever wording; they smuggle false claims about the environment into a tool's response, so the agent's next action looks justified even to guardrails watching for bad behavior. That's a harder problem than ordinary prompt injection, because the defense has to police what data is allowed to say, not just what the model does with it - closer to email provenance checks like SPF and DKIM than to a content filter.\n\nOne flag for readers: the paper names its target agent as \"Gemma 4 31B IT,\" an identifier that does not correspond to any publicly confirmed Google model release at that name or parameter count. Treat that detail as unverified until the authors or Google clarify it.","[\"ai-security\",\"ai-agents\",\"prompt-injection\",\"arxiv-research\"]","2026-08-14T04:00:00.000Z","2026-08-14T21:10:24.632Z","2026-08-14T21:10:35.904Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The draft states 'Google's Gemma 4 31B IT' as a plain fact, but no Gemma 4 release or 31B parameter size is a known, verifiable Google model — flag this identifier for source verification (or note it as unverified) before publishing rather than repeating it uncritically from the arXiv abstract.","resolved","security",[32,33,34,35],"ai-security","ai-agents","prompt-injection","arxiv-research",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.12789",0,{"sections":42},[43,48,51,56,61,66,71,76,81,86,91,96,101,106],{"name":44,"slug":45,"count":46,"latest_published_at":47},"AI","ai",3293,"2026-08-20T04:00:00.000Z",{"name":49,"slug":30,"count":50,"latest_published_at":47},"Security",435,{"name":52,"slug":53,"count":54,"latest_published_at":55},"Policy","policy",210,"2026-08-19T09:32:27.000Z",{"name":57,"slug":58,"count":59,"latest_published_at":60},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":62,"slug":63,"count":64,"latest_published_at":65},"Hardware","hardware",140,"2026-08-19T18:25:42.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":70},"Consumer Tech","consumer-tech",95,"2026-08-18T16:05:00.000Z",{"name":72,"slug":73,"count":74,"latest_published_at":75},"Science","science",90,"2026-08-19T18:41:02.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Software","software",73,"2026-08-18T07:51:50.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Dev Tools","dev-tools",69,"2026-08-18T04:00:00.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"Startups","startups",47,"2026-08-19T19:13:46.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"Gaming","gaming",41,"2026-07-09T04:00:00.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"General","general",33,"2026-08-18T22:18:13.000Z",{"name":102,"slug":103,"count":104,"latest_published_at":105},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":107,"slug":108,"count":109,"latest_published_at":110},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]