[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-security-gateway-drops-ai-agent-attack-rates-to-9-15":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},5821,"security-gateway-drops-ai-agent-attack-rates-to-9-15","Security Gateway Drops AI Agent Attack Rates to 9-15%","ClawSentry cuts a broad agent-security benchmark's attack success rate to 9-15%, though a narrower single test scored under 3%.","A new open-source gateway called ClawSentry tries to stop malicious skills from hijacking AI coding agents before they can do damage.\n\nClawSentry sits in front of agent runtimes like Codex, Claude Code, Kimi CLI, and Gemini CLI without needing changes to those tools themselves. It vets skill packages before first use, then screens actions in real time through a cheap deterministic filter and a rule-based semantic check, escalating only genuinely ambiguous cases to a slower, read-only review agent. It also tracks attackers who get blocked once and then retry through a different tool or reworded prompt later in the same session. Across five agents on the broader SkillsSafety benchmark, this cut attack success rates from an unprotected 33.5-49.7% down to 9.09-15.03%, while legitimate tasks still completed successfully 98.7% of the time. A narrower test against SkillInject attacks on Codex\u002FGPT-5.4 pushed attack success down further, to 2.61% from 39.55%, but that figure applies to one specific attack-and-model pairing, not the broader results.\n\nMost agent guardrails check a single moment, like permissions at install or filtering on one prompt, and miss attacks that wait a turn or swap tools. Spending review effort only on ambiguous cases, while nearly always letting clean tasks through untouched, is the more interesting design choice here. Single-digit-to-low-teens attack success on the wider benchmark is real progress, but it is confinement, not elimination.\n\nWhether the gap between the eye-catching sub-3% number and the broader 9-15% range closes under more adversarial testing will decide if this becomes standard agent middleware or another benchmark-flattering paper with a GitHub repo attached.","[\"ai-security\",\"llm-agents\",\"open-source\",\"benchmarks\"]","2026-08-24T04:00:00.000Z","2026-08-24T06:33:30.148Z","2026-08-24T06:33:42.062Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The headline's 'Under 3%' claim only holds for the narrow SkillInject\u002FCodex-GPT-5.4 test (2.61%); the broader five-agent SkillsSafety benchmark reported in the same body shows 9.09-15.03%, so the headline cherry-picks the best number — qualify the headline\u002Fdek to specify which benchmark the under-3% figure applies to, or lead with the broader range.","resolved","security",[32,33,34,35],"ai-security","llm-agents","open-source","benchmarks",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.21101",0,{"sections":42},[43,48,51,56,61,66,71,76,81,86,91,96,101,106],{"name":44,"slug":45,"count":46,"latest_published_at":47},"AI","ai",3325,"2026-08-24T09:09:31.000Z",{"name":49,"slug":30,"count":50,"latest_published_at":18},"Security",461,{"name":52,"slug":53,"count":54,"latest_published_at":55},"Policy","policy",218,"2026-08-23T19:30:00.000Z",{"name":57,"slug":58,"count":59,"latest_published_at":60},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":62,"slug":63,"count":64,"latest_published_at":65},"Hardware","hardware",145,"2026-08-22T21:25:33.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":70},"Consumer Tech","consumer-tech",95,"2026-08-18T16:05:00.000Z",{"name":72,"slug":73,"count":74,"latest_published_at":75},"Science","science",91,"2026-08-20T10:01:48.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Software","software",73,"2026-08-18T07:51:50.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Dev Tools","dev-tools",69,"2026-08-18T04:00:00.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"Startups","startups",50,"2026-08-22T16:23:09.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"Gaming","gaming",41,"2026-07-09T04:00:00.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"General","general",33,"2026-08-18T22:18:13.000Z",{"name":102,"slug":103,"count":104,"latest_published_at":105},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":107,"slug":108,"count":109,"latest_published_at":110},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]