[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-study-diagnoses-why-security-ai-agents-fail-early":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},5818,"new-study-diagnoses-why-security-ai-agents-fail-early","New Study Diagnoses Why Security AI Agents Fail Early","A new diagnostic method finds security AI agents fail before reaching the skill being tested, and a fix that helped one model backfired on a newer one.","A new study argues that when AI agents fail at multi-step security tasks, we've been measuring the wrong thing.\n\nResearchers built a diagnostic method that inserts checkpoints into long-horizon security tasks, splitting failures that happen before an agent ever reaches the skill being tested from failures at the skill itself. They tested it across four task types: delayed reuse of discovered information, reuse of previously observed state, recovery after a failed strategy, and decisions made after an uncertain outcome. Running the state-reuse task on Gemini 2.5 Flash, they found many failures happened before the model ever gathered the state it needed later, not because it misused the information but because it never collected it. Adding targeted guidance about protocol disambiguation raised how often the model actually observed that state from 65.5% to 95.4% across a 92-seed study.\n\nRerunning the same test on a newer model the paper labels 'Gemini 3.7 Flash' reversed the result: the guidance now hurt observation rates, and whether the model observed the state stopped predicting whether it finished the task at all. That's the finding worth sitting with. The bottleneck that explains a failure in one model generation isn't guaranteed to explain it in the next, which means benchmarks that report only pass or fail are grading the wrong variable.\n\nOne flag for readers: 'Gemini 3.7 Flash' does not match Google's public Gemini numbering, which has moved in whole and half steps (1.0, 1.5, 2.0, 2.5), so treat that name as the paper's own label rather than a confirmed Google release until it's verified.","[\"ai agents\",\"llm security\",\"benchmarking\",\"gemini\"]","2026-08-24T04:00:00.000Z","2026-08-24T05:56:26.010Z","2026-08-24T05:56:37.931Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"\"Gemini 3.7 Flash\" cannot be verified as a real Google model — it breaks the established Gemini versioning pattern (1.0, 1.5, 2.0, 2.5) — so confirm the model name or explicitly attribute it as reported by the source paper before publishing.","resolved","ai",[32,33,34,35],"ai agents","llm security","benchmarking","gemini",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.20563",0,{"sections":42},[43,47,51,56,61,66,71,76,81,86,91,96,101,106],{"name":44,"slug":30,"count":45,"latest_published_at":46},"AI",3325,"2026-08-24T09:09:31.000Z",{"name":48,"slug":49,"count":50,"latest_published_at":18},"Security","security",461,{"name":52,"slug":53,"count":54,"latest_published_at":55},"Policy","policy",218,"2026-08-23T19:30:00.000Z",{"name":57,"slug":58,"count":59,"latest_published_at":60},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":62,"slug":63,"count":64,"latest_published_at":65},"Hardware","hardware",145,"2026-08-22T21:25:33.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":70},"Consumer Tech","consumer-tech",95,"2026-08-18T16:05:00.000Z",{"name":72,"slug":73,"count":74,"latest_published_at":75},"Science","science",91,"2026-08-20T10:01:48.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Software","software",73,"2026-08-18T07:51:50.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Dev Tools","dev-tools",69,"2026-08-18T04:00:00.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"Startups","startups",50,"2026-08-22T16:23:09.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"Gaming","gaming",41,"2026-07-09T04:00:00.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"General","general",33,"2026-08-18T22:18:13.000Z",{"name":102,"slug":103,"count":104,"latest_published_at":105},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":107,"slug":108,"count":109,"latest_published_at":110},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]