[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-ai-security-judges-miss-attacks-they-rate-as-safe":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},8729,"ai-security-judges-miss-attacks-they-rate-as-safe","AI Security Judges Miss Attacks They Rate as Safe","An arXiv paper testing four AI judges used to screen agent security risks finds their confidence often outpaces their actual reliability.","Four AI models built to judge whether an AI agent's input is dangerous turn out to be confidently wrong more often than their scorecards suggest.\n\nA paper posted to arXiv in September 2026, \"Evaluating System One Models for Agent Security Decisions: Reliability, Calibration, and Selective Automation\" (arXiv:2609.33401v2), tested four so-called System One judges - Jev, Laya, Decider, and Bespoke Nimble - against specialized classifiers and other language-model judges. These systems are meant to triage agent interactions: flag prompt injections, score risk, and decide whether to allow, block, or send a request for human review. The paper found that strong overall accuracy and good aggregate calibration can mask failures clustered in specific attack types, including some attacks the models rated \"safe\" with high confidence. Fine-tuned versions of the models did not consistently outperform their base versions, and under the strictest error tolerances tested, the systems automated very few decisions - mostly by blocking more, not by allowing more.\n\nAgent-security setups increasingly lean on exactly these kinds of automated judges to decide what an AI agent can do without a human in the loop. The paper's finding that passing a validation check does not guarantee a model holds its error limits on new test data is the real warning: a judge that looks safe in testing can quietly fail on attacks it has never seen, which are precisely the ones an adversary would use.\n\nStacking judges together is not a clean fix either - the paper notes they catch different attacks from one another but also share each other's high-confidence blind spots, so adding more automated reviewers is not the tidy solution it sounds like.","[\"ai-security\",\"llm-evaluation\",\"prompt-injection\",\"agent-safety\"]","2026-09-30T04:00:00.000Z","2026-09-30T22:18:27.841Z","2026-09-30T22:18:33.845Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Attribute the findings to their actual source — name arXiv as the publication and cite the paper's title\u002FarXiv ID (arXiv:2609.33401v2, posted September 2026) instead of the vague 'a new study,' since none of the claims are currently traceable to a named source or publication.","resolved","ai",[32,33,34,35],"ai-security","llm-evaluation","prompt-injection","agent-safety",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.33401",0,{"sections":42},[43,47,51,55,60,65,69,74,79,83,88,93,98,103],{"name":44,"slug":30,"count":45,"latest_published_at":46},"AI",5567,"2026-10-01T04:00:00.000Z",{"name":48,"slug":49,"count":50,"latest_published_at":46},"Security","security",815,{"name":52,"slug":53,"count":54,"latest_published_at":46},"Policy","policy",430,{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",298,"2026-09-30T21:00:26.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":46},"Science","science",163,{"name":70,"slug":71,"count":72,"latest_published_at":73},"Consumer Tech","consumer-tech",149,"2026-09-30T22:57:11.000Z",{"name":75,"slug":76,"count":77,"latest_published_at":78},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":80,"slug":81,"count":77,"latest_published_at":82},"Software","software","2026-09-30T21:41:11.000Z",{"name":84,"slug":85,"count":86,"latest_published_at":87},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":89,"slug":90,"count":91,"latest_published_at":92},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":94,"slug":95,"count":96,"latest_published_at":97},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":99,"slug":100,"count":101,"latest_published_at":102},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":104,"slug":105,"count":106,"latest_published_at":107},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]