AI/ ai-security · llm-evaluation · prompt-injection · agent-safety

AI Security Judges Miss Attacks They Rate as Safe

An arXiv paper testing four AI judges used to screen agent security risks finds their confidence often outpaces their actual reliability.

Four AI models built to judge whether an AI agent's input is dangerous turn out to be confidently wrong more often than their scorecards suggest.

A paper posted to arXiv in September 2026, "Evaluating System One Models for Agent Security Decisions: Reliability, Calibration, and Selective Automation" (arXiv:2609.33401v2), tested four so-called System One judges - Jev, Laya, Decider, and Bespoke Nimble - against specialized classifiers and other language-model judges. These systems are meant to triage agent interactions: flag prompt injections, score risk, and decide whether to allow, block, or send a request for human review. The paper found that strong overall accuracy and good aggregate calibration can mask failures clustered in specific attack types, including some attacks the models rated "safe" with high confidence. Fine-tuned versions of the models did not consistently outperform their base versions, and under the strictest error tolerances tested, the systems automated very few decisions - mostly by blocking more, not by allowing more.

Agent-security setups increasingly lean on exactly these kinds of automated judges to decide what an AI agent can do without a human in the loop. The paper's finding that passing a validation check does not guarantee a model holds its error limits on new test data is the real warning: a judge that looks safe in testing can quietly fail on attacks it has never seen, which are precisely the ones an adversary would use.

Stacking judges together is not a clean fix either - the paper notes they catch different attacks from one another but also share each other's high-confidence blind spots, so adding more automated reviewers is not the tidy solution it sounds like.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →