A new benchmark puts a hard number on how often AI models fail under jailbreak pressure.
MLCommons released version 1.0 of its Jailbreak Benchmark, a standardized method for testing how eight open-weight AI systems handle attempts to bypass their safety guardrails. Researchers ran 264 seed prompts across eleven hazard categories, pairing a baseline test with an adversarial one that used attacks drawn from the MLCommons Jailbreak Taxonomy. Every response was graded against the AILuminate Assessment Standard. Across all systems and attacks, the unsafe-response rate rose from 11.08% at baseline to 18.65% under jailbreak conditions, a 7.57 point gap the benchmark calls the Resilience Gap.
That gap is the real headline. It is not that models are wildly unsafe by default, it is that phrasing a request as an attack measurably increases the odds of an unsafe answer. The benchmark also found more accessible systems, the kind built to be easy to run or fine-tune, showed a wider Resilience Gap than their more locked-down counterparts, which tracks with the long-standing tradeoff between openness and guardrail durability.
Call it a report card, not a verdict: a reproducible 7.57 point gap is a useful baseline for comparing models, but it also means every one of them still gets worse once someone tries.