AI/ ai-safety · benchmarks · llm-evaluation · research

A Popular AI Safety Score Measures Less Than It Claims

A new audit finds the harmful refusal score in a popular AI safety benchmark blends multiple behaviors into one misleading number.

A widely cited AI safety benchmark may be measuring a lot less than it claims.

Researchers ran a construct validity audit of HELM Safety, the benchmark suite labs use to score how often models refuse dangerous or policy violating prompts. They started with four datasets inside HELM Safety that could plausibly target this harmful refusal trait, but found three of them saturated, meaning most models already score at the ceiling and the numbers stop being informative. That left one dataset, HarmBench, which they ran through two statistical tests: a multidimensional item response theory model and a differential item functioning analysis. The first test found that HarmBench does not track one single trait at all. The second found cases where models from different developers with identical refusal scores answered individual test items differently, though most of that gap faded once questions were grouped by topic.

That matters because HarmBench's score does not stand alone. HELM Safety folds it together with other datasets into one top-line safety number, the kind of figure that ends up in leaderboards, model cards, and procurement decisions. If the underlying dataset already blends unrelated harm behaviors into a single score, averaging it with other datasets just adds another layer of noise on top of noise.

None of this means HELM Safety is useless, but it is a reminder that a safety score is a statistical claim before it is a headline. Treating a single number as proof a model is safer is a bit like judging a car's crash safety from its top speed: related, maybe, but not the same thing.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →