AI/ ai safety · benchmarks · prompt injection · llm research

AI Safety Router Benchmarks Collapse Under Distribution Shift

A new study shows the main benchmark for AI safety routers picks its own comparison baseline, and the advantage nearly vanishes once test conditions shift.

A new paper argues that the leaderboards used to judge AI safety routers have been grading on a curve the routers wrote themselves.

Safety routers send each incoming request to whichever model in a lineup looks safest, then get scored against the best single model available. The researchers found that the major benchmark for this task picks its comparison baseline using the same test data the router is scored on, a shortcut that holds up fine until conditions change. Rerun on held-out categories instead of random splits, that bias explains most of the benefit credited to routing on HELM Safety, and it balloons seven to nine times on the AgentDojo benchmark. Across seven other safety test sets picked in advance, only three of the routing claims survived a preregistered statistical test, and in most cases the router ended up dispatching to the same model an honest baseline would have used anyway.

This matters because a safety router that barely beats a single well-chosen model isn't buying safety, it's buying the appearance of it, while still adding latency and cost. The researchers also found that an attacker who knows which model they're facing can suppress GPT-5.4's acknowledgment of a late prompt injection by nearly 20 points, weakening a defense many systems lean on to catch attacks as they happen.

In other words, the industry has been benchmarking the seatbelt on a car that was already parked.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →