AI/ ai · llm evaluation · compliance · ai safety

Study Finds AI Compliance Models Often Ignore the Rules They Cite

A new study found AI compliance models often reach the same verdict even when the regulatory rule is deleted or reversed.

AI systems built to enforce rules often don't seem to care which rule they're given.

A new study tested five large language models across 20 regulatory and platform-policy domains. Researchers held the facts of each case fixed while swapping, deleting, or negating the governing rule, then checked whether the model's verdict changed or whether its internal representation of "compliant" shifted at all. Across the board, neither moved much. A specialized guard model built for exactly this kind of task, evaluated on an adapted version of its own taxonomy, turned out to be the least rule-sensitive and least accurate of the five, scoring just 51% (barely above a coin flip) while general-purpose models scored 90-92%.

The failure isn't blanket neglect. When deleting the rule actually changes what a model would otherwise get right, the model does track that change closely. The real problem is subtler: high accuracy on compliance questions doesn't prove a model is reasoning from the rule it was handed, and neither better prompting nor direct tinkering with the model's internal representations closed the gap.

That's an inconvenient finding for the compliance-tooling pitch that leans on accuracy as the headline number. Accuracy, it turns out, tells you almost nothing about whether the system read the rulebook or just memorized the shape of the case.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →