AI/ llm-as-judge · ai regulation · benchmarks · fintech compliance

AI Judges Fail the Compliance Stress Test, Study Finds

A new benchmark finds AI compliance judges can be fooled by keyword stuffing, with accuracy collapsing on adversarial financial-promotion tests.

A new benchmark shows the AI judges reviewing financial ads for regulatory compliance can be talked into approving deceptive marketing with the right keywords.

Researchers built Principle-Bench, a set of 168 cryptoasset financial-promotion scenarios mapped to two UK Financial Conduct Authority principles: "fair, clear, and not misleading" and "deliver good outcomes." Each scenario comes in paraphrased, keyword-stuffed, and boundary-perturbed versions, built under a pre-registered rubric. The team tested several judging methods, including keyword counting, three sentence-transformer embedders, an open-weight LLM judge, and a new calibrated assessor called Ceca, across four axes: accuracy, paraphrase robustness, adversarial robustness, and calibration. No single method won on all four axes; a 120-billion-parameter LLM judge that scored best on ordinary inputs saw its accuracy fall from 0.74 to 0.27 once inputs were stuffed with Consumer Duty buzzwords.

Principle-based rules like "fair and clear" resist yes-or-no coding, which is exactly why regulators and compliance teams are starting to lean on LLMs to triage them at scale. The study found a second LLM judge, from a different model family, agreed with the first at only a Cohen's kappa of 0.16 on the adversarial split - evidence the failure sits in the specific model, not the test cases, meaning swapping in another judge will not automatically fix it.

The researchers call this failure mode "compliance theatre": a judge that looks rigorous on friendly inputs and folds the moment someone learns which words trip its filters, which is precisely the scenario a regulator should worry about before signing off on AI graders.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →