AI audit results are usually reported as one number, but that number is more fragile than it looks.
Researchers behind a new project called ASSERT built a measurement pipeline that ties every reported audit rate to a written specification of how it was produced. The tool helps draft a behavioral rubric and test cases, then runs those tests against a GenAI system to generate the rate. In a case study on conversational deception, the team found that the score shifted substantially depending on the dialogue setup, the simulated user, which judge model scored the responses, and where the evidence bar for non-compliance was set. Change any of those inputs and the reported rate moves - sometimes enough to reorder how different systems rank against each other.
That instability matters because audit numbers get used for real decisions: comparing vendors, flagging regressions, and deciding whether a model ships. A single compliance percentage implies a fixed, comparable measurement, but if swapping the judge changes who looks 'safer,' that percentage was never as fixed as it appeared. ASSERT does not resolve the underlying disagreement about what counts as compliant behavior - it just makes the measurement choices explicit enough to argue about.
Audit rates already get cited like settled facts in comparisons and safety reports - a number that comes with its own receipts is at least honest about how much of the score is the model and how much is the ruler.