A new paper spells out exactly how many test queries it takes to catch an AI provider that swapped out the model you paid for.
The researchers tackle third-party challenge-response identity verification, or TP-CRIV: a method that lets an outside auditor check whether a deployed model matches a reference model without getting direct access to it. The catch is that probabilistic models like large language models give different answers every time you ask, even with identical prompts. The paper works out the math connecting that randomness to how reliably a verifier can tell a genuine match from an impostor, expressed as an AUC score. It also shows how to calculate the minimum number of independent challenges and repeated responses needed to hit a target confidence level, then tests the approach on LLMs using open-ended questions.
This matters because AI buyers increasingly have to trust claims they cannot independently confirm: that an API is really running the model advertised, not a cheaper substitute swapped in after a benchmark review. A statistical budget for verification turns a vague promise into a number anyone can check, the same way audits work for cloud uptime or financial statements.
It is a modest, technical fix for a problem that is really about incentives: providers quietly downgrading models to cut costs. No amount of statistics solves that unless someone actually runs the test.