Audio AI models will confidently answer a question they never actually heard - and a new study found a fix for that blind spot.
Researchers tested whether Audio LLMs can judge the reliability of their own speech-to-text transcriptions by simply asking the model to self-assess. The models failed: they almost always rated their own transcriptions as reliable, even when the audio was too degraded to parse correctly. Existing fixes fared little better - speech quality predictors, the model's own generation uncertainty, and transcript-conditioned word error rate (WER) estimation all gave weak signals for catching failures. The breakthrough came from looking inside the model instead of asking it: transcription reliability turned out to be strongly encoded in the audio encoder's internal representations.
Using that discovery, the team built a lightweight reliability predictor that reads those internal representations and classifies a query as reliable or not before the model even generates a response. Unreliable queries can trigger a clarification request instead of a wrong answer, without retraining or modifying the underlying Audio LLM. The predictor hit 81.10% macro-F1 in-domain and 78.09% cross-domain, beating the best existing baselines by more than 10 points in both settings, and its reliability labels even transferred across different Audio LLM families.
It's not a model that hears better - it's a model that finally knows when it didn't, which for now might be the more useful skill.