Speech-recognition AI still can't reliably parse crackly, overlapping police radio traffic - even when researchers let it train on its own guesses.
A new arXiv paper tests whether "pseudo-labeling" can help general-purpose speech-recognition systems, Whisper and Qwen3-ASR, handle Broadcast Police Communication audio from Baltimore and Chicago. Since human-transcribed police audio is scarce and expensive, the researchers had the models generate their own transcripts, then used those noisy guesses to retrain the systems. The usual quality check - the model's internal confidence scores - couldn't distinguish good transcripts from garbage in this domain. The team's fix was to add a separate large language model as a judge, screening out transcripts that don't make contextual sense, and to try training two different models on each other's outputs.
This matters beyond audio-engineering trivia: police radio transcripts are a growing input for research and oversight into how officers make decisions in the field, and unreliable automated transcription could distort that record before a human ever reviews it. The LLM-judge filter did cut errors more aggressively than the built-in confidence scores, which is a real gain. But the researchers found a persistent gap between their best filtered results and what a hand-checked "oracle" filter would produce.
In other words: self-training got the models closer, not close enough. Anyone using off-the-shelf transcription on messy real-world audio, not just police radio, should still budget for a human in the loop.