AI/ audio ai · ai research · noise robustness · self-distillation

Self-Distillation Trick Makes Audio AI More Noise-Resistant

EchoDistill trains audio language models on clean audio as a guide, boosting accuracy on noisy speech without any extra cost at inference time.

AI models that listen to the world still get confused by background noise - a new training method tries to fix that without slowing anything down.

Researchers built EchoDistill, a post-training framework that teaches large audio language models to perform better on noisy audio by learning from how the same model handles a clean version of the same clip. A noisy-input copy of the model generates candidate answers, while a frozen copy processes the clean audio as a reference. The two are aligned through a mix of token-level distillation and consistency training, and only the noisy-trained version ships at inference, so there's no extra compute cost at runtime. Across three different LALM backbones and three audio domains at a harsh -10dB noise level, the method improved average noisy-input accuracy by 1.63 percentage points over the best existing baseline. On Qwen2.5-Omni specifically, noisy-audio accuracy rose from 59.33% to 62.94%, while clean-audio accuracy ticked up too, from 76.56% to 77.56%.

That clean-audio bump matters as much as the noise fix. Plenty of robustness tricks trade accuracy on easy inputs for gains on hard ones; this one apparently doesn't. The researchers also checked their own work by swapping in random, shuffled, or silent audio instead of the matched clean clip - accuracy dropped 3.08 to 6.42 points, which at least confirms the model is leaning on real acoustic evidence rather than pattern-matching text alone.

One caveat worth flagging: the gains show up against additive noise (the kind you simulate by mixing in a noise track) but don't reliably transfer to non-additive distortions like reverb or compression. So call this a fix for one specific, common flavor of bad audio - not a general cure for the messiness of real-world sound.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →