AI models that listen to the world still get confused by background noise - a new training method tries to fix that without slowing anything down.
Researchers built EchoDistill, a post-training framework that teaches large audio language models to perform better on noisy audio by learning from how the same model handles a clean version of the same clip. A noisy-input copy of the model generates candidate answers, while a frozen copy processes the clean audio as a reference. The two are aligned through a mix of token-level distillation and consistency training, and only the noisy-trained version ships at inference, so there's no extra compute cost at runtime. Across three different LALM backbones and three audio domains at a harsh -10dB noise level, the method improved average noisy-input accuracy by 1.63 percentage points over the best existing baseline. On Qwen2.5-Omni specifically, noisy-audio accuracy rose from 59.33% to 62.94%, while clean-audio accuracy ticked up too, from 76.56% to 77.56%.
That clean-audio bump matters as much as the noise fix. Plenty of robustness tricks trade accuracy on easy inputs for gains on hard ones; this one apparently doesn't. The researchers also checked their own work by swapping in random, shuffled, or silent audio instead of the matched clean clip - accuracy dropped 3.08 to 6.42 points, which at least confirms the model is leaning on real acoustic evidence rather than pattern-matching text alone.
One caveat worth flagging: the gains show up against additive noise (the kind you simulate by mixing in a noise track) but don't reliably transfer to non-additive distortions like reverb or compression. So call this a fix for one specific, common flavor of bad audio - not a general cure for the messiness of real-world sound.