An inversion study finds that embeddings from popular audio AI models hold onto more of the original song than their compact size would suggest.
Researchers tested four pretrained audio encoders built for different jobs: VGGish and ConvNeXt, which classify sound; CLAP, which matches audio to text; and EnCodec, which compresses waveforms. They fed each model's frozen embeddings into a shared Stable Audio Open latent diffusion decoder and asked it to rebuild five-second, 44.1-kHz stereo clips from the Million Song Dataset. Reconstruction quality varied by encoder family, and models that exposed finer temporal or spectral detail in their embeddings produced closer rebuilds. Even the most compressed, task-specific embeddings still supported reconstructions that captured identifiable source characteristics and high-level musical content, like melody and instrumentation.
That is a problem for anyone who assumed embeddings are a safe, anonymized stand-in for raw audio. Recommendation systems, content-ID tools, and audio search products all lean on these representations, often treating them as abstract fingerprints rather than compressed copies of the source.
Image and text embeddings have faced similar inversion attacks for years; audio is just catching up to the same uncomfortable lesson - "compressed" is not the same as "unreadable".