Researchers just showed that voice anonymizers leak more than advertised once you test them outside English.
A new study built a multilingual benchmark to test how well attacker speaker-verification systems can unmask people whose voices have been run through anonymization tools. Past work on breaking voice anonymization focused almost entirely on English, so nobody really knew whether those attacks held up elsewhere. The researchers tested two attack styles: ones that match voices by raw acoustic signal, and ones that lean on the actual words being spoken. They also built a new multilingual voice-converted dataset, called MultiVC, to train and evaluate attackers across languages.
The acoustic approach won most of the time, but its lead shrank whenever the anonymized speech left enough spoken content intact to follow clearly. That is an awkward trade-off for anyone building anonymization tools: distort the voice enough to beat acoustic matching, and the speech gets harder to use for transcription or voice assistants; keep it clear, and content-based attackers get an opening. Training on the multilingual dataset helped narrow the gap between languages, though it did not close it.
Most privacy claims for voice anonymization are still benchmarked against a single attacker model tested mainly in English, which is a bit like testing a lock against one burglar who does not speak the local language.