AI/ speech recognition · self-supervised learning · khoisan languages · ai

AI Speech Models Learn to Recognize Click Consonants

Researchers fine-tuned two speech recognition models on rare click sounds from Khoisan languages and found they handled clicks better than ordinary consonants.

Speech recognition models trained mostly on English and other high-resource languages can still learn to spot click consonants, a speech sound family most AI systems never encounter.

Researchers fine-tuned two popular self-supervised speech models, Wav2Vec2 and HuBERT, on recordings from G|ui and West !Xoon, two Khoisan languages that use click consonants as regular parts of words. Both models are pretrained on huge amounts of audio, most of it from languages that don't use clicks at all. After fine-tuning on the click-language data, the models recognized click sounds more accurately than they recognized non-click consonants in the same languages. That held across both models and both languages tested.

It's a useful data point against the assumption that speech AI simply can't generalize past the sounds baked into its training data. If self-supervised pretraining lets a model pick up an entirely unfamiliar phoneme family this well, the bottleneck for supporting more of the world's languages may be fine-tuning data, not model architecture.

Clicks are an extreme case - the sounds don't appear in English, Mandarin, or most of the languages that dominate training sets - which makes this a decent stress test, even if it's just two languages and two models.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →