A new text-to-speech model generates realistic speech in just four steps without ever training against a slower, more accurate teacher model.
Researchers built DriftTTS, a system that converts text into a mel-spectrogram (the audio blueprint speech synthesizers turn into sound) in four quick steps. Most fast speech generators get there by compressing a slower diffusion process or by distilling knowledge from a bigger, slower teacher model. DriftTTS skips both: it trains by matching its own output distribution to the target using a frozen encoder pretrained on the same speech dataset, learning from its own in-progress outputs rather than copying a teacher's answers. Tested on the standard LJSpeech dataset, DriftTTS scored 3.87 dB on a standard audio-accuracy measure called MCD and had a 3.7% word error rate, essentially tied with the four-step version of Matcha-TTS, a leading fast speech model, at 3.85 dB and 3.4%.
In blind listening tests with real people, DriftTTS actually edged out Matcha-TTS, scoring 4.18 out of 5 versus 3.96, trailing real human recordings at 4.22 by a narrow margin. That matters because teacher-based distillation adds a training step, and a dependency, that smaller labs often cannot afford to build or maintain. If a teacher-free method can match or beat distilled systems, it lowers the barrier for building fast, good-sounding voice synthesis.
The code is public on GitHub, so the real test will be whether it holds up outside a single, well-worn dataset like LJSpeech.