AI/ ai · text-to-speech · voice-synthesis · machine-learning

A New Way to Make AI Narrators Sound Less Robotic

Researchers built a prosody-control method that lets text-to-speech systems adjust pitch, energy, and pace clause by clause for more natural narration.

A new text-to-speech system lets writers control exactly how much a voice's pitch, energy, and pace shift from one clause to the next, not just across a whole sentence or paragraph.

Researchers behind a system called SCIC, short for Scope- and Codebook-Aware Instruction Conditioning, built it on a method they call Speaker-Relative Inline Prosody Control. Instead of tagging a whole script with one tone, writers can mark individual clauses with pitch, energy, or speed instructions measured relative to the clause before them, plus separate pause-length tags. The team studied how Alibaba's Qwen3-TTS model encodes audio and found energy information clusters in the earliest layers of its sound codebooks while pitch spreads across a deeper set of layers, so SCIC weights those layers differently depending on which trait it is adjusting. They then fine-tuned the model with GDPO, short for Group-based Direct Preference Optimization, a training step that scores the model's outputs against several goals at once, including how closely the speech follows the clause-level instructions, how clear the words come out, and how closely the voice still matches the target speaker, then adjusts the model to satisfy all of them together instead of trading one off against another.

This targets a real gap. Most instruction-driven TTS tools apply one style setting to an entire clip, which is fine for a 10-second ad read but flat across a 20-minute livestream or audiobook chapter. By making prosody relative and clause-level, SCIC lets a narrator's voice build emphasis, ease off, and pause the way a human host would over a long monologue, without hand-tuning every line, which is the actual use case driving commercial interest in expressive TTS: live shopping streams, AI podcast hosts, and audiobook narration.

The results so far live in a paper and a demo page. There's no word on whether this reaches a commercial voice product, and 'sounds more expressive in our listening tests' is a claim every TTS paper makes before the model meets a real live audience.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →