AI/ voice ai · turn-taking · full-duplex speech · speech models

AI Voice Models Sync Up, But Still Hand Off Slowly

A study pairing two instances of a voice AI model found they synchronize but hand off the conversational floor roughly three times slower than humans do.

Two AI voices held an unscripted conversation with each other, and both were slow to jump in.

Researchers paired two instances of a full-duplex speech model called PersonaPlex-7B and let them talk over a shared audio channel, then compared the timing to human conversations from the Switchboard corpus. The two models did genuinely couple to each other: swapping in a different partner for either model broke the synchronization, so each was responding to its specific counterpart, not just its own habits. But the actual handoff was sluggish. The conversational floor changed hands at a median of 400 to 560 milliseconds, versus 137 milliseconds for humans, and almost none of the swaps landed in the final 120 milliseconds of a partner's turn, the window where humans make about a tenth of their transfers. Delaying one side of the audio channel shifted the response by the same amount and left that run-up window empty.

That points to models reacting after they think someone has stopped talking, rather than anticipating the end of a turn the way people do. That distinction matters more than it sounds, because full-duplex speech models increasingly talk to each other, not just to people, in self-play data generation, simulated agent societies, and model-based evaluation. A half-second hesitation that a human conversation partner would smooth over instead becomes baked into the generated data, unchecked.

Train the next generation of voice assistants on transcripts of two hesitant models waiting each other out, and don't be surprised if the hesitation comes along for the ride.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →