A research team has built a bilingual voice AI that can hold a coherent group conversation for over half an hour, tracking who said what and who's supposed to answer.
The project, called MultiTalk, extends the open-source Moshi framework for full-duplex speech models - systems that listen and talk at the same time, the way people actually do. The team released 57,600 hours of synthetic training data across two sets, MultiTalkPT and MultiTalkFT, built specifically for long, multi-party, English-Chinese conversations with overlapping speech, interruptions, and shifting addressees. They also built MultiTalkBench, a benchmark drawn from real human recordings averaging 32.6 minutes each, testing whether a model can track entities, follow topics, and figure out who's being spoken to. On top of that data, they trained a bilingual Moshi-style model aimed at both languages at once.
Most open-source speech-to-speech models are still tuned for short, two-person exchanges - a far cry from a meeting, a classroom, or a lobby robot fielding several people at once. MultiTalk is a direct attempt to close that gap, and the benchmark gives other researchers a way to test long-context, multi-speaker performance that didn't really exist in open form before now.
The paper says the new model "substantially outperforms" open-source baselines Moshi, MiniCPM-o-4.5, and Qwen3-Omni-30B-A3B-Instruct on MultiTalkBench, but it does not publish the actual score deltas or define the metric behind that claim. That's worth flagging, since the benchmark is the team's own creation - a common setup in AI papers, and one that calls for independent verification rather than a benchmark grading its own homework. The datasets and benchmark are already posted on Hugging Face, so outside labs can test the claim themselves once they get around to it.