[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-a-bilingual-ai-voice-model-learns-to-follow-group-conversations":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},8608,"a-bilingual-ai-voice-model-learns-to-follow-group-conversations","A Bilingual AI Voice Model Learns to Follow Group Conversations","Researchers built a bilingual full-duplex voice model plus a new dataset and benchmark to test how well AI handles long, multi-speaker conversations.","A research team has built a bilingual voice AI that can hold a coherent group conversation for over half an hour, tracking who said what and who's supposed to answer.\n\nThe project, called MultiTalk, extends the open-source Moshi framework for full-duplex speech models - systems that listen and talk at the same time, the way people actually do. The team released 57,600 hours of synthetic training data across two sets, MultiTalkPT and MultiTalkFT, built specifically for long, multi-party, English-Chinese conversations with overlapping speech, interruptions, and shifting addressees. They also built MultiTalkBench, a benchmark drawn from real human recordings averaging 32.6 minutes each, testing whether a model can track entities, follow topics, and figure out who's being spoken to. On top of that data, they trained a bilingual Moshi-style model aimed at both languages at once.\n\nMost open-source speech-to-speech models are still tuned for short, two-person exchanges - a far cry from a meeting, a classroom, or a lobby robot fielding several people at once. MultiTalk is a direct attempt to close that gap, and the benchmark gives other researchers a way to test long-context, multi-speaker performance that didn't really exist in open form before now.\n\nThe paper says the new model \"substantially outperforms\" open-source baselines Moshi, MiniCPM-o-4.5, and Qwen3-Omni-30B-A3B-Instruct on MultiTalkBench, but it does not publish the actual score deltas or define the metric behind that claim. That's worth flagging, since the benchmark is the team's own creation - a common setup in AI papers, and one that calls for independent verification rather than a benchmark grading its own homework. The datasets and benchmark are already posted on Hugging Face, so outside labs can test the claim themselves once they get around to it.","[\"ai\",\"voice ai\",\"open-source\",\"benchmarks\"]","2026-09-30T04:00:00.000Z","2026-09-30T14:37:30.696Z","2026-09-30T14:37:36.625Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The benchmark claim ('outperformed... including Moshi, MiniCPM-o-4.5, and Qwen3-Omni-30B-A3B-Instruct') cites no actual comparison figures or metric definition beyond the source's vague 'substantially outperforms' — either pull specific scores or explicitly note none were disclosed — and the single-sentence caveat-only closing paragraph needs to be expanded into a real concluding thought.","resolved","ai",[30,32,33,34],"voice ai","open-source","benchmarks",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.36903",0,{"sections":41},[42,45,49,53,58,63,68,73,78,82,87,92,97,102],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",5135,{"name":46,"slug":47,"count":48,"latest_published_at":18},"Security","security",788,{"name":50,"slug":51,"count":52,"latest_published_at":18},"Policy","policy",417,{"name":54,"slug":55,"count":56,"latest_published_at":57},"Deals","deals",284,"2026-09-29T21:00:00.000Z",{"name":59,"slug":60,"count":61,"latest_published_at":62},"Hardware","hardware",194,"2026-09-29T13:16:04.000Z",{"name":64,"slug":65,"count":66,"latest_published_at":67},"Science","science",154,"2026-09-28T13:19:18.000Z",{"name":69,"slug":70,"count":71,"latest_published_at":72},"Consumer Tech","consumer-tech",142,"2026-09-29T18:38:03.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":18},"Dev Tools","dev-tools",90,{"name":83,"slug":84,"count":85,"latest_published_at":86},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]