A new benchmark finds today's voice AI assistants still stumble badly once you leave English, especially in Korean and Mandarin.
The arXiv preprint arXiv:2609.35820, posted September 30, 2026, introduces tau-Multilingual, which extends the earlier tau-Voice benchmark to Spanish, Brazilian Portuguese, Hindi, Korean, and Mandarin. Native speakers reviewed both the generated language and the spoken output across 4,500 full-duplex calls and five voice configurations. Spanish, Portuguese, and Hindi stayed within 3.2 points of English task-completion scores, but Korean dropped 14.7 points and Mandarin fell 8.4 points. The paper also breaks down failure modes: Korean systems miss more responses outright, Mandarin systems interrupt callers more often, and both struggle with tools and named entities. Grok topped the task-completion leaderboard but scored lowest on generation quality, which is why the authors report task, interaction, and generation performance as separate scores instead of one blended number.
That separation matters more than it sounds. A single leaderboard number would have let Grok's task-completion lead paper over its weaker speech quality, and it would have hidden that Korean and Mandarin support is not a rounding-error gap but a double-digit one. For any company selling voice agents into East Asian markets, this is the first independently reviewed evidence of how much ground those systems still have to make up.
Most voice-agent benchmarks to date have been English-only, so vendors touting broad language support have had little external checking. This one gives buyers a reason to ask which languages, specifically, before they believe the pitch.