AI/ ai · voice-ai · hindi · speech-tech

First Full-Duplex Voice AI Built for Hindi Conversations

Researchers adapted a duplex speech model to Hindi using 26,000 hours of real conversations, documented so others can reproduce it.

Josh Talks says it built Human-1, the first Hindi voice AI that can listen and talk back at the same time, instead of waiting politely for you to finish a sentence.

The team started with Moshi, an existing full-duplex speech architecture built for English, and adapted it for Hindi. They swapped in a custom Hindi tokeniser and retrained the text-related parameters while keeping Moshi's pretrained audio components intact. Training ran in two stages: a large pretraining pass, then fine-tuning on 1,000 hours of conversational data. The bulk of the learning came from 26,000 hours of real, unscripted Hindi conversations recorded from 14,695 speakers on separate audio channels, which let the model learn actual turn-taking, interruptions, and overlaps rather than scripted back-and-forth.

Full-duplex systems that handle interruptions like a real conversation are still rare even in English, and Josh Talks says this is the first attempt at one for Hindi or any Indian language. That's a meaningful gap: Hindi has hundreds of millions of speakers and a growing voice-AI market, yet most conversational speech research still assumes English. The team describes the method as open and reproducible, meaning other researchers can follow the recipe, though the paper doesn't confirm that code or model weights have actually been released.

A reproducible method on paper is not the same as a product people can use, so the real test is whether anyone outside this team actually builds on it.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →