AI/ voice-ai · speech-recognition · llm-inference · full-duplex-audio

Voice AI Answers Speech Questions Without Transcribing First

DuplexJev skips transcription and still matches it on accuracy, while also picking up speaker tone that transcripts throw away.

A new voice AI system skips transcription entirely and still gets the right answer in about a tenth of a second.

Researchers built DuplexJev, which feeds the hidden states from a speech-recognition encoder into a frozen large language model and has it answer each question with a single token instead of generating full text. Running on an 8-GPU node, the system handles 80 separate decisions across eight spoken utterances in about 0.1 seconds. With a lightweight last-layer connector, spoken question answering lands at 90% accuracy, just one point below the 91% a model gets from reading the transcript. Swap in a more complex cross-attention connector and the model also starts detecting speaker gender (90%, up from 55%) and emotion (90%, up from 28%), with spoken QA accuracy again slipping by only one point, from 83% to 82%.

Full-duplex voice agents - the kind meant to listen and respond without waiting for a person to finish talking - make constant small judgment calls, and autoregressive decoding is too slow for every one of them. Turning those calls into single-token classification instead of generation points to voice assistants that react in real time without the usual tradeoff between speed and picking up tone, not just words. That matters because most speech-to-LLM pipelines still throw away paralinguistic information the moment they transcribe.

The team is also releasing its weights, training recipe, and a bilingual spoken QA benchmark - worth watching for whether other labs can reproduce those numbers outside a controlled 8-GPU test.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →