AI/ voice-ai · ai-benchmarks · hallucinations · speech-ai

New Benchmark Exposes How Voice AI Agents Botch Document Facts

A new benchmark shows voice assistants increasingly make up answers instead of admitting they don't know, as conversations and documents grow longer.

Researchers built a test to catch voice AI agents making things up when asked about documents, and almost every system failed it somewhere.

The benchmark, called DSB-DG, throws 1,636 verified question-and-answer pairs at voice agents, pulled from 50 documents across five professional fields. It checks three specific weak points: whether accuracy drops as documents get longer, whether agents forget document facts over a multi-turn conversation, and whether re-feeding context helps them recover. Researchers tested old-style cascaded systems (speech to text, then a language model, then text to speech) against newer full-duplex voice models like Gemini-Live and GPT-Realtime, plus several open-weight speech-to-speech systems. The cascaded pipeline came out on top for accuracy, with Gemini-Live and GPT-Realtime close behind. Open-weight systems fared worst, with accuracy collapsing sharply once documents got long and facts dropping out across turns.

The detail that matters most: when these systems get something wrong, they rarely say "I don't know." They just generate a confident, unsupported answer. For voice agents being pitched at call centers, insurance lines, and other document-heavy professional work, that is the difference between a system that is occasionally unhelpful and one that is occasionally convincing and wrong.

It is also a useful gut check on the current voice AI pitch: the flashiest, lowest-latency real-time models are not the most trustworthy ones. The clunkier, slower cascaded pipeline still grounds facts better, a reminder that speed and faithfulness remain a tradeoff nobody has fully solved.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →