Voice AI agents get evaluated by three different communities that barely talk to each other, and none of their numbers tell you if the agent actually works.
The paper reviews 38 sources spanning three communities that rarely cite each other: speech-model architecture, turn-taking psycholinguistics, and agentic evaluation. Its central finding is that architecture is a deployment tradeoff, not a verdict. A 2026 enterprise tutorial found no fully self-hostable end-to-end system meets production needs, while a separate chunked-cascade design independently hit state-of-the-art duplex performance. Evaluation itself is also shifting, from grading how good each component sounds toward verifying what the agent actually did on the backend, instead of trusting its own account of events.
That matters because most public voice-agent demos are judged on how natural they sound, not on whether they got the task right or knew when to stay quiet. The paper's multiparty finding, that two-person conversation benchmarks cannot capture reasoning about who should hear what, is a real gap for any assistant meant to sit in a group call rather than a one-on-one demo. A shared standard like TRG, short for Timing, Recovery, Grounded outcome, would let a fully integrated model and a cascaded pipeline get compared on the same terms, something this fragmented field has been missing.
Standards proposals are easy to write and hard to get an industry to adopt, so the real test is whether the next round of voice-agent papers actually reports timing, recovery, and grounded outcomes instead of just latency and a demo reel.