A new benchmark says voice AI agents have a consistency problem: they either get the answer right or handle the conversation well, rarely both.
Researchers built EVA-Bench, an open-source framework for testing voice agents across 213 scenarios in three enterprise domains. Unlike typical text-based chatbot tests, it simulates actual bot-to-bot audio conversations and checks its own simulations for errors before scoring, regenerating a conversation if the simulated user messes up. It scores systems on two fronts: EVA-A for accuracy, EVA-X for overall experience. Across 12 systems spanning all three major voice-agent architectures, not one cleared 0.5 on both metrics at once.
That gap matters because most voice-agent demos are cherry-picked best-case runs. EVA-Bench's multi-trial testing found a median 0.44 gap between a system's peak performance and what it reliably delivers, meaning the smooth demo you saw once is not the call your customer will get on a bad day. Add background noise or an unexpected accent, and performance drops further, by as much as 0.314 depending on the system and architecture.
Enterprises have been rushing to put voice agents on customer-support lines and internal tools on the strength of flashy demos. This benchmark is a reminder that a single good take says little about whether the thing holds up on trial ten, or with a caller who has a regional accent and a bad phone connection.