AI/ voice-agents · ai-evaluation · benchmarks · reliability

New Scoring System Grades Voice AI Agents on Real Failures

A new benchmark measures voice agent reliability by counting failed calls instead of averaging vague quality scores.

Researchers have proposed a standardized way to grade whether AI voice agents actually get the job done.

The protocol, called Inquesto Score (IS), defines voice-agent reliability as the percentage of calls in a fixed, versioned test set that achieve the caller's goal without a functional failure or worse. Instead of blending different metrics into one fuzzy number, it defines explicit failure events and severity levels. Timing failures, like talk-over and delayed responses, are measured directly from audio, while semantic and state-related failures are judged using scenario predicates, tool traces, and a pinned open-model judge. Version 0.1 of the protocol tested a reference voice-agent system across 30 scenarios, three acoustic conditions, and four speaker groups, running 306 calls per agent spread across 13 configurations.

Voice agents now handle things like account verification and transactions, where a bad call has real consequences, not just an annoying chatbot loop. Most current evaluations lean on transcripts and generic accuracy scores that do not capture whether a caller actually got what they needed. IS's insistence on validating its own judges, and on keeping diagnostics like acoustic robustness and identity handling separate from the headline score, is a quiet admission that a lot of industry reliability claims are built on shaky measurement.

The team released the protocol, a reference implementation, and its evaluation records, which is more transparency than most vendors offer. Still, a benchmark built and graded by the same group that designed the reference agent is a first step, not independent proof.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →