AI/ voice-ai · benchmarks · customer-service-ai · llm-evaluation

Voice Agent Benchmark Finds Most Fail Basic Billing Calls

A new benchmark pits fourteen voice AI systems against simulated utility billing calls, and most fail more often than they succeed.

A new benchmark just showed that most voice AI agents can't handle a basic utility billing call.

Researchers built VAmoS Energy, a benchmark of 100 simulated calls about utility billing and payment assistance, based on Pennsylvania's residential billing rules and public household electricity data. Each caller makes two to four requests, and the agent has to work with sixteen tools linked to a stateful Stripe billing twin and the Apache Fineract loan engine, with account access locked until the caller verifies their identity. An LLM-as-a-verifier checks what the agent says and does against the task requirements, agreeing with a code-based verifier on 99.1% of checks in testing. Across fourteen voice stacks tested three times each, task completion ranged from just 17.3% to 44.7% - Grok Voice came out on top, with Gemini 3.8 Live and GPT-Live trailing at roughly the same cost per call.

The benchmark's real finding isn't the leaderboard - it's that background noise wrecks these systems. Adding background television dropped pooled completion from 38.7% to 8.6%, and simulated callers often accepted wrong answers simply because they could hear what the agent said but had no way to verify what it actually changed in the backend. That gap between a confident voice and a correct action is exactly what polished demos tend to hide.

Put plainly: if a billing bot sounds reassuring on a call, that's not the same as it getting your bill right.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →