[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-scoring-system-grades-voice-ai-agents-on-real-failures":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},8235,"new-scoring-system-grades-voice-ai-agents-on-real-failures","New Scoring System Grades Voice AI Agents on Real Failures","A new benchmark measures voice agent reliability by counting failed calls instead of averaging vague quality scores.","Researchers have proposed a standardized way to grade whether AI voice agents actually get the job done.\n\nThe protocol, called Inquesto Score (IS), defines voice-agent reliability as the percentage of calls in a fixed, versioned test set that achieve the caller's goal without a functional failure or worse. Instead of blending different metrics into one fuzzy number, it defines explicit failure events and severity levels. Timing failures, like talk-over and delayed responses, are measured directly from audio, while semantic and state-related failures are judged using scenario predicates, tool traces, and a pinned open-model judge. Version 0.1 of the protocol tested a reference voice-agent system across 30 scenarios, three acoustic conditions, and four speaker groups, running 306 calls per agent spread across 13 configurations.\n\nVoice agents now handle things like account verification and transactions, where a bad call has real consequences, not just an annoying chatbot loop. Most current evaluations lean on transcripts and generic accuracy scores that do not capture whether a caller actually got what they needed. IS's insistence on validating its own judges, and on keeping diagnostics like acoustic robustness and identity handling separate from the headline score, is a quiet admission that a lot of industry reliability claims are built on shaky measurement.\n\nThe team released the protocol, a reference implementation, and its evaluation records, which is more transparency than most vendors offer. Still, a benchmark built and graded by the same group that designed the reference agent is a first step, not independent proof.","[\"voice-agents\",\"ai-evaluation\",\"benchmarks\",\"reliability\"]","2026-09-28T04:00:00.000Z","2026-09-28T18:26:45.743Z","2026-09-28T18:26:52.907Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Fix the call-count figure: the source says 306 calls per agent across 13 configurations (i.e., 306 total, not per configuration) — the draft's 'a306 calls per configuration, spanning 13 setups' inflates the study's actual scale roughly 13x.","resolved","ai",[32,33,34,35],"voice-agents","ai-evaluation","benchmarks","reliability",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.30514",0,{"sections":42},[43,47,52,57,62,67,72,77,82,87,92,97,101,106],{"name":44,"slug":30,"count":45,"latest_published_at":46},"AI",4877,"2026-09-28T16:14:54.000Z",{"name":48,"slug":49,"count":50,"latest_published_at":51},"Security","security",764,"2026-09-28T12:00:00.000Z",{"name":53,"slug":54,"count":55,"latest_published_at":56},"Policy","policy",401,"2026-09-28T15:14:48.000Z",{"name":58,"slug":59,"count":60,"latest_published_at":61},"Deals","deals",268,"2026-09-28T14:00:00.000Z",{"name":63,"slug":64,"count":65,"latest_published_at":66},"Hardware","hardware",191,"2026-09-28T15:45:00.000Z",{"name":68,"slug":69,"count":70,"latest_published_at":71},"Science","science",154,"2026-09-28T13:19:18.000Z",{"name":73,"slug":74,"count":75,"latest_published_at":76},"Consumer Tech","consumer-tech",138,"2026-09-28T14:13:01.000Z",{"name":78,"slug":79,"count":80,"latest_published_at":81},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":83,"slug":84,"count":85,"latest_published_at":86},"Dev Tools","dev-tools",85,"2026-09-28T13:00:00.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"Startups","startups",78,"2026-09-28T15:29:21.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":98,"slug":99,"count":95,"latest_published_at":100},"General","general","2026-09-26T17:02:42.000Z",{"name":102,"slug":103,"count":104,"latest_published_at":105},"Reviews","reviews",30,"2026-09-24T20:07:31.000Z",{"name":107,"slug":108,"count":109,"latest_published_at":110},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]