AI/ ai-agents · voice-interfaces · benchmarks · llms

Talk2Agent Benchmark Measures How Speech Trips Up AI Agents

A new benchmark finds that voice transcription errors can corrupt AI agent instructions, and standard ASR scores fail to catch it.

Researchers have built a benchmark to test whether voice assistants garble the instructions they pass to AI agents.

The team introduced Talk2Agent, a benchmark that evaluates voice interfaces feeding spoken commands to LLM-based computer-use agents. It converts tasks from two existing benchmarks, WildClawBench and OSWorld, into human-spoken versions, then runs them through dedicated speech-recognition models, audio-capable LLMs, contextual biasing, and LLM-based ontology repair. Because actually running long, multi-step computer tasks repeatedly is slow and expensive, the team also built an execution-free evaluation method that checks whether task-critical details survive the voice-to-text step without running the task itself. On 32 hours of real human speech from WildClawBench, that execution-free check tracked actual task completion far better than standard word-error and character-error rate metrics, improving correlation by 0.246.

Voice is becoming a default way to control software agents, but most agent benchmarks still assume clean typed text. A transcription error that swaps a number, a file name, or a constraint can send an otherwise capable agent down the wrong path before it even starts reasoning, and the industry's go-to speech metrics were never built to catch that kind of failure.

Measuring words right has never been the same as getting the task right, and this benchmark is a reminder that voice interfaces for agents need their own yardstick, not a borrowed one from dictation software.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →