AI/ ai · benchmarks · voice-ai · speech-recognition

LEGO Speech AI Holds Steady on Multi-Turn Math Tests

A new benchmark shows LEGO held steady at 77.5 percent whether a problem arrived whole or in shards, while rivals lost 5 to 25 points.

Most voice AI assistants lose accuracy when a problem is revealed piece by piece across a conversation. One proprietary system did not budge.

Researchers released SpeechConversationBench (SCB), a test built from 103 modified GSM8K math problems, to measure how speech-to-speech AI handles information that arrives across multiple turns instead of all at once. Each problem was delivered three ways: as a single spoken prompt, as the same details concatenated into one turn, and split into incremental turns resembling natural conversation. Across four commercial systems, including GPT-4o Realtime, accuracy on the incremental version fell 5.0 to 25.3 percentage points compared with the concatenated version. LEGO, a proprietary pipeline built by SCBX Innovation Lab with explicit conversational context tracking, scored 77.5 percent in all three conditions, while GPT-4o Realtime scored 76.6 percent on the incremental version.

Real conversations rarely arrive in one clean block. People add details, backtrack, and clarify mid-thought, and an assistant that forgets the first half of a request once the second half shows up fails exactly where voice interfaces are supposed to shine. SCB turns that failure mode into a measurable number instead of an anecdote about a bad demo.

A 103-problem math test from one lab is not proof LEGO has solved conversational memory, but it is a flat line next to four systems whose scores all bent downward.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →