AI/ ai · audio-ai · benchmark · reasoning

New Benchmark Shows Top Audio AI Still Can't Reason Well

ReasonAudio tests whether AI can reason about sound, not just match it, and top models are barely passing while humans nearly ace it.

A new benchmark says the best audio AI system still can't reason about sound the way humans do.

Researchers built ReasonAudio, a benchmark that tests whether text-audio retrieval systems can reason about sound rather than just match keywords to noises. It checks four skills: negation, temporal order, which sounds happen together, and how long they last. The test set spans five synthetic subtasks with 1,000 queries over 10,000 composite audio clips, plus a natural subtask with 100 queries over 1,000 real recordings. Across 11 state-of-the-art retrieval systems, the top performer, OmniEmbed-7B, scored just 20.7 out of 100 on the overall benchmark.

That's a rough result for a category of AI often marketed as understanding context, not just recognizing what a dog bark or a car horn sounds like. In a separate, more controlled test designed to strip out simple sound matching, OmniEmbed-7B reached 53.8% accuracy, well behind the 95.6% humans managed on the same task.

OmniEmbed-7B's own generative backbone, Qwen2.5-Omni-7B-Thinker, scored 70.6% on that same controlled test - beating the retrieval system built on top of it, and a reminder that bolting reasoning onto retrieval doesn't guarantee the reasoning survives the trip.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →