A new benchmark says the best audio AI system still can't reason about sound the way humans do.
Researchers built ReasonAudio, a benchmark that tests whether text-audio retrieval systems can reason about sound rather than just match keywords to noises. It checks four skills: negation, temporal order, which sounds happen together, and how long they last. The test set spans five synthetic subtasks with 1,000 queries over 10,000 composite audio clips, plus a natural subtask with 100 queries over 1,000 real recordings. Across 11 state-of-the-art retrieval systems, the top performer, OmniEmbed-7B, scored just 20.7 out of 100 on the overall benchmark.
That's a rough result for a category of AI often marketed as understanding context, not just recognizing what a dog bark or a car horn sounds like. In a separate, more controlled test designed to strip out simple sound matching, OmniEmbed-7B reached 53.8% accuracy, well behind the 95.6% humans managed on the same task.
OmniEmbed-7B's own generative backbone, Qwen2.5-Omni-7B-Thinker, scored 70.6% on that same controlled test - beating the retrieval system built on top of it, and a reminder that bolting reasoning onto retrieval doesn't guarantee the reasoning survives the trip.