Researchers have built a benchmark that actually forces AI to listen and watch at once, and most systems still flunk it.
The new benchmark, OmniReasoningBench, poses 1,150 multiple-choice and open-ended questions split across two tasks: reasoning over video and reasoning beyond it. Both audio and visual evidence are required to answer correctly, which the researchers say most existing benchmarks don't enforce. To train for it, they built a data engine called OmniQA that auto-generates evidence-grounded question-answer pairs with time-stamped clue chains, producing two training sets (OmniReasoning-SFT-112K and OmniReasoning-RL-19K) plus a new training method, Modality-Factored Self-Distillation, that scores a model's reasoning credit separately for audio clues, visual clues, and their interaction. The resulting model, OmniReasoning-30B-A3B, scores 50.0% on OmniVideoBench, a 12.8 point gain over its base model Qwen3-Omni-30B-A3B-Thinking, and 42.5% on OmniReasoningBench, a 9.3 point gain. It also improves on general and long-video benchmarks like Video-MME-v2.
Most models marketed as "omni-modal" process audio and video on largely separate tracks and only fuse them loosely at the language stage, so true cross-modal reasoning rarely gets tested, let alone trained for directly. By building a benchmark that can't be solved from vision or audio alone, this work exposes exactly how shallow that fusion has been and offers a concrete training lever, modality-factored credit assignment, for closing the gap.
A 9-point gain on a benchmark where the model still gets more than half the questions wrong is progress, not a breakthrough - a useful reminder that stacking modalities onto a model is easier than teaching it to actually reason across them.