[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-model-improves-audio-visual-reasoning-still-fails-half-the-time":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},9149,"new-model-improves-audio-visual-reasoning-still-fails-half-the-time","New Model Improves Audio-Visual Reasoning, Still Fails Half the Time","A new open-source model lifts audio-visual joint reasoning scores by 9.3 points on a benchmark it still gets wrong more than half the time.","Researchers have built a benchmark that actually forces AI to listen and watch at once, and most systems still flunk it.\n\nThe new benchmark, OmniReasoningBench, poses 1,150 multiple-choice and open-ended questions split across two tasks: reasoning over video and reasoning beyond it. Both audio and visual evidence are required to answer correctly, which the researchers say most existing benchmarks don't enforce. To train for it, they built a data engine called OmniQA that auto-generates evidence-grounded question-answer pairs with time-stamped clue chains, producing two training sets (OmniReasoning-SFT-112K and OmniReasoning-RL-19K) plus a new training method, Modality-Factored Self-Distillation, that scores a model's reasoning credit separately for audio clues, visual clues, and their interaction. The resulting model, OmniReasoning-30B-A3B, scores 50.0% on OmniVideoBench, a 12.8 point gain over its base model Qwen3-Omni-30B-A3B-Thinking, and 42.5% on OmniReasoningBench, a 9.3 point gain. It also improves on general and long-video benchmarks like Video-MME-v2.\n\nMost models marketed as \"omni-modal\" process audio and video on largely separate tracks and only fuse them loosely at the language stage, so true cross-modal reasoning rarely gets tested, let alone trained for directly. By building a benchmark that can't be solved from vision or audio alone, this work exposes exactly how shallow that fusion has been and offers a concrete training lever, modality-factored credit assignment, for closing the gap.\n\nA 9-point gain on a benchmark where the model still gets more than half the questions wrong is progress, not a breakthrough - a useful reminder that stacking modalities onto a model is easier than teaching it to actually reason across them.","[\"ai\",\"multimodal\",\"benchmarks\",\"research\"]","2026-10-01T04:00:00.000Z","2026-10-01T22:34:02.204Z","2026-10-01T22:34:07.414Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Fix the dek: it says the score rises by 'double digits' while also citing the benchmark where the model 'still errs more than half the time' (OmniReasoningBench, 42.5%), but that benchmark's actual gain per the body is 9.3 percentage points, not a double-digit jump (only the separate OmniVideoBench gain of 12.8 points qualifies) — rewrite the dek so the magnitude claim matches the figure for the same benchmark it's describing.","resolved","ai",[30,32,33,34],"multimodal","benchmarks","research",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.39490",0,{"sections":41},[42,45,49,53,58,63,67,72,77,81,86,91,96,101],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",5572,{"name":46,"slug":47,"count":48,"latest_published_at":18},"Security","security",815,{"name":50,"slug":51,"count":52,"latest_published_at":18},"Policy","policy",430,{"name":54,"slug":55,"count":56,"latest_published_at":57},"Deals","deals",298,"2026-09-30T21:00:26.000Z",{"name":59,"slug":60,"count":61,"latest_published_at":62},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":64,"slug":65,"count":66,"latest_published_at":18},"Science","science",163,{"name":68,"slug":69,"count":70,"latest_published_at":71},"Consumer Tech","consumer-tech",149,"2026-09-30T22:57:11.000Z",{"name":73,"slug":74,"count":75,"latest_published_at":76},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":78,"slug":79,"count":75,"latest_published_at":80},"Software","software","2026-09-30T21:41:11.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":85},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":102,"slug":103,"count":104,"latest_published_at":105},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]