AI/ ai safety · benchmarks · multimodal ai · llm

New Benchmark Finds Omni AI Models Miss Safety Cues

MCBench, a new 1196-scenario test covering vision, audio, and text, shows leading AI models fail to combine multimodal clues into sound safety judgments.

A new benchmark says today's omni AI models still can't reliably tell safe from unsafe when sight, sound, and text collide.

Researchers built MCBench, a test set of 1196 scenarios spanning four safety categories, designed for so-called omni large language models that process video, audio, and text together, not just images and text like earlier safety benchmarks. Each unsafe scenario is paired with an almost identical safe version, so evaluators can check whether a model is actually reasoning about risk or just pattern-matching on obvious red flags. Testing state-of-the-art omni models turned up a consistent weak spot: they do fine when a scenario has a loud visual or acoustic warning sign, but stumble on risks that are subtle or don't involve physical harm. Digging into the models' reasoning traces, the researchers found the problem isn't that models can't read each modality on its own, it's that they fail to combine what they see, hear, and read into one coherent safety judgment.

Most safety benchmarks to date have tested vision-and-text models, which quietly assumes the audio channel is safety-neutral. As voice assistants, video-calling copilots, and robots with microphones move into homes and offices, that assumption gets riskier, and MCBench's minimal-pair design makes it harder for models to fake competence by matching obvious cues instead of reasoning.

It's a reminder that stacking more modalities onto a model doesn't automatically stack up its judgment, sensory input still has to be understood together, not just detected.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →