AI/ multimodal ai · ai benchmarks · gpt-4o · gemini

AI Models Fail to Read Handwriting From Pen Sounds and Video

A new test asks AI to guess words from pen sounds and hand motion alone, and top models scored below 10 percent versus over 80 percent for humans.

A new benchmark just found something GPT-4o and Gemini can't do: read words from the sound and motion of someone writing them.

Researchers built The Unwritten Benchmark, which asks models to identify words written in three different styles using only audio of pen scratches and video of hand movement, with no visible ink and no view of the page itself. Human participants got the letters right, in order, more than 80% of the time. Leading multimodal models, including GPT-4o and Gemini 2.5 Pro, couldn't crack 10%. Oddly, giving the models both audio and video together often made them worse, not better, a "paradoxical fusion effect" the researchers say points to a breakdown in how these systems combine different senses of input.

This isn't a trivia gap. It suggests the multimodal reasoning that headline benchmark scores tout is mostly recognizing what's already visible or audible, not inferring information implied by how two signals unfold together over time. That distinction matters for anything sold as "multimodal understanding," from robotics to accessibility tools that depend on cross-sensory inference.

The benchmark is a reminder that impressive scores on existing tests don't prove models reason like brains do; they prove the tests were the wrong yardstick.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →