AI/ ai · video-generation · neuroscience · research

Video Generating AI Matches Brain Better When Predicting Ahead

A study finds AI video generators predicting future frames resemble human visual cortex activity more than models reconstructing what was seen.

AI models that generate future video frames line up with human brain activity better than models that just reconstruct what was already shown.

Researchers compared fMRI recordings of the visual cortex, taken while people watched video, against the internal representations of two video diffusion models: an autoregressive (AR) model and the non-AR base model it was built on. Representations used to generate upcoming frames matched the visual cortex more closely than representations used to reconstruct frames people had already seen, both within the AR model and when comparing it to its base model. The match for reconstructing observed video clustered in lower-order visual cortex. The match for generating future frames clustered in higher-order visual cortex.

That split matters because the brain doesn't just register what's in front of it, it predicts what's coming next, and this study suggests higher-order visual processing is better explained by prediction than by perception. A behavioral test backed that up: when researchers boosted the contributions of model layers that aligned most closely with the visual cortex, people preferred the resulting videos. That's a real, if narrow, link between how closely a model's internals match brain activity and whether humans find its output more appealing.

It's one study built on two diffusion models, not a general theory of vision, so the leap from better fMRI alignment to AI understands how brains work is exactly the kind of inference a skeptical neuroscientist would want replicated first.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →