AI/ ai · multimodal-models · research · video-data

Feeding Raw Video to an LLM Boosted Its Image Skills Too

Researchers trained a small language model on uncaptioned YouTube clips and saw image and video understanding improve without hurting text performance.

A new study shows raw, uncaptioned video can make language models smarter at reading images - no labels required.

Researchers mid-trained Qwen3-1.7B, a 1.7-billion-parameter open-source language model, on raw clips from YT-Temporal-1B, a large YouTube video dataset, with no captions and no text loss. The model only learned to predict the next visual token from encoded video frames. They then ran the same image-text instruction tuning on both the mid-trained model and an untouched version, isolating the effect of the video step alone. The mid-trained model scored 2.9 points higher on average across four video benchmarks and 5.1 points higher across ten image benchmarks covering perception, documents, and charts, while text performance held steady at 48.9 versus 48.0 across 14 text benchmarks.

Most multimodal training still leans on paired image-text data or captions, both of which cost money and introduce labeling noise. This result suggests raw video - the kind that is abundant and free on the open web - can do some of that work on its own, without degrading the text skills the model already had. The researchers also found that captioning the video did not beat silent next-token prediction, undercutting the assumption that labels are necessary for this kind of gain.

It is a small-scale demonstration - one 1.7B model, one dataset - but it points at a cheaper path to multimodal skill that does not depend on armies of captioners.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →