AI/ ai · video generation · generative ai · research

Video diffusion model closes quality gap without teacher model

A 2B-parameter causal video model skips the usual bidirectional teacher, and a simpler trick called CRP nearly erased the quality gap in controlled tests.

A research team built a video-generating AI that skips the expensive teacher model most rivals rely on, and found a cheap fix for its biggest weakness.

The model is a causal, or autoregressive, video diffusion system: it generates video chunk by chunk, which suits streaming and interactive use, unlike bidirectional models that see the whole clip at once before generating. The researchers trained it starting from an image model, never touching a bidirectional video model at any stage. That is unusual: most teams boost causal models by initializing from, or distilling knowledge out of, a large bidirectional one. The catch with this leaner path: causal models trained on real past frames become so dependent on that history that at inference, when they are feeding on their own generated frames instead of real ones, their own mistakes snowball forward.

The fix, called Conditional Residual Prediction, forces the model to predict each new chunk from the current frame first, then use history only to patch in whatever the present frame could not supply. In separate controlled experiments isolating this technique, CRP nearly closed a 6.14-point VBench gap between a causal model and a bidirectional model trained under identical conditions; that figure comes from those isolated experiments, not from Optica's own benchmark run. Scaled up, the same recipe produced Optica, a 2B-parameter causal model that generates 5-second 480p video and scores 82.78 on VBench after training on roughly 15 million videos, a relatively small dataset for this class of model.

VBench numbers measure benchmark performance, not whether anyone wants to watch the output; the real test comes once outside labs get their hands on the weights.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →