A research team built a video-generating AI that skips the expensive teacher model most rivals rely on, and found a cheap fix for its biggest weakness.
The model is a causal, or autoregressive, video diffusion system: it generates video chunk by chunk, which suits streaming and interactive use, unlike bidirectional models that see the whole clip at once before generating. The researchers trained it starting from an image model, never touching a bidirectional video model at any stage. That is unusual: most teams boost causal models by initializing from, or distilling knowledge out of, a large bidirectional one. The catch with this leaner path: causal models trained on real past frames become so dependent on that history that at inference, when they are feeding on their own generated frames instead of real ones, their own mistakes snowball forward.
The fix, called Conditional Residual Prediction, forces the model to predict each new chunk from the current frame first, then use history only to patch in whatever the present frame could not supply. In separate controlled experiments isolating this technique, CRP nearly closed a 6.14-point VBench gap between a causal model and a bidirectional model trained under identical conditions; that figure comes from those isolated experiments, not from Optica's own benchmark run. Scaled up, the same recipe produced Optica, a 2B-parameter causal model that generates 5-second 480p video and scores 82.78 on VBench after training on roughly 15 million videos, a relatively small dataset for this class of model.
VBench numbers measure benchmark performance, not whether anyone wants to watch the output; the real test comes once outside labs get their hands on the weights.