A team of researchers has published what they call the first comprehensive survey of how to make video-generation AI actually do what it's told.
The paper surveys the growing patchwork of post-training techniques used to fix pretrained video models after the fact, rather than retraining them from scratch. It sorts methods into four buckets: supervised fine-tuning, self-training and distillation, preference- and reward-based training, and techniques applied at inference time rather than during training. The authors also catalog the datasets and benchmarks used to judge these models and flag persistent problems: errors that compound over a video's length, the tangled relationship between motion and appearance, and the difficulty of supervising something as slippery as temporal consistency.
That taxonomy matters because video generation has outpaced the tools to control it. Models have gone from grainy few-second clips to long, high-resolution sequences, but keeping that output physically plausible, safe, and faithful to a prompt is a harder problem than aligning text or image models, since mistakes can accumulate frame by frame. Unlike text generation, where reinforcement learning from human feedback is now a known playbook, video alignment is still assembling its own toolkit from borrowed parts.
The survey's own list of open challenges, including scalable reward design, long-horizon consistency, and safety-aware generation, reads less like a research agenda and more like an admission: the field can now make video that looks convincing, but it still doesn't have a reliable way to make sure that video behaves.