AI-generated video still struggles with a basic trick: remembering what a scene looked like two seconds ago. A new training method called LoGo aims to patch that, at least for the growing class of models that let users steer the camera through a generated scene.
Researchers built LoGo as a post-training step that reshapes how these video models get graded during training. Instead of handing out one overall quality score for an entire clip, as prior techniques do, LoGo pairs a global reward that keeps camera movement and image quality on track with a local, spatially-specific reward that catches when individual objects drift, warp, or vanish as the camera moves. Tested on three different base models against the DL3DV dataset and a new benchmark called TrajectoryBench, built for long, camera-heavy generations, LoGo cut down on object shifting, visual artifacts, and scenes that quietly rearrange themselves mid-shot.
The underlying problem is credit assignment: a single scalar score tells a model to "do better" without saying where it went wrong, which works fine for short clips but falls apart over longer, more complex camera paths. As video models push toward longer generations and more ambitious camera control, that coarseness becomes the bottleneck, not raw generation quality.
It's a sensible fix for a narrow failure mode, not a cure for AI video's broader consistency problems - and like most benchmark wins, it's worth watching whether it holds up outside DL3DV and TrajectoryBench.