AI video editing has turned into its own crowded subfield, and researchers have just published a map of how we got here.
The paper is a systematic review of video editing techniques built on diffusion models, the same family of AI that powers image generators like Stable Diffusion. It starts with the math and the image-editing methods that video approaches borrowed from, then sorts video-specific techniques into groups based on their shared technical roots, tracing how one approach led to the next. The authors also cover more specialized use cases, including editing videos by dragging points around and editing human figures by following a target pose. To settle arguments about which method actually works best, they introduce a new benchmark called V2VBench for head-to-head comparisons.
That benchmark is the real contribution here. Video editing papers have multiplied fast, but most compared themselves against a handful of cherry-picked baselines rather than a shared yardstick. A common benchmark means claims of "state of the art" will finally mean something consistent across papers, not just something true relative to whatever the authors chose to test against.
Surveys like this tend to show up right when a research area is young enough to still be chaotic but old enough that someone needs to clean up the mess - the authors themselves flag unresolved challenges, so treat this as a map of a half-finished territory, not a finished atlas.