Researchers have built an AI model that edits, enhances, and understands video from a single set of plain-English instructions, in seconds rather than hours.
The system, called VIDiff, is built on diffusion models, the same technology behind modern image and video generators. The researchers unified two kinds of jobs into one foundation model: understanding tasks, like picking out an object in a video from a text description, and generative tasks, like editing or enhancing footage. Users give it instructions and get results back in seconds, skipping the lengthy per-clip tuning that typical diffusion-based editors require. To stop long videos from drifting or glitching between edits, the team added an iterative auto-regressive method meant to preserve consistency across an entire clip.
Most video-editing diffusion models are narrow and slow. They're built for short clips and need per-video tuning or lengthy inference runs before producing a result. VIDiff's pitch is consolidation: one model handles both the segmentation step and the editing step, with speed that could make iterative video editing feel closer to adjusting a photo than rendering a project.
The results come from an arXiv preprint, not a shipped product or an independent benchmark, so treat 'seconds' as a lab claim until someone outside the research team puts it through its paces against tools like Runway or Pika.