AI/ video-compression · generative-ai · rectified-flow · research

A New Codec Lets AI Guess the Video Instead of Storing It

Researchers built a zero-shot video codec that sends a pretrained AI model instructions to regenerate footage instead of transmitting pixels.

A new research codec skips storing video entirely and instead ships instructions telling a generative AI model how to redraw it, frame by frame.

The paper, GVCC (Generative Video Codebook Codec), takes a pretrained video generation model and turns it into the decoder. Normally these models use a deterministic step-by-step process called rectified flow to turn noise into video, which leaves no room to smuggle in compressed data. The researchers rewired that process into a mathematically equivalent stochastic version, so the bitstream can transmit the random choices made at each step instead of raw pixels. They built three versions: one that generates video from a text prompt alone, one that extends a single starting frame, and one that fills in the gap between a first and last frame. All three were tested on the seven-sequence UVG dataset, measuring quality against how many of these encoded choices ("atoms") were spent per clip.

This flips the logic of standard video compression. Codecs like H.264 or AV1 squeeze out redundancy and, when bitrate gets too low, the picture turns blurry or blocky. GVCC instead lets the AI hallucinate plausible detail at the low end, which could matter for things like video calls over bad connections, where blur is worse than a slightly reinvented background.

Worth noting: the authors explicitly did not claim their method beats existing codecs at matched bitrates, only that they measured its behavior. Seven test clips and a "zero-shot" label are early-stage territory. A codec is only as good as the model doing the guessing, and nobody has shown yet that the guess holds up against the real thing.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →