A new research model called CaC scans AI generated video to hunt down glitches other detectors miss.
Researchers built CaC, short for Concentrate and Concentrate, a vision-language model that judges anomalies in AI-generated video using a two-step process: it scans the whole clip to find the rough time window where something goes wrong, then zooms in on that window to pinpoint exactly which pixels break down. To train it, the team built what they describe as the first large-scale generated-video dataset with frame-by-frame bounding boxes, timestamped anomaly windows, and labels explaining what went wrong. Training happened in three stages, starting with supervised fine-tuning on single and multiple frames, then reinforcement learning using a two-turn version of Group Relative Policy Optimization, with extra reward signals tied to how well the model's location guesses overlap the real anomaly in time and space.
On benchmarks built for catching subtle anomalies, CaC scored 25.7% more accurately than prior approaches. More notably, when researchers plugged CaC in as a reward signal during video generation, not just as a checker after the fact, it cut anomalies in the output by 11.7% and improved overall quality. That is the more interesting result: a good detector can double as a training signal for better generators.
Text-to-video tools still stumble over things like extra fingers or objects that melt mid-frame, and models like CaC are essentially quality control built to sit inside the training pipeline, not just after it.