A new research framework claims it can flag unusual events in video without being trained on any labeled examples.
Researchers describe Cog-VADU, a training-free system for video anomaly detection. Rather than fine-tuning on curated clips, it runs a large vision-language model through a step-by-step reasoning chain called Chain-of-Anomaly Detection Thought Prompting, or CoADTP. The chain moves across consecutive video segments, carrying forward written rationales so the model keeps something like a running memory of what it has already seen. A second stage then cross-checks those rationales against the raw visual embeddings, which the researchers say keeps predictions consistent and temporally coherent. Tested on multiple public anomaly-detection benchmarks, the team reports competitive zero-shot performance, and found the reasoning approach kept helping even when swapped across different underlying vision-language models.
Most anomaly detectors need labeled footage specific to wherever they will run - a model trained on one parking lot's camera feed does not automatically work on another's. Skipping that training step, while keeping a reasoning trail instead of a black-box score, is the actual pitch: a system that can explain its call and move between contexts without retraining.
Zero-shot, vision-language-model-based anomaly detection has been building for a couple of years, and the hard part has always been time - telling a genuine anomaly apart from ordinary high-motion activity. Whether chained reasoning holds up against messy real-world footage, with bad lighting and worse camera angles, is a question a benchmark table cannot fully answer.