Courts now have a yardstick for catching AI-faked photos submitted as evidence.
Researchers released the CIFAR Synthetic Evidence Corpus, a benchmark built specifically for evidentiary image authentication. It contains 1,505 photos: 720 real, unaltered images and 785 manipulated or fully fabricated ones, spanning surveillance stills, dashcam frames, and ordinary phone photos. The fakes are sorted into three tiers - scene-condition edits, localized edits to a single element, and full fabrications - each produced with contemporary generative tools and tagged with metadata on the generator, prompt template, and manipulation type. The team also benchmarked existing publicly available forensics detectors against the set.
The baseline results are the real story: current detection tools still show error profiles that would be unacceptable in a courtroom, where one missed fake could let fabricated evidence stand and one false alarm could get real evidence thrown out. Earlier forensics benchmarks focused on face-swaps or generic synthetic images; this one is organized around the specific things evidence is used to prove - presence, sequence, causation, identity.
Dashcam footage and surveillance stills have been treated as near-unimpeachable in court for decades. A benchmark showing current tools can't reliably tell real from fake is a quiet admission that assumption no longer holds.