AI/ ai · computer-vision · anomaly-detection · clip

New Scoring Method Cuts False Alarms in AI Defect Detection

A new post-hoc scoring method called TED helps CLIP-based defect detectors tell real flaws from complex normal textures without retraining.

AI quality-inspection tools built on CLIP are good at spotting weirdness. The trouble is they often cannot tell a real defect from a normal part that just looks weird.

Researchers behind a new paper propose TED (Text-Axis Evidence Decomposition), a post-hoc scoring method for CLIP-based anomaly detectors used in defect localization. They found that when these models are adapted with prompts or lightweight add-on modules to spot defects more sensitively, they also start flagging visually complex normal regions as highly anomalous. The paper argues this is not a case of missing defect information - it is a scoring rule that cannot distinguish genuine defect evidence from normal-but-confusing evidence that looks similar. TED fixes this by checking whether an ambiguous response is better explained by real defect patches or by normal patches being mistaken for anomalies, comparing both against the model's own normal-versus-anomaly text prompts, without retraining the backbone or changing the prompts.

The gains land exactly where they matter most: in the hardest cases, where normal-looking complexity fools the detector, average pixel-level localization improved by about 10.9 points, versus 5.0 points in easier settings, according to the paper. For factories and other visual-inspection pipelines leaning on repurposed vision-language models rather than detectors trained from scratch, that is a cheap reliability upgrade that does not require collecting more labeled images from the target domain.

It is a software patch for a hardware-level problem: CLIP was never built to split hairs on a factory floor, and no amount of clever prompting fully closes that gap - TED just makes the gap smaller and cheaper to work around.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →