AI/ ai-safety · clip · multimodal-ai · content-moderation

A New Method Teaches CLIP to Unlearn Harm, Not Everything

ShieldCLIP teaches AI image-text encoders to suppress only the harmful half of a pair, instead of blanking out safe content along with it.

A new technique lets AI models that pair images with text unlearn harmful content without forgetting the safe stuff next to it.

Researchers introduced ShieldCLIP, a method for retraining CLIP-style encoders, the models that link images and text and sit underneath tools like Stable Diffusion and LLaVA. Existing safety datasets pair real safe samples with AI-generated unsafe ones, but label the entire generated pair unsafe even when only the image or only the caption is actually harmful. ShieldCLIP instead tracks the safety status of each modality independently, using a new 195,000-sample dataset called ViSUv2 with separate labels for images and text across 578 concepts. The team tested it on cross-modal search, image generation with Stable Diffusion v1.4 and SDXL, and image captioning with LLaVA, and found it cut harmful outputs while keeping the model's normal performance intact.

Most safety filters work like a blunt instrument: scrub anything near bad content, and you lose fidelity on the good content too. By separating what counts as an unsafe image from what counts as an unsafe caption, ShieldCLIP's approach could make moderation layers less likely to degrade ordinary use cases, a problem that has dogged content filters since generative AI moderation tools first went mainstream.

The code and models are promised for release, but the dataset itself will sit behind a controlled-access protocol, a reminder that even open safety research rarely means anyone can just download the unsafe half.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →