AI/ ai safety · backdoor attacks · model distillation · llm security

Safety Training Via Distillation Can Smuggle In Backdoors

Researchers show a poisoned teacher model can pass hidden backdoors to a safety-trained student with as few as 3 percent poisoned examples.

A new study shows that 'safety training' for AI models can carry a nasty side effect: a hidden backdoor transplanted straight into the student.

Researchers tested on-policy distillation, a technique where a smaller 'student' model is trained to mimic a larger 'teacher' model, including techniques meant to make the student safer. They found that if the teacher is secretly backdoored, that malicious behavior transfers to the student even though the student started out clean. Poisoning just 3% of training data produced an attack success rate of up to 70% on the distilled student. With only 10 poisoned samples and 16 training epochs, the attack success rate still hit 67%, and a common distillation method called top-k KL spread the backdoor faster than the alternative sampled-token KL approach.

Safety distillation has been pitched as a cheap way to pass good behavior down to smaller models, but it quietly assumes the teacher and its training data can be trusted. This research shows that assumption breaks easily and cheaply, which matters for any team importing a third-party teacher model or training pipeline without auditing it line by line.

The researchers' proposed fix, Lazy Defense, only slows the backdoor's spread rather than stopping it, which is a modest patch for what sounds like a structural problem.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →