Security/ ai-security · backdoor-attacks · llm-safety · mixture-of-experts

Researchers Cage LLM Backdoors Inside One Disposable Expert

A new technique lets AI models learn backdoors on purpose, then neutralizes them by shutting off a single quarantined component at deployment.

A new defense trains AI models to grow their backdoors in a cage, then kills the cage at deployment.

Researchers propose Quarantined Expert Shutdown (QES), a method that lets a backdoor form during training but corrals it into one routed "expert" module inside a mixture-of-experts-style setup, built from LoRA branches and lightweight routers. Instead of trying to spot and filter poisoned training data up front, or patch a model after the fact, QES uses routing objectives to pull trigger-conditioned behavior into that single quarantined expert while the rest of the network keeps its normal capabilities. At deployment, defenders simply zero out that expert's routing weight - no scanning inputs for triggers, no retraining required. Across two tasks, three attack types, and four model families, the approach cut attack success rates from 100% down to 0-10% in most cases, while leaving the model's regular performance largely intact.

Most backdoor defenses either try to stop poisoning before it happens or clean up a model that's already compromised - both of which require knowing what a trigger looks like or doing expensive retraining. QES flips that logic: it treats backdoor formation as inevitable and contains it instead, more like isolating an infection than preventing one. That's a meaningfully cheaper failure mode for anyone fine-tuning on data they can't fully vet, which, given how much scraped and crowdsourced data feeds LLM training, is most people.

Still, this is a lab result against known attack types under controlled poisoning. Real attackers adapt, and a defense built around a dedicated quarantine expert is itself a tempting new target to probe.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →