A new defense trains AI models to grow their backdoors in a cage, then kills the cage at deployment.
Researchers propose Quarantined Expert Shutdown (QES), a method that lets a backdoor form during training but corrals it into one routed "expert" module inside a mixture-of-experts-style setup, built from LoRA branches and lightweight routers. Instead of trying to spot and filter poisoned training data up front, or patch a model after the fact, QES uses routing objectives to pull trigger-conditioned behavior into that single quarantined expert while the rest of the network keeps its normal capabilities. At deployment, defenders simply zero out that expert's routing weight - no scanning inputs for triggers, no retraining required. Across two tasks, three attack types, and four model families, the approach cut attack success rates from 100% down to 0-10% in most cases, while leaving the model's regular performance largely intact.
Most backdoor defenses either try to stop poisoning before it happens or clean up a model that's already compromised - both of which require knowing what a trigger looks like or doing expensive retraining. QES flips that logic: it treats backdoor formation as inevitable and contains it instead, more like isolating an infection than preventing one. That's a meaningfully cheaper failure mode for anyone fine-tuning on data they can't fully vet, which, given how much scraped and crowdsourced data feeds LLM training, is most people.
Still, this is a lab result against known attack types under controlled poisoning. Real attackers adapt, and a defense built around a dedicated quarantine expert is itself a tempting new target to probe.