Security/ llm-security · backdoor-attacks · lora · ai-safety

Researchers Strip LoRA Backdoors Without Knowing the Trigger

A new null-space projection method cuts backdoor success rates in fine-tuned LLMs from nearly 100% to under 10% without retraining or clean reference data.

Researchers have found a way to scrub hidden backdoors out of fine-tuned language models without knowing what the backdoor trigger even is.

The method targets models adapted with LoRA, the parameter-efficient fine-tuning technique that lets developers customize a large model by training a small set of extra parameters instead of the whole thing. Most existing backdoor defenses need something the defender rarely has: a known trigger phrase, a clean reference model, or the budget to retrain. This approach skips all three. Instead, it extracts the backdoor's "direction" in the model's internal feature space from careful data analysis, then builds orthogonal null spaces for each layer and projects the LoRA updates onto them, mathematically squeezing the backdoor signal out. In tests, attack success rates dropped from nearly 100% to below 10%, while the model kept both its general abilities and the task-specific skills the adapter was trained for.

This matters because LoRA adapters have become the default way teams customize open models, and that supply chain is exactly where backdoors slip in: download a popular adapter from a hub, fine-tune on it, and you inherit whatever was baked in. Trigger-agnostic cleanup means defenders no longer need to guess the attacker's exact phrase or pay for full retraining, which is the part of backdoor defense that has mostly been theoretical until now.

It is one paper with its own benchmarks, not an industry-wide fix, and whether null-space projection holds up against adaptive attackers who know this defense exists is still an open question.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →