A new alignment method called VaccineBooster cuts harmful outputs after poisoned fine-tuning, but it cannot do everything at once.
Fine-tuning-as-a-service lets customers adapt safety-aligned models to their own data, but that same door lets in a harmful-fine-tuning attack: mix a little harmful data into an otherwise normal training set and the model's guardrails erode. Researchers tested two existing defenses, Vaccine, which hardens the model's internal representations against this drift, and Booster, which simulates harmful updates during training and dampens their effect. They combined both into a single procedure called VaccineBooster and tested it on Llama-2-7B aligned with the BeaverTails dataset, then attacked with poisoned fine-tuning. The combined method produced the lowest OpenAI moderation score, 0.315, among the approaches tested, while Booster alone kept the highest post-attack refusal rate at 50%.
This is a trade-off, not a clean win: reducing flagged harmful content and preserving explicit refusals pull in different directions, so a team picking a defense has to decide which failure mode it can tolerate. For anyone running a fine-tuning-as-a-service product, that is the real design question - not whether to add a defense, but which kind of failure you are willing to live with.
The caveat here matters: this is ten prompts, one unseeded run per setup, not a statistically resolved result, so treat the numbers as a directional signal rather than a benchmark to cite in a sales deck.