New research says the fix for a poisoned language model is correction, not deletion.
A paper titled "Correct, Don't Delete: Mitigating Emergent Misalignment with Corrective Supervision" (arXiv:2609.37624), posted September 30, 2026, tests what happens when a language model is fine-tuned on a mix of bad medical advice and normal chat data. The researchers fine-tuned Qwen2.5-14B-Instruct on that mixture, then took a known slice of the bad rows and either deleted them or swapped in corrected answers to the same prompts. Deleting the bad rows barely changed the model's tendency to give broadly misaligned answers on unrelated topics, a failure mode called emergent misalignment. Replacing them with corrections cut that misalignment rate by about a third and improved the model's answers to medical questions it hadn't seen.
That is a useful, counterintuitive result for anyone building safety pipelines around "find and remove the bad data": removal is the industry's default reflex, and this suggests it is the weaker move. The paper also found that generic realignment training, just throwing more clean chat data at a poisoned model, works less well than the same amount of training on actual corrections, and that instructing the correction-writer to sound careful and harm-avoiding added no measurable benefit over plain fixes.
The results held on a second base model and a second poisoning setup, more replication than most single-paper safety claims get, though it remains one paper on one narrow harm category before anyone rewrites their data-cleaning playbook.