A new patch-like tool called Whiteout can erase your home address or phone number from an AI model's memory without lobotomizing the rest of it.
Researchers built Whiteout to respond to individual takedown requests by overwriting memorized personal data - birth dates, phone numbers, home addresses - with carefully crafted obfuscation samples, rather than deleting broad swaths of training data the way machine unlearning does. They tested it on LLMs of varying sizes and makers, including a widely used OpenAI model. Whiteout blocked disclosure of the targeted personal information while leaving overall model utility and safety mostly intact, beating existing alternatives on both counts. The team also threw black-box jailbreaking attempts and white-box attacks like relearning and quantization at it, and the protections held.
Machine unlearning, the current default fix, tends to overcorrect - stripping out more than the offending data, dulling the model, and remaining vulnerable to being undone through fine-tuning tricks. Whiteout's narrower overwrite approach suggests privacy fixes do not have to trade away usefulness, a tension that has dogged every forget-me request made of a trained model so far. For executives, politicians, and judges, whose details apparently turn up in model outputs, that is a more surgical option than what existed before.
Worth noting: this treats a symptom, not the cause. These models still swallow personal data scraped from the open web during training. Overwriting it after the fact is a patch, not a reason to scrape more carefully.