Researchers have a new way to make a large language model forget specific facts without retraining it at all.
The method, called Nullify, is training-free. Instead of fine-tuning a model to scrub memorized information, it steers the model's internal activations away from sensitive content at inference time. A null-space constraint keeps that steering from touching anything else the model knows, so unrelated answers stay intact. On the TOFU and MUSE unlearning benchmarks, the researchers report that Nullify matches or beats existing fine-tuning-based methods at forgetting target data, while causing almost no loss in general model utility.
That tradeoff, thorough forgetting versus a model that still works, is the central problem in LLM unlearning, and it is why most prior approaches are expensive: retraining or fine-tuning a model to erase one person's data risks degrading everything else it learned. Doing this at inference time, with no weight updates, could make targeted deletion requests cheap enough to run routinely instead of as a rare, costly retraining event.
One catch: steering activations at inference time does not delete anything from the model's weights. The memorized information is still in there, just redirected. Whether that counts as real unlearning in a legal sense is a different question from whether it wins a benchmark.