AI/ ai-alignment · model-training · ai-safety · research

A Cheaper Way to Change What AI Models Believe

A new technique called grafting edits AI beliefs without full retraining, and fixes a glitch where edits bled into unrelated facts.

Researchers have found a cheaper way to change what an AI model believes, without retraining it from scratch.

The standard method, called synthetic document fine-tuning, feeds a model fabricated documents to implant a belief, such as a fake fact or a personality trait. Done early, during pre-training, it works cleanly, but pre-training runs are slow and expensive to redo, so most labs apply the technique later, to an already post-trained model. That shortcut has a cost: the model starts treating unrelated fabricated entities as real too, a glitch the researchers call reality drift, and it degrades the model's existing capabilities. Their fix, called grafting, trains the belief-change adapter on the original pre-trained checkpoint, then adds that weight update onto the finished, post-trained model.

Tested on model families up to 284 billion parameters, including attempts to implant false facts and build intentionally misaligned test models, grafting matched the standard method's ability to install a belief while cutting reality drift and loss of coherence by more than half. That matters because alignment research runs on iteration: right now, every tweak to a model's early training requires redoing the entire expensive post-training process just to check if it worked. A technique that skips that step turns months of testing into a single fine-tuning run.

It is a workaround, not a cure: grafting approximates what a clean pre-training run would produce, it does not replace one.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →