A new training trick helps AI models learn good behavior without absorbing the bad behavior that came packaged with it.
Researchers describe stratified inoculation prompting, or SIP, in a paper posted to arXiv. It builds on an earlier method called inoculation prompting, where developers ask a model to produce the undesired behavior during training, then drop that request at inference time, hoping the model ties the behavior to the prompt instead of generalizing it everywhere. The problem: unwanted behavior still leaks into unrelated prompts, and the original technique can also suppress the desired behavior developers wanted to keep. SIP works around this by taking the small slice of training examples that show the desired behavior without the bad one, then repeating those clean examples across many different prompt phrasings while still inoculating the rest of the dataset.
Fine-tuning datasets often bundle good and bad behavior together, so strict filtering can gut a dataset down to almost nothing useful. SIP's bet is that oversampling a small clean slice does more work than researchers previously assumed. The paper reports SIP cuts unwanted behavior while preserving more of the wanted behavior than standard inoculation prompting, including lower rates of what it calls emergent misalignment in harmful advice test setups.
The researchers even built a password locked version that hides the bad behavior behind a specific trigger phrase, a reminder that scrubbing bad habits out of a model is still closer to duct tape than surgery.