AI/ fine-tuning · reinforcement-learning · ai-research · machine-learning

Study Finds Supervised Fine-Tuning Makes AI Models Forget More

Researchers show supervised fine-tuning drifts toward style bias and forgets faster than reinforcement fine-tuning, even with correct training data.

A new study says fine-tuned AI models don't forget skills because of bad data. They forget because of writing style.

Researchers compared two ways of fine-tuning a model on classification tasks: supervised fine-tuning (SFT), where a model copies a teacher's example answers, and reinforcement fine-tuning (RFT), where a model is rewarded for correct answers regardless of phrasing. Using a simplified linear-softmax model they could analyze mathematically, they split each training update into a semantic component (getting the right answer) and a style component (how that answer is phrased). Both methods update semantics identically, but SFT also pulls the model toward a teacher's particular phrasing habits, even when every teacher example is factually correct. Over repeated training steps, that stylistic pull compounds into measurable forgetting, with a mathematically guaranteed minimum level of semantic error for SFT, while RFT holds at zero.

That explains something AI teams have run into before: fine-tuning on clean, correct data can still make a model worse at the task it was trained for. It also makes a sharper case for reward-based fine-tuning, the family of methods behind RLHF-style post-training, when the goal is preserving existing knowledge rather than matching a particular tone or format.

The catch is that this proof runs on a toy linear model, not a GPT-5-scale transformer, so treat it as a clean hypothesis about why fine-tuning degrades models, not a settled verdict on your own pipeline.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →