Fine-tuning an AI model to be great at one thing usually makes it worse at everything else. A new technique called SFT-as-context tries to fix that without retraining anything.
Here is how it works: instead of relying only on a fine-tuned model's answer, researchers feed that answer as context back into the original, pre-fine-tuned "parent" model, letting it learn from the response on the fly. Tested across 19 parent-and-fine-tuned model pairs and 11 benchmarks, the combined setup stayed within about 2 percentage points of the specialized model's performance on tasks like math competition problems (AIME 2024) and coding (LiveCodeBench), while also staying within roughly 2 percentage points of the parent model's broader general skills. The researchers also found it works across the open-closed divide: pairing a small open-source fine-tuned model with a strong closed-source model beat either one running alone.
This matters because the specialized-versus-general trade-off is a real tax teams pay every time they fine-tune a model for a narrow job, like coding or medical Q&A, and then watch it get worse at everyday reasoning. If a cheap, training-free trick can close most of that gap, it changes the calculus on when fine-tuning is even worth it.
It is one paper, not a shipped product, and running two models to answer one question is not free. Whether the latency and compute cost beats just training a better single model is the question nobody's answered yet.