Merging AI models to save on compute usually costs you accuracy, and a new calibration method claims to close most of that gap.
Model merging combines several fine-tuned task experts into a single model, skipping the need to retrain from scratch or host separate models for each task. The problem is the merged model typically underperforms the individual experts it's built from. Researchers studied this as feature drift: the difference between what the merged model produces and what a dedicated expert would produce on the same input. Their fix, called FeatCal, uses a small calibration set to adjust the merged model's weights layer by layer, with a closed-form calculation instead of gradient descent or extra training.
The numbers matter here because sample efficiency determines whether a calibration method is actually practical. FeatCal hit 85.5% on a CLIP vision benchmark, ahead of the two closest rival methods at 77.0% and 78.8%, and needed only 8 examples per task to reach 82.9%. On a GLUE language benchmark it scored 85.2% versus 82.2%-83.7% for those rivals, while running about four times faster.
For any team trying to avoid running a dozen separate fine-tuned models in production, that's the real pitch: less accuracy tax, less compute, and a calibration step measured in seconds rather than minutes.
Worth remembering this is benchmark performance on CLIP and GLUE, not a live production system, and the baselines it beats - Surgery and ProbSurgery - are themselves recent, narrow research tools rather than industry standards.