AI/ ai · model-merging · machine-learning · computer-vision

AI model merging method nearly matches specialist models

A new method called ReForge merges several specialist AI vision models into one with almost no accuracy loss, though language merging still lags behind.

A new technique called ReForge closes most of the gap between merged AI models and the specialist models they are built from, without extra training.

Researchers describe ReForge as a bilevel optimization framework that treats merging as Bayesian linear regression anchored to a strong prior model, with an inner step producing a closed-form estimate from unlabeled calibration data and an outer step tuning regularization and scaling via Bayesian optimization on a validation set. A data-free variant swaps the calibration activations for task-vector statistics, so it needs no extra data at all. Tested on merges of up to 20 tasks in vision and 5 tasks in language, ReForge beat baseline methods including TA, WUDI-Merging, and TSV across the board. On a 20-task ViT-B/32 vision benchmark, it pushed the previous best method from 77.6% to 82.8% accuracy with calibration data, and to 81.5% without it.

On an eight-task ViT-L/14 vision benchmark, the calibration-assisted version hit 95.1% mean accuracy, within a percentage point of the 95.8% you get from keeping each task's specialist model separate. That is a meaningful result: merging several specialist vision models into one no longer costs much accuracy, which matters for anyone trying to avoid training and serving a pile of separate models. Language merging, by contrast, only went up to 5 tasks in this paper, suggesting the technique is further along for vision than for text.

The code has not shipped yet, so nobody outside the authors' own benchmarks can confirm the gains hold up.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →