AI/ ai · model-merging · reinforcement-learning · vision-language-models

Researchers Show Partial Transplants Boost AI Visual Reasoning

A new technique transfers only the most useful slices of an RL update between AI models, lifting visual-reasoning scores more than copying the whole update.

A new paper argues that copying every bit of a reasoning upgrade from one AI model to another is overkill; the trick is knowing which parts to skip.

Researchers built a method called Selective-RL that transfers reinforcement-learning updates from language models into vision-language models without retraining. Instead of merging entire model checkpoints, the team isolated just the parameter changes created during RL post-training, then kept only the dominant directions of that update while preserving their magnitude. They tested the approach across three model families and five visual-reasoning benchmarks, including one case where a Qwen-based recipient model gained 8.55 percentage points on the MathVision benchmark. Selective-RL beat full-update transfer in 12 of 15 comparisons, and control tests showed the gains were not simply a byproduct of larger updates or arbitrary low-rank trimming.

Model merging has become a cheap substitute for expensive retraining, letting labs bolt reasoning skills from one model onto another with no extra training. This paper suggests that approach has been cruder than it needed to be: full updates carry noise that can actually hurt transfer, and isolating the useful fraction gets better results for less computation.

It is a useful reminder that a bigger merge is not automatically a better one, and the paper's 12-of-15 win rate is a healthy dose of skepticism before anyone declares model merging a solved problem.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →