Teaching an AI model to see can make it worse at thinking.
Most vision-language models (VLMs) are built by bolting a visual module onto a pretrained large language model (LLM) and then aligning the two through multimodal training. That alignment process degrades the reasoning ability the original LLM had before the visual module was added, even though the underlying LLM still retains that ability - the VLM just can't reliably access it anymore. Researchers' fix, called LIFT (Language-side reasonIng Facilitation and Transfer), extracts "reasoning vectors" - the difference in hidden-state activations between a model working through a problem step by step versus answering directly - from the base LLM and injects them into the VLM's language-side activations, without retraining the backbone. Tested on two VLMs across six reasoning benchmarks, vectors pulled from the base LLM consistently beat vectors pulled from the VLM itself.
That result confirms something uncomfortable for VLM builders: the reasoning ability was there in the original language model all along, current alignment methods just bury it, and there was no cheap way to dig it back out. LIFT sidesteps the expensive alternative - retraining or fine-tuning the whole multimodal model - by treating reasoning loss as an activation-steering problem, in the same family as other vector-based techniques used to control LLM behavior. It also reframes why VLMs underperform on reasoning tasks: not because visual grounding conflicts with reasoning, but because the merge process misplaces reasoning the model already knows how to do.
It is a patch, not a cure - the paper calls the recovery "partial," and the code has not shipped yet, so nobody outside the lab can check how far LIFT actually goes.