AI/ ai · multimodal-ai · model-training · distillation

A New Way to Teach AI to See Without Losing Its Reasoning

A new distillation method separates visual grounding from language reasoning, letting AI models see better without reasoning worse.

A new training technique lets AI models learn to see better from images without getting worse at reasoning in text.

Researchers have built a new way to train multimodal AI models - the kind that handle both text and images - called LEGO-OPD. It tackles a flaw in a technique known as on-policy distillation, where a smaller "student" model learns by copying one or more larger "teacher" models. Combining a language teacher with a vision teacher is supposed to give students the best of both: sharp reasoning and sharp eyesight. But feeding a student the vision teacher's full output tangles its grounding knowledge up with its own built-in opinions about language, and cranking up the visual input too far makes the student overweight what it sees at the expense of how it reasons.

That trade-off is a quiet failure mode in a lot of multimodal AI: models that describe an image well but stumble on a word problem, or the reverse. LEGO-OPD splits the two signals mathematically, treating the language teacher's guess as a starting assumption and the vision teacher's input as evidence that updates it, with a built-in dial for how strongly the image should push that update at each step. Tested on Qwen3 models, that separation let the student see better without losing ground on text-only reasoning - beating other distillation setups on both fronts.

It is a plumbing fix, not a new kind of AI - but plumbing fixes are often what separates a demo from a model you would actually ship.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →