A new training technique lets AI models learn to see better from images without getting worse at reasoning in text.
Researchers have built a new way to train multimodal AI models - the kind that handle both text and images - called LEGO-OPD. It tackles a flaw in a technique known as on-policy distillation, where a smaller "student" model learns by copying one or more larger "teacher" models. Combining a language teacher with a vision teacher is supposed to give students the best of both: sharp reasoning and sharp eyesight. But feeding a student the vision teacher's full output tangles its grounding knowledge up with its own built-in opinions about language, and cranking up the visual input too far makes the student overweight what it sees at the expense of how it reasons.
That trade-off is a quiet failure mode in a lot of multimodal AI: models that describe an image well but stumble on a word problem, or the reverse. LEGO-OPD splits the two signals mathematically, treating the language teacher's guess as a starting assumption and the vision teacher's input as evidence that updates it, with a built-in dial for how strongly the image should push that update at each step. Tested on Qwen3 models, that separation let the student see better without losing ground on text-only reasoning - beating other distillation setups on both fronts.
It is a plumbing fix, not a new kind of AI - but plumbing fixes are often what separates a demo from a model you would actually ship.