AI models that "think out loud" before finding an object in an image may be wasting effort - and attention.
Reasoning segmentation asks a multimodal AI to parse a query like "the object the dog is chasing" and then draw a precise mask around the right pixels. Most systems handle this by generating an explicit chain-of-thought in text before localizing the target. A new paper introduces LIRSeg, which drops that text step entirely in favor of a small set of learnable latent tokens, trained first to ground themselves in relevant visual evidence and then refined with reinforcement learning (GRPO) using segmentation-accuracy rewards. Three additional training tricks - selecting the most informative signals, separating exploration from stability updates, and preventing the latent tokens from collapsing into redundant representations - keep those tokens useful. Against the VisionReasoner baseline, LIRSeg posts gIoU accuracy gains of 4.9 points on ReasonSeg, 7.1 points on MUSE, and 4.7 points on MMR, while needing roughly 16 times fewer reasoning tokens.
The result is a data point against the assumption that more visible reasoning always helps. Chain-of-thought prompting has become the default way to make multimodal models look smarter, but padding a model's output with redundant text tokens can crowd out the attention a vision task actually needs. If latent reasoning holds up outside these three benchmarks, it points toward leaner, faster multimodal systems for things like robotics and interactive assistants that can't afford to wait on paragraphs of internal monologue.
The code is only in the paper's supplementary materials for now, and the gains are measured against a single baseline - worth watching for independent replication before anyone calls this the new standard.