A new research harness lets AI models crop, zoom, and compare images instead of just describing them in words.
Researchers introduce Mosaic, a general-purpose multi-image harness that gives multi-modal AI models ten composable image operations to actively construct visual intermediates while reasoning, rather than relying solely on text-based chain-of-thought. They compared five re-representation strategies, some textual and some visual, across existing multi-image benchmarks plus a new benchmark called MosaicBench, built specifically to test fine-grained, grounding-heavy visual reasoning. The results show the better strategy depends heavily on the task: visual re-representation wins on jobs that need precise visual evidence, like hypothesis testing, precision comparisons, and orientation-sensitive reasoning, while tasks about higher-level meaning show smaller, less consistent gains from either approach. Building on that finding, the team trained an 8-billion-parameter model, MosaicAgent-8B, to use Mosaic via reinforcement learning, rewarding only correct answers and proper output format, without any step-by-step demonstrations of how to use the tools.
Most multi-image AI tools today reason entirely in text, describing what's in each photo and comparing the descriptions, which breaks down when the answer hinges on a visual detail no caption fully captures. Mosaic's task-dependent results argue against defaulting to text-only chain-of-thought for visual agents, and the fact that MosaicAgent-8B learned to chain multiple image operations on its own, with no demonstrations and no tool-specific rewards, suggests reinforcement learning can teach genuinely useful tool-use habits rather than just pattern-matched responses.
It's a reminder that teaching a model to look more carefully, rather than just prompting it to think longer, might be the bigger unlock for visual reasoning.