Researchers have built a way for vision language models to read compressed text images, then selectively un-compress only the bits they can't make out.
The system is called LensVLM, an inference framework and post-training recipe built on top of Qwen3.5-9B-Base. Vision language models can render text as images instead of long token sequences, and shrinking the image resolution works as a compression knob. The catch is that accuracy falls apart once characters shrink past what the vision encoder can resolve. LensVLM's fix is to scan the compressed image first, then use learned tools to selectively expand only the relevant regions back to full, readable resolution. Across seven text QA benchmarks, it matches full-text accuracy at 4.3x compression and beats retrieval-based and other compression baselines out to 10.1x, with the gap widening as compression increases.
This matters because context-window costs are a real bottleneck for anyone running VLMs at scale, and "just shrink the image" has been a crude workaround until now. LensVLM's analysis found that training makes the model's visual reading robust to how text is rendered, and that as compression rises the model leans more on the expanded content rather than trying to guess from blurry pixels. It also offers a practical rule of thumb: expand rendered text as text, but expand native documents as high-resolution images when layout carries meaning.
It's a clever patch, not a new law of physics: compression still needs a model smart enough to know what it can't read.