A new scaling law says bigger vision-language models and higher-resolution images do not fix every mistake, and researchers can now predict which ones they will fix.
Researchers tested 26 InternVL and QwenVL vision-language models, with language backbones ranging from 1 billion to 72 billion parameters, across four high-resolution benchmarks using images from 224 pixels up to 8K resolution. They found that whether more scale helps a given question depends on the skill that question requires, and a meaningful share of errors never improve no matter how much compute is applied. The two model families benefited about equally from bigger language backbones, but their gains from more visual tokens diverged sharply between families. The team packaged these findings into what they call the Separable Law, paired with a cost law, to produce a closed-form formula for splitting a compute budget between backbone size and image resolution.
This matters because most teams tune backbone size and image resolution by trial and error, often paying for resolution increases that add cost without fixing the errors actually dragging down performance. A formula that predicts, ahead of deployment, which configuration lands closest to the best feasible result under a fixed budget turns that guesswork into a lookup table. For anyone deploying vision-language models at scale, that is the difference between burning compute on hope and spending it on evidence.
It is a tidy rebuttal to the idea that scale fixes everything - some errors are structural, and no amount of extra pixels or parameters will touch them.