A small AI model just tied a frontier-scale one at predicting kidney injury in ICU patients, using a fraction of the training cost.
Researchers built ViSTA, a lightweight adapter that teaches vision-language models to read patient vital signs and lab trends the way they read charts. Instead of retraining a model's core, ViSTA nudges its existing visual tokens using irregular time-series data while leaving the underlying pretrained model untouched. Tested on the MIMIC-IV clinical dataset, it beat other adaptation methods on all four metrics for predicting acute kidney injury and mortality, across models ranging from 2 billion to 9 billion parameters. The standout number: a 2-billion-parameter model using just 0.516 million trainable parameters scored an AUROC of 0.7376 on kidney injury risk, within a hair of GPT-5.6 Sol's 0.7380 using text input and high reasoning effort.
That gap matters more than it looks. GPT-5.6 Sol is a large model reading text descriptions of patient data with high reasoning effort switched on. ViSTA gets nearly the same accuracy from a model a fraction of the size, trained on a sliver of the parameters, by treating numbers as pictures instead of paragraphs - it also cut trainable parameters by over 90% versus low-rank adaptation on a temporal question-answering task, landing within 2.82 to 4.88 percentage points of that baseline's accuracy.
None of this makes ViSTA ready for a hospital floor. An AUROC in the mid-0.70s is a research benchmark number, not a clinical green light, and prior attempts to bolt clinical time-series data onto language models have stalled at exactly this stage - promising numbers, no deployment.