Picking the right AI encoder for reading CT scans just got a lot cheaper.
Researchers built CheapCT, a lightweight probe that scores a 3D CT image encoder's raw representations without running it through a full language model. They tested it against the standard approach: fine-tuning every candidate encoder end to end and comparing downstream results, a process that burns enormous compute. The comparison ran on two benchmarks, report generation, which grades a full radiology report and mostly reflects disease detection, and a new dataset called MeasureVQA, which the team built to score individual capabilities against segmentation masks and Hounsfield units. Across both, CheapCT's rankings tracked the expensive fine-tuning rankings closely, with rank correlation between 0.90 and 0.97, holding steady even when the probe readout or the language model backbone changed.
That consistency matters because encoder selection is usually the first and most expensive decision in building a medical vision-language model, and today it means fine-tuning every option just to find out which one works. A cheap, reliable stand-in changes who can afford to run that search: smaller labs and hospital research teams, not just organizations with large training budgets. The authors report CheapCT lands on an encoder nearly as good as the best choice while only requiring one full fine-tuning run instead of many.
It is one modality and one team's benchmark, released alongside the code rather than validated by outside labs, so treat this as a promising shortcut rather than a settled best practice.