A new benchmark says the confidence scores guiding AI agents that click around your screen are only trustworthy if you don't swap the model underneath them.
Researchers built Argus, a benchmark that tests 27 uncertainty-quantification methods across four open-weight vision-language models and four GUI-grounding datasets, plus eight methods across three closed-source frontier vendors where internal signals like logits and attention maps aren't exposed. These "computer-use agents" translate a model's visual read of a screen into an actual click, so knowing when a click is likely wrong matters for both safety and for catching bad guesses before they fire. The methods tested ranged from logit-based scores and sampling-consistency checks to hidden-state density estimators and agents simply stating their own confidence in words.
The headline result: a method's ability to rank risky predictions against safe ones holds up well across datasets for the same model, with correlation scores as high as 0.969, but that ranking barely survives a jump to a different model family. Applied to closed-source vendors, the average correlation drops to just 0.08. Calibrating the disk-shaped safety margins drawn around a predicted click does shrink them by 40 to 60 percent, but only when the calibration data matches the real deployment setup; mismatches break the coverage guarantee.
That's a useful corrective for anyone assuming a single confidence metric travels well from one AI agent to the next vendor's model. Uncertainty estimates here look less like a universal safety feature and more like a setting you have to retune every time the underlying model changes.