AI/ ai agents · benchmarks · computer-use · ai evaluation

AI Judges Struggle to Catch Computer Agents Faking Success

AgentHorizon shows the best AI judge gets computer-use task grading right only 80.9% of the time, and tool access can make it worse.

A new benchmark shows that even the best AI judges struggle to tell when a computer-use agent actually finished the job it was asked to do.

Researchers built AgentHorizon, a test set of 1,373 paired instruction-trajectory tasks drawn from 166 hours of human-recorded activity across three operating systems. Each task comes with a trap: a similar-looking instruction swapped in, so a judge has to spot when an agent did something close to the request but not quite the request. The team tested eleven AI judges two ways - feeding them full trajectories of up to 300 screenshots and actions, and letting them inspect the work as coding agents inside five different agent harnesses. The top performer, GPT-5.5, hit 80.9% balanced accuracy on the hardest split.

That matters because companies are increasingly using AI to grade AI agents instead of paying humans to review every action. A judge's blind spots become the system's blind spots, and training pipelines that lean on these scores inherit whatever mistakes slip through. The study also found that giving judges tool access helped some models but hurt open-weight ones - a messy result for anyone hoping "just add more tools" is a reliable fix.

Automated grading was supposed to make agent training cheaper and faster. This benchmark is a reminder that the automated grader needs grading too.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →