AI/ ai-agents · llm-benchmarks · decision-making · ai-research

When AI judges must predict, not just evaluate, they fail

A new study finds the judgment model Jev excels at instant evaluation but stumbles when the correct choice depends on simulating what the input leaves out.

A new benchmark shows the fast judgment model Jev is great at snap verdicts and bad at guessing what happens next.

Researchers ran Jev, a model that scores options with a single probability call and no reasoning text, through reflection tests, one-shot matrix games, the ALFWorld text adventure, and robot-control tasks. It solved 99% of the classic Cognitive Reflection Test's trick questions, cases where the right answer is spelled out in the prompt itself. But its accuracy cratered whenever the correct choice depended on something the input didn't state outright, like an opponent's next move or an unlisted prerequisite step. Told to put a clean knife in the drawer in ALFWorld, Jev grabbed the dirty knife and delivered it straight to the drawer, skipping the sink entirely. In matrix games, it played as though its opponent were acting randomly rather than rationally.

The strange part is that Jev isn't ignorant. Asked point-blank what the opponent would do, or what step comes first, it usually answers correctly. The failure shows up only when a single model call has to both predict the missing piece and judge based on that prediction, a two-step that current judgment models can't pull off in one pass. The researchers' fix is to stop asking the model to guess: feed it a lookahead from actual code or a physics simulator, and Jev's judgment becomes reliable again.

It's a tidy argument against building agents around a single all-purpose brain, and a reminder that the fastest AI component in a pipeline is often the worst one to trust with guesswork it can't see.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →