AI/ ai-agents · terminal-agents · llm-verification · test-time-compute

Verifying Commands First Pushes AI Agent Success to 68%

A new technique verifies candidate terminal commands with a separate AI model before execution, raising one benchmark's success rate from 50% to 68%.

AI agents that run real terminal commands get less reckless when a second model double-checks their work before anything executes.

Researchers built Mid-Harness, a system that samples several candidate shell commands from an agent's existing model, then uses a separate verifier to pick the best one before it runs, without touching the underlying generator or the harness around it. On the TerminalBench-Lite benchmark, sampling more candidate actions from a TMAX-9B model barely helped when the verifier judging them was weak. Swapping in a stronger verifier, GPT-5.6 Sol, raised the agent's Pass@1 success rate from 50.00 percent to 68.03 percent using eight sampled actions. When the TMAX-9B model verified its own candidates instead, comparing pairs of actions against each other outperformed the other verification methods tested, and distilling the stronger verifier's judgments back into TMAX-9B pushed results higher still, even though the command-generating model never changed.

The problem Mid-Harness targets is specific to agents that touch real systems: a single bad command, like installing the wrong package, can corrupt an environment in ways the agent cannot undo, even if a better command was sitting right there in its own output. That is a different failure mode than a chatbot giving a wrong answer you can just regenerate. The paper also found that combining this per-action checking with scaling up full attempt counts beat just generating more full trajectories, at a lower estimated token cost.

It is the same move that improved math-solving models years ago: verify each step instead of only the final answer. Here it is aimed at agents with the ability to actually break their own sandbox, which makes the case for checking before executing a lot more concrete.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →