A new watchdog system can stop AI coding agents from compounding small mistakes into big ones, by reviewing their next move before it runs.
Researchers built HiSentinel, a pair of small language models (0.6 billion and 1.7 billion parameters) trained to look at a coding agent's proposed next action and decide whether to let it proceed, redirect it, or pause the task for a human to step in. The team trained these sentinels using hindsight distillation: a larger, privileged model first judged which past actions actually helped or hurt task completion, using the real recorded outcomes as evidence, and that judgment was distilled into a smaller model that sees only the action about to happen, not the eventual result. To train and test this, the researchers also built a new dataset, SWE-Intervene, which labels real software-engineering action sequences as allowed, redirected, or escalated to a human. On two benchmarks, SWE-bench Verified Mini and Ask or Assume, adding these sentinels raised task completion rates by up to 14% and 10%, respectively, without a big jump in token usage.
The real story here is cost. Coding agents often fail not because they can't code, but because one bad early step, like deleting the wrong file or misreading a test, sends the whole trajectory off course, and recovery is expensive. A cheap, sub-2-billion-parameter checker that catches that before execution is a very different economic proposition than running a second large model as a judge after the fact.
Still, these are benchmark numbers from two specific test sets, not proof that sentinels generalize to the messy, open-ended repositories most engineers actually work in.