AI/ reinforcement-learning · ai-agents · coding-agents

A New Training Trick Teaches AI Agents What Actually Worked

A new reinforcement learning method traces which terminal commands actually drove a coding agent's result, aiming to fix sloppy credit assignment in training.

A new reinforcement learning method tries to stop AI coding agents from getting credit, or blame, for the wrong terminal commands.

Researchers describe a technique called Dependency-Aware Group Policy Optimization (DepGPO) for training "terminal-using" agents, the kind that write and debug code by running shell commands over multiple steps. The problem they identify: existing RL methods score an agent's whole run or individual steps, but never trace which specific commands actually produced the output a task's verifier checks. That means a command that did nothing useful can get rewarded right alongside the one that mattered, while a command that quietly set up the final answer can get ignored. DepGPO instead builds a dependency graph of each command's reads and writes from the execution trace, works backward from whatever the verifier inspects, and routes credit only along that chain. The paper reports better task performance and more stable training across its experiments and ablations on complex terminal tasks.

This matters because credit misassignment is a known weak point in RL-trained coding agents, and it compounds in multi-step terminal work: one stray cd or an ignored pip install error can poison the signal for everything downstream. A dependency-graph approach is a more structural, debuggable fix than the usual answer of just running more rollouts, which is relevant to anyone building the terminal and computer-use agents now showing up in coding tools.

Still, this is one arXiv paper with self-reported benchmarks and no comparison to the terminal agents already shipping in commercial coding tools. A dependency graph built from a clean execution trace is one thing; a real shell full of network calls, background processes, and side effects the verifier never sees is another.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →