Dev Tools/ coding-agents · llm · developer-tools · claude-code

New method teaches coding AI when to actually ask questions

Researchers built a filter that decides which ambiguous coding requests are worth interrupting a developer over, cutting wrong guesses without the nagging.

Coding agents have a guessing problem, and a new method called CONTRA tries to fix it without turning the agent into a chatty pest.

Researchers built CONTRA, a training-free method that figures out which ambiguous parts of a coding request are actually worth stopping to ask about. It first generates a wide set of candidate clarifying questions, then discards ones that do not touch required behavior or are already answered by the original request. For each surviving question, it writes code under two plausible answers and checks whether the two versions actually behave differently on shared inputs - if they do not, the question gets dropped. On a benchmark called ClarifyCodeBench, CONTRA beat the best existing baseline by 13.88 percentage points in F1 across four coding agents, and it outperformed the clarification behavior built into the Claude Code and OpenHands harnesses when using the same underlying LLM; the team also packaged CONTRA as a Claude Code plugin.

Most coding agents resolve ambiguity by picking an answer and moving on, and that guess becomes load-bearing for everything built afterward - the longer it goes unnoticed, the more expensive it is to unwind. CONTRA's execution-based check is the notable part: instead of asking a model whether a question merely sounds important, it runs the candidate code paths and only flags questions where the answer actually changes behavior.

Still, this is one paper's benchmark score, not a field study of developers getting fewer annoying interruptions day to day. Every agent that currently nags you about trivial details, or silently charges ahead on big ones, has the problem CONTRA is trying to solve - whether it cuts real interruption fatigue is a different question than whether it wins on ClarifyCodeBench.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →