AI/ ai · code-generation · benchmarks · requirements

New Benchmark Finds LLMs Bad at Asking Clarifying Questions

ClarifyCodeBench shows top coding models still struggle to ask the right clarifying questions before writing code.

AI coding assistants are great at writing code and apparently terrible at asking what you actually meant.

Researchers built ClarifyCodeBench, an interactive benchmark drawn from real-world programming tasks, to test whether large language models can spot ambiguous or incomplete requirements and ask the right clarifying questions before generating code. The benchmark includes manually annotated ambiguity types, sample clarification questions, and ground-truth answers, plus two new metrics: one that penalizes models for asking too many questions, and one that measures how close they get to the ideal number of clarifying rounds. The team tested six state-of-the-art LLMs. Across the board, the models that wrote the most correct code were not necessarily the ones that asked the best questions.

That's a meaningful split from how the industry usually measures coding assistants, which is almost entirely on whether the final code works. The study found that giving a model more reasoning time improved code correctness but barely helped it notice unclear requirements, and performance fell off fast once a prompt contained multiple ambiguities at once, exactly the messy, multi-part requirements a reader gets from an actual product manager, not a tidy ticket.

Most coding benchmarks still hand the model a perfectly worded spec, which is a bit like judging a translator by sentences nobody actually speaks.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →