AI/ ai · coding-agents · benchmarks · research

Better Documentation Doesn't Help AI Coding Agents Fix Bugs

A new benchmark finds that feeding coding agents better documentation, compact or retrieved, does not improve their ability to resolve real software issues.

A new benchmark suggests all the documentation people feed coding agents might not be doing much good.

Researchers built a roundtrip test that judges a code description's quality by whether code regenerated from it alone passes the original tests. They found completeness, not length, is what makes a description faithful, and used that signal to optimize a prompt that gets descriptions to full fidelity, including on files it had never seen before. They then tested the actual hypothesis behind the project: does better documentation help an agent resolve real repository issues. Across two model families and ten repositories, with a control confirming their setup could detect genuine improvements, the answer was no. When the source code is available, neither compact documentation nor retrieved context beats just handing the agent the issue itself.

That's an awkward result for coding-agent tooling built on the premise that pre-digested context, whether it's summaries, retrieved snippets, or curated docs, makes agents smarter. It points to the bottleneck sitting somewhere other than information access: agents that can already read the source apparently don't need someone to explain it to them first. The researchers do identify a boundary where documentation helps, just not once the underlying code is already in hand.

Good context only helps if the reader, human or model, doesn't already have everything it needs. Sometimes the extra summary is just extra.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →