AI/ ai · dev-tools · benchmarks · documentation

AI agents still can't write documentation like pros

The best AI agent in a new documentation benchmark scored just 47.3 out of 100 against standards technical writers actually enforce.

AI agents still can't write documentation good enough for a human editor to approve without heavy rework.

Researchers built DoGBench, a benchmark of 292 documentation tasks pulled from real pull requests and reported gaps in open-source projects including Helm, PostHog, and Mautic. Each task hands an agent a repository snapshot and a trigger, either a code change or a flagged documentation gap, and asks it to produce an acceptable patch in one try or correctly abstain if no update is needed. Seven agents were tested, and the best one scored 47.3 out of 100 on the 117-item held-out split, graded against rubrics built with the projects' own maintainers. A separate audit of 1,267 patches found the most common failures were incomplete work (45.5%), factual errors (36.6%), and missing conceptual or reference coverage (32.5%); a distinct analysis of agent trajectories found a related pattern, with agents stopping after fixing the first plausible page and leaving other affected docs stale in 30.1% of cases.

Documentation is the part of software work that's easiest to defer and hardest to benchmark, which is exactly why it keeps getting pitched as a job for AI agents that handle the boring stuff. A sub-50 score against maintainer-built rubrics suggests the boring stuff still needs a human who has context the agent doesn't, specifically knowledge of how readers actually use an interface, not just what it does.

Call it a reality check for anyone selling autonomous doc-writing today. The benchmark's authors note a score of 100 doesn't mean matching a human expert, just clearing every requirement for the task.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →