[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-ai-agents-still-cant-write-documentation-like-pros":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},9182,"ai-agents-still-cant-write-documentation-like-pros","AI agents still can't write documentation like pros","The best AI agent in a new documentation benchmark scored just 47.3 out of 100 against standards technical writers actually enforce.","AI agents still can't write documentation good enough for a human editor to approve without heavy rework.\n\nResearchers built DoGBench, a benchmark of 292 documentation tasks pulled from real pull requests and reported gaps in open-source projects including Helm, PostHog, and Mautic. Each task hands an agent a repository snapshot and a trigger, either a code change or a flagged documentation gap, and asks it to produce an acceptable patch in one try or correctly abstain if no update is needed. Seven agents were tested, and the best one scored 47.3 out of 100 on the 117-item held-out split, graded against rubrics built with the projects' own maintainers. A separate audit of 1,267 patches found the most common failures were incomplete work (45.5%), factual errors (36.6%), and missing conceptual or reference coverage (32.5%); a distinct analysis of agent trajectories found a related pattern, with agents stopping after fixing the first plausible page and leaving other affected docs stale in 30.1% of cases.\n\nDocumentation is the part of software work that's easiest to defer and hardest to benchmark, which is exactly why it keeps getting pitched as a job for AI agents that handle the boring stuff. A sub-50 score against maintainer-built rubrics suggests the boring stuff still needs a human who has context the agent doesn't, specifically knowledge of how readers actually use an interface, not just what it does.\n\nCall it a reality check for anyone selling autonomous doc-writing today. The benchmark's authors note a score of 100 doesn't mean matching a human expert, just clearing every requirement for the task.","[\"ai\",\"dev-tools\",\"benchmarks\",\"documentation\"]","2026-10-01T04:00:00.000Z","2026-10-02T00:25:52.378Z","2026-10-02T00:25:58.072Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The third stat in the 1,267-patch audit paragraph is wrong: the source's third failure mode for that audit is 'incomplete conceptual or reference coverage' at 32.5%, but the draft substitutes 30.1%\u002F'stopped after fixing the first doc page,' which is actually from a separate trajectory-pattern analysis — fix the paragraph to either use the audit's real 32.5% figure or clearly attribute the 30.1% stat to the distinct trajectory analysis.","resolved","ai",[30,32,33,34],"dev-tools","benchmarks","documentation",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.39909",0,{"sections":41},[42,45,49,53,58,63,67,72,76,80,85,90,95,100],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",5599,{"name":46,"slug":47,"count":48,"latest_published_at":18},"Security","security",815,{"name":50,"slug":51,"count":52,"latest_published_at":18},"Policy","policy",430,{"name":54,"slug":55,"count":56,"latest_published_at":57},"Deals","deals",298,"2026-09-30T21:00:26.000Z",{"name":59,"slug":60,"count":61,"latest_published_at":62},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":64,"slug":65,"count":66,"latest_published_at":18},"Science","science",163,{"name":68,"slug":69,"count":70,"latest_published_at":71},"Consumer Tech","consumer-tech",149,"2026-09-30T22:57:11.000Z",{"name":73,"slug":32,"count":74,"latest_published_at":75},"Dev Tools",93,"2026-10-01T02:30:48.000Z",{"name":77,"slug":78,"count":74,"latest_published_at":79},"Software","software","2026-09-30T21:41:11.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":84},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]