[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-study-shows-ai-coding-agents-waste-context-on-redundant-evidence":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},7011,"study-shows-ai-coding-agents-waste-context-on-redundant-evidence","Study Shows AI Coding Agents Waste Context on Redundant Evidence","A new arXiv paper finds retrieval tools feed coding agents redundant passages instead of the complete evidence sets their decisions require.","A new paper argues that AI coding assistants are drowning in redundant search results, not missing facts.\n\nThe paper, arXiv:2609.20050, titled 'The Missing Complement: State-Conditioned Minimal Sufficient Evidence for Coding Agents,' says retrieval tools score each passage for relevance but never check whether the resulting set covers everything a decision needs, so a ranker can hand an agent five variations of one fact and call the job done. The authors built SERBench, a benchmark of 500 captured agent states across 45 repositories, to test whether retrieved evidence is complete rather than merely plausible. They also introduce MSS-Complement, a method that uses three semantic calls to assemble a jointly sufficient set of 4-8 source units within a 6,144-token budget. On calibration data, it recovered a complete evidence set for 73.0% of states at five items and 80.6% at eight, compared with 61.4% and 72.4% for a Qwen3 embedding-and-reranking baseline.\n\nA similarity-only control still hit 66.6%, so the gain traces to the set-construction approach rather than extra compute, and on frozen repository source with no curated answer pool the lead held at 5.0 points. On a second benchmark, AMA-Bench, the method answered from a prompt 76.2% smaller than usual while scoring 2.08 points above that benchmark's own memory agent. Most tellingly, pulling one required fact out of an otherwise complete set cost 11 to 12 points of repair-localization precision under two different code-editing executors, suggesting completeness, not volume, is what agents actually lack.\n\nThat double-digit swing from removing a single fact is bigger than most retrieval papers see from doubling context size, but the comparisons here are against the authors' own benchmark and baseline, so the exact percentages deserve independent replication before anyone rewires a production coding agent around them.","[\"ai\",\"coding-agents\",\"retrieval\",\"benchmarks\"]","2026-09-18T04:00:00.000Z","2026-09-19T06:33:54.487Z","2026-09-19T06:34:06.417Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Attribute the findings to the actual arXiv paper (arXiv:2609.20050, 'The Missing Complement: State-Conditioned Minimal Sufficient Evidence for Coding Agents') instead of leaving it as unnamed 'researchers' with no institution or citation.","resolved","ai",[30,32,33,34],"coding-agents","retrieval","benchmarks",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.20050",0,{"sections":41},[42,46,51,56,61,66,71,76,80,85,90,95,100,105],{"name":43,"slug":30,"count":44,"latest_published_at":45},"AI",4114,"2026-09-19T18:33:46.000Z",{"name":47,"slug":48,"count":49,"latest_published_at":50},"Security","security",673,"2026-09-19T17:30:00.000Z",{"name":52,"slug":53,"count":54,"latest_published_at":55},"Policy","policy",345,"2026-09-19T19:22:20.000Z",{"name":57,"slug":58,"count":59,"latest_published_at":60},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":62,"slug":63,"count":64,"latest_published_at":65},"Hardware","hardware",156,"2026-09-19T11:00:00.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":70},"Science","science",129,"2026-09-19T19:45:00.000Z",{"name":72,"slug":73,"count":74,"latest_published_at":75},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":18},"Dev Tools","dev-tools",78,{"name":81,"slug":82,"count":83,"latest_published_at":84},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"General","general",42,"2026-09-18T22:35:10.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":106,"slug":107,"count":108,"latest_published_at":109},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]