[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-framework-sharpens-how-ai-agents-find-bugs-in-code":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},8214,"new-framework-sharpens-how-ai-agents-find-bugs-in-code","New Framework Sharpens How AI Agents Find Bugs in Code","A preprint says a new framework, SemNav, sharply improves how often AI coding agents find the right file, mainly by cutting wasted context.","A new AI coding tool called SemNav says it can find the right file to fix a bug far more often than existing methods, without needing a bigger model.\n\nThe system, described in a preprint posted to arXiv (not yet peer-reviewed, and tested on one open-source model, Gemma 4B, across two benchmarks), tackles a specific bottleneck: LLM agents hunting through a codebase for the file or function tied to a bug report. SemNav pairs a deterministic search step that casts a wide net with an LLM agent that narrows the list as it gathers evidence, using a language server to trace how code pieces connect across files and compact summaries called Semantic Cards that describe what each piece of code does instead of forcing the agent to read raw source. On SWE-bench Lite and PLocBench, this pushed File Hit@10, the odds the correct file lands in the top 10 guesses, from 68.33% to 82.67%. On a separate benchmark, SWE-Explore, SemNav topped all seven evidence-quality metrics tested and lifted downstream issue-resolution rates from 44.00% to 52.33%.\n\nThe real story here isn't the accuracy bump, it's the context savings. Semantic Cards cut the working context agents need by 48.2% compared to reading full source files, which matters more than raw accuracy as agent workflows get billed per token.\n\nThat's the actual bottleneck holding back autonomous coding tools: not smarter models, but cheaper, more targeted ways of feeding them the right slice of a codebase.","[\"ai-agents\",\"code-search\",\"llm-research\",\"dev-tools\"]","2026-09-28T04:00:00.000Z","2026-09-28T17:11:43.887Z","2026-09-28T17:11:50.197Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The closing paragraph reads as a hedge that undercuts the headline's accuracy-boost claim (single-model, two-benchmark, non-peer-reviewed caveat) instead of landing on a resolved conclusion — move the preprint\u002Fscope caveat earlier into the body as context and end with a concrete, non-hedging closer.","resolved","ai",[32,33,34,35],"ai-agents","code-search","llm-research","dev-tools",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.31176",0,{"sections":42},[43,46,50,55,60,65,69,74,79,83,88,93,97,102],{"name":44,"slug":30,"count":45,"latest_published_at":18},"AI",4844,{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",762,{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",399,"2026-09-27T18:39:02.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",265,"2026-09-28T14:00:00.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Hardware","hardware",189,"2026-09-28T10:52:40.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":18},"Science","science",151,{"name":70,"slug":71,"count":72,"latest_published_at":73},"Consumer Tech","consumer-tech",135,"2026-09-26T14:30:00.000Z",{"name":75,"slug":76,"count":77,"latest_published_at":78},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":80,"slug":35,"count":81,"latest_published_at":82},"Dev Tools",84,"2026-09-26T04:20:58.000Z",{"name":84,"slug":85,"count":86,"latest_published_at":87},"Startups","startups",76,"2026-09-25T18:33:59.000Z",{"name":89,"slug":90,"count":91,"latest_published_at":92},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":94,"slug":95,"count":91,"latest_published_at":96},"General","general","2026-09-26T17:02:42.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Reviews","reviews",30,"2026-09-24T20:07:31.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]