[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-method-trains-ais-search-and-reasoning-systems-together":10,"sections":45},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":34,"tags":35,"sources":40,"feedback":44,"feedback_at":22,"cost_usd":44,"total_tokens":44},8474,"new-method-trains-ais-search-and-reasoning-systems-together","New Method Trains AI's Search and Reasoning Systems Together","New research shows training a retriever before the reasoning model, not after, yields bigger gains on multi-hop question answering.","A new reinforcement learning method called BRIDGE trains an AI's search tool and its reasoning model together, instead of just tuning the model and treating search results as a fixed input.\n\nResearchers built BRIDGE to close what they call an \"information-credit gap\": when an AI agent gives a wrong answer because it retrieved bad evidence, most training setups blame the language model, not the retriever. BRIDGE treats retrieval and reasoning as a joint optimization problem, and the team found the two steps are order-sensitive - adjusting the retriever before the reasoning policy produces bigger reward gains than doing it the other way around. They built a memory-efficient method to solve that ordering without heavy compute overhead. Tested on seven open-domain question-answering benchmarks with 3B and 7B parameter models, BRIDGE beat the strongest baseline's multi-hop accuracy by 9.6 and 3.4 exact-match points, respectively.\n\nThe finding that matters here isn't just the accuracy bump - it's the order-sensitivity itself. Most retrieval-augmented agents are trained as if the search tool and the language model are interchangeable pieces tuned in whatever sequence is convenient. This result suggests that sequence is not incidental; it changes how much reward the whole system can extract from training.\n\nThat's a narrow, technical result, not a new AI capability. But if the order-sensitivity finding holds up outside this paper's benchmarks, it's a cheap fix - a training-schedule tweak - that other teams building retrieval-augmented agents could adopt without redesigning their architectures.","[\"agentic ai\",\"retrieval-augmented generation\",\"reinforcement learning\",\"llm research\"]","2026-09-30T04:00:00.000Z","2026-09-30T06:00:43.821Z","2026-09-30T06:00:48.302Z","published",null,[24,30],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The claim that BRIDGE 'scored best on medical QA benchmarks' cites a performance result with no comparison figures or margin (unlike the multi-hop EM-point figures given elsewhere) and never names the baseline it beat on those seven QA benchmarks — add the actual numbers or name the baseline so the claim can be verified.","resolved",{"id":31,"reviewer":26,"round":32,"reason":33,"status":29},"editor-r2",2,"The medical QA claim still lacks the baseline name or comparison numbers the earlier note demanded — merely flagging it as unverifiable doesn't satisfy that; either dig up the actual figures\u002Fbaseline or cut the claim, and don't let the piece's final paragraph be nothing but that caveat — close with a proper wrap-up.","ai",[36,37,38,39],"agentic ai","retrieval-augmented generation","reinforcement learning","llm research",[41],{"name":42,"url":43},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.36505",0,{"sections":46},[47,50,54,58,63,68,73,78,83,88,93,98,103,108],{"name":48,"slug":34,"count":49,"latest_published_at":18},"AI",5028,{"name":51,"slug":52,"count":53,"latest_published_at":18},"Security","security",780,{"name":55,"slug":56,"count":57,"latest_published_at":18},"Policy","policy",417,{"name":59,"slug":60,"count":61,"latest_published_at":62},"Deals","deals",284,"2026-09-29T21:00:00.000Z",{"name":64,"slug":65,"count":66,"latest_published_at":67},"Hardware","hardware",194,"2026-09-29T13:16:04.000Z",{"name":69,"slug":70,"count":71,"latest_published_at":72},"Science","science",154,"2026-09-28T13:19:18.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Consumer Tech","consumer-tech",142,"2026-09-29T18:38:03.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":82},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":84,"slug":85,"count":86,"latest_published_at":87},"Dev Tools","dev-tools",89,"2026-09-29T17:15:00.000Z",{"name":89,"slug":90,"count":91,"latest_published_at":92},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":94,"slug":95,"count":96,"latest_published_at":97},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":99,"slug":100,"count":101,"latest_published_at":102},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":104,"slug":105,"count":106,"latest_published_at":107},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":109,"slug":110,"count":111,"latest_published_at":112},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]