[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-language-models-struggle-to-guess-research-ideas-from-citations":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},5267,"language-models-struggle-to-guess-research-ideas-from-citations","Language Models Struggle to Guess Research Ideas From Citations","A new benchmark shows AI models rarely guess a paper's core idea from its bibliography alone, though teamwork among models lifts accuracy roughly 2.4x.","A new benchmark says most AI models can't guess what a research paper is actually about just from its citation list.\n\nResearchers built Reconstruction, a blind test that hides the seed paper and any literature published around the same time, giving models only an anonymized, frozen bibliography and asking them to guess the paper's core idea. An independent LLM judge then checks whether the guess matches the real idea. Across six scientific domains and 643 papers, seven frontier models nailed the actual idea only about 3 to 15 percent of the time. The researchers then tried a more elaborate setup: four models cross-reviewing each other's guesses and competing in a Swiss-style tournament to pick the best hypothesis, using only the given references and no web search.\n\nThat multi-agent pipeline pushed match rates up to roughly 23 to 42 percent across all six domains - about 2.4 times better than the best single model working alone. It's a real data point in the debate over whether AI can meaningfully assist with hypothesis generation rather than just retrieval or drafting: on this evidence, one model reading a bibliography is closer to guessing than reasoning, while structured collaboration between models closes real ground.\n\nStill, the gap between a 23-42 percent match rate and genuine scientific insight is wide, and this is an unreviewed arXiv draft, not a peer-reviewed verdict.","[\"ai benchmarks\",\"multi-agent systems\",\"llm research\",\"scientific discovery\"]","2026-08-18T04:00:00.000Z","2026-08-18T12:30:18.002Z","2026-08-18T12:30:29.794Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The dek claims multi-agent review 'roughly doubles the hit rate,' but the body states the actual lift is 'roughly 2.4 times' — align the dek's language with the 2.4x figure instead of understating it as doubling.","resolved","ai",[32,33,34,35],"ai benchmarks","multi-agent systems","llm research","scientific discovery",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.16645",0,{"sections":42},[43,47,51,56,61,66,71,76,81,85,90,95,100,105],{"name":44,"slug":30,"count":45,"latest_published_at":46},"AI",3293,"2026-08-20T04:00:00.000Z",{"name":48,"slug":49,"count":50,"latest_published_at":46},"Security","security",435,{"name":52,"slug":53,"count":54,"latest_published_at":55},"Policy","policy",210,"2026-08-19T09:32:27.000Z",{"name":57,"slug":58,"count":59,"latest_published_at":60},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":62,"slug":63,"count":64,"latest_published_at":65},"Hardware","hardware",140,"2026-08-19T18:25:42.000Z",{"name":67,"slug":68,"count":69,"latest_published_at":70},"Consumer Tech","consumer-tech",95,"2026-08-18T16:05:00.000Z",{"name":72,"slug":73,"count":74,"latest_published_at":75},"Science","science",90,"2026-08-19T18:41:02.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":80},"Software","software",73,"2026-08-18T07:51:50.000Z",{"name":82,"slug":83,"count":84,"latest_published_at":18},"Dev Tools","dev-tools",69,{"name":86,"slug":87,"count":88,"latest_published_at":89},"Startups","startups",47,"2026-08-19T19:13:46.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"Gaming","gaming",41,"2026-07-09T04:00:00.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"General","general",33,"2026-08-18T22:18:13.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":106,"slug":107,"count":108,"latest_published_at":109},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]