[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-ai-research-agents-still-lose-to-a-dumb-statistical-baseline":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},6305,"ai-research-agents-still-lose-to-a-dumb-statistical-baseline","AI Research Agents Still Lose to a Dumb Statistical Baseline","A new benchmark finds AI research agents lose to a simple moving average on overall accuracy, though one model narrowly beats it at reacting to fresh evidence.","AI agents built to predict research trends are getting outperformed by a basic moving average.\n\nResearchers introduced Research Attention Prediction (RAP), a benchmark covering 278 AI\u002FML fields across 1,390 episodes. At each test point, an LLM agent searches a time-restricted slice of arXiv and predicts how paper shares across eight frozen research directions will shift over the next six months. The study separates two different measures: compositional accuracy, which scores the full predicted distribution, and future-specific updating, which isolates how well a model reacts to genuinely new information once historical activity is already known. On compositional accuracy, all four models tested performed worse than a simple exponentially weighted moving average (EWMA) baseline, even though enabling web search generally helped each model. On future-specific updating, only GPT-5.5 with search re-enabled slightly beat that same EWMA baseline.\n\nThe gap matters because it undercuts a common assumption about AI research agents: that search access plus a language model should beat dumb statistical extrapolation. Instead, the paper traces part of the failure to a design choice: models that carry forward a running state do better than models asked to forecast directly from scratch, partly because forecast-first approaches pull in less recent evidence during their searches.\n\nThere is a silver lining for anyone building these systems: fine-tuning Qwen3-4B on real outcomes lifted its forecast Spearman correlation by 0.105 on held-out fields it had not seen during training. Grand claims about AI-run research pipelines should probably wait for that kind of tuning to become standard, not the exception.","[\"ai\",\"ai-agents\",\"benchmarks\",\"research\"]","2026-09-11T04:00:00.000Z","2026-09-11T05:27:02.819Z","2026-09-11T05:27:14.746Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Resolve the internal contradiction between 'all four models tested underperformed the EWMA baseline' and 'the best-performing setup, GPT-5.5 with search re-enabled, only slightly beat the dumb baseline' — the source distinguishes compositional-accuracy results (where all four underperform) from a separate future-specific-updating measure (where GPT-5.5+search edges out EWMA), and the draft needs to make that distinction explicit instead of presenting both as the same comparison.","resolved","ai",[30,32,33,34],"ai-agents","benchmarks","research",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.10092",0,{"sections":41},[42,45,49,53,58,63,68,71,76,80,85,90,95,100],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",3507,{"name":46,"slug":47,"count":48,"latest_published_at":18},"Security","security",636,{"name":50,"slug":51,"count":52,"latest_published_at":18},"Policy","policy",338,{"name":54,"slug":55,"count":56,"latest_published_at":57},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":59,"slug":60,"count":61,"latest_published_at":62},"Hardware","hardware",153,"2026-09-09T15:12:32.000Z",{"name":64,"slug":65,"count":66,"latest_published_at":67},"Consumer Tech","consumer-tech",99,"2026-09-09T17:27:33.000Z",{"name":69,"slug":70,"count":66,"latest_published_at":18},"Science","science",{"name":72,"slug":73,"count":74,"latest_published_at":75},"Software","software",75,"2026-09-10T20:41:21.000Z",{"name":77,"slug":78,"count":79,"latest_published_at":18},"Dev Tools","dev-tools",70,{"name":81,"slug":82,"count":83,"latest_published_at":84},"Startups","startups",55,"2026-09-09T23:14:29.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"Gaming","gaming",43,"2026-09-10T12:18:06.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"General","general",41,"2026-09-08T01:57:23.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]