[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-new-benchmark-pits-ai-against-humans-on-training-instincts":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},8955,"new-benchmark-pits-ai-against-humans-on-training-instincts","New Benchmark Pits AI Against Humans on Training Instincts","A new benchmark shows AI can outguess top researchers on training recipes, yet both AI and humans remain clueless about how data shapes results.","Researchers built a test for AI's training instincts, and the results are a mixed bag.\n\nThe ArchitectureIQ benchmark, described in a new arXiv paper, presents synthetic datasets alongside several candidate training recipes and asks the test-taker - human or AI - to predict which recipe will produce the best test score. Frontier language models hit about 76% accuracy, well above the 33% you'd get guessing randomly, and ahead of the best human researcher's 66%. The edge flips on architecture-only questions: top humans scored 65% there, while GPT-6 Astra managed just 38%. The paper also found that giving a weaker model like GPT-4o a distilled, 20-item knowledge base of training lessons closed most of the gap with Claude Opus 5.\n\nThe interesting finding isn't which model won. It's that nobody, model or human, seems to understand how dataset properties should change the right training recipe. The authors call data the real \"dark matter\" of AI research - the thing everyone uses but nobody has a working theory for. More reasoning compute didn't fix this either, suggesting the gap isn't a lack of thinking time but a missing vocabulary for talking about training the way math has one for proofs.\n\nBenchmarks for AI judgment calls keep multiplying, but this one's verdict is refreshingly humble: the machines are good guessers, not oracles, and so is everyone else.","[\"ai benchmarks\",\"llm research\",\"training data\",\"model evaluation\"]","2026-10-01T04:00:00.000Z","2026-10-01T12:41:00.304Z","2026-10-01T12:41:05.515Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Fix the dek: it claims AI 'lags' humans on data intuition, but the body states this blind spot is shared equally by both AI and human researchers, not a human advantage — reword to reflect that it's a mutual weakness, not a human-vs-AI gap.","resolved","ai",[32,33,34,35],"ai benchmarks","llm research","training data","model evaluation",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.39714",0,{"sections":42},[43,46,50,55,60,65,69,74,79,83,88,93,98,103],{"name":44,"slug":30,"count":45,"latest_published_at":18},"AI",5455,{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",805,{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",429,"2026-10-01T02:26:17.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",298,"2026-09-30T21:00:26.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":18},"Science","science",159,{"name":70,"slug":71,"count":72,"latest_published_at":73},"Consumer Tech","consumer-tech",149,"2026-09-30T22:57:11.000Z",{"name":75,"slug":76,"count":77,"latest_published_at":78},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":80,"slug":81,"count":77,"latest_published_at":82},"Software","software","2026-09-30T21:41:11.000Z",{"name":84,"slug":85,"count":86,"latest_published_at":87},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":89,"slug":90,"count":91,"latest_published_at":92},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":94,"slug":95,"count":96,"latest_published_at":97},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":99,"slug":100,"count":101,"latest_published_at":102},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":104,"slug":105,"count":106,"latest_published_at":107},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]