AI/ ai benchmarks · llm research · training data · model evaluation

New Benchmark Pits AI Against Humans on Training Instincts

A new benchmark shows AI can outguess top researchers on training recipes, yet both AI and humans remain clueless about how data shapes results.

Researchers built a test for AI's training instincts, and the results are a mixed bag.

The ArchitectureIQ benchmark, described in a new arXiv paper, presents synthetic datasets alongside several candidate training recipes and asks the test-taker - human or AI - to predict which recipe will produce the best test score. Frontier language models hit about 76% accuracy, well above the 33% you'd get guessing randomly, and ahead of the best human researcher's 66%. The edge flips on architecture-only questions: top humans scored 65% there, while GPT-6 Astra managed just 38%. The paper also found that giving a weaker model like GPT-4o a distilled, 20-item knowledge base of training lessons closed most of the gap with Claude Opus 5.

The interesting finding isn't which model won. It's that nobody, model or human, seems to understand how dataset properties should change the right training recipe. The authors call data the real "dark matter" of AI research - the thing everyone uses but nobody has a working theory for. More reasoning compute didn't fix this either, suggesting the gap isn't a lack of thinking time but a missing vocabulary for talking about training the way math has one for proofs.

Benchmarks for AI judgment calls keep multiplying, but this one's verdict is refreshingly humble: the machines are good guessers, not oracles, and so is everyone else.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →