A new evaluation method quietly makes AI benchmarks a lot less forgiving.
Researchers built HARDEN, a constrained evolutionary search that rewrites benchmark questions into harder versions without changing the expected answer. It mutates inputs along complexity axes tailored to each domain, while built-in checks confirm the new question still means the same thing, still sounds realistic, and can still be executed correctly. The team tested it on three benchmarks - FinQA, PubMedQA, and ContractNLI - across three sizes of Qwen3.5 models, from 35 billion to 397 billion parameters. Accuracy fell by 22.7% on average compared to the standard benchmark versions, and by as much as 49.9% compared to a single pass of HARDEN's own search process, meaning repeated search finds much harder cases than one round alone.
Benchmarks like these already shape how frontier models get ranked and marketed. If a fairly simple evolutionary search can knock a fifth off accuracy on average, and half in the best case, without changing what counts as a correct answer, that says the published scores were measuring something narrower than real-world difficulty. The effect shows up hardest in finance, medicine, and contract analysis - exactly the regulated, high-stakes domains where enterprises are being sold these models as reliable.
A benchmark score, it turns out, is a score against one specific set of questions - not a durable measure of what a model can actually do.