A new study says the fancier your AI heuristic-design pipeline gets, the worse it performs.
Researchers tested ten large language models across three combinatorial optimization problems, comparing frameworks that wrap LLMs in heavily hand-engineered evolutionary scaffolding against a stripped-down setup called SimpleEvol. Most existing systems use the LLM as a narrow tool bolted into fixed roles like crossover or mutation inside a hand-built evolutionary algorithm. The team introduced two new metrics, AHI for how hand-crafted a framework is and ICE for how efficiently it converts model intelligence into better heuristics. Across the board, frameworks with fewer human-authored rules scored higher on ICE, and SimpleEvol, which lets the model run the loop with almost no scaffolding, won by a wide margin.
This is a direct test of the bitter lesson, the idea that general methods scaling with compute beat hand-tuned ones, applied to a corner of AI research that had been quietly ignoring it. If the pattern holds elsewhere, it argues against the industry habit of building elaborate wrapper logic around LLMs and for just giving the model more autonomy and letting scale do the work.
The code is on GitHub, so anyone skeptical of less engineering is more can go check the claim rather than take the abstract's word for it.