A new hill-climbing method just outperformed the elaborate evolutionary search systems that have become the go-to approach for squeezing extra performance out of large language models at test time.
The technique, called Hill Sampling, is almost embarrassingly simple: a frozen LLM repeatedly generates candidate programs, keeps the single best one found so far, and conditions every new attempt on that champion, with no archive, no diversity mechanism, and no updating of model weights. Researchers tested it on three verifiable problems, circle packing, sums and differences of finite sets, and Erdos' minimum-overlap problem, using three open-weight models. It set a new state of the art on circle packing among published methods and beat the AlphaEvolve reference result on the Erdos problem. Both the circle-packing and Erdos results took only hours of wall-clock time on eight H100 GPUs.
The more telling result is what did not work. The team also ran evolution strategies directly on LLM weights, in what they describe as the largest such experiment by parameter count, and found that actually learning the weights performed worse than running the identical method with the learning rate set to zero, essentially random noise. Plain repeated sampling beat both weight-learning setups, and Hill Sampling beat repeated sampling.
AlphaEvolve made headlines by pairing LLMs with elaborate evolutionary scaffolding to find new mathematics; this paper's quieter message is that much of that scaffolding may not have been earning its keep, and a frozen model iterating on its own best guess gets further, faster, and cheaper.