AI/ ai · llm agents · code generation · research

New AI Agent Decides Which Code Candidates Are Worth Testing

EvoAlloc is a resource allocation agent that learns during search, cutting evaluation costs up to 82 percent while boosting final performance.

A new research system learns on the fly which code candidates are worth evaluating, and which ones aren't.

Researchers describe EvoAlloc, a self-evolving resource-allocation agent built for LLM-based program evolution, where search systems generate many candidate programs and score them against test suites. Full evaluations are expensive, so most existing systems use a fixed allocation rule and end up burning compute on weak candidates while starving the promising ones. EvoAlloc instead periodically compiles its own search history into reusable experience, then uses that experience to revise how it doles out evaluation budget. It also runs a counterfactual check, occasionally evaluating candidates it had planned to skip, just to see what the allocator would have missed.

Across coding and agent-harness optimization benchmarks, EvoAlloc reached baseline-level performance with 59-82% fewer full evaluations and 61-89% fewer total LLM tokens, and it beat baselines by 8.7-12.0% when given the same evaluation budget. That is a direct lever on the real bottleneck in LLM-driven code search, which is usually evaluation cost, not model cost.

The numbers come from the paper's own benchmarks, and "self-evolving" is a flattering name for what is functionally a learned scheduling policy, but a scheduler that beats a fixed one by double-digit margins is still worth the compute bill it saves.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →