AI/ evolutionary-algorithms · llm-agents · program-synthesis · ai-research

Self-Evolving Code Gets a Memory for What Actually Worked

A new technique logs which program edits caused which performance changes, letting AI-driven code search converge faster and cheaper.

Researchers have found a way to make AI-driven program search remember its own cause and effect, not just its results.

Most LLM-guided evolutionary search tools work by mutating candidate programs and keeping only the final score, discarding the trail of which edit caused which change. The mutator model then has to guess, from a messy history, what actually helped. A new method called component-aware feedback fixes this by comparing each evaluated program to its parent, isolating exactly which components changed, and logging those changes alongside their metric impact in what the researchers call an attribution memory. It tracks each edit two ways: locally, against the immediate parent, and globally, against the original seed program. Tested on LLM reranking, a pipeline optimization task that trades off answer quality against serving cost, the method hit the best baseline's final quality using a median of one third the search budget, then kept improving to finish 7.2% higher on a held-out quality metric, and under a cost-aware setup it found pipelines that were more accurate while using 11% fewer tokens per query.

This matters because the bottleneck in self-evolving AI systems has never really been the mutation ideas, it has been the mutator's inability to learn from its own history efficiently. That's a bigger deal for smaller, locally run models, which have less capacity to infer cause and effect from cluttered context than frontier models do. Giving search a structured memory of what worked is a cheap fix with an outsized payoff, and it is the kind of unglamorous plumbing improvement that tends to compound across every project built on top of it.

It is a small idea: keep receipts, not just final scores. But most automated code search tools still don't, which says more about how young this field is than about how hard the fix was.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →