AI agents that design algorithms just got better at teaching themselves new tricks.
Researchers have built a training method called Population-Curated Policy Optimization, or PCPO, that lets an AI agent improve by learning from the best algorithms it has already generated, instead of needing a huge library of outside examples. They tested it on a narrow, high-stakes problem: designing learning-rate schedules for chip placement in electronic design automation. Trained on just four chip examples, the resulting model beat established in-context evolutionary methods like OpenEvolve and ShinkaEvolve across sixteen chip test cases, and performed competitively against closed-source frontier models such as GPT-5.5 despite using a much smaller 8-billion-parameter base model. On a separate test, PCPO-designed GPU kernels ran an average of 8.27x faster than PyTorch's default Eager-mode baseline.
Most algorithm-design agents today rely on in-context evolution: feeding an LLM its own past attempts in a prompt and asking for something better. That approach tends to stall once the easy improvements run out. PCPO instead keeps a curated population of past solutions and updates the model's actual parameters, internalizing domain knowledge that no public dataset contains while cutting inference costs through prompt distillation.
The chip and GPU benchmarks are real, but they're also exactly the kind of narrow, well-defined optimization problems that make for a flattering case study. Whether PCPO holds up on messier, less benchmark-friendly design tasks remains unknown.