AI/ ai-coding-agents · llm-research · automation · arxiv

Researchers Let AI Coding Agents Rewrite Their Own Rulebooks

RuleEvolve lets AI coding agents mutate and judge their own rulebooks, beating hand-tuned prompts on correctness, length, and token cost.

AI coding agents can now write their own rulebooks, and early results say the self-authored versions beat the human-written ones.

A new paper describes RuleEvolve, a framework that treats coding rules the way you'd treat code: something to test, mutate, and improve rather than write once and forget. The system keeps a pool of candidate rules, uses an LLM to generate variants of them, then scores each variant with a separate judge model and keeps the winners. Researchers ran it across two coding-agent frameworks, four backbone LLMs, and three benchmarks. RuleEvolve beat both manually engineered rules and existing prompt-optimization baselines on functional correctness, and also produced shorter code at lower token cost.

That's the part worth noting: most prompt-engineering work for agents targets the prompt itself, not the rulebook sitting underneath it. Coding agents typically lean on static style and behavior rules that someone wrote once and rarely revisits. If those rules can self-improve the way model weights do during training, the labor-intensive hand-tuning step that currently gates agent quality gets automated away.

The paper doesn't address how these evolved rules hold up outside benchmark conditions, or whether a judge model reliably catches subtle regressions a human reviewer would flag. Self-improving systems that grade their own homework tend to look great in papers and shakier in production.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →