AI/ ai agents · machine learning · llm research · arxiv

Self-Improving AI Agents Need Curators Tuned to Their Model

A new research framework called EASE shows that AI agents perform better when the system curating their skills adapts to which model is actually using them.

AI agents that write their own playbooks are getting pickier about who reads them.

A new paper introduces EASE, a framework for curating the reusable "skills" that self-evolving AI agents accumulate as they work. The researchers found that skill curators trained around one model's behavior perform noticeably worse when a different model executes those skills later - a problem they call cross-executor degradation. EASE addresses this by training a single curator that tracks an online behavioral profile of whichever model it's paired with, then decides what skills to add, edit, or delete based on that profile. Tested across the ALFWorld, ScienceWorld, and WebShop benchmarks with models ranging from Qwen3-8B and GPT-OSS-120B to unseen models like Kimi K2.6, DeepSeek V4 Flash, and Gemini 3.5 Flash, EASE beat existing skill- and memory-based baselines without any per-model retraining.

Most self-improving agent research assumes the executor model stays fixed, but real deployments swap models often as cheaper or better ones arrive. EASE's results point to a quiet compatibility problem: a skill library tuned for one model can degrade performance when the model underneath it changes, even if nothing else in the pipeline moves. The efficiency gains are real too - the paper reports 34.5-41.0% fewer stored skills and 9.1-14.5% less inference-time token use.

It's still benchmark evidence from three simulated environments, not a live agent fleet, but it's a fair warning that a skill saved by one model isn't automatically useful to the next one you swap in.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →