Teaching an AI agent to write its own instructions does not guarantee those instructions still work on a different task.
Researchers tested five methods that let agents rewrite their own skills, plus a simple one-shot skill, across six benchmarks using the same model, the same agent, and the same train/test split for every method. Of 21 skills that got measurably better on their training tasks, only 5 kept all of that improvement when tested on new tasks. Thirteen kept only part of the gain, and three kept none of it. Reading the skills that failed to generalize, the researchers found a pattern: agents baked in details that should have stayed flexible, like specific column names or output file paths, or turned a one-time bug fix into a blanket rule applied to everything.
That is a problem for anyone betting on agents that learn from experience and then reuse what they learned. An LLM judge could spot bad skills after the fact, matching the actual test rankings 86% of the time, but it could not reliably predict whether a single edit would help before running it. The researchers' fix, called Generalizable Skill Optimization, skips accumulating a fixed skill altogether: it keeps only a general guide and writes a fresh skill for each new task, and it topped all six benchmarks.
In other words, the shortcut everyone wants, an agent that gets smarter by saving its own notes, mostly saves notes that only work for itself.