A new technique lets AI coding agents find and fix the exact instruction that broke a task, instead of guessing.
Researchers describe SkillMorph, a system that compares successful and failed runs of an agent skill - the reusable instruction files that guide code agents through tasks - across repeated trials. It tracks what changed between revision cycles to flag "suspicious" actions, then uses those flagged steps to locate the specific lines in the skill file responsible for the failure before generating a fix. On two benchmarks, SWE-Skills-Bench and CannBot, skills revised this way beat both the originals and four existing skill-evolution methods on trial-level accuracy and execution consistency. An AI operator-development team also used SkillMorph for automated kernel generation, where it produced six skill-revision pull requests that were accepted.
That evidence trail is the real contribution. Most "self-improving agent" tools today rewrite instructions based on a vague summary of what went wrong, which is how you get fixes that patch the wrong thing or quietly break something else. Linking a specific trajectory step to a specific edit is standard practice in regular software debugging - it is new here only because skill files, unlike code, rarely get that kind of scrutiny.
Six accepted pull requests is a promising pilot, not proof this holds up on messier, less structured codebases. And a skill file that edits itself still needs a human reviewer, for the same reason auto-generated code does.