A new framework called SkillMaster trains AI agents to write, edit, and choose their own skills instead of relying on hand-coded rules.
Most agent frameworks treat skills as a library someone else curates - a teacher model, a human, or an external module decides what the agent is allowed to learn. SkillMaster trains the agent to do that job itself. It reviews its own completed task trajectories and decides whether to propose a new skill, update an existing one, or leave it alone. Each candidate edit gets tested against related "probe tasks" to check whether it actually helps before it is kept. A technique the researchers call DualAdv-GRPO scores task-performing decisions and skill-editing decisions separately, so training one does not destabilize the other. On two benchmark environments, ALFWorld and WebShop, agents trained this way beat the best existing methods by 8.8 percent and 9.3 percent on task success rate.
The interesting part is not the score, it is the shift in what the agent is responsible for. Instead of an outside system curating a skill library, the agent diagnoses its own failures, patches its own procedural knowledge, and carries that judgment into future tasks while making only limited edits to its skill bank. That is a step toward agents that improve themselves rather than agents that are improved by someone else's pipeline.
ALFWorld and WebShop are controlled, text-based sandboxes, not the real world, so the harder question - whether self-editing skill banks hold up outside a benchmark - is still open.