A new framework called Prompt2Skill builds and sharpens AI skills using nothing but a plain language task description.
The system takes a natural language prompt, derives a formal task specification from it, then either finds or generates a matching dataset. From there it runs a closed loop of reflective editing, testing and revising the skill until performance stabilizes. The researchers tested the approach across four domains - question answering, reading comprehension, spreadsheet manipulation, and mathematical reasoning - using both open-source and frontier models. Prompt2Skill beat a direct prompting baseline by an average of 10.8 points across all of them.
That gap matters because skill libraries today are mostly hand-written, expensive to produce, and tuned to whichever model version existed when someone wrote them. A skill built for one model's failure modes does not automatically help a newer or smaller one. Earlier automated approaches tried reflection-based tuning too, but still needed a curated, in-distribution training set before they could start - exactly the resource that is often missing for a brand-new task.
Whether a 10.8-point average improvement holds up outside four benchmark domains, or on tasks messier than spreadsheets and math problems, is the open question the paper does not answer.