A team of researchers has turned AI agent 'skills' into code that can actually stop an agent before it goes off the rails.
The new framework, called HASP (Harnessing LLM Agents with Skill Programs), replaces static text tips with what the researchers call Program Functions. These PFs watch for failure-prone states in an agent's task loop, then step in by rewriting the next action or injecting corrective context, rather than just suggesting what to do. HASP works three ways: as live intervention during inference, as structured supervision during training, or as a self-improvement loop where validated, teacher-reviewed PFs keep evolving. In tests on web-search, math-reasoning, and coding tasks, inference-time PFs alone lifted average performance 25% over a standard ReAct agent, and the training-plus-evolution version beat Search-R1 by 30.4%.
That gap between advisory and enforced is the real story here. Most agent-skill systems today amount to a cheat sheet bolted onto a prompt, one the agent can simply ignore mid-task. HASP's pitch is that codified guardrails catch and correct failures as they happen, which matters most on long, multi-step jobs like search chains and multi-file coding runs, where one bad move early on tends to cascade.
Executable guardrails are a sensible idea, but so far the numbers come from the authors' own benchmarks against their own baselines, not independent replication.