A new training method called SkillGym gets AI agents to actually read the instructions before acting, and the fix is dumb-simple: practice with a grader watching.
SkillGym crawls skill guides from around the internet, keeps only the ones whose workflows can be run and checked offline, then uses a builder-reviewer pipeline to turn them into 6,800 difficulty-tuned practice tasks across four task types. Each task ships with a reference solution and an automatic verifier, so there is no guessing about pass or fail. The team collected 19,000 verified successful attempts and used them to finetune models ranging from 2 billion to 122 billion parameters. A 9-billion-parameter Qwen3.5 model trained this way beat an untrained 397-billion-parameter model on two of four benchmarks.
The interesting number here is not the benchmark win, it's the 28-percent-to-96-percent jump in how often trained agents bothered to open the relevant skill file before acting. Most agent failures blamed on "reasoning" are really agents ignoring the documentation sitting right in front of them. That is a cheap, structural problem, and this paper treats it as one: train the retrieval habit, not just the answer.
The part worth watching is whether the improvement generalizes to skills never seen during training, which the researchers say it does. If that holds up outside their own benchmarks, it is a more useful contribution than another leaderboard score - it is a recipe other labs can copy without touching a foundation model at all.