A new method automates the hardest part of training AI agents: building the test environments they learn from.
The approach, called Skill2Env, is detailed in the arXiv paper "Skill2Env: Capability-Oriented Environment Synthesis from Skills for General Agents" (arXiv:2609.33772), posted September 30, 2026. It starts from a "skill" - a packaged bit of domain knowledge, procedures, and tool-use instructions - and works backward to build a full task: objectives, environment facts, information boundaries, and a rubric-based evaluator to grade success. A companion technique called Iterative Task Hardening checks whether a solver agent finds a generated task too easy, then strengthens it. The researchers used 1,500 high-scoring solved trajectories pulled from these generated environments to fine-tune agents, and reported consistent gains across a range of agent benchmarks.
Building executable, appropriately difficult training environments has been a real bottleneck for agent post-training - hand-built benchmarks are expensive, and static ones get memorized fast. A framework that generates environments from existing skill libraries and self-corrects for difficulty could let labs scale agent training the way synthetic data scaled language-model pretraining, without a small army of environment designers.
The paper does not name the institution behind the work or break out which specific benchmarks improved by how much, so treat the 1,500-trajectory result as an encouraging signal, not a verified leaderboard win.