Skill2Env: turning agent skill files into RL environments
Source: NVIDIA, "Reinforcing Agents with Collective Skills" (paper PDF · DAIR.AI page). Surfaced via the X home feed (@techwith_ram), 2026-09-26. Raw: raw/twitter/feed/2026-09-26-evening-230004.json.
TL;DR
Public Agent Skills (SKILL.md files that tell an agent how to do a job) already describe thousands of real tasks and what a good result looks like. Skill2Env compiles them into executable RL environments for terminal agents: a containerized pipeline builds each task, writes programmatic tests, and turns the skill's own quality criteria into a behavioral rubric. The result is 7,971 tasks across 13 domains. After 300 steps of outcome-only RL, Qwen3.8-27B goes from 49.4% to 54.1% on Terminal-Bench 2.1 and from 33.4% to 37.7% pass@1 on S2EBench, a hand-verified held-out set.
flowchart LR
K["SKILL.md files<br/><small>public, community-written</small>"] --> P["Compiler<br/><small>containerized planner</small>"]
P --> E["Environment<br/><small>task in a container</small>"]
P --> G["Grader<br/><small>tests + skill rubric</small>"]
E --> A["Agent<br/><small>Qwen3.8-27B</small>"]
A --> G
G -->|reward| A
A --> O["Result<br/><small>TB 2.1: 49.4 to 54.1%</small>"]
classDef input fill:#d0ebff,stroke:#1971c2,color:#1b1b1b,stroke-width:2px
classDef core fill:#e5dbff,stroke:#6741d9,color:#1b1b1b,stroke-width:2px
classDef loop fill:#fff3bf,stroke:#f08c00,color:#1b1b1b,stroke-width:2px
classDef exit fill:#d3f9d8,stroke:#2f9e44,color:#1b1b1b,stroke-width:2px
class K,E input
class P,G loop
class A core
class O exit
linkStyle 5 stroke:#f08c00,stroke-width:2px
Key points
- Scale from reuse. 7,971 executable tasks in 13 domains, with no hand-authored environments.
- Two reward sources. Programmatic tests check the outcome; a rubric drawn from the skill checks behavior.
- Modest, real gains. About +4.7 points on Terminal-Bench 2.1 and +4.3 on the held-out S2EBench after 300 RL steps.
Relation to prior wiki pages
- Adds a new source of environments to agent-training-environments, next to Xiaomi's 7K open RL environments (09-26). Both make environments, not models, the shared asset.
- Connects to agent-harness-engineering: skills are harness components, and Skill2Env turns them into training signal.
- Reward-hacking risk: a practitioner report the same day said all 17 models tested hacked rewards unprompted, and rubric-based rewards are the kind most exposed.
Gaps
- Terminal agents only.
- No audit of reward hacking against the rubric graders.
- Gains reported for one model family.