agentic-systems · Tier 2

Agent Training Environments

Agent Training Environments

Concept page, created 2026-09-04. An environment for agent training is an executable workspace that can be re-queried into many verifiable tasks and that returns real execution feedback. It is structurally different from a trajectory, which is a single frozen demonstration that can be imitated once and never probed again.

This page exists because three papers in eleven days placed agent capability in the environment corpus rather than in the parameter count, which crosses this wiki's three-paper threshold for declaring a pattern. The distinction between environment and trajectory had been implicit in the harness and self-evolution literature; Terminal-Universe (09-04) states it in one sentence and it deserves its own page.

Why the distinction is load-bearing

A trajectory supports imitation. An environment supports reinforcement learning, verification, and re-querying, which means one environment yields an unbounded number of tasks while one trajectory yields exactly one. The field has an enormous surplus of trajectories, because every deployed coding agent logs them by the millions, and a shortage of environments, because building one has historically meant a human-designed pipeline converting a repository or a webpage into an executable task with a verifier.

That asymmetry sets up two independent scarcity problems, and the 09-04 pair addresses one each:

Problem Symptom Paper
Supply Environments are scarce because construction is expensive Terminal-Universe (09-04)
Difficulty The environments you have are too easy, so RL gets no gradient and the build cost is wasted Environment Evolution (09-04)

These compose and nobody has composed them. Terminal-Universe's output is exactly Environment Evolution's input: recover 37.3k workspaces, then evolve them generation by generation so they stay at the learnable frontier as the model improves. That is the clearest unclaimed experiment on this page.

The three results

Apodex 1.1 (08-25) named the axis. Its claim is that working capability (sustained verifiable progress toward a real objective) is developed along two axes that are explicitly not parameter scale: Environment Scaling, expanding the diversity and verifiability of executable file, search and code environments, and Agentic Coordination Scaling. A 35B model plus a locally deployable 35B Mini reaches the leading performance band on finance, research, math, coding and search. What it did not publish is a reproducible way to obtain the environments.

Terminal-Universe (09-04) supplies one, from an input every lab running a coding agent already has in volume. The insight is that a trajectory is an incomplete but authentic recording of its own environment: the tool-execution history leaks which files existed, what they contained before each edit, what the dependency graph was, and what the test harness expected. Replaying the recorded file operations in reverse restores each touched file to its pre-modification state, yielding a partial workspace; a completion agent then supplies the missing files and dependencies. On the recovered workspace it reconstructs the original intent task and synthesizes new ones, scaled along breadth (cross-workspace queries mined from directional dependency relations between related environments, which reproduces the change-library-A-break-consumer-B pattern that single-repository benchmarks structurally cannot test) and depth (a user agent turns a single-turn query into a multi-round session with iterative feedback). Output: 37.3k task-sufficient environments. SFT on Qwen3.5-27B: +11.9 points on Terminal-Bench 2.1, +13.8 on EvoCode-Bench v2 MT@4.

Environment Evolution (09-04) (Hunyuan Team, Tencent) attacks difficulty, and its critique of the existing fix is the most general idea on this page. Agent-environment co-evolution mines on-policy rollout failures to synthesize harder environments near the model's learnable frontier. That breaks on its own success: a curriculum sampling from the model's current error distribution has a supply problem built in, because the better the model gets the fewer errors there are to mine, and the ones remaining are increasingly idiosyncratic. So the curriculum's information rate falls exactly along the trajectory where it should rise, and the environments it produces encode this policy's weaknesses, which is why co-evolved curricula transfer poorly. The alternative raises difficulty off-policy, deriving three difficulty-increasing directions from the multi-turn learning objective itself and scheduling evolved generations across training, so generation 5 can be pre-built before the model finishes generation 2. Difficulty is validated by measuring rollout success across three architecturally unrelated frontier agents (Hy4 preview, Claude Opus 5, GPT-5.6 Sol) and rises monotonically for all of them. Long-horizon RL: +14.4 points on Qwen3.6-27B, +18.0 on Qwen3.6-35B-A3B.

What the pattern claims

For terminal agents in the 27B-35B scale band, the environment corpus is now the leading capability variable in the published record, and it is the cheapest one to improve. Three papers, three groups, eleven days, gains of 11.9, 13.8, 14.4 and 18.0 points, all without touching parameter count. That is a stronger and more actionable statement than the harness literature's equivalent, because an environment is a training-time asset consumed once, while a harness is an inference-time artifact paid for on every step.

A difficulty axis validated across unrelated frontier models is a more durable object than a policy-calibrated one. This is the most reusable single contribution across the three papers and it is independent of any RL result: it gives the field a way to say an environment is hard without reference to which agent is being trained.

Connections to adjacent pages

On self-evolving-agents.md, Environment Evolution's starvation argument is a third explanation for that page's oldest open result. Evo-Bench's early saturation, where autonomous harness evolution plateaus after a few cycles, has had two prior candidate causes: single-trajectory variance (Mendel Gödel Machine (08-12)) and the availability of self-editing evidence (AutoWorldModel-Bench (08-13), where agents improved an external artifact in 63 of 64 sessions). If a loop conditioned on its own observed failures starves by construction, the plateau is the curriculum rather than the operator or the ability, and the predicted fix is to derive harness-edit directions from the objective rather than from observed failures.

Against self-evolving-agents.md's skill-library cluster, this page describes the cheaper amortization of the same source material. That cluster runs trajectory → skill, distilling what worked into a reusable prose artifact injected at inference; Repo-to-Skill DISCO (09-03) does it from repositories. Terminal-Universe runs trajectory → environment. On SkillZip (08-12)'s accounting, where the recurring half of the bill is the injected artifact, converting a trajectory into an environment is strictly cheaper than converting it into a skill, because the environment is consumed during training and then free.

On Task-CoEvolve (08-25), the pairing suggests a missing method. Task-CoEvolve keeps a validation pool at the frontier by variance-weighted selection, concentrating evaluation where candidate harnesses disagree and matching full-set search quality at 80% fewer evaluations. Environment Evolution keeps a training pool at the frontier by objective-derived generation. Both answer "keep the tasks informative." The unwritten method is variance-weighted scheduling over evolved generations, so training samples the generation where the current policy is most uncertain.

On agent-benchmarks.md, this page carries a leakage risk that page should hold against it. Terminal-Universe mines environments from public trajectories produced by agents driven by benchmark-style prompts, then evaluates on Terminal-Bench 2.1. The training and evaluation distributions share an ancestor and no held-out-provenance split is reported. RealSWE (09-04) makes this sharper by measuring that benchmark-style prompts are 7% of the real distribution (problem-statement-only requests are 88% of real prompts and 7% of benchmarks). A corpus mined from benchmark-driven trajectories inherits the benchmark's prompt distribution.

Open questions

  • Authentic-versus-synthesized ratio. How much of a recovered workspace comes from replaying real file operations and how much from the completion agent's guesses? This is the number that decides whether trajectory mining is "environments for free" or "from-scratch synthesis with a better prior," and it is unreported.
  • What makes an environment "task-sufficient." 37.3k is a headline whose whole quality story lives in the admission filter. No complexity distribution, no verifier-reliability measurement, no estimate of how many carry a task current models actually fail.
  • Off-policy versus on-policy curricula under matched compute, measured late. The decisive experiment is both curricula, same seed environments, same budget, evaluated when co-evolution's failure supply should be thinnest. Not run.
  • Does an off-policy curriculum transfer across policy families? Environment Evolution validates difficulty across three unrelated agents and trains on one model family, so the central claim against co-evolution is demonstrated for difficulty and assumed for learning.
  • Build cost. Both 09-04 papers omit the cost of their own mechanism, and Environment Evolution's omission is the more pointed one because the motivating complaint is wasted build cost. A method justified on that ground owes a build-cost number.
  • Beyond the terminal. Whether the three difficulty directions and the trajectory-inversion trick generalize to browser, GUI and tool-calling environments depends entirely on how tightly they are tied to the terminal's file-operation structure.

Related pages

2026-09-10: trajectory to environment to learned environment

Two HuggingFace arrivals on the same day take opposite positions on how real the training environment needs to be, and the pair defines the axis this page should now track.

T1 makes the environment maximally real. A 122B mixture-of-experts model operates an actual shell in a cloud sandbox for 300+ tool-call turns per task, rewarded by executing each task's own verifier, with a dense process reward counting how many verifiers pass. Terminal-Bench 2.1 goes from 43.8% to 64.0%, and the training corpus is deliberately disjoint from the benchmark so the gain is transfer rather than fitting. Its infrastructure contribution matters beyond agents: rollout routing replay (R3) records the sampler's per-token expert choices at every MoE layer and replays them during training, which together with exact-token-id training cuts the train-inference logprob gap from 0.021 to 0.013. Without it, the router picks different experts at training time than at rollout time and the policy gradient is computed against a different function than the one that acted.

WMRL removes the environment entirely. Its diagnosis is that agent generation batches across a rollout set while environment execution occupies an exclusive sandbox and real machine time, so execution dominates RL cost and becomes the bottleneck as trajectories lengthen. Replace it with a learned world model, correct the resulting biased and noisy rewards with Online Debiasing and Inverse-Variance Denoising (both proven to strictly improve the convergence guarantee), and training accelerates 3-4x while 4B and 9B agents beat open-weight 48B and 120B agents on held-out benchmarks. It transfers to embodied VLA post-training too.

The unresolved question is the divergence between the learned reward and the executed verifier, and neither paper measures it. T1's correctness rests on the environment being real; WMRL's economics rest on it being simulated. That measurement is the most valuable missing experiment on this page.

This also extends Terminal Universe (09-04), which argued that reconstructing executable workspaces from recorded runs beats imitating those runs, since a workspace can be re-solved with new tasks while a recording can only be copied. The progression across two weeks is trajectory → reconstructed environment → learned environment, each step trading fidelity for throughput. And The Pulse #191 (09-10), which reports a CPU shortage caused by agents running tools, supplies the procurement reading: WMRL's 3-4x speedup is a CPU-shortage mitigation, not only an algorithmic one.

2026-09-25: environments have a CPU bill

  • Skill2Env (NVIDIA, surfaced on the X feed via alphaxiv): 3.4K public Agent Skills turned into about 8K executable terminal environments with programmatic tests and behavioral rubrics. 300 RL steps lift Qwen3.8-27B by 4.7 on Terminal-Bench 2.1. Community-written skills as a task distribution closer to real use.
  • The hardware cost of this whole page: the CPU shortage essay names RL environments as a main reason cloud CPUs are scarce. Every environment-scaling result (CodeMidas on this week's Kurate board, Skill2Env) is also a CPU-demand result.

2026-09-28: skills-to-environments, second instance in two days

SkillGym converts human-written skills into 2,756 environments with code checkers and fine-tunes on 8,364 successful trajectories; Qwen3.5-35B-A3B gains 19.10 points on Terminal-Bench 2.1, and the trained model without skills beats the base model with skills in context. With Skill2Env (09-27) (7,971 RL tasks, about +4.7 points), the SKILL.md corpus is now a recognized training-data source. One more independent instance makes it a pattern.

2026-09-29: AutoGym writes the verifier first, and publishes a build cost

AutoGym (Amazon AGI) generates task, environment and verifier together from a domain seed or past trajectories. Its blueprint-first step fixes the solution space and verification spec before building the environment, so every task is solvable by construction and no LLM judge is needed afterwards. Difficulty is set by explicit knobs (topology, depth, obfuscation, distractors): 10% of tasks land in the hard band at low settings, 39% at mid-to-hard, and an active curriculum shifts the knobs as the model improves. It answers this page's open build-cost question with a number: about $100 to $200 per 50-task batch at 20-way parallelism, 86% retained after repair. The Agent World Model baseline was mostly easy for Claude Opus 4.6. This is the third environment-generation result in three days after Skill2Env (09-27) and SkillGym (09-28), which crosses the pattern threshold: environment synthesis is now the default way labs scale agent RL data. No RL training gain was in the captured material, so the curriculum's resistance to the starvation problem raised by Environment Evolution (09-04) is untested.