One sandbox per rollout, or how labs run RL for agents in 2026
Blog post from Hugging Face
Agent reinforcement learning increasingly relies on dedicated, isolated computing environments in which each rollout can run code, browse, edit files, preserve state across long tasks, and receive verifiable rewards. Reviewing reports from major model developers between late 2025 and 2026, the article finds that labs commonly use large fleets of disposable or resumable containers and microVMs, sometimes scaling to hundreds of thousands of concurrent sandboxes, for coding, web research, office software, and specialized tasks such as GPU kernel optimization. Training environments are evolving beyond task simulators to include the full agent harness, either reconstructed as a controllable white-box system or monitored as an unmodified black box, while some labs use agents to generate and validate new tasks. Although RL methods and distillation techniques are increasingly shared, environment design, task data, verification systems, and large-scale orchestration remain largely proprietary and are presented as major sources of competitive advantage. Open-source projects including TRL, OpenEnv, slime, verl, Harbor, AgentENV, and Hugging Face Sandboxes are beginning to provide public alternatives across the task, interface, sandbox, and training layers, enabling smaller teams to experiment with agent training without frontier-lab infrastructure.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| GPT-6 Astra | 3 | No monthly metrics for this publish month. | |||
| Reinforcement learning | 2 | 17 | 7 | 5 | -82% |
| LLM | 1 | 747 | 162 | 79 | -85% |
| OpenClaw | 1 | 11 | 3 | 2 | -94% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.