How Moyai ran 279,577 benchmark machines
Blog post from Boxd
Moyai uses Harbor and boxd virtual machines to run large-scale benchmarks of coding agents, helping its platform identify abnormal agent behavior and real production failures from traces collected through tools such as Langfuse, LangSmith, and OpenTelemetry. Replacing Modal sandboxes, which were costly, difficult to size, and inaccessible for debugging, boxd provides an isolated full Linux KVM VM with Docker and 100 GB of storage for each task, allowing unmodified Docker workloads, SSH access during execution, and automatic cleanup. Since August 2026, Harbor has created 279,577 machines for Moyai across 877 open-source benchmark tasks, with median boot times of 5 milliseconds and a peak of 36,335 machines in one day; the system also executed more than 5 million commands and uploaded 1.9 million files. Moyai plans to expand to more realistic and complex tasks to improve its ability to detect agent failures, while the Harbor boxd environment provider is being prepared for upstream integration.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| AI Agents | 1 | No monthly metrics for this publish month. | |||
| OpenTelemetry | 1 | No monthly metrics for this publish month. | |||
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.