Keeping 100k battles of untrusted agent code in their lane
Blog post from Lambda
In March 2026, Lambda conducted AgentBeats, a month-long AI agent security competition where teams submitted both attacker and defender agents to test their ability to manipulate or resist manipulation of a target LLM. The event involved 48 rounds and nearly 100,000 battles among 22 teams, utilizing up to four NVIDIA HGX H100 GPUs, with a unique infrastructure allowing dynamic GPU allocation without dropping battles. The system was designed to execute untrusted, adversarial code while maintaining fairness by preventing agents from accessing the internet, other battles, or retaining state across rounds. The infrastructure relied on independent per-GPU vLLM replicas and a single Caddy container to manage network traffic and load balance, ensuring a scalable and robust competition environment. Despite not being fully sandboxed, the setup effectively managed the competition's demands by dynamically adjusting resources, maintaining a fair playing field, and keeping operations on schedule. However, it was not equipped to handle sophisticated security threats such as container escapes or registry exploits, which were not deemed necessary for the scale and scope of the event. The competition highlighted the importance of a flexible, secure infrastructure capable of handling adversarial scenarios without compromising the integrity of the results.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Serverless | 7 | 722 | 229 | 93 | -29% |
| LLM | 3 | 6,942 | 1,215 | 234 | +11% |
| AI Agents | 1 | 5,827 | 1,275 | 245 | -5% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.