How attackers are jailbreaking LLMs with CTF framing and how to catch them
Blog post from Sysdig
In a detailed examination of recent cyber threats, Michael Clark, Director of Threat Research at Sysdig, reveals how attackers are manipulating Large Language Models (LLMs) to generate exploit code by framing their requests as Capture-The-Flag (CTF) challenges or Common Vulnerabilities and Exposures (CVE) hunts. This method, which exploits the models' safety training to bypass restrictions, has been observed by the Sysdig Threat Research Team (TRT) across multiple campaigns targeting applications like PraisonAI, LiteLLM, FastGPT, Open-WebUI, and Gotenberg. Attackers use this technique to disguise their requests as legitimate security exercises, prompting LLMs to produce code that can be used in real attacks. The pattern of using CTF framing is consistent among different operators, indicating a shift towards using LLMs for crafting exploits. This approach is significant because it leaves a detectable signature in various request fields, such as user-agent strings and passwords, which defenders can track to identify and mitigate these threats. The CTF framing method not only manipulates the attackers' own tools but has also been adapted to deceive victims' AI agents, highlighting a growing trend in AI-targeted attacks.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 28 | 6,292 | 1,205 | 252 | -36% |
| MCP | 8 | 7,755 | 862 | 214 | 0% |
| AI Coding Assistant | 4 | 2,234 | 577 | 171 | +12% |
| Agent sandbox | 3 | 36 | 13 | 6 | +200% |
| AI Agents | 1 | 6,200 | 1,430 | 272 | +10% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.