Home / Companies / Sysdig / Blog / Post Details
Content Deep Dive

How attackers are jailbreaking LLMs with CTF framing and how to catch them

Blog post from Sysdig

Post Details
Company
Date Published
Author
Michael Clark
Word Count
2,445
Company Posts That Month
10
Language
English
Hacker News Points
-
Post removed?
No
Summary

In a detailed examination of recent cyber threats, Michael Clark, Director of Threat Research at Sysdig, reveals how attackers are manipulating Large Language Models (LLMs) to generate exploit code by framing their requests as Capture-The-Flag (CTF) challenges or Common Vulnerabilities and Exposures (CVE) hunts. This method, which exploits the models' safety training to bypass restrictions, has been observed by the Sysdig Threat Research Team (TRT) across multiple campaigns targeting applications like PraisonAI, LiteLLM, FastGPT, Open-WebUI, and Gotenberg. Attackers use this technique to disguise their requests as legitimate security exercises, prompting LLMs to produce code that can be used in real attacks. The pattern of using CTF framing is consistent among different operators, indicating a shift towards using LLMs for crafting exploits. This approach is significant because it leaves a detectable signature in various request fields, such as user-agent strings and passwords, which defenders can track to identify and mitigate these threats. The CTF framing method not only manipulates the attackers' own tools but has also been adapted to deceive victims' AI agents, highlighting a growing trend in AI-targeted attacks.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 28 6,292 1,205 252 -36%
MCP 8 7,755 862 214 0%
AI Coding Assistant 4 2,234 577 171 +12%
Agent sandbox 3 36 13 6 +200%
AI Agents 1 6,200 1,430 272 +10%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.