Request, Aggregate, Bypass: How Attackers Can Evade LLM Safety Classifiers
Blog post from Crowdstrike
CrowdStrike’s Cyber Superintelligence Lab reports that advanced AI safety classifiers can reliably block direct harmful prompts but may have a structural limitation when evaluating requests independently rather than as part of a sequence. Its research found that harmful objectives can be divided into individually benign, legitimately framed requests and later combined by an unprotected model, enabling results across nine of ten MITRE ATT&CK-aligned offensive security categories tested. The study evaluated roughly 515 direct bypass approaches without finding a successful direct evasion, framing the issue as an architectural gap rather than a failure of classifier accuracy. It notes that Microsoft Research independently described a similar concept, called capability laundering, and argues that defenses should consider cross-request patterns, multi-model workflows, and knowledge transfer between more capable protected models and less restricted systems, while acknowledging that tracking such activity is difficult when queries are spread across providers or local models.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 28 | No monthly metrics for this publish month. | |||
| Real-time | 10 | No monthly metrics for this publish month. | |||
| AI Agents | 7 | No monthly metrics for this publish month. | |||
| AI Guardrails | 4 | No monthly metrics for this publish month. | |||
| Opus 5.5 | 4 | No monthly metrics for this publish month. | |||
| Local AI | 1 | No monthly metrics for this publish month. | |||
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.