Needles and haystacks: Can open-source & flagship models do what Mythos did?
Blog post from Semgrep
The study investigates the capability of several flagship and open-source models to identify security vulnerabilities, specifically those described in the Mythos blog post, and finds that none could successfully discover two key vulnerabilities without significant hints. The research highlights that discovery is much more challenging than verification, akin to the difference between undergraduate and PhD-level work. Through experiments evaluating models like Opus 4.6, GPT 5.4, Gemini 3.1-pro, Deepseek R1-0528, and Qwen 3.6-plus, it was observed that model performance varied significantly based on whether they were analyzing entire files or individual functions. No "magic bullet" was found for full-file assessments, but some models demonstrated better results in function-level analyses. The findings emphasize the importance of model diversity and suggest that using LLMs for vulnerability discovery is promising yet still developing, with iterative improvements ongoing. A notable conclusion is that LLMs paired with deterministic pre-filtering to identify key targets outperform naive whole-file prompts, indicating a strategic shift that could enhance vulnerability detection processes.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 4 | 5,932 | 1,046 | 223 | -2% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.