Home / Companies / Semgrep / Blog / Post Details
Content Deep Dive

We benchmarked A LOT of models, here’s how they compare to Mythos

Blog post from Semgrep

Post Details
Company
Date Published
Author
Isaac Evans, Katie Paxton-Fear, Brenden Noblitt, Seth Jaksik
Word Count
1,612
Company Posts That Month
7
Language
English
Hacker News Points
-
Post removed?
No
Summary

A security benchmarking report evaluates Anthropic’s Mythos model in its Claude Security harness against open- and closed-weight models on 275 human-reviewed insecure direct object reference vulnerabilities across four codebases, using precision, recall, and F1 scores. Mythos achieved 80.0% precision but only 13.9% recall, identifying 20 of 144 confirmed vulnerabilities, placing fifth of 17 raw configurations for precision and fifteenth for recall; GLM 5.2 and GPT-5.6 Terra exceeded it on both metrics, while Claude Opus 5 achieved the highest raw recall at 38.2%. The comparison includes important methodological caveats, since Mythos was tested within a security-specific harness and evaluated with an internal LLM-based judge while other runs used deterministic matching, making model and harness effects difficult to isolate. Tests of GPT-5.6 Sol across three harnesses showed recall varying by 6.5 times, suggesting that prompting, context selection, and supporting tooling can influence results more than the underlying model. The report argues that security vendors should substantiate “Mythos-class” claims with reproducible recall measurements against labeled datasets, clarify evaluation methods and denominators, and continually rebenchmark because model performance and costs can change quickly.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.