GLM-5.3 delivers Opus 4.8-level cybersecurity results at a fraction of the cost
Blog post from Semgrep
Semgrep evaluated Z.ai’s newly released GLM-5.3 and several Grok 4.6 variants on its IDOR vulnerability-detection benchmark, which tests models’ ability to identify missing authorization checks in real open-source code using precision, recall, F1 score, and cost per confirmed finding. Although Z.ai reports that GLM-5.3 has advanced cyber capabilities and strong CyberGym exploitation results, it achieved a 23.8% F1 score in this benchmark, roughly matching Claude Opus 4.8’s 23.6% while costing $0.15 rather than $1.04 per true positive, though it unexpectedly trailed its predecessor GLM-5.2 and will be retested for variance. Grok 4.6 Exacto scored 35.5% F1, approaching Kimi K3 and Claude Opus 4.7 performance at substantially lower cost, while Claude Opus 5 and GPT-5.6 Luna remained the leading models overall at 65.6% and 48.0% F1, respectively. Across most models, precision remained relatively high but recall was low, indicating that flagged vulnerabilities were often valid but that many real flaws went undetected; the results suggest that newer lower-cost models are increasingly competitive economically but have not yet reached current frontier performance in raw detection quality.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.