Home / Companies / Snyk / Blog / Post Details
Content Deep Dive

Benchmarking Secure-and-Functional Remediation and How Snyk Agent Fix Lifts Frontier-Model Fix Rates by over 14%

Blog post from Snyk

Post Details
Company
Date Published
Author
Stephen Thoemmes
Word Count
2,091
Company Posts That Month
11
Language
English
Hacker News Points
-
Post removed?
No
Summary

Snyk reports benchmarking vulnerability remediation across about 150 human-verified JavaScript, Java, and Python code samples, requiring models to produce first-attempt fixes that pass both hidden security tests and functional regression tests. Out-of-the-box Gemini 3.1 Pro, Claude Sonnet 4.6, and Claude Opus 4.6 achieved secure-and-functional fix rates between 72.4% and 74.6%, while adding Snyk Intelligence, which dynamically supplies relevant expert-written remediation examples from a database of more than 35,000 vulnerabilities, raised Sonnet to 82.5% and Opus to 85.4%. The largest reported improvement was in Python, where Opus increased from 64% to 88% with the added security context, supporting Snyk’s argument that targeted security knowledge can matter more than model selection for this task. The benchmark builds on security and software-engineering evaluation approaches such as SWE-bench and Vul4J, but the authors note limitations including its snippet-level scope, three-language coverage, modest sample size, limited repeat runs, and lack of formal statistical error bars for small differences. Snyk plans to expand language coverage and variance analysis, while related VulnBench work examines the reliability and repeatability of AI agents for vulnerability detection.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 2 5,068 1,020 229 -34%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.