DeepSeek V4 Pro is Redefining Security Agent Economics
Blog post from Fireworks AI
Fireworks reports that DeepSeek V4 Pro 0813 performed strongly in CyberGym, a Berkeley-developed benchmark based on real vulnerabilities in open-source projects that requires agents to identify flaws, create working exploits, and patch affected code. Across tested tasks, the company says DeepSeek produced tool calls without refusals or output truncation, entered validation on 97.3% of a 697-task common cohort, and achieved a 53.7% reward rate at an estimated $2.50 per solved task, outperforming GPT-5.5 on both reward rate and cost but trailing Kimi K3’s 68.4% solve rate. The evaluation emphasizes that long, multi-step cybersecurity tasks can expose limitations of models that refuse or fail to generate usable artifacts, with Claude Opus 4.8 showing a low completion rate in the reported comparison. Fireworks argues that DeepSeek’s main weakness is generating initial proof-of-concept exploits rather than applying patches after validation, suggesting that improved prompting or scaffolding could raise its results, and positions open models as a practical option for high-volume vulnerability auditing and remediation workflows.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.