Claude Opus 5: An evaluation review & metrics benchmarks
Blog post from Sonar
Sonar’s evaluation of Claude Opus 5 Thinking on 4,441 Java tasks found that it improved functional correctness over Opus 4.8, achieving an 88.6% pass rate on 544 HumanEval and MBPP tasks with executable tests versus 82.9% for its predecessor. Static analysis showed lower bug density, vulnerability density, and cognitive complexity per line of code, including substantial reductions in blocker-level security issues and reliability findings, while exception handling, API-contract, control-flow, resource-leak, and type-safety issues also declined. However, Opus 5 produced 2.3 times more code and 3.6 times more output tokens, resulting in 2.7 times as many total findings despite improved per-line bug and vulnerability rates. Code smell density, overall issue density, and cyclomatic complexity increased, with collection and generics issues, cryptographic misconfiguration, concurrency defects, and naming or documentation concerns emerging as notable areas for review. The analysis concludes that Opus 5’s gains in correctness and several quality measures are accompanied by greater output volume and maintainability workload, making automated verification especially important.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 2 | 5,068 | 1,020 | 229 | -34% |
| AI Guardrails | 1 | 551 | 150 | 54 | +6% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.