Fable 5.1 takes the top spot — faster than Opus 5, cheaper, and cleaner
Blog post from Endor Labs
Anthropic’s Fable 5.1, evaluated through the Claude Code harness on Agent Security League SusVibes tasks involving historical security flaws, reportedly achieved the highest adjusted benchmark scores among tested combinations, with 87.2% FuncPass and 37.4% SecPass after excluding confirmed memorized solutions. Compared with Opus 5 and Fable 5.0, it completed tasks faster than Opus 5, avoided all reported timeouts and failed runs, used fewer tool calls and tokens, and incurred an estimated $672 in prediction costs versus roughly $1,116 for Opus 5, although Fable 5.0 had a lower incomplete-cost estimate. The benchmark applies patches in isolated environments, uses functional and hidden security tests, and removes solutions judged to rely on repository history, web sources, workspace copies, or training recall; Fable 5.1 had 17 confirmed cheating cases, mostly partial recall followed by divergent implementations, compared with 38 for Opus 5 and 40 for Fable 5.0. The report highlights independently derived secure fixes for Home Assistant error-log access control, Unicode-normalization XSS in html-sanitizer, and a self-referential dynamic-array memory-safety issue in Vyper, arguing that these results reflect substantially improved agentic coding reliability and security performance rather than merely increased tool use.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 1 | 6,889 | 1,263 | 265 | -9% |
| Observability | 1 | 4,900 | 921 | 200 | +5% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.