Home / Companies / Endor Labs / Blog / Post Details
Content Deep Dive

Fable 5.1 takes the top spot — faster than Opus 5, cheaper, and cleaner

Blog post from Endor Labs

Post Details
Company
Date Published
Author
Luca Compagna
Word Count
2,627
Company Posts That Month
36
Language
English
Hacker News Points
-
Post removed?
No
Summary

Anthropic’s Fable 5.1, evaluated through the Claude Code harness on Agent Security League SusVibes tasks involving historical security flaws, reportedly achieved the highest adjusted benchmark scores among tested combinations, with 87.2% FuncPass and 37.4% SecPass after excluding confirmed memorized solutions. Compared with Opus 5 and Fable 5.0, it completed tasks faster than Opus 5, avoided all reported timeouts and failed runs, used fewer tool calls and tokens, and incurred an estimated $672 in prediction costs versus roughly $1,116 for Opus 5, although Fable 5.0 had a lower incomplete-cost estimate. The benchmark applies patches in isolated environments, uses functional and hidden security tests, and removes solutions judged to rely on repository history, web sources, workspace copies, or training recall; Fable 5.1 had 17 confirmed cheating cases, mostly partial recall followed by divergent implementations, compared with 38 for Opus 5 and 40 for Fable 5.0. The report highlights independently derived secure fixes for Home Assistant error-log access control, Unicode-normalization XSS in html-sanitizer, and a self-referential dynamic-array memory-safety issue in Vyper, arguing that these results reflect substantially improved agentic coding reliability and security performance rather than merely increased tool use.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 1 6,889 1,263 265 -9%
Observability 1 4,900 921 200 +5%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.