Home / Companies / Endor Labs / Blog / Post Details
Content Deep Dive

Better models are making agent patch review more expensive, not less

Blog post from Endor Labs

Post Details
Company
Date Published
Author
Robert Haynes
Word Count
996
Company Posts That Month
17
Language
English
Hacker News Points
-
Post removed?
No
Summary

Endor Labs’ Agent Security League evaluated Anthropic’s Claude Fable 5.1 in Claude Code and OpenAI’s GPT-6 Astra in Codex on real historical security-fix tasks, finding that Fable led with 87.2% functional pass rate and 37.4% security pass rate, while Astra achieved 82.1% and 34.1%. Although both models represent substantial improvement over earlier generations, only about 42% of patches that passed functional tests also resolved the underlying security issue, indicating that a green CI suite is not a reliable indicator of secure agent-generated code. A Vyper compiler vulnerability illustrates how an apparently plausible implementation can preserve an exploitable flaw while passing all ordinary tests, because functional tests generally validate intended behavior rather than adversarial conditions. The evaluation also found that inference costs are increasingly driven by repeated reading of accumulated context during tool calls rather than by code output, making excessive turns, timeouts, and exploratory behavior expensive. Fable 5.1 delivered verified secure fixes at about $10 each, compared with roughly $19 for an earlier model, but declining compute costs may increase the more substantial cost of human review, since every functionally valid patch still requires security assessment and engineer review capacity does not scale as cheaply as model output.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.