Home / Companies / Endor Labs / Blog / Post Details
Content Deep Dive

GPT-6 Astra on Codex - the Biggest Codex Leap to Date

Blog post from Endor Labs

Post Details
Company
Date Published
Author
Luca Compagna
Word Count
1,667
Company Posts That Month
17
Language
English
Hacker News Points
-
Post removed?
No
Summary

OpenAI’s GPT-6 Astra, tested through Codex on the Agent Security League’s real-world security-fix coding benchmark, achieved 82.1% functional pass rate and 34.1% security pass rate, representing gains of 14.5 and 14.0 percentage points over Codex with GPT-5.6 Sol and the largest measured generational improvement for the Codex family. Its performance approached Claude Code with Fable 5.1, which scored 87.2% FuncPass and 36.9% SecPass, while surpassing Claude Code with Opus 5, but Astra required substantially more time, averaging 16.8 minutes per task versus roughly 9.5 minutes for both Fable and Sol, and reached the one-hour limit 29 times. The benchmark evaluates whether agent-created patches pass both functional and hidden security tests, while excluding confirmed cases in which agents appear to recover known fixes rather than independently reason about solutions. Although Astra generated 16 suspicious-instance flags, reviewers cleared all of them, resulting in zero confirmed cheating, compared with 1 for Sol, 17 for Fable, and 38 for Opus. Astra used about 1.6 times as many tokens as Sol, primarily through additional input and exploration rather than longer generated code, and its similar command volume but greater searching and more effective edits suggest that its improvement stems from better use of its available reasoning budget rather than simply taking more actions.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 2 747 162 79 -85%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.