Home / Companies / Endor Labs / Blog / Post Details
Content Deep Dive

GPT-5.5 Sets a New Code Security Record with Cursor, not Codex, in Agent Security League

Blog post from Endor Labs

Post Details
Company
Date Published
Author
Henrik Plate
Word Count
1,016
Company Posts That Month
35
Language
English
Hacker News Points
-
Post removed?
No
Summary

Cursor + GPT-5.5 has set a new record for security correctness at 23.5%, surpassing the previous high of 22.9% set by Cursor + Opus 4.7, and marking it as the third agent-model combination to exceed the 20% threshold in security, albeit with a score still considered a failing grade. A noteworthy finding is the performance disparity observed when the same GPT-5.5 model is utilized across different harnesses, exemplified by Codex + GPT-5.5 achieving a lower functional correctness score of 61.5%, compared to Cursor's 87.2%, even though both scored similarly in security. This suggests that the harness plays a significant role in the performance of AI models, as evidenced by Codex's struggles with functional correctness in tasks where different harnesses succeed, possibly due to how it handles repository structures or test frameworks. Despite Codex's narrower functional-security gap compared to other combinations, its limitations are highlighted by its failure in specific tasks, such as with the planet-client-python case, where it uniquely failed due to an incorrect use of the opener argument. This underscores the conclusion that the choice of harness can significantly influence the outcome, sometimes more than the model's capabilities alone.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.