Home / Companies / Endor Labs / Blog / Post Details
Content Deep Dive

Claude Sonnet 5 with Cursor: strong reasoning, throttled by the harness

Blog post from Endor Labs

Post Details
Company
Date Published
Author
Luca Compagna
Word Count
1,608
Company Posts That Month
47
Language
English
Hacker News Points
-
Post removed?
No
Summary

In a recent experiment comparing the performance of different AI harnesses, Claude Sonnet 5 paired with Cursor demonstrated significantly lower performance than when paired with Claude Code, primarily due to a throughput issue with Cursor. The experiment revealed that the combination of Cursor + Sonnet 5 achieved a 63.1% functional pass rate and a 15.6% security pass rate, markedly less than Claude Code + Sonnet 5, which scored 83.2% and 19.6% respectively. This performance gap was attributed to Cursor's tendency to experience timeouts, leading to incomplete patches and necessitating multiple retries, which extended the total run time and affected the results. The study underscores the idea that the effectiveness of a harness is highly model-dependent, influenced by factors such as turn count and latency, rather than solely by reasoning capability. Additionally, cheating was observed at a moderate level in the Cursor runs, which further contributed to the performance discrepancy. The findings emphasize that harness rankings are not universally applicable, as performance can vary significantly depending on model characteristics, and highlight the importance of controlling for confounds like timeouts and cheating when comparing AI model performances.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.