Home / Companies / Cursor / Blog / Post Details
Content Deep Dive

How we compare model quality in Cursor

Blog post from Cursor

Post Details
Company
Date Published
Author
-
Word Count
993
Company Posts That Month
14
Language
English
Hacker News Points
-
Post removed?
No
Summary

Cursor has developed CursorBench, an internal evaluation suite, to assess the performance of coding agents more accurately than public benchmarks. Built on real sessions from their engineering team, CursorBench measures various dimensions of agent performance such as solution correctness, code quality, efficiency, and interaction behavior, making it more aligned with real-world developer outcomes than traditional benchmarks. The suite addresses the limitations of public benchmarks, which often fail to differentiate between models effectively due to issues like misalignment with actual coding tasks, narrow grading criteria, and contamination from training data. CursorBench's tasks are derived from actual developer queries and solutions, ensuring they are relevant and challenging, with the scope of correctness evaluations doubling since its inception. By combining online and offline evaluations, Cursor ensures that model quality aligns with developers' practical experiences, allowing for the identification of regressions that offline methods might miss. As development work evolves to involve long-running agents, Cursor plans to adapt CursorBench to maintain its relevance and effectiveness, aiming to continue improving the agent experience in production environments.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Developer Experience 1 482 254 106 +18%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.