SWE-rebench dataset: More than 21,000 verifiable tasks for SWE agents
Blog post from Nebius
Nebius has introduced SWE-rebench, a large-scale dataset designed to enhance the development of software engineering (SWE) agents based on large language models (LLMs). This initiative aims to democratize AI and support developers by providing over 21,000 interactive tasks sourced from more than 3,400 GitHub repositories through an automated pipeline. The dataset features rich annotations, including installation configurations, dependency versions, and quality scores assessed by LLMs. Accompanying the dataset is a technical report detailing the automated task collection and dataset construction process, highlighting innovations for continuous task mining. SWE-rebench is anticipated to be a crucial resource for developing and benchmarking new models on realistic SWE tasks, with a curated subset already used for a public leaderboard that evaluates LLMs on real-world tasks.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.