Home / Companies / Nebius / Blog / Post Details
Content Deep Dive

SWE-rebench dataset: More than 21,000 verifiable tasks for SWE agents

Blog post from Nebius

Post Details
Company
Date Published
Author
Ibragim Badertdinov
Word Count
224
Company Posts That Month
9
Language
English
Hacker News Points
-
Post removed?
No
Summary

Nebius has introduced SWE-rebench, a large-scale dataset designed to enhance the development of software engineering (SWE) agents based on large language models (LLMs). This initiative aims to democratize AI and support developers by providing over 21,000 interactive tasks sourced from more than 3,400 GitHub repositories through an automated pipeline. The dataset features rich annotations, including installation configurations, dependency versions, and quality scores assessed by LLMs. Accompanying the dataset is a technical report detailing the automated task collection and dataset construction process, highlighting innovations for continuous task mining. SWE-rebench is anticipated to be a crucial resource for developing and benchmarking new models on realistic SWE tasks, with a curated subset already used for a public leaderboard that evaluates LLMs on real-world tasks.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.