Meet SWE-rebench-V2: A multilingual, executable dataset for training Software Engineering Agents
Blog post from Nebius
SWE-rebench-V2 is an expansive dataset designed to enhance the training of autonomous software engineering agents using reinforcement learning by providing access to a large-scale, diverse set of executable tasks. This new iteration addresses the limitation of limited access to open-source data by utilizing a fully automated, language-agnostic pipeline to extract real-world software engineering tasks, resulting in over 32,000 executable tasks complete with pre-built Docker environments and coverage of 20 programming languages, including less commonly used ones like Lua and Scala. The dataset also features more than 100,000 additional tasks derived from pull requests and is designed to facilitate multilingual RL training by enabling research on cross-language reasoning and robust performance beyond Python-centric datasets. Each task includes a pre-configured Docker container for easy reproducibility, and the environments are automatically set up using an interactive agent that resolves dependencies. The tasks are quality-filtered and labeled using large language models (LLMs), with structured metadata that includes method signatures and problem descriptions to provide comprehensive training signals. Accompanying the dataset is a technical report that explains the extraction pipeline, filtering methods, and includes a diagnostic study assessing modern models on these tasks.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.