OpenHands trajectories with Qwen3 Coder 480B
Blog post from Nebius
Reinforcement learning (RL) showcases advanced results in software engineering tasks but faces significant infrastructure challenges, requiring complex MLOps workflows with asynchronous setups and fine-grained experimentation. In contrast, behavioral cloning and model distillation, which involve supervised fine-tuning, offer simpler alternatives. Rejection fine-tuning (RFT) is particularly effective by using successful trajectories from multiple solution attempts, circumventing the infrastructure demands of RL. The research contributes a dataset named nebius/SWE-rebench-openhands-trajectories, containing 67,074 agent trajectories from 1,823 Python repositories on GitHub, generated by the Qwen3-Coder-480B-A35B-Instruct model using the OpenHands framework. RFT checkpoints at two scales, 30B and 235B, show promising performance on benchmarks like SWE-bench Verified, indicating the potential of RFT in capturing high-quality behavior. The work also provides comprehensive documentation to ensure reproducibility in evaluations using the OpenHands framework, emphasizing the importance of detailed configuration management to avoid data leakage and infrastructure instability.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.