Home / Companies / Nebius / Blog / Post Details
Content Deep Dive

OpenHands trajectories with Qwen3 Coder 480B

Blog post from Nebius

Post Details
Company
Date Published
Author
-
Word Count
1,754
Company Posts That Month
11
Language
English
Hacker News Points
-
Post removed?
No
Summary

Reinforcement learning (RL) showcases advanced results in software engineering tasks but faces significant infrastructure challenges, requiring complex MLOps workflows with asynchronous setups and fine-grained experimentation. In contrast, behavioral cloning and model distillation, which involve supervised fine-tuning, offer simpler alternatives. Rejection fine-tuning (RFT) is particularly effective by using successful trajectories from multiple solution attempts, circumventing the infrastructure demands of RL. The research contributes a dataset named nebius/SWE-rebench-openhands-trajectories, containing 67,074 agent trajectories from 1,823 Python repositories on GitHub, generated by the Qwen3-Coder-480B-A35B-Instruct model using the OpenHands framework. RFT checkpoints at two scales, 30B and 235B, show promising performance on benchmarks like SWE-bench Verified, indicating the potential of RFT in capturing high-quality behavior. The work also provides comprehensive documentation to ensure reproducibility in evaluations using the OpenHands framework, emphasizing the importance of detailed configuration management to avoid data leakage and infrastructure instability.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.