Hill-Climbing a SWE Agent: What 1,700 Coding Tasks Taught Kimi K2.7
Blog post from Surge AI
Surge AI reports that reinforcement-learning-only post-training of Kimi K2.7 on 1,700 coding tasks improved its Pass@1 performance across five external coding benchmarks by 4.7 to 20 percentage points, with gains transferring across three different agent harnesses and, in some cases, surpassing reported results from larger models. Analysis of paired trajectories suggests the trained model became more effective not by taking more steps, but by reducing median trajectory lengths while improving four software-engineering behaviors: retaining every specification detail, testing requirements rather than only supported implementation paths, preserving existing functionality against regressions, and constructing independent checks when no reference answer is available. The training data combined repository-change and terminal-deliverable tasks with hidden checks, while a reward function gave partial credit for completed requested behavior but assigned zero reward if existing tests regressed. Although the authors say they have not conducted ablations proving the reward design caused the changes, they argue that the results show high-quality RL environments can teach transferable engineering habits, helping a model move from solving the core task to reliably delivering complete, shippable software.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| AI Model Fine-tuning | 3 | 139 | 28 | 14 | -75% |
| LLM | 2 | 747 | 162 | 79 | -85% |
| AI Agents | 1 | 931 | 231 | 103 | -84% |
| Cost per task | 1 | 10 | 5 | 5 | -84% |
| Reinforcement learning | 1 | 17 | 7 | 5 | -82% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.