Home / Companies / Surge AI / Blog / Post Details
Content Deep Dive

Hill-Climbing a SWE Agent: What 1,700 Coding Tasks Taught Kimi K2.7

Blog post from Surge AI

Post Details
Company
Date Published
Author
-
Word Count
3,100
Company Posts That Month
3
Language
English
Hacker News Points
-
Post removed?
No
Summary

Surge AI reports that reinforcement-learning-only post-training of Kimi K2.7 on 1,700 coding tasks improved its Pass@1 performance across five external coding benchmarks by 4.7 to 20 percentage points, with gains transferring across three different agent harnesses and, in some cases, surpassing reported results from larger models. Analysis of paired trajectories suggests the trained model became more effective not by taking more steps, but by reducing median trajectory lengths while improving four software-engineering behaviors: retaining every specification detail, testing requirements rather than only supported implementation paths, preserving existing functionality against regressions, and constructing independent checks when no reference answer is available. The training data combined repository-change and terminal-deliverable tasks with hidden checks, while a reward function gave partial credit for completed requested behavior but assigned zero reward if existing tests regressed. Although the authors say they have not conducted ablations proving the reward design caused the changes, they argue that the results show high-quality RL environments can teach transferable engineering habits, helping a model move from solving the core task to reliably delivering complete, shippable software.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
AI Model Fine-tuning 3 139 28 14 -75%
LLM 2 747 162 79 -85%
AI Agents 1 931 231 103 -84%
Cost per task 1 10 5 5 -84%
Reinforcement learning 1 17 7 5 -82%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.