DeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and Coding
Blog post from Together AI
Benchmark results on 113 DeepSWE software-engineering tasks find that GPT-5.6 Luna is the stronger standalone model, achieving 67.2% pass@1 versus DeepSeek-V4 Flash 0731’s 53.3%, leading across seven of eight task domains, all tested programming languages, and completing work faster. DeepSeek costs about $0.10 per rollout compared with Luna’s $0.61, however, yielding substantially more solves per dollar and causing regressions in existing test suites less often when it fails. Its relative strengths are query and configuration work and more competitive Rust performance, while JavaScript and reasoning-intensive tasks show large deficits. Despite limited complementary task coverage, using DeepSeek as a first-stage model and escalating failed attempts to Luna produces a reported 78.9% success rate at $0.385 per task, exceeding Luna alone in both accuracy and cost efficiency.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.