Home / Companies / Together AI / Blog / Post Details
Content Deep Dive

DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

Blog post from Together AI

Post Details
Company
Date Published
Author
Together AI
Word Count
1,730
Company Posts That Month
8
Language
English
Hacker News Points
-
Post removed?
No
Summary

A comparison on 113 DeepSWE software-engineering tasks, using four trials per model, found that GPT-5.6 Sol delivered stronger single-attempt performance than DeepSeek V4 Pro 0813, with 72.7% versus 62.8% pass@1, faster median completion times, fewer steps, and higher per-task reliability. DeepSeek V4 Pro 0813, however, cost $0.24 per rollout compared with Sol’s $8.37, achieved a higher pass@4 rate of 88.5% versus 85.8%, and provided far more solved tasks per dollar, making it more suitable for high-volume or retry-tolerant workflows. Sol led across most task domains and programming languages, particularly Python and Go, while Pro slightly outperformed it in Rust and stateful reactivity. Their failures also differed: Sol was more likely to introduce regressions into previously passing tests, while Pro more often produced near misses without breaking the existing suite. The reported optimal deployment strategy is a test-gated cascade that runs Pro first and escalates failed outputs to Sol, achieving 83.0% task resolution at an average cost of $3.35 per task, outperforming either model alone on the combined accuracy-cost measure.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.