DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing
Blog post from Together AI
Benchmark results on 113 DeepSWE software-engineering tasks show that Claude Fable 5 has stronger first-attempt accuracy than DeepSeek V4 Pro 0813, scoring 69.7% versus 62.8% pass@1, but costs about 90 times more per rollout at $21.63 compared with $0.24. With repeated attempts, Pro matches Fable at pass@2 and leads at pass@4, while delivering vastly more solved tasks per dollar and producing failures that are more often near-correct rather than major misses. Fable performs particularly well on Rust, serialization, data modeling, and other exact-contract tasks, whereas Pro leads in TypeScript, stateful reactivity, and concurrency and durability. Because the models succeed on substantially different tasks, with a low per-task correlation of 0.39 and combined coverage of 107 of 113 tasks, the analysis recommends running Pro first and escalating failed, test-verified results to Fable. This cascade reportedly reaches 82.7% accuracy at $8.28 per solved task, exceeding Fable alone’s accuracy while costing less than half as much per task.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.