DeepSeek V4 Pro: Tops SWE-Bench & Cuts Cost per Task by 3x vs. Fable 5
Blog post from Fireworks AI
Fireworks announced availability of DeepSeek V4 Pro 0813 through serverless APIs, dedicated deployments, and training tools, positioning it for long-horizon agent workloads with a 1 million-token context window, native tool calling, and retained reasoning history. In its evaluations against Kimi K3 and Fable 5, DeepSeek led on SWE-Bench Verified at 95.2% and LiveCodeBench at 92.0%, while Fable 5 performed better on Aider Polyglot and matched DeepSeek on Terminal-Bench. Fireworks reports that DeepSeek completed standard coding tasks at substantially lower cost per solved task than Fable 5, though it identified Java as a relative weakness that makes multi-model routing preferable for some codebases. An oracle-routing analysis combining DeepSeek and Fable 5 achieved higher aggregate accuracy and lower estimated cost than Fable 5 alone, although the company notes that oracle routing is a theoretical upper bound rather than a deployable policy. Fireworks also offers fine-tuning through SFT, DPO, and RFT, recommends using DeepSeek’s native harness or a minimal tool setup to avoid performance losses from interface mismatches, and cites internal security tests showing no refusals or output-length truncations across 840 adversarial runs.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Cost per task | 4 | 64 | 45 | 24 | -18% |
| Serverless | 3 | 783 | 217 | 99 | +1% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.