DeepSeek V4 Pro: Tops SWE-Bench & Cuts Cost per Task by 3x vs. Fable 5
Blog post from Fireworks AI
Fireworks announced availability of DeepSeek V4 Pro 0813 through serverless APIs, dedicated deployments, and training tools, positioning it for long-horizon agent workloads with a 1 million-token context window, native tool calling, and retained reasoning history. In its evaluations against Kimi K3 and Fable 5, DeepSeek led on SWE-Bench Verified at 95.2% and LiveCodeBench at 92.0%, while Fable 5 performed better on Aider Polyglot and matched DeepSeek on Terminal-Bench. Fireworks reports that DeepSeek completed standard coding tasks at substantially lower cost per solved task than Fable 5, though it identified Java as a relative weakness that makes multi-model routing preferable for some codebases. An oracle-routing analysis combining DeepSeek and Fable 5 achieved higher aggregate accuracy and lower estimated cost than Fable 5 alone, although the company notes that oracle routing is a theoretical upper bound rather than a deployable policy. Fireworks also offers fine-tuning through SFT, DPO, and RFT, recommends using DeepSeek’s native harness or a minimal tool setup to avoid performance losses from interface mismatches, and cites internal security tests showing no refusals or output-length truncations across 840 adversarial runs.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.