Inference engineering for DeepSeek V4 Pro 0813
Blog post from Baseten
DeepSeek V4 Pro 0813, a 1.7-trillion-parameter MIT-licensed open model, has been released with improved post-training for code generation and agentic tasks while retaining the base architecture of the earlier Preview version. Positioned near GLM-5.2 on Artificial Analysis’ intelligence index and described as less costly per task, it includes first-party benchmarks and an open-source coding harness built around a plugin-first design. Baseten has made the model available through its API and dedicated deployments with zero data retention by default, adapting its inference configuration to changing agentic-coding workloads through tuning of parallelism, KV-cache allocation, and prefill-decode workers. The release also introduces revised input and output formatting specifications and includes a DSpark speculator to support speculative decoding and improve token throughput, though specialized deployments may benefit from a custom-trained speculator. Baseten lists pricing of $1.32 per million uncached input tokens, $0.132 per million cached input tokens, and $3.96 per million output tokens.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Cost per task | 1 | 64 | 45 | 24 | -18% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.