Home / Companies / Baseten / Blog / Post Details
Content Deep Dive

Inference engineering for DeepSeek V4 Pro 0813

Blog post from Baseten

Post Details
Company
Date Published
Author
Model Performance Team
Word Count
672
Company Posts That Month
10
Language
English
Hacker News Points
-
Post removed?
No
Summary

DeepSeek V4 Pro 0813, a 1.7-trillion-parameter MIT-licensed open model, has been released with improved post-training for code generation and agentic tasks while retaining the base architecture of the earlier Preview version. Positioned near GLM-5.2 on Artificial Analysis’ intelligence index and described as less costly per task, it includes first-party benchmarks and an open-source coding harness built around a plugin-first design. Baseten has made the model available through its API and dedicated deployments with zero data retention by default, adapting its inference configuration to changing agentic-coding workloads through tuning of parallelism, KV-cache allocation, and prefill-decode workers. The release also introduces revised input and output formatting specifications and includes a DSpark speculator to support speculative decoding and improve token throughput, though specialized deployments may benefit from a custom-trained speculator. Baseten lists pricing of $1.32 per million uncached input tokens, $0.132 per million cached input tokens, and $3.96 per million output tokens.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Cost per task 1 64 45 24 -18%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.