How to serve DeepSeek-R1 & v3 on NVIDIA GH200 Grace Hopper Superchip (400 tok/sec throughput, 10 tok/sec/query)
Blog post from Lambda
DeepSeek-R1 and v3 are being served on NVIDIA GH200 Grace Hopper Superchip instances, which provide a high throughput of 400 tokens per second. This is made possible by using 12 or 16 GPUs, depending on the required throughput. The model vLLM works better than Aphrodite for DeepSeek right now, and an update has improved its inference speed by roughly 40%. A script is provided to create instances, set up NFS caching, install Python 3.11, create a virtual environment, download models, install VLLM, and serve the model using ray. The guide includes a video showing inference speed with 64 parallel queries.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Serverless | 3 | 577 | 158 | 78 | +5% |
| Kubernetes | 1 | 840 | 160 | 74 | -30% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.