Home / Companies / Lambda / Blog / Post Details
Content Deep Dive

How to serve DeepSeek-R1 & v3 on NVIDIA GH200 Grace Hopper Superchip (400 tok/sec throughput, 10 tok/sec/query)

Blog post from Lambda

Post Details
Company
Date Published
Author
Luke Miles
Word Count
710
Company Posts That Month
5
Language
English
Hacker News Points
-
Post removed?
No
Summary

DeepSeek-R1 and v3 are being served on NVIDIA GH200 Grace Hopper Superchip instances, which provide a high throughput of 400 tokens per second. This is made possible by using 12 or 16 GPUs, depending on the required throughput. The model vLLM works better than Aphrodite for DeepSeek right now, and an update has improved its inference speed by roughly 40%. A script is provided to create instances, set up NFS caching, install Python 3.11, create a virtual environment, download models, install VLLM, and serve the model using ray. The guide includes a video showing inference speed with 64 parallel queries.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Serverless 3 577 158 78 +5%
Kubernetes 1 840 160 74 -30%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.