Home / Companies / Baseten / Blog / Post Details
Content Deep Dive

Testing Llama 3.3 70B inference performance on NVIDIA GH200 in Lambda Cloud

Blog post from Baseten

Post Details
Company
Date Published
Author
Pankaj Gupta, Philip Kiely
Word Count
1,033
Company Posts That Month
5
Language
English
Hacker News Points
-
Post removed?
No
Summary

The NVIDIA GH200 Grace Hopper Superchip is a unique architecture that combines an NVIDIA Hopper GPU with an ARM CPU via NVLink-C2C, promising advantages for AI inference workloads requiring large KV cache allocations. The GH200's high-speed interconnect allows offloading parts of the KV cache to abundant CPU memory, unlocking optimizations like prefix caching and KV cache re-use. In experiments serving Llama 3.3 70B on a single 96GB GH200 GPU, the superchip outperformed an H100 GPU by 32%, with performance gains coming from access to a larger KV cache rather than just higher VRAM bandwidth or identical compute profiles. The results suggest that the GH200 Superchip is well-suited for high-throughput deployments of models that wouldn't fit on standalone GPUs with similar VRAM profiles, and its unique architecture powers the GB200 Grace Blackwell Superchip, which promises to be extremely powerful for model inference and supports multi-node NVLink for serving larger models.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 5 3,220 466 154 -13%
Serverless 3 577 158 78 +5%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.