Home / Companies / Baseten / Blog / Post Details
Content Deep Dive

Using fractional H100 GPUs for efficient model serving

Blog post from Baseten

Post Details
Company
Date Published
Author
Matt Howard, Vlad Shulman, Pankaj Gupta, Philip Kiely
Word Count
1,086
Company Posts That Month
8
Language
English
Hacker News Points
-
Post removed?
No
Summary

The text discusses the use of NVIDIA's Multi-Instance GPU (MIG) feature on H100 GPUs, which allows developers to split a single physical GPU into two or more virtual GPUs, each with its own memory and compute resources. This feature enables efficient model serving for machine learning models by providing equal or better performance compared to A100 GPUs at a 20% lower cost. The fractional H100 GPUs offer advantages such as support for FP8 precision, increased flexibility, and availability of GPUs across cloud providers and regions. The guide provides an overview of how MIG works, the specs of fractional H100 GPUs, and what performance to expect serving models on H100 MIG-based instances.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 6 2,357 311 115 -2%
Real-time 2 2,527 623 172 +6%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.