Introducing: B200s and H200s on Modal
Blog post from Modal
Modal has made Nvidia B200 and H200 GPUs available on demand to all account holders without sales contact or quota requests, with pricing of $6.25 per hour for B200s and $4.54 per hour for H200s. The Blackwell-based B200 offers 180 GB of HBM3e memory, 8 TB/s memory bandwidth, and native FP4 Tensor Core support, while the Hopper-based H200 provides 141 GB of memory and 4.8 TB/s bandwidth, both exceeding the H100’s 80 GB capacity. These larger-memory GPUs can support large mixture-of-experts models that may not fit across eight H100s, while B200s can substantially improve memory-bound inference latency and computational throughput. Modal reports early benchmarks using vLLM and DeepSeek V3 showing B200s delivered 2.5 times faster median time-to-first-token and 1.7 times higher query throughput than H200s under specified conditions, although broader software optimizations are still developing. The platform positions its GPU service around rapid container startup, autoscaling to hundreds of GPUs, usage-based billing, and $30 in monthly free compute.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.