Home / Companies / Modal / Blog / Post Details
Content Deep Dive

How to price serverless GPUs

Blog post from Modal

Post Details
Company
Date Published
Author
-
Word Count
1,613
Company Posts That Month
8
Language
English
Hacker News Points
-
Post removed?
No
Summary

GPU cost decisions for AI workloads often involve trade-offs between serverless capacity, where users pay only while GPUs are active, and long-term reservations, which provide fixed capacity at discounted rates but can leave resources underutilized. Modal presents an interactive cost model centered on the peak-to-average demand ratio, arguing that serverless GPUs can be less expensive when demand spikes substantially above its average level, particularly if that ratio exceeds the discount offered by reservations. The model compares serverless costs based on GPU usage over time with reservation costs based on peak capacity maintained for the full contract period, while noting that inference, training, and agentic-development workloads may have highly variable demand. It acknowledges important limitations, including assumptions of perfect demand forecasting for reservations and instant autoscaling for serverless systems, and notes that organizations may combine reserved baseline capacity with serverless resources for bursts. Beyond direct infrastructure spending, the discussion identifies operational and development factors, such as autoscaling speed, service-level objectives, vendor complexity, and reduced differences between development and production environments, as relevant to GPU procurement choices.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Serverless 26 775 251 99 -24%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.