Fractional GPUs and GPU Rightsizing: Stop Wasting Whole Cards
Blog post from Cast AI
Fractional GPUs, utilizing Multi-Instance GPU (MIG) partitions or time-slicing, offer a method to enhance GPU utilization in Kubernetes clusters by allowing workloads to share parts of a physical GPU, rather than monopolizing the entire card. This approach, in conjunction with GPU rightsizing, addresses the inefficiency highlighted in the Cast AI 2026 State of Kubernetes Optimization Report, which notes an average GPU utilization of just 5% in production environments. MIG provides hardware-based memory and fault isolation, ideal for multi-tenant inference, while time-slicing, which is applicable to any NVIDIA GPU, shares GPU access among pods through software without memory isolation. Rightsizing matches GPU and memory requests to actual usage, reducing waste by ensuring resources align with real workload demands. The Cast AI Workload Autoscaler further automates this rightsizing process by monitoring utilization, generating recommendations, and implementing changes to optimize efficiency. This combined strategy of fractional allocation and rightsizing has proven to elevate GPU utilization rates significantly, thereby reducing infrastructure costs in AI workloads.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Kubernetes | 18 | 1,260 | 165 | 75 | -41% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.