February 2025 Summaries
3 posts from Modal
Filter
Month:
Year:
Post Summaries
Back to Blog
GPU utilization can refer to three distinct metrics for neural-network inference workloads: GPU Allocation Utilization, the share of paid GPU time spent running application code; GPU Kernel Utilization, the share spent executing GPU kernels; and Model FLOP/s Utilization (MFU), the share of a GPU’s theoretical arithmetic throughput used for useful model computation. Allocation utilization is constrained by capacity procurement, overprovisioning, provisioning latency, and operational setup, while kernel utilization falls when data transfers, logging, model loading, CPU-side scheduling, or kernel-launch overhead leave GPUs idle. MFU is the most fundamental but difficult measure, since even continuously active GPUs may be limited by inter-GPU communication, memory bandwidth, low arithmetic intensity, or inefficient kernels rather than floating-point capacity. Suggested improvements include faster and more automated provisioning, profiling traces to identify idle CUDA streams and host bottlenecks, batching or CUDA Graphs to reduce scheduling overhead, optimizing communication and memory use, and relying on high-performance libraries such as CuBLAS, PyTorch, and vLLM. The discussion notes that reported allocation utilization is often below 70% at peak demand, that specialized platforms may target over 90% aggregate allocation utilization, and that large-scale training MFU commonly remains far below theoretical hardware limits.
Feb 24, 2025
3,021 words in the original blog post.
The GPU Glossary, an interlinked collection of concise resources on GPU programming, has been released on GitHub under a Creative Commons BY 4.0 license. Created to address the limited availability of accessible, cross-stack GPU education and reference material, it aims to provide human-friendly documentation for the complexities of high-performance computing. Community interest following its initial release prompted requests for corrections, new sections, eBook conversion, and other contributions, leading its creators to open the source material for easier collaboration and reuse with attribution. The repository includes pre-populated issues for planned additions such as matrix-multiplication kernel examples, coverage of Thread Block Clusters, and a script to compile the glossary into a single Markdown file.
Feb 21, 2025
338 words in the original blog post.
Modal has refreshed its logs dashboard with improved filtering, introduced live container profiling and debug shells for diagnosing running workloads, and made cloud region selection available across all plans to support cost control and data residency needs. Client updates add OIDC authentication for cloud bucket mounts and variable-length arguments for Modal functions, methods, and entrypoints. The company also published a DeepSeek-R1 deployment example, announced general availability for its secure code-execution Sandboxes, described memory snapshotting that can reduce cold-start times by 2.5 times, and promoted a March 6 open-source LLM demo event with Mistral.
Feb 14, 2025
385 words in the original blog post.