GPU Memory Snapshots: Supercharging sub-second startup
Blog post from Modal
Modal has introduced GPU memory snapshots in alpha to reduce cold-start delays for GPU-accelerated serverless workloads by preserving both CPU and GPU state, including model weights in VRAM, CUDA kernels, streams, contexts, memory mappings, and compiled artifacts. Building on its distributed filesystem cache and earlier CPU memory snapshots, the feature uses NVIDIA’s CUDA checkpoint/restore APIs alongside Modal’s gVisor-based checkpointing system to lock active CUDA processes, copy GPU state to host memory, release resources, and later restore the container state on compatible hardware. This removes prior requirements to load models into CPU memory before moving them to GPUs and avoids rerunning expensive operations such as torch.compile after startup. Modal reports cold-boot improvements of up to 10 times, citing examples such as NVIDIA Parakeet decreasing from roughly 20 seconds to 2 seconds and vLLM with Qwen2.5-0.5B-Instruct falling from 45 seconds to 5 seconds. Developers can enable the capability with an experimental option and load models directly onto GPUs within the snapshot lifecycle.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Serverless | 1 | 1,048 | 263 | 99 | +36% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.