Home / Companies / Modal / Blog / December 2025

December 2025 Summaries

2 posts from Modal

Filter
Month: Year:
Post Summaries Back to Blog
Modal describes its GPU reliability system for a globally distributed autoscaling pool that has exceeded 20,000 concurrent GPUs and launched more than four million cloud instances across major providers. The company reports substantial differences in hardware performance, boot reliability, thermal behavior, memory availability, and error rates among anonymized cloud platforms, using internal benchmarks and adjusted pricing to account for these factors. Its approach combines standardized, continuously tested machine images; lightweight checks at instance startup to limit scheduling delays; passive monitoring of logs, ECC errors, temperatures, and hardware slowdowns; and weekly active stress, diagnostic, and interconnect tests for longer-lived machines. Hosts that fail checks are drained and replaced or reinstalled rather than repaired in place, while dashboards and container logs expose GPU health signals to users. Modal also emphasizes support and rapid replacement capacity for failures that evade automation, arguing that GPU reliability remains a major operational challenge compared with CPU reliability.
Dec 28, 2025 1,977 words in the original blog post.
Mistral has launched Mistral 3, a family of open models with strong performance and customization capabilities, supported from day one on Modal, a platform facilitating deployment without complex infrastructure management. Modal offers advanced features such as GPU memory snapshotting, which significantly reduces cold start times for these models by nearly tenfold, from around two minutes to ten seconds. The Mistral 3 suite includes multimodal models with multilingual support, available in various sizes, with the Ministral 3 being particularly optimized for Modal's serverless infrastructure. This optimization makes it appealing for companies seeking a balance of intelligence and computational efficiency. To deploy Mistral 3 models effectively, developers can use Modal's serverless GPUs and distributed file system, which streamline the integration with vLLM servers. Modal's new GPU snapshotting feature, currently in alpha, enables faster cold starts by transferring GPU memory to CPU memory upon initialization, allowing for more cost-effective and responsive deployments.
Dec 02, 2025 521 words in the original blog post.