Home / Companies / Modal / Blog / May 2026

May 2026 Summaries

2 posts from Modal

Filter
Month: Year:
Post Summaries Back to Blog
Modal, a company focused on creating a cloud infrastructure tailored for AI workloads, has recently raised $355 million, bringing its valuation to $4.65 billion. The funding round was led by General Catalyst and Redpoint, with participation from new investors like Menlo, Bain Capital Ventures, and Accel, as well as existing investors who reinforced their commitment. Modal aims to address the limitations of traditional cloud systems for AI by providing a platform that supports a wide range of applications, including elastic inference, reinforcement learning, and dynamic agent runtimes. The company has seen significant growth and adoption of its Sandboxes, isolated environments for running AI-generated code, which are crucial for reinforcement learning and other AI applications. Modal's platform is used by various companies, such as DoorDash and Physical Intelligence, to enhance model ownership, real-time inference, and large-scale batch processing. As part of its future plans, Modal is focusing on improving low-latency inference, expanding the use of Sandboxes, and enhancing the support for agentic development, while continuing to contribute to the open inference stack.
May 21, 2026 1,015 words in the original blog post.
In the age of inference, the demand for large-scale neural network processing has led to the development of serverless computing solutions to handle variable workloads, particularly in AI applications. Modal has engineered a system to optimize the scaling of AI inference workloads on GPUs, reducing the startup time from tens of minutes to mere seconds through several key innovations. These include maintaining cloud buffers of idle GPUs, employing a custom filesystem for lazy loading of container images, and leveraging both CPU and GPU memory snapshotting to expedite process initialization. These advancements enable more efficient use of GPU resources, addressing challenges such as high peak-to-average demand ratios and startup latency, which are critical for maximizing GPU allocation utilization. Modal's approach allows for a truly serverless model, where capacity can be dynamically adjusted in response to demand without the need for over-provisioning, thus enhancing both cost-effectiveness and performance for applications like Reducto's document processing platform. This work not only aims to make AI-driven applications more efficient but also seeks to share insights and collaborate with the broader engineering community to further improve and expand these capabilities.
May 12, 2026 4,960 words in the original blog post.