Home / Companies / Modal / Blog / May 2025

May 2025 Summaries

5 posts from Modal

Filter
Month: Year:
Post Summaries Back to Blog
Modal has made Nvidia B200 and H200 GPUs available on demand to all account holders without sales contact or quota requests, with pricing of $6.25 per hour for B200s and $4.54 per hour for H200s. The Blackwell-based B200 offers 180 GB of HBM3e memory, 8 TB/s memory bandwidth, and native FP4 Tensor Core support, while the Hopper-based H200 provides 141 GB of memory and 4.8 TB/s bandwidth, both exceeding the H100’s 80 GB capacity. These larger-memory GPUs can support large mixture-of-experts models that may not fit across eight H100s, while B200s can substantially improve memory-bound inference latency and computational throughput. Modal reports early benchmarks using vLLM and DeepSeek V3 showing B200s delivered 2.5 times faster median time-to-first-token and 1.7 times higher query throughput than H200s under specified conditions, although broader software optimizations are still developing. The platform positions its GPU service around rapid container startup, autoscaling to hundreds of GPUs, usage-based billing, and $30 in monthly free compute.
May 30, 2025 695 words in the original blog post.
Modal has announced the acquisition of Twirl, a data orchestration company that helps teams develop, test, and deploy data pipelines. Twirl’s founders, Eric Hansander and Rebecka Storm, bring extensive data-platform, analytics, and machine-learning leadership experience from companies including Better, Zettle, Tink, and others. The companies describe their missions as complementary: Modal has built serverless infrastructure from the compute layer upward, while Twirl focused on data transformation and orchestration, with both aiming to simplify work for AI, machine learning, and data teams. Twirl’s team will join Modal’s Stockholm office to help expand its serverless infrastructure and European presence, as Modal broadens beyond AI inference into related use cases such as sandboxed code execution and large-scale batch processing.
May 28, 2025 327 words in the original blog post.
Modal Batch is a new job-processing interface and durable queue system designed to run large-scale, fault-tolerant batch workloads such as embedding documents, preprocessing audio, and preparing training data. Users define a Python Modal Function with optional hardware, image, retry, and storage settings, then use `.spawn_map` to distribute up to one million inputs across thousands of cloud containers, including GPU-enabled instances, with execution guaranteed for up to seven days. Compared with Modal’s earlier `.map` and `.spawn` methods, the service increases queue capacity 500-fold, simplifies concurrent job submission, and provides logs and metrics for individual inputs and aggregate jobs. Modal positions the product as an alternative to managing distributed infrastructure and orchestration systems, highlighting customers including Harvey, which reported a tenfold document-processing speedup, Suno, which uses it to scale GPU audio preprocessing, and Achira, which relies on batching and retries for scientific dataset preparation. Planned additions include function caching, programmatic job controls, and SDK support for JavaScript and other languages.
May 22, 2025 1,042 words in the original blog post.
Modal has upgraded its workspace-wide Dict key-value store with unlimited storage, durable data, a revised expiration policy that retains entries for seven days after their last read or write, and a `skip_if_exists` option for atomic conditional writes. These additions support LRU-like caching because frequently read values remain available, while conditional writes enable distributed locking and request deduplication across concurrent containers. The announcement demonstrates a caching pattern for an expensive web function in which the first request inserts a pending entry, starts the computation, and stores its eventual result, while simultaneous identical requests detect the existing entry and wait for or retrieve the same cached result. This approach can reduce backend load, avoid redundant function calls during traffic bursts, and improve response times, although users may need to update their Modal client to use the new locking-related flag.
May 20, 2025 852 words in the original blog post.
Modal uses a linear-programming-based resource solver to shield customers from volatile GPU markets while providing rapid, predictable access to large-scale compute capacity. The system evaluates real-time demand alongside cloud providers’ prices, availability, hardware performance, regional constraints, and CPU and memory requirements, then determines which instances to acquire or release at the lowest feasible cost. Unlike a simple price-based autoscaler, it must preserve spare capacity so containers can start in seconds despite cloud servers taking minutes to provision, while also handling users’ varied GPU preferences and rapidly changing workloads. Modal relies on Google’s GLOP solver, based on the simplex algorithm, and supplements it with heuristics that prune unnecessary instance options, prevent infeasible optimization problems, and keep decisions fast enough for production use. Background workers enact the solver’s recommendations and feed failed provisioning attempts back into future calculations as updated provider capacity limits, helping the platform exploit temporary price differences and maintain scalable service.
May 07, 2025 1,681 words in the original blog post.