Home / Companies / Modal / Blog / July 2025

July 2025 Summaries

8 posts from Modal

Filter
Month: Year:
Post Summaries Back to Blog
Modal has introduced GPU memory snapshots in alpha to reduce cold-start delays for GPU-accelerated serverless workloads by preserving both CPU and GPU state, including model weights in VRAM, CUDA kernels, streams, contexts, memory mappings, and compiled artifacts. Building on its distributed filesystem cache and earlier CPU memory snapshots, the feature uses NVIDIA’s CUDA checkpoint/restore APIs alongside Modal’s gVisor-based checkpointing system to lock active CUDA processes, copy GPU state to host memory, release resources, and later restore the container state on compatible hardware. This removes prior requirements to load models into CPU memory before moving them to GPUs and avoids rerunning expensive operations such as torch.compile after startup. Modal reports cold-boot improvements of up to 10 times, citing examples such as NVIDIA Parakeet decreasing from roughly 20 seconds to 2 seconds and vLLM with Qwen2.5-0.5B-Instruct falling from 45 seconds to 5 seconds. Developers can enable the capability with an experimental option and load models directly onto GPUs within the snapshot lifecycle.
Jul 30, 2025 1,269 words in the original blog post.
AI code sandboxes are isolated, temporary computing environments designed to safely execute potentially untrusted code, a need that has grown as IDE-integrated LLM agents increasingly generate software for production use. Rooted in earlier isolation technologies such as Unix chroot, Java’s applet sandboxing, and browser-oriented tools like JSFiddle, modern sandboxes commonly use containers or virtual machines with additional protections such as gVisor, which intercepts system calls to limit access to the host system. They help contain coding errors, hallucinated dependencies, malicious prompt-injection outcomes, and other risks that could disrupt shared infrastructure or expose security vulnerabilities. Common applications include background coding agents, automated code reviews, chatbot code interpreters, reinforcement-learning evaluation pipelines, and AI-generated application platforms. Effective sandbox platforms typically emphasize fast startup times, strong kernel-level isolation, elastic scaling, configurable runtime environments, filesystem and memory snapshots, networking controls, and developer-oriented observability tools, while managed services can reduce the operational burden of building these capabilities internally.
Jul 24, 2025 1,990 words in the original blog post.
Open-weight automatic speech recognition models such as NVIDIA’s Parakeet and Canary and Kyutai’s STT have recently approached proprietary services in accuracy while offering high inference speeds and features including multilingual support, timestamps, and voice activity detection. Modal evaluated NVIDIA’s English-focused Parakeet and multilingual Canary models against a proprietary transcription API using a week of ESB benchmark audio, reporting comparable or slightly better error rates alongside configurations that were either more than 100 times faster or up to 200 times cheaper, with one optimized case transcribing a week of audio in about a minute for roughly one dollar. The comparison emphasizes end-to-end throughput, including cold starts, network transfer, and data movement rather than model execution speed alone, and focuses on large-scale batch workloads rather than low-latency streaming use cases. Key engineering approaches included distributing shuffled audio across workers for balanced workloads, sorting recordings by duration within GPU batches to reduce idle processing time, using sufficiently large GPU inference batches, parallelizing file downloads, and empirically selecting GPU types and worker counts based on cost-throughput trade-offs.
Jul 23, 2025 2,301 words in the original blog post.
Self-hosted, open-source large language model inference should be evaluated primarily in dollars per request rather than dollars per token, according to the text, because application users pay for completed interactions rather than token consumption. While token-based pricing suits API providers whose costs scale with input and output size, teams operating their own models must instead connect infrastructure decisions to user workflows, latency expectations, concurrent request volumes, and business value. The text argues that estimating acceptable response times requires considering time to first token and generation speed for typical requests, while capacity planning depends on requests per second and the number of concurrent requests each model replica can handle. Since self-hosting costs are ultimately determined by compute expenses over time and the replicas needed to meet performance targets, token counts become an internal operational metric rather than the main pricing or strategic measure. Framing costs per request is presented as a clearer way to assess whether an LLM feature produces sufficient revenue, conversion, or customer value to justify its infrastructure expense.
Jul 16, 2025 1,187 words in the original blog post.
Modal announced beta support for multi-node GPU training through a clustered decorator that co-schedules GPUs across hosts with 3.2 Tbps InfiniBand RDMA networking, aiming to enable near-linear training scalability. The platform also introduced serverless NVIDIA B200 and H200 GPUs, priced at $6.25 and $4.54 per hour respectively, with claimed LLM inference gains of two to four times over H100s. Modal Client 1.0 emphasizes API stability and includes migration guidance, while recent client updates add read-only volumes, secret support in shell sessions, cron timezones, and timestamped application logs. Additional resources cover benchmark comparisons between SGLang and vLLM, examples for running and optimizing FLUX image-generation models, and Quora’s use of Modal Sandboxes for high-volume code execution on its Poe chatbot platform.
Jul 11, 2025 473 words in the original blog post.
Modal has acquired Jamsocket, a backend platform for building synchronization engines for collaborative applications, and Jamsocket co-founders Paul Butler and Taylor Baldwin will join Modal. Jamsocket’s offerings include Plane, an open-source orchestrator; Y-Sweet, a document store and real-time synchronization backend; and ForeverVM, a service providing stateful Python REPLs for AI code execution. The companies cite their shared technical foundations, longstanding relationship in the New York developer community, and growing demand for AI-centric applications as reasons for the move. Jamsocket products will continue operating during a transition in which their core functionality is integrated into Modal, enabling customers to combine GPU inference, code sandboxes, and real-time backends on a single platform.
Jul 10, 2025 295 words in the original blog post.
Lovable, a no-code application-generation startup that reached $75 million in annual recurring revenue within seven months, uses AI to create, run, and visually edit applications from user prompts. Ahead of a June 2025 promotional event expected to sharply increase traffic, the company sought an alternative to its existing sandbox provider because of concerns about scalability, reliability, and dependence on a single vendor. After considering an in-house AWS and Kubernetes solution, Lovable adopted Modal Sandboxes, reducing its sandbox orchestration code from roughly 15,000 lines to 700 and using encrypted networking tunnels for application communication. During the event, concurrent usage increased 2.5 to threefold, with Modal operating more than one million sandboxes and reaching 20,000 concurrent sandboxes while supporting an estimated 250,000 app creations in 48 hours without requiring platform-team intervention. Lovable now relies on Modal for every application-generation session and plans to explore features such as snapshotting to further improve performance.
Jul 07, 2025 651 words in the original blog post.
Engineers behind qart.codes improved AI-generated artistic QR codes by treating scannability and visual appeal as separate, measurable objectives, prioritizing a 95% scan-rate service-level goal while avoiding aesthetic regressions. Building on ControlNet-guided Stable Diffusion techniques that preserve enough QR-code structure for error correction, they developed automated evaluations using QReader for scan testing and an aesthetic-rating model, then validated those tools against thousands of human judgments. After manually narrowing promising prompts, models, and generation parameters, they ran large-scale offline parameter sweeps and visualized tradeoffs between quality and scannability. When a single generation could not reliably meet the scan-rate target, they applied inference-time compute scaling by generating eight candidates in parallel, evaluating and ranking them by scan success and aesthetics, and displaying the best four. This approach achieved the scan-rate target with under-20-second p95 latency while improving image quality, illustrating how reliable generative-AI applications can be built through iterative eval development, scaled experimentation, and production-time selection rather than relying on compelling but inconsistent demos.
Jul 02, 2025 2,706 words in the original blog post.