Home / Companies / Modal / Blog / June 2026

June 2026 Summaries

9 posts from Modal

Filter
Month: Year:
Post Summaries Back to Blog
Anthropic’s Claude Science is presented as an AI workbench that lets life sciences researchers run computational workflows through conversations with Claude, using local sandboxed execution for smaller tasks and Modal’s cloud infrastructure for GPU-intensive or highly parallel workloads. The integration is designed to simplify workflows that traditionally require managing local machines, HPC clusters, cloud virtual machines, software environments, and job queues, while supporting elastic CPU and GPU capacity, shared storage through Modal Volumes, and reproducible dependencies through Modal Images. Example applications include protein structure prediction, enzyme mutation ranking across multiple models, genome-wide CRISPR screen design, virtual compound screening, and large-scale single-cell analysis. Claude Science is scheduled to launch June 30, and users can connect a Modal workspace through the app’s compute settings; Modal also plans up to $100,000 in compute support for selected academic projects in Anthropic’s 2026 Claude Science Cohort.
Jun 30, 2026 1,028 words in the original blog post.
Modal has introduced Servers, a low-latency HTTP, WebSocket, and gRPC serving option for regionalized, autoscaling application replicas, aimed particularly at interactive workloads such as LLM inference. Unlike Modal Web Functions, which provide built-in queueing and managed request lifecycles, Servers use a lighter request path that prioritizes speed and returns 503 errors when replicas are unavailable, shifting queueing and load-shedding responsibilities to applications. Modal reports reducing median request latency from 39 milliseconds for Web Functions to 6 milliseconds for Servers by avoiding control-plane lookups and network calls in the hot path. Its architecture combines AWS Network Load Balancers, Envoy edge proxies that terminate TLS and normalize traffic to HTTP/2, a custom Rust-based proxy called fprs for domain routing and replica load balancing, and compute-plane workers that relay traffic to user containers. The system maintains cached routing configuration synchronized from Spanner, supports autoscaling based on in-flight requests, provides proxy-level authentication to prevent unnecessary container activation, and includes traffic mirroring for uses such as A/B testing and continual learning. Servers are available through Modal’s SDK and also underpin its Endpoints product for streamlined LLM inference deployment.
Jun 25, 2026 2,361 words in the original blog post.
Modal has launched Auto Endpoints, a deployment offering designed to combine scalable infrastructure with low-latency AI inference while preserving user control over code. The company argues that speculative decoding, particularly when supported by Blackwell GPUs, SGLang, regional Modal Servers, and optimized kernels, is the most consequential latency optimization because decoding typically dominates end-to-end inference time. In a collaboration with AI agent platform Decagon, Modal reduced a voice-agent inference subsystem’s median latency from roughly 290 milliseconds by about 100 milliseconds, ultimately outperforming proprietary providers for the same workload by more than 60 milliseconds. The largest gains came from DFlash speculative decoding models customized through “mid-training” on task-specific synthetic data, which improved token-acceptance rates and removed about 40% of server-side decode latency beyond an already optimized baseline. Modal also describes improving communication routing, GPU prefill kernels, and inference-engine host overhead, while positioning Auto Endpoints as a self-service option for teams seeking both performance and operational flexibility.
Jun 24, 2026 2,040 words in the original blog post.
Modal Auto Endpoints offer a streamlined approach to managing large language model (LLM) inference, enabling teams to maintain control over their inference processes without sacrificing cost-performance or developer efficiency. Unlike traditional proprietary models, Modal emphasizes transparency by providing access to the underlying code, metrics, and performance data, allowing users to optimize and understand their inference engines fully. This service eliminates the need for extensive GPU reservations by using a pay-as-you-go model and leverages a robust autoscaling system to handle varying demand efficiently. The platform includes Modal Servers for ultra-low-latency routing, ensuring reliable performance with minimal overhead, and offers a declarative interface for easy configuration based on workloads and service level objectives (SLOs). By focusing on open-source development and providing comprehensive benchmarking tools, Modal positions itself as a forward-thinking solution, aiming to automate and enhance inference performance continually.
Jun 23, 2026 1,524 words in the original blog post.
Sandbox startup benchmarks often measure only container scheduling and boot time, but production user latency is more strongly affected by application-level initialization such as cloning repositories, installing dependencies, launching services, or loading browsers. Modal distinguishes the lifecycle stages of Created, Scheduled, Started, Ready, and In use, emphasizing that the interval between Started and Ready is frequently the longest and most workload-specific. To reduce perceived latency, it recommends maintaining warm pools of pre-initialized sandboxes that can be assigned immediately when requested, while Directory Snapshots enable project-specific state to be mounted onto otherwise generic pooled environments. Modal’s generally available Readiness Probes allow developers to define readiness through successful shell commands or accessible TCP ports, enabling sandboxes to be added to pools only after initialization is complete and providing a wait_until_ready() mechanism for on-demand workflows. The platform also adds ready events to its dashboard timelines, helping users compare infrastructure startup times with application setup time and identify bottlenecks across the complete startup lifecycle.
Jun 22, 2026 1,481 words in the original blog post.
Speculative decoding is presented as a lossless method for accelerating autoregressive LLM output generation by having a lightweight draft model propose multiple tokens that a larger target model verifies in parallel, with performance depending heavily on how many proposed tokens are accepted. The authors released DFlash draft models for several Qwen models, reporting an additional 5–20% speedup over existing DFlash baselines and claiming that Qwen 3.5 122B-A10B can exceed 1,000 tokens per second on a B200 node at low concurrency. They argue that speculative decoding can yield larger gains than many conventional inference-engine and kernel optimizations, particularly when draft models are customized using application-specific data, while open-source engines such as SGLang and vLLM have improved support for the technique. Using simulated acceptance lengths, a simplified mathematical model, and a hardware roofline model, the discussion shows that greater acceptance lengths can substantially increase throughput, though real-world gains are constrained by target-model load, drafter latency, batch size, model architecture, and hardware behavior. The proposed future direction includes adaptive draft-model training for changing workloads, improved drafter architectures and implementations, possible quality-speed tradeoffs through lossy speculation, and an iterative self-hosted inference cycle in which production data supports evaluation, custom speculators, and model distillation to reduce latency and cost over time.
Jun 19, 2026 3,830 words in the original blog post.
Modal announced regional Function routing in US West, EU West, and Asia Pacific South to reduce latency and support data residency, alongside an alpha VM Sandbox runtime with a real Linux kernel for Docker, eBPF, systemd, cgroups, custom mounts, and improved filesystem I/O. Sandbox updates include outbound domain allowlisting, volume subdirectory mounts, generally available readiness probes, configurable snapshot retention periods, and Named Images for sharing reusable container images across apps. The company also introduced a CLI for managing Modal Agent Skills, expanded role-based access control for Team and Enterprise plans through restricted Environments, and improved workspace limit visibility for GPU, container, and Sandbox usage. Modal released Python SDK 1.5.0 and JavaScript and Go SDK 0.8.0, while highlighting technical posts on serverless GPU startup improvements, reinforcement-learning infrastructure, multimodal inference optimization, and an Anthropic integration that runs Claude Managed Agent tool calls in Modal Sandboxes.
Jun 15, 2026 876 words in the original blog post.
Modal describes a series of contributions to FlashAttention-4 aimed at improving large language model inference, particularly memory-bandwidth-bound token decoding workloads with variable batch sizes, sequence lengths, and paged KV caches. The work focuses on changing parallelism from query-centric execution toward key/value splitting, adding support for irregular memory access through cp.async rather than TMA, and using CuTe DSL to develop specialized kernel variants. Added FP8 attention-input support improved throughput by up to 1.16× while reducing KV-cache memory requirements, while arbitrary KV page-size support improved compatibility and cache efficiency; a follow-up address-generation optimization raised small-page performance by up to 2.40×. Porting split-KV, or Flash-Decoding, increased throughput by up to 4.37× for small query lengths by distributing a query’s KV work across multiple GPU multiprocessors, though it requires a reduction kernel and can introduce small floating-point differences. Other changes reduce unnecessary work for short query sequences, improving single-token decode throughput by up to 3.06×, and extend grouped-query attention packing to irregular head ratios, producing a reported 2.92× improvement for one decode benchmark. The authors argue that flexible tile-level programming models and adaptive choices between memory-access and parallelization strategies are central to future high-performance attention kernels.
Jun 11, 2026 3,284 words in the original blog post.
Modal argues that infrastructure has become the principal bottleneck in reinforcement-learning post-training for large language models, as training requires coordinated multi-node model optimization, high-throughput inference rollouts, and vast numbers of isolated execution environments. It emphasizes that larger open-weight models offer stronger capabilities but make weight synchronization costly, particularly across nodes, while RDMA networking, LoRA, asynchronous training, and delta compression can substantially reduce transfer delays and idle GPU expenses. The company identifies recurring operational challenges including extensive integration code, limited cluster availability, and GPU underuse caused by slow or insufficiently scaled sandbox environments. Modal presents its platform as an abstraction layer for RDMA-connected GPU clusters, fault tolerance, autoscaling, and large-scale sandboxes, allowing teams to focus on rewards, environments, and algorithms rather than infrastructure. It also supports open-source training frameworks such as slime, verl, and OpenRLHF, contributes improvements upstream, and has introduced the experimental open-source Modal Training Gym to simplify defining RL jobs around a model, reward function, and environment with built-in observability and tutorials.
Jun 01, 2026 2,186 words in the original blog post.