Home / Companies / Modal / Blog / May 2026

May 2026 Summaries

6 posts from Modal

Filter
Month: Year:
Post Summaries Back to Blog
Modal has introduced Role-Based Access Control for Team and Enterprise customers to help organizations govern both human and AI agent access as agents increasingly deploy code and manage infrastructure autonomously. The system uses Environments as isolated workspace subdivisions with separate applications, storage, secrets, usage data, and resource limits, supporting structures such as development, staging, production, or dedicated environments for individual developers, changes, customers, and agents. Restricted Environments provide stricter boundaries by making workspace members read-only by default, allowing only explicitly assigned users or service accounts to manage resources, limiting access to logs and secrets, and preventing resources from being accessed outside their environment. RBAC also integrates with audit logs to record activity, and Modal plans to expand the capability with agent spending and time limits, programmatic permissions, resource-level controls, and restricted environments as the default model.
May 27, 2026 535 words in the original blog post.
Modal, a company focused on creating a cloud infrastructure tailored for AI workloads, has recently raised $355 million, bringing its valuation to $4.65 billion. The funding round was led by General Catalyst and Redpoint, with participation from new investors like Menlo, Bain Capital Ventures, and Accel, as well as existing investors who reinforced their commitment. Modal aims to address the limitations of traditional cloud systems for AI by providing a platform that supports a wide range of applications, including elastic inference, reinforcement learning, and dynamic agent runtimes. The company has seen significant growth and adoption of its Sandboxes, isolated environments for running AI-generated code, which are crucial for reinforcement learning and other AI applications. Modal's platform is used by various companies, such as DoorDash and Physical Intelligence, to enhance model ownership, real-time inference, and large-scale batch processing. As part of its future plans, Modal is focusing on improving low-latency inference, expanding the use of Sandboxes, and enhancing the support for agentic development, while continuing to contribute to the open inference stack.
May 21, 2026 1,015 words in the original blog post.
Applied Compute develops custom enterprise AI agents using reinforcement learning to create “Specific Intelligence,” in which models are trained on proprietary company data, evaluated through tailored reward functions, and continually improved from operational use. Founded by contributors to OpenAI’s Codex and o1 efforts, the company has built applications such as DoorDash’s menu-to-storefront onboarding system and Cognition’s software bug-detection agent. Its approach depends on tightly integrated rollout, evaluation, and inference systems, with Modal providing infrastructure for ephemeral, high-fidelity production-system simulations, parallel grading workloads, and GPU-optimized inference. Modal’s fast container startup, isolated sandboxes, caching, retries, and serverless scaling are presented as helping Applied Compute reduce training bottlenecks, maintain reliability at high concurrency, and run complex reinforcement-learning loops. Applied Compute argues that as frontier models become more widely available, companies will increasingly differentiate themselves by owning the post-training processes, evaluation frameworks, proprietary data pipelines, and continual learning systems that adapt AI to their individual operations.
May 20, 2026 895 words in the original blog post.
Modal and Anthropic have introduced an integration that lets Claude Managed Agents execute tool calls in self-hosted Modal Sandboxes while Anthropic continues to host the agent loop and manage session state. The arrangement separates orchestration from code execution, an architecture Modal argues improves security boundaries, observability, failure handling, and scalability. Modal Sandboxes provide customizable container images, rapid cold starts, persistence through volumes and snapshots, usage-based burst pricing, network controls, and support for up to 100,000 concurrent sandboxes with configurable CPU, memory, and GPU resources. Companies including Mason AI, DoorDash, and Blend are evaluating or using related Modal and Claude tooling for secure enterprise automation, merchant-facing agent systems, and engineering support triage across complex environments. Modal has also released CLI-agent and Slack-bot reference implementations to help developers build production deployments using the combined platform.
May 19, 2026 1,027 words in the original blog post.
In the age of inference, the demand for large-scale neural network processing has led to the development of serverless computing solutions to handle variable workloads, particularly in AI applications. Modal has engineered a system to optimize the scaling of AI inference workloads on GPUs, reducing the startup time from tens of minutes to mere seconds through several key innovations. These include maintaining cloud buffers of idle GPUs, employing a custom filesystem for lazy loading of container images, and leveraging both CPU and GPU memory snapshotting to expedite process initialization. These advancements enable more efficient use of GPU resources, addressing challenges such as high peak-to-average demand ratios and startup latency, which are critical for maximizing GPU allocation utilization. Modal's approach allows for a truly serverless model, where capacity can be dynamically adjusted in response to demand without the need for over-provisioning, thus enhancing both cost-effectiveness and performance for applications like Reducto's document processing platform. This work not only aims to make AI-driven applications more efficient but also seeks to share insights and collaborate with the broader engineering community to further improve and expand these capabilities.
May 12, 2026 4,960 words in the original blog post.
A performance investigation of SGLang serving multimodal vision-language models found that its single-threaded scheduler was spending significant CPU time repeatedly reopening CUDA IPC shared-memory handles while processing image features, delaying GPU dispatches. Profiling with py-spy identified the costly PyTorch `_new_shared_cuda` calls within multimodal input hashing, where the same GPU memory pools were reconstructed for each tensor despite remaining unchanged. Developers replaced this repeated bookkeeping with a thread-safe Python dictionary cache that opens each IPC pool handle once and reuses the associated storage. In benchmarks of Qwen2.5-VL-3B-Instruct on an H100 GPU, the optimization raised throughput from 22.2 to 25.7 requests per second, reduced mean time to first token by 13.2%, lowered mean time per output token by 17.2%, and cut mean end-to-end latency by 10.6%, with tail latency improvements as well. The change applies to multimodal models using SGLang’s CUDA IPC transport, is included in SGLang v0.5.10, and can be enabled through the `SGLANG_USE_IPC_POOL_HANDLE_CACHE=1` environment variable.
May 04, 2026 1,362 words in the original blog post.