July 2026 Summaries
8 posts from Modal
Filter
Month:
Year:
Post Summaries
Back to Blog
Hugging Face’s technical timeline identified Modal as third-party infrastructure used during a recent agent intrusion, but Modal stated that its platform and sandbox isolation were not compromised. The incident occurred within a customer-operated application that publicly exposed an unauthenticated endpoint designed to compile and run internet-submitted code inside a Modal Sandbox, with the attacker’s execution limited to that customer’s container and no impact on other customers. Modal emphasized that public unauthenticated access is not the default and recommended authentication, IP allowlists, restricted outbound networking, and treating all user-provided code and inputs as untrusted.
Jul 29, 2026
168 words in the original blog post.
Moonshot has launched Kimi K3, a groundbreaking 2.8 trillion parameter multimodal model designed for extensive token context windows and native vision capabilities. With strategic partnerships with Modal and vLLM, K3 is available on day one with token-based pricing and an Auto Endpoint option for dedicated capacity. Kimi K3 employs a mixture-of-experts transformer architecture, enabling efficient scaling with active experts per token and a 1 million token context window. Its innovative features, such as Kimi Delta Attention and Attention Residuals, provide 2.5 times the scaling efficiency of its predecessor, K2. Moonshot's efforts in quantization-aware training and expert parallelism optimization ensure the model runs across diverse hardware platforms. Modal enhances K3's performance with a custom DFlash speculator, significantly speeding up decode time and boosting throughput on agentic tasks. Available through Modal's platform with a $30 monthly compute credit, Kimi K3 is poised to redefine open AI model capabilities with its impressive performance metrics and accessibility.
Jul 27, 2026
558 words in the original blog post.
Modal and Cognition have launched an alpha integration called modal-devin, an open-source library and CLI that runs Cognition’s Devin AI software engineer in user-controlled Modal Sandboxes while retaining Devin’s cloud-based reasoning. Devin Outposts address workloads that require specialized or private environments, such as GPU-backed ML tasks or data engineering systems with private Snowflake connections, rather than Cognition’s default managed sandbox. The integration uses an Outpost queue to connect Devin Cloud sessions to Modal infrastructure, where an orchestrator dispatches each session to an isolated sandbox configured with a team’s repository, dependencies, internal registries, credentials, and optional GPU resources. Modal provides fast startup times, custom images, workspace snapshots, scale-to-zero behavior, and support for large-scale sandbox deployment, allowing suspended sessions to resume without re-cloning repositories or reinstalling tools. Teams can initialize and customize separate Outposts for different environments with a single command, isolating their deployments, secrets, and sessions; the alpha became available July 21, 2026, for users with uv, Python 3.11 or later, Modal accounts, and Devin organizations enabled for Outposts.
Jul 21, 2026
582 words in the original blog post.
Modal rebuilt its sandbox platform to meet growing demand from AI agents and reinforcement learning workloads, claiming it can now run millions of concurrent sandboxes and create tens of thousands per second, including one million sandboxes in under a minute. The redesign replaces centralized coordination and strongly consistent databases in the critical creation path with horizontally scalable scheduling servers that use cached worker-state data, directly request container creation from workers, and rely on asynchronous Redis streams and durable metadata storage outside the latency-sensitive path. Modal argues that conventional systems such as Kubernetes and its prior Postgres-based architecture face scaling constraints from serialized scheduling, centralized state, and operations proportional to the number of containers or nodes. Developing the new system required extensive backend, worker-management, runtime, networking, observability, and feature rewrites, including changes to mitigate Linux kernel networking contention during large startup bursts. Benchmark results showed median sandbox startup times below half a second, though the company notes a longer latency tail under extreme concurrent starts and plans further container-startup optimizations; the platform is currently available as a beta opt-in before wider deployment.
Jul 16, 2026
1,954 words in the original blog post.
Inkling, released by Thinking Machines, is a general-purpose multimodal model designed to handle text, image, and audio inputs while generating text outputs. It features a mixture-of-experts transformer architecture with 975 billion total parameters, of which 41 billion are active, and employs a 1 million token context window with native audio and vision capabilities, prioritizing breadth over depth. The model is optimized for speed and efficiency through sparse experts and a unique local attention layout, where five out of every six attention layers use sliding window attention, enhancing performance by focusing on recent tokens. This design is supported by DFlash speculation, a technique that advances speculative decoding by generating whole blocks of tokens in parallel, maintaining speed and computational efficiency. Inkling is available on Modal as a Managed Endpoint with token-based pricing, promising improved interactivity and throughput on agentic workloads, and continues to evolve with advancements in local attention and speculative decoding techniques.
Jul 15, 2026
647 words in the original blog post.
Modal is presented as a virtual computer rather than simply an AI infrastructure platform, cloud provider, or inference service, running programs across resources supplied by many cloud providers. Its architecture is compared to that of a conventional computer: containers execute functions, sandboxes, or HTTP servers; container memory and disks support active workloads; object storage provides durable capacity; and caching layers, volumes, and ephemeral disks balance speed and storage scale. Modal converts container image definitions into reusable filesystems, similarly to how conventional systems compile source code into shared executable components, while its runtime schedules and isolates containers across underlying infrastructure. Its Input/Output Plane and Routing Plane handle function, sandbox, and HTTP-server communications, with Tunnels offering more direct TCP or UDP connections for performance-sensitive workloads. By virtualizing compute resources at a higher level than individual cloud platforms, Modal aims to aggregate, multiplex, and isolate demand across clouds while making large-scale compute more accessible to engineering teams.
Jul 09, 2026
1,469 words in the original blog post.
GPU cost decisions for AI workloads often involve trade-offs between serverless capacity, where users pay only while GPUs are active, and long-term reservations, which provide fixed capacity at discounted rates but can leave resources underutilized. Modal presents an interactive cost model centered on the peak-to-average demand ratio, arguing that serverless GPUs can be less expensive when demand spikes substantially above its average level, particularly if that ratio exceeds the discount offered by reservations. The model compares serverless costs based on GPU usage over time with reservation costs based on peak capacity maintained for the full contract period, while noting that inference, training, and agentic-development workloads may have highly variable demand. It acknowledges important limitations, including assumptions of perfect demand forecasting for reservations and instant autoscaling for serverless systems, and notes that organizations may combine reserved baseline capacity with serverless resources for bursts. Beyond direct infrastructure spending, the discussion identifies operational and development factors, such as autoscaling speed, service-level objectives, vendor complexity, and reduced differences between development and production environments, as relevant to GPU procurement choices.
Jul 06, 2026
1,613 words in the original blog post.
Multi-Token Residual Prediction (MRP) is a lightweight module for accelerating diffusion language models, which generate text by progressively unmasking tokens rather than predicting them left to right. Unlike naïve multi-token prediction approaches that attempt to reproduce an entire next-step probability distribution and degrade sharply over repeated steps, MRP predicts the smaller residual change between adjacent denoising-step logits, allowing it to operate accurately across multiple steps while keeping the underlying model frozen. The researchers apply MRP in static decoding as either a speculative drafter that preserves backbone output quality while reaching up to 1.56× throughput in SGLang, or a direct decoder that achieves higher speed with modest accuracy tradeoffs. In dynamic low-threshold decoding, MRP can reassess newly revealed tokens and remask those whose confidence declines after neighboring tokens are considered, recovering substantial accuracy lost through aggressive parallel unmasking, with reported gains of up to 22.6 points on HumanEval. Evaluated across SDAR models ranging from 1.7B to 8B parameters and benchmarks including GSM8K, MATH500, HumanEval, and MBPP, the approach uses a small two- or three-layer transformer and is presented as a flexible inference component that supports both quality-preserving acceleration and quality recovery.
Jul 01, 2026
2,790 words in the original blog post.