February 2026 Summaries
4 posts from Modal
Filter
Month:
Year:
Post Summaries
Back to Blog
Modal presents its cloud platform as infrastructure for AI research workloads requiring scalable GPU access, reproducible environments, and isolation, positioning it as an alternative to local machines and conventional HPC clusters that may face resource limits, queues, and inconsistent performance. The post connects this mission to the author’s earlier work at Weights & Biases and argues that shared infrastructure can create a feedback loop in which AI research improves the systems used to conduct further research. It highlights three projects supported by Modal: KernelBench, a 250-task benchmark and evaluation environment for AI-generated GPU kernels; TTT-Discover, which used test-time training and large-scale GPU benchmarking to produce a contest-winning triangular matrix multiplication kernel; and RL-4-MLE, an early-stage effort to automate parts of the machine-learning research lifecycle through reinforcement learning. Researchers involved in these projects describe Modal’s on-demand GPU fleet, containerized software environments, hardware consistency, and ability to run hundreds of parallel jobs as important for reducing experiment times and improving evaluation reliability.
Feb 25, 2026
1,896 words in the original blog post.
Modal has launched Directory Snapshots in public beta, allowing users to capture and later mount specific Sandbox directories independently of the base image and the rest of a filesystem. Unlike full filesystem snapshots, this granular approach lets system dependencies, application code, and generated artifacts evolve on separate lifecycles, enabling updated base images to be combined with preserved project state. The feature supports workflows such as coding agents that retain user projects while receiving security or runtime updates, warm Sandbox pools that attach project-specific files only when assigned, and faster startup through prioritized loading of frequently accessed directories. Directory Snapshots persist for 30 days after last use and complement Modal’s existing filesystem snapshots for complete environments and memory snapshots for rapidly restoring active process state.
Feb 24, 2026
994 words in the original blog post.
Ramp built Inspect, an internal background coding agent powered by Modal Sandboxes, to provide engineers, product managers, and designers with fast, zero-setup access to full development environments and AI-assisted coding. Each isolated sandbox includes Ramp’s application services, development tools, browser-based visual testing, and integrations with systems such as GitHub, Slack, Buildkite, Sentry, and Datadog, allowing the agent to implement and verify changes end to end. Modal’s filesystem snapshots, refreshed every 30 minutes, enable new sessions to start within seconds from near-current repository states, while distributed functions, queues, and shared storage support hundreds of concurrent, collaborative sessions across multiple client interfaces. Ramp reports that Inspect now initiates roughly half of merged pull requests in its frontend and backend repositories, with more than 80% of Inspect’s own code written by the agent, and expects scalable parallel agent capacity to become increasingly important as coding models improve.
Feb 19, 2026
1,264 words in the original blog post.
Modal Research announces that Z.ai’s open-weight GLM-5, later upgraded on its free endpoint to GLM-5.1, offers frontier-level performance for long-horizon coding agents and systems-engineering tasks under an MIT license, positioning it as an open alternative to recent proprietary models. The post describes GLM-5’s roughly 700 GB FP8 size, mixture-of-experts architecture, sparse-attention approach, and multi-GPU deployment requirements, while reporting internal throughput of 30 to 75 tokens per second per user on eight NVIDIA B200 GPUs using SGLang and specialized DeepSeek kernels. Modal provides reproducible self-hosting code and a free, OpenAI-compatible endpoint available through April with one concurrent request per user, along with commercial options for higher limits or managed deployment. It also explains how to connect the endpoint to OpenCode, OpenClaw, Claude Code through a LiteLLM proxy, and the Vercel AI SDK, emphasizing compatibility with popular agent and application frameworks.
Feb 11, 2026
1,364 words in the original blog post.