April 2026 Summaries
6 posts from Modal
Filter
Month:
Year:
Post Summaries
Back to Blog
AE Studio describes using Modal to train language models for Lean theorem proving with reinforcement learning, comparing Evolution Strategies (ES), which evaluates randomly perturbed model variants and updates toward higher-scoring ones, with the more common Group Relative Policy Optimization (GRPO). The system separates GPU-based proof generation, CPU-based Lean verification, and orchestration, using Modal’s function-specific environments, parallel job mapping, isolated sandboxes, shared model volumes, and secrets management to coordinate thousands of proof attempts while containing verifier failures. ES checkpointing stores only deterministic perturbation seeds and rewards, allowing workers to reconstruct model updates without transferring large weight files. AE Studio reports that its implementation reduced infrastructure code and setup time, estimated lower costs through elastic GPU use, and produced early results in which ES sometimes matched or exceeded GRPO in verified proofs, particularly under limited training data, though performance varied and requires further study. The team plans experiments on hyperparameters, larger models, and ES scaling behavior, while presenting the architecture as applicable to other workflows that combine GPU generation with external verification or testing.
Apr 29, 2026
2,237 words in the original blog post.
OpenAI’s Agents SDK is presented as a framework for building customized internal agent harnesses, and the walkthrough combines it with Modal Sandboxes to create coding agents that can safely execute commands in isolated environments with optional GPU access. Beginning with an unsafe local shell-execution agent, the example moves execution into remote sandboxes, then adds persistent sessions for memory, an orchestrator-subagent model to separate high-level planning from short-lived implementation work, and asynchronous pools that allow multiple agents to run experiments in parallel. The system also introduces agent-status tracking, GPU quotas to control costs, filesystem snapshots that let new agents reuse prepared environments and artifacts, and modular skills that load task-specific instructions without hard-coding them into the general harness. Using MNIST training and OpenAI’s Parameter Golf challenge as examples, the project demonstrates how isolated compute, delegated context, parallel workers, reusable sandbox states, and configurable prompting can be composed into an autonomous research and coding workflow.
Apr 15, 2026
2,756 words in the original blog post.
Modal presents its serverless GPU platform as a way for AI research agents to dynamically choose both the scale and type of computing resources needed for experiments, avoiding the cost of idle clusters while enabling parallel work beyond a single workstation’s capacity. In a demonstration using Claude Code and OpenAI’s Parameter Golf challenge, an agent reportedly conducted 113 experiments over 15 hours and 238 GPU-hours, moving between single-GPU pipeline tests, roughly 40 parallel hyperparameter trials, five simultaneous 8×H100 validation runs, serial debugging, and later large-scale optimization. The agent improved its bits-per-byte score from 1.42 to approximately 1.12 by testing model and training configurations, while resolving a major CPU-based quantization bottleneck by rewriting it for GPU execution. Modal attributes the reported fivefold speedup in core training over an 8×H100 workstation and improved resource efficiency relative to a continuously provisioned 40-GPU cluster to its ability to rapidly provision, scale, and automatically release GPU jobs, sandboxes, storage, and parallel tasks through code-oriented tools and agent guidance.
Apr 14, 2026
1,680 words in the original blog post.
Modal has acquired Butter, bringing founder Erik Dunteman and researcher Raymond Tana onto its Sandbox team to help expand Modal Sandboxes. Butter’s work has focused on agent harness engineering, including computer-use agents, deterministic memory systems, skills, code generation, and bVisor, a lightweight ephemeral sandbox powered by a virtual Linux kernel written in Zig. Modal expects the team’s expertise to strengthen its sandbox product, while Dunteman argues that code generation can improve AI agent systems through greater determinism, auditability, and fine-grained permissions. The acquisition also builds on Dunteman’s prior relationship with Modal, which began during his time cofounding serverless GPU platform Banana and later included a brief contracting role at Modal.
Apr 10, 2026
207 words in the original blog post.
Physical Intelligence (Pi) is developing a general-purpose robotic intelligence system whose Visual-Language-Action models continuously convert camera observations, natural-language instructions, and robot state data into motor commands. To validate each model revision, Pi runs real-world robot evaluations around the clock, requiring scalable low-latency inference across a growing fleet. While Modal Tunnels provide direct, secure TCP access for many latency-sensitive applications, Pi’s continuous control loops were vulnerable to TCP jitter and request stalls, so Pi and Modal developed a QUIC-over-UDP portal with persistent bidirectional connections, NAT traversal, STUN discovery, and UDP hole punching. The system adds roughly 10–15 milliseconds of network overhead while allowing robots without onboard GPUs to use remote, data-center-class hardware. Modal also enables Pi to mount model checkpoints through Modal Volumes, reducing loading time to under 30 seconds, and to deploy inference in regions near robot fleets, supporting expansion without local GPU infrastructure or custom relay systems.
Apr 08, 2026
657 words in the original blog post.
Modal announced availability of NVIDIA’s RTX Pro 6000 Blackwell GPU, offering 96GB of VRAM and strong FP4/FP8 performance for inference and fine-tuning workloads, alongside a dashboard Command Palette for faster navigation. Its Sandbox Filesystem API has entered beta with improved reliability, bidirectional streaming, support for reading files up to 5GB and writing files of any size, and volume synchronization. SDK version 1.4.0 adds richer CLI log retrieval and filtering, a recreate deployment strategy that immediately replaces running containers, minimal images via `modal.Image.from_scratch()`, and OIDC identity tokens for sandbox authentication, while removing compatibility for several older APIs. The update also highlights deployment guidance for Google’s Gemma 4 model, customer stories from Runway, Imbue, and Doppel, and several AI-focused events scheduled in April.
Apr 07, 2026
791 words in the original blog post.