September 2025 Summaries
7 posts from Modal
Filter
Month:
Year:
Post Summaries
Back to Blog
Modal announced an $80 million-plus Series B led by Lux Capital, valuing the company at $1.1 billion and bringing its total funding to $111 million. The company positions its serverless, AI-native infrastructure as an alternative to legacy cloud systems, pooling global GPU and CPU capacity while offering programmable compute, storage, and networking, rapid container startup, low-latency routing, and usage-based pricing. After four years of developing proprietary components including a file system, container runtime, and scheduler, Modal says it supports thousands of customers with AI workloads ranging from model inference and reinforcement learning to transcription, protein folding, and weather forecasting. Its product suite includes inference, secure sandboxes, batch processing, distributed training, and notebooks, with customers and organizations such as Meta using the platform for large-scale AI development. Modal plans to expand toward providing infrastructure for every stage of building and operating production AI systems.
Sep 29, 2025
788 words in the original blog post.
Flash Attention 4 is a CUDA kernel for Transformer attention on Nvidia Blackwell GPUs that reportedly delivers about a 20% improvement over Nvidia cuDNN attention kernels by combining hardware-specific optimizations, asynchronous execution, and numerical techniques. Based on reverse engineering of its released source code, the kernel divides query, key, value, and output tensors into tiles and processes them through a manually coordinated producer-consumer pipeline spanning global memory, shared memory, Tensor Memory, Tensor Cores, CUDA Cores, and specialized warps. Its five warp roles load data, perform matrix multiplications, calculate online softmax normalization, correct prior outputs when scaling changes, and write completed results back to memory, with barriers and buffering used to overlap work and reduce stalls. Key FA4 changes include a software cubic approximation for some exponentials that can reduce pressure on Special Function Units, and a more selective online-softmax rescaling strategy that reportedly reduces correction operations tenfold. The analysis argues that FA4 illustrates a broader shift in GPU programming toward increasingly complex tile-based, warp-specialized asynchronous pipelines, motivating new Nvidia tools and languages intended to make such low-level optimization more manageable.
Sep 26, 2025
4,193 words in the original blog post.
Modal Vibe is a demonstration platform built on Modal Sandboxes to show how AI-generated or “vibe-coded” applications can be securely run and managed at large scale. It consists of a sandbox manager, a web frontend, and a fleet of isolated runtimes, avoiding the operational complexity often associated with orchestrators such as Kubernetes or Nomad. Modal claims its infrastructure can start sandboxes in seconds, launch thousands per second, and support tens of thousands concurrently, with built-in monitoring and dashboards. In a test that included AI code-generation time, the platform created one app in about 30 seconds, 10 in 60 seconds, 100 in 150 seconds, and 1,000 separate sandboxed apps in roughly five minutes, with AI API latency becoming the primary bottleneck at scale. The project’s source code and a public deployment are available, though the company characterizes it as a demo rather than production-ready software and advises users to consult its documentation before deploying similar systems.
Sep 22, 2025
840 words in the original blog post.
After investing in Modal’s Series A and later advising the company, a former MongoDB leader and Pace founder has joined Modal as VP of Sales, motivated by its growth, developer-focused compute platform, and position in the expanding AI infrastructure market. The author describes Modal as a high-eight-figure-revenue company serving thousands of customers with AI and large-scale computing applications, led by founders Erik and Akshat and supported by a culture emphasizing intelligence, hard work, ambition, and kindness. Drawing on experience helping MongoDB scale go-to-market operations from $20 million to more than $1 billion in revenue, the new executive plans to build a scalable, compounding GTM organization to accelerate Modal’s growth. Modal is also expanding its sales, technical pre-sales, partnerships, and marketing teams, seeking ambitious and creative infrastructure go-to-market professionals.
Sep 22, 2025
643 words in the original blog post.
Modal’s latest updates introduce Modal Notebooks for collaborative, browser-based Python development with custom images, volumes, and access to as many as eight B200 GPUs, alongside optional sandbox idle timeouts that automatically end inactive environments. Recent client releases add separate container startup timeouts for functions and classes and an imperative API for managing resources such as volumes, secrets, dictionaries, and queues. The announcement also highlights Zencastr’s use of Modal to scale audio-processing workloads to 1,500 concurrent GPUs, expands Modal’s GPU Glossary with performance concepts, and provides tutorials for fine-tuning Qwen3-14B with Unsloth and improving OpenAI Whisper transcription models. Additional news includes upcoming biotech and voice-AI events in Boston and San Francisco, as well as an investor-produced technical deep dive into Modal’s performance architecture.
Sep 19, 2025
530 words in the original blog post.
Modal Notebooks is a cloud-based, collaborative Jupyter environment designed to provide rapid access to GPUs, custom container images, persistent storage, and modern editor capabilities without the cost of continuously running dedicated development machines. Its architecture runs kernels in isolated Modal Sandboxes and uses a custom kernel shim to translate Jupyter’s ZeroMQ protocol into tunneled HTTP communications, enabling streamed outputs and multi-user access. Fast startup depends on a lazy-loading FUSE filesystem that fetches container files on demand through a tiered content-addressed cache, while scheduling shares a large CPU and GPU pool and automatically pauses idle kernels to reduce costs. Persistent global data access is provided by VolumeFS, a distributed filesystem supporting Modal Volumes, and real-time editing uses operational transformation, Redis Streams, CodeMirror, and durable object storage for large outputs. The product also adds link sharing, Jupyter Widgets, Language Server Protocol support through Pyright, Ruff-based formatting, and AI code-completion experiments, reflecting Modal’s expansion from an SDK-focused platform into a broader interactive development product.
Sep 16, 2025
1,766 words in the original blog post.
Modal Notebooks is now generally available as a collaborative, GPU-enabled Python environment for AI research, data exploration, and experimentation on Modal’s infrastructure. It aims to reduce common cloud-notebook limitations through sub-five-second kernel starts, flexible hardware scaling from small CPU allocations to multiple high-end GPUs, automatic idle shutdown and resumption, and support for custom or curated container images. The platform integrates directly with Modal Volumes, Secrets, Functions, and production runtime tools, while offering real-time shared editing, fine-grained permissions, Jupyter Widgets, language-server assistance, AI completions, and rich visual outputs. During its August beta, more than 5,000 accounts reportedly ran 200,000 code cells, with early users highlighting rapid GPU switching, scalable capacity, and easier collaboration. Modal plans to add memory snapshots, notebook-to-app export, scheduled runs, and edit history, and is offering new users $30 in free credits alongside example notebooks for tasks such as transcription, coding agents, model experimentation, OCR, and embedding visualization.
Sep 09, 2025
1,000 words in the original blog post.