Home / Companies / Modal / Blog / September 2026

September 2026 Summaries

5 posts from Modal

Filter
Month: Year:
Post Summaries Back to Blog
Quail is a query-aware inference engine developed by Modal and Carnegie Mellon University’s Full Stack Data Lab to accelerate AI-SQL workloads, where SQL queries generate prompts from database records to perform tasks such as filtering, classification, and fuzzy joins. Unlike chatbot and coding-agent inference, these workloads can involve millions of mostly independent, relatively low-intelligence requests and often need only a single Boolean token rather than full text generation. Quail uses knowledge of an entire SQL query plan to optimize request ordering, reuse and precisely evict transformer KV caches, avoid decode-phase overhead, reduce CPU bottlenecks, and tailor GPU kernels for prefill-heavy processing. On a complex multi-join query, the authors report throughput above one billion tokens per minute per H100 GPU, more than ten times their vLLM baseline, while their broader AI-SQL benchmark shows a geometric-average speedup of 1.84 times. The system combines conventional database planning techniques with KV-aware join ordering and hardware-based cost estimation, and it relies on open-source components such as sqlglot, PyArrow, FlashAttention, DeepGEMM, and Triton. The authors describe Quail as an early step toward specialized but reusable inference infrastructure, identifying future opportunities in tiered and cross-query caching, larger datasets, prefix-sharing indexes, improved kernel overlap, and potentially adaptive model optimization during query execution.
Sep 24, 2026 4,194 words in the original blog post.
Modal describes how it optimized large-scale inference for coding agents using Moonshot AI’s trillion-parameter Kimi K2.6 model, raising per-replica performance by 2.8 times per user and 5.6 times across users while supporting services that processed hundreds of billions of tokens daily. The company explains that coding-agent workloads are dominated by long, highly overlapping session histories, making key-value cache management, low latency, and high token throughput central challenges. Its approach combined tensor parallelism across four GPUs, customized speculative decoding with a fine-tuned DFlash draft model, memory reductions through FP8 cache quantization and NVFP4 shared experts, and SGLang’s multi-tier HiCache system extending cache storage into CPU memory. Modal also found that scaling required cache-aware, load-aware routing rather than simple session-affinity hashing, because uneven session sizes, concurrent requests within sessions, and cache relocations during scale-ups created overloaded replicas and tail latency. The post emphasizes that inference performance must be evaluated against representative workloads, since output length, cache behavior, and speculative-decoding acceptance rates vary substantially by data, and notes that the resulting methods have been applied to newer models and released through SGLang contributions and Modal endpoint configurations.
Sep 23, 2026 7,266 words in the original blog post.
Modal’s August updates add day-zero Auto Endpoint support for Kimi K3, Qwen 3.8, GLM 5.3, and GLM 5.3 Flash through shared token-priced or dedicated GPU-second-priced deployments, with optimizations intended to improve agentic inference speed. The platform also introduced environment-level budgets, a clearer usage-limits interface, reduced broad-region pricing multipliers, and restricted environments with configurable default roles for stronger access control. Sandbox capabilities expanded with public-alpha sidecars for colocated multi-container workloads, CPU and memory utilization graphs, public-beta VM Sandboxes supporting Linux features such as Docker, eBPF, and systemd, and generally available directory snapshots and a faster filesystem API. Other changes include region-selectable outbound proxies, updated Python, JavaScript, and Go SDKs, a refreshed dashboard with beta light mode, and infrastructure improvements reportedly enabling up to one million concurrent Sandboxes and lower Function I/O latency. Modal also highlighted customer and ecosystem integrations, including support for Botika, Devin Outposts, and Cursor Cloud Agents, alongside its upcoming Runtime conference and community events.
Sep 14, 2026 1,335 words in the original blog post.
Modal is opening a London office as its second European location, expanding on its long-standing engineering presence in Stockholm, where a nearly 20-person team develops core technologies such as its custom file system and Sandboxes product. Led by newly appointed VP EMEA Hugh Killingbeck-Jones, the London office will initially build go-to-market teams serving startups and enterprises, with plans to add engineering and other functions over time. The company views London as Europe’s AI capital and aims to support ambitious European AI businesses in fields including robotics, biotech, health, legal, and fintech with infrastructure for model inference, training, fine-tuning, batch processing, and isolated sandboxes. Modal highlights customers such as Black Forest Labs, Legora, and DoorDash, while also pledging further investment in meeting European regulatory requirements. It is hiring across commercial and engineering roles, planning EMEA meetups and conference appearances, and inviting European companies to adopt its platform.
Sep 02, 2026 628 words in the original blog post.
Botika, a generative-AI company that produces and personalizes fashion imagery for global brands, uses Modal to operate its data processing, model training, research, and production inference systems. The company maintains proprietary foundation models with tens of billions of parameters, processes a 100-terabyte image dataset, and serves roughly 15 production models across several GPU types. After previously managing Kubernetes and GCP Batch infrastructure, Botika adopted Modal in 2023 to reduce the operational work associated with autoscaling, GPU management, cold starts, orchestration, and environment configuration. Its pipeline runs numerous specialized AI models across thousands of concurrent containers, while researchers reportedly increased experiment throughput from one or two runs per day to dozens by launching short jobs and scaling promising results. Botika also uses Modal for multi-node training, reinforcement learning infrastructure, and inference that can automatically handle sudden traffic increases, with CEO Eran Dagan saying the platform has allowed the company to operate with fewer infrastructure-focused staff and avoid relying extensively on traditional cloud services.
Sep 02, 2026 1,102 words in the original blog post.