Home / Companies / Modal / Blog / August 2026

August 2026 Summaries

2 posts from Modal

Filter
Month: Year:
Post Summaries Back to Blog
Qwen3.8-2.4T-A95B, an open-weights model that improves on Qwen 3.7 in coding, workplace, research, and long-horizon tasks, is now available through Modal’s Auto Endpoints and as an OpenAI-compatible Shared Endpoint. Modal worked with Qwen before launch to support the model using SGLang and a custom DFlash speculative decoding model optimized for Qwen3.8’s architecture. The speculator was trained with an emphasis on tool-call-heavy coding, research, and work sequences to improve drafted-token acceptance rates and accelerate inference. The text-only model will be offered for the next month with token-based pricing.
Aug 12, 2026 188 words in the original blog post.
Modal has redesigned its Function I/O plane to reduce network latency and improve performance for globally distributed serverless workloads. Previously, all Function inputs and outputs passed through us-east servers regardless of where clients and containers were located, which could add hundreds of milliseconds for applications such as low-latency European RAG services. The new geographically distributed I/O plane supports at least four regions and lets developers choose a routing region independently from their container region; migration to the new us-east plane has reduced median end-to-end Function Call latency by roughly 80 ms. The architecture queues serialized inputs in Redis, dispatches them to available containers, returns serialized outputs through client polling, and supplies autoscaling, health, observability, and retry mechanisms. Key improvements include moving noncritical work off the request path, reducing shared-storage access, caching metadata, using JWT-based authentication, and rewriting servers in Go for stronger concurrent gRPC performance, while retaining Redis 7.1 after tests found CPU spikes in newer versions. Modal advises users to route requests near clients, balance regional constraints against capacity and cost, keep payloads below 2 MiB to avoid blob-storage transfers, and batch many small inputs to reduce network overhead.
Aug 04, 2026 1,180 words in the original blog post.