Modal's serverless Servers
Blog post from Modal
Modal has introduced Servers, a low-latency HTTP, WebSocket, and gRPC serving option for regionalized, autoscaling application replicas, aimed particularly at interactive workloads such as LLM inference. Unlike Modal Web Functions, which provide built-in queueing and managed request lifecycles, Servers use a lighter request path that prioritizes speed and returns 503 errors when replicas are unavailable, shifting queueing and load-shedding responsibilities to applications. Modal reports reducing median request latency from 39 milliseconds for Web Functions to 6 milliseconds for Servers by avoiding control-plane lookups and network calls in the hot path. Its architecture combines AWS Network Load Balancers, Envoy edge proxies that terminate TLS and normalize traffic to HTTP/2, a custom Rust-based proxy called fprs for domain routing and replica load balancing, and compute-plane workers that relay traffic to user containers. The system maintains cached routing configuration synchronized from Spanner, supports autoscaling based on in-flight requests, provides proxy-level authentication to prevent unnecessary container activation, and includes traffic mirroring for uses such as A/B testing and continual learning. Servers are available through Modal’s SDK and also underpin its Endpoints product for streamlined LLM inference deployment.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 5 | 6,292 | 1,205 | 252 | -36% |
| Real-time | 4 | 6,055 | 1,444 | 270 | -11% |
| Kubernetes | 3 | 2,083 | 321 | 111 | +3% |
| Serverless | 1 | 1,019 | 237 | 96 | -45% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.