How to serve trillions of tokens for trillion-parameter coding agents
Blog post from Modal
Modal describes how it optimized large-scale inference for coding agents using Moonshot AI’s trillion-parameter Kimi K2.6 model, raising per-replica performance by 2.8 times per user and 5.6 times across users while supporting services that processed hundreds of billions of tokens daily. The company explains that coding-agent workloads are dominated by long, highly overlapping session histories, making key-value cache management, low latency, and high token throughput central challenges. Its approach combined tensor parallelism across four GPUs, customized speculative decoding with a fine-tuned DFlash draft model, memory reductions through FP8 cache quantization and NVFP4 shared experts, and SGLang’s multi-tier HiCache system extending cache storage into CPU memory. Modal also found that scaling required cache-aware, load-aware routing rather than simple session-affinity hashing, because uneven session sizes, concurrent requests within sessions, and cache relocations during scale-ups created overloaded replicas and tail latency. The post emphasizes that inference performance must be evaluated against representative workloads, since output length, cache behavior, and speculative-decoding acceptance rates vary substantially by data, and notes that the resulting methods have been applied to newer models and released through SGLang contributions and Modal endpoint configurations.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 3 | 747 | 162 | 79 | -85% |
| AI Model Fine-tuning | 2 | 139 | 28 | 14 | -75% |
| Real-time | 2 | 649 | 155 | 80 | -85% |
| Harness engineering | 1 | 33 | 23 | 14 | -84% |
| Kubernetes | 1 | 956 | 75 | 30 | -73% |
| Local AI | 1 | 15 | 4 | 3 | -94% |
| Observability | 1 | 472 | 102 | 54 | -85% |
| Serverless | 1 | 156 | 54 | 28 | -80% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.