Home / Companies / Modal / Blog / Post Details
Content Deep Dive

How to serve trillions of tokens for trillion-parameter coding agents

Blog post from Modal

Post Details
Company
Date Published
Author
-
Word Count
7,266
Company Posts That Month
5
Language
English
Hacker News Points
4
Post removed?
No
Summary

Modal describes how it optimized large-scale inference for coding agents using Moonshot AI’s trillion-parameter Kimi K2.6 model, raising per-replica performance by 2.8 times per user and 5.6 times across users while supporting services that processed hundreds of billions of tokens daily. The company explains that coding-agent workloads are dominated by long, highly overlapping session histories, making key-value cache management, low latency, and high token throughput central challenges. Its approach combined tensor parallelism across four GPUs, customized speculative decoding with a fine-tuned DFlash draft model, memory reductions through FP8 cache quantization and NVFP4 shared experts, and SGLang’s multi-tier HiCache system extending cache storage into CPU memory. Modal also found that scaling required cache-aware, load-aware routing rather than simple session-affinity hashing, because uneven session sizes, concurrent requests within sessions, and cache relocations during scale-ups created overloaded replicas and tail latency. The post emphasizes that inference performance must be evaluated against representative workloads, since output length, cache behavior, and speculative-decoding acceptance rates vary substantially by data, and notes that the resulting methods have been applied to newer models and released through SGLang contributions and Modal endpoint configurations.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 3 747 162 79 -85%
AI Model Fine-tuning 2 139 28 14 -75%
Real-time 2 649 155 80 -85%
Harness engineering 1 33 23 14 -84%
Kubernetes 1 956 75 30 -73%
Local AI 1 15 4 3 -94%
Observability 1 472 102 54 -85%
Serverless 1 156 54 28 -80%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.