Kimi K2 0905 API Benchmarks: Latency, Throughput & Cost
Blog post from Deepinfra
Kimi K2 0905, developed by Moonshot AI, is an advanced large language model featuring 1 trillion total parameters and a 256k token context window, excelling in agentic coding intelligence and autonomous tasks. The model is available via multiple inference providers, with DeepInfra emerging as the recommended choice due to its balance of low latency (0.53s TTFT), lowest blended price ($0.80 per 1M tokens), and solid throughput (77.7 t/s), making it ideal for most production deployments. Groq offers the fastest generation speed (202.1 t/s) for throughput-intensive applications but at nearly double the cost of DeepInfra. Fireworks provides reliable service with a larger context window, while Novita, although cheaper than Groq and Fireworks, is not ideal for latency-sensitive tasks due to its slower performance. Looking forward, the Kimi K2.5 model presents significant advancements for vision-based inputs and complex workflows, supporting native multimodality and multi-agent orchestration.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Real-time | 4 | 6,296 | 1,346 | 246 | -2% |
| Multi-agent systems | 3 | 460 | 170 | 68 | -20% |
| AI Agents | 2 | 4,430 | 1,100 | 236 | -3% |
| LLM | 2 | 5,932 | 1,046 | 223 | -2% |
| Vector Search | 1 | 1,739 | 413 | 146 | -27% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.