Qwen 3.8 vs. Kimi K3: benchmarks, licenses, and what it takes to run them
Blog post from Featherless
Moonshot’s Kimi K3 and Alibaba’s Qwen3.8 releases represent competing large hybrid-attention mixture-of-experts models with million-token context ambitions, though their publicly available checkpoints differ substantially. Kimi K3 is a 2.8-trillion-parameter multimodal MoE with 104 billion active parameters, a built-in vision encoder, and native 4-bit MXFP4 weights that reduce its storage footprint to roughly 1.4–1.56 TB, while Qwen3.8-Max is an API-only multimodal flagship and the open Qwen3.8-2.4T-A95B checkpoint is text-only, with 2.4 trillion total and 95 billion active parameters; Alibaba also offers the much smaller, Apache 2.0-licensed multimodal Qwen3.8-27B. Both use linear-attention and full-attention layers to make long contexts more practical, but K3 uses more experts and quantization-aware training, whereas Qwen ships primarily at full precision with an FP8 option. Vendor-reported overlapping benchmarks generally place K3 ahead on coding, tool-use, and agent evaluations, while Qwen3.8-Max remains close on reasoning tests and is positioned as stronger for instruction following and long-document retrieval; independent early rankings similarly place K3 slightly ahead among open-weight models. Operating the largest models requires substantial infrastructure, but Qwen3.8-Max’s API pricing is notably lower, particularly for output tokens generated during reasoning, while Featherless provides K3 at up to 256K context and Qwen3.8-27B as a lower-cost deployment option.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| AI Model Fine-tuning | 1 | 516 | 143 | 56 | -47% |
| MCP | 1 | 8,107 | 809 | 199 | -26% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.