Home / Companies / Cursor / Blog / Post Details
Content Deep Dive

Mixture-of-Kittens: our open-source MoE megakernel for NVL72s

Blog post from Cursor

Post Details
Company
Date Published
Author
Stuart Sul, Nash Brown, Henry Wildermuth, William Lin & Federico Cassano
Word Count
5,596
Company Posts That Month
2
Language
English
Hacker News Points
-
Post removed?
No
Summary

Mixture-of-Kittens (MoK) is an open-sourced, highly optimized mixture-of-experts (MoE) training megakernel designed to address the bottleneck in the training of Composer, an agentic coding model, by integrating all MoE communication and computation into a single kernel for NVL72s. MoK emerged from previous efforts to enhance the MoE layer and now powers training across thousands of GPUs by achieving up to 2.37x higher throughput compared to public baselines. The kernel incorporates novel strategies such as pull-based communication to maximize NVLink bandwidth, ring token buffers to eliminate CPU-GPU synchronization, and scheduling overlaps to efficiently manage computation-communication tasks. MoK supports both BF16 and MXFP8 precision modes, with the latter offering faster performance without numerical issues. It is designed to be flexible and modifiable, allowing adaptations for various platforms, and shows significant improvements—up to 41% speedup—over previous implementations in production training stacks, thus lowering barriers to AI research and enabling more efficient model training.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.