To Infinity and Beyond: ThunderKittens Now on NVIDIA Vera Rubin NVL72!
Blog post from Together AI
Together AI’s kernels team evaluated NVIDIA’s Vera Rubin NVL72 platform and extended ThunderKittens to support optimized NVFP4 and FP8 GEMMs, finding that Blackwell-era kernels initially reached only about 42–44% of Rubin’s performance roofline because the faster tensor cores were insufficiently supplied with data. Vera Rubin retains Blackwell’s core programming model but adds a doubled 64-byte K-step for tensor-core MMA instructions, expanded tensor memory of up to 576 columns through exclusive allocations, 328 KiB oversized shared memory, B-side collector support for reducing operand reads, and early A-operand release to accelerate pipeline reuse. By widening MMA instructions, shifting from 1x1 to 2x1 tile layouts to reuse B operands, deepening shared-memory pipelines, tuning CTA cluster configurations and rasterization, using the B-side collector, and applying early A release for large workloads, the team raised its NVFP4 kernel performance to more than 22 PFLOPS and made it competitive with cuBLAS and CuTE DSL. Results were measured on a qualification-sample GPU with CUDA 13.4, and the authors note that software and baseline performance may continue to evolve as Vera Rubin releases mature.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Real-time | 1 | 649 | 155 | 80 | -85% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.