Nebius and PyTorch partner to accelerate frontier MoE training on NVIDIA Blackwell
Blog post from Nebius
In a collaborative effort with PyTorch, Nebius demonstrated up to 41% faster pre-training of DeepSeek-V3 models on NVIDIA Blackwell GPUs, highlighting advancements in model training infrastructure to accommodate evolving architectures like Mixture-of-Experts (MoE) models. The experiments focused on optimizing training performance using TorchTitan on a 256-GPU NVIDIA HGX B200 cluster within Nebius Cloud, leveraging MXFP8 training for improved performance and DeepEP for efficient expert-parallel communication. Utilizing a Nebius Cloud cluster optimized for large-scale AI workloads, the experiments achieved significant throughput increases, confirming MXFP8's equivalent convergence behavior to BF16 through loss-curve validation. Conducted with open-source PyTorch-native tools, the collaboration underscores the importance of integrated hardware, software, and infrastructure innovations in improving the efficiency of large-scale AI model training, providing a reproducible framework for others using Blackwell-based clusters.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Kubernetes | 1 | 2,478 | 412 | 128 | +56% |
| Observability | 1 | 4,660 | 984 | 209 | +14% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.