Nebius and PyTorch partner to accelerate frontier MoE training on NVIDIA Blackwell
Blog post from Nebius
In a collaborative effort with PyTorch, Nebius demonstrated up to 41% faster pre-training of DeepSeek-V3 models on NVIDIA Blackwell GPUs, highlighting advancements in model training infrastructure to accommodate evolving architectures like Mixture-of-Experts (MoE) models. The experiments focused on optimizing training performance using TorchTitan on a 256-GPU NVIDIA HGX B200 cluster within Nebius Cloud, leveraging MXFP8 training for improved performance and DeepEP for efficient expert-parallel communication. Utilizing a Nebius Cloud cluster optimized for large-scale AI workloads, the experiments achieved significant throughput increases, confirming MXFP8's equivalent convergence behavior to BF16 through loss-curve validation. Conducted with open-source PyTorch-native tools, the collaboration underscores the importance of integrated hardware, software, and infrastructure innovations in improving the efficiency of large-scale AI model training, providing a reproducible framework for others using Blackwell-based clusters.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.