Nebius proves bare-metal-class performance for AI inference workloads in MLPerf® Inference v5.1
Blog post from Nebius
Nebius has achieved leading performance results in the MLPerf® Inference v5.1 benchmarks, showcasing its AI systems powered by NVIDIA's high-demand GPUs such as the GB200 NVL72, HGX B200, and HGX H200. These results highlight Nebius' improvements in inference performance, with significant gains in token throughput across various configurations, notably outperforming previous benchmarks with systems like the NVIDIA GB200 NVL72 for large foundational models such as Llama 2 70B and Llama 3.1 405B. The benchmarks underscore Nebius' capability to run AI workloads efficiently in virtualized environments without compromising performance, thanks to their engineering expertise in utilizing NVIDIA hardware and software. The company's achievements illustrate its commitment to providing high-performance, scalable AI infrastructure that remains competitive with top-tier industry standards, offering customers supercomputer-level performance and reliability with the flexibility of a hyperscaler.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.