Home / Companies / Nebius / Blog / Post Details
Content Deep Dive

Nebius proves bare-metal-class performance for AI inference workloads in MLPerf® Inference v5.1

Blog post from Nebius

Post Details
Company
Date Published
Author
Andrey Kuyukov
Word Count
1,768
Company Posts That Month
13
Language
English
Hacker News Points
-
Post removed?
No
Summary

Nebius has achieved leading performance results in the MLPerf® Inference v5.1 benchmarks, showcasing its AI systems powered by NVIDIA's high-demand GPUs such as the GB200 NVL72, HGX B200, and HGX H200. These results highlight Nebius' improvements in inference performance, with significant gains in token throughput across various configurations, notably outperforming previous benchmarks with systems like the NVIDIA GB200 NVL72 for large foundational models such as Llama 2 70B and Llama 3.1 405B. The benchmarks underscore Nebius' capability to run AI workloads efficiently in virtualized environments without compromising performance, thanks to their engineering expertise in utilizing NVIDIA hardware and software. The company's achievements illustrate its commitment to providing high-performance, scalable AI infrastructure that remains competitive with top-tier industry standards, offering customers supercomputer-level performance and reliability with the flexibility of a hyperscaler.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.