Home / Companies / Nebius / Blog / Post Details
Content Deep Dive

MLPerf® Inference v6.1: Previewing NVIDIA Vera Rubin NVL72 and leading full-rack Nebius system with NVIDIA GB300 NVL72 results

Blog post from Nebius

Post Details
Company
Date Published
Author
John Alexander
Word Count
1,003
Company Posts That Month
3
Language
English
Hacker News Points
-
Post removed?
No
Summary

Nebius reported its broadest MLPerf Inference v6.1 submission, benchmarking seven NVIDIA-based system configurations across DeepSeek R1, Qwen3-VL 235B, and gpt-oss 120B in server and offline scenarios. The company earned five first-place results among 20 category entries, led by its 72-GPU GB300 NVL72 rack configuration, which achieved first-place DeepSeek R1 throughput of 603,023 tokens per second in the server scenario and 689,961 in offline, while exceeding 1.1 million tokens per second on gpt-oss 120B. Nebius also submitted preview results for a 36-GPU Vera Rubin NVL72 system, one of only two Vera Rubin-based submissions, reporting 16,427 server tokens per second per GPU for DeepSeek R1. Results indicated near-linear scaling from eight to 72 GB300 GPUs for gpt-oss 120B, performance gains from HGX B200 to B300 systems, and competitive results from an eight-GPU RTX PRO 6000 configuration positioned for cost-efficient inference.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
LLM 1 747 162 79 -85%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.