MLPerf® Inference v6.1: Previewing NVIDIA Vera Rubin NVL72 and leading full-rack Nebius system with NVIDIA GB300 NVL72 results
Blog post from Nebius
Nebius reported its broadest MLPerf Inference v6.1 submission, benchmarking seven NVIDIA-based system configurations across DeepSeek R1, Qwen3-VL 235B, and gpt-oss 120B in server and offline scenarios. The company earned five first-place results among 20 category entries, led by its 72-GPU GB300 NVL72 rack configuration, which achieved first-place DeepSeek R1 throughput of 603,023 tokens per second in the server scenario and 689,961 in offline, while exceeding 1.1 million tokens per second on gpt-oss 120B. Nebius also submitted preview results for a 36-GPU Vera Rubin NVL72 system, one of only two Vera Rubin-based submissions, reporting 16,427 server tokens per second per GPU for DeepSeek R1. Results indicated near-linear scaling from eight to 72 GB300 GPUs for gpt-oss 120B, performance gains from HGX B200 to B300 systems, and competitive results from an eight-GPU RTX PRO 6000 configuration positioned for cost-efficient inference.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 1 | 747 | 162 | 79 | -85% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.