Home / Companies / Nebius / Blog / September 2026

September 2026 Summaries

3 posts from Nebius

Filter
Month: Year:
Post Summaries Back to Blog
Nebius reported its broadest MLPerf Inference v6.1 submission, benchmarking seven NVIDIA-based system configurations across DeepSeek R1, Qwen3-VL 235B, and gpt-oss 120B in server and offline scenarios. The company earned five first-place results among 20 category entries, led by its 72-GPU GB300 NVL72 rack configuration, which achieved first-place DeepSeek R1 throughput of 603,023 tokens per second in the server scenario and 689,961 in offline, while exceeding 1.1 million tokens per second on gpt-oss 120B. Nebius also submitted preview results for a 36-GPU Vera Rubin NVL72 system, one of only two Vera Rubin-based submissions, reporting 16,427 server tokens per second per GPU for DeepSeek R1. Results indicated near-linear scaling from eight to 72 GB300 GPUs for gpt-oss 120B, performance gains from HGX B200 to B300 systems, and competitive results from an eight-GPU RTX PRO 6000 configuration positioned for cost-efficient inference.
Sep 16, 2026 1,003 words in the original blog post.
Nebius has launched the free AI Builder Program to support developers building AI systems with open models and interoperable tools rather than relying on a single vendor. The program provides more than $400 in credits and discounts from Nebius Token Factory, LangChain, Toloka, Tavily, and other partners, along with working-code cookbooks, training through Nebius Academy and NVIDIA, discounted certifications, expert office hours, and a builder community on Discord and at live events. Its approach centers on “open primitives,” enabling users to combine models, inference, orchestration, retrieval, evaluation, coding agents, human labeling, and post-training tools into customizable systems that can be updated as new models emerge. Initial partners include NVIDIA Nemotron, Qwen, MiniMax, Cognition, OpenHands, LangChain, Tavily, Composio, Toloka, LangSmith, Prime Intellect, and Hugging Face OpenEnv, while Nebius supplies infrastructure for hosting, serving, training, and fine-tuning open models.
Sep 10, 2026 786 words in the original blog post.
Nebius submitted MLPerf Storage v3.0 Closed division results for its Enhanced Object Storage service, using standard CPU-only virtual machines and the open-source mlpstorage harness to simulate AI-training I/O over the S3 API. The company reported feeding 768 simulated NVIDIA B200 accelerators for RetinaNet, the highest accelerator count in the round, with 88.4% utilization, 102.83 GiB/s read bandwidth, and 341,881 samples per second; it also achieved the highest Unet3D accelerator count among object-storage submissions, sustaining 114.45 GiB/s and 92.5% utilization across 21 simulated accelerators. For Llama 3.1 checkpointing, aggregate write bandwidth rose from 2.81 GiB/s on one node for an 8B model to 19.46 GiB/s across eight nodes for a 70B model, while read bandwidth increased from 7.14 GiB/s to 38.27 GiB/s. MLPerf Storage evaluates whether storage can maintain minimum accelerator utilization under realistic training workloads, with RetinaNet emphasizing small-file latency and Unet3D emphasizing large-file bandwidth, and Nebius presents its results as evidence that a single production object-storage service can support training and checkpointing workloads at scale.
Sep 01, 2026 841 words in the original blog post.