Home / Companies / Nebius / Blog / April 2026

April 2026 Summaries

4 posts from Nebius

Filter
Month: Year:
Post Summaries Back to Blog
Nebius has reached significant milestones in collaboration with NVIDIA, establishing itself as a leading AI cloud provider by achieving the NVIDIA Exemplar Cloud status on NVIDIA GB300 NVL72 for training workloads. This achievement marks Nebius as one of the few global cloud providers to meet NVIDIA's rigorous performance and reliability standards across multiple GPU generations, including the NVIDIA H200 and Blackwell Ultra GPUs. The Exemplar Cloud initiative assesses cloud providers using real-world AI training workloads, providing a transparent validation of infrastructure performance that is essential for large-scale training decisions. The collaboration between Nebius and NVIDIA encompasses AI factory design, fleet health monitoring, and inference optimizations, with Nebius playing a key role in the Physical AI Data Factory Blueprint alongside major players like Azure. This accomplishment underscores Nebius's commitment to delivering robust AI infrastructure capable of handling demanding workloads, as they prepare for future developments including gigawatt-scale AI factories in the U.S. and Europe and the introduction of NVIDIA Vera Rubin NVL72 in 2026.
Apr 29, 2026 447 words in the original blog post.
Building AI agents within regulated enterprises often faces challenges related to security, compliance, and governance, especially when prototyping with sensitive data. TD SYNNEX has introduced Nebius AI Cloud, a dedicated platform featuring NVIDIA-accelerated infrastructure, as the first AI-Infrastructure-as-a-Service (AI IaaS) in their global portfolio, providing the performance and flexibility needed for enterprise AI. This platform enables secure, air-gapped prototyping using NVIDIA DGX Spark systems and facilitates the transition to scalable deployment without exposing sensitive data. In one practical application within a major health system, AI models were prototyped locally, validating agent performance before moving only the agent artifact to the Nebius AI Cloud. Nebius is recognized as an Exemplar Cloud provider in the NVIDIA ecosystem, supporting advanced AI workloads and optimizing large foundation models. The collaboration with Aible highlights how enterprises can seamlessly transition AI workflows from secure local environments to production on Nebius AI Cloud, maintaining compliance and efficiency.
Apr 16, 2026 625 words in the original blog post.
LK losses are proposed as an alternative training objective to KL divergence for optimizing speculative decoding in large language models (LLMs), aiming to enhance the acceptance rate of draft tokens without computational overhead. This method improves inference throughput across a wide range of model sizes by directly targeting the acceptance rate, which is an important metric in speculative decoding, rather than relying on KL as a proxy measure. Speculative decoding involves a two-stage process where a smaller draft model generates candidate tokens, and a larger target model verifies them, with acceptance rate being crucial for efficiency. LK losses, including a negative log-acceptance objective and a hybrid objective blending KL and TV distance, address the limitations of KL divergence, especially for low-capacity draft models that cannot perfectly match the target distribution. Experiments demonstrate that LK losses consistently outperform KL baselines across different architectures and model sizes, particularly benefiting lower-capacity models and challenging approximation tasks. The approach is scalable, adaptable to various architectures and target models, and has been open-sourced, offering trained models and datasets to the community.
Apr 10, 2026 2,543 words in the original blog post.
Nebius has announced its latest results from the MLPerf Inference v6.0 benchmark, which measures AI inference performance, showcasing their advanced capabilities on NVIDIA's newest Blackwell and Blackwell Ultra GPU platforms. Their submission included systems like the NVIDIA HGX B200, HGX B300, and the GB300 NVL72, which were evaluated using three models: DeepSeek R1, a server and offline inference model; Qwen3-VL 235B, a multimodal model; and gpt-oss 120B, a large open-source language model. Notably, Nebius achieved 10 first-place results out of 16 benchmark submissions, with significant performance demonstrated on the full-rack GB300 NVL72 system utilizing 72 NVIDIA Blackwell Ultra GPUs. The results emphasize Nebius' ability to efficiently support large-scale production deployments of modern AI models, highlighting their continuous optimization of AI infrastructure and collaboration with NVIDIA to maximize hardware potential. The findings illustrate the scalability and performance improvements across successive NVIDIA platforms, reaffirming Nebius' readiness to handle demanding AI workloads and their commitment to enhancing AI cloud platforms for their customers.
Apr 01, 2026 946 words in the original blog post.