Home / Companies / Nebius / Blog / March 2026

March 2026 Summaries

16 posts from Nebius

Filter
Month: Year:
Post Summaries Back to Blog
Nebius, as a founding consortium partner of the Physical AI Leaderboard (PhAIL) by Positronic, is helping to establish a rigorous, real-world benchmark for evaluating robotics AI models, focusing on practical metrics like Units Per Hour (UPH) and Mean Time Between Failures or Assists (MTBF/A) in bin-to-bin order picking tasks. PhAIL aims to provide a transparent and reproducible evaluation process using real hardware, with all runs recorded and available for independent auditing. Positronic developed the evaluation methodology and operates the benchmark, while Nebius contributes its vertically-integrated AI infrastructure, offering managed services and high-performance computing resources to support model fine-tuning and evaluation. The initiative encourages open participation, offering a free dataset and open-source training scripts, with the evaluation hardware being widely accessible. The consortium is inviting new members, including hardware vendors and academic labs, to help shape future metrics, with more details available in Positronic's full blog post.
Mar 31, 2026 477 words in the original blog post.
The Nebius VPN Gateway CLI is an open-source command line tool designed to facilitate secure, private connectivity for teams operating AI workloads on the Nebius AI Cloud. It allows for the creation, configuration, and management of an IPsec VPN gateway, enabling encrypted connections between the Nebius AI Cloud VPC and various external networks, such as on-premises infrastructure or other cloud VPCs. The tool supports both BGP and static routing, providing flexibility for different network configurations and ensuring high availability through multi-tunnel setups with automatic or manual failover. The CLI is designed for technical teams such as platform and network engineers and SREs, offering configuration as code, safe defaults, and operational commands for validation and failover. It employs modern cryptographic standards like AES-256 and SHA-256/384/512 and defaults to IKEv2 for IPsec connectivity, with IKEv1 as a fallback. This ensures secure, reliable, and repeatable connectivity processes that are crucial for AI platforms requiring dependable network interactions.
Mar 27, 2026 768 words in the original blog post.
Nebius AI Cloud's Aether 3.5 update focuses on enhancing scalability, accessibility, and ease of use for AI infrastructure, emphasizing the transition from research to real-world applications in various fields such as robotics, biotech, and scientific computing. The update introduces serverless capabilities that allow for elastic compute without infrastructure management, new NVIDIA RTX PRO 6000 Blackwell GPUs for applied AI workloads, and improved data transfer services for efficient intra-cloud data movement. Additionally, enhancements in cluster configuration and orchestration through Managed Soperator, along with an updated Applications interface, enable streamlined workflows and faster access to necessary tools. Platform-level improvements address security, operational control, and cost management, ensuring efficient governance and observability. These advancements aim to remove barriers between ideas and real-world AI solutions, fostering an environment where teams can experiment, iterate, and deploy AI applications with greater flexibility and control.
Mar 26, 2026 1,477 words in the original blog post.
Nebius is enhancing AI infrastructure accessibility by introducing DevPods, Jobs, and Endpoints, services designed to simplify AI compute through a container-based serverless approach. These services aim to alleviate the complexities of traditional AI environment setups, allowing data scientists and ML engineers to focus on experimentation and model development without the need for permanent clusters or detailed infrastructure management. Serverless computing offers benefits such as on-demand resource provisioning, pay-per-use pricing, and minimized infrastructure visibility, making it an ideal solution for AI workloads. DevPods provide interactive development environments, while Jobs facilitate running finite workloads, and Endpoints enable the deployment of custom models with web accessibility. Although currently in public and private previews, these tools are expected to evolve with additional capabilities like auto-scaling and multi-node support, promising a flexible and elastic AI compute experience on Nebius' robust platform. This evolution aligns with industry standards and is supported by Nebius' control over the infrastructure lifecycle, ensuring secure and high-performance compute environments.
Mar 26, 2026 1,480 words in the original blog post.
NVIDIA has introduced the RTX PRO 6000 Blackwell Server Edition, expanding its GPU lineup to cater to the growing demand for simulation workloads and multimodal AI applications. This new GPU is designed for high-performance computing tasks that require robust single-precision calculations, optimized inference, and advanced visualization, making it ideal for scientific simulations, digital twins, and synthetic data generation. Equipped with 96GB GDDR7 memory, the RTX PRO 6000 Blackwell supports large AI models on a single card, enhancing efficiency and reducing communication overhead in production environments. The GPU's 5th Generation Tensor Cores facilitate FP4 precision for cost-efficient inference, while Multi-Instance GPU (MIG) technology allows partitioning into four isolated instances, optimizing resource allocation. As the successor to the L40S, it offers improved memory capacity, RT Core performance, and single-precision compute, supporting larger models and more flexible configurations. Available on the Nebius platform, it is ready for use in regulated industries, ensuring strong data protection and secure environments for HIPAA-compliant deployments.
Mar 26, 2026 814 words in the original blog post.
In a collaborative effort with PyTorch, Nebius demonstrated up to 41% faster pre-training of DeepSeek-V3 models on NVIDIA Blackwell GPUs, highlighting advancements in model training infrastructure to accommodate evolving architectures like Mixture-of-Experts (MoE) models. The experiments focused on optimizing training performance using TorchTitan on a 256-GPU NVIDIA HGX B200 cluster within Nebius Cloud, leveraging MXFP8 training for improved performance and DeepEP for efficient expert-parallel communication. Utilizing a Nebius Cloud cluster optimized for large-scale AI workloads, the experiments achieved significant throughput increases, confirming MXFP8's equivalent convergence behavior to BF16 through loss-curve validation. Conducted with open-source PyTorch-native tools, the collaboration underscores the importance of integrated hardware, software, and infrastructure innovations in improving the efficiency of large-scale AI model training, providing a reproducible framework for others using Blackwell-based clusters.
Mar 25, 2026 425 words in the original blog post.
On March 10, 2026, a scheduled power infrastructure maintenance by a provider unexpectedly escalated into a broader data center power incident, affecting multiple services in the us-central1 region. Initially planned to have minimal customer impact, the maintenance resulted in several unplanned power failures, leading to a series of disruptions including loss of external VM connectivity, inaccessibility of Managed Kubernetes clusters, and interruptions in public S3 endpoints. The incident unfolded in multiple stages over several hours, with repeated power interruptions complicating recovery efforts. The root cause was linked to the maintenance's deviation from the planned scope, involving incorrect switching and failure of rack ATS units during a temporary power configuration. This incident highlighted the need for improved communication and planning with the power provider, more conservative preparation standards, and the development of a counter-emergency recovery procedure to ensure faster recovery in future incidents. The post-incident action plan includes revising maintenance communication standards, enhancing planning quality for future electrical work, and conducting regular recovery drills to ensure a more efficient and predictable restoration process.
Mar 20, 2026 1,020 words in the original blog post.
AI agents are being implemented in live business workflows where performance, governance, and cost efficiency are as crucial as model quality. At NVIDIA GTC 2026, Nebius and DataRobot, in collaboration with NVIDIA, introduced a validated AI Factory stack designed for production-level agent workloads. This stack, supported on Nebius AI Cloud, combines DataRobot’s Agent Workforce Platform with NVIDIA’s AI infrastructure to provide comprehensive lifecycle management, governance, and cost control for AI agents. Nebius AI Cloud offers the foundational infrastructure for agent workloads, while Nebius Token Factory facilitates serverless model access for token generation. Benchmarking on NVIDIA HGX B200 systems demonstrated high throughput and low latency, validating the stack's efficiency and cost-effectiveness. The stack ensures safe and reliable AI agent operations in production through integrated governance and policy enforcement mechanisms, with further details and demonstrations available at the NVIDIA GTC 2026 event.
Mar 18, 2026 608 words in the original blog post.
Nebius and Eigen AI have partnered to bring optimized open-source AI models to Nebius’s Token Factory, a managed platform for model inference. This collaboration focuses on enhancing models such as DeepSeek, GLM, GPT-OSS, and Qwen by integrating Eigen AI's expertise in model optimization and serving systems with Token Factory's autoscaling and fine-tuning capabilities. The partnership aims to streamline the deployment of open-source models by offering developers API access and managed solutions, reducing the need for custom infrastructure and enabling quick transition from experimentation to production. Eigen AI's optimizations, which include quantization and GPU utilization improvements, have led to leading performance benchmarks for models like GPT-OSS-120B and Qwen3 Coder 480B, ensuring efficient and cost-effective production use. This initiative addresses the challenge of running complex open models at scale, providing a ready-to-use solution that supports enterprise needs without the overhead of managing underlying infrastructure.
Mar 17, 2026 970 words in the original blog post.
AI demos often struggle with real-world complexity due to reliance on fragmented data and the need for multi-step reasoning across systems, where failures are typically caused by orchestration gaps and infrastructure variability rather than model quality. To address these issues, Nexla and Nebius offer a cohesive solution by turning enterprise data into governed, agent-ready data products and ensuring reliable pipeline operation with dedicated computing infrastructure. This approach allows the creation of production-grade agent systems from raw data, as demonstrated in a live presentation at NVIDIA GTC, where Nexla, Nebius, Tripadvisor, and NVIDIA showcased a coordinated, multi-agent workflow for travel planning. The demo involved parsing video content to generate travel itineraries using Tripadvisor's dataset, highlighting the potential of AI systems in practical applications. This collaboration exemplifies the transition from prototype to production-ready AI systems capable of sustained real-time inference.
Mar 17, 2026 316 words in the original blog post.
On February 26, 2026, a power infrastructure fault in the eu-north-1 region caused a brief power interruption, impacting a subset of compute hosts and leading to temporary unavailability of some customer workloads. The incident began with a short circuit in cabling for cooling infrastructure, triggering UPS overcurrent protection and resulting in simultaneous host restarts. While core systems initiated recovery automatically, recovery constraints in compute and storage layers prolonged the restoration process, with some virtual machines experiencing restarts and temporary storage access issues. The root cause was identified as a high overcurrent condition due to the short circuit, which activated UPS protection mechanisms, and recovery efforts were hindered by throughput limitations and storage attachment conflicts. Following the incident, improvements are being implemented to refine electrical protection, enhance recovery throughput, and improve monitoring and alerting systems to better handle future incidents and ensure platform resilience.
Mar 17, 2026 678 words in the original blog post.
The Inference Frontier Program is a new initiative designed to share and recognize the practical innovations of engineers and builders working on AI inference in real-world applications. Launched alongside the Nebius.Build SF event, the program aims to address the challenges of running AI models at scale, such as handling bursty traffic, optimizing runtime, and ensuring stable, cost-effective operations. By inviting submissions from a diverse range of contributors, including solo developers, open-source maintainers, and enterprise teams, the program seeks to foster a community of builders sharing insights on inference optimizations. Through continuous nominations and technical deep dives, the program will highlight successful architectural strategies, offering distribution and networking opportunities for participants. A panel of industry experts will evaluate submissions, ensuring that technical craftsmanship is celebrated, while maintaining sensitivity to proprietary data.
Mar 15, 2026 945 words in the original blog post.
AI Cloud represents a significant advancement in cloud computing, specifically designed to meet the complex demands of modern machine learning (ML) and large language model (LLM) workloads. Unlike traditional cloud platforms, AI Clouds are purpose-built, offering a unified environment where specialized hardware, high-speed networking, and integrated MLOps tools work in concert to facilitate seamless model building, training, and deployment. This infrastructure centers around high-performance GPU clusters interconnected by NVLink and InfiniBand, enhancing scalability and performance for tasks like distributed training and inference. Managed services automate the entire ML lifecycle, allowing developers to focus on models and data rather than infrastructure management. This approach ensures high-speed operations, reproducibility, and operational control, with AI Clouds providing capabilities far beyond mere compute rental, effectively bridging the gap between research and production. As AI projects grow in complexity and scale, AI Clouds like Nebius offer a robust solution by integrating compute, data storage, and MLOps into a single, cohesive platform, allowing teams to efficiently transition from experimentation to deployment without manual configuration or environment switching.
Mar 13, 2026 2,474 words in the original blog post.
NVIDIA Nemotron 3 Super, now accessible on Nebius Token Factory, is a 120 billion parameter hybrid Mixture of Experts (MoE) model designed for multi-agent applications and complex reasoning tasks, featuring 12 billion active parameters per inference step and capable of handling up to 1 million token context length. Optimized for agentic systems, it utilizes a hybrid Transformer–Mamba architecture with MoE routing to enhance compute efficiency while maintaining high reasoning performance. The model is suited for various production use cases, including software development workflows, deep research agents, financial document processing, and cybersecurity analysis. Offering open weights, datasets, and training recipes, it supports multi-token prediction for expedited long-form generation. Nemotron 3 Super can be deployed on Nebius Token Factory through dedicated GPU endpoints with autoscaling throughput and OpenAI-compatible API integration, providing options for EU or US deployment with optional zero-retention inference. This setup allows teams to transition from model access to production deployment without managing GPU clusters, available for deployment via API or testing in the Playground.
Mar 11, 2026 229 words in the original blog post.
OpenClaw has transitioned from being a developer-oriented installation guide to a crucial infrastructure component for AI agent deployment, raising significant questions about its security for production use. As a self-hosted AI agent gateway, it serves as a critical security boundary, managing messaging channels, sandboxed tool execution, and model inference, and has become pivotal due to incidents involving malicious activities and security breaches. The gateway connects with platforms like WhatsApp, Telegram, and Slack, providing a robust routing and session management system that persists conversation histories and supports various integrations, but its open-source, self-hosted nature requires users to manage their own deployment, security, and updates. OpenClaw's architecture, featuring a lightweight Gateway process, supports a flexible skill system, allows for detailed control over tool execution and sandboxing, and uses a Markdown-based memory subsystem, but it requires vigilant security measures, including configuration hardening, access control, and regular security audits, especially given past incidents of skill-based vulnerabilities. As AI agent platforms evolve, the demand for on-device agents is increasing, and OpenClaw positions itself as a flexible, on-premise solution that allows for secure, scalable deployments, integrating with external services for enhanced inference capabilities, while requiring users to actively manage its complex security landscape.
Mar 05, 2026 3,774 words in the original blog post.
SWE-rebench-V2 is an expansive dataset designed to enhance the training of autonomous software engineering agents using reinforcement learning by providing access to a large-scale, diverse set of executable tasks. This new iteration addresses the limitation of limited access to open-source data by utilizing a fully automated, language-agnostic pipeline to extract real-world software engineering tasks, resulting in over 32,000 executable tasks complete with pre-built Docker environments and coverage of 20 programming languages, including less commonly used ones like Lua and Scala. The dataset also features more than 100,000 additional tasks derived from pull requests and is designed to facilitate multilingual RL training by enabling research on cross-language reasoning and robust performance beyond Python-centric datasets. Each task includes a pre-configured Docker container for easy reproducibility, and the environments are automatically set up using an interactive agent that resolves dependencies. The tasks are quality-filtered and labeled using large language models (LLMs), with structured metadata that includes method signatures and problem descriptions to provide comprehensive training signals. Accompanying the dataset is a technical report that explains the extraction pipeline, filtering methods, and includes a diagnostic study assessing modern models on these tasks.
Mar 03, 2026 238 words in the original blog post.