Home / Companies / Nebius / Blog / November 2025

November 2025 Summaries

6 posts from Nebius

Filter
Month: Year:
Post Summaries Back to Blog
Nebius has demonstrated exceptional performance in the MLPerf® Training v5.1 benchmark by utilizing the latest NVIDIA Blackwell and Blackwell Ultra systems, showcasing their capabilities in training next-generation GenAI models. This achievement highlights Nebius' commitment to transparency, collaboration, and excellence in AI training, driven by a combination of hardware innovation and software-level optimization. The company successfully benchmarked several models, including Llama-3.1-8B and FLUX.1, achieving seven first-place results out of nine submissions. Notably, the NVIDIA HGX B300 system exhibited a significant 12.6% reduction in training time compared to the HGX B200, demonstrating its superior speed and stability. Nebius also achieved impressive scalability with the HGX B200 system, showing a ~3.1x speed-up when increasing cluster size from 8 to 32 GPUs, which facilitates faster model training and deployment. As a Reference Platform NVIDIA Cloud Partner, Nebius ensures optimal utilization and stability in large-scale distributed training environments, empowering customers to achieve high GPU utilization in the cloud and accelerate AI development.
Nov 12, 2025 842 words in the original blog post.
AI foundation models, specifically Boltz-2, are revolutionizing drug discovery by rapidly predicting molecular interactions with disease-related proteins, significantly enhancing the speed and accuracy of identifying potential therapeutic compounds. Boltz-2 achieves near-experimental accuracy at a fraction of the computational cost of traditional methods, allowing high-throughput screening of hundreds of thousands of compounds per day when used in high-performance computing environments. The model's integration into Nebius AI Cloud utilizes Managed Kubernetes to orchestrate GPU resources and shared filesystems, facilitating efficient, large-scale biomolecular inference. This setup supports the entire drug discovery pipeline from hit discovery to lead optimization by enabling rapid, scalable, and cost-effective simulations. The framework described includes the deployment of Boltz-2 in a secure, compliant cloud environment, emphasizing the importance of efficient data management and workflow orchestration to support intensive drug discovery workloads.
Nov 12, 2025 1,357 words in the original blog post.
Nebius and Anyscale have formed a partnership to enhance platform integration, making it easier, faster, and more cost-effective for teams to deploy and scale Python and AI workloads using Ray. This collaboration merges Nebius AI Cloud with Anyscale's managed Ray platform, providing a comprehensive solution for developers and platform teams to transition multimodal AI workloads from code to production efficiently. As modern AI increasingly relies on multimodal data, including video, images, and audio, traditional infrastructures face challenges in handling this complexity. Ray offers a cohesive framework for scaling AI workloads across diverse compute environments, but it requires a robust production platform for optimal performance. The integration with Nebius AI Cloud allows teams to deploy production-ready Ray infrastructure effortlessly, offering features like seamless cluster management, developer tools, and observability enhancements. This setup improves performance and efficiency by optimizing resource use and supporting reliable execution on Nebius’s cost-effective Virtual Machines, ultimately enabling high-performance, scalable AI model deployment.
Nov 11, 2025 630 words in the original blog post.
Nebius AI Cloud 3.0, known as "Aether," introduces significant advancements in security, compliance, governance, developer productivity, and performance, aimed at providing enterprise-grade capabilities without hindering AI developers. The platform now supports enhanced security standards such as SOC 2 Type II and ISO 27001, alongside improved secrets management and granular IAM controls, complemented by a refined UI and deeper integration with tools like SkyPilot Server. The UK sees its first deployment with NVIDIA Blackwell Ultra AI infrastructure, offering robust performance for generative AI and foundational model development. Additionally, the new Nebius Token Factory enables enterprises to transform open-source AI checkpoints into production-ready systems, supporting sophisticated models like NVIDIA Nemotron Nano 2 VL. In collaboration with Accenture, Nebius aims to deliver a full-stack sovereign AI cloud, facilitating fast and secure deployment for enterprises and public institutions. The platform also expands its tech capabilities, including improved security measures, multi-tenant IAM support, and integration with Anyscale for large-scale AI workloads, alongside the transition to unified billing rates and enhanced regional service clarity.
Nov 07, 2025 585 words in the original blog post.
Software engineering agents, powered by large language models (LLMs), have rapidly advanced, yet the technical challenges of large-scale experimentation persist due to the need for distributed orchestration beyond single-machine capacity. Nebius’ AI R&D team has focused on building scalable infrastructure to support such experiments, involving the creation of extensive datasets and evaluation pipelines like SWE-bench and SWE-rebench. Their work necessitated using distributed systems such as Kubernetes and TractoAI to manage the orchestration of thousands of agent runs and evaluations, while overcoming challenges related to data processing, container management, and task execution. By leveraging Kubernetes for flexible agent orchestration and TractoAI for efficient data processing and evaluation, the team developed a robust framework that allows for efficient experimentation and sharing with the research community. This infrastructure is designed to facilitate automated and reliable performance measurement of SWE agents, which operate by executing code within containers, much like a human engineer, but with the added complexity of requiring a scalable and distributed backend to handle the vast amounts of data and workload demands. Nebius is now opening this infrastructure to the broader research community, including offering support through their research credits program, to accelerate progress in the field of software engineering agents.
Nov 07, 2025 3,179 words in the original blog post.
Updates to the Nebius Status Board aim to improve transparency during service outages by providing customers with detailed visibility into the availability of Nebius AI Cloud services. The board now features regional views, allowing users to monitor the operational status of platform services in specific geographic regions. This enhancement helps users quickly assess the impact of service degradations or scheduled maintenance by showing which regions and services are affected. In the process of implementing regional views, all incidents before October 27, 2025, have been assigned to the eu-north1 region, impacting its uptime value, while other regions' uptime values are calculated from incidents occurring on or after this date.
Nov 05, 2025 173 words in the original blog post.