Home / Companies / Nebius / Blog / October 2025

October 2025 Summaries

7 posts from Nebius

Filter
Month: Year:
Post Summaries Back to Blog
Nebius AI Studio has introduced NVIDIA Nemotron Nano 2 VL, a compact and efficient multimodal reasoning model, which is designed for tasks like document intelligence and video understanding. Utilizing the hybrid Mamba-Transformer architecture, this model achieves high accuracy without the drawbacks of large models, such as increased cost and latency. The model is part of the NVIDIA Nemotron family, which offers open models, datasets, and recipes, providing transparency and flexibility for developers to create domain-specific AI systems. Nebius AI Studio supports this with a high-performance, OpenAI-compatible inference platform that allows for efficient scaling and deployment. Developers can use Nemotron Nano 2 VL to build applications such as document-intelligent assistants, video summarization tools, and media curation pipelines, benefiting from its low latency and cost-effective deployment on Nebius's infrastructure.
Oct 28, 2025 389 words in the original blog post.
The integration of SkyPilot with Nebius infrastructure since April 2025 has simplified the deployment of AI workloads, enabling developers to provision GPU instances and mount cloud storage with ease. Initially relying on a local API server model, this approach posed limitations for larger teams, prompting the introduction of a Managed SkyPilot API Server on Nebius AI Cloud. This fully managed solution eliminates operational overhead by automating server deployment, enhancing productivity, collaboration, and operational excellence through centralized infrastructure management. It allows asynchronous job execution, seamless workflow integration, and resource sharing, while providing unified visibility and fault tolerance. The managed service targets small to mid-size ML teams, offering one-click deployment and zero operational overhead, contrasting with the complexities of self-hosting. Current capabilities include support for Nebius platforms and Kubernetes clusters, with plans to extend support to additional cloud providers and enhance enterprise integration features. This development underscores a commitment to democratizing enterprise-grade AI infrastructure for teams of all sizes.
Oct 23, 2025 1,374 words in the original blog post.
Nebius has been utilizing the NVIDIA Grace Blackwell platform to enhance data center AI infrastructure by leveraging the fifth-generation NVIDIA NVLink™ scale-up fabric, which significantly improves GPU-to-GPU communication bandwidth. The new NVIDIA GB200 NVL72 architecture supports up to 72 GPUs in a single NVLink domain, enabling efficient AI workload distribution and reducing communication overhead for tasks like pre-training large language models such as the Nemotron-4 340B LLM. This architecture utilizes a combination of Tensor, Pipeline, and Data parallelism to optimize performance, with the NVLink fabric providing high-speed connectivity within racks and InfiniBand used for inter-rack communication. For optimal performance on the GB200 NVL72, workloads must be carefully engineered to maximize the benefits of NVLink connectivity, requiring an understanding of parallelism group creation and communication patterns. Nebius offers insights and assistance for those planning to design workloads for NVIDIA GB200 NVL72 or GB300 NVL72, emphasizing the importance of proper setup to fully leverage the platform's capabilities.
Oct 23, 2025 1,940 words in the original blog post.
Nebius has launched its AI Cloud 3.0, named "Aether," designed to address the challenges enterprises face in scaling AI projects from experiments to business-critical systems, particularly in regulated industries like healthcare and finance. Key features of this release include enhanced compliance certifications, advanced IAM capabilities, and new observability tools that provide granular control and improve collaboration without introducing bureaucratic delays. The platform also boasts significant reliability and performance improvements, such as active health checks and self-healing nodes, alongside faster file storage speeds. Developer experience is enhanced through a refreshed UI, simplified resource allocation, and easier integration with containerized applications and ecosystem platforms. This release reflects Nebius's commitment to evolving its platform based on customer feedback and will be further explored in their upcoming webinar.
Oct 14, 2025 920 words in the original blog post.
Nebius, an AI-native cloud infrastructure provider, has announced the achievement of significant security and compliance milestones, underscoring their commitment to data protection, privacy, and operational resilience. The company has undergone independent third-party audits, verifying that their security controls meet the Service Organization Control (SOC) 2 Type II standards, including HIPAA compliance, and align with the principles of NIS2 and DORA. Additionally, Nebius has obtained ISO 27001 certification and integrated principles from various ISO standards such as ISO 27701, 27018, and 27799, further enhancing their Information Security Management System (ISMS) and Privacy Information Management System (PIMS). These certifications and audits demonstrate Nebius's robust measures in safeguarding data and ensuring business continuity, with rigorous security and privacy controls embedded in their systems. The company's infrastructure is designed to protect data at all stages, supported by stringent software development processes, vulnerability management strategies, and incident reporting protocols. As a trusted partner for AI development and cloud services, Nebius continues to prioritize security-by-default in its operations, providing enterprise-level protection and compliance with international standards, thereby ensuring reliable and secure service delivery across various sectors, including healthcare and finance.
Oct 14, 2025 1,261 words in the original blog post.
Generative video workloads present significant systems engineering challenges, as they demand extensive GPU memory, extended runtimes, and heightened stability compared to text or image generation. Nebius AI Cloud, in collaboration with the Baseten Inference Stack, addresses these challenges by providing a robust infrastructure that supports scalable text-to-video production. This partnership leverages dedicated GPU clusters, elastic provisioning, and intelligent autoscaling to maintain performance and cost-efficiency even during demand spikes, while Baseten's optimized runtime and orchestration features ensure effective resource utilization. The infrastructure prioritizes performance consistency and reliability through SLA-aware autoscaling, seamless integration into existing pipelines, and multi-zone availability across the US and Europe. By combining Nebius' AI cloud foundation with Baseten's advanced inference stack, the system effectively manages latency and operational overhead, making it well-suited for production-grade video generation tasks.
Oct 06, 2025 836 words in the original blog post.
Nebius has announced a significant multi-billion dollar agreement with Microsoft aimed at enhancing their AI cloud infrastructure until 2026, starting with providing dedicated capacity from a new data center in Vineland, New Jersey. The partnership underscores Nebius's commitment to delivering a machine learning-centric stack for developers, complemented by their achievement of NVIDIA Exemplar Cloud Status for effective and reliable performance with NVIDIA H200 GPUs. Nebius has demonstrated superior performance in AI systems through the MLPerf® Inference v5.1 benchmarks, showcasing their advanced capabilities. Additionally, Nebius has expanded its technical documentation, providing comprehensive support for various features, including running containers over VMs, exporting data, and managing data protection. The updated documentation structure aligns with the web console navigation, offering clearer insights into Nebius AI Cloud's capabilities and third-party integrations.
Oct 03, 2025 340 words in the original blog post.