June 2025 Summaries
9 posts from Nebius
Filter
Month:
Year:
Post Summaries
Back to Blog
AI infrastructure is revolutionizing fields like precision medicine, neuroscience, and drug discovery by enabling faster and more scalable analysis of biological data. Companies such as Nebius and NVIDIA are collaborating to provide biotech startups with the necessary tools and infrastructure, including a full-stack AI Cloud with NVIDIA GPU instances and AI software suites like BioNeMo and Parabricks. This infrastructure supports advanced research methods, such as Prima Mente's epigenetic-driven neuroscience models and Simulacra AI's quantum chemistry simulations. Converge Bio is using generative AI to analyze transcriptomic data for precision medicine, demonstrating the technology's potential to create tailor-made therapies. These advancements indicate a new wave in personalized medicine, emphasizing the creation of novel therapies through AI, which is not only analyzing data but also inventing new molecular structures. Through these innovations, AI is helping to overcome operational and financial barriers, enabling biotech teams to focus on scientific discoveries and the development of next-generation medical treatments.
Jun 25, 2025
1,205 words in the original blog post.
Object storage is becoming increasingly popular among IT professionals, cloud architects, and data storage specialists due to its scalability, flexibility, and suitability for handling large volumes of unstructured data, such as photos, videos, and backups. Unlike traditional block storage, which organizes data into fixed-size blocks ideal for structured tasks like databases, object storage organizes data into discrete objects containing the data itself, metadata, and a unique identifier, enabling easier access and management through APIs. The architecture of object storage allows for horizontal scaling, making it an ideal choice for modern cloud environments where data demands grow rapidly. It is cost-effective, as it runs on standard hardware, and features built-in redundancy that ensures data durability and availability. Traditional storage systems like block and file storage are often limited by their hardware and structure, making them less flexible and harder to scale. As data volumes continue to increase and businesses increasingly rely on cloud-based solutions, object storage, with its remote accessibility and adaptability, is positioned as a cornerstone of future data infrastructure, providing a stable and cost-effective solution for modern digital demands.
Jun 20, 2025
1,902 words in the original blog post.
Training AI models involves complex cost considerations that depend on multiple factors, including model architecture, dataset size, infrastructure setup, and utilization efficiency. Costs are primarily driven by compute resources, particularly GPU hours, with storage, networking, and orchestration contributing additional expenses. For instance, training large models like GPT-3 or BLOOM can run into millions of dollars, while smaller models like BERT-Large can still be costly without optimization. Effective cost management requires balancing compute, storage, and network resources, and leveraging strategies such as mixed precision, data parallelism, and efficient checkpointing to maximize utilization and minimize waste. Cloud resources offer scalability and flexibility, but on-premises solutions can be more cost-effective for consistent workloads, leading many teams to adopt a hybrid approach. Overall, aligning infrastructure planning with clear success metrics and robust operational practices is essential for optimizing AI training budgets.
Jun 19, 2025
2,188 words in the original blog post.
Managed Soperator, a fully managed Slurm-on-Kubernetes solution, is now available for self-service, allowing users to quickly set up a Slurm training cluster with NVIDIA GPUs and pre-installed libraries, facilitating immediate machine learning training. Developed as a managed service on the Nebius AI Cloud, it aims to simplify the user experience by automating infrastructure provisioning and configuration, which traditionally required extensive manual setup. Soperator, the core technology behind this solution, is an open-source Kubernetes operator for Slurm, originally released last autumn, and enables rapid deployment of large GPU clusters while ensuring fault tolerance and scalability. With three options available—Managed Soperator, Professional Soperator, and open-source Soperator—the platform caters to different needs, from self-service to customized large-scale installations, each supporting various AI/ML drivers and libraries. This solution is designed to empower AI developers by minimizing operational overhead and focusing on innovative research, with ongoing developments to enhance its features further.
Jun 18, 2025
739 words in the original blog post.
Kubernetes is a powerful container orchestration system that brings order to the complexity of managing AI workloads, from model training to deployment and inference, by automating resource allocation and scaling services based on demand. It abstracts the underlying hardware, enabling infrastructure as code for reproducible environments, and isolates processes in containers, making each pipeline stage independent and conflict-free. By ensuring fault tolerance, observability, and reproducibility, Kubernetes turns sprawling infrastructures into stable, self-managing platforms for AI development and deployment, allowing teams to focus more on models and less on manual recovery. Despite its flexibility, leveraging Kubernetes effectively for AI workloads requires careful attention to resource allocation, observability, and scaling, as well as a mature operational approach to overcome challenges such as GPU management, debugging, and operational complexity. With the right configuration and understanding of AI workload nature, Kubernetes serves as a robust foundation for developing, running, and maintaining AI workloads, simplifying deployment, streamlining resource management, automating scaling, and increasing system resilience.
Jun 16, 2025
2,145 words in the original blog post.
Nebius has introduced SWE-rebench, a large-scale dataset designed to enhance the development of software engineering (SWE) agents based on large language models (LLMs). This initiative aims to democratize AI and support developers by providing over 21,000 interactive tasks sourced from more than 3,400 GitHub repositories through an automated pipeline. The dataset features rich annotations, including installation configurations, dependency versions, and quality scores assessed by LLMs. Accompanying the dataset is a technical report detailing the automated task collection and dataset construction process, highlighting innovations for continuous task mining. SWE-rebench is anticipated to be a crucial resource for developing and benchmarking new models on realistic SWE tasks, with a curated subset already used for a public leaderboard that evaluates LLMs on real-world tasks.
Jun 10, 2025
224 words in the original blog post.
Modern AI workloads necessitate advanced GPU cluster networking technologies to effectively manage the immense computational power required for training complex models. These technologies ensure efficient data exchange and synchronization across vast networks of GPUs, which are crucial for maintaining real-time training and scalability. GPU cluster networks operate over three layers, with intra-server and inter-server communication facilitated by technologies like PCIe, NVLink, and RDMA. Various network topologies, such as fat-tree, torus, and dragonfly, offer different approaches to maximizing efficiency, reducing latency, and ensuring fault tolerance. Effective network management includes optimizing topologies, implementing traffic optimization, and monitoring performance to minimize costs and enhance reliability. Emerging trends in GPU networking, such as photonics-based switches and AI-optimized network architectures, promise further advancements in speed and efficiency, particularly in cloud and edge AI contexts. Nebius provides solutions for setting up and managing GPU clusters, emphasizing the importance of both physical and software-based strategies to achieve optimal performance in AI training environments.
Jun 09, 2025
1,591 words in the original blog post.
NVIDIA GTC Paris and VivaTech 2025 will host Nebius, an AI cloud provider aiming to establish itself as the default choice in Europe, from June 11-13 in Paris, where they will showcase their advancements and research. They offer access to powerful GPU resources and have demonstrated significant AI training performance, evidenced by their achievements in the MLPerf® Training benchmark. Nebius is involved in various projects, such as enhancing DeepSeek R1's performance and supporting Converge Bio's precision medicine developments, while also advancing large language models for molecular drug design. Their technical documentation highlights recent improvements in observability, billing, and Kubernetes management, reflecting their commitment to enhancing infrastructure visibility and reliability. The company also recognizes innovation in AI applications through its AI Discovery Award, spotlighting startups in healthcare and life sciences.
Jun 06, 2025
773 words in the original blog post.
Nebius has announced its first submission of MLPerf® Training v5.0 results, achieving notable performance in training the Llama 3.1 405B model using NVIDIA Hopper GPU clusters interconnected with NVIDIA Quantum-2 InfiniBand networking. This submission underscores Nebius's ability to deliver scalable and predictable training performance, with the benchmarks serving as a reliable measure of AI cloud infrastructure efficiency. The MLPerf® benchmarks provide a standardized framework developed by industry and academic experts to evaluate AI model performance in realistic scenarios, aiding potential customers in making informed decisions. Nebius's achievement, particularly on large clusters with 512 and 1,024 GPUs, highlights near-linear scaling capabilities and cost-effectiveness in distributed model training. The company emphasizes its commitment to continuous improvement and optimization of AI cloud infrastructure, leveraging proprietary software and hardware configurations to maintain high performance and reliability. As an NVIDIA Cloud Partner, Nebius aligns with the latest NVIDIA technologies to enhance its product offerings, ensuring robust GPU utilization and operational simplicity.
Jun 05, 2025
784 words in the original blog post.