Home / Companies / Nebius / Blog / March 2025

March 2025 Summaries

11 posts from Nebius

Filter
Month: Year:
Post Summaries Back to Blog
Managed Service for MLflow has been made publicly available, following its preview launch last September, with improvements in stability to enhance user experience. This service combines core MLOps functionality with cloud deployment, significantly simplifying and accelerating the model development process for data scientists and ML engineers. MLflow is an essential tool for tracking model runs and experiments, streamlining model management, and enhancing cross-team collaboration, ensuring model reproducibility and metadata consistency across various projects, including large-scale generative AI agents. Nebius offers this managed service as a cloud-based solution requiring minimal infrastructure deployment and maintenance, providing a SaaS-like experience where users can simply log in, set up an MLflow cluster, and interact with model artifacts via the MLflow UI. To demonstrate its practical benefits, a guide is available for fine-tuning GenAI models, and users can start by signing up on the platform and ensuring they have a minimum balance of $25 to create an MLflow cluster connected to their training environment.
Mar 31, 2025 333 words in the original blog post.
Nebius has launched a simplified AI development platform by integrating JupyterLab into their Standalone Applications service, eliminating the need for complex infrastructure management like Kubernetes configuration. This fully managed offering, now available on NVIDIA GPUs, allows users to deploy JupyterLab with ease, providing a seamless development experience that supports quick data analysis, prototyping, and model evaluation. JupyterLab, a vital tool for ML engineers and data scientists, is highlighted for its ability to combine code execution, text, and visualizations in a flexible web-based interface. The initial release offers features such as a pre-configured PyTorch framework, persistent storage, and easy package management, though it has some limitations regarding runtime and instance numbers. With a starting price of $2.95 per hour for the NVIDIA H100 GPU configuration, Nebius aims to further enhance this service by adding more AI applications, improving integration, and expanding collaboration features. The goal is to streamline the AI development process, allowing professionals to concentrate on innovative problem-solving and solution-building.
Mar 27, 2025 652 words in the original blog post.
The blog post illustrates the process of fine-tuning large language models (LLMs) like DeepSeek-V3 and Qwen-2.5-72B for domain-specific tasks using Nebius AI Studio, focusing on a function-calling task as a practical example. It guides the reader through each step, from dataset preparation using the ToolACE dataset to model evaluation, and emphasizes the importance of fine-tuning for enhancing model performance in specialized applications. The blog provides a comprehensive walkthrough, including code snippets in a Jupyter notebook, and highlights the cost-effective approach of using LoRA adapters for fine-tuning with the 'Instruct' version of Llama-3.1-8B. The post concludes by demonstrating how the fine-tuned model outperforms the original in various tasks and discusses the advantages of tailoring LLMs to specific needs, ultimately improving their quality on target tasks.
Mar 26, 2025 3,522 words in the original blog post.
Computationally intensive tasks, essential for scientific breakthroughs, can be significantly accelerated by leveraging the power of cloud-based GPUs over traditional CPUs, as demonstrated in a study comparing their performance in bioinformatics workflows. Specifically, the NVIDIA H100 GPU, available on Nebius AI Cloud, showcased an impressive ability to speed up processes such as iterative searches, vector embeddings, clustering, and functional annotation tasks, with performance enhancements ranging from two to 26 times faster than an 8-core Intel Xeon Platinum CPU. This study, inspired by CRISPR-based research, focused on exploring non-CRISPR archaeal defense systems, using deep learning models like ESM-Cambrian for protein embeddings and ML-driven clustering with K-Means. The GPU's parallel processing capabilities were particularly beneficial for tasks involving large datasets, such as dimensionality reduction using UMAP, demonstrating its suitability for complex scientific analyses. The research underscores the role of GPU acceleration in advancing bioinformatics, with a planned webinar to further explore its applications and a detailed implementation guide available on GitHub.
Mar 24, 2025 793 words in the original blog post.
On March 13, 2025, a release of a control plane component in the VPC service's eu-west1 region led to significant networking disruptions for Managed Services for Kubernetes, due to a bug in the network configuration data processing that affected the IP alias mechanism. Although the bug had existed in earlier versions, it was triggered by a new feature in the latest release, complicating efforts to mitigate the incident through rollback procedures. From 13:36 to 18:11 UTC, nearly all pods in the eu-west1 region experienced connectivity loss, while the eu-north1 region remained unaffected. The incident was traced to a bug in the VPC control plane's handling of BGP messages, where route distinguishers were partially ignored, leading to incorrect merging of IP alias announcements. This issue was difficult to detect in testing and canary deployments due to specific conditions required for its manifestation. The response plan includes improving service observability, enhancing incident response capabilities, refining testing and deployment procedures, and introducing alerts on critical network metrics.
Mar 21, 2025 856 words in the original blog post.
Nebius AI Cloud and Metaflow, in collaboration with Outerbounds, offer a robust infrastructure and software stack designed for developing and deploying AI and ML systems efficiently. Nebius AI Cloud provides scalable, high-performance AI infrastructure equipped with NVIDIA GPUs and key services like Compute Cloud and S3-compatible Object Storage, ensuring secure and cost-effective operations. Metaflow, originally developed by Netflix and now open-source, complements this by offering developer-friendly APIs that support a wide range of ML and AI applications, enabling seamless integration with Nebius for comprehensive workflow orchestration and deployment. Outerbounds enhances this ecosystem by offering a unified platform that supports secure, scalable Metaflow workflows across different cloud environments, employing a Bring-Your-Own-Cloud model to optimize costs and performance for enterprise needs. This integrated stack allows teams to iterate rapidly, deploy sophisticated models, and streamline AI development, providing a developer-friendly experience with tools like Torchtune for fine-tuning and comprehensive monitoring and optimization capabilities to ensure reliability and efficiency.
Mar 11, 2025 1,405 words in the original blog post.
Nebius is significantly expanding its computational infrastructure with a new 300 MW region in New Jersey, set to go live this summer in collaboration with AI hosting company DataOne, and a colocation facility in Iceland with Verne, expected to launch in March. This growth brings Nebius's total regions to five, complementing their existing facilities in Finland, France, and Missouri. They are also the Platinum sponsor of NVIDIA GTC 2025, where they will host a booth, offer NVIDIA GPU credits, and present tech talks focusing on building efficient AI cloud platforms and advancing agentic systems. Nebius AI Studio has been rebranded as Nebius Token Factory, now offering fine-tuning capabilities to customize open-source models, as evidenced by Chatfuel's success with Llama-405B models. Managed PostgreSQL has reached general availability, providing a robust cloud-based solution for structured data storage. Nebius has launched Kvax, an open-source Flash Attention implementation for JAX, boasting superior performance in long-context training. Additionally, they have introduced the AI Discovery award for AI-driven startups in healthtech, granting $100,000 in AI Cloud credits. The new Nebius research credits program aims to assist researchers with AI Cloud access, while an event in San Francisco will close the ‘Nebius AI Cloud Unveiled’ series, offering insights into their AI Cloud developments. Updates to their technical documentation include enhanced resources for Managed PostgreSQL, new tutorials, and guidance on performance optimization and monitoring for VMs and storage volumes in the Nebius AI Cloud.
Mar 07, 2025 711 words in the original blog post.
Nebius, in collaboration with DataOne, is constructing a custom-built data center in New Jersey, utilizing innovative energy-efficient designs and aiming for completion within 20 weeks to support AI workloads. This facility will expand up to 300 MW, with an initial commitment of 100 MW by the end of 2025 and flexibility to increase capacity as needed. Concurrently, Nebius is establishing a 10 MW compute cluster in Keflavik, Iceland, powered entirely by renewable energy, with full operational capacity expected by March. Additional expansions include a second phase in a Kansas City data center and existing infrastructure in Finland and France. This expansion strategy is designed to address the growing demands of AI innovators across the U.S. and Europe, providing scalable and reliable compute resources.
Mar 05, 2025 362 words in the original blog post.
Generic AI models often struggle with specialized tasks due to their need for complex prompts and large context windows, but fine-tuning offers a solution by customizing models to meet specific domain requirements. Nebius AI Studio facilitates this customization by allowing users to transform generic models into efficient, specialized solutions through fine-tuning, which enhances accuracy, reduces costs, and ensures consistent outputs. The platform supports over 30 leading models, including the Llama 3 and Qwen series, and offers LoRA fine-tuning and full fine-tuning for models under 20 billion parameters. Designed for production, it features scalable GPU clusters and FlashAttention-3 for optimized training speed, and integrates seamlessly with the OpenAI SDK. Deployment options include serverless, on-demand, or reserved clusters, providing flexibility as models scale. The pricing model is usage-based, charging only for tokens processed during training, and ensures transparency with no hidden fees. Getting started is streamlined, allowing users to select a base model, choose a fine-tuning approach, configure parameters, and deploy their customized model, with a promotional offer available for new users.
Mar 05, 2025 374 words in the original blog post.
Managed Service for PostgreSQL has transitioned from preview to general availability, introducing service pricing and SLAs. This service is particularly suitable for AI workloads, leveraging PostgreSQL's capability to handle large datasets and store vector embeddings using extensions like pgvector. The fully managed service alleviates administrative tasks such as deployment, maintenance, network configuration, and system observability. With general availability, users benefit from an enhanced monitoring dashboard and advanced backup management, allowing more control over data retention and restoration. Previously free during the preview, the service now has a structured pricing model based on server configuration, encouraging users to review and adjust their usage. Existing preview deployments will automatically transition to the paid service, and users can access PostgreSQL through various interfaces like web, CLI, or Terraform.
Mar 04, 2025 504 words in the original blog post.
The rapid growth of AI, particularly in training and running large foundational models, has necessitated the need for powerful computational infrastructures, with GPU clusters being a key solution. These clusters, often comprising dozens to thousands of GPUs, are essential for handling the intense computational demands of Generative AI and large language models (LLMs), which cannot be managed by a single GPU. GPU clusters facilitate parallel computing, breaking down large tasks into smaller operations assigned to interconnected GPUs, thereby accelerating processes such as model training, fine-tuning, and inferencing. The orchestration of these clusters involves both hardware—comprising head and worker nodes equipped with GPUs, CPUs, RAM, and NICs—and software, with tools like Kubernetes and Slurm managing resource allocation and task scheduling. Networking and storage within GPU clusters are crucial for maintaining high performance, with fast data transfer and retrieval speeds required for effective training and checkpointing. Selecting the right GPU cluster involves considering factors such as hardware quality, networking, storage capabilities, costs, and provider offerings, with an emphasis on reliability, power efficiency, and sustainability.
Mar 03, 2025 3,049 words in the original blog post.