Home / Companies / Nebius / Blog / April 2025

April 2025 Summaries

13 posts from Nebius

Filter
Month: Year:
Post Summaries Back to Blog
Apache Spark is an open-source distributed computing platform designed to handle large-scale data processing efficiently, originally developed at the University of California, Berkeley. It surpasses traditional methods like MapReduce by leveraging parallel data processing, enabling it to handle petabytes of data quickly. Spark is based on the resilient distributed dataset (RDD) model, allowing parallel operations and fault tolerance through lineage-based reconstruction. The platform's architecture comprises a driver node, worker nodes, and a cluster manager, facilitating horizontal scaling by distributing tasks across nodes. Spark's ecosystem includes components like Spark SQL for data querying, MLlib for machine learning, and Structured Streaming for real-time data processing. While Spark excels in data preprocessing and integration, its use with large language models (LLMs) is enhanced by its ability to distribute tasks across nodes for scalable text processing and model training. However, challenges such as network overload during data transfers, limited machine learning library support, memory constraints, and complex setup require careful optimization. Managed services like Nebius' Managed Service for Apache Spark help mitigate these challenges by providing automatic resource scaling and simplifying infrastructure management, thus enabling efficient data processing and machine learning workflows.
Apr 30, 2025 2,326 words in the original blog post.
Transferring large volumes of data between S3 buckets can be challenging due to slow speeds and the risk of incomplete transfers. A novel approach using SkyPilot and s5cmd offers a more efficient solution, particularly for specific use cases, by enabling distributed data transfer across multiple machines. This method, however, is more of a niche hack rather than a universal production solution, especially when dealing with extremely large files or numerous small files. The integration of SkyPilot with Nebius AI Cloud allows for the orchestration of distributed compute resources, while s5cmd acts as a fast S3 client, and the use of RAM disk eliminates disk I/O bottlenecks. This method also includes post-transfer verification and supports any S3-compatible storage without vendor lock-in. While Skyplane was originally designed for object storage transfers, it is no longer actively maintained, making the SkyPilot and s5cmd combination a viable alternative. Despite requiring some initial setup, this approach significantly outperforms traditional methods like AWS CLI and AWS DataSync, particularly for cross-cloud transfers.
Apr 24, 2025 1,427 words in the original blog post.
Nebius AI Cloud has introduced Audit Logs as a new feature in preview, aimed at enhancing transparency, accountability, and security within cloud environments by automatically documenting events for compliance and policy enforcement. Unlike regular system logs, audit logs focus on identifying security-related issues, such as unauthorized access, thus creating a comprehensive audit trail that aids in tracking user activity, investigating incidents, and detecting security breaches. This built-in feature eliminates the need for third-party applications, allowing system administrators and security specialists to easily monitor transactions and collect evidence, which is crucial for resolving disputes and protecting against false claims. Access to Audit Logs is restricted to users with admin roles, and events can be filtered by time range, resource ID, and other parameters. Currently, in its preview phase, Audit Logs is offered free of charge to all tenants, and Nebius AI Cloud welcomes user feedback to refine the service for its eventual public release.
Apr 15, 2025 505 words in the original blog post.
Pre-trained AI models are neural networks trained on extensive, diverse datasets to perform specific tasks, such as image recognition and language processing, and can be customized for specialized applications through techniques like transfer learning and fine-tuning. These models offer several advantages, including reduced training times and costs, as they provide a solid foundation that can be adapted rather than built from scratch. However, they also present challenges, such as potential bias inheritance and limited customization for niche tasks. Training these models involves intricate processes, including data collection, cleaning, and selecting suitable model architectures like transformers for NLP tasks or diffusion models for text-to-image applications. The training process demands significant computational resources, including high-performance GPUs and advanced networking options, and requires sophisticated software frameworks for automation and monitoring. Despite the complexities, pre-trained models are widely used across industries, such as healthcare, finance, and technology, for diverse applications like disease diagnosis, fraud detection, and software development. While they abstract many of the challenges of training large-scale models, careful selection and mitigation strategies are necessary to address their limitations and ensure seamless integration into enterprise architectures.
Apr 14, 2025 2,438 words in the original blog post.
AI models are complex systems with billions of parameters and require extensive training on large datasets, which can be costly and time-consuming. Therefore, many teams opt to customize pre-trained models through fine-tuning, a process that involves using smaller, domain-specific datasets to enhance a model's performance for particular tasks while addressing limitations like knowledge cutoff, hallucinations, and bias. Fine-tuning is more efficient than training from scratch, as it starts with a pre-trained model and uses fewer computational resources, making it ideal for specialized applications. Various fine-tuning techniques, such as instruction fine-tuning, parameter-efficient fine-tuning, and transfer learning, allow models to specialize in specific fields like healthcare or finance. A seven-stage pipeline, including data preparation, model initialization, training setup, and monitoring, is recommended to successfully implement fine-tuning. Organizations can leverage platforms like Nebius AI Cloud to facilitate fine-tuning by providing necessary infrastructure and resources, enabling them to innovate and adapt AI capabilities to meet specific business needs effectively.
Apr 14, 2025 2,182 words in the original blog post.
Nebius AI Studio has made fine-tuning widely available, enabling the transformation of generic AI models into specialized solutions tailored to unique needs, with the ability to choose from over 30 leading open-source models and deploy them quickly. The platform now includes a broader model lineup, such as Google's Gemma 3 and DeepSeek models, and has introduced prompt presets for efficient AI workflows. Enhanced rate limits and new integrations with tools like Hugging Face, Helicone, and LlamaIndex have been implemented to improve accessibility and scalability. Additionally, Nebius offers a text-to-image generation service with models optimized for speed and quality, and is expanding its infrastructure globally with new facilities in New Jersey and Iceland to support a growing user base. Upcoming events and a roadmap for Q2 2025 highlight continued advancements in image transformations, editing capabilities, and model optimizations, emphasizing Nebius's commitment to providing a powerful and flexible AI platform.
Apr 11, 2025 1,015 words in the original blog post.
Traditional orchestration tools such as Kubernetes and Slurm pose challenges for machine learning (ML) teams due to their limitations in handling diverse ML workflows beyond training and inference. Kubernetes offers a low-level interface that can complicate development and training processes, while Slurm is optimized for training but lacks support for tasks like development environment management and inference deployment. dstack is an open-source container orchestration platform designed to address these gaps, enabling ML teams to efficiently manage GPU workloads across both cloud and on-premises environments. The platform facilitates the setup and management of development environments, tasks, services, and clusters, providing a more streamlined approach for handling various ML operations. By leveraging dstack, users can configure environments with tools like Nebius, create development environments accessible via desktop IDEs, and execute multi-node tasks with ease. Additionally, dstack's features extend to running services, managing fleets, and handling volumes, making it a comprehensive solution for ML orchestration needs.
Apr 10, 2025 483 words in the original blog post.
Nebius AI Cloud provides a secure platform for storing and processing personal data in compliance with the General Data Protection Regulation (GDPR), although it currently does not support processing of special categories or sensitive personal data as defined by GDPR, HIPAA, and PCI DSS. Personal data, under GDPR, includes any information that can identify an individual, such as names, ID numbers, and location data, while sensitive data requires additional protection and can only be processed under specific legal bases. Nebius is actively working toward HIPAA compliance and plans a third-party audit to enhance its capabilities. The platform integrates a security-by-design approach, embedding security measures such as encryption, data retention policies, continuous monitoring, and incident response protocols from the development stage, applicable across its various services. Nebius also implements a Data Processing Agreement (DPA) with customers to outline its responsibilities as a data processor, ensuring personal data is handled with the highest standards of protection.
Apr 10, 2025 764 words in the original blog post.
Nebius, in collaboration with NVIDIA, is supporting AI-native startups by offering resources and infrastructure to foster innovation in developing generative AI applications. Through a strengthened partnership with NVIDIA Inception, Nebius AI Cloud provides eligible startups with up to $150,000 in cloud credits, technical expertise, and various exclusive benefits aimed at accelerating customer success and reducing the costs associated with AI technology adoption. The Nebius AI Lift program, announced at NVIDIA GTC 2025, is designed to eliminate technical barriers and promote collaboration, enabling startups to focus on delivering transformative AI solutions across various industries. Participants in the program gain access to cutting-edge NVIDIA GPUs, discounted services, and co-marketing opportunities, along with dedicated support to maximize the efficiency and performance of their AI workloads. NVIDIA Inception itself offers a free program to startups at all stages, providing resources, preferred pricing on NVIDIA products, and visibility within the venture capital community to further accelerate their growth and innovation.
Apr 09, 2025 480 words in the original blog post.
Nebius AI Cloud has announced significant updates and expansions, including the deployment of new data center regions such as a dedicated NVIDIA Blackwell-architecture GPU facility in New Jersey and a cluster of NVIDIA H200 GPUs in Iceland. The compute cloud interface now displays all available platforms, facilitating easier project switching, and features an enhanced hardware monitoring system that enables autohealing actions and customer notifications via email, API, or Slackbot. The Managed Service for Kubernetes now integrates with this monitoring system, allowing for automatic resolution of node issues and the introduction of CoreDNS customization and pre-installed GPU and network drivers for faster node provisioning. The Container Registry is now publicly available, offering increased performance and integration ease with CI/CD tools. Slurm-based clusters have improved with node autohealing and enhanced job monitoring, while the data store and network disks offer new attach/detach options. Managed Service for PostgreSQL and MLflow have reached general availability, with enhanced features and integrations. Networking improvements include increased throughput and dedicated private address space, while identity and access management offer more granular access controls. API enhancements and new SDKs for Go and Python have been released, and billing features now include promo codes and improved documentation. Customer support has expanded to Slack, ensuring seamless assistance, with further updates and product releases expected in the next quarter.
Apr 09, 2025 1,317 words in the original blog post.
Nebius shared insights into building AI-compatible cloud infrastructures, highlighting the hardware and software challenges faced, decision-making processes, and architectural strategies implemented. NVIDIA's Adam Grzywaczewski provided guidance on selecting optimal architectures for AI projects, while Grigorii Rochev introduced Soperator, a Kubernetes operator for Slurm that facilitates workload management by integrating Slurm with Kubernetes and offering features like autoscaling. Vasily Pantyukhin shared lessons learned from scaling AI models, aiming to help others avoid common pitfalls. Boris Yangel discussed enhancing agentic systems through test-time computation by combining guided search with agent inference. Meanwhile, Nebius AI Studio's development, focusing on its Inference Service for GenAI models, was elaborated on by Nikita Vdovushkin and Roman Gaev, who offered advice on selecting appropriate inference providers and models.
Apr 08, 2025 286 words in the original blog post.
Nebius announced at NVIDIA GTC 2025 that it will be among the first AI cloud providers to offer the NVIDIA Blackwell Ultra AI factory platform, allowing access to NVIDIA GB300 NVL72-powered instances by the end of 2025. As an ecosystem partner for NVIDIA Dynamo, Nebius will help deploy GenAI efficiently in large-scale distributed environments, significantly boosting throughput on DeepSeek R1. The company has been recognized with a gold medal in the GPU Cloud ClusterMAX™ Rating System by SemiAnalysis, marking it as competitive with major cloud providers. Nebius has been actively engaging with the tech community, offering tech talks and meetups globally, and has made strides in simplifying AI/ML project management through its Standalone Applications service with JupyterLab and a general availability release of Managed MLflow. Additionally, the platform is expanding through partnerships, such as with Outerbounds and SkyPilot, and is continuously updating its documentation and tech resources to enhance user experience.
Apr 07, 2025 629 words in the original blog post.
The integration between Nebius AI Cloud and SkyPilot, an open-source framework, facilitates the execution of AI and batch workloads across various cloud platforms, optimizing GPU availability and reducing costs. This partnership allows users of SkyPilot to access Nebius AI Cloud resources directly, complementing other access methods like API, CLI, and Terraform recipes. SkyPilot offers a streamlined approach to cloud resource management through simple configuration files, making it easy to set up and run jobs on Nebius AI Cloud. The process involves configuring access, setting up the Nebius CLI, and running jobs using YAML configuration files, which specify cloud resources, accelerators, and region. The integration is particularly beneficial for tasks requiring high-performance computing, such as distributed training, due to Nebius AI Cloud's NVIDIA Quantum InfiniBand connections. Documentation and support are available for users looking to leverage SkyPilot's features and provide feedback for further enhancements.
Apr 02, 2025 1,061 words in the original blog post.