December 2024 Summaries
10 posts from Nebius
Filter
Month:
Year:
Post Summaries
Back to Blog
A faulty release in the VM recovery sequence in the eu-north1 region led to 282 virtual machines (VMs) restarting, with 164 experiencing extended downtime that required manual intervention, resulting in customer workload interruptions. The incident was triggered by an issue in the Compute API service, which mistakenly assumed certain VMs were non-operational due to flawed event processing assumptions and lack of optimization for recovery operations, causing some VMs to become stuck. The response involved halting automatic recovery operations, categorizing affected VMs, and applying thoroughly tested mitigation procedures, which included stopping and restarting VMs in controlled batches. The incident underscored the need for improved automation in recovery procedures to enhance response times, and a post-incident action plan was developed focusing on pre-deployment validation, operational improvements, and enhanced communication protocols to prevent future occurrences.
Dec 27, 2024
652 words in the original blog post.
The research blog post discusses the development and improvement of software engineering agents using search-based methods and critic-guided action generators. The research highlights a significant improvement in agent performance over previous models, achieving state-of-the-art results on the SWE-bench Verified benchmark using open-weight models. The study emphasizes the lack of comprehensive datasets for training such agents, prompting the creation of SWE-bench Extra, a richer dataset derived from GitHub repositories. This dataset focuses on issue-solving tasks, crucial for software engineering, and includes extensive data collection, filtering, and validation processes to ensure high-quality training material. The research also details the methodologies used to collect and validate data, including execution-based validation and the application of filtering criteria to maintain dataset quality. Additionally, it underscores the challenges faced in gathering data, such as ensuring stable environments for testing, and highlights the use of TractoAI for scalable data processing. The study concludes with reflections on the improvements achieved and outlines future directions for expanding the scope of the research, including broadening language support and enhancing agent adaptability for environment setups.
Dec 20, 2024
3,677 words in the original blog post.
In October, a series of updates were introduced across various cloud services and platforms, including enhanced GPU support, new monitoring dashboards, and improved cluster management tools. Notably, the Paris region now offers NVIDIA H200 GPUs, and the Compute Cloud platform supports both AMD and NVIDIA configurations. Cluster management saw the addition of SSH access on worker nodes, enroot support without root privileges, and the introduction of Slurm partitions and a REST API for efficient resource scheduling. The Managed Service for Kubernetes integrated load balancer support, node autoscaling, and high availability by default, while the new Container Registry service in preview mode facilitates seamless image management. Data Store updates included filesystem resizing and enhanced object storage metrics, and the Managed Service for PostgreSQL saw improvements in performance metrics and cluster configuration options. MLOps services now offer enhanced MLflow support, with new logging and performance metrics, while the cloud platform features expanded network management options and identity access enhancements. Nebius AI Studio, rebranded as Nebius Token Factory, now features an expanded portfolio of large language models with increased token rate limits, further emphasizing its capability to support AI workloads at scale.
Dec 18, 2024
667 words in the original blog post.
Prompt engineering is a crucial process in the development of AI applications, especially when working with commercial and open-source large language models (LLMs) like GPT and LLaMA. It involves crafting and optimizing prompts, or natural language inputs, to ensure that the LLMs produce relevant and accurate outputs tailored to specific tasks and user expectations. This process not only helps in customizing LLMs for particular use cases but also in preventing inappropriate or unauthorized outputs. Prompt engineering ranges from basic strategies, such as zero-shot and few-shot prompting, to advanced techniques like Retrieval Augmented Generation (RAG) and Program-Aided Language Models (PALM), which integrate LLMs with other technologies for more complex tasks. By understanding and applying these strategies, developers can enhance the performance and reliability of AI systems, ensuring they effectively map user inputs to the LLM's domain and generate optimal outputs.
Dec 17, 2024
2,179 words in the original blog post.
Nebius AI Studio has introduced a range of new AI models and features designed to enhance scalability and performance without rate limits, making them suitable for diverse applications from prototype to production. The platform now includes advanced vision-language models for tasks like image captioning and product recognition, as well as expanded language models such as Meta Llama-3.3-70B-Instruct for multilingual scenarios and dolphin-2.9.2-mixtral-8x22b for conversational AI and coding tasks. Additionally, new embedding models have been added to strengthen retrieval-augmented generation pipelines, ideal for building knowledge bases and semantic search engines. Nebius AI Studio offers a usage-based pricing model for deploying pre-trained LoRA models, eliminating fixed costs and allowing seamless scaling. The infrastructure supports massive batch processing and flexible deployment options, ensuring consistent performance across various use cases, including computer vision, advanced language processing, and RAG implementations. Users can start by logging into the AI Studio, experimenting with models, and integrating them via the comprehensive API, with transparent token-based pricing facilitating cost-effective scaling.
Dec 17, 2024
842 words in the original blog post.
Deploying AI models to production requires continuous performance evaluation, which is traditionally done through human feedback loops but is limited in scalability. To achieve enterprise-grade quality control, organizations should implement AIOps and establish CI/CD/CT pipelines for continuous integration, testing, and deployment. Evaluating and quantifying performance improvements is essential, with various metrics such as perplexity, BLEU, ROUGE, METEOR, and BERTScore aiding in model output assessment. These metrics range from statistical scorers to model-based scorers, each serving different evaluation needs, like translation accuracy or semantic understanding. Responsible AI development emphasizes accountability, transparency, and accuracy, with metrics like SelfCheck GPT, QAG Score, and fairness scores ensuring ethical and reliable AI operation. Additionally, user engagement, speed, cost, and responsible AI metrics are crucial for optimizing model efficiency, justifying costs, enhancing user satisfaction, and maintaining ethical standards. These comprehensive metrics collectively enhance the reliability, effectiveness, and ethical compliance of AI models in production environments.
Dec 13, 2024
2,031 words in the original blog post.
AI model hosting involves deploying a trained AI model to make it accessible to users and applications via an API, with various hosting options influencing performance, cost, and security. The hosting environment is typically complex, involving multiple technology layers such as compute, storage, and orchestration, each managed differently depending on the chosen hosting method. Options range from self-managed on-premises setups, which offer full control but are costly and inflexible, to cloud-based solutions like serverless and managed cloud services, which provide scalability and lower maintenance but may require specific expertise or come with certain limitations. AI Platform as a Service (PaaS) offers an all-in-one managed solution, allowing teams to focus on model development without worrying about infrastructure management. When selecting a hosting option, considerations include inference type (batch or real-time), cost, security, and customizability to ensure the solution aligns with project needs and organizational capabilities.
Dec 11, 2024
1,904 words in the original blog post.
Nebius AI is deploying over 22,000 NVIDIA Blackwell GPUs on its AI-native cloud, marking a significant advancement in hardware capability. This deployment includes the NVIDIA GB200 Grace Blackwell Superchip, which features a reimagined mainframe optimized for large models, and the NVIDIA HGX B200 system, demanding adaptation from users familiar with previous systems. Nebius emphasizes the importance of in-house hardware expertise for maximizing GPU investment and has rewritten its cloud infrastructure to integrate seamlessly with Arm architecture, offering faster storage and efficient multi-node operations. The new systems boast a 25-fold reduction in cost and energy consumption compared to previous models, facilitated by advanced liquid-cooling technology being installed in data centers in Finland and Kansas City. By offering these powerful, Blackwell-powered clusters across Europe and the United States, Nebius addresses latency issues and provides fully integrated solutions with managed Kubernetes and Slurm-based workload orchestration, enhancing the efficiency and predictability of resource-intensive processes for machine learning teams.
Dec 04, 2024
575 words in the original blog post.
Nebius is initiating pre-orders for its NVIDIA Blackwell GPU-powered clusters, specifically the NVIDIA GB200 NVL72 and NVIDIA HGX B200, set to launch in data centers in the United States and Finland by early 2025. These clusters represent a significant advance in generative AI technology, with over 22,000 GPUs to be deployed on the Nebius AI-native cloud. The company is expanding its presence in the US, with a new data center in Kansas City and offices in San Francisco, Texas, and New York. Nebius is also advancing AI-driven software engineering through R&D, focusing on improving large language models (LLMs) using search and learning. Their AI platform, now called Nebius Token Factory, offers competitive pricing for popular models like Llama and Mistral. Additionally, Nebius has detailed processes for building AI applications, showcased their in-house hardware design at the OCP Global Summit, and updated documentation to improve user experience with features like batch inference and billing methods.
Dec 04, 2024
558 words in the original blog post.
In this hands-on tutorial, readers are guided through the process of building a full-stack AI code assistant using Next.js and Nebius AI Studio, which involves creating a code generation system that works across various programming languages and implementing features such as intelligent code reviews and a professional-grade code editor interface. The tutorial emphasizes the use of open-source models and modern web technologies, making the application cost-effective and extensible. The technical stack includes Next.js for the frontend and Nebius AI Studio for AI capabilities, with the latter offering an interactive playground and multiple model options for testing and refining prompts. Key components of the application include the GenerateCode, ReviewCode, and Result components, which allow users to generate code snippets, review code with actionable feedback, and display results. The tutorial also covers setting up Nebius AI, integrating it with the application, and using the OpenAI JavaScript SDK to communicate with AI models, ultimately enabling users to generate and review code snippets efficiently.
Dec 03, 2024
3,715 words in the original blog post.