Home / Companies / Nebius / Blog / June 2026

June 2026 Summaries

10 posts from Nebius

Filter
Month: Year:
Post Summaries Back to Blog
Custom Speculator Training, launched by Nebius Token Factory, allows teams to train workload-specific draft models using their own data, offering a significant improvement over generic models in speculative decoding. This approach enhances throughput and latency by aligning the draft model with the specific traffic patterns of a product, thereby increasing acceptance rates and reducing latency variability under load. The platform provides tools for data preparation, training, and deployment, with preset hyperparameters for ease of use and advanced options for those needing fine-tuning. This innovation is particularly beneficial for teams handling consistent, high-volume workloads, such as AI product teams, enterprise accounts already using speculative decoding, and ML engineers managing open-source models in production. It promises better unit economics and serving behavior control, making it an attractive option for those with predictable workload patterns. The launch is built on ongoing research and is integrated into an existing infrastructure that has been maturing over several months, with a focus on improving acceptance rates and shaping model behavior based on real-world data.
Jun 25, 2026 1,249 words in the original blog post.
Aether 3.6 marks the latest release of a comprehensive AI cloud platform, enhancing security, storage, and user experience based on customer feedback. Key features include Nebius Echo, an AI agent that facilitates natural-language control over cloud environments, and a Managed Service for SkyPilot, simplifying workload management for ML engineers. The update also introduces a unified Notification Center, new security and governance capabilities such as Key Management Service and Workload Identity Federation, as well as improved cost management tools. Storage innovations include the Intelligent storage class for cost-effective data management and enhanced disk functionalities. Other updates feature integration with Datadog Log Management and a Nebius Builder Program offering credits and certifications to promote user engagement and skill development.
Jun 24, 2026 1,699 words in the original blog post.
Nebius Echo is an AI agent integrated into the Nebius cloud console, designed to simplify AI infrastructure management by allowing users to interact with their cloud environment using natural language. It eliminates the need for technical expertise by enabling users to ask questions, retrieve information, and execute resource creation requests without complex commands or separate documentation searches. Building on the existing Nebius MCP Server, Echo facilitates seamless integration with external agent tools, providing a consistent experience across different interfaces. Future developments aim to enhance Echo's capabilities with infrastructure investigation and a native Infrastructure-as-Code (IaC) feature, thereby enabling more complex, context-rich deployments. This evolution towards an agentic-ready AI cloud seeks to democratize access to operational expertise, reducing the time between conceptualizing an idea and deploying it, and ultimately transforming how AI products are built and maintained.
Jun 24, 2026 1,148 words in the original blog post.
Nebius showcased its high-performance AI cloud infrastructure by submitting results to the MLPerf® Training 6.0 benchmark, demonstrating leading performance in training advanced models on NVIDIA Blackwell Ultra systems. Nebius achieved top single-node results for Llama-3.1-8B and GPT-OSS 20B pre-training using the NVIDIA HGX B300 platform, while also delivering competitive multi-node performance across various configurations. The GB300 NVL72 platform excelled at rack scale, performing within a small margin of the fastest submissions. By optimizing hardware and software layers, Nebius ensures reproducible efficiency and high GPU utilization for demanding GenAI workloads. The results highlight Nebius's collaboration with NVIDIA and their commitment to providing robust AI training solutions that drive innovation and breakthroughs in artificial intelligence, benefiting customers in research and development across language, vision, and multimodal AI projects.
Jun 16, 2026 1,091 words in the original blog post.
Nebius AI Cloud has introduced an integration with Datadog, allowing logs from its core services to stream directly into Datadog Log Management, facilitating a more cohesive and efficient investigation process for production AI incidents. This integration covers various services, including Nebius Cloud Compute, Kubernetes, Applications, PostgreSQL, and MLflow, enabling teams to investigate AI workloads without switching between multiple tools and consoles. By consolidating logs within the existing Datadog environment, teams can correlate events and logs across their infrastructure, improving incident response times and maintaining operational workflows. While users can continue using Nebius AI Cloud Observability independently if preferred, the Datadog integration is positioned as an additional tool to enhance visibility and streamline incident investigations.
Jun 15, 2026 573 words in the original blog post.
AI agents often encounter production failures not due to the model itself but because of the supporting system, which may include issues like incorrect data retrieval, lack of current information for queries, unsustainable costs, lack of traceability, and untested behavior. The text discusses a case study involving the development of Sentinel, a regulatory compliance audit agent, to illustrate the process of making an AI agent production-ready by focusing on reliability, observability, and economic feasibility. The study follows Sentinel through four configurations—Prototype, Grounded, Optimized, and Production—each addressing different bottlenecks such as data freshness, inference economics, runtime, and decision-making. The production configuration achieved the best results by ensuring complete traceability and conducting adversarial testing before launch, ultimately focusing on delivering actionable decisions rather than just findings. The narrative emphasizes the importance of infrastructure and continuous improvement for successful large-scale AI deployment, highlighting the Nebius Agents Blueprint as a tool for identifying and optimizing system bottlenecks.
Jun 10, 2026 2,156 words in the original blog post.
In 2026, the focus in AI has shifted from whether agents can function to how they can be run reliably and efficiently in production. The Nebius Agents Blueprint addresses this challenge by providing an open reference architecture for building, operating, and improving AI agents. This blueprint emphasizes that system improvements in areas like retrieval quality, orchestration strategy, grounding, and evaluation are often more impactful than model upgrades. The blueprint comprises six components, including the Nebius Token Factory for inference, LangChain Deep Agents for orchestration, and LangSmith for observability, each independently deployable or integrable into existing systems. A case study demonstrated the blueprint's efficacy by building a regulatory compliance audit agent, showing significant cost reductions and precision improvements by refining system components rather than merely upgrading models. The Nebius Agents Blueprint aims to make AI systems more measurable, economical, and reliable, supporting the industry's goal of creating sustainable, production-ready AI infrastructure.
Jun 10, 2026 996 words in the original blog post.
Nebius and NVIDIA have partnered to introduce two AI blueprints for retail on Nebius infrastructure: the NVIDIA Agentic Commerce Blueprint and the NVIDIA Retail Catalog Enrichment Blueprint, both of which are open-source reference architectures built on NVIDIA NIMs for easy deployment. These blueprints address the growing interest in agentic AI and catalog enrichment among retailers, offering production-ready solutions that can be customized and deployed via Nebius AI Cloud. The collaboration aims to overcome execution challenges in AI implementation, highlighted by the fact that only a small percentage of retailers have fully deployed AI despite significant interest and investment. The Agentic Commerce blueprint utilizes AI protocols to facilitate product discovery and transactions, while the Catalog Enrichment blueprint automates the transformation of raw product images into enriched catalog entries. Nebius provides the necessary AI infrastructure, including GPU compute and managed inference through Token Factory, to support these blueprints, allowing retailers to focus on application development without the need for extensive infrastructure management. This partnership underscores the importance of open-source models and tools in bridging the gap between AI ambition and practical deployment in the retail sector.
Jun 05, 2026 1,172 words in the original blog post.
Firms are transitioning from task-specific models to transaction foundation models, which utilize transformer architectures trained on proprietary transaction data to produce reusable embeddings for various financial applications like fraud detection and credit scoring. NVIDIA's blog discusses this shift, highlighting contributions from institutions such as Revolut, which developed its PRAGMA model with NVIDIA's support. This model, trained on Nebius AI Cloud, demonstrates how infrastructure and architecture combine to enhance pre-training efficiency and fraud detection precision. The Build Your Own Transaction Foundation Model developer example provides a framework for creating transformer embeddings on tabular data using NVIDIA's technology, with options for easy deployment on Nebius AI Cloud. Key challenges include scaling data preparation and ensuring robust infrastructure for large-scale training and real-time inference under stringent service level agreements. Nebius AI Cloud supports the full model lifecycle, offering GPU-accelerated data processing, sustained training capabilities, and managed inference endpoints, all while adhering to data residency and compliance requirements. The PRAGMA model illustrates how advanced infrastructure can support the rapid adaptation of pre-trained models to new tasks without the need for full retraining, showcasing NVIDIA's AI platform's role in enabling production-scale financial intelligence.
Jun 02, 2026 1,131 words in the original blog post.
Nebius Physical AI Workbench addresses the bottleneck in physical AI development by providing a curated platform that integrates and pre-validates essential tools for simulation, synthetic data generation, training, evaluation, and deployment. This workbench streamlines the integration process by offering a unified data layer and utilizing NVIDIA technologies such as Cosmos 3, Isaac Sim, and Isaac GR00T, which are validated for seamless operation. The platform allows teams to focus on creating intelligent systems rather than handling infrastructure, using a headless API and agent-driven workflows. By automating the traditionally labor-intensive tasks of configuring and connecting tools, the Workbench accelerates the lifecycle of Physical AI projects, transforming a traditionally cumbersome process into an efficient, repeatable workflow. Open-source and vendor-neutral, it supports collaboration and customization, enabling developers to rapidly prototype and deploy AI solutions while maintaining control over their data and processes.
Jun 01, 2026 1,375 words in the original blog post.