Home / Companies / Nebius / Blog / Post Details
Content Deep Dive

Introducing Managed Soperator: Your quick access to Slurm training

Blog post from Nebius

Post Details
Company
Date Published
Author
Andrey Kuyukov, Roman Luchkov
Word Count
739
Company Posts That Month
9
Language
English
Hacker News Points
-
Post removed?
No
Summary

Managed Soperator, a fully managed Slurm-on-Kubernetes solution, is now available for self-service, allowing users to quickly set up a Slurm training cluster with NVIDIA GPUs and pre-installed libraries, facilitating immediate machine learning training. Developed as a managed service on the Nebius AI Cloud, it aims to simplify the user experience by automating infrastructure provisioning and configuration, which traditionally required extensive manual setup. Soperator, the core technology behind this solution, is an open-source Kubernetes operator for Slurm, originally released last autumn, and enables rapid deployment of large GPU clusters while ensuring fault tolerance and scalability. With three options available—Managed Soperator, Professional Soperator, and open-source Soperator—the platform caters to different needs, from self-service to customized large-scale installations, each supporting various AI/ML drivers and libraries. This solution is designed to empower AI developers by minimizing operational overhead and focusing on innovative research, with ongoing developments to enhance its features further.

Trends Found in this Post

No tracked trend matches for this post yet.

Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.