Introducing Managed Soperator: Your quick access to Slurm training
Blog post from Nebius
Managed Soperator, a fully managed Slurm-on-Kubernetes solution, is now available for self-service, allowing users to quickly set up a Slurm training cluster with NVIDIA GPUs and pre-installed libraries, facilitating immediate machine learning training. Developed as a managed service on the Nebius AI Cloud, it aims to simplify the user experience by automating infrastructure provisioning and configuration, which traditionally required extensive manual setup. Soperator, the core technology behind this solution, is an open-source Kubernetes operator for Slurm, originally released last autumn, and enables rapid deployment of large GPU clusters while ensuring fault tolerance and scalability. With three options available—Managed Soperator, Professional Soperator, and open-source Soperator—the platform caters to different needs, from self-service to customized large-scale installations, each supporting various AI/ML drivers and libraries. This solution is designed to empower AI developers by minimizing operational overhead and focusing on innovative research, with ongoing developments to enhance its features further.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.