Improving AI cluster observability: New metrics, Grafana dashboards and advanced logging
Blog post from Nebius
Recent improvements to the Nebius AI Cloud have enhanced its observability features, providing users with advanced monitoring metrics through both the web console and API, and the inclusion of out-of-the-box Grafana dashboards. These updates, which allow the upload of custom metrics and logs, are designed to improve visibility and performance management of AI clusters, which are inherently complex and prone to higher failure rates due to their scale. The enhancements aim to increase efficiency, reliability, and quick troubleshooting by offering detailed monitoring of every component within an AI cluster, from compute to networking and storage. The AI Cloud now includes two primary services: Monitoring, which visualizes performance metrics, and Logging, storing information about system events, both of which are crucial for making AI model development more transparent and predictable. The pre-configured Grafana dashboards facilitate effortless visualization of performance data and service logs, and the new logging feature supports the storage and visualization of custom logs. These changes are part of an ongoing effort to make the cloud more transparent and cost-effective, reducing idle compute capacity and optimizing resource utilization, with further enhancements planned to continue improving the platform's observability capabilities.
No tracked trend matches for this post yet.
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.