Home / Companies / Datadog / Blog / Post Details
Content Deep Dive

Monitor TAS and gang scheduling for AI training in Kubernetes

Blog post from Datadog

Post Details
Company
Date Published
Author
David Lentz, Kathy Lin
Word Count
4,465
Company Posts That Month
27
Language
English
Hacker News Points
-
Post removed?
No
Summary

Distributed AI training workloads challenge Kubernetes’s default pod-by-pod scheduling model because workers require both low-latency topology-aware placement and simultaneous startup to avoid communication bottlenecks and idle GPU capacity. The post describes how Kueue provides admission control and topology-aware scheduling by assigning jobs to labeled hardware domains such as hosts, racks, or blocks, while the Coscheduling plugin implements gang scheduling by reserving resources until a job’s required set of pods can bind together. It explains how cluster topology affects GPU communication through NVLink, PCIe, and network fabrics, with cross-rack placement often increasing latency and reducing training throughput, especially for large distributed jobs. Effective validation and troubleshooting require correlating Kueue admission and eviction metrics, PodGroup gang states, workload topology assignments, GPU and network bandwidth, and framework-level throughput metrics. Datadog’s Kubernetes and GPU monitoring tools are presented as a way to centralize these signals, detect queueing, placement, hardware, or assembly issues, and retain diagnostic data for frameworks such as Ray, PyTorch, Kubeflow, and Horovod. Kubernetes is also adding native gang scheduling, workload-aware preemption, and topology-aware scheduling capabilities in versions 1.35 through 1.37, but these remain early-stage or disabled by default, leaving Kueue and Coscheduling as the described production approach.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Kubernetes 40 956 75 30 -73%
Observability 6 472 102 54 -85%
AI Model Fine-tuning 1 139 28 14 -75%
LLM 1 747 162 79 -85%
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.