October 2026 Summaries
1 posts from AI21 Labs
Filter
Month:
Year:
Post Summaries
Back to Blog
AI21 describes replacing manual GPU-allocation negotiations in a shared Google Kubernetes Engine cluster of roughly 10,000 GPUs with Kueue, an open-source Kubernetes-native system for queuing, prioritizing, admitting, and preempting workloads. The previous process relied on internal messaging channels and caused unfair access, idle capacity, fragmented resources, delayed critical jobs, and significant engineering overhead, while native Kubernetes mechanisms lacked workload-level, gang-scheduling, and flexible fair-sharing capabilities. After evaluating alternatives, the team adopted Kueue to manage guaranteed, opportunistic, and on-demand capacity for debug pods, training jobs, and inference deployments, using workload priority, preemptibility, and workload type to route jobs automatically. Production feedback contributed to Kueue features including Admission Fair Sharing, which prioritizes teams with lower historical usage, and Topology Aware Scheduling, which prevents jobs from being admitted when GPUs are available in aggregate but cannot fit required node configurations. The redesigned system also added priority-aware preemption and best-effort lanes for single- and multi-node workloads. Reported results include eliminating manual interventions and partially allocated jobs, reducing GPU fragmentation from 15% to 8%, cutting critical “hero job” starvation from 72 to 12 hours, and reducing time-to-start for high-priority workloads by 83%, while maintaining high fleet utilization and reducing time spent on resource coordination.
Oct 04, 2026
2,825 words in the original blog post.