Technology · Ebook
KAI Scheduler: Fair GPU Scheduling for AI on Kubernetes
by Shriira Press
KAI Scheduler is a Kubernetes-native scheduler, open-sourced by NVIDIA from the Run:ai platform, that solves a stubborn problem: GPUs are expensive and scarce, yet the default scheduler was built for stateless web pods, not distributed AI jobs that must launch every worker at once. This book walks through how KAI fixes that — its operator-driven architecture and once-per-second scheduling cycle, hierarchical queues with quota and fair share, gang scheduling via PodGroups, GPU sharing and fractional GPUs, priorities, preemption and reclaim, bin-packing, consolidation, and topology-aware placement. A closing chapter covers installation, queue setup, framework integrations, and running KAI well at scale, with real components and CRDs throughout.
Contents
- 1Preface
- 2Chapter 1 — Why AI Workloads Need a Different Scheduler
- 3Chapter 2 — Architecture and the Scheduling Cycle
- 4Chapter 3 — Queues, Quota, and Hierarchical Fair Share
- 5Chapter 4 — Gang Scheduling and PodGroups
- 6Chapter 5 — GPU Sharing and Fractional GPUs
- 7Chapter 6 — Priorities, Preemption, and Reclaim
- 8Chapter 7 — Bin-Packing, Consolidation, and Topology-Aware Placement
- 9Chapter 8 — KAI Scheduler in Practice
