Production Multi-node Jobs with Gang Scheduling, K8s, GPUs and RDMA
CNCF [Cloud Native Computing Foundation] via YouTube
Overview
Syllabus
Intro
Deep Learning Applications
AL/DL: Models, Frameworks, Hardware
Trends: Big Data, Larger Models
Sample Multi-GPU Node: DGX-1
Distributed Training Applications Multi-GPU, Multi-node
K8s Challenges & Outline
Kes Orchestration Flow
Sample PyTorch Job Launch
Array Jobs and MPI Operator
SRIOV CNI for K8s Multi-Rail
Gang Scheduling Multi-Node Pods
PodGroup Queue and Manager
Demo
Sample Job Real-Time Telemetry
Sample BERT K8s Scaling
Shared K8s Cluster for Multi-node
Scheduler Dashboard
Summary and Future Work
Taught by
CNCF [Cloud Native Computing Foundation]