Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

YouTube

Beware of Fragmentation - Scheduling GPU-Sharing Workloads with Fragmentation Gradient Descent

USENIX via YouTube

Overview

Save Big on Coursera Plus. 7,000+ courses at $160 off. Limited Time Only!
Explore a 22-minute conference talk from USENIX ATC '23 that addresses the critical issue of GPU underutilization in large tech companies. Dive into the challenges of GPU sharing techniques and the resulting fragmentation problems in large clusters. Learn about a novel approach called Fragmentation Gradient Descent (FGD), which quantifies GPU fragmentation and schedules workloads to minimize its growth. Discover how this innovative method, implemented as a new scheduler in Kubernetes, significantly reduces unallocated GPUs and improves overall utilization. Gain insights into the performance evaluation of FGD using production traces on an emulated cluster of over 6,200 GPUs, and understand its potential to revolutionize GPU resource management in large-scale machine learning environments.

Syllabus

USENIX ATC '23 - Beware of Fragmentation: Scheduling GPU-Sharing Workloads with Fragmentation...

Taught by

USENIX

Reviews

Start your review of Beware of Fragmentation - Scheduling GPU-Sharing Workloads with Fragmentation Gradient Descent

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.