Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

YouTube

SuperBench - Improving Cloud AI Infrastructure Reliability with Proactive Validation

USENIX via YouTube

Overview

Save Big on Coursera Plus. 7,000+ courses at $160 off. Limited Time Only!
Explore a groundbreaking conference talk on improving cloud AI infrastructure reliability through proactive validation. Delve into the innovative SuperBench system, designed to mitigate hidden degradation caused by hardware redundancies in cloud AI environments. Learn about the comprehensive benchmark suite that evaluates individual hardware components and represents real AI workloads. Discover how the Validator component uses machine learning to identify defective components, while the Selector optimizes validation timing and benchmark selection. Examine the impressive results from testbed evaluations and simulations, showcasing SuperBench's ability to significantly increase mean time between incidents. Gain insights into the successful deployment of SuperBench in Azure production, validating hundreds of thousands of GPUs over a two-year period. Understand the critical importance of addressing "gray failures" in cloud AI infrastructure and how SuperBench contributes to enhanced overall reliability for cloud service providers.

Syllabus

USENIX ATC '24 - SuperBench: Improving Cloud AI Infrastructure Reliability with Proactive Validation

Taught by

USENIX

Reviews

Start your review of SuperBench - Improving Cloud AI Infrastructure Reliability with Proactive Validation

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.