Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

YouTube

From Supercomputing to Serving: A Case Study Delivering Cloud Native Foundation Models

CNCF [Cloud Native Computing Foundation] via YouTube

Overview

Save Big on Coursera Plus. 7,000+ courses at $160 off. Limited Time Only!
Explore a conference talk that delves into the challenges and solutions of scaling machine learning infrastructure across multiple cloud providers. Learn how a small team successfully expanded from a single cloud to managing multiple clusters across four cloud providers within six months. Discover unique multi-cloud challenges in supercomputing infrastructure, including cross-cloud networking, capacity and quota management, batch workloads, FinOps, and observability. Gain insights into using Kueue for managing fixed capacity across clouds and understand the current limitations of Kubernetes for HPC workloads. Master the essential requirements for infrastructure teams supporting cloud native foundation models, with special emphasis on handling hardware-coupled software, fixed capacity constraints, and GPU flexibility for batch training workloads and inference.

Syllabus

From Supercomputing to Serving: A Case Study Delivering Cloud Native Foundation Mo... Autumn Moulder

Taught by

CNCF [Cloud Native Computing Foundation]

Reviews

Start your review of From Supercomputing to Serving: A Case Study Delivering Cloud Native Foundation Models

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.