An SLO-Driven Approach to Enhance Kubernetes Cluster Reliability

Overview

Explore an SLO-driven approach to enhance Kubernetes cluster reliability in this conference talk from KubeCon + CloudNativeCon Europe 2021. Delve into the challenges of defining reliability for large-scale Kubernetes clusters and learn how Service Level Objectives (SLOs) can be effectively implemented. Discover the philosophy behind SLO-driven reliability engineering and gain insights from Ant Financial's experience with one of the world's largest Kubernetes clusters. Examine concrete cases and lessons learned in building SLO frameworks, covering aspects such as monitoring, alerting, and tracing. Understand the complexities of defining SLOs for Kubernetes services compared to classic web services, and explore topics including fleet management, EZE SLO design, fine-grained and component SLOs, alerting philosophy, and SLO management.

Syllabus

Thank You to Our Session Recording Sponsor
Outline
Motivation
Fleet Management
General Approach
SLO Approach
SLO Recap
What SRE cares on K8S?
EZE SLO Design
Fine-grained SLO
Component SLO
Overall SLO Graph
Why RatioRate is bad?
Alerting Philosophy
SLO Management

Taught by

CNCF [Cloud Native Computing Foundation]

Reviews

Start your review of An SLO-Driven Approach to Enhance Kubernetes Cluster Reliability

Taught by

Methods to Achieve High SLOs on a Large Scale Kubernetes Cluster

Never Stop Learning.