Overview
Syllabus
Intro
Reinforcement karning: Learning to make decisions
Online vs. Offline (Batch) RL: A Basic View
Outline
Markov Decision Process (MDP)
MDP Example: Deterministic Shortest Path
More General Case: Bellman Equation
Bellman Operator
When Bellman Meets Gauss: Approximate DP
Divergence Example of Tsitsiklis & Van Roy (96)
Does It Matter in Practice?
A Long-standing Open Problem
Linear Programming Reformulation
Why Solving for Fixed Point Directly is Hard?
Addressing Difficulty #2: Legendre-Fenchel Transformation
Reformulation of Bellman Equation
Primal-dual Problems are Hard to Solve
A New Loss for Solving Bellman Equation
Eigenfunction Interpretation
Puddle World with Neural Networks
Conclusions
Taught by
Simons Institute