Overview
Syllabus
Intro
Model-Based Reinforcement Learning
Episodic Reinforcement Learning
Upper Confidence Model-Based RL (UCRL)
The class of deterministic continuous systems . Consider a deterministic system
A Simple Metric-Based RL Algorithm
Doubling Dimension d
Feature space embedding of transition model
The MatrixRL Algorithm
From Feature to Kernel Embedding of Transition Model
A motivating example: MuZero
Assumption of Value-Targeted Regression
Value-Targeted Regression (VTR) for Confidence Set Construction
Full Algorithm of UCRL-VTR
Regret analysis of UCRL-VTR
A Special Case
Taught by
Simons Institute