Optimality and Approximation with Policy Gradient Methods in Markov Decision Processes

Overview

Save Big on Coursera Plus. 7,000+ courses at $160 off. Limited Time Only!

Grab it

Explore the intricacies of policy gradient methods in Markov Decision Processes through this 55-minute lecture by Alekh Agarwal from Microsoft Research Redmond. Delve into optimality and approximation concepts as part of the "Emerging Challenges in Deep Learning" series at the Simons Institute. Examine MDP preliminaries, policy parameterizations, and the policy gradient algorithm, with a focus on softmax parameterization and entropy regularization. Analyze the convergence of entropy-regularized PGA, natural solutions, and proof ideas. Investigate restricted parameterizations, natural policy gradient updates, policy assumptions, and extensions to finite samples. Gain valuable insights into this crucial area of deep learning and reinforcement learning research.

Syllabus

Intro
Questions of interest
Main challenges
MDP Preliminaries
Policy parameterizations
Policy gradient algorithm
Policy gradient example: Softmax parameterization
Entropy regularization
Convergence of Entropy regularized PG
A natural solution
Proof ideas
Restricted parameterizations
A closer look at Natural Policy Gradient • NPG performs the update
Assumptions on policies
Extension to finite samples
Looking ahead

Taught by

Simons Institute

Reviews

Start your review of Optimality and Approximation with Policy Gradient Methods in Markov Decision Processes

Taught by

Average Reward Markov Decision Process - Policy Gradient Algorithms and Regret Analysis

On the Hardness of Reinforcement Learning With Value-Function Approximation

Online Learning in Markov Decision Processes - Part 2

Global Guarantees for Policy Gradient Methods

Learning Decentralized Policies in Multiagent Systems - How to Learn Efficiently

10 Best Deep Learning Courses for 2024

Never Stop Learning.