An exploration of inefficiencies in deep reinforcement learning
File(s)
Author(s)
Richemond, Pierre
Type
Thesis
Abstract
Deep reinforcement learning methods, for all of their impressive empirical success, are known to require tremendous amounts of data and compute. This begs the general question of their inefficiencies. Are there better, more parsimonious ways of implementing these algorithms ? The purpose of this thesis is to use mathematical methods to take a second look at some inefficiencies in reinforcement learning and hopefully somehow improve them.
In the first part of this work, we investigate how to claim better efficiency or robustness in reinforcement learning, often by rethinking core mathematical assumptions. First we re-metrize the policy gradients method to use a trust region metric based on the Wasserstein distance, and prove that under certain general circumstances, performing reinforcement learning in such manner is strictly equivalent to solving the heat equation. Then, we question whether overparameterization, a well-known feature of neural networks design that helps them learn in supervised settings, is actually required for deep reinforcement learning, and thanks to insights from tensor factorization, answer this question empirically by proposing weight-efficient encoder neural network architectures for our policies.
In the second part of this thesis, we examine the question of learning good representations for deep reinforcement learning. Standard exploration methods suffer from the curse of dimensionality, therefore dimensionality reduction methods preceding policy learning steps can be an important means of keeping them efficient. We begin with adapting methods from manifold learning in order to exploit similarities between states and possibly their associated rewards. Then, we introduce Bootstrap Your Own Latent (abbreviated as BYOL), an extremely simple, generic, and end-to-end differentiable self-supervised learning algorithm. The final chapter of this thesis is devoted to further analysis and understanding of BYOL.
In the first part of this work, we investigate how to claim better efficiency or robustness in reinforcement learning, often by rethinking core mathematical assumptions. First we re-metrize the policy gradients method to use a trust region metric based on the Wasserstein distance, and prove that under certain general circumstances, performing reinforcement learning in such manner is strictly equivalent to solving the heat equation. Then, we question whether overparameterization, a well-known feature of neural networks design that helps them learn in supervised settings, is actually required for deep reinforcement learning, and thanks to insights from tensor factorization, answer this question empirically by proposing weight-efficient encoder neural network architectures for our policies.
In the second part of this thesis, we examine the question of learning good representations for deep reinforcement learning. Standard exploration methods suffer from the curse of dimensionality, therefore dimensionality reduction methods preceding policy learning steps can be an important means of keeping them efficient. We begin with adapting methods from manifold learning in order to exploit similarities between states and possibly their associated rewards. Then, we introduce Bootstrap Your Own Latent (abbreviated as BYOL), an extremely simple, generic, and end-to-end differentiable self-supervised learning algorithm. The final chapter of this thesis is devoted to further analysis and understanding of BYOL.
Version
Open Access
Date Issued
2022-08-19
Date Awarded
01/01/2024
Advisor
Guo, Yike
de Montjoye, Yves-Alexandre
Publisher Department
Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
