Deep reinforcement learning in the face of uncertainty
Author(s)
Li, Luchen
Type
Thesis
Abstract
Reinforcement learning is a subtopic of machine learning that models an agent interacting with an environment as a Markov decision process. To learn an optimal and robust policy, the agent needs to have a clear grasp of environmental information and the consequences of its own behaviour. Accounting for the ubiquitous uncertainty is critical to this end especially for real-world practices. In this thesis, I address three sources of uncertainties: the epistemic uncertainty due to insufficient or skewed training data, uncertainty in what can be observed about the environment, i.e. as in partially observable MDPs, and aleatoric uncertainty in the return that stems from intrinsic stochasticity in the MDP dynamics and rewards. To tackle these problems, the following algorithmic building blocks are proposed and tested in this thesis: i) ASTC, a best-first tree-structured search method that targets epistemic uncertainty by directly minimising value estimation uncertainty at the root node. ASTC is guaranteed with monotonic improvement but is applicable only to POMDPs with discrete state spaces; ii) AEHS, another tree exploration mechanism to compensate for the lack of data that allows for both POMDPs with continuous spaces and generic MDPs, albeit at the cost of less efficient search heuristics of expanding the most probable nodes in lieu of uncertainty reduction;
iii) a self-supervised representation learning method leveraging auxiliary tasks to find state representation interpretable in its own right for POMDPs so that they can be solved with the vast choices of more computationally efficient MDP algorithms with respect to the extracted states...
iii) a self-supervised representation learning method leveraging auxiliary tasks to find state representation interpretable in its own right for POMDPs so that they can be solved with the vast choices of more computationally efficient MDP algorithms with respect to the extracted states...
Version
Open Access
Date Issued
2022-05-10
Date Awarded
01/03/2023
License URL
Advisor
Faisal, Aldo
Sponsor
Imperial College London
Publisher Department
Department of Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
