Exploring Restart Distributions
File(s)Tavakoli_RLDM-2019.pdf (264.12 KB)
Accepted version
Author(s)
Tavakoli, Arash
Levdik, Vitaly
Islam, Riashat
Smith, Christopher M
Kormushev, Petar
Type
Conference Paper
Abstract
We consider the generic approach of using an experience memory to help exploration by adapting a restart distribution. That is, given the capacity to reset the state with those corresponding to the agent's past observations, we help exploration by promoting faster state-space coverage via restarting the agent from a more diverse set of initial states, as well as allowing it to restart in states associated with significant past experiences. This approach is compatible with both on-policy and off-policy methods. However, a caveat is that altering the distribution of initial states could change the optimal policies when searching within a restricted class of policies. To reduce this unsought learning bias, we evaluate our approach in deep reinforcement learning which benefits from the high representational capacity of deep neural networks. We instantiate three variants of our approach, each inspired by an idea in the context of experience replay. Using these variants, we show that performance gains can be achieved, especially in hard exploration problems.
Date Issued
2019-07
Date Acceptance
2019-04-12
Citation
4th Multidisciplinary Conference on Reinforcement Learning and Decision Making (RLDM 2019), 2019
Publisher
arXiv
Journal / Book Title
4th Multidisciplinary Conference on Reinforcement Learning and Decision Making (RLDM 2019)
Copyright Statement
© 2020 The Author(s)
Identifier
https://arxiv.org/abs/1811.11298
Source
The Fourth Multidisciplinary Conference on Reinforcement Learning and Decision Making
Subjects
cs.LG
stat.ML
Place of Publication
Montréal, Canada
Publication Status
Published
Start Date
2019-07-07
Coverage Spatial
Montréal, QC, Canada