Diversity-based trajectory and goal selection with hindsight experience replay
File(s) 2108.07887v1.pdf (2.2 MB)
Accepted version
Author(s)
Dai, Tianhong
Liu, Hengyan
Arulkumaran, Kai
Ren, Guangyu
Bharath, Anil Anthony
Type
Conference Paper
Abstract
Hindsight experience replay (HER) is a goal relabelling technique typically
used with off-policy deep reinforcement learning algorithms to solve
goal-oriented tasks; it is well suited to robotic manipulation tasks that
deliver only sparse rewards. In HER, both trajectories and transitions are
sampled uniformly for training. However, not all of the agent's experiences
contribute equally to training, and so naive uniform sampling may lead to
inefficient learning. In this paper, we propose diversity-based trajectory and
goal selection with HER (DTGSH). Firstly, trajectories are sampled according to
the diversity of the goal states as modelled by determinantal point processes
(DPPs). Secondly, transitions with diverse goal states are selected from the
trajectories by using k-DPPs. We evaluate DTGSH on five challenging robotic
manipulation tasks in simulated robot environments, where we show that our
method can learn more quickly and reach higher performance than other
state-of-the-art approaches on all tasks.
used with off-policy deep reinforcement learning algorithms to solve
goal-oriented tasks; it is well suited to robotic manipulation tasks that
deliver only sparse rewards. In HER, both trajectories and transitions are
sampled uniformly for training. However, not all of the agent's experiences
contribute equally to training, and so naive uniform sampling may lead to
inefficient learning. In this paper, we propose diversity-based trajectory and
goal selection with HER (DTGSH). Firstly, trajectories are sampled according to
the diversity of the goal states as modelled by determinantal point processes
(DPPs). Secondly, transitions with diverse goal states are selected from the
trajectories by using k-DPPs. We evaluate DTGSH on five challenging robotic
manipulation tasks in simulated robot environments, where we show that our
method can learn more quickly and reach higher performance than other
state-of-the-art approaches on all tasks.
Date Issued
2021-11-01
Date Acceptance
2021-08-09
Citation
2021, 13033, pp.32-45
Publisher
Springer
Start Page
32
End Page
45
Volume
13033
Copyright Statement
© 2021 Springer Nature Switzerland AG. The final publication is available at Springer via https://doi.org/10.1007/978-3-030-89370-5_3
Identifier
http://arxiv.org/abs/2108.07887v1
Source
18th Pacific Rim International Conference on Artificial Intelligence (PRICAI)
Subjects
cs.LG
cs.LG
cs.RO
Publication Status
Published
Start Date
2021-11-08
Finish Date
2021-11-12
Coverage Spatial
Hanoi, Vietnam
Date Publish Online
2021-11-01
