Diversity-augmented intrinsic motivation for deep reinforcement learning
File(s) diversity_rl_neurocomputing.pdf (3.2 MB)
Accepted version
Author(s)
Dai, Tianhong
Du, Yali
Fang, Meng
Bharath, Anil
Type
Journal Article
Abstract
In many real-world problems, reward signals received by agents are delayed or sparse, which makes it challenging to train a reinforcement learning (RL) agent. An intrinsic reward signal can help an agent to explore such environments in the quest for novel states. In this work, we propose a general end-to-end diversity-augmented intrinsic motivation for deep reinforcement learning which encourages the agent to explore new states and automatically provides denser rewards. Specifically, we measure the diversity of adjacent states under a model of state sequences based on determinantal point process (DPP); this is coupled with a straight-through gradient estimator to enable end-to-end differentiability. The proposed approach is comprehensively evaluated on the MuJoCo and the Arcade Learning Environments (Atari and SuperMarioBros). The experiments show that an intrinsic reward based on the diversity measure derived from the DPP model accelerates the early stages of training in Atari games and SuperMarioBros. In MuJoCo, the approach improves on prior techniques for tasks using the standard reward setting, and achieves the state-of-the-art performance on 12 out of 15 tasks containing delayed rewards.
Date Issued
2022-01-11
Date Acceptance
2021-10-19
Citation
Neurocomputing, 2022, 468, pp.396-406
ISSN
0925-2312
Publisher
Elsevier
Start Page
396
End Page
406
Journal / Book Title
Neurocomputing
Volume
468
Copyright Statement
© 2021 Elsevier Ltd. All rights reserved. This manuscript is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International Licence http://creativecommons.org/licenses/by-nc-nd/4.0/
Identifier
https://www.sciencedirect.com/science/article/pii/S0925231221015265?via%3Dihub
Subjects
Science & Technology
Technology
Computer Science, Artificial Intelligence
Computer Science
Deep reinforcement learning
Curiosity-driven exploration
Determinantal point process
NEURAL-NETWORKS
08 Information and Computing Sciences
09 Engineering
17 Psychology and Cognitive Sciences
Artificial Intelligence & Image Processing
Publication Status
Published
Date Publish Online
2021-10-22
