Understanding the synergies between quality-diversity and deep reinforcement learning
File(s)3583131.3590388.pdf (1.79 MB)
Published version
Author(s)
Lim, Bryan Wei Tern
Flageat, Manon
Cully, Antoine
Type
Conference Paper
Abstract
The synergies between Quality-Diversity (QD) and Deep Reinforcement Learning (RL) have led to powerful hybrid QD-RL algorithms that have shown tremendous potential, and bring the best of both fields. However, only a single deep RL algorithm (TD3) has been used in prior hybrid methods despite notable progress made by other RL algorithms. Additionally, there are fundamental differences
in the optimization procedures between QD and RL which would benefit from a more principled approach. We propose Generalized Actor-Critic QD-RL, a unified modular framework for actor-critic deep RL methods in the QD-RL setting. This framework provides a path to study insights from Deep RL in the QD-RL setting, which is an important and efficient way to make progress in QD-RL. We
introduce two new algorithms, PGA-ME (SAC) and PGA-ME (DroQ) which apply recent advancements in Deep RL to the QD-RL setting, and solve the humanoid environment which was not possible using existing QD-RL algorithms. However, we also find that not all insights from Deep RL can be effectively translated to QD-RL. Critically, this work also demonstrates that the actor-critic models in QD-RL are generally insufficiently trained and performance gains
can be achieved without any additional environment evaluations.
in the optimization procedures between QD and RL which would benefit from a more principled approach. We propose Generalized Actor-Critic QD-RL, a unified modular framework for actor-critic deep RL methods in the QD-RL setting. This framework provides a path to study insights from Deep RL in the QD-RL setting, which is an important and efficient way to make progress in QD-RL. We
introduce two new algorithms, PGA-ME (SAC) and PGA-ME (DroQ) which apply recent advancements in Deep RL to the QD-RL setting, and solve the humanoid environment which was not possible using existing QD-RL algorithms. However, we also find that not all insights from Deep RL can be effectively translated to QD-RL. Critically, this work also demonstrates that the actor-critic models in QD-RL are generally insufficiently trained and performance gains
can be achieved without any additional environment evaluations.
Date Issued
2023-07-12
Date Acceptance
2023-03-31
Citation
GECCO '23: Proceedings of the Genetic and Evolutionary Computation Conference, 2023, pp.1212-1220
Publisher
ACM
Start Page
1212
End Page
1220
Journal / Book Title
GECCO '23: Proceedings of the Genetic and Evolutionary Computation Conference
Copyright Statement
© 2023 Copyright held by the owner/author(s). This work is licensed under a Creative Commons Attribution International 4.0 License (https://creativecommons.org/licenses/by/4.0/)
Source
The Genetic and Evolutionary Computation Conference (GECCO '23)
Publication Status
Published
Start Date
2023-07-15
Finish Date
2023-07-19
Coverage Spatial
Lisbon, Portugal