MAP-elites with descriptor-conditioned gradients and archive distillation into a single policy
File(s)3583131.3590503.pdf (5.64 MB)
Published version
Author(s)
Faldor, Maxence
Chalumeau, Félix
Flageat, Manon
Cully, Antoine
Type
Conference Paper
Abstract
Quality-Diversity algorithms, such as MAP-Elites, are a branch of Evolutionary Computation generating collections of diverse and high-performing solutions, that have been successfully applied to a variety of domains and particularly in evolutionary robotics. However, MAP-Elites performs a divergent search based on random mutations originating from Genetic Algorithms, and thus, is limited to evolving populations of low-dimensional solutions. PGA-MAP-Elites overcomes this limitation by integrating a gradient-based variation operator inspired by Deep Reinforcement Learning which enables the evolution of large neural networks. Although high-performing in many environments, PGA-MAP-Elites fails on several tasks where the convergent search of the gradient-based operator does not direct mutations towards archive-improving solutions. In this work, we present two contributions: (1) we enhance the Policy Gradient variation operator with a descriptor-conditioned critic that improves the archive across the entire descriptor space, (2) we exploit the actor-critic training to learn a descriptor-conditioned policy at no additional cost, distilling the knowledge of the archive into one single versatile policy that can execute the entire range of behaviors contained in the archive. Our algorithm, DCG-MAP-Elites improves the QD score over PGA-MAP-Elites by 82% on average, on a set of challenging locomotion tasks.
Date Issued
2023-07-12
Date Acceptance
2023-03-31
Citation
GECCO '23: Proceedings of the Genetic and Evolutionary Computation Conference, 2023, pp.138-146
ISBN
9798400701191
Publisher
Association for Computing Machinery
Start Page
138
End Page
146
Journal / Book Title
GECCO '23: Proceedings of the Genetic and Evolutionary Computation Conference
Copyright Statement
© 2023 Copyright held by the owner/author(s). This work is licensed under a Creative Commons Attribution International 4.0 License (https://creativecommons.org/licenses/by/4.0/)
License URL
Source
The Genetic and Evolutionary Computation Conference
Subjects
MAP-Elites
Neuroevolution
Policy Gradient
Quality-Diversity
Reinforcement Learning
Publication Status
Published
Start Date
2023-07-15
Finish Date
2023-07-19
Coverage Spatial
Lisbon, Portugal