Quality-diversity actor-critic: learning high-performing and diverse behaviors via value and successor features critics
File(s)grillotti24a.pdf (9.47 MB)
Accepted version
Author(s)
Faldor, Maxence
Grillotti, Luca
Gonzalez Leon, Borja
Cully, Antoine
Type
Conference Paper
Abstract
A key aspect of intelligence is the ability to demonstrate a broad spectrum of behaviors for adapting to unexpected situations. Over the past decade, advancements in deep reinforcement learning have led to groundbreaking achievements to solve complex continuous control tasks. However, most approaches return only one solution specialized for a specific problem. We introduce Quality-Diversity Actor-Critic (QDAC), an off-policy actor-critic deep reinforcement learning algorithm that leverages a value function critic and a successor features critic to learn high-performing and diverse behaviors. In this framework, the actor optimizes an objective that seamlessly unifies both critics using constrained optimization to (1) maximize return, while (2) executing diverse skills. Compared with other Quality-Diversity methods, QDAC achieves significantly higher performance and more diverse behaviors on six challenging continuous control locomotion tasks. We also demonstrate that we can harness the learned skills to adapt better than other baselines to five perturbed environments. Finally, qualitative analyses showcase a range of remarkable behaviors, available at: http://bit.ly/qdac.
Date Issued
2024-07-21
Date Acceptance
2024-05-02
Citation
Proceedings of Machine Learning Research, 2024, 235, pp.16416-16459
ISSN
2640-3498
Publisher
MLResearchPress
Start Page
16416
End Page
16459
Journal / Book Title
Proceedings of Machine Learning Research
Volume
235
Copyright Statement
Copyright 2024 by the author(s) and PMLR.
Identifier
https://maxencefaldor.github.io/
Source
International Conference on Machine Learning
Publication Status
Published
Start Date
2024-07-21
Finish Date
2024-07-27
Coverage Spatial
Vienna, Austria