Efficient and scalable quality-diversity for open-ended robot learning
File(s)
Author(s)
Lim, Bryan Wei Tern
Type
Thesis
Abstract
General-purpose robots can benefit society by assisting and enabling humans to perform various tasks across different environments. While progress has been made towards this goal across disciplines, robots still struggle to obtain the creativity, generality and reliability we want. This is due to the diversity and open-ended nature of tasks and environments in the real-world.
We posit that open-ended problems require open-ended solutions. Open-ended systems continuously invent and generate interesting problems and solutions. In this thesis, we present Quality-Diversity (QD) as a framework for open-ended search and robot learning. We first explore QD algorithms for online learning. Specifically, we combine learning world models and QD search to leverage the synergy between predictive and generative capabilities of deep models and the open-ended creativity of QD. We show this significantly increases the efficiency of QD and enables robots to learn diverse skills from scratch, directly in the real-world in a safe and autonomous manner.
While learning online is important to effectively collect new experiences and adapt, it is also practical to use data, simulation and pre-trained models to build in strong priors into robots offline before deployment. In the second part, we propose methods which highlight QD as effective data generators and for effective learning of diverse priors in the offline setting. To enable this, we present work that scales QD algorithms through tensorization and parallelization. Using hardware accelerators, this results in ∼100× speed up in wall-time through compared to previous implementations, unlocking algorithmic innovation with QD through speed and scale.
Finally, we present a unified perspective of QD algorithms and conclude by highlighting the insights from this thesis provide an algorithmic foundation for open-ended learning systems that can improve the creativity, generality and reliability of embodied AI systems like robots and beyond to any AI system.
We posit that open-ended problems require open-ended solutions. Open-ended systems continuously invent and generate interesting problems and solutions. In this thesis, we present Quality-Diversity (QD) as a framework for open-ended search and robot learning. We first explore QD algorithms for online learning. Specifically, we combine learning world models and QD search to leverage the synergy between predictive and generative capabilities of deep models and the open-ended creativity of QD. We show this significantly increases the efficiency of QD and enables robots to learn diverse skills from scratch, directly in the real-world in a safe and autonomous manner.
While learning online is important to effectively collect new experiences and adapt, it is also practical to use data, simulation and pre-trained models to build in strong priors into robots offline before deployment. In the second part, we propose methods which highlight QD as effective data generators and for effective learning of diverse priors in the offline setting. To enable this, we present work that scales QD algorithms through tensorization and parallelization. Using hardware accelerators, this results in ∼100× speed up in wall-time through compared to previous implementations, unlocking algorithmic innovation with QD through speed and scale.
Finally, we present a unified perspective of QD algorithms and conclude by highlighting the insights from this thesis provide an algorithmic foundation for open-ended learning systems that can improve the creativity, generality and reliability of embodied AI systems like robots and beyond to any AI system.
Version
Open Access
Date Issued
2024-03
Date Awarded
2024-06
Copyright Statement
Creative Commons Attribution NonCommercial Licence
License URL
Advisor
Cully, Antoine
Publisher Department
Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
