On self-supervision and foundation models in robot learning
File(s)
Author(s)
Di Palo, Norman
Type
Thesis
Abstract
Robotics unlocked an era of unprecedented automation in the creation and manufacturing of goods and services. However, the vast majority of robots are currently confined into industrial plants, with a hand-defined behaviour they repeat, and unable to deal with the uncertainties and difficulties of the real world. Equipping robots with the same dexterity, common-sense, and skilful interactions that humans exhibit in their environments, such as homes and offices, could unlock orders of magnitude more tasks where robots could cooperate with humans and substantially reduce the cost and scarcity of many goods and services.
Traditional robotics techniques are insufficient to develop the vast repertoire of skills required for human-like interaction. The recent literature has therefore adopted tools and techniques from artificial intelligence and machine learning, where algorithms are no longer manually designed to perform tasks, but instead a general learning algorithm is employed that can learn to emulate the desired behaviour and demonstrate emergent abilities. This led to the popularisation of the field of robot learning. However, the main current bottleneck is the need for large datasets that these methods generally require to obtain mastery of the desired behaviours.
In this thesis, we focus on a set of techniques and paradigms that can tackle the data hunger of large machine learning models. In particular, we focus on two main families of methods: 1) self-supervised learning and data collection methods, that allow robots to collect data autonomously, strongly reducing the amount of human time needed to learn new tasks, and 2) methods based on existing foundation models, that can transfer the abilities of these models, trained on non-robotics specific data, to robotics applications.
Traditional robotics techniques are insufficient to develop the vast repertoire of skills required for human-like interaction. The recent literature has therefore adopted tools and techniques from artificial intelligence and machine learning, where algorithms are no longer manually designed to perform tasks, but instead a general learning algorithm is employed that can learn to emulate the desired behaviour and demonstrate emergent abilities. This led to the popularisation of the field of robot learning. However, the main current bottleneck is the need for large datasets that these methods generally require to obtain mastery of the desired behaviours.
In this thesis, we focus on a set of techniques and paradigms that can tackle the data hunger of large machine learning models. In particular, we focus on two main families of methods: 1) self-supervised learning and data collection methods, that allow robots to collect data autonomously, strongly reducing the amount of human time needed to learn new tasks, and 2) methods based on existing foundation models, that can transfer the abilities of these models, trained on non-robotics specific data, to robotics applications.
Version
Open Access
Date Issued
2025-02-26
Date Awarded
2026-05-01
Copyright Statement
Attribution-NonCommercial 4.0 International Licence (CC BY-NC)
License URL
Advisor
Johns, Edward
Publisher Department
Department of Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
