Converging pathways: structured visual representations with multi-task learning
File(s)
Author(s)
Liu, Shikun
Type
Thesis
Abstract
Human cognition is inherently adept at concurrently managing multiple tasks, and efficiently leveraging past experiences to acquire new skills. In machine learning, multi-task learning is the learning paradigm that seeks to emulate this cognitive capability by learning shared features from related tasks to improve training efficiency. However, the straightforward application of multi-task learning may lead to sub-optimal performance, owing to the variations in task-specific complexity and distinct training objectives.
In this thesis, we conduct a comprehensive investigation of multi-task learning within the domain of computer vision, resulting in novel solutions in multi-task learning for constructing structured visual representations --- a foundational building block for a wide range of applications, spanning from visual perception and object recognition to vision-based control and reasoning.
Our research includes a spectrum of training strategies and optimisation methods, all intricately designed to enhance the efficacy of structured visual representations. Specifically, we delve into three critical dimensions: i) the design strategy of harnessing multi-task relationships to improve computer vision model performance, ii) the creation of auxiliary tasks to improve computer vision model generalisation, and iii) the utilisation of multi-task knowledge within pre-trained experts to improve open-ended visual reasoning.
We introduce a series of multi-task learning frameworks to address these research questions and showcase their effectiveness in improving the generalisation, training efficiency, and interpretability of computer vision systems. Through this comprehensive exploration of multi-task learning and its implications, we aim to contribute to the development of more intelligent and versatile systems for challenging real-world applications.
In this thesis, we conduct a comprehensive investigation of multi-task learning within the domain of computer vision, resulting in novel solutions in multi-task learning for constructing structured visual representations --- a foundational building block for a wide range of applications, spanning from visual perception and object recognition to vision-based control and reasoning.
Our research includes a spectrum of training strategies and optimisation methods, all intricately designed to enhance the efficacy of structured visual representations. Specifically, we delve into three critical dimensions: i) the design strategy of harnessing multi-task relationships to improve computer vision model performance, ii) the creation of auxiliary tasks to improve computer vision model generalisation, and iii) the utilisation of multi-task knowledge within pre-trained experts to improve open-ended visual reasoning.
We introduce a series of multi-task learning frameworks to address these research questions and showcase their effectiveness in improving the generalisation, training efficiency, and interpretability of computer vision systems. Through this comprehensive exploration of multi-task learning and its implications, we aim to contribute to the development of more intelligent and versatile systems for challenging real-world applications.
Version
Open Access
Date Issued
2023-11
Date Awarded
2024-05
Copyright Statement
Creative Commons Attribution NonCommercial NoDerivatives Licence
Advisor
Davison, Andrew
Johns, Edward
Sponsor
Dyson Ltd.
Publisher Department
Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
