Learning self-supervised representations of audiovisual human-centric data
File(s)
Author(s)
Shukla, Abhinav
Type
Thesis
Abstract
Deep learning models have driven significant progress in audio and audiovisual machine learning but rely on extensive supervision from large labeled datasets. While labels are often intractable to obtain, unlabeled data is significantly more pervasive and ever growing. Self-supervised learning leverages the implicit structure of unlabeled data to learn better representations that generalize better to smaller labeled datasets of interest. This is relevant in domains like emotion and speech recognition with limited labeled data, where training from scratch is infeasible and exhaustive supervision is unavailable.
In this thesis, we study the effectiveness of audiovisual self-supervised learning for learning better unimodal audio representations. The core idea is to leverage the natural synchrony and cross-modal correspondence between audio and video as a source of cross-modal self-supervision. Audio-only self-supervised learning works well, but complementary addition of visual self-supervision produces better representations. We study speech-driven facial reconstruction as a pretext task for audio representation learning in speech and find that the correlation between lip movements in video and speech content in audio; and that between facial expressions in video and speech emotion in audio provides a very rich supervisory signal. We propose new methods for audiovisual self-supervised learning for more general audiovisual content that are based on combining pretext tasks like clustering, invariance-based contrastive learning, and audiovisual sound source localization. We demonstrate strong results on a number of audio recognition tasks in both audiovisual speech data and general audiovisual content. Compared to existing methods, our proposed representations achieve better performance with smaller amounts of pretraining data while getting even better with additional data; are more robust in noisy conditions; and perform better on downstream tasks with limited labeled data. The work provides compelling evidence for the utility of self-supervision in general, and visually guided self-supervision in particular, for learning better representations of audio.
In this thesis, we study the effectiveness of audiovisual self-supervised learning for learning better unimodal audio representations. The core idea is to leverage the natural synchrony and cross-modal correspondence between audio and video as a source of cross-modal self-supervision. Audio-only self-supervised learning works well, but complementary addition of visual self-supervision produces better representations. We study speech-driven facial reconstruction as a pretext task for audio representation learning in speech and find that the correlation between lip movements in video and speech content in audio; and that between facial expressions in video and speech emotion in audio provides a very rich supervisory signal. We propose new methods for audiovisual self-supervised learning for more general audiovisual content that are based on combining pretext tasks like clustering, invariance-based contrastive learning, and audiovisual sound source localization. We demonstrate strong results on a number of audio recognition tasks in both audiovisual speech data and general audiovisual content. Compared to existing methods, our proposed representations achieve better performance with smaller amounts of pretraining data while getting even better with additional data; are more robust in noisy conditions; and perform better on downstream tasks with limited labeled data. The work provides compelling evidence for the utility of self-supervision in general, and visually guided self-supervision in particular, for learning better representations of audio.
Version
Open Access
Date Issued
2023-03
Date Awarded
2024-10
Copyright Statement
Creative Commons Attribution NonCommercial Licence
License URL
Advisor
Pantic, Maja
Petridis, Stavros
Publisher Department
Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)