Audio-visual speech generation and representation learning
File(s)
Author(s)
Kefalas, Triantafyllos
Type
Thesis
Abstract
The progress of Deep Learning has led to the widespread adoption of interfaces and intelligent devices for human-computer interaction. These often rely on audio-visual communication and must interact with a variety of end users across diverse and potentially challenging real-world settings. This requires learning good representations of audio-visual speech, which we explore in this thesis. Progress in many audio-visual speech tasks is constrained by the scarcity of labelled, “in-the- wild” data. To this end, we collect the KAN-AV dataset, which is a large-scale audio and video dataset of speech captured “in-the-wild”. It contains age, kinship and identity annotations, allowing us to investigate challenging problems and particularly the problem of learning cross-modal representations, i.e., comparing and matching samples from different modalities. Crucially, the dataset contains samples across ages for each subject, enabling us to investigate the learning of age-invariant representations. We then turn to the task of speech-driven facial animation and explore a principled approach to combining audio-visual latent representations. We introduce the polynomial fusion layer, formulating a joint representation that includes higher-order interactions. We demonstrate the effectiveness of this approach in experiments with audio-visual speech datasets. The remainder of the thesis investigates the generation of speech from silent videos. We explore pre-training the decoder of a video-to-speech model on large volumes of audio-only data. Concretely, we pre-train audio encoder-decoder models and then fine-tune their decoders for video-to-speech synthesis. We demonstrate that this approach improves the reconstructed speech in both raw waveform and mel spectrogram generation. Finally, we note that previous works either employ silent video inputs only, or video and audio inputs and discard the audio input pathway during inference. We propose a method to include audio and video inputs during both training and inference, by synthesizing the input audio first. Our experiments show that this method outperforms previous approaches.
Version
Open Access
Date Issued
2023-07-14
Date Awarded
01/03/2025
License URL
Advisor
Pantic, Maja
Sponsor
Engineering and Physical Sciences Research Council
Publisher Department
Department of Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
