Deep representation learning for audio-visual speech recognition and synthesis
File(s)
Author(s)
Zinonos, Andreas
Type
Thesis
Abstract
Deep learning has grown rapidly in capability and impact over the past decade, now achieving industry-disrupting results that rival or surpass traditional methods across many domains. While most early progress relied heavily on large labelled datasets, recent advances show that powerful models can be trained by exploiting the abundance of unlabelled data available in the wild. Audio-visual recordings of people speaking are particularly well suited for this paradigm, as they provide rich and naturally aligned cross-modal signals. This thesis investigates how large-scale unlabelled audio-visual data of talking humans can be used to improve two important tasks: speech recognition and lip synchronization. We first extend the self-supervised learning framework RAVEn to multilingual settings, studying how cross-lingual audio-visual pre-training benefits visual speech recognition, especially for languages with limited labelled resources. We then propose BRAVEn, an improved version of RAVEn that more effectively exploits the complementary structure of audio and video through architectural refinements and modality-aware masking, yielding substantial gains in both visual and audio speech recognition, particularly under low-label conditions. Finally, we turn to generative modeling and introduce FlashLips, a two-stage lip synchronization system that produces temporally accurate lip motion at over 100 FPS, while matching the quality of significantly slower iterative diffusion-based approaches. Together, these contributions demonstrate that unlabelled audio-visual data can be leveraged effectively for both discriminative and generative tasks, enabling stronger speech recognition systems and faster, high-fidelity lip-sync generation without the need for expensive manual annotation.
Version
Open Access
Date Issued
2025-12-29
Date Awarded
2026-06-01
Copyright Statement
Attribution-NonCommercial 4.0 International Licence (CC BY-NC)
License URL
Advisor
Pantic, Maja
Publisher Department
Department of Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
