Machine learning methods for audio-visual event analysis
File(s)
Author(s)
Pu, Jie
Type
Thesis
Abstract
This thesis studies machine learning techniques for localizing, separating and recognizing audio-visual events. Until recently, the widely-used methods for analyzing audio-visual events involves either laborious pre-processing/post-processing steps to handle videos, or a huge amount of training data to supervise their learning processes. To overcome these limitations, we aim to develop novel approaches that have one end-to-end framework, and much less dependence on training data than state-of-the-art deep-learning approaches. In particular, we propose novel low-rank and sparse matrix decomposition methods, kernelized matrix decomposition methods and deep neural networks for audio-visual event analysis, with a particular focus on their data-efficiency and computational-efficiency.
First, we investigate the problem of blind audio-visual localization and separation, which aims to localize visual objects associated with an audio signal and simultaneously separate the audio signal from irrelevant audio components. To this end, a novel low-rank and sparse matrix decomposition method is devised, where we use a sparse matrix capturing the correlated components between visual and audio modalities, and hence uncovering the sound source in visual modality and the associated sound in audio modality. After that, we propose a novel kernelized sparse and low-rank decomposition method, which generalizes the linear correlation of our low-rank and sparse model into non-linear ones. This leads to superior performances in the application of active speaker detection and localization in audio-visual recordings.
Given the recent surge of deep learning, we also propose a novel deep low-rank and sparse model, which uses deep neural networks to achieve low-rank and sparse decomposition. It is a further extension of our kernelized matrix decomposition method, where the non-linear correlation between visual and audio modalities is captured by deep networks. Finally, we have worked on the problem of sound event detection and localization. A novel filterbank learning approach is proposed, which takes the raw waveform as input and produces its auditory spectro-temporal representation. In contrast to standard convolutional neural networks that learn all elements of each filter, the proposed filters are parameterized with only two learnable variables. This offers the interpretability of our model after training, and greatly lessen its requirement for training data.
First, we investigate the problem of blind audio-visual localization and separation, which aims to localize visual objects associated with an audio signal and simultaneously separate the audio signal from irrelevant audio components. To this end, a novel low-rank and sparse matrix decomposition method is devised, where we use a sparse matrix capturing the correlated components between visual and audio modalities, and hence uncovering the sound source in visual modality and the associated sound in audio modality. After that, we propose a novel kernelized sparse and low-rank decomposition method, which generalizes the linear correlation of our low-rank and sparse model into non-linear ones. This leads to superior performances in the application of active speaker detection and localization in audio-visual recordings.
Given the recent surge of deep learning, we also propose a novel deep low-rank and sparse model, which uses deep neural networks to achieve low-rank and sparse decomposition. It is a further extension of our kernelized matrix decomposition method, where the non-linear correlation between visual and audio modalities is captured by deep networks. Finally, we have worked on the problem of sound event detection and localization. A novel filterbank learning approach is proposed, which takes the raw waveform as input and produces its auditory spectro-temporal representation. In contrast to standard convolutional neural networks that learn all elements of each filter, the proposed filters are parameterized with only two learnable variables. This offers the interpretability of our model after training, and greatly lessen its requirement for training data.
Version
Open Access
Date Issued
2020-03
Date Awarded
2020-07
Copyright Statement
Creative Commons Attribution NonCommercial NoDerivatives Licence
Advisor
Pantic, Maja
Publisher Department
Department of Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
