Deep learning from uncertainty: exemplified by applications on audio and text
File(s)
Author(s)
Rizos, Georgios
Type
Thesis
Abstract
In this thesis, we explore the potential of using deep learning to first quantify, and then learn from uncertainty towards modifying the importance of data samples during the training process, depending on how likely focusing on their modelling is to improve predictive behaviour. We are specifically interested in the usage of uncertainty as a source of information in its own right, regardless of whether it is due to model parameter stochasticity (model/epistemic) or label disagreement (label/aleatory), and apply our proposed methodologies to the modelling of audio and text data.
We propose a method quantifying an informativeness value [1] (Chapter 3) per annotated time frame of an audio speech affect sample, with the purpose of focusing training on highly informative frames. The informativeness value is dependent, through a separate, trainable model, on various factors of uncertainty about each labelled frame. This method comprises a proof of concept for learning a bespoke, uncertainty-based quantity used for explicitly modifying the supervision process towards improving the model performance on the same data.
We investigate further, and propose an uncertainty-aware, adaptive, and sample- specific variant of label smoothing [2] (Chapter 4) with the aim of improving both the accuracy and calibration performance on two novel animal call detection datasets. Our method relies on the sample-free propagation of uncertainty in a variational Bayesian version of a residual convolutional network with squeeze-and-excitation and multiple head attention model that we show excels in bioacoustic call detection [3].
We investigate uncertainty-aware scientific paper submission modelling [4, 5] (Chapter 5), focusing on acceptance prediction, through a framework that models reviewer recommendation disagreement. We achieve a performance improvement by introducing the auxiliary task of modelling label stochasticity, whether we use the true, empirical label disagreement as supervision, or by attempting to infer it.
We propose a method quantifying an informativeness value [1] (Chapter 3) per annotated time frame of an audio speech affect sample, with the purpose of focusing training on highly informative frames. The informativeness value is dependent, through a separate, trainable model, on various factors of uncertainty about each labelled frame. This method comprises a proof of concept for learning a bespoke, uncertainty-based quantity used for explicitly modifying the supervision process towards improving the model performance on the same data.
We investigate further, and propose an uncertainty-aware, adaptive, and sample- specific variant of label smoothing [2] (Chapter 4) with the aim of improving both the accuracy and calibration performance on two novel animal call detection datasets. Our method relies on the sample-free propagation of uncertainty in a variational Bayesian version of a residual convolutional network with squeeze-and-excitation and multiple head attention model that we show excels in bioacoustic call detection [3].
We investigate uncertainty-aware scientific paper submission modelling [4, 5] (Chapter 5), focusing on acceptance prediction, through a framework that models reviewer recommendation disagreement. We achieve a performance improvement by introducing the auxiliary task of modelling label stochasticity, whether we use the true, empirical label disagreement as supervision, or by attempting to infer it.
Version
Open Access
Date Issued
2023-02
Date Awarded
2024-08
Copyright Statement
Creative Commons Attribution NonCommercial Licence
License URL
Advisor
Schuller, Bjoern
Sponsor
Engineering and Physical Sciences Research Council
Grant Number
2021037
Publisher Department
Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)