Binaural enhancement of speech for hearing assistive devices
File(s)
Author(s)
Tokala, Vikas Dhanapal
Type
Thesis
Abstract
Humans rely on binaural hearing—two ears—to enhance spatial awareness and speech intelligibility, especially in noisy environments. Monaural speech enhancement methods have advanced but often fail to preserve interaural cues, limiting localization and listener comfort. Binaural enhancement addresses this by improving intelligibility while maintaining spatial perception, critical for hearing aids, communication systems, and AR/VR.
This thesis investigates binaural speech enhancement techniques that preserve spatial cues while improving intelligibility. It extends monaural mask-assisted enhancement to binaural signals using STOI-optimal masks and a feed-forward DNN to improve weighted-STOI scores. To avoid artifacts from applying time-frequency (TF) masks in the STFT domain, an Optimally Modified-LSA approach is proposed to enhance perceptual quality.
Recent neural methods in both time and TF domains show promise, with complex-valued networks outperforming real-valued ones by jointly estimating magnitude and phase. Building on this, an end-to-end binaural framework using complex convolutional encoder-decoder networks is introduced. Two models are proposed: one with a complex self-attention transformer block, another with a complex recurrent block. Both estimate complex ratio masks for each ear channel and use a custom loss balancing spatial preservation, intelligibility, and noise reduction. These models outperform existing methods such as binaural TasNet and STOI-optimal masking in isotropic noise.
Additionally, two multichannel methods are explored: a beamformer-guided complex neural network and a multichannel encoder-decoder network, leveraging multiple microphones in modern binaural devices. Both achieve significant gains in intelligibility and noise reduction while preserving spatial cues, validated through simulations, real-world recordings, and MUSHRA listening tests. Finally, an end-to-end sound source localization model is developed to evaluate spatial cue preservation, showing superior performance in noisy conditions compared to traditional localization approaches.
This thesis investigates binaural speech enhancement techniques that preserve spatial cues while improving intelligibility. It extends monaural mask-assisted enhancement to binaural signals using STOI-optimal masks and a feed-forward DNN to improve weighted-STOI scores. To avoid artifacts from applying time-frequency (TF) masks in the STFT domain, an Optimally Modified-LSA approach is proposed to enhance perceptual quality.
Recent neural methods in both time and TF domains show promise, with complex-valued networks outperforming real-valued ones by jointly estimating magnitude and phase. Building on this, an end-to-end binaural framework using complex convolutional encoder-decoder networks is introduced. Two models are proposed: one with a complex self-attention transformer block, another with a complex recurrent block. Both estimate complex ratio masks for each ear channel and use a custom loss balancing spatial preservation, intelligibility, and noise reduction. These models outperform existing methods such as binaural TasNet and STOI-optimal masking in isotropic noise.
Additionally, two multichannel methods are explored: a beamformer-guided complex neural network and a multichannel encoder-decoder network, leveraging multiple microphones in modern binaural devices. Both achieve significant gains in intelligibility and noise reduction while preserving spatial cues, validated through simulations, real-world recordings, and MUSHRA listening tests. Finally, an end-to-end sound source localization model is developed to evaluate spatial cue preservation, showing superior performance in noisy conditions compared to traditional localization approaches.
Version
Open Access
Date Issued
2025-03-19
Date Awarded
01/09/2025
License URL
Advisor
Naylor, Patrick
Brookes, Mike
Sponsor
European Commission
Grant Number
956369
Publisher Department
Department of Electrical and Electronic Engineering
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
