Decoding auditory EEG responses to speech via deep neural networks
File(s)
Author(s)
Thornton, Mike
Type
Thesis
Abstract
Neural speech tracking – the phenomenon by which a listener’s neural activity synchronises with a speech stimulus – is represented in the listener’s electro-encephalogram (EEG). In the continuous speech paradigm, participants listen to extended, naturalistic narratives whilst their EEG is recorded; various aspects of their perception and cognition can subsequently be decoded from these measurements. In this thesis, new methods based on deep learning are employed for the purpose of auditory EEG decoding. Accurate auditory EEG decoders could find clinical applications in the objective diagnosis of hearing disorders, or in cognitively-steered hearing aids which monitor the focus of the listener’s auditory attention and adapt their behaviour accordingly.
We first employed non-linear models based on deep neural networks for the purpose of reconstructing the speech envelope from the listeners’ EEG measurements. These decoders performed significantly better than their linear counterparts, and could also function as auditory attention decoders. However, they did not generalise well between different listeners, a property which would be valuable in translational applications. We then turned to a different decoding problem, the match-mismatch problem, which was part of the ICASSP 2023/2024 Auditory EEG Decoding Signal Processing Grand Challenge series. Inspired by known responses to speech – envelope tracking and speech-related frequency-following responses – we developed highly accurate, challenge-winning decoders, which this time generalised well between participants and even across datasets. Finally, we applied some of these techniques in an auditory attention identification setting using EEG signals recorded from a wearable ear-EEG device. Whilst auditory attention could be identified with significance using this device, the results demonstrate that real-world auditory attention decoding remains a challenging goal. Taken together, our results highlight the strengths and weaknesses of using deep learning for auditory EEG decoding, and highlight important topics for future research.
We first employed non-linear models based on deep neural networks for the purpose of reconstructing the speech envelope from the listeners’ EEG measurements. These decoders performed significantly better than their linear counterparts, and could also function as auditory attention decoders. However, they did not generalise well between different listeners, a property which would be valuable in translational applications. We then turned to a different decoding problem, the match-mismatch problem, which was part of the ICASSP 2023/2024 Auditory EEG Decoding Signal Processing Grand Challenge series. Inspired by known responses to speech – envelope tracking and speech-related frequency-following responses – we developed highly accurate, challenge-winning decoders, which this time generalised well between participants and even across datasets. Finally, we applied some of these techniques in an auditory attention identification setting using EEG signals recorded from a wearable ear-EEG device. Whilst auditory attention could be identified with significance using this device, the results demonstrate that real-world auditory attention decoding remains a challenging goal. Taken together, our results highlight the strengths and weaknesses of using deep learning for auditory EEG decoding, and highlight important topics for future research.
Version
Open Access
Date Issued
2024-10-08
Date Awarded
01/04/2025
License URL
Advisor
Mandic, Danilo
Reichenbach, Tobias
Sponsor
UK Research and Innovation
Grant Number
P/S023283/1
Publisher Department
Department of Computing
Department of Electrical and Electronic Engineering
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
