Deep learning for infant atypical speech analysis
File(s)
Author(s)
Al Futaisi, Najla
Type
Thesis
Abstract
Advancements in speech research increasingly focus on automating processes using Artificial Intelligence (AI), moving away from traditional, labor-intensive methods. This shift facilitates the analysis of atypical speech and language environments in children, aiding the identification of speech disorders and intellectual disabilities. However, the performance of Machine Learning (ML) models heavily depends on the quality of training datasets, making data collection and sharing significant challenges. This research employs Deep Learning (DL) algorithms, including Recurrent Neural Networks (RNN) and transformer networks, to model and classify speech data.
The thesis begins by classifying infant vocalisations, distinguishing between canonical and non- canonical babbling using an Adversarial Multi-Task Learning approach that leverages weakly labeled data. This method effectively addresses category mismatch and data sparsity. Further research explores auditory stimuli exposure, differentiating between infant- and adult-directed speech through Multi-Task Learning and AutoEncoder techniques. Experiments demonstrate that knowledge transfer across datasets improves model performance.
A key focus is the early detection of Autism Spectrum Disorder (ASD) through speech recogni- tion, emphasising the importance of early intervention. Two fine-tuning approaches for transfer learning are evaluated for their ability to classify child speech as Atypically Developing (AT) or Typically Developing (TD), a task known as typicality classification. The models—one using discriminative fine-tuning, the other fine-tuning a Wav2Vec 2.0 model—are also applied to a diagnostic task distinguishing between Pervasive Developmental Disorders (PDD), including ASD, and typical development using 62 minutes of labeled speech data. The findings confirm that deep transfer learning improves classification accuracy but highlight challenges such as category mismatches and class imbalances.
Ethical considerations, especially regarding child data and fairness in AI, are examined. The thesis explores gender bias in ASD classification models, revealing that ML models can reflect biases in training data, underscoring the need for fairness in AI development, particularly for sensitive conditions like ASD.
The thesis begins by classifying infant vocalisations, distinguishing between canonical and non- canonical babbling using an Adversarial Multi-Task Learning approach that leverages weakly labeled data. This method effectively addresses category mismatch and data sparsity. Further research explores auditory stimuli exposure, differentiating between infant- and adult-directed speech through Multi-Task Learning and AutoEncoder techniques. Experiments demonstrate that knowledge transfer across datasets improves model performance.
A key focus is the early detection of Autism Spectrum Disorder (ASD) through speech recogni- tion, emphasising the importance of early intervention. Two fine-tuning approaches for transfer learning are evaluated for their ability to classify child speech as Atypically Developing (AT) or Typically Developing (TD), a task known as typicality classification. The models—one using discriminative fine-tuning, the other fine-tuning a Wav2Vec 2.0 model—are also applied to a diagnostic task distinguishing between Pervasive Developmental Disorders (PDD), including ASD, and typical development using 62 minutes of labeled speech data. The findings confirm that deep transfer learning improves classification accuracy but highlight challenges such as category mismatches and class imbalances.
Ethical considerations, especially regarding child data and fairness in AI, are examined. The thesis explores gender bias in ASD classification models, revealing that ML models can reflect biases in training data, underscoring the need for fairness in AI development, particularly for sensitive conditions like ASD.
Version
Open Access
Date Issued
2023-08
Date Awarded
2024-10
Copyright Statement
Creative Commons Attribution NonCommercial Licence
License URL
Advisor
Schuller, Björn
Sponsor
Oman, Ministry of Higher Education, Research and Innovation
Publisher Department
Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)