When whisper listens to aphasia: advancing robust post-stroke speech recognition
File(s)SANGUEDOLCE_Interspeech.pdf (1.14 MB)
Published version
Author(s)
Sanguedolce, Giulia
Brook, Sophie
Gruia, Dragos
Naylor, Patrick
Geranmayeh, Fatemeh
Type
Conference Paper
Abstract
Despite recent advancements in Automatic Speech Recogni tion (ASR), its accuracy remains low for pathological speech,
thereby limiting AI-based healthcare interventions in such set tings. This work addresses this challenge by fine-tuning Whis per, an ASR known for its ability to capture high-dimensional
features in healthy speech. Using our comprehensive dataset
of patients with stroke, we fine-tuned Whisper and significantly
reduced Word Error Rate (WER), surpassing previous work on
severe aphasia. To demonstrate its generalisability, we tested
the model on a separate database, AphasiaBank, and observed
a lower WER despite variations in dialect, linguistics, and test
protocols. Our result on the AphasiaBank was superior to pre vious ASRs trained on this database, confirming the generalis ability of our approach. These outcomes not only address ASR
limitations in impaired speech but also establish the foundations
for standardised and versatile AI solutions for remote speech
monitoring for timely diagnosis and intervention.
Index Terms: Speech Recognition, Fine-tuning, Pathological
Speech
thereby limiting AI-based healthcare interventions in such set tings. This work addresses this challenge by fine-tuning Whis per, an ASR known for its ability to capture high-dimensional
features in healthy speech. Using our comprehensive dataset
of patients with stroke, we fine-tuned Whisper and significantly
reduced Word Error Rate (WER), surpassing previous work on
severe aphasia. To demonstrate its generalisability, we tested
the model on a separate database, AphasiaBank, and observed
a lower WER despite variations in dialect, linguistics, and test
protocols. Our result on the AphasiaBank was superior to pre vious ASRs trained on this database, confirming the generalis ability of our approach. These outcomes not only address ASR
limitations in impaired speech but also establish the foundations
for standardised and versatile AI solutions for remote speech
monitoring for timely diagnosis and intervention.
Index Terms: Speech Recognition, Fine-tuning, Pathological
Speech
Date Issued
2024-09-01
Date Acceptance
2024-06-04
Citation
Interspeech 2024, 2024, pp.1995-1999
ISSN
2958-1796
Publisher
ISCA
Start Page
1995
End Page
1999
Journal / Book Title
Interspeech 2024
Copyright Statement
© The Author(s).
Identifier
http://dx.doi.org/10.21437/interspeech.2024-2183
Source
Interspeech 2024
Publication Status
Published
Start Date
2024-09-01
Finish Date
2024-09-05
Coverage Spatial
Kos, Greece
Date Publish Online
2024-09-01