WhisperD: Dementia speech recognition and filler word detection with whisper
File(s) 2505.21551v1.pdf (124.64 KB)
Accepted version
Author(s)
Akinrintoyo, Emmanuel
Abdelhalim, Nadine
Salomons, Nicole
Type
Conference Paper
Abstract
Whisper fails to correctly transcribe dementia speech because persons with dementia (PwDs) often exhibit irregular speech patterns and disfluencies such as pauses, repetitions, and fragmented sentences. It was trained on standard speech and may have had little or no exposure to dementia-affected speech. However, correct transcription is vital for dementia speech for cost-effective diagnosis and the development of assistive technology. In this work, we fine-tune Whisper with the open-source dementia speech dataset (DementiaBank) and our in-house dataset to improve its word error rate (WER). The fine-tuning also includes filler words to ascertain the filler inclusion rate (FIR) and F1 score. The fine-tuned models significantly outperformed the off-the-shelf models. The medium-sized model achieved a WER of 0.24, outperforming previous work. Similarly, there was a notable generalisability to unseen data and speech patterns.
Date Issued
2025-08-17
Date Acceptance
2025-08-01
Citation
Proceedings of Interspeech 2025, 2025, pp.1413-1417
ISSN
2958-1796
Publisher
International Speech Communication Association (ISCA)
Start Page
1413
End Page
1417
Journal / Book Title
Proceedings of Interspeech 2025
Copyright Statement
© 2025 ISCA.
Source
Interspeech 2025
Subjects
Dementia
Whisper
ASR
Filler words
Speechto-text
Speech recognition
Disfluency
Fine-tuning
Publication Status
Published
Start Date
2025-08-17
Finish Date
2025-08-21
Coverage Spatial
Rotterdam, The Netherlands
Date Publish Online
2025-08-17
