A study of salient modulation domain features for speaker identification
File(s)210913_APSIPA2021_Paper_Compressed.pdf (1.06 MB)
Published version
Author(s)
McKnight, simon
Hogg, Aidan
Neo, Vincent
Naylor, patrick
Type
Conference Paper
Abstract
This paper studies the ranges of acoustic andmodulation frequencies of speech most relevant for identifyingspeakers and compares the speaker-specific information presentin the temporal envelope against that present in the temporalfine structure. This study uses correlation and feature importancemeasures, random forest and convolutional neural network mod-els, and reconstructed speech signals with specific acoustic and/ormodulation frequencies removed to identify the salient points. Itis shown that the range of modulation frequencies associated withthe fundamental frequency is more important than the 1-16 Hzrange most commonly used in automatic speech recognition, andthat the 0 Hz modulation frequency band contains significantspeaker information. It is also shown that the temporal envelopeis more discriminative among speakers than the temporal finestructure, but that the temporal fine structure still contains usefuladditional information for speaker identification. This researchaims to provide a timely addition to the literature by identifyingspecific aspects of speech relevant for speaker identification thatcould be used to enhance the discriminant capabilities of machinelearning models.
Date Issued
2022-02-03
Date Acceptance
2021-08-31
Citation
2022, pp.705-712
Publisher
IEEE
Start Page
705
End Page
712
Copyright Statement
© 2022 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.
Identifier
https://ieeexplore.ieee.org/abstract/document/9689293
Source
Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC)
Subjects
Science & Technology
Technology
Computer Science, Information Systems
Computer Science, Software Engineering
Engineering, Electrical & Electronic
Computer Science
Engineering
TEMPORAL ENVELOPE
SPEECH RECOGNITION
NORMAL-HEARING
FREQUENCY
REPRESENTATION
PERCEPTION
Publication Status
Published
Start Date
2021-12-14
Finish Date
2021-12-17
Coverage Spatial
Tokyo, Japan
Date Publish Online
2022-02-03