Continuous Estimation of Emotions in Speech by Dynamic Cooperative Speaker Models
File(s)07412670.pdf (1.35 MB)
Accepted version
Author(s)
Mencattini, A
Martinelli, E
Ringeval, F
Schuller, B
Natale, CD
Type
Journal Article
Abstract
Automatic emotion recognition from speech has been recently focused on the prediction of time-continuous dimensions (e.g., arousal and valence) of spontaneous and realistic expressions of emotion, as found in real-life interactions. However, the automatic prediction of such emotions poses several challenges, such as the subjectivity found in the definition of a gold standard from a pool of raters and the issue of data scarcity in training models. In this work, we introduce a novel emotion recognition system, based on ensemble of single-speaker-regression-models (SSRMs). The estimation of emotion is provided by combining a subset of the initial pool of SSRMs selecting those that are most concordance among them. The proposed approach allows the addition or removal of speakers from the ensemble without the necessity to re-build the entire machine learning system. The simplicity of this aggregation strategy, coupled with the flexibility assured by the modular architecture, and the promising results obtained on the RECOLA database highlight the potential implications of the proposed method in a real-life scenario and in particular in WEB-based applications.
Date Issued
2016-02-18
Date Acceptance
2016-02-18
Citation
IEEE Transactions on Affective Computing, 2016, 8 (3), pp.314-327
ISSN
1949-3045
Publisher
Institute of Electrical and Electronics Engineers (IEEE)
Start Page
314
End Page
327
Journal / Book Title
IEEE Transactions on Affective Computing
Volume
8
Issue
3
Copyright Statement
© 2016 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.
Notes
14 pages, to appear (IF: 3.466, 5-year IF: 3.871 (2013))
Publication Status
Published