Prediction-based audiovisual fusion for classification of non-linguistic vocalisations
File(s) predictionbasedavfusion_petridispantic.pdf (11.53 MB)
Accepted version
Author(s)
Petridis, S
Pantic, M
Type
Journal Article
Abstract
Prediction plays a key role in recent computational models of the brain and it has been suggested that the brain constantly makes multisensory spatiotemporal predictions. Inspired by these findings we tackle the problem of audiovisual fusion from a new perspective based on prediction. We train predictive models which model the spatiotemporal relationship between audio and visual features by learning the audio-to-visual and visual-to-audio feature mapping for each class. Similarly, we train predictive models which model the time evolution of audio and visual features by learning the past-to-future feature mapping for each class. In classification, all the class-specific regression models produce a prediction of the expected audio/visual features and their prediction errors are combined for each class. The set of class-specific regressors which best describes the audiovisual feature relationship, i.e., results in the lowest prediction error, is chosen to label the input frame. We perform cross-database experiments, using the AMI, SAL, and MAHNOB databases, in order to classify laughter and speech and subject-independent experiments on the AVIC database in order to classify laughter, hesitation and consent. In virtually all cases prediction-based audiovisual fusion consistently outperforms the two most commonly used fusion approaches, decision-level and feature-level fusion.
Date Issued
2016-01
Date Acceptance
2015-06-03
Citation
IEEE Transactions on Affective Computing, 2016, 7 (1), pp.45-58
ISSN
1949-3045
Publisher
Institute of Electrical and Electronics Engineers (IEEE)
Start Page
45
End Page
58
Journal / Book Title
IEEE Transactions on Affective Computing
Volume
7
Issue
1
Copyright Statement
© 2015 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.
Sponsor
Commission of the European Communities
Commission of the European Communities
Grant Number
611153
688520
Subjects
Science & Technology
Technology
Computer Science, Artificial Intelligence
Computer Science, Cybernetics
Computer Science
Prediction-based fusion
audiovisual fusion
nonlinguistic vocalisation classification
MODULAR NEURAL-NETWORKS
SPEECH RECOGNITION
FACIAL ANIMATION
DRIVEN
IDENTIFICATION
EXPRESSIONS
LAUGHTER
DATABASE
HEAR
HMM
Publication Status
Published
Date Publish Online
2015-06-16
