Emotion filtering at the edge
File(s)1909.08500v1.pdf (3.6 MB)
Working paper
Author(s)
Aloufi, Ranya
Haddadi, Hamed
Boyle, David
Type
Working Paper
Abstract
Voice controlled devices and services have become very popular in the
consumer IoT. Cloud-based speech analysis services extract information from
voice inputs using speech recognition techniques. Services providers can thus
build very accurate profiles of users' demographic categories, personal
preferences, emotional states, etc., and may therefore significantly compromise
their privacy. To address this problem, we have developed a privacy-preserving
intermediate layer between users and cloud services to sanitize voice input
directly at edge devices. We use CycleGAN-based speech conversion to remove
sensitive information from raw voice input signals before regenerating
neutralized signals for forwarding. We implement and evaluate our emotion
filtering approach using a relatively cheap Raspberry Pi 4, and show that
performance accuracy is not compromised at the edge. In fact, signals generated
at the edge differ only slightly (~0.16%) from cloud-based approaches for
speech recognition. Experimental evaluation of generated signals show that
identification of the emotional state of a speaker can be reduced by ~91%.
consumer IoT. Cloud-based speech analysis services extract information from
voice inputs using speech recognition techniques. Services providers can thus
build very accurate profiles of users' demographic categories, personal
preferences, emotional states, etc., and may therefore significantly compromise
their privacy. To address this problem, we have developed a privacy-preserving
intermediate layer between users and cloud services to sanitize voice input
directly at edge devices. We use CycleGAN-based speech conversion to remove
sensitive information from raw voice input signals before regenerating
neutralized signals for forwarding. We implement and evaluate our emotion
filtering approach using a relatively cheap Raspberry Pi 4, and show that
performance accuracy is not compromised at the edge. In fact, signals generated
at the edge differ only slightly (~0.16%) from cloud-based approaches for
speech recognition. Experimental evaluation of generated signals show that
identification of the emotional state of a speaker can be reduced by ~91%.
Date Issued
2019-09-18
Citation
2019
Publisher
arXiv
Copyright Statement
© 2019 The Author(s)
Identifier
http://arxiv.org/abs/1909.08500v1
Subjects
eess.AS
eess.AS
cs.CR
cs.HC
cs.SD
Notes
6 pages, 6 figures, Sensys-ML19 workshop in conjunction with the 17th ACM Conference on Embedded Networked Sensor Systems (SenSys 2019)
Publication Status
Published