Local differential privacy for regret minimization in reinforcement learning
File(s)local_privacy_RL.pdf (2.13 MB)
Accepted version
Author(s)
Garcelon, Evrard
Perchet, Vianney
Pike-Burke, Ciara
Pirotta, Matteo
Type
Conference Paper
Abstract
Reinforcement learning algorithms are widely used in domains where it is desirable
to provide a personalized service. In these domains it is common that user data
contains sensitive information that needs to be protected from third parties. Motivated by this, we study privacy in the context of finite-horizon Markov Decision
Processes (MDPs) by requiring information to be obfuscated on the user side.
We formulate this notion of privacy for RL by leveraging the local differential
privacy (LDP) framework. We establish a lower bound for regret minimization in
finite-horizon MDPs with LDP guarantees which shows that guaranteeing privacy
has a multiplicative effect on the regret. This result shows that while LDP is
an appealing notion of privacy, it makes the learning problem significantly more
complex. Finally, we present an optimistic algorithm that simultaneously satisfies
ε-LDP requirements, and achieves √
K/ε regret in any finite-horizon MDP after
K episodes, matching the lower bound dependency on the number of episodes K.
to provide a personalized service. In these domains it is common that user data
contains sensitive information that needs to be protected from third parties. Motivated by this, we study privacy in the context of finite-horizon Markov Decision
Processes (MDPs) by requiring information to be obfuscated on the user side.
We formulate this notion of privacy for RL by leveraging the local differential
privacy (LDP) framework. We establish a lower bound for regret minimization in
finite-horizon MDPs with LDP guarantees which shows that guaranteeing privacy
has a multiplicative effect on the regret. This result shows that while LDP is
an appealing notion of privacy, it makes the learning problem significantly more
complex. Finally, we present an optimistic algorithm that simultaneously satisfies
ε-LDP requirements, and achieves √
K/ε regret in any finite-horizon MDP after
K episodes, matching the lower bound dependency on the number of episodes K.
Date Issued
2021-12-07
Date Acceptance
2021-09-28
Citation
Advances in neural information processing systems, 2021, pp.1-13
ISSN
1049-5258
Publisher
NeurIPS
Start Page
1
End Page
13
Journal / Book Title
Advances in neural information processing systems
Copyright Statement
© 2021 The Author(s)
Identifier
https://proceedings.neurips.cc/paper/2021/hash/580760fb5def6e2ca8eaf601236d5b08-Abstract.html
Source
Advances in Neural Information Processing Systems
Subjects
1701 Psychology
1702 Cognitive Sciences
Publication Status
Published
Start Date
2021-12-07
Finish Date
2021-12-14
Coverage Spatial
Virtual-only Conference
Date Publish Online
2021-12-07