Do androids dream of electric fences? Safety-aware reinforcement learning with latent shielding
File(s) paper_50.pdf (847.38 KB)
Published version
Author(s)
He, Chloe
Gonzalez Leon, Borja
Belardinelli, Francesco
Type
Conference Paper
Abstract
The growing trend of fledgling reinforcement learning sys-
tems making their way into real-world applications has been
accompanied by growing concerns for their safety and ro-
bustness. In recent years, a variety of approaches have been
put forward to address the challenges of safety-aware rein-
forcement learning; however, these methods often either re-
quire a handcrafted model of the environment to be pro-
vided beforehand, or that the environment is relatively simple
and low-dimensional. We present a novel approach to safety-
aware deep reinforcement learning in high-dimensional envi-
ronments called latent shielding. Latent shielding leverages
internal representations of the environment learnt by model-
based agents to “imagine” future trajectories and avoid those
deemed unsafe. We experimentally demonstrate that this
approach leads to improved adherence to formally-defined
safety specifications.
tems making their way into real-world applications has been
accompanied by growing concerns for their safety and ro-
bustness. In recent years, a variety of approaches have been
put forward to address the challenges of safety-aware rein-
forcement learning; however, these methods often either re-
quire a handcrafted model of the environment to be pro-
vided beforehand, or that the environment is relatively simple
and low-dimensional. We present a novel approach to safety-
aware deep reinforcement learning in high-dimensional envi-
ronments called latent shielding. Latent shielding leverages
internal representations of the environment learnt by model-
based agents to “imagine” future trajectories and avoid those
deemed unsafe. We experimentally demonstrate that this
approach leads to improved adherence to formally-defined
safety specifications.
Date Acceptance
2021-12-08
Citation
CEUR Workshop Proceedings, 3087, pp.1-19
ISSN
1613-0073
Publisher
CEUR Workshop Proceedings
Start Page
1
End Page
19
Journal / Book Title
CEUR Workshop Proceedings
Volume
3087
Copyright Statement
© 2022 for this paper by its authors. Use permitted under
Creative Commons License Attribution 4.0 International (CC BY 4.0) https://creativecommons.org/licenses/by/4.0/
Creative Commons License Attribution 4.0 International (CC BY 4.0) https://creativecommons.org/licenses/by/4.0/
License URL
Source
Workshop on Artificial Intelligence Safety 2022 (SafeAI 2022)
Publication Status
Published
Start Date
2022-02-28
Finish Date
2022-03-01
Coverage Spatial
Virtual
