Enabling risk-aware Reinforcement Learning for medical interventions through uncertainty decomposition
File(s) 2109.07827.pdf (1.47 MB)
Accepted version
Author(s)
Festor, Paul
Luise, Giulia
Komorowski, Matthieu
Faisal, Aldo
Type
Conference Paper
Abstract
Reinforcement Learning (RL) is emerging as tool
for tackling complex control and decision-making
problems. However, in high-risk environments
such as healthcare, manufacturing, automotive or
aerospace, it is often challenging to bridge the gap
between an apparently optimal policy learned by
an agent and its real-world deployment, due to the
uncertainties and risk associated with it. Broadly
speaking RL agents face two kinds of uncertainty,
1. aleatoric uncertainty, which reflects randomness or noise in the dynamics of the world, and 2.
epistemic uncertainty, which reflects the bounded
knowledge of the agent due to model limitations
and finite amount of information/data the agent
has acquired about the world. These two types
of uncertainty carry fundamentally different implications for the evaluation of performance and
the level of risk or trust. Yet these aleatoric and
epistemic uncertainties are generally confounded
as standard and even distributional RL is agnostic
to this difference. Here we propose how a distributional approach (UA-DQN) can be recast to
render uncertainties by decomposing the net effects of each uncertainty . We demonstrate the
operation of this method in grid world examples
to build intuition and then show a proof of concept application for an RL agent operating as a
clinical decision support system in critical care.
for tackling complex control and decision-making
problems. However, in high-risk environments
such as healthcare, manufacturing, automotive or
aerospace, it is often challenging to bridge the gap
between an apparently optimal policy learned by
an agent and its real-world deployment, due to the
uncertainties and risk associated with it. Broadly
speaking RL agents face two kinds of uncertainty,
1. aleatoric uncertainty, which reflects randomness or noise in the dynamics of the world, and 2.
epistemic uncertainty, which reflects the bounded
knowledge of the agent due to model limitations
and finite amount of information/data the agent
has acquired about the world. These two types
of uncertainty carry fundamentally different implications for the evaluation of performance and
the level of risk or trust. Yet these aleatoric and
epistemic uncertainties are generally confounded
as standard and even distributional RL is agnostic
to this difference. Here we propose how a distributional approach (UA-DQN) can be recast to
render uncertainties by decomposing the net effects of each uncertainty . We demonstrate the
operation of this method in grid world examples
to build intuition and then show a proof of concept application for an RL agent operating as a
clinical decision support system in critical care.
Date Issued
2021-07-23
Date Acceptance
2021-07-03
Citation
2021
Copyright Statement
© The Author(s).
Source
ICML2021 workshop on Interpretable Machine Learning in Healthcare
Publication Status
Published
Start Date
2021-07-23
Finish Date
2021-07-23
Coverage Spatial
Online
