Variational inference for data-efficient model learning in POMDPs
File(s)
Author(s)
Tschiatschek, Sebastian
Arulkumaran, Kai
Stühmer, Jan
Hofmann, Katja
Type
Working Paper
Abstract
Partially observable Markov decision processes (POMDPs) are a powerful
abstraction for tasks that require decision making under uncertainty, and
capture a wide range of real world tasks. Today, effective planning approaches
exist that generate effective strategies given black-box models of a POMDP
task. Yet, an open question is how to acquire accurate models for complex
domains. In this paper we propose DELIP, an approach to model learning for
POMDPs that utilizes amortized structured variational inference. We empirically
show that our model leads to effective control strategies when coupled with
state-of-the-art planners. Intuitively, model-based approaches should be
particularly beneficial in environments with changing reward structures, or
where rewards are initially unknown. Our experiments confirm that DELIP is
particularly effective in this setting.
abstraction for tasks that require decision making under uncertainty, and
capture a wide range of real world tasks. Today, effective planning approaches
exist that generate effective strategies given black-box models of a POMDP
task. Yet, an open question is how to acquire accurate models for complex
domains. In this paper we propose DELIP, an approach to model learning for
POMDPs that utilizes amortized structured variational inference. We empirically
show that our model leads to effective control strategies when coupled with
state-of-the-art planners. Intuitively, model-based approaches should be
particularly beneficial in environments with changing reward structures, or
where rewards are initially unknown. Our experiments confirm that DELIP is
particularly effective in this setting.
Date Issued
2018-05-23
Citation
2018
Identifier
http://arxiv.org/abs/1805.09281v1
Subjects
stat.ML
stat.ML
cs.LG
