Representation Balancing MDPs for Off-Policy Policy Evaluation
File(s)representation_balancing_model_for_OPPE.pdf (620.81 KB)
Submitted version
Author(s)
Type
Conference Paper
Abstract
We study the problem of off-policy policy evaluation (OPPE) in RL. In contrast
to prior work, we consider how to estimate both the individual policy value and average policy value accurately. We draw inspiration from recent work in causal reasoning, and propose a new finite sample generalization error bound for value estimates from MDP models. Using this upper bound as an objective, we develop a learning algorithm of an MDP model with a balanced representation, and show that our approach can yield substantially lower MSE in a common synthetic domain and on a challenging real-world sepsis management problem.
to prior work, we consider how to estimate both the individual policy value and average policy value accurately. We draw inspiration from recent work in causal reasoning, and propose a new finite sample generalization error bound for value estimates from MDP models. Using this upper bound as an objective, we develop a learning algorithm of an MDP model with a balanced representation, and show that our approach can yield substantially lower MSE in a common synthetic domain and on a challenging real-world sepsis management problem.
Date Issued
2018-12-03
Date Acceptance
2018-09-17
Citation
2018
Copyright Statement
© 2018 The Author(s)
Identifier
https://nips.cc/
Source
Thirty-second Annual Conference on Neural Information Processing Systems (NIPS)
Start Date
2018-12-03
Finish Date
2018-12-08
Coverage Spatial
Montreal
Date Publish Online
2018-12-03