Model-Based Reinforcement Learning with Continuous States and Actions
File(s)esann2008_errata.pdf (90.58 KB) es2008-8.final.pdf (817.43 KB)
Supporting information
Published version
Author(s)
Deisenroth, Marc P
Rasmussen, Carl E
Peters, Jan
Type
Conference Paper
Abstract
Finding an optimal policy in a reinforcement learning (RL) framework with continuous state and action spaces is challenging. Approximate solutions are often inevitable. GPDP is an approximate dynamic programming algorithm based on Gaussian process (GP) models for the value functions. In this paper, we extend GPDP to the case of unknown transition dynamics. After building a GP model for the transition dynamics, we apply GPDP to this model and determine a continuous-valued policy in the entire state space. We apply the resulting controller to the underpowered pendulum swing up. Moreover, we compare our results on this RL task to a nearly optimal discrete DP solution in a fully known environment.
Date Issued
2008-04
Citation
Proceedings of the 16th European Symposium on Artificial Neural Networks (ESANN 2008), 2008, pp.19-24
ISBN
2-930307-08-0
Start Page
19
End Page
24
Journal / Book Title
Proceedings of the 16th European Symposium on Artificial Neural Networks (ESANN 2008)
Copyright Statement
© 2008 ESANN
Description
22.10.13 KB. Ok to add the published version to spiral. ESANN
Source
ESANN 2008
Notes
review: Comments from reviewer 1 I think Gaussian processes in RL are a very interesting topic. The description of the algorithm is a bit sloppy. Are tilde X and tilde U really sets (i.e., unordered and not containing double elements) or rather sequences? Are they based on a single episode or several episodes (or are the samples not based on episodes)? Let’s presume the index of x_k refers to a timestep. This would imply that there are several x_k in the set/sequence? If x_y is unique, lines 5 and 6 make little sense (for a fixed k: for all x_k, for all u_k). If there are more than one episode in tilde X and tilde U (otherwise I do not understand lines 5 and 6), why is there not an "for all x_N in tilde X" before line 2, i.e. why is there only one terminal state? The notation "...simmathcal GP.." - adaptation of the GP parameters based on the assigned data points - has to be introduced. I am scared by the scaling of the algorithm. From the training set, a model for f (also an GP?), a GP for V, two GPs for pi and a GP for Q(x,.) for every x in tilde X is estimated. Doesn’t the data "wear out"? Overfitting? Where do the priors come from? The empirical evaluation is not convincing, but ESANN has strict space restrictions. I appreciate the comparison with the solution obtained by dynamic programming. Typos: p.1,2nd para.: "with in" -> "within" p.3 "Thatn" Alg.1: "(.)" missing in lines 3 and 13? Comments from reviewer 2 Well written. Comments from reviewer 3 The paper extends the use of Gaussian processes in RL to the case of unknown system dynamics and uses GP consequently for all modeled elements, which is in my opinion a methodically very appealing approach. The drawback of the approach seems to be that many GP are needed to be trained by the whole algorithm, which leads to high computational costs. The experiment are described in a very appropriate way, and are detailed enough to easily allow to repeat the experiments independently. The own results are compared to a standard DP (using the true model) as ground truth. The algorithm detailed in the paper is confusing, it is not clear whether the index k counts observations from a single episodes or runs through several consecutive episodes. page 1, first sentence of second paragraph: "Model-based policy iteration with in continuous state and action spaces based on value function evaluation using GPs is presented in [3]." "with" is unnecessary "in reinforcement learning system dynamics are often unknown" - even tough naturally an important statement for the paper it is repeated a bit often. page 3, last sentence in "Example: Learning the dynamics": "... and the mean of the model f is smaller thatn 0.04" wants to be "... and the mean of the model f is smaller than 0.04" page 5, first full sentence: "Thus, we use of the SE covariance function"- I think something is missing here. page 6, Figure2" Unfortunately the final version of the paper will have to be black and white. timestamp: 2007.11.14
Place of Publication
Bruges, Belgium
Start Date
2008-04-23
Finish Date
2008-04-25
Coverage Spatial
Bruges, Belgium