Gaussian Process Dynamic Programming
File(s)neurocomputing2009.pdf (1.66 MB)
Accepted version
Author(s)
Deisenroth, Marc P
Rasmussen, Carl E
Peters, Jan
Type
Journal Article
Abstract
Reinforcement learning (RL) and optimal control of systems with continuous states and actions require approximation techniques in most interesting cases. In this article, we introduce Gaussian process dynamic programming (GPDP), an approximate value-function based RL algorithm. We consider both a classic optimal control problem, where problem-specific prior knowledge is available, and a classic RL problem, where only very general priors can be used. For the classic optimal control problem, GPDP models the unknown value functions with Gaussian processes and generalizes dynamic programming to continuous-valued states and actions. For the RL problem, GPDP starts from a given initial state and explores the state space using Bayesian active learning. To design a fast learner, available data has to be used efficiently. Hence, we propose to learn probabilistic models of the a priori unknown transition dynamics and the value functions on the fly. In both cases, we successfully apply the resulting continuous-valued controllers to the under-actuated pendulum swing up and analyze the performances of the suggested algorithms. It turns out that GPDP uses data very efficiently and can be applied to problems, where classic dynamic programming would be cumbersome.
Date Issued
2009-03
Citation
Neurocomputing, 2009, 72 (7-9), pp.1508-1524
ISSN
0925-2312
Publisher
Elsevier
Start Page
1508
End Page
1524
Journal / Book Title
Neurocomputing
Volume
72
Issue
7-9
Copyright Statement
© 2009 Elsevier Ltd All rights reserved. NOTICE: this is the author’s version of a work that was accepted for publication in Neurocomputing. Changes resulting from the publishing process, such as peer review, editing, corrections, structural formatting, and other quality control mechanisms may not be reflected in this document. Changes may have been made to this work since it was submitted for publication. A definitive version was subsequently published in NEUROCOMPUTING, Vol.:72, Issue:7-9, (2009)] DOI: http://dx.doi.org/10.1016/j.neucom.2008.12.019.
Description
25.06.13 KB. Ok to add accepted version to spiral, Elsever say ok while mandate not enforced.
Identifier
http://mlg.eng.cam.ac.uk/marc/publications/neurocomputing2009_preprint.pdf
Notes
citeseerurl: http://www.science-direct.com/science?_ob=MImg&_imagekey=B6V10-4VC0YDJ-2-11&_cdi=5660&_user=1495569&_orig=browse&_coverDate=03%2F31%2F2009&_sk=999279992&view=c&wchp=dGLbVlW-zSkWb&md5=e6894442349ef0ff2bd4026d3d620d6c&ie=/sdarticle.pdf timestamp: 2008.06.16