Approximate Dynamic Programming with Gaussian Processes
File(s)acc2008_final.pdf (390.1 KB)
Accepted version
Author(s)
Deisenroth, Marc P
Peters, Jan
Rasmussen, Carl E
Type
Conference Paper
Abstract
In general, it is difficult to determine an optimal closed-loop policy in nonlinear control problems with continuous-valued state and control domains. Hence, approximations are often inevitable. The standard method of discretizing states and controls suffers from the curse of dimensionality and strongly depends on the chosen temporal sampling rate. In this paper, we introduce Gaussian process dynamic programming (GPDP) and determine an approximate globally optimal closed-loop policy. In GPDP, value functions in the Bellman recursion of the dynamic programming algorithm are modeled using Gaussian processes. GPDP returns an optimal state-feedback for a finite set of states. Based on these outcomes, we learn a possibly discontinuous closed-loop policy on the entire state space by switching between two independently trained Gaussian processes. A binary classifier selects one Gaussian process to predict the optimal control signal. We show that GPDP is able to yield an almost optimal solution to an LQ problem using few sample points. Moreover, we successfully apply GPDP to the underpowered pendulum swing up, a complex nonlinear control problem.
Date Issued
2008-06
Citation
Proceedings of the 2008 American Control Conference (ACC 2008), 2008, pp.4480-4485
ISBN
978-1-4244-2078-0
ISSN
0743-1619
Start Page
4480
End Page
4485
Journal / Book Title
Proceedings of the 2008 American Control Conference (ACC 2008)
Copyright Statement
© 2008 AACC
Description
22.10.13 KB. Ok to add the author version to spiral.
Source
American Control Conference 2008
Notes
timestamp: 2007.09.12
Place of Publication
Seattle, WA, USA
Start Date
2008-06-11
Finish Date
2008-06-13
Coverage Spatial
Seatle, WA