Combining Local and Global Direct Derivative-free Optimization for Reinforcement Learning
File(s) cait-2012-0021 (1).pdf (1003.67 KB)
Published version
Author(s)
Leonetti, Matteo
Kormushev, Petar
Sagratella, Simone
Type
Journal Article
Abstract
We consider the problem of optimization in policy space for reinforcement learning. While a plethora of methods have been applied to this problem, only a narrow category of them proved feasible in robotics. We consider the peculiar characteristics of reinforcement learning in robotics, and devise a combination of two algorithms from the literature of derivative-free optimization. The proposed combination is well suited for robotics, as it involves both off-line learning in simulation and on-line learning in the real environment. We demonstrate our approach on a real-world task, where an Autonomous Underwater Vehicle has to survey a target area under potentially unknown environment conditions. We start from a given controller, which can perform the task under foreseeable conditions, and make it adaptive to the actual environment.
Date Issued
2012
Date Acceptance
2013-03-22
Citation
International Journal of Cybernetics and Information Technologies, 2012, 12
ISSN
1311-9702
Publisher
De Gruyter
Start Page
53
End Page
65
Journal / Book Title
International Journal of Cybernetics and Information Technologies
Volume
12
Issue
3
Copyright Statement
© 2013 The Authors. Creative Commons Attribution Non-Commercial No Derivatives License
Identifier
http://kormushev.com/papers/Leonetti_CIT-2012.pdf
Subjects
Reinforcement learning
policy search
derivative-free optimization
robotics
autonomous underwater vehicles
Publication Status
Published
Publisher URL
Article Number
3
