Model-based policy iterations for nonlinear systems via controlled Hamiltonian dynamics
Author(s)
Sassano, Mario
Mylvaganam, Thulasi
Astolfi, Alessandro
Type
Journal Article
Abstract
The infinite-horizon optimal control problem for nonlinear systems is studied. In the context of model-based, iterative learning strategies we propose an alternative definition and construction of the temporal difference error arising in Policy Iteration strategies. In such architectures the error is computed via the evolution of the Hamiltonian function (or, possibly, of its integral) along the trajectories of the closed-loop system. Herein the temporal difference error is instead obtained via two subsequent steps: first the dynamics of the underlying costate variable in the Hamiltonian system is steered by means of a (virtual) control input in such a way that the stable invariant manifold becomes externally attractive. Then, the distance-from-invariance of the manifold, induced by approximate solutions, yields a natural candidate measure for the policy evaluation step. The policy improvement phase is then performed by means of standard gradient descent methods
that allows to correctly update the weights of the underlying functional approximator. The above architecture then yields an iterative (episodic) learning scheme based on a scalar, constant reward at each iteration, the value of which is insensitive to the length of the episode, as in the original
spirit of Reinforcement Learning strategies for discrete-time systems. Finally, the theory is validated by means of a numerical simulation involving an automatic flight control problem.
that allows to correctly update the weights of the underlying functional approximator. The above architecture then yields an iterative (episodic) learning scheme based on a scalar, constant reward at each iteration, the value of which is insensitive to the length of the episode, as in the original
spirit of Reinforcement Learning strategies for discrete-time systems. Finally, the theory is validated by means of a numerical simulation involving an automatic flight control problem.
Date Issued
2023-05-01
Date Acceptance
2022-05-13
Citation
IEEE Transactions on Automatic Control, 2023, 68 (5), pp.2683-2698
ISSN
0018-9286
Publisher
Institute of Electrical and Electronics Engineers
Start Page
2683
End Page
2698
Journal / Book Title
IEEE Transactions on Automatic Control
Volume
68
Issue
5
Copyright Statement
© 2023 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.
Publication Status
Published
Date Publish Online
2022-08-17