Policy evaluation in distributional LQR
File(s)Distributional LQR.pdf (1.34 MB)
Accepted version
Author(s)
Type
Journal Article
Abstract
Distributional reinforcement learning (DRL) enhances the understanding of the effects of the randomness in the environment by letting agents learn the distribution of a random return, rather than its expected value as in standard reinforcement learning. Meanwhile, a challenge in DRL is that the policy evaluation typically relies on the representation of the return distribution, which needs to be carefully designed. In this paper, we address this challenge for the special class of DRL problems that rely on a discounted linear quadratic regulator (LQR), which we call distributional LQR. Specifically, we provide a closed-form expression for the distribution of the random return, which is applicable for all types of exogenous disturbance as long as it is independent and identically distributed (i.i.d.). We show that the variance of the random return is bounded if the fourth moment of the exogenous disturbance is bounded. Furthermore, we investigate the sensitivity of the return distribution to model perturbations. While the proposed exact return distribution consists of infinitely many random variables, we show that this distribution can be well approximated by a finite number of random variables. The associated approximation error can be analytically bounded under mild assumptions. When the model is unknown, we propose a model-free approach for estimating the return distribution, supported by sample complexity guarantees. Finally, we extend our approach to partially observable linear systems. Numerical experiments are provided to illustrate the theoretical results.
Date Issued
2025-11-01
Date Acceptance
2025-06-01
Citation
IEEE Transactions on Automatic Control, 2025, 70 (11), pp.7477-7492
ISSN
0018-9286
Publisher
Institute of Electrical and Electronics Engineers (IEEE)
Start Page
7477
End Page
7492
Journal / Book Title
IEEE Transactions on Automatic Control
Volume
70
Issue
11
Copyright Statement
Copyright © 2025 IEEE. This is the author’s accepted manuscript made available under a CC-BY licence in accordance with Imperial’s Research Publications Open Access policy (www.imperial.ac.uk/oa-policy)
License URL
Publication Status
Published
Date Publish Online
2025-06-02