Sample-efficiency in multi-batch reinforcement learning: the need for dimension-dependent adaptivity
File(s) ICLR_2024_Submission.pdf (713.44 KB)
Accepted version
Author(s)
Johnson, Emmeran
Pike-Burke, Ciara
Rebeschini, Patrick
Type
Conference Paper
Abstract
We theoretically explore the relationship between sample-efficiency and adaptivity in reinforcement learning. An algorithm is sample-efficient if it uses a number of queries n to the environment that is polynomial in the dimension d of the problem. Adaptivity refers to the frequency at which queries are sent and feedback is processed to update the querying strategy. To investigate this interplay, we employ
a learning framework that allows sending queries in K batches, with feedback being processed and queries updated after each batch. This model encompasses the
whole adaptivity spectrum, ranging from non-adaptive ‘offline’ (K “ 1) to fully adaptive (K “ n) scenarios, and regimes in between. For the problems of policy evaluation and best-policy identification under d-dimensional linear function approximation, we establish Ωplog log dq lower bounds on the number of batches K required for sample-efficient algorithms with n “ Oppolypdqq queries. Our results
show that just having adaptivity (K ą 1) does not necessarily guarantee sample-efficiency. Notably, the adaptivity-boundary for sample-efficiency is not between offline reinforcement learning (K “ 1), where sample-efficiency was known to not be possible, and adaptive settings. Instead, the boundary lies between different regimes of adaptivity and depends on the problem dimension.
a learning framework that allows sending queries in K batches, with feedback being processed and queries updated after each batch. This model encompasses the
whole adaptivity spectrum, ranging from non-adaptive ‘offline’ (K “ 1) to fully adaptive (K “ n) scenarios, and regimes in between. For the problems of policy evaluation and best-policy identification under d-dimensional linear function approximation, we establish Ωplog log dq lower bounds on the number of batches K required for sample-efficient algorithms with n “ Oppolypdqq queries. Our results
show that just having adaptivity (K ą 1) does not necessarily guarantee sample-efficiency. Notably, the adaptivity-boundary for sample-efficiency is not between offline reinforcement learning (K “ 1), where sample-efficiency was known to not be possible, and adaptive settings. Instead, the boundary lies between different regimes of adaptivity and depends on the problem dimension.
Date Acceptance
2024-01-16
Publisher
ICLR
Copyright Statement
© 2024 The Author(s).
Source
International Conference on Learning Representations (ICLR 2024)
Publication Status
Accepted
Start Date
2024-05-07
Finish Date
2024-05-11
Coverage Spatial
Vienna, Austria
