Offline reinforcement learning with behavioral supervisor tuning
File(s)IJCAI_2024 (1).pdf (1.07 MB)
Accepted version
Author(s)
Srinivasan, Padmanaba
Knottenbelt, William
Type
Conference Paper
Abstract
Offline reinforcement learning (RL) algorithms are
applied to learn performant, well-generalizing poli-
cies when provided with a static dataset of inter-
actions. Many recent approaches to offline RL
have seen substantial success, but with one key
caveat: they demand substantial per-dataset hyper-
parameter tuning to achieve reported performance,
which requires policy rollouts in the environment
to evaluate; this can rapidly become cumbersome.
Furthermore, substantial tuning requirements can
hamper the adoption of these algorithms in prac-
tical domains. In this paper, we present TD3 with
Behavioral Supervisor Tuning (TD3-BST), an al-
gorithm that trains an uncertainty model and uses
it to guide the policy to select actions within the
dataset support. TD3-BST can learn more effec-
tive policies from offline datasets compared to pre-
vious methods and achieves the best performance
across challenging benchmarks without requiring
per-dataset tuning.
applied to learn performant, well-generalizing poli-
cies when provided with a static dataset of inter-
actions. Many recent approaches to offline RL
have seen substantial success, but with one key
caveat: they demand substantial per-dataset hyper-
parameter tuning to achieve reported performance,
which requires policy rollouts in the environment
to evaluate; this can rapidly become cumbersome.
Furthermore, substantial tuning requirements can
hamper the adoption of these algorithms in prac-
tical domains. In this paper, we present TD3 with
Behavioral Supervisor Tuning (TD3-BST), an al-
gorithm that trains an uncertainty model and uses
it to guide the policy to select actions within the
dataset support. TD3-BST can learn more effec-
tive policies from offline datasets compared to pre-
vious methods and achieves the best performance
across challenging benchmarks without requiring
per-dataset tuning.
Date Issued
2024-08-03
Date Acceptance
2024-04-16
Citation
Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, 2024, pp.4929-4937
ISBN
978-1-956792-04-1
Publisher
International Joint Conferences on Artificial Intelligence
Start Page
4929
End Page
4937
Journal / Book Title
Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence
Copyright Statement
Copyright © 2024 International Joint Conferences on Artificial Intelligence This is the author’s accepted manuscript made available under a CC-BY licence in accordance with Imperial’s Research Publications Open Access policy (www.imperial.ac.uk/oa-policy)
License URL
Identifier
https://www.ijcai.org/proceedings/2024/0545
Source
33rd International Joint Conference on Artificial Intelligence (IJCAI 2024)
Publication Status
Published
Start Date
2024-08-03
Finish Date
2024-08-09
Coverage Spatial
Jeju, South Korea
Date Publish Online
2024-08-03