Learning robust reward machines from noisy labels
File(s) kr_submission_noisy_rm-8.pdf (1.6 MB)
Accepted version
Author(s)
Type
Conference Paper
Abstract
This paper presents PROB-IRM, an approach that learns
robust reward machines (RMs) for reinforcement learning
(RL) agents from noisy execution traces. The key aspect
of RM-driven RL is the exploitation of a finite-state ma-
chine that decomposes the agent’s task into different sub-
tasks. PROB-IRM uses a state-of-the-art inductive logic programming framework robust to noisy examples to learn RMs from noisy traces using the Bayesian posterior degree of beliefs, thus ensuring robustness against inconsistencies. Pivotal for the results is the interleaving between RM learning and policy learning: a new RM is learned whenever the RL agent generates a trace that is believed not to be accepted by the current RM. To speed up the training of the RL agent, PROB-IRM employs a probabilistic formulation of reward shaping that uses the posterior Bayesian beliefs derived from the traces. Our experimental analysis shows that PROB-IRM can learn (potentially imperfect) RMs from noisy traces and exploit them to train an RL agent to solve its tasks successfully. Despite the complexity of learning the RM from noisy traces, agents trained with PROB-IRM perform comparably to agents provided with handcrafted RMs.
robust reward machines (RMs) for reinforcement learning
(RL) agents from noisy execution traces. The key aspect
of RM-driven RL is the exploitation of a finite-state ma-
chine that decomposes the agent’s task into different sub-
tasks. PROB-IRM uses a state-of-the-art inductive logic programming framework robust to noisy examples to learn RMs from noisy traces using the Bayesian posterior degree of beliefs, thus ensuring robustness against inconsistencies. Pivotal for the results is the interleaving between RM learning and policy learning: a new RM is learned whenever the RL agent generates a trace that is believed not to be accepted by the current RM. To speed up the training of the RL agent, PROB-IRM employs a probabilistic formulation of reward shaping that uses the posterior Bayesian beliefs derived from the traces. Our experimental analysis shows that PROB-IRM can learn (potentially imperfect) RMs from noisy traces and exploit them to train an RL agent to solve its tasks successfully. Despite the complexity of learning the RM from noisy traces, agents trained with PROB-IRM perform comparably to agents provided with handcrafted RMs.
Date Acceptance
2024-07-25
Citation
Proceedings of the 21st International Conference on Principles of Knowledge Representation and Reasoning (KR 2024)
Publisher
IJCAI Organization
Journal / Book Title
Proceedings of the 21st International Conference on Principles of Knowledge Representation and Reasoning (KR 2024)
Copyright Statement
Subject to copyright. This paper is embargoed until publication. Once published the author’s accepted manuscript will be made available under a CC-BY License in accordance with Imperial’s Research Publications Open Access policy (www.imperial.ac.uk/oa-policy).
License URL
Source
21st International Conference on Principles of Knowledge Representation and Reasoning (KR 2024)
Publication Status
Accepted
Start Date
2024-11-02
Finish Date
2024-11-08
Coverage Spatial
Hanoi, Vietnam
Rights Embargo Date
10000-01-01
Date Publish Online
2024-11-02
