Learning reward machines in cooperative multi-agent tasks
File(s)Chapter Author Accepted.pdf (1.14 MB)
Accepted version
Author(s)
Ardon, Leo
Furelos-Blanco, Daniel
Russo, Alessandra
Type
Conference Paper
Abstract
This paper presents a novel approach to Multi-Agent Reinforcement Learning (MARL) that combines cooperative task decomposition with the learning of Reward Machines (RMs) encoding the structure of the sub-tasks. The proposed method helps deal with the non-Markovian nature of the rewards in partially observable environments and improves the interpretability of the learnt policies required to complete a cooperative task. The RMs associated with the sub-tasks are learnt in a decentralised manner and then used to guide the behaviour of each agent in a team acting towards a common goal. By doing so, the complexity of a cooperative multi-agent problem is reduced, allowing for more effective learning. The results suggest that our approach is a promising direction for future research in cooperative MARL, especially in complex and partially observable environments.
Date Issued
2024-03-29
Date Acceptance
2023-05-29
Citation
Autonomous Agents and Multiagent Systems. Best and Visionary Papers, 2024, pp.43-59
ISBN
9783031562549
ISSN
0302-9743
Publisher
Springer Nature Switzerland
Start Page
43
End Page
59
Journal / Book Title
Autonomous Agents and Multiagent Systems. Best and Visionary Papers
Copyright Statement
Copyright © 2024 Springer-Verlag. This version of the article has been accepted for publication, after peer review (when applicable) and is subject to Springer Nature’s AM terms of use, but is not the Version of Record and does not reflect post-acceptance improvements, or any corrections. The Version of Record is available online at: http://dx.doi.org/10.1007/978-3-031-56255-6_3
Identifier
http://dx.doi.org/10.1007/978-3-031-56255-6_3
Source
AAMAS 2023 Workshops
Publication Status
Published
Start Date
2023-05-29
Finish Date
2023-06-02
Coverage Spatial
London, UK
Rights Embargo Date
2025-03-29
Date Publish Online
2024-03-30