Stability of multi-agent learning in competitive networks: delaying the onset of chaos
File(s)Stability_and_Chaos-3.pdf (703.64 KB)
Accepted version
Author(s)
Hussain, Aamal
Belardinelli, Francesco
Type
Conference Paper
Abstract
The behaviour of multi agent learning in competitive network games is often studied within the context of zero sum games, in which convergence guarantees may be obtained. However, outside of this class the behaviour of learning is known to display complex behaviours and convergence cannot be always guaranteed. Nonetheless, in order to develop a complete picture of the behaviour of multi agent learning in competitive settings, the zero sum assumption must be lifted.
Motivated by this we study the Q Learning dynamics, a popular model of exploration and exploitation in multi agent learning, in competitive network games. We determine how the degree of competition, exploration rate and network connectivity impact the convergence of Q Learning. To study generic competitive games, we parameterise network games in terms of correlations between agent payoffs and study the average behaviour of the Q Learning dynamics across all games drawn from a choice of this parameter. This statistical approach establishes choices of parameters for which Q Learning dynamics converge to a stable fixed point. Differently to previous works, we find that the stability of Q Learning is explicitly dependent only on the network connectivity rather than the total number of agents. Our experiments validate these findings and show that, under certain network structures, the total number of agents can be increased without increasing the likelihood of unstable or chaotic behaviours.
Motivated by this we study the Q Learning dynamics, a popular model of exploration and exploitation in multi agent learning, in competitive network games. We determine how the degree of competition, exploration rate and network connectivity impact the convergence of Q Learning. To study generic competitive games, we parameterise network games in terms of correlations between agent payoffs and study the average behaviour of the Q Learning dynamics across all games drawn from a choice of this parameter. This statistical approach establishes choices of parameters for which Q Learning dynamics converge to a stable fixed point. Differently to previous works, we find that the stability of Q Learning is explicitly dependent only on the network connectivity rather than the total number of agents. Our experiments validate these findings and show that, under certain network structures, the total number of agents can be increased without increasing the likelihood of unstable or chaotic behaviours.
Date Issued
2024-03-24
Date Acceptance
2023-12-09
Citation
Proceedings of the AAAI Conference on Artificial Intelligence, 2024, 38 (16), pp.17435-17443
ISSN
2159-5399
Publisher
Association for the Advancement of Artificial Intelligence (AAAI)
Start Page
17435
End Page
17443
Journal / Book Title
Proceedings of the AAAI Conference on Artificial Intelligence
Volume
38
Issue
16
Copyright Statement
© 2024, Association for the Advancement of Artificial
Intelligence (www.aaai.org). All rights reserved.
Intelligence (www.aaai.org). All rights reserved.
Identifier
http://dx.doi.org/10.1609/aaai.v38i16.29692
Source
Thirty-Eighth AAAI Conference on Artificial Intelligence
Publication Status
Published
Start Date
2024-02-20
Finish Date
2024-02-27
Coverage Spatial
Vancouver, Canada