The impact of exploration on convergence and performance of multi-agent Q-learning dynamics
File(s)hussain23a.pdf (847.12 KB)
Published version
Author(s)
Hussain, A
Belardinelli, F
Paccagnan, D
Type
Conference Paper
Abstract
Understanding the impact of exploration on the behaviour of multi-agent learning has, so far, benefited from the restriction to potential, or network zero-sum games in which convergence to an equilibrium can be shown. Outside of these classes, learning dynamics rarely converge and little is known about the effect of exploration in the face of non-convergence. To progress this front, we study the smooth Q-Learning dynamics. We show that, in any network game, exploration by agents results in the convergence of Q-Learning to a neighbourhood of an equilibrium. This holds independently of whether the dynamics reach the equilibrium or display complex behaviours. We show that increasing the exploration rate decreases the size of this neighbourhood and also decreases the ability of all agents to improve their payoffs. Furthermore, in a broad class of games, the payoff performance of Q-Learning dynamics, measured by Social Welfare, decreases when the exploration rate increases. Our experiments show this to be a general phenomenon, namely that exploration leads to improved convergence of Q-Learning, at the cost of payoff performance.
Date Issued
2023
Date Acceptance
2023-07-23
Citation
Proceedings of Machine Learning Research, 2023, 202, pp.14178-14202
ISSN
2640-3498
Publisher
Proceedings of Machine Learning Research
Start Page
14178
End Page
14202
Journal / Book Title
Proceedings of Machine Learning Research
Volume
202
Copyright Statement
Copyright
2023 by the author(s).
2023 by the author(s).
Source
International Conference on Machine Learning
Publication Status
Published
Start Date
2023-07-23
Finish Date
2023-07-29
Coverage Spatial
Honolulu, Hawaii, USA
Date Publish Online
2023