Asymptotic convergence and performance of multi-agent Q-learning dynamics
File(s)
Author(s)
Hussain, Aamal
Type
Thesis
Abstract
In this thesis we study the Q-Learning
dynamics, a classical model which describes the behaviour of agents who explore their space of
actions whilst simultaneously aiming to maximise their rewards. Our goal is to describe the
behaviour of Q-Learning in as wide an array of settings as possible, and to understand the factors
which contribute towards its convergence.
We first show that Q-Learning can converge to a unique stable strategy, called the
Quantal Response Equilibrium (QRE), in arbitrary normal form games. Next, we study games where interactions between agents are constrained by a network. We determine a number of sufficient
conditions, depending on the game and network structure, which guarantee that agent strategies
converge to a unique QRE.
Next, we perform a statistical analysis of the dynamics of Q-Learning in network games. Here, we
parameterise network games in terms of correlations between agent payoffs and study the average
behaviour of the Q-Learning dynamics. We find that the stability of Q-Learning is explicitly dependent only on the
network connectivity rather than the total number of agents.
We then directly study non-stationary behaviours of Q-Learning. We show that Q-Learning dynamics
converges to within a neighbourhood of a QRE in network games. The size of this neighbourhood is
controlled by the connectivity of the network, exploration rates and payoff structure. Similarly, we
consider competitive network games and show that Q-Learning always converges to a neighbourhood of
the unique QRE of a 'nearby' network zero sum game.
Finally, we study the payoff performance of Q-Learning dynamics. Here, we
present a number of cases which suggest that Q-Learning agents may asymptotically achieve a higher Social Welfare
by following a learning dynamic which does not converge to an equilibrium solution..
dynamics, a classical model which describes the behaviour of agents who explore their space of
actions whilst simultaneously aiming to maximise their rewards. Our goal is to describe the
behaviour of Q-Learning in as wide an array of settings as possible, and to understand the factors
which contribute towards its convergence.
We first show that Q-Learning can converge to a unique stable strategy, called the
Quantal Response Equilibrium (QRE), in arbitrary normal form games. Next, we study games where interactions between agents are constrained by a network. We determine a number of sufficient
conditions, depending on the game and network structure, which guarantee that agent strategies
converge to a unique QRE.
Next, we perform a statistical analysis of the dynamics of Q-Learning in network games. Here, we
parameterise network games in terms of correlations between agent payoffs and study the average
behaviour of the Q-Learning dynamics. We find that the stability of Q-Learning is explicitly dependent only on the
network connectivity rather than the total number of agents.
We then directly study non-stationary behaviours of Q-Learning. We show that Q-Learning dynamics
converges to within a neighbourhood of a QRE in network games. The size of this neighbourhood is
controlled by the connectivity of the network, exploration rates and payoff structure. Similarly, we
consider competitive network games and show that Q-Learning always converges to a neighbourhood of
the unique QRE of a 'nearby' network zero sum game.
Finally, we study the payoff performance of Q-Learning dynamics. Here, we
present a number of cases which suggest that Q-Learning agents may asymptotically achieve a higher Social Welfare
by following a learning dynamic which does not converge to an equilibrium solution..
Version
Open Access
Date Issued
2024-01
Date Awarded
2024-09
Copyright Statement
Creative Commons Attribution Licence
License URL
Advisor
Belardinelli, Francesco
Paccagnan, Dario
Publisher Department
Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)