Diameter-independent reinforcement learning for the control of state-dependent queues
File(s) main-cr.pdf (801.28 KB)
Accepted version
Author(s)
Sheldon, Matthew
Casale, Giuliano
Type
Journal Article
Abstract
The regret of general-purpose average-cost reinforcement learning algorithms on queue control problems
often varies exponentially with the number of states. This is due to a dependence on the diameter D, equal
to the largest hitting time between any two states under the most favorable policy, which is exponential in
many queue control problems. In this paper, we present an algorithm for the control of a state-dependent
Markovian queue. Under one of two sufficient conditions, the proposed algorithm achieves a regret bound
of ˜O(κ2S1.5 (N+ +N−)T), where S is the number of states, N+ is the number of available arrival rates at each state, N− is the number of available service rates at each state, and κ is the ratio between the maximal and minimal transition rate bounds. In contrast to prior work on diameter-independent model based learning in queues, we do not assume that the transition rates follow a specific parametrization, and
allow for general state-dependent rewards.
often varies exponentially with the number of states. This is due to a dependence on the diameter D, equal
to the largest hitting time between any two states under the most favorable policy, which is exponential in
many queue control problems. In this paper, we present an algorithm for the control of a state-dependent
Markovian queue. Under one of two sufficient conditions, the proposed algorithm achieves a regret bound
of ˜O(κ2S1.5 (N+ +N−)T), where S is the number of states, N+ is the number of available arrival rates at each state, N− is the number of available service rates at each state, and κ is the ratio between the maximal and minimal transition rate bounds. In contrast to prior work on diameter-independent model based learning in queues, we do not assume that the transition rates follow a specific parametrization, and
allow for general state-dependent rewards.
Date Acceptance
2026-08-07
Citation
Performance evaluation (Print)
ISSN
0166-5316
Publisher
Elsevier
Journal / Book Title
Performance evaluation (Print)
Copyright Statement
Copyright This paper is embargoed until publication. Once published the Version of Record (VoR) will be available on immediate open access.
License URL
Publication Status
Accepted
