Learning for optimized provisioning of network resources
File(s)
Author(s)
Chen, Zheyu
Type
Thesis
Abstract
This thesis focuses on developing novel learning techniques for resource provisioning in communication networks with the goal of identifying the optimal resource provisioning decisions. In this context, both the static and dynamic problem formulations are considered to investigate the impact of various resource restrictions and time evolution. Initially, we explore the static problem formulation that models the resource provisioning problems at different time instances as individually constrained optimization problems. To address the computational challenges of repeatedly solving these problems due to the time evolution (i.e., the changing system parameters), two Coupled Long Short-Term Memory networks (CLSTM) are proposed. Extensive numerical experiments using Alibaba datasets validate that the CLSTM can quickly produce the optimal solution over a range of system parameters in a few iterations during the inference (e.g., achieving 98% mean relative accuracy after 67 iterations), significantly reducing CPU time and iteration consumption compared with conventional gradient-based iterative algorithms. To incorporate time evolution and system dynamics, the thesis considers the dynamic problem formulation. As communication networks often demonstrate periodic temporal characteristics (e.g., periodic task arrivals), we model the resource provisioning problem as a periodic Markov Decision Process (pMDP). A model-free deep-learning method with a multi-policies solution framework is developed to address the lack of existing model-free algorithms for pMDPs. Evaluation results demonstrate the proposed technique outperforms the state-of-the-art deep reinforcement learning algorithm by 32% on average. To further improve the computational efficiency of applying the proposed multi-policies solution, we introduce the multi-MDP framework that allows a single policy to be used for non-uniform time durations, thus reducing the number of required policies. Our experimental results using the practical task arrivals from Alibaba reveal that the task utilities obtained using the proposed framework of sequential policies can yield an average improvement of 23% over those from a single-policy RL solution.
Version
Open Access
Date Issued
2024-03
Date Awarded
2024-09
Copyright Statement
Creative Commons Attribution NonCommercial Licence
License URL
Advisor
Leung, Kin
Publisher Department
Department of Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
