Hierarchical model-based deep reinforcement learning for trading
File(s)
Author(s)
Millea, Adrian
Type
Thesis
Abstract
The kernel of this thesis posits that a hierarchical reinforcement learning (RL) system, with specialized agents, some of them being deep RL (DRL) agents, trained on distinct market data or employing diverse decision-making models, will surpass single-agent approaches in adapting to varying market conditions, optimizing risk-adjusted returns, and maintaining robust performance across different financial markets without frequent retraining.
We first develop a single-asset approach consisting of a hierarchical reinforcement learning (RL) architecture which employs different low-level agents to act in the environment, i.e. the market. The highest level agent selects among a group of different specialised agents, which then act on the market. We add two simple decision-making models to the set of low-level agents, which diversify the pool of available agents, thus increasing overall behaviour flexibility. We also use a risk-based reward at the high-level which transforms the overall problem into a risk--return optimization. This shows significant reduction in risk while minimally reducing profits.
We then devise a hierarchical decision-making architecture for portfolio optimization. At the highest level a DRL agent selects among a number of discrete actions, representing low-level agents. For the low-level agents, we use a set of Hierarchical Risk Parity (HRP) and Hierarchical Equal Risk Contribution (HERC) models with different hyperparameters, which all run in parallel, off-market (in a simulation). The modelling resembles a statefull, non-stationary, multi-arm bandit, where the performance of the individual arms changes with time and is assumed to be dependent on the recent history. We perform experiments on the cryptocurrency market (117 assets), on the stock market (46 assets) and on the foreign exchange market (28 pairs) showing the excellent robustness and performance of the overall system. Moreover, we eliminate the need for retraining and are able to deal with large testing sets successfully.
We first develop a single-asset approach consisting of a hierarchical reinforcement learning (RL) architecture which employs different low-level agents to act in the environment, i.e. the market. The highest level agent selects among a group of different specialised agents, which then act on the market. We add two simple decision-making models to the set of low-level agents, which diversify the pool of available agents, thus increasing overall behaviour flexibility. We also use a risk-based reward at the high-level which transforms the overall problem into a risk--return optimization. This shows significant reduction in risk while minimally reducing profits.
We then devise a hierarchical decision-making architecture for portfolio optimization. At the highest level a DRL agent selects among a number of discrete actions, representing low-level agents. For the low-level agents, we use a set of Hierarchical Risk Parity (HRP) and Hierarchical Equal Risk Contribution (HERC) models with different hyperparameters, which all run in parallel, off-market (in a simulation). The modelling resembles a statefull, non-stationary, multi-arm bandit, where the performance of the individual arms changes with time and is assumed to be dependent on the recent history. We perform experiments on the cryptocurrency market (117 assets), on the stock market (46 assets) and on the foreign exchange market (28 pairs) showing the excellent robustness and performance of the overall system. Moreover, we eliminate the need for retraining and are able to deal with large testing sets successfully.
Version
Open Access
Date Issued
2023-03
Date Awarded
2024-03
Copyright Statement
Creative Commons Attribution ShareAlike Licence
License URL
Advisor
Edalat, Abbas
Sponsor
Engineering and Physical Sciences Research Council
Grant Number
EP/L016796/1
Publisher Department
Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
