Representation learning for efficient reinforcement learning
File(s)
Author(s)
Pritz, Paul J
Type
Thesis
Abstract
Reinforcement-learning (RL) agents learn from often unstructured, incomplete or complex observations collected from their environment.
In traditional RL approaches, the burden of interpreting these observations falls on the value-function and policy models, which use rewards as the only learning signal.
Simultaneous representation learning, i.e., interpreting observations, and RL, therefore, becomes a challenging task in complex or partially observable environments.
This thesis improves upon traditional approaches by proposing a separation of representation learning and RL.
We cast these as separate tasks and propose learning representations using more meaningful signals than rewards alone, such as observed states, actions and environment dynamics.
These representations are then used in RL as a downstream task.
First, we focus on the single-agent domain.
RL algorithms in this context often struggle to learn strong policies in complex environments with many states and actions.
To alleviate this problem, we propose an embedding methodology, which jointly learns representations for states and actions.
These representations capture structure in the environment and can be used to compress the original state and action spaces, reducing the complexity of the environment.
We use a model of the environment and state information instead of reward as a learning signal for representation learning.
These representations are then used in policy-gradient-based RL algorithms, thereby combining elements of model-free and model-based RL.
We demonstrate the validity of the approach theoretically and empirically.
Second, we turn to the multi-agent domain, where we also propose a separation of representation learning and RL.
However, here we use representation learning to infer information missing from agents' local observations.
We first train a belief model using complete environment information, which is then used by a fully decentralised state-based RL algorithm using local observations only.
We construct partially observable environments and empirically demonstrate our approach's efficacy compared to relevant benchmarks.
In traditional RL approaches, the burden of interpreting these observations falls on the value-function and policy models, which use rewards as the only learning signal.
Simultaneous representation learning, i.e., interpreting observations, and RL, therefore, becomes a challenging task in complex or partially observable environments.
This thesis improves upon traditional approaches by proposing a separation of representation learning and RL.
We cast these as separate tasks and propose learning representations using more meaningful signals than rewards alone, such as observed states, actions and environment dynamics.
These representations are then used in RL as a downstream task.
First, we focus on the single-agent domain.
RL algorithms in this context often struggle to learn strong policies in complex environments with many states and actions.
To alleviate this problem, we propose an embedding methodology, which jointly learns representations for states and actions.
These representations capture structure in the environment and can be used to compress the original state and action spaces, reducing the complexity of the environment.
We use a model of the environment and state information instead of reward as a learning signal for representation learning.
These representations are then used in policy-gradient-based RL algorithms, thereby combining elements of model-free and model-based RL.
We demonstrate the validity of the approach theoretically and empirically.
Second, we turn to the multi-agent domain, where we also propose a separation of representation learning and RL.
However, here we use representation learning to infer information missing from agents' local observations.
We first train a belief model using complete environment information, which is then used by a fully decentralised state-based RL algorithm using local observations only.
We construct partially observable environments and empirically demonstrate our approach's efficacy compared to relevant benchmarks.
Version
Open Access
Date Issued
2025-04-11
Date Awarded
01/01/2026
License URL
Advisor
Leung, Kin
Publisher Department
Department of Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
