Offline reinforcement learning: in pursuit of perfect policies from imperfect data
File(s)
Author(s)
Srinivasan, Padmanaba
Type
Thesis
Abstract
Intelligent agents can solve tasks in dynamic environments without the need for human supervision. Many environments are so large and complex that specifying exactly how an agent should behave is challenging. Instead, an agent must learn through trial and error from environment feedback; this learning paradigm is reinforcement learning (RL). Rather than being explicitly told how to behave, RL agents improve their policy by interacting with the environment and learning which actions maximize the long-term utility of a reward function in a sequential decision-making process. The sample efficiency of an algorithm describes how performance scales with data. Maximizing efficiency is critical as environments grow; with it, the data needed grows exponentially due to the curse of dimensionality.
An alternative RL paradigm explores how to learn without any direct environment interaction. Given historical data of (imperfect) demonstrators attempting to complete a task, offline RL aims to recover an optimal policy in a setting that requires maximal sample efficiency.
This thesis presents three novel algorithms for offline RL. First, we design an algorithm that expands on a minimalist approach to scale to multimodal datasets without the tuning complexity that recent methods increasingly assume. Second, we extend an existing model-free method to derive an effective and efficient offline model-based RL algorithm. Finally, we reframe offline policy improvement as a regression problem that aligns the policy towards high-reward, in-sample behaviors. We evaluate our algorithm in off-policy and on-policy settings and show that our method can enjoy the stability benefits of on-policy learning while remaining competitive with both our off-policy variant and other baselines.
An alternative RL paradigm explores how to learn without any direct environment interaction. Given historical data of (imperfect) demonstrators attempting to complete a task, offline RL aims to recover an optimal policy in a setting that requires maximal sample efficiency.
This thesis presents three novel algorithms for offline RL. First, we design an algorithm that expands on a minimalist approach to scale to multimodal datasets without the tuning complexity that recent methods increasingly assume. Second, we extend an existing model-free method to derive an effective and efficient offline model-based RL algorithm. Finally, we reframe offline policy improvement as a regression problem that aligns the policy towards high-reward, in-sample behaviors. We evaluate our algorithm in off-policy and on-policy settings and show that our method can enjoy the stability benefits of on-policy learning while remaining competitive with both our off-policy variant and other baselines.
Version
Open Access
Date Issued
2025-01-19
Date Awarded
01/07/2025
License URL
Advisor
Knottenbelt, William
Publisher Department
Department of Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
