Synaptic plasticity models for reinforcement learning
File(s)
Author(s)
Zannone, Sara
Type
Thesis
Abstract
Animal learning is based on a process of trial and error. This is a fundamental observation in behavioural psychology, but also the core idea of Reinforcement Learning (RL). As such, Reinforcement Learning provides a suitable theoretical framework to study animal learning, but little is known about how Reinforcement Learning is achieved at the level of the neural networks in our brain. In general, it is believed that salient information is stored in neural network by altering the connections between neurons, a process called synaptic plasticity. Here, we study how known plasticity models can give rise to Reinforcement Learning algorithms in simple models of hippocampal networks. \\
First, we present experimental results from synaptic plasticity experiments under the control of neuromulation: specifically dopamine and acetylcholine. Then, we incorporate these observations in a network model of a navigation task. We find that dopamine is responsible for reward-modulated learning, whereas acetylcholine affects exploration patterns and allows flexible learning in changing environments. Separately, we propose a biologically plausible plasticity rule that can learn a predictive representation of the environment, a concept borrowed from RL theory which is called the Successor Representation. We then analyse the learning dynamics of our network from a normative perspective and relate it to Temporal Difference (TD) methods, well known RL algorithms. This relationship with TD learning provides us with a powerful tool to study the role of various biological parameters in relationship to the RL algorithm. It allows us to propose a new mechanism for learning on behavioural timescales through STDP, and to postulate novel functional roles for neuronal activity patterns called replays. We also find that altering the timescale of the neuronal activity adds flexibility to the representation of the state space, which is useful for subgoal discovery and hierarchical reasoning.
First, we present experimental results from synaptic plasticity experiments under the control of neuromulation: specifically dopamine and acetylcholine. Then, we incorporate these observations in a network model of a navigation task. We find that dopamine is responsible for reward-modulated learning, whereas acetylcholine affects exploration patterns and allows flexible learning in changing environments. Separately, we propose a biologically plausible plasticity rule that can learn a predictive representation of the environment, a concept borrowed from RL theory which is called the Successor Representation. We then analyse the learning dynamics of our network from a normative perspective and relate it to Temporal Difference (TD) methods, well known RL algorithms. This relationship with TD learning provides us with a powerful tool to study the role of various biological parameters in relationship to the RL algorithm. It allows us to propose a new mechanism for learning on behavioural timescales through STDP, and to postulate novel functional roles for neuronal activity patterns called replays. We also find that altering the timescale of the neuronal activity adds flexibility to the representation of the state space, which is useful for subgoal discovery and hierarchical reasoning.
Version
Open Access
Date Issued
2019-10
Date Awarded
2020-03
Copyright Statement
Creative Commons Attribution NonCommercial Licence
Advisor
Clopath, Claudia
Sponsor
Engineering and Physical Sciences Research Council
Publisher Department
Bioengineering
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)