Continual learning in deep eural network by using a kalman optimiser
File(s) 1905.08119v3.pdf (531.07 KB)
Working paper
Author(s)
Li, Honglin
Enshaeifar, Shirin
Ganz, Frieder
Barnaghi, Payam
Type
Working Paper
Abstract
Learning and adapting to new distributions or learning new tasks sequentially
without forgetting the previously learned knowledge is a challenging phenomenon
in continual learning models. Most of the conventional deep learning models are
not capable of learning new tasks sequentially in one model without forgetting
the previously learned ones. We address this issue by using a Kalman Optimiser.
The Kalman Optimiser divides the neural network into two parts: the long-term
and short-term memory units. The long-term memory unit is used to remember the
learned tasks and the short-term memory unit is to adapt to the new task. We
have evaluated our method on MNIST, CIFAR10, CIFAR100 datasets and compare our
results with state-of-the-art baseline models. The results show that our
approach enables the model to continually learn and adapt to the new changes
without forgetting the previously learned tasks.
without forgetting the previously learned knowledge is a challenging phenomenon
in continual learning models. Most of the conventional deep learning models are
not capable of learning new tasks sequentially in one model without forgetting
the previously learned ones. We address this issue by using a Kalman Optimiser.
The Kalman Optimiser divides the neural network into two parts: the long-term
and short-term memory units. The long-term memory unit is used to remember the
learned tasks and the short-term memory unit is to adapt to the new task. We
have evaluated our method on MNIST, CIFAR10, CIFAR100 datasets and compare our
results with state-of-the-art baseline models. The results show that our
approach enables the model to continually learn and adapt to the new changes
without forgetting the previously learned tasks.
Date Issued
2019-05-24
Citation
2019
Publisher
arXiv
Copyright Statement
© 2019 The Author(s).
Identifier
http://arxiv.org/abs/1905.08119v3
Subjects
cs.LG
cs.LG
stat.ML
Notes
accepted by ICML workshop
Publication Status
Published online
