Caffe barista: brewing caffe with FPGAs in the training loop
File(s) 2006.13829v1.pdf (462.54 KB)
Working paper
Author(s)
Vink, Diederik Adriaan
Rajagopal, Aditya
Venieris, Stylianos I
Bouganis, Christos-Savvas
Type
Working Paper
Abstract
As the complexity of deep learning (DL) models increases, their compute
requirements increase accordingly. Deploying a Convolutional Neural Network
(CNN) involves two phases: training and inference. With the inference task
typically taking place on resource-constrained devices, a lot of research has
explored the field of low-power inference on custom hardware accelerators. On
the other hand, training is both more compute- and memory-intensive and is
primarily performed on power-hungry GPUs in large-scale data centres. CNN
training on FPGAs is a nascent field of research. This is primarily due to the
lack of tools to easily prototype and deploy various hardware and/or
algorithmic techniques for power-efficient CNN training. This work presents
Barista, an automated toolflow that provides seamless integration of FPGAs into
the training of CNNs within the popular deep learning framework Caffe. To the
best of our knowledge, this is the only tool that allows for such versatile and
rapid deployment of hardware and algorithms for the FPGA-based training of
CNNs, providing the necessary infrastructure for further research and
development.
requirements increase accordingly. Deploying a Convolutional Neural Network
(CNN) involves two phases: training and inference. With the inference task
typically taking place on resource-constrained devices, a lot of research has
explored the field of low-power inference on custom hardware accelerators. On
the other hand, training is both more compute- and memory-intensive and is
primarily performed on power-hungry GPUs in large-scale data centres. CNN
training on FPGAs is a nascent field of research. This is primarily due to the
lack of tools to easily prototype and deploy various hardware and/or
algorithmic techniques for power-efficient CNN training. This work presents
Barista, an automated toolflow that provides seamless integration of FPGAs into
the training of CNNs within the popular deep learning framework Caffe. To the
best of our knowledge, this is the only tool that allows for such versatile and
rapid deployment of hardware and algorithms for the FPGA-based training of
CNNs, providing the necessary infrastructure for further research and
development.
Date Issued
2020-06-18
Citation
2020
Publisher
arXiv
Copyright Statement
© 2020 The Author(s)
Sponsor
Huawei Technologies Co. Ltd
Identifier
http://arxiv.org/abs/2006.13829v1
Grant Number
YBN2017070060
Subjects
cs.DC
cs.DC
cs.LG
Notes
Published as short paper at FPL2020
Publication Status
Published
