Recurrent neural networks with column-wise matrix-vector multiplication on FPGAs
File(s)tvlsi21zq_cr05.pdf (1.45 MB)
Accepted version
Author(s)
Type
Journal Article
Abstract
This article presents a reconfigurable accelerator for REcurrent Neural networks with fine-grained cOlumn-Wise matrix-vector multiplicatioN (RENOWN). We propose a novel latency-hiding architecture for recurrent neural network (RNN) acceleration using column-wise matrix-vector multiplication (MVM) instead of the state-of-the-art row-wise operation. This hardware (HW) architecture can eliminate data dependencies to improve the throughput of RNN inference systems. Besides, we introduce a configurable checkerboard tiling strategy which allows large weight matrices, while incorporating various configurations of element-based parallelism (EP) and vector-based parallelism (VP). These optimizations improve the exploitation of parallelism to increase HW utilization and enhance system throughput. Evaluation results show that our design can achieve over 29.6 tera operations per second (TOPS) which would be among the highest for field-programmable gate array (FPGA)-based RNN designs. Compared to state-of-the-art accelerators on FPGAs, our design achieves 3.7-14.8 times better performance and has the highest HW utilization.
Date Issued
2021-12-29
Date Acceptance
2021-11-20
Citation
IEEE Transactions on Very Large Scale Integration (VLSI) Systems, 2021, 30 (2), pp.227-237
ISSN
1063-8210
Publisher
Institute of Electrical and Electronics Engineers (IEEE)
Start Page
227
End Page
237
Journal / Book Title
IEEE Transactions on Very Large Scale Integration (VLSI) Systems
Volume
30
Issue
2
Copyright Statement
© 2021 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.
Identifier
https://ieeexplore.ieee.org/document/9664799
Subjects
Science & Technology
Technology
Computer Science, Hardware & Architecture
Engineering, Electrical & Electronic
Computer Science
Engineering
Logic gates
Computer architecture
Recurrent neural networks
Throughput
Optimization
Field programmable gate arrays
Microprocessors
Hardware Accelerator
long short-term (LSTM)
recurrent neural network (RNN)
MEMORY
Computer Hardware & Architecture
0805 Distributed Computing
0906 Electrical and Electronic Engineering
1006 Computer Hardware
Publication Status
Published
Date Publish Online
2021-12-29