Efficient weight reuse for large LSTMs
File(s)asap19zq.pdf (873.85 KB)
Accepted version
Author(s)
Type
Conference Paper
Abstract
Long Short-Term Memory (LSTM) networks havebeen deployed in speech recognition, natural language processingand financial calculations in recent years, and are beginning to beused in systems where low latency and low power are required.To meet such requirements, we propose a stall-free hardwarearchitecture by reorganising the order of operations in an LSTMsystem and develop a unique blocking-batching strategy to reusethe LSTM weights fetched from external memory to optimise thebenefits of on-chip memory with a limited size for a large machinelearning model. Evaluation results show that our architecture canachieve up to 20.8 GOPS/W, which would be among the highestfor FPGA designs targeting LSTM systems with weights stored inoff-chip memory. Comparing to the state-of-the-art design usingoff-chip memory to store the weights, we achieve 1.65 timeshigher performance-per-watt efficiency and 1.60 times higherperformance-per-DSP efficiency. When compared with CPU andGPU implementation, our novel hardware architecture is 23.7and 1.3 times faster while consuming 208 and 19.2 times lowerenergy respectively, which shows that our approach contributesto high performance and low power FPGA-based LSTM systems.
Date Issued
2019-09-05
Date Acceptance
2019-05-10
Citation
2019 IEEE 30th International Conference on Application-specific Systems, Architectures and Processors (ASAP), 2019
Publisher
IEEE
Journal / Book Title
2019 IEEE 30th International Conference on Application-specific Systems, Architectures and Processors (ASAP)
Copyright Statement
© 2019 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.
Source
The 30th IEEE International Conference on Application-specific Systems, Architectures and Processors
Publication Status
Published
Start Date
2019-07-15
Finish Date
2019-07-17
Coverage Spatial
New York, USA