Enabling fast uncertainty estimation: accelerating bayesian transformers via algorithmic and hardware optimizations
File(s)dac22hf3_final_bayesatt.pdf (1.11 MB)
Accepted version
Author(s)
Fan, Hongxiang
Ferianc, Martin
Luk, Wayne
Type
Conference Paper
Abstract
Quantifying the uncertainty of neural networks (NNs) has been
required by many safety-critical applications such as autonomous
driving or medical diagnosis. Recently, Bayesian transformers have
demonstrated their capabilities in providing high-quality uncer-
tainty estimates paired with excellent accuracy. However, their
real-time deployment is limited by the compute-intensive attention
mechanism that is core to the transformer architecture, and the
repeated Monte Carlo sampling to quantify the predictive uncer-
tainty. To address these limitations, this paper accelerates Bayesian
transformers via both algorithmic and hardware optimizations. On
the algorithmic level, an evolutionary algorithm (EA)-based frame-
work is proposed to exploit the sparsity in Bayesian transformers
and ease their computational workload. On the hardware level,
we demonstrate that the sparsity brings hardware performance
improvement on our optimized CPU and GPU implementations.
An adaptable hardware architecture is also proposed to acceler-
ate Bayesian transformers on an FPGA. Extensive experiments
demonstrate that the EA-based framework, together with hardware
optimizations, reduce the latency of Bayesian transformers by up to
13, 12 and 20 times on CPU, GPU and FPGA platforms respectively,
while achieving higher algorithmic performance.
required by many safety-critical applications such as autonomous
driving or medical diagnosis. Recently, Bayesian transformers have
demonstrated their capabilities in providing high-quality uncer-
tainty estimates paired with excellent accuracy. However, their
real-time deployment is limited by the compute-intensive attention
mechanism that is core to the transformer architecture, and the
repeated Monte Carlo sampling to quantify the predictive uncer-
tainty. To address these limitations, this paper accelerates Bayesian
transformers via both algorithmic and hardware optimizations. On
the algorithmic level, an evolutionary algorithm (EA)-based frame-
work is proposed to exploit the sparsity in Bayesian transformers
and ease their computational workload. On the hardware level,
we demonstrate that the sparsity brings hardware performance
improvement on our optimized CPU and GPU implementations.
An adaptable hardware architecture is also proposed to acceler-
ate Bayesian transformers on an FPGA. Extensive experiments
demonstrate that the EA-based framework, together with hardware
optimizations, reduce the latency of Bayesian transformers by up to
13, 12 and 20 times on CPU, GPU and FPGA platforms respectively,
while achieving higher algorithmic performance.
Date Issued
2022-07-01
Date Acceptance
2022-02-21
Citation
DAC '22: Proceedings of the 59th ACM/IEEE Design Automation Conference, 2022, pp.325-330
ISBN
9781450391429
Publisher
ACM / IEEE
Start Page
325
End Page
330
Journal / Book Title
DAC '22: Proceedings of the 59th ACM/IEEE Design Automation Conference
Copyright Statement
© 2022 ACM. This is the author's version of the work. It is posted here by permission of ACM for your personal use. Not for redistribution. The definitive version was published in DAC '22: Proceedings of the 59th ACM/IEEE Design Automation Conference (01 Jul 2022) https://dl.acm.org/doi/abs/10.1145/3489517.3530451
Source
Design Automation Conference (DAC) 2022
Publication Status
Published
Start Date
2022-07-10
Finish Date
2022-07-14
Coverage Spatial
San Francisco, USA
Date Publish Online
2022-08-23