Solving the global atmospheric equations through heterogeneous reconfigurable platforms
File(s)trets14lgan1.pdf (790.36 KB)
Accepted version
Author(s)
Type
Journal Article
Abstract
One of the most essential and challenging components in climate modeling is the atmospheric model. To solve multiphysical atmospheric equations, developers have to face extremely complex stencil kernels that are costly in terms of both computing and memory resources. This article aims to accelerate the solution of global shallow water equations (SWEs), which is one of the most essential equation sets describing atmospheric dynamics. We first design a hybrid methodology that employs both the host CPU cores and the field-programmable gate array (FPGA) accelerators to work in parallel. Through a careful adjustment of the computational domains, we achieve a balanced resource utilization and a further improvement of the overall performance. By decomposing the resource-demanding SWE kernel, we manage to map the double-precision algorithm into three FPGAs. Moreover, by using fixed-point and reduced-precision floating point arithmetic, we manage to build a fully pipelined mixed-precision design on a single FPGA, which can perform 428 floating-point and 235 fixed-point operations per cycle. The mixed-precision design with four FPGAs running together can achieve a speedup of 20 over a fully optimized design on a CPU rack with two eight-core processorsand is 8 times faster than the fully optimized Kepler GPU design. As for power efficiency, the mixed-precision design with four FPGAs is 10 times more power efficient than a Tianhe-1A supercomputer node.
Date Issued
2015-03-01
Date Acceptance
2014-03-01
Citation
ACM Transactions on Reconfigurable Technology and Systems, 2015, 8 (2)
ISSN
1936-7414
Publisher
Association for Computing Machinery (ACM)
Journal / Book Title
ACM Transactions on Reconfigurable Technology and Systems
Volume
8
Issue
2
Copyright Statement
© ACM, 2015. This is the author's version of the work. It is posted here by permission of ACM for your personal use. Not for redistribution. The definitive version was published in ACM Transactions on Reconfigurable Technology and Systems, March 2015 https://dx.doi.org/10.1145/2629581
Publication Status
Published