Design and Initial Performance of a High-level Unstructured Mesh Framework on Heterogeneous Parallel Systems
File(s) Parallel Computing_39_11_2013.pdf (1.05 MB)
Accepted version
Author(s)
Type
Journal Article
Abstract
OP2 is a high-level domain specific library framework for the solution of unstructured mesh-based applications. It utilizes source-to-source translation and compilation so that a single application code written using the OP2 API can be transformed into multiple parallel implementations for execution on a range of back-end hardware platforms. In this paper we present the design and performance of OP2’s recent developments facilitating code generation and execution on distributed memory heterogeneous systems. OP2 targets the solution of numerical problems based on static unstructured meshes. We discuss the main design issues in parallelizing this class of applications. These include handling data dependencies in accessing indirectly referenced data and design considerations in generating code for execution on a cluster of multi-threaded CPUs and GPUs. Two representative CFD applications, written using the OP2 framework, are utilized to provide a contrasting benchmarking and performance analysis study on a number of heterogeneous systems including a large scale Cray XE6 system and a large GPU cluster. A range of performance metrics are benchmarked including runtime, scalability, achieved compute and bandwidth performance, runtime bottlenecks and systems energy consumption. We demonstrate that an application written once at a high-level using the OP2 API is easily portable across a wide range of contrasting platforms and is capable of achieving near-optimal performance without the intervention of the domain application programmer.
Editor(s)
Hollingsworth, J
Date Issued
2013-09-24
Citation
Parallel Computing, 2013, n/a, pp.n/a-
ISSN
0167-8191
Publisher
Elsevier
Start Page
n/a
End Page
692
Journal / Book Title
Parallel Computing
Volume
n/a
Issue
11
Copyright Statement
Copyright © 2013 Elsevier. All rights reserved. NOTICE: this is the author’s version of a work that was accepted for publication in Parallel Computing. Changes resulting from the publishing process, such as peer review, editing, corrections, structural formatting, and other quality control mechanisms may not be reflected in this document. Changes may have been made to this work since it was submitted for publication. A definitive version was subsequently published in Parallel Computing, 39(11), 2013. DOI:10.1016/j.parco.2013.09.004.
Identifier
PII: S0167-8191(13)00116-6
Subjects
OP2; Domain Specific Language; Active Library; Unstructured mesh; GPU; Heterogeneous systems; Energy consumption of parallel systems
Publication Status
Accepted
