Precise event sampling-based data locality tools for AMD multicore architectures
File(s)2023-03-29-AsAccepted-CPE-22-2441.R1_Proof_hi.pdf (1.59 MB)
Accepted version
Author(s)
Sasongko, Muhammad Aditya
Chabbi, Milind
Kelly, Paul
Unat, Didem
Type
Journal Article
Abstract
We propose COMDETECTIVE+, an inter-thread communication analyzer, and
REUSETRACKER+, a reuse distance analyzer, that leverage the hardware features
in AMD processors to support low-overhead profiling. Both tools employ the
instruction-based sampling (IBS) facility and debug registers in AMD processors
to detect inter-thread communication and data reuse. Different from prior arts,
COMDETECTIVE+ differentiates the communication into true and false sharing, and
REUSETRACKER+ measures reuse distance in private and shared caches by also considering cache line invalidation with low overhead. Both tools can attribute the
communications and reuses to source code lines. To our knowledge these tools are two
of the few profiling tools designed specifically for AMD x86 architectures using IBS.
Our tools are timely and relevant considering the rise in numbers of AMD processor
based data centers and HPC systems. We perform experiments to evaluate the accuracy and overheads of the proposed tools on an AMD machine with two-socket EPYC
7352 processors. COMDETECTIVE+ exhibits high accuracy while introducing 5.14×
runtime and 1.4× memory overheads. REUSETRACKER+ also displays high accuracy,
which is 95%, with 11.76×runtime and 1.46× memory overheads. These overheads are
much lower than the overheads of existing simulators and code instrumentation-based
tools. Lastly, we demonstrate the usage of the tools by having COMDETECTIVE+ and
REUSETRACKER+ facilitate the code refactoring of two data mining benchmarks to
improve their performance by up to 29%.
REUSETRACKER+, a reuse distance analyzer, that leverage the hardware features
in AMD processors to support low-overhead profiling. Both tools employ the
instruction-based sampling (IBS) facility and debug registers in AMD processors
to detect inter-thread communication and data reuse. Different from prior arts,
COMDETECTIVE+ differentiates the communication into true and false sharing, and
REUSETRACKER+ measures reuse distance in private and shared caches by also considering cache line invalidation with low overhead. Both tools can attribute the
communications and reuses to source code lines. To our knowledge these tools are two
of the few profiling tools designed specifically for AMD x86 architectures using IBS.
Our tools are timely and relevant considering the rise in numbers of AMD processor
based data centers and HPC systems. We perform experiments to evaluate the accuracy and overheads of the proposed tools on an AMD machine with two-socket EPYC
7352 processors. COMDETECTIVE+ exhibits high accuracy while introducing 5.14×
runtime and 1.4× memory overheads. REUSETRACKER+ also displays high accuracy,
which is 95%, with 11.76×runtime and 1.46× memory overheads. These overheads are
much lower than the overheads of existing simulators and code instrumentation-based
tools. Lastly, we demonstrate the usage of the tools by having COMDETECTIVE+ and
REUSETRACKER+ facilitate the code refactoring of two data mining benchmarks to
improve their performance by up to 29%.
Date Issued
2023-11-01
Date Acceptance
2023-03-14
Citation
Concurrency and Computation: Practice and Experience, 2023, 35 (24), pp.1-18
ISSN
1532-0626
Publisher
Wiley
Start Page
1
End Page
18
Journal / Book Title
Concurrency and Computation: Practice and Experience
Volume
35
Issue
24
Copyright Statement
Copyright © 2023 Owner. This is the peer reviewed version of the following article: : Sasongko MA, Chabbi M, Kelly PHJ, Unat D. Precise event sampling-based data locality tools for AMD
multicore architectures. Concurrency Computat Pract Exper. 2023;e7707. doi: 10.1002/cpe.7707, which has been published in final form at https://doi.org/10.1002/cpe.7707. This article may be used for non-commercial purposes in accordance with Wiley Terms and Conditions for Use of Self-Archived Versions. This article may not be enhanced, enriched or otherwise transformed into a derivative work, without express permission from Wiley or by statutory rights under applicable legislation. Copyright notices must not be removed, obscured or modified. The article must be linked to Wiley’s version of record on Wiley Online Library and any embedding, framing or otherwise making available the article or pages thereof by third parties from platforms, services and websites other than Wiley Online Library must be prohibited.
multicore architectures. Concurrency Computat Pract Exper. 2023;e7707. doi: 10.1002/cpe.7707, which has been published in final form at https://doi.org/10.1002/cpe.7707. This article may be used for non-commercial purposes in accordance with Wiley Terms and Conditions for Use of Self-Archived Versions. This article may not be enhanced, enriched or otherwise transformed into a derivative work, without express permission from Wiley or by statutory rights under applicable legislation. Copyright notices must not be removed, obscured or modified. The article must be linked to Wiley’s version of record on Wiley Online Library and any embedding, framing or otherwise making available the article or pages thereof by third parties from platforms, services and websites other than Wiley Online Library must be prohibited.
Identifier
https://onlinelibrary.wiley.com/doi/full/10.1002/cpe.7707
Publication Status
Published
Article Number
ARTN e7707
Date Publish Online
2023-04-03