Toward Better Simulation of MPI Applications on Ethernet/TCP Networks
File(s)bedaride_pmbs2013.pdf (927.68 KB)
Accepted version
Author(s)
Type
Conference Paper
Abstract
imulation and modeling for performance predic-
tion and profiling is essential for developing and maintaining
HPC code that is expected to scale for next-generation exascale
systems, and correctly modeling network behavior is essential for
creating realistic simulations. In this article we describe an im-
plementation of a flow-based hybrid network model that accounts
for factors such as network topology and contention, which are
commonly ignored by other approaches. We focus on large-scale,
Ethernet-connected systems, as these currently compose 37.8%
of the TOP500 index, and this share is expected to increase
as higher-speed 10 and 100GbE become more available. The
European Mont-Blanc project to study exascale computing by de-
veloping prototype systems with low-power embedded devices will
also use Ethernet-based interconnect. Our model is implemented
within SMPI, an open-source MPI implementation that connects
real applications to the SimGrid simulation framework. SMPI
provides implementations of collective communications based
on current versions of both OpenMPI and MPICH. SMPI and
SimGrid also provide methods for easing the simulation of large-
scale systems, including shadow execution, memory folding, and
support for both online and offline (i.e., post-mortem) simulation.
We validate our proposed model by comparing traces produced
by SMPI with those from real world experiments, as well as
with those obtained using other established network models.
Our study shows that SMPI has a consistently better predictive
power than classical LogP-based models for a wide range of
scenarios including both established HPC benchmarks and real
applications.
tion and profiling is essential for developing and maintaining
HPC code that is expected to scale for next-generation exascale
systems, and correctly modeling network behavior is essential for
creating realistic simulations. In this article we describe an im-
plementation of a flow-based hybrid network model that accounts
for factors such as network topology and contention, which are
commonly ignored by other approaches. We focus on large-scale,
Ethernet-connected systems, as these currently compose 37.8%
of the TOP500 index, and this share is expected to increase
as higher-speed 10 and 100GbE become more available. The
European Mont-Blanc project to study exascale computing by de-
veloping prototype systems with low-power embedded devices will
also use Ethernet-based interconnect. Our model is implemented
within SMPI, an open-source MPI implementation that connects
real applications to the SimGrid simulation framework. SMPI
provides implementations of collective communications based
on current versions of both OpenMPI and MPICH. SMPI and
SimGrid also provide methods for easing the simulation of large-
scale systems, including shadow execution, memory folding, and
support for both online and offline (i.e., post-mortem) simulation.
We validate our proposed model by comparing traces produced
by SMPI with those from real world experiments, as well as
with those obtained using other established network models.
Our study shows that SMPI has a consistently better predictive
power than classical LogP-based models for a wide range of
scenarios including both established HPC benchmarks and real
applications.
Date Issued
2014-09-30
Citation
Lecture Notes in Computer Science, 2014, (8551), pp.158-181
Publisher
Springer
Start Page
158
End Page
181
Journal / Book Title
Lecture Notes in Computer Science
Issue
8551
Copyright Statement
© Springer International Publishing Switzerland 2014. The final publication is available at Springer via http://dx.doi.org/10.1007/978-3-319-10214-6_8
Description
16.01.15 KB. Ok to add accepted version to spiral, subject to 12 months embargo
Identifier
http://hal.in2p3.fr/file/index/docid/919507/filename/smpi_pmbs13.pdf
Source
4th International Workshop on Performance Modeling, Benchmarking and Simulation of High Performance Computing Systems
Start Date
2013-11-17
Finish Date
2013-11-21
Coverage Spatial
Denver, Colorado, USA