TY - GEN
T1 - Stream Support in MPI Without the Churn
AU - Schuchart, Joseph
AU - Gabriel, Edgar
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2025.
PY - 2025
Y1 - 2025
N2 - Accelerators have become a corner stone of parallel computing, ranging from scientific computing to artificial intelligence. At the application level, accelerators are controlled by submitting work into a stream, from which the work is executed by the hardware. Vendor-specific communication libraries such as NCCL and RCCL have integrated support for submitting communication operations onto a stream to enable ordering of communication and work on streams. It is safe to assume that stream-based computing will remain relevant for the foreseeable future. MPI has yet to catch up to this reality and prior proposals involved extensions of MPI that would incur significant additions to the API. In this work, we explore alternatives that involve only minor additions to the standard to enable the integration of MPI operations with compute stream. Our additions include i) associating streams with communication objects, ii) blocking streams until completion, and iii) synchronizing streams while progressing MPI operations. Our API is agnostic of the type of stream, reuses existing communication procedures and semantics, and enables integration with graph capturing. We provide a proof-of-concept implementation and show that stream integration of MPI operations can be beneficial.
AB - Accelerators have become a corner stone of parallel computing, ranging from scientific computing to artificial intelligence. At the application level, accelerators are controlled by submitting work into a stream, from which the work is executed by the hardware. Vendor-specific communication libraries such as NCCL and RCCL have integrated support for submitting communication operations onto a stream to enable ordering of communication and work on streams. It is safe to assume that stream-based computing will remain relevant for the foreseeable future. MPI has yet to catch up to this reality and prior proposals involved extensions of MPI that would incur significant additions to the API. In this work, we explore alternatives that involve only minor additions to the standard to enable the integration of MPI operations with compute stream. Our additions include i) associating streams with communication objects, ii) blocking streams until completion, and iii) synchronizing streams while progressing MPI operations. Our API is agnostic of the type of stream, reuses existing communication procedures and semantics, and enables integration with graph capturing. We provide a proof-of-concept implementation and show that stream integration of MPI operations can be beneficial.
KW - accelerators
KW - GPU
KW - HIP
KW - MPI
KW - streams
UR - https://www.scopus.com/pages/publications/85206118979
U2 - 10.1007/978-3-031-73370-3_4
DO - 10.1007/978-3-031-73370-3_4
M3 - Conference contribution
AN - SCOPUS:85206118979
SN - 9783031733697
T3 - Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
SP - 56
EP - 72
BT - Recent Advances in the Message Passing Interface - 31st European MPI Users’ Group Meeting, EuroMPI 2024, Proceedings
A2 - Blaas-Schenner, Claudia
A2 - Niethammer, Christoph
A2 - Haas, Tobias
PB - Springer Science and Business Media Deutschland GmbH
T2 - 31st European MPI Users’ Group Meeting, EuroMPI 2024
Y2 - 25 September 2024 through 27 September 2024
ER -