TY - GEN
T1 - Experiences with Porting the FLASH Code to Ookami, an HPE Apollo 80 A64FX Platform
AU - Feldman, Catherine
AU - Michalowicz, Benjamin
AU - Siegmann, Eva
AU - Curtis, Tony
AU - Calder, Alan
AU - Harrison, Robert
N1 - Publisher Copyright:
© 2022 ACM.
PY - 2022/1/11
Y1 - 2022/1/11
N2 - We present initial experiences with running the community simulation code FLASH, developed at the University of Chicago for multi-scale multi-physics applications, on Ookami, a technology testbed featuring the A64FX processor developed by Fujitsu. Our effort focused largely on running FLASH "right out of the box"to see which combinations of compilers and software implementations (e.g. MPI) allowed the code to run with minimal modification. FLASH was one application in a larger effort to deploy Ookami; it served as a test for different versions of newly installed software, and as a cornerstone for the FAQ page of the Ookami website. We report on our results with different compilers and other software, along with our initial scaling results and attempts to utilize the A64FX's SVE instructions and NUMA architecture. We found that FLASH readily ran with different compilers and MPI implementations, and showed the expected good scaling with no turning. However, more work must be done to fully take advantage of the A64FX's architectural features and produce a significant speedup for FLASH on Ookami.
AB - We present initial experiences with running the community simulation code FLASH, developed at the University of Chicago for multi-scale multi-physics applications, on Ookami, a technology testbed featuring the A64FX processor developed by Fujitsu. Our effort focused largely on running FLASH "right out of the box"to see which combinations of compilers and software implementations (e.g. MPI) allowed the code to run with minimal modification. FLASH was one application in a larger effort to deploy Ookami; it served as a test for different versions of newly installed software, and as a cornerstone for the FAQ page of the Ookami website. We report on our results with different compilers and other software, along with our initial scaling results and attempts to utilize the A64FX's SVE instructions and NUMA architecture. We found that FLASH readily ran with different compilers and MPI implementations, and showed the expected good scaling with no turning. However, more work must be done to fully take advantage of the A64FX's architectural features and produce a significant speedup for FLASH on Ookami.
KW - A64FX
KW - astrophysics
KW - high-performance computing
KW - porting software
UR - https://www.scopus.com/pages/publications/85124032589
U2 - 10.1145/3503470.3503478
DO - 10.1145/3503470.3503478
M3 - Conference contribution
AN - SCOPUS:85124032589
T3 - ACM International Conference Proceeding Series
SP - 72
EP - 77
BT - Proceedings of International Conference on High Performance Computing in Asia-Pacific Region Workshops, HPCAsia 2022
PB - Association for Computing Machinery
T2 - 2022 International Conference on High Performance Computing in Asia-Pacific Region Workshops, HPCAsia 2022
Y2 - 11 January 2022 through 14 January 2022
ER -