TY - GEN
T1 - A64FX performance
T2 - 2021 IEEE International Conference on Cluster Computing, Cluster 2021
AU - Bari, Md Abdullah Shahneous
AU - Chapman, Barbara
AU - Curtis, Anthony
AU - Harrison, Robert J.
AU - Siegmann, Eva
AU - Simakov, Nikolay A.
AU - Jones, Matthew D.
N1 - Publisher Copyright:
©2021 IEEE.
PY - 2021
Y1 - 2021
N2 - We examine the performance of scientific and engineering kernels on the Fujitsu A64FX processor, both outof-the-box using various toolchains and with processor-specific optimizations. While nearly all applications port with little to no modification, significant performance variation is observed between the multiple tool chains. This variation depends heavily upon characteristics of the application (most notably its use of mathematical functions) and is also constrained by the most performant toolchains having limited support for recent language standards. As expected, high performance demands that a kernel is vectorized, multi-threaded, and localizes memory references. Detailed optimizations, including use of intrinsics, are also examined to understand performance gaps and what is necessary to attain peak performance. This article employs the Ookami computer technology testbed funded by the American National Science Foundation. The system provides researchers worldwide with access to 176 Fujitsu A64FX compute nodes as well as other state-of the-art technology.
AB - We examine the performance of scientific and engineering kernels on the Fujitsu A64FX processor, both outof-the-box using various toolchains and with processor-specific optimizations. While nearly all applications port with little to no modification, significant performance variation is observed between the multiple tool chains. This variation depends heavily upon characteristics of the application (most notably its use of mathematical functions) and is also constrained by the most performant toolchains having limited support for recent language standards. As expected, high performance demands that a kernel is vectorized, multi-threaded, and localizes memory references. Detailed optimizations, including use of intrinsics, are also examined to understand performance gaps and what is necessary to attain peak performance. This article employs the Ookami computer technology testbed funded by the American National Science Foundation. The system provides researchers worldwide with access to 176 Fujitsu A64FX compute nodes as well as other state-of the-art technology.
KW - High-performance computing
UR - https://www.scopus.com/pages/publications/85124041963
U2 - 10.1109/Cluster48925.2021.00106
DO - 10.1109/Cluster48925.2021.00106
M3 - Conference contribution
AN - SCOPUS:85124041963
T3 - Proceedings - IEEE International Conference on Cluster Computing, ICCC
SP - 711
EP - 718
BT - Proceedings - 2021 IEEE International Conference on Cluster Computing, Cluster 2021
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 7 September 2021 through 10 September 2021
ER -