Sustaining Exascale Performance: Lessons from HPL and HPL-MxP on Aurora Kazushige Goto∗ , Huda Ibeid∗ , Kalyan Kumaran§ , Servesh Muralidharan§ , Anthony-Trung Nguyen∗ , Aditya Nishtala∗ ∗ Intel Corporation, Santa Clara, CA, USA § Argonne National Laboratory, Lemont, IL, USA
arXiv:2604.09517v1 [cs.DC] 10 Apr 2026
{kazushige.goto, huda.ibeid, anthony.d.nguyen, aditya.nishtala}@intel.com {kumaran, servesh}@anl.gov
Abstract—Sustaining exascale performance in production requires engineering choices and operational practices that emerge only under real deployment constraints and demand coordination across system layers. This paper reports experience from three successive campaigns running HPL and HPL-MxP on Aurora, an Intel-based exascale system featuring the first large-scale deployment of Intel discrete GPUs, CPU-attached network interfaces, and the largest production Slingshot-11 interconnect. Aurora progressed from 0.585 EF/s on 5,439 nodes to 1.01 EF/s on 9,234 nodes in FP64 HPL, while HPL-MxP reached 11.64 EF/s, an 11.5× speedup over FP64 enabled by mixed-precision arithmetic and Intel AMX acceleration. We identify and classify by role at production scale the system-level choices that sustained these results, including deterministic locality-aware resource mapping, explicit CPU-GPU pipelining, mixed-precision orchestration, and a hybrid P2P/collective resilience strategy introduced after synchronization stalls at scale. While some observations are Auroraspecific, the broader lessons are likely to apply to tightly coupled heterogeneous systems at extreme scale. Index Terms—Exascale computing, HPL, HPL-MxP, Aurora, mixed-precision computing, heterogeneous CPU–GPU systems, performance engineering, large-scale deployment
I. I NTRODUCTION Sustaining exascale performance on a new heterogeneous platform is as much an operational challenge as a computational one. Full-system benchmark runs must navigate hardware variability, transient network events, and a combinatorial space of runtime parameters, all within narrow windows of system availability. When the target platform simultaneously introduces a new GPU architecture, CPU-attached network interfaces, and a multi-NIC node topology without precedent at this scale, established tuning practices from prior systems may not transfer, and practitioners must revisit assumptions about communication, locality, and resource management. Aurora [1], [2], an exascale system deployed at the Argonne Leadership Computing Facility, exemplifies these challenges. Aurora integrates over 10,000 compute nodes, each combining Intel® Xeon® CPU Max Series processors with on-package high-bandwidth memory [3], Intel® Data Center GPU Max Series accelerators interconnected via Xe-Link [4], [5], and the HPE Slingshot-11 interconnect [6]–[8]. This architecture represents the largest production deployment of Slingshot to date and the first large-scale use of Intel’s discrete GPUs, supported by a new software ecosystem based on Intel oneAPI [9]. While
offering substantial performance potential, this heterogeneous design introduces challenges in communication behavior, locality management, CPU–GPU interaction, and robustness at full-system scale. In this work, we report experience from sustaining exascale performance with HPL and HPL-MxP on Aurora. We describe the system-level choices across communication, resource mapping, CPU–GPU overlap, and mixed-precision execution that were required in practice to sustain performance at near fullsystem scale. HPL measures sustained FP64 performance and defines the TOP500 list [10], while HPL-MxP combines reduced-precision arithmetic with iterative refinement to reflect emerging scientific and AI workloads [11]. Together, these benchmarks stress compute, memory, and interconnect subsystems at extreme scale, providing insight into both peak performance and system stability. At near full-system scale, Aurora sustained 1.01 EF/s on FP64 HPL and 11.64 EF/s on mixed-precision HPLMxP across more than 9,000 nodes, corresponding to 78.8% parallel scaling efficiency when normalized to single-node performance. While prior work demonstrated sustained exascale performance on GPU-centric systems such as Frontier [12], [13], Aurora differs in several important architectural respects. In particular, CPUs integrate on-package highbandwidth memory (HBM) [14] and support Advanced Matrix Extensions (AMX) [15]; network interfaces are attached to CPUs rather than GPUs; and each node exposes more network interfaces than accelerators, enabling communication specialization. At system scale, communication is provided by the HPE Slingshot-11 interconnect organized as a highradix Dragonfly topology, which shapes collective behavior and sensitivity to transient network events. This work does not propose new numerical algorithms for LU factorization. Instead, it reports the system-level choices required to sustain performance on Aurora’s heterogeneous architecture and the practical lessons learned from productionscale deployment. While motivated by Aurora’s design, many of the observations extend beyond a single system and may inform performance engineering for a broader class of current and future heterogeneous HPC platforms. This work makes the following contributions: • Reports sustained exascale performance on an Intel-based
heterogeneous system, achieving 1.01 EF/s in FP64 HPL and 11.64 EF/s in mixed-precision HPL-MxP, and documents the progression of results across three deployment campaigns (SC23, ISC24, SC24). • Describes the system-level choices spanning communication behavior, resource mapping, CPU–GPU overlap, and precision management that collectively sustained performance in practice, and classifies each by its estimated impact (Enabling, Correctness, Critical, or Incremental) based on engineering assessment at production scale. • Documents communication strategies that supported robustness and scalability at extreme scale, including a hybrid point-to-point/collective approach, introduced after observing production stalls, for synchronization-critical panel exchanges, and phase-specific NIC partitioning that isolated latency-sensitive from bandwidth-sensitive traffic on multi-NIC nodes. • Discusses the role of NUMA- and PCIe-aware rank, GPU, and NIC placement on CPU-attached network architectures, where misalignment introduces cross-socket UPI traffic, and describes how explicit CPU–GPU pipeline management through multiple OpenCL queues was required to sustain GPU utilization. • Highlights the practical benefit of exploiting BF16 arithmetic on both Aurora CPUs and GPUs, and identifies Intel AMX acceleration as a directly attributable enabler of the ∼10% HPL-MxP improvement between ISC24 and SC24. • Identifies operational challenges and lessons learned from multi-hour production runs at near full-system scale, including network variability, system reliability, and resilience strategies, and discusses which observations are likely Aurora-specific versus generalizable. II. R ELATED W ORK A. HPL Optimization on Exascale Systems As GPU-accelerated systems have become dominant at leadership scale, recent work has focused on restructuring HPL to maximize accelerator residency while mitigating latencysensitive phases and communication bottlenecks. Prior work describes the design and optimization of rocHPL for exascale accelerated nodes (e.g., Frontier), emphasizing GPU-resident trailing updates while retaining CPU panel factorization and pivoting, together with CPU time-sharing and communication-hiding strategies to improve strong scaling [13]. Complementary analyses study HPL performance and tuning on Frontier in depth, highlighting panel factorization and its interaction with synchronization and communication as key limiters of efficiency at extreme scale [16]. B. Mixed-Precision Benchmarks: HPL-AI and HPL-MxP HPL-MxP was introduced as a mixed-precision benchmark that combines low-precision LU factorization with nonstationary iterative refinement based on GMRES, and also addresses scalable input-matrix data generation and scaling effects on backward error [17]. Prior work demonstrates how
FP16 tensor-core arithmetic can accelerate mixed-precision iterative refinement while preserving numerical stability, providing key motivation for mixed-precision LU and refinement–based benchmarks on GPU-accelerated systems [18]. At system scale, prior studies demonstrate that mixedprecision Linpack-style benchmarks can be scaled to systems such as Summit and Frontier, showing that communication, synchronization, and algorithmic coordination increasingly dominate as concurrency grows [19]. Beyond GPU-centric platforms, large-scale mixed-precision Linpack performance has also been demonstrated on CPU-based systems such as Fugaku, emphasizing architecture-aware optimizations and performance considerations distinct from accelerator-focused approaches [20]. In contrast to prior work that primarily focuses on GPUcentric execution models and peak throughput, this paper reports deployment experience on a CPU-attached, multi-NIC exascale architecture. We emphasize practical lessons from host-mediated communication, deterministic NIC mapping, and resilience strategies for sustained performance during long-duration, full-system runs. III. BACKGROUND A. Benchmarks 1) HPL: High Performance Linpack (HPL) measures sustained floating-point performance by solving a dense linear system Ax = b, where A is an N × N double-precision matrix with random entries [21]. HPL has long served as the standard for the TOP500 list [10], which reports both the sustained performance Rmax achieved during execution and the theoretical peak performance Rpeak derived from hardware specifications. Achieving sustained performance requires optimized numerical kernels and efficient utilization of the communication fabric. a) Algorithmic Structure: HPL computes the LU factorization of A with partial pivoting, followed by triangular solves. The operation count is [22] FLOPtotal = 32 N 3 + O(N 2 ), and the sustained performance is measured as 2
N 3 + O(N 2 ) . Runtime For parallel execution, HPL employs a two-dimensional block-cyclic data distribution over a P × Q MPI process grid with block size N B. Each iteration of the algorithm operates on a single panel of width N B, with the following phases [22]: 1) Panel Factorization (PFACT): factorization of the leading panel; latency- and communication-sensitive. 2) Panel Broadcast (BCAST): distribution of the factored panel across the process grid. 3) Row Swapping (SWAP): pivot permutations to ensure numerical stability. 4) Triangular Solve (DTRSM): update of the current panel row using the triangular factor. Rmax = 3
25GB/s
25GB/s
200Gb
200Gb
HPE Cassini NIC 2x Port
25GB/s
200Gb
200Gb
32GB/s
32GB/s 64GB/s
PCIe Switch 64GB/s
64GB/s
Intel Data Center GPU Max Series 64GB Stack 0 HBM2e Stack 1
Stack 1
64GB HBM2e
Intel Data Center GPU Max Series 64GB HBM2e
Stack 0
Stack 1
28GB/s Per Link between Stacks
Intel Data Center GPU Max Series 64GB HBM2e
Stack 0
Stack 1
64GB HBM2e
Intel Data Center GPU Max Series Stack 0
5) Matrix Update (DGEMM): update of the trailing submatrix using dense matrix–matrix multiplications, which dominate the total floating-point cost for large problems. To improve scalability, modern implementations employ pipelining and lookahead, which can often overlap DTRSM with DGEMM, reducing its impact on the critical path. b) Performance Model: At each of the N/N B panel steps, the effective runtime is determined by the slowest operation among compute- and communication-bound phases [22]:
64GB HBM2e
Stack 1
64GB HBM2e
Intel Data Center GPU Max Series Stack 0
64GB HBM2e
Stack 1
64GB HBM2e
64GB HBM2e 64GB HBM2e
64GB/s
64GB/s
64GB/s
64GB/s 64GB/s 32GB/s
32GB/s
32GB/s
HPE Cassini NIC 2x Port 200Gb 200Gb
HPE Cassini NIC 2x Port 200Gb 200Gb
25GB/s
25GB/s
25GB/s
Intel Xeon CPU Max Series 64GB HBM2e
PCIe Switch 32GB/s
64GB HBM2e
UPI 48GB/s
Fig. 1. HPL cost model. Early iterations are dominated by compute-bound DGEMM, while later stages are increasingly limited by communication- and latency-bound operations (PFACT, BCAST, SWAP).
Stack 0
64GB HBM2e
8x35GB/s Intel Xeon CPU Max Series
64GB/s 64GB/s
Intel Data Center GPU Max Series
512GB DDR5
HPE Cassini NIC 2x Port
32GB/s
32GB/s
25GB/s
8x35GB/s
512GB DDR5
25GB/s
2
3
·N B TDGEMM ∼ N RDGEMM ,
Argonne Leadership N ·N B 2Computing Facility
Fig. 2. Aurora Exascale Compute Blade (ECB).
TPFACT ∼ RPFACT ,
·N B TBCAST , TSWAP ∼ N BWnet .
Here, RDGEMM and RPFACT denote the sustained floating-point rates of the corresponding kernels, while BWnet is the effective per-process network bandwidth. The total runtime is then approximated as X Ttotal = max TDGEMM , TPFACT + TBCAST , TSWAP . panels
For sufficiently large problem sizes, DGEMM typically dominates early iterations due to its cubic complexity, while later iterations become limited by communication- and latency-bound phases such as PFACT, BCAST, and SWAP, as illustrated in Figure 1. c) Mapping to Heterogeneous Architectures: On heterogeneous systems such as Aurora, compute-intensive kernels like DGEMM are executed on GPUs. Panel factorizations (PFACT), row swaps (SWAP), triangular solves (TRSM), and MPI coordination are executed on CPUs, which also orchestrate GPU kernel launches and manage overlap. Achieving high sustained performance requires efficient distribution of work across CPUs and GPUs, fine-grained pipelining to overlap computation with communication, and topology-aware process mapping to minimize communication bottlenecks.
2) HPL-MxP: HPL-MxP is a mixed-precision variant of HPL designed to reflect emerging AI and scientific applications [11]. It accelerates the solution of dense linear systems by executing most arithmetic operations in reduced precision, with iterative refinement restoring double-precision accuracy. While the algorithm retains the structure of HPL, factorizations and trailing matrix updates are performed in low precision (FP32, FP16, or BF16), whereas residual evaluation and corrections are computed in double precision (FP64). Pivoting is omitted to reduce communication overhead at scale, with stability restored through iterative refinement. These modifications increase GPU throughput and reduce memory footprint. In our Aurora implementation, matrices are stored in BF16 format, substantially reducing memory requirements and enabling larger problems to fit in GPU memory. Although reduced precision may increase the number of iterative refinement steps, the resulting gains in GPU throughput and utilization outweigh this additional cost. On Aurora, GPUs perform the reduced-precision factorization and trailing matrix updates using BF16-optimized kernels, while CPUs compute residuals and apply refinement in double precision, enabling efficient utilization of both CPUs and GPUs.
S1 S2
Service Rack 8
S3 N8
C1
C1
C1
Service Rack 1
D8
D64
DAOS Dragonfly Group 1
DAOS Dragonfly Group 8
S4
N8
D56
D1
S2
S1
N1
S3
S3 S4
N8
S4
S1
N1
S2
N1
Service Racks 1 - 8 Dragonfly Group 175
D# – Racks with 16 DAOS servers each
C8
Rack 1 Dragonfly Group 1
N[1-8] – Compute Nodes 1 to 8 per chassis
6
C8
Rack 2 Dragonfly Group 2
C[1-8] – Chassis 1 to 8 per rack
Argonne Leadership Computing Facility
S1 S2 S3
N8
1 Local In Rack Link / 28 Per Switch 2 Global Compute To Compute Group Link / 165 Per Compute Group
S4
S1
1 Local In Chassis Link / 4 Per Switch
S3 S4
S3
N8
S4
N8
N1
S2
N1
S2
N1
S1
Dragonfly Group 167 - 174
2 Edge Link / 16 Per Switch
2 Global Compute To Daos Group Link / 8 Per Compute Group Distributed To 8 DAOS Groups
C8
24 Global Links Between Each Daos Group
Rack 166 Dragonfly Group 166
S[1-4] – Rosetta Switches 1 to 4 per chassis
32 Local Links within Each Daos Group 8 Global Links Between Daos And Service Group 2 Global Compute To Service Group Link / 2 Per Compute Group Round Robin To 8 Service Groups
Fig. 3. Aurora’s 1-D Dragonfly interconnect topology.
B. Aurora Architecture Aurora consists of 10,624 compute nodes interconnected with HPE’s Slingshot-11 network [6], organized in a Dragonfly topology designed for scalability and high global bandwidth [7]. Each node, referred to as an Exascale Compute Blade (ECB), integrates CPUs, GPUs, and high-speed network interfaces to support large-scale heterogeneous applications. Figure 2 shows the layout of an ECB. Each blade contains two Intel® Xeon Max Series CPUs (Sapphire Rapids) [3] and six Intel® Data Center GPU Max Series accelerators (Ponte Vecchio) [4], [5], delivering up to 145 TFLOP/s of doubleprecision peak performance per node. Each CPU integrates 64 GB of on-package HBM2e and 512 GB of DDR5 memory [14], and supports Intel’s Advanced Matrix Extensions (AMX) for accelerating low-precision matrix operations [15]. The two CPUs are interconnected via three Intel Ultra Path Interconnect (UPI) links, which maintain cache coherence and support data exchange between the CPUs. Each CPU connects to three GPUs via PCIe Gen5 x16 links, each delivering 64 GB/s of bandwidth. Within the GPU complex, Xe-Link provides an all-to-all interconnect among the six GPUs, with 28 GB/s per link. Each GPU integrates 128 Xe-Cores across two Xe-Stacks, with 64 GB of HBM2e per stack (128 GB total per GPU), and supports FP64, FP32, BF16, FP16, and INT arithmetic. Network connectivity is provided by eight HPE Cassini NICs, connected through PCIe switches that fan out Gen5 x16 lanes to Gen4 x16 endpoints. Each NIC delivers 200 Gbps and supports adaptive routing for congestion management. At the system level, Aurora’s 10,624 compute nodes are organized into 166 compute groups, along with eight I/O groups
and one service group. Each compute group corresponds to a cabinet housing 32 switches in an all-to-all configuration. Global connectivity between compute groups is provided by two dedicated global links per pair, forming a one-dimensional Dragonfly topology. Figure 3 shows the network layout. The interconnect delivers 2.12 PB/s of aggregate injection bandwidth and 0.69 PB/s of global bisection bandwidth. This architecture, with heterogeneous memory systems, tightly integrated CPU–GPU interconnects, and a highperformance network fabric, offers substantial performance potential while presenting significant challenges for achieving sustained exascale efficiency. IV. E XPERIMENTAL S ETUP A. Software Environment All benchmark experiments are conducted on Aurora using a pre-production software stack that includes Intel’s oneAPI [9] and an Aurora-tuned MPICH [23]. Dense linear algebra kernels are provided by the Intel oneAPI Math Kernel Library (oneMKL) [24], which offers optimized BLAS routines for both CPU and GPU execution. The HPL implementation is based on Netlib HPL v2.3 [22] with Aurora-specific enhancements, while HPL-MxP employs BF16 matrix storage, FP64 iterative refinement, and mixed-precision optimizations. B. Benchmark Parameters Table I summarizes the configurations used for the largestscale runs. Parameters such as block size (N B) and process grid dimensions (P × Q) are tuned to balance GPU arithmetic intensity, communication overhead, and overlap between CPU and GPU tasks. For both HPL and HPL-MxP, we select
the largest N B values that fit in GPU memory. Because HPL-MxP employs BF16 matrix storage rather than FP64, it enables block sizes up to four times larger than HPL. Problem sizes (N ) are chosen to saturate available memory, while process grid dimensions are selected to maintain balanced communication patterns at scale. HPL employs 6 MPI ranks per node (PPN=6), with each rank bound to a GPU/NIC pair, while HPL-MxP employs 2 MPI ranks per node (PPN=2), with each rank managing 3 GPUs through a thread-based runtime. This mapping reflects the distinct communication and computation characteristics of the two benchmarks. Further discussion of tuning trade-offs is provided in Section V. TABLE I B ENCHMARK CONFIGURATIONS FOR LARGE - SCALE RUNS ON AURORA . Parameter Nodes Processes per node (PPN) Process grid P × Q Block size N B Problem size N Precision Look-ahead depth
HPL 9,234 6 162 × 342 384 28,773,888 FP64 1
HPL-MxP 9,500 2 152 × 125 1,536 57,693,696 BF16 + FP32 + FP64 1
C. Performance Metrics We report performance using parallel scaling efficiency, which captures how effectively the system scales with increasing node count. Parallel scaling efficiency is defined as ηscale =
Rmax (N ) , N · Rmax (1)
where Rmax (N ) is the achieved performance on N nodes and Rmax (1) is the achieved performance on a single node. This metric reflects the efficiency with which performance scales as the system size increases, independent of theoretical peak hardware capability. V. S YSTEM -L EVEL C HOICES FOR S USTAINED P ERFORMANCE Sustaining exascale performance on Aurora required coordinating decisions across multiple system layers: communication behavior, resource mapping, CPU–GPU interaction, and numerical execution. Aurora’s architecture, with CPU-attached NICs, multiple network interfaces per node, heterogeneous memory, and a new GPU software stack, makes these decisions particularly consequential, as misalignment at any layer can limit throughput or destabilize multi-hour runs. This section describes the configurations and strategies that proved important in practice, organized by system layer. These factors interact and were tuned jointly during production-scale deployment on 9,500+ nodes, where fullsystem allocations competed with acceptance testing and science workloads. As is common in exascale deployment practice [12], [13], we report the validated configuration and the reasoning behind each choice, rather than isolated singlevariable ablations.
A. Algorithmic–Level Choices Algorithmic decisions determine arithmetic intensity, memory footprint, and the balance between CPU and GPU workloads. In practice, the most consequential choices centered on storage format, block size, and process mapping. HPL is restricted to FP64 storage, which constrains memory capacity, while HPL-MxP permits reduced-precision storage. In both cases, the problem size is chosen to saturate memory, and the MPI process grid is tuned to balance communication across Aurora’s Slingshot-11 topology (Table I). In HPL, CPUs perform panel factorization (PFACT), row swaps (SWAP), MPI communication, and control logic, while GPUs accelerate DGEMM computations. Intel’s OpenCL runtime manages GPU kernel launches and CPU–GPU data transfers, enabling pipelined overlap of CPU-managed tasks with GPU kernels. CPU operations rely on single-threaded Intel oneMKL BLAS, while GPU computations use a modified oneMKL DGEMM kernel tuned for Aurora. HPL-MxP retains this overall structure but replaces FP64 computation with a mixed-precision approach in which CPU-based iterative refinement restores FP64 accuracy. The storage format directly shapes feasible block sizes (N B) and overall problem size (N ) because both are bounded by GPU memory capacity. With FP64 storage, HPL quickly exhausts memory, forcing smaller N B values and limiting N . HPL-MxP instead stores the matrix in BF16, reducing the footprint by up to 4×. This relaxation enables substantially larger block sizes and nearly 2× larger problems. Larger blocks increased arithmetic intensity, while reduced-precision computation delivered higher throughput on both GPUs and CPUs, sustaining performance at scale. The two benchmarks adopt different process layouts to balance computation and communication. HPL uses 6 MPI ranks per node (PPN=6), each bound to a dedicated GPU/NIC pair, maximizing per-rank bandwidth and reducing contention. This configuration supported communication-intensive panel factorizations alongside GPU updates and provided high aggregate bandwidth. In contrast, HPL-MxP uses 2 MPI ranks per node (PPN=2), with each rank managing 3 GPUs through a threadbased runtime. With larger block sizes and no pivoting, HPLMxP incurs lower communication overhead but increases the cost of CPU-based TRSM, as larger blocks make TRSM solves more expensive. A coarser process mapping alleviated this bottleneck by reducing CPU contention and improving TRSM throughput, while the reduced communication requirements of HPL-MxP made the bandwidth trade-off acceptable. B. Software and Runtime Configuration Software and runtime choices span both system-level configuration and application-level communication strategies. At the system level, MPI configuration and transport-layer parameters shaped the efficiency, robustness, and stability of collective communication at extreme node counts. At the application level, resilience mechanisms mitigated transient network instabilities that disrupted multi-hour runs at scale.
1) MPI Configuration and Capabilities: MPI runtime configuration is a key factor in scalability. Aurora’s MPI environment is based on the open-source MPICH [23], optimized for the HPE Slingshot/Cassini interconnect [25]. This implementation offers a broad set of capabilities for extreme-scale performance, including: Collectives: High-radix algorithms [26] reduce latency and improve scaling. • Concurrency: Multi-threaded progress supports both point-to-point and collective communication [27]. • GPU-aware communication: Transfers of GPU-resident buffers via Level Zero IPC over XeLinks, direct NIC access to GPU memory (GPU Direct RDMA), and GDRCopy for low-latency small messages. • Datatype handling: Yaksa [28] enables efficient packing and unpacking of non-contiguous buffers. • Scalable startup: PMIx [29] combined with HPE Cray Parallel Application Launch (PALS) [30].
•
These capabilities establish a flexible baseline for Aurora applications. In HPL and HPL-MxP, however, all inter-rank communication uses host-resident buffers managed by CPUs. Consequently, GPU-aware MPI features (Level Zero IPC, GPU Direct RDMA, GDRCopy) are disabled. This choice reduced overhead and improved robustness and scaling efficiency in our large-scale runs. 2) Transport Layer Configuration – libfabric/cxi: Beyond MPI, transport-layer configuration strongly influences scalability. Aurora employs the libfabric communication stack with the cxi provider for its Slingshot-11 interconnect. At scale, performance depends on how the transport layer is tuned to handle high message rates and overall communication volume. Through iterative tuning at scale, we arrived at a configuration organized around three principles: provision sufficient queue resources, enforce deterministic mapping, and minimize monitoring overhead. Following these principles, the runtime configuration is adjusted as follows. Provisioning resources: Ensured sufficient transport capacity for collective-heavy workloads. – FI_PROVIDER=cxi selects the CXI provider. – FI_CXI_DEFAULT_CQ_SIZE=131072 enlarges completion queues to prevent CQ overflows during high-volume collective operations. • Enforcing deterministic mapping: Provided stable NIC assignment and minimized run-to-run jitter. – MPIR_CVAR_ODD_EVEN_CLIQUES=1 improves logical rank ordering for balanced utilization. – MPIR_CVAR_CH4_OFI_ENABLE_STRIPING=0 disables endpoint striping. – MPIR_CVAR_CH4_OFI_ENABLE_HASHING=0 disables hashing-based endpoint selection. • Reducing monitoring overhead: Removed unnecessary polling and registration costs when GPU-aware MPI is not enabled. •
– FI_MR_CACHE_MONITOR=memhooks leverages lightweight memory hooks for MR cache coherence. – FI_MR_ZE_CACHE_MONITOR_ENABLED=0 disables Level Zero MR monitoring. – MPIR_CVAR_ENABLE_GPU=0 explicitly confirms host-only communication. Enlarging completion queues avoided overflow during long multi-hour runs, deterministic NIC mapping reduced jitter, and disabling GPU-related monitoring eliminated unnecessary polling and registration costs. Collectively, these settings supported low-latency, low-variability communication at full system scale. 3) Communication Resilience – Handling Flap Events: At near full-system scale, transient network events such as link flaps and localized congestion can trigger MPI-level timeouts and retries. Large-scale validation studies on Aurora have shown that these events are unavoidable at near full-system scale, particularly for collective-heavy workloads [8]. Since collectives such as MPI_Allreduce enforce strict global synchronization, a delay on a single rank can stall an entire process column. To mitigate this, the implementation is extended with an alternative reduction path based on point-to-point (P2P) exchanges. The first panel column is communicated via direct P2P transfers rather than MPI_Allreduce, since this column is most sensitive to synchronization. Subsequent columns revert to optimized collectives to preserve efficiency. This hybrid approach allowed incremental progress while limiting the propagation of flap-induced delays. Although P2P requires additional communication rounds, the modest overhead was outweighed by improved resilience in practice, sustaining throughput during multi-hour runs. In our experience, this was a prerequisite for sustained exascale performance. Without it, single-rank instabilities could compromise global progress. C. Hardware–Aware Configuration Several choices exploited Aurora’s architectural features, targeting NUMA locality, PCIe transfer concurrency between CPUs and GPUs, NIC assignment, and Intel AMX instructions for accelerating low-precision compute. 1) NUMA–Aware Process Mapping: Each MPI process was mapped such that its CPU cores, GPUs, NIC, and memory allocations were co-located on the same CPU, preserving locality across computation, communication, and memory accesses. Each Aurora CPU exposes two NUMA nodes, one backed by HBM2e and the other by DDR memory. For largescale HPL runs, the global matrix resides in DDR due to capacity constraints, as problem sizes exceed the available HBM capacity. DDR is therefore used as the primary memory for CPU-managed data structures and bulk storage. HBM2e is used during panel factorization, where tiles required for TRSM and row-exchange operations are staged from DDR into HBM to benefit from higher bandwidth. While this supports efficient execution of CPU-side panel operations, overall runtime is dominated by GPU-accelerated GEMM.
TABLE II S UMMARY OF SYSTEM - LEVEL CHOICES , OBSERVED ROLES , AND ESTIMATED IMPACT CLASSIFICATION . Category Algorithmic design MPI configuration Transport layer Resilience NUMA locality CPU–GPU overlap NIC assignment CPU acceleration
Configuration Choice BF16 storage with iterative refinement Block size and grid tuning Disable GPU-aware MPI Enlarged CQs and deterministic NIC mapping Hybrid P2P/collective strategy Rank, memory, GPU, and NIC pinning Overlapped D2H/H2D transfers Phase-specific NIC mapping AMX-accelerated BF16 GEMM
Observed Role at Scale Reduced memory footprint; enabled ∼ 2× larger problems Supported GPU throughput and overlap Reduced overhead and variability Avoided CQ overflow; reduced jitter Sustained throughput during flap events Reduced remote memory traffic Hid PCIe latency; sustained GPU utilization Reduced contention; increased bandwidth Improved CPU-side refinement efficiency
Impact Class Enabling Critical Incremental Correctness Correctness Critical Critical Incremental Critical
Impact classes reflect engineering assessment, not controlled ablation. Enabling: required to run the benchmark in its intended mode. Correctness: runs fail or hang without this configuration. Critical: expected to have large performance impact based on phase analysis. Incremental: expected to have modest performance impact.
CPU affinity is enforced with taskset or numactl, memory placement with numactl (e.g., --membind), and GPU binding with device-affinity masks (e.g., ZE_AFFINITY_MASK). This mapping ensures that CPU execution, GPU operations, memory accesses, and NIC traffic remain within the same CPU-local memory domain, which reduces UPI traffic between CPUs. By minimizing remote memory access and intra-node contention, this configuration supported lower latency and more predictable performance for CPU-managed operations in HPL and HPL-MxP. 2) Overlapping CPU–GPU PCIe Data Transfers with Computation: On Aurora, HPL and HPL-MxP are executed as a pipelined sequence of GPU kernels, device-to-host (D2H) transfers, CPU computation with MPI communication, host-todevice (H2D) transfers, and subsequent GPU kernels. Without careful scheduling, PCIe latency would stall this pipeline and reduce GPU utilization. To mitigate this, the benchmarks explicitly overlap CPU–GPU data transfers with computation. This strategy exploits the availability of multiple DMA copy engines on Aurora GPUs, allowing communication and computation to proceed in parallel. Concurrency is achieved through explicit queue management. Within each process, the benchmarks manage OpenCL command queues directly. Each queue issues a single operation at a time, and concurrency is enabled by issuing independent transfers and kernels across multiple compute and copy queues. Panel data is transferred from GPU to CPU (D2H) in parallel across multiple copy engines, while CPU-toGPU (H2D) updates are concurrently scheduled on available copy engines to sustain bidirectional throughput. OpenCL events synchronize these operations so that kernels execute on already-resident data while upcoming transfers progress in the background. This design sustained a continuous execution pipeline where GPUs remained highly utilized, PCIe latency was effectively hidden, and computation proceeded without idle periods. In practice, overlapping transfers with computation was essential for mitigating PCIe bottlenecks at scale. 3) Phase–Specific NIC Assignment: Each Aurora node includes eight NICs (four per CPU). To reduce contention and preserve locality, HPL and HPL-MxP assign NICs to com-
munication phases in a phase-specific manner. Row swapping operations (SWAP), which are latency-sensitive and occur on every panel factorization, are bound to the first NIC (NIC0) on each CPU, while broadcasts (BCAST), which are network bandwidth-sensitive, are distributed across the remaining three NICs (NICs 1–3). This separation isolated SWAP traffic from BCAST traffic, reducing contention within the same CPU. Binding is enforced at communicator creation through MPICH multi_nic_pref_nic hints, which provide explicit control over NIC selection. This mapping supported consistent NIC usage across runs, reduced intra-CPU contention, and improved performance robustness. At scale, isolating SWAP to a dedicated NIC and distributing BCAST across the others increased effective bandwidth, lowered variability, and sustained communication throughput during the most communication-intensive phases of HPL and HPL-MxP. 4) Exploiting Intel AMX for Low-Precision Compute: Aurora’s Xeon Max CPUs support Advanced Matrix Extensions (AMX), which provide hardware acceleration for lowprecision matrix operations such as BF16 GEMMs [15]. In HPL-MxP, AMX is leveraged in two contexts. It accelerates CPU-based iterative refinement and low-precision panel factorization. Both stages benefit from higher throughput on GEMM operations, reducing time in CPU-managed phases that would otherwise limit scalability. Intel oneMKL transparently dispatches AMX-accelerated kernels where supported, ensuring that applications exploit these hardware features without requiring code modifications. At scale, AMX acceleration improved efficiency in CPU-side computation and helped sustain mixed-precision performance across thousands of nodes. D. Lessons Learned Across algorithmic, software, runtime, and hardware layers, several practical lessons emerged from our deployment experience. While some are specific to Aurora’s architecture, others reflect broader challenges likely to arise on any heterogeneous exascale system. We summarize these below, noting where our observations appear to generalize and where they may be system-specific. Communication benefits from application-aware tuning. Aurora’s MPICH and libfabric/CXI stack provides well-tuned
defaults that work well for a wide range of applications. However, at 9,000+ nodes with collective-heavy workloads, we found that application-specific tuning made a meaningful difference. Disabling GPU-aware MPI (since HPL and HPLMxP communicate through host buffers), enlarging completion queues, and enforcing deterministic NIC mappings collectively reduced overhead and variability. The specific settings are Aurora/Slingshot-11-specific. However, the general principle is likely to apply on other large-scale systems as well: collective-heavy applications at extreme scale often benefit from explicit transport and NIC configuration rather than relying on defaults. Resilience requires proactive, application-level strategies. Multi-hour runs on 9,000+ nodes are exposed to transient network events (flap events, localized congestion, timeoutinduced retries) that are rare at small scale but become routine at full system size. Standard MPI collectives propagate these delays globally: a single slow rank can stall an entire process column. We addressed this with a hybrid strategy that uses point-to-point exchanges for the most synchronizationsensitive panel column and reverts to optimized collectives for subsequent columns. This was not part of the original design but was introduced after observing stalls during production runs. The specific mechanism is tied to HPL’s panel broadcast structure, but the broader lesson, that applications running multi-hour jobs at extreme scale need explicit resilience strategies beyond what the MPI runtime provides, applies to any tightly coupled exascale workload. Locality-aware resource mapping is essential, especially on CPU-attached network architectures. On Aurora, network interfaces are attached to CPUs rather than GPUs, and each node exposes eight NICs across two CPU sockets. Misalignment between rank placement, GPU affinity, memory binding, and NIC assignment introduces cross-socket traffic over UPI and can destabilize performance at scale. We found that colocating each MPI rank’s CPU cores, GPU, memory, and NIC within the same NUMA domain was necessary for predictable scaling. The severity of this effect is likely architecturedependent; systems with GPU-attached NICs face different locality tradeoffs. However, the general principle that resource mapping decisions become more consequential at extreme scale holds broadly. Explicit CPU–GPU overlap is necessary to sustain throughput. On Aurora, HPL and HPL-MxP execute as a pipelined sequence of GPU kernels, PCIe transfers, CPU computation, and MPI communication. Without explicit management of this pipeline through multiple OpenCL command queues and copy engines, PCIe latency would stall GPU execution. We found that achieving high GPU utilization required careful scheduling of D2H and H2D transfers in parallel with computation, rather than relying on runtime-level overlap. This observation is not unique to Aurora; any system where CPU-managed communication shares the critical path with GPU computation will face similar pipeline management challenges. Mixed precision delivers large gains but requires end-toend orchestration. Switching from FP64 to BF16 storage
enabled an 11.5 × speedup (1.01 EF/s to 11.64 EF/s), but realizing this gain required coordinating block sizes, process grids, CPU–GPU workload balance, and iterative refinement parameters. For example, larger block sizes enabled by BF16 storage increased arithmetic intensity but also increased CPUside TRSM cost, requiring a coarser process grid (PPN=2 vs. PPN=6 for HPL) to rebalance. Exploiting Intel AMX on CPUs to accelerate BF16 GEMMs in refinement and low-precision factorization was a key enabler between ISC24 (10.6 EF/s) and SC24 (11.64 EF/s). The Aurora-specific detail is AMX and the particular CPU–GPU work split. However, the broader lesson is increasingly relevant as mixed-precision methods become standard in both scientific computing and AI: mixed-precision acceleration at scale requires tuning across the full software stack, not just swapping data types. While these lessons are drawn from HPL and HPL-MxP on Aurora, several extend to other workloads and platforms. Communication resilience and topology-aware mapping are relevant to any collective-heavy application at extreme scale, including stencil codes and graph analytics. CPU–GPU overlap and locality tuning apply to workloads with pipeline parallelism, and mixed-precision orchestration reflects trends in scientific AI where reduced-precision acceleration is balanced with high-precision refinement. We do not claim that every observation transfers directly, but the recurring theme is likely to hold broadly: sustained exascale performance requires coordinating decisions across system layers, not optimizing any single layer in isolation. Table II summarizes the systemlevel choices and their observed roles at scale. VI. R ESULTS AND D EPLOYMENT E XPERIENCE This section reports sustained HPL and HPL-MxP performance on Aurora at near full-system scale and discusses the operational challenges encountered during production runs. Table III summarizes the progression of results across deployment campaigns. We present scaling behavior and phaselevel analysis for FP64 HPL and mixed-precision HPL-MxP, followed by a discussion of the system-level challenges that shaped our deployment experience. TABLE III D EPLOYMENT PROGRESSION OF HPL AND HPL-M X P ON AURORA . Campaign SC23 ISC24 SC24
HPL Nodes EF/s 5,439 0.585 9,234 1.01 9,234 1.01
HPL-MxP Nodes EF/s – – 9,234 10.6 9,500 11.64
Key Change Initial large-scale Full-scale HPL + MxP AMX acceleration
A. HPL: Sustained FP64 Exascale Performance Aurora achieved 1.012 EF/s on 9,234 nodes, corresponding to 78.8% parallel scaling efficiency relative to single-node performance. Compared to earlier large-scale experiments, efficiency remained stable as the system scaled from 5,439 to 9,234 nodes, spanning both SC23 and ISC24 milestones. As shown in Figure 4, Aurora sustained high per-node throughput even at near full system size.
1e9
600
1.4
500
Performance (GF/s)
1.2
400 Time (ms)
1.0 0.8
300
0.6
200
0.4 0.2
100 00
Fig. 4. HPL performance on Aurora, scaling from 5,439 to 9,234 nodes with sustained exascale performance.
The phase breakdown of the 9,234-node exascale run is shown in Figure 5. The x-axis shows progress through the LU factorization, while the left y-axis reports per-phase time in milliseconds and the right y-axis reports performance in GF/s. GEMM (purple) dominates the early portion of the run, keeping the benchmark compute-bound during the initial stages. As factorization progresses and the trailing submatrix shrinks, per-step GEMM time decreases. SWAP (green) begins to dominate runtime once GEMM drops, marking the transition from compute-bound to communication-bound execution. The system transitions from compute-bound to communication-bound execution when the GEMM curve falls below the SWAP curve. This crossover aligns with the drop in sustained performance indicated by the dashed orange line. The process grid parameters (P, Q) were chosen to extend the GEMM-dominated region, keeping the run compute-bound as long as possible. This tuning, together with the communication and runtime choices described in Section V, sustained efficiency and reduced the impact of communication bottlenecks. The phase breakdown also provides indirect evidence of relative impact among the system-level choices described in Section V. SWAP accounts for a substantial fraction of per-step time once the run transitions from compute-bound to communication-bound execution, establishing an upper bound on the system-level impact of transport-layer and NICassignment tuning. Conversely, GEMM dominance during the early, highest-throughput phase confirms that GPU kernel efficiency and CPU–GPU overlap are the primary determinants of peak sustained performance. Between ISC24 and SC24, AMX enablement was the primary configuration change, corresponding to a ∼10% improvement in sustained HPL-MxP performance (10.6 to 11.64 EF/s, Table III), making it the one factor whose contribution can be directly attributed. B. HPL-MxP: Mixed-Precision Acceleration Aurora sustained 11.64 EF/s on 9,500 nodes with HPLMxP, representing an 11.5× speedup over FP64 HPL. At ISC24, Aurora delivered 10.6 EF/s, and by SC24, further
Pfact Bcast Swap
20
TRSM GEMM Perf
40 60 Matrix fraction
80
0.0 100
Fig. 5. Phase breakdown of the ∼1 EF/s HPL run on 9,234 nodes. GEMM (purple) dominates in the early stages, while SWAP (green) becomes dominant as the trailing submatrix shrinks. The system transitions from compute-bound to communication-bound execution when GEMM time falls below SWAP, which aligns with the decline in the estimated performance line (dashed orange).
tuning improved sustained performance to 11.64 EF/s. A key enabler between ISC24 and SC24 was the utilization of AMX units on the CPUs, which accelerated iterative refinement and reduced overhead while maintaining solver accuracy. The LU factorization employed mixed-precision arithmetic (BF16/FP32) on GPUs to increase arithmetic intensity, followed by FP64 iterative refinement on CPUs to restore accuracy. AMX acceleration improved the throughput of CPU-side GEMMs in both refinement and low-precision factorization, reducing the time spent in CPU-managed work. As shown in Figure 6, the per-step performance profile exhibits a brief warm-up transient at the start, sustains high performance across the early stages, and then declines gradually as the trailing submatrix shrinks and communication becomes a larger fraction of the work.
Fig. 6. HPL-MxP performance on Aurora at 9,500 nodes. The curve shows a short warm-up followed by a long high-performance plateau and a gradual decline as the factorization progresses, with a final result of 11.64 EF/s, representing an order-of-magnitude speedup over FP64 HPL.
Together, these results confirm that Aurora sustained exascale throughput in both FP64 HPL and mixed-precision
HPL-MxP. The system-level choices described in Section V, collectively supported stable performance across hours-long production runs at near full system scale. Although systemwide reliability ultimately depends on underlying hardware and job management infrastructure, Aurora consistently delivered exascale performance when stable resources were available. C. Operational Challenges at Scale Despite sustained exascale performance, large-scale runs on Aurora faced challenges beyond application-level choices. These challenges stem from the extreme scale of the system, where millions of components must operate in synchrony for hours. Two factors were especially consequential in our experience: system reliability, which determined whether jobs could complete without interruption, and network variability, which influenced how predictably performance was delivered across thousands of nodes. 1) System Reliability: At exascale, reliability is strained by the sheer number of components operating concurrently. With over 63,000 GPUs and 21,000 CPUs, Aurora processes more than 350,000 unique failure events per week from sources including machine checks, PCIe errors, GPU driver events, and firmware logs [31]. Historical data shows that mean time between failures decreases steadily with system scale, from days on petascale systems to hours on current exascale platforms [31]. Even with robust hardware, faults such as node failures, memory errors, or storage issues become more probable during hours-long production runs. For tightly coupled applications like HPL and HPL-MxP, a single fault can cascade to job termination. Failures on Aurora span permanent, transient, and intermittent categories [31]. Intermittent faults, stemming from aged or degraded components, are particularly challenging because they are difficult to distinguish from transient events unless they repeatedly occur on the same component. Beyond individual component failures, correlated events caused by power fluctuations, cooling issues, or filesystem disruptions can affect thousands of nodes simultaneously. Aurora addresses these risks with hardware-level RAS features, continuous monitoring, and an automated failure management framework that reduced mean time to repair by 84× compared to manual servicing during acceptance testing [31]. The system employs fine-grained multi-strike repair policies driven by statistical properties of failure reoccurrence, enabling automated recovery for the majority of failure signatures. The job scheduler and system management software further isolate failing components and recycle healthy resources, improving overall availability. In our deployment experience, successful full-scale runs required pre-screening nodes and scheduling during periods of system stability, as not every attempt completed successfully. Nevertheless, when millions of devices interact at scale, occasional failures are unavoidable. While resilience mechanisms reduce their frequency and impact, the risk of node loss or
job termination remained an inherent challenge for sustained exascale runs. 2) Network Variability: At near full system scale, HPL and HPL-MxP runs were sensitive to transient network events. Aurora’s Slingshot-11 interconnect provides adaptive routing and congestion management, which generally sustained high throughput across thousands of nodes. However, brief instabilities such as flap events or localized congestion manifested as timeouts and retries at the MPI layer. When amplified by collective operations, these disruptions appeared as short stalls or increased variability during communication-bound phases. To mitigate such effects, large-scale jobs are preceded by validation tests, including MPI collectives (e.g., MPI_Alltoall) and GPCNet [32], to identify lowperforming nodes or unstable links. In production runs, hybrid collective/P2P communication strategies (Section V-B) reduced sensitivity to localized disruptions by limiting the scope of their impact. Despite these measures, network variability remained an inherent challenge at exascale. While application- and systemlevel strategies reduced its impact, occasional stalls and throughput fluctuations were unavoidable when millions of components communicated synchronously for hours at a time. VII. C ONCLUSION This paper reported experience from sustaining exascale performance with HPL and HPL-MxP on Aurora, an Intelbased heterogeneous system featuring the first large-scale deployment of Intel discrete GPUs, CPU-attached networking, and the largest production deployment of the Slingshot-11 interconnect. At near full-system scale, Aurora sustained 1.01 EF/s in FP64 HPL on 9,234 nodes and 11.64 EF/s in mixedprecision HPL-MxP on 9,500 nodes, with 78.8% parallel scaling efficiency relative to single-node performance. The 11.5× speedup from FP64 to mixed precision reflected the combined effect of BF16 storage, reduced-precision computation on both CPUs and GPUs, and iterative refinement, with Intel AMX acceleration on CPUs serving as a key enabler between ISC24 (10.6 EF/s) and SC24 (11.64 EF/s). Rather than proposing new numerical algorithms, this work described the system-level choices that collectively sustained performance in practice: communication tuning and resilience strategies for multi-hour runs, deterministic localityaware resource mapping across CPUs, GPUs, and NICs, explicit CPU–GPU pipelining, and mixed-precision orchestration. These choices were not developed in isolation but tuned jointly during production-scale deployment, where interactions across system layers shaped the final configuration. Our deployment experience highlighted operational challenges that are routine at exascale but invisible at moderate scale: transient network events, system reliability constraints, and run-to-run variability. These required both applicationlevel mitigations (such as hybrid collective/P2P communication) and operational practices (such as pre-run node validation and scheduling during periods of system stability).
While some observations are specific to Aurora’s architecture, particularly its CPU-attached NICs, multi-NIC nodes, and AMX-capable processors, several lessons appear broadly applicable. Communication resilience, locality-aware resource mapping, explicit CPU–GPU overlap, and end-to-end mixedprecision orchestration are likely to remain important on any tightly coupled heterogeneous system at extreme scale. We hope this experience report provides practical guidance for performance engineering on current and future heterogeneous HPC platforms. ACKNOWLEDGMENT The authors would like to thank colleagues for their valuable contributions to this work. Specific names will be added following the completion of the double-blind review process. AI-generated text assistance was used during the preparation of this manuscript. Specifically, an AI writing assistant (Anthropic Claude) was used to help refine and improve the clarity and presentation of the text in the paper. All technical content, experimental results, and scientific claims were produced entirely by the authors. The authors reviewed and verified all AI-assisted text for accuracy. R EFERENCES [1] B. S. Allen, J. Anchell, V. Anisimov, T. Applencourt, A. Bagusetty, R. Balakrishnan, R. Balin, S. Bekele, C. Bertoni, C. Blackworth, R. Bustamante, K. Canada, J. Carrier, C. Chan-nui, L. C. Cheney, T. Childers, P. Coffman, S. Coghlan, M. D’Mello, M. Emani, K. G. Felker, S. Foreman, O. Franza, L. Gao, M. Garcı́a, M. Garzarán, B. Gerofi, Y. Ghadar, N. Gupta, K. Harms, V. Hatanpää, B. Holland, C. Holohan, B. Homerding, K. Hossain, L. Huot, H. Ibeid, J. A. Insley, S. Jayanthi, H. Jiang, W. Jiang, X.-Y. Jin, J. Kim, C. Knight, K. Kumaran, J. Kwack, T. Leggett, B. Lenard, C. Lewis, N. Liber, J. Lombardi, R. M. Loy, Y. Luo, B. Lusch, N. Mahadevan, V. A. Mateevitsi, G. McPheeters, R. Milner, V. A. Morozov, S. Muralidharan, T. Musta, M. Nagar, V. Narayana, M. Ngom, A.-T. Nguyen, N. Nichols, A. Nishtala, J. C. Osborn, M. E. Papka, S. Parker, S. S. Patel, A. C. Pope, S. Raghunanda, E. Rangel, P. M. Rich, S. Rizzi, K. Rowe, V. Sastry, A. Scovel, F. Simini, H. S. Som, P. Steinbrecher, R. Stevens, X. Tian, P. Upton, T. Uram, A. K. Vasan, Álvaro Vázquez-Mayagoitia, K. Velusamy, B. Videau, V. Vishwanath, B. Whitney, T. J. Williams, M. Woodacre, S. Zeltner, G. Zheng, and H. Zheng, “Aurora: Architecting argonne’s first exascale supercomputer for accelerated scientific discovery,” 2025. [Online]. Available: https://arxiv.org/abs/2509.08207 [2] Aurora. [Online]. Available: https://www.alcf.anl.gov/aurora [3] N. Nassif, A. O. Munch, C. L. Molnar, G. Pasdast, S. V. Lyer, Z. Yang, O. Mendoza, M. Huddart, S. Venkataraman, S. Kandula, R. Marom, A. M. Kern, B. Bowhill, D. R. Mulvihill, S. Nimmagadda, V. Kalidindi, J. Krause, M. M. Haq, R. Sharma, and K. Duda, “Sapphire rapids: The next-generation intel xeon scalable processor,” in 2022 IEEE International Solid-State Circuits Conference (ISSCC), vol. 65, 2022, pp. 44–46. [Online]. Available: https://doi.org/10.1109/ISSCC42614.2022.9731107 [4] W. Gomes, A. Koker, P. Stover, D. Ingerly, S. Siers, S. Venkataraman, C. Pelto, T. Shah, A. Rao, F. O’Mahony, E. Karl, L. Cheney, I. Rajwani, H. Jain, R. Cortez, A. Chandrasekhar, B. Kanthi, and R. Koduri, “Ponte vecchio: A multi-tile 3d stacked processor for exascale computing,” in 2022 IEEE International Solid-State Circuits Conference (ISSCC), vol. 65, 2022, pp. 42–44. [Online]. Available: https://10.1109/ISSCC42614.2022.9731673 [5] D. Blythe, “ XeHPC Ponte Vecchio ,” in 2021 IEEE Hot Chips 33 Symposium (HCS). Los Alamitos, CA, USA: IEEE Computer Society, Aug. 2021, pp. 1–34. [Online]. Available: https://doi.ieeecomputersociety.org/10.1109/HCS52781.2021.9567038
[6] “Hpc slingshot launched into network space,” in Cray User Group 2022 (CUG2022) Proceedings, May 2022. [Online]. Available: https://cug. org/proceedings/cug2022 proceedings/includes/files/pap121s2-file1.pdf [7] J. Kim, W. J. Dally, S. Scott, and D. Abts, “Technology-driven, highly-scalable dragonfly topology,” in 2008 International Symposium on Computer Architecture, 2008, pp. 77–88. [Online]. Available: https://doi.com/10.1109/ISCA.2008.19 [8] H. Ibeid, A.-T. Nguyen, A. Nishtala, P. Sakarda, L. Kaplan, N. Mahadevan, M. Woodacre, V. Anisimov, K. Kumaran, J. Kwack, V. Morozov, S. Muralidharan, and S. Parker, “Scaling mpi applications on aurora,” 2025. [Online]. Available: https://arxiv.org/abs/2512.04291 [9] Intel oneAPI. [Online]. Available: https://www.intel.com/content/www/ us/en/developer/tools/oneapi/overview.html [10] TOP500 list. [Online]. Available: https://www.top500.org/lists/top500/ [11] HPL-MxP. [Online]. Available: https://hpl-mxp.org/results.md [12] S. Atchley, C. Zimmer, J. Lange, D. Bernholdt, V. Melesse Vergara, T. Beck, M. Brim, R. Budiardja, S. Chandrasekaran, M. Eisenbach, T. Evans, M. Ezell, N. Frontiere, A. Georgiadou, J. Glenski, P. Grete, S. Hamilton, J. Holmen, A. Huebl, D. Jacobson, W. Joubert, K. Mcmahon, E. Merzari, S. Moore, A. Myers, S. Nichols, S. Oral, T. Papatheodore, D. Perez, D. M. Rogers, E. Schneider, J.-L. Vay, and P. K. Yeung, “Frontier: Exploring exascale,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, ser. SC ’23. New York, NY, USA: Association for Computing Machinery, 2023. [Online]. Available: https://doi.org/10.1145/3581784.3607089 [13] N. Chalmers, J. Kurzak, D. Mcdougall, and P. Bauman, “Optimizing high-performance linpack for exascale accelerated architectures,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, ser. SC ’23. New York, NY, USA: Association for Computing Machinery, 2023. [Online]. Available: https://doi.org/10.1145/3581784.3607066 [14] H. Ibeid, V. Narayana, J. Kim, A. Nguyen, V. Morozov, and Y. Luo, “Performance analysis of hpc applications on the aurora supercomputer: Exploring the impact of hbm-enabled intel xeon max cpus,” pp. 1–11, 2025. [Online]. Available: https://doi.org/10.23919/ISC.2025.11018301 [15] Optimizing machine learning (ml) models with intel® advanced matrix extensions (intel® amx). [Online]. Available: https://www.intel.com/content/dam/www/central-libraries/us/en/ documents/2022-12/optimizing-ml-models-with-amx-brief.pdf [16] H. Lu, M. Matheson, N. Chalmers, A. Kashi, N. Malaya, and F. Wang, “Insights from optimizing hpl performance on exascale systems: A comparative analysis of panel factorization,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, ser. SC ’25. New York, NY, USA: Association for Computing Machinery, 2025, p. 397–410. [Online]. Available: https://doi.org/10.1145/3712285.3759875 [17] J. Dongarra and P. Luszczek, “Hpl-mxp benchmark: Mixedprecision algorithms, iterative refinement, and scalable data generation,” The International Journal of High Performance Computing Applications, Sep. 2025. [Online]. Available: http://dx.doi.org/10.1177/10943420251382476 [18] A. Haidar, S. Tomov, J. Dongarra, and N. J. Higham, “Harnessing gpu tensor cores for fast fp16 arithmetic to speed up mixed-precision iterative refinement solvers,” in SC18: International Conference for High Performance Computing, Networking, Storage and Analysis, 2018, pp. 603–613. [Online]. Available: https://doi.com/10.1109/SC.2018.00050 [19] H. Lu, M. Matheson, V. Oles, A. Ellis, W. Joubert, and F. Wang, “Climbing the summit and pushing the frontier of mixed precision benchmarks at extreme scale,” in SC22: International Conference for High Performance Computing, Networking, Storage and Analysis, 2022, pp. 1–15. [Online]. Available: https://doi.com/10.1109/SC41404.2022. 00083 [20] S. Kudo, K. Nitadori, T. Ina, and T. Imamura, “ Prompt Report on ExaScale HPL-AI Benchmark ,” in 2020 IEEE International Conference on Cluster Computing (CLUSTER). Los Alamitos, CA, USA: IEEE Computer Society, Sep. 2020, pp. 418–419. [Online]. Available: https: //doi.ieeecomputersociety.org/10.1109/CLUSTER49012.2020.00058 [21] J. J. Dongarra, P. Luszczek, and A. Petitet, “The linpack benchmark: past, present and future,” Concurrency and Computation: practice and experience, vol. 15, no. 9, pp. 803–820, 2003. [Online]. Available: https://doi.org/10.1002/cpe.728 [22] A. Petitet, R. C. Whaley, J. Dongarra, and J. Cleary, “HPL - a portable implementation of the high-performance linpack benchmark
for distributed-memory computers,” University of Tennessee, Tech. Rep. UT-CS-01-448, 2001. [Online]. Available: http://www.netlib.org/ benchmark/hpl/ [23] Y. Guo, K. Raffenetti, H. Zhou, P. Balaji, M. Si, A. Amer, S. Iwasaki, S. Seo, G. Congiu, R. Latham, L. Oden, T. Gillis, R. Zambre, K. Ouyang, C. Archer, W. Bland, J. Jose, S. Sur, H. Fujita, D. Durnov, M. Chuvelev, G. Zheng, A. Brooks, S. Thapaliya, T. Doodi, M. Garzaran, S. Oyanagi, M. Snir, and R. Thakur, “Preparing MPICH for exascale,” The International Journal of High Performance Computing Applications, vol. 39, no. 2, pp. 283–305, 2025. [Online]. Available: https://doi.org/10.1177/10943420241311608 [24] Intel(R) oneAPI Math Kernel Library (oneMKL). [Online]. Available: https://www.intel.com/content/www/us/en/developer/tools/oneapi/ onemkl.html [25] K. Raffenetti, A. Amer, L. Oden, C. Archer, W. Bland, H. Fujita, Y. Guo, T. Janjusic, D. Durnov, M. Blocksome, M. Si, S. Seo, A. Langer, G. Zheng, M. Takagi, P. Coffman, J. Jose, S. Sur, A. Sannikov, S. Oblomov, M. Chuvelev, M. Hatanaka, X. Zhao, P. Fischer, T. Rathnayake, M. Otten, M. Min, and P. Balaji, “Why is mpi so slow? analyzing the fundamental limits in implementing mpi-3.1,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, ser. SC ’17. New York, NY, USA: Association for Computing Machinery, 2017. [Online]. Available: https://doi.org/10.1145/3126908.3126963 [26] T. Doodi, N. Islam, G. Zheng, R. Kalidas, A. Langer, and G. Maria, “High radix collective algorithms,” in Proceedings of EuroMPI, 2021. [Online]. Available: https://doi.org/10.1007/978-3-031-29927-8 31 [27] R. Zambre, A. Chandramowliswharan, and P. Balaji, “How I learned to stop worrying about user-visible endpoints and love MPI,” in Proceedings of the 34th ACM International Conference on Supercomputing, ser. ICS ’20. New York, NY, USA: Association for Computing Machinery, 2020. [Online]. Available: https://doi.org/10. 1145/3392717.3392773 [28] High-Performance Data Type Engine. [Online]. Available: https: //www.yaksa.org [29] Application programming interface for exascale systems. [Online]. Available: https://pmix.github.io/ [30] HPE Parallel Application Launch Service (PALS). [Online]. Available: https://support.hpe.com/hpesc/public/docDisplay?docId= a00117940en us&page=Parallel Application Launch Service PALS. html&docLocale=en US [31] Y. Levitt, R. Barella, S. Zeltner, T. Musta, L. Cheney, G. Espinosa, O. Franza, and B. Gerofi, “Fine-grained automated failure management for extreme-scale gpu accelerated systems,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, ser. SC ’25. New York, NY, USA: Association for Computing Machinery, 2025, p. 1073–1084. [Online]. Available: https://doi.org/10.1145/3712285.3759883 [32] S. Chunduri, T. Groves, P. Mendygral, B. Austin, J. Balma, K. Kandalla, K. Kumaran, G. Lockwood, S. Parker, S. Warren, N. Wichmann, and N. Wright, “Gpcnet: designing a benchmark suite for inducing and measuring contention in hpc networks,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, ser. SC ’19. New York, NY, USA: Association for Computing Machinery, 2019. [Online]. Available: https://doi.org/10.1145/3295500.3356215