ConceptioArchivearXiv CS
arXiv CSopen access

Characterizing the Impact of Congestion in Modern HPC Interconnects

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
clouddistributedcomputingparallelcomputing
distributed computing, parallel computing, cloud

Characterizing the Impact of Congestion in Modern HPC Interconnects Lorenzo Piarulli∗ , Marco Faltelli† , Dirk Pleiter‡ , Karthee Sivalingam‡§ , Dancheng Zhang§ , Kexue Zhao§ , Matteo Turisini¶ , Francesco Iannone† , Aldo Artigiani§ , Daniele De Sensi∗ ∗ Sapienza University of Rome, {piarulli, desensi}@di.uniroma1.it † ENEA, {marco.faltelli, francesco.iannone}@enea.it ‡ OEHI, University of Groningen, [email protected]

arXiv:2604.11432v1 [cs.DC] 13 Apr 2026

§ Huawei, {karthee.sivalingam, zhangdancheng, zhaokexue, aldo.artigiani}@huawei.com ¶ CINECA, [email protected]

Abstract—High-performance computing (HPC) systems increasingly support both scalable AI training and large-scale simulation workloads. Both typically rely heavily on collective communication operations. On modern supercomputers, however, network congestion has emerged as a major limitation, driven by heterogeneous traffic patterns resulting from diverse workload mixes. As system scale and active users continue to grow, understanding how today’s interconnect technologies respond to congestion is essential for establishing realistic performance expectations and informing future system design. This paper presents a comprehensive characterization of congestion behavior across four major HPC fabrics: EDR InfiniBand, HDR InfiniBand, NDR InfiniBand, Cray Slingshot, and emerging Ethernet fabrics. These fabrics span high-performance proprietary interconnects as well as adaptive Ethernet-based designs aligned with emerging standards such as Ultra Ethernet. We evaluate their responses to both steady congestion and a wide range of bursty patterns that vary in duration, intensity, and pause length, capturing the bursty communication typical of AI workloads. Our study covers multiple scales, examining how congestion manifests differently as system size increases and identifying scale-dependent behaviors that influence collective performance. By analyzing the challenges that arise under these controlled stress conditions, we aim to provide a practical overview of congestion issues and possible optimizations. The insights derived from this evaluation can guide researchers and HPC architects in designing more effective congestion-control mechanisms and network load-balancing strategies. Index Terms—congestion control, load balance, collective operations, MPI

I. I NTRODUCTION Supercomputers are rapidly increasing in both scale and performance, enabling new advances in scientific computing and artificial intelligence. As system sizes grow and user demand increases, multiple jobs are often executed concurrently, with workloads running either on dedicated or shared compute nodes. Moreover, with the growth of the number of endpoints, new network topologies started to become popular that can be expected to be more susceptible to congestion. Regardless of the allocation strategy, all applications ultimately rely on a shared communication infrastructure. When several jobs simultaneously traverse common network components,

links may become oversubscribed and switch buffers can fill, leading to network congestion. Since communication is a critical phase of both HPC and AI workloads, congestion directly translates into performance degradation and reduced system efficiency. Although congestion can affect any HPC system, its manifestation and impact depend on several factors, including network topology, system scale, traffic patterns, communication primitives, fabric technology, and the congestion-mitigation mechanisms available. To address these challenges, a wide range of solutions has been proposed. These include end-host–based and network-level congestion control mechanisms that detect congestion signals and throttle injection rates, as well as other network-level techniques, such as adaptive routing and load balancing, that mitigate congestion. The problem is further complicated by dynamic, multitenant workloads and allocation policies optimized for overall system utilization. As a result, network resources are frequently shared, and congestion must be detected and mitigated at runtime. Modern congestion-control and load-balancing mechanisms attempt to address this challenge by dynamically reallocating flows, adjusting routing decisions, or reducing traffic injection rates at the sources, thereby improving traffic distribution and limiting packet loss. Therefore, the ability to effectively mitigate congestion is tightly coupled with the interconnect technology deployed by an HPC facility. The more capable a technology is at handling diverse congestion scenarios, the better the system can support concurrent multi-tenant workloads with heterogeneous communication patterns. However, modern supercomputing environments rely on a wide variety of interconnects, each with distinct architectural designs, routing strategies, and congestion-management mechanisms. In recent years, alongside the strong performance and widespread adoption of NVIDIA’s InfiniBand, several Ethernet-based solutions have gained increasing attention. This trend has been driven both by the emergence of the Ultra Ethernet Consortium [1] and by the demonstrated effectiveness of Ethernet-based technologies such as HPE/Cray’s

Slingshot, which now account for a significant fraction of the aggregate performance of systems listed in the TOP500 ranking [2]. Given the diversity of interconnect technologies and congestion-management solutions deployed in modern supercomputing systems, it is essential to systematically examine their limitations, characterize congestion effects, and correlate these effects with specific architectural and protocollevel features. A comprehensive and comparative analysis is therefore required to understand how different fabrics behave under realistic congestion conditions. In this work, we analyze multiple state-of-the-art interconnect technologies, including EDR, HDR, and NDR InfiniBand, Cray Slingshot, and Network Scale Load Balance (NSLB) enabled Ethernet fabrics. We evaluate these technologies under a range of controlled congestion scenarios generated using collective communication patterns with both persistent and bursty traffic intensities, designed to emulate realistic HPC and AI workloads. Our evaluation spans both production-scale supercomputers and research testbeds. Specifically, we analyze CINECA’s Leonardo, ENEA’s CRESCO8, and LUMI systems at both small and large scales, and we additionally evaluate emerging Network Scale Load Balance NSLB-capable Ethernet fabrics deployed in the HAICGU and Nanjing laboratory clusters. Through this evaluation, we characterize the behavior of modern supercomputing networks under congestion, identify their distinguishing features, and assess the effectiveness of their congestion-mitigation mechanisms. II. BACKGROUND Congestion manifests in distinct forms, each requiring a specific mitigation strategy. As messages traverse multiple switches between source and destination, the physical location of the bottleneck dictates the appropriate response. When congestion develops at an intermediate switch, load balancing (or adaptive routing) can effectively alleviate the pressure by rerouting traffic along alternative paths to avoid resource collisions. Conversely, if congestion occurs at an edge switch directly connected to the endpoint, load balancing becomes ineffective because all possible network paths must eventually converge on that specific switch to reach the destination. In such scenarios, congestion control mechanisms are essential to throttle the injection rate at the source. These dynamics are closely tied to specific communication patterns: AlltoAll and permutation communications frequently stress intermediate switches due to global bandwidth demands, whereas Incast patterns, where multiple nodes simultaneously target a single receiver, typically lead to congestion at the edge. In our analysis, we consider five different systems (described in Table I) based on the three main existing interconnect technologies: Ethernet, InfiniBand, and Slingshot, which we describe in the following. A. Ethernet Ethernet’s popularity in large-scale clusters stems from its open, standards-based ecosystem, which fosters vendor competition, ensures broad interoperability, and significantly

reduces the total cost of ownership. By avoiding vendor lockin and leveraging mature management tools, organizations can scale infrastructure while closing the performance gap with specialized fabrics through technologies like RDMA over Converged Ethernet (RoCE) and the Ultra Ethernet Consortium (UEC) [1]. Specifically, the Ultra Ethernet Specification v1.0 [3] introduces the Ultra Ethernet Transport (UET), which brings HPC-grade capabilities, such as multipath packet spraying, out-of-order delivery, and local link-layer retransmissions, to the standard IP ecosystem to mitigate tail latency. Congestion Control: While traditionally a ”best-effort” medium, Ethernet has evolved to support the lossless operation required by RDMA through Priority Flow Control (PFC), which prevents packet loss by pausing specific traffic classes at the link level, though it can trigger deadlock or congestion spreading [4]. To mitigate these risks, modern fabrics employ proactive congestion control typically based on Explicit Congestion Notification (ECN), where switches mark IP headers upon exceeding queue thresholds so that the source can throttle its injection rate before buffers overflow. Building on this, the congestion control algorithm Data Center Quantized Congestion Notification (DCQCN) couples ECN marking with Quantized Congestion Notification (QCN) style feedback to rapidly restore fairness [4], while TIMELY utilizes fine-grained RTT gradients to bound queuing delay without switch-side marking [5]. More advanced schemes like HPCC leverage in-network telemetry to compute near-optimal rate updates, achieving fast convergence under large Incast [6]. Within the Ultra Ethernet framework, congestion control leverages both receiver- and sender-based algorithms [7], and the Huawei Cloud Engine switches (CE8850 and CE9855) analyzed in this work implement AI ECN to dynamically adjust ECN thresholds based on the workload [8]. Load Balancing: Ethernet routing strategies range from congestion-oblivious methods like Equal-Cost Multi-Path (ECMP) [9], which frequently suffers from hash collisions and link underutilization [10]–[13], to congestion-aware dynamic adaptation. While schemes such as MPTCP [14], PLB [15], FlowBender [16], Flowlet Switching [17], Flowcell [18], and Flowcut [19] improve efficiency by splitting traffic into subflows or flowlets, they often introduce state overhead and potential packet reordering [20]. More granular packet-level approaches like OPS [21] and MPRDMA [20] maximize path utilization through spraying but require robust out-of-order delivery handling. To address these challenges, the Huawei Cloud Engine switches examined in this work utilize Network Scale Load Balance (NSLB), a decentralized, real-time mechanism integrated into the network controller [22]. By employing a flow matrix to compute collision-free uplink assignments for each (source edge, destination edge) pair, NSLB ensures that concurrent flows from the same source are distributed across distinct uplinks, effectively minimizing collisions and optimizing fabric throughput [22].

B. InfiniBand InfiniBand is a tightly integrated interconnect fabric that has evolved over multiple generations to provide the high bandwidth and low latency required for traditional HPC and AI workloads. While its specialized architecture offers superior performance metrics, it can lead to vendor lock-in and lacks the broad interoperability found in commodity Ethernet ecosystems. Congestion Control: Unlike Ethernet’s reactive pausebased mechanisms, InfiniBand employs a hardware-level, hopby-hop credit-based flow control that inherently avoids buffer overflows by ensuring a sender only transmits when the receiver has available buffer space. However, this can lead to backpressure that propagates through the fabric, potentially creating ”victim” flows. To mitigate this, InfiniBand utilizes an end-to-end closed-loop mechanism where congested switch ports probabilistically mark packets with Forward Explicit Congestion Notification (FECN). The destination then returns a Backward Explicit Congestion Notification (BECN), or Congestion Notification Packets (CNPs), to the source, which throttles its injection rate via per-flow inter-packet delay. This makes the fabric’s performance highly sensitive to the precise tuning of marking thresholds and recovery rates [23], [24]. Load Balancing: While InfiniBand was historically built on deterministic forwarding, modern implementations provide Adaptive Routing (AR) to dynamically manage link utilization. In an AR-enabled fabric, switches independently select among groups of equivalent egress ports based on real-time port load. This dynamic selection is orchestrated by the Subnet Manager (SM), which is responsible for configuring AR groups and implementing topology-specific routing policies, such as those optimized for fat-trees or Dragonfly+, to maximize bisection bandwidth and minimize contention [25]. C. Slingshot HPE/Cray Slingshot follows a distinct design philosophy by merging the high-performance characteristics of specialized HPC fabrics like InfiniBand with the openness and interoperability of the Ethernet ecosystem. By optimizing RDMA over Ethernet (RoCE), Slingshot achieves performance levels comparable to proprietary interconnects while remaining compatible with standard data-center infrastructure. This approach has catalyzed a broader industry trend where vendors seek to leverage existing Ethernet-based management tools and expertise without sacrificing the low-latency requirements of large-scale AI and HPC clusters. Congestion Control: Slingshot implements a hardwarebased congestion management system designed to provide granular flow isolation and prevent the ”congestion spreading” typical of traditional lossless Ethernet. Unlike Priority Flow Control (PFC), which operates on broad traffic classes and can lead to head-of-line blocking, Slingshot tracks every packet and provides feedback for thousands of individual flows. This allows the fabric to identify and throttle only the specific sources contributing to a bottleneck, ensuring that ”victim” flows remain unaffected. By maintaining low, stable queuing

delays even under heavy Incast conditions, Slingshot provides the performance predictability required for tightly coupled parallel workloads [26], [27]. Load Balancing: Slingshot couples fine-grained adaptive routing with global congestion awareness to optimize path selection, particularly in Dragonfly deployments. Switches dynamically evaluate multiple minimal and non-minimal paths by estimating real-time path congestion and considering path length, allowing for high bisection bandwidth utilization. This mechanism is especially effective in multi-tenant environments, as it can steer traffic away from transient hotspots and minimize tail-latency interference caused by competing workloads. Furthermore, Slingshot can guarantee in-order delivery of packets even in the presence of adaptive routing by draining active flows before changing their path [19]. III. M ETHODOLOGY We developed a custom experimental pipeline that injects controlled congestion while executing collective communication operations and simultaneously benchmarking their performance. The standard way to create congestion is to split the allocated nodes into victims and aggressors, and then measure the victims while the aggressors generate continuous, communication-intensive traffic [26], [28]. We started from this idea, but then developed a more flexible approach capable of producing multiple congestion scenarios. In addition to steady, continuous congestion, we also inject bursty congestion, varying both the burst duration and the idle gap between bursts to emulate realistic, time-varying contention. By implementing various congestion profiles, we can simulate a spectrum of network conditions and analyze the system’s response to each scenario. While the aggressors are running, the victims perform 1,000 iterations of the benchmarked collectives, all of which are recorded. We then discard the first 100 iterations to account for network warmup. The final execution time used to compare congested collectives against uncongested ones is computed by taking the mean of the remaining iterations. The following chapter describes the methods used to carry out this analysis in detail. A. Congestion Injection In our approach, the aggressors continuously launch collectives without interruption in an endless loop, thereby generating communication noise. The victims, instead, perform the benchmarked collectives for a fixed number of iterations. Group creation is performed by interleaving the nodes in increasing order: the first node is assigned to the victims’ group, the second to the aggressors’ group, the third again to the victims’ group, and so on, alternating until all nodes are allocated. This ensures a balanced distribution of aggressors and victims across the allocated set of nodes, maximizing network resources sharing and, thus, congestion. The benchmark is first executed on the victims without congestion to establish a baseline. Then, a congestion job and the benchmark are run in parallel on the aggressors and victims, respectively,

Memory

Communication

AlltoAll Memory Reduction

Communication

AllReduce Fig. 1. Comparison of time distribution between AlltoAll and AllReduce operations.

to measure and analyze the effects of congestion. Attackers are provided with two types of collectives for noise injection: AlltoAll and Incast. The first is used to send as many messages as possible to all nodes, creating a general state of network noise. The latter, instead, focuses the traffic on a single node, attempting to generate congestion at a specific edge switch.

B. Selected Collectives Our analysis is based on collective communication primitives, which are the de facto methodology for distributed computing in HPC systems. These collectives are typically provided by MPI libraries such as NCCL, RCCL, Open MPI and MPICH, based on the system library availability. A wide range of collectives exists, each serving a different purpose: some are purely communication-oriented (e.g., AllGather, AlltoAll, and Broadcast), while others also include computation through reduction operations (e.g., AllReduce and ReduceScatter) Our initial experiments considered the Open MPI [29] AllReduce collective, which showed up to a 25% bandwidth loss compared to AlltoAll [30]. A follow-up analysis with a custom ring AllReduce (separating ReduceScatter and AllGather) indicated that performance was mainly limited by reduction costs and memory handling (initial buffer setup and memcpy operations), rather than by network communication (Figure 1). Since these overheads also affect other collectives and can dominate execution time, we exclude computation-based collectives (AllReduce and ReduceScatter) from the rest of this work. We therefore focus on communication-only collectives (AlltoAll and AllGather) and measure only their communication phase, enabling a cleaner characterization of network behavior and congestion effects. To eliminate the uncertainty introduced by the MPI libraries’ dynamic algorithm selection and ensure a fair comparison across different software stacks, we did not use the default MPI collective implementations. We implemented custom ring AllGather and linear AlltoAll algorithms via standard MPI send/recv primitives. This approach ensured we used the same communication pattern across all systems, since they rely on different MPI libraries (Cray MPI on LUMI, Open MPI on Leonardo/CRESCO8). Moreover, our custom implementation allowed us to remove memory-handling overheads such as malloc and memcpy of temporary buffers, ensuring that measurements focus strictly on network-level communication latency and reducing the impact of differences in node and memory architectures.

C. Steady Congestion Injection The first method injects steady, continuous congestion using the AlltoAll and Incast collectives. By repeatedly executing the same collective pattern over an extended time window, we create a quasi-stationary stress condition in which queue occupancy, link utilization, and contention hotspots remain stable (or oscillate within a narrow range) for most of the run. This “steady-state” regime is useful because it exposes the behavior of routing, arbitration, and flow-control mechanisms once initial transients have decayed. In particular, persistent congestion gives time for closed-loop mechanisms (e.g., ECN/credit-based backpressure and rate adaptation) to converge, allowing us to observe whether the fabric reaches a stable operating point (bounded queues and fair throughput) or instead exhibits pathological effects such as sustained head-ofline blocking, victim-flow throughput collapse, or congestion spreading. The AlltoAll collective primarily exercises transit contention by simultaneously activating many source–destination pairs whose shortest paths overlap on intermediate links and core/aggregation switches. When multiple minimal routes exist, this condition is well suited to evaluate load balancing and adaptive routing, because congestion can often be mitigated by shifting traffic to alternative paths with lower instantaneous occupancy. Conversely, Incast concentrates traffic from many senders onto a single receiver (or a small set of receivers), creating a fan-in bottleneck at the destination NIC and the adjacent edge (ToR) switch. Since the limiting resource is typically the receiver-side egress and buffering rather than a congested transit link, rerouting cannot eliminate the hotspot. As a result, Incast congestion isolates the effectiveness of end-host and switch-level congestion control (rate reduction, marking thresholds, and recovery dynamics), and highlights whether the fabric can prevent queue build-up and backpressure from propagating to unrelated flows. D. Bursty Congestion Injection The second method generates congestion in bursts rather than continuously. Each burst consists of one or more consecutive collective invocations, and we vary both the burst duration (i.e., number of collectives per burst or total burst time) and the inter-burst idle interval. These parameters control how far congestion can develop within device buffers and switch queues before traffic stops, and how much time the network has to drain queues and restore steady throughput before the next burst. By sweeping multiple burst/idle combinations, we capture regimes ranging from “microbursts” that may only transiently fill shallow buffers to longer bursts that can trigger feedback-based throttling and potentially affect subsequent communication phases. Bursty congestion is inherently intermittent and therefore stresses detection and reaction latency: if congestion onset is faster than the control loop (marking, feedback, and sender rate adaptation), the system may only respond after the burst has already completed, leaving residual queueing delay and backpressure that impacts other traffic. Moreover, each burst

effectively reintroduces a transient, forcing the network to repeatedly transition between uncongested and congested states; this exposes overshoot/undershoot behavior, recovery speed, and stability under repeated excitation. This pattern closely resembles communication phases in production HPC and AI workloads, such as distributed deep learning, where gradient aggregation (e.g., AllReduce) follows each optimization step and produces periodic spikes in network utilization rather than a constant load.

Fig. 2. Bursty congestion injection visualization, bursty aggressor on the bottom and victim on the top.

E. Evaluation Environments The fabrics have been evaluated under many different architectures, node counts, and topologies. This allowed us to understand congestion under a variety of scenarios. The systems considered were CINECA’s Leonardo, ENEA’s CRESCO8, LUMI, Huawei AI and Computing at Goethe University (HAICGU), and the HPC lab in Nanjing (Nanjing). Leonardo, CRESCO8, and LUMI are supercomputers ranked in the TOP500 [2], offering large-scale executions and production environments with many active users. On these systems, we cannot fully control job allocations, as they depend on the current usage of the computers. These systems enforce different policies on the maximum allocation size; therefore, we evaluated them at multiple scales during testing. For comparison and continuity, in this paper we report the 256 node results as maximum allocations, since it is a common scale across the platforms and the maximum available on Leonardo. HAICGU and the Nanjing lab, in contrast, are research HPC systems designed for evaluating new architectures. They are small, with a maximum of 10 nodes for HAICGU and 8 nodes for Nanjing. Despite their small size, they provide the ability to test the emerging Ethernet fabrics CE8850 and CE9855, and to compare it with an EDR InfiniBand system composed of the same TaiShan nodes. Leonardo is composed of two partitions: GPU-accelerated (Booster) and CPU-only partition (Data Centric General Purpose) [31]. We use the Booster partition which employs BullSequana X2135 “Da Vinci” nodes featuring a single Intel Xeon Platinum 8358 CPU (32 cores, 2.60 GHz) paired with 4× NVIDIA A100 GPUs (64 GB HBM2e) and 512

GB RAM(8×64 GB DDR4) per node. All compute nodes are interconnected via NVIDIA Mellanox HDR InfiniBand with a Dragonfly+ topology. Every node includes 2 dual port HDR100 boards delivering 400Gb/s per node. CRESCO8 [32] is the new Tier-0 HPC facility at ENEA. This cluster comprises 760 CPU nodes and 17 GPU nodes. For the massive scale purposes of this paper, we chose to run our experiments on the CPU partition. Each CPU node is composed of 2×64 Intel Xeon Platinum 8592+ cores, 512 GB RAM and a dual-port Mellanox-NDR ConnectX-7 NIC that delivers 200 Gb/s. Nodes are connected through a 1.67:1 blocking Fat-Tree topology. LUMI uses HPE Cray EX nodes in a multi-partition architecture based on AMD CPUs and GPUs. We selected the GPU partition, which consists of 2978 GPU-CPU nodes, each equipped with one AMD EPYC “Trento” CPU, four AMD Instinct MI250X accelerators, and 512 GB of system DDR4 memory. The accelerators provide a total of 512 GB of HBM2e memory per node. All nodes are interconnected through the Cray Slingshot fabric using 4×200 Gb/s links in a Dragonfly topology, enabling scalable communication across the system [33]. HAICGU uses TaiShan 200 (Model 2280) nodes with dual Kunpeng 920 CPUs (ARMv8 AArch64, 64 cores, 2.6GHz) and 128GB RAM (16×8GB DIMMs). It is maintained by Open Edge and HPC initiative (OEHI) [34], has two 10node partitions with 2 leaf switches each: one using Mellanox MSB7890-ES2F EDR switches and the other RoCE-based CE8850, both at 100GE. Nanjing uses TaiShan 200 (Model 2280) nodes with dual Kunpeng 920 CPUs (ARMv8 AArch64, 64 cores, 2.6GHz) and 128GB RAM (16×8GB DIMMs). It includes 8 nodes on a RoCE partition built with 4 CE9855 switches with NSLB, arranged in a two-leaf, two-spine 200GE topology. This design allows NSLB to leverage multiple path configurations for load balancing.

F. Analysis Structure To build the results section progressively, we start from the least challenging scenario and move toward the most demanding one. Specifically, we first analyze steady congestion at small scale (Sec.IV), where contention is limited and congestion effects are typically easier to mitigate. Although small-scale setups are not the most common setting in which severe congestion emerges, they are still useful to (i) characterize baseline congestion behavior in controlled scenarios and (ii) provide an initial, clean evaluation of emerging Ethernetbased fabrics before turning to large-scale (Sec.V) and bursty (Sec.VI) regimes, which constitute the main focus of this paper.

TABLE I S UMMARY OF EVALUATED ENVIRONMENTS AND INTERCONNECT CONFIGURATIONS . System

Partition Nodes 3456

Interconnect

Compute node (CPU/GPU)

Memory

Link rate / node

Topology

HDR InfiniBand

512 GB

NDR InfiniBand

512 GB

400 Gb/s (2×dualport) 200 Gb/s (dual-port)

Dragonfly+

760

BullSequana X2135 “Da Vinci”, Intel Xeon Platinum 8358 Intel Xeon Platinum 8592+

2978

Cray Slingshot

AMD EPYC “Trento”

512 GB

Dragonfly

HAICGU

10

TaiShan 200 (2280), Kunpeng 920

128 GB

Single switch

Nanjing lab

8

EDR InfiniBand/RoCE RoCE-NSLB

800 Gb/s (4×200 Gb/s links) 100 GE;

TaiShan 200 (2280), Kunpeng 920

128 GB

200 GE

2-spine/2-leaf

Leonardo (CINECA) CRESCO8 (ENEA) LUMI (CSC)

IV. S TEADY C ONGESTION AT S MALL S CALE A. HAICGU and Nanjing Since prior work reported a throughput drop for the CloudEngine (CE) switches under AlltoAll communication [30], we first compare the two CE-based Ethernet testbeds under uncongested conditions to further investigate the reasons for this drop. On HAICGU (CE8850), RoCE cannot sustain throughput with large messages due to a recurring sawtooth pattern for AlltoAll and AllGather on vectors larger than 16 MiB (Figure 3). On the same system (thus using the same compute nodes) InfiniBand remains stable, and we can thus attribute such behavior to an unstable feedback-based congestion-control response [35]. In contrast, the Nanjing CE9855-based fabric maintains stable throughput across message sizes and does not exhibit the oscillations seen on the CE8850 (Figure 3), achieving variability comparable to that of the HAICGU InfiniBand partition. This indicates that the issue has been resolved in later generations of CE switches. For these reasons, we exclude the HAICGU system, based on CE8850 switch, from the following analysis due to its unstable behavior.

Fig. 3. 4 nodes HAICGU sawtooth behavior on 128 MiB messages with AllGather collective

1.67:1 blocking Fat-Tree

On the Nanjing system, based on the CE9855 Ethernet switch, we instead also analyze the impact of congestion on network performance. On this testbed, we can also explicitly enable and disable NSLB to isolate the contribution of network load balancing, and see how the system would behave without load balancing. We run an AlltoAll as victim workload on four nodes, and another AlltoAll as aggressor on the other four nodes. Nodes are allocated to the victim and the aggressor in an interleaved way. Figure 4 shows the results of the analysis. Without congestion, the victim AlltoAll reaches a peak bandwidth of 180 Gb/s. When NSLB is active, there is no performance drop even in presence of congestion (the two lines perfectly overlap). On the other hand, when disabling NSLB, in presence of congestion the performance drops to 120 Gb/s, showing the effectiveness of NSLB. B. Leonardo, CRESCO8, and LUMI At the same scale, Leonardo, CRESCO8, and LUMI exhibit standard uncongested behavior, without signaling particular drops or anomalies. Even in presence of congestion, they handle steady congestion with limited performance impact. For Leonardo and LUMI, the Dragonfly+/Dragonfly topology is particularly effective for small allocations, since nodes are often distributed across groups, enabling multiple paths between endpoints; as a result, load balancing can easily deflect traffic and resolve congestion on intermediate switches. Despite its different topology, CRESCO8 exhibits similarly robust behavior, and its blocking topology does not noticeably affect performance at this scale. Overall, these results indicate that, when operating at small scale on sufficiently large systems, steady congestion is often mitigated by path diversity and available network capacity, reducing the sensitivity to the specific topology or congestion-control mechanism. V. S TEADY C ONGESTION AT L ARGE S CALE

Observation 1. Even without explicit congestion injection, a system can still exhibit congestion-related effects: if an application’s communication rate is sufficiently high, the congestion control may be unable to sustain it, leading to throughput drops and instability even when no other workloads are present.

After the small-scale study, we move to a more challenging regime by increasing the allocation up to 256 nodes, while keeping steady and persistent congestion. In this section, we focus only on production supercomputers (Leonardo, CRESCO8, and LUMI), since the two Ethernet-based systems have no more than 10 nodes per partition. At 256 nodes, the experiment spans a non-negligible fraction of each machine:

Fig. 4. Steady NSLB analysis in a AlltoAll congestion with 4 victims and 4 aggressor nodes

7.39% of Leonardo’s Booster partition, 8.60% of LUMI’s GPU partition, and 33.68% of the CRESCO8 CPU partition. These results are shown in Figure 5, where, for each system, we report two heatmaps: one showing the data from an AllGather running together with an AlltoAll aggressor, and the other showing the data from an AllGather run with an Incast aggressor. On each heatmap, we report on the x-axis the number of nodes (half used by the victim and half by the aggressor), and on the y-axis the size of the vector used as an input by the victim AllGather. The value in each cell represents the ratio between the average uncongested and congested runtimes of the victim (i.e., the higher the better). It is worth noting that in some cases this might be slightly greater than 1. I.e., due to temporary variability sometimes the congested execution, if not strongly affected by congestion, might be slightly faster (but still within the 10%) than the uncongested execution. A. CRESCO8 AlltoAll Aggressor: Figure 5 (top left) highlights that CRESCO8 remains close to baseline under steady AlltoAllinduced congestion up to 32 nodes, but a clear degradation emerges once allocations reach 64 nodes and beyond. For a few vector sizes, the value drops to as low as 0.45. This means that, when running with congestion, the application performance can be just the 45% of the uncongested case, indicating that the fabric can no longer fully absorb the persistent congestion. The impact of congestion is even higher when using the all-to-all as a victim (not shown in the plot due to space constraints), with slowdowns up to 5×. Incast Aggressor: The effect is similar under steady Incast (Figure 5, bottom left), with performance dropping to 60% of the uncongested case. However, the impact of congestion is larger when using the all-to-all as a victim (not shown in the plot due to space constraints), with slowdowns up to 25×. B. Leonardo AlltoAll Aggressor: Figure 5 (top center) shows a markedly different behavior on Leonardo. Under steady AlltoAll background traffic, performance remains close to (and in a few cases slightly above) the uncongested baseline across all tested scales: most entries stay around ∼0.95–1.05, with only minor

localized slowdowns (e.g., ≈0.82–0.85) on 64 nodes. This indicates that, for intermediate-switch contention, Leonardo’s fabric provides sufficient path diversity and adaptive mechanisms to deflect traffic and preserve throughput, at least for the scales considered in our analysis. Incast Aggressor: The picture changes drastically with steady Incast (Figure 5, bottom center). While 8-node runs are still essentially unaffected, degradation appears already at 16 nodes and becomes severe at 32–64 nodes, where the ratio between uncongested and congested runtime for several vector sizes collapse to 0.2 (i.e., more than a 5× slowdown). This is caused by edge-localized congestion: the bottleneck is concentrated near the destination leaf switch, so alternative paths offer limited relief, and the performance is dominated by congestion control quality and buffering dynamics. C. LUMI AlltoAll Aggressor: LUMI exhibits the most robust behavior among the evaluated systems. As shown in Figure 5 (top right), under steady AlltoAll background traffic, the performance remains essentially unchanged across all tested scales, with performance in the congested case always within 5% of the uncongested case. This indicates that intermediateswitch contention is effectively managed by the fabric, and that load balancing preserves near-baseline performance even as the allocation grows. Incast Aggressor: More importantly, LUMI is also resilient to steady Incast (Figure 5, bottom right). Unlike Leonardo and CRESCO8, where Incast triggers large collapses at moderate node counts, LUMI maintains ratios close to 1.0 across message sizes and node counts, with only minor fluctuations (generally within a few percent). This suggests that edgelocalized hot spots are handled effectively by the congestion control, which correctly identifies the victim flows and protect them from congestion induced by the Incast aggressor. Overall, these results show that, within 8–256 nodes, LUMI sustains performance even in the challenging edge congestion regime, providing stable throughput largely independent of both vector size and congestion pattern. D. Summary Overall, the three supercomputers exhibit different behaviors, depending on their respective interconnect technologies and topologies. CRESCO8 deploys a blocking fat-tree, and while congestion does not significantly affect it below 32 nodes, performance degrades sharply as the allocation grows, in part also due to its tapered topology. This is particularly evident when using an AlltoAll aggressor, especially since we are using a large fraction of the system (256 out of 760 nodes). On the other hand, although Leonardo also relies on a tapered network (based on a Dragonfly+ topology), there are more opportunities to route traffic around hotspots since we are using a smaller fraction of the system. However, it suffers Incast aggressors more than CRESCO8, highlighting issues related to congestion control. In contrast, LUMI combines the Cray Slingshot interconnect with a Dragonfly topology

(a) CRESCO8

(b) Leonardo

(c) LUMI

Fig. 5. Ratio between uncongested and congested runtimes on CRESCO8, Leonardo and LUMI, from 16 to 256 nodes, and vectors ranging from from 8 bytes to 16 MiB.

(a) CRESCO8

(b) Leonardo

(c) LUMI

Fig. 6. Ratio between uncongested and congested runtime of 512 bytes, 32KiB and 2MiB AllGather victim, on Leonardo, CRESCO8 and LUMI, for 64 nodes under AlltoAll and Incast aggressors.

and shows the most stable behavior: in the presence of both AlltoAll and Incast congestion, the performance remains near baseline across scales, suggesting that its routing and congestion-management mechanisms handle both intermediate- and edge-localized steady congestion more effectively.

Observation 2. Although CRESCO8 and Leonardo use similar network technology, they respond differently to different types of congestion. Alltoall congestion has a higher impact on CRESCO8, whereas Incast has a higher impact on Leonardo. This can be attributed in part to the smaller number of nodes in CRESCO8, in part to different network generation (older on Leonardo than on CRESCO8), and in part to adaptive routing and congestion control tuning or design differences. Lastly, it is worth noting that congestion (especially generated by Incast) affects performance even at moderate allocation sizes (on the order of ∼ 2% of a partition).

VI. B URSTY C ONGESTION AT L ARGE S CALE

After characterizing steady congestion, we move to a more realistic and challenging bursty regime, where contention is injected intermittently with configurable burst duration and idle gaps, as described in Sec. III-D. Bursty congestion is particularly challenging because it requires the fabric to operate in a dynamic environment: rather than converging once to a steady-state, the system must rapidly adapt to changes. We focus in this section on large-scale allocations, since the steady congestion experiments showed that this is where congestion has the highest impact. The results of this analysis on 64 nodes are shown in Fig. 6, which displays multiple 3×3 heatmaps, one for each aggressor type and victim vector size. Each of those heatmaps show the victim slowdowns (ratio between uncongested and congested runtimes), when varying the burst pause and the burst length. Moreover, we show similar results on 128 nodes on CRESCO8 in Fig. 7, and 256 nodes on LUMI in Fig. 8.

Fig. 7. Ratio between uncongested and congested runtime. 128 nodes AllGather execution analysis on CRESCO8 under AlltoAll and Incast congestion ranging from 8 byte vectors to 16 MiB.

A. CRESCO8 AlltoAll aggressor: the performance degradation caused by bursty intermediate congestion is comparable to that of steady congestion at the same node count. The application’s performance drops to 70% of its uncongested baseline on 64 nodes (Fig. 6) and 128 nodes (Fig. 7). Incast aggressor: Conversely, Incast bursts are, in most cases, less tolerated on 64 nodes (Fig. 6), with many configurations characterized by a ratio of approximately 0.08 (i.e., 12× slower than the uncongested case). On 128 nodes (Fig. 7), performance is less affected, presumably because on a higher node count, the congestion tree originating from the Incast destination can spread over a higher number of switches, and thus the average queue length will be smaller.

B. Leonardo AlltoAll Aggressor: Consistently with what observed in the steady case, under AlltoAll congestion, the system is largely resilient across most burst configurations: congested performance remains close to the uncongested baseline for small and medium message sizes, with only localized slowdowns appearing in the most aggressive regimes characterized by short gaps between bursts. Incast Aggressor: The picture changes markedly for Incast bursts. Here, performance degradation becomes widespread, especially when having short gaps between bursts, regardless of the burst length. Even in this case, a limited idle time prevents the system from recovering and dynamically adapting to congestion. On the other hand, long gaps between bursts, Leonardo experiences only a 15% performance drop on short bursts, but a higher performance drops on longer bursts.

Observation 3. Bursty congestion exposes the limits of reactive congestion handling: while modern HPC fabrics can often tolerate intermediate-switch contention even under bursty workloads, edge-intensive congestion remains the dominant challenge. In addition, short idle gaps between bursts are especially harmful, leaving insufficient time for queues to drain and for endpoints and control mechanisms to react. Consequently, performance is bounded by edge buffering and endpoint rate control, for which additional path diversity provides limited mitigation. C. LUMI AlltoAll Aggressor: On both on 64 (Figure 6) and 256 nodes (Figure 8) LUMI exhibits the highest resilience to bursty congestion among the evaluated systems. Across the explored burst configurations, throughput remains close to the uncongested baseline for almost all message sizes, indicating that transient congestion rarely escalates into sustained performance loss. Even for aggressive burst schedules (short gaps and longer bursts), the system rapidly absorbs and dissipates contention. Incast Aggressor: Similar considerations apply for Incast aggressors, regardless of scale, burst length, and burst idle gaps. Observation 4. LUMI shows that Slingshot-based fabrics can effectively handle both intermediate and edge contention under dynamic, bursty congestion, maintaining near-baseline throughput across a wide range of burst configurations. D. Summary Results show that congestion resilience is not a single property of an interconnect, but an interaction between where contention occurs (intermediate fabric vs. edge-links), its pattern (steady vs. bursty), and the specific congestion-management mechanisms available at endpoints and in-network. High path diversity and adaptive routing (e.g., Dragonfly-like fabrics)

are effective at absorbing intermediate contention typical of AlltoAll-like patterns, often preserving near-ideal throughput, unless a high fraction of the system is used (as we have seen on CRESCO8). However, edge-intensive congestion remains the dominant challenge across platforms: it stress endpoint injection control and shallow leaf buffering, and can trigger sharp throughput collapse already at moderate allocation sizes (on the order of ∼2% of a partition). Bursty congestion further exposes the limits of reactive handling: performance becomes highly sensitive to the duty cycle, since short idle gaps do not provide enough time for buffers to drain and for rate control to converge. In this sense, the best-performing systems are not simply those with more connectivity, but those that combine path diversity with robust endpoint throttling and fast, stable congestion response. While interconnect topologies may appear comparable, performance diverges based on the specific network generation and the congestion control and adaptive routing mechanisms. The design and fine-tuning of these algorithms are critical; they determine whether a system can maintain throughput or be affected by congestion under the unpredictable demands of multi-tenant workloads. In production environments, these challenges are even more exacerbated by inter-job interference from multiple co-running applications and the highly unpredictable, stochastic nature of real-world traffic. While our methodology employs periodic bursts with fixed and recurring durations, representing a relatively ’favorable’ or ’easy’ scenario for network recovery, our side-by-side comparison reveals that even these regular patterns expose critical weak spots in modern adaptive routing and congestion control (CC) algorithms. The fact that these mechanisms struggle to stabilize the fabric under controlled conditions suggests that their efficacy in chaotic, productiongrade environments remains a major concern. Observation 5. Physical topology alone does not dictate how a system behaves under saturation; rather, congestion resilience is determined by the combined impact of the interconnect technology, generation, and topology. Ultimately, the system’s performance depends on how these architectural features interact with the specific design and tuning of congestion control and adaptive routing algorithms. This multi-factored dependence explains why there is no single correlation between fabric layout and its ability to withstand network pressure. VII. R ELATED W ORK The impact of network congestion and performance variability has been a significant focus of recent HPC research, ranging from empirical characterization to the development of system-level mitigation strategies. A. Characterization of Network Noise and Variability Extensive research has focused on quantifying the impact of network noise using benchmarks centered on MPI collective

operations [36]–[39]. These studies establish a baseline for how variability affects communication performance in architectures like Cray Aries. Recent comparative studies have also explored the performance gap between on-premise and cloudbased HPC environments, utilizing small-scale measurements and large-scale simulations to demonstrate how congestion can degrade collective performance as systems scale [40]. Additionally, researchers have investigated the sensitivity of in-network collective offloading (where collective operations are processed within the switch) to background traffic and network contention [41]. B. Job Placement and Adaptive Routing Strategies A significant body of work explores how job allocation and placement strategies influence a network’s sensitivity to congestion across various topologies, including Dragonfly [42] and fat-trees [43], [44]. These studies have identified that while randomized allocation can alleviate hotspots for communication-intensive workloads, contiguous placement is often superior for jobs with lower communication demands [45], [46]. Beyond placement, the role of adaptive routing in performance stability has been closely examined, particularly in low-diameter networks [47]. Because routing decisions are typically made based on local switch visibility rather than a global network state, switches may select suboptimal, longer paths during periods of contention, which can result in significant performance drops depending on the application traffic pattern [47], [48]. Various telemetry tools have also been developed to aggregate hardware counters to help visualize and diagnose these network-wide congestion events [49]–[51]. C. Fabric-Specific Congestion Studies Existing research often provides deep-dive analyses of specific hardware architectures. For instance, detailed studies have characterized the congestion management systems of the Slingshot and Aries fabrics [26], while other work has examined the impact of Incast patterns on large-scale InfiniBand systems like Leonardo, identifying how persistent traffic pressure reduces effective bandwidth [52]. D. Previous Works Limitations While the aforementioned studies were foundational, they often reflect the technological landscape of their time or focus on single architectures in isolation. Our work differs in several key aspects. On one side, whereas much of the existing literature relies on simulations or legacy technologies (e.g., Cray Aries), this study provides a side-by-side experimental comparison of modern InfiniBand, Slingshot, and Ethernet ecosystems using actual production hardware. On the other side, by evaluating these fabrics under both steady-state and bursty congestion, scenarios often overlooked in individual fabric studies, we provide a comprehensive view of interconnect resilience that is currently missing.

Fig. 8. Ratio between uncongested and congested runtime. 256 nodes AllGather execution analysis on LUMI under AlltoAll and Incast congestion ranging from 8 byte vectors to 16 MiB.

VIII. D ISCUSSION AND C ONCLUSION Through a rigorous side-by-side evaluation of diverse interconnect architectures, this work identifies critical limitations in state-of-the-art congestion control and adaptive routing mechanisms. By employing a flexible methodology that subjects fabrics to both steady and bursty traffic profiles, we isolate weak spots where traditional mitigation strategies fail. Specifically, our results demonstrate that while path diversity effectively resolves intermediate switch contention at small scales, edge congestion remains a persistent bottleneck across several modern fabrics. Furthermore, our analysis of bursty traffic highlights that the network duty cycle—the relationship between burst intensity and recovery intervals—is the primary determinant of whether a fabric can maintain stability or falls into recurrent congestion cycles. These findings underscore that congestion management remains an open research challenge, suggesting that future system designs and procurement metrics should prioritize endpoint and leaf-link control over raw peak bandwidth. More broadly, our analysis reveals that future fabrics must address the increasing communication bottleneck in AI and HPC workloads, which is exacerbated by congestion. Emerging standards such as Ultra Ethernet leverage standardized packet spraying to reduce the impact of congestion and mitigate tail latency. Meanwhile, other efforts are moving towards centralized adaptive load balancers specialized for LLM training. However, while centralized approaches can minimize congestion effects, their effectiveness at large scale or under highly dynamic traffic patterns remains limited. As future work, we plan to extend the analysis to broader scales and incorporate application-driven traces to bridge controlled microbenchmarks with production workload behavior. ACKNOWLEDGMENTS We thank Matteo Marcelletti for the support with part of the code used for the benchmarks. This work is supported by the European Union’s Horizon Europe under grant 101175702 (NET4EXA), by Sapienza University Grants ADAGIO and

D2QNeT (Bando per la ricerca di Ateneo 2023 and 2024). We acknowledge ISCRA for awarding this project access to the LEONARDO supercomputer, owned by the EuroHPC Joint Undertaking, hosted by CINECA (Italy). We acknowledge the EuroHPC Joint Undertaking for awarding this project access to the EuroHPC supercomputer LUMI, hosted by CSC (Finland) and the LUMI consortium through a EuroHPC Regular Access call. The CRESCO8 computing resources and related technical support used for this work were provided by the CRESCO/ENEAGRID HPC and its staff [32]. CRESCO/ENEAGRID HPC infrastructure is funded by ENEA, the Italian National Agency for New Technologies, Energy and Sustainable Economic Development, and by Italian and European research programmes. Furthermore, we gratefully acknowledge the Goethe University of Frankfurt for hosting the HAICGU cluster. R EFERENCES [1] U. E. Consortium, “Ultra ethernet,” 2024, https://ultraethernet.org/. [2] TOP500 Project, “TOP500: Ranking of the World’s 500 Fastest Supercomputers,” https://top500.org/, 2025, accessed July 22, 2025. [3] “Ultra ethernet specification version 1.0,” Ultra Ethernet Consortium, Technical Specification, 2024, available from the Ultra Ethernet Consortium. [4] Y. Zhu, H. Eran, D. Firestone, C. Guo, M. Lipshteyn, Y. Liron, J. Padhye, S. Raindel, M. H. Yahia, and M. Zhang, “Congestion Control for LargeScale RDMA Deployments,” in Proceedings of the ACM SIGCOMM 2015 Conference (SIGCOMM ’15), London, United Kingdom, 2015, pp. 523–536. [5] R. Mittal, V. T. Lam, N. Dukkipati, E. Blem, H. Wassel, M. Ghobadi, A. Vahdat, Y. Wang, D. Wetherall, and D. Zats, “TIMELY: RTTbased Congestion Control for the Datacenter,” in Proceedings of the ACM SIGCOMM 2015 Conference (SIGCOMM ’15), London, United Kingdom, 2015, pp. 537–550. [6] Y. Li, R. Miao, H. H. Liu, Y. Zhuang, F. Feng, L. Tang, Z. Cao, M. Zhang, F. Kelly, M. Alizadeh, and M. Yu, “HPCC: High Precision Congestion Control,” in Proceedings of the ACM SIGCOMM 2019 Conference (SIGCOMM ’19), Beijing, China, 2019, pp. 44–58. [7] D. omitted for double-blind reviewing, “Under submission.” [8] Huawei Support, “Ai ecn threshold of lossless queues,” https://support.huawei.com/enterprise/en/doc/EDOC1100420118/7ade444e/aiecn-threshold-of-lossless-queues, 2024, accessed: 2025-12-20. [9] C. Hopps, “Analysis of an Equal-Cost Multi-Path Algorithm,” RFC 2992, Nov. 2009. [Online]. Available: https://www.ietf.org/rfc/rfc2992.txt

[10] M. Al-Fares, S. Radhakrishnan, B. Raghavan, N. Huang, and A. Vahdat, “Hedera: dynamic flow scheduling for data center networks,” in Proceedings of the 7th USENIX Conference on Networked Systems Design and Implementation, ser. NSDI’10. USA: USENIX Association, 2010, p. 19. [11] M. Alizadeh, T. Edsall, S. Dharmapurikar, R. Vaidyanathan, K. Chu, A. Fingerhut, V. T. Lam, F. Matus, R. Pan, N. Yadav, and G. Varghese, “Conga: distributed congestion-aware load balancing for datacenters,” in Proceedings of the 2014 ACM Conference on SIGCOMM, ser. SIGCOMM ’14. New York, NY, USA: Association for Computing Machinery, 2014, p. 503–514. [Online]. Available: https://doi.org/10.1145/2619239.2626316 [12] A. Gangidi, R. Miao, S. Zheng, S. J. Bondu, G. Goes, H. Morsy, R. Puri, M. Riftadi, A. J. Shetty, J. Yang, S. Zhang, M. J. Fernandez, S. Gandham, and H. Zeng, “Rdma over ethernet for distributed training at meta scale,” in Proceedings of the ACM SIGCOMM 2024 Conference, ser. ACM SIGCOMM ’24. New York, NY, USA: Association for Computing Machinery, 2024, p. 57–70. [Online]. Available: https://doi.org/10.1145/3651890.3672233 [13] T. Hoefler, D. Roweth, K. Underwood, R. Alverson, M. Griswold, V. Tabatabaee, M. Kalkunte, S. Anubolu, S. Shen, M. McLaren, A. Kabbani, and S. Scott, “Data center ethernet and remote direct memory access: Issues at hyperscale,” Computer, vol. 56, no. 7, pp. 67–77, 2023. [14] C. Raiciu, S. Barre, C. Pluntke, A. Greenhalgh, D. Wischik, and M. Handley, “Improving datacenter performance and robustness with multipath tcp,” SIGCOMM Comput. Commun. Rev., vol. 41, no. 4, p. 266–277, aug 2011. [Online]. Available: https://doi.org/10.1145/2043164.2018467 [15] M. A. Qureshi, Y. Cheng, Q. Yin, Q. Fu, G. Kumar, M. Moshref, J. Yan, V. Jacobson, D. Wetherall, and A. Kabbani, “Plb: congestion signals are simple and effective for network load balancing,” in Proceedings of the ACM SIGCOMM 2022 Conference, ser. SIGCOMM ’22. New York, NY, USA: Association for Computing Machinery, 2022, p. 207–218. [Online]. Available: https://doi.org/10.1145/3544216.3544226 [16] A. Kabbani, B. Vamanan, J. Hasan, and F. Duchene, “Flowbender: Flow-level adaptive routing for improved latency and throughput in datacenter networks,” in Proceedings of the 10th ACM International on Conference on Emerging Networking Experiments and Technologies, ser. CoNEXT ’14. New York, NY, USA: Association for Computing Machinery, 2014, p. 149–160. [Online]. Available: https://doi.org/10.1145/2674005.2674985 [17] E. Vanini, R. Pan, M. Alizadeh, P. Taheri, and T. Edsall, “Let it flow: Resilient asymmetric load balancing with flowlet switching,” in 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17). Boston, MA: USENIX Association, Mar. 2017, pp. 407–420. [Online]. Available: https://www.usenix.org/conference/nsdi17/technicalsessions/presentation/vanini [18] K. He, E. Rozner, K. Agarwal, W. Felter, J. Carter, and A. Akella, “Presto: Edge-based load balancing for fast datacenter networks,” in Proceedings of the 2015 ACM Conference on Special Interest Group on Data Communication, ser. SIGCOMM ’15. New York, NY, USA: Association for Computing Machinery, 2015, p. 465–478. [Online]. Available: https://doi.org/10.1145/2785956.2787507 [19] T. Bonato, D. De Sensi, S. Di Girolamo, A. Bataineh, D. Hewson, D. Roweth, and T. Hoefler, “ Flowcut Switching: High-Performance Adaptive Routing With In-Order Delivery Guarantees ,” IEEE Transactions on Networking, no. 01, pp. 1–14, Dec. 2025. [Online]. Available: https://doi.ieeecomputersociety.org/10.1109/TON.2025.3636209 [20] Y. Lu, G. Chen, B. Li, K. Tan, Y. Xiong, P. Cheng, J. Zhang, E. Chen, and T. Moscibroda, “Multi-path transport for rdma in datacenters,” in Proceedings of the 15th USENIX Conference on Networked Systems Design and Implementation, ser. NSDI’18. USA: USENIX Association, 2018, p. 357–371. [21] A. Dixit, P. Prakash, Y. C. Hu, and R. R. Kompella, “On the impact of packet spraying in data center networks,” in 2013 Proceedings IEEE INFOCOM, 2013, pp. 2130–2138. [22] W. Wang, F. Chen, P. Cao, L. Shan, T. Wu, and H. Wen, “Network load balancing technologies for intelligent computing centers,” Communications of Huawei Research, no. Issue 9, pp. 13–22, 2025. [23] E. G. Gran and S.-A. Reinemo, “InfiniBand Congestion Control: Modelling and Validation,” in OMNeT++ 2011 (Workshop at SIMUTools), Barcelona, Spain, 2011.

[24] F. Alali, F. Mizero, M. Veeraraghavan, and J. M. Dennis, “A Measurement Study of Congestion in an InfiniBand Network,” in 2017 Network Traffic Measurement and Analysis Conference (TMA), 2017, pp. 1–9. [25] J. Rocher-González, E. G. Gran, S.-A. Reinemo, T. Skeie, J. EscuderoSahuquillo, P. J. Garcı́a, and F. J. Q. Flor, “Adaptive routing in infiniband hardware,” in 2022 22nd IEEE/ACM International Symposium on Cluster, Cloud and Internet Computing (CCGrid), 2022, pp. 463–472. [26] D. De Sensi, S. Di Girolamo, K. H. McMahon, D. Roweth, and T. Hoefler, “An in-depth analysis of the slingshot interconnect,” in SC20: International Conference for High Performance Computing, Networking, Storage and Analysis, 2020, pp. 1–14. [27] D. Roweth, “Hpe slingshot launched into network space,” Cray User Group (CUG) Proceedings, 2022. [28] S. Chunduri, T. Groves, P. Mendygral, B. Austin, J. Balma, K. Kandalla, K. Kumaran, G. Lockwood, S. Parker, S. Warren, N. Wichmann, and N. J. Wright, “Gpcnet: Designing a benchmark suite for inducing and measuring contention in hpc networks,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC ’19). New York, NY, USA: Association for Computing Machinery, 2019. [29] “Open MPI: Open Source High Performance Computing,” https://www.open-mpi.org/, Associated with Software in the Public Interest, 2025, accessed: 2025-09-02. [30] L. Pichetti, D. De Sensi, K. Sivalingam, S. Nassyr, D. Cesarini, M. Turisini, D. Pleiter, A. Artigiani, and F. Vella, “Benchmarking ethernet interconnect for hpc/ai workloads,” in Proceedings of the SC ’24 Workshops of the International Conference on High Performance Computing, Network, Storage, and Analysis, ser. SCW ’24. IEEE Press, 2025, p. 869–875. [Online]. Available: https://doi.org/10.1109/SCW63240.2024.00124 [31] M. Turisini, G. Amati, and M. Cestari, “Leonardo: A pan-european pre-exascale supercomputer for hpc and ai applications,” 2023. [Online]. Available: https://arxiv.org/abs/2307.16885 [32] F. Iannone, F. Ambrosino, G. Bracco, M. De Rosa, A. Funel, G. Guarnieri, S. Migliori, F. Palombi, G. Ponti, G. Santomauro, and P. Procacci, “CRESCO ENEA HPC clusters: a working example of a multifabric GPFS Spectrum Scale layout,” in 2019 International Conference on High Performance Computing Simulation (HPCS), 2019, pp. 1051–1052. [33] LUMI Consortium, “LUMI supercomputer,” https://lumisupercomputer.eu/, 2024, accessed: 2025-01. [34] “Open edge and hpc initiative,” https://www.open-edge-hpcinitiative.org/, 2025, accessed: 2025-08-28. [35] D.-M. Chiu and R. Jain, “Analysis of the increase and decrease algorithms for congestion avoidance in computer networks,” Computer Networks and ISDN Systems, vol. 17, no. 1, pp. 1–14, 1989. [Online]. Available: https://www.sciencedirect.com/science/article/pii/0169755289900196 [36] T. Groves, Y. Gu, and N. J. Wright, “Understanding performance variability on the aries dragonfly network,” in 2017 IEEE International Conference on Cluster Computing (CLUSTER), Sept 2017, pp. 809–813. [37] T. Hoefler, T. Schneider, and A. Lumsdaine, “The impact of network noise at large-scale communication performance,” in 2009 IEEE International Symposium on Parallel Distributed Processing, May 2009, pp. 1–8. [38] S. Chunduri, K. Harms, S. Parker, V. Morozov, S. Oshin, N. Cherukuri, and K. Kumaran, “Run-to-run variability on xeon phi based cray xc systems,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, ser. SC ’17. New York, NY, USA: ACM, 2017, pp. 52:1–52:13. [Online]. Available: http://doi.acm.org/10.1145/3126908.3126926 [39] D. Skinner and W. Kramer, “Understanding the causes of performance variability in hpc workloads,” in IEEE International. 2005 Proceedings of the IEEE Workload Characterization Symposium, 2005., Oct 2005, pp. 137–149. [40] D. De Sensi, T. De Matteis, K. Taranov, S. Di Girolamo, T. Rahn, and T. Hoefler, “Noise in the clouds: Influence of network performance variability on application scalability,” Proc. ACM Meas. Anal. Comput. Syst., vol. 6, no. 3, Dec. 2022. [41] D. De Sensi, E. Costa Molero, S. Di Girolamo, L. Vanbever, and T. Hoefler, “Canary: Congestion-aware in-network allreduce using dynamic trees,” Future Generation Computer Systems, vol. 152, pp. 70–82, 2024. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0167739X23003850

[42] B. Prisacari, G. Rodriguez, P. Heidelberger, D. Chen, C. Minkenberg, and T. Hoefler, “Efficient task placement and routing of nearest neighbor exchanges in dragonfly networks,” in Proceedings of the 23rd International Symposium on High-performance Parallel and Distributed Computing, ser. HPDC ’14. New York, NY, USA: ACM, 2014, pp. 129– 140. [Online]. Available: http://doi.acm.org/10.1145/2600212.2600225 [43] S. D. Pollard, N. Jain, S. Herbein, and A. Bhatele, “Evaluation of an interference-free node allocation policy on fat-tree clusters,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage, and Analysis, ser. SC ’18. Piscataway, NJ, USA: IEEE Press, 2018, pp. 26:1–26:13. [Online]. Available: http://dl.acm.org/citation.cfm?id=3291656.3291691 [44] A. Bhatele and L. V. Kalé, “Quantifying network contention on large parallel machines,” Parallel Processing Letters, vol. 19, no. 04, pp. 553–572, 2009. [Online]. Available: https://doi.org/10.1142/S0129626409000419 [45] X. Yang, J. Jenkins, M. Mubarak, R. B. Ross, and Z. Lan, “Watch out for the bully! job interference study on dragonfly network,” in SC ’16: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, Nov 2016, pp. 750–760. [46] X. Wang, M. Mubarak, X. Yang, R. B. Ross, and Z. Lan, “Trade-off study of localizing communication and balancing network traffic on a dragonfly system,” in 2018 IEEE International Parallel and Distributed Processing Symposium (IPDPS), May 2018, pp. 1113–1122. [47] D. De Sensi, S. Di Girolamo, and T. Hoefler, “Mitigating network noise on dragonfly networks through application-aware routing,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, ser. SC ’19. New York, NY, USA: ACM, 2019, pp. 16:1–16:32. [Online]. Available: http://doi.acm.org/10.1145/3295500.3356196 [48] S. A. Smith, C. E. Cromey, D. K. Lowenthal, J. Domke, N. Jain, J. J. Thiagarajan, and A. Bhatele, “Mitigating inter-job interference using adaptive flow-aware routing,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, ser. SC ’18, 2018. [49] J. M. Brandt, E. Froese, A. C. Gentile, L. Kaplan, B. A. Allan, and E. J. Walsh, “Network performance counter monitoring and analysis on the cray xc platform.” 5 2016. [50] A. Bhatele, N. Jain, Y. Livnat, V. Pascucci, and P. Bremer, “Analyzing network health and congestion in dragonfly-based supercomputers,” in 2016 IEEE International Parallel and Distributed Processing Symposium (IPDPS), May 2016, pp. 93–102. [51] R. E. Grant, K. T. Pedretti, and A. Gentile, “Overtime: A tool for analyzing performance variation due to network interference,” in Proceedings of the 3rd Workshop on Exascale MPI, ser. ExaMPI ’15. New York, NY, USA: ACM, 2015, pp. 4:1–4:10. [Online]. Available: http://doi.acm.org/10.1145/2831129.2831133 [52] D. De Sensi, L. Pichetti, F. Vella, T. De Matteis, Z. Ren, L. Fusco, M. Turisini, D. Cesarini, K. Lust, A. Trivedi, D. Roweth, F. Spiga, S. Di Girolamo, and T. Hoefler, “Exploring gpu-to-gpu communication: Insights into supercomputer interconnects,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC’24), Nov. 2024.

Record · ID 10310 · SHA-256 fdbfbcffa847070b
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.