JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
1
Switching Efficiency: A Novel Framework for Dissecting AI Data Center Network Efficiency
arXiv:2604.14690v1 [cs.NI] 16 Apr 2026
Niangen Ye, Jiawen Zhu, Baojun Chen, Dong Wang, Jiang Sun, Weiqiang Sun, Senior Member, IEEE, and Weisheng Hu, Member, IEEE
Abstract—Communication is pivotal in LLM training, and a thorough analysis of the communication efficiency of AI data center (AIDC) network is essential for guiding the design of these capital-intensive clusters. However, conventional metrics are inadequate for such analysis, as they do not directly link network activity to computational progress and lack granularity to diagnose the impact of different network design patterns. To address this, we introduce a metric framework, the Switching Efficiency Framework, whose core metric — Switching Efficiency (η) — quantifies computationally effective data throughput per unit switching capacity. We further decompose η into three factors — Data, Routing Efficiency, and Port Utilization to facilitate analysis of distinct communication bottlenecks. Using this metric framework, we demonstrate how the symmetric, distributed switching of 3D-Torus and the centralized, hierarchical switching of Rail-Optimized architecture align with sparse or imbalanced LLM training traffic, and show that Allto-All traffic from Mixture-of-Experts models severely degrades their port utilization and routing efficiency. Our analysis also demonstrates how key design choices — such as adjusting switching resource allocation, expanding server size, adopting in-network computing, and multi-plane design — positively influence distinct facets of communication efficiency. Ultimately, the Switching Efficiency Framework provides an analytical tool for analyzing efficiency bottlenecks, thereby informing the design of future-generation AIDC networks. Index Terms—AI data center networks, Communication efficiency, Efficiency metric
I. I NTRODUCTION The rapidly growing size of Large Language Models (LLMs) is driving the construction of hyperscale AIDCs, where tens or even hundreds of thousands of GPUs are interconnected by a high-speed network [1–6]. Critically, with the growth in both model training volume and per-GPU computational capability outpacing the interconnect capabilities, communication has emerged as the principal performance bottleneck [7–10]. This reality elevates the communication efficiency of the AIDC network to the central determinant of training throughput and system scalability for these multibillion-dollar clusters [11, 12]. To improve communication efficiency, existing studies have explored both algorithmic and architectural approaches. The former [13–19] focuses on developing communication libraries N. Ye, J. Zhu, W. Sun, and W. Hu are with the State Key Laboratory of Photonics and Communications, Shanghai Jiao Tong University, Shanghai, China. B. Chen is with Pengcheng Laboratory, China. W. Dong and J. Sun are with Department of Fundamental Network Technology, China Mobile Research Institute, Beijing, China. Corresponding author: Weiqiang Sun (E-mail: [email protected]).
and scheduling policies to minimize communication time during LLM training in existing network architectures. The latter [1, 20–30] focuses on proposing domain-specific architectures to adapt to the structured nature (sparse and imbalanced) of the traffic in LLM training [25, 29, 31]. As these approaches broaden the design space, it becomes increasingly critical to quantitatively assess how different design trade-offs impact the efficiency of AIDC networks, which is crucial in guiding the resource-efficient design of future million-GPU clusters. However, it is challenging to quantify the efficiency of AIDC networks in a way that disentangles design trade-offs to provide actionable guidance. First, in LLM training, network communication is a means to advance computation, not an end in itself. Consequently, the evaluation of communication efficiency should measure the network’s direct contribution to computation. For instance, consider the All-Reduce operation in data parallelism of LLM training. As illustrated in Fig. 1, an implementation of All-Reduce without In-Network Computing (INC) would necessitate transmitting multiple intermediate data, which would be discarded after reduction to “computationally effective data” for subsequent neural network computation [32]. In contrast, an approach using INC delivers only the reduced result to the endpoints [33] and thus exhibits higher efficiency. In this context, an efficiency evaluation that cannot distinguish between these two methods — that is, one that fails to measure the extent to which network resources are utilized to generate computationally effective data — provides an incomplete view. communication data generation
network communication
GPU 1
data reduction GPU 1
subsequent neural network computation
⋯ ⋯ GPU 2
GPU 2
⋯ ⋯ GPU 3
data awaiting reduced
⋯ ⋯
intermediate data in communication
GPU 3
computationally effective data
Fig. 1: Intermediate data versus computationally effective data in implementing All-Reduce without in-network computing. Second, a prevailing design paradigm for AIDC networks is
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
2
to tailor the topology to traffic patterns. An evaluation intended to guide network evolution should therefore quantify how different design trade-offs within that philosophy impact communication efficiency. Consider the 3D-Torus and Rail-Optimized architectures prevalent in AIDCs [2, 6, 23]. As illustrated in Fig. 2, 3D-Torus matches the sparse communication patterns in PTD (Pipeline/Tensor/Data) parallelism by using low-radix switches (6 switching ports) within each accelerator, a design choice that avoids costly high-radix switches [23, 26]. However, its symmetric, uniformly distributed bandwidth allocation mismatches the imbalance in LLM training traffic. Conversely, the Rail-Optimized architecture matches the imbalanced traffic by creating high-bandwidth “scale-up” domains (HBD) for intensive local communication (e.g., TP) and a “scale-out” fabric for lightweight remote communication (e.g., DP, PP) [34, 35]. Yet, its non-blocking connectivity design may be an over-provision for the sparse traffic. Crucially, while tailored for PTD parallelism, both design patterns are challenged by the all-to-all traffic in training Mixture-of-Experts (MoE) models [4, 36]. Given this background, an efficiency evaluation is ambiguous if it cannot separately quantify which efficiency dimensions of a design align with the specific traffic characteristics, as it would fail to provide clear guidance on how to evolve the architecture to resolve specific bottlenecks. DP
GPU3
GPU5
GPU7
GPU2
GPU4
GPU6
GPU8
y sit ar nce Sp ala b Im
Im b Sp ala ar nc sit e y
PP
TP
GPU1
(a) Communication in PTD-Parallelism.
SW
SW
SW
SW
(Scale Out)
SW
SW
SW
SW
HBD (Scale Up)
(b) 3D-Torus: distributed, symmetric switching capacity.
SW
SW
HBD (Scale Up)
(c) Rail-Optimized: centralized, hierarchical switching capacity.
Fig. 2: Selective alignment of topology with different LLM training traffic features in designing AIDC networks. Addressing these challenges requires developing an evaluation framework based on metrics precisely defined to capture the nature of the AIDC network. Here we argue that the design of this evaluation framework and its core metric should be based on the following principles: P1: Measure computationally effective data throughput. To link the communication efficiency of the AIDC network with its goal of advancing neural network computation, the core metric needs to quantify the throughput of computationally effective data that is immediately usable for subsequent neural network calculations. This application-level
focus avoids merely tracking network “busyness” and instead assesses its direct contribution to training progress. P2: Couple traffic with the configured switching capacity in network topology. To quantitatively guide the trafficpattern-driven network design, the core metric needs to assess how effectively an architecture — which physically manifests as a specific distribution of switching capacity in the topology — is aligned with the structured traffic features of LLM training. By doing so, this metric enables comparison of key design trade-offs, such as different resource allocation strategies under different topologies from a resource utilization perspective, and thus promotes resource-efficient network design. P3: Enable diagnostic, fine-grained analysis. To further pinpoint how different design trade-offs impact communication efficiency in networks, the framework should also be able to dissect the macro-level efficiency into fine-grained factors. By enabling such fine-grained granularity, the evaluation framework turns these metrics from a simple efficiency benchmark to an analytical tool that reveals the sources of communication inefficiency, thereby providing better guidance for network design. In this study, building on the above principles, we introduce the Switching Efficiency Framework, designed to analyze the communication efficiency of AIDC networks. The framework is centered on its core metric, Switching Efficiency (η), defined as the computationally effective data throughput of an LLM training workload to the total switching resources of the network. We then decompose η into three fine-grained factors — Data, Routing Efficiency, and Port Utilization (γ, δ, θ) — each designed to isolate a specific bottleneck. Finally, we extend these metrics by introducing weighted efficiency that accounts for diverse LLM training workloads, enabling the analysis of long-term efficiency of production AIDCs that encounter various training workloads throughout their lifecycle. Based on this framework, we analyze the communication efficiency of two network design patterns prevalent in hyperscale AIDCs, represented by the 3D-Torus [4, 23] and the Rail-Optimized architecture [1–3, 5, 6]. The analysis begins with the efficiency dissection across a diverse suite of dense and MoE model training workloads under fixed network architectures. We then extend the analysis to investigate how different network design options influence efficiency. Together, these analyses demonstrate the capacity of the Switching Efficiency Framework to both benchmark current architectures and inform the design direction of future AIDC networks. The major contributions of this work include the following: • We introduce a hierarchical metric framework, comprising a top-level metric (η) to quantify overall efficiency and a set of fine-grained factors (γ, δ, θ) to dissect the root inefficiency, providing actionable guidance for network design. • We quantitatively reveal how the design patterns of 3DTorus and Rail-Optimized effectively align with the sparse or imbalanced traffic of PTD-parallelism, yet are challenged by the all-to-all traffic in training MoE models. • We clarify how different network design options, including switching resource allocation, server size, adoption of INC, and multi-plane design, impact the efficiency of port utilization, routing, and effective data reception.
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
II. L ITERATURE R EVIEW A. Studies of Traffic Patterns and Communication Inefficiencies in AIDC Networks Many studies have empirically analyzed communication characteristics during LLM training to understand the sources of communication inefficiency in AIDC networks. Investigating the traffic patterns in LLM training. Preliminary works investigated how different parallelism strategies influence the types and volumes of collective communication [46, 47]. Further research delved deeper by analyzing the spatial distribution of traffic to characterize its properties at a more granular level [25, 29, 31]. Specifically, the analyses in [25, 31] exposed the sparse and imbalanced nature of LLM training traffic in PTD hybrid parallelism, which directly motivated the design of the sparsity-aware Rail-only topology [25]. Similarly, the analysis in [29] revealed the locality of Allto-All traffic within Expert Parallelism (EP) for MoE models, leading to the proposal of MixNet, a hybrid optical-electrical architecture. While these studies provide foundational comprehension into LLM traffic characteristics, our work builds upon them by introducing a hierarchical metric framework to systematically analyze how different network design patterns align with these salient traffic patterns, aiming to derive actionable guidance for network design. Identifying communication inefficiencies in production. [1, 5, 40] point out that using ECMP routing policies can lead to hash collisions for LLM training traffic, particularly in the presence of link failures or when multiple tasks are training simultaneously. They exhibit that adopting a dual-plane RailOptimized architecture [1], or employing traffic engineering and optimizing congestion control strategies [5] can mitigate hash collisions. ByteDance [2] identifies congestion control failure, link failure, and NCCL synchronization timeouts as common failure types that cause network inefficiency in its large-scale GPU clusters, and presents a tool to identify the
3
causes of failures quickly. Collectively, these studies reveal a critical class of inefficiencies that arise from operational dynamics in production systems. Our work, on the other hand, introduces a metric framework to pinpoint the intrinsic efficiency of the network architecture itself, aiming at guiding the design of these capital-intensive clusters. B. Existing Metrics for Analyzing AIDC Network Design The analysis of AIDC network design is a multifaceted endeavor that draws upon a range of established metrics. As summarized in Table I, existing metrics provide a basis for evaluating performance, topology, and cost. However, when it comes to analyzing how different design trade-offs impact the intrinsic efficiency of AIDC networks for LLM training, they exhibit the following limitations. Performance-oriented metrics offer limited diagnostic value for network design. High-level, “black-box” performance metrics such as collective operation bandwidth or iteration time reduce system performance to a monolithic score. While they can effectively report a symptom — for instance, a performance bottleneck — they cannot analyze its origin, be it an inefficient communication algorithm, a traffic-topology mismatch, or a resource allocation issue. Even a seemingly granular “white-box” metric such as bandwidth utilization is deceptive because it fails to distinguish effective data transfers from wasteful traffic on congested multi-hop paths, consequently quantifying network busyness rather than useful throughput. Static topological metrics are traffic-agnostic, making them poor indicators of workload-specific performance. These metrics, such as bisection bandwidth or network diameter, characterize the physical graph and are independent of the traffic it supports. However, this abstraction overlooks the highly structured and non-uniform traffic patterns intrinsic to LLM training. Consequently, such metrics cannot quantify the critical alignment between a topology and its intended
TABLE I: Key Metrics for Analyzing Network Architectures Design in AIDC Category
Description
Typical Metrics
References
End-to-End Performance
Quantifies the system-level performance from an application perspective.
Model FLOPs Utilization
[2, 25, 30]
Sample Processing Rate
[2, 23, 37]
Iteration Time
[2, 25, 29]
Collective Op. Bandwidth
[1, 21, 23]
Network Throughput
[1, 27, 38]
Bandwidth Utilization
[39–41]
Latency
[20, 37, 40]
Bisection Bandwidth
[23, 42, 43]
Network Scale
[1, 42, 44]
Network Diameter
[21, 43]
Fault Explosion Radius
[30]
Cost & Cost-Efficiency
[21, 29, 45]
Power & Power-Efficiency
[25, 30, 45]
Communication Performance
Topological Characteristic
Capital Expenditure
Evaluates the effectiveness and capacity of network fabric in supporting communication operations.
Describes the static, graph-theoretic properties of the physical layout of the network architecture.
Measures the capital investment or energy footprint of a network architecture.
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
4
workload. For example, a network may exhibit an acceptable high bisection bandwidth in the global domain, yet still create a severe bottleneck for intense, localized communication (e.g., traffic within a TP group) if its local connectivity is inefficient. These metrics, assessing only the static property, would fail to identify this traffic-specific bottleneck. Capital metrics, while crucial for deployment, are illsuited for technical performance analysis. These metrics are extrinsic economic attributes, rather than the intrinsic physical properties characterizing the system behavior. Consequently, they offer limited insight into resolving performance bottlenecks, which are rooted in physical rather than commercial limitations. Collectively, these limitations underscore the need for a specific set of metrics designed for AIDC networks. In this article, we aim to address these limitations by proposing an evaluation framework with novel metrics, the details of which are provided in the subsequent sections. III. D ESIGN T HE S WITCHING E FFICIENCY F RAMEWORK Following the principles we have established, this section formally introduces the Switching Efficiency Framework, a metrics set for analyzing the communication efficiency of AIDC networks for LLM training.
a data rate of Rp . For a network A, consider a set of communication primitives O executed by the LLM training workload w over a time duration T . If each primitive i ∈ O involves Ni participants and yields an effective data volume of ∆Di , the switching efficiency η of network A for the workload w is defined as: P i∈O ∆Di (1) η= P T p∈P Rp where Ni · D Ni · NDi · (Ni − 1) N · D i Ni ∆Di = Ni · D D N i · Ni · (Ni − 1) · kr Ni · NDi · (Ni − 1) · 1
Guided by P1 and P2, we define η as the ratio of the total effective data (i.e., data that is directly usable for neural network computation after the completion of communication primitives, see Fig. 1) throughput to the network’s aggregate switching capacity. Formally, let P be the set of all egress electrical switch ports, where each egress port p ∈ P has
p∈RS
p∈RL
p∈RSV
p∈RG
in which RS , RL , RSV , and RG denote the sets of switch ports at Spine and Leaf layers, and within servers and GPUs, respectively. Here, the effective data volume ∆Di is defined as the increment of data that is directly usable to continue computations of the neural network, as illustrated in Fig. 3. Specifically,
GPU 1
GPU 1
GPU 1
GPU 1
GPU 1
GPU 1
GPU 1
GPU 1
GPU 2
GPU 2
GPU 2
GPU 2
GPU 2
GPU 2
GPU 2
GPU 2
All-Gather comm.
Point-to-Point comm. GPU 3
GPU 3
(a) ∆D = 3 · D GPU 1
copy
copy
×k copy
GPU 3 copy
GPU 3
(b) ∆D = 3 · D · (3 − 1) 3
×k copy
GPU 2
Reduce-Scatter comm.
GPU 3
All-to-All All-to-All Dispatch dispatch (k routed routed (k expert) experts)
×k copy
(e) ∆D = 3 · D · (3 − 1) · kr 3
(2)
in which D denotes the size of the full gradient or activation tensor on each GPU that the communication primitive acts on, and kr denotes the number of routed experts. And for a two-layer Rail-Optimized architecture, X X X X X Rp = Rp + Rp + Rp + Rp (3) p∈P
A. Designing the Core Metric: Switching Efficiency
Point-to-Point All-Gather Reduce-Scatter All-Reduce All-to-All dispatch All-to-All combine
GPU 3
All-Reduce comm. GPU 3
(c) ∆D = 3 · D 3
GPU 1
GPU 1
×k
×k
GPU 2
GPU 2
×k
×k
GPU 3
GPU 3
×k
×k
GPU 3
GPU 3
(d) ∆D = 3 · D GPU 1 ×k
All-to-All All-to-All Dispatch combine (k (k routed routed expert) experts)
reduce
GPU 2 ×k
reduce
GPU 3 ×k
reduce
(f) ∆D = 3 · D · (3 − 1) · 1 3
Fig. 3: The effective data volume (∆D ) for each communication primitive. Assuming uniform token distribution for All-to-All dispatch and All-to-All combine. Grid-shaped shards represent the data that did not undergo network communication.
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
the exact calculation of ∆Di under different primitives can be categorized as: 1. Non-reduction primitives (Point-to-Point, All-Gather, All-to-All dispatch): All received data is directly usable for subsequent neural network computation. Therefore, ∆Di equals the aggregate newly received data volume. For instance, in All-to-All dispatch, Ni GPU receives Ni · NDi · (Ni − 1) · kr unique shards, yielding ∆Di = Ni · NDi · (Ni − 1) · kr . 2. Reduction primitives (Reduce-Scatter, All-Reduce, Allto-All combine): Data is usable only after reduction. Therefore, ∆Di is calculated based solely on the final reduced output. For instance, in All-to-All combine, although its raw communication volume might match that of All-to-All dispatch (without INC), each GPU ultimately retains only its final reduced shards, giving ∆Di = Ni · NDi · (Ni − 1) · 1. P On the other hand, p Rp is defined as the aggregate data rate of all electrical switch ports in the network A, including those within GPUs capable of data forwarding, as exemplified by the two-layer Rail-Optimized architecture in Fig. 4.
5
as follows: P P recv g∈G Dg i∈O ∆Di P · η= P recv g∈G Dg p∈P Rp T | {z } | {z }
γ:Data Efficiency µ:Network Efficiency
P recv P RT P ∆Di g Dg p 0 fp (t)dt i P = P recv · P R T · fp (t)dt g Dg p Rp T p 0 | {z } | {z } {z } | γ:Data Efficiency
δ:Routing Efficiency
(4)
θ:Port Utilization
where the composite metric µ (Network Efficiency) provides a network-level perspective on efficiency, and the fine-grained factors pinpoint the specific type of efficiency bottleneck, as elaborated below: 1. Data Efficiency (γ): This factor isPdefined as the ratio of the total effective data increment P ( ∆Di ) to the total data volume received by all GPUs ( Dgrecv ). It highlights the inefficiency stemming from the redundancy in the implementation of communication. Essentially, it reveals data redundancy (rbyte ) caused by redundant reception within the implementation of communication primitives: P ∆D γ = P i∈O recvi g∈G Dg (5) 1 1 P P = = recv Dg − i ∆Di 1 + rbyte 1+ g P i ∆D i
Fig. 4: In a two-layer AIDC network, switching resources are distributed in switches, on server boards, and in GPUs. In essence, Switching Efficiency (η) directly measures how effectively a network architecture can utilize its total provisioned switching capacity to transfer computationally useful data. P By relating the workload-specific effective data throughput ( ∆D ) to the architecture-defined aggregate switching i P resource ( Rp ), this metric serves as a top-level indicator of the fundamental alignment between the workload and the network design.
B. Decomposing Switching Efficiency for Fine-Grained Bottleneck Analysis Although η provides a holistic score, it encompasses the effects of communication operation, network topology, and resource configuration. To enable a more granular, diagnostic capability (P3), we decompose η to facilitate fine-grained bottleneck analysis by isolating and quantifying the efficiency of different contributing factors. Formally, let Dgrecv be the data volume received by GPU g (Let G denote the set of GPUs) during time T , and fp (t) be the forwarding rate of egress port p at time t. η is decomposed
where rbyte represents the redundant bytes that must be received to yield one byte of effective data. A low γ indicates inefficient implementation of communication primitives, exemplified by operations like All-to-All combine and AllReduce without INC, where a high rbyte arises as GPUs receive extra intermediate data that is discarded after reduction. 2. Routing Efficiency (δ): This factor is defined as the ratio of the total data volume received by all GPUs to the Raggregate data volume forwarded by all switch ports P T ( fp (t)dt). It indicates how well the traffic flows can be 0 routed in the network topology. Equivalently, δ is the inverse of the volume-weighted average number of switch-forwarding actions n̄fwd of the data flows (assuming that the total data sent equals the total data received for all GPUs, and that each switch satisfies flow conservation): P recv g∈G Dg δ=P RT p∈P 0 fp (t)dt P (6) 1 1 f ∈F Vf P P = = = f ∈F (Vf ·nf ) n̄fwd P f ∈F (Vf · nf ) f ∈F Vf
where F is the set of all data flows, Vf is the volume of a flow f , and nf is its number of switch-forwarding actions. The term n̄fwd is thus the volume-weighted switch-forwarding count. A low δ value, therefore, directly reveals the case of topological mismatch that forces significant data volume to undergo multiple forwarding actions, a situation corresponding to a high n̄fwd .
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
6
3. Port Utilization (θ): This factor is defined as the ratio of the aggregate data forwarding rate on all switch P ports to the aggregate switching capacity time product ( Rp T ). It assesses the utilization efficiency of provisioned resources on the network topology by revealing efficiency across both the spatial (θspatial ) and temporal (θtemporal ) dimensions (Letting Pactive ⊆ P be the set of active ports with non-zero forwarding rates, i.e., fp (t) > 0): RT
P
p∈P
θ=
0
fp (t)dt
C. Generalizing Efficiency Metrics for Long-term Analysis To further assess production AIDCs that typically encounter heterogeneous LLM training workloads with distinct traffic patterns in their life cycle, we generalize the definition of the efficiency metrics for a single workload to a workload profile. Formally, consider a set of K distinct training workloads, denoted as W = {w1 , w2 , . . . , wK } that occur within an observation window TH , the generalized form of Switching Efficiency η is defined as:
P
p∈P Rp T
P
Rp p∈P · = P active p∈P Rp
(7)
RT
P
p∈Pactive
T
fp (t)dt 0
P
p∈Pactive Rp
PK P
where θspatial represents the fraction of total deployed port capacity that is activated, and θtemporal represents the intensity of that activated capacity utilized for data forwarding. A low θ value thus indicates inefficient spatial or temporal resource configuration. For instance, an inappropriate allocation of dedicated resources for different communication phases (e.g., providing insufficient bandwidth for TP but excessive for PP) would degrade the utilization of alternately activated network subsystems, dually manifests as a reduced θspatial within each phase and a lowered θtemporal across the entire iteration. Table II demonstrates the application of the fine-grained metrics in analyzing principal efficiency bottlenecks within typical scenarios, highlighting how each metric reveals specific inefficiencies arising from the implementation of communication primitives, network topologies, or resource allocations. TABLE II: Bottleneck examples in AIDC identified by the fine-grained metrics Metrics
Examples of Identified Bottleneck
Data
Redundant reception in reductions without INC:†
Efficiency (γ)
• Reduce-Scatter (Recv: n − 1 byte, Eff: 1 byte): 1 rbyte = n − 2 → γ = n−1
• All-Reduce (Recv: 2(n − 1) byte, Eff: n byte): n → γ = 2(n−1) rbyte = n−2 n
• A2A combine (Recv: (n − 1)kr byte, Eff: n − 1 byte): rbyte = kr − 1 → γ = k1
k=1
η̄ =
= θspatial · θtemporal
TH
i∈Owk ∆Di,wk
(8)
P
p∈P Rp
Correspondingly, the generalized form of the efficiency metrics γ̄, δ̄, θ̄, and µ̄ can be derived analogously by applying the decomposition logic in Eq. (4). Specifically, assuming that workloads occur sequentially PK without overlap, such that TH = k=1 Tk = Ttotal where Tk is the duration of workload wk , Eq. (8) can be reformulated into a more concise and intuitive form: PK P η̄ = =
=
k=1
TH K X k=1 K X
Tk TH
i∈Owk ∆Di,wk
P
p∈P Rp
P
i∈Owk ∆Di,wk
Tk
P
p∈P Rp
! =
K X Tk k=1
Ttotal
· ηk
(9)
λk · η k
k=1
This simplified form decouples the analysis from a specific observation window TH , allowing the overall efficiency to be computed as a weighted sum of individual efficiencies ηk using the time-based weights λk . This facilitates the theoretical analysis and comparison of network architectures under standardized or projected operational conditions within the life cycle of the AIDCs. Nevertheless, in scenarios with concurrent workloads or resource contention, direct measurement via Eq. (8) is required to maintain accuracy.
r
Routing Efficiency (δ)
Traffic requires multi-forwardings within the network: • Inter-Pod traffic over Leaf-Spine-Core-Spine-Leaf: n̄fwd ≈ 5 → δ ≈ 15 • Inter-server All-to-All traffic over Leaf-Spine-Leaf: n̄fwd ≈ 3 → δ ≈ 13 • All-to-All traffic along one dimension on 3D-Torus: n̄fwd ∝ kr → δ ∝ k1 r
Port Utilization (θ)
Allocation of one-dimension switching port resources on a 3D-Torus to TP/EP traffic in 3D parallelism: • TP phase (utilizes bidirectional link via ring algorithm): θspatial = 26 , θtemporal ≈ 1 → θ ≈ 13 • EP phase (utilizes unidirectional link w/o ring algorithm): θspatial = 16 , θtemporal ≈ 1 → θ ≈ 16
† n: number of communication participants; k : number of routed experts. The r
received bytes and the effective bytes are counted across n GPUs; see [31, 48] and Eq. (2) for the detailed formulas for each primitive.
D. The Switching Efficiency Framework: A Hierarchical Metric Set for AIDC Network Analysis Collectively, the metrics defined in this section constitute the Switching Efficiency Framework, providing a hierarchical structure for top-down network communication efficiency analysis. At the holistic level, η and µ offer a top-level assessment of communication effectiveness. The fine-grained metrics (γ, δ, θ) then facilitate a diagnostic analysis, isolating bottlenecks related to the communication algorithm, network topology, and resource allocation. Finally, the generalized metrics (η̄, µ̄, γ̄, δ̄, θ̄) extend these analytical capabilities to long-term efficiency assessments under diverse production workloads. Table III summarizes the components of this framework and their analysis focus.
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
TABLE III: Components in Switching Efficiency Framework Level
Metrics
Analysis Focus
Holistic
η
Effectiveness of end-to-end communication
µ
Effectiveness of network-level communication
γ
Data redundancy in communication algorithm
δ
Data forwarding count imposed by topology
θ
Utilization of switching resource
η̄, µ̄, γ̄, δ̄, θ̄
Long-term efficiency under diverse workloads
Fine-grained
Long-term
7
Step 3: Scale Workload Parameters. For each parallelism configuration, a specific LLM workload is specified by scaling model architectural parameters. The number of layers, hidden dimension size, number of experts (for MoE model), and batch size are set to be respectively proportional to the parallelism degrees p, t, e, and the DP degree of attention layer d. The scaling coefficients are derived from the two typical model architectures: GPT-3 (for dense model) and DeepSeek-V3 (for MoE model), ensuring that the specified workloads serve as scaled analogues of these seminal architectures. The overall procedure to specify the workloads is summarized in Table IV. TABLE IV: Summary of Workload Specification
IV. A NALYZING C OMMUNICATION E FFICIENCY WITH THE S WITCHING E FFICIENCY F RAMEWORK UNDER LLM TRAINING WORKLOADS
Having established the theoretical foundation of the switching efficiency framework, this section demonstrates its practical utility. Based on the simulation setup detailed below, this section applies the proposed switching efficiency metric framework to dissect the communication efficiency of the production network architectures under various LLM workloads and network design parameters.
Parameter
Configuration Setting
Model Types
Dense models and Mixture-of-Experts (MoE) models.
Parallelism Strategy
• Dense: 3D parallelism (DP, PP, TP) with parallelism degrees denoted as (d, p, t). • MoE: A decoupled approach combining 2D (DP, PP) for Attention layers (denoted as (d, p)) and 3D (EDP, PP, EP) for MoE layers (denoted as (de , p, e)).
Parallelism Configuration
For a cluster of size N , configurations are enumerated as: • Dense: All tuples (d, p, t) where d · p · t = N . • MoE: All tuples (de , p, e) where de · p · e = N . The DP degree for Attention layers (d) is then determined by d = de · e, which ensures that d · p = N .
Parallelism Constraints
N . • PP degree (p): 2 ≤ p ≤ 16 N • EP degree (e): 2 ≤ e ≤ 32 (with 2e routed experts). • TP degree (t): t ≤ GPUs per server (e.g., 8).
Workload Scaling
Model parameters are scaled proportionally to parallelism degrees: • # Layers ∝ PP degree (p) • Hidden Dimension ∝ TP degree (t) • # Experts ∝ EP degree (e) • Batch Size ∝ DP degree of Attention layers (d)
Reference Models
Scaling coefficients are derived from typical models: GPT-3 (for dense model) and DeepSeek-V3 (for MoE model).
A. Simulation Setup 1. Workload Specification. A suite of workloads is specified based on hybrid parallelism strategies for both dense and MoE models. Specifically, for dense models, we employ a 3D (DP, PP, TP) strategy, with their parallelism degrees denoted by the tuple (d, p, t). For MoE models, we adopt an approach similar to that of the training of DeepSeek-V3, which replaces TP with EP and decouples DP between Attention and MoE layers, resulting in two configurations: a 2D (DP, PP) strategy for Attention layers with degrees (d, p), and a 3D (EDP, PP, EP) strategy for MoE layers with degrees (de , p, e). Based on these strategies, a set of workload configurations is then specified through the following procedure: Step 1: Enumerate Parallelism Configurations. All valid parallel configurations for a cluster of size N (total GPUs) are first enumerated. For dense models, a valid configuration is represented by a tuple (d, p, t) such that d · p · t = N . For MoE models, initial tuples (de , p, e) for the MoE layer are first identified such that de · p · e = N . The Attention layer’s DP degree d is then determined by the relationship d = de · e, ensuring d · p = N . Step 2: Apply Practical Constraints. To ensure realism, several constraints derived from established training practice are imposed on these configurations. The size of PP is conN strained between 2 (to avoid overly shallow models) and 16 (to prevent excessively deep pipelines). Similarly, the size of EP is constrained between 2 (with 2e routed experts to ensure the N presence of All-to-All communication) and 32 (to avoid the formation of impractically large EP groups). Additionally, following typical hardware placement constraints, the maximum TP size is limited to the number of GPUs within a single highbandwidth domain — i.e., one server in the Rail-Optimized architecture (e.g., 8 GPUs).
2. Traffic Modeling. For each workload instance, its traffic demands are modeled as a sequence of communication primitives, where the parallel strategy dictates the primitive type and the model hyperparameter determines the traffic volume [31, 48]. Furthermore, similar to the approach in Megatron [49–51] and DeepSpeed [52], data transfers for TP and DP are implemented as the Reduce-Scatter and All-Gather primitives using the Ring algorithm [32]. Data transfers for PP are implemented as Point-to-Point communication between GPUs in adjacent pipeline stages. Finally, the All-to-All communication for EP is modeled as a uniform traffic within each expert group, a simplification similar to the analysis in Rail-Only [25]. To isolate the analysis on communication efficiency, GPU computation time is excluded from our study. 3. Network Architectures. Our analysis focuses on two architectural paradigms prevalent in hyper-scale AI data centers: the 3D-Torus and the Rail-Optimized architecture. The network configurations are specified in two parts: a baseline configuration for direct efficiency dissection (Section IV-B), and a series of network parameter variation cases for investigating specific design trade-offs using our metrics (Section IV-C).
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
8
for the Rail-Optimized, multi-plane designs (single-, dual-, quad-, and octa-plane) analogous to the HPN architecture [1] are also included in this scalability analysis.
• Baseline Network Configurations: 3D-Torus. Each GPU is equipped with a 6-port switching chip. And assume the topology is reconfigurable via optical circuit switches (OCS), such that the communication within each parallel dimension in 3D parallelism is mapped to a dedicated physical dimension of the 3D-Torus. Rail-Optimized. An 8-GPU server configuration featuring a single-level intra-server switching fabric and an inter-server switching fabric (64-port switches), where the link bandwidth ratio (tiered bandwidth ratio) of the intra- to inter-server switching network is set to 9:1 (similar to the NVIDIA DGX H100 system [53]). The cluster scale is set to 4096 for both architectures unless otherwise specified. • Network Parameters: Case 1: Tiered Bandwidth Ratio. Ratio ranging from 1:1 to 17:1 is evaluated for both architectures. For the 3D-Torus, we model this tiered ratio by increasing the port count along one dimension of the on-GPU switches, creating a high-bandwidth link that is dedicated to a traffic-intensive parallel dimension: EP for MoE models, or TP (or PP, when TP size is 1) for dense models. Case 2: Server Size. Server size ranging from 8 to 256GPU is evaluated in Rail-Optimized architecture, with the assumption that a single-level intra-server switch fabric is maintained as the server size scales. Case 3: In-Network Computing. INC applied for the AllReduce (substituting the All-Gather & Reduce-Scatter primitives) within the TP phase by the intra-server switch of Rail-Optimized is evaluated. Our analysis focuses on this practical application, excluding two scenarios where practical implementation of INC is challenged: its application in 3D-Torus, where INC’s tree-based communication pattern is incompatible with the distributed switch fabric [54], and its use for the All-to-All combine in MoE workloads, which is challenged by the fine-grained reduction scope [55]. Case 4: Cluster Scale. Cluster size scales from 512 to 65,536 GPUs is evaluated for both architectures. In particular,
1
.tor: = 0:64
/tor: = 1:00
3tor: = 0:32
.opt: = 0:64
/opt: = 0:96
3opt: = 0:51
B. Efficiency Dissection of Baseline Network Architectures Fig. 5 presents the dissection of 3D-Torus and RailOptimized architectures on a 4096-GPU cluster, proceeding from a fine-grained diagnosis of data, routing, and resource utilization efficiency (γ, δ, θ) to a macroscopic assessment of overall performance (η, µ) under diverse dense and MoE model workloads. 1) Observations from the fine-grained Data, Routing, and Port Utilization Efficiency: 1a. Data Efficiency (γ) reveals that redundant reception of data is a considerable bottleneck that constrains the endto-end communication efficiency. In particular, the redundant data reception rbyte in implementations of Reduce-Scatter/AllReduce and All-to-All combine without in-network computing, leads to lower γ for workloads featuring more reductionintensive communication (e.g., those with larger TP groups or larger EP groups). In aggregate, both the 3D-Torus and RailOptimized exhibit low γ̄ of 0.64 for dense model workloads (Fig. 5a) and 0.53 for MoE model workloads (Fig. 5b). This data redundancy imposes a hard ceiling on end-to-end communication efficiency. 1b. Routing Efficiency (δ) highlights the challenge of MoE models for existing network architectures and underscores the deficiency of 3D-Torus in this scenario. Specifically, multihop transmissions (i.e., large n̄fwd ) would result in a small δ, such as the cases of All-to-All traffic for MoE workloads in a 3D-Torus or traffic traversing a Leaf-Spine-Leaf path in a Rail-Optimized with an EP size > 8. Overall, these cases cause δ̄ on the 3D-Torus to plummet from 1.00 for dense workloads — where its reconfigurability provides perfectly matched single-hop paths for PTD parallelism — to 0.05 for MoE workloads. For the Rail-Optimized, it decreases from
1
.tor: = 0:53
/tor: = 0:05
3tor: = 0:17
.opt: = 0:53
/opt: = 0:41
3opt: = 0:21
3D-Tor. / Rail-Opt.
.
.
3D-Tor. / Rail-opt. 0.5
TP1 TP2 TP4 TP8
0 0
0.2
0.4
0 0.2 0.4
0 0
0.6
0.8
0.8
1 1
2opt: = 0:32
7tor: = 0:32
2tor: = 0:004
7opt: = 0:49
Rail-opt.
Rail-Opt.
0.4
2
0
0.25
0.5
7
(a) Dense model training workloads.
EP32 EP64 EP128
0
0.1
0.6
0.6
0.8
2opt: = 0:046
0.2
2
0 0.2 0.4
0.8 1 1
/
3
3D-Tor.
0.2
0.4
/
3D-Tor.
0
0.2
0.6
3 2tor: = 0:21
EP2 EP4 EP8 EP16
0.5
0.3
7tor: = 0:008
0.4 0
0.1
7opt: = 0:085
0.2
0.3
0.4
0.5
7
(b) MoE model training workloads.
Fig. 5: Efficiency dissection of 3D-Torus vs. Rail-Optimized for (a) dense and (b) MoE workloads on a 4096-GPU cluster. The top plot provides diagnosis of Data (γ), Routing (δ), and Utilization (θ) efficiency. The bottom plot presents the macroscopic assessment of Switching (η) and Network (µ) Efficiency. Dashed lines link identical workloads across the two architectures.
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
9
0.96 — where dominant TP traffic is confined within a server — to 0.41. This sharp decline, particularly for the 3D-Torus, reveals their topological inadequacy for handling large-scale All-to-All traffic in MoE workloads. 1c. Port Utilization (θ) reveals the superiority of RailOptimized and highlights the challenges posed by MoE models in terms of resource utilization efficiency. To be specific, the Rail-Optimized’s tiered bandwidth design has an advantage over the 3D-Torus’s symmetric design, as its high bandwidth intra-server subnetworks achieve higher θspatial when processing intensive localized traffic (e.g., TP traffic within a server). This leads to a superior θ̄ on dense model workloads (0.51 vs. 0.32). However, this advantage diminishes for MoE workloads, as cross-server All-to-All traffic underutilizes the fabric within the server, lowering θtemporal . In this scenario, both systems face a considerable challenge, with their θ̄ values dropping sharply to 0.21 and 0.17 respectively. 2) Synthesizing fine-grained factors by holistic efficiency metrics: The interplay of these fine-grained factors is synthesized by the holistic metrics (η and µ), leading to the following findings: 2a. The Rail-Optimized architecture exhibits higher efficiency. Except for marginal cases at TP1, the Rail-Optimized design consistently achieves higher switching efficiency (η) across the remaining parallelism configurations. The generalized holistic metric η̄ quantifies this overall advantage: RailOptimized outperforms 3D-Torus by 1.6x in dense models (0.32 vs. 0.21), and dramatically extends this lead to 11.5x in MoE models (0.046 vs. 0.004). 2b. MoE Workloads severely degrade efficiency in both architectures. Both the network architectures suffer a sharp decline in switching efficiency (η) when shifting from dense model to MoE model workloads. Overall, η̄ of the RailOptimized, drops by 85.6% from 0.32 to 0.046. The degradation is more acute for the 3D-Torus, where η̄ falls by 98.1%,
from 0.21 to 0.004. 2c. Data redundancy limits overall switching efficiency. A persistent gap is observed between network efficiency (µ̄) and the overall switching efficiency (η̄). Overall, exemplified by Rail-Optimized architecture, this efficiency gap causes a 35% efficiency loss for dense workloads (µ̄ = 0.49 → η̄ = 0.32) and a 46% loss for MoE workloads (µ̄ = 0.085 → η̄ = 0.046), underscoring a fundamental limitation that network-level optimizations alone cannot overcome. C. Efficiency Analysis of Varying Network Design Parameters The previous subsection applied the Switching Efficiency Framework for a top-down analysis of communication efficiency in fixed network architectures. This subsection extends its application to analyze how network parameters — including tiered bandwidth ratio, server size, in-network computing, and cluster scale — influence communication efficiency. 1) Impact of the Tiered Bandwidth Ratio on Switching Resource Utilization: To analyze the impact of the tiered bandwidth ratio on switching resource utilization, θ̄ is measured, revealing that strategically tuning the tiered bandwidth ratio leads to workload-dependent efficiency gains. For dense workloads (Fig. 6a), the Rail-Optimized architecture experiences only a marginal gain from bandwidth ratio tuning, with its overall efficiency (TP-All) rising slightly from θ̄ = 0.51 at baseline to 0.52 at optimum in our simulation configuration. In contrast, the 3D-Torus benefits far more significantly. At its optimal configuration, its overall port utilization (TP-All) increases by 78%, rising from a baseline of 0.32 to 0.57. This substantial gain results from creating a highbandwidth axis for the dominant parallel dimension (TP, or PP when TP=1), which enables the architecture to accommodate the traffic hierarchy. For MoE workloads (Fig. 6b), the architectures show opposing responses to bandwidth hierarchy tuning. The 3D-
Rail-Optimized
0.6
Rail-Optimized
0.6
(0.40 + 0.1) 0.5
(0.51)
0.5
(0.52)
0.4
3
3
0.4 0.3
TP1
TP2
TP4
TP8
TP-All
0.2 0.1
0.1
0 3
5
7
9 11 (baseline) (opt.)
13
15
17
0.5
(0.57)
3 (0.32)
0.2
3
EP8 5
EP16 7
EP32
9 (baseline)
EP64 11
13
EP128
EP-All
15
17
3D-Torus (0.44 + 0.1) + 0.1
0.4
0.4 0.3
EP4
0.6
0.6 0.5
EP2 1 (opt.)
3D-Torus
0.7
(0.21 + 0.1)
0.3
0.2
1
3
+ 0.1
0.3 (0.17 + 0.1) 0.2
TP1
TP2
TP4
TP8
0.1
TP-All
EP2
0.1
EP4
EP8
EP16
EP32
EP64
EP128
EP-All
15
17 (opt.)
0 1 (baseline)
3
5 7 9 11 (opt.) Tiered Bandwidth Ratio
13
15
(a) θ̄ vs. bandwidth ratio (dense models).
17
1 (baseline)
3
5
7
9
11
13
Tiered Bandwidth Ratio
(b) θ̄ vs. bandwidth ratio (MoE models).
Fig. 6: Impact of tiered bandwidth ratio on generalized Port Utilization (θ̄), evaluated across dense and MoE model workloads with different TP and EP sizes. For clarity, the EP-All curves are vertically shifted by +0.1.
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
10
0.7 1
EP2 EP32
0.6
EP4 EP64
EP8 EP128
EP16 EP-All
0.6
(0.99)
0.9
(0.58) 0.5
(0.58)
0.5
0.8
0.4 0.4
7
/
3
0.7
0.3
0.6 0.3 0.5
0.2
(0.41)
(0.21)
0.4
EP2 EP32
EP4 EP64
EP8 EP128
EP16 EP-All
(0.09)
0.2
EP2 EP32
EP4 EP64
EP8 EP128
EP16 EP-All
0.3 8
16
32
64
Server Size
(a) δ̄ vs. server size.
128
256
8
16
32
64
128
256
0.1
0 8
16
Server Size
32
64
128
256
Server Size
(b) θ̄ vs. server size.
(c) µ̄ vs. server size.
Fig. 7: Influence of server size on communication efficiency under MoE training workload, with trends for generalized (a) Routing Efficiency (δ̄), (b) Port Utilization (θ̄), and (c) Network Efficiency (µ̄).
communication efficiency. As shown in Fig. 8, γ̄ reaches a near-perfect 0.99 across all evaluated TP sizes (note that the Reduce-Scatter operation in DP communication still incurs redundant data reception). This high efficiency stems from the application of INC to the dominant TP communication, which eliminates redundant data reception inherent in the reduction process, thereby raising its data efficiency to the ideal value of 1. The improvement in data efficiency, in turn, boosts the overall η̄ (TP-All) from 0.32 to 0.45, demonstrating the gains in end-to-end communication efficiency from applying INC.
0.58
0.64
0.99
0.99
w TP INC
0.99 0.67
0.99
.
1
0.99
w/o INC
0.5
0.45
TP-All
0.32
0.34
0.38
0.4
0.38
0.6
TP8
0.51
TP4
0.32
TP2
0.47
0
2
Torus benefits significantly, with its overall port utilization (EP-All) increasing by 159% from a baseline θ̄ of 0.17 to 0.44 at the optimal configuration by dedicating a high-speed axis to EP communication. In contrast, for the Rail-Optimized architecture, a higher ratio is counterproductive. Its overall θ̄ (EP-All) drops from 0.40 at the 1:1 ratio to 0.21 at the baseline: when EP size exceeds server size, dominant all-toall traffic is forced onto lower-bandwidth inter-server links, underutilizing the high-bandwidth intra-server fabric. 2) Influence of Server Size on Network-Level Efficiency for MoE model workloads: To assess the influence of server size on network-level communication efficiency for MoE workloads, whose dominant traffic (All-to-All traffic) regularly spans servers in practice, network-level metrics δ̄, θ̄, and µ̄ are leveraged to quantify this influence, showing that a larger server size in the RailOptimized leads to considerable efficiency improvements in the training of MoE models. As shown in Fig. 7, when the server size increases to match a specific EP size, the corresponding δ̄ and θ̄ for that workload become approximately 2 to 3 times their baseline values, as the All-to-All communication in EP is now entirely contained within a single serve, which eliminates the multi-level switch forwarding and removes the utilization bottleneck formed by bandwidth disparity between intra- and inter-server. In the general case, increasing the server size allows a greater portion of traffic to be resolved via single-level switches forwarding, rather than requiring multi-level forwarding, which promotes a rise in δ̄. But the reduced utilization of the top-level switches in turn lowers θ̄. µ̄ balances these two aspects, exhibiting a more stable trend. Overall (EP-All), as the server size increases from 8 to 256, µ̄ rises sharply from 0.09 to 0.58, demonstrating the substantial efficiency gains from enlarging the server size. 3) Efficiency Improvement from In-Network Computing for Dense Model Workloads: To quantify the benefits of INC for dense model workloads, γ̄ is measured to isolate its effect on eliminating redundant reception, and η̄ is measured to analyze the overall impact, the results show that the application of INC eliminates redundant data reception, which translates to the increase in end-to-end
0.2
0
TP2
TP4
TP8
TP-All
Fig. 8: Efficiency improvement from the In-network computing under dense model training workloads in the RailOptimized architecture (note that the TP=1 case is omitted, as it involves no TP communication). 4) Impact of Network Architecture Scale on Overall Communication Efficiency: To investigate architectural scalability, η̄ is employed to track overall communication efficiency across a wide range of cluster sizes, revealing that different architectures exhibit distinct efficiency degradation trends as the cluster size grows. Under dense model workloads (Fig. 9a), 3D-Torus maintains a persistently low η̄ of 0.21. This is because, while 3DTorus provides a single-hop topology for the PTD-parallelism traffic — a property that remains constant as it scales —
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
11
0.4
0.2
0.35
0.15
3D-Torus Rail-Optimized (single-plane) Rail-Optimized (dual-plane) Rail-Optimized (quad-plane) Rail-Optimized (octa-plane)
2
0.25
2
0.45
0.3
0.25
0.2 512
0.1
3D-Torus Rail-Optimized (single-plane) Rail-Optimized (dual-plane) Rail-Optimized (quad-plane) Rail-Optimized (octa-plane)
1024
2048
4096
8192
16384
0.05
32768
65536
Cluster Scale
(a) η̄ vs. cluster scale (dense models).
0 512
1024
2048
4096
8192
16384
32768
65536
Cluster Scale
(b) η̄ vs. cluster scale (MoE models).
Fig. 9: Generalized Switching Efficiency (η̄) of 3D-Torus and Rail-Optimized architectures as cluster scales from 512 to 65,536 GPUs under (a) dense and (b) MoE model training workloads.
its symmetric fabric imposes a uniform bandwidth allocation, which is ill-suited to the imbalanced LLM training traffic. In contrast, while Rail-Optimized architectures exhibit a stepwise decline in efficiency, multi-plane designs delay this degradation. For instance, η̄ of the single-plane architecture drops from 0.37 at 2048 GPUs to 0.32 at 4096 GPUs, reflecting the addition of a new switching layer that incurs inefficiency in both δ̄ and θ̄ (due to potential increased forwarding hops or idleness of upper-layer switches). The octa-plane design, however, sustains this high efficiency of η̄ ≈ 0.37 up to 16384 GPUs, as multi-plane design delays the need to add switching layers to support cluster expansion. Under the MoE model workloads (Fig. 9b), all architectures exhibit a continuous decline in η̄ when EP sizes scale with cluster size. For the 3D-Torus, η̄ declines rapidly from 0.033 to below 0.001 for scaling from 512 to 16384 GPUs, because larger EP groups amplify the routing inefficiency of the Allto-All traffic on its neighbor-only connectivity topology. For the Rail-Optimized architectures, η̄ also declines continuously, as growing EP groups force more All-to-All traffic onto inefficient, multi-hop inter-server links. Despite this overall decline, multi-plane designs prove more scalable, with the octa-plane design, for instance, maintaining a slight advantage over the single-plane as the cluster scales from 4096 to 16384 GPUs (η̄ ranging from 0.053 to 0.046 vs. 0.046 to 0.039). V. I NSIGHTS FOR F UTURE N ETWORK E VOLUTION Fig. 5 reveals that communication efficiency, across both holistic and fine-grained metrics, is highly dependent on communication-intensive strategies like TP and EP. Given that the communication arising from these strategies is typically mapped to the High-Bandwidth Domain (HBD, e.g., high bandwidth NVLink within the server), and that these HBDs constitute the majority of the network switching resources, we argue that co-optimizing the design and usage of HBDs is a pivotal pathway to improving communication efficiency. Empowering the server-scale HBD in Rail-Optimized. For the Rail-Optimized architecture, expanding the server size to encompass the entire All-to-All communication within an EP
group effectively boosts both port utilization (θ) and routing efficiency (δ) for MoE model training, as shown in Fig. 7. Similarly, equipping intra-server switches with INC capabilities prevents GPUs from receiving redundant, unreduced data during TP, leading to a significant improvement in data efficiency (γ), as shown in Fig. 8. Both the server-level enhancements yield more profound efficiency returns compared to adjusting the bandwidth resource allocation (Fig. 6) or adopting a multiplane switching plane design (Fig. 9). Curating workloads for the neighbor-only connectivity HBD of 3D-Torus. The 3D-Torus equips each accelerator with lowradix switches to form a distributed switching topology [26], forming an HBD that spans the entire cluster but offers limited connectivity flexibility. However, executing workloads that incorporate 3D hybrid parallelism on a 3D-Torus requires a oneto-one mapping between the orthogonal parallel dimensions and the network’s physical dimensions [23], which leads to inefficiency due to bandwidth fragmentation, as quantified by θ. Furthermore, mapping the All-to-All traffic of EP to a single dimension of the Torus necessitates multi-hop transmissions, resulting in routing inefficiency, as quantified by δ. In the absence of switch chips that permit flexible bandwidth allocation across dimensions (an alternative explored in Fig. 6), specializing workloads becomes a necessary strategy to enhance efficiency. For instance, dedicating a single Torus network exclusively to model parallelism would allow its high-volume data transfers to leverage switch ports across more network dimensions and reduce the communication diameter, a practice exemplified by Google in training its Gemini models [4, 56]. VI. C ONCLUSION AND F UTURE W ORK This paper introduces the Switching Efficiency Framework, a hierarchical metrics set for top-down analysis of how effectively network switching capacity is converted into computationally useful data throughput in LLM training. By shifting the evaluation paradigm from monolithic performance benchmarking to diagnostic analysis, this work provides actionable guidance for the resource-efficient design of future AIDCs. Applying this framework, we revealed the (mis)alignment
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
between the design patterns of mainstream architectures such as 3D-Torus and Rail-Optimized, and the structured traffic patterns they support. Moreover, the fine-grained metrics enabled us to diagnose the root causes of their specific efficiency bottlenecks and identify potential optimization pathways, thereby showcasing the practical utility of the proposed framework. VII. ACKNOWLEDGEMENTS This work was supported in part by the National Key Research and Development Project of China under Grant 2024YFB2908301 and in part by the National Natural Science Foundation of China (NSFC) under Grant 62331017. (Corresponding author: Weiqiang Sun.) R EFERENCES [1] K. Qian, Y. Xi, J. Cao, J. Gao, Y. Xu, Y. Guan, B. Fu, X. Shi, F. Zhu, R. Miao, et al., “Alibaba HPN: A Data Center Network for Large Language Model Training,” in Proceedings of the ACM SIGCOMM 2024 Conference, ser. ACM SIGCOMM ’24, New York, NY, USA: Association for Computing Machinery, 2024, pp. 691–706. [2] Z. Jiang, H. Lin, Y. Zhong, Q. Huang, Y. Chen, Z. Zhang, Y. Peng, X. Li, C. Xie, S. Nong, et al., “MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs,” in 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24), 2024, pp. 745–760. [3] A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, et al., The Llama 3 Herd of Models, 2024. arXiv: 2407.21783 [cs]. [4] G. Comanici, E. Bieber, M. Schaekermann, I. Pasupat, N. Sachdeva, I. Dhillon, M. Blistein, O. Ram, D. Zhang, E. Rosen, et al., Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities, 2025. arXiv: 2507. 06261. [5] Q. Meng, H. Zheng, Z. Zhang, C. Lao, C. Huang, B. Li, Z. Zhu, H. Lu, W. Dang, Z. Lin, et al., “Astral: A Datacenter Infrastructure for Large Language Model Training at Scale,” in Proceedings of the ACM SIGCOMM 2025 Conference, São Francisco Convent Coimbra Portugal: ACM, 2025, pp. 609–625. [6] xAI, Colossus, https://x.ai/colossus. [7] Z. Zhang, C. Chang, H. Lin, Y. Wang, R. Arora, and X. Jin, “Is Network the Bottleneck of Distributed Training?” In Proceedings of the Workshop on Network Meets AI & ML, ser. NetAI ’20, New York, NY, USA: Association for Computing Machinery, 2020, pp. 8–13. [8] E. Erdil and D. Schneider-Joseph, Data movement limits to frontier model training, 2024. arXiv: 2411.01137 [cs]. [9] K. Benyahya, A. G. Diaz, J. Liu, V. Lyutsarev, M. Pantouvaki, K. Shi, S. Y. Siew, H. Ballani, T. Burridge, D. Cletheroe, et al., “Mosaic: Breaking the Optics versus Copper Trade-off with a Wide-and-Slow Architecture and MicroLEDs,” in Proceedings of the ACM SIGCOMM 2025 Conference, ser. SIGCOMM ’25, New York, NY, USA: Association for Computing Machinery, 2025, pp. 234–247. [10] Y. Wei, T. Hu, C. Liang, and Y. Cui, “Communication Optimization for Distributed Training: Architecture, Advances, and Opportunities,” IEEE Network, vol. 39, no. 3, pp. 241–248, 2025, ISSN: 1558-156X. [11] B. Cottier, R. Rahman, L. Fattorini, N. Maslej, T. Besiroglu, and D. Owen, The rising costs of training frontier AI models, 2025. arXiv: 2405.21015 [cs]. [12] K. F. Pilz, J. Sanders, R. Rahman, and L. Heim, Trends in AI Supercomputers, 2025. arXiv: 2504.16026 [cs]. [13] M. Cowan, S. Maleki, M. Musuvathi, O. Saarikivi, and Y. Xiong, “MSCCLang: Microsoft Collective Communication Language,” in Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2, Vancouver BC Canada: ACM, 2023, pp. 502–514. [14] A. Shah, V. Chidambaram, M. Cowan, S. Maleki, M. Musuvathi, T. Mytkowicz, J. Nelson, O. Saarikivi, and R. Singh, “TACCL: Guiding Collective Algorithm Synthesis using Communication Sketches,” in 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23), 2023, pp. 593–612.
12
[15] B. Li, X. Wang, J. Wang, Y. Liu, Y. Gong, H. Lu, W. Dang, W. Zhang, X. Huang, M. Chen, et al., “TCCL: Co-optimizing Collective Communication and Traffic Routing for GPU-centric Clusters,” in Proceedings of the 2024 SIGCOMM Workshop on Networks for AI Computing, Sydney NSW Australia: ACM, 2024, pp. 48–53. [16] D. D. Sensi, T. Bonato, D. Saam, and T. Hoefler, “Swing: Short-cutting Rings for Higher Bandwidth Allreduce,” in 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24), 2024, pp. 1445–1462. [17] H. Song, “In-Network AllReduce Optimization with Virtual Aggregation Trees,” in Proceedings of the 2024 SIGCOMM Workshop on Networks for AI Computing, ser. NAIC ’24, New York, NY, USA: Association for Computing Machinery, 2024, pp. 54–60. [18] X. Zhao, Z. Zhang, and C. Wu, “AdapCC: Making Collective Communication in Distributed Machine Learning Adaptive,” in 2024 IEEE 44th International Conference on Distributed Computing Systems (ICDCS), Jersey City, NJ, USA: IEEE, 2024, pp. 25–35. [19] L. Zhao, S. Maleki, Z. Yang, H. Pourreza, and A. Krishnamurthy, ForestColl: Throughput-Optimal Collective Communications on Heterogeneous Network Fabrics, 2025. arXiv: 2402.06787 [cs]. [20] J. Dong, Z. Cao, T. Zhang, J. Ye, S. Wang, F. Feng, L. Zhao, X. Liu, L. Song, L. Peng, et al., “EFLOPS: Algorithm and System Co-Design for a High Performance Distributed Training Platform,” in 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA), 2020, pp. 610–622. [21] T. Hoefler, T. Bonato, D. De Sensi, S. Di Girolamo, S. Li, M. Heddes, J. Belk, D. Goel, M. Castro, and S. Scott, “HammingMesh: A network topology for large-scale deep learning,” in Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis, ser. SC ’22, Dallas, Texas: IEEE Press, 2022, pp. 1–18. [22] W. Wang, M. Khazraee, Z. Zhong, M. Ghobadi, Z. Jia, D. Mudigere, Y. Zhang, and A. Kewitsch, “TopoOpt: Co-optimizing Network Topology and Parallelization Strategy for Distributed Training Jobs,” in 20th USENIX Symposium on Networked Systems Design and Implementation, NSDI 2023, Boston, MA, April 17-19, 2023, M. Balakrishnan and M. Ghobadi, Eds., USENIX Association, 2023, pp. 739–767. arXiv: 2202.00433. [23] N. Jouppi, G. Kurian, S. Li, P. Ma, R. Nagarajan, L. Nai, N. Patil, S. Subramanian, A. Swing, B. Towles, et al., “TPU v4: An Optically Reconfigurable Supercomputer for Machine Learning with Hardware Support for Embeddings,” in Proceedings of the 50th Annual International Symposium on Computer Architecture, ser. ISCA ’23, New York, NY, USA: Association for Computing Machinery, 2023, pp. 1–14. [24] NVIDIA, “NVIDIA DGX SuperPOD: Next Generation Scalable Infrastructure for AI Leadership Reference Architecture Featuring NVIDIA DGX B200 — NVIDIA DGX SuperPOD: Next Generation Scalable Infrastructure for AI Leadership Reference Architecture Featuring NVDIA DGX B200,” NVIDIA Corporation, Tech. Rep., 2024. [25] W. Wang, M. Ghobadi, K. Shakeri, Y. Zhang, and N. Hasani, “Railonly: A Low-Cost High-Performance Network for Training LLMs with Trillion Parameters,” in IEEE Symposium on High-Performance Interconnects, HOTI 2024, Albuquerque, NM, USA, August 21-23, 2024, IEEE, 2024, pp. 1–10. arXiv: 2307.12169. [26] Y. Zu, A. Ghaffarkhah, H.-V. Dang, B. Towles, S. Hand, S. Huda, A. Bello, A. Kolbasov, A. Rezaei, D. Du, et al., “Resiliency at Scale: Managing Google’s TPUv4 Machine Learning Supercomputer,” in 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24), 2024, pp. 761–774. [27] H. Liao, B. Liu, X. Chen, Z. Guo, C. Cheng, J. Wang, X. Chen, P. Dong, R. Meng, W. Liu, et al., UB-Mesh: A Hierarchically Localized nDFullMesh Datacenter Network Architecture, 2025. arXiv: 2503.20377 [cs]. [28] Z. Yan, D. Li, L. Chen, D. Xiong, K. Gao, Y. Zhang, R. Yan, M. Zhang, B. Zhang, Z. Jiang, et al., “From ATOP to ZCube: Automated Topology Optimization Pipeline and A Highly Cost-Effective Network Topology for Large Model Training,” in Proceedings of the ACM SIGCOMM 2025 Conference, ser. SIGCOMM ’25, New York, NY, USA: Association for Computing Machinery, 2025, pp. 861–881. [29] X. Liao, Y. Sun, H. Tian, X. Wan, Y. Jin, Z. Wang, Z. Ren, X. Huang, W. Li, K. F. Tse, et al., “MixNet: A Runtime Reconfigurable OpticalElectrical Fabric for Distributed Mixture-of-Experts Training,” in Proceedings of the ACM SIGCOMM 2025 Conference, ser. SIGCOMM ’25, New York, NY, USA: Association for Computing Machinery, 2025, pp. 554–574. [30] C. Shou, G. Liu, H. Nie, H. Meng, Y. Zhou, Y. Jiang, W. Lv, Y. Xu, Y. Lu, Z. Chen, et al., “InfiniteHBD: Building Datacenter-
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
Scale High-Bandwidth Domain for LLM with Optical Circuit Switching Transceivers,” in Proceedings of the ACM SIGCOMM 2025 Conference, ser. SIGCOMM ’25, New York, NY, USA: Association for Computing Machinery, 2025, pp. 1–23. [31] W. Li, X. Liu, Y. Li, Y. Jin, H. Tian, Z. Zhong, G. Liu, Y. Zhang, and K. Chen, “Understanding Communication Characteristics of Distributed Training,” in Proceedings of the 8th Asia-Pacific Workshop on Networking, ser. APNet ’24, New York, NY, USA: Association for Computing Machinery, 2024, pp. 1–8. [32] P. Patarasuk and X. Yuan, “Bandwidth optimal all-reduce algorithms for clusters of workstations,” Journal of Parallel and Distributed Computing, vol. 69, no. 2, pp. 117–124, 2009, ISSN: 0743-7315. [33] B. Klenk, N. Jiang, G. Thorson, and L. Dennison, “An In-Network Architecture for Accelerating Shared-Memory Multiprocessor Collectives,” in 2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA), 2020, pp. 996–1009. [34] “2022 Fast Inter-GPU Communication with NCCL for Deep Learning Training, and More (a Magnum IO session) — GTC Digital Spring 2022 — NVIDIA On-Demand,” NVIDIA. [35] NVIDIA, Scaling Deep Learning Training: Fast Inter-GPU Communication with NCCL, https://www.nvidia.com/en-us/on-demand/session/ gtcspring23-s51111/. [36] A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, et al., DeepSeek-V3 Technical Report, 2024. arXiv: 2412.19437. [37] M. Khani, M. Ghobadi, M. Alizadeh, Z. Zhu, M. Glick, K. Bergman, A. Vahdat, B. Klenk, and E. Ebrahimi, “SiP-ML: High-bandwidth optical network interconnects for machine learning training,” in Proceedings of the 2021 ACM SIGCOMM 2021 Conference, ser. SIGCOMM ’21, New York, NY, USA: Association for Computing Machinery, 2021, pp. 657– 675. [38] Y. Jiang, Y. Zhu, C. Lan, B. Yi, Y. Cui, and C. Guo, “A Unified Architecture for Accelerating Distributed {DNN} Training in Heterogeneous {GPU/CPU} Clusters,” in 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20), 2020, pp. 463–479. [39] Z. Zhang, C. Chang, H. Lin, Y. Wang, R. Arora, and X. Jin, “Is Network the Bottleneck of Distributed Training?” In Proceedings of the Workshop on Network Meets AI & ML, Virtual Event USA: ACM, 2020, pp. 8–13. [40] A. Gangidi, R. Miao, S. Zheng, S. J. Bondu, G. Goes, H. Morsy, R. Puri, M. Riftadi, A. J. Shetty, J. Yang, et al., “RDMA over Ethernet for Distributed Training at Meta Scale,” in Proceedings of the ACM SIGCOMM 2024 Conference, Sydney NSW Australia: ACM, 2024, pp. 57–70. [41] T. Bonato, A. Kabbani, A. Ghalayini, M. Papamichael, M. Dohadwala, L. Gianinazzi, M. Khalilov, E. Achermann, D. De Sensi, and T. Hoefler, REPS: Recycled Entropy Packet Spraying for Adaptive Load Balancing and Failure Mitigation, 2025. [42] N. Blach, M. Besta, D. D. Sensi, J. Domke, H. Harake, S. Li, P. Iff, M. Konieczny, K. Lakhotia, A. Kubicek, et al., “A High-Performance Design, Implementation, Deployment, and Evaluation of The Slim Fly Network,” in 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24), 2024, pp. 1025–1044. [43] Y. Feng, T. Chen, Y. Wei, S. Shen, S. Wang, W. Li, K. Ma, and T. Hoefler, RailX: A Flexible, Scalable, and Low-Cost Network Architecture for Hyper-Scale LLM Training Systems, 2025. arXiv: 2507.18889 [cs]. [44] X. Han, Y. Lv, S. Zhao, Z. Liu, X. Liu, and X. Wang, LumosCore: Highly Scalable LLM Clusters with Optical Interconnect, 2025. arXiv: 2411.01503 [cs]. [45] H. Liu, R. Urata, K. Yasumura, X. Zhou, R. Bannon, J. Berger, P. Dashti, N. Jouppi, C. Lam, S. Li, et al., “Lightwave Fabrics: At-Scale Optical Circuit Switching for Datacenter and Machine Learning Systems,” in Proceedings of the ACM SIGCOMM 2023 Conference, New York NY USA: ACM, 2023, pp. 499–515. [46] S. Cheng, J.-L. Lin, M. Emani, S. Raskar, S. Foreman, Z. Xie, V. Vishwanath, and M. T. Kandemir, “Thorough Characterization and Analysis of Large Transformer Model Training At-Scale,” Proceedings of the ACM on Measurement and Analysis of Computing Systems, vol. 8, no. 1, pp. 1–25, 2024, ISSN: 2476-1249. [47] Q. Anthony, B. Michalowicz, J. Hatef, L. Xu, M. Abduljabbai, A. Shafi, H. Subramoni, and D. K. Panda, “Demystifying the communication characteristics for distributed transformer models,” in 2024 IEEE Symposium on High-Performance Interconnects (HOTI), IEEE, 2024, pp. 57–65. [48] C. Jin, Z. Jiang, Z. Bai, Z. Zhong, J. Liu, X. Li, N. Zheng, X. Wang, C. Xie, Q. Huang, et al. “MegaScale-MoE: Large-Scale CommunicationEfficient Training of Mixture-of-Experts Models in Production.” arXiv: 2505.11432 [cs], pre-published.
13
[49] M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro. “Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.” arXiv: 1909.08053, pre-published. [50] D. Narayanan, M. Shoeybi, J. Casper, P. LeGresley, M. Patwary, V. Korthikanti, D. Vainbrand, P. Kashinkunti, J. Bernauer, B. Catanzaro, et al., “Efficient large-scale language model training on GPU clusters using megatron-LM,” in International Conference for High Performance Computing, Networking, Storage and Analysis, SC 2021, St. Louis, Missouri, USA, November 14-19, 2021, B. R. de Supinski, M. W. Hall, and T. Gamblin, Eds., ACM, 2021, p. 58. [51] V. Korthikanti, J. Casper, S. Lym, L. McAfee, M. Andersch, M. Shoeybi, and B. Catanzaro, “Reducing activation recomputation in large transformer models,” in Proceedings of Machine Learning and Systems, vol. 5, 2023, pp. 341–353. [52] S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He, “ZeRO: Memory optimizations toward training trillion parameter models,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, SC 2020, Virtual Event / Atlanta, Georgia, USA, November 9-19, 2020, C. Cuicchi, I. Qualters, and W. T. Kramer, Eds., IEEE/ACM, 2020, p. 20. arXiv: 1910.02054 [cs]. [53] A. Ishii and R. Wells, “The Nvlink-Network Switch: Nvidia’s Switch Chip for High Communication-Bandwidth Superpods,” in 2022 IEEE Hot Chips 34 Symposium (HCS), Cupertino, CA, USA: IEEE, 21, 2022, pp. 1–23. [54] E. Ding, C. Ouyang, and R. Singh, Photonic Rails in ML Datacenters, 2025. arXiv: 2507.08119. [55] C. Zhao, C. Deng, C. Ruan, D. Dai, H. Gao, J. Li, L. Zhang, P. Huang, S. Zhou, S. Ma, et al., “Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures,” in Proceedings of the 52nd Annual International Symposium on Computer Architecture, ser. ISCA ’25, New York, NY, USA: Association for Computing Machinery, 2025, pp. 1731–1745. [56] Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M. Dai, Anja Hauth, Katie Millican, et al., Gemini: A Family of Highly Capable Multimodal Models, 2025. arXiv: 2312.11805.