Conceptio › Archive › arXiv CS
arXiv CSopen access

CCL-Bench 1.0: A Trace-Based Benchmark for LLM Infrastructure

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
distributed-systemsinternetnetworkingprotocols
networking, internet, protocols, distributed systems

CCL-Bench 1.0: A Trace-Based Benchmark for LLM Infrastructure

arXiv:2605.06544v1 [cs.DC] 7 May 2026

Eric Ding∗ Byungsoo Oh Bhaskar Kataria Kaiwen Guo Jelena Gvero Abhishek Vijaya Kumar Arjun Devraj Lindsey Bowen Atharv Sonwane Emaad Manzoor Rachee Singh Cornell University

Abstract Evaluative claims about LLM infrastructure —“workload X is fastest on hardware Y with software Z”—depend on a complex configuration space spanning hardware accelerators, interconnect bandwidth, software frameworks, parallelism plans, and communication libraries. Current infrastructure evaluation benchmarks publish a small set of end-to-end numbers that do not explain why one configuration outperforms another. We present CCL-Bench, a trace-based benchmark that addresses the limitations of existing benchmarks by recording reusable evidence for every ML workload. Each contributed data point in CCL-Bench packages an execution trace, a YAML workload card, and the launch scripts. We have developed a community-extensible toolkit to compute fine-grained compute, memory, and communication efficiency metrics from this evidence. Using CCL-Bench, we surface three claims that summary-statistic benchmarks cannot support: (i) higher compute-communication overlap can coincide with longer training step time and reveal inefficient parallelization choices, (ii) doubling TPU interconnect bandwidth yields a much higher end-to-end improvement in step time than doubling GPU interconnect bandwidth on small and medium workloads, and (iii) the best-tuned configuration on one training framework can run up to 3× slower than the best-tuned configuration on a peer framework on identical hardware.

1

Introduction

Several broad audiences rely on LLM infrastructure benchmarks to make critical decisions. Hardware vendors who design accelerators (e.g., GPUs, TPUs [21], Trainium [18], and MTIA [35]) need to ensure that representative workloads perform well on their hardware. Software framework and library developers who build training engines, serving engines, communication libraries, and compilers that run on accelerators (e.g., Megatron-LM [66], vLLM [30], NCCL [48], XLA [62]) need to localize where their software spends time on a given hardware platform. Production operators who deploy LLM training jobs or serving endpoints (e.g., AWS [4], Azure [39]) need to find configurations that meet latency or throughput budgets while minimizing costs. Limitations of Current Benchmarks. Current benchmarks are insufficient for three reasons. First, evaluation metrics are frozen at experiment time. Benchmarks like MLPerf [60, 34], LLM-Perf [24], and other vendor leaderboards publish a small set of end-to-end summary numbers. To understand the same workload along a new evaluation axis, like communication breakdown [36], compute utilization [31], straggler severity [16], or memory-transfer overhead [61], requires another full experiment. Second, current benchmarks provide limited insights for performance enhancement. While current benchmarks support performance comparisons of existing software-hardware stacks [63, 14], they provide limited insights on where improvements should be made to close a given performance gap. Third, tuning effort is invisible. A reported win on public leaderboards often reflects a well-tuned engine against a stock baseline, and the ranking can change once both sides are tuned. ∗ [email protected]

Preprint.

Run Script

Workload card Workload.model: llama-3.1-8b Workload.phase: inference Framework.name: vllm Framework.parallelism.tp: 2 Hardware.xpu.model: a100_gpu …

Trace { “ph”: “x”, “cat”: “kenel”, “name”: “…” “ts”: 2507623347709.573, “dur”: 2.464, “args”: { “External id”: 30, “queued”: 0, … } }, …

CCL-Bench Comm. Metrics

Analysis layer E2E Metrics

XPU Metrics

Run Scripts

Ranking 1. MSCCL++ 2. NCCL

Config optimization

Metric

Memory Metrics

CCLSearch

Deploy Config

Evidence layer Workload cards

SELECT * FROM workload_cards WHERE …model = 'llama-3.1-8b' AND …phase = 'inference' AND framework_name = 'vllm’ … ORDER BY ttft ASC;

Rank 0

What-if analysis

Trace Pool Rank 1

A

B

BW1

BW2

Figure 1: CCL-Bench overview: a standardized trace plus workload card recording each run, a metric toolkit computing fine-grained metrics, and downstream analysis and optimization plug-ins.

Outcome vs. Explanation. The common thread across these limitations is that current benchmarks report only outcomes but not explanations. An outcome tells the reader which infrastructure combination is fastest on a given workload. An explanation tells the reader which component contributed to the win, whether the ranking would survive a different metric or configuration, and where the next bottleneck lies. Without an execution record rich enough to support explanation, every new evaluation metric requires a new experiment, often on hardware the reader does not own. We present CCL-Bench, a trace-based benchmark for LLM infrastructure that records fine-grained execution evidence sufficient to explain observed performance. CCL-Bench has two components. • An evidence-based schema. Every CCL-Bench submission includes three artifacts: an execution trace [56] that records operators, kernels, communication events, timestamps, and per-rank activity; a YAML workload card that specifies the workload (model family, batch size, sequence length, step count) and names the full infrastructure under test (hardware architecture, training or serving engine, parallelism configuration, collective library tuning, compiler); and the run scripts that launched the experiment. Together, these artifacts let a reader attribute measured performance to specific infrastructure components without rerunning the workload. • A community-extensible metric toolkit. CCL-Bench includes an open-source library of analysis tools that are portable across XPU types. A tool is any function that takes a trace and a workload card and returns a scalar metric. The current library covers MFU, compute utilization, compute-communication overlap, memory-transfer overhead, end-to-end step time, and a novel hardware-resource utility metric. New tools can be contributed and applied retroactively. Evaluations enabled by CCL-Bench. CCL-Bench supports two evaluation axes. Software Infrastructure comparisons fix the workload and the hardware architecture, vary one software system component, and match the rest of the execution plan (e.g., NCCL [46] vs. MSCCL++ [38] on the same GPU cluster running Llama-3.1-8B). Hardware Infrastructure comparisons fix the workload and the XPU budget and let each platform run its architecture-native deployment (e.g., PyTorch with NCCL on GPU vs. MaxText [3] with XLA on TPU). CCL-Bench’s protocol of trace collection lets a reader attribute an observed performance gap to compute, memory, or collectives. Beyond these comparisons, CCL-Bench supports three use cases that current benchmarks do not enable. • Post-hoc metric extension. A new evaluation metric can apply retroactively to every compatible trace in the CCL-Bench pool. To demonstrate this, we add a communication traffic-volume metric to explain a performance gap observed between two benchmark entries. • Trace-driven what-if analysis. Empirical traces in CCL-Bench can feed popular distributed ML simulators [76, 73] to estimate how iteration time changes under different interconnect bandwidth, topology, or collective algorithm assumptions. Grounding the simulation in measured kernel timings avoids the calibration gap of purely simulation-based performance models. • Automated configuration optimization. CCL-Bench integrates CCL-Search, an LLM-agentbased optimizer that automates configuration tuning on any hardware–software infrastructure. Given a workload, a target infrastructure, and an optimization objective (e.g., minimize step time, or balance latency against accelerator cost), CCL-Search iteratively proposes configurations, executes them on hardware, collects traces, and evaluates the objective using the metric toolkit to refine its next proposal. The resulting configuration policy and every intermediate trial are recorded as CCL-Bench benchmark entries, so the tuning effort is verifiable and reproducible. Technical claims. Using CCL-Bench, we test technical hypotheses and make the following claims: • Higher compute-communication overlap does not always reduce LLM training step time, because the parallelism choices that increase overlap can also increase collective traffic volume. 2

• Doubling TPU interconnect bandwidth yields up to 100× higher end-to-end step-time improvement than doubling GPU interconnect bandwidth on small and medium workloads, while large GPU training workloads benefit more from a larger scale-up domain. • A training framework with its best-found parallelism configuration can run up to 3× slower on a peer framework on identical hardware with the same workload. We make CCL-Bench’s workload cards and metric toolkits publicly available at https://github. com/cornell-sysphotonics/ccl-bench. CCL-Bench is hosted at https://cclbench.ai/.

2

Benchmarking Infrastructure for Large Language Models

Infrastructure

CCL-Bench evaluates the infrastructure that executes LLM workloads. This section defines how the software, hardware and LLM workload compose (Figure 2). Workload specifies the computation task to be perWorkload formed, i.e., the model family and size, quantization, Model family, phase, batch size, sequence length the phase (training or inference), the batch size, and Software the sequence length. CCL-Bench targets open-weight Compiler, communication library, framework models that can be hosted by a third party for evaluation on their infrastructure. The workload determines Hardware Platform the arithmetic computation of an experimental run. Accelerator, memory, interconnect, topology Software. The software maps a workload onto a hardware platform. It comprises, from lowest to high- Figure 2: LLM infrastructure evaluated by CCLBench. The workload defines what computation est level of (1) compiler support for graph capture, must happen. The infrastructure—hardware platoperator fusion, and kernel generation (e.g., TorchIn- form and software—determines how it is executed. ductor [7] and Triton [68]), (2) the communication library (e.g., NCCL, MSCCL++, XLA collectives [52]) and (3) the training or serving framework together with its parallelism plan (e.g., TorchTitan [33], Megatron-LM, vLLM, SGLang, or MaxText). The framework is configured with Data Parallelism (DP), Tensor Parallelism (TP), Pipeline Parallelism (PP), Expert Parallelism (EP) degrees, micro-batch size, pipeline schedule, and memory-management policies like activation recomputation, ZeRO stage, offload, and cudagraph [41, 61, 58, 59, 55]. Hardware platform is the physical environment on which the workload executes. This includes the accelerator type and model (GPU, TPU, or other XPU), the memory hierarchy, the server-scale interconnect (e.g., NVLink [51], UALink [69]), and the datacenter-level network topology (e.g., Torus [28], Dragonfly [29], Rail-optimized [57]). Infrastructure. Infrastructure is the combination of hardware platform and software. It is the object of study in CCL-Bench: every evaluative claim names a workload and a piece of infrastructure. 2.1

Existing Benchmarks

Machine learning benchmarks such as LLM-Perf and MLPerf [60, 34, 63, 14, 27, 8, 24, 15, 78, 80, 1] run full models and report scalar summaries that are used for performance rankings. They are useful for understanding current infrastructure’s performance via specific metric suites (e.g., query latency in InferenceX [63], energy consumption in ML.Energy [14]). Parameterized studies [74] broaden the metric or workload surface but remain coupled to the metrics selected before execution. Infrastructure providers have published component-level benchmark artifacts with new releases. Hardware vendors report product-specific results with limited metrics [19, 20, 44, 45, 47, 5, 6, 43]. Communication libraries focus on collective microbenchmarks and tuning studies [49, 46, 25, 38]. Serving engines including vLLM and SGLang release throughput and latency benchmarks scoped to individual releases [71, 72, 64, 65]. These artifacts expose useful LLM-specific metrics (TTFT, TPOT, collective bandwidth) but compare against selected baselines on a single platform. Their logging, metric calculation, and tuning records are not standardized across contributors. Automated tuning strategies for parallelism and communication [79, 37, 26] address part of the reproducibility gap but require significant re-engineering to port across software engines or hardware platforms. Trace-based systems. One family of trace-based systems treats the trace as the workload itself, providing compact, replayable objects for co-design [40], serving request streams [75], or scheduling and provisioning studies [67, 54]. A second family uses traces as evidence for analyzing the original execution. Holistic Trace Analysis (HTA) [36] exemplifies this approach, computing kernel 3

breakdowns, communication-computation overlap, and trace diffs over Kineto [56] traces. CCLBench follows this second direction and extends it into a benchmarking framework with standardized workloads, comparison axes, and a shared trace pool for cross-infrastructure metric analysis. Simulators and performance models. Simulators for model architecture, network topology, and hardware [76, 73, 10], along with tools for training/inference configuration planning and deployment [2, 12, 77, 9, 79], enable design-space exploration without running every configuration on physical hardware. Their outputs complement but do not substitute for evidence grounded in observed executions, and they depend on empirical calibration that trace-based benchmarks can provide. 2.2

Limitations of Existing Benchmarks

Existing LLM infrastructure benchmarks face three challenges that limit their diagnostic value and leave them functioning only as scoreboards. Benchmark artifacts are metric-specific. Suites like MLPerf [60], LLM-Perf [24], and vendor leaderboards answer only the metrics chosen by the benchmark owner, but not metrics introduced later by a library developer, hardware designer, or operator. These suites are scoreboards that define what is measured, but they do not expose a reusable artifact from which new metrics or bottleneck attributions can be derived post-hoc. Follow-up questions about communication breakdown [36], compute utilization [31], or memory-transfer overhead [61] require rerunning the workload and laborious tuning because the published result does not contain the evidence needed to compute them. Cross-referencing different benchmarks is also not possible due to the large configuration space. Rankings do not identify the bottleneck. End-to-end benchmark scores are useful for comparisons, but they provide little guidance for improving one stack or investing in the next hardware revision. After observing a performance gap, a user needs to know whether the limiting factor is collective communication, memory bandwidth, placement of parallelism domains, or saturation of a particular network fabric. Scalar benchmark reports discard the per-rank timeline [34, 60, 63] and configuration context needed to make that attribution. Configuration search is outside the benchmark record. LLM performance is a result of environment-specific choices like TP, DP, PP, micro-batch size, collective-library settings, compiler options, and memory policies [79, 46]. Existing benchmarks often report the final configuration, but not the search path, failed trials, or objective used to choose it. This makes a published win difficult to reproduce and makes it hard to judge whether two systems were compared at similarly optimized operating points. Requiring every participant to perform this tuning manually also raises the cost of contributing new benchmark entries.

3

Table 1: Workload-card fields (subset).

CCL-Bench overview Role

CCL-Bench (Figure 1) records evidence, instead of final evaluation metrics, from LLM workloads. Each experiment contributes to CCL-Bench an execution trace, a workload card describing the workload and infrastructure under test, and the scripts that launched it. Anyone can later compute new metrics on the same evidence without rerunning the experiment, and submit those metrics as separately versioned tools to our open-source repository. 3.1

What CCL-Bench records

Template field

workload.model. Workload model_family Workload workload.model.phase Workload workload.data.batch_size Workload workload.data.seq_len workload.hardware. Architecture xpu_spec.model workload.hardware. Architecture network_topo.topology Model-executor. System framework.name Model-executor. System model_plan_parallelization Model-executor. System communication_library.name Model-executor. System protocol_selection

Example llama-3.1-8b training 4 8192 nvidia_a100 slingshot torchtitan DP shard=2, TP=4, PP=2 NCCL rocev2, p2p

Trace. CCL-Bench collects LLM execution traces in JSON format, including Kineto traces [56] and JAX/XLA profiler traces collected via XProf [53]. Kineto, the tracing backend used by the PyTorch Profiler, supports NVIDIA and AMD GPUs, Intel XPU, and HPU backends [56], while the JAX Profiler/XProf provides trace collection for JAX workloads on CPU, GPU, and TPU through XLA runtime events [53]. For each iteration on every rank, these traces record the operators the model executes, the kernels and communication events each operator launches across one or more streams, and the timestamp of every event, supporting both end-to-end performance metrics and fine-grained trace-level breakdowns. Auxiliary traces, like 4

Nsight Systems traces [50], further provide hardware-specific counters not exposed by Kineto. We provide details about the trace format and collection overhead in Appendix A. Workload card. The YAML card records the workload, the architecture, and the software stack so that the trace becomes interpretable evidence. Provenance fields (version, description, hf_url, trace_url, contributor) identify the experiment and its author. Workload fields record phase, MoE (Mixture-of-expert model [17]) flag, model family, floating point precision, iteration count recorded, parameter and layer counts, batch size, sequence length, and dataset. Architecture fields record XPU model, device counts, driver version, and network topology and bandwidth. System fields record the framework, compiler, parallelism plan (DP/TP/PP/CP/EP and pipeline microbatch), and communication library and protocol. Metric source fields list the trace types and any auxiliary artifacts the metric tools need. Table 1 shows representative fields; the full schema is in Appendix B. Run scripts are required in a submission for reproducibility. We discuss how contributors collect traces to capture accurate infrastructure performance for different workload types in Appendix C. 3.2

How CCL-Bench analyzes evidence

A tool i in CCL-Bench is a function from a workload card and a trace to a scalar metric fi : W × T −→ R, where W is the space of workload cards and T is the space of execution traces. A suite of n tools applied tothe same evidence yields a performance profile m(w, τ ) = f1 (w, τ ), f2 (w, τ ), . . . , fn (w, τ ) ∈ Rn , where (w, τ ) ∈ W × T is a single submission. No single scalar captures infrastructure behavior; the vector m jointly characterizes compute efficiency, memory throughput, and communication overhead. Toolkit. The shipping toolkit covers four categories. Efficiency reports average step time and Model FLOPs utilization (MFU) [13]. Compute reports compute unit coverage and primary kernel timespan. Memory reports host-device bandwidth and memory-transfer overhead. Communication reports collective bandwidth, communication fraction (amount of time spent in communication), and compute-communication overlap. The toolkit also reports MoE fraction for mixture-of-experts models and TTFT/TPOT for inference. Appendix D catalogs the full set with details for MFU, memory-transfer overhead, and compute-communication overlap. 3.3

How users interact and contribute

Evidence interface. A contributor runs a workload from CCL-Bench’s rolling workload set on the infrastructure under test, collects traces, and uploads them with a workload card and run scripts via the web UI. Rule-based checks verify metadata correctness—minimum iteration count, complete environment reporting, and schema validity. CCL-Bench then applies the current tool suite and publishes the entry once the submission and metric results pass review. Tool interface. A contributor submits a tool along with its category, supported trace types, and expected output. Maintainers check compatibility before admission. Once admitted, CCL-Bench applies the tool to every compatible trace already in the pool, so a new metric immediately covers all existing workloads without rerunning any experiment. If a tool is found to be incorrect, CCL-Bench retracts the version, flags all affected entries, and recomputes those metrics against the corrected tool. Contributors can also challenge an existing entry with a counter-trace or a new tool. The dispute is resolved on the public submission thread (e.g., a GitHub pull request). Result access. The leaderboard is a two-dimensional table—submissions as rows, metrics as columns. Entries can be clustered by workload—with selectable rows for cross-infrastructure comparison. A website screenshot is attached (Figure 8 in Appendix E). Workload cards, run scripts, and the tool suite are publicly available in the CCL-Bench GitHub repository. Access to raw traces is gated on contributing a trace for the same workload, incentivizing pool growth.

4

Case studies and insights from CCL-Bench

We use CCL-Bench to evaluate LLM infrastructure across two hardware platforms and multiple software configurations. The first is the NERSC Perlmutter supercomputer [42], where each node contains four A100 GPUs connected by NVLink 3.0 at 300 GB/s unidirectional bandwidth (scale-up domain), and nodes connect through a Slingshot-11 fabric at 200 Gbps (scale-out domain). The second is Google TPU v6e [22], where eight chips form a node and 32 nodes form a pod (scale-up domain) in a 2D torus topology with 100 GB/s unidirectional inter-chip interconnect (ICI) bandwidth. Software 5

WL1 Qwen 4B infer WL2 Llama 8B infer

WL3 DSK 16B infer WL4 Llama 8B train

WL5 DSK 16B train

Step time (s)

101

EP=8 EP=4 0.0

100 10 1 10 20 MFU (%)

0 25 50 75 Comp-comm overlap (%) (a)

EP=8 EP=4

0

Compute Exposed comm 3.67s

Overlapped comm Idle

11.98s 5.0 7.5 10.0 12.5 Step time (s) AllReduce ReduceScatter AllGather AllToAll 7.9 GB 29.1 GB 10 20 30 Collective traffic (GB) 2.5

(b)

Figure 3: (a) MFU vs. step time and compute-comm overlap vs. step time on A100 GPU (Perlmutter). Each point is one CCL-Bench entry. (b) Step-time and communication traffic-volume break-down for WL5 DeepSeekV3-16B EP=4 and EP=8 MoE training runs (TP=4, DP=2, PP=1).

infrastructure spans communication libraries (NCCL, MSCCL++, XLA collectives) and training and serving engines (Megatron-LM, TorchTitan, MaxText, vLLM, SGLang). Table 4 in Appendix F lists the workload suite we use to put the infrastructure combinations under test. We show how CCL-Bench enables making technical claims that cannot be derived from coarser LLM infrastructure benchmarks. Cross-system and cross-architecture comparisons can be found in Appendix G. 4.1

Perspective of a framework and library developer

A framework developer tuning a distributed LLM workload typically optimizes for two metrics: MFU (model FLOPs utilization) and compute-communication overlap. Both metrics are commonly used as proxies for lower step time under a fixed workload and hardware configuration: higher MFU indicates that more of the hardware’s peak compute is being converted into model FLOPs, while higher compute-communication overlap reduces exposed communication on the critical path [13, 23]. CCL-Bench’s trace pool lets us test whether these assumptions hold across configurations. Figure 3(a, left) confirms the first assumption. Across training and inference workloads in CCLBench, higher MFU generally yields lower step time. Inference workloads on vLLM and SGLang reach higher MFU than training workloads on TorchTitan, which require further configuration tuning to achieve comparable utilization. The second assumption, however, does not hold. Figure 3(a, right) shows that for DeepSeek-V3-16B MoE training (WL5), higher compute-communication overlap coincides with worse step time. A traditional summary-statistic benchmark would report only the step time, but CCL-Bench’s execution traces help us diagnose the reason for this counterintuitive finding. We develop new metric tools for the diagnosis. First, we attribute the step time to compute vs. exposed communication vs. overlapped communication. Second, we capture the traffic volume of different collectives in each step. We apply these tools to the existing traces without rerunning any experiment. We find that experiments with a higher degree of compute–communication overlap had used a small expert-parallelism (EP) degree relative to the total GPU count (e.g., EP = 4 on 8 GPUs), and experiments with lower overlap had used a larger EP degree (e.g., EP = 8). While compute–communication breakdown (Figure 3(b)) shows that communication dominates the step time in both configurations, the traffic-volume breakdown shows that workloads with smaller EP domains generate significantly more R EDUCE S CATTER and A LL G ATHER traffic. This increase arises because a smaller EP domain replicates experts across the data-parallelism domain, requiring fully-sharded data parallelism to gather weights and synchronize gradients. Although these collectives can be partly overlapped with computation, the additional communication volume outweighs the overlap benefit and increases overall step time. This finding can motivate developers to explore expert and data parallelism co-optimization opportunities, balancing expert A LLT OA LL communication with data parallelism collectives. The workload trace collection and post-hoc metric extension in CCL-Bench enables this finding. Claim 1: Higher computation-communication overlap does not always imply lower LLM training step latency; it may instead reveal inefficient parallelization choices.

6

4.2

Perspective of a hardware vendor

A hardware designer evaluating whether to invest in faster interconnects between accelerators needs to estimate how much end-to-end performance a bandwidth upgrade would buy for a given workload. Traditionally, computer architects use simulators to answer such questions [40, 76, 73]. However, these simulators rely entirely on analytical models and lack grounding in real execution behavior. CCL-Bench bridges this gap by combining empirical traces with simulation frameworks. From traces to simulation. CCL-Bench traces record per-kernel timing, collective types, group IDs, and buffer sizes for every rank. We convert each trace into a Chakra [40] execution graph that captures the dependencies between computation and communication across ranks. Compute kernels become COMP_NODEs with measured durations; collectives become COMM_COLL_NODEs parameterized by type, message size, and communicator group; and idle intervals become gap nodes that preserve the original timeline. Edges enforce CUDA/XLA stream order and rank-local sequencing (Figure 4). We feed these execution graphs to Astra-Sim [76], a distributed ML CCL-Bench 4 Calculate utility metrics workload simulator that models netUtility Metrics Rank 0 work topologies with configurable Simulation gen_gpu_chakra_et( ) bandwidth, latency, and collective alRank 1 Trace pool gorithms. Using Astra-Sim, we replay gen_tpu_chakra_et( ) C A B Comp Comm Workload cards the computation nodes at their empirinode node … BW cal durations and re-simulate the comsimulation and Select target workload 2 Convert to Chakra ET trace format 3 Performance 1 what-if analysis munication nodes under a target network configuration, producing a new Figure 4: CCL-Bench trace-driven what-if pipeline: empirical end-to-end step time estimate. This traces feed a network simulator that estimates step time and the pipeline preserves the measured com- utility of doubling a hardware resource. pute behavior of the original run while letting the network parameters vary. We build conversion pipelines for both GPU (Kineto) and TPU (XLA) traces. We use this pipeline to study how improvements in interconnect bandwidth impact iteration time for workloads run on GPUs vs. TPUs. Appendix H includes additional results. Utility metric. For a hardware resource r (e.g., scale-out network bandwidth, scale-up network bandwidth, or the size of scale-up domain—where all XPUs are interconnected by high-bandwidth scale-up network), let T be the baseline step time and T2× (r) be the simulated step time after doubling r while holding the workload and all other infrastructure parameters fixed (Figure 4). We define the utility of resource r as the resulting fractional step-time improvement: T − T2× (r) Utility(r) = × 100%. (1) T A small value indicates that the workload gets little benefit from upgrading r, while a large value indicates a bottleneck where the doubling buys meaningful end-to-end speedup. The metric answers an operator-facing upgrade question. Given the workloads already running on the current cluster, which infrastructure upgrade is most worth the spend? GPU vs. TPU bandwidth utility. Figure 5 reports utility for representative workloads on Perlmutter A100 (GPU, three axes: scale-out bandwidth, scale-up bandwidth, scale-up domain size) and TPUv6e (TPU, ICI bandwidth). In GPU-based cluster architectures, the best upgrade is workload-dependent. Doubling scale-out bandwidth helps communication-heavy multi-node training (up to 28.7%) but does not improve single-node inference. Doubling the scale-up domain size provides the largest gains for WL4 and WL5 (53.9% and 46.7%) by shifting collective traffic from the slower scale-out fabric onto the faster scale-up fabric. Larger training workloads (WL6 at 128 GPUs, WL7 at 256 GPUs) show low utility (3.9% and 3.4%) because their step time is compute-bound. TPU workloads, by contrast, draw high utility from ICI bandwidth doubling for both training and inference, up to 102.57× the matched GPU scale-up bandwidth utility for inference and up to 22.82× for training, due to TPU’s lower baseline bandwidth (100GBps ICI vs. 300GBps NVLink) and lower network connectivity (torus vs. fully-connected topology). The contrast indicates that TPU ICI is a more useful investment per bandwidth doubling on small and medium workloads, while GPU clusters benefit more from scale-up domain expansion at model-training scale. 7

Utility (%)

GPU Scale-out BW

GPU Scale-up BW

GPU Scale-up domain

TPU ICI BW

101 100 10 1 WL1 Qwen 4B infer WL2 Llama 8B infer WL3 DSK 16B infer WL4 Llama 8B train WL5 DSK 16B train WL6 DSK 16B train WL7 DSK 236B train

Figure 5: Utility metrics for scale-out bandwidth, scale-up/ICI bandwidth, and scale-up domain size on Perlmutter A100 Slingshot and TPUv6e Torus. We are not able to run WL6,7 on TPUs due to resource limits.

Claim 2: Doubling TPU ICI bandwidth yields higher end-to-end utility than doubling GPU scale-up bandwidth for small workloads (up to 100× better for inference and 22× for training). For medium-to-large GPU training workloads, investing in a larger scale-up domain is the more useful upgrade than improving the bandwidth. 4.3

Perspective of a production operator

A production operator deploying an LLM training job must choose parallelism degrees, micro-batch sizes, activation checkpointing, etc., before launching a run. This deployment configuration space is large, interdependent, and framework-specific (i.e., a setting that performs well on one training engine may perform poorly on another). In practice, operators rely on heuristic guidelines and manual iteration, and the tuning effort behind a published benchmark result is rarely recorded or reproducible. We build an agentic configuration optimizer, CCL-Search, that leverages CCL-Bench’s evidence and analysis layers (e.g., traces, metrics, and prior runs) to progressively discover high-quality configurations through guided experimentation, while recording the entire search process as reproducible benchmark artifacts. How CCL-Search works. execute() generate_config() Trace CCL-Bench Config. CCL-Search is an LLM-agentAnalysis layer CCL-Search compute_metrics() … based optimizer that iteratively XPU Metrics E2E Metrics Policy Metrics explores the configuration space update_history() Evidence layer for a distributed LLM workload Workload Run Trace (Figure 6). The user specifies an update_policy() cards Scripts Pool optimization objective as one or a combination of CCL-Bench Figure 6: CCL-Search loop. update_policy is an LLM call. Other metrics (e.g., training step steps are local/tool execution. time, inference throughput, or a composite objective) and defines the configuration knobs the agent may vary. Like the ADRS framework [11], CCL-Search maintains a policy program, generate_config, that maps workload and infrastructure context to a concrete configuration (e.g., DP, TP, PP, EP degrees, micro-batch size, activation checkpointing). The policy is represented as a programmatic template that is iteratively rewritten by the LLM. Each iteration proceeds in four steps: (1) generate_config proposes a configuration, (2) the configuration is executed on the target hardware, (3) CCL-Bench collects a trace and computes the objective score via compute_metric, and (4) update_policy invokes an LLM with the score and the full execution history to revise generate_config for the next round. Unlike black-box optimizers, CCL-Search enables the LLM to reason over structured configuration semantics (e.g., TP/DP tradeoffs, pipeline depth) and trace-level execution feedback, allowing it to avoid invalid, resource-constrained, or inefficient configurations and adapt its strategy based on observed execution behavior. Moreover, every explored configuration produces a complete CCLBench benchmark entry—workload card, trace, and score—so the entire search history is preserved for analysis, comparison, and public submission. Framework-specific tuning optima. Figure 7(a) shows CCL-Search benchmarking two training engines, TorchTitan and Megatron-LM, on Llama-3.1-8B (WL4, 16 Perlmutter GPUs) over 15 iterations across five knobs (TP, DP, PP, micro-batch size, activation checkpointing). The key finding is that each framework’s optimum lies at a different point in configuration space. TorchTitan reaches its best step time of 1.50 s at TP = 1, DP = 4, PP = 4, while Megatron-LM reaches 0.44 s at TP = 4, DP = 1, PP = 4—a 3.4× gap. Applying TorchTitan’s best configuration to Megatron-LM yields 1.3 s, which is 15% faster than TorchTitan but still 3× slower than Megatron-LM’s own optimum. Thus, def policy (workload, environment) -> workload_config: tp = environment.hardware.count_per_node dp = 1 pp = 1 micro_batch = workload.batch_size / pp # Further tuning return (tp, dp, pp, micro_batch)

8

TorchTitan: Step time

15 10 5 0

1.50 s 0

2

4

6

8

Iteration

10

12

14

15 10 5 0

0.44 s 0

2

4

6

8

Iteration

10

12

14

Config choices

Config choices

0 1 2 3 4 5 6 7 8 9 10 11 12 13 14

0 1 2 3 4 5 6 7 8 9 10 11 12 13 14

TP 4 2 4 4 4 4 4 2 4 2 8 2 2 2 1 TP 4 4 2 2 2 2 2 2 1 2 2 1 4 8 4 DP 4 8 4 4 2 1 2 4 2 4 1 2 2 2 4 DP 4 4 8 8 4 4 2 2 4 2 1 2 1 1 1 PP 1 1 1 1 2 4 2 2 2 2 2 4 4 4 4 PP 1 1 1 1 2 2 4 4 4 4 8 8 4 2 4 MBS 2 4 4 1 4 4 2 1 4 4 4 4 2 1 2 MBS 2 4 4 2 2 4 2 4 2 1 1 1 1 1 2 Act Ckpt F F F F F F F F T T T T T T T Act Ckpt F F T F F F F F F F F F F F F

Iteration

Megatron-LM: GPU Eff. + Step Time

Trial Running best

20

Step time (s)

Step time (s)

20

Megatron-LM: Step time

25

Trial Running best

Step time (s)

25

Iteration

(a)

2.75 2.50 2.25 2.00 iter 7 1.75 1.50 1.25 11

TP DP PP EP MBS Act Ckpt

4 4 1 4 2 T 0

iter 0

iter 3 iter 8 1.69

12

13

1.50

14

15

16

Num GPUs (TP × DP × PP)

2 2 1 2 2 T 1

4 2 1 2 2 T 2

Config choices 2 2 3 2 1 T 3

2 1 3 1 1 T 4

1 2 3 2 1 T 5

2 4 1 4 1 T 6

Iteration

2 2 3 2 1 F 7

2 2 3 2 2 F 8

17 2 1 1 2 3 3 1 2 2 2 T T 9 10

(b)

Figure 7: CCL-Search results. Bottom panels show explored TP, DP, PP, micro-batch size, and activationcheckpointing choices. Grayed-out boxes indicate failed runs. (a) Step-time objective: CCL-Search reduces step time by up to 8× on TorchTitan and 19× on Megatron-LM within 15 iterations using Score = −avg_step_time. (b) Composite objective: CCL-Search finds the Pareto-frontier using Score = w × T0 /T + (1 − w) × N0 /N on Megatron-LM. T is average step time, N = num. XPUs, and w = 0.5.

configuration policies do not transfer across frameworks. An operator who tunes on one engine and assumes the same settings will work on another leaves significant performance on the table. Composite objectives. Figure 7(b) shows CCL-Search optimizing for a composite objective, Score = w × T0 /T + (1 − w) × N0 /N , which balances step time (T ) against accelerator count (N ) relative to a baseline (T0 , N0 ), with w = 0.5. This formulation rewards configurations that simultaneously reduce latency and hardware cost. On DeepSeek-V3-16B (WL5), the agent identifies Pareto-optimal configurations that trade off latency against hardware cost, enabling cost-aware infrastructure comparisons. Additional runs in Appendix I confirm CCL-Search’s stability across repeated trials. Since CCL-Search records every trial as a CCL-Bench benchmark entry, the tuning process becomes reproducible. An operator on a different cluster can rerun the same policy, inspect the resulting traces, and verify whether the published optimum transfers to their environment. Claim 3: Deploying the same workload on different training engines requires framework-specific tuning. A configuration that is optimal for one framework can be up to 3× slower than the best-found configuration on another.

5

Discussion and Limitations

Coverage. CCL-Bench 1.0 covers small-to-medium open-source models on GPU and TPU for training and batch inference. Future versions will add larger models, additional accelerators (e.g., Trainium), online serving with request-level latency, and accuracy-affecting optimizations (quantization, sparsity, speculative decoding). Exhaustive coverage of the workload–infrastructure cross product is impossible by construction, so CCL-Bench adopts a rolling, theme-driven submission model (e.g., communication backends, interconnect fabrics, MoE collectives) on top of a workloadand infrastructure-agnostic trace protocol that lets the suite extend without schema changes. Tools and plug-ins. The same trace artifact supports post-hoc metric extension and downstream simulation-based what-if analysis. Building new trace-facing simulator plug-ins, or wiring CCLBench into existing performance models, is a natural next step [32]. Trace storage at scale. A single multi-rank trace can reach 10 GB, and 500+ accelerator runs approach 100 GB. Scaling the leaderboard will require a submission-selection policy that prioritizes diversity across models, platforms, and parallelism, together with compressed or sampled trace formats that preserve the fields the metric tools consume. Additional discussion, including handling of proprietary information in traces, is in Appendix J. 9

6

Conclusion

CCL-Bench is a first step toward trace-based LLM infrastructure evaluation. Originating from a university class project, CCL-Bench has collected over 100 workloads, spanning over 7 model architectures, 7 frameworks, and 3 hardware environments. The benchmarking platform features a community-extensible metric toolkit, a trace-replay pathway to hardware simulators for what-if analysis, and CCL-Search, a benchmarking protocol for automating infrastructure configuration optimization. We provide case studies that demonstrate CCL-Bench surfaces findings that coarser, headline-number benchmarks cannot. We invite the community to contribute matched-workload traces and analysis tools to broaden cross-system and cross-architecture coverage.

References [1] Robert Adolf, Saketh Rama, Brandon Reagen, Gu-Yeon Wei, and David Brooks. Fathom: Reference workloads for modern deep learning methods. In 2016 IEEE International Symposium on Workload Characterization (IISWC), pages 1–10. IEEE, 2016. [2] Amey Agrawal, Nitin Kedia, Jayashree Mohan, Ashish Panwar, Nipun Kwatra, Bhargav S Gulavani, Ramachandran Ramjee, and Alexey Tumanov. Vidur: A large-scale simulation framework for llm inference. Proceedings of Machine Learning and Systems, 6:351–366, 2024. [3] AI Hypercomputer. MaxText: A Simple, Performant and Scalable JAX LLM. https://github.com/ AI-Hypercomputer/maxtext, 2026. Accessed: 2026-05-05. [4] Amazon Web Services. Amazon web services (aws). https://aws.amazon.com/, 2026. Accessed: 2026-05-06. [5] AMD. Accelerating generative AI: How AMD Instinct GPUs delivered breakthrough efficiency and scalability in MLPerf Inference v5.1. https://www.amd.com/en/blogs/2025/ accelerating-generative-ai-how-instinct-gpus-delivered.html, 2025. Accessed 2026-0426. [6] AMD. Accelerating AI training: How AMD Instinct MI350 Series GPUs delivered breakthrough performance and efficiency in MLPerf Training v5.1. https://www.amd.com/en/blogs/2025/ accelerating-ai-training.html, 2025. Accessed 2026-04-26. [7] Jason Ansel, Edward Yang, Horace He, Natalia Gimelshein, Animesh Jain, Michael Voznesensky, Bin Bao, Peter Bell, David Berard, Evgeni Burovski, Geeta Chauhan, Anjali Chourdia, Will Constable, Alban Desmaison, Zachary DeVito, Elias Ellison, Will Feng, Jiong Gong, Michael Gschwind, Brian Hirsh, Sherlock Huang, Kshiteej Kalambarkar, Laurent Kirsch, Michael Lazos, Mario Lezcano, Yanbo Liang, Jason Liang, Yinghai Lu, C. K. Luk, Bert Maher, Yunjie Pan, Christian Puhrsch, Matthias Reso, Mark Saroufim, Marcos Yukio Siraichi, Helen Suk, Michael Suo, Phil Tillet, Eikan Wang, Xiaodong Wang, William Wen, Shunting Zhang, Xu Zhao, Keren Zhou, Richard Zou, Ajit Mathews, Gregory Chanan, Peng Wu, and Soumith Chintala. PyTorch 2: Faster machine learning through dynamic python bytecode transformation and graph compilation. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2, pages 929–947, 2024. doi: 10.1145/3620665.3640366. LLM Leaderboard: Comparison of over 100 AI Models. [8] Artificial Analysis. artificialanalysis.ai/leaderboards/models, 2026. Accessed: 2026-05-05.

https://

[9] Jehyeon Bang, Yujeong Choi, Myeongwoo Kim, Yongdeok Kim, and Minsoo Rhu. vTrain: A simulation framework for evaluating cost-effective and compute-optimal large language model training. In 2024 57th IEEE/ACM International Symposium on Microarchitecture (MICRO), pages 153–167. IEEE, 2024. doi: 10.1109/MICRO61859.2024.00021. [10] Archit Bansal, Danny Stoll, Maciej Janowski, Arber Zela, and Frank Hutter. Jahs-bench-201: A foundation for research on joint architecture and hyperparameter search. Advances in Neural Information Processing Systems, 35:38788–38802, 2022. [11] Audrey Cheng, Shu Liu, Melissa Pan, Zhifei Li, Bowen Wang, Alex Krentsel, Tian Xia, Mert Cemri, Jongseok Park, Shuo Yang, et al. Barbarians at the gate: How ai is upending systems research. arXiv preprint arXiv:2510.06189, 2025. [12] Jaehong Cho, Minsu Kim, Hyunmin Choi, Guseul Heo, and Jongse Park. Llmservingsim: A hw/sw co-simulation infrastructure for llm inference serving at scale. In 2024 IEEE International Symposium on Workload Characterization (IISWC), pages 15–29. IEEE, 2024.

10

[13] Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. Journal of machine learning research, 24(240):1–113, 2023. [14] Jae-Won Chung, Jeff J Ma, Ruofan Wu, Jiachen Liu, Oh Jun Kweon, Yuxuan Xia, Zhiyu Wu, and Mosharaf Chowdhury. The ml. energy benchmark: Toward automated inference energy measurement and optimization. arXiv preprint arXiv:2505.06371, 2025. [15] Cody Coleman, Deepak Narayanan, Daniel Kang, Tian Zhao, Jian Zhang, Luigi Nardi, Peter Bailis, Kunle Olukotun, Christopher Ré, and Matei Zaharia. DAWNBench: An end-to-end deep learning benchmark and competition. In NIPS ML Systems Workshop, 2017. [16] Arjun Devraj, Eric Ding, Abhishek Vijaya Kumar, Robert Kleinberg, and Rachee Singh. Efficient allreduce with stragglers, 2025. URL https://arxiv.org/abs/2505.23523. [17] William Fedus, Barret Zoph, and Noam Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research, 23(120):1–39, 2022. [18] Xinwei Fu, Zhen Zhang, Haozheng Fan, Guangtai Huang, Mohammad El-Shabani, Randy Huang, Rahul Solanki, Fei Wu, Ron Diamant, and Yida Wang. Distributed training of large language models on aws trainium. In Proceedings of the 2024 ACM Symposium on Cloud Computing, pages 961–976, 2024. [19] Google Cloud. Introducing Cloud TPU v5p and AI Hypercomputer. https://cloud.google.com/blog/products/ai-machine-learning/ introducing-cloud-tpu-v5p-and-ai-hypercomputer, 2023. Accessed 2026-04-26. From LLMs to image generation: Accelerate inference workloads [20] Google Cloud. with AI Hypercomputer. https://cloud.google.com/blog/products/compute/ ai-hypercomputer-inference-updates-for-google-cloud-tpu-and-gpu, 2025. Accessed 2026-04-26. [21] Google Cloud. Tensor processing units (tpus). https://cloud.google.com/tpu?hl=en, 2026. Accessed: 2026-05-06. [22] Google Cloud. Cloud TPU v6e. https://cloud.google.com/tpu/docs/v6e, 2026. Accessed: 202605-06. [23] Sayed Hadi Hashemi, Sangeetha Abdu Jyothi, and Roy H. Campbell. TicTac: Accelerating distributed deep learning with communication scheduling. arXiv preprint arXiv:1803.03288, 2018. URL https: //arxiv.org/abs/1803.03288. [24] Hugging Face Optimum Team. LLM-Perf Leaderboard. https://huggingface.co/spaces/optimum/ llm-perf-leaderboard, 2024. Accessed: 2026-05-05. [25] Changho Hwang, Peng Cheng, Roshan Dathathri, Abhinav Jangda, Saeed Maleki, Madan Musuvathi, Olli Saarikivi, Aashaka Shah, Ziyue Yang, Binyang Li, Caio Rocha, Qinghua Zhou, Mahdieh Ghazimirsaeed, Sreevatsa Anantharamu, and Jithin Jose. MSCCL++: Rethinking GPU communication abstractions for AI inference. In ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS), 2026. [26] Zhihao Jia, Matei Zaharia, and Alex Aiken. Beyond data and model parallelism for deep neural networks. In MLSys, 2019. [27] Yinsicheng Jiang, Yao Fu, Yeqi Huang, Ping Nie, Zhan Lu, Leyang Xue, Congjie He, Man-Kit Sit, Jilong Xue, Li Dong, et al. Moe-cap: Benchmarking cost, accuracy and performance of sparse mixture-of-experts systems. arXiv preprint arXiv:2412.07067, 2024. [28] Norman P. Jouppi, George Kurian, Sheng Li, Peter Ma, Rahul Nagarajan, Lifeng Nai, Nishant Patil, Suvinay Subramanian, Andy Swing, Brian Towles, Cliff Young, Xiang Zhou, Zongwei Zhou, and David Patterson. TPU v4: An optically reconfigurable supercomputer for machine learning with hardware support for embeddings. In Proceedings of the 50th Annual International Symposium on Computer Architecture, 2023. doi: 10.1145/3579371.3589350. [29] John Kim, Wiliam J Dally, Steve Scott, and Dennis Abts. Technology-driven, highly-scalable dragonfly topology. ACM SIGARCH Computer Architecture News, 36(3):77–88, 2008. [30] Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th symposium on operating systems principles, pages 611–626, 2023.

11

[31] Cheng Li, Abdul Dakkak, Jinjun Xiong, Wei Wei, Lingjie Xu, and Wen-Mei Hwu. XSP: Acrossstack profiling and analysis of machine learning models on GPUs. In Proceedings of the 34th IEEE International Parallel and Distributed Processing Symposium (IPDPS), pages 326–327. IEEE, 2020. doi: 10.1109/IPDPS47924.2020.00042. [32] Mingyu Liang, Hiwot T Kassa, Wenyin Fu, Brian Coutinho, Louis Feng, and Christina Delimitrou. Lumos: Efficient performance modeling and estimation for large-scale llm training. Proceedings of Machine Learning and Systems, 7, 2025. [33] Wanchao Liang, Tianyu Liu, Less Wright, Will Constable, Andrew Gu, Chien-Chin Huang, Iris Zhang, Wei Feng, Howard Huang, Junjie Wang, et al. Torchtitan: One-stop pytorch native solution for production ready llm pre-training. arXiv preprint arXiv:2410.06511, 2024. [34] Peter Mattson, Christine Cheng, Gregory Diamos, Cody Coleman, Paulius Micikevicius, David Patterson, Hanlin Tang, Gu-Yeon Wei, Peter Bailis, Victor Bittorf, et al. Mlperf training benchmark. Proceedings of Machine Learning and Systems, 2:336–349, 2020. [35] Meta. MTIA v1: Meta’s first-generation AI inference accelerator. https://ai.meta.com/blog/ meta-training-inference-accelerator-AI-MTIA/, 2023. Accessed: 2026-05-04. [36] Meta Platforms, Inc. Holistic trace analysis. https://hta.readthedocs.io/en/latest/, 2024. Accessed 2026-04-22. [37] Xupeng Miao, Yujie Wang, Youhe Jiang, Chunan Shi, Xiaonan Nie, Hailin Zhang, and Bin Cui. Galvatron: Efficient transformer training over multiple GPUs using automatic parallelism. In VLDB, 2023. [38] Microsoft. MSCCL++: A GPU-driven communication stack for scalable AI applications. https: //github.com/microsoft/mscclpp, 2026. Accessed 2026-04-26. [39] Microsoft. Microsoft azure. https://azure.microsoft.com/en-us, 2026. Accessed: 2026-05-06. [40] MLCommons. Chakra. https://mlcommons.org/working-groups/research/chakra/, 2026. Accessed 2026-04-22. [41] Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley, Mostofa Patwary, Vijay Anand Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, Amar Phanishayee, and Matei Zaharia. Efficient large-scale language model training on gpu clusters using megatron-lm. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, 2021. doi: 10.1145/3458817.3476209. [42] National Energy Research Scientific Computing Center. Perlmutter architecture. https://docs.nersc. gov/systems/perlmutter/architecture/, 2026. Accessed: 2026-05-06. [43] NVIDIA. H100 GPUs Set Standard for Gen AI in Debut MLPerf Benchmark. https://blogs.nvidia. com/blog/generative-ai-debut-mlperf/, 2023. Accessed 2026-04-26. [44] NVIDIA. NVIDIA Blackwell takes pole position in latest MLPerf inference results. https://blogs. nvidia.com/blog/blackwell-mlperf-inference/, 2025. Accessed 2026-04-26. [45] NVIDIA. NVIDIA Blackwell Ultra sets the bar in new MLPerf inference benchmark. https://blogs. nvidia.com/blog/mlperf-inference-blackwell-ultra/, 2025. Accessed 2026-04-26. [46] NVIDIA. Understanding NCCL tuning to accelerate GPU-toGPU communication. https://developer.nvidia.com/blog/ understanding-nccl-tuning-to-accelerate-gpu-to-gpu-communication/, 2025. Accessed 2026-04-26. [47] NVIDIA. NVIDIA wins every MLPerf Training v5.1 benchmark. https://blogs.nvidia.com/blog/ mlperf-training-benchmark-blackwell-ultra/, 2025. Accessed 2026-04-26. [48] NVIDIA. Nvidia collective communication library (NCCL) documentation. https://docs.nvidia. com/deeplearning/nccl/user-guide/docs/, 2026. Accessed: 2026-05-04. [49] NVIDIA. NCCL Tests. https://github.com/NVIDIA/nccl-tests, 2026. Accessed 2026-04-26. [50] NVIDIA. NVIDIA Nsight Systems. https://docs.nvidia.com/nsight-systems/, 2026. Accessed: 2026-05-06. [51] NVIDIA. NVIDIA NVLink and NVLink Switch. https://www.nvidia.com/en-us/data-center/ nvlink/, 2026. Accessed: 2026-05-06.

12

[52] OpenXLA Project. XLA Operation Semantics. https://openxla.org/xla/operation_semantics, 2026. Accessed: 2026-05-06. [53] OpenXLA Project. Profiling jax computations with xprof. https://openxla.org/xprof/jax_ profiling, 2026. Accessed: 2026-05-05. [54] Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, and Ricardo Bianchini. Splitwise: Efficient generative llm inference using phase splitting. In 2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA), pages 118–132. IEEE, 2024. Accelerating PyTorch with CUDA Graphs. https://pytorch.org/blog/ [55] PyTorch. accelerating-pytorch-with-cuda-graphs/, 2021. Accessed: 2026-05-05. [56] PyTorch Foundation. Libkineto readme. https://github.com/pytorch/kineto/blob/main/ libkineto/README.md, 2026. Accessed 2026-04-23. [57] Kun Qian, Yongqing Xi, Jiamin Cao, Jiaqi Gao, Yichi Xu, Yu Guan, Binzhang Fu, Xuemei Shi, Fangbo Zhu, Rui Miao, et al. Alibaba hpn: A data center network for large language model training. In Proceedings of the ACM SIGCOMM 2024 Conference, pages 691–706, 2024. [58] Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. Zero: Memory optimizations toward training trillion parameter models. In SC20: international conference for high performance computing, networking, storage and analysis, pages 1–16. IEEE, 2020. [59] Samyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith, and Yuxiong He. Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning. In Proceedings of the international conference for high performance computing, networking, storage and analysis, pages 1–14, 2021. [60] Vijay Janapa Reddi, Christine Cheng, David Kanter, Peter Mattson, Guenther Schmuelling, Carole-Jean Wu, Brian Anderson, Maximilien Breez, et al. MLPerf inference benchmark. In ISCA, 2020. [61] Jie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, and Yuxiong He. ZeRO-Offload: Democratizing Billion-Scale model training. In 2021 USENIX Annual Technical Conference (USENIX ATC 21), pages 551–564. USENIX Association, July 2021. ISBN 978-1-939133-23-6. URL https://www.usenix.org/conference/atc21/presentation/ ren-jie. [62] Amit Sabne. Xla: Compiling machine learning for peak performance. 2020. [63] SemiAnalysis. InferenceX: Inference benchmarking. https://inferencex.semianalysis.com/ inference, 2026. Accessed 2026-04-26. [64] SGLang Team. SGLang bench serving guide. https://github.com/sgl-project/sglang/blob/ main/docs/developer_guide/bench_serving.md, 2026. Accessed 2026-04-26. [65] SGLang Team. SGLang documentation. https://docs.sglang.io/, 2026. Accessed 2026-04-26. [66] Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. Megatron-lm: Training multi-billion parameter language models using model parallelism. arXiv preprint arXiv:1909.08053, 2019. [67] Jovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Josep Torrellas, and Esha Choukse. Dynamollm: Designing llm inference clusters for performance and energy efficiency. In 2025 IEEE International Symposium on High Performance Computer Architecture (HPCA), pages 1348–1362. IEEE, 2025. [68] Philippe Tillet, Hsiang-Tsung Kung, and David Cox. Triton: an intermediate language and compiler for tiled neural network computations. In Proceedings of the 3rd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages, pages 10–19, 2019. [69] UALink Consortium. Ultra Accelerator Link (UALink) 200G 1.0 Specification. ualinkconsortium.org/specification/, 2026. Accessed: 2026-05-06.

https://

[70] vLLM Project. vLLM Bench Latency. https://docs.vllm.ai/en/stable/cli/bench/latency/. Accessed: 2026-05-05. [71] vLLM Team. vLLM v0.6.0: 2.7x throughput improvement and 5x latency reduction. https://vllm.ai/ blog/perf-update, 2024. Accessed 2026-04-26.

13

[72] vLLM Team. vLLM benchmarks. https://docs.vllm.ai/en/v0.14.1/api/vllm/benchmarks/, 2026. Accessed 2026-04-26. [73] Xizheng Wang et al. SimAI: Unifying architecture design and performance tuning for large-scale large language model training with scalability and precision. NSDI, 2025. [74] Yu Wang, Gu-Yeon Wei, and David Brooks. A systematic methodology for analysis of deep learning hardware and software platforms. Proceedings of Machine Learning and Systems, 2:30–43, 2020. [75] Yuxin Wang, Yuhan Chen, Zeyu Li, Xueze Kang, Yuchu Fang, Yeju Zhou, Yang Zheng, Zhenheng Tang, Xin He, Rui Guo, et al. Burstgpt: A real-world workload dataset to optimize llm serving systems. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2, pages 5831–5841, 2025. [76] William Won, Taekyung Shi, Dhruvit Ajith, Saeed Sudarshan, Adarsh Ravichandran, and Tushar Krishna. ASTRA-sim2.0: Modeling hierarchical networks and disaggregated systems for large-model training at scale. arXiv preprint arXiv:2303.14006, 2023. [77] Hengrui Zhang, August Ning, Rohan Baskar Prabhakar, and David Wentzlaff. LLMCompass: Enabling efficient hardware design for large language model inference. In 2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA), pages 1080–1096. IEEE, 2024. doi: 10.1109/ISCA59077. 2024.00082. [78] Wei Zhang, Wei Wei, Lingjie Xu, Lingling Jin, and Cheng Li. Ai matrix: A deep learning benchmark for alibaba data centers. arXiv preprint arXiv:1909.10562, 2019. [79] Lianmin Zheng, Zhuohan Li, Hao Zhang, Yonghao Zhuang, Zhifeng Chen, Yanping Huang, Yida Wang, Yuanzhong Xu, Danyang Zhuo, Eric P Xing, Joseph E Gonzalez, and Ion Stoica. Alpa: Automating interand intra-operator parallelism for distributed deep learning. In OSDI, 2022. [80] Hongyu Zhu, Mohamed Akrout, Bojian Zheng, Andrew Pelegris, Amar Jayarajan, Amar Phanishayee, Bianca Schroeder, and Gennady Pekhimenko. Benchmarking and analyzing deep neural network training. In 2018 IEEE International Symposium on Workload Characterization (IISWC), pages 88–100. IEEE, 2018.

A

Trace collection

CCL-Bench accepts two families of trace format, JSON traces and NSYS traces, with JSON traces as the preferred source for general metrics. Chrome-trace JSON files (traceEvents array) are produced on both GPU (Kineto/CUDA backend) and TPU (XLA backend). We prioritize this format because it is portable across accelerator families, exposes the operator, kernel, communication, rank, stream, and timestamp structure needed by most CCL-Bench metrics. GPU and TPU Chrome-trace JSON files share the same container format but differ in several aspects of their event schema: • Timestamps. GPU events carry ts/dur wall-clock timestamps in microseconds (CUDA host clock); TPU events use device_offset_ps / device_duration_ps in picoseconds (XLA device clock), requiring unit conversion before any cross-platform time comparison. • FLOPs reporting. XLA records model_flops directly in trace-event arguments, so the MFU tool can sum device-reported FLOPs; GPU traces carry no equivalent field, so FLOPs must be estimated from the workload-card model architecture. • Step boundaries. GPU traces delimit iterations with ProfilerStep#N spans emitted by the PyTorch profiler; TPU traces use $core.py:331 step (or equivalent XLA step markers) as iteration boundaries. • Collective naming. GPU collectives appear as NCCL kernel names (e.g., ncclAllReduce_...); TPU collectives appear as XLA HLO category strings (e.g., all-reduce, all-gather) attached to device-side events. • Memory-transfer events. GPU traces expose explicit memcpy/memset CUDA kernel records in separate streams; TPU traces encode DMA as paired copy-start.k / copy-done.k async events identified by device-side timestamps. 14

• Stream model. GPU traces interleave events across multiple CUDA streams (compute, communication, memory copy); TPU traces typically present a single device execution timeline per chip, with host annotations on a separate thread. NSYS traces are NVIDIA Nsight Systems SQLite databases exported from nsys profile. They provide NVIDIA-specific CUDA, CUPTI, and interconnect details that are useful for hardwarecounter metrics, but they are less general across accelerators and are treated as auxiliary evidence when a metric needs information not present in the Torch trace. JSON trace can be collected with low profiling overhead in routine benchmark runs. To show this, we train a 1.88B test model using tensor parallelism on one machine with Intel Xeon Gold 6438M2 CPU and 2 NVIDIA L40 GPUs, and find that the tracing overhead is only 0.22% (averaged over 6 steps), with a coefficient of variation of 2.67%.

B

Workload-card and trace schema Table 2: Workload-card schema fields from the YAML template. Role Provenance

Workload model Model architecture Workload data Hardware network Hardware XPU Executor framework Parallelization

Communication

Metric source

C

Template field

Example

1; https:// version; description; hf_url; trace_url; contributor; contact huggingface.co/... training; false; model_fwd_bwd_pass; workload.model.phase; moe; granularity; model_family; precision; llama-3.1-8b; bf16 epochs; iteration 16380544000; workload.model.model_arch.num_params; num_params_embedding; 28; 16; 128 num_layers; num_heads; head_dim 4; 8192; workload.data.batch_size; seq_len; input_len; output_len; dataset 1024; 128; c4 slingshot; workload.hardware.network_topo.topology; bandwidth_gbps [200, 2000] GPU; nvidia_a100; workload.hardware.xpu_spec.type; model; total_count; 16; 4; cuda_12.4 count_per_node; workload.hardware.driver_version torchtitan; Model-executor.framework.name; compiler_tool_selection plain_pytorch 1; 2; 4; Model-executor.model_plan_parallelization.dp_replicate; 2; 1; 1; 1 dp_shard; tp; pp; cp; ep; pp_mb NCCL; 2.14.3; NCCL_IB_QPS...; Model-executor.communication_library.name; version; env; [rocev2, p2p] Model-executor.protocol_selection [nsys, json]; metric_source.traces; metrics_specific_trace [memory_trace]

Run scripts and trace collection method.

Contributors attach the launch scripts and environment definitions, and run at least five training iterations (steps) per workload in steady state (after 5 initial warm-up iterations that allow CUDA kernels to compile and memory to settle [55]). For inference, contributors first send a small number of warm-up batches so that the serving engine completes CUDA graph capture and KV-cache initialization [70], then collect traces that cover both the prefill pass (prompt processing) and at least 128 steady-state decode steps (autoregressive token generation). The iteration count enters the workload card so downstream analyses can quantify the variance behind a published number. Contributors are not required to perform full workload training or fine-tuning, as CCL-Bench does not collect accuracy metrics.

D

Metric tool catalog

Each CCL-Bench metric is implemented as a self-contained Python tool under tools/<metric>/. The dispatcher (tools/main.py) reads the workload card’s metric_source.traces field and invokes the appropriate backend for each trace format. Table 3 lists all currently supported metrics. 15

Table 3: CCL-Bench 1.0 metric catalog. “Dir” indicates whether a higher (↑) or lower (↓) value is preferred. P = PyTorch/Kineto JSON (GPU) or XLA profiler JSON (TPU); Pgpu = GPU-only Kineto; N = Nsight Systems SQLite. Metric key

Label

Unit Dir Src

Model execution avg_step_time

Step Time

s

↓

P

mfu

MFU

%

↑

P

ttft

TTFT

s

↓

P

tpot

TPOT

s

↓

P

Compute mean_sm_coverage

SM Coverage

%

↑

%

↓

Pgpu , N Average SM occupancy across all GPU kernels, weighted by kernel duration. P, N Fraction of total GPU/TPU time in the single most time-consuming kernel or HLO category. P, N Fraction of GPU time in compute-bound kernels (SM coverage >70% and duration >10 µs). P, N Fraction of GPU/TPU time in memory-bound kernels (SM coverage <50%). Pgpu , N Fraction of GPU time in MoE-specific kernels (expert compute, routing, dispatch/combine).

dominant_kernel_concentration Top Kernel compute_bound_fraction

Compute-Bound

%

↑

memory_bound_fraction

Memory-Bound

%

↓

moe_fraction

MoE Fraction

%

↓

Memory average_memory_bandwidth

Avg Mem BW

GB/s

↑

P, N

memory_transfer_overhead

Mem Transfer OH

%

↓

P, N

Communication communication_fraction

Comm Fraction

%

↓

P, N

total_communication_time

Comm Time

s

↓

P, N

compute_comm_overlap

Overlap

%

↑

P, N

load_imbalance_ratio

Load Imbalance

ratio

↓

P, N

straggler

Straggler

ratio

↓

Pgpu

bw_allgather

AllGather BW

GB/s

↑

P

bw_allreduce

AllReduce BW

GB/s

↑

P

bw_reducescatter

ReduceScatter BW

GB/s

↑

P

bw_alltoall

AllToAll BW

GB/s

↑

P

bw_peertopeer

P2P BW

%

↑

N

Utility scale_up_bw_utility

Scale-Up BW Utility

%

↑

P

16

Description Average wall-clock time per engine iteration (full fwd+bwd for training; decode step for inference). Model FLOPs Utilization: observed FLOP/s ÷ (peak FLOP/s × device count). Time to first token: latency of the prefill pass. Inference only. Time per output token: mean decode-step latency after the prefill. Inference only.

Sustained bandwidth during memory-copy operations (total bytes / total copy duration). Fraction of step time in exposed DMA/memcpy intervals with no concurrent compute. See §D.2. Fraction of GPU/TPU time in collectivecommunication kernels (NCCL, XLA collectives). Average per-step communication time with compute-overlapped intervals subtracted. Fraction of communication time that executes concurrently with compute kernels. See §D.3. Max-to-min GPU active time across ranks; 1.0 = perfectly balanced. Straggler delay (max − min)/ max across ranks per NCCL collective. Median effective AllGather bandwidth; algorithm factor (N −1)/N applied per group size. Median effective AllReduce bandwidth; algorithm factor 2(N −1)/N applied per group size. Median effective ReduceScatter bandwidth; algorithm factor (N −1)/N applied per group size. Median effective AllToAll bandwidth (expert dispatch/combine in MoE). NVLink/PCIe TX utilization during peer-to-peer transfers (pipeline parallelism). Step-time improvement from doubling intra-node (scale-up) bandwidth, via Astra-sim trace-replay simulation. Near 0% = compute-bound or scaleup-insensitive; high = intra-node fabric is on the critical path.

D.1

Model FLOPs utilization: calculation

Definition. Model FLOPs Utilization (MFU) reports the fraction of peak accelerator FLOPs used by the model computation: observed model FLOPs/s MFU = × 100. (2) peak FLOPs/s across all XPUs The tool reads the workload card to obtain the trace type, hardware model, accelerator count, workload phase, sequence length, batch size, and model architecture fields. The denominator is Fpeak × world_size, where Fpeak is the per-XPU BF16 peak FLOP/s for the declared hardware model. PyTorch trace backend (GPU JSON). For PyTorch-profiler JSON traces, the tool estimates the model FLOPs per token from the workload-card model architecture. Let Nact be the active per-token parameter count (or total parameter count if the active count is unavailable), Nemb the embedding parameter count, L the number of layers, H the head dimension, Q the number of attention heads, and S the sequence length. The per-token FLOP estimate is  2(Nact − Nemb ) + 4LHQS, inference, ftoken = (3) 6(Nact − Nemb ) + 12LHQS, training. With batch size B and step time Tstep , observed model FLOPs/s = ftoken ×

BS . Tstep

(4)

XLA / TPU trace backend. For TPU traces, XLA records model_flops in trace-event arguments. The tool sums positive model_flops values and divides by the unioned active duration of events that carry those FLOPs, yielding active FLOP/s. Because the XLA profiler reports model FLOPs summed across all chips in the run, the denominator is the per-chip peak FLOP/s multiplied by xpu_spec.total_count. D.2

Memory transfer overhead: derivation

Definition. The memory transfer overhead reports the fraction of execution time spent in exposed memory copies — transfers that execute while no compute kernel is running concurrently. Transfers that are pipelined with compute add no wall-time overhead and are therefore excluded. Concurrent transfers are unioned before subtraction so that overlapping copies are not double-counted. Let union(S) denote the total duration of the interval union of a set of intervals S, and let pure(M, C) = union(M) \ union(C) denote the DMA time with no concurrent compute. GPU / NSYS backend. Let M be all CUDA memcpy records from the CUPTI CUPTI_ACTIVITY_KIND_MEMCPY table and K all GPU kernel records from CUPTI_ACTIVITY_KIND_KERNEL. pure(M, K) × 100, (5) Mem-OHnsys = Ttrace where Ttrace is the wall-clock span from the earliest to the latest activity across M ∪ K. PyTorch trace backend (GPU and TPU). Both GPU (Kineto/CUDA) and TPU (XLA) produce Chrome-trace JSON files and share the same algorithmic structure. For each trace file r, let Mr be the set of memory-transfer intervals and Cr the compute intervals; the metric is X pure(Mr , Cr ) Mem-OHtorch =

r

X

× 100.

(6)

Tr

r

Summing per-file numerators and denominators avoids mixing device clocks across ranks or chips. The two architectures differ only in how Mr , Cr , and Tr are extracted from the trace: GPU (Kineto). Events are complete intervals (ph="X") with ts (start, µs) and dur (duration, µs). Mr contains kernel events whose name matches copy-like patterns (e.g., memcpy, memset, copy_kernel, dma). Cr contains all remaining non-communication kernels. Tr is the kernel-activity span [min ts, max(ts + dur)] for rank r. 17

TPU (XLA). Device-side ops are annotated with device_offset_ps (device clock start, ps) and device_duration_ps (duration, ps). Mr is built from async copy-start.k / copy-done.k event pairs using their device_offset_ps timestamps as interval endpoints. Cr contains leaf XLA compute ops (events with both device_offset_ps and device_duration_ps), excluding high-level jit_* container spans that encompass entire model sub-graphs and dependency-wait synchronization barriers. Tr = Tstep × 106 converts the total $core.py:331 step duration (µs) to picoseconds. Because XLA aggressively pipelines DMA with compute, the non-overlapped fraction is typically near 0 % for well-optimized workloads. D.3

Compute-communication overlap: derivation

Definition. The compute-communication overlap reports the fraction of collective-communication time that runs concurrently with compute kernels. Let Kcomm and Kcomp be the sets of communication and compute kernel intervals within a step window, and let union(S) denote the total duration of the merged interval union of set S. The per-step overlap ratio is union(Kcomm ) ∩ union(Kcomp ) Overlap = × 100. (7) union(Kcomm ) Communication kernels are identified by matching kernel names against NCCL and XLA collective patterns (e.g., ncclAllReduce, all-gather); all remaining GPU/TPU kernels are treated as compute. The final metric averages the per-step ratio across inner steps (first and last steps excluded to avoid warm-up and cool-down skew) and then across ranks. GPU JSON backend. Within each step window, GPU kernel events (ph="X", cat="kernel") are split into Kcomm and Kcomp by name, each set is interval-merged, and the intersection is computed. Overlap is then averaged over inner steps per rank and across ranks. XLA / TPU backend. Comm events are matched with a strict regex that excludes broadcast, which in XLA HLO denotes a local data-layout operation rather than a network collective. Compute events are identified by XLA HLO category keywords (dot, convolution, gemm, fusion, custom-call). Nsight Systems backend. All GPU kernel intervals are read from CUPTI_ACTIVITY_KIND_KERNEL. Because NVTX step ranges are not always present in NSYS traces, the metric is computed globally over the full trace rather than per step. Intervals are partitioned per device before merging to avoid falsely counting simultaneous per-device communication as overlap.

E

CCL-Bench user interface

18

Figure 8: CCL-Bench user interface. The interface provides a dashboard for performance ranking, workload comparison, and trace upload.

F

Selected workloads

ID

Table 4: Selected workload suite for CCL-Bench 1.0. Model Phase Batch Sequence/input length

WL1 WL2 WL3 WL4 WL5 WL6 WL7

Qwen3-4B Llama-3.1-8B DeepSeek-MoE-16B Llama-3.1-8B DeepSeek-V3-16B DeepSeek-V3-16B DeepSeek-V3-236B

Inference Inference Inference Training Training Training Training

128 128 128 4 8 64 64

G

Extended cross-system and cross-architecture study

G.1

Pairwise cross-system examples

1024 input 1024 input 1024 input 512 sequence 1024 sequence 2048 sequence 1024 sequence

Figure 9 compares NCCL and MSCCL++ on matched Perlmutter A100 vLLM inference runs. MSCCL++ lowers communication fraction for Qwen3-4B (WL1, TP4) from 49.5% to 46.5% and 19

MSCCL++

NCCL

TPOT (ms)

Communication (%)

NCCL

60 40 20 0

Qwen3-4B (WL1) TP4/B128

Llama-3.1-8B (WL2) DeepSeek-MoE (WL3) TP4/B128 TP4/EP4/B128

80 60 40 20 0

Qwen3-4B (WL1) TP4/B128

MSCCL++

Llama-3.1-8B (WL2) DeepSeek-MoE (WL3) TP4/B128 TP4/EP4/B128

vLLM

SGLang

50 25 0

vLLM

TPOT (ms)

TTFT (ms)

75

Llama-3.1-8B (WL2)

Qwen3-4B DeepSeek-MoE-16B (WL1) (WL3)

SGLang

75 50 25 0

Llama-3.1-8B (WL2)

Qwen3-4B DeepSeek-MoE-16B (WL1) (WL3)

Communication (%)

Figure 9: Example cross-system comparisons on Qwen3-4B (WL1), Llama-3.1-8B (WL2), and DeepSeek-MoE-16B (WL3). Each pair holds workload, hardware, framework, TP/EP degree, and accelerator count fixed while varying the communication stack.

vLLM

60

SGLang

40 20 0

Llama-3.1-8B (WL2)

Qwen3-4B DeepSeek-MoE-16B (WL1) (WL3)

Figure 10: Example cross-system comparisons on Qwen3-4B (WL1), Llama-3.1-8B (WL2), and DeepSeek-MoE-16B (WL3). Each pair holds workload, hardware, communication stack, TP/EP degree, and accelerator count fixed while varying the framework. for DeepSeek-MoE-16B (WL3, TP4/EP4) from 53.7% to 46.4%. The latency effect is workloaddependent: TPOT improves for WL1 (43.6 ms to 42.9 ms) and WL3 (81.2 ms to 76.0 ms), but regresses for Llama-3.1-8B (WL2) from 36.6 ms to 38.4 ms. Figure 10 compares vLLM and SGLang on A100s while keeping communication backend, parallelism, and other configuration fixed. vLLM is faster on WL1 and WL2, while SGLang is faster on WL3 and also lowers WL3 communication fraction (38.8% vs. 53.7%). The CCL-Bench secondary metrics reveal the mechanism behind each case. Dense inference (WL1/WL2): vLLM faster due to higher compute efficiency. For Qwen3-4B and Llama-3.1-8B, the dominant collective is A LL R EDUCE after each attention and MLP layer. vLLM achieves substantially higher MFU than SGLang on both workloads (WL1: 17.1% vs. 8.6%; WL2: 28.4% vs. 15.5%), while communication fraction is similar across frameworks (WL1: 49.5% vs. 53.1%; WL2: 37.0% vs. 41.3%). The communication bottleneck is identical. The gap comes from vLLM’s more efficient compute kernels, reflecting more aggressive CUDA graph capture and attention kernel tuning. MoE inference (WL3): SGLang faster due to more efficient A LLT OA LL. For DeepSeekMoE-16B (EP4, TP4), the bottleneck shifts to A LLT OA LL for expert token dispatch and combine. vLLM exposes 53.7% of step time as pure communication. SGLang cuts exposed communication to 38.8%—a 15 percentage point reduction—through a more efficient A LLT OA LL implementation, improving wall time (76.1 ms vs. 81.2 ms). Figure 11 reports the TPU-only tensor-parallel sweep: for Llama-3.1-8B batch-128/input-1024 inference, moving from TP1 to TP4 reduces average step time from 222 ms to 85 ms while increasing communication fraction from effectively zero to 19.3%. TP8 raises the communication fraction to 26.9% and regresses step time to 98 ms. G.2

GPU vs. TPU: A Cross-Architecture Overview

Figure 12 provides a holistic, architecture-level view of A100 GPU (Perlmutter/Slingshot) versus TPU v6e (Torus) across five representative workloads (WL1–WL3 inference, WL4–WL5 training), averaging over all CCL-Bench leaderboard entries that fall within each workload task. The four panels compare step time, MFU, CPU-chip memory bandwidth, and communication fraction for the same workload tasks across the two platforms. 20

Communication (%)

Step time (ms)

250 225 200 175 150 125 100

30 25 20 15 10 5 0

75 TP1

TP2

TP4

TP8

TP1

TP2

TP4

TP8

Figure 11: Supplementary TPU v6e tensor-parallel sweep for Llama-3.1-8B batch-128/input-1024 inference.

101

TPU v6e (Torus) 60

100

40

MFU (%)

Step time (s)

A100 GPU (Perlmutter)

10 1 10 2

20 0

WL-14B WL2B WL3E WL4B WL5B 3 fer ama-8fer SK-Mofer ama-8ain SK-16ain n e Qw in Ll in D in Ll tr D tr

WL-14B WL2B WL3E WL4B WL5B 3 fer ama-8fer SK-Mofer ama-8ain SK-16ain n e Qw in Ll in D in Ll tr D tr

102 101

Comm frac. (%)

CPU-chip BW (GB/s)

100 75 50 25 0

WL-14B WL-28B WL3oE WL-48B WL56B M a a 3 -1 n r r r n wen infe Llam infe DSK infe Llam trai DSK trai

WL-14B WL-28B WL3oE WL-48B WL56B M a a 3 r r r n n fe m fe K fe m i SK-1 in Qwe in Lla in DS in Lla tra D tra

Q

Figure 12: GPU vs. TPU holistic comparison across WL1–WL5. Each bar is the average over all CCL-Bench entries for that workload task and hardware platform.

Compute utilization. TPU v6e achieves higher MFU than A100 on all five workloads: 2.1× higher for WL1 (36.3% vs. 16.9%), 1.5× for WL2 (41.1% vs. 27.7%), 4.3× for WL3 (22.6% vs. 5.2%), 10.4× for WL4 (50.9% vs. 4.9%), and 1.6× for WL5 (7.8% vs. 4.9%). For training (WL4/WL5), GPU MFU is suppressed by higher communication overhead from the scale-out network. Step time. TPU v6e achieves 2.0× lower step time on WL1 and WL2 inference: 23 ms vs. 46 ms on WL1 and 21 ms vs. 43 ms on WL2, consistent with its higher MFU. For WL3, TPU reports a higher average step time (1.09 s vs. 0.12 s) despite higher MFU. For training, TPU is 3.0× faster on WL4 and 39.9× faster on WL5, reflecting the GPU’s larger communication fraction on these workloads. CPU-chip memory-transfer bandwidth. GPU training entries (WL4/WL5) show much higher CPU-chip transfer bandwidth (240–484 GB/s) than TPU (45–140 GB/s), reflecting larger host-device transfer volume in the recorded multi-GPU training traces. For inference, CPU-chip bandwidth is closer across architectures: 12.5 vs. 16.2 GB/s on WL1, 10.8 vs. 21.3 GB/s on WL2, and 14.6 vs. 40.2 GB/s on WL3. Communication fraction. GPU entries spend a larger fraction of traced time in communication on every workload shown: 37.9% vs. 14.7% on WL1, 36.4% vs. 15.7% on WL2, 36.4% vs. 0.9% on WL3, 35.2% vs. 16.9% on WL4, and 93.1% vs. 2.0% on WL5. 21

H

Additional result from trace-based simulation Compute + idle timeline Exposed comm Overlapped comm Configuration metadata 12.17s, comm 39.0% Baseline Scenario SU SU BW SO BW Hidden 10.63s, -1541ms Balanced Baseline 4 300 GB/s 200 Gbps 53% Balanced 32 1800 GB/s 400 Gbps 79% 10.58s, -1596ms Aggressive Aggressive 128 3600 GB/s 800 Gbps 81% 0.0 2.5 5.0 7.5 10.0 12.5 Step time (s)

Figure 13: Cluster architecture sweep for WL7 (TP=4, PP=8, DP=8, EP=32) DeepSeek-V3-236B, varying scale-up domain size, scale-up bandwidth, and scale-out bandwidth together. The scale-up (SU) topology is fully-connected. The scale-out (SO) topology is switch-based. We apply the simulation pipeline to the largest trace we collected, DeepSeek-V3-236B (WL7) on 256 GPUs, to understand the workload behavior under different network conditions. Figure 13 shows that having additional network resources (scale-up domain size, scale-up bandwidth, and scale-out bandwidth) has a diminishing effect in increasing the overlapped portion of communication (hidden by compute). This calls for the importance of reducing network latency in different infrastructure layers: framework collective operation scheduling, collective library algorithm, RDMA protocol, XPU-NIC architecture, and network fabric design.

I

CCL-Search additional details

CCL-Search uses claude-opus-4-6 via the Anthropic Messages API that forces the model to submit a generate_config Python policy via a structured tool-use call each iteration. The system prompt and full execution history (configs, scores, errors) are passed as user-turn context on every iteration. In our experiment, the average LLM policy-synthesis time is around 33s per iteration. Figure 14 shows CCL-Search’s support for composite objectives: using Score = w × T0 /T + (1 − w) × N0 /N , the benchmark finds configurations that balance step time against XPU count for WL4 Llama-3.1-8B workload task, complementing the result for WL5 DeepSeek-V3-16B (Figure 7(b)). Figure 15 shows another configuration searching of WL5 workload, using a different seed policy (TP=4, DP=1, PP=3, EP=1, microbatch size=4, no activation checkpointing) from the experiment in Figure 7(b) (TP=4, DP=4, PP=1, EP=4, microbatch size=2, activation checkpointing). The final policies only differ in micro-batch sizes.

22

Megatron-LM: GPU Eff. + Step time iter 0 Step time (s)

Step time (s)

0.35 0.32 0.31 0.30 0.29 0.28

iter 4 iter iter 911

4

Megatron-LM: GPU Eff. + Step Time

iter 3

2.00 2.502.76

10.00 5.00 2.00 0.45

6

8

1.50

10

12

14

Num GPUs (TP × DP × PP)

3.0 2.5 2.0 1.5 1.0 0.5 0.0

16

iter 8 iter 5

iter 0

iter 4 iter 7 2.00 1.93iter 9

8

10

12

14

1.50

16

Num GPUs (TP × DP × PP)

Config choices

Config choices

TP 4 2 4 4 2 2 2 2 4 2 1 2 1 1 2 DP 1 1 1 2 2 2 2 2 4 2 4 4 4 4 2 PP 3 3 1 1 3 3 3 3 1 3 3 1 3 3 1 EP 1 1 1 2 2 2 1 2 4 2 2 2 4 2 2 MBS 4 2 1 1 2 1 2 4 2 4 4 4 4 4 1 Act Ckpt F T T T T T T T T F F T T T T

TP 4 1 2 4 4 2 1 4 1 4 4 4 2 4 2 4 2 DP 4 1 2 2 1 1 1 1 1 1 1 1 1 1 2 1 1 PP 1 1 1 1 1 1 2 1 2 1 1 1 1 1 1 1 1 MBS 2 1 4 2 2 1 1 4 1 1 4 1 1 4 2 4 1 Act Ckpt F T T T T T T T T T T F T F T F T

0 1 2 3 4 5 6 7 8 9 10 11 12 13 14

Iteration

0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16

Iteration

Figure 15: Composite objective: Score = w × T0 /T + (1 − w) × N0 /N . WL5 DeepSeek-V316B on Perlmutter. This run uses a different seed policy compared to Figure 7(b). w = 0.5

Figure 14: Composite objective: Score = w × T0 /T + (1 − w) × N0 /N . WL4 Llama-3.1-8B on Perlmutter. w = 0.5

Optimizer prompt. The system prompt used in CCL-Search configuration optimization agent is listed below. You are a configuration optimization agent for LLM infrastructure (CCL-Bench). Your goal: write and iteratively refine a Python function ‘generate_config‘ that maps workload cards and environment descriptors to configuration key-value pairs optimizing the user-defined performance objective. ## Function signature ‘‘‘python def generate_config(workload: dict, environment: dict) -> dict: """ Args: workload: Workload card fields –model_family, phase, batch_size, seq_len, num_heads, num_layers, precision, moe (bool), config_space (list of tunable dimensions with valid choices), run_script, trace_dir. environment: Hardware/software descriptor –gpu_model, gpu_memory_gb, total_gpus, gpus_per_node, intra/inter_node_bandwidth_gbps, framework, framework_version. Returns: dict of configuration key-value pairs matching config_space keys, e.g.: {"tp": 4, "dp": 8, "pp": 1, "micro_batch_size": 4, "activation_checkpointing": True} """ ‘‘‘ ## Context you receive - The current ‘generate_config‘ source policy. - Full execution history: every config tried, its measured metrics, and its score. Use this to understand what worked, what failed, and why. - A summary table of all runs with scores. ## Workflow 1. Analyse the history –which configs performed best, which failed and why. 2. Submit an improved ‘generate_config‘ via ‘submit_config‘. 3. The new function is executed immediately; results appear next iteration. ## Scoring score = weighted sum of CCL-Bench metrics (lower is better when ‘minimize‘). Priority 1 –fix errors/timeouts. Priority 2 –improve the score. ## Design guidance for adaptive policies

23

Write general, model-aware, environment-aware logic –NOT a fixed config dict. The policy should reason about the workload and hardware to pick the right parallelism strategy: ### Key workload fields to use: - ‘workload["model_family"]‘ –model name (e.g., "llama-3.1-8b", "deepseek-v2-lite") - ‘workload["moe"]‘ –True if Mixture-of-Experts model - ‘workload["num_layers"]‘ –number of transformer layers (affects PP choices) - ‘workload["num_params"]‘ –total parameter count (affects memory requirements) - ‘workload["batch_size"]‘ –global batch size - ‘workload["seq_len"]‘ –sequence length - ‘workload["config_space"]‘ –list of tunable dimensions with valid choices ### Key environment fields to use: - ‘environment["gpu_memory_gb"]‘ –GPU memory (e.g., 40 for A100-40GB) - ‘environment["total_gpus"]‘ –total GPUs available - ‘environment["gpus_per_node"]‘ –GPUs per node (e.g., 4) - ‘environment["inter_node_bandwidth_gbps"]‘ –inter-node bandwidth (affects TP/DP/EP tradeoffs) ### Parallelism reasoning principles: 1. **TP (tensor parallelism):** Keep TP within a node (tp <= gpus_per_node). Higher TP reduces per-GPU memory but adds allreduce communication. On slow interconnects, minimize cross-node TP. 2. **PP (pipeline parallelism):** PP must divide num_layers evenly. PP adds pipeline bubble overhead proportional to (PP-1)/num_microbatches. Use PP to reduce memory when TP alone isn’t enough. 3. **DP (data parallelism):** Scales throughput but adds allreduce for gradients. On slow interconnects, DP across nodes is expensive. 4. **EP (expert parallelism, MoE only):** EP must divide num_experts and dp. Distributes experts across GPUs. More EP = less per-GPU expert memory but more alltoall communication for token routing. On slow interconnects, alltoall is very expensive. 5. **Memory estimation:** Total training memory per GPU ≈ (num_params ×12 bytes) / (tp ×pp) + activation_memory. Must fit in gpu_memory_gb. Use activation_checkpointing=True if tight. 6. **micro_batch_size:** Larger = better GPU utilization but more activation memory. Start with 1, try 2, then 4. If OOM, reduce. 7. **tp ×dp ×pp** does NOT need to equal total_gpus. Using fewer GPUs can be more efficient if the model fits. ### MoE-specific guidance: - MoE models have sparse experts –only a subset are active per token - EP distributes experts across GPUs, reducing memory but adding alltoall communication - On slow interconnects (< 100 Gbps), alltoall for expert routing dominates step time - EP should divide both num_experts and dp - With EP, allgather/reducescatter for FSDP parameter sync can dominate –check if reducing DP or increasing EP helps ### Example adaptive policy structure: ‘‘‘python def generate_config(workload, environment): gpu_mem = environment.get("gpu_memory_gb", 40) gpus_per_node = environment.get("gpus_per_node", 4) total_gpus = environment.get("total_gpus", 16) is_moe = workload.get("moe", False) num_layers = workload.get("num_layers", 32) # Parse config space for valid choices config_space = {d["key"]: d["choices"] for d in workload.get("config_space", [])} # Start with TP fitting within a node tp = min(gpus_per_node, max(config_space.get("tp", [1]))) # Estimate memory and adjust parallelism # ... model-specific logic ... return {"tp": tp, "dp": dp, "pp": pp, ...} ‘‘‘ ## Important constraints - Each entry in config_space is a dict: ‘{"key": "tp", "type": "int", "choices": [1,2,4], "description": "..."}‘. Always use ‘dim["key"]‘ (not ‘dim["name"]‘) to get the dimension name. - Each iteration should incorporate lessons learned from the history. - You MUST call ‘submit_config‘ exactly once per iteration.

J

Broader Impacts

By providing a shared, reproducible substrate for benchmarking training and inference systems, CCL-Bench lowers the barrier for academic researchers, hardware vendors, and software developers to compare infrastructure choices in a rigorous and transparent way. This can accelerate progress on efficiency, reduce duplicated evaluation effort across organizations, and surface bottlenecks that headline-number benchmarks obscure. This democratizing effect is particularly valuable for academic groups and smaller companies that cannot reproduce the scale of large cloud providers. A notable limitation is that traces can contain sensitive information. Profiling data of open-sourced models may reveal cluster topology, operator sequences, or configuration choices that organizations 24

consider proprietary information. As a result, some contributors may be unwilling or legally unable to share raw traces publicly. CCL-Bench can partially address this by supporting metric-only submissions, where a contributor runs the analysis toolkit locally and submits only the computed metric values without the underlying trace. However, metric-only submissions sacrifice the auditability guarantee that a trace provides, so the community must balance openness against sensitivity on a caseby-case basis. Establishing clearer norms and tooling for trace anonymization (e.g., operator-name redaction, timing perturbation) remains an open problem that we hope future work will address. CCL-Bench does not train models, generate content, or interact with end users, so it does not raise the risks associated with generative AI misuse.

25

Record · ID 168235 · SHA-256 63b2d92e993af638
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.