ConceptioArchivearXiv CS
arXiv CSopen access

Tools-CC-Bench: a Benchmark Suite for Collective Communication with Compression in HPC and AI Workloads

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
clouddistributed-computingparallel-computing
distributed computing, parallel computing, cloud

Tools-CC-Bench: a Benchmark Suite for Collective Communication with Compression in HPC and AI Workloads Haozhe Fan∗‡§ , Wei Wang†§ , Xingchen Liu∗ , Man Liu∗ , Xingjian Tian∗ , Haoquan Long∗ , Zedong Liu∗ , Daran Sun∗ , Jinwu Yang∗ , Bo Yang† , Jie Liu† , Yonggang Che† , Hairui Zhao∗ , Guangming Tan∗ , Dingwen Tao∗¶

arXiv:2609.08739v1 [cs.DC] 8 Sep 2026

∗ Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China

{liuxingchen232, liuman24, tianxingjian25, longhaoquan25}@mails.ucas.ac.cn {sundaran24s, yangjinwu24z, zhaohairui, tgm, taodingwen}@ict.ac.cn, {captain.liu77}@gmail.com † College of Computer Science and Technology, National University of Defense Technology, Changsha, China {weiwang_jsjxy, yb, liujie, ygche}@nudt.edu.cn ‡ School of Computer Science, Nanjing University, Nanjing, China {231180018}@smail.nju.edu.cn

Abstract Distributed HPC and LLM workloads increasingly require efficient communication for scalability, yet growing data movement has become a major performance bottleneck. Communication compression can reduce this overhead and complement execution-level optimizations, but its benefits remain difficult to assess because existing benchmarks lack support for diverse backends, realistic datasets, applicationspecific accuracy metrics, and overlap-induced resource contention. We present CC-Bench, a lightweight, extensible, and application-oriented benchmark suite for evaluating communication compression under realistic execution conditions. CC-Bench uses declarative application-environment modeling to decouple profiling logic from communication libraries, datasets, and fidelity metrics, enabling portable cross-library evaluation. It further combines function-level interception and hardware counter monitoring to characterize per-phase latency, hardware utilization, numerical fidelity, and computation interference. With representative datasets from HPC and LLM workloads, CC-Bench evaluates three compression-enabled communication libraries on CPU and GPU clusters, revealing accuracy-performance trade-offs and bottlenecks to guide practical deployment and optimization.

1. Introduction Modern HPC and LLM workloads increasingly exceed the capacity of a single device. Applications ranging from molecular dynamics simulations with tens of billions of atoms to AI models with hundreds of billions of parameters rely on distributed systems to scale computation across many devices [7, 9, 12, 31, 35, 50, 54, 56]. Communication is fundamental to this scaling, enabling data exchange and synchronization. However, communication demand grows rapidly with workload scale, while network bandwidth § These authors contributed equally to this work. ¶ Corresponding author: Dingwen Tao, [email protected].

CPU Backend Open MPI

MVAPICH

GPU Backend NCCL

RCCL

Customize Existing Bench dataset Coarse-Grained Limited Quality Performance Metric

Real Workloads Requirements Real Datasets

Specific Metric

End-to-End Performance

(a) Existing Benchmarks

CPU Backend Real dataset &customize Specific metric ApplicationSpecific Metrics

GPU Backend CC-Bench

Full-Stack Profiling Infrastructure Phase-Level Breakdown

Real Workloads Requirements Phase Breakdown

Hardware Profiling

End-to-End Performance

(b) CC-Bench (Ours)

Figure 1: Benchmarking compression-supported communication libraries using our bench and previous methods.

remains comparatively constrained. Consequently, communication overhead has become a major bottleneck for workload scalability and end-to-end performance [44]. Existing efforts to mitigate communication overhead mainly fall into two categories. One line of work improves communication execution efficiency through topologyaware optimization [33, 51, 60], optimized communication algorithms [8, 14, 15], and communication resource management [10, 11, 29, 59, 61], facilitated by mature profiling and analysis tools [20, 24, 34]. Another line reduces the volume of transmitted data through compression. Communication compression is attractive because it is broadly applicable, offers substantial performance benefits, and complements execution-level optimizations. It has therefore become increasingly effective in real-world applications [21, 30, 39], and is now supported by a growing number of communication libraries [37, 40, 42, 58]. However, despite its promise and growing research interest, communication compression remains difficult to evaluate in realistic workloads due to the lack of efficient benchmarking tools. As a result, its practicality and performance characteristics remain poorly understood, limiting effective optimization and large-scale deployment for communication-intensive workloads. Traditional communication benchmarks, such as OSU Micro-Benchmarks and XCCL-Tests [49, 53], primarily measure communication efficiency but provide limited support for communication compression. Evaluating compression in real applications requires jointly characterizing com-

pression ratio, throughput, application-specific accuracy, and interference with overlapped computation. Although recent efforts have added compression support, they remain limited in three aspects [6], as shown in Fig. 1(a). First, they are largely MPI-only and difficult to extend, failing to cover the diverse communication libraries used by modern applications. Second, communication compression is datadependent: compression ratio and numerical fidelity vary across datasets, while accuracy requirements differ across applications. However, current benchmarks provide limited datasets and accuracy metrics, making it difficult to assess compression methods across workloads. Third, compression may contend with application computation during overlap, while existing benchmarks only report coarsegrained latency and throughput for isolated communication compression. This limits their ability to reveal end-to-end speedups and performance bottlenecks in real applications. These mismatches motivate lightweight, extensible, and application-oriented benchmarks that evaluate communication compression under realistic execution conditions without requiring full application deployment. Developing such a benchmark suite, however, introduces several challenges: 1) Heterogeneous communication backends across applications hinder unified cross-library evaluation. 2) Opaque communication compression implementations complicate fine-grained performance attribution. 3) Diverse application data characteristics, accuracy requirements, and execution patterns render fixed benchmark configurations insufficient for realistic end-to-end workloads. To address these challenges, we present CC-Bench1 , a lightweight, customizable, and application-oriented benchmark suite for evaluating communication compression across diverse communication backends. As shown in Fig. 1(b), CC-Bench enables systematic communication compression analysis through application-oriented configuration and runtime performance profiling, enabling realistic workload emulation during benchmark execution. In addition, CC-Bench incorporates communication datasets from representative applications, enabling flash cross-domain evaluation of communication compression. To the best of our knowledge, CC-Bench is the first open-source benchmark suite for comprehensive evaluation of communication compression. Our contributions are summarized as follows: • We introduce an application-oriented configuration module that decouples profiling logic from communication libraries, datasets, and error metrics through unified abstractions, supports user-defined integration of customized components, and enables declarative configuration to substantially simplify cross-application evaluation. • We develop a full-stack profiling module that combines function-level interception with hardware counter monitoring for systematic characterization of communication compression. The framework captures per-phase latency, records function traces and monitors hardware utilization under varying resource interference, exposes 1. https://anonymous.4open.science/status/CC-Bench

performance breakdowns and overlap-induced contention, and quantifies numerical fidelity across datasets. • We incorporate built-in representative datasets spanning scientific computing and LLM training/inference, enabling convenient evaluation of communication compression methods across diverse workloads. • We use CC-Bench to evaluate three compressionsupported communication libraries on CPU and GPU clusters across HPC and AI workloads. The evaluation characterizes accuracy, performance trade-offs, and overhead breakdowns under diverse computational conditions, providing practical deployment and optimization insights. These results show that CC-Bench bridges basic benchmark evaluation and full-application profiling for communication compression. The remainder of this paper is organized as follows. Sec. 2 reviews communication workloads, libraries, compression techniques, and benchmark gaps. Sec. 3 presents the fundamental goals and design principles of our benchmark. Sec. 4 illustrates the architecture, workflow and coverage of CC-Bench along with a usage demonstration. Sec. 5 presents evaluations of compression-supported communication libraries on both CPU and GPU HPC clusters, and Sec. 6 discusses future directions.

2. Background This section covers communication in HPC and AI workloads (Sec. 2.1), collective communication libraries (Sec. 2.2), compression techniques (Sec. 2.3), and the limitations of existing benchmarks (Sec. 2.4).

2.1. Communication for HPC and AI Workloads As application scales continue to grow, a single device is increasingly insufficient to sustain modern workloads. Distributed execution therefore relies on inter-device communication to coordinate computation and synchronize intermediate states across multiple devices [5]. Two mainstream classes of distributed workloads—scientific computing and AI—exhibit markedly different communication patterns. In distributed scientific computing, applications traditionally run on CPU clusters using process-based parallelism, where CPU-side collective primitives such as AllGather, AllReduce, and ReduceScatter are used to exchange intermediate results and global states [17]. These communication operations are typically coarse-grained, with each transfer ranging from hundreds of megabytes to tens of gigabytes [23]. Recently, HPC workloads have increasingly offloaded computation to GPUs, leading to heterogeneous CPU–GPU execution that requires runtime switching between CPU- and GPU-based communication paths [57]. AI workloads, including AI-for-Science, large-scale model training, and inference services, are primarily executed on GPU clusters and employ multiple parallelization strategies, such as data parallelism (DP), tensor parallelism (TP), and pipeline parallelism (PP) [52]. Different strategies induce distinct communication patterns. TP mainly relies on intra-node AllReduce to aggregate partial activations,

while PP uses point-to-point communication to transfer activations across pipeline stages on different nodes. These communications are often only tens of megabytes per operation, but can occur thousands of times within a single iteration. In contrast, DP primarily uses ReduceScatter and AllGather to synchronize gradients and model parameters, where individual transfers can reach several to tens of gigabytes but occur only a few times per iteration.

2.2. Collective Communication Libraries To support increasingly complex communication patterns, distributed applications rely on collective communication libraries that are tightly coupled with the underlying hardware platform. In high-performance scientific computing, where workloads primarily execute on CPU clusters, the Message Passing Interface (MPI) remains the de facto communication standard [17]. As scientific computing workloads are increasingly offloaded to GPUs, CUDA-aware MPI has become widely adopted to enable direct GPU memory access and efficient communication across heterogeneous CPU–GPU systems [57]. MPI provides a rich set of collective communication primitives, including AllReduce, ReduceScatter, AllGather, AlltoAll, and their vector variants, organized through the communicator abstraction. Modern MPI implementations, such as MVAPICH and Open MPI [18, 46, 49], map these primitives onto high-performance interconnects (e.g., InfiniBand, OmniPath, and Ethernet) using topology- and message-aware communication algorithms, while also supporting CUDA-aware extensions. Modern AI workloads, including AI for Science, LLM training, and inference, predominantly execute on GPU clusters and rely on GPU-specific collective communication libraries to exchange data efficiently across GPUs. Representative libraries include the NVIDIA Collective Communication Library (NCCL) [48] and the ROCm Communication Collectives Library (RCCL) [2]. NCCL provides optimized collective primitives for multi-GPU and multinode communication within NVIDIA GPU and networking ecosystems, delivering high-throughput and low-latency data movement over intra-node NVLink/PCIe interconnects and inter-node InfiniBand (IB)/RoCE fabrics. RCCL extends the NCCL design to AMD GPU ecosystems.

2.3. Communication Compression Communication compression accelerates distributed data exchange by reducing the volume of transmitted data. It is commonly categorized into lossless and lossy approaches, whose effectiveness is inherently data-dependent: achievable compression ratio and throughput vary with data distribution, while lossy methods additionally introduce accuracy trade-offs [55]. Lossless compression (e.g., DietGPU [32], LZ4 [13], and MANS [27]) preserves numerical fidelity but typically provides limited compression ratios and throughput, making it suitable for precision-sensitive applications under network bandwidth constraints [32]. In contrast, lossy compression (e.g., SZ [36], SDP4Bit [30], and PRISM [41]) trades numerical precision for substantially

TABLE 1: Comparison of existing benchmarks with CC-Bench. Performance Profiling: coarse- and fine-grained metrics; Flexibility and Extensibility: communication backends, datasets, and quality metrics switching and integration; Workload Characteristics: realworld datasets and computation interference. Benchmark

Performance Profiling

Flexibility

Extensibility

Workload Characteristics

Vendor (N/R/I) CommBench OMB OMB-Compr

⃝ ⃝ ⃝ ⃝

× ✓ ✓ ⃝

× × × ×

× ⃝ × ⃝

CC-Bench (Ours)

✓: Achieved;

×: Not supported;

⃝: Partial

higher compression ratios and throughput, and is therefore widely adopted in error-tolerant applications and highspeed network environments [30]. Prior studies have shown that lossy communication compression can significantly improve distributed application performance with negligible accuracy degradation [3]. To facilitate practical deployment, communication compression has been integrated into collective communication libraries for scientific computing and AI workloads. In scientific computing, ZCCL [26] incorporates the SZ lossy compressor into MPI and co-designs compression and communication pipelines, while gZCCL [25] extends the design to GPUs. For AI workloads, UCCL-ZIP [42] integrates the DietGPU lossless compressor [32] into UCCL [43] via pipelining and operator fusion; COCCL [40] supports customizable compression operators on NCCL; NCCLZ [58] combines quantization with lossless entropy coding; and ZipCCL [37] optimizes compression for Mixture-of-Experts (MoE) training. These libraries primarily target either scientific computing or AI training, with limited support for AIfor-Science applications. We benchmark them systematically in Sec. 5.

2.4. Limitations of Existing Benchmarks Several benchmarks partially address the evaluation of compression-aware collective communication, as summarized in Tab. 1. Vendor benchmarks, including NVIDIA NCCL Tests [47], AMD RCCL Perftest [1], and Intel MPI Benchmarks [28], measure backend-specific collective latency and bandwidth.CommBench [22] benchmarks realistic communication behavior across MPI, NCCL, RCCL, and OneCCL on hierarchical HPC systems. OSU MicroBenchmarks (OMB) [19] provide portable MPI communication benchmarks, while OMB-Compr [6] extends OMB with ZFP-based compressed collectives and reports overall latency and reconstruction error. Despite their contributions, existing benchmarks share several limitations in evaluating compression-aware communication: Limited Extensibility. Existing benchmarks provide limited flexibility in switching communication libraries, datasets, and error metrics. Most either support only a single communication library or require recompilation to switch among supported libraries, while offering no mechanism for integrating new libraries. Moreover, since compression introduces data dependencies, benchmarking must account for diverse datasets and corresponding error

metrics; however, prior work lacks support for customizable dataset and metric evaluation. Incomplete Performance Analysis. Existing benchmarks provide only coarse-grained latency across message sizes, lacking fine-grained profiling of communication/compression overhead breakdown, domain-specific metrics and hardware utilization. This limits the ability to identify performance bottlenecks in communication compression, resulting in a gap between benchmarked and end-to-end application performance. Lack of Datasets. Existing benchmarks largely overlook diverse real-world application datasets despite their importance and limited availability. This restricts the evaluation of communication compression across heterogeneous data characteristics. Consequently, users must manually curate application-specific datasets for realistic benchmarking, substantially increasing evaluation cost.

bounded by computation, memory bandwidth, or interconnect saturation, users can estimate performance bottlenecks when phase decomposition is unavailable. Realistic Workload Conditioning. We inject real-world datasets (e.g., scientific datasets and LLM communication traces) and computation interference profiles into the benchmark. This dimension targets operational realism: it quantifies how compression behaves under conditions resembling production deployments, including interference from co-located workloads, unlike micro-benchmark settings where such effects are invisible. Together, these four dimensions provide a systematic evaluation methodology—from macro-level stack comparison down to micro-level resource attribution—that captures the multifaceted performance profile of communication compression in a single, configurable benchmark suite.

3. Methodology of CC-Bench

3.2. Design Principles

We present our benchmark goals (Sec. 3.1) and design principles (Sec. 3.2).

3.1. Benchmark Goals CC-Bench bridges the gap left by prior work (Tab. 1) as a practical, diagnostic tool for compression performance in data-dependent, application-relevant settings. It answers three questions: (1) What latency–throughput–accuracy trade-offs arise across compressor–backend combinations, and how do they vary with workload characteristics? (2) Where do bottlenecks reside once compression is deployed—in codec throughput, network bandwidth, or interference-induced contention? (3) Do conclusions from isolated micro-benchmarks hold under realistic conditions with real data, computation interference, and applicationspecific accuracy requirements? We structure the evaluation along four dimensions, each tied to one of these goals. Cross-stack Sensitivity. We sweep a matrix of compressors and communication backends, measuring latency, throughput, and error metrics such as Mean Absolute Error (MAE). This dimension targets design-space coverage: by exposing performance and accuracy variations across combinations, it helps users identify deviation-performance trade-offs and optimal stack for a given workload class without manual benchmarking. Fine-grained Phase Decomposition. We decompose each collective call into compress, send, recv, and decompress stages using lightweight runtime instrumentation. This dimension targets bottleneck localization: by revealing whether the dominant cost lies in compression throughput, network transfer, codec or insufficient overlap. It guides optimization efforts to the right component. Hardware-level Introspection. To address intractable conditions when function decomposition is unavailable, we enable collection of hardware utilization timelines (GPU SM/DRAM, NIC bandwidth, PCIe throughput) synchronized with operation call timeline. This dimension targets resource contention diagnosis: by exposing whether throughput is

In order to achieve our goals for evaluation, CC-Bench is motivated to be lightweight, application-oriented and heterogeneously flexible. It is designed around five principles that directly address the limitations identified in Sec. 2.4. Extensible and Wide-coverage Design. CC-Bench addresses the limited backend and compressor coverage of existing benchmarks through a unified abstraction layer that decouples communication, compression, and profiling into independently swappable components. Users can integrate any compressor into any communication backend via uniform interfaces. The suite also provides a configurable dataset loader and a lightweight LD_PRELOAD-based injection system that requires no recompilation, enabling users to easily sweep across backends, compressors, datasets, and runtime parameters with minimal effort. Full-stack Bottleneck Decomposition. To move beyond coarse-grained latency and throughput, CC-Bench exposes where time is actually spent. On the software side, wrapper-bassed profiling isolates compress, decompress, send, recv, and overlap overhead. On the hardware side, daemons concurrently capture utilization metrics (CPU, GPU, memory, NIC), revealing whether the bottleneck lies in computation, memory, or communication. This dual-layer approach diagnoses contention and overlap inefficiencies that aggregate metrics alone cannot expose, directly addressing the limitations of isolated communication benchmarks. Data-dependent Accuracy Assessment. Recognizing compression is inherently data-dependent, CC-Bench provides a pluggable metric framework supporting domain-specific criteria. To facilitate evaluation across diverse data modalities, the suite ships with representative datasets — including scientific fields and prevalent LLM traces — alongside domain-appropriate metrics such as PSNR for climate fields and SSIM or error quantiles for LLM tensors. This enables systematic characterization of how data modality and accuracy requirements reshape compression-quality trade-offs.

Deviation Calculator Test Binary (Allreduce.c, overlap.c, ...)

Compression Backends (SZ3/ZFP/SDP4 Bit)

Data loader

Data

Sec 4.2

Sec 4.3

RUN ENV

Perf Wrapper Layer

Deviation Metrics (PSNR/SSIM)

Data Analyzer

Collective Communication Backends (MPI/XCCL)

CCL Wrapper Layer

Target Perf Funcs

libperf.so Compression Wrapper Layer

Stress Daemon

Real-world datasets Hurricane/ NYX/ KV grads

wrapper functions

Perf Daemon

Abstract Device Interface(PCIE,RDMA,RNIC,etc.) Hardware Layer (CPUs, GPUs, NICs etc.) User Config

Application-Oriented Configuration Module

Profiling Module

Figure 2: CC-Bench architecture. The three swappable preload layers and daemons form the runtime for the test binaries.

4. Design of CC-Bench In this section, we first give an overview of our benchmark framework(Sec. 4.1). Then, we present the details of each component in our benchmark framework and supported plugins(Sec. 4.2–4.4). Finally, we provide some examples of how to use our benchmark(Sec. 4.5).

4.1. Overview Architecture. Fig. 2 illustrates CC-Bench’s architecture, which consists of an application-oriented configuration module and a full-stack profiling module. The former uses communication and compression wrappers, a configurable dataset loader with pluggable deviation metrics, and a stress daemon to emulate realistic application environments. The latter adds a function-profiling wrapper and performance daemons whose hardware samples and traces are postprocessed by the data analyzer. The benchmark extends to more nodes simply by editing the host list. Workflow. Fig. 3 exhibits CC-Bench’s five-stage (Config, Customize, Build, Run, Analyze) workflow. ① The Configuration stage centralizes all parameters in modular JSONC files. ② Users provide wrapper code for libraries with inconsistent interfaces, forming the three wrapper layers. ③ A single build script parses configurations, compiles wrappers, links backends, and emits a scheduler-aware run script. ④ The run phase activates daemons and executes in two phases for precise latency and breakdown measurement. ⑤ Analysis instruments compress/send/recv/decompress stages with hardware monitoring, from CSV metrics to overlap-efficiency and hardware-utilization heatmaps.

4.2. Application-Oriented Configuration Module Communication Wrapper. When target communication backends provide non-standard entry functions, this layer intercepts standard collective APIs (e.g., MPI_AllReduce, nccl_AllReduce) and reimplements them using compressed send/recv primitives or delegates to compression-aware libraries (e.g. ZCCL, UCCL-Zip) as shown in Fig. 6(a), where we provide an instance of

integrating ZCCL functions into standard MPI collectives. It maintains collective semantics while operating on compressed data, remaining compressor-agnostic. Compression Wrapper. Similar to the previous layer, this wrapper allows weak-symbol defaults for compression entry points such as compress/decompress to be overridden by user-supplied codec shared libraries. This design enables users to intuitively substitute compression for certain backends, even when the backend has a built-in or hardcoded compressor. Our bench currently ships with wrappers for SZ3, ZFP for ZCCL backend and DietGPU, SDP4Bit wrappers for COCCL collectives. Dynamic Deviation Calculation. A dynamic deviation metric calculator ensures wide coverage across usage fields. As illustrated in Fig. 4, users can implement custom metric functions (e.g., SSIM or domain-specific error measures) following a predefined template (see Fig. 6(c)) and register them via a script run which automatically compiles them into a shared library. The framework then dynamically loads this library at runtime, seamlessly injecting the user-defined logic into the evaluation workflow without modifying core code. This design enables flexible, extensible metric integration across diverse application domains. Stress Daemons. In real-world HPC and AI clusters, background interference rarely manifests as sustained, full-saturation load. Instead, it typically arises from colocated jobs, system daemons, periodic synchronization routines, or network background traffic–all of which exhibit intermittent, bursty patterns with alternating active and idle phases. As illustrated in Fig. 5, our stress daemons emulate such realistic contention through a configurable duty-cycle mechanism: they repeatedly execute lightweight compute loops during busy phases and remain idle during rest phases, with the busy/idle ratio specified by the user. We deliberately approximate interference with stress daemons rather than embedding full application stacks, since the latter would conflict with CC-Bench’s lightweight, portable design.

4.3. Full-stack Profiling Module Profiling Wrapper. The final layer of interception before function call. By bracketing the target function, users can trace per-call timestamps, message sizes, and compression ratios by utilizing the perf notedown helpers provided by our bench as shown in Fig. 6(b). Activated only during decomposition phases, it provides traces which can further reveal macro-level information including overlap . Auxiliary Tests. Apart from collective operations tests, The pingpong test records point-to-point latency/bandwidth across both intra-node and inter-node configurations. Given the linear or near-linear relationship between message size and latency, pingpong results enable interpolation-based estimation of non-blocking communication overhead in downstream tests. The overlap test measures computationcommunication interference by dispatching compute and communication kernels concurrently on two GPU streams (or CPU threads) and comparing against serial execution. While the overlap test does not directly feed into other tests,

User Space Userconfig /*.jsonc

Build System

Wrapper Codes User Libs

① Environment Setup & Custom Code

Config Parser

Integrated Libraries Compressors, CCLs

Build Engine Compilation & Link

Wrappers (.so)

Run Script Generator

Run Scripts local/cluster

User Choices Environment variables ② JSONC Configuration

③ Automated Build & Generation

Run Environment

Data & Results

Phase 1 (Baseline) output value test LD_PRELOAD + execution comm. comp. raw bw/latency

Output data

Phase 2 (Decomposition) LD_PRELOAD comm. comp. perf

hardware data test + execution operation perf

④ Two-Phase Execution & Performance Collection

Performance analysis Deviation analysis & heat map

overlap

⑤ Analysis

Figure 3: CC-Bench workflow: ① Users configure benchmarks through JSONC files and ② provide custom wrapper code. ③ The build stage compiles all wrappers and generates execution scripts; ④ the two-phase run collects performance and profiling data; ⑤ the analysis stage produces multi-metric reports. psnr.c ssim.c

… User Register Command

Metrics List C file Compile & Update

Dynamic injection

Validation.so

0

Data Analyzer deviation

deviation metrics std_output test_output

100ms 200ms 300ms 400ms 500ms 600ms 700ms 800ms 900ms 1000ms

NCCL AllReduce Pipeline

… time

CPU Stress

GPU Stress

time …

Test Binaries

time

Figure 4: A user-friendly, lightweight design for dynamically plugging deviation metrics into the evaluation workflow.

Figure 5: Overview of the stress daemon working process, intermittently spinning for a given portion of loops in the background.

it provides critical insight into how seemingly independent workloads interfere with each other when co-located. Perf Daemons. Poll hardware counters (CPU, GPU, InfiniBand, PCIe etc.) and log timestamped samples. Similar to the deviation metrics system, they are plugged into the framework through a unified registration and dynamic loading mechanism, enabling easy extension. Sampling runs at a low frequency in a separate process and thus imposes negligible overhead on the measured communication path. Analysis Layer. Consumes profiling wrapper traces and hardware samples to produce multi-dimensional performance views. Three statistical tools serve different granularities: draw_hardware plots hardware-counter heatmaps from daemon CSVs and supports time-range filtering, per-metric selection, and node isolation.draw_perf_function draws per-rank function-trace heatmaps from function traces of a specific rank with the chosen metrics (e.g. duration, compression ratio) and an aggregate mode. function_analyzer classifies calls into communication vs. compression, estimates async durations via pingpong interpolation, and reports overlap efficiency and idle time per phase. Together they trace a bottleneck from macro utilization down to per-call overlap. Overall, these components expose overlap efficiency, per-function latency, hardware utilization, and their inter-dependencies.

and GPU compressors such as SZ3 [36], ZFP [16, 38], DietGPU [32], SDP4Bit [30], TAH-Quant [21], and additional COCCL plugins, covering error-bounded, fixed-rate, lossless, and quantization-based compression paths. In terms of communication backends, we compare uncompressed MPI/NCCL baselines against compression-aware libraries including ZCCL [25], COCCL [40], and UCCL-Zip [42] under a unified benchmark interface. Our data modality is highlighted: we capture data-dependent compression behavior across scientific datasets (Hurricane ISABEL [45], NYX [4], CESM-ATM [62]), LLM KV-cache traces, Llama3 8B activation values, weights and gradients, spanning scientific computing, LLM inference/training, and controlled synthetic distributions. Finally, our variety of evaluation metrics link performance, fidelity, and profiling breakdowns to workloadspecific constraints through latency, throughput, compression quality metrics, hardware traces, phase breakdowns, and configurable rank-offset schemes.

4.4. Benchmark Coverage CC-Bench covers the main degrees of freedom needed to evaluate compression-enhanced collective communication. Our supported collective operations include Broadcast, Reduce, AllReduce, Scatter, Gather, AllGather, AlltoAll, ReduceScatter and vector collectives, enabling measurement of collective latency/bandwidth, point-to-point baselines, and slowdown under communication-computation overlap. Our currently integrated compression wrappers span CPU

4.5. Benchmark Usage Fig. 3 shows the overall CC-Bench workflow: users configure benchmark inputs, customize optional extensions, build the run environment, execute tests, and analyze outputs. Fig. 6 complements this workflow by giving concrete examples of user-facing code formats and commands. Setup. Setup corresponds to the config, customize, and build stages in Fig. 3. Users flexibly select the execution environment, communication library, compressor, dataset, message sizes, background load, and accuracy metrics through modular JSONC configuration files. If a backend, profiler, daemon, or metric is not built in, users provide the lightweight extension formats shown in Fig. 6(a)–(c). The build command in Fig. 6(d) then compiles the selected extensions and emits a self-contained run script. Run & Analysis. Run and analysis corresponds to the final two stages in Fig. 3. Users execute the generated script

(a) Comm. wrapper 1 int MPI_Allreduce(sbuf, rbuf, count, dtype, op, comm) 2 { 3 // Run custom compression-accelerated version 4 int ret = MPI_Allreduce_ZCCL_RI2_mt_oa_record(...); 5 if (ret != MPI_SUCCESS) // fallback 6 ret = PMPI_Allreduce(sbuf, rbuf, count, dtype, op, comm); 7 return ret; 8 } (b) Perf. wrapper 10 int MPI_Allreduce(sbuf, rbuf, count, dtype, op, comm) 11 { 12 perf_notedown(); // timestamp before 13 int ret = MPI_Allreduce(sbuf, rbuf, count, dtype, op, comm); 14 perf_notedown(); // timestamp after 15 return ret; 16 } (c) Deviation Code 18 int compute_mse(const void *buf1, const void *buf2, 19 MPI_Aint count, MPI_Datatype dtype) { 20 // Compute element-wise squared error 21 // ... 22 return result; Implementation 23 } (d) Build command $ ./scripts/build_script.sh [--rebuild-bench] (e) Run command $ # Local $ bash scripts/run/run_benchmark.sh $ # Interactive $ salloc -N 2 -n 8 && bash scripts/run/run_benchmark.srun.sh $ # Batch $ sbatch scripts/run/run_benchmark.slurm

datasets, LLM workload tensors, and synthetic distributions. Scientific inputs include Hurricane ISABEL atmospheric fields, NYX cosmology fields, and CESM-ATM climate fields, covering 2D/3D arrays, multi-field simulations, and different entropy characteristics [62]. LLM inputs include simulated KV-cache traces for inference and Llama 3 8B activations, gradients, and weights for training-oriented communication. Quality Metrics. For an original vector x = {xi }ni=1 and reconstructed vector x̂ = {x̂i }ni=1 , CC-Bench reports pointwise and application-oriented quality metrics: 1) Mean P Absolute Error (MAE) = 1n ni=1 |xi − x̂i |, 2) Mean Squared P Error (MSE) = 1n ni=1 (xi− x̂i )2 , 3) Peak Signal-to-Noise 2 Ratio (PSNR) = 10 log10 max_val , 4) Cosine Similarity MSE

P

(CosSim) = qP

xi x̂i iq

P 2 , 5) Relative Error (RE) = i x̂i sizecompressed 6) Compression Rate (CR) = size . original 2 i xi

∥x−x̂∥2 , ∥x∥2

Performance and profiling results include latency(L), effective bandwidth(BW), slowdown(S) = Lload /Lbaseline , phase duration, SM and memory bandwidth utilization. Deployment

Figure 6: Example codes for deploying CC-Bench. Blue denotes custom C code templates, while gray denotes bash commands for setup and run.

as in Fig. 6(e), producing latency, throughput, deviation metrics, per-rank traces, and hardware logs. Post-processing then reports phase breakdowns, overlap efficiency, and visualizations; hardware-utilization heatmaps provide a complementary bottleneck view when function-level interception is unavailable.

5. Evaluation We evaluate CC-Bench as a diagnostic benchmark for compression-enhanced collective communication. Using AllReduce as the representative primitive, the evaluation process aims to answer the three questions crucial to resolving current limitations (see Sec. 3.1).

5.1. Basic Benchmark Setup Platform. We evaluate CPU benchmarks on a cluster with two Intel Xeon Platinum 8358P CPUs (32 cores each) per node, using Open MPI 4.1.4 and 200G InfiniBand, with 16 tasks per node across 16 nodes. For GPU benchmarks, we use a GPU cluster with two AMD EPYC 7402 CPUs (24 cores each) and eight NVIDIA A800 SXM4 80 GB GPUs per node, with CUDA 12.8, NCCL 2.27, and 8 tasks per node across 2 nodes. Benchmark Library. We adopt plain MPI and NCCL as uncompressed baselines. Other setups including ZCCL+SZx, ZCCL+SZ3, UCCL+DietGPU, COCCL+SDP4Bit and COCCL+TAH-Quant stand for diverse compressionaccelerated communication schemes. Datasets. To capture the data-dependent behavior of communication compression, CC-Bench uses scientific

5.2. Quality Evaluation CC-Bench treats compression quality as workloadspecific instead of a unified error metric. We present fidelity and compression ratios for typical CPU scientific datasets and GPU LLM workloads. CPU tests adopt PSNR to measure reconstruction quality across Hurricane ISABEL, CESMATM and NYX. GPU experiments leverage cosine similarity and relative error for Llama 3 8B activations, gradients, weights and KV-cache traces. Tab. 2(a) shows that CPU compression quality is strongly dataset-dependent. For Hurricane ISABEL, ZCCL+SZx provides slightly higher PSNR than ZCCL+SZ3 (66.60 versus 66.31 dB) while also producing a lower compression rate (0.104 versus 0.117). For CESM-ATM, the qualityrate tradeoff reverses: ZCCL+SZ3 reaches 75.07 dB, but its compression rate is 1.033, whereas ZCCL+SZx reports 56.19 dB at a much lower rate of 0.104. NYX is a harder TABLE 2: Quality metrics for CPU and GPU backend-compressor configurations. CPU rows report PSNR (dB, higher is better) and compression rate; GPU rows report cosine similarity(CosSim), relative error(RE), and compression rate(CR). (a) CPU backend-compressor Backend

Metric

Hurricane

CESM

NYX

ZCCL+SZx

PSNR CR

66.6 0.104

56.2 0.104

-127.3 1.026

ZCCL+SZ3

PSNR CR

66.3 0.117

75.1 1.033

-12.7 –

(b) GPU backend-compressor Backend

Metric

Act.

Grad.

Weight

KV

COCCL+SDP4Bit

CosSim RE CR

0.992 86.710 0.133

0.992 0.669 0.133

0.999 0.903 0.133

0.991 3.85e4 0.133

COCCL+TAH-Quant

CosSim RE CR

0.414 6.81e5 0.500

0.452 19.1 0.500

0.410 8.39 0.500

0.427 6.96e6 0.500

UCCL+DietGPU

CosSim RE CR

1.000 0 0.846

1.000 0 0.851

1.000 0 0.844

1.000 0 0.855

Figure 7: CPU cross backend-compressor performance on scientific datasets. Latency incorporates overhead from all stages including codec and communication.

Figure 8: GPU cross backend-compressor performance on LLM workloads. Latency incorporates overhead from all stages including codec, transfer and communication.

case for the tested compressors, with negative PSNR values and an SZx compression rate above one. Insight 1.1: A lower compression rate does not necessarily mean that performance will be better. The same holds on the GPU platfrom as observed in Tab. 2(b). The lossless UCCL-DietGPU achieves full accuracy with a lower compression ratio and uncompetitive thoughput,while interestingly SDP4bit significantly outperforms TAH-Quant in both cosine similarity and relative error while preserving a higher compression ratio of over 7.5. It can be also observed that COCCL+SDP4bit generally yield high relative errors, while cosine similarities remain promising, indicating that the reconstructed values are close to zero. These results clearly show that a compressor selected by a single dataset or metric can be potentially misleading.

Insight 1.2: Failure on a single parameter does not necessarily indicate poor performance.

5.3. Performance Evaluation CC-Bench evaluates performance at three levels: overall latency for complete backend-compressor stacks(Sec. 5.3.1), phase-level breakdown for bottleneck localization(Sec. 5.3.2), and slowdown under background computation for realistic resource contention(Sec. 5.3.3). This organization separates whether a compressed collective is faster, why it is faster or slower, and whether that conclusion remains valid when application kernels share the same hardware resources. 5.3.1. Cross-combination Performance. This combinatorial experiment incorporates multiple backend-compressor stacks across two platforms, measuring each as a deployable communication path including codec cost, transfer time, and reconstruction error. On the CPU cluster, we compare

Default

============================================================== Function Analysis Report -- rank 0 ==============================================================

Comm

Comp

Decomp Comm

35657m𝒔

(union)

------------ Per-function duration breakdown ---------------------------------------------------------------Function Count Total(us) Avg(us) -------------------------------------------------------- Communication -MPI_Irecv 19635 470185.161 23.946 MPI_Isend 19635 16357.855 0.833 -- Compression -ZCCL_compress_mt 19712 19706789.000 999.736 ZCCL_decompress_mt 39347 15912163.000 404.406

(a) Output of CC-Bench

Decomp

( 1.36%) ( 99.89%) ( 1.26%)

36105m𝒔

Timeline (us): Total wall-clock: 94795526.710 (idle excluded from %) Communication time: 486543.016 Compression time: 35618375.121 Overlap (comm+comp): 447624.136 Verification: comm + comp - overlap = 35657294.002 (union = effective time = 100%)

Overlap

Comp

(b) Overlap Breakdown

Figure 9: The output of CC-Bench and CPU overlap breakdown for communication and compression phases.

uncompressed MPI with ZCCL+SZx and ZCCL+SZ3 over NYX, Hurricane ISABEL, and CESM-ATM. On the GPU cluster, we compare uncompressed NCCL with UCCL+DietGPU, COCCL+SDP4Bit, and COCCL+TAH-Quant on KV-cache traces and Llama 3 8B activations, gradients, and weights. We record and observe the trade-offs between aggregate latencies and MAE of each combination. Fig. 7 reports CPU latency and MAE across the three scientific datasets. Compression is not a faster replacement for raw communication: MPI remains the lowest-latency choice in tested sizes, while compressed collectives add codec overhead that varies by dataset and error bound. For example, on Hurricane ISABEL, the best compressed latency at 64 MB with ZCCL+SZx (CR=10−3 ) is 210,698.25 µs, compared with 161,494.36 µs for uncompressed MPI. Fig. 8 shows the corresponding GPU comparison. The LLM results are interpreted according to the natural message-size regimes of different tensor types, following distinct message-size regimes: gradients and weights which represent data-parallel synchronization traffic are relatively large (≥256MB), while KV-cache and activations which communicate in inference or tensor parallelism are small to medium (≤512MB). We therefore sweep sizes accordingly. For activation and KV-cache messages, COCCL quantization can sharply reduce latency at selected small sizes. At 8 KB, the best COCCL path achieves 36.36–40.33× speedup, and at 512 KB it achieves 21.34–21.45× speedup. The benefit is highly non-monotonic, however. At 8 MB, both COCCL paths exhibit a latency spike of 0.89–0.91 s while NCCL remains around 0.48 ms, indicating a size-specific overhead in the compressed path. For larger activation and KVcache messages within the practical range, COCCL+SDP4Bit becomes competitive again, reaching 7.45× speedup on 512 MB activations and 5.98× and 3.60× speedups on 256 MB and 512 MB KV-cache messages, respectively. For gradients and weights, the large-message regime shows a more conservative but more deployment-relevant trend. On gradients, COCCL+SDP4Bit is slower than NCCL at the 256 MB boundary and remains close to parity at 768 MB, but becomes faster at 1.00 GB and 1.25 GB, with 1.18× and 1.41× speedups. On weights, COCCL+SDP4Bit outperforms NCCL at 512 MB, 768 MB, and 1.00 GB, with

1.52×, 1.36×, and 1.92× speedups, respectively, although it is slower at 256 MB and 1.25 GB. Across all workloads, UCCL+DietGPU preserves zero MAE but is usually slower than NCCL, especially for medium and large messages, suggesting that lossless compression overhead and backend compatibility dominate the saved transfer time on this platform. Among the two lossy COCCL paths, SDP4Bit consistently has lower MAE than TAH-Quant, with MAE around 1.04 × 10−2 on gradients, 5.06 × 10−4 –5.37 × 10−4 on weights, up to 3.34 × 10−2 on activations, and 1.63– 2.64 on KV-cache messages of at least 8 MB. TAH-Quant is occasionally faster at some tiny-message points, but it consistently incurs larger reconstruction error. Insight 2.1: The selection of a compressor or backend should only be made once the quality gain, data reduction, execution overhead and compatibility with hardware have been considered together. 5.3.2. Performance Breakdown. The performancebreakdown experiment explains where total time is spent. Here we demonstrate the per-operation breakdown for ZCCL+SZx, since CC-Bench intercepts its communication functions such as MPI_Isend and codec functions such as SZx calls with minimal effort using its customized profiling wrappers to. The post-run analyzer then separates compression, decompression, send/recv, wrapper overhead, and overlap time, as reported in Fig. 9. On the CPU platform, the overlap breakdown shows that communication is mostly hidden by the codec workflow rather than appearing as exposed transfer time. The profiled communication time is 486,543 µs, but only 38,919 µs remains after overlap analysis. In particular, 447,289 µs overlaps with compression and 335 µs overlaps with decompression, meaning that approximately 92.0% of communication time is concurrent with codec work. The exposed timeline is therefore dominated by compression-only and decompression-only regions, which explains why optimizing only the communication backend or overlap mechanism would have limited impact for this configuration. At a finer granularity, the analysis output also reveals notable asymmetries existing between operations of the same category. While compression incurs higher per-call latency, its lower invocation count (19,712 vs. 39,347 for decompress) results in comparable total execution time between the two phases. This suggests that the communication scheduler may benefit from prioritizing overlap of the more frequent decompression operations with ‘Irecv‘, rather than treating compression and decompression symmetrically. Insight 2.2: Codecs dominate the ZCCL timeline, with communication largely hidden behind its operations – decompression by volume, compression by weight. It is also worth noting that for many GPU backends such as UCCL-zip, whose compression and communication entry points are less directly separable , CC-Bench can resort to reading hardware counters including SM utiliza-

Figure 10: Slowdown of COCCL+SDP4Bit on KV-cache. Slowdown (%) denotes latency increase relative to the no-load baseline; percentages indicate injected background load intensity.

tion and network-device utilization over successive time windows. These time-resolved hardware states also provide complementary signals for locating bottlenecks through exposing overlap and underutilization of components. 5.3.3. Performance Slowdown. The slowdown experiment evaluates whether clean-environment performance holds when communication compression shares GPU resources with application computation. We use CCBench to inject controlled background load, evaluate COCCL+SDP4Bit on KV-cache traces, sweep GPU load from 0% to 70%, and record the resulting latency degradation. As shown in Fig. 10, slowdown generally increases as GPU background load becomes heavier, but the effect is also message-size dependent. The largest-message and high-load cases consistently show substantial latency degradation, while very small messages and several intermediate-size points are non-monotonic. In particular, low background load can occasionally reduce measured latency, reflecting launch overhead, scheduling noise, and the sensitivity of short collectives to runtime variance. These results show that compressed-communication latency measured in isolation can overestimate practical speedup when compression kernels contend with application kernels. Insight 2.3: Background computation can erode compression-enhanced communication throughput.

5.4. Insight Discussion Communication compression should be evaluated as a workload-conditioned communication path, not as an isolated codec property. The evaluation directly addresses the first question in Sec. 3.1 on how performance-accuracy trade-offs vary across stacks and data. As observed in Insight 1 and Insight 2.1, higher compression ratios improve communication efficiency at the cost of increased encoding overhead and potential fidelity loss. However, tradeoff trends shift substantially with workload characteristics—low-entropy HPC data tend to be compression-ratio-driven, while high-entropy training workloads are often accuracy-driven. Compression is therefore a workload-conditioned optimization, not a universal replacement; users should apply CC-Bench to full backend-compressor-data combinations rather than any single dimension in isolation. Performance bottleneck location requires explaining why a stack is fast or slow. The finer the decomposition of the overall operation, the more precisely

users can locate what throttles macro efficiency. Insight 2.2 reveals through phase decomposition that the bottleneck for ZCCL is codec work, not communication itself. Following this logic, CC-Bench provides profiling coverage across the entire granularity spectrum—from coarse module-level (communication vs. compression) down to function-level (compress, decompress, send, recv), and further to functioninternal parameters such as compression ratio or per-call latency. This coarse-to-fine breakdown lets users narrow the bottleneck from communication-versus-computation down to the specific operation and its internal characteristics. CC-Bench thus bridges clean micro-benchmarks and fullapplication profiling. CC-Bench also reflects end-to-end behavior by preserving context around communication. Full applications do not execute communication compression in isolation: LLM training and inference overlap collectives, compression kernels, and application kernels on the same GPU resources. Insight 2.3 demonstrates that this interference is non-negligible by showing that shared GPU computation interference can substantially reduce compression-enhanced communication performance. CC-Bench approximates this setting by preserving the full communication path and injecting configurable GPU background load—for instance, 30% to simulate a contended inference server, or 70% to model training with colocated computation.

6. Conclusion and Future Work This paper presents CC-Bench, a lightweight benchmark suite for compression-enhanced collective communication. Its core contribution is twofold: a flexible and easy-touse configuration workflow for execution environments, communication libraries, compressors, accuracy metrics, and real application datasets; and a profiling path that breaks benchmark runs into software phases and hardwareutilization evidence. Our evaluation shows that latency, throughput, and accuracy must be considered jointly, that bottlenecks shift with overlap and interference, and that isolated microbenchmark conclusions do not always hold under realistic conditions. Future work extends CC-Bench along three directions. First, error propagation profiling tracks compression error variation across iterative compression-decompression steps in collectives. Second, the plugin framework integrates new compressors, hardware backends and I/O/power monitoring daemons. Third, lightweight proxies replay collective traces in mini-apps and training loops, bridging benchmark tests and practical application behaviors.

Acknowledgments This work was supported by the National Key Research and Development Program of China (Grant No. 2025YFB3003702), the National Natural Science Foundation of China (Grant No. T2125013), and the Innovation Funding of ICT, CAS (Grant No. E461050). The experiments were performed on the robotic AIScientist platform of Chinese Academy of Sciences.

References [1]

Advanced Micro Devices, Inc., “RCCL Tests: Benchmark executables for RCCL,” https://github.com/ROCm/rccl-tests, 2026.

[2]

Advanced Micro Devices, Inc., “Rccl: Rocm communication collectives library,” https://github.com/ROCm/rccl, 2019.

[3]

D. Alistarh, D. Grubic, J. Li, R. Tomioka, and M. Vojnovic, “Qsgd: Communication-efficient SGD via gradient quantization and encoding,” in Advances in Neural Information Processing Systems, I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, Eds., vol. 30. Curran Associates, Inc., 2017.

[4]

A. S. Almgren, J. B. Bell, M. J. Lijewski, Z. Lukić, and E. Van Andel, “Nyx: A massively parallel amr code for computational cosmology,” The Astrophysical Journal, vol. 765, no. 1, p. 39, feb 2013. [Online]. Available: https://doi.org/10.1088/0004-637X/765/1/39

[5]

T. Ben-Nun and T. Hoefler, “Demystifying parallel and distributed deep learning: An in-depth concurrency analysis,” ACM Computing Surveys (CSUR), vol. 52, no. 4, pp. 1–43, 2019.

[6]

M. A. Bender, S. A. Weil, N. Dryden, F. Pruvost, and W. D. Jones, “OMB-Compr: An extension to OSU micro benchmarks for collective compression error measurement,” in Proceedings of the Practice and Experience in Advanced Research Computing 2025: Computational Science and Engineering Innovation in the Age of AI, 2025.

[7]

T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al., “Language models are few-shot learners,” Advances in neural information processing systems, vol. 33, pp. 1877–1901, 2020.

[8]

Z. Cai, Z. Liu, S. Maleki, M. Musuvathi, T. Mytkowicz, J. Nelson, and O. Saarikivi, “Synthesizing optimal collective algorithms,” in Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming, 2021, pp. 62–75.

[9]

X. Cao, T. Başar, S. Diggavi, Y. C. Eldar, K. B. Letaief, H. V. Poor, and J. Zhang, “Communication-efficient distributed learning: An overview,” IEEE Journal on Selected Areas in Communications, vol. 41, no. 4, pp. 851–873, 2023.

[10] C. Chen, X. Li, Q. Zhu, J. Duan, P. Sun, X. Zhang, and C. Yang, “Centauri: Enabling efficient scheduling for communication-computation overlap in large model training via communication partitioning,” in Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3, 2024, pp. 178–191. [11] S. Cheng, S. Lin, L. Diao, H. Wu, S. Wang, C. Si, Z. Liu, X. Zhao, J. Du, W. Lin et al., “Concerto: Automatic communication optimization and scheduling for large-scale deep learning,” in Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1, 2025, pp. 198–213. [12] A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann et al., “Palm: Scaling language modeling with pathways,” Journal of Machine Learning Research, vol. 24, no. 240, pp. 1–113, 2023. [13] Y. Collet, “Lz4: Extremely fast compression algorithm,” GitHub repository, 2011. [14] M. Cowan, S. Maleki, M. Musuvathi, O. Saarikivi, and Y. Xiong, “Mscclang: Microsoft collective communication language,” in Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2, 2023, pp. 502–514.

[17] M. P. I. Forum, “Mpi: A message passing interface,” in Supercomputing ’93:Proceedings of the 1993 ACM/IEEE Conference on Supercomputing, 1993, pp. 878–883. [18] E. Gabriel, G. E. Fagg, G. Bosilca, T. Angskun, J. J. Dongarra, J. M. Squyres, V. Sahay, P. Kambadur, B. Barrett, A. Lumsdaine, R. H. Castain, D. J. Daniel, R. L. Graham, and T. S. Woodall, “Open mpi: Goals, concept, and design of a next generation mpi implementation,” in Recent Advances in Parallel Virtual Machine and Message Passing Interface, D. Kranzlmüller, P. Kacsuk, and J. Dongarra, Eds. Berlin, Heidelberg: Springer Berlin Heidelberg, 2004, pp. 97–104. [19] R. L. Graham et al., “The OSU micro-benchmarks,” 2005, ohio State University benchmark suite. [Online]. Available: https://mvapich.cse.ohio-state.edu/benchmarks/ [20] Y. Gu, F. Wang, J. Fu, Z. Sun, Q. Zhang, H. Zhao, X. Liu, Y. Tian, W. Huang, Z. Liu, Y. Chen, J. Yang, Y. Zhou, Q. Zhao, H. Li, T. Wang, F. Yu, Z. Wang, G. Tan, and D. Tao, “Ccl-d: A high-precision diagnostic system for slow and hang anomalies in large-scale model training,” in Proceedings of the 31st ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming, ser. PPoPP ’26, 2026, pp. 425–438. [Online]. Available: https://doi.org/10.1145/3774934.3786429 [21] G. He, Y. Cao, Y. He, T. Bai, K. Yuan, and B. Yuan, “Tah-quant: Effective activation quantization in pipeline parallelism over slow network,” arXiv preprint arXiv:2506.01352, 2025. [22] M. Hidayetoglu, S. G. de Gonzalo, C. Pearson, T. Bicer, D. K., J. Dayal, B. Ren, S. Madireddy, R. Kettimuthu, I. Foster, and W.-M. Hwu, “CommBench: Micro-benchmarking hierarchical networks with multigpu, multi-nic nodes,” in Proceedings of the 38th ACM International Conference on Supercomputing (ICS), 2024, pp. 426–436. [23] T. Hoefler, T. Schneider, and A. Lumsdaine, “Characterizing the influence of system noise on large-scale applications by simulation,” in SC ’10: Proceedings of the 2010 ACM/IEEE International Conference for High Performance Computing, Networking, Storage and Analysis, 2010, pp. 1–11. [24] H. Hu, “dpro: A performance profiling tool for distributed systems,” https://figshare.com/articles/software/dpro/19165622, 2022, mIT License. [25] J. Huang, S. Di, X. Yu, Y. Zhai, J. Liu, Y. Huang, K. Raffenetti, H. Zhou, K. Zhao, X. Lu, Z. Chen, F. Cappello, Y. Guo, and R. Thakur, “gZCCL: Compression-accelerated collective communication framework for gpu clusters,” in Proceedings of the 38th ACM International Conference on Supercomputing, 2024, pp. 437–448. [26] J. Huang, S. Di, X. Yu, Y. Zhai, Z. Zhang, J. Liu, X. Lu, K. Raffenetti, H. Zhou, K. Zhao, K. A. Alharthi, Z. Chen, F. Cappello, Y. Guo, and R. Thakur, “Zccl: Significantly improving collective communication with error-bounded lossy compression,” ArXiv, vol. abs/2502.18554, 2025. [Online]. Available: https://api.semanticscholar.org/CorpusID:276617948 [27] W. Huang, J. Yang, S. Yin, H. Li, Y. Gu, Z. Liu, X. M. Jing, Z. Wei, S. Fu, H. Hu, G. Tan, and D. Tao, “Mans: Efficient and portable ANS encoding for multi-byte integer data on CPUs and GPUs,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, ser. SC ’25, 2025, code available at https://github.com/hpdps-group/MANS. [Online]. Available: https://doi.org/10.1145/3712285.3759825 [28] Intel Corporation, “Intel MPI Benchmarks (IMB),” https://github.com/ intel/mpi-benchmarks, 2026.

[15] D. De Sensi, T. Bonato, D. Saam, and T. Hoefler, “Swing: Short-cutting rings for higher bandwidth allreduce,” in 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24), 2024, pp. 1445–1462.

[29] A. Jangda, J. Huang, G. Liu, A. H. N. Sabet, S. Maleki, Y. Miao, M. Musuvathi, T. Mytkowicz, and O. Saarikivi, “Breaking the computation and communication abstraction barrier in distributed machine learning workloads,” in Proceedings of the 27th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, 2022, pp. 402–416.

[16] J. Diffenderfer, A. L. Fox, J. A. Hittinger, G. Sanders, and P. G. Lindstrom, “Error analysis of zfp compression for floating-point data,” SIAM Journal on Scientific Computing, vol. 41, no. 3, pp. A1867–A1898, 2019.

[30] J. Jia, C. Xie, H. Lu, D. Wang, H. Feng, C. Zhang, B. Sun, H. Lin, Z. Zhang, X. Liu et al., “Sdp4bit: Toward 4-bit communication quantization in sharded data parallelism for llm training,” Advances in Neural Information Processing Systems, vol. 37, pp. 8734–8759, 2024.

[31] Z. Jiang, H. Lin, Y. Zhong, Q. Huang, Y. Chen, Z. Zhang, Y. Peng, X. Li, C. Xie, S. Nong, Y. Jia, S. He, H. Chen, Z. Bai, Q. Hou, S. Yan, D. Zhou, Y. Sheng, Z. Jiang, H. Xu, H. Wei, Z. Zhang, P. Nie, L. Zou, S. Zhao, L. Xiang, Z. Liu, Z. Li, X. Jia, J. Ye, X. Jin, and X. Liu, “Megascale: Scaling large language model training to more than 10,000 gpus,” 2024. [Online]. Available: https://arxiv.org/abs/2402.15627 [32] J. Johnson, “Dietgpu: Gpu-based lossless compression for numerical data,” https://github.com/facebookresearch/dietgpu, 2025, gitHub repository. [33] H. Kim, J. Ryu, and J. Lee, “Tccl: Discovering better communication paths for pcie gpu clusters,” in Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3, 2024, pp. 999–1015. [34] S. Lee and J. Lee, “Collective communication performance evaluation for distributed deep learning training,” Applied Sciences, vol. 14, no. 12, p. 5100, 2024. [35] F. Liang, Z. Zhang, H. Lu, V. C. M. Leung, Y. Guo, and X. Hu, “Communication-efficient large-scale distributed deep learning: A comprehensive survey,” 2024. [Online]. Available: https://arxiv.org/abs/2404.06114 [36] X. Liang, K. Zhao, S. Di, S. Li, R. Underwood, A. M. Gok, J. Tian, J. Deng, J. C. Calhoun, D. Tao, Z. Chen, and F. Cappello, “Sz3: A modular framework for composing prediction-based error-bounded lossy compressors,” IEEE Transactions on Big Data, vol. 9, no. 2, pp. 485–498, 2023. [37] W. Lin, X. Pan, R. Fan, S. Shi, and X. Chu, “Zipccl: Efficient lossless data compression of communication collectives for accelerating llm training,” 2026. [Online]. Available: https://arxiv.org/abs/2604.27844 [38] P. Lindstrom, “Fixed-rate compressed floating-point arrays,” IEEE transactions on visualization and computer graphics, vol. 20, no. 12, pp. 2674–2683, 2014. [39] M. Liu, X. Liu, X. Tian, B. Lu, S. Lyu, S. Yin, W. Huang, Z. Wei, H. Zhao, G. Tan et al., “Taco: Efficient communication compression of intermediate tensors for scalable tensor-parallel llm training,” arXiv preprint arXiv:2604.24088, 2026, code available at https://github.com/ hpdps-group/COCCL/tree/TACO. [40] X. Liu, H. Kong, H. Zhao, S. Lyu, Z. Wei, M. Liu, X. Tian, L. Zhao, Z. Chen, F. Wang, Z. Chen, Z. Wang, G. Tan, and D. Tao, “Coccl: A collective communication library supporting easy integration and configuration of customized compression for scalable llm training,” in Proceedings of the 31st ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming, ser. PPoPP ’26, Sydney, NSW, Australia, 2026, pp. 384–397, code available at https://github.com/hpdps-group/coccl. [Online]. Available: https://doi.org/10.1145/3774934.3786432 [41] B. Lu, Z. Liu, H. Zhao, D. Luo, W. Huang, Y. Gu, J. Liu, G. Tan, and D. Tao, “Prism: An efficient GPU-based lossy compression framework for progressive data retrieval with multi-level interpolation,” in Proceedings of the 31st ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming, ser. PPoPP ’26, 2026, code available at https://github.com/hpdps-group/PRISM. [Online]. Available: https://doi.org/10.1145/3774934.3786438 [42] S. Ma, C. L. Lao, Z. Xu, Z. Wang, Z. Mao, D. Meng, J. Zhen, J. Wu, I. Stoica, Y. Wang, and Y. Zhou, “Uccl-zip: Lossless compression supercharged gpu communication,” 2026. [Online]. Available: https://arxiv.org/abs/2604.17172 [43] Z. Mao, Y. Zhang, C. Cui, Z. Huang, K. You, Z. Chen, Z. Xu, Z. Gu, S. Shenker, C. Raiciu et al., “Uccl-ep: Portable expert-parallel communication,” arXiv preprint arXiv:2512.19849, 2025. [44] D. Narayanan, M. Shoeybi, J. Casper, P. LeGresley, M. Patwary, V. Korthikanti, D. Vainbrand, P. Kashinkunti, J. Bernauer, B. Catanzaro, A. Phanishayee, and M. Zaharia, “Efficient large-scale language model training on gpu clusters using megatron-lm,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, ser. SC ’21. New York, NY, USA: Association for Computing Machinery, 2021. [Online]. Available: https://doi.org/10.1145/3458817.3476209

[45] National Center for Atmospheric Research (NCAR), “IEEE visualization 2004 contest: Hurricane isabel dataset,” http://vis.computer.org/ vis2004contest/data.html, 2004, accessed: 2024-05. [46] Network-Based Computing Laboratory, “MVAPICH: Mpi over infiniband, omni-path, ethernet/iwarp, and roce,” http://mvapich.cse. ohio-state.edu, 2024, the Ohio State University. [47] NVIDIA Corporation , “Nccl tests: Performance benchmarks for collective operations,” https://github.com/NVIDIA/nccl-tests, 2026. [48] NVIDIA Corporation, “Nccl: Nvidia collective communications library,” https://developer.nvidia.com/nccl, 2017. [49] D. K. Panda, H. Subramoni, C.-H. Chu, and M. Bayatpour, “The mvapich project: Transforming research into high-performance mpi library for hpc community,” Journal of Computational Science, vol. 52, p. 101208, 2021. [50] E. Prašnikar, M. Ljubič, A. Perdih, and J. Borišek, “Machine learning heralding a new development phase in molecular dynamics simulations,” Artificial intelligence review, vol. 57, no. 4, p. 102, 2024. [51] A. Shah, V. Chidambaram, M. Cowan, S. Maleki, M. Musuvathi, T. Mytkowicz, J. Nelson, O. Saarikivi, and R. Singh, “{TACCL}: Guiding collective algorithm synthesis using communication sketches,” in 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23), 2023, pp. 593–612. [52] M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro, “Megatron-lm: Training multi-billion parameter language models using model parallelism,” 2020. [Online]. Available: https://arxiv.org/abs/1909.08053 [53] M. Shroff and R. A. Van De Geijn, “Collmark: Mpi collective communication benchmark,” in International Conference on Supercomputing. Citeseer, 2000, p. 10. [54] S. Smith, M. Patwary, B. Norick, P. LeGresley, S. Rajbhandari, J. Casper, Z. Liu, S. Prabhumoye, G. Zerveas, V. Korthikanti, E. Zhang, R. Child, R. Y. Aminabadi, J. Bernauer, X. Song, M. Shoeybi, Y. He, M. Houston, S. Tiwary, and B. Catanzaro, “Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model,” 2022. [Online]. Available: https://arxiv.org/abs/2201.11990 [55] D. Tao, S. Di, Z. Chen, and F. Cappello, “Significantly improving lossy compression for scientific data sets based on multidimensional prediction and error-controlled quantization,” in 2017 IEEE International Parallel and Distributed Processing Symposium (IPDPS), 2017, pp. 1129–1139. [56] H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample, “Llama: Open and efficient foundation language models,” 2023. [Online]. Available: https://arxiv.org/abs/2302.13971 [57] H. Wang, S. Potluri, M. Luo, A. K. Singh, S. Sur, and D. K. Panda, “Mvapich2-gpu: optimized gpu to gpu communication for infiniband clusters,” Computer Science-Research and Development, vol. 26, no. 3-4, pp. 257–266, 2011. [58] J. Wang, Z. Ye, and X. Yu, “Ncclz: Compression-enabled gpu collectives with decoupled quantization and entropy coding,” 2026. [Online]. Available: https://arxiv.org/abs/2605.12396 [59] S. Wang, J. Wei, A. Sabne, A. Davis, B. Ilbeyi, B. Hechtman, D. Chen, K. S. Murthy, M. Maggioni, Q. Zhang et al., “Overlap communication with dependent computation via decomposition in large deep learning models,” in Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1, 2022, pp. 93–106. [60] W. Wang, M. Khazraee, Z. Zhong, M. Ghobadi, Z. Jia, D. Mudigere, Y. Zhang, and A. Kewitsch, “{TopoOpt}: Co-optimizing network topology and parallelization strategy for distributed training jobs,” in 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23), 2023, pp. 739–767.

[61] S. Zhang, N. Zheng, H. Lin, Z. Jiang, W. Bao, C. Jiang, Q. Hou, W. Cui, S. Zheng, L.-W. Chang et al., “Comet: Fine-grained computationcommunication overlapping for mixture-of-experts,” in Eighth Conference on Machine Learning and Systems. [62] K. Zhao, S. Di, X. Liang, S. Li, D. Tao, J. Bessac, Z. Chen, and F. Cappello, “SDRBench: Scientific data reduction benchmark for lossy compressors,” in Proceedings of the 1st International Workshop on Big Data Reduction (IW-BDR), held in conjunction with IEEE BigData, Dec. 2020, pp. 1–10.

Appendix A. Artifact Appendix A.1. Abstract This artifact contains the full source code of CC-Bench, which is available at CCBench_IISWC_AE. This appendix gives instructions to regenerate the experiments described in section 5 of the paper.

A.2. Artifact check-list (meta-information) • Algorithm: Lossy compression in collective communication • Program: C, CUDA • Compilation: gcc, mpicc, nvcc • Data set: Hurricane ISABEL, Llama 3 8B gradients. • Run-time environment: Linux; Open MPI; NCCL; CUDA • Hardware: CPU: 2×Xeon Platinum 8358P per node × 2

nodes, 10G InfiniBand; GPU: 2×EPYC 7402, 4×A800 SXM4 80GB per node × 2 nodes, 200GbE+ Infiband; • Metrics: Compression Rate, MAE, MSE, PSNR, CosSim, Relative Error; latency, bandwidth, Slowdown Percentage • Output: terminal metric results, per-function traces • Experiments: domain-specific error(E1), quality– performance trade-off(E2), Interference slowdown(E3), overhead breakdown(E4) • How much disk space required: ≈25GB • Time needed to prepare workflow: ≈ 15 minutes • Time needed to complete experiments: ≈ 15 minutes • Data licenses: MIT-License (Hurricane ISABEL) • Workflow automation: Semi (Shell script + manual run) • Archived: https://doi.org/10.5281/zenodo.21849825

A.3. Description Our experiments coincide with the main contributions stated in our paper as follows: C1. We introduce an application-oriented configuration module that decouples profiling logic from communication libraries, datasets, and error metrics. C2. We develop a full-stack profiling module that combines function-level interception with hardware counter monitoring. C3. We incorporate built-in representative datasets spanning scientific computing and LLM training/inference. Contribution

Supporting Artifact

Supporting Experiments

C1 C2 C3

A A A

E1, E2 E3, E4 E1, E2

A.3.1. How to access. The artifact is available at https: //github.com/konnyakucstdio/CCBench_IISWC_AE.git. A frozen version is archived at Zenodo with DOI: https: //doi.org/10.5281/zenodo.21849825. A.3.2. Hardware dependencies. see Hardware section in A.2.

A.3.3. Software dependencies. • OS: Ubuntu 22.04+(GPU) or CentOS7+(CPU) • Compiler: GCC ≥ 9.0(GPU), 7.3.1(CPU); • Python: ≥ 3.8(GPU), 2.7(CPU) • MPI: Open MPI 4.1.4 • CUDA: 12.8 with NCCL 2.27 • Compression libraries:SZx[1] , SDP4bit[2] • Communication libraries: CoCCL[3] , ZCCL[4] we recommend using our locally revised libraries[5] due to bugs in the original version. A.3.4. Data sets. • Hurricane: from https://sdrbench.github.io/. • Llama 3 8B: gradients captured from runs of Llama 3 8B. This will be provided in our link[5] .

A.4. Installation Compulsory if running first time on new platform. 1 Clone the repository, datasets and libs:

git clone https://github.com/konnyakucstdio/ CCBench_IISWC_AE.git cd CCBench_IISWC_AE ./scripts/download_data.sh 2 Set up the toolchain via module load. If modules are unavailable, manually export the environment variables following the example in scripts/export_env.sh.

# On GPU: module load cuda/12.8 openmpi/4.1.5_cuda12.8 ucx /1.12.1_cuda12.8 nccl/2.27_cuda12.8 # On CPU: module load compiler/dtk/22.04.2 compiler/devtoolset /7.3.1 mpi/hpcx/gcc-7.3.1 3 Set the communication_arch field to mpi or nccl in userconfig/config_in_jsonc/bench_basic_config, then build:

./scripts/build_script.sh --rebuild-bench

A.5. Experiment workflow GPU platform. To reproduce experiments E1 and E2: 0 Replace the cross_alloc_nodes in userconfig/ config_in_jsonc/job_config.jsonc with the nodes available for run. 1 Build the benchmark

CPU platform. To reproduce E1, E2 and E4 on CPU: 0 Same as 0 on GPU platform. 1 Build and prepare the pingpong baseline (all-in-one).: ./scripts/build_ae.sh 4 mpi 2 Run the E1 & E2 benchmark:

./scripts/run/run_benchmark.mpirun.sh 3 For E4, analyze the function traces:

python scripts/statistic/function_analyzer.py

A.6. Evaluation and expected results Outputs for experiments can be found in dir ccbench_reports. Experiments 1 and 2 should be like: Size: 536870912 bytes (134217728 elements) User: avg=11100.49 us,min=11071.08 us,max=11135.39 us Bandwidth: avg= 96.73 GB/s, min= 96.43 GB/s, max= 96.99 GB/s Correct: NO cos_sim: 9.813335e-01 mae: 3.090251e-04 ...

This should be consistent with the results in Table 2, Figure 7 and Figure 8 in our paper. For experiment 3: === AE_E3 summary: latency (us) / bandwidth (GB/s) per GPU stress level === stress% avg_us avg_GBps vs-baseline 0% 3739.50 143.59 1.00x 10% 15485.43 34.69 0.24x ...

This should be similar to Figure 10 in our paper. For experiment 4: Timeline (us): Total wall-clock: 20812851.140 Communication time: 7055756.000 ( 33.90%) Compression time: 4272003.499 ( 20.53%) Overlap (comm+comp): 4272003.499 ( 20.53%) ...

This should be similar to Figure 9 in our paper.

./scripts/build_ae.sh 1 nccl 2 Run the benchmark:

./scripts/run/run_benchmark.mpirun.sh

Likewise,To reproduce experiment E3: ./scripts/build_ae.sh 3 nccl ./scripts/AE_E3.sh

A.7. Experiment customization CC-Bench supports extensive customization through eight independent JSONC configuration files and implementing backend, compressor, profiling kernels and deviation metrics. Users can explore more by toggling all files under the userconfig directory.

Record · ID 667979 · SHA-256 b985042b3710ca31
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.