Conceptio › Archive › arXiv CS
arXiv CSopen access

DBLP: Phase-Aware Bounded-Loss Transport for Burst-Resilient Distributed ML Training

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
distributed-systemsinternetnetworkingprotocols
networking, internet, protocols, distributed systems

DBLP: Phase-Aware Bounded-Loss Transport for Burst-Resilient Distributed ML Training Zechen Ma* , Zixi Qu* , Jinyan Yi* , David Lin, Yashar Ganjali

arXiv:2605.01989v1 [cs.LG] 3 May 2026

Department of Computer Science, University of Toronto, Toronto, Canada {zma, zixiqu, jyi, davidlin, yganjali}@cs.toronto.edu * Equal contribution

Abstract—Distributed machine learning (ML) training has become a necessity with the prevalence of billion to trillionparameter-scale models. While prior work has improved training efficiency from the ML perspective at the application layer, it often fails to address transient congestion events at the network layer that introduce severe tail latency and training-time variability, thereby undermining the quality of service (QoS) of distributed ML training systems. Existing network optimizations treat all gradients equally and thus fail to integrate sufficient model-training insights into communication protocol design. In this paper, we present Dynamic Bounded-Loss Protocol (DBLP), a burst-resilient, training-phase-aware, and hardwareagnostic transport protocol that incorporates model-level tolerance properties into gradient communication. By dynamically adjusting gradient loss tolerance across training phases, DBLP reduces overall training time and mitigates tail-latency collapse during transient high-loss events (i.e., microbursts). Compared to the current state-of-the-art solution (baseline), DBLP tolerates significantly higher loss while achieving comparable test accuracy, and reduces end-to-end training time by an average of 24.4% and a maximum of 33.9%. At microburst events, DBLP achieves up to 5.88× single-round communication latency speedups over the baseline, preventing burst-induced taillatency spikes and maintaining stable training performance. Index Terms—AI Infrastructure, Distributed DNN Training, Customized Transport Protocol Design, Tail Latency Control

I. I NTRODUCTION Recent advancements in deep learning have been accompanied by rapid growth in model size. Neural networks with billions to trillions of parameters have emerged in natural language processing [1] and computer vision [2] tasks. These models consistently demonstrate that scaling up the number of parameters leads to improved performance. However, the ever-increasing model size necessitates distributed training. Distributively training trillion-parameter-scale models introduces a key scalability challenge: the overhead of gradient communication. Communication remains a bottleneck that has been widely observed and reported in [3], [4]. For instance, COMET [5] found that communication accounted for 47% of the execution time in the forward pass. To mitigate this, prior work has explored techniques such as overlapping communication with computation to partially hide latency [6], as well as gradient compression [7], quantization [8], sparsification [9], and low-rank gradient compression [10]. While these solutions effectively reduce the average latency, the tail latency (often caused by packet loss, network congestion, and transient traffic bursts) remains largely unaddressed [11]. Such tail behavior

introduces significant training-time variability and undermines the quality of service (QoS) of distributed deep neural network (DNN) training [12]. Based on these observations, we identify five design properties that a modern gradient communication protocol should satisfy: 1 Order-independence: the protocol should not rely on in-order delivery guarantees. Gradients should be applied as soon as they arrive. 2 Bounded-loss tolerance: the system should leverage the model’s intrinsic tolerance to bounded gradient loss, thereby avoiding excessive retransmissions that contribute to long-tail latency. 3 Burst-resilience: since stochastic and short-lived microbursts are inevitable congestion events in data centers [13]–[15], the protocol should maintain robust performance during microburst events. 4 Phase adaptivity: as a model’s sensitivity to gradient loss varies across training phases, the communication layer should dynamically adapt its tolerance level over time. 5 Hardware agnosticism: the protocol should minimize reliance on configurable network infrastructure, enabling deployment across commodity clusters and cloud environments where switch-level configuration is infeasible. Prior work, MLT, proposed a customized network transport protocol that optimizes both average and tail latency in distributed DNN training [11]. They leverage inter-packet dependency and priority-based packet queuing and dropping on switches to achieve properties 1 and 2 , which violates 5 . Property 3 is claimed to be satisfied; however, our microburst experiments show that an implementation of MLT without hardware support fails to sustain its communication latency and goodput during burst events. Furthermore, 4 remains unaddressed. The importance of 4 phase adaptivity is supported by prior work. [16] demonstrates that early training phases are particularly sensitive to corrupted inputs, leading to irreversible performance degradation. Building on this insight, later work [17] introduces gradient-norm-based critical-level detection and applies it to gradient compression and batchsize scheduling. These results suggest that communication reliability decisions should be informed by training dynamics. To satisfy all five desired properties in distributed training communication, we present a novel transport protocol Dynamic Bounded-Loss Protocol (DBLP). DBLP natively supports out-of-order packet delivery over unreliable channels for gradient transmission. It exploits model-level loss tolerance

by transmitting only a subset of packets within a controlled threshold. DBLP further leverages fluctuations in gradient norms to dynamically adjust this threshold across training phases, in response to the model’s varying sensitivity to gradient loss. Unlike previous work [11] that relies on specialized switch configurations, DBLP is a lightweight software-based solution that requires no hardware modifications. Our experiments demonstrate that DBLP preserves model accuracy while improving training efficiency and stability. When trained for the same number of epochs, DBLP tolerates up to 40% additional gradient loss compared to the baseline while achieving comparable evaluation accuracy. Across all evaluated models, DBLP reduces end-to-end training time by an average of 24.4% and a maximum of 33.9% over the baseline. Under microburst events, DBLP achieves up to 5.88× latency speedups within a single round of bursty communication, delivering up to 1.91× lower tail latency and 1.56× lower average latency than the baseline. These results show that DBLP mitigates burst-induced tail-latency collapse and improves the QoS of distributed ML training systems. II. BACKGROUND AND M OTIVATION A. Communication As a Bottleneck 1) Distributed DNN Training Basics: DNN models learn features over iterative training epochs. In each iteration, a partition of the data (i.e., mini-batch) is propagated through the model from the input to the output layers via forward propagation. The loss is then computed, and gradients are derived through backward propagation. Finally, the model parameters are updated using these gradients. With the emergence and prevalence of billion to trillionparameter-scale models [1], model training is no longer feasible on a single accelerator, which necessitates the use of distributed parallelism strategies. Prior work has explored a range of approaches, including data parallelism [18], tensor parallelism [19], pipeline parallelism [3], [20], expert parallelism [21]. As our focus is on the protocol behavior of DBLP, we adopt data parallelism in this work and leave the exploration of other parallelism schemes to future work. Data parallelism partitions the training data across nodes with each node maintaining a complete replica of the model. In data-parallel training, each worker node computes gradients locally and must periodically synchronize them across nodes to ensure consistent model updates. This synchronization can be performed through collective or centralized routines. Recent work [22] pointed out that MoE training involves both decentralized (i.e., inter-host and intra-host all-to-all) and centralized (i.e., intra-host gather and scatter) communication patterns. As a result, we built a centralized all-reduce architecture that supports reduce and broadcast operations between workers and the server. 2) Communication Overhead: Communication between nodes demands adequate and significant network support, as the size of gradient exchanges per iteration can range from MB to GB [23]. Many prior studies have reported and acknowledged the substantial communication overhead [6],

[23]–[26]. For instance, Pipedream observed that training over 32 GPUs resulted in 90% of total training time being spent on communication [3]. Similarly, ByteScheduler found that the end-to-end performance improvement does not scale linearly with the number of GPUs due to communication bottlenecks [23]. COMET further reinforced this observation: 47% of their training time was consumed by communication [5]. B. Existing Solutions MLT [11] identified that existing solutions for reducing the communication overhead, such as gradient compression [7], quantization [8], sparsification [9], low-rank gradient compression [10], only reduced the average latency. However, a more critical factor that contributes to the training time is the tail-latency metric [27]. To address this, MLT proposed the concept of bounded-loss tolerance: persistently losing a fraction of gradients does not significantly affect the model convergence rate. Let the bounded-loss tolerance be p, where p ∈ [0, 1). Training runs that receive all gradients versus 1 − p of total gradients can both achieve the same target accuracy, given the same number of training epochs. With these observations in mind, we now ask the following questions: if a model has a degree of tolerance to gradient loss, will the loss be a fixed constant throughout training as suggested by MLT? Could models have varying degree of tolerance to gradient loss across training phases? C. Preliminary Experiment Results To answer these questions, we set up a preliminary experiment. We trained DenseNet169 [28] on CIFAR-10 [29] for nine epochs. We divided the training into three phases: the first three epochs were the early phase; the middle three epochs were the middle phase; and the last three epochs were the final phase. We manually simulated network packet drops by randomly zeroing out a fraction of received gradients. We repeated this experiment five times. There was no loss in the first run. In the second run, we dropped 40% of the gradients during the middle and the final phases. We applied the same setting from the second run for the third experiment, but instead dropped 80% of the gradients. For the last two runs, we dropped 40% and 80% of the gradients only during the early phase. The drop rates of 40% and 80% were chosen arbitrarily to represent moderate and aggressive loss scenarios, respectively. The result is shown in Figure 1. When we drop 40% to 80% of gradients at the middle and final training phases, the model converges at roughly the same rate as when there are no drops. However, dropping gradients at the early stage significantly slows down the convergence rate, resulting in a noticeable accuracy gap. Findings from [16] are consistent with our preliminary results. Their experiments revealed that when a neural network was trained on corrupted images during early epochs, the resulting performance loss could not be recovered even with subsequent clean training. [30] also demonstrated that gradients converge to a small subspace in the early stages

Fig. 1: Preliminary Results on DenseNet169

of training. Authors from [31] pointed out that sparse and trainable sub-networks emerge during the early stages as well. These findings prompt us to purposefully ask: is early phase the only stage that deserves the highest priority with the least gradient dropping tolerance? Could some sporadic iterations in other phases require a low tolerance to gradient loss as well? D. Critical Learning Regime (CLR) We answer the questions above in this subsection: the beginning of model training is not the only important phase. Certain conditions can trigger the model to enter a critical period that necessitates low tolerance to gradient loss. The concept of a critical phase, introduced in [32], can be detected at any moment using the top eigenvalues of the Hessian. [17] further proposed a simpler yet effective gradientnorm based detection that is computationally cheap and yields comparable results to previous work. They scheduled the gradient compression between a high ratio lhigh and a low ratio llow . We employ these insights in designing our transport-layer protocol. III. D ESIGN OF DBLP In this section, we present core components of DBLP: an enhanced protocol that supports bounded-loss-tolerant transmission (III-A); a mechanism for dynamically and automatically adapting the loss-tolerant threshold (III-B). A. Bounded-Loss-Tolerant Transmission Reliable transmissions such as TCP guarantee in-order packet delivery. When packet drops occur, TCP keeps retransmitting packets until all have arrived, which can cause longtail latency. The in-order nature also contradicts property 1 : packets should be allowed to arrive out of order. On the other hand, unreliable best-effort services such as UDP do not provide any guarantees whatsoever. As a result, semi-reliable transmission was proposed in [11]. Important metadata are exchanged before training begins through reliable channels (e.g., TCP). Gradient payloads are

divided into chunks and sent through unreliable channels (e.g., UDP). During a round of communication, the sender and receiver synchronize to check progress using three control signals: probe, bitmap, and stop. The sender sends the probe signal to the receiver after sending all the gradient packets. The receiver replies with the bitmap signal along with the bitmap data that indicates delivered and missing packets. When the number of received packets out of the total reaches the bounded threshold, the receiver sends the stop signal to the sender, ending a communication cycle. Control signals (and bitmap data) are guaranteed to be delivered through reliable channels. UDP inherently allows out-of-order packet arrival. Prior work [11] divides gradients into groups and tags sub-gradients with individual IDs and offsets. This way, upon arrival, the gradients can be rearranged and the original tensor can be reconstructed. However, this methodology collapses when a packet experiences significant delay during occasional network bursts [33]. In our experiments, we indeed encountered a scenario where a packet from round k arrived at the end host during round k + 1. There is no mechanism to distinguish packets belonging to different communication cycles. Therefore, we add a new application-layer header field, an unsigned counter that increments by each communication, and we pack it with the gradient in the UDP payload. By checking the value of the counter, the end host can uniquely distinguish packets from different communication cycles and will drop any packet that does not belong to the current round (i.e., stale data that are no longer useful). B. Dynamically Changing Loss Tolerance When packet drops occur, both senders and receivers synchronize through bitmaps. Both MLT and DBLP will retransmit the lost packets based upon the most up-to-date bitmap received. Retransmission stops once the ratio of missing packets is equal to or less than the loss value p. MLT maintains a fixed value of p throughout the entire training and argues that different models have different degrees of tolerance. For example, losing 0.8% gradients while training EfficientNetB0 [34] achieves the same test accuracy and requires the same number of training iterations as when there is no loss at all. In contrast, ResNet50 [35] can tolerate up to 2.4% without any accuracy degradation. Inspired by Accordion [17] and our preliminary result in Figure 1, we implement DBLP with an adaptive bounded-loss tolerance. Instead of dividing the training equally into three regions (early, middle, final), DBLP only has critical regions and non-critical regions. We use relative drop in gradient magnitude to detect the activation of critical regimes (the terms regime and region are used interchangeably throughout this paper): |∥Gprev ∥ − ∥Gcurr ∥| ≥η ∥Gprev ∥

(1)

where Gprev is the gradient from the previous training iteration, and Gcurr is the gradient from the current iteration.

We set η = 0.5 as the detection threshold, following the value suggested in the Accordion paper [17]. Past literature claimed that larger gradients are generally more important than smaller ones [11]. Therefore, we leverage L2-norm to emphasize the contribution of large gradients while attenuating the influence of small ones. By using L2-norm, small gradients yield proportionally smaller magnitudes and thus naturally receive lower weights. The resulting norm is therefore dominated by large gradients, which are found to be more crucial. This makes it an ideal signal for detecting critical learning regimes during the model’s training process. Once a critical phase is detected, DBLP enters a critical region with a lower tolerance Plow for a predefined number of iterations. We set Plow to be the same as p in MLT. In other words, EfficientNetB0 has Plow = 0.8% and ResNet50 has Plow = 2.4%. Outside of the critical region, DBLP automatically adapts to a higher tolerance Phigh . IV. I MPLEMENTATION Algorithm 1 DBLP Send Input: Data to send 1: Split data into chunks 2: Initialize bitmap ← 0 3: while no stop signal do 4: for each unsent chunk do 5: if stop signal arrives then 6: return 7: end if 8: Send chunk 9: end for 10: Send probe signal P through TCP channel 11: if upon receiving the bitmap signal (B) then 12: bitmap ← bitmap data 13: else if upon receiving the stop signal (S) then 14: return 15: end if 16: end while

We implement DBLP in Python and use PyTorch for model training. As DBLP does not rely on any hardware configurations, it offers a suite of software APIs that enable researchers to seamlessly integrate it into their own frameworks. The list of APIs and the implementation details are presented below. A. Metadata Before training, DBLP performs an initial synchronization across all nodes to exchange the following metadata: (1) the total number of chunks per communication cycle and (2) the tensor layer names. Since these metadata values remain constant throughout training, a single synchronization over a reliable transmission channel is sufficient. During training, each UDP packet includes a custom header embedded in the data payload, consisting of three fields: (1) a sequence number that uniquely identifies the packet within a single communication round and allows the end host to rearrange out-of-order packets; (2) a counter value that distinguishes packets across multiple communication rounds; and (3) the length of the data.

Algorithm 2 DBLP_RECV Output: Data received up to the current tolerance threshold 1: bitmap ← 0 2: stop ← false 3: while not stop do 4: if a new chunk arrives via UDP channel then 5: Perform chunk integrity check 6: if the chunk is valid then 7: Accept the chunk and update bitmap 8: end if 9: end if 10: if Probe P is received then 11: if loss-tolerance threshold is satisfied then 12: Send stop signal S 13: stop ← true 14: continue 15: else 16: Send bitmap signal B with bitmap 17: end if 18: else 19: if loss-tolerance threshold is satisfied then 20: Send stop signal S 21: stop ← true 22: continue 23: end if 24: end if 25: end while 26: Unreceived chunks remain 0-filled. 27: Reconstruct the original data. 28: return reconstructed data

Algorithm 3 Server Input: Number of workers N Input: Phigh and Plow (CLR upper/lower tolerance) 1: for each worker i ∈ {1, . . . , N } do 2: Spawn a thread for worker i 3: Accept a TCP connection from worker i 4: step ← 0 5: while the worker connection is alive do 6: ∇wi ← DBLP_RECV() 7: Wait until all workers’ gradients arrive d ← mean(∇w1 , ∇w2 , . . . , ∇wN ) 8: ∇w 9: if step mod DBLPf req = 0 then 10: Reset loss-tolerance threshold to Phigh 11: if CLR is detected then 12: Set loss-tolerance threshold to Plow 13: end if 14: end if d 15: DBLP_SEND(∇w) 16: step ← step + 1 17: end while 18: end for

Algorithm 4 Worker Input: Server TCP socket address 1: Connect to the server 2: for each epoch do 3: for each mini-batch do 4: ∇w ← forward and backward pass 5: DBLP_SEND(∇w) d ← DBLP_RECV() 6: ∇w d 7: w ← w − α∇w 8: end for 9: end for

M ODEL

DATASET

PARAMETERS (M)

BASELINE p (%)

DBLP Plow (%)

DBLP Phigh (%)

E FFICIENT N ET B0 R ES N ET 50 A LEX N ET GPT-2-S

CIFAR-10 CIFAR-100 CIFAR-100 W IKI T EXT-2

5 25 61 125

0.8 2.4 0.8 0.8

0.8 2.4 0.8 0.8

40.8 42.4 40.8 40.8

TABLE I: Models and their respective bounded-loss tolerances for baseline and DBLP

B. DBLP Send At each training step, DBLP partitions the tensor into gradient chunks. The size of a chunk and the length of the customized metadata exactly fill up a UDP packet payload. Together with the UDP header, they form a MTU-sized packet. The total number of packets sent within one step depends on the size of the model. Once all gradients have been sent, DBLP transmits a probe signal P to the receiver via TCP. The receiver responds with a signal B accompanied by bitmap data, which the sender uses to retransmit only the lost packets. This process repeats until the sender receives a stop signal S from the receiver. A sender monitors the reliable channel for stop signals, and it makes sure to immediately terminate sending upon the arrival of a stop signal. When a stop signal arrives before a sender finishes sending, we call it an early stop case. The pseudocode is shown in Algorithm 1. C. DBLP Receive At each training step, DBLP initializes an empty bitmap and sets the stop signal to false. Upon receiving a valid chunk, DBLP updates the bitmap accordingly. When a probe signal P arrives, the receiver replies with the most up-to-date bitmap. Once the tolerance threshold is reached, the stop signal is sent to the sender. The pseudocode is presented in Algorithm 2. D. Centralized All-Reduce Architecture DBLP achieves the reduce and broadcast operations via a multithreaded server and three-worker setup. The goal of our evaluation is to validate the phase-aware transport policy of DBLP. Specifically, we focus on DBLP’s ability to distinguish critical from non-critical learning phases, adapt loss tolerance accordingly, and maintain robust communication performance under transient microbursts. A three-worker setup is sufficient to exercise and observe this protocol behavior in a controlled manner. Before training begins, the server spawns three threads, each dedicated to a worker, accepts their incoming connections, and allocates unique UDP port numbers for communication. Evaluating DBLP at larger scales is left for future work. At each training step, the server performs the reduce operation by collecting gradients from all workers and computing their mean. If the current gradient triggers the CLR detection mechanism, the server adjusts the bounded-losstolerance threshold to Plow ; otherwise, it sets it to Phigh . The tolerance threshold is initialized to Plow , as the start of training is indisputably critical [16]. Finally, the server broadcasts averaged gradients to all workers for model updates.

The pseudocode for the server and the worker is presented in Algorithms 3 and 4. E. CLR Parameters For the threshold η that determines whether the model is entering a critical region, we fix it at η = 0.5 for all of our experiments, following the suggestion from [17]. For every fixed round of iterations, denoted as DBLPf req , DBLP compares the gradient norm from the previous iteration with the current one, and determines if the model is entering a critical phase. V. E VALUATION A. Experimental Setup 1) Testbed: We evaluate DBLP on a testbed consisting of 3 worker nodes and 1 server node, following the centralized allreduce architecture. For image classification tasks, each worker node has 1 NVIDIA RTX A4500 GPU, 4 CPU cores, and 32 GB of memory. For language model benchmark, each worker node has 1 NVIDIA RTX 4090 GPU, 32 CPU cores, and 128 GB of memory. The network topology is treated as a black box to our system, and DBLP is designed to adapt to any settings. 2) Models and Datasets: EfficientNetB0 [34] on CIFAR10 [29]; ResNet50 [35] and AlexNet [36] on CIFAR-100 [29]; GPT-2-S [37] on WikiText-2 [38]. 3) Baseline and Metrics: We compare DBLP with a baseline across all experiments. For the baseline, we adopt a fixed bounded-loss tolerance p, following the suggested value from [11], which was shown to maintain model performance. For DBLP, we set the loss tolerance to the same p as the baseline within the critical learning regime (CLR). Outside the CLR, DBLP adaptively increases its tolerance to p + k, where the value of k is based on our empirical observations in Figure 1. We choose k = 40% as our preliminary experiments reveal that dropping 40% gradients in non-critical phases does not noticeably affect convergence. We believe that setting k equal to 40% serves two purposes: (1) DBLP is able to absorb microbursts (V-B). With its adaptive loss-tolerance design, DBLP is sufficiently robust to bursts while achieving significantly lower end-to-end training time (V-C); (2) DBLP converges to the same evaluation accuracy with the same training epochs as the baseline (V-D). Setting k too low would provide insufficient headroom to absorb bursts, undermining the core benefit of adaptive tolerance. Conversely, setting k too high (e.g., k = 80%) would be overkill, as the

tolerance threshold could exceed the packet loss rate caused by bursts [13], [39], and the amount of lost gradient information could become too large. The list of models along with their corresponding tolerance values for the baseline and DBLP is shown in Table I. Exploring the sensitivity of k (i.e., the value of Phigh ) is left for future work.

training happens outside the CLR as shown in Figure 2, the probability that a microburst collides with DBLP’s CLR is low.

B. Microburst Results

We also train all models without manually injecting microbursts. The results, shown in Figure 5 and Table V, demonstrate that DBLP consistently outperforms the baseline across all models. We normalize the training time such that DBLP is set to 1. For EfficientNetB0, DBLP achieves a 23.86% training time speedup under microbursts and a 17.44% speedup without bursts. The two short-lived stochastic microbursts constitute only 0.215% of total communication cycles, yet result in a 6.42% latency degradation for the baseline. For ResNet50, DBLP achieves a 19.78% reduction in training time under microbursts, and a 17.51% reduction in the absence of bursts. DBLP also achieves training time speedups of 33.9% on AlexNet, and 33.6% on GPT-2-S.

The majority of congestion events in data centers are transient [13]–[15]. Over 90% of packet loss is caused by short-lived microbursts [13], [40]. [39] observed events of sudden network congestion where packet loss rate exceeds 50% within a short period. Therefore, we simulated microbursts by manually injecting a 70% packet loss rate over the unreliable transmission channel for one iteration. We chose 70% as it is both a reasonably high loss threshold reported in prior work [39] and sufficiently high to stress-test DBLP’s 40.8% and 42.4% tolerance ceilings. Specifically, we train EfficientNetB0 for 10 epochs (i.e., 930 iterations) and ResNet50 for 15 epochs (i.e., 1395 iterations), and compare DBLP against the baseline. The baseline employs a constant loss tolerance, whereas DBLP adopts a 40% higher tolerance in the non-critical learning region (non-CLR). We manually inject microbursts twice for EfficientNetB0 and three times for ResNet50. The occurrence of the microbursts is chosen at random. The results are shown in Figure 2. For each iteration, we record the send latency from workers to the server. For EfficientNetB0, at the first iteration of epoch 3 (iter. 279) and 7 (iter. 651), the baseline exhibits a huge latency spike because of the injected microbursts, whereas DBLP maintains its gradient communication time at a level comparable to the no-burst case. Table II summarizes the microburst (i.e., spike) latencies, tail latencies, and average latencies for DBLP and the baseline. At the two microburst iterations, DBLP delivers 4.68× and 5.88× faster communication. DBLP also achieves 1.91× lower tail latency and 1.56× lower average latency. A CDF curve depicting the long tail of the baseline is presented in Figure 3. For ResNet50, the latency spikes appear at the first iteration of epoch 4 (iter. 372), 7 (iter. 651), and 13 (iter. 1209). DBLP, however, continues to demonstrate robustness to microbursts without exhibiting any latency spikes. Table III summarizes the statistics for ResNet50 using the same metrics, and a CDF curve is presented in Figure 4. Despite [11] claiming that their bounded-loss-tolerance gradient communication scheme resolves the tail-latency issue, we find that in scenarios without switch-level support and under microbursts, a solution that deploys a fixed small tolerance value collapses. In contrast, DBLP, a fully software-based protocol, demonstrates latency resilience against bursty packet loss and accuracy robustness to high gradient loss tolerance, owing to its dynamic phase-adaptive property. We present the evaluation accuracy results for this microburst experiment in Table IV. DBLP exhibits only 1.55% and 0.76% test accuracy degradations. Since the majority of

C. Training Time Speedups

D. Evaluation Accuracy Results We present the evaluation accuracy performance of DBLP versus the baseline in Figure 6. Despite operating under a substantially higher gradient loss rate (i.e., 40% more) and missing significantly more information across gradient transmission in non-critical phases, DBLP maintains nearly the same accuracy as the baseline across all models. For EfficientNetB0, DBLP achieves 99.59% of the baseline accuracy. For ResNet50, DBLP achieves 99.42%. Both DBLP and the baseline converge to approximately the same accuracy on AlexNet. For GPT-2-S, we use perplexity as the evaluation metric, where lower values indicate better performance. Both converge at the same rate except that the baseline displays a slightly lower perplexity at the end, as shown in Figure 6d. The test accuracy/perplexity values for all models are shown in Table VI. It is worth noting that sudden and intensive packet loss events (i.e., microbursts) do not necessarily imply that the baseline training will converge to a lower accuracy. The total number of gradient packets required for transmission remains the same. The amount of time to deliver all packets to the receiver increases dramatically, yet the end host eventually receives all the data. While we observe pronounced tail latency spikes in the baseline that significantly increases its training time, the baseline still receives and processes 99.2% of gradient data at every training iteration (because of the fixed 0.8% tolerance threshold for EfficientNetB0 as an example). This indicates that the final test accuracy converges within a stable range and may fluctuate within that range, which explains why some evaluation accuracies in the microburst experiments are higher than the no-burst results.

(a) EfficientNetB0: fixed 0.8% tolerance

(b) EfficientNetB0: CLR: 0.8%, non-CLR: 40.8%

(c) ResNet50: fixed 2.4% tolerance

(d) ResNet50: CLR: 2.4%, non-CLR: 42.4%

Fig. 2: Send Latency Comparison Under Microbursts

Latency

DBLP

Baseline

Speedup of DBLP over Baseline

Microburst 1 (Iter. 279) Microburst 2 (Iter. 651) Tail Average

0.4621 0.3619 1.1305 0.5907

2.1629 2.1293 2.1629 0.9236

4.68× 5.88× 1.91× 1.56×

TABLE II: Send Latency Measurements (EfficientNetB0): latencies of DBLP and baseline are in seconds

Latency

DBLP

Baseline

Speedup of DBLP over Baseline

Microburst 1 (Iter. 372) Microburst 2 (Iter. 651) Microburst 3 (Iter. 1209) Tail Average

2.3017 2.4367 2.6885 5.6665 3.1966

10.1180 10.0079 10.0118 10.1180 4.8519

4.40× 4.11× 3.72× 1.79× 1.52×

TABLE III: Send Latency Measurements (ResNet50)

VI. F UTURE W ORK

1.0

A. Bounded-loss Cross-DC Collectives

0.8

CDF

0.6 0.4 0.2 0.0

DBLP Baseline 0.25

0.50

0.75

1.00 1.25 1.50 Send Latency (s)

1.75

2.00

2.25

Fig. 3: CDF Curve: Microburst, EfficientNetB0

Recent efforts in NCCL have introduced topology-aware optimizations for cross-datacenter training, reducing traffic over slower inter-DC links and improving communication efficiency [41]. However, these systems still assume exact collective semantics and do not account for how a model’s sensitivity to communication errors can change during training. One possible direction is to combine topology awareness with training-aware transport, enabling DBLP to dynamically adjust reliability across datacenter boundaries through its boundedloss and phase-adaptive mechanisms. This could help mitigate tail latency and synchronization stalls in cross-DC training while maintaining convergence. B. Network Scheduling

1.0

We discover that model training is highly sensitive only during the critical learning regime (CLR). This insight suggests a potential opportunity for datacenter network scheduling: allocating higher priority and network bandwidth to ML training flows that are within their CLRs. Such scheduling policies could improve overall cluster efficiency by dynamically prioritizing the most convergence-critical communication flows at any given time.

0.8

CDF

0.6 0.4 0.2 0.0

DBLP Baseline 2

4

6 Send Latency (s)

8

10

Fig. 4: CDF Curve: Microburst, ResNet50 Model

DBLP

Baseline

EfficientNetB0 ResNet50

70.78% 82.78%

72.33% 83.54%

TABLE IV: Microburst Evaluation Accuracy Comparison Experiment Type

DBLP

Baseline

Microburst (EfficientNetB0) EfficientNetB0 Microburst (ResNet50) ResNet50 AlexNet GPT-2

1.00 1.00 1.00 1.00 1.00 1.00

1.2386 1.1744 1.1978 1.1751 1.3393 1.3357

TABLE V: Training Time Comparison (Normalized) Model

DBLP

Baseline

EfficientNetB0 ResNet50 AlexNet GPT-2-S

71.01% 82.74% 70.10% 1.6463

71.30% 83.22% 69.97% 1.5981

TABLE VI: Test Accuracy/Perplexity Comparison

C. Choosing Parameter Values DBLP includes several hyperparameters: DBLP tolerance threshold Plow and Phigh ; detection threshold η; detection frequency DBLPf req ; duration (equal to DBLPf req ) for which CLR remains active once triggered. We determine these values based on either the recommendations from prior work [17] or our own empirical observations that yielded the best results. A formal mathematical tuning guidance is out of the scope of this work, and thus further investigation into parameter tuning strategies is left to the community. VII. C ONCLUSION As distributed training continues to scale, communication reliability can no longer be treated as a purely networklayer concern independent of model behavior. Our study shows that the interaction between training dynamics and transport decisions fundamentally shapes system performance under realistic datacenter conditions. In particular, transient congestion events and microbursts expose the limitations of static reliability policies that treat all gradients and all training phases uniformly. DBLP demonstrates that communication protocols for distributed model training can be designed to be phase-aware rather than phase-agnostic. By coupling bounded-loss transmission with critical learning regime detection, DBLP adapts reliability in response to training sensitivity, allowing the system to relax retransmissions when safe and enforce stricter delivery when necessary. This alignment between learning dynamics and transport control enables substantial reductions in tail latency when microbursts occur and in training time while preserving model convergence. Our results suggest that

DBLP

Baseline

Training Time (DBLP normalized to 1)

1.5

1.0

0.5

0.0 Microburst (EfficientNetB0) EfficientNetB0

Microburst (ResNet50)

ResNet50

AlexNet

Model Type

Fig. 5: Training Time Comparison

(a) EfficientNetB0

(b) ResNet50

(c) AlexNet

(d) GPT-2-S

Fig. 6: Evaluation Accuracy Results Across Models

GPT2

incorporating training-phase awareness into transport-layer decisions offers a practical and scalable path toward improving QoS in large-scale ML systems. R EFERENCES [1] J. A. et al, “Gpt-4 technical report,” 2024. [Online]. Available: https://arxiv.org/abs/2303.08774 [2] H. Lu, W. Liu, B. Zhang, B. Wang, K. Dong, B. Liu, J. Sun, T. Ren, Z. Li, H. Yang, Y. Sun, C. Deng, H. Xu, Z. Xie, and C. Ruan, “Deepseek-vl: Towards real-world vision-language understanding,” 2024. [Online]. Available: https://arxiv.org/abs/2403.05525 [3] D. Narayanan, A. Harlap, A. Phanishayee, V. Seshadri, N. R. Devanur, G. R. Ganger, P. B. Gibbons, and M. Zaharia, “Pipedream: generalized pipeline parallelism for dnn training,” in Proceedings of the 27th ACM Symposium on Operating Systems Principles, ser. SOSP ’19. New York, NY, USA: Association for Computing Machinery, 2019, p. 1–15. [Online]. Available: https://doi.org/10.1145/3341301.3359646 [4] J. Fei, C.-Y. Ho, A. N. Sahu, M. Canini, and A. Sapio, “Efficient sparse collective communication and its application to accelerate distributed deep learning,” in Proceedings of the 2021 ACM SIGCOMM 2021 Conference, ser. SIGCOMM ’21. New York, NY, USA: Association for Computing Machinery, 2021, p. 676–691. [Online]. Available: https://doi.org/10.1145/3452296.3472904 [5] S. Zhang, N. Zheng, H. Lin, Z. Jiang, W. Bao, C. Jiang, Q. Hou, W. Cui, S. Zheng, L.-W. Chang, Q. Chen, and X. Liu, “COMET: Fine-grained computation-communication overlapping for mixture-of-experts,” in Eighth Conference on Machine Learning and Systems, 2025. [Online]. Available: https://openreview.net/forum?id=fGgQS5VW09 [6] A. Jayarajan, J. Wei, G. A. Gibson, A. Fedorova, and G. Pekhimenko, “Priority-based parameter propagation for distributed dnn training,” ArXiv, vol. abs/1905.03960, 2019. [Online]. Available: https://api.sema nticscholar.org/CorpusID:85461415 [7] Y. Lin, S. Han, H. Mao, Y. Wang, and W. J. Dally, “Deep gradient compression: Reducing the communication bandwidth for distributed training,” ArXiv, vol. abs/1712.01887, 2017. [Online]. Available: https://api.semanticscholar.org/CorpusID:38796293 [8] D. Alistarh, D. Grubic, J. Li, R. Tomioka, and M. Vojnovic, “Qsgd: Communication-efficient sgd via gradient quantization and encoding,” 2017. [Online]. Available: https://arxiv.org/abs/1610.02132 [9] J. Wangni, J. Wang, J. Liu, and T. Zhang, “Gradient sparsification for communication-efficient distributed optimization,” 2017. [Online]. Available: https://arxiv.org/abs/1710.09854 [10] T. Vogels, S. P. Karimireddy, and M. Jaggi, “Powersgd: Practical low-rank gradient compression for distributed optimization,” ArXiv, vol. abs/1905.13727, 2019. [Online]. Available: https://api.semanticscholar. org/CorpusID:173188890 [11] H. Wang, H. Tian, J. Chen, X. Wan, J. Xia, G. Zeng, W. Bai, J. Jiang, Y. Wang, and K. Chen, “Towards domain-specific network transport for distributed dnn training,” in Proceedings of the 21st USENIX Symposium on Networked Systems Design and Implementation, ser. NSDI’24. USA: USENIX Association, 2024. [12] E. Warraich, O. Shabtai, K. Manaa, S. Vargaftik, Y. Piasetzky, M. Kadosh, L. Suresh, and M. Shahbaz, “Optireduce: Resilient and tail-optimal allreduce for distributed deep learning in the cloud,” 2025. [Online]. Available: https://arxiv.org/abs/2310.06993 [13] S. Ghorbani, B. Godfrey, Y. Ganjali, and A. Firoozshahian, “Micro load balancing in data centers with drill,” in Proceedings of the 14th ACM Workshop on Hot Topics in Networks, ser. HotNets-XIV. New York, NY, USA: Association for Computing Machinery, 2015. [Online]. Available: https://doi.org/10.1145/2834050.2834107 [14] T. Benson, A. Akella, and D. A. Maltz, “Network traffic characteristics of data centers in the wild,” in Proceedings of the 10th ACM SIGCOMM Conference on Internet Measurement, ser. IMC ’10. New York, NY, USA: Association for Computing Machinery, 2010, p. 267–280. [Online]. Available: https://doi.org/10.1145/1879141.1879175 [15] S. Kandula, S. Sengupta, A. Greenberg, P. Patel, and R. Chaiken, “The nature of data center traffic: measurements & analysis,” in Proceedings of the 9th ACM SIGCOMM Conference on Internet Measurement, ser. IMC ’09. New York, NY, USA: Association for Computing Machinery, 2009, p. 202–208. [Online]. Available: https://doi.org/10.1145/1644893.1644918

[16] A. Achille, M. Rovere, and S. Soatto, “Critical learning periods in deep neural networks,” 2019. [Online]. Available: https://arxiv.org/abs/ 1711.08856 [17] S. Agarwal, H. Wang, K. Lee, S. Venkataraman, and D. Papailiopoulos, “Accordion: Adaptive gradient communication via critical learning regime identification,” 2020. [Online]. Available: https://arxiv.org/abs/ 2010.16248 [18] Y. Jiang, Y. Zhu, C. Lan, B. Yi, Y. Cui, and C. Guo, “A unified architecture for accelerating distributed DNN training in heterogeneous GPU/CPU clusters,” in 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20). USENIX Association, Nov. 2020, pp. 463–479. [Online]. Available: https://www.usenix.org/conference/osdi20/presentation/jiang [19] M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro, “Megatron-lm: Training multi-billion parameter language models using model parallelism,” arXiv preprint arXiv:1909.08053, 2019. [20] Y. Huang, Y. Cheng, A. Bapna, O. Firat, M. X. Chen, D. Chen, H. Lee, J. Ngiam, Q. V. Le, Y. Wu, and Z. Chen, GPipe: efficient training of giant neural networks using pipeline parallelism. Red Hook, NY, USA: Curran Associates Inc., 2019. [21] D. Lepikhin, H. Lee, Y. Xu, D. Chen, O. Firat, Y. Huang, M. Krikun, N. Shazeer, and Z. Chen, “{GS}hard: Scaling giant models with conditional computation and automatic sharding,” in International Conference on Learning Representations, 2021. [Online]. Available: https://openreview.net/forum?id=qrwe7XHTmYb [22] X. Liao, Y. Sun, H. Tian, X. Wan, Y. Jin, Z. Wang, Z. Ren, X. Huang, W. Li, K. F. Tse, Z. Zhong, G. Liu, Y. Zhang, X. Ye, Y. Zhang, and K. Chen, “Mixnet: A runtime reconfigurable optical-electrical fabric for distributed mixture-of-experts training,” in Proceedings of the ACM SIGCOMM 2025 Conference, ser. SIGCOMM ’25. New York, NY, USA: Association for Computing Machinery, 2025, p. 554–574. [Online]. Available: https://doi.org/10.1145/3718958.3750465 [23] Y. Peng, Y. Zhu, Y. Chen, Y. Bao, B. Yi, C. Lan, C. Wu, and C. Guo, “A generic communication scheduler for distributed dnn training acceleration,” in Proceedings of the 27th ACM Symposium on Operating Systems Principles, ser. SOSP ’19. New York, NY, USA: Association for Computing Machinery, 2019, p. 16–29. [Online]. Available: https://doi.org/10.1145/3341301.3359642 [24] H. Zhang, Z. Zheng, S. Xu, W. Dai, Q. Ho, X. Liang, Z. Hu, J. Wei, P. Xie, and E. P. Xing, “Poseidon: an efficient communication architecture for distributed deep learning on gpu clusters,” in Proceedings of the 2017 USENIX Conference on Usenix Annual Technical Conference, ser. USENIX ATC ’17. USA: USENIX Association, 2017, p. 181–193. [25] Y. Li, J. Park, M. Alian, Y. Yuan, Z. Qu, P. Pan, R. Wang, A. G. Schwing, H. Esmaeilzadeh, and N. S. Kim, “A network-centric hardware/algorithm co-design to accelerate distributed training of deep neural networks,” in Proceedings of the 51st Annual IEEE/ACM International Symposium on Microarchitecture, ser. MICRO-51. IEEE Press, 2018, p. 175–188. [Online]. Available: https://doi.org/10.1109/MICRO.2018.00023 [26] J. Fei, C.-Y. Ho, A. N. Sahu, M. Canini, and A. Sapio, “Efficient sparse collective communication and its application to accelerate distributed deep learning,” in Proceedings of the 2021 ACM SIGCOMM 2021 Conference, ser. SIGCOMM ’21. New York, NY, USA: Association for Computing Machinery, 2021, p. 676–691. [Online]. Available: https://doi.org/10.1145/3452296.3472904 [27] D. Zats, T. Das, P. Mohan, D. Borthakur, and R. Katz, “Detail: reducing the flow completion time tail in datacenter networks,” in Proceedings of the ACM SIGCOMM 2012 Conference on Applications, Technologies, Architectures, and Protocols for Computer Communication, ser. SIGCOMM ’12. New York, NY, USA: Association for Computing Machinery, 2012, p. 139–150. [Online]. Available: https://doi.org/10.1145/2342356.2342390 [28] G. Huang, Z. Liu, L. van der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), July 2017. [29] A. Krizhevsky, “Learning multiple layers of features from tiny images,” University of Toronto, 2009. [30] G. Gur-Ari, D. A. Roberts, and E. Dyer, “Gradient descent happens in a tiny subspace,” 2018. [Online]. Available: https: //arxiv.org/abs/1812.04754 [31] J. Frankle, G. K. Dziugaite, D. M. Roy, and M. Carbin, “Stabilizing the lottery ticket hypothesis,” 2020. [Online]. Available: https: //arxiv.org/abs/1903.01611

[32] S. Jastrz˛ebski, Z. Kenton, N. Ballas, A. Fischer, Y. Bengio, and A. Storkey, “On the relation between the sharpest directions of DNN loss and the SGD step length,” in International Conference on Learning Representations, 2019. [Online]. Available: https: //openreview.net/forum?id=SkgEaj05t7 [33] S. Gallenmüller, F. Wiedner, J. Naab, and G. Carle, “Ducked tails: Trimming the tail latency of(f) packet processing systems,” in 2021 17th International Conference on Network and Service Management (CNSM), 2021, pp. 537–543. [34] M. Tan and Q. V. Le, “Efficientnet: Rethinking model scaling for convolutional neural networks,” 2020. [Online]. Available: https: //arxiv.org/abs/1905.11946 [35] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 770–778. [36] A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems, F. Pereira, C. Burges, L. Bottou, and K. Weinberger, Eds., vol. 25. Curran Associates, Inc., 2012. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2012/file/c3 99862d3b9d6b76c8436e924a68c45b-Paper.pdf [37] A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learners,” 2019. [38] S. Merity, C. Xiong, J. Bradbury, and R. Socher, “Pointer sentinel mixture models,” 2016. [Online]. Available: https://arxiv.org/abs/1609 .07843 [39] S. Wang, R. Han, and X. Wang, “A coloring-based packet loss rate measurement scheme on network nodes,” Electronics, vol. 13, no. 23, 2024. [Online]. Available: https://www.mdpi.com/2079-9292/13/23/46 92 [40] T. Benson, A. Anand, A. Akella, and M. Zhang, “Understanding data center traffic characteristics,” SIGCOMM Comput. Commun. Rev., vol. 40, no. 1, p. 92–99, Jan. 2010. [Online]. Available: https://doi.org/10.1145/1672308.1672325 [41] T. Gillis, M. Mubarak, and M. Nicely. (2025, Jul.) Nccl deep dive: Cross data center communication and network topology awareness. NVIDIA Technical Blog. [Online]. Available: https: //developer.nvidia.com/blog/nccl-deep-dive-cross-data-center-communi cation-and-network-topology-awareness/

Record · ID 155212 · SHA-256 97fbce07e9f5dec6
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.