ConceptioArchivearXiv CS
arXiv CSopen access

ResiHP: Taming LLM Training Failures with Dynamic Hybrid

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
clouddistributed-computingparallel-computing
distributed computing, parallel computing, cloud

arXiv:2605.06374v1 [cs.DC] 7 May 2026

ResiHP: Taming LLM Training Failures with Dynamic Hybrid Parallelism Tenghui Ma∗

Jihu Guo∗

Wei Gao†

[email protected] Fudan University Shanghai AI Laboratory

[email protected] Fudan University Shanghai AI Laboratory

[email protected] Hong Kong University of Science and Technology

Sitian Lu

Zhisheng Ye

Hanjing Wang

[email protected] Shanghai Jiao Tong University Shanghai AI Laboratory

[email protected] Independent Researcher

[email protected] Shanghai AI Laboratory

Dahua Lin [email protected] The Chinese University of Hong Kong

Abstract Hybrid parallelism underpins large-scale LLM training across tens of thousands of GPUs. At such scale, hardware failures on individual devices lead to performance skew across devices, diminishing overall training efficiency. Existing resilient systems overlook sequence length variability in datasets and device performance skew under hybrid parallelism. As a result, (1) iteration time fluctuations induced by sequence length variability can trigger spurious failslow detections, and (2) failures are mitigated through individual adaptations in hybrid parallelism, leading to unnecessary detection overhead and inefficient resilient training. To respond, this paper presents ResiHP, a resilient system that enables robust failure detection and fine-grained adaptation for hybrid parallel training. First, we develop a Detector to accurately identify failures. In particular, it employs a workload-aware execution time predictor that disentangles failures from iteration time fluctuations while remaining lightweight for online detection. Second, we design a Scheduler that dynamically adapts parallelism group sizes, model partitioning, and workload scheduling policies to improve training efficiency under failures. Experiments show that ResiHP improves training throughput by 1.04–4.39× compared with state-of-the-art resilient training systems under diverse failure scenarios in a 256-GPU cluster.

1

INTRODUCTION

Training ever large language models (LLMs) imposes unprecedented demands on computational resources [5, 7, 34, 35, 50]. At today’s scale, sustaining high throughput requires hybrid parallelism that combines data parallelism (DP) [42], tensor parallelism (TP) [38], pipeline parallelism (PP) [18], and others [19, 25, 27]. However, as cluster scale grows, hardware failures become statistically inevitable [7, 53]. These failures commonly appear as fail-stop failures, where devices abruptly terminate due to catastrophic faults such as GPU HBM errors [7, 29, 46, 53], and fail-slow failures, where

devices remain operational but degrade in performance and act as stragglers [6, 15, 31]. Despite their different manifestations, both fail-stop and failslow introduce device performance skew, which impairs training efficiency. We define device performance skew as failure-induced heterogeneity in effective compute and/or communication rates across devices1 . Fail-stop failures force devices offline [21, 40, 53], reducing the number of active devices in a parallel group (e.g., DP, PP, or TP) and thus lowering its effective service rate [10]. Meta reports that fail-stop failures wasted approximately 178,000 GPU hours during the training of OPT-175B [53]. Fail-slow failures reduce the computation and communication rates of devices [15], triggering a cascading slowdown that originates in TP groups, propagates as bubbles across PP stages, and amplifies global synchronization delays at the DP boundary. Recent measurements [48] show that 59.2% of large-scale training jobs (≥ 512 GPUs) encounter fail-slow failures, increasing average job completion time by 34.59%. Overall, fail-stop and fail-slow introduce significant device performance skew that severely impairs training efficiency. Prior work [10, 21, 40, 48] generally structures fail-stop and failslow failure mitigation as a two-stage protocol: (1) failure detection, followed by (2) system-level adaptation to failures. Fail-stop failures can be identified by periodically collecting execution status from devices, where a device is marked as failed if status collection times out [21, 40] or reports explicit error signals [10]. Detecting fail-slow failures is challenging due to the absence of explicit failure indicators [15, 48]. The state-of-the-art fail-slow detection approach [48] relies on variations in iteration time as a proxy signal to identify candidate fail-slow failures, followed by validation to localize and confirm the degraded devices. Yet, the iteration time correlates not only with device performance but also with workloads. Real-world datasets often have diverse sequence lengths, as in the open-source GitHub dataset. Even after applying sequence packing [26, 34] to equalize input lengths, the computation workloads can still vary 1 Device performance skew differs from stragglers [29]: skew characterizes the underly-

∗ Both authors contributed equally to this work. † Corresponding author.

ing compute/communication rate heterogeneity, whereas stragglers are an executionlevel symptom that may arise from skew or from non-failure factors such as workload imbalance.

across iterations [11, 44, 45, 55] due to the quadratic complexity of self-attention with respect to sequence length [43]. As a result, workload variability across iterations leads to time fluctuations, which render failure detection prone to false fail-slow positives, thereby incurring unnecessary validation overhead. Prior resilient systems [10, 21, 47, 48] adapt to failures by tuning individual dimensions of hybrid parallelism, resulting in suboptimal training efficiency. ReCycle [10] focuses solely on PP-level workload migration to tolerate fail-stop failures. Oobleck [21] and Greyhound [48] refine workload redistribution across DP groups to balance execution time. Adaptra [47] optimizes PP-level workload scheduling to alleviate fail-slow effects. However, individual optimization in hybrid parallelism fails to address device performance skew efficiently, resulting in workload imbalance (§ 3.2). Moreover, they conservatively exclude entire TP groups even when only a subset of devices within a TP group suffer from fail-stop failures, causing hardware waste. Overall, prior resilient systems [10, 21, 47, 48] fail to jointly adapt hybrid parallelism to device performance skew, resulting in workload imbalance or low resource utilization. These gaps motivate accurate failure detection and progressive system-level adaptation in hybrid parallelism. Accurate failure detection requires identifying both fail-stop and fail-slow failures in the presence of iteration-time fluctuations caused by sequence length variability, while remaining lightweight to support online per-iteration detection. Fine-grained system-level adaptation in hybrid parallelism requires progressively adapting along the TP, PP, and DP dimensions to counter the propagation and amplification of failures. (1) TP-dimension challenge. Excluding an entire affected TP group results in severe hardware waste, whereas selectively excluding failed devices to salvage healthy ones introduces complex inter-TP-group communication. (2) PP-dimension challenge. Failures exacerbate workload imbalance across PP groups, creating extensive bubbles or stalling the entire pipeline, significantly degrading overall training efficiency. (3) DP-dimension challenge. Any remaining imbalance manifests as severe delays at global DP synchronization. Balancing replica completion times must be tightly coordinated with TP and PP adaptations. To address these challenges, we present ResiHP, a resilient training system that achieves robust failure detection and efficient systemlevel adaptation. For failure detection, we design a lightweight Detector that accurately identifies both fail-stop and fail-slow failures (§5). To detect fail-stop failures, the Detector employs a lightweight heartbeat mechanism to periodically collect heartbeat signals from all devices and marks devices as failed upon heartbeat loss [21]. To detect fail-slow failures, the Detector adopts online time series analysis [1] on recorded iteration times to identify device performance degradation. Specifically, ResiHP employs an execution time predictor to filter out iteration-time fluctuations caused by sequence length variability, thereby avoiding spurious detections and unnecessary validation while enabling highly accurate and efficient failure identification. For system-level adaptation, we design a Scheduler that progressively mitigates the device performance skew introduced by failures across TP, PP, and DP dimensions. (1) TP dimension: The Scheduler reconfigures TP group sizes to preserve healthy devices whenever possible and improve the effective throughput of affected TP groups. Additionally, it eliminates redundant communication between TP groups of varying

sizes to improve communication efficiency (§6.1). (2) PP dimension: The Scheduler adaptively repartitions the model to balance iteration time among PP groups. Moreover, it reorders workload execution to efficiently overlap communication and computation (§6.2). (3) DP dimension: Guided by the TP and PP adaptations, the Scheduler finally schedules micro-batches across DP groups to balance their execution time (§6.3). Overall, we make the following contributions in this paper. • We present ResiHP, a novel framework for resilient LLM training that tames failures with dynamic hybrid parallelism and maximizes throughput. • ResiHP utilizes an execution time predictor to factor out time fluctuations by sequence length variability, thereby improving detection accuracy and efficiency. • ResiHP effectively restores training resources and throughput by leveraging fine-grained, system-level adaptation in hybrid parallelism to mitigate failure-induced imbalances. • We implement and evaluate ResiHP with variants of LLaMA 2 [41] and Qwen 2.5 [51] under diverse failure scenarios in a cluster of 256 A100 GPUs. Experimental results show that ResiHP achieves approximately 99.4% failure detection accuracy and improves throughput by 1.04–4.39× over the baselines [10, 21, 47, 48].

2 BACKGROUND AND MOTIVATION 2.1 Fail-Stop and Fail-Slow Failures Under hybrid parallelism, failures on individual devices introduce device performance skew within and across parallel groups due to inherent synchronization [23, 34]. In LLM training, hardware failures primarily manifest as two distinct categories: fail-stop and fail-slow. We categorize fail-stop and fail-slow failure cases based on an analysis of prior studies [6, 7, 17, 39] as summarized in Table 1. Fail-stop Failures refer to deterministic events, such as CUDA errors, NVLink failures, or out-of-memory (OOM) errors, that interrupt the hardware execution immediately. Fail-slow Failures refer to gray failures where a hardware unit remains functional but exhibits reduced efficiency. Unlike fail-stop failures, which trigger immediate termination, fail-slow failures are insidious because they allow the hardware to continue running. Fail-slow failures are often induced by factors such as GPU thermal throttling, HBM3 performance degradation, or network jitter. Observations. Table 1 reveals that fail-stop and fail-slow are fundamentally intertwined rather than isolated phenomena. For example, memory and network issues appear as fail-slow when they manifest as performance degradation. Yet, they can escalate into fail-stop once error thresholds are exceeded, timeouts are triggered, or components become unavailable [6, 31, 36, 49]. The shared root causes tightly entangle fail-slow and fail-stop into a coupled failure regime. Motivation. These observations necessitate a training system that efficiently handles device performance skew caused by both failslow and fail-stop failures to preserve training efficiency.

2.2

Iteration-time Fluctuations

In LLM training, iteration time is inherently confounded by input sequence lengths [28, 29, 44, 55]. Figure 1 shows that even

ResiHP: Taming LLM Training Failures with Dynamic Hybrid Parallelism

Table 1: Summary of root causes for fail-stop and fail-slow failures in distributed training.

Fail-stop

Hardware: Memory Error (e.g., OOM and ECC errors), Network Error (e.g., RoCE, NVLink, NIC errors), Node Failure, SSD Storage Error. Software: Data race, Buggy error handling, Indefinite blocking, or loops.

Wasting 178,000 GPU hours [53] ∼10% of training time wasted [7]

Fail-slow

Hardware: Memory pressure, Network degradation (e.g., RoCE, NVLink, NIC issues), CPU contention, Power instability, and Thermal interface anomalies. Software: Data corruption, Buggy internal checker.

34.59% ACT increase [48] and up to 45% GPU underutilization [29]

Packed seq. of micro-batch

Reported Impact

Attn. mask 𝑖

0 1 2 0 1 0

𝑖 F and B of micro-batch 𝒊

0

1

0

1

2

0 2

1

1 2

2

2

Iteration B

0

0

1 0

2 0

1

0

1

2

0 1

2

1

2

2

1

2 Iteration A

Figure 1: Illustration of the impact of sequence length variability on iteration-time fluctuations. with sequence packing [26, 38, 44, 45], the quadratic cost of selfattention [43] causes substantial computation variability across micro-batches. For example, the attention computation cost of one contiguous 4K-token sequence is about four times that of a packed input of four independent 1K-token sequences. Such variability alters micro-batch execution time, disrupts tightly aligned pipeline schedules, and creates pipeline bubbles due to inter-stage dependencies. Ultimately, it appears as iteration-time fluctuations, which can be misinterpreted as the device performance skew. Motivation. Robust failure detection must account for workload variations to avoid misinterpreting iteration-time fluctuations as device performance skew.

2.3

30 20 10

12.00

1.00 3.00 0 Failure TP

PP

16.00

DP

Failure Propagation Path

Additional GPU time (Normalized)

Root Causes

Number of additionally affected devices

Category

30

25.43 19.13

20 10

4.75 1.00 0 Failure TP

PP

DP

Failure Propagation Path

Figure 2: Failure amplification across TP, PP, and DP under a fail-slow injection on LLaMA 2-13B with (𝑇 𝑃, 𝐷𝑃, 𝑃𝑃) = (4, 2, 4).

3

LIMITATIONS OF EXISTING SOLUTIONS

Existing solutions fall short in two respects in improving training efficiency under failures. First, sequence length variability interferes with failure detection. Second, they lack progressive adaptation in hybrid parallelism.

3.1

High-Overhead Detection

Detecting fail-slow failures requires inferring anomalies from indirect signals such as iteration time [48]. Although existing detectors can accurately identify iteration-time anomalies, they often fail to distinguish workload-induced fluctuations from device performance skew, especially in long-sequence training, leading to false positives and unnecessary validation overhead.

Failure Amplification Effect

In hybrid-parallel training, failures usually first manifest within TP. Because TP ranks synchronize frequently within each layer, a single crashed or slow rank can immediately disable or delay its entire TP group. If left unmitigated, this disruption propagates to PP as a degraded or unavailable pipeline stage, and eventually stalls peer DP replicas at global synchronization, amplifying the degradation across the job. To quantify this amplification, we inject a fail-slow failure that halves the speed of one GPU while training LLaMA 2-13B with (𝑇 𝑃, 𝐷𝑃, 𝑃𝑃) = (4, 2, 4). We measure the number of additionally affected devices and the additional idle GPU time. As shown in Figure 2 (left), one degraded GPU delays 3 additional GPUs in its local TP group, 12 more across the pipeline, and the remaining 16 GPUs at DP synchronization. Figure 2 (right) further shows that additional idle GPU time increases by 4.75× at TP, 19.13× at PP, and 25.43× at DP, relative to the slowdown duration of the faulty GPU. These results show that a localized failure is substantially amplified as it propagates through the hybrid-parallel hierarchy, eventually affecting the entire 32-GPU job. Motivation. Effective failure mitigation should intervene as early as possible along the failure propagation path, before the effects spread from TP to PP and DP.

3.2

Inefficient System Adaptation

Prior resilient systems typically optimize individual dimensions in hybrid parallelism [10, 21, 47, 48]. Lacking a progressive adaptation mechanism that coordinates across TP, PP, and DP, they suffer from the following critical limitations: Resource Wastage within TP Groups. When a fail-stop failure occurs within a TP group, prior resilient systems [10, 21, 47, 48] conservatively exclude the entire group, even if only a single device has failed. As a result, healthy devices are unnecessarily discarded, leading to severe resource wastage and underutilization. Inter-DP Imbalance after Workload Migration. As shown in Figure 3(a), ReCycle [10] tolerates fail-stop failures by migrating workloads at the PP level. When some device in DP0 fails, ReCycle transfers its workloads to DP1, which preserves training progress but introduces significant workload imbalance across DP groups. Intra-DP Imbalance after Workload Redistribution. As shown in Figure 3(b), Greyhound [48] mitigates inter-DP imbalance by redistributing workloads across DP groups. When DP0 suffers failslow failures, its workloads take longer to execute than those in DP1. Greyhound therefore reduces the batch size assigned to DP0

GPU F/stop

Workloads

DP

Failure report

PP Imbalance

Migration Imbalance

TP

Time

1 Training Job

F/slow

TP

Time

Time

Scheduler (Sec.6) Execution plan

4

2

Detector (Sec.5)

3

Heartbeats

GPU Pool Profiling&results

(a) Inter-DP Imbalance after Migration (b) Intra-DP Imbalance after Redistribution

Figure 4: ResiHP Overview.

Figure 3: Adapting individual dimensions in hybrid parallelism leads to severe workload imbalance and resource wastage under failures.

aggregates these signals to maintain the active device set and trigger a fail-stop decision upon the absence of several consecutive heartbeats. At the inter-node level, a central coordinator exclusively tracks the status of these node-local monitors, centrally aggregating their fail-stop decisions. By localizing the raw heartbeat stream and centralizing only the failure decisions, the global monitoring overhead scales gracefully with the number of nodes rather than individual devices, significantly reducing the overhead and complexity of communication across large clusters.

and offloads the remaining samples to DP1 to balance iteration time across DP groups. However, this redistribution introduces workload imbalance among PP groups within a DP group, as illustrated by PP0 and PP1 in DP0. This limitation indicates that effective adaptation requires cross-dimensional coordination to align workloads with device performance skew and improve resource utilization.

4

OVERVIEW

Figure 4 presents the overall architecture of ResiHP. It primarily consists of two key components: the Scheduler and the Detector. The Scheduler orchestrates the distributed training job, dictates progressive system adaptations, and implements the hybrid-parallel execution plan. Meanwhile, the Detector continuously performs lightweight and accurate online failure diagnosis across the cluster. 1 the Scheduler ingests the Job Launch. Upon job submission (○), training configuration to generate an initial execution plan that determines the optimal hybrid-parallel setup and initial workload placement. It then provisions the required computing resources 2 from the GPU pool and dispatches the execution plan (○). Online Monitoring. During training, workers in the GPU pool continuously stream heartbeats and runtime profiling results to the 3 The Detector analyzes these runtime signals to Detector (○). accurately identify failures (§5). Confirmed failures are summarized 4 into failure reports and promptly sent to the Scheduler (○). System-level Adaptation. Upon receiving a failure report, the Scheduler generates a progressive adaptation strategy based on the current cluster topology and surviving system resources (§6). It seamlessly reconfigures the parallelism dimensions and redistributes workloads across the active GPU pool, allowing the training process to resume with high efficiency. Overall, this decoupled design preserves training semantics while hiding complex failure mitigation within the system backend.

5

FAILURE DIAGNOSIS

This section describes how Detector identifies fail-stop and failslow failures during training.

5.1

Heartbeat-based fail-stop detection

To detect fail-stop failures, Detector employs a lightweight, hierarchical two-level heartbeat mechanism. At the intra-node level, each worker periodically reports compact liveness signals paired with their local training progress. A dedicated node-local monitor

5.2

Workload-Aware Fail-Slow Detection

To detect fail-slow failures, we analyze the iteration-time series to identify change points and invoke the validation phase to confirm the detection result, following the approach of Greyhound [48]. However, iteration-time fluctuations caused by workload variations can trigger false alarms, incurring unnecessary overhead and perturbing the time series for subsequent detections. To avoid this overhead and interference, once a change point is identified, we analytically estimate the expected healthy iteration time under the current workload and pipeline configuration. Micro-Batch Time Prediction. We first model the execution time of each micro-batch by separating its linear and quadratic computation components. A standard Transformer layer consists mainly of MLP operations and self-attention. For a packed micro-batch with a fixed token budget 𝑁 , MLP computation scales linearly with 𝑁 ; since 𝑁 is fixed across micro-batches, the MLP execution time remains relatively stable. In contrast, self-attention has quadratic complexity. With sequence packing, multiple documents with Í lengths {𝑙 1, 𝑙 2, . . . , 𝑙𝑘 }, where 𝑘𝑖=1 𝑙𝑖 = 𝑁 , are concatenated with block-diagonal attention masks to prevent cross-document attention. Therefore, the attention cost is proportional not to 𝑁 2 , but to Í𝑘 2 𝑖=1 𝑙𝑖 . Based on this property, we model the expected micro-batch execution time as 𝑘 ∑︁ 𝑇MB ≈ 𝛼𝑁 + 𝛽 𝑙𝑖2, (1) 𝑖=1

where 𝛼 and 𝛽 capture hardware- and model-specific costs profiled during an initial warm-up phase. Iteration-time Prediction. To derive the expected healthy iteration time, Detector uses a lightweight DAG-based analytical simulator that follows the exact pipeline schedule. For each micro-batch 𝑚 at pipeline stage 𝑠, the computation is decomposed into 𝐹𝑚,𝑠 , 𝐵𝑚,𝑠 , and 𝑊𝑚,𝑠 , which denote the Forward (F), Backward-Activation (B), and Backward-Weight (W) chunks, respectively [33, 37]. The execution time of each chunk is estimated using the micro-batch time predictor. We formulate the pipeline schedule as a DAG G = (V, E). Each vertex 𝑣 ∈ V represents a computation chunk, such as 𝐹𝑚,𝑠 , with execution cost 𝑇cost (𝑣). The directed edges in E encode two scheduling constraints. First, data dependencies ensure that a chunk

ResiHP: Taming LLM Training Failures with Dynamic Hybrid Parallelism

Layer Partition

Pipeline Workload Schedule

Stage 0 𝐿! 𝐿" 𝐿# 𝐿$ 0 Stage 1 𝐿% 𝐿& 𝐿' 𝐿( F/slow 𝒕 Stage 2 𝐿) 𝐿* 𝐿"!𝐿""

0

1

5/2𝒕

0

1

2 1

2

Repartition Stage 0 𝐿! 𝐿" 𝐿# 𝐿$ 𝐿% 0 1 2 Stage 1 𝐿& 𝐿' F/slow 4/5𝒕 0 1 Stage 2 𝐿( 𝐿) 𝐿* 𝐿"!𝐿"" 4/5𝒕 0

3 2 3 1 2 3

3 2

3

GPU

3

𝐿+

Layer 𝑖

Speed up

𝑖

Workload 𝑖

𝑖

Workload 𝑖 w/ Fail-slow

Time

Figure 5: Alleviating PP imbalance via layer repartition. can start only after its input tensor or gradient becomes available. For example, 𝐹𝑚,𝑠 depends on 𝐹𝑚,𝑠 −1 , and the corresponding edge carries the point-to-point communication time 𝑇P2P . Similar edges model gradient dependencies during the backward pass. Second, resource dependencies encode that each device executes only one chunk at a time and follows the order specified by the pipeline schedule. Thus, consecutive chunks assigned to the same stage are connected by zero-weight edges, ensuring that the next chunk starts only after the previous one completes. Given this DAG, Detector computes the earliest start time of each vertex by topological traversal:  𝑡 start (𝑣) = max 𝑡 start (𝑢) + 𝑇cost (𝑢) + 𝑇edge (𝑢, 𝑣) , (2)

preserve attention-head divisibility, communicator layout compatibility, and efficient kernel support [38], candidate TP degrees are restricted to powers of two. Thus, the feasible TP-degree set is: K = {𝑘 | 𝑘 min ≤ 𝑘 ≤ |𝐺 ′ |, 𝑘 = 2𝑞 , 𝑞 ∈ Z ≥0 }.

Throughput-Aware Subgroup Selection. For each feasible degree 𝑘 ∈ K, Scheduler constructs a candidate subgroup 𝑆𝑘 ⊆ 𝐺 ′ by greedily selecting the top-𝑘 devices ranked by normalized throughput 𝑝𝑖 , where 𝑝𝑖 is measured relative to the healthy peak of device 𝑖. This strategy naturally prioritizes healthy devices and includes the fastest fail-slow devices when necessary. After generating all candidate subgroups, Scheduler selects the optimal one. Because TP relies on tightly synchronized collectives, the effective speed of a TP group is bottlenecked by its slowest member. Meanwhile, a larger TP degree increases aggregate compute throughput by distributing computation across more devices. To balance parallel scale-out against straggler penalties, Scheduler selects the subgroup 𝑆 ∗ that maximizes the estimated aggregate throughput:

𝑆 ∗ = arg max

𝑢 ∈pred(𝑣)

where pred(𝑣) denotes the set of predecessor chunks that must complete before 𝑣 can start, 𝑡 start (𝑢) is the start time of predecessor 𝑢, 𝑇cost (𝑢) is its execution time, and 𝑇edge (𝑢, 𝑣) is the edge cost from 𝑢 to 𝑣, such as P2P communication time for data dependencies and zero for resource-ordering edges. This recurrence states that a chunk can start only after all its predecessors have completed and any required communication has finished; hence its start time is determined by the latest satisfied dependency. The expected healthy iteration time is then the completion time of the final sink vertex, i.e., the critical-path length of the DAG. Finally, if the observed iteration time of a change point exceeds the predicted healthy time by more than 25%, Detector triggers the subsequent validation phase. Otherwise, it treats the iteration as benign, removes the point from the time series, and skips validation.

6

HYBRID-PARALLEL SCHEDULING

To mitigate the amplified impact of failures, the Scheduler performs progressive adaptations across the TP, PP, and DP.

6.1

Selective Device Exclusion Within Affected TP Groups

When failures occur within a TP group, Scheduler dynamically reconfigures the affected group by selectively excluding failed or severely degraded devices, rather than conservatively removing the entire TP group. The goal is to preserve training continuity while maximizing the effective throughput of the reconfigured TP group. Candidate TP-Degree Generation. Let 𝐺 denote the original TP group, let 𝐹 stop be the set of fail-stop devices in 𝐺, and let 𝐺 ′ = 𝐺 \ 𝐹 stop be the remaining executable device set. Scheduler first determines the feasible range of candidate TP degrees. The maximum executable TP degree is bounded by |𝐺 ′ |. The lower bound 𝑘 min is dictated by device memory limits, i.e., the minimum TP degree required for each device to hold its model shards. To

(3)

𝑆𝑘 , 𝑘 ∈ K

  𝑘 · min 𝑝𝑖 .

(4)

𝑖 ∈𝑆𝑘

Due to the power-of-two constraint, the selected subgroup may still leave some healthy or moderately degraded devices unassigned. Scheduler keeps these devices online as node-local standby devices, allowing the system to reuse them for subsequent intra-node failures. Furthermore, the adaptation results in heterogeneous TP degrees, which severely complicates the point-to-point (P2P) communication between different TP groups. To address this, we design a symmetric mapping rule to efficiently orchestrate inter-TP-group communication, the details of which are deferred to the P2P communication optimization in §7. In summary, selective exclusion isolates severe failures and salvages partial TP computation capacity. However, the remaining throughput heterogeneity across TP groups can propagate to PP, where it turns the affected stage into a straggler and necessitates subsequent PP-level compensation.

6.2

Layer Repartition to Alleviate PP Imbalance

To alleviate the straggler introduced after TP adaptation, Scheduler assigns fewer layers to the straggling PP stages and evenly redistributes the excess layers across the remaining stages. Scheduler adjusts the number of layers assigned to each device to alleviate workload imbalance caused by fail-slow failures. Figure 5 illustrates a pipeline where a fail-slow failure on Stage 1 increases the perworkload computation time. This slowdown propagates to Stage 0 and Stage 2 as pipeline bubbles, forcing the execution time of all PP groups to synchronize with the degraded PP group. To reduce this imbalance, Scheduler repartitions layers across PP groups. Specifically, it reduces the number of layers on the PP group with fail-slow failures and reallocates them to healthy PP groups. Figure 5 shows the resulting layer repartition, changing the number of layers per PP group from (4, 4, 4) to (5, 2, 5). This repartition mitigates execution time imbalance.

Table 2: Notation used in DP workload migration. Symbol

Meaning

D S 𝑑, 𝑑 ′ 𝑖, 𝑗, 𝑡 𝑑→𝑑 ′ 𝑥𝑖,𝑗 𝑇makespan (𝑑) 𝑀𝑑 ′ ,𝑖 , 𝐶𝑑 ′ ,𝑖 𝑃𝑑,𝑖,𝑡 𝑑 min, 𝑑 max 𝛿 𝐹 𝑗,𝑖,𝑑

Set of DP groups Set of PP stages, S = {0, 1, . . . , 𝐼 − 1} Source and executor DP groups PP stage, micro-batch, and scheduling time slot 1 iff micro-batch 𝑗 of stage-𝑖 from 𝑑 runs on 𝑑 ′ Completion time of DP group 𝑑 Memory footprint and capacity of stage 𝑖 on 𝑑 ′ Progress of stage 𝑖 in DP group 𝑑 at time slot 𝑡 Slowest and fastest DP groups for stage 𝑖 Progress-imbalance threshold F chunks of micro-batch 𝑗 at stage 𝑖 in 𝑑

6.3

Algorithm 1: Progress-Aware Workload Migration while unfinished workloads exist do Advance pipeline schedules and update dependencies; foreach PP stage 𝑖 ∈ S do 1 Identify slow/fast replicas // ○ Compute 𝑃𝑑,𝑖,𝑡 for all 𝑑 ∈ D; 𝑑 min ← arg min𝑑 𝑃𝑑,𝑖,𝑡 , 𝑑 max ← arg max𝑑 𝑃𝑑,𝑖,𝑡 ; 2 Generate migration plan // ○ if (𝑑 min, 𝑖) is fail-stop or 𝑃𝑑max ,𝑖,𝑡 − 𝑃𝑑min ,𝑖,𝑡 > 𝛿 then 𝑗 ← NextPending(𝑑 min, 𝑖); 3 Migrate if memory is feasible // ○ if 𝑗 ≠ ⊥ and MemoryFeasible( 𝑗, 𝑖, 𝑑 max ) then Migrate stage-𝑖 workload of 𝑗 to 𝑑 max ; Update 𝑃𝑑min ,𝑖,𝑡 and 𝑃𝑑max ,𝑖,𝑡 ;

DP Adaptation via Progress-Aware Workload Migration

After TP and PP adaptations, residual execution skew may still remain across DP groups due to heterogeneous TP configurations or unresolved pipeline bubbles. To better align DP completion times, Scheduler dynamically migrates micro-batch workloads across DP groups at the granularity of individual PP stages. We formulate this fine-grained migration as a constrained makespan minimization problem and solve it using an online progress-aware heuristic. Problem Formulation. A micro-batch normally executes all its stages within its source DP group. To tolerate failures or mitigate stragglers, the Scheduler may migrate the workloads of one stage to the corresponding stage in another DP group. The migration 𝑑→𝑑 ′ . The objective is to minimize the decision is represented by 𝑥𝑖,𝑗 iteration time, i.e., the maximum completion time across DP groups: min max 𝑇makespan (𝑑)

(5)

X 𝑑∈D

subject to dependency and resource constraints. Scheduling Constraints. The migration plan must satisfy three constraints. (1) Execution completeness. Each stage of each microÍ 𝑑→𝑑 ′ = 1. (2) Dependency batch is executed exactly once: 𝑑 ′ ∈ D 𝑥𝑖,𝑗 preservation. If stage 𝑖 of a micro-batch from 𝑑 is migrated to 𝑑 ′ , activations and gradients must be exchanged between 𝑑 and 𝑑 ′ so that the adjacent stages on the original pipeline can continue execution. (3) Memory capacity. The memory footprint on the destination stage must not exceed its capacity, i.e., 𝑀𝑑 ′ ,𝑖 ≤ 𝐶𝑑 ′ ,𝑖 at any time, including live activations that remain until the corresponding backward computation completes. Progress-Aware Heuristic Solver. Solving the global migration problem as a mixed-integer program is too expensive for online training. Instead, the Scheduler uses a progress-aware heuristic with an analytical pipeline simulator, as shown in Algorithm 1 and Figure 6. The computation of each micro-batch is decomposed into Forward (F), Backward-Activation (B), and Backward-Weight (W) 1 Scheduler maintains chunks. To quantify execution progress ○, 𝑃𝑑,𝑖,𝑡 for each stage and DP group. Specifically, 𝑃𝑑,𝑖,𝑡 counts the number of forward workloads completed by stage 𝑖 in DP group 𝑑, including both local and migrated micro-batches. At each scheduling iteration, Scheduler advances the execution state and identifies

the slowest and fastest DP groups for each stage: 𝑑 min = arg min 𝑃𝑑,𝑖,𝑡 , 𝑑∈D

𝑑 max = arg max 𝑃𝑑,𝑖,𝑡 .

(6)

𝑑∈D

As shown in Figure 6(b), at time slot 𝑇2 , the progress metrics for stage 0 are 𝑃0,0,2 = 𝑃2,0,2 = 2 and 𝑃1,0,2 = 1. This identifies DP1 as the straggler (𝑑 min ), while DP0 and DP2 tie for the maximum progress (𝑑 max ). Guided by the detected failures and measured progress, the 2 Scheduler evaluates the following migration decisions ○: • Fail-slow load balancing. The Scheduler triggers workload migration only when the progress gap exceeds a predefined threshold 𝛿 (i.e., 𝑃𝑑max ,𝑖,𝑡 − 𝑃𝑑min ,𝑖,𝑡 > 𝛿). As illustrated in Figure 6(b), at time 𝑇2 , when 𝛿 = 0, the observed progress gap is 2 − 1 = 1 > 0. Consequently, stage 0 in both DP0 and DP2 are eligible migration destinations for the pending chunk 𝐹 5,0,1 from the degraded DP1. • Fail-stop eviction. If one stage encounters a fail-stop failure, it can no longer execute its remaining workloads. The Scheduler therefore migrates its pending micro-batches to healthy peer stages in other DP groups. In Figure 6(a), stage 2 of DP1 crashes, so the pending stage-2 workloads of DP1 are migrated to peer stages in other DP groups. When DP0 and DP2 have the same progress at this stage, both are valid destinations. Each migration is first simulated and finalized only if the destination stage satisfies the memory constraint at the projected arrival time 3 By traversing all stages and tracking their progress globally, ○. the Scheduler handles fail-stop and fail-slow failures in a unified manner and performs fine-grained stage-level migration subject to memory feasibility. This global yet lightweight heuristic greedily reduces inter-DP progress gaps and balances DP completion times with negligible scheduling overhead.

7

IMPLEMENTATION & OPTIMIZATION

We have implemented ResiHP in approximately 9k lines of Python code based on our internal optimized LLM training framework, similar to Megatron-LM[34]. Detector. Detector employs two decoupled runtime mechanisms to identify fail-stop and fail-slow failures. For immediate fail-stop

ResiHP: Taming LLM Training Failures with Dynamic Hybrid Parallelism

1 2 3 4 5

7

Time

20

19

18

17

16

15

14

13

12

11

10

9

8

7

6

5

𝑀 1 2 3 4 5 Pipeline of DP2 8 9 6 10 11 8 8 9 9 6 10 6 11 10 11 8 9 10 11 8 8 9 9 10 10 11 11 8 9 5 8 10 7 5 8 9 11 10 7 11 5 9 10 7 11 8 8 8 9 9 9 10 10 10 11 11 11

26

25

24

1 2 3 4 5

𝑀

Pipeline of DP1 4 7 4 4 7 4 5 6 7 4 4 5 5 6 6 7 7 4 5 6 4 7 5 6 7 4 4 5 5 4 6 6 7 7 5 6 7

4

1 2 3 4 5

z F, B, W of pipeline for micro-batch z of DP 2

3

𝑀

7

z

2

1 2 3 4 5

1

7

23

22

21

20

19

18

17

16

15

14

13

1

Stage 3

12

Stage 2

11

Stage 1

6

𝑀

Pipeline of DP2 Inter-DP imbalance 𝑀 8 9 10 11 8 8 9 9 10 10 11 11 8 9 10 11 8 8 9 9 10 10 11 11 8 9 10 5 8 11 5 8 7 9 5 10 9 7 10 11 7 11 8 8 8 9 9 9 10 10 10 11 11 11 10

Stage 0

9

Stage 3

8

Stage 2

7

F/stop

Pipeline of DP1 4 5 6 7 4 4 5 5 6 4 5 6 7 4 4 5 5 6 6 7 7 4 5 4 6 5 7 6 7 4 4 5 5 4 5 6 6 6 7 7 7

6

Stage 0 F/slow Stage 1

5

Stage 3

Failed Stage GPU 𝑀 1 2 3 4 5 Pipeline of DP0 0 1 5 2 3 0 0 5 5 1 2 1 3 2 3 0 1 2 3 0 0 1 1 2 2 3 3 0 4 1 6 0 2 4 0 1 3 6 2 3 4 1 2 6 3 0 0 0 1 1 1 2 2 2 3 3 3

4

Stage 2

z

x F, B, W of pipeline for micro-batch x of DP 0 y F, B, W of pipeline for micro-batch y of DP 1 Pipeline of DP0 0 1 2 3 0 0 1 1 2 2 3 3 0 1 2 3 0 0 1 1 2 2 3 3 0 4 1 2 0 6 4 0 3 1 4 2 1 6 2 3 6 3 0 0 0 1 1 1 2 2 2 3 3 3

3

Stage 1

x

y

2

Stage 0

x

y

Time

(b) Workload Migration by Scheduler of ResiHP

(a) Schedule of ReCycle

Figure 6: (a) ReCycle [10] migrates failed-stage workloads without considering stage-level progress, causing inter-DP imbalance. (b) Scheduler migrates pending workloads to faster peer stages under memory constraints, thereby reducing imbalance and shortening iteration time. TP

TP G0

G1

G2

G0

G3

GPU

G1

G2

G1

G2

Scatter G0 PP

P2P

×2

PP PP

Exclusion

Fail-stop ❌ G4

G6

G5 TP

G7

(a) W/o P2P Comm Optimization

All-Gather

P2P

×1 ❌ G4

Layer Params & Opt. States

G5

G6

G6

Copy

Transfer

Transferred Layer

F/slow

F/stop

PP

G7

PP

F/stop

G7

TP

(b) With P2P Comm Optimization

Figure 7: Eliminating redundant P2P transfers after dynamic TP reconfiguration. Left: without scatter/gather, identical tensors are sent repeatedly (2× top-down and 4× bottom-up) across nodes over InfiniBand. Right: with scatter/gather, tensors are first scattered or gathered over the faster intra-node NVLink/NVSwitch fabric, and only one copy is sent across nodes over InfiniBand, reducing communication traffic.

detection, we use a TCP-based side channel: a CPU agent on each node maintains a persistent TCP connection to a centralized controller. When a node crashes, the controller detects the socket disconnection and broadcasts a fail-stop notification to trigger system reconfiguration. Detector preserves Greyhound’s validation-based fail-slow detection criterion while introducing a workload-aware filtering step prior to validation. Scheduler. The Scheduler receives failure diagnostics from the Detector, including failure location and severity, and generates an adaptation plan that specifies layer partitioning and workload scheduling. During reconfiguration, we use torch.distributed to destroy stale communication groups and rebuild them while excluding ranks affected by fail-stop failures; the model is then reshaped across the reconstructed groups according to the new partitioning plan. For runtime execution, Scheduler emits a sequence

F/stop

F/stop

DP

Scatter

All-Gather ❌ G4 G5

Drop

TP

G3 1×

TP shard

DP

G3

F/slow

F/stop TP PP

F/stop

(a)Reconfigure without interrupting the training process.

F/stop

(b) Reconfigure when all replicates encounter F/stop.

F/stop

(c) Reconfigure for Layer Repartition.

Figure 8: Recovery of optimizer and parameters during reconfiguration. ResiHP copies or transfers states from surviving replicas to rebuild dropped TP shards, support layer repartition, and recover training states when all replicas in a stage encounter fail-stop failures. of primitives, such as Forward, Backward, Send, and Recv, which are parsed and executed by a lightweight worker-side interpreter. Portability. The ideas of the Detector and Scheduler are portable to other hybrid-parallel frameworks with analogous control over communication groups and execution schedules, while communicator reconstruction and primitive dispatch are framework-specific. P2P Communication Optimization. During P2P communication between adjacent pipeline stages, GPUs within the same TP group send and receive identical tensors. As shown in Figure 7(a), without scatter/gather optimization, the same tensor set is transmitted twice in the top-down direction and four times in the bottomup direction over InfiniBand, causing substantial cross-node communication redundancy. To reduce this redundancy, we build on Megatron-LM’s scatter/gather optimization [34]. This optimization requires identical TP degrees for both the sender and receiver in P2P communication. However, dynamic TP reconfiguration (§6.1) for handling fail-stop failures may leave communicating peers with

Table 3: Models and 3D parallelism settings. Scale

LLaMA 2 [41] Qwen 2.5 [51] (TP,DP,PP) #GPUs

Small Medium Large XLarge

7B 13B 30B 70B

7B 14B 32B 72B

(4, 2, 2) (4, 2, 4) (4, 2, 8) (4, 4, 16)

16 32 64 256

Table 4: Prediction accuracy of the micro-batch time predictor (MTP) and iteration-time predictor (ITP), reported as Mean Absolute Percentage Error (MAPE). Model

Seq.

MBs

Sched.

MTP

ITP

Qwen 2.5-7B Qwen 2.5-14B LLaMA 2-13B

8K 16K 32K

4 8 8

1F1B[33] ZBH[37] 1F1B

1.19% 1.58% 1.21%

2.81% 5.06% 4.89%

heterogeneous TP degrees, violating this requirement. To preserve scatter/gather in this setting, we introduce new P2P communication rules. As shown in Figure 7(b), the sender first scatters tensors into 𝑁 equal-sized chunks, where 𝑁 is the larger TP degree of the two peers, and sends the chunks to the corresponding GPUs over InfiniBand. The receiver then reconstructs the final tensor via a faster intra-node NVLink/NVSwitch all-gather, so that identical tensors are sent only once across nodes, thereby reducing InfiniBand traffic and improving communication efficiency. Optimizer State and Parameter Recovery. After a failure is detected, the runtime enters an online reconfiguration phase to reconstruct communication groups, model states, and optimizer states while preserving training progress. As shown in Figure 8(a), ResiHP first excludes failed DP replicas and rebuilds the communication groups, allowing healthy replicas to continue execution. Once the current iteration completes, the latest committed parameters and optimizer states are synchronized across the reconfigured groups. Figure 8(b) illustrates the case in which different DP replicas encounter fail-stop failures. In this case, training must pause, and ResiHP falls back to the persistent states from the last completed iteration to reconstruct the missing states and parameters. Figure 8(c) shows recovery under layer repartition: ResiHP migrates parameters and optimizer states together with reassigned layers, and if the TP degree also changes, it dynamically reshards the transferred states to match the target TP layout using optimized P2P communication. Through this system-level state remapping, ResiHP preserves training semantics, maintains training progress, and sustains convergence even under frequent failures.

8

EVALUATION

We evaluate the Detector and the Scheduler to answer three questions: (1) How accurately does the Detector identify fail-stop and fail-slow failures across various models and parallelism settings (§8.2)? (2) How effective is the Scheduler at mitigating fail-stop and fail-slow failures under different failure scenarios and parallelism settings (§8.3)? (3) How well does ResiHP preserve training efficiency in large-scale real-world failure scenarios (§8.4)?

Table 5: Avg. number of false alarms (FA), overhead for one false alarm, and fail-slow detection accuracy over traces. Results are reported as ResiHP vs. Greyhound [48]. Model Size, Seq. Small, 8K Medium,16K Medium, 32K

8.1

FA

Overhead

Accuracy

0 / 3.7 0.1 /5.2 0.3 / 8.7

34ms / 2.24s 45ms / 3.28s 49ms / 3.72s

1.00 / 1.00 1.00 / 1.00 0.98 / 0.98

Experimental Setup

Testbed Configuration. We conduct our evaluation on a cluster comprising 32 nodes, each equipped with 8 NVIDIA A100 GPUs connected via NVSwitch. The nodes are interconnected via 200Gbps HDR InfiniBand. We utilize our internal optimized framework to train a set of LLaMA 2 and Qwen 2.5 models of various sizes and parallel strategies. Table 3 details the optimized hybrid parallelism strategies for the allocated set of healthy GPUs. The testbed runs CUDA version 12.2 and NCCL version 2.20.1. Failure injection. We evaluate the resilience of our system using deterministically injected fail-stop and fail-slow failures. To emulate fail-stop failures, we manually terminate a subset of workers during training, forcing the system to resume execution with the remaining available devices. To simulate computational fail-slow failures, we employ nvidia-smi to lock the GPU SM frequency. To inject communication fail-slow failures, we initiate side-channel communication jobs that create network bandwidth contention, thereby reducing the available bandwidth on specific network links. Baselines. We compare ResiHP with four representative baselines: Greyhound [48], Adaptra [47], ReCycle [10], and Oobleck [21]. They cover three failure-handling capabilities. Fail-slow detection and mitigation. Greyhound detects fail-slow failures from anomalous iteration-time increases and mitigates them by redistributing micro-batches across DP groups according to processing speed. Adaptra optimizes PP-level workload scheduling to mitigate communication-related fail-slow failures. Fail-stop tolerance. ReCycle tolerates fail-stop failures by rerouting micro-batches from failed ranks to DP peers in the same pipeline stage to execute. Oobleck recovers by switching to a precomputed pipeline template that uses fewer nodes. Mixed-failure handling. By integrating Greyhound’s fail-slow detection and mitigation, we build strengthened versions of ReCycle and Oobleck that can handle both fail-slow and fail-stop failures. Metrics. For detection, we report the Mean Absolute Percentage Error (MAPE) of the micro-batch and iteration-time prediction, as well as detection accuracy and average false alarms over the entire job. To evaluate the effectiveness of Scheduler in handling failures, we report end-to-end throughput in samples per second (samples/s) across training iterations. Unless otherwise stated, each aggregate result is averaged over multiple independent runs, and error bars denote 95% confidence intervals computed from perrun means [13, 20]. For single-run temporal traces, we report the observed trajectory without confidence intervals.

8.2

Detector Accuracy

Time Prediction Accuracy. We first assess the micro-batch and iteration-time predictors across three representative training setups

ResiHP: Taming LLM Training Failures with Dynamic Hybrid Parallelism

Table 6: Average Throughput (samples/s) with increasing fail-stop failure frequency (– denotes aborted training). Models Fail-stop Frequency Fault-free Oobleck[21] ReCycle[10] ResiHP

LLaMA 2-7B 2h 1h 30m

LLaMA 2-13B 2h 1h 30m

Throughput (samples/s) Throughput (samples/s)

0

5 0

1.81x 1.30-1.30x 1x

W

LLaMA 2-7B

1.0

2.47x 3.31x 2.01x 2.66x 0.5 1.40x 1x 1x1.45x

M # Failure Severity

1.32x 1x1.06-1.09x

W

Qwen 2.5-7B 2h 1h 30m

Qwen 2.5-14B 2h 1h 30m

Qwen 2.5-32B 2h 1h 30m

8.22 8.05 5.05 7.07 5.89 5.59 6.04 4.59 – 6.47 4.79 – 4.48 3.31 – 5.37 4.37 – 4.74 3.65 – 4.47 3.42 – 4.16 3.94 – 4.91 4.66 – 3.48 3.36 – 4.37 4.23 – 3.73 3.58 – 3.58 3.47 – 7.55 6.48 5.46 7.95 7.23 5.95 4.80 4.38 3.53 6.57 5.95 5.62 5.59 5.06 4.45 4.80 4.23 3.86

Fail-slow 5

LLaMA 2-30B 2h 1h 30m

S

0.0

Qwen 2.5-7B 1.76x 1.44x 1.21x 1x

2.13x 5.0 1.70x 1.32x 1x 2.5

M # Failure Severity

S

0.0

×101

Adaptra Greyhound LLaMA 2-13B

1.66x 1.30-1.33x 1x

W

W

3.20x 2.40x 5.0 1.80x 2.23x 1.37x 1.39x 1x 2.5 1x

M # Failure Severity

1.48x 1.19-1.26x 1x

ResiHP

S

0.0

Qwen 2.5-14B 1.83x 1.41x 1.20x 1x

M # Failure Severity

1.52x 1.13-1.20x 1x

W

LLaMA 2-30B

2.02x 2.51x 1.31-1.45x 1.72x 1.36x 1x 1x

M # Failure Severity

S

Qwen 2.5-32B

2.23x 5.0 1.74x 1.29x 1x 2.5

S

0.0

1.51x 1.74x 1.99x 1.13-1.26x 1.22-1.26x 1.42x 1x 1.20x 1x 1x

W

M # Failure Severity

S

Figure 9: Effectiveness of the Scheduler in mitigating various fail-slow severities. varying in sequence lengths, model sizes, and pipeline schedules. As shown in Table 4, the micro-batch time predictor achieves a MAPE of 1.19%–1.58%, while the iteration-time predictor achieves 2.81%– 5.06%. Detector accurately estimates healthy execution times even under substantial workload variability, providing a reliable baseline to filter out the benign iteration-time fluctuations. Failure Detection Accuracy and Overhead. To evaluate failslow detection accuracy, we follow Greyhound [48] and launch multiple short training jobs. In approximately half of the jobs, we inject a persistent fail-slow at a random iteration after the first 50 iterations and before the last 50 iterations, leaving sufficient time for detector warm-up and failure response. Table 5 shows that the Detector substantially reduces false alarms and the associated overhead relative to Greyhound. Across all evaluated setups, the Detector incurs only 34–49 ms of false-alarm overhead, compared with 1.24–1.72 s for Greyhound. This reduction comes from filtering benign workload-induced spikes before invoking the validation phase. In addition, the Detector achieves over 99.0% accuracy in identifying fail-slow anomalies and 99.6% accuracy for fail-stop failures, demonstrating its effectiveness and robustness in failure diagnosis. This demonstrates that the Detector reliably covers fail-stop and persistent fail-slow failures, while failures without liveness or timing signatures are outside the current scope.

8.3

Scheduler Effectiveness

Comparison against ReCycle[10] and Oobleck[21] under Failstop Failures. To evaluate ResiHP’s adaptability across a wide spectrum of deployment scenarios, we vary the fail-stop frequency from once every 2 hours to once every 30 minutes across 4–16h

training sessions (scaled by model size), following prior work [2, 10, 21]. Workers are monotonically terminated over time, leaving only 50% of the initial cluster intact in the 30m scenario. Table 6 shows that ReCycle already suffers throughput drops of up to 49.4% at the 2h frequency due to severe inter-DP imbalance induced by workload migration. Oobleck performs better initially, but still degrades by up to 44.2% at the 1h frequency due to imbalanced heterogeneous pipelines and high reconfiguration latency. Both baselines abort training under the severe 30m frequency, as they conservatively discard entire TP groups after intra-group failures and cannot recover when all DP replicas of a pipeline stage fail. In contrast, ResiHP sustains training even in the 30m scenario. Across the comparable 2h and 1h cases, it achieves 1.22–1.82× and 1.07– 1.51× throughput speedups over ReCycle and Oobleck, respectively, by progressively adapting TP, PP, and DP to salvage fragmented resources and maintain workload balance. Comparison against Greyhound [48] and Adaptra [47] under Fail-slow Failures. To evaluate ResiHP’s ability to handle realistic fail-slow scenarios of varying severities, we inject weak (W), medium (M), and severe (S) fail-slow failures. Following experimental setups in prior work [47, 48], these fail-slow categories induce unmitigated throughput drops of roughly 35%, 55%, and 70% relative to the fault-free baseline, respectively. As illustrated in Figure 9, ResiHP improves the overall throughput of the degraded baseline by 1.32–3.31×, translating to 1.18–2.30× and 1.22–1.46× speedups over Adaptra and Greyhound, respectively. By consistently realigning workloads to match heterogeneous device speeds, ResiHP maintains high training efficiency across diverse pipeline configurations and fail-slow severities.

Throughput (samples/s) Throughput (samples/s)

5.0 2.5

2.21x 1.97x 1.70x 1x

0.0

5.0 2.5 0.0

ReCycle LLaMA 2-7B

2h

4.39x 3.16x 1-1.02x

3.68x 5 1-1.04x

1h 30min # Failure Frequency 1.87x 1.59x 0.99-1x

1x

1-1.01x

1.61x 4

1h 30min # Failure Frequency

2.24x 1.88x 1.66x 1x

0

Qwen 2.5-7B

1.33-1.62x

2h

Strengthened ReCycle Strengthened Oobleck LLaMA 2-13B

2

2h

2h

0

1h 30min # Failure Frequency

1.84x 1.46-1.61x 1x

0

4

3.20x 2.42x 2.45x 2 1-1.01x 0.99-1.01x

Qwen 2.5-14B

1-1.01x

1.78-1.87x

1-1.04x

2h

2.32x 1.86x 0.99-1x

0.97-1x

1.88x

1h 30min # Failure Frequency

Qwen 2.5-32B

4

2.00x 1.74x 1-1.02x

ResiHP LLaMA 2-30B

1.75x

2 0

1h 30min # Failure Frequency

1.49-1.54x 1-1.02x

2h

1.56-1.68x 1.48x 0.99-1.04x 1-1.01x

1h 30min # Failure Frequency

30B 13B 7B 0.0

Device Exclusion

1.47x 1.82x 1.01x 0.18x

0.5

1.0

Layer Repartition 0.17x0.24x 0.22x 1.03x

1.5

2.0

Workload Migration 1.16x

2.5

Throughput (normalized to ReCycle baseline)

3.0

Figure 11: Performance breakdown for handling mixed failures across LLaMA 2-7B, 13B, and 30B models, with failure frequencies of 2h, 1h, and 30 min, respectively. Handling mixed Failures. To evaluate effectiveness under complex and realistic failure scenarios, we alternately inject fail-stop failures and medium-severity fail-slow failures during training. Figure 10 shows that ResiHP improves throughput by 1.48–4.39×, 1.22–4.32×, and 1.04–3.57× over ReCycle, strengthened ReCycle, and strengthened Oobleck, respectively. Notably, strengthened ReCycle provides negligible benefit over its vanilla counterpart in these mixed-failure scenarios. This is because devices already degraded by fail-slow failures may also receive additional workloads from crashed DP peers, turning them into severe stragglers that throttle end-to-end throughput. Performance Ablation. We quantify the contribution of each ResiHP component by incrementally enabling selective device exclusion, adaptive layer repartition, and workload migration. Figure 11 reports each component’s throughput contribution normalized to ReCycle. Among the three components, selective device exclusion provides the largest gain because it directly preserves useful TP computation capacity and reduces resource waste under failures. Adaptive layer repartition contributes less because its uniform application across DP replicas limits flexibility. In contrast, workload migration is more fine-grained, dynamically rebalancing workloads across DP replicas according to pipeline progress and residual performance skew. To further explain how ResiHP mitigates the failure amplification in Figure 2, we isolate its impact at each propagation level. ResiHP reduces the delay to 0.64× at the fail-slow node, 2.03× at TP, 8.72× at PP, and 11.14× at DP, showing that progressive adaptation limits failure propagation across dimensions and preserves high training efficiency. Validation of Fail-stop Handling on Convergence. To ensure ResiHP preserves training dynamics, we trained LLaMA 2-7B for

Training Loss

Model Size

Figure 10: Effectiveness of the Scheduler in handling mixed failures.

10

F/stop

Fault-free

ResiHP

F/slow

F/stop

F/slow

1000

1500

2000

5 0

500

Iteration

2500

Figure 12: LLaMA 2-7B training loss: Fault-free baseline (orange) and recovered trajectory by ResiHP (blue).

2,500 iterations under a fault-free baseline and ResiHP with injected failures. Figure 12 shows that both loss curves tightly overlap, sharing identical decay trends and final loss values. While a transient loss spike occurs during the third injected failure, the model immediately recovers because ResiHP focuses on the system level without altering the mathematical semantics of training. Thus, ResiHP can maintain strict convergence and final training quality without disrupting the overall loss trajectory. System Overhead. The system overhead consists of three main components: the Detector identifying failures, the Scheduler generating an adaptation plan, and the reconfiguration process modifying the 3D parallelism dimensions. As Figure 13 demonstrates, the Detector overhead is negligible, increasing the per-iteration time by only 1.2–1.5%. The warm-up profiling for the linear FLOPsto-time model is a one-time cost which is not included here. The Scheduler’s planning overhead scales with model size due to larger PP degrees but remains minimal, taking just 1.44s for the 32B model (less than half a training iteration). Communication group reconstruction overhead is bounded to under 2s across all three models. Lastly, layer transfer overhead during reconfiguration scales with model size and transfer volume, but can be effectively amortized over long-running training.

8.4

Large-Scale Evaluation

To evaluate ResiHP at scale, we train LLaMA 2-70B on 256 NVIDIA A100 GPUs with (𝑇 𝑃, 𝐷𝑃, 𝑃𝑃) = (4, 4, 16). We use a dynamic training scenario with recurring failures and re-joins to evaluate end-toend detection and mitigation.

1

46.2%

0

1.2-4.1%

7B

Time (s)

100.0% Comm. Rebuild Avg. Iter. Time 100.0% 61.2% 86.9% 45.4%

14.6% 1.5%

14B

Model Size

1.5%

32B

15

Time (s)

2

Detector Scheduler 100.0%

3

32B 14B 7B

10 5 0

0

2

4

6

8

Num of Transfered Layers

10

Figure 13: Left: Overhead of ResiHP across Qwen 2.5 models. Right: Overhead of layer transfer during reconfiguration across Qwen 2.5 models. Detection. We first evaluate the Detector. As shown in Figure 14, throughput drops align closely with failure occurrences, and the subsequent rebounds indicate successful failure detection and mitigation. Detector identifies all failures within 2–3 training iterations despite high failure concurrency. When fail-stop and fail-slow failures co-occur, Detector prioritizes fail-stop detection to restore training first; after Scheduler reconfigures the job, Detector resumes detecting the remaining fail-slow failures. Failure Handling. Figure 14 shows that strengthened ReCycle becomes increasingly ineffective as mixed failures accumulate over time, especially around iterations 2000 and 6000. This is because it may reassign workloads from failed DP peers to devices already degraded by fail-slow failures, further turning them into severe bottlenecks that significantly slow the entire cluster. By combining accurate detection with effective mitigation, ResiHP substantially reduces the impact of both fail-stop and fail-slow failures throughout training. In terms of average end-to-end throughput, ResiHP achieves 1.39× and 1.11× speedups over strengthened ReCycle and strengthened Oobleck, respectively.

9

RELATED WORK

System Failure Analysis. Failure analysis has been widely studied in cloud services [4, 8, 9], operating systems [54], and storage systems [14, 31]. LLM training differs by using thousands of expensive GPUs under tightly synchronized execution, where a single failed or slow component can stall the entire job and recovery often requires coordinated cluster-wide reconfiguration. Silent data corruptions (SDCs) represent another important failure class. Unlike fail-stop and fail-slow failures, SDCs may not affect device liveness or execution time, and thus require different signals. Prior works suggest lightweight checks on critical layers during LLM inference [39] and training-level anomalies such as loss spikes or parameter drift [32]. Extending ResiHP with an SDC-specific detector based on these signals could reuse our system-level adaptations to isolate suspicious devices or stages, and we leave this to future work. Data Heterogeneity. Variable-length sequences in LLM training induce workload variability across micro-batches and can further cause load imbalance under parallel execution [12, 22, 43, 44]. Prior work mitigates such heterogeneity through dynamic scheduling [22], adaptive parallelism [44], balanced data assignment [12], and data reordering [55]. However, these techniques do not explicitly stabilize the iteration-time series for failure detection, and residual fluctuations may still trigger spurious alarms. Resilient and Efficient LLM Training. Prior systems improve resilience under failures, preemptions, and resource changes [10, 21, 24, 40, 52]. These systems mainly target fail-stop failures or

Failures (number) Throughput (samples/s)

ResiHP: Taming LLM Training Failures with Dynamic Hybrid Parallelism

6

ResiHP

S_Recycle

S_Oobleck

Fail-slow

Fail-stop

2000

4000

4 2

0 10 0

0

6000

Iteration

8000

10000

Figure 14: Evaluation of ResiHP for a 256 A100 GPU training with fail-stop and fail-slow failures. elasticity events, whereas ResiHP also diagnoses and mitigates failslow performance degradation through progressive adaptations. Some systems automatically search for efficient parallel training plans [2, 3, 16, 30, 56]. ResiHP can benefit from these parallelism optimization strategies to further improve training efficiency.

10

CONCLUSION

This paper presents ResiHP, a system that automatically detects and mitigates both fail-slow and fail-stop failures in large-scale LLM training. ResiHP combines a workload-aware execution time predictor that enables accurate fail-slow detection. Also, ResiHP orchestrates a scheduler that jointly adapts hybrid parallelism via dynamic TP reconfiguration, layer repartition, and adaptive progress-aware workload migration. Our evaluation demonstrates the efficiency and scalability of ResiHP.

References [1] Diego Agudelo-España, Sebastian Gomez-Gonzalez, Stefan Bauer, Bernhard Schölkopf, and Jan Peters. 2020. Bayesian online prediction of change points. In Conference on uncertainty in artificial intelligence. PMLR, 320–329. [2] Sanjith Athlur, Nitika Saran, Muthian Sivathanu, Ramachandran Ramjee, and Nipun Kwatra. 2022. Varuna: scalable, low-cost training of massive deep learning models. In Proceedings of the Seventeenth European Conference on Computer Systems. 472–487. [3] Zhenkun Cai, Xiao Yan, Kaihao Ma, Yidi Wu, Yuzhen Huang, James Cheng, Teng Su, and Fan Yu. 2022. TensorOpt: Exploring the Tradeoffs in Distributed DNN Training With Auto-Parallelism. IEEE Transactions on Parallel and Distributed Systems 33, 8 (2022), 1967–1981. [4] Mike Chow, Yang Wang, William Wang, Ayichew Hailu, Rohan Bopardikar, Bin Zhang, Jialiang Qu, David Meisner, Santosh Sonawane, Yunqi Zhang, et al. 2024. { ServiceLab } : Preventing tiny performance regressions at hyperscale through { Pre-Production } testing. In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24). 545–562. [5] DeepSeek-AI, Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei Li, H. Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J. L. Cai, Jian Liang, Jianzhong Guo, Jiaqi Ni, Jiashi Li, Jiawei Wang, Jin Chen, Jingchang Chen, Jingyang Yuan, Junjie Qiu, Junlong Li, Junxiao Song, Kai Dong, Kai Hu, Kaige Gao, Kang Guan, Kexin Huang, Kuai Yu, Lean Wang, Lecong Zhang, Lei Xu, Leyi Xia, Liang Zhao, Litong Wang, Liyue Zhang, Meng Li, Miaojun Wang, Mingchuan Zhang, Minghua Zhang, Minghui Tang, Mingming Li, Ning Tian, Panpan Huang, Peiyi Wang, Peng Zhang, Qiancheng Wang, Qihao Zhu, Qinyu Chen, Qiushi Du, R. J. Chen, R. L. Jin, Ruiqi Ge, Ruisong Zhang, Ruizhe Pan, Runji Wang, Runxin Xu, Ruoyu Zhang, Ruyi Chen, S. S. Li, Shanghao Lu, Shangyan Zhou, Shanhuang Chen, Shaoqing Wu, Shengfeng Ye, Shengfeng Ye, Shirong Ma, Shiyu Wang, Shuang Zhou, Shuiping Yu, Shunfeng Zhou, Shuting Pan, T. Wang, Tao Yun, Tian Pei, Tianyu Sun, W. L. Xiao, Wangding Zeng, Wanjia Zhao, Wei An, Wen Liu, Wenfeng Liang, Wenjun Gao, Wenqin Yu, Wentao Zhang, X. Q. Li, Xiangyue Jin, Xianzu Wang, Xiao Bi, Xiaodong Liu, Xiaohan Wang, Xiaojin Shen, Xiaokang Chen, Xiaokang Zhang, Xiaosha Chen, Xiaotao Nie, Xiaowen Sun, Xiaoxiang Wang, Xin Cheng, Xin Liu, Xin Xie, Xingchao Liu, Xingkai Yu, Xinnan Song, Xinxia Shan, Xinyi Zhou, Xinyu Yang, Xinyuan Li, Xuecheng Su, Xuheng Lin,

Y. K. Li, Y. Q. Wang, Y. X. Wei, Y. X. Zhu, Yang Zhang, Yanhong Xu, Yanhong Xu, Yanping Huang, Yao Li, Yao Zhao, Yaofeng Sun, Yaohui Li, Yaohui Wang, Yi Yu, Yi Zheng, Yichao Zhang, Yifan Shi, Yiliang Xiong, Ying He, Ying Tang, Yishi Piao, Yisong Wang, Yixuan Tan, Yiyang Ma, Yiyuan Liu, Yongqiang Guo, Yu Wu, Yuan Ou, Yuchen Zhu, Yuduan Wang, Yue Gong, Yuheng Zou, Yujia He, Yukun Zha, Yunfan Xiong, Yunxian Ma, Yuting Yan, Yuxiang Luo, Yuxiang You, Yuxuan Liu, Yuyang Zhou, Z. F. Wu, Z. Z. Ren, Zehui Ren, Zhangli Sha, Zhe Fu, Zhean Xu, Zhen Huang, Zhen Zhang, Zhenda Xie, Zhengyan Zhang, Zhewen Hao, Zhibin Gou, Zhicheng Ma, Zhigang Yan, Zhihong Shao, Zhipeng Xu, Zhiyu Wu, Zhongyu Zhang, Zhuoshu Li, Zihui Gu, Zijia Zhu, Zijun Liu, Zilin Li, Ziwei Xie, Ziyang Song, Ziyi Gao, and Zizheng Pan. 2025. DeepSeek-V3 Technical Report. arXiv:2412.19437 [cs.CL] https://arxiv.org/abs/2412.19437 [6] Gen Dong, Yu Hua, Yongle Zhang, Zhangyu Chen, and Menglei Chen. 2025. Understanding and detecting fail-slow hardware failure bugs in cloud systems. In Proceedings of the 2025 USENIX Conference on Usenix Annual Technical Conference (Boston, MA, USA) (USENIX ATC ’25). USENIX Association, USA, Article 66, 16 pages. [7] Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv e-prints (2024), arXiv–2407. [8] Yu Gan, Mingyu Liang, Sundar Dev, David Lo, and Christina Delimitrou. 2021. Sage: Leveraging ml to diagnose unpredictable performance in cloud microservices. arXiv preprint arXiv:2112.06263 (2021). [9] Yu Gan, Yanqi Zhang, Kelvin Hu, Dailun Cheng, Yuan He, Meghna Pancholi, and Christina Delimitrou. 2019. Seer: Leveraging big data to navigate the complexity of performance debugging in cloud microservices. In Proceedings of the twentyfourth international conference on architectural support for programming languages and operating systems. 19–33. [10] Swapnil Gandhi, Mark Zhao, Athinagoras Skiadopoulos, and Christos Kozyrakis. 2024. Recycle: Resilient training of large dnns using pipeline adaptation. In Proceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles. 211–228. [11] Wei Gao, Yuheng Zhao, Dakai An, Tianyuan Wu, Lunxi Cao, Shaopan Xiong, Ju Huang, Weixun Wang, Siran Yang, Wenbo Su, et al. 2025. Rollpacker: Mitigating long-tail rollouts for fast, synchronous rl post-training. arXiv preprint arXiv:2509.21009 (2025). [12] Hao Ge, Junda Feng, Qi Huang, Fangcheng Fu, Xiaonan Nie, Lei Zuo, Haibin Lin, Bin Cui, and Xin Liu. 2025. ByteScale: Communication-Efficient Scaling of LLM Training with a 2048K Context Length on 16384 GPUs. In Proceedings of the ACM SIGCOMM 2025 Conference. 963–978. [13] Andy Georges, Dries Buytaert, and Lieven Eeckhout. 2007. Statistically rigorous Java performance evaluation. In Proceedings of the 22nd Annual ACM SIGPLAN Conference on Object-Oriented Programming Systems and Applications. [14] Haryadi S Gunawi, Riza O Suminto, Russell Sears, Casey Golliher, Swaminathan Sundararaman, Xing Lin, Tim Emami, Weiguang Sheng, Nematollah Bidokhti, Caitie McCaffrey, et al. 2018. Fail-slow at scale: Evidence of hardware performance faults in large production systems. ACM Transactions on Storage (TOS) 14, 3 (2018), 1–26. [15] Haryadi S. Gunawi, Riza O. Suminto, Russell Sears, Casey Golliher, Swaminathan Sundararaman, Xing Lin, Tim Emami, Weiguang Sheng, Nematollah Bidokhti, Caitie McCaffrey, Deepthi Srinivasan, Biswaranjan Panda, Andrew Baptist, Gary Grider, Parks M. Fields, Kevin Harms, Robert B. Ross, Andree Jacobson, Robert Ricci, Kirk Webb, Peter Alvaro, H. Birali Runesha, Mingzhe Hao, and Huaicheng Li. 2018. Fail-Slow at Scale: Evidence of Hardware Performance Faults in Large Production Systems. ACM Trans. Storage 14, 3, Article 23 (Oct. 2018), 26 pages. doi:10.1145/3242086 [16] Jihu Guo, Tenghui Ma, Wei Gao, Peng Sun, Jiaxing Li, Xun Chen, Yuyang Jin, and Dahua Lin. 2025. AdaPtis: Reducing Pipeline Bubbles with Adaptive Pipeline Parallelism on Heterogeneous Models. arXiv preprint arXiv:2509.23722 (2025). [17] Qinghao Hu, Zhisheng Ye, Zerui Wang, Guoteng Wang, Meng Zhang, Qiaoling Chen, Peng Sun, Dahua Lin, Xiaolin Wang, Yingwei Luo, Yonggang Wen, and Tianwei Zhang. 2024. Characterization of large language model development in the datacenter. In Proceedings of the 21st USENIX Symposium on Networked Systems Design and Implementation (Santa Clara, CA, USA) (NSDI’24). USENIX Association, USA, Article 39, 21 pages. [18] Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Mia Xu Chen, Dehao Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Yonghui Wu, and Zhifeng Chen. 2019. GPipe: efficient training of giant neural networks using pipeline parallelism. Curran Associates Inc., Red Hook, NY, USA. [19] Sam Ade Jacobs, Masahiro Tanaka, Chengming Zhang, Minjia Zhang, Shuaiwen Leon Song, Samyam Rajbhandari, and Yuxiong He. 2023. DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models. arXiv:2309.14509 [cs.LG] https://arxiv.org/abs/2309.14509 [20] Raj Jain. 1991. The Art of Computer Systems Performance Analysis: Techniques for Experimental Design, Measurement, Simulation, and Modeling. Wiley. [21] Insu Jang, Zhenning Yang, Zhen Zhang, Xin Jin, and Mosharaf Chowdhury. 2023. Oobleck: Resilient distributed training of large models using pipeline templates. In Proceedings of the 29th Symposium on Operating Systems Principles. 382–395.

[22] Chenyu Jiang, Zhen Jia, Shuai Zheng, Yida Wang, and Chuan Wu. 2024. DynaPipe: Optimizing multi-task training through dynamic pipelines. In Proceedings of the Nineteenth European Conference on Computer Systems. 542–559. [23] Ziheng Jiang, Haibin Lin, Yinmin Zhong, Qi Huang, Yangrui Chen, Zhi Zhang, Yanghua Peng, Xiang Li, Cong Xie, Shibiao Nong, et al. 2024. { MegaScale } : Scaling large language model training to more than 10,000 { GPUs } . In 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24). 745–760. [24] Xueze Kang, Guangyu Xiang, Yuxin Wang, Hao Zhang, Yuchu Fang, Yuhang Zhou, Zhenheng Tang, Youhui Lv, Eliran Maman, Mark Wasserman, et al. 2025. ElasWave: An Elastic-Native System for Scalable Hybrid-Parallel Training. arXiv preprint arXiv:2510.00606 (2025). [25] Vijay Anand Korthikanti, Jared Casper, Sangkug Lym, Lawrence McAfee, Michael Andersch, Mohammad Shoeybi, and Bryan Catanzaro. 2023. Reducing activation recomputation in large transformer models. Proceedings of Machine Learning and Systems 5 (2023), 341–353. [26] Mario Michael Krell, Matej Kosec, Sergio P. Perez, and Andrew Fitzgibbon. 2022. Efficient Sequence Packing without Cross-contamination: Accelerating Large Language Models without Impacting Performance. arXiv:2107.02027 [cs.CL] https://arxiv.org/abs/2107.02027 [27] Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. 2020. Gshard: Scaling giant models with conditional computation and automatic sharding. arXiv preprint arXiv:2006.16668 (2020). [28] Haoyang Li, Fangcheng Fu, Sheng Lin, Hao Ge, Xuanyu Wang, Jiawen Niu, Jinbao Xue, Yangyu Tao, Di Wang, Jie Jiang, and Bin Cui. 2025. Hydraulis: Balancing Large Transformer Model Training via Co-designing Parallel Strategies and Data Assignment. Proc. ACM Manag. Data 3, 6, Article 337 (Dec. 2025), 30 pages. doi:10.1145/3769802 [29] Jinkun Lin, Ziheng Jiang, Zuquan Song, Sida Zhao, Menghan Yu, Zhanghan Wang, Chenyuan Wang, Zuocheng Shi, Xiang Shi, Wei Jia, et al. 2025. Understanding Stragglers in Large Model Training Using What-if Analysis. arXiv preprint arXiv:2505.05713 (2025). [30] Zhiqi Lin, Youshan Miao, Guanbin Xu, Cheng Li, Olli Saarikivi, Saeed Maleki, and Fan Yang. 2024. Tessel: Boosting Distributed Execution of Large DNN Models via Flexible Schedule Search. In HPCA. [31] Ruiming Lu, Erci Xu, Yiming Zhang, Fengyi Zhu, Zhaosheng Zhu, Mengtian Wang, Zongpeng Zhu, Guangtao Xue, Jiwu Shu, Minglu Li, and Jiesheng Wu. 2023. PERSEUS: a fail-slow detection framework for cloud storage systems. In Proceedings of the 21st USENIX Conference on File and Storage Technologies (Santa Clara, CA, USA) (FAST’23). USENIX Association, USA, Article 4, 15 pages. [32] Jeffrey Jian Ma, Hengzhi Pei, Leonard Lausen, and George Karypis. 2025. Understanding silent data corruption in LLM training. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 20372–20394. [33] Deepak Narayanan, Aaron Harlap, Amar Phanishayee, Vivek Seshadri, Nikhil R Devanur, Gregory R Ganger, Phillip B Gibbons, and Matei Zaharia. 2019. PipeDream: Generalized pipeline parallelism for DNN training. In Proceedings of the 27th ACM symposium on operating systems principles. 1–15. [34] Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley, Mostofa Patwary, Vijay Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, Amar Phanishayee, and Matei Zaharia. 2021. Efficient large-scale language model training on GPU clusters using megatronLM. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (St. Louis, Missouri) (SC ’21). Association for Computing Machinery, New York, NY, USA, Article 58, 15 pages. doi:10.1145/3458817.3476209 [35] OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff Belgum, Irwan Bello, Jake Berdine, Gabriel Bernadett-Shapiro, Christopher Berner, Lenny Bogdonoff, Oleg Boiko, Madelaine Boyd, Anna-Luisa Brakman, Greg Brockman, Tim Brooks, Miles Brundage, Kevin Button, Trevor Cai, Rosie Campbell, Andrew Cann, Brittany Carey, Chelsea Carlson, Rory Carmichael, Brooke Chan, Che Chang, Fotis Chantzis, Derek Chen, Sully Chen, Ruby Chen, Jason Chen, Mark Chen, Ben Chess, Chester Cho, Casey Chu, Hyung Won Chung, Dave Cummings, Jeremiah Currier, Yunxing Dai, Cory Decareaux, Thomas Degry, Noah Deutsch, Damien Deville, Arka Dhar, David Dohan, Steve Dowling, Sheila Dunning, Adrien Ecoffet, Atty Eleti, Tyna Eloundou, David Farhi, Liam Fedus, Niko Felix, Simón Posada Fishman, Juston Forte, Isabella Fulford, Leo Gao, Elie Georges, Christian Gibson, Vik Goel, Tarun Gogineni, Gabriel Goh, Rapha Gontijo-Lopes, Jonathan Gordon, Morgan Grafstein, Scott Gray, Ryan Greene, Joshua Gross, Shixiang Shane Gu, Yufei Guo, Chris Hallacy, Jesse Han, Jeff Harris, Yuchen He, Mike Heaton, Johannes Heidecke, Chris Hesse, Alan Hickey, Wade Hickey, Peter Hoeschele, Brandon Houghton, Kenny Hsu, Shengli Hu, Xin Hu, Joost Huizinga, Shantanu Jain, Shawn Jain, Joanne Jang, Angela Jiang, Roger Jiang, Haozhun Jin, Denny Jin, Shino Jomoto, Billie Jonn, Heewoo Jun, Tomer

ResiHP: Taming LLM Training Failures with Dynamic Hybrid Parallelism

Kaftan, Łukasz Kaiser, Ali Kamali, Ingmar Kanitscheider, Nitish Shirish Keskar, Tabarak Khan, Logan Kilpatrick, Jong Wook Kim, Christina Kim, Yongjik Kim, Jan Hendrik Kirchner, Jamie Kiros, Matt Knight, Daniel Kokotajlo, Łukasz Kondraciuk, Andrew Kondrich, Aris Konstantinidis, Kyle Kosic, Gretchen Krueger, Vishal Kuo, Michael Lampe, Ikai Lan, Teddy Lee, Jan Leike, Jade Leung, Daniel Levy, Chak Ming Li, Rachel Lim, Molly Lin, Stephanie Lin, Mateusz Litwin, Theresa Lopez, Ryan Lowe, Patricia Lue, Anna Makanju, Kim Malfacini, Sam Manning, Todor Markov, Yaniv Markovski, Bianca Martin, Katie Mayer, Andrew Mayne, Bob McGrew, Scott Mayer McKinney, Christine McLeavey, Paul McMillan, Jake McNeil, David Medina, Aalok Mehta, Jacob Menick, Luke Metz, Andrey Mishchenko, Pamela Mishkin, Vinnie Monaco, Evan Morikawa, Daniel Mossing, Tong Mu, Mira Murati, Oleg Murk, David Mély, Ashvin Nair, Reiichiro Nakano, Rajeev Nayak, Arvind Neelakantan, Richard Ngo, Hyeonwoo Noh, Long Ouyang, Cullen O’Keefe, Jakub Pachocki, Alex Paino, Joe Palermo, Ashley Pantuliano, Giambattista Parascandolo, Joel Parish, Emy Parparita, Alex Passos, Mikhail Pavlov, Andrew Peng, Adam Perelman, Filipe de Avila Belbute Peres, Michael Petrov, Henrique Ponde de Oliveira Pinto, Michael, Pokorny, Michelle Pokrass, Vitchyr H. Pong, Tolly Powell, Alethea Power, Boris Power, Elizabeth Proehl, Raul Puri, Alec Radford, Jack Rae, Aditya Ramesh, Cameron Raymond, Francis Real, Kendra Rimbach, Carl Ross, Bob Rotsted, Henri Roussez, Nick Ryder, Mario Saltarelli, Ted Sanders, Shibani Santurkar, Girish Sastry, Heather Schmidt, David Schnurr, John Schulman, Daniel Selsam, Kyla Sheppard, Toki Sherbakov, Jessica Shieh, Sarah Shoker, Pranav Shyam, Szymon Sidor, Eric Sigler, Maddie Simens, Jordan Sitkin, Katarina Slama, Ian Sohl, Benjamin Sokolowsky, Yang Song, Natalie Staudacher, Felipe Petroski Such, Natalie Summers, Ilya Sutskever, Jie Tang, Nikolas Tezak, Madeleine B. Thompson, Phil Tillet, Amin Tootoonchian, Elizabeth Tseng, Preston Tuggle, Nick Turley, Jerry Tworek, Juan Felipe Cerón Uribe, Andrea Vallone, Arun Vijayvergiya, Chelsea Voss, Carroll Wainwright, Justin Jay Wang, Alvin Wang, Ben Wang, Jonathan Ward, Jason Wei, CJ Weinmann, Akila Welihinda, Peter Welinder, Jiayi Weng, Lilian Weng, Matt Wiethoff, Dave Willner, Clemens Winter, Samuel Wolrich, Hannah Wong, Lauren Workman, Sherwin Wu, Jeff Wu, Michael Wu, Kai Xiao, Tao Xu, Sarah Yoo, Kevin Yu, Qiming Yuan, Wojciech Zaremba, Rowan Zellers, Chong Zhang, Marvin Zhang, Shengjia Zhao, Tianhao Zheng, Juntang Zhuang, William Zhuk, and Barret Zoph. 2024. GPT-4 Technical Report. arXiv:2303.08774 [cs.CL] https://arxiv.org/abs/2303.08774 [36] Biswaranjan Panda, Deepthi Srinivasan, Huan Ke, Karan Gupta, Vinayak Khot, and Haryadi S. Gunawi. 2019. IASO: A Fail-Slow Detection and Mitigation Framework for Distributed Storage Services. In 2019 USENIX Annual Technical Conference (USENIX ATC 19). USENIX Association, Renton, WA, 47–62. https: //www.usenix.org/conference/atc19/presentation/panda [37] Penghui Qi, Xinyi Wan, Guangxing Huang, and Min Lin. 2024. Zero bubble (almost) pipeline parallelism. In The Twelfth International Conference on Learning Representations. [38] Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019. Megatron-lm: Training multi-billion parameter language models using model parallelism. arXiv preprint arXiv:1909.08053 (2019). [39] Yu Sun, Zhu Zhu, Cherish Mulpuru, Roberto Gioiosa, Zhao Zhang, Bo Fang, and Lishan Yang. 2025. Ft2: First-token-inspired online fault tolerance on critical layers for generative large language models. In Proceedings of the 34th International Symposium on High-Performance Parallel and Distributed Computing. 1–14. [40] John Thorpe, Pengzhan Zhao, Jonathan Eyolfson, Yifan Qiao, Zhihao Jia, Minjia Zhang, Ravi Netravali, and Guoqing Harry Xu. 2023. Bamboo: Making Preemptible Instances Resilient for Affordable Training of Large DNNs. In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23). USENIX Association, Boston, MA, 497–513. https://www.usenix.org/conference/ nsdi23/presentation/thorpe [41] Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288 (2023). [42] Leslie G. Valiant. 1990. A bridging model for parallel computation. Commun. ACM 33, 8 (Aug. 1990), 103–111. doi:10.1145/79173.79181 [43] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2023. Attention Is All You Need. arXiv:1706.03762 [cs.CL] https://arxiv.org/abs/1706.03762 [44] Yujie Wang, Shiju Wang, Shenhan Zhu, Fangcheng Fu, Xinyi Liu, Xuefeng Xiao, Huixia Li, Jiashi Li, Faming Wu, and Bin Cui. 2025. FlexSP: Accelerating Large Language Model Training via Flexible Sequence Parallelism. arXiv:2412.01523 [cs.DC] https://arxiv.org/abs/2412.01523 [45] Zheng Wang, Anna Cai, Xinfeng Xie, Zaifeng Pan, Yue Guan, Weiwei Chu, Jie Wang, Shikai Li, Jianyu Huang, Chris Cai, Yuchen Hao, and Yufei Ding. 2025. WLB-LLM: workload-balanced 4D parallelism for large language model training. In Proceedings of the 19th USENIX Conference on Operating Systems Design and Implementation (Boston, MA, USA) (OSDI ’25). USENIX Association, USA, Article 43, 17 pages. [46] BigScience Workshop, Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, et al. 2022. Bloom: A 176b-parameter open-access multilingual

language model. arXiv preprint arXiv:2211.05100 (2022). [47] Tianyuan Wu, Lunxi Cao, Hanfeng Lu, Xiaoxiao Jiang, Yinghao Yu, Siran Yang, Guodong Yang, Jiamang Wang, Lin Qu, Liping Zhang, et al. 2025. Adaptra: Straggler-Resilient Hybrid-Parallel Training with Pipeline Adaptation. arXiv preprint arXiv:2504.19232 (2025). [48] Tianyuan Wu, Wei Wang, Yinghao Yu, Siran Yang, Wenchao Wu, Qinkai Duan, Guodong Yang, Jiamang Wang, Lin Qu, and Liping Zhang. 2025. { GREYHOUND } : Hunting { Fail-Slows } in { Hybrid-Parallel } Training at Scale. In 2025 USENIX Annual Technical Conference (USENIX ATC 25). 731–747. [49] Yifan Xiong, Yuting Jiang, Ziyue Yang, Lei Qu, Guoshuai Zhao, Shuguang Liu, Dong Zhong, Boris Pinzur, Jie Zhang, Yang Wang, Jithin Jose, Hossein Pourreza, Jeff Baxter, Kushal Datta, Prabhat Ram, Luke Melton, Joe Chau, Peng Cheng, Yongqiang Xiong, and Lidong Zhou. 2024. SuperBench: improving cloud AI infrastructure reliability with proactive validation. In Proceedings of the 2024 USENIX Conference on Usenix Annual Technical Conference (Santa Clara, CA, USA) (USENIX ATC’24). USENIX Association, USA, Article 51, 16 pages. [50] An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jing Zhou, Jingren Zhou, Junyang Lin, Kai Dang, Keqin Bao, Kexin Yang, Le Yu, Lianghao Deng, Mei Li, Mingfeng Xue, Mingze Li, Pei Zhang, Peng Wang, Qin Zhu, Rui Men, Ruize Gao, Shixuan Liu, Shuang Luo, Tianhao Li, Tianyi Tang, Wenbiao Yin, Xingzhang Ren, Xinyu Wang, Xinyu Zhang, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yinger Zhang, Yu Wan, Yuqiong Liu, Zekun Wang, Zeyu Cui, Zhenru Zhang, Zhipeng Zhou, and Zihan Qiu. 2025. Qwen3 Technical Report. arXiv:2505.09388 [cs.CL] https://arxiv.org/abs/2505.09388 [51] An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jianxin Yang, Jin Xu, Jingren Zhou, Jinze Bai, Jinzheng He, Junyang Lin, Kai Dang, Keming Lu, Keqin Chen, Kexin Yang, Mei Li, Mingfeng Xue, Na Ni, Pei Zhang, Peng Wang, Ru Peng, Rui Men, Ruize Gao, Runji Lin, Shijie Wang, Shuai Bai, Sinan Tan, Tianhang Zhu, Tianhao Li, Tianyu Liu, Wenbin Ge, Xiaodong Deng, Xiaohuan Zhou, Xingzhang Ren, Xinyu Zhang, Xipin Wei, Xuancheng Ren, Xuejing Liu, Yang Fan, Yang Yao, Yichang Zhang, Yu Wan, Yunfei Chu, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, Zhifang Guo, and Zhihao Fan. 2024. Qwen2 Technical Report. arXiv:2407.10671 [cs.CL] https://arxiv.org/abs/2407.10671 [52] Zhisheng Ye, Wei Gao, Qinghao Hu, Peng Sun, Xiaolin Wang, Yingwei Luo, Tianwei Zhang, and Yonggang Wen. 2024. Deep Learning Workload Scheduling in GPU Datacenters: A Survey. ACM Comput. Surv. 56, 6, Article 146 (Jan. 2024), 38 pages. [53] Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2022. OPT: Open Pre-trained Transformer Language Models. arXiv:2205.01068 [cs.CL] https://arxiv.org/abs/ 2205.01068 [54] Shenglin Zhang, Yongxin Zhao, Xiao Xiong, Yongqian Sun, Xiaohui Nie, Jiacheng Zhang, Fenglai Wang, Xian Zheng, Yuzhi Zhang, and Dan Pei. 2024. Illuminating the gray zone: Non-intrusive gray failure localization in server operating systems. In Companion Proceedings of the 32nd ACM International Conference on the Foundations of Software Engineering. 126–137. [55] Zili Zhang, Yinmin Zhong, Yimin Jiang, Hanpeng Hu, Jianjian Sun, Zheng Ge, Yibo Zhu, Daxin Jiang, and Xin Jin. 2025. DistTrain: Addressing Model and Data Heterogeneity with Disaggregated Training for Multimodal Large Language Models. In Proceedings of the ACM SIGCOMM 2025 Conference (São Francisco Convent, Coimbra, Portugal) (SIGCOMM ’25). Association for Computing Machinery, New York, NY, USA, 24–38. doi:10.1145/3718958.3750472 [56] Lianmin Zheng, Zhuohan Li, Hao Zhang, Yonghao Zhuang, Zhifeng Chen, Yanping Huang, Yida Wang, Yuanzhong Xu, Danyang Zhuo, Joseph E. Gonzalez, and Ion Stoica. 2022. Alpa: Automating Inter- and Intra-Operator Parallelism for Distributed Deep Learning. In OSDI.

Record · ID 168264 · SHA-256 5e503a8612c6de18
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.