HyDra: Demystifying and Taming Dynamic Context
Parallelism at Production Scale Zihao Fan†‡ , Yunzhuo Liu‡ , Bo Jiang† , Changgang Zheng§ , Lin Zheng‡ , Ray Ying‡ , Key Zhang‡ † Shanghai Jiao Tong University ‡ Tencent Hy § Nanjing University
arXiv:2609.34318v1 [cs.DC] 28 Sep 2026
Abstract
attention (GQA) [1], where the few KV heads keep the exchange cheap. Mainstream models now use multi-head latent attention (MLA) [6–8], which materializes K and V for every head during training, so the same exchange moves far more data and performance degrades sharply at large scale. What remains is the DCP shipped in Megatron-Core (Mcore) [24], which is today the only implementation usable at production scale. Its scheduler, however, sizes the CP group of each sample from a memory budget, namely the largest number of tokens a rank can hold. Memory grows linearly with sequence length while attention computation grows quadratically, so a CP size chosen to fit memory leaves the computation imbalanced, and the imbalance worsens as sequences get longer. We confirm this on a 256K-context training job, running Mcore DCP across more than 11K GPUs in our model production (§3). Per-rank token counts are comparable, yet microbatch execution times differ by 5× and CP-group loads by 2.4×. This imbalance leads to a 46.0% pipeline (PP) bubble and a data-parallel (DP) bubble that adds up to 12.8% to every iteration. Planning costs another 9%, because the scheduler rescans the entire pool of sequences on every iteration. We present HyDra, a load-driven DCP system that keeps both planning and execution efficient at production scale. (1) For execution, HyDra derives one per-rank load target for the whole iteration from the quadratic attention computation, and sizes every sequence’s CP degree to meet it, pulling every rank toward the same load and shrinking both the PP and DP bubble. The target is the lowest one the longest sequence can meet, which also makes it the easiest for short sequences to reach. Long sequences reach that target only at large CP degrees, which far exceed the upper bound of attention-head count in Ulysses. Ring does support such large CP degrees, but its per-rank traffic does not shrink as the degree grows, so communication becomes the bottleneck, preventing long sequences from reaching the target. HyDra therefore factorizes each CP degree into an inner Ulysses group inside a shallow outer ring, and overlaps the exposed all-to-alls with computation. A higher degree now lowers both computation and communication per rank, so even the longest sequences can reach the load target. (2) For planning, HyDra derives these quantities in closed form and places sequences with heaps. Iteration time is a graph of waits, each a maximum over ranks, so it has no closed form. Balance removes those maxima and leaves one unknown, the largest CP degree allowed, in a convex trade-off between the
Long-context training runs on sequences whose lengths span orders of magnitude, and dynamic context parallelism (DCP) gives each sequence its own CP degree. Existing DCP systems either do not scale or perform poorly on mainstream models, leaving Megatron-Core (Mcore) DCP as the only option at production scale. Mcore DCP, however, sizes each degree to fit memory, which grows linearly with length while attention grows quadratically, so comparable token counts hide unequal computation. In our production 256K-context training job on more than 11K GPUs under Mcore DCP, per-rank microbatch times differ by up to 5× at similar token counts. The skew leads to a 46% pipeline bubble and a 13% data-parallel bubble. We present HyDra, a scalable load-driven DCP system. Its scheduler balances computation by pulling every rank toward one load target, and balance in turn makes that target solvable in closed form. It places sequences with lazy heaps rather than whole-pool scans. That balance asks for CP degrees larger than memory requires, so its CP engine nests an inner Ulysses group in a shallow outer ring, letting a higher degree lower computation and communication together. Evaluation at both scales shows consistent gains. On a 512-GPU testbed, HyDra raises throughput over Mcore DCP by 1.18× on average at 32K context and 2.48× at 256K. On a 2,048-GPU production job, it shrinks the pipeline bubble from 36% to 14%, cuts scheduling time by 2.3×, and raises throughput by 1.10–1.43× (avg. 1.25×) over Mcore DCP and 1.33–1.90× (avg. 1.59×) over static CP.
1
Introduction
Long-context capabilities are increasingly important for frontier language models [2–5, 27, 29, 43]. The data used for longcontext training, however, do not form a uniform stream of long sequences. Real corpora mix short documents with a small fraction of very long ones, so sequence lengths span orders of magnitude [13, 41, 42]. Dynamic context parallelism (DCP) is an effective way to train on such data, because it adapts the context-parallel (CP) degree [15, 28] to each sequence instead of fixing it for the whole job [13, 41]. In practice, however, every existing DCP system falls short in a different way [13, 17, 23, 24, 41]. The first limitation is scalability. FlexSP [41], Hydraulis [23], and DCP [17] all plan every iteration by solving a global optimization problem, which becomes prohibitively expensive at large scale. The second limitation is performance. ByteScale [13] does scale, but its ring-based CP [28] is designed for grouped-query 1
Fan et al. Code / Progr. (43.7%)
2
Background
2.1
5D Parallelism for Distributed LLM Training
Percentage (%)
pipeline bubble and the per-MB overhead. HyDra therefore solves that trade-off in closed form, and the load target and MB count follow from the degree by substitution. Placement is the only cost left, and a lazy heap per CP degree turns its whole-pool scans into logarithmic lookups. We implement HyDra in approximately 3K LoC of Python on Megatron-LM [37]. On a 512-GPU testbed, it improves end-to-end throughput over Mcore DCP by 1.09–1.25× (avg. 1.18×) at 32K context and 1.42–3.08× (avg. 2.48×) at 256K, keeping the compute stream busy over 76% of each iteration at both lengths. In production 256K-context training on 2,048 GPUs, it improves throughput by 1.10–1.43× (avg. 1.25×) over Mcore DCP and 1.33–1.90× (avg. 1.59×) over static CP. It shrinks the pipeline bubble from 36.4% to 13.6% of each iteration and cuts scheduling time by 2.3×, rising to 11.6× at 16K GPUs, while its training-loss curves track static CP over roughly 8,000 iterations. In simulation, it averages 1.28× over the state-of-the-art (SOTA) ByteScale [13] on a GQA model and 1.53× on a communication-heavy MLA model. This paper makes the following contributions: • We characterize production Mcore DCP on more than 11K GPUs, where up to a 46.0% PP bubble, a 12.8% DP bubble, and a 9% scheduling stall dominate: balancing tokens instead of attention leaves MB times 5× and CPgroup loads 2.4× apart. • We propose HyDra, a load-driven DCP system that collapses the PP and DP bubbles with one load target. HyDra co-designs planning with execution, so the CP degrees that target demands stay affordable at production scale. • We design a scheduler that derives the load target in closed form and places sequences with lazy heaps rather than whole-pool scans. We break the ring–Ulysses dilemma by running a Ulysses group inside a shallow ring, so a larger degree lowers computation and communication together. • We implement HyDra in Megatron-LM and evaluate it at 32K and 256K context: on a 512-GPU testbed it raises throughput over Mcore DCP by 1.09–1.25× (avg. 1.18×) at 32K and 1.42–3.08× (avg. 2.48×) at 256K, on 2,048 production GPUs by 1.10–1.43× (avg. 1.25×) over Mcore DCP and 1.33–1.90× (avg. 1.59×) over static CP while cutting scheduling time by 2.3×; in simulation to 40K GPUs it leads the SOTA baseline by 1.28–1.53×.
21.0 20
Web / Knowledge (40.2%)
22.3
15.8
15.1
12.4
Structured (2%) Forum / QA (1.7%)
10
Instruction / SFT (0.09%)
7.2 3.2
0
STEM / Math (12.4%)
Topic by token
64
128 256 512
1K
2K
4K
1.6 0.8 0.3 0.1 0.1
8K 16K 32K 64K 128K256K
Document length (tokens, unpacked)
Figure 1. Profile of the 256K-context training dataset. through all-to-all [11, 21]; and data parallelism (DP) replicates the model, shards the batch, and synchronizes gradients [26, 35]. Context parallelism (CP) shards each sequence across ranks, so each rank holds only a fraction of its activations and long contexts become feasible [10, 15, 22, 25, 28]. Because attention couples all tokens, CP circulates key and value blocks around a ring or switches sequence and head shardings through all-to-all [15, 28]. This work fixes TP, PP, and EP at their deployment settings and jointly manages CP and DP to place variable-length sequences across ranks. 2.2
Long-Context Data Characteristics
Long-context corpora have highly skewed sequence-length distributions [13, 41]. In the benchmark dataset used for our measurements (Figure 1), lengths span four orders of magnitude up to 256K tokens, with more than 85% of sequences below 2K and fewer than 0.6% above 32K. The rare long ones are what teach long-context ability, so every batch mixes them with abundant short ones [29]. To fill each context window, systems pack short sequences together and mask attention across their boundaries [9, 20, 38], while ragged-tensor compilation and mini-sequence execution cut padding and non-attention memory further [12, 30]. Packing equalizes token counts, not attention work, Í which scales as 𝑖 O (𝑠𝑖2 ), so packed samples of the same size can differ widely in computation. 2.3
Dynamic Context Parallelism
DCP gives every sequence its own CP degree instead of one degree for the whole job [13, 17, 41]. Mcore ships a production scheduler that sizes each degree from a token budget, the most tokens a rank can hold, so every sequence goes to the fewest ranks that fit it [24]. Figure 2 shows the rule on four ranks with a 4K budget, where a 16K sequence 1 × 16K
2 × 8K 4K
Sorted Seq.
Large-scale LLM training spans five parallelism dimensions, which planners combine under model, memory, and topology constraints [16, 18, 32, 37, 44]. Tensor parallelism (TP) shards operators within a node [36, 37]; pipeline parallelism (PP) splits layers into microbatched stages, trading bubbles against activation memory [14, 31]; expert parallelism (EP) shards mixture-of-experts (MoE) weights and routes tokens
4 × 4K 4K … …
HDP
… … MB 0
MB 1
MB 2
Figure 2. DCP under a 4K per-rank token budget. 2
DP 0 DP 1 DP 2 DP 3
R ep l
ica s
HyDra: Demystifying and Taming Dynamic Context Parallelism at Production Scale
PP Bubbles
PP Bubbles
…
P
u PB
bb
le P
u PB
bb
PP Bubbles
le
P
u PB
bb
Bu DP
bb
le
le
…
3
8.0
7.8
7.4 8
3.8
3.6 0.69
0.39
4.1
3.9 0.21
2.6 0.03
0.10
4.1 2.6 0.03
3.8
3.9 0.21
3.6 0.64
2
4
6.2 19
5.4 22
4.8 12
6.8 47
45
6.2 77
27 5.2
4.9
175
6.4 98
92 4.8
392
206
6.1
10
0.10
0.0
100
0.38
3.0
5.0
671 6.4
6.4
367
1000
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 Microbatch
1 0.1 0.01
out around it. Under DCP each MB packs a different mix of variable-length sequences at a different CP degree, and in this example the heaviest takes more than 5× as long to execute as the lightest. We call this skew inter-MB imbalance; its footprint is the PP bubble, the green gaps in the figure, which consume 45.97% of this pipeline, nearly half of every iteration. A balanced schedule at 𝑝𝑝=9 with three interleaved chunks would bubble only about 10% at this MB count, so skew inflates it 4.6×. The imbalance traces to a mismatch between how load is partitioned and how it scales. Figure 4 profiles each MB in this iteration with two per-rank quantities. The first is the token count, which grows linearly with length and is what the scheduler equalizes. The second is the attention intensity Í of a rank, I = 𝑖 𝑠𝑖2 /𝑅, the quadratic work of the sequences {𝑠𝑖 } in the MB divided over the 𝑅 ranks that carry them. The two diverge sharply. Token counts stay of the same order, averaging around 5K, while intensity spans four orders of magnitude, from ≈0.1 M to ≈1,000 M. Equal tokens need not imply equal computation, which has motivated FLOP-aware packing and scheduling [13, 42]; our point is that the skew survives production DCP, whose scheduler equalizes tokens and memory per rank but never the quadratic attention cost.
Characterizing DCP at Production Scale
Measurement Setup and Methodology
All measurements in this section are collected from a training job of Hy4-preview [40], an open-weight MoE model with 770B total and 49B activated parameters, running on more than 11K GPUs at 256K context length. The model is parallelized with 𝑡𝑝 = 2, 𝑝𝑝 = 9, 𝑒𝑝 = 32, while the remaining ranks form the DP and CP dimensions managed by Mcore DCP [24], whose CP communication uses Ulyssesstyle all-to-all [15]. Traces are collected with Argus [46], a lightweight, always-on tracing system. 3.2
6.0
Intensity / rank
Figure 4. Per-rank profile of each MB.
In this section, we measure Mcore DCP at production scale and identify its major bottlenecks. 3.1
Tokens / rank
9.0
4.8
Tokens Per Rank (k)
spans all four, an 8K sequence takes two, and a 4K sequence stays on one. Short sequences thus avoid the redundant communication a job-wide degree would force on them. CP and DP partition the same batch in two directions, CP within one sequence and DP across sequences, so a static job organizes its ranks as a DP×CP mesh sized for the longest sequence. Per-sequence degrees leave no fixed mesh, so the DP×CP ranks become one pool, the hybrid data-parallel (HDP) domain, from which each group is cut at placement time [13, 24]. In Figure 2, the same four ranks serve as one group of four in MB 0, as two groups of two in MB 1, and as four groups of one in MB 2. In our production training, Mcore DCP improves throughput over static CP by about 8% on 32K data and 39% on 256K data. The budget, however, counts tokens, and Section 3 shows that this leaves quadratic attention work badly imbalanced.
Intensity Per Rank (M)
Figure 3. Training timeline of one iteration under Mcore DCP, across nine PP stages and DP replicas.
Inter-MB Imbalance
Figure 3 plots one iteration across all nine pipeline stages, with the forward and backward passes of every MB over three interleaved chunks and the closing optimizer step. The heaviest MB bounds the iteration. Iteration time follows the critical path through this schedule: every stage must process every MB, so the longest-running one recurs stage after stage and dictates the iteration latency. Figure 3 traces one iteration of this job, where the heaviest MB, the first in the schedule, stretches every stage and idle time fans
Insight 1: Under Mcore DCP, per-rank tokens are balanced yet attention computation intensity varies by orders of magnitude, opening a 5× MB execution-time gap and widening the PP bubble from its 10% floor to 45.97%. MBs must therefore be balanced by attention computation, not tokens alone. 3
0
2
Mic
4
6
rob a
tch
8
320 10
0