O NLINE D RAFT C O -T RAINING FOR S PECULATIVE D ECODING IN L ARGE -S CALE , L ONG -C ONTEXT RL P OST-T RAINING
arXiv:2609.07108v1 [cs.LG] 7 Sep 2026
Zili Wang Zhaopeng Qiu Yuekai Zhang Shuang Yu Junjie Lai NVIDIA {ziliw, alexq, yuekaiz, shuangy, julienl}@nvidia.com
A BSTRACT Abstract. Speculative decoding accelerates rollout generation, which dominates the cost of reinforcement learning (RL) post-training. Online co-training can further increase the draft’s accuracy, yielding greater speedups. However, scaling this approach to co-training on large models with long contexts poses two obstacles: (1) branch attention is unsupported by standard causal context-parallel (CP) implementations, and (2) target features span across pipeline-parallel (PP) stages. We address both with an end-to-end system for large-scale online draft co-training. For CP, we extend packed, load-balanced zigzag ring attention by merging rank-local branch attention with causal main-sequence attention. For PP, TapChannel transports intermediate target features across stages via a separate path, leaving the pipeline schedule unaffected. Experiments demonstrate that co-trained drafts closely track the policy baseline while delivering substantial rollout and end-to-end speedups across model scales up to 122B. Our CP design achieves strong scaling at 256K tokens with significant memory savings over prior work, and our PP transport incurs modest overhead. Code can be found here. 1
I NTRODUCTION
Reinforcement learning (RL) post-training has become a standard paradigm for building reasoning and agentic Large Language Models (LLMs) (Shao et al., 2024; Yu et al., 2025; GLM-5-Team et al., 2026). The wall-clock time of RL post-training is often dominated by rollout generation. Speculative decoding (SD) mitigates this through parallel verification (Leviathan et al., 2023; Chen et al., 2023): a small draft model generates multiple draft tokens, and the target policy verifies them in parallel, accelerating autoregressive generation without changing the output distribution. Recent RL frameworks have therefore integrated SD into their rollout engines, demonstrating that rollout speedups can translate into end-to-end training acceleration (Zhang et al., 2026; Chen et al., 2026b; Liu et al., 2025; Iso et al., 2026; Kim et al., 2026; veRL Team, 2026; Zhu et al., 2025). Online co-training can further extend the draft model’s acceptance length as the RL policy evolves, yielding greater rollout speedups (Zhao, 2024; Zhang et al., 2026; Chen et al., 2026b; Wang et al., 2026a). However, scaling this approach to large-model, long-context RL post-training is non-trivial. Advanced drafts such as EAGLE-3, DFlash, and
DSpark introduce branch-structured attention and consume intermediate hidden states from the target model (Li et al., 2026c; Chen et al., 2026a; Cheng et al., 2026). Standard causal context parallelism (CP) does not support this branch structure, while under pipeline parallelism (PP), the required target features may reside on stages remote from the draft. Consequently, the draft cannot simply inherit the system’s existing parallelism configuration. We address this by designing two mechanisms that extend the system’s existing CP and PP parallelism configuration to support large-scale draft co-training. For CP, each branch query attends to two key sets: a causal prefix of the main sequence and the keys of the branch tokens. We compute the causal component with packed zigzag-ring attention, handle the branch-local component on the owning rank, and merge the two results. This unified mechanism supports EAGLE-3, DFlash, and DSpark. For PP, we introduce TapChannel, which transports intermediate target features across distributed PP stages on a side path independent of pipeline communication, leaving the pipeline schedule unchanged. Together, these two mechanisms yield a complete system design that makes online draft co-training practical for RL post-training on large models with long contexts.
Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training
Experiments across draft families, targets at various scale, and single- and multi-turn RL workloads demonstrate stable learning and scalable system performance. Co-trained drafts closely track the baseline in reward and accuracy, yielding 1.50–1.88× end-to-end speedups. Our CP implementation outperforms USP by up to 2.9× in latency with a 2.7× reduction in per-GPU memory, achieving scaling at 256K tokens, while TapChannel enables online co-training with modest PP overhead. Our contributions are summarized as follows: • Branch attention under CP. We decompose draft branch attention into a causal main-sequence component and a rank-local branch component, and merge both results to recover the attention output. This unified mechanism supports EAGLE-3, DFlash, and DSpark within a packed, load-balanced zigzag-ring execution. • Out-of-schedule target-feature transport under PP. We introduce TapChannel, which delivers intermediate target features across distributed PP stages without entering or altering the pipeline schedule. • End-to-end draft co-training at scale. We integrate our CP and PP mechanisms into NeMo-RL (nem, 2025) framework, enabling online draft co-training on large models with long contexts. We evaluate our system across three draft families, target model sizes from 8B to 122B, and both single- and multi-turn RL workloads, measuring learning stability, rollout speedups, and system overhead.
2
R ELATED W ORK
Speculative decoding for RL rollouts. Speculative decoding accelerates generation by letting a lightweight draft propose several tokens and using the large target model to verify them in parallel; rejection sampling preserves the target model’s output distribution, so the orocedure does not change the output distribution (Leviathan et al., 2023; Chen et al., 2023). Recent systems adapt speculative decoding to the rollout engine in different ways. Nemo-RL (nem, 2025) integrate MTP and EAGLE-3 drafting into both synchronous and asynchronous RL rollouts and study how deployment and draft configurations affect end-to-end training speed (Iso et al., 2026). SPEC-RL avoids a learned drafter altogether: it reuses response segments from the preceding policy iteration as speculative prefixes and verifies them under the current policy, retaining on-policy samples while exploiting cross-iteration similarity (Liu et al., 2025). EfficientRollout instead constructs a quantized self-drafter from the target, adjusts speculative length according to observed acceptance, and enables speculation only when the runtime is likely to benefit (Kim et al., 2026). These works establish that speculative decoding can accelerate RL in practice,
but primarily optimize draft sourcing, serving configuration, and verification on the rollout side. Our focus is complementary: continuously training target-feature-conditioned drafts inside the distributed policy learner. Draft adaptation under evolving RL policies. A fixed draft loses alignment with an evolving policy; adapting it during RL recovers acceptance length and speedup. FastGRPO combines online draft learning with concurrencyaware configuration (Zhang et al., 2026); ReSpec dynamically selects speculative parameters and distills the policy into the drafter (Chen et al., 2026b). Parallel work adapts MTP modules rather than separate draft models: MTPRL uses a parameter-sharing MTP module with advantageaware optimization (Wang et al., 2026a); OCC derives an adaptive coefficient to balance auxiliary loss against policy update (Wang et al., 2026b); Bebop trains directly for acceptance using total-variation objectives and studies pre-RL adaptation (Li et al., 2026b). These methods improve draft adaptation at the objective level; we address the systems challenge of making co-training practical under CP and PP. Target-feature-conditioned and parallel draft architectures. Draft architectures trade proposal quality for sequential cost. Blockwise parallel decoding introduced multiple future-token predictors with prefix validation (Stern et al., 2018). MTP generalizes this as an auxiliary objective (Gloeckle et al., 2024); FastMTP reduces cost with a shared head and recursive conditioning (Cai et al., 2025). EAGLE-3 conditions a separate drafter on fused target-layer features and uses TTT to expose multi-step errors (Li et al., 2026c). DFlash produces candidate blocks in parallel via block diffusion (Chen et al., 2026a); DSpark combines a parallel backbone with a Markov head for dependency restoration (Cheng et al., 2026); Domino similarly separates parallel modeling from sequential dependency (Huang et al., 2026a). These architectures differ in training branches, attention masks, and feature interfaces. Integrating these draft models into current distributed training systems remains a challenge. Large-model training combines tensor, pipeline, and other parallelism (Shoeybi et al., 2019; Yan et al., 2026). As for long-context training, RingAttention distributes long sequences and overlaps KV communication with computation (Liu et al., 2024). These methods support causal attention training but do not handle draft-specific branch attention or cross-stage feature routing. P-EAGLE parallelizes EAGLE with structured masks (Hui et al., 2026); LongSpec addresses longcontext inference with bounded KV cache and positional adaptation (Yang et al., 2026). SpecForge develops target–draft decoupling and hybrid-parallel TTT training, including sequence-parallel branch attention (Li et al., 2026a). Our setting differs from these works. The target is an evolv-
Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training
c0
packed main sequence · zigzag chunks c1
c2
c3
c4
c5
c6
c7
branch rank 3 (no communication)
rank 0
rank 1
rank 2
rank 3
main-seq K/V c0 + c7
main-seq K/V c1 + c6
main-seq K/V c2 + c5
main-seq K/V c3 + c4
branches
branches
branches
branches
K/V circulate · C ring steps
(a) K/V ring over zigzag chunks; branches stay on their rank
one attention call on rank r queries stay resident
main-seq part C ring steps
branch part 1 local step
merge output
(b) Merging the two parts
Figure 1. Illustration of branch attention under context parallelism. (a) The causal prefix is zigzag-sharded across CP ranks; branch-local keys stay local. (b) Each rank merges the ring attention result with local branch attention via Eq. (2).
ing RL policy trained with existing CP and PP parallelism. Unlike previous work that modifies the parallel layout to accommodate draft training, we keep the target’s topology unchanged and instead adapt the system around it.
3
M ETHOD
We first describe the online co-training procedure (§3.1) and then present the two mechanisms that support it: branch attention under context parallelism (§3.2) and TapChannel under pipeline parallelism (§3.3). 3.1
Online Draft Co-Training in RL Post-Training
Let θ and ϕ denote the policy and draft parameters. The draft is trained on the same policy rollout tokens, with a stop-gradient applied to the target features (Hθ ) collected from intermediate policy layers. The joint objective is L(θ, ϕ) = LRL (θ) + λ Ldraft ϕ; x, sg(Hθ (x)) , (1) where x denotes the rollout tokens and sg denotes stopgradient. The draft model is instantiated as a submodule on the policy’s last pipeline stage and updated jointly along with the policy model. 3.2
Branch Attention under Context Parallelism
Standard causal context parallelism does not directly support the branch-structured attention introduced by draft training. SpecForge (Li et al., 2026a) addresses EAGLE-3 TTT (Train-Time Test) setting but relies on sequential ring sharding, which yields imbalanced causal workloads, and imposes constraints on the Ulysses dimension that can be restrictive for draft KV heads. As shown in Figure 1, our design organizes each branch query with two key sets: a causal prefix of the main sequence, sharded across CP ranks, and a small set of branch-local keys, kept on the rank that owns the branch’s anchor. The two key sets are attended to
independently and then merged. The main-sequence component follows the same packed zigzag-ring attention as standard CP. For each ring step, the main-sequence K/V circulate across ranks while queries remain local; the locally computed branch component is then merged via Eq. (2). The merge follows the standard online-softmax reduction used between ring-attention steps: ℓ = log eℓm + eℓb , O = eℓm −ℓ Om + eℓb −ℓ Ob , (2) where (Om , ℓm ) and (Ob , ℓb ) correspond to main-sequence and branch, respectively. O denotes the attention output and ℓ denotes the log-sum-exp for the main-sequence and branch-local components. Our method supports draft families with different branch structures. (1) EAGLE-3 (Li et al., 2026c) employs TrainingTime Test (TTT): during training, it simulates multi-step autoregressive generation by feeding its own previous predicted hidden state back as inputs over multiple steps. This creates a separate branch at every draft position, each attending to a growing context of previously generated tokens within the draft. (2) DFlash (Chen et al., 2026a) uses block diffusion language model to generate an entire block of tokens in a single forward pass. DSpark (Cheng et al., 2026) extends DFlash with a lightweight Markov head that refines the block left-to-right with a low-rank, previous-token-conditioned bias, restoring causal dependencies among block positions at small additional cost. Despite their different generation strategies, all three architectures share the same requirement: each branch, whether a TTT position or a block position, must attend to both the causal prefix of the main sequence and its branch-local KV context. Communication cost and overlap. For a CP degree C and N main-sequence tokens evenly split across ranks, each rank holds N/C tokens. During the forward ring, each rank sends its local K and V to the other C − 1 ranks, yielding a
Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training
pipeline P2P
TapChannel (RDMA)
stage 0
stage 1
stage 2
stage 3
low tap
mid tap
no tap to send
draft model
(outside the schedule)
from stage 0 from stage 1
pre-allocated slots stage 3 GPU
Figure 2. Illustration of TapChannel feature fan-in under PP=4. Tap-producing stages deliver target features to the draft stage over an out-of-schedule path, bypassing the standard adjacent-stage P2P communication (gray). Cross-node sources use GPUDirect RDMA; colocated sources use CUDA IPC.
per-rank outbound volume of N dkv b, (3) C where dkv is the per-token K/V width and b the bytes per element. Since branch-local K/V stay on their anchor-owning ranks, this cost is independent of the number or depth of branches. The backward pass replays the same K/V ring and accumulates gradients locally, incurring no additional branch-dependent traffic. fwd VCP = 2(C − 1)
We overlap communication with attention at ring-step granularity: the exchange for the next shard is issued before the current attention step and waited on only at the subsequent boundary. The exposed overhead at step r is thus (r) (r) ECP = max t(r) t (4) comm attn Since attention computation scales quadratically with context length while communication scales linearly, longer contexts increase the overlap headroom. Conversely, aggressive strong scaling shortens the local context and can eventually expose communication as the bottleneck. 3.3
TapChannel under Pipeline Parallelism
To synchronize this producer-consumer handshake without interfering with pipeline schedule, each slot carries a sequence stamp that both the source and the draft increment on each write and read. For colocated sources (CUDA IPC), the draft clears the stamp after reading. For cross-node sources (dedicated NCCL communicator with GPUDirect RDMA), the stamp primarily orders operations; buffer reuse is managed separately by bounded in-flight sends and receiver-side CUDA events. Communication cost and overlap. Let n be the number of token rows in a microbatch and S the set of source stages that send features to the draft stage. Let ds be the per-token feature dimension produced by source s. For the first stage, ds equals the input embedding width; for later stages, it equals the hidden size. Let b be the bytes per element. The feature payload delivered per microbatch is X VTap = bn ds . (5) s∈S
Each feature is transferred directly once, regardless of the number of PP hops between its source and the draft; features produced on the draft stage itself require no transfer.
Under pipeline parallelism, the layers producing the taps (target hidden states) span multiple stages, while the draft resides only on the last stage. Standard pipeline communication only connects adjacent stages and cannot deliver these non-adjacent features. Since taps require no return path, TapChannel transports them on a side path independent of the pipeline schedule (Figure 2).
The pipeline schedule naturally provides a slack window between the time a source stage produces its tap and the draft stage’s forward for the same microbatch. If the tap transfer completes within this window, it adds no extra latency. For microbatch m, let ∆s,m denote this slack for source s, and τs,m the transfer time. The residual rendezvous delay is given by:
TapChannel implements this side path with a per-source mailbox on the draft stage. Each source has a pre-allocated buffer slot in the draft stage’s memory. After a source finishes its policy forward for a microbatch, it writes the resulting taps into its slot; the draft stage reads that slot right before its own forward for the same microbatch.
ETap = max max (τs,m , ∆s,m ) ,
(m)
s∈S
(6)
Only the portion of a transfer that exceeds the slack incurs visible overhead. We empirically characterize these overheads and their implications for end-to-end training performance in Section 4.5.
Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training Baseline Reward
0.8
w/ DSpark ×10−4
0.0
60
KL Divergence
Accuracy (%)
0.2
50 40
−0.2
30
−0.4
20
Acceptance Length
Training-Inference KL
7.0
70
0.4
6.8 6.6 6.4 6.2 6.0
Rollout Throughput
×104
End-to-End Step Time 450
3.5
3.5
400
3.0
2.5
2.0
3.0
s/step
Tokens/s
Accepted Tokens / Verification
w/ DFlash
Val Accuracy
0.6
Reward
w/ EAGLE-3
2.5
350 300
2.0
250
1.5
200 150
0
50
100
150
200
0
50
Step
100
Step
150
200
0
50
100
150
200
Step
Figure 3. Experimental results of Qwen3-8B online co-training with EAGLE-3 (yellow), DFlash (red) and DSpark (blue) on DAPOMath17K, including reward, accuracy evaluated on AIME2024, KL divergence, accepted length, throughput and end-to-end step time.
4
E XPERIMENTS
We evaluate the correctness and efficiency of our draft cotraining implementation. Specifically, we verify: (1) that our implementation preserves the RL learning trajectory; (2) speculative decoding gains across draft architectures, target scales, and single/multi-turn tasks; (3) CP attention efficiency at long sequences; and (4) the incremental cost of PP-side co-training. 4.1
ment within NVIDIA NeMo Gym, designed for complex task execution in simulated office scenarios. It is based on the WorkBench benchmark (Styles et al., 2024), a simulated workplace tool-use environment with email, calendar, CRM, project management, and analytics databases. Experiments of larger models, including Nemotron-3.5Lightning-30B-A3B (TP/PP/CP/EP=2/2/2/8), Qwen3.5122B-A10B (TP/PP/CP/EP=4/4/2/16), and GPT-OSS 120B (TP/PP/CP/EP=4/4/2/16), are conducted on GB200 GPU.
Experimental Setup
Models. We evaluate three representative, advanced draft families: EAGLE-3 (Li et al., 2026c), DFlash (Chen et al., 2026a), and DSpark (Cheng et al., 2026). Qwen3-8B (Team, 2025) serves as the target model for the three draft families. We also evaluate on larger models, including Qwen3.5-35BA3B and Qwen3.5-122B-A10B (Team, 2026) with DFlash, Nemotron-3.5-Lightning-30B-A3B (NVIDIA, 2025) with DSpark, and GPT-OSS-120B with DFlash, to test scalability and generality. Draft models are from official checkpoints.
Tasks and training. We conduct experiments in NemoRL (nem, 2025) with GRPO (Shao et al., 2024) on DAPOMath-17K (Yu et al., 2025) and evaluate on AIME 2024 (AoPS, 2024), with 4,096 input and 16,384 response on H100 GPU (for Qwen3-8B, TP/PP/CP=2/2/2 and for Qwen3.5-35B-A3B, TP/PP/CP/EP=2/2/2/8). For multi-turn evaluation, we adapt the NeMo Gym Workplace Assistant (Styles et al., 2024) on Qwen3-8B with sequence length of 32,768 with tool feedback. Workplace Assistant is a multi-turn, agentic tool-use environ-
Metrics. We report training reward, validation accuracy, and training–inference KL divergence, which measures the per-token KL divergence between the log-probability distributions produced by the training and inference backends on the generated responses. This metric quantifies the numerical consistency between the two backends under identical weights; a near-zero value confirms that the implementation correctly synchronizes the inference engine with the training policy, preserving the on-policy assumption critical for GRPO. For speculative decoding, we report acceptance length, which is the mean number of tokens accepted per verification. Rollout throughput is also reported, with rollout speedup relative to the baseline and end-to-end speedup w.r.t the total policy-update time. 4.2
Verifying Policy Learning under Speculative Decoding
We select Qwen3-8B as the target for all three draft models. We compare four runs: baseline without neither online draft co-training nor rollour speculative decoding, and online cotraining with EAGLE-3, DFlash, and DSpark. As shown
Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training Baseline Reward
w/ EAGLE-3
w/ DFlash
Acceptance Length 7
0.80
0.60 0.55 0.50 0.45
140 5
3.25 3.00 2.75 2.50
0
10
20
30
40
4 3
120 110
2
100
2.25 1
2.00
0.40
130
s/step
0.65
6
3.50
Tokens/s
Accepted Tokens / Verify
Reward
0.70
End-to-End Step Time 150
3.75
0.75
×10
w/ DSpark
Rollout Throughput
4
0
10
Step
20
30
40
90 0
10
20
Step
30
40
0
10
Step
20
30
40
Step
Figure 4. Qwen3-8B online co-training with EAGLE-3 (yellow), DFlash (red), and DSpark (blue) on NeMo Gym Workplace Assistant. Table 1. End-to-end performance across target scales and draft families. Target Model
Draft Model
Accepted Length ↑
Rollout Speedup ↑
E2E Training Speedup ↑
Qwen3-8B
EAGLE-3 DFlash DSpark
2.28 3.45 3.63
1.63× 2.23× 2.18×
1.50× 1.88× 1.83×
Qwen3.5-35B-A3B Nemotron-3.5-Lightning-30B-A3B
DFlash DSpark
4.58 2.65
1.50× 1.19×
1.46× 1.16×
Qwen3.5-122B-A10B GPT-OSS-120B
DFlash DFlash
4.78 3.80
1.72× 1.48×
1.35× 1.19×
4.3
Performance across Drafts and Model Scales
Table 1 summarizes end-to-end results across three draft families and targets from 8B to 122B. Co-trained drafts reach 2.28–4.78 acceptance length, yielding 1.19–2.23× rollout speedup and 1.16–1.88× end-to-end training speedup. Comparing the three draft models, DFlash and DSpark consistently outperform EAGLE-3 in acceptance length, suggesting better speedup. Notably, while the larger MoE targets (Qwen3.5-122B and GPT-OSS-120B) achieve high acceptance lengths, their end-to-end speedups are lower, since each verification forward invokes sparse routing more expert compute (Huang et al., 2026b). For linearattention, since the verification step itself is comparatively less expensive, the relative speedup from speculation is inherently smaller (Wang et al., 2026c). Multi-Turn Workloads Workplace Assistant interleaves multiple model turns with tool calls and environment delays. As shown in Figure 4, in our four-way comparison on this multi-turn workload, all configurations achieve consistent rewards, while acceptance length improves monotonically across training. However, the end-to-end speedup (1.25–1.43×) falls well below the rollout-phase speedup
USP (Ulysses × ring) attention, padded batches
Fwd+bwd (ms)
CP = 2 600
400
563
400
2.9×
200 0 30 20 10 0
196 u = 2 u = 1 Ours r=1 r=2
26.3
Zigzag ring attention, packed batches (ours)
CP = 4
801
800
Peak HBM/GPU (GB)
in Figure 3, the reward, validation, and training-inference KL divergence results confirm that speculative decoding preserves the RL learning trajectory. The co-trained runs closely track the baseline learning trajectory across all three kind of draft models.
0
2.7× 8.3 u = 2 u = 1 Ours r=1 r=2
10 5 0
CP = 8
479
300 200
287
2.3×
200
15
22.2
427
126 u = 4 u = 2 u = 1 Ours r=1 r=2 r=4
13.2 12.7
11.3
4.2 u = 4 u = 2 u = 1 Ours r=1 r=2 r=4
254
287
146
1.5×
100 0 8 6
2.7×
216
99 u = 8 u = 4 u = 2 u = 1 Ours r=1 r=2 r=4 r=8
6.6
6.3
6.4
4
2.7×
2 0
5.6
2.1 u = 8 u = 4 u = 2 u = 1 Ours r=1 r=2 r=4 r=8
Figure 5. Comparison of packed zigzag-ring attention (ours) vs. padded USP attention under the same workload. Top: latency of the forward+backward pass; bottom: per-GPU peak memory. Each bracket marks the advantage over the best USP configuration.
(1.75–2.23×), because rollout accounts for only 55.8% of the step time—tool execution and environment latency sit inside this phase and are unreachable by faster decoding. Against the single-turn results in Table 1, where the same drafts reach 1.50–1.88×, this comparison isolates what the multi-turn structure costs. 4.4
Context-Parallel Attention Performance
We benchmark our packed zigzag-ring attention against the USP implementation from SpecForge (Li et al., 2026a) under matched GPU counts. The comparison focuses on EAGLE-3 TTT (3 passes) at the attention-operator level. We report forward/backward latency and per-GPU peak HBM.
CP=2
CP=8
120
DSpark ( = 7, N = 64)
104 102
102
103 102
Peak HBM/GPU (GB)
Fwd+bwd (ms)
CP=4
DFlash ( = 15, N = 64)
26
16K 32K 64K 128K 256K
101
16K 32K 64K 128K 256K
101
24
22
22
22
20
20
20
16K 32K 64K 128K 256K
Sequence length
Sequence length
16K 32K 64K 128K 256K
Sequence length
Figure 6. Scaling of draft-training attention over CP=1, 2, 4, 8. Each group fixes the global sequence length and reports forward+backward latency (top) and peak HBM (bottom).
Table 2. Pipeline-parallel overhead for Qwen3-8B, averaged over first 10 policy update steps. DSpark achieves the largest end-to-end speedup (1.85×) with 14.6% update-time overhead. Rollout
Improvement
Accepted ↑ Time Speedup ↑ Time Overhead ↓ Tap wait E2E ↑
Draft None EAGLE-3 DFlash DSpark
— 1.89 2.84 3.37
329.2 238.7 229.2 156.8
— 1.38× 1.44× 2.10×
31.5 42.4 35.8 36.1
— 34.3% 13.6% 14.6%
— 0.37 0.57 0.54
1.00× 1.31× 1.37× 1.85×
Analysis Figure 5 compares our packed zigzag attention against USP under the same variable-length workload (longest 20,480 tokens). USP pads the batch to 2.25× the real token count; our packed implementation avoids this overhead. At CP=2, 4, and 8, packed zigzag outperforms the best USP variant by 2.9×, 2.3×, and 1.5× in latency, with 2.7× lower per-GPU peak memory.
Long-context scaling. We evaluate our implementation across CP∈ {1, 2, 4, 8} at fixed global lengths up to 256K tokens (Figure 6). For TTT attention, latency drops from 17.7 s at CP=1 to 2.35 s at CP=8 (7.5×, 94% parallel efficiency), and per-GPU memory falls nearly linearly from 53.2 GB to 7.5 GB. Block-draft attention is launch-bound rather than FLOP-bound, which works in its favor: it stays cheap even at CP=8 (≤55 ms from 16K to 256K), yet scales 6.9× at 256K when the workload is substantial. Context parallelism thus rides the policy’s layout nearly for free. 4.5
Pipeline-Parallel Overhead
We measure the incremental cost of adding TapChannel to an existing PP policy, comparing the full method against the same pipeline-parallel policy with draft training disabled. We evaluate three configurations on Qwen3-8B: EAGLE-3, DFlash, and DSpark. We compare against a host-staging baseline that stages taps through pinned host memory.
Host-staged TapChannel (one-sided)
100
109 77
80 60
49
40 20 0
16K 32K 64K 128K 256K
22
22 16K 32K 64K 128K 256K
Fan-in latency (ms)
CP=1
EAGLE-3 TTT
10
23 2.3
7
4
E3-8B 4K tok
E3-8B 16K tok
DF-8B 4K tok
13 DF-8B 16K tok
16
3
DF-35B 4K tok
10 DF-35B 16K tok
Compute slowdown (%)
Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training 100 80
+88.1
+83.2
+85.0
+82.5
-0.1 Stage 1
+0.1
+1.6
-0.1 Stage 0
Stage 2
Draft stage
60 40 20 0
Figure 7. TapChannel vs. host staging on a PP=4 fan-in. Left: fan-in latency; right: slowdown during transfers, measured as the time increased when transfers run concurrently. One-sided writes complete 4.5–8.5× faster than host staging, leave source stages within noise, and cost the receiving draft stage only 1.6% in HBM contention; host staging slows every rank by 83–88%.
Transport microbenchmark. We first measure TapChannel’s raw transport cost on a PP=4 fan-in using the real tap layouts from the EAGLE-3 and DFlash checkpoints. Results are shown in Figure 7. One-sided writes achieve 27–39 GB/s and complete the fan-in 4.5–8.5× aster than staging through pinned host memory. Critically, TapChannel does not sit on the critical path: side-stream transfers introduce negligible interference on source stages and incur only 1.6% contention on the receiving draft stage, whereas host staging slows every rank by >80%.
Full-run overhead. We report the PP overhead results in Table 2. Draft co-training adds modest time overhead (block drafts under 15%, EAGLE-3 at 34% due to its TTT passes), while speculative rollout reduces generation time by 28–52%, yielding net speedups of 1.31–1.85×. The Tap wait column further shows why this overhead does not translate into proportional slowdown: the draft stage waits only 0.4–0.6s per policy update for taps to arrive, or 1.5–2.2% of optimization time, confirming that rendezvous cost is largely overlapped by normal pipeline scheduling.
5
C ONCLUSION
We presented a system for co-training target-featureconditioned speculative drafts under CP and PP in longcontext RL post-training. Two mechanisms enable this: packed zigzag-ring attention for branch-structured CP and TapChannel for cross-stage feature transport outside the pipeline schedule, supporting online co-training of EAGLE3, DFlash, and DSpark. Experiments show that all three drafts preserve the RL learning trajectory while delivering consistent end-to-end speedups across target scales; the CP decomposition scales at long sequences with lower memory usage, and the PP overhead is modest. Future work includes tailoring speculative decoding to sparse MoE and linearattention models, where its current benefits are limited.
Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training
R EFERENCES Nemo rl: A scalable and efficient post-training library. https://github.com/NVIDIA-NeMo/RL, 2025. GitHub repository. AoPS. Aime 2024 dataset, 2024. URL https: //artofproblemsolving.com/wiki/index. php/2024_AIME_I,II. Cai, Y., Liang, X., Wang, X., Ma, J., Liang, H., Luo, J., Zuo, X., Duan, L., Yin, Y., and Chen, X. Fastmtp: Accelerating llm inference with enhanced multi-token prediction. arXiv preprint arXiv:2509.18362, 2025. Chen, C., Borgeaud, S., Irving, G., Lespiau, J.-B., Sifre, L., and Jumper, J. Accelerating large language model decoding with speculative sampling. arXiv preprint arXiv:2302.01318, 2023. Chen, J., Liang, Y., and Liu, Z. Dflash: Block diffusion for flash speculative decoding. arXiv preprint arXiv:2602.06036, 2026a. Chen, Q., Liu, Z., Sun, P., Li, S., Wang, G., Liu, Z., Wen, Y., Feng, S., and Zhang, T. Respec: Towards optimizing speculative decoding in reinforcement learning systems. Proceedings of Machine Learning and Systems, 8:367– 379, 2026b. Cheng, X., Yu, X., Shao, C., Li, J., Xiong, Y., Qian, Y., Zhu, J., Ma, S., Zhang, X., Ye, J., et al. Dspark: Confidencescheduled speculative decoding with semi-autoregressive generation. arXiv preprint arXiv:2607.05147, 2026. GLM-5-Team, :, Zeng, A., Lv, X., Hou, Z., Du, Z., Zheng, Q., Chen, B., Yin, D., Ge, C., Huang, C., Xie, C., Zhu, C., Yin, C., Wang, C., Pan, G., Zeng, H., Zhang, H., Wang, H., Chen, H., Zhang, J., Jiao, J., Guo, J., Wang, J., Du, J., Wu, J., Wang, K., Li, L., Fan, L., Zhong, L., Liu, M., Zhao, M., Du, P., Dong, Q., Lu, R., Shuang-Li, Cao, S., Liu, S., Jiang, T., Chen, X., Zhang, X., Huang, X., Dong, X., Xu, Y., Wei, Y., An, Y., Niu, Y., Zhu, Y., Wen, Y., Cen, Y., Bai, Y., Qiao, Z., Wang, Z., Wang, Z., Zhu, Z., Liu, Z., Li, Z., Wang, B., Wen, B., Huang, C., Cai, C., Yu, C., Li, C., Hu, C., Zhang, C., Zhang, D., Lin, D., Yang, D., Wang, D., Ai, D., Zhu, E., Yi, F., Chen, F., Wen, G., Sun, H., Zhao, H., Hu, H., Zhang, H., Liu, H., Zhang, H., Peng, H., Tai, H., Zhang, H., Liu, H., Wang, H., Yan, H., Ge, H., Liu, H., Chu, H., Zhao, J., Wang, J., Zhao, J., Ren, J., Wang, J., Zhang, J., Gui, J., Zhao, J., Li, J., An, J., Li, J., Yuan, J., Du, J., Liu, J., Zhi, J., Duan, J., Zhou, K., Wei, K., Wang, K., Luo, K., Zhang, L., Sha, L., Xu, L., Wu, L., Ding, L., Chen, L., Li, M., Lin, N., Ta, P., Zou, Q., Song, R., Yang, R., Tu, S., Yang, S., Wu, S., Zhang, S., Li, S., Li, S., Fan, S., Qin, W., Tian, W., Zhang, W., Yu, W., Liang, W., Kuang, X., Cheng, X., Li, X., Yan, X.,
Hu, X., Ling, X., Fan, X., Xia, X., Zhang, X., Zhang, X., Pan, X., Zou, X., Zhang, X., Liu, Y., Wu, Y., Li, Y., Wang, Y., Zhu, Y., Tan, Y., Zhou, Y., Pan, Y., Zhang, Y., Su, Y., Geng, Y., Yan, Y., Tan, Y., Bi, Y., Shen, Y., Yang, Y., Li, Y., Liu, Y., Wang, Y., Li, Y., Wu, Y., Zhang, Y., Duan, Y., Zhang, Y., Liu, Z., Jiang, Z., Yan, Z., Zhang, Z., Wei, Z., Chen, Z., Feng, Z., Yao, Z., Chai, Z., Wang, Z., Zhang, Z., Xu, B., Huang, M., Wang, H., Li, J., Dong, Y., and Tang, J. Glm-5: from vibe coding to agentic engineering, 2026. URL https://arxiv.org/abs/2602.15763. Gloeckle, F., Idrissi, B. Y., Rozière, B., Lopez-Paz, D., and Synnaeve, G. Better & faster large language models via multi-token prediction. arXiv preprint arXiv:2404.19737, 2024. Huang, J., Zhang, Y., Zhang, Q., Lin, H., Xu, H., and Zhang, L. Domino: Decoupling causal modeling from autoregressive drafting in speculative decoding. arXiv preprint arXiv:2605.29707, 2026a. Huang, Z., Zhu, L., Zhan, Z., Hu, T., Mao, W., Yu, X., Liu, Y., and Zhang, T. Moesd: Unveil speculative decoding’s potential for accelerating sparse moe. Advances in Neural Information Processing Systems, 38:125276–125311, 2026b. Hui, M., Huang, X., Salas, J. C., Sun, Y., Pemberton, N., Song, X., Khetan, A., and Karypis, G. P-eagle: Paralleldrafting eagle with scalable training. arXiv preprint arXiv:2602.01469, 2026. Iso, H., Mitra, T., Mondal, S., Shafipour, R., Elango, V., Kong, T., Huang, Y., Na, S., Putterman, I., Chislett, B., et al. Accelerating rl post-training rollouts via system-integrated speculative decoding. arXiv preprint arXiv:2604.26779, 2026. Kim, M., Lee, M., Oh, S., Galim, K., Kim, D., Hooper, C., Singh, H., Gholami, A., Koo, H. I., and Kang, W. Efficientrollout: System-aware self-speculative decoding for rl rollouts. arXiv preprint arXiv:2606.18967, 2026. Leviathan, Y., Kalman, M., and Matias, Y. Fast inference from transformers via speculative decoding. In International Conference on Machine Learning, pp. 19274– 19286. PMLR, 2023. Li, S., Wang, C., Zhu, Y., Wang, Y., Yin, F., Shi, S., Chen, Y., Dong, X., Chen, Q., Pan, J., et al. Specforge: A flexible and efficient open-source training framework for speculative decoding. arXiv preprint arXiv:2603.18567, 2026a. Li, Y., Jiang, H., Xu, Y., Yang, J., Zhang, Y., Cao, Y., Shen, Y., Zhou, F., Men, R., Zhang, J., et al. Breaking entropy bounds: Accelerating rl training via mtp with rejection sampling. arXiv preprint arXiv:2606.12370, 2026b.
Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training
Li, Y., Wei, F., Zhang, C., and Zhang, H. Eagle-3: Scaling up inference acceleration of large language models via training-time test. Advances in Neural Information Processing Systems, 38:136737–136756, 2026c.
Wang, Z., Chai, J., Chen, L., Wang, X., Xiang, S., and Yin, G. Joint training of multi-token prediction in reinforcement learning via optimal coefficient calibration. arXiv preprint arXiv:2605.28184, 2026b.
Liu, B., Wang, A., Min, Z., Yao, L., Zhang, H., Liu, Y., Zeng, A., and Su, J. Spec-rl: Accelerating on-policy reinforcement learning via speculative rollouts. 2025.
Wang, Z., Han, X., Yang, Z., Liu, F., Li, X., Gu, R., Zhong, S., and Tian, C. Specla: Efficient speculative decoding for linear-attention models. arXiv preprint arXiv:2607.16673, 2026c.
Liu, H., Zaharia, M., and Abbeel, P. Ringattention with blockwise transformers for near-infinite context. In International Conference on Learning Representations, volume 2024, pp. 3992–4008, 2024. NVIDIA. Nvidia nemotron 3: Efficient and open intelligence, 2025. URL https://arxiv.org/abs/ 2512.20856. White Paper.
Yan, Z., Bai, H., Yao, X., Liu, D., Liu, T., Liu, H., Li, P., Wu, E., Fan, S., Tao, L., et al. Scalable training of mixtureof-experts models with megatron core. arXiv preprint arXiv:2603.07685, 2026.
Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Bi, X., Zhang, H., Zhang, M., Li, Y., Wu, Y., et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024.
Yang, P., Du, C., Zhang, F., Wang, H., Pang, T., Du, C., and An, B. Longspec: Long-context lossless speculative decoding with efficient drafting and verification. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 1826–1844, 2026.
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B. Megatron-lm: Training multibillion parameter language models using model parallelism. arXiv preprint arXiv:1909.08053, 2019.
Yu, Q., Zhang, Z., Zhu, R., Yuan, Y., Zuo, X., Yue, Y., Dai, W., Fan, T., Liu, G., Liu, L., et al. Dapo: An open-source llm reinforcement learning system at scale. arXiv preprint arXiv:2503.14476, 2025.
Stern, M., Shazeer, N., and Uszkoreit, J. Blockwise parallel decoding for deep autoregressive models. Advances in Neural Information Processing Systems, 31, 2018.
Zhang, Y., Lv, N., Wang, T., and Dang, J. Fastgrpo: Accelerating policy optimization via concurrency-aware speculative decoding and online draft learning. In International Conference on Learning Representations, volume 2026, pp. 59620–59636, 2026.
Styles, O., Miller, S., Cerda-Mardini, P., Guha, T., Sanchez, V., and Vidgen, B. Workbench: a benchmark dataset for agents in a realistic workplace setting. arXiv preprint arXiv:2405.00823, 2024. Team, Q. Qwen3 technical report, 2025. URL https: //arxiv.org/abs/2505.09388. Team, Q. Qwen3.5: Towards native multimodal agents, February 2026. URL https://qwen.ai/blog? id=qwen3.5. veRL Team. Multi-token prediction in verl. https://verl.readthedocs.io/en/latest/advance/mtp.html, 2026. Accessed: 2026-02-15. Wang, K., Zeng, A., Du, Z., Hu, Y., Zhang, B., Wang, X., Tang, J., and Zhang, J. MTP-RL: Acceleration of reinforcement learning rollouts with policy-aligned multi-token prediction. In Liakata, M., Moreira, V. P., Zhang, J., and Jurgens, D. (eds.), Findings of the Association for Computational Linguistics: ACL 2026, pp. 37530–37542, San Diego, California, United States, July 2026a. Association for Computational Linguistics. ISBN 979-8-89176-395-1. doi: 10.18653/v1/2026.findings-acl. 1871. URL https://aclanthology.org/2026. findings-acl.1871/.
Zhao, C. Power up speculative decoding in reinforcement learning. https: //github.com/zhaochenyang20/ Awesome-ML-SYS-Tutorial/blob/main/ rlhf/slime/spec/readme-en.md, 2024. Part of the Awesome-ML-SYS-Tutorial repository. Zhu, Z., Xie, C., Lv, X., and slime Contributors. slime: An llm post-training framework for rl scaling. https:// github.com/THUDM/slime, 2025. GitHub repository. Corresponding author: Xin Lv.