2026-09-29
Q WEN G YRE: An Elastic Reinforcement Learning Framework for Training XLong-Horizon Agents Weiqi Wang1,2,* Yi Zhang1 Jiemin Jiang1
Yuxin Zhou1,* Yuyan Luo1 Wentao Yao1
Mouxiang Chen1,* Zhiyu Yin1 Chujie Zheng1
Siyuan Zhang1,3 Chencan Wu1 JianWei Zhang1,†
1 Alibaba Token Hub, Alibaba Group 2 University of Science and Technology of China 3 Tsinghua University
Large language model (LLM) agents increasingly undertake eXtreme-long (XLong) horizon tasks, where a single execution can span hours, hundreds of model–environment interactions, and nearly 1M tokens per rollout. Applying online reinforcement learning (RL) to such executions poses two fundamental challenges: (1) severe execution variance and prolonged rollout delays cause massive GPU idling; and (2) complex non-linear branching generates massive trajectory redundancy, crippling training efficiency. To address these, we presents Q WEN G YRE, an end-to-end framework for XLong-horizon online RL. Q WEN G YRE elastically reallocates GPUs between rollout and training without interrupting live executions, while its trajectory processor reconstructs branching histories, scores partial progress, and deduplicates redundant paths to bound training costs. Scaled to our flagship model, Qwen 3.8 2.4T, with 700K tokens per rollout, Q WEN G YRE yields a 6.0% absolute gain on NL2RepoBench (52.5% → 58.5%) in 48 steps. Across our evaluations on diverse domains of training datasets, Q WEN G YRE delivers up to 1.85× and 1.78× speedups over Colocate and Async, respectively.
(a) Elastic Scheduler
(b) Trajectory Processor
Colocate
Branching agentic trajectories
0
Step 3
50
Step 2
Step 1
100
Model call Tool response New context
GPU utilization (%)
arXiv:2609.33848v1 [cs.LG] 27 Sep 2026
Abstract
Async 100
Step 1
50
Step 2
Expand all paths
Step 3
0
QwenGyre 100
Step 1
Step 2 Step 3
Rollout
40
60
Training
C D F
A
B
C
A
B
D
Training samples
0 20
B B E
Next burst
50 0
A A A
80
≤ k paths
100
Training targets
Idle
Masked
Figure 1: Q WEN G YRE addresses two challenges of XLong-horizon agentic RL. (a) The Elastic Scheduler reallocates idle rollout capacity to training while continuing to serve ongoing executions. (b) The Trajectory Processor converts branching executions into bounded training samples while masking repeated targets in shared prefixes.
* Equal contribution. † Corresponding author.
1
1
Introduction
while preventing the sheer volume of paths and tokens from overwhelming the training pipeline.
Large language model (LLM) agents increasingly solve software engineering tasks that are long-horizon: completing one task requires an extended sequence of interdependent decisions and environment interactions over a persistent workspace. A representative class of such tasks is issue resolution (e.g., SWE-bench (Jimenez et al., 2024; Deng et al., 2026)), which localizes a single fix within an existing repository and spans tens of interactions. On these workloads, online reinforcement learning (RL) has demonstrated promising gains by iteratively refining policies through environment interactions (Luo et al., 2025a; Golubev et al., 2025; Song et al., 2026b; Du et al., 2026b).
In this paper, we present Q WEN G YRE, an end-to-end framework for online RL through black-box agent harnesses at XLong-horizons. Its design separates the lifetimes of harness executions, GPU roles, and training samples. Harness state persists across GPU role changes, while recorded model calls are materialized into training samples on demand. To address Challenge 1, Q WEN G YRE introduces an elastic scheduler, which adjusts GPU allocation between rollout and training according to the amount of unfinished rollout work while ensuring that model calls from ongoing harness executions continue to be served. Centralized data-parallel training allows newly available nodes to join an ongoing batch as the training pool grows. Streaming training distributes work according to each participating data-parallel group’s progress, helping the groups finish the same batch at similar times to minimize idle gaps.
As foundation models rapidly advance in reasoning and context capabilities, the paradigm is shifting toward fully autonomous, end-to-end software development, such as repository generation (Ding et al., 2026), system reimplementation (Yang et al., 2026), and codebase-wide refactoring (Hong et al., 2026). These workloads extend long-horizon execution to multi-hour, project-scale runs. A single stateful rollout spans hundreds of dependent model–environment interactions and processes around 1M tokens across its context. We refer to this regime as XLong-horizon. Scaling RL to XLong-horizons introduces fundamental challenges on both system scheduling and trajectory modeling.
To address Challenge 2, Q WEN G YRE introduces a trajectory processor, which organizes each execution’s recorded model calls into a trajectory tree with shared prefixes, preserving each output’s original conditioning context. Task-specific evaluation scores preserved work after termination, including assessable partial progress after timeout. Once valid rewards are available, Q WEN G YRE selects a bounded set of trajectories per execution by role priority to control training cost. Training uses only policy-generated targets (Gallouédec & Rasul, 2026), counts shared targets once, and averages token losses within each execution to avoid overweighting executions with more paths.
Challenge 1: Extreme rollout variance and severe GPU underutilization at XLong-horizons. Online RL conventionally adopts a colocate design that alternates the GPU pool between rollout and training (Sheng et al., 2025; Mei et al., 2025). At XLong-horizons, however, rollouts exhibit severe long-tail variance; without general support for checkpointing workspace state, training remains blocked until all live executions end, leaving GPUs idle on a few trailing tasks. Alternative async deployments statically partition GPUs into dedicated pools (Fu et al., 2025). Nevertheless, at XLong-horizons, the wait for a ready batch can far exceed the duration of a training step, leaving training GPUs idle while rigid partitions prevent either stage from utilizing the other’s spare capacity. The challenge is therefore to reallocate GPUs dynamically as demand changes, using idle capacity for training while maintaining inference service for ongoing rollouts.
We evaluate Q WEN G YRE on NL2RepoBench (Ding et al., 2026), DeepSWE (Huang et al., 2026), and TerminalBench (Merrill et al., 2026) using Qwen 3.6 122B. We also apply Q WEN G YRE to XLong RL training of our flagship model, Qwen 3.8 2.4T, on NL2RepoBench. We compare against Colocate and Async baselines using equal GPU budgets under the same staleness constraints. Across the reported configurations, Q WEN G YRE achieves endto-end speedups of 1.38 to 1.78 times over Async and 1.21 to 1.85 times over Colocate while matching baseline training scores. On NL2RepoBench, Q WEN G YRE improves Qwen 3.8 2.4T’s score from approximately 52.5 % to 58.5 % over 48 training steps.
Challenge 2: Complex trajectory branching and prohibitive training overhead at XLong-horizons. The linear agent loop used in simple agent designs (Yao et al., 2023; Yang et al., 2024) does not survive at XLonghorizons: sustaining hours of execution typically relies on black-box harnesses that routinely compact saturated histories, delegate sub-tasks to sub-agents with isolated contexts, and retry failed execution paths (He et al., 2026b; Song et al., 2026a). As a result, a single task execution no longer yields a sequential trajectory, but an intricate, non-linear graph of divergent context paths. This non-linearity leads to an explosion in sample and token volume: expanding all these paths into individual training samples introduces severe redundancy across shared prefixes and incurs prohibitive training overhead. The challenge is thus to reconstruct valid training representations from such complex, branching executions
In summary, this paper makes the following contributions: • We identify the scheduling and trajectory-processing challenges of online RL for XLong agent executions. • We design an elastic scheduler that reallocates GPUs between rollout and training as demand changes, lets newly available GPUs join an ongoing training batch, and balances training work across them to reduce idle time. • We develop a trajectory processor that preserves each model output’s original context and evaluates partial work after timeouts. It selects a limited number of trajectories per execution, counts shared outputs once,
2
group’s scalar rewards, avoiding a learned value function (Shao et al., 2024). In our pipeline, each training step consumes a batch of B distinct rollout groups and advances the policy version by one. To amortize the cost of weight publication, the RL framework may perform µ consecutive training steps between publications; we call this interval a burst. Each burst consumes µB rollout groups and publishes the weights at its final step, as shown in Figure 2(a).
(a) Bursts and training steps Prev. burst
Current burst
Training step
···
Next burst
Policy publication
×μ
B queries
···
(b) From queries to trajectories Query
Rollout group
Training trajectories Sample
Rollout 1
··· Qu ery
···
Sample
Rollout i
···
Ji trajectories
···
Sample
Rollout N
Harness-driven agentic RL An agentic rollout is an end-to-end task execution in which a model interacts with an environment through a harness. The harness maintains task state, invokes tools according to model outputs, and constructs context for subsequent model calls. We call each model call–response interaction a model calls, each tool invocation a tool call, and the full task lifecycle, including retries, a rollout execution. Figure 2(c) shows an execution that branches at tool calls and the trajectories selected from it for training.
J1 trajectories
JN trajectories
(c) Trajectory tree of rollout i Model call
Tool call/text response subagent-call
main-agent trajectory
New context (root) trigger compaction
In our black-box setting, the harness accesses the model through a proxy provided by the RL framework. The framework records model calls and receives task-level evaluation results without changing the harness’s execution logic (He et al., 2026b; Song et al., 2026a). This interface provides no general mechanism to checkpoint and resume a live execution. If a model call cannot be served, the harness may stall or time out, changing the task outcome for reasons unrelated to the policy’s decisions. Scheduling must therefore preserve inference access throughout the execution. Requests can be routed to different rollout instances while the harness continues running.
···
summary-call
post-compaction root
sub-agent trajectories
···
trigger compaction
···
Figure 2: Hierarchy of training bursts, rollout groups, and trajectories. (a) Each burst performs µ consecutive training steps, then publishes the policy; each step consumes B query groups. (b) Each query produces N independent rollouts, and rollout i yields Ji training trajectories. (c) A rollout includes main-agent and sub-agent trajectories, with compaction producing summary branches. Dashed links reinsert summaries into post-compaction contexts.
The XLong execution profile. Recent software engineering benchmarks evaluate agents on repository construction, program reimplementation, and migration across entire codebases (Ding et al., 2026; Yang et al., 2026; Desai et al., 2026; Hong et al., 2026). We focus on X Long executions that last hours, span hundreds of dependent model–environment interactions, and process around 1M input and output tokens across model calls. This lifecycle volume counts repeated contexts at each request and can exceed the context limit of an individual request.
and averages token losses within each execution. • We demonstrate Q WEN G YRE on Qwen 3.8 2.4T, improving NL2RepoBench passrate from 52.5% to 58.5% in 48 training steps. Across evaluated workloads, Q WEN G YRE achieves up to 1.85× and 1.78× end-toend speedups over Colocate and Async, respectively.
2
Background and Motivation
2.1
X Long Agentic RL Workloads
These workloads exhibit both long execution tails and clustered terminations. Execution durations vary across queries and within each rollout group: generation length and tool latency affect the time per interaction (Zhang et al., 2026b). Over hundreds of dependent interactions, these differences can leave some executions running long after their peers finish. Since group-relative advantages require all N rewards, a straggler can delay training on its entire group. Executions with similar start times and a common timeout budget may expire within a narrow interval, producing clusters of execution terminations that can sharply reduce inference demand.
We first describe the group-relative RL pipeline and the execution characteristics of XLong agentic workloads. We then examine their implications for rollout–training scheduling and training data construction.
2.2
Group-relative RL pipeline. For each query, N independent rollouts sampled from the policy form a rollout group (Figure 2(b)). Advantages are computed from the
Rollout–Training Scheduling
RL frameworks organize rollout and training through two scheduling policies. A data dispatch policy controls 3
when and how much rollout work enters the pipeline. A GPU placement policy allocates GPUs between rollout and training in space and time. These policies can be combined to form different framework designs: Async, for example, can use either continuous or boundary dispatch.
motivates the black-box interface of Section 2.1, which exposes only model calls and task-level evaluation results. Using these records for group-relative RL requires valid outcome rewards and explicit choices of training targets and loss weights. Execution structure. An execution can contain multiple trajectories. A trajectory is a linear path of model calls in which each successive call extends the preceding call’s recorded context and response. Compaction and other context edits can rewrite an existing history (He et al., 2026b). Calls may remain causally related across these changes without forming one append-only history. A single agent or role can therefore contribute multiple trajectories. Concatenating requests across these boundaries as one history can give responses contexts they did not have during generation. Across hundreds of interactions, repeated context changes produce varying numbers of trajectories with different lengths.
Data dispatch policy and staleness. Under boundary dispatch, each rollout group enters the dispatch queue when admitted and retains its slot until training consumes it. Let φ ≥ 0 denote the extra dispatch in units of B rollout groups. With µ training steps per burst, the target queue depth is ( φ + µ) B groups. The queue is filled to this depth at startup and replenished with µB new groups after each burst’s weight publication. We measure a group’s scheduling staleness as d = vt − vd , with policy versions indexed by training steps. Here, vd is the version of the latest published policy when the group is admitted, and vt is the policy version immediately before the step that consumes it.
Reward validity. An execution’s reward can fail to support training in two distinct ways. First, an execution may reach its deadline before completing the task (Desai et al., 2026): its captured trajectories describe a partial attempt, and assigning all partial attempts the same failure reward obscures differences in task progress. Second, infrastructure or evaluator failures may leave no valid assessment at all; a missing assessment is distinct from a valid zero reward and cannot be filled in as one. Even after every execution in a group has terminated, the group may therefore lack the N valid sample rewards needed for advantage estimation.
The target queue depth Qb and steady-state mean scheduling staleness satisfy Qb = ( φ + µ) B, (1) µ−1 E[ d ] = φ + . (2) 2 The φ term accounts for extra dispatched groups, while µ −1 2 averages the training-step positions within a burst, indexed from zero. GPU placement policy and utilization. Async assigns fixed GPU pools to rollout and training, allowing the two stages to overlap (Fu et al., 2025; Zhong et al., 2025). Long rollout latency requires high concurrency to sustain throughput, but increasing dispatch depth also raises scheduling staleness (Equation 2). Limiting dispatch depth can therefore leave training GPUs waiting for data, while the fixed split prevents training from using spare rollout capacity as executions finish (Figure 1, Async).
Training semantics. A task-level outcome reward evaluates an execution of the original task. Figure 2 illustrates the varying trajectory counts: each query has N executions, while execution i contributes Ji selected trajectories. A flat average of trajectory losses gives executions with more selected paths greater aggregate weight (He et al., 2026b). Within an execution, the same averaging can give many short auxiliary paths more aggregate weight than a long main-agent path. A shared response can likewise receive extra weight if each copy contributes to the loss. The RL objective must therefore specify which roles, trajectories, and tokens contribute to the loss, together with the normalization and weighting rules.
Colocate alternates a shared GPU pool between rollout and training, as supported by HybridFlow and ReaL (Sheng et al., 2025; Mei et al., 2025). Because the pool switches as a unit, training must wait for the rollout tail even when groups are ready and GPUs are idle (Figure 1, Colocate). More steps per burst amortize switching and weight-publication costs, but increase scheduling staleness for later steps. 2.3
3
From Harness Executions to Training Data
Q WEN G YRE Overview
Q WEN G YRE supports online RL through unmodified black-box agent harnesses at XLong-horizons. Its architecture consists of two cooperating parts (Figure 3). The elastic scheduler (Section 4) reallocates GPUs between rollout and training while allowing live harness executions to continue. The trajectory processor (Section 5) constructs training samples from recorded model calls, retaining execution records independently of the samples admitted for optimization.
At XLong-horizons, context management and sub-agent delegation routinely break the append-only history of a simple agent loop (Yao et al., 2023; Yang et al., 2024). Harnesses compact histories as context windows fill and maintain separate contexts for sub-agents. The harnesses used for these executions are also used in deployment, and the most capable, such as Claude Code and Codex, are closed and evolve rapidly; reimplementing their control flow inside an RL framework is impractical and would train the policy under a harness different from the one it is deployed with. This
Elastic scheduler. GPU resources comprise elastic cells and an optional standalone rollout pool. Cells share a 4
their designated outcomes are valid. The buffer in Figure 3 stores execution records and associated rewards and packs selected paths from ready groups on demand into micro-steps carrying loss masks and weights. Training cells consume this stream, and the core performs one optimizer update per batch.
Waterlevel drop
rollout
train
train
rollout
Rollout Waterlevel
Core Training
Satellite Training
Satellite Rollout
Pull micro-steps
Multi-turn generation Push records
Sections 4 and 5 detail the elastic scheduling and trajectory processing planes, respectively.
4
Trajectory Processor
Sandbox
Evaluator
Repo
Tool call
Elastic Scheduler
The elastic scheduling plane coordinates role transitions and streaming training to use available GPUs while serving ongoing rollouts.
Model request
Reward
Buffer
Standalone Rollout
Rollout dispatch
Burst end
Elastic Scheduler
Harness
4.1
Topology-Aware Resource Organization
Cell layout and roles. Q WEN G YRE uses Megatron (Shoeybi et al., 2019) for training and SGLang (Zheng et al., 2024) for rollout. Its K elastic cells share a training parallel layout; an optional standalone pool remains in rollout. We favor small cells for fine-grained resource reallocation, subject to parallelism and topology constraints. Tensor, pipeline, expert, and context parallel groups remain within each cell, with placement favoring high-bandwidth interconnects.
Figure 3: Elastic scheduling and data flow in Q WEN G YRE. The sandbox hosts the harness and repository, linked by tool calls; Solid and dashed control links denote burst-boundary actions and waterlevel feedback, respectively. Double lines show data flow; the buffer combines trajectory storage and stream packing.
training parallel layout; each can serve rollouts or host a complete training actor replica. The black-box proxy, as part of the trajectory processor forwards model calls to rollout engines and reroutes them during role changes, while harnesses and workspaces remain outside GPU workers. The elastic scheduler dispatches rollout work at startup and burst boundaries and tracks queued and running executions as the rollout waterlevel. As this level falls, cells move to training when the remaining rollout capacity covers outstanding work. The core cell holds the authoritative weights and alone maintains optimizer state. Satellite cells pull its parameter snapshot and can join the ongoing batch before the rollout tail completes.
Cells follow a fixed order, with the core cell first and satellite cells thereafter. The core holds the authoritative weights and is the sole optimizer owner. Rollout capacity. Rollout capacity is measured as the number of concurrent rollout executions allowed by Q WEN G YRE. Let Ce and Cs ≥ 0 denote the capacities of each elastic cell in rollout mode and the standalone pool, respectively, with Cs = 0 when the standalone pool is absent. With k (t) cells assigned to training, the available rollout capacity is CR (t) = Cs + K − k (t) Ce . (3)
Trajectory processor. The proxy records exact modelrequest tokens and behavior log-probabilities in trajectory trees that share common prefixes and preserve each output’s original conditioning context. The evaluator scores each execution’s preserved workspace, including assessable partial progress after timeout, and sends the reward and its validity status to the buffer. At training admission, Q WEN G YRE computes group-relative advantages from the original execution rewards and selects at most Jmax trajectories per execution by role priority. Selected trajectories inherit their execution’s advantage. Only policy-generated tokens are eligible targets, and shared targets contribute once. Averaging losses over selected trainable tokens within each execution, then over executions in the batch, prevents additional trajectories from increasing an execution’s aggregate weight.
Boundary dispatch. To implement boundary dispatch (Section 2.2), the scheduler tracks three cumulative execution counts: d(t) increases when executions are dispatched into the rollout queue, p(t) increases when they launch, and f (t) increases when they terminate. Thus d(t) − p(t) counts queued executions, p(t) − f (t) counts running executions, and 0 ≤ f (t) ≤ p(t) ≤ d(t). Each launched execution increments f (t) exactly once on completion, failure, timeout, or cancellation, including surplus executions under oversampling. Model-request retries remain within the same execution and leave all counters unchanged. Queued entries record the published policy version at dispatch; harnesses and sandboxes are created at launch.
The two parts connect through a shared stream of training micro-steps, allowing the core to begin a batch before all samples are materialized. Execution termination lowers the scheduling waterlevel; a group becomes ready for training only after all its executions have closed and
In execution units, the initial dispatch is d(0) = ( φ + µ) NB, with µNB new executions dispatched after each burst’s weight publication. Whenever there is spare rollout capacity (Equation 3) and engines are ready to serve, the scheduler immediately fills available slots by
4.2
5
Boundary Dispatch and Waterlevel Control
n5
Roll Roll
n7
Roll
Pub.
T→R
Opt T→R
Opt
n6
Train
Policy publication
Roll
Train
T→R
n4
Reduce
Train
Pull
Roll
R→T
n3
If returning to rollout
Train
Broadcast
Roll
Reduce
n2
Train optimizer state + authoritative weights
Idle
Roll
Step #2
Pull
Roll
n1
R→T
n0
R→T
Step #1
Roll Roll
Roll Roll Roll Roll Roll Roll
Burst
Figure 4: Illustrative timeline of a two-step burst. The core (n0–n1) performs both optimizer updates. Satellite cells join ongoing steps after pulling its current parameters; n6–n7 stay in rollout during this burst. Gradient reduction precedes each update; a parameter broadcast follows the nonfinal update, and rollout policy publication follows the final update. The core’s final role transition and local publication (Pub.) occur only when it returns to rollout.
launching
min d(t) − p(t), CR (t) − p(t) − f (t)
tool execution state remain in place. The source cell first stops accepting new requests. The scheduler reserves destination capacity for affected executions and updates their routes, then cancels outstanding requests on the source engines. The proxy transparently retries these requests at the destination engines. The scheduler prioritizes available standalone engines, followed by engines in rollout cells in descending cell order. This ordering favors destinations expected to remain available longer, reducing repeated rerouting of the same execution within a burst.
(4)
queued executions and leaving any excess queued. The launches in Equation 4 use the existing quota, so d(t) remains fixed within the burst. Waterlevel control and training supply. The waterlevel w(t) = d(t) − f (t) is nonincreasing within a burst. The next cell can join training only if the remaining rollout capacity covers these outstanding executions: w(t) ≤ CR (t) − Ce .
KV-cache migration. KV-cache migration can reduce prefix recomputation during request rerouting. Before a request is canceled on the source engine, the source pins the corresponding cache blocks. Once the request is canceled and cache writes have stopped, the destination can read these blocks through RDMA. The source releases the cache only after the destination acknowledges receipt and takeover.
(5)
Equation 5 is paired with a supply check that counts ready, unassigned groups, capped by the number of groups the current batch still needs. This count must cover at least a fraction ρmin ∈ (0, 1] of the batch’s B groups continuously for τ before a transition begins. The timer resets after each transition and at each batch start. Cells transition serially, core first and then satellites by index, and remain in training throughout the burst, forming a growing prefix.
Role-state management. After request rerouting and cache handoff are complete, the cell offloads rollout parameters and releases KV caches and temporary buffers to reclaim GPU memory, then restores training parameters and buffers. Backend processes and communication groups remain alive across role changes.
Burst-End Weight Sync and Core Retention. After the burst’s final optimizer update, all training satellites return to rollout and restore their SGLang engines. The core then publishes the updated weights to all satellite and standalone rollout engines, preparing them for the next burst.
4.4
Centralized dynamic data parallelism. To let satellites join an ongoing batch, the core exposes buffers holding its authoritative weights through Mooncake’s Transfer Engine during its transition from rollout to training for the burst’s first step, and before each subsequent step. While training on the current batch is still in progress, satellites can join in the predefined cell order. Each satellite registers with the scheduler, pulls the core’s weights via RDMA, and begins training without interrupting the core (Figure 4). The weights remain fixed during gradient accumulation, so joining satellites read a consistent snapshot. The controller closes the joining window before the optimizer update.
After synchronization, the core remains in training only if at least ρmin B ready, unassigned groups are available for the next batch and the remaining rollout capacity covers outstanding work plus the pending µNB executions: w(t) + µNB ≤ Cs + (K − 1)Ce .
(6)
If either the supply check or Equation 6 fails, the core returns to rollout and reshards its parameters into the local SGLang (Figure 4). 4.3
Streaming Training with Dynamic Membership
Harness-Preserving Role Transitions
Request rerouting. During a role transition, the proxy reroutes model calls while the harness, workspace, and 6
After the entire batch has been processed, the core aggregates and normalizes gradients from all training cells. It then applies gradient clipping and performs a single optimizer update. After each nonfinal update, it broadcasts the new parameters to the current satellites, which remain in training for the next step.
a record of each model call. Each record preserves the exact input and output token IDs, the output tokens’ behavior log-probabilities, and associated metadata. To maintain these records across harness message processing, the proxy also caches original tool-call payloads by call ID. Before matching a subsequent request’s history, it restores these payloads while preserving the incoming IDs for tool-response association. This avoids spurious history mismatches caused by tool-call reformatting. The harness keeps its existing message interface, while training consumes the recorded tokens and associated metadata. Each output retains its original conditioning context even if later compaction removes it from the harness’s history.
For K ordered cells, we prebuild K − 1 sets of communication groups, one for each prefix of 2, . . . , K cells. Within each set, groups connect ranks holding corresponding model shards across cells. Each collective uses the set matching its participating prefix. Streaming training. Streaming training uses two buffers, shown as the Buffer in Figure 3. The trajectory buffer stores execution records and associated rewards by rollout group. The stream packing buffer prepares training inputs on demand. When its buffered trajectories run low, it pulls a ready group from the trajectory buffer, prioritizing the oldest policy version. It materializes the group’s selected trajectories into token rows and reorders the rows for efficient packing.
Prefix sharing through trajectory trees Q WEN G YRE organizes the input/output tokens and associated metadata in the TITO request records described above into a trajectory tree for each execution, following the prefix representation used in black-box agent training (Song et al., 2026a). Each node is built from a request record and stores its incremental input, recorded output, and associated metadata. After the tool-call restoration described above, the framework matches an incoming request’s serialized context against existing paths and reuses the tokens stored in the matched nodes. Only new input content is encoded, and the new request record is appended as a node. Diverging contexts create branches; requests with no matching prefix attach to the root. This organization captures the different context paths introduced by compaction and sub-agent calls (Figure 2(c)), while retaining independently generated outputs as separate nodes.
The packing buffer emits training inputs in small microsteps, each providing data for a cell’s data-parallel ranks. Cells atomically claim micro-steps from the packing buffer’s shared queue as compute capacity becomes available, receiving disjoint work. Earlier or faster cells consume more micro-steps, helping participating cells finish the batch at similar times. The packing buffer selects the batch’s B groups incrementally and closes the stream only after enqueuing all their micro-steps. Each cell processes incoming micro-steps with a continuous one-forward-one-backward (1F1B) pipeline. All stages within a cell use the same signal for input availability. When data are temporarily unavailable, the stages complete outstanding communication and pause at consistent scheduling points, preserving pipeline state. After the stream closes and the queue is exhausted, the pipeline drains its remaining work. Each cell that processes micro-steps therefore incurs one startup and one final drain per batch.
5
Candidate trajectories remain in this shared representation until selection, avoiding full-history expansion for every path. Selected paths are materialized from the stored token segments and their aligned metadata. 5.2
Scoring preserved work At execution closure, Q WEN G YRE fixes the committed policy-generation request record and identifies any available artifact state to assess. Q WEN G YRE evaluates the preserved work under the task’s scoring rule after an execution terminates, including after timeout. An unfinished execution can receive a valid score when its completed work is assessable. For these assessed outcomes, the data plane does not derive rewards from elapsed time, trajectory length, or termination status. Using executable checks or task-specific agentic evaluation (Zhuge et al., 2024), the evaluator returns a score, its validity status, and a reference to the assessed state.
Trajectory Processor
The trajectory processing plane constructs training samples from harness execution records and task outcomes. During rollout, Q WEN G YRE preserves the exact modelrequest tokens and each output’s original conditioning context using token-in, token-out (TITO) (Gallouédec & Rasul, 2026) (Section 5.1). After an execution terminates, Q WEN G YRE evaluates its preserved work, including assessable partial progress after timeout, under the task’s scoring rule to obtain an execution-level reward (Section 5.2). At training admission, Q WEN G YRE selects a bounded set of paths per execution from ready groups by role priority and fixes target masks so shared targets contribute once (Section 5.3). 5.1
Partial Scoring
Preserving the scope of an assessment An XLong execution may evaluate intermediate artifacts, revise them, and abandon branches. Each assessment remains associated with its execution, artifact version, and available branch reference; later assessments do not overwrite these associations. A critique used by a subsequent policy request remains part of its input, but its evaluator provenance excludes it from policy training targets. The
TITO and Trajectory Trees
Request records with TITO Q WEN G YRE implements TITO in the black-box proxy (Section 4.3) to maintain 7
designated task assessment supplies one reward for each scored execution; intermediate assessments remain associated with the states they evaluate.
highest-priority class that still has a leaf with unmasked trainable tokens, and draws one leaf from it with probability proportional to its unmasked count. The drawn trajectory’s targets are fixed to the tokens still unmasked at this draw; Q WEN G YRE then masks its entire root-toleaf path in the tree. This keeps a shared prefix from being trained more than once and leaves already-fixed targets untouched, while the lowered counts reweight the remaining leaves for the next draw. Admission repeats until no unmasked trainable token remains or Jmax trajectories are admitted, yielding Ji ≤ Jmax trajectories from execution i (Figure 2), each inheriting the execution’s group-relative advantage.
Handling failures An evaluator failure leaves the assessment unresolved; evaluation may be retried against the same preserved state within the execution’s overall time budget. If this budget expires without a valid reward, the recorded trajectories are excluded from training. An unresolved assessment is distinct from a valid zero score, which remains usable for training. For rare complete failures that leave no assessable work, such as image-pull failures, Q WEN G YRE inserts a dummy trajectory with zero reward as a failure placeholder. 5.3
The admitted trajectories retain their execution identities, recorded contexts and behavior log-probabilities, and fixed target masks. Materialization and packing preserve these annotations for the execution-level objective.
Trajectory Sampling
The trajectory tree records all context in the execution, which Q WEN G YRE turns into training data in two steps: it first masks every token that should not be trained, and then converts the tree into individual training trajectories.
Experimental Results
6.1
Experimental Setup
Tasks. We evaluate Q WEN G YRE on three agentic workloads: NL2RepoBench (Ding et al., 2026), DeepSWE (Huang et al., 2026), and TerminalBench (Merrill et al., 2026). We construct training data internally for all three workloads. For benchmark evaluation, we follow each benchmark’s prescribed evaluation protocol. NL2RepoBench covers the XLong setting of building software repositories from natural-language specifications; during training, we use a separate Qwen 3.7-Max evaluator to score the generated code using independently written tests. DeepSWE covers software engineering in existing repositories; during training, we use repository tests for evaluation. TerminalBench comprises multi-turn terminal tasks; during training, we use task-specific verifiers for evaluation. All workloads use Claude Code 2.1.220 (Anthropic, 2026) to execute shell commands and manipulate files in containerized environments.
Trainable tokens are defined by provenance. A node’s tokens are eligible training targets only if the policy produced them during rollout. TITO (Section 5.1) records whether each token was emitted by our SGLang engines or supplied from outside. System prompts, user prompts, tool results, and assistant-role content the harness injects rather than samples from the policy will be masked and carry no training signal: they are retained as conditioning context only. Sampling trajectories from the tree. Candidate trajectories correspond to the leaves of the trajectory tree: a leaf together with the prefix path from the root defines one trajectory. At XLong-horizons a single execution can produce many candidate trajectories, and their number may vary widely across executions. Simply retaining all trajectories may lead to a decrease in training efficiency, so Q WEN G YRE retains at most Jmax training trajectories per execution and fills this budget in order of expected importance. To rank the candidates, Q WEN G YRE classifies each leaf by the role of the request sequence it terminates. For a Claude Code harness these classes are, for the main agent and for each sub-agent, its taskdriving trajectory and its summary trajectory, ordered by the priority main > main-summary > sub-agent > sub-agent-summary.
6
Models. We use Qwen 3.6 122B and Qwen 3.8 2.4T. For Qwen 3.6 122B, the NL2RepoBench experiments use 32 nodes, organized into eight cells of four nodes each by default, with an inference concurrency limit of 144 per cell. The cell-granularity ablation also evaluates four 8-node cells under the same 32-node budget. DeepSWE and TerminalBench follow the same protocol, using 24 nodes organized into six cells with an inference concurrency limit of 192 per cell. For Qwen 3.8 2.4T, the NL2RepoBench traces use four cells of 48 nodes each, with a total inference concurrency limit of 384. Unless otherwise stated, we use no dedicated standalone rollout nodes.
(7)
The ordering in Equation 7 follows two principles. The main agent’s turns are what the task-level reward most directly credits, whereas sub-agent trajectories contribute only through delegated subtasks; and within each agent, turns that advance the task rank above the summary trajectory, which manages context rather than acting on the task.
RL algorithm. All methods use our adaptation of Group Sequence Policy Optimization (GSPO) (Zheng et al., 2025), combining token-level importance weighting with execution-level clipping. Token importance weights are capped at 5, and the sequence-level clipping bounds are [0.995, 1.005]. Each group contains 16 independent rollout executions; each update consumes 32
Admission starts from the provenance mask above and draws trajectories one at a time. Each step takes the
8
(a) Query time
(b) Rollout execution 8%
Share
10%
(c) Tree tokens
Harness timeout (6.25%) Overall timeout (0.87%)
6%
(d) Tool calls Total Decode
15%
0% 0
1
2
3
≥4
0% 0
1
Duration (h)
2
≥4
3
20%
2%
5%
2%
30% 4%
10%
4%
5%
(e) Tree leaves
6%
0% 0
0.5
Duration (h)
1.0
≥ 1.5
10%
0%
Tokens (M)
0
400
800
≥ 1.2k
0% 0
Calls per rollout
5
10
15 ≥ 20
Leaves per tree
Figure 5: NL2RepoBench execution distributions with Qwen 3.6 122B at E[d] = 1.5 over steps 1–48. (a) Query durations, measured as the longest of each query’s 16 rollout execution durations. (b) Rollout execution durations. (c) Total (solid) counts input and output tokens per rollout tree; Decode (dashed) counts generated output tokens. Both count shared prefixes once. The smoothed curves show shares per 50k tokens. (d) Tool calls per rollout execution. (e) Leaves per rollout tree. In (b), pink and purple mark harness and overall timeouts, respectively; executions with both flags appear only in purple. Final bins labeled ≥ collect values at or above the indicated thresholds. (a) Roles during the first three bursts
(b) Training time share
Burst 1 Step 1
Burst 2 Step 2 Step 3
Step 4
Burst 3 Step 5 Step 6
Step 7
Step 8 Step 9
45.36% 39.86% 33.71% 27.01% 19.60% 13.49% 5.12%
0 1
Cell
2 3 4 5 6
0%
7 0
2
4
6
8
Elapsed time (h)
10
12
rollout
train
0%
20%
40%
60%
Figure 6: Execution profile of Q WEN G YRE on NL2RepoBench with Qwen 3.6 122B at E[d] = 1.5. (a) Cell roles during the first three bursts (steps 1–9). (b) Each cell’s recorded training time as a fraction of the full observation window. Cell 7 remains in rollout throughout this interval.
groups for NL2RepoBench or 48 for DeepSWE and TerminalBench. The default trajectory cap is Jmax = 5 per rollout execution. All methods use Adam (Kingma & Ba, 2015) with a constant learning rate of 10−6 .
6.2
End-to-End Performance
NL2RepoBench on Qwen 3.6 122B Figure 5 shows mean durations of 1.93 hours per rollout execution and 2.96 hours per query; 9.51% of queries take at least four hours. Figure 5(b) shows harness and overall timeout rates of 6.25% and 0.87%, respectively. Harness timeouts indicate that task execution has timed out; the results are still scored and used for training. Failed evaluations can be retried within the overall execution time budget. An overall timeout exhausts this budget before a valid reward is obtained, so the execution is discarded from training.
Baselines. We compare Q WEN G YRE with two scheduling baselines. Async assigns separate, fixed GPU pools to rollout and training, allowing training steps to consume ready rollout groups while other executions continue. It uses boundary dispatch with a target queue depth, refilling consumed groups’ slots after weight publication. Colocate alternates a shared GPU pool between rollout and training. Once the admitted rollout executions finish, the entire pool switches to training and processes the completed rollout groups in a burst of consecutive training steps.
Figure 7(a–d) shows that Q WEN G YRE achieves 1.42– 1.53× speedups over Async and 1.36–1.47× over Colocate after 48 training steps at both mean scheduling staleness targets. Figure 7(e,f) shows comparable trainingscore curves across methods, indicating that faster training is achieved while maintaining task performance.
We compare methods at common target values of mean scheduling staleness E[d], defined in Section 2.2. All three methods use boundary dispatch; we choose the extra-dispatch parameter φ and the number of training steps per burst µ using Equation 2. The original Colocate cannot dispatch extra groups without interrupting the harness, so it uses φ = 0. For example, at E[d] = 1.5, Async can use µ = 1 with φ = 1.5, Colocate uses µ = 4 with φ = 0, and Q WEN G YRE can use µ = 3 with φ = 0.5.
The role timeline in Figure 6(a) shows how Q WEN G YRE dynamically allocates resources between training and inference. Cells enter training at different points while other cells remain in rollout. Figure 6(b) shows how training time shares vary across cells over the full observation window. This staggered allocation allows training to begin with a subset of cells and subsequently expand. 9
(a) [d] = 1
(b) [d] = 1.5
(a) Query time
(b) Rollout execution
40%
200
30% 20%
0% 36
48
1
12
Training step
Speedup
1.75
48
36
48
1
0.60
0.60
0.55
0.55
0.50 0.45
1
12
24
36
48
0.50 0.45
1
12
24
36
48
1
Training step QwenGyre
4
1.940
1.831
1.784 1.784
1.195
1.205
1.212 1.213
12
24
1.50 1.25
12
24
36
48
1
36
48
Training step (f) Eval score
0.80 0.76
24
36
24
36
48
0
12
24
36
48
Training step Async
Colocate
Figure 8: NL2RepoBench with Qwen 3.8 2.4T. (a) Measured query durations, using the longest rollout execution wall time per query. (b) Measured rollout execution durations. The final bin collects durations ≥ 4.45 h. In (b), pink and purple mark harness and overall timeouts, respectively. (c,d) Step times and cumulative speedups over 48 steps. Arrows in (d) label speedups at selected checkpoints in matching colors. (e,f) Training and evaluation scores.
48
Training step Async
0.54 0.52
12
QwenGyre
12
0.56
Training step
2 48
3
1.75
1.00
1
3
36
100
0.72
4
24
2
2.00
Training step (e) Train score
(h)
12
1
Duration (h) (d) Speedup 2.25
1
5
1
0
200
48
Training step
(g)
4
0.58
Training step
Max. |Δlog p|
36
0.40
0.40
6
24
3
300
Training step (f) Train score Train score
Train score
Training step (e) Train score
12
Train score
24
2
400
12 24 36 48 1.556 1.503 1.467 1.423 1.442 1.462 1.455 1.475
1.25 12
1
Duration (h) (c) Step time
1.50
1
0% 0
(d) 12 24 36 48 1.576 1.554 1.541 1.528 1.357 1.390 1.390 1.361
2.00
36
Training step
(c) 2.25
24
Speedup
24
Eval score
12
Step time (min)
1
4% 2%
10%
100
Harness timeout (9.37%) Overall timeout (4.63%)
6%
Share
300
Share
Step time (min)
400
Colocate
Figure 7: NL2RepoBench results with Qwen 3.6 122B at E[d] = 1 (left) and E[d] = 1.5 (right). Speedup is cumulative baseline time divided by cumulative Q WEN G YRE time. In (c,d), in-plot tables give speedups at selected checkpoints: Async in the upper row and Colocate in the lower row, with matching curve colors. Max. |∆ log p| is the maximum absolute difference between current-policy and recorded rollout-policy log-probabilities over unmasked response tokens.
hours for Async and 91.47 hours for Colocate, yielding end-to-end speedups of 1.78× and 1.21×, respectively. The concentration of query durations near the timeout limit helps explain the difference in gains over the two baselines. It suggests fewer opportunities for Q WEN G YRE to reclaim cells early and expand training gradually, consistent with the smaller gain over Colocate, which also shares its resource pool between training and rollout. Async’s fixed partition leaves resources reserved for training unavailable to rollout, even when training is idle. By dynamically allocating these resources to rollout, Q WEN G YRE can increase inference capacity during long executions, yielding a larger gain over Async.
It uses resources that a fixed Async partition leaves assigned to rollout and avoids requiring Colocate’s poolwide phase transition before training can proceed. The average switching times are 8.52 s for rollout to training and 3.46 s for training to rollout. Both are negligible compared with rollout executions lasting hours.
Figure 8(e) shows that Q WEN G YRE’s training scores match those of Async and Colocate. Its evaluation passrate improves from 52.48% to 58.54% over 48 training steps (Figure 8(f)).
NL2RepoBench on Qwen 3.8 2.4T On NL2RepoBench, Qwen 3.8 2.4T has a higher timeout rate than Qwen 3.6 122B, even though we extend the timeout limit. Figure 8(a,b) shows the query and rollout execution duration distributions. Among queries with complete duration records, 61.1% take at least four hours, compared with 9.51% for Qwen 3.6 122B. In 38.2% of these queries, the longest rollout execution terminates with an overall timeout.
DeepSWE and TerminalBench on Qwen 3.6 122B DeepSWE and TerminalBench share the same end-toend evaluation protocol: Qwen 3.6 122B, 24 training steps, and all three methods (Q WEN G YRE, Async, and Colocate) at target mean scheduling staleness values E[d] = 1 and 1.5. Table 2 reports the resulting cumulative speedups.
As shown in Figure 8(c,d), Q WEN G YRE completes 48 training steps in 75.42 hours, compared with 134.55
Across the reported DeepSWE and TerminalBench con10
Table 1: Scheduling ablations over the first 12 NL2RepoBench training steps with Qwen 3.6 122B at E[d] = 1.5. A0, C0, and E0 reuse the Async, Colocate, and Q WEN G YRE base configurations from Section 6.2. Arrows show the evolution paths; E1 and E2 branch independently from E0. Switch is – for fixed partitions, Global for switching all nodes together, Coarse for switching the Colocate pool together, and Fine for switching cells independently. Standalone nodes remain in rollout. Allocation lists training/rollout (T/R) or Colocate/standalone (C/S) node counts; cell entries give cells × nodes per cell. Total is the time for all 12 steps in hours; Time / E0 normalizes it to E0. Lower is better. ID
Evolution
Stream
Switch
( φ, µ)
Allocation
Total (h)
Time / E0
A0 A1 A2 A3
Original Async A0 with adjusted partition A0 + multi-step burst A0 + streaming
Disabled Disabled Disabled Enabled
– – – –
(1.5, 1) (1.5, 1) (0.5, 3) (1.5, 1)
16 T / 16 R 8 T / 24 R 16 T / 16 R 16 T / 16 R
25.08 27.01 32.58 22.18
1.556 1.676 2.021 1.376
C0 C1 C2 C3
Original Colocate C0 + standalone rollout nodes A3 + role switching; C1 + streaming C2 with adjusted allocation
Disabled Disabled Enabled Enabled
Global Coarse Coarse Coarse
(0, 4) (0.5, 3) (0.5, 3) (0.5, 3)
32 C / 0 S 28 C / 4 S 28 C / 4 S 16 C / 16 S
23.23 22.56 19.23 20.69
1.442 1.400 1.193 1.284
E0 E1 E2
C2 + fine-grained elastic allocation E0 with one step per burst E0 with four 8-node cells
Enabled Enabled Enabled
Fine Fine Fine
(0.5, 3) (1.5, 1) (0.5, 3)
8 cells × 4 8 cells × 4 4 cells × 8
16.12 17.87 16.47
1.000 1.109 1.022
Table 2: 24-step speedup of Q WEN G YRE over each baseline with Qwen 3.6 122B. The ( φ, µ) column gives Q WEN G YRE’s extra dispatch in batch units and training steps per burst. Async uses ( φ, µ) = (E[d], 1); Colocate uses ( φ, µ) = (0, 2E[d] + 1), following Equation 2. Dataset
E[ d ]
( φ, µ)
vs. Async
vs. Colocate
DeepSWE DeepSWE
1 1.5
(0.5, 2) (0.5, 3)
1.569 1.572
1.610 1.824
TerminalBench TerminalBench
1 1.5
(0.5, 2) (0.5, 3)
1.383 1.430
1.809 1.849
ness state. Although extra dispatch mitigates the impact of rollout stragglers, fewer training steps are completed per burst, limiting the end-to-end improvement over C0 to just 2.9%. Adding streaming to C1 (C2) reduces time by 14.8%. However, all 28 colocated nodes must still switch together, so streaming training can begin only late in the burst. Increasing standalone capacity (C3) shrinks the pool that must switch together, but also leaves fewer nodes for training. The resulting longer training phase delays burst completion and the next dispatch, increasing end-to-end time by 7.6% relative to C2. The fundamental limitation is the coarse switching granularity, which prevents training from fully using the idle capacity that emerges as colocated rollouts finish at different times.
figurations, Q WEN G YRE provides consistent end-to-end speedups over both Async and Colocate. 6.3
Ablation Study Combining streaming and fine-grained allocation. E0 builds on C2 by replacing its fixed split of 28 colocated nodes and four standalone rollout nodes with eight independently switching 4-node cells, while retaining streaming and three training steps per burst. Training can thus start on a subset of cells and expand as data arrives while other cells continue rollout, yielding a 1.19× speedup over C2.
Table 1 traces two paths toward Q WEN G YRE: adding streaming training to Async and adding standalone rollout capacity to Colocate. Adding streaming to Async. Reallocating nodes from training to rollout (A0 to A1) increases end-to-end time by 7.7%. Although training nodes wait for rollout data, executions finish in bursts; the smaller training pool consumes these bursts more slowly, delaying weight publication and replacement rollouts under boundary dispatch. Longer training bursts (A2) increase time by 29.9% relative to A0 because the fixed rollout pool cannot supply them efficiently. Streaming (A3) instead reduces time by 11.6% by starting training on completed queries before a full batch is ready. However, streaming captures only part of the potential gain. The fundamental limitation remains Async’s fixed resource partition: it cannot reallocate idle GPUs between training and rollout as their demands change.
E1 retains E0’s streaming and cell allocation but uses one training step per burst, with extra dispatch increased from 0.5 to 1.5 to preserve scheduling staleness. Although E1 takes 10.9% longer than E0, it still achieves a 1.24× speedup over A3, which also uses streaming and one training step per burst, and a 1.08× speedup over C2. These results show that fine- grained allocation provides substantial gains even without multi-step bursts; longer bursts offer an additional benefit. Streaming makes rollout data available for training earlier, while fine-grained allocation supplies training resources as that data arrives.
Adding standalone rollout nodes to Colocate. while the colocated pool switches to training, preserving har-
Cell granularity. E2 builds on E0 by regrouping the 11
same 32 nodes from eight 4-node cells into four 8-node cells, retaining streaming and ( φ, µ) = (0.5, 3). It completes the 12 training steps in 16 hours 28 minutes (16.47 hours), only 2.2% longer than E0’s 16.12 hours. In E0, rollout fully hides training, with 20.85% slack remaining. E2’s coarser cell granularity exposes a small amount of training on the critical path, while most training remains overlapped with rollout. This explains the modest impact of reducing the number of cells from eight to four in this setting.
out and trainer pools, using bounded staleness, multiple policy versions, or fixed-weight trajectory groups to control training semantics (Fu et al., 2025; Han et al., 2025; Li et al., 2026; Hu et al., 2026; Chen et al., 2026b). SAO and SPO++ further adapt policy optimization to asynchronously arriving agent trajectories (Hou et al., 2026; Ruan et al., 2026). Elastic rollout–training scheduling. Recent RL frameworks dynamically adjust the resource boundary between rollout and training. TideRL adapts rollout and reference execution using readiness signals (Ren et al., 2026b); BiDiRL allows either side of an asynchronous pipeline to borrow idle resources from the other (Tan et al., 2026); and DynaResize reallocates GPUs between rollout and training pools using communicator reuse and state staging (Du et al., 2026a). Libra combines a global resource planner with an elastic hybrid pool whose workers can join training as additional data-parallel replicas (Chen et al., 2026a). Other RL frameworks obtain elasticity from complementary phases across RL jobs, online-serving clusters, or preemptible rollout resources (Wu et al., 2026a; Gao et al., 2026b; Wu et al., 2026b), while AReaL-DTE reduces cross-role policy synchronization through sparse weight transfer (Peng et al., 2026).
Trajectory sampling budget. Sampling only the main trajectory per execution is insufficient: it yields lower training scores and larger gradient norms over the first 42 training steps (Figure 9). Increasing the budget (Section 5.3) to five trajectories matches the scores obtained by sampling all available trajectories, with similarly small gradient norms. Over the same window, mean forward–backward time per step with caps of one and five is 41.4% and 74.8% of the uncapped baseline, respectively. A cap of five therefore preserves comparable training scores while reducing forward–backward time by 25.2%. (a) Grad. norm
(b) Train score 0.56
Train score
Grad. norm
10−1 10−2 10−3
0.48
Black-box harnesses and structured trajectories. Harness-native RL frameworks connect opaque agent runtimes to policy optimization by instrumenting or proxying model calls. Agent Lightning decouples agent execution from training, while Polar and LEGO-RL preserve rollout-time token identity under harnesscontrolled context processing (Luo et al., 2025b; He et al., 2026b; Xu et al., 2026a; Du et al., 2026b). ClawGym II reconstructs captured requests as shared-prefix trees and adapts PPO and GRPO so that each shared node contributes training signal only once (Song et al., 2026a); its retained paths share a reward derived from the final workspace evaluation.
0.40 0.32
1
12
24
36
42
1
Training step Max traj 1
12
24
36
42
Training step Max traj 5
Uncapped
Figure 9: Trajectory-budget ablation for Q WEN G YRE on NL2RepoBench with Qwen 3.6 122B over training steps 1–42. (a) Gradient norm on a logarithmic scale. (b) Mean training score. Max traj 1 samples only the main trajectory per execution; Max traj 5 admits up to five trajectories; “Uncapped” uses all available leaf trajectories.
7
Related work exploits execution structure at other layers. BPO, IAPO, and MileGPO use branches, dependency graphs, or intermediate milestones for finer-grained credit assignment (He et al., 2026a; Ren et al., 2026a; Qian et al., 2026), whereas psRL and related frameworks reuse shared prefixes after sample construction (Yu et al., 2026; Wang et al., 2025a). These techniques address trajectory capture, credit assignment, or prefix computation individually. Q WEN G YRE instead couples readiness-driven actor reconfiguration with a trajectory processor preserving role, branch, environment-state, and evaluator provenance throughout XLong rollout collection, admission, and training materialization.
Related Work
Agentic RL execution. General-purpose LLM RL frameworks such as HybridFlow/veRL, OpenRLHF, ReaL, and ROLL compose rollout, reference, reward, and policy-update workers under different placement strategies (Sheng et al., 2025; Hu et al., 2025; Mei et al., 2025; Wang et al., 2025b). Agent-oriented runtimes extend these abstractions with multi-turn tools, containerized environments, asynchronous dispatch, and trajectory-level scheduling (Jiang et al., 2026; Wang et al., 2025c; Zhang et al., 2025; Cao et al., 2025; Zhang et al., 2026a;b).
8
Synchronous RL frameworks preserve explicit policyupdate boundaries but expose rollout-length skew. Recent work mitigates this tail through scheduling, lengthaware packing, partial-rollout continuation, and tail isolation (Qin et al., 2026; Gao et al., 2026a; Zhou et al., 2025; Xu et al., 2026b). Asynchronous RL frameworks instead stream trajectories between disaggregated roll-
Limitations
GPU budget and cell granularity. Elastic scheduling needs multiple independently switchable cells while maintaining enough rollout capacity to serve unfinished executions. Each cell must accommodate the model’s memory and training parallelism requirements, so a very small GPU budget may not support this organi12
zation. Such deployments are generally impractical for the resource-intensive XLong workloads targeted here. Within our evaluated regime, a modest number of cells is sufficient: under the same 32-node budget, regrouping eight 4-node cells (E0) into four 8-node cells (E2) increases execution time by only 2.2% (Table 1). E0 fully hides training behind rollout, while E2 exposes a small amount of training on the critical path. Most of the overlap benefit is therefore retained with four cells in this setting. This result concerns cell granularity at a fixed budget; efficiency under substantially smaller total GPU budgets remains unevaluated.
Kaiwen Chen, Xin Tan, Jingzong Li, and Hong Xu. Libra: Efficient resource management for agentic RL posttraining, 2026a. URL https://arxiv.org/abs/2606.0 3077. Rongjian Chen, Jianmin Hu, Kejiang Ye, and Minxian Xu. RolloutPipe: Overlapping pipelined rollout and training in disaggregated on-policy LLM reinforcement learning. In Proceedings of the 2026 International Conference on Cognitive Computing (ICCC 2026), 2026b. URL https://arxiv.org/abs/2606.26997. Xiang Deng, Jeff Da, Edwin Pan, Yannis Yiming He, Charles Ide, Kanak Garg, Niklas Lauffer, Andrew Park, Chetan Rane, Karmini Sampath, Maya Krishnan, Srivatsa Kundurthy, Sean Hendryx, Zifan Wang, Chen Bo Calvin Zhang, Noah Jacobson, Bing Liu, and Brad Kenstler. SWE-Bench Pro: Can AI agents solve longhorizon software engineering tasks? In Proceedings of the 43rd International Conference on Machine Learning, volume 306 of Proceedings of Machine Learning Research. PMLR, 2026. URL https://icml.cc/virtual/2026/p oster/61047.
Shuffling with streaming training. With multiple training steps per burst (µ > 1), streaming training assigns ready rollout groups incrementally to successive updates. The current design therefore cannot globally shuffle data across all minibatches before the burst starts. Supporting such a shuffle would require waiting for the complete burst’s data before its first update, delaying training and sacrificing some overlap with rollout. This constrains training procedures that depend on globally randomized minibatch assignment. Single-step bursts (µ = 1) avoid the cross-minibatch split while retaining elastic allocation and streaming. At matched scheduling staleness, E1 takes 10.9% longer than E0, but still achieves a 1.24× speedup over streaming Async (A3) and a 1.08× speedup over Colocate with streaming (C2) (Table 1). Thus, multi-step bursts provide an additional efficiency gain, while single-step bursts remain a practical option. These ablations quantify execution time; the effect of sample ordering on learning quality has not been isolated.
9
Rishi Desai, Jesse Hu, Joan Cabezas, Neel Harsola, Pratyush Shukla, Roey Ben Chaim, Adnan El Assadi, Omkaar Mukund Kamath, Fenil Faldu, Prannay Hebbar, Jiankai Sun, Yiyuan Li, Pramod Srinivasan, Ishan Gupta, Christopher Settles, Daniel Wang, Derek Chen, Pranav Raja, Albert Liu, Marek Šuppa, Nevasini Sasikumar, Luyang Kong, Erik Quintanilla, Xiangyi Li, Ivan Bercovich, and Steven Dillmann. SWE-Marathon: Can agents autonomously complete ultra-long-horizon software work?, 2026. URL https: //arxiv.org/abs/2606.07682. Jingzhe Ding, Shengda Long, Changxin Pu, Ge Zhang, Huan Zhou, Hongwan Gao, Xiang Gao, Chao He, Yue Hou, Fei Hu, Zhaojian Li, Weiran Shi, Zaiyuan Wang, Daoguang Zan, Chenchen Zhang, Xiaoxu Zhang, Qizhi Chen, Xianfu Cheng, Bo Deng, Qingshui Gu, Kai Hua, Juntao Lin, Pai Liu, Mingchen Li, Minghao Li, Xuanguang Pan, Zifan Peng, Yujia Qin, Yong Shan, Zhewen Tan, Haoran Wang, Tom Tang, Weihao Xie, Yishuo Yuan, Jiayu Zhang, Yunfei Zhao, He Zhu, Liya Zhu, Chenyang Zou, Ming Ding, Jianpeng Jiao, Jiaheng Liu, Liam Liu, Qian Liu, Chongyang Tao, Jian Yang, Tong Yang, Zhaoxiang Zhang, Xinjie Chen, and Wenhao Huang. NL2Repo-Bench: Towards longhorizon repository generation evaluation of coding agents. In Proceedings of the 43rd International Conference on Machine Learning, volume 306 of Proceedings of Machine Learning Research. PMLR, 2026. URL https://icml.cc/virtual/2026/poster/60772.
Conclusion
We presented Q WEN G YRE, which combines elastic scheduling and trajectory processing for online RL through unmodified harnesses at XLong-horizons. It reallocates GPUs while preserving live executions and constructs bounded training samples with original contexts and balanced execution-level contributions. We demonstrate these capabilities through XLong RL training of our flagship model, Qwen 3.8 2.4T. Across three workloads, Q WEN G YRE achieves up to 1.78× speedup over Async and 1.85× over Colocate under equal GPU budgets and matched scheduling staleness, with comparable training scores.
References Anthropic. Claude Code. Official software repository, 2026. URL https://github.com/anthropics/claude -code.
Hanlin Du, Zhiyuan Yan, Yungang Bao, and Sa Wang. DynaResize: Runtime GPU reallocation for disaggregated LLM post-training, 2026a. URL https: //arxiv.org/abs/2607.22614.
Shiyi Cao, Dacheng Li, Fangzhou Zhao, Shuo Yuan, Sumanth R. Hegde, Connor Chen, Charlie Ruan, Tyler Griggs, Shu Liu, Eric Tang, Richard Liaw, Philipp Moritz, Matei Zaharia, Joseph E. Gonzalez, and Ion Stoica. SkyRL-Agent: Efficient RL training for multiturn LLM agent, 2025. URL https://arxiv.org/abs/ 2511.16108.
Yiming Du, Yuxin Jiang, Tao Yuan, Jianbo Dai, Shaowei Wang, Jierun Chen, Chaofan Tao, Xianzhi Yu, Lifeng Shang, Kam-Fai Wong, Xiaohui Li, and Haoli Bai. LEGO-RL: Harness-native reinforcement learning for coding agents, 2026b. URL https://arxiv.org/abs/ 2608.17393. 13
Wei Fu, Jiaxuan Gao, Xujie Shen, Chen Zhu, Zhiyu Mei, Chuyi He, Shusheng Xu, Guo Wei, Jun Mei, Jiashu Wang, Tongkai Yang, Binhang Yuan, and Yi Wu. AReaL: A large-scale asynchronous reinforcement learning system for language reasoning. In Advances in Neural Information Processing Systems, volume 38, pp. 36256–36282. Curran Associates, Inc., 2025. doi: 10.5 2202/085713-1218. URL https://proceedings.neur ips.cc/paper_files/paper/2025/hash/33c00862bfa 29ac72ecf630a41e19352-Abstract-Conference.html.
Deyao Hong, Yizhe Chi, Wenyi Li, Xiaoqiu Wang, Mingju Gao, Kaisen Yang, Bingxiang He, Youjie Zheng, Calvin Xiao, and Qinhuai Na. SWE Refactor Bench: Can coding agents complete a long-horizon, wholerepository stack migration?, 2026. URL https://arxi v.org/abs/2608.23564. Zhenyu Hou, Yujiang Li, Jie Tang, and Yuxiao Dong. Single-rollout asynchronous optimization for agentic reinforcement learning, 2026. URL https://arxiv.or g/abs/2607.07508.
Quentin Gallouédec and Kashif Rasul. Agentic RL: Token-in, token-out done right. Hugging Face Blog, 2026. URL https://huggingface.co/blog/huggingf ace/tito.
Jian Hu, Xibin Wu, Wei Shen, Jason Klein Liu, Weixun Wang, Songlin Jiang, Haoran Wang, Hao Chen, Bin Chen, Wenkai Fang, Xianyu, Yu Cao, Haotian Xu, and Yiming Liu. OpenRLHF: A ray-based easy-to-use, scalable and high-performance RLHF framework. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pp. 656–666, Suzhou, China, 2025. Association for Computational Linguistics. doi: 10.18653/v1/2025.emnlp -demos.48. URL https://aclanthology.org/2025.em nlp-demos.48/.
Wei Gao, Yuheng Zhao, Dakai An, Tianyuan Wu, Lunxi Cao, Shaopan Xiong, Ju Huang, Weixun Wang, Siran Yang, Wenbo Su, Jiamang Wang, Lin Qu, Bo Zheng, and Wei Wang. RollPacker: Taming long-tail rollouts for RL post-training with tail batching. In 23rd USENIX Symposium on Networked Systems Design and Implementation (NSDI 26), pp. 849–866, Renton, WA, 2026a. USENIX Association. ISBN 978-1-939133-54-0. URL https://www.usenix.org/conference/nsdi26/p resentation/gao-wei.
Tianhao Hu, Xiangcheng Liu, Yuchun Miao, Youshao Xiao, Hongyu Zang, Yang Zheng, Xuan Huang, Jinrui Ding, Yufei Zhang, Yu Yang, Yi-Kai Zhang, Yueqing Sun, Chengcheng Han, Xiandi Ma, Wei Wang, Qi Gu, Yerui Sun, Yuchen Xie, and Xunliang Cai. DORA: A scalable asynchronous reinforcement learning system for language model training, 2026. URL https://ar xiv.org/abs/2604.26256.
Wei Gao, Yuheng Zhao, Dilxat Muhtar, Dakai An, Xuchun Shang, Tianyuan Wu, Lunxi Cao, Shaopan Xiong, Weixun Wang, Ju Huang, Teng Ma, Siran Yang, Jiamang Wang, Lin Qu, Bo Zheng, and Wei Wang. ROSE: Rollout on serving GPUs via cooperative elasticity for agentic RL, 2026b. URL https: //arxiv.org/abs/2605.06534.
Wenqi Huang, Charley Lee, Leonard Tng, and Serena Ge. DeepSWE: Measuring frontier coding agents on original, long-horizon engineering tasks, 2026. URL https://arxiv.org/abs/2607.07946.
Alexander Golubev, Maria Trofimova, Sergei Polezhaev, Ibragim Badertdinov, Maksim Nekrashevich, Anton Shevtsov, Simon Karasik, Sergey Abramov, Andrei Andriushchenko, Filipp Fisin, Sergei Skvortsov, and Boris Yangel. Training long-context, multi-turn software engineering agents with reinforcement learning, 2025. URL https://arxiv.org/abs/2508.03501.
Dongfu Jiang, Yi Lu, Zhuofeng Li, Zhiheng Lyu, Ping Nie, Haozhe Wang, Alex Su, Hui Chen, Kai Zou, Chao Du, Tianyu Pang, and Wenhu Chen. VerlTool: Towards holistic agentic reinforcement learning with tool use. Transactions on Machine Learning Research, 2026. URL https://arxiv.org/abs/2509.01055.
Zhenyu Han, Ansheng You, Haibo Wang, Kui Luo, Guang Yang, Wenqi Shi, Menglong Chen, Sicheng Zhang, Zeshun Lan, Chunshi Deng, Huazhong Ji, Wenjie Liu, Yu Huang, Yixiang Zhang, Chenyi Pan, Jing Wang, Xin Huang, Chunsheng Li, and Jianping Wu. AsyncFlow: An asynchronous streaming RL framework for efficient LLM post-training, 2025. URL https://arxiv.org/abs/2507.01663.
Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. SWE-bench: Can language models resolve real-world GitHub issues? In International Conference on Learning Representations, pp. 54107–54157, 2024. URL https: //proceedings.iclr.cc/paper_files/paper/2024/f ile/edac78c3e300629acfe6cbe9ca88fb84-Paper-Con ference.pdf.
Bowei He, Yankai Chen, Xiaokun Zhang, and Xue Liu. Branching policy optimization: Sandbox-native language agent reinforcement learning. In World Artificial Intelligence Conference Academic (WAICA), 2026a. URL https://arxiv.org/abs/2607.14171. Main Conference.
Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In The Third International Conference on Learning Representations, 2015. URL http s://arxiv.org/abs/1412.6980. Haoyang Li, Sheng Lin, Fangcheng Fu, Yuming Zhou, Xiaodong Ji, Yanfeng Zhao, Lefeng Wang, Jie Jiang, and Bin Cui. StaleFlow: Staleness-aware data management for mitigating data skewness in fully disaggregated RL post-training, 2026. URL https://arxiv.org/abs/ 2601.12784. To appear at ACM SIGMOD 2027.
Zhiyuan He, Siwei Zhang, Zhiwen Zhou, Yuqing Yang, Yu Kang, Yuge Zhang, Luna K. Qiu, Tin Yan Tsui, Jiahang Xu, and Chong Luo. Agent Lightning v1.0: Towards harnessed agentic RL, 2026b. URL https: //arxiv.org/abs/2608.17528.
14
John D. C. Little. A proof for the queuing formula: L = λW. Operations Research, 9(3):383–387, 1961. doi: 10.1 287/opre.9.3.383. URL https://doi.org/10.1287/op re.9.3.383.
Bo Qian, Yuting Wu, Shuang Zeng, Huaiyu Wan, Dalin Zhang, and Jiqiang Liu. MileGPO: Milestone inference with local evidence for graph-based policy optimization of long-horizon LLM agents, 2026. URL https://arxiv.org/abs/2608.19803.
Michael Luo, Naman Jain, Jaskirat Singh, Sijun Tan, Ameen Patel, Qingyang Wu, Alpay Ariyak, Colin Cai, Tarun Venkat, Shang Zhu, Ben Athiwaratkun, Manan Roongta, Ce Zhang, Li Erran Li, Raluca Ada Popa, Koushik Sen, and Ion Stoica. DeepSWE: Training a fully open-sourced, state-of-the-art coding agent by scaling RL. Together AI research blog, July 2025a. URL https://www.together.ai/blog/deepswe.
Ruoyu Qin, Weiran He, Weixiao Huang, Yangkun Zhang, Yikai Zhao, Bo Pang, Xinran Xu, Yingdi Shan, Yongwei Wu, and Mingxing Zhang. Seer: Online context learning for fast synchronous LLM reinforcement learning. In 20th USENIX Symposium on Operating Systems Design and Implementation (OSDI 26), pp. 883–901, Seattle, WA, 2026. USENIX Association. ISBN 978-1-93913355-7. URL https://www.usenix.org/conference/os di26/presentation/qin.
Xufang Luo, Yuge Zhang, Zhiyuan He, Zilong Wang, Siyun Zhao, Dongsheng Li, Luna K. Qiu, and Yuqing Yang. Agent Lightning: Train ANY AI agents with reinforcement learning, 2025b. URL https://arxiv. org/abs/2508.03680.
Bo Ren, Yirong Mao, Yi Yang, and Wenhui Que. IAPO: Influence-aware policy optimization for credit assignment in multi-turn service agents, 2026a. URL https: //arxiv.org/abs/2608.24588.
Zhiyu Mei, Wei Fu, Kaiwei Li, Guangju Wang, Huanchen Zhang, and Yi Wu. ReaL: Efficient RLHF training of large language models with parameter reallocation. In Proceedings of Machine Learning and Systems, volume 7. MLSys, 2025. URL https://proceedings.mlsys.org/ paper_files/paper/2025/hash/3b3889d313ba9476c1 2c2d77ea66b24f-Abstract-Conference.html.
Yanyu Ren, Xizheng Wang, Xiao Liu, Bowen Lv, Hanchen Zhang, Shudan Zhang, Hanyu Lai, Shuai Wang, Li Chen, Dan Li, and Jie Tang. TideRL: Boosting agentic RL goodput with readiness-aware scheduling, 2026b. URL https://arxiv.org/abs/2608.10402. Kai Ruan, Jinghao Lin, Qianshan Wei, Ziqi Zhou, and Zihe Huang. SPO++: Stream-aligned policy optimization for asynchronous agentic RL, 2026. URL https://arxiv.org/abs/2608.24870.
Mike Merrill, Alexander Shaw, Nicholas Carlini, Boxuan Li, Harsh Raj, Ivan Bercovich, Lin Shi, Jeong Shin, Thomas Walshe, E. Kelly Buchanan, Junhong Shen, Guanghao Ye, Haowei Lin, Jason Poulos, Maoyu Wang, Marianna Nezhurina, Di Lu, Orfeas Menis Mastromichalakis, Zhiwei Xu, Zizhao Chen, Yue Liu, Robert Zhang, Leon Liangyu Chen, Anurag Kashyap, Jan-Lucas Uslu, Jeffrey Li, Jianbo Wu, Minghao Yan, Song Bian, Vedang Sharma, Ke Sun, Steven Dillmann, Akshay Anand, Andrew Lanpouthakoun, Bardia Koopah, Changran Hu, Etash Guha, Gabriel Dreiman, Jiacheng Zhu, Karl Krauth, Li Zhong, Niklas Muennighoff, Robert Amanfu, Shangyin Tan, Shreyas Pimpalgaonkar, Tushar Aggarwal, Xiangning Lin, Xin Lan, Xuandong Zhao, Yiqing Liang, Yuanli Wang, Zilong (Ryan) Wang, Changzhi Zhou, David Heineman, Hange Liu, Harsh Trivedi, John Yang, Junhong Lin, Manish Shetty, Michael Yang, Nabil Omi, Negin Raoof, Shanda Li, Terry Yue Zhuo, Wuwei Lin, Yiwei Dai, Yuxin Wang, Wenhao Chai, Shang Zhou, Dariush Wahdany, Ziyu She, Jiaming Hu, Zhikang Dong, Yuxuan Zhu, Sasha Cui, Ahson Saiyed, Arinbjörn Kolbeinsson, Christopher Rytting, Ryan Marten, Yixin Wang, Jenia Jitsev, Alex Dimakis, Andy Konwinski, and Ludwig Schmidt. Terminal-Bench: Benchmarking agents on hard, realistic tasks in command line interfaces. In International Conference on Learning Representations, volume 2026, pp. 40903–40986, 2026. URL https://proceedings.iclr.cc/paper_files/paper/ 2026/file/444a3737adaee10d86ad2ef5f74468e6-Pap er-Conference.pdf.
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo. DeepSeekMath: Pushing the limits of mathematical reasoning in open language models, 2024. URL https: //arxiv.org/abs/2402.03300. Guangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu, Wang Zhang, Ru Zhang, Yanghua Peng, Haibin Lin, and Chuan Wu. HybridFlow: A flexible and efficient RLHF framework. In Proceedings of the Twentieth European Conference on Computer Systems, pp. 1279–1297. ACM, 2025. doi: 10.1145/3689031.3696075. URL https://doi.org/10.1145/3689031.3696075. Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. Megatron-LM: Training multi-billion parameter language models using model parallelism, 2019. URL https://arxiv.org/abs/1909.08053. Huatong Song, Fei Bai, Ming Yang, Renyuan Li, Jia Deng, Jujie He, Zhange Zhang, Daixuan Cheng, Yan Xing, Qi Yun, Xuxing Chen, Danyang Li, Feng Chang, Chuan Hao, Ran Tao, Jian Yang, Bryan Dai, Wayne Xin Zhao, Mingjie Tang, and Ji-Rong Wen. ClawGym II: Exploring black-box RL on agent harness, 2026a. URL https://arxiv.org/abs/2608.16798. Huatong Song, Lisheng Huang, Shuang Sun, Jinhao Jiang, Ran Le, Daixuan Cheng, Guoxin Chen, Yiwen Hu, Zongchao Chen, Yiming Jia, Wayne Xin Zhao, Yang Song, Tao Zhang, and Ji-Rong Wen. SWEMaster: Unleashing the potential of software engineering agents via post-training, 2026b. URL https: //arxiv.org/abs/2602.03411.
Yingqi Peng, Jiawei Zhang, Wenhao Zhou, Ruida Xu, Ran Yan, Wei Dong, Yi Gao, Zhiqiang Ding, Tongkai Yang, and Binhang Yuan. AReaL-DTE: Sparse policyweight transfer for online agentic reinforcement learning, 2026. URL https://arxiv.org/abs/2608.00455. 15
Zhiqiang Tan, Maoxin Wang, Sijie Wang, Yiming Yin, Qiang Wang, Xiaowen Chu, and Shaohuai Shi. Bidirectional resource scheduling for disaggregated and asynchronous RL post-training, 2026. URL https: //arxiv.org/abs/2607.09207.
Li, Ye Bai, Jun Zhang, Lei Chen, Yue Li, and Mingchen Wan. TailSieve: Partial-rollout-guided tail routing for LLM rollouts, 2026b. URL https://arxiv.org/abs/ 2608.22788. John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. SWE-agent: Agent-computer interfaces enable automated software engineering. In Advances in Neural Information Processing Systems, volume 37, pp. 50528– 50652. Curran Associates, Inc., 2024. doi: 10.52202/0 79017-1601. URL https://proceedings.neurips.cc /paper_files/paper/2024/hash/5a7c947568c1b1328 ccc5230172e1e7c-Abstract-Conference.html.
Jinghui Wang, Shaojie Wang, Yinghan Cui, Xuxing Chen, Chao Wang, Liang Huang, Can Tang, Xiaojiang Zhang, Junyi Peng, Li Wan, Haotian Zhang, and Bin Chen. Tree training: Accelerating agentic LLMs training via shared prefix reuse, 2025a. URL https://arxiv.org/ abs/2511.00413. Weixun Wang, Shaopan Xiong, Gengru Chen, Wei Gao, Sheng Guo, Yancheng He, Ju Huang, Jiaheng Liu, Zhendong Li, Xiaoyang Li, Zichen Liu, Haizhou Zhao, Dakai An, Lunxi Cao, Qiyang Cao, Wanxi Deng, Feilei Du, Yiliang Gu, Jiahe Li, Xiang Li, Mingjie Liu, Yijia Luo, Zihe Liu, Yadao Wang, Pei Wang, Tianyuan Wu, Yanan Wu, Yuheng Zhao, Shuaibing Zhao, Jin Yang, Siran Yang, Yingshui Tan, Huimin Yi, Yuchi Xu, Yujin Yuan, Xingyao Zhang, Lin Qu, Wenbo Su, Wei Wang, Jiamang Wang, and Bo Zheng. Reinforcement learning optimization for large-scale learning: An efficient and user-friendly scaling library, 2025b. URL https://arxiv.org/abs/2506.06122.
John Yang, Kilian Lieret, Jeffrey Ma, Parth Thakkar, Dmitrii Pedchenko, Sten Sootla, Emily McMilin, Pengcheng Yin, Rui Hou, Gabriel Synnaeve, Diyi Yang, and Ofir Press. ProgramBench: Can language models rebuild programs from scratch?, 2026. URL https://arxiv.org/abs/2605.03546. Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. ReAct: Synergizing reasoning and acting in language models. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/for um?id=WE_vluYUL-X.
Zihan Wang, Kangrui Wang, Qineng Wang, Pingyue Zhang, Linjie Li, Zhengyuan Yang, Xing Jin, Kefan Yu, Minh Nhat Nguyen, Licheng Liu, Eli Gottlieb, Yiping Lu, Kyunghyun Cho, Jiajun Wu, Fei-Fei Li, Lijuan Wang, Yejin Choi, and Manling Li. RAGEN: Understanding self-evolution in LLM agents via multi-turn reinforcement learning, 2025c. URL https://arxiv.org/abs/2504.20073.
Mianjie Yu, Zizhao Mo, Huanyu Qu, Zhirong Qian, Huanle Xu, Cen Li, Zifeng Zhao, Zhi Zhou, Jinhua Zhou, Jun Xie, and Chengzhong Xu. psRL: Efficient training for agentic AI via training-time prefix sharing, 2026. URL https://arxiv.org/abs/2608.25683. Hanchen Zhang, Xiao Liu, Bowen Lv, Xueqiao Sun, Bohao Jing, Iat Long Iong, Zhenyu Hou, Zehan Qi, Hanyu Lai, Yifan Xu, Rui Lu, Hongning Wang, Jie Tang, and Yuxiao Dong. AgentRL: Scaling agentic reinforcement learning with a multi-turn, multi-task framework, 2025. URL https://arxiv.org/abs/2510 .04206.
Tianyuan Wu, Lunxi Cao, Yining Wei, Wei Gao, Yuheng Zhao, Dakai An, Shaopan Xiong, Zhiqiang Lv, Ju Huang, Siran Yang, Yinghao Yu, Jiamang Wang, Lin Qu, and Wei Wang. Weave: Efficient Co-Scheduling for disaggregated RL Post-Training. In 20th USENIX Symposium on Operating Systems Design and Implementation (OSDI 26), pp. 809–827, Seattle, WA, 2026a. USENIX Association. ISBN 978-1-939133-55-7. URL https://www.usenix.org/conference/osdi26/prese ntation/wu-tianyuan.
Lei Zhang, Mouxiang Chen, Ruisheng Cao, Jiawei Chen, Fan Zhou, Yiheng Xu, Jiaxi Yang, Zeyao Ma, Liang Chen, Changwei Luo, Kai Zhang, Fan Yan, KaShun Shum, Jiajun Zhang, Zeyu Cui, Feng Hu, Junyang Lin, Binyuan Hui, and Min Yang. MegaFlow: Large-scale distributed orchestration system for the agentic era, 2026a. URL https://arxiv.org/abs/2601.07526.
Yongji Wu, Xueshen Liu, Haizhong Zheng, Juncheng Gu, Beidi Chen, Z. Morley Mao, Arvind Krishnamurthy, and Ion Stoica. RLBoost: Harvesting preemptible cloud resources for cost-efficient reinforcement learning on LLMs. In 23rd USENIX Symposium on Networked Systems Design and Implementation (NSDI 26), pp. 1323– 1339, Renton, WA, 2026b. USENIX Association. ISBN 978-1-939133-54-0. URL https://www.usenix.org/c onference/nsdi26/presentation/wu-yongji.
Zili Zhang, Yinmin Zhong, Chengxu Yang, Chao Jin, Bingyang Wu, Xinming Wei, Yuliang Liu, and Xin Jin. Heddle: A distributed orchestration system for agentic RL rollout, 2026b. URL https://arxiv.org/abs/2603 .28101.
Binfeng Xu, Hao Zhang, Shaokun Zhang, Songyang Han, Mingjie Liu, Jian Hu, Shizhe Diao, Zhenghui Jin, Yunheng Zou, Michael Demoret, Jan Kautz, and Yi Dong. Polar: Agentic RL on any harness at scale, 2026a. URL https://arxiv.org/abs/2605.24220.
Chujie Zheng, Shixuan Liu, Mingze Li, Xiong-Hui Chen, Bowen Yu, Chang Gao, Kai Dang, Yuqiong Liu, Rui Men, An Yang, Jingren Zhou, and Junyang Lin. Group sequence policy optimization, 2025. URL https://ar xiv.org/abs/2507.18071.
Tianqi Xu, Lu Lv, Haoyang Huang, Wenjie Huang, Zhanming Shen, Yuhao Shen, Baolin Zhang, Xinyi Hu, Shuang Ge, Jun Dai, Tianyu Liu, Suorong Yang, Zhikai
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark 16
Barrett, and Ying Sheng. SGLang: Efficient execution of structured language model programs. In Advances in Neural Information Processing Systems, volume 37, pp. 62557–62583. Curran Associates, Inc., 2024. doi: 10.52202/079017- 2000. URL https: //doi.org/10.52202/079017-2000. Yinmin Zhong, Zili Zhang, Xiaoniu Song, Hanpeng Hu, Chao Jin, Bingyang Wu, Nuo Chen, Yukun Chen, Yu Zhou, Changyi Wan, Hongyu Zhou, Yimin Jiang, Yibo Zhu, and Daxin Jiang. StreamRL: Scalable, heterogeneous, and elastic RL for LLMs with disaggregated stream generation, 2025. URL https://arxiv.org/ab s/2504.15930. Yuzhen Zhou, Jiajun Li, Yusheng Su, Gowtham Ramesh, Zilin Zhu, Xiang Long, Chenyang Zhao, Jin Pan, Xiaodong Yu, Ze Wang, Kangrui Du, Jialian Wu, Ximeng Sun, Jiang Liu, Qiaolin Yu, Hao Chen, Zicheng Liu, and Emad Barsoum. APRIL: Active partial rollouts in reinforcement learning to tame long-tail generation, 2025. URL https://arxiv.org/abs/2509.18521. Mingchen Zhuge, Changsheng Zhao, Dylan Ashley, Wenyi Wang, Dmitrii Khizbullin, Yunyang Xiong, Zechun Liu, Ernie Chang, Raghuraman Krishnamoorthi, Yuandong Tian, Yangyang Shi, Vikas Chandra, and Jürgen Schmidhuber. Agent-as-a-judge: Evaluate agents with agents, 2024. URL https://arxiv.org/ab s/2410.10934.
17
A
Analysis of Dispatch Policies
training or the next publication. Thus Qc counts only unfinished groups; ready groups wait outside these slots. Equal depths do not imply equal total unconsumed populations. Figure 10 contrasts the resulting slot occupancy and replenishment times.
We first establish exact staleness identities, then derive step times under a timing model and compare dispatch policies. Results that also require the ready-rate approximation state it explicitly. Counts and rates use rollout groups; subscripts b and c denote boundary and continuous dispatch, respectively.
Training ends Publish v = 1
v=0
Slot 1
Table 3: Notation for dispatch and step-time analysis.
G1
Meaning
B, µ φ Qb
Groups per update; updates per burst. Extra boundary dispatch in units of B groups. Boundary depth after replenishment: all ( φ + µ) B unconsumed groups. Continuous depth: unfinished groups only. Published version at admission; trainer version immediately before consumption. Staleness vt − vd ; denotes its boundary mean in Section A.5. Mean ready, unconsumed groups at publication. Mean time from admission to readiness, including rollout and evaluation. Active training rate (groups per unit time). Mean update interval, including data waiting (µ = 1). Mean admission phase between publications under continuous dispatch. Boundary completion offset from the start of the step before the consuming step, averaged and divided by Tb .
Qc vd , vt d R L vT Tb , Tc α β
A.1
G5, vd = 1
G2 held G3
Slot 3
Qc = 3
(b) Continuous dispatch Training ends Publish v = 1
v=0
Slot 1 Slot 2
G4, vd = 0
G1
G5, vd = 0
G2 G3
Slot 3 Rollout / eval.
Ready / selected
Figure 10: Illustrative dispatch timelines for B = 2, µ = 1, and Qb = Qc = 3. Rows track logical slots. G1/G2 train together once both are ready. Dashed lines mark training completion and publication of version 1. (a) Boundary dispatch then admits G4/G5 under version 1. (b) Continuous dispatch admits each replacement as soon as its predecessor becomes ready, under version 0. Gold segments retain ready or selected groups; rollouts reaching the right edge continue beyond the plotted interval.
Assumption A.1 (Population accounting). Batch size B, burst size µ, and dispatch depths are fixed. Each admitted group is eventually consumed exactly once. Queue populations and relative policy ages are stationary, with finite expected total age of unconsumed groups at publication.
Model and Accounting Assumptions
Each optimizer update consumes B distinct groups and advances the trainer’s version by one. A burst contains µ updates and publishes weights after its final update (Sections 2.1 and 4.2). As in Section 2.2, scheduling staleness is d = vt − vd , (8)
The counting results require neither FIFO nor a particular latency distribution. Finite-window balances include a correction for the change in outstanding staleness. For long-run empirical means, this change per consumed group must vanish; bounded group counts alone do not suffice.
where vd is the latest published version at admission and vt is the trainer’s version immediately before consumption. The consuming update is excluded; the metric tracks admission timing, not each model call’s behaviorpolicy lag.
A.2
Boundary dispatch. Each admitted group retains its dispatch slot until training consumes it. The queue is filled at startup; thereafter, the µB slots freed within a burst are refilled together after its weight publication, using the updated policy. Rollout replenishment therefore remains coupled to training completion and weight publication. The target depth includes unfinished rollouts and ready or selected training groups, and is restored to Qb = ( φ + µ) B.
G4, vd = 1
G1 held G2
Slot 2
Symbol
Qb = 3
(a) Boundary dispatch
Exact Staleness Identities
Theorem A.1 (Boundary staleness). Under Assumption A.1, boundary dispatch has group-weighted mean scheduling staleness Eb [ d ] =
Qb µ+1 µ−1 − = φ+ . B 2 2
(10)
Proof. Successive updates within a burst leave Qb − B, Qb − 2B, . . . , Qb − µB surviving groups. Each survivor gains one unit of age, adding µQb − Bµ(µ + 1)/2 in total. Replacements enter only after publication, at age zero. Consumed staleness equals this added age minus the change in outstanding age. In stationarity the latter has mean zero. Dividing by the µB consumed groups proves the result.
(9)
Continuous dispatch. In the asynchronous setting, continuous dispatch fully decouples rollout replenishment from training. A group releases its slot once rollout and evaluation finish. A replacement starts immediately using the latest published policy, without waiting for
18
Corollary A.2 (Colocate without extra dispatch). Setting φ = 0 in Theorem A.1 gives
By Little’s law (Little, 1961), continuous dispatch has ready rate νR = Qc /L. Subsequent expressions use Qc and L directly.
µ−1 . (11) 2 Corollary A.3 (Async with one step per burst). Setting µ = 1 in Theorem A.1 gives E[ d ] =
E[ d ] =
Qb − B = φ. B
Assumption A.3 (Ready-rate approximation). After reserving a burst’s groups, older ready backlog is negligible. The completion rate during training is close to its long-run mean Qc /L. Corollary A.6 (Approximate ready count and staleness). Under Assumptions A.1–A.3, the publication-time ready count and mean continuous staleness are approximated by
(12)
In particular, φ = 0 and µ = 1 give zero staleness regardless of rollout duration: wall-clock waiting does not advance the version. Completion order and GPU placement affect when updates occur, but leave the counting identity unchanged.
R≃ Ec [ d ] ≃
For continuous dispatch, let R be the mean ready, unconsumed group count at weight publication. It is sampled once per publication, not averaged over wall-clock time. Theorem A.4 (Continuous staleness). Under Assumption A.1, continuous dispatch has Ec [d] =
Qc + R µ − 1 + . B 2
(13)
A.4
Mean Step Time for Async
Set µ = 1. Let Tb and Tc be the mean intervals between completed updates, including data waiting. The active training duration B/v T is fixed, so these also equal the mean intervals between training starts.
Qc µ−1 + . (14) B 2 For µ = 1 and equal depths Qc = Qb , subtracting the two identities gives Ec [ d ] =
Completion phase. For each boundary-dispatched group, measure completion from the start of the step preceding the one that consumes it. Define β as the mean offset divided by Tb . For equal step lengths, one marks the consuming training start and one half the midpoint. For variable lengths, this is a ratio of means, not a mean of fractions. Groups ready before the preceding start have negative offsets.
(15)
The R = 0 case describes the idealized resumable Colocate reference: rollout stops when µB groups are ready, and unfinished groups pause while those groups train. The live-harness constraint in Section 2.1 can instead require draining rollouts, whose publication-time population must be counted separately. A.3
(17)
Assumption A.3 does not follow from Qc /L < v T . Training starts when enough data are ready, so its completion rate can differ from the rate at arbitrary times. Timeout clusters and persistent ready backlog can invalidate the approximation.
Corollary A.5 (Consequences for continuous dispatch). With no ready carryover at publication, R = 0 gives
R . B
Qc µ − 1 µQc . + + B 2 Lv T
(16)
Proof. A burst lasts µB/v T . Under Assumption A.3, multiplying this duration by Qc /L estimates the groups that complete and remain ready at publication. Reserved groups have already been consumed. Substitution into Theorem A.4 gives the staleness estimate.
Proof. Measure age against the latest published version; new groups enter at zero age in this coordinate. Each publication advances the version by µ and adds mean age µ( Qc + R) to survivors. In stationarity, training removes the same published-version age per burst. Actual staleness also includes positions within the burst, which contribute Bµ(µ − 1)/2 in total. Dividing by µB gives the identity.
Ec [ d ] − E b [ d ] = 1 +
Qc µB , L vT
Theorem A.7 (Continuous-dispatch step time). Under Assumptions A.1 and A.2 with µ = 1, continuous dispatch has mean step time BL Tc = . (18) Qc
Timing Model and Ready-Rate Approximation Proof. Continuous flow balance gives Tc = B/( Qc /L).
The remaining timing results concern the Async reference. Assumption A.2 (Timing model). Each admitted group starts rollout immediately. Durations through rollout and evaluation are independent draws from a common distribution under both policies, with finite mean L independent of depth. The trainer reserves µB ready groups before a burst and executes its updates at fixed throughput v T , while rollout continues. Dispatch and publication overheads are negligible. Mean waiting and step times are finite, and Qc /L < v T .
Theorem A.8 (Boundary-dispatch step time). Under Assumptions A.1 and A.2 with µ = 1, boundary dispatch has mean step time Tb =
L + B/v T . Qb /B − 1 + β
(19)
Proof. For boundary dispatch, every update consumes B groups, so mean ready waiting is (1 − β) Tb . A slot 19
remains occupied for the rollout duration L, this ready wait, and training time B/v T . Little’s law therefore gives B B Qb = L + (1 − β) Tb + . Tb vT
Corollary A.10 (Approximate depth relation). Under Assumptions A.1–A.3 with µ = 1, matching interpolated staleness gives B(d + α) Qc ≃ . (22) 1 + B/( Lv T )
Solving for Tb gives the result.
Proof. Substitute R ≃ BQc /( Lv T ) into Theorem A.9 and collect the terms in Qc .
A later completion phase shortens the time a finished group holds its boundary slot. A.5
Theorem A.11 (Exact step-time comparisons). Under Assumptions A.1 and A.2 with µ = 1, equal interpolated staleness gives
Comparison at Matched Staleness
Tc d+β = . Tb d + α − R/B + Qc /( Lv T )
Continue with µ = 1. Admissions between publications share the same integer version vd , although continuous dispatch can admit groups later in that interval. For each admission, adding the fraction of the interval already elapsed to vd defines an interpolated dispatch coordinate. Subtracting this fraction from a group’s staleness gives its interpolated staleness; the consuming trainer version remains discrete. Figure 11 illustrates the two metrics for the same group. v=1
v=0
Policy
+1
Tc d+β . = Tb d − R/B + Qc /( Lv T )
Train
+1
+1
=3
Tc BL d + β = Tb Qc L + B/v T B(d + β) = Qc 1 + B/( Lv T )
+1
+1
= 2.5
omit
+0.5
Interpolated Rollout / eval.
Ready / selected
=
Corollary A.12 (Approximate time comparisons). If Assumption A.3 also holds, equal mean interpolated staleness gives Tc /Tb ≃ (d + β)/(d + α). At equal mean discrete staleness d > 0, the ratio is approximately 1 + β/d.
Let α be the mean admission fraction under continuous dispatch, with 0 ≤ α < 1. Boundary admissions have fraction zero. In this subsection, reuse d for the boundary policy’s mean staleness. At equal mean interpolated staleness, the mean discrete stalenesses are d for boundary dispatch and d + α for continuous dispatch. The two counting identities therefore give Qb − 1, B
d+α =
Qc + R . B
Proof. The boundary depth gives the exact step time Tb =
(20)
BL Qc BL 1 + B/( Lv T ) ≃ B(d + α) L 1 + B/( Lv T ) = d+α L + B/v T = . d+α
Tc =
The phase α is measured, not assumed to equal one half.
Qc = B(d + α) − R.
L + B/v T L + B/v T = . Qb /B − 1 + β d+β
For continuous dispatch, substitute the depth estimate from Corollary A.10:
Theorem A.9 (Depths at equal interpolated staleness). Under Assumption A.1 with µ = 1, matching the mean interpolated staleness at d requires Qb = B ( d + 1),
d+β . Qc /B + Qc /( Lv T )
For interpolated staleness, Theorem A.9 gives Qc /B = d + α − R/B. For discrete staleness, the counting identities instead give Qc /B = d − R/B. Substituting these relations proves both expressions. Their denominators are positive in the compared systems.
Train
Figure 11: Discrete and interpolated staleness on a shared timeline (µ = 1). The group is admitted halfway between publications 0 and 1 and consumed at vt = 3. With equally spaced publications shown for illustration, the discrete row counts 1 + 1 + 1 = 3, while the interpolated row excludes the pre-admission fraction, giving 0.5 + 1 + 1 = 2.5. Training continues beyond the right edge.
d=
(24)
Proof. Since Qb = B(d + 1), the boundary denominator is Qb /B − 1 + β = d + β. Theorems A.7 and A.8 therefore give
v=3
Ready
Rollout / eval.
Discrete
Matching mean discrete staleness at d > 0 instead gives
Consume (vt = 3)
Admit
Group
v=2
(23)
(21)
Proof. Rearrange the two balances in Equation 20. Eliminating d also gives Qc = Qb − (1 − α) B − R.
Only the depth substitution uses the ready-rate approximation; the remaining steps are algebraic. Taking the
20
ratio gives Tc ( L + B/v T )/(d + α) ≃ Tb ( L + B/v T )/(d + β) L + B/v T d + β = d + α L + B/v T d+β = . d+α
(25)
The common factor L + B/v T cancels because both policies share L and v T under Assumption A.2. For equal discrete staleness, the continuous mean is d rather than d + α. The same calculation gives Tc ≃ ( L + B/v T )/d, hence Tc /Tb ≃ (d + β)/d = 1 + β/d. Matching phases α = β thus gives Tc ≃ Tb under the interpolated comparison. For β ≥ 0, the discrete comparison’s approximate ratio is at least one. Both conclusions require Assumption A.3; matching staleness alone does not imply equal step times. Scheduling implications. Continuous dispatch maintains Qc unfinished groups, while boundary dispatch’s unfinished population falls between publications. Its steadier rollout load comes with the equal-depth staleness cost in Corollary A.5. Elastic placement can stabilize per-instance concurrency under boundary dispatch by reducing the rollout pool as unfinished work falls and moving released cells to training (Section 4.2), preserving its staleness accounting. Scope. The counting identities concern groupweighted mean staleness; worst-case and tokenweighted guarantees require separate analysis. Drops, reuse, or varying B and µ require new accounting. Filtering within a group leaves the identities unchanged if each group still counts once. Resource saturation, concurrency-dependent latency, and data stalls within updates fall outside Assumption A.2. Approximate depth and time comparisons additionally require Assumption A.3.
21
B
RL Algorithm Details
Loss normalization. Using the token losses in Equation 28, weighting each trajectory by its target share Lr /Li and averaging over its targets gives
This appendix specifies the GSPO adaptation summarized in Section 6.1, including its execution-level clipping, token-level importance correction, loss normalization, and training hyperparameters.
Lr 1 1 ℓu (θ ) = · ∑ ∑ ℓ u ( θ ), L L L r i i u∈T r ∈R u∈T
Li ( θ ) = ∑
i
(29)
i
Equation 29 gives every admitted target the same normalization coefficient 1/Li ; importance weighting and clipping do not change this denominator. Executions with Li = 0 contribute zero loss. A batch averages over all B queries and their N original rollout executions:
Training configuration. We use Group Sequence Policy Optimization (GSPO) (Zheng et al., 2025) with tokenlevel importance correction. Each rollout group contains N = 16 independent rollout executions. Each training step consumes a batch of B = 32 distinct rollout groups for NL2RepoBench or B = 48 for DeepSWE and TerminalBench. The remaining RL hyperparameters are shared across all experiments.
L(θ ) =
All configurations use Adam (Kingma & Ba, 2015) with a constant learning rate of 10−6 , β 1 = 0.9, β 2 = 0.95, ϵ = 10−15 , weight decay 0.01, and gradient clipping at norm 1.0. KL and entropy regularization are disabled.
1 B N ∑ Li(q,k) (θ ), B N q∑ =1 k =1
(30)
where i (q, k ) is query q’s k-th execution. Equation 30 preserves each execution’s aggregate normalization weight regardless of how many trajectories it contributes. For a fixed admitted target set Ti , changing which selected path carries a shared target leaves the execution’s sequence ratio, clipping decision, and loss unchanged. Sampling can change Ti itself; the invariance concerns the assignment and ordering of a fixed set of targets. Packing preserves these per-execution weights.
Execution-level targets. For execution i, let Ri be its admitted trajectories. For each r ∈ Ri , the target set Tr is fixed to the trainable tokens still unmasked when r is drawn (Section 5.3). SOn-tree masking makes these sets disjoint. Let Ti = r∈Ri Tr , Lr = |Tr |, and Li = |Ti | = ∑r∈Ri Lr . The sequence used for importance-ratio computation and clipping is this entire execution’s set of admitted, deduplicated trainable tokens. Each token retains its original generation context. Tokens excluded by provenance, role, or the Jmax cutoff contribute neither to the sequence ratio nor to the loss. Importance correction and clipping. Let yu be target token u, cu its recorded conditioning context, and proll u its recorded behavior probability, reused without recomputation. Let sg denote stop-gradient. For Li > 0, the token importance ratio and detached execution-level sequence ratio are πθ (yu | cu ) , proll u " !# 1 si (θ ) = sg exp log ρu (θ ) . Li u∑ ∈T
r
ρu (θ ) =
(26)
i
The detached sequence ratio in Equation 26 is used for monitoring and determines a shared clipping mask according to execution i’s group-relative advantage Âi : 0, Âi > 0 and si (θ ) > 1.005, mi (θ ) = 0, Âi < 0 and si (θ ) < 0.995, (27) 1, otherwise. The token-specific importance weight is ρ̄u (θ ) = sg(min{ρu (θ ), 5}). These weights may differ within an execution; si affects the loss only through the common mask in Equation 27. The token loss, upper-bounded by 100, is ℓu (θ ) = min 100, −mi (θ )ρ̄u (θ ) Âi log πθ (yu | cu ) . (28)
22
C
Case Study: One Execution, Multiple Training Trajectories
C.3
One response attempts four Agent calls, but the harness admits only two concurrent sub-agents. Two tool results report the concurrency limit. Subsequent calls start the remaining work, but the first FDW attempt is interrupted after a 53-node path. Later, a local test run reports OK for 259 tests, with one skipped, yet an import check cannot find the FDW, dbt, pipeline, and test-support packages. The main agent launches another attempt, which contributes a separate 202-node path. Rejection, partial execution, and retry thus affect the record differently within one rollout.
We examine execution f3ad38e5, in which an agent uses Claude Code to build a Python analytics SDK from a natural-language specification. The task requires nine packages, including catalog services, data frames, Flight connectivity, and pipeline utilities, implemented from scratch in a shared workspace. Figure 12 combines its prefix structure with recorded messages to illustrate the trajectory-processing problem of Section 5. C.1
Execution and Context Structure
The leaf tagged is_final_main ends with renewed delegation: it identifies the last recorded main trajectory, not successful completion. An interrupted sub-agent likewise does not determine the execution’s reward. Section 5.2 requires the designated evaluator to assess the preserved workspace; a group becomes trainable only after its executions close with valid outcomes. The visualization supplies no designated task reward. The recovery sequence also motivates preserving harness and workspace state during the request rerouting of Section 4.3.
The trace reports 636 proxy requests and 1,347 unique message nodes across ten candidate trajectories. The main agent produces two task-driving trajectories and one summary. Five sub-agent invocations produce six task-driving trajectories and one summary: the catalog sub-agent compacts its history, while the FDW work item has two attempts. These visualization nodes count messages, distinct from the data plane’s request increments and trainable tokens. Main and sub-agent contexts have separate roots because their system prompts differ. Delegation can be identified separately by matching an Agent call’s prompt to a sub-agent’s initial user message after removing harness reminders; the export also identifies a sub-agent billing flag. Causal parentage therefore differs from contextprefix sharing. Concatenating the two agents’ conversations would give outputs conditioning messages absent during generation. C.2
Delegation, Interruption, and Outcome Validity
C.4
Applying the Training Admission Rule
The priority classes are two main, one main-summary, six sub-agent, and one sub-agent-summary. For illustration, suppose Jmax = 4 and each path has eligible policygenerated targets. The rule in Section 5.3 admits both main trajectories, the main-summary trajectory, and one sub-agent trajectory. Selection within each class is proportional to remaining trainable tokens, with shared targets masked after every draw. The figure’s message counts cannot determine these token probabilities.
Compaction and Shared Prefixes
Both the main agent and the catalog sub-agent compact their contexts. After their system prompts, they have shared prefixes of 168 and 193 messages, respectively. Each prefix forks into a task-driving assistant response and a summarization request–response pair. Continuation after compaction starts a separate path under the system prompt, with a new user message containing the summary.
A flat mean over all ten leaf losses would allocate eight equal coefficients to delegated or summary paths and repeat shared targets. In Q WEN G YRE, selected paths inherit the original execution’s group-relative advantage. Token losses are averaged over disjoint targets within each execution, then over the original rollout executions. Additional compaction or delegation changes the admitted targets without creating independent rewards or increasing the execution’s aggregate normalization weight.
The ten complete paths contain 53–293 message nodes each. Expanding them independently produces 1,716 node occurrences, compared with 1,347 shared nodes. The 369 extra occurrences repeat shared prefixes and system prompts; their token volume depends on message lengths. If a task-driving leaf and its summary leaf are both selected, their shared policy outputs must contribute loss once. Section 5.3’s on-tree masking retains the prefix as context while suppressing repeated targets. Provenance also matters: the original summary can be a target if the rollout policy generated it, whereas its reinsertion into the continuation’s user prompt is conditioning input. TITO preserves the original token contexts and distinguishes these two occurrences.
23
Execution f3ad38e5 System context
Main system prompt 1 node
SYS[0] “… helps users with software engineering tasks.”
636 proxy requests · 1,347 unique message nodes · 10 candidate trajectories Prefix or continuation
Main prefix
Continuation after a fork
168 nodes
USR[1] “You MUST write all code from scratch using your own coding abilities.”
Main trajectory
1 node · path 170
ASST[169] Write: "file_path": "/workspace/acloud_sdk/catalog/entity.py"
Main summary
2 nodes · path 171
USR[170] “CRITICAL: Respond with TEXT ONLY. Do NOT call any tools.” ASST[171] “… with 9 Python packages based on an exhaustive spec in start.md.”
Main after compaction (last main)
171 nodes · path 172
USR[172] “This session is being continued from a previous conversation that ran out of context.” TOOL[857] “Concurrent subagent limit reached. You can run 2 subagents at once.” TOOL[1036] “[Request interrupted by user for tool use]” TOOL[1142] “Ran 259 tests in 0.029s”; “OK (skipped=1)” TOOL[1144] “FAIL acloud_fdw -> ModuleNotFoundError No module named 'acloud_fdw'” ASST[1145] Agent: “Build fdw, dbt, pipelines, test_support”
Sub-agent prompt
Catalog prefix
SYS[226] “You are an agent for Claude Code …”
USR[229] “You are contributing to an original, from-scratch reimplementation …”
1 node
193 nodes
Catalog trajectory
1 node · path 195
ASST[560] “Now let me write the workspace service.”
Catalog summary
2 nodes · path 196
ASST[562] “Let me create a thorough summary of the conversation so far, …”
Catalog after compaction
292 nodes · path 293
USR[563] “The summary below covers the earlier portion of the conversation.”
Frames
140 nodes · path 141
Flight
122 nodes · path 123
FDW attempt 1 (interrupted)
52 nodes · path 53
USR[861] “Build these FOUR packages.”
FDW attempt 2
201 nodes · path 202
USR[1146] “Build these FOUR packages from scratch (none exist yet).” ASST[1346] Write: "file_path": "/workspace/tests/test_pipelines.py" Figure 12: Prefix graph of execution f3ad38e5 with real message excerpts. Rectangles represent collapsed context segments; right-angle edges encode prefix sharing. Each parent is top-aligned with its first child, and other children are stacked below. Node counts refer to each segment; path counts the complete root-to-leaf context. Dashed rectangles identify summary leaves. Excerpts carry their original roles and node IDs; ellipses mark omissions. The interruption at TOOL[1036] remains in the main context where it was received.
24