Conceptio › Archive › arXiv CS
arXiv CSopen access

Conduit: An Experience Data Plane for Distributed Reinforcement Learning

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
clouddistributed-computingparallel-computing
distributed computing, parallel computing, cloud

arXiv:2609.24456v1 [cs.DC] 21 Sep 2026

Conduit: An Experience Data Plane for Distributed Reinforcement Learning Sitong Zhang

Tuo Shi

Mario Di Francesco

Aalto University [email protected]

Shenzhen University of Advanced Technology [email protected]

Aalto University [email protected]

Zeke Wang

Bo Zhao

Zhejiang University [email protected]

Aalto University [email protected] Experience-path handling

Abstract

20.3%

Distributed reinforcement learning (RL) scales training by parallelizing actors and learners around an Experience Buffer. As RL workloads grow, however, the buffer becomes more than a replay queue: it is the storage substrate of a largecapacity, latency-critical experience path that every iteration traverses to move, transform, sample, and batch experiences before learner updates can begin. Existing RL systems embed this path inside framework control flow or expose it as a request-driven buffer service, leaving experience placement fixed and experience-path work difficult to schedule independently as a runtime-level optimization target. We present Conduit, a framework-agnostic runtime that exposes RL experience management as an explicit systems optimization problem. At its core is the Experience Data Plane (EDP), a runtime abstraction that separates RL experiencehandling semantics from framework-specific execution logic by exposing experience ingestion, experience placement, and experience delivery as explicit control points. Built on EDP, Conduit introduces capacity-constrained, bandwidth-aware placement, which distributes experience state across CPU/GPU memory tiers and nodes under heterogeneous interconnect and device-memory constraints, and latency-aware scheduling, which controls when experience-path handling runs to reduce exposed experience-path latency while preserving RL semantics. Integrated with RLlib without changing its framework execution logic, Conduit reduces exposed experiencepath latency by up to 97% and end-to-end iteration latency by up to 38%, scales to 1,024 GPUs, and preserves convergence.

1

0

50.2% 10

20

30.5% 26.2% 0

100

Actor rollout

29.5%

Learner update

Experience-path handling dominates latency. 40 Off-policy (DQN) in RLlib with 8 actors & 8 learners.

30

43.3%

Experience-path handling is a major latency contributor. 400 On-policy (PPO) in RLlib with 8 actors & 8 learners.

200

Fig. 1: Experience-path latency is a major component of iteration time in distributed RL: 50.2% in off-policy Deep Q-Network (DQN) and 26.2% in on-policy proximal policy optimization (PPO) across representative workloads on an 8ˆA100 cluster.

Experience-path handling includes two types of operations: Data transfer: data movement in actor→buffer→learner path (e.g., CPU-GPU, memcpy,...) Data processing: data batch computation (e.g., concatenation, reshaping,...)

...

Actor Rollout

raw rollout data Experience Ingestion

Experience Buffer Experience Placement

...

training batches

Learner Update

Experience Delivery

Experience Data Plane DQN in RLlib (config: 8 actors, 8 learners): 33 operations =

×8 +

×25 per-iteration

PPO in RLlib (config: 8 actors, 8 learners): 27 operations =

×7 +

×20 per-iteration

Fig. 2: Experience-path handling drives latency through transfer and processing operations. Existing systems embed this handling in their actor/learner workflow; Conduit instead abstracts it into an explicit Experience Data Plane.

the storage substrate of a large-capacity, latency-critical experience path. Modern experiences carry large visual, multimodal, or long-context payloads, and each iteration repeatedly moves, transforms, samples, and batches these heavy payloads before learner update can begin. Despite extensive effort to optimize actors and learners, distributed RL systems still overlook experience-path latency: the repeated work that converts raw rollouts into trainingready batches, i.e., experience-path handling. Fig. 1 shows that this latency accounts for 50.2% of iteration time in off-policy DQN [27] and 26.2% in on-policy PPO [35]. Fig. 2 shows where this latency comes from: experience-path handling runs dozens of transfer and processing operations every iteration (33 for DQN, 27 for PPO in RLlib).

Introduction

Distributed reinforcement learning (RL) increasingly operates at scale, with large observations and heavy per-sample payloads [3, 4, 9, 17–19, 37, 39, 47, 48, 51]. In a typical distributed RL system, actors generate experiences, learners update models, and an Experience Buffer bridges the two [23, 25, 29, 53]. As workloads grow, however, this buffer becomes 1

The experience path is also capacity-constrained: large buffers often exceed a single GPU and must be distributed across CPUs, GPUs, and nodes on heterogeneous fabrics. Existing systems, however, treat capacity management, placement, and scheduling as local framework tactics rather than as a runtime problem. Experience buffers built into RL frameworks such as RLlib [23], MSRL [53], and SRL [25] embed handling within actor/learner execution at fixed locations (see Fig. 2 for the RLlib case); service-based systems such as Reverb [2] and Gear [44] are driven by incoming requests rather than proactively scheduling critical-path work to reduce end-to-end latency. Both designs treat the experience buffer as an internal framework component or service endpoint, rather than exposing it as a runtime interface for optimization. Our key insight is that distributed RL needs precisely this missing abstraction: an explicit runtime interface for managing and optimizing experience data, which we call the Experience Data Plane (EDP). RL frameworks retain control over which experiences to store and consume, while EDP manages the systems-level decisions of where those experiences reside and when ingestion and delivery occur. We introduce Conduit to realize EDP as a frameworkagnostic experience-buffer runtime that exposes experience ingestion, placement, and delivery as explicit system-level control points (Fig. 2) while preserving framework execution logic. Built on EDP, Conduit introduces two mechanisms: capacity-constrained, bandwidth-aware placement, which distributes the Experience Buffer across CPU and GPU memory under heterogeneous interconnect bandwidths and devicememory limits; and latency-aware scheduling, which determines when ingestion and delivery run. Conduit makes the following contributions: (1) Experience Data Plane as a runtime abstraction. We introduce the Experience Data Plane (EDP), a runtime abstraction that exposes experience ingestion, experience placement, and experience delivery as explicit control points rather than framework-internal side effects. EDP decouples RL experience semantics from framework-specific execution logic, creating an optimization surface for data residency, capacity management, and scheduling in distributed RL. (§4) (2) Capacity-constrained, bandwidth-aware placement. We design a placement mechanism that distributes the Experience Buffer across CPU/GPU memory tiers and nodes by jointly accounting for actor Ñ buffer Ñ learner path costs, heterogeneous interconnect bandwidths, and device-memory constraints. This allows Conduit to support buffers larger than a single GPU while reducing data-movement cost on heterogeneous fabrics. (§5) (3) Latency-aware scheduling of experience-path handling. We design a scheduling mechanism that decides when

experience ingestion and delivery run, and at what granularity, to reduce exposed experience-path latency while preserving on-policy freshness and off-policy replay semantics. This allows Conduit to overlap experience-path handling with rollout and learner update when semantics permit, while amortizing overhead through batching and pipelining. (§6) We implement Conduit and integrate it with RLlib [23]—a state-of-the-art reinforcement-learning framework—without changing its execution logic. We further show that Conduit generalizes beyond RLlib: it integrates into SRL [25], a distributed RL framework with a different, streaming-dataflow execution model (Appendix B), and extends to large language model (LLM) post-training with the Verl [37] framework (Appendix A). Across on-policy and off-policy RL workloads, Conduit reduces both exposed experience-path latency (up to 97%) and end-to-end iteration latency (up to 38%), and scales efficiently to 1,024 GPUs. These results establish EDP as a practical optimization surface for distributed RL: a runtime abstraction that decouples capacity, placement, and scheduling decisions from framework execution logic.

2

Background and Motivation

This section motivates treating experience-path handling as an independently optimizable runtime component on the actor Ñ Experience Buffer Ñ learner path. §2.1 frames distributed RL through this path, §2.2 summarizes its responsibilities, §2.3 states the three challenges that motivate Conduit, and §2.4 reviews existing approaches. 2.1

Actor Ñ Experience Buffer Ñ Learner Data Path

In RL, an agent learns a policy (typically a deep neural network, DNN) to act in an environment [38]. Each training iteration executes the current policy to generate rollouts (experience tuples such as state, action, reward, and next state) and then performs learner update to train the policy DNN on those experiences. In practice, these tuples are rarely flat records: modern experiences increasingly contain large visual, multimodal, or long-context payloads augmented with termination flags and auxiliary metadata (e.g., policy_info, episode_ID) [41, 42]. To scale training, modern RL systems parallelize actor rollout and learner update [24] by running many actors and learners, bridged by an Experience Buffer that ingests rollouts and serves training-ready batches (Fig. 3). The actor Ñ Experience Buffer Ñ learner connection forms a large-capacity, latency-critical experience path. Every iteration must repeatedly move, transform, sample, and batch these heavy payloads into training-ready data under strict algorithmic constraints (e.g., freshness requirements in onpolicy RL and reuse via replay in off-policy RL). To understand how capacity demands and execution latency along 2

Tab. 1: Comparison of existing experience buffer implementations for distributed RL. Category Built-in Experience Buffers

System RLlib [23] MSRL [53]

Decoupling Embedded module

SRL [25]

Placement Fixed (CPU) Fixed (CPU/GPU)

Scheduling

Fixed (pinned CPU)

Reactive (fixed one-step overlap)

Reactive (inline in actor/learner loops)

Experience Buffer Services

Reverb [2] Gear [44]

Process-decoupled service

Fixed (CPU) Fixed (pinned CPU & GPU-driven access)

Reactive (request-driven)

Experience Management Layer

Conduit

Multi-dimensional Decoupling (system, algorithm, hardware)

Adaptive (capacity-constrained, bandwidth-aware)

Proactive (schedulable execution timing & processing granularity)

Actors Policy Model

action state

Environment

Experience Buffer Experience Ingestion

Experience Delivery

Learners Value Model

Experience placement controls where and how experience state resides across distributed resources (e.g., CPU memory, pinned host memory, GPU memory, or distributed across multiple nodes). As a logical knob, placement does not change the RL semantics of what experience ingestion or delivery must accomplish, but it dictates how capacity is distributed across devices. It determines where data-plane operators read/write records and which transfer edges are exercised. Because every experience must traverse this path, placement directly governs whether large datasets can fit in available memory, as well as the cost of cross-device data movement on heterogeneous fabrics. (ii) Physical execution realizes experience ingestion and delivery as operator graphs with two operator types: data transfer moves rollouts/batches across devices, memory domains, and nodes, while data processing transforms them into learner-consumable layouts. Transfer operators include CPUØGPU copies, GPUØGPU transfers, inter-node send/recv, and shard movement via collectives [5, 43]. Processing operators include trajectory assembly, concat/ stack, reshape/ slice, sampling, shuffling, and batching, optionally with padding/ packing for variable-length sequences [38]. Thus, ingestion and delivery define the logical phases, transfer and processing operators define physical execution, and experience placement determines capacity and locality, shaping the dominant transfer costs on the actor Ñ Experience Buffer Ñ learner path in Fig. 3.

Policy Model

sync weights actor rollout

experiences

learner update

Fig. 3: Distributed deep reinforcement learning architecture.

this path shape end-to-end training efficiency, the next subsection characterizes its logical responsibilities and connects them to the underlying operator graph. 2.2

Experience Path Responsibilities

Experience-path handling spans two complementary levels: (i) a logical view that defines what the path provides (ingestion, delivery, and placement), and (ii) a physical execution plan that defines how this logical view is realized as operator graphs (transfer and processing operators). (i) Logical view comprises two functional operations, experience ingestion and experience delivery, and one physical control point, experience placement. Experience ingestion collects actor rollouts into bufferresident experiences by running data-plane operators that materialize and organize incoming records. In on-policy training, ingestion assembles step-level tuples into trajectories and applies preprocessing [30, 35] (e.g., boundary detection, padding/truncation, normalization, and computing returns/advantages) before forming batches for the current iteration. In off-policy training, it inserts transitions into a persistent replay dataset and maintains indices and metadata for future sampling and eviction [15, 34]. Experience delivery extracts experiences from the buffer and produces training-ready batches for the learner by running operators that select, assemble, and transform records under algorithmic constraints [15, 34]. In on-policy training, delivery consumes freshly ingested experiences within the same iteration by partitioning trajectories into mini-batches and applying lightweight batch transformations (e.g., shuffling and padding [35]). In off-policy training, it samples from a persistent replay dataset, materializes the selected transitions, and performs batch/layout transformations (e.g., concatenation and reshaping) required by the learner [28].

2.3

Experience-Path Challenges and Opportunities

Existing RL frameworks face three challenges in managing the experience path. Each challenge motivates the design of the Experience Data Plane (EDP), which turns it into an optimization opportunity. Challenge 1: Coupled execution serializes handling onto the critical path. Most distributed RL frameworks embed ingestion and delivery inside actor/learner execution: ingestion is tied to actor rollout, and delivery is tied to learner update. Although dependencies only constrain operators within a batch, hard-wiring these operator graphs into fixed execution loops turns them into global barriers that block pipelining across batches (e.g., ingesting D𝑖 while generating D𝑖 `1 , or delivering D𝑖 `1 while updating on D𝑖 ; Fig. 5), forcing handling latency onto the critical path. 3

GPU7 100GB/s GPU5 400GB/s GPU6 GPU4 200GB/s GPU0 GPU2 GPU1

GPU3

(a) Non-uniform GPU fabric.

Destination rank

CPU Cores 0 - 64 72GB/s

7

102 107 146 147 144 147 40 33

6

77 105 144 147 146 146 22

5

142 147 105 107 40 33 144 146

4

142 146 74 107 3

3

142 146 51 3 106 107 146 146

2

145 146 2

1

52 3 144 146 146 147 105 107

0

2

53 145 146 141 146 80 107

0

1

150

(a) SOTA: Sequential execution Actor D1 rollout

41

2 3 4 5 Source rank

6

D3

D4

...

Learner update

D4

Delivery D1

D1

D2

D3

D4

100

Ingestion

42 146 146

54 81 106 144 146

D2 D1

D2

D3

D2

D3

D4

...

(b) Overlapped execution w/o breaking data dependency 50

Actor D1 D2 D3 D4 rollout Ingestion

0

7

Learner update

... latency reduction

D1 D2 D3 D4 latency reduction

Delivery D1 D2 D3 D4

D1 D2 D3 D4

...

(c) Batched exectution w/o breaking data dependency

(b) GPUØGPU latency heatmap.

Actor rollout D1 D2 D3 D4

Fig. 4: Non-uniform intra-node transfers on an 8-GPU AMD MI250X node. (a) interconnect topology; (b) GPUØGPU latency variation for a 5 GB transfer, making placement performance-critical.

Ingestion

D1,2

Learner update

... latency reduction D3,4

data dependent

D1 D2 D3 D4 latency reduction

D1 (D3) and D2 (D4) are batched

Delivery D1,2 D3,4

...

data independent

Fig. 5: Operator dependencies (D1–D4 are batches): (a) sequential, (b) overlapped via timing, (c) overlapped and batched via joint timing and granularity.

Opportunity: Decoupling. By decoupling ingestion and delivery from framework execution, the EDP turns them into first-class objects that the system controls directly, granting three forms of autonomy: execution autonomy, to run them asynchronously and pipeline across batches when dependencies allow; semantic autonomy, to serve on-policy freshness and off-policy replay under one interface; and hardware autonomy, to detach experience placement from fixed framework bindings (e.g., CPU or a specific GPU). Challenge 2: Fixed placement either overflows memory or pays for slow fabric transfers. Where experience data resides determines both capacity and transfer cost, and the two pull against each other. Capacity binds at scale: in visual RL with 128ˆ128–256ˆ256 RGB inputs [11, 13, 32], a 106 -sample buffer reaches 48–192 GB, exceeding a single GPU. But spilling elsewhere can be costly. Fig. 4 illustrates this non-uniformity within a single node1 : GPU pairs differ in bandwidth (100–400 GB/s) and CPUØGPU is slower (72 GB/s) (Fig. 4a), so a 5 GB GPUØGPU transfer spans 40– 147 ms, a 3.7ˆ spread (Fig. 4b). A fixed choice (always CPU or a fixed GPU) thus either overflows memory or repeatedly pays for slow transfers. Opportunity: Placement. The EDP’s hardware autonomy lets Conduit treat placement as a runtime decision rather than a fixed binding. Through capacity-constrained, bandwidthaware placement, Conduit distributes the buffer across CPU, GPU, and node tiers to fit memory limits while steering data onto fast edges—scaling capacity beyond a single device and avoiding the slowest fabric edges at the same time. Challenge 3: Synchronous handling fully exposes its latency and wastes per-invocation overhead. Experience handling is far from free: ingestion and delivery perform nontrivial processing (trajectory assembly, batching, sampling, reshaping, shuffling, padding). Run synchronously between rollout and update, this work lands entirely on the critical path, where it becomes a major component of iteration time (Fig. 1); executed one unit at a time, its fixed per-invocation overhead is also paid repeatedly.

Ingestion

D1 Batch size 32

D1,2 64

D1-4 128

Delivery

D1-8 256

D1-16 512

Fig. 6: Larger operator batches reduce per-transition processing cost for both ingestion and delivery.

Opportunity: Scheduling. The EDP’s execution autonomy lets Conduit schedule experience-path handling instead of running it inline. Because dependencies bind operators only within a batch (Fig. 5), two knobs reduce cost. Timing overlaps ingestion for D𝑖 with rollout for D𝑖 `1 , and delivery for D𝑖 `1 with the update on D𝑖 (Fig. 5b); granularity batches trajectories per invocation to amortize overhead (Fig. 5c). Both help: on RLlib [23], increasing the operator batch from 32 to 512 transitions cuts amortized ingestion from 0.20 to 0.085 ms/transition (2.4ˆ) and delivery from 0.38 to 0.21 ms (1.8ˆ) (Fig. 6). Summary. Together, addressing these three challenges yields Conduit’s optimization opportunities—decoupling, placement, and scheduling—which it targets end-to-end on the actor Ñ Experience Buffer Ñ learner data path. 2.4

Limitations of Existing RL Experience Buffers

Existing RL buffer designs fall into two structural categories: framework-internal modules and standalone service endpoints (Tab. 1). In both cases, experience handling remains reactive, placement-fixed, and tightly bound to framework execution, preventing systematic optimization across decoupling, placement, and scheduling. Built-in experience buffers (embedded and reactive). RL frameworks such as RLlib [23] and SRL [25] implement the experience buffer as an internal framework component, so experience ingestion and delivery execute inline with actor rollout and learner update. This hard-wires experience handling into framework-specific execution logic, preventing principled pipelining across batches (see Fig. 5). Built-in

1We profile these fabrics on an AMD MI250X supercomputer

4

Experience Data Plane (§4)

experience buffers are also placement-fixed (CPU or pinned host memory), failing to manage capacity or optimize transfer costs across distributed GPU resources. MSRL [53] improves task placement across actors and learners via a fragmented dataflow graph, which is complementary. However, it does not expose an explicit runtime surface for the experience path itself, leaving experience handling reactive and tied to task execution. Experience buffer services (throughput-scaled, but not latency-optimized). Reverb [2] and Gear [44] decouple the experience buffer at the process level. However, they treat the buffer as a service endpoint rather than as a runtime control surface. They retain fixed placement (CPU resident) and request-driven execution. Reverb stores experience data in pageable CPU memory. Gear stores it in pinned CPU memory while using GPUs to leverage DMA and RDMA reads, rather than dynamically placing experience data across GPU boundaries. By ignoring the joint optimization of capacityconstrained placement and proactive scheduling, these systems focus on insertion/sampling throughput but do not directly optimize exposed experience-path latency. Summary. Prior work treats the experience buffer as an internal component or external service rather than as an explicit Experience Data Plane, motivating Conduit.

3

Actor

Experience Ingestion

Experience Buffer

Experience Placement

...

...

Experience Delivery

Capacity-Constrained, Bandwidth-aware Placement (§5)

Latency-Aware Scheduling (§6) Processing granularity

...

D1,D2 D1 D2 D3 D4

batched

... D3,D4 ... ... D1 D2 D3 D4

D1,D2

D3,D4

... ...

overlapped

Learner

Execution timing

Fig. 7: Conduit workflow: the Experience Data Plane exposes three control points—experience ingestion, placement, and delivery—around the Experience Buffer.

@conduit.register_ingestion(input="Actor", output="ExperienceBuffer") def ingestion(actor_in, exp_out): d = actor_in.read() # read from Actor d = fw_ingestion_ops(d) # recall framework-defined ingestion logic exp_out.write(d) # write to Experience Buffer @conduit.register_delivery(input="ExperienceBuffer",output="Learner") def delivery(exp_in, learn_out): d = exp_in.read() # read from Experience Buffer d = fw_delivery_ops(d) # recall framework-defined delivery logic learn_out.write(d) # write to Learner

Register framework-defined algorithm-specific ingestion/delivery logic (fw_*_ops)

Fig. 8: Conduit integration interface: frameworks register ingestion/delivery logic via buffer hooks; Conduit controls placement and scheduling.

Overview of Conduit

Conduit is a runtime system that realizes the Experience Data Plane (EDP) for the large-capacity, latency-critical experience path. It separates RL experience handling from framework-specific execution by exposing experience ingestion, experience placement, and experience delivery as runtime control points. Conduit integrates with existing distributed RL frameworks by intercepting buffer interfaces; on-policy freshness and off-policy replay semantics bound feasible scheduling and placement, so Conduit optimizes capacity and latency within these bounds. As shown in Fig. 7, the EDP exposes three control points around the persistent Experience Buffer—experience ingestion, placement, and delivery—refining the rigid actor Ñ Experience Buffer Ñ learner execution into a fine-grained, independently controllable experience path. Actors push raw rollout payloads asynchronously ( 1 ); Conduit ingests them into the Experience Buffer ( 2 ) and delivers training-ready batches ( 3 ); learners then pull these staged batches for model updates ( 4 ). Because ingestion and delivery run as operators decoupled from actor and learner execution, they become pipelineable constraints: Conduit can overlap ingestion with actor-environment interaction and delivery with policy DNN training, reducing exposed experience-path latency. Conduit integrates with existing RL frameworks through a low-intrusion interception layer (Fig. 8). Frameworks register their algorithm-specific ingestion and delivery logic via standard buffer hooks, while Conduit owns execution

timing, granularity, and capacity distribution. This separation lets frameworks keep defining which experiences are stored and consumed, while Conduit controls where they reside and when experience-path handling executes, leaving framework logic unchanged. Conduit addresses the capacity demands and data movement costs of large-scale RL through capacity-constrained, bandwidth-aware placement. Using measured interconnect bandwidths and device-memory limits, Conduit systematically distributes the Experience Buffer across CPU and GPU memory tiers. This enables Conduit to host experience datasets that exceed any single GPU’s memory by distributing them across devices, while actively avoiding the slowest transfer edges on heterogeneous, non-uniform fabrics (§5). Finally, Conduit implements latency-aware scheduling for experience-path handling. It dynamically determines when experience ingestion and delivery execute (execution timing) and how much data they process per invocation (execution granularity). This cost-based scheduling lets Conduit overlap experience-path handling with actor rollout and learner update whenever semantics allow, amortizing overheads through batching and minimizing exposed experience-path latency on the critical path (§6). 5

... Actor() Ingestion() ... Actor() Ingestion() (a) Coupled Ingestion: Sequential execution (off/on-policy)

Actor()

Actor() ... Actor()

Actor(ging)

Ingestion()

Delivery()

... Actor(g ) ing

Delivery(gdel)

...

... ...

Ingestion(ging)

(b) Decoupled Experience Ingestion enables overlap in off-policy (ging ≥ 1)

Ingestion(ging) ...

... ... Learner() Delivery() Learner() Delivery() (a) Coupled Delivery: Sequential execution (off/on-policy) ...

Delivery(gdel)

... Learner() ... Learner()

Ingestion(ging)

(b) Decoupled Experience Delivery enables overlap in off-policy (gdel ≥ 1)

(c) Decoupled Experience Ingestion enables overlap in on-policy (0 < ging <1)

Learner()

... Delivery(gdel) ... Learner(gdel)

...

Learner(gdel)

(c) Decoupled Experience Delivery enables overlap in on-policy (0 < gdel < 1)

Fig. 9: Decoupling ingestion (delivery) from actor (learner) execution. Actors push rollouts asynchronously, while the ingestion operator pulls data at configurable granularity, enabling batching and pipelining (left). The delivery operator prepares trainingready batches ahead of time, freeing learners from synchronous data-preparation waits (right).

4

Experience Data Plane

into framework-specific actor logic. Off-policy training tolerates mild staleness, so ingestion may batch rollouts with max u, where 𝑔max caps staleness and memory 𝑔ing P t1, . . . , 𝑔ing ing overhead. On-policy training requires fresh samples: if a rollout unit feeds 𝑀 mini-batch updates, ingestion can pipeline only at mini-batch granularity with 𝑔ing P t 𝑀1 , 𝑀2 , . . . , 1u.

We formalize Conduit’s decoupling through the Experience Data Plane (EDP), which breaks the tightly coupled actor Ñ Experience Buffer Ñ learner dependency by turning experience ingestion, placement, and delivery into first-class asynchronous runtime operators rather than inline steps in framework execution loops. This section focuses on the timing-related controls: ingestion, which governs how rollouts enter the buffer (§4.1), and delivery, which governs how training-ready batches leave (§4.2); placement (where data resides) is deferred to §5. This decoupling underpins Conduit’s capacity-constrained, bandwidth-aware placement (§5) and latency-aware scheduling (§6). 4.1

4.2

Decoupling Experience Delivery

Concept. Experience delivery separates experience-path handling from learner execution. As an explicit runtime operator, it lets EDP prepare training-ready batches ahead of time while learners asynchronously pull staged batches. Users specify only standard training configuration, and Conduit derives valid control choices without manual tuning. We define delivery granularity 𝑔del as the amount of training data prepared per delivery invocation, measured in multiples or fractions of one training batch. Enabling overlap and pipelining. Without decoupled delivery, each learner update waits for the next batch, leaving little scheduling freedom. With EDP boundaries, delivery runs proactively and stages batches, while 𝑔del sets the work per invocation: when 𝑔del ě 1, delivery can prefetch multiple batches, overlapping learner updates with delivery of upcoming batches; when 0 ă 𝑔del ă 1, delivery feeds minibatch chunks, enabling fine-grained intra-iteration pipelining (right side of Fig. 9). Unifying off-policy and on-policy via 𝑔del . As with ingestion, the feasible range of 𝑔del captures semantic constraints at the EDP boundary. For off-policy training, delivery may max u, where prefetch multiple batches with 𝑔del P t1, 2, . . . , 𝑔del max 𝑔del bounds staleness and memory overhead. For on-policy training, delivery preserves freshness within the current iteration and provides mini-batches sequentially, so 𝑔del “ 𝑀1 .

Decoupling Experience Ingestion

Concept. Experience ingestion is the control point that separates actor execution from experience-path handling. By making ingestion an explicit runtime operator, EDP lets actors push raw payloads asynchronously while ingestion pulls and processes data at its own pace. Users only provide standard RL settings (e.g., rollout configuration and whether training is on/off-policy), and Conduit derives valid experience-path control choices automatically. We capture this knob using an ingestion granularity 𝑔ing , defined as how much rollout data one ingestion invocation processes, measured in multiples (or fractions) of one rollout unit.2 Enabling overlap and pipelining. Without decoupled ingestion, each rollout unit must be ingested immediately after it is generated, creating a strict one-to-one dependency between actor execution and experience handling. With decoupled ingestion, actors push asynchronously, and 𝑔ing determines how ingestion runs: (i) when 𝑔ing ě 1, ingestion can batch multiple rollout units and process them together; (ii) when 0 ă𝑔ing ă 1, ingestion can pipeline by processing a fraction of the current rollout while actor–environment interaction continues. This turns tight per-rollout dependencies into pipelineable constraints, enabling overlap without changing RL semantics (left side of Fig. 9). Unifying off-policy and on-policy via 𝑔ing . The feasible range of 𝑔ing encodes algorithm semantics (freshness vs. reuse) in a runtime boundary rather than hard-coding them

4.3

Configuration and Correctness

Users do not tune buffer granularities directly. Instead, Conduit derives feasible settings from RL configurations that frameworks already expose: whether training is off-policy or on-policy, and (for on-policy) the mini-batch count 𝑀 per update. These inputs define safe batching and mini-batch pipelining ranges that preserve replay and freshness semantics. Within this safe space, Conduit selects execution mode

2 A rollout unit is the RL algorithm-defined volume generated by one envi-

ronment interaction phase, e.g., a rollout batch. 6

and effective granularity automatically; users may optionally cap batching to limit staleness or memory overhead, but sensible defaults typically suffice. These same bounds also preserve correctness: EDP changes when and where experience-path handling runs, not what data it produces, and overlap/granularity choices are drawn only from spaces that preserve on-policy freshness and off-policy replay semantics. Formal invariants are summarized in Appendix C. Overall, EDP turns framework execution loops into pipelineable constraints, making experience-path operations independently schedulable while preserving RL semantics.

copy) and remotely with probability 7{8 (GPU Ñ GPU peerto-peer), estimated from buffer sharding and sampling behavior, the expected bandwidth 𝐵𝑊 p𝑝 Ñ GPUq reflects this mixture and yields an accurate latency estimate. 5.2

𝑝 ˚ “ arg min 𝑇transfer p𝑝q

5 Capacity-Constrained, Bandwidth-Aware Experience Data Placement

𝑝PP

s.t. 𝐶 buf ¨ 𝑥 ď 𝑀𝑒𝑚p𝑝q. (1)

Configuration transparency. Users do not tune placement directly. Conduit derives actor/learner device types and 𝐶 buf from the RL configuration, profiles 𝑀𝑒𝑚p¨q and effective 𝐵𝑊 p¨q at initialization, and solves Eq. (1) automatically. Users may optionally restrict candidates, but defaults typically suffice. Although Eq. (1) defines the target optimum, exhaustive search is impractical: effective bandwidths vary with contention and concurrency, and device-subset candidates can grow as Op2𝑛 q for 𝑛 devices. Conduit therefore uses an approximate, hardware-aware strategy (Alg. 1) that searches a small candidate set. Approximate solution. Conduit chooses placement at startup and during occasional reconfiguration, off the training critical path. It constructs P using two heuristics. (i) Cross-node balance: in multi-node runs, Conduit distributes capacity across nodes to avoid inter-node hotspots, since inter-node links are slower than intra-node GPUØGPU fabrics. (ii) Intra-node representatives: within a node, Conduit considers CPU pinned placement and shared GPU placements over GPU subsets of increasing size, preferring high-bandwidth ranks on non-uniform fabrics. With 𝐺 ď 8 GPUs per node, enumerating these subsets remains efficient. Given this reduced P, Conduit scans candidates (Alg. 1, lines 2–5) and selects the feasible placement with the smallest 𝑇transfer p𝑝q. Complexity. Each 𝑇transfer p𝑝q evaluation is constant-time, so selection costs Op|P|q. In practice, |P| is small and bounded (e.g., at most 2𝐺 GPU subsets per node with 𝐺 ď 8), and placement runs outside the critical path.

Building on the Experience Data Plane (§4), Conduit treats experience placement as an explicit optimization that reduces actor Ñ Experience Buffer Ñ learner overhead. For each candidate placement, it (i) estimates transfer latency from measured effective bandwidths (§5.1) and (ii) selects the lowest-latency placement that remains feasible under memory constraints (§5.2). Since experience sizes can grow midrun (e.g., richer observations or longer sequences), an initial placement may become infeasible; Conduit then performs (iii) online migration driven by a pre-computed placement map, avoiding runtime re-profiling (§5.3). 5.1

Placement Optimization

Given the transfer model above, Conduit selects the placement that minimizes transfer latency subject to buffer capacity. Let 𝐶 buf be the required buffer capacity, 𝑥 the per-sample size, and 𝑀𝑒𝑚p𝑝q the profiled memory available under placement 𝑝. Among candidate placements P, Conduit solves:

Transfer-Latency Model

An Experience Buffer placement 𝑝 induces two critical-path transfers: (i) moving new experiences from actors into the buffer (experience ingestion), and (ii) moving sampled training data from the buffer to learner GPUs (experience delivery). We model per-sample transfer latency as the sum 𝑥 𝑥 of these two legs: 𝑇transfer p𝑝q “ 𝐵𝑊 pactor ` 𝐵𝑊 p𝑝 Ñ Ñ𝑝 q GPUq , where 𝑥 is per-sample size and 𝐵𝑊 p¨q is the effective bandwidth of the path. Effective bandwidth is obtained from microbenchmarks on the deployment stack and captures realized throughput under the hardware topology and runtime implementation [21, 22, 45]. To account for locality (intra- vs. inter-node and GPU-rank dependence), 𝐵𝑊 p¨q is looked up from measured bandwidth tables indexed by endpoint type and topology. We consider four placement options: CPU pageable, CPU pinned, single GPU, and shared GPU. These cover the two factors that govern transfer cost on heterogeneous fabrics: memory tier (pageable/pinned CPU vs. GPU) and device distribution (single device vs. sharded across multiple devices). Together with the actor-side device type, each placement uniquely determines both legs, making 𝑇transfer p𝑝q directly computable. Example. Consider a node with 8 GPUs wherein actors run on CPU, the Experience Buffer is sharded across GPUs, and learners run on GPUs. Then actor Ñ 𝑝 uses CPU Ñ GPU bandwidth, while 𝑝 Ñ GPU uses GPU Ñ GPU bandwidth. If a learner samples locally with probability 1{8 (device-local

5.3

Online Placement Migration

The placement optimizer selects the initial placement, but in some RL workloads the per-sample size 𝑥 grows during training as observations, sequences, or auxiliary fields expand [37, 52]. When growth makes the current placement infeasible, Conduit switches via a pre-computed placement map to the next feasible low-latency placement, avoiding full online re-optimization. At runtime, migration briefly pauses ingestion and delivery, transfers experience state using a pre-solved plan, and 7

Actor Rollout

Ingestion

Learner Update

Delivery overlap sync latency

Strategies

Scenarios

multi-granularity latency

off-policy

(a)

(b)

(c)

(d)

the sync overhead outweighs the latency reduction from overlap

overlaping = 0, ging ≥1 no overlap + multi-granularity D1

D1,2

D2

D3

D4

D1,2

the sync overhead outweighs the latency reduction from overlap the latency reduction from overlap outweighs the sync overhead

D1,2 D1

D2

D2

D3

D3,4

D4

Tdel (gdel)

overlapdel = 1, gdel ≥1 overlap + multi-granularity

D3,4 D4

D3

D1

Tupdate

Trollout Ting(ging) overlaping = 1, gingest ≥1 overlap + multi-granularity

the latency reduction from overlap outweighs the sync overhead

overlapdel = 0, gdel ≥1 no overlap + multi-granularity

D3,4

D1,2 D1

Tsync,ing

D3,4 D2

D3

Tsync,del

D4

on-policy overlapdel = 0, gdel =1 no overlap + single-granularity

overlaping = 0, ging =1 no overlap + single-granularity D1

D1

D1

overlaping = 1, 0<ging <1 overlap + sub-granularity D11

D11

D12

D13

D12

D13

D14

overlapdel = 1, 0<gdel<1 overlap + sub-granularity D11

D14

D1

D12

D13

D14

D11

D12

D13

D14

(overlaping = 0, 0<ging <1) ,(overlaping = 1, ging =1), (overlapdel = 0, 0<gdel<1) and (overlapdel = 1, gdel =1 ) are not suitable for on-policy

Fig. 10: Ingestion /delivery-path modes: overlap ˆ granularity across iterations (D1–D4) and mini-batches (D11 –D14 ).

Two semantic-free controls. Conduit schedules both operators with two knobs: (i) Timing, controlled by 𝑜𝑣𝑒𝑟𝑙𝑎𝑝 ing and 𝑜𝑣𝑒𝑟𝑙𝑎𝑝 del P t0, 1u, decides whether ingestion/delivery overlaps with rollout/update; and (ii) Granularity, controlled by 𝑔ing and 𝑔del (§4), decides how much work each invocation performs. Feasible granularities are constrained by Ging and Gdel derived from standard RL configuration (§6.3), keeping scheduling algorithm-agnostic. Fig. 10 instantiates these knobs as execution modes. Nonoverlapped timing corresponds to (a) and (c), while overlap corresponds to (b) and (d). Along granularity, 𝑔 ą 1 batches units in (a) and (b), 𝑔 “ 1 runs single-unit in (c), and 0 ă 𝑔 ă 1 pipelines mini-batches in (d). Together, these knobs cover the common RL settings—off-policy sequential/overlapped and on-policy sequential/pipelined—under a unified scheduling interface.

Algorithm 1: Bandwidth-aware placement (linear scan): select the placement with minimum transfer latency. Input: Actor-side device type, required buffer capacity 𝐶 buf , candidates P, profiled 𝐵𝑊 p¨q and 𝑀𝑒𝑚p¨q, per-sample size 𝑥. Output: Selected placement 𝑝 ˚ ˚ 1 𝑝 Ð None; 𝑇min Ð `8; 2 foreach 𝑝 P P do 3 if 𝐶 buf ¨ 𝑥 ď 𝑀𝑒𝑚p𝑝q then 4 𝑇 Ð 𝑇transfer p𝑝q; 5 if 𝑇 ă 𝑇min then 𝑇min Ð 𝑇 , 𝑝 ˚ Ð 𝑝; 6

return 𝑝 ˚ ;

resumes both operators under the new placement. Since the EDP decouples actor and learner execution, this adaptation does not stall either side of the training loop. Appendix D gives the migration algorithm and transfer-plan formulation.

6

Latency-Aware Scheduling of Experience Ingestion and Delivery

6.2

We model one RL iteration as 𝑇iter “ 𝑇ing-path ` 𝑇del-path , where 𝑇ing-path (resp. 𝑇del-path ) is the exposed latency of the ingestion (resp. delivery) path. Let 𝑇rollout and 𝑇update denote the execution time of one actor rollout and one learner update. Let 𝑇ing p𝑔q and 𝑇del p𝑔q denote the wall-clock time of one ingestion/delivery invocation that processes 𝑔 rollout units/training batches (including transfer and processing under placement 𝑝 ˚ from §5). Granularity and timing. Processing more data per invocation amortizes fixed per-invocation overheads, so we use per-unit costs 𝑇¯ing p𝑔q “ 𝑇ing p𝑔q{𝑔 and 𝑇¯del p𝑔q “ 𝑇del p𝑔q{𝑔. Overlap hides part of ingestion/delivery behind rollout/update, but adds synchronization overhead. Given a scheduling mode, the exposed ingestion-path time ` ¯ing p𝑔ing q ` 𝑜𝑣𝑒𝑟𝑙𝑎𝑝 ing ¨ 𝑇sync,ing ´ is 𝑇ing-path “ 𝑇 ` 𝑇 rollout ˘ ℎ𝑖𝑑𝑒 ing , and the exposed delivery-path time ˘is 𝑇del-path “ ` 𝑇update `𝑇¯del p𝑔del q`𝑜𝑣𝑒𝑟𝑙𝑎𝑝 del¨ 𝑇sync,del ´ℎ𝑖𝑑𝑒 del . The hidden terms ℎ𝑖𝑑𝑒 ing and ℎ𝑖𝑑𝑒 del capture the overlapable portion of

With placement fixed (§5), Conduit further reduces exposed experience-path latency by scheduling experience ingestion and delivery. Because the EDP (§4) makes these operators independently runnable, Conduit controls when they run (timing: sequential vs. overlapped) and how much each invocation processes (granularity: fractional vs. batched). We outline these opportunities (§6.1), model their latency impact (§6.2), and formulate a cost-based optimizer for the latencyminimizing schedule (§6.3). 6.1

Iteration-Latency Model

Scheduling Opportunities

Each RL iteration traverses two EDP paths: the ingestion path (actor rolloutÑexperience ingestion) and the delivery path (experience deliveryÑlearner update). The EDP makes ingestion and delivery schedulable rather than inline with rollout/update, enabling cross-iteration batching (off-policy) or intra-iteration pipelining (on-policy) to reduce latency. 8

operator latency under off-policy batching and on-policy pipelining; Appendix E details these cases. This compact model matches Fig. 10: larger 𝑔 lowers amortized cost, while overlap hides latency at the cost of synchronization.

total) and eight AMD MI250X GPUs (64 GB per GPU die), running SUSE Linux Enterprise Server 15 SP5. We use the A100 cluster for overall efficiency (§7.2), placement (§7.3), scheduling (§7.4), joint adaptive control (§7.5), and convergence (§7.7). We use the MI250X supercomputer for scalability experiments (§7.6) to demonstrate hardware generality and leverage its larger GPU pool (up to 1,024 GPUs). Baselines. We compare Conduit against the following three baselines: RLlib [23], an open-source RL framework whose experience buffer is embedded in actor/learner control flow with default CPU placement and request-driven sampling; Gear [44], an RL buffer service that stores trajectories in pinned host memory and uses GPU-driven zero-copy DMA/RDMA to accelerate transfers; and Reverb [2], a distributed buffer service that hosts buffers in CPU memory for scalable ingestion. We also compare against two other RL frameworks, Verl [37] (LLM post-training) and SRL [25], in Appendices A and B. RL algorithms. We evaluate three representative RL algorithms: Deep Q-Network (DQN) [27] and soft actor-critic (SAC) [12] as off-policy workloads, and proximal policy optimization (PPO) [35] as an on-policy workload. Appendix A additionally studies Group Relative Policy Optimization (GRPO) [36] in an LLM post-training pipeline. Environments. We evaluate five environments that span different kinds of experience-path pressure: MountainCar [8], a classic control task with small fixed-size observations; Multi-agent CartPole [40], a cooperative multi-agent task that multiplies per-step experience volume; Meta-World [50] („48 KB per sample), a 50-task robotic manipulation benchmark with visual observations; Open X-Embodiment [32] (OXE, „192 KB per sample), a large-scale multi-embodiment robotics dataset aggregating 60+ real-robot sources; and a configurable synthetic stress test following [2], whose transition dimension varies from 128 to 20,000 for the placement, scheduling, and adaptive-control studies. Metrics. We use end-to-end iteration latency—including actor rollout, experience-path handling, and learner update—as the primary metric. We also measure the exposed experiencepath latency, which captures the portion of ingestion and delivery work that is not hidden by overlap with rollout or learner update within each iteration.

6.3 Cost-Based Scheduling Conduit minimizes iteration latency by jointly choosing overlap and granularity for both operators: min 𝑇ing-path p𝑜𝑣𝑒𝑟𝑙𝑎𝑝 ing, 𝑔ing q ` 𝑇del-path p𝑜𝑣𝑒𝑟𝑙𝑎𝑝 del, 𝑔del q s.t. 𝑜𝑣𝑒𝑟𝑙𝑎𝑝 ing, 𝑜𝑣𝑒𝑟𝑙𝑎𝑝 del P t0, 1u, 𝑔ing P Ging, 𝑔del P Gdel .

(2)

Configuration transparency. Users do not tune overlap flags or granularities directly. Conduit derives Ging and Gdel from RL configuration—integer granularities for off-policy replay and fractional mini-batch steps for on-policy training— optionally respecting user caps. It then profiles 𝑇rollout , 𝑇update , and operator costs once, enumerates the resulting finite candidate set, and selects the lowest-latency mode under Eq. (2). §7.5 shows it automatically selects the best configuration. Complexity. Conduit solves Eq. (2) in constant time, Op1q. It enumerates all 4 ¨ |Ging | ¨ |Gdel | candidates (two binary overlap flags), scoring each in constant time from profiled costs; this count is fixed by the RL configuration and independent of workload scale, so the search runs once at initialization (and only re-runs if workload or resources change).

7

Evaluation

We evaluate Conduit along the three dimensions of the Experience Data Plane introduced in §2: (i) decoupling, (ii) capacity-constrained, bandwidth-aware placement, and (iii) latency-aware scheduling. After outlining the experimental setup (§7.1), the evaluation answers the following questions: (1) What is Conduit’s overall performance, and how does it reduce exposed experience-path latency in state-of-the-art distributed RL frameworks? (§7.2) (2) How well does Conduit’s capacity-constrained, bandwidthaware placement adapt to workloads under heterogeneous interconnect and device-memory limits? (§7.3) (3) How effectively does Conduit’s latency-aware scheduling reduce exposed experience-path latency? (§7.4) (4) How does Conduit’s auto-optimizer compare against tuned manual baselines? (§7.5) (5) What is the scalability performance of Conduit? (§7.6) (6) Does Conduit preserve RL convergence? (§7.7)

7.2 7.1

Experimental Setup

Overall Effectiveness and Efficiency

We first evaluate Conduit’s end-to-end benefit by integrating it into RLlib, a production-grade distributed RL framework: Conduit drop-in replaces RLlib’s built-in experience buffer without modifying actor/learner control flow. This directly exercises the decoupling dimension, validating that the Experience Data Plane is structurally independent of framework execution logic. We report both end-to-end iteration latency and exposed experience-path latency, and

Our experiments use the following setup. Testbeds. We conduct experiments on two clusters: (1) an NVIDIA A100 cluster—two nodes, each equipped with eight NVIDIA A100 80 GB GPUs, two Intel Xeon Platinum 8342 CPUs (48 cores in total), and 2 TB RAM, running Red Hat Enterprise 9.5; and (2) an AMD MI250X supercomputer—each node containing eight AMD EPYC 7A53 CPUs (512 cores 9

RLlib

150 100

42%

50 0

RLlib + CONDUIT

37% 71%

71% 94%

MC -CP MW XE MW XE DQN- DQN SAC- SAC-O PPO- PPO-O

97%

End-to-end

600

RLlib

RLlib + CONDUIT

400

24%

200 0

12%

Exposed experience-path latency (ms, log)

Exposed experience-path

Latency per iteration (ms)

Latency per iteration (ms)

200

24%

38%

30%

38%

MC -CP MW XE MW XE DQN- DQN SAC- SAC-O PPO- PPO-O

Fig. 11: Conduit’s per-iteration latency reduction over RLlib: exposed experience-path (left), end-to-end (right). Workloads: MC=MountainCar, CP=MA-CartPole, MW=MetaWorld, OXE=Open X-Embodiment

Single GPU Shared GPU

Bandwidth-dominated

103

CONDUIT

Capacity-dominated 127×

102

4.1×

101

128 1000 2000 5000 10000 20000

Transition dimension

Fig. 12: Placement ablation: Conduit vs. four placements (Reverb=CPU pageable, Gear=CPU pinned with zero-copy access, single GPU, shared GPU). Missing points denote OOM.

Conservative (all-GPU) Greedy (single-GPU) CONDUIT d = 128 d = 1000 d = 5000 d = 10000 d = 20000

iteration CumulativeEnd-to-end latency (s) latency (ms, log)

omit Gear and Reverb, which are standalone buffer services without compatible drop-in hooks. We test Conduit across three algorithms (DQN, SAC, PPO) and four environments: DQN on MountainCar and Multi-agent CartPole, and SAC and PPO on Meta-World and OXE (large visual observations that stress both capacity and bandwidth), using eight actors and eight learners. Fig. 11 reports exposed experience-path latency (left) and end-to-end iteration latency (right) for each (algorithm, environment) pair. Across all workloads, integrating Conduit with RLlib reduces exposed experience-path latency—by up to 71% for DQN, 71% for SAC, and 97% for PPO. These reductions translate into end-to-end improvements of up to 38% (DQN), 38% (SAC), and 30% (PPO). Two EDP levers drive these reductions: scheduling hides experience-path handling behind rollout and update, while bandwidth-aware placement distributes the large buffer across devices to keep transfer costs low. For off-policy DQN and SAC the savings reach end-to-end latency, as replay adds cross-iteration overlap. For on-policy PPO, Conduit cuts experience-path latency sharply, but end-to-end gains remain bounded since freshness limits batching and the residual rollout/update time dominates the critical path. Insights. Conduit’s end-to-end benefit scales with the experience path’s share of the iteration: the larger that share, the greater the gain from EDP’s placement and scheduling optimizations. 7.3

Reverb Gear

OOM

102 10

replacement (30 ms) replacement (30 ms) 1.9×

0

0

10

OOM

20 30 Iteration number

Fig. 13: Online migration as transition dimension 𝑑 grows, vs. Greedy (single GPU) and Conservative (all-GPU) baselines. Top: iteration latency (log scale); bottom: cumulative latency.

zero-copy access) instantiate, alongside single-GPU and sharedGPU placement. In the bandwidth-dominated regime (128– 5,000), all placements are feasible; single-GPU placement is fastest by eliminating CPU–GPU transfers, so Conduit adopts it, reducing latency by up to 130ˆ over Reverb and 2.9ˆ over Gear at 𝑑“5,000. Beyond 5,000, single GPU exceeds per-GPU memory and Conduit switches to shared GPU, which dominates via fast GPU–GPU links („10ˆ faster than CPU–GPU), leaving Conduit 127ˆ faster than Reverb and 4.1ˆ faster than Gear at 𝑑“20,000. These ratios primarily reflect memory-tier placement: both baselines keep experience CPU-resident, and the gap quantifies what bandwidthaware, GPU-resident placement buys. Online migration. Fig. 13 traces this adaptive process over time: as transition dimension grows from 128 to 20,000, Conduit automatically migrates from single GPU to shared GPU when the capacity boundary is crossed. We compare against two baselines that bracket the trade-off. Greedy selects the lowest-latency placement for the initial workload (single GPU per learner); it is fastest while it fits (52–103 ms) but exhausts GPU memory at 𝑑ě10,000. Conservative shards the buffer across all 8 GPUs from the start to guarantee feasibility at peak size; it stays feasible throughout but pays multi-GPU coordination overhead at every dimension (151–403 ms).

Bandwidth-Aware Placement

Using the Synthetic Environment, we sweep transition dimension from 128 to 20,000 with buffer capacity 106 and batch size 512. This sweep covers both axes of the placement dimension: for small transitions, the buffer fits within one GPU and placement simply optimizes bandwidth (bandwidthdominated regime); as transitions grow, the buffer exceeds per-GPU memory and capacity becomes the constraint—only placements with sufficient aggregate memory remain viable (capacity-dominated regime). Conduit vs. state-of-the-art buffer services. Fig. 12 compares Conduit against the four placements that Reverb (CPU pageable) and Gear (CPU pinned, with GPU-driven 10

CONDUIT g=8

102 101 100 32

64

128

256

Batch size

512

1024

103

CONDUIT g=2 CONDUIT g=4

Reverb Gear

CONDUIT g=8

102 101 100

128

512

1000 2000 5000 10000

Transition dimension

Exposed experience-path latency (ms, log)

CONDUIT g=2 CONDUIT g=4

Exposed experience-path latency (ms, log)

Exposed experience-path latency (ms, log)

Reverb Gear

CONDUIT placement-only (g=1, no overlap) CONDUIT placement+overlap (g=1)

Reverb Gear

102 100 128

512

1000

2000

5000

10000

Exposed experience-path latency (ms, log)

Transition dimension Fig. 14: Impact of granularity 𝑔 under varying batch sizes (left) and transition dimensions (middle) for 𝑔=𝑔ing =𝑔del ; the right panel shows the impact of execution timing 𝑜𝑣𝑒𝑟𝑙𝑎𝑝 under varying transition dimensions.

Conduit combines the strengths of both: it matches Greedy while feasible, then absorbs two brief „30 ms replacement spikes (at 𝑑“10,000 and 𝑑“20,000) and holds iteration latency at 146 and 266 ms in the two largest segments, 34–46% below Conservative once Greedy has already failed. Each replacement reallocates the buffer under the new shard layout, reusing freed HBM from the caching allocator so the spike stays in the tens of milliseconds. Over the full 40-iteration trace, Conduit accumulates 1.9ˆ less cumulative latency than Conservative, while Greedy cannot complete the run. Insights. Capacity and bandwidth jointly determine placement: Conduit’s cost model selects the lowest-latency feasible placement at each operating point, and the same model drives online migration when workload statistics shift, keeping placement aligned throughout training. 7.4

Manual configurations

103

CONDUIT

102 101 100 10 1

128

256

512

1000 2000 5000 10000

Transition dimension

Fig. 15: Auto vs. manual over the joint placement ˆ overlap ˆ granularity space. Gray: 32 manual configs; red star: Conduit’s auto pick.

overlap) suffer 27ˆ and 4.4ˆ growth in experience-path latency as transition dimension increases. By contrast, Conduit’s placement-only mode grows only 2.5ˆ (3.12 to 7.83 ms) by holding experience handling in fast memory; adding overlap then hides most of this residual behind rollout and update, further reducing the exposed latency. Insights. Conduit’s scheduling wins by combining overlap and granularity: overlap hides ingestion/delivery behind rollout and update, while 𝑔 amortizes per-invocation overheads.

Latency-Aware Scheduling

This section isolates the benefit of Conduit’s latency-aware scheduling. We study how the two control knobs—execution granularity (𝑔) and execution timing (𝑜𝑣𝑒𝑟𝑙𝑎𝑝)—affect exposed experience-path latency under increasing workload pressure, and compare against Reverb and Gear (Fig. 14). Execution granularity. Using the Synthetic Environment, we fix timing to sequential execution (𝑜𝑣𝑒𝑟𝑙𝑎𝑝 “ 0) and scale workload along two dimensions: (i) batch size 32–1,024 at fixed transition dimension 1,000, and (ii) transition dimension 128–10,000 at fixed batch size 512. The left and middle panels of Fig. 14 report Conduit’s exposed experience-path latency at 𝑔 P t2, 4, 8u against both baselines. Larger 𝑔 amortizes per-invocation overheads (kernel launch, inter-process communication, metadata, interrank synchronization) and enables batched GPU throughput (§6.2): on the batch sweep, exposed latency drops from 1.87 ms (𝑔“2) to 0.57 ms (𝑔“8) at batch 1,024. The gain then reverses on harder workloads: at dimension 10,000 the optimum shifts back to 𝑔“4 (2.33 ms) and 𝑔“8 regresses to 3.67 ms, as the per-invocation chunk grows large enough that managing it outweighs the amortization benefit. Execution timing. We next isolate timing by varying transition dimension from 128 to 10,000 and comparing four configurations: Reverb, Gear, Conduit placement-only (𝑔“1, 𝑜𝑣𝑒𝑟𝑙𝑎𝑝“0), and Conduit with overlap (𝑔“1, 𝑜𝑣𝑒𝑟𝑙𝑎𝑝“1). The right panel of Fig. 14 shows that Reverb and Gear (no

7.5

Joint Adaptive Control: Auto vs Manual

This subsection evaluates whether Conduit’s auto-optimizer correctly navigates the joint placementˆscheduling design space. Using the Synthetic Environment, we sweep 32 manual configurations (4 placements ˆ 2 overlap settings ˆ 4 granularities) across transition dimensions 128 to 10,000, and compare them against Conduit’s automatic choice. Fig. 15 shows that the cloud span widens from 82.9 ms at dimension“ 128 to 888.5 ms at dimension“ 10,000, and that the lowest manual point is reached by different configurations across dimensions: single GPU with overlap and 𝑔“8 at dimension“ 128, and single GPU with overlap and 𝑔“2 at dimension“ 10, 000. Conduit’s automatic choice sits on the lower envelope because the profile-driven cost model ranks candidates correctly at every operating point. Appendix F give further details of the cost model’s accuracy. Insights. Conduit’s auto-optimizer matches the best manual point on every workload without re-tuning, while any fixed configuration must compromise across workloads. 11

2

4

#GPUs

8

16

101

150 200

70 60 Ideal CONDUIT 50 CONDUIT w/o adaptive placement 8 16 32 64 128 256 512 1024 256 512 1024

#GPUs

4000

8000

PPO on Meta-World

4000 2000

RLlib + CONDUIT RLlib

10

20

30

40

50

Iteration number (c) Reward vs. training iteration.

(b) MI250X supercomputer: 8–1,024 GPUs at batch size 32,768.

Fig. 16: Strong scaling of Conduit with GPU count on an NVIDIA A100 cluster and an AMD MI250X supercomputer (end-to-end iteration time, log scale).

DQN on MountainCar

100

RLlib + CONDUIT RLlib

1.44× faster to 8000 iters

400

800

150 200

0

Wall clock time (seconds)

(b) Reward vs. wall-clock time.

PPO on Meta-World

1.25× faster to 50 iters

4000 2000 0

RLlib + CONDUIT RLlib

0

100

200

Wall clock time (seconds)

(d) Reward vs. wall-clock time.

Fig. 17: Convergence analysis for Conduit on RLlib.

placement further improves scaling by avoiding slow ranks under non-uniform intra-node links. 7.7

7.6

0

Iteration number (a) Reward vs. training iteration.

(a) A100 cluster: 1–16 GPUs at batch size 2,048. 103 80

102

RLlib + CONDUIT RLlib

Episode reward mean

1

CONDUIT Ideal

DQN on MountainCar

100

Episode reward mean

Episode reward mean

Reverb Gear

101

Episode reward mean

End-to-end iteration latency (ms, log)

End-to-end iteration latency (ms, log)

102

Convergence and Correctness Analysis

We evaluate (Fig. 17) whether Conduit changes RL training behavior by comparing the learning curves of RLlib with and without Conduit on DQN/MountainCar (off-policy) and PPO/Meta-World (on-policy), each with eight actors and eight learners. The left column reports episode reward versus training iteration; Conduit closely matches the baseline trend on both tasks, indicating no observable impact on convergence. The right column reports the same metric versus wall-clock time; Conduit reaches comparable reward levels sooner. This time-to-quality gain comes from reducing exposed experience-path latency, which shortens iteration time without changing algorithm semantics. Together with the invariant-preserving EDP boundary in §4, this result supports semantics preservation across the evaluated on-policy and off-policy settings.

Scalability Analysis

We evaluate Conduit’s strong-scaling behavior on two platforms (Fig. 16). On the A100 cluster, we fix the total rollout/training batch size to 2,048 and scale from 1 to 16 A100 GPUs, comparing against two buffer-service baselines, Gear and Reverb. On the MI250X supercomputer, we scale Conduit from 8 to 1,024 MI250X GPUs under a fixed workload of 32,768 and compare against Conduit-w/o-adaptiveplacement, which disables adaptive placement. Fig. 16a shows that Conduit scales more efficiently than both buffer-service baselines. From 1 to 16 GPUs, Conduit reduces end-to-end iteration latency from 95 ms to 15 ms (84% reduction), remaining consistently below both Gear and Reverb at every scale. At 16 GPUs, Conduit achieves 57% lower latency than Gear (15 vs. 35 ms) and 65% lower than Reverb (15 vs. 43 ms). All systems deviate from ideal linear scaling at larger GPU counts due to synchronization overhead, but Conduit maintains the smallest gap. Fig. 16b shows that Conduit scales into the supercomputer regime. From 8 to 1,024 GPUs, Conduit cuts end-toend iteration latency from 614 ms to 50 ms (92%), and beats Conduit-w/o-adaptive-placement by 9–21% at all counts (e.g., 110 vs. 140 ms at 64 GPUs, 50 vs. 55 ms at 1,024). This gap reflects adaptive placement on the MI250X’s non-uniform intra-node fabric (links of roughly 100–400 GB/s): Conduit prefers higher-bandwidth GPU ranks (0/2/4/6 in Fig. 4a) over slower ones (1/3/5/7). We omit Gear and Reverb here, as they lack AMD support [1, 10]; this experiment thus shows hardware generality and adaptive placement at scale rather than a head-to-head comparison. Insights. Conduit scales across both NVIDIA and AMD platforms, staying faster than baselines at every scale. It tracks the ideal trend at small scales and keeps a smaller gap at large scales by reducing exposed experience-path latency on the critical path. On the MI250X supercomputer, adaptive

8

Related Work

Replay buffers and distributed RL systems. Large RL systems decouple actors and learners via replay buffers and actor/learner parallelism—Ape-X [15], importance-weighted actor-learner architectures (IMPALA) [7], and SEED RL [6]. Other work scales multi-GPU distributed reinforcement learning (DRL) via finer-grained GPU sharing [46], and reinforcement learning from human feedback (RLHF)/large-model stacks optimize model execution and orchestration (RLlib [23], Verl [37], PUZZLE [20]), but these efforts target compute rather than the experience path. Replay-buffer services (Reverb [2]) provide APIs and sampling but treat the buffer as passive storage driven by actor/learner loops. Conduit is complementary: it treats the experience path itself as the optimization target, cutting exposed latency without changing algorithm logic. Decoupling and pipelining for RL training. Asynchronous actor–learner execution improves throughput (A3C [26], IMPALA [7]), but experience-path handling often runs inline 12

or as fixed background behavior, limiting control over when it runs and how much it processes. Conduit exposes this control explicitly through EDP. Data movement and placement in heterogeneous training stacks. Prior work reduces training overhead by optimizing placement and movement across heterogeneous memory and network fabrics [21, 22, 45], and recent RL systems extend this to placing computation—models and parallelism—across heterogeneous GPUs (HetRL [14]). RL exacerbates these costs because transient experiences repeatedly traverse the actor Ñ Experience Buffer Ñ learner path; Conduit targets this path with capacity-constrained, bandwidth-aware Experience Buffer placement. Summary. Across these lines of work, the experience path remains embedded in framework execution or behind passive services; Conduit instead exposes it as an explicit, independently optimizable runtime surface.

9

[7] Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Vlad Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, et al. 2018. IMPALA: Scalable distributed deep-RL with importance weighted actor-learner architectures. In International Conference on Machine Learning. PMLR, 1407–1416. [8] Farama Foundation. 2025. Gymnasium: Classic Control. https: //gymnasium.farama.org/environments/classic_control/. [9] Wei Fu, Jiaxuan Gao, Xujie Shen, Chen Zhu, Zhiyu Mei, Chuyi He, Shusheng Xu, Guo Wei, Jun Mei, Jiashu Wang, et al. 2025. AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning. arXiv preprint arXiv:2505.24298 (2025). [10] Google DeepMind. 2023. Reverb Issue #120: AMD GPU Support. https: //github.com/google-deepmind/reverb/issues/120. Accessed: 2026-0129. [11] Jiayuan Gu, Fanbo Xiang, Xuanlin Li, Zhan Ling, Xiqiang Liu, Tongzhou Mu, Yihe Tang, Stone Tao, Xinyue Wei, Yunchao Yao, et al. 2023. ManiSkill2: A unified benchmark for generalizable manipulation skills. In International Conference on Learning Representations. [12] Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. 2018. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In International Conference on Machine Learning. PMLR, 1861–1870. [13] Nicklas Hansen, Hao Su, and Xiaolong Wang. 2024. TD-MPC2: Scalable, robust world models for continuous control. In International Conference on Learning Representations. [14] Yongjun He, Shuai Zhang, Jiading Gai, Xiyuan Zhang, Boran Han, Bernie Wang, Huzefa Rangwala, and George Karypis. 2026. HetRL: Efficient Reinforcement Learning for LLMs in Heterogeneous Environments. In Proceedings of Machine Learning and Systems (MLSys). arXiv:2512.12476. [15] Dan Horgan, John Quan, David Budden, Gabriel Barth-Maron, Matteo Hessel, Hado Van Hasselt, and David Silver. 2018. Distributed prioritized experience replay. arXiv preprint arXiv:1803.00933 (2018). [16] Jian Hu, Xibin Wu, Wei Shen, Jason Klein Liu, Weixun Wang, Songlin Jiang, Haoran Wang, Hao Chen, Bin Chen, Wenkai Fang, et al. 2025. OpenRLHF: A Ray-based Easy-to-use, Scalable and High-performance RLHF Framework. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. 656– 666. [17] Zhe Huang. 2022. Distributed reinforcement learning for autonomous driving. Carnegie Mellon University. [18] Matthias Jasny, Tobias Ziegler, Tim Kraska, Uwe Roehm, and Carsten Binnig. 2020. DB4ML—An in-memory database kernel with machine learning support. In Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data. 159–173. [19] Linus Jern, Valter Uotila, Cong Yu, and Bo Zhao. 2025. Agent-Q: Fine-Tuning Large Language Models for Quantum Circuit Generation and Optimization. 2025 IEEE International Conference on Quantum Computing and Engineering (QCE) 01 (2025), 1621–1632. https://api. semanticscholar.org/CorpusID:277786843 [20] Kinman Lei, Yuyang Jin, Mingshu Zhai, Kezhao Huang, Haoxing Ye, and Jidong Zhai. 2024. PUZZLE: Efficiently Aligning Large Language Models through Light-Weight Context Switch. In 2024 USENIX Annual Technical Conference (USENIX ATC 24). USENIX Association, 127–140. [21] Ang Li, Shuaiwen Leon Song, Jieyang Chen, Jiajia Li, Xu Liu, Nathan R Tallent, and Kevin J Barker. 2019. Evaluating modern GPU interconnect: PCIe, NVLink, NV-SLI, NVSwitch and GPUDirect. IEEE Transactions on Parallel and Distributed Systems 31, 1 (2019), 94–110. [22] Ang Li, Shuaiwen Leon Song, Jieyang Chen, Xu Liu, Nathan Tallent, and Kevin Barker. 2018. Tartan: Evaluating modern GPU interconnect via a multi-GPU benchmark suite. In 2018 IEEE International Symposium on Workload Characterization (IISWC). IEEE, 191–202. [23] Eric Liang, Zhanghao Wu, Michael Luo, Sven Mika, Joseph E Gonzalez, and Ion Stoica. 2021. RLlib Flow: Distributed reinforcement learning is

Conclusion

We presented Conduit, a framework-agnostic runtime that makes the experience path an independently optimizable stage on the actor Ñ Experience Buffer Ñ learner path. Conduit realizes the Experience Data Plane (EDP), exposing experience ingestion, placement, and delivery as explicit runtime control points that run asynchronously while preserving on-policy freshness and off-policy replay semantics. Guided by measured bandwidths and a compact latency model, it applies capacity-constrained, bandwidth-aware placement and latency-aware scheduling, reducing exposed experiencepath latency and improving end-to-end training across distributed RL workloads.

References [1] BigRL Team. 2023. GEAR: GPU-centric Experience Replay System. https://github.com/bigrl-team/gear. Accessed: 2026-01-29. [2] Albin Cassirer, Gabriel Barth-Maron, Eugene Brevdo, Sabela Ramos, Toby Boyd, Thibault Sottiaux, and Manuel Kroiss. 2021. Reverb: A framework for experience replay. arXiv preprint arXiv:2102.04736 (2021). [3] Yuming Chen, Jiangyan Feng, Haodong Zhang, Lijun Gong, Feng Zhu, Rui Zhao, Qibin Hou, Ming-Ming Cheng, and Yibing Song. 2025. ReAligning Language to Visual Objects with an Agentic Workflow. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net. https://openreview. net/forum?id=MPJ4SMnScw [4] Pramod Chunduri, Jaeho Bang, Yao Lu, and Joy Arulraj. 2022. Zeus: Efficiently localizing actions in videos using reinforcement learning. In Proceedings of the 2022 International Conference on Management of Data. 545–558. [5] Daniele De Sensi, Lorenzo Pichetti, Flavio Vella, Tiziano De Matteis, Zebin Ren, Luigi Fusco, Matteo Turisini, Daniele Cesarini, Kurt Lust, Animesh Trivedi, et al. 2024. Exploring GPU-to-GPU communication: Insights into supercomputer interconnects. In SC24: International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, 1–15. [6] Lasse Espeholt, Raphaël Marinier, Piotr Stanczyk, Ke Wang, and Marcin Michalski. 2019. SEED RL: Scalable and efficient deep-RL with accelerated central inference. arXiv preprint arXiv:1910.06591 (2019). 13

[39] Zheyue Tan, Mustapha Abdullahi, Tuo Shi, Huining Yuan, Zelai Xu, Chao Yu, Boxun Li, and Bo Zhao. 2025. EARL: Efficient Agentic Reinforcement Learning Systems for Large Language Models. CoRR abs/2510.05943 (2025). doi:10.48550/ARXIV.2510.05943 arXiv:2510.05943 [40] Jordan Terry, Benjamin Black, Nathaniel Grammel, Mario Jayakumar, Ananth Hari, Ryan Sullivan, Luis S Santos, Clemens Dieffendahl, Caroline Horsch, Rodrigo Perez-Vicente, et al. 2021. PettingZoo: Gym for multi-agent reinforcement learning. Advances in Neural Information Processing Systems 34 (2021), 15032–15043. [41] The Ray Team. [n. d.]. Episodes — RLlib Documentation. Ray Documentation. https://docs.ray.io/en/latest/rllib/single-agent-episode.html Accessed: 2026-01-17. [42] The TensorFlow Authors. [n. d.]. tf_agents.trajectories.Trajectory. TensorFlow Agents API Documentation. https://www.tensorflow.org/ agents/api_docs/python/tf_agents/trajectories/Trajectory Accessed: 2026-01-17. [43] Sathish S Vadhiyar, Graham E Fagg, and Jack Dongarra. 2000. Automatically tuned collective communications. In SC’00: Proceedings of the 2000 ACM/IEEE Conference on Supercomputing. IEEE, 3–3. [44] Hanjing Wang, Man-Kit Sit, Congjie He, Ying Wen, Weinan Zhang, Jun Wang, Yaodong Yang, and Luo Mai. 2023. GEAR: A GPU-centric experience replay system for large reinforcement learning models. In International Conference on Machine Learning. PMLR, 36380–36390. [45] Qi Wang, Ludmila Cherkasova, Jun Li, and Haris Volos. 2016. Interconnect emulator for aiding performance analysis of distributed memory applications. In Proceedings of the 7th ACM/SPEC on International Conference on Performance Engineering. 75–83. [46] Yuke Wang, Boyuan Feng, Zheng Wang, Guyue Huang, Tong Geng, Ang Li, and Yufei Ding. 2025. GMI-DRL: Empowering Multi-GPU DRL with Adaptive-Grained Parallelism. In 2025 USENIX Annual Technical Conference (USENIX ATC 25). USENIX Association, 89–103. [47] Zhengtong Yan, Valter Uotila, and Jiaheng Lu. 2023. Join order selection with deep reinforcement learning: fundamentals, techniques, and challenges. Proceedings of the VLDB Endowment 16, 12 (2023), 3882–3885. [48] Zhewei Yao, Reza Yazdani Aminabadi, Olatunji Ruwase, Samyam Rajbhandari, Xiaoxia Wu, Ammar Ahmad Awan, Jeff Rasley, Minjia Zhang, Conglong Li, Connor Holmes, et al. 2023. DeepSpeed-Chat: Easy, fast and affordable RLHF training of ChatGPT-like models at all scales. arXiv preprint arXiv:2308.01320 (2023). [49] Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Weinan Dai, Tiantian Fan, Gaohong Liu, Lingjun Liu, et al. 2025. DAPO: An open-source LLM reinforcement learning system at scale. arXiv preprint arXiv:2503.14476 (2025). [50] Tianhe Yu, Deirdre Quillen, Zhanpeng He, Ryan Julian, Karol Hausman, Chelsea Finn, and Sergey Levine. 2020. Meta-World: A benchmark and evaluation for multi-task and meta reinforcement learning. In Conference on Robot Learning. PMLR, 1094–1100. [51] Ruiqi Zhang, Guang Chen, Jing Hou, Zhijun Li, and Alois Knoll. 2022. PIPO: Policy optimization with permutation-invariant constraint for distributed multi-robot navigation. In 2022 IEEE International Conference on Multisensor Fusion and Integration for Intelligent Systems (MFI). IEEE, 1–7. [52] Yinmin Zhong, Zili Zhang, Bingyang Wu, Shengyu Liu, Yukun Chen, Changyi Wan, Hanpeng Hu, Lei Xia, Ranchen Ming, Yibo Zhu, et al. 2025. Optimizing RLHF training for large language models with stage fusion. In 22nd USENIX Symposium on Networked Systems Design and Implementation (NSDI 25). 489–503. [53] Huanzhou Zhu, Bo Zhao, Gang Chen, Weifeng Chen, Yijie Chen, Liang Shi, Yaodong Yang, Peter Pietzuch, and Lei Chen. 2023. MSRL: Distributed reinforcement learning with dataflow fragments. In 2023 USENIX Annual Technical Conference (USENIX ATC 23). 977–993.

a dataflow problem. Advances in Neural Information Processing Systems 34 (2021), 5506–5517. [24] Zhihong Liu, Xin Xu, Peng Qiao, and Dongsheng Li. 2024. Acceleration for deep reinforcement learning using parallel and distributed computing: A survey. Comput. Surveys 57, 4 (2024), 1–35. [25] Zhiyu Mei, Wei Fu, Jiaxuan Gao, Guangju Wang, Huanchen Zhang, and Yi Wu. 2023. SRL: Scaling distributed reinforcement learning to over ten thousand cores. arXiv preprint arXiv:2306.16688 (2023). [26] Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. 2016. Asynchronous methods for deep reinforcement learning. In International Conference on Machine Learning. PMLR, 1928– 1937. [27] Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller. 2013. Playing Atari with Deep Reinforcement Learning. arXiv preprint arXiv:1312.5602 (2013). [28] Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al. 2015. Human-level control through deep reinforcement learning. Nature 518, 7540 (2015), 529– 533. [29] Philipp Moritz, Robert Nishihara, Stephanie Wang, Alexey Tumanov, Richard Liaw, Eric Liang, Melih Elibol, Zongheng Yang, William Paul, Michael I. Jordan, and Ion Stoica. 2018. Ray: A Distributed Framework for Emerging AI Applications. In 13th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2018, Carlsbad, CA, USA, October 8-10, 2018, Andrea C. Arpaci-Dusseau and Geoff Voelker (Eds.). USENIX Association, 561–577. https://www.usenix.org/conference/ osdi18/presentation/nishihara [30] Daniel Eugênio Neves, Lucila Ishitani, and Zenilton Kleber Goncalves do Patrocinio Junior. 2024. Advances and challenges in learning from experience replay. Artificial Intelligence Review 58, 2 (2024), 54. [31] Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems 35 (2022), 27730–27744. [32] Abby O’Neill, Abdul Rehman, Abhiram Maddukuri, Abhishek Gupta, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, Ajinkya Jain, et al. 2024. Open X-Embodiment: Robotic learning datasets and RT-X models. In 2024 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 6892–6903. [33] Qwen. 2025. Qwen2.5-Math-7B. https://huggingface.co/Qwen/Qwen2. 5-Math-7B. [34] Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver. 2016. Prioritized Experience Replay. In 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedings, Yoshua Bengio and Yann LeCun (Eds.). http://arxiv.org/abs/1511.05952 [35] John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347 (2017). [36] Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Yang Wu, et al. 2024. DeepSeekMath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300 (2024). [37] Guangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu, Wang Zhang, Ru Zhang, Yanghua Peng, Haibin Lin, and Chuan Wu. 2025. HybridFlow: A flexible and efficient RLHF framework. In Proceedings of the Twentieth European Conference on Computer Systems. 1279–1297. [38] Richard S. Sutton and Andrew G. Barto. 2018. Reinforcement Learning: An Introduction (2nd ed.). MIT Press.

14

A

Using Conduit for LLM Post-Training

Tab. 2: Conduit on SRL: per-iteration end-to-end and exposed experience-path latency (ms).

Conduit targets traditional distributed RL, but its EDP is defined over the experience path and is independent of framework execution logic. We show that it also generalizes to RL-based LLM post-training by integrating it into Verl [37]. Ingestion and delivery costs. In LLM post-training, the work that turns rollouts into training-ready batches is substantial: each sample requires reward (and, for actor-critic methods, value) computation [37, 52], and batches are variablelength sequences carrying auxiliary fields (e.g., rewards and masks) [16, 31]. Conduit abstracts this work as its two experience-path operators, ingestion and delivery, both of which lie on the critical path. Integration with Verl. Verl [37] is a state-of-the-art RLHF framework whose per-iteration dataflow forms an experience path: each iteration generates rollouts and then turns them—through log-prob, reward, and advantage computation— into the training batch consumed by the update. Conduit maps this post-generation processing onto experience ingestion and the hand-off of the training batch to the learner onto experience delivery (Tab. 3). By exposing these as independently schedulable operators, the EDP lets Conduit overlap ingestion work that Verl otherwise runs serially on the critical path: here, reward depends only on the generated sequences, so Conduit runs it concurrently with the remaining post-generation processing rather than afterward, without changing Verl’s reward logic, loss, or compute engines. Setup and results. We evaluate this integration by running GRPO [36] on the DAPO-Math-17k [49] dataset with the Qwen2.5-Math-7B [33] model for the task of generating Python code to solve math problems on an 8-GPU NVIDIA A100 node (§7.1). Integrating Conduit reduces exposed experience-path latency from 42 s to 20 s per iteration (52% reduction), which translates into a 4% end-to-end reduction, as LLM rollout and gradient update dominate iteration time; at industrial post-training scale, even single-digit per-iteration savings accumulate into substantial wall-clock and cost savings.

B

End-to-end Workload

+Conduit

Exp-path SRL

+Conduit

DQN–Qbert 275.8 191.3 (Ó31%) 106.6 PPO–Meta-World 235.9 193.6 (Ó18%) 63.5

18.6 (Ó83%) 23.2 (Ó63%)

SRL

learner uses SRL’s default policy network. Tab. 2 reports per-iteration latency. Integrating Conduit reduces exposed experience-path latency by 83% for DQN and 63% for PPO, translating into 31% and 18% lower end-to-end iteration latency, respectively. We observe a similar off-policy vs. onpolicy trend as in RLlib (§7.2): replay permits batching and cross-iteration overlap, so DQN sees larger end-to-end gains; for PPO, rollout and update dominate iteration time to begin with, and freshness constraints further limit how much of the remaining experience-path latency can be hidden.

C

Correctness Invariants

Conduit preserves RL training correctness by maintaining two invariants. Invariant 1: Training-data integrity. Conduit preserves the semantic integrity of experience tuples along the actor Ñ Experience Buffer Ñ learner path. (1) The Experience Data Plane introduces lightweight boundaries but does not change contents; it only stages data for decoupled execution (§4). (2) Bandwidth-aware placement changes only where experiences reside (CPU/GPU, local/remote), not what they contain, so tuple semantics and batch composition are preserved (§5). (3) Any processing during experience ingestion and experience delivery is algorithm-defined and registered through Conduit’s interface (§3); Conduit schedules and executes these functions without altering their behavior. Invariant 2: Algorithm-specific freshness constraints. Conduit enforces the freshness and replay semantics required by the RL algorithm. For on-policy training, updates must consume samples from the current policy iteration, so Conduit allows only within-iteration mini-batch pipelining and excludes cross-iteration reuse by construction (§4). For off-policy training, replay across iterations is permitted, so Conduit may batch and prefetch samples while bounding staleness through the EDP configurations (§4). Replay state such as priorities [34] is updated by algorithm-defined callbacks (Invariant 1(3)), whose buffer dependencies Conduit leaves intact. In all cases, the scheduler chooses overlap and granularity only from these feasible spaces, ensuring optimization never violates correctness (§6).

Using Conduit on SRL

We further verify this by integrating Conduit into SRL [25]. Integration with SRL. SRL [25] is a state-of-the-art distributed RL framework that splits RL training across actor, policy, and trainer workers operating concurrently via streaming dataflows. Its experience-path, however, remains embedded: it resides in pinned CPU memory and runs reactively with a fixed one-step overlap (Tab. 1). Conduit dropin replaces it: SRL’s workers retain their streaming control flow, while Conduit takes over experience placement and the timing and granularity of ingestion and delivery. Setup and results. We run off-policy DQN on Qbert [27] and on-policy PPO on Meta-World [50] on the NVIDIA A100 cluster (§7.1), with eight actors and eight learners; each

D

Placement-Migration Details

We detail the online placement migration introduced in §5.3: the runtime procedure (Alg. 2) and the transfer-plan formulation that minimizes migration makespan. 15

Tab. 3: Conduit mapped onto Verl’s per-iteration experience path (GRPO): post-generation processing (log-prob, reward, advantage) is ingestion; the hand-off of the training batch to the update workers is delivery. # Verl stage 1 generate_sequences (rollout) 2 old / ref log-prob 3 reward, KL penalty 4 compute_advantage 5 dispatch batch to update workers

Algorithm 2: Online placement migration. Input: Placement map M, pre-solved transfer plans F, active placement 𝑝 ˚ , per-sample size 𝑥. Output: Updated active placement 𝑝 ˚ ˚ 1 𝑝 new Ð Mp𝑥q ; // Oplog 𝑅q lookup ˚ ˚ ˚ 2 If 𝑝 new “ 𝑝 then return 𝑝 ; // no migration needed ˚ ˚ 3 t𝑓𝑖 𝑗 u Ð Fp𝑝 , 𝑝 new q ; // pre-computed plan 4 Pause ingestion and delivery operators; 1 5 foreach p𝑑𝑖 , 𝑑 𝑗 q with 𝑓𝑖 𝑗 ą 0 in parallel do 6 Transfer 𝑓𝑖 𝑗 data: 𝑑𝑖 Ñ𝑑 1𝑗 on dedicated stream;

Role actor (upstream)

EDP

ingestion

˚ 𝑝 ˚ Ð 𝑝 new ; Resume ingestion and delivery under 𝑝 ˚ ; ˚ 9 return 𝑝 ;

7

delivery

8

6 update_actor

learner (downstream)

Tab. 4: Cost-model accuracy per workload dimension 𝑑 of Fig. 15 (32 configurations each). For every dimension, the predicted ordering of the top-8 configurations exactly matches the measured ordering.

During the offline profiling phase, Conduit sweeps a discretized range of per-sample sizes and solves the placement scan once per point, producing a compact placement map M. At runtime, migration resolves Mp𝑥q, briefly pauses ingestion and delivery, transfers data to the new placement, and resumes under the updated placement. The EDP boundaries keep actors and learners decoupled while the buffer reallocates. A placement specifies a device set and the data volume held on each device. Migrating from placement 𝑝 (device set 𝐷) to 𝑝 1 (device set 𝐷 1 ) requires deciding how much data each source device sends to each target device. Let 𝑓𝑖 𝑗 denote the volume sent from 𝑑𝑖 P 𝐷 to 𝑑 1𝑗 P 𝐷 1 ; since transfers execute in parallel on dedicated streams, the migration makespan is the slowest individual transfer. Conduit minimizes this makespan: min max

𝑓𝑖 𝑗 ě0 𝑖,𝑗

Best 𝑇iter (ms) Measured

128 256 512 1,000 2,000 5,000 10,000

17.1 18.3 25.4 26.4 30.0 42.6 48.6

Predicted Top-8 ranking match 20.1 20.8 29.9 33.9 35.8 43.2 58.9

✓ ✓ ✓ ✓ ✓ ✓ ✓

Thus, in off-policy batching (𝑔 ě 1), overlap can hide the full per-unit operator cost, while in on-policy fractional pipelining (0 ă 𝑔 ă 1), only the unexposed fraction scales with p1 ´ 𝑔q.

ř ř 𝑓𝑖 𝑗 1 𝑗 𝑓𝑖 𝑗 “ 𝑉𝑖 , 𝑖 𝑓𝑖 𝑗 “ 𝑉 𝑗 , (3) 1 s.t. 𝐵𝑊 p𝑑𝑖 Ñ𝑑 𝑗 q

F

where 𝑉𝑖 is the data volume on 𝑑𝑖 and 𝑉𝑗1 is the required volume on 𝑑 1𝑗 . Since the placement map has only a small number of regions, Conduit pre-computes transfer plans for the expected transitions and reuses them online.

E

𝑑

Cost-Model Accuracy

Conduit’s optimizer selects configurations by argmin over cost-model predictions, so rank fidelity matters more than pointwise accuracy. We evaluate both on the 224 manual configurations of Fig. 15 (seven workload dimensions ˆ 32 configurations each), predicting each configuration’s 𝑇iter with the model of §6.2 and comparing against measurement (Tab. 4). The comparison here uses 𝑇iter , the optimizer’s objective in Eq. (2), rather than the exposed experiencepath latency plotted in Fig. 15. Pointwise, predictions follow measurements with near-perfect linear correlation (Pearson 𝑟 “0.990) and low absolute error (MAPE 20.8%). Rank-wise, the model is stronger still: for every dimension, the predicted ordering of the eight lowest-latency configurations exactly matches the measured ordering, so the optimizer always picks the truly best configuration.

Iteration-Latency Model Details

We expand the iteration-latency model introduced in §6.2, making explicit the hidden-latency terms behind its compact exposed-latency form. For ingestion and delivery, overlap can hide at most the shorter of the neighboring compute stage and the per-unit operator cost: ` ˘ ℎ𝑖𝑑𝑒 ing “ 𝛾p𝑔ing q ¨ min 𝑇rollout, 𝑇¯ing p𝑔ing q , (4) ` ˘ ℎ𝑖𝑑𝑒 del “ 𝛾p𝑔del q ¨ min 𝑇update, 𝑇¯del p𝑔del q . (5) The factor 𝛾p𝑔q captures how much of that bound is practically overlapable: # 1, 𝑔ě1 𝛾p𝑔q “ (6) 1 ´ 𝑔, 0 ă 𝑔 ă 1. 16

Record · ID 1028653 · SHA-256 c05c66eba7440173
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.