From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention Sichang Su1 , Benjamin Yang2 , Zhiyun Deng1 , Boyuan Liang3 , Yip Fun Yeung2 , Zelin Wang2 , and Lingfeng Sun2 https://destiny000621.github.io/PARTS/
Insert 2nd
Close case & place
PARTS: Policy Adaptation with RL on Targeted Subtasks Frozen base policy
Base policy
✓
✗
✗
✗
Full-task RL
one policy
✗
✗
✗
✓
RL0
RL1
RL2
PARTS
✓
Residual RL policy
✗ ✗
✓
action
local reward
Success ✓ /✗ verifier
local reset
Insert 1st
✓ Tasks
arXiv:2609.21788v1 [cs.RO] 18 Sep 2026
Approach case Open & hold case
base policy reward
bottleneck subtask reset from full task start
residual RL policies local reset
Earbud insertion
LEGO sorting
Cable unplug & plug
Fig. 1. PARTS overview. A pretrained policy completes most subtasks of a long-horizon task but fails at a few bottlenecks; full-task success requires every subtask to succeed. Full-task RL fine-tunes a single policy using a sparse reward delivered only upon full-task success. Each bottleneck failure triggers an episode restart, so progressively fewer rollouts reach later subtasks (narrowing bar). In contrast, PARTS retains the frozen base policy for subtasks it already handles reliably and trains one residual policy per bottleneck, with a local reward for every attempt.
Abstract— A pretrained robot foundation policy may execute most of a long-horizon task yet repeatedly fail at a few critical subtasks. Collecting additional full-task demonstrations for supervised fine-tuning (SFT) requires operators to repeat behaviors the policy already performs well. Reinforcement learning (RL) fine-tuning offers a promising path to bridge this gap, but existing approaches struggle to solve long-horizon tasks using only sparse rewards. We present PARTS (Policy Adaptation with RL on Targeted Subtasks), a real-world subtask RL framework that concentrates practice at these bottlenecks while allowing training rollouts to proceed with minimal human intervention. The frozen pretrained policy supplies nominal actions throughout execution, while agent-generated selectors and success verifiers activate residual corrections and provide local outcome rewards. These rewards support learning from successful subtasks even when complete-task successes are scarce. Training combines online RL with success-reweighted retraining, and each retrained residual policy is redeployed to collect further experience. Humans identify bottlenecks during setup and perform physical resets when needed. On bimanual YAM and single-arm Franka tasks, PARTS improves completetask success from 32% to 61% and from 50% to 95%, respectively, using tens of minutes of real-world RL rollouts per task on average. Compared with existing real-world RL fine-tuning methods, PARTS raises full-task success by more than 25% under the same robot-rollout budget while requiring less human involvement.
1 UT Austin. 2 Autel US. 3 UC Berkeley.
Corresponding authors: Sichang Su ([email protected]) and Lingfeng Sun ([email protected]).
I. INTRODUCTION Pretrained robot policies offer useful behaviors for adapting to new manipulation tasks. Recent vision-language-action (VLA) models and world-action models (WAMs) draw on large robot datasets, visual and semantic knowledge, and video prediction to produce increasingly capable policies [1], [2], [3], [4]. A target task can nevertheless demand changes in grasp strategy, contact behavior, spatial arrangement, or coordination between successive actions. Supervised finetuning (SFT) on additional demonstrations can address these differences. For a long-horizon task, collecting complete demonstrations repeatedly incurs human effort even for behaviors that already work. We study how to make better use of a pretrained policy’s existing capabilities while learning the changes needed for reliable target-task execution. Our starting observation is that failures in the tasks we study concentrate at a few consequential subtasks. These bottlenecks include high-precision operations, such as earbud insertion, and preparatory subtasks whose terminal states affect subsequent execution. For example, a robot may successfully pick up an earbud but hold it in a pose that makes insertion difficult. Adaptation may therefore target both precisioncritical motions and earlier actions that establish suitable grasps or placements, while retaining the reliable behaviors of the pretrained policy. Real-world RL provides a way to improve existing behaviors through interaction, including by learning residual
corrections to a fixed policy [5]. If learning uses only full-task success, an improved intermediate behavior can still receive no positive reward when a later step fails. Repeated early failures also reduce opportunities to practice later subtasks. Local episodes with verifiable outcomes can provide successful experience before complete-task execution becomes reliable, provided the initial policy supports productive local exploration. Prior work uses planned subgoals for online adaptation [6] and refines selected task phases [7], [8]. Our focus is on organizing such local learning to adapt a pretrained policy across the bottlenecks of a long-horizon physical task. We train on bottleneck subtasks and measure progress by success of the complete task. Local practice must integrate with full-task execution. Each subtask requires reachable entry states and a success criterion that captures readiness for the next stage. The system must manage continuation, retries from the current state, and physical reset requests, while selecting bottlenecks for improvement and deciding when to deploy updated policies. The framework must therefore coordinate episode supervision, task execution, and policy improvement with minimal human involvement. We present PARTS (Policy Adaptation with RL on Targeted Subtasks), a real-world subtask RL framework for adapting pretrained robot policies, as shown in Fig. 1. A frozen base policy supplies nominal actions throughout execution, and lightweight, bounded residual policies modify those actions within selected bottlenecks. Given humanidentified bottlenecks, coding agents construct executable policy selectors and success verifiers; VLMs act as languageconditioned visual perception modules. The selectors govern residual activation, and the verifiers supply sparse local outcome rewards. Online RL improves the residuals with newly collected robot experience. Success-reweighted retraining emphasizes successful local experience in the accumulated data, and the retrained policies are deployed to collect further rollouts. Retraining thus affects both the policy and the experience available for further learning. Training rollouts require no human action corrections, switching decisions, or outcome labels. Humans specify bottlenecks, assist with setup, and perform physical resets when needed. We instantiate our framework using π0.5 as the backbone foundation model, a state-of-the-art generalist policy that has demonstrated strong performance across diverse manipulation tasks [1]. We evaluate PARTS on two bimanual long-horizon tasks with YAM and one single-arm long-horizon task with Franka; earbud and cable insertion require millimeter-level precision, while LEGO sorting involves bricks that are out of distribution for the base policy. Full-task success increases from 32% to 61% on YAM and from 50% to 95% on Franka, requiring only tens of minutes of real-world RL rollouts per task on average. Compared with existing real-world RL finetuning methods, PARTS raises full-task success by more than 25% while requiring less human involvement. Our contributions are: • A formulation for adapting pretrained robot policies through RL at selected bottlenecks, using local outcome
TABLE I Human involvement during RL data collection, as published for the baselines and PARTS. ✓: autonomous; ✗: requires a human. Rollout: corrective interventions or handoff decisions during RL rollouts; Reward: success labeling; Reset: scene restoration.
Method
Rollout
Reward
Reset
DSRL [9] EXPO-FT [10] RLT [7]
✓ ✗ ✗∗
✗ ✓† ✗
✗ ✗ ✗
PARTS (ours)
✓
✓‡
✓‡
∗ A human hands control from the VLA to the RL policy in every episode; corrective interventions are optional. † Rule-based detectors, with a human override key in the released code. ‡ Automatic on LEGO and cable; on
earbuds a human checks the verifier’s labels, and a human resets when the robot cannot restore the scene.
supervision while evaluating reliability over complete long-horizon execution. • PARTS, a framework connecting executable subtask supervision, residual RL, and success-reweighted retraining with redeployment, enabling training rollouts without human corrections, switching decisions, or reward labels. • Experiments on the bimanual YAM and single-arm Franka platforms showing that PARTS boosts full-task success even from weak initial policies under a limited real-world rollout budget. II. RELATED WORK A. Real-World RL in Long-Horizon Tasks Unlike sim-to-real RL pipelines [11], [12], which train with RL in simulation before hardware deployment, realworld RL updates a policy through physical interaction, where data collection, resets, and unproductive exploration are costly. Prior systems improve efficiency through demonstrations [13], [14], residual control [15], [5], learned reward functions [16], and human intervention [17]. Long-horizon tasks introduce an additional challenge because rewards are often sparse, prior work decomposes them into shorterhorizon goals or reusable skills, composing demonstrated primitives [6], discovering skills from prior experience [18], [19], or sharing parameters across related tasks [20]. More recently, pretrained VLA policies have become increasingly capable of executing complex tasks, shifting the problem from learning behaviors from scratch toward improving already capable policies [21], [1], [2]. Online RL has consequently been used to further adapt pretrained VLA policies to real-world tasks [22], [7], [10]. However, these methods emphasize global policy improvement. PARTS instead adopts failure-localized policy repair, concentrating real-world RL practice on the few bottleneck subtasks that limit longhorizon success while retaining already reliable behaviors. Several recent systems also restrict learning to part of a task [23], [24], [25]. PARTS likewise targets base-policy failures, but delimits bottlenecks through executable contracts rather than action variance or a fixed contact phase.
Target subtask identification and setup
Real-world RL on targeted bottlenecks
Human: identify bottlenecks from failures ✗
✗
Offline demos
✗
Task decomposition
Bottleneck selection
Subtask 1
Subtask 1
Subtask 2
Bottleneck 1
Subtask 3
Subtask 3
Subtask 4
Bottleneck 2
Reset auto / human
SFT
human feedback
Frozen base policy
Coding agents
nominal
at Success verifier
residual
Policy selector
Auto-reset
evidence
rk
update θk
Success verifier
VLM
sparse local reward rk
Residual policy k
Policy selector
Human-specified contracts
Real robot
Online TD3+BC
Replay buffer
actor + twin critics
per bottleneck 𝒟k
Full-task policy inference during evaluation Full-task Base
Base + residual
selector
…
…
redeploy Base
Retrain successes + ρ failures
verifier
Fig. 2. PARTS architecture. Left: Humans identify bottlenecks from base-policy failures and specify contracts. Coding agents implement selection, verification, and reset programs. Right: Residual policies correct the frozen base policy. Local rewards and per-bottleneck replay support online TD3+BC and retraining on all successes plus a fraction ρ of failures. Retrained policies resume online learning. At evaluation, PARTS enables the switch between the base and residual policies to complete the full task.
Automatic verifiers provide local outcome rewards instead of intervention-derived rewards. B. Human Supervision for Real-World Policy Adaptation Human supervision is commonly used in real-world policy learning to guide exploration and prevent costly failures through corrective actions [26], policy takeover or switching [17], [7], and success/failure labeling [27], [7]. Although effective, these approaches rely on human monitoring or intervention during task execution, limiting their scalability. Recent systems automate parts of this supervision: UniIntervene [28] detects value degradation and retrieves recovery behaviors, and Robot Trains Robot [29] automates protection, failure detection, scheduling, and resets, in both cases for a policy that practices the whole task. PARTS instead confines practice to bottleneck subtasks, so that outcome labeling and resets become local problems handled by executable verifiers and reset programs, and RL rollouts need no human corrections or switching decisions. III. METHOD A. Problem Statement We consider fine-tuning a pretrained VLA policy πVLA (π0.5 in this work) with real-world RL. πVLA is a generalist trained on a diverse demonstration corpus and conditioned on a language instruction [1]. Like most modern VLAs, it uses action chunking: at each replan it predicts a chunk Āt = (āt , . . . , āt+C−1 ) of C future actions and executes a prefix of E ≤ C actions before predicting again [1]. Observations ot consist of multi-view RGB images and the proprioceptive state pt . For tasks on which πVLA exhibits limited zero-shot performance, we assume access to a small offline dataset of expert demonstrations, Dexp , obtained through human teleoperation and from open-source datasets. SFT on Dexp yields the task-specific base policy, denoted π0 , from which all RL methods in this paper start. The target tasks are long-horizon. We model an episode as a sequence of N subtasks σ1 , . . . , σN , for example opening a
case, inserting a first earbud, and inserting a second. Subtask i is entered from a set of states Ei ⊂ S and ends when its outcome criterion ϕi : S → {0, 1} is evaluated at exit or a time limit expires. The only reward is a sparse binary signal R ∈ {0, 1} for task completion, provided by a success verifier or a human at the end of the episode; no dense shaping is available. Because the task succeeds only if every subtask succeeds, this terminal reward factorizes as R=
N Y
ϕi ,
(1)
i=1
although a learner that observes only R never sees the individual ϕi . If the base policy completes subtask i with probability pi from the states it typically enters, then Pr[R = Q 1] ≈ i pi , so a few subtasks with low pi bound completetask success even when the remaining stages are reliable. We call these bottleneck subtasks and write K ⊆ {1, . . . , N } for their index set, K = |K|. A bottleneck may begin before the visible failure. When the terminal state of σi−1 determines the difficulty of σi , the bottleneck extends into σi−1 , whose success criterion requires a configuration suitable for executing σi . For example, case preparation succeeds only when the lid is open and the case is held in an orientation that facilitates earbud insertion. Bottlenecks may form an ordered chain or a set of alternatives, and the same bottleneck may recur within one episode, such as a grasp that repeats for every object. The entry distribution µi of subtask i is induced by executing σ1 , . . . , σi−1 from the task’s initial conditions; a subtask may also be entered from a restaged distribution µ̂i prepared by a reset procedure, which need not equal µi . The objective is to maximize the full-task success rate. This setting is hard for two reasons. First, the reward (1) is sparse over a long horizon, and when the base policy rarely completes the task, most rollouts return no signal at all. Second, real-robot training time is limited, so a budget spent uniformly over the task leaves few attempts at the bottlenecks.
B. Policy Adaptation with RL on Targeted Subtasks Fig. 2 summarizes our recipe for adapting a pretrained policy to a long-horizon task with real-world RL. The core idea is to spend robot interaction only where the pretrained policy fails. Fine-tuning the whole task online would revisit subtasks the base policy already performs and would receive the sparse terminal reward (1) only when every stage succeeds. Instead, we keep the VLA frozen: it executes every subtask, supplies reference action chunks and visual features, and defines the neighborhood in which small residual policies may act. We first fine-tune the VLA on task demonstrations and identify the bottlenecks where it still fails. For each bottleneck we then train a lightweight residual actor-critic with online TD3+BC, bounded to a few action dimensions around the VLA’s reference chunk and rewarded by that subtask’s own outcome, so that each attempt yields a usable signal. Executable programs authored by coding agents determine when to activate a residual policy, whether an attempt has succeeded, and how to reset, enabling training rollouts without human corrections or manual policy switching. Periodic success-reweighted retraining consolidates the rare successes and redeploys the residual to collect further experience. At inference, the fixed residuals are activated at their entries and hand control back to the base policy at their exits. This design turns real-world RL into targeted refinement of a few behaviors while the rest of the task retains the reliability of the pretrained model. a) Residual RL: Each time the frozen base policy is queried, it supplies the nominal chunk Āt . For bottleneck k ∈ K, a residual actor fθk (zt , pt , Āt ) outputs a normalized residual chunk Uk,t ∈ [−1, 1]C×dk , conditioned on visual features zt extracted by the base policy, proprioception pt , and the nominal chunk. A binary coordinate mask Mk and physical bounds Bk , specified per bottleneck (Sec. IIIC), restrict the correction. Writing gt ∈ {0} ∪ K for the active bottleneck, with zero denoting nominal execution, the command is ! X at = C āt + 1[gt = k] Bk ⊙ Mk ⊙ uk,t . (2) k∈K
Here uk,t is the current row of Uk,t after temporal smoothing, embedded in the robot’s action coordinates, and C applies the command constraints. The mask confines RL to the action dimensions relevant to a bottleneck, such as gripper openness or end-effector translation, while every other dimension follows the base policy. Untrained actors output a zero residual, and exploration adds Gaussian noise to the normalized residual before clipping. The residual learning formulation [5], [7] requires only nominal action chunks and an observation representation, so it is agnostic to the base policy’s architecture. b) TD3+BC learner: Each bottleneck owns a residual actor, twin critics, and a replay buffer Dk whose transitions store the RL state s = (z, p), the nominal chunk Ā, the executed residual chunk U , the chunk’s per-step rewards rτ , whose sum over an attempt equals the local reward
rk , and the terminal flag d; we drop the bottleneck index below. We train with chunk-level TD3 [30] and behaviorcloning regularization [31]. The critics Qj (s, U ) estimate the value of a residual chunk in a state and are trained by temporal-difference learning over the C-step chunk. With bars denoting target networks, Gaussian target noise ϵ, the coordinate mask M , and a minibatch B ⊂ Dk , U ′ = M ⊙ clip fθ̄ (s′ , Ā′ ) + ϵ, −1, 1 , y= LQ =
C−1 X τ =0 2 X
γ τ rτ + (1 − d)γ C min Qψ̄j (s′ , U ′ ), j=1,2
(3)
EB (Qψj (s, U ) − y)2 .
j=1
The actor maximizes the critic’s value while staying close to corrections that worked: it clones the executed residuals of successful episodes B + and, optionally, penalizes nonzero residuals of failed episodes B − , which anchors failed atb denoting the tempts back to the nominal policy. With U actor’s masked prediction, b )] Lπ = − λQ EB [Qψ1 (s, U b − U ∥2 ⟩B+ + λ− ⟨∥U b ∥2 ⟩B− , + λ+ ⟨∥U
(4)
where the weights λ balance value improvement, imitation, and anchoring. Because conditioning on Ā and regularizing toward executed residuals can let the actor copy rather than improve, we apply reference dropout, zeroing Ā for a random subset of each batch [7]. C. Agentic Scaffolding A human identifies the bottlenecks K from real-robot evaluations of π0 and writes a contract for each: its entry set Ek , the correctable action coordinates and their bounds, its outcome criterion ϕk , and a motion budget, with a handful of subtask demonstrations. Each bottleneck is then supervised by its own local reward rk = ϕk ; the complete-task outcome R serves only for evaluation. Turning a contract into repeatable practice requires attempts that start at reachable states, end with a trustworthy label, and can be repeated. PARTS implements these functions as three executable programs (a policy selector, a success verifier, and a reset policy) that run alongside the control loop. Coding agents (Claude Fable 5 and GPT-6 Astra) author these programs using the contracts, subtask demonstrations, robot interfaces, and recorded rollouts [32], [33], [34]. Humans review the programs’ decisions on labeled episodes, and the agents revise the code accordingly. Once deployed, the programs run as ordinary code and do not generate motor commands. Table I summarizes the human involvement that remains during RL data collection, compared with the baselines. Their perceptual evidence comes from promptable segmentation with SAM3 [35], which supplies object masks and locations for geometric predicates, and from a VLM (Gemini-3.7-flash) that answers asynchronous queries about specified object states, as language models have been used to resolve partially observable task state [36].
Earbud insertion | bimanual YAM
LEGO sorting | bimanual YAM
open & hold case → insert first earbud → insert second earbud → close case & place
right-arm grasp / left-arm grasp → transport → drop into color bin · repeated for ten bricks
Cable unplug & plug | single-arm Franka FR3
grasp connector → unplug → move to destination router → align & insert → release
Fig. 3. Tasks on bimanual YAM and single-arm Franka FR3. Colored borders mark states produced by a targeted residual policy, in the color of its subtask name in the row header. Top: earbud insertion, where inserting earbuds requires high precision. Middle: LEGO sorting, where the grasp of each small brick is the recurring bottleneck because the brick size is out-of-distribution for the base policy. Bottom: cable insertion requires precise manipulation.
Policy Selector. The selector outputs gt from images, proprioception, and motion predicates under the task structure of the contracts, either an ordered sequence of residuals or a choice among those whose entry conditions hold, in both cases allowing repeated activation when an entry recurs. It activates residual k only when the observations support that bottleneck’s entry conditions and otherwise leaves the base policy in control. Success Verifier. The verifier produces the local reward. Once human-defined event gates are satisfied, such as the gripper opening after an insertion, it tests the contract’s postcondition, requiring persistence across observations when a stable outcome matters, and returns one for success and zero for failure. An uncertain verdict requests a human label, and a human can override an automatic label. During evaluation, verified subtask success triggers a handoff from the residual policy to the base policy, allowing full-task execution to continue and subsequent residual policies to be activated as needed. Auto-Reset Policy. After each terminal label, the reset policy decides whether to retry or reset: a failed attempt retries directly if the subtask’s starting conditions still hold, whereas a successful attempt that changed them requires a reset, which the robot performs when feasible, for example by lowering and releasing a grasped object, and a human performs otherwise.
starting conditions rather than to the task’s initial state, so consecutive attempts see nearly the same scene configuration; online updates can then overfit to that configuration and fail under the state variation that the preceding nominal behavior produces at evaluation. PARTS therefore periodically retrains each residual policy on a success-reweighted copy of its replay and redeploys it [37]. The curated dataset keeps every successful episode and a uniformly sampled fraction ρ of the failed ones, ek = D+ ∪ Sampleρ (D− ), which raises the share of D k k rewarded experience across all restagings while retaining some failures as negatives. Fresh actor and critic networks are trained on this fixed dataset, and an operator selects a candidate checkpoint. The selected checkpoint and curated replay buffer initialize a new online run, where updates resume as new rollouts are collected. Retraining thus changes the policy used for subsequent data collection, rather than serving solely as a final policy extraction step.
D. Success-Reweighted Retraining and Redeployment
A. Tasks and Setup
Online updates provide a weak learning signal for a bottleneck whose base success is low: most attempts fail, so rewarded transitions are rare in the replay buffer. A second problem arises from how local practice is reset. To save reset time, attempts are often restaged to the subtask’s
We evaluate PARTS on three long-horizon tasks across two real-robot platforms (Fig. 3). • Earbud insertion (bimanual YAM). The robot opens a charging case, holds it in the left gripper, inserts two earbuds with the right arm, and closes the case. The
IV. Real-World Experiments Our experiments address three questions. Q1. Does targeted subtask RL improve a pretrained VLA on complete long-horizon tasks? Q2. Under a matched robot-rollout budget, how does PARTS compare with RL fine-tuning methods that train on the full task? Q3. Does success-reweighted retraining matter?
task is hard because it is a dependent chain: a poor holding pose propagates to both insertions; each slot’s clearance is small relative to the positioning error of bimanual coordination; and the gripper partly occludes the slot, so a seated earbud and one resting on the case look alike until release, which complicates both control and outcome assessment. The base policy reaches the second insertion only after earlier successes, so fulltask rollouts rarely practice it. The bottlenecks are case preparation (RL0), first insertion (RL1), and second insertion (RL2). • LEGO sorting (bimanual YAM). The robot sorts ten bricks into three color bins within 150 s. The SFT demonstrations come from the public ABC-130k dataset [38], whose bricks are larger than ours, so the nominal gripper often closes too little to retain a brick even from a well-placed approach. Since the grasp recurs ten times per episode, a modest per-grasp failure rate consumes the time budget. The bottleneck is the grasp by either arm, and the base policy supplies approach, transport, and placement. • Cable unplug-and-plug (single-arm Franka FR3). The robot unplugs a cable from a source router and inserts it into a destination router. Insertion demands alignment within the port tolerance followed by a decisive contact motion, and small offsets catch the connector on the housing. Insertion is evaluated in the states produced by the preceding extraction. The bottleneck is alignment and insertion. These tasks involve grasping, repositioning, object handovers, and alignment, and span 20–120 s (approximately 600–3,600 control steps at 30 Hz). Bottleneck subtasks typically last 3–15 s (90–450 control steps). Base policies and contracts. Humans identify the bottlenecks from real-robot evaluations of the base policies, provide a few subtask demonstrations to delimit them, and specify each contract. Reward and reset protocols depend on how hard the outcome is to recognize and how hard the scene is to restore (Table I). For the LEGO and cable tasks, a lifted brick or a seated connector is visually unambiguous, so the automatic success verifier supplies every reward. For earbuds, a seated earbud and one resting on the case look alike to a VLM without task-specific post-training, so a human checks the verifier’s labels. Resets are automated when the robot can restore the initial conditions, as in the LEGO task, where it lowers and releases a grasped brick onto the table. Human assistance is required when a reset exceeds the robot’s hardware capabilities, such as extracting a seated earbud from its case, or when the scene cannot be restored autonomously, such as after an earbud falls to the floor. Evaluation. We evaluate each method on 20 complete episodes per task, starting from the full-task’s initial states. To capture partial completion and improvements that binary full-task success can obscure in long-horizon tasks, we additionally report a normalized progress score. LEGO progress is the mean fraction of the ten bricks sorted before the
deadline; because the task repeats one pick-and-place ten times, we report only this score for LEGO, as a binary outcome would reduce to whether a single brick is placed. Earbud progress assigns one third per completed stage, with success requiring all three stages. Cable progress assigns one half each for unplugging and plugging, with success requiring both. We also report per-stage success within these same episodes as the fraction in which each bottleneck’s outcome criterion is satisfied, independently of other stages’ outcomes. For example, the second earbud may be seated correctly even if the first is misplaced inside the case. B. Baselines We compare PARTS to the frozen base policy and to three RL methods that improve a pretrained VLA from real-world experience. For fair comparison, every RL method starts from the same π0.5 -SFT base policy, trains with the same amount of robot rollout time, and receives no corrective teleoperation or DAgger-style [39] human interventions during RL rollouts. All baselines use human resets and human reward labeling; on LEGO the label is the progress score, the fraction of bricks sorted in the episode. SFT: We report the performance of the base policy after SFT on task demonstrations, before any online interaction. • DSRL [9]: DSRL learns an online RL policy in the latent noise space of the frozen VLA, steering action generation by selecting the noise fed to the VLA’s action generator. Exploration is thereby confined to actions the VLA can generate. We run DSRL on the full task. • EXPO-FT [10]: EXPO-FT learns an edit policy that modifies the VLA’s action chunks under Q-guidance while continuing to update the π0.5 backbone. Unlike PARTS, it adapts the base policy itself and operates on the full task. Its published regime additionally uses human interventions and automatic reward detectors, which we do not use. • RLT [7]: like PARTS, RLT freezes the VLA and trains a small actor-critic that refines the VLA’s reference action chunks, conditioned on the VLA’s representation. Unlike PARTS, RLT has a single RL phase entered by a human-selected VLA-to-RL handoff, after which the RL policy controls the rest of the episode with no handback to the VLA. We follow this procedure on earbuds and cable. It cannot express LEGO sorting, whose grasp bottleneck recurs for every brick and requires repeated switching between the VLA and the residual, so on LEGO we run RLT as full-task RL. RLT’s optional corrections are not used. •
Matched rollout budget. Each baseline receives the same RL robot-rollout time as PARTS on the same task, counted as elapsed robot time during RL rollouts, excluding physical resets and pauses for RL updates. DSRL and EXPO-FT spend this budget on full-task rollouts, RLT spends it from its handoff onward on earbuds and cable and on the full task on LEGO, and PARTS spends it inside the bottlenecks.
Full-task evaluation
π0.5 + SFT
RLT PARTS (ours) PARTS w/o retraining dark: binary whole-task success
Earbuds
Cable
progress + binary success
progress + binary success
LEGO progress score only 97.5 95
100
50
48.3
45
54
40
37.5 25
44.5
48
50
20
15
10
82 70
65 50
41.7
40
85
82.5
75
66.7
0
0
0
Per-stage success
DSRL EXPO-FT light: progress score
Earbuds
Earbuds
Earbuds
Cable
Cable
Open & hold case
Insert 1st
Insert 2nd
unplug
insert
100
95
90 80
100 100
100 100
95
90
80 50
50
35 25
40
40 30
70
65
60
55 45
50
30
0
20
20 0
0
Fig. 4. PARTS significantly raises success across tasks with different difficulty levels and outperforms RL fine-tuning baselines under a matched robotrollout budget. Top: full-task evaluation over 20 episodes per method; for each method the light bar is the progress score and the dark bar is binary whole-task success. Earbud progress credits one third per completed stage and success requires all three; cable progress credits one half each for unplugging and insertion and success requires both. LEGO is a repeated pick-and-place task over ten bricks, so we report only its progress score, the fraction of bricks sorted before the deadline, which is more informative than a binary outcome over the whole task. The ablation is reported for binary success only. Bottom: per-stage success measured within the same episodes, for the three earbud bottlenecks and the two cable stages. Hatched bars are PARTS without successreweighted retraining.
The comparison therefore controls robot data-collection time, while wall-clock and compute costs can differ. Implementation details. We use official DSRL and EXPO-FT code on Franka and reimplement both for YAM, changing only the robot interfaces. RLT is reproduced from the paper, as no official code is available. C. Experimental Results Q1: PARTS improves over the base VLA policy. Fig. 4 shows that PARTS improves performance in full-task evaluations across all three long-horizon tasks and both robot platforms, including settings with weak base policies. On the relatively easy cable task, insertion success nearly doubles after only 29 min of RL rollouts. This training efficiency stems from our system design: because the base policy already unplugs reliably, RL is concentrated on insertion, directing online interaction toward the subtask that limits success. On the most challenging earbud task, which requires precise bimanual manipulation, PARTS achieves four times the base policy’s full-task success rate. Per-stage results show substantial gains in both insertions, whose success rates improve by 20-35 percentage points over the base policy. For LEGO, the longest task at 120 s, correcting gripper closure alone increases progress from 54% to 82% with only 17 min of RL rollouts. The same two grasp residual policies are reused across all ten bricks, allowing each learned correction to address repeated occurrences of its bottleneck throughout the task. Q2: PARTS outperforms RL fine-tuning baselines under the same rollout budget. Given the same robot
rollout time, PARTS consistently outperforms all baselines across tasks of varying difficulty. It concentrates learning and repeated practice on bottleneck subtasks, whereas the baselines spend part of their rollout budget re-executing subtasks that the base policy already handles reliably. Moreover, the baselines use sparse terminal rewards that provide outcome feedback only at the end of each 20–120 s episode. Low base-policy success rates leave the replay buffer with few successful trajectories, making critic learning difficult under sparse positive feedback and long credit-assignment horizons. In contrast, PARTS uses local rewards over 3– 15 s subtask windows with initial success rates above 25%, providing more frequent positive feedback and shortening the credit-assignment horizon. EXPO-FT ultimately underperforms the base policy, occasionally degrading previously reliable behaviors such as cable unplugging. This suggests that backbone updates driven by sparse full-task rewards can disrupt reliable behaviors in non-bottleneck subtasks. RLT and DSRL are stronger baselines, but neither significantly improves on the SFT policy on average under limited realworld rollout budgets. DSRL constrains exploration to behaviors generated by the frozen VLA, promoting stable learning but potentially limiting further improvements. RLT focuses on critical-phase adaptation, but still underperforms PARTS. Its reported system targets a single critical phase following a VLA-to-RL handoff, whereas PARTS coordinates repeated transitions between the base policy and specialized residual policies to address multiple bottlenecks in the LEGO and earbud tasks. Q3: success-reweighted retraining matters. We ablate
retraining on the earbud task, where it was applied to RL1 and RL2 only. Both start from low base success, so their local rewards are sparse, and both are restaged to RL0’s success state between attempts. Restaging removes the variation that full-task execution introduces into the relative pose of the held case and the inserting gripper, so consecutive attempts see similar configurations, and a residual updated online alone can settle into corrections that fit that configuration but not the states produced by the preceding nominal behavior at evaluation. RL0 was not retrained: its base success is already high and its reset is the task’s initial condition, so its attempts already cover the evaluation distribution. Removing successreweighted retraining reduces both subtask and full-task success, as shown in the hatched bars of Fig. 4. Therefore, retraining fresh networks on the curated episode set can help escape local optima, providing a key mechanism underlying PARTS’s success. V. CONCLUSION We presented PARTS, a real-world subtask RL framework that adapts a pretrained robot policy to a long-horizon task by concentrating practice on the few subtasks where the policy fails, with minimal human intervention during training. Across bimanual and single-arm platforms, PARTS substantially improves full-task success over the base policy and outperforms existing RL fine-tuning methods under matched robot rollout budgets, while requiring fewer recurring forms of human involvement. Future work will pursue full autonomy through agents that identify bottlenecks from policy failures, generate dense subtask rewards, and construct recovery phases. These capabilities could automate setup decisions, accelerate learning when successful outcomes are rare, and enable the robot to practice recovery behaviors after subtask failures, reducing reliance on human resets. References [1] K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. R. Equi, C. Finn, N. Fusai, M. Y. Galliker, D. Ghosh, L. Groom, K. Hausman, b. ichter, S. Jakubczak, T. Jones, L. Ke, D. LeBlanc, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, A. Z. Ren, L. X. Shi, L. Smith, J. T. Springenberg, K. Stachowicz, J. Tanner, Q. Vuong, H. Walke, A. Walling, H. Wang, L. Yu, and U. Zhilinsky, “π0.5 : a vision-language-action model with open-world generalization,” in Proceedings of The 9th Conference on Robot Learning, ser. Proceedings of Machine Learning Research, J. Lim, S. Song, and H.-W. Park, Eds., vol. 305. PMLR, 27–30 Sep 2025, pp. 17–40. [Online]. Available: https://proceedings.mlr.press/v305/black25a.html [2] Physical Intelligence, B. Ai, A. Amin, R. Aniceto, A. Balakrishna, et al., “π0.7 : A steerable generalist robotic foundation model with emergent capabilities,” arXiv preprint arXiv:2604.15483, 2026. [Online]. Available: https://arxiv.org/abs/2604.15483 [3] S. Ye, Y. Ge, K. Zheng, S. Gao, S. Yu, et al., “World action models are zero-shot policies,” arXiv preprint arXiv:2602.15922, 2026. [Online]. Available: https://arxiv.org/abs/2602.15922 [4] H. Li, L. Sun, Y. Hu, D. Ta, J. Barry, G. Konidaris, and J. Fu, “NovaFlow: Zero-shot manipulation via actionable flow from generated videos,” in IEEE International Conference on Robotics and Automation (ICRA), 2026. [5] L. Ankile, Z. Jiang, R. Duan, G. Shi, P. Abbeel, and A. Nagabandi, “Residual off-policy rl for finetuning behavior cloning policies,” arXiv preprint arXiv:2509.19301, 2025.
[6] K. Fang, P. Yin, A. Nair, and S. Levine, “Planning to practice: Efficient online fine-tuning by composing goals in latent space,” in IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2022. [Online]. Available: https://arxiv.org/abs/2205.08129 [7] C. Xu, J. T. Springenberg, M. Equi, A. Amin, A. Esmail, S. Levine, and L. Ke, “Rl token: Bootstrapping online rl with vision-languageaction models,” arXiv preprint arXiv:2604.23073, 2026. [8] H. Zheng, Y. Yang, K. Ma, S. Xu, T. Xie, G. Li, X. Wang, Y. Ma, S. Liu, Y. Mao, et al., “Torl-vla: Tactile guided online reinforcement learning for contact-rich manipulation,” arXiv preprint arXiv:2606.09337, 2026. [9] A. Wagenmaker, M. Nakamoto, Y. Zhang, S. Park, W. Yagoub, A. Nagabandi, A. Gupta, and S. Levine, “Steering your diffusion policy with latent space reinforcement learning,” arXiv preprint arXiv:2506.15799, 2025. [10] P. Dong, K.-H. Hung, T. Gao, D. Sadigh, and C. Finn, “Expoft: Sample-efficient reinforcement learning finetuning for visionlanguage-action models,” arXiv preprint arXiv:2605.25477, 2026. [11] L. Shi, S. Chen, F. Gao, Y. Chen, K. Chen, T. Zhang, H. Zang, J. Zhou, W. Zhang, C. Yu, et al., “Beyond imitation: Reinforcement learning-based sim-real co-training for vla models,” arXiv preprint arXiv:2602.12628, 2026. [12] H. Zang, M. Wei, S. Xu, Y. Wu, Z. Guo, Y. Wang, H. Lin, L. Shi, Y. Xie, Z. Xu, et al., “Rlinf-vla: A unified and efficient framework for vla+ rl training,” arXiv preprint arXiv:2510.06710, 2025. [13] K. Lei, H. Li, D. Yu, Z. Wei, L. Guo, Z. Jiang, Z. Wang, S. Liang, and H. Xu, “Performant robotic manipulation with real-world reinforcement learning,” Science Robotics, vol. 11, no. 116, p. eaed6267, 2026. [14] J. Luo, Z. Hu, C. Xu, S. Gadipudi, A. Sharma, R. Ahmad, S. Schaal, C. Finn, A. Gupta, and S. Levine, “Serl: A software suite for sample-efficient robotic reinforcement learning,” in Towards Generalist Robots: Learning Paradigms for Scalable Skill Acquisition@ CoRL2023, 2024. [15] T. Johannink, S. Bahl, A. Nair, J. Luo, A. Kumar, M. Loskyll, J. A. Ojea, E. Solowjow, and S. Levine, “Residual reinforcement learning for robot control,” in 2019 international conference on robotics and automation (ICRA). IEEE, 2019, pp. 6023–6029. [16] O. Biza, T. Weng, L. Sun, K. Schmeckpeper, T. Kelestemur, Y. J. Ma, R. Platt, J.-W. van de Meent, and L. L. S. Wong, “On-robot reinforcement learning with goal-contrastive rewards,” in IEEE International Conference on Robotics and Automation (ICRA), 2025. [17] J. Luo, C. Xu, J. Wu, and S. Levine, “Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning,” Science Robotics, vol. 10, no. 105, p. eads5033, 2025. [18] Y. Zhu, P. Stone, and Y. Zhu, “Bottom-up skill discovery from unsegmented demonstrations for long-horizon robot manipulation,” IEEE Robotics and Automation Letters, vol. 7, no. 2, pp. 4126–4133, 2022. [19] L. X. Shi, J. J. Lim, and Y. Lee, “Skill-based model-based reinforcement learning,” arXiv preprint arXiv:2207.07560, 2022. [20] L. Sun, H. Zhang, W. Xu, and M. Tomizuka, “PaCo: Parametercompositional multi-task reinforcement learning,” in Advances in Neural Information Processing Systems (NeurIPS), 2022. [21] M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, et al., “Openvla: An open-source vision-language-action model,” arXiv preprint arXiv:2406.09246, 2024. [22] Y. Guo, J. Zhang, X. Chen, X. Ji, Y.-J. Wang, Y. Hu, and J. Chen, “Improving vision-language-action model with online reinforcement learning,” in 2025 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2025, pp. 15 665–15 672. [23] Y. Xu, L. Ma, G. Zuo, W. Zhang, H. Ding, and L. Zhu, “Mori: Mixture of rl and il experts for long-horizon manipulation tasks,” arXiv preprint arXiv:2604.10165, 2026. [24] Z. Chen, Y. Han, Y. Shao, H. Liu, C. Xu, X. Chen, Y. Mu, and W. Lian, “Bora: Bridging offline reinforcement learning and online residual adaptation for real-world dexterous vla models,” arXiv preprint arXiv:2605.30226, 2026. [25] O. Hausdörfer, L. Schwarz, G. Marko, C. Dietz, T. Class, L. Hofer, J. Y.-J. Li, J. Hechtl, R. Römer, and A. P. Schoellig, “Data and learning where it matters for contact-rich manipulation,” arXiv preprint arXiv:2607.15982, 2026. [26] C. Celemin, G. Maeda, J. Ruiz-del Solar, J. Peters, and J. Kober, “Reinforcement learning of motor skills using policy search and human
corrective advice,” The International Journal of Robotics Research, vol. 38, no. 14, pp. 1560–1580, 2019. [27] M. B. I. Pamies, M. T. Villasevil, Z. Wang, S. Desai, P. Agrawal, and A. Gupta, “Autonomous robotic reinforcement learning with asynchronous human feedback,” in 7th Annual Conference on Robot Learning, 2023. [28] H. Deng, Y. Gao, Y. Lin, H. Liu, Z. Wu, and Z. Wang, “Uniintervene: Agentic intervention for efficient real-world reinforcement learning,” arXiv preprint arXiv:2606.12372, 2026. [29] K. Hu, H. Shi, Y. He, W. Wang, C. K. Liu, and S. Song, “Robot trains robot: Automatic real-world policy adaptation and learning for humanoids,” arXiv preprint arXiv:2508.12252, 2025. [30] S. Fujimoto, H. Hoof, and D. Meger, “Addressing function approximation error in actor-critic methods,” in International conference on machine learning. Pmlr, 2018, pp. 1587–1596. [31] S. Fujimoto and S. S. Gu, “A minimalist approach to offline reinforcement learning,” Advances in neural information processing systems, vol. 34, pp. 20 132–20 145, 2021. [32] J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. Florence, and A. Zeng, “Code as policies: Language model programs for embodied control,” in 2023 IEEE International conference on robotics and automation (ICRA). IEEE, 2023, pp. 9493–9500. [33] K. Elmaaroufi, J. Svegliato, S. Kalade, G. Schelle, S. A. Seshia, and M. Zaharia, “Rho: Your coding agent is secretly a roboticist,” arXiv preprint arXiv:2606.16458, 2026. [34] M. Fu, J. Yu, K. El-Refai, E. Kou, H. Xue, H. Huang, W. Xiao, G. Wang, F.-F. Li, G. Shi, et al., “Cap-x: A framework for benchmarking and improving coding agents for robot manipulation,” arXiv preprint arXiv:2603.22435, 2026. [35] N. Carion, L. Gustafson, Y.-T. Hu, S. Debnath, R. Hu, D. Suris CollVinent, C. Ryali, K. V. Alwala, H. Khedr, A. Huang, et al., “Sam 3: Segment anything with concepts,” in International conference on learning representations, vol. 2026, 2026, pp. 138 846–138 923. [36] L. Sun, D. K. Jha, C. Hori, S. Jain, R. Corcodel, X. Zhu, M. Tomizuka, and D. Romeres, “Interactive planning using large language models for partially observable robotics tasks,” in IEEE International Conference on Robotics and Automation (ICRA), 2024. [37] M. S. Mark, A. Sharma, F. Tajwar, R. Rafailov, S. Levine, and C. Finn, “Offline retraining for online rl: Decoupled policy learning to mitigate exploration bias,” arXiv preprint arXiv:2310.08558, 2023. [38] A. Allshire, H. G. Singh, R. Singh, A. Rashid, H. Choi, D. McAllister, J. Yu, Y. Chen, H. Huang, P. Abbeel, X. Chen, R. Duan, P. Isola, J. Malik, F. Shentu, G. Shi, P. Wu, and A. Kanazawa, “Scalable behavior cloning with open data, training, and evaluation,” 2026. [Online]. Available: https://arxiv.org/abs/2606.27375 [39] M. Kelly, C. Sidrane, K. Driggs-Campbell, and M. J. Kochenderfer, “Hg-dagger: Interactive imitation learning with human experts,” in 2019 International Conference on Robotics and Automation (ICRA). IEEE, 2019, pp. 8077–8083.