Conceptio › Archive › arXiv CS
arXiv CSopen access

SynthDemo-RL: Breaking the Zero-Reward Barrier in VLA Adaptation with LLM-Guided Synthetic Demonstrations

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

SynthDemo-RL: Breaking the Zero-Reward Barrier in VLA Adaptation with LLM-Guided Synthetic Demonstrations

arXiv:2609.21650v1 [cs.RO] 18 Sep 2026

Hiroaki Kingetsu1 , Hiroaki Kurihara1 , Kaoru Yokoo1 , Kenji Fukumizu2,1 , and Manohar Kaul1 Abstract— Fine-tuning Vision-Language-Action (VLA) models commonly relies on human teleoperation demonstrations, while reinforcement learning (RL) with sparse binary rewards faces an exploration challenge when successful trajectories are rarely sampled. We propose SynthDemo-RL, a teacher-student framework in which an automated teacher converts simulatorprivileged state into successful manipulation trajectories, a VLA student is distilled from them by supervised fine-tuning (SFT), and PPO with binary task-success rewards refines the student. We study reward coverage, the fraction of tasks for which at least one success is observed under the fixed evaluation protocol, as a complement to the average success rate. On LIBERO-PRO, a public benchmark of perturbed LIBERO tasks for which no demonstrations exist, 27 of 57 scored tasks are at exactly 0% success for a π0.5 policy fine-tuned on the original LIBERO tasks. Direct PPO from this policy, under the same PPO recipe and the same RL compute as SynthDemo-RL’s refinement stage, rescues 10 of these 27 tasks and leaves 17 at 0%. SynthDemoRL, with 50 synthesized trajectories per task and no new human demonstrations, rescues all 27 and reaches average success rates of 97.8% and 97.1% on the Position and Task axes of LIBERO-PRO, respectively. On standard LIBERO, the same pipeline reaches 96.0% with no human demonstrations, within 1.7 points of π0.5 trained on 50 human demonstrations per task. We further validate the pipeline on RoboTwin 2.0 and verify that trajectories from a policy trained in a MuJoCo twin execute open-loop on a physical robot.

I. I NTRODUCTION Vision-Language-Action (VLA) models [1] have shown strong generalization in robotic manipulation through internet-scale vision-language pretraining. Adapting them to specific downstream tasks, however, commonly relies on imitation learning through supervised fine-tuning (SFT) on human teleoperation demonstrations, which fundamentally limits the scalability of VLA-based systems across tasks and environments. Recent work shows that reinforcement learning (RL) can lift VLA policies far above their SFT baseline, even from a single demonstration [2]–[4], but it still assumes humancollected demonstrations. When successful trajectories are rarely sampled, RL with sparse binary rewards may struggle to discover rewarding behavior. This zero-reward regime is our starting point. Demonstrations are the classical way out [5], and unlike reward shaping they need no per-task reward design [6], [7]. If an automated procedure can solve a task even occasionally, its successful trajectories can be distilled into the policy as an initialization for sparse-reward RL. We therefore distinguish 1 Fujitsu Limited, Kawasaki, Japan. 2 The Institute of Statistical Mathematics, Tokyo, Japan.

Correspondence: [email protected]

average success rate from reward coverage, the fraction of tasks with at least one observed success. Average success measures how well a policy performs; coverage measures on how many tasks RL has anything to build on. A high average can hide tasks with no observed success, and those are exactly where sparse-reward RL fails. We propose SynthDemo-RL, a three-stage framework for training VLA manipulation policies without new human demonstrations (Fig. 1). An automated teacher with simulator-privileged state collects successful trajectories. A pretrained grasp predictor proposes grasps and a scripted controller executes a waypoint sequence (approach, grasp, lift, place). When an attempt fails, an LLM inspects the failure log and images. It then revises the waypoint sequence and execution parameters such as contact depth and lateral retreat. The teacher prioritizes coverage over optimality and keeps trying until every task yields successful trajectories. The teacher is too slow to run online, so SFT distills its trajectories into a VLA student. PPO with binary tasksuccess rewards then refines the student past its teacher. We demonstrate that synthetic demonstrations can initialize sparse-reward RL at VLA scale without collecting new human demonstrations for the target tasks. SFT achieves nonzero success on every task, and RL substantially improves the resulting policies. In Sec. IV-C we study human-demonstration-free adaptation on LIBERO-PRO [8], a public benchmark that systematically perturbs LIBERO tasks (e.g., moving initial object placements or redefining task goals); such perturbations severely degrade policies fine-tuned on human demonstrations, and no demonstrations exist for the perturbed variants. Our contributions are: • SynthDemo-RL, a pipeline that combines an automated teacher with VLA distillation and sparse-reward PPO. The teacher converts simulator-privileged state into successful manipulation trajectories with no human demonstrations or task-specific training. • Reward coverage, the fraction of tasks with at least one observed success, as a diagnostic complementary to average success. On perturbed LIBERO-PRO, distilling synthesized trajectories rescues all 27 tasks with no observed success before adaptation, while the PPOmatched sparse-reward control leaves most of them at 0%. Among tasks with nonzero SFT success, RL lifts even weak initializations to high final success. • An analysis of synthetic SFT limitations. Synthetic trajectories differ systematically from human trajectories in motion statistics and frozen-base flow-matching

Fig. 1: SynthDemo-RL: VLA task adaptation without new human demonstrations. Human demonstrations exist only for the original task (green marker). LIBERO-PRO asks for the other bowl (red), and a VLA fine-tuned on the original tasks has no observed successes on 27 of the 57 perturbed tasks under the evaluation protocol. An automated teacher solves the perturbed task from simulator-privileged state, SFT distills its successful trajectories into the VLA student, and PPO refines the student past its teacher. Bottom: per-task success (Table III). Direct PPO leaves 20 of the 57 tasks with no observed success after adaptation (17 unrescued of the 27, plus 3 driven to 0%), while ours leaves none.

loss. SFT on regenerated demonstrations with reduced measured differences remains well below humandemonstration SFT. II. R ELATED W ORK

model. LLMs have served as automated tuning agents for rewards and task specifications [6], [7]. Our LLM tuner instead revises waypoint sequences and execution parameters inside the teacher to raise data-collection success.

A. Reinforcement Learning for VLA Fine-Tuning

C. Teacher-Student Policy Distillation

SimpleVLA-RL [2], VLA-RL [3], RLinf-VLA [4], and VLA-RFT [9] refine VLA policies with online RL. They differ in policy interface, infrastructure, and whether rollouts use a simulator or a learned world model, but all assume human-collected SFT data for initialization. The binary tasksuccess reward is the common choice in this line [2], [4], [10]. Our contribution is to show that synthesized demonstrations can initialize sparse-reward RL at VLA scale on target task distributions without human demonstrations.

Distilling non-learned or privileged teachers into deployable students is well established: Neural MP [17] distills a sampling-based motion planner into a policy that surpasses its teacher, and privileged-teacher training is standard in locomotion [18]. RPD [19] distills a VLA teacher into an RL student but queries the teacher online at every step, and the teacher itself was trained on demonstrations. Our teacher requires no task-specific training, and distillation is offline: the teacher generates data once and is discarded.

B. Automated Data Generation and Learning without Human Demonstrations

III. M ETHOD : S YNTH D EMO -RL

MimicGen [11] and RoboCasa [12] expand demonstration datasets through retargeting but require source demonstrations of the target behavior. Our teacher generates target-task demonstrations without such sources. Ha et al. [13] generate task plans with an LLM, solve them with a samplingbased planner under success verification, and distill into a language-conditioned policy. SynthDemo-RL is closest to this line in structure. We extend this synthesis-and-distillation approach to VLA adaptation, using the distilled policy as the initialization for sparse-reward RL. Among approaches that aim to remove demonstrations entirely, ReWiND [14] still needs initial demonstrations for its reward model. SelfImproving VLA [15] needs an initial policy that already succeeds, and GigaBrain [16] needs a proprietary world

A. Overview The goal is to learn a VLA policy πθ that maximizes task success rate in a simulation environment without collecting new human demonstrations for the target tasks. The teacher generates D = {(τi , ri )}N i=1 , where N is the number of recorded trials, τi is an observation-action trajectory, and ri ∈ {0, 1} is the outcome of the simulator’s success checker. The successful trajectories D+ = {τi | ri = 1} are used to train πθSFT . PPO with binary task-success rewards then refines it into πθRL . B. Automated Teacher Pipeline The teacher takes a language instruction and access to the simulator state (ground-truth object poses, metric depth,

instance segmentation, camera parameters, and the endeffector pose). It executes end-effector actions and records the resulting observation–action trajectory τ . It is a scripted executor around two learned components, a pretrained grasp generator and an LLM tuner. 1) Goal Specification and Object Grounding: The task definition identifies the manipulated object and its goal region, and the simulator provides their ground-truth poses. The object’s instance mask and metric depth are backprojected through the camera parameters into a 3D point cloud, which is the geometric input to grasping. 2) Grasp and Placement: A neural grasp predictor (GraspGen [20]) proposes 6-DoF grasps on the object point cloud. Candidates are filtered to the object’s axis-aligned bounding box and post-processed to use a top-down approach by default with a tunable contact depth (Sec. III-B.4). The placement pose is computed from the goal object’s groundtruth top-surface position with a clearance offset. 3) Closed-Loop Execution: The LLM tuner (GPT5.5 [21]) configures the waypoint sequence within each manipulation primitive, including its stages and their order, as well as execution parameters. A pick-and-place plan may include pre-grasp, grasp, lift, pre-place, place, and retreat, while a drawer-opening plan may use a contact-sweep-retreat sequence. A proportional servo tracks these waypoints with per-step position correction, which matters during contactrich grasp phases where open-loop execution diverges. 4) LLM-Guided Tuning: After failed attempts, the LLM tuner revises the waypoint sequence and execution parameters, including contact depth, approach angle, lift and lateral retreat, and placement offset. It receives the failure stage and reason from the execution log, observation images of the failed attempt, a success library of configurations that worked on other tasks (ranked by relevance), and a failure library of configurations already tried on this task. From these it proposes a revised configuration for the next attempt. The success library transfers configurations across geometrically similar tasks; the failure library prevents revisiting dead ends. 5) Data Collection and Filtering: During execution the teacher records multi-camera observations (third-person and wrist), 7-DoF actions, and proprioceptive state. Postexecution filtering retains only successful episodes and removes no-op transitions, and the result is exported to LeRobot format. C. Offline Policy Distillation The teacher needs ∼1.7 minutes per episode (measured on LIBERO-Object) for grasp inference and closed-loop planning, incompatible with 10–20 Hz control, so we distill its behavior into a VLA policy that predicts action chunks directly from observations. The student is π0.5 [22], a flowmatching VLA that predicts continuous action chunks from third-person and wrist images, proprioceptive state, and the instruction. We fine-tune on the successful teacher trajectories D+ , starting from the released π0.5 base checkpoint for standard LIBERO and, for LIBERO-PRO adaptation, from π0.5 -LIBERO, our own human-demonstration SFT of that

base (row (d) of Table I). Both settings train one policy per suite (on LIBERO-PRO, one per suite and perturbation axis, six conditions in total) on 50 successful episodes per task. LIBERO-PRO adaptation uses only trajectories synthesized on each perturbed task, with no original-LIBERO replay. Training follows the standard flow-matching objective and the openpi recipe: 30,000 steps, batch size 32, peak learning rate 5 × 10−5 , 8 MI300X GPUs. Sec. IV-F examines the SFT gap and the effect of modifying the teacher’s motion characteristics. D. RL Refinement We apply PPO [23] through the RLinf framework [4] to the SFT-initialized student, with the binary task-success reward and no shaping. This keeps the optimization target identical to the evaluation metric and needs no per-task reward engineering. To obtain a tractable likelihood for PPO, we use Flow-SDE [10]: Gaussian noise at each denoising step yields transitions with closed-form densities, whose logdensities sum to the sampled denoising-path log-probability used in the PPO ratio. Each MDP step executes one action chunk, the sparse reward arrives at episode termination, and GAE [24] operates over chunk steps using a value head on the π0.5 backbone, with no separate critic model. Each iteration collects eight rounds of rollouts, each running 64 parallel environments for 240 chunk steps. We use three denoising steps with exploration noise 0.5 and five-action chunks of 7-DoF actions. Each PPO update runs one epoch over a global batch of 2048 chunk steps (micro batch 128) with γ = 0.99, GAE λ = 0.95, clip 0.2, dual clip 3.0, value clip 0.2, and no KL penalty or entropy bonus. The actor and value head use learning rates of 5 × 10−6 and 1 × 10−4 with gradient clip 1.0. RL serves two purposes: it recovers from the biases of the teacher’s scripted execution by visiting states the teacher never produced, and it lets the student surpass the teacher’s own reliability, as also observed for distilled planners [17], [19]. IV. E XPERIMENTS A. Experimental Setup 1) Benchmarks: LIBERO [25] is a single-arm manipulation benchmark (Franka Panda, MuJoCo). We evaluate on the three suites targeting single-stage manipulation: LIBEROSpatial (10 tasks, spatial relations), LIBERO-Object (10 tasks, object identity), and LIBERO-Goal (9 tasks, goal variation).1 LIBERO-PRO [8] applies systematic perturbations to these suites; we use the two axes under which fine-tuned VLAs degrade most: Position, which reassigns objects to alternative placement regions while keeping goals and instructions unchanged, and Task, which rewrites goal predicates and instructions while keeping the object set (e.g., 1 We restrict evaluation to single-stage manipulation, which excludes LIBERO-Long and LIBERO-Goal Task 3. Task 3 chains drawer opening with a placement; the teacher supports both primitives individually but does not implement their sequencing (Sec. V). It is excluded from all of our runs, in training and scoring.

“open the middle drawer” → “open the bottom drawer”).2 We use the officially distributed perturbed task definitions and evaluation initial states, frozen with content hashes before any experiment. Our Position-Goal and Task-Goal evaluations score 9 and 8 tasks, respectively, and each task is evaluated over 50 trials. RoboTwin 2.0 [26] is a dualarm benchmark in SAPIEN. We use its single-arm Agilex Piper configuration on 4 tasks as a cross-simulator, crossembodiment test (bimanual coordination is out of scope). 2) Methods Compared: We compare (a) SynthDemoTeacher (the automated teacher of Sec. III-B), (b) SynthDemo-SFT (teacher data plus SFT of the π0.5 student), and (c) SynthDemo-RL (ours: teacher data plus SFT plus RL). Methods (a)–(c) use zero human demonstrations. On LIBERO, we further compare (d) π0.5 (human) (the same student fine-tuned by us on 50 human teleoperation demonstrations per task, one policy per suite) and (d’) π0.5 (human) + PPO (row (d) refined under the same RL recipe as (c)). Rows (d, d’) are run by us under our protocol. Row (d) is trained per suite from the released π0.5 base on the 29 scored tasks and is distinct from the publicly released π0.5 -LIBERO checkpoint, a single policy trained jointly on all four suites including LIBERO-Long and then evaluated suite by suite. The two are not comparable. On RoboTwin only, we quote the published (e) OpenVLA-OFT and (e’) SimpleVLA-RL results [2], which use a different policy backbone. Row (e) is SFT on the benchmark’s generated expert demonstrations, and row (e’) adds RL. 3) Evaluation Protocol and Initial-State Provenance: Success is determined by the simulator’s built-in success checker. We report per-suite averages of per-task success rates and, on LIBERO-PRO, reward coverage, the fraction of tasks with at least one observed success in the reported evaluation trials. Our LIBERO-PRO results use three SFT training seeds, and the RL stage starts from the seed whose average success rate is the median across the six suite-axis conditions. For the per-task results, the coverage-sensitivity analysis, and the correlation between initialization and final success, SynthDemo-SFT is evaluated at the checkpoints used to initialize RL. Each seed is evaluated over 50 trials per task, for which the binomial standard error is about 7 points for a task at 50% success and 2.4 points at 97%, and about 2.2 and 0.8 points for a 10-task suite mean. On standard LIBERO, synthesis uses initial states disjoint from the benchmark’s fixed array of 50 evaluation initial states per task, and the human demonstrations are likewise a disjoint sample from the same distribution (no synthesis or human-demonstration initial state coincides with an evaluation state on any of the three suites). Rows (a)–(d) therefore share initial-state provenance, and Sec. IV-B compares them at matched policy, training budget, and episode 2 We exclude Task-Goal Task 0 from all aggregates of our runs. The unchanged layout leaves a plate and a bowl in the bottom drawer’s swept path, so opening it shoves them aside (∼9 and 15 cm in all 50 evaluation states), and the evaluated rollouts satisfy the success checker despite this displacement. The perturbation thus changes the required physical behavior, not only the instruction as the Task axis intends, and we judge the instance ill-posed. The Task axis scores 28 tasks.

TABLE I: Success rates (%) on LIBERO (50 trials/task). Row (a): teacher single-attempt success during data generation. Retained demonstrations are 100% successful by construction. Rows (b, c): fixed final checkpoints under the prescribed training schedules. Rows (d, d’): our humandemonstration reference (50 demos/task). Goal is the 9-task subset of Sec. IV-A.1. Method

Spatial

Object

Goal

Avg.

Zero human demonstrations (ours) (a) SynthDemo-Teacher (1-attempt) (b) SynthDemo-SFT (c) SynthDemo-RL (ours)

79.9 55.3 96.2

98.0 69.8 98.6

43.0 45.1 93.3

73.6 56.7 96.0

Human-demonstration reference (d) π0.5 + human-demo SFT (d’) π0.5 + human-demo SFT + PPO

97.8 95.8

99.0 99.0

96.2 97.3

97.7 97.4

count. Both RL arms draw rollouts from the official initial states, as is standard in LIBERO RL work. On LIBEROPRO, training initial states for synthesis and RL come from seeded environment resets, with an automated check that rejects any state coinciding with the officially distributed set, which is reserved for evaluation. B. Standard LIBERO Table I asks whether the pipeline can reach humandemonstration-level performance without any human demonstrations; rows (b) and (d) differ only in the source of the 50 SFT episodes per task (Sec. IV-A.3). SynthDemo-RL (c) reaches 96.0% without human demonstrations, within 1.7 points of both human-demonstration references. Two observations set up the rest of the paper. First, distillation alone remains well below human-demonstration SFT at matched policy, budget, and episode count (Sec. IVF). Second, RL rather than distillation closes the gap. RL gains about 40 points over SFT and surpasses the teacher’s single-attempt success on every suite, most markedly on Goal (93.3 vs 43.0). Generating one retained demonstration costs 0.4–3.4 minutes depending on the suite. C. Adaptation to Perturbed Tasks without New Human Demonstrations Using LIBERO-PRO, we test whether a VLA fine-tuned on the original tasks can adapt to shifted task distributions without new human demonstrations. We adapt π0.5 -LIBERO, i.e. our per-suite row (d) of Table I, with trajectories synthesized directly on each perturbed task. The teacher retains 50 successful trajectories per task, and SFT uses these only (no original-LIBERO replay, an assumption tested below). As the sparse-reward control, direct PPO starts from the same policy and uses the same PPO recipe as SynthDemo-RL. The direct-PPO control and SynthDemo-RL use identical PPO iteration counts and environment interactions within each suite-axis condition, averaging 188 iterations (23.1 million chunk steps) across the six conditions. The only difference is the target-task synthesis and SFT that precede RL in SynthDemo-RL. For

Task perturbations the rewritten instruction is supplied during training and evaluation.3 Table II reports the results. Before adaptation, 27 of the 57 scored perturbed tasks are at exactly 0%. The teacher yields at least one successful trajectory on every one of the 57 tasks. On average the tuner produced 3.3 further distinct execution configurations per task beyond the initial one (SD 6.3 over the 57 tasks). Without LLM revision, rerunning each suite’s frozen initial configuration (25 attempts per task) leaves 12 of the 57 tasks with no successful trajectory and gives task-mean success of 54.4% (Position) and 47.6% (Task). Table II reports teacher success rates using each task’s final tuned configuration (65.6% and 71.0%); attempt counts vary per task because generation stops once 50 successful trajectories are retained. The 12 tasks involve articulated drawers and a stove knob, placements onto a stove, rack, or cabinet top, or flat and elongated containers. The tuned configurations rescue these tasks by switching to rim or pinch grasps and adjusting grasp depth and lateral retreat. SynthDemo-RL exceeds direct PPO by 28.7/55.9 points on Task/Position. Sparse-reward exploration does find some tasks on its own: over its RL budget the control rescues 10 of the 27 tasks with no observed success before adaptation. But 17 of the 27 still have no observed successes after control training (13 of 15 on Position, 4 of 12 on Task). SynthDemoSFT, by contrast, makes every one of the 27 nonzero before any RL. This coverage result holds across all three SFT seeds, each achieving at least one success on every one of the 57 tasks in 50 evaluation trials per task. Coverage depends on the trial budget. Had only 10 of the 50 recorded trials per task been run, the expected number of covered tasks (averaged over all 10-trial subsets) would be 26.6 for π0.5 -LIBERO, 35.4 for the control, 53.6 for SynthDemo-SFT, and 57 for SynthDemo-RL, out of 57. Even when coverage requires at least five successes in 50 trials, SynthDemo-SFT covers 53 of 57 tasks and SynthDemo-RL covers all 57. For context, Table II also quotes published LIBEROPRO results. SPARK and Pigey use LLM-based inferencetime planning or orchestration, whereas CounterAlign applies counterfactual offline RL to LIBERO demonstrations. Across the 57 tasks and both RL arms (114 task–arm pairs, where the initialization is π0.5 -LIBERO for the control and SynthDemo-SFT for ours), whether the initialization has zero or positive observed success is strongly associated with the post-RL result (point-biserial r = 0.71): zerosuccess initializations, all of which are control-arm tasks, end at 26.0% on average, against 91.7% for initializations with at least one observed success. Among SynthDemo-SFT initializations on the Position axis, a weak one (0 < SFT < 50%) ends at 97.0% and a strong one (≥ 50%) at 98.4%, 3 As released, the official evaluation feeds the pre-perturbation instruction while scoring the perturbed goal, a mismatch also reported independently in the benchmark’s issue tracker. It measures blind instruction-following, so our Task-axis numbers are not comparable to the benchmark’s released numbers; the quoted rows of Table II also feed the rewritten instruction (note ‡) and are unaffected.

a 1.4-point difference against the ∼91-point gap between zero-success and weak initializations. This correlation does not isolate the effect of coverage: task difficulty and the synthesis-plus-SFT intervention also vary, and SynthDemoRL rates are near ceiling. Table III gives the per-task view: the control also drives 3 tasks with nonzero success before adaptation to 0%, whereas SynthDemo-SFT loses none, and all tasks retain observed successes after PPO. SFT shows forgetting on previously solved target tasks. Where π0.5 -LIBERO already solved some tasks, distilling perturbation-specific data costs performance on those tasks: on Task-Spatial, five of ten tasks drop after SFT (e.g., 92 → 20, 100 → 34). RL largely closes these regressions: after RL, all five are back within 10 points of their original values (task 6: 92 → 20 → 92). On Position-Spatial, the four tasks the original policy solved at 96 to 100% end within 8 points. Replay does not explain the gain. To test originalLIBERO replay, we compared Position-axis SFT on 50 synthetic episodes per task against the same recipe with 20 original episodes per task added. Mean success was 57.7% without replay and 51.5% with, so replay did not improve target adaptation in this comparison. D. RoboTwin 2.0 To test the pipeline beyond LIBERO’s simulator and embodiment, we apply it to four single-arm tasks with the Agilex Piper arm in RoboTwin 2.0. Table IV reports success once over 100 held-out trials per task. SynthDemo-RL improves the SFT average from 55.5% to 74.3%, numerically close to SimpleVLA-RL. The gain is largest on beat hammer, where SynthDemo-SFT is weakest (32%), but the student remains below its teacher on two tasks. For place cup, RL did not improve the strong SFT policy. E. Physical Executability on a Real Robot We assess the physical executability of trajectories generated by a twin-trained SynthDemo-RL policy through openloop execution on hardware; closed-loop sim-to-real transfer is outside our scope (Sec. V). The robot is a Trossen WidowX AI stationary station (right arm, 30 Hz, overhead and wrist RGB) with a MuJoCo twin of its workspace. Using the same teacher and SFT-plus-PPO recipe, we train the π0.5 student (LoRA adapters) entirely in the twin, with no fine-tuning on real-robot data. The twin’s renders composite the simulated arm and objects over the real camera backgrounds and a 3D Gaussian Splatting reconstruction of the room. In preliminary closed-loop hardware trials the same checkpoint did not succeed, and an observation-sensitivity analysis suggests that the rendered appearance of the foreground (arm and plate) contributes to this gap. We therefore test trajectory executability with open-loop execution, which uses no real-camera observations for policy feedback. The policy runs closed-loop in the twin from a matched initial state, and the resulting action sequence is executed on the robot without replanning, with objects placed by hand at

TABLE II: Task success rates (%) and reward coverage on LIBERO-PRO (50 trials/task for our runs). For evaluated policies, “Coverage” counts tasks per axis with at least one observed success in the reported trials (teacher: at least one successful synthesis; success rates use each task’s final tuned configuration). SFT success rates are averaged over three training seeds, and ± denotes the across-seed standard deviation of the axis average. Every seed achieves at least one success on all 57 tasks in 50 evaluation trials per task. Position

Task

Method

Spatial

Object

Goal†

Avg.

Coverage

Spatial

Object

Goal†

Before adaptation π0.5 -LIBERO

Avg.

Coverage

48.4

16.8

24.7

30.0

14/29

51.2

10.8

27.0

29.7

16/28

56.0 66.0 60.0

43.4 54.0 51.0

40.0 44.0 41.0

46.5 54.7 50.7

– – –

72.4 80.0 63.0

36.4 54.0 26.0

14.0 22.0 46.0

40.9 52.0 45.0

– – –

Sparse-reward control, matched RL compute π0.5 -LIBERO + PPO 50.8

33.6

41.3

41.9

15/29

79.2

52.4

73.5

68.4

22/28

SynthDemo adaptation SynthDemo-Teacher (1-attempt) SynthDemo-SFT SynthDemo-RL (ours)

68.0 56.5 99.0

68.9 54.8 98.0

65.6 56.2±1.3 97.8

29/29 29/29 29/29

60.1 53.3 95.0

70.6 66.3 99.6

82.4 42.7 96.8

71.0 54.1±1.9 97.1

28/28 28/28 28/28

Existing approaches (quoted)‡ SPARK [27] Pigey [28] CounterAlign [29]

60.0 57.3 96.4

† For our runs, Goal scores 9 tasks on Position (Task 3 excluded) and 8 on Task (Tasks 0 and 3 excluded, Sec. IV-A.1); quoted rows score 10. ‡ Quoted

F. Analysis The synthetic training set retains only successful episodes and matches the human demonstrations in policy, training budget, episode count, and initial-state provenance. Even so, SFT reaches a three-suite average of 56.7%, versus 97.7% for human demonstrations (Table I). Synthetic and human demonstrations differ in motion statistics and frozen-base loss. Table V quantifies the difference. Kinematic statistics separate the two sources completely: saturated-action frames and jerk norm reach δ ≈ 1, driven by the vertical axis. This servo signature is directly visible in Fig. 3, and it is not simply speed (velocity separates far more weakly, δ = 0.58). Under the frozen π0.5 base checkpoint, flow-matching loss is systematically higher for synthetic episodes (δ = 0.89 on Spatial, 0.99 on Object, and 0.52 on Goal), and saturation and jerk separate with δ ≥ 0.99 on all three suites. This indicates greater prediction error under the flow-matching objective, not a direct estimate of trajectory likelihood. Robometer-4B [31], a

ana

urn

Method

ban

ret

dis tra cto r

the twin’s nominal positions. Four conditions share one workspace and one success criterion (Fig. 2): stack (the red block onto the plate), banana (a novel object), return (the reversed placement goal), and distractor (a blue block on the usual pick spot while the instruction names the red one). Human teleoperation demonstrations exist for stack only, so the other three are the demonstration-free regime of this paper, now on hardware. Open-loop execution succeeds in all 20 trials per condition. As context, the humandemonstration reference (trained on real images and run closed-loop) reaches 50% on stack and 0% on the three conditions for which it has no demonstrations. The rows differ in training inputs and execution mode, so the table documents executability rather than a controlled comparison.

sta ck

results follow their source protocols (own evaluation harnesses; 10, 50, and 100 trials per task for Pigey, SPARK, and CounterAlign; the rewritten instruction on the Task axis, as in our runs) and are contextual references, not controlled comparisons.

Human-demo reference, π0.5 SFT, closed-loop SynthDemo-RL, no human demos, open-loop

50.0 100.0

0.0 100.0

0.0 100.0

0.0 100.0

Fig. 2: Physical executability on a real robot. Top: real camera views (red box) and the twin’s hybrid renders of the same views. Middle: the four conditions in the twin. Bottom: hardware success (%) over 20 trials per condition (target object released and at rest on its goal region). Ours executes twin rollouts open-loop, the reference runs closedloop, so the rows are not a controlled comparison. Human demonstrations exist for stack only.

pretrained vision-language reward model that estimates perframe progress from video and instructions, assigns lower progress monotonicity to synthetic episodes on Spatial and Goal (both p < 0.02), with similar final predicted progress. On Spatial, within-task DTW diversity shows little observed difference between synthetic and human demonstrations (δ = −0.02, p = 0.28), which does not point to a diversity deficit. The SFT gap persists after humanizing the teacher. Constraining the teacher’s servo velocity and rotation limits to human-measured levels and regenerating the dataset (a

TABLE III: Per-task success (%) on LIBERO-PRO (50 trials/task). Original: π0.5 -LIBERO before adaptation. Ctrl: direct PPO at matched RL compute. Ours: SynthDemo SFT, then PPO under the same RL compute as Ctrl (blue rows). Goal omits Task 3. Task-Goal Task 0 is excluded and not scored (ill-posed, see the footnote in Sec. IV-A.1). Gray cells: the 27 tasks with 0% Original success. Red cells: 0% success. Bold: best result per task. Suite

0

1

2

3

4

5

6

7

8

9

Position perturbations Spatial Original Spatial Ctrl PPO Spatial Ours (SFT) Spatial Ours (+RL)

Method

96 98 62 92

30 22 42 100

100 100 64 98

0 0 62 96

60 90 38 94

98 98 96 100

0 0 18 96

0 0 66 100

100 100 28 92

0 0 74 96

Object Object Object Object

Original Ctrl PPO Ours (SFT) Ours (+RL)

30 100 96 100

98 98 32 92

0 0 18 100

2 0 96 100

0 0 94 100

0 0 94 100

34 100 28 100

0 0 96 100

0 4 24 98

4 34 6 100

Goal Goal Goal Goal

Original Ctrl PPO Ours (SFT) Ours (+RL)

0 0 78 96

66 90 8 98

0 0 16 100

– – – –

0 0 94 100

0 0 6 94

0 0 78 94

88 100 90 100

68 98 92 100

0 84 74 100

Task perturbations Spatial Original Spatial Ctrl PPO Spatial Ours (SFT) Spatial Ours (+RL)

0 2 68 96

90 100 60 100

0 0 42 94

12 100 46 96

0 100 46 98

98 98 64 88

92 96 20 92

96 98 58 96

24 98 72 100

100 100 34 90

Object Object Object Object

Original Ctrl PPO Ours (SFT) Ours (+RL)

0 100 12 100

96 100 100 100

0 36 76 100

0 100 10 100

0 0 100 100

0 0 90 100

0 88 72 100

12 0 32 100

0 0 92 98

0 100 54 98

Goal Goal Goal Goal

Original Ctrl PPO Ours (SFT) Ours (+RL)

– – – –

20 0 26 100

24 88 22 94

– – – –

8 22 4 100

2 100 98 100

24 92 22 100

90 100 100 100

48 98 20 80

0 88 50 100

me am

a2b

86.0 56.0 66.0

80.0 32.0 75.0

79.3 55.5 74.3

Baselines (RoboTwin expert-generated demonstrations) (e) OpenVLA-OFT [30] 77.3 28.1 (e’) +SimpleVLA-RL (w/ demo) [2] 94.2 61.2

37.5 45.3

28.1 87.5

42.8 72.1

bea

Zero human demonstrations (ours) (a) SynthDemo-Teacher (b) SynthDemo-SFT (c) SynthDemo-RL

th

57.0 38.0 60.0

Method

ce pla

an mo ve c

94.0 96.0 96.0

ce pla

cup

r

TABLE IV: Success rates (%) on four RoboTwin 2.0 singlearm tasks. Rows (a)–(c) use 50 trajectories per task synthesized by our teacher. Rows (e, e’) use an OpenVLAOFT backbone trained on RoboTwin’s expert demonstrations (1,000 per task for SimpleVLA-RL), with results quoted from [2]. The comparison is contextual, not controlled.

Avg.

two-line generator-configuration change) gives column S† : saturation vanishes and the frozen-base loss falls slightly below the human value. The outcome moves far less: averaged over three seeds, regenerated-data SFT improves success from 55.3% to 60.2%, closing 4.9 of the 42.5-point Spatial gap to human-demonstration SFT (97.8%). Eight of ten tasks improve, and the regenerated episodes are 24% longer. This improvement is consistent with teacher-induced motion artifacts contributing to the SFT gap, but the regeneration does not isolate their effect or explain the remaining gap. Trajectory shape and grasp semantics remain possible con-

TABLE V: How far the synthetic demonstrations are from human ones, and what closing that distance buys. Per-episode medians on LIBERO-Spatial for the original synthetic data (S), the humanized-teacher regeneration (S† ), and human demonstrations (H). Cliff’s δ compares S against H (n = 864 episodes per source; Mann–Whitney p < 10−48 for every row except within-task DTW diversity, p = 0.28). The frozen-base loss scores each dataset under the unmodified π0.5 base checkpoint with identical normalization: a descriptive measure of mismatch with the frozen model. Measure Action saturation (frac. frames) vertical axis only Jerk norm Velocity norm Frozen π0.5 base FM loss Within-task DTW diversity Episode length (steps) SFT success (%)

S

S†

H

δ

0.158 0.498 0.125 0.118 0.123 0.199 124

0.000 – 0.087 – 0.090 – 153

0.027 0.094 0.069 0.099 0.093 0.200 123

1.00 1.00 0.997 0.58 0.89 −0.02 –

55.3±5.1

60.2±1.5

97.8

–

tributors; the Object-suite gap concentrates on three boxshaped objects that fit the enforced top-down grasp poorly. Under the fixed training schedule RL raises SynthDemoSFT from 55.3% to 96.2% on Spatial (Table I). Fig. 3 illustrates accompanying changes in policy behavior. SFT retains the teacher’s saturation signature, while RL changes the trajectory geometry: the detour ratio closes 41% of the SFT-to-human gap, while the per-frame ∆z distribution moves away from the human profile on all ten tasks (W1 distance). The 41-point gain therefore occurs without making the measured vertical-action distribution more human-like. V. L IMITATIONS AND F UTURE W ORK Our claim is human-demonstration-free adaptation with a simulator in the loop, not simulator-free learning: the teacher consumes ground-truth simulator state, and RL requires an environment to roll out in, so a target-domain simulation (as in Sec. IV-E) is a prerequisite. Replacing the teacher’s privileged inputs with learned perception, or the RL simulator with a learned world model [9], are untested extensions. SFT without rehearsal can reduce success on previously solved target tasks, although subsequent RL largely recovers these losses (Sec. IV-C). The teacher executes one primitive per task. Multi-step manipulation (e.g., LIBEROLong) would require sub-goal decomposition and sequencing of the existing primitives, which we have not implemented or evaluated. Closed-loop sim-to-real transfer remains future work (Sec. IV-E). The LLM tuner relies on a proprietary API. VI. C ONCLUSION We presented SynthDemo-RL, a framework that adapts VLA manipulation policies to new task distributions without new human teleoperation demonstrations, and studied reward coverage as a diagnostic of the distilled initialization. On the 57 scored perturbed LIBERO-PRO tasks, SynthDemo-SFT

vertical action Δz

H (human demo) S (teacher demo)

SFT policy

RL policy

1 0 −1 1 0 −1 1 0 −1 1 0 −1

H

S

SFT

RL

0

20

40

60 80 timestep

100

120

140

Fig. 3: From teacher signature to policy behavior. Top: end-effector paths on one LIBERO-Spatial task: human demonstrations (H), teacher demonstrations (S), and successful rollouts before (SFT) and after RL, overlaid thin. The episode closest to each source’s median detour ratio (path length over start–end displacement) is bold. SFT and RL share the initial state, and all panels share one crop. Bottom: vertical action of the bold episodes; gray bands mark saturation (|a| ≥ 0.9375). RL straightens the geometry (median detour ratio 3.7 → 3.2 vs 2.5 for human) but keeps the teacher’s saturation signature: median saturation 0.126 (SFT) and 0.133 (RL) against 0.158 (S) and 0.027 (H).

rescues all 27 tasks with no observed success before adaptation. Subsequent PPO raises success to 97.8% (Position) and 97.1% (Task), while the PPO-matched sparse-reward control rescues 10 of the 27 and leaves 17 at 0%. The pipeline reaches 96.0% on standard LIBERO without collecting new human demonstrations and carries over to RoboTwin 2.0. Its twin-trained trajectories execute open-loop on a physical robot. These findings suggest that demonstration synthesis should prioritize reward coverage, at least one successful trajectory on every task, and leave raising the average success rate to sparse-reward RL. ACKNOWLEDGMENTS OpenAI tools (GPT-5.5 and Codex) assisted with automated teacher tuning and manuscript refinement. R EFERENCES [1] M. J. Kim et al., “OpenVLA: An open-source vision-language-action model,” in CoRL, 2024. [2] H. Li et al., “SimpleVLA-RL: Scaling VLA training via reinforcement learning,” in The Fourteenth International Conference on Learning Representations, 2026. [3] G. Lu et al., “VLA-RL: Towards masterful and general robotic manipulation with scalable reinforcement learning,” arXiv preprint arXiv:2505.18719, 2025. [4] H. Zang et al., “RLinf-VLA: A unified and efficient framework for reinforcement learning of vision-language-action models,” arXiv preprint arXiv:2510.06710, 2025. [5] A. Rajeswaran et al., “Learning complex dexterous manipulation with deep reinforcement learning and demonstrations,” in Robotics: Science and Systems (RSS), 2018.

[6] Y. J. Ma et al., “Eureka: Human-level reward design via coding large language models,” ICLR, 2024. [7] Y. Wang et al., “RoboGen: Towards unleashing infinite data for automated robot learning via generative simulation,” ICML, 2024. [8] X. Zhou et al., “Libero-pro: Towards robust and fair evaluation of vision-language-action models beyond memorization,” arXiv preprint arXiv:2510.03827, 2025. [9] H. Li et al., “VLA-RFT: Vision-language-action reinforcement finetuning with verified rewards in world simulators,” arXiv preprint arXiv:2510.00406, 2025. [10] K. Chen et al., “πrl : Online rl fine-tuning for flow-based visionlanguage-action models,” arXiv preprint arXiv:2510.25889, 2025. [11] A. Mandlekar et al., “MimicGen: A data generation system for scalable robot learning using human demonstrations,” in Conference on Robot Learning (CoRL), 2023. [12] S. Nasiriany et al., “RoboCasa: Large-scale simulation of everyday tasks for generalist robots,” in Robotics: Science and Systems (RSS), 2024. [13] H. Ha, P. Florence, and S. Song, “Scaling up and distilling down: Language-guided robot skill acquisition,” in Conference on Robot Learning (CoRL), 2023. [14] J. Zhang et al., “RewiND: Language-guided rewards teach robot policies without new demonstrations,” in 9th Annual Conference on Robot Learning, 2025. [15] W. Xiao et al., “Self-improving vision-language-action models with data generation via residual RL,” arXiv preprint arXiv:2511.00091, 2025. [16] A. Ye et al., “GigaBrain-0: A world model-powered vision-languageaction model,” arXiv preprint arXiv:2510.19430, 2025. [17] M. Dalal, J. Yang, R. Mendonca, Y. Khaky, R. Salakhutdinov, and D. Pathak, “Neural MP: A generalist neural motion planner,” arXiv preprint arXiv:2409.05864, 2024. [18] J. Lee, J. Hwangbo, L. Wellhausen, V. Koltun, and M. Hutter, “Learning quadrupedal locomotion over challenging terrain,” Science Robotics, vol. 5, no. 47, 2020. [19] T. Jülg, W. Burgard, and F. Walter, “Refined policy distillation: From vla generalists to rl experts,” 2025. [20] A. Murali et al., “Graspgen: A diffusion-based framework for 6-dof grasping with on-generator training,” 2025. [21] OpenAI, “Introducing GPT-5.5,” https://openai.com/index/ introducing-gpt-5-5/, 2026, accessed 2026-09-09. [22] Physical Intelligence, K. Black et al., “π0.5 : a vision-languageaction model with open-world generalization,” arXiv preprint arXiv:2504.16054, 2025. [23] J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” arXiv preprint arXiv:1707.06347, 2017. [24] J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel, “Highdimensional continuous control using generalized advantage estimation,” in ICLR, 2016. [25] B. Liu et al., “LIBERO: Benchmarking knowledge transfer for lifelong robot learning,” in NeurIPS, 2023. [26] T. Chen et al., “Robotwin 2.0: A scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation,” arXiv preprint arXiv:2506.18088, 2025. [27] B. Grant, A. Rothenberg, L. Senning, Z. Chua, Z. Patterson, and P. Wang, “Sequential planning via anchored robotic keypoints,” arXiv preprint arXiv:2606.30613, 2026. [28] L. Galanti, D. Shah, and T. Dao, “Addressing the orchestration gap in generalist robots via physical agency,” arXiv preprint arXiv:2607.21725, 2026. [29] H. Kondoh, K. Ota, A. Kanezaki, and Y.-H. Wu, “CounterAlign: Counterfactual supervision for Vision-Language-Action models,” arXiv preprint arXiv:2608.21740, 2026. [30] M. J. Kim, C. Finn, and P. Liang, “Fine-tuning vision-language-action models: Optimizing speed and success,” in Robotics: Science and Systems (RSS), 2025. [31] A. Liang et al., “Robometer: Scaling general-purpose robotic reward models via trajectory comparisons,” arXiv preprint arXiv:2603.02115, 2026.

Record · ID 1006910 · SHA-256 4ee15b083ac0810a
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.