Preprint
P REDACTOR : P REDICTIVE ACTION D IFFUSION FOR S TEERABLE O NBOARD H UMANOID C ONTROL
arXiv:2609.24840v1 [cs.RO] 21 Sep 2026
Lei Ye1,2,3 , Haibo Gao1,* , Yitang Li3,4 , Peng Xu1,* , Zetong Jing3 Junhan Sun3 , Fanrong Dong3 , Ziqi Han3 , Xue Wang3 , Jianhua Sun5,* Cewu Lu2,5 , Hao Zhao3,4 , Liang Ding1,* 1 Harbin Institute of Technology 2 Shanghai Innovation Institute 3 RoboParty Lab 4 Tsinghua University 5 Shanghai Jiao Tong University * Corresponding authors: Haibo Gao, Peng Xu, Jianhua Sun, and Liang Ding
(a) Joy Steering Locomotion
(b) Behavioral Reaction to External Interference
Cmd: Walk Robot hesitate when disturbing
Cmd: Running
Cmd: Walking
Cmd: Stand Robot react to human interaction
(c) Semantic Interpotion
Stand
Raise Hand
Stand
Running Past
Future
Text Cmd Joy Cmd
Task Interface
······
Stand
State Walk forward
Bow
Run forward
Jump
Action
PredActor
Raise hand
PREDICTIVE ACTION DIFFUSION
Time
Walk
Jog
Squat Down
Walk
Squat Down
Walk
Stand
(d) Text Command for Robot Control
Figure 1: Qualitative demonstrations of PredActor. (a) Joystick steering redirects walking and running; (b) PredActor responds to external interference during walk and stand commands; (c) semantic interpolation connects standing with raising a hand and with running; and (d) text commands trigger stand, walk, jog, and squat sequences.
A BSTRACT Diffusion models offer flexible motion generation, but translating this flexibility into feedback-responsive humanoid control remains challenging. Hierarchical systems steer motion through references that may exceed a separate tracker’s capabilities, leaving recovery and physical execution largely to the tracker. Actiononly diffusion generates actions directly but lacks an explicit future-state trajectory for test-time motion objectives. Joint state–action diffusion provides this representation, yet representative controllers often depend on privileged full-body states, and support for learned behavior selection and test-time motion steering remains fragmented. We present PredActor, a predictive action diffusion policy that brings these complementary steering capabilities into one directly executed policy using proprioceptive observations. Conditioned on proprioceptive history and optional task context, PredActor jointly generates executable actions and an internal future-state trajectory. Classifier-free guidance strengthens text1
Preprint
conditioned behavior, while classifier guidance steers predicted states toward testtime objectives. Only actions are executed, without a separate motion-reference tracker or externally estimated full-body states as policy inputs. In simulation, PredActor reaches all 15 destination targets and achieves a text retrieval score of 0.580, compared with 0.373 for conditional action diffusion, with similar observed disturbance survival. To make this guided policy practical onboard, rolling denoising and computation-preserving runtime optimizations reduce the complete callback to 16.790 ms median and 19.383 ms p95 on a Jetson Orin NX, both below the 20 ms control period. We deploy PredActor on a Unitree G1; evaluations across simulation and physical hardware demonstrate text-conditioned motion, disturbance response, joystick control, and semantic interpolation. Project page: https://masteryip.github.io/predactor.github.io/.
1
I NTRODUCTION
Humanoid policies need to follow task commands while adapting to physical feedback. Advances in motion tracking and behavior foundation models have expanded robot skills and task interfaces (Luo et al., 2023; RoboParty Lab Team, 2026; Luo et al., 2026; Chen et al., 2026a; Li et al., 2025a; Zeng et al., 2025). Diffusion models offer flexible conditioning for diverse motion and action generation (Tevet et al., 2023; Chi et al., 2023). These capabilities raise a question: how can a policy guide future motion while directly generating executable actions?
CFG
CG
Kinematic Diffusion
S0
(b) Close Loop Track
S1
⋯
(c) Action Diffusion
CFG
CFG
Sn
S0
S1
⋯
Action Diffusion
Sn
Tracking Policy
O
Robot
Tracking Policy
indirect
A0 O
Robot
··· An direct action
Robot
CFG
O
CG
Joint diffusion
PredActor
S-1 S0 ··· Sn
S-1 S0 ··· Sn
A-1 A0 ··· An
A-1 A0 ··· An
direct action
indirect
O
(e) PredActor
CG
Kinematic Diffusion
A0 A1 A0
(d) Joint Diffusion
State Estimator
(a) Open Loop Track
Robot
direct action
O
Robot
Figure 2: Paradigms of diffusion policies: (a,b) generator–tracker; (c) action-only; (d,e) joint state– action diffusion. O: observation; CG/CFG: method-specific guidance. Diffusion control paradigms differ in how motion generation informs action execution (Figure 2). Kinematic generators supply references to trackers. Open-loop variants operate without robot feedback (Xie et al., 2026; Zhao et al., 2026; Wang et al., 2026), whereas closed-loop variants incorporate robot observations (Tevet et al., 2025; Chen et al., 2026b). Both may generate dynamically infeasible motion outside tracker support. Moreover, they don’t show behavioral reaction under interference and reaction and recovery remain largely with the tracker, while separate planning and control clocks complicate synchronization and replanning (Huang et al., 2025b; Gu et al., 2026). Action-only diffusion bypasses reference tracking by generating actions directly (Chi et al., 2023; Truong et al., 2024; Huang et al., 2025a; Høeg et al., 2024). Its predictions, however, contain no explicit future-state trajectory for classifier guidance (CG). Joint state–action diffusion provides this predictive representation (Huang et al., 2025b; Carroll et al., 2026; Zhang et al., 2026), but representative controllers often condition on privileged fullbody states. That dependence makes sim-to-real transfer rely on a deployable state estimator, while the representation’s complementary interfaces remain unevenly supported: classifier-free guidance (CFG) strengthens learned behavior conditions, whereas classifier guidance (CG) steers predicted motion toward test-time objectives, yet CFG is uncommon among representative joint controllers (Huang et al., 2025b; Liao et al., 2026). The open question is therefore functional: can one directly executed policy use deployable observations to combine learned behavior selection and test-time motion steering while retaining closed-loop disturbance robustness? We introduce PredActor, a predictive action diffusion policy within this joint formulation (Figure 2(e)). It jointly denoises future states and actions from proprioceptive history and optional task context. Predicted states remain internal and provide the representation for CG; conditional and null
2
Preprint
predictions support CFG. Only the selected action is issued to the joint controller, without a separate motion-reference tracker. This functional gap is accompanied by a deployment gap: the computation of joint guided diffusion has left few complete onboard demonstrations. PredActor uses proprioceptive history as input, a partial-observation setting also studied by SCDP (Carroll et al., 2026), and future-state trajectories as training targets. To broaden training coverage, perturbed teacher rollouts supply recovery data, and a DAgger-style procedure aggregates teacher labels at learner-visited states (Laskey et al., 2017; Ross et al., 2011). Iterative sampling also delays feedback; rolling denoising and computationpreserving runtime optimizations address this constraint, with their onboard effect evaluated later. In simulation, PredActor reaches all 15 destination targets and achieves 0.580 text retrieval versus 0.373 for conditional action diffusion, with similar observed push survival (0.535 versus 0.564). On Jetson Orin NX, the complete callback reaches 16.790 ms median and 19.383 ms p95, below the 20 ms control period at both percentiles. To our knowledge, PredActor is the first joint state–action diffusion policy deployed entirely on a Unitree G1’s onboard Jetson Orin NX for 50 Hz control. Across simulation and physical G1 evaluation, demonstrations cover text commands, disturbance response, joystick steering, and semantic interpolation. Our contributions are (1) a predictive action diffusion policy for steerable humanoid control, using only proprioceptive observations to simplify deployment and supporting both CG and CFG for task steering while retaining direct action execution; (2) a data and training pipeline with a motion library, automatic task labels, asynchronous tracker rollouts, and learner-state aggregation; (3) the first fully onboard deployment of a joint state–action diffusion policy on a Unitree G1’s Jetson Orin NX for 50 Hz control, supported by rolling inference, runtime optimization, and delay compensation; and (4) an editable task-control interface and open-source training-todeployment implementation.
2
R ELATED W ORK
Kinematic Motion Generation and Tracking. Physics-trained trackers execute diverse reference motions (Luo et al., 2023; RoboParty Lab Team, 2026; Luo et al., 2026; Chen et al., 2026a), while diffusion generators provide multimodal, commandable kinematic synthesis (Tevet et al., 2023). Their complementarity motivates hierarchical systems that separate kinematic planning from physical execution (Xie et al., 2026; Zhao et al., 2026; Wang et al., 2026). Open-loop variants do not return robot feedback to the generator; CLoSD and ReactiveBFM instead return executed state to later planning windows (Tevet et al., 2025; Chen et al., 2026b). Both forms retain the same interface limitation: references may violate dynamics or tracker support, and separate planning/control clocks create synchronization and replanning delay. Feedback mitigates drift but does not intrinsically address dynamic-feasibility during generation. Generalist BFMs broaden skills and task interfaces but retain the separation between motion generation and physical execution (Li et al., 2025a; Zeng et al., 2025; Shao et al., 2025). Diffusion Policies for Humanoid Control. Diffusion control also spans policies that map observations and task conditions directly to actions (Chi et al., 2023; Truong et al., 2024; Huang et al., 2025a). Action-only diffusion, however, predicts no accompanying state trajectory for state objectives at test time. Joint state–action diffusion restores this structure and enables state-based shaping of executable actions (Janner et al., 2022; Huang et al., 2025b; Liao et al., 2026). Representative systems in Table 1 emphasize different capabilities and steering interfaces. For guidance, classifier-free guidance mixes conditional and null predictions, whereas classifier guidance adds gradients from a test-time objective (Ho & Salimans, 2021). Diffuse-CLoC and BeyondMimic emphasize classifier guidance (Huang et al., 2025b; Liao et al., 2026); SCDP addresses partial observation but reports neither classifier guidance, classifier-free guidance, nor rolling reuse; and SCRIPT adds language conditioning and post-training but is evaluated in simulation (Carroll et al., 2026; Zhang et al., 2026). Beyond modeling and conditioning, deployment requires efficient closed-loop inference. Replanning with overlapping diffusion horizons can add latency and induce cross-tick disagreement (e.g., fixed action chunks delay feedback and ensembling can average incompatible modes) (Chi et al., 2023; Zhao et al., 2023). Fast samplers reduce denoising steps (Song et al., 2021; Lu et al., 2022; Song et al., 2023), and streaming methods reuse partially denoised horizons across ticks (Høeg
3
Preprint
Table 1: Capabilities and deployment of generative control systems. Full hardware specifications, inference timings, and source evidence appear in Table 4 and Appendix A.1. Model settings
Task capability
Deployment
Method
Paradigm
Proprio.
CFG
CG
Text
Robot
Inference GPU
Control rate
MotionBricks + SONIC ARDY + SONIC TextOp ReactiveBFM CLoSD DiffuseLoco Diffuse-CLoC BeyondMimic SCDP SCRIPT
Generator + tracker Open-loop track Open-loop track Closed-loop track Closed-loop track Action diffusion Joint diffusion Latent joint diffusion + dec. Joint diffusion Joint diffusion
⋄ ⋄ ⋄ ⋄ – ✓ – ⋄ ✓ –
– ⋄ ✓ – ✓ – – – – ✓
– – – – – – ✓ ✓ – –
– ⋄ ✓ ✓ ✓ – – – – ✓
✓ – ✓ ✓ – ✓ – ✓ ✓ –
Orin (on) RTX 4090b RTX 4090 (off)c RTX 4090 (off) – RTX 4060M (on) RTX 4060 (sim) RTX 4060M (on) RTX 5090 (off) –
50 Hz trackera – 50 Hz tracker 50 Hz N/A 30 Hz target N/A 25 Hz 50 Hz N/A
PredActor
Joint diffusion
✓
✓
✓
✓
✓
Orin NX (on)
50 Hz
Symbols: ✓: supported; ⋄: partial/modular; ×: explicitly unsupported; –: unreported/unverified; N/A: not applicable. Abbreviations: Proprio.: proprioceptive input; dec.: decoder; Robot: physical deployment; on/off: onboard/offboard; sim: simulation; M: Mobile. Rates refer to robot control, not generation throughput. a MotionBricks uses a non-diffusion generator; 50 Hz is from SONIC, not separately reported for the composition. b ARDY generator-demo GPU; robot-stack rate unreported. c TextOp external generator GPU; onboard tracker.
et al., 2024; Zhang et al., 2024; Chen et al., 2024), but joint state–action prediction additionally requires aligned schedules and reuse at compatible noise levels. PredActor is also a joint state– action policy. It combines learned commands and predicted-state objectives under proprioceptive conditioning, exposes both classifier-free and classifier guidance within one directly executed policy (Table 1), and uses per-stream, per-position scheduling with schedule-matched cross-tick reuse to keep state–action horizons aligned. Our deployment study fixes a sampling budget and measures computation-preserving optimizations end-to-end in the onboard callback. Appendix A records the evidence.
3
P REDACTOR
PredActor augments action diffusion with an internal future-state trajectory for task guidance. It is a joint state–action policy: predicted states expose future motion to steering objectives, and actions are executed directly. We first define this representation and its CG/CFG interfaces (Figure 3), then explain how teacher trajectories supervise it (Figure 4), and finally describe rolling inference and onboard execution (Figure 5). Detailed tensor layouts, schedule derivations, and system interfaces appear in Appendix B. 3.1
P ROBLEM F ORMULATION AND A RCHITECTURE
At control time t, let ot ∈ Rdo denote the proprioceptive observation available on the robot, phys s(t) ∈ Rds its full-body motion state, and a(t) ∈ Rda the action issued to the joint controller. Conditioned on an observation history ot−ℓ+1:t of length ℓ and optional task context z = [z1 , . . . , zn ] of n task tokens, PredActor predicts clean trajectories s0 ∈ Rh×ds and a0 ∈ Rh×da over h horizon positions in model coordinates. Superscripts denote diffusion noise levels (0 is clean), and subscripts i ∈ {1, . . . , h} index horizon positions. Given corrupted trajectories, the denoiser Fθ predicts: (ŝ0 , â0 ) = Fθ (sks , aka , ks , ka ; ot−ℓ+1:t , z), (1) where ks , ka ∈ {0, . . . , T }h specify per-position noise levels, T is the maximum diffusion level, and sks , aka are the corresponding noisy inputs. Full-body states supply training targets, while deployment requires only proprioceptive observations and optional task context. Predicted states remain internal guidance variables; only the selected action is issued to the joint controller, without a separate motion-reference tracker. Feature definitions and the physical-to-model state representation are detailed in Appendix B.1. Figure 3 follows conditioning, denoising, and sampling from top to bottom. The encoder maps the observation-history tokens (shown as o1 , . . . , oℓ ) and task tokens z1 , . . . , zn into condition memory 4
Preprint
PredActor Architecture CONDITIONING
Proprio Observation
Objective Steering (CG)
Task Conditioning (optional CFG)
⋯
⋯
⋯
MLP/Transformer Encoder
Condition Memory
⋯
⋯
Enc
c₁ c₁c₁
Clean Prediction Demux
Enc
Cross Attn Causal Attn Self Attn
c₁ c₁c₁
DENOISING
⋯
Interleaving Mux
Noise Level
DENOISING
Diffusion Transformer State/Act Denoise
c₁ c₁c₁
Time
Dec
⋯
Dec
⋯
SAMPLING
⋯
Denoise Update
Pos+ Noise embed
Noise Scheduler
Guidance Update
Figure 3: Joint predictive architecture of PredActor. Observation history and task commands condition a transformer that jointly denoises interleaved future state and action tokens. Optional steering modifies the predicted state while passing the predicted action through unchanged to the denoising update; only an action is executed. C = [c1 , . . . , cm ], with m encoded tokens. Separate projections embed noisy state and action tokens si , ai with their noise levels and horizon positions. We adopt Diffuse-CLoC’s interleaved state–action token pairs and asymmetric self-attention pattern (Huang et al., 2025b): state queries attend to all state tokens but not action tokens, while action queries attend causally to both streams. PredActor extends this backbone with cross-attention to C, allowing the denoising tokens to query encoded proprioceptive history and task context. Separate heads produce clean predictions ŝ0 , â0 (shown as ŝi , âi ). The noise scheduler selects per-position target levels k′s , k′a , and the denoising update returns both streams to the next sampling step (Section 3.3). At the application level, task conditioning and objective steering provide complementary interfaces. Conditional, null, and partial-condition training supports classifier-free guidance (CFG), which combines conditional and null predictions to strengthen learned task behavior, such as text-conditioned motion (Ho & Salimans, 2021). Steering commands g = [g1 , . . . , gr ] specify r inputs to a differentiable objective J (ŝ0 ; g). Classifier guidance (CG) uses its state gradient ∇ŝ0 J to modify the predicted state before the reverse update (Huang et al., 2025b); subsequent action queries can attend to the adjusted state, while CG leaves the action-head output unchanged at the current update. Thus CFG strengthens learned task conditions and CG steers future motion toward test-time objectives. Both are optional, and only the resulting action is executed. 3.2
DATA AND T RAINING
Figure 4(a) constructs a collection-ready kinematic library from prompted motion generation and AMASS motions, with BABEL annotations and automatic task labeling supplying text commands and task terms. The adaptive sampler selects motion references for the asynchronous rollout pipeline in panel (b). A frozen tracking teacher executes these references under observation noise, action noise, and external pushes to form an initial dataset D0 of observation histories, future states, teacher actions, and task labels. To cover learner-induced deviations, DAgger-style training alternates student rollout collection, teacher trajectory shooting, and diffusion reconstruction (Ross et al., 2011). At round j, the frozen student πθj visits origins Rj ; from each origin ξ, a separate teacher branch follows the aligned motion reference to produce a state–action target window WE (ξ), paired with the student’s observed history and task label. The lower timeline depicts these teacher continuations
5
Preprint
(a) Motion & Auto Labeliing Prompted Kimodo Gen
(b) Diffusion Dagger Training Aync Tracker Rollout Collection Teacher Obs
Async Env Rollout
Motion Lib
AMASS
Obs Noise
Auto Labeler Task Terms Text Cmd
BABEL
External Push
Diffusion Student
Diffusion Student Behavior Clone
Student Obs Motion Traj
······
Collection Ready Motion Lib
Action OU Noise
Tracker Teacher
Adaptive Sampler (Hard Mining)
Task Cond Label
Combined Dataset
Student Rollout
+
Teacher Shooting
Epoch Collecting Collected Dataset
Figure 4: Data and training pipeline. Generated and annotated motion references drive asynchronous tracker rollouts. The resulting deployable observations and task labels condition joint state–action training. Teacher trajectory shooting supplies trajectory targets at student-visited states, which are aggregated for subsequent training. branching from student-visited states; the collected windows enter the combined dataset on the right, and the updated student is used in the next collection round: Dj+1 = Dj ∪ {WE (ξ) : ξ ∈ Rj }, θj+1 ≈ arg min L(θ; Dj+1 ).
(2)
θ
Both initial training and aggregated replay reconstruct independently corrupted state and action windows using X h h X ∥∆â0i − ∆a0i ∥22 wia ∥â0i − a0i ∥22 + ρ L(θ; D) = E i=2
i=1
+ γs
h X
(3)
wis ∥ŝ0i − s0i ∥22 ,
i=1
where the expectation covers training windows, noise, and diffusion levels; wia , wis weight horizon positions, ∆a0i = a0i − a0i−1 , and ρ, γs weight action-difference regularization and state reconstruction. The three terms supervise actions, action differences, and future states. The teacher supplies training targets only and is absent at deployment; dataset partitioning, branch isolation, and window alignment are detailed in Appendix B.2. 3.3
ROLLING I NFERENCE AND D EPLOYMENT
Figure 5(a) illustrates the receding-horizon loop used to fit repeated denoising within the control period. Training samples noise levels independently, whereas rolling inference carries a partially denoised state–action horizon across control ticks. After each action is issued, the buffers shift and the latest proprioceptive observation conditions the next denoising call. For either trajectory stream u ∈ {s, a}, the DDIM update below converts a clean prediction into a reverse-sampling state at the target noise level: √ q √ ′ ut − ᾱt û0 , ut = ᾱt′ û0 + 1 − ᾱt′ − σt2 ϵ̂t + σt ϵ, (4) ϵ̂t = √ 1 − ᾱt where t′ < t, ϵ ∼ N (0, I), ᾱt is the cumulative noise schedule, and the DDIM variance for this jump is r r 1 − ᾱt′ ᾱt σt = η 1− . 1 − ᾱt ᾱt′ 6
Preprint
The evaluated deterministic profiles set η = 0, so σt = 0 and the random-noise term vanishes. The runtime can therefore carry intermediate samples across control steps instead of restarting the full horizon. Each horizon position can take its own DDIM jump. PredActor represents the source and target levels by matrices S, S ′ ∈ ZK×h , with each entry broadcast over the feature dimension at its horizon position; state and action streams may use different matrices. The runtime reuses the saved intermediate sample whose source level matches the next required position, adds fresh noise only to newly exposed tail positions, and denoises the shifted horizon under the latest observation. This schedule-matched reuse preserves every sample’s diffusion coordinate rather than copying a clean prediction into an arbitrary noise level. In Figure 5(b), the delay coordinate u = (nobs − 1) + δ/∆t selects adjacent clean action slots for interpolation, compensating actuation delay in the same horizon coordinates used for planning. For onboard execution, we export the diffusion actor and control interface to C++. The evaluated Orin profile uses two-step DDIM; all-QKV packing, cached condition embeddings and DDIM schedules, stacked-axis WBG, and LibTorch inference mode reduce repeated work in the callback. These operation-level optimizations are evaluated end-to-end in Section 4.2. The same deployed interface supports text commands, semantic interpolation, and whole-body steering, while Appendix B.3 specifies the export, safety, and timing boundaries.
4
E XPERIMENTS
We ask three questions: Q1, can PredActor combine CG destination steering and CFG text control while retaining disturbance recovery? Q2, can exact guided inference meet the 20 ms onboard budget? Q3, can the same policy support simulated interfaces and physical G1 execution? The policy controls 29 actuated joints at 50 Hz from proprioceptive history and predicts a 20-step state– action horizon. Full protocols, metrics, coverage, and uncertainty treatment are in Appendix C. 4.1
S TEERING AND C LOSED -L OOP C ONTROL
Setup. We compare DP w TextCFG, CLOC, CLOC w TextCFG, PredActor w/o TextCFG, PredActor, and ARDY+SONIC in one calibrated MuJoCo plant using matched push, destination, and text protocols (Table 2). The five diffusion policies use the common training contract described in Appendix C.2; ARDY+SONIC is the separately trained generator–tracker reference. Results. PredActor attains 0.535 ± 0.062 push survival, close to DP w TextCFG (0.564 ± 0.142), and reaches all 15 destination targets with 0.588 ± 0.005 m final error. It achieves 0.993 ± 0.017 text survival and 0.580 ± 0.049 retrieval over 45 trials, versus 0.988 ± 0.028 and 0.373 ± 0.090 for DP w TextCFG. Its 1.832 ± 0.115 ms actor latency and 12.603 ± 0.426 jerk RMS are comparable to DP w TextCFG (15.357 ± 0.514 ms, 12.654 ± 0.676) and far below CLOC w TextCFG (11.134 ± (a) Accelerated Inference
(b) Delay Compensation
Receding Horizon DDIM
noise level
Joint Transformer
Ramp Noise Schedule
Packed all-QKV
Denoise Delay coordinate
Packed WBG
Interpolate clean slots
Precomputed Mask
t DDIM update Cached Schedule
Receding Horizon DDIM
Robot Command
next observation -> shift horizon -> replan
Figure 5: Rolling inference and delay compensation. Schedule-matched samples are shifted and replanned after each observation. A delay coordinate interpolates adjacent clean action slots before execution. 7
Preprint
Table 2: Six-policy comparison of robustness, navigation, and text control. Means ± sample SD; brackets show evaluated/requested cells. Within each numeric column, cell color runs from the worst value through a white midpoint to the best value, respecting the direction of the metric. Method DP w TextCFG
Basic Jerk RMS ↓ Latency ↓ 3 −3 (10 s ) (ms) 12.654 15.357
Perturbation Survival ↑ 0.564
Survival ↑ N/E
+/-0.676 [5/5]
+/-0.142 [100/100]
[0/15]
[0/15]
0.187
0.000
+/-0.514 [5/5]
283.057
7.534
0.333
+/-18.571 [5/5]
+/-0.345 [5/5]
+/-0.070 [100/100]
CLOC w TextCFG
105.074
11.134
0.330
+/-10.863 [5/5]
+/-0.624 [5/5]
+/-0.081 [100/100]
PredActor w/o TextCFG
11.764
1.763
0.566
+/-0.272 [5/5]
+/-0.218 [5/5]
+/-0.097 [100/100]
12.603
1.832
0.535
+/-0.426 [5/5]
+/-0.115 [5/5]
+/-0.062 [100/100]
21.274
N/E
0.534
+/-12.741 [5/5]
[0/5]
+/-0.070 [100/100]
CLOC
PredActor ARDY + SONIC
Point Navigation Last valid err. ↓ Min valid err. ↓ Arrival ↑ (m) (m) N/E N/E N/E [0/15]
+/-0.125 [15/15] +/-0.000 [15/15]
0.847 0.665 1.000 1.000
[0/45]
3.570
N/E
N/E
N/E
[0/45]
[0/45]
[0/45]
4.518
2.592
0.673
0.582
N/E
+/-1.939 [15/15]
+/-0.947 [15/15]
+/-0.038 [45/45]
+/-0.048 [28/45*]
[0/45]
3.613
3.546
N/E
N/E
N/E
+/-0.416 [15/15]
+/-0.451 [15/15]
[0/45]
[0/45]
[0/45]
0.588
0.588
0.993
0.580
+/-0.005 [15/15]
+/-0.005 [15/15]
+/-0.017 [45/45]
+/-0.049 [45/45]
[0/45]
0.656
0.656
1.000
0.237
0.511
+/-0.080 [2/15*]
+/-0.080 [2/15*]
0.500
+/-0.000 [2/15*] +/-0.707 [2/15*]
Admit. ↑ N/E
+/-0.308 [15/15]
1.000
+/-0.000 [15/15] +/-0.000 [15/15]
Retrieval ↑ 0.373 +/-0.090 [45/45]
4.187
0.000
+/-0.366 [15/15] +/-0.000 [15/15]
Survival ↑ 0.988 +/-0.028 [45/45]
+/-0.695 [15/15]
0.400
+/-0.115 [15/15] +/-0.279 [15/15]
[0/15]
Text Command
N/E
+/-0.000 [23/45*] +/-0.198 [23/45*] +/-0.099 [23/45]
Bold marks the numeric optimum at the displayed precision, including conditional cells and ties. N/E denotes a capability boundary and NR denotes an unavailable recording; neither is zero-imputed. An asterisk marks a conditional subset excluded from full-coverage comparisons. Push survival averages 100 cells (20 paired direction/seed clusters × five forces), with SD across the 20 cluster means. Other rollout SDs use five seed means. Latency is synchronized, post-warmup RTX 4090 actor time amortized per action at B = 5. Point errors exclude one verified post-terminal autoreset sample; text retrieval uses complete windows beginning at least 3 s after command activation. (a) Navigation & Perturbation
(b) Text Command Retrival DP w TextCFG 45/45 scored | 1,586 win.
CLOC w TextCFG 28/45 scored | 942 win.
share
Eval Motions 1.0
stand
0
jog
−1
Pelvis Position x (m)
PredActor (Proposed) 45/45 scored | 1,600 win.
0.0
kick
jump
raise hand
boxing
jog
run
squat down
walk
kick
stand
jump
kick
5
raise hand
4
boxing
3
jog
2
0.5
jump
r=0.6
run
1
N/E
raise hand
squat down
0
squat down boxing
walk
−3
CLOC CLOC w TextCFG PredActor w/o TextCFG PredActor (Proposed) ARDY + SONIC
run
stand
−2
Requested
Pelvis Position y (m)
walk
ARDY + SONIC 23/45 scored | 805 win.
share
Requested
ARDY admissibility 23/45; 22 N/E
1.0
stand
stand
3/5
walk
walk
4/5
jog
jog
4/5
run
4/5
squat down
0/5
boxing
boxing
1/5
raise hand
raise hand
5/5
jump
jump
run squat down
N/E
0.5
1/5 1/5
kick
Raw predicted
kick
jump
raise hand
boxing
run
squat down
jog
walk
stand
kick
jump
raise hand
boxing
run
squat down
jog
walk
stand
kick
0.0
0
0.5
1.0
Accepted
Raw predicted
Figure 6: MuJoCo evidence for steering, recovery, and text control across six policies. (a) Representative destination trajectories in the common protocol frame (top; the star is the target and the circle marks the 0.6 m arrival threshold) and direction-matched pelvis tilt under a 300 N push (bottom; lines are means, bands show sample SD, and crosses mark early terminals) across 20 paired direction/seed cells per policy. (b) Per-class raw MotionCLIP top-1 retrieval for each evaluable policy, example evaluation motions, and ARDY+SONIC reference admissibility by requested class. Missing cells remain distinct from zero; 22/45 ARDY references are rejected before SONIC rollout. Table 2 aggregates all 100 push cells per policy and reports each selected metric, denominator, and capability boundary. 0.624 ms, 105.074 ± 10.863). Figure 6 and Table 2 report per-class panels and denominators; incomplete ARDY+SONIC and CLOC subsets are not ranked against full-coverage rows. These results show that internal future-state prediction combines CG steering and CFG behavior selection with competitive smoothness, latency, and tested disturbance performance. 4.2 I NFERENCE ACCELERATION ON J ETSON O RIN Setup. We benchmark batch-one FP32 guided inference and the complete callback on one Jetson Orin NX. The sampler sweep and staged graph, structural, and runtime measurements are defined in Appendix D.1; the timing benchmark produces no actuation commands. Results. Figure 7(a,d) shows that two-step NS-DDIM gives the best tested speed–robustness tradeoff. The selected implementation reduces actor p50 from 21.461 to 15.930 ms along the structural path, and inference mode brings the remeasured actor to 13.175 ms (Table 3). The complete callback is 16.790/19.383 ms at p50/p95; 593/600 calls meet 20 ms, with output drift below 1.073 × 10−6 . 8
Preprint
Table 3: Cumulative acceleration on Orin. Frozen actor CUDA timing except the final live callback; NS-DDIM2, B=1 FP32, two forwards, WBG on, eager finite checks. p50
p95
Saved
> 20 ms
Sequential WBG; other switches off + stacked WBG + DDIM precomputation + all-QKV packing + condition cache
21.461 18.315 17.442 16.331 15.930
22.036 19.677 18.347 17.372 16.728
– 3.147 0.872 1.111 0.401
600/600 8/600 0/600 0/600 0/600
Same structure, runtime remeasurement + inference mode (final actor)
16.234 13.175
16.912 13.400
– 3.058
1/600 0/600
Final live callback-to-host
16.790
19.383
–
7/600
Composition / boundary
Times in ms: medians of three block quantiles (200 calls/block). Saved: preceding-row p50 decrease within one suite, not additive factor effects. Rules mark suite/boundary changes; Appendix D.1 reports factorial effects and ranges. Callback misses 7/600 deadlines: median feasibility, not hard real time.
(a) DDIM Steps - Timing
(b) Inference Acceleration Procedure
80
21.85
60 40 20 ms budget
20 0
20
stacked WBG
-3.54ms 18.31 selected implementation −2.39 ms
16
3
5
NS-DDIM steps
(d) DDIM Steps - Push Survival
10
P0
P1
P2
P3
13.18
C0 P0 Seq
21.46
21.23
20.67
20.52
C0 P0 Stack
18.31
17.97
18.58
17.75
C0 P1 Seq
19.08
19.32
19.67
18.68
C0 P1 Stack
17.44
16.23
17.18
16.33
C1 P0 Seq
20.99
20.38
20.72
21.11
C1 P0 Stack
18.19
17.79
17.89
17.42
C1 P1 Seq
19.18
19.26
19.18
18.80
C1 P1 Stack
16.63
16.10
16.39
15.93*
none
self
cross
all
P4
QKV packing
Optimization Precedure
(e) Inference Acceleration Contribution Stacked WBG axes DDIM schedule precomputation 1.672 [1.477, 1.698] All-QKV packing 0.775 [0.256, 0.799] Command-embedding cache 0.190 [0.151, 0.347] Inference mode (P3−P4) 3.058 [2.812, 3.492]
0
1
2
(f) Onboard In-loop Inference Timing 13.18
E0
2.845 [2.627, 2.868]
Push Survival
inference mode
−3.06 ms
12 2
16.23
15.93
optimized 13.18
1
(c) Inference Acceleration Ablation
24
WBG disabled WBG enabled
Actor p50 (ms)
Actor p50 (ms)
100
3
Median actor time saved (ms)
p95 13.40 13.84
E1
p95 16.22 16.79
E2
7/600 callbacks exceeded 20 ms
12
16
p95 19.38
20
Latency (ms)
(g) G1 Onboard Deploy (Jetson Orin NX)
Figure 7: Two-step inference meets the 50 Hz p95 budget on Jetson Orin and enables onboard G1 deployment. (a) Sampler-budget sweep. (b) Measured path through P0 –P4 : original TorchScript actor, cache-capable TorchScript re-export, factorially selected exact configuration, runtimesuite remeasurement, and LibTorch inference mode. (c) Exact two-forward implementation matrix. Within this panel, C0/C1 and P0/P1 denote condition caching and schedule precomputation disabled/enabled, while Seq/Stack denotes sequential/stacked WBG axes. (d) Push survival across the tested sampler budgets. (e) Matched implementation contributions, which are nonadditive. (f) The p50–p95 intervals for E0 (actor with fixed input), E1 (actor after new robot-state receipt and conversion), and E2 (complete callback through host-action production). Each primary endpoint uses three blocks of 200 retained calls; 593/600 E2 measurements finish within 20 ms, with p95 at 19.383 ms. (g) Qualitative frames from onboard G1 deployment with Jetson Orin NX. Thus the optimized export is feasible for 50 Hz onboard control at the measured p95, while the seven misses leave no hard-real-time guarantee.
9
Preprint
4.3
O NBOARD AND S IMULATION D EMONSTRATIONS
We evaluate the same conditioned export in simulation and on the physical Unitree G1–Jetson Orin NX stack. Figure 1(a,c) covers simulated joystick steering and semantic interpolation, while Figure 1(b,d) shows physical disturbance response and staged text commands. The (g) panel of Figure 7 is the direct evidence of the exported policy running onboard on the G1 with Jetson Orin NX; the Figure 1 sequences demonstrate functionality. The G1 remains upright without visible jitter or falls, and the scope and protocol are specified in Appendix C.3.
5
L IMITATIONS AND C ONCLUSION
PredActor uses joint state–action diffusion to make direct humanoid action generation steerable. Its internal future-state trajectory provides a target for test-time guidance, while proprioceptive history and task conditions support feedback-conditioned execution and learned text commands. In the evaluated simulation protocol, the selected policy reaches 15/15 destination targets and achieves 0.580 text retrieval versus 0.373 for conditional action diffusion, with similar observed disturbance survival. The result is a policy that combines predictive steering and direct execution without a separate motion-reference tracker at inference. On Jetson Orin, exact two-step sampling plus graph, structural, and runtime optimization reduces the actor median from 21.846 ms at P0 to 13.175 ms at P4 . The complete callback endpoint E2 reaches 16.790/19.383 ms at p50/p95, and 593/600 measurements complete within the 20 ms period. This establishes p95 timing feasibility for the evaluated policy configuration and isolates seven slight tail-latency overruns. Simulation demonstrations show joystick steering and semantic interpolation; qualitative Unitree G1 sequences show onboard text-command and disturbance-response execution without visible jitter or falls. Under the evaluated task, checkpoint, and hardware conditions, PredActor demonstrates steerable onboard humanoid control through predictive action diffusion. The present evidence covers selected checkpoints and task interfaces. Architecture comparisons vary complete policy configurations, so isolating the contribution of internal prediction requires further controlled ablations. The DAgger-style aggregation procedure provides supervision at learnervisited states; a matched behavioral study is needed to quantify its benefit. Broader physical trials with explicit success and recovery denominators, together with additional tail-latency characterization, are the next steps toward wider deployment. AI USE STATEMENT Generative AI tools assisted with conceptual framing, experiment-design review, literature discovery and summarization, manuscript editing, and code-based figures and scripts. We verified literature claims against primary sources, inspected generated code and artifacts, rebuilt the manuscript, and reviewed all AI-assisted content. Generative AI output was not used as experimental evidence. The authors take responsibility for the final text, claims, and artifacts.
R EFERENCES Anurag Ajay, Yilun Du, Abhi Gupta, Joshua B. Tenenbaum, Tommi S. Jaakkola, and Pulkit Agrawal. Is conditional generative modeling all you need for decision making? In International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id= sP1fo2K9DFG. Milo Carroll, Tianhu Peng, Lingfan Bao, Chengxu Zhou, and Zhibin Li. SCDP: Learning humanoid locomotion from partial observations via mixed-observation distillation, 2026. URL https: //arxiv.org/abs/2603.09574. Boyuan Chen, Diego Martı́ Monsó, Yilun Du, Max Simchowitz, Russ Tedrake, and Vincent Sitzmann. Diffusion forcing: Next-token prediction meets full-sequence diffusion. In Advances in Neural Information Processing Systems 37, NeurIPS 2024, pp. 24081–24125. Neural Information Processing Systems Foundation, Inc. (NeurIPS), 2024. doi: 10.52202/079017-0759. URL https://doi.org/10.52202/079017-0759.
10
Preprint
Maiyue Chen, Kaihui Wang, Bo Zhang, Xihan Ma, Zhiyuan Yang, Yi Ren, Qijun Huang, Zihao Zhu, Yucheng Wang, and Zhizhong Su. HoloMotion-1 technical report, 2026a. URL https: //arxiv.org/abs/2605.15336. Xiao Chen, Weishuai Zeng, Xiaojie Niu, Zirui Wang, Jianan Li, Huayi Wang, Furui Xu, Jiahe Chen, Weixiang Zhong, Lihe Ding, Kailin Li, Jiangmiao Pang, Tai Wang, Tianfan Xue, and Jingbo Wang. ReactiveBFM: Reactive closed-loop motion planning towards universal humanoid wholebody control, 2026b. URL https://arxiv.org/abs/2606.30362. Cheng Chi, Siyuan Feng, Yilun Du, Zhenjia Xu, Eric Cousineau, Benjamin Burchfiel, and Shuran Song. Diffusion policy: Visuomotor policy learning via action diffusion. In Robotics: Science and Systems XIX, RSS2023. Robotics: Science and Systems Foundation, July 2023. doi: 10. 15607/rss.2023.xix.026. URL https://doi.org/10.15607/rss.2023.xix.026. Zhaoyuan Gu, Yipu Chen, Zimeng Chai, Alfred Cueva, Thong Nguyen, Yifan Wu, Huishu Xue, Minji Kim, Isaac Legene, Fukang Liu, KyoungMok Kim, Ayan Barula, Yongxin Chen, and Ye Zhao. REFINE-DP: Diffusion policy fine-tuning for humanoid loco-manipulation via reinforcement learning. IEEE Robotics and Automation Letters, 11(10):11617–11624, Oct 2026. doi: 10.1109/lra.2026.3723742. URL https://doi.org/10.1109/lra.2026.3723742. Tairan He, Zhengyi Luo, Wenli Xiao, Chong Zhang, Kris Kitani, Changliu Liu, and Guanya Shi. Learning human-to-humanoid real-time whole-body teleoperation. In 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pp. 8944–8951. IEEE, Oct 2024. doi: 10.1109/iros58592.2024.10801984. URL https://doi.org/10.1109/ iros58592.2024.10801984. Tairan He, Zhengyi Luo, Xialin He, Wenli Xiao, Chong Zhang, Weinan Zhang, Kris M. Kitani, Changliu Liu, and Guanya Shi. OmniH2O: Universal and dexterous human-to-humanoid wholebody teleoperation and learning. In Proceedings of the 8th Conference on Robot Learning, volume 270 of Proceedings of Machine Learning Research, pp. 1516–1540. PMLR, 06–09 Nov 2025. URL https://proceedings.mlr.press/v270/he25b.html. Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. In NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications, 2021. doi: 10.48550/arxiv.2207.12598. URL https://openreview.net/forum?id=qw8AKxfYbI. Sigmund H. Høeg, Yilun Du, and Olav Egeland. Streaming diffusion policy: Fast policy synthesis with variable noise diffusion models, 2024. URL https://arxiv.org/abs/2406. 04806. Xiaoyu Huang, Yufeng Chi, Ruofeng Wang, Zhongyu Li, Xue Bin Peng, Sophia Shao, Borivoje Nikolic, and Koushil Sreenath. DiffuseLoco: Real-time legged locomotion control with diffusion from offline datasets. In Proceedings of the 8th Conference on Robot Learning, volume 270 of Proceedings of Machine Learning Research, pp. 1567–1589. PMLR, 06–09 Nov 2025a. URL https://proceedings.mlr.press/v270/huang25a.html. Xiaoyu Huang, Takara Truong, Yunbo Zhang, Fangzhou Yu, Jean Pierre Sleiman, Jessica Hodgins, Koushil Sreenath, and Farbod Farshidian. Diffuse-CLoC: Guided diffusion for physics-based character look-ahead control. ACM Transactions on Graphics, 44(4):1–12, July 2025b. doi: 10.1145/3731206. URL https://doi.org/10.1145/3731206. Michael Janner, Yilun Du, Joshua B. Tenenbaum, and Sergey Levine. Planning with diffusion for flexible behavior synthesis. In Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pp. 9902–9915. PMLR, 17–23 Jul 2022. URL https://proceedings.mlr.press/v162/janner22a.html. Zhengyao Jiang, Yingchen Xu, Nolan Wagener, Yicheng Luo, Michael Janner, Edward Grefenstette, Tim Rocktäschel, and Yuandong Tian. H-GAP: Humanoid control with a generalist planner. In International Conference on Learning Representations, volume 2024, pp. 17774–17791, 2024. URL https://openreview.net/forum?id=LYG6tBlEX0.
11
Preprint
Michael Kelly, Chelsea Sidrane, Katherine Driggs-Campbell, and Mykel J. Kochenderfer. HGDAgger: Interactive imitation learning with human experts. In 2019 International Conference on Robotics and Automation (ICRA), pp. 8077–8083. IEEE, May 2019. doi: 10.1109/icra.2019. 8793698. URL https://doi.org/10.1109/icra.2019.8793698. Michael Laskey, Jonathan Lee, Roy Fox, Anca Dragan, and Ken Goldberg. DART: Noise injection for robust imitation learning. In Proceedings of the 1st Annual Conference on Robot Learning, volume 78 of Proceedings of Machine Learning Research, pp. 143–156. PMLR, 13–15 Nov 2017. URL https://proceedings.mlr.press/v78/laskey17a.html. Yitang Li, Zhengyi Luo, Tonghe Zhang, Cunxi Dai, Anssi Kanervisto, Andrea Tirinzoni, Haoyang Weng, Kris Kitani, Mateusz Guzek, Ahmed Touati, Alessandro Lazaric, Matteo Pirotta, and Guanya Shi. BFM-Zero: A promptable behavioral foundation model for humanoid control using unsupervised reinforcement learning, 2025a. URL https://arxiv.org/abs/2511. 04131. Yixuan Li, Yutang Lin, Jieming Cui, Tengyu Liu, Wei Liang, Yixin Zhu, and Siyuan Huang. CLONE: Closed-loop whole-body humanoid teleoperation for long-horizon tasks, 2025b. URL https://arxiv.org/abs/2506.08931. Qiayuan Liao, Takara E. Truong, Xiaoyu Huang, Yuman Gao, Guy Tevet, Koushil Sreenath, and C. Karen Liu. BeyondMimic: From motion tracking to versatile humanoid control via guided diffusion. Science Robotics, 11(117), Aug 2026. doi: 10.1126/scirobotics.adx8924. URL https://doi.org/10.1126/scirobotics.adx8924. Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu. Dpm-solver: A fast ODE solver for diffusion probabilistic model sampling in around 10 steps. In Advances in Neural Information Processing Systems 35, NeurIPS 2022, pp. 5775–5787. Neural Information Processing Systems Foundation, Inc. (NeurIPS), 2022. doi: 10.52202/068431-0418. URL https://doi.org/10.52202/068431-0418. Zhengyi Luo, Jinkun Cao, Alexander Winkler, Kris Kitani, and Weipeng Xu. Perpetual humanoid control for real-time simulated avatars. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 10895–10904. IEEE, Oct 2023. doi: 10.1109/iccv51070.2023.01000. URL https://doi.org/10.1109/iccv51070.2023.01000. Zhengyi Luo, Ye Yuan, Tingwu Wang, Chenran Li, Fernando Castañeda, Sirui Chen, Zi-Ang Cao, Jiefeng Li, David Minor, Qingwei Ben, Jinhyung Park, David Sami, Zi Wang, Xingye Da, Runyu Ding, Cyrus Hogg, Lina Song, Edy Lim, Eugene Jeong, Tairan He, Haoru Xue, Wenli Xiao, Simon Yuen, Jan Kautz, Yan Chang, Umar Iqbal, Linxi “Jim” Fan, and Yuke Zhu. SONIC: Supersizing motion tracking for natural humanoid whole-body control. Science Robotics, 11 (117), Aug 2026. doi: 10.1126/scirobotics.aed4592. URL https://doi.org/10.1126/ scirobotics.aed4592. Allen Z. Ren, Justin Lidard, Lars L. Ankile, Anthony Simeonov, Pulkit Agrawal, Anirudha Majumdar, Benjamin Burchfiel, Hongkai Dai, and Max Simchowitz. Diffusion policy policy optimization. In International Conference on Learning Representations, volume 2025, pp. 77288–77329, 2025. URL https://openreview.net/forum?id=mEpqHvbD2h. RoboParty Lab Team. MimicLite: Efficient and effective general humanoid motion tracking. https://github.com/EGalahad/mimic-lite, 2026. Technical report: https:// github.com/Roboparty/MimicLite/blob/main/mimic-lite.pdf. Ralf Römer, Alexander von Rohr, and Angela Schoellig. Diffusion predictive control with constraints. In Proceedings of the 7th Annual Learning for Dynamics & Control Conference, volume 283 of Proceedings of Machine Learning Research, pp. 791–803. PMLR, 04–06 Jun 2025. URL https://proceedings.mlr.press/v283/romer25a.html. Stephane Ross, Geoffrey Gordon, and Drew Bagnell. A reduction of imitation learning and structured prediction to no-regret online learning. In Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, volume 15 of Proceedings of Machine Learning Research, pp. 627–635, Fort Lauderdale, FL, USA, 11–13 Apr 2011. PMLR. URL https://proceedings.mlr.press/v15/ross11a.html. 12
Preprint
Carmelo Sferrazza, Dun-Ming Huang, Xingyu Lin, Youngwoon Lee, and Pieter Abbeel. HumanoidBench: Simulated humanoid benchmark for whole-body locomotion and manipulation. In Robotics: Science and Systems XX, RSS2024. Robotics: Science and Systems Foundation, July 2024. doi: 10.15607/rss.2024.xx.061. URL https://doi.org/10.15607/rss.2024. xx.061. Yiyang Shao, Bike Zhang, Qiayuan Liao, Xiaoyu Huang, Yuman Gao, Yufeng Chi, Zhongyu Li, Sophia Shao, and Koushil Sreenath. LangWBC: Language-directed humanoid whole-body control via end-to-end learning. In Robotics: Science and Systems XXI, RSS2025. Robotics: Science and Systems Foundation, June 2025. doi: 10.15607/rss.2025.xxi.065. URL https: //doi.org/10.15607/rss.2025.xxi.065. Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In International Conference on Learning Representations, 2021. doi: 10.48550/arxiv.2010.02502. URL https://openreview.net/forum?id=St1giarCHLP. Yang Song, Prafulla Dhariwal, Mark Chen, and Ilya Sutskever. Consistency models. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pp. 32211–32252. PMLR, 23–29 Jul 2023. URL https: //proceedings.mlr.press/v202/song23a.html. Chen Tessler, Yunrong Guo, Ofir Nabati, Gal Chechik, and Xue Bin Peng. MaskedMimic: Unified physics-based character control through masked motion inpainting. ACM Transactions on Graphics, 43(6):1–21, Nov 2024. doi: 10.1145/3687951. URL https://doi.org/10. 1145/3687951. Chen Tessler, Yifeng Jiang, Xue Bin Peng, Erwin Coumans, Yi Shi, Haotian Zhang, Davis Rempe, Gal Chechik, and Sanja Fidler. ProtoMotions3: An open-source framework for humanoid simulation and control. https://github.com/NVLabs/ProtoMotions/, 2025. URL https://github.com/NVLabs/ProtoMotions. Guy Tevet, Brian Gordon, Amir Hertz, Amit H Bermano, and Daniel Cohen-Or. MotionCLIP: Exposing human motion generation to CLIP space. In European Conference on Computer Vision (ECCV), pp. 358–374. Springer, Springer, 2022. doi: 10.1007/978-3-031-20047-2 21. URL https://doi.org/10.1007/978-3-031-20047-2_21. Guy Tevet, Sigal Raab, Brian Gordon, Yonatan Shafir, Daniel Cohen-Or, and Amit H. Bermano. Human motion diffusion model. In International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=SJ1kSyO2jwu. Guy Tevet, Sigal Raab, Setareh Cohan, Daniele Reda, Zhengyi Luo, Xue Bin Peng, Amit H. Bermano, and Michiel van de Panne. CLoSD: Closing the loop between simulation and diffusion for multi-task character control. In International Conference on Learning Representations, volume 2025, pp. 46506–46520, 2025. URL https://openreview.net/forum? id=pZISppZSTv. Takara Everest Truong, Michael Piseno, Zhaoming Xie, and C. Karen Liu. PDP: Physics-based character animation via diffusion policy. In SIGGRAPH Asia 2024 Conference Papers, SA ’24, pp. 1–10. ACM, Dec 2024. doi: 10.1145/3680528.3687683. URL https://doi.org/10. 1145/3680528.3687683. Tingwu Wang, Olivier Dionne, Michael De Ruyter, David Minor, Davis Rempe, Kaifeng Zhao, Mathis Petrovich, Ye Yuan, Chenran Li, Zhengyi Luo, Brian Robison, Xavier Blackwell, Bernardo Antoniazzi, Xue Bin Peng, Yuke Zhu, and Simon Yuen. MotionBricks: Scalable realtime motions with modular latent generative model and smart primitives. ACM Transactions on Graphics, 45(4):1–22, July 2026. doi: 10.1145/3811334. URL https://doi.org/10. 1145/3811334. Zhendong Wang, Max Li, Ajay Mandlekar, Zhenjia Xu, Jiaojiao Fan, Yashraj Narang, Linxi Fan, Yuke Zhu, Yogesh Balaji, Mingyuan Zhou, Ming-Yu Liu, and Yu Zeng. One-step diffusion policy: Fast visuomotor policies via diffusion distillation. In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, 13
Preprint
pp. 63399–63416. PMLR, 13–19 Jul 2025. URL https://proceedings.mlr.press/ v267/wang25ba.html. Weiji Xie, Jiakun Zheng, Jinrui Han, Jiyuan Shi, Weinan Zhang, Chenjia Bai, and Xuelong Li. TextOp: Real-time interactive text-driven humanoid robot motion generation and control, 2026. URL https://arxiv.org/abs/2602.07439. Yanjie Ze, Zixuan Chen, João Pedro Araújo, Zi-ang Cao, Xue Bin Peng, Jiajun Wu, and C. Karen Liu. TWIST: Teleoperated whole-body imitation system. In Proceedings of The 9th Conference on Robot Learning, volume 305 of Proceedings of Machine Learning Research, pp. 2143–2154. PMLR, 27–30 Sep 2025. URL https://proceedings.mlr.press/v305/ ze25a.html. Weishuai Zeng, Shunlin Lu, Kangning Yin, Xiaojie Niu, Minyue Dai, Jingbo Wang, and Jiangmiao Pang. Behavior foundation model for humanoid robots, 2025. URL https://arxiv.org/ abs/2509.13780. Jingyan Zhang, Han Liang, Ruichi Zhang, Bin Li, Juze Zhang, Xin Chen, Jingya Wang, Lan Xu, and Jingyi Yu. SCRIPT: Scalable diffusion policy with multi-stage training for language-driven physics-based humanoid control, 2026. URL https://arxiv.org/abs/2605.22894. Zhikai Zhang, Jun Guo, Chao Chen, Jilong Wang, Chenghuai Lin, Yunrui Lian, Han Xue, Zhenrong Wang, Maoqi Liu, Jiangran Lyu, Huaping Liu, He Wang, and Li Yi. Track any motions under any disturbances, 2025. URL https://arxiv.org/abs/2509.13833. Zihan Zhang, Richard Liu, Rana Hanocka, and Kfir Aberman. TEDi: Temporally-entangled diffusion for long-term motion synthesis. In Special Interest Group on Computer Graphics and Interactive Techniques Conference Conference Papers, SIGGRAPH ’24, pp. 1–11. ACM, July 2024. doi: 10.1145/3641519.3657515. URL https://doi.org/10.1145/3641519. 3657515. Kaifeng Zhao, Mathis Petrovich, Haotian Zhang, Tingwu Wang, Siyu Tang, and Davis Rempe. ARDY: Autoregressive diffusion with hybrid representation for interactive human motion generation, 2026. URL https://arxiv.org/abs/2607.08741. Tony Z. Zhao, Vikash Kumar, Sergey Levine, and Chelsea Finn. Learning fine-grained bimanual manipulation with low-cost hardware. In Robotics: Science and Systems XIX, RSS2023. Robotics: Science and Systems Foundation, July 2023. doi: 10.15607/rss.2023.xix.016. URL https: //doi.org/10.15607/rss.2023.xix.016.
A
E XTENDED R ELATED -W ORK C OVERAGE
A.1
D EPLOYMENT H ARDWARE AND T IMING S OURCES
Table 4 expands the main-text comparison with the full capability matrix, inference hardware, and timing boundaries. It separates inference hardware from training hardware and robot control rates from generator throughput. MotionBricks reports Jetson Orin deployment on G1 with 10 Hz replanning and 5 ms generator latency (Section 7.4), without specifying the Orin variant or a complete-loop latency (Wang et al., 2026). Its resource-sharing discussion suggests that the generator and tracker share onboard compute, although their individual hardware assignments and the tracker rate are not explicitly specified. The table’s 50 Hz tracker entry is sourced from its adopted SONIC controller (Section 3.5), not an independently reported MotionBricks measurement (Luo et al., 2026). ARDY’s Section 4.2 reports an RTX 4090 interactive generator with 33/63 ms latency for 4/10 diffusion steps; these measurements do not characterize an ARDY–SONIC robot control loop (Zhao et al., 2026). The two generator–tracker compositions have separate capability rows. MotionBricks steers motion through keyframes and root-trajectory constraints supplied by smart primitives (Sections 5–6). Its masked-token generator does not report conditional–null prediction mixing (CFG) or sampling-time objective gradients (CG); the CFG discussion in Section 7.2 concerns comparison baselines. These entries do not imply a lack of controllability. ARDY explicitly supports text and spatial conditioning with CFG (Section 3.5); these are marked as generator-level capabilities in the SONIC composition. 14
Preprint
Table 4: Policy capabilities and deployment configurations of generative control systems. Top: representation and task interfaces. Bottom: inference hardware, location, and reported execution rates; component timings are not complete-loop measurements. Model settings Method
Paradigm
MotionBricks + SONIC ARDY + SONIC TextOp ReactiveBFM CLoSD DiffuseLoco Diffuse-CLoC BeyondMimic SCDP SCRIPT
Task capability
Deployment
Proprio.
Rolling reuse
CFG
CG
Text
Joy/Goal
Robot
Generator + tracker Open-loop track Open-loop track Closed-loop track Closed-loop track Action diffusion Joint diffusion Latent joint diffusion + dec. Joint diffusion Joint diffusion
⋄ ⋄ ⋄ ⋄ – ✓ – ⋄ ✓ –
– – – – – – ✓ – – –
– ⋄ ✓ – ✓ – – – – ✓
– – – – – – ✓ ✓ – –
– ⋄ ✓ ✓ ✓ – – – – ✓
✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓
✓ – ✓ ✓ – ✓ – ✓ ✓ –
PredActor
Joint diffusion
✓
✓
✓
✓
✓
✓
✓
Method
Inference CPU
Inference GPU
Compute location
Rate / timing (reported boundary)
MotionBricks + SONIC
Orin SoCa
Orin integrateda
G1 / Jetson Orin
ARDY + SONIC TextOp
– i7-13700c
RTX 4090b RTX 4090c
ReactiveBFM
i9-13900K
RTX 4090
Workstationb External generator; onboard tracker Offboard workstation
50 Hz tracker∗ ; 10 Hz replanning; 5 ms generator 33 / 63 ms generator (4 / 10 steps)b 6.25 Hz generator; 50 Hz tracker
CLoSD DiffuseLoco Diffuse-CLoC BeyondMimic
– i7-13700H – – (CPU decoder)
– RTX 4060 Mobile RTX 4060 RTX 4060 Mobile
Simulation only Onboard Go1 mini-PC Simulation only Onboard mini-PC (v4)
SCDP SCRIPT
– –
RTX 5090 –
Remote workstation Simulation only
50 Hz controller; async planner (19.3 ms) 11.4 ms for a 2 s motion plan 30 Hz control target; 116.5 Hz inference 77.41 Hz inference (selected profile) 25 Hz control; async diffusion about 20 ms 50 Hz control; 105 Hz inference Physical control: N/A
PredActor
Orin NX SoC
Orin NX integrated Fully onboard G1
50 Hz control; callback p95 19.383 ms
Symbols: ✓: supported; ⋄: partial/modular; ×: explicitly unsupported; –: unreported/unverified; N/A: not applicable. Capability definitions: Diffusion paradigms follow Figure 2; MotionBricks uses a non-diffusion generator + tracker. Proprio. means proprioceptive input; dec. means decoder; Robot denotes physical deployment. Proprioceptive input permits task commands but excludes privileged physical-state input; ⋄ marks tracker-level input in generator–tracker systems. Rolling reuse requires explicit reuse of overlapping state/action predictions or their partially denoised representations, not just replanning. CFG requires conditional–null prediction mixing; ⋄ marks ARDY generator-level CFG/text support without separate validation in the SONIC composition. Hardware notes: ∗ 50 Hz is sourced from MotionBricks’ adopted SONIC tracker, not independently reported for the composition. a Orin variant, CPU model, and full-stack timing are unreported. b ARDY generator-demo benchmark; robot-stack hardware and rate are unreported. c TextOp external workstation; tracker CPU is unreported. Inference hardware is not inferred from training hardware. Timing boundary: PredActor’s callback covers received robot-state messages through host-action production; 593/600 calls meet 20 ms. The timing benchmark does not actuate the robot. Component timings are not full-loop measurements. Sources: Appendix A.1.
15
Preprint
The ballet and locomotion videos under Humanoid Robot Control on its official project page show simulated G1 tracking, not physical deployment. We therefore leave its physical-deployment entry unmarked. Both systems support spatial steering, which does not by itself establish CG. Their proprioceptive-input marks refer to the tracker, not robot feedback to the generator. TextOp runs its generator on an external i7-13700/RTX 4090 workstation at 6.25 Hz and its tracker onboard at 50 Hz (Section III-E and Appendix B-F) (Xie et al., 2026). ReactiveBFM uses an offboard i9-13900K/RTX 4090 workstation, a 50 Hz controller, and an asynchronous planner; the reported 19.3 ms planner and 5.9 ms controller timings are component measurements (Deployment and Appendix B.5) (Chen et al., 2026b). SCDP uses a remote RTX 5090 workstation for 50 Hz G1 control and reports 105 Hz inference throughput (Section II-D) (Carroll et al., 2026). For BeyondMimic, we use arXiv v4, Supplement S4 (Diffusion): the diffusion-stage setup uses 25 Hz rather than the standalone tracker’s 50 Hz, with onboard RTX 4060 Mobile inference and a synchronous CPU VAE decoder (Liao et al., 2026). Its approximately 20 ms asynchronous diffusion latency is not a complete-loop measurement. Earlier versions (v1, v2, and v3, Section VI-B) instead describe offboard inference without stating this 25 Hz setting. The table follows v4, not the earlier deployment description. We classify its latent state–action model as joint diffusion, rather than a kinematic generator followed by a reference tracker. DiffuseLoco reports onboard Go1 deployment on an i7-13700H/RTX 4060 Mobile mini-PC; its 30 Hz control requirement differs from the 116.5 Hz TensorRT inference benchmark (Appendix E and Figure 16) (Huang et al., 2025a). Diffuse-CLoC evaluates simulated characters and reports RTX 4060 inference, with 77.41 Hz for the selected profile (Section V-B and Table II) (Huang et al., 2025b). CLoSD’s 11.4 ms measurement generates a two-second motion plan, not a physical-robot callback (Section 4.1) (Tevet et al., 2025); SCRIPT evaluates simulated humanoids rather than onboard robot execution (Section 5.1) (Zhang et al., 2026). Unreported hardware and rates are marked “–”, even when training hardware is stated; N/A denotes physical control rates for simulation-only systems. PredActor combines proprioceptive conditioning, CG/CFG, direct joint-action execution, and fully onboard G1 inference on Jetson Orin NX. Its measured callback runs from received robot-state messages to host-action production, with p50/p95 of 16.790/19.383 ms and 593/600 calls within 20 ms. Physical execution is demonstrated separately; the timing benchmark issues no actuation commands. This distinguishes the measured boundary without claiming a hardware-normalized speed advantage over other systems. A.2
M OTION -C ONTROL L ANDSCAPE
The literature spans several control paradigms that are easy to conflate. A motion tracker receives a reference trajectory or motion representation and produces physically executed actions. PHC established scalable perpetual tracking and recovery for simulated avatars Luo et al. (2023). Recent robot trackers emphasize different bottlenecks: MimicLite reduces the reported training cost of a deployable G1 tracker RoboParty Lab Team (2026); SONIC, released as GEAR-SONIC, scales model capacity, motion data, and compute while exposing several downstream motion interfaces Luo et al. (2026); HoloMotion-1 uses a hybrid motion corpus and temporal mixture-of-experts model for zeroshot whole-body tracking Chen et al. (2026a); and Any2Track couples a general tracker with a history-conditioned dynamics adapter for disturbances Zhang et al. (2025). ProtoMotions3 is instead an open simulation, learning, and deployment framework that can train tracking policies; it is not one additional controller architecture Tessler et al. (2025). These trackers and tools strengthen the execution side of generator–tracker stacks, but do not themselves select a long-horizon task behavior without a reference or higher-level input. A.3
R EACTIVE P LANNER –T RACKER C OMPARISONS
ReactiveBFM distinguishes open-loop cascades from closed-loop compositions with physical-state feedback; those labels describe the evaluated composition and do not imply that every component replans from physical state (Chen et al., 2026b). A second family broadens the policy interface. BFM-Zero learns a shared latent space for motion, goal, and reward prompts using unsupervised reinforcement learning Li et al. (2025a). 16
Preprint
BFM4Humanoid combines masked online distillation with a conditional variational autoencoder to support multiple whole-body control modes Zeng et al. (2025). LangWBC is a non-diffusion comparison: reinforcement learning, policy distillation, and a conditional variational autoencoder yield one language-conditioned policy that outputs actions directly Shao et al. (2025). These are unified, promptable, or direct controllers rather than a kinematic planner followed by an independently trained tracker. They demonstrate broader task interfaces, but their papers do not constitute matched tests of PredActor’s state–action co-diffusion or rolling runtime design. Teleoperation systems occupy another role. H2O, OmniH2O, TWIST, and CLONE map live human inputs to whole-body robot control, with CLONE specifically adding closed-loop global-position correction for long-horizon teleoperation He et al. (2024; 2025); Ze et al. (2025); Li et al. (2025b). They are important sources of control interfaces and demonstrations, but are not autonomous motion generators in the sense used here. HumanoidBench is likewise an evaluation suite for locomotion and whole-body manipulation, not a controller Sferrazza et al. (2024). This taxonomy motivates comparisons along reference dependence, task conditioning, generated variables, deploy-time observations, guidance, and closed-loop compute rather than a single overloaded “end-to-end” label. A.4
T RAJECTORY G ENERATIVE M ODELS FOR C ONTROL
Diffusion control methods differ in what they generate. Action-only policies denoise action chunks conditioned on observations Chi et al. (2023); this formulation extends to offline legged locomotion Huang et al. (2025a) and physics-based character control Truong et al. (2024). Joint trajectory models also represent future states: Diffuser performs return- or constraint-guided planning Janner et al. (2022), while Decision Diffuser uses conditional state-sequence generation with inverse dynamics to recover actions Ajay et al. (2023). Outside diffusion, H-GAP learns a discrete latent state–action trajectory model and uses model-predictive planning for generalist simulated humanoid control Jiang et al. (2024). PredActor follows the joint state–action diffusion route: future states support guidance within the model that generates executable actions. Recent humanoid systems provide closer comparisons. BeyondMimic guides a latent diffusion model of human motion primitives toward test-time objectives and transfers the resulting controller to hardware Liao et al. (2026). SCDP generates interleaved future states and actions from onboard observation histories, using privileged future trajectories for supervision and partial observations at deployment Carroll et al. (2026). SCRIPT jointly represents action, state, and text streams in a diffusion transformer and adds reinforcement-learning post-training for language-driven simulated humanoids Zhang et al. (2026). Within this landscape, PredActor combines proprioceptive-history conditioning with learned-command CFG and future-state CG, then implements these interfaces within an onboard control loop. Generated variables, policy inputs, guidance mechanisms, and hardware evidence therefore form separate comparison axes in Table 1. A.5
E FFICIENT AND T EMPORALLY C ONSISTENT D IFFUSION I NFERENCE
Real-time control constrains the cost of iterative sampling. DDIM introduces a non-Markovian implicit process that permits deterministic subsampling without retraining Song et al. (2021); DPMSolver uses dedicated high-order integration for diffusion ODEs Lu et al. (2022); and consistency models learn direct mappings from noise to data Song et al. (2023). Robot-specific distillation similarly trades extra training for fewer evaluations: OneDP distills a pretrained diffusion policy into a one-step generator Wang et al. (2025). These approaches accelerate or replace the sampler, but do not specify how successive control horizons should share computation. Temporal reuse addresses that second issue. Streaming Diffusion Policy carries forward a partially denoised action trajectory whose near-term entries are cleaner than its tail Høeg et al. (2024). TEDi entangles diffusion time with motion time in a recursively advanced buffer for long-horizon motion synthesis Zhang et al. (2024), while Diffusion Forcing assigns per-token noise levels to combine next-token prediction with full-sequence diffusion Chen et al. (2024). Action chunking with temporal ensembling smooths overlapping predictions Zhao et al. (2023), and constrained diffusion adds feasibility projections at sampling time Römer et al. (2025). PredActor’s (K, h) noise matrix provides one implementation of per-position schedules, DDIM jumps, and rolling-buffer reuse. The individual acceleration and streaming mechanisms are not introduced here.
17
Preprint
A.6
I MITATION L EARNING , AGGREGATION , AND P OLICY I MPROVEMENT
Behavior cloning suffers from covariate shift because learner actions change the future state distribution. DAgger queries the expert on learner-induced states and aggregates the resulting labels Ross et al. (2011). DART instead perturbs demonstrations to collect recovery behavior without executing the novice during collection Laskey et al. (2017), whereas HG-DAgger gates human intervention using a learned uncertainty threshold Kelly et al. (2019). These methods separate three design choices that are sometimes conflated: who visits the state, who supplies the label, and whether the collected sample is appended to replay. Diffusion policies can also be improved with reward optimization rather than label aggregation. DPPO supplies a policy-gradient recipe for fine-tuning diffusion policies Ren et al. (2025), and REFINE-DP jointly fine-tunes a diffusion planner and low-level controller for humanoid locomanipulation Gu et al. (2026). PredActor’s DAgger extension belongs to the aggregation family: learner-visited rollout origins seed isolated teacher continuations, whose trajectory labels enter combined replay for continued diffusion training. This procedure is distinct from offline noise injection and reinforcement-learning fine-tuning. A.7
C ONDITIONING AND H UMANOID C ONTROL I NTERFACES
Humanoid objectives may be expressed as velocity commands, sparse body targets, motion clips, language, or teleoperation. MotionCLIP aligns text and human motion embeddings Tevet et al. (2022); CLoSD couples text-conditioned motion diffusion to physics-based control Tevet et al. (2025); and TextOp streams text changes through a motion generator and a low-level tracking policy Xie et al. (2026). MaskedMimic treats diverse partial motion descriptions—including text, keyframes, objects, and paths—as masks for one physics-based controller Tessler et al. (2024). Whole-body teleoperation systems such as H2O, OmniH2O, and TWIST instead map live human motion to humanoid control He et al. (2024; 2025); Ze et al. (2025). PredActor exposes each conditioning source through a shared runtime interface. Text embeddings, joystick commands, semantic interpolation, and whole-body guidance can therefore reuse the same exported policy interface. The experiments evaluate behavioral execution and control-period timing separately from interface availability.
B
E XTENDED M ETHOD AND S YSTEM D ETAILS
The following details separate PredActor’s predictive representation from its training and execution machinery. Future states are internal variables for guidance, proprioceptive histories supply policy inputs, and selected actions are issued to the joint controller. Teacher supervision, rolling schedules, and runtime optimizations implement this policy without introducing a separate motion-generation and tracking hierarchy at deployment. B.1 B.1.1
P REDICTIVE S TATE –ACTION R EPRESENTATION P ROPRIOCEPTIVE INPUTS AND INTERNAL FUTURE STATES
Table 5 separates the deployable observation from the physical state predicted by PredActor. The selected policy uses the 93-dimensional observation; the broader 96-dimensional profile additionally includes base linear velocity. The 192-dimensional physical state is mapped by the fixed emphasis projection to a 384-dimensional model representation. The action a(t) contains 29 joint-controller commands, converted to PD targets by the runtime. Optional task context z encodes the requested task; the text-conditioned setting uses a 512-dimensional semantic embedding. Normalization statistics are fitted on the training split and fixed at test time. The model directly predicts clean trajectories and is optimized with reconstruction losses on both streams. PredActor does not contain a separate deterministic observation-to-state module. In the evaluated setting, deployment initializes the entire state stream from Gaussian noise and reconstructs it jointly with actions through conditional co-diffusion. The reconstructed state trajectory remains internal; only an action is sent to the robot. These predictions are used for guidance, not
18
Preprint
Table 5: Observation and physical-state components for the selected G1 profile. The optional baselinear-velocity row produces the broader 96-dimensional observation and is excluded from the selected 93-dimensional input. The emphasis projection changes representation dimension, not the underlying physical state. Component
Dimension
Description
Deployable observation ot Joint positions Joint velocities Projected gravity Base angular velocity Previous action Selected observation total Base linear velocity
29 29 3 3 29 93 +3
Robot joint coordinates Robot joint velocities Gravity in the base frame IMU angular velocity Last joint command Used by the evaluated profile Optional; yields the 96-D profile
Physical state s(t) Local body positions Local body linear velocities Root position Root orientation Root linear velocity Root angular velocity Physical-state total Emphasis-projected state
90 90 3 3 3 3 192 384
30 bodies × 3 coordinates 30 bodies × 3 velocities Local character frame Local rotation representation Local character frame Local character frame dphys s Model representation ds
presented as a separately validated full-body state estimator. Onboard signal processing remains part of observation construction. B.1.2
I NTERLEAVED CONDITIONAL TRANSFORMER
State and action inputs are projected separately, concatenated with sinusoidal embeddings of their tok tok tok respective noise levels, and interleaved as [stok 1 , a1 , . . . , sh , ah ]. Learned positional embeddings retain horizon order. A transformer decoder applies self-attention over this sequence and cross-attention to encoded condition tokens. In the reference attention mask, state queries attend to all state tokens but not action tokens; action queries attend causally to state and action tokens. Separate output heads predict clean states and clean actions. The current reference backbone uses two decoder layers, four attention heads, and a 256-dimensional embedding; these are configuration choices rather than architectural requirements. A fixed, profile-driven state emphasis transform E can be applied before the backbone and its pseudoinverse E † after prediction. The evaluated transform concatenates a weighted symmetric random projection with an identity component. It is specified before training and is not learned by gradient descent. B.2
DATASET C ONSTRUCTION AND DAGGER T RAINING
B.2.1
K INEMATIC LIBRARY AND INITIAL ROLLOUTS
Figure 4(a) combines prompted motion generation and annotated motion sources into a common kinematic library with task labels. A tracking teacher executes these references under observation noise, action noise, and external perturbations. Asynchronous rollouts store synchronized observation histories, future body states, teacher actions, and task context to form the initial offline dataset D0 . Complete rollout trajectories are split into training and evaluation sets before extracting fixed-length windows; normalization statistics are fitted only on the training split. Thus, no rollout contributes windows to both partitions. The teacher tracker is a training-time data source, not a component of PredActor’s inference path.
19
Preprint
B.2.2
T EACHER TRAJECTORY SHOOTING AND AGGREGATION
The aggregation update is given in Equation 2. The teacher continuation, not the student’s future actions, supplies the target trajectory. This extends the training distribution to policy-visited states without changing the diffusion objective or introducing reward optimization. The horizon-teacher procedure in Figure 4(b) provides joint trajectory supervision for Equation 2. At round j, a frozen student advances the live simulator. A rollout origin ξ contains the simulator state, student observation/action history, and motion reference phase needed to start an aligned teacher continuation. The tracking teacher is queried in a separate simulation branch with the same origin and reference. State and random-number-generator checks ensure that the query does not advance or alter the live student trajectory. Thus teacher actions label the branch; they are not interventions in the live rollout. The window constructor WE retains the observed student-history prefix and appends the teacher continuation. Past actions are the student’s executed actions paired with their pre-transition states; the action at the branch origin and subsequent actions come from the teacher. State targets follow the same temporal alignment. Task labels are aligned with the reference, and incomplete continuations are marked so that only complete training windows enter replay. Stored-window prefix lengths are configuration details and do not change the horizon-position convention i = 1, . . . , h. Each aggregate is used to continue diffusion reconstruction training before the next collection round. B.2.3
J OINT RECONSTRUCTION OBJECTIVE
Equation 3 is used for both initial training and aggregated replay. Its expectation covers windows from D, sampled noise, and diffusion levels. The action-difference term compares consecutive predicted and target actions, while the state reconstruction term supervises the internal future-state trajectory used for guidance. Aggregation changes the replay distribution without replacing this reconstruction objective. B.3
TASK AND D EPLOYMENT I NTERFACE
Task-specific providers inject raw condition keys on every step, and a composite provider distributes updates to its children. A condition composer then assembles the history consumed by the policy. This division allows observation-only and semantic/command-conditioned actors to share the same policy interface without hard-coding a particular input source. The exported artifact contains a TorchScript backbone, diffusion schedule, normalizer statistics, actor configuration, and deployment configuration. The C++ runtime normalizes observations and conditions, runs the actor, unnormalizes the action trajectory, and selects index N − 1 as the PD target. An asynchronous UDP receiver latches the latest external condition under a mutex, so condition transport is decoupled from the control loop. A Passive–Ready–Running state machine gates policy actions and falls back to Passive on fall or torque faults. B.4
E VALUATED D EPLOYMENT P ROFILE
The evaluated G1 Orin profile is a batch-one FP32 conditioned PredActor policy with two-step deterministic DDIM, packed self- and cross-attention projections, cached command embeddings, a precomputed schedule, stacked-axis WBG, eager finite checks, and LibTorch inference mode. For each level ti , the reduced backbone predicts clean action and state trajectories once, (â0i , ŝ0i ) = Fθ (ai , si , ti , cenc ).
(5)
When guidance is active, the same state prediction is adjusted by an analytic whole-body cost and reused in the DDIM update, s̃0i = ŝ0i − G(ŝ0i ; uvx , uvy , uvz ),
(ai−1 , si−1 ) = D(ai , si , â0i , s̃0i ; ti , ti−1 ).
(6)
The same clean state prediction is reused by guidance, so the evaluated path performs one backbone evaluation per diffusion step, or two in total. The older duplicate-backbone implementation executes four forwards and changes the guidance equation; it is not an exact acceleration baseline. The actor also performs encoding, guidance, scheduler updates, and boundary checks, so operation counts do 20
Preprint
not determine complete-loop speedups. Actor and callback timings, their acceptance thresholds, and the nonadditive component diagnostics are reported in Section 4.2 and Appendix D.1.
C
D ETAILED E XPERIMENTAL P ROTOCOL
C.1
O UTCOME D EFINITIONS AND M ISSINGNESS
Survival and recovery. Architecture-study push survival is the observed fraction of the 650-step horizon, so an early terminal contributes its observed duration rather than a binary zero; it is not the fraction of cells that finish. A terminal is triggered by root height below 0.55 m, absolute pelvis tilt at least 0.3 rad, or base angular speed at least 1 rad/s. Strict recovery requires 25 sustained post-force steps inside all three gates. In the architecture study, early failures remain in every denominator. Destination and text outcomes. Destination arrival is first entry within 0.6 m of the target. Final and minimum error use horizontal root-to-target distance. Text retrieval is raw top-1 MotionCLIP prompt retrieval over complete windows beginning at least 3 s after command activation. A trial may therefore contribute to survival but lack an eligible retrieval window. Generated references outside the tracker converter’s accepted joint range are retained as non-evaluable rather than assigned failure or zero. Timing. The timing endpoints in Figure 7 are E0 (actor execution with a fixed input tensor), E1 (receipt and conversion of a new robot state followed by actor execution), and E2 (the complete callback through production of the host action tensor). These boundaries are measured directly; their percentiles are not added. The 20 ms line is the 50 Hz target used for percentile and deadlineattainment reporting. C.2
A RCHITECTURE P ROTOCOL AND C OVERAGE
The five diffusion policies use the same 50-epoch budget, diffusion-loss formulation, and 7,600episode dataset. The dataset contains 50 teacher-rollout episodes for each of 152 G1-retargeted motions from the ACCAD subset of AMASS. The comparison preserves each method’s architecture, inference procedure, conditioning mechanism, and method-specific settings. One frozen checkpoint is evaluated per policy, so the reported uncertainty measures rollout variation; training-seed variation lies outside the reported intervals. ARDY+SONIC remains a separately trained generator–tracker baseline outside this shared diffusion-training contract. The architecture study uses one calibrated MuJoCo plant with whole-body guidance disabled. Push evaluation applies 100, 200, 300, 400, or 500 N horizontally for 0.5 s during walking. Each magnitude uses the same 20 paired direction/seed draws across policies, giving 100 requested cells per policy and 700 total. Destination following uses three target draws for each of five seeds and runs for at most 20 s. Text evaluation crosses nine prompts with the five seeds. Actor latency is synchronized on one exclusive RTX 4090, measured after warm-up, and reported per action from a fixed policy input repeated to batch five. Basic walking jerk uses the paired pre-push nominal segment. Table 6: Architecture-study coverage. Counts are evaluated/requested cells. Conditional subsets are not ranked against complete coverage. Policy
Push
Conditioned action diffusion Joint diffusion Conditioned joint diffusion PredActor w/o TextCFG Conditioned PredActor (checkpoint A) PredActor (selected checkpoint) ARDY+SONIC
100/100 100/100 100/100 100/100 100/100 100/100 100/100
21
Destination Text survival 0/15 15/15 15/15 15/15 15/15 15/15 2/15
45/45 0/45 45/45 0/45 45/45 45/45 23/45
Preprint
20 paired direction/seed clusters per force | ribbons: 95% paired-cluster bootstrap CI
a Survival duration fraction
b Sustained recovery fraction
1.00
Fraction
0.75
0.50
0.25
0.00 100
200
300
400
500
100
200
Push magnitude (N) Action Diff. + Cond. Joint Diff.
300
400
500
Push magnitude (N) Joint Diff. + Cond. PRDP
Cond. PRDP (checkpoint A) HoffMan (PRDP)
ARDY + SONIC
Figure 8: Architecture survival and recovery over all tested push magnitudes. Each point contains 20 paired direction/seed cells; ribbons are descriptive 95% paired-direction cluster-bootstrap intervals from 20,000 replicates. Lines guide the eye. The aggregate values in Table 2 retain all five forces within each of the 20 paired clusters. C.3
D EPLOYMENT D EMONSTRATION S COPE
The demonstration suite spans simulation and hardware. Joystick steering and semantic interpolation are evaluated in simulation (Figure 1(a,c)); external interference and staged text commands are demonstrated on the physical Unitree G1 (Figure 1(b,d)). Figure 7(g) provides additional onboard frames from the G1–Jetson Orin NX stack. The deployed actor uses the observation, condition, export, and safety interfaces described in Appendix B.3. The selected physical sequences show integrated execution without visible jitter or falls. These sequences provide qualitative evidence of stable onboard execution; future matched trials with synchronized safety telemetry and explicit denominators can quantify success, recovery, and failure rates.
D
C OMPLEMENTARY R ESULTS AND L IMITATIONS
D.1
O RIN S AMPLER AND I MPLEMENTATION A BLATIONS
The sampler sweep in Table 7 uses P0 , the original TorchScript actor before the acceleration changes evaluated here. Rolling DDPM has 20 schedule levels but executes 15 updates in the deployed rolling window. With legacy WBG it evaluates the backbone twice per update. The corrected twostep WBG path executes two backbone forwards in total. Figure 7(a) reports the timing sweep, and Figure 7(d) separately measures push survival across the tested sampler budgets; together they identify DDIM-2 as the best tested trade-off between inference speed and disturbance survival. Table 7: P0 actor timing on Jetson Orin. Values are actor CUDA p50 in milliseconds, each the median of three block medians. Sampler NS-DDIM NS-DDIM NS-DDIM NS-DDIM NS-DDIM Rolling DDPM
Steps WBG off WBG on 1 2 3 5 10 15
9.725 16.148 22.915 36.001 68.091 121.229
12.221 21.846 29.946 46.948 92.011 235.216
The exact factorial study crosses four attention-projection layouts with condition-embedding reuse, DDIM schedule precomputation, and sequential versus stacked WBG axes. All 32 valid combinations use the correct two-forward guidance equation. Across matched valid cells, all-QKV packing
22
Preprint
saves 0.775 ms with a range of [0.256, 0.799] ms across the three block-level effects; condition reuse saves 0.190 ms, with a range of [0.151, 0.347] ms; precomputation saves 1.672 [1.477, 1.698] ms, and stacked WBG saves 2.845 [2.627, 2.868] ms. These conditional effects depend on the other settings and must not be summed. The cumulative path in Table 3 reports individual matched composition differences rather than these matrix-wide conditional effects. Cross-attention-only packing is not consistently faster. An implementation that duplicated backbone execution changes the guidance equation and is excluded from the exact ablation. The optimization path also crosses two controlled boundaries. Stage P1 is a TorchScript re-export that supports condition caching, so its 18.315 ms median is not a within-graph effect relative to the P0 value of 21.846 ms. The factorially selected exact configuration P2 measures 15.930 ms. The same configuration measures 16.234 ms as P3 in the runtime suite; the difference reflects the measurement suite rather than an implementation change. Enabling LibTorch inference mode at P4 reduces the median by 3.058 ms to 13.175 ms. Table 8 collects all 32 valid structural compositions as 16 sequential/stacked WBG pairs, all eight runtime controls, and matched component effects. Frozen actor and live callback timings remain separate; detailed per-boundary records are retained in the source data.
23
Preprint
Table 8: Comprehensive exact Orin implementation ablation. NS-DDIM2, B=1 FP32, WBG on, two backbone forwards. Times are milliseconds. A. All 32 structural compositions (16 paired rows). QKV
C
P
Seq. p50
Seq. p95
Stk. p50
Stk. p95
SavedWBG
none self cross all none self cross all none self cross all none self cross all
0 0 0 0 0 0 0 0 1 1 1 1 1 1 1 1
0 0 0 0 1 1 1 1 0 0 0 0 1 1 1 1
21.461 21.233 20.669 20.522 19.082 19.315 19.674 18.685 20.985 20.384 20.717 21.112 19.184 19.258 19.177 18.803
22.036 22.260 22.016 20.812 20.631 20.292 20.656 19.374 21.948 21.196 21.823 22.008 19.492 20.695 19.492 19.467
18.315 17.972 18.583 17.751 17.442 16.232 17.180 16.331 18.188 17.789 17.888 17.423 16.631 16.104 16.387 15.930
19.677 18.376 18.885 18.873 18.347 17.116 17.580 17.372 19.359 19.430 18.414 18.262 17.572 16.825 16.720 16.728
3.147 3.261 2.086 2.771 1.640 3.083 2.494 2.354 2.798 2.595 2.829 3.688 2.553 3.154 2.790 2.873
C: condition cache; P: DDIM precomputation (0/1 = off/on). Seq./Stk.: sequential/stacked WBG. SavedWBG = Seq. p50 − Stk. p50 in the same cell; inference off, eager checks. B. All eight runtime configurations (separate suite). Structure
I
Checks
Actor p50
p95
Saved
Host p50
> 20 ms
base base base base all all all all
0 0 1 1 0 0 1 1
eager defer eager defer eager defer eager defer
18.266 18.301 15.401 14.924 16.234 15.374 13.175 13.183
19.104 19.286 15.651 15.201 16.912 15.932 13.400 13.609
0.000 -0.035 2.865 3.342 2.032 2.892 5.090 5.083
21.042 20.943 19.451 18.873 18.855 18.446 16.790 16.836
600/600 600/600 49/600 61/600 60/600 90/600 7/600 13/600
Base/all: QKV packing, cache and precompute all off/on; WBG stacked. I: inference mode. Saved compares frozen actor p50 to the first row (18.266 ms). Host = live callback-to-host; last column counts host deadline misses. Deferred checks change failure-detection semantics. C. Conditional frozen-actor savings, not additive. Component
Matched change
QKV packing QKV packing QKV packing Condition cache DDIM precompute WBG axes Inference mode
none → self none → cross none → all off → on off → on sequential → stacked off → on (selected structure)
Pairs/block
Saved
Block-effect range
8 8 8 16 16 16 1
0.221 0.172 0.775 0.190 1.672 2.845 3.166
[0.003, 0.555] [-0.100, 0.363] [0.256, 0.799] [0.151, 0.347] [1.477, 1.698] [2.627, 2.868] [2.812, 3.492]
Each primary cell uses three rotated blocks of 200 retained calls; p50/p95 are medians of block quantiles. Panel C gives median matched block effects and their min–max range, not confidence intervals. Its inference-mode paired effect (3.166 ms) differs from the 3.058 ms difference of endpoint medians in Table 3. Bold marks the selected structure/runtime. Its fresh actor p50/p95 is 13.843/16.221 ms; callback p95 is 19.383 ms, with 7/600 overruns: no hard-real-time guarantee. Sampler controls remain in Table 7; the complete per-boundary source data retain block ranges, actor misses and live structural timings omitted here. The 32 duplicate-backbone cells are excluded as an execution bug; no cross-suite savings are inferred.
24
Preprint
D.2
T IMING D ECOMPOSITION
Table 9 records the median component diagnostics most relevant to the deployment boundary. The two backbone and WBG spans are inclusive traced CUDA regions. Their medians cannot be added to each other or to host spans. Tracing inflates the final actor span to 40.090 ms, whereas an independent unprofiled fixed-input pass measures 12.198 ms; neither replaces the primary 13.175 ms E0 endpoint. Table 9: Diagnostic component p50 timing in milliseconds. P0
Component
P4 Boundary
State fetch/copy 0.002 0.001 host Pre-FK and FK 0.757 0.779 host Observation construction 0.919 0.997 host Condition packing 0.083 0.082 host Action selection/copy 0.234 0.238 host Backbone forward 1 18.053 13.056 traced CUDA Backbone forward 2 18.253 13.852 traced CUDA WBG gradient 1 3.095 2.562 traced CUDA WBG gradient 2 3.030 2.648 traced CUDA Actor-loop residual 8.054 2.474 traced CUDA
D.3
C APABILITY, C OVERAGE , AND I NTERPRETATION L IMITS
Architecture study. The seven policies are fixed checkpoints. Push statistics use 20 paired direction/seed clusters across all five forces; point, text, and nominal-walk statistics use five rollout seeds; latency uses five timing seeds. Means and sample standard deviations quantify the selected policies across those rollout or timing repeats; training-seed variation lies outside the reported intervals. Non-evaluable cases mark capability boundaries. ARDY+SONIC values condition on convertercompatible references and retain their partial-coverage denominators. MotionCLIP retrieval measures representation agreement. Timing study. Timing applies to one batch-one FP32 artifact on one Jetson Orin. The benchmark measures the offline callback-to-host-action boundary from received robot-state messages; all actuation counters remain zero. The complete-callback p95 is below 20 ms, and 593/600 calls meet the target; the seven observed overruns define the remaining tail-latency margin. Physical closed-loop behavior is documented by the deployment demonstrations. Graph bytes, thermal state, runtime version, and implementation settings are part of the measured system; a different export or device requires requalification. Real-robot demonstrations and scope. The displayed Unitree G1 sequences verify hardware execution of text-command and disturbance-response interfaces without visible jitter or falls. Together with the simulation interfaces, they provide qualitative deployment evidence for the tested architecture, checkpoint, hardware, and tasks. Matched trials with explicit denominators and uncertainty can extend this evidence to quantitative success, recovery, and robustness rates. The predicted state serves as an internal future-state guidance representation; calibration as a state estimate lies outside its evaluated role.
25