ConceptioArchivearXiv CS
arXiv CSopen access

AHA-WAM:Asynchronous Horizon-Adaptive World-Action Modeling with Observation-Guided Context Routing

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
artificialintelligenceknowledgerepresentationreasoning
artificial intelligence, reasoning, knowledge representation

AHA-WAM:Asynchronous Horizon-Adaptive World-Action Modeling with Observation-Guided Context Routing Jisong Cai1,2∗ Long Ling1,3∗ Shiwei Chu1 Zhongshan Liu3 Jiayue Kang1 Zhixuan Liang4,2 Wenjie Xu3 Yinan Mao3 Weinan Zhang1,2 Xiaokang Yang1 Ru Ying3 Ran Zheng3 Yao Mu1,2†

arXiv:2606.09811v1 [cs.RO] 8 Jun 2026

1

Shanghai Jiao Tong University 2 Shanghai AI Laboratory 3 Baidu AI Cloud 4 The University of Hong Kong

Project Page: https://serene-sivy.github.io/aha-wam/ Abstract: World-action models have emerged as a promising paradigm for robot manipulation, jointly modeling visual scene dynamics and actions to inject physical priors into policy learning. However, existing world-action models couple world prediction and action execution at the same temporal resolution, forcing the world branch to model near-term frame variations that are redundant and weakly informative. We posit that strictly binding world prediction and action execution to the same temporal rhythm may underutilize the potential of the video branch for embodied control. Therefore, we propose AHA-WAM, an Asynchronous Horizon-Adaptive World-Action Model built on a dual Diffusion Transformer (DiT) architecture that reorganizes world-action modeling around this temporal asymmetry. AHA-WAM instantiates the video DiT as a low-frequency world planner that maintains rolling key-value memory over past observations and exposes reusable layerwise latent context encoding long-horizon scene evolution, while a high-frequency action DiT executes short action chunks in closed loop by querying this context through layerwise joint attention. To support asynchronous execution, we introduce horizon-adaptive offset training and Observation-Guided Video-Context Routing (OVCR), which together let the action expert exploit long-horizon world context while remaining responsive to realtime execution state without rerunning the video DiT. Experiments on RoboTwin and real-world manipulation tasks show that AHA-WAM achieves state-of-the-art performance without any robot-data pretraining, attaining 92.80% average success on RoboTwin and 78.3% success across 4 real-world tasks, while reaching 24.17 Hz closed-loop control with a 4.59× speedup over Fast-WAM. Keywords: Robot Learning, Embodied Manipulation, World-Action Model

1

Introduction

Robotic manipulation requires policies that understand not only the current scene, but also how the scene may evolve under the robot’s actions. Recent vision-language-action (VLA) models have advanced robot control by scaling imitation learning with large vision-language model backbones, yet action labels provide relatively sparse supervision for the underlying physical dynamics of manipulation. World-action models (WAMs) address this by coupling action prediction with dense video-based world modeling, learning how actions co-evolve with visual scene dynamics to inject physical priors into control and offer a more generalizable, and transferable policy representation. * Equal contribution. †

Corresponding author. Email: [email protected]. This work was done during Long Ling’s and Jiayue Kang’s internship at Shanghai Jiao Tong University.

t4

t0

World Planner

t8

t12

long-horizon latent planning t0

Action Expert high-frequency closed-loop control

t1

t2

t3

t4

t5

t6

t7

t8

t9

t10

t11

t12

t13

a0

a1

a2

a3

a4

a5

a6

a7

a8

a9

a10

a 11

a12

Organize Desktop

Store Plate

Prepare Soy milk

Fold Towel

Naive WAM

Real-world Motus 21.67

FastWAM 68.33 π05 76.67

RoboTwin 2.0

AHA-WAM

π05 79.75 Motus 87.84

AHA-WAM 78.33

0

FastWAM, 91.83

20 Motus, 0.54

85

90

80

100

Frequency (Hz) AHA-WAM 24.17

AHA-WAM 92.80 80

60

FastWAM 5.26

LingBot-VA 92.24 75

40

AHA-WAM-Flash 56.95

95

0

20

40

60

Figure 1: Overview of AHA-WAM. AHA-WAM connects past observations, future-oriented world planning, and fast closed-loop action execution: a slow world planner maintains reusable memory and planning context, while a fast action expert adapts that context to the latest observation before predicting short action chunks. Despite this promise, current WAMs leave an important design space underexplored: how can the video branch more effectively empower robot control? Existing approaches either explicitly roll out future frames for inverse-dynamics decoding or joint world-action modeling, while others reuse the video branch as a latent encoder alongside the action branch. Both designs assume that world prediction and action execution share the same short horizon, which forces the world branch to spend capacity on dense adjacent-frame variations that are often highly correlated and only weakly informative for control. We argue that this coupling reflects a structural mismatch in temporal abstraction: the video world model can better serve embodied control by forming a temporally extended latent plan over future visual states, whereas the action model should remain tightly coupled to the real-time control loop, incorporating the latest observation to issue timely closed-loop corrections. We propose AHA-WAM, an Asynchronous Horizon-Adaptive World-Action Model that reorganizes world-action modeling around this temporal asymmetry. AHA-WAM instantiates the video Diffusion Transformer (DiT) as a low-frequency world planner that maintains a rolling key-value (KV) memory over past observations and exposes reusable layerwise latent context encoding long-horizon scene evolution, amortizing expensive world-model computation across multiple action steps. In parallel, a high-frequency action DiT executes short action chunks in closed loop by querying this context through layerwise joint attention. This design preserves the benefits of joint world-action modeling while aligning each branch with the temporal scale at which it’s most informative. In AHA-WAM, horizon-adaptive refers not to online horizon adjustment, but to a horizon-decoupled formulation where the world and action branches are assigned different temporal horizons according to their functional roles. A key challenge introduced by asynchronous execution is that the planner context may become stale or misaligned with the current action chunk as the executor runs ahead. To address this, we introduce Observation-Guided Video-Context Routing (OVCR), which constructs chunk-specific latent queries from the latest visual observations to route and update the planner’s layerwise video context before it is consumed by the action DiT, giving each action prediction an observation-conditioned view of the slow planner state without feeding dense visual tokens into the high-frequency action branch. Proprioceptive feedback enters the action DiT directly, while visual feedback is injected 2

indirectly through the routed video context, keeping the reused planner context responsive to realtime execution state without recomputation. We further introduce horizon-adaptive offset training to expose the model to diverse planner-executor phase relationships, and equip the video DiT with rolling KV memory to connect past observations with future plans. We evaluate AHA-WAM on RoboTwin 2.0 and real-world manipulation tasks. Though without robot-data pretraining, AHA-WAM achieves 92.80% average success across 50 RoboTwin 2.0 tasks, reaching state-of-the-art-level performance among strong VLA and WAM baselines. Across four real-world tasks covering deformable manipulation, long-horizon organization, fine-grained tool use, and spatial generalization, AHA-WAM achieves 78.3% average success, demonstrating robust deployment performance beyond simulation. Across four dimensions of real-world out-ofdistribution evaluation, AHA-WAM matches π0.5 in exhibiting the most limited performance degradation, suggesting stronger robustness to distribution shifts. Benefiting from the asynchronous inference schedule, ODE distillation and cuda optimizations, AHA-WAM reaches up to 56.9 Hz closedloop control frequency, corresponding to a 10.82× speedup over Fast-WAM. In summary, our contributions are threefold: • First, we propose AHA-WAM, an asynchronous horizon-adaptive world-action model that decouples slow video-DiT world planning from fast action-DiT closed-loop execution, introducing horizon-adaptive offset training to support arbitrary planner-executor phase relationships. • Second, we develop Observation-Guided Video-Context Routing (OVCR), which dynamically constructs chunk-specific latent video context from current observations to keep asynchronous planner context aligned with real-time execution state. • Third, we validate AHA-WAM across simulation, real-world manipulation, and latency benchmarks, demonstrating state-of-the-art-level manipulation performance without large-scale robotdata pretraining and substantially improved inference efficiency over existing world-action model baselines.

2

Related Work

Generalist robot manipulation policies. Generalist robot policies aim to learn broadly applicable manipulation skills from large-scale diverse demonstrations [1–4]. RT-1 [5] and RT-2 [6] established the vision-language-action (VLA) paradigm by absorbing heterogeneous robot trajectories and transferring web-scale vision-language knowledge into robotic control. Subsequent works [7–15] further improved generalization by integrating flow-matching and diffusion-based action heads [16, 17] with large pretrained backbones, establishing Diffusion Transformers (DiTs) [18] as expressive and scalable policy architectures. However, these methods remain primarily actioncentric, with physical scene dynamics only implicitly captured through demonstrations, which motivates the world-action modeling paradigm that AHA-WAM extends. World models for robot control. World models [19–22] augment robot policies with predictive visual dynamics beyond action-only imitation. Existing video-based robot world models can be broadly grouped into two lines. The first follows an imagine-then-act paradigm: methods such as UniPi [23], Seer [24], and Video Prediction Policy [25] first predict future visual states or predictive visual representations, and then recover actions through inverse dynamics or a policy conditioned on the predicted future, introducing latency in closed-loop deployment. The second line performs joint world-action modeling, where future visual dynamics and actions are modeled within a shared generative architecture [26–31]. But all couple world prediction and action execution at the same temporal resolution, spending capacity on short-horizon adjacent frames that are redundant for control. AHA-WAM addresses this by decoupling the video DiT as a low-frequency long-horizon planner from the action DiT as a high-frequency closed-loop executor. Dual-system and asynchronous robot policies. Recent robot policies have explored dual-system or asynchronous designs to combine deliberative computation with reactive control. RoboDual [32] uses a VLA generalist to provide task understanding and coarse action guidance for an efficient dif3

Cross Attention Residual Update

Slow

Video DiT

Fast

Action DiT

(World Planner)

(Action Expert)

f1 f2 f�

�1

1 1 a1 a2

�2

2 2 a1 a2

×L

×L Cross Attention & FFN

Cross Attention & FFN Layer-wise Joint Attention

AdaLN

Video Branch Action Chunk1 Action Chunk2

AdaLN

Video Context Routing

Self Attntion

K/V

Joint Attention Training

QKV AdaLN

VAE Encoder

f1

QKV Obs-guided query

�1

1 a1 1 a2

�2

2 a2 1 a2

AdaLN

Vision Encoder

State Encoder Inference

Rolling K/V Memory

Obs-guided query

Updated K / V

Figure 2: AHA-WAM architecture and attention mask. AHA-WAM decouples world planning and action execution into a slow video-DiT planner and a fast action-DiT executor. The video branch is trained with a fully causal mask to learn temporal dynamics. For each action update, the latest observation on first queries and updates the video DiT’s K/V states through OVCR, producing updated planner context that is consumed by the action DiT through layerwise joint attention. fusion specialist, while Reactive Diffusion Policy [33] combines a slow latent diffusion policy with a fast tactile feedback pathway for contact-rich manipulation. AsyncVLA [34] runs a large foundation model asynchronously to provide delayed semantic guidance, which is refined by a lightweight edge adapter for fast execution. While these methods validate the effectiveness of dual-system execution, their slow branch primarily provides guidance, feedback, or representations for the action policy. AHA-WAM instead realizes the slow-fast split within joint world-action modeling: the video DiT and action DiT operate asynchronously yet remain coupled through layerwise joint attention.

3

Asynchronous Horizon-Adaptive World-Action Modeling

World-action models couple action prediction with learned visual dynamics, but existing formulations typically organize the video branch and the action branch under the same short-horizon execution rhythm. We instead formulate world-action modeling as a two-timescale generation problem. Figure 2 illustrates the overall architecture. AHA-WAM reorganizes WAM inference into an asynchronous world-action coupling framework while preserving a dual-DiT architecture: a lowfrequency video planner produces reusable long-horizon planner context, and a high-frequency action expert consumes this context for closed-loop action denoising. Here, “horizon-adaptive” denotes robustness to arbitrary planner–executor phase offsets under decoupled temporal horizons, rather than online horizon selection. In other words, AHA-WAM does not dynamically choose the video or action horizon at test time; instead, it trains the executor to consume long-horizon planner context under variable action-start offsets, making the asynchronous interface robust to the phase misalignment induced by streaming deployment. This design is supported by three complementary mechanisms. Observation-Guided Video-Context Routing (OVCR) adapts the cached planner context to the latest observation without rerunning the video DiT. Horizon-adaptive offset training exposes the action expert to the planner–executor phase shifts induced by asynchronous streaming, while rolling K/V memory extends the video planner’s 4

temporal receptive field over past observations. Finally, real-time inference optimizations further accelerate the high-frequency action stream for closed-loop deployment. 3.1

Dual-DiT Planner–Executor Architecture

Problem setting. We consider language-conditioned visuomotor policy learning from visual observations and proprioceptive robot states. At control time t, the policy receives a visual observation history Otv , a proprioceptive state st , and a language instruction l, and predicts an executable action chunk At = {at , . . . , at+ha −1 } of horizon ha . A standard action-centric policy directly models v πθ (At | O≤t , st , l).

(1)

World-action models augment this policy with visual dynamics learning by jointly modeling future v visual evolution and action generation. Let Zt:t+h denote future video latents over a visual planv ning horizon hv . Rather than updating video modeling and action generation under the same short horizon, AHA-WAM decouples their temporal roles: the video branch models a longer horizon hv , while the action branch predicts executable chunks of horizon ha , with ha < hv . Model architecture. AHA-WAM is built on a dual-DiT architecture consisting of a video DiT world planner and an action DiT executor. Visual observations are encoded by the pretrained VAE, while language embeddings from the text encoder condition both branches. The action branch is a lightweight action expert DiT with the same layer depth as the video DiT, enabling layerwise interaction between the two branches. The video DiT takes visual latent tokens as input and is trained to predict future video latents over the longer planning horizon. The action DiT receives noisy action tokens and proprioceptive tokens, and denoises the action chunk under closed-loop robot-state feedback. Visual feedback for the highfrequency action branch is not introduced by naively concatenating dense visual tokens; instead, visual information is mediated through the planner video context and later adapted by OVCR in Section 3.2. Layerwise planner–executor coupling. The central interface between the two DiTs is the layerwise planner video context. Given the visual observation context and language instruction, the video planner produces   L Cτp = Kτp,ℓ , Vτp,ℓ ℓ=1 , (2) where p denotes the planner branch, τ indexes the latest planner refresh and ℓ indexes the transformer layer. This context is a latent world-plan representation exposed by one video-DiT forward and reused by multiple subsequent action-DiT forwards. It differs from the rolling K/V memory in Section 3.4, which stores historical video states across planner refreshes inside the video planner. During training, the video branch predicts future video latents under a fully causal video mask, encouraging the video DiT to learn forward scene dynamics while shaping the planner context with visual dynamics supervision. The action branch is masked from attending to future video tokens so that during inference, the future-video prediction path can be removed. Thus, future-video prediction serves as a world-modeling training signal, while planner video context serves as the inference-time interface to the action executor. At each action update, the raw planner context Cτp is first adapted by OVCR into a chunk-specific e tp,ℓ , Vetp,ℓ )}L . The action DiT then denoises the current action chunk through context Cetp = {(K ℓ=1 layerwise joint attention:  h i h i a,ℓ e p,ℓ H̄ta,ℓ = Attn Qa,ℓ , Vta,ℓ ; Vetp,ℓ , (3) t , Kt ; Kt a,ℓ a,ℓ where Qa,ℓ are the action-DiT query, key, and value projections, and H̄ta,ℓ is the t , Kt , and Vt planner-conditioned action hidden state. This coupling preserves WAM-style interaction between visual dynamics and action generation, while amortizing expensive video-DiT computation across multiple high-frequency action updates.

5

Joint world-action training objective. AHA-WAM is trained with a joint flow-matching objective over action chunks and future video latents. For a target variable y, either an action chunk At or v future video latents Zt:t+h , we sample Gaussian noise ϵ ∼ N (0, I) and a flow time ρ ∈ (0, 1), and v construct the interpolated sample yρ = (1 − ρ)y + ρϵ. (4) The model predicts the corresponding velocity field: i h 2 LFM (y) = Ey,ϵ,ρ ∥fθ (yρ , ρ, Otv , st , l) − (ϵ − y)∥2 .

(5)

We instantiate this objective for action generation and video co-training: La = LFM (At ),

v Lv = LFM (Zt:t+h ). v

(6)

The final objective is L = La + λLv ,

(7)

where λ balances action learning and video dynamics learning. At deployment, AHA-WAM removes explicit future-frame decoding: the video DiT only refreshes planner video context, and the action DiT uses this context to generate closed-loop action chunks. 3.2

Observation-Guided Video-Context Routing

Asynchronous execution reuses one planner video context for multiple action chunks, which amortizes video-DiT computation but creates a context-alignment problem: before the next planner refresh, the robot state and visual scene may have changed. Observation-Guided Video-Context Routing (OVCR) addresses this by using the latest visual observation to convert the shared planner context into a chunk-specific context. Instead of feeding dense visual tokens into the high-frequency action DiT, OVCR compresses them into a small set of routing queries that select and edit the relevant planner context before action denoising. For each action chunk at time t, we split the aligned observation context into visual tokens Xtv and proprioceptive tokens Xts . The proprioceptive input is compact and directly tied to the instantaneous robot state, so a lightweight encoder maps it to a state token which is directly provided to the action DiT. Visual feedback is injected indirectly through context routing. Given Q learnable base queries B ∈ RQ×d , OVCR constructs observation-guided routing queries by attention pooling over the current visual tokens: Ztq = Attn (B, fv (Xtv ), fv (Xtv )) , (8) where fv is a lightweight visual projection module. The resulting queries provide compact, observation-conditioned slots for retrieving chunk-relevant information from the planner video context. p,ℓ L Let Cτp(t) = {(Kτp,ℓ (t) , Vτ (t) )}ℓ=1 denote the latest available planner context. For each layer ℓ, OVCR first reads planner features using the routing queries and then predicts residual key–value updates:     p,ℓ Rtℓ = Attn Ztq,ℓ , Kτp,ℓ (∆Ktp,ℓ , ∆Vtp,ℓ ) = gψℓ Rtℓ , Ztq,ℓ , (9) (t) , Vτ (t) ,

where gψℓ is a lightweight layerwise router. The chunk-specific planner context is produced by a gated residual update: e tp,ℓ = K p,ℓ + αtℓ ∆Ktp,ℓ , K τ (t)

p,ℓ ℓ Vetp,ℓ = Vτp,ℓ (t) + αt ∆Vt ,

(10)

e tp,ℓ , Vetp,ℓ )}L is then consumed by the where αtℓ is a learned gate. The adapted context Cetp = {(K ℓ=1 action DiT through the layerwise joint attention in Eq. 3. OVCR turns planner-context reuse from static caching into observation-conditioned retrieval and adaptation. Thus, the video DiT remains outside the per-update critical path, while each action chunk still receives a planner representation aligned with the latest visual evidence. 6

Our asynchrony spans both training and inference.

t4

t0

t0

Horizon-Adaptive Offset δ

t8

t12

t1

t2

t3

t4

t5

t6

t7

t8

t9

t10

t11

t12

t13

a0

a1

a2

a3

a4

a5

a6

a7

a8

a9

a10

a11

a12

Horizon-Adaptive Offset Training

Figure 3: Horizon-adaptive offset training. We randomly shift the action-chunk grid by δ ∈ Inference finished [0, ha ) inside the video planning horizon, so the action executor learns to consume planner context and get obs under different phaset offsets induced by deployment. t asynchronous t t K/V K/V K/V 4

0

8

12

Inference

3.3

Horizon-Adaptive Offset Training t t t t t 0

1

2

3

4

t5

t6

t7

t8

t9

t10

t11

t12

t13

a a a a a a a a a a a Asynchronous streaming achanges the relative temporal phase between the slow planner anda the fast executor. If training always (b) uses a fixed alignment between the video planning window and the Asynchronous Dual-stream Deployment action chunk, the action DiT may overfit to a single planner–executor phase and become brittle when reusing planner context at intermediate phases during deployment. We therefore introduce horizon-adaptive offset training to expose the executor to the phase shifts induced by asynchronous inference. 0

1

2

3

4

5

6

7

8

9

10

11

12

Let hv denote the video planning horizon and ha denote the action chunk horizon, with ha < hv . As shown in Figure 3, for each training segment starting at time τ , the video planner models future video latents over [τ, τ + hv ). Instead of aligning the first action chunk to the planner start, we sample δ ∼ U {0, 1, . . . , ha − 1},

(11)

and shift the action-chunk grid by δ within the planner horizon. The action objective is then evaluated over offset-aligned chunks:   La = Eδ LFM Aδτ , (12) where Aδτ denotes an action chunk starting from an offset-aligned position inside the planner horizon, and LFM is defined in Eq. 5. Since the planner–executor alignment is periodic with the action chunk size, sampling δ ∈ [0, ha ) covers all chunk-level phases encountered when one planner context is reused across multiple executor updates. Thus, horizon-adaptive offset training teaches the action DiT to consume long-horizon planner context under variable action-start offsets, while OVCR handles observation-conditioned context selection at each phase. 3.4

Rolling Planner Memory

Since the video DiT operates as a low-frequency planner, it should preserve historical scene information across planner refreshes rather than relying only on the current observation. This is important for long-horizon manipulation, where completed subgoals, displaced objects, or previously observed states can provide essential context. We therefore maintain a fixed-size FIFO rolling K/V memory inside the video planner. For each layer ℓ, the memory stores historical video states from recent planner refreshes:  Mτℓ = FIFO Mτℓ−1 ∪ {(Kτp,ℓ , Vτp,ℓ )} ,

(13)

where τ indexes planner refreshes and the memory window size is fixed. At the next refresh, the video DiT attends to this memory when producing the new planner video context Cτp = {(Kτp,ℓ , Vτp,ℓ )}L ℓ=1 . This memory is internal to the slow video planner and is not directly consumed by the action DiT. It extends the planner’s temporal receptive field before the next planner context is produced, while the high-frequency executor still interacts only with the latest OVCR-adapted planner video context. 7

3.5

Streaming Inference and Real-Time Optimization

The asynchronous planner-executor architecture naturally leads to a streaming inference schedule. At deployment, the video planner and the action executor are executed as two non-blocking streams, where AHA-WAM removes the video DiT from the per-update critical path, but the action branch must still run at the closed-loop control frequency. We therefore optimize the action-chunk inference path, measuring the end-to-end latency Lchunk of one action update, including image encoding, planner-context access, OVCR routing, and action denoising. Video-DiT prefill is executed asynchronously, so Lchunk directly determines the action-update frequency reported in Section 4.5. CUDA acceleration. We compile the repeated action-phase computation into a static deployment path by executing the action DiT, memory and context modules, and VAE encoder through TensorRT and CUDA-graph capture where possible, while removing redundant loop-invariant computations and repeated buffer copies from the denoising hot path. These changes do not alter the model architecture, weights, or sampling procedure, but reduce the 10-step action inference latency from 415.77 ms in PyTorch eager execution to 41.37 ms. ODE distillation. On top of the CUDA-accelerated 10-step path, we construct AHA-WAM-Flash by distilling the action sampler from 10 denoising steps to 2 steps while keeping the same observation, planner-context, and OVCR interface. The video DiT is frozen during distillation, and the student is trained from sampled intermediate states to directly predict the teacher’s final action output, with sampling biased toward noisier states to support aggressive step reduction. ODE distillation further reduces Lchunk from 41.37 ms to 17.56 ms. Additional CUDA ablations, distillation schedules, and step-latency tradeoffs are provided in Appendix D.

4

Experiments

We evaluate AHA-WAM in simulation and on real robots to test whether asynchronous world-action modeling preserves manipulation performance while improving closed-loop efficiency. We cover RoboTwin 2.0 [35] performance, component ablations, real-world deployment, generalization analysis and inference latency. Specifically, our experiments ask five questions including capability, mechanism, deployability, robustness, and efficiency. Q1. Can our asynchronous WAM preserve or improve simulation performance while faster? (Sec. 4.2) Q2. How important is each component for AHA-WAM? (Sec. 4.3) Q3. Can AHA-WAM achieve reliable real-robot performance? (Sec. 4.4) Q4. How robust is AHAWAM under previously unseen real-world task conditions? (Sec. 4.4) Q5. How much closed-loop latency reduction does AHA-WAM provide? (Sec. 4.5) 4.1

Experimental Setup

Model and training. AHA-WAM is implemented as a dual-DiT policy with an explicitly asynchronous planner–executor interface. We use the pretrained Wan2.2-5B video model to initialize the world-planning branch, including the video DiT, text encoder, and video VAE. Following Fast-WAM [30], the high-frequency action executor adopts a compact DiT architecture with hidden dimension da = 1024, corresponding to approximately 1.02B parameters. In addition to the 4.99B-parameter video planner and the 1.02B-parameter action DiT, AHA-WAM introduces 1.22B parameters for rolling K/V memory and context-routing modules, giving a total instantiated model size of about 7.23B parameters. For the temporal configuration of AHA-WAM, the video branch operates over a long planning horizon of hv = 64, while the action branch predicts short executable chunks with horizon ha = 16. During asynchronous inference, the video planner maintains reusable layerwise K/V context, which is augmented by a FIFO rolling memory over at most 6 historical observation frames. The action executor adapts this context through OVCR, using 32 observation-guided routing queries for each action chunk. 8

Table 1: RoboTwin 2.0 average success on 50 tasks. Method

Robo. P.T.

Clean (%)

Rand. (%)

Avg. (%)

π0 π0.5 ABot-M0 Motus from Wan2.2 Motus LingBot-VA Fast-WAM

✓ ✓ ✓ ✗ ✓ ✓ ✗

65.92 82.74 81.20 77.56 88.66 92.90 91.88

58.40 76.76 80.40 77.00 87.02 91.50 91.78

62.16 79.75 80.80 77.28 87.84 92.20 91.83

AHA-WAM-Flash AHA-WAM

✗ ✗

90.48 93.40

89.92 92.20

90.20 92.80

Training uses flow matching for both world modeling and action prediction, with logit-normal sampling of noise times. For action generation, we use 10 denoising steps at inference and set CFG to 1.0. All models are trained with AdamW, a learning rate of 1 × 10−4 , weight decay 0.01, cosine scheduling, mixed precision, and gradient clipping. Latency is reported on a single NVIDIA RTX 5090D GPU. Appendix A lists additional implementation details, including AHA-WAM-Flash distillation and other training configurations. Evaluation scope. We evaluate AHA-WAM in both RoboTwin 2.0 simulation [35] and real-world experiment. The simulation benchmark tests multi-task manipulation under clean and randomized scenes, while the real-world experiments test deployment on physical bimanual tasks. Across both settings, we report task success rate as the primary metric and additionally analyze closed-loop inference latency to quantify deployment efficiency. 4.2

RoboTwin Simulation

We evaluate AHA-WAM on RoboTwin 2.0 [35], a bimanual manipulation benchmark with 50 dualarm tasks covering diverse skills. We follow the multi-task training setup of [27, 29, 30]: for each task, the training set contains 50 demonstrations in the clean setting and 500 demonstrations in the randomized setting. Models are trained in a multi-task manner with a global batch size of 512. Each method is tested over 100 trials per task in both the clean and randomized settings, and we report task-averaged success rates. We compare against competitive VLA and WAM baselines, including π0 [2], π0.5 [9], ABot-M0 [36], Motus [27], LingBot-VA [29], and Fast-WAM [30]. Appendices B and E provide detailed protocols, baseline settings, and per-task success rates. Results and analysis. Table 1 summarizes the RoboTwin results: AHA-WAM achieves 93.40% success in the clean setting and 92.20% under randomized evaluation, yielding an average success rate of 92.80%. Without robot data pretraining, AHA-WAM improves over Fast-WAM by 0.97 percentage points, showing that our asynchronous design not only increases the closed-loop control frequency but also preserves and improves performance. AHA-WAM also exceeds LingBot-VA, the strongest baseline with large-scale robot-data pretraining, by 0.60 points. The AHA-WAM-Flash variant retains 90.20% average success, incurring only a modest performance drop while further cutting the latency of AHA-WAM by more than half as analyzed in Section 4.5. 4.3

Ablation Studies on RoboTwin

We ablate the components that make asynchronous world-action modeling effective. Naive-Async decouples the video and action branches, but directly reuses the latest planner context without rolling K/V memory or OVCR. We then add each mechanism separately and jointly.

Table 2: RoboTwin 2.0 ablations. Variant Fast-WAM Naive-Async + KV Memory + OVCR AHA-WAM

Clean (%) Rand. (%) Avg. (%) 91.88 88.64 91.40 91.52 93.40

91.78 88.56 90.62 91.42 92.20

91.83 88.60 91.01 91.47 92.80

Table 2 shows that asynchronous execution alone is insufficient. Naive-Async drops from 91.83% to 88.60%, indicating that faster action up9

dates cannot compensate for stale or phase-misaligned planner context. Adding rolling K/V memory recovers performance to 91.01%, showing that persistent planner states help stabilize the reused video context across refreshes. The gain is moderate, which is expected because most RoboTwin tasks are short-to-medium horizon and keep task-relevant objects visible. OVCR further improves success to 91.47%, suggesting that observation-conditioned context adaptation is the more direct remedy for asynchronous planner–executor mismatch. The full AHA-WAM reaches 92.80%, confirming that memory and routing are complementary: memory preserves temporal context, while OVCR aligns it with the current execution state.

4.4

Real-World Robot Experiments

Setup. We evaluate on a bimanual AgileX Piper platform with an ego-view RGB camera, and policies uses only head-view RGB observations, proprioceptive states and language instructions as input. We consider four tasks—Fold Towel, Organize Desktop, Prepare Soy Milk, and Store Plate— covering deformable manipulation, long-horizon rearrangement, fine-grained tool use, contact-rich control, and spatial generalization. For each task, we collect approximately 120 episodes on average. Appendix C provides task details, data scale, execution procedures, and generalization variants. Implementation. Since Fast-WAM and AHA-WAM are not pretrained on any robot data by default, we pretrain Fast-WAM and AHA-WAM on the selected RoboCOIN subset [37], containing 24,600 trajectories and approximately 165 hours of robot data for a fair and stable deployment comparison. Both models are finetuned on the same task-specific demonstrations. Motus and Fast-WAM incur large naive deployment latency, which leads to sparse action updates and unstable closed-loop behavior. We deploy them with an RTC-style non-blocking execution scheme [38] and action interpolate between predicted chunks, enabling smoother control under high inference latency. AHAWAM is deployed with its asynchronous planner–executor interface, where the executor performs high-frequency closed-loop updates using the latest observation and reused planner context. Results. Figure 4 reports real-world performance under original settings and generalization shifts. In the original settings, AHA-WAM achieves 78.33% success, clearly outperforming the WAM baselines Motus (21.67%) and Fast-WAM (68.33%), while matching the strong generalist VLA baseline π0.5 (76.67%). Under generalization shifts, π0.5 obtains the highest success rate, while AHA-WAM remains second in success and achieves the highest progress score (35.00). This suggests that, despite not relying on π0.5 -scale generalist pretraining, AHA-WAM attains comparable real-world robustness and delivers the strongest deployment performance among WAM-based baselines. These results indicate that asynchronous world-action modeling improves deployment stability: long-horizon planner context supports task progress, while high-frequency execution enables timely physical correction.

4.5

Latency and Control Frequency

Finally, we compare closed-loop inference latency in Tab. 3. Fast-WAM is reported with official latency, while the remainings are measured on a single NVIDIA RTX 5090D GPU. Appendix D gives protocol and optimization details.

Table 3: Inference latency and frequency. Method

Lat. (ms) Freq. (Hz) Speedup

Motus 1866.10 Fast-WAM 190.00 AHA-WAM 41.37 AHA-WAM-Flash 17.56

0.54 5.26 24.17 56.95

0.10× 1.00× 4.59× 10.82×

Compared with Fast-WAM, AHA-WAM reduces latency from 190.00 ms to 41.37 ms, improving the closed-loop rate from 5.26 Hz to 24.17 Hz. With the distilled sampler, AHA-WAM-Flash further reaches 17.56 ms and 56.95 Hz, yielding a 10.82× speedup while keeping the same asynchronous planner–executor interface. 10

Original

Store-plate

Success Motus

Original

Organize-desktop

Lighting Variation

Original

Object Configuration Shift

21.7% 76.7%

Pi05

68.3%

FastWAM

78.3%

AHA-WAM

Score 18.75

Motus Pi05

39

FastWAM AHA-WAM

37.5

33.3%

86.7% 80.0% 93.3%

38

Pi05 FastWAM AHA-WAM

53.3% 46.7% 73.3%

6.7%

80.0% 60.0% 80.0%

Fold-towel

Generalization

Success Motus

13.3%

13.3%

66.7% 46.7% 46.7%

Prepare-soy-milk

Original

Texture Generalization

46.7%

40.0%

Novel Environment

Original

16.7% 55.0% 46.7% 53.3%

Score Motus

19

Pi05

33.25

FastWAM

31.5

AHA-WAM

35

93.3% 73.3% 86.7%

Motus

66.7% 60.0% 53.3%

Pi05

FastWAM

0.0%

46.7% 60.0% 53.3%

0.0%

33.3% 33.3% 40.0%

AHA-WAM

Figure 4: Real-world task success rates and scores. Success and 0–3 task scores are computed over 30 trials; scoring criteria are in Appendix C.

5

Conclusion

We presented AHA-WAM, an asynchronous horizon-adaptive world-action model that decouples a low-frequency video DiT world planner from a high-frequency action DiT closed-loop executor. Observation-Guided Video-Context Routing and horizon-adaptive offset training make this asynchronous interface effective by adapting reused planner context to each current observation and handling arbitrary planner-executor phase offsets. Experiments on RoboTwin and real-world manipulation tasks demonstrate strong performance and substantially reduced closed-loop latency, showing that learned visual dynamics can improve robot control without compromising control frequency. Limitations and future work. The planner update frequency, video horizon, and action chunk size introduce temporal hyperparameters whose optimal allocation may depend on task dynamics and embodiment. Future work could strengthen the slow planner with longer-horizon prediction and richer scene representations, while systematic evaluation on dedicated long-horizon benchmarks would better quantify how the asynchronous design scales with task complexity. More broadly, AHA-WAM opens a new design space for WAMs. Once the video branch is decoupled from the high-frequency control loop, it can afford more expensive computation without directly increasing action latency. The asynchronous interface makes these extensions particularly attractive: the video branch can become more predictive, deliberative, or physically grounded, while the action branch preserves the fast closed-loop rate needed for deployment. We believe this separation between slow scalable world planning and fast reactive action execution is a promising direction for building more capable, efficient, and deployable world-action models.

Acknowledgments We would like to express our sincere gratitude to the Baidu AI Cloud Baige Team for their exceptional technical support and for providing access to the state-of-the-art Baidu AIHC platform. We specifically appreciate the platform’s powerful capabilities in delivering efficient training acceleration and enabling ultra-low-latency inference, which were instrumental in optimizing our system’s performance during evaluation. The advanced system optimizations and robust distributed infrastructure provided by this team were crucial in accelerating our experiments and validating the scalability of our proposed methods.

11

References [1] A. Brohan, Y. Chebotar, C. Finn, K. Hausman, A. Herzog, D. Ho, J. Ibarz, A. Irpan, E. Jang, R. Julian, et al. Do as i can, not as i say: Grounding language in robotic affordances. In CoRL, 2023. [2] K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, et al. π0 : A vision-language-action flow model for general robot control. arXiv preprint arXiv:2410.24164, 2024. [3] T. Yang, G. Chen, Y. Chen, Z. Liang, Y. Liu, Z. Chen, C. Xu, H. Liang, J. Pang, Y. Mu, et al. HiVLA: A visual-grounded-centric hierarchical embodied manipulation system. arXiv preprint arXiv:2604.14125, 2026. [4] Z. Liang, Y. Mu, H. Ma, M. Tomizuka, M. Ding, and P. Luo. Skilldiffuser: Interpretable hierarchical planning via skill abstractions in diffusion-based task execution. In CVPR, 2024. [5] A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al. RT-1: Robotics transformer for real-world control at scale. In RSS, 2023. [6] B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid, et al. RT-2: Vision-language-action models transfer web knowledge to robotic control. In CoRL, 2023. [7] M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. P. Foster, P. R. Sanketi, Q. Vuong, et al. OpenVLA: An open-source vision-language-action model. In CoRL, 2025. [8] J. Zheng, J. Li, Z. Wang, D. Liu, X. Kang, Y. Feng, Y. Zheng, J. Zou, Y. Chen, J. Zeng, et al. Xvla: Soft-prompted transformer as scalable cross-embodiment vision-language-action model. arXiv preprint arXiv:2510.10274, 2025. [9] Physical Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, M. Y. Galliker, D. Ghosh, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, D. LeBlanc, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, A. Z. Ren, L. X. Shi, L. Smith, J. T. Springenberg, K. Stachowicz, J. Tanner, Q. Vuong, H. Walke, A. Walling, H. Wang, L. Yu, and U. Zhilinsky. π0.5 : A vision-languageaction model with open-world generalization. arXiv preprint arXiv:2504.16054, 2025. [10] S. Liu, L. Wu, B. Li, H. Tan, H. Chen, Z. Wang, K. Xu, H. Su, and J. Zhu. Rdt-1b: a diffusion foundation model for bimanual manipulation. In ICLR, 2025. [11] M. J. Kim, C. Finn, and P. Liang. Fine-tuning vision-language-action models: Optimizing speed and success. arXiv preprint arXiv:2502.19645, 2025. [12] Q. Bu, Y. Yang, J. Cai, S. Gao, G. Ren, M. Yao, P. Luo, and H. Li. Univla: Learning to act anywhere with task-centric latent actions. arXiv preprint arXiv:2505.06111, 2025. [13] Z. Liang, Y. Li, T. Yang, C. Wu, S. Mao, T. Nian, L. Pei, S. Zhou, X. Yang, J. Pang, et al. Discrete diffusion vla: Bringing discrete diffusion to action decoding in vision-language-action policies. In ICML, 2026. [14] J. Liu, H. Chen, P. An, Z. Liu, R. Zhang, C. Gu, X. Li, Z. Guo, S. Chen, M. Liu, et al. Hybridvla: Collaborative diffusion and autoregression in a unified vision-language-action model. arXiv preprint arXiv:2503.10631, 2025. [15] J. Chen, W. Song, P. Ding, Z. Zhou, H. Zhao, F. Tang, D. Wang, and H. Li. Unified diffusion vla: Vision-language-action model via joint discrete denoising diffusion process. arXiv preprint arXiv:2511.01718, 2025. 12

[16] C. Chi, Z. Xu, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song. Diffusion policy: Visuomotor policy learning via action diffusion. IJRR, 2025. [17] M. Janner, Y. Du, J. Tenenbaum, and S. Levine. Planning with diffusion for flexible behavior synthesis. In ICML, 2022. [18] W. Peebles and S. Xie. Scalable diffusion models with transformers. In ICML, 2023. [19] S. Gao, W. Liang, K. Zheng, A. Malik, S. Ye, S. Yu, W.-C. Tseng, Y. Dong, K. Mo, C.-H. Lin, et al. Dreamdojo: A generalist robot world model from large-scale human videos. arXiv preprint arXiv:2602.06949, 2026. [20] D. Hafner et al. DreamerV3: Mastering diverse domains through world models. arXiv preprint arXiv:2301.04104, 2023. [21] T. Wan, A. Wang, B. Ai, B. Wen, C. Mao, C.-W. Xie, D. Chen, F. Yu, H. Zhao, J. Yang, et al. Wan: Open and advanced large-scale video generative models. arXiv preprint arXiv:2503.20314, 2025. [22] T. Seedance, D. Chen, L. Chen, X. Chen, Y. Chen, Z. Chen, Z. Chen, F. Cheng, T. Cheng, Y. Cheng, et al. Seedance 2.0: Advancing video generation for world complexity. arXiv preprint arXiv:2604.14148, 2026. [23] Y. Du, S. Yang, B. Dai, H. Dai, O. Nachum, J. Tenenbaum, D. Schuurmans, and P. Abbeel. Learning universal policies via text-guided video generation. NeurIPS, 2023. [24] Y. Tian, S. Yang, J. Zeng, P. Wang, D. Lin, H. Dong, and J. Pang. Predictive inverse dynamics models are scalable learners for robotic manipulation. In ICLR, 2025. [25] Y. Hu, Y. Guo, P. Wang, X. Chen, Y.-J. Wang, J. Zhang, K. Sreenath, C. Lu, and J. Chen. Video prediction policy: A generalist robot policy with predictive visual representations. In ICML, 2025. [26] S. Ye, Y. Ge, K. Zheng, S. Gao, S. Yu, G. Kurian, S. Indupuru, Y. L. Tan, C. Zhu, J. Xiang, et al. World action models are zero-shot policies. arXiv preprint arXiv:2602.15922, 2026. [27] H. Bi, H. Tan, S. Xie, Z. Wang, S. Huang, H. Liu, R. Zhao, Y. Feng, C. Xiang, Y. Rong, et al. Motus: A unified latent action world model. arXiv preprint arXiv:2512.13030, 2025. [28] M. J. Kim, Y. Gao, T.-Y. Lin, Y.-C. Lin, Y. Ge, G. Lam, P. Liang, S. Song, M.-Y. Liu, C. Finn, et al. Cosmos policy: Fine-tuning video models for visuomotor control and planning. arXiv preprint arXiv:2601.16163, 2026. [29] L. Li, Q. Zhang, Y. Luo, S. Yang, R. Wang, F. Han, M. Yu, Z. Gao, N. Xue, X. Zhu, Y. Shen, and Y. Xu. Causal world modeling for robot control. arXiv preprint arXiv:2601.21998, 2026. [30] T. Yuan, Z. Dong, Y. Liu, and H. Zhao. Fast-WAM: Do world action models need test-time future imagination? arXiv preprint arXiv:2603.16666, 2026. [31] A. Ye, B. Wang, C. Ni, G. Huang, G. Zhao, H. Li, H. Li, J. Li, J. Lv, J. Liu, et al. Gigaworldpolicy: An efficient action-centered world–action model. arXiv preprint arXiv:2603.17240, 2026. [32] Q. Bu, H. Li, L. Chen, J. Cai, J. Zeng, H. Cui, M. Yao, and Y. Qiao. Towards synergistic, generalized, and efficient dual-system for robotic manipulation. arXiv preprint arXiv:2410.08001, 2024. [33] H. Xue, J. Ren, W. Chen, G. Zhang, F. Yuan, G. Gu, H. Xu, and C. Lu. Reactive diffusion policy: Slow-fast visual-tactile policy learning for contact-rich manipulation. In ICRA 2025 Workshop, 2025. 13

[34] N. Hirose, C. Glossop, D. Shah, and S. Levine. Asyncvla: An asynchronous vla for fast and robust navigation on the edge. arXiv preprint arXiv:2602.13476, 2026. [35] T. Chen, Z. Chen, B. Chen, Z. Cai, Y. Liu, Z. Li, Q. Liang, X. Lin, Y. Ge, Z. Gu, et al. Robotwin 2.0: A scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation. arXiv preprint arXiv:2506.18088, 2025. [36] Y. Yang, S. Zeng, T. Lin, X. Chang, D. Qi, J. Xiao, H. Liu, R. Chen, Y. Chen, D. Huo, et al. Abot-m0: Vla foundation model for robotic manipulation with action manifold learning. arXiv preprint arXiv:2602.11236, 2026. [37] S. Wu, X. Liu, S. Xie, P. Wang, X. Li, B. Yang, Z. Li, K. Zhu, H. Wu, Y. Liu, et al. RoboCOIN: An open-sourced bimanual robotic data collection for integrated manipulation. arXiv preprint arXiv:2511.17441, 2025. [38] K. Black, M. Galliker, and S. Levine. Real-time execution of action chunking flow policies. NeurIPS, 2026.

14

A

Implementation Settings

This appendix records implementation details that complement the experimental setup in Section 4.1. The main text describes the benchmark-level settings and the asynchronous planner– executor schedule; here we summarize the concrete implementation settings used for the RoboTwin runs. We omit logging paths, checkpoint paths, and inactive legacy annealing options. Model architecture. Table 4 provides the key architecture and hyperparameter settings for AHAWAM on RoboTwin, including the Video-DiT planner, Action-DiT executor, and OVCR interface. The instantiated model contains approximately 4.99B parameters in the video branch, 1.02B parameters in the action DiT, and 1.22B parameters in the memory and context-routing modules. During asynchronous inference, the action path receives the latest observation and proprioceptive state while querying the most recently available planner context through OVCR. Table 4: Implementation settings for AHA-WAM on RoboTwin. Component

Subcomponent

Setting

Backbone Observation cameras Image resolution History frames Video/action frequency ratio Video RoPE stride Routed transformer layers Attention heads / head dimension

Wan2.2-TI2V-5B video expert 3 views: head, left wrist, right wrist 384 × 320 6 8 8 30 24 / 128

Action-DiT

Action horizon / chunk size Number of action chunks State / proprio dimension Action dimension Training objective

64 actions / 16 actions 4 14 14 Action prediction with planner-context conditioning

OVCR

Method name Query source Queries per chunk Target context Granularity

Observation-Guided Video-Context Routing Visual observation context 32 Causal video K/V cache Per action chunk, per transformer layer

Optimizer / learning rate / weight decay Learning-rate schedule / warmup Global batch size / epochs Dataloader workers Video/action train timesteps Video/action train and inference shift

AdamW / 1 × 10−4 / 1 × 10−2 Cosine / first 5% of training 512 / 5 16 1000 5.0

Default action denoising steps / CFG

10 / 1.0

Component settings Video-DiT

Hyperparameter settings Training

Inference

ODE distillation. Table 5 lists the hyperparameter settings used for ODE-distilled fast sampling. It records the distillation target, teacher–student denoising schedules, and the main optimization settings; dataset paths, checkpoint paths, logging paths, and other run-specific bookkeeping are omitted. Table 5: Hyperparameter settings for ODE-distilled fast sampling. Item

Setting

Distillation target Teacher denoising steps Student denoising steps Distilled timesteps Teacher capture schedule Prediction parameterization Global batch size Learning rate Optimizer / weight decay Learning-rate schedule Epochs Local batch size / gradient accumulation Dataloader workers

Action ODE / flow trajectory 16 2 5,000 0, 1, 2, 4, 8, 12, 16 Flow 512 2 × 10−5 AdamW / 1 × 10−2 Cosine 5 16 / 4 16

15

Training configuration. The optimizer, scheduler, batch-size, epoch, and default inference settings for AHA-WAM are summarized in Table 4, while the corresponding ODE-distillation settings are summarized in Table 5. We use flow matching for both video prediction and action prediction, with noise times sampled from a logit-normal distribution; when both branches are optimized jointly, the video loss and action loss are weighted equally. Horizon-adaptive offset training follows Section 3.3, so each action chunk is trained to consume planner context under randomized planner–executor phase offsets. The distilled AHA-WAM-Flash variant uses the reduced-step sampler described in Appendix D.

B

RoboTwin Evaluation Details

This appendix provides the benchmark and baseline details omitted from the main RoboTwin discussion in Section 4.2. Benchmark protocol. RoboTwin 2.0 is evaluated with the AgileX embodiment and contains 50 dual-arm manipulation tasks spanning a broad range of bimanual skills. Following the multi-task setting used by prior WAM baselines, AHA-WAM is trained with 50 clean demonstrations and 500 randomized demonstrations per task, totaling 2,500 clean and 25,000 randomized demonstrations. Each task is evaluated with 100 clean episodes and 100 randomized episodes, and we report task-averaged success rates. The randomized setting introduces visual and scene-level variations, providing a stronger robustness test than clean evaluation. Baseline comparison. We compare AHA-WAM with representative VLA and WAM baselines. Fast-WAM is the closest baseline because it also uses a video DiT inside a world-action model, while Motus and LingBot-VA represent unified world-action modeling architectures. π0 , π0.5 , and ABot-M0 provide comparisons to strong generalist VLA policies. When available, we include both embodied-pretrained and Wan2.2-initialized variants to separate the effect of robot-data pretraining from the model architecture. In Table 1, Embodied PT. indicates whether the method uses embodied robot-data pretraining before RoboTwin 2.0 training. We report official published RoboTwin 2.0 results for external baselines when available.

C

Real-World Task Execution and Scoring Criteria

Figure 5 provides a step-by-step illustration of the four real-world manipulation tasks used in our evaluation. Each task is decomposed into three ordered subtask steps, which are executed sequentially during a rollout and directly define the partial-progress score used in Figure 4. For each model and task, success rate and score are computed over 30 independent trials. Success is recorded as a binary outcome, while the score is assigned on a 0–3 scale: a score of 0 indicates no meaningful task progress, and scores of 1, 2, and 3 indicate completion of subtask steps respectively. The task is counted as success when the rollout reaches score 3. This unified execution-and-scoring protocol makes the evaluation more informative than a binary success label alone, because it distinguishes early failures from trials that complete most of the required manipulation sequence. Together, these four tasks cover complementary real-world control challenges: rigid-object placement, deformable manipulation, long-horizon multi-object organization, and fine-grained tool use. Figure 5 and Table 6 therefore illustrate the task execution process and the basis for the 0–3 scoring criteria.

D

Inference Speedup Details

Measurement protocol. We report Lchunk , the end-to-end latency of a single action-chunk inference call. This includes image encoding, incremental planner-context access, OVCR context routing, and Action DiT denoising. Video-DiT prefill runs asynchronously with action inference; we therefore use Lchunk as the primary latency metric for closed-loop action frequency and separately 16

Subtask Steps (Executed in Sequence)

Task Spatial generalization

3

Store Plate

Deformable object manipulation

Fold Towel Long-horizon organization

Organize Desktop Fine-grained contact-rich tool use

Spoon Powder

Figure 5: Comprehensive illustration of the real-world task execution process. Each row corresponds to one task type, and the numbered panels show the ordered subtask steps executed in sequence. Table 6: Summary of real-world task objectives, partial-progress scoring criteria, and evaluated control challenges. All tasks use a unified 0–3 score, where score 0 indicates no meaningful progress and score 3 indicates full task completion. Task

Objective

Store Plate

Move a plate into a plate rack.

Score 1 2 3

Fold Towel

Fold a towel and place it into a basket.

Organize Desktop

Clear multiple desktop objects into their target containers.

Prepare soy milk powder using a cup and spoon.

The left hand securely grasps the plate. The plate is handed over from the left hand to the right hand. The right hand places the plate into the rack.

Evaluated challenges Bimanual transfer, spatial placement, and pose control.

1 2 3

Both hands fold the towel forward. The right hand folds the towel once more. The right hand deposits the folded towel into the basket.

Deformable-object handling and robustness to shape changes.

1

The left hand places left-side objects, such as the banana and carrot, onto the plate. The right hand places the pen and tool into the basket. The right hand places the plastic bag roll into the basket.

Multi-object sequencing, arm-specific selection, and subgoal tracking.

The right hand moves the cup to the workspace center. The left hand scoops soy milk powder with the spoon and pours it into the cup. The left hand stirs the mixture and returns the spoon.

Fine-grained tool use, contact-rich manipulation, and bimanual coordination.

2 3

Prepare Soy Milk

Completion criterion

1 2 3

track prefill throughput to ensure that planner context can be refreshed in the background. All latency values are measured in milliseconds on RoboTwin 2.0 tasks unless otherwise specified. The first warmup episode is discarded, and remaining per-episode means are averaged. Both baseline and optimized pipelines use bf16 precision, so the reported speedups do not come from reducing numerical precision. Table 7: Latency notation used in the inference speedup analysis. Symbol

Definition

Lchunk Lprefill

End-to-end latency of one action-chunk inference Latency of one asynchronous Video-DiT prefill call

17

CUDA acceleration. Table 8 reports the cumulative optimization path from PyTorch eager inference to the optimized CUDA deployment. The main deployment optimizations are grouped into three categories. First, graph-level static capture compiles backbone modules including the Action DiT, memory/context modules, and VAE encoder into TensorRT engines and replays the fixed denoising loop with CUDA Graphs, reducing Python dispatch and redundant kernel launch overhead. Second, selective torch.compile is applied to the Video-DiT prefill path, where full trace export is difficult due to Python-side control flow. Third, hot-path redundancy elimination hoists chunk-level computations outside the denoising loop and skips repeated host-to-device copies for tensors that remain fixed across denoising steps. Table 8: Cumulative CUDA acceleration ablation on RTX 5090D. CG denotes CUDA Graph. Stage 0 1 2 3 4 5 6 7

Optimization PyTorch eager baseline + Action DiT TRT + CG + Memory/context TRT + CG + Video-DiT prefill compile + Hot-path redundancy elimination + Runtime upgrade (A → B) + TRT video-KV skip-copy + VAE encoder TRT

Lchunk (ms)

Lprefill (ms)

Runtime

415.77 ± 0.33 83.87 ± 0.51 71.37 ± 0.15 71.45 ± 0.14 50.37 ± 0.27 47.72 ± 0.08 45.77 ± 0.03 41.37 ± 0.03

61.15 ± 0.26 – – 34.59 ± 0.08 – – – –

A A A A A B B B

Runtime A: PyTorch 2.7.1 + cu128 / Triton 3.3.1 / TensorRT 10.16.1.11. Runtime B: PyTorch 2.12.0 + cu130 / Triton 3.7.0 / TensorRT 10.16.1.11.

Prefill compilation. The goal of prefill optimization is to make planner-context refresh available fast enough for asynchronous deployment. Table 9 compares compile modes. Although reduce-overhead gives the lowest prefill latency, it increases Lchunk ; we therefore use the default mode. Table 9: Video-DiT prefill compile mode comparison. Configuration

Lchunk (ms)

Lprefill (ms)

71.37 71.45 75.08

57.15 ± 0.14 34.59 ± 0.08 25.70 ± 0.22

Eager (no compile) compile (default) compile (reduce-overhead)

Hot-path streamlining. The denoising loop contains computations that depend only on chunklevel inputs, such as condition embeddings, positional encodings, and references to planner-context K/V tensors. Hoisting these computations outside the 10-step loop reduces Lchunk from 71.45 ms to 63.25 ms. Removing redundant recursive state traversals over modules already in evaluation mode further reduces latency to 50.37 ms. The TensorRT wrapper also reuses static input buffers for video-KV tensors that remain unchanged within a chunk, reducing latency by another 1.95 ms in the 10-step setting. VAE encoder validation. We validate the VAE encoder TensorRT engine against a PyTorch eager reference using both random inputs and real RoboTwin frames. As shown in Table 10, the TensorRT output remains numerically close to the reference and produces no NaN or Inf values. ODE distillation. After CUDA acceleration establishes the optimized 10-step path, ODE distillation reduces the number of action denoising steps needed by the sampler. For AHA-WAM-Flash, we freeze the video DiT and distill only the action denoising path so that the student preserves the same planner-context and OVCR interface as the 10-step teacher. The teacher produces a 16-step denoising trajectory, indexed from the noisy state 0 to the final denoised state 16, and we select trajectory 18

Table 10: VAE encoder numerical validation against PyTorch eager reference. Metric

Random inputs (N = 1000)

Real frames (N = 1000)

0.999974 0.999897 0.031–0.078 1.24% 0

0.999941 0.999589 0.023–0.086 1.31% 0

Cosine similarity (mean) Cosine similarity (min) Max absolute deviation (min–max) Relative max deviation (mean) NaN/Inf count

anchors {0, 1, 2, 4, 8, 12, 16}. During training, the student randomly samples a non-final anchor as the starting state and directly predicts the teacher’s final denoised action state using a regression loss. We sample more frequently from states near the noisy end of the trajectory, because accurate prediction from high-noise states is the key requirement for reducing the number of inference denoising steps. Table 11: Latency of the CUDA-accelerated action sampler under different ODE-distilled denoising step counts. Denoising steps

Lchunk (ms)

Frequency (Hz)

Latency vs. 10-step

1 2 4 10

14.67 17.56 23.45 41.37

68.3 56.9 42.6 24.2

−64.60% −57.50% −43.30% –

19

E

Per-Task RoboTwin Success Rates

Table 12 reports per-task success rates on RoboTwin 2.0 for AHA-WAM and the selected baselines. All values are percentages, with clean and randomized evaluation reported separately. Table 12: Per-task success rates on RoboTwin 2.0 under clean and randomized evaluation settings. Task

AHA-WAM

AHA-WAM-Flash

Fast-WAM

LingBot-VA

π0.5

Motus

Clean

Rand.

Clean

Rand.

Clean

Rand.

Clean

Rand.

Clean

Rand.

Clean

Rand.

Adjust Bottle Beat Block Hammer Blocks Ranking RGB Blocks Ranking Size Click Alarmclock Click Bell Dump Bin Bigbin Grab Roller Handover Block Handover Mic Hanging Mug Lift Pot Move Can Pot Move Pillbottle Pad Move Playingcard Away Move Stapler Pad Open Laptop Open Microwave Pick Diverse Bottles Pick Dual Bottles Place A2B Left Place A2B Right Place Bread Basket Place Bread Skillet Place Burger Fries Place Can Basket Place Cans Plasticbox Place Container Plate Place Dual Shoes Place Empty Cup Place Fan Place Mouse Pad Place Object Basket Place Object Scale Place Object Stand Place Phone Stand Place Shoe Press Stapler Put Bottles Dustbin Put Object Cabinet Rotate QRcode Scan Object Shake Bottle Shake Bottle Horizontally Stack Blocks Three Stack Blocks Two Stack Bowls Three Stack Bowls Two Stamp Seal Turn Switch

100% 100% 100% 98% 100% 100% 99% 100% 96% 99% 85% 100% 78% 100% 100% 88% 98% 40% 89% 98% 100% 96% 89% 91% 99% 81% 100% 100% 91% 100% 100% 93% 81% 96% 94% 98% 91% 98% 83% 91% 90% 94% 100% 100%

100% 94% 98% 98% 100% 100% 92% 100% 85% 99% 78% 100% 83% 99% 100% 77% 97% 46% 85% 99% 92% 92% 94% 87% 98% 94% 100% 96% 99% 100% 87% 82% 80% 87% 93% 96% 96% 99% 96% 89% 90% 90% 100% 100%

100% 100% 100% 97% 100% 100% 95% 100% 97% 94% 67% 100% 89% 96% 98% 75% 98% 9% 83% 96% 96% 98% 89% 87% 94% 83% 98% 100% 92% 100% 96% 79% 89% 89% 96% 98% 94% 98% 83% 89% 88% 85% 100% 100%

100% 93% 100% 91% 100% 100% 95% 100% 87% 92% 53% 100% 94% 98% 100% 63% 98% 23% 87% 98% 98% 96% 98% 85% 98% 79% 100% 94% 89% 98% 94% 81% 85% 77% 90% 94% 98% 100% 77% 85% 90% 91% 100% 100%

100% 99% 100% 94% 100% 100% 97% 100% 95% 99% 58% 100% 90% 100% 100% 77% 98% 62% 80% 100% 95% 93% 91% 90% 96% 71% 99% 96% 94% 100% 96% 83% 89% 90% 90% 97% 96% 90% 95% 94% 93% 89% 100% 100%

100% 97% 100% 98% 100% 100% 96% 100% 81% 100% 62% 100% 88% 99% 100% 64% 100% 45% 85% 96% 93% 99% 93% 93% 99% 69% 96% 100% 88% 100% 96% 89% 88% 97% 94% 99% 99% 97% 90% 89% 89% 92% 100% 100%

90% 96% 99% 94% 99% 100% 89% 100% 99% 94% 40% 100% 94% 99% 100% 91% 92% 82% 89% 100% 97% 97% 97% 95% 97% 81% 100% 99% 94% 100% 99% 93% 91% 96% 99% 97% 98% 85% 87% 85% 96% 96% 100% 100%

94% 98% 98% 96% 100% 100% 96% 100% 78% 96% 28% 99% 97% 99% 99% 79% 94% 86% 82% 99% 93% 95% 95% 90% 95% 84% 99% 97% 89% 100% 93% 96% 88% 95% 96% 97% 98% 82% 91% 87% 91% 91% 97% 99%

100% 96% 92% 49% 98% 99% 92% 100% 66% 98% 18% 96% 51% 84% 96% 56% 90% 34% 81% 93% 87% 87% 77% 85% 94% 62% 94% 99% 75% 100% 87% 60% 80% 86% 91% 81% 92% 87% 84% 80% 89% 72% 99% 99%

99% 93% 85% 26% 89% 66% 97% 100% 57% 97% 17% 85% 55% 61% 84% 42% 96% 77% 71% 63% 82% 84% 64% 66% 87% 62% 84% 95% 75% 99% 85% 39% 76% 80% 85% 81% 93% 83% 79% 79% 87% 65% 97% 99%

89% 95% 99% 75% 100% 100% 95% 100% 86% 78% 38% 96% 34% 93% 100% 83% 95% 95% 90% 96% 88% 91% 91% 86% 98% 81% 98% 98% 93% 99% 91% 66% 81% 88% 98% 87% 99% 93% 81% 88% 89% 67% 100% 100%

93% 88% 97% 63% 100% 100% 91% 100% 73% 63% 38% 99% 74% 96% 96% 85% 91% 91% 91% 90% 79% 87% 94% 83% 98% 76% 94% 99% 87% 98% 87% 68% 87% 85% 97% 86% 97% 98% 79% 71% 73% 66% 97% 98%

99% 100% 79% 100% 98% 70%

98% 99% 82% 100% 90% 74%

100% 100% 81% 94% 81% 53%

98% 100% 83% 98% 83% 65%

95% 100% 80% 92% 90% 61%

97% 100% 81% 98% 94% 59%

99% 100% 86% 94% 96% 44%

98% 98% 83% 98% 97% 45%

91% 97% 77% 95% 79% 62%

76% 100% 71% 96% 55% 54%

91% 100% 79% 98% 93% 84%

95% 98% 87% 98% 92% 78%

Average

93.40% 92.20% 90.48% 89.92%

91.88% 91.78% 92.90% 91.50% 82.74% 76.76% 88.66% 87.02%

20

Record · ID 267680 · SHA-256 b7035e82cbba6f7c
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.