Agile-WAM: An Agile Tactile World Action Model for Contact-Rich Robot Control
arXiv:2609.20761v1 [cs.RO] 17 Sep 2026
Hanchu Zhou1,⋆ , Brendan Lynch2 , Raman Goyal2 , Dechen Gao1 , Begum Kasap2 , Boqi Zhao1 , Junshan Zhang1
Fig. 1: Agile-WAM, a tactile World Action Model, takes multi-modal inputs (including visual and tactile observations) and performs joint action–vision–tactile prediction through a direct vision-tactile-to-action flow-matching (FM) backbone. To provide effective supervision for future latent prediction, Agile-WAM introduces multi-horizon multimodal prediction to enable joint prediction of visual and tactile latent over different horizons. By capturing the distinct temporal dynamics of the two modalities, Agile-WAM aligns physical dynamics modeling better with action generation and enables more precise, fine-grained manipulation. Abstract— World Action Models (WAMs) advance beyond conventional visuomotor policies by jointly predicting future world states and robot actions, enabling the policy to learn physical dynamics that support effective control. However, recent tactile WAMs often rely on large-scale pretrained generative backbones to capture contact-rich physical dynamics, which limit their inference efficiency and flexible deployment. In this paper, we present Agile-WAM, an agile tactile World Action Model for contact-rich robot control. Agile-WAM encodes visual and tactile observations into a shared latent that serves as the source of a direct vision-tactile-to-action flow-matching process, which can jointly generate latent representations of action 1 Hanchu Zhou, Dechen Gao, Boqi Zhao, and Junshan Zhang are with University of California, Davis, One Shields Avenue, Davis, CA, USA.
[email protected] 2 Raman Goyal, Begum Kasap, and Brendan Lynch are with Analog Devices, San Jose, CA, USA. [email protected] (Corresponding author: Brendan Lynch) ⋆ Work done during an internship at Analog Devices.
chunks and future visual/tactile latents. A key observation is that vision and tactile signals evolve at inherently different timescales: adjacent visual frames are often highly similar, whereas tactile signals can change abruptly upon contact. We therefore introduce multi-horizon multimodal prediction in Agile-WAM, which provides supervision for visual latent at a larger temporal offset while predicting the tactile latent in the next frame to capture fine-grained contact dynamics. Across nine simulated and five real-world contact-rich manipulation tasks, Agile-WAM demonstrates strong and robust performance, outperforming the strongest baseline in success rate while maintaining low inference latency. In particular, in five real-world experiments, Agile-WAM yields a relative gain of 29.4% in overall success rates while achieving inference latency of 11.9 ms. These results demonstrate that multimodal WAM can be achieved with an agile architecture suitable for precise and high-frequency robot control. More details are available on our project page.
I. INTRODUCTION World Action Models (WAMs) have recently emerged as a promising paradigm for robot learning by coupling physical dynamics prediction with action generation. In contrast to conventional imitation learning or reinforcement learning paradigms [1], WAMs learn how the world will evolve and which actions should be executed jointly [2], [3], [4], [5]. By predicting future observations from consecutive frames, the WAM model is encouraged to capture dense spatiotemporal information about motion and interaction in the physical scene. Jointly modeling these future states, together with robot actions, further fosters the learned representation to connect physical dynamics evolution with action control. In this way, WAMs acquire a structured look-ahead capability and can produce actions that are more consistent with the anticipated evolution of the physical world. Clearly, physical interaction is not fully observable from vision alone. This limitation is particularly pronounced in contact-rich manipulation, where visually similar observations may correspond to substantially different physical states, especially when the end effector contacts a rigid object with different levels of force. These tasks often require delicate and precise action adjustments despite exhibiting only subtle visual differences. Tactile sensing provides a complementary channel that captures local contact, force, and slip, and has consequently become increasingly important for learning precise manipulation policies [6], [7]. Recent studies have begun to incorporate tactile sensing into world-action modeling by extending large pretrained video generation models with tactile observations and actions [8], [9], [10]. These early attempts demonstrate that explicitly modeling tactile dynamics can benefit contactrich manipulation. However, reliance on large pretrained backbones introduces substantial computational and memory overhead during inference, which limits deployment on robotic platforms with constrained onboard resources and demanding high control frequencies. This gives rise to a fundamental question: Is it possible to develop an agile tactile-centric world-action model that enables flexible deployment across robotic platforms, while retaining the predictive benefits of world models? To tackle this challenge, we first note that visual and tactile observations evolve at inherently different time scales, in the following sense: Visual appearance typically changes smoothly and slowly, making adjacent visual frames highly redundant, whereas tactile force field (TacFF) signals can change abruptly upon contact. It would be naive to apply an identical prediction horizon to both visual and tactile modalities, which would otherwise fail to account for their distinct temporal characteristics: Near-term visual observations often exhibit high redundancy, providing limited supervision for world modeling, while long-horizon tactile prediction may be less precise due to the rapid evolution of contact signals. We address this mis-alignment through multi-horizon multimodal prediction. Specifically, visual prediction is supervised at a larger temporal offset, where the observation
exhibits more noticeable changes, encouraging the model to capture meaningful visual dynamics. In contrast, tactile prediction targets the next adjacent observation, allowing the model to track rapid contact changes without skipping finegrained tactile transitions. This separate supervision mechanism better matches the temporal characteristics of each modality and supports precise contact-rich manipulation. With this insight, we introduce Agile-WAM, an agile tactile World Action Model for contact-rich robot control. Rather than finetuning a large video foundation model, Agile-WAM employs a vision-tactile-to-action flowmatching backbone trained from scratch using visual observations and TacFF signals. The direct transition from vision-tactile to action obviates the need for repeated visual conditioning during the flow, thus significantly reducing the complexity [11]. As illustrated in fig. 1, the model first encodes both visual and tactile modalities into fused latent representations, and then jointly predicts future visual observations and future tactile observations using multihorizon multimodal prediction, as well as action chunks. Specifically, the fused latent representation, which captures both visual and tactile information, serves as the source distribution. A vision-tactile-to-action flow-matching network then learns a velocity field that directly transports this source representation toward a target latent composed of three parts: the action latent, future visual latent, and future tactile latent. We supervise these components with modality-specific objectives, applying the action loss in the action space and the visual and tactile prediction losses in their respective latent spaces. This joint objective encourages the policy representation to capture both visual evolution and local contact dynamics, thereby better aligning action generation with future dynamics and enabling fine-grained actions required for contact-rich manipulation. Meanwhile, the visiontactile-to-action flow-matching backbone eliminates the need for costly conditioning mechanisms and enables a direct mapping from observations to actions and future observations using a compact backbone, thereby improving action precision and inference efficiency. We evaluate Agile-WAM on nine simulated and five realworld contact-rich manipulation tasks, with focus on its performance and inference efficiency. Extensive ablations further isolate the contributions of tactile sensing and the multi-horizon multimodal prediction. The main contributions of this paper are summarized as follows: • We introduce Agile-WAM, an agile tactile World Action Model that encodes visual and tactile multi-modal observations into a shared latent space, which directly drives a vision-tactile-to-action flow-matching model to jointly predict future visual observations, TacFF observations, and robot actions. In particular, the visiontactile-to-action mechanism substantially reduces the computational overhead by removing cross-attention, thus improving inference efficiency. • We propose multi-horizon multimodal prediction that supervises visual and tactile signals at different temporal
offsets, capitalizing their inherent physical dynamics across distinct timescales, thus improving action generation for contact-rich manipulation. • We evaluate Agile-WAM on nine simulated and five real-world contact-rich tasks. In particular, in five realworld experiments, Agile-WAM yields a relative gain of 29.4% in overall success rates while achieving inference latency of 11.9 ms. Comprehensive ablation studies demonstrate the effectiveness of each design choice and model component. II. R ELATED W ORK A. Tactile Integration in Robot Learning Tactile sensing provides complementary physical information on contact surface that is difficult to infer from vision alone. Recent work has developed transferable tactile representations across different sensors and modalities through self-supervised learning, multimodal alignment, and sensorinvariant representation learning [12], [13], [14]. Beyond representation learning, tactile and force feedback have also been incorporated directly into manipulation policies to improve contact-rich control, with recent approaches emphasizing force-aware policy learning and fast tactile feedback for reactive manipulation [7], [6]. These studies demonstrate the importance of tactile feedback for physical interaction, but most use tactile signals primarily as additional policy inputs or reactive feedback. In contrast, Agile-WAM explicitly models the future evolution of tactile observations together with visual observations and actions, allowing tactile dynamics to directly participate in predictive policy learning. B. World Action Models for Robotics World Action Models extend conventional visuomotor policies by jointly learning future world evolution and robot actions, encouraging the policy to capture physical dynamics that are useful for control. Recent approaches have explored joint observation–action modeling [2], [3], [4]. More recent work further scales this paradigm with large generative backbones, demonstrating that jointly modeling future observations and actions can improve physical generalization and support closed-loop robot control [5]. While most existing approaches focus primarily on visual observations, concurrent studies have begun incorporating tactile signals into predictive world-action modeling for contact-rich manipulation [8], [9]. However, these tactile extensions largely build upon large pretrained backbones. Agile-WAM instead studies an agile alternative that jointly models visual, tactile, and action dynamics with an efficient architecture designed for high-frequency robot control and low-resource deployment. C. Flow Matching for Generative Model and Robot Control Flow matching learns a continuous vector field that transports samples between source and target distributions, enabling high-quality generation with efficient inference and making it well suited for continuous action generation [15], [16]. In robot learning, flow-based policies have been applied to action generation from visual and 3D observations, while
recent methods further improve inference efficiency by reducing the number of flow-matching steps required for action generation [17], [18], [19]. Flow matching has also been successfully scaled to large generalist robot policies, demonstrating its effectiveness for learning complex and multimodal action distributions [20]. More recent work develops noisefree flow-matching policies by replacing Gaussian noise with informative representations derived from observations or previous actions, achieving strong control performance with substantially improved inference efficiency [11], [19]. Inspired by this direction, Agile-WAM adopts vision-tactileto-action direct flow matching as an efficient backbone and generalizes it beyond action generation to jointly model future visual and tactile observations. III. METHODOLOGIES Agile-WAM aims to capture multimodal world dynamics in a shared latent space to support high-quality action generation. In this section, we describe how this objective is realized with a lightweight vision-tactile-to-action flow matching backbone and several key design choices. In section III-A, we formulate the direct mapping from current observations to joint action-future predictions. In section IIIB, we present the overall architecture of Agile-WAM and its training objectives. In section III-C, we introduce multihorizon multimodal prediction designed to better capture the distinct dynamics of different modalities. A. From Observation to Action–Vision–Tactile Joint Prediction We formulate robot control as joint action and future latent prediction. At environment step t, the robot receives a multimodal observation Ot = Otvis , Ottac , qt , where Otvis denotes the visual observation, Ottac denotes the tactile force field (TacFF), and qt denotes robot’s proprioceptive states. Rather than learning only an action generation policy, our policy π at:t+H , ztfuture | Ot jointly models an action chunk and the future evolution of both modalities, where at+H denotes an action chunk with prediction horizon H, of which only the first h actions are executed before replanning. vis tac ztfuture = zt+h , zt+h comprises latent representations vis tac of future observations. We emphasize that jointly predicting future observations and actions encourages the learned representation to capture subtle physical dynamics that are directly relevant to contact-rich robot control, rather than relying only on the simple mapping from observation to action. To this end, we develop this joint prediction algorithm using vision-tactile-to-action flow matching. Traditional flowbased policies typically start from a random Gaussian noise and repeatedly inject observations through conditioning modules, such as cross-attention. Instead, our backbone uses the latent representation of the current multimodal observation directly as the sole source of the flow, eliminating the need for an explicit conditioning mechanism during generation. This substantially simplifies the generation process and reduces computational overhead, leading to improved inference efficiency [11].
Fig. 2: Overview of Agile-WAM. Agile-WAM encodes visual, tactile, and proprioceptive observations into a shared latent that serves as the source of a lightweight vision-tactile-to-action flow-matching backbone. The flow jointly generates latent representations of the action chunk and future visual and tactile observations. An action autoencoder provides a structured action latent space, while multi-horizon multimodal prediction uses a longer horizon for vision and a short horizon for tactile feedback to capture their distinct temporal dynamics. The model is trained end-to-end with flow-matching, action reconstruction and generation, and visual/tactile latent prediction losses. We use learned encoders to encode the image and TacFF observations respectively, and fuse them into observation latent z0 as the source of the flow. The target of the flow is divided into the action and future multimodal observations: z1 = z1a , z1vis , z1tac . The source and target are constructed to have the same dimensionality, as required by flow matching. For a flow time τ ∈ [0, 1], we define the linear interpolation zτ = (1 − τ )z0 + τ z1 , whose target velocity is z1 − z0 . A lightweight flow network vθ learns the vision-tactile-to-action velocity field through h i 2 LFM = Eτ,z0 ,z1 ∥vθ (zτ , τ ) − (z1 − z0 )∥2 . Since sensory information is already embedded in the source z0 , vθ does not require an additional observationconditioning module during ODE integration. At inference time, the current observation is encoded once into z0 , after which we solve dzτ = vθ (zτ , τ ), zτ =0 = z0 , dτ from τ = 0 to τ = 1. This produces the predicted joint latent ẑ1 = ẑ1a , ẑ1vis , ẑ1tac , from which the action chunk is decoded and executed. B. The Design of Agile-WAM The overall architecture of Agile-WAM is illustrated in fig. 2. The model consists of modality-specific encoders, an action autoencoder, and a lightweight flow-matching network. a) Multimodal observation encoding: At the input side, the learned visual encoder and tactile encoder extract representations from the current RGB image Otvis and TacFF Ottac signal respectively. Their features are concatenated with proprioceptive states qt and subsequently fused into the source latent z0 through linear projection. In this way, visual and tactile information is incorporated once at the beginning of the flow, allowing the subsequent generation process to
operate directly in the fused latent space without repeatedly injecting conditions. b) Latent action representation: Inspired by VITA [11], we adopt the same design on action decoding that provides supervision on action space to prevent the collapse of action generation. We therefore introduce an action encoder Ea and decoder Da to construct a structured latent action space: z1a = Ea (at:t+H ),
ãt:t+H = Da (z1a ).
The action autoencoder is trained end-to-end with the policy using LAE = ∥at:t+H − Da (Ea (at:t+H ))∥1 . (1) During inference, however, the action decoder receives the ODE-generated latent ẑ1a rather than the encoder-generated target z1a . To align the decoder with the action latent generated via flow matching, we additionally decode the ODEgenerated latent during training to provide supervision on action space: Lact = ∥at:t+H − Da (ẑ1a )∥1 .
(2)
Gradients from this objective are propagated through the action decoder and the ODE integration process, directly anchoring the generated latent to executable ground-truth actions. Together, Eqs. (1) and (2) stabilize the learned action representation while reducing the discrepancy between training-time target latents and inference-time generated latents. C. Multi-horizon Multimodal Prediction A key challenge in multimodal world modeling is that different sensing modalities evolve at different temporal scales. Consecutive visual observations often exhibit substantial redundancy because scene appearance changes relatively smoothly. In contrast, tactile signals can vary sharply within a short period when the robot establishes contact or encounters resistance. Consequently, imposing the same prediction horizon on both modalities can lead to mismatched
learning signals: a short visual horizon provides only trivial supervision due to the limited changes between nearby frames, whereas a long tactile horizon may blur the finegrained contact dynamics essential for precise control. To account for these heterogeneous dynamics, we introduce multi-horizon multimodal prediction. Instead of predicting vision and tactile latents at the same future step, their target latents are constructed as vis tac z1vis = Evis Ot+h , z1tac = Etac Ot+h vis tac In our design, tactile prediction focuses on the next future frame, i.e., htac = 1, while visual prediction uses a longer temporal offset hvis = h that matches the executed action length at inference. The longer visual horizon encourages the model to capture meaningful scene evolution rather than collapsing to duplicate current frames, whereas the shorter tactile horizon preserves abruptly changing contact information. After solving the flow ODE, the predicted joint latent is decomposed into action, visual, and tactile components. Multimodal future prediction is supervised in latent space: 2
Lvis = ẑ1vis − z1vis 2 ,
2
Ltac = ẑ1tac − z1tac 2
Unlike pixel-level reconstruction, latent prediction provides a compact learning objective that encourages the shared flow representation to capture task-relevant evolution of both modalities without requiring expensive high-dimensional observation generation. Together, the complete training objective is LAgile-WAM = λFM LFM + λAE LAE + λact Lact + λvis Lvis + λtac Ltac , where λFM , λAE , λact , λvis , and λtac control the relative contributions of the five objectives. Through this joint optimization, Agile-WAM learns a compact flow representation that simultaneously captures multimodal world modeling and generates high-quality action sequences for contact-rich manipulation. IV. EXPERIMENTS We evaluate Agile-WAM on nine simulated and five real-world contact-rich manipulation tasks. The simulation experiments are conducted using ManiFeel [21], a comprehensive benchmark for tactile manipulation policy learning that provides realistic tactile simulation and challenging tasks in which tactile feedback is essential. The platform uses a 7-DoF Franka Emika Panda robot equipped with a TacFF sensor on its gripper. The TacFF sensor has a resolution of 10 × 14, with each sensing point measuring force magnitude and directions along the x- and y-axes, resulting in a 10 × 14 × 3 tactile observation. For each task, we use 20–50 demonstrations from the official dataset for training. Following the benchmark’s standard setup, only the wrist camera of 256 × 256 is used during training and evaluation. Because the wrist camera view is frequently occluded as the gripper interacts with the target object, this setting highlights the importance of tactile feedback.
For the real-world experiments, we use a 7-DoF Flexiv Rizon 4 robot to evaluate five challenging tasks: Gear Assembly, Peg Insertion, Internet Cable Insertion, Internet Cable Unplugging, and Power Plug Insertion, as illustrated in fig. 3. Visual observations are captured using an Intel RealSense D405 RGB-D wrist camera at a resolution of 320 × 240 and a frame rate of 30 FPS. Tactile feedback is acquired at 30 Hz using an Analog Devices 32 × 32 piezoresistive pressure sensor mounted on the gripper. For each task, we collect 50 expert demonstrations through teleoperation using a Meta Quest 3 headset.
Fig. 3: Real-world experiments include five challenging tasks: Gear Assembly, Peg Insertion, Internet Cable Insertion, Internet Cable Unplugging, and Power Plug Insertion. A. Experiment Settings Backbones. We use an ImageNet-pretrained ResNet-18 as the visual encoder and an MLP-based autoencoder to encode and reconstruct action sequences. Because the TacFF signal has a spatial structure analogous to that of an image, we employ a separate ResNet-18, trained from scratch, as the tactile encoder to extract contact-related features. For flow matching, we adopt a lightweight MLP to parameterize the velocity field. During inference, we integrate the learned ODE using an explicit Euler solver with 6 steps to generate action and future latents. Baselines. We compare Agile-WAM with state-of-the-art vision-based and tactile-aware policies, including Diffusion Policy with visual observations (DP), Diffusion Policy with visual and tactile observations (DP-VT) [21], Tactile-WAM [9], and the vision-only (VITA) [11] and vision–tactile variants of VITA (VITA-VT). For a fair comparison, we reproduce Tactile-WAM by incorporating its core component, the Tactile Asymmetric Attention mechanism, into the tactile diffusion-policy implementation provided by ManiFeel. We additionally construct the VITA-VT by augmenting the original vision-only VITA with tactile observations using the same multimodal fusion strategy as Agile-WAM. The performance and efficiency comparisons are presented in section IV-B. Real-World Tasks. We design five challenging real-world manipulation tasks that require tactile feedback for precise object alignment and correction. As illustrated in fig. 3, the tasks are: Gear Assembly: The gear set is 3D-printed using the same assets as the corresponding ManiFeel task. The robot must
TABLE I: Success rates comparison on simulation tasks
Task
Agile-WAM
VITA-VT
VITA
Tactile-WAM
DP-VT
DP
Bulb Screw Gear Assembly Power Plug Insertion Peg Reorientation Peg Insertion USB Insertion Ball Sorting Object Search Nut Bolt Threading
92.67±2.49 70.67±2.49 59.33±0.94 42.00±1.63 49.33±0.94 62.00±5.89 92.00±0.00 50.67±2.49 92.67±2.49
87.33±0.94 68.67±1.89 63.33±3.27 28.67±0.94 47.33±2.49 59.33±9.43 90.00±3.27 36.00±2.83 92.00±2.83
71.33±0.94 68.33±0.47 54.67±3.40 41.33±4.11 44.00±1.67 55.33±4.11 72.00±7.12 40.00±3.27 86.67±4.71
4.00±0.00 68.00±1.63 61.33±2.49 34.67±2.49 40.00±9.80 44.67±8.06 45.33±0.94 33.33±2.49 13.33±2.49
4.67±2.49 58.67±6.18 58.00±3.27 28.00±2.83 20.67±13.20 38.00±2.83 74.67±5.25 13.33±0.94 72.00±3.27
5.33±0.94 58.67±1.89 53.67±1.25 22.00±1.63 21.33±18.86 41.33±5.25 38.67±3.40 14.67±3.40 4.67±0.94
TABLE II: Success rate comparison on real-world tasks. Policy
Gear Assembly
Peg Insertion
Power Plug Insertion
Agile-WAM VITA-VT VITA
0.80 (16/20) 0.55 (11/20) 0.75 (15/20)
0.55 (11/20) 0.60 (12/20) 0.25 (5/20)
0.65 (13/20) 0.35 (7/20) 0.25 (5/20)
Policy
Internet Cable Insertion
Internet Cable Unplugging
0.30 (6/20) 0.15 (3/20) 0.20 (4/20)
0.85 (17/20) 0.75 (15/20) 0.20 (4/20)
Agile-WAM VITA-VT VITA
insert the middle gear onto its shaft and then perform slight rotation until its teeth properly engage with the neighboring gears. Peg Insertion: The peg and base are 3D-printed from the ManiFeel assets. The robot must align the peg with the hole and insert it successfully. Internet Cable Insertion: The robot inserts a standard Ethernet cable with a locking tab into the port. It must carefully align the port and push it in until it locks. Internet Cable Unplugging: The robot must grasp the Ethernet cable, depress the locking tab to release the connector, and then pull the cable out of the port. Power Plug Insertion: The robot must align and insert a two-prong power plug into a surge protector. As the plug approaches the outlet, it occludes most of the wrist-camera view, making tactile feedback particularly important. Training and Evaluation. In both simulation and realworld settings, all methods use an action chunk horizon H = 16 and execute h = 8 actions at each inference step. We use 6 ODE integration steps for VITA, VITA-VT, and Agile-WAM, and 100 denoising steps for DP and DP-VT. During training, we evaluate each policy every 500 training steps with 50 evaluation rollouts in simulation. Policies are trained for 40k–100k steps to ensure convergence of the success rate. For each task, we report the highest success rate achieved during training, averaged over three random seeds. In the real-world experiments, each policy is evaluated for 20 episodes per task. All training and evaluation experiments can be performed on a single NVIDIA RTX 4090 GPU. B. Performance In this section, we highlight the performance of Agile-WAM that it matches or outperforms the state-of-
the-art methods with respect to success rate. Additionally, we demonstrates its competitive inference efficiency against baseline while retain the benefit of world modeling. 1) Success Rates: We evaluate Agile-WAM on nine simulation tasks and five real-world tasks against state-of-theart baselines. Simulation and real-world success rates are reported in section IV and table II, respectively. By leveraging both learned world dynamics and tactile feedback, Agile-WAM exhibits effective corrective behaviors in both simulation and real-world experiments. In the real-world demonstrations, we deliberately include trajectories in which the object is not precisely installed on the first attempt and the robot must perform corrective motions to complete the task. With the limited view of the wrist camera, these corrections are particularly challenging, as the policy must maintain sufficient contact to acquire informative tactile feedback while avoiding damage to the objects or gripper. Benefiting from tactile signal and learned dynamics, Agile-WAM efficiently learns these delicate behaviors and actively attempts recovery when the initial insertion fails, as illustrated in fig. 4.
Fig. 4: Agile-WAM effectively learns corrective behaviors in real-world experiments. When the initial attempt fails to align the object with the target position, Agile-WAM can still recover under an obstructed visual view by gently maintaining contact with the surrounding surface and using tactile feedback to search for the correct insertion position. This recovery capability substantially improves the success rate over baselines, which often fail to correct the misalignment. For example, during peg insertion, the wrist camera view becomes partially occluded by the peg as it approaches the bottom, making it difficult to determine visually whether the
peg is aligned with the hole or pressed against its edge. In such cases, Agile-WAM responds to tactile feedback by sliding the peg along the surface toward the hole until successful insertion. Similarly, during gear assembly, the middle gear can become stuck against the other gears and fail to engage properly. Agile-WAM detects this stagnation and performs corrective twisting motions until the gear becomes aligned and fits into place. In contrast, VITA relies solely on visual observations and therefore has limited awareness of contact states. VITA-VT incorporates tactile observations but often fails to learn stable recovery behaviors, tending to become stuck or push blindly after an unsuccessful initial attempt. These results demonstrate that joint action–future modeling encourages the policy representation to capture both visual evolution and local contact dynamics, enabling the fine-grained corrective actions required for contact-rich manipulation. 2) Inference Efficiency: We compare the inference efficiency of Agile-WAM with several baselines. All measurements are conducted on a single NVIDIA RTX 4090 in FP32 with a batch size of 1. Each method generates an action chunk with a horizon of 16, and the reported inference time is averaged over 50 runs. We follow the same sampling settings as in the simulation experiments, using 6 sampling steps for VITA and Agile-WAM, and 100 denoising steps for DP, DP-VT and TAAM. Inference time is measured endto-end from observation input to action generation, while the corresponding control frequency represents the maximum frequency supported by policy inference. In practical deployment, the actual control frequency may additionally be limited by factors such as the robot control interface and sensor sampling rate. Additionally, for easier comparison, we normalize the metric of VITA to 1× and report the relative values for the other policies. The results in table III show that Agile-WAM maintains low inference latency comparable to VITA while achieving a control frequency up to 40× higher than DP. Although Agile-WAM shares the same lightweight flowmatching backbone as VITA, it additionally incorporates tactile encoding and a larger latent space for future prediction. These additions introduce little inference overhead due to the compact model design. Although Agile-WAM operates on a larger latent space, the additional computation is efficiently parallelized on the GPU, resulting in inference latency comparable to VITA. As a result, Agile-WAM remains an agile yet capable World Action Model, making it suitable for highfrequency robot control and potentially for deployment on computationally constrained edge platforms. C. Ablation of World Modeling We investigate the benefits of world modeling by ablating the prediction losses and the modalities involved in future prediction. As shown in fig. 5(a), Agile-WAM outperforms the baseline without joint action–future modeling. As joint prediction requires a larger flow-matching latent space than the non-predictive baseline, we further disable the worldmodeling loss terms Lvis and Ltac of Agile-WAM while
TABLE III: Inference efficiency comparison. Inference Time Model
Control Frequency
Value (ms)
Relative
Value (Hz)
Relative
Agile-WAM
10.35±0.14
1.07×
96.62
0.94×
VITA VITA-VT DP DP-VT TAAM
9.71±0.11 10.52±0.14 408.59±0.47 420.34±0.36 636.34±0.90
1.00× 1.08× 42.08× 43.29× 65.53×
102.99 95.06 2.45 2.38 1.57
1.00× 0.92× 0.02× 0.02× 0.02×
Fig. 5: Ablation studies of Agile-WAM. (a) Joint action– future modeling improves performance beyond the gain from increased latent capacity. (b) Jointly modeling both future visual and tactile observations provides the strongest performance under multimodal input. (c) Multi-horizon multimodal prediction better matches the different temporal characteristics of vision and tactile signals. (d) We vary the visual and tactile prediction loss weights to explore their respective influence on policy performance. keeping the enlarged latent for prediction head unchanged. This comparison isolates the effect of world modeling from the additional model capacity. The results show that the improvement mainly comes from joint action–future modeling rather than the larger latent space. By jointly learning actions and future observations, the policy can better capture how its actions affect subsequent states, leading to higher-quality actions that are consistent with the expected future evolution. We further study the contribution of predicting different observation modalities. In contact-rich manipulation, vision and tactile sensing provide complementary feedback, and modeling their future states can impose different forms of alignment on action learning. We therefore compare different combinations of visual and tactile prediction when both modalities are provided as input, as well as visual prediction when only vision is available. As shown in fig. 5(b), visual world modeling improves performance in the vision-only setting. When both vision and tactile observations are used as input, however, removing the prediction of either modality degrades performance, while Agile-WAM achieves the best result by jointly predicting both future vision and tactile observations. These results suggest that world modeling is
most effective when the predicted modalities are consistent with the observation modalities available to the policy. With multimodal input, predicting the future of only one modality provides incomplete supervision of the environment dynamics, whereas jointly modeling both modalities encourages the policy to learn a stronger correspondence among multimodal observations, actions, and future states. D. Ablation of Prediction Horizon We analyze the importance of multi-horizon multimodal prediction by varying the prediction horizons for vision and tactile observations. Visual observations typically evolve smoothly over time, making adjacent-frame prediction highly redundant and potentially too trivial to provide effective supervision to learn visual dynamics. In contrast, tactile signals can change abruptly upon contact, so short-horizon prediction is better suited to capturing fine-grained contact dynamics. We therefore compare our multi-horizon prediction design against two alternatives: next-frame prediction for both vision and tactile, and the same longer prediction horizon for both modalities. As shown in fig. 5(c), using the same prediction horizon for vision and tactile consistently degrades performance. A long prediction horizon for tactile can overlook high-frequency contact changes, whereas nextframe visual prediction provides limited learning signal and fails to capture precise visual dynamics due to the strong redundancy between adjacent frames. Consistent with the observations in section IV-C, ineffective prediction of either modality weakens the alignment between future modeling and multimodal observations, ultimately degrading the quality of the jointly generated actions. V. CONCLUSIONS This paper presented Agile-WAM, an agile tactile World Action Model that couples multimodal world prediction with action generation for contact-rich manipulation. By using the fused observation latent as the source of a visiontactile-to-action flow, Agile-WAM jointly generates action, visual, and tactile representations without relying on a large generative backbone or complex conditioning modules. Its multi-horizon multimodal prediction mechanism provides supervision matched to the temporal characteristics of each modality: longer-horizon visual prediction captures meaningful evolution, whereas next-step tactile prediction preserves rapid contact dynamics. Evaluations on nine simulated and five real-world tasks demonstrate strong performance across diverse contact-rich interactions. At the same time, Agile-WAM maintains low inference latency and a high control frequency. Overall, these results demonstrate that tactile World Action Model can be both effective and computationally efficient, making Agile-WAM well suited for precise, high-frequency robot manipulation. R EFERENCES [1] D. Gao, H. Wang, H. Zhou, N. Ammar, S. Mishra, A. Moradipari, I. Soltani, and J. Zhang, “In-ril: Interleaved reinforcement and imitation learning for policy fine-tuning,” arXiv preprint arXiv:2505.10442, 2025.
[2] Y. Guo, Y. Hu, J. Zhang, Y.-J. Wang, X. Chen, C. Lu, and J. Chen, “Prediction with action: Visual policy learning via joint denoising process,” Advances in Neural Information Processing Systems, vol. 37, pp. 112 386–112 410, 2024. [3] S. Li, Y. Gao, D. Sadigh, and S. Song, “Unified Video Action Model,” in Proceedings of Robotics: Science and Systems, LosAngeles, CA, USA, June 2025. [4] C. Wan, K. Wang, Y. Si, P. Zhang, and M. Li, “Worldagen: Unified state-action prediction with test-time world model training,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 40, no. 22, 2026, pp. 18 584–18 592. [5] S. Ye, Y. Ge, K. Zheng, S. Gao, S. Yu, G. Kurian, S. Indupuru, Y. L. Tan, C. Zhu, J. Xiang et al., “World action models are zeroshot policies,” arXiv preprint arXiv:2602.15922, 2026. [6] H. Xue, J. Ren, W. Chen, G. Zhang, Y. Fang, G. Gu, H. Xu, and C. Lu, “Reactive diffusion policy: Slow-fast visual-tactile policy learning for contact-rich manipulation,” arXiv preprint arXiv:2503.02881, 2025. [7] J. J. Liu, Y. Li, K. Shaw, T. Tao, R. Salakhutdinov, and D. Pathak, “Factr: Force-attending curriculum training for contact-rich policy learning,” arXiv preprint arXiv:2502.17432, 2025. [8] H. Yuan, W. Yi, Z. Zhang, W. Chen, Y. Mo, J. Yin, X. Li, X. Zeng, C. Wen, C. Lu et al., “Vtam: Video-tactile-action models for complex physical interaction beyond vlas,” arXiv preprint arXiv:2603.23481, 2026. [9] S. Wu, L. You, J. Zhu, Y. Liu, H. Kaixiang, C. Yonghang, J. Li, C. Zhang, J. Liu, Z. Hengshuo et al., “Tactile-wam: Touch-aware world action model with tactile asymmetric attention,” arXiv preprint arXiv:2606.26663, 2026. [10] Y. Lou, Y. Ye, Y. Fu, J. Cen, X. Chi, Y. Lyu, P. Jia, S. Han, Z. Lu, and S. Zhang, “Dream-tac: A unified tactile world action model for contact-rich robot manipulation,” arXiv preprint arXiv:2606.08737, 2026. [11] D. Gao, B. Zhao, A. Lee, I. Chuang, H. Zhou, H. Wang, Z. Zhao, J. Zhang, and I. Soltani, “Vita: Vision-to-action flow matching policy,” in International Conference on Learning Representations, vol. 2026, 2026, pp. 100 903–100 933. [12] R. Feng, J. Hu, W. Xia, T. Gao, A. Shen, Y. Sun, B. Fang, and D. Hu, “Anytouch: Learning unified static-dynamic representation across multiple visuo-tactile sensors,” in International Conference on Learning Representations, vol. 2025, 2025, pp. 31 265–31 285. [13] H. Gupta, Y. Mo, S. Jin, and W. Yuan, “Sensor-invariant tactile representation,” in International Conference on Learning Representations, vol. 2025, 2025, pp. 96 515–96 538. [14] R. Feng, Y. Zhou, S. Mei, D. Zhou, P. Wang, S. Cui, B. Fang, G. Yao, and D. Hu, “Anytouch 2: General optical tactile representation learning for dynamic tactile perception,” arXiv preprint arXiv:2602.09617, 2026. [15] Y. Lipman, R. T. Chen, H. Ben-Hamu, M. Nickel, and M. Le, “Flow matching for generative modeling,” arXiv preprint arXiv:2210.02747, 2022. [16] F. Zhang and M. Gienger, “Affordance-based robot manipulation with flow matching,” arXiv preprint arXiv:2409.01083, 2024. [17] E. Chisari, N. Heppert, M. Argus, T. Welschehold, T. Brox, and A. Valada, “Learning robotic manipulation policies from point clouds with conditional flow matching,” arXiv preprint arXiv:2409.07343, 2024. [18] X. Hu, B. Liu, X. Liu, and Q. Liu, “Adaflow: Imitation learning with variance-adaptive flow-based policies,” Advances in Neural Information Processing Systems, vol. 37, pp. 138 836–138 858, 2024. [19] Q. Zhang, Z. Liu, H. Fan, G. Liu, B. Zeng, and S. Liu, “Flowpolicy: Enabling fast and robust 3d flow-based policy via consistency flow matching for robot manipulation,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 39, no. 14, 2025, pp. 14 754– 14 762. [20] K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter et al., “π0 : A visionlanguage-action flow model for general robot control,” arXiv preprint arXiv:2410.24164, 2024. [21] Q. K. Luu, P. Zhou, Z. Xu, Z. Zhang, Q. Qiu, and Y. She, “Manifeel: Benchmarking and understanding visuotactile manipulation policy learning,” arXiv preprint arXiv:2505.18472, 2025.