ConceptioArchivearXiv CS
arXiv CSopen access

Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control Jihoon Hong⋆ , Julian Skifstad⋆ , Qiyue Dai, Alice Chan, Glen Chou Georgia Institute of Technology {jhong392, jskifstad3, qdai41, ichan30, chou}@gatech.edu ⋆ Equal contribution § Code

Å Video

Abstract: World Action Models (WAMs) enable semantically- and physicallyinformed control but are brittle under distribution shift. In this work, we use mechanistic interpretability to study how robustness-relevant perturbations are represented in WAM activation space. Comparing activations across successful and unsuccessful rollouts, we find some WAM architectures exhibit low-dimensional linear separability for robustness-critical features, while others do not. This motivates the use of contrastive activation directions for training-free WAM steering. We also show that local linearity in WAM activation dynamics enables efficient feedback steering via model-based optimal control, yielding World-Action Linear Quadratic Regulator (WA-LQR), a minimally-invasive reduced-order LQR controller. Via mechanistic evaluations, we predict strong steerability in the CosmosPolicy and DiT4DiT models but weak steerability in LingBot-VA, consistent with steering intervention results. On Cosmos-Policy and DiT4DiT, WA-LQR generalizes contrastive directions to new tasks and improves robustness to camera, gripper, and visual-noise perturbations over unsteered and prompt steering baselines. Keywords: mechanistic interpretability, world action models, optimal control

𝐆𝐫𝐢𝐩𝐩𝐞𝐫 𝐏𝐨𝐬𝐢𝐭𝐢𝐨𝐧

𝐂𝐚𝐦𝐞𝐫𝐚 𝐎𝐫𝐢𝐞𝐧𝐭𝐚𝐭𝐢𝐨𝐧

𝐆𝐚𝐮𝐬𝐬𝐢𝐚𝐧 𝐍𝐨𝐢𝐬𝐞

𝐞𝐫𝐮𝐥𝐢𝐚𝐅 𝐬𝐬𝐞𝐜𝐜𝐮𝐒

𝐝𝐞𝐫𝐞𝐞𝐭𝐒 𝐑𝐐𝐋-𝐀𝐖 𝐭𝐨𝐍

arXiv:2607.14943v1 [cs.RO] 16 Jul 2026

€ Website

Figure 1: WA-LQR makes World Action Models more robust to perturbations including gripperposition changes, camera-orientation shifts, and Gaussian sensor noise. In these Cosmos-Policy examples from LIBERO-10, the yellow and green boxes mark the two objects that must be placed in the basket. Without steering, the WAM fails; with WA-LQR, it succeeds.

1

Introduction

Foundation models have advanced robot learning through policies that generalize across objects, tasks, and environments. While VLA models map observations and instructions directly to actions, World Action Models (WAMs) additionally couple action prediction with future-state mod-

eling through video-generative backbones [1–5]. This makes WAMs a promising route toward dynamically-coherent robot policies. However, they remain brittle under out-of-distribution shifts, including camera changes, robot initial-state perturbations, and visual corruption [5]. Because such nuisance factors are unavoidable in deployment and cannot be exhaustively covered with training data, it is essential to improve robustness without extensive new data collection or slow retraining. In this paper, we ask if the internal activations of WAMs reveal why they fail under such perturbations, and if these representations can be used to improve robustness without finetuning. We study perturbations to camera position, initial gripper position, and Gaussian image noise. Using mechanistic interpretability (MI), we compare activations from nominal and perturbed rollouts to test for simple linear geometric activation space structure. By evaluating the degree of linear separability across models, we find this structure is perturbation- and architecture-dependent, suggesting that some WAMs contain steerable representations of robustness-critical features, while others may not. By leveraging the degree of linear separability as a predictor of steerability, we can identify models that are well suited to training-free WAM activation steering, in which contrastive examples are used to construct directions that distinguish nominal from perturbed behavior. We use these directions for open-loop and closed-loop robustness interventions. In particular, we introduce World-Action Linear Quadratic Regulator (WA-LQR), a reduced-order optimal-control method that projects activations into a low-dimensional contrastive subspace and uses local linear dynamics to synthesize a closed-loop LQR steering controller. Unlike open-loop activation addition, WA-LQR adapts online, steering only when activations deviate from the target feature strength while penalizing large perturbations. We evaluate our mechanistic predictions and steering method on Cosmos-Policy [2], DiT4DiT [6], and LingBot-VA [7]. Our low-dimensional linear separability analysis predicts strong steerability for Cosmos-Policy and DiT4DiT and weak steerability for LingBot-VA, which empirical intervention results validate. On Cosmos-Policy, contrastive directions transfer across LIBERO tasks, and WA-LQR improves robustness to camera, gripper, and visual-noise perturbations over non-steered, prompt-steering, and several open-loop steering baselines. These results show that MI can diagnose WAM robustness and guide inference-time interventions. Our contributions are: • We conduct a mechanistic analysis of WAM activations under robustness-relevant perturbations, identifying when nuisance features exhibit low-dimensional linear separability. • We show that steerability is architecture-dependent: Cosmos-Policy [2] and DiT4DiT [6] exhibit clear linear structure across multiple perturbations, whereas LingBot-VA [7] exhibits substantially weaker separability. Moreover, we show that separability is strongly correlated with steerability, making it a useful diagnostic for identifying models that are amenable to activation steering. • We leverage these insights to construct contrastive activation directions for training-free WAM steering. First, we adapt open-loop activation addition techniques from the LLM literature to WAMs. We then show that the local linearity of the diffusion transformer dynamics in a reduced WAM activation space enables the efficient closed-form synthesis of closed-loop steering controllers via WA-LQR. To the best of our knowledge, these are the first open- and closed-loop activation steering methods for WAMs. • We evaluate WA-LQR on robustness benchmarks, showing that MI-based steerability predictions strongly correlate with intervention outcomes and that WA-LQR improves robustness on CosmosPolicy across multiple perturbation types.

2

Related Work

WAMs. Foundation models have enabled general-purpose robotic policies that map observations to actions [1, 8–14]. VLA models are often reactive action predictors and do not explicitly model how the world evolves under robot interventions, limiting long-horizon reasoning and robustness under distribution shift [5, 15]. WAMs address this by coupling action generation with future-state prediction, often adapting video generative backbones to robotic data [2–4, 16]. Existing WAMs include inverse-dynamics-style models, where predicted futures are decoded into actions [17–21], and unified video-action models, where future states and actions are generated jointly [2–4]. We study both families and show that steerable, robot-relevant activation structure is architecture-dependent. 2

Robustness of Action Models. Despite progress in general-purpose robot policies, robustness remains a major deployment obstacle. Benchmarks such as VLATest, COLOSSEUM, and LIBEROPlus expose brittleness to camera viewpoint, robot initial state, object layout, lighting, background texture, sensor noise, and language phrasing [22–26]. Prior studies suggest that robustness gains often require data diversity, wrist-camera observations, RL post-training, or robustness-oriented finetuning rather than arising from inherently stable representations [27–30]. WAMs aim to improve robustness with future-state modeling and video-based priors [2–4, 19, 21, 31], but still fail under shifts such as camera viewpoint and robot initial-state changes [5]. This motivates inference-time methods that improve robustness without exhaustive perturbation data or costly retraining. MI and Activation Steering. MI identifies internal representations that modulate model behavior [32, 33]. In LLMs, many semantic features appear approximately linear in activation space, motivating activation steering: inference-time hidden-state modifications that change behavior without retraining [34–38]. Most methods compute contrastive directions from examples with and without a target concept, then add or transform activations to steer behaviors [39–46]. Because these interventions are often layer-local and open-loop, recent work models activations as dynamical systems and uses feedback for more targeted steering [47–51]. Activation steering has also begun to extend to image and video generators [52–55], where semantics are distributed across text, spatial, temporal, timestep, and layer representations. WAMs add a further challenge: their activations affect not only generated visual content, but also action-relevant predictions that determine robot behavior. Relatively little work applies MI to robotic foundation models, and existing studies focus on VLAs. Prior work identifies steerable VLA activation directions [56], formalizes feature observability and controllability [57], finetunes task-relevant attention heads [58], discovers SAE-based motion primitives [59], measures causal reliance on visual regions [60], and uses activation injection, probes, sparse latents, or conceptor subspaces to analyze and steer VLA behavior [61–63]. In contrast, we study WAMs, whose DiT-style video backbones jointly encode future visual states, action dynamics, and control outputs. We show that WAM steerability is architecture-dependent and develop feedback-based steering that treats WAM inference as a reduced-order dynamical system.

3

Preliminaries and Problem Statement

Linear Quadratic Regulator (LQR) The LQR problem (1) [64] seeks a controller that minimizes a quadratic state-control cost (1) for linear time-varying dynamics (1b) H−1 X  ⊤ min L := zk⊤ Qk zk + u⊤ (1a) k Rk uk + zH QH zH {uk }H−1 k=1

k=1

subject to zk+1 = Ak zk + Bk uk , ∀k = 1, . . . , H − 1. (1b) where Qk ⪰ 0 penalizes state error and Rk ≻ 0 penalizes control effort. The optimal policy has closed form u∗k = −Kk zk , where Kk is obtained efficiently via Riccati recursions. Given a set of nominal setpoints {(z̄k , ūk )}k=1,··· ,H , (1) can be generalized to penalize deviations δzk := zk − z̄k and δuk := uk − ūk from the setpoints. By modifying (1) to H−1 X  ⊤ min δzk⊤ Qk δzk + δu⊤ (2a) k Rk δuk + δzH QH δzH {δuk }H−1 k=1

k=1

subject to δzk+1 = Ak δzk + Bk δuk , k = 1, . . . , H − 1, (2b) ∗ the solution admits a closed-form tracking controller uk := ūk − Kk δzk . When the dynamics zk+1 = f (zk , uk ) are non-linear, similar formulations to (2) are possible by letting Ak = ∇zk f (zk , uk ) and Bk = ∇uk f (zk , uk ), and approximating δzk+1 ≈ Ak δzk + Bk δuk assuming local linearity. WAM Architectures While the two types of WAMs vary in architecture, they both build on Diffusion Transformer (DiT) based video generation models [65]. Starting from a random latent action representation x̂T ∼ N (0, I), these models are used to gradually denoise it to a clean xout by x̂t−1 = STEP(x̂t , t, M (x̂t , t, h)), xout = DEC(x̂1 ), (3) 3

Cosmos-Policy

First

Best

Last

First

Best

Last

Condition positive negative

PC3

PC3

PC2 PC1

Block 0

DiT4DiT

avg loss: 0.0000

PC2 PC1

(a)

Block 4

avg loss: 0.0000

Block 0

avg loss: 0.5501

PC3

PC2 PC1

Block 0

avg loss: 0.2626

(b)

Block 21

avg loss: 0.0854

PC2 PC1

PC3

PC2 PC1

Block 8

avg loss: 0.2626

Block 27

avg loss: 0.3930

PC3

PC2 PC1

Block 15

avg loss: 0.0000

PC3

PC2 PC1

PC3

PC2 PC1

PC3

PC2 PC1

Block 27

avg loss: 0.1645

PC3

PC2 Block 0

avg loss: 0.8877

PC3

PC2 PC1

Block 2

avg loss: 0.0000

PC3

PC1

PC3

PC2 PC1

Block 15

avg loss: 0.2626

Figure 2: On Task 0 of LIBERO-10 [22], we evaluate (a) activations corresponding to noiseperturbed and clean inputs on the first, best intermediate, and final DiT block residual stream for Cosmos-Policy 2B and DiT4DiT [2, 6]. Activations are projected onto the top three principal components, with the reported SVM (gray) and hinge loss. (b) Repeated for camera perturbations. where h is the embedding of the task prompt p, M is the model, STEP is a choice of ODE solver, and DEC decodes latent actions. Typically, M is a sequence of L DiT layers ϕ(l) for l = 0, · · · , L − 1: x0,t := Tok(x̂t ),

xl+1,t = ϕ(l) (xl,t , t, h),

M (x̂t , t, h) := Detok(xL,t ).

(4)

After block L, the WAM output is detokenized and passed to the model scheduler (3). We treat the scheduler transition in (3) as the mechanism that chains together T independent L-horizon withindenoising-timestep controllers. In this work, we will perform activation steering by adding control inputs inside the DiT blocks. Let ul,t denote these perturbations. The steered WAM block is xl+1,t = fl,t (xl,t , ul,t ) := ϕ(l) (xl,t , t, h) + ul,t .

(5)

Problem Statement In this paper, we use MI to identify robustness-relevant features in WAM activation space and use these features for training-free robustness improvement via activation steering. Problem 1: Linear feature discovery in WAM activation space (Sec. 4). Given a WAM, determine whether a target nuisance feature is represented by approximately linear directions in WAM activation space. Concretely, for each layer l and denoising timestep t, we seek low-dimensional projections Pl,t and linear directions vl,t such that the projected activations of ξ + and ξ − are separable. Problem 2: Training-free robustness steering (Sec. 5). Given a WAM, synthesize an open- or closedloop inference-time steering policy that modifies activations and modulates the features discovered in Problem 1.

4

A Mechanistic Study for Interpreting WAMs

A central assumption in the activation steering literature is that model activation spaces exhibit simple, interpretable geometric structure, across domains including LLMs, VLAs, and video generation models [41, 42, 51, 55, 56]. Specifically, most steering algorithms assume linear semantic feature directions, as determined by the emergent linear separability of activations corresponding to semantically contrastive inputs [43, 66]. However, the existence of this structure in WAMs is not guaranteed, despite the steerable foundation model backbone. Indeed, prior work has shown that such representations are fragile to the finetuning process undergone by robotics foundation models [56, 67]. Setup We study the emergent geometry in robustness features for WAMs, specifically CosmosPolicy 2B, DiT4DiT, and LingBot-VA [2, 6, 7]. We consider contrastive datasets related to three key sources of sensitivity in WAM manipulation tasks: perturbation to initial gripper position, initial camera position, and corruption of camera inputs with Gaussian noise. We collect activations corresponding to each dataset through a full model forward pass. That is, for prompts p+ ∈ D+ , 4

p− ∈ D− where, e.g., D+ = {Clean inputs}, D− = {Noised inputs}, we collect DiT actip p vations xt,ℓ+ and xt,ℓ− for all t, ℓ. Typically, DiT activations have a shape of (F, H, W, D), for some token frame, height, width and hidden dimension F < H ≤ W ≪ D, making direct storage intractable for a substantial sample size. Thus, for this section, we consider an average pooling over the token position, resulting in a summarized activation x̄ ∈ RD for all sampled activations. We also perform an average pooling over robot action chunk timesteps; see Sec. 5.1 for details. − Given contrastive sets of activations {x̄+ k }k and {x̄k }k , we define a contrastive direction − dk = x̄+ k − x̄k .

(6)

Inspired by [66], we consider a simple geometric evaluation of the separability of contrastive datasets via principal component analysis (PCA) conducted on the set of contrastive directions, {dk }k . With the objective of enabling linear feature-based steering, we seek to identify linear separability of contrastive datasets. Note that in the high dimensionality of the original mean-pooled system, N samples with N ≪ D are trivially linearly separable, rendering analysis in the full-dimensional space uninformative. Hence, we consider a low-dimensional approximation of the data. In fact, we find in a later section (Sec. 6.1) that the contrastive directions are well summarized by as few as three principal components. Projecting the contrastive activations x̄+ and x̄− onto the top three PCs, we construct a quantitative metric for the linear separability of the pairs of contrastive datasets. In particular, we fit a linear support vector machine (SVM) to the contrastive activations in three dimensions, and measure the classification loss as the average hinge loss per sample [68], 1 X max(0, 1 − yi (wT x̄i + b)), (7) loss(D+ , D− ) = N i∈[N ]

where {x | wT x + b = 0} denotes the SVM hyperplane and yi ∈ {−1, 1}. Intuitively, this measures how well the three-dimensional linear SVM distinguishes the two datasets, where perfect classification yields a loss of 0, and random classification yields a loss of 1.

Applying the same analysis to the action inference activations for LingBot-

Tasks 6&0

First Cond. / Task task 6 positive task 6 negative task 0 positive task 0 negative

Best PC3

PC2 Block 0 avg loss: 0.0000

Tasks 6&1

PC1

PC2 Block 0 avg loss: 0.0000

Tasks 6&4

PC1

PC1

PC2 Block 0 avg loss: 0.0000

PC1

PC2 Block 0 avg loss: 0.0000

PC1

PC2 Block 1 avg loss: 0.0000

PC1

PC2 Block 2 avg loss: 0.0000

PC1

PC2 Block 1 avg loss: 0.0000

PC2 Block 27 avg loss: 0.1012

PC3

PC1

PC3

PC1

PC2 Block 27 avg loss: 0.1068

PC3

PC3

PC3

PC1

PC2 Block 1 avg loss: 0.0000

PC3

PC3

PC3

PC1

Last PC3

PC3

Tasks 6&7

Results We perform this analysis across the perturbations for all 10 tasks of the LIBERO-10 dataset [22], and visualize the results for CosmosPolicy in Fig. 2 (further evaluations are provided in App. E, including for DiT4DiT [6]). Notably, we find that the separability of different features is highly task-dependent, indicating that the latent representation corresponding to the same perturbation is not always shared across different scenes and tasks. A key observation of this study is that certain combinations of tasks share feature representations corresponding to the same input perturbation – see Fig. 3 for an example with Cosmos-Policy. Using these clustered features, we are able to construct steering objectives which generalize across tasks, as discussed in the following sections. This observation enables the activation steering formulation in Sec. 5.

PC2 Block 27 avg loss: 0.1829

PC3

PC1

PC2 Block 27 avg loss: 0.0309

Figure 3: Pairwise separability for Cosmos-Policy under Gaussian noise corruption. Across task pairs, high linear separability and shared feature clusters suggest reusable representations that can be exploited for activation steering. 5

Figure 4: Activations from camera-perturbed and clean LingBot-VA inputs at the first, best intermediate, and final transformer block residual streams. We observe much weaker separability than in the Cosmos-Policy setting. VA yields substantially milder results. Across tasks and perturbations, we observe little to no separability as reported by the SVM loss, as well as qualitatively by inspection (see Fig. 4).

5

Activation Steering for Robustifying WAMs

Motivated by the MI results of Sec. 4, we give an overview of our steering method. We first construct contrastive vectors isolating the desired robustness feature and discuss a simple open-loop method for steering WAMs with the contrastive vectors (Sec. 5.1). We also propose a reduced-order LQR-based steering approach (Sec. 5.2) to enable scalable control-theoretic steering for WAMs. 5.1 Open-Loop Activation Addition from Contrastive Vectors Given paired inputs (ξn+ , ξn− ) that differ primarily in a desired feature, e.g., nominal versus perturbed camera pose or clean versus corrupted observations, we compute contrastive activation directions by subtracting hidden states from the two forward passes (App. A). For layer l, denoising timestep t, and action-chunk timestep τ ∈ [Ha ], this gives (n)

(ξ + )

(ξ − )

n n dl,t,τ := xl,t,τ − xl,t,τ .

(8)

We thus describe the simplest contrastive steering method: activation addition (ActAdd) [41]. ActAdd assumes that the target feature is roughly linear in activation space. A steering vector is formed by averaging contrastive directions (8) over pairs/action-chunk positions al,t :=

N Ha 1 XX (n) d . N Ha n=1 τ =1 l,t,τ

(9)

At inference time, ActAdd steers by adding al,t with strength γ: xl,t,τ ← xl,t,τ + γal,t ,

∀τ ∈ [Ha ].

(10)

The same direction is applied across all action-chunk timesteps, with γ setting steering strength: positive values push toward ξ + , while negative values reverse the effect. ActAdd is training-free and weight-preserving, but open-loop: it ignores the current activation, WAM dynamics, and whether the feature is already at the desired strength, which can cause oversteering and action degradation. 5.2

Steering Latent World Activation Dynamics via WA-LQR

WA-LQR generalizes ActAdd with feedback control, adapting the T2V steering method of [55] to WAMs. Instead of adding a fixed γal,t , it projects contrastive vectors into a latent feature space, measures deviation from a latent setpoint, and computes a minimum-cost LQR intervention. Thus, it preserves ActAdd’s training-free nature while adapting interventions to the realized WAM activation. Dimensionality Reduction As in T2V models [55], full activation-space LQR is infeasible due to the high dimensionality of WAM activations. We instead assume (and validate in Sec. 6) that robustness-relevant factors lie largely in a low-dimensional contrastive subspace. For each layerdenoising pair (l, t), we construct this subspace by pooling contrastive directions (8) over prompt pairs and action-chunk timesteps, then applying streaming randomized singular value decomposition (n) (SVD) [69] to the matrix whose rows are {dl,t,τ : n ∈ [N ], τ ∈ [Ha ]}. This yields a compact 6

⊤ orthonormal basis Vl,t ∈ Rdx ×dz with dz ≪ dx . Define the projection matrix Pl,t := Vl,t ∈ dz ×dx R , which is shared across robot action-chunk timesteps τ ∈ [Ha ]. For each (l, t, τ ), define the latent activation zl,t,τ := Pl,t xl,t,τ ∈ Rdz . Let Al,t and Bl,t denote the raw Jacobians of ∂f the controlled block dynamics in (5) along a nominal trajectory: Al,t := ∂xl,t x̄ ,ū , Bl,t := l,t

l,t

∂fl,t ∂u x̄l,t ,ūl,t . The reduced latent dynamics are

el,t δzl,t,τ + B el,t δul,t,τ , δzl+1,t,τ ≈ A

l = 0, . . . , L − 1,

(11)

el,t := Pl+1,t Al,t P ⊤ ∈ Rdz ×dz and B el,t := Pl+1,t Bl,t ∈ Rdz ×du . Neither Al,t nor Bl,t is where A l,t el,t and B el,t are computed efficiently using Jacobian-vector materialized explicitly; products with A products (JVPs) or vector-Jacobian products (VJPs). Defining Feature Setpoints We define feature setpoints in the latent space for robustness-relevant WAM features. We first compute a latent contrastive direction averaged over action chunk indices ezl,t :=

N Ha + − 1 XX (ξn ) (ξn ) Pl,t (xl,t,τ − xl,t,τ ), N Ha n=1 τ =1

z vl,t :=

ezl,t . ∥ezl,t ∥2

(12)

z z ⊤ For a realized latent activation zl,t,τ , the feature strength is βl,t,τ := (vl,t ) zl,t,τ . We set the desired z,∗ z feature strength as βl,t := λ∥el,t ∥2 , where λ controls steering strength. The feature tracking error z,∗ z ⊤ z for action-chunk timestep τ is αl,t,τ := βl,t − (vl,t ) zl,t,τ , with δzl,t,τ := −αl,t,τ vl,t . Thus, δzl,t,τ is the latent tracking error that, if corrected, would bring the chunk-τ activation to the desired feature setpoint along the contrastive direction.

Reaching Feature Setpoints via WA-LQR Using the latent dynamics in (11), we compute LQR controllers that steer WAM activations toward the feature setpoints. Unlike the T2V setting [55], we do not solve one T ×L-step LQR over both transformer blocks and scheduler transitions. Instead, for each chunk index τ and denoising timestep t, we solve an independent L-step LQR over the transformer blocks: L−1 X

min

{δul,t }L−1 l=0

s.t.

 ⊤ chunk ⊤ δzl,t Ql,t δzl,t + δu⊤ (τ )δul,t + δzL,t QL,t δzL,t l,t Rl,t

l=0

el,t δzl,t + B el,t δul,t , δzl+1,t = A

(13)

l = 0, . . . , L − 1.

The control penalty incorporates an action-decay schedule over robot action chunk indices. Let τ ∈ {0, . . . , Ha − 1} index the action chunks executed over time. For chunk τ , we set r(τ ) := min(Rfinal , Rinit exp(τ /τR )), where Rinit is the initial steering penalty, Rfinal is a large saturation value, and τR controls the decay rate. The LQR control-cost matrix at chunk τ is then chunk Rl,t (τ ) := r(τ )Idu , which is directly used in LQR (13). Since r(τ ) increases with τ and saturates at Rfinal , steering is strongest for early chunks and gradually decays as the robot proceeds. A larger τR slows this growth, causing more chunks to be steered before saturation. In our experiments, Rfinal is chosen large enough that saturated chunks receive negligible steering. Notably, (13) can be efficiently solved in O(Ld3z ) time [70] on the CPU or O(log L · log2 dz ) on the GPU [71]. For each action chunk index τ and denoising timestep t, solving (13) yields gains Kl,t,τ ∈ Rdu ×dz , which are used only within the corresponding denoising pass. At inference time, the controller is z u∗l,t,τ := ūl,t,τ − Kl,t,τ δzl,t,τ = ūl,t,τ + Kl,t,τ αl,t,τ vl,t .

(14)

When ūl,t,τ = 0, the intervention magnitude is proportional to the online feature-tracking error for robot action chunk τ . The chunk-specific policy Kl,t,τ is then applied to the activation at the corresponding action-chunk index. The full WA-LQR procedure chains the T per-timestep controllers through WAM inference. For action chunk index τ and denoising timestep t, we apply the L feedback gains {K0,t,τ , . . . , KL−1,t,τ } inside the DiT blocks, as in (5). At each chunk index τ , the steered block output xL,t is then passed to the scheduler, which produces x̂t−1 and initializes the 7

next denoising pass, yielding T chained L-step LQR controllers. Compared with open-loop ActAdd, WA-LQR adapts to the realized latent activation at each layer and denoising timestep, steering only when the WAM deviates from the desired robustness feature setpoint.

6

Results

In this section, we evaluate whether the mechanistic structure identified above can be used to improve WAM robustness through activation steering. We first assess the validity of the assumptions made by WA-LQR to justify its applicability to WAM steering (Sec. 6.1). Next, we compare openloop activation addition (ActAdd) and closed-loop WA-LQR across Cosmos-Policy, DiT4DiT, and LingBot-VA under OOD perturbations to camera orientation, initial gripper position, and camera noise (Sec. 6.2). Overall, the results show that steering is effective when feature-relevant activations are linearly separable, improving success rates by up to 41%. WA-LQR further improves robustness while helping avoid oversteering in settings where open-loop control is less reliable. See App. D for further experimental details. 6.1

Activation Properties

𝟗𝟏 = 𝐤𝐧𝐚𝐑@𝟗𝟗.𝟎

𝟔 = 𝐤𝐧𝐚𝐑@𝟗.𝟎 𝟑 = 𝐤𝐧𝐚𝐑@𝟓.𝟎

𝐞𝐜𝐧𝐚𝐢𝐫𝐚𝐕 𝐝𝐞𝐧𝐢𝐚𝐥𝐩𝐱𝐄

𝟏.𝟎 We first empirically verify the assumptions that enable LQR-based steering of 𝟎.𝟖 latent activations: 1) preservation of the 𝟎.𝟔 information in the contrastive vectors in the latent subspace and 2) linearity 𝟎.𝟒 of the WAM dynamics within the sub𝟎.𝟐 space. We first study contrastive information preservation. In Fig. 6 we show 𝟎.𝟎 𝟏𝟎 𝟐𝟎 𝟑𝟎 𝟒𝟎 𝟓𝟎 𝟔𝟎 that the variance of contrastive vectors 𝐑𝐚𝐧𝐤 projected to a 64-dimensional latent + − (ξn ) (ξn ) Figure 6: Cumulative variance of latent contrastive vecsubspace, Pl,t (xl,t,τ −xl,t,τ ), is mostly tors for camera orientation perturbation on Cosmos-Policy captured by the top few dimensions. 2B [2] explained by the top-k singular vectors. This is caused by rapid decay of singular values, which suggests that the influence of perturbations such as the change in camera orientation on activations manifests in a significantly lower dimensional subspace compared to raw activations.

To assess dynamics linearity, we observe that WAMs are locally linear (aligned with recent discoveries in LLM [51] and T2V [55] models), which allows them to be effectively steered using LQR.   T Fig. 5(a) shows the cosine similarity and the magnitude ratio between Pl+1,t ϕ(l) xl,t + Pl,t ϵ, t, h and Pl+1,t ϕ(l) (xl,t , t, h) + Ãl,t ϵ for random perturbation ϵ whose norm is proportional to that of ∥xl,t ∥. The cosine similarities and magnitude ratios remain close to 1 throughout inference, demon𝟏.𝟎

𝐲𝐭𝐢𝐫𝐚𝐥𝐢𝐦𝐢𝐒 𝐞𝐧𝐢𝐬𝐨𝐂

𝟏.𝟎

𝟎.𝟗𝟖

𝟎.𝟖

𝟎.𝟗𝟔 𝟎.𝟗𝟒 𝟎.𝟗𝟐

𝐓𝐚𝐬𝐤 𝟎 𝐓𝐚𝐬𝐤 𝟐 𝐓𝐚𝐬𝐤 𝟒 𝐓𝐚𝐬𝐤 𝟓 𝐓𝐚𝐬𝐤 𝟗

𝟏.𝟐𝟎 𝟏.𝟏𝟓 𝟏.𝟏𝟎 𝟏.𝟎𝟓 𝟏.𝟎𝟎 𝟎.𝟗𝟓

𝐥𝐭

𝟔𝟎

( , )=( , )

𝐥𝐭

( , )=(

𝟏𝟒,𝟐)

𝟎.𝟔

𝐨𝐢𝐭𝐚𝐑 𝐞𝐝𝐮𝐭𝐢𝐧𝐠𝐚𝐌

𝟎.𝟒 𝟎.𝟐

𝐥𝐭

( , )=(

𝐥𝐭

(, )

𝐚

𝟐𝟓,𝟒)

𝐑𝐚𝐧𝐝𝐨𝐦

𝟎.𝟎

𝐛

( )

( )

Figure 5: (a) The cosine similarity/magnitude ratio between the linear approximation using Ãl,t and actual latent activation, under change in camera orientation. (b) Overlap between subspaces spanned by 16 top right singular vectors of Ãl,t matrices, obtained from 25 random inputs across 5 tasks. 8

Table 1: LIBERO-10 [22] success rates on Cosmos-Policy 2B [2]. Task i → Task j denotes all Pl,t , Ãl,t , B̃l,t matrices, ezl,t vectors are computed with task i, and used to steer task j. 30 trials/task. Perturbation Tasks No Prompt ActAdd WA-LQR Steering Steering [41] [72] Task 0 → Task 0 33.3%±8.6% 40.0%±8.9% 43.3%±9.0% 60.0%±8.9% Task 0 → Task 2 63.3%±8.8% 60.0%±8.9% 73.3%±8.1% 70.0%±8.4% Camera Task 0 → Task 4 10.0%±5.5% 6.7%±4.6% 16.7%±6.8% 13.3%±6.2% 56.7%±9.0% 70.0%±8.4% 56.7%±9.0% 83.3%±6.8% Orientation Task 0 → Task 5 Task 0 → Task 9 66.7%±8.6% 56.7%±9.0% 56.7%±9.0% 70.0%±8.4% Average 46.0%±4.1% 46.7%±4.1% 49.3%±4.1% 59.3%±4.0% Task 1 → Task 1 46.7%±9.1% 46.7%±9.1% 66.7%±8.6% 60.0%±8.9% Task 1 → Task 2 83.3%±6.8% 83.3%±6.8% 80.0%±7.3% 100.0%±0.0% Initial Task 1 → Task 3 46.7%±9.1% 53.3%±9.1% 56.7%±9.0% 60.0%±8.9% Gripper Task 1 → Task 7 80.0%±7.3% 63.3%±8.8% 60.0%±8.9% 83.3%±6.8% Position Task 1 → Task 9 50.0%±9.1% 56.7%±9.0% 53.3%±9.1% 60.0%±8.9% Average 61.3%±4.0% 60.7%±4.0% 63.3%±3.9% 72.7%±3.6% Task 6 → Task 0 3.3%±3.3% 0.0%±0.0% 43.3%±9.0% 33.3%±8.6% Task 6 → Task 1 53.3%±9.1% 0.0%±0.0% 70.0%±8.4% 73.3%±8.1% Camera Task 6 → Task 4 10.0%±5.5% 0.0%±0.0% 80.0%±7.3% 53.3%±9.1% Gaussian Task 6 → Task 6 36.7%±8.8% 3.3%±3.3% 76.7%±7.7% 73.3%±8.1% Noise Task 6 → Task 7 30.0%±8.4% 0.0%±0.0% 66.7%±8.6% 60.0%±8.9% Average 26.7%±3.6% 0.7%±0.7% 67.3%±3.8% 58.7%±4.0% strating that the latent dynamics are well captured by a first-order approximation. Furthermore, we find that the matrices Ãl,t are highly similar across different inputs. Motivated by the concentration of variance on the top few singular vectors, we visualize in Fig. 5(b) the overlap between the subspaces spanned by the top 16 right singular vectors of Ãl,t over 25 random inputs from 5 different tasks, measured by the mean squared cosine of principal angles (as inspired by a similar metric in [51]). While not shown here, the left singular vectors also demonstrate significant overlap, enabling the reuse of Ãl,t computed from a single input for steering model behavior on other tasks. 6.2

Steering to Improve Robustness

Hinge Loss

We leverage the mechanistic insights from 1.0 r = 0.63 Sec. 4 to inform steering objectives designed WA-LQR 0.8 ActAdd to improve robustness to OOD perturbations 0.6 in Cosmos-Policy 2B [2], DiT4DiT [6], and 0.4 LingBot-VA [7], using WA-LQR (Sec. 5.2) and a variant of simple activation addition adapted 0.2 to WAMs (Sec. 5.1) [41]. Specifically, we seek 0.0 to improve success rate under perturbations to 0.2 initial gripper position, camera orientation, and 20 0 20 40 60 Gaussian noise corruption in the image data, Success rate (Steered Unsteered) applied to LIBERO-10 evaluation tasks [22]. Figure 7: Hinge loss vs. steering performance According to the mechanistic analysis, it is inacross tasks and models, with reported line of best feasible to construct linear feature directions fit (red) and correlation coefficient. that apply for all scenes and tasks (see App. E). Thus, we consider groups of activations which we empirically find have shared representation as suggested by our analysis in Sec. 4. For baselines that do not involve finetuning, we compare to the original model under perturbation and a baseline that decomposes the prompt into subtasks, inspired by [72].

9

Results The grouping and steering results for Table 2: Success rates on DiT4DiT [6]. Cosmos-Policy are summarized in Tab. 1, and Tasks No Steer. ActAdd WA-LQR we provide qualitative examples of unsteered T0→T0 53.3%±9.1% 60.0%±8.9% 66.7%±8.6% Initial T0→T1 69.6% 76.7%±7.7% 70.0%±8.4% and steered trajectory rollouts in Fig. 1. We Gripper T0→T2 70.0%±8.4% 63.3%±8.8% 73.3%±8.1% Position T0→T3 70.0%±8.4% 63.3%±8.8% 76.7%±7.7% provide a full set of snapshots from qualiAvg. 65.7%±4.3% 65.8%±4.3% 71.7%±4.1% tative unsteered and steered rollouts in Fig. T1→T0 16.7%±6.8% 40.0%±8.9% 73.3%±8.1% Camera 8. WA-LQR is the most effective method in T1→T1 20.0%±7.3% 63.3%±8.8% 43.3%±9.0% Gaussian T1→T2 10.0%±5.5% 26.7%±8.1% 30.0%±8.4% improving robustness to camera orientation and Noise Avg. 15.6%±3.8% 43.3%±5.2% 48.9%±5.3% gripper position perturbations, with both steering methods outperforming the unsteered or Table 3: LingBot-VA [7] success rates. prompt-steering baselines. As one exception, Tasks No Steer. ActAdd WA-LQR ActAdd is more effective than WA-LQR on T0→T0 65.0%±10.7% 35.0%±10.7% 70.0%±10.2% T0→T2 55.0%±11.1% 45.0%±11.1% 45.0%±11.1% camera Gaussian noise corruption, which we Camera T0→T4 15.0%±8.0% 10.0%±6.7% 30.0%±10.2% Orientation T0→T5 70.0%±10.2% 75.0%±9.7% 60.0%±11.0% hypothesize results from the strong separability T0→T9 35.0%±10.7% 40.0%±11.0% 50.0%±11.2% of this feature, enabling simple open-loop conAvg. 48.0%±5.0% 41.0%±4.9% 51.0%±5.0% trol directly in activation space to be viable. We T1→T1 65.0%±10.7% 75.0%±9.7% 75.0%±9.7% T1→T2 70.0%±10.3% 70.0%±10.3% 80.0%±8.9% also show in Sec. C that ActAdd is highly senInitial T1→T3 70.0%±10.3% 85.0%±8.0% 75.0%±9.7% Gripper sitive to the strength parameter γ, highlighting T1→T7 80.0%±8.9% 60.0%±11.0% 80.0%±8.9% Position T1→T9 75.0%±9.7% 70.0%±10.3% 65.0%±10.7% the benefit of the closed-loop steering provided Avg. 72.0%±4.5% 72.0%±4.5% 75.0%±4.3% by WA-LQR, which adaptively modulates the T6→T0 35.0%±10.7% 15.0%±8.0% 20.0%±8.9% steering magnitude based on alignment with T6→T1 90.0%±6.7% 60.0%±11.0% 95.0%±4.9% Camera T6→T4 75.0%±9.7% 45.0%±11.1% 75.0%±9.7% Gaussian the desired robustness feature direction. On T6→T6 85.0%±8.0% 55.0%±11.1% 70.0%±10.3% Noise T6→T7 10.0%±6.7% 15.0%±8.0% 20.0%±8.9% DiT4DiT (Tab. 2), ActAdd yields substantial Avg. 59.0%±4.9% 38.0%±4.9% 56.0%±5.0% improvements on Gaussian noise but does not change performance on initial gripper position, while WA-LQR outperforms all baselines and ActAdd on Gaussian noise and initial gripper position. Across evaluations, we find that prompt steering is ineffective, either matching unsteered performance or strongly degrading performance. Task groupings did not emerge on LingBot-VA due to the poor alignment results in Sec. 4, so we instead map the groups discovered for Cosmos-Policy onto LingBot-VA. The results are summarized in Tab. 3 (more results are in App. B). Consistent with Sec. 4, we do not observe major robustness improvements from steering for LingBot-VA. In fact, ActAdd often degrades performance due to over-steering, while WA-LQR avoids this because of its closed-loop modulation of steering magnitude. We summarize the relationship between separability and steering performance in Fig. 7, where we observe a negative correlation between hinge loss and robustness improvements through steering, supporting the use of hinge loss as a predictor for steering success. Overall, these results suggest that hinge loss is an effective predictor of steering success, that both open-loop and closedloop activation steering are effective in improving success rate across models with low separability loss, and that closed-loop steering can be more effective than open-loop perturbations.

7

Discussion, Limitations, and Conclusion

We investigate mechanistic interpretability and activation steering in WAMs. Our analysis reveals clear linear structure in feature-relevant activations under OOD perturbations for Cosmos-Policy [2] and DiT4DiT [6], but not for LingBot-VA [7]. We further find that linear separability loss is a strong predictor of steering performance. Motivated by these observations, we adapt activation addition to the WAM setting and introduce WA-LQR, a novel activation steering framework that exploits locally linear dynamics within a feature-relevant activation subspace to improve OOD robustness without finetuning, achieving success-rate improvements of up to 41%. Overall, our results suggest that mechanistic control of hidden model representations is promising for improving WAM robustness. Limitations. A key limitation of WA-LQR is its limited applicability across tasks and models. To our knowledge, there is currently no interpretable method for predicting when tasks or environments share transferable representations, requiring mechanistic analysis on a per-setting basis. We also observe strong dependence on model architecture. These results highlight the need to better understand how steerability emerges in robotics foundation models and motivate the design of WAMs that are both steerable and able to preserve representations from their base foundation models. 10

Camera pert. Gaussian pert.

WA-LQR “put both the alphabet soup and the tomato sauce in the basket” Unsteered WA-LQR “put the white mug on the plate and put the chocolate pudding to the right of the plate”

Cosmos-Policy

Gripper pert.

Unsteered

Unsteered

WA-LQR

Unsteered WA-LQR “turn on the stove and put the moka pot on it” DiT4DiT

Gripper pert.

Gaussian pert.

“put both the cream cheese box and the butter in the basket”

Unsteered

WA-LQR

Gripper pert.

Unsteered WA-LQR “put both the alphabet soup and the tomato sauce in the basket”

Unsteered WA-LQR “put both the alphabet soup and the tomato sauce in the basket”

LingBot-VA

Gaussian pert.

Camera pert.

“put the black bowl in the bottom drawer of the cabinet and close it”

Unsteered WA-LQR “put both the cream cheese box and the butter in the basket”

Figure 8: Snapshots from rollouts where steering enables task success despite unsteered failure, across perturbations, LIBERO-10 tasks, and WAM architectures. Each rollout shows six equally spaced snapshots, with time increasing from left to right. 11

References [1] M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, et al. Openvla: An open-source vision-language-action model. arXiv preprint arXiv:2406.09246, 2024. [2] M. J. Kim, Y. Gao, T.-Y. Lin, Y.-C. Lin, Y. Ge, G. Lam, P. Liang, S. Song, M.-Y. Liu, C. Finn, et al. Cosmos policy: Fine-tuning video models for visuomotor control and planning. arXiv preprint arXiv:2601.16163, 2026. [3] S. Ye, Y. Ge, K. Zheng, S. Gao, S. Yu, G. Kurian, S. Indupuru, Y. L. Tan, C. Zhu, J. Xiang, et al. World action models are zero-shot policies. arXiv preprint arXiv:2602.15922, 2026. [4] A. Ye, B. Wang, C. Ni, G. Huang, G. Zhao, H. Li, H. Li, J. Li, J. Lv, J. Liu, et al. Gigaworldpolicy: An efficient action-centered world–action model. arXiv preprint arXiv:2603.17240, 2026. [5] Z. Zhang, Z. Li, B. Rahmati, R. H. Yang, Y. Ma, A. Rasouli, S. Pakdamansavoji, Y. Wu, L. Zhang, T. Cao, et al. Do world action models generalize better than vlas? a robustness study. arXiv preprint arXiv:2603.22078, 2026. [6] T. Ma, J. Zheng, Z. Wang, C. Jiang, A. Cui, J. Liang, and S. Yang. Dit4dit: Jointly modeling video dynamics and actions for generalizable robot control. arXiv preprint arXiv:2603.10448, 2026. [7] L. Li, Q. Zhang, Y. Luo, S. Yang, R. Wang, F. Han, M. Yu, Z. Gao, N. Xue, X. Zhu, Y. Shen, and Y. Xu. Causal world modeling for robot control. arXiv preprint arXiv:2601.21998, 2026. [8] A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al. Rt-1: Robotics transformer for real-world control at scale. arXiv preprint arXiv:2212.06817, 2022. [9] B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid, et al. Rt-2: Vision-language-action models transfer web knowledge to robotic control. In Conference on Robot Learning, pages 2165–2183. PMLR, 2023. [10] O. M. Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu, et al. Octo: An open-source generalist robot policy. arXiv preprint arXiv:2405.12213, 2024. [11] A. O’Neill, A. Rehman, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, A. Jain, et al. Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0. In 2024 IEEE International Conference on Robotics and Automation (ICRA), pages 6892–6903. IEEE, 2024. [12] K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, et al. π0 : A vision-language-action flow model for general robot control. arXiv preprint arXiv:2410.24164, 2024. [13] M. J. Kim, C. Finn, and P. Liang. Fine-tuning vision-language-action models: Optimizing speed and success. arXiv preprint arXiv:2502.19645, 2025. [14] K. Pertsch, K. Stachowicz, B. Ichter, D. Driess, S. Nair, Q. Vuong, O. Mees, C. Finn, and S. Levine. Fast: Efficient action tokenization for vision-language-action models. arXiv preprint arXiv:2501.09747, 2025. [15] B. Hou, G. Li, J. Jia, T. An, X. Guo, S. Leng, H. Geng, Y. Ze, T. Harada, P. Torr, et al. World model for robot learning: A comprehensive survey. arXiv preprint arXiv:2605.00080, 2026.

12

[16] S. Wang, J. Shi, Z. Fu, X. He, F. Liu, C. Yang, Y. Zhou, Z. Fei, J. Gong, J. Fu, et al. World action models: The next frontier in embodied ai. arXiv preprint arXiv:2605.12090, 2026. [17] Y. Du, S. Yang, B. Dai, H. Dai, O. Nachum, J. Tenenbaum, D. Schuurmans, and P. Abbeel. Learning universal policies via text-guided video generation. Advances in neural information processing systems, 36:9156–9172, 2023. [18] Y. Wen, J. Lin, Y. Zhu, J. Han, H. Xu, S. Zhao, and X. Liang. Vidman: Exploiting implicit dynamics from video diffusion model for effective robot manipulation. Advances in Neural Information Processing Systems, 37:41051–41075, 2024. [19] Y. Hu, Y. Guo, P. Wang, X. Chen, Y.-J. Wang, J. Zhang, K. Sreenath, C. Lu, and J. Chen. Video prediction policy: A generalist robot policy with predictive visual representations. arXiv preprint arXiv:2412.14803, 2024. [20] Y. Jia, J. Liu, S. Liu, R. Zhou, W. Yu, Y. Yan, X. Chi, Y. Guo, B. Shi, and S. Zhang. Video2act: A dual-system video diffusion policy with robotic spatio-motional modeling. arXiv preprint arXiv:2512.03044, 2025. [21] L. Li, Q. Zhang, Y. Luo, S. Yang, R. Wang, F. Han, M. Yu, Z. Gao, N. Xue, X. Zhu, et al. Causal world modeling for robot control. arXiv preprint arXiv:2601.21998, 2026. [22] B. Liu, Y. Zhu, C. Gao, Y. Feng, Q. Liu, Y. Zhu, and P. Stone. Libero: Benchmarking knowledge transfer for lifelong robot learning. Advances in Neural Information Processing Systems, 36:44776–44791, 2023. [23] X. Li, K. Hsu, J. Gu, K. Pertsch, O. Mees, H. R. Walke, C. Fu, I. Lunawat, I. Sieh, S. Kirmani, et al. Evaluating real-world robot manipulation policies in simulation. arXiv preprint arXiv:2405.05941, 2024. [24] Z. Wang, Z. Zhou, J. Song, Y. Huang, Z. Shu, and L. Ma. Vlatest: Testing and evaluating vision-language-action models for robotic manipulation. Proceedings of the ACM on Software Engineering, 2(FSE):1615–1638, 2025. [25] W. Pumacay, I. Singh, J. Duan, R. Krishna, J. Thomason, and D. Fox. The colosseum: A benchmark for evaluating generalization for robotic manipulation. arXiv preprint arXiv:2402.08191, 2024. [26] S. Fei, S. Wang, J. Shi, Z. Dai, J. Cai, P. Qian, L. Ji, X. He, S. Zhang, Z. Fei, et al. Libero-plus: In-depth robustness analysis of vision-language-action models. arXiv preprint arXiv:2510.13626, 2025. [27] J. Zhou, K. Ye, J. Liu, T. Ma, Z. Wang, R. Qiu, K.-Y. Lin, Z. Zhao, and J. Liang. Exploring the limits of vision-language-action manipulations in cross-task generalization, 2025. URL https://arxiv.org/abs/2505.15660. [28] J. Liu, F. Gao, B. Wei, X. Chen, Q. Liao, Y. Wu, C. Yu, and Y. Wang. What can rl bring to vla generalization? an empirical study. Advances in Neural Information Processing Systems, 38: 97121–97151, 2026. [29] S. Tan, K. Dou, Y. Zhao, and P. Krähenbühl. Interactive post-training for vision-languageaction models. arXiv preprint arXiv:2505.17016, 2025. [30] Y. Guo, J. Zhang, X. Chen, X. Ji, Y.-J. Wang, Y. Hu, and J. Chen. Improving vision-languageaction model with online reinforcement learning. In 2025 IEEE International Conference on Robotics and Automation (ICRA), pages 15665–15672. IEEE, 2025. [31] T. Yuan, Z. Dong, Y. Liu, and H. Zhao. Fast-wam: Do world action models need test-time future imagination? arXiv preprint arXiv:2603.16666, 2026.

13

[32] L. Bereska and S. Gavves. Mechanistic interpretability for AI safety - a review. Transactions on Machine Learning Research, 2024. ISSN 2835-8856. URL https://openreview.net/ forum?id=ePUVetPKu6. Survey Certification, Expert Certification. [33] L. Sharkey, B. Chughtai, J. Batson, J. Lindsey, J. Wu, L. Bushnaq, N. Goldowsky-Dill, S. Heimersheim, A. Ortega, J. Bloom, et al. Open problems in mechanistic interpretability. arXiv preprint arXiv:2501.16496, 2025. [34] N. Elhage, T. Hume, C. Olsson, N. Schiefer, T. Henighan, S. Kravec, Z. Hatfield-Dodds, R. Lasenby, D. Drain, C. Chen, R. Grosse, S. McCandlish, J. Kaplan, D. Amodei, M. Wattenberg, and C. Olah. Toy models of superposition. (arXiv:2209.10652), 2022. doi:10.48550/ arXiv.2209.10652. URL http://arxiv.org/abs/2209.10652. arXiv:2209.10652. [35] K. Park, Y. J. Choe, and V. Veitch. The linear representation hypothesis and the geometry of large language models. arXiv preprint arXiv:2311.03658, 2023. [36] S. Marks and M. Tegmark. The geometry of truth: Emergent linear structure in large language model representations of true/false datasets. In First Conference on Language Modeling, 2024. URL https://openreview.net/forum?id=aajyHYjjsk. [37] A. Zou, L. Phan, S. Chen, J. Campbell, P. Guo, R. Ren, A. Pan, X. Yin, M. Mazeika, A.K. Dombrowski, et al. Representation engineering: A top-down approach to ai transparency. arXiv preprint arXiv:2310.01405, 2023. [38] B. W. Lee, I. Padhi, K. N. Ramamurthy, E. Miehling, P. Dognin, M. Nagireddy, and A. Dhurandhar. Programming refusal with conditional activation steering. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/ forum?id=Oi47wc10sm. [39] S. Dathathri, A. Madotto, J. Lan, J. Hung, E. Frank, P. Molino, J. Yosinski, and R. Liu. Plug and play language models: A simple approach to controlled text generation. In International Conference on Learning Representations, 2020. URL https://openreview.net/forum? id=H1edEyBKDS. [40] K. Li, O. Patel, F. Viégas, H. Pfister, and M. Wattenberg. Inference-time intervention: Eliciting truthful answers from a language model. Advances in Neural Information Processing Systems, 36:41451–41530, 2023. [41] A. M. Turner, L. Thiergart, G. Leech, D. Udell, U. Mini, and M. MacDiarmid. Activation addition: Steering language models without optimization. 2024. [42] N. Rimsky, N. Gabrieli, J. Schulz, M. Tong, E. Hubinger, and A. Turner. Steering llama 2 via contrastive activation addition. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15504–15522, 2024. [43] A. Arditi, O. Obeso, A. Syed, D. Paleka, N. Panickssery, W. Gurnee, and N. Nanda. Refusal in language models is mediated by a single direction. Advances in Neural Information Processing Systems, 37:136037–136083, 2024. [44] P. Rodriguez, A. Blaas, M. Klein, L. Zappella, N. Apostoloff, X. Suau, et al. Controlling language and diffusion models by transporting activations. In International Conference on Learning Representations, volume 2025, pages 89812–89855, 2025. [45] Z. Wu, A. Arora, Z. Wang, A. Geiger, D. Jurafsky, C. D. Manning, and C. Potts. Reft: Representation finetuning for language models. Advances in Neural Information Processing Systems, 37:63908–63962, 2024. [46] H. M. Vu and T. M. Nguyen. Angular steering: Behavior control via rotation in activation space. (arXiv:2510.26243), Oct. 2025. doi:10.48550/arXiv.2510.26243. URL http: //arxiv.org/abs/2510.26243. arXiv:2510.26243. 14

[47] A. Bhargava, C. Witkowski, S.-Z. Looi, and M. Thomson. What’s the magic word? a control theory of llm prompting. arXiv preprint arXiv:2310.04444, 2023. [48] L. Kong, H. Wang, W. Mu, Y. Du, Y. Zhuang, Y. Zhou, Y. Song, R. Zhang, K. Wang, and C. Zhang. Aligning large language models with representation editing: A control perspective. (arXiv:2406.05954), Nov. 2024. doi:10.48550/arXiv.2406.05954. URL http://arxiv.org/ abs/2406.05954. arXiv:2406.05954. [49] E. Cheng and C. A. Alonso. Linearly controlled language generation with performative guarantees. (arXiv:2405.15454), Sept. 2025. doi:10.48550/arXiv.2405.15454. URL http: //arxiv.org/abs/2405.15454. arXiv:2405.15454. [50] D. V. Nguyen, H. M. Vu, N. Y. Pham, L. Zhang, and T. M. Nguyen. Activation steering with a feedback controller. (arXiv:2510.04309), Oct. 2025. doi:10.48550/arXiv.2510.04309. URL http://arxiv.org/abs/2510.04309. arXiv:2510.04309. [51] J. Skifstad, X. A. Yang, and G. Chou. Local linearity of llms enables activation steering via model-based linear optimal control. arXiv preprint arXiv:2604.19018, 2026. [52] P. Rodriguez, M. Klein, E. Gualdoni, V. Maiorca, A. Blaas, L. Zappella, M. Cuturi, and X. Suau. Lineas: End-to-end learning of activation steering with a distributional loss. arXiv preprint arXiv:2503.10679, 2025. [53] S. Facchiano, S. Saravalle, M. Migliarini, E. De Matteis, A. Sampieri, A. Pilzer, E. Rodolà, I. Spinelli, L. Franco, and F. Galasso. Video unlearning via low-rank refusal vector. arXiv preprint arXiv:2506.07891, 2025. [54] Y. Ekin and Y. Gandelsman. The unreasonable effectiveness of text embedding interpolation for continuous image steering. arXiv preprint arXiv:2603.17998, 2026. [55] J. Hong, A. Chan, Q. Dai, J. Skifstad, and G. Chou. Activation steering of video generation models via reduced-order linear optimal control. 2026. [56] B. Häon, K. C. Stocking, I. Chuang, and C. Tomlin. Mechanistic interpretability for steering vision-language-action models. In J. Lim, S. Song, and H.-W. Park, editors, Proceedings of The 9th Conference on Robot Learning, volume 305 of Proceedings of Machine Learning Research, pages 2743–2762. PMLR, 27–30 Sep 2025. URL https://proceedings.mlr. press/v305/haon25a.html. [57] H. Buurmeijer, C. A. Alonso, A. Swann, and M. Pavone. Observing and controlling features in vision-language-action models. arXiv preprint arXiv:2603.05487, 2026. [58] C. Mitra, Y. Luo, R. Saravanan, D. Niu, A. Pai, J. Thomason, T. Darrell, A. Anwar, D. Ramanan, and R. Herzig. Mechanistic finetuning of vision-language-action models via few-shot demonstrations. arXiv preprint arXiv:2511.22697, 2025. [59] A. Swann, L. McGranahan, H. Buurmeijer, M. Kennedy III, and M. Schwager. Sparse autoencoders reveal interpretable and steerable features in vla models. arXiv preprint arXiv:2603.19183, 2026. [60] H. Zhang, M. Xu, A. Dhafer, S. Yue, H. Dong, and Z. D. Hao. Embodied interpretability: Linking causal understanding to generalization in vision-language-action models. arXiv preprint arXiv:2605.00321, 2026. [61] B. Grant, X. Zhao, and P. Wang. Not all features are created equal: A mechanistic study of vision-language-action models. arXiv preprint arXiv:2603.19233, 2026. [62] M. A. Khan, N. Boskov, F. M. Anwar, and M. A. Khan. Controlling vision–language–action policies through sparse latent directions. In Mechanistic Interpretability Workshop at NeurIPS 2025. 15

[63] M. M. Miao, S. Kim, B. Yang, and L. Ungar. Contrastive conceptor activation steering (coast): Unlocking vision-language-action models through hidden states. arXiv preprint arXiv:2605.17144, 2026. [64] J. P. Hespanha. Linear systems theory. Princeton university press, 2018. [65] W. Peebles and S. Xie. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF international conference on computer vision, pages 4195–4205, 2023. [66] S. Marks and M. Tegmark. The geometry of truth: Emergent linear structure in large language model representations of true/false datasets. arXiv preprint arXiv:2310.06824, Aug. 2024. doi:10.48550/arXiv.2310.06824. URL http://arxiv.org/abs/2310.06824. [67] C. Huang, M. M. Zhang, R. Azarcon, G. Chou, and Z. Kira. Maps: Preserving vision-language representations via module-wise proximity scheduling for better vision-language-action generalization. arXiv preprint arXiv:2511.19878, Nov. 2025. doi:10.48550/arXiv.2511.19878. URL http://arxiv.org/abs/2511.19878. [68] V. Kecman. Support Vector Machines – An Introduction, page 1–47. Springer, Berlin, Heidelberg, 2005. ISBN 9783540323846. doi:10.1007/10984697 1. URL https://doi.org/10. 1007/10984697_1. [69] N. Halko, P.-G. Martinsson, and J. A. Tropp. Finding structure with randomness: Probabilistic algorithms for constructing approximate matrix decompositions. SIAM review, 53(2):217–288, 2011. [70] J. B. Rawlings, D. Q. Mayne, M. Diehl, et al. Model predictive control: theory, computation, and design, volume 2. Nob Hill Publishing Madison, WI, 2020. [71] J. Fang and G. Chou. Safe large-scale robust nonlinear mpc in milliseconds via reachabilityconstrained system level synthesis on the gpu. arXiv preprint arXiv:2604.07644, 2026. [72] Z. Wang, Y. Chen, Y. Liu, J. Ye, P. Chen, C. Lu, S. Liu, B. Yu, and J. Jia. Vpvla: Visual prompting as an interface for vision-language-action models. arXiv preprint arXiv:2603.22003, May 2026. doi:10.48550/arXiv.2603.22003. URL http://arxiv.org/ abs/2603.22003. arXiv:2603.22003.

16

Appendices In the following, we provide an overview of our appendices. In App. A, we describe how contrastive vectors are constructed for Gaussian noise, camera perturbations, and gripper-position perturbations, including the definitions of the desirable and undesirable activation sets used for steering. In App. B, we provide supplemental experimental results for LingBot-VA, including additional WA-LQR evaluations on camera orientation, initial gripper-position, and Gaussian-noise perturbations in both the action and video modules. In App. C, we provide a parameter sensitivity analysis on ActAdd. In App. D, we provide additional experimental details for the mechanistic study and for the model evaluations, including activation collection, dimensionality reduction, SVM fitting, and LQR steering implementation details. In App. E, we provide the complete low-dimensional separability results across all LIBERO-10 tasks for all models, including both task-level and pairwise separability plots. In App. F, we provide full model separation plots across layers for Cosmos-Policy and DiT4DiT.

A

Contrastive Vectors

To construct contrastive vectors we define a desirable or target set of inputs D+ , with rollouts emblematic of behavior we seek to induce via steering, and undesirable set of inputs D− , which are perturbed by some nuisance and result in failed rollouts. The definition of this set is consistent between models, for each perturbation. For Gaussian noise, we consider noised and clean camera images, i.e., D+ = {Clean inputs} and D− = {Noised inputs}. For pairs of inputs in ξ + ∈ D+ and ξ − ∈ D− , we perform a model forward pass on each input and collect their corresponding activations x+ , x− . Similarly, the camera perturbation contrastive input sets are defined as D+ = {No perturbation} and D− = {Perturbed camera view}. The corresponding activations x+ , x− are collected from a forward pass of the model. For gripper position perturbation, the definition of D+ and D− differs slightly. Rather than defining the sets as unperturbed and perturbed, respectively, we let D+ be inputs corresponding to successful rollouts and D− be inputs corresponding to unsuccessful inputs, both under gripper perturbation. This is to account for an observed variable sensitivity to different perturbations, i.e., directly including perturbed inputs in the negative dataset would result in many successful rollouts in the negative activations, resulting intuitively in “steering away” from desired behavior. Note, however, that we do not observe a meaningful difference in the mechanistic analysis when making the distinction between successful and unsuccessful rollouts in our dataset configuration (see App. D.1). − Given sets of contrastive vectors {x+ k } and {xk }, we compute the contrastive direction simply as the difference between pairs of positive and negative activations, with different pooling and processing as described in Sec. 5.

B

Supplemental Experimental Results on LingBot-VA

Although we show some results of WA-LQR on LingBot in Table 3, we conducted a more comprehensive evaluation of WA-LQR applied in both the action module and the video module separately. Specifically, relative to Table 3, we further evaluate on Gaussian noise perturbations, on more task transfers, and on steering in the video module (denoted “(Video)” in Table 4). To allow a fair comparison with Cosmos, we evaluate our method on LingBot-VA on the same set of tasks and perturbations as shown in Table 4. We did not observe significant effectiveness of steering, and this is supported by the result of our mechanistic analysis on the extracted activations from LingBotVA’s action and video modules, where it is shown that there is poor linear separability (with high numerical classification losses, especially relative to Cosmos), as presented in Fig 15, 16, and 17.

17

Table 4: LIBERO-10 [22] success rates on LingBot-VA [7]. Task i → Task j denotes all Pl,t , Ãl,t , B̃l,t matrices, ezl,t vectors are computed with task i, and used to steer task j. 20 trials/task. Perturbation

Camera Orientation

Initial Gripper Position

Camera Gaussian Noise

Tasks Task 0 → Task 0 Task 0 → Task 2 Task 0 → Task 4 Task 0 → Task 5 Task 0 → Task 9 Average Task 1 → Task 1 Task 1 → Task 2 Task 1 → Task 3 Task 1 → Task 7 Task 1 → Task 9 Average Task 6 → Task 0 Task 6 → Task 1 Task 6 → Task 4 Task 6 → Task 6 Task 6 → Task 7 Average

No ActAdd Ours Ours Steering (Action) (Video) 65.0%±10.7% 35.0%±10.7% 70.0%±10.3% 65.0%±10.7% 55.0%±11.1% 45.0%±11.1% 45.0%±11.1% 55.0%±11.1% 15.0%±8.0% 10.0%±6.7% 30.0%±10.3% 10.0%±6.7% 70.0%±10.3% 75.0%±9.7% 60.0%±11.0% 65.0%±10.7% 35.0%±10.7% 40.0%±11.0% 50.0%±11.2% 60.0%±11.0% 48.0%±5.0% 41.0%±4.9% 51.0%±5.0% 51.0%±5.0% 65.0%±10.7% 75.0%±9.7% 75.0%±9.7% 70.0%±10.3% 70.0%±10.3% 70.0%±10.3% 80.0%±8.9% 70.0%±10.3% 70.0%±10.3% 85.0%±8.0% 75.0%±9.7% 85.0%±8.0% 80.0%±8.9% 60.0%±11.0% 80.0%±8.9% 80.0%±8.9% 75.0%±9.7% 70.0%±10.3% 65.0%±10.7% 70.0%±10.3% 72.0%±4.5% 72.0%±4.5% 75.0%±4.3% 75.0%±4.3% 35.0%±10.7% 15.0%±8.0% 20.0%±8.9% 65.0%±10.7% 90.0%±6.7% 60.0%±11.0% 95.0%±4.9% 20.0%±8.9% 75.0%±9.7% 45.0%±11.1% 75.0%±9.7% 100.0%±0.0% 85.0%±8.0% 55.0%±11.1% 70.0%±10.3% 75.0%±9.7% 10.0%±6.7% 15.0%±8.0% 20.0%±8.9% 20.0%±8.9% 59.0%±4.9% 38.0%±4.9% 56.0%±5.0% 56.0%±5.0%

18

C

ActAdd Sensitivity Analysis

We perform a sensitive analysis of the effectiveness of ActAdd on improving the performance of Cosmos-Policy 2B [2] under Gaussian noise perturbation in the sensor input. The steering vectors are acquired from Task06 and used to steer the model over 5 tasks. As shown in Tab. 5, ActAdd demonstrates a sharp peak in success rates at a specific hyperparameter (γ = 0.1), and this performance boost quickly vanishes with smaller or larger γ. Table 5: Performance across different values of γ under sensor input perturbed with Gaussian noise. γ 0.0 0.02 0.05 0.075 0.1 0.15 0.2 0.25 0.5

Mean 26.6% 36.2% 50.8% 52.6% 67.4% 32.6% 0.0% 0.0% 0.0%

Task06 37% 37% 57% 70% 77% 13% 0% 0% 0%

Task00 3% 17% 37% 37% 43% 10% 0% 0% 0%

19

Task01 53% 67% 73% 70% 70% 17% 0% 0% 0%

Task04 10% 13% 27% 43% 80% 80% 0% 0% 0%

Task07 30% 47% 60% 43% 67% 43% 0% 0% 0%

D

Experimental Details

D.1

Mechanistic Study

To analyze the geometry of each model’s latent activations, we operate directly in the full activation space, using mean-pooled activations across token positions as described in Sec. 4. For CosmosPolicy, which utilizes a unified single-backbone architecture, we consider the activations across all latent frames. For LingBot-VA and DiT4DiT, which adopt inverse-dynamics style architectures with distinct DiTs for different modalities, we consider the activations from the action-generation DiT. This is consistent with our steering formulation (App. D.2-D.3). To collect contrastive examples for each perturbation type, we conduct a similar procedure to the contrastive vector collection process in App. A. That is, we consider successful vs. unsuccessful rollout as D+ and D− , respectively, under camera or gripper position perturbation. For Gaussian noise, we let D+ = {Clean inputs}, and D− = {Noised inputs}. We also evaluated the gripper and camera perturbation in terms of D+ = {Without perturbation}, and D− = {With perturbation}, but did not observe a meaningful difference. Given D+ and D− , we compute contrastive directions, and conduct Principal Component Analysis on the set of contrastive directions, as described in Sec. 4. To fit the support vector machine (SVM), we initialize the separating {x|ŵT x + b̂ = 0} hyperplane as the zero vector and perform gradient descent with the loss function: X lossSVM (ŵ, b̂) = (1/2)ŵ⊤ ŵ + C · max(0, 1 − yi (ŵT x̄i + b̂)), (15) i∈[N ]

where C is some constant (in our experiments, we set C = 10). Note the second term in Eq. 15 resembles the hinge loss in Eq. 7, evaluated on the intermediate hyperplane parameterized by ŵ, b̂. The SVM is always computed in the three-dimensional subspace defined by the top three principal components of the contrastive directions. D.2

Cosmos-Policy Evaluations on LIBERO-10

Cosmos-Policy is a diffusion-based world-action model with a diffusion transformer backbone, where it jointly denoises a robot action chunk and future state predictions (future proprioception, wrist image, and third-person image). In our method, steering is applied at all latent outputs except for the action chunks. To construct contrastive vectors, activations are collected from all 28 transformer blocks across 5 denoising timesteps for matched pairs of clean and perturbed observations of the same task and scene. For each layer and timestep, a randomized SVD is run on the paired matrix to produce a rank64 basis to reduce dimensionality. The compact subspace is used to find the linearized dynamics that approximates how activations propagate from one layer to the next. The LQR gains are precomputed via backward Riccati recursion, where the control cost grows exponentially with the action chunk index. An LQR hyperparameter search is conducted before settling on the best-performing parameters for all other tasks of the same scene. During inference time, the projected activation error relative to the nominal trajectory is used to inject a correction through the precomputed gains. D.3

LingBot-VA Evaluations on LIBERO-10

LingBot-VA is a mixture-of-transformers (MoT) based world-action model. It first uses a video module to predict future visual frames, which are then passed to a lightweight action module with a smaller hidden dimension to generate robot actions. In our experiments, we apply our method to either the video module or the action module while keeping the other component unchanged. For both modules, we collect activations from the outputs of all 30 transformer blocks. Since the video and action modules use different denoising schedules, we select module-specific denoising timesteps. For the action module, we use timesteps (t ∈ 0, 10, 20, 30, 40). For the video module, we 20

use timesteps (t ∈ 0, 4, 9, 14, 19). These timesteps are chosen to cover the corresponding denoising trajectory of each module. To reduce dimensionality, we partition the 30 transformer layers into three groups: layers 0–9, 10–19, and 20–29. For each layer partition and denoising timestep, we pool activation differences of contrastive pairs from all layers within the corresponding partition and compute a shared rank64 SVD basis. The produced compact subspace for each layer group and timestep pair is used by our LQR injector during evaluation. At evaluation time, the LQR injector follows the same layer partitioning and timestep mapping. For each activation, we project it onto the corresponding low-dimensional SVD subspace, compute the LQR correction in this reduced space, and map the correction back to the full activation space before applying it to the model. D.4

DiT4DiT Evaluations on LIBERO-10

Similar to LingBot-VA, DiT4DiT is a mixture-of-transformers (MoT) based world-action model with distinct video and action modules. In our experiments, we apply our method to the action module, and do not intervene on the video generation module. Otherwise, our evaluation procedure structurally follows Cosmos-Policy on all token positions, intervening across all 4 diffusion timesteps and 16 model layers. The dimensionality reduction procedure also follows the other models, with action generation layers partitioned into three groups: layers 0-5, 6-10, and 11-15. For each layer pool and diffusion timestep, we pool contrastive vectors and analogously construct a 64-dimensional SVD basis, which defines the projection onto the low-dimensional latent space. Jacobian construction, LQR, and control synthesis follow the same procedure as the other two models.

21

E

Complete Low-Dimensional Separability Results

We report the feature separation for three considered robustness features: perturbation to initial gripper position, initial camera position, and corruption of camera inputs with Gaussian noise, for all tasks in the LIBERO-10 dataset [22]. The results are summarized in Figs. 9-14. As reported in Sec. 4, we observe linear separation across tasks and perturbations in Cosmos-Policy (Figs. 9-11), but do not observe such separation in LingBot-VA (Figs. 12-17). For Cosmos-Policy, where we observe meaningful separation, we also present pairwise separability plots (Fig. 18-20), generated by aggregating the activations for pairs of tasks and otherwise following the same procedure as before.

22

First

Best

First

Last

Best

Last

Condition

PC3

PC3

PC2

PC3

PC2

PC1

Task 5

negative

PC3

PC2

PC1

PC3

PC2

PC1

PC3

PC2

PC1

PC2

PC1

PC1

Block 27

Block 0

Block 2

Block 27

avg loss: 0.1645

avg loss: 0.0000

avg loss: 0.0000

avg loss: 0.4458

PC3

PC3

PC2

PC3

PC2

PC1

Task 6

Block 2 avg loss: 0.0000

Task 1

Block 0 avg loss: 0.0000

PC3

PC2

PC1

PC3

PC2

PC1

PC3

PC2

PC1

PC2

PC1

PC1

Block 27

Block 0

Block 2

Block 27

avg loss: 0.1715

avg loss: 0.0000

avg loss: 0.0000

avg loss: 0.1190

PC3

PC3

PC2

PC3

PC2

PC1

Task 7

Block 2 avg loss: 0.0000

Task 2

Block 0 avg loss: 0.0000

PC3

PC2

PC1

PC3

PC2

PC1

PC3

PC2

PC1

PC2

PC1

PC1

Block 27

Block 0

Block 2

Block 27

avg loss: 0.1712

avg loss: 0.0000

avg loss: 0.0000

avg loss: 0.0787

PC3

PC3

PC2

PC3

PC2

PC1

Task 8

Block 1 avg loss: 0.0000

Task 3

Block 0 avg loss: 0.0000

PC3

PC2

PC1

PC3

PC2

PC1

PC3

PC2

PC1

PC2

PC1

PC1

Block 1

Block 27

Block 0

Block 1

Block 27

avg loss: 0.0000

avg loss: 0.0000

avg loss: 0.0000

avg loss: 0.0000

avg loss: 0.0702

Task 4

Block 0 avg loss: 0.0000

PC3

PC3

PC2

PC3

PC2

PC1

Task 9

Task 0

positive

PC3

PC2

PC1

PC3

PC2

PC1

PC3

PC2

PC1

PC2

PC1

PC1

Block 0

Block 1

Block 27

Block 0

Block 1

Block 27

avg loss: 0.0000

avg loss: 0.0000

avg loss: 0.2411

avg loss: 0.0000

avg loss: 0.0000

avg loss: 0.0183

Figure 9: Cosmos-Policy noise corruption separation for all LIBERO-10 tasks

First

Best

First

Last

Best

Last

Condition

PC3

PC3

PC2

PC3

PC2

PC1

Task 5

negative

PC3

PC2

PC1

PC3

PC2

PC1

PC3

PC2

PC1

PC2

PC1

PC1

Block 27

Block 0

Block 12

Block 27

avg loss: 0.6671

avg loss: 0.4680

avg loss: 0.3881

avg loss: 0.6923

PC3

PC3

PC2

PC3

PC2

PC1

Task 6

Block 19 avg loss: 0.0979

Task 1

Block 0 avg loss: 0.8049

PC3

PC2

PC1

PC3

PC2

PC1

PC3

PC2

PC1

PC2

PC1

PC1

Block 27

Block 0

Block 20

Block 27

avg loss: 0.5929

avg loss: 0.6096

avg loss: 0.2609

avg loss: 0.5963

PC3

PC3

PC2

PC3

PC2

PC1

Task 7

Block 17 avg loss: 0.1119

Task 2

Block 0 avg loss: 0.7786

PC3

PC2

PC1

PC3

PC2

PC1

PC3

PC2

PC1

PC2

PC1

PC1

Block 27

Block 0

Block 12

Block 27

avg loss: 0.7934

avg loss: 0.7709

avg loss: 0.1952

avg loss: 0.4725

PC3

PC3

PC2

PC3

PC2

PC1

Task 8

Block 1 avg loss: 0.3670

Task 3

Block 0 avg loss: 0.6296

PC3

PC2

PC1

PC3

PC2

PC1

PC3

PC2

PC1

PC2

PC1

PC1

Block 21

Block 27

Block 0

Block 10

Block 27

avg loss: 0.6263

avg loss: 0.8211

avg loss: 0.6446

avg loss: 0.6777

avg loss: 0.7774

Task 4

Block 0 avg loss: 0.6728

PC3

PC3

PC2 PC1

PC3

PC2 PC1

Task 9

Task 0

positive

PC3

PC2

PC3

PC2

PC1

PC1

PC3

PC2 PC1

PC2 PC1

Block 0

Block 16

Block 27

Block 0

Block 2

Block 27

avg loss: 0.6275

avg loss: 0.1449

avg loss: 0.5281

avg loss: 0.6799

avg loss: 0.4546

avg loss: 0.8781

Figure 10: Cosmos-Policy camera perturbation separation for all LIBERO-10 tasks.

23

First

Best

First

Last

Best

Last

Condition

PC3

PC3

PC2

PC3

PC2

PC1

Task 5

negative

PC3

PC2

PC1

PC3

PC2

PC1

PC3

PC2

PC1

PC2

PC1

PC1

Block 27

Block 0

Block 25

Block 27

avg loss: 0.2864

avg loss: 0.0829

avg loss: 0.0415

avg loss: 0.0616

PC3

PC3

PC2

PC3

PC2

PC1

Task 6

Block 6 avg loss: 0.2885

Task 1

Block 0 avg loss: 0.3878

PC3

PC2

PC1

PC3

PC2

PC1

PC3

PC2

PC1

PC2

PC1

PC1

Block 27

Block 0

Block 26

Block 27

avg loss: 0.1084

avg loss: 0.2048

avg loss: 0.1345

avg loss: 0.2082

PC3

PC3

PC2

PC3

PC2

PC1

Task 7

Block 25 avg loss: 0.0998

Task 2

Block 0 avg loss: 0.1746

PC3

PC2

PC1

PC3

PC2

PC1

PC3

PC2

PC1

PC2

PC1

PC1

Block 27

Block 0

Block 25

Block 27

avg loss: 0.1718

avg loss: 0.3548

avg loss: 0.2738

avg loss: 0.3276

PC3

PC3

PC2

PC3

PC2

PC1

Task 8

Block 26 avg loss: 0.1106

Task 3

Block 0 avg loss: 0.1679

PC3

PC2

PC1

PC3

PC2

PC1

PC3

PC2

PC1

PC2

PC1

PC1

Block 25

Block 27

Block 0

Block 25

Block 27

avg loss: 0.1365

avg loss: 0.1351

avg loss: 0.6899

avg loss: 0.6883

avg loss: 0.6973

Task 4

Block 0 avg loss: 0.1397

PC3

PC3

PC2

PC3

PC2

PC1

Task 9

Task 0

positive

PC3

PC2

PC1

PC3

PC2

PC1

PC3

PC2

PC1

PC2

PC1

PC1

Block 0

Block 3

Block 27

Block 0

Block 21

Block 27

avg loss: 0.3931

avg loss: 0.3329

avg loss: 0.4239

avg loss: 0.1921

avg loss: 0.1429

avg loss: 0.1595

Figure 11: Cosmos-Policy gripper perturbation separation for all LIBERO-10 tasks.

First

Best

First

Last

Best

Last

Condition

PC3

PC3

PC2

PC3

PC2

PC1

Task 5

negative

PC3

PC2

PC1

PC3

PC2

PC1

PC3

PC2

PC1

PC2

PC1

PC1

Block 29

Block 0

Block 18

Block 29

avg loss: 0.8571

avg loss: 0.9715

avg loss: 0.9612

avg loss: 0.9655

PC3

PC3

PC2

PC3

PC2

PC1

Task 6

Block 16 avg loss: 0.8550

Task 1

Block 0 avg loss: 0.8571

PC3

PC2

PC1

PC3

PC2

PC1

PC3

PC2

PC1

PC2

PC1

PC1

Block 29

Block 0

Block 3

Block 29

avg loss: 0.8571

avg loss: 0.9104

avg loss: 0.8889

avg loss: 0.9200

PC3

PC3

PC2

PC3

PC2

PC1

Task 7

Block 19 avg loss: 0.8571

Task 2

Block 0 avg loss: 0.8571

PC3

PC2

PC1

PC3

PC2

PC1

PC3

PC2

PC1

PC2

PC1

PC1

Block 29

Block 0

Block 1

Block 29

avg loss: 0.8462

avg loss: 0.7310

avg loss: 0.6923

avg loss: 0.6923

PC3

PC3

PC2

PC3

PC2

PC1

Task 8

Block 2 avg loss: 0.8462

Task 3

Block 0 avg loss: 0.8770

PC3

PC2

PC1

PC3

PC2

PC1

PC3

PC2

PC1

PC2

PC1

PC1

Block 20

Block 29

Block 0

Block 13

Block 29

avg loss: 0.9189

avg loss: 0.9189

avg loss: 0.9546

avg loss: 0.9372

avg loss: 0.9478

Task 4

Block 0 avg loss: 0.9271

PC3

PC3

PC2 PC1

PC3

PC2 PC1

Task 9

Task 0

positive

PC3

PC2

PC3

PC2

PC1

PC1

PC3

PC2 PC1

PC2 PC1

Block 0

Block 6

Block 29

Block 0

Block 11

Block 29

avg loss: 0.8859

avg loss: 0.8571

avg loss: 0.8571

avg loss: 0.8571

avg loss: 0.8571

avg loss: 0.8571

Figure 12: LingBot-VA camera perturbation separation for all LIBERO-10 tasks (Action Module).

24

First

Best

First

Last

Best

Last

Condition

PC3

PC3

PC2

PC3

PC2

PC1

Task 5

negative

PC3

PC2

PC1

PC3

PC2

PC1

PC3

PC2

PC1

PC2

PC1

PC1

Block 29

Block 0

Block 15

Block 29

avg loss: 0.8777

avg loss: 0.8571

avg loss: 0.8213

avg loss: 0.8571

PC3

PC3

PC2

PC3

PC2

PC1

Task 6

Block 10 avg loss: 0.8530

Task 1

Block 0 avg loss: 0.8597

PC3

PC2

PC1

PC3

PC2

PC1

PC3

PC2

PC1

PC2

PC1

PC1

Block 29

Block 0

Block 16

Block 29

avg loss: 0.8571

avg loss: 0.9046

avg loss: 0.8778

avg loss: 0.9012

PC3

PC3

PC2

PC3

PC2

PC1

Task 7

Block 6 avg loss: 0.8571

Task 2

Block 0 avg loss: 0.8571

PC3

PC2

PC1

PC3

PC2

PC1

PC3

PC2

PC1

PC2

PC1

PC1

Block 29

Block 0

Block 4

Block 29

avg loss: 0.9300

avg loss: 0.6856

avg loss: 0.6400

avg loss: 0.6400

PC3

PC3

PC2

PC3

PC2

PC1

Task 8

Block 1 avg loss: 0.9288

Task 3

Block 0 avg loss: 0.9303

PC3

PC2

PC1

PC3

PC2

PC1

PC3

PC2

PC1

PC2

PC1

PC1

Block 12

Block 29

Block 0

Block 6

Block 29

avg loss: 0.8235

avg loss: 0.8235

avg loss: 0.8767

avg loss: 0.8738

avg loss: 0.9464

Task 4

Block 0 avg loss: 0.8235

PC3

PC3

PC2

PC3

PC2

PC1

Task 9

Task 0

positive

PC3

PC2

PC1

PC3

PC2

PC1

PC3

PC2

PC1

PC2

PC1

PC1

Block 0

Block 15

Block 29

Block 0

Block 1

Block 29

avg loss: 0.9474

avg loss: 0.9474

avg loss: 0.9474

avg loss: 0.7746

avg loss: 0.7753

avg loss: 0.7843

Figure 13: LingBot-VA gripper perturbation separation for all LIBERO-10 tasks (Action Module).

First

Best

First

Last

Best

Last

Condition

PC3

PC3

PC2

PC3

PC2

PC1

Task 5

negative

PC3

PC2

PC1

PC3

PC2

PC1

PC3

PC2

PC1

PC2

PC1

PC1

Block 29

Block 0

Block 25

Block 29

avg loss: 0.9982

avg loss: 0.9470

avg loss: 0.9375

avg loss: 0.9375

PC3

PC3

PC2

PC3

PC2

PC1

Task 6

Block 14 avg loss: 0.9881

Task 1

Block 0 avg loss: 0.9907

PC3

PC2

PC1

PC3

PC2

PC1

PC3

PC2

PC1

PC2

PC1

PC1

Block 29

Block 0

Block 13

Block 29

avg loss: 0.9189

avg loss: 0.9072

avg loss: 0.8889

avg loss: 0.8889

PC3

PC3

PC2

PC3

PC2

PC1

Task 7

Block 27 avg loss: 0.9189

Task 2

Block 0 avg loss: 0.9441

PC3

PC2

PC1

PC3

PC2

PC1

PC3

PC2

PC1

PC2

PC1

PC1

Block 29

Block 0

Block 21

Block 29

avg loss: 0.8571

avg loss: 0.9375

avg loss: 0.9375

avg loss: 0.9375

PC3

PC3

PC2

PC3

PC2

PC1

Task 8

Block 5 avg loss: 0.8571

Task 3

Block 0 avg loss: 0.8571

PC3

PC2

PC1

PC3

PC2

PC1

PC3

PC2

PC1

PC2

PC1

PC1

Block 18

Block 29

Block 0

Block 7

Block 29

avg loss: 0.9874

avg loss: 0.9882

avg loss: 0.9741

avg loss: 0.9697

avg loss: 0.9804

Task 4

Block 0 avg loss: 0.9911

PC3

PC3

PC2 PC1

PC3

PC2 PC1

Task 9

Task 0

positive

PC3

PC2

PC3

PC2

PC1

PC1

PC3

PC2 PC1

PC2 PC1

Block 0

Block 6

Block 29

Block 0

Block 27

Block 29

avg loss: 0.8235

avg loss: 0.8235

avg loss: 0.8235

avg loss: 0.9015

avg loss: 0.8889

avg loss: 0.8889

Figure 14: LingBot-VA noise corruption separation for all LIBERO-10 tasks (Action Module).

25

First

Best

Last

Condition

Task 0

positive negative PC3

PC3

PC2

PC3

PC2

PC1

PC2

PC1

PC1 Block 14

Block 29

avg loss: 0.6073

avg loss: 0.7372

Task 2

Block 0 avg loss: 0.8168

PC3

PC3

PC2

PC3

PC2

PC1

PC2

PC1

PC1 Block 15

Block 29

avg loss: 0.8549

avg loss: 0.6008

avg loss: 0.6602

Task 4

Block 0

PC3

PC3

PC2

PC3

PC2

PC1

PC2

PC1

PC1 Block 21

Block 29

avg loss: 0.8784

avg loss: 0.4518

avg loss: 0.7223

Task 5

Block 0

PC3

PC3

PC2

PC3

PC2

PC1

PC2

PC1

PC1 Block 28

Block 29

avg loss: 0.8876

avg loss: 0.7135

avg loss: 0.7155

Task 9

Block 0

PC3

PC3

PC2 PC1

PC3

PC2 PC1

PC2 PC1

Block 0

Block 7

Block 29

avg loss: 0.8542

avg loss: 0.7163

avg loss: 0.7455

Figure 15: LingBot-VA camera orientation perturbation separation for LIBERO-10 task 0, 2, 4, 5, 9 (Video Module).

26

First

Best

Last

Condition

Task 6

positive negative PC3

PC3

PC2

PC3

PC2

PC1

PC2

PC1

PC1 Block 26

Block 29

avg loss: 0.7879

avg loss: 0.7879

Task 0

Block 0 avg loss: 0.7879

PC3

PC3

PC2

PC3

PC2

PC1

PC2

PC1

PC1 Block 28

Block 29

avg loss: 0.9958

avg loss: 0.9820

avg loss: 0.9859

Task 1

Block 0

PC3

PC3

PC2

PC3

PC2

PC1

PC2

PC1

PC1 Block 5

Block 29

avg loss: 0.8889

avg loss: 0.8889

avg loss: 0.8889

Task 4

Block 0

PC3

PC3

PC2

PC3

PC2

PC1

PC2

PC1

PC1 Block 24

Block 29

avg loss: 0.9773

avg loss: 0.9140

avg loss: 0.9589

Task 7

Block 0

PC3

PC3

PC2 PC1

PC3

PC2 PC1

PC2 PC1

Block 0

Block 3

Block 29

avg loss: 0.8276

avg loss: 0.8276

avg loss: 0.8276

Figure 16: LingBot-VA noise corruption separation for LIBERO-10 task 0, 1, 4, 6, 7 (Video Module).

27

First

Best

Last

Condition

Task 1

positive negative PC3

PC3

PC2

PC3

PC2

PC1

PC2

PC1

PC1 Block 16

Block 29

avg loss: 0.6207

avg loss: 0.6207

Task 2

Block 0 avg loss: 0.6207

PC3

PC3

PC2

PC3

PC2

PC1

PC2

PC1

PC1 Block 28

Block 29

avg loss: 0.9326

avg loss: 0.9134

avg loss: 0.9132

Task 3

Block 0

PC3

PC3

PC2

PC3

PC2

PC1

PC2

PC1

PC1 Block 15

Block 29

avg loss: 0.9249

avg loss: 0.8848

avg loss: 0.8960

Task 7

Block 0

PC3

PC3

PC2

PC3

PC2

PC1

PC2

PC1

PC1 Block 21

Block 29

avg loss: 0.9189

avg loss: 0.8584

avg loss: 0.9032

Task 9

Block 0

PC3

PC3

PC2 PC1

PC3

PC2 PC1

PC2 PC1

Block 0

Block 10

Block 29

avg loss: 0.8815

avg loss: 0.8211

avg loss: 0.8425

Figure 17: LingBot-VA gripper perturbation separation for LIBERO-10 task 1, 2, 3, 7, 9 (Video Module).

28

First

Best

Last

Cond. / Task

Tasks 6&0

task 6 positive task 6 negative task 0 positive

PC3

PC3

PC3

task 0 negative

PC2

PC2

PC1

PC2

PC1

PC1 Block 1

Block 27

avg loss: 0.0000

avg loss: 0.1068

Tasks 6&1

Block 0 avg loss: 0.0000

PC3

PC3

PC2

PC3

PC2

PC1

PC2

PC1

PC1 Block 1

Block 27

avg loss: 0.0000

avg loss: 0.1012

Tasks 6&4

Block 0 avg loss: 0.0000

PC3

PC3

PC2

PC3

PC2

PC1

PC2

PC1

PC1 Block 2

Block 27

avg loss: 0.0000

avg loss: 0.0000

avg loss: 0.1829

Tasks 6&7

Block 0

PC3

PC3

PC2 PC1

PC3

PC2 PC1

PC2 PC1

Block 0

Block 1

Block 27

avg loss: 0.0000

avg loss: 0.0000

avg loss: 0.0309

Figure 18: Pairwise separability for Cosmos-Policy with Gaussian noise corruption.

29

First

Best

Last

Cond. / Task

Tasks 1&2

task 1 positive task 1 negative task 2 positive

PC3

PC3

PC3

task 2 negative

PC2

PC2

PC1

PC2

PC1

PC1 Block 23

Block 27

avg loss: 0.1275

avg loss: 0.2020

Tasks 1&3

Block 0 avg loss: 0.1973

PC3

PC3

PC2

PC3

PC2

PC1

PC2

PC1

PC1 Block 25

Block 27

avg loss: 0.1207

avg loss: 0.1379

Tasks 1&7

Block 0 avg loss: 0.1471

PC3

PC3

PC2

PC3

PC2

PC1

PC2

PC1

PC1 Block 25

Block 27

avg loss: 0.2340

avg loss: 0.1592

avg loss: 0.1827

Tasks 1&9

Block 0

PC3

PC3

PC2 PC1

PC3

PC2 PC1

PC2 PC1

Block 0

Block 25

Block 27

avg loss: 0.1836

avg loss: 0.1498

avg loss: 0.1799

Figure 19: Pairwise separability for Cosmos-Policy with gripper position perturbation.

30

First

Best

Last

Cond. / Task

Tasks 0&2

task 0 positive task 0 negative task 2 positive

PC3

PC3

PC3

task 2 negative

PC2

PC2

PC1

PC2

PC1

PC1 Block 21

Block 27

avg loss: 0.2191

avg loss: 0.3805

Tasks 0&4

Block 0 avg loss: 0.6072

PC3

PC3

PC2

PC3

PC2

PC1

PC2

PC1

PC1 Block 22

Block 27

avg loss: 0.1272

avg loss: 0.3306

Tasks 0&5

Block 0 avg loss: 0.5156

PC3

PC3

PC2

PC3

PC2

PC1

PC2

PC1

PC1 Block 24

Block 27

avg loss: 0.6505

avg loss: 0.1662

avg loss: 0.4288

Tasks 0&9

Block 0

PC3

PC3

PC2 PC1

PC3

PC2 PC1

PC2 PC1

Block 0

Block 21

Block 27

avg loss: 0.6539

avg loss: 0.1858

avg loss: 0.4459

Figure 20: Pairwise separability for Cosmos-Policy with camera position perturbation.

31

F

Full Model Separation

For completeness, we report the full mechanistic separation plots across all layers in the following. Task 0

Block 0

DiT blocks 0 15

Block 1

Block 2

Block 3

Condition positive negative

PC3

PC3

PC2 PC1

PC2 PC1

loss: 0.0000

Block 4

Block 5

Block 6

Block 8

Block 9

Block 12

Block 11

Block 14

loss: 0.0469

Block 15

PC3

PC2 PC1

PC2 PC1

loss: 0.0000

PC3

PC2

PC3

PC2 PC1

Block 13

loss: 0.0000

loss: 0.0000

PC3

loss: 0.0000

PC3

PC1

PC2

Block 10

PC2 PC1

loss: 0.0000

PC3

PC1

loss: 0.0000

PC3

PC2 PC1

Block 7

PC2 PC1

loss: 0.0000

PC3

loss: 0.0000

PC3

PC2 PC1

PC2 PC1

loss: 0.0000

PC3

PC2 loss: 0.0000

PC3

PC2 PC1

loss: 0.0000

PC3

PC1

PC3

PC3

PC2 PC1

loss: 0.2298

loss: 0.1853

Figure 21: Cosmos-Policy Task 0 noise corruption

32

PC2 PC1

loss: 0.0835

Task 0

Block 16

DiT blocks 16 27

Block 17

Block 18

Block 19

Condition positive negative

PC3

PC3

PC2 PC1

PC2 PC1

loss: 0.0261

Block 20

Block 21

Block 22

Block 24

Block 25

loss: 0.0000

PC2

Block 26

loss: 0.0000

Block 27

PC3

PC2 PC1

PC3

PC1

loss: 0.0200

PC3

PC2 PC1

Block 23

PC2 PC1

loss: 0.0000

PC3

loss: 0.0000

PC3

PC2 PC1

PC2 PC1

loss: 0.0039

PC3

PC2 loss: 0.0027

PC3

PC2 PC1

loss: 0.0348

PC3

PC1

PC3

PC3

PC2 PC1

loss: 0.0063

loss: 0.1182

Figure 22: Cosmos-Policy Task 0 noise corruption

33

PC2 PC1

loss: 0.1645

Task 1

Block 0

DiT blocks 0 15

Block 1

Block 2

Block 3

Condition positive negative

PC3

PC3

PC2 PC1

PC2 PC1

loss: 0.0000

Block 4

Block 5

Block 6

Block 8

Block 9

Block 12

Block 11

Block 14

loss: 0.0000

Block 15

PC3

PC2 PC1

PC2 PC1

loss: 0.0000

PC3

PC2

PC3

PC2 PC1

Block 13

loss: 0.0000

loss: 0.0000

PC3

loss: 0.0000

PC3

PC1

PC2

Block 10

PC2 PC1

loss: 0.0000

PC3

PC1

loss: 0.0000

PC3

PC2 PC1

Block 7

PC2 PC1

loss: 0.0000

PC3

loss: 0.0000

PC3

PC2 PC1

PC2 PC1

loss: 0.0000

PC3

PC2 loss: 0.0000

PC3

PC2 PC1

loss: 0.0000

PC3

PC1

PC3

PC3

PC2 PC1

loss: 0.1949

loss: 0.4516

Figure 23: Cosmos-Policy Task 1 noise corruption

34

PC2 PC1

loss: 0.1000

Task 1

Block 16

DiT blocks 16 27

Block 17

Block 18

Block 19

Condition positive negative

PC3

PC3

PC2 PC1

PC2 PC1

loss: 0.0002

Block 20

Block 21

Block 22

Block 24

Block 25

loss: 0.0047

PC2

Block 26

loss: 0.0000

Block 27

PC3

PC2 PC1

PC3

PC1

loss: 0.0000

PC3

PC2 PC1

Block 23

PC2 PC1

loss: 0.0000

PC3

loss: 0.0000

PC3

PC2 PC1

PC2 PC1

loss: 0.0069

PC3

PC2 loss: 0.0000

PC3

PC2 PC1

loss: 0.0000

PC3

PC1

PC3

PC3

PC2 PC1

loss: 0.0492

loss: 0.2189

Figure 24: Cosmos-Policy Task 1 noise corruption

35

PC2 PC1

loss: 0.1715

Task 0

Block 0

DiT blocks 0 15

Block 1

Block 2

Block 3

Condition positive negative

PC3

PC3

PC2 PC1

PC2 PC1

loss: 0.8049

Block 4

Block 5

Block 6

Block 8

Block 9

Block 12

Block 11

Block 14

loss: 0.7869

Block 15

PC3

PC2 PC1

PC2 PC1

loss: 0.7719

PC3

PC2

PC3

PC2 PC1

Block 13

loss: 0.3500

loss: 0.7584

PC3

loss: 0.7763

PC3

PC1

PC2

Block 10

PC2 PC1

loss: 0.7442

PC3

PC1

loss: 0.7564

PC3

PC2 PC1

Block 7

PC2 PC1

loss: 0.7485

PC3

loss: 0.7814

PC3

PC2 PC1

PC2 PC1

loss: 0.7876

PC3

PC2 loss: 0.7832

PC3

PC2 PC1

loss: 0.8044

PC3

PC1

PC3

PC3

PC2 PC1

loss: 0.4208

loss: 0.4855

Figure 25: Cosmos-Policy Task 0 camera perturbation

36

PC2 PC1

loss: 0.3126

Task 0

Block 16

DiT blocks 16 27

Block 17

Block 18

Block 19

Condition positive negative

PC3

PC3

PC2 PC1

PC2 PC1

loss: 0.1219

Block 20

Block 21

Block 22

Block 24

Block 25

loss: 0.1614

PC2

Block 26

loss: 0.1455

Block 27

PC3

PC2 PC1

PC3

PC1

loss: 0.1322

PC3

PC2 PC1

Block 23

PC2 PC1

loss: 0.1166

PC3

loss: 0.0980

PC3

PC2 PC1

PC2 PC1

loss: 0.1020

PC3

PC2 loss: 0.1698

PC3

PC2 PC1

loss: 0.2819

PC3

PC1

PC3

PC3

PC2 PC1

loss: 0.3327

loss: 0.9824

Figure 26: Cosmos-Policy Task 0 camera perturbation

37

PC2 PC1

loss: 0.6671

Task 1

Block 0

DiT blocks 0 15

Block 1

Block 2

Block 3

Condition positive negative

PC3

PC3

PC2 PC1

PC2 PC1

loss: 0.7785

Block 4

Block 5

Block 6

Block 8

Block 9

Block 12

Block 11

Block 14

loss: 0.7620

Block 15

PC3

PC2 PC1

PC2 PC1

loss: 0.7148

PC3

PC2

PC3

PC2 PC1

Block 13

loss: 0.3978

loss: 0.5879

PC3

loss: 0.6694

PC3

PC1

PC2

Block 10

PC2 PC1

loss: 0.6774

PC3

PC1

loss: 0.6497

PC3

PC2 PC1

Block 7

PC2 PC1

loss: 0.6185

PC3

loss: 0.7157

PC3

PC2 PC1

PC2 PC1

loss: 0.7031

PC3

PC2 loss: 0.7269

PC3

PC2 PC1

loss: 0.6822

PC3

PC1

PC3

PC3

PC2 PC1

loss: 0.3679

loss: 0.6062

Figure 27: Cosmos-Policy Task 1 camera perturbation

38

PC2 PC1

loss: 0.6071

Task 1

Block 16

DiT blocks 16 27

Block 17

Block 18

Block 19

Condition positive negative

PC3

PC3

PC2 PC1

PC2 PC1

loss: 0.1164

Block 20

Block 21

Block 22

Block 24

Block 25

loss: 0.9066

PC2

Block 26

loss: 0.3127

Block 27

PC3

PC2 PC1

PC3

PC1

loss: 0.2669

PC3

PC2 PC1

Block 23

PC2 PC1

loss: 0.2586

PC3

loss: 0.1451

PC3

PC2 PC1

PC2 PC1

loss: 0.1945

PC3

PC2 loss: 0.2587

PC3

PC2 PC1

loss: 0.1119

PC3

PC1

PC3

PC3

PC2 PC1

loss: 0.9851

loss: 0.7562

Figure 28: Cosmos-Policy Task 1 camera perturbation

39

PC2 PC1

loss: 0.5930

Best

Task 0

First

Last

PC3

PC3

PC2

PC3

PC2

PC1

PC2

PC1

PC1 Block 4

Block 15

avg loss: 0.8877

avg loss: 0.0000

avg loss: 0.0000

Task 1

Block 0

PC3

PC3

PC2

PC3

PC2

PC1

PC2

PC1

PC1 Block 5

Block 15

avg loss: 0.7908

avg loss: 0.0000

avg loss: 0.4720

Task 2

Block 0

PC3

PC3

PC2

PC3

PC2

PC1

PC2

PC1

PC1

Block 0

Block 4

Block 15

avg loss: 0.7742

avg loss: 0.0000

avg loss: 0.0000

Task 3

Figure 29: Feature separability: DiT4DiT Gaussian noise corruption for tasks 0-2. PC3

PC3

PC2

PC2

PC1

PC2

PC1

PC1 Block 8

Block 15

avg loss: 0.0000

avg loss: 0.3054

Task 4

Block 0 avg loss: 0.8395

PC3

PC3

PC2

PC3

PC2

PC1

k 5

PC3

PC2

PC1

PC1

Block 0

Block 14

Block 15

avg loss: 0.8239

avg loss: 0.0240

avg loss: 0.2146

PC3

40

PC3

PC3

Best

Task 0

First

Last

PC3

PC3

PC2

PC3

PC2

PC1

PC2

PC1

PC1 Block 5

Block 15

avg loss: 0.7222

avg loss: 0.5830

avg loss: 0.6357

Task 1

Block 0

PC3

PC3

PC2

PC3

PC2

PC1

PC2

PC1

PC1 Block 3

Block 15

avg loss: 0.3751

avg loss: 0.4193

Task 2

Block 0 avg loss: 0.8488

PC3

PC3

PC2

PC3

PC2

PC1

PC2

PC1

PC1 Block 8

Block 15

avg loss: 0.6472

avg loss: 0.6676

Task 3

Block 0 avg loss: 0.7213

PC3

PC3

PC2

PC3

PC2

PC1

PC2

PC1

PC1

Block 0

Block 9

Block 15

avg loss: 0.5951

avg loss: 0.5951

avg loss: 0.5951

Task 4

Figure 30: Feature separability: DiT4DiT gripper perturbation for tasks 0-3. PC3

PC3

PC2

PC3

PC2

PC1

PC2

PC1

PC1 Block 14

Block 15

avg loss: 0.5197

avg loss: 0.3090

avg loss: 0.3284

Task 5

Block 0

PC3

PC3

PC3

41 PC2 PC1

PC2 PC1

Block 0

PC2 PC1

Block 13

Block 15

Task 0

Block 0

DiT blocks 0 15

Block 1

Block 2

Block 3

Condition positive negative

PC3

PC3

PC2 PC1

PC2 PC1

loss: 0.8877

Block 4

Block 5

Block 6

Block 8

Block 9

Block 12

Block 11

Block 14

loss: 0.0124

Block 15

PC3

PC2 PC1

PC2 PC1

loss: 0.0203

PC3

PC2

PC3

PC2 PC1

Block 13

loss: 0.0000

loss: 0.0000

PC3

loss: 0.0133

PC3

PC1

PC2

Block 10

PC2 PC1

loss: 0.0006

PC3

PC1

loss: 0.0000

PC3

PC2 PC1

Block 7

PC2 PC1

loss: 0.0000

PC3

loss: 0.7490

PC3

PC2 PC1

PC2 PC1

loss: 0.7761

PC3

PC2 loss: 0.0000

PC3

PC2 PC1

loss: 0.8995

PC3

PC1

PC3

PC3

PC2 PC1

loss: 0.0000

loss: 0.0000

Figure 31: DiT4DiT Task 0 noise corruption

42

PC2 PC1

loss: 0.0000

Task 1

Block 0

DiT blocks 0 15

Block 1

Block 2

Block 3

Condition positive negative

PC3

PC3

PC2 PC1

PC2 PC1

loss: 0.7908

Block 4

Block 5

PC2

Block 6

Block 8

Block 9

Block 12

Block 11

PC2 PC1

loss: 0.3965

Block 14

loss: 0.4297

Block 15

PC3

PC3

PC2 PC1

PC3

PC2 PC1

loss: 0.2992

PC3

PC2

loss: 0.1486

PC3

Block 13

loss: 0.3947

PC2 PC1

Block 10

PC2 PC1

loss: 0.1936

PC3

loss: 0.0563

PC3

PC2

PC1

Block 7

PC2 PC1

loss: 0.0000

PC3

PC1

loss: 0.6874

PC3

PC2 PC1

PC2 PC1

loss: 0.7052

PC3

loss: 0.0033

PC3

PC2 PC1

loss: 0.7985

PC3

PC1

PC3

PC3

PC2 PC1

loss: 0.4285

PC2 PC1

loss: 0.3617

loss: 0.4720

Figure 32: DiT4DiT Task 1 noise corruption

First

Best

Last

Cond. / Task

Tasks 1&2

task 1 positive task 1 negative task 2 positive

PC3

PC3

PC3

task 2 negative

PC2

PC2

PC1

PC2

PC1

PC1

Block 0

Block 5

Block 15

avg loss: 0.8109

avg loss: 0.0000

avg loss: 0.0113

Tasks 1&0

Figure 33: Pairwise Gaussian noise corruption feature separation on DiT4DiT. PC3

PC3

PC3

43 PC2 PC1

PC2 PC1

PC2 PC1

Block 0

Block 4

Block 15

avg loss: 0.8542

avg loss: 0.0000

avg loss: 0.0000

First

Best

Last

Cond. / Task

Tasks 1&0

task 1 positive task 1 negative task 0 positive

PC3

PC3

PC3

task 0 negative

PC2

PC2

PC1

PC2

PC1

PC1 Block 13

Block 15

avg loss: 0.7503

avg loss: 0.3592

avg loss: 0.3732

Tasks 1&2

Block 0

PC3

PC3

PC2

PC3

PC2

PC1

PC2

PC1

PC1 Block 4

Block 15

avg loss: 0.3167

avg loss: 0.4142

Tasks 1&3

Block 0 avg loss: 0.7431

PC3

PC3

PC2 PC1

PC3

PC2 PC1

PC2 PC1

Block 0

Block 5

Block 15

avg loss: 0.8160

avg loss: 0.4485

avg loss: 0.4959

Figure 34: Pairwise gripper perturbation feature separation on DiT4DiT.

44

Record · ID 373394 · SHA-256 e11eb6aacf2a81fe
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.