ConceptioArchivearXiv CS
arXiv CSopen access

DLAM: Distributional Latent Actions with Temporal Constraints

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
artificialintelligenceknowledgerepresentationreasoning
artificial intelligence, reasoning, knowledge representation

DLAM: Distributional Latent Actions with Temporal Constraints Zuojin Tang1,2 Feifan Luo1 Haoyun Liu3,5,2 Botai Yuan4,2 Dekang Qi2 Ronghan Chen2 Yandan Yang2 Tong Lin6,2 Xinyuan Chang2 Mu Xu2 Bin Liu7 De Ma1 Zhiheng Ma5 1

Zhejiang University 2 Amap, Alibaba Group 3 Nanjing University 4 Shanghai Jiaotong University 5 Shenzhen University of Advanced Technology 6 Xi’an Jiaotong University 7 Embodied Intelligence General Platform Laboratory, Chery Auto

arXiv:2607.27138v1 [cs.RO] 29 Jul 2026

Abstract Vision-language-action (VLA) models remain constrained by scarce action-labeled robot data, whereas action-free videos offer abundant observations of physical change. Latent action models can extract such priors, but reconstruction-trained codes may predict future observations without the structure required for joint generation with robot actions. Existing structured methods add temporal constraints but retain deterministic transition points, so residual errors in locally inferred transitions may propagate and compound under recursive composition. We introduce DLAM, a distributional latent-action model that represents each transition as a diagonal Gaussian. Reconstruction conditioned on the reference frame grounds the mean in observed visual change, while normalized composition and reversal over equal-gap triplets constrain both the mean and dimension-wise variance. Variance composition uses a lightweight shared-correlation coefficient to account for dependence between adjacent transitions that share an intermediate frame, whereas reversal negates the mean and preserves the variance. For downstream policy learning, we freeze the encoder and train a flow-matching policy to jointly generate mean transition sequences and robot actions. On held-out transitions, DLAM learns more temporally consistent latent dynamics than existing latent-action baselines and achieves stronger direct and cumulative reconstruction on held-out videos. Under the same controlled π0 transfer protocol, it also improves policy performance on MetaWorld MT50, LIBERO, and real-world manipulation tasks. Controlled ablations show that normalized mean constraints account for most of the reconstruction gain, while learned variance and correlation-aware composition provide complementary improvements in downstream control.

Introduction Vision-language-action (VLA) models provide a unified framework for visual perception, language understanding, and action generation in robot policies. Recent systems learn diverse manipulation skills from large robot datasets (Zitkovich et al. 2023; Kim et al. 2024; Black et al. 2024; O’Neill et al. 2024), but their scalability remains limited by the cost of action-labeled demonstrations. Action-free videos, by contrast, provide abundant observations of physical change and offer a complementary source of transition priors for downstream control (Chen et al. 2026; Li et al. 2026a; Shen et al. 2026).

LAM

ALAM

DLAM (Ours)

ordinary latent action space (no algebraic structure)

deterministic structured latent action space

distributional structured latent action space

dispersion-aware

covariance-aware distribution-space constraints

𝒃 𝒛𝒂𝒃

𝒂 No Compositionality or Reversibility 𝑧!" ≠ 𝑧!# + 𝑧#" | 𝑧!# ≠ −𝑧#!

no rule for composition no rule for reversal

?

?

𝒛𝒃𝒄

𝒛𝒃𝒂 𝒛𝒄𝒂

𝒒𝒄𝒃 𝒒𝒃𝒂

𝒄

Point Composition & Reversal

𝒂

𝒄

𝒒𝒄𝒂

Distributional Composition

𝑧!" = 𝑧!# + 𝑧#" | 𝑧!# = −𝑧#!

same mean, different posterior dispersion

𝒃

𝒒𝒃𝒂

𝜇!" =

(𝜎!" )$=

𝜇!# + 𝜇#"

2 (𝜎!# )$ + (𝜎#" )$+2𝜌𝜎!# ⊙ 𝜎#" 2

Distributional Reversal 𝜇!# = −𝜇#"

𝜎!# = 𝜎#!

Figure 1: From point-valued to distribution-valued transitions. Reconstruction-only LAMs do not explicitly organize temporal relations. ALAM adds composition and reversal constraints to deterministic transition points. DLAM applies these relations to Gaussian posteriors, constraining both the transition center and its dimension-wise dispersion.

Latent action models (LAMs) infer transitions without action labels, typically using reconstruction to ground the latent in observed change (Bruce et al. 2024; Ye et al. 2024; Bu et al. 2025). However, reconstruction alone does not guarantee a representation suitable for policy generation. The latent may encode camera motion, background dynamics, or appearance changes that improve frame prediction but are only weakly related to control (Zhang et al. 2025a). Such a transition can be predictive without being well suited to joint generation with robot actions. Structured LAMs impose composition, reversal, inverse, or cycle constraints (Wei et al. 2026; Tang et al. 2026a; Li et al. 2026b), but still represent each transition as a deterministic point. Each relation thus acts on a single estimate, whose residual errors may compound under recursive long-horizon composition. We instead represent each transition as a diagonal Gaussian, allowing temporal constraints to supervise both its mean and dimension-wise variance. Reconstruction grounds the mean in observed change, while the variance provides an auxiliary signal to the shared encoder. We propose DLAM, a distributional latent-action model

for action-free video. Given a frame pair, an encoder predicts the transition mean and dimension-wise variance, while a source-conditioned decoder reconstructs the target using only the mean. For an equally spaced triplet, normalized composition matches the directly encoded transition with the composition of two adjacent transitions in both mean and variance. As shown in Figure 1, a lightweight sharedcorrelation coefficient accounts for their dependence when composing variances. Reversal negates the mean while preserving the variance. These constraints are local and do not assume an exact global group structure. To transfer DLAM to downstream policies, we freeze the encoder and use its mean transition sequences as auxiliary generative targets alongside robot actions. A flow-matching policy jointly generates both trajectories, requiring neither the predicted variance, a latent-to-action decoder, nor a new VLA backbone. We evaluate DLAM at the representation, reconstruction, and policy levels. Compared with existing latent-action baselines, DLAM achieves consistently lower composition and reversal residuals on held-out transitions, less sensitive to increasing temporal span under the common normalized probe. Normalized temporal probes further show that these residuals are less sensitive to increasing recursive span. DLAM also improves direct and cumulative reconstruction by 3.45 dB and 1.17 dB, respectively, with consistent gains across complementary perceptual measures. Under the controlled π0 transfer setting, the learned means achieve 87.6% average success on MetaWorld MT50, 99.0% on LIBERO, and 73.8% across four real-world tasks. A controlled ablation further shows higher downstream success when introducing normalized mean constraints, learned variance, and correlation-aware composition. Our contributions are: • We extend structured latent actions from deterministic points to diagonal-Gaussian transitions, allowing temporal constraints to supervise both the mean and variance. • We formulate normalized composition and reversal for Gaussian transitions, including correlation-aware composition of dimension-wise variances. • We improve direct and cumulative reconstruction and transfer only the learned posterior means to joint flowmatching VLA learning. Controlled ablations further isolate the complementary contributions of normalized mean constraints, learned variance, and correlation-aware composition to downstream control.

Related Work Vision-Language-Action and World Action Models VLA models build on large-scale vision-language representations for robot control. Representative systems include autoregressive models such as RT-2 and OpenVLA, together with related large-scale VLA approaches (Zitkovich et al. 2023; Kim et al. 2024; Tang et al. 2025; Wang et al. 2026; Yuan et al. 2026a). Diffusion Policy and π0 instead generate continuous action trajectories through diffusion or flow matching (Chi et al. 2025; Black et al. 2024). Cross-embodiment datasets further broaden the tasks and robot morphologies available for policy learning (O’Neill

et al. 2024). Several methods also use future visual information to guide action generation. Prediction with Action jointly denoises future observations and actions, while videoprediction policies condition control on predicted visual features (Guo et al. 2024; Hu et al. 2024). World action models more directly connect representations of visual dynamics with executable actions (Tang et al. 2026b; Chen et al. 2026; Li et al. 2026a; Yuan et al. 2026b). Our work follows this direction but focuses on the transition representation learned from action-free videos while keeping the downstream VLA backbone unchanged.

Latent Action Models from Videos Latent action models infer action-like variables directly from observation sequences. Genie learns discrete latent actions for interactive world generation, while LAPA transfers videoderived latent actions to robot policies (Bruce et al. 2024; Ye et al. 2024). UniVLA, AdaWorld, and CLAM further study task-aware and continuous latent actions for crossembodiment robot learning (Bu et al. 2025; Gao et al. 2025; Liang et al. 2025). Reconstruction-based latents may, however, encode visual changes that help prediction but are weakly related to control (Zhang et al. 2025a). AC-LAM, ALAM, and RotVLA address this issue by introducing explicit temporal structure (Wei et al. 2026; Tang et al. 2026a; Li et al. 2026b). Despite using different operators, these methods represent each transition deterministically. Stochastic video-prediction methods also use latent variables, primarily to model multiple plausible futures (Babaeizadeh et al. 2018; Denton and Fergus 2018). In contrast, DLAM uses a diagonal-Gaussian representation for the transition between two observed frames and imposes temporal constraints on both its mean and variance. The variance provides an additional training signal rather than calibrated uncertainty, and downstream transfer uses only the learned mean.

Method Overview DLAM extends deterministic latent-action learning with diagonal-Gaussian transition tokens. Source-conditioned reconstruction grounds each transition mean in the observed visual change, while equal-gap triplets apply composition and reversal constraints to both the mean and variance, providing an additional training signal beyond deterministic mean constraints. For downstream control, the frozen encoder provides only mean targets, which are jointly generated with robot actions (Figure 2).

Distributional Latent-Action Pretraining Let Ot denote the video frame at time t. We sample equally spaced triplets (Oa , Ob , Oc ) with a < b < c and b − a = c − b = k, where k is the frame gap. The forward set Tfwd = {(a, b), (b, c), (a, c)} contains two adjacent transitions and their direct counterpart, while the backward pair (b, a) provides reversal supervision. No action, language, or proprioceptive labels are required. j K For an ordered pair (Oi , Oj ), let Zij = {Zi,κ }κ=1 denote the K transition tokens inferred from source frame Oi and

Current Time: T Third-view

Future Sequence:

T

T+1

Wrist-view

T+2

T+H-1

Attention Mask

T+H State

···

Latent action token

“Gather the fruit on the table into the basket.”

···

Latent action

State: T

Gemma

···

State

GT Action

+ Noise

Latent Action & Action Expert

Shared Attention

xS

Joint Flow Matching

Third-view feature & latent action Wrist-view feature & latent action

···

Action token

Tokenizer

T · · ·T+H-1 ···

Latent action token

SigLIP

Action token

DLAM (Our Pretrained)

Trainable

Frozen

···

Denoise

S: Add/Denoise Steps H: Action Horizon

Figure 2: Downstream transfer of DLAM. The frozen encoder extracts distributional transitions from consecutive demonstration frames. Their posterior means are projected and interleaved with robot-action tokens. Conditioned on current visual observations and language, the policy jointly generates both streams with flow matching. Future frames are used only to construct training targets, and only robot actions are executed. target frame Oj . The conditional posterior factorizes across QK j token slots: qij = qϕ (Zij | Oi , Oj ) = κ=1 qi,κ . Each factor is a diagonal Gaussian:   j j j 2 qi,κ = N µi,κ , diag (σi,κ ) . (1) j Here κ indexes the token slot, d is the token width, and µi,κ ∈ j d d R and σi,κ ∈ R+ are its mean and element-wise standard j 2 deviation. Thus, diag((σi,κ ) ) is the diagonal covariance, and N denotes a Gaussian distribution. The relational encoder Eϕ , parameterized by ϕ, processes the two frames with K learnable queries and predicts [µij , ℓeij ] = Eϕ (Oi , Oj ), where both outputs have shape K × d. The raw log-variance ℓeij is converted to the standard deviation by   ℓ j = clip(ℓej , ℓmin , ℓmax ), σ j = exp 1 ℓ j . (2) i

i

i

2 i

The clipping and exponential are applied element-wise, with ℓmin , ℓmax as fixed bounds. Because both endpoints are observed and the reconstruction path uses only the mean, the predicted standard deviation is used to define the diagonal Gaussian and provide an additional training signal through the prior and temporal constraints. It is not interpreted as calibrated uncertainty about an unseen future. For each (i, j) ∈ Tfwd , a source-conditioned decoder Dω reconstructs the target frame from the source frame and b i,j = Dω (Oi , µ j ). The reconstacked transition means: O i struction loss is X 2 1 b i,j − Oj , Lrec = O (3) P |Tfwd | 2 (i,j)∈Tfwd

where P is the number of scalar pixel and channel values in one frame, and the norm is applied to its flattened representation. Conditioning on Oi provides the scene content and

encourages µij to describe the visual change toward Oj . No posterior sample is used in this reconstruction path. We regularize each forward posterior toward the factorized QK standard-Gaussian prior p = κ=1 N (0, Id ), where 0 and Id are the d-dimensional zero vector and identity matrix: X 1 DKL (qij ∥p). (4) Lprior = |Tfwd | (i,j)∈Tfwd

Here DKL denotes the Kullback–Leibler divergence. Because both distributions factorize, the joint KL decomposes into the sum of K token-wise Gaussian KL terms. We apply the free-nats floor before the final loss reduction.

Temporal Constraints on Mean and Variance An equal-gap triplet provides adjacent posteriors qab and qbc together with the directly encoded posterior qac . Composition and reversal are applied independently to corresponding token slots. For clarity, we fix a slot κ and omit its index from the mean and variance equations below. b c Let Zab ∼ qa,κ and Zbc ∼ qb,κ . We model the pair (Zab , Zbc ) as jointly Gaussian with these marginals and define their normalized pairwise composition as Za⇝c =

Zab + Zbc √ . 2

(5)

Because the adjacent transitions share intermediate frame Ob , we allow a lightweight coupling between them. A learned scalar r ∈ R defines ρ = ρmax tanh(r), where the fixed bound 0 < ρmax < 1 ensures |ρ| < 1. We share ρ across examples, token slots, and latent dimensions, and assume the diagonal cross-covariance  Cov(Zab , Zbc ) = diag ρ σab ⊙ σbc , (6)

Figure 3: Deterministic and diagonal-Gaussian transitions. A PCA projection of held-out transitions encoded by trained ALAM and DLAM checkpoints. ALAM returns one point per transition, whereas DLAM predicts a mean together with an element-wise standard deviation. This visualization shows the difference in representation; it is not an uncertainty-calibration or cross-model geometric comparison. where Cov denotes cross-covariance and ⊙ denotes elementwise multiplication. The resulting composed mean and variance are µab + µbc √ , (7) 2 (σ b )2 + (σbc )2 (σ ac )2 = a + ρ σab ⊙ σbc . (8) 2 The overbar distinguishes the composed mean and variance from the directly encoded µac and (σac )2 . It defines a pairwise normalized relation rather than an associative composition law. Restoring the token index, these quantities define  c c c q a,κ = N µa,κ , diag (σ a,κ )2 . (9) µac =

c Thus, Za⇝c,κ ∼ q a,κ , and the full composed posterior facQK c c . torizes as q a = κ=1 q a,κ Setting ρ = 0 recovers independent variance propagation. √ In this case, the 1/ 2 factor preserves unit variance when the two inputs are independent standard Gaussians. We apply this operator once to the equal-gap relation k + k → 2k; it defines a pairwise normalized relation rather than an associative composition law. Reversal changes the direction of a transition by negating its mean while preserving its variance:    R N µ, diag (σ)2 = N −µ, diag (σ)2 . (10)

The operator is applied independently to all token slots. Applying it twice recovers the original posterior. We use it as a forward–backward consistency relation between qab and qba , not as an exact inverse under the normalized composition above. To compare two factorized K-token posteriors q and q ′ , let µ, µ′ , ℓ, ℓ′ ∈ RK×d denote their stacked means and logvariances. We define the mean-and-log-variance discrepancy ∥µ − µ′ ∥2F + λℓ ∥ℓ − ℓ′ ∥2F , (11) Kd where ∥ · ∥F is the Frobenius norm and λℓ controls the weight of log-variance matching. For the composed posc terior, ℓa = log((σ ac )2 ) is computed element-wise. The temporal-constraint losses are D(q, q ′ ) =

Lcomp = D(qac , q ac ),

(12)

Lrev = D(qab , R[qba ]). (13) Composition matches the direct and composed posteriors, while reversal matches the forward posterior with the transformed backward posterior. Let S = {rec, prior, comp, rev}. The complete pretraining objective is X LDLAM = λs Ls , (14) s∈S

where λs weights the loss indexed by s. The shared transition encoder jointly predicts and optimizes the mean and log-variance, giving the variance terms an additional training signal beyond reconstruction and the mean-based temporal constraints. Reconstruction and downstream policy transfer use only the resulting mean.

Transfer to World Action Modeling After pretraining, we discard the reconstruction decoder and freeze the transition encoder. Let Itm denote the frame at time t from camera view m ∈ {1, . . . , M }, where M is the number of views. Starting from the current demonstration time T , the frozen encoder Fϕ extracts transitions from H consecutive frame pairs. For h = 0, . . . , H − 1, m m m (µm (15) h , ℓh ) = Fϕ (IT +h , IT +h+1 ). Here Fϕ includes the encoder and log-variance clipping in Eq. (2); h indexes a step within the H-step target sequence, and both outputs have shape K × d. Only µm h is used as the downstream latent target; the log-variance is not passed to the policy. We use joint flow matching (Lipman et al. 2023; Tang et al. 2026a) to generate the view-specific latent trajectories together with the executable robot-action trajectory: M X u Ltransfer = λu LFM + λm Lm (16) FM . m=1

Here LuFM is the flow-matching loss for the executable action stream, Lm FM is the loss for the latent stream of view m, and λu , λm are their weights. Future demonstration frames and the frozen encoder are used only to construct targets during policy training. At inference time, the policy generates both streams, but only the robot-action stream is executed.

CALVIN

Method

1,795,025 (28.6%) sampling weight: 200

RoboTurk

134,496 (2.1%)

NYU Door

sampling weight: 10

15,221 (0.2%) sampling weight: 5

Cable Routing

42.3%

VIOLA

15,223 (0.2%)

6.27M

sampling weight: 20

training samples

60,987 (1.0%) sampling weight: 3

1.1% 0.6% 1.1% 1.1%

AutoLab UR5

JaCo-Play

2.1%

4.2% 1.1%

74,883 (1.2%)

sampling weight: 5

4.2%

10.6%

31.7%

52,715 (0.8%) sampling weight: 20

TACO-Play

126,298 (2.0%) sampling weight: 5

TOTO

236,632 (3.8%)

BridgeData V2

sampling weight: 5

579,207 (9.2%)

Fractal

Type Size Easy Med Hard V-Hard Avg.

RT-2 AR RoboTron Mani AR

7B 4B

75.5 35.3 30.7 85.5 67.7 76.7

15.2 81.0

39.2 77.7

GR-1 PAD Evo-1

VA 0.2B 76.6 35.3 46.0 VA – 81.8 65.1 56.7 VA 0.8B 89.2 76.8 77.2

44.0 87.2 79.2

50.5 72.7 80.6

π0.5 π0 SmolVLA π0 +ALAM

FM FM FM LA

3B 3B 2B 3B

68.2 71.8 87.1 89.3

41.7 41.7 70.0 85.0

28.0 30.0 64.0 82.0

43.8 47.9 68.2 85.0

π0 +DLAM

LA

3B

90.3 84.8 84.0

91.3

87.6

37.3 48.2 51.8 83.6

sampling weight: 50

3,182,825 (50.7%) sampling weight: 150 OUTER RING — available training-sample share INNER RING — normalized sampling probability

Figure 4: Action-free pretraining mixture. The outer ring shows the available training-sample share, whereas the inner ring shows the normalized sampling probability over the 11 video sources.

Experiments Experimental Setup Action-free pretraining. All latent-action models are pretrained on the same mixture of 11 action-free robot-video datasets, drawn mainly from Open X-Embodiment (O’Neill et al. 2024) and CALVIN (Mees et al. 2022). Figure 4 shows the sampling weights. We use the same visual tokenizer, Transformer capacity, source-conditioned decoder, data order, and training budget for every controlled variant. Training runs for 57 epochs on 64 AMD MI308X GPUs with AdamW, a peak learning rate of 10−4 , weight decay of 10−4 , and a per-device batch size of 64. We set λrec = 1, λprior = 0.005, λcomp = λrev = 0.05, and λℓ = 0.1. Policy transfer. We transfer each frozen transition encoder to the same π0 policy (Black et al. 2024), which uses PaliGemma-2B as the vision-language backbone and Gemma-300M as the action expert (Gemma Team et al. 2024). For each third-person and wrist-camera view, H + 1 demonstration frames provide H transition-mean targets. Only the policy backbone, action expert, and projection layers are updated. We train all variants with AdamW using a learning rate of 5 × 10−5 , weight decay of 10−4 , and a per-device batch size of 32. More implementation details are provided in the Technical Supplement. Benchmarks and baselines. On MetaWorld (McLean et al. 2025), we train one policy across 50 tasks, grouped into Easy, Medium, Hard, and Very Hard tiers. LIBERO (Liu et al. 2023) evaluates manipulation generalization through its Spatial, Object, Goal, and Long suites. We further evaluate on a Piper 6-DoF arm using four real-world tasks: insert cylinder, insert cube, arrange flowers, and hang cup. For transition representation and reconstruction, we compare with a reconstruction-only LAM and deterministic

Table 1: Results on MetaWorld MT50. Avg. is the macroaverage over the four difficulty tiers, computed before rounding the displayed tier values. AR: autoregressive; FM: flow matching; VA: video-augmented; LA: latent action.

ALAM (Tang et al. 2026a). The MetaWorld table also includes RT-2 (Zitkovich et al. 2023), RoboTron Mani (Yan et al. 2024), GR-1 (Wu et al. 2023), PAD (Guo et al. 2024), Evo-1 (Lin et al. 2025), π0.5 (Physical Intelligence et al. 2025), and SmolVLA (Shukor et al. 2025). On LIBERO, we additionally report OpenVLA (Kim et al. 2024), CoTVLA (Zhao et al. 2025), DreamVLA (Zhang et al. 2025b), OneWM-VLA (Tang et al. 2026b), GR00T-N1 (Bjorck et al. 2025), LAPA (Ye et al. 2024), UniVLA (Bu et al. 2025), and JALA (Luo et al. 2026).

Evaluating Transition Representations Evaluation protocol. We sample held-out windows from 11 datasets using k = 10 frames per segment, evaluating reconstruction up to 5k and temporal relations up to 10k. Direct reconstruction decodes the transition between the endpoints, whereas cumulative reconstruction composes adjacent k-frame transitions left to right before decoding. For temporal diagnostics, z denotes the posterior mean for DLAM and the deterministic embedding for point-valued baselines. We use the common length-aware probe h

1 X Ch (z1:h ) = √ zi , h i=1

(17)

which matches DLAM’s training rule at h = 2. For h > 2, it measures per-length consistency rather than associativity. To reduce sensitivity to model-specific latent magnitudes, we define mean |x − y| . (mean |x| + mean |y|) + ϵ 2

Rsym (x, y) = 1

(18)

µ µ We report Rcomp = Rsym (zdirect , Ch (z1:h )) and Rrev = Rsym (za→b , −zb→a ). Spans from 3k to 10k lie beyond the directly supervised k +k → 2k relation. Metrics are computed per window and then averaged; these mean-based probes do not test correlation-aware variance composition beyond the supervised span.

LAM

ALAM

OOD region ( ¸ 3k)

DLAM

Paired clip-bootstrap 95% CI (A only, n = 32)

Direct reconstruction (A.1) Scale-Normalized Composition Residual

(a) PSNR"

(b) SSIM"

(c) LPIPS#

0.95 4

0.20

PSNR (dB)

30

¹ Rcomp #

3 2 1

4:04

25

0.90

0:057

0.85

20

0.15 0.10

0:062

0.80 0.05

0 k

2k

3k

4k

5k

k

Temporal span t

2k

3k

4k

5k

k

Temporal span t

2k

3k

4k

5k

Temporal span t

Cumulative reconstruction

(d) PSNR"

(A.2) Scale-Normalized Reversal Residual

(e) SSIM" 0.20

30

PSNR (dB)

4 ¹ Rrev #

3 2

0.90 25

0.15

0.85 1:08

20

0:036

0.80

1 0

(f) LPIPS#

0.95

0:040

0.10 0.05

k k

2k

3k

4k

5k

6k

7k

8k

9k

10k

2k

3k

4k

Temporal span t

Temporal span t

(A) Temporal-relation span degradation

5k

k

2k

3k

4k

5k

Temporal span t

k

2k

3k

4k

5k

Temporal span t

(B) Absolute reconstruction quality

Figure 5: Held-out transition diagnostics and reconstruction. A. Scale-normalized mean composition and reversal residuals over increasing temporal spans. The shaded region marks spans beyond the directly supervised k +k → 2k composition relation. B. Absolute reconstruction quality across endpoint spans for direct and cumulative decoding. Method

Type Size Spatial Object Goal Long Avg.

OpenVLA CoT-VLA

AR AR

7B 7B

84.7 87.5

88.4 91.6

79.2 53.7 76.5 87.6 69.0 83.9

DreamVLA VA 0.4B OneWM-VLA VA 3B

97.5 98.2

94.0 99.6

89.5 89.5 92.6 99.0 95.1 98.0

GR00T N1 π0 π0.5

FM FM FM

2B 3B 3B

94.4 96.8 98.8

97.6 98.8 98.2

93.0 90.6 93.9 95.8 85.2 94.1 98.0 92.4 96.9

LAPA UniVLA JALA π0 +ALAM

LA LA LA LA

7B 9B 3B 3B

87.4 96.5 96.0 99.2

91.2 96.8 98.2 99.6

90.0 95.6 97.4 99.0

π0 +DLAM

LA

3B

99.6

99.8

99.6 97.1 99.0

65.4 92.0 96.0 94.4

83.5 95.2 96.9 98.1

Table 2: Success rates on LIBERO. Avg. is the mean over the four suites, computed before rounding the displayed suite values. AR: autoregressive; FM: flow matching; VA: videoaugmented; LA: latent action.

over progressively longer temporal spans, using the lengthaware probe in Eq. (17). Across the unseen 3k–10k range, residuals for DLAM stay low and degrade only mildly with span, while those for ALAM increase markedly. This pattern is consistent with improved approximate temporal consistency under the reported diagnostic, though it reflects behavior on the evaluated spans rather than a general guarantee at arbitrary horizons. Direct and cumulative reconstruction. We next evaluate whether the learned transition mean remains informative for reconstruction across larger endpoint spans. Across the 3k– 5k spans, DLAM achieves average PSNR values of 29.14 dB for direct reconstruction and 22.40 dB for cumulative reconstruction, improving over ALAM by 3.45 dB and 1.17 dB, respectively (Figure 5(B)). The same trend holds for the perceptual metrics: DLAM reduces LPIPS by 45.6% for direct reconstruction and by 26.1% after cumulative composition, with consistent improvements in SSIM. These results show that the learned mean retains information useful for endpoint reconstruction across the evaluated spans.

Downstream Policy Learning Latent transition visualization. Figure 3 illustrates the difference between the learned representations. For a fair comparison, we use checkpoints trained for 57 epochs under the same protocol for both methods. ALAM maps each transition to a single point, whereas DLAM predicts a mean and a diagonal variance vector. This visualization shows the difference in representation and is not an uncertainty-calibration analysis. Temporal-span diagnostics. Figure 5(A) reports composition and reversal residuals of the learned latent transitions

Table 1 reports the results on MetaWorld MT50. Under the controlled π0 transfer setting, DLAM achieves 87.6% average success, improving the deterministic transition baseline by 2.6 percentage points and the action-only policy by 39.7 points. The largest tier-level gain over ALAM occurs on Very-Hard, where success increases from 82.0% to 91.3%. As shown in Table 2, DLAM reaches 99.0% average success on LIBERO, with improvements across all four suites. The largest gain is observed on LIBERO-Long, where it improves the deterministic baseline by 2.7 percentage points. Figure 6

Insert Cylinder

Insert Cube

Arrange Flowers

Hang Cup

Figure 6: Real-world manipulation. Task-level success rates on a Piper 6-DoF arm for insert cylinder, insert cube, arrange flowers, and hang cup. We compare action-only and latent-action-augmented policies under the same training and evaluation. Reconstruction Quality

Scale-Normalized Relations

Variant

PSNR ↑

SSIM ↑

LPIPS ↓

µ Rcomp ↓

No temporal relations Matched mean-only (σ = 1, ρ = 0) Learned variance (ρ = 0) Full DLAM

21.236 22.109 22.084 22.400

0.7894 0.8057 0.8067 0.8086

0.1755 0.1644 0.1637 0.1623

1.2071 1.1808 1.1777 1.1662

Downstream Control

µ Rrev ↓

Avg.Success (%) ↑

1.0314 1.0371 1.0341 1.0010

76.6 82.1 85.3 87.6

Table 3: Controlled ablation of DLAM. Held-out reconstruction and relation results are averaged over 3k–5k temporal spans. Reconstruction uses cumulative decoding. Downstream performance reports the macro-average success rate on MetaWorld.

further evaluates the same transfer strategy on four real-world manipulation tasks. π0 + DLAM achieves the highest success rate on every task and reaches an average of 73.8%, compared with 63.8% for π0 + ALAM, 53.8% for π0.5 , and 40.0% for π0 . In particular, it improves over the deterministic transition baseline by 10 percentage points on each task. Together, these results show that pretraining with constraints on transition means and variances produces mean representations that transfer effectively under the reported simulation and real-world evaluation protocols. They establish the utility of the complete training formulation without attributing the gains to calibrated uncertainty, learned variance alone, or the temporal-span diagnostic.

Ablation Studies We compare four configurations using the same data, architecture, reconstruction objective, and training budget. No temporal relations retains the Gaussian encoder and prior but removes composition and reversal. Matched mean-only fixes σ = 1 and ρ = 0 while retaining the normalized mean composition, reversal constraint, and mean-prior penalty of the full model. Learned variance (ρ = 0) restores the predicted variance but propagates adjacent transitions independently, whereas the full DLAM additionally learns the shared coupling coefficient ρ. As shown in Table 3, the normalized mean constraints recover most of the reconstruction gain, increasing cumulative PSNR from 21.236 to 22.109 dB and downstream success from 76.6% to 82.1%. Learning the variance with ρ = 0 maintains comparable reconstruction quality, slightly reduces both relation residuals, and further improves success to 85.3%. Introducing the shared coupling coefficient yields the best results across all metrics: full DLAM reaches 22.400 dB PSNR, composition and reversal residuals

of 1.1662 and 1.0010, and 87.6% downstream success, an 11.0-point improvement over the variant without temporal relations. These results isolate the complementary contributions of normalized mean constraints, learned variance, and correlation-aware composition.

Conclusion We presented DLAM, a distributional latent-action model that represents visual transitions as diagonal Gaussians, applying normalized composition and reversal constraints to both the means and the dimension-wise variances. Under a common scale-normalized probe, the learned means show improved temporal consistency and retain more information for both direct and cumulative reconstruction. Controlled ablations reveal that normalized mean constraints drive most of this reconstruction improvement, while learned variance and correlation-aware composition contribute complementary gains in downstream control. The full formulation performs best across reconstruction quality, relational diagnostics, and policy transfer. In downstream learning, the frozen encoder contributes only posterior-mean transition sequences as auxiliary flow-matching targets, allowing the learned representation to enhance control without altering the policy backbone or execution interface. Limitations. DLAM imposes local constraints on equalgap triplets, leaving long-horizon generalization open. The variance objective may admit near-constant solutions. Since downstream transfer uses only posterior means, variance serves as an auxiliary training signal rather than calibrated uncertainty. The shared correlation may also miss contextor dimension-dependent dependencies.

References Babaeizadeh, M.; Finn, C.; Erhan, D.; Campbell, R. H.; and Levine, S. 2018. Stochastic Variational Video Prediction. In International Conference on Learning Representations. Bjorck, J.; Castañeda, F.; Cherniadev, N.; Da, X.; Ding, R.; Fan, L.; Fang, Y.; Fox, D.; Hu, F.; Huang, S.; et al. 2025. Gr00t n1: An open foundation model for generalist humanoid robots. arXiv preprint arXiv:2503.14734. Black, K.; Brown, N.; Driess, D.; Esmail, A.; Equi, M.; Finn, C.; Fusai, N.; Groom, L.; Hausman, K.; Ichter, B.; et al. 2024. π0 : A Vision-Language-Action Flow Model for General Robot Control. arXiv preprint arXiv:2410.24164. Bruce, J.; Dennis, M.; Edwards, A.; Parker-Holder, J.; Shi, Y.; Hughes, E.; Lai, M.; Mavalankar, A.; Steigerwald, R.; Apps, C.; et al. 2024. Genie: Generative Interactive Environments. In Proceedings of the 41st International Conference on Machine Learning. Bu, Q.; Yang, Y.; Cai, J.; Gao, S.; Ren, G.; Yao, M.; Luo, P.; and Li, H. 2025. UniVLA: Learning to Act Anywhere with Task-Centric Latent Actions. arXiv preprint arXiv:2505.06111. Chen, R.; Yang, Y.; Tang, Z.; Huo, D.; Lin, T.; Wu, H.; Liu, H.; Chen, Y.; Zheng, L.; Yuan, B.; et al. 2026. ABot-M0. 5: Unified Mobility-and-Manipulation World Action Model. arXiv preprint arXiv:2607.00678. Chi, C.; Xu, Z.; Feng, S.; Cousineau, E.; Du, Y.; Burchfiel, B.; Tedrake, R.; and Song, S. 2025. Diffusion Policy: Visuomotor Policy Learning via Action Diffusion. The International Journal of Robotics Research, 44(10–11): 1684–1704. Denton, E.; and Fergus, R. 2018. Stochastic Video Generation with a Learned Prior. In Proceedings of the 35th International Conference on Machine Learning, 1174–1183. Gao, S.; Zhou, S.; Du, Y.; Zhang, J.; and Gan, C. 2025. Adaworld: Learning adaptable world models with latent actions. arXiv preprint arXiv:2503.18938. Gemma Team; Mesnard, T.; Hardin, C.; Dadashi, R.; Bhupatiraju, S.; Pathak, S.; Sifre, L.; Rivière, M.; Kale, M. S.; Love, J.; et al. 2024. Gemma: Open models based on gemini research and technology. arXiv preprint arXiv:2403.08295. Guo, Y.; Hu, Y.; Zhang, J.; Wang, Y.-J.; Chen, X.; Lu, C.; and Chen, J. 2024. Prediction with Action: Visual Policy Learning via Joint Denoising Process. In Advances in Neural Information Processing Systems, volume 37, 112386–112410. Hu, Y.; Guo, Y.; Wang, P.; Chen, X.; Wang, Y.-J.; Zhang, J.; Sreenath, K.; Lu, C.; and Chen, J. 2024. Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations. arXiv preprint arXiv:2412.14803. Kim, M. J.; Pertsch, K.; Karamcheti, S.; Xiao, T.; Balakrishna, A.; Nair, S.; Rafailov, R.; Foster, E.; Lam, G.; Sanketi, P.; et al. 2024. OpenVLA: An Open-Source VisionLanguage-Action Model. arXiv preprint arXiv:2406.09246. Li, L.; Zhang, Q.; Luo, Y.; Yang, S.; Wang, R.; Han, F.; Yu, M.; Gao, Z.; Xue, N.; Zhu, X.; Shen, Y.; and Xu, Y. 2026a. Causal World Modeling for Robot Control. arXiv preprint arXiv:2601.21998.

Li, Q.; Gong, X.; Li, X.; Li, P.; Zhou, Q.; Ye, H.; Zhou, J.; and Mu, Y. 2026b. RotVLA: Rotational Latent Action for VisionLanguage-Action Model. arXiv preprint arXiv:2605.13403. Liang, A.; Czempin, P.; Hong, M.; Zhou, Y.; Biyik, E.; and Tu, S. 2025. CLAM: Continuous Latent Action Models for Robot Learning from Unlabeled Demonstrations. arXiv preprint arXiv:2505.04999. Lin, T.; Zhong, Y.; Du, Y.; Zhang, J.; Liu, J.; Chen, Y.; Gu, E.; Liu, Z.; Cai, H.; Zou, Y.; et al. 2025. Evo-1: Lightweight vision-language-action model with preserved semantic alignment. arXiv preprint arXiv:2511.04555. Lipman, Y.; Chen, R. T. Q.; Ben-Hamu, H.; Nickel, M.; and Le, M. 2023. Flow Matching for Generative Modeling. In International Conference on Learning Representations. Liu, B.; Zhu, Y.; Gao, C.; Feng, Y.; Liu, Q.; Zhu, Y.; and Stone, P. 2023. Libero: Benchmarking knowledge transfer for lifelong robot learning. Advances in Neural Information Processing Systems, 36: 44776–44791. Luo, H.; Wang, Y.; Zhang, W.; Yuan, H.; Feng, Y.; Xu, H.; Zheng, S.; and Lu, Z. 2026. Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the Wild. arXiv:2602.21736. McLean, R.; Chatzaroulas, E.; McCutcheon, L.; Röder, F.; Yu, T.; He, Z.; Zentner, K.; Julian, R.; Terry, J. K.; Woungang, I.; Farsad, N.; and Castro, P. S. 2025. Meta-World+: An Improved, Standardized, RL Benchmark. In The Thirtyninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track. Mees, O.; Hermann, L.; Rosete-Beas, E.; and Burgard, W. 2022. Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks. IEEE Robotics and Automation Letters, 7(3): 7327–7334. O’Neill, A.; Rehman, A.; Maddukuri, A.; Gupta, A.; Padalkar, A.; Lee, A.; Pooley, A.; Gupta, A.; Mandlekar, A.; Jain, A.; et al. 2024. Open x-embodiment: Robotic learning datasets and rt-x models. In 2024 IEEE International Conference on Robotics and Automation (ICRA), 6892–6903. IEEE. Physical Intelligence; Black, K.; Brown, N.; Darpinian, J.; Dhabalia, K.; Driess, D.; Esmail, A.; Equi, M.; Finn, C.; Fusai, N.; et al. 2025. π0.5 : a Vision-Language-Action Model with Open-World Generalization. arXiv preprint arXiv:2504.16054. Shen, Q.; Zhang, S.; Liao, Y.; Li, Q.; Tan, Z.; Wang, S.; Yan, S.; and Wang, X. 2026. World Action Models: A Survey. arXiv preprint arXiv:2606.20781. Shukor, M.; Aubakirova, D.; Capuano, F.; Kooijmans, P.; Palma, S.; Zouitine, A.; Aractingi, M.; Pascal, C.; Russi, M.; Marafioti, A.; et al. 2025. Smolvla: A vision-language-action model for affordable and efficient robotics. arXiv preprint arXiv:2506.01844. Tang, Z.; Hu, B.; Zhao, C.; Ma, D.; Pan, G.; and Liu, B. 2025. VLASCD: A visual language action model for simultaneous chatting and decision making. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 9223–9243.

Tang, Z.; Liu, H.; Chang, X.; Wu, C.; Huo, D.; Yang, Y.; Liu, B.; Cai, Z.; Xiong, F.; Xu, M.; et al. 2026a. ALAM: Algebraically Consistent Latent Action Model for VisionLanguage-Action Models. arXiv preprint arXiv:2605.10819. Tang, Z.; Yuan, S.; Bai, X.; Jing, Z.; Ma, D.; Pan, G.; and Liu, B. 2026b. One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy. arXiv preprint arXiv:2605.07931. Wang, Q.; Li, M.; Guan, J.; Ye, J.; Xie, S.; Liu, Y.; Chen, J.; Liang, Z.; Zhang, J.; Hu, X.; et al. 2026. Qwen-vla: Unifying vision-language-action modeling across tasks, environments, and robot embodiments. arXiv preprint arXiv:2605.30280. Wei, H.; Chen, X.; Zhang, C.; Pearce, T.; Chen, J.; Lamb, A.; Zhao, L.; and Bian, J. 2026. Learning Additively Compositional Latent Actions for Embodied AI. arXiv preprint arXiv:2604.03340. Wu, H.; Jing, Y.; Cheang, C.; Chen, G.; Xu, J.; Li, X.; Liu, M.; Li, H.; and Kong, T. 2023. Unleashing large-scale video generative pre-training for visual robot manipulation. arXiv preprint arXiv:2312.13139. Yan, F.; Liu, F.; Zheng, L.; Zhong, Y.; Huang, Y.; Guan, Z.; Feng, C.; and Ma, L. 2024. RoboTron-Mani: All-in-One Multimodal Large Model for Robotic Manipulation. arXiv preprint arXiv:2412.07215. Ye, S.; Jang, J.; Jeon, B.; Joo, S.; Yang, J.; Peng, B.; Mandlekar, A.; Tan, R.; Chao, Y.-W.; Lin, B. Y.; et al. 2024. Latent Action Pretraining from Videos. arXiv preprint arXiv:2410.11758. Yuan, H.; Liang, Z.; Chen, A.; Wang, Y.; Li, H.; Lin, P.; Huang, Y.; Lei, Z.; Zhang, T.; Zhang, J.; et al. 2026a. Qwenrobotmanip technical report: Alignment unlocks scale for robotic manipulation foundation models. arXiv preprint arXiv:2606.17846. Yuan, T.; Dong, Z.; Liu, Y.; and Zhao, H. 2026b. Fast-wam: Do world action models need test-time future imagination? arXiv preprint arXiv:2603.16666. Zhang, C.; Pearce, T.; Zhang, P.; Wang, K.; Chen, X.; Shen, W.; Zhao, L.; and Bian, J. 2025a. What Do Latent Action Models Actually Learn? arXiv preprint arXiv:2506.15691. Zhang, W.; Liu, H.; Qi, Z.; Wang, Y.; Yu, X.; Zhang, J.; Dong, R.; He, J.; Lu, F.; Wang, H.; et al. 2025b. Dreamvla: a vision-language-action model dreamed with comprehensive world knowledge. arXiv preprint arXiv:2507.04447. Zhao, Q.; Lu, Y.; Kim, M. J.; Fu, Z.; Zhang, Z.; Wu, Y.; Li, Z.; Ma, Q.; Han, S.; Finn, C.; et al. 2025. Cot-vla: Visual chainof-thought reasoning for vision-language-action models. In Proceedings of the Computer Vision and Pattern Recognition Conference, 1702–1713. Zitkovich, B.; Yu, T.; Xu, S.; Xu, P.; Xiao, T.; Xia, F.; Wu, J.; Wohlhart, P.; Welker, S.; Wahid, A.; et al. 2023. RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control. In Proceedings of the 7th Conference on Robot Learning, volume 229 of Proceedings of Machine Learning Research, 2165–2183.

Record · ID 411071 · SHA-256 bc0f5a739d8e2ba3
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.