DAGR: State-Conditioned Goal Representations via Difference-Aware Goal Cross-Attention Xing Lei1 Wenyan Yang2 Xuetao Zhang1∗ Donglin Wang3 1 Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University 2 Department of Electrical Engineering and Automation, Aalto University 3 School of Engineering, Westlake University [email protected]
arXiv:2607.13731v1 [cs.LG] 15 Jul 2026
Abstract Goal-conditioned reinforcement learning hinges on how the goal is encoded. Contrastive, metric, temporal-distance, and information-theoretic encoders differ in objective. They still share one trait. None of them sees the current state. Such a stateindependent embedding cannot mark which part of the goal still needs action. The policy must then recover that cue by inverting both encoders. We propose DAGR. It refines the static embedding of any late-fusion encoder into a state-conditioned one through multi-scale gated cross-attention. A near-identity gated residual preserves the base representation. Difference-aware Goal Cross-Attention then biases the attention scores using a per-token state-goal discrepancy map. On OGBench, DAGR improves navigation. Our ablations trace the gain to the gated residual, not to the difference bias that names the method. On manipulation and puzzle tasks it matches or falls below the base. DAGR is a structured refinement, not a universal improvement. Code is available at https://github.com/leixingxing1/DAGR.
1
Introduction
Goal-conditioned reinforcement learning (GCRL) trains agents that reach any specified goal state [Schaul et al., 2015, Plappert et al., 2018, Liu et al., 2022], with applications spanning navigation [Shah et al., 2021a, Hirose et al., 2023], manipulation [Nair et al., 2018, Kalashnikov et al., 2018, Fang et al., 2022], and multi-task control [Lynch et al., 2020, Jang et al., 2022]. In the offline setting [Lange et al., 2012, Levine et al., 2020], the learned goal representation governs whether the policy generalizes to unseen state-goal pairs [Yang et al., 2022b, Ma et al., 2022b, Park et al., 2025]. Recent methods disagree sharply on how to shape it. Contrastive learning [Eysenbach et al., 2022, Zheng et al., 2024b], metric learning [Ma et al., 2023, Wang et al., 2023, Park et al., 2024], temporaldistance modeling [Park et al., 2026], and information-theoretic compression [Alemi et al., 2017, Shah et al., 2021b] pull the embedding in different directions. They nonetheless agree on one structural point. The state encoder ψ and the goal encoder ϕ are trained apart, and their outputs meet only inside the policy and value heads [Park et al., 2025]. We call the resulting ϕ(g) a state-independent goal representation. That independence keeps these objectives tractable and their guarantees clean [Park et al., 2026], but it costs something. Because ϕ(g) never sees the current state, it cannot say which aspects of the goal still call for action, and the policy must recover that cue by inverting both encoders and comparing their outputs. The bottom panel of Figure 1 makes this concrete. One maze goal hands every state the same embedding, and therefore the same hint, no matter which way the agent must move. We formalize the cost as an information bound (Proposition 4.1) and through Rademacher composition (Proposition 4.2). ∗
Corresponding author.
Preprint. Under review.
State �
Goal �
State �
Goal �
State �
Goal � Dual Goal Rep
Dual Goal Rep
�(�)
�(�)
Concat �� = [�; �] Policy �(�|�� )
Value �(�� )
Policy �(�|�� )
Value �(�� )
(b) Dual Goal Representation ✓ Flexible Representation Learning ✗ Goal Embedding is State-independent
Dual goal representation (late fusion)
�2
�1
�(�|�)
Concat �� = [�; �(�)]
(a) Original ✓ Captures Pixel-level Correspondence ✗ Limited Representation Flexibility
�3
DGCA: Difference-aware Goal Cross-attention
�4
Concat �� = [�; �(�|�)]
Value �(�� )
Policy �(�|�� )
(c) DAGR (Ours) ✓ State-conditioned Goal Embedding ✓ Flexible Representation Learning
DAGR (state-conditioned refinement, ours)
�3
�5 Condition on �
�2
�1
�
�(�): Same hint at every state -- 'goal is over there'
�4
�5
�
�(�|�): Different hint at each �� -- 'go this way next'
Figure 1: From a state-independent goal embedding to a state-conditioned one. The top panel contrasts the data flows of late fusion, Dual [Park et al., 2026], and DAGR. The bottom panel shows the consequence in a shared maze with the same g and s1 , . . . , s5 . Under Dual, ϕ(g) supplies the same hint at every si . Under DAGR, the hint reflects the discrepancy between si and g. Formally, ′ ∗ ′ ′ ∗ ′ DAGR defines the goal representation φ∨ DAGR (g | s)(s ) = d (s , g)·∆s,g (s ), where d (s , g) denotes the optimal temporal distance from s′ to g and ∆s,g : S → [0, 1] weights the relevance of each s′ to the current state-goal mismatch. This augments the Dual functional φ∨ (g)(s′ ) = d∗ (s′ , g) with a state-dependent term, which is the property we call difference-aware. We approximate ∆s,g by the per-token discrepancy map of multi-scale D GCA in Section 4.2. What the figure isolates, and what our ablations confirm carries the gain, is the state-conditioning itself. The particular form of ∆s,g is a design choice on top of it. We propose DAGR, a module that refines the static ϕ(g) of any late-fusion encoder into a stateconditioned ϕ(g | s) through cross-attention, as sketched in the top panel of Figure 1. Refining a representation that already carries guarantees is delicate, so two choices keep it safe. The first is a multi-scale gated residual, initialized so that ϕ(g | s) starts out equal to ϕ(g). The downstream value and policy losses then shape a state-conditioned perturbation of the base, and they open the gate only where doing so pays. The second is an attention rule that adds a learnable nonnegative bias, derived from a per-token state-goal difference map, to the usual similarity scores. Section 4 gives the architecture. It also shows that DAGR preserves the sufficiency and the noise invariance of any base representation (Theorems 4.3 and 4.4), that the added approximation error stays under a gate-controlled bound (Theorem 4.5), and that on discrepancy-structured tasks it admits a tighter sample-complexity upper bound than late fusion (Proposition 4.6). We evaluate DAGR on OGBench [Park et al., 2025], layered on Dual [Park et al., 2026] with GCIVL [Park et al., 2025] downstream. On state-based navigation it improves the success rate on every task but one, and the gains are largest on the mazes where Dual is weakest. On visual navigation it is the strongest of the six goal-representation baselines we compare against, so the directional signal survives the pixel 2
encoder. The single exception is PointMaze, whose state is low-dimensional enough that the encoder barely compresses, and Proposition 4.6 predicts a null result exactly there. On manipulation and on the discrete puzzles the outcome is mixed. DAGR matches the base on some tasks and falls below it on others, and the tasks where it regresses are the ones that violate the structural condition of Definition 3. Our ablations then ask which part of the module produces the gains that do appear, and the answer is not the part the method is named after. The gated residual carries the improvement, whereas removing the difference bias leaves navigation performance intact. We treat this as a finding rather than an inconvenience, and we state it in the main text. Contributions. (i) We isolate the state-independence of the goal encoder as a design axis orthogonal to the choice of representation objective (Figure 2), and quantify its cost with an information bound and a Rademacher composition bound. (ii) We propose DAGR, a module that state-conditions any late-fusion goal encoder while provably retaining sufficiency, exogenous noise invariance, and a gate-controlled bound on the added value error. (iii) We state a structural condition on the task that predicts in advance where state-conditioning helps and where its optimum collapses back to the base, and both predictions hold on OGBench. (iv) We attribute the gain component by component, and report that the difference bias, which gives the module its name, contributes an inductive bias at initialization and little beyond it.
2
Related Work
Offline Goal-Conditioned RL. Offline GCRL trains goal-reaching policies from fixed datasets [Lange et al., 2012, Levine et al., 2020]. Existing paradigms include behavioral cloning [Lynch et al., 2020, Ghosh et al., 2021], value-based estimation [Kumar et al., 2020, Kostrikov et al., 2022, Park et al., 2025], hierarchical or subgoal-based planning [Park et al., 2023, Ahn et al., 2025, Zhou and Kao, 2025, Lei et al., 2025, Giammarino and Qureshi, 2026], decision transformer [Lei et al., 2026], contrastive value learning [Eysenbach et al., 2022, Myers et al., 2024, Zheng et al., 2024b], explicit subgoal planning [Eysenbach et al., 2019, Wang et al., 2024b], dual optimization [Ma et al., 2022a, Sikchi et al., 2024], and generative trajectory modeling [Jain and Ravanbakhsh, 2024, Bao et al., 2026]. OGBench [Park et al., 2025] shows that no single method dominates across all task categories. DAGR is orthogonal to this axis, refining the goal representation consumed by any such downstream algorithm rather than proposing a new policy objective. objective on ϕ no rep-learning with rep-learning
Goal Representation Learning. Several families of objectives have been proposed for learning goal embeddings that generalize across unseen state-goal combinations [Park et al., 2026, Kim et al., 2026]. Metric-based methods align geometric and temporal distances (VIP [Ma et al., 2023], HILP [Park et al., 2024], QRL [Wang et al., 2023] and its quasimetric extensions [Myers et al., 2025a, Zheng et al., 2026]). Temporal-distance methods fit the embedding directly to a value or distance signal (TRA [Myers et al., 2025b], Dual [Park et al., 2026]). Information-theoretic methods compress the goal representation via a variational information bottleneck [Tishby et al., 2000, Alemi et al., 2017], applied to GCRL by Shah et al. [2021b] and Park et al. [2023]. Self-predictive methods use bootstrapped temporal consistency in the form of BYOL-γ [Lawson et al., 2025], building on the self-predictive representation framework of Schwarzer et al. [2021]. All such methods produce a state-independent ϕ(g). DAGR is orthogonal and composable, as it operates on the output ϕ(g) rather than on the objective that produces it.
VIP, TRA, HILP, BYOL-γ, Dual rep-learning on ϕ(g), state-independent
DAGR (ours) rep-learning on ϕ(g), refined to ϕ(g | s), state-conditional
UVFA
BVN, FiLM, SCRL
state-independent ϕ(g), no rep-learning loss
state-conditional ϕ(s, g), no rep-learning loss
state-independent ϕ(g)
state-conditional ϕ(s, g)
goal-embedding signature
Figure 2: Two design axes for goal encoders in offline GCRL. Horizontal: whether ϕ is state-conditional (ϕ(s, g) or ϕ(g | s)) or state-independent (ϕ(g)). Vertical: whether ϕ is trained with a representation-learning objective. Among these, DAGR is the only one that Two axes prior work conflates. Goal representation learn- satisfies both. Quasimetric methods act ing shapes ϕ through an auxiliary objective [Ma et al., on a third orthogonal axis (constraining 2023, Myers et al., 2025b, Park et al., 2024, Lawson et al., the value-function class) and are omitted 2025, Park et al., 2026] but keeps it a function of g alone. here. A separate axis, often conflated with the first, asks whether ϕ is state-conditional. UVFA [Schaul et al., 2015] keeps ϕ(g) state-independent, whereas BVN [Yang et al., 2022a] uses ϕ(s, g) and shows 3
it to be strictly more expressive. FiLM [Perez et al., 2018] and SCRL [Zheng et al., 2024a] also condition on s but introduce no representation-learning objective. DAGR is closest in spirit to BVN on this axis but differs in three respects. First, it refines a representation-learned base ϕ(g) rather than training ϕ(s, g) monolithically. Second, it reduces to the base encoder at initialization through a gated residual that preserves sufficiency and noise invariance (Theorems 4.3 and 4.4). Third, it biases attention with an explicit per-token difference map. A third orthogonal line constrains the value-function class to be quasimetric [Pitis et al., 2020, Wang and Isola, 2022b,a, Wang et al., 2023], exploiting state-goal geometry on the value head rather than on ϕ. Figure 2 summarizes these axes. Cross-Attention in Computer Vision and RL. Cross-attention has become a standard mechanism for selective information flow across modalities in computer vision [Carion et al., 2020, Jaegle et al., 2021, Alayrac et al., 2022], with recent variants modifying the attention rule itself such as Gated Attention [Qiu et al., 2026]. In RL, it underlies trajectory modeling [Chen et al., 2021, Reed et al., 2022], in-context value inference [Xu et al., 2026], and an expanding body of work on visionlanguage-action models [Brohan et al., 2023b,a, Team et al., 2024, Kim et al., 2024, Black et al., 2024, Wang et al., 2024a, Huang et al., 2025, Zhong et al., 2026]. FiLM [Perez et al., 2018, Hill et al., 2020] applies feature-wise modulation in goal-conditioned RL but is spatially uniform. To our knowledge, DAGR is the first to introduce a learnable difference bias into goal-conditioned cross-attention and to deploy it at multiple scales over a temporal-distance goal representation.
3
Preliminaries
3.1
Offline Goal-Conditioned RL
A goal-conditioned MDP (GCMDP) [Kaelbling, 1993] is a tuple M = (S, A, G, p, r, γ) with goal space G ⊆ S, sparse reward r(s, Pg) = 1[s = g], and discount γ ∈ (0, 1). A policy π : S ×G → ∆(A) induces value V π (s, g) = Eπ [ t γ t r(st , g) | s0 = s] with optimum V ∗ (s, g). We define the optimal temporal distance d∗ (s, g) = logγ V ∗ (s, g), which equals the shortest-path length in deterministic environments. The offline dataset D = {(si , ai , s′i )}N i=1 is fixed. 3.2
Late Fusion and the State-Independence of ϕ(g)
We focus on the late-fusion paradigm shared by current goal-representation methods (cf. Figure 1, left). A state encoder ψ : S → Rds and a goal encoder ϕ : G → Rd are trained independently, and the policy is parameterized as πθ (a | ψ(s), ϕ(g)). The defining property of late fusion is that the goal encoder ϕ takes no state input. For any fixed goal g, the same vector ϕ(g) is supplied to the policy and value heads regardless of the current state s. The policy network must therefore infer at each step which components of ϕ(g) are currently actionable. DAGR replaces ϕ(g) by a state-conditioned ϕ(g | s) that exposes this information explicitly at the representation level.2 Definition 1 (Sufficient Goal Representation [Park et al., 2026]). A representation ϕ : G → Rd is ∗ sufficient for optimal control if there exists π ∗ : S × Rd → ∆(A) such that V π (·|s,ϕ(g)) (s, g) = ∗ V (s, g) for all (s, g) ∈ S × G. 3.3
Multi-Head Cross-Attention
q ×d Given Q ∈ Rn√ , K ∈ Rnk ×d , V ∈ Rnk ×dv , cross-attention is CrossAttn(Q, K, V ) = ⊤ softmax(QK / d) V [Vaswani et al., 2017]. Multi-head attention computes H such operations in parallel with distinct projections WhQ , WhK , WhV and concatenates them through W O . The softmax weights measure similarity, so attention concentrates on tokens that resemble the query. For goalconditioned control, the policy must act on the components in which state and goal disagree rather than on the components in which they agree. This mismatch between the inductive bias of standard cross-attention and the requirements of goal-conditioned control motivates the difference-bias term introduced in Section 4.
2 Throughout the paper, ϕ(g | s) denotes a deterministic function ϕ : G × S → Rd , with the vertical bar reading “ϕ of g, parameterized by s.” This is a notational convention and should not be read as a conditional distribution.
4
�(�)
�(�)
�
Fine Scale
Medium Scale
Coarse Scale
Token Projection
Token Projection
Token Projection
� = ��
s_tok
s_tok, g_tok
� = ��
s_tok
s_tok, g_tok
Difference Map
Difference Map
�
MH-DGCA
MH-DGCA (Eq. 17)
attn_out
�
(Eq. 17)
attn_out
� = ��
s_tok
s_tok, g_tok
Difference Map �
MH-DGCA (Eq. 17)
attn_out
Gated Residual Block
Gated Residual Block
Gated Residual Block
level₀
level₁
level₂
(Eq. 18-19)
(Eq. 18-19)
(Eq. 18-19)
Learnable Fusion (Eq. 16)
�(�|�)
State-Aware Goal Rep
Figure 3: Multi-scale D GCA architecture. Three scale levels with token counts Tℓ ∈ {16, 8, 4} project the flat encoder outputs ψ(s) and ψg (g) into aligned pseudo-token spaces. Each level computes a difference map ∆(ℓ) , runs multi-head D GCA with ϕ(g) as the query, and produces a gated residual update. Outputs are combined through learnable fusion weights to yield ϕDAGR (g | s).
4
Difference-Aware Goal Representations
We begin with the core idea of DAGR, following the exposition pattern of the Dual representation [Park et al., 2026]. Recall from Section 3 that the optimal temporal distance d∗ (s, g) = logγ V ∗ (s, g) corresponds in deterministic environments to the shortest-path length from s to g. The Dual representation characterizes a goal g by the set of optimal temporal distances d∗ (·, g) from all other states. In a discrete state space S = {s1 , . . . , sK }, this corresponds to the vector ⊤ φ∨ (g) = d∗ (s1 , g), d∗ (s2 , g), . . . , d∗ (sK , g) , (1) and in general to the functional φ∨ (g)(s′ ) = d∗ (s′ , g), which depends on g alone. DAGR extends this construction with a state-dependent weighting: ∗ ⊤ ∗ φ∨ , (2) DAGR (g | s) = d (s1 , g) · ∆s,g (s1 ), . . . , d (sK , g) · ∆s,g (sK ) ′ ∗ ′ ′ in the discrete case, and to the functional φ∨ DAGR (g | s)(s ) = d (s , g) · ∆s,g (s ) in general. Here ′ ∆s,g : S → [0, 1] is large when s is informative about the difference between s and g and small when s′ is irrelevant. We call φ∨ DAGR the difference-aware goal representation of the GCMDP M. In continuous environments we approximate ∆s,g by the per-token difference map of multi-scale Difference-aware Goal Cross-Attention in Section 4.2.
The rest of this section unpacks this construction. We first show what a state-independent goal representation cannot encode (Section 4.1), then specify the single-scale (Section 4.2) and multi-scale (Section 4.3) DGCA blocks that realize ∆s,g , and conclude with the theoretical properties of the resulting representation (Section 4.4). The joint training procedure with the downstream offline GCRL algorithm is deferred to Appendix A. 4.1
The Representation-Level Bottleneck of Late Fusion
To make precise what is missing in a state-independent ϕ(g), let S, G, and A∗ denote the random variables for state, goal, and optimal action. Under late fusion, conditioned on S = s the optimal action A∗ is determined by π ∗ (· | s, g), which depends on g only through ϕ(g). This yields the Markov chain G → ϕ(G) → A∗ given S = s, from which the data processing inequality yields the following bound. Proposition 4.1 (Information Bottleneck of State-Independent Goal Representations). For any state-independent ϕ, I(A∗ ; ϕ(G) | S) ≤ I(A∗ ; G | S), 5
with equality if and only if ϕ(G) retains all goal information relevant to A∗ given S. The equality condition fails precisely when the optimal action depends on a joint property of (S, G) that ϕ cannot resolve without access to S. A canonical example is when π ∗ (a | s, g) = f (s − g), the relational case. Even when ϕ is information-theoretically sufficient, the downstream policy still pays a sample-complexity price. Realizing π ∗ from inputs (ψ(s), ϕ(g)) requires the network to implicitly invert both encoders before computing the difference. This composition is strictly harder than receiving s − g directly. Proposition 4.2 (Late Fusion Admits a Looser Complexity Bound on Relational Tasks). Suppose π ∗ (a | s, g) = f (s − g). Then realizing π ∗ from the late-fusion inputs (ψ(s), ϕ(g)) requires composing f with measurable inverse selections of ψ and of ϕ. The composition property of Rademacher complexity [Bartlett and Mendelson, 2002] then yields an upper bound on the policy class that exceeds the corresponding bound under direct access to s − g by a factor Lψ−1 Lϕ−1 ≥ 1. We are precise about what this does and does not say. It compares two upper bounds. It does not establish that the true Rademacher complexity of the realized function class is larger under late fusion. The proposition motivates the architecture. It does not prove that late fusion must be worse. Both observations point to the same fix: surface state-conditioned goal information at the representation level rather than asking the policy to recover it. This motivates the module we introduce next. Full proofs are in Appendices D.1 and D.2. 4.2
Single-Scale Difference-Aware Goal Cross-Attention (D GCA)
We aim to construct a transformation that takes the static ϕ(g) ∈ Rd , the state encoding ψ(s) ∈ Rds , and a separate goal-image encoding ψg (g) ∈ Rds , and returns a state-conditioned representation ϕ(g | s) ∈ Rd . The transformation should satisfy two properties. First, at initialization it should reduce to the identity map on ϕ(g), so that training begins from the base representation. Second, it should be able to route information from state features that disagree with the goal into the refined representation. This routing is trainable end to end via the downstream policy and value losses. The building block is a multi-head cross-attention in which the goal supplies the query and the state supplies the keys and values, with an additive bias derived from a per-token discrepancy map. Figure 3 shows the architecture. Three ingredients realize it. The first is a token decomposition. Let T denote the pseudo-token count, H the number of heads, dk the per-head key dimension, and dm := Hdk . With separately learned projections Ws , Wg ∈ R(T dm )×ds , stok = Reshape Ws ψ(s) , gtok = Reshape Wg ψg (g) ∈ RT ×dm . (3) The two projections are untied, which aligns the token spaces position-wise and renders the per-token (t) (t) difference stok − gtok meaningful. The second ingredient is the difference map. A normalized ℓ2 difference measures the mismatch at each token position, ∆t =
(t)
(t)
(t′ )
(t′ )
∥stok − gtok ∥2
∈ [0, 1],
maxt′ ∥stok − gtok ∥2 + ε
∆ = (∆1 , . . . , ∆T ),
(4)
with ε = 10−8 for numerical stability. The maximum is taken within the same (s, g) pair, so ∆ ranks tokens by relative mismatch and is invariant to the absolute scale of state-goal differences. A value near one marks a position at which state and goal still differ, and a value near zero marks a position at which they agree. The third ingredient is the attention rule that uses ∆ as an additive bias in logit space. Definition 2 (Difference-Aware Goal Cross-Attention). Given a goal query Q ∈ R1×dk , state keys and values K, V ∈ RT ×dk , and a difference map ∆ ∈ [0, 1]T as defined in Equation (4), QK ⊤ DGCA(Q, K, V, ∆) = softmax √ + ζ(λ) · ∆ V ∈ R1×dk , (5) dk where ζ : R → R≥0 is defined by ζ(x) := softplus(x) = log(1 + ex ), and λ ∈ R is a learnable scalar parameter. We initialize λ0 = −5, which yields ζ(λ0 ) = log(1 + e−5 ) ≈ 0.0067. 6
The reparameterization through ζ keeps the bias non-negative, so higher token-wise discrepancy only increases the corresponding attention logit. At initialization, ζ(λ0 ) is negligible and D GCA reduces to standard cross-attention. As λ grows during training, the bias additively boosts the logits of tokens with high ∆t . A per-head λh lets different heads specialize at different bias intensities. The multi-head version applies H parallel heads with projections WhQ ∈ Rd×dk for the goal query and WhK , WhV ∈ Rdm ×dk for the state tokens, ϕ(g)WhQ (stok WhK )⊤ √ headh = softmax + ζ(λ )∆ stok WhV , attn_out = Concath (headh ) W O , h dk (6) with W O ∈ R(Hdk )×d and a per-head λh that permits different heads to adopt different bias intensities. The attention output is integrated into the goal representation through a gated residual block, followed by layer normalization and a feed-forward network with its own gated residual, x = ϕ(g) + σ(αattn ) ⊙ attn_out, x ← LayerNorm(x), (7) out = x + σ(αffn ) ⊙ FFN(x), out ← LayerNorm(out), (8) Here σ is the elementwise sigmoid, ⊙ the Hadamard product, and FFN a two-layer network with GELU activation. The gate vectors αattn , αffn ∈ Rd are initialized componentwise to −5, which yields σ(α) ≈ 0.0067. Vector-valued gates permit each coordinate of the goal representation to be modulated independently. We write DGCA-BlockT (ϕ(g), ψ(s), ψg (g)) for the output of Equation (8). The gates start nearly closed, so ϕ(g | s) ≈ ϕ(g) and the temporal-distance structure of the base survives. They open during training only where doing so reduces the downstream losses, which yields an automatic curriculum from late fusion to state-conditioned refinement. 4.3
Multi-Scale Extension
A single token count T commits the module to one analysis granularity. In practice, the spatial scale at which state-goal discrepancies are informative varies across tasks. Navigation often hinges on coarse global offsets, while manipulation may depend on a small object-level mismatch. We therefore apply L independent D GCA blocks at decreasing token counts T1 > T2 > · · · > TL , each with its own token projections and attention parameters, and combine them through learnable fusion weights, levelℓ = DGCA-BlockTℓ ϕ(g), ψ(s), ψg (g) ,
ϕDAGR (g | s) =
L X
softmax(w)ℓ · levelℓ .
ℓ=1
(9) The fusion logits w ∈ RL are initialized to zero so that all scales contribute equally at the start. We use L = 3 with (T1 , T2 , T3 ) = (16, 8, 4) throughout the paper. Fine scales (large T ) carry local information, coarse scales (small T ) aggregate global structure, and the data-driven w trades off between them based on the downstream loss. DAGR is trained jointly with the downstream offline GCRL algorithm through standard value and policy losses propagated through multi-scale D GCA (MS-D GCA), with the full procedure (Algorithm 1) given in Appendix A. 4.4
Theoretical Properties
We state five properties. Three are safety guarantees that hold on any task. Two tie the value of refinement to the following structural condition. Definition 3 (Discrepancy Structure). A goal-conditioned task with optimal policy π ∗ has discrepancy structure if there exists a measurable function h : S × S → Z and a measurable function ρ : Z → ∆(A) such that π ∗ (a | s, g) = ρ h(s, g) for all (s, g) ∈ S × G, (10) and such that h is not a function of g alone. We call h the discrepancy map. A task admitting (10) with h a function of g alone is called goal-only. Throughout, we treat ϕ, ψ, and ψg as fixed and consider the map ϕ(g) 7→ ϕDAGR (g | s) for any fixed s ∈ S. Full statements and proofs are deferred to Appendix D. A base representation that is already sufficient stays sufficient after refinement, because each scale block is an injective perturbation of the identity. 7
Theorem 4.3 (Sufficiency Preservation). Assume ϕ is sufficient (Definition 1), that the residual perturbation δℓ of each scale is Lℓ -Lipschitz in ϕ(g) with ∥σ(α)∥∞ Lℓ < 1, and that the in-block normalization is injective on the range of ϕ. Then ϕDAGR (· | s) is sufficient. The contraction condition holds on every run we instrument (Figure 5). The injectivity condition does not. Equations (7) and (8) normalize the output of each sub-block, and LayerNorm discards the mean and the scale of its input. Theorem 4.3 therefore certifies the gate and the attention but not the normalization we place after them. We flag the gap here and return to it in Section 6, where the same architectural detail turns out to be our leading explanation for the manipulation regression. Noise invariance carries over more cleanly. Every node of the MS-D GCA graph touches the goal observation only through ϕ and ψg , so their invariance propagates. Theorem 4.4 (Noise Invariance Preservation). Suppose goal observations decompose as og = (xg , ϵg ) with ϵg task-irrelevant, and suppose the base encoders satisfy ϕ(og ) = ϕ(xg ) and ψg (og ) = ψg (xg ) for all (xg , ϵg ). Then ϕDAGR (og | s) = ϕDAGR (xg | s) for every ϵg and every s ∈ S. The third property bounds the worst-case cost of the added expressiveness. Theorem 4.5 (Approximation Error Bound). Let f : Rds × Rd → R be LV -Lipschitz in its second argument, and suppose ∥ϕDAGR (g | s) − ϕ(g)∥2 ≤ B holds uniformly over (s, g) ∈ S × G. Define εbase := sup(s,g)∈S×G |V ∗ (s, g) − f (ψ(s), ϕ(g))| and V̂ (s, g) := f (ψ(s), ϕDAGR (g | s)). Then sup
|V ∗ (s, g) − V̂ (s, g)| ≤ εbase + LV · B.
(11)
(s,g)∈S×G
At initialization the gates are near zero, so B ≈ 0 and the bound reduces to that of the base. It grows only as the optimizer opens the gates. We claim no more than that the upper bound is gate-controlled, and Remark 4.8 shows why that is less than it sounds. The remaining two properties tie the value of refinement to Definition 3. Proposition 4.6 (Tighter Sample-Complexity Bound on Discrepancy-Structured Tasks). Suppose the task has discrepancy structure (10) with discrepancy map h, and suppose that the pair (ψ, ϕDAGR ) realizes h in the sense that there exists an Lh -Lipschitz h̃ : Rds × Rd → Z with h̃(ψ(s), ϕDAGR (g | s)) = h(s, g), while no Lh -Lipschitz function on (ψ(s), ϕ(g)) realizes h. Let π ∗ = ρ ◦ h with ρ Lipschitz. The standard Rademacher composition bound [Bartlett and Mendelson, 2002] then yields data-dependent upper bounds Unlate and UnDAGR on Rn of the policy class that satisfy UnDAGR ≤ Lh · Lρ · Rn (Hρ ),
Unlate ≤ Lψ−1 · Lϕ−1 · Lh · Lρ · Rn (Hρ ),
(12)
where Lψ−1 , Lϕ−1 ≥ 1 are the Lipschitz constants of any measurable selections inverting ψ and ϕ, and Hρ denotes the policy class realizing ρ. The DAGR bound is therefore tighter by a multiplicative factor Lψ−1 Lϕ−1 . The proof extends Proposition 4.2 from h(s, g) = s − g to arbitrary discrepancy maps, and the intuition is unchanged. DAGR exposes h at the representation level, while late fusion must recover it by inverting both encoders. The bound rests on a realizability assumption that we do not verify directly, and our gains are consistent with it but do not establish it. Appendix C states the assumption precisely and names the probe experiment that would settle it. Proposition 4.7 (Reduction to Base on Goal-Only Tasks). If the task is goal-only in the sense of Definition 3, that is, π ∗ (a | s, g) = ρ(h̄(g)) for some h̄ depending only on g, and if the base representation ϕ is sufficient (Definition 1) for π ∗ with ρ ◦ h̄ ∈ Flate , then the global minimum of any value or policy loss in FDAGR that achieves the base error εbase is attained at gate parameters α, λh → −∞, equivalently at B = 0 in Theorem 4.5 and at ϕDAGR (g | s) ≡ ϕ(g). The two propositions together mark the regime where DAGR formally helps and the regime where its best behavior is to reduce to the base. The next remark asks whether it does. Remark 4.8 (What the Bound Does and Does Not Explain). Read together, Proposition 4.7 and Theorem 4.5 predict that a goal-only task pays a penalty proportional to its converged gate. That prediction fails on our data, and we report the failure rather than force the fit. On Cube-Double the gates stay below 0.01 (Figure 6), so B and the certified surcharge LV · B are both small, yet the success rate falls by 25 points. Table 7 shows the mismatch from the other side. Opening the gates at initialization 8
Table 1: Success rates (%) on state-based OGBench with GCIVL. Mean ± std over 8 seeds. The Orig, VIB, VIP, TRA, BYOL-γ, and Dual columns are reproduced from Park et al. [2026], Table 1. DAGR = Dual + MS-D GCA (ours). Orange = best, underline = second best. Environment Orig VIB VIP TRA BYOL-γ Dual DAGR Navigation pointmaze-medium-navigate pointmaze-large-navigate antmaze-medium-navigate antmaze-large-navigate antmaze-giant-navigate humanoidmaze-medium-navigate humanoidmaze-large-navigate antsoccer-arena-navigate
78±8 52±6 71±4 16±3 0±0 27±3 3±0 47±4
69±13 50±7 68±4 9±3 0±0 24±2 3±1 34±4
0±1 0±0 31±5 9±2 0±0 7±3 1±0 2±1
3±6 1±2 22±15 22±12 0±0 21±3 2±1 8±2
37±7 22±12 39±5 11±5 0±0 18±5 2±1 11±4
76±7 46±6 75±4 28±11 0±0 29±3 3±2 31±3
87±8 41±7 95±1 82±3 4±2 83±3 62±4 58±4
Manipulation cube-single-play cube-double-play scene-play
52±3 35±5 46±3
90±3 33±3 58±1
40±7 3±2 23±6
40±5 7±2 46±6
51±11 6±4 44±9
89±3 60±4 72±6
87±2 35±2 59±2
Discrete Reasoning puzzle-3x3-play puzzle-4x4-play
5±1 14±1
14±3 6±3
3±1 1±1
5±1 10±3
0±0 1±2
5±1 23±3
5±1 14±3
Average
34±1
35±2
9±1
15±2
19±2
41±2
55
raises B by two orders of magnitude and costs only 8 points. A large B therefore hurts less than a small one, which no monotone reading of Theorem 4.5 allows. The bound still does its job, since it certifies that the refinement cannot blow up at any gate value we observe. It is simply not the mechanism behind the regression, and we do not present it as one. Section 6 names the mechanism.
5
Experiments
We evaluate DAGR on OGBench [Park et al., 2025] along three axes. First, whether state-conditioned refinement improves goal reaching across diverse tasks. Second, whether the difference bias contributes beyond standard cross-attention. Third, on which tasks it does not help, and why. Setup. We evaluate on 13 state-based tasks and 7 visual variants. The state-based tasks comprise navigation tasks (PointMaze, AntMaze, HumanoidMaze, and AntSoccer), manipulation tasks (CubeSingle, Cube-Double, and Scene), and discrete-reasoning tasks (Puzzle-3x3 and Puzzle-4x4). The base goal representation is Dual [Park et al., 2026] with bilinear (inner-product) parameterization, and the downstream offline GCRL algorithm is GCIVL [Park et al., 2025]. DAGR uses H = 4 heads with dk = 64, scales (T1 , T2 , T3 ) = (16, 8, 4), FFN hidden dimension 256, and one D GCA block per scale. All gate parameters and difference scalings are initialized to α0 = λ0 = −5, and fusion logits to zero. We use eight seeds for state-based tasks and four seeds for visual tasks. Full hyperparameters are in Table 5 of Appendix E. The Orig, VIB, VIP, TRA, BYOL-γ, and Dual numbers are reproduced from Park et al. [2026] under their own tuned settings. DAGR uses a single fixed configuration across all tasks rather than per-task tuning. The reported gains are therefore not attributable to task-specific hyperparameter search. 5.1
Main Results: State-Based Tasks
Table 1 reports the state-based results. The pattern is highly structured. On every navigation task except pointmaze-large, DAGR attains the best or tied-best score, with the largest absolute gains where Dual is weakest. These tasks fit Definition 3 with h(s, g) = πxy (s) − πxy (g), the position offset. Proposition 4.6 therefore gives a tighter sample-complexity upper bound on these tasks, consistent with the early-training separation we observe in Figure 8. On manipulation the picture is mixed. Cube-Single matches Dual. Cube-Double and Scene fall below it. Neither of the latter two fits Definition 3. On Cube-Double the optimal action turns on a binary choice of which cube to move first. On Scene it turns on a multi-step ordering decision. Neither 9
Table 2: Success rates (%) on visual OGBench with GCIVL. Mean ± std over 4 seeds. The Orig, VIB, VIP, TRA, BYOL-γ, and Dual columns are reproduced from Park et al. [2026], Table 2. The DAGR column is ours, and the bottom average row is our own aggregation. Orange = best, underline = second best. Environment
Orig
VIB
VIP
TRA
BYOL-γ
Dual
DAGR (Ours)
Navigation visual-antmaze-medium-navigate visual-antmaze-large-navigate
66±4 26±5
18±9 5±2
30±7 9±1
48±4 13±3
32±5 9±4
78±4 40±4
90±2 52±2
Manipulation visual-cube-single-play visual-cube-double-play visual-scene-play
53±4 9±2 25±2
18±19 0±0 6±3
39±6 0±0 4±1
31±24 3±2 15±6
35±8 2±1 10±8
58±5 9±2 26±5
44±2 11±4 27±3
Discrete Reasoning visual-puzzle-3x3-play visual-puzzle-4x4-play
22±2 65±4
0±0 0±0
0±0 0±0
0±0 0±0
0±0 0±0
0±0 0±0
0±0 0±0
Avg. (all 7 tasks) Avg. (excl. puzzle)
38 36
7 9
12 16
16 22
13 18
30 42
32 45
factors through a discrepancy map h(s, g), so Proposition 4.7 applies, and it predicts that the best DAGR can do on these tasks is to match Dual by closing its gates. The gates do stay nearly closed. Figure 6 shows σ(α) below 0.01 throughout. What Proposition 4.7 does not predict, and what we observe, is a 25-point drop. Remark 4.8 works through why the gate-controlled bound of Theorem 4.5 cannot supply the missing 25 points. We therefore report the regression as an open failure of our account rather than as a confirmation of it. Section 6 gives our leading hypothesis. On puzzle the gap to Dual is essentially zero, because no late-fusion representation method we evaluate reaches a non-trivial success rate at all (Appendix F.4). Full training curves are in Figure 8. 5.2
Main Results: Visual Tasks
Table 2 reproduces the state-based pattern on navigation. DAGR is highest on both Visual-AntMaze variants, so the directional signal survives the pixel encoder. Visual manipulation splits. Cube-Double and Scene match Dual, whereas Cube-Single falls below it, and Section 5.3 identifies this as the one task on which state-conditioning helps but our attention rule does not. Visual-Puzzle remains at zero for every late-fusion method, an encoder-level bottleneck we examine in Appendix F.4. Curves are in Figure 9. 5.3
Ablation: Standard Cross-Attention vs. D GCA
Table 3 compares Dual, Dual with a standard cross-attention module (+CA, single-scale, gate at σ(0) = 0.5, no difference bias), and Dual with the full D GCA. On navigation, D GCA clearly beats +CA, and the gap is widest where Dual is weakest. The two differ in two ways at once, since D GCA adds the difference bias and it closes the gate at initialization. Table 6 separates them. Removing the difference bias from D GCA leaves AntMaze-Large at 84.4, which is not below the full model at 82.5. Removing the gated residual drops it to 59.3. The gate, not the bias, is what +CA is missing. We state this plainly. The mechanism that names our method is not the mechanism that produces our navigation gains. Visual-Cube-Single is the sharpest counterexample in the paper, and we treat it as one. +CA reaches 80 there against 58 for Dual and 44 for D GCA. Two readings follow, and they point in opposite directions. State-conditioning helps a great deal on this task, and the +22 points +CA gains over Dual is the largest single-task gain any module produces anywhere in this paper, so the failure is not on the state-conditional axis. Yet the two ingredients that separate D GCA from +CA are jointly harmful here, and together they cost 36 points. We cannot separate their contributions with the ablations we have. One untested account is that single-object visual goal reaching is a template-matching problem, for which similarity-maximizing attention is the correct bias and a discrepancy-maximizing one is not. On Cube-Double and Scene both variants fall below Dual, consistent with Section 5.1. 10
Table 3: Standard cross-attention vs. D GCA, both applied to Dual. Success rate (%). Dual numbers from Park et al. [2026]. Environment Dual +CA +D GCA
5.4
State-Based antmaze-medium antmaze-large humanoidmaze-medium humanoidmaze-large antsoccer-arena cube-double scene
75±4 28±11 29±3 3±2 31±3 60±4 72±6
84±6 33±4 41±7 8±2 44±2 26±5 57±1
95±1 82±3 83±3 62±4 58±4 35±2 59±2
Visual visual-antmaze-medium visual-antmaze-large visual-cube-single visual-cube-double visual-scene visual-puzzle-3x3
78±4 40±4 58±5 9±2 26±5 0±0
80±14 41±9 80±3 10±2 13±2 0±0
90±2 52±2 44±2 11±4 27±3 0±0
Component Attribution and Scope
Table 6 of Appendix F isolates each component of MS-D GCA. Three results matter here. Removing the gated residual drops AntMaze-Large from 82.5 to 59.3, which confirms that the near-identity initialization protects the base temporal-distance structure. Removing the difference bias leaves it at 84.4, within one standard deviation of the full model. Removing the FFN drops it to 57.7. The gate and the per-token projection therefore carry the navigation gain, and the difference bias does not. Table 9 of Appendix F explains why. The learned ζ(λ) remains at its initialization value on five of six tasks, so the bias never activates at convergence. Figure 5 shows the same for the gates, which stay below 0.01 throughout. DAGR therefore operates as a small, gated perturbation of ϕ(g), and the difference bias supplies an inductive prior at initialization rather than a converged contribution. Table 7 confirms the protective reading of the gate. Initializing α0 = 0 rather than −5 opens the gates from the first gradient step and costs 8 points on Cube-Double. Table 8 of Appendix F delimits the scope. Every late-fusion representation method, ours included, scores zero on Visual-Puzzle, whereas vanilla GCIVL under early fusion does not. The bottleneck lies in the encoder rather than in ϕ, since the IMPALA CNN pools away the pixel-level correspondence between the state and goal images before ϕ(g) is computed. A post-encoder module can only reweight what the encoder preserves.
6
Conclusion and Discussion
DAGR refines a static goal embedding into a gated, multi-scale, state-conditioned one. It preserves sufficiency and noise invariance, tightens the sample-complexity bound on discrepancy-structured tasks, and improves goal reaching on every navigation task that Definition 3 covers. The ablations attribute that improvement to the gated residual rather than to the difference bias the method is named for. The regression on Cube-Double and Scene is architectural. Equations (7) and (8) normalize each subblock output, so a closed gate discards the norm of the base embedding, on which the bilinear temporal distance of Dual depends. Navigation reads direction alone and is unaffected, whereas manipulation is not. Table 6 supports this, since removing the normalization recovers Cube-Double while leaving AntMaze-Large within noise. Pre-normalizing each sub-block would restore exact identity and close the injectivity gap in Theorem 4.3. Two limitations remain. The token decomposition is not object-aware, which plausibly compounds the manipulation deficit, and the Visual-Puzzle failure lies in the encoder rather than the goal representation, where hierarchical planning [Park et al., 2023] is complementary. 11
References Hongjoon Ahn, Heewoong Choi, Jisu Han, and Taesup Moon. Option-aware temporally abstracted value for offline goal-conditioned reinforcement learning. arXiv preprint arXiv:2505.12737, 2025. Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. Flamingo: A visual language model for few-shot learning. In Neural Information Processing Systems (NeurIPS), 2022. Alexander A Alemi, Ian Fischer, Joshua V Dillon, and Kevin Murphy. Deep variational information bottleneck. In International Conference on Learning Representations (ICLR), 2017. Erdemt Bao, Xing Lei, and Jun Chen. Nftr: From provable mode-averaging to geodesic subgoal selection in offline goal-conditioned rl. arXiv preprint arXiv:2607.07855, 2026. Peter L Bartlett and Shahar Mendelson. Rademacher and gaussian complexities: Risk bounds and structural results. Journal of machine learning research, 3(Nov):463–482, 2002. Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, et al. π0 : A vision-language-action flow model for general robot control. arXiv preprint arXiv:2410.24164, 2024. Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Xi Chen, Krzysztof Choromanski, Tianli Ding, Danny Driess, Avinava Dubey, Chelsea Finn, et al. Rt-2: Vision-language-action models transfer web knowledge to robotic control. arXiv preprint arXiv:2307.15818, 2023a. Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Joseph Dabis, Chelsea Finn, Keerthana Gober, Karol Hausman, Alex Herzog, Jasmine Hsu, et al. Rt-1: Robotics transformer for real-world control at scale. In Robotics: Science and Systems (RSS), 2023b. Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European Conference on Computer Vision (ECCV), 2020. Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Misha Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch. Decision transformer: Reinforcement learning via sequence modeling. In Neural Information Processing Systems (NeurIPS), 2021. Thomas M Cover. Elements of information theory. John Wiley & Sons, 1999. Benjamin Eysenbach, Ruslan Salakhutdinov, and Sergey Levine. Search on the replay buffer: Bridging planning and reinforcement learning. In Neural Information Processing Systems (NeurIPS), 2019. Benjamin Eysenbach, Tianjun Zhang, Ruslan Salakhutdinov, and Sergey Levine. Contrastive learning as goal-conditioned reinforcement learning. In Neural Information Processing Systems (NeurIPS), 2022. Kuan Fang, Jonathan Tompson, Nicholas Rhinehart, and Sergey Levine. Planning goals for exploration. In International Conference on Learning Representations (ICLR), 2022. Dibya Ghosh, Abhishek Gupta, Ashwin Reddy, Justin Fu, Coline Manon Devin, Benjamin Eysenbach, and Sergey Levine. Learning to reach goals via iterated supervised learning. In International Conference on Learning Representations, 2021. Vittorio Giammarino and Ahmed H Qureshi. Goal reaching with eikonal-constrained hierarchical quasimetric reinforcement learning. In The Fourteenth International Conference on Learning Representations, 2026. Felix Hill, Andrew Lampinen, Rosalia Schneider, Stephen Clark, Matthew Botvinick, James L McClelland, and Adam Santoro. Environmental drivers of systematicity and generalization in a situated agent. In International Conference on Learning Representations (ICLR), 2020. Noriaki Hirose, Dhruv Shah, Sanjay Sridhar, and Sergey Levine. Lelan: Learning latent navigation maps. In International Conference on Robotics and Automation (ICRA), 2023. 12
Jie Huang et al. Early fusion of vision and language for robot manipulation. arXiv preprint, 2025. Andrew Jaegle, Felix Gimeno, Andy Brock, Andrew Zisserman, Oriol Vinyals, and Joao Carreira. Perceiver: General perception with iterative attention. In International Conference on Machine Learning (ICML), 2021. Vineet Jain and Siamak Ravanbakhsh. Learning to reach goals via diffusion. In International Conference on Machine Learning, pages 21170–21195. PMLR, 2024. Eric Jang, Alex Irpan, Mohi Khansari, Daniel Kappler, Frederik Ebert, Corey Lynch, Sergey Levine, and Chelsea Finn. Bc-z: Zero-shot task generalization with robotic imitation learning. In Conference on Robot Learning (CoRL), 2022. Leslie Pack Kaelbling. Learning to achieve goals. In International Joint Conference on Artificial Intelligence (IJCAI), 1993. Dmitry Kalashnikov, Alex Irpan, Peter Pastor, Julian Ibarz, Alexander Herzog, Eric Jang, Deirdre Quillen, Ethan Holly, Mrinal Kalakrishnan, Vincent Vanhoucke, and Sergey Levine. Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation. In Conference on Robot Learning (CoRL), 2018. Junseok Kim, Dohyeong Kim, Mineui Hong, and Songhwai Oh. Compositional transduction with latent analogies for offline goal-conditioned reinforcement learning. arXiv preprint arXiv:2605.20609, 2026. Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, et al. Openvla: An open-source vision-language-action model. arXiv preprint arXiv:2406.09246, 2024. Ilya Kostrikov, Ashvin Nair, and Sergey Levine. Offline reinforcement learning with implicit q-learning. In International Conference on Learning Representations (ICLR), 2022. Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine. Conservative q-learning for offline reinforcement learning. In Neural Information Processing Systems (NeurIPS), 2020. Sascha Lange, Thomas Gabel, and Martin Riedmiller. Batch reinforcement learning. Reinforcement learning: State-of-the-art, pages 45–73, 2012. Dieter Lawson et al. Self-predictive goal representations with byol-γ for offline goal-conditioned rl. In International Conference on Learning Representations (ICLR), 2025. Xing Lei, Wenyan Yang, Kaiqiang Ke, Shentao Yang, Xuetao Zhang, Joni Pajarinen, and Donglin Wang. Gchr: Goal-conditioned hindsight regularization for sample-efficient reinforcement learning. arXiv preprint arXiv:2508.06108, 2025. Xing Lei, Jincheng Wang, Xuetao Zhang, and Donglin Wang. QHyer: Q-conditioned hybrid attentionmamba transformer for offline goal-conditioned RL. In Forty-third International Conference on Machine Learning, 2026. Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu. Offline reinforcement learning: Tutorial, review, and perspectives on open problems. arXiv preprint arXiv:2005.01643, 2020. Minghuan Liu, Menghui Zhu, and Weinan Zhang. Goal-conditioned reinforcement learning: Problems and solutions. arXiv preprint arXiv:2201.08299, 2022. Corey Lynch, Mohi Khansari, Ted Xiao, Vikash Kumar, Jonathan Tompson, Sergey Levine, and Pierre Sermanet. Learning latent plans from play. In Conference on Robot Learning (CoRL), 2020. Yecheng Jason Ma, Jason Yan, Dinesh Jayaraman, and Osbert Bastani. How far i’ll go: Offline goalconditioned reinforcement learning via f -advantage regression. arXiv preprint arXiv:2206.03023, 2022a. Yecheng Jason Ma, Jason Yan, Dinesh Jayaraman, and Osbert Bastani. Offline goal-conditioned reinforcement learning via f -advantage regression. In Neural Information Processing Systems (NeurIPS), 2022b. 13
Yecheng Jason Ma, Shagun Sodhani, Dinesh Jayaraman, Osbert Bastani, Vikash Kumar, and Amy Zhang. Vip: Towards universal visual reward and representation via value-implicit pre-training. In International Conference on Learning Representations (ICLR), 2023. Vivek Myers, Chongyi Zheng, Anca Dragan, Sergey Levine, and Benjamin Eysenbach. Learning temporal distances: Contrastive successor features can provide a metric structure for decisionmaking. In International Conference on Machine Learning (ICML), 2024. Vivek Myers, Bill Zheng, Benjamin Eysenbach, and Sergey Levine. Offline goal-conditioned reinforcement learning with quasimetric representations. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025a. Vivek Myers, Chongyi Zheng, Anca Dragan, Sergey Levine, and Benjamin Eysenbach. Temporal representation alignment for offline goal-conditioned reinforcement learning. In International Conference on Learning Representations (ICLR), 2025b. Ashvin V Nair, Vitchyr Pong, Murtaza Dalal, Shikhar Bahl, Steven Lin, and Sergey Levine. Visual reinforcement learning with imagined goals. In Neural Information Processing Systems (NeurIPS), 2018. Seohong Park, Dibya Ghosh, Benjamin Eysenbach, and Sergey Levine. Hiql: Offline goal-conditioned rl with latent states as actions. In Neural Information Processing Systems (NeurIPS), 2023. Seohong Park, Tobias Kreiman, and Sergey Levine. Foundation policies with hilbert representations. In International Conference on Machine Learning (ICML), 2024. Seohong Park, Kevin Frans, Benjamin Eysenbach, and Sergey Levine. Ogbench: Benchmarking offline goal-conditioned rl. In International Conference on Learning Representations (ICLR), 2025. Seohong Park, Deepinder Mann, and Sergey Levine. Dual goal representations. In The Fourteenth International Conference on Learning Representations, 2026. Ethan Perez, Florian Strub, Harm De Vries, Vincent Dumoulin, and Aaron Courville. Film: Visual reasoning with a general conditioning layer. In AAAI Conference on Artificial Intelligence, 2018. Silviu Pitis, Harris Chan, Kiarash Jamali, and Jimmy Ba. An inductive bias for distances: Neural nets that respect the triangle inequality. In International Conference on Learning Representations, 2020. Matthias Plappert, Marcin Andrychowicz, Alex Ray, Bob McGrew, Bowen Baker, Glenn Powell, Jonas Schneider, Josh Tobin, Maciek Chociej, Peter Welinder, et al. Multi-goal reinforcement learning: Challenging robotics environments and request for research. 2018. Zihan Qiu, Zekun Wang, Bo Zheng, Zeyu Huang, Kaiyue Wen, Songlin Yang, Rui Men, Le Yu, Fei Huang, Suozhi Huang, et al. Gated attention for large language models: Non-linearity, sparsity, and attention-sink-free. Advances in Neural Information Processing Systems, 38:100092–100118, 2026. Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Gomez Colmenarejo, Alexander Novikov, Gabriel Barth-Maron, Mai Gimenez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, et al. A generalist agent. Transactions on Machine Learning Research, 2022. Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver. Universal value function approximators. In International Conference on Machine Learning (ICML), 2015. Max Schwarzer, Ankesh Anand, Rishab Goel, R Devon Hjelm, Aaron Courville, and Philip Bachman. Data-efficient reinforcement learning with self-predictive representations. In International Conference on Learning Representations (ICLR), 2021. Dhruv Shah, Benjamin Eysenbach, Gregory Kahn, Nicholas Rhinehart, and Sergey Levine. Ving: Learning open-world navigation with visual goals. In International Conference on Robotics and Automation (ICRA), 2021a. 14
Dhruv Shah, Benjamin Eysenbach, Nicholas Rhinehart, and Sergey Levine. Relmogen: Leveraging motion generation in reinforcement learning for mobile manipulation. In International Conference on Robotics and Automation (ICRA), 2021b. Harshit Sikchi, Qinqing Zheng, Amy Zhang, and Scott Niekum. Dual rl: Unification and new methods for reinforcement and imitation learning. In International Conference on Learning Representations (ICLR), 2024. Octo Model Team, Dibya Ghosh, Homer Walke, Karl Pertsch, Kevin Black, Oier Mees, Sudeep Dasari, Joey Hejna, Tobias Xu, Jianlan Luo, et al. Octo: An open-source generalist robot policy. arXiv preprint arXiv:2405.12213, 2024. Naftali Tishby, Fernando C Pereira, and William Bialek. The information bottleneck method. arXiv preprint physics/0004057, 2000. Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Neural Information Processing Systems (NeurIPS), 2017. Lirui Wang, Xinlei Zhao, Jialiang Liu, Herke van Hoof, and Pieter Abbeel. Scaling proprioceptivevisual learning with heterogeneous pre-trained transformers. In Neural Information Processing Systems (NeurIPS), 2024a. Mianchu Wang, Keiran Paster, Jimmy Ba, and Sheila Agrawal. Go-plan: Goal-conditioned offline reinforcement learning by planning with learned models. In International Conference on Machine Learning (ICML), 2024b. Tongzhou Wang and Phillip Isola. Improved representation of asymmetrical distances with interval quasimetric embeddings. In NeurIPS 2022 Workshop on Symmetry and Geometry in Neural Representations, 2022a. Tongzhou Wang and Phillip Isola. On the learning and learnability of quasimetrics. In International Conference on Learning Representations, 2022b. Tongzhou Wang, Antonio Torralba, Phillip Isola, and Amy Zhang. Optimal goal-reaching reinforcement learning via quasimetric learning. In International Conference on Machine Learning (ICML), 2023. Qiushui Xu, Yu-Hao Huang, Yushu Jiang, Wenliang Zheng, Lei Song, Jinyu Wang, and Jiang Bian. In-context compositional q-learning for offline reinforcement learning. In The Fourteenth International Conference on Learning Representations, 2026. Ge Yang, Zhang-Wei Hong, and Pulkit Agrawal. Bi-linear value networks for multi-goal reinforcement learning. In International Conference on Learning Representations, 2022a. Rui Yang, Yiming Lu, Wenzhe Li, Hao Sun, Meng Fang, Yali Du, Xiu Li, Lei Han, and Chongjie Zhang. Rethinking goal-conditioned supervised learning and its connection to offline rl. In International Conference on Learning Representations (ICLR), 2022b. Bill Zheng, Vivek Myers, Benjamin Eysenbach, and Sergey Levine. Scaling goal-conditioned reinforcement learning with multistep quasimetric distances. In The Fourteenth International Conference on Learning Representations, 2026. Chongyi Zheng, Benjamin Eysenbach, Homer Rich Walke, Patrick Yin, Kuan Fang, Ruslan Salakhutdinov, and Sergey Levine. Stabilizing contrastive RL: Techniques for robotic goal reaching from offline data. In The Twelfth International Conference on Learning Representations, 2024a. Chongyi Zheng, Benjamin Eysenbach, Homer Walters, Ruslan Salakhutdinov, and Sergey Levine. Contrastive difference predictive coding. In International Conference on Learning Representations (ICLR), 2024b. Linqing Zhong, Yi Liu, Yifei Wei, Ziyu Xiong, Maoqing Yao, Si Liu, and Guanghui Ren. Acot-vla: Action chain-of-thought for vision-language-action models. arXiv preprint arXiv:2601.11404, 2026. John L Zhou and Jonathan C Kao. Flattening hierarchies with policy bootstrapping. arXiv preprint arXiv:2505.14975, 2025. 15
Contents of Appendix A Training Procedure
17
B Notation
18
C On the Realizability Assumption of Proposition 4.6
19
D Theoretical Proofs
19
D.1 Proof of Proposition 4.1 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
19
D.2 Proof of Proposition 4.2 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
19
D.3 Proof of Theorem 4.3 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
20
D.4 Proof of Theorem 4.4 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
20
D.5 Proof of Theorem 4.5 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
21
D.6 Proof of Proposition 4.6 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
21
D.7 Proof of Proposition 4.7 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
22
E Experimental Details
22
E.1 Environment Details . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
22
E.2 Hyperparameters . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
22
E.3 Network Architecture Details . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
22
E.4 Computational Resources and Overhead Measurements . . . . . . . . . . . . . . .
24
F Additional Experimental Results
24
F.1
Component Ablation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
24
F.2
Sensitivity Analyses . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
24
F.3
Where the Module Looks: Attention and Gate Analysis . . . . . . . . . . . . . . .
24
F.4
Why Visual-Puzzle Remains at Zero . . . . . . . . . . . . . . . . . . . . . . . . .
24
F.5
Computational Overhead . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
25
F.6
Per-Task Learned ζ(λ) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
25
F.7
Per-Task Fusion Weights . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
26
F.8
Gate and λ Evolution on Additional Tasks . . . . . . . . . . . . . . . . . . . . . .
26
F.9
Learning Curves . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
26
16
A
Training Procedure
DAGR is trained jointly with the downstream offline GCRL algorithm. The base representation loss Lrep produces ϕ(g) without cross-attention, exactly as in the underlying Dual method, so the temporal-distance structure of ϕ relied on by Theorems 4.3 and 4.4 is not perturbed. Only the value, Q, and policy losses propagate gradients through MS-D GCA. Its parameters comprise the per-scale to(ℓ) (ℓ) Q,(ℓ) K,(ℓ) V,(ℓ) ken projections {Ws , Wg } (Equation (3)), attention matrices {Wh , Wh , Wh , W O,(ℓ) } (ℓ) (ℓ) (ℓ) (Equation (6)), gate vectors {αattn , αffn } and per-head difference scalars {λh } (Equations (5), L (7) and (8)), per-scale FFN parameters, and fusion logits w ∈ R (Equation (9)). Each call MS-D GCA(ϕ(g), ψ(s), ψg (g)) proceeds in four steps. Step (i) projects ψ(s), ψg (g) into pseudotokens via Equation (3). Step (ii) computes the difference map via Equation (4). Step (iii) runs multi-head D GCA and the gated residual block via Equations (6) to (8) at each of L scales. Step (iv) fuses the scale outputs via Equation (9). The near-zero gate initialization ensures ϕ̃ ≈ ϕ(g) at step zero. Gates open only when doing so reduces the downstream losses, which realizes the late-fusion-to-refinement curriculum discussed in Section 4.2 and bounded by Theorem 4.5. Algorithm 1 DAGR layered on top of the Dual goal representation [Park et al., 2026] with GCIVL [Kostrikov et al., 2022, Park et al., 2025] as the downstream offline GCRL algorithm. Lines highlighted in orange mark the differences from a plain Dual + GCIVL baseline. Require: Dataset D, encoders ϕ, ψ, ψg , MS-D GCA module, value VηV , Q-network QηQ , policy πθ , AWR temperature β, expectile τ , discount γ, target rate τtgt . (ℓ) (ℓ) (ℓ) 1: Initialize αattn , αffn ← −5 (Equations (7) and (8)), λh ← −5 (Equation (5)), w ← 0 (Equation (9)). 2: for each gradient step do 3: Sample (s, a, s′ ) ∼ D and future goal g (goal-sampling distribution in Table 5). 4: Update ϕ via Lrep on ϕ(g) alone, with MS-D GCA not in the graph. 5: ϕ̃ ← MS-D GCA(ϕ(g), ψ(s), ψg (g)) and ϕ̃′ ← MS-D GCA(ϕ(g), ψ(s′ ), ψg (g)) // ϕ̃ ≈ ϕDAGR (g | s) 6: yV ← r(s, g) + γ V̄ (ψ(s′ ), ϕ̃′ ) using EMA target V̄ . 7: Update VηV on LV = Lτ Q(ψ(s), ϕ̃, a) − V (ψ(s), ϕ̃) with Lτ (u) = |τ − 1[u < 0]| u2 [Kostrikov et al., 2022]. 2 8: Update QηQ on LQ = Q(ψ(s), ϕ̃, a) − yV . 9: A ← Q(ψ(s), ϕ̃, a) − V (ψ(s), ϕ̃) (stop-gradient). 10: Update πθ on Lπ = − exp(A/β) log πθ (a | ψ(s), ϕ̃). 11: Soft-update target: V̄ ← (1 − τtgt )V̄ + τtgt VηV . 12: end for 13: return πθ , ϕ, ψ, ψg , MS-D GCA.
17
B
Notation
Symbol
Table 4: Notation used in the proofs and not introduced in the main text. Description
General notation ϕ(g | s) φ∨ (g), φ∨ DAGR (g | s) d∗ (s, g) ∆s,g (·)
Deterministic state-parameterized goal encoding; the bar denotes parameter dependence, not a conditional distribution. Equivalently ϕ(g; s) or ϕs (g). Dual and DAGR goal functionals from S to R (Equations (1) and (2)); the latter is ∆s,g -weighted version of the former. Optimal temporal distance logγ V ∗ (s, g) (Section 3); equals shortest-path length in deterministic environments. State-dependent weighting S → [0, 1] in the DAGR functional (Equation (2)); approximated by the per-token discrepancy map ∆t of Equation (4).
Discrepancy structure (Definition 3) h(s, g), ρ Discrepancy map S × S → Z and corresponding policy factor Z → ∆(A) such that π ∗ (a | s, g) = ρ(h(s, g)). Z Latent discrepancy space; concrete instances include Rk (e.g., h(s, g) = πxy (s) − πxy (g) for maze navigation). “goal-only” Task satisfying Equation (10) with h a function of g alone, that is, π ∗ (a | s, g) = ρ(h̄(g)) for some h̄. Information-theoretic quantities (Propositions 4.1, 4.2 and 4.6) S, G, A∗ Random variables for state, goal, and optimal action. I(X; Y | Z) Conditional mutual information between X and Y given Z. X→Y →Z Markov chain, equivalently X ⊥⊥ Z | Y . Rn (·) Rademacher complexity of a function class on n samples; DAGR Rlate denote the complexities under late-fusion and n and Rn DAGR inputs respectively (Equation (12)). Lg Lipschitz constant of a function g. Lψ−1 , Lϕ−1 Lipschitz constants of implicit encoder inverses required by late-fusion to recover (s, g) (Proposition 4.6). Sufficiency and noise invariance (Theorems 4.3 and 4.4) Fℓ , δℓ Per-scale block map Fℓ (ϕ) = ϕ + δℓ (ϕ, s) and its residual perturbation. cℓ Contraction constant of scale ℓ, cℓ = ∥σ(α)∥∞ · Lℓ < 1. Recovers Inverse of the map ϕ(g) 7→ ϕDAGR (g | s) on its image. Goal observation split into a task-relevant component xg and og = (xg , ϵg ) an exogenous noise component ϵg ∈ E. Approximation and conditional improvement (Theorem 4.5 and Propositions 4.6 and 4.7) V̂ (s, g) Approximated value V̂ (s, g) = f (ψ(s), ϕDAGR (g | s)). Base approximation error, perturbation radius ∥ϕDAGR (g | εbase , B, LV s) − ϕ(g)∥2 ≤ B, and Lipschitz constant of f in its second argument. Flate , FDAGR Function classes induced by late-fusion and DAGR encoders, respectively. The realizability inclusion Flate ⊆ FDAGR (proven inline in Appendix D.7) underlies the global-optimum analysis of Proposition 4.7.
18
C
On the Realizability Assumption of Proposition 4.6
Proposition 4.6 assumes a realizability gap. There exists an Lh -Lipschitz h̃ with h̃(ψ(s), ϕDAGR (g | s)) = h(s, g), and no Lh -Lipschitz function on (ψ(s), ϕ(g)) realizes h. We do not verify this, and we are explicit about what our experiments can and cannot say. The gains of Section 5.1 are consistent with the assumption. They are equally consistent with two alternatives that we cannot rule out, namely the added capacity of the module and a better-conditioned optimization path. Proposition 4.6 therefore motivates the architecture rather than explaining the result, and the ablation of Appendix F.1 is what constrains which part of the architecture does the work. The assumption is directly testable. Freeze the trained encoders, sample state-goal pairs with groundtruth h(s, g) = πxy (s) − πxy (g), and fit two probes of matched capacity and matched spectral-norm budget, one on (ψ(s), ϕ(g)) and one on (ψ(s), ϕDAGR (g | s)). A lower held-out error for the second at equal Lipschitz budget is exactly the gap the proposition assumes. We identify this as the most direct open experiment our theory calls for.
D
Theoretical Proofs
This appendix contains complete proofs for all results stated in Section 4. We restate each result before proving it. All notation follows Table 4. D.1
Proof of Proposition 4.1
Proposition D.1 (Restated). For any state-independent ϕ, I(A∗ ; ϕ(G) | S) ≤ I(A∗ ; G | S), with equality if and only if ϕ(G) retains all goal information relevant to A∗ given S. Proof. Because ϕ does not depend on S, conditioned on S = s the optimal action A∗ is determined by π ∗ (· | s, g), which under the late-fusion architecture depends on g only through ϕ(g). This gives the Markov chain G → ϕ(G) → A∗ conditioned on S = s. The data processing inequality [Cover, 1999] states that for any Markov chain X → Y → Z, I(X; Z) ≤ I(X; Y ). Applying this with X = A∗ , Y = G, Z = ϕ(G) conditioned on S = s gives I(A∗ ; ϕ(G) | S = s) ≤ I(A∗ ; G | S = s). Taking expectation over S yields I(A∗ ; ϕ(G) | S) ≤ I(A∗ ; G | S). Equality holds if and only if A∗ ⊥⊥ G | (ϕ(G), S), that is, ϕ(G) is a sufficient statistic for G with respect to A∗ conditional on S. For state-independent ϕ, the sufficient-statistic condition fails whenever the optimal action A∗ depends on a joint property of (S, G) that ϕ cannot resolve without access to S, which yields strict inequality in the data processing bound. D.2
Proof of Proposition 4.2
Proposition D.2 (Restated). Suppose π ∗ (a | s, g) = f (s − g). Then realizing π ∗ from the late-fusion inputs (ψ(s), ϕ(g)) requires composing f with measurable inverse selections of ψ and of ϕ. The composition property of Rademacher complexity [Bartlett and Mendelson, 2002] then yields an upper bound on the policy class that exceeds the corresponding bound under direct access to s − g by a factor Lψ−1 Lϕ−1 ≥ 1. Proof. Under late fusion, πθ (a | s, g) = hθ (ψ(s), ϕ(g)) with fixed encoders ψ and ϕ. To realize π ∗ = f ◦subtract from these inputs, hθ must invert ψ to recover s, invert ϕ to recover g, compute s−g, and apply f . The composition property of Rademacher complexity [Bartlett and Mendelson, 2002] gives Rn (g1 ◦ g2 ) ≤ Lg2 · Rn (g1 ) for Lg2 -Lipschitz g2 . Adding the inversion layers introduces extra Lipschitz factors Lψ−1 , Lϕ−1 in front of Rn (f ), increasing the bound on the effective complexity by a multiplicative factor relative to the case where s − g is directly available. A representation that exposes s − g to the policy therefore admits a strictly smaller upper bound on sample complexity for realizing the same π ∗ . 19
D.3
Proof of Theorem 4.3
We first establish a key lemma. Lemma D.3 (Injectivity of the Gated Residual Mapping). Under the assumptions of Theorem 4.3, for any fixed s ∈ S and any scale ℓ, the map ϕ(g) 7→ D GCA-BlockTℓ (ϕ(g), ψ(s), ψg (g)) is injective. Consequently, ϕ(g) is recoverable from the pair (s, ϕDAGR (g | s)). Proof. From Equations (7) and (8), the block output factors as Fℓ (ϕ) := D GCA-BlockTℓ (ϕ, ψ(s), ψg (g)) = ϕ + δℓ (ϕ, s), where δℓ collects the contributions of the cross-attention, the two gated residuals, the layer normalizations, and the FFN, modulated by σ(αattn ) and σ(αffn ). The assumed Lipschitz constant Lℓ of δℓ in ϕ, together with ∥σ(α)∥∞ · Lℓ < 1, implies ∥δℓ (ϕ1 , s) − δℓ (ϕ2 , s)∥ ≤ cℓ ∥ϕ1 − ϕ2 ∥ with cℓ < 1. Hence ∥Fℓ (ϕ1 ) − Fℓ (ϕ2 )∥ ≥ ∥ϕ1 − ϕ2 ∥ − ∥δℓ (ϕ1 , s) − δℓ (ϕ2 , s)∥ ≥ (1 − cℓ ) ∥ϕ1 − ϕ2 ∥ > 0 whenever ϕ1 ̸= ϕ2 . Therefore Fℓ is injective. P For the multi-scale composite (Equation (9)), ϕDAGR (g | s) = ℓ softmax(w)ℓ · Fℓ (ϕ(g)). Each Fℓ (ϕ) = ϕ + δℓ (ϕ, s) is a contraction perturbation of the identity with constant cℓ < 1. A direct calculation gives ⟨Fℓ (ϕ1 ) − Fℓ (ϕ2 ), ϕ1 − ϕ2 ⟩ = ∥ϕ1 − ϕ2 ∥2 + ⟨δℓ (ϕ1 , s) − δℓ (ϕ2 , s), ϕ1 − ϕ2 ⟩ ≥ (1 − cℓ )∥ϕ1 − ϕ2 ∥2 , so Fℓ (ϕ1 ) − Fℓ (ϕ2 ) has a strictly positive inner product with ϕ1 − ϕ2 whenever ϕ1 ̸= ϕ2 , that is, the two vectors lie in a common open half-space. A convex combination of such vectors with strictly positive softmax weights also has strictly positive inner product with ϕ1 − ϕ2 , hence cannot vanish, so ϕDAGR (g1 | s) = ϕDAGR (g2 | s) forces ϕ(g1 ) = ϕ(g2 ). Given (s, ϕDAGR (g | s)), ϕ(g) can be recovered by the contraction-mapping iteration X ϕ(k+1) = ϕDAGR (g | s) − softmax(w)ℓ · δℓ (ϕ(k) , s), ℓ
which converges by the Banach fixed-point theorem since maxℓ cℓ < 1. Theorem D.4 (Restated). Under the assumptions of Theorem 4.3, ϕDAGR (· | s) is sufficient: there exists π̃ such that V π̃(·|s,ϕDAGR (g|s)) (s, g) = V ∗ (s, g). ∗
Proof. By the assumed sufficiency of ϕ, there exists π ∗ : S × Rd → ∆(A) with V π (·|s,ϕ(g)) (s, g) = V ∗ (s, g). By Lemma D.3, the map ϕ(g) 7→ ϕDAGR (g | s) is invertible on its image with inverse Recovers . Define π̃(a | s, z) := π ∗ (a | s, Recovers (z)). Then for any trajectory generated by π̃(· | st , ϕDAGR (g | st )), at each step π̃(· | st , ϕDAGR (g | st )) = π ∗ (· | st , Recoverst (ϕDAGR (g | st ))) = π ∗ (· | st , ϕ(g)). The value attained by π̃ therefore matches V ∗ (s, g). D.4
Proof of Theorem 4.4
Theorem D.5 (Restated). Under encoder noise invariance (ϕ(o) = ϕ(x), ψ(o) = ψ(x), ψg (o) = ψg (x)), ϕDAGR (og | s) = ϕDAGR (xg | s) for every ϵg ∈ E. Proof. Write og = (xg , ϵg ). The base encoder ϕ satisfies ϕ(og ) = ϕ(xg ) by assumption. We note that for the Dual representation [Park et al., 2026] this property follows from [Park et al., 2026, Theorem 3.2], so when DAGR is layered on top of Dual the assumption is satisfied without additional work. We now trace the MS-D GCA computation graph and verify that every intermediate quantity is independent of ϵg . The encoder outputs satisfy ϕ(og ) = ϕ(xg ) and ψg (og ) = ψg (xg ) by assumption, while ψ(s) is independent of og . (ℓ)
(ℓ)
The difference map ∆(ℓ) is a fixed function of stok and gtok , both already ϵg -independent. 20
(ℓ)
The queries Qh = ϕ(og )WhQ = ϕ(xg )WhQ are ϵg -independent. The keys Kh = stok WhK and values √ (ℓ) Vh = stok WhV depend only on ψ(s). Therefore the attention logits Qh Kh⊤ / dk + ζ(λh )∆(ℓ) are ϵg -independent, and so is attn_out. The gated residual in Equation (7) computes ϕ(og ) + σ(αattn ) ⊙ attn_out = ϕ(xg ) + σ(αattn ) ⊙ attn_out, again ϵg -independent. The layer normalization and the FFN are deterministic functions of their inputs, so the independence propagates. The second gated residual is analogous. Each levelℓ is therefore ϵg -independent. The fusion weights softmax(w) are parameters, hence independent of og . Therefore ϕDAGR (og | s) =
L X
softmax(w)ℓ · levelℓ = ϕDAGR (xg | s).
ℓ=1
D.5
Proof of Theorem 4.5
Theorem D.6 (Restated). Let f : Rds × Rd → R be LV -Lipschitz in its second argument, and suppose ∥ϕDAGR (g | s) − ϕ(g)∥2 ≤ B holds uniformly over (s, g) ∈ S × G. Define εbase := sup(s,g)∈S×G |V ∗ (s, g) − f (ψ(s), ϕ(g))| and V̂ (s, g) := f (ψ(s), ϕDAGR (g | s)). Then sup
|V ∗ (s, g) − V̂ (s, g)| ≤ εbase + LV · B.
(s,g)∈S×G
Proof. Fix (s, g) and write Vϕ (s, g) = f (ψ(s), ϕ(g)). By the triangle inequality, |V ∗ (s, g) − V̂ (s, g)| ≤ |V ∗ (s, g) − Vϕ (s, g)| + |Vϕ (s, g) − V̂ (s, g)|. The first term is bounded by εbase by definition. For the second term, the Lipschitz assumption gives |f (ψ(s), ϕ(g)) − f (ψ(s), ϕDAGR (g | s))| ≤ LV ∥ϕ(g) − ϕDAGR (g | s)∥2 ≤ LV · B. Taking the supremum over (s, g) yields the claim. Remark D.7 (Behavior at initialization). At the start of training, σ(α0 ) = σ(−5) ≈ 0.007 and the fusion weights are uniform. The norm of the attention contribution is bounded by ∥σ(α0 )∥∞ · ∥W O ∥op · ∥Vh ∥∞ , which makes B0 ≈ 0, and the additional error LV · B0 is negligible. The bound grows during training only because the optimizer increases the gate values in response to reductions in the downstream losses, so B grows only when it pays off. D.6
Proof of Proposition 4.6
Proposition D.8 (Restated). Under the assumptions of Proposition 4.6, the data-dependent Rademacher upper bounds Unlate and UnDAGR on the policy class satisfy UnDAGR ≤ Lh · Lρ · Rn (Hρ ),
Unlate ≤ Lψ−1 · Lϕ−1 · Lh · Lρ · Rn (Hρ ),
so that the DAGR upper bound is tighter by a factor Lψ−1 Lϕ−1 ≥ 1. Proof. By the discrepancy structure assumption, π ∗ (a | s, g) = ρ(h(s, g)) where ρ is Lρ -Lipschitz and h admits an Lh -Lipschitz realization. Any policy realizing π ∗ on the respective encoder inputs computes ρ ◦ h as a function of those inputs. In the DAGR case, h̃(ψ(s), ϕDAGR (g | s)) = h(s, g) with h̃ being Lh -Lipschitz by assumption, so the policy network realizes ρ ◦ h̃. The composition property of Rademacher complexity [Bartlett and Mendelson, 2002] gives UnDAGR ≤ Lh · Lρ · Rn (Hρ ). In the late-fusion case, (ψ, ϕ) does not admit any Lh -Lipschitz realization of h by assumption. Hence any hlate : Rds × Rd → Z with hlate (ψ(s), ϕ(g)) = h(s, g) must factor through measurable inverse selections ψ −1 , ϕ−1 , yielding hlate = h ◦ (ψ −1 × ϕ−1 ). The Lipschitz constant of hlate is at most Lh · max(Lψ−1 , Lϕ−1 ), where Lψ−1 , Lϕ−1 ≥ 1 are the Lipschitz constants of the inverse selections 21
(they are ≥ 1 because ψ, ϕ are non-expansive reductions of dimension in any realistic offline GCRL encoder, hence their inverses are non-contracting). Applying the composition bound: Unlate ≤ Lψ−1 · Lϕ−1 · Lh · Lρ · Rn (Hρ ). The ratio Unlate /UnDAGR = Lψ−1 · Lϕ−1 ≥ 1, with strict inequality whenever ψ or ϕ strictly reduces dimension, which is the case for every encoder considered in this paper. D.7
Proof of Proposition 4.7
Proposition D.9 (Restated). If the task is goal-only (π ∗ (a | s, g) = ρ(h̄(g))) and the base representation ϕ is sufficient for π ∗ with ρ ◦ h̄ ∈ Flate , then any value or policy loss minimizer in FDAGR that attains εbase is attained at B = 0, equivalently at ϕDAGR (g | s) ≡ ϕ(g). Proof. By assumption there exists f ∗ ∈ Flate with f ∗ (ψ(s), ϕ(g)) = ρ(h̄(g)) and downstream value error εbase . We first show f ∗ is realizable in FDAGR . Setting all gate parameters αattn , αffn → −∞ componentwise yields σ(αattn ), σ(αffn ) → 0 in Equations (7) and (8), so each D GCA-Block reduces to the identity out = ϕ(g). The multi-scale fusion of Equation (9) then yields ϕDAGR (g | s) = ϕ(g) for all (s, g), and f ∗ (ψ(s), ϕDAGR (g | s)) = f ∗ (ψ(s), ϕ(g)) achieves εbase . This corresponds to B = 0 in Theorem 4.5. Now consider any DAGR configuration with B > 0, that is, ϕDAGR (g | s) = ϕ(g) + δ(s, g) with ∥δ∥2 > 0 for some (s, g). By the goal-only assumption, V ∗ (s, g) = f ∗ (ψ(s), ϕ(g)) depends on g only through ϕ(g). For any f ∈ FDAGR that is LV -Lipschitz in its second argument: |V ∗ (s, g) − f (ψ(s), ϕ(g) + δ(s, g))| ≥ |V ∗ (s, g) − f (ψ(s), ϕ(g))| − LV ∥δ(s, g)∥2 . The first term on the right is at least εbase in the worst case over (s, g). Therefore any minimizer attaining error exactly εbase must satisfy LV ∥δ(s, g)∥2 = 0 uniformly, that is, δ ≡ 0 and B = 0.
E
Experimental Details
E.1
Environment Details
We follow the experimental setup of Park et al. [2026] on OGBench [Park et al., 2025]. The locomotion suite uses PointMaze, AntMaze, and HumanoidMaze with medium, large, and giant variants of increasing layout complexity and path length. The manipulation suite uses a 6-DoF UR5e robot arm on Cube (Single requires placing one cube at a target, Double requires coordinating two), Scene (multi-object interaction with a cube, drawer, window, and button locks, where evaluation tasks chain up to eight atomic behaviors), and Puzzle (a Lights-Out variant in which the 3 × 3 grid admits 29 = 512 button states and the 4 × 4 grid admits 216 = 65,536). Visual variants render the same tasks as 64 × 64 RGB images, with the arm made transparent and colors adjusted for full observability. E.2
Hyperparameters
Table 5 lists all hyperparameters used for the DAGR module, the Dual base representation, GCIVL training, the visual encoder, goal sampling, and the value and policy networks across all experiments reported in this paper. E.3
Network Architecture Details
Base representation network. The Dual representation uses a GCBilinearValue network with separate state (ψ) and goal (ϕ) encoders, each consisting of a 3-layer MLP with hidden dimensions (512, 512, 512), GELU activations, and layer normalization. The temporal distance is parameterized as d(s, g) = ψ(s)⊤ ϕ(g) with output dimension 256. MS-DGCA module. The module receives three inputs, namely the goal representation ϕ(g) ∈ R256 , the state encoder output ψ(s) ∈ R512 for visual tasks (or Rds for state-based tasks), and the goal encoder output ψg (g) ∈ R512 . At each of the three scale levels, ψ(s) and ψg (g) are independently projected into Tℓ pseudo-tokens of dimension dm = 256 via separate learned linear projections. The 22
Table 5: Hyperparameters for all experiments. Hyperparameter Value DAGR Module (MS-D GCA) Scale levels L Token counts (T1 , T2 , T3 ) Attention heads H per level Head dimension dk Model dimension dm = H · dk FFN hidden dimension Gate initialization α0 Difference scaling initialization λ0 Fusion logit initialization w0 D GCA blocks per scale
3 (16, 8, 4) 4 64 256 256 −5 −5 0 1
Dual Representation Representation type Goal representation dimension Representation hidden dims Representation expectile (state) Representation expectile (visual)
bilinear (inner product) 256 (512, 512, 512) 0.9 0.7
GCIVL Training Learning rate Optimizer Batch size (state / visual) Discount γ Target network update rate τ Value expectile AWR temperature β Training steps (state / visual) Seeds (state / visual)
3 × 10−4 Adam 1024 / 256 0.99 0.005 0.9 10.0 106 / 5 × 105 8/4
Visual Encoder Architecture Stack sizes Residual blocks per stack MLP hidden dims Layer normalization Image augmentation probability
IMPALA-small (16, 32, 32) 1 (512, ) True 0.5
Goal Sampling Value: current state pcur Value: geometric future pgeom Value: random prand Actor: trajectory future ptraj
0.2 0.5 0.3 1.0
Value and Policy Networks Hidden dimensions Activation Layer normalization
(512, 512, 512) GELU True
difference map, attention, gated residual, and FFN then proceed as in Section 4.2. The three scale outputs are combined through softmax-weighted fusion as in Equation (9). Downstream networks. The value and actor are standard MLPs with hidden dimensions (512, 512, 512). They receive the concatenation of ψ(s) and the DAGR-enhanced goal representation ϕ(g | s). The value network uses an ensemble of two heads. The actor outputs a Gaussian with constant standard deviation. 23
Table 6: Component ablation on Cube-Double (state) Table 7: Sensitivity of DAGR on Cubeand AntMaze-Large (state), 8 seeds. Double-Play to token count T , attention Variant Cube-Double AntMaze-Large heads H, and gate initialization α0 . DAGR (full)
35.4±5.5
82.5±4.5
w/o gated residual w/o diff. bias (ζ(λ) = 0) w/o multi-scale (L = 1) w/o layer norm w/o FFN L = 2 blocks per scale
21.5±3.9 35.4±3.0 37.4±5.5 36.4±10.8 33.1±2.5 31.0±3.5
59.3±10.6 84.4±2.1 72.4±8.5 80.5±3.9 57.7±7.4 85.1±1.9
E.4
T
Succ.
H
1 38.5±3.1 1 4 37.4±3.6 2 8 36.4±3.6 4 16 38.1±4.1 8 32 37.2±5.9
Succ.
α0
Succ.
43.5±5.0 −5 36.8±7.9 39.8±5.1 −2 33.2±3.9 32.5±8.2 0 28.4±2.8 33.1±3.0
Computational Resources and Overhead Measurements
All experiments were conducted on NVIDIA A100 GPUs. State-based experiments require approximately three hours per seed and visual experiments approximately five hours per seed. Empirical overhead measurements taken on a single A100 are summarized below: per-step training time on the state-based Cube-Double agent increases from 17.1 ms (Dual) to 31.4 ms (DAGR), and per-step training time on the visual Cube-Double agent increases from 47.6 ms to 73.8 ms. Inference time per call increases from 1.6 ms to 4.0 ms (state-based) and from 4.1 ms to 5.9 ms (visual). The overhead is non-trivial in relative terms but the absolute training budget remains practical.
F
Additional Experimental Results
F.1
Component Ablation
Table 6 reports the full component ablation. Section 5.4 discusses it. Two entries are not covered there. Removing the multi-scale extension costs 10 points on AntMaze-Large but slightly helps Cube-Double, and stacking two D GCA blocks per scale offers no benefit over one. F.2
Sensitivity Analyses
Table 7 sweeps three parameters on Cube-Double. Section 5.4 discusses the gate initialization. The token count has no effect across the tested range, which is consistent with fine spatial decomposition not driving the manipulation behavior. The head count declines mildly, and H = 1 exceeds our default H = 4. We retain H = 4, because the configuration is held fixed across all twenty tasks and tuning it on the one task where the module underperforms would be the wrong trade. Our reported numbers therefore understate what per-task tuning would attain. F.3
Where the Module Looks: Attention and Gate Analysis
Figure 4 compares the attention distributions of standard cross-attention and D GCA on four CubeDouble samples against the corresponding difference maps. Both are nearly one-hot. D GCA attends to a different token than cross-attention on every sample, so the bias does shift the attended position, yet neither method selects the argmax of ∆. On this task the difference map therefore acts as a soft regularizer that displaces the head from the similarity-maximizing token, rather than as a literal pointer to the misaligned region. Figure 5 plots the gate and ζ(λ) trajectories that Section 5.4 discusses. F.4
Why Visual-Puzzle Remains at Zero
Table 8 supports the encoder-level account in Section 5.4. Two remedies would apply. Crossattention inside the CNN [Huang et al., 2025] preserves the correspondence that pooling destroys, and hierarchical planning sidesteps the requirement for it. Both are complementary to DAGR, since they act upstream of ϕ. 24
Attention Comparison: CA vs DGCA (cube-double-play-v0, sample 0) Difference map Δ
1.0
0.5
0.1
0
2
4
6
8
10 12 14
T=4 Δt
1.0
0.5
0.0
0
1
2
0.0
0.2
0.1
0
2
4
6
8
0.0
10 12 14
1.00
1.00
0.75
0.75
0.50
0.50
0.25
0.25
0.00
3
0
1
2
0
2
6
8
2
4
6
8
10 12 14
0.5
0.0
0.1
0
2
4
6
8
10 12 14
0
1
2
0.0
3
0
1
2
1.00
0.75
0.75
0.50
0.50
0
2
4
6
8
10 12 14
0.25
0.00
3
0.0
1.00
0.25 0
1
2
0.00
3
0.2
1
2
3
4
5
6
7
0.0
0.2
0
1
2
Token index
0
1
2
3
0
2
0.0
7
0
1
2
3
4
5
6
0.5
0.2
0.0
7
0
1
2
Token index
3
4
5
6
7
0.0
0.2
0
1
2
Token index
3
4
5
6
7
0.0
0
1
2
Token index
3
(b) Sample 1
Standard CA attention (head 0)
6
8
10 12 14
0.5
1
2
3
DGCA attention (head 0)
0.2
0.2
0.1
0.1
0.0
0.0
0
2
4
6
8
10 12 14
1.00
1.00
0.75
0.75
0.50
0.50
0.25
0.25
0.00
0.00
0
1
2
3
0
2
4
6
8
Difference map Δ
1.0
0.0
10 12 14
Standard CA attention (head 0)
0.5
0
2
4
6
8
10 12 14
1.0
0.5
0.2
0.1
0.1
0
2
4
6
8
10 12 14
0
1
2
3
0
1
2
1.00
0.75
0.75
0.50
0.50
6
7
0
2
4
6
8
10 12 14
0.25
0.00
3
0.0
1.00
0.25 0.0
5
DGCA attention (head 0)
0.2
0.0
4
Token index
Attention Comparison: CA vs DGCA (cube-double-play-v0, sample 3)
4
0
6
(a) Sample 0
1.0
0.0
5
0.4
Attention Comparison: CA vs DGCA (cube-double-play-v0, sample 2)
0.5
0.0
4
Token index
Difference map Δ
1.0
3
T=16 Δt
0
0.4 T=8 Δt
0.4
T=4 Δt
T=8 Δt
0.5
0.0
T=16 Δt
0.1
0
DGCA attention (head 0) 0.2
1.0 0.4
T=4 Δt
0.5
0.0
10 12 14
1.0
0.00
3
4
1.0
0
1
2
0.00
3
0
1
2
3
1.0
1.0
0.5
0.0
0
1
2
3
4
5
6
7
0.4
0.4
0.2
0.2
0.0
0.0
0
1
Token index
2
3
4
5
6
7
0.4 T=8 Δt
T=8 Δt
Standard CA attention (head 0)
0.2 T=16 Δt
T=16 Δt
0.2
0.0
Attention Comparison: CA vs DGCA (cube-double-play-v0, sample 1) Difference map Δ
DGCA attention (head 0)
T=4 Δt
1.0
Standard CA attention (head 0)
0
Token index
1
2
3
4
5
6
0.5
0.0
7
0.4
0.2
0
1
2
3
4
5
6
7
0.0
0.2
0
Token index
Token index
1
2
3
4
5
6
7
0.0
0
1
Token index
(c) Sample 2
2
3
4
5
6
7
Token index
(d) Sample 3
Figure 4: Attention versus difference map on four Cube-Double samples. Each panel shows three rows (token counts T ∈ {16, 4, 8}), with ∆ in the first column (left bars), standard CA attention in the middle, and D GCA attention on the right. Both CA and D GCA place near-all attention mass on a single token. D GCA chooses a different token than CA on every sample, but neither systematically aligns with the argmax of ∆. On this manipulation task, attention concentration is sharp, but the choice of token is not driven by the difference map alone. Gate and λ evolution (averaged over 4 seeds)
(a) Attention Gate
(b) FFN Gate
T=16 T=4 T=8
0.010
T=16 T=4 T=8
0.008 0.006
0.0045 0.2
0.012
ζ(λh )
σ(αffn )
σ(αattn )
0.0060
0.0050
0.008
0.014
0.0065
0.0055
(c) Diff Scaling (T=16)
0.016
0.0070
0.4
0.6
0.8
Training Step
1.0 1e6
0.2
0.4
0.6
Training Step
0.8
1.0 1e6
0.007
0.006
Head 0 Head 1 0.2
Head 2 Head 3 0.4
0.6
Training Step
0.8
1.0 1e6
Figure 5: Gate and difference-scaling evolution on AntMaze-Large. Both σ(αattn ) and σ(αffn ) stay below a small threshold throughout the entire training run, and per-head ζ(λh ) values remain within a narrow band. The state-conditioned refinement therefore operates as a small but consistent perturbation of ϕ(g) rather than as a wholesale replacement. F.5
Computational Overhead
MS-D GCA roughly doubles per-step training time on Cube-Double and adds about 55% on VisualCube-Double, with a similar relative increase at inference time. The absolute budget remains practical. Detailed timings are in Appendix E. F.6
Per-Task Learned ζ(λ)
The learned ζ(λ) values stay at or barely above initialization on five of six tasks, with the only substantial deviation on Scene. The naive reading that manipulation needs more difference bias than navigation is not supported: AntMaze-Large and Cube-Double share essentially the same value. This 25
Table 8: Visual-Puzzle: representation-level methods are uniformly zero, whereas early fusion alone (Park et al. [2026], Table 3) already achieves substantial success. Hierarchical planning (HIQL) is also far above zero. Method Puzzle-3×3 Puzzle-4×4 Late Fusion (Dual) Dual + CA Dual + DAGR Early Fusion (Orig) HIQL [Park et al., 2023]
0 0 0 22 73
0 0 0 65 60
Table 9: Per-task average ζ(λ) (mean across H = 4 heads and L = 3 scales at the end of training, one seed per task). Initial value ζ(λ0 ) = ζ(−5) ≈ 0.0067. Task Average ζ(λ) Navigation (AntMaze-Large) Single-Object Manipulation (Cube-Single) Multi-Object Manipulation (Cube-Double) Scene Arrangement Visual Navigation (Visual-AntMaze-Large) Visual Manipulation (Visual-Cube-Double)
0.0066 0.0227 0.0066 0.0592 0.0074 0.0069
is consistent with the component ablation result that removing the difference bias does not eliminate the navigation gains. Practically, the role of ∆ in MS-D GCA is best read as a structured inductive bias on the cross-attention scoring at initialization rather than as a heavily optimized contribution at convergence. F.7
Per-Task Fusion Weights
Table 10: Learned fusion weights softmax(w) at the end of training. Fine = T1 = 16, Medium = T2 = 8, Coarse = T3 = 4. Initial weights are uniform at 1/3 ≈ 0.333. Task Fine (T1 ) Medium (T2 ) Coarse (T3 ) Navigation (AntMaze-Large) Single-Object Manipulation (Cube-Single) Multi-Object Manipulation (Cube-Double) Scene Arrangement Visual Navigation (Visual-AntMaze-Large) Visual Manipulation (Visual-Cube-Double)
0.360 0.423 0.317 0.224 0.683 0.374
0.257 0.216 0.312 0.144 0.137 0.328
0.383 0.361 0.371 0.632 0.180 0.299
Fusion weights tell a less clean story than the design intuition in Section 4.3 alone would suggest. Scene drifts strongly toward the coarse scale, Visual Navigation toward the fine scale, and the remaining tasks stay close to uniform. We take this as evidence that the optimal analysis granularity is genuinely data-driven and that the role of the multi-scale architecture is to let the model select rather than to enforce a particular bias. F.8
Gate and λ Evolution on Additional Tasks
F.9
Learning Curves
Figures 8 and 9 show the full training trajectories on the state-based and visual OGBench suites. The two views tell consistent stories, and together they sharpen the picture from the aggregate tables. On the state-based suite (Figure 8), the navigation panels share a clear shape. DAGR departs from the two baselines early, typically within the first fifth of training, and the gap widens or stabilizes rather than closing. The most pronounced separation occurs on HumanoidMaze-large, AntMaze-large, and AntSoccer-arena, where the Dual baseline plateaus low and DAGR continues to climb. Standard CA 26
Gate and λ evolution (averaged over 4 seeds) (a) Attention Gate
(b) FFN Gate
0.008
0.006
T=16 T=4 T=8 0.2
0.008
(c) Diff Scaling (T=16)
T=16 T=4 T=8
0.008
ζ(λh )
0.010
σ(αffn )
σ(αattn )
0.010
0.006
0.007 0.006
0.004 0.005 0.4
0.6
0.8
Training Step
1.0 1e6
0.2
0.4
0.6
Training Step
0.8
1.0 1e6
Head 0 Head 1 0.2
Head 2 Head 3 0.4
0.6
0.8
Training Step
1.0 1e6
Figure 6: Gate and difference-scaling evolution on Cube-Double. The same pattern as on AntMazeLarge (Figure 5, main text) holds: all gates and per-head ζ(λh ) values move only marginally from their initial values throughout training. Gate and λ evolution (averaged over 4 seeds) (a) Attention Gate
0.0055 0.0050
0.012
0.010
(c) Diff Scaling (T=16)
T=16 T=4 T=8
0.008
ζ(λh )
0.0060
σ(αattn )
(b) FFN Gate T=16 T=4 T=8
σ(αffn )
0.0065
0.007
0.008
0.0045
0.006
0.0040 100000 200000 300000 400000 500000
Training Step
100000 200000 300000 400000 500000
Training Step
Head 0 Head 1
Head 2 Head 3
100000 200000 300000 400000 500000
Training Step
Figure 7: Gate and difference-scaling evolution on Visual-Cube-Double. The pattern is qualitatively identical to the state-based case. tracks Dual closely on most navigation tasks, sometimes with a modest lift. Introducing a state-goal interaction is therefore not sufficient on its own. What separates DAGR from CA is the gated residual, as Table 6 shows, and the curves confirm that the separation opens from the very first evaluation rather than emerging late. On PointMaze-medium and PointMaze-large the state space is small enough that the late-fusion bottleneck does not bite, and all three methods converge to similar curves. The manipulation panels (Cube-Single, Cube-Double, Scene) show no consistent ordering among the three methods, with curves overlapping inside the confidence band. This matches the aggregate table and is now shown to be stable across training rather than an artifact of the final checkpoint. The puzzle panels (Puzzle-3x3, Puzzle-4x4) stay flat near zero for all three methods throughout training, with no sign of a late breakthrough. We attribute this outcome to encoder-level rather than representation-level limitations in Appendix F.4. On the visual suite (Figure 9), the navigation behavior of Figure 8 reproduces: on both VisualAntMaze variants, DAGR pulls away from Dual and CA within the first quarter of training and the separation persists to the end. Visual manipulation curves cluster tightly, with no clear separation in either direction across the three methods, again consistent with the aggregate result. Visual-Puzzle remains at zero throughout training for every method, which is the strongest evidence we have that the bottleneck on these tasks is not something a goal-representation refinement, however structured, can fix. Two cross-cutting observations emerge from comparing the two figures. First, the gains from DAGR appear early and grow with training rather than emerging only at convergence. This is consistent with the gated-residual design that opens slowly. As the gates lift and B in Theorem 4.5 grows, the value approximation gains room to improve when the state-conditioned signal actually helps. Second, where DAGR does not improve over Dual the behavior splits into two regimes. On the navigation tasks with small state spaces the three curves overlap. On Cube-Double and Scene DAGR settles below Dual from early in training and stays there, which rules out a late-stage optimization artifact and points at the architectural cause identified in Section 6.
27
DAGR
pointmaze-medium-navigate
50 25
0.4
0.6
0.8
Training Step
25
0.2
0.4
0.6
0.8
Training Step
25 0
0.0
0.2
0.4
0.6
0.8
Training Step
25
0.0
0.2
0.4
0.6
0.8
Training Step
0.8
25
0.4
0.6
0.8
0
0.0
0.2
0.4
0.6
Training Step
0.8
1.0 1e6
1.0 1e6
25
0.0
0.2
0.4
0.6
0.8
1.0 1e6
0.8
1.0 1e6
0.8
1.0 1e6
cube-single-play
75 50 25
0.0
0.2
0.4
0.6
Training Step
puzzle-3x3-play
100
25
0.8
50
0
scene-play
50
0.6
Training Step
1.0 1e6
75
0.4
humanoidmaze-medium-navigate
100
50
0.2
0.2
75
0
1.0 1e6
75
0.0
0.0
75 50 25 0
0.0
0.2
0.4
0.6
Training Step
puzzle-4x4-play
100
Success Rate (%)
1.0 1e6
0.6
Training Step
Success Rate (%)
Success Rate (%)
50
0.4
25
Training Step
antsoccer-arena-navigate
100
75
0
0.2
50
100
25
0.0
75
0
1.0 1e6
50
0
1.0 1e6
cube-double-play
100
0.8
Training Step
Success Rate (%)
Success Rate (%)
50
0.6
antmaze-giant-navigate
100
75
0.4
75
0
1.0 1e6
humanoidmaze-large-navigate
100
0.2
Training Step
Success Rate (%)
Success Rate (%)
50
0.0
0.0
100
75
0
0
1.0 1e6
antmaze-large-navigate
100
25
Success Rate (%)
0.2
50
Success Rate (%)
0.0
75
Success Rate (%)
0
antmaze-medium-navigate
100
Success Rate (%)
75
CA
pointmaze-large-navigate
100
Success Rate (%)
Success Rate (%)
100
Dual
75 50 25 0
0.0
0.2
0.4
0.6
Training Step
0.8
1.0 1e6
Figure 8: Learning curves on the thirteen state-based OGBench tasks. DAGR (red), Dual (gray), and standard CA (blue) trained for 106 steps with four seeds; shaded band is the 95% confidence interval. Navigation tasks (top two rows) show a clean separation between DAGR and the two baselines that opens early in training and persists. Manipulation tasks (bottom-left) show no consistent ordering, with Dual ahead on Cube-Double and Scene, and puzzle tasks (bottom-right) stay near zero for all three methods.
28
DAGR
visual-antmaze-medium-navigate
50 25
0
1
2
3
4
Training Step
1
2
3
4
0
1
5 1e5
2
3
4
Training Step
Success Rate (%)
Success Rate (%)
25
0
25
50 25
0
1
2
3
Training Step
50 25
0
4
5 1e5
1
2
3
4
Training Step
5 1e5
visual-puzzle-3x3-play
100
75
0
75
0
5 1e5
visual-scene-play
100
50
Training Step
75 50 25 0
0
1
2
3
Training Step
4
5 1e5
visual-puzzle-4x4-play
100
Success Rate (%)
5 1e5
75
0
50
0
visual-cube-double-play
100
75
Success Rate (%)
0
visual-cube-single-play
100
Success Rate (%)
75
CA
visual-antmaze-large-navigate
100
Success Rate (%)
Success Rate (%)
100
Dual
75 50 25 0
0
1
2
3
Training Step
4
5 1e5
Figure 9: Learning curves on the seven visual OGBench tasks. Same color scheme as Figure 8, trained for 5 × 105 steps with four seeds. Visual-AntMaze (left two panels) reproduces the navigation pattern observed in the state-based setting, with DAGR separating from both baselines within the first quarter of training. Visual manipulation panels track each other closely. Visual-Puzzle tasks (rightmost two panels) remain at zero for every method.
29