ConceptioArchivearXiv CS
arXiv CSopen access

DADiff: Diffusion-Driven Cross-Domain Policy Adaptation for Reinforcement Learning

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

DAD IFF: Diffusion-Driven Cross-Domain Policy Adaptation for Reinforcement Learning

arXiv:2607.16090v1 [cs.LG] 17 Jul 2026

Hanyang Chen1 , Anirudh Satheesh2 , Longchao Da1 and Hua Wei1 Abstract— Transferring policies across domains poses a vital challenge in reinforcement learning, due to the dynamics mismatch between the source and target domains. In this paper, we consider the setting of online dynamics adaptation, where policies are trained in the source domain with sufficient data, while only limited interactions with the target domain are allowed. There are a few existing works that address the dynamics mismatch by employing domain classifiers, value-guided data filtering, or representation learning. Instead, we study the domain adaptation problem from a generative modeling perspective. Specifically, we introduce DAD IFF, a diffusion-based framework that leverages the discrepancy between source and target domain generative trajectories in the generation process of the next state to estimate the dynamics mismatch. Both reward modification and data selection variants are developed to adapt the policy to the target domain. We also provide a theoretical analysis to show that the performance difference of a given policy between the two domains is bounded by the generative trajectory deviation. More discussions on the applicability of the variants and the connection between our theoretical analysis and the prior work are further provided. We conduct extensive experiments in environments with various shifts to validate the effectiveness of our method. The results demonstrate that our method provides superior performance compared to existing approaches, effectively addressing the dynamics mismatch. We provide the code of our method at https://github.com/ hanyang-chen/DADiff-release.

I. I NTRODUCTION Reinforcement learning (RL) has shown strong potential in complex decision-making tasks, but training directly in the real-world environment (target domain) is often restricted by safety, cost, and limited interaction budgets. An alternative strategy is to train policies in a surrogate environment (source domain), such as a simulator, and then transfer them to the target domain. But due to the dynamics mismatch between the source and target domains, directly transferring the policy often leads to performance degradation, which is a critical challenge in the sim-to-real problem [1], [2]. One solution to this transfer problem is known as online dynamics adaptation [3], [4], where policies are trained with abundant sourcedomain data and only limited interactions in the target domain. In this setting, the state space, action space, and reward function remain consistent across domains, while the transition dynamics differ. Compared with solutions such as domain randomization [5]–[7] or simulator calibration [8], online dynamics adaptation does not require access to highfidelity simulators or prior knowledge of target dynamics, 1 Hanyang Chen, Longchao Da, and Hua Wei are with Arizona State University, {hchen478, longchao, hua.wei}@asu.edu 2 Anirudh Satheesh is with the University of Maryland, College Park,

[email protected]

and can therefore be applied in situations where such information is unavailable. Existing online dynamics adaptation methods, including classifier-based approaches [9], value-guided filtering [3], and representation learning [10], capture dynamics discrepancy from different perspectives: classifiers provide coarse distinctions between domains, value-guided methods depend on the modeling of forward predictions, and representation learning relies on assumptions of invariant latent structures across domains. When the domains are complex or stochastic, a key challenge that remains is to develop an approach capable of capturing dynamics discrepancy in a more finegrained and distributional manner. The generative modeling perspective provides a potential direction. Generative models, such as diffusion models [11], [12], have demonstrated strong capability in representing complex distributions. When state transitions are viewed as a conditional generative process, the mismatch between source and target domains can be interpreted as a discrepancy between their respective generative processes. Specifically, the multi-step sampling procedure in diffusion models and flow matching methods produces several latent states, which construct a generative trajectory, serving as structured signals of source–target dynamics deviation. These latent states allow the discrepancy to be captured not only at the next-state level but also along the entire trajectory. Intuitively, if the source and target domains follow different dynamics, their trajectories will diverge at multiple steps, a phenomenon we term generative trajectory deviation. This notion provides a fine-grained view of dynamics discrepancy by revealing how divergence accumulates along the trajectory, rather than relying solely on local or aggregated comparisons. Our theoretical analysis further connects trajectory deviation to performance guarantees, providing motivation for algorithmic design. Building on this perspective, we introduce DAD IFF, a diffusion-based framework for online dynamics adaptation. DAD IFF leverages latent states in diffusion models to measure generative trajectory deviation between source and target domains, and exploits this deviation in two complementary ways: (i) DAD IFF-modify, which adjusts source-domain rewards with deviation-based penalties, and (ii) DAD IFFselect, which filters source-domain data based on deviation before value function updates. We further discuss the applicability of these variants to different tasks, highlight the advantages of our method compared to prior work, and establish a connection between our analysis and the theoretical guarantee of prior work. Empirical results in environments

with various shifts show the superior performance of our method compared to existing algorithms. II. R ELATED W ORKS a) Domain Adaptation in RL: Generalizing RL policies to diverse environments is critical for real-world deployment, where transition dynamics [9], [13], state or action spaces [14], [15] may be different. To address domain adaptation, prior work falls under three categories: (i) domain randomization that randomizes transition dynamics to expose agents to many environment configurations [16], [17], (ii) meta-learning to few-shot adapt to many environments [18], [19], and (iii) expert demonstrations of target environments through imitation learning [20], [21]. However, these approaches are either computationally expensive or require hard-to-obtain demonstrations. With only limited targetdomain data, some works perform reward modifications to transition to the target domain by using transition classifiers [9], [22] or reward augmentations [4], [23]. Data selection methods [3], [24] have also been used to filter out part of the source-domain transitions and train policies on both source and target domain data. When the domains are complex or stochastic, a key challenge that remains is to develop an approach capable of capturing the dynamics discrepancy. Our method explores this challenge from a generative modeling perspective by measuring the generative trajectory deviation between the source and target domains. b) Diffusion Models in RL: Diffusion models [11], [12] have been extensively used for generating effective decisionmaking policies in several domains, such as reinforcement learning [25] and robotics [26]. Specifically, they are widely leveraged to synthesize data for offline RL [27], facilitate planning and action generation in multi-task scenarios [28], and enhance the representational capacity of learned RL policies [29]. In addition, diffusion models have also been extended to the multi-agent settings [30]. In the field of domain adaptation, they are utilized to augment the targetdomain data in order to boost the performance of offline RL policies [31]. However, the introduction of synthesizers may lead to extra computational costs, and the quality of synthesized data is hard to guarantee. In contrast, we choose to directly estimate the dynamics discrepancy by multiple latent states from diffusion models instead of generating more synthetic data. III. P RELIMINARIES a) Online Dynamics Adaptation: We consider two Markov Decision Processes (MDPs), denoted as Msrc = (S, A, Psrc , r, γ) and Mtar = (S, A, Ptar , r, γ) for the source domain and target domain, respectively. The state space S, action space A, reward function r : S × A → R and discount factor γ ∈ [0, 1] are consistent across both domains, while the transition dynamics Psrc and Ptar differ. The goal of online dynamics adaptation is to learn a policy π that achieves high performance in the target domain Mtar , utilizing sufficient data from the source domain and only limited interactions from the target domain. In addition,

we specify a domain M and define the probability that a π policy π encounters a state s at time step t as PM,t (s). Therefore, the normalized probability that a policy π visits a state-action pair (s, a) in thePdomain M can be repre∞ π sented as ρπM (s, a) := (1 − γ) t=0 γ t PM,t (s)π(a|s). The expected return of a policy π in M is defined as ηM (π) = E(s,a)∼ρπM [r(s, a)]. We assume the reward are bounded by |r(s, a)| ≤ rmax , ∀s ∈ S, a ∈ A. b) Diffusion Models: Diffusion models [11], [12] are a family of generative models that learn to generate samples from a target distribution. We mainly focus on the denoising diffusion probabilistic model (DDPM) [12] in this paper. DDPM consists of a forward process and a reverse process. The forward process is regarded as a Markov chain that gradually adds noise to data, transforming a clean data point x0 into Gaussian noise, which is formulated as follows, p p (1) xk = 1 − βk xk−1 + βk ϵ, ϵ ∼ N (0, I), where xk is the noisy data at diffusion timestep k, βk is the noise schedule, and ϵ is Gaussian noise. To simplify the forward process, we can directly sample the noisy data at diffusion timestep k as follows, √ √ xk = ᾱk x0 + 1 − ᾱk ϵ, ϵ ∼ N (0, I), (2) Qk where αk = 1 − βk and ᾱk = i=1 αi . The reverse process learns to denoise the noisy data step by step, which is formulated as follows, 1 xk−1 = √ αk

 xk − √

 r βk 1 − ᾱk−1 ϵθ (xk , k) + βk ϵ, 1 − ᾱk 1 − ᾱk

ϵ ∼ N (0, I),

(3) where ϵθ (xk , k) is a noise model that estimates the noise from the noisy data point xk . The noisy data points {xk }K k=0 form a generative trajectory from the initial noisy data xK to the clean data x0 . The training objective of the noise model is formulated as follows,   √ √ Ldiff = Ex0 ,ϵ,k ||ϵ − ϵθ ( ᾱk x0 + 1 − ᾱk ϵ, k)||2 . (4) IV. M ETHODOLOGY In this section, we first introduce a theoretical analysis to demonstrate the connection between the dynamics mismatch and the generative trajectory mismatch. Then, we present our diffusion-based method, DAD IFF, which measures the generative trajectory deviation from the perspective of diffusion models and adapts the learned policy to the target domain. The overview of our method is shown in Figure 1. A. Theoretical Analysis Before introducing the theoretical analysis, we first provide the definition of a generative trajectory, which is crucial for the analysis. For clarity, we denote the next state s′ as s′0 . Definition 4.1: (Generative trajectory.) Specify a domain M with transition dynamics PM (s′0 |s, a). There is a generative trajectory for the next state s′0 consisting of K auxiliary variables {s′k }K k=1 , referred to as latent states. These latent states form a Markov chain from the initial latent state s′K to the next state s′0 conditioned on the state–action pair (s, a).

Source

Target

theoretical guarantee of PAR is provided in Section VI-A. B. Domain Adaptation with Diffusion Theorem 4.2 provides a theoretical guarantee linking the performance difference of a policy π to the generative trajectory, thereby motivating a careful design of latent states in the trajectory. In this section, we adopt the formulation of DDPM to better characterize the dynamics discrepancy. We first redeclare the reverse process of DDPM in a reparameterized form to describe the latent state transition in domain M as follows,

...

...

1 s′k−1 = √ αk

Fig. 1: Illustration of DAD IFF. This figure visualizes the generative trajectories in the source and target domains. The deviation d(s, a, s′ ) is measured by the discrepancy dk of each latent state s′k in the source and target domain generative trajectories.

  r βk 1 − ᾱk−1 s′k − √ βk ϵ, ϵM (s′k , s, a, k) + 1 − ᾱk 1 − ᾱk

(6) where ϵM (s′k , s, a, k) is the noise from the latent state s′k in domain M. It indicates that the latent state transition follows a Gaussian distribution, i.e., PM (s′k−1 | s′k , s, a)

Remark. The Markov-chain definition enables the transition dynamics to be decomposed into multiple conditional i.e., PM (s′0 |s, a) = R QK probabilities, ′ ′ ′ PM (sK |s, a) k=1 PM (sk−1 |sk , s, a)ds′1:K . In this way, the next state s′0 can be viewed as being generated step by step with latent states, forming a generative trajectory. The discrepancy of such generative trajectories across domains provides a natural estimation of the dynamics discrepancy. We construct generative trajectories in both source and target domains, starting from the same initial latent state s′K , and derive Theorem 4.2 to establish the connection between the dynamics mismatch and the generative trajectory mismatch. The detailed proof is provided in Appendix VII-B. Theorem 4.2: (Performance bound controlled by generative trajectory discrepancy.) Denote Msrc and Mtar as the source and target domains with different dynamics, respectively. The performance difference of any policy π evaluated in Msrc and Mtar can be bounded as below, ηMsrc (π) − ηMtar (π) ≤ √  q 2γrmax Eρπsrc EPsrc [DKL (Psrc (s′K | s, a) ∥ Ptar (s′K | s, a))] + (1 − γ)2 | {z } (a): initial latent state deviation

v "K # u u X  2γrmax tE ′ ′ , s, a) ∥ P ′ ′ , s, a)  . π  E D P (s | s (s | s ρ P KL src tar src k−1 k k−1 k src (1 − γ)2 k=1 | {z } (b): latent state transition mismatch

(5) Remark. This bound indicates that the performance difference of a policy π between the source and target domains is controlled by the initial latent state deviation term (a) and the latent state transition mismatch term (b). Since the generative trajectories in both the source and target domains share the same initial latent state s′K , term (a) vanishes, leaving term (b) as the sole determinant of the performance difference. In other words, as long as the generative trajectories are similar in the source and target domains, the performance difference is small, and vice versa. We note that PAR [10] can be considered as a special case of Theorem 4.2 when K = 1. A discussion on the connection between our analysis and the

ϵ ∼ N (0, I),

∼N

!  1 − ᾱ βk 1  ′ k−1 sk − √ βk I . ϵM (s′k , s, a, k) , √ αk 1 − ᾱk 1 − ᾱk

(7) According to Theorem 4.2, the performance difference of a policy π across domains is determined by the latent state transition mismatch term (b). Therefore, we can estimate the generative trajectory deviation d(s, a, s′ ) with the defined distribution of latent state transition in Equation 7 as follows, d(s, a, s′ ) =

=

K X k=1 K X k=1

DKL (Psrc (s′k−1 |s′k , s, a)||Ptar (s′k−1 |s′k , s, a)) βk 2 ∥ϵsrc (s′k , s, a, k) − ϵtar (s′k , s, a, k)∥ . 2(1 − ᾱk−1 )αk

(8) We derive this equation by computing the KL divergence between two Gaussian distributions. Notably, as the state transition tuple (s, a, s′ ) comes from the source domain, the noise ϵsrc (s′k , s, a, k) estimated in the reverse process must be consistent with the noise used in the forward process to generate the latent state s′k , which indicates ϵsrc (s′k , s, a, k) = ϵ with ϵ ∼ N (0, I). Besides, we introduce a noise model ϵθtar (s′k , s, a, k), trained with target-domain data, to estimate the noise in the target domain. The training objective is formulated as follows, Lnoise = E(s,a,s′ )∼Dtar ,ϵ,k

h

i √ √ 2 ϵ − ϵθtar ( ᾱk s′0 + 1 − ᾱk ϵ, s, a, k) .

(9) This objective mirrors the standard DDPM training loss, but conditions on (s, a) to capture dynamics in the target domain. For the latent state s′k in Equation 8, there are two ways to obtain it: (i) by iteratively applying the reverse process in Equation 6, and (ii) by sampling √ directly from √ the forward process of DDPM, i.e., s′k = ᾱk s′0 + 1 − ᾱk ϵ with ϵ ∼ N (0, I). Specifically, the first way requires sequential sampling across all steps to generate the entire generative trajectory, which is computationally expensive. In contrast, the second way can produce all latent states in parallel, yielding a much more efficient implementation. Therefore, we choose to obtain the latent state s′k via the forward process in our method. Finally, the deviation d(s, a, s′ ) can

be practically estimated as follows, d(s, a, s′ ) =

K X k=1

Algorithm 1: Domain Adaptation with DAD IFF

√ √ βk 2 ϵ − ϵθtar ( ᾱk s′0 + 1 − ᾱk ϵ, s, a, k) , 2(1 − ᾱk−1 )αk ϵ ∼ N (0, I).

(10) We further introduce two variants based on SAC [32] to utilize the deviation d(s, a, s′ ), including reward modification and data selection, since we find that baselines adopting these two techniques exhibit complementary advantages in different tasks, which is shown in Section V-B. We analyze the possible reason for this phenomenon from the reward distribution aspect in Section VI-B. The details of DAD IFF variants are provided as follows. a) Reward modification.: We refer to this variant as DAD IFF-modify. It adopts the deviation d(s, a, s′ ) as a reward penalty to modify the reward function in the source domain, i.e., ′

rmod (s, a, s ) = r(s, a, s ) − λd(s, a, s ),

(11)

where λ is a penalty coefficient to balance the original reward and the penalty. The objective function for training the value function gives,   Lcritic = E(s,a,rmod ,s′ )∼Dsrc ∪Dtar (Qϕ − T Qϕ )2 , (12) where Dtar and Dsrc are the datasets from the target and source domains, respectively, Qϕ is the value function, and T is the Bellman operator. b) Data selection.: We refer to this variant as DAD IFFselect. We select fixed percentage data with the lowest deviation d(s, a, s′ ) from a batch of source domain data. The selected data is then used to update the value function. We formulate the objective function of the value function as follows,   Lcritic = E(s,a,r,s′ )∼Dtar (Qϕ − T Qϕ )2 +   (13) E(s,a,r,s′ )∼Dsrc ω(s, a, s′ )(Qϕ − T Qϕ )2 , where ω(s, a, s′ ) = 1(d(s, a, s′ ) < dξ% ), 1 is the indicator function, and dξ% denotes the lowest ξ-quantile deviation in the batch. For both variants, the objective function of the policy π is formulated as: Lactor = E(s,a,r,s′ )∼Dsrc ∪Dtar [− mini=1,2 Qϕi (s, a) + τ log π(a|s)] ,

(14) where τ is the entropy temperature coefficient, and i denotes the value function index. We provide the pseudocode of DAD IFF in Algorithm 1. V. E XPERIMENTS A. Experimental Setup We conduct experiments in four environments (ant, hopper, halfcheetah, walker) from Gym MuJoCo [33], [34]. The source domain is set as the original environment, while the target domain is set as the environment with shifts in kinematics, morphology, friction, or gravity. Kinematic shifts restrict joint rotation ranges, morphology shifts reduce

Input: Source domain Msrc , target domain Mtar , and target domain interaction frequency F Initialization: Policy π, value function {Qϕi }i=1,2 , target value function {Qϕ′ }i=1,2 , noise model ϵθtar , replay i buffers {Dsrc , Dtar }, penalty coefficient λ, data selection ratio ξ, batch size N 1 for j = 1, 2, . . . do 2 Collect (ssrc , asrc , rsrc , s′src ) from Msrc , store in Dsrc 3 if j mod F = 0 then 4 Collect (star , atar , rtar , s′tar ) from Mtar , store in Dtar

7 8 9

Sample N transitions from Dtar , train model ϵθtar via Eq. 9 Sample N transitions from Dsrc , compute d(ssrc , asrc , s′src ) via Eq. 10 if using reward modification then Modify source domain rewards via Eq. 11 Update value functions Qϕi by minimizing Eq. 12

10 11 12

else if using data selection then Select ξ-quantile data from Dsrc by d(ssrc , asrc , s′src ) Update value functions Qϕi by minimizing Eq. 13

13 14

Update actor π by minimizing Eq. 14 Update target value functions Qϕ′

5 6

i

limb sizes, friction shifts modify the friction coefficient, and gravity shifts adjust gravitational acceleration. Kinematic and morphology configurations follow PAR [10], while friction and gravity shifts follow ODRL [35] at a level of 0.5. We compare our method with the following baselines: DARC [9], which trains domain classifiers to estimate the dynamics discrepancy and modifies the reward function in the source domain; VGDF [3], which uses a valueguided data filtering method to select data from the source domain; PAR [10], which trains encoders to estimate the representation discrepancy and modifies the reward function in the source domain; SAC-IW, which estimates the dynamics discrepancy as an importance sampling term for value function; SAC-tune, which fine-tunes the policy in the target domain for 105 environmental steps; SAC-tar [32], which is the vanilla SAC trained in the target domain with 105 environmental steps; Oracle [32], which is the vanilla SAC trained in the target domain with 1M environmental steps. We implement all algorithms based on the official code of ODRL [35] and follow the hyperparameters in the original paper. We allow all algorithms to interact with the source domain for 1M environmental steps and the target domain for 105 environmental steps, i.e., the target domain interaction frequency F = 10. All algorithms are trained with five random seeds. B. Adaptation Performance Evaluation We conduct experiments on sixteen tasks with diverse shifts to evaluate the adaptation performance of DAD IFF and baselines. The results are summarized in Figure 2. Overall, DAD IFF exhibits consistently strong performance, demonstrating superior or competitive performance against all baselines in the majority of tasks. While some existing methods, such as VGDF, PAR, or SAC-tune, occasionally reach competitive results in specific tasks, their performance fluctuates significantly across different tasks. In contrast,

DADiff-modify

DADiff-select

ant (broken hips)

PAR

VGDF

DARC

halfcheetah (broken back thigh)

4000

SAC-IW

SAC-tar

SAC-tune

hopper (broken joints)

6000

3000

4000

2000

2000

1000

Oracle

walker (broken right foot) 4000

Return

3000 2000 0 0 0.0

0.2

0.4

0.6

0.8

Environment Steps

1.0 1e5

0.0

Return

0.2

0.4

0.6

0.8

Environment Steps

1.0 1e5

0 0.0

halfcheetah (no thighs)

0.2

0.4

0.6

0.8

Environment Steps

1.0 1e5

2000

2500

1000

0 0.2

0.4

0.6

0.8

Environment Steps

1.0 1e5

0.0

0.2

0.4

0.6

0.8

Environment Steps

1.0 1e5

1000 0 0.0

0.2

0.4

0.6

0.8

Environment Steps

1.0 1e5

0.2

0.4

0.6

0.8

Environment Steps

1.0 1e5

0 0.0

0.2

0.4

0.6

0.8

Environment Steps

1.0 1e5

0.2

0.4

0.6

0.8

Environment Steps

1.0 1e5

hopper (gravity) 3000

0.2

0.4

0.6

0.8

Environment Steps

1.0 1e5

0.0

0.2

0.4

0.6

0.8

Environment Steps

1.0 1e5

0.6

0.8

1.0 1e5

0.8

1.0 1e5

walker (gravity)

1000

0

0

0.4

Environment Steps

2000

1000

2500

0.2

3000

2000

5000

0.0

4000

7500 2000

0.0

1.0 1e5

0 0.0

halfcheetah (gravity)

0

0.8

2000

10000

1000

0.6

walker (friction)

1000

ant (gravity) 3000

0.4

Environment Steps

4000

0 0.0

0.2

2000

2500

0

0.0

hopper (friction)

5000

2000

1.0 1e5

3000

7500

4000

0.8

2000

halfcheetah (friction) 10000

6000

0.6

walker (no right thigh)

0

ant (friction)

8000

0.4

Environment Steps

4000

2000

0.0

0.2

3000

5000

0

0.0

hopper (big head) 3000

7500

4000

Return

1000

0

ant (short feet)

Return

2000

0 0.0

0.2

0.4

0.6

0.8

Environment Steps

1.0 1e5

0.0

0.2

0.4

0.6

Environment Steps

Fig. 2: Adaptation performance under kinematic, morphology, friction, and gravity shifts (from top to bottom). The solid curves and the shaded regions denote the mean and standard deviation over five random seeds, respectively. DADiff demonstrates superior or highly competitive performance against all baselines in the majority of tasks. DADiff-modify

DADiff-select

PAR

100 50 0

VGDF

15

200 150

DARC

Runtime Comparison

Runtime (h)

GPU Memory (MB)

GPU Memory Comparison

10 5 0

Fig. 3: GPU memory and runtime comparisons on the halfcheetah (broken back thigh) task. In the GPU memory comparison, DAD IFF-modify and DAD IFF-select exhibit slightly higher GPU memory cost compared to PAR and DARC. In the runtime comparison, VGDF requires 3× more training time than other methods due to its model-based approach.

DAD IFF maintains stable and superior adaptation performance across a wide range of shift types. We further discuss the performance of two variants of DAD IFF, DAD IFFmodify and DAD IFF-select, respectively. a) Reward modification variant.: The reward modification variant of our method, DAD IFF-modify, demonstrates strong and consistent performance across diverse tasks. As shown in Figure 2, it surpasses other reward modification baselines, including PAR, DARC, and SAC-IW, in most tasks and achieves performance comparable to oracle-level methods. On average, DAD IFF-modify improves by 8.7% across all sixteen tasks, with the largest gain of 42.3%

on the halfcheetah (broken back thigh). In addition to its performance advantages, we observe that our method incurs a slight increase in GPU memory usage compared to PAR and DARC due to latent state generation, as shown in Figure 3. This modest increase, however, contributes positively to adaptation performance by enabling better discrepancy estimation, thus representing a favorable trade-off between computational cost and effectiveness. To further explore the performance of DAD IFF-modify in stochastic environments, we provide an experiment in Section VI-A. b) Data selection variant.: In Figure 2, the data selection variant, DAD IFF-select, proves to be a highly effective alternative by achieving competitive performance against top baselines in tasks where reward modification methods falter. Specifically, in the halfcheetah (no thighs), hopper (big head), and hopper (friction) tasks, reward modification methods exhibit poor performance. In contrast, DAD IFFselect achieves results that are highly competitive with the top-performing baseline, VGDF. This indicates that in certain tasks, directly filtering for transitions with low dynamics mismatch is a more effective strategy than modifying rewards. We analyze the possible reason in Section VI-B. Furthermore, while VGDF demonstrates top-tier performance in these tasks, it carries significant trade-offs. Since VGDF is a

λ=0 λ=0.5

λ=0.01 λ=1.0

C. Parameter Study

VI. D ISCUSSIONS A. Connection between DAD IFF and PAR We explore the connection between PAR and our method from a theoretical perspective. The performance bound of our method is controlled by the generative trajectory discrepancy in Theorem 4.2. We consider a special case, where the number of latent states in the trajectory is K = 1. Instead of considering latent states in the generative trajectory, we take s′1 as a latent representation and introduce the one-toone representation mapping assumption in PAR [10], which assumes that there exists a one-to-one mapping for each state-action pair (s, a) and its latent representation s′1 . In this setting, the state-action pair (s, a) in Equation 5 can

λ=0.1 λ=5.0

walker (no right thigh) 4000 3000

4000 2000 2000

1000

0

0 0.0

0.2

0.4

0.6

0.8

1.0 1e5

Environment Steps

0.0

0.2

0.4

0.6

0.8

Environment Steps

1.0 1e5

(a) Penalty coefficient λ. ξ%=0% ξ%=50%

ξ%=10% ξ%=75%

ξ%=25% ξ%=100%

halfcheetah (broken back thigh)

Return

walker (no right thigh) 4000

6000

3000

4000

2000 2000

1000

0

0 0.0

0.2

0.4

0.6

0.8

1.0 1e5

Environment Steps

0.0

0.2

0.4

0.6

0.8

Environment Steps

1.0 1e5

(b) Data ratio ξ%. K=10

K=50

K=100

halfcheetah (broken back thigh) 6000

Return

The performance of DAD IFF is influenced by several key hyperparameters. To better understand their roles, we conducted a series of experiments across different tasks. The results on halfcheetah (broken back thigh) and walker (no right thigh) are presented in Figure 4. a) Penalty Coefficient λ.: λ controls the scale of reward penalty in DAD IFF-modify. As shown in Figure 4a, we evaluate the performance of DAD IFF-modify across multiple values of λ. We find that a worse performance is often shown in the setting λ = 0, where no penalty is adopted for rewards. It demonstrates the necessity of reward modification. Meanwhile, the results also indicate that the optimal value of λ is task-dependent, and there could be multiple values that yield good performance for a specific task. For instance, in the halfcheetah (broken back thigh) task, both λ = 0.5 and λ = 5.0 achieve the best performance. A poorly chosen λ can significantly degrade performance, highlighting the importance of tuning this coefficient. b) Data Selection Ratio ξ%.: ξ% controls the ratio of source domain data to retain in DAD IFF-select. As shown in Figure 4b, we evaluate the performance of DAD IFFselect across multiple values of ξ%. Similar to the penalty coefficient, the optimal value of ξ% is task-dependent. We find that both too much (ξ% = 100%) and too little ((ξ% = 0%)) source data can lead to suboptimal performance. As retaining too much source data may introduce transitions with significant dynamics mismatch, while retaining too little may result in insufficient data for effective learning. c) Diffusion Timesteps K.: K controls the number of diffusion timesteps used to measure the discrepancy in both DAD IFF-modify and DAD IFF-select. We provide the results of DAD IFF-modify in Figure 4c. The results shows that performance improves up to K = 100. Increasing K further to 200 causes a decline, likely due to the limited capacity of the noise model, which may struggle to accurately estimate noise across too many timesteps.

λ=0.05 λ=2.0

halfcheetah (broken back thigh) 6000

Return

model-based approach, it takes significantly longer to train by more than 3×, as shown in Figure 3. On the other hand, DAD IFF-select is able to match or exceed the performance of VGDF on such environments while maintaining comparable efficiency to similar model-free baselines.

K=200

walker (no right thigh) 4000 3000

4000

2000 2000

1000

0

0 0.0

0.2

0.4

0.6

0.8

Environment Steps

1.0 1e5

0.0

0.2

0.4

0.6

0.8

Environment Steps

1.0 1e5

(c) Diffusion timesteps K. Fig. 4: Parameter study. The solid curves and the shaded regions denote the mean and standard deviation over five random seeds, respectively.

be all replaced by the corresponding latent representation s′1 . Therefore, the performance bound can be rewritten as follows, ηMsrc (π) − ηMtar (π) ≤ √ q  2γrmax ′ |s′ )||P ′ |s′ ))] . π E E [D (P (s (s ρ P KL src tar src 0 1 0 1 (1 − γ)2 src

(15) We further introduce a conclusion proven in PAR [10]: DKL (Psrc (s′1 |s′0 )||Ptar (s′1 |s′0 )) = DKL (Psrc (s′0 |s′1 )||Ptar (s′0 |s′1 )) + H(s′src ) − H(s′tar ).

(16) Therefore, the performance bound can be rewritten as follows, ηMsrc (π)−ηMtar (π) ≤ √  q 2γrmax ′ |s′ )||P ′ |s′ ))] + π E [D (P (s (s E ρ P KL src tar src 1 0 1 0 (1 − γ)2 src √ hp i 2γrmax ′ ) − H(s′ )] . π E E [H(s ρ P tar src src (1 − γ)2 src

(17) This performance bound is consistent with the performance bound of PAR, which indicates that PAR can be considered as a special case of our method. However, the one-toone representation mapping assumption may not hold in practice, especially in stochastic environments, which limits

TABLE I: Adaptation performance under stochastic dynamics controlled by the standard deviation parameter ς. Average return and standard deviation over five random seeds are reported. The best results are in bold, and performance change relative to the deterministic setting (ς = 0.0) is shown in parentheses. Environment

ς

DAD IFF-modify

PAR

hopper (broken joints)

0.00 0.01 0.02 0.03

2582.1±251.6 2591.0±159.2 (↑0.34%) 2515.9±101.8 (↓2.57%) 2574.2±280.6 (↓0.31%)

2623.1±105.2 2398.3±297.8 (↓8.57%) 2328.7±302.9 (↓11.22%) 2406.1±455.7 (↓8.27%)

0.00 0.01 0.02 0.03

3390.4±464.4 2879.3±688.9 (↓15.08%) 2812.5±934.6 (↓17.05%) 3176.8±796.4 (↓6.30%)

2943.3±546.7 2373.8±1072.4 (↓19.35%) 2825.8±466.6 (↓3.99%) 1613.9±878.7 (↓45.17%)

walker (broken right foot)

Original 10.0

Processed

halfcheetah (no thighs)

hopper (big head) 6

5.0

Reward

Reward

7.5

2.5 0.0

which utilizes diffusion models to measure the dynamics discrepancy and performs either reward modification or data selection to adapt to the target domain. Extensive experiments demonstrate that our method outperforms existing baselines in tasks with various shifts. APPENDIX A. Useful Lemmas Lemma 7.1: (Telescoping lemma.) Denote M1 = (S, A, P1 , r, γ) and M2 = (S, A, P2 , r, γ) as two MDPs with the same state and action spaces but different transition dynamics P1 and P2 . The performance difference of a policy π evaluated in M1 and M2 can be expressed as:   γ π π (s′ )] − Es′ ∼P2 [VM (s′ )] ηM1 (π) − ηM2 (π) = 1−γ EρπM (s,a) Es′ ∼P1 [VM 2 2 1

4

Proof. Please see Lemma 4.3 in SLBO [36] for a detailed proof.

2 0

−2.5 DADiff-modify

DADiff-select

DADiff-modify

DADiff-select

Fig. 5: Reward distribution comparison between the source-domain rewards before processing (Original) and after modification or selection (Processed).

the application of PAR. In contrast, our method does not rely on this assumption and can handle more general scenarios. We validate this point in environments with stochastic dynamics. Noises with different standard deviation ς are introduced to the actions to simulate stochastic dynamics, and two tasks with kinematic shifts, hopper (broken joints) and walker (broken right foot), are considered. We evaluate the performance of DAD IFF-modify and PAR, which is presented in Table I. Notably, our method maintains robust performance even as the standard deviation ς increases, while PAR’s performance degrades significantly. We believe the decrease in PAR’s performance is due to its reliance on oneto-one representation assumptions, which may not hold in stochastic settings. B. Reward Distribution Analysis We further examine the reasons behind the superior performance of DAD IFF-select, in contrast to the severe failure of DAD IFF-modify on halfcheetah (no thighs) and hopper (big head) tasks, as illustrated in Figure 2. Specifically, we analyze the reward distributions of source-domain data after modification or selection. The results are presented in Figure 5. We find that DAD IFF-select generates a higher distribution in the low-reward region compared to DAD IFFmodify on both tasks. This suggests that the low-reward data may play a crucial role in these tasks, which can effectively guide the policy to avoid undesirable states and actions. VII. C ONCLUSION This work explores the problem of online dynamics adaptation in reinforcement learning from a generative modeling perspective. We first theoretically analyze the performance bound of a policy in the source and target domains, which is controlled by the generative trajectory discrepancy. Based on this analysis, we propose a novel method, DAD IFF,

B. Proof of Theorem 4.2 Theorem 7.2: (Performance bound controlled by generative trajectory discrepancy.) Denote Msrc and Mtar as the source and target domains with different dynamics, respectively. The performance difference of any policy π evaluated in Msrc and Mtar can be bounded as below, ηMsrc (π) − ηMtar (π) ≤ √  q 2γrmax Eρπsrc EPsrc [DKL (Psrc (s′K | s, a) ∥ Ptar (s′K | s, a))] + 2 (1 − γ) | {z } (a): initial latent state deviation

v "K # u u X  2γrmax t ′ ′ ′ ′  π Eρsrc EPsrc DKL Psrc (sk−1 | sk , s, a) ∥ Ptar (sk−1 | sk , s, a)  . (1 − γ)2 k=1 {z } | (b): latent state transition mismatch

π (s) estimates the expected Proof. As the value function VM return of a policy π starting from state s in domain M, and π the rewards are bounded, we have |VM (s)| ≤ rmax /(1 − γ), ∀s. By using Lemma 7.1, we have: γ ηMsrc (π) − ηMtar (π) = Eρπ [EPsrc [r(s, a)] − EPtar [r(s, a)]] 1 − γ src # "Z Z γ π π = Eρπ Psrc (s′0 |s, a)Vtar (s′0 ) − Ptar (s′0 |s, a)Vtar (s′0 )ds′0 ′ 1 − γ src s′0 s0 "Z # γ ′ ′ π (P (s |s, a) − P (s |s, a)) Vtar (s′0 ) ds′0 ≤ Eρπ src 0 tar 0 src 1−γ s′0 "Z # γrmax ′ ′ ′ π ≤ E P (s |s, a) − P (s |s, a)ds src tar ρ 0 0 0 src (1 − γ)2 s′ "Z 0 # γrmax ′ ′ ′ π = E P (s |s, a) − P (s |s, a)ds src tar ρ 0:K 0:K 0:K (1 − γ)2 src s′0:K   2γrmax ′ ′ = Eρπ DTV (Psrc (s0:K |s, a)||Ptar (s0:K |s, a)) (1 − γ)2 src √ q  2γrmax ≤ Eρπ DKL (Psrc (s′0:K |s, a)∥Ptar (s′0:K |s, a)) src 2 (1 − γ) (a) v " # u √ ′ u P (s |s, a) 2γrmax tEPsrc log src 0:K  = Eρπ src (1 − γ)2 Ptar (s′0:K |s, a) v " # u √ K u X Psrc (s′k−1 |s′k , s, a) Psrc (s′K |s, a) 2γrmax  π tEP = E log + log ρ src src (1 − γ)2 Ptar (s′K |s, a) k=1 Ptar (s′k−1 |s′k , s, a) (b) q    2γrmax ′ ′ π ≤ Eρsrc EPsrc DKL (Psrc (sK |s, a)||Ptar (sK |s, a)) + 2 (1 − γ) v " K # u √ u X 2γrmax ′ ′ , s, a)||P ′ ′ , s, a))  π tEP E D (P (s |s (s |s src tar ρ KL src k−1 k k−1 k src (1 − γ)2 k=1 √

(c)

where DTV (P ||Q) is the total variation distance between two distributions P and Q, the step (a) holds by Pinsker’s inequality, the step (b) holds by the Markov property, and the step (c) holds by the subadditivity of the square root function. The proof shows that the performance difference can be controlled by the distributional divergence of latent states in generative trajectories. ACKNOWLEDGMENT The work was partially supported by NSF award #2442477, #2550203 and #2536297. The views and conclusions in this paper should not be interpreted as representing any funding agencies. R EFERENCES [1] W. Zhao, J. P. Queralta, and T. Westerlund, “Sim-to-real transfer in deep reinforcement learning for robotics: a survey,” in 2020 IEEE symposium series on computational intelligence (SSCI). IEEE, 2020, pp. 737–744. [2] L. Da, J. Turnau, T. P. Kutralingam, A. Velasquez, P. Shakarian, and H. Wei, “A survey of sim-to-real methods in rl: Progress, prospects and challenges with foundation models,” arXiv preprint arXiv:2502.13187, 2025. [3] K. Xu, C. Bai, X. Ma, D. Wang, B. Zhao, Z. Wang, X. Li, and W. Li, “Cross-domain policy adaptation via value-guided data filtering,” Advances in Neural Information Processing Systems, vol. 36, pp. 73 395– 73 421, 2023. [4] J. Lyu, C. Bai, J. Yang, Z. Lu, and X. Li, “Cross-domain policy adaptation by capturing representation mismatch,” arXiv preprint arXiv:2405.15369, 2024. [5] X. B. Peng, M. Andrychowicz, W. Zaremba, and P. Abbeel, “Sim-toreal transfer of robotic control with dynamics randomization,” in 2018 IEEE international conference on robotics and automation (ICRA). IEEE, 2018, pp. 3803–3810. [6] B. Mehta, M. Diaz, F. Golemo, C. J. Pal, and L. Paull, “Active domain randomization,” in Conference on Robot Learning. PMLR, 2020, pp. 1162–1176. [7] A. Curtis, E. Li, M. Noseworthy, N. Gothoskar, S. Chitta, H. Li, L. P. Kaelbling, and N. E. Carey, “Flow-based domain randomization for learning and sequencing robotic skills,” in Forty-second International Conference on Machine Learning, 2025. [8] Y. Chebotar, A. Handa, V. Makoviychuk, M. Macklin, J. Issac, N. Ratliff, and D. Fox, “Closing the sim-to-real loop: Adapting simulation randomization with real world experience,” in 2019 International Conference on Robotics and Automation (ICRA). IEEE, 2019, pp. 8973–8979. [9] B. Eysenbach, S. Asawa, S. Chaudhari, S. Levine, and R. Salakhutdinov, “Off-dynamics reinforcement learning: Training for transfer with domain classifiers,” arXiv preprint arXiv:2006.13916, 2020. [10] J. Lyu, C. Bai, J. Yang, Z. Lu, and X. Li, “Cross-domain policy adaptation by capturing representation mismatch,” in Proceedings of the 41st International Conference on Machine Learning, 2024, pp. 33 638–33 663. [11] J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli, “Deep unsupervised learning using nonequilibrium thermodynamics,” in International conference on machine learning. pmlr, 2015, pp. 2256–2265. [12] J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” Advances in neural information processing systems, vol. 33, pp. 6840–6851, 2020. [13] Z. Xue, Q. Cai, S. Liu, D. Zheng, P. Jiang, K. Gai, and B. An, “State regularized policy optimization on data with dynamics shift,” Advances in neural information processing systems, vol. 36, pp. 32 926–32 937, 2023. [14] Y. Ge, A. Macaluso, L. E. Li, P. Luo, and X. Wang, “Policy adaptation from foundation model feedback,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 19 059–19 069. [15] K.-C. Pan, M. Chen, Y.-D. Huang, X. Liu, and P.-C. Hsieh, “Crossdomain reinforcement learning under distinct state-action spaces via hybrid q functions.”

[16] R. B. Slaoui, W. R. Clements, J. N. Foerster, and S. Toth, “Robust visual domain randomization for reinforcement learning,” arXiv preprint arXiv:1910.10537, 2019. [17] Y. Jiang, C. Li, W. Dai, J. Zou, and H. Xiong, “Variance reduced domain randomization for reinforcement learning with policy gradient,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 2, pp. 1031–1048, 2023. [18] A. Nagabandi, I. Clavera, S. Liu, R. S. Fearing, P. Abbeel, S. Levine, and C. Finn, “Learning to adapt in dynamic, real-world environments through meta-reinforcement learning,” arXiv preprint arXiv:1803.11347, 2018. [19] Z. Wu, Y. Xie, W. Lian, C. Wang, Y. Guo, J. Chen, S. Schaal, and M. Tomizuka, “Zero-shot policy transfer with disentangled task representation of meta-reinforcement learning,” arXiv preprint arXiv:2210.00350, 2022. [20] D. S. Raychaudhuri, S. Paul, J. Vanbaar, and A. K. Roy-Chowdhury, “Cross-domain imitation from observations,” in International conference on machine learning. PMLR, 2021, pp. 8902–8912. [21] A. Fickinger, S. Cohen, S. Russell, and B. Amos, “Crossdomain imitation learning via optimal transport,” arXiv preprint arXiv:2110.03684, 2021. [22] Y. Guo, Y. Wang, Y. Shi, P. Xu, and A. Liu, “Off-dynamics reinforcement learning via domain adaptation and reward augmented imitation,” Advances in Neural Information Processing Systems, vol. 37, pp. 136 326–136 360, 2024. [23] L. L. P. Van, H. T. Tran, and S. Gupta, “Policy learning for offdynamics rl with deficient support,” arXiv preprint arXiv:2402.10765, 2024. [24] X. Wen, C. Bai, K. Xu, X. Yu, Y. Zhang, X. Li, and Z. Wang, “Contrastive representation for data filtering in cross-domain offline reinforcement learning,” arXiv preprint arXiv:2405.06192, 2024. [25] B. Kang, X. Ma, C. Du, T. Pang, and S. Yan, “Efficient diffusion policies for offline reinforcement learning,” Advances in Neural Information Processing Systems, vol. 36, pp. 67 195–67 212, 2023. [26] C. Chi, Z. Xu, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song, “Diffusion policy: Visuomotor policy learning via action diffusion,” The International Journal of Robotics Research, vol. 44, no. 10-11, pp. 1684–1704, 2025. [27] C. Lu, P. Ball, Y. W. Teh, and J. Parker-Holder, “Synthetic experience replay,” Advances in Neural Information Processing Systems, vol. 36, pp. 46 323–46 344, 2023. [28] H. He, C. Bai, K. Xu, Z. Yang, W. Zhang, D. Wang, B. Zhao, and X. Li, “Diffusion model is an effective planner and data synthesizer for multi-task reinforcement learning,” Advances in neural information processing systems, vol. 36, pp. 64 896–64 917, 2023. [29] Y. Wang, L. Wang, Y. Jiang, W. Zou, T. Liu, X. Song, W. Wang, L. Xiao, J. Wu, J. Duan, et al., “Diffusion actor-critic with entropy regulator,” Advances in Neural Information Processing Systems, vol. 37, pp. 54 183–54 204, 2024. [30] Z. Zhu, M. Liu, L. Mao, B. Kang, M. Xu, Y. Yu, S. Ermon, and W. Zhang, “Madiff: Offline multi-agent learning with diffusion models,” Advances in Neural Information Processing Systems, vol. 37, pp. 4177–4206, 2024. [31] L. L. P. Van, M. H. Nguyen, D. Kieu, H. Le, H. T. Tran, and S. Gupta, “Dmc: Nearest neighbor guidance diffusion model for offline crossdomain reinforcement learning,” arXiv preprint arXiv:2507.20499, 2025. [32] T. Haarnoja, A. Zhou, K. Hartikainen, G. Tucker, S. Ha, J. Tan, V. Kumar, H. Zhu, A. Gupta, P. Abbeel, et al., “Soft actor-critic algorithms and applications,” arXiv preprint arXiv:1812.05905, 2018. [33] E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control,” in 2012 IEEE/RSJ international conference on intelligent robots and systems. IEEE, 2012, pp. 5026–5033. [34] G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba, “Openai gym,” arXiv preprint arXiv:1606.01540, 2016. [35] J. Lyu, K. Xu, J. Xu, J.-W. Yang, Z. Zhang, C. Bai, Z. Lu, X. Li, et al., “Odrl: A benchmark for off-dynamics reinforcement learning,” Advances in Neural Information Processing Systems, vol. 37, pp. 59 859–59 911, 2024. [36] Y. Luo, H. Xu, Y. Li, Y. Tian, T. Darrell, and T. Ma, “Algorithmic framework for model-based deep reinforcement learning with theoretical guarantees,” arXiv preprint arXiv:1807.03858, 2018.

Record · ID 381733 · SHA-256 d6a195e8349cb089
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.