JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
1
COOPO: Cyclic Offline-Online Policy Optimization Algorithm
arXiv:2605.18675v1 [cs.LG] 18 May 2026
Qisai Liu∗† , Zhanhong Jiang§ , Joshua R. Waite§ , Aditya Balu§ , Cody Fleming∗ , Soumik Sarkar∗†§ ,
agent first learning from an offline dataset and then interacting with an environment for further improvement. While conceptually appealing, emerging offline-to-online RL methods endure catastrophic forgetting [11] that shows significant performance degradation when transitioning from offline learning to online fine-tuning. Since there is a distributional shift between offline and online data, the knowledge learned offline is overwritten during online fine-tuning. In cyber-physical systems (CPS) such as autonomous driving, additive manufacturing, and energy grids, physical constraints and safety limitations make large-scale online exploration infeasible. Traditional reinforcement learning lacks guarantees when deployed in real-world applications due to its sample inefficiency and concerns about robustness. From a control perspective, such systems ensure stability and safety under limited data and uncertain dynamics, as well as resilience against cyber and communication disruptions that can compromise sensing or actuation. Recent studies have shown that data poisoning, unknown disturbances, and model manipulation can Index Terms—Offline-online leaning, Policy Optimization, Rein- mislead learning-based controllers, highlighting the necessity forcement Learning, Sample complexity, Markov decision process for algorithms that remain stable under distribution shifts or adversarial perturbations. To mitigate this issue, a recent work proposed an optimistic I. I NTRODUCTION exploration and meta adaptation (OEMA) [9] to stabilize Reinforcement learning (RL) has demonstrated compelling online exploration and reduce the distributional shift during performance and remarkable success in numerous fields over the offline-to-online transition, thus improving the sample the past decade, such as robotics [1], [2], manufacturing [3], efficiency. A more straightforward method is to blend offline medical imaging [4], and reasoning with large language models data with online interactions in the replay buffer by using (LLMs) [5]. In widespread online RL approaches such as off-policy techniques [10], [12], even without offline RL preProximal Policy Optimization (PPO) [6] and Soft Actor- training. While reducing catastrophic performance drops, these Critic (SAC) [7], an agent learns the optimal policy through methods put the efficacy and necessity of contemporary RL interactions with an environment. However, the requirement fine-tuning methods into question. To validate online RL of exploration computationally poses the issue of prohibitive fine-tuning, a more recent work [13] proposed Warm Start sample complexity. Fortunately, in many real-world cases, a rich Reinforcement Learning (WSRL) to employ a warmup phase abundance of logged data collected from experts is accessible at the beginning of online fine-tuning by utilizing a small to practitioners such that offline RL [8] provides a promising number of rollouts from the offline pre-trained policy, bridging alternative to learn polices. Without directly interacting with the distribution mismatch. Even if WSRL does not retain offline the environment, offline RL tends to learn a policy exclusively data, it is unclear how long the warm up phase should be for from a fixed dataset of pre-collected experiences. Nevertheless, different environments. Additionally, most existing methods due to the need to mitigate data drift during learning, offline have focused primarily on improving the performance when RL usually suffers from policy suboptimality, limiting its transitioning from offline pre-training to online fine-tuning, applicability in unseen tasks. The dilemma of online and offline ignoring the sample efficiency itself in fine-tuning. Thus, a RL has recently motivated work [9], [10] to combine these two question naturally arises: Can we develop an efficient offlinelines of research, whereby an optimal policy is achieved by an to-online RL algorithm to mitigate catastrophic forgetting and reduce online interactions with the environment? ∗ Department of Mechanical Engineering, Iowa State University, Ames, IA Contributions. To answer this question, we introduce COOPO 50011, USA. † Department of Computer Science, Iowa State University, Ames, IA 50011, (Cyclic Offline-Online Policy Optimization), a unified frameUSA. work that resolves the critical limitations of hybrid rein‡ Department of Industrial and Manufacturing Systems Engineering, Iowa forcement learning distribution drift during offline-to-online State University, Ames, IA 50011, USA. transitions and catastrophic forgetting of offline knowledge § Translational AI Center, Iowa State University, Ames, IA 50011, USA. Corresponding author: Soumik Sarkar (e-mail: [email protected]) during online fine-tuning. By innovatively looping between Abstract—Offline reinforcement learning struggles with distributional shift and constrained performance due to static dataset limitations, while online RL demands prohibitive environment interactions. The recent advent of hybrid offlineto-online methods bridges these paradigms but suffers from distributional shift during transitions and catastrophic forgetting of offline knowledge. We introduce COOPO (Cyclic Offline-Online Policy Optimization), a generalized framework that repeatedly cycles between constrained offline training and online finetuning. Each cycle first anchors the policy to the dataset via KL-regularized advantage-weighted offline updates to minimize distributional shift and then fine-tunes it online using any policy optimization for stable exploration. Crucially, periodically returning to offline training eliminates forgetting and drift while maximizing dataset reuse. The cyclic behavior also helps reduce the online environment interactions. Theoretically, COOPO achieves better online sample efficiency, surpassing pure online RL, with guaranteed performance improvement under standard coverage assumptions. Extensive D4RL benchmarks demonstrate COOPO reduces online interactions versus state-of-the-art hybrids while improving final returns. This looped synergy sets new efficiency and performance standards for adaptive RL.
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
Kullback-Liebler (KL)-regularized offline optimization and online adaptation (as shown in Figure 1), COOPO dynamically anchors policies to the static dataset during offline phases, constraining distributional shift via advantage-weighted updates. It interleaves short online exploration bursts with periodic returns to offline training, eliminating forgetting while maximizing dataset reuse. In the meantime, it improves both learning stability and defense against data or policy drift in safety-critical CPS environments. COOPO formalizes a cyclic synergy where offline phases provide stability and online phases enable adaptation, transforming the coverage limitations of offline RL and sample inefficiency of online RL into complementary strengths. This architecture sets a new paradigm for sample-efficient and stable adaptive learning with theoretical and empirical guarantees. When there is only one cycle, COOPO degenerates to regular offline-to-online RL. The main contributions can be summarized as: 1) We propose a generalized cyclic framework alternating KL-regularized offline training and adaptive online fine-tuning with periodic dataset anchoring to prevent distributional shift and catastrophic forgetting. 2) We theoretically show the performance difference bound, presenting per-cycle return increase under data coverage. Additionally, COOPO achieves better online sample efficiency, surpassing pure online RL algorithms. Please see Table I for method comparison. 3) We show that COOPO achieves competitive or superior performance compared to state-ofthe-art baselines on the widely-used D4RL benchmarks and demonstrate its effectiveness in the offline-to-online RL setting.
2
Return Offline Learning Cyclic Offline Learning
W/ Data Retention COOPO (Ours) WSRL Offline RL/Fine-tuning
Forgetting
Step
Fig. 1: Cyclic offline-online policy optimization: Traditional methods pre-train offline and then fine-tune online, but a single offline phase often leads to catastrophic forgetting and distributional shift. While warm-up phases can partially mitigate this, they fall short in preserving long-term stability. COOPO addresses this by leveraging cyclic synergy between offline re-training and online fine-tuning, repeatedly anchoring the policy to the offline dataset to enhance robustness and prevent forgetting. II. K EY R ELATED W ORK
Online RL. Online RL typically includes two categories, i.e., on-policy and off-policy algorithms. The key difference lies in how these methods update their policies. On-policy ones [6] update policies using data collected by their current behavior policies, ignoring any data generated by history behavior policies and leading to high sample complexity. In contrast, off-policy methods [7] allow policies to learn from Control-Theoretic and CPS Perspective. From a controlexperience produced by prior policies, resulting in high sample theoretic viewpoint, the proposed COOPO framework can efficiency. However, when real-world problems are complex, be interpreted as a closed-loop adaptive control mechanism such as transportation [14] and biological systems [15], the operating over hybrid data and physical dynamics. In traditional requirement of massive online interactions still makes online control, system stability is maintained by ensuring that feedback RL inapplicable to solving them. updates stay within a bounded region of operation. COOPO Offline RL. Different from online RL, offline RL is inaccessible achieves a similar effect through its cyclic KL-regularized to any online environment and learns policies solely from updates, which restrict policy divergence between consecutive pre-collected data. A central challenge in offline RL is value cycles, thus enforcing a form of discrete-time stability in policy overestimation, such that pessimistic updating strategies have space. Even under imperfect sensing or cyber perturbations, been developed to address the distribution shift problem [16], COOPO’s repeated realignment to the offline dataset functions [17]. Model-free offline RL adopts behavior-constrained apas a corrective stabilizing term that regularly updates or corrects proaches to regularize the learned policy to stay close to the its estimate of the system’s internal state in a robust and faultdata collecting policy [18]–[20]. A recent work found that tolerant control system. the inherent conservatism of some on-policy algorithms can help offline RL methods attenuate the overestimation and develop behavior PPO [21]. When turning to model-based TABLE I: Complexity Comparison RL, existing methods apply a parameterized model to estimate states and then update policies in a pessimistic way [22], [23]. Method Sample Complexity Though distributional shift is to some extent resolved, offline RL 4 methods cannot always ensure decent performance, particularly Online PPO Õ( Hε2 ς ) when the data quality is poor, significantly underperforming c21 Offline AWAC Õ( ε2 ) online RL approaches. Additionally, evaluating learned policies 4 c21 1-cycle Hybrid Õ( ε2 + Hε2 ς ) is also challenging if the online environment is unavailable. 3 c2 COOPO Õ( ε12 log 1ε + Hε2 ς log 1ε ) Offline-to-Online RL: The emerging offline-to-online RL has drawn considerable attention due to the synergy of H: the episode length, ς > 0: policy class advantages from offline and online RL. The mixed setting param, ε > 0: the accuracy. c1 > 0: a was introduced in [24] to learn the policy offline first from the constant related to offline learning. pre-collected dataset and then online via interactions with the environment. More recently, a calibrated Q-learning algorithm
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
was shown to be effective for online fine-tuning [25], leading to high sample efficiency. Different from value-based methods, [26] unified online and offline deep RL with multi-step onpolicy optimization without introducing extra conservatism or regularization. They leveraged diverse ensemble policies to address the distributional shift issue. Other methods, such as adaptive policy learning framework [27], policy-only RL fine-tuning [28], ensemble-based offline-to-online RL [29], and policy re-evaluation and value alignment [30], have been researched to bridge the technical gaps in offline-to-online settings. Another line of research does not require any offline pretraining, instead directly integrates offline datasets with the online training [10], [12], [31], which gives rise to questions about the efficacy of RL fine-tuning. However, a more recent work [13] has proposed a simple yet efficient fine-tuning method without retention of offline data by having a warmup phase at the beginning. To further address exploration challenges in longhorizon and sparse-reward tasks, Q-chunking [32] adopts action chunking by directly running RL in a “chunked" action space instead of a single action to maximize the sample efficiency. However, the compounding error, along with a long horizon, may negatively affect the performance. Also, determining the action length depends highly on empirical tasks, requiring manual tuning efforts. III. P RELIMINARIES Markov Decision Process. This work considers a finitehorizon Markov Decision Process (MDP) with discounted reward defined by the tuple M = (S, A, P, r, d0 , γ), where S and A represent the sets of states and actions respectively, P : S × A × S → [0, 1] is the transition probability function, r : S × A → R is the reward function, d0 is the initial state distribution of environment, and γ ∈ [0, 1] is the discount factor. A stochastic policy expressed by π(a|s) : S → A, defines a mapping from the state s to a probability distribution over actions a. Denote by h a specific time step and by H the horizon length in a trajectory. We then also define the π stationary PHstate hdistribution under the policy π as d (s) = (1 − γ) h=0 γ p(sh = s), where p signifies the probability. Reinforcement learning aims to choose a policy that is able to maximize PHthe expected discounted cumulative rewards J(π) = Eτ ∼π [ h=0 γ h r(sh , ah )], where τ ∼ π signifies a trajectory sampled according to s0 ∼ d0 , ah ∼ π(·|sh ), and sh+1 ∼ p(·|sh , ah ). Additionally, we define state value function Pthe H of the policy π as V π (s) = Eτ ∼π [ h=0 γ h r(sh , ah )|s0 = s], the state-action value function, i.e., Q-function, as Qπ (s, a) = PH t Eτ ∼π [ h=0 γ r(sh , ah )|s0 = s, a0 = a], and the critical advantage function as Aπ (s, a) = Qπ (s, a) − V π (s). This intuitively assesses how much better it is by taking action a than the average. Online RL. We focus on on-policy policy optimization algorithms, which estimate policy gradients and apply firstorder stochastic gradient ascent. PPO, one of the most popular choices, is favored for its strong performance, simplicity, and theoretical grounding via a policy improvement lower bound. It constrains policy updates using a clipping heuristic, leading
3
to widely used PPO-Clip variant [33]. At each update, the following objective is optimized: PO πh LP actor (θ) = Eπold [min(ch (θ)Â (s, a), (1) clip(ch (θ), 1 − ι, 1 + ι)Âπh (s, a))], h |sh ) where ch (θ) = ππoldθ (a (ah |sh ) , clip(o, q, v) = min(max(o, q), v). Âπh (s, a) is an estimator of the advantage function at time step h. The clipping function in PPO ensures the probability ratio between current and new policies stays within [1 − ι, 1 + ι], enforcing stable updates. guarantees a lower bound on the surrogate loss. In practice, using a small learning rate and many time steps helps PPO approximate this objective reliably. While PPO is used in COOPO, other trust-region methods (e.g., TRPO [34]) or off-policy algorithms like SAC and DDPG [35] can also be used for online fine-tuning. Offline RL. In offline RL, the agent can only access a precollected dataset with transitions D = {(st , at , sj+1 , rj )N j=1 }, generated by an unknown behavior policy πβ . Unlike online RL, offline RL aims to learn an optimal policy directly from a static dataset D. A key challenge is the distributional shift between the learned policy π and the behavior policy πβ . When π chooses actions not well represented in D, it can lead to inaccurate value estimates. To address this, a behavior constraint is commonly applied to keep π close to πβ , typically in the following form: maxπ Es∼D [Eπ(a|s) [Q(s, a)]], s.t.DKL (π||πβ ) ≤ δ, (2) where DKL (·, ·) is the Kullback-Leibler divergence, and δ > 0 is a threshold. To solve the above optimization problem, we adopt the Advantage Weighted Actor-Critic (AWAC) algorithm [24], which trains an off-policy critic and an actor under an implicit policy constraint. Following the actor-critic framework, it alternates between evaluating the policy via Qπ and improving it by maximizing a weighted objective based on the estimated advantages. With abuse of notation, we denote by Qπt (s, a) the estimated Q-function at time step t (this is episode herein). Since optimizing Qπt (s, a) is equivalent to optimizing Aπt (s, a) [24], Eq. 3 can be rewritten as πt+1 = argmaxπ Ea∼π(·|s) [Aπt (s, a)] (3) s.t.DKL (π(·|s)||πβ (·|s)) ≤ δ To this end, a non-parametric analytical solution is obtained for the actor π, subsequently followed by projecting it into the parametric policy class. We first introduce the Lagrangian by augmenting the loss in Eq. 3 with a multiplier λ > 0 such that L(π, λ) = Ea∼π(·|s) [Aπt (s, a)]+λ(δ−DKL (π(·|s)||πβ (·|s))). (4) Pertaining to closed-form solutions in [24], [36], we can 1 attain π ∗ (a|s) = Z(s) πβ (a|s)exp( λ1 Aπt (s, a)), where Z(s) = R π(a|s)exp( λ1 Aπt (s, a)) is a partition function. We then a parameterize the policy with a deep neural network θ ∈ Rn , and project the non-parametric solution into policy space. Mathematically, this is via minimizing the KL divergence between πθ and π ∗ such that argminθ Es∼D [DKL (π ∗ (·|s)||πθ (·|s))]. With some mathematical manipulations as in [36], the parameter update at time step t + 1 is expressed: 1 θt+1 = argmaxθ E(s,a)∼D [logπθ (a|s)exp( Aπt (s, a))]. (5) λ Updating θ amounts to the maximum likelihood estimation
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
Offline Dataset, D
{(s, a, r, s’)}
4
Replay Buffer, B
V, π {(s, a, r, s’)}
s, r Cyclic Offline-Online
Update π, Q, V Learning via Offline RL
a
Rollout
Update π, V
Fine-tuning via Online RL
Fig. 2: Schematic Diagram of COOPO: COOPO cyclically alternates KL-regularized offline training and trust-region online fine-tuning to eliminate distribution drift and catastrophic forgetting. When there is only one cycle, COOPO degenerates to vanilla offline-to-online RL. where the labels are accessible by weighting the state-action pairs observed in the dataset by the predicted advantages. This can simply be done by sampling (s, a) from D. Though in our work, offline training is done through AWAC, other offline RL algorithms can alternatively be used to replace AWAC, possibly with some minor algorithmic adjustments. A key requirement in cyber-physical systems is that the closed-loop controller [37] evolves in a stable and predictable manner, avoiding large or abrupt changes that may violate safety constraints. In classical adaptive control [38], this requirement is typically enforced by bounding the incremental change of controller parameters or by maintaining a sufficient stability margin. In COOPO, an analogous stability mechanism naturally emerges from the KL-regularization imposed in the offline policy update. Due to each offline iteration, COOPO restricts the divergence between the updated policy πk+1 and the reference policy πk through a KL trust region: DKL (πk+1 ∥ πk ) ≤ δ. (6) This constraint ensures that the new controller remains close to the previous one. Moreover, by Pinsker’s inequality [39], this directly bounds on the total-variation distance between the two policies: q q 1 δ (7) ∥πk+1 − πk ∥TV ≤ D (π ∥π ) ≤ KL k+1 k 2 2. Thus, policy changes across offline–online cycles remain uniformly bounded, resulting in a smooth and safe evolution of the closed-loop behavior. IV. P ROPOSED M ETHOD We start with the algorithmic framework in the sequel. Relevant notations will be defined alongside the analysis and the proofs are deferred to Appendix. COOPO. Algorithm 1 outlines the core steps of COOPO. While COOPO can theoretically adopt any RL algorithms, actorcritic methods are preferred in practice for smoother offlineto-online transitions. The actor and critic, initially trained offline, are fine-tuned online for performance gains. Unlike conventional approaches that stop after online fine-tuning risking distributional shift and catastrophic forgetting. COOPO cyclically returns to the offline dataset, anchoring the policy
to supported states and preserving prior knowledge. As shown in Figure 2, each cycle includes: Offline Phase: Optimize policies using KL-regularized advantage-weighted updates, minimizing distributional shift via a loss combining importanceweighted behavior cloning and KL penalties; Online Phase: Fine-tune policies via any exploratory policy gradient method (e.g., PPO, TRPO, SAC), ensuring stable exploration with clipping mechanisms or divergence constraints. KL-regularized AWAC with value function. We parameterize the Q-function and value-function with ϕ and ψ, initializing them from the models used in online fine-tuning. Lines 4–11 in Algorithm 1 outline the offline updates, which largely follow AWAC, with a key difference: we introduce an additional value model (Line 8) to support a seamless transition to online learning. This is essential since our online phase uses PPO, which relies on a value function. This also enables us to off calculate advantage Âπθe (s, a)) in a slightly different way (Line 9) from that used in [24]. Particularly, we also incorporate an explicit KL divergence penalty Es∼D [DKL (πθoff ||πθoffe )(·|s)] to constrain the current policy πθoffe and the next learned policy πθoff (Line 10), directly controlled by a coefficient λ. While Eq. 5, implicitly constrains the learned policy π to stay close to the behavior policy πβ , KL-regularized AWAC explicitly enforces a trust region, ensuring updates remain near previous policies and dataset-supported actions. Unlike the implicit constraint in Eq. 5 where λ only scales advantages, explicit KL control reduces sensitivity to hyperparameters and better prevents distributional shift and catastrophic forgetting. Moreover, incorporating a value function Vψoff helps avoid out-of-distribution (OOD) action estimates, similar to IQL, enabling more robust learning from static datasets. Transition to online RL. Line 12 shows the critical transition from offline to online settings, involving the value and policy models. Concretely, the last iterates of KL-regularized AWAC of ψ and θ are the initialization of online RL, for which we resort to the widespread PPO algorithm (Lines 14-20). In Line 21, the cyclic models at the k-th cycle are updated by using fine-tuned value and policy models. As the parametric Q model is not fine-tuned in the online phase, the cyclic Q model is obtained by directly adopting the last iterate of Q model in the offline phase. Non-trivial synergy in COOPO. COOPO cyclically integrates KL-regularized AWAC and trust-region PPO to address limitations of naive hybrids. To maintain critic stability across cycles, COOPO freezes the offline-trained Q-function during the online PPO phase. We empirically observed that jointly updating the critic online increases variance and value drift due to distribution mismatch between offline and online states. Freezing the critic ensures a consistent advantage baseline and aligns with findings in prior offline-to-online RL literature, where maintaining a stable value estimator is essential for preventing policy collapse during online fine-tuning. The KL constraint anchors the policy to the offline dataset, reducing distributional drift, while PPO’s clipped updates enable stable online fine-tuning. The KL constraint plays a dual role: (i) it bounds inter-cycle policy deviation, thereby ensuring stable improvement at each step, and (ii) it mitigates overfitting to the offline dataset by preventing aggressive policy updates
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
5
that diverge from known-safe behaviors. This mechanism is Algorithm 1 COOPO particularly important in CPS settings, where abrupt controller 1 Input: Offline dataset D = {(s, a, s′ , r)j }, cycle K, offline epochs E, online episodes T , KL coefficient λ, horizon H changes can violate safety margins. Repeated offline phases 2 Initialize: πθ1 , Qϕ1 , Vψ1 progressively average advantage errors and stabilize training. 3 Output: πθK This closed-loop alternation forms a feedback mechanism: each 4 for k = 1, . . . , K do off off cycle resets policy divergence with KL penalties and refines 5 Qoff ϕ1 ← Qϕk , πθ1 ← πθk , Vψ1 ← Vψk policy using online signals. This synergy lets AWAC stabilize Offline learning 6 for e = 1, . . . , E do PPO’s exploration and PPO guide AWAC updates. COOPO 7 Sample batch (s, a, s′ , r) ∼ D also amortizes offline data cost across cycles, thus improving ′ ′ 8 y ← r(s, a) + γEs′ ,a′ Qoff ϕe (s , a ) efficiency. Theoretically, it guarantees performance bounds and off 9 ϕe+1 ← arg minϕ ED (Qϕ (s, a) − y)2 logarithmic cycle complexity; empirically, it achieves strong 10 ψe+1 ← arg minψ ED (Vψoff (s) − y)2 sample efficiency on benchmarks, demonstrating its effective off off 11 Âπθe (s, a) ← Qoff ϕ (s, a) − Vψ (s) design. off To isolate the effect of the static offline dataset, online rollouts are intentionally excluded during the offline update. This design ensures that the offline update has enough contribution into the current policy by using the better data : when the online policy becomes locally trapped or drifts due to limited exploration, the offline step pulls the agent back toward high-quality behaviors supported by the dataset. By avoiding the accumulation of mixed with the online samples, COOPO prevents dataset degradation and maintains a stable reference distribution, which give a clearer attribution of performance gains to the cyclic update mechanism. Incorporating replayed online samples remains feasible and may further accelerate adaptation, which we leave for future work. Theoretical Analysis. To start with the analysis, we denote by πk := πθk the policy at the start of cycle k. To ease the notation, we also denote by πk+1/2 and πk+1 the policies after the offline phase and online phase of cycle k, respectively. The state distribution in the offline dataset D is defined as ρβ (s). In both offline and online phases, the advantage function plays a central role in evaluating predicted actions, With abuse of notations, we still adopt Aπk (s, a) (or Âπk (s, a)) as the advantage (or estimated advantage) throughout the rest of paper. In a rigorous sense, the notations of advantage functions differ slightly as in Algorithm 1, but we unify them for the ease of analysis. Analogously, we also use the same value to upper bound advantages in these two phases, i.e., ϵk = maxs,a |Aπk (s, a)|. Throughout the analysis, O(·) represents the standard big O notation for complexity. Õ(·) notation simplifies O(·) notation by ignoring logarithmic factors. In what follows, we present a key assumption on the distribution shift coefficient C, which helps characterize the bound.
12
1 πθ θe+1 ← arg max Es,a∼D log πθoff (a|s) exp  e (s, a) θ λ h i − λEs∼D DKL πθoff (·|s)∥πθoff (·|s) e
19
Vψon ← Vψoff , πθon ← πθoff 1 1 E E Online fine-tuning for t = 1, . . . , T do Collect trajectories Rt = {τi } using πθon t Compute reward-to-go R̂on h within horizon H π Compute advantages Âhθt using Vψon t Update policy parameterh θt+1 via Eq. (1) i PH 2 on 1 ψt+1 ← arg minψ ERt H h=0 (Vψ (sh ) − R̂h )
20 21 22
Qϕk+1 ← Qoff ϕE Vψk+1 ← Vψon T πθk+1 ← πθon T
13 14 15 16 17 18
Lemma 1. With KL-AWAC objective function πnew = argmaxπ Es,a∼D [logπ(a|s) · w(s, a)] − λEs∼D old (·|s)], [DKL (π(·|s)||π where w(s, a) = exp
Aπ
(8)
old
(s,a) λ
is the advantage weight,
and λ > 0 is a hyperparameter to control the regularization length, we have the following relationship Ea∼π∗ (·|s) [Aπold (s, a)] = λDKL (π ∗ ||πold )(s) + λlogZ(s), (9) Aπold (s,a) where Z(s) = Ea∼πold (·|s) exp . λ
Proof. The KL divergence term can expand to DKL (π(·|s)||πold (·|s) = Ea∼π(·|s) [logπ(a|s) − logπold (a|s)]. Substituting it into the objective yields η(π) = Es,a∼D [logπ(a|s) · w(s, a)] (10) − λEa∼π(·|s) [logπ(a|s) − logπold (a|s)]. ∗ We now solve the optimal policy π by taking the functional Assumptionπ 1. There exists a constant 1 ≤ C < ∞ such that derivative of η(π) with respect to π(·|s) and set to zero: k (s) w(s, a) C = maxs dρβ (s) if for all state s with dπk (s) > 0, ρβ (s) > 0. dη = − λ(logπ(a|s) − logπold (a|s) + 1) + ν(s) = 0, dπ π(s|a) Assumption 1 bounds the distance between the visitation (11) R distributions of the behavior policy πβ and any learned policy where ν(s) us a Lagrangian multiplier enforcing π(a|s)da = πk , capturing the data coverage to ensure their closeness, similar 1. Solving for π(a|s), we have Aπold (s, a) ν(s) + λ to single policy concentrability in prior work [40], [41]. In π ∗ (a|s) ∝ πold (a|s)exp( )exp(− ). (12) Algorithm 1, the offline learning algorithm is KL-regularized λ λ Z(s): we then have AWAC, instead of AWAC, which leads to Eq. 4. Hence, we Define the normalization constant as 1 Aπold (s, a) ∗ will show under the KL-regularized AWAC, such a relationship π (a|s) = πold (a|s)exp , (13) Z(s) λ still holds. We present an auxiliary technical lemma to reveal πold this. where Z(s) = Ea∼πold (·|s) exp A λ(s,a) This resembles
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
the same formula in the analysis in [36]. We then take the log on both sides of π ∗ (a|s), obtaining Aπold (s, a) logπ ∗ (a|s) = logπold (a|s) + − logZ(s). (14) λ Now we compute the KL divergence: DKL (π ∗ ||πold )(s) = Ea∼π∗ (·|s) [logπ ∗ (a|s) − logπold (a|s)] πold A (s, a) = Ea∼π∗ (·|s) − logZ(s) . λ (15) Rearranging to isolate the expected advantage completes the proof.
6
Corollary 1 in [43] J(π ′ ) − J(π) ≥ Es∼dπ [Ea∼π′ (·|s) [
1 Aπ (s, a) 1−γ
(21) ′ 2γϵπ ′ max D (π ||π)(s)]]. s TV (1 − γ)2 ′ Setting π = πk+1/2 and π = πk yields the following relationship: 1 J(πk+1/2 ) − J(πk ) ≥ Es∼dπk [Ea∼πk+1/2 [ Aπk (s, a) 1−γ 2γϵk ζk ]]. − (1 − γ)2 (22) Substituting Eq. 19 into Eq. 22, we have 1 J(πk+1/2 ) − J(πk ) ≥ (λEs∼dπk [DKL (πk+1/2 ||πk )(s)] 1−γ 2γϵk ζk We next state the lower bound for the difference between − ϵadv ) − E π [E ]], s∼d k a∼πk+1/2 [ k (1 − γ)2 πk and πk+1 in COOPO. (23) Theorem 1. (Performance Difference Bound) Let Assumption 1 which is the bound after the offline phase. We next proceed hold. The advantage error in the offline phase is defined such for the online phase, where we use PPO to update πk+1/2 to πk πk πk+1 . The PPO update is designed to ensure a trust region. that ϵadv k = maxs,a |Â (s, a) − A (s, a)| for any policy πk . Hence, for any cycle k, two consecutive policies πk , πk+1 Following Theorem 1 in [34] and Proposition 1 in [43], we obtain the following relationship produced by COOPO, satisfy the following: J(πk+1 ) − J(πk+1/2 ) ≥ λ J(πk+1 ) − J(πk ) ≥ Goff (16) k − Λk , (24) 4ϵk γ 1−γ − maxs DT V (πk+1 ||πk+1/2 )(s). 2 where (1 − γ) π Goff Combining Eqs. 23 and 24 obtains the desirable result. k = Es∼d k [DKL (πk+1/2 ||πk )(s)], ϵadv 4ϵk γαk2 2γϵk ζk k + Es∼dπk [Ea∼πk+1/2 [ ]] + , 2 1−γ (1 − γ) (1 − γ)2
−
Theorem 1 suggests that the error term Λk is primarily dictated by the advantage estimation error, the total variation distance between πk and πk+1/2 , and the total variation distance αk = between π ζk = maxs DT V (πk+1/2 ||πk )(s), k+1/2 and πk+1 . The first one can be connected maxs DT V (πk+1 ||πk+1/2 )(s). DT V (·||·) is the total variation with KL divergence as it is in the offline phase with the distance. exact constraint in Eq. 3 by using the Pinsker’s inequality [44]. However, the second one cannot be connected to KL divergence, Proof. We divide the proof into two parts: offline phase requiring an assumption to upper bound it. In this context, we (KL-reguarlized AWAC) and offline phase (PPO). Recalling follow the practice from [34] such that we use a constant α > 0. Lemma 1, we obtain the following relationship: Eq. 16 implies the policy performance difference between two Es∼dπk [Ea∼πk+1/2 [Âπk (s, a)]] (17) consecutive steps is attributed to differences in both offline = Es∼dπk [λ · DKL (πk+1/2 ||πk )(s) + λlogZ(s)], and online phases. To further quantify the cycle complexity where Z(s) = Ea∼πk [exp( λ1 Âπk (s, a))]. Now we relate this for K, we assume that the advantage estimation error ϵadv k is to the true advantage such that bounded by ϵ̄adv > 0 and that the maximum advantage value Es∼dπk [Ea∼πk+1/2 [Aπk (s, a)]] = Es∼dπk [Ea∼πk+1/2 [Âπk (s, a)]] ϵk is bounded by ϵ̄ > 0. We also define the suboptimality as ∆k = J(π ∗ )−J(πk ), where π ∗ is the optimal policy. In Eq. 16, − Es∼dπk [Ea∼πk+1/2 [Âπk (s, a) − Aπk (s, a)]]. off (18) we observe the positive gain from offline phase, Gk , which is assumed to be at least propositional to the current suboptimality: The second term of the last equality can be bounded by ϵadv 1−γ k Goff k ≥ κ∆k , where 0 < κ ≤ λ . We briefly justify why this such that π is a reasonable assumption in this work. When the policy is Es∼dπk [Ea∼πk+1/2 [A k (s, a)]] far from optimal (large ∆k ), the offline phase should yield adv ≥ Es∼dπk [λ · DKL (πk+1/2 ||πk )(s) + λlogZ(s)] − ϵk . significant improvements (large Goff k . Conversely, when the (19) policy is near optimal (small ∆ ), improvements are harder k Note that logZ(s) ≥ 0 because exp( λ1 Âπk (s, a)) ≥ 0 and by to come by and thus Goff is small. Hence, Goff k k scales with Jensen’s inequality. To construct the performance improvement, ∆k . It holds empirically when the dataset D covers states with we leverage a well-known Performance Difference Lemma [42] positive advantages relative to πk . While sparse rewards or poor as follows: coverage can violate this, optimistic exploration during online 1 J(πk+1/2 )−J(πk ) = Es∼dπk+1/2 [Ea∼πk+1/2 [Aπk (s, a)]]. phases may likely restore validity by adding diverse rollouts to 1−γ (20) D and guiding it to include high-advantage transitions. With But this expectation is under dπk+1/2 instead of dπk . We then these in hand, the cycle complexity is stated in the following connect these distributions using the following inequality from main result. Λk =
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
Theorem 2. (Cycle Complexity) With definitions in Theorem 1, and ϵadv ≤ ϵ̄adv , ϵk ≤ ϵ̄, αk ≤ α, Goff k k ≥ κ∆k , COOPO converges to the neighborhood of π ∗ in a linear converλκ K ) ∆0 + Λ(1−γ) gence rate, i.e., ∆K ≤ (1 − 1−γ λκ , where adv
√
2
2δϵ̄ 4ϵ̄γα ϵ̄ Λ = 1−γ + Cγ (1−γ)2 + (1−γ)2 .
Proof. Based on Eq. 16, we have λκ ∆k + Λk . (25) 1−γ The last inequality is due to the assumption Goff k ≥ κ∆k . For Λk , based on the conditions, it is immediately obtained that ϵ̄adv 4ϵ̄α2 γ ϵ̄adv 4ϵ̄k αk2 γ k ≤ . ≤ , 1−γ 1 − γ (1 − γ)2 (1 − γ)2 For the second term in Λk , we would like to bound it with the KL divergence constraint in Eq. 3. However, the expectation is over the distribution dπk (s), instead of the dataset distribution ρD (s) as in the offline objective. Based on Assumption 1, we know that, for any function f (s) ≥ 0 Z J(π ∗ ) − J(πk+1 ) ≤ J(π ∗ ) − J(πk ) −
Es∼dπk [f (s)] = dπk (s)f (s)ds Z Z (26) dπk (s) f (s)ds ≤ C ρD (s)f (s)ds. = ρD (s) ρD (s) Thus, Es∼dπk [f (s)] ≤ CEs∼ρD [f (s)]. This results in 2γϵk ζk ]] ≤ Es∼dπk [Ea∼πk+1/2 [ (1 − γ)2 (27) 2γϵk ζk ]]. ≤ CEs∼ρD [Ea∼πk+1/2 [ (1 − γ)2 Pinsker’s inequality leads to r DKL (πk+1/2 ||πk )(s) ζk ≤ , 2 which attains √ 2δCγϵ̄ 2γϵk ζk ]] ≤ . (28) Es∼dπk [Ea∼πk+1/2 [ (1 − γ)2 (1 − γ)2 Hence, the following relationship holds λκ ∆k+1 ≤ (1 − )∆k + Λ. (29) 1−γ Applying the last inequality iteratively from k = 0, ..., K completes the proof.
Theorem 2 implies the linear convergence to the neighborhood of π ∗ , which typically appears in the gradient-based optimization with smooth and strongly convex objectives. However, in this work, we do not have such an assumption for the objective. Suppose that the desirable accuracy is ε for COOPO, i.e., ∆K ≤ ε. Using the conclusion from Theorem 2, λκ K ) ∆0 ≤ 2ε and Λ(1−γ) ≤ 2ε , it is obtained that (1 − 1−γ λκ ελκ which is equivalent to Λ ≤ 2(1−γ) . Therefore, we have
7
phase involves training for E epochs on a fixed dataset D of size |D|, resulting in E|D| offline samples per cycle and a total of KE|D| offline samples over K cycles. Since the same dataset is reused, the key question is how to set E. Offline training must be sufficiently accurate to keep error terms in the performance bound under control, particularly the advantage estimation error ϵ̄adv , which which must be small enough to ensure Λ to be O(ε). From standard supervised learning convergence, if we use firstorder stochastic optimization method (which we have used in COOPO), thep error after E epochs satisfies the relationship [45]: ϵ̄adv = O(1/ E|D|), from which we can obtain the offline samples per cycle. In each cycle, the online phase collects T samples such that the total online samples over K cycles is KT . In this phase, we run a trust-region for a few episodes to control the error in the improvement. Particularly, we require that the online phase does not degrade the policy significantly and improves upon the offline policy. From the online complexity of policy gradient methods [46], [47], they p ensure that the online suboptimality εon satisfies εon = Õ( H 3 ς/T ), where H := 1/(1 − γ) is the effective horizon and ς > 0 is a policy class parameter. For example, ς = Õ(|S||A|) if the parametric model for the policy is neural network. Based on this, we can obtain the online sample complexity per cycle. We are now ready to present the main result to characterize the total sample complexity in the following. Theorem 3. (Total Sample Complexity) Suppose that the offline p error satisfies ϵ̄advp = O(1/ E|D|) and that the online error satisfies εon = Õ( H 3 ς/T ). With the cycle complexity from Theorem 2, the total sample complexity incurred by COOPO is (1−γ)c21 (1−γ)H 3 ς 1 1 Õ λκε2 log ε + λκε2 log ε by setting εon = O(ε) and ϵ̄adv = O(ε), where c1 > 0 is a constant on offline setting. Proof. We have known that in the offline phase, the sample processed per cycle is E|D|. As ϵ̄adv = O( √ 1 ), we can E|D|
obtain ϵ̄adv ≤ √ c1
E|D|
. Therefore, it is immediately attained that
c2
E ≥ (ϵ̄adv )12 |D| . Hence, the per cycle offline sample complexity cycle is Coff =
c21 (ϵ̄adv )2 .
Additionally, for the online phase, the q 3 online error is required to satisfy εon = Õ( HT ς ) such that 3
cycle Con = Õ( Hε2 ς ). To ensure that Λ = Õ(ε), we let εon = O(ε) on adv and ϵ̄ = O(ε), which lead to the following total sample complexity cycle cycle Ctotal = KCoff + KCon 2 3 log(2∆0 /ε) 1−γ 2∆0 (30) K ≥ log(1/(1− κλ ) ≈ κλ log( ε ). This shows the cycle 1−γ 2∆0 c1 H ς 1−γ = log O + Õ . 2 2 λκ ε ε ε complexity is logarithmic in 1ε , which is favored in general. ελκ The last equality is because of the cycle complexity obtained To ensure the relationship Λ ≤ 2(1−γ) and the per-cycle gain in Theorem 2. The desirable result is immediately obtained by κλ 1−γ to be positive and bounded away from zero, practically simplifying the bound. speaking, if the offline dataset is good (high coverage leads to C → 1) and the online phase is stable, these two are expected Theorem 3 reveals the total sample complexities over K to hold. cycles of offline and online fine-tuning, learning (1 − γ)c21 1 (1 − γ)H 3 ς 1 So far, we have quantified the cycle complexity K of Õ log + log . 2 2 COOPO’s cyclic framework, but the total sample complexity λκε ε λκε ε} | {z } | {z dependent on both offline and online phases remains to of f line online be determined. We now present a result that characterizes Surprisingly, the offline sample complexity is independent COOPO’s key sample complexity. In each cycle, the offline of |D| as the data reuse amortizes the cost. Both offline
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
8
method can likely be beneficial for the performance improvement, we tune them manually in this work. TABLE III: Hyperparameters
Fig. 3: Mujoco environments used in this work for evaluation: Half-cheetah, Hopper, and Walker. and online phases scale with Õ( ε12 log 1ε ), which matches the complexity bound in [48]. The bound reveals that although KL regularization appears only in the offline phase, it affects sample complexity in both phases. A large λ reduces cost but leads to conservative update and lower κ, while a small λ accelerates learning but increases online sample use, highlighting a key trade-off between efficiency and learning performance. From Table I, we know that pure online PPO has sample 4 complexity of Õ( Hε2 ς ), COOPO’s online phase reduces by Hλκ factor (1−γ)log(1/ε) . Also, COOPO outperforms PPO in terms of horizon scaling with only H 3 , which makes it suitable for long horizon (H ≫ 1). This stems from the offline warm start. Theorem 3 also implies that online term dominates in low-ε regime, consistent with empirical results. V. N UMERICAL R ESULTS We evaluate COOPO on simulators to assess: (1) its performance against offline and offline-to-online RL baselines; (2) its asymptotic performance; (3) the role of offline–online cycles; (4) its ability to reduce online interaction compared to PPO; and (5) the impact of the multiplier λ in offline learning. Experimental details are provided in the sequel. Benchmark anvironments. Figure 3 shows the environments we resort to validate COOPO in this work, including Halfcheetah, Hopper, and Walker. While demonstrated in MuJoCo locomotion tasks, the cyclic offline-online paradigm generalizes to CPS domains such as adaptive robotic control, energy management, or manufacturing scheduling, where real-time policy refinement must go along with safety and latency constraints. Please see [49] for more information about these environments. Model architecture. In this work, we follow the standard setup as widely used in RL domain to parameterize the actor and critic networks. Specifically, they are all multi-layer perceptron (MLP) models. The architecture is shown in Table II. TABLE II: MLP model architecture Parameter
Value
# of layers # of hidden units per layer Activation function
4 256 ReLU
Hyperparameters. The hyperparameter setting is in Table III. Note that in this context we summarize key hyperparameters in RL setting. Though a hyperparameter optimization
Hyperparameter.
Value
Optimizer Learning rate Batch size (off) Batch size (on) # of epochs (off) size of dataset (off) # of epochs (on) # of episodes (on) # of cycles Initialization Episode length (on) buffer size (on) discount factor beta KL weight
Adam 3e-4 512 64 100 1M 5 5 500 Random 1M 512 0.99 0.99 0.05
Computing Infrastructure. All experiments were conducted on a workstation equipped with an NVIDIA GeForce RTX 4090 GPU and an Intel(R) Core(TM) i7-14700 CPU, with each run utilizing approximately 1.4 GiB of GPU memory. The system was running Ubuntu 22.04.4 LTS (64-bit) (Distributor ID: Ubuntu, Release: 22.04). Offline datasets were obtained from the D4RL benchmark suite, while online training was performed using Gym environments (version 0.23.1). The implementation leveraged standard deep learning and reinforcement learning frameworks, including PyTorch (version 2.4.1), and other supporting libraries such as NumPy and Matplotlib. Comparative study. We compare COOPO with multiple offline and offline-to-online algorithms. For offline RL, the methods encompass Behavior Cloning (BC) [50], Onestep RL [51], TD3+BC [52], Conservative Q-Learning (CQL) [16] and Implicit Q-Learning (IQL) [17]. For offline-to-online RL, we also include multiple popular baseline methods that have been adopted in many works: Cal-QL [25], Advantage Weighted Actor-Critic (AWAC) [24], Uni-O4 [26], Online Decision Transformer (ODT) [53], PEX [54], and Off2ON [55]. We first evaluate whether COOPO outperforms or remains competitive with offline and offline-to-online approaches. As shown in Table IV COOPO achieves top performance on 7 out of 9 tasks, surpassing all one-step methods like IQL and Onestep RL, which often yield conservative policies due to behavior constraints. COOPO also outperforms pure offline RL methods, highlighting the benefit of online finetuning for effective exploration. Compared to offline-to-online baselines, COOPO consistently excels across most environments, remaining competitive with Uni-O4 on select tasks. These gains stem from its cyclic synergy between offline and online learning, which mitigates distributional shift and catastrophic forgetting. Figure 4 further confirms COOPO’s superior asymptotic performance demonstrated in HalfCheetah compared to AWAC and IQL, aligning with trends in Figure 1
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
9
TABLE IV: Reward values of various algorithms on the D4RL locomotion benchmarks using medium, medium-replay, and medium-expert datasets. Each value is the mean return over seeds, averaged across the final evaluation episodes. Dataset
BC
Onestep RL
TD3+BC
CQL
IQL
Cal-QL
AWAC
Uni-O4
ODT
PEX
Off2ON
COOPO (Ours)
halfcheetah-medium hopper-medium walker2d-medium halfcheetah-medium-replay hopper-medium-replay walker2d-medium-replay halfcheetah-medium-expert hopper-medium-expert walker2d-medium-expert
4233.3 885.8 2748.3 1967.7 308.5 1037.3 5192.7 1269.7 4936.6
4828.4 995.6 2987.6 2053.5 1620.1 1972.6 8949.3 2484.6 5189.1
4822.7 990.6 3052.0 2423.5 1016.0 3252.7 8674.5 2359.7 5056.0
4371.2 977.7 2651.3 2477.9 1580.7 3073.5 8762.8 2534.6 4996.3
4520.6 1092.3 2718.4 3368.4 913.2 2076.4 8278.3 2201.3 5033.0
4754.7 1606.1 3083.6 1795.2 714.8 2271.8 8782.9 2596.7 5102.7
4531.1 967.3 2910.2 2189.6 624.2 1075.8 3976.5 1339.3 4643.2
5277.5 1740.8 3289.3 2911.2 1197.7 3072.3 8859.3 2683.2 5421.3
4235.4 781.7 2625.7 2087.6 698.8 1380.2 3566.1 1556.3 3040.0
5084.2 937.8 2922.4 2995.8 360.3 2503.7 2778.3 2196.7 3205.1
3878.4 1625.0 2413.6 2747.4 322.1 553.6 4333.0 2697.0 3934.8
4947.1 2014.7 3570.4 4602.5 1993.8 3491.2 9242.2 3575.8 5042.8
0 500 1000
1500 1000 500 0
1000 1250 1500 1750 2000 2250 2500 2750
Step
3000
Online Offline
2000
500
1500
4000
0
5(a): HalfCheetah
2000
3500 3000
1000
AWAC COOPO IQL
0 0
20000
40000
60000
80000
Gradient Steps
100000 120000 140000
Fig. 4: Average return vs. gradient steps on HalfCheetah for AWAC, COOPO, and IQL. COOPO achieves the best final performance with improved sample efficiency, while AWAC and IQL converge faster but plateau earlier.
Episode Reward
Average Return
2500
Online Offline
Episode Reward
5000
1000
Episode Reward
and emphasizing the benefits of data reuse and stable policy refinement.
250
500
750
1000
Step
1250
1500
1750
5(b): Hopper Online Offline
2500 2000 1500 1000 500 0 0
1000
2000
Step
3000
4000
5(c): Walker2D
Fig. 5: Illustration of the training dynamics of COOPO on three D4RL medium datasets. The red solid lines represent the online training segments, while the blue dashed lines indicate the offline updates bridging the gaps between online phases. For clarity, the figure is using a single training seed from each environment.
Offline and Online Training Phases in COOPO. Figure 5 illustrates the training dynamics of our proposed method COOPO, which alternates between offline and online training phases, on the HalfCheetah, Hopper, and Walker2D tasks using improves performance, particularly in the low-ε regime near the D4RL medium datasets. In each plot, the redred solid lines optimality, supporting the theoretical insights from Theorem 3. represent online training segments, while the blueblue dashed This highlights the greater impact of online updates in the lines indicate offline updates between online training phases. low-ε regime, as predicted by the sample complexity analysis. Across all three environments, the offline updates consistently Conversely, when the frequency of online fine-tuning is fixed, improve policy performance. Although there may be rare increasing offline training epochs boosts performance by cases where offline updates momentarily hinder online training, anchoring the policy more effectively to the static dataset, overall they help the online phase achieve higher rewards. thus mitigating distributional shift and catastrophic forgetting. The results highlight the importance of COOPO’s alternating Interestingly, when offline epochs increase (e.g., from E = 75 architecture between offline and online training phases to to E = 100) while online episodes decrease (e.g., T = 15 to T = 5), performance improves early in training but plateaus achieve better performance and more efficient learning. Ablation studies. We now examine the interplay between later, indicating diminishing returns from offline updates offline and online training phases in each cycle. As shown in alone. In contrast, increasing online fine-tuning leads to more Figure 6 and outlined in Algorithm 1, each COOPO cycle uses sustained gains, as seen when comparing T = 20, E = 50 with a small, fixed number of online and offline rollout episodes (T = T = 5, E = 100. These findings demonstrate COOPO’s capac5 to 25, E = 25 to 100), depending on the task. This bounded ity to reduce online environment interaction by reusing offline sampling budget reduces total environment interactions by data. However, if offline training is too sparse, even frequent approximately 50% compared to PPO while still allowing online fine-tuning may not be sufficient, which is evident when comparing T = 20, E = 25 with T = 5, E = 100. Therefore, meaningful policy refinement. We evaluate performance across various combinations of to deploy COOPO effectively, it is beneficial to emphasize offline epochs (E) and online fine-tuning steps (T ). When offline training in the early (high-ε) phase and prioritize online the offline training epochs are fixed at 100, increasing the fine-tuning in the later (low-ε) phase. frequency of online fine-tuning (e.g, T = 5 vs. T = 25) Figure 7 compares the number of trajectories required by
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
10
COOPO and PPO to achieve high performance. COOPO consistently attains higher rewards using significantly fewer trajectories, highlighting its sample efficiency. In contrast, PPO suffers from high sample complexity due to its on-policy nature. These results underscore the effectiveness of COOPO’s offline data reuse and the benefits of the proposed looped synergy. Moreover, they empirically validate the theoretical claim from Theorem 3 that COOPO reduces the required online interactions by a constant factor. We now answer the last question of how the key λ value in the offline learning affects the overall performance.
stable and consistent policy updates compared to the more variable results of the other two cases. Limitations. While COOPO enhances sample efficiency, its effectiveness is constrained by the quality and coverage of the offline dataset poorly curated data may trap policies in local optima, limiting online gains. The approach also struggles with sparse-reward environments due to limited exploration during brief online phases, hindering novel discovery. Additionally, computational overhead from repeated offline retraining increases wall-clock time, and static datasets cannot dynamically incorporate online experience, capping adaptability in nonstationary tasks.
5000
VI. C ONCLUSIONS
Reward
4000 3000 2000
T=5, E=100 T=15, E=75 T=20, E=25 T=20, E=50 T=20, E=100 T=25, E=100 80 100
1000 0 0
20
40
Cycle
60
Fig. 6: Training performance across different online-offline cycle combinations on the HalfCheetah environment. Increasing offline cycles generally improves stability and reduces the need for extensive online updates.
Evaluation Reward
5000
COOPO PPO
4000 3000 2000
R EFERENCES
1000 0 0
250
500
750
1000
1250
Trajectory Number
1500
1750
2000
Fig. 7: Training performance against trajectory number between COOPO vs. PPO on the HalfCheetah environment.
509.3
500
4000
549.9
544.9 482.6
400
3000
Variance
Episode Reward
5000
2000 =1 =3 =9
1000 0 0
COOPO introduces a cyclic synergy between offline and online reinforcement learning, obtaining stable policy improvement under limited interaction budgets. By alternating KLregularized offline optimization with trust-region online finetuning, the framework mitigates distributional drift, prevents catastrophic forgetting, and provides a stabilizing corrective mechanism when online exploration becomes trapped or noisy. Theoretically, COOPO offers guarantees on bounded reduced online interaction and a performance lower bound, which provides a principled foundation for reliable deployment in cyber-physical systems. Empirically, COOPO achieves consistently strong returns and sample efficiency across diverse benchmarks, outperforming recent state-of-the-art offline-toonline RL approaches. These results highlight COOPO as a practical and robust solution for safe and adaptive learning in real-world CPS.
20
40
Cycles
60
80
300 200
279.1 193.9
Final Average
100 100
8(a): Training curves for different λ values. Shaded regions show reward ranges across seeds.
0
=1
=3
=9
8(b): Final reward variance vs. average variance across seeds for each λ.
Fig. 8: Performance comparison in the HalfCheetah environment for different λ values, averaged over multiple seeds. Figure 8 (a) shows that all λ values eventually reach high episode rewards, but smaller λ (e.g., λ = 1) leads to faster convergence, while larger (e.g., λ = 9) results in more conservative updates. Figure 8 (b) reveals that λ = 3 strikes a balance by achieving relatively fast convergence with the lowest final and average variance across seeds, indicating more
[1] L. Brunke, M. Greeff, A. W. Hall, Z. Yuan, S. Zhou, J. Panerati, and A. P. Schoellig, “Safe learning in robotics: From learning-based control to safe reinforcement learning,” Annual Review of Control, Robotics, and Autonomous Systems, vol. 5, no. 1, pp. 411–444, 2022. [2] B. Singh, R. Kumar, and V. P. Singh, “Reinforcement learning in robotic applications: a comprehensive survey,” Artificial Intelligence Review, vol. 55, no. 2, pp. 945–990, 2022. [3] C. Li, P. Zheng, Y. Yin, B. Wang, and L. Wang, “Deep reinforcement learning in smart manufacturing: A review and prospects,” CIRP Journal of Manufacturing Science and Technology, vol. 40, pp. 75–101, 2023. [4] S. K. Zhou, H. N. Le, K. Luu, H. V. Nguyen, and N. Ayache, “Deep reinforcement learning in medical imaging: A literature review,” Medical image analysis, vol. 73, p. 102193, 2021. [5] D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi et al., “Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,” arXiv preprint arXiv:2501.12948, 2025. [6] J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” arXiv preprint arXiv:1707.06347, 2017. [7] T. Haarnoja, A. Zhou, K. Hartikainen, G. Tucker, S. Ha, J. Tan, V. Kumar, H. Zhu, A. Gupta, P. Abbeel et al., “Soft actor-critic algorithms and applications,” arXiv preprint arXiv:1812.05905, 2018. [8] R. F. Prudencio, M. R. Maximo, and E. L. Colombini, “A survey on offline reinforcement learning: Taxonomy, review, and open problems,” IEEE Transactions on Neural Networks and Learning Systems, 2023. [9] S. Guo, L. Zou, H. Chen, B. Qu, H. Chi, P. S. Yu, and Y. Chang, “Sample efficient offline-to-online reinforcement learning,” IEEE Transactions on Knowledge and Data Engineering, vol. 36, no. 3, pp. 1299–1310, 2023. [10] Y. Song, Y. Zhou, A. Sekhari, J. A. Bagnell, A. Krishnamurthy, and W. Sun, “Hybrid rl: Using both offline and online data can make rl efficient,” arXiv preprint arXiv:2210.06718, 2022.
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021
[11] Y. Luo, J. Kay, E. Grefenstette, and M. P. Deisenroth, “Finetuning from offline reinforcement learning: Challenges, trade-offs and practical solutions,” arXiv preprint arXiv:2303.17396, 2023. [12] P. J. Ball, L. Smith, I. Kostrikov, and S. Levine, “Efficient online reinforcement learning with offline data,” in International Conference on Machine Learning. PMLR, 2023, pp. 1577–1594. [13] Z. Zhou, A. Peng, Q. Li, S. Levine, and A. Kumar, “Efficient online reinforcement learning fine-tuning need not retain offline data,” arXiv preprint arXiv:2412.07762, 2024. [14] Q. Wang and C. Tang, “Deep reinforcement learning for transportation network combinatorial optimization: A survey,” Knowledge-Based Systems, vol. 233, p. 107526, 2021. [15] E. O. Neftci and B. B. Averbeck, “Reinforcement learning in artificial and biological systems,” Nature Machine Intelligence, vol. 1, no. 3, pp. 133–143, 2019. [16] A. Kumar, A. Zhou, G. Tucker, and S. Levine, “Conservative q-learning for offline reinforcement learning,” Advances in neural information processing systems, vol. 33, pp. 1179–1191, 2020. [17] I. Kostrikov, A. Nair, and S. Levine, “Offline reinforcement learning with implicit q-learning,” arXiv preprint arXiv:2110.06169, 2021. [18] Y. Wu, G. Tucker, and O. Nachum, “Behavior regularized offline reinforcement learning,” arXiv preprint arXiv:1911.11361, 2019. [19] S. Fujimoto, D. Meger, and D. Precup, “Off-policy deep reinforcement learning without exploration,” in International conference on machine learning. PMLR, 2019, pp. 2052–2062. [20] D. Ghosh, A. Ajay, P. Agrawal, and S. Levine, “Offline rl policies should be trained to be adaptive,” in International Conference on Machine Learning. PMLR, 2022, pp. 7513–7530. [21] Z. Zhuang, K. Lei, J. Liu, D. Wang, and Y. Guo, “Behavior proximal policy optimization,” arXiv preprint arXiv:2302.11312, 2023. [22] R. Kidambi, A. Rajeswaran, P. Netrapalli, and T. Joachims, “Morel: Model-based offline reinforcement learning,” Advances in neural information processing systems, vol. 33, pp. 21 810–21 823, 2020. [23] T. Yu, G. Thomas, L. Yu, S. Ermon, J. Y. Zou, S. Levine, C. Finn, and T. Ma, “Mopo: Model-based offline policy optimization,” Advances in Neural Information Processing Systems, vol. 33, pp. 14 129–14 142, 2020. [24] A. Nair, A. Gupta, M. Dalal, and S. Levine, “Awac: Accelerating online reinforcement learning with offline datasets,” arXiv preprint arXiv:2006.09359, 2020. [25] M. Nakamoto, S. Zhai, A. Singh, M. Sobol Mark, Y. Ma, C. Finn, A. Kumar, and S. Levine, “Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning,” Advances in Neural Information Processing Systems, vol. 36, pp. 62 244–62 269, 2023. [26] K. Lei, Z. He, C. Lu, K. Hu, Y. Gao, and H. Xu, “Uni-o4: Unifying online and offline deep reinforcement learning with multi-step on-policy optimization,” arXiv preprint arXiv:2311.03351, 2023. [27] H. Zheng, X. Luo, P. Wei, X. Song, D. Li, and J. Jiang, “Adaptive policy learning for offline-to-online reinforcement learning,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, no. 9, 2023, pp. 11 372–11 380. [28] W. Xiao, J. Liu, Z. Zhuang, R. Suo, S. Lyu, and D. Wang, “Efficient online rl fine tuning with offline pre-trained policy only,” arXiv preprint arXiv:2505.16856, 2025. [29] K. Zhao, Y. Ma, J. Liu, J. Hao, Y. Zheng, and Z. Meng, “Improving offlineto-online reinforcement learning with q-ensembles,” in ICML Workshop on New Frontiers in Learning, Control, and Dynamical Systems, 2023. [30] Q.-W. Luo, M.-K. Xie, Y. Wang, and S.-J. Huang, “Optimistic critic reconstruction and constrained fine-tuning for general offline-to-online rl,” Advances in Neural Information Processing Systems, vol. 37, pp. 108 167–108 207, 2024. [31] R. Huang, D. Li, C. Shi, C. Shen, and J. Yang, “Augmenting online rl with offline data is all you need: A unified hybrid rl algorithm design and analysis,” arXiv preprint arXiv:2505.13768, 2025. [32] Q. Li, Z. Zhou, and S. Levine, “Reinforcement learning with action chunking,” arXiv preprint arXiv:2507.07969, 2025. [33] N.-C. Huang, P.-C. Hsieh, K.-H. Ho, and I.-C. Wu, “Ppo-clip attains global optimality: Towards deeper understandings of clipping,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 11, 2024, pp. 12 600–12 607. [34] J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, “Trust region policy optimization,” in International conference on machine learning. PMLR, 2015, pp. 1889–1897. [35] T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, “Continuous control with deep reinforcement learning,” arXiv preprint arXiv:1509.02971, 2015.
11
[36] X. B. Peng, A. Kumar, G. Zhang, and S. Levine, “Advantage-weighted regression: Simple and scalable off-policy reinforcement learning,” arXiv preprint arXiv:1910.00177, 2019. [37] H. K. Khalil, Nonlinear Systems. Prentice Hall, 2002. [38] K. J. Åström and B. Wittenmark, Adaptive Control, 2nd ed. AddisonWesley, 1995. [39] T. M. Cover, Elements of information theory. John Wiley & Sons, 1999. [40] T. Xie, N. Jiang, H. Wang, C. Xiong, and Y. Bai, “Policy finetuning: Bridging sample-efficient offline and online reinforcement learning,” Advances in neural information processing systems, vol. 34, pp. 27 395– 27 407, 2021. [41] P. Rashidinejad, B. Zhu, C. Ma, J. Jiao, and S. Russell, “Bridging offline reinforcement learning and imitation learning: A tale of pessimism,” Advances in Neural Information Processing Systems, vol. 34, pp. 11 702– 11 716, 2021. [42] S. Kakade and J. Langford, “Approximately optimal approximate reinforcement learning,” in Proceedings of the nineteenth international conference on machine learning, 2002, pp. 267–274. [43] J. Achiam, D. Held, A. Tamar, and P. Abbeel, “Constrained policy optimization,” in International conference on machine learning. PMLR, 2017, pp. 22–31. [44] A. A. Fedotov, P. Harremoës, and F. Topsoe, “Refinements of pinsker’s inequality,” IEEE Transactions on Information Theory, vol. 49, no. 6, pp. 1491–1498, 2003. [45] G. Garrigos and R. M. Gower, “Handbook of convergence theorems for (stochastic) gradient methods,” arXiv preprint arXiv:2301.11235, 2023. [46] S. Qiu, Z. Yang, J. Ye, and Z. Wang, “On finite-time convergence of actor-critic algorithm,” IEEE Journal on Selected Areas in Information Theory, vol. 2, no. 2, pp. 652–664, 2021. [47] R. Yuan, R. M. Gower, and A. Lazaric, “A general sample complexity analysis of vanilla policy gradient,” in International Conference on Artificial Intelligence and Statistics. PMLR, 2022, pp. 3332–3380. [48] T. Xu, Z. Wang, and Y. Liang, “Improving sample complexity bounds for (natural) actor-critic algorithms,” Advances in Neural Information Processing Systems, vol. 33, pp. 4358–4369, 2020. [49] E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for modelbased control,” in 2012 IEEE/RSJ international conference on intelligent robots and systems. IEEE, 2012, pp. 5026–5033. [50] F. Torabi, G. Warnell, and P. Stone, “Behavioral cloning from observation,” arXiv preprint arXiv:1805.01954, 2018. [51] B. Eysenbach, M. Geist, S. Levine, and R. Salakhutdinov, “A connection between one-step rl and critic regularization in reinforcement learning,” in International Conference on Machine Learning. PMLR, 2023, pp. 9485–9507. [52] S. Fujimoto and S. S. Gu, “A minimalist approach to offline reinforcement learning,” Advances in neural information processing systems, vol. 34, pp. 20 132–20 145, 2021. [53] Q. Zheng, A. Zhang, and A. Grover, “Online decision transformer,” in international conference on machine learning. PMLR, 2022, pp. 27 042–27 059. [54] H. Zhang, W. Xu, and H. Yu, “Policy expansion for bridging offline-toonline reinforcement learning,” arXiv preprint arXiv:2302.00935, 2023. [55] S. Lee, Y. Seo, K. Lee, P. Abbeel, and J. Shin, “Offline-to-online reinforcement learning via balanced replay and pessimistic q-ensemble,” in Conference on Robot Learning. PMLR, 2022, pp. 1702–1712.