Conceptio › Archive › arXiv CS
arXiv CSopen access

GEM-MPC: Balancing Exploration and Exploitation through Expert-Guided Planning

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

GEM-MPC: Balancing Exploration and Exploitation through Expert-Guided Planning

arXiv:2609.21735v1 [cs.LG] 18 Sep 2026

Álvaro Serra-Gomez and Thomas Moerland Abstract— Effective exploration in high-dimensional continuous control remains a central challenge in reinforcement learning. Planning-based methods address this by combining online planning with learned policies and value functions, but their components can become misaligned during training: learned sampling policies may diverge from planner behavior, while planning distributions stored in replay become stale as the model and value function evolve. Reanalysis can refresh these targets, but at substantial computational cost. We propose GEM-MPC, an MPPI-based reinforcement learning method that improves the interaction between planning and learning. GEM-MPC uses MPPI to combine a policy trained to clone the planner with a KL-regularized policy that explores around it, providing complementary exploitation and guided exploration within planning. We further introduce Gated Prior Distillation, which selectively learns from stored planning distributions only when they provide a better target than the current prior, reducing the impact of stale planning data without requiring full reanalysis. Across continuous-control benchmarks, GEMMPC consistently outperforms existing planning-based baselines under lower computational budgets.

I. I NTRODUCTION Exploration in high-dimensional continuous action spaces remains a major challenge in reinforcement learning. Planning-based methods address this challenge by combining shallow online search, such as Model Predictive Path Integral control (MPPI) [1], with learned sampling policies and bootstrap value functions [2], [3]. The sampling policy, πθs , directs the planner toward promising regions of the action π space, while the action-value function, QθQθs , extends the effective planning horizon. However, learning the sampling policy independently from the planner creates a mismatch between πθs and the MPPIinduced planning distribution π P . This mismatch degrades value function estimation, reducing sampling efficiency and limiting final performance. Recent methods address it by imitating the planner distribution [4] or using it as a prior for policy regularization [5], [6]. PO-MPC [7] further interprets this interaction as KLregularized reinforcement learning, where the sampling policy explores high-value actions around a policy prior, πθp , that acts as a proxy for the planner distribution. This formulation provides effective guided exploration, but does not fully resolve exploitation. The sampling policy may concentrate on only a subset of the local maxima of the action-value function around the learned prior, without necessarily preserving Both authors are with the Leiden Institute of Advanced Computer Science, Leiden University, 2311 EZ Leiden, The Netherlands.

{a.serra-gomez,t.m.moerland}@liacs.leidenuniv. nl

Fig. 1. Overview of GEM-MPC. Two independent MPPI instances compute distinct planning distributions, from which the higher-scoring plan is selected according to its estimated return. Both instances share the learned environment model and are biased by the sampling policies πs and πbc together with their corresponding bootstrap action-value functions. The selected planning distribution is then used to update the learned sampling policies. The regularized sampling policy πs is trained with KL-regularized RL, using a regularization policy learned with a support-covering objective over stored planning distributions, while the cloned policy πbc is trained with a mode-seeking behavior-cloning objective that concentrates on highdensity planner behavior.

useful modes represented by the learned prior itself. It is therefore well suited for exploration around the planner, while stable exploitation depends on maintaining a highquality planning samples. Learning a reliable planner representation heavily depends on the planning distributions stored in the replay buffer. Unlike transition data, which retains valid information about the environment’s transition function, planning data depends on the model and value function at the time of generation. As these get updated throughout training, stored planning targets may become stale and introduce noisy or harmful behavior. Reanalyze and Lazy reanalyze methods [8], [4] address this problem by recomputing planning solutions for stored environment transitions. However, such methods are computationally expensive and do not necessarily solve the problem, since a newly recomputed target is not necessarily better: an inaccurate or overconfident value function may replace a useful high-entropy historical target with a planning distribution that commits to an incorrect region of the action space. This work proposes Gated distillation with Expert Mixtures for Model Predictive Control (GEM-MPC), an MPPI-based RL method that automatically balances exploration and exploitation while providing an alternative to Reanalyzing strategies. Building upon PO-MPC [7], this

paper uses MPPI to seamlessly combine two policies: a greedy policy that clones planner, and an exploratory policy that explores around it. The former is trained through Behavior Cloning (BC), while the latter is trained through KLregularized RL. To avoid harmful behavior being distilled from stale planning distributions, we introduce a simple gating mechanism that distills a stored planning distribution only when it provides a better target than the current prior without requiring full reanalysis for every replayed state [4]. In summary, our contributions are: • Expert-augmented planning: a mixture-of-experts formulation that combines both a cloned policy, and an exploratory policy learned through KL-regularized RL to balance exploration and exploitation within planning. • Gated Prior Distillation: a theoretically motivated gating mechanism that filters stale or harmful planning targets, resulting in planning distribution representation at substantially lower computational cost than reanalysis. We show that the proposed method consistently improves existing planning-based reinforcement learning baselines [7], [4] under lower computational budgets. An overview of our method’s structure is presented in Fig. 1. II. R ELATED W ORK Model-based reinforcement learning combines environment models with policy and value learning to improve sequential decision-making [9]. Our work focuses on methods that use online planning, i.e. MPPI [1], both to select actions and to guide learning. MPPI-based MBRL methods exploit a reciprocal relationship between planning and learning. In TDMPC and TD-MPC2, learned policies provide trajectory proposals, while value functions estimate returns beyond the planning horizon [2], [3]. Conversely, planning can provide supervision for policy learning. BMPC learns a policy by imitating the planner’s action distribution [4], while PO-MPC uses a learned representation of this distribution as a prior for KL-regularized policy optimization [7]. Similar approaches follow similar guidelines underscoring the importance of policy learning and planning alignment [5] to avoid sampling out-of-distribution state-action couples. Exploration and exploitation during planning. Sampling-based planning must concentrate its search on promising actions while retaining sufficient diversity to discover alternatives. Entropy regularization and learned behavior priors offer mechanisms for shaping policy exploration [10], [11]. Other approaches modify trajectory evaluation through uncertainty penalties [12], pessimistic value estimates to avoid overestimation in out-of-distribution state-action couples [13], or maintain independent trajectory optimizers to reduce susceptibility to local optima, as in DecentCEM [14]. Building on PO-MPC, our method combines a planner-cloned policy for exploitation with a KL-regularized policy for guided exploration. Separate

MPPI branches preserve their complementary sampling biases during planning. Learning from stored planning results. Planning distributions stored in replay can become outdated as models, policies, and value functions evolve. Reanalysis addresses this problem by recomputing planning targets for previously collected experience, as in MuZero Reanalyse and BMPC’s lazy reanalyze [8], [4]. A related line of work accounts for differences in stored target quality through value-based weighting of planner alignment using the current Q-function [6], [15]. Our work follows an alternative path, similar to demonstration Q-filters in Model-free RL, which apply imitation only when a demonstrated action is judged better than the policy’s action [16], [17]. Gated Prior Distillation follows this selective-imitation perspective, using value-based comparisons to determine which stored planner distributions are not detrimental when learning the planning distribution. This allows planning experience to be reused selectively without recomputing each target. III. P RELIMINARIES We consider a discrete-time sequential decision-making problem over a finite horizon of T steps, modeled as a Markov decision process (MDP) (S, A, p, r, γ). Here, S and A denote the state and action spaces, respectively, p(· | s, a) is the transition distribution, encompassing both stochastic and deterministic dynamics, r(s, a) is the immediate reward for taking action a in state s, and γ ∈ [0, 1) is the discount factor. A policy π(· | s) specifies a distribution over actions in state s. The objective is to find a policy that maximizes the expected discounted return: J(π) = Es0 ∼ρ0 , at ∼π(·|st ),

"T −1 X

st+1 ∼p(·|st ,at )

# t

γ r(st , at ) ,

t=0

where ρ0 denotes the initial state distribution. MPPI-based Reinforcement Learning. Recent methods in model-based reinforcement learning integrate online planning directly into the learning loop, with approaches based on Model Predictive Path Integral Control (MPPI) [1] proving particularly effective in high-dimensional robotic control [2], [3]. These methods learn a latent model of the environment, (Ŝ, A, p̂, r̂, γ), where observations or states are mapped to a latent representation z = hθh (s) ∈ Ŝ, and planning is performed using learned reward and transition models r̂(z, a) = rθr (z, a) and p̂ = pθd [18]. MPPI is a sample-based stochastic planning method that iteratively refines a distribution over finite-horizon action sequences according to their predicted returns. Let H denote the planning horizon, and ā0:H−1 = (ā0 , . . . , āH−1 ) the mean of the current open-loop action distribution. At each planning iteration, MPPI samples M perturbed action sequences, (i)

(i)

at = āt + ϵt ,

(i)

ϵt ∼ N (0, σt2 I),

(1)

and rolls them out through the learned latent dynamics,   (i) (i) (i) zt+1 ∼ pθd · | zt , at . (2) The resulting trajectories τi are evaluated using their predicted returns: H−1 X  (i) (i)  R(τi ) = r̂ zt , at . (3)

maximize its action-value function while being regularized to generate trajectories that lie within the planner’s trajectory distribution. For that reason, a learned policy that approximates the planner is used as regularizer. The sampling policy is updated by maximizing the following objective function: π

exp (R(τi )/λ) , wi = P j exp (R(τj )/λ)

− λKL[πθs (· | zt ) ∥ πreg (· | zt )],

where the update may be restricted to the K highest-return trajectories. After a fixed number of iterations, the optimized distribution over the first action defines the planning policy, πP (· | z) = N (ā0 , σ02 I).

(6)

Only the first action is executed before replanning from the next state, yielding a receding-horizon feedback policy. The mean sequence can additionally be warm-started by shifting the solution from the previous decision step by one time step. While standard MPPI samples from a broad Gaussian proposal, recent model-based RL methods additionally generate part of the candidate trajectories using a learned sampling policy πθs [3], [4]. This biases the finite planning budget toward action sequences that are already expected to achieve high return, while the higher-variance MPPI proposal preserves exploration outside the current learned policy. The learned sampling policy is also used to estimate returns beyond the finite planning horizon. In particular, π an action-value function QθQθs bootstraps the terminal latent state of each rollout, yielding the H-step estimate (i)

Q(z0 , a0:H ) =

H−1 X t=0

π

where the regularized Q-value function Qθ̃ θs Q following Bellman recursive expression:

(4)

where λ > 0 controls the concentration of the weighting distribution. The parameters of the action distribution are then updated using the weighted perturbations, v u PK K (i) u X wi (ϵt )2 (i) , (5) āt ← āt + wi ϵt , σt ← t i=1 PK i=1 wi i=1

    π (i) (i) (i) (i) γ t rθr zt , at +γ H QθQθs zH , aH .

(7) Thus, planning operates over a mixture of trajectories proposed by the learned sampling policy and by the broader MPPI distribution. The learned policy concentrates samples in regions of the action space that are already predicted to be useful, whereas the MPPI proposal retains broader coverage. MPPI subsequently reweights these candidate trajectories according to their predicted H-step returns, combining policyguided sampling with online trajectory optimization. The method proposed by this paper is built PO-MPC [7], which unifies a subset of MPPI-based algorithms [3], [4] by framing sampling policy learning as KL-regularized RL. Under this framework, the sampling policy is trained to

(8)

Q

t=0

MPPI assigns exponentially larger weight to trajectories with higher predicted returns,

,λ

J(π) =Ea∼πθs [Qθ̃ θs (zt , at )] ,λ

satisfies the

" π ,λ Qθ̃ θs (zt , at ) = Est+1 ∼p(·|st ,at ), Q a∼πθs (·|zt+1 )

+γ

r(zt , at )

π ,λ Qθ̃ θs (zt+1 , a) − λ log Q

πθs (a | zt+1 ) πreg (a | zt+1 )

(9) !!#

Note that, as remarked in [7], there are many ways to train the regularization policy. This results in different inductive biases being embedded into the sampling policy, resulting in faster convergence or enhanced exploration during planning. We shall leverage this property in the following section. As in PO-MPC, the rest of this work borrows world model learning from TD-MPC2, and maintains the architecture choices for the encoder, policies and action-value functions. IV. M ETHOD An overview of our method is presented in Fig. 1. This work addresses two limitations of MPPI-based reinforcement learning: the exploration-exploitation trade-off during planning and the use of stale planning distributions for learning sampling policies. Prior approaches typically bootstrap MPPI with a single sampling policy and its corresponding action-value function, thereby biasing planning toward either exploration or exploitation depending on how the sampling policy is trained. In high-dimensional, non-convex action-value landscapes, this can lead the planner to overcommit to local optima or remain around more stable but suboptimal solutions. We instead combine exploratory and exploitative sampling policies, together with their corresponding action-value functions, to provide complementary proposals during planning. However, this requires access to the current planning policy. A separate challenge arises when distilling the resulting planning distributions into learned sampling policies. Ideally, these policies would be trained against the current planner, but obtaining up-to-date planning targets requires re-planning from replayed states. Prior methods either learn directly from stored planning distributions, which become stale as the policies and value functions evolve, or mitigate this mismatch through reanalysis, i.e., re-planning transitions from the replay buffer at substantial computational cost. We instead introduce a theoretically motivated gating criterion that selectively distills stored planning targets according to their action-value under the current learned policies, providing a lower-cost alternative to reanalysis.

A. Expert-augmented Planning via Decentralized MPPI Biased Trajectory Sampling. Candidate trajectories are generated from a three-component sampling scheme composed of the planning distribution πP and two biasing policies πθs and πθbc . If a candidate is sampled from πP with probability (1 − p) and from πi , i ∈ {θs , θbc } each with equal probability p2 , the induced trajectory distribution is π̂s (τ ) = (1 − p)πP (τ ) + p2 πθs (τ ) + p2 πθbc (τ ), where each component denotes the trajectory distribution induced by the corresponding action-sequence sampling mechanisms and the environment dynamics. Assuming access to previously computed planning distributions stored in the replay buffer, the aim of the first policy, the regularized sampling policy: πθs , is to provide a distribution that maximizes the action value function while remaining close to the planner distribution. Therefore, it is trained using KL-regularized RL, as in PO-MPC [7], where the policy updates are regularized using a regularization policy, πθreg , that keeps πθs trajectories close to the planner’s following Equation 8. While the regularization policy should be equal to the planner distribution, this one is constantly changing because it depends on the current state of the bootstrap action-value functions, which are also constantly being updated. For that reason, we use a learned regularization policy that clones the planning distributions. In order not to prematurely dismiss part of the support of stored MPPI samples, which might turn out to contain high local maxima of the action-value function, the regularization policy is trained to clone the stored planner samples while discouraging assigning low probability to regions with high planner density. This is done by minimizing the following forward KL divergence: h i J(θreg ) = E(s,πP )∼D KL[πP (· | zt ) ∥ πθreg (· | zt )] , (10) where D is the data stored in the replay buffer, and πP is the stored planning distribution. On the contrary, the goal of the second policy, the cloned sampling policy πθbc , is to clone closely the most likely behavior of the planner. This is why it is trained to concentrate on high-density regions of the planning distributions stored within the replay buffer: h i J(θbc ) = E(s,πP )∼D KL[πθbc (· | zt ) ∥ πP (· | zt )] . (11) The application of both distillation objectives to potentially stale planning targets is governed by the gating mechanism introduced in Sec. IV-B. Trajectory evaluation. Typically, simulated trajectories are evaluated using the modeled H-step action-value function, where H is the planning horizon (Eq. 7). This implies using the modeled cumulative reward plus a bootstrap action value function of the learned policy is used to bias planning. Assuming the sampling policy is followed past the planning horizon, the bootstrap action-value function

is an estimate of the expected value of the trajectory past the planning horizon. Instead, our method bootstraps each trajectory with the mean of the action-value function of each policy, πθs and πθbc , evaluated in actions sampled from their respective distributions:

(i) Q(z0 , a0:H ) =

H−1 X

s,(i)

γ t rθr (zt , at

)

t=0

+

s,(i)

γ H π θs πθ (i) bc,(i) (QθQ (zH , aH ) + QθQbc (zH , aH )). s bc 2 (12) s,(i)

Where at and at are sampled according to πθs and πθbc . Once the planning horizon limit is reached, this is equivalent, in expectation, to randomly choosing one out of the two policies and using it until reaching the end of the episode. The intuition behind our approach is to evaluate each trajectory by the potential cumulative reward it may gather past the planning horizon under the assumption that either one of the learned sampling policies is followed. Decentralized MPPI. To obtain a coherent plan while preserving the complementary biases induced by πθs and πθbc , we adopt a strategy analogous to Decentralized CEM [14], while retaining the exponential weighting used by MPPI. We maintain two independent MPPI proposal distributions over action sequences, one associated with each sampling policy. At each MPPI iteration, each branch i ∈ s, bc receives a budget of N/2 candidate trajectories. Of these, N/2 − nbias action sequences are sampled from the current MPPI proposal of that branch, while nbias are generated using its corresponding learned sampling policy πθi . Candidate trajectories are simulated and evaluated using Eq. 12. The highest-scoring top-k candidates are then used to independently update each branch according to the exponential MPPI aggregation in Eq. 5. After K planning iterations, each branch yields a candidate plan given by the mean of its final action-sequence distribution. We evaluate both candidate plans using Eq. 12 and execute the first action of the higher-scoring plan. Only the parameters of the selected planning distribution are stored in the replay buffer together with the resulting transition. B. Gated Planning Distillation This work learns policies that approximate the planner’s conditioned action distribution, both to regularize πθs and to train the cloned policy πθbc . Prior work pursuing similar objectives typically assumes access to the planner’s current distribution. In our setting, satisfying this assumption would require re-planning from every sampled state, as the planning distribution depends on the current biasing distributions and bootstrap action-value functions. Recomputing these targets throughout training is therefore computationally prohibitive. A cheaper alternative is to distill planning distributions stored in the replay buffer. However, these targets become

stale as the policies and value functions evolve. Recent methods mitigate this issue through partial or periodic reanalysis [4], [8], but doing so still incurs substantial computational overhead. Rather than assuming access to an up-to-date planner, we study when stale planning targets remain useful for training the regularization policy πθreg and the cloned sampling policy πθbc . Related work [6] similarly recognizes that planning targets should not contribute equally, and reweights them using a normalized exponential function of their recorded episodic returns. We instead consider whether a planning target should contribute to distillation at all. To this end, we introduce Gated Planning Distillation, which selectively distills stored planning distributions according to their value under the current learned policies. Although preferentially distilling higher-value planning targets is intuitive, our gating criterion follows from the objectives governing the two sampling policies. Specifically, we show that the cloned sampling policy affects an upper bound on the planning objective, while the regularization policy affects an analogous bound on the KL-regularized policy objective. These results motivate gating stored planning targets according to whether their distillation is expected to improve the corresponding bound.

Eq. 13 is tight at the optimal q [1], increasing this quantity raises the maximum achievable value of the KL-regularized objective in Eq. 13. Thus, a biasing policy with lower expected exponential return can impose a lower ceiling on the achievable objective of the planner, whereas improving this quantity relaxes that limitation. Since πθbc is trained by directly cloning planning distributions stored in the replay buffer, the preceding result motivates the first component of our gating mechanism. Each replay transition stores the action aP executed by the planner, corresponding to the mean of the resulting planning distribution, together with its covariance matrix. Rather than indiscriminately cloning every stored planning distribution through Eq. 11, we only distill a replayed target when it is estimated to increase the expected exponential-return term in Eq. 14, thereby avoiding targets estimated to lower the maximum achievable planning objective. Since computing the corresponding trajectory-level expectation exactly would require additional rollouts and is computationally expensive, we introduce the following tractable approximation:

Effects in planning: Cloned Sampling Policy. As shown in [1], given an arbitrary sampling policy over action sequences that acts as prior distribution from which to sample trajectories, τ := (z0 , a0 , z1 , a1 , . . .), with induced trajectory distribution π̂s (τ ), the planning objective function is upper bounded by:

where ai ∼ πθbc (·|z). The left-hand side is a Monte Carlo estimate of the expected exponential action-value under the current cloned policy, whereas the right-hand side evaluates the same surrogate at the action taken by the planner at the transition. We use Qπθbc (z, a) as a tractable approximation of the expected return conditioned on (z, a), replacing the trajectory-level exponential-return criterion with an action-value-based gating rule.

J(q) = Eτ ∼q(τ ) [R(τ )] − λKL[q ∥ π̂s ] ≤ −1

λ log Eτ ∼π̂s [exp(λ

(13)

R(τ ))]

Assuming candidate trajectories are generated from a two-component sampling scheme composed of the planning distribution πP and a biasing policy πi , i ∈ {s, bc}, as commonly done in MPPI-based RL methods [3], [5]. If a candidate is sampled from πP with probability (1 − p) and from πi with probability p, the induced trajectory distribution is π̂s (τ ) = (1 − p)πP (τ ) + pπi (τ ). Then,    1 R(τ ) = (14) Eπ̂s exp λ       1 1 (1 − p)EπP exp R(τ ) + pEπi exp R(τ ) , λ λ where R(τ ) is the cumulative reward gathered along the trajectory τ . In practice p is implemented as a fixed proportion of samples being taken from πi and (1 − p) from πP . For fixed p > 0, and holding the πP term fixed, the right-hand side of Eq. 14 is monotonically increasing in Eτ ∼πi [exp(λ−1 R(τ ))]. Since the logarithm is also monotonically increasing for λ > 0, and the variational bound in

Nbc 1 X exp(λ−1 Qπθbc (z, ai )) ≤ exp(λ−1 Qπθbc (z, aP )), Nbc i=1 (15)

Effects in planning: KL-regularized Sampling policy. A similar implication holds for the KL-regularized sampling policy. Here, stale planning targets do not affect the planning objective directly; instead, they shape the regularization policy πθreg , which determines the variational ceiling of the KL-regularized objective optimized by πθs . Consequently, distilling poor planning targets into πθreg can reduce the maximum achievable value of the regularized policy objective. Starting from the KL-regularized policy objective, we obtain an analogous variational bound: Ea∼πθs [Qπθs (z, a)] − λKL[πθs |πθreg ]

(16) −1

≤ λ log Eπθreg [exp(λ

πθs

Q

(z, a))].

Thus, the maximum achievable value of the KLregularized objective depends explicitly on the regularization policy πθreg . For a fixed Qπθs , increasing the right-hand side of Eq. 16 raises the maximum achievable value of the KL-regularized objective, although it does not necessarily imply an improvement in the resulting policy πθs . Conversely, decreasing this

Fig. 2. Performance comparison in 14 state-based high-dimensional control tasks from HumanoidBench [19]. Interquartile Mean (IQM) of 5 runs; shaded areas are the 95% bootstrap Confidence Intervals (CI). In the top left, we visualize the Aggregate IQM with stratified bootstrap CI across all tasks except for Reach, which has a different return range. GEM-MPC either matches or outperforms other baselines across tasks.

quantity imposes a lower ceiling on the objective optimized by πθs . We therefore apply an analogous gating mechanism as in Eq. 15 when distilling the planning distribution into πθreg , retaining a replayed target only when its exponential action-value surrogate is at least as large as the corresponding expectation under the current regularization policy: Nreg 1 X exp(λ−1 Qπθs (z, ai )) ≤ exp(λ−1 Qπθs (z, aP )), Nreg i=1 (17)

where ai ∼ πθreg (·|s). Note that the action-value function in Eq. 17 is that of the KL-regularized sampling policy, Qπθs . Unlike Eq. 13, whose expectation is defined over complete trajectories, Eq. 16 is defined over single-step actions conditioned on the current state. This allows the exponentialvalue term to be evaluated directly using Qπθs (z, a), without introducing the trajectory-level return surrogate required for the cloned sampling policy. C. Method Summary. Our method addresses two limitations of MPPI-based reinforcement learning: the exploration-exploitation tradeoff induced by relying on a single sampling policy during planning, and the instability caused by learning from stale planning distributions stored in replay. We address the first by augmenting MPPI with two complementary sampling policies: a KL-regularized policy πθs that remains close to the planner while maximizing return, and a mode-seeking cloned policy πθbc that captures the planner’s most likely behavior. Together with their corresponding bootstrap action-value functions, these policies provide complementary exploratory

and exploitative trajectory proposals, which are optimized independently through a decentralized MPPI procedure and compete to determine the executed plan. We address the second limitation through Gated Planning Distillation. Rather than indiscriminately cloning stored planning distributions or recomputing them through costly reanalysis, we update the cloned policy πθbc and the planner proxy πθreg only when a stored planning action improves the upper-bounds that depend on them (Eqs. 13 and 16). This Plan→Evaluate→Gate loop allows past planning experience to be reused selectively as the learned action-value functions evolve, retaining useful planner information while discarding stale targets that could restrict subsequent planning or policy performance. The resulting framework therefore combines complementary exploration and exploitation at planning time with a computationally efficient alternative to re-analysis for learning from replayed planning distributions. V. E XPERIMENTS Experimental Setup. We evaluate different configurations of the proposed framework (GEM-MPC ) on the 14 continuous control tasks from HumanoidBench locomotion suite [19]. These tasks are high-dimensional and cover a diverse range of continuous control challenges, including sparse reward, locomotion with high-dimensional state and action space (A ∈ R61 ). All the experiments are run on a partition of an NVIDIA A100 GPU configured as a 4g.40GB Multi-Instance GPU (MIG), for 1e6 timesteps. In this work, we build upon the official JAX [20] implementation of PO-MPC [7]. As in PO-MPC, we inherit all architectural and pipeline choices from TD-MPC2, including data gathering through random walks for 1e4

TABLE I A BLATIONS OF GEM-MPC COMPARED WITH PO-MPC [7]. I NTERQUARTILE M EAN (IQM) OF 5 RUNS WITH 95% BOOTSTRAP CI. Method PO-MPC w/o LR PO-MPC

Lazy Re-analyze

MoE

GPD

Aggregate IQM

Training Time (h)

✗ ✓

✗ ✗

✗ ✗

697.0 [354.5, 918.5] 728.0 [431.3, 918.5]

6.4 ± 0.2 14.0 ± 0.8

GEM-MPC (Ours)

✗

✓

✓

762.5 [436.2, 930.1]

9.2 ± 0.4

GEM-MPC w/o GPD Only πθs Only πθbc

✗ ✗ ✗

✓ ✗ ✗

✗ ✓ ✓

738.9 [391.6, 929.0] 605.0 [273.8, 868.0] 667.5 [315.8, 905.2]

8.3 ± 0.3 7.0 ± 0.2 6.8 ± 0.2

total environment steps and the pre-training phase of the dynamic model and bootstrap action value functions. The architecture of all KL-regularized and bootstrap action value functions follow the same design. For reproducibility, our implementation and hyperparameters are available at https://anonymous.4open.science/r/GEM-MPC-1E79.

these gains do not require the computational cost of Lazy Reanalyze: GEM-MPC reduces training time from 14.0 to 9.2 hours while attaining higher aggregate performance. These results suggest that selectively exploiting useful past planning distributions provides a more favorable performancecompute trade-off than periodically recomputing them.

Baselines. We empirically support the claims in this work by comparing GEM-MPC with other algorithms within the state of the art in MPPI-based RL, namely TD-MPC2 [3], BMPC [4] and PO-MPC [7] with, and without, reanalyzing. Note that BMPC also uses Lazy Re-analyze. We also explore ablations of GEM-MPC by studying the effect of omitting any of the design choices introduced in Section IV. Table I provides an overview on the contributions of our work and their effects. For our experiments, we employ the baseline implementations in JAX [7], [21], which are either implemented by the original authors or in collaboration with them, since they reproduce the results from the original paper while increasing the computation speed. We evaluate our method under the same hyperparamers as PO-MPC, except for those related only to GEM-MPC .

B. Balancing Exploration and Exploitation

A. Results The objective of this section is to test GEM-MPC from three perspectives. First, we show our method either surpasses or matches in performance the other baselines. Second, we make an empirical study over the effect of each of our contributions introduced in Section IV, comparing each ablation against PO-MPC, with and without Lazy Reanalyze, showing that their combination improves upon the state of the art both in terms of higher performance at lower training times. Finally, we further ablate our method, empirically analyzing each of the mechanisms introduced as part of each contribution. Figure 2 compares GEM-MPC against the baselines and illustrates the effect of Lazy Re-analyze in PO-MPC. Overall, combining Expert-augmented Planning (MoE) with Gated Planning Distillation (GPD) matches or improves upon the baselines across the high-dimensional HumanoidBench tasks, achieving the highest aggregate IQM of 762.5, compared with 728.0 for PO-MPC with Lazy Re-analyze. The largest gains are observed in Slide and Stair, where GEM-MPC reaches substantially higher returns, while its learning curves also exhibit greater stability in several tasks. Importantly,

We investigate whether the performance gains of GEMMPC arise from separating exploration and exploitation across sampling policies with distinct objectives. Starting from PO-MPC, where a single sampling policy must balance both behaviors through its regularization objective, we progressively disentangle these roles and evaluate different combinations of KL objectives: mode-seeking behavior for exploitation and support-covering behavior for broader exploration. Our results in Table I show that assigning these objectives to separate policies yields stronger and more stable performance, allowing the KL-regularized sampling policy to optimize over a broader action support while the cloned policy exploits high-probability planning behavior. Decentralized MPPI further improves robustness by maintaining separate planning proposals, consistent with improved robustness in tasks with multiple local optima, e.g. Sit Hard. C. Effects of Gating Planning Data We next analyze the behavior of Gated Planning Distillation throughout training. Table I shows the impact of gating on performance. During training, the proportion of samples satisfying each of Eq. 15 and 17 is observed to drop rapidly below 50% of the sampled replay data and steadily decline over training until reaching around 20%. Despite training on substantially fewer planning targets, the gated variant consistently achieves higher performance than indiscriminate distillation. This suggests that a large fraction of stored planning distributions becomes uninformative, or potentially detrimental, as the planner evolves, and that selectively discarding such targets can improve learning. VI. D ISCUSSION AND C ONCLUSION Summary of Findings. Across 14 HumanoidBench tasks, GEM-MPC consistently improves upon the baselines. As shown in Table I, using both πθs and πθbc during planning already yields gains over PO-MPC, even without Gated Planning Distillation. Enabling GPD further widens this gap while remaining a more computationally efficient

alternative to Lazy Re-analyze. Our mixture-of-experts ablations show that jointly using the exploratory and cloned sampling policies outperforms relying on either policy alone, supporting the role of complementary sampling objectives during planning. In turn, the GPD ablations show that selectively distilling replayed planning targets according to Eqs. 15 and 17 improves performance over indiscriminate cloning. Together, these results support the two main design choices of GEM-MPC : combining complementary sampling policies improves the quality and robustness of planning, while Gated Planning Distillation provides a tractable mechanism for filtering stale planning targets and retaining those that remain useful for learning the cloned and regularization policies. Limitations and Future Work. Our first limitation concerns Gated Planning Distillation, which relies on learned action-value estimates to determine which replayed planning targets remain useful. Errors in these estimates may therefore affect the gating decisions, and our criteria should be interpreted as tractable surrogates rather than exact tests of target usefulness. A second limitation is the additional memory and computation required relative to approaches using a single sampling policy, as our method maintains an additional policy and action-value function and evaluates trajectories under both. Finally, our method relies on Gaussian policy and planning distributions, which may be restrictive in high-dimensional or multimodal action spaces where high-value regions can exhibit more complex structure. More expressive policy classes, such as flow-matching policies that approximate Boltzmann distributions [22], provide a promising direction for representing richer sampling distributions and potentially reducing the need for multiple specialized components. ACKNOWLEDGMENT The authors used ChatGPT to assist generating code for plotting the results. The authors reviewed and verified the resulting code, and take full responsibility for the final manuscript, implementation, and reported results. R EFERENCES [1] G. Williams, P. Drews, B. Goldfain, J. M. Rehg, and E. A. Theodorou, “Information-Theoretic Model Predictive Control: Theory and Applications to Autonomous Driving,” IEEE Transactions on Robotics, vol. 34, no. 6, pp. 1603–1622, 2018. [2] N. A. Hansen, H. Su, and X. Wang, “Temporal Difference Learning for Model Predictive Control,” International Conference on Machine Learning, pp. 8387–8406, 2022. [3] N. Hansen, H. Su, and X. Wang, “TD-MPC2: Scalable, Robust World Models for Continuous Control,” in International Conference on Learning Representations (ICLR), 2024. [4] Y. Wang, H. Guo, S. Wang, L. Qian, and X. Lan, “Bootstrapped Model Predictive Control,” in The Thirteenth International Conference on Learning Representations, 2025. [5] H. Lin, P. Wang, J. Schneider, and G. Shi, “TD-M(PC)2 : Improving Temporal Difference MPC Through Policy Constraint,” arXiv preprint arXiv:2502.03550, 2025. [6] G. Zhan, L. Wang, X. Zhang, J. Gao, M. Tomizuka, and S. E. Li, “Bootstrap Off-policy with World Model,” in The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. [Online]. Available: https://openreview.net/forum?id=zNqDCSokDR

[7] A. Serra-Gomez, D. J. Ornia, D. Tirumala, and T. M. Moerland, “A KL-regularization framework for learning to plan with adaptive priors,” in Forty-third International Conference on Machine Learning, 2026. [Online]. Available: https://openreview.net/forum?id=zO8vzSGgTn [8] J. Schrittwieser, T. K. Hubert, A. Mandhane, M. Barekatain, I. Antonoglou, and D. Silver, “Online and Offline Reinforcement Learning by Planning with a Learned Model,” in Advances in Neural Information Processing Systems, A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan, Eds., 2021. [Online]. Available: https://openreview.net/forum?id=HKtsGW-lNbw [9] T. M. Moerland, J. Broekens, A. Plaat, and C. M. Jonker, “Modelbased reinforcement learning: A survey,” Found. Trends Mach. Learn., vol. 16, no. 1, p. 1–118, Jan. 2023. [10] T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Offpolicy maximum entropy deep reinforcement learning with a stochastic actor,” in Int. conf. on machine learning. Pmlr, 2018, pp. 1861–1870. [11] D. Tirumala, A. Galashov, H. Noh, L. Hasenclever, R. Pascanu, J. Schwarz, G. Desjardins, W. M. Czarnecki, A. Ahuja, Y. W. Teh, and N. Heess, “Behavior Priors for Efficient Reinforcement Learning,” Journal of Machine Learning Research, vol. 23, no. 221, pp. 1–68, 2022. [Online]. Available: http://jmlr.org/papers/v23/20-1038.html [12] T. Evers, C. Meo, W. Bohmer, J. Dauwels, and Y. Oren, “Efficienttdmpc: Improved mpc objectives for sample-efficient continuous control,” arXiv preprint arXiv:2605.16692, 2026. [13] W.-D. Chang, M. Henaff, B. Amos, G. Dudek, and S. Fujimoto, “The surprising difficulty of search in model-based reinforcement learning,” in Forty-third International Conference on Machine Learning, 2026. [14] Z. Zhang, J. Jin, M. Jagersand, J. Luo, and D. Schuurmans, “A simple decentralized cross-entropy method,” in Advances in Neural Information Processing Systems, A. H. Oh, A. Agarwal, D. Belgrave, and K. Cho, Eds., 2022. [Online]. Available: https://openreview.net/forum?id=IQIY2LASzYx [15] Z. Zhuang, D. Shi, R. Suo, X. He, H. Zhang, T. Wang, S. Lyu, and D. Wang, “Tdmpbc: Self-imitative reinforcement learning for humanoid robot control,” arXiv preprint arXiv:2502.17322, 2025. [16] A. Nair, B. McGrew, M. Andrychowicz, W. Zaremba, and P. Abbeel, “Overcoming exploration in reinforcement learning with demonstrations,” in 2018 IEEE international conference on robotics and automation (ICRA). IEEE, 2018, pp. 6292–6299. [17] Z. Wang, A. Novikov, K. Zolna, J. S. Merel, J. T. Springenberg, S. E. Reed, B. Shahriari, N. Siegel, C. Gulcehre, N. Heess, and N. de Freitas, “Critic regularized regression,” in Advances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, Eds., vol. 33. Cur. Assoc., Inc., 2020, pp. 7768–7778. [18] M. Bhardwaj, S. Choudhury, and B. Boots, “Blending MPC & Value Function Approximation for Efficient Reinforcement Learning,” in International Conference on Learning Representations, 2021. [19] C. Sferrazza, D.-M. Huang, X. Lin, Y. Lee, and P. Abbeel, “HumanoidBench: Simulated Humanoid Benchmark for Whole-Body Locomotion and Manipulation,” in Proceedings of Robotics: Science and Systems, Delft, Netherlands, July 2024. [20] J. Bradbury, R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, G. Necula, A. Paszke, J. VanderPlas, S. WandermanMilne, and Q. Zhang, “JAX: composable transformations of Python+NumPy programs,” 2018. [21] S. Flandermeyer, “bmpc-jax: Jax/flax implementation of BMPC,” https://github.com/ShaneFlandermeyer/bmpc-jax, 2024, accessed: 2025-08-28. [22] H. Zhong, Z. Li, X. Wang, and L. Huang, “Reparameterization flow policy optimization,” in Forty-third International Conference on Machine Learning, 2026.

Record · ID 1006871 · SHA-256 0edf41b3a5ba2992
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.