LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Manith Adikari1,4∗ , Bei Peng2 , Samuele Vinanzi3 , Angelo Cangelosi1,4 1
Department of Computer Science, University of Manchester 2 School of Computer Science, University of Sheffield 3 School of Computing & Digital Technologies, Sheffield Hallam University 4 Centre for Robotics & AI, University of Manchester
arXiv:2607.29559v1 [cs.AI] 31 Jul 2026
Abstract Reinforcement Learning (RL) systems are typically trained using a single, well-specified scalar reward function. However, real-world decision-making tasks often involve multiple, competing objectives, such as performance versus efficiency, where ground-truth reward functions are difficult to specify or inaccessible. While Multi-Objective RL (MORL) addresses such trade-offs by modeling rewards as vectors, existing approaches typically assume access to a well-specified reward function for each objective, inheriting the same challenges faced by single-objective RL. Meanwhile, Preference-based RL (PbRL) has shown great potential in solving complex tasks without access to a pre-defined reward function through reward learning from human feedback, yet has largely been studied in single-objective settings. In this work, we bridge this gap with LEMUR: Learning to Align with Multi-Objective Reinforcement Learning with Preference feedback, a novel framework where an agent interactively learns from the preferences of multiple humans to learn optimal multi-objective policies. Our approach jointly learns policies and multiple objective-specific reward models from human feedback, enabling agents to effectively balance competing objectives during learning. We evaluate LEMUR on a variety of benchmark multi-objective tasks, and empirical results demonstrate its superior performance over baseline methods. Our method presents a promising direction for solving multi-objective decision-making tasks without pre-defined reward functions.
1
Introduction
Reinforcement Learning (RL) has achieved remarkable success in training autonomous agents, from game-playing (Mnih et al. 2013) to robotics (Tan et al. 2018). By formalizing learning as the maximization of cumulative rewards through trial and error learning (Sutton and Barto 1998), RL provides a natural framework for training autonomous ‘goalseeking’ agents (McCarthy 1997). However, standard RL relies on two critical assumptions: that the goal can be represented by a single scalar reward, and that this reward function is well-specified. In practice, these assumptions rarely hold in complex, real-world domains. Real-world tasks often involve multiple, competing objectives (Dulac-Arnold, Mankowitz, and Hester 2019), such as balancing speed versus safety in autonomous driving ∗
Corresponding Author: [email protected]
(Wang et al. 2026), or throughput versus energy efficiency in robotics (Huang et al. 2022; Kouritem et al. 2022). The field of Multi-Objective Reinforcement Learning (MORL) addresses this by modeling rewards as vectors to find a set of Pareto-optimal policies (Roijers et al. 2013): policies where no objective’s expected returns can be increased without decreasing the returns of other objectives. However, existing MORL approaches typically assume that the ground-truth reward function for each objective is accessible and manually specified (Hayes et al. 2022b). Thus, they inherit the same challenges of reward specification from the single-objective RL domain, now across multiple objectives. Manually designing a reward function that can achieve an adequate balance between competing objectives is challenging and can lead to oversimplification, which results in suboptimal policies (Knox et al. 2012) and possible reward exploitation (Amodei et al. 2016). To avoid complicated reward engineering, Preference-based RL (PbRL) learns reward models directly from human feedback (Christiano et al. 2017), which leads to better alignment between the system’s behavior and human preferences. Despite progress in reward learning for single-objective RL, this problem remains largely unexplored in the MORL setting. While some works consider reward learning in MORL, they are typically limited to Large Language Model (LLM) post-training settings or narrow natural language-based tasks (Bakker et al. 2022; Rame et al. 2023; Yang et al. 2024), and do not study the joint learning of policies and rewards in environments with multiple, competing objectives. Reward specification is already a major bottleneck in scaling single-objective RL settings, and becomes even more critical in MORL. Two key challenges arise. First, the agent must optimize multiple objectives whose underlying reward functions are complex, implicit, or inaccessible. Second, even when human feedback is available, collapsing multi-criteria reward signals into a single scalar obscures the very tradeoff structure that MORL is designed to optimize (Vamplew et al. 2011; Roijers et al. 2013; Sorensen et al. 2024). Consider training a robotic system: non-expert operators can reliably judge task-level success, such as whether objects were grasped or placed correctly, while expert operators are required to assess fine-grained criteria such as grasp stability, long-term wear, or safety-related objectives. Collapsing such heterogeneous feedback into a single reward risks conflating
distinct objectives and diluting expert signals, motivating the need for separate, objective-specific reward models learned from the appropriate source of feedback. This creates a crucial gap: How can agents learn optimal trade-offs between conflicting objectives when the reward functions are unknown and must be inferred from feedback? To address this gap, we introduce LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback, a framework for learning to balance multiple objectives without pre-defined reward functions. Our approach jointly learns policies and multiple objectivespecific reward models from preference feedback, enabling agents to effectively balance competing objectives during learning. By removing the assumption of known reward functions and explicitly modeling multiple objectives, our method tackles a key challenge to scaling MORL in realworld, human-aligned domains. To summarize, our main contributions are: • We propose a novel framework, LEMUR, which learns multiple objective-specific reward models from preference feedback and then optimizes policies against these reward models using multi-objective reinforcement learning. This enables agents to effectively solve multiobjective decision-making tasks without access to predefined reward functions, while naturally accommodating heterogeneous sources of feedback, where different annotators may hold expertise over different objectives. • Our extensive experiments demonstrate that LEMUR outperforms baselines across a range of benchmark multiobjective environments, namely high-dimensional continuous control tasks, and further show LEMUR’s robustness to label noise, constrained feedback budgets, and scaling to more objectives.
2
Preliminaries
Multi-Objective Reinforcement Learning. We formulate the problem as a Multi-Objective Markov Decision Process (MOMDP) (White 1982), defined by the tuple ⟨S, A, T , γ, r⟩. Here, S and A denote state and action spaces respectively, T the transition dynamics, and γ ∈ [0, 1) the discount factor. Unlike standard RL, the reward is a vector r(s, a) ∈ Rm comprising m distinct objectives. The agent’s goal is to maximize P∞ the expected discounted vector returns J(π) = Eπ [ t=0 γ t r(st , at )]. A policy π maps states to action distributions. Since no single policy typically maximizes all objectives in MORL, the agent learns a set of policies Π representing optimal trade-offs, defined by Pareto dominance (Hayes et al. 2022a). A policy π dominates π ′ (denoted J(π) ≻ J(π ′ )) if it is superior in at least one objective and no worse in others. The solution set is the Pareto Frontier F = {π ∈ Π | ∄π ′ ∈ Π : J(π ′ ) ≻ J(π)}. Soft Actor-Critic (SAC). SAC (Haarnoja et al. 2018) is an off-policy actor-critic algorithm grounded in the maximum entropy framework. It augments standard single-objective RL with an entropy term to encourage exploration. The agent P aims to maximize J(π) = Eπ [ t γ t (rt + αH(π(·|st )))]. Reward Learning from Preference Feedback. We follow the standard Preference-based RL (PbRL) framework,
which uses ‘latent rewards’ as a proxy for values (Christiano et al. 2017). Each human’s reward function is learned independently using their respective preference feedback. Similar to prior work, we use preferences over pairs of trajectory segments, (σ 0 , σ 1 ). An expert teacher (e.g., human) provides a label y ∈ {0, 1, 0.5} to indicate their preference. To learn a reward function r̂ parameterized by ψ, we employ the Bradley-Terry model (Bradley and Terry 1952), modeling P the preference probability as P [σ 1 ≻ σ 0 ; ψ] = exp r̂(σ 1P ) P exp r̂(σ 1 )+exp r̂(σ 0 ) . Given a dataset D, the reward function is trained by minimizing the cross-entropy loss: h i LCE (ψ) = −ED (1−y) log P [σ 0 ≻ σ 1 ]+y log P [σ 1 ≻ σ 0 ] . (1) The learned reward function can then be used to update the policy with any RL algorithm to maximize expected returns.
3
Problem Setup
In this section, we present our formulation of Multi-Objective Reinforcement Learning (MORL) without pre-defined reward functions for multiple, conflicting objectives. Latent Reward Vector. In standard MORL, the reward function is a known vector r(s, a) ∈ Rm . However, in our work, the agent does not have access to the ground-truth reward function. Instead, we assume the existence of m distinct multiple, conflicting human users (objectives), where each dimension ri corresponds to the latent reward function of the i-th specific user. Since these rewards are inaccessible, we must approximate them. We define a parameterized reward vector r̂ψ (s, a) = [r̂ψ1 (s, a), . . . , r̂ψm (s, a)]T , where each component is a reward model learned from human preference feedback (Christiano et al. 2017). Multi-Objective RL Optimization. The agent’s goal is to maximize the expected discounted vector returns. We adopt the most prevalent MORL formalism called the utility-based approach (Roijers et al. 2013), where we define a scalarizam tion function fw (r) = w⊤ r, where w P∈ R is a preference weight vector on the simplex (i.e., wi = 1). Thus, we define an optimal MORL agent to be the policies belonging to the Convex Coverage Set (CCS), the subset of F optimal for linearly scalarized preferences (Roijers et al. 2013). To summarize, the MORL agent learns these policies by maximizing the expected returns of the learned latent reward vector linearly scalarized by the weight w, which determines the trade-offs between objectives: "∞ # X J(π) = Eπ γ t wT r̂ψ (st , at ) . (2) t=0
4
LEMUR
LEMUR (Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback), illustrated in Figure 1 proceeds in three stages: (1) unsupervised pre-training, where the agent explores via intrinsic rewards to collect diverse experiences (Section 4.1); (2) reward learning, where multiple human teachers are queried for preference feedback to train objective-specific reward models
Figure 1: Illustration of our framework LEMUR: (1) Unsupervised Pre-training for the MORL agent to explore and collect diverse experiences via maximizing state entropy H(s). (2) Reward learning from Preference feedback, where each reward model is learned separately from the preferences queried from each teacher. The reward models are used to dynamically relabel the state-action pairs as a reward vector for each objective (i.e., each teacher’s preferences). (3) Multi-Objective RL agent denoted by πϕ uses each of the trained reward models to do multi-objective policy optimization to maximize the expected vector rewards.
(Section 4.2); and (3) multi-objective RL training against the learned reward models (Section 4.3). Stages 2 and 3 repeat, continually improving both the reward models and the multi-objective policies. Full pseudocode is provided in Appendix B.
4.1
Unsupervised Pre-training
Standard PbRL suffers from uninformative queries caused by the limited coverage of random initialization. To generate informative queries, LEMUR employs an unsupervised pretraining phase driven by intrinsic motivation (Lee, Smith, and Abbeel 2021). We encourage exploration by maximizing state entropy, approximated via a particle-based k-nearest neighbors (k-NN) estimator (Liu and Abbeel 2021). The intrinsic reward rint (st ) = log(∥st − skt ∥) is the normalized distance to the k-th nearest neighbor in B, and the agent maxPT imizes Jint (ϕ) = Eπϕ [ t=0 γ t rint (st )]. This populates the buffer with diverse behaviors, accelerating the subsequent multi-objective reward learning.
4.2
Reward learning of Multiple Objectives from Preferences
A core challenge in our setup is that the agent does not have access to the ground-truth rewards, and the m conflicting objectives are characterized instead by the conflicting preferences of m humans. Following the PbRL formulation in Section 2, LEMUR learns a separate reward model r̂ψj (s, a) per teacher, trained by minimizing the cross-entropy loss between the model’s predictions and that teacher’s labels (Equation 1). Weight-Conditioned Reward Models. Rather than learning each teacher’s reward in isolation, we condition every objective-specific model on the shared objective space. Each teacher j is assigned a reward model r̂ψj (s, a), a lightweight
MLP predicting the full objective vector, whose scalar utility is obtained by projecting onto that teacher’s preference anchor, r̂j (s, a) = a⊤ j r̂ψj (s, a). This couples the m learned models to a common vector-reward structure, so that a policy conditioned on w reads a consistent per-teacher utility at inference, and the reward models remain directly comparable as the number of conflicting teachers grows. We deliberately adopt this simple architecture, as in prior approaches (Mu, Luan, and Jia 2025); we find it sufficient to recover strong compromise policies while keeping reward learning fast enough to remain in the loop with online policy optimization. Query Sampling Strategy. We sample trajectory pairs (σ 0 , σ 1 ) uniformly at random from the buffer B, so that queries span the diverse state-action distributions explored by all policies. While more sophisticated disagreementbased strategies exist, uniform sampling offers simplicity and avoids bias toward particular regions of the objective space during early training.
4.3
Multi-Objective RL Training
Given the parameterized reward vector r̂ψ (s, a), LEMUR trains the MORL agent to maximize expected latent vector rewards (Equation 2). For policy optimization we leverage MORL/D, a state-of-the-art Multi-Objective Soft ActorCritic (MO-SAC) algorithm (Felten, Talbi, and Danoy 2024) which learns a set of independent SAC policies and applies an evolutionary strategy for policy search. This off-policy choice is deliberate: reusing past experience from the replay buffer is essential for sample efficiency under a limited human feedback budget. Weight Vector Initialization and Adaptation. The scalarization weight vectors {w} determine the trade-offs between objectives. Given w ∈ Rm , we optimize for policies using SAC (Haarnoja et al. 2018) on the scalarized reward of
Equation 2. We employ a Pareto Simulated Annealing (PSA) approach similar to (Felten, Talbi, and Danoy 2024), adapting the weight vector in response to the current policies and their distance to non-dominated solutions, allowing the agent to focus training on feasible regions of the objective space while maintaining policy diversity. Cooperation via Shared Buffer. To facilitate information exchange across policies learning different trade-offs, all policies store and sample from a common replay buffer B. This enables policies to learn from diverse experiences collected under different preference weightings, improving sample efficiency, a critical consideration given the limited human feedback budget. Relabeling of Vector Rewards. Combining off-policy RL with a reward function learned from preferences introduces non-stationarity: as the reward model is updated with new feedback, rewards associated with past transitions in the buffer become stale, destabilizing learning. To address this, LEMUR employs a vector reward relabeling strategy inspired by prior work (Lee, Smith, and Abbeel 2021). Rather than storing rewards, we store only the transitions, and compute vector rewards on the fly when a batch is sampled, using the most up-to-date reward models. This ensures the agent always trains on updated rewards, synchronizing policy and reward learning while preserving the sample efficiency of our off-policy approach.
5
Experiments
Our experiments address three questions: (1) Can LEMUR learn multi-objective policies that balance multiple learned reward models from multiple teachers? (2) How does LEMUR compare to existing baselines on multi-objective benchmarks? (3) Does explicitly learning multiple reward models for conflicting feedback outperform aggregating feedback into a single reward model? For all experiments, we report the mean across five random seeds with standard error. Additional implementation details are reported in Appendices D & E. Benchmark Environments & Setup. We evaluate LEMUR on high-dimensional environments from the MORL-Generalization benchmark (Teoh, Varakantham, and Vamplew 2025). Following standard practice in PbRL (Lee et al. 2021; Christiano et al. 2017), we use scripted teachers that generate feedback according to the components of the ground-truth vector reward, enabling quantitative evaluation; the ground-truth rewards remain inaccessible to the agent, which must jointly learn the conflicting preferences and optimize to find a balance. Our main experiments use two conflicting teachers, the fundamental version of the problem; Section 5.1 demonstrates scaling to more objectives. We evaluate on: MO-Lunarlander, where Teacher A rewards precise, stable landings and Teacher B prioritizes fuel conservation; MO-Hopper and MO-Cheetah, continuous control locomotion tasks where Teacher A prefers fast locomotion and Teacher B prefers slow, energy-efficient gaits; and MO-MetaWorld (Drawer-Close), a robotic-manipulation task from the Meta-World suite (Yu et al. 2020). Meta-World tasks are natively single-objective; we convert Drawer-Close into a two-objective task by pairing the native task-progress
reward (Teacher A) with a control-effort penalty (Teacher B), mirroring the reward decomposition standard in the MORL benchmark suite (Teoh, Varakantham, and Vamplew 2025). Full environment details are given in Appendix G. Baselines. We compare against five baselines spanning distinct strategies for preference aggregation and learning from multiple objectives: (1) a Utilitarian agent, a single SAC agent optimizing the arithmetic mean of the independently learned rewards; (2) Naive data pooling, which trains one monolithic reward model on all conflicting feedback, akin to standard Reinforcement Learning from Human Feedback (RLHF); (3) MORAL (Peschl et al. 2022), which recovers per-teacher rewards via Adversarial Inverse Reinforcement Learning (AIRL) and learns a scalarization over them; (4) PbMORL (Mu, Luan, and Jia 2025), a recent preference-based multi-objective method learning a weightconditioned vector reward from pairwise feedback; and (5) FPbRL (Siddique, Sinha, and Cao 2023), which aggregates learned per-teacher rewards through a Generalized Gini Welfare scalarization to optimize for fairness. We additionally report an (6) Oracle trained on ground-truth rewards as an upper bound. Unless otherwise noted, every baseline shares LEMUR’s interactive learning loop, reward-model architecture, pre-training stage, teacher weight vectors, query budget, and environment-step budget; the primary distinction lies in how conflicting reward signals are aggregated and optimized. Where a baseline’s original policy optimizer would disadvantage it in our environments, we adapt in the baseline’s favour; all deviations are disclosed in Appendix E.
5.1
Results & Analysis
Figure 2 presents the learning curves for all methods across the benchmark environments. Across every environment, LEMUR is the method that most closely tracks the Oracle on both objectives simultaneously, effectively recovering policies that balance multiple, conflicting objectives. The Utilitarian and Naive agents remain flat with suboptimal returns throughout, supporting our hypothesis that aggregating conflicting reward signals into a single scalar degrades performance in such multi-objective settings. While MORAL improves early in LunarLander and Hopper, it fails to sustain this progress: in MO-Hopper its returns peak mid-training and then steadily declines. This is consistent with the known failure of out-of-distribution optimization due to static reward models; MORAL infers its per-teacher rewards offline from expert demonstrations via AIRL and holds them fixed. LEMUR instead does online policy optimization and addresses non-stationarity through vector reward relabeling. PbMORL and FPbRL, which both learn vector rewards from preference feedback, perform better than the aggregation baselines, yet still fall short of LEMUR. We attribute this to the assumptions both methods inherit. PbMORL trains a single weight-conditioned reward model over pooled feedback, implicitly assuming all preferences originate from one single teacher; under conflicting teachers the pooled model must average over conflicting labels, degrading the reward signal, most visibly in MO-Hopper and MO-Cheetah where it consistently trails LEMUR on both objectives. FPbRL
Episode Returns (Objective One)
LunarLander
Cheetah
70
1200
5000
1000
4000
80
800
110
20
40
60
200
1000
500
0
0
0
2
4
6
1400
4500
1200
4000
20
40
60
Training Steps (×10³)
0
4
1000
800
2500
800
600
2000
0
80
2
4
6
Ground Truth (Oracle) FPbRL
0
1
3
4
5
2
3
4
5
200
500
Training Steps (×10 )
2
400
1000
0
1
600
1500
0
0
1400 1200
200 0
2
3000
400 80
0
3500
1000 70
1500 1000
80
60
2000
2000
400
0
2500
3000
600
90
MetaWorld
3000
6000
1400
100
Episode Returns (Objective Two)
Hopper
1600
60
0
2
Training Steps (×10 )
LEMUR
Utilitarian
MORAL Naive
0
4
Training Steps (×10 )
PbMORL
Figure 2: Learning curves on all benchmark environments: MO-LunarLander, MO-Hopper, MO-Cheetah, and MO-MetaWorld. Curves depict the true objective returns (inaccessible to the agent), averaged across five seeds, with shaded regions representing standard error. LEMUR (blue) most closely tracks the Oracle (red) on both objectives simultaneously. Baselines that aggregate conflicting feedback (Naive, Utilitarian) fail to make progress, while the external baselines (MORAL, PbMORL, FPbRL) learn but consistently trail LEMUR. 100
LEMUR
Success Rate (%)
Environment
Ground Truth (Oracle) PbMORL FPbRL Naive Utilitarian MORAL
80
60
MO-LunarLander MO-Hopper MO-HalfCheetah MO-MetaWorld
40
Hypervolume (↑)
Sparsity (↓)
LEMUR
PbMORL
LEMUR PbMORL
1.10 × 104 3.67 × 106 4.86 × 107 2.15 × 106
1.09 × 104 2.24 × 106 4.78 × 107 1.43 × 106
134.5 294.7 294.7 436.3
7.8 1731.7 2191.0 5564.7
20
0
0
1
2
3
Training Steps (×10 )
4
5
Figure 3: Task Success Rate (%) learning curves on MetaWorld (Drawer Close Task).
preserves the vector structure but employs a fixed Generalized Gini welfare scalarization a priori, converging to a single welfare-optimal policy rather than a set of trade-offs: it achieves reasonable returns on MetaWorld, but fails to achieve task success (shown in Figure 3) nor make progress on the other environments. LEMUR avoids both failure modes by maintaining objective-specific reward models and adapting the trade-off online. On MO-MetaWorld, Figure 3 reports the task success rate: LEMUR reaches closest to the Oracle, while PbMORL plateaus. Multi-Objective Metrics. We evaluate LEMUR using standard multi-objective metrics (Hayes et al. 2022a). Hypervolume (HV) measures the volume of objective space, rewarding policies that are both high-performing and broadly spread (Teoh, Varakantham, and Vamplew 2025); Sparsity
Table 1: Hypervolume and Sparsity for LEMUR and PbMORL. Full results are in Appendix F.1.
(SPS) measures the average distance between policies along the front, with lower values indicating more uniform coverage (Teoh, Varakantham, and Vamplew 2025). As summarised in Table 1, LEMUR attains the highest or comparable Hypervolume across all four environments while achieving markedly lower sparsity than PbMORL, indicating that it learns policies that is both higher-performing and more uniformly distributed over the trade-off space. MORAL is excluded from these set-based metrics, as its single-objective policy optimization against a scalarized reward recovers only one solution rather than a front. Full results are reported in Appendix F.1. Reward Model Alignment. We additionally evaluate the learned reward models directly against the ground-truth teacher rewards, following established PbRL evaluation practice (Lee et al. 2021). We report Spearman rank correlation, which measures how accurately the learned reward models rank individual states compared to the teachers’ ground-truth
reward; the Trajectory Alignment Coefficient (TAC) (Muslimani et al. 2025), which compares rankings over whole trajectories rather than individual transitions. Table 2 shows that LEMUR’s reward models recover their teachers’ preference orderings with consistently strong correlation, and outperform both PbMORL and FPbRL (Appendix F.2). Environment
Spearman (ρ ↑)
TAC (↑)
MO-Hopper MO-HalfCheetah MO-MetaWorld
0.945 ± 0.002 0.710 ± 0.005 0.520 ± 0.007
0.856 ± 0.014 0.898 ± 0.041 0.347 ± 0.026
7000 6000 5000 4000 3000 2000 1000 0
Episode Returns (Objective One)
Episode Returns (Objective One)
Table 2: LEMUR reward model alignment, reporting Spearman rank correlation and Trajectory Alignment Coefficient (TAC). Per-metric comparisons against PbMORL and FPbRL are in Appendix F.2.
6000
Episode Returns (Objective Two)
Episode Returns (Objective Two)
5000 4000 3000 2000 1000
4000 3000 2000
Episode Returns (Objective Three)
0
2000
1000
0
1
2
3
Training Steps (×10 )
4
LEMUR
Ground-Truth Reward (Oracle)
3000 2000 1000 0
5
Episode Returns (Objective Four)
Episode Returns (Objective Three)
5000
1000
0
0
7000 6000 5000 4000 3000 2000 1000 0
2000
1000
0
0
1
2
Training Steps (×10 )
3
4
Figure 4: LEMUR scalability to higher-dimensional objective spaces. Episode returns for (a) 3-objective and (b) 4objective tasks, comparing policies trained with LEMUR (blue) versus ground-truth oracle rewards (orange). Scaling to More Objectives. We next vary the teacher configuration on MO-Cheetah. Figure 4 extends LEMUR to three and four conflicting teachers: the learned-reward policies closely track the ground-truth oracle across all objectives, demonstrating the per-teacher decomposition scales without modification, each additional objective adding one reward model.
5.2
Ablation Studies
To validate the components of LEMUR, we conduct ablation studies on the high-dimensional MO-Cheetah domain; Figure 5 visualizes the learning curves. Impact of Shared Buffer, Relabeling, and Pre-training. Figure 5(a) isolates the contributions of the shared replay buffer, vector reward relabeling, and unsupervised pre-
training. Disabling the shared buffer causes the most severe degradation. With the shared buffer intact, removing relabeling alone produces a modest but consistent drop relative to full LEMUR, confirming that recomputing rewards under the current models stabilizes learning against reward nonstationarity. Removing pre-training accelerates the earliest phase of training, but converges to lower final returns, indicating that the diverse initial buffer ultimately yields better reward models and policies. Robustness to Label Noise. Figure 5(b) corrupts a fraction of teacher labels (flipping preferences with probability up to 15%), following PbRL benchmarking protocol (Lee et al. 2021). LEMUR degrades gracefully: performance is essentially unaffected up to 10% noise, and at 15% the agent still learns effective compromise policies on both objectives, albeit with slower convergence and higher variance, indicating tolerance to levels of annotator error. Impact of Feedback Budget. Figure 5(c) varies the total query budget from 260 to 5,200 per teacher. Performance improves with budget, and larger budgets learn faster; notably, even 260 total learns adequate policies on both objectives. This feedback efficiency is particularly important in preference-based RL, where human queries are limited. Reward Model Ablation. To verify that LEMUR’s gains stem from its weight-conditioned reward model rather than the surrounding pipeline, we re-run LEMUR replacing this component with an ensemble of three unconditioned MLPs (Christiano et al. 2017; Lee, Smith, and Abbeel 2021), holding all else fixed. The weight-conditioned variant converges to 6,812 ± 39 and 4,404 ± 22 on the two objectives, against 4,556 ± 369 and 2,902 ± 245 for the ensemble. (For full results, refer to Appendix C.7). Additional Experiments. In Appendices C.1-C.7, we demonstrate that LEMUR accommodates changing teachers/objectives mid-training, varying levels of conflict between teachers, and also non-stationary preferences without reinitialization. Query ablations reveal that segment length is important to performance, and we verify that LEMUR maintains performance even when teachers are in agreement with overlapping preferences.
6
Related Work
Reward Learning from Preference Feedback. Designing reward functions is a primary bottleneck in scaling RL, as manual crafting is impractical and can induce unsafe behavior (Amodei et al. 2016); prior work instead learns rewards from demonstrations (Ng and Russell 2000; Abbeel and Ng 2004), language (Lin et al. 2022), or human feedback (Christiano et al. 2017). Preference-based RL (PbRL) learns rewards from pairwise comparisons (Christiano et al. 2017; Lee, Smith, and Abbeel 2021); popularized as RLHF (Stiennon et al. 2020; Ouyang et al. 2022), it typically trains a single reward model, aggregating diverse feedback into one scalar (Ouyang et al. 2022) and failing to capture the multi-objective nature of human values (Sorensen et al. 2024). Offline variants train rewards on fixed datasets before policy optimization (Shin, Dragan, and Brown 2023), but static models suffer distribution shift as the policy diverges from the offline coverage, causing reward exploitation (Gao, Schulman, and Hilton
No Relabel No Shared Buffer LEMUR No Shared Buffer, No Relabel No Pretrain
Episode Returns (Objective Two)
Episode Returns (Objective One)
6000 5000 4000
0% Noise 5% Noise 10% Noise 15% Noise
6000 5000
6000 5000
4000
4000
3000
3000
2000
2000
2000
1000
1000
1000
0
0
0
4000
4000
4000
3000
3000
3000
2000
2000
2000
1000
1000
1000
0
0
1
2
3
Training Steps (×10 )
(a) Buffer Relabeling
4
5
0
260 Total Queries 1,300 Total Queries 3,900 Total Queries 5,200 Total Queries
7000
3000
0
1
2
3
Training Steps (×10 )
4
5
6
(b) Noisy Labels
0
0
1
2
3
4
5
Training Steps (×10 )
6
7
8
(c) Query Budget
Figure 5: Ablation studies on MO-Cheetah evaluating the impact of (a) the shared buffer, vector reward relabeling, and unsupervised pre-training, (b) noisy teacher labels, and (c) varying the total query budget per teacher on agent returns for both objectives; by default LEMUR uses 3900 queries (green). The results are averaged over multiple runs across five seeds.
2023; Ye et al. 2024). Online, iterative RLHF mitigates this by collecting feedback alongside policy optimization (Dong et al. 2024; Gao, Schulman, and Hilton 2023; Ye et al. 2024), which is more critical still in MORL, where the agent must span a space of diverse policies (Hayes et al. 2022a); offline MORL instead presupposes specified rewards or adequate dataset coverage (Yuan et al. 2024; Zhu, Dang, and Grover 2023). LEMUR circumvents this by jointly and interactively optimizing both the reward models and the multi-objective policy online. Multi-Objective Reinforcement Learning (MORL). MORL learns a set of policies approximating the Pareto frontier (Roijers et al. 2013; Hayes et al. 2022b), via single-policy scalarization, weight-conditioned, or multi-policy methods; multi-task and Meta-RL are closely related (Chen et al. 2019; Abdolmaleki et al. 2020; Sener and Koltun 2018). Most of this literature assumes a vector of ground-truth reward functions (Hayes et al. 2022b), which is impractical in complex, real-world tasks. Reward-free MORL (Chen et al. 2026) relaxes this only partially, using reward-free exploration as an auxiliary objective while still assuming an extrinsically specified ground-truth reward. LEMUR extends the MORL paradigm to the setting where the objectives are never observed and must be inferred directly from preferences. Learning & Alignment with Diverse Objectives. Standard RLHF fails to capture the pluralistic nature of human values (Sorensen et al. 2024), and existing remedies rely on manual aggregation functions (Rodriguez-Soto et al. 2023), expensive consensus datasets (Tessler et al. 2024), static offline learning (Bakker et al. 2022), or model heterogeneous feedback as hidden context without optimizing the trade-off between preferences (Siththaranjan, Laidlaw, and Hadfield-Menell 2024). Our closest baselines learn rewards for multiple objectives but inherit strong coherence assumptions: MORAL (Peschl et al. 2022) requires expert demonstrations and freezes AIRL-learned rewards (Fu, Luo, and Levine 2018), PbMORL (Mu, Luan, and Jia 2025) assumes a single teacher, and FPbRL (Siddique, Sinha, and
Cao 2023) fixes a welfare scalarization a priori. LEMUR instead learns objective-specific reward models from separate feedback streams and jointly optimizes a multi-objective policy online via vector reward relabeling, without offline pre-training or expert demonstrations.
7
Conclusion & Limitations
We propose LEMUR, a framework for Multi-Objective RL in domains where reward functions are unknown and must be inferred from the conflicting preferences of multiple teachers. LEMUR jointly learns the objectives and the policies that balance them, without expert demonstrations, pre-defined rewards, or a priori aggregation rules that existing methods require. Across multi-objective RL control and robotic manipulation benchmark environments, LEMUR outperforms aggregation baselines and recent preference-based multiobjective methods, and remains robust to label noise, reduced feedback budgets, and scaling to additional objectives. Several directions for future work are as follows. Our evaluation uses scripted teachers, standard practice in PbRL for controlled and reproducible comparison (Lee et al. 2021); our noise ablations suggest the framework tolerates the label error real annotators exhibit, and a human study is the natural next validation. We adopt linear scalarization, and extending to non-linear scalarization would allow the framework to learn policies in non-convex regions of the front (Roijers et al. 2013; Hayes et al. 2022b). Finally, active querying (Akrour, Schoenauer, and Sebag 2012) offers a route to further reducing the number of queries, which is a promising path for deploying preference-based MORL in real-time.
References Abbeel, P.; and Ng, A. Y. 2004. Apprenticeship learning via inverse reinforcement learning. In Proceedings of the twenty-first international conference on Machine learning, 1. Abdolmaleki, A.; Huang, S.; Hasenclever, L.; Neunert, M.; Song, F.; Zambelli, M.; Martins, M.; Heess, N.; Hadsell, R.;
and Riedmiller, M. 2020. A distributional view on multiobjective policy optimization. In International conference on machine learning, 11–22. PMLR. Akrour, R.; Schoenauer, M.; and Sebag, M. 2012. April: Active preference learning-based reinforcement learning. In Joint European conference on machine learning and knowledge discovery in databases, 116–131. Springer. Amodei, D.; Olah, C.; Steinhardt, J.; Christiano, P.; Schulman, J.; and Mané, D. 2016. Concrete problems in AI safety. arXiv preprint arXiv:1606.06565. Bahlous-Boldi, R.; Puri, I.; Shenfeld, I.; Kumar, A.; Damani, M.; Risi, S.; Khattab, O.; Hong, Z.-W.; and Agrawal, P. 2026. Vector Policy Optimization: Training for Diversity Improves Test-Time Search. arXiv preprint arXiv:2605.22817. Bai, Y.; Jones, A.; Ndousse, K.; Askell, A.; Chen, A.; DasSarma, N.; Drain, D.; Fort, S.; Ganguli, D.; Henighan, T.; et al. 2022. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862. Bakker, M.; Chadwick, M.; Sheahan, H.; Tessler, M.; Campbell-Gillingham, L.; Balaguer, J.; McAleese, N.; Glaese, A.; Aslanides, J.; Botvinick, M.; et al. 2022. Finetuning language models to find agreement among humans with diverse preferences. Advances in neural information processing systems, 35: 38176–38189. Bowling, M.; Martin, J. D.; Abel, D.; and Dabney, W. 2023. Settling the reward hypothesis. In International Conference on Machine Learning, 3003–3020. PMLR. Bradley, R. A.; and Terry, M. E. 1952. Rank analysis of incomplete block designs: I. the method of paired comparisons. Biometrika, 39(3/4): 324–345. Chen, X.; Ghadirzadeh, A.; Björkman, M.; and Jensfelt, P. 2019. Meta-learning for multi-objective reinforcement learning. In 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 977–983. IEEE. Chen, Y.-T.; Hung, W.; Wu, B.-S.; Hong, Z.-W.; and Hsieh, P.-C. 2026. A Reward-Free Viewpoint on Multi-Objective Reinforcement Learning. In The Fourteenth International Conference on Learning Representations. Christiano, P. F.; Leike, J.; Brown, T.; Martic, M.; Legg, S.; and Amodei, D. 2017. Deep Reinforcement Learning from Human Preferences. In Guyon, I.; Luxburg, U. V.; Bengio, S.; Wallach, H.; Fergus, R.; Vishwanathan, S.; and Garnett, R., eds., Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc. Ding, L.; Zhang, J.; Clune, J.; Spector, L.; and Lehman, J. 2024. Quality Diversity through Human Feedback: Towards Open-Ended Diversity-Driven Optimization. In Forty-first International Conference on Machine Learning. Dong, H.; Xiong, W.; Pang, B.; Wang, H.; Zhao, H.; Zhou, Y.; Jiang, N.; Sahoo, D.; Xiong, C.; and Zhang, T. 2024. RLHF Workflow: From Reward Modeling to Online RLHF. Transactions on Machine Learning Research. Dulac-Arnold, G.; Mankowitz, D.; and Hester, T. 2019. Challenges of Real-World Reinforcement Learning.
Felten, F.; Talbi, E.-G.; and Danoy, G. 2024. Multi-Objective Reinforcement Learning Based on Decomposition: A Taxonomy and Framework. Journal of Artificial Intelligence Research, 79: 679–723. Fickinger, A.; Zhuang, S.; Hadfield-Menell, D.; and Russell, S. 2020. Multi-principal assistance games. arXiv preprint arXiv:2007.09540. Fu, J.; Luo, K.; and Levine, S. 2018. Learning Robust Rewards with Adverserial Inverse Reinforcement Learning. In International Conference on Learning Representations. Gao, L.; Schulman, J.; and Hilton, J. 2023. Scaling laws for reward model overoptimization. In International Conference on Machine Learning, 10835–10866. PMLR. Gunjal, A.; Wang, A.; Lau, E.; Nath, V.; He, Y.; Liu, B.; and Hendryx, S. M. 2025. Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains. In NeurIPS 2025 Workshop on Efficient Reasoning. Haarnoja, T.; Zhou, A.; Abbeel, P.; and Levine, S. 2018. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In International conference on machine learning, 1861–1870. Pmlr. Hayes, C. F.; Rădulescu, R.; Bargiacchi, E.; Källström, J.; Macfarlane, M.; Reymond, M.; Verstraeten, T.; Zintgraf, L. M.; Dazeley, R.; Heintz, F.; Howley, E.; Irissappane, A. A.; Mannion, P.; Nowé, A.; Ramos, G.; Restelli, M.; Vamplew, P.; and Roijers, D. M. 2022a. A practical guide to multiobjective reinforcement learning and planning. Auton. Agent. Multi. Agent. Syst., 36(1). Hayes, C. F.; Rădulescu, R.; Bargiacchi, E.; Källström, J.; Macfarlane, M.; Reymond, M.; Verstraeten, T.; Zintgraf, L. M.; Dazeley, R.; Heintz, F.; Howley, E.; Irissappane, A. A.; Mannion, P.; Nowé, A.; Ramos, G.; Restelli, M.; Vamplew, P.; and Roijers, D. M. 2022b. A Practical Guide to Multi-Objective Reinforcement Learning and Planning. Autonomous Agents and Multi-Agent Systems, 36(1): 26. ArXiv:2103.09568 [cs]. Huang, S.; Abdolmaleki, A.; Vezzani, G.; Brakel, P.; Mankowitz, D. J.; Neunert, M.; Bohez, S.; Tassa, Y.; Heess, N.; Riedmiller, M.; et al. 2022. A constrained multi-objective reinforcement learning framework. In Conference on Robot Learning, 883–893. PMLR. Ibarz, B.; Leike, J.; Pohlen, T.; Irving, G.; Legg, S.; and Amodei, D. 2018. Reward learning from human preferences and demonstrations in atari. Advances in neural information processing systems, 31. Kirk, R.; Mediratta, I.; Nalmpantis, C.; Luketina, J.; Hambro, E.; Grefenstette, E.; and Raileanu, R. 2024. Understanding the Effects of RLHF on LLM Generalisation and Diversity. In The Twelfth International Conference on Learning Representations. Knox, W. B.; Glass, B. D.; Love, B. C.; Maddox, W. T.; and Stone, P. 2012. How humans teach agents: A new experimental perspective. International Journal of Social Robotics, 4(4): 409–421. Kouritem, S. A.; Abouheaf, M. I.; Nahas, N.; and Hassan, M. 2022. A multi-objective optimization design of indus-
trial robot arms. Alexandria Engineering Journal, 61(12): 12847–12867. Lee, K.; Smith, L.; Dragan, A.; and Abbeel, P. 2021. B-Pref: Benchmarking Preference-Based Reinforcement Learning. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1). Lee, K.; Smith, L. M.; and Abbeel, P. 2021. PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training. In Meila, M.; and Zhang, T., eds., Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, 6152–6163. PMLR. Lin, J.; Fried, D.; Klein, D.; and Dragan, A. 2022. Inferring Rewards from Language in Context. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8546–8560. Liu, H.; and Abbeel, P. 2021. Behavior from the void: Unsupervised active pre-training. Advances in Neural Information Processing Systems, 34: 18459–18473. Ma, Y. J.; Hejna, J.; Fu, C.; Shah, D.; Liang, J.; Xu, Z.; Kirmani, S.; Xu, P.; Driess, D.; Xiao, T.; Bastani, O.; Jayaraman, D.; Yu, W.; Zhang, T.; Sadigh, D.; and Xia, F. 2025. Vision Language Models are In-Context Value Learners. In Yue, Y.; Garg, A.; Peng, N.; Sha, F.; and Yu, R., eds., International Conference on Learning Representations, volume 2025, 33984–34009. McCarthy, J. 1997. What is Artificial Intelligence? Stanford University. Also available in a 2007 version at http://wwwformal.stanford.edu/jmc/whatisai.pdf. Mnih, V.; Kavukcuoglu, K.; Silver, D.; Graves, A.; Antonoglou, I.; Wierstra, D.; and Riedmiller, M. A. 2013. Playing Atari with Deep Reinforcement Learning. CoRR, abs/1312.5602. Mu, N.; Luan, Y.; and Jia, Q.-S. 2025. Preference-based multi-objective reinforcement learning. IEEE Transactions on Automation Science and Engineering. Munos, R.; Valko, M.; Calandriello, D.; Gheshlaghi Azar, M.; Rowland, M.; Guo, Z. D.; Tang, Y.; Geist, M.; Mesnard, T.; Fiegel, C.; Michi, A.; Selvi, M.; Girgin, S.; Momchev, N.; Bachem, O.; Mankowitz, D. J.; Precup, D.; and Piot, B. 2024. Nash Learning from Human Feedback. In Salakhutdinov, R.; Kolter, Z.; Heller, K.; Weller, A.; Oliver, N.; Scarlett, J.; and Berkenkamp, F., eds., Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, 36743–36768. PMLR. Muslimani, C.; Johnstonbaugh, K.; Chandramouli, S.; Booth, S.; Knox, W. B.; and Taylor, M. E. 2025. Towards Improving Reward Design in RL: A Reward Alignment Metric for RL Practitioners. In Reinforcement Learning Conference. Ng, A. Y.; and Russell, S. J. 2000. Algorithms for Inverse Reinforcement Learning. In Proceedings of the Seventeenth International Conference on Machine Learning, 663–670. Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al.
2022. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35: 27730–27744. Pásztor, B.; Buening, T. K.; and Krause, A. 2025. Stackelberg Learning from Human Feedback: Preference Optimization as a Sequential Game. In NeurIPS 2025 Workshop: Second Workshop on Aligning Reinforcement Learning Experimentalists and Theorists. Peschl, M.; Zgonnikov, A.; Oliehoek, F. A.; and Siebert, L. C. 2022. MORAL: Aligning AI with Human Norms through Multi-Objective Reinforced Active Learning. In Proceedings of the 21st International Conference on Autonomous Agents and Multiagent Systems, 1038–1046. Pierrot, T.; Richard, G.; Beguir, K.; and Cully, A. 2022. Multi-objective quality diversity optimization. In Proceedings of the genetic and evolutionary computation conference, 139–147. Rafailov, R.; Sharma, A.; Mitchell, E.; Manning, C. D.; Ermon, S.; and Finn, C. 2023. Direct preference optimization: Your language model is secretly a reward model. Advances in neural information processing systems, 36: 53728–53741. Rame, A.; Couairon, G.; Dancette, C.; Gaya, J.-B.; Shukor, M.; Soulier, L.; and Cord, M. 2023. Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards. Advances in Neural Information Processing Systems, 36: 71095–71134. Rodriguez-Soto, M.; Rodriguez-Aguilar, J. A.; LopezSanchez, M.; and Nowé, A. 2023. Multi-objective reinforcement learning for guaranteeing alignment with multiple values. In 2023 Adaptive and Learning Agents Workshop at AAMAS. Roijers, D. M.; Vamplew, P.; Whiteson, S.; and Dazeley, R. 2013. A survey of multi-objective sequential decisionmaking. Journal of Artificial Intelligence Research, 48(1): 67–113. Ross, S.; Gordon, G.; and Bagnell, D. 2011. A reduction of imitation learning and structured prediction to no-regret online learning. In Proceedings of the fourteenth international conference on artificial intelligence and statistics, 627–635. JMLR Workshop and Conference Proceedings. Sener, O.; and Koltun, V. 2018. Multi-task learning as multiobjective optimization. Advances in neural information processing systems, 31. Shen, W. F.; Qiu, X.; Whitehouse, C.; Alazraki, L.; Goel, S.; Barbieri, F.; Willi, T.; Mathur, A.; and Leontiadis, I. 2026. Rethinking Rubric Generation for Improving LLM Judge and Reward Modeling for Open-ended Tasks. arXiv preprint arXiv:2602.05125. Shin, D.; Dragan, A. D.; and Brown, D. S. 2023. Benchmarks and Algorithms for Offline Preference-Based Reward Learning. ArXiv:2301.01392 [cs]. Siddique, U.; Sinha, A.; and Cao, Y. 2023. Fairness in Preference-based Reinforcement Learning. In ICML 2023 Workshop The Many Facets of Preference-Based Learning. Silver, D.; Singh, S.; Precup, D.; and Sutton, R. S. 2021. Reward is enough. Artificial intelligence, 299: 103535.
Siththaranjan, A.; Laidlaw, C.; and Hadfield-Menell, D. 2024. Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF. In The Twelfth International Conference on Learning Representations. Skalse, J. M. V.; and Abate, A. 2023. The Reward Hypothesis is False. Sorensen, T.; Moore, J.; Fisher, J.; Gordon, M. L.; Mireshghallah, N.; Rytting, C. M.; Ye, A.; Jiang, L.; Lu, X.; Dziri, N.; Althoff, T.; and Choi, Y. 2024. Position: A Roadmap to Pluralistic Alignment. In Salakhutdinov, R.; Kolter, Z.; Heller, K.; Weller, A.; Oliver, N.; Scarlett, J.; and Berkenkamp, F., eds., Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, 46280–46302. PMLR. Stiennon, N.; Ouyang, L.; Wu, J.; Ziegler, D.; Lowe, R.; Voss, C.; Radford, A.; Amodei, D.; and Christiano, P. F. 2020. Learning to summarize with human feedback. Advances in neural information processing systems, 33: 3008–3021. Sutton, R. S.; and Barto, A. G. 1998. Reinforcement learning: an introduction. Adaptive computation and machine learning. Cambridge, Mass: MIT Press. ISBN 978-0-26219398-6. Tan, J.; Zhang, T.; Coumans, E.; Iscen, A.; Bai, Y.; Hafner, D.; Bohez, S.; and Vanhoucke, V. 2018. Sim-to-Real: Learning Agile Locomotion For Quadruped Robots. Preprint. Teoh, J.; Varakantham, P.; and Vamplew, P. 2025. On Generalization Across Environments In Multi-Objective Reinforcement Learning. In The Thirteenth International Conference on Learning Representations. Tessler, M. H.; Bakker, M. A.; Jarrett, D.; Sheahan, H.; Chadwick, M. J.; Koster, R.; Evans, G.; Campbell-Gillingham, L.; Collins, T.; Parkes, D. C.; Botvinick, M.; and Summerfield, C. 2024. AI can help humans find common ground in democratic deliberation. Science, 386(6719): eadq2852. Umer, M.; Mohsin, M. A.; Bilal, A.; Chaudhry, A.; Haupt, A.; Koyejo, S.; Fox, E.; and Cioffi, J. M. 2026. General Preference Reinforcement Learning. arXiv preprint arXiv:2605.18721. Vamplew, P.; Dazeley, R.; Berry, A.; Issabekov, R.; and Dekker, E. 2011. Empirical evaluation methods for multiobjective reinforcement learning algorithms. Machine learning, 84(1): 51–80. Vamplew, P.; Smith, B. J.; Källström, J.; Ramos, G.; Rădulescu, R.; Roijers, D. M.; Hayes, C. F.; Heintz, F.; Mannion, P.; Libin, P. J.; et al. 2022. Scalar reward is not enough: A response to silver, singh, precup and sutton (2021). Autonomous Agents and Multi-Agent Systems, 36(2): 41. Van Seijen, H.; Fatemi, M.; Romoff, J.; Laroche, R.; Barnes, T.; and Tsang, J. 2017. Hybrid reward architecture for reinforcement learning. Advances in neural information processing systems, 30. Wang, Z.; Rahmani, S.; Cornelisse, D.; Sarkar, B.; Goldie, A. D.; Foerster, J. N.; and Whiteson, S. 2026. Learning to Drive in New Cities Without Human Demonstrations. In Workshop on Simulation for Autonomous Driving.
White, D. 1982. Multi-objective infinite-horizon discounted Markov decision processes. Journal of mathematical analysis and applications, 89(2): 639–647. Wu, Z.; Hu, Y.; Shi, W.; Dziri, N.; Suhr, A.; Ammanabrolu, P.; Smith, N. A.; Ostendorf, M.; and Hajishirzi, H. 2023. Finegrained human feedback gives better rewards for language model training. Advances in Neural Information Processing Systems, 36: 59008–59033. Xu, J.; Tian, Y.; Ma, P.; Rus, D.; Sueda, S.; and Matusik, W. 2020. Prediction-guided multi-objective reinforcement learning for continuous robot control. In International conference on machine learning, 10607–10616. PMLR. Yang, R.; Pan, X.; Luo, F.; Qiu, S.; Zhong, H.; Yu, D.; and Chen, J. 2024. Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment. In Forty-first International Conference on Machine Learning. Yang, R.; Sun, X.; and Narasimhan, K. 2019. A generalized algorithm for multi-objective reinforcement learning and policy adaptation. Advances in neural information processing systems, 32. Ye, C.; Xiong, W.; Zhang, Y.; Dong, H.; Jiang, N.; and Zhang, T. 2024. Online Iterative Reinforcement Learning from Human Feedback with General Preference Model. In The Thirty-eighth Annual Conference on Neural Information Processing Systems. Yu, T.; Quillen, D.; He, Z.; Julian, R.; Hausman, K.; Finn, C.; and Levine, S. 2020. Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning. In Conference on robot learning, 1094–1100. PMLR. Yuan, Y.; Zheng, Z.; Dong, Z.; and Hao, J. 2024. MODULI: Unlocking Preference Generalization via Diffusion Models for Offline Multi-Objective Reinforcement Learning. ArXiv:2408.15501 [cs]. Zhi-Xuan, T.; Carroll, M.; Franklin, M.; and Ashton, H. 2025. Beyond Preferences in AI Alignment: T. Zhi-Xuan et al. Philosophical Studies, 182(7): 1813–1863. Zhou, Z.; Liu, J.; Shao, J.; Yue, X.; Yang, C.; Ouyang, W.; and Qiao, Y. 2024. Beyond one-preference-fits-all alignment: Multi-objective direct preference optimization. In Findings of the Association for Computational Linguistics: ACL 2024, 10586–10613. Zhu, B.; Dang, M.; and Grover, A. 2023. Scaling ParetoEfficient Decision Making via Offline Multi-Objective RL. In The Eleventh International Conference on Learning Representations.
A
Appendix Extended Related Works
Vector Rewards and the Limits of Scalar Reward. A key premise in modern RL is the reward hypothesis: that goals can be adequately captured by maximizing a single scalar reward (Silver et al. 2021). A growing body of work contests the sufficiency of this framing, arguing that many objectives of interest cannot always be expressed by scalar reward (Vamplew et al. 2022; Skalse and Abate 2023; Bowling et al. 2023; Van Seijen et al. 2017) and are more naturally represented as vectors (Vamplew et al. 2022; Skalse and Abate 2023). Vector rewards and Pareto-optimal policy sets are common in multi-objective decision-making (Roijers et al. 2013; Hayes et al. 2022b), where general policy optimization recovers the Pareto front by conditioning a single policy-network on a sampled weight vector. This literature, however, largely assumes the ground-truth vector reward function is given. The reward-free viewpoint of (Chen et al. 2026) relaxes part of this assumption by using preference-guided exploration as an auxiliary task, yet still requires the ground-truth multiobjective reward during training. In contrast, LEMUR avoids these assumptions. Instead, it posits that true objectives are never directly observed and must be learned, in this case from preference feedback. Thus, the agent jointly learns reward models and an optimized policy that balances its various components. Multi-Dimensional and Fine-Grained Preference Learning. Standard RLHF distills human comparisons into a single scalar reward (Christiano et al. 2017; Ouyang et al. 2022), which can conflate distinct criteria and collapse the multiobjective structure of human values (Sorensen et al. 2024). A line of work therefore decomposes feedback into finergrained components: Fine-Grained RLHF (Wu et al. 2023) attaches rewards to localized segments and categories, while Rewarded Soups (Rame et al. 2023) and MODPO (Zhou et al. 2024) obtain a separate signal per objective and expose a Pareto family through weight interpolation or scalarized preference optimization. These methods combine per-objective signals post hoc rather than jointly learning the rewards and a trade-off policy, and are developed almost exclusively for LLM post-training. A related direction forgoes the reward model entirely, optimizing the policy directly from preference data (e.g., DPO (Rafailov et al. 2023)); such methods are effective but train on static or periodically refreshed preferences, forgoing the online, interactive feedback and policy optimization, ideal for reliable reward learning under policy improvement (Christiano et al. 2017; Lee, Smith, and Abbeel 2021; Gao, Schulman, and Hilton 2023; Bai et al. 2022; Ross, Gordon, and Bagnell 2011; Ibarz et al. 2018; Stiennon et al. 2020). A further perspective treats heterogeneous feedback as hidden context, learning a distribution over reward functions (Siththaranjan, Laidlaw, and Hadfield-Menell 2024); this captures which preferences are present but does not, on its own, optimize an explicit compromise between them. LEMUR differs from all three: it learns an explicit, objectivespecific reward model per feedback stream online, and jointly optimizes a set of policies over the resulting vector reward
in continuous control. A complementary line moves beyond the Bradley-Terry preference model assumption itself. General preference models represent preferences with a richer (e.g., skew-symmetric, k-dimensional) structure that admits intransitive cycles, optimizing policies directly against this structure rather than a scalar reward (Umer et al. 2026). Related game-theoretic formulations cast alignment as a Nash or Stackelberg equilibrium over preferences (Munos et al. 2024; Pásztor, Buening, and Krause 2025). Intransitivity is a different failure mode of scalar rewards from the one we study; it concerns the shape of preferences rather than the presence of multiple objectives, and we regard bridging it with multi-objective preference optimization as promising future work. In the LLM setting, rubric-based rewards have similarly been used to supply multi-dimensional supervision for reinforcement fine-tuning, though their reliability is sensitive to rubric coverage and correlated criteria (Gunjal et al. 2025; Shen et al. 2026). These directions are largely orthogonal to LEMUR, which targets online vector-reward learning and trade-off policy optimization. Connections to Diversity and Foundation-Model PostTraining. A contemporary line of work observes that scalar RL post-training induces entropy collapse, eroding the solution diversity required by inference-time search (Kirk et al. 2024). Most directly, Vector Policy Optimization (BahlousBoldi et al. 2026) samples scalarizations over the reward simplex and trains a policy to output a set of solutions spanning the Pareto front, improving downstream best-of-N search. Such methods share LEMUR’s premise that collapsing a vector reward into a scalar discards useful structure. They differ from our setting in three respects: (i) they target LLM test-time search, seeking a diverse candidate pool for a downstream selector, whereas LEMUR seeks a policy that compromises between objectives during deployment; (ii) they assume the reward components are known and observable (e.g., per-test-case correctness), whereas LEMUR must learn them from feedback; and (iii) they operate on natural language generation rather than continuous control. We therefore treat this literature as motivating context rather than comparable works, and leave the study of train-time multi-objective learning from feedback to improve test-time diversity (Ding et al. 2024; Pierrot et al. 2022) as future work. Pluralistic Alignment and Multiple Principals. LEMUR’s multi-teacher formulation connects to pluralistic alignment, which holds that a single reward model cannot represent the plurality of human values (Sorensen et al. 2024) and that an agent should instead balance multiple objectives to reach a compromise across them. A closely related framing casts this as serving multiple principals: a single agent acting on behalf of several stakeholders must respect the preferences of each, turning alignment into the problem of representing and trading off competing principals rather than satisfying one (Fickinger et al. 2020). How to reconcile these principals is itself contested. Naively aggregating preferences via RLHF has been shown to behave as a Borda count over latent objectives, with limited normative justification (Siththaranjan, Laidlaw, and Hadfield-Menell 2024), and others question whether aggregation is the right
frame at all, proposing non-aggregative alternatives for reconciling plural values (Zhi-Xuan et al. 2025). Practical efforts that do aggregate rely on curated consensus datasets or hand-specified aggregation functions (Bakker et al. 2022; Rodriguez-Soto et al. 2023; Tessler et al. 2024), both costly to obtain a priori and brittle when preferences shift during deployment. Rather than collapsing principals to a consensus in advance or assuming their reward functions are known, LEMUR preserves each principal’s objective as a separate learned reward model and recovers an explicit compromise through multi-objective optimization.
▷ Pre-training
4: for each iteration do 5: if iteration mod K == 0 then ▷ Reward learning 6: for each teacher j ∈ {1, . . . , m} do 7: Sample queries (σ 0 , σ 1 ) ∼ B 8: Dj ← Dj ∪ {(σ 0 , σ 1 , yj )}M n=1 9: Update r̂ψj on Dj via cross-entropy loss 10: end for 11: end if
Collect transitions (st , at , st+1 ) with πϕ and store in B
for each gradient step do ▷ Multi-objective policy optimization 14: Sample batch (s, a, s′ ) ∼ B 15: Relabel vector rewards: rw = w⊤ r̂ψ (s, a) 16: Update actor and critic using MO-SAC 17: end for 18: end for 13:
Episode Returns (Objective One) Episode Returns (Objective Two)
2000 1000
Additional Experiments
Adding a Teacher Mid-Training
A practical deployment of preference-based MORL is unlikely to have a fixed, known set of stakeholders at the outset: new preference sources appear over time. A framework that requires retraining from scratch whenever a teacher joins is therefore of limited practical use. Because LEMUR maintains one weight-conditioned reward model per teacher and couples them only through the shared MORL/D population, adding a teacher requires instantiating a single new reward model and extending the vector reward, leaving the existing reward models and the trained policy population intact. We test this directly: a run begins with the standard twoteacher MO-Cheetah setup and a third teacher is introduced
3000 2000 1000 0
0
1
2
3
4
5
Training Steps (×10 )
6
7
8
Figure 6: Adding a third teacher mid-training (MO-Cheetah). A third teacher is introduced at 4 × 105 steps (dotted line) into an already-training two-teacher run. The new objective (bottom) is learned from scratch while the two existing objectives (top, middle) are preserved, so no retraining from scratch is required. Mean ± std over five seeds.
at 4 × 105 environment steps, with training continuing uninterrupted from the existing population. Figure 6 shows the outcome. The newly added Objective Three is optimized from a cold start and rises steeply once its teacher joins, while Objectives One and Two, already near convergence at the changepoint, are retained rather than degraded, settling into a marginally adjusted equilibrium that accommodates the new objective. Crucially, the framework absorbs the new preference source online, without reinitialising either the reward models or the policy population, demonstrating that the cost of adding a teacher is incremental rather than a full retraining cycle.
C.2
C
2000
0
Require: Feedback frequency K, queries per session M , number of objectives m 1: Initialize policy ϕ and reward models {ψj }m j=1 2: Initialize shared buffer B ← ∅ and preference datasets Dj ← ∅
C.1
3000
3000
Algorithm 1: LEMUR
12:
4000
0
LEMUR Pseudocode
3: B ← Explore(πϕ , r int )
5000
1000
Episode Returns (Objective Three)
B
LEMUR (2 3 Teachers) Teacher C Added
6000
Non-Stationary Preferences
Human preferences are not static, and a teacher may revise its trade-off during the course of training. Since LEMUR re-queries every teacher and relabels the shared replay buffer at each feedback interval, a revised preference propagates into the learned reward and hence into the policy objective without any special-case handling. To test this, Teacher A’s weight vector is altered at 2 × 105 steps mid-training. Figure 7 shows that all three objectives continue to improve monotonically across the changepoint: the policy adapts to the revised utility rather than collapsing or plateauing, and the widening variance band immediately after the change reflects the transient period during which the reward models are being re-fit to the new preference before the population re-converges.
LEMUR (Non-Stationary Preferences) Preference Change
4000 3000 2000
6000
1000
5000
Episode Returns (Objective One)
Episode Returns (Objective One)
5000
0
Episode Returns (Objective Three)
6000 5000 4000
2000 1000 0 0.0
3000
0.5
1.0
1.5
2000 1000 0
0
1
2
3
Training Steps (×10 )
4
5
Figure 7: Non-stationary preferences (MO-Cheetah). Teacher A’s preference vector is changed at 2×105 steps (dotted line). Returns continue to improve across the changepoint, indicating that the online re-query and buffer-relabeling loop tracks the revised utility. Mean ± std over five seeds.
2.0
2.5
Training Steps (×10 )
3.0
3.5
4.0
4.5
Figure 8: Effect of query segment length (MO-Cheetah). Ground-truth return per objective for preference queries of length 50 (default), 35 and 1 transition, at a fixed budget of 300 queries per teacher. Shading is ±1 std; the length-35 and length-1 arms are single seeds and carry a nominal band. Shorter segments learn faster early but plateau by ∼2 × 105 steps, while length 50 overtakes them and continues improving.
Ablation: Query Length
We also tested if which segments are queried matters at a fixed budget and length. We compared LEMUR’s default uniform sampler against an entropy-based sampler that selects the most uncertain segment pairs from a 10× candidate pool (Figure 9). Entropy-based selection yields a modest gain in performance. Uniform sampling recovers most of the achievable return without requiring uncertainty estimates, candidate pools, or extra forward passes. Combined with our querylength findings, this shows that under a fixed budget, how much behavior each query covers matters significantly more than which specific segments are chosen.
Overlapping Preferences
Our main results consider teachers whose anchors are genuinely conflicting. A natural question is whether the machinery required to resolve conflict imposes a cost when the teachers largely agree. We therefore repeat the experiment with overlapping (near-aligned) teacher weight vectors, comparing LEMUR against the ground-truth oracle upper bound
Episode Returns (Objective One)
Ablation: Query Sampling Strategy
6000
Standard Uniform Sampling Entropy-Based Sampling
5000 4000 3000 2000 1000 0 4000
Episode Returns (Objective Two)
LEMUR elicits preferences over trajectory segments, determining the behavioral context teachers see per comparison. At a fixed budget of 300 queries per teacher, we tested segment lengths of 50, 35, and 1 transition (Figure 8). LEMUR maintains performance even with shorter segments of feedback. The improvement in performance due to longer segments is potentially due to more context for reward learning (Lee et al. 2021; Lee, Smith, and Abbeel 2021).
C.5
2000
3000
7000
C.4
3000
0
1000 0
C.3
4000
1000
2000
Episode Returns (Objective Two)
Episode Returns (Objective Two)
3000
Query Length 50 Query Length 35 Query Length 1
3000 2000 1000 0 0.0
0.5
1.0
1.5
2.0
2.5
Training Steps (×10 )
3.0
3.5
4.0
4.5
Figure 9: Uniform vs. entropy-based query sampling (MO-Cheetah). Ground-truth return per objective for the default uniform sampler and an entropy-based sampler that scores a 10× candidate pool by reward-model uncertainty. Shading is ±1 std; the entropy arm is a single seed and carries a nominal band. Entropy-based selection is consistently ahead, but by a small margin relative to the effect of query length (Figure 8).
6000
Episode Returns (Objective One)
D
LEMUR (Overlapping) Ground Truth Oracle (Overlapping)
7000 5000 4000 3000 2000 1000 0
Episode Returns (Objective Two)
6000 5000 4000 3000 2000 1000 0
0
1
2
3
4
Training Steps (×10 )
5
6
7
Figure 10: Overlapping preferences (MO-Cheetah). With near-aligned teacher anchors, LEMUR converges close to the ground-truth oracle on both objectives, showing that the method degrades gracefully when teachers largely agree. Mean ± std over five seeds.
under the same overlap condition. Figure 10 shows that LEMUR tracks the oracle closely on both objectives, converging to a comparable final return with only a modest lag in sample efficiency. This indicates that the weight-conditioned formulation degrades gracefully toward the single-preference case, when there is little conflict to resolve.
C.6
Ablation: Varying Levels of Teacher Conflict
This ablation isolates the degree of teacher disagreement. Holding MO-Cheetah, the MORL/D backbone, explorer, and query budget fixed, we sweep the anchors from mildly to severely opposed and train LEMUR alongside an Oracle given the same anchors at each level. Comparing against a per-level Oracle separates degradation caused by the trade-off becoming harder from degradation caused by reward learning failing under disagreement. LEMUR remains competitive across the sweep (Figure 11).
C.7
Ablation: Weight-Conditioned Reward Model vs. Reward Ensemble
LEMUR’s central architectural choice is to condition each teacher’s reward model on a preference weight vector, rather than learning an ensemble of unconditioned reward models as in the earlier formulation. To isolate the contribution of this choice we hold everything else fixed, the same MORL/D backbone, explorer, teacher anchors, and query budget, and vary only the reward model. Figure 12 shows a substantial and consistent gap on both objectives: the weight-conditioned model converges to 6,812 ± 39 and 4,404 ± 22, against 4,556 ± 369 and 2,902 ± 245 for the ensemble. The ensemble also exhibits an order-of-magnitude larger standard error, indicating greater seed-to-seed instability.
LEMUR Implementation Details
Training Details. We utilize the MORL/D algorithm (Felten, Talbi, and Danoy 2024), a Multi-Objective Soft Actor-Critic (MO-SAC) implementation from the MORL-Generalization benchmark (Teoh, Varakantham, and Vamplew 2025). A summary of the hyperparameters is provided in Tables 3 and 4. In the initial Exploration phase (Stage 0), an intrinsicmotivation explorer bootstraps a replay buffer of environment transitions, run for 50,000 timesteps on LunarLander, 80,000 on Hopper, and 100,000 on Cheetah and MetaWorld. This buffer is shared identically with every baseline, so no method receives more exploration data than another. In the Reward Learning phase (Stage 1), each of the m teachers trains its own weight-conditioned reward model: a 2-layer MLP with 256 hidden units that takes the stateaction pair concatenated with a preference weight vector w, trained by Bradley–Terry cross-entropy over pairwise trajectory-segment comparisons of length H = 50. Rather than querying each teacher only at its own fixed anchor wj , weights are sampled from a Dirichlet distribution centred on that anchor with concentration κ = 30, so a single model generalises across a neighbourhood of the anchor instead of memorising one point on the simplex. At each iteration of reward-learning and policy-optimization, the per-teacher query budgets are M = 200 (LunarLander), 500 (Hopper), and 300 (Cheetah, MetaWorld), each trained for 100 epochs with Adam optimizer. For Multi-Objective Policy Optimization (Stages 2-3), a population of 6 MO-SAC policies is trained on the mdimensional learned vector reward, coupled by a shared replay buffer, weighted-sum scalarization, PSA weight adaptation, and weight transfer between neighbouring policies (neighbourhood size 2). Learning rates and exchange frequencies are adapted per environment (Table 4) while the population size is held constant at 6 across all environments. Crucially, reward learning does not terminate after Stage 1: every feedback_interval steps the pipeline re-queries all teachers, performs 30 online reward-update epochs, and relabels the shared replay buffer with the updated reward models, so the policy and the reward model co-adapt throughout training.
E
Baselines Implementation Details
To ensure a fair comparison, all baselines share LEMUR’s Stage-0 explorer, scripted teacher weight vectors, query budget M , query length, interaction frequency K, reward-model capacity, and total environment-step budget, and are evaluated and logged under identical protocols (Tables 3 and 4). The primary distinction between methods therefore lies in how conflicting reward signals are aggregated and optimized, not in the data or compute they receive. Each baseline retains its own defining reward-learning mechanism: MORAL’s adversarial AIRL rewards, PbMORL’s weight-conditioned Bradley–Terry model, and FPbRL’s GGF welfare model. Where a method’s original policy optimizer would place it at an unfair disadvantage in our environments, we adapt in the baseline’s favour, upgrading
Medium Conflict Level (w=0.6/0.4) Episode Returns (Objective One) Episode Returns (Objective Two)
Hard Conflict Level (w=0.75/0.25)
Medium Conflict Level (w=0.6/0.4) Medium Conflict Level Oracle
6000 5000
6000 5000
Harder Conflict Level (w=0.9/0.1)
Hard Conflict Level (w=0.75/0.25) Hard Conflict Level Oracle
6000 5000
4000
4000
4000
3000
3000
3000
2000
2000
2000
1000
1000
1000
0
0
0
4000
4000
4000
3000
3000
3000
2000
2000
2000
1000
1000
1000
0 0.0
0.5
1.0
1.5
2.0
2.5
3.0
Training Steps (×10 )
3.5
4.0
4.5
0 0.0
0.5
1.0
1.5
2.0
2.5
3.0
Training Steps (×10 )
3.5
4.0
4.5
0 0.0
Harder Conflict Level (w=0.9/0.1) Harder Conflict Level Oracle
0.5
1.0
1.5
2.0
2.5
3.0
Training Steps (×10 )
3.5
4.0
4.5
Figure 11: Varying levels of teacher conflict (MO-Cheetah). Each column is one conflict level, set by the teacher preference anchors: medium (w = [0.6, 0.4]/[0.4, 0.6]), hard ([0.75, 0.25]/[0.25, 0.75]) and harder ([0.9, 0.1]/[0.1, 0.9]). Rows give each objective’s ground-truth return. Solid lines are LEMUR, dashed the ground-truth-reward Oracle under the identical configuration; shading is ±1 std across seeds, with a common y-scale per row.
Hyperparameter
Value
Hyperparameter
Value
Explorer Timesteps Explorer Policy / Q LR
Exploration (Stage 0) 50,000 (LunarLander), 80,000 (Hopper), 100,000 (Cheetah, MetaWorld) 3 × 10−4
Explorer Batch Size Optimizer
128 Adam
Reward Model Reward Learning Rate Queries per Teacher (M ) Reward Epochs Dirichlet Concentration κ
Reward Learning (Stage 1) Weight-conditioned MLP (2-layer) 5e-4 (LL), 3e-4 (Hopper), 2.5e-4 (Cheetah), 1.5e-4 (MetaWorld) 200 (LL), 500 (Hopper), 300 (Cheetah, MetaWorld) 100 (All Environments) 30.0 (All Environments)
Hidden Dim Reward Batch Size Query Length (H) Loss Optimizer
256 64 50 Bradley–Terry CE Adam
Online Update Epochs
30
Buffer Relabeling
True
Online Reward Updates (Stages 2–3)
Table 3: Hyperparameters for Exploration (Stage 0) and Reward Learning (Stage 1).
Hyperparameter
Value
Hyperparameter
Value
Algorithm MORL Timesteps Policy LR Q LR Population Size Exchange Frequency Weight Adaptation Scalarization
MORL/D (MO-SAC) 100,000 (LL), 750,000 (Hopper, Cheetah), 500,000 (MetaWorld) 5e-5 (LL), 3e-4 (Hopper, Cheetah), 1e-4 (MetaWorld) 5e-5 (LL), 1e-4 (Hopper, MetaWorld), 3e-4 (Cheetah) 6 (All Environments) 5,000 (LL), 10,000 (Hopper, Cheetah, MetaWorld) PSA Weighted Sum (ws)
Batch Size Target Entropy Scale Optimizer Discount Factor γ Shared Buffer Weight Transfer Neighborhood Size
512 (LL, MetaWorld), 1024 (Hopper, Cheetah) 0.3 Adam 0.99 True True 2
Table 4: Hyperparameters for Multi-Objective Policy Optimization (Stages 2–3).
Hyperparameter
Value
Hyperparameter
Value
Reward Model Query Budget (M ) Reward LR / Epochs Optimizer (policy)
PbMORL (Mu, Luan, and Jia 2025) Weight-conditioned MLP Hidden Dim Query Length (H) 300 (total, split across teachers) 2.5e-4 / 100 Reward Batch Size Envelope (LunarLander), MORL/D (Others) Population Size
256 50 64 6
Reward Model Query Budget (M ) / Length Policy Algorithm PPO Steps / Iter PPO Update Epochs GAE λ
FPbRL (Siddique, Sinha, and Cao 2023) GGF welfare MLP (K-dim) Hidden Dim 300 / 50 Reward LR / Epochs PPO (continuous), SAC-discrete (LunarLander) Policy LR 2,048 PPO Minibatches 10 PPO Clip Coef 0.95 GGF Weights
256 2.5e-4 / 100 3e-4 32 0.2 Decreasing on sorted utilities
Table 5: Hyperparameters for the PbMORL and FPbRL baselines.
Hyperparameter
Value
Hyperparameter
Value
Stage 0 & 1: Expert Collection & AIRL Reward Inference Expert Timesteps 1 × 106 Demo Episodes Expert Net Arch. 256, 256 Expert Batch Size Expert Policy / Q LR 3 × 10−4 Expert Buffer Size Expert α / τ 0.2 / 0.005 Expert Learning Starts AIRL Generator Timesteps AIRL Hidden Dim 32 AIRL Disc. Update Interval 1,024 AIRL Gen. Net Arch. Optimizer AIRL Gen. Buffer / Batch 300,000 / 256
50 256 1 × 106 10,000 200,000 256, 256 Adam
Table 6: MORAL (Peschl et al. 2022): hyperparameters for expert demonstration collection and adversarial AIRL reward training.
Hyperparameter
Value
Hyperparameter
Policy Algorithm Policy / Q LR Batch Size Discount Factor γ
Stage 2: Online Policy Optimization with Active Learning SAC (all environments; see App. E) Active Query Budget 3 × 10−4 Query Interval 256 Scalarization Posterior 0.99 Query Strategy
Value
5,000 200 Bradley-Terry Volume removal
Table 7: MORAL (Peschl et al. 2022): hyperparameters for online policy optimization (Stage 2). Note we optimize with SAC rather than the original work’s PPO; see Appendix E.
Weight-Conditioned (LEMUR) Reward Ensemble
Episode Returns (Objective One)
6000 5000 4000 3000 2000 1000 0
Episode Returns (Objective Two)
4000 3000 2000 1000 0
0
1
2
3
4
Training Steps (×10 )
5
6
7
Figure 12: Reward-model ablation (MO-Cheetah). The weight-conditioned reward model substantially outperforms the reward-ensemble variant on both objectives under an identical optimizer, explorer, and query budget. Mean ± std over five seeds.
the optimizer rather than reproducing the paper verbatim, so that reported gaps reflect differences in reward learning and preference aggregation, the object of study, rather than differences in policy optimization. Most notably, MORAL is optimized with SAC rather than the PPO used in the original work, substantially improving its sample efficiency on our continuous-control tasks. All such deviations are disclosed below. Utilitarian Agent. This baseline imposes an a priori scalarization on the objectives. Like LEMUR, it learns m distinct reward models {r̂ψj }m j=1 , one per teacher, but rather than learning a set of trade-off policies it trains a single SAC agent to optimize their arithmetic mean: m
1 X rutil (s, a) = r̂ψ (s, a). m j=1 j
(3)
Learning rates, batch sizes, and buffer sizes are identical to LEMUR, with the population size set to 1 (standard singleobjective SAC). Naive Data Pooling. This baseline aggregates conflicting preferences at the data level, mimicking standard RLHF applied to heterogeneous feedback. Rather than maintaining separate datasets Dj per objective, all feedback Sm tuples are stored in a single monolithic dataset Dpool = j=1 Dj , over which one reward model r̂pool is trained to minimize the cross-entropy loss. The policy is then trained with standard SAC to maximize rnaive (s, a) = r̂pool (s, a).
(4)
As with the Utilitarian baseline, we use the environmentspecific hyperparameters in Table 4 with the population size fixed to 1.
PbMORL (Mu, Luan, and Jia 2025). The original PbMORL assumes a single, internally-consistent teacher whose preferences are valid under any sampled scalarization weight, and therefore has no mechanism for multiple, independentlyopinionated teachers. We retain its weight-conditioned mdimensional reward model, trained with the Bradley–Terry cross-entropy objective over pairwise trajectory comparisons, and extend it to our setting in the two most natural ways: PbMORL-naive pools all teachers’ preferences into one shared weight-conditioned model, while PbMORLutilitarian trains a separate model per teacher and combines them at inference by a fixed uniform average. Each teacher answers every query under its own fixed weight vector, exactly as in LEMUR. The original paper pairs its reward model with Envelope Q-learning (Yang, Sun, and Narasimhan 2019), a discrete, value-based method. We retain this paper-faithful pairing on LunarLander, but Envelope’s Q-network requires an argmax over actions and is structurally inapplicable to continuous control; on Hopper, HalfCheetah, and MetaWorld we therefore substitute MORL/D (Felten, Talbi, and Danoy 2024), the same population-based optimizer LEMUR uses. This is the only structural substitution, and it equalizes the optimizer between PbMORL and LEMUR on those environments, isolating the reward-learning strategy as the sole difference. FPbRL (Siddique, Sinha, and Cao 2023). We implement FPbRL’s K-dimensional welfare reward model, trained from Generalized Gini Welfare (GGF)-based preferences over the same K = 2 scripted teachers across which LEMUR and the other baselines compromise. At each policy-update iteration, the learned K-dimensional reward is scalarized by the GGF weight assignment, which places the largest weight on the currently worst-off objective, refreshed from a running per-objective return estimate, the standard practical form of Siddique et al.’s fair policy gradient. Unlike MORAL, we retain PPO for continuous-action environments, matching the original paper: FPbRL’s fair policy gradient is formulated for an on-policy optimizer, and substituting an off-policy method would alter the method’s semantics rather than simply strengthen it. On LunarLander we use SAC-discrete, as no discrete PPO implementation is used elsewhere in our pipeline. In addition to the per-teacher returns reported for all methods, we log FPbRL’s own fairness metrics (welfare, coefficient of variation, and minimum objective), so that it is also evaluated on the criterion it is explicitly designed to optimize. MORAL (Peschl et al. 2022). We reimplement MORAL’s pipeline in three stages: (i) two expert policies are trained on the ground-truth per-teacher scalarized reward and used to collect demonstrations; (ii) for each teacher, an AIRL (Fu, Luo, and Levine 2018) discriminator is trained adversarially against a live generator, rather than against a fixed pool of shuffled expert transitions, which would render training non-adversarial, yielding a per-teacher learned reward g(s); and (iii) an active-learning wrapper scalarizes the two frozen AIRL rewards with a Bradley-Terry weight posterior, updated online via volume-removal preference queries, which a single-objective policy optimizes. Since the original
E.1
MORAL Baseline: Learned Weights and Optimiser Choice
Learned preference weight
MORAL was evaluated only on grid-world tasks, we adopt the AIRL hyperparameters and network architectures of Fu et al. (Fu, Luo, and Levine 2018) for our high-dimensional MuJoCo experiments; these are detailed in Tables 6 and 7. We deliberately preserve MORAL’s defining contribution, adversarial IRL reward learning combined with an activelyqueried scalarization posterior, and do not replace it with our own preference-based reward model, as doing so would no longer constitute a MORAL baseline. We do, however, strengthen its policy optimization: the original work uses PPO, whereas we optimize with SAC (SAC-continuous, or SAC-discrete on LunarLander) across all environments. This off-policy upgrade materially improves MORAL’s sample efficiency under an identical environment-step budget and matches the backbone family used by LEMUR, ensuring MORAL is not penalized for an on-policy optimizer choice unrelated to its reward-learning contribution.
0.600 0.575 0.550 0.525 0.500 0.475 0.450 0.425 0.400
Objective One Weight Objective Two Weight
0
1
2
3
4
Training Steps (×10 )
5
6
7
Figure 13: MORAL’s learned scalarisation weights in (MO-Cheetah). The two components of MORAL’s activequery scalarisation posterior over training.
Learned scalarisation weights. MORAL maintains a posterior over a scalarisation weight vector, refined through active queries, and it is this posterior, rather than a per-teacher reward decomposition, that carries its notion of whose preferences the policy is serving. Figure 13 tracks both components over training.
F F.1
Additional Results
Multi-Objective RL Metrics
Evaluating a multi-objective agent requires assessing the set of policies it recovers rather than any single return, so we adopt two metrics standard in the MORL literature (Hayes et al. 2022a; Roijers et al. 2013). Let P = {J(π1 ), . . . , J(πP )} denote the set of objective-value vectors attained by the learned policy population. Hypervolume (HV). Given a reference point z dominated by all solutions, Hypervolume is the volume of the region
MORAL (SAC) MORAL (PPO)
Episode Returns (Objective One)
100 200 300 400 100
Episode Returns (Objective Two)
SAC vs. PPO. As described in Appendix E, we optimize the MORAL baseline with SAC rather than the PPO used in the original work, on the grounds that an on-policy optimizer would disadvantage the baseline for reasons unrelated to its reward-learning contribution. This ablation verifies that the substitution is genuinely favourable to MORAL and therefore that our reported comparison is conservative. Figure 14 confirms this: holding MORAL’s adversarial AIRL reward learning and active-query scalarisation posterior fixed and varying only the policy optimizer, the SAC variant dominates PPO on both objectives throughout training, and the gap widens as training proceeds. The PPO variant additionally displays a pronounced sawtooth characteristic of on-policy updates (shown here under heavy smoothing). Reporting MORAL with SAC therefore strengthens the baseline relative to a faithful reproduction of the original paper, and the weight collapse in Figure 13 is a property of MORAL’s scalarization posterior rather than a symptom of the optimizer.
200 300 400 500
0
1
2
3
4
Training Steps (×10 )
5
6
7
Figure 14: MORAL optimizer ablation (MO-Cheetah). With MORAL’s reward learning held fixed, SAC outperforms the original work’s PPO on both objectives, confirming that our SAC upgrade strengthens the baseline. Curves are heavily smoothed to expose the trend beneath PPO’s on-policy oscillation. Mean ± std over five seeds.
dominated by P and bounded by z, [ HV(P, z) = Λ [z, p] ,
(5)
p∈P
where Λ is the Lebesgue measure. It simultaneously rewards solutions that are high-performing (pushing the front outward) and diverse (covering more of the objective space), and is the most widely adopted MORL quality indicator because it is the only common metric strictly monotonic with Pareto dominance: any set that dominates another is guaranteed a higher score (Roijers et al. 2013). We use the reference point supplied by the MORL-Generalization benchmark (Teoh, Varakantham, and Vamplew 2025) so that values are comparable across methods within an environment; absolute magnitudes are not comparable across environments, since they depend on both the reference point and the reward scale. Sparsity (SPS). Sparsity measures how evenly solutions are distributed along the recovered front. Sorting the |P| solutions by each objective i and writing P̃i (k) for the k-th value, SPS(P) =
m |P|−1 2 X X 1 P̃i (k) − P̃i (k + 1) , (6) |P| − 1 i=1 k=1
with lower values indicating more uniform coverage (Xu et al. 2020). Sparsity must be read alongside Hypervolume rather than independently: a degenerate front that collapses to a single solution reports a low, apparently favourable value despite failing to cover the objective space, which is why we report both. This is the case for FPbRL, whose fixed welfare scalarization converges to a single policy, so no front is recovered and sparsity is undefined (−) on three of four environments. Environment
LEMUR
PbMORL
FPbRL
Hypervolume (↑) LunarLander 1.10 × 104 ± 2.46 × 102 1.09 × 104 ± 2.87 × 102 5.05 × 103 ± 5.75 × 102 Hopper 3.67 × 106 ± 2.01 × 105 2.24 × 106 ± 0.000 7.19 × 104 ± 2.04 × 104 HalfCheetah 4.86 × 107 ± 9.45 × 105 4.78 × 107 ± 0.000 7.73 × 105 ± 2.91 × 105 6 5 6 5 MetaWorld-DrawerClose 2.15 × 10 ± 6.49 × 10 1.43 × 10 ± 4.43 × 10 2.79 × 106 ± 3.39 × 105 Sparsity (↓) LunarLander Hopper HalfCheetah MetaWorld-DrawerClose
134.5 294.7 294.7 436.3
7.8 1731.7 2191.0 5564.7
− − − 0.000
Table 8: Hypervolume and Sparsity across benchmark environments. FPbRL’s fixed welfare scalarization fails to recover a set of trade-off policies, converging instead to a single solution: sparsity is therefore undefined (−) where no front exists, and its near-zero value on MetaWorld-DrawerClose reflects this collapse rather than uniform coverage. PbMORL is deterministic under our protocol on Hopper and HalfCheetah, hence zero variance.
ground-truth teacher utilities, following PbRL benchmarking practice (Lee et al. 2021). Let r̂ψj denote teacher j’s learned reward model and rj (s, a) = wj⊤ r(s, a) its ground-truth utility, where r is the environment’s native vector reward and wj teacher j’s scripted weight vector. All metrics lie in [−1, 1], higher is better, and are averaged across the m teachers. Per-state correlation. Over states sampled from evaluation rollouts, Spearman rank correlation measures how faithfully the learned reward orders individual transitions. It is our primary metric because preference-based rewards are identifiable only up to a positive monotone transform, making an order-preserving measure the appropriate notion of correctness. We also report Pearson correlation on raw values, which is stricter in penalising any nonlinear distortion; following prior reward-evaluation work we refer to this as the Value-Order Correlation (VOC) (Ma et al. 2025). Trajectory and policy ranking. Per-state correlation can be high even when a reward model induces the wrong ordering over whole behaviours, which is what the policy ultimately optimizes. We therefore roll out each policy in the MORL/D population and compare the induced rankings using Kendall’s τ -b over trajectory returns, and the Trajectory Alignment Coefficient (TAC) of (Muslimani et al. 2025), computed identically but on discounted returns so that alignment is weighted by the same temporal discounting the agent optimizes. Trajectory VOC additionally averages the withintrajectory correlation between learned and ground-truth perstep reward sequences, capturing whether the model tracks the shape of the signal within an episode rather than only across episodes. Results. Tables 9-11 report all metrics. LEMUR attains the strongest per-state correlations on every environment, with the largest margins on Hopper and MetaWorld-DrawerClose where the teachers’ anchors are most opposed, consistent with per-teacher decomposition mattering most under genuine conflict. Trajectory-level metrics are more mixed: FPbRL attains a higher Kendall τ and TAC on Hopper despite substantially weaker per-state correlation, reflecting that its welfare scalarization orders whole behaviours consistently even where the underlying reward is poorly calibrated. Reporting both families is what makes this distinction visible, and we recommend the same practice for future work in this setting. Metric Spearman (ρ ↑) Pearson (r ↑) VOC (↑) VOC (traj) (↑) Kendall τ (policy) (↑) Traj-Alignment Coeff (↑)
LEMUR
PbMORL
FPbRL
0.710 ± 0.005 0.863 ± 0.003 0.863 ± 0.003 0.858 ± 0.003 0.939 ± 0.030 0.898 ± 0.041
0.705 ± 0.000 0.853 ± 0.000 0.853 ± 0.000 0.851 ± 0.000 0.933 ± 0.000 0.858 ± 0.000
0.512 ± 0.089 0.493 ± 0.065 0.493 ± 0.065 0.494 ± 0.066 0.567 ± 0.233 0.633 ± 0.233
Table 9: Reward-Model Evaluation Metrics : HalfCheetah.
F.2
Reward Model Evaluation Metrics
Policy return alone cannot distinguish a reward model that has genuinely recovered a teacher’s utility from one merely correlated with it on the visited state distribution. We therefore evaluate the learned reward models directly against the
G
Benchmark Environment Details
Table 12 summarises the native multi-objective structure of each benchmark environment and the scripted
Metric Spearman (ρ ↑) Pearson (r ↑) VOC (↑) VOC (traj) (↑) Kendall τ (policy) (↑) Traj-Alignment Coeff (↑)
LEMUR
PbMORL
FPbRL
0.520 ± 0.007 0.589 ± 0.020 0.589 ± 0.020 0.340 ± 0.035 0.359 ± 0.002 0.347 ± 0.026
0.231 ± 0.000 0.335 ± 0.000 0.335 ± 0.000 0.339 ± 0.000 0.291 ± 0.000 0.300 ± 0.000
0.085 ± 0.074 0.345 ± 0.005 0.345 ± 0.005 0.358 ± 0.010 −0.167 ± 0.100 0.167 ± 0.167
Table 10: Reward-Model Evaluation Metrics: MetaWorldDrawerClose. Metric Spearman (ρ ↑) Pearson (r ↑) VOC (↑) VOC (traj) (↑) Kendall τ (policy) (↑) Traj-Alignment Coeff (↑)
LEMUR
PbMORL
FPbRL
0.945 ± 0.002 0.951 ± 0.002 0.951 ± 0.002 0.947 ± 0.004 0.914 ± 0.002 0.856 ± 0.014
0.527 ± 0.000 0.532 ± 0.000 0.532 ± 0.000 0.606 ± 0.000 0.291 ± 0.000 0.370 ± 0.000
0.648 ± 0.094 0.716 ± 0.073 0.716 ± 0.073 0.714 ± 0.073 0.933 ± 0.000 0.933 ± 0.000
Table 11: Reward-Model Evaluation Metrics: Hopper.
teacher weight vectors used throughout our experiments. In every case, teacher j’s ground-truth utility is the linear scalarisation rj (s, a) = wj⊤ r(s, a) of the environment’s native vector reward r, and it is this quantity that validation/gt_teacher_{a,b}_return tracks during training. The weight vectors are chosen to be genuinely conflicting: each teacher places its largest weight on a different objective, so no single policy can simultaneously maximise both utilities. MO-LunarLander. A four-objective variant of the classic LunarLander domain with a discrete action space, taken from the MORL-Generalization benchmark (Teoh, Varakantham, and Vamplew 2025). The native reward vector comprises the shaping term (progress toward the landing pad), the mainengine fuel cost, the side-engine fuel cost, and the terminal landing/crash outcome. Teacher A weights the two enginecost terms asymmetrically against Teacher B, producing a fuel-allocation conflict on top of a shared landing objective. MO-Hopper. A three-objective continuous-control locomotion task in which the native reward vector is [ vx , h, −c ∥a∥2 ]: forward velocity, hop height, and a negated energy/control cost. Teacher A ([0.8, 0.1, 0.1]) strongly prefers fast locomotion, while Teacher B ([0.3, 0.5, 0.2]) prefers a higher, more energy-efficient and stable gait. MO-Cheetah. A two-objective continuouscontrol task whose native reward vector is [ reward_forward, reward_ctrl ], i.e. forward velocity against control cost. The teacher anchors [0.6, 0.4] and [0.4, 0.6] place opposing emphasis on velocity versus energy efficiency. MO-MetaWorld (Drawer-Close). Meta-World tasks are natively single-objective, providing only a dense taskprogress reward. We convert Drawer-Close into a twoobjective task by pairing this native reward with a controleffort penalty, yielding the vector reward [ rtask , −λ∥a∥22 ] with control-cost weight λ = 0.1 and a fixed horizon of 500
steps. This mirrors the forward-reward/control-cost decomposition standard to the MuJoCo suite, but applied to a more complex robot-manipulation domain. Teacher A is rewarded by task completion, Teacher B by smooth, energy-efficient actuation. Environment
Actions
m
Native Objectives
Teacher Weights (wA ; wB )
MO-LunarLander MO-Hopper MO-Cheetah MO-MetaWorld (Drawer-Close)
Discrete Continuous Continuous Continuous
4 3 2 2
shaping, main-engine cost, side-engine cost, landing forward velocity, jump height, energy cost forward velocity, control cost task progress, control effort
[0.6, 0.3, 0.05, 0.05]; [0.6, 0.05, 0.3, 0.05] [0.8, 0.1, 0.1]; [0.3, 0.5, 0.2] [0.6, 0.4]; [0.4, 0.6] [0.6, 0.4]; [0.4, 0.6]
Table 12: Benchmark environments: native objective decomposition and scripted teacher weight vectors. m denotes the dimensionality of the native vector reward.
H
Compute Resources
In all experiments, we use 12 CPUs and a single GPU, of type either NVIDIA A100 or L40. Training in all environments takes approximately three to five hours on average.