Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap Yahan Lua , Dongyang Xiaa , Nursen Aydinb,∗ , Shadi Sharif Azadeha
arXiv:2609.24750v1 [math.OC] 21 Sep 2026
a
Department of Transport & Planning, Delft University of Technology, The Netherlands b Warwick Business School, University of Warwick, UK ∗: Corresponding author
The growing demand for real-time, data-driven decision-making in complex and dynamic systems is placing increasing pressure on traditional Operational Research (OR) methodologies. Reinforcement learning (RL) has emerged as a complementary approach, offering strong learning and computational capabilities for sequential decision-making in dynamic and uncertain environments. Recent research shows an increasing interest in integrating RL with OR to address dynamic decision-making problems, enhance heuristic and exact methods for combinatorial optimization, and support the development of digital replicas of operational systems. The overarching goal across these efforts is to leverage the learning capabilities of RL to strengthen traditional OR algorithms, improving solution quality, computational efficiency, and robustness. Given the diversity of integration approaches and application settings, there is a clear need for a systematic and technically detailed review of how RL empowers OR methods. To address this gap, this paper presents a structured review of three key roles that RL plays in empowering OR: (i) solving sequential decision-making problems in dynamic environments, (ii) serving as an end-to-end solution method or as a component integrated within heuristic and exact OR methods for combinatorial optimization problems, and (iii) facilitating extended reality analysis through integration with digital twin systems. We critically synthesize recent advances across these roles, highlighting their advantages, implementation requirements, limitations, and challenges. Finally, based on these insights, we outline a roadmap for future research to further advance the methodological and practical integration of RL and OR. Key words : Transportation, Reinforcement learning, Sequential decision-making, Digital twins, Heuristics and exact algorithms
1.
Introduction
Operational Research (OR) has long provided the foundational framework for supporting decisionmaking at the strategic, tactical, and operational levels. Advances in information systems, together with increasing urbanization, mobility, digitalization, and the rise of customer-centric services, have created a pressing need for anticipatory, real-time decision-making. While classical OR methods address these needs to a significant extent, processing continuous information streams and optimizing operations under highly dynamic and uncertain conditions have become increasingly challenging. 1
2
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
Reinforcement learning (RL) offers a data-driven framework capable of modeling dynamic behaviors and learning sequential decision-making policies in complex environments. Due to its flexibility and efficiency in tackling large-scale problems, RL algorithms have attracted growing attention in recent years and have demonstrated strong performance in high-dimensional Markov decision processes. Beyond sequential control, RL has also shown potential to enhance OR methodologies by improving exact algorithms (e.g., branch-and-cut, cutting-plane generation), and design smarter heuristic algorithms to solve complex optimization problems. Furthermore, RL can support emerging technologies by enabling the creation of digital replicas of complex systems, allowing for performance evaluation and testing under various conditions. In recent years, several review papers have summarized advances in RL. These include general surveys from a computer-science perspective (Kaelbling et al. 1996, Levine et al. 2020, Murphy 2025), reviews of multi-agent RL (Buşoniu et al. 2010, Hu et al. 2024), and domain-specific overviews in areas such as continuous control (Recht 2019), building control systems (Wang and Hong 2020), industrial process control (Nian et al. 2020), and robotics (Singh et al. 2022). Other studies examine RL from a safety perspective in applications such as autonomous driving and power systems (Gu et al. 2024, Su et al. 2025), while several reviews focus on deep RL methods more broadly (Mousavi et al. 2018, Arulkumaran et al. 2017, Li 2018, Wang et al. 2020, Ladosz et al. 2022) or in specialized fields such as fluid mechanics (Garnier et al. 2021) and healthcare (Yu et al. 2021). Additionally, dedicated reviews on model-free (Shakya et al. 2023) and model-based RL (Luo et al. 2024) further enrich this landscape. Despite this growing body of work, existing surveys focus primarily on RL algorithms, often within a single application area or algorithmic family, and do not examine how RL can be integrated with OR methodologies to improve decisionmaking and optimization performance. The interaction between RL and OR, especially the use of RL to enhance, support, or extend traditional OR algorithms, remains insufficiently explored. Within the broader field of machine learning, only a few studies review the integration of learning techniques with OR algorithms (Talbi 2016, Song et al. 2019, Bengio et al. 2021, Karimi-Mamaghan et al. 2022). For instance, Bengio et al. (2021) provide a methodological review of machine learning (ML) for combinatorial optimization, focusing on how learning components can be embedded within OR algorithms. Their taxonomy contrasts supervised/unsupervised learning with RL and describes integration patterns ranging from end-to-end learning to algorithm configuration and solver guidance. Karimi-Mamaghan et al. (2022) examine the use of ML to design components of metaheuristics, with RL treated as a relatively small subset compared with supervised and unsupervised learning. Wu et al. (2025) provide a tutorial on using deep reinforcement learning to solve OR problems via end-to-end learning, with an emphasis on foundational concepts and key background knowledge. Despite these important contributions, the literature still lacks a dedicated
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
3
and systematic overview of RL-OR integration, as well as a clear discussion of the associated research opportunities. Against this backdrop, two major gaps remain. First, existing review papers do not systematize how RL enhances heuristics, exact methods, and commercial solvers nor clarify the mechanisms through which RL improves algorithmic performance. Second, there is no unified taxonomy that maps the roles of RL, such as end-to-end decision making, guidance of OR algorithms, coupling with digital twins, to concrete integration points in OR workflows, including algorithm selection, initialization, operator learning, parameter tuning, and bound tightening. This paper addresses these gaps by providing a comprehensive review of how RL empowers OR. The work most closely related to ours is given by Bengio et al. (2021), which surveys ML for combinatorial optimization problems but treats RL as one element within a broader landscape. In contrast, our study provides a focused and in-depth analysis of how RL enhances OR models and workflows beyond classical combinatorial optimization. Before outlining our contributions, we clarify the scope of OR problems considered in this review. Classical combinatorial optimization problems operate under static inputs and aim to construct a one-shot solution satisfying all constraints. Sequential decision-making problems, by contrast, involve dynamic environments in which the system evolves between decisions and future realizations are unknown at the time of action. RL is naturally suited for such sequential settings and, in combinatorial optimization, it is mainly used either as an end-to-end constructive method or as a learned component that guides existing OR algorithms. This distinction motivates the taxonomy we adopt. To ensure analytical depth, we center our review on transport-related studies. However, the proposed taxonomy is applicable beyond the transportation domain. Our contributions are five-fold. First, we introduce a unified taxonomy that organizes the literature around three key roles of RL in enhancing OR: (i) directly solving sequential decision-making problems, (ii) serving as an end-to-end and independent solution method or as a component within heuristic and exact methods for combinatorial optimization problems, and (iii) enabling extended reality analysis through integration with digital twins. For the second role, we further provide a clear summary of the functions and integration points of RL within various OR algorithms. Second, we synthesize the main categories of RL algorithms for solving sequential decision-making problems with different structural characteristics. Third, we systematize how RL can guide both heuristic and exact optimization methods, and discuss when RL improves solution quality and convergence relative to traditional OR baselines. Fourth, we review RL-in-the-loop digital replicas used for training, accelerating the solution process of RL algorithms, and optimizing the digital twin operations. Fifth, we propose a roadmap that indicates when to favor different types of endto-end RL algorithms or hybrid RL-OR methods, and we highlight key opportunities, challenges, and future research directions in this interdisciplinary area.
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
4
The remainder of the paper is organized as follows. Section 2 provides background on RL. Section 3 outlines the review methodology. Section 4 focuses on the use of RL for sequential decisionmaking processes. Section 5 reviews state-of-the-art methods combining RL with OR algorithms where RL serves as a tool for solving combinatorial optimization problems. Section 6 examines applications that integrate RL with digital twins. Finally, Section 7 concludes the paper.
2.
RL framework
RL is a subfield of machine learning focused on learning decision-making policies through repeated interaction and feedback. It has been widely applied in areas such as automatic control, robotics, transportation logistics, and game playing. For a comprehensive introduction, we refer to Sutton and Barto (2018). In this section, we provide a concise overview of key RL concepts (Section 2.1) and the main categories of RL algorithms (Section 2.2). 2.1.
Key concepts of RL
RL models decision-making as an interaction between an agent and an environment. The key elements include a state, an agent, a policy, a reward signal, a value function, and optionally a model of the environment. As indicated in Figure 1, at each decision-making stage t, the agent observes a state St , selects an action At according to a policy π, receives a reward Rt+1 , and the environment transitions to a new state St+1 . The objective is to learn a policy that maximizes long-term reward.
Agent
reward Rt action At
state St Rt+1 Environment St+1 Figure 1
Illustration of the general RL setting.
Specifically, the policy determines which action the agent takes in each state. The value function estimates the expected long-term return from a state or state-action pair under a given policy. A model, when present, predicts state transitions and rewards. The agent is not explicitly told which actions to take; rather, it must learn through trial and error which actions yield high long-term returns. Consequently, actions influence not only immediate reward but also the sequence of future states and rewards.
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
5
RL offers a flexible framework for learning policies that optimize long-term performance in stochastic and dynamic environments without requiring an explicit analytical model. It is particularly effective in high-dimensional, sequential, and structurally complex decision-making problems, and can be combined with traditional optimization or control methods. Despite these advantages, RL faces several challenges. Designing an informative reward function is often difficult, especially in settings where rewards are sparse and provided only at task completion or when a feasible solution is identified. This limits learning signals and may lead the agent to fail to learn or converge to suboptimal strategies. In addition, RL methods must carefully balance exploration of new actions with exploitation of actions already known to be effective. To avoid poor outcomes, agents may behave overly conservative, while high-quality strategies often require extensive exploration, which can be costly in practice. 2.2.
RL algorithms
RL algorithms can be broadly categorized into model-free and model-based approaches, depending on whether the agent has access to (or learns) an explicit model of the environment. 2.2.1.
Model-free RL algorithms Model-free algorithms learn the value of actions directly
through extensive trial-and-error interactions with the environment. A simple analogy is navigating an unfamiliar city without a map. A traveler gradually identifies efficient routes by receiving positive feedback for good choices and negative feedback for poor ones. In a similar way, model-free methods improve their behaviour based solely on observed outcomes. Depending on their learning objective, they can be grouped into three categories: (i) Value-based RL algorithms. These algorithms estimate the value of state-action pairs and select actions by choosing those with the highest estimated values. The policy is derived indirectly by acting greedily with respect to the learned value function. Representative algorithms include Q-learning (Watkins and Dayan 1992), Deep Q-Network (DQN) (Mnih et al. 2013), Double DQN (Van Hasselt et al. 2016), Dueling DQN (Wang et al. 2016), and Value-based Augmented Proximal Policy Optimization (VAPO) (Yue et al. 2025). (ii) Policy-Based RL algorithms. These algorithms learn a policy function directly by estimating the probability of taking an action in a given state. Instead of evaluating individual actions, they focus on maximizing the expected return of the entire policy. This is typically achieved through gradient-based updates, increasing the likelihood of effective actions. Examples include Trust Region Policy Optimization (TRPO) (Schulman et al. 2017a) and Proximal Policy Optimization (PPO) (Schulman et al. 2017b). (iii) Actor-Critic RL algorithms. These algorithms combine the advantages of value-based and policy-based approaches. The actor proposes actions according to a policy network, while
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
6
the critic evaluates them using a value function. Both components are trained simultaneously to improve decision-making efficiency and stability. Representative algorithms include Advantage Actor-Critic (A2C) (Konda and Tsitsiklis 1999), Deep Deterministic Policy Gradient (DDPG) (Lillicrap et al. 2015), Asynchronous Advantage Actor-Critic (A3C) (Mnih et al. 2016), Soft ActorCritic (SAC) (Haarnoja et al. 2018), and Twin Delayed DDPG (TD3) (Fujimoto et al. 2018). 2.2.2.
Model-based RL algorithms Model-based RL algorithms first learn an explicit
model of the environment, capturing the transition dynamics and reward function. That is, they learn how actions taken in a given state lead to new states and rewards. Once such a model is learned, it can be used to simulate future trajectories and plan actions that maximize long-term performance. This approach is similar to studying a city map to understand how locations are connected and then planning the most efficient route without physically exploring every option. Algorithms in this category typically involves two components: (i) model learning, where an environment simulator is trained using collected data to replicate state transitions and rewards, and (ii) policy planning, where planning or search methods such as trajectory rollouts or tree search are used within the learned model to identify optimal actions. Representative algorithms include Dyna-Q (Sutton 1991), Model-based Value Expansion (MVE) (Feinberg et al. 2018), Stochastic Ensemble Value Expansion (STEVE) (Buckman et al. 2019), Model-Based Policy Optimization (MBPO) (Janner et al. 2019), and Robust Adversarial ModelBased Offline Reinforcement Learning (RAMBO-RL) (Rigter et al. 2022). By simulating interactions within the learned model, these algorithms reduce reliance on costly real-world exploration and often accelerate learning.
3.
Taxonomy and review methodology
In this section, we first present the taxonomy of our review and then outline the search strategies and inclusion criteria. 3.1.
Taxonomy
The purpose of this review is to examine how RL contributes to OR across domains of transportation logistics, energy systems, and healthcare. As illustrated in Figure 2, we frame the integration of RL into OR through three complementary angles: (i) RL for solving sequential decision-making problems. These problems involve dynamic systems in which critical information (e.g., demand, disturbances, opponent behavior) is revealed over time. The objective is to optimize cumulative system performance, measured by criteria such as discounted reward, finite-horizon return, or long-run average reward. Such problems are typically formulated as Markov Decision Processes and solved using RL algorithms.
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
7
(ii) RL as an optimization tool for combinatorial optimization problems (COPs). In COPs, the objective is to construct a discrete solution object (e.g., a route, schedule, assignment, or facility set) based on complete knowledge of the problem instance at the outset. Time is not dynamic in this setting but serves as a modeling device to construct static solutions. RL can be employed as a standalone solver, integrated with heuristic methods, or embedded into exact algorithms to enhance optimization performance. (iii) RL embedded in digital twins. RL also supports OR through its integration with digital twin systems, where it can leverage high-fidelity simulations for training, accelerate computational procedures, and support the optimization of operational decisions within virtual replicas of realworld systems.
RL
Sequential decision-making problems with dynamics
Tailoring resolution methodologies
Integration with digital twins: Training and optimization
End-to-end solver for COPs RL-enhanced heuristic algorithms RL-enhanced exact algorithms * RL: Reinforcement Learning ** COPs: Combinatorial Optimization Problems Figure 2
Taxonomy on the use of RL in OR and creating a digital replica of reality.
We include studies that satisfy the following conditions: (a) the study addresses an optimization or control problem within OR-relevant domains (e.g., transportation, logistics, energy systems, healthcare), and (b) the study explicitly involves an RL component. For DT-related work, we include studies where the digital twin serves as an operational decision-support system, enabling optimization, control, or performance evaluation typical of OR applications. A formal OR model is not required, but the study must define operational objectives and constraints and evaluate them through the DT. 3.2.
Review methodology
To ensure comprehensive coverage, we conducted literature searches across major scientific databases, including Scopus, Web of Science Core Collection, IEEE Xplore, ScienceDirect, SpringerLink, and INFORMS PubsOnline. In addition, we performed a venue-restricted search of leading
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
8
OR journals (e.g., INFORMS journals, European Journal of Operational Research (EJOR), and Transportation Research Parts A–F) as well as top-tier AI conferences (NeurIPS, ICML, ICLR, AAAI) to capture online-first and early-access papers not yet indexed in databases. We adopted a layered keyword search strategy that combines three mandatory blocks and an optional refinement block: Query = (A RL) ∧ (B applications) ∧ (C OR methods). A (RL) includes RL algorithms. Terms are grouped by families as follows: • Generic & model-based RL family. “reinforcement learning” ∨ “deep reinforcement learning” ∨ “model-free reinforcement learning” ∨ “model-based reinforcement learning” ∨ “model-based value expansion”. • Value-based RL family. “value-based reinforcement learning” ∨ “Q-learning” ∨ “deep QNetwork” ∨ “double deep Q-network” ∨ “dueling deep Q-network”. • Policy-based RL family. “policy-based reinforcement learning” ∨ “trust region policy optimization” ∨ “proximal policy optimization” ∨ “value-based augmented proximal policy optimization” ∨ “deep deterministic policy gradient” ∨ “twin delayed deep deterministic policy gradient”. • Actor-critic RL variants. “actor-critic” ∨ “advantage actor-critic” ∨ “asynchronous advantage actor-critic” ∨ “soft actor-critic”. B (applications) consists of research domains. Domains are grouped by themes as follows: • Transport & logistics. “transport”
∨
“logistics”∨
“vehicle
routing”
∨
“scheduling”
∨
“timetabling” ∨ “line planning” ∨ “network design”. • Operations management. “inventory control” ∨ “assortment” ∨ “pricing” ∨ “revenue management” ∨ “facility location” ∨ “resource allocation”. • Healthcare & energy. “healthcare” ∨ “power system” ∨ “energy”. C (OR methods). OR methods are grouped by categories as follows: • Mathematical programming & decomposition. “operational research” ∨ “mixed-integer programming” ∨ “integer programming” ∨ “linear programming” ∨ “combinatorial optimization” ∨ “column generation” ∨ “Benders decomposition” ∨ “Lagrangian relaxation” ∨ “decomposition”. • Branching & cutting families. “branch-and-price” ∨ “branch-and-cut” ∨ “cutting plane” ∨ “branch and bound”. • Constraint-based family. “constrained (or relaxed) decision diagrams” ∨ “constraint programming”. • Solvers and generic. “CPLEX” ∨ “GUROBI” ∨ “Markov decision process” ∨ “MDP” ∨ “digital twin” ∨ “digital replica”.
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
9
We followed PRISMA-style stages for systematic review: Identification, Screening, Eligibility, and Inclusion. The inclusion criteria are as follows: (i) the paper includes an explicit RL component, (ii) there is a clear OR connection through models, algorithms, or applications, and (iii) the paper was published between 2016 and 2026.
4.
RL: A dynamic sequential decision-making tool
In this section, we review studies that apply RL algorithms to solve sequential decision-making problems in dynamic environments. These problems are typically formulated as Markov Decision Processes (MDPs, e.g., Powell 2007), where the aim is to learn a policy that selects actions over time to maximize cumulative rewards based on state observations and feedback from the environment. We classify the reviewed studies based on the RL algorithms they adopt and annotate each work using a set of orthogonal labels that capture key modeling and implementation aspects. We first analyze the literature across different RL algorithm families and discuss the problem characteristics suited to each, and then outline key challenges and future research directions. Specifically, Section 4.1 reviews studies that solve MDPs using value-based methods. Section 4.2 focuses on policy-based RL algorithms. Section 4.3 surveys actor–critic variants. Section 4.4 discusses model-based RL approaches. and Section 4.5 outlines cross-cutting challenges and future research directions. Given the uneven volume of literature, our coverage follows a consistent convention: for Sections 4.1 and 4.3, where the number of studies is large, we synthesize cross-study patterns and cite representative works; for Sections 4.2 and 4.4, where the literature is relatively limited, we briefly discuss each study for completeness. 4.1.
Solving MDPs with value-based RL algorithms.
Given the large number of studies in the category of applying value-based RL algorithms to MDPs, we synthesize cross-study patterns rather than reviewing related studies individually. Table 1 summarizes state-of-the-art studies, reporting domain/task, problem signature, the employed RL algorithm, training mode, and baselines. Representative applications include railway scheduling (e.g., Semrov et al. 2016, Wang et al. 2025, Yu and Hyland 2025), crowdsourced urban delivery (e.g., Ahamed et al. 2021), same-day delivery with vehicles and drones (e.g., Chen et al. 2022, 2023), and inventory control (e.g., Oroojlooyjadid et al. 2022, Mao et al. 2025), which typically feature discrete actions, high observability, and mostly episodic horizons. Overall, value-based methods are most frequently used in resource allocation and scheduling, urban delivery and crowdsourcing, inventory and credit control, and digital marketing-settings with predominantly discrete operational decisions (e.g., replenishment, dispatch/retention, promotion selection). Most studies assume full observability, with only a few relying on partial or binary observations. Time horizons are typically finite or episodic (daily or period-based), with
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
10
occasional infinite-horizon discounted models in inventory control. The underlying dynamics commonly incorporate randomness or non-stationarity, such as demand variation, stochastic travel times, perishability, user heterogeneity, and shifting reward structures. These are addressed through period-by-period learning or restart strategies. Training is conducted both online and offline, the latter often using historical logs or simulation, consistent with the off-policy nature of value-based methods. Baseline comparisons range from rule-based and heuristic policies to behavioral policies and human benchmarks, with fewer comparisons against simulation-only strategies. In conclusion, sequential decision-making problems with discrete actions, high observability, and episodic or discrete-time formulations are well suited to value-based RL algorithms. Transportation planning and inventory control can also benefit from offline learning pipelines, a setting where value-based methods are particularly effective. Table 1
Classification of papers studying value-based RL algorithms for sequential decision-making.
Publications Domain / Task
Problem signature
Railway / Train rescheduling Crowdsourced urban delivery Same-day delivery with vehicles and drones Beer game / Multi-echelon inventory
A: discrete; O: full; Horizon: Semrov et al. (2016) episodic; Dyn: disturbances A: discrete; O: full; Ahamed et al. (2021) Horizon: episodic A: discrete; O: full; Chen et al. (2022) Horizon: daily; Dyn: stochastic A: discrete; O: partial/local; Oroojlooyjadid et al. (2022) Horizon: finite; Dyn: lead-time lags A: discrete; O: full; Synchromodal logistics / Horizon: finite episodic; Guo et al. (2022) Shipment matching Dyn: dynamic and stochastic travel times Digital marketing / A: discrete; O: full; Wang et al. (2023) Sequential personalized Horizon: 30-day episodic; promotions Dyn: stochastic users A: discrete; O: full; Chen et al. (2023) Same-day delivery Horizon: daily; Dyn: stochastic A: binary; O: fully observed; Credit limit Horizon: episodic over customers; Alfonso-Sánchez et al. (2024) adjustment Dyn: stochastic A: discrete; O: fully observed; Reusable resource Horizon: finite; Dyn: Wang et al. (2025) allocation & pricing non-stationary reward Yu and Hyland (2025) Mao et al. (2025)
Transportation / Strategic planning
A: continuous; O: partially observed
Nonstationary MDPs / Inventory control
A: discrete; O: full; Horizon: episodic
RL method
Training
Baselines
Q-learning
Online
FIFO; Random walk
DQN
Offline
Heuristics
DQN
Offline
PFA
DQN
Online
Base-stock rules; Behavioral policies
Q-learning
Offline
MA
DDDQN
Online
DQN
Offline
Double Q-learning
Offline
Alternative strategies
Value-based episodic RL
Online
Full-information greedy oracle; ε-greedy variants
Mass policies; Myopic; Heuristics; DQN variants Bucket policy; Reserved-vehicle policy
Human reactive; Simulator Random; Value-only policy Q-Learning UCB; RestartQ-UCB Simulator ε-greedy DQN
A = action type; O = observation; Dyn = dynamics. Use “N/A” if not specified. FIFO = First-In- First-Out, DQN = Deep Q-Network, PFA = Policy Function Approximation, MA = Myopic approach, DDDQN = Double Dueling DQN, RestartQ-UCB = Restarted Q-learning with Upper Confidence Bounds.
4.2.
Solving MDPs with policy-based RL algorithms.
Table 2 summarizes advancements in solving MDPs using policy-based RL algorithms. These methods have been applied across diverse domains including information retrieval (Wei et al. 2017), inventory lot-sizing (Dehaybe et al. 2024), container port truck dispatching (Jin et al. 2024), urban
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
11
rail transit rescheduling (Ying et al. 2024), job-shop scheduling (Monaci et al. 2024), and equipment maintenance (Verleijsdonk et al. 2024). As can be observed from Table 2, most applications involve real-time decision-making problems. Action types are predominantly discrete, with a few examples involving continuous actions such as inventory control. Observability is generally full. Most problems feature finite or episodic horizons, while some inventory and maintenance settings operate under infinite-horizon discounting. Except for learning-to-rank problems, which are nearly static, most studied environments involve randomness or non-stationarity arising from factors such as demand variations, schedule disruptions, or system degradation. Training strategies vary accordingly. Online learning is used for scheduling and inventory applications, offline learning is applied to ranking tasks, and simulation-based rollouts are employed in maintenance problems. The baseline algorithms also vary widely, ranging from learning-to-rank methods, approximate dynamic programming, heuristics, and handcrafted rules to deep reinforcement learning (DRL), DRL-based hyper-heuristics, metaheuristics, distributed PPO, and commercial solvers. Table 2
Classification of papers studying policy-based RL algorithms for sequential decision-making.
Publications Domain / Task
Problem signature
Information retrieval / A: discrete; O: full; Wei et al. (2017) Learning to rank Horizon: finite episodic; Dyn: static A: continuous; O: fully observed; Horizon: Dehaybe et al. (2024) Inventory lot-sizing rolling discounted; Dyn: non-stationary Container port A: discrete; O: fully observed; Jin et al. (2024) truck dispatching Horizon: finite; Dyn: stochastic Urban rail / A: discrete; O: fully observed; Ying et al. (2024) Disruption Horizon: finite; Dyn: non-stationary rescheduling Job-shop A: discrete; O: full; Monaci et al. (2024) scheduling Horizon: episodic A: discrete per engineer; O: full; Horizon: Verleijsdonk et al. (2024) DTMPA infinite; Dyn: stochastic degradation
RL method Training
Baselines
Policy gradient
Offline
PPO
Online
PPO
Online
Learning-to-rank algorithms ADP; Heuristic Handcrafted rule; DRL-HH
Multi-agent Online PPO PPO API with DCL
Metaheuristics; distributed PPO
CPLEX; DRL In-simulation Ranking rollouts heuristics
Online
A = action type; O = observation; Dyn = dynamics. Use “N/A” if not specified. PPO = Proximal Policy Optimization, ADP = Approximate Dynamic Programming, DRL = Deep Reinforcement Learning, DRL-HH = Deep Reinforcement Learning-based Hyper-Heuristic, DTMPA = Dynamic Traveling Multi-Maintainer with Alerts, API = Approximate Policy Iteration, DCL = Deep Controlled Learning.
We briefly discuss representative studies that span ranking versus control, discrete versus continuous action spaces, and offline versus online training.Wei et al. (2017) formulate ranking as an episodic Markov decision process and employ policy-gradient methods trained offline on logged user interactions. The action space is discrete, corresponding to ranking positions, the system is fully observable, and benchmarks include classical learning-to-rank algorithms; the RL formulation explicitly optimizes long-term engagement signals rather than myopic relevance. Dehaybe et al. (2024) consider continuous ordering decisions under non-stationary demand in a rolling-discounted setting, using PPO trained online in simulation. The problem features continuous (or mixed) actions with full observability, comparisons are made against ADP/heuristic baselines, illustrating
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
12
policy-based RL’s suitability for continuous-control inventory problems. Ying et al. (2024) adopt a multi-agent PPO framework for real-time rescheduling under stochastic disruptions. Their setting involves discrete operational decisions with full observability, online training in a simulator, and baselines including metaheuristics and distributed PPO variants, highlighting policy-based RL in time-critical operational control. In summary, policy-based RL is often preferred in scenarios characterized by high randomness, online or rolling decision-making, and continuous or mixed action spaces. These methods directly optimize the policy objective, demonstrate robustness to sparse or delayed rewards, and can be continuously improved through simulation. They are particularly suitable for problems with continuous or hybrid action spaces, explicit policy constraints or action masking requirements, and multi-agent interactions that involve cooperation or competition. 4.3.
Solving MDPs with actor-critic RL algorithms.
Given the breadth and heterogeneity of studies employing actor–critic RL algorithms to solve MDPs, we focus on distilling recurring patterns rather than reviewing individual studies one by one. Table 3 summarizes recent work in this area. As shown, these methods are most often applied in settings with continuous or mixed action spaces, and in environments characterized by stochastic dynamics, multi-agent interactions, or safety-critical constraints. Application domains include railway scheduling and timetabling, ride-hailing dispatch, autonomous mobility-on-demand systems, large-scale network control, safety-constrained continuous control, and manufacturing resource allocation. A comprehensive review of these studies reveals three key trends. First, algorithm choice is closely tied to the action space: Deep Deterministic Policy Gradient (DDPG), Soft Actor-Critic (SAC), and Actor-Critic-Proximal Policy Optimization (AC-PPO) are applied to continuous or mixedaction problems, while Multi-Agent Actor-Critic (MAA2C) and multi-agent DDPG are adopted in decentralized decision-making settings. Second, training methods reflect system accessibility. Online learning is prevalent in operational problems that require real-time adaptation, simulators are used when real-world interaction is costly or risky, and imagined rollouts are leveraged to enhance data efficiency. Third, benchmark baselines remain diverse, ranging from heuristic rules and greedy strategies to metaheuristics, classical optimization tools (e.g., CPLEX), MPC, and various PPO variants. In summary, actor-critic RL algorithms are particularly effective for exploring high-dimensional action spaces, adapting to non-stationary dynamics through continual policy updates, and enabling coordination in multi-agent settings via decentralized actors with a shared or centralized critic. Safe or constrained variants can incorporate penalties or action masks during learning, offering practical advantages in transportation, mobility and process control applications.
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
Table 3
13
Classification of papers studying actor-critic RL algorithms for sequential decision-making.
Publications Domain / Task
Problem signature
A: continuous; O: fully observed; Horizon: finite; Dyn: non-stationary A: large discrete; O: full; Horizon: finite; Dyn: multi-state A: discrete; O: partial; Horizon: finite; Multi-agent A: continuous policy over zones; Ride-sourcing O: aggregated; Horizon: finite Railway / A: discrete; O: local (actors), Train timetabling global (critic); Horizon: episodic A: continuous embedding; O: fully Urban rail / Coordinated operations observed; Horizon: finite AMoD / A: discrete; O: fully observed; Request assignment Horizon: finite episodic; & rejection Dyn: stochastic arrivals A: mainly continuous; Large-scale network O: local/partial; Horizon: finite episodic; control / Traffic, power, Dyn: stochastic exogenous inputs and pandemic, CACC non-stationary interactions A: continuous; O: fully observed; Safe RL (Safety Gym) / M: learned dynamics ensemble; Continuous control Horizon: episodic (truncated H); Dyn: stochastic transitions & costs Manufacturing / A: continuous quantities; O: fully observed; Dynamic resource matching Horizon: finite; Dyn: stochastic demand Seru production / A: discrete; O: full; Dynamic worker allocation Horizon: finite; Dyn: deterministic
Metro rescheduling / Ying et al. (2020) Headway control Dynamic selective Liu et al. (2020) maintenance Disaster response / Lee and Lee (2021) DSPA Zhu et al. (2021) Li et al. (2022) Ying et al. (2022) Enders et al. (2023)
Ma et al. (2024)
Jayant and Bhatnagar (2022)
Panda et al. (2024) yu Li et al. (2025)
RL method
Training
Baselines
DDPG
Online
DE; PSO; GA
Actor-Critic
Online
DP
Multi-agent Actor-Critic
Online
Actor-Critic
Simulator
MAA2C
Simulator
FCFS; Naive policy; Oracle upper bound Multi-agent RL variants GA; PSO
Multi-agent DDPG Onlin
GA; Distributed DDPG
Multi-agent SAC
Simulator
Greedy; MPC
Multi-agent
Online
PPO variants
Actor-Critic
Imagined rollouts PPO-Lagrangian; CPO; + real env data safe-LOOP (MBRL)
DKDDPG
Online
AC-PPO
Simulator
Exact (small); DKQL; DQN; DPG; DDPG Standard AC; PPO; Heuristic rules.
A = action type; O = observation; Dyn = dynamics. Use “N/A” if not specified. DDPG = Deep Deterministic Policy Gradient, DE = Differential Evolution, PSO = Particle Swarm Optimization, GA = Genetic Algorithm, DP = Dynamic Programming, DSPA = Decentralized Selective Patient Admission, FCFS = First-Come First-Serve, MAA2C = Multi-Agent Actor-Critic, SAC = Soft Actor-Critic, MPC = Model Predictive Control, DKDDPG = DDPG with Domain-Knowledge Q Penalty, DKQL = Domain Knowledge-informed Q-learning, DQN = Deep Q-Network, DPG = Deterministic Policy Gradient, AC = Actor-Critic.
4.4.
Solving MDPs with model-based RL algorithms.
Model-based RL algorithms require learning a transition and reward model, but many OR problems are high-dimensional, stochastic, mixed-action, and heavily constrained, making such model learning especially difficult. As a result, research employing model-based RL algorithms remains relatively limited. Table 4 summarizes the state-of-the-art research in this area. Clavera et al. (2018) propose a Model-Based Meta-Policy Optimization (MB-MPO) RL algorithm for fully observed, episodic continuous control problems. Their method fits an ensemble of dynamics models to capture uncertainty, generates imagined rollouts on these models, and updates the policy using trust-region gradients. Only a small amount of real interaction is used to periodically refit the models and incorporate new data. Lecarpentier and Rachelson (2019) introduce Risk Averse Tree Search (RATS) for non-stationary MDPs. The approach constructs a snapshot model at each epoch and performs risk-averse closed-loop tree search with replanning at every step under infinitehorizon discounting and full observability. Kidambi et al. (2020) develop Model-Based Offline Reinforcement Learning (MOReL) for offline continuous control. Their method learns an ensemble dynamics model with uncertainty estimates from logged data, constructs a pessimistic MDP that routes uncertain transitions to a low-value absorbing state, and performs policy improvement entirely within this model without new interactions.
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
14 Table 4
Classification of papers studying model-based RL algorithms for sequential decision-making. Publications Domain / Task
Problem signature
RL method Training
A: continuous; O: fully observed; M: learned dynamics ensemble; Clavera et al. (2018) Continuous control MB-MPO Horizon: episodic A: discrete; O: full; Horizon: infinite discounted; Robust planning in Lecarpentier and Rachelson (2019) RATS non-stationary MDPs NSMDP with Lipschitz evolution; snapshot model available each epoch A: cont.; O: full; Offline data; Offline Model: learned dynamics ensemble; MOReL Kidambi et al. (2020) continuous control P-MDP
Baselines
DDPG; TRPO; Imagined rollouts PPO; ACKTR; on learned models ME-TRPO; MB-MPC Closed-loop tree search
DP-snapshot; DP-NSMDP
Offline
BCQ; BEAR; BRAC; naive MBRL
A = action type; O = observation; M = model; Dyn = dynamics. Use “N/A” if not specified. MB-MPO = Model-Based Meta-Policy-Optimization, DDPG = Deep Deterministic Policy Gradient, TRPO = Trust Region Policy Optimization, ACKTR = Actor-Critic using Kronecker-Factored Trust Region, ME-TPRO = Model-Ensemble Trust-Region Policy Optimization, MB-MPC = Model-Based Model Predictive Control, NSMDP = Non-Stationary MDP, DP-snapshot = Dynamic Programming-snapshot, DP-NSMDP = Dynamic Programming-Non-Stationary MDP, P-MDP = Pessimistic MDP, MOReL = Model-Based Offline Reinforcement Learning, BCQ = Batch-Constrained Q-learning, BEAR = Bootstrapping Error Accumulation Reduction, BRAC = Behavior Regularized Actor-Critic, MBRL = Model-Based Reinforcement Learning.
In summary, MB-MPO improves sample efficiency through imagined data, RATS enables riskrobust planning under model drift, and MOReL curbs model-exploitation in offline settings. Collectively, these approaches are particularly well suited to domains in which real-world interaction is costly or risky, offering safer and more sample-efficient alternatives to purely model-free methods while enabling an explicit treatment of uncertainty, constraints, and non-stationarity. 4.5.
Discussion & future research directions
We now consolidate the main insights across RL algorithm families, uncover their fundamental limitations, and point to promising directions for future research. Insights across RL algorithm categories. The analyzes above reveal that value-based methods are well-suited to environments with discrete action spaces, full observability, and finite or segmented decision horizons. In contrast, settings characterized by strong stochasticity, rolling online control, continuous or hybrid action spaces, or multi-agent interactions tend to favor policy-based or actor-critic approaches. When interactions are costly or risky, the environment is non-stationary, safety constraints must be observed, or historical logs are available, model-based RL algorithms become a more appropriate choice. Regarding the performance-cost trade-offs, model-based and offline methods typically achieve higher sample efficiency, but this comes with substantial training and engineering overhead due to model learning, simulator construction, and uncertainty quantification. These approaches also face scalability challenges, such as high-dimensional encoder-decoder architectures, long-horizon credit assignment, and sparse or delayed reward signals. Overall, model-free and model-based RL represent two distinct paradigms that differ fundamentally in how they interact with and learn from the environment. In model-free RL frameworks,
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
15
the agent directly maps observations to actions based on past experiences. It learns to act by trial and error, updating its behavior using reward prediction errors. This makes model-free methods effective in highly complex or unpredictable environments where modeling transitions is infeasible or unnecessary and where real-time decision-making can be supported through abundant interactions or reliable simulators.Model-based RL, in contrast, constructs an internal model of state transitions and rewards, enabling planning through simulated trajectories. These methods excel in structured or deterministic environments, where long-term predictions are more reliable. Modelbased RL algorithms are often favored when sample efficiency is critical, or prior knowledge about the system’s structure is available. They allow agents to learn faster with fewer interactions, albeit at the cost of higher training and modeling complexity. Ultimately, the choice between model-free and model-based RL algorithms depends on the nature of the environment and practical considerations such as data availability, computational resources, and operational requirements such as safety, interpretability, or planning depth. Limitations and future research directions. Model-free RL algorithms have made remarkable progress in recent years. However, their low sample efficiency presents a significant limitation, that is, they typically require extensive real-world interactions with the environment to collect sufficient training data and learn effective policies. This constraint hinders their applicability to real-world domains where interactions are costly, time-consuming, or risky, confining their practical use largely to settings with accessible and reliable simulators. Moreover, model-free methods lack the ability to simulate hypothetical future trajectories, relying solely on real interactions rather than imagined rollouts. This limitation not only reduces the efficiency of exploration but also restricts the agent’s ability to anticipate long-term consequences. For challenges specific to value-based model-free RL algorithms, we refer to Park et al. (2024). As for model-based RL algorithms, despite their potential to improve sample efficiency and enable planning through simulated rollouts, they also face important limitations. First, they rely on accurate modeling of environment dynamics. In systems with complex, discontinuous, or stochastic behavior, learning such models can be difficult and may outweigh potential benefits, making these algorithms less effective. Second, model-based RL introduces additional sources of estimation error: not only from value function approximation but also from the learned dynamics model. When both are imperfect, compounded errors can significantly degrade policy quality. Third, these algorithms are highly sensitive to model accuracy. An inaccurate model can mislead the agent during planning, resulting in poor generalization or unsafe decisions, especially in long-horizon tasks where prediction errors accumulate. Therefore, reliable model learning is essential, yet often challenging in highdimensional or partially observable environments.
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
16
Looking forward, an important future research direction is the development of hybrid frameworks that enable flexible coordination between model-free and model-based RL algorithms within a single decision process, thereby fully combining their advantages. Specifically, one direction worth exploring is the development of centralized decision-making architectures that can dynamically coordinate the use of model-free or model-based RL algorithms based on the characteristics of the problem, such as decision-making stages, uncertainty, and dynamics. Such centralized coordination frameworks could allow agents to exploit model-based planning for rapid learning, and then switch to model-free strategies for robustness and policy generalization, or alternate between the two to improve both computational efficiency and solution quality.
5.
RL: Enhancer of solution methods for COPs
In this section, we focus on the research stream in which RL serves as a tool to enhance OR solution methods. A growing body of work treats RL as an end-to-end, independent solution method for solving COPs. Here, COPs refer to problems where the goal is to construct a static solution (e.g., a route or schedule) in a one-shot fashion, rather than to learn a policy for dynamic, sequential decision-making environments. This line of research leverages the exploration capabilities of RL to learn a mapping from problem inputs to solutions and to directly search the solution space. In addition, RL can be integrated with heuristic or exact algorithms to tailor solution methodologies for COPs, typically by guiding parts of the search process, such as selecting promising neighborhoods or cutting planes. Section 5.1 reviews studies in which RL serves as an end-to-end and independent solution method, Section 5.2 examines RL-enhanced heuristic algorithms, and Section 5.3 discusses RL-enhanced exact algorithms. In each subsection, we analyze the relevant literature, summarize the associated challenges and limitations, and propose potential directions for future research. 5.1.
RL as an end-to-end and independent solution method
When RL is used as an end-to-end, independent solution method for COPs, it learns a direct mapping from problem instances to feasible solutions, which are typically decoded using greedy, sampling-based, or beam-search strategies. We first review classic COPs such as the Traveling Salesman Problem (TSP), Vehicle Routing Problem (VRP), and Minimum Vertex Cover Problem (MVCP) in Section 5.1.1, as early research on end-to-end RL methods primarily focused on these problems. We then examine applications in more domain-specific combinatorial optimization contexts in Section 5.1.2. Finally, we discuss key challenges and potential future research directions in Section 5.1.3. In this research context, terms such as Pointer Networks and Transformers refer to the neural network architectures used to parameterize policies or value functions. These choices are orthogonal to the underlying RL algorithms (value-based, policy-based, actor-critic) and to whether the
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
17
method is model-free or model-based. These architectures are frequently adopted in this line of research, as they naturally align with the structural properties of classic COPs with routing decisions and support the construction of high-quality solutions in an end-to-end manner. Specifically, pointer networks (see, Vinyals et al. 2015) are sequence-to-sequence models whose decoder “points” to input items, making them particularly suitable for constructing permutations (e.g., routing or scheduling). Transformers (see, Parmar et al. 2018, Kool et al. 2019) use self-attention mechanisms to encode variable-size sequences or sets and support autoregressive decoding, offering improved scalability and performance on large-scale problem instances. 5.1.1.
Literature analysis of classic COPs Focusing on RL as an end-to-end, independent
solution method for solving classic COPs, Table 5 classifies representative studies along four dimensions: (i) RL paradigm and architecture (e.g., as policy gradient, value-based, actor-critic; Pointer Networks, attention/Transformers); (ii) Solved COPs; (iii) Baselines, i.e., benchmark algorithms used for comparison (e.g., Christofides, OR-Tools, Concorde); and (iv) Problem scale. As shown in Table 5, graph-structured problems, especially TSP and VRP, are the most frequently studied, and classic heuristics are commonly used as baselines. Across these studies, three recurring patterns emerge: (i) routing problems often use policy gradient methods with pointer or Transformer decoders, while general graph problems prefer value-based methods with graph neural network (GNN) encoders; (ii) recent studies shift from pointer networks to Transformer architectures as problem scale increases; (iii) baseline selection varies across domains, making it difficult to compare results across studies due to the lack of standardized evaluation protocols. Table 5
Classification of papers studying RL as an end-to-end and independent solution method for solving classic COPs.
Publications RL paradigm / architecture Bello et al. (2017) Nazari et al. (2018) Kool et al. (2019) Dai et al. (2017) Barrett et al. (2020) Li et al. (2021) Jin et al. (2023)
Solved COPs
Baselines
Problem scale
Christofides, OR-Tools, Policy gradient; Pointer Network TSP, Knapsack 20–100 nodes and Concorde (exact) Policy gradient; Pointer-like seq2seq VRP Classic heuristics and OR-Tools 10–100 nodes OR-Tools, classical heuristics, Policy gradient; Transformer TSP, VRP 20–100 nodes Concorde (exact), and GUROBI Value-based; GNN TSP, MVCP, MAXCUT Approximation and heuristics 50–100 nodes An RL-based heuristic Value-based; GNN + vertex flipping Graph-structured COPs 20–200 vertices and a greedy algorithm Policy gradient; Pointer Network MOTSP NSGA-II, MOEA/D, MOGLS 20–500 nodes Policy gradient; Transformer pointer TSP Existing deep-learning algorithms 20–500 nodes
Notes: TSP = Traveling Salesman Problem, VRP = Vehicle Routing Problem, GNN = Graph neural networks, MVCP = Minimum Vertex Cover Problem, MAXCUT = Max-Cut Problem, MOTSP = Multiobjective Traveling Salesman Problem.
We now turn to analyze these studies in detail. As one of the earliest efforts, Bello et al. (2017) introduced a pioneering neural combinatorial optimization framework that applies RL to solve the TSP. Leveraging a pointer network trained via policy gradients, their algorithm learns to generate
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
18
high-quality tours by directly minimizing tour length. They also include the Knapsack Problem to demonstrate generalizability. This work was the first to show the potential of an end-to-end, trainable RL-based framework for COPs, inspiring substantial follow-up research. Subsequent studies extend this framework to VRPs and other large-scale COPs. For instance, Nazari et al. (2018) adapted pointer-like sequence models to solve VRPs with both static and dynamic elements, while Dai et al. (2017) combined RL with graph embeddings to address a range of graph-structured COPs (e.g., Minimum Vertex Cover Problem, Max-Cut Problem, and TSP) by learning a greedy construction policy that acts as a metaheuristic. Li et al. (2021) proposed a deep RL method for the multi-objective TSP. The multi-objective optimization problem is decomposed into several single-objective subproblems, where each is formulated as an RL decision process under a given weight vector. The trained model functions as a black-box metaheuristic that generates solutions without per-instance reoptimization. Transformer-based methods further improve scalability and sample efficiency (e.g., Kool et al. 2019). More recently, Jin et al. (2023) proposed an end-to-end solution method named Pointerformer, which can solve instances with up to several hundred nodes. Beyond TSP, Barrett et al. (2020) developed a deep Q-learning approach for graph-structured COPs that integrates GNN encoders with local-search-inspired mechanisms (e.g., vertex flipping and reward shaping). Applied to the Max-Cut problem, their method outperforms prior RL-based baselines and can also be combined with other search heuristics. 5.1.2.
Literature analysis of solving other COPs Beyond classic COPs, a growing body
of research applies RL to other combinatorial problems in domain-specific contexts. Table 6 classifies related studies along four dimensions: (i) Research domain; (ii) Addressed problem; (iii) RL paradigm and architecture; and (iv) Baselines, i.e., benchmark algorithms used for comparison. Table 6
Classification of papers studying RL as an end-to-end and independent solution method for solving domain-specific COPs. Publications Domain
Ma et al. (2025) yu Li et al. (2025) Yu et al. (2026) Su and Yang (2025) Ding et al. (2025) Meng et al. (2025) Tian et al. (2025) Teusch et al. (2025) Wang et al. (2025) Vanvuchelen et al. (2024) Teck et al. (2025)
Supply chain Public Transport Manufacturing Public transport Electric robots Disaster response Logistics Urban mobility planning Public transport Healthcare supply chain Robotic mobile fulfillment
Problem
RL paradigm / architecture
Baselines
Product design change Headway optimization Job-shop scheduling Route-frequency design Charging-robot scheduling Volunteer management Service scheduling Facility location Train timetabling Lateral transshipment Inventory optimization
Actor-critic / Bi-level SAC Value-based / DDQN PPO / CNN PPO / Action masking Policy-gradient / Transformer Actor-Critic / critic heads Actor-critic / Transformer Value-based / DDQN Multi-agent actor-critic / DMARL PPO / Continuous-action PPO / actor-critic
DRL and heuristics Defender-only DDQN State-of-the-art RL methods Heuristics and Q-learning GUROBI, heuristics, and DRL Heuristic Heuristics and RL baselines Simulation and PPO Heuristics and RL baselines Heuristics GUROBI and RL baselines
Notes: SAC = Soft Actor-Critic, DRL = Deep Reinforcement Learning, DDQN = Double DQN, PPO = Proximal Policy Optimization, CNN = Convolutional Neural Network, DMARL = Distributed Multi-agent Reinforcement Learning.
In supply chains, Ma et al. (2025) modeled product-design change as a bilevel joint optimization problem and solve it end-to-end with a bilevel DRL method, benchmarking against conventional
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
19
DRL and heuristic baselines. In manufacturing, Yu et al. (2026) employed PPO with a CNNbased image state representation to address the dynamic job-shop scheduling problem. RL-based algorithmic frameworks are also applied to areas related to the public transport network design (e.g., Su and Yang 2025), charging and scheduling for electric robots or vehicles (e.g., Ding et al. 2025), healthcare (e.g., Vanvuchelen et al. 2024, Meng et al. 2025), logistics (e.g., Tian et al. 2025), public transport timetabling (e.g., Wang et al. 2025), urban mobility planning (e.g., Teusch et al. 2025), inventory optimization (e.g., Teck et al. 2025). In summary, across various domains, actor-critic and PPO variants are commonly used for tasks involving continuous or mixed control, such as charging robots and inventory management. In contrast, value-based methods like DDQN are typically applied to discrete design and location decisions. Benchmark methods vary considerably across studies, including heuristics, simulators, GUROBI, and RL baselines, highlighting the need for standardized evaluation protocols. Several studies benchmark primarily against other RL methods; incorporating OR baselines, such as ORTools, CP-SAT, or domain-specific heuristics, would strengthen the evidence for the effectiveness of the proposed RL-based methods. 5.1.3.
Discussion & future research directions When RL is used as an end-to-end, inde-
pendent solution method for classic and domain-specific COPs, three commonalities emerge. First, routing problems typically use policy-gradient decoders with Pointer or Transformer architectures, while general graph problems such as the Minimum Vertex Cover Problem and Max-Cut Problem often pair value-based methods with GNN encoders. Domains with continuous or mixed control frequently adopt actor-critic or PPO. Second, in routing studies, larger instance sizes are usually addressed with Transformer architectures rather than Pointer networks. Third, benchmark solution methods remain heterogeneous across studies, mixing classical heuristics, simulators, commercial solvers, and RL baselines, which complicates fair comparison. Despite rapid progress, several challenges remain. First, RL-based standalone solvers often struggle with scalability and generalization, particularly for large-scale instances (e.g., thousands of nodes) and typically require retraining when scale, topology, costs, or constraints change. Second, training times grow quickly with instance size. Third, most studies focus on simplified problems, with real-world constraints and uncertainties often overlooked. Fourth, inconsistent benchmarking and the lack of standardized evaluation protocols hinder comparability. The limited scalability of the RL methods arises mainly from two issues. First, typical encoders rely on O(n2 ) attention or high-order message passing, and decoders use step-by-step construction, both of which scale poorly and lead to long training times. Second, feasibility is rarely enforced as a hard constraint, and classical heuristics are not integrated into the generation step, making it difficult for the RL model
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
20
to learn stable and near-optimal construction rules. Additionally, training is usually performed on a fixed input distribution, preventing transfer across scales or objective coefficients and limiting the reuse of learned policies or value functions. To address these limitations and challenges, future research could (i) develop more efficient encoding and decoding schemes to reduce training time; (ii) combine RL with heuristic or exact algorithms to improve computational efficiency; (iii) address problems that better reflect real-world complexity; (iv) enhance generalization by employing dataset randomization and systematically evaluating transfer performance across different distributions; and (v) establish unified evaluation benchmarks that incorporate robust OR tools and problem-specific heuristics. 5.2.
RL-enhanced heuristic algorithms
In this subsection, we analyze algorithms that integrate RL with heuristics. We then outline limitations, challenges, and potential directions for future research. 5.2.1.
Literature analysis In this stream of research, RL plays five main roles in algorithms
that combine RL with heuristics. Table 7 summarizes the literature where RL is incorporated into heuristic methods for solving COPs, outlining the role of RL, the RL technique applied, the heuristic component, and the OR problems considered. The studies can be grouped into five categories according to the role of RL: (i) RL-driven cooperative search; (ii) RL constructs initial solutions; (iii) RL selects the most promising heuristic; (iv) RL selects the most promising operators or neighborhood; and (v) RL helps to tune intensities/quantities online. (i) RL-driven cooperative search. In this algorithmic framework, multiple agents run distinct metaheuristics/local searches and cooperate via RL-guided information sharing. RL adapts agents’ behaviors (e.g., which peer/solution pattern to trust or exchange) so that the overall search escapes local minima and balances diversification/intensification. For example, Martin et al. (2016) propose a general multi-agent cooperative search framework where each agent runs a different metaheuristic/local-search configuration. Agents communicate asynchronously and adapt via a cooperation protocol based on RL and pattern matching: they identify good patterns and share them, while RL adjusts how agents exploit shared information. (ii) RL constructs initial solutions. In this research line, RL constructs an initial solution that is then refined by heuristics (e.g., Deudon et al. 2018, Brammer et al. 2022, Li et al. 2024a). RL thus provides high-quality starting points and reduces the workload of the post-improvement phase. For instance, Deudon et al. (2018) proposed an algorithm that integrates a policy gradientbased RL with a local search heuristic. A neural network (i.e., Pointer Network with Attention) is used to generate many initial solutions and select the best one for the TSP. Subsequently, a 2-opt local search heuristic is applied as a post-processing step to further refine the solution and reduce
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
Table 7
21
Classification of papers studying the combinations of RL with heuristic algorithms for solving COPs. Publications Detaied roles of RL Martin et al. (2016)
Benlic et al. (2017) Deudon et al. (2018) Mosadegh et al. (2020) Lamghari and Dimitrakopoulos (2020) Ma et al. (2021) De Meijer and Sotirov (2021) Alicastro et al. (2021) Brammer et al. (2022) Zhang et al. (2022) Li et al. (2022) Kallestad et al. (2023) Zhang et al. (2023) Karimi-Mamaghan et al. (2023) Li et al. (2024a) Li et al. (2024b) Wu et al. (2024) Lu et al. (2024) Cui and Yuan (2024) Zou et al. (2024) Zhang et al. (2025) Rolim et al. (2025) Zhao and Hifi (2025) An et al. (2026)
Coordinate multi-agent cooperative search Help selecting the perturbation Help generating initial solutions RL selects heuristics Selects low-level heuristics Help selecting the neighborhood Constructs initial solutions (SDP-guided rounding) Help selecting the neighborhood Help generating initial solutions RL selects parameterised low-level heuristics Help selecting the neighborhood Help selecting operators Help selecting operators Help selecting perturbation operatos and strength Guides construction of initial solutions and steer subsequent search Help selecting operators Guide TS to focus on promising add/drop moves Help selecting the neighborhood Help selecting operators Help selecting the neighborhood Hyper-heuristic controller that adjusts the number of non-dominated solutions Help selecting the neighborhood Help selecting the neighborhood Help selecting the neighborhood
RL method
OR algorithm
RL-driven cooperation
Heuristics
PFSP, CVRP
MAB Pointer Network Tabular Q-learning Bandit (choice-function) DQN
VSP TSP SMMALSP SMPSP DPDP
Probability learning PPO DQN
BLS LS SA 27 heuristics LS SDP-based rounding LS LS Heuristics with rules TS ALNS ALNS
AGAP COPs SFTRP
Tabular Q-learning
IG
PFSP
Tabular Q-learning
LS
CDAP
Q-learning Learning-automata (probability-matrix) Probability learning Tabular Q-learning Tabular Q-learning
GA
FAP Clustered Orienteering Problem IUCP PSP LLRP Formation and Scheduling Optimization PBSP k-CMBCP CBSP
Tabular Q-learning Tabular Q-learning PPO DDQN
TS LS GA VND
Tabular Q-learning
GA
Q-learning Probability learning MAB
SLS, ALNS VNS VNS
OR problem
QCCP MSP PFSP COPs
Notes: PFSP = Permutation Flowshop Scheduling Problem, CVRP = Capacitied Vehicle Routing Problem, MAB = Multi-Armed Bandit, BLS = Breakout Local Search, VSP = Vertex Separator Problem, LS = Local Search, TSP = Traveling Salesman Problem, SA = Simulated Annealing, SMMALSP = Stochastic Mixed-Model Assembly Line Sequencing Problem, SMPSP = Stochastic open-pit Mine Production Scheduling Problem, DQN = Deep Q-Network, DPDP = Dynamic Pickup and Delivery Problem, SDP = Semidefinite Programming, QCCP = Quadratic Cycle Cover Problem, MSP = Machine Scheduling Problem, PPO = Proximal Policy Optimization, DDQN = Double Deep Q Network, TS = Tabu Search, AGAP = Airport Gate Assignment Problem, ALNS = Adaptive Large Neighborhood Search, SFTRP = Synchromodal Freight Transport Re-planning Problem, IG = Iterated Greedy, CDAP = Cross-dock Door Assignment Problem, IUCP = Maximum Independent Union of Cliques Problem, GA = Genetic Algorithm, PSP = Production Scheduling Problem, VND = Variable Neighborhood Descent, LLRP = Latency Location Routing Problem, SLS = Stochastic Local Search, VNS = Variable Neighborhood Search, k-CMBCP = k-clustering Minimum Biclique Completion Problem, CBSP = Customized Bus Scheduling Problem.
the tour length. To solve the Cross-dock Door Assignment Problem, Li et al. (2024a) proposed a hybrid algorithm combining RL and a local search algorithm: a Q-learning agent constructs an initial solution, which is then refined through local search in an iterative loop of “Q-learning-based construction” to “improvement by local search” to “reward update”. (iii) RL selects the most promising heuristic. RL can be used to select the most promising heuristics (see, Mosadegh et al. 2020, Lamghari and Dimitrakopoulos 2020). Mosadegh et al. (2020) proposed a Hyper Simulated Annealing framework, where a tabular Q-learning algorithm is embedded to select the most suitable heuristic online during the search process. Here, RL acts as
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
22
a controller that dynamically chooses which heuristic to apply at each iteration based on learned Q-values. (iv) RL selects the most promising operators or neighborhood. In this type of algorithm, RL helps a heuristic select the the most promising operator or neighborhood in each iteration (e.g., Benlic et al. 2017, Ma et al. 2021, Alicastro et al. 2021, Li and Ni 2022, Kallestad et al. 2023, Zhang et al. 2023, Li et al. 2024b, Wu et al. 2024, Cui and Yuan 2024, Zou et al. 2024, Lu et al. 2024, Rolim et al. 2025, Zhao and Hifi 2025, An et al. 2026). Specifically, to solve the Dynamic Pickup and Delivery Problem (DPDP), Ma et al. (2021) employed RL as an intelligent scheduler for a local search heuristic, using the Deep Q-Network to dynamically select the most promising local perturbation at each decision point. This enhances global search capability and improves convergence quality. For COPs in general, Kallestad et al. (2023) developed a selection hyperheuristic framework that integrates DRL into the Adaptive Large Neighborhood Search (ALNS) algorithm, where DRL replaces ALNS’s adaptive weights and is used to choose which destroy/repair operator, or an extra deterministic heuristic, to apply at each iteration. For the Clustered Orienteering Problem, Wu et al. (2024) designed an algorithm combining RL and tabu search (TS). TS serves as the main search engine, while RL updates a probability matrix that guides neighborhood evaluation, filtering out unpromising neighbors and reinforcing the search. (v) RL helps tuning intensities/quantities online. In line of work, RL tunes heuristic intensities or quantities online (e.g., perturbation strength, number of tasks removed, or parameters of low-level heuristics), allowing the search to modulate diversification and intensification on the fly (see, Zhang et al. 2022, Karimi-Mamaghan et al. 2023, Zhang et al. 2025). For instance, Zhang et al. (2022) proposed a DRL-based hyper-heuristic framework for COPs under uncertainty. The approach augments traditional hyper-heuristics with a data-driven heuristic selection module, where DRL is used to select parameter-controlled low-level heuristics to improve performance under uncertainty across diverse problem domains. For the Permutation Flowshop Scheduling Problem, Karimi-Mamaghan et al. (2023) embedded Q-learning into an Iterated Greedy (IG) algorithm, where the perturbation operators and their intensities (i.e., the number of jobs removed and reinserted per iteration) are selected online during the search process. In this framework, RL does not directly construct solutions but guides the perturbation phase of IG. The reward function combines local improvements with progress toward the global best solution. Zhang et al. (2025) proposed a multi-objective cooperative co-evolution algorithm combining RL and a genetic algorithm to tackle planning in a hybrid seru system (i.e., a production mode combining seru cells with a flow line). A hypervolume-based Q-learning agent serves as a hyper-heuristic controller that adaptively adjusts the number of cooperative non-dominated solutions participating in coevolution at each period.
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
5.2.2.
23
Discussion & future research directions Across the five research streams, the RL
selects the most promising operators or neighborhood category currently dominates. These methods are relatively easy to integrate with existing metaheuristics (e.g., VNS/ALNS/GA/LS/TS) and often yield noticeable improvements in computational efficiency and/or solution quality. By contrast, the RL-driven cooperative search, RL constructs initial solutions, and RL selects the most promising heuristic streams are less common in the literature, as they require deeper integration and are more sensitive to reward shaping and feasibility safeguards. The RL helps tuning intensities/quantities online stream is gaining traction, yet it remains underexplored beyond small and narrowly defined parameter sets. Specifically, only a limited number of studies have explored the RL-driven cooperative search method, in which each agent runs a metaheuristic. Its effectiveness has only been demonstrated on the Permutation Flowshop Scheduling Problem (PFSP) and the Capacitated Vehicle Routing Problem (CVRP). Its limitations are largely due to the strategy of aggregating locally promising solution fragments from each agent. This approach is less suitable for problems that involve long decision sequences or complex route configurations. Moreover, using only a few agents results in marginal performance gains over single-agent methods, while employing many agents considerably increases communication and computational costs. To address these challenges, future research could explore integrating a broader range of heuristics and designing pattern-discovery mechanisms that coordinate more effectively across different algorithms. It would also be valuable to investigate strategies for extracting and filtering higher-level structural features from candidate solutions, rather than relying solely on local fragments. Finally, developing more efficient concurrent implementations and communication strategies, along with systematic evaluations of how solution quality scales with the number of agents, would contribute to improving scalability and overall performance. In the domain where RL constructs initial solutions that are then refined by heuristics, current research is largely limited to relatively simple problems, such as single-vehicle orienteering (Li et al. 2024a). Future research could extend these methods to larger scales and more complex variants, including multi-vehicle orienteering with capacity and time-window constraints. Within the limited literature on RL for heuristic selection and for online tuning of intensities or quantities, existing approaches are mostly problem-specific, and their generalizability to broader families of COPs remains uncertain. Future research should establish cross-domain benchmarks, evaulate transferability accross instance sizes and input distributions, and compare RL-augmented methods against non-RL heuristics under consistent computational budgets. In the field of using RL to select the most promising operators or neighborhoods, several recurring limitations have been identified. One commonly cited problem is the lack of quantifiable optimality
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
24
gaps for the solutions obtained. However, this is a limitation of heuristic methods rather than of RL itself. Another frequently noted concern is that RL training times can be relatively long and may even exceed the time required to solve the problem using heuristics alone. Nevertheless, incorporating RL typically yields higher-quality solutions. Many existing studies also highlight that RL-based approaches for operator or neighborhood selection often lack generalization capability. When applied to new problem classes, these methods usually require retraining, and their performance cannot be guaranteed. This challenge stems from the nature of heuristics themselves, as effective operators and neighborhood structures are generally problem-specific and do not transfer reliably across different COP families. In summary, while RL-enhanced heuristic algorithms often deliver higher solution quality compared with heuristics alone, they also inherit several fundamental limitations from heuristic frameworks. Optimality guarantees remain elusive, generalization to unseen instances or new problem classes is often limited, and retraining is usually required when problem characteristics such as scale, distribution, or constraint structure change. Importantly, these shortcomings are largely attributable to the heuristic backbone rather than to RL itself. An interesting future research direction is to investigate whether RL can help mitigate these long-standing limitations, particularly by improving the generalization behavior of heuristic methods and enabling the creation of more transferable operator- or neighborhood-selection strategies. 5.3.
RL-enhanced exact algorithms
We summarize and analyze the literature that integrates RL with various types of exact algorithms and then discuss key challenges and future research directions. 5.3.1.
Literature analysis A wide range of exact algorithms have been developed in the
OR field, including the Alternating Direction Method of Multipliers (ADMM), Branch and Bound (B&B), Branch and Cut, Branch and Price, Benders Decomposition, Constrained (or Relaxed) Decision Diagrams (DD), and Constraint Programming (CP) to tackle combinatorial optimization problems. Commercial solvers such as GUROBI and CPLEX include implementations of many of these methods. Recent studies have explored integrating RL with exact algorithms or commercial solvers to improve computational efficiency and solution quality. Unlike traditional RL approaches that directly generate solutions for COPs, this research stream focuses on using RL to steer, tune, or augment exact algorithmic frameworks. We classify the relevant literature based on the type of exact algorithm integrated with RL and summarize representative studies in Table 8. (i) RL integrated with commercial solvers. Several studies combine RL with commercial solvers, where the solver handles subproblems exactly while RL focuses on high-level policy decisions (e.g., Yan et al. 2023, Liu et al. 2025, Harsha et al. 2025, Chohlas-Wood et al. 2025). Yan et al.
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
Table 8
25
Classification of papers studying the combinations of RL with exact algorithms for solving COPs.
Publications Roles of RL
OR algorithm
Cappart et al. (2019) Learn DD ordering
DD bounding Cutting-plane, Tang et al. (2020) Select Gomory cuts Branch-and-Cut Ichnowski et al. (2021) Tune ADMM parameters ADMM Cappart et al. (2021) Learn CP branching CP Cappart et al. (2022) Learn DD ordering DD bounding and B&B Yan et al. (2023) Learn SARSA(∆) policy CPLEX Tassel et al. (2023) Learn dispatching strategy CP Li et al. (2024a) VFA look-ahead Branch and Cut GUROBI Liu et al. (2025) Generate solver hints Harsha et al. (2025) MILP-optimized actor CPLEX (B&B) Open-source solver Chohlas-Wood et al. (2025) Bandit + LP policy
OR problem Max-Cut and MIS General IP classes QP General COPs MIS CODP Job-Shop Scheduling DOP-rd 0–1 IP Inventory Replenishment Resource Allocation
Notes: DDs = Decision Diagrams, MIS = Maximum Independent Set, IP = Integer Programming, ADMM = Alternating Direction Method of Multipliers, QP = Quadratic Programming, CP = Constraint Programming, SARSA = State–Action–Reward–State–Action, CODP = Charging and Order Dispatch Problem, VFA = Value Function Approximation, DOP-rd = Orienteering Problem with Stochastic and Dynamic Release Dates, B&B = Branch and Bound, LP = Linear Programming.
(2023) proposed a model-based RL framework for the charging and order dispatching problem in EV-based ride-hailing systems. At each decision epoch, a sample average approximation model is solved using CPLEX to evaluate short-term plans, which are then embedded into a SARSA(∆) policy for long-term control. Liu et al. (2025) developed a multi-agent RL framework for the fleet relocation problem of shared autonomous electric vehicles, where agents in the RL algorithm generate relocation preferences fed into a 0-1 integer programming model solved by GUROBI. This approach tightly integrates RL-driven decision guidance with exact optimization. Harsha et al. (2025) introduced a solution framework that combines deep policy iteration with mathematical programming, using neural networks to approximate the value function while exact optimization determines the actions. (ii) RL with Decision Diagram-based exact algroithms. Decision Diagrams (DDs), such as relaxed or restricted DDs, are graph-based structures that compactly encode feasible regions or objective bounds. The quality of DD-based bounds is highly sensitive to the variable ordering. Cappart et al. (2019) formulated DD construction as a sequential decision process and used Q-learning to learn variable orderings that produce tighter bounds for Max-Cut and Maximum Independent Set problems. The agent in the Q-learning algorithm selects the next variable based on the current DD state, with rewards linked to bound improvement. Building on this, Cappart et al. (2022) applied deep RL to learn effective variable orderings that improve both primal and dual bounds. This framework was further integrated into a full-fledged branch-and-bound algorithm, and the results showed that optimization bounds can be significantly enhanced through the use of deep RL algorithm.
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
26
(iii) RL with the Alternating Direction Method of Multipliers. ADMM is widely used for convex optimization problems such as quadratic programming (QP). Ichnowski et al. (2021) introduced RLQP, which uses deep RL to compute a policy that adapts the internal parameters of a QP solver to speed up the computation. Their framework significantly accelerates convergence compared to traditional fixed-parameter baselines. Here, RL acts as a meta-controller while the solver handles exact computations. (iv) RL with Branch and Cut and cutting-plane algroithms. Branch and Cut methods integrate branch-and-bound search with cutting-plane generation, where selecting effective cuts (e.g., Gomory cuts) is key to performance. A growing body of literature has explored how to enhance the Branch and Cut algorithm by controlling cutting-plane selection through RL, see, Tang et al. (2020), Li et al. (2024c). Tang et al. (2020) proposed a deep RL policy that selects Gomory cuts during cutting-plane iterations. Embedding this policy into Branch and Cut resulted in fewer nodes and faster convergence across several integer programming instances. The RL component is trained using an attention-based representation and evolution strategies. (v) RL with Constraint Programming algorithms. In this research stream, RL is typically used to learn the branching strategy in Constraint Programming (CP) (see, Cappart et al. 2021, Tassel et al. 2023). For instance, Cappart et al. (2021) formulated COPs in a unified dynamic programming-based framework and trained policies with Deep Q-Network or Proximal Policy Optimization methods to guide branching decisions, which were then embedded into a CP solver. 5.3.2.
Discussion & future research directions The literature summarized in Table 8 indi-
cates that RL–exact algorithm integration is particularly effective for large-scale problems where decisions such as branching, variable ordering, or cut selection strongly influence convergence but are difficult to manage through fixed, analytically derived rules. Three settings are especially well suited to RL integration: (i) when search, cutting, or sequencing decisions play a central role in the computational efficiency of the exact algorithms, such as branching in B&B, cut selection in Branch and Cut, or variable ordering in DDs; (ii) when parameter tuning in exact algorithms is challenging, for example, the update rules in ADMM that affect convergence speed, RL can be used to learn adaptive tuning strategies that enhance algorithmic performance; (iii) when a hierarchical frameworks are adopted, where the high-level RL component generates decisions, guidance, or priorities, and the low-level optimization module such as MILP, CP, or QP performs the final optimization. Such structures are common in transportation planning, scheduling, and fleet management. Within these integrated frameworks, RL typically plays one of the following four roles. First, RL generates high-level decisions (such as preferences or priorities), that are then passed to commercial solvers for refinement. Second, RL provides value functions or policies to guide look-ahead decisions, which are subsequently optimized by exact algorithms such as Branch-and-Cut, CP, or QP.
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
27
Third, RL learns algorithmic strategies such as branching, pruning, or cut selection. For example, guiding CP branching or Gomory cut selection to improve convergence behavior. Fourth, RL tunes parameters of exact algorithms, such as dynamically adjusting ADMM parameters to reduce the number of iterations and computation time. Despite promising results, several limitations remain. First, learned branching, cutting, or sequencing strategies often generalize poorly. Their effectiveness tends to diminish or even become counterproductive when applied to problems with different sizes, structures, or distributions. Second, commercial solvers already include sophisticated heuristic and presolve routines for branching, ordering, and cut selection strategies. Achieving meaningful improvements using RL typically requires substantial training effort. Third, explainability remains a considerable challenge, as the rationale behind learned strategies is difficult to interpret and validate. Fourth, the problems tested so far are relatively simple. More realistic and operationally relevant scenarios, such as multi-hub systems, multiple vehicles, batch arrivals, partial charging, or uncertainty in service and charging times, have not yet been thoroughly addressed. Future research should therefore aim to enhance both the generalizability and explainability of RL-enhanced exact algorithms, particularly in frameworks that leverage the complementary strengths of learning and optimization. In parallel, more complex and practical problem settings should be incorporated to test the robustness and scalability of these approaches, such as the applications in the multi-hub location problem, partial charging problem, and integrated optimization problem of routing and charging.
6.
RL: Facilitator for extended reality analysis
Digital twins (DTs) map physical systems into real-time digital representations, thereby enhancing decision-making through simulation, monitoring, and optimization. Research on integrating RL with DTs is growing rapidly, particularly for extending existing systems or designing new ones in dynamic and data-rich environments. In Section 6.1, we analyze the state-of-the-art advancements at this intersection, and in Section 6.2, we discuss the associated challenges and future research directions. 6.1.
Literature analysis
Table 9 presents a classification of the literature integrating RL with DTs. As shown, the integration has attracted attention across a wide range of domains, such as scheduling in air-ground and manufacturing systems, traffic signal control, autonomous driving, and resource allocation. Relevant studies can be broadly categorized into three streams: (i) DTs serving as virtual training environments for RL; (ii) DTs accelerating RL by reducing required real-world interactions; and (iii) RL optimizing the operations of DTs themselves.
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
28
Table 9 Publications Sun et al. (2022) Yan et al. (2022) Liu et al. (2023) Tang et al. (2023) Kamal et al. (2024) Schlappa et al. (2024) Wu et al. (2021) Zhang et al. (2024) Wang et al. (2024) Park et al. (2022) Xu et al. (2024)
Classification of papers studying the integration of RL and DTs.
Domain Scheduling in air-ground networks Scheduling in manufacturing systems Function virtualization migration Task assignment Traffic signal control Control of waste incineration plants Autonomous driving Resource management Resource allocation Production control Mapping mechanism
Roles of DT and RL RL uses DTs as training environments RL uses DTs as training environments RL uses DTs as training environments RL uses DTs as training environments RL uses DTs as training environments RL uses DTs as training environments DTs accelerate RL DTs accelerate RL RL optimizes operations of DTs RL optimizes operations of DTs RL optimizes operations of DTs
RL method DIFL Double-layer Q-learning DPPO Deep Q-learning MADDPG Deep Q-network Actor-critic Deep Q-network MADDPG Tabular Q-learning PPO
Notes: DIFL = Dynamic Incentive for Federated Learning, MADDPG = Multiagent Deep Deterministic Policy Gradient, PPO = Proximal Policy Optimization, DPPO = Deep RL based on the Distributed Proximal Policy Optimization.
(i) DTs as virtual training environments for RL. In the first stream, DTs replicate physical systems or processes to create realistic virtual environments for the training of RL algorithms. This stream leverages the strong mapping capabilities of DTs to represent physical entities digitally, enabling extensive training without incurring the costs, risks, or time associated with experimentation in real-world settings. Representative applications include dynamic scheduling in air–ground integrated networks (Sun et al. 2022), adaptive scheduling in manufacturing systems (Yan et al. 2022), task assignment in multi-unmanned aerial vehicle systems (Tang et al. 2023), and traffic signal control for CO2 emission reduction (Kamal et al. 2024). Among representative studies, Liu et al. (2023) designed a DT-enabled network that captures the real-time dynamics of an IoT environment to optimize network function virtualization migration, where a deep RL algorithm is trained within the DT to make energy-efficient migration decisions. Similarly, Schlappa et al. (2024) proposed a DT-based framework for the optimal control of waste incineration plants, in which a data-driven DT serves as the learning environment for RL-based operational policies. (ii) DTs accelerate RL. The second stream focuses on using DTs to accelerate RL. Although research in this area remains limited, a common approach is to use DTs to approximate the dynamics of the real-world environment, thereby reducing the need for costly online interactions. In the context of autonomous driving, Wu et al. (2021) proposed a DT-enabled RL framework in which the DT is used to model the transition dynamics of the physical driving scenarios. This predictive environment enables the RL agent to train more efficiently by reducing the number of interactions with the real system. In the domain of network slicing for resource management, Zhang et al. (2024) designed a DT-enhanced deep RL framework. The DT is constructed using historical data to learn both the transition dynamics and the reward function of the real environment. By enabling RL agents to interact with this predictive virtual space rather than the physical system, the framework significantly reduces training costs and improves the generalization of learned policies.
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
29
(iii) RL optimizes DT operations. Lastly, we present a representative line of work from the third research stream, where RL is employed to optimize the operations of DTs. In this stream, RL is integrated into the management of DT synchronization, fidelity, and update efficiency. These studies exemplify how RL can enhance the operational performance of large-scale DT systems operating in complex and time-varying environments. For instance, in the field of dynamic platoon digital twin networks, Wang et al. (2024) investigated the problem of resource allocation, where the DT itself is the optimization target. The authors formulate a high-order Markov decision process that captures the update dynamics of DTs across multiple vehicles and propose a decentralized multi-agent deep deterministic policy gradient algorithm. In the context of production control in a re-entrant job shop, Park et al. (2022) designed a control framework that integrates DTs and horizontal coordination with RL-based production control to improve manufacturing efficiency and responsiveness. In the energy domain, Xu et al. (2024) proposed a DRL-based, data-driven mapping mechanism for Internet of Energy systems. By formulating the DT construction process as a Markov decision process, they train an RL agent to optimize the mapping from physical sensor data to virtual states, minimizing the deviation between the DT and its physical counterpart. 6.2.
Challenges and future research directions
The integration of RL and DTs has emerged as a promising direction for enabling data-driven decision-making in complex, dynamic environments. Existing studies primarily focus on using DTs to support RL training and leveraging RL to enhance DT operations. These studies highlight the mutual reinforcement between DTs and RL, opening up new avenues for future research in intelligent systems across domains such as transportation, energy, and manufacturing. In this section, we first provide a guideline for researchers on when the integration of RL and DTs is most appropriate. We then discuss key technical challenges and conclude with several future research directions based on the research gaps identified in Table 9. 6.2.1.
When to use RL with DTs When DTs are used as virtual training environments for
RL, or to accelerate RL, the most suitable applications typically exhibit at least one of the following characteristics: (i) real-world exploration is risky, costly, or slow (e.g., waste-incineration plants, air–ground networks, autonomous driving); (ii) effective decision-making requires large volumes of safe rollouts to learn high-frequency control or dispatch policies,allowing RL agents to learn through trial and error within the DT; (iii) the system features frequent task dynamics and strict decision-time constraints; and/or (iv) robustness to rare events and extreme disturbances is critical, requiring exposure to a wide range of scenarios. When DTs are used specifically to accelerate RL, an additional motivation is to further reduce computational cost while improving generalization.
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
30
When optimizing the operations of DTs using RL (e.g., determining update frequency, simulation fidelity, mapping/assimilation plans, or bandwidth/computational resource allocation within datadriven pipelines), applicability varies. This approach is mainly suited for the following problems: (i) Real-time resource allocation within data-driven pipelines (communication, computation, storage); (ii) Data mapping or simulation operations under finite cost and latency constraints; (iii) Online adjustment of dataset calibration/assimilation windows. To employ DTs as RL training grounds or to accelerate RL, the DT must accurately reproduce key input–output behaviors at the relevant decision granularity and expose agents to stochasticity beyond a single nominal scenario. In practice, this requires curating multi-scenario datasets that randomize demand, failures, schedule changes, and exogenous signals, and partitioning them into training, validation, and test sets (e.g., by time or scenario) to prevent data leakage. For acceleration-focused applications, reward definitions in the DT should align with real-world operational objectives. When using RL to optimize DT operations, key requirements include: (i) exposing adjustable control knobs (e.g., update rates, fidelity levels, computational budgets, and assimilation windows); (ii) defining DT-specific performance metrics such as reconstruction/prediction error, trustworthiness, timeliness, and cost; and (iii) enabling observability of system states (e.g., workload, queue lengths, latency, and channel conditions). 6.2.2.
Technical challenges Despite growing interest, several key challenges persist across
the three application streams. (i) Many studies train and evaluate models entirely within DT environments, without systematically characterizing the sources, magnitude, or performance impact of discrepancies between DTs and their physical counterparts. In some cases, DTs are treated as theoretical constructs that overlook real-world deviations, whereas in others, simulation biases are acknowledged but not quantified, limiting the credibility of reported performance. (ii) Purely data-driven DTs demand large volumes of high-quality data and extensive hyperparameter tuning, while RL training itself entails significant computational overhead. Although offline training is feasible and can reduce real-world interaction costs, the overall computational burden remains high, constraining practical deployment. (iii) In multi-agent settings, the integration of edge computing and communication networks introduces latency and resource constraints that distort the observation–action feedback loop and degrade control performance. (iv) Transferability across network topologies, operating regimes, and equipment configurations is rarely assessed rigorously; several studies explicitly point to limited generalizability beyond the training domain. (v) Human-RL comparisons often suffer from evaluation bias due to inconsistent definitions of reward functions, KPIs, and telemetry metrics, making it difficult to conduct fair or reproducible assessments. A cross-cutting challenge in both using DTs to accelerate RL and using RL to optimize DT operations is the inherently interdisciplinary nature of these systems. Successful implementation
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
31
necessitates integrating expertise across multiple fields, including modeling, data engineering, DT development, and RL algorithm design. 6.2.3.
Future research directions When integrating RL with DTs, existing frameworks can
be extended to address the aforementioned challenges by incorporating the following aspects: (i) establishing systematic methods to quantify discrepancies between DTs and real-world systems; (ii) prioritizing offline and model-based RL to improve sample efficiency; (iii) explicitly modeling communication and computational latency, as well as bandwidth constraints, particularly in multi-agent scenarios; (iv) adopting risk-sensitive objectives to mitigate tail-end failures; and (v) leveraging inverse RL and preference learning to align with implicit human objectives, while standardizing reward definitions through offline policy evaluation to reduce comparison bias. For the stream in which RL is used to optimize DT operations, an additional promising direction could be developing multi-agent RL approaches that jointly consider communication, computation, fidelity, and assimilation trade-offs. These may involve adjusting update rates, selecting appropriate simulation granularity, and managing resource budgets. Another interesting direction is to implement risk-aware control strategies that account for DT-centric performance metrics while respecting resource and security constraints. Finally, methods that maximize information gain, such as actively scheduling data mapping and assimilation, represent an additional avenue for advancing DT–RL integration.
7.
Conclusion
This paper has presented a unified taxonomy for how RL empowers OR and provided a comprehensive and technical review of the integration between the two fields. We structured the literature around three core roles of RL within OR. First, RL serves as a direct solution method for sequential decision-making problems. Second, RL functions as an end-to-end solver or as an enhancement to heuristic and exact methods for COPs. Third, RL can be combined with digital twins to support planning, learning, and operational optimization. These roles were aligned with a conceptual taxonomy and a practical decision map to guide method selection. Our analysis highlights the settings where RL provides the greatest benefits. These gains are most evident in high-dimensional, uncertain, or non-stationary environments, and in applications that require frequent or real-time decisions under tight operational constraints. In such settings, RL can enhance exact algorithms by learning branching, cutting, variable-ordering, or parameter tuning strategies, and it can strengthen heuristic methods by identifying promising operators, neighborhood structures, and search intensities. When integrated with digital twins, RL further facilitates safe and scalable experimentation, although effective deployment depends critically on rigorous calibration to reduce gaps between simulated and physical systems.
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
32
Overall, RL and OR are strongly complementary. Their integration provides a promising pathway to more efficient and higher-quality decision-making in complex systems and opens broad opportunities for rigorous research and impactful applications. Nevertheless, several important challenges remain. RL’s low sample efficiency and limited training stability restrict its broad applicability. Generalisation and transfer across problem sizes, topologies, and distributions remain weak. Explanations and theoretical guarantees for learned strategies are often lacking. Benchmarks and evaluation protocols remain fragmented. Digital twins require reliable data and careful model validation, while multi-agent implementations introduce latency and bandwidth constraints. To address these challenges and advance the field, we outline several priorities. First, establishing standardized benchmarks and unified evaluation protocols is essential for fair comparison and reproducibility. Second, future research should develop hybrid frameworks that combine model-free and model-based RL algorithms to solve sequential decision-making problems with high-quality solutions and strong computational efficiency. Third, efforts should focus on strengthening the generalizability and explainability of RL-enhanced exact algorithms, with the long-term goal of building unified frameworks that coordinate learning and optimization for large-scale real-time operations. Fourth, in the context of digital twins, prioritizing offline and model-based RL offers a promising direction for improving sample efficiency and scalability.
References Ahamed, T., Zou, B., Farazi, N.P., Tulabandhula, T., 2021. Deep reinforcement learning for crowdsourced urban delivery. Transportation Research Part B: Methodological 152, 227–257. doi:10.1016/j.trb. 2021.08.015. Alfonso-Sánchez, S., Solano, J., Correa-Bahnsen, A., Sendova, K.P., Bravo, C., 2024. Optimizing credit limit adjustments under adversarial goals using reinforcement learning. European Journal of Operational Research 315, 802–817. doi:10.1016/j.ejor.2023.12.025. Alicastro, M., Ferone, D., Festa, P., Fugaro, S., Pastore, T., 2021. A reinforcement learning iterated local search for makespan minimization in additive manufacturing machine scheduling problems. Computers & Operations Research 131, 105272. doi:10.1016/j.cor.2021.105272. An, X., Li, X., Zhang, B., 2026.
Flexible scheduling of customized bus for green mega-events: A
distributionally robust optimization approach.
Computers & Operations Research 185, 107249.
doi:10.1016/j.cor.2025.107249. Arulkumaran, K., Deisenroth, M.P., Brundage, M., Bharath, A.A., 2017. Deep reinforcement learning: A brief survey. IEEE Signal Processing Magazine 34, 26–38. doi:10.1109/MSP.2017.2743240. Barrett, T., Clements, W., Foerster, J., Lvovsky, A., 2020. Exploratory combinatorial optimization with reinforcement learning. Proceedings of the AAAI Conference on Artificial Intelligence 34, 3243–3250. doi:10.1609/aaai.v34i04.5723.
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
33
Bello, I., Pham, H., Le, Q.V., Norouzi, M., Bengio, S., 2017. Neural combinatorial optimization with reinforcement learning. URL: https://arxiv.org/abs/1611.09940, arXiv:1611.09940. Bengio, Y., Lodi, A., Prouvost, A., 2021. Machine learning for combinatorial optimization: A methodological tour d’horizon. European Journal of Operational Research 290, 405–421. doi:10.1016/j.ejor.2020. 07.063. Benlic, U., Epitropakis, M.G., Burke, E.K., 2017. A hybrid breakout local search and reinforcement learning approach to the vertex separator problem. European Journal of Operational Research 261, 803–818. doi:10.1016/j.ejor.2017.01.023. Brammer, J., Lutz, B., Neumann, D., 2022. Permutation flow shop scheduling with multiple lines and demand plans using reinforcement learning. European Journal of Operational Research 299, 75–86. doi:10.1016/j.ejor.2021.08.007. Buckman, J., Hafner, D., Tucker, G., Brevdo, E., Lee, H., 2019. Sample-efficient reinforcement learning with stochastic ensemble value expansion. URL: https://arxiv.org/abs/1807.01675, arXiv:1807.01675. Buşoniu, L., Babuška, R., De Schutter, B., 2010. Multi-agent reinforcement learning: An overview. Springer Berlin Heidelberg, Berlin, Heidelberg. pp. 183–221. doi:10.1007/978-3-642-14435-6_7. Cappart, Q., Bergman, D., Rousseau, L.M., Prémont-Schwarz, I., Parjadis, A., 2022. Improving variable orderings of approximate decision diagrams using reinforcement learning. INFORMS Journal on Computing 34, 2552–2570. doi:10.1287/ijoc.2022.1194. Cappart, Q., Goutierre, E., Bergman, D., Rousseau, L.M., 2019. Improving optimization bounds using machine learning: Decision diagrams meet deep reinforcement learning. Proceedings of the AAAI Conference on Artificial Intelligence 33, 1443–1451. doi:10.1609/aaai.v33i01.33011443. Cappart, Q., Moisan, T., Rousseau, L.M., Prémont-Schwarz, I., Cire, A.A., 2021. Combining reinforcement learning and constraint programming for combinatorial optimization. Proceedings of the AAAI Conference on Artificial Intelligence 35, 3677–3687. doi:10.1609/aaai.v35i5.16484. Chen, X., Ulmer, M.W., Thomas, B.W., 2022. Deep q-learning for same-day delivery with vehicles and drones. European Journal of Operational Research 298, 939–952. doi:10.1016/j.ejor.2021.06.021. Chen, X., Wang, T., Thomas, B.W., Ulmer, M.W., 2023. Same-day delivery with fair customer service. European Journal of Operational Research 308, 738–751. doi:10.1016/j.ejor.2022.12.009. Chohlas-Wood, A., Coots, M., Zhu, H., Brunskill, E., Goel, S., 2025. Learning to be fair: A consequentialist approach to equitable decision making. Management Science doi:10.1287/mnsc.2022.00345. Clavera, I., Rothfuss, J., Schulman, J., Fujita, Y., Asfour, T., Abbeel, P., 2018.
Model-based rein-
forcement learning via meta-policy optimization, in: Billard, A., Dragan, A., Peters, J., Morimoto, J. (Eds.), Proceedings of The 2nd Conference on Robot Learning, PMLR. pp. 617–629. https://proceedings.mlr.press/v87/clavera18a.html.
URL:
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
34
Cui, W., Yuan, B., 2024. A hybrid genetic algorithm based on reinforcement learning for the energy-aware production scheduling in the photovoltaic glass industry. Computers & Operations Research 163, 106521. doi:10.1016/j.cor.2023.106521. Dai, H., Khalil, E., Zhang, Y., Dilkina, B., Song, L., 2017. Learning combinatorial optimization algorithms over graphs, in: Guyon, I., Luxburg, U.V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., Garnett, R. (Eds.), Advances in Neural Information Processing Systems, Curran Associates, Inc.
URL: https://proceedings.neurips.cc/paper_files/paper/2017/file/
d9896106ca98d3d05b8cbdf4fd8b13a1-Paper.pdf. De Meijer, F., Sotirov, R., 2021. Sdp-based bounds for the quadratic cycle cover problem via cutting-plane augmented lagrangian methods and reinforcement learning. INFORMS Journal on Computing 33, 1262–1276. doi:10.1287/ijoc.2021.1075. Dehaybe, H., Catanzaro, D., Chevalier, P., 2024. Deep reinforcement learning for inventory optimization with non-stationary uncertain demand. European Journal of Operational Research 314, 433–445. doi:10. 1016/j.ejor.2023.10.007. Deudon, M., Cournut, P., Lacoste, A., Adulyasak, Y., Rousseau, L.M., 2018. Learning heuristics for the tsp by policy gradient, in: van Hoeve, W.J. (Ed.), Integration of Constraint Programming, Artificial Intelligence, and Operations Research, Springer International Publishing, Cham. pp. 170–181. Ding, Y., Deng, M., Ke, G.Y., Shen, Y., Zhang, L., 2025. Scheduling intelligent charging robots for electric vehicle: A deep reinforcement learning approach. Transportation Research Part E: Logistics and Transportation Review 200, 104090. doi:10.1016/j.tre.2025.104090. Enders, T., Harrison, J., Pavone, M., Schiffer, M., 2023. Hybrid multi-agent deep reinforcement learning for autonomous mobility on demand systems, in: Matni, N., Morari, M., Pappas, G.J. (Eds.), Proceedings of The 5th Annual Learning for Dynamics and Control Conference, PMLR. pp. 1284–1296. URL: https://proceedings.mlr.press/v211/enders23a.html. Feinberg, V., Wan, A., Stoica, I., Jordan, M.I., Gonzalez, J.E., Levine, S., 2018. Model-based value estimation for efficient model-free reinforcement learning. URL: https://arxiv.org/abs/1803.00101, arXiv:1803.00101. Fujimoto, S., van Hoof, H., Meger, D., 2018. Addressing function approximation error in actor-critic methods. URL: https://arxiv.org/abs/1802.09477, arXiv:1802.09477. Garnier, P., Viquerat, J., Rabault, J., Larcher, A., Kuhnle, A., Hachem, E., 2021. A review on deep reinforcement learning for fluid mechanics. Computers & Fluids 225, 104973. doi:10.1016/j.compfluid. 2021.104973. Gu, S., Yang, L., Du, Y., Chen, G., Walter, F., Wang, J., Knoll, A., 2024. A review of safe reinforcement learning: Methods, theories, and applications. IEEE Transactions on Pattern Analysis and Machine Intelligence 46, 11216–11235. doi:10.1109/TPAMI.2024.3457538.
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
35
Guo, W., Atasoy, B., Negenborn, R.R., 2022. Global synchromodal shipment matching problem with dynamic and stochastic travel times: a reinforcement learning approach. Annals of Operations Research 350, 63–94. doi:10.1007/s10479-021-04489-z. Haarnoja, T., Zhou, A., Abbeel, P., Levine, S., 2018.
Soft actor-critic: Off-policy maximum entropy
deep reinforcement learning with a stochastic actor. URL: https://arxiv.org/abs/1801.01290, arXiv:1801.01290. Harsha, P., Jagmohan, A., Kalagnanam, J., Quanz, B., Singhvi, D., 2025. Deep policy iteration with integer programming for inventory management. Manufacturing & Service Operations Management 27, 369– 388. doi:10.1287/msom.2022.0617. Hu, K., Li, M., Song, Z., Xu, K., Xia, Q., Sun, N., Zhou, P., Xia, M., 2024. A review of research on reinforcement learning algorithms for multi-agents. Neurocomputing 599, 128068. doi:10.1016/j. neucom.2024.128068. Ichnowski, J., Jain, P., Stellato, B., Banjac, G., Luo, M., Borrelli, F., Gonzalez, J.E., Stoica, I., Goldberg, K., 2021. Accelerating quadratic optimization with reinforcement learning, in: Ranzato, M., Beygelzimer, A., Dauphin, Y., Liang, P., Vaughan, J.W. (Eds.), Advances in Neural Information Processing Systems, Curran Associates, Inc.. pp. 21043–21055. URL: https://proceedings.neurips.cc/paper_files/ paper/2021/file/afdec7005cc9f14302cd0474fd0f3c96-Paper.pdf. Janner, M., Fu, J., Zhang, M., Levine, S., 2019. When to trust your model: Model-based policy optimization, in: Wallach, H., Larochelle, H., Beygelzimer, A., d'Alché-Buc, F., Fox, E., Garnett, R. (Eds.), Advances in Neural Information Processing Systems, Curran Associates, Inc. URL: https://proceedings. neurips.cc/paper_files/paper/2019/file/5faf461eff3099671ad63c6f3f094f7f-Paper.pdf. Jayant, A.K., Bhatnagar, S., 2022.
Model-based safe deep reinforcement learning via a constrained
proximal policy optimization algorithm, in: Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., Oh, A. (Eds.), Advances in Neural Information Processing Systems, Curran Associates, Inc.. pp. 24432–24445. URL: https://proceedings.neurips.cc/paper_files/paper/2022/file/ 9a8eb202c060b7d81f5889631cbcd47e-Paper-Conference.pdf. Jin, J., Cui, T., Bai, R., Qu, R., 2024. Container port truck dispatching optimization using real2sim based deep reinforcement learning. European Journal of Operational Research 315, 161–175. doi:10.1016/ j.ejor.2023.11.038. Jin, Y., Ding, Y., Pan, X., He, K., Zhao, L., Qin, T., Song, L., Bian, J., 2023. Pointerformer: Deep reinforced multi-pointer transformer for the traveling salesman problem. Proceedings of the AAAI Conference on Artificial Intelligence 37, 8132–8140. doi:10.1609/aaai.v37i7.25982. Kaelbling, L.P., Littman, M.L., Moore, A.W., 1996. Reinforcement learning: A survey. Journal of artificial intelligence research 4, 237–285. doi:10.1613/jair.301.
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
36
Kallestad, J., Hasibi, R., Hemmati, A., Sörensen, K., 2023. A general deep reinforcement learning hyperheuristic framework for solving combinatorial optimization problems. European Journal of Operational Research 309, 446–468. doi:10.1016/j.ejor.2023.01.017. Kamal, H., Yánez, W., Hassan, S., Sobhy, D., 2024. Digital-twin-based deep reinforcement learning approach for adaptive traffic signal control. IEEE Internet of Things Journal 11, 21946–21953. doi:10.1109/ JIOT.2024.3377600. Karimi-Mamaghan, M., Mohammadi, M., Meyer, P., Karimi-Mamaghan, A.M., Talbi, E.G., 2022. Machine learning at the service of meta-heuristics for solving combinatorial optimization problems: A state-ofthe-art. European Journal of Operational Research 296, 393–422. doi:10.1016/j.ejor.2021.04.032. Karimi-Mamaghan, M., Mohammadi, M., Pasdeloup, B., Meyer, P., 2023. Learning to select operators in meta-heuristics: An integration of q-learning into the iterated greedy algorithm for the permutation flowshop scheduling problem. European Journal of Operational Research 304, 1296–1330. doi:10.1016/ j.ejor.2022.03.054. Kidambi, R., Rajeswaran, A., Netrapalli, P., Joachims, T., 2020. reinforcement
learning,
in:
Larochelle,
H.,
Ranzato,
M.,
Morel: Model-based offline
Hadsell,
R.,
Balcan,
M.,
Lin,
H. (Eds.), Advances in Neural Information Processing Systems, Curran Associates, Inc.. pp. 21810–21823.
URL: https://proceedings.neurips.cc/paper_files/paper/2020/file/
f7efa4f864ae9b88d43527f4b14f750f-Paper.pdf. Konda, V., Tsitsiklis, J., 1999. Actor-critic algorithms, in: Solla, S., Leen, T., Müller, K. (Eds.), Advances in Neural Information Processing Systems, MIT Press. URL: https://proceedings.neurips.cc/ paper_files/paper/1999/file/6449f44a102fde848669bdd9eb6b76fa-Paper.pdf. Kool, W., van Hoof, H., Welling, M., 2019. Attention, learn to solve routing problems!, in: International Conference on Learning Representations. URL: https://openreview.net/forum?id=ByxBFsRqYm. Ladosz, P., Weng, L., Kim, M., Oh, H., 2022. Exploration in deep reinforcement learning: A survey. Information Fusion 85, 1–22. doi:10.1016/j.inffus.2022.03.003. Lamghari, A., Dimitrakopoulos, R., 2020. Hyper-heuristic approaches for strategic mine planning under uncertainty. Computers & Operations Research 115, 104590. doi:10.1016/j.cor.2018.11.010. Lecarpentier, E., Rachelson, E., 2019. Non-stationary markov decision processes, a worst-case approach using model-based reinforcement learning, in: Wallach, H., Larochelle, H., Beygelzimer, A., d'AlchéBuc, F., Fox, E., Garnett, R. (Eds.), Advances in Neural Information Processing Systems, Curran Associates, Inc.
URL: https://proceedings.neurips.cc/paper_files/paper/2019/file/
859b00aec8885efc83d1541b52a1220d-Paper.pdf. Lee, H.R., Lee, T., 2021. Multi-agent reinforcement learning algorithm to solve a partially-observable multiagent problem in disaster response. European Journal of Operational Research 291, 296–308. doi:10. 1016/j.ejor.2020.09.018.
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
37
Levine, S., Kumar, A., Tucker, G., Fu, J., 2020. Offline reinforcement learning: Tutorial, review, and perspectives on open problems. arXiv:2005.01643. yu Li, G., Chow, A.H., Ying, C., 2025. Robust optimization for adaptive bus service scheduling with adversarial reinforcement learning under demand uncertainties. Transportation Research Part C: Emerging Technologies 178, 105222. doi:10.1016/j.trc.2025.105222. Li, K., Zhang, T., Wang, R., 2021. Deep reinforcement learning for multiobjective optimization. IEEE Transactions on Cybernetics 51, 3103–3114. doi:10.1109/TCYB.2020.2977661. Li, M., Hao, J.K., Wu, Q., 2022. Learning-driven feasible and infeasible tabu search for airport gate assignment. European Journal of Operational Research 302, 172–186. doi:10.1016/j.ejor.2021.12.019. Li, M., Hao, J.K., Wu, Q., 2024a. A flow based formulation and a reinforcement learning based strategic oscillation for cross-dock door assignment. European Journal of Operational Research 312, 473–492. doi:10.1016/j.ejor.2023.07.014. Li, W., Ni, S., 2022. Train timetabling with the general learning environment and multi-agent deep reinforcement learning. Transportation Research Part B: Methodological 157, 230–251. doi:10.1016/j. trb.2022.02.006. Li, X., An, X., Zhang, B., 2024b. Minimizing passenger waiting time in the multi-route bus fleet allocation problem through distributionally robust optimization and reinforcement learning. Computers & Operations Research 164, 106568. doi:10.1016/j.cor.2024.106568. Li, Y., 2018. Deep reinforcement learning. arXiv:1810.06339. Li, Y., Archetti, C., Ljubić, I., 2024c. Reinforcement learning approaches for the orienteering problem with stochastic and dynamic release dates. Transportation Science 58, 1143–1165. doi:10.1287/trsc.2022. 0366. Lillicrap, T.P., Hunt, J.J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., Wierstra, D., 2015. Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971 . Liu, C., Wang, Z., Liu, Z., Huang, K., 2025. Multi-agent reinforcement learning framework for addressing demand-supply imbalance of shared autonomous electric vehicle. Transportation Research Part E: Logistics and Transportation Review 197, 104062. doi:10.1016/j.tre.2025.104062. Liu, Q., Tang, L., Wu, T., Chen, Q., 2023. Deep reinforcement learning for resource demand prediction and virtual function network migration in digital twin network. IEEE Internet of Things Journal 10, 19102–19116. doi:10.1109/JIOT.2023.3281678. Liu, Y., Chen, Y., Jiang, T., 2020. Dynamic selective maintenance optimization for multi-state systems over a finite horizon: A deep reinforcement learning approach. European Journal of Operational Research 283, 166–181. doi:10.1016/j.ejor.2019.10.049. Lu, Z., Gao, J., Hao, J.K., Yang, P., Zhou, L., 2024. Learning driven three-phase search for the maximum independent union of cliques problem. Computers & Operations Research 164, 106549. doi:10.1016/ j.cor.2024.106549.
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
38
Luo, F.M., Xu, T., Lai, H., Chen, X.H., Zhang, W., Yu, Y., 2024. A survey on model-based reinforcement learning. Science China Information Sciences 67, 121101. doi:10.1007/s11432-022-3696-5. Ma, C., Li, A., Du, Y., Dong, H., Yang, Y., 2024. Efficient and scalable reinforcement learning for large-scale network control. Nature Machine Intelligence 6, 1006–1020. doi:10.1038/s42256-024-00879-7. Ma, Y., Hao, X., Hao, J., Lu, J., Liu, X., Xialiang, T., Yuan, M., Li, Z., Tang, J., Meng, Z., 2021.
A hierarchical reinforcement learning based optimization framework for large-scale
dynamic pickup and delivery problems, in: Ranzato, M., Beygelzimer, A., Dauphin, Y., Liang, P., Vaughan, J.W. (Eds.), Advances in Neural Information Processing Systems, Curran Associates, Inc.. pp. 23609–23620. URL: https://proceedings.neurips.cc/paper_files/paper/2021/file/ c6a01432c8138d46ba39957a8250e027-Paper.pdf. Ma, Y., Xia, X., Liu, P., Zhang, C., 2025. Bilevel joint optimization for product design changes with a resilient supply chain based on deep reinforcement learning. International Journal of Production Economics , 109791doi:10.1016/j.ijpe.2025.109791. Mao, W., Zhang, K., Zhu, R., Simchi-Levi, D., Başar, T., 2025. Model-free nonstationary reinforcement learning: Near-optimal regret and applications in multiagent reinforcement learning and inventory control. Management Science 71, 1564–1580. doi:10.1287/mnsc.2022.02533. Martin, S., Ouelhadj, D., Beullens, P., Ozcan, E., Juan, A.A., Burke, E.K., 2016. A multi-agent based cooperative approach to scheduling and routing. European Journal of Operational Research 254, 169– 178. doi:10.1016/j.ejor.2016.02.045. Meng, Q., Feng, B., Yu, G., 2025. Dynamic volunteer assignment: Integrating skill diversity, task variability and volunteer preferences. Transportation Research Part E: Logistics and Transportation Review 197, 104068. doi:10.1016/j.tre.2025.104068. Mnih, V., Badia, A.P., Mirza, M., Graves, A., Lillicrap, T.P., Harley, T., Silver, D., Kavukcuoglu, K., 2016. Asynchronous methods for deep reinforcement learning. arXiv:1602.01783. Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., Riedmiller, M., 2013. Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602 . Monaci, M., Agasucci, V., Grani, G., 2024. An actor-critic algorithm with policy gradients to solve the job shop scheduling problem using deep double recurrent agents. European Journal of Operational Research 312, 910–926. doi:10.1016/j.ejor.2023.07.037. Mosadegh, H., Fatemi Ghomi, S., Süer, G., 2020. Stochastic mixed-model assembly line sequencing problem: Mathematical modeling and q-learning based simulated annealing hyper-heuristics. European Journal of Operational Research 282, 530–544. doi:10.1016/j.ejor.2019.09.021. Mousavi, S.S., Schukat, M., Howley, E., 2018. Deep reinforcement learning: An overview, in: Bi, Y., Kapoor, S., Bhatia, R. (Eds.), Proceedings of SAI Intelligent Systems Conference (IntelliSys) 2016, Springer International Publishing, Cham. pp. 426–440. doi:10.1007/978-3-319-56991-8_32.
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
39
Murphy, K., 2025. Reinforcement learning: An overview. arXiv:2412.05265. Nazari, M., Oroojlooy, A., Snyder, L., Takáč, M.T., 2018.
Reinforcement learning for solving
the vehicle routing problem, in: Bengio, S., Wallach, H., Larochelle, H., Grauman, K., CesaBianchi, N., Garnett, R. (Eds.), Advances in Neural Information Processing Systems, Curran Associates, Inc.
URL: https://proceedings.neurips.cc/paper_files/paper/2018/file/
9fb4651c05b2ed70fba5afe0b039a550-Paper.pdf. Nian, R., Liu, J., Huang, B., 2020. A review on reinforcement learning: Introduction and applications in industrial process control. Computers & Chemical Engineering 139, 106886. doi:10.1016/j.compchemeng. 2020.106886. Oroojlooyjadid, A., Nazari, M., Snyder, L.V., Takáč, M., 2022. A deep q-network for the beer game: Deep reinforcement learning for inventory optimization. Manufacturing & Service Operations Management 24, 285–304. doi:10.1287/msom.2020.0939. Panda, S.K., Xiang, Y., Liu, R., 2024. Dynamic resource matching in manufacturing using deep reinforcement learning. European Journal of Operational Research 318, 408–423. doi:10.1016/j.ejor.2024.05.027. Park, K.T., Jeon, S.W., Noh, S.D., 2022.
Digital twin application with horizontal coordination for
reinforcement-learning-based production control in a re-entrant job shop. International Journal of Production Research 60, 2151–2167. doi:10.1080/00207543.2021.1884309. Park, S., Frans, K., Levine, S., Kumar, A., 2024. Is value learning really the main bottleneck in offline rl? arXiv:2406.09329. Parmar, N., Vaswani, A., Uszkoreit, J., Kaiser, L., Shazeer, N., Ku, A., Tran, D., 2018. Image transformer, in: Dy, J., Krause, A. (Eds.), Proceedings of the 35th International Conference on Machine Learning, PMLR. pp. 4055–4064. URL: https://proceedings.mlr.press/v80/parmar18a.html. Powell, W.B., 2007. Approximate dynamic programming: Solving the curses of dimensionality. volume 703. John Wiley & Sons. Recht, B., 2019. A tour of reinforcement learning: The view from continuous control. Annual Review of Control, Robotics, and Autonomous Systems 2, 253–279. doi:10.1146/annurev-control-053018-023825. Rigter, M., Lacerda, B., Hawes, N., 2022.
Rambo-rl: Robust adversarial model-based offline
reinforcement learning, in: Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., Oh, A. (Eds.), Advances in Neural Information Processing Systems, Curran Associates, Inc.. pp. 16082–16097.
URL: https://proceedings.neurips.cc/paper_files/paper/2022/file/
6691c5e4a199b72dffd9c90acb63bcd6-Paper-Conference.pdf. Rolim, G.A., Tomazella, C.P., Nagano, M.S., 2025. On the integration of reinforcement learning and simulated annealing for the parallel batch scheduling problem with setups. European Journal of Operational Research 326, 220–233. doi:10.1016/j.ejor.2025.04.042.
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
40
Schlappa, M., Hegemann, J., Spinler, S., 2024. Optimizing control of waste incineration plants using reinforcement learning and digital twins. IEEE Transactions on Engineering Management 71, 3076–3087. doi:10.1109/TEM.2022.3201434. Schulman, J., Levine, S., Moritz, P., Jordan, M.I., Abbeel, P., 2017a. Trust region policy optimization. URL: https://arxiv.org/abs/1502.05477, arXiv:1502.05477. Schulman, J., Wolski, F., Dhariwal, P., Radford, A., Klimov, O., 2017b. Proximal policy optimization algorithms. URL: https://arxiv.org/abs/1707.06347, arXiv:1707.06347. Semrov, D., Marsetič, R., Žura, M., Todorovski, L., Srdic, A., 2016. Reinforcement learning approach for train rescheduling on a single-track railway. Transportation Research Part B: Methodological 86, 250–267. doi:10.1016/j.trb.2016.01.004. Shakya, A.K., Pillai, G., Chakrabarty, S., 2023. Reinforcement learning algorithms: A brief survey. Expert Systems with Applications 231, 120495. doi:10.1016/j.eswa.2023.120495. Singh, B., Kumar, R., Singh, V.P., 2022. Reinforcement learning in robotic applications: a comprehensive survey. Artificial Intelligence Review 55, 945–990. doi:10.1007/s10462-021-09997-9. Song, H., Triguero, I., Özcan, E., 2019. A review on the self and dual interactions between machine learning and optimisation. Progress in Artificial Intelligence 8, 143–165. doi:10.1007/s13748-019-00185-z. Su, T., Wu, T., Zhao, J., Scaglione, A., Xie, L., 2025. A review of safe reinforcement learning methods for modern power systems. Proceedings of the IEEE 113, 213–255. doi:10.1109/JPROC.2025.3584656. Su, Y., Yang, H., 2025. Enhancing feeder bus service coverage with multi-agent reinforcement learning: A case study in hong kong. Transportation Research Part E: Logistics and Transportation Review 196, 103997. doi:10.1016/j.tre.2025.103997. Sun, W., Xu, N., Wang, L., Zhang, H., Zhang, Y., 2022. Dynamic digital twin and federated learning with incentives for air-ground networks. IEEE Transactions on Network Science and Engineering 9, 321–333. doi:10.1109/TNSE.2020.3048137. Sutton, R.S., 1991. Dyna, an integrated architecture for learning, planning, and reacting. SIGART Bull. 2, 160–163. doi:10.1145/122344.122377. Sutton, R.S., Barto, A.G., 2018. Reinforcement learning: An introduction. 2nd ed., MIT Press, Cambridge, MA. Talbi, E.G., 2016. Combining metaheuristics with mathematical programming, constraint programming and machine learning. Annals of Operations Research 240, 171–215. doi:10.1007/s10479-015-2034-y. Tang, X., Li, X., Yu, R., Wu, Y., Ye, J., Tang, F., Chen, Q., 2023. Digital-twin-assisted task assignment in multi-uav systems: A deep reinforcement learning approach. IEEE Internet of Things Journal 10, 15362–15375. doi:10.1109/JIOT.2023.3263574.
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
41
Tang, Y., Agrawal, S., Faenza, Y., 2020. Reinforcement learning for integer programming: Learning to cut, in: III, H.D., Singh, A. (Eds.), Proceedings of the 37th International Conference on Machine Learning, PMLR. pp. 9367–9376. URL: https://proceedings.mlr.press/v119/tang20a.html. Tassel, P., Gebser, M., Schekotihin, K., 2023. An end-to-end reinforcement learning approach for job-shop scheduling problems based on constraint programming. Proceedings of the International Conference on Automated Planning and Scheduling 33, 614–622. doi:10.1609/icaps.v33i1.27243. Teck, S., am, T.S.P., Rousseau, L.M., Vansteenwegen, P., 2025. Deep reinforcement learning for the realtime inventory rack storage assignment and replenishment problem. European Journal of Operational Research 327, 606–622. doi:10.1016/j.ejor.2025.05.008. Teusch, J., Saavedra, B.N., Scherr, Y.O., Müller, J.P., 2025. Strategic planning of geo-fenced micro-mobility facilities using reinforcement learning. Transportation Research Part E: Logistics and Transportation Review 194, 103872. doi:10.1016/j.tre.2024.103872. Tian, R., Chang, L., Sun, Z., Zhao, G., Lu, X., 2025. Ptb: A deep reinforcement learning method for flexible logistics service combination problem with spatial-temporal constraint. Transportation Research Part E: Logistics and Transportation Review 195, 103978. doi:10.1016/j.tre.2025.103978. Van Hasselt, H., Guez, A., Silver, D., 2016. Deep reinforcement learning with double q-learning. Proceedings of the AAAI Conference on Artificial Intelligence 30. doi:10.1609/aaai.v30i1.10295. Vanvuchelen, N., De Boeck, K., Boute, R.N., 2024. Cluster-based lateral transshipments for the zambian health supply chain. European Journal of Operational Research 313, 373–386. doi:10.1016/j.ejor. 2023.08.005. Verleijsdonk, P., Van Jaarsveld, W., Kapodistria, S., 2024. Scalable policies for the dynamic traveling multimaintainer problem with alerts. European Journal of Operational Research 319, 121–134. doi:10. 1016/j.ejor.2024.05.049. Vinyals, O., Fortunato, M., Jaitly, N., 2015.
Pointer networks, in: Cortes, C., Lawrence, N., Lee,
D., Sugiyama, M., Garnett, R. (Eds.), Advances in Neural Information Processing Systems, Curran Associates, Inc.
URL: https://proceedings.neurips.cc/paper_files/paper/2015/file/
29921001f2f04bd3baee84a12e98098f-Paper.pdf. Wang, D., Wu, J., Chang, X., Yin, H., 2025. Distributed multi-agent reinforcement learning approach for energy-saving optimization under disturbance conditions. Transportation Research Part E: Logistics and Transportation Review 200, 104180. doi:10.1016/j.tre.2025.104180. Wang, H.n., Liu, N., Zhang, Y.y., Feng, D.w., Huang, F., Li, D.s., Zhang, Y.m., 2020. Deep reinforcement learning: a survey. Frontiers of Information Technology & Electronic Engineering 21, 1726–1744. doi:10. 1631/FITEE.1900533. Wang, L., Liang, H., Mao, G., Zhao, D., Liu, Q., Yao, Y., Zhang, H., 2024. Resource allocation for dynamic platoon digital twin networks: A multi-agent deep reinforcement learning method. IEEE Transactions on Vehicular Technology 73, 15609–15620. doi:10.1109/TVT.2024.3414447.
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
42
Wang, W., Li, B., Luo, X., Wang, X., 2023. Deep reinforcement learning for sequential targeting. Management Science 69, 5439–5460. doi:10.1287/mnsc.2022.4621. Wang, Z., Hong, T., 2020. Reinforcement learning for building controls: The opportunities and challenges. Applied Energy 269, 115036. doi:10.1016/j.apenergy.2020.115036. Wang, Z., Schaul, T., Hessel, M., Hasselt, H., Lanctot, M., Freitas, N., 2016. Dueling network architectures for deep reinforcement learning, in: Balcan, M.F., Weinberger, K.Q. (Eds.), Proceedings of The 33rd International Conference on Machine Learning, New York, New York, USA. pp. 1995–2003. doi:10. 5555/3045390.3045601. Watkins, C.J., Dayan, P., 1992. Q-learning. Machine learning 8, 279–292. doi:10.1007/BF00992698. Wei, Z., Xu, J., Lan, Y., Guo, J., Cheng, X., 2017. Reinforcement learning to rank with markov decision process, in: Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval, Association for Computing Machinery, New York, NY, USA. p. 945–948. doi:10.1145/3077136.3080685. Wu, J., Huang, Z., Hang, P., Huang, C., De Boer, N., Lv, C., 2021. Digital twin-enabled reinforcement learning for end-to-end autonomous driving, in: 2021 IEEE 1st International Conference on Digital Twins and Parallel Intelligence (DTPI), pp. 62–65. doi:10.1109/DTPI52967.2021.9540179. Wu, Q., He, M., Hao, J.K., Lu, Y., 2024. An effective hybrid evolutionary algorithm for the clustered orienteering problem. European Journal of Operational Research 313, 418–434. doi:10.1016/j.ejor. 2023.08.006. Wu, Y., Bukhsh, Z., Zhang, Y., 2025. mization: A Tutorial.
Deep Reinforcement Learning for Combinatorial Opti-
Technical Report.
doi:https://research.tue.nl/en/publications/
deep-reinforcement-learning-for-combinatorial-optimization-a-tuto/. Xu, S., Guan, X., Peng, Y., Liu, Y., Cui, C., Chen, H., Ohtsuki, T., Han, Z., 2024. Deep reinforcement learning based data-driven mapping mechanism of digital twin for internet of energy. IEEE Transactions on Network Science and Engineering 11, 3876–3890. doi:10.1109/TNSE.2024.3390797. Yan, P., Yu, K., Chao, X., Chen, Z., 2023. An online reinforcement learning approach to charging and order-dispatching optimization for an e-hailing electric vehicle fleet. European Journal of Operational Research 310, 1218–1233. doi:10.1016/j.ejor.2023.03.039. Yan, Q., Wang, H., Wu, F., 2022. Digital twin-enabled dynamic scheduling with preventive maintenance using a double-layer q-learning algorithm. Computers & Operations Research 144, 105823. doi:10. 1016/j.cor.2022.105823. Ying, C., Chow, A.H., Chin, K.S., 2020. An actor-critic deep reinforcement learning approach for metro train scheduling with rolling stock circulation under stochastic demand. Transportation Research Part B: Methodological 140, 210–235. doi:10.1016/j.trb.2020.08.005.
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
43
Ying, C., Chow, A.H., Nguyen, H.T., Chin, K.S., 2022. Multi-agent deep reinforcement learning for adaptive coordinated metro service operations with flexible train composition. Transportation Research Part B: Methodological 161, 36–59. doi:10.1016/j.trb.2022.05.001. Ying, C., Chow, A.H., Yan, Y., Kuo, Y.H., Wang, S., 2024. Adaptive rescheduling of rail transit services with short-turnings under disruptions via a multi-agent deep reinforcement learning approach. Transportation Research Part B: Methodological 188, 103067. doi:10.1016/j.trb.2024.103067. Yu, C., Liu, J., Nemati, S., Yin, G., 2021. Reinforcement learning in healthcare: A survey. ACM Computing Surveys 55. doi:10.1145/3477600. Yu, H., Gu, W., Tang, N., Guo, Z., 2026. A deep reinforcement learning approach for dynamic job-shop scheduling problem considering time variable and new job arrivals. Computers & Operations Research 185, 107263. doi:10.1016/j.cor.2025.107263. Yu, J., Hyland, M.F., 2025. Interpretable state-space model of urban dynamics for human-machine collaborative transportation planning. Transportation Research Part B: Methodological 192, 103134. doi:10.1016/j.trb.2024.103134. Yue, Y., Yuan, Y., Yu, Q., Zuo, X., Zhu, R., Xu, W., Chen, J., Wang, C., Fan, T., Du, Z., Wei, X., Yu, X., Liu, G., Liu, J., Liu, L., Lin, H., Lin, Z., Ma, B., Zhang, C., Zhang, M., Zhang, W., Zhu, H., Zhang, R., Liu, X., Wang, M., Wu, Y., Yan, L., 2025. Vapo: Efficient and reliable reinforcement learning for advanced reasoning tasks. arXiv:2504.05118. Zhang, Y., Bai, R., Qu, R., Tu, C., Jin, J., 2022. A deep reinforcement learning based hyper-heuristic for combinatorial optimisation with uncertainties. European Journal of Operational Research 300, 418–427. doi:10.1016/j.ejor.2021.10.032. Zhang, Y., Negenborn, R.R., Atasoy, B., 2023. Synchromodal freight transport re-planning under service time uncertainty: An online model-assisted reinforcement learning. Transportation Research Part C: Emerging Technologies 156, 104355. doi:10.1016/j.trc.2023.104355. Zhang, Z., Huang, Y., Zhang, C., Zheng, Q., Yang, L., You, X., 2024. Digital twin-enhanced deep reinforcement learning for resource management in networks slicing. IEEE Transactions on Communications 72, 6209–6224. doi:10.1109/TCOMM.2024.3395698. Zhang, Z., Yu, Y., Qi, X., Lu, Y., Li, X., Kaku, I., 2025. Multi-objective cooperative co-evolution algorithm with hypervolume-based q-learning for hybrid seru system. European Journal of Operational Research 324, 839–854. doi:10.1016/j.ejor.2025.02.025. Zhao, J., Hifi, M., 2025. Reinforcement learning-enhanced variable neighborhood search strategies for the k-clustering minimum biclique completion problem. Computers & Operations Research 178, 107008. doi:10.1016/j.cor.2025.107008. Zhu, Z., Ke, J., Wang, H., 2021. A mean-field markov decision process model for spatial-temporal subsidies in ride-sourcing markets. Transportation Research Part B: Methodological 150, 540–565. doi:10.1016/ j.trb.2021.06.014.
44
Lu et al.: Reinforcement Learning in Operational Research: A Technical Review and Practical Roadmap
Zou, Y., Hao, J.K., Wu, Q., 2024. A reinforcement learning guided hybrid evolutionary algorithm for the latency location routing problem. Computers & Operations Research 170, 106758. doi:10.1016/j. cor.2024.106758.