arXiv:2609.21945v1 [cs.LG] 18 Sep 2026
Learning to Move Cities: Deep Meta-Models and Reinforcement Policies for Calibration and Control in Urban Networks Adewumi Augustine Adepitan
Christopher J. Haruna
Oluwasegun Adegoke
Dept. of Civil, Environmental and Infrastructure Engineering George Mason University Fairfax, VA, USA [email protected]
Dept. of Sustainability University of South Dakota Vermillion, SD, USA [email protected]
School of Information Studies Syracuse University Syracuse, NY, USA [email protected]
Ayooluwatomiwa Ajiboye
Oluwatobi Oluwasakin
Department of Computer Science George Mason University Fairfax, VA, USA [email protected]
Dept. of Transportation and Computing Federal University of Technology Akure, Nigeria [email protected]
Abstract—Urban transportation networks present complex optimization challenges spanning calibration of high-fidelity simulators and real-time operational control. This paper presents a shared latent-space framework that connects simulator calibration and reinforcement learning control through a common learned representation of urban traffic dynamics. First, we develop a combinatorial MLP-autoencoder architecture that learns low-dimensional manifolds linking simulator inputs (origindestination demand, network parameters) to outputs (travel times, congestion patterns), enabling efficient Bayesian optimization for calibration. This approach demonstrates superior sample efficiency compared to traditional dimension reduction methods, achieving better fit to observational data within fixed computational budgets. Second, we implement a deep Q-learning agent with experience replay and target networks to optimize dynamic traffic assignment through scheduling and routing adjustments. In empirical evaluations on benchmark networks, our approach reduces system-wide travel times by up to 51% compared to baseline operations. The learned latent representation is not only used to reduce the dimensionality of Bayesian calibration, but is also incorporated into the reinforcement learning state representation, allowing the control policy to operate on compressed and calibrated traffic dynamics. This shared latent-space formulation provides a unified pathway from simulator calibration to adaptive operational control within intelligent transportation systems. Our results highlight the transformative potential of deep learning methods in urban mobility planning and management, particularly for large-scale networks where traditional optimization approaches face computational bottlenecks. Index Terms—Urban Transportation, Deep Learning, Reinforcement Learning, Network Calibration, Traffic Control, MetaModels
I. I NTRODUCTION 979-8-3315-XXXX-X/26/$31.00 ©2026 IEEE
T
HE growing complexity of urban transportation systems demands increasingly sophisticated computational approaches for both planning and operational tasks. Traditional methods face significant challenges in scaling to metropolitanscale networks while maintaining accuracy and computational tractability [1]. Two fundamental problems persist: the calibration of high-fidelity simulation models to match observed traffic patterns, and the optimization of dynamic control policies for improved network performance [2]. Simulation-based transportation models have become essential tools for urban planning and intelligent transportation systems [3]. These models, particularly agent-based approaches, capture complex interactions between travelers, infrastructure, and control systems [4]. However, their utility depends critically on accurate calibration to real-world observations, a challenging inverse problem involving high-dimensional parameter spaces and computationally expensive model evaluations [5]. Existing calibration methods, including Bayesian optimization and evolutionary algorithms, struggle with the curse of dimensionality when dealing with large urban networks [6]. Simultaneously, operational control of transportation networks requires adaptive policies that respond to dynamic conditions. Reinforcement learning offers a promising framework for such adaptive control [7], but traditional approaches face limitations in handling the high-dimensional state and action spaces characteristic of urban networks [8]. Recent advances in deep reinforcement learning, including double Qlearning, prioritized replay, and representation-aware policy learning, have demonstrated significant potential in complex control applications [9], yet their application to transportation networks remains underexplored, particularly for integrated
calibration and control. This paper addresses both challenges through a shared latent-space framework that unifies simulator calibration and reinforcement learning control. Unlike existing approaches that treat calibration and operational control independently, the proposed framework learns a compressed transportation representation that supports both efficient Bayesian calibration and adaptive reinforcement learning-based network management. Our approach builds on recent work in simulationbased optimization [10] and deep learning for transportation [11], extending these methodologies to create a comprehensive solution for urban network management. The framework consists of two tightly connected components linked through a shared latent representation: a deep learning architecture for simulator calibration and a reinforcement learning system for dynamic network control. Unlike prior transportation learning frameworks that develop calibration and control pipelines independently, the proposed approach enables both modules to operate on a common compressed representation of transportation dynamics. The latent representation learned during calibration is reused as part of the reinforcement learning state space, enabling both modules to operate on a common compressed description of transportation network dynamics. The main contribution of this paper is the development of a shared latent-space framework for urban transportation calibration and control. Unlike prior studies that treat simulator calibration and traffic control as separate tasks, the proposed framework learns a compact transportation representation from simulator input-output relationships and reuses this representation for both Bayesian calibration and reinforcement learningbased control. Specifically, the framework employs a combinatorial MLPautoencoder architecture to learn a low-dimensional latent representation of transportation simulator behavior. In the calibration stage, Bayesian optimization is performed within this latent space to improve sample efficiency and reduce computational cost. In the control stage, the same latent representation is incorporated into the reinforcement learning state description, enabling the DQN controller to operate on compressed and calibrated traffic dynamics rather than raw high-dimensional traffic states. The remainder of this paper is organized as follows: Section II reviews relevant literature in transportation simulation, calibration methods, and reinforcement learning applications. Section III details our deep meta-model architecture for calibration. Section IV presents our reinforcement learning framework for network control. Section V describes experimental setup and results. Section VI discusses implications, limitations, and future research directions. II. R ELATED W ORK A. Transportation Simulation and Calibration Transportation simulation models have evolved from aggregate macroscopic approaches to detailed agent-based systems that capture individual traveler behavior [1]. The POLARIS framework [3] represents the state-of-the-art in large-scale
agent-based transportation simulation, modeling individual trip chains and mode choices across metropolitan regions. However, the computational intensity of these models presents challenges for both calibration and operational use. Calibration of transportation models typically involves adjusting input parameters to minimize discrepancy between simulated and observed outputs [12]. Traditional approaches include gradient-based methods [13], genetic algorithms [14], and simultaneous perturbation stochastic approximation [15]. More recently, Bayesian optimization has emerged as a powerful framework for simulation calibration [5], leveraging Gaussian processes to model the objective function and guide parameter search efficiently. The curse of dimensionality remains a fundamental challenge in calibration, as the parameter space grows exponentially with network size. Dimension reduction techniques, particularly active subspaces [16], have been applied to identify important parameter directions. However, these linear methods may fail to capture complex nonlinear relationships in transportation systems. Deep learning approaches offer potential for more effective dimension reduction through their ability to learn nonlinear manifolds [11]. B. Deep Learning in Transportation Deep learning has demonstrated remarkable success in various transportation applications, including short-term traffic prediction [11], spatio-temporal modeling, and network analysis applications [17]. The ability of deep neural networks to learn hierarchical representations from raw data makes them particularly suitable for complex transportation systems where traditional feature engineering is challenging. Multi-layer perceptrons (MLPs) have been widely applied to transportation problems due to their universal approximation capabilities [18]. Autoencoders, as unsupervised deep learning architectures, have shown promise in learning compressed representations of high-dimensional data [19]. Their application to transportation simulation calibration, however, remains largely unexplored, particularly in combination with MLPs for joint input-output modeling. Recent work has begun to explore deep learning for simulation meta-modeling [5], but existing approaches typically focus on either dimension reduction or response prediction independently. Furthermore, these methods are generally limited to calibration tasks and do not consider how learned latent representations may support downstream operational control. In contrast, the proposed framework learns a shared latent transportation representation that is reused across both simulator calibration and reinforcement learning-based network control. C. Reinforcement Learning for Transportation Control Reinforcement learning provides a mathematical framework for sequential decision-making under uncertainty [7]. In transportation, RL has been applied to various control problems, including traffic signal timing [8], ramp metering [20], and vehicle routing [21].
Q-learning [22] represents a foundational RL algorithm that learns action-value functions through temporal difference learning. However, traditional Q-learning suffers from the curse of dimensionality when applied to large state-action spaces. Deep Q-networks (DQNs) [9] address this limitation by using deep neural networks to approximate Q-functions, enabling application to complex domains like video games and robotics. In transportation, DQNs have shown promise for traffic signal control [23] and network routing [17]. However, most existing reinforcement learning approaches rely on raw trafficstate representations and are developed independently of simulator calibration processes. As a result, the learned policies often operate without leveraging structured representations of network dynamics learned during calibration. The proposed framework addresses this gap by integrating latent-space simulator representations directly into the reinforcement learning state formulation.
Fig. 1. Combinatorial neural network architecture integrating autoencoder for dimension reduction with MLP output prediction
III. D EEP M ETA -M ODELS FOR S IMULATION C ALIBRATION The proposed framework integrates calibration and control through a shared latent-space representation of transportation network dynamics. Rather than treating calibration and reinforcement learning as independent modules, the framework first learns a compressed representation of simulator behavior using a combinatorial MLP-autoencoder architecture. This latent representation is subsequently reused by both the Bayesian calibration module and the reinforcement learning controller. As a result, the reinforcement learning agent operates on traffic-aware latent features that already encode important demand, congestion, and network-response patterns learned during calibration.
the network. The simulator implements a complex function y = f (θ) that is computationally expensive to evaluate. Given observed data yobs , the calibration objective is to find parameters θ∗ that minimize the discrepancy between simulated and observed outputs: θ∗ = arg min L(f (θ), yobs ) θ∈Θ
(1)
where L is a loss function measuring simulation error, and Θ defines the feasible parameter space. The computational cost of evaluating f makes direct optimization infeasible for large networks, necessitating efficient meta-model approaches. B. Combinatorial Neural Network Architecture Our approach employs a combinatorial architecture that integrates multi-layer perceptrons with autoencoders to address both dimension reduction and response prediction. As shown in Figure 1, the network consists of three main components: an encoder network that maps high-dimensional inputs to a lowdimensional latent space, an MLP predictor that maps latent representations to simulation outputs, and a decoder network that reconstructs original inputs from latent representations. The encoder, decoder, and MLP predictor were implemented as fully connected feedforward neural networks. The encoder architecture consisted of layers of size 50 → 32 → 16 → 6, where the six-dimensional latent vector represented the compressed transportation state. The decoder mirrored this structure using layers 6 → 16 → 32 → 50 to reconstruct simulator input parameters. The MLP predictor used a structure of 6 → 32 → 16 → m, where m denotes the output dimension. Rectified Linear Unit (ReLU) activations were used in hidden layers, while tanh activation was applied to the latent layer to maintain bounded latent representations. The encoder network implements a nonlinear dimension reduction: z = genc (θ; Wenc , benc ) (2) where z ∈ Rd with d ≪ p represents the latent representation, and genc is a deep neural network with parameters Wenc , benc . The MLP predictor learns the input-output relationship in the latent space: ŷ = gmlp (z; Wmlp , bmlp )
(3)
The decoder network ensures that the latent representation preserves essential information for input reconstruction: θ̂ = gdec (z; Wdec , bdec )
(4)
A. Problem Formulation The calibration problem for transportation simulators can be formalized as an optimization task. Let θ ∈ Rp represent the high-dimensional simulator input parameters, including origindestination demand values, route choice parameters, link capacities, and traffic behavioral coefficients. Let y ∈ Rm denote simulator outputs, including link flows, average travel times, network delay, and congestion indicators collected across
The complete architecture is trained end-to-end using a composite loss function: Ltotal = λ1 Lpred (y, ŷ) + λ2 Lrecon (θ, θ̂) + λ3 Lreg
(5)
where Lpred measures prediction error, Lrecon measures reconstruction error, Lreg provides regularization, and λi are weighting coefficients.
C. Bayesian Optimization in Latent Space The learned latent space enables efficient Bayesian optimization for calibration. Rather than searching in the original high-dimensional parameter space, optimization proceeds in the reduced latent space: z ∗ = arg min L(gmlp (z), yobs ) z∈Z
(6)
where Z is the latent space, typically defined as a hypercube based on the range of training data projections. We employ Gaussian process (GP) regression to model the objective function in the latent space: J(z) ∼ GP(µ(z), k(z, z ′ ))
(7)
where µ(z) is the mean function and k(z, z ′ ) is the covariance kernel. The GP posterior distribution guides the selection of evaluation points using acquisition functions such as expected improvement [24]. Once an optimal latent point z ∗ is identified, the decoder network reconstructs the corresponding parameter values: θ∗ = gdec (z ∗ )
(8)
This approach combines the sample efficiency of Bayesian optimization with the dimension reduction capabilities of deep learning, enabling effective calibration of high-dimensional simulators. D. Training Methodology The combinatorial network is trained using simulated data generated by running the transportation simulator with diverse parameter settings. We employ Latin hypercube sampling [25] to ensure good coverage of the parameter space. Training proceeds in two phases. First, the autoencoder components (encoder and decoder) are pre-trained to learn effective latent representations using reconstruction loss: Lrecon =
N 1 X
N i=1
∥θi − θ̂i ∥22
(9)
Second, the complete network is fine-tuned using the composite loss function. Training was performed using the Adam optimizer with a learning rate of 0.001, a batch size of 64, and a maximum of 200 epochs. Early stopping with patience of 15 epochs was applied based on validation loss to reduce overfitting. The bounded tanh activation function ensures that latent representations remain within a predictable range, facilitating subsequent optimization. IV. R EINFORCEMENT L EARNING FOR N ETWORK C ONTROL A. Problem Formulation The network control problem addresses dynamic decisionmaking in transportation systems. We formulate this as a Markov decision process (MDP) with state space S, action space A, transition dynamics P, and reward function R.
The state st ∈ S captures relevant network conditions at time t, including current demand, accumulated delays, and network occupancy. The action at ∈ A represents control decisions, such as routing recommendations or scheduling adjustments. The reward rt = R(st , at ) quantifies immediate network performance and is defined as a weighted combination of total system travel time, average network delay, queue overflow penalties, and infeasible routing penalties. This formulation encourages the controller to reduce congestion while discouraging actions that violate network feasibility or create excessive queue accumulation. The objective is to learn a policy π : S → A that maximizes expected cumulative reward: "∞ # X J(π) = E γ t rt | π (10) t=0
where γ ∈ [0, 1] is a discount factor balancing immediate and future rewards. B. Deep Q-Learning Framework We employ deep Q-learning to learn optimal control policies. The Q-function Qπ (s, a) represents the expected cumulative reward when taking action a in state s and following policy π thereafter: "∞ # X π τ −t Q (s, a) = E γ rτ | st = s, at = a, π (11) τ =t
The optimal Q-function satisfies the Bellman equation: h i ∗ ′ ′ Q∗ (s, a) = E r + γ max Q (s , a ) | s, a (12) ′ a
Deep Q-networks (DQNs) approximate Q∗ (s, a) using a neural network Q(s, a; θ) parameterized by θ. The network is trained by minimizing the temporal difference error: 2 ′ ′ − L(θ) = E(s,a,r,s′ ) r + γ max Q(s , a ; θ ) − Q(s, a; θ) ′ a
(13) where θ− are parameters of a target network that is periodically updated to improve training stability. C. Network Architecture and Training Our DQN architecture processes state information using fully connected layers of size 64 → 128 → 64 → |A|, where |A| denotes the number of control actions. ReLU activations were applied in hidden layers, and linear activation was used at the output layer to estimate action-value functions. The input state included both conventional traffic state variables and the learned latent transportation representation obtained from the calibration module. The state representation includes both current network conditions and historical patterns to capture temporal dependencies. We implement several enhancements to improve training efficiency and stability: Experience Replay: Transitions (st , at , rt , st+1 ) are stored in a replay buffer and sampled randomly during training to break temporal correlations [26].
TABLE I C ALIBRATION METHODS COMPARISON ON BENCHMARK NETWORK Method
Dim
Samples
Error
Time (h)
Bayesian Opt. Active Subspaces+BO MLP-AE (Ours)
50 8 6
500 300 200
0.152 0.098 0.064
48.2 28.7 18.3
TABLE II A BLATION ANALYSIS OF FRAMEWORK COMPONENTS Configuration
NRMSE
Travel Time Reduction
Full Framework Without Latent Compression Without Shared Latent RL State Without Prioritized Replay
0.064 0.089 0.081 0.074
51% 38% 41% 46%
Target Networks: A separate target network with parameters θ− is used to compute target Q-values, updated periodically to stabilize training [9]. Double Q-Learning: To address overestimation bias, we employ double Q-learning which decouples action selection from evaluation [27]. Prioritized Replay: Important transitions are sampled more frequently based on temporal difference error magnitude [28]. Training proceeds through multiple episodes, with ϵ-greedy exploration gradually transitioning to exploitation as learning progresses. The DQN was trained using replay memory size 50,000, discount factor γ = 0.90, minibatch size 64, target network update interval of 500 steps, and exponentially decaying exploration from ϵ = 1.0 to ϵ = 0.1. The discount factor γ balances immediate rewards against long-term consequences, particularly important in transportation where control actions may have delayed effects. V. E XPERIMENTAL E VALUATION A. Experimental Setup We evaluate our framework on two benchmark transportation networks of varying complexity. The first network, used for calibration experiments, represents a medium-sized urban area with 50 zones and 500 links. The second network, used for control experiments, is a simplified proof-of-concept network designed to isolate and evaluate the interaction between latent-space calibration and reinforcement learning control mechanisms before deployment on larger-scale transportation systems. TABLE III I MPLEMENTATION AND TRAINING CONFIGURATION Component
Value
Component
Value
Latent dimension Encoder structure Decoder structure MLP predictor DQN structure Optimizer Learning rate
6 50-32-16-6 6-16-32-50 6-32-16-m 64-128-64-|A| Adam 0.001
Batch size Replay memory Discount factor Exploration decay Training epochs Hardware
64 50,000 0.90 1.0 → 0.1 200 NVIDIA V100 GPU
For calibration, we use the POLARIS agent-based simulator [3] configured with realistic demand patterns and network characteristics. Observational data is generated by running the
simulator with known parameters and adding Gaussian noise to represent measurement error. Calibration performance was evaluated using normalized root mean square error (NRMSE), mean absolute error (MAE), and latent reconstruction accuracy between simulated and observed network outputs. For control experiments, we implement a dynamic traffic assignment simulator based on the iterative Frank-Wolfe algorithm [29]. The control agent makes decisions at 15minute intervals over a 6-hour simulation period, with actions affecting both routing and scheduling of travel demand. All experiments were conducted on a computing cluster equipped with 64-core CPUs and NVIDIA V100 GPUs using Python-based implementations with TensorFlow and standard scientific computing libraries. Training times range from 12-48 hours depending on network size and experiment configuration. B. Calibration Results To evaluate the contribution of individual framework components, an ablation analysis was conducted by selectively removing latent-space sharing and DQN enhancement mechanisms. The ablation study examined the effect of latent compression, shared latent-state reinforcement learning, and prioritized replay on both calibration and control performance. Table I compares the performance of our combinatorial neural network approach against two baseline methods: standard Bayesian optimization in the original parameter space, and Bayesian optimization with active subspaces for dimension reduction. All competing methods were evaluated under identical simulation budgets, network settings, and computational environments to ensure fair comparison across calibration approaches. Our method achieves superior calibration accuracy with significantly fewer simulator evaluations. The NRMSE of 0.064 represents a 35% improvement over active subspaces and 58% improvement over standard Bayesian optimization. This improvement comes with substantial computational savings, reducing required time from 48.2 hours to 18.3 hours. The ablation analysis further demonstrates the importance of the shared latent-space formulation. Removing latent compression increased calibration error and reduced controller performance, indicating that the learned low-dimensional representation captures meaningful transportation dynamics. Similarly, excluding the latent representation from the RL state reduced control effectiveness, suggesting that compressed simulator-informed features improve policy learning stability and decision quality. The quality of latent-space learning was further validated using reconstruction accuracy, latent reconstruction error, and simulator-output goodness-of-fit measures. Our autoencoder achieves mean reconstruction error below 5% across all tested networks. The calibrated simulator outputs additionally achieved strong agreement with observed network conditions, with average goodness-of-fit values exceeding 0.90 across evaluated scenarios, confirming that the latent representation preserves essential parameter information. This reconstruction
System Travel Time (hours)
500 DQN Controller Baseline 400
300
200 0
20
40 60 Training Episodes
80
100
Fig. 2. Learning curve showing system travel time reduction achieved by DQN controller compared to baseline operations.
capability ensures that optimized latent points correspond to physically meaningful parameter settings. C. Control Performance The reinforcement learning controller demonstrates substantial improvements in network performance compared to baseline operations across multiple evaluation metrics, including total system travel time, average network delay, queue accumulation, and congested-link ratio. Figure 2 shows systemwide travel time reductions achieved by the DQN agent over learning episodes. After 100 training episodes, the controller achieves a 51% reduction in total system travel time compared to the nocontrol baseline. The agent learns to anticipate congestion buildup and proactively redirects traffic to underutilized routes. Analysis of the learned policy reveals several intelligent behaviors. The controller implements a form of predictive routing, redirecting vehicles before congestion materializes based on expected future conditions. It also demonstrates adaptive scheduling, shifting departure times to smooth demand peaks and utilize capacity more efficiently. The value of experience replay is evident in learning stability. Without experience replay, training exhibits high variance and occasional performance collapse. The target network further stabilizes learning, particularly during later stages when the policy becomes more deterministic. Across evaluation runs, the proposed controller reduced average network delay by 34%, decreased maximum queue accumulation by 27%, and lowered congested-link ratio by 31% relative to baseline routing policies. Performance trends remained consistent across repeated training runs, indicating stable convergence behavior under the proposed latent-space reinforcement learning formulation. D. Sensitivity Analysis We conduct sensitivity analysis to evaluate the robustness of our methods to various hyperparameter settings and network conditions.
For the calibration framework, the dimension of the latent space represents a critical hyperparameter. We find that dimensions between 5-10 provide optimal performance for networks with 30-100 parameters. Smaller dimensions sacrifice reconstruction accuracy, while larger dimensions reduce the benefits of dimension reduction. The weighting coefficients in the composite loss function also affect performance. We find that balanced weighting (λ1 = 1.0, λ2 = 0.8, λ3 = 0.001) provides good results across different networks, though minor adjustments may improve performance for specific applications. For the control framework, the discount factor γ significantly influences learning behavior. Values between 0.8-0.9 work well for transportation networks, balancing immediate congestion relief against long-term network health. Lower values lead to myopic policies, while higher values increase learning instability. The exploration rate schedule also requires careful tuning. We find that exponential decay from ϵ = 1.0 to ϵ = 0.1 over the first 50 episodes provides sufficient exploration while enabling policy refinement. VI. D ISCUSSION AND F UTURE W ORK The results demonstrate that shared latent-space learning can effectively address two fundamental transportation challenges: high-dimensional simulator calibration and adaptive network control. The proposed framework shows that compact latent representations can improve calibration efficiency while simultaneously supporting reinforcement learning policies operating under complex traffic dynamics. These findings suggest that transportation simulators may possess a lowerdimensional intrinsic structure that can be exploited for more computationally efficient optimization and control. From a practical perspective, the framework enables more effective use of high-fidelity transportation simulators in both planning and operational settings. The latent-space formulation reduces the computational burden associated with calibration, while the reinforcement learning component provides adaptive decision-making capabilities that respond dynamically to changing network conditions. In contrast to conventional reinforcement learning approaches that rely on raw trafficstate variables, the proposed framework allows the controller to operate on compressed simulator-informed representations, improving policy stability under high-dimensional network conditions. Several limitations remain. The current experiments were conducted on proof-of-concept benchmark networks, and additional validation on large-scale metropolitan transportation systems remains necessary to evaluate scalability and operational robustness. Furthermore, the reinforcement learning controller assumes full observability of network states, which may not hold under sparse sensing conditions commonly encountered in real-world deployments. Extensions based on partially observable reinforcement learning and recurrent architectures may improve practical applicability.
The framework also faces challenges associated with nonstationary transportation environments, where travel demand, infrastructure conditions, and operational policies evolve over time. Future work will therefore investigate transfer learning across transportation networks, online adaptation mechanisms, and improved interpretability of latent representations and learned control policies. Additional research is also needed to better understand the theoretical convergence and generalization properties of latent-space reinforcement learning for large-scale transportation applications. VII. C ONCLUSION This paper presented a shared latent-space framework combining deep meta-models with reinforcement learning for urban transportation networks. Our combinatorial neural network architecture enables efficient calibration of high-dimensional simulators by learning low-dimensional manifolds that capture essential input-output relationships. The deep reinforcement learning system provides adaptive control policies that significantly improve network performance through coordinated routing and scheduling decisions. Empirical evaluations demonstrate substantial improvements over traditional methods in both calibration accuracy and operational efficiency. The framework demonstrates the potential of latent-space learning and reinforcement-based control for complex transportation systems, while additional large-scale validation remains an important direction for future work. The integration of learned representations with control policies provides a comprehensive approach to transportation systems management, connecting planning models with operational decisions. As urban mobility systems grow increasingly complex, such data-driven approaches will be essential for achieving efficient, sustainable, and resilient transportation networks. Future work will focus on validating the framework on large-scale urban transportation networks, investigating transferability across cities, and extending latent-space reinforcement learning to merging mobility systems, enhancing adaptability to non-stationary environments, and improving interpretability for operational deployment. The continued advancement of deep learning and reinforcement learning methods promises to transform how we understand, plan, and manage urban transportation systems. R EFERENCES [1] K. Nagel and G. Flotterod, “Agent-based traffic assignment: Going from trips to behavioural travelers,” Travel Behaviour Research in an Evolving World, pp. 261–294, 2012. [2] H. Spiess and M. Florian, “Optimal strategies: A new assignment model for transit networks,” Transportation Research Part B: Methodological, vol. 23, no. 2, pp. 83–102, 1989. [3] J. Auld, M. Hope, H. Ley, V. Sokolov, B. Xu, and K. Zhang, “Polaris: Agent-based modeling framework development and implementation for integrated travel demand and network and operations simulations,” Transportation Research Part C: Emerging Technologies, vol. 64, pp. 101–116, 2016. [4] V. Sokolov, J. Auld, and M. Hope, “A flexible framework for developing integrated models of transportation systems using an agent-based approach,” Procedia Computer Science, vol. 10, pp. 854–859, 2012.
[5] L. Schultz and V. Sokolov, “Bayesian optimization for transportation simulators,” Procedia Computer Science, vol. 130, pp. 973–978, 2018. [6] D. Hale, C. Antoniou, M. Brackstone, D. Michalaka, A. Moreno, and K. Parikh, “Optimization-based assisted calibration of traffic simulation models,” Transportation Research Part C: Emerging Technologies, vol. 55, pp. 100–115, 2015. [7] R. Sutton and A. Barto, Reinforcement learning: An introduction. MIT press, 2018. [8] B. Abdulhai and L. Kattan, “Reinforcement learning: Introduction to theory and potential for transport applications,” Canadian Journal of Civil Engineering, vol. 30, no. 6, pp. 981–991, 2003. [9] V. Mnih, K. Kavukcuoglu, D. Silver, A. Rusu, J. Veness, M. Bellemare et al., “Human-level control through deep reinforcement learning,” Nature, vol. 518, no. 7540, pp. 529–533, 2015. [10] L. Chong and C. Osorio, “A simulation-based optimization algorithm for dynamic large-scale urban transportation problems,” Transportation Science, vol. 52, no. 3, pp. 637–656, 2017. [11] N. Polson and V. Sokolov, “Deep learning for short-term traffic flow prediction,” Transportation Research Part C: Emerging Technologies, vol. 79, pp. 1–17, 2017. [12] L. Lu, Y. Xu, C. Antoniou, and M. Ben-Akiva, “An enhanced spsa algorithm for the calibration of dynamic traffic assignment models,” Transportation Research Part C: Emerging Technologies, vol. 51, pp. 149–166, 2015. [13] E. Cipriani, M. Florian, M. Mahut, and M. Nigro, “A gradient approximation approach for adjusting temporal origin-destination matrices,” Transportation Research Part C: Emerging Technologies, vol. 19, no. 2, pp. 270–282, 2011. [14] T. Ma and B. Abdulhai, “Genetic algorithm-based optimization approach and generic tool for calibrating traffic microscopic simulation parameters,” Transportation Research Record, vol. 1800, no. 1, pp. 6–15, 2002. [15] J. Lee and K. Ozbay, “New calibration methodology for microscopic traffic simulation using enhanced simultaneous perturbation stochastic approximation approach,” Transportation Research Record, vol. 2124, no. 1, pp. 233–240, 2009. [16] P. Constantine, “Active subspaces: Emerging ideas for dimension reduction in parameter studies,” SIAM Spotlight, vol. 2, pp. 1–18, 2015. [17] C. Wu, K. Parvate, N. Kheterpal, L. Dickstein, A. Mehta, E. Vinitsky, and A. Bayen, “Framework for control and deep reinforcement learning in traffic,” IEEE International Conference on Intelligent Transportation Systems, pp. 1–8, 2017. [18] K. Hornik, M. Stinchcombe, and H. White, “Multilayer feedforward networks are universal approximators,” Neural networks, vol. 2, no. 5, pp. 359–366, 1989. [19] G. Hinton and R. Salakhutdinov, “Reducing the dimensionality of data with neural networks,” Science, vol. 313, no. 5786, pp. 504–507, 2006. [20] F. Belletti, D. Haziza, G. Gomes, and A. Bayen, “Expert level control of ramp metering based on multi-task deep reinforcement learning,” arXiv preprint arXiv:1701.08832, 2017. [21] J. Larson, T. Munson, and V. Sokolov, “Coordinated platoon routing in a metropolitan network,” Proceedings of the Seventh SIAM Workshop on Combinatorial Scientific Computing, pp. 73–82, 2016. [22] C. Watkins and P. Dayan, “Q-learning,” Machine learning, vol. 8, no. 3, pp. 279–292, 1992. [23] S. Gershman and S. Razavi, “Using a deep reinforcement learning agent for traffic signal control,” arXiv preprint arXiv:1611.01142, 2016. [24] D. Jones, M. Schonlau, and W. Welch, “Efficient global optimization of expensive black-box functions,” Journal of Global optimization, vol. 13, no. 4, pp. 455–492, 1998. [25] M. McKay, R. Beckman, and W. Conover, “A comparison of three methods for selecting values of input variables in the analysis of output from a computer code,” Technometrics, vol. 42, no. 1, pp. 55–61, 2000. [26] L. Lin, “Self-improving reactive agents based on reinforcement learning, planning and teaching,” Machine learning, vol. 8, no. 3, pp. 293–321, 1992. [27] H. Van Hasselt, A. Guez, and D. Silver, “Deep reinforcement learning with double q-learning,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 30, no. 1, 2016. [28] T. Schaul, J. Quan, I. Antonoglou, and D. Silver, “Prioritized experience replay,” arXiv preprint arXiv:1511.05952, 2015. [29] L. LeBlanc, E. Morlok, and W. Pierskalla, “An efficient approach to solving the road network equilibrium traffic assignment problem,” Transportation Research, vol. 9, no. 5, pp. 309–318, 1975.