© 2026 IEEE. Personal use of this material is permitted. However, permission to use this material for any other purposes must be obtained from the IEEE by sending a request to [email protected].
1
Adaptive and Resilient Dual-Layer Resource Slicing for Hovering Aerial Backhaul Networks
arXiv:2609.32798v1 [cs.MA] 26 Sep 2026
Chuan-Chi Lai, Member, IEEE, and Jen-Hsiang Li
Abstract—This paper investigates adaptive and resilient duallayer resource slicing in hovering aerial agent (HAA)-assisted backhaul networks for heterogeneous 5G/6G services, including enhanced mobile broadband (eMBB), ultra-reliable and lowlatency communications (URLLC), and massive machine-type communications (mMTC). To address the complex coupling of this dual-layer architecture in non-stationary environments, we propose the resilient adaptive priority orchestration enhanced twin delayed deep deterministic policy gradient (RAPO-TD3) framework. We introduce a novel double soft-max projection mechanism to map the continuous action space into physically feasible bandwidth distributions, ensuring strict constraint adherence. Additionally, a resilient adaptive priority orchestration (RAPO) mechanism is embedded to safeguard mission-critical URLLC latency. Crucially, we establish a rigorous mathematical foundation proving that our framework ensures Lipschitz continuity and satisfies the Robbins-Monro conditions for stable asymptotic convergence. Extensive simulations under nonstationary traffic demonstrate that our RAPO-TD3 framework achieves superior performance relative to PPO, DDPG, and traditional solvers. Notably, via the RAPO mechanism, our approach maintains URLLC satisfaction levels closely approaching theoretical optima even during 500% demand surges. Furthermore, scalability evaluations indicate that sub-millisecond execution latencies strictly satisfy the 1 ms URLLC budget, demonstrating the performance efficacy of our proposed framework. Index Terms—HAA-assisted backhaul, network slicing, RAPOTD3, URLLC, priority optimization
I. I NTRODUCTION
T
HE evolution of sixth-generation (6G) wireless networks demands global coverage, intelligent sensing, and programmable networking [1], [2], [3]. To realize this vision, 6G architectures must support heterogeneous services with conflicting requirements [4], [5], [6]. These include enhanced mobile broadband (eMBB) for high-data-rate applications, ultra-reliable and low-latency communications (URLLC) for mission-critical millisecond-level constraints, and massive machine-type communications (mMTC) for hyper-dense Internet of Things connectivity [7], [8]. However, terrestrial inThis research was supported by the National Science and Technology Council, Taiwan, under Grant Nos. NSTC 114-2221-E-194-062- and NSTC 115-2221-E-194-042-MY2. This work was also partially supported by the Advanced Institute of Manufacturing with High-tech Innovations (AIM-HI) from the Featured Areas Research Center Program within the framework of the Higher Education Sprout Project by the Ministry of Education (MOE) in Taiwan. (Corresponding author: Chuan-Chi Lai.) Chuan-Chi Lai is with the Department of Communications Engineering, National Chung Cheng University, Minxiong Township, Chiayi County 621301, Taiwan, and also with the Advanced Institute of Manufacturing with Hightech Innovations (AIM-HI), National Chung Cheng University, Minxiong Township, Chiayi County 621301, Taiwan (e-mail: [email protected]). Jen-Hsiang Li is with the Department of Information Engineering and Computer Science, Feng Chia University, Taichung 407102, Taiwan.
frastructure deployment is frequently impeded by geographical barriers and protracted construction cycles, limiting capacity expansion during temporary hotspots or post-disaster recovery [9]. Similarly, traditional backhaul solutions lack the reconfigurability to adapt to rapid spatial-temporal traffic shifts. Consequently, hovering aerial agent (HAA)-assisted backhaul networks have emerged as a promising architecture [10], [11]. Leveraging rapid deployment capabilities, HAA fleets form an agile three-dimensional topology, acting as both backhaul and access nodes to extend coverage and guarantee quality of service (QoS) in dynamic environments [12], [13]. Despite these advantages, HAA systems are inherently constrained by limited radio frequency bandwidth and finite battery capacity [14], [15]. The primary challenge is providing differentiated eMBB, URLLC, and mMTC services within this restricted airspace. Coordinating these resources while maintaining strict isolation and service-specific satisfaction remains a critical optimization problem in 6G architectures. Existing backhaul management predominantly relies on offline planning and static slicing, pre-partitioning spectral resources into fixed proportions [12], [16]. Although effective under quasi-static loads, these approaches lack the agility to handle non-stationary traffic surges during large-scale events or emergency missions. Aligning with the 3GPP management and orchestration framework [17], transitioning toward autonomous and adaptive slicing paradigms is essential to handle highly volatile demand. Without this adaptability, unmanaged eMBB spikes may starve mMTC connections or push URLLC latencies beyond survival thresholds, causing critical failures. To protect mission-critical links under resource scarcity, traditional frameworks often utilize hard preemption strategies that abruptly terminate lower-priority connections, inducing severe packet dropouts and system oscillations. Ultimately, these mechanisms lack the real-time granularity and structural differentiability needed to gracefully balance diverse QoS metrics within volatile aerial environments. Deep reinforcement learning (DRL) provides a potent framework for adaptive slicing through its intrinsic online interaction and autonomous policy tuning capabilities [18], [19], [20]. Nevertheless, conventional DRL models, such as deep deterministic policy gradient (DDPG), struggle with high-dimensional continuous action spaces where core-layer slice ratios and HAA-layer sub-slice allocations are hierarchically coupled and jointly constrained [21]. During exploration, standard continuous-control algorithms frequently generate unconstrained outputs violating physical bandwidth simplex boundaries, which compromises resource feasibility and destabilizes policies. Furthermore, hard-threshold reward functions
2
often cause sparse gradients and convergence instability. To address these limitations, this study proposes the resilient adaptive priority orchestration enhanced twin delayed deep deterministic policy gradient (RAPO-TD3) framework. Building upon the baseline TD3 algorithm (which employs twinQ networks to suppress overestimation bias, delayed actor updates for stability, and target policy smoothing to mitigate oscillations [18], [22]), our approach integrates a double softmax projection technique with differentiable penalty shaping. This structural design ensures decisions strictly adhere to feasible resource regions while preserving continuous gradients for efficient learning. This research deploys this intelligent architecture for duallayer network slicing in HAA-assisted backhaul systems. By incorporating prioritized experience replay (PER) [23] alongside our newly developed resilient adaptive priority orchestration (RAPO) mechanism, the proposed solution enhances operational resilience against stochastic traffic bursts without brittle, non-differentiable hard preemption strategies. • Adaptive Dual-Layer Architecture and Differentiable Projection: We formulate a hierarchical slicing model for aerial backhaul to maximize heterogeneous service satisfaction. A novel double soft-max projection transforms the unconstrained action space into feasible bandwidth allocations, guaranteeing simplex constraint adherence during training and execution without heuristic boundary rules. • Resilient Adaptive Priority Orchestration Mechanism: The core RAPO mechanism is embedded within a differentiable penalty-driven reward function utilizing adaptive smooth gating. This provides dense gradient propagation, enabling the agent to dynamically shift resource focus and resiliently safeguard URLLC reliability under extreme workloads. • Theoretical Convergence and Stability Guarantees: We establish a rigorous mathematical foundation for algorithmic stability. We formally prove that the continuously differentiable design ensures action space boundedness and Lipschitz continuity, satisfying the Robbins-Monro conditions [24] to mathematically guarantee asymptotic convergence to a stable local optimum without divergent oscillations. • Performance and Scalability Verification: Extensive simulations under non-stationary traffic demonstrate that RAPO-TD3 achieves superior performance relative to DDPG and PPO baselines. By integrating the proposed Double Soft-max Projection mechanism, the RAPO-TD3 framework enables URLLC satisfaction levels closely approaching theoretical optima during 500% traffic surges, restricting steady-state constraint violation penalties to approximately 0.08 compared with 3.3 in unconstrained baselines, while strictly maintaining sub-millisecond execution latencies. II. R ELATED W ORK Resource management in aerial-assisted networks has evolved significantly with the integration of network slicing
and advanced machine learning techniques. This section reviews relevant literature across three key domains: network slicing paradigms, aerial deployment optimization, and deep reinforcement learning frameworks. A. Network Slicing Paradigms Network slicing provides the essential flexibility and scalability required to adapt to evolving service demands, facilitating rapid application deployment and efficient service scaling. However, the coordinated resource allocation across heterogeneous slices, encompassing eMBB for high throughput, URLLC for low latency, and mMTC for massive connectivity, remains a formidable challenge. Extensive research has investigated resource management strategies tailored for these diverse slices. Early studies in [25] formulated the resource slicing problem over licensed and unlicensed bands as a mathematical optimization model to maximize the aggregate utility of heterogeneous services.To manage the coexistence of eMBB and URLLC services, the integration of slicing with physical layer technologies, such as non-orthogonal multiple access (NOMA), was demonstrated in [26] as an essential approach to balancing high throughput with stringent latency constraints. Beyond fundamental allocation, network slicing has been conceptualized as a critical backbone for IoT connectivity in smart cities, with automated scheduling schemes proposed in [27] to satisfy diverse sensing task requirements. Current research predominantly focuses on end-to-end (E2E) resource coordination and service layer orchestration. For instance, cloud-native container technology has been explored to establish flexible slicing service chains between softwarized HAA nodes and terrestrial core networks [12]. In the context of aerial access and backhaul, slice-aware joint optimization of deployment and spectrum allocation has been shown to significantly reduce resource consumption, particularly when aerial nodes are limited [28]. Recent advancements have further extended slicing paradigms to emerging fields, such as integrated sensing and communication (ISAC) [3] and large-scale IoT frameworks for 6G intelligent infrastructures [1]. Collectively, achieving fine-grained bandwidth and power allocation while ensuring QoS and security isolation remains a pivotal issue in the evolution of slicing technologies. B. HAA Deployment and Backhaul Resource Management The proliferation of HAA technology has facilitated the development of three-dimensional network topologies that integrate aerial access with backhaul connectivity. While HAAs enable rapid deployment for emergency response or temporary events, they are governed by stringent physical constraints, including finite bandwidth, limited power, and battery endurance. Additionally, backhaul links in these environments are highly susceptible to environmental fluctuations and lineof-sight obstructions. Initial research in this domain focused primarily on frontend access optimization. For instance, stochastic geometry analysis has been employed to reveal the sensitivity of uplink interference to mMTC success rates, emphasizing the critical
3
balance between deployment density and power control [9]. Surveys on the intersection of 6G and artificial intelligence have further positioned HAAs as portable edge inference nodes, highlighting their utility for latency-sensitive services within the Internet of Vehicles (IoV) and smart city infrastructures [11]. Regarding the coordination of multiple aerial nodes, an adaptive and fair deployment approach was proposed in [29] to balance offload traffic in multi-HAA cellular networks, ensuring equitable resource distribution among heterogeneous ground users. Recent trends have shifted toward the joint design of access and backhaul integration. A dual-scale approach was proposed in [12], where HAA positions and backhaul capacities are optimized offline, while slice resources are dynamically allocated online using DRL to satisfy service-level requirements. In the realm of integrated access and backhaul (IAB), cooperative parallel resource allocation mechanisms, such as CPReal, have been introduced to manage both bursty and non-bursty traffic profiles [13]. Furthermore, issues regarding fairness in multioperator architectures [14] and real-time trade-offs between eMBB and URLLC traffic in multi-hop relay networks [8] have been investigated. Federated learning has also been explored to facilitate decentralized slicing decisions under multi-operator sharing and open radio access network (Open RAN) frameworks [15]. C. DRL-based Resource Management Traditional optimization and heuristic algorithms often struggle with the high-dimensional and non-stationary environments of 5G/6G networks. Consequently, DRL has been widely adopted for wireless resource management due to its ability to learn optimal policies through trial-and-error interactions. Integrating deep neural networks (DNNs) with reinforcement learning has enabled systems to transcend the dimensionality constraints inherent in traditional methods [18], with the trade-off between exploration and exploitation serving as a critical factor for convergence quality [30]. In network slicing, predictive-assisted DRL has demonstrated potential by achieving significant latency reductions through real-time optimization of power and user access [19]. For heterogeneous networks with mobile edge computing (MEC), multi-agent TD3 (MATD3) algorithms have been successfully implemented to maximize spectral efficiency [20]. To enhance service orchestration, federated DRL frameworks have been proposed to coordinate multiple xAPPs in Open RAN [31], while dual-granularity RAN slicing strategies have been developed to concurrently maximize QoS and resource utilization across the entire slice lifecycle [16]. Advanced architectures, such as hierarchical DRL, have been introduced to coordinate control between the RAN and the core layer, effectively reducing backhaul signaling overhead [21]. In HAA-assisted scenarios, the synergy between DRL and energy prediction models has proven effective for joint trajectory and bandwidth scheduling [12]. Recent literature in 2026 has further emphasized the necessity of handling complex service coexistence and dynamic adaptations. For instance, advanced soft actor-critic architectures have been
proposed to manage the stringent coexistence of URLLC traffic with distributed learning services through intelligent device selection [32]. Similarly, enhanced TD3 algorithms featuring dynamic normalization and adaptive weights have been successfully applied to balance delay and energy consumption in aerial collaborative computing [33]. These latest developments lay the foundation for the dual-layer continuous control framework proposed in this study. D. Summary of Research Gaps and Motivation A comprehensive comparison of the proposed framework with existing literature is summarized in Table I. Despite the significant advancements reviewed in the previous subsections, including the very recent progress in DRL-driven service coexistence [32] and adaptive TD3 offloading [33], several critical research gaps remain unaddressed. As categorized in Table I under the Slicing Hierarchy column, most existing literature treats HAAs merely as extensions of terrestrial base stations, predominantly focusing on single-layer resource division between the base station (BS) and HAAs. Such simplified models often overlook the two-tier coupling mechanism between core-layer slice proportions and HAA-layer sub-slice allocations, failing to capture the complex interdependencies within a multi-layer backhaul network. To bridge these gaps, this paper proposes a dual-layer framework that first performs slicing at the core layer and subsequently subdivides these resources at the HAA layer for mMTC, URLLC, and eMBB services. Our work distinguishes itself from existing studies in the following key aspects: • Algorithmic Robustness: We combine the stability of the TD3 algorithm with the sample efficiency of PER. This combination is specifically tailored to handle the stringent latency constraints of URLLC in dynamic and non-stationary HAA backhaul contexts. • Constraint-Aware Design: Unlike existing DRL-based slicing literature that often bypasses or simplifies actionspace constraints, we introduce a double soft-max projection mechanism. This provides a theoretically grounded solution to the action-space constraint problem, ensuring that resource allocations remain physically feasible while maintaining continuous gradients for efficient learning. • Comprehensive QoS Support: In contrast to works focused on a single performance metric, our framework simultaneously optimizes rates, latency, and success rates for heterogeneous services. This multi-objective optimization is reflected in the QoS Metrics column of Table I, demonstrating a more holistic approach to 6G service requirements. • Adaptive Priority Execution: To overcome the limitations of offline static slicing, we introduce the RAPO mechanism. This provides an adaptive framework to gracefully scale down best-effort services and protect mission-critical links under severe resource scarcity, an operational flexibility largely absent in prior studies. By addressing the two-tier coupling and the action-space constraint simultaneously, the proposed RAPO-TD3 framework armed with RAPO offers a robust, continuous-control solution for the next generation of aerial backhaul networks.
4
TABLE I C OMPARISON OF R ELATED L ITERATURE Paper / Perspective
Slicing Hierarchy
[25] (2018)
Single-layer (Core)
[13] (2023) [22] (2023)
IAB Integrated Single-layer (HAA)
[10] (2024)
Hybrid Air-Ground
[32] (2026)
Single-layer (Edge)
Device Selection, Bandwidth
[33] (2026)
Three-layer (Collaborative) Dual-layer: Core + HAA Sub-slice
Computing, Energy Bandwidth (Access + Backhaul)
Proposed Scheme
Resource Type Bandwidth (Licensed / Unlicensed) Bandwidth, Backhaul Slot Bandwidth, HAA Position Bandwidth (Access + Backhaul)
QoS Metrics Aggregate Utility URLLC Latency SLA Satisfaction (Rate) eMBB Throughput and Coverage URLLC Latency, Convergence Time Delay, Energy Consumption Rate, Latency, and Success Rate
Algorithm Game Theory / Distributed Algo. MATD3 DQN Genetic-based Deployment BSAC DN-TD3 RAPO-TD3
III. S YSTEM M ODEL AND P ROBLEM F ORMULATION As depicted in Fig. 1, the studied system operates under a stratified dual-layer slicing framework consisting of a centralized macro slicing orchestrator (MSO) and an aerial radio access network. The MSO partitions the total bandwidth among HAAs on a macroscopic time scale based on aggregate demands, while each HAA subsequently distributes its allocated share among three local service slices on a sub-millisecond execution scale. This hierarchical structure mathematically decouples the spatial-temporal optimization scales, enabling finegrained, real-time adaptation to non-stationary traffic while mitigating the state-action dimensionality explosion inherent in flat centralized frameworks. Fig. 1. HAA-assisted backhaul network with dual-layer slicing architecture.
A. Network Architecture We consider a 6G-oriented HAA-assisted downlink backhaul network comprising a macro base station hosting the MSO and a set of 𝑁 HAAs indexed by 𝑛 = 1, 2, . . . , 𝑁. To support diverse QoS requirements, the network is partitioned into three primary network slices: eMBB, URLLC, and mMTC. To mitigate terrestrial non-line-of-sight propagation blockages from urban terrain obstructions, HAAs are deployed at strategic altitudes to maintain clear line-of-sight links with the macro base station. This architecture adopts a hierarchical control paradigm, operating the macro base station strictly as a centralized macro orchestrator to avoid high-frequency signal penetration losses and minimize core network computational overhead. Furthermore, by employing infrastructuregrade tethered HAAs, the system not only secures continuous power supply through physical tethers, thereby eliminating transient battery depletion limitations, but also gains high structural stability against aerial wind disturbances. This robust physical anchoring guarantees stable line-of-sight backhaul links, justifying their deployment as reliable quasi-stationary relays. During backhaul service delivery, the HAAs maintain a quasi-stationary profile, treating their spatial coordinates as fixed boundary parameters. Building on this stable operational assumption, aerodynamic flight paths, three-dimensional trajectory optimization, and altitude adjustments are decoupled from the resource allocation problem. Furthermore, propulsion power consumption and flight energy limitations are likewise isolated from the operational action space. This structural
configuration ensures that the dual-layer slicing framework can focus entirely on multi-service queueing dynamics and spectrum efficiency during the localized execution window.
B. Heterogeneous Air-to-Ground Channel Model with Spatial Abstraction To characterize the spatial heterogeneity of the aerial transport stratum, we consider a network layout where 𝑁 HAAs are deployed at a constant operational altitude 𝐻haa . The horizontal distance between the centralized macro base station and the 𝑛-th HAA is denoted by 𝑟 𝑛 , yielding a deterministic q 2 + 𝑟2. three-dimensional propagation distance of 𝑑 𝑛 = 𝐻haa 𝑛 Following the standard specifications for dense urban environments [34], the line-of-sight (LoS) probability between the centralized orchestrator and the 𝑛-th HAA is dynamically governed by the spatial elevation angle 𝜑 𝑛 = arctan(𝐻haa /𝑟 𝑛 )× 180 𝜋 , expressed as: 𝑃LoS (𝜑 𝑛 ) =
1 , 1 + 𝑎 env exp [−𝑏 env (𝜑 𝑛 − 𝑎 env )]
(1)
where 𝑎 env and 𝑏 env represent environment-specific constants that dictate the steepness and the offset of the propagation probability curve corresponding to the local urban building densities [34]. Consequently, the macro-scale geometric path
5
loss 𝑃𝐿 𝑛 (in linear scale) is formulated as the expectation over LoS and non-line-of-sight (NLoS) states: 𝑃𝐿 𝑛 = 𝑃LoS (𝜑 𝑛 )10
𝑃𝐿LoS (𝑑𝑛 ) 10
+ [1 − 𝑃LoS (𝜑 𝑛 )] 10
𝑃𝐿NLoS (𝑑𝑛 ) 10
, (2) where the components 𝑃𝐿 LoS (𝑑 𝑛 ) and 𝑃𝐿 NLoS (𝑑 𝑛 ) incorporate free-space path loss alongside shadowing mid-scale attenuations at a designated carrier frequency 𝑓𝑐 : 𝑃𝐿 LoS (𝑑 𝑛 ) = 20 log10 (𝑑 𝑛 ) + 20 log10 ( 𝑓𝑐 ) + 𝐶FSPL + 𝜂LoS , (3) 𝑃𝐿 NLoS (𝑑 𝑛 ) = 20 log10 (𝑑 𝑛 ) + 20 log10 ( 𝑓𝑐 ) + 𝐶FSPL + 𝜂NLoS (4) with 𝐶FSPL denoting the free-space reference constant derived from the speed of light. The parameters 𝜂LoS and 𝜂NLoS signify the average additional attenuation factors for LoS and NLoS links, respectively, which explicitly capture the structural penetration and obstacle reflection losses inherent to the 3D urban geometry [34]. To guarantee compliance with reinforcement learning convergence boundaries without undermining physical rigor, the instantaneous channel gain ℎ 𝑛 (𝑡) at a discrete time slot 𝑡 governing the backhaul transport capability is modeled as a composite framework merging deterministic path loss and fastvarying small-scale fading: ℎ 𝑛 (𝑡) = Ω𝑛 (𝑡) · 𝜁 𝑛 ,
(5)
where 𝑡 indexes the discrete operational intervals, Ω𝑛 (𝑡) models the instantaneous small-scale multi-path fading profile governed by a Rayleigh distribution [35], and 𝜁 𝑛 represents the normalized large-scale spatial structural parameter mapping the path-loss heterogeneity onto the unified agent observation space: −1/2 𝑃𝐿 (6) 𝜁 𝑛 = Í 𝑛 −1/2 . 𝑁 1 𝑃𝐿 𝑗 𝑗=1 𝑁 C. Fronthaul Abstraction and Demand Modeling
A major architectural feature of the proposed hierarchical management plane is the functional decoupling of lowertier edge fronthaul access from upper-tier downlink backhaul transport provisioning. The system operates over a discretetime horizon indexed by 𝑡 = 1, . . . , 𝑇, where each slot corresponds to a centralized macro control period 𝑇𝐶 . At the beginning of each interval 𝑡, the MSO collects state information, executes global resource slicing orchestration, and disseminates decisions for localized enforcement. Crucially, the index 𝑡 denotes this macro-scale orchestration interval; once the sub-slice ratios g𝑖 (𝑡) are determined at the start of slot 𝑡, each HAA executes these assigned bandwidth proportions continuously across localized, sub-millisecond physical resource block scheduling loops throughout the entire duration of 𝑇𝐶 . This structure strictly isolates the global optimization step from high-frequency fronthaul execution, which mathematically resolves time-scale separation without conflating the indices. We assume that microscopic user equipment association, intra-cell proportional fair scheduling, and physical-layer radio
resource block tiling are executed autonomously at the network edge by decentralized base station components embedded within each HAA. From the operational perspective of the centralized MSO, these dynamic edge radio dynamics are systematically compressed and filtered. The decentralized edge mechanisms map the instantaneous local user traffic states into an aggregated, multi-service exogenous downlink traffic demand profile vector for each HAA 𝑛 at control period 𝑡, denoted as: ⊤ 𝚫𝑛 (𝑡) = Δ𝑛,𝑒 (𝑡), Δ𝑛,𝑢 (𝑡), Δ𝑛,𝑚 (𝑡) , (7)
where Δ𝑛,𝑒 (𝑡), Δ𝑛,𝑢 (𝑡), and Δ𝑛,𝑚 (𝑡) represent the aggregated downlink traffic demands for eMBB, URLLC, and mMTC, respectively. Simultaneously, each HAA reports its instantaneous backhaul channel measurement |ℎ 𝑛 (𝑡)| along with this aggregated demand vector 𝚫𝑛 (𝑡) back to the MSO via the dedicated backhaul control channel. To quantify the traffic pressure experienced by each HAA within every control period 𝑇𝐶 across the three heterogeneous service types, the downlink traffic demand is mathematically characterized by the total generated data volume expressed in bits. Specifically, the demand observed at the 𝑛-th HAA for service type 𝑖 ∈ {𝑒, 𝑢, 𝑚} is parameterized as the summation of downlink user-plane packet arrivals routed from the core network: Õ 𝑉𝑘,𝑖 (𝑡), (8) Δ𝑛,𝑖 (𝑡) = 𝑘 ∈ K𝑛,𝑖
where K𝑛,𝑖 denotes the set of active user devices belonging to service type 𝑖 associated with HAA 𝑛, and 𝑉𝑘,𝑖 (𝑡) represents the discrete data volume generated for user 𝑘 during the current control interval. To guarantee compliance with physical causality, the traffic demand vector 𝚫𝑛 (𝑡) observed by the centralized orchestrator at the exact start of control period 𝑡 represents the accumulated traffic volume that has arrived from the physical core network and populated the transmission buffers during the immediately preceding interval. Therefore, the centralized orchestrator operates strictly on materialized, deterministic queue states awaiting backhaul transport transmission rather than predicting future stochastic traffic bursts within slot 𝑡, ensuring complete mathematical causality in realtime execution. To eliminate observation bias caused by structural spatial differences in the user population densities across distinct HAAs, the framework normalizes the absolute bit load into a dimensionless demand ratio: Δ𝑛,𝑖 (𝑡) , 𝑖 ∈ {𝑒, 𝑢, 𝑚}. (9) 𝛿 𝑛,𝑖 (𝑡) = Í 𝑗 ∈ {𝑒,𝑢,𝑚} Δ𝑛, 𝑗 (𝑡) Here, 𝛿 𝑛,𝑖 (𝑡) explicitly defines the fractional load contribution of service-type-𝑖 demand observed Í at HAA 𝑛, satisfying the probability simplex constraint 𝑖∈ {𝑒,𝑢,𝑚} 𝛿 𝑛,𝑖 (𝑡) = 1. The complete vectorized demand state mapping is thus defined as δ𝑛 (𝑡) = [𝛿 𝑛,𝑒 (𝑡), 𝛿 𝑛,𝑢 (𝑡), 𝛿 𝑛,𝑚 (𝑡)] ⊤ . This aggregated vector follows non-stationary distribution boundaries parameterized by macro diurnal tidal trends and stochastic Poisson burst multipliers. By capturing the timevarying multi-service load at each HAA through this macro abstraction, the system accurately feeds the DRL controller.
6
where 𝑓𝑛,𝑖 (𝑡) ∈ [0, 1]. Suppose the total available bandwidth is 𝐵total . Then, the bandwidth allocated to service type 𝑖 at HAA 𝑛 during control period 𝑡 is given by:
MSO
𝐵𝑛,𝑖 (𝑡) = 𝑓𝑛,𝑖 (𝑡) · 𝐵total ,
∀𝑖 ∈ {𝑒, 𝑢, 𝑚}.
(15)
E. QoS Satisfaction Model
Fig. 2. The proposed dual-layer resource slicing architecture orchestrated by the macro slice orchestrator (MSO). The core-layer allocation partitions the total wireless backhaul bandwidth into service-specific macro-slices (𝑆𝑒 (𝑡 ), 𝑆𝑢 (𝑡 ), and 𝑆𝑚 (𝑡 )). Subsequently, the HAA-layer sub-slice allocation dynamically distributes these isolated resources across spatially deployed HAAs (e.g., 𝑔𝑛,𝑒 (𝑡 ) for the eMBB slice) to serve localized traffic demands.
Isolating the centralized state-space from high-frequency fronthaul user mutations successfully bypasses the dimensionality curse, ensuring that the global core network management loop remains scalable, stable, and mathematically tractable. D. Dual-Layer Slice Resource Allocation Model As shown in Fig. 2, let 𝑆𝑖 (𝑡) denote the macro-layer slice ratio allocated to service type 𝑖 ∈ {𝑒, 𝑢, 𝑚} at control period 𝑡. The vector S(𝑡) = [𝑆 𝑒 (𝑡), 𝑆 𝑢 (𝑡), 𝑆 𝑚 (𝑡)] must satisfy the following constraints: 𝑆𝑖 (𝑡) ≥ 0, ∀𝑖 ∈ {𝑒, 𝑢, 𝑚}, Õ 𝑆𝑖 (𝑡) = 1.
(10) (11)
𝑖∈ {𝑒,𝑢,𝑚}
Here, constraint (10) ensures that each slice ratio is nonnegative, while constraint (11) guarantees that the total allocated slice ratios sum to unity, reflecting the complete partitioning of available resources at the macro layer. At the HAA layer, each HAA 𝑛 further divides its allocated macro-layer slice 𝑆𝑖 (𝑡) among its locally served users of service type 𝑖. Let 𝑔𝑛,𝑖 (𝑡) denote the HAA-layer sub-slice ratio for service type 𝑖 at HAA 𝑛 during control period 𝑡. For each service type 𝑖 ∈ {𝑒, 𝑢, 𝑚}, the vector g𝑖 (𝑡) = [𝑔1,𝑖 (𝑡), 𝑔2,𝑖 (𝑡), . . . , 𝑔 𝑁 ,𝑖 (𝑡)] must satisfy: 𝑔𝑛,𝑖 (𝑡) ≥ 0, ∀𝑛 = 1, 2, . . . , 𝑁, 𝑁 Õ 𝑔𝑛,𝑖 (𝑡) = 1, ∀𝑖 ∈ {𝑒, 𝑢, 𝑚}.
(12) (13)
𝑛=1
Here, constraint (12) ensures that each HAA’s sub-slice ratio is non-negative, while constraint (13) guarantees that the total sub-slice ratios across all HAAs for each service type sum to unity, reflecting the complete distribution of macro-layer slices at the aerial layer. With (11) and (13), the effective resource proportion allocated to service type 𝑖 at HAA 𝑛 during control period 𝑡 can be expressed as: 𝑓𝑛,𝑖 (𝑡) = 𝑆 𝑖 (𝑡) · 𝑔 𝑛,𝑖 (𝑡),
∀𝑖 ∈ {𝑒, 𝑢, 𝑚},
(14)
Let 𝑁 denote the number of HAAs and 𝑡 denote the index of the control time slot. For HAA 𝑛 (𝑛 = 1, 2, . . . , 𝑁), the eMBB throughput 𝑅𝑛,𝑒 (𝑡) during time slot 𝑡 is calculated using the Shannon capacity theorem as: 𝑃𝑛 |ℎ 𝑛 (𝑡)| 2 , (16) 𝑅𝑛,𝑒 (𝑡) = 𝐵𝑛,𝑒 (𝑡) log2 1 + 𝑁0 𝐵𝑛,𝑒 (𝑡)
where 𝐵𝑛,𝑒 (𝑡) is the bandwidth allocated to eMBB at HAA 𝑛, 𝑃𝑛 denotes the constant directional downlink transmit power allocated by the macro base station to the backhaul link of HAA 𝑛, |ℎ 𝑛 (𝑡)| 2 is the instantaneous channel gain, and 𝑁0 is the noise power spectral density. To accurately capture the transmission dynamics of short packets in mission-critical scenarios, the URLLC transmission rate 𝑅𝑛,𝑢 (𝑡) is modeled via the finite-blocklength approximation [36] as: " 𝑅𝑛,𝑢 (𝑡) = 𝐵𝑛,𝑢 (𝑡) log2 1 + Γ𝑛,𝑢 (𝑡) −
s
# 𝑉 (Γ𝑛,𝑢 (𝑡)) 𝑄 −1 (𝜖𝑢 ) , 𝐿 𝑛,𝑢 ln 2
(17)
2
𝑛 (𝑡 ) | where Γ𝑛,𝑢 (𝑡) = 𝑃𝑁𝑛0 |ℎ 𝐵𝑛,𝑢 (𝑡 ) represents the instantaneous signalto-noise ratio (SNR), 𝑉 (Γ𝑛,𝑢 (𝑡)) = 1 − (1 + Γ𝑛,𝑢 (𝑡)) −2 denotes the channel dispersion, 𝐿 𝑛,𝑢 is the URLLC packet blocklength, 𝑄 −1 (·) signifies the inverse Q-function, and 𝜖𝑢 bounds the decoding error probability. To establish the cross-layer mapping with the user-plane traffic volume 𝑉𝑘,𝑢 (𝑡) introduced in (8), the parameter 𝐿 𝑛,𝑢 represents the standardized physical transportblock size into which this arriving data volume 𝑉𝑘,𝑢 (𝑡) is segmented for backhaul transmission. Unlike the volatile, time-varying traffic volume 𝑉𝑘,𝑢 (𝑡), the physical blocklength 𝐿 𝑛,𝑢 is a time-invariant protocol configuration predetermined during network initialization to eliminate dynamic packetformatting signaling overhead and guarantee deterministic processing delays. Furthermore, by characterizing the edge buffer of each HAA as an M/M/1 queueing system [37], the total URLLC latency 𝐷 𝑛,𝑢 (𝑡) is rigorously formulated to represent the complete system sojourn time, which inherently encompasses both the queueing wait time and the physical transmission delay:
𝐷 𝑛,𝑢 (𝑡) =
1 , 𝜇 𝑛,𝑢 (𝑡) − 𝜆 𝑛,𝑢 (𝑡)
(18)
𝑅𝑛,𝑢 (𝑡 ) denotes the corresponding service 𝐿𝑛,𝑢 𝑁pkt,𝑛,𝑢 (𝑡 ) rate, and 𝜆 𝑛,𝑢 (𝑡) = represents the average packet 𝑇𝐶 Δ (𝑡 ) arrival rate during the control period, with 𝑁pkt,𝑛,𝑢 (𝑡) = 𝑛,𝑢 𝐿𝑛,𝑢
where 𝜇 𝑛,𝑢 (𝑡) =
signifying the number of standardized physical packets generated at HAA 𝑛 for URLLC services within time slot 𝑡.
7
The mMTC success rate 𝐺 𝑛,𝑚 (𝑡) at HAA 𝑛 during time slot 𝑡 is defined as the fraction of successfully received packets 𝑚 within the maximum tolerable delay 𝑇max and maximum retransmissions 𝐾max : 𝐺 𝑛,𝑚 (𝑡) =
1 𝑁pkt,𝑛,𝑚 (𝑡)
𝑁pkt,𝑛,𝑚 Õ (𝑡 )h
i
𝑚 𝑇 𝑗 ≤ 𝑇max ∧ RTX 𝑗 ≤ 𝐾max ,
𝑗=1
(19) Δ (𝑡 ) where 𝑁pkt,𝑛,𝑚 (𝑡) = 𝑛,𝑚 is the total number of physical 𝐿𝑛,𝑚 packets generated at HAA 𝑛 for mMTC services, 𝑇 𝑗 is the delay of packet 𝑗, and RTX 𝑗 is the number of retransmissions for packet 𝑗. To contextualize the scale relationship, the packet𝑚 < 𝑇 , meaning that the level latency threshold satisfies 𝑇max 𝐶 microscopic transmission and retransmission processes of all packets generated within slot 𝑡 are fully completed at the HAA edge plane before the MSO initiates the next macro control period. Structurally, the total delay 𝑇 𝑗 is analyzed as the summation of the over-the-air transmission delay, propagation delay, and accumulated backoff waiting times across multiple attempts. Furthermore, the retransmission counter RTX 𝑗 is modeled via a standard automatic repeat request (ARQ) protocol mechanism. Each decoding failure at the receiving end triggers a retransmission attempt and increments RTX 𝑗 by 1. If successful decoding is not achieved within RTX 𝑗 ≤ 𝐾max 𝑚 , the packet is or if the cumulative delay exceeds 𝑇 𝑗 > 𝑇max permanently dropped, thereby reducing the aggregated success rate 𝐺 𝑛,𝑚 (𝑡). To quantify whether each service meets or violates its service level agreement (SLA) threshold, we convert the continuous QoS indicators into normalized satisfaction levels in the interval [0, 1]. For HAA 𝑛 at control period 𝑡, the satisfaction functions corresponding to the observed throughput, latency, and success rate are defined as: ( 1, 𝑅𝑛,𝑒 (𝑡) ≥ 𝑅min,𝑒 , 𝑟 𝑛,𝑒 (𝑡) = (20) 0, 𝑅𝑛,𝑒 (𝑡) < 𝑅min,𝑒 , ( 1, 𝐷 𝑛,𝑢 (𝑡) ≤ 𝐷 max,𝑢 , 𝑟 𝑛,𝑢 (𝑡) = (21) 0, 𝐷 𝑛,𝑢 (𝑡) > 𝐷 max,𝑢 , ( 1, 𝐺 𝑛,𝑚 (𝑡) ≥ 𝐺 min,𝑚 , 𝑟 𝑛,𝑚 (𝑡) = (22) 0, 𝐺 𝑛,𝑚 (𝑡) < 𝐺 min,𝑚 . Here, 𝑅min,𝑒 denotes the minimum acceptable eMBB rate, 𝐷 max,𝑢 is the maximum tolerable URLLC latency, and 𝐺 min,𝑚 is the minimum acceptable packet success rate for mMTC. Finally, the global time-average satisfaction of each service type is evaluated after resource allocation is completed over 𝑇 consecutive periods. The satisfaction values are aggregated across all 𝑁 HAAs and normalized by the total number of observations 𝑁𝑇, yielding the average performance over the entire observation window expressed as follows: 𝑇
𝑁
𝑟¯𝑒 =
1 ÕÕ 𝑟 𝑛,𝑒 (𝑡), 𝑁𝑇 𝑡=1 𝑛=1
𝑟¯𝑢 =
1 ÕÕ 𝑟 𝑛,𝑢 (𝑡), 𝑁𝑇 𝑡=1 𝑛=1
𝑇
(23)
𝑁
(24)
𝑇
𝑟¯𝑚 =
𝑁
1 ÕÕ 𝑟 𝑛,𝑚 (𝑡). 𝑁𝑇 𝑡=1 𝑛=1
(25)
F. Problem Formulation In this study, we model the resource allocation between the macro layer and the aerial layer as a two-tier optimization problem. The macro layer determines the allocation proportions of the three service types (eMBB, URLLC, and mMTC) from the overall resource pool, while the HAA layer further refines the resource proportions for each service type according to the macro-layer decision. Specifically, the macro-layer ratio 𝑆𝑖 (𝑡) determines the global portion assigned to each service type, and the HAAlayer ratio 𝑔𝑛,𝑖 (𝑡) further distributes the corresponding resources among HAAs from (10) to (15). This multi-level allocation ensures a reasonable distribution of resources for each service type while accounting for the practical demands and capabilities of HAAs. Based on this, we incorporate the satisfaction functions 𝑟 𝑛,𝑒 , 𝑟 𝑛,𝑢 , and 𝑟 𝑛,𝑚 defined from (16) to (19), and examine whether the QoS requirements are satisfied. This work aims to jointly maximize the overall satisfaction of the three service types while ensuring that all allocation ratios satisfy the physical boundaries. Therefore, the proposed optimization framework is formulated as: Õ 𝑟¯𝑖 , (P1) max S(𝑡 ),g𝑖 (𝑡 )
s.t.
∀𝑖∈ {𝑒,𝑢,𝑚}
(10), (11), (12), (13).
where S(𝑡) is the macro-layer ratio vector and g𝑖 (𝑡) is the HAA-layer ratio vector. Note that the total backhaul capacity of the centralized macro base station is implicitly bounded by the total available bandwidth 𝐵total . Because the dual-layer resource orchestration constraints strictly enforce Í𝑁 Í 𝑛=1 𝑖∈ {𝑒,𝑢,𝑚} 𝑓 𝑛,𝑖 (𝑡) = 1, the aggregate throughput of the entire network is mathematically capped by the physical spectrum limit, preventing any backhaul resource violation or unfeasible allocation over the transport stratum. The objective function in (P1) maps the heterogeneous service performance indicators from (20) to (22) onto a unified evaluation axis to maximize SLA compliance across all HAAs. However, the non-stationary 6G traffic dynamics and strict dual-layer resource coupling render (P1) a nonlinear fractional multi-dimensional knapsack problem, which is strictly NP-hard. Furthermore, the threshold-triggered SLA boundaries introduce severe non-smoothness and discontinuity into the objective space, invalidating standard gradient-based optimization solvers. Traditional numerical optimization methods, such as branch-and-bound or iterative convex relaxations, are physically prohibitive in this context. They require perfect deterministic environmental foresight and incur excessive execution latencies that violate the millisecond-level coherence time of URLLC services. Consequently, these traditional solvers are excluded from the scope of this study, as they cannot be deployed for online real-time execution in non-stationary aerial environments.
8
IV. T HE P ROPOSED RAPO-TD3 F RAMEWORK To solve the formulated non-convex MINLP problem (P1), we propose an intelligent resource management framework based on the TD3 algorithm [38]. This choice is motivated by the capability of TD3 to mitigate policy overestimation bias in continuous action spaces, which is critical for maintaining training stability in high-dimensional HAA environments. A. DRL Component Definitions We define the state space, action space, and reward function to capture the dynamics of the cognitive backhaul network. 1) State Space Design: To provide the TD3 agent with complete network state feedback, we transition from a basic environment snapshot to a comprehensive 61-dimensional state vector s𝑡 . The design philosophy of s𝑡 is to encapsulate immediate physical-layer dynamics, strategic network memory, SLA boundaries, and temporal regularities. The global state vector is structured as: 𝑁 𝑁 , ψ𝑡 , (26) s𝑡 = {o𝑛,𝑡 } 𝑛=1 , S(𝑡 − 1), {g𝑛 (𝑡 − 1)} 𝑛=1
where 𝑁 denotes the number of HAAs. The state vector consists of three functional modules: • Multi-Dimensional Environmental Observations: For each HAA indexed by 𝑛 = 1, . . . , 𝑁, an 11-dimensional observation vector o𝑛,𝑡 encapsulates localized network dynamics. To ensure scale uniformity, it incorporates the dimensionless multi-service demand ratios 𝛿 𝑛,𝑒 (𝑡), 𝛿 𝑛,𝑢 (𝑡), and 𝛿 𝑛,𝑚 (𝑡) comprising the vectorized demand state δ𝑛 (𝑡), alongside the composite instantaneous backhaul channel gain |ℎ 𝑛 (𝑡)|. The remaining seven dimensions encode static service level agreement boundaries to contextualize volatile traffic loads against contract thresholds, preventing feature dominance and gradient saturation. To preserve the fundamental Markov property without exponentially expanding the state dimensionality via recurrent networks, the single-step feedback loop serves as a mathematically sufficient statistic to guarantee policy continuity. • Hierarchical Decision Feedback: To handle interdependencies within the dual-layer architecture, the agent tracks its historical decisions. This 15-dimensional component consists of the preceding macro-layer slice ratios S(𝑡 − 1) (3 dimensions) and the individual HAA-layer sub-slice weights g𝑛 (𝑡 − 1) for 𝑛 = 1, . . . , 𝑁 (12 dimensions). Rather than utilizing a longer temporal history window or recurrent networks which would exponentially expand the state dimensionality and violate the fundamental Markov property of the optimization framework, this single-step 1-indexed feedback loop serves as a mathematically sufficient statistic. It provides the agent with necessary strategic memory to ensure policy continuity and smooth incremental adjustments, effectively preventing tracking oscillations during real-time execution. • Cyclic Temporal Context: To model 24-hour periodic traffic variations, a 2-dimensional cyclic temporal feature ψ𝑡 = [sin(2𝜋𝑡/24), cos(2𝜋𝑡/24)] is integrated. This trigonometric encoding enables proactive slice priority adaptation before anticipated peak traffic surges.
2) Action Space and Two-Tier Projection: The action vector a𝑡 represents the multi-stage decision-making of the resource management framework, encompassing both macro-layer provisioning and HAA-layer spatial distribution. The continuous action space features a total dimensionality of 3 + 3𝑁, expressed as: a𝑡 = zcore (𝑡), z1 (𝑡), . . . , z 𝑁 (𝑡) , (27) where zcore (𝑡) ∈ R3 corresponds to the raw outputs for the macro-layer slices, and z𝑛 (𝑡) = [𝑧 𝑛,𝑒 , 𝑧 𝑛,𝑢 , 𝑧 𝑛,𝑚 ] denotes the distribution weights for service types at the 𝑛-th HAA. To ensure that the agent actions strictly adhere to the bandwidth capacity constraints in (11) and (13) without sacrificing the differentiability of the network, we implement a double soft-max projection mechanism. This nested transformation converts the raw actor network outputs into physically feasible allocation ratios through the following stages: • Macro-Layer Slice Projection: The first three elements zcore (𝑡) are processed via a soft-max function to obtain the global slice ratios 𝑆𝑖 (𝑡) for eMBB, URLLC, and mMTC services. To prevent resource starvation and preserve dense gradient propagation, we apply a linear scaling: 𝑆𝑖 (𝑡) = (1 − 3𝜌) · soft-max(𝑧𝑖 (𝑡)) + 𝜌,
(28)
where 𝑧𝑖 (𝑡) denotes the continuous action value within zcore (𝑡), and 𝜌 represents the minimum reserved slice ratio. This stage ensures structural compliance with constraint (11). • HAA-Layer Distribution Projection: For each HAA 𝑛, the agent determines the relative spatial weights across the three services. By applying a local soft-max operation over z𝑛 (𝑡), the framework standardizes the spatial priorities and ensures that the sub-slice ratios 𝑔 𝑛,𝑖 (𝑡) for each service type sum to unity across all HAAs, thus satisfying constraint (13). • Deterministic Resource Mapping: The final bandwidth 𝐵𝑛,𝑖 (𝑡) allocated to slice 𝑖 at UAV 𝑛 is determined by (15) using the effective resource proportion 𝑓𝑛,𝑖 (𝑡) from (14). To clarify the technical necessity of the Double Soft-max Projection layer, we explicitly analyze the dual-layer coupling constraints that cause standard independent projections to fail. Let u(𝑡) denote the macro-tier allocation vector across HAAs and v𝑛 (𝑡) represent the micro-tier service slice distribution vector localized within HAA 𝑛. The joint resource orchestration action must satisfy a rigid hierarchical conservation law, where the absolute physical capacity allocated to a specific network slice is bounded by the multiplicative product of these two vectors. If standard independent softmax layers were utilized, the policy network would output separate vectors lacking physical scaling coherence. Enforcing the joint boundary via post-hoc clipping or rule-based truncation introduces non-differentiable operations that break the continuous policy gradient flow, leading to severe training instability. By contrast, the Double Soft-max Projection layer implements a nested conditionally dependent tensor mapping that explicitly enforces the joint dual-layer simplex constraint by construction. This specialized mathematical architecture preserves continuous differentiability across both resource strata, allowing exact
9
policy gradients to backpropagate seamlessly through the duallayer parameters to accelerate convergence.
Therefore, the system penalty 𝑃system (𝑡) is purely dedicated to penalizing severe URLLC latency violations, ensuring that mission-critical links are safeguarded:
B. Reward Function Design and RAPO Mechanism
𝑃system (𝑡) = 𝜛
To guide the agent toward an effective policy that balances heterogeneous service satisfaction with physical feasibility, we design a reward function characterized by differentiable utility shaping and the RAPO mechanism. This approach addresses the sparse-gradient issues inherent in binary QoS indicators while ensuring strict adherence to mission-critical constraints under volatile workloads. 1) Differentiable Satisfaction Functions: To facilitate effective gradient-based learning, we replace conventional binary QoS indicators with differentiable, piecewise-linear satisfaction functions 𝑟ˆ𝑛,𝑖 (·) ∈ [0, 1]. This design allows the agent to receive continuous feedback even when the SLA requirements are not fully met. • eMBB Satisfaction: For eMBB services, the satisfaction
level is driven by the achieved data rate 𝑅𝑛,𝑒 (𝑡). The utility is mapped between the minimum survival rate 𝑅min,𝑒 and the target request rate 𝑅req,𝑒 as follows: 0, 𝑅𝑛,𝑒 (𝑡 ) −𝑅min,𝑒 𝑟ˆ𝑛,𝑒 (𝑡) = 𝑅req,𝑒 −𝑅min,𝑒 , 1,
𝑅𝑛,𝑒 (𝑡) ≤ 𝑅min,𝑒 , 𝑅min,𝑒 < 𝑅𝑛,𝑒 (𝑡) ≤ 𝑅req,𝑒 , 𝑅𝑛,𝑒 (𝑡) > 𝑅req,𝑒 . (29)
• URLLC Satisfaction: Given the latency-sensitive nature
of URLLC, the satisfaction function is defined by the observed delay 𝐷 𝑛,𝑢 (𝑡). The utility remains at its peak for delays below the target 𝐷 req,𝑢 and drops to zero once the delay exceeds the maximum tolerable threshold 𝐷 max,𝑢 : 0, 𝐷max,𝑢 −𝐷𝑛,𝑢 (𝑡 ) 𝑟ˆ𝑛,𝑢 (𝑡) = 𝐷max,𝑢 −𝐷req,𝑢 , 1,
𝐷 𝑛,𝑢 (𝑡) > 𝐷 max,𝑢 , 𝐷 req,𝑢 < 𝐷 𝑛,𝑢 (𝑡) ≤ 𝐷 max,𝑢 ,
1, 𝐺 (𝑡 ) −𝐺min,𝑚 , 𝑟ˆ𝑛,𝑚 (𝑡) = 𝐺𝑛,𝑚 req,𝑚 −𝐺min,𝑚 0,
𝐺 𝑛,𝑚 (𝑡) ≥ 𝐺 req,𝑚 , 𝐺 min,𝑚 ≤ 𝐺 𝑛,𝑚 (𝑡) < 𝐺 req,𝑚 ,
𝐷 𝑛,𝑢 (𝑡) ≤ 𝐷 req,𝑢 . (30)
• mMTC Satisfaction: For mMTC services, the satisfaction
level is determined by the packet success rate 𝐺 𝑛,𝑚 (𝑡). The utility reflects the reliability of massive device connectivity between the minimum acceptable level 𝐺 min,𝑚 and the target reliability 𝐺 req,𝑚 :
𝐺 𝑛,𝑚 (𝑡) < 𝐺 min,𝑚 . (31)
2) Global Reward and Differentiable Penalties: While conventional approaches rely on heuristic penalty terms to discourage boundary violations, they often introduce nondifferentiable artifacts. In contrast, our double soft-max projection layer enforces simplex constraints by construction. Consequently, the formulation inherently obviates the need for explicit bandwidth boundary penalties, thereby preserving the gradient continuity required for stable convergence.
𝑁 Õ 𝐷 𝑛,𝑢 (𝑡) − 𝐷 max,𝑢 + 𝑛=1
𝐷 max,𝑢
(32)
where 𝜛 signifies the URLLC delay penalty coefficient regulating the scaling severity of delay violations, and (·) + denotes the positive part function. Without loss of generality, 𝜛 is parameterized as 1.0 to serve as the normalized reference baseline for the penalty domain, thereby enabling the reward weight vectors ωbase and ωalert to be calibrated systematically relative to this unified penalty scaling. The global reward 𝑟 (𝑡) observed by the agent at time slot 𝑡 is then formulated as the weighted aggregate of slice satisfactions minus the timevarying penalty: 𝑟 (𝑡) =
𝑁 Õ
Õ
𝜔𝑖 (𝑡) · 𝑟ˆ𝑛,𝑖 (𝑡) − 𝜒(𝑡) · 𝑃system (𝑡),
(33)
𝑛=1 𝑖∈ {𝑒,𝑢,𝑚}
where 𝜔𝑖 (𝑡) represents the dynamic service weights governed by the RAPO mechanism, and 𝜒(𝑡) is the dynamic penalty scaling factor. To prioritize exploration in early training while ensuring strict constraint adherence during convergence, we implement a cyclical penalty factor 𝜒(𝑡) defined as: 𝐸 𝑃current 𝜒(𝑡) = 𝜒init + 𝜒grow × 𝐸 𝑃total 𝐸 𝑃current , (34) × 1 + 𝜒osc · sin 2𝜋 · 𝑇osc where 𝐸 𝑃current and 𝐸 𝑃total are the current and total episode indices, respectively. The parameters { 𝜒init , 𝜒grow , 𝜒osc , 𝑇osc } control the baseline, growth rate, oscillation amplitude, and oscillation period. 3) Resilient Adaptive Priority Orchestration Mechanism: A distinctive architectural feature of our framework is the RAPO mechanism governed by an adaptive smooth gating structure. This enables the agent to shift its priorities fluidly under extreme network pressure without inducing training oscillations. Based on the real-time penalty magnitude, the system autonomously transitions between two conceptual operational modes: • Balanced Mode: When the network is stable (𝑃system (𝑡) ≤
𝑃thresh ), the system adopts the base weight vector ωbase to maximize eMBB utility while maintaining baseline URLLC stability. • Alert Mode: Upon detecting significant QoS violations (𝑃system (𝑡) > 𝑃thresh ), the system prioritizes the alert weight vector ωalert to strictly safeguard ultra-reliable transmissions. Instead of an abrupt, non-differentiable mode switch, the dynamic weight vector ω(𝑡) = [𝜔 𝑒 (𝑡), 𝜔𝑢 (𝑡), 𝜔 𝑚 (𝑡)] ⊤ is derived by seamlessly blending the base and alert mode vectors: ω(𝑡) = (1 − 𝜉 (𝑡)) · ωbase + 𝜉 (𝑡) · ωalert ,
(35)
10
where 𝜉 (𝑡) ∈ (0, 1) serves as the adaptive interpolation factor. To ensure a smooth transition, 𝜉 (𝑡) is modeled via a sigmoid function evaluated against the instantaneous system penalty: 1 . 𝜉 (𝑡) = (36) 𝑃system (𝑡 ) − 𝑃thresh 1 + exp − 𝜅
Here, 𝑃thresh represents the critical penalty threshold that triggers the alert state, and 𝜅 acts as the temperature parameter controlling the steepness of the mode transition. This mathematical design enables an elastic priority adaptation where best-effort services are dynamically and gracefully scaled down to safeguard the mission-critical URLLC slices during extreme traffic demand surges.
C. RAPO-TD3 Implementation Details To ensure reliable resource management in the nonstationary aerial backhaul environment, we detail the implementation of the proposed RAPO-TD3 framework. Our approach mitigates the overestimation bias inherent in standard actor-critic methods by maintaining twin critic networks, 𝑄 𝜙1 and 𝑄 𝜙2 , and adopting the minimum of their estimates for value updates. Furthermore, to strike a critical balance between global architectural stability and localized edge node exploration, we introduce a selective exploration noise mechanism. Specifically, the Ornstein-Uhlenbeck (OU) exploration noise is exclusively injected into the HAA-layer spatial actions z𝑛 (𝑡), while the macro-layer ratios zcore (𝑡) are strictly generated by deterministic policy outputs. This ensures that the MSO does not destabilize the entire network core while local HAAs explore optimal spatial distributions. Finally, we replace uniform sampling with 𝛼-prioritized experience replay [23]. This mechanism prioritizes transitions with larger temporal difference (TD) errors, which correspond to rare critical states where QoS requirements are violated. During execution, the RAPO mechanism dynamically calibrates the multi-service priority weights to compute the global reward, seamlessly guiding the agent to safeguard URLLC constraints. The integration of the double soft-max projection, selective noise, and the RAPO mechanism is designed so that every evaluated action satisfies the mathematical simplex bounds while accelerating convergence. The complete algorithmic procedure is detailed in Algorithm 1. D. Complexity Analysis The computational and structural complexity of the proposed hierarchical slicing framework is analyzed across two distinct phases: offline centralized training and online decentralized execution. During the online execution phase, the centralized MSO strictly relies on the trained actor network 𝜋 𝜃 to generate real-time allocation decisions. Let 𝐿 𝑎 denote the number of fully connected layers in the actor network, and 𝑈𝑙 represent the number of neurons in the 𝑙-th layer, where 𝑈0 and 𝑈 𝐿𝑎 correspond to the dimensions of the state vector s𝑡 and the action vector a𝑡 , respectively. The asymptotic time complexity for generating a single operational decision is bounded
Algorithm 1: RAPO-TD3 for Dual-Layer Slicing in HAA Backhaul Networks Input : Initial state s1 , target update rate 𝜏, delayed update interval 𝑑, PER parameters 𝛼, 𝛽, 𝜖per , hyperparameters 𝜎tgt , 𝑐 clip , learning rates 𝜂 𝑎 , 𝜂𝑐 , discount factor 𝛾, maximum episodes 𝐸 𝑃total , total time steps 𝑇, mini-batch size 𝑀 Output: Optimized actor network 𝜋 𝜃 for resource slicing 1 Initialize critic networks 𝑄 𝜙1 , 𝑄 𝜙2 and actor 𝜋 𝜃 with random weights; ′ ′ ′ 2 Initialize target networks 𝜙 ← 𝜙1 , 𝜙 ← 𝜙2 , 𝜃 ← 𝜃; 1 2 3 Initialize PER buffer D; 4 for 𝑒 = 1 to 𝐸 𝑃total do 5 Observe initial state 𝑁 , S(0), {g (0)} 𝑁 , ψ }; s1 = {{o𝑛,1 } 𝑛=1 𝑛 1 𝑛=1 6 for 𝑡 = 1 to 𝑇 do 7 Select base action a𝑡 = 𝜋 𝜃 (s𝑡 ); // Selective Exploration Noise 8 Inject OU noise ǫ exclusively into HAA-layer sub-slice actions z𝑛 (𝑡); // Double Soft-max Projection 9 Stage-1: Map Í zcore (𝑡) to macro-layer ratios 𝑆𝑖 (𝑡) such that 𝑖 𝑆𝑖 (𝑡) = 1; 10 Stage-2: Map z𝑛 (𝑡) to spatial distribution weights 𝑔𝑛,𝑖 (𝑡); 11 Calculate final bandwidth fraction: 𝑓𝑛,𝑖 (𝑡) = 𝑆𝑖 (𝑡) · 𝑔𝑛,𝑖 (𝑡); 12 Calculate physical bandwidth: 𝐵 𝑛,𝑖 (𝑡) = 𝑓𝑛,𝑖 (𝑡) · 𝐵total ; // RAPO Evaluation & Execution 13 Execute 𝐵 𝑛,𝑖 (𝑡), compute dynamic priority weights via the RAPO mechanism, and observe global reward 𝑟 (𝑡) by (33) and next state s𝑡+1 ; // PER Storage & Priority Update 14 Store transition (s𝑡 , a𝑡 , 𝑟 (𝑡), s𝑡+1 ) in D with maximum priority; // Network Update Stage with Twin-Critic 15 Sample mini-batch of 𝑀 Ítransitions from D with probability 𝑃(𝑏) = 𝑝 𝑏𝛼 / 𝑗 𝑝 𝛼𝑗 ; 16 Generate smoothed target noise ǫ̃ ∼ clip(N (0, 𝜎tgt ), −𝑐 clip , 𝑐 clip ); 17 Compute target action: ã ← clip(𝜋 𝜃 ′ (s𝑡+1 ) + ǫ̃, alow , ahigh ); 18 𝑦 ← 𝑟 𝑏 + 𝛾 min 𝑗=1,2 𝑄 𝜙′𝑗 (s𝑡+1 , ã); 19 Update critics with learning rate 𝜂𝑐 by minimizing weightedÍMSE loss 𝑀 1 2 L= 𝑀 𝑏=1 𝑤 𝑏 (𝑦 − 𝑄 𝜙 𝑗 (s𝑏 , a𝑏 )) ; 20 if 𝑡 (mod 𝑑) == 0 then 21 Update actor 𝜃 with learning rate 𝜂 𝑎 using sampled policy gradient ∇ 𝜃 𝐽; 22 Soft update target networks: 𝜃 ′ ← 𝜏𝜃 + (1 − 𝜏)𝜃 ′ ; 23 end 24 Update priority 𝑝 𝑏 ← |𝑦 − 𝑄 𝜙1 (s𝑏 , a𝑏 )| + 𝜖per for sampled transitions; 25 end 26 end
Í 𝐿𝑎 −1 by O( 𝑙=0 𝑈𝑙 𝑈𝑙+1 ). This forward-propagation calculation requires merely a fraction of a millisecond, mathematically satisfying the strict execution constraints mandated by URLLC services. Conversely, the offline training phase incurs a higher computational footprint. During each gradient update step, the framework computes the forward and backward passes for
11
both the actor and the twin critic networks using a minibatch of size 𝑀. Furthermore, the integration of prioritized experience replay utilizing a sum-tree data structure introduces a sampling and priority-updating overhead of O(𝑀 log |D|), where |D| is the maximum capacity of the replay buffer. Since this training complexity is exclusively processed offline by high-performance computing clusters at the core network, it remains completely decoupled from real-time aerial operations. Regarding the signaling overhead, the MSO requires state gathering from 𝑁 HAAs and subsequent action broadcasting at the beginning of each control period. The dimension of the gathered state vector scales linearly with the number of agents, yielding a communication complexity of O(𝑁). This strict linear scaling ensures that the signaling overhead does not exponentially explode as the aerial fleet expands, efficiently addressing the scalability requirements for dense 6G network deployments. E. Theoretical Convergence and Stability Analysis Although providing a global convergence proof for deep reinforcement learning with non-linear neural network approximators remains an open mathematical challenge, we establish the theoretical stability and convergence properties of the proposed RAPO-TD3 framework. The convergence guarantees of our architecture are primarily driven by the properties introduced by the double soft-max projection and the RAPO mechanism. Property 1 (Action and Reward Boundedness). Let A denote the transformed action space and R denote the reward space. The double soft-max projection restricts the executed bandwidth allocations to the joint probability simplex, i.e., Í𝑁 Í 𝑛=1 𝑖∈ {𝑒,𝑢,𝑚} 𝑓 𝑛,𝑖 (𝑡) = 1, with 𝑓 𝑛,𝑖 (𝑡) ∈ [0, 1]. Consequently, the action space A is compact and bounded. Furthermore, as defined in (29) through (31), the piecewise-linear satisfaction functions yield utilities bounded within [0, 1], and the weight vector ω(𝑡) is convexly interpolated. Therefore, the global reward 𝑟 (𝑡) is bounded by 𝑟 min ≤ 𝑟 (𝑡) ≤ 𝑟 max , preventing diverging returns and bounding the temporal difference (TD) error variance during experience replay. Property 2 (Lipschitz Continuity of the Objective Landscape). A fundamental requirement for stable gradient descent in deterministic policy gradient frameworks is the Lipschitz continuity of the objective function [39]. Conventional approaches with binary QoS penalties or heuristic hard-switching rules introduce step-function discontinuities, causing infinite gradient variances. In our framework, both the double soft-max projection and the RAPO sigmoid gating function (36) are continuously differentiable (C 1 functions) across the entire stateaction space. Let 𝐽 (𝜃) be the expected return objective of the actor network. Following standard optimization theory [40], the gradients of the transformations with respect to the preactivation outputs z(𝑡) are bounded by the properties of the soft-max derivative (Jacobian matrix norms are bounded by 1) 1 and the sigmoid derivative (bounded by 4𝜅 ). Consequently, the overall reward function is 𝐿 𝑟 -Lipschitz continuous, satisfying: |𝑟 (s, a1 ) − 𝑟 (s, a2 )| ≤ 𝐿 𝑟 ||a1 − a2 || 2 ,
∀a1 , a2 ∈ A,
(37)
where 𝐿 𝑟 > 0 is the Lipschitz constant of the reward function. Building upon the compact state-action boundary and the smooth, gradient-bounded objective landscape, we formulate the stability guarantee of the proposed learning architecture. These fundamental properties collectively eliminate the risks of gradient explosion and unbounded Q-value divergence, paving the way for the algorithmic stability theorem. Theorem 1 (Algorithmic Stability Guarantee). For the proposed RAPO-TD3 framework, the variance of the sampled policy gradient ∇ 𝜃 𝐽 is bounded. When coupled with the twincritic overestimation clipping and a prioritized sampling probability 𝑃(𝑏) > 0 for all transitions in D, the Robbins-Monro conditions for stochastic approximation [24] are satisfied. This implies asymptotic convergence to a stable local optimum policy under standard decaying learning rate assumptions without divergent oscillations. Proof. To establish asymptotic convergence, we verify that the proposed architecture satisfies the Robbins-Monro conditions for stochastic approximation [24]. First, let 𝑄 𝜋 (s, a) denote the true action-value function under policy 𝜋 𝜃 . Based on the boundedness established in Property 1, the instantaneous reward is bounded within [𝑟 min , 𝑟 max ]. Given the infinite-horizon discounted return with a discount factor 𝛾 ∈ (0, 1), the target value 𝑦 and the critic estimates 𝑄 𝜙 𝑗 (s, a) are bounded within the compact interval [𝑟 min /(1 − 𝛾), 𝑟 max /(1 − 𝛾)], eliminating the risk of numerical divergence. Second, the deterministic policy gradient is formulated as ∇ 𝜃 𝐽 = E[∇ 𝜃 𝜋 𝜃 (s)∇a 𝑄 𝜙 𝑗 (s, a)| a= 𝜋 𝜃 (s) ]. Under Property 2, the reward function is 𝐿 𝑟 -Lipschitz continuous, which, under the standard assumption of smooth environmental transition dynamics, inherently bounds the spatial derivative of the parameterized critic network such that k∇a 𝑄 𝜙 𝑗 (s, a) k 2 ≤ 𝐿 critic < ∞, where 𝐿 critic represents the uniform Lipschitz constant of the critic network gradient. Simultaneously, since the actor network 𝜋 𝜃 (s) is constructed using fully connected layers with smooth activation functions over a compact parameter space, its structural gradient is also bounded by k∇ 𝜃 𝜋 𝜃 (s) k 2 ≤ 𝐿 actor < ∞, where 𝐿 actor is the corresponding bounding constant. During training, transitions are drawn via prioritized sampling rather than uniform distribution. By applying Importance Sampling (IS) weights to correct the non-uniform sampling bias, the stochastic policy gradient estimate remains mathematically unbiased. Because the norm of the sample gradient mapping is bounded by the product of the Lipschitz constants, the variance of this unbiased estimator computed over a minibatch of size 𝑀 scales inversely with 𝑀 and is structurally upper-bounded by a finite constant variance 𝜎 2 : i (𝐿 h 2 2 critic 𝐿 actor ) ≤ 𝜎 2 < ∞. (38) E ∇ˆ 𝜃 𝐽 − ∇ 𝜃 𝐽 2 ≤ 𝑀 Let 𝜂 𝑎,𝑡 denote the decaying learning rateÍat training step 𝑡. By enforcing the standard step-size criteria ∞ 𝑡=1 𝜂 𝑎,𝑡 = ∞ and Í∞ 2 𝜂 < ∞, the bounded gradient variance ensures that the 𝑡=1 𝑎,𝑡 stochastic noise cancels out asymptotically. Furthermore, the twin-critic clipping mechanism in TD3 minimizes the positive overestimation bias, ensuring that the policy updates do not
12
overshoot into unstable regions of the action simplex. Since the prioritized experience replay mechanism guarantees a non-zero selection probability 𝑃(𝑏) > 0 for all historical transitions, the tracking sequence cannot entrap in empty gradient zones. Consequently, according to the stochastic approximation theorem, the actor network parameter vector 𝜃 converges to a stationary local optimum satisfying ∇ 𝜃 𝐽 = 0 with probability 1. V. S IMULATION R ESULTS AND A NALYSIS A. Simulation Setup and Traffic Profile The simulation environment is designed to evaluate the performance of the proposed RAPO-TD3 framework in a dynamic aerial backhaul scenario. To accurately capture the complex inter-cell interference and multi-agent resource contention without succumbing to the exponential state-space explosion typical of deep reinforcement learning, we model a dense interference sub-cluster comprising 𝑁 = 4 HAAs. This configuration is consistent with standard 3GPP deployment scenarios for aerial vehicles [41] and macro-cell sectoring interference models [42]. It provides a mathematically sound baseline for validating decentralized execution, as the previously proven O(𝑁) computational complexity guarantees linear scalability to larger fleet deployments. The detailed communication parameters, QoS thresholds, and reinforcement learning hyperparameters are consolidated in Table II. To mathematically evaluate the adaptive priority optimization capability governed by the proposed RAPO framework, we implement two non-stationary multi-service traffic models, as illustrated in Fig. 3. The simulation evaluates the system under the following distinct workload behaviors: • 24-Hour Periodic Traffic: As shown in Fig. 3(a), the arrival rates for eMBB, URLLC, and mMTC slices follow a non-stationary Poisson process with time-varying intensities. As depicted in the relative demand profile, the eMBB and mMTC traffic exhibit a clear diurnal cycle, peaking during daytime hours to simulate intense humancentric urban activity, while the URLLC traffic maintains a lower but strictly bounded baseline demand. • Burst URLLC Traffic: To simulate mission-critical emergencies, we introduce stochastic data bursts into the environment. A burst event occurs with an anomaly probability of 𝑃burst = 0.20 evaluated at each operational hour. Once triggered, the localized workload surge persists for a deterministic duration of 𝑇burst = 0.2 hours, which corresponds to 12 minutes of continuous peak pressure. During these anomalous intervals, the URLLC packet arrival rate is instantaneously scaled by a factor of 5.0, representing a severe 500% demand surge that forces the MSO to execute real-time slice priority adaptation. Fig. 3(b) depicts a sample realization of this bursty traffic pattern, capturing the extreme multi-order spikes over the 24-hour temporal evaluation window. To comprehensively demonstrate the effectiveness of the proposed RAPO-TD3 framework in handling dual-layer resource slicing and physical boundary constraints, we benchmark its performance against two state-of-the-art continuouscontrol reinforcement learning baselines, namely PPO and
TABLE II C ONSOLIDATED S IMULATION AND H YPERPARAMETERS Parameter Description
Symbol
Default Value
Network Environment Number of HAAs Total System Bandwidth Centralized Macro Control Period MBS Downlink Power per HAA Noise Power Spectral Density Carrier Frequency HAA Operational Altitude Dense Urban Environment Constants FSPL Bulk Normalization Constant Additional Path Loss (LoS, NLoS) Small-Scale Fading Profile
𝑁 𝐵total 𝑇𝐶 𝑃𝑛 𝑁0 𝑓𝑐 𝐻haa 𝑎env , 𝑏env 𝐶FSPL 𝜂LoS , 𝜂NLoS Ω𝑛 (𝑡 )
4 10 MHz 100 ms 23 dBm 1 × 10−9 W/Hz 2.5 GHz 120 m 12.08, 0.11 -147.55 1.0 dB, 20.0 dB Rayleigh (𝑠 = 1.0)
Service Requirements eMBB Rate Threshold (Min, Req) 𝑅min,𝑒 , 𝑅req,𝑒 10 Mbps, 50 Mbps URLLC Delay Threshold (Req, Max) 𝐷req,𝑢 , 𝐷max,𝑢 2 ms, 4 ms mMTC Success Rate (Min, Req) 𝐺min,𝑚 , 𝐺req,𝑚 0.90, 0.99 URLLC Packet Size & Error Bound 𝐿𝑛,𝑢 , 𝜖𝑢 3000 bits, 10−5 mMTC Packet Size 𝐿𝑛,𝑚 256 bits 𝑚 mMTC Time Constraint 𝑇max 20 ms mMTC Maximum Retransmissions 𝐾max 1 URLLC Burst Probability 𝑃burst 0.20 per hour URLLC Burst Duration 𝑇burst 0.2 hours RAPO-TD3 Hyperparameters Actor/Critic Learning Rate Discount Factor Target Update Rate Replay Buffer Capacity Mini-batch Size PER Parameters (Exponent, Const) Policy Smoothing (std, clip) Delayed Actor Update Interval Minimum Reserved Slice Ratio Total Training Episodes Total Steps per Episode
𝜂 𝑎 , 𝜂𝑐 𝛾 𝜏 |D| 𝑀 𝛼, 𝜖per 𝜎tgt , 𝑐clip 𝑑 𝜌 𝐸 𝑃total 𝑇
5 × 10−5 , 5 × 10−4 0.99 0.005 3 × 105 256 0.6, 10−3 0.1, 0.3 2 0.0 500 200
ωbase , ωalert 𝑃thresh 𝜅 𝜛 𝜒init , 𝜒grow 𝜒osc , 𝑇osc
[7, 8, 2], [8, 9, 1] 0.05 0.05 1.0 0.2, 0.8 0.1, 30 Episodes
Reward & RAPO Settings Base / Alert Mode Weights Initial Penalty Threshold Penalty Adaptation Factor URLLC Delay Penalty Coefficient Penalty Factor (Init, Grow) Penalty Oscillation (Amp, Period)
DDPG. The structural characteristics of these baseline schemes are outlined as follows: Baseline: This scheme implements the PPO algorithm, which functions as an on-policy actor-critic framework utilizing a clipped surrogate objective function to bound policy updates. In this evaluation, PPO represents an unconstrained reward-maximization approach that lacks the specialized hierarchical projection architecture, thereby serving as a primary reference to evaluate the necessity of embedding structural constraints directly into the action space. • DDPG Baseline: This scheme implements the DDPG algorithm, which optimizes a deterministic policy through the gradients of a single critic network. This baseline operates without target policy smoothing or delayed actor updates, allowing us to isolate and verify the algorithmic efficacy of the twin-critic overestimation mitiga• PPO
3.5
2 1.5 1 0.5 0 0
2
4
6
8
10
12
14
16
18
20
2.6
2.4
2.2
22
Hour of Day
DDPG Baseline PPO Baseline Proposed RAPO-TD3 0
eMBB URLLC mMTC URLLC (with surges)
6 5 4 3 2 1
URLLC QoS Satisfaction Ratio
7
2
4
6
8
10
12
200
300
400
500
DDPG Baseline PPO Baseline Proposed RAPO-TD3
0.4 0.3 0.2 0.1 0
0
100
14
16
18
20
22
Hour of Day
(b) Burst URLLC Traffic Fig. 3. Two different multi-service traffic demand patterns in the simulation environment: (a) 24-Hour Periodic Traffic, and (b) Burst URLLC Traffic.
200
300
400
500
Training Episode
(a) Convergence of Total Reward
(b) Constraint Violation Penalty
1.05
1.01 DDPG Baseline PPO Baseline Proposed RAPO-TD3
1
0.95
0.9
0 0
100
0.5
Training Episode
(a) 24-Hour Periodic Traffic
Relative Demand
2.8
mMTC QoS Satisfaction Ratio
2.5
Total Cumulative Reward
Relative Demand
3
eMBB URLLC mMTC
3
System Constraint Violation Penalty
13
0
100
200
300
400
500
Training Episode
(c) URLLC QoS Satisfaction Ratio
DDPG Baseline PPO Baseline Proposed RAPO-TD3
1.005 1 0.995 0.99 0.985 0.98
0
100
200
300
400
500
Training Episode
(d) mMTC QoS Satisfaction Ratio
Fig. 4. Convergence performance and constraint satisfaction dynamics of different DRL algorithms under the normal 24-hour periodic traffic scenario: (a) Total cumulative reward, (b) Constraint violation penalty, (c) URLLC QoS satisfaction ratio, (d) mMTC QoS satisfaction ratio.
tion mechanisms inherent to the TD3 architecture under highly volatile environments. B. Convergence Performance and Constraint Satisfaction To validate the learning efficiency and the constraint adherence capabilities of the proposed framework, we first evaluate the algorithmic convergence under the normal 24-hour periodic traffic scenario. The performance of the proposed RAPO-TD3 algorithm is benchmarked against two widely adopted continuous-control DRL algorithms: DDPG and PPO. The training dynamics, including the cumulative reward, system penalty, and mission-critical satisfaction ratios over 500 episodes, are illustrated in Fig. 4. An initial observation of the total cumulative reward in Fig. 4(a) reveals that the PPO baseline numerically achieves a competitive raw reward score. However, the system penalty dynamics in Fig. 4(b) expose a critical operational flaw. Because PPO lacks the structural double soft-max projection to accurately navigate the coupled multi-tier conservation laws, it fails to optimize the spatial distribution effectively. This structural deficiency causes PPO to over-allocate resources to eMBB services at the expense of mission-critical slices, triggering severe URLLC latency violations and resulting in penalty spikes reaching up to 0.35. Conversely, the proposed RAPO-TD3 framework converges to a stable, constraint-aware operating point. Driven by the architectural double soft-max projection and the dynamic penalty scaling factor 𝜁 (𝑡), the RAPO-TD3 agent is able to navigate the complex non-convex action space without relying on constraint-violating shortcuts. As depicted in Fig. 4(b), the system penalty for RAPO-TD3 is aggressively suppressed to a near-zero level after the initial exploration phase, maintaining an exceptionally low and safe footprint compared to the baseline algorithms.
The effectiveness of the proposed constraint-aware design is further supported by the results shown in Fig. 4(c), which tracks the URLLC QoS satisfaction ratio. While the PPO baseline frequently drops below the mandatory reliability threshold, declining to approximately 0.92 during peak traffic hours, the proposed RAPO-TD3 framework consistently safeguards the latency-sensitive URLLC slices. Furthermore, Fig. 4(d) explicitly illustrates the convergence trajectory of the mMTC QoS satisfaction ratio under this baseline scenario to address the holistic multi-service evaluation. In complete alignment with the reward and penalty dynamics, the proposed RAPOTD3 framework tightly tracks the near-optimal 1.0 satisfaction boundary throughout the entire training process with minor statistical fluctuations. Conversely, the PPO baseline displays pronounced performance degradation, dropping down to a satisfaction level of 0.985 near training episodes 225 and 450. This unstable behavior further substantiates that unconstrained reward-maximization schemes compromise best-effort slice stability even under non-bursty diurnal workloads, whereas our constraint-aware architecture guarantees stable multi-service isolation and steady-state protocol compliance. C. Resilience and Resource Allocation Under Extreme URLLC Traffic Surges To fully evaluate the operational resilience and stresstolerance limits of the proposed framework during missioncritical emergencies, we expose the trained policies to a nonstationary bursty traffic profile. This scenario introduces acute 500% surges in URLLC packet arrival rates, serving as a rigorous benchmark to test the dynamic priority orchestration agility of the MSO. The transient tracking of multi-service sat-
System Constraint Violation Penalty
14
URLLC QoS Satisfaction Ratio
1.05 1 0.95 0.9 0.85 DDPG Baseline PPO Baseline Proposed RAPO-TD3
0.8 0.75
0
100
200
300
400
500
3.5 DDPG Baseline PPO Baseline Proposed RAPO-TD3
3 2.5 2 1.5 1 0.5 0
0
100
(a) URLLC QoS Satisfaction Ratio
400
500
1.05
0.9 0.8 0.7 0.6 0.5 DDPG Baseline PPO Baseline Proposed RAPO-TD3
0.4 0
100
200
300
400
500
Training Episode
(c) eMBB QoS Satisfaction Ratio
mMTC QoS Satisfaction Ratio
eMBB QoS Satisfaction Ratio
300
(b) Constraint Violation Penalty
1
0.3
200
Training Episode
Training Episode
1
0.95
0.9
0.85
DDPG Baseline PPO Baseline Proposed RAPO-TD3 0
100
200
300
400
500
Training Episode
(d) mMTC QoS Satisfaction Ratio
Fig. 5. Dynamic resilience and resource allocation performance of different DRL algorithms under the extreme URLLC traffic surge scenario: (a) URLLC QoS satisfaction ratio, (b) Constraint violation penalty, (c) eMBB QoS satisfaction ratio, and (d) mMTC QoS satisfaction ratio.
isfaction ratios, constraint penalties, and resource allocations over 500 episodes is consolidated in Fig. 5. As illustrated in Figs. 5(a) and 5(b), the PPO baseline exposes the structural vulnerability of unconstrained rewardmaximization strategies. Driven by the impulse to maximize raw generalized utility, the PPO agent consistently overallocates resources to eMBB services, maintaining an allocation profile near 0.98 as depicted in Fig. 5(c). When stochastic URLLC traffic spikes occur, this behavior triggers severe resource contention. Consequently, the PPO policy incurs sharp constraint violation penalties that spike above 3.0 and 3.3 around episodes 320 and 400, respectively. These boundary breaches cause the URLLC QoS satisfaction ratio to experience significant degradation, dropping to a low of 0.77. Such operational deficits indicate that standard continuous-control algorithms fail to safeguard mission-critical communication links under non-stationary pressure. Conversely, the DDPG baseline exhibits an overly conservative learning trajectory. Although it appears to maintain acceptable URLLC satisfaction with a low penalty footprint (Figs. 5(a) and 5(b)), its structural deficiency is clearly unveiled in Fig. 5(c). Due to severe policy underestimation bias, the DDPG agent fails to effectively optimize the multi-service action space, with its normalized eMBB resource allocation stagnating below 0.75 even after 500 episodes of training. This conservative suboptimality severely underutilizes the available aerial backhaul capacity, rendering it impractical for highthroughput cooperative networking. In sharp contrast, the proposed RAPO-TD3 framework armed with the RAPO mechanism learns a constraint-aware high-performing policy. As illustrated in Figs. 5(a) and 5(b), the proposed framework undergoes a transient penalty spike
reaching 1.35 and a corresponding satisfaction dip to 0.85 near episode 25, reflecting initial structural adjustments under extreme environmental volatility. Following this brief exploratory adaptation, the system rapidly stabilizes. Driven by the dynamic weight vector ω(𝑡) generated by the adaptive smooth gating structure, the agent seamlessly scales down best-effort mMTC resource shares to accommodate the emergency surges. Beyond episode 50, the proposed RAPO-TD3 framework maintains a stable eMBB allocation above 0.93 while simultaneously shielding the latencysensitive URLLC slice, keeping its satisfaction ratio tightly hovering at the near-optimal 1.0 boundary with near-zero system penalties. This resource allocation elasticity is comprehensively validated in Fig. 5(d), which tracks the dynamic trajectory of the mMTC QoS satisfaction ratio. The RAPOTD3 framework suffers only an isolated transient dip to 0.942 at episode 22 during the initial training phase, after which it rapidly recalibrates and preserves the near-optimal 1.0 satisfaction boundary. Meanwhile, the unconstrained PPO baseline undergoes pronounced service drops decreasing sharply to 0.89 and 0.87 near episodes 320 and 400, and the DDPG baseline remains locked in a highly inefficient and conservative operational zone. Our framework is able to circumvent these suboptimalities, demonstrating how the RAPO online tuning loop effectively secures dual-layer isolation under extreme 6G backhaul anomalies. D. Ablation Study and Architectural Efficacy To isolate and quantify the specific performance gains contributed by the twin-critic architecture and validate the necessity of mitigating overestimation bias in volatile aerial environments, we conduct an ablation study. We introduce the 1Q-Optimized baseline, a structural variant that degrades the RAPO-TD3 framework by utilizing only a single critic network and removing the delayed policy smoothing mechanisms, while preserving identical state-action spaces and prioritized replay configurations. The comparative training dynamics over 500 episodes under the extreme traffic surge scenario are illustrated in Fig. 6. A rigorous analysis of the total cumulative reward in Fig. 6(a) demonstrates the superior learning efficiency of the proposed RAPO-TD3 framework. Armed with the twin-critic mechanism, the proposed framework achieves rapid convergence, reaching a high-utility stable plateau of approximately 2.95 within the first 50 episodes and maintaining consistent policy stability throughout the remaining training duration. Conversely, the 1Q-Optimized baseline exhibits a severely impeded learning trajectory, requiring nearly 350 episodes to approach a sub-optimal reward of 2.8. More critically, around episode 400, the 1Q-Optimized baseline undergoes a significant performance degradation, with its cumulative reward declining toward 2.65. This instability provides empirical evidence of unmitigated overestimation bias, where a single critic network consistently overvalues poor resource-slicing actions, leading to propagated gradient errors that eventually destabilize the actor policy. The operational implications of this architectural deficiency are further illuminated by the system constraint violation
15
1.4
2.8 2.7 2.6 2.5 DDPG Baseline 1Q-Optimized Baseline Proposed RAPO-TD3
2.4 2.3
0
100
200
300
400
Training Episode
(a) Total Cumulative Reward
500
Constraint Violation Penalty
Total Cumulative Reward
3 2.9
DDPG Baseline 1Q-Optimized Baseline Proposed RAPO-TD3
1.2 1 0.1
0.8 0.6
Steady-State Volatility
0.05
0.4 0 0.2 0
0
100
200
460
480 300
400
500
TABLE III S CALABILITY A NALYSIS OF THE P ROPOSED F RAMEWORK HAAs (𝑁 ) 2 4 6 8
Avg. Steady-State Reward 2.9495 2.8795 2.7922 2.5556
Inference Time (ms) 0.1260 0.1326 0.1365 0.1442
500
Training Episode
(b) Constraint Violation Penalty
Fig. 6. Ablation study evaluating the structural efficacy of the twin-critic design under the extreme traffic scenario: (a) Total cumulative reward dynamics, and (b) Constraint violation penalty with a steady-state volatility inset.
penalty depicted in Fig. 6(b). Following the initial exploration phase, both the proposed framework and the 1Q-Optimized variant successfully suppress large-scale boundary breaches. However, a deeper examination of the steady-state volatility inset (spanning episodes 450 to 500) reveals distinct behavioral patterns. The proposed RAPO-TD3 framework exhibits minor boundary-tracking oscillations that peak boundedly near 0.08, whereas the 1Q-Optimized baseline maintains a lower penalty ripple. When evaluated jointly with the reward profiles from Fig. 6(a), this lower penalty profile indicates that the 1QOptimized agent has trapped itself within an overly conservative, suboptimal operational region that fails to fully exploit the available backhaul bandwidth capacity. In contrast, the proposed twin-critic design provides highly accurate value estimations, empowering the MSO to effectively adapt to and utilize the physical resource capacity frontier to maximize service utility while keeping localized SLA violations strictly constrained within a safe, negligible footprint. E. Scalability Analysis and Real-Time Execution Feasibility To evaluate the scalability and practical deployment feasibility of the proposed RAPO-TD3 framework, we conduct an additional experiment under varying network scales (𝑁 ∈ {2, 4, 6, 8}). The objective is to verify whether the computational efficiency and convergence stability can be maintained as the dense interference cluster expands. As demonstrated in Table III, the proposed architecture is highly scalable and yields two critical insights. First, while the framework maintains stable convergence across all scales, the steady-state average reward exhibits a marginal decline as 𝑁 increases. This phenomenon is fundamentally driven by physical-layer realities: a higher number of HAAs within a single frequency-reuse cluster intensifies co-channel interference, inherently degrading the overall Signal-to-Interference-plusNoise Ratio (SINR) and tightening the feasible multi-objective allocation space. Second, the inference latency experiences a slight, strictly linear increase from 0.1260 ms to 0.1442 ms. This is attributed to the O(𝑁) expansion of the stateaction dimensionality, which linearly increases the floatingpoint operations in the neural network’s input and output layers. Importantly, this computational overhead consumes only a negligible fraction of the stringent 1 ms URLLC latency
budget, leaving ample time for physical-layer transmission and queuing. These results theoretically and empirically suggest that the primary bottleneck in massive HAA deployments is physical interference rather than algorithmic scalability, indicating the stable applicability of our framework. VI. C ONCLUSION In this paper, we have proposed a stable dual-layer network slicing framework for HAA-assisted backhaul networks using a RAPO-TD3 deep reinforcement learning approach. To orchestrate heterogeneous service requirements under nonstationary traffic, we introduced a double soft-max projection mechanism and the RAPO scheme to guarantee physical constraint compliance and URLLC resilience. Furthermore, we presented a theoretical analysis showing that the proposed architecture satisfies the Robbins-Monro conditions and ensures Lipschitz continuity, supporting asymptotic convergence under the stated assumptions. Extensive simulations and scalability analyses under varying network dimensions demonstrate that our framework achieves better performance than existing baselines and achieves sub-millisecond inference latency. This computational efficiency consumes only a negligible fraction of the 1 ms URLLC latency budget while maintaining stable steady-state rewards against physical co-channel interference. Future work will extend this framework to accommodate unanchored small-scale UAVs by jointly optimizing 3D spatial trajectory mechanics, energy-aware hovering dynamics, and end-to-end fronthaul-backhaul slicing via decentralized federated learning paradigms for 6G coordination. R EFERENCES [1] W. Rafique, J. Rani Barai, A. O. Fapojuwo, and D. Krishnamurthy, “A survey on beyond 5G network slicing for smart cities applications,” IEEE Commun. Surveys Tuts., vol. 27, no. 1, pp. 595–628, Feb. 2025. [2] S. Zhang, “An overview of network slicing for 5G,” IEEE Wireless Commun., vol. 26, no. 3, pp. 111–117, Jun. 2019. [3] Y. Guan, Q. Song, T. Chen, W. Qi, L. Guo, and A. Jamalipour, “Slicingaware aerial networks for integrated sensing and communication: 3D placement and adaptive allocation of resources,” IEEE Veh. Technol. Mag., vol. 19, no. 2, pp. 79–88, Jun. 2024. [4] S. Tul Muntaha, M. Hafeez, Q. Z. Ahmed, F. A. Khan, Z. D. Zaharis, and P. I. Lazaridis, “RAN slicing with joint resource allocation for a multitenant–multi-service system,” IEEE Trans. on Cogn. Commun. Netw., vol. 11, no. 3, pp. 1927–1939, Jun. 2025. [5] S. K. Taskou, M. Rasti, and E. Hossain, “End-to-end resource slicing for coexistence of eMBB and URLLC services in 5G-advanced/6G networks,” IEEE Trans. Mobile Comput., vol. 23, no. 7, pp. 8015–8032, Jul. 2024. [6] H. Peng, L.-C. Wang, and Z. Jian, “Data-driven spectrum partition for multiplexing URLLC and eMBB,” IEEE Trans. on Cogn. Commun. Netw., vol. 9, no. 2, pp. 386–397, Apr. 2023. [7] A. Filali, Z. Mlika, S. Cherkaoui, and A. Kobbane, “Dynamic SDNbased radio access network slicing with deep reinforcement learning for URLLC and eMBB services,” IEEE Trans. Netw. Sci. Eng., vol. 9, no. 4, pp. 2174–2187, July-Aug. 2022.
16
[8] M. Tian, C. Li, Y. Hui, B. Chen, W. Yue, Y. Fu, and Z. Han, “An intelligent coexistence strategy for eMBB/URLLC traffic in multi-UAV relay networks via deep reinforcement learning,” IEEE Trans. Wireless Commun., vol. 23, no. 10, pp. 13 424–13 439, Oct. 2024. [9] Y. Liu, Q. Wang, H.-N. Dai, Y. Fu, N. Zhang, and C. C. Lee, “UAVassisted wireless backhaul networks: Connectivity analysis of uplink transmissions,” IEEE Trans. Veh. Technol., vol. 72, no. 9, pp. 12 195– 12 207, Sep. 2023. [10] M. Tian, C. Li, Y. Hui, N. Cheng, W. Yue, Y. Fu, and Z. Han, “Ondemand multiplexing of eMBB/URLLC traffic in a multi-UAV relay network,” IEEE Trans. Intell. Transp. Syst., vol. 25, no. 6, pp. 6035– 6048, Jun. 2024. [11] R. Chataut, M. Nankya, and R. Akl, “6G networks and the AI revolution—exploring technologies, applications, and emerging challenges,” Sensors, vol. 24, no. 6, Mar. 2024. [12] H. Cao, N. Kumar, L. Yang, M. Guizani, and F. R. Yu, “Resource orchestration and allocation of E2E slices in softwarized UAVs-assisted 6G terrestrial networks,” IEEE Trans. Netw. Service Manag., vol. 21, no. 1, pp. 1032–1047, Feb. 2024. [13] M. Yu, Y. Pi, A. Tang, and X. Wang, “Coordinated parallel resource allocation for integrated access and backhaul networks,” Computer Networks, vol. 222, p. 109533, Feb. 2023. [14] E. Ataeebojd, M. Rasti, and M. Latva-Aho, “Network selection and resource allocation for coexistence of eMBB and URLLC services in a 6G multi-band HetNet,” IEEE Trans. Green Commun. Netw., vol. 9, no. 3, pp. 1179–1194, Sep. 2025. [15] A. Abouaomar, A. Taik, A. Filali, and S. Cherkaoui, “Federated deep reinforcement learning for open RAN slicing in 6G networks,” IEEE Commun. Mag., vol. 61, no. 2, pp. 126–132, Feb. 2023. [16] J. Mei, X. Wang, K. Zheng, G. Boudreau, A. B. Sediq, and H. AbouZeid, “Intelligent radio access network slicing for service provisioning in 6G: A hierarchical deep reinforcement learning approach,” IEEE Trans. Commun., vol. 69, no. 9, pp. 6063–6078, Sep. 2021. [17] 3rd Generation Partnership Project (3GPP), “Management and orchestration; architecture framework,” 3rd Generation Partnership Project (3GPP), Tech. Spec. 28.533, 2024. [18] K. Arulkumaran, M. P. Deisenroth, M. Brundage, and A. A. Bharath, “Deep reinforcement learning: A brief survey,” IEEE Signal Process. Mag., vol. 34, no. 6, pp. 26–38, Nov. 2017. [19] Y. Cai, P. Cheng, Z. Chen, M. Ding, B. Vucetic, and Y. Li, “Deep reinforcement learning for online resource allocation in network slicing,” IEEE Trans. Mobile Comput., vol. 23, no. 6, pp. 7099–7116, Jun. 2024. [20] G. Chen, X. Mu, H. Liang, Q. Zeng, and Y.-D. Zhang, “Distributed RAN slicing based on MATD3 joint with evolutionary game assisted user association for MEC-enabled HetNets,” IEEE Trans. Wireless Commun., vol. 24, no. 1, pp. 260–276, Jan. 2025. [21] H. Kang, X. Chang, J. Mišić, V. B. Mišić, J. Fan, and Y. Liu, “Cooperative UAV resource allocation and task offloading in hierarchical aerial computing systems: A MAPPO-based approach,” IEEE Internet Things J., vol. 10, no. 12, pp. 10 497–10 509, Jun. 2023. [22] L. Bellone, B. Galkin, E. Traversi, and E. Natalizio, “Deep reinforcement learning for combined coverage and resource allocation in UAV-aided RAN-slicing,” in The 19th International Conference on Distributed Computing in Smart Systems and the Internet of Things (DCOSS-IoT), Pafos, Cyprus, 2023. [23] T. Schaul, J. Quan, I. Antonoglou, and D. Silver, “Prioritized experience replay,” in The 4th International Conference on Learning Representations (ICLR), San Juan, Puerto Rico, 2016. [24] H. Robbins and S. Monro, “A stochastic approximation method,” The Annals of Mathematical Statistics, vol. 22, no. 3, pp. 400–407, 1951. [25] Y. Xiao, M. Hirzallah, and M. Krunz, “Distributed resource allocation for network slicing over licensed and unlicensed bands,” IEEE J. Sel. Areas Commun., vol. 36, no. 10, pp. 2260–2274, Oct. 2018. [26] E. N. Tominaga, H. Alves, R. D. Souza, J. Luiz Rebelatto, and M. Latvaaho, “Non-orthogonal multiple access and network slicing: Scalable coexistence of eMBB and URLLC,” in IEEE 93rd Vehicular Technology Conference (VTC2021-Spring), Helsinki, Finland, 2021. [27] F. Zhou, P. Yu, L. Feng, X. Qiu, Z. Wang, L. Meng, M. Kadoch, L. Gong, and X. Yao, “Automatic network slicing for IoT in smart city,” IEEE Wireless Commun., vol. 27, no. 6, pp. 108–115, Dec. 2020. [28] A. Coelho, J. Rodrigues, H. Fontes, R. Campos, and M. Ricardo, “An algorithm for placing and allocating communications resources based on slicing-aware flying access and backhaul networks,” IEEE Access, vol. 10, pp. 128 923–128 942, Dec. 2022. [29] C.-C. Lai, Bhola, A.-H. Tsai, and L.-C. Wang, “Adaptive and fair deployment approach to balance offload traffic in multi-UAV cellular
networks,” IEEE Trans. Veh. Technol., vol. 72, no. 3, pp. 3724–3738, Mar. 2023. [30] P. Ladosz, L. Weng, M. Kim, and H. Oh, “Exploration in deep reinforcement learning: A survey,” Information Fusion, vol. 85, pp. 1–22, Sep. 2022. [31] J. A. Hurtado Sánchez, K. Casilimas, and O. M. Caicedo Rendon, “Deep reinforcement learning for resource management on network slicing: A survey,” Sensors, vol. 22, no. 8, Apr. 2022. [32] M. Ganjalizadeh, H. S. Ghadikolaei, D. Gündüz, and M. Petrova, “BSAC-CoEx: Coexistence of URLLC and distributed learning services via device selection,” IEEE Trans. Netw. Service Manag., vol. 23, pp. 1406–1421, 2026. [33] X. Song, B. Zhang, Z. Fan, R. Li, and S. Xu, “Dynamic normalization TD3-based task offloading for UAV-assisted collaborative computing,” IEEE Trans. Netw. Service Manag., vol. 23, pp. 2762–2777, 2026. [34] A. Al-Hourani, S. Kandeepan, and S. Lardner, “Optimal LAP altitude for maximum coverage,” IEEE Wireless Commun. Lett., vol. 3, no. 6, pp. 569–572, Dec. 2014. [35] W. Khawaja, I. Guvenc, D. W. Matolak, U.-C. Fiebig, and N. Schneckenburger, “A survey of air-to-ground propagation channel modeling for unmanned aerial vehicles,” IEEE Commun. Surveys Tuts., vol. 21, no. 3, pp. 2361–2391, thirdquarter 2019. [36] G. Durisi, T. Koch, and P. Popovski, “Toward massive, ultrareliable, and low-latency wireless communication with short packets,” Proc. IEEE, vol. 104, no. 9, pp. 1711–1726, Sep. 2016. [37] C. She, C. Yang, and T. Q. S. Quek, “Cross-layer optimization for ultrareliable and low-latency radio access networks,” IEEE Trans. Wireless Commun., vol. 17, no. 1, pp. 127–141, Jan. 2018. [38] S. Fujimoto, H. van Hoof, and D. Meger, “Addressing function approximation error in actor-critic methods,” in The 35th International Conference on Machine Learning (ICML), Stockholmsmässan, Stockholm, 2018. [39] D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller, “Deterministic policy gradient algorithms,” in The 31st International Conference on Machine Learning (ICML), Beijing, China, 2014. [40] S. Boyd and L. Vandenberghe, Convex Optimization. Cambridge University Press, 2004. [41] 3rd Generation Partnership Project (3GPP), “Study on Enhanced LTE Support for Aerial Vehicles,” 3rd Generation Partnership Project (3GPP), Tech. Rep. 36.777, 2017. [42] 3rd Generation Partnership Project (3GPP), “Study on channel model for frequencies from 0.5 to 100 ghz (release 14),” 3rd Generation Partnership Project (3GPP), Tech. Rep. 38.901 V16.1.0, 2017.
Chuan-Chi Lai (Member, IEEE) received the Ph.D. degree in Computer Science and Information Engineering from the National Taipei University of Technology, Taiwan, in 2017. He held research and faculty positions at National Chiao Tung University and Feng Chia University prior to his current role. Since 2024, he has been an Assistant Professor with the Department of Communications Engineering, National Chung Cheng University, Chiayi, Taiwan. His research interests include mobile edge computing, UAV networks, and AI for wireless communications. Dr. Lai was a recipient of the Postdoctoral Researcher Academic Research Award from the NSTC, Taiwan, in 2019, and Best Paper Awards at WOCC (2018, 2021) and ICUFN (2015).
Jen-Hsiang Li received the bachelor’s degree in Information Technology from Overseas Chinese University, Taichung, Taiwan, in 2023. In 2023, he joined the Applied Intelligent Communication and Computing Lab. Currently, he is pursuing his master’s degree in Information Engineering and Computer Science at Feng Chia University, Taichung, Taiwan. His research interests include radio resource management, network slicing, and reinforcement learning.