Hierarchical Cooperative MARL for Joint Downlink PRB and Power Allocation in a 5G System Alireza Ebrahimi Dorcheh, Tolunay Seyfi, Ryan Barker, Fatemeh Afghah
arXiv:2605.02149v1 [cs.NI] 4 May 2026
Holcombe Department of Electrical and Computer Engineering Clemson University, USA { alireze, tseyfi, rcbarke, fafghah}@clemson.edu Abstract—Efficient 5G downlink radio resource management requires jointly optimizing user scheduling and transmit-power allocation under time-varying wireless conditions. This is challenging in OFDMA because PRB assignment is combinatorial, power allocation is continuous, and performance depends on channel evolution, link adaptation, and long-term fairness. We propose a hierarchical cooperative multi-agent reinforcement learning framework with staged curriculum training for joint downlink PRB and power allocation in a physically grounded 5G environment. Simulations use Sionna for system-level modeling and Sionna RT for wireless scene construction and mobilityaware ray-traced channels. The task is decomposed into two stages: a PRB agent learns user-level resource shares, mapped to exact PRB assignments by a deterministic channel-aware quota resolver, and a power agent allocates base-station power across users and assigned PRB-symbol resources. The framework runs in a cross-layer loop with adaptive modulation and coding, HARQ feedback, outer-loop link adaptation, and a fairnessaware reward using smoothed throughput and Jain’s fairness index. A three-phase curriculum improves stability by training PRB allocation, power control, and joint fine-tuning. Under matched channel realizations, comparisons with an equal-power PF scheduler and two ablations isolating the learned PRB and power-control components show both components improve throughput distribution over PF, while the full PRB and power controller achieves the largest cell-throughput gain with only a modest reduction in Jain’s fairness index. Index Terms—5G, resource allocation, power allocation, reinforcement learning, Sionna.
I. I NTRODUCTION Downlink radio resource management remains central in 5G New Radio (NR). The next-generation NodeB (gNB) schedules physical resource blocks (PRBs), each spanning 12 subcarriers within a bandwidth part, while physical downlink shared channel (PDSCH) resource allocation and modulationand-coding-scheme (MCS) selection follow standardized procedures [1], [2]. The gNB must jointly select users, assign PRBs, and allocate transmit power under base-station (BS) power constraints [3]. These decisions are coupled: favorable PRBs can be wasted by poor power allocation, whereas aggressive power concentration can raise instantaneous rate but reduce fairness. The challenge further increases with channel variation, mobility, hybrid automatic repeat request (HARQ) feedback, and link adaptation [2], [4]. Proportionalfair (PF) scheduling is an interpretable baseline for balancing channel quality and long-term service, but it is hand-crafted and often paired with fixed or equal-power transmission [5]–[7]. Joint PRB/subcarrier and power allocation has been studied through optimization and learning [8]–[12]. Yet many methods cast the task as one large decision, causing highdimensional actions, unstable training, limited interpretability, and reliance on simplified channels that miss realistic geometry and dynamics. We address these limitations with a hierarchical cooperative multi-agent reinforcement learning This material is based upon work supported by the National Science Foundation under Grant Numbers CNS-2202972, CNS-2318726, and CNS2232048.
(MARL) framework built on Sionna system-level simulation and Sionna RT scene-aware channels. The PRB agent learns user-level shares instead of a full PRB map; a deterministic channel-aware quota resolver converts these shares into exact assignments. Conditioned on this schedule, a factorized power agent allocates each user’s total power and shapes it over assigned PRB-symbol resources. This keeps actions compact while preserving scheduling–power coupling. A key strength of the proposed framework is that it is evaluated in a physically grounded street-canyon scenario with mobile users and Sionna RT ray traced propagation. Its cross layer loop links scheduling and power decisions to signal-tointerference-plus-noise ratio (SINR), adaptive MCS, HARQ, and long term throughput fairness. The reward combines cell throughput and Jain’s fairness index. Empirically, the learned PRB and power control components each improve efficiency over PF with equal power, and their joint use provides the largest throughput gain with modest fairness reduction. To improve optimization stability, we adopt a staged curriculumbased training procedure. The PRB policy is trained first, the power policy is then trained on top of the learned scheduling behavior, and both policies are finally fine-tuned jointly. This training strategy aligns naturally with the hierarchical structure of the control problem and reduces the difficulty of simultaneous exploration over scheduling and power-control behavior. The main contributions of this paper are summarized as follows: • We formulate joint downlink PRB/power allocation as hierarchical cooperative MARL with sequential PRB and power agents optimizing a common throughput–fairness objective. • We implement the system-level simulator in Sionna with Sionna RT for physically grounded, mobility-aware, raytraced channels. • We introduce a deterministic quota-based PRB resolver that maps learned user shares to feasible channel-aware PRB assignments. • We design a factorized power policy for inter-user power partitioning and intra-user PRB-symbol power shaping. • We use staged curriculum training and matched-channel ablations for stable learning and fair comparison with PF, PRB-only, and power-only controllers. The remainder presents the system model, proposed framework, simulation methodology, results, and conclusion. II. R ELATED W ORK A. Joint Optimization of Radio Resources Joint radio resource optimization is well studied when frequency assignment and power control interact. For example, [8] formulates joint PRB/power allocation for eMBB/URLLC coexistence in 5G C-RAN, and [9] extends this direction to 5G H-CRAN with cross-layer interference and energy-efficiency objectives. These works show that PRB/RB assignment and
power allocation should not be optimized independently because frequency gains can be lost under poor power splits, while aggressive power concentration can reduce fairness. Fairness-aware resource optimization has also been studied in heterogeneous systems. [13] considers joint user association and resource allocation with multi-level fairness, showing that differentiated fairness can be built directly into distributed radio-resource optimization. Although not downlink PRB scheduling with symbol-level power shaping, it supports our use of an explicit throughput–fairness objective. Reinforcement learning is attractive for high-dimensional or non-convex online allocation. Prior work applies RL to joint URLLC power/resource allocation [10], OFDM subcarrier/power allocation [11], decomposed IAB scheduling/resource allocation [12], and O-RAN slice-level PRB allocation, where DORA uses PPO with deterministic intra-slice scheduling to reduce complexity [14]. Together, these works motivate structured learning-based control when monolithic action spaces are impractical. Taken together, these studies strongly motivate our formulation: joint scheduling and power control matter, fairness-aware objectives matter, and decomposition is often necessary for tractable learning. However, most prior works in this stream either remain optimization-driven, operate at subcarrier or association level rather than downlink PRB-level scheduling with intra-user power shaping, or do not explicitly consider a sequential hierarchical policy in which user-level PRB shares are first inferred and then resolved into exact PRB assignments before power is allocated. B. Digital-Twin-Based and Physically Grounded Resource Management Digital twins (DTs) and physically grounded wireless environments are increasingly used for radio access network (RAN) control control because scheduling and power decisions depend on geometry, mobility, propagation, and protocol-state evolution, which abstract channels often oversimplify. For example, [15] develops a DT-enabled framework for site-specific radio resource management in a NextG aerial corridor, combining high-fidelity ray tracing with deep reinforcement learning (DRL) for BS association and beam selection. [16] studies uplink RB scheduling and power allocation in a DT-integrated open radio access network (ORAN) Internet-of-Drones network with geo-referenced 3D DTs, GPU-accelerated propagation, and RL. These works reinforce the value of realistic, site-specific evaluation for resource-block and power decisions. DT-assisted learning also appears in edge offloading and synchronization: [17] uses DTassisted RL for microservice offloading and bandwidth allocation, while [18] studies continual RL for DT synchronization through resource-block allocation and device scheduling. Although outside 5G downlink PHY/MAC PRB-power scheduling, they demonstrate adaptive control under dynamics that static analytical models cannot easily capture. Adjacent sitespecific RAN work further supports realistic wireless-control evaluation: [19] combines Bayesian optimization, DRL, and Sionna ray tracing for mobility management, while REAL demonstrates closed-loop O-RAN PRB allocation using an OSC near-RT RIC, srsRAN, and a PPO-based xApp under GNU Radio channel impairments [20]. Thus, prior work motivates both joint scheduling–power optimization and realistic closed-loop evaluation, but existing DT/O-RAN studies mainly address BS association, beam selection, slicing, uplink interference, mobility, synchronization, or edge offloading rather than the downlink PRB–power task studied here. Overall, although prior studies provide important building blocks,
they do not capture the full combination addressed in this paper: joint downlink PRB allocation and power control, a hierarchical sequential learning structure, physically grounded mobility-aware ray-traced channels, and a closed-loop crosslayer execution with HARQ, link adaptation, and matchedchannel benchmarking. This combination is important because realistic geometry, mobility, and protocol state fundamentally shape the throughput–fairness tradeoff induced by PRB and power decisions. III. S YSTEM M ODEL AND P ROBLEM F ORMULATION A. Network and Resource Model We consider a single-cell downlink OFDMA system in which one base station (BS) serves U mobile user equipments (UEs) over a time-varying wireless channel. The downlink resource grid contains B physical resource blocks (PRBs) per slot and L data OFDM symbols in each slot. Let t denote the slot index, b ∈ {1, . . . , B} the PRB index, ℓ ∈ {1, . . . , L} the OFDM symbol index, and u ∈ {1, . . . , U } the user index. The scheduler determines whether PRB b in slot t is assigned to user u through the binary variable xt,b,u ∈ {0, 1}, where xt,b,u = 1 means that PRB b is assigned to user u for the entire slot. Since the downlink is orthogonal across users, each PRB can be assigned to at most one user in a slot: U X
xt,b,u ≤ 1,
∀t, b.
(1)
u=1
The same PRB assignment is applied across the L data symbols of the slot, while the transmit power may vary across symbols within an assigned PRB. Let pt,ℓ,b,u ≥ 0 denote the transmit power allocated to user u on PRB b and symbol ℓ. The BS operates under a total slot-level power constraint: L X B X U X
pt,ℓ,b,u ≤ Pmax ,
∀t.
(2)
ℓ=1 b=1 u=1
To ensure consistency between scheduling and power allocation, power can be allocated only on scheduled PRBs: 0 ≤ pt,ℓ,b,u ≤ xt,b,u Pmax ,
∀t, ℓ, b, u.
(3)
B. Sionna System-Level Model and Sionna RT Channel Model The overall system-level simulation is implemented using Sionna, while Sionna RT is used to generate scene-aware, mobility-dependent channel realizations. Let ht,ℓ,b,u denote the effective channel coefficient for user u on PRB b and symbol ℓ at slot t. Since the present setting is single-cell and orthogonal, the received signal quality is primarily noiselimited, and the corresponding SINR can be written as γt,ℓ,b,u =
pt,ℓ,b,u |ht,ℓ,b,u |2 , N0
(4)
where N0 is the effective noise power, and a corresponding rate proxy is r̂t,ℓ,b,u = log2 (1 + γt,ℓ,b,u ) .
(5)
However, the controller does not optimize this proxy alone. Instead, each action is executed through a cross-layer simulation loop in Sionna with adaptive modulation and coding, HARQ feedback, and outer-loop link adaptation (OLLA). Therefore, the delivered throughput depends not only on the instantaneous channel and power allocation, but also on the current state of the link-adaptation process.
Fig. 1: Proposed hierarchical control loop. Sionna RT generates mobility-aware channels. The PRB agent outputs quotas that a deterministic resolver maps to PRBs; the power agent then performs inter-user sharing and intra-user shaping. The PHY/MAC loop with HARQ, OLLA, and MCS adaptation returns throughput/fairness rewards and next observations. C. Fairness-Aware Long-Term Objective The goal of this subsection is to define the scalar throughput–fairness objective used to evaluate each slot-level allocation. Let Ru (t) denote the achieved throughput of user u at slot t. To avoid overly myopic behavior, the environment tracks an exponentially smoothed throughput for each user: Tu (t) = (1 − β)Tu (t − 1) + βRu (t),
(6)
where β ∈ (0, 1] is the smoothing factor. To account for both efficiency and service balance, we use Jain’s fairness index over the smoothed user throughputs: P 2 U T (t) u u=1 Jt = PU , (7) U u=1 Tu2 (t) + ϵ where ϵ > 0 is a small constant for numerical stability. The cell-throughput term is normalized as ! PU u=1 Tu (t) T̃t = clip , 0, 1 , (8) Tnorm and the instantaneous throughput–fairness objective is defined as gt = (1 − α)T̃t + αJt , (9) where α ∈ [0, 1] controls the throughput–fairness tradeoff. D. Joint PRB and Power Allocation Problem The BS seeks a sequential control rule that jointly determines the slot-level PRB assignment variables {xt,b,u } and the symbol-level power-allocation variables {pt,ℓ,b,u } over time so as to maximize the expected long-term throughput–fairness objective: "H−1 # X t max Eµ δ gt (10) µ
s.t.
t=0
xt,b,u ∈ {0, 1}, U X xt,b,u ≤ 1,
(11) ∀t, b,
(12)
0 ≤ pt,ℓ,b,u ≤ xt,b,u Pmax , ∀t, ℓ, b, u, L X B X U X pt,ℓ,b,u ≤ Pmax , ∀t.
(13)
u=1
ℓ=1 b=1 u=1
(14)
Here, δ ∈ (0, 1] is the discount factor and µ denotes a generic sequential control rule, which will be instantiated by the hierarchical MARL controller in the next section. The formulation above should be viewed as a general sequential control problem. Conventional online optimization is difficult because the integer PRB variables and continuous power variables are coupled through a stateful PHY/MAC loop. A monolithic RL formulation would also be unwieldy: the agent would need to observe channel summaries together with throughput, HARQ/OLLA, and MCS states, while directly outputting a mixed action containing the full PRB map and the symbol-level power tensor. This motivates the hierarchical MARL design in the next section, which separates user-level PRB sharing from conditioned power control while optimizing the same objective. Figure 1 provides an overview of the end-to-end control loop that connects the Sionna RT scene generation, hierarchical decision making, and cross-layer PHY/MAC execution considered in this work. IV. P ROPOSED H IERARCHICAL C OOPERATIVE MARL F RAMEWORK A. Hierarchical Sequential Decision Structure Directly learning a joint PRB-and-power allocation policy is difficult because the action space mixes combinatorial scheduling decisions with continuous power control. To address this, we decompose the problem into two cooperative sequential stages within each slot, as illustrated in Fig. 1. The PRB agent first determines how the available PRBs should be shared across users, and a deterministic resolver then maps these user-level quotas into exact PRB assignments. Conditioned on the resolved schedule, the power agent allocates the BS transmit-power budget across users and shapes that power over the assigned resources. The resulting allocation is executed in the Sionna-based PHY/MAC loop, whose throughput, fairness, HARQ, and link-adaptation outcomes are fed back to the controller. We refer to the framework as cooperative MARL because two distinct policies, with different observations and action spaces, act sequentially within each slot and are trained to optimize the same long-term objective. In the MARL implementation, the slot-level objective in (9) is used as the common reward, i.e., rt ≜ gt . Although execution is staged, the two agents cooperate through the shared environment and common reward.
Formally, let sprb and spow denote the observations of the t t PRB and power agents at slot t. The control process follows sprb → aprb → Xt → spow → apow → Pt , t t t t
(15)
where Xt is the PRB assignment and Pt is the resulting power-allocation tensor. B. PRB Allocation via Quota Learning and Deterministic Resolution Rather than outputting a full combinatorial PRB map, the PRB agent outputs a compact user-level vector zt = [zt,1 , . . . , zt,U ],
(16)
which is converted into normalized PRB shares and integer quotas as ezt,u qt,u = PU , zt,v v=1 e
B̄t,u = qt,u B,
U X
(17) The integer quotas Bt,u are obtained by flooring B̄t,u and applying largest-remainder correction. Given these quotas, a deterministic channel-aware resolver assigns exact PRB indices. Since power is selected only after scheduling, the resolver ranks PRBs using the channel-only score L X |ht,ℓ,b,u |2 . (18) ψt,b,u = ℓ=1
Let Btavail be the set of unassigned PRBs and Utact = {u : Bt,u > 0} the users with remaining quota. The resolver cycles over active users and assigns each user its best available PRB, b∈Btavail
xt,b⋆t (u),u = 1,
(19)
then removes b⋆t (u) from Btavail and decrements Bt,u until all quotas are exhausted. The resulting PRB map is reused across all L data symbols, while symbol-level powers are determined by the power agent. Thus, the policy learns how much spectrum each user receives, and the resolver determines which PRBs are assigned, guaranteeing feasibility without learning a full PRB map. C. Factorized Power Allocation After the PRB allocation is determined, the power agent allocates the BS power budget in a structured way. Its action is factorized into a user-level weight vector wt = [wt,1 , . . . , wt,U ] and a shaping-coefficient vector κt = [κt,1 , . . . , κt,U ]. The user-level weights are first converted into normalized power shares: ewt,u ηt,u = PU , (20) wt,v v=1 e which define the per-user power budgets tot Pt,u = ηt,u Pmax ,
(b) Radio map
Bt,u = B.
u=1
b⋆t (u) = arg max ψt,b,u ,
(a) Ray-tracing scene
Fig. 2: Sionna RT evaluation environment. (a) Ray-tracing visualization of the urban scene with a rooftop BS (red) and four vehicle-mounted receivers (black). (b) Corresponding radio map showing the spatial distribution of received power over the same environment.
(21)
PU
tot so that u=1 Pt,u = Pmax . The second component controls how each user’s power budget is distributed across its scheduled PRB-symbol resources. Let Bt,u = {b : xt,b,u = 1} (22)
denote the set of PRBs assigned to user u in slot t, and let ρt,ℓ,u (m) denote the PRB with rank m after sorting Bt,u in
descending order of |ht,ℓ,b,u |2 for symbol ℓ. We then define the exponential shaping weights as ωt,ℓ,ρt,ℓ,u (m),u = exp(−κt,u (m − 1)) ,
(23)
for m = 1, . . . , |Bt,ℓ,u |. The resulting power allocated to user u on symbol ℓ and PRB b is ωt,ℓ,b,u . PB ′ ′ ′ ℓ′ =1 b′ =1 xt,b ,u ωt,ℓ ,b ,u
tot pt,ℓ,b,u = xt,b,u Pt,u PL
(24)
When κt,u ≈ 0, the power budget of user u is distributed nearly uniformly over its scheduled PRB-symbol resources, whereas larger values of κt,u increasingly concentrate power on the stronger channel-ranked resources. The resulting symbol-by-PRB power tensor is finally expanded uniformly over the 12 subcarriers within each scheduled PRB to obtain the RE-level power tensor used by the PHY abstraction. D. Observation Design and Curriculum Training Both agents use compact cross-layer observations rather than raw channel tensors. Per-user features include channelquality summaries, previous allocation share, smoothed throughput, HARQ success history, and MCS state. The PRB agent observes them before scheduling; the power agent observes them after PRB resolution. Both agents are trained using PPO. Since joint exploration over scheduling and power control can be unstable, we adopt a staged curriculum-based training procedure. In the first phase, only the PRB agent is trained while power is assigned using a simple baseline rule. In the second phase, the power agent is trained on top of the learned PRB behavior. In the final phase, both agents are fine-tuned jointly. This strategy matches the hierarchy of the control problem and improves training stability in the Sionna-based environment. The intermediate PRB-only policy obtained after the first phase is also used as an ablation in evaluation, while a separate power-only ablation is trained by fixing the PRB scheduler to PF and learning only the power policy. V. S IMULATION S ETUP AND E VALUATION M ETHODOLOGY A. Sionna-Based Simulation Environment We evaluate a physically grounded single-cell 5G downlink using Sionna for system-level simulation and Sionna RT for scene construction, mobility-aware propagation, and ray-traced channels. Fig. 2 shows a rooftop BS serving
Algorithm 1 Training and Inference of the Proposed Hierarchical Controller 1: Initialize PRB policy πprb and power policy πpow 2: for phase = 1, 2, 3 do 3: if phase = 1 then 4: Train πprb only; use equal-power transmission 5: else if phase = 2 then 6: Freeze πprb and train πpow 7: else 8: Fine-tune πprb and πpow jointly 9: end if 10: for each PPO iteration do 11: Reset the environment 12: for each slot t do 13: Observe sprb and sample aprb t t prb 14: Convert at into quotas {Bt,u }U u=1 15: Resolve exact PRB assignment Xt using the deterministic channel-aware resolver 16: if phase = 1 then 17: Apply equal-power allocation and execute the PHY step 18: else 19: Construct spow from the resolved schedule t 20: Observe spow and sample apow t t 21: Compute inter-user power shares and intra-user exponential shaping 22: Build the RE-level power tensor Pt and execute the PHY step 23: end if 24: Receive common reward rt and update throughput, HARQ, and MCS trackers 25: end for 26: Update the trainable policy or policies using PPO 27: end for 28: end for 29: Inference: For each slot, apply πprb , deterministic PRB resolution, πpow , and the PHY execution without exploration
four vehicle-mounted receivers in a street-canyon environment whose radio map captures geometry-dependent receivedpower variation. The carrier frequency is 3.5 GHz with 30 kHz subcarrier spacing, 51 PRBs, 12 subcarriers per PRB, and 12 data OFDM symbols per 14-symbol slot; two symbols are reserved for signaling/control. The BS power is 40 dBm, the UE noise figure is 7 dB, and the slot duration is 0.5 ms. User mobility follows predefined trajectories and speeds, producing geometry-driven channel evolution rather than purely synthetic fading. B. Cross-Layer Execution At every slot, the environment first obtains the current channel realization from Sionna RT. The PRB agent then selects user-level resource shares, which are converted into exact PRB assignments through the deterministic quota-based resolver. Conditioned on this schedule, the power agent determines how the total BS power budget is distributed across users and across their assigned REs. This interaction is summarized in Fig. 1, which highlights the sequential coupling between scene-aware channel generation, hierarchical control, and PHY/MAC feedback. The resulting scheduling and power tensors are then passed to the Sionna-based system-level simulation loop, which performs effective SINR computation, adaptive MCS selection,
TABLE I: Quantitative comparison under matched-channel evaluation. Throughput is reported in Mbps, and ∆T̄ is the mean throughput gain relative to PF. Scheme T̄ ∆T̄ T10 T50 J¯ J50 PF (Baseline) 75.67 – 63.87 80.10 0.9233 0.9337 PRB Agent 76.98 +1.73% 65.94 81.20 0.9193 0.9311 Power Agent 77.77 +2.77% 67.89 81.48 0.9168 0.9294 PRB+Power 79.90 +5.58% 76.14 81.70 0.9101 0.9248
HARQ feedback, and outer-loop link adaptation. The environment updates the throughput, reliability, and allocationhistory trackers used to construct the next state and common reward. This ensures that the learned controller is evaluated in a closed-loop wireless system rather than through a static rate model. C. Comparison Schemes and Metrics We compare four schemes under the same evaluation protocol. The first is the PF baseline with equal-power transmission, denoted PF (Baseline). For each PRB b in slot t, the PF scheduler selects u⋆t,b = arg max u
ψt,b,u , Tu (t) + ϵ
(25)
PL 2 where ψt,b,u = ℓ=1 |ht,ℓ,b,u | is the slot-level channelquality score and Tu (t) is the smoothed throughput defined earlier. Thus, PF favors users with strong instantaneous channels while preventing persistent domination by already wellserved users. Power is then distributed uniformly over the scheduled PRB-symbol resources. To isolate the contribution of each learned component, we also evaluate two ablation variants. The PRB Agent uses the PRB policy trained in the first curriculum phase and applies uniform power over the scheduled resources. The Power Agent uses the PF scheduler for PRB assignment and applies a separately trained power policy conditioned on the PF schedule. Finally, PRB+Power denotes the proposed hierarchical controller after joint finetuning of both agents. Performance is evaluated using empirical CDFs and summary statistics of cell throughput and Jain’s fairness index under the matched-channel protocol. We report the mean, median, and 10th-percentile cell throughput, together with the mean and median Jain’s fairness index, so that both average behavior and lower-tail performance can be compared across schemes. D. Matched-Channel Evaluation Protocol To ensure a fair comparison, all comparison schemes are evaluated on the same cached sequence of Sionna RT channel realizations. Before evaluation run, the internal PHY/MAC states of all systems are reset so that HARQ memory, linkadaptation state, and related historical variables start from identical conditions. This matched-channel protocol is important because it removes channel randomness as a source of performance variation. Any observed difference among the compared schemes can therefore be attributed more directly to the quality of the control policy itself. The same sequential execution logic used during training is preserved during evaluation, with the PRB decision applied first and the power decision applied second.
modest Jain-fairness reduction, showing that hierarchical decomposition can translate coordinated scheduling and power adaptation into a better throughput–fairness operating point. Future work will extend the framework to multi-cell, queueaware, and advanced multi-antenna settings. R EFERENCES
(a) Cell throughput
(b) Jain’s fairness index
Fig. 3: Empirical CDFs under the matched-channel evaluation protocol: (a) cell throughput and (b) Jain’s fairness index. The PRB Agent and Power Agent curves isolate the effects of learned scheduling and learned power allocation, respectively, while PRB+Power denotes the proposed joint controller.
VI. R ESULTS AND D ISCUSSION Fig. 3 and Table I compare the PF equal-power baseline, the two learned ablations, and the proposed PRB+Power controller under the matched-channel evaluation protocol. Both ablations improve the throughput distribution relative to PF: the PRB Agent increases the mean cell throughput by 1.73%, while the Power Agent increases it by 2.77%. The full PRB+Power controller achieves the largest gain, improving the mean cell throughput from 75.67 Mbps to 79.90 Mbps, corresponding to a 5.58% improvement over PF. The Jain’s fairness CDF in Fig. 3b shows the expected throughput–fairness tradeoff. PF maintains the strongest fairness behavior, while the learned variants slightly shift the distribution toward lower fairness values. However, the degradation remains modest relative to the throughput improvement, especially for the full PRB+Power controller. This indicates that the proposed method improves cell efficiency primarily by exploiting favorable PRB and power-allocation opportunities, while still preserving a high level of long-term service balance. The lower-tail throughput also improves substantially. The 10th-percentile cell throughput increases from 63.87 Mbps under PF to 76.14 Mbps with PRB+Power, giving a 19.21% gain. This indicates that the proposed controller improves not only average cell efficiency, but also the lower-throughput operating region. The Jain’s fairness results show the expected throughput–fairness tradeoff: the median fairness index decreases from 0.9337 under PF to 0.9248 with PRB+Power, which is a relative reduction of only 0.95%. Overall, the ablation results strengthen the main claim of the paper. The PRB-only and power-only variants each outperform PF in throughput, confirming that learned scheduling and learned power allocation are individually useful. Their combination achieves the best throughput distribution, especially in the lower tail, while preserving a high fairness level in the considered Sionna RT-based 5G environment. VII. C ONCLUSION This paper presented a hierarchical cooperative MARL framework for joint downlink PRB and power allocation in a Sionna/Sionna RT 5G environment. The design combines quota-based PRB allocation, deterministic channel-aware resolution, and factorized power control within a HARQ/linkadaptation loop. Under matched-channel evaluation against PF equal-power scheduling and two ablations, PRB+Power delivers the strongest throughput improvement with only
[1] 3GPP, “NR; Physical Channels and Modulation,” 3rd Generation Partnership Project (3GPP), Technical Specification TS 38.211, 2025, Release 18, Version 18.7.0. [2] 3GPP, “NR; Physical Layer Procedures for Data,” 3rd Generation Partnership Project (3GPP), Technical Specification TS 38.214, 2025, Release 17, Version 17.14.0. [3] 3GPP, “NR; Base Station (BS) Radio Transmission and Reception,” 3rd Generation Partnership Project (3GPP), Technical Specification TS 38.104, 2025, Release 18, Version 18.10.0. [4] 3GPP, “NR; Medium Access Control (MAC) Protocol Specification,” 3rd Generation Partnership Project (3GPP), Technical Specification TS 38.321, 2025, Release 18. [5] F. P. Kelly, A. K. Maulloo, and D. K. H. Tan, “Rate control for communication networks: Shadow prices, proportional fairness and stability,” Journal of the Operational Research Society, vol. 49, no. 3, pp. 237–252, 1998. [6] A. Jalali, R. Padovani, and R. Pankaj, “Data throughput of cdma-hdr: A high efficiency-high data rate personal communication wireless system,” in Proc. IEEE Vehicular Technology Conference (VTC), 2000, pp. 1854– 1858. [7] H. J. Kushner and P. A. Whiting, “Convergence of proportional-fair sharing algorithms under general conditions,” IEEE Transactions on Wireless Communications, vol. 3, no. 4, pp. 1250–1259, 2004. [8] M. Setayesh, S. Bahrami, and V. W. S. Wong, “Joint PRB and power allocation for slicing eMBB and URLLC services in 5G C-RAN,” in Proc. IEEE Global Communications Conference (GLOBECOM), 2020. [9] Q. Cheng, K. Li, P. Zhu, J. Li, Y. Jiang, and D. Wang, “Joint resource block and power allocation for eMBB and URLLC coexistence in 5G HCRAN,” in Proc. International Conference on Wireless Communications and Signal Processing (WCSP), 2023. [10] M. Elsayed and M. Erol-Kantarci, “Reinforcement learning-based joint power and resource allocation for URLLC in 5G,” in Proc. IEEE Global Communications Conference (GLOBECOM), 2019. [11] X. Li, W. Zhou, H. Zhang, J. Zhao, D. Zhao, and Z. Dong, “Joint subcarrier and power allocation in mobile scenario of the OFDM systems based on deep reinforcement learning,” in Proc. International Conference on Computer, Communication and Control Systems (ICCCS), 2023, pp. 209–214. [12] J. Kim, Y. Jeon, J. Lee, M. Lee, and T. Kwon, “Joint scheduling and resource allocation based on reinforcement learning in integrated access and backhaul networks,” ICT Express, vol. 11, no. 3, pp. 536–541, 2025. [13] J. Jang, H. Lyu, D. J. Love, and H. J. Yang, “Joint optimization of user association and resource allocation for load balancing with multi-level fairness,” arXiv preprint arXiv:2505.08573, 2025. [14] A. E. Dorcheh, T. Seyfi, and F. Afghah, “Dora: Dynamic o-ran resource allocation for multi-slice 5g networks,” in 2025 IEEE Middle East Conference on Communications and Networking (MECOM), 2025, pp. 1–6. [15] P. Tarafder, Z. Hassan, I. Ahmed, D. B. Rawat, K. Hasan, and C. Pu, “Digital-twin empowered deep reinforcement learning for site-specific radio resource management in nextg wireless aerial corridor,” arXiv preprint arXiv:2602.03801, 2026, submitted for possible publication to IEEE. [16] M. Elloumi, M. Z. Hassan, and G. Kaddoum, “Uplink radio resource block and power coordination in open ran-digital twin-integrated multicell internet of drone networks,” TechRxiv, Feb. 2026, posted on 2 Feb 2026; preprint. [17] X. Chen, J. Cao, Z. Liang, Y. Sahni, and M. Zhang, “Digital twinassisted reinforcement learning for resource-aware microservice offloading in edge computing,” in 2023 IEEE 20th International Conference on Mobile Ad Hoc and Smart Systems (MASS), 2023. [18] H. Tong, M. Chen, J. Zhao, Y. Hu, Z. Yang, Y. Liu, and C. Yin, “Continual reinforcement learning for digital twin synchronization optimization,” IEEE Transactions on Mobile Computing, vol. 24, no. 8, pp. 6843–6857, 2025. [19] M. Benzaghta, S. Ammar, D. López-Pérez, B. Shihada, and G. Geraci, “Data-driven cellular mobility management via bayesian optimization and reinforcement learning,” arXiv preprint arXiv:2505.21249, 2025. [20] R. Barker, A. E. Dorcheh, T. Seyfi, and F. Afghah, “Real: Reinforcement learning-enabled xapps for experimental closed-loop optimization in oran with osc ric and srsran,” in 2025 IEEE International Conference on Communications Workshops (ICC Workshops), 2025, pp. 389–395.