Conceptio › Archive › arXiv CS
arXiv CSopen access

QoS-Aware RACH Preamble Slicing via Quota-Projected Branching Deep Reinforcement Learning

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
distributed-systemsinternetnetworkingprotocols
networking, internet, protocols, distributed systems

1

QoS-Aware RACH Preamble Slicing via Quota-Projected Branching Deep Reinforcement Learning Jiulin Guo† , Student Member, IEEE, Jiahan Xu† , Jiashuo Zhang† , Heng Yang , Member, IEEE,

arXiv:2609.08199v1 [cs.NI] 8 Sep 2026

Yizhen Sun , Yutong Xie , Shanshan Li , Zhenyu Liu , and Lei Zhang

Abstract—Quality-of-service (QoS)-aware random access requires adaptive allocation of a finite random access channel (RACH) preamble budget across heterogeneous traffic and access procedures. This paper proposes QP-BD3QN-RACH, a quotaprojected branching deep reinforcement learning controller for mixed two-step (2RA) and four-step (4RA) contention-based random access. Four action branches correspond to the delaysensitive and delay-tolerant 2RA/4RA preamble pools. A branching dueling Double DQN selects pool-specific multipliers, and deterministic quota projection converts them to nonnegative integer allocations that preserve the preamble budget. With five actions per branch, the controller represents 625 pre-projection branch-action tuples using 20 branch-action outputs. Evaluation covers five arrival loads, cross-method comparison under nominal seed 42, six-seed sensitivity of QP-BD3QN-RACH, and targeted ablations. Across the five-load grid, its mean direction-aligned differences relative to four comparators are positive: 5.74 to 8.21 percentage points for success/collision, 1.23 to 1.92 percentage points for fallback, 0.35 to 0.68 percentage points for blocking, and 0.128 to 0.456 decision intervals for successful-access delay. Load-wise results exhibit metric-dependent tradeoffs, particularly under intermediate and overload conditions. Index Terms—Action branching, contention-based random access, deep reinforcement learning, massive IoT, preamble slicing, QoS-aware access control, quota projection, random access channel.

I. I NTRODUCTION ASSIVE Internet of Things (IoT) networks must accommodate dense sporadic access while sustaining timely network entry for latency-sensitive services. In New Radio (NR), contention-based random access (CBRA) allows unscheduled user equipment (UE) to initiate access by selecting

M

This work was supported in part by the Natural Science Foundation of Liaoning Province under Grant 20180520022, Grant 2024MSLH359 and Grant 2025MSLH547; in part by the Scientific Research Fund of Liaoning Provincial Education Department under Grant LJ212410142021, Grant LJ222410142060 and Grant LJ212510142045; and in part by the UniversityIndustry Collaborative Education Program of the Ministry of Education under Grant 231106665072727 and Grant 231106665030834. (Corresponding authors: Heng Yang; Shanshan Li.) Jiulin Guo, Jiahan Xu, Jiashuo Zhang, Heng Yang, Yizhen Sun, Yutong Xie, Zhenyu Liu, and Lei Zhang are with the School of Information Science and Engineering, Shenyang University of Technology, Shenyang 110870, Liaoning, China (e-mail: [email protected]; [email protected]; [email protected]; [email protected]; [email protected]; [email protected]; [email protected]; [email protected]). Shanshan Li is with the School of Mechanical Engineering, Shenyang University of Technology, Shenyang 110870, Liaoning, China (e-mail: [email protected]). † Jiulin Guo, Jiahan Xu, and Jiashuo Zhang contributed equally.

random access channel (RACH) preambles, with contention resolution (CR) incorporated into the access procedure when multiple UEs select the same preamble. The 3GPP serviceaccessibility specification defines service-level requirements for heterogeneous services [1]. The NR architecture specifies two-step random access (2RA) and four-step random access (4RA) procedures [2], while the NR medium access control (MAC) specification defines contention-resolution and retransmission behavior [3]. When heterogeneous human-tohuman (H2H) and machine-to-machine (M2M) arrivals share a finite preamble budget, repeated contention can propagate into fallback, backlog accumulation, and blocking. These characteristics motivate adaptive preamble allocation across traffic classes and access procedures. Accordingly, a central line of RACH research has focused on adaptive resource control at the MAC layer. Turan et al. learn class-specific access class barring (ACB) rates for radio access network slices [4]; Lee et al. combine preamble expansion with resource-allocation waiting to mitigate access failures [5]; Althumali et al. adapt service-class preamble partitions according to traffic load [6]; and Chowdhury and De prioritize ACB using queue occupancy [7]. Gedikli et al. further use deep reinforcement learning (DRL) to allocate preamble subsets among service classes [8], while Piao and Lee jointly configure 2RA/4RA resources for heterogeneous IoT traffic [9]. Together, these studies establish dynamic RACH resource partitioning as an effective means of accommodating heterogeneous access demands. Beyond direct MAC-layer resource partitioning, access performance has also been improved through signaling, receiver processing, and alternative access structures. Jiang et al. introduce super-preambles for grant-free massive-MIMO access [10]; Jang et al. classify collided preambles and timingadvance values using deep neural networks [11]; He and Ren exploit access-point clustering for pilot-collision resolution [12]; and Khairy et al. optimize non-orthogonal multiple access (NOMA) transmission probabilities in dense multicell IoT networks [13]. Liva and Polyanskiy survey unsourced multiple access as a coding-oriented framework for massive random access [14]. More recent studies combine NOMA, 2RA, ACB, and contextual bandits for low-latency access [15], employ reconfigurable-intelligent-surface-assisted grant-free transmission opportunities [16], or learn delayoriented sensing-free access strategies [17]. The present study

2

addresses adaptive allocation of a finite preamble budget for mixed 2RA/4RA CBRA. The time-varying nature of traffic, contention, and backlog further motivates learning-based RACH control. TelloOquendo et al. learn the ACB barring rate under mixed H2H/M2M traffic [18]; Pacheco-Paramo et al. apply double deep reinforcement learning to dynamic ACB [19] and subsequently extend the control space to include randomaccess-opportunity periodicity for delay management [20]. Bui and Pham jointly tune the barring factor and barring time with a dueling deep Q-network [21]. Jiang et al. coordinate learned ACB, backoff, and distributed-queuing controls [22], whereas Bai et al. employ a branching actor–critic for intelligent preamble selection [23]. Fan et al. combine hierarchical ACB/backoff with multi-agent DRL for prioritized access [24]; Liu et al. apply a deep Q-network (DQN) to dynamic H2H/M2M RACH allocation [25]; and Elmeligy et al. optimize priority-aware preamble-selection probabilities using a multi-armed bandit [26]. These studies establish access feedback as a basis for online adaptation, with control variables centered primarily on barring, access timing, and preamble selection. The learning architecture must additionally accommodate incomplete network information and multidimensional discrete decisions. Mnih et al. establish DQN as a deep valuelearning framework [27]; Van Hasselt et al. introduce Double DQN (DDQN) to mitigate value overestimation [28]; and Wang et al. develop dueling state-value/action-advantage decomposition [29]. Building on these foundations, Wang et al. formulate dynamic multichannel access as a partially observable Markov decision process and adapt DQN to time-varying channel dynamics [30]. Tavakoli et al. address multidimensional discrete control through action branching with perdimension value heads [31]. Nisioti and Thomos formulate adaptive irregular-repetition slotted ALOHA as a decentralized partially observable Markov decision process and accelerate Q-learning through virtual experience [32]. Du et al. subsequently survey learning-based massive access and resource allocation as components of intelligent wireless-network management [33], while Sohaib et al. apply a branching dueling Q-network to dynamic multichannel random access [34]. These developments provide the methodological basis for observation-driven, scalable value learning over structured access decisions. Taken together, the literature identifies three requirements pertinent to the present problem: adaptive allocation of a finite RACH resource, learning from aggregate access feedback, and scalable representation of multidimensional discrete decisions. QP-BD3QN-RACH addresses these requirements by associating one branch with each of four preamble pools defined by quality-of-service (QoS) class—delay-sensitive (DS) or delay-tolerant (DT)—and access mode—2RA or 4RA—and by projecting the selected branch multipliers to a feasible integer quota vector. Branch-wise value learning supports action selection, quota projection enforces the executable preamble budget, and subsequent access outcomes provide feedback for the next decision interval. The resulting architecture therefore couples the value-function representation directly to the con-

trollable RACH allocation dimensions. This paper makes the following contributions. • A finite-budget preamble-slicing formulation is established for mixed 2RA/4RA CBRA with four controllable QoS/access-mode pools. • QP-BD3QN-RACH couples branch-wise value decisions with deterministic quota projection, producing nonnegative integer preamble allocations that preserve the CBRA budget. • A procedure-level contention-resolution model is paired with explicit numerator/denominator definitions for success, collision, fallback, blocking, and successful-access delay. • Cross-load method comparison, six-seed sensitivity analysis for QP-BD3QN-RACH, and targeted ablations characterize performance and design sensitivity under a common simulation framework. The remainder of this paper is organized as follows. Section II formulates the system and control problem, Section III presents QP-BD3QN-RACH, Section IV reports the simulation evaluation, and Section V summarizes the findings and limitations. II. S YSTEM M ODEL AND P ROBLEM F ORMULATION This section develops the CBRA system model, procedurelevel contention accounting, performance measures, and the partially observed preamble-slicing problem. Table I summarizes the main notation, and Fig. 1 illustrates the scenario interpretation and access process. A. Network, Traffic, and Decision-Interval Model Consider one gNB serving H2H and M2M UEs over 𝑁CBRA CBRA preambles. Each simulation epoch comprises 𝑇 RACH decision intervals indexed by 𝑡 ∈ {1, . . . , 𝑇 }. The epoch index 𝑒 is suppressed in the interval-level model and included explicitly in cross-epoch aggregation. The active UE set in an interval is U𝑡 = N𝑡 ∪ L𝑡 ,

N𝑡 ∩ L𝑡 = ∅,

(1)

where N𝑡 contains new arrivals and L𝑡 contains UEs carried over from preceding intervals. Each UE 𝑢 is characterized by a QoS class 𝑐 𝑢 ∈ {DS, DT}, a user type ℎ𝑢 ∈ {H2H, M2M}, and a current access mode 𝑚 𝑢,𝑡 ∈ {2RA, 4RA}. The user-type and QoS labels are distinct model attributes; their evaluated joint distribution is specified in Section IV. The proposed controller allocates preambles over four QoS/access-mode groups G = {DS2, DT2, DS4, DT4}, with 𝑔(𝑢, 𝑡) =

 DS2,      DT2,   DS4,     DT4, 

𝑐 𝑢 = DS, 𝑚 𝑢,𝑡 = 2RA, 𝑐 𝑢 = DT, 𝑚 𝑢,𝑡 = 2RA, 𝑐 𝑢 = DS, 𝑚 𝑢,𝑡 = 4RA, 𝑐 𝑢 = DT, 𝑚 𝑢,𝑡 = 4RA.

(2)

(3)

3

TABLE I M AIN N OTATION U SED T HROUGHOUT THE M ANUSCRIPT Notation

Description

Notation

Description

𝑁CBRA B U𝑡 L𝑡 𝐴𝑢,𝑡 𝑚𝑢,𝑡 𝐹2 , 𝐹4 𝑁𝑡act 𝑁𝑡succ 𝑁𝑡fb 𝑁𝑡res q𝑡 𝑛𝑡rem 𝜓𝑢,𝑡 𝝆 s𝑡 𝝌𝑡 𝜃, 𝜃− 𝜂 bs 𝜂 bfb bsucc 𝐷 Λ 𝑀, 𝑡warm 𝑛step , 𝑛upd

Number of CBRA preambles Branch set of QP-BD3QN-RACH Active UE set at decision interval 𝑡 Backlog carried into decision interval 𝑡 ACB allow indicator of UE 𝑢 Current access mode of UE 𝑢 Failure thresholds for 2RA fallback and 4RA blocking Active UE count Post-CR successful-UE count Two-step fallback count UEs resolved from contended buckets Integer preamble quota vector Residual preamble count after flooring Preamble selected by UE 𝑢 Baseline preamble-share vector Aggregate controller observation Latent UE-level simulator state Online and target network parameters Success per attempt Fallback per active 2RA UE Delay per successful UE Evaluated arrival-load set Mini-batch size and global warm-up boundary Interaction and update counters

G W N𝑡 I𝑡 𝑝𝑡 𝐵𝑢,𝑡 𝑔 (𝑢, 𝑡 ) 𝑁𝑡att 𝑁𝑡coll 𝑁𝑡blk 𝑁𝑡unres a𝑡 𝑖 𝜋 Π 𝑁CBRA (·) 𝑟𝑡 𝛾 D 𝜂 bc 𝜂 bblk Ttail Z 𝑂 𝑈

QoS/access-mode preamble-control groups Candidate branch multiplier set New-arrival UE set ACB-allowed attempted UE set Exogenous ACB allow probability Four-step blocking indicator QoS/access-mode group of UE 𝑢 Attempted UE count Post-CR unresolved-UE count Four-step blocking count UEs unresolved in contended buckets Branch action vector Preamble index Observation-based control policy Quota-projection operator Scalar reward Discount factor Replay memory Collision per attempt Blocking per active 4RA UE Tail set of epoch–interval pairs Seed set for proposed-method diagnostics Value-network output dimension Target-network update period

The proposed controller assigns quotas by QoS and current access mode; user type enters the traffic model and subgroup statistics. Fig. 1 summarizes the modeled access sequence from arrivals and ACB gating to preamble contention, procedurelevel resolution, fallback, blocking, and backlog carryover. A gNB-side controller updates the four preamble quotas once per decision interval using aggregate traffic, preceding quota allocation, recent access outcomes, and backlog information. Assumption 1 (Controller information). Before currentinterval arrivals are introduced, the controller receives pending-UE counts by QoS/access mode and by access mode/user type, preceding quota shares, the nominal arrival load, and recent access outcomes. Action selection precedes the introduction of new UEs.

B. ACB Gate and Attempt Process Before preamble selection, each active UE passes through an ACB gate. Let 𝐴𝑢,𝑡 ∈ {0, 1} denote the access decision, with Pr( 𝐴𝑢,𝑡 = 1) = 𝑝 𝑡 ,

0 ≤ 𝑝 𝑡 ≤ 1,

(4)

where 𝑝 𝑡 is an exogenous ACB allow probability. The set of UEs admitted to the current access attempt is I𝑡 = {𝑢 ∈ U𝑡 : 𝐴𝑢,𝑡 = 1},

(5)

𝑁𝑡act = |U𝑡 |,

(6)

The simulator assigns each arriving UE an initial access mode according to the traffic-generation rule specified in Section IV. For overhead accounting, each 2RA attempt is assigned a nominal message cost of two and each 4RA attempt a cost of four: ( 2, 𝑚 𝑢,𝑡 = 2RA, msg (7) 𝐶𝑢,𝑡 = 4, 𝑚 𝑢,𝑡 = 4RA. The aggregate nominal message cost in interval 𝑡 is ∑︁ msg msg 𝐶𝑢,𝑡 . 𝐶𝑡 =

(8)

𝑢∈ I𝑡

This quantity summarizes nominal procedure overhead for the diagnostic analysis. Each UE maintains a per-mode counter initialized to zero at arrival. The counter increases once per unsuccessful decision interval, including ACB denial and unresolved preamble contention, and resets upon successful resolution or 2RA fallback. In the threshold tests below, 𝑓𝑢,𝑡 denotes the value after the current outcome is recorded and before any fallback reset. The mode and blocking updates are ( 4RA, 𝑚 𝑢,𝑡 = 2RA, 𝑓𝑢,𝑡 ≥ 𝐹2 , 𝑚 𝑢,𝑡+1 = (9) 𝑚 𝑢,𝑡 , otherwise, 𝐵𝑢,𝑡+1 = 1{𝑚 𝑢,𝑡 = 4RA, 𝑓𝑢,𝑡 ≥ 𝐹4 },

(10)

where 𝐹2 and 𝐹4 are the per-mode unsuccessful-interval thresholds. The resulting event counts are

and 𝑁𝑡att = |I𝑡 |.

C. Two-Step and Four-Step Procedure Dynamics

The evaluated value of 𝑝 𝑡 is specified with the simulation configuration in Section IV.

𝑁𝑡fb = {𝑢 ∈ U𝑡 : 𝑚 𝑢,𝑡 = 2RA, 𝑚 𝑢,𝑡+1 = 4RA} , ∑︁ 𝑁𝑡blk = 𝐵𝑢,𝑡+1 . 𝑢∈ U𝑡

(11) (12)

4

Network and traffic setting DT_4RA Pool (delay-tolerant, 4RA)

0

1

2

···

···

12

CBRA procedure and accounting DT_2RA Pool (delay-tolerant, 2RA)

M2M Sensor UEs

13

14

···

···

···

25

13 preambles

Pool-specific preamble selection Active UEs (from cell)

Collision example (same preamble)

DS_2RA preamble bucket

2RA Access Lane

Selected preamble

MsgB

MsgA

Winner UE (Success)

13 preambles

DT_2RA preamble bucket

ACB Gate

Contention Resolution (one-winner UE identity resolution

DS_4RA preamble bucket

(admission control before access)

CR Losers (same preamble)

4RA Access Lane Msg1

Msg2

Msg3

Msg4

DT_4RA preamble bucket

gNB Post-CR loser

DS_4RA Pool (delay-sensitive, 4RA)

DS_2RA Pool (delay-sensitive, 2RA)

26

27

···

···

···

39

··· ··· 53

14 preambles

14 preambles

H2H Smartphone UEs

40

41

Current mode: 2RA

···

False

Remain 2RA

Finite CBRA resources NCBRA = 54

Current mode: 4RA

Fallback to 4RA

Backlog Loop (Unresolved / Active UEs) Backlog Queue (Unresolved/Active UEs waiting for access)

True

False

True

Remain 4RA

Blocking

•

Re-evaluate ACB gate

•

Re-select preamble

•

Re-attempt access

•

Repeat until success,

Removed

fallback/blocking

Fig. 1. System model for QoS-aware CBRA preamble slicing with H2H/M2M arrivals, ACB gating, 2RA/4RA access, procedure-level contention resolution, fallback, blocking, and backlog carryover. The displayed 13/14/14/13 quotas illustrate one feasible preamble allocation.

Successfully resolved and blocked UEs leave the active population; all remaining UEs form the next backlog: L𝑡+1 = {𝑢 ∈ U𝑡 : 𝑢 ∉ S𝑡succ , 𝐵𝑢,𝑡+1 = 0},

(13)

where S𝑡succ is the set of successfully resolved UEs. D. Procedure-Level Contention-Resolution Accounting For a nonempty quota pool, each admitted UE independently selects one preamble uniformly from that pool. Using disjoint pool indices within {1, . . . , 𝑁CBRA }, let 𝜓𝑢,𝑡 denote the selected preamble and define P𝑡 ,𝑖 = {𝑢 ∈ I𝑡 : 𝜓𝑢,𝑡 = 𝑖}.

(14)

Assumption 2 (Contention-resolution accounting). A singleton preamble bucket contributes one resolved UE. For each non-singleton bucket, one UE is recorded as resolved and the remaining UEs are recorded as unresolved for the current attempt. The accounting operates at the random-access procedure level; physical-layer capture and multi-user decoding are not modeled.

{𝑖 : |P𝑡 ,𝑖 | = 1} ,

(15)

{𝑖 : |P𝑡 ,𝑖 | ≥ 2} .

(16)

Under Assumption 2, 𝑁𝑡res = 𝑁𝑡cb , ∑︁ 𝑁𝑡unres = 𝑖:| P𝑡,𝑖 | ≥2

sing

𝑁𝑡succ = 𝑁𝑡

+ 𝑁𝑡res ,

𝑁𝑡coll = 𝑁𝑡unres .

(17) |P𝑡 ,𝑖 | − 1 . 

(18)

(19) (20)

Here, 𝑁𝑡coll denotes UEs remaining unresolved after the procedure-level contention accounting. When every attempted group has a nonempty pool, each attempted UE is recorded once as resolved or unresolved, giving 𝑁𝑡att = 𝑁𝑡succ + 𝑁𝑡coll .

(21)

E. Performance Metrics and Aggregation For an analysis set T of epoch–interval pairs (𝑒, 𝑡), define Í (𝑒,𝑡 ) ∈ T 𝑥 𝑒,𝑡 b 𝑅 𝑥/𝑦 (T ) = Í , (22) (𝑒,𝑡 ) ∈ T 𝑦 𝑒,𝑡 for nonnegative quantities with a positive total denominator. Numerators and denominators are pooled separately across the selected intervals. Interval-level ratios used in observations and rewards are assigned zero when their denominators vanish. 1) Primary Metrics: Let 𝑁𝑡act,𝑚 = {𝑢 ∈ U𝑡 : 𝑚 𝑢,𝑡 = 𝑚} ,

Let sing 𝑁𝑡 = 𝑁𝑡cb =

The success and collision counts used throughout the manuscript are therefore

𝑚 ∈ {2RA, 4RA}, (23)

Let 𝜏𝑢 be the arrival interval of UE 𝑢. Successful resolution in interval 𝑡 gives 𝑑𝑢,𝑡 = 𝑡 − 𝜏𝑢 , measured in decision intervals. Success in the arrival interval has zero delay, and waiting during ACB denial contributes to the elapsed delay. Modespecific delay summaries group successful UEs by initial access mode. The successful-user delay sum is ∑︁ 𝐷 succ = 𝑑𝑢,𝑡 . (24) 𝑡 𝑢∈ S𝑡succ

5

The primary metrics are b𝑁 succ /𝑁 att (T ), 𝜂bs (T ) = 𝑅 b𝑁 coll /𝑁 att (T ), 𝜂bc (T ) = 𝑅

b𝑁 fb /𝑁 act,2RA (T ), 𝜂bfb (T ) = 𝑅 b𝑁 blk /𝑁 act,4RA (T ), 𝜂bblk (T ) = 𝑅

bsucc (T ) = 𝑅 b𝐷 succ /𝑁 succ (T ). 𝐷

(25a) (25b) (25c) (25d) (25e)

Equation (21) further gives 𝜂bs (T ) + 𝜂bc (T ) = 1

(26)

The pending subsets L𝑡 ,𝑔 and L𝑡 ,𝑚,ℎ group UEs by QoS/current access mode and by current access mode/user type, respectively. Their normalized counts form  pend x𝑡 = (|L𝑡 ,𝑔 |/𝐻𝜆 )𝑔∈ G , (35)  (|L𝑡 ,𝑚,ℎ |/𝐻𝜆 ) 𝑚,ℎ , |L𝑡 |/𝐻𝜆 . Group entries follow DS2, DT2, DS4, DT4. Mode/type entries follow 2RA–H2H, 2RA–M2M, 4RA–H2H, 4RA–M2M. The preceding global outcomes are  act perf x𝑡 = 𝜂s,𝑡 −1 , 𝜂c,𝑡 −1 , 𝜂blk,𝑡 −1 , (36)  𝜂fb,𝑡 −1 , 𝐷¯ 𝑡 −1 /10 ,

for horizons containing at least one attempted UE and nonempty pools for all attempted groups. The evaluated configurations satisfy this condition. 2) Diagnostic Metrics: Active and attempted access pressure per preamble are Í act (𝑒,𝑡 ) ∈ T 𝑁 𝑒,𝑡 , (27) b 𝜌act (T ) = |T |𝑁CBRA Í att (𝑒,𝑡 ) ∈ T 𝑁 𝑒,𝑡 b 𝜌att (T ) = . (28) |T |𝑁CBRA

act where 𝜂s,𝑡 = 𝑁𝑡succ /max(𝑁𝑡act , 1) and 𝐷¯ 𝑡 = succ succ 𝐷 𝑡 /max(𝑁𝑡 , 1). The other rate entries use their interval-level denominators from Section II-E. DS-specific outcomes form  DS4,ℎ DS2,ℎ xDS 𝑡 = (𝜂blk,𝑡 −1 ) ℎ , (𝜂fb,𝑡 −1 ) ℎ , (37)  ¯ DS2,ℎ ( 𝐷¯ DS4,ℎ 𝑡 −1 /10) ℎ , ( 𝐷 𝑡 −1 /10) ℎ ,

b𝑁 cont,UE /𝑁 att (T ). 𝜂bcont (T ) = 𝑅

where x𝑡 = (𝑥 𝑡 ,𝑔 )𝑔∈ G and 𝑝 𝑡 is the preceding logged ACB probability. The six blocks contain 9, 4, 1, 1, 5, and 8 entries, respectively. At epoch initialization, pending counts and preceding metric entries, including the previous-ACB feature, are zero; quota features reflect the static allocation. 2) Four-Branch Action Space: The branch set follows the executable quota dimensions:

Let 𝑁𝑡cont,UE denote the number of attempted UEs belonging to non-singleton preamble buckets before procedure-level resolution. The corresponding contention pressure is

Mean post-interval backlog over the tail window is ∑︁ 1 b 𝐿 tail = |L𝑒,𝑡+1 |. |Ttail |

(29)

(30)

(𝑒,𝑡 ) ∈ Ttail

Here L𝑒,𝑡+1 is the backlog after interval 𝑡 of epoch 𝑒; L𝑒,𝑇+1 records the final-interval backlog before the next epoch resets. Nominal message cost is reported per attempt and per successful UE: batt

b𝐶 msg /𝑁 att (T ), 𝐶 (T ) = 𝑅 succ b (T ) = 𝑅 b𝐶 msg /𝑁 succ (T ). 𝐶

with ℎ ordered as H2H, M2M. Subgroup rates use currentmode active populations; subgroup delays use successful UEs grouped by initial mode. The flattened observation is  pend quota s𝑡 = x𝑡 , x𝑡 , 𝜆/100, (38)  prev perf 𝑝 𝑡 , x𝑡 , xDS ∈ R28 , 𝑡 quota

These quantities are used as secondary diagnostics for access pressure, backlog, contention, and signaling overhead. F. Partially Observed Preamble-Slicing Problem 1) Controller Observation: Let 𝝌𝑡 denote the latent UElevel simulator state containing the variables required for access-mode evolution, failure histories, realized contention, and backlog evolution. Under Assumption 1, the controller receives only the aggregate observation s𝑡 . The policy is therefore written as 𝜋(a𝑡 | s𝑡 ), (33) with no requirement that s𝑡 constitute a complete Markov state. Let 𝜆 denote the nominal number of new UEs per interval and set 𝐻𝜆 = max{𝑁CBRA , 5𝜆, 1}, (34) quota prev 𝑥 𝑡 ,𝑔 = 𝑞 𝑡 ,𝑔 /𝑁CBRA .

prev

B = G.

(39)

Each branch selects one multiplier from

(31) (32)

quota

giving

W = {0.50, 0.75, 1.00, 1.25, 1.50},

(40)

a𝑡 = (𝑎 𝑡 ,𝑏 ) 𝑏∈ B ∈ A = W | B | .

(41)

The four branches therefore define (4) 𝑁tuple = |W| | B | = 54 = 625

(42)

pre-projection branch-action tuples. 3) Quota Projection: Let 𝝆 = (𝜌 𝑏 ) 𝑏∈ B denote the fixed baseline preamble-share vector, with ∑︁ 𝜌 𝑏 ≥ 0, 𝜌 𝑏 = 1. (43) 𝑏∈ B

For branch 𝑏, the selected multiplier produces the score 𝑧 𝑡 ,𝑏 = 𝑎 𝑡 ,𝑏 𝜌 𝑏 .

(44)

Since all elements of W are positive, 𝑏∈ B 𝑧 𝑡 ,𝑏 > 0, and the normalized real-valued quota is 𝑧 𝑡 ,𝑏 . (45) 𝑞˜ 𝑡 ,𝑏 = 𝑁CBRA Í 𝑗 ∈ B 𝑧𝑡 , 𝑗 Í

6

Initial integer quotas are obtained by flooring:   𝑞 𝑡(0) ,𝑏 = 𝑞˜ 𝑡 ,𝑏 , leaving

𝑛rem = 𝑁CBRA − 𝑡

∑︁

𝑞 𝑡(0) ,𝑏

(46) (47)

𝑏∈ B

preambles to be assigned. Let ≺ B order the branches as DS2, DT2, DS4, DT4. For ℓ = 1, . . . , 𝑛rem 𝑡 , the residual allocation selects 𝑏★ℓ ∈

arg min

𝑏∈ B:𝑧𝑡,𝑏 >0

𝑞 𝑡(ℓ,𝑏−1) 𝑧 𝑡 ,𝑏

,

(48)

(49)

(𝑛rem )

,

𝑏 ∈ B,

(50)

or, equivalently, q𝑡 = Π 𝑁CBRA (a𝑡 , 𝝆).

𝑏∈ B

Proof. Flooring produces nonnegative integer quotas whose sum is at most 𝑁CBRA . The subsequent 𝑛rem single-unit assign𝑡 ments increase the aggregate quota by exactly 𝑛rem 𝑡 , yielding a final sum of 𝑁CBRA . Under the proposed four-pool formulation, each branch controls one QoS/access-mode quota. 4) Reward and Objective: The scalar reward 𝑟 𝑡 is a dimensionless training score combining access success, collision, fallback, blocking, delay, and DS-specific penalties:

+ 𝑤 blk 𝜂blk,𝑡 + 𝑤 fb 𝜂fb,𝑡 + 𝑤 d4

𝐷¯ 4,𝑡 10

sen 𝛿d,𝑡 𝐷¯ 2,𝑡 sen sen + 𝑤 sb 𝛿blk,𝑡 + 𝑤 sf 𝛿fb,𝑡 + 𝑤 sd . (53) 10 10 Here 𝜂s2,𝑡 and 𝜂s4,𝑡 use current-mode active populations. The delay summaries 𝐷¯ 2,𝑡 and 𝐷¯ 4,𝑡 group successful UEs by initial access mode. Collision, fallback, and blocking use the intervallevel counterparts of the ratios in Section II-E. DS-specific rates use the corresponding current-mode active populations, and DS-specific delays retain initial-mode grouping. The DSspecific terms are

+ 𝑤 d2

DS4,H2H DS4,M2M sen 𝛿blk,𝑡 = 0.5 𝜂blk,𝑡 + 0.5 𝜂blk,𝑡 , DS2,H2H DS2,M2M sen 𝛿fb,𝑡 = 0.5 𝜂fb,𝑡 + 0.5 𝜂fb,𝑡 ,

sen 𝛿d,𝑡 = 0.3 𝐷¯ DS4,H2H + 0.2 𝐷¯ DS4,M2M 𝑡 𝑡 DS2,H2H ¯ + 0.3 𝐷 𝑡 + 0.2 𝐷¯ DS2,M2M . 𝑡

(54a) (54b)

(54c)

The default coefficient vector is wdefault = (1.50, 1.50, −0.50, −0.60, −0.60,

− 0.25, −0.40, −1.00, −0.75, −0.40),

𝑎 𝑡 ,𝑏 ∈ W,

∀𝑏 ∈ B,

𝑏∈ B

(56b) (56c) (56d) (56e)

The environment transition follows the ACB, access-mode, contention-resolution, and backlog dynamics defined in the preceding subsections.

(51)

Proposition 1 (Quota feasibility). The projection Π 𝑁CBRA returns an integer quota vector satisfying ∑︁ 𝑞 𝑡 ,𝑏 = 𝑁CBRA . (52) 𝑞 𝑡 ,𝑏 ∈ Z ≥0 ,

𝑟 𝑡 = 𝑤 s4 𝜂s4,𝑡 + 𝑤 s2 𝜂s2,𝑡 + 𝑤 c 𝜂c,𝑡

s.t.

𝑡=1

q𝑡 = Π 𝑁CBRA (a𝑡 , 𝝆).

The final allocation is 𝑞 𝑡 ,𝑏 = 𝑞 𝑡 ,𝑏𝑡

𝜋

𝑞 𝑡 ,𝑏 ∈ Z ≥0 , ∀𝑏 ∈ B, ∑︁ 𝑞 𝑡 ,𝑏 = 𝑁CBRA ,

with ties resolved by ≺ B , and updates

𝑞 𝑡(ℓ,𝑏) = 𝑞 𝑡(ℓ,𝑏−1) + 1{𝑏 = 𝑏★ℓ }.

ordered according to the terms in (53). Delay terms are scaled by 1/10 before weighting. The coefficient vector is fixed across the main comparisons; the reward-profile ablation is specified in Section IV. Performance comparison uses the primary metrics in Section II-E. The controller seeks an observation-based policy maximizing discounted return over one 𝑇-interval epoch: " 𝑇 # ∑︁ 𝑡 −1 max 𝐽 (𝜋) = E 𝜋 𝛾 𝑟𝑡 (56a)

(55)

III. T HE QP-BD3QN-RACH C ONTROLLER QP-BD3QN-RACH maps the aggregate observation s𝑡 to one multiplier decision for each of the four executable preamble-quota dimensions defined in Section II-F. As illustrated in Fig. 2, a shared encoder and branch-wise dueling value heads support action selection, while the deterministic operator Π 𝑁CBRA converts the selected multipliers to a budgetfeasible integer quota vector. The resulting CBRA outcomes provide the reward and next observation used for replaybased learning. This section specifies the branch-wise value representation, action selection, Double DQN update, training procedure, and output dimensionality. A. Branching Dueling Value Representation and Action Selection The controller processes s𝑡 with two fully connected hidden layers and rectified linear unit (ReLU) activations. Each layer has width 128 in the default configuration, and Adam minimizes the mini-batch loss. The scalar value/action-advantage decomposition follows the dueling architecture of Wang et al. [29]. The shared representation with one action branch per decision dimension follows the action-branching architecture introduced by Tavakoli et al. [31]. The shared encoder produces h𝑡 = 𝑓 𝜃 (s𝑡 ),

(57)

from which the network estimates a scalar value 𝑉 𝜃 (s𝑡 ) and a branch-specific advantage 𝐴 𝜃 ,𝑏 (s𝑡 , 𝑎 𝑏 ) for each 𝑏 ∈ B and 𝑎 𝑏 ∈ W. The corresponding branch action value is 1 ∑︁ 𝐴 𝜃 ,𝑏 (s𝑡 , 𝑎 ′ ). 𝑄 𝜃 ,𝑏 (s𝑡 , 𝑎 𝑏 ) = 𝑉 𝜃 (s𝑡 ) + 𝐴 𝜃 ,𝑏 (s𝑡 , 𝑎 𝑏 ) − |W| ′ 𝑎 ∈W (58) The greedy action of branch 𝑏 is 𝑎★𝑡,𝑏 = arg max 𝑄 𝜃 ,𝑏 (s𝑡 , 𝑎 𝑏 ). 𝑎𝑏 ∈ W

(59)

7

Observation

Shared Encoder (neural network trunk)

Dueling Head (branching)

Access pressure

Branch action selector

Value Stream •

•

• • •

•

DS-2RA

Recent / previous diagnostics

DT-2RA DS-4RA

MLP/FC layers (activation: ReLU)

Output: state vector

DT-4RA

Output: encoded feature

Output: branch values

Transition record

Replay memory

0.75

1.00

1.25

1.50

CBRA procedure (simulator)

Reward and metric computation

Epoch metrics

Fixed ACB allow gate

Composite rewards (success, fallback, collision, blocking, delay)

Per-epoch metrics

Pool shares

Advantage Branches

•

Branch choices

DS_2RA 0.50

Quota shares

Quota Projection

DT_2RA 0.50

0.75

1.00

1.50

Integer quotas

Pool contention

1.25

1.50

Fixed CBRA preamble budget

One-winner post-CR accounting

1.25

1.50

Per-pool integer quotas

Fallback / blocking

Output: pool quotas

Backlog update

1.25

DS_4RA 0.50

0.75

0.50

0.75

1.00

DT_4RA 1.00

Output: selected multipliers

Mini-batch update

Online network selects

Target network evaluates

Raw logs

Network update

Output: metrics

Run Manifest (hyperparameters, seeds, ...)

Artifacts (csv logs, plots)

Target sync

Fig. 2. Control and learning architecture of QP-BD3QN-RACH. A shared encoder maps the aggregate observation to branch-wise action values; selected branch multipliers are projected to a feasible integer quota vector; and subsequent CBRA outcomes provide reward and next-observation feedback for replaybased Double DQN updates.

A single exploration draw 𝜉𝑡 ∼ Uniform(0, 1) determines whether the complete branch vector follows exploration or greedy selection. Conditional on exploration, each branch independently samples one action 𝑎 rnd 𝑡 ,𝑏 uniformly from W. The executed action vector is ( rnd (𝑎 𝑡 ,𝑏 ) 𝑏∈ B , 𝜉𝑡 < 𝜀, a𝑡 = (60) (𝑎★𝑡,𝑏 ) 𝑏∈ B , 𝜉𝑡 ≥ 𝜀. The selected branch vector is converted to the executable allocation by q𝑡 = Π 𝑁CBRA (a𝑡 , 𝝆), (61) using the projection defined in Section II-F. Thus, value estimation is factorized across the four decision dimensions, while quota execution remains coupled through the common preamble budget. B. Branch-Specific Double DQN Update Action selection and target evaluation follow the Double DQN separation introduced by Van Hasselt et al. [28]. Among the temporal-difference target constructions examined for action branching, QP-BD3QN-RACH uses the branchspecific form described by Tavakoli et al. [31]. For a replay transition (s𝑡 , a𝑡 , 𝑟 𝑡 , s𝑡+1 , 𝑑𝑡 ), with 𝑑𝑡 = 1{𝑡 = 𝑇 } marking the end of an epoch, the online network first selects (62) 𝑎 +𝑡+1,𝑏 = arg max 𝑄 𝜃 ,𝑏 (s𝑡+1 , 𝑎 𝑏 ). 𝑎𝑏 ∈ W

The target network then evaluates the selected action: 𝑦 𝑡 ,𝑏 = 𝑟 𝑡 + 𝛾(1 − 𝑑𝑡 )𝑄 𝜃 − ,𝑏 (s𝑡+1 , 𝑎 +𝑡+1,𝑏 ).

(63)

The branch-wise temporal-difference error is 𝛿𝑡 ,𝑏 = 𝑦 𝑡 ,𝑏 − 𝑄 𝜃 ,𝑏 (s𝑡 , 𝑎 𝑡 ,𝑏 ), 1 ∑︁ 2 𝛿 . |B| 𝑏∈ B 𝑡 ,𝑏

Input: Training epochs 𝐸 , decision intervals per epoch 𝑇 , branch set B , candidate set W , baseline shares 𝝆 , total preambles 𝑁CBRA , exploration probability 𝜀 , discount factor 𝛾 , mini-batch size 𝑀 , warm-up boundary 𝑡warm , target-update period 𝑈 Output: Trained online-network parameters 𝜃 1: Initialize 𝜃 and set 𝜃 − ← 𝜃 . 2: Initialize replay memory D and counters 𝑛step ← 0, 𝑛upd ← 0. 3: for 𝑒 = 1, . . . , 𝐸 do 4: Reset the simulator to an empty backlog, static quotas, and zero preceding metrics. 5: for 𝑡 = 1, . . . , 𝑇 do 6: Receive aggregate observation s𝑡 . 7: During replay warm-up, sample each branch action independently and uniformly from W ; thereafter, select a𝑡 using (59)–(60). 8: Compute q𝑡 = Π 𝑁CBRA (a𝑡 , 𝝆) . 9: Execute the CBRA decision interval under q𝑡 and obtain 𝑟𝑡 , s𝑡+1 , and terminal indicator 𝑑𝑡 . 10: Log the raw metric numerators and denominators. 11: Store (s𝑡 , a𝑡 , 𝑟𝑡 , s𝑡+1 , 𝑑𝑡 ) in D . 12: Set 𝑛step ← 𝑛step + 1. 13: if 𝑛step > 𝑡warm and | D | ≥ 𝑀 then 14: Sample a mini-batch from D and minimize the mini-batch average of (65). 15: Set 𝑛upd ← 𝑛upd + 1. 16: if 𝑛upd is a multiple of 𝑈 then 17: Set 𝜃 − ← 𝜃 . 18: end if 19: end if 20: end for 21: end for

For mini-batch training, 𝐿 TD 𝑡 (𝜃) is averaged over the sampled transitions. Parameter updates begin after the global interaction count exceeds 𝑡 warm and the replay memory contains at least 𝑀 transitions. The target parameters 𝜃 − are synchronized with the online parameters every 𝑈 gradient updates. Algorithm 1 summarizes the interaction and learning cycle. Online and target parameters, replay memory, and global counters persist across epochs; the simulator is reset at each epoch boundary. Raw metric numerators and denominators are logged during simulation and aggregated in Section IV according to (22).

(64)

C. Action-Output Dimensionality For 𝑛 𝐵 decision dimensions with 𝐾 candidate actions per dimension, the Cartesian branch-action set contains

(65)

𝑁tuple = 𝐾 𝑛𝐵

and the transition loss is 𝐿 TD 𝑡 (𝜃) =

Algorithm 1 Training Procedure for QP-BD3QN-RACH

(66)

8

pre-projection tuples. Explicit enumeration of the same symmetric tuple set requires 𝐾 𝑛𝐵 action-dependent value outputs, whereas the branching representation requires 𝑂 branch = 𝑛 𝐵 𝐾

(67)

action-dependent outputs in addition to the shared scalar value stream. For QP-BD3QN-RACH, 𝑛 𝐵 = 4 and 𝐾 = 5, giving 625 pre-projection tuples and 20 branch-action outputs. The eightbranch configuration gives 58 = 390,625 tuples and 40 branchaction outputs. These counts describe the symmetric fivechoice branch grid before quota projection; multiple tuples may map to the same integer quota allocation. The Flat DDQN and Flat D3QN comparators use the separate 108action catalog specified in Section IV. IV. S IMULATION E VALUATION A. Evaluation Configuration and Evidence Sets The evaluation follows the CBRA model and performance measures defined in Section II. Each decision interval introduces a fixed number of new UEs, after which the active population evolves through ACB admission, preamble contention, procedure-level resolution, fallback, blocking, and backlog carryover. The evaluated arrival-load set is Λ = {10, 20, 50, 100, 200}, where each value denotes the number of new UEs introduced per decision interval. Because unresolved UEs remain active through backlog carryover, the active and attempted populations can exceed the nominal arrival load. Performance measures are recomputed from their accumulated numerators and denominators according to (22). For each run, cross-load, seed-sensitivity, and ablation summaries pool all 100 decision intervals from epochs 401–500, so Ttail = {401, . . . , 500} × {1, . . . , 100}. The simulator starts each epoch with an empty backlog and static quotas, while the trainable controllers retain their learned parameters and replay memory. The evaluation comprises three complementary analyses. Cross-load comparison covers the five controller configurations over Λ under nominal seed 42. Seed sensitivity evaluates QP-BD3QN-RACH over the same load set using Z = {42, 23, 16, 15, 8, 4}. The ablation study examines branchaction, reward, warm-up, target-update, and hidden-width variants at loads 20, 100, and 200 under nominal seed 42. Table II summarizes the simulation configuration and the default learning parameters of the trainable controllers. For the trainable controllers, the first 2000 interactions use uniform action sampling; subsequent decisions use the fixed exploration probability listed in Table II. The same traffic-generation, ACB, access-mode, and contention-resolution model is used throughout the evaluation. B. Comparator Configuration Table III defines the five controller configurations used for cross-load comparison and structural analysis. Static ACB uses the fixed ACB probability and static preamble allocation. Flat DDQN and Flat D3QN use the same fixed ACB gate and the same 108-element joint action catalog over the four

TABLE II S IMULATION AND D EFAULT T RAINING C ONFIGURATION Parameter

Value

CBRA preambles Arrival loads Nominal comparison seed Seed-sensitivity set M2M probability DS probability H2H data-amount variable M2M data-amount variable Initial access mode Four-pool static quotas (DS2/DT2/DS4/DT4) 2RA/4RA failure thresholds ACB allow probability Default reward profile Training epochs Decision intervals per epoch Analysis window Warm-up interactions Learning rate Post-warm-up exploration probability Mini-batch size Discount factor Replay capacity Target-update period Hidden width

54 10, 20, 50, 100, 200 42 42, 23, 16, 15, 8, 4 0.8 0.3 for H2H; 0.1 for M2M Discrete uniform on {2, . . . , 8} Discrete uniform on {1, 2, 3} 2RA if data amount ≤ 2; otherwise 4RA 14/13/14/13 𝐹2 = 6, 𝐹4 = 11 0.5 Coefficients in (55) 500 100 Final 100 epochs 2000 3 × 10 −5 0.1 64 0.90 5000 500 gradient updates 128

preamble pools. Flat D3QN adds the dueling value–advantage decomposition to the Double DQN architecture. The catalog is the Cartesian product of the multiplier sets {1, 1.2, 1.4, 1.6} for DS2, {0.8, 1, 1.2} for DT2 and DT4, and {1, 1.2, 1.4} for DS4, giving 4 × 33 = 108 actions. QP-BD3QN-RACH selects one of five multipliers on each of four branches and applies quota projection before execution, yielding 625 pre-projection branch-action tuples. B-D3QNRACH-8 assigns quotas to eight QoS/access-mode/user-type contention groups and performs preamble selection separately within each group. Four-pool quota totals are also recorded for reporting. The four-pool learning controllers use 28 observation features; the eight-group controller uses 35. Together, the configurations compare static and learned four-pool allocation with finer eight-group contention partitioning. C. Training Behavior and Cross-Load Performance Fig. 3 compares training reward in the upper row and validation composite score in the lower row across the five arrival loads under nominal seed 42. The five plotted epoch positions are 100, 200, 300, 400, and 500. The validation composite applies the weighted access-quality terms of the training objective to the validation trajectory and provides a consistent view of checkpoint evolution across training. Let 𝑚 denote QP-BD3QN-RACH. For comparator 𝑏, metric 𝑚 denote the corresponding tail-window 𝑘, and load 𝜆, let 𝑅 𝑘,𝜆 value. Define 𝜎𝑘 = 1 for higher-is-better metrics and 𝜎𝑘 = −1 for lower-is-better metrics. The direction-aligned difference is   𝑚 𝑏 Δ𝑚−𝑏 = 𝜎 𝑅 − 𝑅 (68) 𝑘 𝑘,𝜆 𝑘,𝜆 𝑘,𝜆 , and its average over the evaluated load grid is 1 ∑︁ 𝑚−𝑏 Δ̄𝑚−𝑏 = Δ . 𝑘 |Λ| 𝜆∈Λ 𝑘,𝜆

(69)

9

TABLE III C ONTROLLER C ONFIGURATIONS U SED IN THE E VALUATION Method

Role

Value representation

Action representation

Static ACB Flat DDQN Flat D3QN QP-BD3QN-RACH B-D3QN-RACH-8

Non-learning comparator Learning baseline Learning baseline Proposed controller Structural comparator

Fixed control Double DQN Dueling Double DQN Branching Dueling Double DQN Branching Dueling Double DQN

Fixed ACB and static preamble allocation Single head over 108 joint actions Single head over the same 108 joint actions Four branches with five actions per branch Eight branches with five actions each over eight contention groups

Score

Reward

Flat DDQN 𝜆 = 10

1.2 1.1 1 1.2 1.1 1 0.9

Flat D3QN

𝜆 = 20

−0.1 −0.2 −0.3 −0.4 −0.5

0.8 0.6 0.4 0.8 0.6 0.4 100

300 Epoch

100

500

300 Epoch

QP-BD3QN-RACH 𝜆 = 50

500

−0.2 −0.3 −0.4 −0.5

B-D3QN-RACH-8 𝜆 = 100

−0.6

−0.88

−0.65

−0.9

−0.7

100

300 Epoch

−0.65 −0.7 −0.75

500

𝜆 = 200

−0.92

100

300 Epoch

−0.9 −0.92 −0.94

500

100

300 Epoch

500

Fig. 3. Training reward (upper row) and validation composite score (lower row) across the five arrival loads under nominal seed 42. Markers occupy the five plotted epoch positions 100–500. Static ACB Success

Rate

0.4

0.2 10

20

50 100 200

Delay 4

0.08 0.06

3

0.04

0.05

0.2

B-D3QN-RACH-8

Blocking

Fallback

0.1

0.6

0.4

QP-BD3QN-RACH

0.15

0.8

0.6

Flat D3QN

Collision 1

0.8

0

Flat DDQN

2

0.02 0

10

20

50 100 200

10

20 50 100 200 New arrivals

10

20

50 100 200

1

10

20

50 100 200

Fig. 4. Cross-load performance of the five controller configurations under nominal seed 42.

Positive values therefore indicate better performance by QPBD3QN-RACH after accounting for the direction of the metric.

TABLE IV M EAN D IRECTION -A LIGNED D IFFERENCES ACROSS F IVE L OADS : R ATES IN P ERCENTAGE P OINTS AND D ELAY IN D ECISION I NTERVALS

Fig. 4 shows the five performance measures across the full load range. Success per attempt and collision per attempt follow the complementarity relation in (21) and consequently exhibit mirrored trends. From loads 10 through 100, QPBD3QN-RACH records the highest success per attempt and the corresponding lowest collision per attempt. It also records the lowest two-step fallback across all five loads and the lowest successful-access delay at loads 10, 20, 100, and 200.

Comparator

The comparison also reveals two distinct operating points. At load 50, B-D3QN-RACH-8 records the lowest successfulaccess delay, while QP-BD3QN-RACH records the highest success and the lowest collision, fallback, and blocking. At load 200, Static ACB records the highest success and the lowest collision and four-step blocking, while QP-BD3QNRACH records the lowest two-step fallback and successfulaccess delay. The cross-load curves therefore show how the relative balance among successful resolution, fallback, blocking, and delay changes as contention increases.

Static ACB Flat DDQN Flat D3QN B-D3QN-RACH-8

Success/Collision

Fallback

Blocking

Delay

5.74 6.18 5.95 8.21

1.37 1.34 1.23 1.92

0.35 0.42 0.43 0.68

0.456 0.329 0.323 0.128

Equations (68) and (69) transform the five load-wise comparisons in Fig. 4 to a common better-positive scale and average them over Λ. Table IV reports 100Δ̄𝑚−𝑏 in percentage 𝑘 points for rate measures and Δ̄𝑚−𝑏 in decision intervals for 𝑘 delay. Success and collision have identical direction-aligned differences because of (21). Table IV shows positive mean differences for all four comparators across the evaluated load grid, with the largest success/collision and fallback differences occurring against B-D3QN-RACH-8. Fig. 5 reports the means with descriptive 95% percentile-bootstrap intervals obtained from 2000 resamples, with replacement, of the five load-wise paired

10

Success

Collision

Mean with descriptive 95% bootstrap interval Blocking Fallback

Delay

Static ACB Flat DDQN Flat D3QN B-D3QN-RACH-8 0

0.05

0.1

0.15

0

0.05

0.1

0.15

0 0.01 0.02 0.03 Direction-aligned difference

0

0.01

0.02

−0.2 0

0.2 0.4 0.6

Fig. 5. Direction-aligned differences from each comparator under nominal seed 42. Markers show five-load means; horizontal bars show descriptive 95% percentile-bootstrap intervals of the means. Positive values favor QP-BD3QN-RACH. Rate differences are unscaled; delay differences are measured in decision intervals. Seeds

Rate

Success

Six-seed mean with 95% bootstrap interval Blocking Fallback

Collision

Delay

0.08

0.8

0.8

0.12

0.4

0.4

0.06

0

0 10

20

50

100

0.04

20

50

100

200

2

0

0 10

200

3

10

20

50

100

200

1 10

20

50

100

200

10

20

50

100

200

New arrivals

Fig. 6. Seed sensitivity of QP-BD3QN-RACH across six seeds and five arrival loads. Gray curves show individual seeds; green diamonds show their mean, with vertical bars denoting 95% percentile-bootstrap intervals of the mean.

differences. Its rate axes show unscaled differences, with 0.01 corresponding to one percentage point. D. Seed Sensitivity and Ablation Analysis Seed sensitivity of QP-BD3QN-RACH is evaluated over Z = {42, 23, 16, 15, 8, 4}. Let 𝑅 𝑘,𝜆,𝑧 denote the tail-window value of QP-BD3QN-RACH for metric 𝑘, load 𝜆, and seed 𝑧. The six-seed mean is 1 ∑︁ 𝑅¯ 𝑘,𝜆 = 𝑅 𝑘,𝜆,𝑧 . (70) |Z|

As shown in Fig. 7, narrowing the branch-action range shifts the operating balance toward selected high-load success, collision, and blocking measures, while the default branch grid gives stronger low- and mid-load success, fallback, and delay behavior. The full-penalty reward, shorter warm-up, faster target update, and wider hidden layer produce more localized changes across the three loads. These load- and metricdependent tradeoffs support retaining the default five-choice branch grid, mild-penalty reward, 2000-interaction warm-up, 500-update target period, and hidden width 128 for the main comparison.

𝑧∈Z

Fig. 6 reports the six seed curves and their mean. At each load, vertical error bars show 95% percentile-bootstrap intervals of the mean, obtained from 2000 resamples with replacement of the six seed-level values. Seed variation is comparatively small at loads 10, 20, 100, and 200, whereas load 50 exhibits the widest spread, particularly in success/collision and delay. The ablation analysis evaluates the “Branch mild” grid {0.75, 0.875, 1, 1.125, 1.25} and the “Branch narrow” grid {0.875, 0.9375, 1, 1.0625, 1.125}, together with the fullpenalty reward, warm-up 1000, target-update 200, and hiddenwidth 256 variants at loads 20, 100, and 200. The fullpenalty profile retains the success coefficients and doubles every penalty coefficient in (55). For variant 𝑣, define   𝑣 − 𝑅 def (71) Δ𝑣−def = 𝜎𝑘 𝑅 𝑘,𝜆 𝑘,𝜆 . 𝑘,𝜆 Values above zero favor the variant under the metric-direction convention, whereas values below zero favor the default configuration.

E. Mechanism Diagnostics and Representation Structure The aggregate performance trends can be related to the underlying access pressure through Fig. 8. The figure compares attempted UEs per preamble, pre-resolution contention pressure, mean post-interval backlog, and nominal message cost per successful UE across controller configurations and arrival loads. Together with Fig. 4, these quantities describe how access pressure, post-interval backlog, and nominal signaling cost vary across arrival loads. Fig. 9 reports 2RA/4RA success and delay, together with DT-2RA and M2M-2RA success. Success panels use current access mode and the corresponding active-user denominators; delay panels group successful UEs by initial access mode. Fig. 10 reports tail-window Spearman associations between aggregate access/composition variables and projected quota shares, reward, and failure pressure. Finally, Table V complements the behavioral diagnostics with the value-representation dimensions derived in (66) and (67). For the five-choice grid, QP-BD3QN-RACH represents

11

𝜆 = 20

Branch mild Branch narrow Full penalty Warm-up Target update Hidden width

𝜆 = 100

𝜆 = 200

−4.13

−4.13

−0.55

−0.06

−0.14

−1.36

−1.36

−1.33

+0.07

−0.36

+0.05

+0.05

−0.14

+0.08

−6.10

−6.10

−0.63

−0.12

−0.20

−2.09

−2.09

−1.93

+0.07

−0.51

+0.14

+0.14

−0.12

+0.13

−0.55

+0.03

+0.03

+0.02

−0.00

+0.00

−2.09

−2.09

−1.20

−0.22

+0.08

−0.18

−0.18

−0.19

−0.01

+0.03

+1.17

+1.17

+0.13

+0.00

+0.04

−1.49

−1.49

−0.87

−0.15

+0.06

−0.03

−0.03

−0.05

+0.01

−0.02

+0.01

+0.01

+0.01

−0.00

+0.00

+0.66

+0.66

+0.35

+0.14

−0.12

−0.03

−0.03

−0.11

+0.03

−0.13

+0.34

+0.34

+0.17

−0.03

+0.02

+0.04

+0.04

+0.02

+0.03

−0.05

−0.01

−0.01

−0.07

+0.01

Suc.

Col.

Fb.

Blk.

Dly.

Suc.

Col.

Fb.

Blk.

Dly.

Suc.

Col.

Fb.

−0.34

−0.02

Blk.

Dly.

Fig. 7. Direction-aligned ablation differences relative to the default QP-BD3QN-RACH configuration at loads 20, 100, and 200. Rate entries are percentagepoint differences; delay entries are differences in decision intervals. Values are rounded to two decimal places. Static ACB

Flat DDQN

Attempts/preamble

Attempt pressure

Flat D3QN

20 10

0.5

0 20

50

100

200

10

20

50

100

B-D3QN-RACH-8

Post-interval backlog

Contention

1

10

QP-BD3QN-RACH

Nominal cost/success

2,000

80

1,000

40

0 10 200 New arrivals

20

50

100

200

0

10

20

50

100

200

Fig. 8. Cross-load access diagnostics: attempted UEs per preamble, pre-resolution contention pressure, mean post-interval backlog, and nominal message cost per successful UE. Flat DDQN 4RA success

2RA success

0.4 Rate

0.2 0

20

100

200

0

2

2 20

100

DT-2RA success

3

3

0.2

0.1

B-D3QN-RACH-8

Initial-4RA delay

4

0.4

0.3

QP-BD3QN-RACH

Initial-2RA delay

200

20

100

200 20 New arrivals

100

200

0.3

0.3

0.2

0.2

0.1

0.1

0

20

100

M2M-2RA success

0.4

200

0

20

100

200

Fig. 9. Selected access-mode and traffic measures at loads 20, 100, and 200 for Flat DDQN, QP-BD3QN-RACH, and B-D3QN-RACH-8. Success rates use current-mode active populations; delays use successful UEs grouped by initial access mode.

DS Share

−0.50

−0.24

−0.73

+0.76

+0.86

−0.88

2RA Share

−0.63

+0.06

−0.78

+0.87

+0.92

−0.93

Contention

+0.61

−0.10

+0.82

−0.82

−0.99

+0.98

+0.62

−0.11

+0.82

−0.83

−0.98

+0.99

+0.63

−0.12

+0.82

−0.83

−0.98

+1.00

+0.62

−0.11

+0.82

−0.83

−0.98

+0.99

DS-2

DT-2

DS-4

DT-4

Reward

Failure

Backlog Attempt Active

Fig. 10. Tail-window Spearman associations between aggregate access/composition variables and projected quota shares, reward, and failure pressure for QP-BD3QN-RACH. Values are rounded to two decimal places.

625 pre-projection branch-action tuples with 20 branch-action outputs. The eight-branch configuration uses 40 outputs for 390,625 tuples, while explicit enumeration of the symmetric four-pool Cartesian reference requires 625 action-dependent outputs. V. C ONCLUSION This paper presented QP-BD3QN-RACH, a quota-projected branching reinforcement-learning controller for finite-budget QoS-aware CBRA preamble slicing. The four action branches

TABLE V ACTION -VALUE O UTPUT C OUNTS FOR THE F IVE -C HOICE B RANCH G RID Representation Symmetric flat four-pool QP-BD3QN-RACH B-D3QN-RACH-8

Explicit outputs

Cartesian tuples

625 20 40

625 625 390,625

follow the four executable QoS/access-mode quota dimensions, while deterministic projection converts the selected multipliers to feasible integer allocations. Across the evaluated load grid, QP-BD3QN-RACH achieves favorable mean direction-aligned differences against all four comparators and exhibits load-dependent performance gains in success, collision, fallback, blocking, and successful-access delay. At load 200, Static ACB achieves higher success and lower collision and four-step blocking, while QP-BD3QN-RACH maintains lower two-step fallback and successful-access delay. The evaluated four-pool configuration combines budget-feasible allocation with a compact action-value representation and loaddependent performance tradeoffs. The current evidence is bounded by the adopted simula-

12

tion and comparison design. Cross-method comparisons use nominal seed 42, and the six-seed analysis concerns QPBD3QN-RACH. The flat and four-branch learning controllers share the same aggregate observation but use different action representations, while B-D3QN-RACH-8 additionally employs a finer user-type-specific observation and contention-group partition. The CBRA model performs procedure-level contention resolution and does not include physical-layer capture, signal-to-interference-plus-noise ratio (SINR), power control, multi-user decoding, contention-free access, or MsgA/Msg3 decoding. The load-paired intervals and Spearman coefficients provide descriptive summaries of the evaluated data. Future work should extend the evaluation to multi-seed cross-method comparisons, broader traffic and arrival processes, external baseline reproductions, and channel-aware access models. R EFERENCES [1] 3GPP, “Technical Specification Group Services and System Aspects; Service accessibility,” 3rd Generation Partnership Project, Technical Specification (TS) 22.011, 2025, version 18.6.0. [Online]. Available: https://www.3gpp.org/dynareport/22011.htm [2] ——, “Technical Specification Group Radio Access Network; NR; NR and NG-RAN Overall Description; Stage 2,” 3rd Generation Partnership Project, Technical Specification (TS) 38.300, 2025, version 18.6.0. [Online]. Available: https://www.3gpp.org/dynareport/38300.htm [3] ——, “Technical Specification Group Radio Access Network; NR; Medium Access Control (MAC) protocol specification,” 3rd Generation Partnership Project, Technical Specification (TS) 38.321, 2025, version 18.6.0. [Online]. Available: https://www.3gpp.org/dynareport/38321.htm [4] A. Turan, M. Koseoglu, and E. A. Sezer, “Reinforcement learning based adaptive access class barring for RAN slicing,” in 2021 IEEE International Conference on Communications Workshops. IEEE, 2021, pp. 1–6. [5] B.-H. Lee, H.-S. Lee, S. Moon, and J.-W. Lee, “Enhanced random access for massive-machine-type communications,” IEEE Internet of Things Journal, vol. 8, no. 8, pp. 7046–7064, 2021. [6] H. Althumali, M. Othman, N. K. Noordin, and Z. M. Hanapi, “Prioritybased load-adaptive preamble separation random access for QoSdifferentiated services in 5G networks,” Journal of Network and Computer Applications, vol. 203, p. 103396, 2022. [7] M. R. Chowdhury and S. De, “Queue-aware access prioritization for massive machine-type communication,” IEEE Internet of Things Journal, vol. 9, no. 17, pp. 15 858–15 873, 2022. [8] A. M. Gedikli, M. Koseoglu, and S. Sen, “Deep reinforcement learning based flexible preamble allocation for RAN slicing in 5G networks,” Computer Networks, vol. 215, p. 109202, 2022. [9] Y. Piao and T. J. Lee, “Integrated 2–4 step random access for heterogeneous and massive IoT devices,” IEEE Transactions on Green Communications and Networking, vol. 8, no. 1, pp. 441–452, 2024. [10] H. Jiang, D. Qu, J. Ding, and T. Jiang, “Multiple preambles for high success rate of grant-free random access with massive MIMO,” IEEE Transactions on Wireless Communications, vol. 18, no. 10, pp. 4779– 4789, 2019. [11] H. S. Jang, H. Lee, T. Q. S. Quek, and H. Shin, “Deep learning-based cellular random access framework,” IEEE Transactions on Wireless Communications, vol. 20, no. 11, pp. 7503–7518, 2021. [12] Y. He and G. Ren, “Cluster-aided collision resolution random access in distributed massive MIMO systems,” IEEE Internet of Things Journal, vol. 9, no. 13, pp. 11 453–11 463, 2022. [13] S. Khairy, P. Balaprakash, L. X. Cai, and H. V. Poor, “Data-driven random access optimization in multi-cell IoT networks using NOMA,” IEEE Transactions on Wireless Communications, vol. 21, no. 7, pp. 4938–4953, 2022. [14] G. Liva and Y. Polyanskiy, “Unsourced multiple access: A coding paradigm for massive random access,” Proceedings of the IEEE, vol. 112, no. 9, pp. 1214–1229, 2024. [15] D. Nie, W. Yu, C. H. Foh, and Q. Ni, “A NOMA-enhanced two-step RACH procedure for low-latency access in 5G networks,” IEEE Internet of Things Journal, vol. 12, no. 9, pp. 11 568–11 580, 2025.

[16] J. C. Marinello Filho, T. Abrão, E. Hossain, and A. Mezghani, “Grantfree random access for RIS-aided machine-type communication,” IEEE Transactions on Wireless Communications, vol. 24, no. 9, pp. 7794– 7808, 2025. [17] H. Zhang, X. Zhao, and L. Dai, “Delay-optimal random access: A learning framework,” IEEE Transactions on Communications, vol. 74, pp. 3059–3073, 2026. [18] L. Tello-Oquendo, D. Pacheco-Paramo, V. Pla, and J. Martinez-Bauset, “Reinforcement learning-based ACB in LTE-A networks for handling massive M2M and H2H communications,” in 2018 IEEE International Conference on Communications. IEEE, 2018, pp. 1–7. [19] D. Pacheco-Paramo, L. Tello-Oquendo, V. Pla, and J. Martinez-Bauset, “Deep reinforcement learning mechanism for dynamic access control in wireless networks handling mMTC,” Ad Hoc Networks, vol. 94, p. 101939, 2019. [20] D. Pacheco-Paramo and L. Tello-Oquendo, “Delay-aware dynamic access control for mMTC in wireless networks using deep reinforcement learning,” Computer Networks, vol. 182, p. 107493, 2020. [21] A.-T. H. Bui and A. T. Pham, “Deep reinforcement learning-based access class barring for energy-efficient mMTC random access in LTE networks,” IEEE Access, vol. 8, pp. 227 657–227 666, 2020. [22] N. Jiang, Y. Deng, A. Nallanathan, and J. Yuan, “A decoupled learning strategy for massive access optimization in cellular IoT networks,” IEEE Journal on Selected Areas in Communications, vol. 39, no. 3, pp. 668– 685, 2021. [23] J. Bai, H. Song, Y. Yi, and L. Liu, “Multiagent reinforcement learning meets random access in massive cellular internet of things,” IEEE Internet of Things Journal, vol. 8, no. 24, pp. 17 417–17 428, 2021. [24] W. Fan, P. Fan, and Y. Long, “Joint delay-energy optimization for multi-priority random access in machine-type communications,” IEEE Transactions on Wireless Communications, vol. 23, no. 2, pp. 1416– 1431, 2024. [25] X. Liu, H. Yang, S. Li, Z. Liu, and X. Lian, “Enhanced RACH optimization in IoT networks: A DQN approach for balancing H2H and M2M communications,” Internet of Things, vol. 28, p. 101433, 2024. [26] A. O. Elmeligy, I. Psaromiligkos, and M. Au, “Preamble selection probability optimization in RACH: A multi-armed bandits approach,” IEEE Open Journal of the Communications Society, vol. 6, pp. 10 761– 10 780, 2025. [27] V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski et al., “Human-level control through deep reinforcement learning,” Nature, vol. 518, no. 7540, pp. 529–533, 2015. [28] H. Van Hasselt, A. Guez, and D. Silver, “Deep reinforcement learning with double Q-learning,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 30, no. 1, 2016, pp. 2094–2100. [Online]. Available: https://ojs.aaai.org/index.php/AAAI/article/view/10295 [29] Z. Wang, T. Schaul, M. Hessel, H. Van Hasselt, M. Lanctot, and N. De Freitas, “Dueling network architectures for deep reinforcement learning,” in Proceedings of the 33rd International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, vol. 48. PMLR, 2016, pp. 1995–2003. [Online]. Available: https://proceedings.mlr.press/v48/wangf16.html [30] S. Wang, H. Liu, P. H. Gomes, and B. Krishnamachari, “Deep reinforcement learning for dynamic multichannel access in wireless networks,” IEEE Transactions on Cognitive Communications and Networking, vol. 4, no. 2, pp. 257–265, 2018. [31] A. Tavakoli, F. Pardo, and P. Kormushev, “Action branching architectures for deep reinforcement learning,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 32, no. 1, 2018, pp. 4131–4138. [32] E. Nisioti and N. Thomos, “Fast q-learning for improved finite length performance of irregular repetition slotted aloha,” IEEE Transactions on Cognitive Communications and Networking, vol. 6, no. 2, pp. 844–857, 2020. [33] J. Du, C. Jiang, J. Wang, Y. Ren, and M. Debbah, “Machine learning for 6G wireless networks: Carrying forward enhanced bandwidth, massive access, and ultrareliable/low-latency service,” IEEE Vehicular Technology Magazine, vol. 15, no. 4, pp. 122–134, 2020. [34] M. Sohaib, J. Jeong, and S.-W. Jeon, “Dynamic multichannel access via multi-agent reinforcement learning: Throughput and fairness guarantees,” IEEE Transactions on Wireless Communications, vol. 21, no. 6, pp. 3994–4008, 2022.

Record · ID 667939 · SHA-256 9fbe07fd900e75ae
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.