ConceptioArchivearXiv CS
arXiv CSopen access

ARIADNE: AI-RAN Informed Link Adaptation in Digital Twin Network Environments

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
distributedsystemsprotocols
networking, internet, protocols, distributed systems

This paper has been accepted for publication at European Wireless 2026. This is the authors’ accepted version of the article. The final version published by IEEE is: M. Tsampazi, N. N. Santhi, N. Perrotta, F. Dressler, T. Melodia, “ARIADNE: AI-RAN Informed Link Adaptation in Digital Twin Network Environments,” in Proc. of European Wireless, Rimini, Italy, June 2026.

ARIADNE: AI-RAN Informed Link Adaptation in Digital Twin Network Environments Maria Tsampazi∗ , Neagin Neasamoni Santhi∗ , Nicole Perrotta∗ , Falko Dressler§ , Tommaso Melodia∗ ∗ Institute for Intelligent Networked Systems, Northeastern University, Boston, MA, U.S.A.

arXiv:2605.29772v1 [cs.NI] 28 May 2026

E-mail: {tsampazi.m, neasamonisanthi.n, perrotta.nic, t.melodia}@northeastern.edu § School of Electrical Engineering and Computer Science, TU Berlin, Germany E-mail: {dressler}@ccs-labs.org Abstract—Artificial Intelligence (AI)-powered Radio Access Network (RAN) networks have attracted significant attention from both industry and academia. Meanwhile, Digital Twins offer a safe playground for experimenting with AI/Machine Learning (ML)-based solutions for advanced AI-RAN research. By enabling the testing of online algorithms before deployment on the RAN, they reduce costs and safety risks associated with physical field testing. In this article, we propose ARIADNE, an online Reinforcement Learning (RL)-based module that seamlessly integrates with SIONNA and is tasked with performing link adaptation. We explore different design choices and demonstrate how ARIADNE can surpass industry-standard and state-of-the-art methods by achieving up to 11% and 20% improvements in Spectral Efficiency, respectively. Finally, we show that RL learns a Modulation and Coding Scheme (MCS) selection strategy that diverges from Outer Loop Link Adaptation (OLLA), exhibiting either more conservative or more aggressive behavior depending on the configuration, a trend further corroborated by training offline on 5th generation (5G) over-the-air (OTA) measurements. Index Terms—5G/6G, Adaptive Modulation and Coding, RL, SIONNA

I. I NTRODUCTION Recent years have seen collaboration among leading industry entities such as Nokia and NVIDIA to advance AI-RAN research [1]. AI-native 6G is expected to emerge from simulation, with Digital Twins playing a key role in the train–simulate– deploy–optimize lifecycle [2]. Indeed, Digital Twins are envisioned to enable faster innovation by allowing the evaluation of “what-if” scenarios for future technologies that are not yet available in hardware [3]. By reducing the “concept-to-live” cycle, dense urban environments and complex 6G use cases can be simulated with greater accuracy, supporting the design and delivery of NextG systems while accelerating the adoption and deployment of advanced AI solutions [2]. In this context, NVIDIA’s SIONNA Digital Twin environment [4] provides an all-in-one platform for wireless research, enabling the execution of system-level simulations over raytraced channels. In particular, SIONNA-SYS allows the onthe-fly integration of plugin AI/ML solutions, facilitating the testing of key features and capabilities essential to AI-native 6G research. Simulated, controlled environments often provide a practical first step for developing AI/ML solutions before conducting OTA experiments, allowing for algorithm refinement and parameter optimization. At the same time, a trending topic that has gained widespread attention in the research community is the optimization of the This article is based upon work supported by the U.S. National Science Foundation under Grant CNS-2112471.

Next Generation Node Base (gNB)’s Medium Access Control (MAC)-layer scheduler. This involves the optimal allocation of Physical Resource Blocks (PRBs) and selection of the MCS index, as well as the tuning of power control mechanisms (e.g., 3GPP-compliant open and closed-loop power control) that adjust transmit power based on predefined Signal-to-Noise-Ratio (SNR) targets to meet Quality of Service (QoS) requirements for different slices and verticals. Importantly, by jointly optimizing power control and MCS selection, Spectral Efficiency can be significantly improved, demonstrating the benefits of a coordinated approach for enhanced network management. A. Related Work Numerous works have focused on RL for MCS selection [5], ranging from offline approaches [6] to contextual Multi-Armed Bandit (MAB) solutions [7], [8]. The authors in [9] introduce GrGym, an RL-based framework for MCS selection in WiFi, which leverages a custom gym environment. In [10], the authors formulate an RL framework where a low discount factor enforces myopic MCS selection based on the current channel state. The industry standard for link adaptation is Outer Loop Link Adaptation (OLLA) [11], and modifications involving RL have been introduced for intelligent on-the-fly adaptation [12]. More recently, the authors in [13] propose a gradient descent approach with a learning rate that self-adapts online through knowledge distillation. However, none of the works mentioned above focuses on the integration of an AI module on the fly within high-fidelity Digital Twins. B. Contributions Motivated by the suitability of system-level simulations for digital twin environments, we propose ARIADNE, a framework that leverages NVIDIA’s platform SIONNA to integrate an RL-module for AI-RAN networks, enabling learning-based link adaptation through adaptive MCS selection on channels simulated via SIONNA ray tracing. Our module seamlessly integrates with the SIONNA-SYS platform, enabling direct comparison with state-of-the-art link adaptation algorithms (such as the industry-standard OLLA [11] and NVIDIA’s proposed link adaptation scheme, entitled Self-Adaptive Link Adaptation (SALAD) [13]). In this article, we train and evaluate ARIADNE on a variety of high-fidelity ray-traced channels generated with SIONNA, and we demonstrate that online RL outperforms both OLLA and SALAD in MCS selection in terms of performance, as measured by the achieved Spectral Efficiency. Finally, we evaluate the suitability of RL on offline

5G data collected from an OTA 5G testbed, demonstrating its effectiveness on real-world measurements. II. ARIADNE: L EARNING L INK A DAPTATION VIA R EINFORCEMENT L EARNING Our RL agent operates at the Fast Link Adaptation level, where decisions (i.e., the selection of MCS levels for all User Equipments (UEs)) are enforced on a per-slot basis. This requires precise knowledge of the channel quality. However, Channel Quality Information (CQI) reports are subject to feedback delay [14]. In addition, CQI is typically reported over a wide band, while transmission may occur over a narrower band. As a result, the true Signal to Interference plus Noise Ratio (SINR) is not directly observable and must be estimated online at each slot using only past feedback. Consequently, link adaptation requires jointly inferring the channel quality and selecting an appropriate MCS [13]. Therefore, for the current slot’s channel quality, we consider two estimation modes: an Oracle mode that establishes an upper bound using SIONNA’s simulated environment to calculate the SINR, and a predictor mode for SINR estimation that leverages previous CQI reports. Markov Decision Process (MDP) Formulation and Temporal Resolution. We formulate the MCS selection problem as an MDP defined by the tuple (S, A, R, P, β). Critical to our design is a 1 : 1 mapping between the RL environment and the 5G New Radio (NR) Physical (PHY)-layer, where each discrete environment step corresponds exactly to one 5G NR time slot.

(u)

∆offset,t : A normalized asymmetric feedback accumulator updated by +0.1 on ACK and −0.9 on NACK.1 (u) 1 • āt : The running ACK rate over the full episode. B. Action Space •

The agent performs link adaptation for the 5G NR Physical Downlink Shared Channel (PDSCH). The action at ∈ {0, 1, . . . , 28}U is a vector of MCS indices, one per UE, as defined in the 5G NR MCS Table 1 [15]. The mapping is performed directly by the policy network as follows in (3):   (1) (U ) at = π(st ) = mt , . . . , mt . (3) C. Reward We aim to capture the fundamental MCS selection tradeoff. Higher MCS indices result in higher Spectral Efficiency, but simultaneously increase the BLER and, consequently, the probability of decoding failure. Conversely, a conservative MCS lowers the BLER but also reduces the Spectral Efficiency. To target the practical throughput–reliability tradeoff, we maximize a reward that captures the effective Spectral Efficiency, which is directly proportional to system goodput. Instantaneous Reward. At each environment step (i.e., each 5G NR time slot), the agent selects an MCS index for each UE, and the environment subsequently returns a scalar reward based on the successfully delivered Spectral Efficiency in that slot. The instantaneous reward rt at slot t is defined as the total achieved Spectral Efficiency across all UEs, as expressed in (4):

A. State Space At each time slot t, the agent observes a state vector st ∈ R , where U is the number of UEs. The state is composed of per-UE feature vectors: h i (1) (U ) st = st , . . . , st , (1) (Ncqi +Nharq +5)·U

rt =

U X

n o (u) (u) (u) Q(mt ) · Rc (mt ) · ⊮ Ht = ACK ,

(4)

u=1 (u)

(u)

where Q(mt ) and Rc (mt ) denote the modulation order and the effective code rate, respectively, associated with the MCS (u) index mt selected for UE u at slot t. The term ⊮{·} denotes (u) where each per-UE component st is given by the indicator function, which evaluates to 1 if the HARQ h i (u) (u) (u) (u) (u) (u) (u) (u) (u) st = γ̃post,t , γ̃eff,t−k , ht−k , m̃t−1 , B̂t , ∆offset,t , āt , outcome Ht for the u-th UE is a successful acknowledgment (ACK), and 0 otherwise. (2) Expected Reward. In the SIONNA-SYS simulator, the and the respective components are defined as follows: HARQ feedback is generated by the PHY Abstraction module (u) (u) • γ̃post,t : An estimate of the current slot’s normalized postas a Bernoulli outcome with success probability Pr(ACKt = equalization SINR, computed as the mean SINR across 1 | s , a ) = 1 − BLER(m(u) , γ (u) ), where the BLER t t t t allocated resource elements after Linear Minimum Mean (u) depends on the selected MCS index mt and the effective Square Error (LMMSE) equalization. In Oracle mode, (u) SINR γt computed from the post-equalization SINR at slot this is derived from the current channel realization; in t. Given that the nominal Spectral Efficiency is defined as predictor mode, it is estimated from past measurements. (u) (u) (u) SEnom,t ≜ Q(mt )·Rc (mt ) for a chosen MCS, the expected (u) • γ̃eff,t−k , k = 1, . . . , Ncqi : The Ncqi most recent effective per-slot reward is given in (5): normalized SINR values reported by the PHY-layer AbU X straction. (u) (u)  (u) (u) (u) E[r | s , a ] = Q(mt ) · Rc (mt ) · 1−BLER(mt , γt ) . t t t • ht−k , k = 1, . . . , Nharq : A window of Nharq prior Hy{z } | u=1 brid Automatic Repeat reQuest (HARQ) outcomes, where SEnom (mt ) (5) h ∈ {−1, 0, 1} denotes unscheduled, NACK, and ACK, Reward Interpretation. The expectation in (5) evaluates the respectively. (u) effective Spectral Efficiency, which is directly proportional to • m̃t−1 : The normalized previously selected MCS index. (u) (u) • B̂t : The sliding-window Block Error Rate (BLER) es1 Both ∆(u) are derived from HARQ feedback; the former offset,t and āt timate over the last N scheduled slots, calculated as bler weights NACK outcomes 9× more heavily than ACK, making it more sensitive P P (u) (u) to link failures than the symmetric ACK rate āt . B̂t = NACK / Scheduled.

Ray Tracing Engine 1. Multipath propagation 2. Reflection, diffraction etc. 3. ITU-R materials

Bridging RT with Signal Processing 1. RT paths converted to CFR on OFDM grid 2. MIMO-OFDM channel matrix

compute_sinr 𝐇 predict_sinr() List of SINR predictors:

end ΜΑC-Layer Link Adaptation SIONNA-SYS

ARIADNE Episode Start: H =pool[randint()] (k)

for i=1:N_slots 𝜋) 𝑠" (MlpPolicy)

end

𝑟"

env.step(m( ) MCS Enforcement

𝑠"#$

OLLA

SALAD

% & 𝑚#$(

𝑠"

PPO Update

𝑚"

_get_obs()

%

Decision Tree Random Forest Kalman Filter OCO

' harq $ {ℎ"#! }!%&

end Bernoulli BLER Calc.

i)PFSchedulerSUMIMO ii)downlink_fair_power _control iii)RZFPrecodedChannel iv)LMMSEPostEqualizati onSINR &

)

SINR history buffer

𝛾post, #$% %'(

MCS Decode

&

𝛾eff, #

EESM

Fig. 1. End-to-end execution of MCS selection in SIONNA. The channel matrix is generated from a pool of ray-traced channel realizations to ensure diversity. The current SINR is either directly used or estimated from previous observations. The post-equalization SINR is then provided as input to the PHY Abstraction. The state st is observed by the ARIADNE framework, which selects the corresponding MCS index.

20

40

gNB

UE1

gNB

UE 2

UE1

UE 2

gNB

gNB

UE3

a) Simple Street Canyon

2 UE

UE 3

c) Étoile, Paris, France

b) Munich, Germany

Digital Replica of Real-World

Simple Street Canyon

Munich, Germany

gNB gNB

Étoile, Paris, France

# decoded bits

PHY Abstraction SIONNA-SYS T.B. size Calc.

0

1 UE

• • • •

dB 20

-40

UE 3

iii)paths.cfr()

SIONNA-RT

Physics-based 3D Geometry 1. ‘simple_street_canyon’ 2. ‘etoile’ 3. ‘munich’

gNB MAC Layer Processing SIONNA-SYS

𝑯 * 𝑇 for i=1:N_slots

ii)PathSolver() !

SINR Coverage Map

Channel Generation SIONNA-RT 𝑎 , 𝜏!

i)load_scene()

SIONNA-SYS-LEVEL

for i=1:N_channel_realizations

Topology (3D Scene)

PF Scheduler

Power Control

User Mobility

Sionna RT Channel

RZF Precoding + LMMSE Equalization

post-eq SINR

PHY Abstraction Effective SINR MCS

BLER Tables HARQ

Link Adaptation Channel Quality

OLLA Δ-SINR offset

SALAD

GD SINR est.

ARIADNE π(s) → MCS

slot++

the system goodput. Our RL agent is tasked with maximizing the immediate per-slot reward, which over an episode is Fig. 2. System-level simulations using SIONNA over ray-traced channels for equivalent to maximizing the time-average effective through- link adaptation with OLLA, ARIADNE, and SALAD. put. Therefore, our reward formulation naturally captures the A. System Architecture and Link Adaptation Workflow fundamental link adaptation tradeoff, with the RL agent learnThe end-to-end pipeline is shown in Fig. 1. At the beginning ing a policy that selects the MCS to maximize the product of each simulation, a 3D scene (e.g., Urban Street Canyon, SEnom · (1 − BLER), thereby learning the rate–reliability Munich, or Étoile in Fig. 2) is loaded into SIONNA’s ray tracer tradeoff through trial and error. to compute the channel frequency response Ht per slot. At each Adaptive BLER Penalty. Although the reward in (4) im- slot, the simulator executes proportional-fair scheduling, power plicitly penalizes BLER through zero Spectral Efficiency on control, and Regularized Zero-Forcing (RZF) precoding with NACK, this signal alone may not be sufficient to prevent persis- LMMSE equalization, yielding post-equalization SINR values. tently aggressive MCS selection. Link adaptation mechanisms Link adaptation then selects an MCS index based on the SINR such as OLLA target a fixed BLER level (typically around and HARQ feedback. This is the only component that differs 10% [11], [13], [15]), balancing throughput and reliability. across methods (Fig. 2). OLLA maintains a per-UE SINR offset To incorporate a similar notion of reliability control, we ad- updated via ACK/NACK feedback and selects the highest MCS ditionally augment the reward with a per-UE penalty weight under a target BLER, using the latest SINR report and HARQ (u) λt , governed by an integral controller that increases when outcome. SALAD primarily relies on ACK/NACK feedback, the cumulative BLER exceeds a target τ and remains zero estimating the SINR online via recursive updates and adapting otherwise, as defined in (6)–(7): i the BLER target through a feedback control loop. ARIADNE X h (u) (u) (u) (u) rt = SEnom,t ⊮{Ht = ACK}−λt ⊮{Ht = NACK} , observes a window of past CQI and HARQ feedback, together with a predicted SINR, and directly maps this state to MCS u (6) indices via a trained policy network. Finally, PHY Abstracwhere tion maps the selected MCS and post-equalization SINR to a h X i (u) (u) λt = max 0, kE ⊮{Hi = NACK} − τ . (7) BLER, from which the HARQ outcome is sampled, yielding the achieved Spectral Efficiency. All steps except link adaptation i≤t (u) are identical across methods, ensuring that any performance Hi ̸=∅ difference is solely due to the MCS selection strategy. (u) When λt = 0 for all u ∈ {1, . . . , U }, (6) reduces to (4). B. Implementation and Training We train the agent using Proximal Policy Optimization D. Transition Dynamics (PPO) [16] as implemented in Stable-Baselines3 [17] within The environment transitions P (st+1 |st , at ) are determined a custom Gymnasium environment [18] built on top of the by the SIONNA-SYS simulator as detailed in Section III. SIONNA system-level simulator, with ARIADNE handling agent–environment interaction. The policy is parameterized by III. S YSTEM M ODEL with 64 hidden units We compare three link adaptation strategies, namely OLLA, a two-layer Multilayer Perceptron (MLP) −4 per layer, a learning rate of 3 × 10 , a clip range of ε = 0.2, SALAD, and ARIADNE. All operate within the same and an entropy coefficient of 0.05 to encourage exploration, SIONNA-SYS pipeline, and the decision-making component resulting in low computational complexity suitable for gNB resides within the gNB’s MAC-layer scheduler. Independent of deployment. We set the discount factor β = 0, treating the the specific algorithm employed, this reflects a centralized arproblem as a contextual bandit. This reflects the observation chitecture in which the gNB collects CQI and HARQ feedback from all UEs, executes the link adaptation algorithm internally, that MCS selection is primarily driven by instantaneous channel and assigns MCS indices to scheduled UEs for the upcoming state, and temporal credit assignment for future slots does not improve performance. In time-varying fading channels, channel transmission. The per-slot pipeline is illustrated in Fig. 2.

Parameter

Setup A

Setup B

3 10

1 1

CQI window (Ncqi ) HARQ window (Nharq )

le

ac Or

DT

KF

I RF CQ CO LLA LAD d O O SA

0 10 0 0. 12

49

41

0.

01

0.

0. 1

10 6

12 0. 1

03

0. 1

0.10

Target

0. 0

3

0.15

0. 0

2. 8

3.

24 5

0.20 94 Mean BLER

88 3. 64 5 3. 63 0 3. 61 3 3. 60 9 3. 60 7

4

3. 6

Mean SE [bps/Hz]

A. Impact of Perfect and Imperfect Channel State Information We begin our experimental evaluation by quantifying the impact of various SINR estimators on system performance, where the predicted current slot’s normalized post-equalization (u) SINR, denoted as γ̃post,t , serves as input to the RL model, as described in Section II-A. Performance is evaluated in terms of the achieved mean Spectral Efficiency (SE) and mean BLER. In detail, when using the Oracle as input to ARIADNE’s RL state, we rely on the ground-truth SINR at the current slot, as reported by SIONNA-SYS. When using one of the predictors, namely Decision Tree (DT) [19], Kalman Filter (KF) [20], Random Forest (RF) [21], and Online Convex Optimization (OCO) [14], we replace the Oracle with the predicted SINR in order to assess how closely the resulting performance matches that of the Oracle, which serves as an upper bound. We then compare all the aforementioned RL-based methods against state-of-theart baselines, namely OLLA and SALAD. Finally, we also discard the current slot’s predicted SINR as input to the RL model (denoted as delayed CQI (dCQI) in the plots) and instead rely solely on past CQI observations, evaluating the resulting performance across all methods.

0.05 0

le

ac Or

DT

KF

I RF CQ CO LLA LAD d O O SA

Fig. 3. Mean Spectral Efficiency and BLER for all RL methods and baselines under Setup A from Table I.

Regarding the inputs used by each SINR predictor2 , KF, DT, and RF leverage past SINR observations, KF and OCO incor2 We do not aim to directly compare the predictors. Instead, we assess whether the performance of the RL agent is affected by the predictor choice, focusing on the robustness of the approach. The KF is a constant-velocity filter with HARQ-driven bias correction, adaptive process noise, and innovation gating. The DT and RF are trained online on past SINR observations, and retrained every Nslots . The OCO follows the ACK/NACK-only feedback setting of [14], using mirror descent with Fixed-Share expert mixing over 12 experts defined on the grid η ∈ {0.5, 1, 2, 3}, β ∈ {0, 0.15, 0.3}, with α = 0.

TABLE II P ERFORMANCE COMPARISON WITH RESPECT TO THE O RACLE . Method

SE [bps/Hz]

∆SE vs. Oracle [%]

BLER

Oracle RL DT RL KF RL RF RL dCQI RL OCO RL OLLA SALAD

3.688 3.645 3.630 3.613 3.609 3.607 3.245 2.894

0.0 -1.17 -1.57 -2.03 -2.14 -2.20 -12.01 -21.53

0.112 0.103 0.101 0.106 0.100 0.120 0.041 0.049

In Fig. 3, we observe that ARIADNE, regardless of the predictor used, consistently outperforms OLLA and SALAD. Even without prediction of the current slot’s SINR, denoted as dCQI in the plots, the RL agent relying only on past CQI reports outperforms OLLA by 10% in terms of spectral efficiency, achieving a mean spectral efficiency of 3.609 bps/Hz compared to 3.245 bps/Hz, at the cost of a higher mean BLER of 0.1, whereas OLLA achieves a BLER of 0.041. Finally, OLLA outperforms SALAD by approximately 11%, while achieving similar BLER values. For the configuration of SALAD, we use the default parameters as specified in [13].3 It is worth noting that SALAD may require further parameter tuning for the specific scenario at hand, and therefore the parameterization used here may not be optimal for the current setup. This further highlights the advantage of RL in adapting to changing channel conditions and dynamically adjusting its policy. UE 1

Relative Frequency

TABLE I O BSERVATION S IZE CONFIGURATIONS FOR ARIADNE.

porate the most recent HARQ feedback, while DT and RF rely (u) on the HARQ accumulator (i.e., ∆offset,t ) to capture feedback trends. In contrast, in our implementation, OCO does not rely on explicit SINR history and instead operates using only HARQ outcomes and the previously selected MCS index [14].

0.85

Relative Frequency

temporal correlation decreases rapidly, limiting the reliability of long-horizon credit assignment. A myopic policy (i.e., β = 0) is therefore well suited to this setting. In our setup, we assume three mobile UEs while the ray-traced channels correspond to SIONNA’s Simple Street Canyon scenario. IV. E XPERIMENTAL E VALUATION ARIADNE observes a configurable window of Ncqi past effective SINR reports and Nharq prior HARQ outcomes, as highlighted in Section II-A. We evaluate two configurations, as shown in Table I. In Setup A, the observation window leverages Ncqi = 3 reports and Nharq = 10 outcomes, providing the agent with temporal context. In Setup B, the observation is limited to a single past SINR report and a single HARQ outcome. Finally, all experimental results reported below have been collected over multiple experiments and channel realizations.

0.85

0

Oracle DT KF RF OCO dCQI

5

10 15 20 25 MCS Index

UE 2

0.2 0

5

10 15 20 25 MCS Index

1

5 10 15 20 25 MCS Index

Oracle DT KF RF OCO dCQI

0.5 0

5 10 15 20 25 MCS Index UE 3

OLLA SALAD

0.5 0

1

UE 2

UE 1 OLLA SALAD

UE 3 Oracle DT KF RF OCO dCQI

0.4

0

1

OLLA SALAD

0.5

5

10 15 20 25 MCS Index

0

5

10 15 20 25 MCS Index

Fig. 4. Distribution of Selected MCS Indices across Methods and UEs under Setup A from Table I.

In Table II, we provide a summary of how closely all methods perform relative to the Oracle when used as input to ARIADNE. We observe that DT achieves a Spectral Efficiency closest to the Oracle, with only ∼ 1% lower performance, reporting a mean Spectral Efficiency of 3.645 bps/Hz compared to 3.688 bps/Hz for the Oracle, followed by KF, whose Spectral Efficiency is reported at 3.630 bps/Hz. All RL methods achieve a Spectral Efficiency close to the Oracle, within ∼ 2%. Finally, OLLA reports a mean Spectral Efficiency that is ∼ 12% lower compared 3 SALAD hyperparameters follow the reference implementation [22]: learning rate ε = 1.2, bias score threshold ρ = 0.25, score window T = 15, probing probability pprobe = 0.15, probing BLER target τ probe = 0.95, and integral gain kE = 0.05.

ke = 0.5

ke = 0.025 1

ke = 0.1

ke = 0

0.75

0.75

0.75

0.5

CDF

0.5

CDF

OLLA 1

CDF

1

0.25

0.25

0.25

0 2.4 2.8 3.2 3.6 SE [bps/Hz]

4

0

0.04

0.1 BLER

0.16

SALAD

Target

0.5

0

16

20 MCS

Fig. 5. CDFs of Spectral Efficiency, BLER, and MCS under Setup A from Table I, leveraging the DT predictor. TABLE III I MPACT OF kE ON M EDIAN S PECTRAL E FFICIENCY, BLER, AND MCS kE SE [bps/Hz] BLER MCS 0 0.025 0.1 0.5

3.547 3.520 3.421 3.342

0.096 0.072 0.047 0.037

21 20 19 18

OLLA SALAD

3.131 2.882

0.041 0.049

17 16

Mean SE [bps/Hz]

ke = 0

ke = 0.025

0

1

ke = 0.1

2

3

OLLA

SALAD

Target

4

5

6

7

8

5

6

7

8

Channel Index

0.15

Mean BLER

ke = 0.5

4 3.6 3.2 2.8 2.4

0.1

0.05 0

0

1

2

3

4

Channel Index

3.8 3.6 3.4

3.645

3.604

Mean BLER

Fig. 6. Spectral Efficiency and BLER under Setup A from Table I, using the DT predictor across multiple channel realizations and experiments. Mean SE [bps/Hz]

to the Oracle, while SALAD underperforms all methods by ∼ 20%. However, all RL methods report a mean BLER of 0.1, indicating the tradeoff between high throughput and reliability. In contrast, OLLA and SALAD maintain the BLER well below the 10% target, at the cost of lower Spectral Efficiency. Fig. 4 shows the relative frequency distribution of selected MCS indices across different SINR predictor methods and UEs. Notably, the mean SINR observed per UE is ∼ 26 dB for UE 1, ∼ 6 dB for UE 2, and ∼ 28 dB for UE 3. We observe that, for UE 1 and UE 3, all RL methods select the highest MCS indices, indicating agreement across methods in high-SINR regimes. In contrast, for UE 2, which experiences the lowest SINR, the RL agent adopts a more aggressive strategy, selecting MCS indices between 5 and 15, with most selections concentrated around 8–12. On the other hand, OLLA and SALAD predominantly select low MCS indices, concentrated below 10, indicating a more conservative approach. B. Impact of the Adaptive BLER Penalty We evaluate the impact of different kE values, as defined in (6)–(7). Since the previous results do not explicitly constrain the BLER and yield values above 10% for all predictor modes and the Oracle, in contrast to SALAD and OLLA, which maintain BLER below 0.1, we vary kE ∈ {0, 0.025, 0.1, 0.5}. Here, kE = 0 corresponds to no penalty, kE = 0.025 provides a moderate penalty, and kE = 0.1 and kE = 0.5 impose increasingly strong penalties on the median BLER. The resulting performance is summarized in Table III and further illustrated in Fig. 5. As expected, increasing kE leads to more conservative MCS selection, reducing Spectral Efficiency while improving reliability. In particular, kE = 0 achieves the highest Spectral Efficiency but at the cost of a higher BLER, while larger values of kE progressively reduce the median BLER at the expense of Spectral Efficiency. Among all configurations, kE = 0.1 provides the best tradeoff, achieving high Spectral Efficiency while consistently maintaining the BLER below the 10% target (Fig. 5). Finally, OLLA and SALAD remain the most conservative.

Target

0.12

0.103

0.111

0.1

B B A A tup tup tup tup Se Se Se Se Fig. 7. Mean Spectral Efficiency and BLER for all RL methods and baselines under Setup A and B from Table I, leveraging the DT predictor.

Finally, in Fig. 6, we observe the mean Spectral Efficiency across each channel realization, averaged over multiple experiments. We observe that RL with kE = 0.1 consistently meets the 10% BLER target while achieving one of the highest Spectral Efficiency values among all considered methods, outperforming the industry-standard OLLA. In contrast, RL with kE = 0 and kE = 0.025 achieves higher Spectral Efficiency but often reaches or exceeds the BLER target. Based on these results, we identify kE = 0.1 as the recommended operating point, as it consistently meets the 10% BLER target while achieving high Spectral Efficiency outperforming both OLLA and SALAD. C. Impact of Observation Window Size We now proceed by evaluating how different observation window sizes impact the mean performance of the RL agent. Once again, for the current slot’s SINR prediction, we leverage the DT, as it achieved the closest performance to the Oracle in terms of Spectral Efficiency in the previous evaluation. The parameterization for this step is given in Table I. As shown in Fig. 7, a smaller observation window (Setup B) results in a slightly lower mean Spectral Efficiency of 3.604 bps/Hz compared to 3.645 bps/Hz achieved by Setup A. Although both Spectral Efficiency values are comparable, Setup A achieves a lower BLER (0.103) than Setup B (0.111). Both values remain above the 10% BLER target, which is expected since the adaptive penalty reward is not leveraged in these experiments. Overall, Setup A is preferred, as it achieves slightly higher Spectral Efficiency while maintaining a lower BLER, enabling the RL agent to select actions without significantly exceeding the BLER target. V. S ITE -S PECIFIC T RAINING AND ROBUSTNESS ACROSS C HANNEL S CENARIOS Fig. 8 shows the training reward evolution across three different channel scenarios in SIONNA. The RL agent consistently improves its performance and converges in all considered environments. This indicates that the learned policy adapts to varying channel characteristics.

Reward

1,000 800 Simple Street Canyon

600

0

100

Munich

Étoile

200 300 Episode

400

500

Fig. 8. Training reward over episodes.

1,600 1,200 800 0

10 20 30 FQI Iteration

22

21

19 20

18

16 17

15

13 14

12

10 20 30 FQI Iteration

OAI Scheduler FQI Policy 11

0

60 40 20 0

MCS Index 60 40 20 0

OAI Scheduler FQI Policy 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27

100

Percentage [%]

150

Percentage [%]

Avg. Q-value

Avg. Q-value

VI. F UTURE I NTEGRATION OF RL ON OTA 5G T ESTBEDS To examine whether reward-driven MCS selection extends beyond simulation, we train an Fitted Q-Iteration (FQI) agent [23] on OTA measurements collected from a 5G testbed [24], using both Downlink (DL) and Uplink (UL) data. The state consists of CQI, Reference Signal Received Power (RSRP), and instantaneous BLER; the reward is a proxy for achieved throughput derived from logged measurements; and the Q-function is an RF with γ = 0.5. As shown in Fig. 9, Q-function targets stabilize across Bellman iterations, and the learned policy shifts MCS selections away from the scheduler’s preferred indices. Since these datasets reflect only the OpenAirInterface (OAI) OLLA scheduler’s choices, closedloop evaluation is left for future work. The fact that both this offline agent and the online PPO agent independently deviate from OLLA suggests that reward-driven MCS selection tends to identify operating points that rule-based schedulers do not.

MCS Index

Fig. 9. Offline FQI results on OTA data for UL (top) and DL (bottom). Left: Q-value convergence across Bellman iterations. Right: MCS distribution of the learned policy vs. the OAI scheduler. VII. C ONCLUSIONS AND F UTURE W ORK

We developed ARIADNE, a framework that seamlessly integrates with SIONNA, leveraging online RL for MCS selection. The use of system-level simulators, which enable the on-thefly integration of AI modules, has facilitated the exploration of various approaches, spanning reward formulation and observation window design, to study the impact of different design choices on link adaptation. Although the SIONNA integration provides a practical first step toward the design of RL-based link adaptation, future work will focus on the deployment of ARIADNE on OTA 5G testbeds. Preliminary results on both DL and UL 5G data demonstrate the potential of learning-based approaches in real-world scenarios. R EFERENCES [1] NVIDIA Corporation, “NVIDIA and Nokia to Pioneer the AI Platform for 6G — Powering America’s Return to Telecommunications Leadership,” NVIDIA Newsroom, Press Release, Oct. 2025, [Available Online]: https: //nvidianews.nvidia.com/news/nvidia-nokia-ai-telecommunications. [2] Nokia, “Nokia launches Nokia RAN Digital Twin to turbo-charge AI-native 6G, powered by NVIDIA Aerial Omniverse Digital Twin,” Nokia Corporate Blog, Feb. 2026, [Available Online]: https: //www.nokia.com/blog/nokia-launches-nokia-ran-digital-twin-to-turbocharge-ai-native-6g-powered-by-nvidia-aerial-omniverse-digital-twin/.

[3] M. Polese, L. Bonati, S. D’Oro, P. Johari, D. Villa, S. Velumani, R. Gangula, M. Tsampazi, C. P. Robinson, G. Gemmi et al., “Colosseum: The open RAN digital twin,” IEEE Open Journal of the Communications Society, 2024. [4] J. Hoydis, S. Cammerer, F. Ait Aoudia, M. Nimier-David, A. Keller, A. Vem, M. Stark, and T. O’Shea, “Sionna: An Open-Source Library for Next-Generation Physical Layer Research,” arXiv preprint arXiv:2203.11854, 2022. [5] N. Khedhri and M. Najar, “Adaptive modulation selection in wireless communications: A comparative study of reinforcement learning, deep learning, deep reinforcement learning, and traditional policies,” in International Wireless Communications and Mobile Computing (IWCMC). IEEE, 2025, pp. 1622–1625. [6] S. Peri, A. Russo, G. Fodor, and P. Soldati, “Offline reinforcement learning and sequence modeling for downlink link adaptation,” in International Conference on Machine Learning for Communication and Networking (ICMLCN). IEEE, 2025, pp. 1–7. [7] S. K. Pulliyakode and S. Kalyani, “Reinforcement learning techniques for outer loop link adaptation in 4G/5G systems,” arXiv preprint arXiv:1708.00994, 2017. [8] V. Saxena, J. Jaldén, J. E. Gonzalez, M. Bengtsson, H. Tullberg, and I. Stoica, “Contextual multi-armed bandits for link adaptation in cellular networks,” in Workshop on Network Meets AI & ML, 2019, pp. 44–49. [9] A. Zubow, S. Rösler, P. Gawłowicz, and F. Dressler, “GrGym: When GNU radio goes to (AI) gym,” in Proc. of the 22nd International Workshop on Mobile Computing Systems and Applications, 2021, pp. 8–14. [10] J. P. Leite, P. H. P. de Carvalho, and R. D. Vieira, “A flexible framework based on reinforcement learning for adaptive modulation and coding in OFDM wireless systems,” in Wireless Communications and Networking Conference (WCNC). IEEE, 2012, pp. 809–814. [11] K. I. Pedersen, G. Monghal, I. Z. Kovacs, T. E. Kolding, A. Pokhariyal, F. Frederiksen, and P. Mogensen, “Frequency domain scheduling for OFDMA with limited and noisy channel feedback,” in IEEE 66th Vehicular Technology Conference, 2007, pp. 1792–1796. [12] P. Kela, T. Höhne, T. Veijalainen, and H. Abdulrahman, “Reinforcement learning for delay sensitive uplink outer-loop link adaptation,” in Joint European Conference on Networks and Communications & 6G Summit. IEEE, 2022, pp. 59–64. [13] R. Wiesmayr, L. Maggi, S. Cammerer, J. Hoydis, F. A. Aoudia, and A. Keller, “SALAD: Self-adaptive link adaptation,” arXiv preprint arXiv:2510.05784, 2025. [14] L. Maggi, B. Bonev, R. Wiesmayr, S. Cammerer, and A. Keller, “SINR Estimation under Limited Feedback via Online Convex Optimization,” arXiv preprint arXiv:2603.02061, 2026. [15] 3rd Generation Partnership Project, “Nr; physical layer procedures for data (3gpp ts 38.214),” 2023. [16] J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal Policy Optimization Algorithms,” 2017, [Available Online]: https://arxiv.org/abs/1707.06347. [17] A. Raffin, A. Hill, A. Gleave, A. Kanervisto, M. Ernestus, and N. Dormann, “Stable-Baselines3: Reliable Reinforcement Learning Implementations,” GitHub Repository, 2021, [Available Online]: https://github. com/DLR-RM/stable-baselines3. [18] Farama Foundation, “Gymnasium: A Standard Interface for Reinforcement Learning Environments,” GitHub Repository, 2023, [Available Online]: https://github.com/Farama-Foundation/Gymnasium. [19] J. R. Quinlan, “Induction of decision trees,” Machine learning, vol. 1, no. 1, pp. 81–106, 1986. [20] R. E. Kalman, “A New Approach to Linear Filtering and Prediction Problems,” Transactions of the ASME–Journal of Basic Engineering, vol. 82, pp. 35–45, 1960. [21] L. Breiman, “Random Forests,” Machine Learning, vol. 45, 2001. [22] NVlabs, “SALAD: Self-Adaptive Link Adaptation — Reference Implementation,” GitHub Repository, NVlabs, 2025, [Available Online]: https://github.com/NVlabs/salad/blob/main/notebooks/meet salad.ipynb. [23] D. Ernst, P. Geurts, and L. Wehenkel, “Tree-based batch mode reinforcement learning,” Journal of Machine Learning Research, vol. 6, 2005. [24] D. Villa, I. Khan, F. Kaltenberger, N. Hedberg, R. S. da Silva, S. Maxenti, L. Bonati, A. Kelkar, C. Dick, E. Baena et al., “X5G: An open, programmable, multi-vendor, end-to-end, private 5G O-RAN testbed with NVIDIA ARC and OpenAirInterface,” IEEE Transactions on Mobile Computing, 2025.

Record · ID 238546 · SHA-256 63c7be671f71852f
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.