Conceptio › Archive › arXiv CS
arXiv CSopen access

AI-Native Open RAN: A Roadmap from xApps and rApps to Autonomous Network Agents

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
distributed-systemsinternetnetworkingprotocols
networking, internet, protocols, distributed systems

AI-Native Open RAN: A Roadmap from xApps and rApps to Autonomous Network Agents Ryan Barker, Alireza Ebrahimi Dorcheh, Tolunay Seyfi, Mohammad Raihan Uddin, Alireza Mohammadhosseini, Julia Boone, Stephen Streit, Drew Schlesener, Fatemeh Afghah

arXiv:2609.15704v1 [cs.NI] 14 Sep 2026

Holcombe Department of Electrical and Computer Engineering, Clemson University, Clemson, SC, USA or manage a model over seconds to hours. A near-RT xApp can adjust slicing, power, mobility, or interference decisions on an approximately 10 ms to 1 s loop. Scheduler-local logic and distributed applications (dApps) are proposed for decisions below 10 ms, where transport to a centralized RIC is already too slow [3]. In this paper, we use the term autonomous network agent to denote an intelligent network function that observes semantically defined network state, reasons over current and predicted conditions, selects actions within explicitly authorized control boundaries, coordinates with other network functions, and adapts or relinquishes authority when its operating assumptions are no longer supported. Such an agent therefore extends beyond a task-specific xApp or rApp: autonomy requires not only decision capability, but also generalization, coordination, assurance, and bounded actuation across control timescales. Despite this breadth of applications, many existing approaches learn a narrowly defined task from a prescribed set of observations and objectives. A policy trained for network slicing does not automatically provide useful knowledge for mobility management or interference mitigation, even when these functions depend on overlapping network conditions. Changes in traffic, topology, propagation, or available telemetry can also require adaptation within the same task [4]. MetaI. I NTRODUCTION learning, domain-shift-approached and coordinated training, Open Radio Access Network (Open RAN) changes the address aspects of this problem, but often retain separate radio access network from a tightly integrated appliance into a representations and training procedures for individual functions. programmable system assembled from disaggregated functions, As AI becomes more widely integrated into the RAN, the cost open interfaces, cloud platforms, and independently devel- of developing and maintaining these specialized models raises oped control applications. The O-RAN Alliance architecture a broader question: how much knowledge can be shared across makes this programmability operational through the service network tasks and deployments? Foundation models offer a potential basis for such reuse management and orchestration framework, the non-real-time RAN Intelligent Controller (non-RT RIC), the near-real-time through broad pretraining followed by adaptation to downRIC (near-RT RIC), and E2-connected radio nodes. These stream tasks. In wireless systems, this could involve learning components create distinct places to analyze long-horizon data, representations of radio signals, channels, or network behavior train and govern models, and execute closed-loop actions. that support multiple prediction and control functions [5]. Surveys of the architecture have documented the resulting Language and multimodal models offer a complementary interfaces, algorithms, security questions, and deployment capability by incorporating service requirements, operational opportunities, while broader 6G analyses identify Open RAN as knowledge, and application context into network decisions. a practical substrate for increasingly software-defined cellular Their value, however, must be established through demonstrated transfer, reduced adaptation cost, and improved downstream systems [1], [2]. The central research opportunity is not simply to place performance under realistic deployment constraints. This paper examines AI-RAN from that perspective, using artificial intelligence (AI) inside the RIC. It is to determine Open RAN as the architectural setting for deploying and which intelligence belongs at each control timescale, what state it can observe, what actions it is authorized to issue, and what evaluating intelligence. We first organize existing approaches evidence is required before those actions reach a live radio by their prediction, control, coordination, and adaptation roles, system. A non-RT rApp may forecast demand, generate policy, relating each role to its observations, actions, and execution timescale. We then examine the limitations of task-specific This work was supported, in part, at Clemson University by the State of learning and assess how foundation models could support South Carolina through funding for the Battelle Savannah River Alliance reusable wireless representations and broader network context. Workforce Development Program and NVIDIA Academic Development Grant. Finally, we identify the data, adaptation, and experimental It is also funded by the National Science Foundation under Grant Numbers CNS-2202972, CNS- 2318726, and CNS-2232048. requirements needed to evaluate this transition. Abstract—Open Radio Access Networks (O-RAN) have emerged as a transformative paradigm for future wireless systems by introducing openness, virtualization, disaggregation, and programmable intelligence through the RAN Intelligent Controller (RIC). The availability of standardized interfaces and near-realtime control loops has created unprecedented opportunities for integrating artificial intelligence (AI) into radio access network management and optimization. Over the past several years, a broad range of AI techniques have been proposed to address key O-RAN challenges such as radio resource management, network slicing, traffic prediction, mobility management, interference mitigation, and spectrum sharing. Despite significant progress, existing solutions often remain task-specific, require extensive retraining, and exhibit limited generalization across deployment environments and network conditions. This paper presents a comprehensive review of AI-enabled ORAN systems and provides a unifying perspective on the evolution of intelligence in wireless networks. We first examine the O-RAN architecture and the role of intelligence within near-real-time and non-real-time RIC frameworks. We then develop a taxonomy of AI approaches for O-RAN, covering machine learning, deep reinforcement learning (DRL), digital-twin-assisted optimization, and emerging foundation-model-based architectures. Index Terms—AI-native RAN (AI-RAN), Open RAN (ORAN), RAN Intelligent Controller (RIC), xApp, rApp, dApp, reinforcement learning.

Fig. 1. Control and learning hierarchy for AI-native Open RAN. Intelligence moves from long-horizon policy and model lifecycle management toward nearRT control, real-time radio logic, and inline execution as deadlines tighten. MEC objectives and assurance evidence must be bound to the same state and action contracts.

KPIs. Deployability therefore depends not only on latency but also on whether the required observations and actions are exposed by the RAN. Faster functions can move closer to the radio. dApps execute alongside O-CU/O-DU functions to shorten the path between observation, inference, and actuation [3], with demonstrations in real-time AI control [21], spectrum classification [22], and ISAC inference [23]. Table I summarizes representative mappings among tasks, observations, execution environments, and interfaces. These mappings distinguish algorithmic performance from deployability. An operational controller must obtain the required telemetry, execute within its deadline, and translate its output into supported RAN actions. Measurements of AI inference inside a Near-RT RIC xApp demonstrate why the complete control loop matters beyond inference latency alone [24]. OpenRAN Gym [25], OAIC [26], and broader testing frameworks [27] provide platforms for evaluating these system-level requirements.

II. I NTELLIGENCE ACROSS THE O-RAN C ONTROL S TACK

C. Closed-Loop Control with Reinforcement Learning

A. The O-RAN Architecture

Reinforcement learning (RL) naturally models closed-loop O-RAN control as repeated interaction between network state and RAN actions. Early work applied RL to dynamic resource allocation across O-RAN slices, establishing the basic loop in which observed resource state and service demand determine an allocation whose outcome informs subsequent decisions [28]. More recent work embeds this loop in operational O-RAN systems. REAL integrates an RL-enabled xApp with the ORAN Software Community (OSC) Near-RT RIC and srsRAN for experimental closed-loop slicing [29]. DORA uses online Proximal Policy Optimization (PPO) for dynamic multi-slice resource allocation in an OpenAirInterface-based environment [30]. Other implementations connect DRL resource allocation to executable OpenAirInterface control [31] and adapt near-real-time slicing to channel, traffic, mobility, and QoS conditions [32]. Deployability ultimately depends on the complete control loop. Observation intervals determine reaction speed, action spaces determine available control authority, and rewards must balance throughput, latency, fairness, efficiency, and Service Level Agreement (SLA) requirements. Communication and inference delays further create a mismatch between the observed state and the state in which an action is applied. Evaluation should therefore combine learning metrics with endto-end latency, action frequency, signaling overhead, robustness, and network performance. The progression from slice optimization [28] through experimental and online RL control [29]–[32] to integrated traffic-aware control [33], [34] reflects the broader transition toward continuously operating O-RAN control loops.

The O-RAN architecture extends the 3GPP next generation (NG)-RAN through disaggregation, open interfaces, cloudnative deployment, and programmable control. As illustrated in Fig. 1, the RAN is decomposed into the O-RAN Centralized Unit (O-CU), Distributed Unit (O-DU), and Radio Unit (ORU), while the Service Management and Orchestration (SMO), Non-Real-Time RAN Intelligent Controller (Non-RT RIC), and Near-Real-Time RAN Intelligent Controller (Near-RT RIC) provide management and programmable intelligence [6]. The O-CU may be divided into control-plane and userplane functions, with F1 connecting the O-CU and O-DU and NG connecting the NG-RAN to the 5G Core [7]. The O-DU connects to the O-RU through the Open Fronthaul using the 7-2x functional split [8]. Above these functions, the SMO manages O-RAN functions and cloud infrastructure through O1 and O2 [9]. The Non-RT RIC supports long-horizon optimization, policy management, data processing, and AI/ML lifecycle functions. rApps interact with this framework through R1, while A1 conveys policies, enrichment information, and AI/ML guidance toward the Near-RT RIC [10]. The NearRT RIC hosts xApps and interacts with E2 Nodes through E2 Service Models that define available measurements and control operations [11]. Together, these components separate data collection, policy coordination, decision making, and RAN actuation across distinct control timescales. B. Intelligent Control Functions in O-RAN O-RAN distributes intelligence across rApps in the Non-RT RIC, xApps in the Near-RT RIC, and emerging dApps colocated with RAN functions. Placement depends on the observations, actions, and latency required by each task. For near-real-time control, E2 Service Model for Key Performance Measurements (E2SM-KPM) provides cell- and UE-level measurements while E2 Service Model for RAN Control (E2SM-RC) exposes supported control services, allowing xApps to observe network state and apply actions supported by the E2 Node. This framework supports resource allocation and slicing [12]– [15], traffic steering and connection management [16], [17], hierarchical control [18], [19], and energy management [20]. Some tasks require specialized observations beyond standard

III. G ENERALIZATION AND R EUSABLE I NTELLIGENCE FOR AUTONOMOUS O-RAN RL controllers trained under one set of traffic, channel, mobility, topology, or service conditions may degrade when deployed under another. Generalization concerns preserving useful behavior under such distribution shifts rather than merely adapting within conditions represented during training. Existing O-RAN approaches address this problem through transfer- and meta-learning, robust and distributed learning, environment diversification, transferable representations, and uncertaintyaware adaptation.

A. Transfer Learning and Policy Reuse

F. Uncertainty-Aware and Model-based Adaptation Generalizable controllers should also recognize unfamiliar Transfer learning reuses policies, parameters, value estimates, conditions. Bayesian RL incorporates uncertainty into action or representations rather than retraining from random initialselection and has been applied to joint O-RAN and Multi-access ization. Deep transfer RL has been applied to radio and cache Edge Computing (MEC) orchestration [52]. Model-based RL allocation [35], while O-RAN studies use policy reuse and complements this capability by learning network dynamics distillation for slicing [36], repositories of trained agents for for planning and safer adaptation, including model-based safe new service configurations [37], and policy reuse across nonslicing designed to reduce SLA violations [45]. Uncertainty and terrestrial operating conditions [38]. Its effectiveness depends learned dynamics can therefore trigger conservative behavior, on similarity between source and target environments, since adaptation, policy replacement, or higher-level intervention large shifts can produce negative transfer. when deployment conditions depart from training experience. G. Representation-based Generalization Generalization also depends on how network state is repreMeta-learning instead trains controllers to adapt rapidly to sented. Attention-based slicing emphasizes relevant portions new tasks. Meta-RL accelerates adaptation across changing O- of network state [4], while predictive methods incorporate RAN conditions [39], multi-task initialization supports unseen learned traffic dynamics into RL decisions [53]. More recent slicing objectives [40], and hierarchical meta-RL combines approaches introduce semantic context. LLM-augmented DRL rapid adaptation with decomposed resource management [41]. incorporates contextual information [54], prompt tuning adapts These benefits still depend on training-task diversity, since this representation to resource management [55], and ORANnarrow meta-training may not support changes in topology, GUIDE combines retrieval, learned prompting, and RL [56]. mobility, scale, or objectives. Wireless foundation models extend this direction through reusable features across wireless tasks [5], while ORANSight2.0 explores foundation models for O-RAN knowledge and C. Robust Optimization and Safe Generalization reasoning [57]. These approaches shift generalization from task-specific Robust learning seeks policies that remain effective as conditions change. Sharpness-aware optimization reduces policies tied to observations toward representations that capture sensitivity to policy perturbations in multi-agent O-RAN relationships among radio conditions, traffic, resources, service resource management [42]. SafeSlice [43], robust SLA-aware requirements, and network context. The next challenge is control [44], and model-based safe RL [45] additionally seek whether such representations can support multiple control to preserve operational constraints during distribution shifts. functions and ultimately provide a common abstraction not MORPH broadens training across multiple environments rather only for observing the RAN but also for acting upon it. than optimizing a Physical Resource Block (PRB) controller for H. Foundation Models, World Models, and Reusable Network a single network realization [46]. Together, these approaches Intelligence emphasize that useful generalization requires both performance Reusable representations alone do not provide general preservation and safe adaptation. control. Most RL controllers still operate within predefined state and action spaces, while transfer and meta-learning generally D. Federated, Multi-Agent, and Population-Based Learning assume compatible tasks or policies. The remaining challenge is to connect representations learned across environments and Distributed learning broadens the experience available dur- network functions to controllers with different observations, ing training. Federated RL shares model knowledge across objectives, and actuation capabilities. heterogeneous local conditions without exchanging complete The emergence of wireless foundation models (WFMs) datasets and has been applied to O-RAN slicing and resource marks a paradigm shift from fragmented, task-specific deep allocation [14], [47]. Multi-agent approaches decompose re- learning toward unified, adaptable architectures for nextsource allocation and coordinated control among interacting generation (6G) networks. Pre-trained on large, unlabeled radio policies [13], [18], [19]. Population-based methods further frequency (RF) datasets spanning raw in-phase/quadrature (IQ) diversify policies through federated neuroevolution [48] and samples, channel state information (CSI), and spectrograms evolutionary RL [49]. Their generalization benefit ultimately de- using self-supervised learning [58], [59], these architectures pends on whether training exposes the policies to meaningfully extract universal physical-layer representations. Adopting foundifferent conditions. dation models in telecommunications bridges long-standing operational silos, unifying core physical-layer operations like channel estimation, beam prediction, and signal identification E. Digital Twins and Environment Diversification with integrated sensing and communication (ISAC) and agentic Digital twins and programmable testbeds expose controllers network orchestration. to broader variations in traffic, propagation, mobility, topology, This adoption is of conventional supervised models that interference, and resources without unsafe exploration on require expensive data labeling and suffer acute performance production infrastructure. Proposed RAN twins [50], Colos- degradation in dynamic, out-of-distribution wireless environseum [51], and OpenRAN Gym [25] support repeatable training ments. By contrast, WFM decouples universal feature extraction and validation. Reliable transfer nevertheless requires both from downstream task execution. This enables powerful zeroenvironment diversity and sufficient fidelity to deployment shot and few-shot transfer learning, allowing networks to conditions. adapt to unseen propagation environments or emerging tasks B. Meta-Learning and Rapid Adaptation

with minimal parameter fine-tuning [60]. Ultimately, adopting itself. Adapting video bitrate and inference model complexity foundation models provides the computational efficiency, gen- changes transmission and execution demand while preserving eralizability, and multi-task scalability required to realize truly an application-specific accuracy requirement [70]. Fine-grained autonomous, AI-native wireless ecosystems. DNN partitioning further distributes execution across devices Transferable representations provide reusable state abstrac- and edge servers, making intermediate-feature transmission tions, but they do not by themselves capture how the network and heterogeneous computing capabilities part of the latency evolves under candidate control actions. Model-free policies trade-off [76]. These decisions couple where an application similarly select actions without explicitly predicting their runs with how it is executed. consequences. World models complement these approaches by Allocated computing capacity, however, does not directly learning network dynamics, allowing a decision to be evaluated imply predictable service time. On shared accelerators, GPU in imagination before execution [61] and recent works learn operator characteristics and interference between concurrent latent dynamics over KPI and resource trajectories to support workloads affect execution latency, motivating fine-grained planning without online trials [62], [63]. scheduling rather than a uniform computing-cycle abstraction [77]. Deadline-sensitive MEC additionally requires visibility IV. M ULTI -ACCESS E DGE C OMPUTING into request progress and mechanisms to adjust CPU allocation The transition toward autonomous O-RAN intelligence and GPU priorities at runtime [71]. The generalization problem extends control beyond radio resources to the applications from Section III consequently extends to changes in model they serve. MEC introduces service placement, application configuration, execution placement, and hardware contention execution, and computing contention into this decision process. as well as wireless channels. A transferable policy must retain Low packet latency alone does not ensure timely application an adequate representation of these changing constraints. completion: transmission, queueing, and execution consume a shared deadline budget [69]–[71]. We therefore organize C. Cross-Layer Coordination and Service Assurance Application-aware scheduling provides a bridge between MEC research along three complementary problem dimensionsservice placement and continuity, task offloading and appli- service objectives and RAN actuation. Application context cation adaptation, and resource allocation and execution-and and channel estimates can be combined to anticipate frame then consider cross-layer coordination and end-to-end service demand and prioritize radio grants according to approaching assurance as a synthesis dimension. As in Sections II and III, deadlines [69]. Extending this approach across the application the relevant distinctions are the information available to each and network allows resource information and latency budgets to inform both radio and GPU scheduling [70]. Such coordination controller, its actuation scope, and its operating timescale. exposes dependencies that are obscured when transmission and A. Architecture, Placement, and Service Continuity execution are optimized independently. Explicit cross-layer exchange is one coordination approach; The MEC architecture separates hosts, comprising the platform and virtualized infrastructure, from system-level local inference is another. Request activity can be inferred at orchestration responsible for host selection and application the RAN from standard 5G control signals, while probing lifecycle management [72]. In 5G, local user plane routing and application lifecycle instrumentation provide the edge through an appropriately selected User Plane Function (UPF) with estimates of the remaining deadline budget. Independent provides connectivity to edge application servers, comple- schedulers can then adapt radio resources, CPU allocation, mented by application-server discovery and traffic-steering GPU priorities, and early dropping without direct RAN–edge or relocation procedures [73]. These mechanisms extend the coordination [71]. The resulting architectural choice concerns service environment around the RIC architecture in Section II. how much state is exchanged, how accurately unobserved MEC application management, 5G session control, and RAN progress can be inferred, and whether information remains control retain distinct responsibilities even when their functions useful when an action takes effect. Neither explicit coordination nor local inference removes the need to account for observation share infrastructure. Service placement determines where application instances error and control delay. and their state reside; resource provisioning determines the capacity available to execute their workloads. Joint orchestra- D. Implications for Autonomous O-RAN Agents tion formulations connect hosting locations, O-RAN functional Hierarchical coordination has already been investigated in splits, routing, and computing allocations to infrastructure cost O-RAN: operator objectives can guide the selection of lowerand delay [52]. Placement must also accommodate demand level xApps through hierarchical reinforcement learning [78], variation: robust formulations distinguish advance resource while proposed dApp architectures extend control to RANreservation from subsequent placement and workload adapta- local execution under tighter timing constraints [21]. Recent tion, accounting for spatial and temporal demand correlations agentic frameworks further place intent interpretation and model [74]. This introduces a planning problem that extends beyond governance in longer-timescale control loops, coupled with the instantaneous radio state. faster control and inference functions [79]. These precedents provide an architectural basis for discussing MEC integration; B. Offloading, Adaptation, and Heterogeneous Execution the hierarchy itself should not be regarded as a new contribution Task offloading connects application demand to communica- of this roadmap. tion and computing budgets. Learning-based policies can jointly For MEC, the additional requirement is to connect these decide whether tasks execute locally or are offloaded to a MEC control layers to service placement and application execution. server, while allocating communication and computational Joint O-RAN/MEC orchestration and adaptive placement resources under differentiated latency and energy objectives formulations establish relevant provisioning decisions [52], [74], [75]. For AI services, this action space extends to the workload while learning-based task-offloading and resource allocation

TABLE I R EPRESENTATIVE M APPING OF I NTELLIGENT RAN F UNCTIONS TO O-RAN A PPLICATIONS RAN function Resource allocation

Representative observations

Application

Traffic, CQI, throughput, resource xApp utilization Network slicing Slice traffic, latency, throughput, xApp/rApp PRB utilization Traffic steering UE radio measurements, throughput, xApp cell load Connection UE–cell measurements, load, inter- xApp management ference rApp/xApp Intent-driven orchestra- Service intent and network state tion Energy management Load, traffic history, utilization rApp/xApp Spectrum sensing Spectrum/PHY observations, interfer- xApp/dApp ence state Localization SRS channel estimates, radio mea- xApp surements Fast PHY/MAC control PHY state, I/Q, channel information dApp

Interface/service requirement Representative work E2 telemetry and RAN control Actor–critic [12], team learning [13] E2 control and A1 policy

Federated DRL [14], ORANSlice [64], VR slicing [15]

E2SM-KPM and E2SM-RC

Traffic steering [16], hierarchical control [18]

E2 monitoring and control

GNN/RL xApp [17]

Non-RT policy and near-RT Hierarchical orchestration [19] control Policy and near-RT control DRL energy saving [20], ScalO-RAN [65] E2 or lower-layer data exposure SenseORAN [66], LibIQ [22] SRS-oriented E2 service model OAI/FlexRIC localization [67] Low-latency RAN interface

dApps [3], [21], ISAC inference [23], [68]

policies [75] and application-aware MEC systems expose com- incoming KPI windows toward that subspace before inference, plementary mechanisms for coordinating radio and execution and using the projection residual as an anomaly signal [90], deadlines [70], [71]. Building on this evidence, we identify [91]. Evaluation on Colosseum ColO-RAN telemetry shows that the coordination of placement, application configuration, and such sanitization can substantially recover policy behavior for runtime scheduling as a design requirement for autonomous detectable triggers, while performance degrades when malicious RAN–MEC operation. Application lifecycle information and perturbations lie within the learned safe subspace [91], [92]. accelerator controls must be integrated through the relevant This limitation illustrates a broader principle for autonomous MEC and runtime mechanisms [71], [72], rather than inferred O-RAN: anomaly mitigation should inform and bound control to be available merely because a RAN control interface exists. authority rather than establish unconditional trust. Receiver feedback forms an important trust boundary because The research question is therefore how to compose these capabilities under an end-to-end service objective. For the reports such as the Channel Quality Indicator (CQI), Rank Indiroadmap considered here, we propose evaluating this composi- cator (RI), Precoding Matrix Indicator (PMI), and Hybrid Aution under workload shifts, mobility, stale observations, and tomatic Repeat reQuest (HARQ) directly influence scheduling resource failures, with explicit checks on control authority and and link adaptation [93]. Prior work shows that channel infordefined handling of infeasible actions. These are evaluation mation can be cross-checked using acknowledgement/negativerequirements of our synthesis, not guarantees established by acknowledgement (ACK/NACK) feedback [94], while decepthe cited systems. They connect MEC service assurance to the tive RI signaling can distort Proportional Fair (PF) scheduler behavior and disadvantage honest users through its effect on trust and control-security questions addressed in Section V. scheduler memory [95], [96]. These results show that authentiV. Z ERO T RUST S ECURITY IN O PEN RAN cated or protocol-valid feedback should not automatically be treated as trustworthy state. For autonomous O-RAN control, The same interoperable interfaces that enable ORAN proindependently observable evidence should therefore be used to grammability also widen its threat surface [80]. A1 policy, assess reported state and to restrict the authority of a policy E2 telemetry and control, O1/O2 management, and thirdwhen the evidence is inconsistent. party xApps introduce trust boundaries alongside inherited These results establish a progression from authenticating 5G New Radio vulnerabilities [1], [81]. SNI5GECT demonactors to validating evidence and finally bounding control strates pre-authentication injection, software attacks can comauthority. Lifecycle governance and cross-site analysis can promise insufficiently protected packetized fronthaul, and remain in non-RT control, reversible adaptation can operate sidelink and location-aware attacks expose additional radioin the near-RT layer, and only deadline-critical checks should local threats [82]–[84]. Such attacks can corrupt control inputs move toward local execution [3], [80]. Deployment therefore or selectively deny service without compromising the 5G requires fresh evidence, explicit UE and bearer scope, bounded Core, making deadline and service failures relevant security actions, and tested fallback behavior. Security evaluation must outcomes. likewise account for observation age and complete observationZero Trust couples continuous verification with least- to-actuation delay rather than inference latency alone [97]. privilege authorization [85]. Interface protections provide a VI. D ISCUSSION AND O PEN R ESEARCH Q UESTIONS transport baseline, while OZTrust and XRF restrict xApp communication and service access [81], [86]–[88]. Yet authenticated The reviewed literature suggests that progress toward auactors and protected transport do not establish that subsequent tonomous Open RAN depends not only on increasingly capable telemetry is semantically truthful or bound to the current UE models, but on whether intelligence can generalize across and bearer context [80]. operating conditions, compose across control functions and This distinction becomes critical for learning-enabled control. timescales, act within evidence-supported authority, and satisfy A backdoored deep reinforcement learning (DRL) xApp can measurable deployment constraints. From this perspective, five behave normally until poisoned observations activate harmful research challenges define the transition from task-specific actions [89]. ORAN-DEFEND applies subspace-based saniti- xApps and rApps toward autonomous network agents. zation to frozen black-box O-RAN policies by using trusted The first is generalization with bounded authority. Transfer trajectories to identify a safe observation subspace, projecting learning, meta-learning, robust optimization, multi-environment

training, and foundation model representations reduce adaptation cost, but none eliminates distribution shift. Future evaluations should therefore characterize a transfer envelope over traffic, channel conditions, topology, software configuration, available telemetry, and execution environment. Controllers should detect operation outside this envelope and reduce their action authority or invoke a validated fallback rather than extrapolate with unchecked confidence. The second challenge is composition across intelligent functions. Multiple rApps, xApps, dApps, and MEC agents may produce individually reasonable decisions that compete for the same radio or computing resources. Learning-based coordination alone does not resolve this runtime problem. Resources such as PRBs, power, handover parameters, accelerator capacity, and model placement require explicit scope, priority, and validity intervals. Arbitration must also distinguish conflicting policies from malicious or untrusted inputs because both can produce unsafe actions but require different recovery mechanisms. Authentication should therefore establish origin without implying semantic validity or unrestricted control authority [80]. The third challenge is evidence-conditioned control authority. Increasing model capability does not imply that a controller should exercise unrestricted authority. An autonomous agent may operate with stale observations, unfamiliar network conditions, conflicting policies, unavailable actuators, or evidence whose integrity cannot be independently established. Future systems therefore need mechanisms that translate the quality and freshness of available evidence into explicit control boundaries: which actions are permitted, over which resources and users, for how long, and with what fallback when confidence degrades. This shifts Zero Trust from authentication alone toward a runtime relationship between evidence, confidence, and actuation authority. The fourth challenge is experimentally credible autonomy. Simulation, emulation, software-defined-radio systems, and live O-RAN platforms preserve different aspects of network behavior and should support correspondingly different claims. Evaluation should separate four relevant clocks. The radio clock governs scheduling and retransmission, the control clock spans observation through actuation, the application clock determines whether completed service retains value, and the learning clock governs model adaptation and promotion. Reporting inference latency alone can therefore make a controller appear deployable even when its observations are stale or its actions miss radio or application deadlines. Relevant measurements include state age, end-to-end control latency, actuation success, deadline success, radio and compute utilization, and recovery under failed or untrusted inputs [97]. The fifth challenge is semantic interoperability. Portable models are insufficient when implementations disagree about the meaning of observations and actions. A deployable controller should carry a machine-readable contract describing input semantics, normalization, update rate, action units, target scope, dependencies, validated operating envelope, and fallback behavior. Model and policy versions should remain bound to the service-model schema and software stack on which they were validated. O-RAN interfaces provide communication interoperability, but autonomous operation additionally requires preservation of these state and action semantics across O1, A1, E2, local APIs, and application services [1]. These requirements also define a practical path toward

increasing autonomy. Controllers can progress from offline evaluation to shadow operation and then to bounded control over selected cells, slices, or services. Promotion should depend on stored evidence and support rollback, while execution should produce an audit trail connecting observed state, selected policy, issued action, outcome, and fallback status. Security evaluation should similarly distinguish algorithmic recovery from service recovery because success in KPI replay or slot-level simulation does not establish end-to-end containment under live timing constraints [91], [96], [97]. Foundation models fit naturally into this architecture, but their near-term role is more compelling in representation, planning, and coordination than in direct fast-timescale actuation. They can translate human intent and network knowledge into policies, retrieve operational context, select or configure specialist models, and provide reusable RF representations [55]– [57], [60]. Near-real-time and lower-layer actions should remain grounded through typed interfaces and bounded controllers whose timing and behavior can be validated. The roadmap toward autonomous O-RAN is therefore not a progression from small models to larger models, nor from xApps to a single monolithic agent. It is a progression from isolated task-specific controllers toward transferable and predictive intelligence that can coordinate specialized functions across timescales while acquiring, exercising, and relinquishing control authority according to current evidence and validated system constraints. VII. C ONCLUSION AI-native Open RAN is evolving from isolated optimization applications toward coordinated network agents operating across multiple control timescales. The reviewed literature traces this progression through adaptive and generalizable RL, semantic representations, closed-loop xApps, MEC and radio-compute orchestration, accelerated execution, and security assurance. Across these developments, increasing autonomy requires more than increasingly capable models. It requires transferable state representations, explicit action semantics, measurable timing, bounded authority, and validated fallback behavior. rApps, xApps, dApps, MEC agents, and inline functions can therefore operate as complementary components of a common control architecture rather than independent models competing for resources and authority. Foundation models can broaden representation, reasoning, and coordination, while time-critical actuation remains grounded in bounded and verifiable controllers. Progress toward autonomous Open RAN will ultimately depend on whether intelligence can generalize across deployments while preserving interoperability, timing guarantees, and control authority supported by current evidence. R EFERENCES [1] M. Polese, L. Bonati, S. D’Oro, S. Basagni, and T. Melodia, “Understanding O-RAN: Architecture, interfaces, algorithms, security, and research challenges,” IEEE Communications Surveys & Tutorials, vol. 25, no. 2, pp. 1376–1411, 2023. [2] M. Polese, M. Dohler, F. Dressler, M. Erol-Kantarci, R. Jana, R. Knopp, and T. Melodia, “Empowering the 6G cellular architecture with open RAN,” IEEE Journal on Selected Areas in Communications, vol. 42, no. 2, pp. 245–262, 2024. [3] S. D’Oro, M. Polese, L. Bonati, H. Cheng, and T. Melodia, “dapps: Distributed applications for real-time inference and control in O-RAN,” IEEE Communications Magazine, vol. 60, no. 11, pp. 52–58, 2022. [4] F. Lotfi, F. Afghah, and J. Ashdown, “Attention-based open RAN slice management using deep reinforcement learning,” in Proc. IEEE Global Communications Conference, 2023, pp. 6328–6333.

[5] S. Alikhani, G. Charan, and A. Alkhateeb, “Lwm: A pre-trained wireless foundation model for universal feature extraction,” in 2025 IEEE International Conference on Machine Learning for Communication and Networking (ICMLCN), 2025, pp. 1–6. [6] O-RAN ALLIANCE, “O-RAN Architecture Description,” O-RAN ALLIANCE, Tech. Rep. O-RAN.WG1.OAD, 2026, o-RAN Technical Specification. [7] 3GPP, “NG-RAN; Architecture Description,” 3rd Generation Partnership Project (3GPP), Tech. Rep. TS 38.401, 2026, 3GPP Technical Specification. [8] O-RAN ALLIANCE, “O-RAN Open Fronthaul Control, User and Synchronization Plane Specification,” O-RAN ALLIANCE, Tech. Rep. O-RAN.WG4.CUS, 2026, o-RAN Technical Specification. [9] ——, “O-RAN Operations and Maintenance Architecture and O2 Interface Specifications,” O-RAN ALLIANCE, Tech. Rep., 2026, o-RAN Technical Specifications. [10] ——, “O-RAN A1 Interface: General Aspects and Principles,” O-RAN ALLIANCE, Tech. Rep. O-RAN.WG2.A1GAP, 2025, o-RAN Technical Specification. [11] ——, “O-RAN E2 General Aspects and Principles,” O-RAN ALLIANCE, Tech. Rep. O-RAN.WG3.E2GAP, 2025, o-RAN Technical Specification. [12] M. Kouchaki and V. Marojevic, “Actor-critic network for O-RAN resource allocation: xapp design, deployment, and analysis,” in Proc. IEEE Globecom Workshops, 2022. [13] H. Zhang, H. Zhou, and M. Erol-Kantarci, “Team learning-based resource allocation for open radio access network (O-RAN),” in Proc. IEEE International Conference on Communications (ICC), 2022. [14] ——, “Federated deep reinforcement learning for resource allocation in O-RAN slicing,” in Proc. IEEE Global Communications Conference, 2022, pp. 958–963. [15] A. Casparsen, B. Soret, J. J. Nielsen, and P. Popovski, “Near real-time data-driven adaptive RAN slicing of VR traffic in ORAN,” IEEE Open Journal of the Communications Society, vol. 6, pp. 4624–4637, 2025. [16] A. Lacava, M. Polese, R. Sivaraj, R. Soundrarajan, B. S. Bhati, T. Singh, T. Zugno, F. Cuomo, and T. Melodia, “Programmable and customized intelligence for traffic steering in 5G networks using open RAN architectures,” IEEE Transactions on Mobile Computing, vol. 23, no. 4, pp. 2882–2897, 2024. [17] O. Orhan, V. N. Swamy, T. Tetzlaff, M. Nassar, H. Nikopour, and S. Talwar, “Connection management xapp for O-RAN RIC: A graph neural network and reinforcement learning approach,” in Proc. IEEE International Conference on Machine Learning and Applications, 2021. [18] M. A. Habib, H. Zhou, P. E. Iturria-Rivera, Y. Ozcan, M. Elsayed, M. Bavand, R. Gaigalas, and M. Erol-Kantarci, “Machine learningenabled traffic steering in O-RAN: A case study on hierarchical learning approach,” IEEE Communications Magazine, vol. 63, no. 1, pp. 100–107, 2025. [19] M. A. Habib, H. Zhou, P. E. Iturria-Rivera, M. Elsayed, M. Bavand, R. Gaigalas, Y. Ozcan, and M. Erol-Kantarci, “Intent-driven intelligent control and orchestration in O-RAN via hierarchical reinforcement learning,” in Proc. IEEE 20th International Conference on Mobile Ad Hoc and Smart Systems (MASS), 2023, pp. 55–61. [20] M. Bordin, A. Lacava, M. Polese, S. Satish, M. A. Nittoor, R. Sivaraj, F. Cuomo, and T. Melodia, “Design and evaluation of deep reinforcement learning for energy saving in open RAN,” in Proc. IEEE 22nd Consumer Communications & Networking Conference (CCNC), 2025. [21] A. Lacava, L. Bonati, N. Mohamadi, R. Gangula, F. Kaltenberger, P. Johari, S. D’Oro, F. Cuomo, M. Polese, and T. Melodia, “dapps: Enabling real-time AI-based open RAN control,” Computer Networks, vol. 269, p. 111342, 2025. [22] F. Olimpieri, N. Giustini, A. Lacava, S. D’Oro, T. Melodia, and F. Cuomo, “LibIQ: Toward real-time spectrum classification in O-RAN dapps,” in Proc. 23rd Mediterranean Communication and Computer Networking Conference (MedComNet), 2025, pp. 1–6. [23] M. Polese, R. Gangula, and T. Melodia, “Enabling programmable inference and ISAC at the 6G RAN edge with dapps,” arXiv preprint arXiv:2603.29146, 2026. [24] L. Obiuwevwi, K. J. Rechowicz, S. Jayarathna, S. H. Bouk, F. Afrin, C. N. Barati, N. Moghim, V. Nanou, M. E. Rahman, and S. Shetty, “Enabling real-time AI in O-RAN: Deploying and measuring AI inside a near-rt RIC xapp,” arXiv preprint arXiv:2607.01583, 2026. [25] L. Bonati, M. Polese, S. D’Oro, S. Basagni, and T. Melodia, “OpenRAN Gym: AI/ML development, data collection, and testing for O-RAN on PAWR platforms,” Computer Networks, vol. 220, p. 109502, 2023. [26] P. S. Upadhyaya, N. D. Tripathi, J. D. Gaeddert, and J. H. Reed, “Open AI cellular (OAIC): An open source 5G O-RAN testbed for design and testing of AI-based RAN management algorithms,” IEEE Network, vol. 37, no. 5, pp. 7–15, 2023. [27] B. Tang, V. K. Shah, V. Marojevic, and J. H. Reed, “AI testing framework for next-g O-RAN networks: Requirements, design, and research opportunities,” IEEE Wireless Communications, vol. 30, no. 1, pp. 70–77, 2023.

[28] N. F. Cheng, T. Pamuklu, and M. Erol-Kantarci, “Reinforcement learning based resource allocation for network slices in O-RAN midhaul,” in Proc. IEEE 20th Consumer Communications & Networking Conference (CCNC), 2023, pp. 140–145. [29] R. Barker, A. E. Dorcheh, T. Seyfi, and F. Afghah, “REAL: Reinforcement learning-enabled xapps for experimental closed-loop optimization in O-RAN with OSC RIC and srsRAN,” in Proc. IEEE International Conference on Communications Workshops, 2025, pp. 389–395. [30] A. E. Dorcheh, T. Seyfi, and F. Afghah, “DORA: Dynamic ORAN resource allocation for multi-slice 5G networks,” arXiv preprint arXiv:2509.07242, 2025. [31] O. Sever, O. Salan, I. Hokelek, and A. Gorcin, “A practical demonstration of DRL-based dynamic resource allocation xapp using OpenAirInterface,” arXiv preprint arXiv:2501.05879, 2025. [32] P. Yan, J. Lu, H. Zeng, and Y. T. Hou, “Near-real-time resource slicing for QoS optimization in 5G O-RAN using deep reinforcement learning,” IEEE/ACM Transactions on Networking, vol. 34, pp. 1596–1611, 2026. [33] U. Sharma and Y. Liu, “Protocol-faithful hybrid traffic steering in open RAN via adaptive reinforcement learning,” Computer Networks, p. 112642, 2026. [34] J. Groen, Z. Yang, D. Muruganandham, M. Belgiovine, L. Ying, and K. R. Chowdhury, “From classification to optimization: Slicing and resource management with TRACTOR,” Computer Communications, vol. 257, p. 108625, 2026. [35] H. Zhou, M. Erol-Kantarci, and H. V. Poor, “Learning from peers: Deep transfer reinforcement learning for joint radio and cache resource allocation in 5G RAN slicing,” IEEE Transactions on Cognitive Communications and Networking, vol. 8, no. 4, pp. 1925–1941, 2022. [36] A. M. Nagib, H. Abou-Zeid, and H. S. Hassanein, “Safe and accelerated deep reinforcement learning-based O-RAN slicing: A hybrid transfer learning approach,” IEEE Journal on Selected Areas in Communications, vol. 42, no. 2, pp. 310–325, 2024. [37] S. Aleyadeh, I. Tamim, and A. Shami, “Transfer learning-accelerated network slice management for next generation services,” Computer Communications, vol. 228, p. 107937, 2024. [38] M. M. H. Qazzaz, S. A. R. Zaidi, A. A. Al-Hameed, A. Salama, and D. McLernon, “xapp empowered resource management for non-terrestrial users in 5G O-RAN networks,” IEEE Transactions on Machine Learning in Communications and Networking, vol. 4, pp. 794–812, 2026. [39] F. Lotfi and F. Afghah, “Meta reinforcement learning approach for adaptive resource optimization in O-RAN,” in Proc. IEEE Wireless Communications and Networking Conference, 2025, pp. 1–6. [40] B. Zeng and X. Niu, “Multi-task meta-initialized DQN for fast adaptation to unseen slicing tasks in O-RAN,” PLOS ONE, vol. 20, no. 10, p. e0330226, 2025. [41] F. Lotfi and F. Afghah, “Meta-hierarchical reinforcement learning for scalable resource management in O-RAN,” arXiv preprint arXiv:2512.13715, 2025. [42] F. Lotfi, H. Rajoli, and F. Afghah, “Task-specific sharpness-aware O-RAN resource management using multi-agent reinforcement learning,” IEEE Transactions on Machine Learning in Communications and Networking, vol. 4, pp. 98–113, 2026. [43] A. M. Nagib, H. Abou-Zeid, and H. S. Hassanein, “SafeSlice: Enabling SLA-compliant O-RAN slicing via safe deep reinforcement learning,” arXiv preprint arXiv:2503.12753, 2025. [44] N. M. Yungaicela, V. Sharma, and S. Scott-Hayward, “RSLAQ: A robust SLA-driven 6G O-RAN QoS xapp using deep reinforcement learning,” IEEE Transactions on Mobile Computing, vol. 25, no. 8, pp. 11 554– 11 570, 2026. [45] A. Tuerxun and A. Nakao, “Safe RAN slicing in O-RAN: Minimizing SLA violations via model-based reinforcement learning,” in Proc. IEEE 23rd Consumer Communications & Networking Conference (CCNC), 2026. [46] A. E. Dorcheh, T. Seyfi, R. Barker, and F. Afghah, “MORPH: Multienvironment orchestrated reinforcement learning for PRB handling in O-RAN,” arXiv preprint arXiv:2605.01128, 2026. [47] R. Mudi and H. Elbiaze, “Federated reinforcement learning-based resource allocation in O-RAN slicing for metaverse,” in Proc. IEEE International Conference on Communications (ICC), 2025. [48] M. Kouchaki, A. S. Abdalla, and V. Marojevic, “Federated neuroevolution O-RAN: Enhancing the robustness of deep reinforcement learning xapps,” arXiv preprint arXiv:2506.12812, 2025. [49] F. Lotfi, O. Semiari, and F. Afghah, “Evolutionary deep reinforcement learning for dynamic slice management in O-RAN,” in Proc. IEEE Globecom Workshops, 2022, pp. 227–232. [50] I. Vilà, O. Sallent, and J. Pérez-Romero, “On the design of a network digital twin for the radio access network in 5G and beyond,” Sensors, vol. 23, no. 3, p. 1197, 2023. [51] M. Polese, L. Bonati, S. D’Oro, P. Johari, D. Villa, S. Velumani, R. Gangula, M. Tsampazi, C. P. Robinson, G. Gemmi, A. Lacava, S. Maxenti, H. Cheng, and T. Melodia, “Colosseum: The open RAN digital twin,” arXiv preprint arXiv:2404.17317, 2024.

[52] F. W. Murti, S. Ali, and M. Latva-aho, “A bayesian framework of deep reinforcement learning for joint O-RAN/MEC orchestration,” IEEE Open Journal of the Communications Society, vol. 5, pp. 7685–7700, 2024. [53] F. Lotfi and F. Afghah, “Open RAN LSTM traffic prediction and slice management using deep reinforcement learning,” in Proc. 57th Asilomar Conference on Signals, Systems, and Computers, 2023, pp. 646–650. [54] F. Lotfi, H. Rajoli, and F. Afghah, “LLM-augmented deep reinforcement learning for dynamic O-RAN network slicing,” in Proc. IEEE International Conference on Communications, 2025, pp. 3827–3832. [55] ——, “Prompt-tuned LLM-augmented DRL for dynamic O-RAN network slicing,” arXiv preprint arXiv:2506.00574, 2025. [56] ——, “ORAN-GUIDE: RAG-driven prompt learning for LLM-augmented reinforcement learning in O-RAN network slicing,” arXiv preprint arXiv:2506.00576, 2025. [57] P. Gajjar and V. K. Shah, “ORANSight-2.0: Foundational LLMs for O-RAN,” IEEE Transactions on Machine Learning in Communications and Networking, vol. 3, pp. 903–920, 2025. [58] S. Alikhani, G. Charan, and A. Alkhateeb, “Lwm: A pre-trained wireless foundation model for universal feature extraction,” in 2025 IEEE International Conference on Machine Learning for Communication and Networking (ICMLCN). IEEE, 2025, pp. 1–6. [59] F. Zhou, C. Liu, H. Zhang, W. Wu, Q. Wu, T. Q. Quek, and C.-B. Chae, “Spectrumfm: A foundation model for intelligent spectrum management,” IEEE Journal on Selected Areas in Communications, 2025. [60] M. R. Uddin, T. Seyfi, and F. Afghah, “RFPrompt: Prompt-based expert adaptation of the large wireless model for modulation classification,” arXiv preprint arXiv:2605.03279, 2026. [61] D. Ha and J. Schmidhuber, “World models,” arXiv preprint arXiv:1803.10122, vol. 2, no. 3, p. 440, 2018. [62] F. Rezazadeh et al., “Agentic world modeling for 6G: Near-real-time generative state-space reasoning,” arXiv preprint arXiv:2511.02748, 2025. [63] Y. Zhou et al., “World-model-based adaptive network slicing,” in Proc. IEEE GLOBECOM, 2025. [64] H. Cheng, S. D’Oro, R. Gangula, S. Velumani, D. Villa, L. Bonati, M. Polese, G. Arrobo, C. Maciocco, and T. Melodia, “ORANSlice: An open-source 5G network slicing platform for O-RAN,” in Proc. 30th Annual International Conference on Mobile Computing and Networking (ACM MobiCom), 2024, pp. 1–6. [65] S. Maxenti, S. D’Oro, L. Bonati, M. Polese, A. Capone, and T. Melodia, “ScalO-RAN: Energy-aware network intelligence scaling in open RAN,” in Proc. IEEE INFOCOM, 2024. [66] G. R. Muns, P. S. Upadhyaya, U. Demir, N. Stephenson, N. Soltani, V. K. Shah, and K. R. Chowdhury, “SenseORAN: O-RAN-based radar detection in the CBRS band,” IEEE Journal on Selected Areas in Communications, vol. 42, no. 2, pp. 326–338, 2024. [67] N. Bouknana, M. Ahadi, F. Kaltenberger, and R. Schmidt, “An O-RAN framework for AI/ML-based localization with OpenAirInterface and FlexRIC,” in Proc. Wireless On-Demand Network Systems and Services Conference (WONS), 2026. [68] D. Villa, M. Belgiovine, N. Hedberg, M. Polese, C. Dick, and T. Melodia, “Programmable and GPU-accelerated edge inference for real-time ISAC on NVIDIA Aerial testbed,” arXiv preprint arXiv:2512.06493, 2026. [69] D. Xu, A. Zhou, G. Wang, H. Zhang, X. Li, J. Pei, and H. Ma, “Tutti: Coupling 5g RAN and mobile edge computing for latency-critical video analytics,” in Proceedings of the 28th Annual International Conference on Mobile Computing and Networking, ser. MobiCom ’22. Association for Computing Machinery, 2022, pp. 729–742. [70] J. Yi, G. Lee, M. Jeong, S. Shin, D. Kim, and Y. Lee, “Towards endto-end latency guarantee in MEC live video analytics with app–RAN mutual awareness,” in Proceedings of the 23rd Annual International Conference on Mobile Systems, Applications and Services, ser. MobiSys ’25. Association for Computing Machinery, 2025. [71] X. Zhang and D. Kim, “Enabling SLO-aware 5g multi-access edge computing with SMEC,” in Proceedings of the 23rd USENIX Symposium on Networked Systems Design and Implementation, ser. NSDI ’26. USENIX Association, 2026, pp. 1759–1776. [72] ETSI, “Multi-access edge computing (MEC); framework and reference architecture,” European Telecommunications Standards Institute, ETSI Group Specification GS MEC 003 V3.2.1, Apr. 2024. [73] 3GPP, “5g system enhancements for edge computing; stage 2,” 3rd Generation Partnership Project, 3GPP Technical Specification TS 23.548 V18.9.0, Apr. 2025. [74] J. Cheng, D. T. A. Nguyen, and D. T. Nguyen, “Robust dynamic edge service placement under spatio-temporal correlated demand uncertainty,” IEEE Transactions on Services Computing, vol. 18, no. 4, pp. 2372–2387, 2025. [75] A. Ebrahimi and F. Afghah, “Intelligent task offloading: Advanced MEC task offloading and resource management in 5g networks,” in 2025 IEEE Wireless Communications and Networking Conference (WCNC). IEEE, 2025, pp. 1–6.

[76] H. Li, X. Li, Q. Fan, Q. He, X. Wang, and V. C. M. Leung, “Distributed DNN inference with fine-grained model partitioning in mobile edge computing networks,” IEEE Transactions on Mobile Computing, 2024. [77] F. Strati, X. Ma, and A. Klimovic, “Orion: Interference-aware, finegrained GPU sharing for ML applications,” in Proceedings of the Nineteenth European Conference on Computer Systems, ser. EuroSys ’24. Association for Computing Machinery, 2024, pp. 1075–1092. [78] M. A. Habib, H. Zhou, P. E. Iturria-Rivera, M. Elsayed, M. Bavand, R. Gaigalas, Y. Ozcan, and M. Erol-Kantarci, “Intent-driven intelligent control and orchestration in O-RAN via hierarchical reinforcement learning,” in 2023 IEEE International Conference on Machine Learning and Applications (ICMLA). IEEE, 2023, pp. 55–61. [79] H. Navidan, M. Cheraghinia, J. Fontaine, M. Seif, E. D. Poorter, H. V. Poor, I. Moerman, and A. Shahid, “Toward autonomous O-RAN: A multi-scale agentic AI framework for real-time network control and management,” IEEE Network, 2026. [80] R. Barker, J. Boone, T. Seyfi, A. E. Dorcheh, C. P. Robinson, J. Boccuzzi, and F. Afghah, “Zero-trust in open RAN: A survey of identity continuity, mission-critical radio security, and learning-enabled control assurance,” 2026, submitted to IEEE Communications Surveys & Tutorials. [81] O-RAN ALLIANCE, “O-RAN ALLIANCE security update 2026,” 2026. [82] S. Luo, M. Garbelini, S. Chattopadhyay, and J. Zhou, “SNI5GECT: A practical approach to inject aNRchy into 5G NR,” in Proc. 34th USENIX Security Symposium, 2025, pp. 5385–5404. [83] J. Xing, S. Yoo, X. Foukas, D. Kim, and M. K. Reiter, “On the criticality of integrity protection in 5G fronthaul networks,” in Proc. 33rd USENIX Security Symposium, 2024, pp. 4463–4479. [84] S. Erni, M. Kotuliak, R. Baker, I. Martinovic, and S. Capkun, “GLaDoS: Location-aware denial-of-service of cellular networks,” in Proc. 34th USENIX Security Symposium, 2025, pp. 5307–5325. [85] M. Keshavarz, A. Shamsoshoara, F. Afghah, and J. Ashdown, “A realtime framework for trust monitoring in a network of unmanned aerial vehicles,” in IEEE INFOCOM 2020 - IEEE Conference on Computer Communications Workshops (INFOCOM WKSHPS), 2020, pp. 677–682. [86] J. Groen, S. D’Oro, U. Demir, L. Bonati, D. Villa, M. Polese, T. Melodia, and K. Chowdhury, “Securing O-RAN open interfaces,” IEEE Transactions on Mobile Computing, vol. 23, no. 12, pp. 11 265–11 277, 2024. [87] H. Jiang, H. Chang, S. Mukherjee, and J. Van der Merwe, “OZTrust: An O-RAN zero-trust security system,” in Proc. IEEE Conference on Network Function Virtualization and Software Defined Networks (NFVSDN), 2023, pp. 129–134. [88] T. O. Atalay, S. Maitra, D. Stojadinovic, A. Stavrou, and H. Wang, “Securing 5G OpenRAN with a scalable authorization framework for xapps,” in Proc. IEEE INFOCOM, 2023, pp. 1–10. [89] A. Lacava, S. Maxenti, L. Bonati, S. D’Oro, A. Oprea, T. Melodia, and F. Restuccia, “How to poison an xapp: Dissecting backdoor attacks to deep reinforcement learning in open radio access networks,” Computer Networks, vol. 273, p. 111727, 2025. [90] S. K. Bharti, X. Zhang, A. Singla, and X. Zhu, “Provable defense against backdoor policies in reinforcement learning,” in Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 14 704–14 714. [91] M. R. Uddin, F. Lotfi, T. Seyfi, and F. Afghah, “ORAN-DEFEND: Subspace detection and sanitization of backdoor DRL xapps in open RAN,” arXiv preprint arXiv:2607.06647, 2026, submitted to MECOM 2026. [92] M. Polese, L. Bonati, S. D’Oro, S. Basagni, and T. Melodia, “ColO-RAN: Developing machine learning-based xapps for open RAN closed-loop control on programmable experimental platforms,” IEEE Transactions on Mobile Computing, vol. 22, no. 10, pp. 5787–5800, 2023. [93] 3GPP, “NR; physical layer procedures for data,” 3rd Generation Partnership Project, Tech. Rep. TS 38.214 V18.6.0, Release 18, 2025. [94] R. Wiesmayr, L. Maggi, S. Cammerer, J. Hoydis, F. Aït Aoudia, and A. Keller, “SALAD: Self-adaptive link adaptation,” arXiv preprint arXiv:2510.05784v2, 2026. [95] S. Timilsina and M. Krunz, “Tampering attacks on CSI feedback reports in 5G new radio,” in Proc. IEEE Military Communications Conference (MILCOM), 2025, pp. 1005–1010. [96] R. Barker, J. Boone, T. Seyfi, A. Ebrahimi, E. Farquhar, B. Davis, J. Boccuzzi, and F. Afghah, “PRB starvation at slot time via rank indicator inflation: Attacking proportional fair memory in the scheduler,” 2026, submitted to ICEET 2026. [97] R. Barker, T. Seyfi, A. E. Dorcheh, J. Boone, F. Afghah, and J. Boccuzzi, “AtlasRAN: Timing-aware evaluation of open-source 5G platforms for integrated wireless testbeds,” in 2026 IEEE Military Communications Conference (MILCOM), Workshop 10 (WS10), 2026.

Record · ID 919299 · SHA-256 ad2aa91b37550605
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.