Conceptio › Archive › arXiv CS
arXiv CSopen access

AIIM: Adaptive Inter-cell Interference Mitigation for Heterogeneous Multi-vendor 5G O-RAN Networks

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
distributed-systemsinternetnetworkingprotocols
networking, internet, protocols, distributed systems

AIIM: Adaptive Inter-cell Interference Mitigation for Heterogeneous Multi-vendor 5G O-RAN Networks Samuel Reinders, Alireza Ebrahimi Dorcheh, Ryan Barker, Tolunay Seyfi, Fatemeh Afghah

arXiv:2605.01112v1 [cs.NI] 1 May 2026

Department of Electrical and Computer Engineering, Clemson University, Clemson SC USA {sreinde,alireze,rcbarke,tseyfi,fafghah}@clemson.edu Abstract—Inter-cell interference is a persistent issue in dense 5G deployments, especially in heterogeneous Open Radio Access Network (O-RAN) environments where coordination between base stations is limited. This paper presents AIIM, an adaptive inter-cell interference mitigation xApp for the O-RAN near-realtime RAN Intelligent Controller (near-RT RIC) that performs coordinated physical resource block (PRB) allocation across multiple base stations under diverse traffic demands and channel conditions. Unlike prior studies that rely primarily on simulation or fully hardware-centric testbeds, AIIM is developed and evaluated in a full-stack O-RAN system built on srsRAN, Open5GS, and O-RAN Software Community (ORAN-SC), and deployed on a hybrid experimental platform that simultaneously combines software defined radio (SDR)-based and virtual gNodeBs (gNBs) and user equipment (UEs). This design preserves realistic PHY-layer interactions while substantially improving scalability, reproducibility, and cost-effectiveness for multi-cell interference experiments. AIIM explicitly models overlapping PRB regions across neighboring cells and learns coordinated allocation policies that adapt to per-user QoS demand and pathloss variation across the network. Experimental results show that AIIM improves QoS satisfaction and reduces interference-induced PRB loss relative to proportionalfair scheduling baselines while maintaining comparable aggregate network throughput. These results demonstrate the promise of scalable, learning-driven O-RAN control for practical interference management in heterogeneous multi-gNB 5G networks.1 Index Terms—5G, O-RAN, Inter-Cell Interference, Reinforcement Learning, Resource Allocation, xApp.

I. Introduction Fifth-generation (5G) cellular networks support various services with diverse QoS requirements, while dense deployments are increasingly used to meet rising traffic demands [1]. In such environments, users experience highly variable channel conditions and interference depending on their location relative to neighboring base stations. These effects become more pronounced in complex heterogeneous deployments like multi-vendor neutral host networks (NHNs) and private 5G networks that commonly use shared or unlicensed spectrum, such as the Citizens Broadband Radio Service (CBRS) band [2], [3]. In these settings, competition or privacy concerns restrict intercell coordination and scheduling, increasing resource contention and inter-cell interference. For instance, without coordination, neighboring base stations may operate on the same carrier frequencies and independently schedule This material is based upon work supported by the National Science Foundation under Grant Numbers CNS-2202972, CNS2318726, and CNS-2232048. 1 A video demonstration of the running system can be found at https://github.com/sireinders/AIIM-Multi-gNB-Interference.git.

the same physical resource block (PRBs) during the same transmission interval [1]. Although Orthogonal FrequencyDivision Multiple Access (OFDMA) preserves orthogonality within a cell, it does not prevent co-channel interference (CCI) across cells. This is particularly harmful for users near cell edges, where competing signals from adjacent base stations (gNBs) can significantly reduce throughput, increase block error rate (BLER), and violate quality of service (QoS) requirements. These challenges motivate interference mitigation mechanisms that improve spectrum utilization while enabling fair coordination across complex deployments [1]. The Open Radio Access Network (O-RAN) framework provides a promising foundation for addressing these challenges in multi-vendor and heterogeneous deployments. ORAN disaggregates traditional RAN functions and implements them in software, enabling more flexible network control under changing wireless conditions. By defining open and standardized interfaces, O-RAN enables the integration of intelligent control mechanisms through RAN Intelligent Controllers (RICs), allowing artificial intelligence (AI) and machine learning (ML) workflows to be embedded directly into the network control loop [4]. The near-Real Time RIC (near-RT RIC) plays a central role in this architecture, providing control and optimization capabilities on timescales ranging from 10 ms to 1 s. Through the deployment of third-party xApps managed by the near-RT RIC, fine-grained and adaptive control of RAN behavior across base stations from different vendors is possible [5]. These capabilities make the near-RT RIC a suitable platform for implementing learning-based interference mitigation strategies. Reinforcement learning (RL) has emerged as an effective paradigm for learning complex control policies in 5G and beyond networks, including slice-aware resource allocation, dynamic spectrum sharing, and RAN-level optimization [6]–[8]. In RL-based approaches, agents observe key performance metrics (KPMs) from the wireless environment and receive feedback based on their actions in the form of rewards, enabling them to iteratively learn policies that optimize long-term performance objectives. Metrics such as aggregate throughput, spectral efficiency, transmission queue length, and block error rate (BLER) provide valuable insight into overall network behavior and serve as informative reward signals for learning-based interference mitigation. Existing work on interference mitigation often emphasizes either simulation-scale evaluation or hardware realism, but rarely both within a deployable O-RAN control loop. This paper presents, a near-RT RIC xApp for coordinated cross-cell PRB allocation in heterogeneous 5G ORAN networks, and validates it on a hybrid platform that

combines software defined radio (SDR)-based and virtual gNBs and user equipment (UEs). By explicitly modeling shared PRB regions and adapting allocation decisions to user demand and pathloss, the proposed system enables scalable yet realistic experimental evaluation of inter-cell interference mitigation in multi-gNB deployments. The main contributions of this paper are as follows: • We design AIIM, a near-RT RIC xApp for coordinated inter-cell interference mitigation through cross-cell PRB allocation in heterogeneous 5G O-RAN networks. • We develop a hybrid experimental platform that combines SDR-based and virtual gNBs and UEs, enabling a practical balance between PHY-layer realism, scalability, and reproducibility. • We formulate inter-cell interference mitigation as a learning-based control problem over shared PRB regions using user demand and pathloss as state information. • We implement the full-stack system using srsRAN, Open5GS, and ORAN-SC, and experimentally compare AIIM against proportional-fair and learning-based baselines. II. Related Work Interference management has been a core research topic throughout the evolution of cellular networks. Higher user numbers have necessitated spatial and spectral re-use resulting in dense network deployments that, in certain cases, sacrifice network performance when neighboring cells unknowingly contend for resources. For instance, CCI occurs when adjacent cells reuse the same frequency resources simultaneously [9]. A common approach to this challenge focuses on equitable inter-cell PRB management. By defining a high-level policy that is aware of the network structure, PRBs can be allocated to cells in a fixed manner through frequency planning, or dynamically, using communication over the network backhaul such as enhanced inter-cell interference coordination [10] and coordinated multi-point (CoMP) transmission [11]. Although these techniques can reduce interference, they either sacrifice robustness or require strict synchronization and substantial backhaul signaling, making them less ideal for real-time PRB control [1], [12]. As an alternative, RL-based resource allocation algorithms have become popular for real-time, dynamic optimization of 5G O-RAN networks. REAL [6] and DORA [13] demonstrate how RL-enabled xApps can be deployed on an open-source near-RT RIC to perform closed-loop slice-level PRB scheduling for meeting diverse user QoS requirements. Similarly, ORANSlice [14] provides an opensource platform for slice-level resource management within O-RAN architectures. Several recent works apply these learning-based approaches to interference mitigation. The ChARM framework [7] demonstrates the feasibility of real-time, datadriven cross-technology spectrum sharing with a deep neural network (DNN) model that leverages O-RAN control loops to reconfigure network parameters and avoid interference. The xDiff model [15] introduces a diffusionbased RL approach for collaborative inter-cell interference mitigation (ICIM) by manipulating the per-user proportional fair (PF) metrics used by the resource schedulers of each cell. Evaluated on an over-the-air (OTA) testbed consisting of a single centralized unit (CU), three

SDR distributed units (DUs), and ten commercial smartphones, xDiff outperforms state-of-the-art ICIM techniques by minimizing delay for users while achieving comparable throughput and BLER. Conversely, Qualcomm’s Interference-Aware Intelligent Scheduling (IAIS) framework [16] uses an offline-trained Long Short-Term Memory (LSTM) model to drive intelligent link adaptation and modulation and coding scheme (MCS) selection to minimize interference within an OTA industrial Internet-ofThings (IIoT) factory automation testbed. MLCIMO [17] proposes an ML-based classification model for nested macro/micro-cell deployments that characterizes interference levels and drives user-offloading to maximize network performance in a fully simulated environment. Existing interference research and 5G O-RAN testbeds each have a critical limitation. First, to the best of our knowledge, there are no learning-based ICIM algorithms investigating multi-vendor 5G O-RAN deployments with different user service profiles like Enhanced Mobile Broadband (eMBB) and Massive Machine Type Communications (mMTC). Secondly, existing 5G O-RAN testbeds typically assume a single device-type in their network architectures. Namely, industry-leading large-scale wireless experimental platforms such as Colosseum [18], AERPAW [19], and OAIC [20] are limited to software-only or purely SDR-based architectures at one time. This limits the scalability of realistic wireless experiments to the size of the hardware testbed. Further, access to these platforms are limited and hardware resources are shared, making it difficult to conduct comprehensive experimentation in a timely manner. Fully simulated environments have been proposed to alleviate accessibility and scaling concerns; however, they neglect or omit important PHY-layer effects and protocol timing constraints that are required for properly emulating complex cellular deployments. Mixed hardware/software device architectures offer a practical compromise that allow realistic, low-cost, reproducible experimentation within standard laboratory environments while scaling user density and traffic diversity. In this work, we propose an adaptive interference mitigation algorithm implemented within a multi-vendor ORAN environment consisting of simultaneous SDR-based and virtual base stations and user devices with diverse QoS profiles. By combining real RF hardware with scalable virtual users, the proposed architecture enables costeffective experimentation while capturing realistic PHYlayer hardware behavior. This design allows RL-based resource allocation and ICIM algorithms to be evaluated under diverse user demand and channel scenarios. III. System Model A. Interference Modeling In this section, we describe the system model and the assumptions adopted in this paper for interference analysis. We consider an O-RAN network composed of three vendor’s base stations, each with their own radio unit (RU), distributed unit (DU), and centralized unit (CU). The DUs and CUs connect to a shared NRT-RIC via the E2 interface for metric reporting and network control. The open fronthaul interface (FH) and open midhaul interface (F1) connect the RU to DU and DU to CU, respectively. The virtual environment layout is illustrated in Fig. 1. Following the New-Radio (NR) standard, we define 52

B. Interference-Aware PRB Allocation RL Agent The inter-cell interference mitigation (ICIM) xApp employs an offline-trained proximal policy optimization (PPO) agent comprised of separate actor and critic networks, each having two fully-connected layers of 64 neurons. PPO uses on-policy learning to sample actions and actively refine its policy. PPO was selected due to its stable policy updates and strong performance in complex, dynamic environments. Finalized hyperparameters for the PPO model were determined through a parameter sweep of learning rate, gamma, ent coeff, n steps, batch size, and n epochs, in which we selected the parameters that achieved the strongest performance on validation scenarios without overfitting. The PPO and DQN model parameters are summarized in Table I. Fig. 1. Virtual multi-vendor system model for per-cell PRB allocation including user slices and RU relative locations.

Fig. 2. Frequency-domain PRB distribution for gNB2 (blue, virtual), gNB1 (pink, SDR-based), and gNB3 (orange, virtual).

PRBs per gNB, 10 MHz channel bandwidth, and 15 kHz subcarrier spacing in FR1. Vendor A’s gNB supports an eMBB slice, while Vendors B and C support two slices, eMBB and mMTC. For analytical tractability, we consider inter-cell interference between users connected to different base stations based on overlapping frequency PRB assignments. If two separate base stations allocate the same frequency PRB to different users, the entire downlink bandwidth corresponding to that PRB is lost for both users. In our experiments, we do not consider partial-PRB overlap. The frequency-domain PRB overlap between the gNBs is defined in Fig. 2. There are a total of 156 PRBs, of which, there are 116 unique frequency PRBs. The impact of interference is realized within the system through the following process. First, the algorithm determines peruser PRB allocations based on their respective allocation strategies defined in Section IV. We assume that each gNB allocates their private (non-overlapping) PRBs to users before their shared (overlapping) PRBs. Therefore, the number of shared PRBs allocated by a single gNB is equal to the total PRBs it allocated subtracted by its private PRBs. The PRBs lost to interference (PRBL ) are calculated using equation (1) for each shared PRB region the gNB has PRBL = max(0, (PRBS1 + PRBSX ) − PRBO )

(1)

where PRBS1 and PRBSX represent the number of shared PRBs allocated by gNB1 and its neighbor (gNB2 or gNB3), respectively, and PRBO denotes the total available PRBs in that specific overlapping frequency region. This value is then subtracted from the original PRB allocations to reflect effective throughput degradation for the users.

TABLE I Model Parameters Parameter 2 learning rate gamma ent coef n steps batch size n epochs exploration fraction exploration initial eps exploration final eps

Value 3e-4 0.99 0.01 512 16 3 0.1 1.0 0.05

PPO ✓ ✓ ✓ ✓ ✓ ✓ χ χ χ

DQN ✓ ✓ χ χ ✓ χ ✓ ✓ ✓

The RL state space S consists of concatenated user state vectors st = (cuser , λt , Lt ) at time step t, where the components are defined as: • cuser ∈ C: QoS category, where C = {ceMBB , cmMTC }. • λt ∈ R≥0 : Throughput demand in Bytes per second (Bps). • Lt ∈ R: Channel pathloss in decibels (dB). The complete state space is thus defined by the Cartesian product S = C × R≥0 × R. During preliminary testing, this lightweight state design enabled faster policy convergence and better reward performance than those that included historic demand or resource allocations of users. We hypothesize that the agent struggles to learn high-level network trends when overwhelmed with dense, highlydependent per-user metrics. The action space A defines a discrete PRB splits for each shared PRB region. This is done to enable dynamic and collaborative bandwidth sharing between base stations without redundant backhaul signaling. Within a single shared region, a lower-priority gNB selects the number of PRBs it will defer to usage by a higher priority gNB at ∈ A, where A = {0, 1, 2, ..., K}. K is the maximum number of overlapping PRBs. By selecting an action, the RL agent determines each cell’s share of the total PRBs. After this, individual base stations distribute their PRBs proportionally between users based on their demand. In our system, gNB1 is configured with the highest priority. This approach prevents action space explosion inherent to fine-grained per-user PRB assignments as user counts scale, while maintaining flexibility to network conditions. The reward function is designed in such a way for the agent to generate an effective policy that dynamically adjusts resource allocations across multiple cells as to 1 2 Remaining parameters set to StableBaseline3 default values.

satisfy per-user QoS demands and 1 achieve efficient spectrum utilization amidst inter-cell interference conditions. Within the environment, let N denote the set of all cells and U represent the set of all UEs. An overlapping PRB region between two cells A and B is denoted as OA,B with the set of all overlapping regions written as O. For UE i ∈ U , we use the clipped ratio of average achieved throughput λA to requested throughput λR in Bytes to represent the per-user satisfaction reward as rs (t): rs (t) = min(

λA , 1) (λR + ϵ)

(2)

where ϵ = 1 × 10−6 . This ratio-based throughput reward design consistently outperformed preliminary methods that used a sigmoid function or PRB-based rewards which caused the agent to settle into sub-optimal policies due to saturating rewards or coarse reward signals. Our approach discourages the agent from resource over-allocation, while accounting for the wireless channel and providing a smooth gradient for learning. During our initial testing, we found that relying exclusively on user-centric rewards induced myopic greedy policies. Instead of a globally equitable strategy, cells with fewer users or higher interference were starved of shared resources entirely. To combat this, we incorporated a spectrum utilization reward ru (t) for each OA,B that can be written as: n∗ϕ∗α ru (t) = (3) χ where n is the total number of users within the network, ϕ is a configurable fairness parameter, χ is the size of OA,B in PRBs, and α is the number of shared PRBs allocated to the higher priority gNB. We select ϕ to be 0.05 in our experiments, after values of 0.1, 0.025 and 0.01 produced 13%, 3%, and 7% higher QoS violations in evaluation scenarios. With this utilization reward, the agent learned to better share resources between vendor’s cells and avoid selfish behaviors while maintaining overall network performance in highly dynamic user scenarios. The total reward for the agent across all users and overlapping PRB regions is shown as: X X r(t) = rs (t) + ru (t) (4) i∈U

j∈O

C. Training Environment and Traffic Model The reinforcement learning agent was trained offline in a Gymnasium-based environment designed to emulate the multi-cell interference scenarios observed in the testbed. Training was performed over 10,000 episodes with 100 steps per episode. At the beginning of each episode, user positions and bearing are randomly initialized within the coverage area of their serving base station. User mobility follows the 3GPP Extended Vehicular A (EVA) channel model with a velocity of 10 m/s [21]. User demand profiles are randomly sampled from ranges corresponding to realistic eMBB and mMTC service classes for every step. Table II summarizes the traffic profiles used during training. Demand ranges were selected based on achievable downlink throughput observed on the experimental platform while producing a comprehensive range of interference conditions (none - high) and maintaining UE connectivity. Minimum connectivity requirements

TABLE II Service Profiles Parameter SDR eMBB (UE0) VIR eMBB (UE1,3) VIR mMTC (UE2,4)

Demand Range 1.5MB - 3.5MB 907kB - 3.1MB 37.5kB - 62.5kB

Frequency PRB (1/s) Range 1 20 - 49 1 12 - 44 4 5

were determined empirically during preliminary testing. A minimum allocation of five PRBs was required to maintain stable UE connectivity, corresponding to a minimum PRB allocation ratio of 10% within the RIC resource control policy. IV. Experimental Evaluation A. Testbed Architecture The experimental platform integrates the srsRAN 5G stack, the Open5GS core network, and the O-RAN Software Community (OSC) RIC to create a hybrid multigNB and multi-UE environment for evaluating PRB resource allocation under multi-cell co-channel interference. The architecture combines SDR-based and virtual network elements to enable scalable experimentation while maintaining realistic PHY-layer interactions. The system consists of the following components: • Open5GS Core: Provides 5G core network functions for UE registration (AMF), session management (SMF), and user-plane forwarding (UPF). • ORAN-SC Near-RT RIC: Hosts the ICIM xApp responsible for coordinating PRB allocation decisions across multiple gNBs via the E2 interface. • srsRAN gNB: Implements a monolithic gNB architecture with RU, CU, and DU functionality. Both SDRbased and virtual ZeroMQ (ZMQ)-based gNB instances are deployed. • srsRAN UE: Provides both SDR-based and virtual user equipment implementations supporting LTE and 5G Non-Stand Alone (NSA) connectivity. • GNU Radio Blocks: Used to multiplex and route I/Q streams between virtual users and gNB instances using ZMQ sockets. The testbed was composed of two physical hosts (Host 1 and Host 2), each connected to an Ettus B210 SDR. Host 1 executed the Open5GS core network and ORAN-SC NearRT RIC within Docker containers, while the SDR-based base station (gNB1), two virtual base stations (gNB2 and gNB3), and four virtual srsUEs instances (UE1-UE4) were executed as processes on the same host. Host 2 executed a single SDR-based UE process (UE0) connected to its local B210 SDR. GNU Radio flowgraphs were used to route and multiplex I/Q samples for the virtual UEs, enabling successful registration with the Access and AMF, and Protocol Data Unit (PDU) session establishment through the Session Management Function (SMF). The trained ICIM agent was deployed as an xApp running on the ORAN-SC Near-RT RIC integrated with the srsRAN stack. Downlink traffic was generated using iPerf3 and performance metrics were reported to the Near-RT RIC via the E2 interface. A high-level overview of the testbed architecture is shown in Fig. 3.

Non-Slice Aware PF (NSA-PF): Each DU independently allocates PRBs without inter-cell coordination. • Slice Aware Constant-Allocation PF (SA-CAPF): Non-overlapping PRBs are statically partitioned among gNBs while slice-level demand is aggregated across the network. PRBs are evenly distributed among users belonging to the same slice. • Slice-Aware Variable-Allocation PF (SA-VAPF): Non-overlapping PRBs are statically partitioned among gNBs with PRBs assigned according to the standard rate-to-average-rate ratio utility while respecting slice-level PRB guardrails (Nsmin , Nsmax ). • Deep Q-Learning Network (DQN): An off-policy value-based reinforcement learning algorithm that replaces PPO and uses a 1-dimensional action space. •

Fig. 3. Testbed architecture with hardware connections and software processes.

B. Experimental Configuration The experimental testbed operates in the n3 frequency band, which uses frequency division duplexing (FDD). The downlink and uplink center frequencies were configured at 1842.5 MHz and 1747.5 MHz, respectively. Due to limitations in GNU Radio-based multi-UE emulation with srsRAN, direct RF-level pathloss modeling is not feasible. Instead, large-scale fading effects were emulated using a calibrated throughput scaling model based on a reference pathloss value of PLref = 83.3dB. The calibrated throughput λcal is defined as:    λcal = λach · max gmin , min 1, e−κ(PLUE −PLref,dB ) (5) λach represents the throughput achieved on the testbed before pathloss scaling, gmin = 0.5 ensures a minimum scaling factor to maintain UE connectivity, κ = 5.5 × 10−2 is a decay factor controlling signal attenuation, and PLUE represents the pathloss calculated from the UE position and downlink channel frequency. C. Evaluation Methodology To evaluate performance, we compared our model against four baselines using two performance metrics: • Total QoS violations across all users • Total PRBs lost due to inter-cell interference QoS violations occur when a user’s achieved throughput is less than it requested. These metrics jointly capture user-level performance and overall spectral efficiency of the network. Evaluation was conducted over 5 episodes consisting of 70 environment steps each, for a total of 350 evaluation steps. While this sample size is constrained by the realtime execution requirements of the hardware-in-the-loop (HITL) testbed, the randomized initialization of UE positions and traffic demands in each episode ensures that the results capture a statistically diverse range of interference scenarios. During each step, downlink traffic was generated using iPerf3 for 15 seconds. Achieved throughput values were reported to the Near-RT RIC via the E2 interface and computed as the average of 10 throughput samples per user before pathloss scaling. User and demand scenarios remained consistent between evaluated models for a fair comparison. Each model was evaluated independently in continuous testing blocks. D. Baseline Algorithms The proposed PPO-based interference mitigation model was compared against four baseline approaches:

E. Evaluation Results The resulting PRBs lost to interference and user QoS violations across all the evaluation scenarios are displayed in Figs. 4(a) and 4(b). As expected, the NSA-PF algorithm is the only baseline that experiences PRB losses due to inter-cell interference. This happens because NSA-PF allows overlapping PRB usage across neighboring cells without coordination. In contrast, the other approaches prevent interference by statically or dynamically partitioning overlapping spectrum regions. The maximum effective PRBs loss in the system is 80 PRBs, which occurs when all three gNBs simultaneously allocate 52 PRBs and maximize over-utilization of the shared spectrum regions. During evaluation, NSA-PF experienced an average number of 7.08 lost PRBs with a standard deviation of 11.32. This variation across steps is caused by the diverse combinations of traffic profiles and user positions, which produce different interference patterns. Since each PRB equates to approximately 550 kbps of achievable throughput when MCS is at its maximum, losing seven PRBs results in an average throughput degradation of approximately 3.85 Mbps or roughly 6% of the maximum aggregate network throughput. This result clearly illustrates the performance impact of interference and the need for effective mitigation strategies in dynamic networks. While SA-VA-PF and SA-CA-PF effectively eliminate interference by partitioning spectrum resources between neighboring cells, they remain inflexible to changing user demands and channel conditions. The PPO and DQN models provide more robust solutions to interference by interpreting path-loss and user demand to dynamically adjust resource allocation decisions. This leads to better QoS satisfaction in complex multi-cell environments. Specifically, the PPO and DQN models reduce total QoS violations by 7.6% and 4.7%, respectively, when compared to the best-performing non-learning baseline algorithm. Table III summarizes the average aggregate network throughput achieved by each model, which are similar across evaluation scenarios. PPO and DQN models achieve TABLE III Average Aggregate Network Throughput Model NSA-PF Tput 42.69 [Mbps]

SA-CA-PF

SA-VA-PF

DQN

PPO

43.92

52.72

50.31

50.18

(a) Total PRBs lost to inter-cell interference

(b) Total user QoS violations

Fig. 4. Performance metrics across PPO, DQN, SA-VA-PF, SA-CA-PF, and NSA-PF.

slightly lower aggregate throughput than SA-VA-PF because they allocate additional PRBs to users experiencing higher pathloss in order to satisfy QoS requirements. This behavior improves fairness across users by ensuring reliable service even for users located near cell edges or experiencing unfavorable channel conditions. Overall, these results indicate that learning-based interference mitigation strategies outperform static or semistatic rule-based approaches in dynamic multi-cell environments. PPO achieves the strongest performance due to its stable policy updates and improved exploration capabilities in stochastic environments, enabled by clipped policy gradients and on-policy optimization. V. Conclusion In this work, we proposed AIIM, an offline-trained ICIM xApp for heterogeneous multi-vendor 5G O-RAN networks. AIIM leverages reinforcement learning to determine per-cell resource allocation actions that minimize user QoS violations and avoid losing PRBs to interference, while maintaining near-optimal aggregate network throughput. AIIM was implemented on a mixed 5G testbed composed of virtual and hardware base stations and users, confirming its applicability toward real-world deployments. This enables future work that can leverage the high scalability of such systems while balancing implementation costs. Experimental results indicate that learning-based algorithms like AIIM’s PPO model as well as the DQN baseline model readily outperform traditional static and semi-static rule-based interference mitigation techniques across unique user placement and demand scenarios. These findings highlight the potential of the O-RAN framework to support scalable interference mitigation solutions for heterogeneous multi-vendor networks. References [1] Y. Zhou et al., “An overview on intercell interference management in mobile cellular networks: From 2G to 5G,” in 2014 IEEE International Conference on Communication Systems, 2014, pp. 217–221. [2] M. I. Rochman et al., “Neutral-hosts in the shared mid-bands: Addressing indoor cellular performance,” in 2025 IEEE International Symposium on Dynamic Spectrum Access Networks (DySPAN). IEEE, May 2025, p. 1–9. [Online]. Available: http://dx.doi.org/10.1109/DySPAN64764.2025.11115898 [3] A. Tusha et al., “A comprehensive analysis of secondary coexistence in a real-world cbrs deployment,” in 2024 IEEE International Symposium on Dynamic Spectrum Access Networks (DySPAN), 2024, pp. 79–87. [4] M. Polese et al., “Understanding O-RAN: Architecture, interfaces, algorithms, security, and research challenges,” IEEE Communications Surveys & Tutorials, vol. 25, no. 2, pp. 1376– 1411, 2023. [5] A. S. Abdalla et al., “Toward next generation open radio access networks: What O-RAN can and cannot do!” IEEE Network, vol. 36, no. 6, pp. 206–213, 2022.

[6] R. Barker et al., “REAL: Reinforcement learning-enabled xApps for experimental closed-loop optimization in O-RAN with OSC RIC and srsRAN,” in 2025 IEEE International Conference on Communications Workshops (ICC Workshops). IEEE, Jun. 2025, p. 389–395. [Online]. Available: http: //dx.doi.org/10.1109/ICCWorkshops67674.2025.11162144 [7] L. Baldesi, F. Restuccia, and T. Melodia, “ChARM: NextG spectrum sharing through data-driven real-time O-RAN dynamic control,” in IEEE INFOCOM 2022 - IEEE Conference on Computer Communications, 2022, pp. 240–249. [8] F. Lotfi, H. Rajoli, and F. Afghah, “Prompt-tuned LLMaugmented DRL for dynamic O-RAN network slicing,” 2025. [Online]. Available: https://arxiv.org/abs/2506.00574 [9] F. Qamar et al., “Interference management issues for the future 5G network: A review - telecommunication systems,” May 2019. [Online]. Available: https://link.springer.com/article/10. 1007/s11235-019-00578-4 [10] S. Deb et al., “Algorithms for enhanced inter-cell interference coordination (eICIC) in LTE HetNets,” IEEE/ACM Transactions on Networking, vol. 22, no. 1, pp. 137–150, 2014. [11] D. Lee et al., “Coordinated multipoint transmission and reception in LTE-advanced: deployment scenarios and operational challenges,” IEEE Communications Magazine, vol. 50, no. 2, pp. 148–155, 2012. [12] Q. Cui et al., “Evolution of limited-feedback CoMP systems from 4G to 5G: CoMP features and limited-feedback approaches,” IEEE Vehicular Technology Magazine, vol. 9, no. 3, pp. 94–103, 2014. [13] A. E. Dorcheh, T. Seyfi, and F. Afghah, “DORA: Dynamic O-RAN resource allocation for multi-slice 5G networks,” 2025. [Online]. Available: https://arxiv.org/abs/2509.07242 [14] H. Cheng et al., “ORANSlice: An open-source 5G network slicing platform for O-RAN,” 2024. [Online]. Available: https: //arxiv.org/abs/2410.12978 [15] P. Yan, H. Zeng, and Y. T. Hou, “xDiff: Online diffusion model for collaborative inter-cell interference management in 5G ORAN,” IEEE Transactions on Networking, vol. 34, pp. 1363– 1376, 2026. [16] B. Akgun et al., “Interference-aware intelligent scheduling for virtualized private 5G networks,” IEEE Access, vol. PP, pp. 1– 1, 01 2024. [17] D. Anand, M. A. Togou, and G.-M. Muntean, “A machine learning-based approach for interference mitigation to enhance QoS and QoE in 5G O-RAN networks,” in 2024 IEEE 35th International Symposium on Personal, Indoor and Mobile Radio Communications (PIMRC), 2024, pp. 1–6. [18] L. Bonati et al., “Colosseum: Large-scale wireless experimentation through hardware-in-the-loop network emulation,” in 2021 IEEE International Symposium on Dynamic Spectrum Access Networks (DySPAN), 2021, pp. 113–122. [19] A. Panicker et al., “AERPAW emulation overview and preliminary performance evaluation,” Elsevier Computer Communications, vol. 194, pp. 158–169, 2021. [20] H. Zhang et al., “Open AI Cellular (OAIC): An open source 5G O-RAN testbed for design and testing of AI-based RAN management algorithms,” IEEE Network, vol. 37, no. 6, pp. 102–109, 2023. [21] 3rd Generation Partnership Project (3GPP), “Study on Scenarios and Requirements for Next Generation Access Technologies,” 3rd Generation Partnership Project (3GPP), Technical Report (TR) 38.913, 2020, release 16. [Online]. Available: https://www.3gpp.org/ftp/Specs/archive/38 series/38.913/

Record · ID 155224 · SHA-256 dfacee6dce1b5fbc
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.