ConceptioArchivearXiv CS
arXiv CSopen access

Explanation-Guided Federated Deep Reinforcement Learning for Joint Resource Allocation and Scheduling in 6G in-X Subnetworks

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
distributed-systemsinternetnetworkingprotocols
networking, internet, protocols, distributed systems

Explanation-Guided Federated Deep Reinforcement Learning for Joint Resource Allocation and Scheduling in 6G in-X Subnetworks Ramoni O. Adeogun

arXiv:2609.24102v1 [eess.SP] 21 Sep 2026

Department of Electronic Systems, Aalborg University, Denmark Email: [email protected] Abstract—Sixth-generation (6G) wireless systems are envisioned as networks of networks, integrating diverse in-X subnetworks that provide localized, high-performance connectivity. Ensuring reliable communication in dense deployments, such as industrial robots and vehicles, is challenging due to dynamic interference and strict performance requirements. Traditional radio resource management (RRM) methods have limitations, prompting the need for AI-based solutions. In this paper, we address the challenges of dynamic resource allocation and scheduling in 6G in-X subnetworks supporting applications with heterogeneous characteristics by proposing a novel framework that combines Multi-Agent Reinforcement Learning (MARL), Federated Learning (FL), and Explainable AI (XAI). Our solution is designed to improve the reliability, robustness, and transparency of resource management and intra-subnetwork scheduling while ensuring data privacy and fairness across multiple co-existing subnetworks. Unlike existing works, our approach considers a realistic scenario with multiple devices per subnetwork, thereby offering a more comprehensive and scalable solution. The proposed explainable RL framework enables agents to collaboratively optimize channel allocation and scheduling without the need to share raw data, preserving the privacy of each participating subnetwork. Extensive simulations based on 3GPP scenarios demonstrate the effectiveness of our approach, showing significant improvements in performance, and transparency over existing solutions. Index Terms—6G, resource allocation, RL, FL, subnetworks, XAI

I. I NTRODUCTION Sixth generation (6G) wireless systems are expected to operate as networks of networks integrating subnetworks with heterogeneous types of supported services, varying complexity, coverage, and operating spectrum [1]. 6G in-X subnetworks [2], [3] are short-range, low-power cells envisioned to provide highly localized, high-performance wireless connectivity at the edge of the 6G network of networks. These subnetworks are intended to replace wired infrastructure for communication services with data rate, latency, and reliability requirements that existing wireless technologies cannot support in entities such as robots, production modules, vehicles, and classrooms. Subnetworks are expected to support diverse applications with demanding communication requirements. For example, subnetworks installed within an industrial robot can support highspeed, closed-loop robot control applications with a fraction of a millisecond control cycle time and reliability above

99.9999%. Also, classroom subnetworks can enable immersive augmented reality (XR)-based education activities [1]. The critical nature of these applications supported by 6G subnetworks underscores the importance of ensuring that the required communication requirements are consistently met, regardless of the environment. This becomes especially crucial in dense scenarios, such as subnetworks deployed inside vehicles at a busy intersection, where the potential for high levels of interference could lead to disruption of communication. Consequently, techniques for managing dynamic interference via optimization of radio resource utilization are essential to guarantee the seamless operation of subnetworks in such dense scenarios [4]. This has indeed been the focus of active research within the last few years, see e.g., [5]–[7]. While classical heuristic algorithms for resource management such as distributed greedy selection, centralized graph coloring, minimum SINR guarantee, and sequential iterative subband allocation (SISA) [8] have been studied for subnetworks, there appears to be a consensus that AI-based methods are essential to cope with the requirements for 6G. Since dynamic radio resource management can be modeled as a sequential decision problem, reinforcement learning (RL) based solutions for 6G RRM appear quite popular, see e.g., [9]–[11]. In the context of 6G in-X subnetworks, the focus has been on Multi-agent RL (MARL) [4], [5], [12]. Despite the potential of AI-based resource allocation in existing works, the lack of transparency and trust in the decisions made by an AI model (or RL agent) represents a major bottleneck to the deployment of such solutions in practical systems. Also, privacy and fairness are important to prevent AI models from being discriminatory or biased [13]. This is particularly relevant for resource allocation in scenarios where multiple independent 6G in-X subnetworks co-exist over a limited set of shared resources. To overcome these challenges, we conjecture that a combination of MARL [5], federated learning [14], and explainable AI (XAI) techniques [15] can guarantee the reliability, robustness, and transparency of dynamic channel allocation decisions in 6G in-X subnetworks without while preserving the privacy and performance of the system. While MARL provides the framework for solving the dynamic resource allocation problem, FL ensures that MARL agents can be trained collaboratively without explicitly sharing raw measurements between the participating nodes and a central server translating to a guarantee

of the privacy of the nodes. On the other hand, XAI provides a set of procedures and methods to reveal the inner operations of AI algorithms and their potential strengths, weaknesses, and behaviors. This paper addresses three key challenges in dense 6G in-X subnetworks: (i) joint resource allocation and scheduling, (ii) privacy-preserving distributed learning, and (iii) explainable decision-making. The main contributions are as follows: • We propose a novel AI-based solution for joint resource allocation and scheduling in dense deployments of 6G inX subnetworks based on a combination of MARL, FL, and XAI. To the best of our knowledge, this is the first work on joint resource allocation and scheduling in the context of 6G in-X subnetworks. Unlike existing works on RRM for subnetworks, e.g., [6], [12] which make the unrealistic assumption of a single device per subnetwork, this paper considers a more general and realistic set up with multiple devices in each subnetwork. • We design an explainable RL framework for training agents to perform joint channel allocation and intrasubnetwork scheduling. • We perform extensive simulations using 3GPP channel models and benchmark performance of the proposed scheme with state-of-the-art solutions.

where Yx is the value of a two-dimensional Gaussian random field at the device’s or AP’s location, and dcorr is the correlation distance. We assume that transmission within each subnetwork occurs over a single sub-band and that each sub-band is further divided into L(L ≤ Mn ; ∀n) resource units (RUs). Devices within each subnetwork are then scheduled over the L RUs. With these assumptions, each transmission with a subnetwork may experience both intra- and inter-subnetwork interference. At slot, t, the received signal-to-noise-plus-interference ratio (SINR) on the link between the nth AP and its mth device can therefore be expressed as k [t] ξn,n,m P z z 2 n′ ∈Jnn′ ξn,n′ ,m′ [t] + n′ ∈Inn′ ξn,n′ ,m′ [t] + σ (4) where Jnn′ and Inn′ denotes the set of all devices and AP generating intra- and inter-subnetwork interference on the z subband, respectively. The term, σ 2 denotes the noise power calculated as a function of the bandwidth of each RU. Assuming a single antenna at both the APs and devices and considering the finite-block length approximation, the achieved rate can then be expressed as s Vnm [t] −1 ζnm [t] ≈ log2 (1 + Υnm [t]) − Q (η) log e, (5) Cℓ

Υznm [t] = P

II. S YSTEM M ODEL AND P ROBLEM F ORMULATION A. System Model In this paper, we consider the deployment of N 6G in-X subnetworks in an indoor factory hall for supporting industrial control operations. The subnetworks are indexed by the set N = {1, 2, · · · , N }. Each subnetwork consists of a controller collocated with a single access point (AP) for coordinating transmission to/from its associated devices i.e., sensor and actuators). We assume the nth subnetwork supports Mn randomly distributed devices. The devices in the nth subnetwork are indexed with mn ∈ Mn = {1, 2, · · · , Mn }. We assume that a total bandwidth, Bs , which is partitioned into Z subbands is available and that the number of sub-bands is much less than the number of subnetworks, i.e., Z << N . We index the available sub-bands with z ∈ {1, 2, · · · , Z}. Denoting the transmit power as P , the received signal strength (RSS) on the link between the nth AP from the mth device in the uth subnetwork can be expressed as k ξn,u,m [t] = P |hkn,u,m [t]|2 Γkn,u,m ϑn,u,m ,

where J0 (·), Ts and fd denote the zeroth order Bessel function of the first kind, the slot duration, and the maximum Doppler frequency, respectively. The path-loss component, Γkn,u,m is expressed as Γkn,z,m = 2 −α c dn,u,m /16π 2 fk2 , where dn,u,m is the link distance, c ≈ 3 × 108 ms−1 denotes the speed of light, fk and α are the carrier frequency of channel k and path-loss exponent, respectively. The shadowing fading component is computed using     d   1 − e − dn,u,m corr q , (3) ϑn,u,m = ln  d  (Yn + Yu,m )  √ − dn,u,m corr 2 1+e

(1)

where Γkn,u,m , hkn,u,m [t], and ϑn,u,m denote the pathloss, the small scale gain and log-normal shadowing, respectively. The small scale gain, hkn,u,m [t], is modelled as p (2) hkn,u,m [t] = ρhkn,u,m [t − 1] + 1 − ρ2 ϱkn,u,m , where ρ is the autocorrelation coefficient and ϱkn,u,m is an iid complex Gaussian variable. The autocorrelation coefficient is modeled as ρ = J0 (2πfd Ts )

where η denotes the decoding error probability, Q is the complementary Gaussian cumulative distribution function, and V denotes the channel dispersion which is defined as 1 Vnm [t] = 1 − . (6) (1 + Υnm [t])2 B. Control Operation Characteristics Unlike recent works on resource allocation for 6G in-X subnetworks, we consider more realistic heterogeneous characteristics of the control operations within each subnetwork and the associated communication characteristics. We consider the following key assumptions about the control operations supported by the devices in each subnetwork: • Packet size: The packet size of sensors is determined by their function. For instance, while sensors such as light, temperature, contact, proximity, and ultrasonic typically require a small packet size (in the order of a few bytes), others such as image sensors, audio sensors, and

environmental sensors like air quality sensors require a larger packet size (in the order of kilobytes to megabytes). We, therefore, consider these two classes of devices within each subnetwork. The packet size for device m transmission in subnetwork, n is then denoted as βmn = βsmall · I(U ≤ p) + βlarge · I(U > p),

(7)

where U ∼ U (0, 1) is a uniform random variable, p(0 ≤ p ≤ 1 is the probability that a device has packet size, βsmall and I is the indicator function that equals 1 if the condition is true and 0 otherwise. • Packet periodicity: Similar to the works in [6], [12], we consider deterministic periodic traffic with different periodicity depending on the requirements of the control operation supported by the device. We denote the period for the mth device in subnetwork n as Tmn . • Survival time: Another important characteristic of control applications is the survival time defined as the number of slots over which a control operation can continue without successful reception of anticipated packets. The survival time of the control operation supported by device m in the nth subnetwork is denoted as τmn . With the above control operation characteristics, intra- and inter-subnetwork interference becomes inevitable and must be properly managed to guarantee stringent communication requirements. C. Problem Formulation We consider an RRM problem involving a fully distributed joint selection of sub-bands and intra-subnetwork scheduling. Considering the control characteristics in Section II-B, the optimization problem can be formulated as that of minimizing the probability that the control operation fails or equivalently the probability that the burst error length, Lburst (i.e., the number of consecutive transmission opportunities in which a device fails to successfully deliver its packet) exceeds the survival time for all devices and formally be written as ( Mn )N mn P: min Prob[Lburst (a, vn ) > τmn ] {a,v}

s.t.

m=1

n=1

|vn | ≤ L ∀n

(8)

where a = [a1 · · · aN ]; an ∈ {1, 2, · · · , Z} ∀n is a vector of indices of the sub-band selected by all subnetworks and v contains the device scheduling decisions of all subnetworks. Note that minimizing the probability of control-operation failure as in (8) directly improves communication reliability for time-sensitive industrial applications. III. P ROPOSED E XPLAINABLE DRL M ETHOD FOR R ESOURCE A LLOCATION To solve the multi-objective optimization problem in (8), we proposed an explainable federated DRL (EFDRL) for joint sub-band selection and intra-subnetwork scheduling in this paper. The suggested EFDRL framework is depicted along with an illustration of the federated training procedure

in Figure 1. EFDRL combines explanation-guided learning (EGL) with federated multi-agent DRL to obtain enhanced performance and interoperability of the agent’s decisions. The different components of the proposed approach are described in the sequel. A. Action Space The decisions to be taken by the agents involve two main components viz: sub-band selection and scheduling to minimize inter- and intra-subnetwork interference, respectively. The action space for the nth subnetwork is then formed by combining the set of all available sub-bands and the set containing all possible combinations of L out of its Mn devices. B. State Space To aid the agents in learning to take joint sub-band selection and intra-subnetwork scheduling decisions, the local observation for each subnetwork is defined as a combination of the current wireless link characteristics (i.e., channel gain ben tween the AP and all its associated devices), {ξn,n,m (t)}M m=1 , the control system status (current estimated time to failure), TTF 1 n {τnm (t)}M and the sense aggregate interference power m=1 on all sub-bands, {Inz (t)}Z z=1 . In addition, the current action, an (t − 1) and reward value are included as part of the local state. The local observation for subnetwork n is then defined as h Mn TTF z Z n sn [t] = {ξn,n,m (t)}M m=1 , {τnm (t)}m=1 , {In (t)}z=1 , an (t − 1), rn (t − 1)]

T

(9)

C. Explanation Guided Reward Signal Inspired by the works in [14], [16], we design a composite reward comprising a QoS and XAI denoted for the n subnetwork as rnQoS and rnXAI , respectively. The composite reward, which is used as feedback during the training, is given as rn (t) = rnQoS (t) + rnXAI (t).

(10)

Considering the optimization problem in (8), the QoS reward for the subnetwork n is designed as Mn 1 X rnm (t), rn (t) = Mn i=1

(11)

where rnm (t) denotes the reward received by the agent in the nth subnetwork due to the success or otherwise of the link between sensor m and the AP and is defined as ( +αs if transmission is succesful rn (t) = . −αf otherwise exp (βTTFn (t)) (12) In (12), αs/f (αs/f > 0) is a constant reward/penalty scaling factor, and β(0 < β ≤ 1) is a parameter that controls the decay of the exponential function depending on the time to 1 Time-to-Failure (TTF) represents the remaining number of transmission opportunities before the survival time threshold of device mmm in subnetwork nnn is violated.

Next State

Subnetwork Access point Device Coordinator

Composite Reward

RL Agent Action

Replay buffer

Subnetwork Environment

QoS Reward

Reward Generator

Minibatch

RL Explainer

Soft Attributes

Entropy Mapper

XAI Reward

(b). Multi-subnetwork FL

(a). Explainable RL framework

Fig. 1. Illustration of the framework for explainable reinforcement learning (a) and multi-subnetwork federated learning (b).

failure, TTFnm (t) of the control operation supported by the sensor m. To characterize the XAI component of the reward, we adopt the multiplicative inverse of the Shannon entropy, i.e., rnXAI (t) =

1 , maxu Hu

(13)

with the entropy, Hu expressed as Hu = −

Z X

pz,u log(pz,u ).

(14)

z=1

In (14), pz,u denotes the probability distribution of the statesfeatures and is expressed for the uth sample in the experience replay buffer of the agent in subnetwork n as pn,z,u = PZ

exp{|ηn,z,u |}

z ′ =1 exp{|ηn,z ,u |}

,

(15)

where ηn,z,u denotes the SHAP value computed for state variable z of sample u in the experience replay buffer. D. Policy Representation To model the mapping between the state measurements and the dynamic joint channel selection and scheduling decision at each subnetwork, we propose to use a multi-agent proximal policy optimization (MAPPO) [17] algorithm with federated training. We denote the proposed algorithm as F-MAPPO. In F-MAPPO, two separate networks are trained for each subnetwork: an actor network and a value function network (called a critic). Without loss of generality, we assume that all agents share the actor and critic network. We denote the actor and critic network as πθ and Vϕ , respectively. The actornetwork learns to map the agent observations to a categorical distribution over the actions and is trained to maximize the objective function [17]

TABLE I S IMULATION PARAMETERS Parameter Number of subnetworks Devices per subnetwork Number of subchannels Number of resource units Subnetwork radius Minimum device-controller distance Minimum controller distance Channel bandwidth Carrier frequency Clutter type Clutter element size Clutter density Shadowing standard deviation Maximum transmission power Noise power spectral density Noise figure Factory area Correlation distance Sample time Robot speed Target transmission rate Survival times Traffic periodicity

Value 10 4 3 2 0.5 m 0.3 m 1.0 m 40 MHz 6 GHz Dense 2m 60% 7.2 1W -174 dBm/Hz 5 dB 20 × 20 m2 5m 0.05 s 5 m/s Randomized (0.4 or 1 Mbps) Randomized (2-5 steps) Randomized (1-10 steps)

n a generalized advantage estimation [18] method, and rθ,b is defined as πθ (anb |snb ) n rθ,b = . (17) πθold (πθ (anb |snb )

The critic network learns to minimize the function  |B| N 2 1 XX L(ϕ) = max Vϕ (snb ) − R̂b , |Bn |N b=1 n=1  2  clip (Vϕ (snb ), Vϕold (snb ) − ϵ, Vϕold (snb ) + ϵ) − R̂b , (18) where R̂b denotes the discounted reward calculated over data b from the experience replay buffer.

|B|

L(θ) =

N   1 XX IV. P ERFORMANCE E VALUATION n n min rθ,b Anb , clip rθ,b , 1 − ϵ, 1 + ϵ Anb |Bn |N A. Simulation Setup n=1 b=1

|B|

N

σ XX + S [πθ (snb )] , |Bn |N n=1

(16)

b=1

where |B| and σ denote the batch size and the entropy coefficient, respectively. The term S represents the policy entropy, Anb denotes the advantage, which is computed using

We consider a network with N = 10 subnetworks, each comprising a single controller serving as the access point (AP) for multiple devices within the subnetwork. The subnetworks are uniformly distributed within a rectangular area 20 m × 20 m, corresponding to a high deployment density of 25, 000 subnetworks/km2 . Subnetworks move in the area

Fig. 2. Decrease in probability of failure relative to random subband allocation with random scheduling with M = 2.

Fig. 3. Decrease in probability of failure relative to random subband allocation with random scheduling with M = 4.

following a restricted random direction mobility with a velocity v = 2 m/s, resulting in a Doppler frequency of fd = 40 Hz at the considered carrier frequency of 6 GHz. A total bandwidth B = 120 MHz, which is divided into K = 3 subbands, each consisting of Z = 2 resource units (RUs) is considered. We utilized a fixed transmit power Ptx = 0 dBm. In the simulation, each subband is divided into 2 resource units (RUs), enabling the concurrent transmission of up to 2 devices within each subnetwork at every time instant. At each transmission slot, Zs (1 ≤ Zs ≤ Z) devices are scheduled on the RUs within the subband allocated to each subnetwork. Devices are associated with varying target rates, randomly assigned from a low-rate and high-rate distribution, reflecting the diverse communication demands of industrial applications. Each device has a survival time randomly selected between 2 and 6 slots, representing critical deadlines for successful transmission. Other simulation parameters are presented in Table I. B. Benchmarks The performance of the proposed method is compared to the following state-of-the-art benchmarks: • SISA - PF: This approach combine the Subband Iterative Scheduling Algorithm (SISA) with Proportional Fair scheduling. This approach balances fairness and efficiency by leveraging SISA for subband allocation while optimizing scheduling to prioritize users with favorable channel conditions relative to their time to failure. • SISA - RR: This method uses SISA for subband allocation but employs a Round Robin (RR) scheduler approach

Fig. 4. Mean absolute SHAP values for the features in the state space.

to ensure fairness among users. Each device is served in turn, irrespective of channel conditions or requirements. • Random - PF: Subbands are randomly allocated to subnetworks. Proportional Fair scheduling is then applied to perform intra-subnetwork scheduling. • Random - RR: Subbands are assigned randomly, and devices are scheduled via round robin. • Random - Random: Both subband allocation and scheduling are completely random. This serves as a baseline to assess the performance of other algorithms. • SISA - GA: This method combines SISA with an idealized scheduler that has perfect knowledge of all network parameters. This provides an upper bound on performance. C. Results The F-MAPPO-based approach for joint subband allocation and intra-subnetwork scheduling is trained over 3, 000 episodes, each consisting of 200 steps with an interval of 0.05 s between steps. During training, the agent demonstrates convergence after approximately 2, 000 episodes. The trained agent is then deployed across each subnetwork to perform joint subband selection and scheduling during the execution phase, which spans 500 episodes. The performance of the proposed scheme is evaluated based on the probability of failure, defined as the likelihood that the burst error length at any device exceeds the survival time. To quantify the improvement resulting from a subband allocation and/or scheduling approach, we compute the percentage decrease of the probability of failure (PDPF) of the different schemes relative to the Random Random baseline which performs both subband allocation and scheduling randomly. Figure 2 presents the PDPF for various approaches when the number of devices equals the number of RUs per unit, i.e., M = Z = 2. As expected, all schemes based on random subband allocation show no improvement in PDPF. This outcome is due to the scheduling decision, which permits the transmission of all devices, as the number of available RUs matches the number of devices. Conversely, methods utilizing SISA for subband allocation achieve a significant improvement in the probability of failure, approximately 29%. The proposed F-MAPPO method offers a comparable gain of 28.96%, resulting in similar performance but with substantially lower complexity.

The PDPF is shown in Figure 3 for the case where the number of devices per subnetwork is M = 4 and the number of RUs is Z = 2, corresponding to a scenario where both subband allocation and scheduling decisions impact overall performance. The figure demonstrates that using a roundrobin (RR) scheduler is detrimental to transmission performance, regardless of the scheme used for subband allocation. The performance degradation due to RR scheduling is approximately 22% and 18% with random and SISA-based subband allocation, respectively. This is expected, as RR scheduling ignores the survival time requirement, potentially delaying transmissions that are critical for devices with tighter deadlines. The figure also reveals that survival time-aware PF scheduling and Genuine Aided (GA) schedulers result in significant performance improvements, even when random subband allocation is used. Overall, the SISA-GA and FMAPPO schemes offer the best performance, with a reduction in the probability of device transmission failure of 46.67% and 45.52%, respectively. Figure 4 presents a bar chart of the mean absolute SHAP values, highlighting the relative importance of various features in the state space within the model’s decision-making process for joint scheduling and subband allocation. As illustrated, features such as interference power and time to failure exhibit the highest SHAP values, underscoring their critical role in shaping the agent’s decisions. This aligns with domain knowledge, as mutual interference between subnetworks directly affects the efficiency of subband allocation, while time-to-failure is essential for prioritizing devices appropriately in scheduling. The figure also reveals that intra-subnetwork channel gain and past decisions have relatively lower impact compared to interference power and time to failure. This suggests that the model uses these features to refine its predictions, but with less emphasis than the more influential features. In contrast, features like past achieved rate show comparatively lower SHAP values, indicating a minimal direct influence on decisions, although they may contribute indirectly through interactions with other features. V. C ONCLUSION In this paper, we proposed a novel AI-based framework for dynamic resource allocation and scheduling in dense deployments of 6G in-X subnetworks. By integrating MultiAgent Reinforcement Learning (MARL), Federated Learning (FL), and Explainable AI (XAI), our approach effectively addresses key challenges such as privacy, transparency, and scalability. Unlike existing solutions, our method supports multiple devices per subnetwork, providing a more realistic and flexible resource management strategy that can adapt to the complexities of 6G networks. The proposed framework enables collaborative optimization without the need for raw data sharing, ensuring the privacy of individual subnetworks. Additionally, the use of XAI techniques enhances the interpretability of the decision-making process, allowing stakeholders to better understand and trust the model’s decisions. Our extensive simulations, based on 3GPP scenarios, demonstrate

that the proposed approach outperforms existing baselines for joint subband allocation and scheduling, which combine stateof-the-art solutions for subband allocation and scheduling. The results confirm that our framework not only improves performance but also provides a scalable and transparent solution for managing resource allocation. R EFERENCES [1] G. Berardinelli, R. Adeogun, E. M. Vitucci, S. Giannoulis et al., “Boosting short-range wireless communications in entities: the 6g-shine vision,” in IEEE Future Networks World Forum 2023. IEEE, 2023. [2] R. Adeogun, G. Berardinelli, P. E. Mogensen, I. Rodriguez, and M. Razzaghpour, “Towards 6g in-x subnetworks with sub-millisecond communication cycles and extreme reliability,” IEEE Access, vol. 8, pp. 110 172–110 188, 2020. [3] G. Berardinelli, P. Baracca, R. O. Adeogun, S. R. Khosravirad, F. Schaich, K. Upadhya, D. Li, T. Tao, H. Viswanathan, and P. Mogensen, “Extreme communication in 6g: Vision and challenges for ‘inx’subnetworks,” IEEE Open Journal of the Communications Society, vol. 2, pp. 2516–2535, 2021. [4] R. Adeogun and G. Berardinelli, “Distributed channel allocation for mobile 6g subnetworks via multi-agent deep q-learning,” in IEEE WCNC. IEEE, 2023, pp. 1–6. [5] X. Du, T. Wang, Q. Feng, C. Ye, T. Tao, L. Wang, Y. Shi, and M. Chen, “Multi-agent reinforcement learning for dynamic resource management in 6g in-x subnetworks,” IEEE Transactions on Wireless Communications, vol. 22, no. 3, pp. 1900–1914, 2022. [6] R. Adeogun, G. Berardinelli, I. Rodriguez, and P. Mogensen, “Distributed dynamic channel allocation in 6g in-x subnetworks for industrial automation,” in IEEE GC Wkshps. IEEE, 2020, pp. 1–6. [7] R. Adeogun, G. Berardinelli, and P. Mogensen, “Learning to dynamically allocate radio resources in mobile 6g in-x subnetworks,” in IEEE PIMRC. IEEE, 2021, pp. 959–965. [8] D. Li, S. R. Khosravirad, T. Tao, and P. Baracca, “Advanced frequency resource allocation for industrial wireless control in 6g subnetworks,” in 2023 IEEE Wireless Communications and Networking Conference (WCNC), 2023, pp. 1–6. [9] N. Khan, S. Coleri, A. Abdallah, A. Celik, and A. M. Eltawil, “Explainable and robust artificial intelligence for trustworthy resource management in 6g networks,” IEEE Communications Magazine, 2023. [10] S. Kavaiya, “Learn with curiosity: A hybrid reinforcement learning approach for resource allocation for 6g enabled connected cars,” Mobile Networks and Applications, pp. 1–11, 2023. [11] Q. Wang, Y. Liu, Y. Wang, X. Xiong, J. Zong, J. Wang, and P. Chen, “Resource allocation based on radio intelligence controller for open ran towards 6g,” IEEE Access, 2023. [12] R. Adeogun and G. Berardinelli, “Multi-agent dynamic resource allocation in 6g in-x subnetworks with limited sensing information,” Sensors, vol. 22, no. 13, p. 5062, 2022. [13] A.-D. Marcu, S. K. G. Peesapati, J. M. Cortes, S. Imtiaz, and J. Gross, “Explainable artificial intelligence for energy-efficient radio resource management,” in 2023 IEEE Wireless Communications and Networking Conference (WCNC). IEEE, 2023, pp. 1–6. [14] S. Roy, H. Chergui, and C. Verikoukis, “Explanation-guided fair federated learning for transparent 6G RAN slicing,” IEEE Transactions on Cognitive Communications and Networking, vol. 10, no. 6, pp. 2269– 2281, 2024. [15] N. S. Ruiz, S. K. G. Peesapati, and J. M. Cortes, “Xai-assisted radio resource management: Feature selection and shap enhancement,” in ICC 2023-IEEE International Conference on Communications. IEEE, 2023, pp. 2516–2521. [16] F. Rezazadeh, H. Chergui, and J. Mangues-Bafalluy, “Explanationguided deep reinforcement learning for trustworthy 6G RAN slicing,” in 2023 IEEE ICC Workshops. IEEE, 2023, pp. 1026–1031. [17] C. Yu, A. Velu, E. Vinitsky, J. Gao, Y. Wang, A. Bayen, and Y. Wu, “The surprising effectiveness of ppo in cooperative multi-agent games,” Advances in Neural Information Processing Systems, vol. 35, pp. 24 611– 24 624, 2022. [18] J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel, “Highdimensional continuous control using generalized advantage estimation,” in International Conference on Learning Representations (ICLR), 2016.

Record · ID 1028630 · SHA-256 822aeab9283f115b
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.