1
Spectral- and Energy-efficient Multi-BS Multi-RIS Pinching-antenna Systems: A GNN-based Approach
arXiv:2605.01307v1 [eess.SP] 2 May 2026
Changpeng He, Yang Lu, Senior Member, IEEE, Wei Chen, Senior Member, IEEE, Bo Ai, Fellow, IEEE, Arumugam Nallanathan, Fellow, IEEE and Zhiguo Ding, Fellow, IEEE
Abstract—This paper investigates coordinated downlink transmission in a multi-base station (multi-BS) multi-reconfigurable intelligent surface (multi-RIS)-assisted pinching-antenna (PA) system, where each user equipment (UE) is associated with a single BS and each BS is equipped with movable PAs deployed on parallel waveguides. We formulate sum rate (SR) and energy efficiency (EE) maximization problems by jointly optimizing PA placement, RIS phase shifts, transmit beamforming, and BSUE association under constraints of inter-PA spacing, power budget, and unit-modulus phase shift. To address the resulting highly coupled mixed-variable problem, we propose a threestage graph neural network (GNN) that integrates heterogeneous and homogeneous graph representations and is trained end-toend in an unsupervised manner. Extensive numerical results demonstrate that the proposed three-stage GNN consistently outperforms representative system and learning baselines, generalizes well to unseen numbers of UEs, RISs, and BSs, and maintains millisecond-level inference time. Besides, the results validate the effectiveness of the proposed design from both system and architectural perspectives. Moreover, PAs are shown to enhance SR and EE, and the performance gain is enlarged with increasing number of PAs. Index Terms—PA system, RIS, BS-UE association, graph neural network.
I. I NTRODUCTION The sixth generation (6G) wireless network is expected to support immersive services, native intelligence, and stringent spectral- and energy-efficiency requirements in highly dynamic environments, making the conventional paradigm of optimizing only transceivers increasingly inadequate [1]. To enlarge the controllable degrees of freedom (DoFs) of wireless propagation, reconfigurable intelligent surfaces (RISs) have attracted extensive attention as a low-power means of shaping channels via programmable reflections [2]–[5]. RISaided transmission can improve coverage, spectral efficiency, energy efficiency (EE) , and physical-layer security [6]–[9]. Nevertheless, RISs mainly provide environment-side control, Changpeng He and Yang Lu are with the State Key Laboratory of Advanced Rail Autonomous Operation, and also with the School of Computer Science and Technology, Beijing Jiaotong University, Beijing 100044, China (e-mail: [email protected], [email protected]). Wei Chen and Bo Ai are with the School of Electronics and Information Engineering, Beijing Jiaotong University, Beijing 100044, China (e-mail: [email protected], [email protected]). Arumugam Nallanathan is with the School of Electronic Engineering and Computer Science, Queen Mary University of London, London and also with the Department of Electronic Engineering, Kyung Hee University, Yongin-si, Gyeonggi-do 17104, South Korea (e-mail: [email protected]). Zhiguo Ding is with the School of Electrical and Electronic Engineering (EEE), Nanyang Technological University, Singapore 639798 ([email protected]).
and their performance is still constrained by cascaded path loss and the fixed geometry of active arrays, especially in blocked or cell-edge scenarios. Recently, flexible-antenna technologies such as movable antennas (MAs), fluid antenna systems (FASs), and pinchingantenna (PA) systems (also known as PASS) have emerged as a complementary paradigm for wireless channel reconfiguration [10]–[13]. By enabling adaptive repositioning of radiating elements, these architectures exploit the spatial variation of wireless channels within a prescribed aperture and create additional DoFs beyond conventional digital beamforming. For example, MA/FAS studies have established field-response based channel models, quantified the gains of antenna position adaptation, and demonstrated substantial benefits in multi-user and multiple-input multiple-output (MIMO) systems [14]–[17]. PA systems further introduce low-loss dielectric waveguides and movable PAs, making it possible to radiate signals from different positions along extended waveguides and thus enhance line-of-sight accessibility and aperture reconfigurability [12], [13]. This motivates the integration of PA systems with RISs: PA systems offer transmitter-side spatial reconfiguration, while RISs provide environment-side wave control; together they can potentially deliver more robust coverage and higher efficiency than either mechanism alone. Their interplay is also nontrivial, as the gains brought by antenna repositioning and environmental reflection may be complementary or partially overlapping depending on the propagation geometry [18]. The coordinated design of a multi-base station (multi-BS) multi-RIS-assisted PA system is, however, highly challenging. To avoid the stringent synchronization and signaling overhead required by coherent multi-BS transmission, each user equipment (UE) is served by a single BS, and the introduced BSUE association serves as an additional degree of freedom for interference management and load balancing. Consequently, system performance depends on the joint optimization of PA positions, baseband beamforming, RIS phase shifts, and BSUE association, all of which are tightly coupled through the effective channels. The resulting SR- and EE-maximization problems involve mixed continuous, unit-modulus, and binary variables under power and spacing constraints. Compared with classical beamforming design, the combination of multi-BS coordination and multi-variable optimization renders conventional iterative algorithms substantially more computationally demanding and harder to converge, with significantly degraded scalability for large antenna arrays, dense RIS deployments, or fast-varying network topologies [9], [19], [20]. Learning-based wireless optimization provides a promising
2
alternative. Prior to graph-based models, general deep architectures such as multilayer perceptrons (MLPs) and convolutional neural networks (CNNs) had already been explored for beamforming and related transmission design problems [21]–[24]. When the problem dimension is fixed and the input admits a regular grid-like representation, these models are effective. However, they do not exploit the relational structure among network entities such as BSs and UEs, and their transfer capability is therefore often limited when the number of network entities or the interference topology changes. In contrast, graph neural networks (GNNs) are naturally suited to wireless systems because propagation, interference, and cooperation relations can be represented by graphs, enabling structure-aware policy learning with favorable scalability and permutation-equivariant inductive bias [25], [26]. Recent studies have demonstrated the effectiveness of GNNs for beamforming and wireless network design [27], [28]. For flexibleantenna systems, staged GNN frameworks have also started to emerge for FAS and PA systems [29]–[32]. Nevertheless, existing learning-based PA systems/RISs work mainly focus on the single-cell scenarios, and omit BS-UE association. A unified framework for coordinated multi-BS multi-RISassisted PA systems is still lacking. To fill this gap, this paper studies coordinated downlink transmission in a multi-BS multi-RIS-assisted PA system and develops a three-stage GNN for end-to-end unsupervised optimization. The main contributions are summarized as follows. We formulate a multi-BS multi-RIS-assisted PA systems serving multiple UEs, based on which, we formulate SR and EE maximization problems by jointly optimizing PA placement, RIS phase design, transmit beamforming, and BS-UE association under practical power, spacing, and unit-modulus constraints. The SR and EE maximization problems are solved by a deep learning model in a unified manner. • We propose a three-stage GNN architecture composed of ChanGNN, BeamGNN, and AssocGNN. By combining heterogeneous and homogeneous graph representations, the proposed model explicitly captures BS-UE, BS-RIS, RIS-UE, intra-BS, and intra-UE interactions. Moreover, feasibility-preserving readout mechanisms are designed for all decision variables, including spacingbased placement, unit-modulus RIS phase normalization, per-BS power normalization, and differentiable BS-UE association via Gumbel-Softmax. • Extensive numerical results demonstrate that the proposed framework consistently outperforms both system and model baselines. The results further show that the proposed method generalizes well to unseen numbers of UEs, BSs, and RISs while maintaining millisecond-level inference times. In addition, the ablation study confirms the importance of message passing, residual fusion, and complex-valued mappings in the proposed architecture, and the impact of key system parameters on system performance is also illustrated.
•
Notation: The following mathematical notations and symbols are used throughout this paper. a and A stand for a
column vector and a matrix (or tensor), respectively. The sets of real numbers, and n-by-m real matrices are denoted by R, and Rn×m , respectively. The sets of n-dimensional complex column vectors and n-by-m complex matrices are denoted by Cn and Cn×m , respectively. For a complex number a, |a| denotes its modulus and Re(a) denotes its real part. For a vector a, ∥a∥ is the Euclidean norm. For a matrix A, AH and ∥A∥ denote its conjugate transpose and Frobenius norm, respectively. [a]i , [A]i,j , and [A]i,: denote the i-th element of vector a, the i-th row and the j-th column element of the matrix A, and the i-th row vector of the matrix A, respectively. {ai } denotes the set of elements for all the admissible i. diag(·) and Concat(·) denote the diagonalization and concatenation operations, respectively. CN(·) denotes the circularly symmetric complex Gaussian distribution. II. R ELATED W ORK The literature most relevant to this work can be grouped into three categories: RISs, PA systems, and GNN-based wireless optimization. A. RIS-aided Wireless Communications RIS-enabled wireless transmission has been extensively investigated over the past few years. Foundational works established the signal model, reflection mechanism, and fundamental advantages of RIS-assisted smart radio environments [2]–[5]. Subsequent studies considered energy-efficient design, secure transmission, and multicell coordination [6]– [9]. These works consistently showed that jointly optimizing active beamforming and passive reflection can significantly improve system performance. Beyond pure beamforming design, association-aware RIS cellular optimization has also been studied. For example, Liu et al. jointly optimized BS-RIS-UE association together with active and passive beamforming for RIS-aided cellular networks through a successive-access and alternating-optimization framework [33]. However, the above studies assume fixed active antenna arrays and therefore only exploit environment-side reconfiguration. B. Pinching-antenna Systems PA has recently emerged as a representative flexible-antenna architecture based on low-loss dielectric waveguides and movable pinching elements [12], [13]. Existing studies on PA systems have progressed from principle and modeling toward optimization and intelligent design. On the optimization side, Bereyhi et al. studied downlink MIMO beamforming with PA systems and jointly optimized the digital precoder and activated PA positions via fractional programming [34]. Xu et al. further investigated joint transmit and pinching beamforming in multi-user PA systems and proposed both an optimizationbased majorization-minimization and penalty dual decomposition (MM-PDD) solver and a learning-based Karush-KuhnTucker-guided dual learning (KDL)-Transformer framework [35]. On the learning side, Xie et al. formulated PA systems as a bipartite graph and used a graph attention network (GAT)-based model to jointly optimize antenna placement
3
and power allocation for EE maximization [30]. Guo et al. proposed GPASS, where one sub-GNN learns PA positions and another sub-GNN learns transmit beamforming, thus yielding a staged GNN framework tailored to pinching beamforming [31]. More recently, the RIS-assisted downlink PA systems in [32] incorporated RIS phase design into a three-stage GNN for PA positioning, RIS configuration, and beamforming. Despite these advances, current learning frameworks for PA systems are still mainly centered on single-BS settings and do not jointly handle coordinated multi-BS transmission and BS-UE association.
...
...
...
...
...
...
RIS 2
BS 2 UE 3
h 2,1 H1,2
UE 2 UE 1
h1,1
C. GNN-based Wireless Optimization GNN-based wireless optimization has developed from general graph learning principles to increasingly coupled and structured transmission design. Early studies established why GNNs are well matched to wireless resource allocation, by showing that many wireless optimization problems exhibit graph structure and permutation equivariance [25], [26]. Representative applications include multi-user multiple-input single-output (MU-MISO) beamforming, where Li et al. directly mapped channel state information (CSI) graphs to beamforming vectors for SR maximization [27], and unmanned aerial vehicle (UAV) communications, where Wang et al. used a two-stage GNN to jointly handle UAV placement and transmission design [28]. Multi-stage GNN design has also appeared in flexible-antenna systems; for example, He et al. proposed a two-stage GNN for FAS, where antenna-position inference and beamforming inference are learned sequentially [29]. More recent works have started to address coupled multivariable optimization in RIS systems. Le et al. designed a heterogeneous GNN for joint active and passive beamforming in distributed STAR-RIS-assisted MU-MISO systems [36], while Liu et al. proposed a heterogeneous GNN with RIS association and beamforming updates for multi-RIS multi-user mmWave systems [37]. These studies strongly motivate graph-based learning for structured wireless optimization. Nevertheless, existing GNN frameworks have not yet unified transmitter-side spatial reconfiguration, environment-side wave control, active beamforming, and BS-UE association within one coordinated PA-RIS architecture, which is precisely the gap addressed in this paper. III. S YSTEM M ODEL AND P ROBLEM D EFINITION As illustrated in Fig. 1, we consider the coordinated downlink transmission for a multi-RIS-assisted PA system, where B PA-BSs serve K single-antenna UEs with the assistance of R RISs within a rectangular region spanning a length of D along the x-axis and S along the y-axis. Each BS is equipped with N parallel waveguides, and K ≤ N holds. The waveguides of BS b are positioned at a height of Hb , extend along the x-axis, and span a length of Cb . Each waveguide hosts M movable PAs. We denote the position of PA m on P P = (xP waveguide n of BS b as ψb,n,m b,n,m , yb,n , Hb ). UE k U U U is located at ψk = (xk , yk , 0) within the S × D rectangular region. RIS r, comprising L reflecting elements, is located R R at ψrR = (xR r , yr , zr ). For clarity, the sets of UEs, BSs,
f1,1 H1,1
RIS 1
...
...
...
...
...
...
BS 1
Fig. 1. Illustration of multi-BS multi-RIS-assisted PA systems for downlink multi-user communications.
RISs, PAs per waveguide are denoted as K ≜ {1, . . . , K}, B ≜ {1, . . . , B}, R ≜ {1, . . . , R}, and M ≜ {1, . . . , M } respectively. A. Channel Model The system comprises four types of channels: 1) inwaveguide propagation; 2) the PA-RIS link, 3) the RIS-UE link, and 4) the direct PA-UE link. They are detailed as follows. 1) In-Waveguide Propagation: Denote the pinching beam(M N )×N forming matrix of BS b by Gb ({xP , which b,n,m }) ∈ C is given by Gb xP = (1) b,n,m n o gb,1 xP ··· 0 b,1,m .. .. .. , . . . n o P 0 · · · gb,N xb,N,m where " 2π P P P −ψb,n,1 − ζ+j λ ∥ψb,n,0 ∥ g ,..., gb,n xb,n,m = e
e
#T 2π P P − ζ+j λ −ψb,n,M ∥ψb,n,0 ∥ g
(2)
∈ CM ,
P P with ψb,n,0 = (xP b,n,0 , yb,n , Hb ) denoting the location of the feed point for waveguide n of BS b, and λg = λ/neff denoting the guided wavelength, where ζ denotes the in-waveguide attenuation coefficient, λ denotes the free-space wavelength and neff denotes the effective refractive index of the dielectric waveguide. Notably, gb,n ({xP b,n,m }) captures the in-waveguide attenuation. 2) PA-RIS Link: The channel between M PAs on waveguide n of BS b and RIS r is modeled by (3), where α denotes
4
"s Hb,r,n xP = b,n,m
β0 lb,r xP b,n,1 , . . . , P ∥ψrR − ψb,n,1 ∥α
the path loss exponent, β0 denotes the channel gain at the L reference distance of 1 m, and lb,r (xP b,n,m ) ∈ C denotes the small-scale fading component from PA m on waveguide n of BS b to L reflecting elements, i.e., r r κ LoS P 1 NLoS = l l lb,r xP x + , (4) b,n,m b,n,m 1 + κ b,r 1 + κ b,r where κ denotes the Rician factor, and P lLoS b,r xb,n,m = h iT 2π 2π 1, e−j λ ∆φb,r,n,m , · · · , e−j λ (L−1)∆φb,r,n,m ,
(5)
where φb,r,n,m ∈ [0, 2π) denotes the angle-of-departure (AoD) from PA m on waveguide n of BS b to RIS r, ∆ denotes the element separation, and lNLoS ∼ CN(0, IL ) contains the nonb,r line-of-sight (NLoS) coefficients. The channel matrix from all PAs of BS b to RIS r is expressed as Hb,r xP = (6) b,n,m P N L×M N Concat Hb,r,n xb,n,m ∈C . n=1 3) RIS-UE Link: The channel between RIS r and UE k is denoted by hr,k ∈ CL , given by s ! r r β0 κ 1 hLoS + hNLoS , hr,k = 1 + κ r,k 1 + κ r,k ∥ψrR − ψkU ∥α (7) NLoS where hLoS denote the LoS and NLoS small-scale r,k and hr,k P fading components with similar definitions to lLoS b,r (xb,n,m ) and NLoS lb,r . 4) PA-UE Link: The direct PA-UE link from all PAs of BS b to UE k is modeled by fb,k xP = (8) b,n,m T P T T fb,1,k xb,1,m , . . . , fb,N,k xP ∈ CM N , b,N,m
where
s
# β0 lb,r xP ∈ CL×M b,n,M P ∥ψrR − ψb,n,M ∥α
binary elements {ub,k } given by ( 1, UE k is associated with BS b, (10) ub,k = 0, otherwise. P where b∈B ub,k = 1 holds. The UE receives the direct signal from its associated BS and the reflections from all RISs. The transmitted signal of each PA is a phase-shifted replica of the signal from the feed point of its waveguide. The signal emitted by PAs of BS b intended for UE k is given by wb,k ub,k sk ∈ CM N , sb,k = Gb xP (11) b,n,m where sk ∈ C represents the information symbol for UE k with E[|sk |2 ] = 1 and wb,k ∈ CN denotes the baseband beamforming vector of BS b for UE k. The received signal at UE k is given by (12), where the diagonal matrix Φr ≜ diag ejϕr,1 , ejϕr,2 , ..., ejϕr,L ∈ CL×L (13) denotes the phase-shift matrix of RIS r with ϕr,l ∈ [0, 2π), and nk ∼ CN(0, σk2 ) denotes the additive white Gaussian noise (AWGN). The achievable rate at UE k is given by (14). Our objective is to jointly optimize PA positions, baseband beamforming vectors, RIS reflecting coefficients, and BS-UE association to maximize the sum rate (SR) or EE: X P1 : max Rk xP b,n,m , {wb,k } , {Φr } , U P k∈K {xb,n,m },U, {wb,k },{Φr }
(15a) n o Rk xP b,n,m , {wb,k } , {Φr } , U k∈K P P P2 : max 2 {xPb,n,m },U, b∈B k∈K ub,k ∥wb,k ∥ + PC P
{wb,k },{Φr }
(15b) s.t. 0 ≤ xP b,n,m ≤ Cb , ∀b, n, m, P xb,n,m − xP b,n,m−1 ≥ ∆min , X
fb,n,k xP = (9) b,n,m T 2π 2π P P √ −j ∥ψkU −ψb,n,1 √ −j ∥ψkU −ψb,n,M ∥ ∥ ηe λ ηe λ ∈ CM , , . . . , P P ∥ψkU − ψb,n,1 ∥ ∥ψkU − ψb,n,M ∥ 2
2
with η = c /(4πfc ) , where c is the speed of light and fc is the carrier frequency.
B. BS-UE Association and Problem Formulation We assume that each UE is associated with exactly one BS, and define the corresponding variables to optimize, i.e., the association matrix, which is defined by U ∈ RB×K with
(3)
2
k∈K
(15c) ∀b, n, ∀m > 1,
ub,k ∥wb,k ∥ ≤ Pmax , ∀b,
(15d) (15e)
ϕr,l ∈ [0, 2π), ∀r, l, X ub,k = 1, ∀k,
(15g)
ub,k ∈ {0, 1}, ∀k, b,
(15h)
b∈B
(15f)
where PC denotes the constant circuit power, ∆min > 0 denotes the minimum spacing between any two adjacent PAs, and Pmax denotes the total power budget. Notably, the considered Problems P1 and P2 are challenging to solve due to complexly coupled variables and integer variables. Specifically, they deviate significantly from the tractable forms amenable to conventional convex optimization approaches. Instead, we propose a unified deep learning-based approach to obtain near-optimal solutions to Problems P1 and
5
! yek =
P X H H hr,k Φr Hb,r xP fb,k xb,n,m + b,n,m
X
X
|
P xb,n,m wb,k′ ub,k′ sk′ + nk
{z
}
e H ({xP ≜h b,k b,n,m },{Φr })
P xb,n,m , {wb,k } , {Φr } , U = n o n o 2 P P P eH u h x , {Φ } G x w b,k r b b,k b,k b,n,m b,n,m b∈B log2 1 + P n o n o 2 P H ′e xP xP wb,k′ + σk2 b∈B k′ ∈K\{k} ub,k hb,k b,n,m , {Φr } Gb b,n,m Rk
{ rR }
G2
G1
(19) (20) (21) (22)
{xbP,n , m }
(23) (13)
{Φ r }
Graph Representation
{ kU }
(14)
BeamGNN (Stage 2) Pre-processing (24)
Graph Representation
ChanGNN (Stage 1)
{ bP,n ,0 }
(12)
k′ ∈K
r∈R
b∈B
Gb
HAL
G3
G4
G5
U
(33)
(34)
{ p b , k } G8
G9
HZM Learning (26)
{ pb, k }
Initialization (31)
{ b ,k }
Graph Representation
{w b ,k }
HZM Learning (26)
AssocGNN (Stage 3)
{ b ,k }
{ p b,k }
GAL
(29)
(30)
G6
G7
FL
Fig. 2. Structure of the proposed three-stage GNN. Stage 1 (ChanGNN) employs CHAL and CFL to jointly map location features to PA positions {xP b,n,m } (from BS node outputs) and RIS phase shift matrices {Φr } (from RIS node outputs). Stage 2 (BeamGNN) uses two parallel branches of CGAL and CFL to map effective channel features to hybrid coefficients {αb,k } and power allocation {e pb,k }. Stage 3 (AssocGNN) employs CGAL and CFL to map beamforming gain features to the association matrix U via Gumbel-Softmax. Finally, {wb,k } is recovered via HZM.
P2 . IV. T HREE -S TAGE GNN FOR J OINT O PTIMIZATION To exploit the spatial relationships and interference structure inherent in the considered multi-BS multi-RIS PA system, we propose a three-stage GNN that maps the locations of BSs, RISs, and UEs to PA positions, RIS phase shifts, beamforming vectors, and BS-UE association coefficients, with the objective of maximizing the system SR or EE. The proposed model employs three types of layers: the Complex Heterogeneous Graph Attention Layer (CHAL), the Complex Graph Attention Layer (CGAL), and the Complex Fully-Connected Layer (CFL), which will be introduced in detail later. A. Graph Representation and Overall Framework The proposed model is a three-stage GNN, as illustrated in Figure 2. To facilitate effective feature extraction, we model the considered system as distinct corresponding to each of the three stages.
1) Stage 1: A heterogeneous graph G1 = (V1 , E1 ) is constructed as shown in Figure 3(a), where V1 = VB ∪VU ∪VR with |VB | = B, |VU | = K, |VR | = R denotes BS, UE, and RIS nodes, respectively. These nodes are featured by P their locations, i.e., {ψb,n,0 }, {ψkU } and {ψrR }. The edge set E1 comprises six types of fully-connected directed edges, representing B × K BS-UE pairs, B × K UE-BS pairs, B × R BS-RIS pairs, B × R RIS-BS pairs, and R × K RIS-UE pairs, and R × K UE-RIS pairs, respectively. Stage 1 adopts a model termed, channel GNN (ChanGNN), to jointly learn PA positions {xP b,n,m } (from BS node outputs) and RIS phase shift matrices {Φr } (from RIS node outputs) over the defined graph. 2) Stage 2: A homogeneous directed graph G2 = (V2 , E2 ) is constructed as illustrated in Figure 3(b), where V2 contains K × B nodes, each representing a BS-UE link. Each node is characterized by the effective channel (cf. (24)), computed with the obtained {xP b,n,m } and {Φr }. The edge set E2 comprises BK(K + B − 2)/2 bidirectional edges that char-
6
as BS Node
Fig. 3. An example of graph representation with K = 3, B = 2 and R = 2. Intra-user edges
edges
B rows
acterize both inter-BS and intra-BS relationships. Specifically, the inter-BS edges for UE k connect all pairs of nodes (k, b) and (k, b′ ) with b ̸= b′ , capturing inter-BS competition for associating with the UE; while the intra-BS edges for BS b connect all pairs of nodes (k, b) and (k ′ , b) with k ̸= k ′ , K columns capturing inter-UE interference within the BS. Stage 2 adopts a model, termed beamforming GNN (BeamGNN), to learn B × K beamforming vectors {wb,k } over the defined graph, where each node feature is mapped to a corresponding beamforming vector. 3) Stage 3: The modeled graph is also based on G2 = (V2 , E2 ), but each node is characterized by the beamforming gain (cf. (31)). Stage 3 adopts a model, termed association GNN (AssocGNN), to learn the BS-UE association matrix U from beamforming gain features. The three GNNs are sequentially stacked and jointly trained, with the detailed processes of each stage provided as follows.
B. Stage 1: ChanGNN P The input of ChanGNN is {ψb,n,0 }, {ψkU } and {ψrR }, and P it yields {xb,n,m } based on the updated features of BS nodes and {Φr } based on the updated features of RIS nodes in a node-wise manner. 1) Node Feature Initialization: The initial feature vectors are defined per node type as B P V0 b,: = Concat {ψb,n,0 }n , (16) U V0 k,: = ψkU , V0R r,: = ψrR .
2) Feasibility-Guaranteed PA Placement: To guarantee that the learned PA positions satisfy constraints (15c) and (15d), the auxiliary inter-antenna spacing variables are introduced as
Notably, the reformulation allows ChanGNN to first yield {δb,n,m } and then reconstruct {xP b,n,m }. 3) CHAL and CFL in ChanGNN: We employ G1 CHALs to extract heterogeneous relational features from the location inputs. Denote the output feature matrices of BS and RIS BS RIS nodes of the G1 -th CHAL as VG and VG , respectively. 1 1 BS Then, we employ G2 CFLs and G3 CFLs to map VG and 1 RIS ′′ ′′ VG1 to VG2 and VG3 , respectively, which are further used RIS Node to generate {xP b,n,m } and {Φr }, respectively. UE Node
edges
Inter-BS
(a) Graph representation of hetero- (b) Graph representation of homogegeneous graph G1 for Stage 1. neous graph G2 for Stage 2 and 3.
(18)
BS-RIS Edge
RIS Node
m′ ∈M
δb,n,m′ ≤ δb,max , ∀b, n, m.
RIS-BS Edge
B rows
BS-RIS Edge
RIS-BS Edge
UE Node
Intra-BS edges
K columns
Intra-BS
BS Node
X
δb,n,m ≥ 0,
a) PA Position Output: For BS b, we first obtained δ b,n,m ′′ from VG as 2 ′′ δ b,n,m = Re VG . (19) 2 b,(n−1)M +m To ensure (18), we apply the sigmoid activation σ(·) and numerical scaling to δ b,n,m : δb,n,m = δb,max σ(δ b,n,m ),
δb,n,m :=
δb,n,m , δb,n,m δb,max , P ′ m′ ∈M δb,n,m
X
(20) δb,n,m′ ≤ δb,max
m′ ∈M
.
otherwise (21)
Then, PA positions are recovered via Xm xP δb,n,m′ + (m − 1)∆min , ∀b, n, m, (22) b,n,m = ′ m =1
satisfying constraints (15c) and (15d). b) RIS Phase Shift Output: To satisfy constraint (15f), ′′ each element of VG is normalized as 3 ′′ VG3 r,l jϕr,l (23) e = , ∀r, l, ′′ VG 3 r,l and Φr is constructed following (13).
C. Stage 2: BeamGNN The input and output of BeamGNN are the effective channel features computed with the obtained {xP b,n,m } and {Φr } and the beamforming vectors {wb,k }, respectively. 1) Pre-processing: With the obtained {xP b,n,m } and {Φr }, the effective channel for BS-UE node (b, k) is computed as P P H bH h xb,n,m , {Φr } = fb,k xb,n,m + (24) b,k X P P H hr,k Φr Hb,r xb,n,m Gb xb,n,m . r∈R
P δb,n,m = xP b,n,m − xb,n,m−1 − ∆min , ∀b, n, ∀m > 1, (17)
The feature of BS-UE node (b, k) is initialized with the stacked effective channel: P bH [V0′ ](b−1)K+k,: = h xb,n,m , {Φr } ∈ CBK×N . (25) b,k
with maximum available spacing δb,max = Cb −(M −1)∆min . Then, constraints (15c) and (15d) are equivalently expressed
2) HZM Learning: To reduce output dimensionality, we adopt the hybrid zero-forcing and maximum ratio transmission
δb,n,1 = xP b,n,1 ,
7
(HZM) learning [38], decomposing each beamforming vector as √ wb,k = pb,k wb,k (αb,k ), ∥wb,k (αb,k )∥2 = 1, (26) where pb,k ∈ R+ is the allocated power and αb,k ∈ [0, 1] is the hybrid coefficient. The unit-norm direction vector is wb,k (αb,k ) =
b q h + (1 − αb,k ) bb,k αb,k ∥qb,k b,k ∥ ∥hb,k ∥ q
b h ∥hb,k ∥
,
(27)
αb,k ∥qb,k + (1 − αb,k ) bb,k b,k ∥
H −1 where qb,k is the k-th column of Qb ≜ ZH with b (Zb Zb ) i h bH ; h bH ; . . . ; h b H ∈ CK×N . (28) Zb ≜ h b,1 b,2 b,K
Notably, the HZM learning reduces the required output dimension from KBN complex scalars to 2KB real scalars, enhancing both expressiveness and training efficiency. 3) CGAL and CFL in BeamGNN: Two parallel branches are employed to separately learn {αb,k } and {pb,k }. For hybrid coefficients, G4 CGAL layers followed by G5 CFL layers construct the mapping from updated features of ′′ ∈ CBK×1 , and the sigmoid activation BS-UE nodes to VG 5 is applied: ′′ , ∀b, k. (29) αb,k = σ Re [VG ] 5 (b−1)K+k,: For power allocation, G6 CGAL layers followed by G7 CFL ′′ ∈ layers construct the mapping from node features to VG 7 BK×1 C , yielding unconstrained power values: ′′ , ∀b, k. (30) peb,k = Pmax σ Re [VG ] 7 (b−1)K+k,: b b,k } are then recovUnconstrained beamforming vectors {w ered following (26) with {αb,k } and {e pb,k }. D. Stage 3: AssocGNN The input of AssocGNN is the (unconstrained) beamformb b,k }, and its output ing gains computed with the obtained {w is U. 1) Initial Node Features: For BS-UE node (b, k), its node feature is initialized by the beamforming gain: bH w [V0′ ](b−1)K+k,: = h b,k b b,k
2
, ∀k, b.
(31)
2) CGAL and CFL in AssocGNN: G8 CGAL layers fol′′ lowed by G9 CFL layers map V0′ to VG , from which the 9 association logits are extracted as ′′ ℓb,k = Re [VG ] , ∀k, b. (32) 9 (b−1)K+k,: 3) Differentiable Association via Gumbel-Softmax: To enable end-to-end differentiable training while satisfying constraints (15g) and (15h), the Gumbel-Softmax [39] is employed. During the training stage, we set exp ((ℓb,k + gb,k ) /τgs ) , ′ ′ b′ ∈B exp ((ℓb ,k + gb ,k ) /τgs )
ub,k = P
(33)
where gb,k ∼ Gumbel(0, 1) i.i.d. and τgs > 0 is the temperature. During the inference stage, the hard assignment
ub,k = 1[b = arg maxb′ ℓb′ ,k ] is used, satisfying (15g) and (15h) exactly. 4) Feasibility-Guaranteed Power Allocation: To enforce constraint (15e) for BS b, a per-BS normalization is applied to peb,k : pb,k := peb,k ,
(34) X
ub,k′ peb,k′ ≤ Pmax
k′ ∈K
, ∀b, peb,k Pmax , otherwise ′ eb,k ′ k′ ∈K ub,k p P such that k∈K ub,k ∥wb,k ∥2 ≤ Pmax holds for all BSs. Thus, the beamforming vectors {wb,k } are recovered following (26) with {αb,k } and {pb,k }. P
E. Detailed Processes of CHAL, CGAL, and CFL This subsection elaborates on the three layer types employed in our model, namely the CHAL, CGAL, and CFL. 1) Complex Heterogeneous Graph Attention Layer: The CHAL performs feature extraction via message passing over heterogeneous graph edges. To enrich representational capacity, a two-tier hierarchical attention mechanism is adopted: node-level attention captures the relative importance of neighboring nodes along a given meta-path, while semantic-level attention consolidates information across all meta-paths [40]. For the g-th (g ∈ {1, . . . , G}) CHAL, let VgΨ ∈ CMΨ ×Sg denote the output feature matrix for nodes of type Ψ ∈ {BS, UE, RIS}, where MΨ is the number of nodes of type Ψ and Sg is the per-node feature dimension. The input to Ψ the g-th CHAL is {Vg−1 }Ψ , where the initial features V0BS , UE RIS V0 , and V0 are exactly those defined in the Stage 1 input construction. Node-Level Attention. Each CHAL employs D parallel attention heads. For an ordered edge type connecting node types Ψ and Φ (Ψ, Φ ∈ {BS, UE, RIS}), the node-level attention coefficient matrix of the d-th head in the g-th layer, AΨ,Φ g,d ∈ RMΨ ×MΦ , is computed via (35), where RLeakyReLU(·) is Ψ the real-valued LeakyReLU activation, Wg,d ∈ CSg−1 ×Sg is Ψ,Φ 2Sg the feature transformation matrix, ag,d ∈ C is the attention weight vector, and NiΨ,Φ denotes the set of type Φ neighbors of the i-th type Ψ node. The d-th head aggregates weighted type Φ neighbor features for the i-th type Ψ node as X Ψ,Φ Ψ,Φ Φ Φ xg,i,d = CReLU [Ag,d ]i,j [Vg−1 ]j,: Wg,d , (36) j∈NiΨ,Φ
where CReLU(·) denotes the complex ReLU activation [41]. The D head outputs are then concatenated to form the intermediate feature matrix VgΨ,Φ ∈ CMΨ ×Sg : Ψ,Φ Ψ,Φ Ψ,Φ Ψ,Φ Vg = Concat x , x , . . . , x (37) g,i,1 g,i,2 g,i,D . i,: Semantic-Level Attention. To weight the contribution of each meta-path, semantic-level attention coefficients BgΨ,Φ ∈ R are computed as in (38), where Rtanh(·) is the real tanh activation, Wg ∈ CSg ×Sg and qg ∈ CSg are learnable
8
Φ T Ψ Ψ Φ exp RLeakyReLU Re aΨ,Φ Concat V W , V W g−1 i,: g−1 j,: g,d g,d g,d [AΨ,Φ g,d ]i,j = P Ψ,Φ T Ψ Ψ , VΦ Φ Vg−1 W W g−1 g,d g,d k∈NΨ,Φ exp RLeakyReLU Re ag,d Concat i,: k,:
(35)
T Ψ,Φ q · Rtanh Re V W g g g i∈VΨ i,: h BgΨ,Φ = i P P Ψ,Φ′ 1 T Vg Wg i∈VΨ qg · Rtanh Re Φ′ ∈NΨ exp |VΨ |
(38)
i
exp
1 |VΨ |
P
i,:
[Ag ]k,k′ = P
′ ′ exp RLeakyReLU Re aTg Concat [Vg−1 ]k,: Wg′ , [Vg−1 ]k′ ,: Wg′
k′′ ∈Nk exp
′ ′ RLeakyReLU Re aTg Concat [Vg−1 ]k,: Wg′ , [Vg−1 ]k′′ ,: Wg′
parameters, VΨ is the set of type Ψ nodes, and NΨ is the set of node types adjacent to type Ψ nodes. The output VgΨ is then obtained via weighted fusion across meta-paths with a residual connection to alleviate oversmoothing: X Ψ cΨ VgΨ = CReLU BgΨ,Φ VgΨ,Φ + Vg−1 Wg , (39) Φ∈NΨ
c Ψ ∈ CSg−1 ×Sg is the learnable residual matrix. where W g 2) Complex Graph Attention Layer: The CGAL captures inter-node interactions through the attention mechanism over a homogeneous graph, whereby each node assigns adaptive importance weights to its neighbors [42]. ′ ′ For the g-th CGAL with input features Vg−1 ∈ CBK×Sg−1 , the attention coefficient′ from′ node k ′ to node′ k is given by (40), where Wg′ ∈ CSg−1 ×Sg and ag ∈ C2Sg are learnable parameters, RLeakyReLU(·) denotes the real LeakyReLU activation, and Nk denotes the set of neighbors of node k. The updated node features are produced by aggregating attention-weighted neighbor representations: X ′ ′ ′ ′ Vg k,: = CReLU [A ] V W g k,k g−1 k′ ,: g . ′ k ∈Nk
(41) A residual connection is introduced to stabilize training and mitigate over-smoothing: ′ ′ ′ c ′ , (42) Vg k,: := Vg′ k,: + Vg−1 Wg + [V0′ ]k,: W g k,: ′
′
′
′
′
c ′ ∈ CS0 ×Sg are learnable where Wg ∈ CSg−1 ×Sg and W g residual matrices. 3) Complex Fully-Connected Layer: The CFL maps the embedding produced by the CHALs or CGALs to the target vector via feedforward operations adapted for complex-valued data. ′′ For the f -th CFL with input Vf′′−1 ∈ CMΨ ×Sf −1 , the output is computed as Vf′′ = CReLU Vf′′−1 Wf′′ + Bf , (43) ′′
′′
′′
where Wf′′ ∈ CSf −1 ×Sf and Bf ∈ CMΨ ×Sf are learnable weight matrices. Each CFL is followed by a complex batch
(40)
normalization layer [41] to improve convergence and reduce overfitting.
F. Unsupervised Loss Function Given input node location features, the proposed GNN sequentially produces the complete solution {xP b,n,m , {Φr }, {wb,k }, U}, enabling end-to-end unsupervised training by directly optimizing system utility. Denote all learnable parameters as Θ, encompassing the parameters of all CHALs, CGALs, and CFLs. The proposed GNN solves both P1 and P2 in a unified framework, differing only in the loss function. For Problem P1: LT (Θ) = T
1X T t=1 P
(t) k∈K Rk
1 n o . (44) xP b,n,m , {Φr }, {wb,k }, U Θ
For Problem P2: LT (Θ) = T
(t) (t) u ∥w ∥2 + PC b∈B n k∈K b,k b,k (t) xP k∈K Rk b,n,m , {Φr }, {wb,k }, U
1X T t=1 P
P
P
Θ
o . (45)
Notably, there are no penalty terms associated with constraints in (44) and (45), as all constraints are guaranteed to be satisfied. As summarized in Table I, the constraints in (15c)–(15h) are enforced by feasibility-preserving mappings rather than by soft penalties. Remark 1. (Scalability with the numbers of UEs, BSs and RISs) Among the learnable parameters in Θ, only the bias term Bf , is nominally dependent on the numbers of UEs, BSs, and RISs through its row dimension. Nevertheless, the MΨ row vectors of Bf can be set identically for any value of MΨ . Hence, all computations in the proposed GNN can be implemented independently of MΨ . By parameter sharing, proposed GNN is scalable to the numbers of UEs, BSs and RISs, ensuring that proposed GNN is acceptable to unseen problem sizes during both training and test phases.
9
TABLE I C ONSTRAINT ENFORCEMENT MECHANISMS IN THE PROPOSED THREE - STAGE GNN.
Constraint in (15)
Stage
(15c) 0 ≤ xP b,n,m ≤ Cb
Stage 1 Re-parameterize PA positions by auxiliary spacing variables δb,n,m , apply sigmoid scaling and per-waveguide normalization, and then reconstruct xP b,n,m Stage 1 Same spacing-based feasibility-preserving reconstruction as above Stage 3 Apply per-BS power normalization after association is obtained Stage 1 Normalize each complex output to unit modulus and construct Φr accordingly Stage 3 Use Gumbel-Softmax during training and hard one-hot assignment during inference Stage 3 Use argmax-based hard assignment at inference; differentiable relaxation is used only for training
P (15d) xP b,n,m − xb,n,m−1 ≥ ∆min
(15e)
2 k∈K ub,k ∥wb,k ∥ ≤ Pmax
P
(15f) ϕr,l ∈ [0, 2π) (15g)
P
b∈B ub,k = 1
(15h) ub,k ∈ {0, 1}
Enforcement mechanism
V. N UMERICAL R ESULTS This section provides numerical results to evaluate the proposed three-stage GNN. A. Simulation Setting 1) Simulation Scenario: The simulation scenario follows Figure 1. We consider B ∈ {1, 2, 3, 4} PA-BSs and R ∈ {1, 2, 3, 4} RISs deployed to serve K ∈ {2, . . . , 6, 10, . . . , 18} single-antenna UEs within a rectangular region of length D m and width S m. Each BS is equipped with N ∈ {8, 16, 20} parallel waveguides each hosting M ∈ {2, 3, 4, 5, 6} movable PAs, and each RIS comprises L ∈ {16, 64} reflecting elements. The total power budget is Pmax = 10 W, the circuit power is PC = 5 W, and the noise power is σk2 = −60 dBm. The simulation parameters are summarized in Table II. UEs are distributed uniformly within [0, D] × [0, S]. The B BSs are placed on a uniform grid within the same region, with each BS’s N waveguides offset evenly in the y-direction with spacing ∆wg and placed at height Hb , with feed points at the left end of each waveguide. The R RISs are similarly placed on a uniform grid at height Hb /2. Specifically, given B BSs (or R RISs), the grid dimensions Nrow and Ncol are jointly determined by solving (Nrow , Ncol ) =
arg min p,q: pq≥B, p,q∈Z+
q D − , p S
(46)
so that the grid aspect ratio Ncol /Nrow best matches the region aspect ratio D/S. The grid-center coordinates are then assigned as D S xc = c − 12 , yr = r − 21 , (47) Ncol Nrow where c ∈ {1, . . . , Ncol } and r ∈ {1, . . . , Nrow } index the grid columns and rows, respectively. 2) Baselines: The proposed system with jointly optimized PA positions, RIS phase shifts, beamforming vectors, and BSUE association is denoted Proposed GNN. To evaluate the
Key equations (17)–(22)
(17)–(22) (34) (13), (23) (33) (33)
TABLE II S IMULATION PARAMETERS . Parameter Number of BSs Number of RISs Number of waveguides per BS Number of PAs per waveguide Number of UEs Number of RIS elements Serving area Waveguide height Waveguide length Waveguide y-spacing Power budget Circuit power Noise power Carrier frequency Free-space wavelength Minimum PA spacing Effective refractive index In-waveguide attenuation Rician factor Path loss exponent Channel gain at 1 m
Value B ∈ {1, 2, 3, 4} R ∈ {1, 2, 3, 4} N ∈ {8, 16, 20} M ∈ {2, 3, 4, 5, 6} K ∈ {2, . . . , 6, 10, . . . , 18} L ∈ {16, 64} S = D ∈ {30, 40, 50, 60, 70} m Hb = 5 m C = 10 m ∆wg = 0.7 m Pmax = 10 W PC = 5 W σk2 = −60 dBm fc = 6 GHz λ = c/fc = 0.05 m ∆min = 0.1 m neff = 1.4 ζ = 0.0046 κ = 3 dB α = 2.8 β0 = −20 dB
contribution of each component, the following baselines are considered. System baselines: • No-RIS PA: The case without RIS assistance, where {Φr } is removed and the remaining other variables are jointly learned. • Fixed-PA: The case where PA positions are fixed at equal spacing ∆min on each waveguide, with {Φr }, {wb,k }, and U jointly learned. • No-RIS Fixed-PA: The case without RIS assistance and with fixed PA positions, serving as a conventional fixedposition antenna baseline. • Random-U: The BS-UE association matrix U is generated randomly at each inference, while {xP b,n,m }, {Φr }, and {wb,k } are jointly learned, isolating the gain from association optimization.
10
Model baselines: • MLP: A basic feedforward MLP with HZM learning that directly maps node location features to the complete solution, without exploiting graph structure. • Single HAN [40]: A single-stage complex heterogeneous graph attention network (HAN) with HZM learning, in which the UE nodes output {wb,k } and U, the RIS nodes output {Φr }, and the BS nodes output {xP b,n,m }. • HAN: A three-stage complex HAN with HZM learning, using the same node output design for UE, RIS, and BS nodes. • GAT [42]: A four-stage complex GAT with HZM learning, in which the system is modeled as a fully-connected homogeneous graph with K nodes. The input includes only UE positions, while the outputs {Φr } and {xP b,n,m } are obtained by averaging the embedding over all nodes [32]. 3) Training and Dataset: All learnable parameters are initialized using the Kaiming normal initialization method [43] with an initial learning rate of 5 × 10−5 . A multi-step learning rate scheduler adaptively reduces the learning rate during training. The Adam optimizer [44] is used for gradient-based updates over 50 epochs with a batch size of 128. An early stopping mechanism monitors validation performance and retains the parameter set yielding the best validation metric. Each training sample is generated by independently drawing K UE locations uniformly at random within the rectangular region, while BS and RIS locations are fixed across all samples according to the uniform grid placement described above. The NLoS fading components lNLoS ∼ CN(0, IL ) and hNLoS ∼ b,r r,k CN(0, IL ) are independently redrawn for each sample. Two data categories are constructed: a primary dataset of 100, 000 samples split into training, validation, and test subsets in an 8:1:1 ratio, and a generalization dataset of 10, 000 samples with unseen system configurations used exclusively for testing. 4) Computer Configuration: All models are trained and tested under Python 3.10 with PyTorch 1.11.0 on a computer equipped with an Intel Xeon Gold 6278C CPU and an NVIDIA Tesla V100 GPU (32 GB memory).
framework lies in jointly exploiting both types of reconfigurability rather than relying on only one of them. The RandomU baseline is also informative. In the single-BS case, it essentially coincides with the proposed GNN since association degenerates to a trivial decision. By contrast, in the multi-BS case, random association leads to a clear performance loss, which confirms that the proposed model effectively matches each UE to a more favorable serving BS. Table IV validates the effectiveness of the model design. The significant performance gain of the GNN methods over the MLP indicates that the data contain rich graph-topological information, which can be effectively exploited by GNNbased models. In the single-BS single-RIS scenario, GAT achieves performance comparable to that of the proposed GNN, showing that the proposed GNN is at least as effective as architectures specifically tailored to homogeneous-graph scenarios. In the multi-BS multi-RIS setting, however, neither GAT based on homogeneous graph representation nor HAN based on heterogeneous graph representation can match the performance of the proposed GNN. This is because the homogeneous graph representation adopted by GAT cannot effectively exploit the heterogeneous information among the three node types from the outset, while HAN, after leveraging heterogeneous information, still struggles to distinguish inter-BS competition and inter-UE interference. These results confirm the effectiveness of the proposed GNN in jointly incorporating both heterogeneous and homogeneous graph representations. Another important observation from Tables III and IV is that the proposed GNN exhibits scalability with respect to the number of UEs, which alleviates the burden of retraining deep learning models to accommodate dynamic wireless environments during deployment. Specifically, both EE and SR increase approximately linearly with the number of UEs, while exhibiting a tendency toward saturation. Furthermore, the scalability becomes more limited as the problem size grows larger. In addition, the inference times remain at the millisecond level as the problem size grows, showing that the proposed GNN is suitable for online deployment with limited latency budget.
B. Performance Comparison under Varying Problem Sizes Tables III and IV compare different methods under varying problem sizes and mismatched training/test UE numbers. The proposed GNN consistently delivers the strongest overall EE and SR performance, which verifies the effectiveness of the proposed graph construction and stage-wise decomposition in jointly coordinating PA placement, RIS configuration, active beamforming, and BS-UE association. Table III further demonstrates the effectiveness of the individual system components. Comparing the proposed GNN with No-RIS PA and No-RIS Fixed-PA shows that RIS reconfiguration and PA mobility are complementary. RISs provide environment-side controllability, whereas movable PAs improve transmitter-side spatial matching to the UE geometry. Once either component is removed, the performance degrades; when both are removed, the degradation becomes more evident. This confirms that the advantage of the proposed
C. Scalability with Respect to the Numbers of BSs and RISs Table V evaluates the scalability of the proposed GNN across different numbers of BSs and RISs. The diagonal entries achieve the best performance, as the training and test graph topologies are matched. Nevertheless, the off-diagonal results show that the proposed GNN still exhibits scalability to unseen (B, R) settings, especially when the topology mismatch is moderate. This demonstrates that the proposed architecture captures reusable structural relations among BSs, RISs, and UEs, rather than memorizing only one specific deployment scenario. At the same time, Table V also reveals that scalability across (B, R) is more challenging than generalization across the UE number K. Changing the numbers of BSs and RISs alters not only the node count, but also the coordination topology, the interference pattern, and the feasible association space.
11
TABLE III P ERFORMANCE COMPARISON OF DIFFERENT SYSTEM BASELINES UNDER VARYING PROBLEM SIZES .
B
1
R
1
N
8
L
M Ktr Kte
64
4
2 3 4 5 6
12
10 11 12 13 14
16
14 15 16 17 18
6
Inference time
4
4
16
16
6
Inference time
4
4
20
16
6
Inference time
Proposed GNN EE SR 4.12 31.55 5.81 44.69 7.29 56.74 8.48 67.28 9.38 76.16 0.81 ms 17.10 155.92 18.10 167.17 18.84 177.59 19.35 186.41 19.30 193.07 2.00 ms 22.41 209.95 23.07 219.80 23.54 228.47 23.63 239.75 23.17 234.11 2.56 ms
No-RIS PA EE SR 3.62 28.72 5.10 40.51 6.34 51.27 7.33 60.67 8.03 68.31 0.76 ms 15.98 149.02 16.89 160.10 17.61 170.02 17.91 178.32 17.83 184.09 1.97 ms 21.12 201.96 21.71 211.38 22.09 219.46 22.17 229.55 21.63 224.01 2.54 ms
No-RIS Fixed-PA EE SR 3.12 25.60 4.27 35.55 5.22 44.44 5.95 52.04 6.45 58.23 0.57 ms 13.66 143.13 14.11 153.21 14.44 162.25 14.48 169.95 14.06 175.26 1.73 ms 18.03 191.69 18.36 199.86 18.45 207.47 18.13 213.57 17.36 217.09 2.28 ms
Random-U EE SR 4.12 31.55 5.81 44.69 7.29 56.74 8.48 67.28 9.38 76.16 0.72 ms 14.31 127.49 15.04 136.10 15.54 143.78 15.84 149.59 15.74 153.19 1.64 ms 18.75 184.54 19.23 192.70 19.52 199.57 19.48 203.99 18.97 205.26 2.07 ms
KTr /KTe : Value of K in the training/test set.
19
Therefore, large topology mismatches lead to more visible performance degradation, which is particularly pronounced for EE.
180 175
E. Impact of Number of PAs Figure 4 shows that increasing the number of PAs improves both EE and SR for all compared schemes. The gain is largest when moving from a small number of PAs to a moderate one, and then gradually saturates, which is consistent with the diminishing marginal spatial DoF offered by additional elements. Moreover, the performance gap between the movablePA schemes and the fixed-PA counterpart becomes more visible as M increases, which indicates that a larger set of PAs allows the optimizer to exploit PA mobility more effectively.
Sum Rate (bit/s/Hz)
D. Architectural Ablation Table VI presents the ablation study on the main architectural components of the proposed GNN. Removing message passing leads to only a moderate performance loss, suggesting that the reconstructed effective channels already provide informative structural priors, while message passing further enriches the extraction of graph-topological information. In contrast, removing the residual connection causes a much more pronounced degradation, since deep graph layers are more vulnerable to over-smoothing without residual fusion. When the CFL in each stage is removed individually, the performance degrades in every case, and the degradation becomes even more significant when they are removed simultaneously. This confirms that the complex embedding decoding process is crucial in every stage of the joint optimization.
Energy Efficiency (bit/J/Hz)
18
17
16
15
14
170 165 160 155 150
Proposed GNN Fixed-PA 13
Proposed GNN Fixed-PA
145 2
3
4
5
6
2
Number of PAs M
3
4
5
6
Number of PAs M
Fig. 4. Impact of number of PAs on (B, R, N, L, K, S, D) = (4, 4, 16, 16, 12, 30, 30).
EE
and
SR,
with
VI. C ONCLUSION For coordinated downlink transmission in a multi-BS multiRIS-assisted PA system, we have presented an unsupervised three-stage GNN architecture, composed of ChanGNN, BeamGNN, and AssocGNN in cascade, for jointly optimizing
12
TABLE IV P ERFORMANCE COMPARISON OF DIFFERENT MODEL BASELINES UNDER VARYING PROBLEM SIZES .
B
1
R
1
N
8
L
M Ktr Kte
64
4
2 3 4 5 6
12
10 11 12 13 14
16
14 15 16 17 18
6
Inference time
4
4
16
16
6
Inference time
4
4
20
16
6
Inference time
Proposed GNN EE SR 4.12 31.55 5.81 44.69 7.29 56.74 8.48 67.28 9.38 76.16 0.81 ms 17.10 155.92 18.10 167.17 18.84 177.59 19.35 186.41 19.30 193.07 2.00 ms 22.41 209.95 23.07 219.80 23.54 228.47 23.63 239.75 23.17 234.11 2.56 ms
MLP EE SR × × × × 3.85 43.14 × × × × 0.03 ms × × × × 8.68 154.35 × × × × 0.21 ms × × × × 11.14 200.26 × × × × 0.25 ms
Single HAN EE SR 2.10 31.51 2.97 44.69 3.78 56.58 4.48 67.30 5.09 76.32 0.68 ms 13.14 142.95 13.80 153.19 14.16 162.29 14.37 169.85 14.17 174.83 0.83 ms 18.18 196.00 18.61 204.72 18.83 212.44 18.79 217.90 18.38 220.78 0.91 ms
HAN EE SR 2.12 31.51 3.00 44.75 3.81 56.70 4.52 67.37 5.12 76.38 1.19 ms 14.27 121.18 15.01 128.67 15.52 135.43 15.81 140.45 15.66 142.81 1.43 ms 18.76 160.50 19.29 166.19 19.50 171.47 19.44 173.84 18.99 173.68 1.58 ms
GAT EE SR 4.12 31.51 5.80 44.65 7.25 56.70 8.45 67.26 9.35 76.26 0.43 ms 14.26 146.52 15.02 156.45 15.52 166.29 15.78 174.07 15.66 179.03 0.83 ms 18.67 196.67 19.14 205.27 19.41 212.83 19.37 218.43 18.84 221.70 1.10 ms
× represents “not applicable”.
TABLE V S CALABILITY WITH RESPECT TO THE NUMBERS OF BS S AND RIS S .
EE (bit/J/Hz) (BTe , RTe ) (BTr , RTr ) (1, 1) (2, 2) (3, 3) (4, 4)
(1, 1)
(2, 2)
(3, 3)
(4, 4)
11.23 9.46 2.36 1.86
11.38 20.35 14.99 2.80
11.43 18.50 21.61 14.25
11.48 18.23 21.67 23.54
SR (bit/s/Hz) (BTe , RTe ) (BTr , RTr ) (1, 1) (2, 2) (3, 3) (4, 4)
(1, 1)
(2, 2)
(3, 3)
(4, 4)
168.42 113.59 106.15 108.16
184.59 194.13 181.99 159.85
194.55 214.50 215.10 210.08
199.61 227.05 227.91 228.47
(BTr , RTr )/(BTe , RTe ): Value of B and R in the training/test set.
PA positions, RIS phase shifts, transmit beamforming, and BS-UE association. By integrating heterogeneous and homogeneous graph representations, the proposed framework to effectively captures the structured interactions among BSs, RISs, and UEs. In addition, feasibility-preserving output mechanisms ensure valid PA placement, RIS phase normalization, power allocation, and BS-UE association throughout the end-to-end learning process. Extensive numerical results
TABLE VI A BLATION EXPERIMENT WITH (B, R, N, L, M, K, S, D) = (4, 4, 16, 16, 6, 12, 30, 30).
MP ✓ × ✓ ✓ ✓ ✓ ✓
RD ✓ ✓ × ✓ ✓ ✓ ✓
CFL1 ✓ ✓ ✓ × ✓ ✓ ×
CFL2 ✓ ✓ ✓ ✓ × ✓ ×
CFL3 ✓ ✓ ✓ ✓ ✓ × ×
EE 18.84 17.36 8.87 13.26 14.46 14.57 8.79
SR 177.59 172.08 128.30 170.23 155.80 144.74 132.61
MP/RD/CFL1/CFL2/CFL3: message passing/residual/CFL in Stage 1/CFL in Stage 2/CFL in Stage 3.
demonstrate the proposed framework’s superior performance over representative system and model baselines, together with favorable scalability and millisecond-level inference time. To the best of our knowledge, effective transmission design for PA systems under the challenging multi-BS multi-RIS setting considered in this work has not yet been reported in the open literature. Therefore, the proposed GNN provides a practically viable and pioneering solution for future wireless communications. R EFERENCES [1] C.-X. Wang et al., “On the road to 6G: Visions, requirements, key technologies, and testbeds,” IEEE Commun. Surv. Tutorials, vol. 25, no. 2, pp. 905–974, Second Quart., 2023.
13
[2] M. Di Renzo, A. Zappone, M. Debbah, M.-S. Alouini, C. Yuen, J. de Rosny, and S. Tretyakov, “Smart radio environments empowered by reconfigurable intelligent surfaces: How it works, state of research, and the road ahead,” IEEE J. Sel. Areas Commun., vol. 38, no. 11, pp. 2450– 2525, Nov. 2020. [3] E. Basar, M. Di Renzo, J. de Rosny, M. Debbah, M.-S. Alouini, and R. Zhang, “Wireless communications through reconfigurable intelligent surfaces,” IEEE Access, vol. 7, pp. 116753–116773, 2019. [4] Q. Wu, S. Zhang, B. Zheng, C. You, and R. Zhang, “Intelligent reflecting surface-aided wireless communications: A tutorial,” IEEE Trans. Commun., vol. 69, no. 5, pp. 3313–3351, May 2021. [5] Y. Liu, X. Liu, X. Mu, T. Hou, J. Xu, M. Di Renzo, and N. Al-Dhahir, “Reconfigurable intelligent surfaces: Principles and opportunities,” IEEE Commun. Surv. Tutorials, vol. 23, no. 3, pp. 1546–1577, Third Quart., 2021. [6] C. Huang, A. Zappone, G. C. Alexandropoulos, M. Debbah, and C. Yuen, “Reconfigurable intelligent surfaces for energy efficiency in wireless communication,” IEEE Trans. Wireless Commun., vol. 18, no. 8, pp. 4157–4170, Aug. 2019. [7] X. Guan, Q. Wu, and R. Zhang, “Intelligent reflecting surface assisted secrecy communication: Is artificial noise helpful or not?” IEEE Wireless Commun. Lett., vol. 9, no. 6, pp. 778–782, Jun. 2020. [8] X. Yu, D. Xu, and R. Schober, “Robust and secure wireless communications via intelligent reflecting surfaces,” IEEE J. Sel. Areas Commun., vol. 38, no. 11, pp. 2637–2652, Nov. 2020. [9] C. Pan et al., “Multicell MIMO communications relying on intelligent reflecting surfaces,” IEEE Trans. Wireless Commun., vol. 19, no. 8, pp. 5218–5233, Aug. 2020. [10] L. Zhu, W. Ma, and R. Zhang, “Movable antennas for wireless communication: Opportunities and challenges,” IEEE Commun. Mag., vol. 62, no. 6, pp. 114–120, Jun. 2024. [11] W. K. New et al., “A tutorial on fluid antenna system for 6G networks: Encompassing communication theory, optimization methods and hardware designs,” IEEE Commun. Surv. Tutorials, vol. 27, no. 4, pp. 2325– 2377, Fourth Quart., 2025. [12] Y. Liu, H. Jiang, X. Xu, Z. Wang, J. Guo, C. Ouyang, X. Mu, Z. Ding, A. Nallanathan, G. K. Karagiannidis, and R. Schober, “Pinching-antenna systems (PASS): A tutorial,” IEEE Trans. Commun., vol. 74, pp. 48814918, 2026. [13] Z. Ding, R. Schober, and H. V. Poor, “Flexible-antenna systems: A pinching-antenna perspective,” IEEE Trans. Commun., vol. 73, no. 10, pp. 9236–9253, Oct. 2025. [14] L. Zhu, W. Ma, and R. Zhang, “Modeling and performance analysis for movable antenna enabled wireless communications,” IEEE Trans. Wireless Commun., vol. 23, no. 6, pp. 6234–6250, Jun. 2024. [15] L. Zhu, W. Ma, B. Ning, and R. Zhang, “Movable-antenna enhanced multiuser communication via antenna position optimization,” IEEE Trans. Wireless Commun., vol. 23, no. 7, pp. 7214–7229, Jul. 2024. [16] W. Ma, L. Zhu, and R. Zhang, “MIMO capacity characterization for movable antenna systems,” IEEE Trans. Wireless Commun., vol. 23, no. 4, pp. 3392–3407, Apr. 2024. [17] W. Mei, X. Wei, B. Ning, Z. Chen, and R. Zhang, “Movable-antenna position optimization: A graph-based approach,” IEEE Wireless Commun. Lett., vol. 13, no. 7, pp. 1853–1857, Jul. 2024. [18] X. Wei, W. Mei, Q. Wu, Q. Jia, B. Ning, Z. Chen, and J. Fang, “Movable antennas meet intelligent reflecting surface: Friends or foes?” IEEE Trans. Commun., vol. 73, no. 11, pp. 12756–12770, Nov. 2025. [19] E. Björnson, M. Bengtsson, and B. Ottersten, “Optimal multiuser transmit beamforming: A difficult problem with a simple solution structure,” IEEE Signal Process. Mag., vol. 31, no. 4, pp. 142–148, Jul. 2014. [20] Y. Shi, L. Lian, Y. Shi, Z. Wang, Y. Zhou, L. Fu, L. Bai, J. Zhang, and W. Zhang, “Machine learning for large-scale optimization in 6G wireless networks,” IEEE Commun. Surv. Tutorials, vol. 25, no. 4, pp. 2088–2132, Fourth Quart., 2023. [21] C. Zhang, P. Patras, and H. Haddadi, “Deep learning in mobile and wireless networking: A survey,” IEEE Commun. Surv. Tutorials, vol. 21, no. 3, pp. 2224–2287, Third Quart., 2019. [22] T. Lin and Y. Zhu, “Beamforming design for large-scale antenna arrays using deep learning,” IEEE Wireless Commun. Lett., 2019, doi: 10.1109/LWC.2019.2943466. [23] J. M. J. Huttunen, D. Korpi, and M. Honkala, “DeepTx: Deep learning beamforming with channel prediction,” IEEE Trans. Wireless Commun., vol. 22, no. 3, pp. 1855-1867, Mar. 2023. [24] J.-M. Kang, “Deep learning enabled multicast beamforming with movable antenna array,” IEEE Wireless Commun. Lett., vol. 13, no. 7, pp. 1848–1852, Jul. 2024.
[25] Y. Shen, Y. Shi, J. Zhang, and K. B. Letaief, “Graph neural networks for scalable radio resource management: Architecture design and theoretical analysis,” IEEE J. Sel. Areas Commun., vol. 39, no. 1, pp. 101–115, Jan. 2021. [26] Y. Shen, J. Zhang, S. H. Song, and K. B. Letaief, “Graph neural networks for wireless communications: From theory to practice,” IEEE Trans. Wireless Commun., vol. 22, no. 5, pp. 3554–3569, May 2023. [27] Y. Li, Y. Lu, B. Ai, O. A. Dobre, Z. Ding, and D. Niyato, “GNN-based beamforming for sum-rate maximization in MU-MISO networks,” IEEE Trans. Wireless Commun., vol. 23, no. 8, pp. 9251–9264, Aug. 2024. [28] Q. Wang, Y. Lu, W. Chen, B. Ai, Z. Zhong, and D. Niyato, “GNNenabled optimization of placement and transmission design for UAV communications,” IEEE Trans. Veh. Technol., vol. 73, no. 11, pp. 16789– 16802, Nov. 2024. [29] C. He, Y. Lu, W. Chen, B. Ai, K.-K. Wong, and D. Niyato, “Graph neural network enabled fluid antenna systems: A two-stage approach,” IEEE Trans. Veh. Technol., vol. 74, no. 10, pp. 16625–16629, Oct. 2025. [30] X. Xie, Y. Lu, and Z. Ding, “Graph neural network enabled pinching antennas,” IEEE Wireless Commun. Lett., 2025, doi: 10.1109/LWC.2025.3584919. [31] J. Guo, Y. Liu, and A. Nallanathan, “GPASS: Deep learning for beamforming in pinching-antenna systems (PASS),” arXiv preprint arXiv:2502.01438, 2025. [32] C. He, Y. Lu, Y. Xu, C.-Y. Chi, B. Ai, and A. Nallanathan, “RISassisted downlink pinching-antenna systems: GNN-enabled optimization approaches,” arXiv preprint arXiv:2511.20305, 2025. [33] S. Liu, P. Ni, R. Liu, Y. Liu, M. Li, and Q. Liu, “BS-RIS-user association and beamforming designs for RIS-aided cellular networks,” in Proc. IEEE/CIC Int. Conf. Commun. China (ICCC), Xiamen, China, 2021, pp. 563-568. [34] A. Bereyhi, S. Asaad, C. Ouyang, Z. Ding, and H. V. Poor, “Downlink beamforming with pinching-antenna assisted MIMO systems,” in Proc. IEEE Int. Conf. Commun. Workshops (ICC Workshops), Montreal, QC, Canada, 2025, pp. 1-6. [35] X. Xu, X. Mu, Y. Liu, and A. Nallanathan, “Joint transmit and pinching beamforming for pinching antenna system (PASS): Optimization-based or learning-based?,” IEEE Trans. Wireless Commun., vol. 25, pp. 1144911464, 2026. [36] H. A. Le, T. Van Chien, and W. Choi, “Graph neural network-based active and passive beamforming for distributed STAR-RIS-assisted multiuser MISO systems,” IEEE Trans. Commun., vol. 73, no. 10, pp. 92999312, Oct. 2025. [37] M. Liu, C. Huang, A. Alhammadi, M. Di Renzo, M. Debbah, and C. Yuen, “Beamforming design and association scheme for multi-RIS multi-user mmWave systems through graph neural networks,” IEEE Trans. Wireless Commun., vol. 24, no. 9, pp. 7940-7954, Sept. 2025. [38] W. Guo et al., “Hybrid MRT and ZF learning for energy-efficient transmission in multi-RIS-assisted networks,” IEEE Trans. Veh. Technol., vol. 73, no. 8, pp. 12247-12251, Aug. 2024. [39] E. Jang, S. Gu, and B. Poole, “Categorical reparameterization with Gumbel-Softmax,” in Proc. ICLR, 2017, pp. 1920-1931. [40] X. Wang et al., ‘Heterogeneous graph attention network,” in Proc. ACM WWW, pp. 2022-2032, 2019. [41] C. Trabelsi, O. Bilaniuk, Y. Zhang, D. Serdyuk, S. Subramanian, J. F. Santos, S. Mehri, N. Rostamzadeh, Y. Bengio, and C. J Pal, “Deep complex networks,” in Proc. ICLR, 2018, pp.1-19. [42] P. Velickovic, G. Cucurull, A. Casanova, A. Romero, P. Lio, and Y. Bengio, ‘Graph attention networks,” in Proc. ICLR, pp. 1–12, 2018. [43] K. He, X. Zhang, S. Ren, and J. Sun, “Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification,” in Proc. ICCV, pp. 1026-1034, 2015. [44] D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Proc. ICLR, pp. 1–15, Feb. 2015.