ConceptioArchivearXiv CS
arXiv CSopen access

CALO: Constraint-Aware Learning Optimization for Joint Resource Allocation in Double-Active RIS-Assisted Wireless Networks

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
distributedsystemsprotocols
networking, internet, protocols, distributed systems

IEEE TRANSACTIONS ON COGNITIVE COMMUNICATIONS AND NETWORKING

1

CALO: Constraint-Aware Learning Optimization for Joint Resource Allocation in Double-Active RIS-Assisted Wireless Networks

arXiv:2606.30803v1 [cs.NI] 29 Jun 2026

Alaa S. Arabiyat and Mohammad J. Abdel-Rahman, Senior Member, IEEE

Abstract—Double-active reconfigurable intelligent surface (RIS)-assisted wireless systems can improve coverage and achievable rate in blockage-dominated environments. Still, their joint resource allocation is challenging due to the coupling among RIS placement, amplification power allocation, and reflecting-element assignment. The resulting problem is linearly constrained, nonconvex, and involves both continuous and discrete variables, making conventional iterative solvers such as block coordinate descent (BCD) computationally expensive for real-time deployment. This paper proposes a constraint-aware learning optimization (CALO) framework for data-driven joint resource allocation in double-active RIS-assisted networks. CALO reformulates the decision variables into grouped fractional representations and maps them to physical resources through constraint-preserving transformations, ensuring that distance, power, and elementbudget constraints are satisfied by construction. A straightthrough estimator is incorporated to enable differentiable learning over discrete reflecting-element assignments, while a regretdriven hinge objective uses the BCD solution as a reference and encourages performance improvement beyond solver imitation. Simulation results show that CALO achieves 100% feasibility across all tested configurations, improves the achievable rate over BCD in both urban and rural scenarios, and reduces online inference time by orders of magnitude. These results demonstrate the effectiveness of structure-aware learning for feasible and realtime optimization in active multi-RIS wireless systems. Index Terms—Double-active RIS, resource allocation, constraint-aware learning, learning-based optimization, nonconvex optimization.

I. I NTRODUCTION

T

HE sixth generation (6G) of wireless communication systems is expected to support unprecedented requirements in terms of data rate, latency, energy efficiency, reliability, and network adaptability. These requirements are driven by emerging applications such as digital twins [1], [2], holographic communications [3], [4], and ultra-reliable low-latency communication (URLLC) [5], [6]. To meet these demands, higher frequency bands, including sub-THz and THz bands, have been widely investigated due to their abundant spectral resources. However, communication at these frequencies is highly vulnerable to severe path loss, blockage, and strong line-of-sight (LoS) dependency, which limits coverage and link reliability in practical deployments. A. S. Arabiyat is with the Department of Data Science, Princess Sumaya University for Technology, Amman 11941, Jordan (email:[email protected] ). M. J. Abdel-Rahman is with the Department of Data Science, Princess Sumaya University for Technology, Amman 11941, Jordan. He is also with the Department of Electrical and Computer Engineering, Virginia Tech, Blacksburg, VA 24061, USA (e-mail: [email protected]). This work was supported by Jordan Scientific Research and Innovation Support Fund (SRISF) under Grant ICT/1/21/2021.

Reconfigurable intelligent surfaces (RISs) have emerged as a promising technology to address these propagation challenges by enabling programmable control of the wireless environment. By adjusting the phase and amplitude response of incident electromagnetic waves, RISs can enhance signal strength, improve coverage, and increase spectral and energy efficiency [7], [8]. While early RIS studies mainly considered passive single-surface deployments, recent research has moved toward more advanced architectures, including active RISs, multi-RIS systems, and double-reflection links. These architectures provide additional degrees of freedom for controlling signal propagation and can offer substantial performance gains over single-RIS systems [9]–[11]. Among these architectures, double-active RIS systems are particularly attractive in blockage-dominated environments, where the direct transmitter-receiver link and single-reflection paths may be unavailable. By employing two active surfaces, the system can exploit a cascaded double-reflection link while compensating for severe attenuation through signal amplification. However, these benefits introduce a more challenging resource-allocation problem. The achievable rate depends jointly on RIS placement, amplification power allocation, and reflecting-element assignment. These variables are strongly coupled through the cascaded channel and are subject to strict physical constraints, including total distance, total amplification power, and total number of reflecting elements. The resulting optimization problem is non-convex, mixed continuous-discrete, and linearly constrained. Classical optimization methods, such as alternating optimization (AO) and block coordinate descent (BCD), address this problem by decomposing it into tractable subproblems and iteratively updating the decision variables [12], [13]. Although such methods can provide high-quality reference solutions, they require repeated optimization for every system configuration. Their computational complexity increases with the number of RIS elements, system parameters, and optimization blocks, making them difficult to deploy in real-time or dynamic wireless environments. Learning-based optimization has recently been explored as an alternative to iterative wireless resource allocation. Neural networks can approximate the mapping from system parameters to resource-allocation decisions and provide low-latency inference after offline training. However, existing learningbased approaches still face two major limitations. First, many methods treat resource allocation as a black-box prediction problem and do not guarantee strict feasibility with respect to equality, budget, or integer constraints. Second, supervised learning approaches often imitate solutions generated by classical solvers, which may cause the learned model to inherit the

IEEE TRANSACTIONS ON COGNITIVE COMMUNICATIONS AND NETWORKING

suboptimality of the reference algorithm rather than improve upon it. To address these limitations, this paper proposes a constraint-aware learning optimization (CALO) framework for joint resource allocation in double-active RIS-assisted wireless systems. Instead of directly predicting unconstrained physical variables, CALO reformulates the problem using fractional grouped decision variables. These variables are mapped to physical resources through constraint-preserving transformations, ensuring that the placement, power-allocation, and element-assignment constraints are satisfied by construction. Discrete reflecting-element allocation is handled through a straight-through estimator (STE), enabling end-to-end training despite the presence of integer-valued variables. Moreover, a regret-driven hinge objective is introduced to use BCD as a performance reference while encouraging the learned solution to exceed the baseline achievable rate. From a broader perspective, the proposed framework is not limited to a single double-active RIS configuration. Rather, it represents a learning-based solver for a class of linearly constrained non-convex wireless resource-allocation problems with grouped continuous and discrete variables. By embedding the structure of the optimization problem into the neural architecture, CALO combines the efficiency of single-shot inference with strict feasibility guarantees and direct performance optimization. Our Contributions: After formulating the joint optimization of RIS placement, amplification power allocation, and reflecting-element assignment in double-active RIS-assisted wireless systems as a grouped, linearly constrained, nonconvex optimization problem with mixed continuous and discrete variables, our main contributions of this paper can be summarized as follows: Feasibility-by-construction neural parameterization: We develop a fractional resource-allocation parameterization based on grouped normalization, where the predicted variables are mapped to physical resources while satisfying the distance, power, and element-budget constraints by construction. • End-to-end handling of discrete element allocation: We incorporate an STE to enable differentiable learning over integer reflecting-element assignments, allowing continuous and discrete decision variables to be optimized within a unified neural framework. • Regret-driven learning beyond solver imitation: We introduce a hinge-style regret objective that uses BCDgenerated solutions as reference points while penalizing the model only when its achievable rate falls below the baseline, thereby encouraging performance improvement rather than direct imitation.

We show through extensive simulations that the proposed CALO framework achieves strictly feasible solutions, statistically significant achievable-rate gains over BCD, and ordersof-magnitude reduction in online inference time across heterogeneous deployment scenarios. Paper Organization: The remainder of this paper is organized as follows. Section II reviews the related literature.

2

Section III presents the system model and problem formulation. Section IV details the proposed CALO framework. Section V reports the simulation results and performance analysis. Finally, Section VI concludes the paper. II. R ELATED W ORK RISs have emerged as a key technology for beyond-5G and 6G networks by enabling programmable control of the wireless propagation environment. Early studies mainly considered passive single-RIS architectures, where transmit beamforming and RIS phase-shift design are the main optimization variables. However, passive RISs suffer from severe multiplicative path loss in cascaded transmitter-RIS-receiver links, especially in high-frequency, long-distance, and blockagedominated scenarios. To mitigate this limitation, active RIS architectures have been proposed, where RIS elements provide signal amplification in addition to phase control. While active RISs can compensate for propagation loss, they introduce amplification noise, hardware complexity, and power-budget constraints [14], [15]. Therefore, power allocation becomes a key design dimension in active RIS systems [16]. More recently, multi-RIS and cooperative RIS architectures have been investigated to exploit additional reflection paths and spatial degrees of freedom in complex environments [17], [18]. In double-active RIS systems, this coupling becomes more pronounced because two active surfaces jointly determine the cascaded signal path, amplification noise, power distribution, and reflecting-element allocation. Existing works have studied deployment optimization, active element allocation, and total power-constrained designs [19], [20], providing the main architectural foundation for the system considered in this paper. Resource allocation in RIS-assisted systems is generally non-convex due to the coupling among beamforming, RIS coefficients, deployment variables, power allocation, user association, and hardware constraints. This challenge is further amplified in active and multi-RIS systems, where power budgets, noise enhancement, multiple reflecting surfaces, and distributed channels introduce additional coupled variables. Conventional optimization methods, including AO, BCD, weighted minimum mean-square error (WMMSE), successive convex approximation, semidefinite relaxation, and manifold optimization, remain dominant for solving such non-convex wireless design problems [21], [22]. For example, WMMSEbased designs have been applied to multi-RIS-assisted cellfree networks [17], while robust and secure STAR-RIS systems commonly rely on iterative optimization to handle imperfect CSI, secrecy constraints, and coupled beamforming variables [23]. Although these methods provide strong reference solutions, they require repeated per-instance optimization, which limits scalability and real-time applicability, particularly in active and double-active RIS systems [19], [20]. Learning-based optimization has recently emerged as an alternative for reducing online computational complexity in wireless resource allocation. Neural models can learn mappings from system parameters to resource-allocation decisions, enabling fast inference after offline training. Recent surveys highlight the growing role of supervised learning, reinforcement learning, deep reinforcement learning, and graph-based

IEEE TRANSACTIONS ON COGNITIVE COMMUNICATIONS AND NETWORKING

learning in wireless resource management [24], [25]. In RISassisted systems, learning-based methods have been applied to phase-shift design, beamforming, channel estimation, and resource allocation. However, many existing approaches follow black-box prediction or policy-learning paradigms and do not explicitly guarantee feasibility with respect to equality, powerbudget, integer, or hardware constraints. Moreover, supervised models trained only to imitate AO/BCD solutions may inherit the suboptimality of the reference solver. Recent works comparing optimization and machine-learning solutions for advanced RIS architectures highlight the promise of learningbased methods, but also reveal the need for stronger integration of problem structure into the learning model [23], [26]. In contrast, the proposed CALO framework embeds the linear resource-allocation constraints directly into the neural parameterization. It maps grouped fractional variables through simplex-preserving transformations to satisfy placement, power-allocation, and reflecting-element budget constraints by construction, handles discrete element allocation using a straight-through estimator, and adopts a regret-driven hinge objective that encourages performance improvement beyond BCD imitation. Hence, CALO is tailored to linearly constrained non-convex RIS resource-allocation problems involving both continuous and discrete variables. III. S YSTEM M ODEL AND P ROBLEM F ORMULATION A. General Problem Structure and Motivation The framework developed in this paper targets a class of wireless resource allocation problems that can be expressed as structured linearly constrained non-convex optimization programs with coupled continuous and discrete decision variables. A generic form of this class can be written as: max

f (z; u) X s.t. zj = Bm ,

(1)

z

∀m ∈ M,

(2)

j∈Gm

zj ≥ 0,

∀j,

(3)

zk ∈ Z,

∀k ∈ I,

(4)

where z denotes the vector of decision variables, u collects the system parameters, {Gm }m∈M represents groups of variables subject to linear budget constraints, and I denotes the subset of integer-valued variables. The objective function f (z; u) is generally non-convex due to multiplicative coupling among variables, cascaded signal transformations, and nonlinear dependence on system parameters. The constraints in (2) define a set of affine budget constraints that induce a structured feasible region, where decision variables are organized into groups sharing limited resources such as power, distance, or hardware elements. This class of problems frequently arises in advanced wireless systems, where multiple interdependent resources must be jointly optimized under physical constraints. Their nonconvexity, variable coupling, and mixed decision spaces make conventional optimization methods computationally demanding, especially in large-scale or real-time settings. The double-active RIS-assisted system considered in this paper is a representative instance of the problem class defined

3

Fig. 1. Double-active RIS-assisted wireless communication system.

in (1)–(4). The joint optimization of RIS placement, amplification power allocation, and reflecting-element assignment gives rise to grouped variables under linear resource constraints and a strongly coupled non-convex objective induced by multistage signal propagation and amplification. This formulation exposes a structured optimization template that facilitates feasibility-preserving parameterizations and scalable learningbased inference, allowing the proposed framework to serve as a general solver for a broader class of linearly constrained non-convex optimization problems. B. System Model We consider a downlink wireless communication system assisted by two cooperative active RISs, as illustrated in Fig. 1. A base station (BS) equipped with M antennas communicates with a single-antenna user through a double-reflection link formed by RIS 1 and RIS 2. This system setup follows the model in [27], while being adopted here as a representative structure to expose the underlying optimization challenges. 1) Network Geometry: The BS, RIS 1, RIS 2, and the user are deployed along a horizontal line. Let x1 , x2 , and x3 denote the horizontal distances of the BS → RIS 1, RIS 1 → RIS 2, and RIS 2 → user links, respectively, satisfying x1 + x2 + x3 = D, where D is the total BS-user distance. Both RISs are deployed at a fixed height H. The total number of reflecting elements is N , which is partitioned as N1 +N2 = N, N1 , N2 ∈ N. We focus on a challenging scenario where the direct BSuser link and single-reflection paths are blocked, and only the cascaded BS → RIS 1 → RIS 2 → user link is available, as commonly considered in multi-RIS systems [27]. 2) Channel Model: To provide clear structural insights, we assume that all channels have a direct LoS, following [27]. The BS-to-RIS 1 channel is HB = hB ar aH t , where hB = √ β −j 2π λ d1 represents the path-loss and phase shift, with d = e 1 d1 p x21 + H 2 . Similarly, the inter-RIS channel HI ∈ CN2 ×N1 1×N2 and the RIS 2-to-user channel hH are modeled using U ∈C the same LoS structure. 3) Active RIS Model: Each RIS performs both reflection and amplification. The reflection matrix of RIS i ∈ {1, 2} is

IEEE TRANSACTIONS ON COGNITIVE COMMUNICATIONS AND NETWORKING

 Θi = diag ai ejϕi,1 , . . . , ai ejϕi,Ni , where ai is the amplification factor and ϕi,n is the phase shift of the n-th element. Due to hardware constraints and under the LoS assumption, all elements within each RIS share a common amplification factor, as justified in [27]. The amplification process introduces noise ni ∼ CN (0, δi2 I), i ∈ {1, 2} at RIS i, while the receiver noise is n0 ∼ CN (0, δ02 ). The BS applies a beamforming vector w ∈ CM satisfying ∥w∥2 ≤ 1. 4) Received Signal Model: The received signal at the user is expressed as: y = hH U a2 Θ2 (HI a1 Θ1 HB ws + HI a1 Θ1 n1 + n2 ) + n0 , where s denotes the transmitted symbol with power PB . This expression captures the cascaded signal propagation and the accumulation of amplification noise across RIS stages. 5) Achievable Rate: The signal-to-noise ratio (SNR) at the user can be written as: 2 PB hH U a2 Θ2 HI a1 Θ1 HB w . γ= 2 H 2 2 δ1 ∥hU a2 Θ2 HI a1 Θ1 ||2 + δ22 ∥hH U a2 Θ2 || + δ0 The corresponding achievable rate is R = log2 (1 + γ). 6) Remarks on Model Generality: Although the above model follows a LoS assumption for analytical tractability, the formulation captures the essential characteristics of multistage RIS-assisted communication systems. The structure can be extended to more general channel models without altering the underlying optimization framework. C. Problem Formulation Let P1 and P2 denote the amplification power allocated to RIS 1 and RIS 2, respectively, subject to a total amplification power budget. The objective is to maximize the achievable downlink rate by jointly optimizing the BS beamforming vector, RIS configurations, amplification factors, element allocation, power allocation, and placement variables. Accordingly, the joint optimization problem is formulated as: (P1) s.t.

max

log2 (1 + γ)

w,Θ,a,N,P,X 2

(5)

∥w∥ ≤ 1,

Θi = diag ai e

(6) jϕi,1

, . . . , ai e

jϕi,Ni



,

i ∈ {1, 2}, (7)

N1 + N2 = N,

N1 , N2 ∈ N,

(8)

P1 + P2 = PF ,

P1 , P2 ≥ 0,

(9)

xi ≥ 0,

(10)

x1 + x2 + x3 = D, a21 P1 (·) ≤ P1 ,

(11)

a22 P2 (·) ≤ P2 .

(12)

where Θ = {Θ1 , Θ2 }, a = {a1 , a2 }, N = {N1 , N2 }, P = {P1 , P2 }, and X = {x1 , x2 , x3 }. The functions P1 (·) and P2 (·) represent the received signal power at RIS 1 and RIS 2, respectively, and are directly derived from the signal model. These constraints ensure that the amplification power at each RIS does not exceed its allocated budget. Problem P1 is highly challenging due to several intrinsic properties: • Non-convex objective: The achievable rate depends on the SNR, which involves multiplicative coupling across beamforming, RIS configurations, amplification, and placement variables.

4

Mixed discrete-continuous variables: The element allocation variables {N1 , N2 } are integer-valued, while the remaining variables are continuous. • Strong variable coupling: All decision variables jointly affect both signal power and noise propagation due to the cascaded structure of the system. • Coupled constraints: The amplification constraints depend on multiple variables simultaneously, making the feasible region highly non-trivial. These characteristics render P1 a linearly constrained nonconvex optimization problem with mixed variables. Although P1 involves multiple coupled decision variables, it admits a structured decomposition that significantly reduces its effective dimensionality. In particular, following the approach in [27], it can be shown that for any given element allocation, power allocation, and placement, the optimal BS beamforming vector, RIS phase shifts, and amplification factors can be obtained in closed form. •

D. Problem Reformulation and Structural Reduction 1) Closed-Form Elimination of Auxiliary Variables: Following [27], for any feasible element allocation N, amplification power allocation P, and placement X, the auxiliary variables, namely the BS beamforming vector w, RIS phase shifts Θ1 , Θ2 , and amplification factors a1 , a2 ,admit closedform optimal solutions. t , M) αt (θB , w∗ = t ∥αt (θB , M )∥       t t r r  ∗ −j ∠ αH t (θI ,ϑI ,N1 ) n +∠ αr (θB ,ϑB ,N1 ) n 1 1 , Θ1 n1 = e       r r  ∗ −j ∠ hH U n +∠ αr (θI ,ϑI ,N2 ) n 2 2 , Θ2 n2 = e s P1 ∗ a1 = 2 + δ2 ) , N1 (M PB ηB 1

s a∗2 =

2) P2 (δ12 + M PB ηB 2 (δ 2 + N P η 2 )) , N2 (δ12 (δ22 + P1 ηI2 ) + M PB ηB 1 1 I 2

where

β , ηI = , ηU = . d1 d2 d3 These results imply that w, Θ1 , Θ2 , a1 , and a2 can be treated as auxiliary variables and eliminated via substitution. 2) Equivalent SNR Expression: Substituting the optimal auxiliary variables into the received signal model yields an equivalent SNR expression: 2 2 2 ηI ηU N12 N22 M PB ηB γ= 2 η2 2 . N δ 2 + 2 2 U + δ0 N1 N2 δ12 ηI2 ηU 2 2 a a a2 ηB =

β

β

1

1 2

Further substituting a∗1 and a∗2 leads to a reduced-form expression: M PB β 3 γ = E1 , E2 E3 E4 E5 E6 N1 N2 + N2 P2 + N1 N2 P1 + N1 N2 P1 + N1 N2 P2 + N1 N2 P1 P2 where E1 , . . . , E6 are constants determined by system parameters (path loss, transmit power, and noise levels).

IEEE TRANSACTIONS ON COGNITIVE COMMUNICATIONS AND NETWORKING

Stage 1

Stage 2

Stage 3

Input Features

MLP Predictor

Grouped Softmax Mapping

timization with a direct learning-based solution. Specifically, CALO learns a mapping: Fθ : u → (X, P, N),

Backpropagation Update neural network weights

Stage 4

Stage 5

Physical Variable Mapping

STE for Discrete Variable

Compute Objective

Margin-Based Loss

Fig. 2. Overall proposed learning framework. The neural network predicts grouped fractional allocations, which are transformed through softmax-based heads, mapped into physical variables, and refined using an STE for discrete variables. The objective and margin-based loss are then computed, and the network parameters are updated via backpropagation.

3) Reduced Problem Formulation: Therefore, P1 is equivalently reduced to: (P2)

max

γ(N, P, X)

s.t.

N1 + N2 = N,

N1 , N2 ∈ N,

(14)

P1 + P2 = PF ,

P1 , P2 ≥ 0,

(15)

xi ≥ 0.

(16)

N,P,X

(13)

x1 + x2 + x3 = D,

5

This reformulation significantly reduces the dimensionality of the problem while preserving its non-convex nature. 4) Structural Interpretation: The reduced problem reveals a grouped structure where decision variables are partitioned into three independent resource allocation groups: Element allocation {N1 , N2 }, power allocation {P1 , P2 }, and placement {x1 , x2 , x3 }. Each group is constrained by a linear equality (budget constraint), forming a simplex domain. This structure aligns with the generic formulation introduced in Section II-A. 5) Limitations of Conventional Optimization: The reference solution applies BCD to iteratively optimize each variable group. However, such approaches suffer from high computational complexity, limited scalability, and a lack of real-time applicability in dynamic wireless environments. 6) Transition to Learning-Based Reformulation: To overcome these limitations, we reformulate the decision variables using fractional representations with simplex constraints. This transformation guarantees feasibility by construction and converts P2 into an unconstrained optimization problem over simplex domains. This reformulation enables the design of a learning-based framework that directly predicts optimal resource allocation in a single forward pass. IV. CALO F RAMEWORK

(17)

where u denotes the system parameters, including D, N , PF , and channel-related variables. The mapping is parameterized by a neural network and produces a high-quality feasible solution in a single forward pass. As illustrated in Fig. 2, the framework consists of a neural predictor that outputs unconstrained logits, which are transformed via grouped softmax operations into fractional variables defined over simplex domains. These structured transformations enforce feasibility by construction and map the fractional variables into the corresponding physical decision variables. Discrete variables are handled using an STE, enabling end-to-end differentiable training. The resulting variables are then evaluated through the system model to compute the achievable rate, which is used to define a performancedriven loss. Unlike conventional black-box learning approaches, CALO explicitly incorporates the structure of the underlying optimization problem into the model design. This ensures feasibility, enables direct optimization of the system performance, and avoids reliance on iterative solvers. As a result, CALO replaces computationally expensive iterative optimization with a single-shot inference mechanism, enabling efficient and realtime resource allocation in RIS-assisted wireless systems. B. Challenges in Learning-Based Optimization Despite their potential, the application of learning-based methods to structured resource allocation problems, such as P2, faces several fundamental challenges. Constraint awareness: Deep neural networks are unconstrained function approximators and do not inherently satisfy problem-specific feasibility conditions. In the considered problem, the decision variables are subject to strict linear equality constraints. Conventional approaches rely on penalty terms or projection steps, which may lead to infeasible or suboptimal solutions. Absence of optimal labels: The underlying problem is non-convex, and globally optimal solutions are generally intractable. Existing methods such as BCD provide only suboptimal solutions, limiting the applicability of supervised learning approaches that depend on high-quality ground truth labels. Discrete decision variables: The presence of discrete variables, such as reflecting-element allocation, introduces nondifferentiability and prevents direct application of gradientbased optimization. This challenge is further exacerbated by the coupling between discrete and continuous variables. These challenges motivate the design of the proposed CALO framework, which addresses them through constraint-aware parameterization, differentiable handling of discrete variables, and performance-driven learning.

A. Framework Overview To solve the reduced structured non-convex optimization problem (P2), we propose the constraint-aware learning optimization (CALO) framework, which replaces iterative op-

C. Constraint-Aware Parameterization A central challenge in P2 is ensuring that the predicted solutions satisfy the linear equality constraints (14)-(16) together

IEEE TRANSACTIONS ON COGNITIVE COMMUNICATIONS AND NETWORKING

with non-negativity and integer requirements. To address this, CALO adopts a structure-aware parameterization that transforms the constrained variables into fractional representations defined over simplex domains. 1) Fractional Representation: We introduce three groups of fractional variables: α = [α1 , α2 , α3 ], β = [β1 , β2 ], and γ = [γ1 , γ2 ], such that: 3 X αi = 1, αi ≥ 0, (18) i=1 2 X i=1 2 X

βi = 1,

βi ≥ 0,

(19)

γi = 1,

γi ≥ 0.

(20)

6

After rounding, a small mismatch may occur. To enforce exact feasibility, a simple correction is applied: N2 = N − N1 , ensuring that (14) holds exactly. 4) Discussion: The proposed approach enables differentiable learning over discrete variables without resorting to combinatorial search or relaxation techniques. Although STE provides a biased gradient approximation, it has been shown to be effective in practice for learning problems involving discrete operations. In this work, it allows joint optimization of continuous and discrete variables within a unified end-to-end framework while preserving feasibility. E. Regret-Based Learning Objective

i=1

The physical variables are obtained through: xi = αi D,

i = 1, 2, 3,

Pi = βi PF ,

i = 1, 2,

Ni = round(γi N ),

i = 1, 2.

(21) (22) (23)

2) Feasibility Guarantee: By construction, the simplex constraints in (18) and (19), together with the mappings in (21) and (22), ensure that the equality constraints in (16) and (15) are strictly satisfied. For the discrete variables, the rounding operation in (23) enforces integer feasibility. Any residual mismatch is resolved through a simple correction step, ensuring that (14) holds exactly. 3) Neural Parameterization via Softmax: To enforce the simplex constraints, the fractional variables are obtained via softmax transformations applied to the neural network outputs: ezβ,i ezγ,i ezα,i αi = P zα,k , βi = P zβ,k , γi = P zγ,k . (24) ke ke ke 4) Discussion: The proposed parameterization transforms P2 into an unconstrained learning problem over simplex domains, eliminating the need for projection or penaltybased methods while guaranteeing feasibility by construction. Moreover, preserving the grouped structure enables the model to effectively exploit the problem decomposition, leading to accurate and efficient learning. D. Discrete Variable Handling via STE A key challenge in (P2) is the presence of discrete decision variables, namely the reflecting-element allocation {N1 , N2 }, which must take integer values. This introduces non-differentiability and prevents direct application of gradient-based optimization. 1) Forward Mapping: Using the fractional parameterization in (23), the continuous representation of the discrete variables is obtained as Ñi = γi N, i = 1, 2, followed by a rounding operation Ni = round(Ñi ). 2) STE: To enable gradient-based training, the rounding i operation is approximated using STE, where ∂N ≈ 1, such ∂ Ñi ∂L ∂L that ∂ Ñ = ∂Ni . This allows gradients to propagate through i the discrete variables, enabling end-to-end learning. 3) Constraint Consistency: Since the fractional variables satisfy (20), the continuous variables obey Ñ1 + Ñ2 = N.

A key challenge in P2 is the absence of globally optimal ground truth solutions due to its non-convex nature. Existing methods provide only suboptimal solutions, denoted by Rref , which serve as performance baselines rather than optimal labels. Instead of relying on supervised learning, we adopt a performance-driven objective that directly optimizes the achievable rate. Let Rpred denote the rate obtained from the predicted solution. The regret is defined as Rref − Rpred . The learning objective is formulated as a hinge-based regret minimization: Lreg = max (0, Rref − Rpred ) .

(25)

This objective penalizes the model only when it underperforms the reference solution and imposes zero loss otherwise. As a result, the model is encouraged to match or exceed the performance of the baseline rather than imitate it. The gradient with respect to the network parameters θ is given by ( −∇θ Rpred , if Rpred < Rref , ∇θ Lreg = (26) 0, otherwise. This ensures that updates are applied only when the predicted solution is suboptimal, driving the model toward improved performance. Unlike conventional losses, such as MSE, which force imitation of suboptimal labels, the proposed objective enables the model to discover superior solutions. 1) Discussion: The proposed formulation shifts the learning objective from solution imitation to direct performance optimization. By leveraging suboptimal solutions as references, the model is encouraged to surpass classical optimization methods while remaining fully compatible with end-to-end gradient-based training. F. Learning Framework and Model Architecture The CALO framework is illustrated in Fig. 3, where all components are integrated into a unified end-to-end learning pipeline. The model maps system parameters directly to feasible decision variables while optimizing the achievable rate. 1) Input and Neural Prediction: Let s denote the input vector capturing system parameters, including D, PF , and channel-related variables. The input is normalized and fed into a neural network fθ (·), which produces grouped logits corresponding to placement, power allocation, and element allocation variables, i.e., [zα , zβ , zγ ] = fθ (s).

IEEE TRANSACTIONS ON COGNITIVE COMMUNICATIONS AND NETWORKING

Input Layer

Hidden Layer 1

Hidden Layer n

Output Logits

Neural Predictor

7

Fractional Variables

Physical Variables Variable Recovery

Constraint Satisfaction zx1 zx2

αx1 Softmax

zx3

αx2

D

x1

Objective Evaluation Rref

x2

×

x3

αx3 PF

···

zp1

αp1 ×

Softmax

.. .

.. .

zp2

P1

αp2

Analytical Rate Evaluation Rpred

Hinge Loss max(0, Rref − Rpred )

P2 N

γn1

zn1 zn2

N1 ×

Softmax γn2

STE N2

Only N1 , N2 are discrete and require STE.

.. .

Backpropagation: Update Neural Network Weights

Fig. 3. Architecture of the proposed neural framework. The multilayer perceptron produces grouped logits that are transformed by separate softmax heads into normalized fractional allocations. These fractions are mapped into physical decision variables using the system constraints. The discrete reflector-count branch is handled by an STE, after which the predicted achievable rate is evaluated and compared against a reference solution through a hinge-style loss.

2) Structured Variable Construction: The logits are transformed into fractional variables via softmax mappings as defined in (24). These fractional variables are then mapped to physical decision variables using the deterministic transformations in (21)–(23), ensuring feasibility by construction. Discrete variables are handled using the STE mechanism introduced earlier. 3) Objective Evaluation: The resulting decision variables are evaluated through the system model to compute the achievable rate, Rpred = F(X, P, N, s), where F(·) denotes the analytical rate expression. 4) End-to-End Training: The model is trained using the regret-based loss defined in (25). Gradients propagate through the system model, parameterization layers, and neural network, enabling end-to-end optimization. 5) Discussion: The proposed framework integrates constraint-aware parameterization, discrete variable handling, and performance-driven learning into a single differentiable architecture. This enables direct optimization of the system objective while guaranteeing feasibility and avoiding iterative solvers, resulting in an efficient and scalable solution for RIS-assisted wireless systems. V. S IMULATION R ESULTS A. Experimental Setup The experiments are conducted under a unified and controlled environment to ensure fair comparison across all considered models. A single dataset is generated using BCD, which serves as a strong suboptimal baseline. This ensures consistency in supervision and enables a reliable evaluation of the proposed learning-based framework against the same reference solutions. The dataset is constructed by simulating a wide range of wireless system configurations, capturing variations in key parameters including D, N , and PF . Such diversity is essential to promote generalization and to evaluate

the robustness of the learned models across heterogeneous deployment scenarios. All experiments are executed on a personal computing platform equipped with an 8-core ARM-based processor (architecture: aarch64) and 8 GB of RAM. Due to the absence of GPU acceleration, both training and inference are performed entirely on CPU resources, highlighting the practicality of the proposed approach under limited computational capabilities. The implementation is developed in Python (version 3.11.6), using TensorFlow (version 2.21.0) as the primary deep learning framework. Supporting libraries include NumPy (version 1.26.0) and Pandas (version 2.1.3) for numerical operations and data handling, while Scikit-learn (version 1.3.1) is utilized for preprocessing and dataset partitioning. Statistical validation and result analysis are carried out using the R programming language (version 4.5.0). All neural network models are trained using the Adam optimizer with a fixed learning rate of 10−3 over 20 epochs. The Rectified Linear Unit (ReLU) activation function is employed in all hidden layers to effectively capture non-linear relationships in the structured input space. These training settings are kept consistent across all architectures, while variations are introduced only in the number of hidden layers, the number of neurons per layer, and the batch size, allowing for a controlled investigation of architectural design choices. All experiments are conducted under fixed random seeds and identical training conditions. B. Experimental Design 1) Training and Testing Datasets: The dataset used in this study is generated using BCD. The input variables are sampled from predefined random distributions to ensure sufficient coverage of the system parameter space. For each generated sample, BCD is executed to obtain the corresponding optimal design variables and the achievable data rate, which serve as

IEEE TRANSACTIONS ON COGNITIVE COMMUNICATIONS AND NETWORKING

TABLE I PARAMETER DISTRIBUTIONS FOR TRAINING AND TESTING DATASETS UNDER DIFFERENT SCENARIOS

Var. Description D M H PF PB N β δ12 δ22 δ02

Type

Training Dataset

Urban Scenario

Rural Scenario

Total horizontal Continuous U (50, 200) U (20, 80) U (150, 300) distance (m) Number of Discrete Ud (2, 17) Ud (4, 9) Ud (2, 7) BS antennas Height of Continuous U (1, 6) U (2, 4) U (4, 7) RISs (m) Total RIS Continuous U (20, 28) U (20, 26) U (22, 28) power (dBm) BS transmit Continuous U (25, 35) U (28, 32) U (30, 34) power (dBm) Total reflecting Discrete Ud (50, 1500) Ud (50, 750) Ud (750, 2000) elements Channel gain Continuous N (−30, 2) N (−32, 2) N (−28, 2) (dB) Noise power at RIS1 Continuous N (−70, 3) N (−75, 2) N (−68, 2) (dBm) Noise power at RIS2 Continuous N (−70, 3) N (−75, 2) N (−68, 2) (dBm) Noise power at user Continuous N (−80, 3) N (−85, 2) N (−78, 2) (dBm)

reference solutions. The training dataset consists of 10,000 samples generated from broad parameter ranges to capture diverse system configurations. Empirically, this dataset size was found sufficient to ensure stable convergence of all considered models, with no significant performance gains observed when increasing the number of training samples further. To rigorously evaluate the generalizability of CALO, two distinct and unseen testing datasets are constructed, corresponding to urban and rural deployment scenarios. Each testing dataset contains 2,000 samples and is generated using scenario-specific parameter distributions that reflect realistic propagation environments. The adopted parameter distributions are designed to reflect both general training conditions and scenario-specific characteristics, capturing variations in propagation distance, power budgets, and channel conditions commonly encountered in RIS-assisted wireless systems, as detailed in Table I. The corresponding output variables (x1 , x2 , x3 , P1 , P2 , N1 , N2 ) are obtained by solving each instance using BCD, ensuring that all reference solutions satisfy the problem constraints and represent high-quality feasible operating points. The urban scenario is characterized by shorter transmission distances and more severe propagation conditions, reflected by lower values of D and higher attenuation levels. In contrast, the rural scenario represents more open environments with longer transmission distances and relatively lower attenuation. These distinct distributions allow for evaluating the robustness of the proposed models under heterogeneous deployment conditions. Both testing datasets are generated independently from the training dataset, ensuring that evaluation is performed on unseen data distributions. This design enables a rigorous assessment of the model’s ability to generalize beyond the conditions observed during training and to outperform the

8

TABLE II T RAINING P ERFORMANCE ACROSS D IFFERENT A RCHITECTURES (I NDUCTION P HASE ) Layers 4

6

8

16

Units Batch 64 16 128 512 64 64 128 512 64 256 128 512 64 512 128 512

Trainable Params 1111

21959

465159

3949063

Train Loss 0.230 0.448 1.742 0.075 0.147 0.563 0.021 0.039 0.144 0.026 0.048 0.183

Val Loss 0.005 0.267 1.546 0.003 0.003 0.355 0.003 0.003 0.006 0.003 0.003 0.052

BCD baseline under diverse system configurations. All models are evaluated on identical testing datasets. For each scenario, performance metrics are computed over all samples and include the achievable data rate, the average achievable data rate gain with respect to the BCD baseline, and constraint feasibility. This evaluation protocol enables a systematic assessment of (i) the extent to which the proposed models can surpass the BCD solution, (ii) their ability to preserve feasibility under varying system configurations, and (iii) their generalization capability across heterogeneous deployment environments. 2) Evaluation of the Proposed Solution under Different Architectures: This subsection evaluates the effectiveness of CALO in addressing the underlying non-convex optimization problem. Specifically, the evaluation focuses on: (i) the ability to satisfy the problem constraints through learningbased parameterization, (ii) the performance improvement over the BCD solution, and (iii) the computational efficiency of the learned models in enabling fast inference. To identify a suitable model architecture, multiple neural network configurations are evaluated by varying the number of layers, neurons per layer, and batch sizes, as summarized in Table II. The considered architectures range from shallow networks with limited capacity to deeper models with a significantly higher number of trainable parameters. For each configuration, both training and validation losses are monitored to assess convergence behavior and generalization capability. Although several architectures achieve comparably low validation loss values, it is observed that validation loss alone is not a sufficient criterion for selecting the best-performing model in this problem. This is primarily due to the nature of the learning objective, where the loss function does not directly capture the true optimization goal, namely the achievable data rate under strict feasibility constraints. As a result, models with similar validation loss may exhibit significantly different performance in terms of constraint satisfaction and rate optimization during testing. Therefore, instead of relying solely on validation loss, the final model selection is based on a comprehensive evaluation that includes feasibility, achievable data rate performance, and computational efficiency. This ensures that the selected architecture not only generalizes well but also effectively solves the underlying non-convex opti-

IEEE TRANSACTIONS ON COGNITIVE COMMUNICATIONS AND NETWORKING

TABLE III P ERFORMANCE E VALUATION UNDER D IFFERENT A RCHITECTURES (U RBAN S CENARIO ) Layers 4

6

8

16

Units Batch Avg. Inference Time (s) Feasibility Avg. Rate Gain 64 0.00009 100% 0.425 16 128 0.00009 100% 0.404 512 0.00008 100% 0.369 64 0.00010 100% 0.407 64 128 0.00010 100% 0.426 512 0.00010 100% 0.383 64 0.00017 100% 0.392 256 128 0.00014 100% 0.426 512 0.00014 100% 0.431 64 0.00030 100% 0.411 512 128 0.00030 100% 0.421 512 0.00030 100% 0.423

TABLE IV P ERFORMANCE E VALUATION UNDER D IFFERENT A RCHITECTURES (RURAL S CENARIO ) Layers 4

6

8

16

Units Batch Avg. Inference Time (s) Feasibility Avg. Rate Gain 64 0.00008 100% 0.312 16 128 0.00008 100% 0.288 512 0.00008 100% 0.280 64 0.00009 100% 0.284 64 128 0.00009 100% 0.301 512 0.00009 100% 0.292 64 0.00014 100% 0.282 256 128 0.00014 100% 0.309 512 0.00014 100% 0.320 64 0.00029 100% 0.305 512 128 0.00029 100% 0.311 512 0.00029 100% 0.295

mization problem. The results for all considered architectures are summarized in Tables III and IV for urban and rural scenarios, respectively. All evaluated models strictly satisfy the problem constraints with zero numerical violation across samples. In P all test P particular, the equality constraints x = D, i i i Ni = N , P and P = P are inherently enforced through the proi F i posed softmax-based parameterization. As shown in Tables III and IV, all evaluated architectures consistently achieve a feasibility rate of 100% across both scenarios. This property is critical for practical deployment, as infeasible solutions would render the system design invalid regardless of their performance in terms of data rate. In terms of achievable data rate performance, the results indicate that the architecture with 8 layers and 256 neurons, trained with a batch size of 512, achieves the highest average rate gain in both urban and rural scenarios. Specifically, this configuration yields the best performance despite not exhibiting the lowest validation loss during the training phase. This observation further supports the earlier argument that validation loss alone is not a reliable indicator of the true optimization performance in this problem. Moreover, the results demonstrate that increasing model complexity beyond a certain point does not necessarily lead to performance improvement. For instance, deeper architectures with significantly higher numbers of parameters do not consistently outperform the selected configuration. This highlights

9

that the proposed learning framework is capable of effectively solving the underlying non-convex optimization problem without requiring excessively deep or computationally expensive models. From a computational perspective, the proposed learningbased approach offers a significant advantage over the BCD algorithm. The average inference time per sample for the neural network models is on the order of 10−4 seconds, as shown in Tables III and IV. In contrast, the average time required by the BCD algorithm to compute a suboptimal solution is approximately 0.0768 seconds for the urban scenario and 0.3373 seconds for the rural scenario. 3) Statistical Testing and Comparative Performance: This subsection compares CALO with the BCD baseline across urban and rural environments using paired achievable-rate differences. Since both methods are evaluated on the same testing samples, the comparison is naturally paired. Therefore, the statistical analysis is performed on the per-sample rate (i) (i) (i) (i) difference d(i) = RCALO − RBCD , where RCALO and RBCD denote the achievable rates obtained by CALO and BCD for the i-th test sample, respectively. This paired formulation directly measures the improvement achieved by CALO over BCD while reducing the effect of sample-to-sample variability. Figs. 4a and 4b show the achievable-rate distributions obtained by CALO and BCD in the urban and rural scenarios, respectively. The boxplots indicate a consistent upward shift of CALO relative to BCD. The statistical testing procedure is designed to answer three questions: whether a parametric paired test is appropriate, whether the observed improvement is statistically significant, and whether the magnitude of improvement is practically meaningful. First, we apply the Anderson–Darling test to assess the normality of the paired differences, which is a standard goodness-of-fit approach for evaluating distributional assumptions [28]. This step is required because a paired ttest assumes approximate normality of the paired differences. The Anderson–Darling test rejects normality in both scenarios, with p < 2.2×10−16 . Therefore, we avoid a parametric paired t-test and use the Wilcoxon signed–rank test, which is a rankbased nonparametric test suitable for paired comparisons when the normality assumption is not satisfied [29], [30]. The Wilcoxon signed–rank test is applied to examine whether CALO provides a positive location shift over BCD. The hypotheses are: H0 : median(d) = 0

vs.

H1 : median(d) > 0,

where d = RCALO −RBCD . For both urban and rural environments, the Wilcoxon signed–rank test yields p < 2.2 × 10−16 , leading to rejection of H0 . This result provides strong statistical evidence that CALO achieves a positive paired achievablerate improvement over BCD. To complement the p-values, we report the Hodges– Lehmann (HL) median shift with 95% confidence intervals, which provides a robust estimate of the typical paired improvement [30]. We √ also report the standardized Wilcoxon effect size r = Z/ n, where Z is the normal approximation of the signed-rank statistic and n is the number of paired samples. Effect-size measures are important because they quantify the

10

2.0

35

Density

Achievable rate (bps/Hz)

IEEE TRANSACTIONS ON COGNITIVE COMMUNICATIONS AND NETWORKING

30

1.5 1.0 0.5

25

0.0 BCD

0.0

CALO 8-256

0.4

0.6

0.8

1.0

DR = RCALO - RBCD (bps/Hz)

Model (a) Urban scenario.

(a) Urban scenario.

32.5

3

30.0

Density

Achievable rate (bps/Hz)

0.2

27.5

2 1

25.0 0 BCD

0.0

CALO 8-256

0.2

0.4

0.6

0.8

1.0

DR = RCALO - RBCD (bps/Hz)

Model (b) Rural scenario.

(b) Rural scenario.

Fig. 4. Distribution of the achievable data rate obtained by the BCD baseline and the selected CALO model across the testing samples under (a) urban and (b) rural deployment scenarios.

Fig. 5. Distribution of the rate difference (RCALO − RBCD ) under urban and rural environments.

(MRI) as: N

strength of the improvement rather than only its statistical significance [31]. In the urban scenario, the HL shift is 0.4088 bps/Hz with a 95% confidence interval of [0.3996, 0.4181] bps/Hz. In the rural scenario, the HL shift is 0.3025 bps/Hz with a 95% confidence interval of [0.2972, 0.3078] bps/Hz. In both cases, the standardized Wilcoxon effect size is very large, with r ≈ 0.8661. The rank-biserial correlation is approximately one, and the common-language effect size is about 0.9995, indicating that nearly all paired comparisons favor CALO. Figs. 5a and 5b show the distributions of the paired rate differences (RCALO − RBCD ). The distributions are concentrated on positive values, which is consistent with the Wilcoxon test results and confirms that the improvement is not driven by a small number of isolated samples. Although the relative gains are modest, they are practically meaningful because throughput scales with bandwidth. If R denotes the achievable rate per unit bandwidth, the end-to-end throughput is T = R × B, where B is the system bandwidth. Therefore, an incremental gain ∆R leads to an absolute throughput increase of ∆T = ∆R × B, which can become significant over wide bandwidths and when aggregated across users, time, and slices. To provide a normalized interpretation of the gain, we compute the Mean Relative Improvement

1 X MRI = N i=1

(i)

(i)

RCALO − RBCD

!

× 100 (i) RBCD The resulting MRI values are approximately 1.29% in the urban scenario and 0.98% in the rural scenario. Table V summarizes the statistical and performance results. Overall, the statistical analysis confirms that the observed gains are not due to random fluctuations or isolated test samples. Instead, CALO provides a consistent positive shift over the BCD baseline in both urban and rural environments while maintaining 100% feasibility. Combined with the elimination of per-instance iterative optimization at inference, these results support the main claim that CALO achieves feasible, statistically significant, and practically meaningful performance improvements with real-time applicability. C. Performance Analysis Under Varying System Parameters To further assess the robustness and generalization capability of CALO, this subsection examines its behavior under varying system parameters. In particular, we evaluate the impact of key design variables, including D, PF , and N , on the achievable data rate. For each parameter, controlled experiments are conducted by varying the parameter of interest while keeping the remaining variables fixed on the values listed in Table VI, thereby isolating its effect on the system

IEEE TRANSACTIONS ON COGNITIVE COMMUNICATIONS AND NETWORKING

Metric Rural Urban Feasibility (%) 100 100 Avg. rate gain (bps/Hz) 0.320 0.431 Anderson–Darling p-value < 2.2 × 10−16 < 2.2 × 10−16 Wilcoxon signed–rank p-value < 2.2 × 10−16 < 2.2 × 10−16 HL shift (bps/Hz) 0.3025 0.4088 Wilcoxon effect size r 0.8661 0.8661 Mean Relative Improvement (MRI) 0.98% 1.29%

32

Achievable rate (bps/Hz)

TABLE V CALO VS . BCD ACROSS RURAL AND URBAN ENVIRONMENTS .

11

BCD CALO 4-16 CALO 6-64 CALO 8-256 CALO 16-512

31 30 29 28 27 26 25 50

100

150

200

250

300

Distance (D) (meters)

TABLE VI D EFAULT SYSTEM PARAMETERS .

behavior. This analysis provides insights into how the proposed learning-based solution adapts to different operating conditions of the underlying non-convex optimization problem. 1) Impact of D: To investigate the effect of propagation distance on system performance, D is varied while keeping the remaining system parameters fixed on the values in Table VI. Fig. 6 illustrates the achievable data rate and the corresponding rate gain over the BCD baseline as a function of D. As shown in Fig. 6a, the achievable data rate decreases monotonically with increasing distance for all considered methods. This behavior is expected due to the increased path loss associated with larger transmitter–receiver separation, which reduces the effective received signal power and consequently the SNR. Despite this degradation, CALO consistently achieves higher data rates compared to the BCD baseline across the entire range of distances. More importantly, Fig. 6b shows that the achievable-rate gain remains positive and relatively stable, indicating that the learned solution maintains its superiority even under unfavorable propagation conditions. Furthermore, among the evaluated architectures, the configuration with 8 layers and 256 neurons demonstrates the most consistent and highest performance gain across all distances. This observation reinforces that an appropriately designed architecture can effectively capture the underlying structure of the non-convex optimization problem without requiring excessive model complexity. Overall, these results confirm that CALO, not only improves performance under nominal conditions but also generalizes well across varying propagation regimes, making it suitable for practical wireless deployment scenarios. 2) Impact of N : To evaluate the influence of the number of reflecting elements on system performance, N is varied while keeping the remaining system parameters fixed on the values in Table VI. Fig. 7 presents the achievable data rate and the corresponding rate gain over the BCD baseline as a function of the number of reflecting elements. As shown in Fig. 7a, the achievable data rate increases with N for all considered methods. This behavior is attributed to the enhanced passive beamforming gain provided by the RIS, where a larger number

DR = RCALO - RBCD (bps/Hz)

(a) Achievable rate vs. distance. Parameter Value Parameter Value Parameter Value Parameter Value M 4 H 2m N 750 D 100 m PF 26 dBm PB 30 dBm β −30 dB δ02 −80 dBm 2 2 δ1 −70 dBm δ2 −70 dBm

CALO 4-16 CALO 6-64 CALO 8-256 CALO 16-512

0.4 0.3 0.2 0.1 0.0 50

100

150

200

250

300

Distance (D) (meters) (b) Achievable-rate gain vs. distance. Fig. 6. Performance versus distance.

of reflecting elements enables more precise signal reflection and constructive combining at the receiver, thereby improving the effective channel gain. However, the rate improvement exhibits a diminishing return as N becomes large. This saturation effect is expected, as the incremental contribution of additional reflecting elements decreases once the dominant propagation paths are sufficiently reinforced and the system approaches its performance limits. Importantly, CALO consistently outperforms the BCD baseline across the entire range of N . As illustrated in Fig. 7b, the achievable-rate gain remains positive and stable, indicating that the learning-based solution effectively captures the relationship between the number of reflecting elements and the optimal resource allocation. Furthermore, the architecture with 8 layers and 256 neurons again demonstrates the most favorable trade-off between performance and model complexity, achieving the highest and most consistent gains across all values of N . This result further confirms that the proposed framework does not require excessively deep architectures to approximate high-quality solutions for the underlying non-convex optimization problem. This behavior indicates that the proposed model effectively learns the nonlinear relationship between RIS size and system performance, which is difficult to capture using traditional optimization methods. Overall, these findings highlight the ability of the proposed approach to generalize across different RIS configurations and to maintain robust performance gains

32 29 26

BCD CALO 4-16 CALO 6-64 CALO 8-256 CALO 16-512

23

12

Achievable rate (bps/Hz)

Achievable rate (bps/Hz)

IEEE TRANSACTIONS ON COGNITIVE COMMUNICATIONS AND NETWORKING

20

29.4 29.2 29.0 28.8

BCD CALO 4-16 CALO 6-64 CALO 8-256 CALO 16-512

28.6 28.4

0

250

500

750

1000 1250 1500 1750 2000

0.0

Reflecting elements (N)

0.3 0.2 CALO 4-16 CALO 6-64 CALO 8-256 CALO 16-512

0.1 0.0 0

250

500

750 1000 1250 1500 1750 2000

Reflecting elements (N) (b) Achievable-rate gain vs. number of reflecting elements.

0.4

0.6

0.8

1.0

(a) Achievable rate vs. total RIS power.

DR = RCALO - RBCD (bps/Hz)

DR = RCALO - RBCD (bps/Hz)

(a) Achievable rate vs. number of reflecting elements.

0.4

0.2

Total RISs power (Pf) (watts)

0.4 0.3 0.2 CALO 4-16 CALO 6-64 CALO 8-256 CALO 16-512

0.1 0.0 0.0

0.2

0.4

0.6

0.8

1.0

Total RISs power (Pf) (watts) (b) Achievable-rate gain vs. total RIS power.

Fig. 7. Performance versus the number of reflecting elements.

Fig. 8. Performance versus total RIS power budget.

even as the system scales in size. 3) Impact of PF : For the sensitivity analysis of PF , the total RIS power budget is converted from dBm to watts before evaluating the analytical rate expression. To investigate the effect of the total RIS power budget, PF is varied while keeping the remaining system parameters fixed on the values in Table VI. Fig. 8 illustrates the achievable data rate and the corresponding gain over the BCD baseline as a function of the total RIS power. As shown in Fig. 8a, the achievable rate increases monotonically with PF for all considered methods. This behavior is fundamentally different from the passive RIS case, as the active RIS is capable of amplifying the reflected signals, thereby directly enhancing the end-to-end SNR. Specifically, increasing PF enables stronger amplification at the RIS elements, which improves the effective cascaded channel gain and compensates for the severe path loss typically associated with multi-hop RIS-assisted communication. As a result, the system transitions from a reflection-limited regime at low PF to an amplification-assisted regime at higher power budgets. However, similar to the behavior observed with the number of reflecting elements, the rate improvement exhibits a diminishing return as PF increases. This saturation effect arises because the system becomes progressively constrained by other limiting factors, such as receiver noise and residual interference, reducing the marginal benefit of additional RIS power.

Importantly, CALO consistently outperforms the BCD baseline across the entire range of PF , as shown in Fig. 8b. The gain remains positive and relatively stable, indicating that the learning-based model effectively captures the nonlinear coupling between RIS power allocation and system performance. This observation further confirms that optimizing RIS power is a critical factor in active RIS systems, as it directly controls the balance between signal amplification and noise enhancement. Among all configurations, the architecture with 8 layers and 256 neurons again achieves the best performance, demonstrating its ability to accurately model the complex interaction between amplification, resource allocation, and channel conditions without requiring excessive model depth. These results highlight a key advantage of the proposed approach: its capability to efficiently exploit the additional degrees of freedom introduced by active RISs. In contrast to conventional optimization methods, which may struggle with the increased complexity of power-dependent variables, the proposed learning-based framework maintains robust performance and scalability across different power regimes. VI. C ONCLUSION In this paper, we proposed a constraint-aware learning optimization (CALO) framework for linearly constrained nonconvex resource allocation, with a focus on double-active RIS-assisted wireless systems. CALO reformulates the original problem using fractional parameterization over simplex

IEEE TRANSACTIONS ON COGNITIVE COMMUNICATIONS AND NETWORKING

domains, enabling the distance, power, and element-budget constraints to be satisfied by construction. Discrete reflectingelement allocation is handled through a straight-through estimator, while a regret-driven objective allows the model to improve beyond BCD-based reference solutions without relying on globally optimal labels. Simulation results show that CALO achieves 100% feasibility across all tested configurations without projection or penalty-based corrections. The proposed framework also provides statistically significant achievable-rate gains over the BCD baseline in both urban and rural scenarios, as validated using nonparametric statistical testing. Moreover, by replacing iterative optimization with a single forward pass, CALO reduces online inference time by orders of magnitude, supporting real-time applicability. Sensitivity analysis under varying propagation distance, number of reflecting elements, and RIS power budget further confirms the robustness and generalization capability of the proposed framework. Overall, the results demonstrate that embedding optimization structure into the learning process can produce feasible, scalable, and high-performance solvers for complex wireless resource-allocation problems. Future work will extend CALO to more general channel models, multi-user and multi-cell scenarios, and dynamic environments with time-varying system parameters. Incorporating uncertainty-aware and robust learning mechanisms is also an important direction for improving practical deployment.

R EFERENCES [1] G. Araniti, C. Garibotto, A. Jose, F. Marcello, V. Pilloni, A. Sciarrone, C. Suraci, P. Zema, and M. Zerbino, “Medical digital twin and IoT devices: Monitoring patients through human-centered technologies,” IEEE Internet Things Mag., vol. 5, no. 1, pp. 1–7, 2025. [2] B. Hu, H. Liu, J. Du, M. López-Benı́tez, C. Wu, X. Chu, and D. Niyato, “MADQN-enhanced computation offloading and resource allocation for 6G low-altitude economy vehicular networks,” IEEE Trans. Cogn. Commun. Netw., vol. 11, no. 3, pp. 1–1, 2025. [3] C.-X. Wang, X. You, X. Gao, X. Zhu, Z. Li, C. Zhang, H. Wang, Y. Huang, Y. Chen, H. Haas, J. S. Thompson, E. G. Larsson, M. D. Renzo, W. Tong, P. Zhu, X. Shen, H. V. Poor, and L. Hanzo, “On the road to 6G: Visions, requirements, key technologies, and testbeds,” IEEE Commun. Surv. Tutor., vol. 25, no. 2, pp. 905–974, 2023. [4] M. H. Alsamh, A. Hawbani, S. Kumar, and S. H. Alsamhi, “Multisensory metaverse-6G: A new paradigm of commerce and education,” IEEE Access, vol. 12, pp. 75 657–75 677, 2024. [5] A. Pradhan, S. Das, M. J. Piran, and Z. Han, “A survey on physical layer security of ultra/hyper reliable low latency communication in 5G and 6G networks: Recent advancements, challenges, and future directions,” IEEE Access, vol. 12, pp. 112 320–112 353, 2024. [6] W. Saad, M. Bennis, and M. Chen, “A vision of 6G wireless systems: Applications, trends, technologies, and open research problems,” IEEE Network, vol. 34, no. 3, pp. 134–142, 2020. [7] Q. Wu and R. Zhang, “Towards smart and reconfigurable environment: Intelligent reflecting surface aided wireless network,” IEEE Commun. Mag., vol. 58, no. 1, pp. 106–112, 2020. [8] S. E. Chakkaravarthy, D. Rayaroth, V. B. Kumaravelu, T. S. Jayaraman, V. P. G. Sivabalan, V. C. Thirumavalavan, R. Venkatesan, A. Murugadass, and A. L. Imoize, “Reconfigurable intelligent surfaces for 6G: A comprehensive overview and electromagnetic analysis,” in Reconfigurable Intelligent Surfaces for 6G and Beyond Wireless Networks, 2025, pp. 71–112. [9] Y. Chen, Q. Wu, G. Chen, and W. Chen, “Spatial multiplexing oriented channel reconfiguration in multi-IRS aided MIMO systems,” IEEE Trans. Veh. Technol., vol. 74, no. 6, pp. 9840–9845, 2025.

13

[10] Q. Sun, Y. Wu, X. Chen, and J. Zhang, “SLNR-based joint RISUE association and beamforming design for multi-RIS aided wireless communications,” IEEE Trans. Veh. Technol., vol. 73, no. 6, pp. 8660– 8670, 2024. [11] Y. Han, S. Zhang, L. Duan, and R. Zhang, “Double-IRS aided MIMO communication under LoS channels: Capacity maximization and scaling,” IEEE Trans. Commun., vol. 70, no. 4, pp. 2820–2837, 2022. [12] Y.-F. Liu, T.-H. Chang, M. Hong, Z. Wu, A. M.-C. So, E. A. Jorswieck, and W. Yu, “A survey of recent advances in optimization methods for wireless communications,” IEEE J. Sel. Areas Commun., vol. 42, no. 11, pp. 2992–3031, 2024. [13] M. Razaviyayn, T. Huang, S. Lu, M. Nouiehed, M. Sanjabi, and M. Hong, “Nonconvex min-max optimization: Applications, challenges, and recent theoretical advances,” IEEE Signal Process. Mag., vol. 37, no. 5, pp. 55–66, 2020. [14] M. Ahmed, S. Raza, A. Amin Soofi, F. Khan, W. Ullah Khan, S. Zain Ul Abideen, F. Xu, and Z. Han, “Active reconfigurable intelligent surfaces: Expanding the frontiers of wireless communication–A survey,” IEEE Commun. Surv. Tutor., vol. 27, no. 2, pp. 839–869, Apr. 2025. [15] P. Gavriilidis, D. Mishra, B. Smida, E. Basar, C. Yuen, and G. C. Alexandropoulos, “Active reconfigurable intelligent surfaces: Circuit modeling and reflection amplification optimization,” IEEE Open J. Commun. Soc., vol. 6, pp. 5692–5710, 2025. [16] K. Zhi, C. Pan, H. Ren, K. K. Chai, and M. Elkashlan, “Active RIS versus passive RIS: Which is superior with the same power budget?” IEEE Commun. Lett., vol. 26, no. 5, pp. 1150–1154, 2022. [17] X. Pan, Z. Zheng, X. Huang, and Z. Fei, “WMMSE-based joint transceiver design for multi-RIS-assisted cell-free networks using hybrid CSI,” IEEE Trans. Wireless Commun., vol. 24, no. 9, pp. 7654–7669, 2025. [18] Q. Xue, J. Mu, F. Wei, M. Hua, and Q. Chen, “Joint user association and cooperative beamforming for multi-IRSs aided mmWave communication systems,” Digit. Commun. Netw., vol. 12, no. 2, pp. 252–261, 2026. [19] Z. Kang, C. You, and R. Zhang, “Double-active-IRS aided wireless communication: Deployment optimization and capacity scaling,” IEEE Wireless Communications Letters, vol. 12, no. 11, pp. 1821–1825, 2023. [20] X. Li, C. You, Z. Kang, Y. Zhang, and B. Zheng, “Double-active-IRS aided wireless communication with total amplification power constraint,” IEEE Commun. Lett., vol. 27, no. 10, pp. 2817–2821, 2023. [21] Y.-F. Liu, T.-H. Chang, M. Hong, Z. Wu, A. M.-C. So, E. A. Jorswieck, and W. Yu, “A survey of recent advances in optimization methods for wireless communications,” IEEE J. Sel. Areas Commun., vol. 42, no. 11, pp. 2992–3031, Nov. 2024. [22] M. Ahmed, F. Xu, Y. Lyu, A. A. Soofi, Y. Li, F. Khan, W. U. Khan, M. Sheraz, T. C. Chuah, and M. Deng, “RIS-driven resource allocation strategies for diverse network environments: A comprehensive review,” Trans. Emerg. Telecommun. Technol., vol. 36, no. 6, 2025. [23] S. Pala, K. Singh, O. Taghizadeh, C. Pan, O. A. Dobre, and T. Q. Duong, “Robust and secure multi-user STAR-RIS-aided communications: Optimization versus machine learning,” IEEE Trans. Commun., vol. 73, no. 9, pp. 7517–7534, 2025. [24] D. G. S. Pivoto, F. A. P. de Figueiredo, C. Cavdar, G. R. d. L. Tejerina, and L. L. Mendes, “A comprehensive survey of machine learning applied to resource allocation in wireless communications,” IEEE Commun. Surv. Tutor., vol. 28, pp. 1986–2053, 2026. [25] Y. Dai, L. Lyu, N. Cheng, M. Sheng, J. Liu, X. Wang, S. Cui, L. Cai, and X. Shen, “A survey of graph-based resource management in wireless networks–Part II: Learning approaches,” IEEE Trans. Cogn. Commun. Netw., vol. 11, no. 4, pp. 2101–2122, 2025. [26] Z. Liu, Y. Li, Y.-C. Wu, and Y. Gong, “Learning to optimize resource allocation in dynamic wireless environments: Embracing the new while engaging the old,” IEEE Trans. Wireless Commun., vol. 24, no. 9, pp. 7346–7359, 2025. [27] X. Li, C. You, Z. Kang, Y. Zhang, and B. Zheng, “Double-active-IRS aided wireless communication with total amplification power constraint,” IEEE Commun. Lett., vol. 27, no. 10, pp. 2817–2821, 2023. [28] R. B. D’Agostino and M. A. Stephens, Eds., Goodness-of-Fit Techniques. New York, NY, USA: Marcel Dekker, 1986. [29] W. J. Conover, Practical Nonparametric Statistics, 3rd ed. New York, NY, USA: John Wiley & Sons, 1999. [30] M. Hollander, D. A. Wolfe, and E. Chicken, Nonparametric Statistical Methods, 3rd ed. Hoboken, NJ, USA: John Wiley & Sons, 2014. [31] R. J. Grissom and J. J. Kim, Effect Sizes for Research: Univariate and Multivariate Applications, 2nd ed. New York, NY, USA: Routledge, 2012.

Record · ID 324858 · SHA-256 04638a1c23214ba4
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.