1
Multi-Block Attention for Efficient Channel Estimation in IRS-Assisted mmWave MIMO
arXiv:2605.15032v1 [eess.SP] 14 May 2026
Mehrdad Momen-Tayefeh1, Mehrshad Momen-Tayefeh2, Maryam Sabbaghian1∗
Abstract—Intelligent Reflecting Surfaces (IRSs) are a promising technology for enhancing the spectral and energy efficiency of millimeter-wave (mmWave) multiple-input multiple-output (MIMO) systems. In these systems, accurate channel estimation remains challenging due to the passive nature of IRS elements and the high pilot overhead in large-scale deployments. This paper presents a deep learning-based Multi-Block Attention (MBA) framework for efficient cascaded channel estimation in IRS-assisted mmWave MIMO systems that utilize orthogonal frequency division multiplexing (OFDM). First, we show the optimality of the discrete Fourier transform (DFT) and Hadamard matrices as phase configurations for least squares (LS) estimation. To reduce training overhead, we selectively deactivate IRS elements and compensate for induced feature loss using a two-stage architecture: (i) a Convolutional Attention Network (CAN) for spatial correlation recovery and (ii) a Complex Multi-Convolutional Network (CMN) for noise suppression. The MBA architecture mitigates error propagation through attentionguided feature refinement and denoising. Simulation results indicate that the MBA method reduces pilot overhead by up to 87% compared to the LS estimator. Additionally, at signal-to-noise ratios of 10 dB, our proposed method achieves approximately 51% lower normalized mean squared error (NMSE) than leading methods. It also maintains low computational complexity and adapts effectively to various propagation environments. Index Terms—Intelligent reflecting surfaces, channel estimation, deep learning, MIMO, millimeter-wave.
I. I NTRODUCTION NTELLIGENT reflecting surfaces (IRSs) have emerged as a promising solution for enhancing coverage, link quality, and throughput in future wireless communication systems [1]. An IRS consists of a planar array of passive elements capable of dynamically controlling the phase of incident signals to reconfigure the wireless propagation environment and reflect signals toward desired directions [2]. Unlike active relays, IRSs operate passively, offering energy-efficient and cost-efficient enhancements without requiring additional radiofrequency (RF) chains [3], [4]. In IRS-assisted systems, accurate knowledge of the cascaded channel state information (CSI), representing the combined base station (BS)-IRS and IRS-mobile user (MU) links,
I
1 Mehrdad Momen-Tayefeh and Maryam Sabbaghian are with the School of Electrical and Computer Engineering, University of Tehran, Tehran, Iran (e-mail: [email protected]; [email protected]). 2 Mehrshad Momen-Tayefeh is with the Department of Computer Engineering, Sharif University of Technology, Tehran, Iran (e-mail: [email protected]). *Corresponding Author: Maryam Sabbaghian (e-mail: [email protected]). 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media. This is the author’s version of the work. It is posted here for your personal use. Not for redistribution. The definitive Version of Record was published in IEEE Transactions on Communications, DOI: 10.1109/TCOMM.2025.3618696.
is essential for configuring the IRS phase shifts to maximize signal alignment and system throughput [5]. However, acquiring accurate CSI is particularly challenging due to the passive nature of IRSs and the lack of baseband processing capabilities. This issue is further exacerbated in IRS-assisted orthogonal frequency division multiplexing (OFDM) systems, where the cascade channels over multiple subcarriers must be estimated concurrently [6]. Therefore, developing efficient channel estimation techniques that can accurately recover the cascade channel across subcarriers with minimal pilot overhead is crucial for implementing IRS-assisted millimeter-wave (mmWave) multiple-input multiple-output (MIMO) OFDM systems. A. Prior Works Channel estimation in IRS-assisted wireless systems has received considerable attention in recent years. However, existing approaches face persistent challenges in balancing estimation accuracy, training overhead, and computational efficiency. Early IRS-assisted channel estimation methods relied on sequential estimation by activating IRS elements one at a time, as in [7]. While simple, this approach introduces prohibitive training overhead due to the large number of IRS elements. To mitigate this, Jensen et al. proposed a sub-surface grouping strategy coupled with minimum variance unbiased estimation [8], later extended to frequency-selective channels in [9]. For wideband single-input single-output (SISO) systems, Zheng et al. employed the least squares (LS) method for estimating frequency-selective channels in IRS-aided OFDM scenarios [10]. Alternatively, minimum mean square error (MMSE) estimators under Gaussian assumptions offer closed-form solutions but often lack robustness in practical environments [11]. Channel estimation methods utilize the inherent sparsity in both the angular and delay domains of mmWave channels to minimize pilot overhead. In their work, J. He et al. introduced atomic norm minimization techniques for the super-resolution estimation of angle-of-departure (AoD) and angle-of-arrival (AoA) in [12]. This approach was later extended to OFDM systems in [13]. Additionally, structured matrix recovery and compressed sensing (CS) methods have been proposed in [14], [15]. However, these methods have high computational complexity, particularly when applied to large IRS deployments. To overcome these challenges, low-rank matrix completion methods for multi-user cascaded channel estimation were proposed in [16]. Although effective in principle, this approach suffers from high sensitivity to inaccuracies in rank estimation,
2
which can significantly degrade performance. In the wideband context, the authors in [17] developed orthogonal matching pursuit (OMP) and SBL-based algorithms, which exploit joint time–angle sparsity for SIMO-OFDM systems. However, their computational complexity scales cubically with the number of IRS elements and quadratically with the number of BS antennas, limiting scalability. More recently, a low-rank sparse tensor recovery method was proposed in [18] for IRS-assisted mmWave OFDM systems. This approach models the cascaded channel as a lowrank sparse tensor in the angle-delay domain and reconstructs it via manifold optimization. Although this reduces pilot overhead and improves accuracy, it requires accurate prior knowledge and remains computationally intensive. Moreover, most CS-based techniques involve solving iterative optimization problems, often requiring high-resolution dictionaries that significantly increase complexity and estimation time. Deep learning (DL) has emerged as a powerful alternative for IRS-aided channel estimation due to its data-driven modeling capability and potential for real-time inference [19], [20], [21]. Xu et al. in [22] proposed a DL model combining an ODE-enhanced recurrent neural network for temporal interpolation and an ODE-inspired feedforward network for spatial extrapolation to estimate time-varying IRS-assisted channels from partial pilot observations. To enhance estimation while reducing hardware dependency, Liu et al. [23] incorporated additional RF chains at the IRS and utilized denoising convolutional neural networks (DnCNNs) for estimation. Gong et al. [24] expanded this idea by equipping IRS elements with ADCs and deploying a deep neural network (DNN) to infer channel states. While these methods improve accuracy, they introduce high hardware costs and complexity due to the use of active IRS components. Similarly, Zhang et al. [25] proposed a two-stage DNN approach for channel reconstruction and beam selection, requiring active IRS elements. Mao et al. [26] presented RS-OMP, which integrates residual learning with OMP. This was further extended by Abdallah et al. in [27] through DD-FF and DD-FS models, which exploit angular sparsity for both frequency-flat and frequencyselective channels. While effective in reducing pilot overhead, these methods degrade significantly when the number of pilots is much smaller than the number of IRS elements. In [28], a super-resolution DNN approach was proposed using cascaded SRCNN and DnCNN networks for IRS-assisted MIMO-OFDM systems. Though it achieves high estimation accuracy, it still requires many time slots equal to the number of IRS elements, limiting efficiency gains. In [29], hybriddriven learning strategies were introduced for CSI acquisition in IRS-aided wideband systems. The authors proposed the DARLAMP network for passive IRS, which integrates modeldriven RLAMP with denoising and attention mechanisms, and further extended it to hybrid IRS through the lightweight MDA-RLAMP network. Generative adversarial networks (GAN) based approaches have also been explored. For instance, the authors in [30] introduced a GAN model for narrowband channel estimation. However, it struggles to reduce training overhead in large-scale
scenarios and cannot be extended to frequency-selective systems. Similarly, [31] used conditional GANs for narrowband fading to enhance the accuracy of the cascade channel. In summary, while prior works have demonstrated significant advancements across classical, CS-based, and DL-based channel estimation techniques, achieving an optimal trade-off between accuracy, overhead, and complexity remains an open challenge, particularly in wideband scenarios. These limitations motivate the development of our proposed deep learning framework, which aims to address these issues holistically. B. Contributions of the Paper While IRS-assisted channel estimation has attracted considerable attention, many existing works rely on oversimplified models and overlook critical challenges such as pilot overhead, computational complexity, limited generalization across varying channel conditions, and estimation accuracy. To bridge these gaps, we propose a novel multi-block attention (MBA) framework that jointly optimizes estimation accuracy, training efficiency, and generalization. Our design is specifically tailored to the structural properties of cascaded IRS channels in wideband OFDM systems. This ensures that both slotdependent and subcarrier-dependent variations are thoroughly captured. The main contributions of this work are summarized as follows: • Pilot Overhead Reduction via Intelligent IRS Element Deactivation: We propose a novel strategy to reduce training overhead in IRS-assisted mmWave MIMOOFDM systems by selectively deactivating a subset of IRS phase shifters (PSs). While deactivation may disrupt spatial correlation and weaken cascade channel observability [32], we address these effects using a DL–based reconstruction framework that recovers lost features. This approach is scalable and reduces pilot costs without compromising estimation accuracy. • Multi-Block Axial Attention-Based Deep Learning Architecture: We design an innovative two-stage deep neural network named MBA that integrates (i) a convolutional attention network (CAN) to reconstruct disrupted spatial correlations and (ii) a complex multi-convolutional network (CMN) for denoising. Unlike conventional architectures, MBA explicitly targets error mitigation and feature recovery from partially observed data. Moreover, our complexity analysis shows that the model scales linearly with IRS size, providing a scalable and efficient alternative to computationally expensive CS-based methods. Simulation results further validate the MBA’s superior performance and strong generalization capability across varying channel distributions. To the best of our knowledge, this is the first IRS channel estimation framework that leverages axial attention to reconstruct deactivated elements, combined with a two-stage recoverydenoising design to suppress error propagation directly. This approach is crucial for efficiency as it avoids the quadratic computational cost of applying standard selfattention to the entire flattened Nt × M channel matrix, a cost that would be prohibitive for large IRS arrays.
3
Optimal Phase Shift Design for LS Estimation: We derive the optimal IRS phase shift configuration for LS estimation in mmWave MIMO systems, demonstrating that DFT and Hadamard matrices minimize channel estimation’s mean squared error (MSE). We analytically show that the LS estimator’s MSE increases linearly with the IRS size, which is extremely inefficient in large-scale IRS deployments. • Error Propagation Mitigation in Two-Stage DNNs: We conduct a theoretical analysis of error propagation in two-stage deep learning models, specifically within the context of IRS-assisted channel estimation [33]. We show how estimation errors accumulate and derive conditions under which our MBA model effectively suppresses this propagation. Both analysis and simulations confirm that MBA achieves significantly lower MSE than LS and state-of-the-art DL baselines.
•
C. Organization and Notation The remainder of this paper is organized as follows. Section II introduces the system model. In Section III, we explain the proposed two-stage DNN architecture. Section IV derives the optimal phase configuration for IRS cascade channel estimation. This is followed by theoretical analysis and complexity analysis of the proposed method. Simulation results are presented in Section V. Finally, Section VI concludes the paper. Notation: We employ boldface uppercase letters (A) to denote matrices, boldface lowercase letters (a) for vectors, and lowercase letters (a) for scalars. The symbols (.)T , (.)H , and (.)−1 signify the transpose, Hermitian, and inverse of matrices, respectively. The notation kAkF and tr{A} represent the Frobenius norm and trace of matrix A, respectively. Furthermore, CM×N denotes the complex space of dimension M × N . Specifically, diag(a) represents a diagonal matrix whose diagonal elements are the components of vector a. Lastly, E [.], shows the statistical expectation operator. II. S YSTEM M ODEL We consider an IRS-assisted MIMO OFDM system operating in time-division duplex (TDD) mode. We assume that each transmission comprises T time slots, where each time slot corresponds to one OFDM symbol (block). The total duration of these time slots is shorter than the channel coherence time, so the channel is assumed to be quasi-static. Furthermore, due to the severe path loss and high susceptibility to blockages in mmWave communications, we assume that the direct link between the BS and MU is completely obstructed, rendering IRS-assisted transmission essential for reliable connectivity. As presented in Fig. 1, the BS is equipped with Nt antennas, while each user, with a single antenna, communicates with the BS through an IRS consisting of M passive elements. Using orthogonal pilot signals, the BS estimates user channels. Without loss of generality, we focus on single-user channel estimation for simplicity. The communication channel between the IRS and the BS is represented by G ∈ CNt ×M , and the channel between the
Figure 1. Schematic of IRS-assisted communication systems, where the direct channel between the BS and the MU is obstructed.
MU and the IRS is denoted by hr ∈ CM×1 . The IRS reflection matrix at the tth time slot is defined as Φ(t) = diag(v(t) ) ∈ CM×M , where v(t) ∈ CM×1 is the IRS phase-shift vector: (t) (t) (t) (t) (t) T (t) (1) v(t) = β1 ejθ1 , β2 ejθ2 , . . . , βM ejθM .
Here, the ith IRS element at the tth time slot is characterized (t) by its phase shift θi ∈ [0, 2π) and amplitude reflection (t) (t) coefficient βi ∈ [0, 1]. Specifically, βi = 0 indicates that (t) the element is deactivated, while βi = 1 corresponds to full reflection [4]. Each IRS element independently controls its phase shift and amplitude. A. Channel Model The channel considered in this paper is a multipath channel whose frequency response on each OFDM subcarrier needs to be estimated. To emulate practical propagation conditions, we adopt the clustered delay line (CDL) channel model specified in the 3GPP TR 38.901 standard based on extensive real-world measurement campaigns conducted in various environments and frequencies. The CDL model effectively captures multipath clustering, angular spreads, and delay dispersion, which are the dominant propagation characteristics in IRS-assisted OFDM systems. Following the guidelines in [34] and the modeling in [35], [29], the IRS-assisted channel is described in the delay domain as follows: r LBS-IRS Nt M X G(τ ) = αi δ (τ − τi ) aBS (φBS i ) LBS-IRS i=1 (2) IRS IRS ×aH IRS (φi , θi ),
hr (τ ) =
r
M LIRS-MU
LX IRS-MU j=1
IRS αj δ(τ − τj )aIRS (φIRS j , θj ), (3)
where LBS-IRS and LMU-IRS present the number of paths between the BS and IRS and between the MU and IRS, respectively. The delay and average power gain of ith path are denoted by τi and αi , respectively where αi ∼ CN (0, σα2 ). The angles φIRS and θiIRS represent the azimuth and elevation i
4
at the IRS. Similarly, φBS i denotes the azimuth AoA or AoD at the BS. Furthermore, the IRS and BS array response vectors are denoted by aIRS and aBS , respectively. In particular, the array response vector of a uniform linear array (ULA) with NULA half-wavelength spaced elements is given by: T 1 1, ejπ sin(φ) , ..., ejπ(N −1)d sin(φ) , (4) aULA (φ) = √ NULA and the array response vector of a half-wavelength spaced uniform planar array (UPA) with NUPA = Nx × Ny elements is given by: 1 1, . . . , ejπ(nx sin(φ) sin(θ)+ny cos(θ) , aUPA (φ, θ) = √ NUPA T jπ((Nx −1) sin(φ) sin(θ)+(Ny −1) cos(θ) , ...,e
(5) where nx and ny index the horizontal and vertical antenna elements. Based on the channel models in (2) and (3), the channel frequency response of the k th subcarrier for the BS–IRS and IRS–user channels is given by: r LBS-IRS Nt M X k H IRS IRS Gk = αi e−j2πτi fs K aBS (φBS i )aIRS (φi , θi ), LBS-IRS i=1 (6) r LX IRS-MU M k IRS hrk = αj e−j2πτj fs K aIRS (φIRS j , θj ), (7) LIRS-MU j=1
where fs denotes the sampling rate, and K represents the total number of subcarriers. B. Cascaded Channel Model In the uplink training phase, the single-antenna user trans(t) mits a pilot signal sk at the tth time slot on the k th subcarrier. (t) The received signal at the BS, yk ∈ CNt ×1 , is expressed as follow: (t) (t) (t) yk = Gk Φ(t) hrk sk + nk (t) (t) = Gk diag(hrk )v(t) sk + nk (t) (t) , Hcsk v(t) sk + nk , (t)
(8)
where nk ∈ CNt ×1 is the additive noise vector at the BS, whose entries are independent and identically distributed circularly symmetric complex Gaussian variables with zero mean and variance σn2 . The cascaded channel matrix Hcsk , Gk diag(hrk ) ∈ CNt ×M captures the joint effect of the BS–IRS channel Gk and the IRS–user channel hrk . Downlink channels are acquired through uplink channel estimation by exploiting channel reciprocity in TDD systems. To estimate Hcsk , the system requires sufficient observations with varying IRS phase shift configurations. Specifically, during B consecutive training slots, the IRS applies distinct phase(t) shift vectors v(t) , while the MU transmits pilot symbols sk . After collecting these B measurements, the BS reconstructs the cascaded channel matrix. A summary of the notation is provided in Table I.
TABLE I S YSTEM M ODEL N OTATIONS Notation Nt M Gk ∈ CNt ×M hrk ∈ CM ×1 Hcsk ∈ CNt ×M Φt ∈ CM ×M vt ∈ CM ×1 Ψ ∈ CM ×B Ŷk ∈ CNt ×B N̂k ∈ CNt ×B
Definition Number of antennas at the BS. Number of elements in the IRS. BS-IRS channel matrix for the k th subcarrier. IRS-MU channel vector for the k th subcarrier. Cascade channel matrix (BS-IRS-MU) for the k th subcarrier. Diagonal reflection matrix of the IRS at time slot t IRS control vector at time slot t. IRS control vectors over B time slots. Received signal matrix over B time slots in the k th subcarrier. The noise matrix over B time slots in the k th subcarrier.
C. Conventional LS Estimator for Cascaded Channel The received signal matrix over B time slots in the k th subcarrier can be presented as: (9) Ŷk = Hcsk Ψ + N̂k , where Ψ , v(1) , . . . , v(B) ∈ CM×B is the IRS phase matrix (1) (B) ∈ CNt ×B is the across B time slots, N̂k , n̂k , . . . , n̂k (1) (B) equivalent noise matrix, and Ŷk , ŷk , . . . , ŷk ∈ CNt ×B denotes the matrix of received signals. Each column is given by: 1 (t) (t) (t) (t) ŷk = (t) yk (sk )∗ = Hcsk v(t) + n̂k , (10) pk (t)
where pk denoting the pilot signal power. The training dataset encompasses data from all subcarriers, eliminating the need to retrain the MBA model for each individual sub-carrier. The model is trained once and subsequently applied across all sub-carriers. As a result, runtime and computational complexity increase linearly with the number of sub-carriers. Using the conventional LS estimator, the cascaded channel Hcsk is then obtained as follows: −1 (11) . ĤLSk = Ŷk × ΨH ΨΨH
It is important to note that the use of the LS estimator, as defined in (11), is valid only under the condition B ≥ M , where B denotes the number of pilot symbols and M represents the number of IRS elements. This constraint becomes particularly significant when the total number of available time slots T (i.e., OFDM symbols) is smaller than M . In such cases, channel estimation using (11) becomes impractical, as the channel may vary across different coherence intervals, violating the quasi-static assumption. Furthermore, the LS estimator is highly sensitive to noise, which can lead to substantial estimation errors. These limitations highlight the necessity of developing more efficient cascaded channel estimation techniques that can reduce training overhead while maintaining estimation accuracy. III. P ROPOSED M ETHOD This section presents the two-stage MBA framework, detailing its architecture, key components, and application to IRSassisted channel estimation.
5
(a)
(b)
(c)
(d)
Figure 2. Comparison of IRS activation patterns, where yellow represents activated elements, and gray represents deactivated ones. (a) Column-wise activation; (b) Row-wise activation; (c) Random scheme activation; (d) Proposed scheme.
A. Activation Pattern for IRS A key challenge in IRS-assisted channel estimation is the high training overhead resulting from a large number of passive elements. To address this, we selectively deactivate certain IRS elements, effectively reducing the dimensionality of the cascaded channel and enabling more efficient channel estimation using an LS estimator. In this framework, the deactivation or activation of the ith (t) IRS element in tth time slot is achieved by setting βi = 0 or (t) βi = 1, respectively. Let B denote the number of activated elements. Since M − B columns of Hcsk , which correspond to deactivated elements, do not contribute to the received signal, we can eliminate them and use the reduced dimension cascade channel between MU and BS denoted by H̃csk ∈ CNt ×B . The received signal over B time slots in k th subcarrier is then given by: Ỹk = H̃csk Ψ̃ + N̂k , (12) where Ψ̃ ∈ CB×B represents the IRS matrix associated with the B activated PS elements. Hence, the LS estimation k th H̃csk is obtained as: H H −1 ˆ Ψ̃Ψ̃ . (13) H̃ LSk = Ỹk Ψ̃
Although deactivating IRS elements effectively reduces pilot overhead, it disrupts spatial correlations and introduces NLOS effects [32]. Thus, we develop a deep learning-based recovery mechanism using the MBA framework to address this issue. Figure 2 illustrates four IRS activation patterns. The proposed scheme ensures that at least one fully active column is maintained while deactivated elements are distributed evenly, preserving local spatial correlation. B. Dataset Construction To construct the dataset, CSI samples are generated for each of the K subcarriers using the channel model in Section II, with channel estimation performed independently per subcarrier. For subcarrier k, the reduced-dimension channel H̃csk is first obtained via the LS estimator in (13). Zero columns are then inserted at the deactivated PS indices to recover the missing entries, yielding the augmented estimate ĤAugk . Since Hcsk is complex-valued, its real and imaginary parts are separated and concatenated into a real-valued tensor of size (Nt , M, 2). Repeating this process over all K subcarriers produces a dataset in which each sample corresponds to the channel of a single subcarrier. This dataset is subsequently
used to train the MBA model to refine the noisy augmented estimate ĤAugk toward the true channel Hcsk . C. Attention Mechanism As stated earlier, to overcome the serious issue of huge overhead in traditional channel estimation approaches, we deactivate some IRS elements at the cost of disrupting the spatial correlation. To solve this problem, we employ an attention mechanism as part of our proposed channel estimation method for the following reasons: Spatial correlation compensation: The attention mechanism can capture long-range spatial correlation by dynamically reweighting features [36]. This can improve the estimation accuracy in dynamic environments [36]. • Adaptability: The attention mechanism can model global dependencies. Thus, they can effectively handle scattered paths and non-stationary channels. • Parallelism: Another key advantage of these mechanisms, specifically self-attention, is parallelism. Unlike sequential processing, self-attention allows parallel computation, significantly accelerating training and inference. •
The self-attention mechanism that we incorporate relies on three key components: Key (K), Query (Q), and Value (V). The input X ∈ RNt ×M is first transformed using three multilayer perceptrons (MLPs) to generate the corresponding weights in the original attention mechanism [37]. To reduce complexity, we replace traditional fully connected layers in self-attention with multi-convolutional blocks (MB), which better exploit spatial correlations in IRS-based channel estimation. Compared to MLPs, convolutional neural networks (CNNs) require less training data, exhibit stronger noise robustness, and are better suited for structured inputs [38]. While attention mechanisms are effective for modeling long-range dependencies, most IRS-based tasks primarily rely on local correlations, making CNNs a more practical choice. In our design, convolutional layers improve efficiency and reduce self-attention complexity. In the proposed model, to address the challenge of missing channel information resulting from the deactivation of IRS elements, we employ a specialized deep learning mechanism known as axial attention [39]. The primary goal of this mechanism is to intelligently reconstruct the deactivated elements by learning the spatial relationships and interdependencies from the active ones.
6
The process begins with the initial, incomplete channel estimate, which is structured as a three-dimensional tensor, denoted as X. This tensor has dimensions of Nt × M × 2, where Nt represents the number of antennas at BS, M is the total number of reflecting elements on the IRS, and the final dimension of size 2 separates the real and imaginary parts of the complex channel values for processing by the neural network. Next, the model transforms the input tensor X into three distinct tensors: Q, K, and V. These tensors are generated by projecting the input X through learned weight matrices (Wq , Wk , Wv ), as shown below: Q = XWq ,
K = XWk ,
V = XWv .
Figure 3. Structure of the Attention Block (AB), combining a self-attention mechanism and a multi-convolutional block (MB) to enhance feature extraction.
(14)
Each resulting tensor has dimensions Nt ×M ×d, where d is a learned feature dimension. The core of the mechanism lies in calculating attention scores to determine the focus each IRS element should place on every other element. A relevance score is computed by comparing the Query of one element with the Key of all other elements via a dot product, [QKT ]n = Qn KTn . Importantly, this calculation is carried out independently for each of the Nt BS antennas. This approach concentrates the model’s learning capacity along the IRS element dimension (M ), precisely where information was lost. The resulting raw √ scores are scaled by 1/ d for numerical stability and then passed through a Softmax function. This function converts the scores into a set of weights that sum to √ one, yielding the final attention tensor S = Softmax(QKT / d) with dimensions Nt × M × M . To formalize the IRS indexing, let i = (r, c) denote the ith IRS element located at the rth row and cth column of the 2D planar IRS. Accordingly, an entry Sn,i,j quantifies the importance of the element at position j = (r′ , c′ ) in reconstructing the element at position i = (r, c) for the channel link associated with the nth BS antenna. Finally, these attention weights are used to generate a refined and complete channel estimate. The attention tensor S is applied to the Value tensor (V). This process enables a deactivated element, initially represented by zeros, to be reconstructed using an informed combination of features from the most relevant active elements. The final output, O, is obtained by O = SV. As a result, the attention mechanism is computed as follows: QKT (15) Attention (X) = Softmax √ V. d The resulting tensor O represents the reconstructed channel with the missing information filled in. This axial approach is significantly more computationally efficient than standard attention mechanisms, ensuring the model remains scalable for large-scale IRS deployments. Finally, the schematic representation of the attention mechanism is shown in Fig. 3. D. Multi-Block Attention (MBA) As presented in Fig. 4, the proposed MBA framework comprises two primary components: (i) the CAN for restoring spatial correlations and (ii) the CMN for mitigating noise. The CAN model is initially trained in isolation, after which its parameters are frozen before training the CMN module.
The CAN architecture comprises two attention blocks (ABs) reconstructing inactive channels by capturing the missing spatial correlations. Each attention block Ai (i = 1, 2) is composed of a convolutional layer, a ReLU activation, and a self-attention mechanism, as formulated below: Ai = Iai + Conv(Attention(Iai )),
(16)
where Iai denotes the input to the attention mechanism. The output representation from the CAN module is mathematically expressed as follows: P1 = ReLU Conv(ĤAugk ) , A1 = AB1 (P1 ) + P1 , (17) P2 = ReLU Conv(A1 ) , A2 = AB2 (P2 ) + P2 , ĤCANk = ĤAugk − ReLU(Conv(A2 )), where ĤAugk refers to the channel matrix with augmented zero columns described in Section III.B, and P1 , P2 serve as the inputs to AB1 and AB2 , respectively. The second component, CMN, suppresses noise and residual distortions in the channel estimates produced by CAN. It is designed with a deep-layered architecture incorporating residual connections to counteract vanishing gradient issues during training. The CMN module consists of two convolutional blocks CBi (i = 1, 2), each defined as: CBi (Ii ) = BN(Conv(PReLU(BN(Conv(Ii ))))) + Ii ,
(18)
where BN denotes batch normalization, Ii is the input, and the output of each CB is added to its input as a residual connection. The full output of the CMN module is computed as: OCMN = Conv BN(Conv(CB2 (CB1 ( (19) PReLU(Conv(ĤCANk )))))) + Conv(ĤCANk ),
where OCMN denotes the refined channel estimation output. The final output of the MBA model is obtained by sequentially applying CAN and CMN, ensuring robust reconstruction and denoising. This process is described as: (20) ĤMBAk = OCMN = FCMN FCAN (ĤAugk ) ,
where FCAN (·) and FCMN (·) represent the learned transformations of the CAN and CMN networks, respectively. This
7
By substituting (11) into the objective function, we obtain: i h JLS = E ||Hcs − (Hcs Ψ + N) Ψ† ||2F h i (23) = E ||NΨ† ||2F o n = tr (ΨΨH )−1 ΨE NH N ΨH (ΨΨH )−1 .
Since each of the Nt rows of N has a covariance of σn2 IM , we have E[NH N] = Nt σn2 IM . Therefore, in equation (23), by substituting this result, we obtain: o n JLS = Nt σn2 tr (ΨΨH )−1 ΨIM ΨH (ΨΨH )−1 o n (24) = Nt σn2 tr (ΨΨH )−1 . The optimization problem is now reformulated as follows: o n min tr (ΨΨH )−1 Ψ
s.t.
Figure 4. Overview of the proposed MBA framework.
hierarchical design guarantees structural consistency and robustness to noise in the estimated channels. To optimize both networks, we adopt the MSE loss. The loss functions for CAN and CMN are defined as: D
2 (i) 1 X H(i) , LCAN = csk − ĤCANk 2D i=1 F
(21)
D
2 1 X (i) H(i) , LCMN = csk − OCMNk 2D i=1 F
(22)
where D is the total number of training samples. IV. T HEORETICAL A NALYSIS This section presents two key theoretical contributions to IRS-assisted channel estimation. First, we derive the optimal phase for the IRS elements for the channel estimation. Second, we analyze the error propagation phenomenon in twostage DNNs and theoretically demonstrate how our proposed architecture effectively suppresses this detrimental effect. A. On the optimality of IRS matrix As previously mentioned, it is essential to determine the phase of the PS in the IRS for each time slot during the estimation stage. The phase of the PS in the LS estimator, as indicated in (11), plays a crucial role in the accuracy of channel estimation. Therefore, optimizing the IRS phase shifts in each time slot is vital for minimizing the MSE in the following optimization problem: h i LS min JLS = E ||Hcs − Ĥcs ||2F Ψ
s.t.
|Ψ(i, j)| = 1,
i, j = 1, 2, ..., M.
|Ψ(i, j)| = 1,
i, j = 1, 2, ..., M.
Minimizing the trace of the inverse requires the eigenvalues of ΨΨH to be uniformly distributed. The unimodular constraint requires |Ψ(i, j)| = 1, implying Ψ(i, j) = ejθij , with elements on the unit circle [40]. The optimal Ψ leverages equiangular tight frame (ETF) properties, which minimize the trace of the inverse of Gram unit-norm n matrices under o constraints. Thus, minimizing tr (ΨΨH )−1 for a unimodular Ψ involves constructing Ψ as an ETF, ensuring uniform eigenvalues of ΨΨH . Therefore, we choose Ψ as follows: ΨΨH = IM .
(25)
Two well-known examples of an ETF are the DFT and the Hadamard matrix. The Hadamard matrix is particularly advantageous when the PSs of the IRS are quantized, as it maintains a structured and unimodular design. When the DFT matrix is selected as the Ψ matrix, the MSE for the LS estimator can be expressed as follows: JLS = M Nt σn2 .
(26)
As we can see, the LS estimator’s MSE scales linearly with IRS size due to the accumulation of noise across the elements. B. Error propagation analysis The conventional LS estimator provides a simple initial channel estimate but is highly sensitive to noise, particularly when IRS elements are deactivated during estimation. This leads to substantial estimation error, denoted as JLS . To address this issue, we introduce the MBA framework, a twostage DL approach that enhances estimation robustness more effectively than single-stage models [33]. The first stage, the CAN, refines the LS estimate using self-attention to recover lost spatial features. Its output is modeled as follows: ĤCAN = ĤLS + λCAN (Hcs − ĤLS ) + εCAN ,
(27)
where λCAN is the feature recovery gain, and εCAN denotes the residual error due to imperfect recovery. This structure is analogous to classical error correction models, where the
8
network learns to reconstruct missing information based on prior estimates. The term Hcs − ĤLS represents the estimation error, guiding CAN to refine rather than relearn the channel. If λCAN = 1, CAN ideally recovers the complete channel, while λCAN ≤ 0 implies no improvement or degradation. By substituting the LS estimation model into (27), the residual error after CAN processing is given by εCAN = (1 − λCAN )N̂.
(28)
In this study, the NMSE serves as the performance metric; therefore, the feature recovery gain of CAN is defined as: λCAN =
NMSELS − NMSECAN , NMSELS
(29)
where the numerator quantifies the reduction in estimation error, and the normalization ensures scale invariance across different configurations. This approach aligns with the denoising autoencoder framework proposed in [41], which measures robustness improvement via reductions in reconstruction error. Following CAN processing, the refined estimate ĤCAN is fed into the CMN module, which further suppresses residual noise using deep convolutional filtering. The output of CMN is modeled as follows: ĤCMN = ĤCAN + λCMN (Hcs − ĤCAN ) + εCMN ,
(30)
where λCMN captures the noise suppression gain, and εCMN is the residual error. Similar to CAN, the residual error after CMN can be expressed as: εCMN = (1 − λCMN )εCAN = (1 − λCMN )(1 − λCAN )N̂. (31) The noise suppression gain λCMN is defined analogously to (29) as follow: NMSECAN − NMSECMN . λCMN = NMSECAN
(32)
This expression follows the noise reduction principles observed in deep convolutional denoising networks [42]. The overall estimation error after both CAN and CMN stages can thus be characterized as follows: i h JMBA = E kHcs − ĤCMN k2F h i = E k(1 − λCMN )(1 − λCAN )N̂k2F (33) h i 2 2 2 = (1 − λCMN ) (1 − λCAN ) E kN̂kF = (1 − λCAN )2 (1 − λCMN )2 JLS .
This formulation shows that as λCAN , λCMN → 1, the error JMBA → 0, indicating near-optimal channel reconstruction. Simulations validate this trend, demonstrating the superior performance of MBA over LS in terms of estimation accuracy. Moreover, increased pilot signal availability drives λCAN and λCMN toward unity.
C. Computational Complexity Analysis Evaluating computational complexity is critical for assessing the real-time feasibility of IRS-based channel estimation techniques. This issue is crucial in wireless systems, where applications demand extremely low latencies. Here, we analyze the computational complexity of our proposed MBAbased channel estimation algorithm. Following the analysis in [43], the order of complexity for a convolutional layer and batch normalization is given by Cconv = O XY NI NO S 2 , CBN = O(XY NI ) ,
where X and Y denote the dimensions of the output feature map, in our case, X = Nt and Y = M correspond to the number of BS antennas and the number of IRS elements, respectively. The term S denotes the convolutional kernel size, and NI and NO represent the convolutional layer’s input and output channel numbers. In addition to convolutional modules, attention mechanisms also play a role in the overall computational cost. The complexity of a single attention block can be expressed as CAttention = O(M Nt NI ) , where NI represents the feature dimension used for query, key, and value projections [37]. Combining the convolutional layers, batch normalization, and attention blocks can systematically derive the complexities of the CAN, CMN, and MB modules. Consequently, the total computational complexity of the MBA estimator across all K subcarriers is given by LAB h X X M Nt NI,i CMBA = O K Nt M s2j nj−1 nj + j∈CAN
X
i=1
X
s2j nj−1 nj + Nt M s2j nj−1 nj + 6Nt M j∈MB j∈CMN
! i ,
where sj denotes the convolutional kernel size of the j th layer, nj is the number of channels, LAB is the number of attention blocks, and NI,i is the input dimension of the ith attention block. Here, the notation j ∈ CAN (or j ∈ CMN, j ∈ MB) indicates that the summation is performed over all convolutional layers belonging to the respective module. It is important to emphasize that this complexity scales linearly with Nt , M , and K. Unlike iterative optimizationbased methods whose costs grow with the pilot length or number of iterations, the MBA method maintains a fixed per-subcarrier cost. This results in improved computational efficiency compared to state-of-the-art approaches, thereby reducing channel estimation time and making MBA especially attractive for latency-sensitive wireless communication scenarios. V. S IMULATION R ESULTS This section presents comparative simulations for MBA and state-of-the-art channel estimation methods. To have a realistic scenario, we consider 3GPP urban micro (UMi) with the CDL model.
9
TABLE II C OMPARISON OF NMSE PERFORMANCE FOR DIFFERENT ACTIVATION PATTERNS WITH AN IRS. Number of Activated Elements B = 24 B = 48
Column Pattern SNR=20 dB SNR=10 dB 0.031148 0.044473 0.001802 0.015167
Row Pattern SNR=20 dB SNR=10 dB 0.029091 0.044093 0.001798 0.015053
Random Pattern SNR=20 dB SNR=10 dB 0.024032 0.036665 0.000826 0.004651
Proposed Pattern SNR=20 dB SNR=10 dB 0.011532 0.0166701 0.000160 0.0014953
A. Channel Model and Simulation Parameters CAN B = 24, SNR = 10 dB CAN B = 24, SNR = 20 dB CAN B = 48, SNR = 10 dB CAN B = 48, SNR = 20 dB
100
10-1
NMSE
We adopt the 3GPP TR 38.901 Release 17 CDL model to simulate a mmWave channel in a UMi scenario at a carrier frequency of fc = 28 GHz [34]. The BS is equipped with the ULA featuring Nt = 16 transmit antennas and an IRS composed of a M = 12 × 12 UPA with passive reflecting elements. Each channel realization includes LIRS-BS = 4 multipath components for the IRS-BS link and LMU-IRS = 10 for the MU-IRS link. The root-mean-square delay spread (DS), denoted as XDS , follows a log-normal distribution in the UMi scenario. The parameters for it are derived as follows:
10-2
10-3
µlog10 XDS = −0.24 log10 (1 + fc ) − 6.83, σlog10 XDS = 0.16 log10 (1 + fc ) + 0.28,
10-4
0
50
100
150
200
250
(a)
where fc is expressed in GHz. Additionally, Path delays are generated as follows:
CMN B = 24, SNR = 10 dB CMN B = 24, SNR = 20 dB CMN B = 48, SNR = 10 dB CMN B = 48, SNR = 20 dB
100
τl = −rτ · XDS · log(Ul ),
We compare the MBA against several existing estimation techniques, including LS, SMJCE [14], RS-OMP [26], DDFS [27], SRDnNet [28], DA-RLAMP [29], and SFCNN [35].
10-1
NMSE
where rτ = 2.1 and Ul ∼ U(0, 1) is a uniform random variable. Furthermore, all angular spreads depend on the frequency and have log-normal distributions, as detailed in Table 7.5-6 of the UMi scenario in [34]. We trained all models on the same dataset across various SNR levels. The same channel data was used to generate received pilot signals at various SNRs, which were then employed for training. The test set was generated similarly. In total, 80,000 channel samples were synthesized, with 50,000 for training, 20,000 for validation, and 10,000 for testing. We set a batch size of 64 for MBA model training and employ the Adam optimizer for both CAN and CMN, with parameters β1 = 0.9 and β2 = 0.999 [44]. The CAN model adopts a learning rate scheduler that reduces the learning rate by 0.6 every 150 epoch, starting from 0.0002, while the CMN model uses a fixed learning rate of 0.0001. All deep learning-based methods, including our proposed model, were also trained using an NVIDIA Tesla T4 GPU and implemented with the PyTorch framework. As previously mentioned, we employ the NMSE as the performance metric, defined as follows [14]: # " kHcs − ĤMBA k2F (34) NMSE = E kHcsk2F
300
Epochs
10-2
10-3
10-4
0
10
20
30
40
50
60
70
80
90
100
Epochs
(b) Figure 5. Convergence of the proposed model in terms of NMSE over epochs. (a) CAN model; (b) CMN model.
B. Analysis of MBA We begin by analyzing the convergence behavior of the CAN and CMN models, which are the components of the MBA, separately. Figure 5 presents the NMSE convergence of the two models under SNRs of 10 dB and 20 dB on the test set. The CAN model converges within 100 epochs for B = 24, while increasing the number of pilot signals to B = 48 extends the training time to approximately 300 epochs. In contrast, the CMN model exhibits rapid and stable convergence across all cases, indicating strong robustness in noise suppression. To assess the impact of IRS activation strategies, we evaluate MBA under various IRS element deactivation patterns. As
10
TABLE III 101
ANALYSIS OF FEATURE RECOVERY AND NOISE SUPPRESSION GAIN
B = 12 B = 18 B = 24 B = 36 B = 48 B = 72
λCAN SNR=10 dB SNR=20 dB 0.5175 0.5535 0.8345 0.8541 0.9508 0.9544 0.9871 0.9905 0.9950 0.9988 0.9954 0.9993
λCMN SNR=10 dB SNR=20 dB 0.1980 0.2521 0.4350 0.5121 0.6028 0.6972 0.6555 0.7903 0.6741 0.8935 0.7072 0.9150
100
10-1
NMSE
Training Signal
10-2
40
10-3
LS MBA
35 10-4 30
-5
SMJCE [14] DA-RLAMP [29] RS-OMP [26] SFCNN [35] DD-FS [27] LS SRDnCNN [28] MBA
0
5
10
15
20
25
30
15
20
25
30
SNR (dB)
MSE
25
(a) 20 101 15
10
100
5 10-1 64
100
144
196
256
IRS size (M)
NMSE
0 36
10-2
Figure 6. MSE comparison between the LS estimator and the MBA model for different IRS sizes. 10-3
shown in Table II, our proposed activation scheme achieves the lowest NMSE. This is due to preserving a higher degree of spatial correlation compared to both random and structured deactivation baselines. Given these results, the proposed pattern is adopted in the rest of the simulations. As explained in the theoretical analysis in Section IV.B, the parameters λCAN and λCMN characterize the feature recovery and noise suppression capabilities of CAN and CMN, respectively. Table III presents these values for different pilot signal counts at SNRs of 10 dB and 20 dB. As observed, both λCAN and λCMN approaches 1 as the number of pilot signals increases, which correlates with a reduction in channel estimation error. Figure 6 illustrates the MSE performance of both the LS estimator and the proposed MBA model across different IRS sizes. For this experiment, the number of training signals for MBA is fixed at B = 48. The results confirm that while the MSE of the LS estimator increases linearly with IRS size, the MBA model consistently achieves lower MSE due to its superior feature recovery and denoising capabilities. This demonstrates that our model not only enhances estimation accuracy but also reduces the dependence on extensive pilot training, making it suitable for large-scale IRS-assisted systems. C. Performance Comparison of Different Estimation Schemes To evaluate the effectiveness of the proposed method, we compare the NMSE across different SNR levels with several
10-4 -5
SMJCE [14] SFCNN [35] DA-RLAMP [29] RS-OMP [26] LS SRDnCNN [28] DD-FS [27] MBA
0
5
10
SNR (dB)
(b) Figure 7. NMSE comparison of different methods for Nt = 16 and M = 144: (a) B = 36; (b) B = 48.
state-of-the-art baseline methods, as shown in Fig. 7. The evaluation considers pilot lengths of B = 36 and B = 48 for methods capable of operating under reduced pilot overhead. As expected, the LS estimator is highly noise-sensitive, leading to severe NMSE degradation at low SNRs. It is important to highlight that LS, SFCNN, and SRDnNet require the number of pilots (B) to exceed the number of IRS elements (M ). Therefore, in Fig. 7, where M = 144, these methods do not perform properly when B = 36 or 48. The corresponding performance curves for these methods are generated for B = 144 and are presented solely for comparative purposes. Despite the substantially greater pilot overhead associated with LS, SFCNN, and SRDnNet (144 as opposed to 36 or 48), the MBA algorithm consistently demonstrates superior performance. The RS-OMP, SRDnNet, DD-FS, and DA-RLAMP methods demonstrate stronger noise robustness than the SMJCE technique by exploiting learning-based mechanisms for effective noise suppression. In particular, the DD-FS model further enhances estimation accuracy by incorporating a dual-domain
11
D. Effect of Pilot Overhead on Estimation Accuracy and Model Generalization This subsection analyzes the impact of the number of pilot signals on channel estimation performance and evaluates the generalization capabilities of various estimation techniques, including SMJCE, RS-OMP, DD-FS, DA-RLAMP, and the proposed MBA model. Specifically, we investigate how the number of pilot signals affects the NMSE at two SNR levels: 10 dB and 20 dB. As shown in Fig. 8, increasing the number of pilot signals consistently enhances estimation accuracy across all methods evaluated. The CS-based approaches, including DD-FS, RS-OMP, DARLAMP, and SMJCE, exhibit noticeable NMSE reductions as the number of pilots increases from B = 72 to B = 144, with SMJCE showing the most pronounced improvement due to the availability of more measurements that enhance CS recovery. In addition, the accuracy of the DA-RLAMP approach is poor with few training signals; however, the accuracy improves significantly once the number of pilots increases, with steeper NMSE reductions between B = 48 to B = 72 and B = 72 to B = 144. At an SNR of 10 dB, the proposed MBA method achieves substantially lower NMSE, outperforming DD-FS by 51%, RS-OMP by 76%, DA-RLAMP by 62%, and SMJCE by 90%. To further assess the robustness and generalization ability of the estimation methods, we evaluate their performance across different propagation environments. Here, generalization refers to the ability of a model trained under one scenario to maintain reliable performance when applied to unseen environments. According to the 3GPP TR 38.901 specification [34], the propagation characteristics, such as power delay profiles and angular spreads, vary significantly between the UMi and urban macro (UMa) scenarios. To this end, all models, including MBA, DD-FS, DA-RLAMP, and RS-OMP, are trained using channel data generated under the UMi scenario and then tested on the UMa scenario without any adaptation or prior knowledge of the UMa channel statistics. As illustrated in Fig. 9, the proposed MBA model and DARLAMP demonstrate superior generalization performance. Since RS-OMP employs an initial channel estimation stage using CS-based algorithms to generate inputs for its DNN, it
100 RS-OMP: 10 dB SMJCE: 10 dB DD-FS: 10 dB DA-RLAMP: 10 dB MBA: 10 dB LS: 10 dB
NMSE
10-1
SMJCE: 20 dB RS-OMP: 20 dB DD-FS: 20 dB DA-RLAMP: 20 dB MBA: 20 dB LS: 20 dB
10-2
10-3
10-4
12 18 24
36
48
72
144
Training overhead (B)
Figure 8. Impact of the number of training signals on NMSE performance with: Nt = 16, M = 144.
100 LS ( UMa ) RS-OMP ( UMa ) SMJCE ( UMa ) DD-FS ( UMa ) DA-RLAMP ( UMa ) MBA ( UMa )
10-1
NMSE
sparsity structure. As illustrated in Fig. 7(b), increasing the number of pilot signals significantly enhances the performance of the proposed MBA model, especially at SNR levels above 10 dB. Comparing SRDnNet and DD-FS, we observe that increasing the pilot number from 36 to 48 leads to a more substantial performance gain for DD-FS. This improvement is attributed to the availability of more measurements, enabling DD-FS to surpass SRDnNet under these conditions. Furthermore, the DA-RLAMP approach exhibits strong dependence on the number of BS measurements; increasing the pilot signals from 36 to 48 significantly improves its estimation accuracy. However, beyond approximately 15 dB SNR, DA-RLAMP reaches a performance plateau, unlike other methods that continue to improve.
LS ( UMi ) SMJCE ( UMi ) RS-OMP ( UMi ) DD-FS ( UMi ) DA-RLAMP ( UMi ) MBA ( UMi )
10-2
10-3
10-4
12 18 24
36
48
72
144
Training overhead (B)
Figure 9. Generalization analysis of different models in UMa and UMi scenarios with: Nt = 16, M = 144.
also achieves a reasonable level of generalization. In contrast, the DD-FS method suffers from more severe performance degradation than RS-OMP. Similar to RS-OMP, DA-RLAMP leverages a pre-estimation stage, enabling its DNN model to generalize well across different scenarios. Nevertheless, the estimation accuracy of DA-RLAMP remains considerably lower than that of the proposed MBA, with reductions of approximately 73% in the UMi scenario and about 62% in the UMa scenario. E. Effect of the Number of BS Antennas and IRS Size on Channel Estimation Performance Figure 10 illustrates the impact of the number of BS antennas on NMSE under two pilot overhead settings: B = 36 and B = 48, with the SNR fixed at 20 dB. The results show that increasing the number of BS antennas has a limited effect on estimation accuracy. A modest improvement is observed when the number of antennas increases from 4 to
12
TABLE IV C OMPUTATIONAL COMPLEXITY ORDER OF DIFFERENT APPROACHES .
100 SMJCE: B = 36 RS-OMP: B = 36 DA-RLAMP: B = 36 DD-FS: B = 36 MBA: B = 36 LS
NMSE
10-1
SMJCE: B = 48 RS-OMP: B = 48 DA-RLAMP: B = 48 DD-FS: B = 48 MBA: B = 48 SRDnCNN
Method MBA RS-OMP SMJCE SRDnNet DA-RLAMP SFCNN DD-FS LS
10-2
Complexity Order CMBA ≈ O(2.5 × 105 Nt M K) CRS-OMP ≈ O((4 × 105 Nt M + Nt M B)K) CSMJCE ≈ O(((M 2 (LBS-IRS + 2B))K) CSRDnNet ≈ O(13.7 × 105 Nt M K) CDA-RLAMP ≈ O(Nt M × (48BK + 9.5 × 106 )) CSFCNN ≈ O(3 × 105 Nt M K) CDD-FS = O(B(Nt + M ) + 2.6 × 105 Nt M K) CLS = O(Nt M 2 K)
Estimation Time 41.77 ms 281.12 ms 2512.32 ms 55.43 ms 63.87 ms 10.54 ms 66.43 ms 16.59 ms
10-3
10-4 4
9
16
25
36
49
64
Number of BS Antennas ( Nt )
Figure 10. Effect of the number of BS antennas on NMSE performance: M = 144
creases, the performance of DD-FS and DA-RLAMP degrades significantly, revealing its limited scalability and adaptability. This degradation is even more pronounced in DA-RLAMP, particularly for larger IRS sizes. In contrast, the NMSE performance of LS and SRDnNet remains relatively stable as the IRS size increases. This indicates that these methods implicitly distribute pilot resources proportionately to the IRS size, maintaining consistent performance across different configurations.
101
10
0
DA-RLAMP: B = 36 SMJCE: B = 36 DD-FS: B = 36 RS-OMP: B = 36 MBA: B = 36 LS
SMJCE: B = 48 DA-RLAMP: B = 48 DD-FS: B = 48 RS-OMP: B = 48 MBA: B = 48 SRDnCNN
F. Computational Complexity and Processing Time Analysis
NMSE
10-1
10-2
10-3
10-4
36
64
100
144
196
256
IRS size (M)
Figure 11. Impact of the number of elements in the IRS on NMSE performance: Nt = 16
16, particularly for DL-based methods. This is likely because smaller input dimensions limit the effectiveness of feature extraction. Overall, these results suggest that the IRS size has a more pronounced influence on cascaded channel estimation performance than the number of BS antennas. Figure 11 presents the effect of IRS size on NMSE performance under the same pilot overhead settings and SNR. As the number of IRS elements increases, more pilot signals are required to preserve estimation accuracy. This trend reveals a key limitation of CS-based methods, where larger IRS sizes necessitate proportional dictionary expansion, significantly increasing computational complexity. A comparison between DD-FS and the proposed MBA model shows that DD-FS and DA-RLAMP slightly outperform MBA when the IRS size is B = 48. This observation is consistent with the fact that DD-FS performs better when the number of pilot signals exceeds the number of IRS elements. However, as the IRS size increases or the number of pilots de-
Computational complexity plays a critical role in channel estimation, directly impacting estimation latency and the feasibility of real-time deployment. Table IV presents a comparative analysis of the computational complexity for the proposed MBA model and several benchmark methods. Notably, the complexity of MBA, as well as that of SFCNN, DA-RLAMP, and SRDnCNN, scales linearly with the number of BS antennas (Nt ), the number of IRS elements (M ), and the number of subcarriers (K), making them well-suited for large-scale deployments. In contrast, methods such as RS-OMP, DD-FS, and SMJCE exhibit computational burdens that additionally depend on the number of pilot symbols (B), resulting in significantly higher complexity under more intensive training conditions. To assess the practical impact of computational complexity, we fixed the pilot overhead at B = 48 and measured the average estimation time over 10,000 test samples across all subcarriers. The proposed MBA framework achieves significant computational savings, reducing runtime by about 85% compared to RS-OMP, 37% compared to DD-FS, 34% compared to DA-RLAMP, and 24% compared to SRDnCNN, while still delivering higher estimation accuracy than all fast baseline methods. Although SFCNN and LS achieve faster processing, their practicality is limited since they require pilot overhead equal to or larger than the number of IRS elements. VI. C ONCLUSION This work introduces a deep learning-based framework for efficient cascaded channel estimation in IRS-assisted mmWave MIMO systems. We propose an IRS element deactivation strategy that selectively reduces training dimensionality while preserving spatial coherence to address the significant pilot overhead associated with large IRS deployments. We analytically demonstrate that DFT and Hadamard matrices provide optimal phase configurations for minimizing estimation error
13
in LS estimation. To mitigate the challenges of spatial decorrelation and noise amplification arising from deactivation, we develop the MBA architecture comprising two modules: the CAN for feature restoration and the CMN for denoising. The MBA framework is analytically shown to suppress error propagation and exhibits linear complexity with respect to IRS and BS dimensions. Extensive simulations conducted under standardized mmWave channel models demonstrate that the proposed method can estimate the cascade channel approximately 24% faster than state-of-the-art methods due to its lower complexity. Additionally, it achieves up to 87% less pilot overhead compared to LS estimation. Furthermore, the MBA model effectively generalizes to different deployment scenarios, confirming its scalability and practical viability. Future research will focus on extending this work to multi-IRS and distributed IRS architectures, aiming for more dynamic and flexible configurations. R EFERENCES [1] Q. Wu, S. Zhang, B. Zheng, C. You, and R. Zhang, “Intelligent reflecting surface-aided wireless communications: A tutorial,” IEEE Transactions on Communications, vol. 69, no. 5, pp. 3313–3351, 2021. [2] S. Gong, X. Lu, D. T. Hoang, D. Niyato, L. Shu, D. I. Kim, and Y.-C. Liang, “Toward smart wireless communications via intelligent reflecting surfaces: A contemporary survey,” IEEE Communications Surveys & Tutorials, vol. 22, no. 4, pp. 2283–2314, 2020. [3] M. Di Renzo, K. Ntontin, J. Song, F. H. Danufane, X. Qian, F. Lazarakis, J. De Rosny, D.-T. Phan-Huy, O. Simeone, R. Zhang, M. Debbah, G. Lerosey, M. Fink, S. Tretyakov, and S. Shamai, “Reconfigurable intelligent surfaces vs. relaying: Differences, similarities, and performance comparison,” IEEE Open Journal of the Communications Society, vol. 1, pp. 798–807, 2020. [4] Q. Wu and R. Zhang, “Intelligent reflecting surface enhanced wireless network via joint active and passive beamforming,” IEEE Transactions on Wireless Communications, vol. 18, no. 11, pp. 5394–5409, 2019. [5] M. A. ElMossallamy, H. Zhang, L. Song, K. G. Seddik, Z. Han, and G. Y. Li, “Reconfigurable intelligent surfaces for wireless communications: Principles, challenges, and opportunities,” IEEE Transactions on Cognitive Communications and Networking, vol. 6, no. 3, pp. 990–1002, 2020. [6] B. Zheng and R. Zhang, “Intelligent reflecting surface-enhanced OFDM: Channel estimation and reflection optimization,” IEEE Wireless Communications Letters, vol. 9, no. 4, pp. 518–522, 2020. [7] D. Mishra and H. Johansson, “Channel estimation and low-complexity beamforming design for passive intelligent surface assisted MISO wireless energy transfer,” in ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, UK, May 2019, pp. 4659–4663. [8] T. L. Jensen and E. D. Carvalho, “An optimal channel estimation scheme for intelligent reflecting surfaces based on a minimum variance unbiased estimator,” in ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain, May 2020, pp. 5000–5004. [9] Y. Yang, B. Zheng, S. Zhang, and R. Zhang, “Intelligent reflecting surface meets OFDM: Protocol design and rate maximization,” IEEE Transactions on Communications, vol. 68, no. 7, pp. 4522–4535, 2020. [10] B. Zheng, C. You, and R. Zhang, “Intelligent reflecting surface assisted multi-user OFDMA: Channel estimation and training design,” IEEE Transactions on Wireless Communications, vol. 19, no. 12, pp. 8315– 8329, 2020. [11] H. Alwazani, A. Kammoun, A. Chaaban, M. Debbah, M.-S. Alouini et al., “Intelligent reflecting surface-assisted multi-user MISO communication: Channel estimation and beamforming design,” IEEE Open Journal of the Communications Society, vol. 1, pp. 661–680, 2020. [12] J. He, H. Wymeersch, and M. Juntti, “Channel estimation for RISaided mmWave MIMO systems via atomic norm minimization,” IEEE Transactions on Wireless Communications, vol. 20, no. 9, pp. 5786– 5797, 2021. [13] M. Liu, T. Lin, Y. Zhu, and Y.-J. A. Zhang, “Sparse channel estimation for IRS-assisted millimeter wave MIMO OFDM systems,” IEEE Transactions on Communications, vol. 72, no. 10, pp. 6553–6568, 2024.
[14] J. Chen, Y.-C. Liang, H. V. Cheng, and W. Yu, “Channel estimation for reconfigurable intelligent surface aided multi-user mmWave MIMO systems,” IEEE Transactions on Wireless Communications, vol. 22, no. 10, pp. 6853–6869, 2023. [15] K. F. Masood, J. Tong, J. Xi, J. Yuan, and Y. Yu, “Inductive matrix completion and root-MUSIC-based channel estimation for intelligent reflecting surface (IRS)-aided hybrid MIMO systems,” IEEE Transactions on Wireless Communications, vol. 22, no. 11, pp. 7917–7931, 2023. [16] H. Chung, S. Hong, and S. Kim, “Efficient multi-user channel estimation for RIS-aided mmWave systems using shared channel subspace,” IEEE Transactions on Wireless Communications, vol. 23, no. 8, pp. 8512– 8527, 2024. [17] F.-S. Tseng, H. Chung, and P.-H. Wu, “Efficient channel estimation for millimeter-wave RIS-assisted SIMO-OFDM systems with reduced complexity and pilot overhead,” IEEE Transactions on Vehicular Technology, vol. 73, no. 10, pp. 14 809–14 821, 2024. [18] Z. Huang, C. Liu, Y. Song, and Y. Fu, “A low-rank sparse tensor recovery-based channel estimation scheme for RIS-assisted mmWave OFDM systems,” IEEE Transactions on Vehicular Technology, pp. 1– 16, 2025. [19] Y. C. Eldar, A. Goldsmith, D. Gündüz, and H. V. Poor, Machine learning and wireless communications. Cambridge University Press, 2022. [20] M. Chen, U. Challita, W. Saad, C. Yin, and M. Debbah, “Artificial neural networks-based machine learning for wireless networks: A tutorial,” IEEE Communications Surveys & Tutorials, vol. 21, no. 4, pp. 3039– 3071, 2019. [21] T. O’Shea and J. Hoydis, “An introduction to deep learning for the physical layer,” IEEE Transactions on Cognitive Communications and Networking, vol. 3, no. 4, pp. 563–575, 2017. [22] M. Xu, S. Zhang, J. Ma, and O. A. Dobre, “Deep learning-based time-varying channel estimation for RIS assisted communication,” IEEE Communications Letters, vol. 26, no. 1, pp. 94–98, 2022. [23] S. Liu, Z. Gao, J. Zhang, M. Di Renzo, and M.-S. Alouini, “Deep denoising neural network assisted compressive channel estimation for mmWave intelligent reflecting surfaces,” IEEE Transactions on Vehicular Technology, vol. 69, no. 8, pp. 9223–9228, 2020. [24] T. Gong, S. Zhang, F. Gao, Z. Li, and M. Di Renzo, “Deep learningbased channel extrapolation for hybrid RIS-aided mmWave systems with low-resolution ADCs,” IEEE Transactions on Wireless Communications, vol. 23, no. 10, pp. 14 408–14 420, 2024. [25] S. Zhang, S. Zhang, F. Gao, J. Ma, and O. A. Dobre, “Deep learning optimized sparse antenna activation for reconfigurable intelligent surface assisted communication,” IEEE Transactions on Communications, vol. 69, no. 10, pp. 6691–6705, 2021. [26] Z. Mao, X. Liu, and M. Peng, “Channel estimation for intelligent reflecting surface assisted massive MIMO systems—a deep learning approach,” IEEE Communications Letters, vol. 26, no. 4, pp. 798–802, 2022. [27] A. Abdallah, A. Celik, M. M. Mansour, and A. M. Eltawil, “RISaided mmWave MIMO channel estimation using deep learning and compressive sensing,” IEEE Transactions on Wireless Communications, vol. 22, no. 5, pp. 3503–3521, 2023. [28] W. Shen, Z. Qin, and A. Nallanathan, “Deep learning for superresolution channel estimation in reconfigurable intelligent surface aided systems,” IEEE Transactions on Communications, vol. 71, no. 3, pp. 1491–1503, 2023. [29] S. Zheng, S. Wu, C. Jiang, W. Zhang, and X. Jing, “Hybrid driven learning for channel estimation in intelligent reflecting surface aided millimeter wave communication,” IEEE Transactions on Wireless Communications, vol. 23, no. 6, pp. 5801–5815, 2024. [30] M. Haider, I. Ahmed, A. Rubaai, C. Pu, and D. B. Rawat, “GANbased channel estimation for IRS-aided communication systems,” IEEE Transactions on Vehicular Technology, vol. 73, no. 4, pp. 6012–6017, 2024. [31] M. Ye, H. Zhang, and J.-B. Wang, “Channel estimation for intelligent reflecting surface aided wireless communications using conditional GAN,” IEEE Communications Letters, vol. 26, no. 10, pp. 2340–2344, 2022. [32] S. Noh, K. Seo, Y. Sung, D. J. Love, J. Lee, and H. Yu, “Joint direct and indirect channel estimation for RIS-assisted millimeter-wave systems based on array signal processing,” IEEE Transactions on Wireless Communications, vol. 22, no. 11, pp. 8378–8391, 2023. [33] C. M. Bishop and N. M. Nasrabadi, Pattern recognition and machine learning. Springer, 2006, vol. 4, no. 4. [34] ETSI, “TR 138 901 v17.0.0 (2022-04): 5G; study on channel model for frequencies from 0.5 to 100 GHz (3GPP TR 38.901 version 17.0.0 release 17),” ETSI, Technical Report, 2022.
14
[35] P. Dong, H. Zhang, G. Y. Li, I. S. Gaspar, and N. NaderiAlizadeh, “Deep CNN-based channel estimation for mmWave massive MIMO systems,” IEEE Journal of Selected Topics in Signal Processing, vol. 13, no. 5, pp. 989–1000, 2019. [36] J. Gao, M. Hu, C. Zhong, G. Y. Li, and Z. Zhang, “An attention-aided deep learning framework for massive MIMO channel estimation,” IEEE Transactions on Wireless Communications, vol. 21, no. 3, pp. 1823– 1835, 2022. [37] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems, vol. 30, 2017. [38] D. Rolnick, A. Veit, S. Belongie, and N. Shavit, “Deep learning is robust to massive label noise,” 2018. [Online]. Available: https://arxiv.org/abs/1705.10694 [39] J. Ho, N. Kalchbrenner, D. Weissenborn, and T. Salimans, “Axial attention in multidimensional transformers,” arXiv preprint arXiv:1912.12180, 2019. [40] R. Dujardin and C. Favre, “The dynamical manin–mumford problem for plane polynomial automorphisms,” Journal of the European Mathematical Society, vol. 19, no. 11, pp. 3421–3465, 2017. [41] P. Vincent, H. Larochelle, Y. Bengio, and P.-A. Manzagol, “Extracting and composing robust features with denoising autoencoders,” in Proceedings of the 25th International Conference on Machine Learning, ser. ICML ’08. New York, NY, USA: Association for Computing Machinery, July 2008, p. 1096–1103. [Online]. Available: https://doi.org/10.1145/1390156.1390294 [42] K. Zhang, W. Zuo, Y. Chen, D. Meng, and L. Zhang, “Beyond a gaussian denoiser: Residual learning of deep CNN for image denoising,” IEEE Transactions on Image Processing, vol. 26, no. 7, pp. 3142–3155, 2017. [43] K. He and J. Sun, “Convolutional neural networks at constrained time cost,” in 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA, June 2015, pp. 5353–5360. [44] D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” CoRR, vol. abs/1412.6980, 2014. [Online]. Available: https://api.semanticscholar.org/CorpusID:6628106