Conceptio › Archive › arXiv CS
arXiv CSopen access

FreqSpaNet: Frequency and Spatial Learning of SFPF for Physical Layer Hardware Integrity Detection

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

FREQSPANET: FREQUENCY AND SPATIAL LEARNING OF SFPF FOR PHYSICAL LAYER HARDWARE INTEGRITY DETECTION Xiaoxuan Huang, Jinlong Xu, YiZhe Wang, Meng Zhang, Xian Li, Yuying Bian

arXiv:2609.17491v1 [cs.LG] 15 Sep 2026

ABSTRACT Unauthorized hardware replacement can preserve a wireless device’s logical identity while altering its physical implementation, posing a challenge to hardware integrity verification. Spatio-frequency polarization fingerprints (SFPFs) capture device-dependent responses across multiple frequencies and directions, but their frequency and spatial dimensions exhibit different structural dependencies. We propose FreqSpaNet, an SFPF representation learning network for open set hardware anomaly detection. A frequency branch captures local variations among neighboring frequencies, while a geometryaware spatial branch models directional relationships using angular information. The two representations are combined through adaptive fusion, and complementary pretraining further captures shared information while preserving the distinct characteristics of the frequency and spatial representations. Experiments show that FreqSpaNet achieves a mean AUROC of 96.31%, 9.05 points above the baseline. Results under seven hardware replacement scenarios further verify the effectiveness of FreqSpaNet. Index Terms— physical layer security, hardware integrity detection, open set detection, frequency–spatial learning, polarization fingerprint 1. INTRODUCTION Unauthorized replacement of a wireless device’s hardware module may preserve its communication functions and logical identity while altering its physical implementation, creating a hardware integrity threat. Such hardware changes perturb the device-dependent radiation characteristics of the transmitted signal, which in turn modify its polarization response [1]. Polarization fingerprint (PF) characterizes this response through the complex relation between two orthogonally polarized received components[1, 2, 3]. SFPF extends PF by organizing these responses over multiple frequencies and directions, yielding a structured frequency–spatial representation. Neighboring frequencies tend to show locally correlated responses, whereas different directions capture complementary spatial information. Moreover, the effect of a hardware change is generally nonuniform across frequencies and directions. Learning an effective representation from SFPF never-

theless presents three challenges. First, its two dimensions exhibit different structural dependencies. Second, hardwareinduced changes are often nonuniform across frequencies and directions, so informative local responses may be weakened by early global aggregation. Third, the relationships among directions are determined by their physical angular separation rather than by image grid or sequence positions. Generic CNNs and vision Transformers apply largely homogeneous operators or flatten the frequency–spatial structure [4, 5], making them poorly matched to these SFPF-specific characteristics. Motivated by these SFPF characteristics, we propose FreqSpaNet, a SFPF representation learning framework that jointly exploits frequency and spatial information while retaining the complete angle–frequency structure. Its dualbranch encoder uses local convolution to capture variations across frequencies and angular-aware attention to model dependencies among directions. The two branch representations are then adaptively fused for each input sample. FreqSpaNet is further trained through frequency–spatial complementary pretraining, which infers masked SFPF responses while learning both the information shared by the two branches and the characteristics specific to each branch. The resulting classifier outputs are calibrated and fused to detect hardware changes. Experiments show that FreqSpaNet outperforms the evaluated open set systems built on CNN, Transformer, and MAE backbones [4, 5, 6], and remains robust under seven hardware-replacement scenarios. 2. PROPOSED FREQSPANET 2.1. Network Overview Figure 1 illustrates the architecture of FreqSpaNet. An input SFPF is denoted by X ∈ R2×A×F ×T , which organizes the complex polarization responses measured over A directions and F frequencies. The first dimension contains the real and imaginary components, and T denotes the number of sampling points. We use A = 301, F = 9, and T = 1024. The direction set is D = {(θa , ϕa )}A a=1 , where θa and ϕa are the elevation and azimuth of direction a. For each direction–frequency pair, a shared SFPF encoder maps xa,f ∈ R2×T to a token ha,f ∈ Rd . The resulting tensor H ∈ RA×F ×d is processed by two parallel branches. The

frequency branch captures local variations across neighboring frequencies, while the spatial branch models dependencies among directions. Their outputs, zf and zs , are adaptively fused into the SFPF representation z, which is mapped to the logits o ∈ RC of the C enrolled devices. FreqSpaNet is first pretrained through frequency–spatial complementary learning and then fine-tuned for enrolled-device classification. At inference, four scores derived from the classifier outputs are calibrated and fused to detect hardware changes. 2.2. Frequency–Spatial Encoding and Adaptive Fusion The shared SFPF encoder applies two 1D convolutions, GELU activations, global pooling, and layer normalization independently to each xa,f . This produces one token for every measured direction–frequency response before branchspecific modeling. The frequency branch adds a frequency embedding to each token and applies two residual depthwise 1D convolution blocks [7]. Their local kernels capture response variations across neighboring frequencies. Attention pooling first aggregates the frequency tokens within each direction and then combines all directions to produce zf . The spatial branch operates on the direction tokens at each frequency. Direction (θa , ϕa ) is represented by ua = [sin θa cos ϕa , sin θa sin ϕa , cos θa ]T , together with the sine and cosine values of its two angles. A MLP maps this coordinate descriptor to a direction embedding. Let Dij = arccos(uT i uj ) denote the angular separation between directions i and j. The attention weights of head m are   (m) (m)T Q K (m) √ + gb (D) , (1) P(m) = softmax dm where Q(m) and K(m) are the query and key matrices, dm is (m) the feature dimension of head m, and gb maps the angularseparation matrix D to attention biases. Thus, the spatial branch models interactions between directions using both their learned response features and their angular separation. Pooling over directions and frequencies gives zs . The relative contribution of the two branches can vary among samples. We therefore derive two variation statis2 tics from the response amplitude Ra,f,t = (XRe,a,f,t + 2 1/2 XIm,a,f,t ) . The frequency variation statistic sf is obtained by averaging |Ra,f +1,t − Ra,f,t | over all pairs of adjacent frequencies, directions, and sampling points. For the spatial statistic, we construct an angular neighborhood graph G = (V, E) by treating each measured direction as a vertex and connecting neighboring directions on the acquisition grid. Each edge (i, j) is weighted by wij = exp(−Dij /γ), where γ controls the decrease with angular separation. The statistic ss is the normalized weighted mean of |Ri,f,t − Rj,f,t | over graph edges, frequencies, and sampling points. A fusion network takes zf , zs , and the normalized pair (sf , ss ) as input and produces weights αf and αs satisfying

αf + αs = 1. The final representation is z = go ([αf zf + αs zs ; zf ⊙ zs ; |zf − zs |]) ,

(2)

where go is the fusion multilayer perceptron. The three components retain the branch contribution, cross-branch agreement, and complementary information, respectively. 2.3. Pretraining and Open Set Detection During complementary pretraining, a masked view XM and a noisy view XN of each SFPF are separately processed by the FreqSpaNet encoder with shared parameters. The FreqSpaNet encoder comprises the shared SFPF encoder, the frequency and spatial branches, and the adaptive fusion module, producing the fused representations zM and zN for the two views. A reconstruction decoder combines zM with the frequency and direction embeddings to recover the masked responses [6, 8]. The reconstruction error on the masked responses defines Lrec . Meanwhile, a shared projection head maps zM and zN , and the resulting representations are optimized using the supervised contrastive loss Lcon [9]. For the masked view, one shared projection head maps zf and zs to cf and cs , while two separate projection heads map them to pf and ps . The former pair represents information common to both branches, whereas the latter pair represents information unique to the frequency and spatial branches. We define Lcom = 1 − cos(cf , cs ) to align their common information and Lpri = | cos(pf , ps )| to reduce redundancy between their unique information. The pretraining objective is Lpre = λr Lrec + λt Lcon + λc Lcom + λp Lpri , (3) where λr , λt , λc , and λp weight the four loss terms. After pretraining, the decoder and projection heads are removed. The pretrained FreqSpaNet encoder is then jointly fine-tuned with a linear enrolled classifier, and the parameters of both are updated. During fine tuning, SFPFs collected from devices of different brands and models from the enrolled devices are used for outlier exposure [10, 11]. No hardware-replacement samples are used during training. The fine-tuning objective is Lft = Lce + λu Luni + λe Leng , where Lce is the classification loss, Luni encourages uniform predictions for auxiliary outliers, and Leng separates enrolled and auxiliary samples. For open set detection, let pc = softmax(o)c be the predicted probability of class c. Four anomaly PC scores are extracted from o: the energy E = −Te log c=1 exp(oc /Te ), maximum-probability PCuncertainty 1 − maxc pc , normalized entropy H = − c=1 pc log pc / log C, and probabilitymargin uncertainty 1 − (p(1) − p(2) ). Here, Te is the energy temperature and p(1) ≥ p(2) are the two largest class probabilities. The resulting score vector r(X) is standardized using known-device validation statistics, giving b r(X), and fused as a(X) = σ(wTb r(X) + b). The fusion parameters w and b are

FreqSpaNet Encoder and Open-Set Decision Adaptive Fusion Frequency & Spatial Branches

Input & Shared Encoder 𝑋 ∶ 2 × A × F × T(Re / Im)

Frequency branch

Re

…

frequencies

Im

…

𝑓1

Shared SFPF Encoder Conv1D GELU Global LN Pool …

𝛼𝑓 , 𝛼𝑠

Weighted sum

separation 2× biased Direction Frequency 𝜶𝒇 𝒛𝒇 + 𝜶𝒔 𝒛𝒔 MLP matrix D Transformer pooling pooling

…

𝑷(𝒎) = 𝒔𝒐𝒇𝒕𝒎𝒂𝒙

𝑸(𝒎) 𝑲(𝒎) ᵀ 𝒅𝒎

𝒛𝒔

+ 𝒈𝒃 ᵐ) 𝑫

Linear classifier …

Branch-weight MLP + softmax

𝑧𝑠

y

Open-Set Decision

𝑠𝑠

Product 𝒛𝒇 ⊙ 𝒛𝒔

Abs. diff. 𝒛𝒇 − 𝒛𝒔

⊙

−

Concat

…

Logits

FreqSpaNet Encoder

𝑧𝑀

𝑧𝑓 𝑧𝑠 𝑧𝑁

Reconstruction Decoder (+ frequency & direction embeddings) Contrastive projection head (shared) Shared projection head Frequency projection head Spatial projection head

𝑐𝑓 , 𝑐𝑠 𝑝𝑓 𝑝𝑠

𝑳𝒄𝒐𝒏 𝑳𝒄𝒐𝒎 𝑳𝒑𝒓𝒊

…

C

Anomaly score 𝑎 𝑋

𝒛 = 𝒈ₒ 𝜶𝒇 𝒛𝒇 + 𝜶𝒔 𝒛𝒔 ; 𝒛𝒇 ⊙ 𝒛𝒔 ; 𝒛𝒇 − 𝒛𝒔

𝑳𝒓𝒆𝒄

2

Energy Confidence Entropy Margin 𝐻 𝐸 deficit deficit Known-val standardization

Fusion MLP

෡𝑴 𝑿

1

Class probability p

1

0

Anomaly

Training Procedure FreqSpaNet Encoder Shared weights

+𝜖

𝑧𝑓

𝑓n

z

𝐻∶ A × F × 𝑑 (angle × frequency tokens)

Noisy view 𝑋𝑁

Frequency Direction pooling pooling 𝒛𝒇

Spatial branch

x

Masked view 𝑋𝑀

2× residual DW-Conv1D

𝑠𝑓

Pretrained objective

𝑳𝒑𝒓𝒆 = 𝝀𝒓 𝑳𝒓𝒆𝒄 + 𝝀𝒕 𝑳𝒄𝒐𝒏 + 𝝀𝒄 𝑳𝒄𝒐𝒎 + 𝝀𝒑 𝑳𝒑𝒓𝒊

Yes 𝒂 𝑿 No > 𝝉?

Enrolled identity

Enrolled samples auxiliary outlier samples …

…

Linear classifier

…

Loss: 𝑳𝒇𝒕 = 𝑳𝒄𝒆 + 𝝀𝒖 𝑳𝒖𝒏𝒊 + 𝝀𝒆 𝑳𝒆𝒏𝒈

Fig. 1. Architecture and learning pipeline of FreqSpaNet. learned from known validation samples and auxiliary outliers without access to the final hardware anomalies. The threshold τ is defined by a 5% false-alarm rate on the known-device validation set. A sample with a(X) > τ is reported as a hardware anomaly; otherwise, it is assigned to the enrolled class with the largest probability. 3. EXPERIMENTS 3.1. Experimental setup The experiment uses ten wireless devices enrolled in their original hardware configurations. A USRP X310 receiver and an orthogonal dual-polarized antenna are placed 3 m from the device under test, while a motorized pan–tilt platform rotates the device and the receiver remains fixed. Measurements span 913–917 MHz in 0.5 MHz steps, and 1024 complex sampling points are collected for each angle–frequency pair. We use θ ∈ [0◦ , 60◦ ] and ϕ ∈ [0◦ , 120◦ ], both with 5◦ steps, yielding 301 physical directions and nine frequency points per SFPF. Five enrolled devices are selected as attack targets, each paired with non-reused replacement hardware and evaluated under all combinations of antenna (A), digital/baseband (D), and RF-front-end (R) replacement, including A, D, R, A+D, A+R, D+R, and A+D+R. For each enrolled device, 600 normal SFPFs at 20 dB are used for training and 100 additional samples are reserved for validation. Auxiliary SFPFs comprise 241 samples from each of three devices whose brands and models are disjoint from those of the enrolled devices and final-test hardware, and are used only for outlier exposure and learning the logistic score fusion weights. Hardware replacement samples are reserved exclusively for final

evaluation. At each SNR from 0 to 20 dB in 1 dB increments, the test set contains 241 normal samples per enrolled device and 241 anomalous samples per target device under each hardware replacement scenario. 3.2. Comparison with generic open set systems We compare FreqSpaNet with open set systems built on ResNet-18 [4], ViT-B/16 and ViT-Tiny/16 [5], and MAEB/16 [6]. The evaluated combinations include maximum softmax probability (MSP) [12], energy scoring [11], OpenMax [13], outlier exposure (OE) [10], and Deep-SVDD [14]. Figure 2 compares FreqSpaNet with representative generic open set systems. FreqSpaNet achieves a mean AUROC of 96.31%, exceeding the best evaluated generic baseline on SFPF, ResNet-18+MSP 87.26%, by 9.05 percentage points. It also maintains the highest AUROC at every tested SNR, indicating that separately modeling the dependencies along the frequency and spatial dimensions provides more reliable anomaly ranking than generic backbones. FreqSpaNet achieves a mean unknown class F1 score of 89.93% over the full 0–20 dB range and consistently outperforms the compared methods from 13 to 20 dB. These results support the use of SFPF-specific frequency and spatial modeling rather than treating SFPF as a generic image or token sequence. 3.3. Robustness to hardware-replacement scenarios As shown in Fig. 3, antenna replacement is the most challenging case: over 0–5 dB, it yields a mean AUROC of 85.30% and a mean FPR95 of 24.87%, whereas the other scenarios achieve mean FPR95 values between 12.47% and 16.95%.

100 90 80 70 60 50 40 30 20 10 0

0

2

4

MAE-B/16 + MSP ResNet-18 + Energy ViT-Tiny/16 + Energy

6

8

10

12

SNR (dB)

MAE-B/16 + OE ResNet-18 + MSP ViT-Tiny/16 + MSP

14

16

18

MAE-B/16 + Deep-SVDD MAE-B/16 + OpenMax ViT-B/16 + MSP FreqSpaNet

Unknown F1-score (%)

AUROC (%)

MAE-B/16 + Deep-SVDD MAE-B/16 + OpenMax ViT-B/16 + MSP FreqSpaNet

20

100 90 80 70 60 50 40 30 20 10 0

0

2

4

MAE-B/16 + MSP ResNet-18 + Energy ViT-Tiny/16 + Energy

6

8

(a)

10

12

SNR (dB)

MAE-B/16 + OE ResNet-18 + MSP ViT-Tiny/16 + MSP

14

16

18

20

(b)

Fig. 2. Comparison with representative generic open set systems over 0–20 dB. (a) AUROC; (b) unknown-class F1-score. R A+D+R

A+D

A A+R

0

2

4

6

8

10

12

SNR (dB)

14

16

18

20

100 90 80 70 60 50 40 30 20 10 0

D D+R

R A+D+R

A+D

A A+R 100 90 80 70 60 50 40 30 20 10 0

D D+R

R A+D+R

A+D

FPR95 (%) #

100 90 80 70 60 50 40 30 20 10 0

D D+R

Unknown F1-score (%)

AUROC (%)

A A+R

0

2

4

6

(a)

8

10

12

SNR (dB)

(b)

14

16

18

20

0

2

4

6

8

10

12

SNR (dB)

14

16

18

20

(c)

Fig. 3. FreqSpaNet under seven hardware-replacement scenarios. (a) AUROC; (b) unknown-class F1-score; (c) FPR95.

Table 1. Ablation results averaged over 0–20 dB (%). Variant AUROC ↑ FPR95 ↓ Bal. acc. ↑ Frequency only 93.44 14.42 81.12 Spatial only 94.60 14.69 79.57 w/o direction encoding 92.36 31.89 84.15 w/o angular bias 95.39 21.53 86.79 Fixed fusion 93.89 26.13 85.70 w/o reconstruction 94.11 25.38 85.96 w/o contrastive 87.59 45.61 80.09 w/o common–private 92.10 38.52 85.28 FreqSpaNet (full) 96.31 12.30 88.54

Nevertheless, from 15 to 20 dB, the averages over all seven scenarios reach 99.33% AUROC, 95.45% unknown-class F1score, and 2.32% FPR95. These results demonstrate that FreqSpaNet can reliably detect both single- and multi-module hardware replacements. 3.4. Ablation study Table 1 shows that both branches are necessary: using either branch alone reduces balanced accuracy by more than 7 percentage points. Removing direction encoding, angular separation bias, or adaptive fusion also degrades the overall performance, confirming the effectiveness of geometry-aware spatial modeling and adaptive branch fusion. The pretraining

ablations further show that reconstruction, contrastive learning, and common–private decomposition all contribute to the final performance. 4. CONCLUSION We presented FreqSpaNet, an SFPF-specific representation learning network for open set hardware anomaly detection. By separately modeling local frequency dependencies and the angular relationships among directions, FreqSpaNet preserves the distinct characteristics of the frequency and spatial dimensions and integrates them through sampleadaptive fusion. Complementary pretraining further improves the learned representation. Experiments over 0–20 dB show that FreqSpaNet achieves 96.31% mean AUROC and 89.93% mean unknown-class F1-score. Results under seven hardware-replacement scenarios and ablation studies further demonstrate the effectiveness of the proposed frequency– spatial representation learning framework. 5. REFERENCES [1] Jinlong Xu, Dong Wei, and Weiqing Huang, “Polarization fingerprint: A novel physical-layer authentication in wireless IoT,” in Proc. IEEE Int. Symp. World

of Wireless, Mobile and Multimedia Networks (WoWMoM), 2022, pp. 434–443. [2] Jinlong Xu and Dong Wei, “Polarization fingerprintbased LoRaWAN physical layer authentication,” IEEE Transactions on Information Forensics and Security, vol. 18, pp. 4593–4608, 2023. [3] Jinlong Xu, Dong Wei, and Weiqing Huang, “Specific emitter identification via spatial characteristic of polarization fingerprint,” in Proc. IEEE Symp. Computers and Communications (ISCC), 2022, pp. 1–7. [4] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, “Deep residual learning for image recognition,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), 2016, pp. 770–778. [5] Alexey Dosovitskiy et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” in Int. Conf. Learning Representations (ICLR), 2021. [6] Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick, “Masked autoencoders are scalable vision learners,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition (CVPR), 2022, pp. 16000–16009. [7] François Chollet, “Xception: Deep learning with depthwise separable convolutions,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), 2017, pp. 1251–1258. [8] Zhicheng Huang, Xiaojie Jin, Chengze Lu, Qibin Hou, Ming-Ming Cheng, Dongmei Fu, Xiaohui Shen, and Jiashi Feng, “Contrastive masked autoencoders are stronger vision learners,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 4, pp. 2506–2517, 2024. [9] Prannay Khosla et al., “Supervised contrastive learning,” in Advances in Neural Information Processing Systems, 2020, vol. 33, pp. 18661–18673. [10] Dan Hendrycks, Mantas Mazeika, and Thomas Dietterich, “Deep anomaly detection with outlier exposure,” in Int. Conf. Learning Representations (ICLR), 2019. [11] Weitang Liu, Xiaoyun Wang, John D. Owens, and Yixuan Li, “Energy-based out-of-distribution detection,” in Advances in Neural Information Processing Systems, 2020, vol. 33, pp. 21464–21475. [12] Dan Hendrycks and Kevin Gimpel, “A baseline for detecting misclassified and out-of-distribution examples in neural networks,” in Int. Conf. Learning Representations (ICLR), 2017.

[13] Abhijit Bendale and Terrance E. Boult, “Towards open set deep networks,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), 2016, pp. 1563– 1572. [14] Lukas Ruff, Robert A. Vandermeulen, Nico Görnitz, Lucas Deecke, Shoaib A. Siddiqui, Alexander Binder, Emmanuel Müller, and Marius Kloft, “Deep one-class classification,” in Proc. 35th Int. Conf. Machine Learning (ICML), 2018, vol. 80 of Proceedings of Machine Learning Research, pp. 4393–4402.

Record · ID 919376 · SHA-256 1809cf11371e87d6
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.