ConceptioArchivearXiv CS
arXiv CSopen access

RF-LEGO: Modularized Signal Processing-Deep Learning Co-Design for RF Sensing via Deep Unrolling

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
clouddistributedcomputingparallelcomputing
distributed computing, parallel computing, cloud

RF-LEGO: Modularized Signal Processing-Deep Learning Co-Design for RF Sensing via Deep Unrolling Luca Jiang-Tao Yu

Chenshu Wu

The University of Hong Kong [email protected]

The University of Hong Kong [email protected]

arXiv:2604.10183v1 [cs.DC] 11 Apr 2026

ABSTRACT Wireless sensing, traditionally relying on signal processing (SP) techniques, has recently shifted toward data-driven deep learning (DL) to achieve performance breakthroughs. However, existing deep wireless sensing models are typically end-to-end and task-specific, lacking reusability and interpretability. We propose RF-LEGO, a modular co-design framework that transforms interpretable SP algorithms into trainable, physics-grounded DL modules through deep unrolling. By replacing hand-tuned parameters with learnable ones while preserving core processing structures and mathematical operators, RF-LEGO ensures modularity, cascadability, and structure-aligned interpretability. Specifically, we introduce three deep-unrolled modules for critical RF sensing tasks: frequency transform, spatial angle estimation, and signal detection. Extensive experiments using realworld data for Wi-Fi, millimeter-wave, UWB, and 6G sensing demonstrate that RF-LEGO significantly outperforms existing SP and DL baselines, both standalone and when integrated into multiple downstream tasks. RF-LEGO pioneers a novel SP-DL co-design paradigm for wireless sensing via deep unrolling, shedding light on efficient and interpretable deep wireless sensing solutions. Our code is available at https://github.com/aiot-lab/RF-LEGO.

CCS CONCEPTS • Human-centered computing → Ubiquitous and mobile computing; • Networks → Mobile networks.

KEYWORDS Deep Unrolling, Wireless Sensing, Signal Processing, Deep Learning, RF Signals ACM Reference Format: Luca Jiang-Tao Yu, Chenshu Wu. 2026. RF-LEGO: Modularized Signal Processing-Deep Learning Co-Design for RF Sensing via

This work is licensed under a Creative Commons Attribution 4.0 International License. MobiCom ’26, October 26–30, 2026, Austin, TX, USA © 2026 Copyright held by the owner/author(s). ACM ISBN 979-8-4007-2505-0/26/10. https://doi.org/10.1145/3795866.3796683

Deep Unrolling. In The 32nd Annual International Conference on Mobile Computing and Networking (MobiCom ’26), October 26–30, 2026, Austin, TX, USA. ACM, New York, NY, USA, 16 pages. https: //doi.org/10.1145/3795866.3796683

1

INTRODUCTION

Over the past decade, wireless sensing has transformed the wireless industry across Wi-Fi, millimeter-wave/UWB radars, 6G networks, LoRa, and even GPS. As AI continues to revolutionize various applications, the field of wireless sensing is undergoing a significant shift from traditional model-centric signal processing [37, 67, 72] toward data-heavy deep learning [49, 66, 75, 77]. Early efforts in Deep Wireless Sensing (DWS) [44, 76] exploit purely data-driven approaches by leveraging existing neural network architectures like CNNs, RNNs, and Transformers, with recent attempts to explore large language models (LLMs) to process wireless signals directly [27, 69, 71]. Albeit inspiring, these models face significant challenges related to data scarcity and model interpretability, which in turn hinder generalization, reuse, and deployment efficiency. To overcome these issues, researchers have resorted to signal-informed models through signal processing-deep learning (SP-DL) co-design, and remarkable advances have been achieved: (i) direct embedding of signal processing blocks to inject physical priors into neural networks [14]; (ii) RF-intrinsic basis-inspired design, e.g., phase-aware encoder [11, 64]; (iii) DSP-inspired parameterizations and mechanisms, like Fourier-based initialization [74] and STFT-like operations [29, 65]; and (iv) RF-specific architectures with math-interpretable design [15, 70]. Despite the progress, existing designs still suffer from major limitations: 1) Limited Reusability: Most models are trained end-to-end for a specific downstream task, such as gesture recognition or vital sign monitoring [10, 14, 29, 34, 74, 78]. This undermines their reusability, making it difficult to reuse or swap individual modules when tasks or environments vary. Unlike in computer vision or natural language processing, few DWS models are reused in subsequent research. 2) Lack of Interpretability: Although some prior works incorporate signal processing inspirations in network design [34, 64], they lack structure-aligned interpretability; specifically, they do not maintain a clear stage-by-stage signal processing diagram with well-defined input/output contracts

MobiCom ’26, October 26–30, 2026, Austin, TX, USA

S1

S2

Neural Network

S3

Signal Processing

Deep Unrolling S1

RF-LEGO

Deep Learning

RF-LEGO Module

S2

Interpretability l

Modularity S1

RF-LEGO Module

Yu and Wu

e

g

o

Cascadability S3

RF-LEGO Module

RF-LEGO Module

Figure 1: Core principles of RF-LEGO: Modularity, Cascadability, and Interpretability. RF-LEGO bridges the gap between classical SP and DL for RF sensing via deep unrolling.

and semantically meaningful intermediate outputs in standard signal domains. Recently, deep unrolling techniques have emerged as a promising approach to SP-DL co-design by unrolling classical algorithms in a learnable framework. Deep unrolling replaces fixed algorithmic coefficients with learnable yet physics-constrained parameters while preserving operator structure and physical input semantics [40], reserving the operator interpretability of signal processing while leveraging the parameter learnability of deep learning. These compelling advantages inspire us to embrace deep unrolling for modularized SP-DL co-design, promoting both reusability and interpretability for DWS models. While deep unrolling has been successful in numerous classical algorithms [30, 33, 34, 39, 46, 50], modularized unrolling in DWS faces unique challenges, as summarized below: Different signal processing operators impose distinct proximal structures and constraints; non-differentiable steps require smooth, bounded surrogates to prevent back-propagation collapse; and true modularity demands preserving the classical input-output and complex-valued semantics so blocks remain plug-and-play and calibratable across platforms. We propose RF-LEGO, as shown in Fig. 1, the first modularized SP-DL co-design approach for RF sensing based on deep unrolling. By turning each algorithmic iteration into a differentiable, trainable layer, RF-LEGO maintains mathematically grounded SP interfaces while replacing handcrafted configurations with learnable parameters, delivering reusable SP-DL modules with structure-aligned interpretability. The design of RF-LEGO centers around three key principles: • Modularity: Our modules unroll classical algorithms, mirroring their input/output semantics while introducing a minimal number of trainable parameters. By doing so, each module will be plug-and-play across different pipelines, datasets, and applications without requiring retraining. • Cascadability: The modular nature of these blocks allows them to be flexibly combined with each other or integrated into existing methods, offering reusable "LEGO bricks" for effortlessly building sophisticated learning pipelines.

• Interpretability: By converting algorithmic iterations into differentiable and trainable layers, the proposed blocks preserve SP-aligned operator structures and semantically meaningful intermediate outputs. While this design does not ensure strict model interpretability of the neural networks and the learned parameters, it offers intra-block structure and inter-block pipeline interpretability as they are precisely aligned with classical SP pipelines. Among various signal processing techniques involved in RF sensing, we primarily focus on the three most fundamental components, i.e., frequency transformation, spatial angle estimation, and signal detection, which have been widely used in most of the existing literature [16, 28, 44, 49, 59, 63– 65, 67, 73–76, 78, 79]. Correspondingly, we present the following three deep unrolled algorithms in RF-LEGO, which can already enable many RF sensing applications. We leave it to future work in the community to contribute more RFLEGO blocks for other useful techniques. ■ RF-LEGO FT: Frequency Transformation (FT) is a fundamental processing step in almost all RF sensing applications. However, the classical Cooley-Tukey FFT [13] has fixed basis functions, rendering it inherently rigid and unable to adapt to signal-dependent artifacts like spectral leakage, which can obscure weak targets in noisy environments. To overcome this, we choose to deep unroll Bluestein’s Algorithm implementation [3], which factorizes the transform into chirp multiplications and a single convolution. This exposes a clean insertion point for a compact, complex-valued learnable operator, allowing for data-driven suppression of spectral leakage while preserving the intrinsic semantics of RF signals. ■ RF-LEGO Beamformer: RF sensing often performs a spatial transformation to obtain the angle spectrum, a crucial step to separate multiple targets. Super-resolution methods like MUSIC [48], relying on subspace decomposition, are notoriously brittle in practice, failing under challenging conditions like coherent multipath and low SNR conditions. Furthermore, it faces unstable gradients during backpropagation, often causing training to fail if unrolled. In RF-LEGO, we instead investigate the unrolled LASSO beamformer [38] using ADMM (Alternating Direction Method of Multipliers) [6]. This iterative framework is inherently better suited for real-world conditions. While the ADMM solver using rigid parameters and fixed update rules exhibits limited performance and slow convergence in complex scenarios [62], our unrolled design addresses this by making these coefficients trainable while leaving the essential steering model unchanged. ■ RF-LEGO Detector: Many classical approaches, like the commonly used Constant False Alarm Rate (CFAR) family [23, 47, 57, 60], have been developed for signal detection, an essential step for sensing applications. The core operations of CFAR detectors, however, are non-differentiable and hence

RF-LEGO

cannot be directly unrolled. For example, operators like order statistics, used for robust noise estimation, are incompatible with gradient-based backpropagation. We therefore propose an unrolled design with the same workflow as classical CFAR, yet recast its adaptive noise estimation as a compact and differentiable state space model [22], which acts as a smooth surrogate and allows the detector to learn the adaptive dynamics. The design not just circumvents the tedious manual parameter configurations in classical CFAR, but also improves robustness across diverse environments. There are existing works to improve these classical algorithms with deep learning, which, however, are mostly end-to-end models lacking interpretability [34, 39, 74]. In contrast, we propose novel techniques to unroll classical algorithms as learnable blocks without changing the original signal processing structure. Note that our designs are primarily within the context of wireless sensing, rather than for generic applications like the original algorithms. We conduct extensive real-world experiments to evaluate RF-LEGO, with a main focus on validating the effectiveness of individual unrolled techniques, the performance of applying RF-LEGO techniques in tandem with other existing models, and the performance of downstream tasks such as trajectory tracking, vital sign monitoring, and human activity recognition. We evaluate on both self-collected data and public datasets, totaling approximately 3M frames across various environments and different modalities (including Wi-Fi, millimeter-wave, UWB, and 6G). Our experimental results demonstrate that RF-LEGO consistently outperforms classical SP baselines and achieves competitive performance compared with prior DL baselines, while offering reusable learnable components for broader applications. RF-LEGO pioneers a new paradigm for building reusable and interpretable SP-DL co-design for wireless sensing via deep unrolling, inspiring new research opportunities in the community to build generalizable and efficient DWS solutions. In summary, our main contributions are as follows: • We present RF-LEGO, the first modular co-design framework that bridges rigid-but-interpretable signal processing and adaptive-but-opaque deep learning via deep unrolling, yielding reusable and interpretable learning modules for the wireless sensing community. • We propose three deep unrolled, physics-grounded modules, i.e., RF-LEGO FT for robust frequency transformation, RF-LEGO Beamformer for angle estimation, and RF-LEGO Detector for adaptive signal detection. • We conduct extensive evaluation using real-world mmWave, UWB, and Wi-Fi signals, demonstrating the effectiveness of RF-LEGO techniques either when using individually or integrating into existing networks.

MobiCom ’26, October 26–30, 2026, Austin, TX, USA

a Operator Unrolling S1

S2

Trainable Operator

O2

b Iterative-Opt Unrolling

S3

S1 Ia

Ib

S2 Ic

S3

Trainable Iterative Block

Figure 2: Deep Unrolling. S: Signal Processing Block; O: Unrolled Trainable Operator; I: Unrolled Trainable Iterative Block.

2

A PRIMER ON DEEP UNROLLING

Most deep learning models are purely data-driven, and their learned structures are difficult to interpret. End-to-end networks learn task-specific mappings (e.g., regression or classification) entirely through backpropagation over the highdimensional parameters whose individual roles are opaque [40]. By contrast, classical signal processing is interpretable at both the pipeline and block levels because it is derived from physical models and domain priors, i.e., the intrinsics of signals. However, traditional methods suffer from parameter rigidity, meaning that their performance hinges on experttuned hyperparameters that often need to be recalibrated upon environment changes. Deep unrolling transforms signal processing techniques into structured neural architectures, combining the adaptability of deep learning with the interpretability and physics of classical methods. Broadly, as illustrated in Fig. 2, it falls into two categories as follows: • Operator Unrolling. Keep the signal processing pipeline, but replace a classical block with a drop-in trainable operator that mirrors its mathematics. Fixed coefficients (e.g., kernels, filters, thresholds) become a compact set of physics-constrained learnable parameters, while complexvalued semantics and intermediate states remain exposed. The result is a plug-and-play block with strong domain alignment, interpretability, and easy composition with neighboring stages. • Iterative-Optimization Unrolling. Keep the classical block but open its inner loop: each solver’s iterative block becomes a network layer with learnable step sizes, preconditioners, and proximal or shrinkage maps. This turns an algorithm into a structured, trainable stack that preserves the structure-alignedinterpretability. Together, these two categories provide a concise design vocabulary: either swap in a trainable operator that respects the classical interface, or unfold the operator’s iterative routine into a stable, interpretable, and data-driven network. Subsequent sections instantiate these patterns for specific RF sensing blocks. Deep unrolling has achieved remarkable advances in various domains, with successful applications in unrolling classical algorithms for sparse coding [62], Kalman filtering [46, 50], and image restoration [30]. However, its

MobiCom ’26, October 26–30, 2026, Austin, TX, USA

Yu and Wu

potential has not yet been fully explored for the unique challenges of DWS. RF-LEGO pioneers this direction by proposing a framework for deep unrolled RF sensing.

(a) Bluestein’s Algorithm FT

𝒃∗ 𝒓

𝒃∗

Time-frequency transformation is a cornerstone of wireless sensing for extracting range and Doppler information. Classical methods such as FFT rely on fixed bases, making the spectrum vulnerable to significant sidelobe leakage. In challenging scenarios with multipath fading, this artifact can obscure weaker signals or even submerge the main lobe. Their coefficients are scene independent, so robustness and adaptivity degrade across hardware and environments in RF sensing. The widely-used Cooley-Tukey implementation is highly optimized for the common 2𝑁 case, but its multi-stage factorization offers no localized, structure-preserving learning handle. Alternatively, Bluestein’s Algorithm recasts the transformation as a convolution with a fixed chirp, which is efficient but not adaptive to signal-dependent artifacts. Preprocessing with filters and heuristic parameter adjustments provides limited gains, while advanced decompositions such as the Hilbert-Huang transform [25] increase computational cost and are susceptible to mode mixing. Meanwhile, many end-to-end models simply adopt vanilla architectures that ignore both the temporal-frequency domain gap and the complex-valued nature of RF signals [74]. Below, we outline our proposed design to address these limitations. Signal Model. Let 𝒔 ∈ C𝑁 denote a clean complex-baseband 𝑁 −1 , and let 𝒏 ∈ C𝑁 be additive signal with samples {𝑠𝑛 }𝑛=0 noise. The received signal is 𝒓 = 𝒔 + 𝒏. Let 𝔉(·) be the discrete Fourier Transformation. The frequency-domain sam𝑁 −1 satisfy ples {R [𝑘]}𝑘=0 𝑁 −1 ∑︁

𝑛

𝒔𝑛 𝑒 − 𝑗2𝜋𝑘 𝑁 + 𝔉(𝒏) [𝑘],

(1)

𝑛=0

where 𝑘 = 0, . . . , 𝑁 − 1, and the noise term perturbs both amplitude and phase. We do not impose a specific distribution on 𝒏 and treat 𝔉(𝒏) through its statistics. To expose a structure that preserves the classical transformation while enabling learning, we adopt Bluestein’s Algorithm implementation, which is rewritten as a convolution with a known 2 2 2 chirp. Using the identity: 𝑛 · 𝑘 = − (𝑘 −𝑛) + 𝑛2 + 𝑘2 , we obtain 2 𝑘2

𝑛2

𝑛2

R [𝑘] = 𝑒 − 𝑗𝜋 𝑁 · (𝒔𝑛 𝑒 − 𝑗𝜋 𝑁 ⊛ 𝑒 𝑗𝜋 𝑁 ) [𝑘] + 𝔉(𝒏) [𝑘], 2

Elementwise Product

𝒃

z*

*

Convolution

𝒃∗

. 𝓡

Target

(b) RF-LEGO FT

3 RF-LEGO DESIGN 3.1 RF-LEGO FT

R [𝑘] = 𝔉(𝒔 + 𝒏) [𝑘] =

.

z* Conjugate

(c) Results Bluestein’s Algorithm FT vs. RF-LEGO FT

2

(2)

with 𝒂𝑛 = 𝒔𝑛 𝑒 − 𝑗𝜋𝑛 /𝑁 , 𝒃 𝑛 = 𝑒 𝑗𝜋𝑛 /𝑁 , and ⊛ denoting discrete convolution. This factorization is algebraically equivalent to the discrete Fourier Transformation, retains the complexvalued semantics, and isolates a single fixed-kernel convolution as illustrated in Fig. 3(a).

𝒓

.

z*

𝒃

z*

Complex CNN

𝒃∗

. 𝓡$ Nulling

Submerge

High-SNR

Reduced Leakage

Low-SNR

Clear Target

Clutter

Reduced Interference

# Frequency Bin

Figure 3: RF-LEGO FT. (a) Bluestein’s Algorithm implementation of Fourier transformation via convolution, where 𝒃 represents the chirp signal and 𝒛 ∗ denotes its conjugate. (b) RF-LEGO FT replaces the fixed convolution with a learnable convolutional layer.

Deep Unrolling. The convolutional form in Eqn. (2) reveals a practical handle: the discrete Fourier Transformation can be executed as a single convolution with a known chirp. In real scenarios, however, clutter and other non-stationary, signal-dependent disturbances limit what a fixed chirp can achieve. We therefore keep the outer chirp multiplications intact and make only the convolution learnable, yielding a frequency-transformation block that remains structurally identical to the classical operator and preserves complexvalued architecture. Concretely, deep learning toolchains implement cross-correlation rather than flipped-kernel convolution as follows: (Í 𝑁 −1 𝒂 [𝑘 − 𝑛] · 𝒃 [𝑛], (SP) (𝒂 ⊛ 𝒃) [𝑘] ≜ Í𝑛=0 . (3) 𝑁 −1 𝑛=0 𝒂 [𝑘 + 𝑛] · 𝒃 [𝑛], (DL) We exploit the mathematical equivalence: a correlation with a filter is identical to a convolution with the flipped filter if trainable [2]. Thus, we replace the signal processing convolution by a hierarchical learnable complex filter applied through correlation, with weights initialized from the chirp sequence 𝒃. This preserves the algebraic blueprint of Bluestein’s Algorithm while exposing a compact set of trainable coefficients. We instantiate the learnable operator as a shallow complex-valued convolutional neural network that sits exactly at the convolution site of Bluestein’s Algorithm FT. The network uses smooth activations to provide mild nonlinearity and is initialized from signal processing bases to ensure a stable start and domain consistency. An optional, lightweight nulling head applies a soft, trainable shrinkage at the output to attenuate residual leakage. Our ablation study (§5.4) shows that the main gains stem from the unrolled convolution itself, with the nulling head acting as a stabilizer rather than the primary source of improvement. This design keeps interpretability, retains complex-valued semantics, and introduces data-driven adaptation where classical methods are rigid. In practice, it improves spectral suppression and reduces leakage while remaining plug-and-play with other modules. As visualized in Fig. 3 on the mmWave range frequency axis, RF-LEGO FT effectively suppresses sidelobes in

RF-LEGO

low-SNR conditions and isolates targets submerged in clutter where the classical algorithm fails.

MobiCom ’26, October 26–30, 2026, Austin, TX, USA (a) LASSO Beamformer

𝒓

(c) Results Iterative Block

𝒗(𝒕) Init

𝒛(𝒕)

Primal Update 𝑨𝐇 𝑨, 𝝆

(𝒕#𝟏)

𝒙

(𝒕#𝟏)

Shrinkage

𝒛

LASSO vs. RF-LEGO Beamformer Target@-30°

Variable Update

High-SNR

Target@0°

Low-SNR

𝒗(𝒕#𝟏)

(b) RF-LEGO Beamformer

𝒗(𝒕) 𝒓

Iterative Block Gate

Init

𝒈

𝒛(𝒕)

3.2

RF-LEGO Beamformer

Beamforming converts array measurements into an angular spectrum for estimating the angle of arrival. Many signal processing techniques have been proposed, including the classical delay and sum (DAS) and Bartlett [56], Capon [7] beamformers, and subspace methods such as Multiple Signal Classification (MUSIC) [48]. Built upon certain signal models, these methods suffer from different limitations. For example, the most widely used MUSIC method requires the number of sources and assumes a clear separation between the signal and noise subspaces. Moreover, in practice, these methods rely on strong modeling assumptions and manual tuning, which limits robustness across hardware and environments. Recent attempts exploit deep neural networks for better performance. Yet existing works are mostly end-to-end models for specific tasks. They often do not provide an explicit angle spectrum [74]. Furthermore, some models incorporate operations that are ill-suited for gradient-based training; for instance, methods based on the MUSIC algorithm require eigenvalue decomposition, which is known to have unstable gradients that can destabilize the training process [39, 45]. These drawbacks motivate a tightly coupled unrolled design that keeps the classical steering model and explicit angle spectrum, avoids unstable subspace factorization, and learns only a compact set of physics-constrained solver coefficients that are otherwise hand-tuned. We recast beamforming as sparse spectral estimation on a discretized set of angles and adopt the Least Absolute Shrinkage and Selection Operator (LASSO) formulation [38] to enforce a sparse angle spectrum consistent with the steering model. Based on this formulation, we then employ iterative optimization unrolling to unroll the LASSO beamformer as an improved Alternating Direction Method of Multipliers (ADMM) solver [6]. The unrolled structure keeps the classical operator and complex-valued inputs intact while converting a compact set of physics-constrained coefficients, such as step sizes, preconditioning weights, and shrinkage levels, into learnable parameters. This design outperforms conventional methods by removing reliance on subspace factorization and yielding robust performance under low SNR and coherent sources, while retaining an interpretable structure and a reusable interface like the classical methods. Our detailed signal model and unrolled design are as follows. Signal Model. Consider a uniform linear array (ULA) with 𝑀 elements spaced at 𝜆/2. The array observes 𝐾 far-field sources from directions {𝜽 0, 𝜽 1, . . . , 𝜽 𝐾 −1 } in spatially white

(𝒕) Primal Update 𝑾(𝒕) , 𝜼(𝒕)

(𝒕#𝟏)

𝒙

Gated Shrinkage

𝒛(𝒕#𝟏) Variable Update

Strong Clutters

Target@-10°

Main lobe Declined

Target@10°

𝒗(𝒕#𝟏)

Figure 4: RF-LEGO Beamformer. (a) The classical LASSO solver using ADMM. (b) The unrolled architecture with learnable parameters. Here, 𝒙 (𝒕 ) represents the estimated sparse angular spectrum, 𝒛 (𝒕 ) is the auxiliary variable, 𝒗 (𝒕 ) is the dual variable, and 𝒈 (𝒕 ) denotes the learned gate that dynamically balances the update between historical and current estimates.

complex Gaussian noise 𝒏 ∼ CN (0, 𝜎 2 𝑰 ). The received snapshot 𝒓 ∈ C𝑀 is modeled as 𝐾 −1 ∑︁ 𝒓= 𝒔 [𝑘]𝒂(𝜽 𝑘 ) + 𝒏, (4) 𝑘=0

where 𝒔 ∈ C𝐾 collects the complex source responses and 2𝜋 𝒅

𝑀 −1 . 𝒂(𝜽 𝑘 ) ∈ C𝑀 is the steering vector 𝒂(𝜽 𝑘 ) T = [𝑒 − 𝑗 𝜆 𝑖 sin 𝜽 𝑘 ] 𝑖=0 𝑁 Our goal is to estimate a sparse angle spectrum 𝒙ˆ ∈ C on a predefined grid {𝜽 0, . . . , 𝜽 𝑁 −1 } by solving   1 ˆ𝒙 = arg min ∥𝒓 − 𝑨𝒙 ∥ 22 + 𝜏 ∥𝒙 ∥ 1 , (5) 2 𝒙

where 𝑨 = [𝒂(𝜽 0 ), . . . , 𝒂(𝜽 𝑁 −1 )] ∈ C𝑀 ×𝑁 and 𝜏 > 0 controls sparsity. We adopt ADMM to decouple data fidelity and sparsity by introducing 𝒛 ∈ C𝑁 and enforcing 𝒙 = 𝒛: 1 min ∥𝒓 − 𝑨𝒙 ∥ 22 + 𝜏 ∥𝒛 ∥ 1 s.t. 𝒙 = 𝒛. (6) 𝒙,𝒛 2 The ADMM updates read 𝒙 (𝑡 +1) = (𝑨H 𝑨 + 𝜌𝑰 ) −1 (𝑨H 𝒓 + 𝜌 (𝒛 (𝑡 ) − 𝒗 (𝑡 ) )), 𝒛 (𝑡 +1) = R𝜏/𝜌 (𝒙 (𝑡 +1) + 𝒗 (𝑡 ) ), 𝒗

(𝑡 +1)

=𝒗

(𝑡 )

+𝒙

(𝑡 +1)

−𝒛

(𝑡 +1)

(7) ,

with element-wise soft threshold R(0) = 0 and 𝒖 𝜏 R𝜏/𝜌 (𝒖) = max(0, |𝒖| − ). |𝒖| 𝜌

(8)

In practice, fixed update rules Eqn. (7) and a fixed threshold Eqn. (8) may over-shrink true components and slow convergence in coherent-source regimes, and the matrix inverse can be costly without structure [43]. Deep Unrolling. To overcome these issues while iteratively optimizing the spatial spectrum with ADMM, as shown in Fig. 4, we preserve the steering model and the classical input and output interface, and unroll the solver into a trainable network that learns only coefficients typically fixed in classical methods. Inspired by the iterative gated recurrent unit (GRU) [12] architecture, we add a lightweight gate that connects 𝒛 (𝑡 ) and 𝒛 (𝑡 +1) , so each update blends the fresh

MobiCom ’26, October 26–30, 2026, Austin, TX, USA

Yu and Wu (a) CFAR Detector

(c) Results

!

UWB Data Threshold &

'

(b) RF-LEGO Detector Feedforward Mat # !

Input Mat "

SUM

Σ

(

State Mat '

Output Mat &

$%

Target

High-SNR

Target

Strong Interference

CFAR Detector vs. RF-LEGO Detector Detected w/ Few False Alarm

Probability

Testing "" (·) Selection "! (·)

Normalized Value

candidate with the previous iterate. This acts as learned relaxation or momentum and is known to improve stability and speed in unrolled sparse solvers [62]. We also replace the fixed soft-threshold with a learnable shrinkage level and parameterize step sizes to be positive via a softplus mapping, following evidence that learned thresholds and steps accelerate convergence and enhance robustness [19, 26]. Finally, a learnable diagonal preconditioner 𝑾 (𝑡 ) keeps the linear solver inexpensive and provides data-adaptive weighting without resorting to subspace factorization [6]. The resulting solver removes reliance on brittle eigen-decompositions and, by learning an adaptive update strategy, achieves robust angular resolution even under challenging conditions. At each iteration, the RF-LEGO Beamformer updates read

Target Masked by Interference

Detected w/ High-Confidence Detection

Detected w/ Interference Suppressed

# ToF Bin

Figure 5: RF-LEGO Detector. (a) Classical CFAR uses a fixed sliding window. (b) RF-LEGO unrolls the logic into a state space model, where 𝒓 is the input signal vector, 𝒛 represents the learned state vector, and 𝒔ˆ is the detection vector derived from the state.

At each step, a specific sample, denoted 𝒓 𝑛 , is designated as the cell under test, while its surrounding samples, excluding a guard region, form a local neighborhood 𝒓 𝑔𝑢𝑎𝑟𝑑 . The detection output 𝒔 ∈ R𝑁 can then be expressed as:

𝒙 (𝑡 +1) = (𝑾 (𝑡 ) + 𝜼 (𝑡 ) 𝑰 ) −1 (𝑨H 𝒓 + 𝜼 (𝑡 ) (𝒛 (𝑡 ) − 𝒗 (𝑡 ) )), 𝒈 (𝑡 ) = 𝜎 (𝑾 𝑔 𝒛 (𝑡 ) + 𝑼 𝑔 𝒗 (𝑡 ) ), 𝒛 (𝑡 +1) = 𝒈 (𝑡 ) ⊙ (𝒙 (𝑡 +1) + 𝒗 (𝑡 ) ) + (1 − 𝒈 (𝑡 ) ) ⊙ 𝒛 (𝑡 ) ,

𝒔𝑛 = CFAR(𝒓 𝑛 ) = 𝑪𝒈 1 (𝒓 𝑛 ) + 𝒈 2 (𝒓 𝑛 ),

(9)

𝒗 (𝑡 +1) = 𝒗 (𝑡 ) + 𝒙 (𝑡 +1) − 𝒛 (𝑡 +1) , where 𝑾 (𝑡 ) = diag(𝒘 (𝑡 ) ) is a learnable diagonal preconditioner, 𝜼 (𝑡 ) = softplus(𝜼˜ (𝑡 ) ) > 0 enforces stable steps, and 𝒈 (𝑡 ) ∈ [0, 1] 𝑁 blends new and historical estimates.

3.3

(10)

RF-LEGO Detector

Detection underpins localization, tracking, and activity recognition. Classical matched detectors and Bayesian detectors [58] require prior signal knowledge and degrade in dynamic, non-cooperative settings. Non-coherent energy and cyclostationary detectors (e.g., autocorrelation) offer alternatives but adapt poorly [61]. CFAR and its variants, i.e., Cell Averaging (CA-CFAR) [60], Order Statistic (OS-CFAR) [47], (Greatest Of) GO-CFAR [23], and (Smallest Of) SO-CFAR [57], adjust thresholds using surrounding samples, yet still demand expert tuning and remain sensitive to heterogeneous clutter, limiting reliability in complex scenes. Recently, neural CFARlike detectors report high accuracy with little manual tuning [34, 52], but they are often black boxes with limited physical relevance for RF sensing [61] and are typically task-specific. There are already unrolled approaches for CFAR detection, such as CNN-based CFAR [34], CNN-LSTM for maritime radar, OTFS-CFAR with neural components [52], CFARNet [15], and VAMP-CFAR [68]. However, these prior works often face a critical dilemma: they either over-simplify the noise estimation model to ensure differentiability, which limits robustness in complex clutter, or they insert black-box neural components that undermine the very interpretability that unrolling aims to preserve. Signal Model. The CFAR-based detection algorithms operate by systematically applying selection and hypothesis testing across a sliding window over the input signal 𝒓 ∈ R𝑁 .

where 𝒈 1 (·) represents a selection operator that estimates the local noise level. The selection operator can be linear (e.g., CA-CFAR) or non-linear (e.g., OS-CFAR). 𝒈 2 (·) serves as a testing operator that compares this estimate to the signal in the test cell. Typically, the testing operator takes a linear form, i.e., 𝒈 2 (𝒓 𝑛 ) = 𝐵𝒓 𝑛 , where 𝑩 is a linear transformation matrix and often simplifies to 𝑩𝒓 𝑛 = −𝒓 𝑛 . This formulation captures the essence of CFAR: a dynamically adapted thresholding process anchored in local statistics. Again, its performance is sensitive to the choice of parameters such as the threshold multiplier 𝑪 and the configuration of the guard and training cells. These choices require manual tuning and are often non-trivial, particularly in complex environments where the improper configuration can significantly impair detection reliability. Deep Unrolling. Our key insight to unroll CFAR is that it performs adaptive thresholding driven by temporally estimated noise, which aligns naturally with a state space model (SSM) architecture [20, 22]. Moreover, an SSM offers minimal structured memory to track clutter or drift over time without altering the decision diagram of CFAR. We thus unroll CFAR as a discrete SSM, preserving the workflow and operating-point control while learning the noise/clutter dynamics. We therefore unroll CFAR’s selection-testing routine into a discrete SSM, yielding a trainable detector with interpretability, which preserves the CFAR workflow. Rather than applying the fixed operators 𝒈 1 and 𝒈 2 to the neighborhood samples 𝒓 guard , we introduce a latent state 𝒛𝑛 that evolves under learned dynamics. This hidden state will not affect the main information diagram but additionally aggregates higher-order statistics and other context and acts as an unobserved variable set, enriching the representation of the system’s internal behavior [35, 46]. Fig. 5 shows the architecture: the CFAR workflow is preserved, while the matrices

RF-LEGO

MobiCom ’26, October 26–30, 2026, Austin, TX, USA

are trainable. The unrolled model is 𝒛𝑛 = 𝑨𝒛𝑛−1 + 𝑩𝒙 𝑛 ,

𝒔𝑛 = 𝑪𝒛𝑛 + 𝑫𝒙 𝑛 ,

(11)

where 𝑨 governs state evolution, 𝑩 and 𝑪 project inputs and states, and 𝑫 is a feedforward term that plays the role of the CFAR testing operator 𝒈 2 (·). This formulation replaces the selection 𝒈 1 (·) with the learned state 𝒛𝑛 and the fixed test with a linear readout over the state and current input. This discrete model can be viewed as the trapezoidal discretization of the continuous-time system: 𝑑𝒛 (𝑡) = 𝑨𝒛 (𝑡) + 𝑩𝒙 (𝑡), 𝒔 (𝑡) = 𝑪𝒛 (𝑡) + 𝑫𝒙 (𝑡), (12) 𝑑𝑡 which reduces to Eqn. (11) by absorbing the step size Δ𝑡 into the learnable parameters. To numerically solve this ordinary differential equation, we adopt the Trapezoidal Rule to approximate the integral between steps 𝑡𝑛−1 and 𝑡𝑛 = 𝑡𝑛−1 + Δ𝑡: Δ𝑡 [𝑨𝒛𝑛 + 𝑨𝒛𝑛−1 + 𝑩𝒙 𝑛 + 𝑩𝒙 𝑛−1 ] 2  (13)  ⇒ 𝒛𝑛 = (𝑰 − Δ𝑡2 𝑨) −1 (𝑰 + Δ𝑡2 𝑨)𝒛𝑛−1 + Δ𝑡 𝑩𝒙 𝑛 ,

𝒛𝑛 − 𝒛𝑛−1 ≈

which simplifies to the form in Eqn. (11) by absorbing the time-step Δ𝑡 into the learnable parameters. Therefore, our network preserves physical fidelity to the underlying continuous dynamics while allowing flexible, data-driven adaptation via trainable matrices 𝑨, 𝑩, 𝑪, 𝑫. As illustrated in Fig. 5, the design preserves the CFAR structure of selection and testing, adds memory to store latent information, and introduces trainable components. In practice, the detector adapts to environmental variation, generalizes better across unseen scenarios, and remains interpretable because each component maps to a clear physical role.

3.4

Remarks

RF-LEGO pioneers deep unrolled SP-DL co-design for RF sensing. We present three unrolled algorithms for the most fundamental processing in sensing applications, i.e., frequency transform, spatial spectrum estimation, and detection. The three proposed modules can be used separately or combined together, like LEGO blocks. For example, one can perform RF-LEGO FT and then RF-LEGO Detector for breathing rate estimation. They can also be flexibly reused and integrated into prior and emerging sensing solutions, since the unrolled modules preserve a similar interface, just like the classical signal processing techniques. We leave it as future work to unroll the above three techniques in different ways and unroll other important techniques in RF sensing.

4

IMPLEMENTATION

A key advantage is that each RF-LEGO module is trained on synthetic data, yet, as our experiments show, it generalizes robustly to diverse real-world scenarios. We synthesize 30k

frames for each module that mimic real measurements while exposing controlled variability. RF-LEGO FT. We first construct clean frequency-domain spectra, then inject typical artifacts (e.g., spectral leakage) and additive white Gaussian noise across 5-40 dB SNR. An inverse Fourier Transformation converts these spectra to noisy time-domain signals for training. The default frequency transform points are 256. RF-LEGO Beamformer. We simulate a uniform linear array with a steering dictionary. We draw a random set of sources with random directions and complex amplitudes to form a clean array snapshot, then add white Gaussian noise. The ground truth is a sparse vector on a discrete angle grid whose nonzero entries correspond to the dictionary indices nearest the true directions. By default, the array has 8 antennas, 10 ADMM layers, and the steering grid spans from -60 to 60 degrees at 1-degree resolution. RF-LEGO Detector. We synthesize one-dimensional signals by superimposing 1-5 target peaks on unit-variance white Gaussian noise. Peaks use Hann or Hamming shapes with varying widths and are scaled to achieve 5-40 dB SNR. Labels are binary masks marking the peak locations. Each sample has a length of 128. All modules are implemented in PyTorch and trained on a single NVIDIA GeForce RTX 4090. We use AdamW with a learning rate 1 × 10−3 and a weight decay of 0.01, batch size 512, and an 80/20 train-validation split; dropout of 0.2 is applied within modules. Losses are task-aligned: cosine similarity for RF-LEGO FT and RF-LEGO Beamformer, and binary cross-entropy for RF-LEGO Detector. Computation/size costs are: RF-LEGO FT (0.02 GFLOPs, 0.8M), Beamformer (0.83 GFLOPs, 0.6M), and Detector (0.05 GFLOPs, 0.4M).

5

EVALUATION

Our evaluation benchmarks individual modules against signal processing, deep learning, and loose coupling baselines, validates plug-and-play cascadability in existing pipelines, and analyzes structure-aligned interpretability.

5.1

Experiment Design

5.1.1 Experiment Setup. To evaluate the modularity of RFLEGO across diverse RF-based platforms, we utilize a range of IoT devices, including mmWave radar, UWB radar, and Wi-Fi. Evaluation experiments are tailored to the distinct features of each sensor’s electromagnetic properties and modulation techniques. • mmWave. IQ-modulated FMCW raw signals are captured from the TI IWR1843 [54] using the DCA1000EVM evaluation board [53], operating within the 77-81 GHz band. • UWB. Impulse-based time-of-flight (ToF) signals are collected using the Novelda XeThru X4A02 ultra-wideband

MobiCom ’26, October 26–30, 2026, Austin, TX, USA a

Indoor Open Space Target

3.5m

Equipment

Sensor

b

Workspace

c

d

Refuge Room

Tables & Chairs

8.0m

Yu and Wu

Vacant Space 2.3m

Corridor

e

Construction Corridor

f

Open Area

g

Lab

Equipment

h

Indoor Open Space

i

Outdoor Square

Construction Waste

Metal Window Bars

7.8m

FoV = 120 deg

>10m

>10m

Velocity

Velocity

5.0m

Figure 6: Experimental scenarios. (a-f) Range experiments; (g-h) Doppler experiments; (g,i) Angle experiments.

radar [42], which operates at a center frequency of 7.29 GHz and a frame rate of 100 Hz. • Wi-Fi. Wi-Fi signals are acquired using a commercial router and an Intel AX200 NIC. The transmitter antennas send packets to two receiver antennas, enabling the extraction of channel state information (CSI). To accurately benchmark our algorithms, we selected a metal plate as an ideal point-like reflector. Its large radar crosssection not only ensures a strong and unambiguous echo for simplified ground truth acquisition but also minimizes confounding variables arising from complex target geometries. As Fig. 6 shows, we design distinct, controlled experiments across nine scenarios for the three key sensing dimensions: Range: The target is positioned at multiple distances from the sensor to capture distance-dependent signal returns. Doppler: The plate is mounted on a programmable linear rail and driven at precisely controlled velocities, with the sensor fixed at one end of the rail to record the resulting frequency shifts. Angle: To evaluate performance at varying distances, the target is placed on circular tracks with 3 m and 5 m radii (corresponding to scenarios g and i, respectively) and rotated through a sweep of azimuth angles relative to the sensor. For each modality and each scenario, we collect approximately 36k samples spanning range, Doppler, and angle, and every sequence is labeled with its exact measurement.

5.1.2 Evaluation Metrics. To quantify the performance of the three RF-LEGO modules, we use a suite of metrics that assess their core capabilities in spectral analysis, parameter estimation, and target detection. Peak-to-Side Lobe Ratio (PSLR): Measures the ratio between the main lobe’s peak and the highest side lobe, quantifying spectral leakage suppression. A higher value indicates better isolation of strong signals from weaker adjacent. Peak-to-Average Power Ratio (PAPR): Evaluates spectrum sparsity by comparing the peak power to the average power. A higher PAPR implies more concentrated signal energy and a cleaner, less cluttered spectrum. Mean Absolute Error (MAE): Computes the absolute error between the estimated target and its ground truth value. A lower MAE indicates higher estimation accuracy for target parameters such as range, Doppler, or angle. Detection Rate (DR): Represents the percentage of correctly identified targets, reflecting the model’s ability to resolve distinct objects.

5.1.3 Baselines. We compare RF-LEGO against various signal processing, deep learning and loose coupling baselines. • Signal Processing Baselines: We compare our three modules against their most fundamental and direct signal processing counterparts: the conventional FFT, the iterative LASSO beamformer, and the CFAR detector, respectively. • Deep Learning Baselines: (i) CubeLearn: [74] Represents the purely data-driven approach, which learns features directly from raw signals. For a fair comparison, we adapt it as a learned frequency transform front-end by removing its final classifier. (ii) DA-MUSIC: [39] Represents a hybrid approach that integrates a deep neural network into the classical MUSIC pipeline to improve beamforming, serving as a key baseline for data-driven angle estimation. (iii) DL-CFAR: [34] Represents the component-replacement approach, where the noise estimation stage of classical CFAR is replaced with a ResNet-based module, enhancing a specific component in a purely neural network manner. • Loose Coupling Baselines (LC): We cascade a fixed SP front-end with a deep learning back-end, where the neural layers are configured to match the parameter count of the corresponding RF-LEGO modules for a fair comparison.

5.2

Experimental Evaluation

5.2.1 Modularity Evaluation. To validate the modularity of RF-LEGO, we conduct a series of rigorous module-swap control experiments. In this paradigm, we keep the classical signal processing pipeline intact, replacing only a single block with its RF-LEGO counterpart to cleanly isolate and quantify the contribution of each component. First, as Fig. 7 illustrates, the RF-LEGO FT module demonstrates a marked performance improvement over the conventional FT in frequency transformation tasks across all tested modalities. By suppressing spectral leakage through its learned, data-driven filter, the module achieves an average PSLR of around 24 dB and 27 dB and PAPR of approximately 16 dB and 23 dB for range and Doppler spectra, respectively. This superior spectral quality directly translates into more robust peak detection, significantly reducing range and velocity estimation errors (MAE). As shown in Fig. 7, this corresponds to an average MAE reduction of approximately 10% for range and 20% for Doppler compared to the signal processing baseline. Compared with the deep learning baseline, CubeLearn [74], RF-LEGO FT achieves a comparable accuracy, while offering the distinctive interpretability. And it

RF-LEGO

Scenario a

MobiCom ’26, October 26–30, 2026, Austin, TX, USA

Scenario b

Scenario c

Scenario d

Scenario e

Scenario f

Scenario g

mmWave

Scenario h

Scenario g

UWB

Scenario h

Scenario g

Wi-Fi

Scenario h

Figure 7: The results of RF-LEGO Range FT of mmWave and RF-LEGO Doppler FT of mmWave, UWB, and Wi-Fi. Scenario a

Scenario g

Scenario b

Scenario c mmWave

Scenario i Scenario d

Scenario e

UWB

Wi-Fi

Scenario f

Figure 8: The results of RF-LEGO Beam- Figure 9: The results of RF-LEGO Detec- Figure 10: Cascadability evaluation of RFLEGOvs.SP and LC baselines. tor of UWB ToF signals. former of mmWave.

consistently outperforms the loose coupling baseline, due to the structured SP-DL co-design. Subsequently, the RF-LEGO Beamformer demonstrates a significant leap in spatial angle estimation over the classical LASSO baseline, as in Fig. 8. On mmWave signals across scenarios g and i, it reduces MAE from 4.23 to 1.35 degrees by 68%. Although DA-MUSIC [39] also offers similar accuracy, RF-LEGO achieves this competitive performance with two key advantages. First, it retains pipeline interpretability by outputting a physically meaningful angular spectrum. Second, its ADMM-based structure avoids the reliance on eigenvalue decomposition, an operation central to the MUSIC algorithm. This is a critical design choice, as eigenvalue decomposition is known to have unstable gradients within deep learning frameworks, often complicating or destabilizing the training process for models like DA-MUSIC [45]. Despite the strong performance of the LC baseline, RF-LEGO Beamformer still achieves a lower mean angular error. Finally, Fig. 9 shows the performance of RFLEGO Detector for target detection using UWB ToF signals. We compare RF-LEGO Detector against both the traditional CFAR family and the end-to-end DL-CFAR model [34] on the UWB dataset. At a matched false alarm rate of 10−3 , our RF-LEGO Detector achieves a considerably higher DR than classical CFAR and a competitive performance compared to the DL-CFAR baseline. Crucially, RF-LEGO Detector allows for explicit control over the detector’s operating points, e.g., the trade-off between detection and false alarm rates, an important feature for building practical systems in realworld RF sensing. Moreover, RF-LEGO Detector performs on par with the loose coupling baseline even under challenging conditions, while maintaining the co-designed SP structure. In summary, these modularity evaluations prove RF-LEGO as drop-in modules that consistently surpass its classical SP counterparts, the loose coupling baseline, and prior DL methods across different metrics.

5.2.2 Cascadability Evaluation. We also demonstrate that RF-LEGO components are cascadable, composing seamlessly into high-performance pipelines with compounding benefits. We test this by constructing two hybrid pipelines: (i) RF-LEGO FT plus RF-LEGO Detector for range/velocity estimation, and (ii) RF-LEGO Beamformer plus RF-LEGO Detector for angle estimation. We benchmark these against classical pipelines, cascaded loose-coupling pipelines, and cascaded end-to-end baselines. We use the Detection Rate (DR) of the ground truth target bin at a fixed false alarm rate of 10−3 as the primary metric. As shown in Fig. 10, the cascaded RF-LEGO pipelines achieve higher DR than classical pipelines, with improvements of 27.6% (range), 27.3% (Doppler), and 40.2% (angle) on mmWave data, and further gains of an average of 28.2% and 14.7% over the cascaded deep learning and loose coupling baselines. Similar advantages are observed on UWB and Wi-Fi data. Results confirm RF-LEGO outperforms standalone methods while enabling reusable, cascadable complex sensing. 5.2.3 Interpretability Analysis. RF-LEGO ensures structurealigned interpretability through the SP-structured module design. This structure interpretability allows each trainable module to be directly inspected and reasoned about using principled, SP-style analysis of intermediate results with classical criteria and metrics. For instance, we can calculate PSLR/PAPR for RF-LEGO FT outputs in Fig. 7, quantifying spectral quality (e.g., sidelobes) in the same way commonly used in conventional spectral analysis via signal processing. Moreover, the structure-aligned interpretability supports flexible and robust cascading when composing multi-module pipelines. Thanks to the interpretable structures with trainable modules, RF-LEGO-enhanced pipelines are more accurate and more resilient to inter-module error propagation, as demonstrated in Fig. 10 and Fig. 12.

MobiCom ’26, October 26–30, 2026, Austin, TX, USA

5.3

Yu and Wu

Table 1: Ablation study.

Microbenchmarks

5.3.1 Performance on Public Datasets. We evaluate RF-LEGO across multiple online datasets: UWCR [17, 18]: The dataset contains real urban and parkinglot driving scenes with synchronized radar-camera data. We use it to assess RF-LEGO Range FT, Doppler FT, and Beamformer. The ground truth is provided by radar-camera fusion annotations derived from vision labels aligned to radar frames with manual verification. OPERAnet [4]: The dataset offers indoor Wi-Fi CSI with synchronized Kinect and camera data. We align skeleton trajectories to the sensor frame to obtain radial-velocity ground truth and evaluate RF-LEGO Doppler FT. UWB-Context [5]: The dataset contains residential multistatic UWB CIRs. We convert CIR to ToF to evaluate RFLEGO Doppler FT and Detector. Ground truth 2D positions are provided by a synchronized, independent active UWB. DeepSense 6G [1]: The dataset features large-scale realworld 60 GHz millimeter-wave communication measurements, i.e., 6G, collected via a phased array by beam sweeping. We utilize the received power vectors obtained during the scan to assess RF-LEGO Detector for robust beam selection tasks. The ground truth is determined by mmWave radar. 5.3.2 Impact of Data Fine-tuning. By default, RF-LEGO modules are trained exclusively on synthetic data. For our main experiments, these pre-trained models are then directly evaluated on real-world data without any fine-tuning. In this experiment, we fine-tune them with a varying fraction of real data per modality from 0-100%. As shown in Fig. 11(c), fine-tuning with real-data further enhances RF-LEGO: the MAE decreases for Range FT (about 18% from 0% to 100%), Doppler FT on mmWave, UWB, and Wi-Fi (about 15-25%), and the Beamformer (about 15%), while the DR of Detector rises by about 5-6% at a fixed order of magnitude of FAR. The gains are monotonic and begin to plateau after roughly 60-80%, indicating effective low-shot adaptation with diminishing returns at larger fractions. This suggests a practical deployment workflow where a small amount of in-domain data can be used for finetuning to recover most gains. 5.3.3 Impact of Multiple Targets. To assess the effect of multiple targets, we conduct mmWave experiments with 2-4 targets placed at random ranges of 1-4 m and random azimuths of -60 to 60 degrees. We collect approximately 8k samples across 27 configurations. As Fig. 11(b) shows, RF-LEGO consistently outperforms classical signal processing baselines across all target counts and metrics, indicating robustness to mutual interference and clutter without fine-tuning. 5.3.4 Interface Sensitivity under Cascading. Fig. 12 evaluates error propagation by injecting noise of different SNR levels between cascaded mmWave sensing modules and measuring

Module / Metric

Default

③a

③b

Range FT / PSLR [dB] Doppler FT / PSLR [dB] Beamformer / MAE [deg]

23.27 27.86 1.35

20.16 21 2.1

22.08 25.64 2.23

22.9 26.3 1.8

23.38 27.3 5.7

Detector / DR@FAR [10-3] [email protected] 0.89@8 0.88@10 [email protected] [email protected]

the normalized detection rate. As expected, noise degrades performance for all pipelines; Nevertheless, RF-LEGO consistently exhibits slower performance decay than pure signal processing pipelines, indicating stronger resilience to upstream errors. Moreover, robustness improves as more RFLEGO modules are integrated, suggesting that the unrolled, physics-constrained operators remain compatible under cascading and mitigate cumulative error amplification. Overall, these results confirm that RF-LEGO ’s cascadability is not merely due to downstream tolerance, but stems from enhanced error resilience at each unrolled module. 5.3.5 Inference Latency on Edge Devices. We also report the inference latency breakdown of each RF-LEGO module on representative edge platforms under default settings. As detailed in Tab. 2, RF-LEGO achieves fast performance on the Jetson Orin Nano by incurring additional computation while improving task performance. While it remains practical on the Raspberry Pi 4, the microcontroller-class ESP32-P4 presents the most significant constraints. For practical deployment, effective model compression can be further performed for extremely resource-constrained hardware.

5.4

Ablation Study

We evaluate the unrolled design of each module, and the results are depicted in Tab. 1. RF-LEGO FT ① w/o activation: removing nonlinearities degrades PSLR by 3.1 dB and 6.9 dB for Range and Doppler FT, respectively, confirming that mild nonlinearity suppresses leakage beyond pure signal processing convolution. ② w/o nulling: performance is comparable to the main model, indicating gains mainly come from the unrolled architecture rather than the nulling head. ③ # points: using 128 (a) or 512 (b) points is on par with the default 256, showing good generalizability. RF-LEGO Beamformer ① w/o iteration connection: removing the explicit link to the previous iterate increases MAE (from 1.35 to 2.10 degrees), highlighting the benefit of iterative memory under challenging conditions. ② w/o gating: replacing the learnable gate with a fixed soft-threshold further worsens MAE to 2.23 degrees, showing that the dataadaptive step is key to robustness. ③ # antennas: reducing the ULA from 8 to (a) 6 causes moderate degradation, while

RF-LEGO

MobiCom ’26, October 26–30, 2026, Austin, TX, USA

DR@FAR[10!" ]=0.98

a

c

b

UWB ! RF-LEGO Doppler FT Wi-Fi ! RF-LEGO Doppler FT UWB ! RF-LEGO Detector

100.0%

3 95.0% 2

DR

MAE [cm or cm/s or deg]

mmWave ! RF-LEGO Range FT mmWave ! RF-LEGO Doppler FT mmWave ! RF-LEGO Beamformer

90.0%

1 +0% +20% +40% +60% +80% +100% Real-world Data

85.0%

Figure 11: Microbenchmarks. (a) Performance on public datasets. (b) Impact of multiple targets. (c) Impact of data fine-tuning. Table 2: Inference latency on different edge platforms. The results are formatted as RF-LEGO/ SP [ms].

Device

FT

Beamformer

Detector

2.60 / 0.73 8.33 / 1.93 81.68 / 3.80

6.20 / 1.74 30.86 / 4.59 122.46 / 7.76

Jetson Orin Nano 0.76 / 0.03 Raspberry Pi 4 3.08 / 0.17 ESP32-P4 31.07 / 0.36

a b

100 100 80 60 80 40 No Noise 60

Performance [%]

Performance [%]

+noise

40

30 20 SNR [dB]

10

+noise

+noise Module °! RF-LEGO Detector FrontFront Module °°°!°° RF-LEGO Detector

0

No Noise

40

30 20 SNR [dB]

10

0

Figure 40 12: Effect of interface sensitivity under cascading. No Noise

40

30 20 SNR [dB]

10

0

No Noise

40

30 20 SNR [dB]

10

reducing to (b) 4 sharply degrades accuracy to 5.7 degrees, matching the expected resolution loss. RF-LEGO Detector ① w/o activation: detection rate drops slightly while false alarm rate increased by approximately 2.8×, indicating that nonlinear capacity helps suppress false alarms in co-design. ② w/o trainable state: setting the state matrix to a fixed identity matrix, i.e., 𝑨 = 𝑰 , reduces DR to 0.88 and raises FAR by an order of magnitude, showing the state stores useful latent information. ③ input length: using (a) 64 or (b) 256 bins yields results close to default (128).

5.5

Module Behavior Analysis

Compared with Bluestein’s Algorithm FT, whose signal processing convolution kernel has a uniform response across kernel bins, the RF-LEGO FT learns a data-driven convolution kernel. As shown in Fig. 13(a), this manifests as a clear reweighting pattern over kernel bins, enabling the FT module to adapt the intermediate convolution to non-ideal measurements and thus outperform the fixed-kernel baseline in downstream performance. Fig. 13(b) shows the evolution of the learned gate 𝑔 (𝑡 ) across unrolled iterative blocks. The gate values indicate an iteration-dependent focus between the previous iterate and the newly-updated estimate in the update candidate, i.e., 𝒙 (𝑡 +1) + 𝒗 (𝑡 ) in Eqn. (9). Specifically, larger gates in early blocks encourage stronger incorporation of the update candidate, while smaller gates in later blocks emphasize stability, reflecting a learned coarse-tofine scheduling that improves convergence behavior. Fig.

c

Figure 13: Module behavior analysis. (a) RF-LEGO FT learns a non-uniform convolution kernel beyond Bluestein’s Algorithm FT. (b) RF-LEGO Beamformer learns adaptive gate schedules across unrolled iterations. (c) RF-LEGO Detector learns a structured state matrix 𝑨, unlike memoryless CFAR.

Bluestein’s Algorithm FT RF-LEGO RF-LEGO FT LASSOLASSO Beamformer RF-LEGO RF-LEGO Beamformer Algorithm FT FT Beamformer Beamformer Front Module: Bluestein’s +noise Module °! CFAR FrontFront Module °°°!°° CFAR

RF-LEGO Beamformer

0

13(c) compares the learned state matrix 𝑨 with the memoryless CFAR baseline. While CFAR corresponds to a null state transition, RF-LEGO Detector learns a structured 𝑨 and meanwhile maintains the stable structural pattern inherited from its initialization. This is desirable because it preserves a constrained state evolution while still allowing data-driven adaptation, yielding a detector with memory modeling beyond local statistics.

6

CASE STUDY

We evaluate RF-LEGO’s impact by replacing SP modules in three pipelines, as shown in Fig. 14, while keeping downstream components fixed. For tracking, RF-LEGO replaces the baseline’s [21] FFT, MUSIC, and CFAR feeding an extended Kalman filter. For vital signs, we compare breathing rates against [55] using identical phase analysis. For activity recognition [32], RF-LEGO replaces the SP front-end of the baseline [31] classifier.

6.1

Trajectory Tracking

6.1.1 Experimental Design. As shown in Fig. 15(a-b), we place the mmWave radar at the corner of a 4m×4m room instrumented with an OptiTrack motion capture system [41] to obtain ground truth human trajectories. The dataset contains 65 designated trajectories spanning 13 categories, including straight, turn, circle, and 5 random trajectories. We compare two pipelines in Fig. 14, keeping pre-processing, tracking, and hyperparameters identical (both processing and Kalman filter) to ensure a fair test. We use Absolute Trajectory Error (ATE) and Relative Trajectory Error (RTE) as the metric,

MobiCom ’26, October 26–30, 2026, Austin, TX, USA

Yu and Wu

Feature Extraction

Classification

CFAR Detection

Tracking

Activities Raw Data

Range FFT

Doppler FFT

MUSIC Beamforming

a

OptiTrack Motion Capture Systems

b

c Sensor

Pre-processing

RF-LEGO Range FT

RF-LEGO Doppler FT

RF-LEGO Beamformer

RF-LEGO Detector

Sensor

Trajectories

MUSIC Beamforming

Phase Demodulation

Sensor Breathing Belt

RF-LEGO 4.0m x 4.0m

Vital Signs Range FFT

e

d

Marker

Infant Simulator

Breathing Estimation

Peak Detection

Breathing Belt

Sensor

3.5 m

CDF

Figure 15: Case study scenarios. (a-b) Trajectory tracking; (c-e) Figure 14: Pipelines for trajectory tracking, vital sign monitor- Vital sign monitoring for infant simulator and human adults. ing, and human activity recognition with RF-LEGO modules vs.classical baselines. 1.0 0.5 0.0 Signal Processing

Infant Simulator

Human Adult

OptiTrack (Ground Truth)

Figure 17: Human trajectory tracking for letters RF-LEGO. Colors fade from light (start) to dark (end). Respiration Rate [BPM]

Figure 16: CDF of ATE and RTE.

RF-LEGO

Ground Truth

Signal Processing

RF-LEGO

60 Time [s]

100

7-Way Classification

SP RF-LEGO

0

5 MAE [BPM]

Figure 18: CDF of breathing MAE. 13-Way Classification

30 20 10 0

20

40

80

Figure 19: Performance for an infant sim- Figure 20: Natural breathing monitoring Figure 21: Accuracy for 7-way and 13-way ulator and the human adult. classification across various settings. by a human adult.

where ATE indicates the RMSE of the predicted trajectory and the ground truth trajectory, and RTE measures the RMSE between both trajectories over a time interval. And we use a 1-second time window for evaluation. 6.1.2 Evaluation. The RF-LEGO-integrated pipeline demonstrates superior performance in the trajectory tracking task. Quantitative results are presented in Fig. 16. Compared to the traditional SP baseline, the RF-LEGO pipeline significantly reduces both error metrics. For instance, the median of ATE for RF-LEGO is about 0.3 meters, a 40% error reduction from the traditional baseline. Similar gains are observed for RTE, further confirming RF-LEGO’s high precision and stability throughout the tracking process. Fig. 17 visualizes example tracking trajectories. As seen, the RF-LEGO trajectories (red) closely align with the ground truth (gray). In contrast, the trajectories from the SP pipeline (blue) exhibit noticeable drift and greater deviation, especially when tracing complex letter shapes. This indicates that by providing a higher-quality, lower-noise point cloud, RF-LEGO enhances the overall tracking robustness and precision.

6.2

Vital Sign Monitoring

6.2.1 Experimental Design. As shown in Fig. 15(c-e), we evaluate breathing in two setups: an infant simulator with programmable breathing rates and human adults instrumented

with a respiratory belt for ground truth. The simulator is tested at {20, 40, 50, 60, 80, 100} BPM. Human trials span {6, 8, 10, 12, 15, 18, 20} BPM, plus spontaneous resting sessions with one and two participants. Each condition is recorded for 6 minutes. We compare the two pipelines in Fig. 14, holding pre-processing, demodulation, and the rate estimator fixed to ensure a fair comparison. We measure the MAE of BPM as the evaluation metric. 6.2.2 Evaluation. Fig. 18 presents the CDF of the MAE for breathing rate estimation, aggregating the results from both our single and multiple person experiments of scenarios d and e. The result reveals that RF-LEGO achieves an 80%tile MAE of less than 2 BPM, while the traditional baseline yields 5× higher error of 10 BPM. Fig. 19 further details the performance on both an infant simulator and human subjects. Whether for the various breathing frequencies of the infant simulator (ranging from 20 to 100 BPM) or the human adult (ranging from 6 to 20 BPM), As seen, RF-LEGO consistently outperforms the baseline in all cases, achieving a respective MAE reduction of 52.7% and 56.8% compared with the SP baseline for the infant simulator and human adults. Fig. 20 illustrates an example trace of an adult’s natural breathing, where our RF-LEGO-enhanced solution tracks the breathing rates accurately and continuously, while the baseline experiences drifts and artifacts.

RF-LEGO

6.3

Human Activity Recognition

6.3.1 Experimental Design. We further validate the end-toend effectiveness and cascadability of RF-LEGO on more complex sensing tasks, i.e., human activity recognition using the MCD-Gesture dataset [31], a cross-domain mmWave benchmark spanning 6 environments, 25 users, and 5 locations (i.e., user positions relative to the radar), with 6 predefined gestures and 7 additional labeled activities. We evaluate two tasks: 7-way classification (merging the 7 non-gesture activities into a single Non-gesture class, following the dataset protocol) and 13-way classification (recognizing all gestures and activities). The baseline adopts a standard signal processing front-end that forms radar cube representations for a neural classifier [32]. In the RF-LEGO-integrated pipeline, we replace only the corresponding upstream operators with RFLEGO modules, while keeping the classifier backbone identical across methods. We consider four evaluation settings via train/test splits: in-domain (random 4:1 split), cross-user (15 users for training, 10 for testing), cross-environment (3 environments for training, 3 for testing), and cross-location (3 locations for training, 2 for testing). 6.3.2 Evaluation. Fig. 21 summarizes the recognition accuracy across the four settings. RF-LEGO consistently improves over the signal processing baseline in both tasks, indicating that the learned RF front-end produces features that are directly usable by a downstream classifier, proving RF-LEGO’s cascadability with neural blocks as well. Quantitatively, RF-LEGO achieves a higher mean accuracy of 93.92% compared with 91.88% on 7-way classification, and 84.15% compared with 79.31% on the more challenging 13-way classification. The larger gain in 13-way recognition suggests that RF-LEGO is beneficial when finer-grained class boundaries amplify the impact of front-end feature quality, while remaining compatible with the same neural back-end.

7

RELATED WORK

Deep Wireless Sensing. Recent advances in deep wireless sensing (DWS) have seen the successful application of largescale models to various tasks [10, 16, 24, 28, 49, 63, 73, 75, 76, 78, 79]. However, their black-box nature often overlooks the physical properties of RF signals, motivating a surge in signal processing-deep learning (SP-DL) co-design. Key strategies include directly embedding SP blocks into networks [14], developing phase-aware complex-valued models [11, 64], adopting DSP-inspired mechanisms [29, 65, 74], and designing interpretable, complex-valued transformers [70]. Deep Unrolling. Deep unrolling, also known as algorithm unrolling, is a principled SP-DL co-design methodology for building interpretable neural networks by mapping traditional iterative algorithms into the hierarchical structure of a deep neural network [40]. The seminal work in this domain

MobiCom ’26, October 26–30, 2026, Austin, TX, USA

is the Learned Iterative Shrinkage-Thresholding Algorithm (LISTA) [19], which successfully unrolled the ISTA algorithm for sparse coding into a feed-forward network. This foundational concept has since been successfully applied across diverse domains, from image processing tasks like restoration and deblurring [9, 30, 51] to RF sensing with models such as KalmanNet [46] and DA-MUSIC [39]. Its applicability further extends to solving differential equations [8, 36], demonstrating its versatility. RF-LEGO pioneers modular SPDL co-design via deep unrolling for trustworthy RF sensing.

8

DISCUSSION

While our work demonstrates that RF-LEGO successfully bridges the gap between classical signal processing and deep learning, RF-LEGO also reveals a limitation that merits further investigation: information alienation. Although the unrolled modules perform exceptionally well on downstream tasks, their learned internal representations may diverge significantly from the clean, canonical outputs with clear physical definitions produced by traditional algorithms. This indicates that our interpretability claim is primarily structurealigned: we preserve well-defined input/output contracts and semantically meaningful intermediate outputs, but we do not guarantee full fidelity or interpretability of all internal learned states within each block. A key future direction is balancing physical fidelity with task optimization to achieve fully interpretable, high-performance sensing.

9

CONCLUSIONS

In this paper, we address a fundamental dilemma in RF sensing that forces a choice between rigid, interpretable signal processing pipelines and adaptable but opaque deep learning models. We introduce RF-LEGO, a novel signal processing-deep learning co-design approach that resolves this trade-off by unrolling cornerstone algorithms, including frequency transform, beamforming, and detection, into trainable, physically-grounded neural operators. Our extensive evaluations demonstrate that these modules not only outperform their traditional counterparts in robustness and accuracy but also compose seamlessly into end-to-end pipelines, significantly enhancing downstream applications. By providing a library of interpretable, reusable LEGO bricks, RF-LEGO opens up a new paradigm of designing flexible, robust, and trustworthy solutions for RF sensing.

ACKNOWLEDGMENTS This work was supported by NSFC under grant No. 62222216, Hong Kong RGC ECS No. 27204522, GRF No. 17212224 and No. 17211725, and CRF No. C5002-23Y. We thank the anonymous Reviewers and Shepherd for their insightful feedback.

MobiCom ’26, October 26–30, 2026, Austin, TX, USA

REFERENCES [1] Ahmed Alkhateeb, Gouranga Charan, Tawfik Osman, Andrew Hredzak, Joao Morais, Umut Demirhan, and Nikhil Srinivas. 2023. DeepSense 6G: A large-scale real-world multi-modal sensing and communication dataset. IEEE Communications Magazine 61, 9 (2023), 122– 128. [2] Maurice Bellanger. 2000. Digital processing of signals: theory and practice. John Wiley & Sons. [3] Leo Bluestein. 2003. A linear filtering approach to the computation of discrete Fourier transform. IEEE Transactions on Audio and Electroacoustics 18, 4 (2003), 451–455. [4] Mohammud J Bocus, Wenda Li, Shelly Vishwakarma, Roget Kou, Chong Tang, Karl Woodbridge, Ian Craddock, Ryan McConville, Raul Santos-Rodriguez, Kevin Chetty, et al. 2022. OPERAnet, a multimodal activity recognition dataset acquired from radio frequency and visionbased sensors. Scientific data 9, 1 (2022), 474. [5] Mohammud J Bocus and Robert Piechocki. 2022. A comprehensive ultra-wideband dataset for non-cooperative contextual sensing. Scientific Data 9, 1 (2022), 650. [6] Stephen Boyd, Neal Parikh, Eric Chu, Borja Peleato, Jonathan Eckstein, et al. 2011. Distributed optimization and statistical learning via the alternating direction method of multipliers. Foundations and Trends® in Machine learning 3, 1 (2011), 1–122. [7] Jack Capon. 1969. High-resolution frequency-wavenumber spectrum analysis. Proc. IEEE 57, 8 (1969), 1408–1418. [8] Ricky TQ Chen, Yulia Rubanova, Jesse Bettencourt, and David K Duvenaud. 2018. Neural ordinary differential equations. Advances in neural information processing systems 31 (2018). [9] Yunjin Chen and Thomas Pock. 2016. Trainable nonlinear reaction diffusion: A flexible framework for fast and effective image restoration. IEEE transactions on pattern analysis and machine intelligence 39, 6 (2016), 1256–1272. [10] Zhe Chen, Tianyue Zheng, Chao Cai, and Jun Luo. 2021. MoVi-Fi: Motion-robust vital signs waveform recovery via deep interpreted RF sensing. In Proceedings of the 27th annual international conference on mobile computing and networking. 392–405. [11] Guoxuan Chi, Zheng Yang, Chenshu Wu, Jingao Xu, Yuchong Gao, Yunhao Liu, and Tony Xiao Han. 2024. RF-diffusion: Radio signal generation via time-frequency diffusion. In Proceedings of the 30th Annual International Conference on Mobile Computing and Networking. 77–92. [12] Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio. 2014. Empirical evaluation of gated recurrent neural networks on sequence modeling. arXiv preprint arXiv:1412.3555 (2014). [13] James W Cooley and John W Tukey. 1965. An algorithm for the machine calculation of complex Fourier series. Mathematics of computation 19, 90 (1965), 297–301. [14] Shuya Ding, Zhe Chen, Tianyue Zheng, and Jun Luo. 2020. RF-net: A unified meta-learning framework for RF-enabled one-shot human activity recognition. In Proceedings of the 18th Conference on Embedded Networked Sensor Systems. 517–530. [15] Tzvi Diskin, Yiftach Beer, Uri Okun, and Ami Wiesel. 2024. CFARNet: Deep Learning for Target Detection with constant false alarm rate. Signal Processing 223 (2024), 109543. [16] Laura Dodds, Tara Boroushaki, Kaichen Zhou, and Fadel Adib. 2025. Non-Line-of-Sight 3D Object Reconstruction via mmWave Surface Normal Estimation. (2025). [17] Xiangyu Gao, Youchen Luo, Guanbin Xing, Sumit Roy, and Hui Liu. 2022. Raw ADC Data of 77GHz MMWave radar for Automotive Object Detection. https://doi.org/10.21227/xm40-jx59

Yu and Wu [18] Xiangyu Gao, Guanbin Xing, Sumit Roy, and Hui Liu. 2021. RAMPCNN: A Novel Neural Network for Enhanced Automotive Radar Object Recognition. IEEE Sensors Journal 21, 4 (2021), 5119–5132. https: //doi.org/10.1109/JSEN.2020.3036047 [19] Karol Gregor and Yann LeCun. 2010. Learning fast approximations of sparse coding. In Proceedings of the 27th international conference on international conference on machine learning. 399–406. [20] Albert Gu, Isys Johnson, Karan Goel, Khaled Saab, Tri Dao, Atri Rudra, and Christopher Ré. 2021. Combining recurrent, convolutional, and continuous-time models with linear state space layers. Advances in neural information processing systems 34 (2021), 572–585. [21] Arjun Gupta, Dashiell Kosaka, Edwin Pan, Jingning Tang, Ruihao Yao, and Sanjay Patel. 2019. OpenRadar: A Toolkit for Prototyping mmWave Radar Applications. arXiv preprint arXiv:1912.12395 (2019). [22] James D Hamilton. 1994. State-space models. Handbook of econometrics 4 (1994), 3039–3080. [23] V. Gregers Hansen. 1973. Constant false alarm rate processing in search radars. In IEE Conference on Radar, Present and Future. 325–332. [24] Weiying Hou and Chenshu Wu. 2024. Rfboost: Understanding and boosting deep wifi sensing via physical data augmentation. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 8, 2 (2024), 1–26. [25] Norden E Huang, Zheng Shen, Steven R Long, Manli C Wu, Hsing H Shih, Quanan Zheng, Nai-Chyuan Yen, Chi Chao Tung, and Henry H Liu. 1998. The empirical mode decomposition and the Hilbert spectrum for nonlinear and non-stationary time series analysis. Proceedings of the Royal Society of London. Series A: mathematical, physical and engineering sciences 454, 1971 (1998), 903–995. [26] Daisuke Ito, Satoshi Takabe, and Tadashi Wadayama. 2019. Trainable ISTA for sparse signal recovery. IEEE Transactions on Signal Processing 67, 12 (2019), 3113–3125. [27] Sijie Ji, Xinzhe Zheng, and Chenshu Wu. 2024. Hargpt: Are llms zeroshot human activity recognizers?. In 2024 IEEE International Workshop on Foundation Models for Cyber-Physical Systems & Internet of Things (FMSys). IEEE, 38–43. [28] Haowen Lai, Gaoxiang Luo, Yifei Liu, and Mingmin Zhao. 2024. Enabling visual recognition at radio frequency. In Proceedings of the 30th Annual International Conference on Mobile Computing and Networking. 388–403. [29] Shuheng Li, Ranak Roy Chowdhury, Jingbo Shang, Rajesh K Gupta, and Dezhi Hong. 2021. Units: Short-time fourier inspired neural networks for sensory time series classification. In Proceedings of the 19th ACM conference on embedded networked sensor systems. 234–247. [30] Yuelong Li, Mohammad Tofighi, Junyi Geng, Vishal Monga, and Yonina C Eldar. 2020. Efficient and interpretable deep blind image deblurring via algorithm unrolling. IEEE Transactions on Computational Imaging 6 (2020), 666–681. [31] Yadong Li, Dongheng Zhang, Jinbo Chen, Jinwei Wan, Dong Zhang, Yang Hu, Qibin Sun, and Yan Chen. 2022. DI-Gesture: DomainIndependent and Real-Time Gesture Recognition with MillimeterWave Signals. In GLOBECOM 2022 - 2022 IEEE Global Communications Conference. 5007–5012. https://doi.org/10.1109/GLOBECOM48099.20 22.10001175 [32] Yadong Li, Dongheng Zhang, Jinbo Chen, Jinwei Wan, Dong Zhang, Yang Hu, Qibin Sun, and Yan Chen. 2023. Towards DomainIndependent and Real-Time Gesture Recognition Using mmWave Signal. IEEE Transactions on Mobile Computing 22, 12 (2023), 7355–7369. https://doi.org/10.1109/TMC.2022.3207570 [33] Qiushi Liang, Yeyue Cai, Jianhua Mo, and Meixia Tao. 2025. CFARNet: Learning-Based High-Resolution Multi-Target Detection for Rainbow Beam Radar. arXiv preprint arXiv:2505.10150 (2025).

RF-LEGO [34] Chia-Hung Lin, Yu-Chien Lin, Yue Bai, Wei-Ho Chung, Ta-Sung Lee, and Heikki Huttunen. 2019. DL-CFAR: A novel CFAR target detection method based on deep learning. In 2019 IEEE 90th Vehicular Technology Conference (VTC2019-Fall). IEEE, 1–6. [35] Wing-Kuen Ling. 2010. Nonlinear digital filters: analysis and applications. Academic Press. [36] Zichao Long, Yiping Lu, Xianzhong Ma, and Bin Dong. 2018. Pdenet: Learning pdes from data. In International conference on machine learning. PMLR, 3208–3216. [37] Sheng Lyu and Chenshu Wu. 2024. ASE: Practical Acoustic Speed Estimation Beyond Doppler via Sound Diffusion Field. arXiv preprint arXiv:2412.20142 (2024). [38] Dmitry Malioutov, Müjdat Cetin, and Alan S Willsky. 2005. A sparse signal reconstruction perspective for source localization with sensor arrays. IEEE transactions on signal processing 53, 8 (2005), 3010–3022. [39] Julian P Merkofer, Guy Revach, Nir Shlezinger, Tirza Routtenberg, and Ruud JG Van Sloun. 2023. DA-MUSIC: Data-driven DoA estimation via deep augmented MUSIC algorithm. IEEE Transactions on Vehicular Technology 73, 2 (2023), 2771–2785. [40] Vishal Monga, Yuelong Li, and Yonina C Eldar. 2021. Algorithm unrolling: Interpretable, efficient deep learning for signal and image processing. IEEE Signal Processing Magazine 38, 2 (2021), 18–44. [41] Inc. (DBA OptiTrack) NaturalPoint. 2025. OptiTrack — Motion Capture Systems. https://www.optitrack.com/. [42] Novelda AS. 2018. XeThru X4 Radar User Guide. https://github.c om/novelda/Legacy- Documentation/blob/master/ApplicationNotes/XTAN-13_XeThruX4RadarUserGuide_rev_a.pdf. Application Note XTAN-13, Rev. A. [43] Carl M O’Brien. 2016. Statistical learning with sparsity: the lasso and generalizations. (2016). [44] Muhammed Zahid Ozturk, Chenshu Wu, Beibei Wang, Min Wu, and KJ Ray Liu. 2023. Radio SES: mmwave-based audioradio speech enhancement and separation system. IEEE/ACM Transactions on Audio, Speech, and Language Processing 31 (2023), 1333–1347. [45] PyTorch Core Team. 2024. torch.linalg.eig — PyTorch Documentation. https://pytorch.org/docs/stable/generated/torch.linalg.eig.html. Accessed: 2025-04-15. [46] Guy Revach, Nir Shlezinger, Xiaoyong Ni, Adria Lopez Escoriza, Ruud JG Van Sloun, and Yonina C Eldar. 2022. KalmanNet: Neural network aided Kalman filtering for partially known dynamics. IEEE Transactions on Signal Processing 70 (2022), 1532–1547. [47] Hermann Rohling. 2007. Radar CFAR thresholding in clutter and multiple target situations. IEEE transactions on aerospace and electronic systems 4 (2007), 608–621. [48] Ralph Schmidt. 1986. Multiple emitter location and signal parameter estimation. IEEE transactions on antennas and propagation 34, 3 (1986), 276–280. [49] Fei Shang, Panlong Yang, Dawei Yan, Sijia Zhang, and Xiang-Yang Li. 2024. LiquImager: Fine-grained Liquid Identification and Container Imaging System with COTS WiFi Devices. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 8, 1, Article 15 (March 2024), 29 pages. https://doi.org/10.1145/3643509 [50] Sharad Singhal and Lance Wu. 1988. Training multilayer perceptrons with the extended Kalman algorithm. Advances in neural information processing systems 1 (1988). [51] Oren Solomon, Regev Cohen, Yi Zhang, Yi Yang, Qiong He, Jianwen Luo, Ruud JG van Sloun, and Yonina C Eldar. 2019. Deep unfolded robust PCA with application to clutter suppression in ultrasound. IEEE transactions on medical imaging 39, 4 (2019), 1051–1063. [52] Long Tan, Weijie Yuan, Xiaoqi Zhang, Kecheng Zhang, Zhongjie Li, and Yonghui Li. 2024. DNN-Based Radar Target Detection With OTFS. IEEE Transactions on Vehicular Technology (2024).

MobiCom ’26, October 26–30, 2026, Austin, TX, USA [53] Texas Instruments. 2025. DCA1000EVM Evaluation Module. https: //www.ti.com/tool/DCA1000EVM. Accessed: June 2, 2025. [54] Texas Instruments. 2025. IWR1843BOOST Evaluation Module. https: //www.ti.com/tool/IWR1843BOOST. Accessed: June 2, 2025. [55] Texas Instruments. 2025. TI mmWave Labs - Vital Signs Measurement (version 1.2). Texas Instruments. Accessed: Sep. 2, 2025. Available at: https://e2e.ti.com/cfs-file/__key/communityserver-discussionscomponents-files/1023/vitalSigns_5F00_lab_5F00_user_5F00_guide _5F00_v1.2UPDATE.pdf. [56] HL Van Trees et al. 2002. Optimum array processing. [57] GV Truck. 1977. Range resolution of targets using automatic detectors. NASA STI/Recon Technical Report N 78 (1977), 21348. [58] Harry L Van Trees. 2004. Detection, estimation, and modulation theory, part I: detection, estimation, and linear modulation theory. John Wiley & Sons. [59] Fengyu Wang, Feng Zhang, Chenshu Wu, Beibei Wang, and KJ Ray Liu. 2020. ViMo: Multiperson vital sign monitoring using commodity millimeter-wave radio. IEEE Internet of Things Journal 8, 3 (2020), 1294–1307. [60] M Weiss. 2007. Analysis of some modified cell-averaging CFAR processors in multiple-target situations. IEEE Trans. Aerospace Electron. Systems 1 (2007), 102–114. [61] John Wright and Yi Ma. 2022. High-dimensional data analysis with low-dimensional models: Principles, computation, and applications. Cambridge University Press. [62] Kailun Wu, Yiwen Guo, Ziang Li, and Changshui Zhang. 2020. Sparse coding with gated learned ISTA. In International conference on learning representations. [63] Dawei Yan, Panlong Yang, Fei Shang, Feiyu Han, Yubo Yan, and XiangYang Li. 2025. Pushing the Limits of WiFi-Based Gait Recognition Towards Non-Gait Human Behaviors . IEEE Transactions on Mobile Computing 01 (Feb. 2025), 1–17. https://doi.org/10.1109/TMC.2025.3 540863 [64] Zheng Yang, Yi Zhang, Kun Qian, and Chenshu Wu. 2023. {SLNet}: A Spectrogram Learning Neural Network for Deep Wireless Sensing. In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23). 1221–1236. [65] Shuochao Yao, Ailing Piao, Wenjun Jiang, Yiran Zhao, Huajie Shao, Shengzhong Liu, Dongxin Liu, Jinyang Li, Tianshi Wang, Shaohan Hu, et al. 2019. Stfnets: Learning sensing signals from the time-frequency perspective with short-time fourier neural networks. In The World Wide Web Conference. 2192–2202. [66] Luca Jiang-Tao Yu, Running Zhao, Sijie Ji, Edith CH Ngai, and Chenshu Wu. 2025. USpeech: Ultrasound-Enhanced Speech with Minimal Human Effort via Cross-Modal Synthesis. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 9, 2 (2025), 1–31. [67] Feng Zhang, Chenshu Wu, Beibei Wang, Min Wu, Daniel Bugos, Hangfang Zhang, and KJ Ray Liu. 2019. SMARS: Sleep monitoring via ambient radio signals. IEEE Transactions on Mobile Computing 20, 1 (2019), 217–231. [68] Haoyun Zhang, Jianghong Han, Xueqian Wang, Gang Li, and XiaoPing Zhang. 2025. Parameter Convergence Detector Based on VAMP Deep Unfolding: A Novel Radar Constant False Alarm Rate Detection Algorithm. arXiv preprint arXiv:2504.09912 (2025). [69] Haopeng Zhang, Yili Ren, Haohan Yuan, Jingzhe Zhang, and Yitong Shen. 2025. Wi-chat: Large language model powered wi-fi sensing. arXiv preprint arXiv:2502.12421 (2025). [70] Xie Zhang, Yina Wang, and Chenshu Wu. 2025. Unlocking Interpretability for RF Sensing: A Complex-Valued White-Box Transformer. arXiv preprint arXiv:2507.21799 (2025).

MobiCom ’26, October 26–30, 2026, Austin, TX, USA [71] Yuwei Zhang, Kumar Ayush, Siyuan Qiao, A Ali Heydari, Girish Narayanswamy, Maxwell A Xu, Ahmed A Metwally, Shawn Xu, Jake Garrison, Xuhai Xu, et al. 2025. SensorLM: Learning the Language of Wearable Sensors. arXiv preprint arXiv:2506.09108 (2025). [72] Yi Zhang, Weiying Hou, Zheng Yang, and Chenshu Wu. 2023. {VeCare}: Statistical acoustic sensing for automotive {In-Cabin} monitoring. In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23). 1185–1200. [73] Mingmin Zhao, Yonglong Tian, Hang Zhao, Mohammad Abu Alsheikh, Tianhong Li, Rumen Hristov, Zachary Kabelac, Dina Katabi, and Antonio Torralba. 2018. RF-based 3D skeletons. In Proceedings of the 2018 Conference of the ACM Special Interest Group on Data Communication. 267–281. [74] Peijun Zhao, Chris Xiaoxuan Lu, Bing Wang, Niki Trigoni, and Andrew Markham. 2023. Cubelearn: End-to-end learning for human motion recognition from raw mmwave radar signals. IEEE Internet of Things Journal 10, 12 (2023), 10236–10249. [75] Running Zhao, Jiangtao Yu, Hang Zhao, and Edith CH Ngai. 2023. Radio2text: Streaming speech recognition using mmwave radio signals. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 7, 3 (2023), 1–28. [76] Running Zhao, Luca Jiang-Tao Yu, Tingle Li, Zhihan Jiang, Chenwei Zhang, Chenshu Wu, Hang Zhao, and Edith CH Ngai. 2025. SPACE: Speaker Adaptation for Acoustic Eavesdropping using mmWave Radio Signals. IEEE Transactions on Mobile Computing (2025). [77] Xiaopeng Zhao, Zhenlin An, Qingrui Pan, and Lei Yang. 2023. Nerf2: Neural radio-frequency radiance fields. In Proceedings of the 29th Annual International Conference on Mobile Computing and Networking. 1–15. [78] Tianyue Zheng, Zhe Chen, Shujie Zhang, Chao Cai, and Jun Luo. 2021. MoRe-Fi: Motion-robust and fine-grained respiration monitoring via deep-learning UWB radar. In Proceedings of the 19th ACM conference on embedded networked sensor systems. 111–124. [79] Yue Zheng, Yi Zhang, Kun Qian, Guidong Zhang, Yunhao Liu, Chenshu Wu, and Zheng Yang. 2019. Zero-effort cross-domain gesture recognition with Wi-Fi. In Proceedings of the 17th annual international conference on mobile systems, applications, and services. 313–325.

Yu and Wu

Record · ID 10322 · SHA-256 1959bb4b44637e61
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.