Conceptio › Archive › arXiv CS
arXiv CSopen access

StarBOA: Real-Time Mamba State-Space Unrolling for Sparse Radar Micro-Doppler in ISAC Networks

Mustafa Bora Çelik et al. · arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
distributed-systemsinternetnetworkingprotocols
networking, internet, protocols, distributed systems

PREPRINT. UNDER REVIEW.

1

StarBOA: Real-Time Mamba State-Space Unrolling for Sparse Radar Micro-Doppler in ISAC Networks

arXiv:2609.33408v1 [cs.LG] 27 Sep 2026

Mustafa Bora Çelik, Ceren Çelik, and Orhan Gazi

Abstract—In Integrated Sensing and Communications (ISAC), radar sensing must operate under chirp subsampling with up to 90% missing data. An attention-based baseline, limited to a 52 ms buffer, collapses toward maximum uniform entropy (H = 2.584 bits) as sparsity increases, failing to capture long-range gait-cycle context. We propose StarBOA, which replaces attention with a causal Mamba state-space model that updates incrementally on a per-window basis without re-scanning past reconstructions. By maintaining a persistent state, StarBOA integrates over 100× more temporal history at no additional per-step computational cost. StarBOA outperforms the baseline’s published results across all sparsity levels, with SSIM gains increasing from +0.0379 at 50% missing data to +0.2472 at 90%. Each window is processed in 1.53 ms with zero lookahead, demonstrating efficient causal reconstruction under extreme chirp subsampling. Index Terms—ISAC, Micro-Doppler Radar, Compressive Sensing, Algorithm Unrolling, State-Space Models, Mamba, Real-Time Edge Processing.

I. I NTRODUCTION

I

SAC is a foundational technology for next-generation 6G cellular systems, merging high-rate wireless data transmission with radar sensing over shared spectrum and hardware platforms [1]–[3]. To maximize communication data throughput, base stations surrender up to 90% of transmission timefrequency resource blocks to downlink data packets. As a consequence, radar sensing is confined to scarce, intermittently gathered chirp pulses (10% duty cycle). As highlighted in recent compressive sensing radar studies [4], [5], reconstructing high-fidelity micro-Doppler (mD) spectrograms from such heavily undersampled measurements is vital: pulse dropouts create severe velocity aliasing and noise artifacts that obscure subtle kinematic Doppler signatures essential for human activity recognition (HAR). To reconstruct sparse radar returns, deep algorithm unrolling has emerged as a powerful paradigm, unrolling iterative compressive sensing solvers into neural network layers [6], [7]. In this context, the Single Thresholding with Attention Refinement (STAR) framework [4] combines Learned Iterative Hard Thresholding (LIHT) with temporal attention and solution refinement, showing promising reconstruction under mild subsampling (50% missing chirps). This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible. Corresponding author: Mustafa Bora Çelik (e-mail: [email protected]). Code and reproducible models are publicly available at https://github.com/BOA-clk/StarBOA.

A. Limitations of Attention Under Severe Sparsity Despite its effectiveness under mild subsampling, STAR’s non-parametric attention mechanism exhibits fundamental limitations as pulse missingness escalates. STAR computes correlation weights directly in the frequency domain without learnable projection matrices: αi [t] = Softmax( √1K y[t − i]T ỹ[t]) for i ∈ {1, . . . , Np }. Under severe 90% missingness, the candidate spectrum ỹ[t] is dominated by compressive sensing noise (SNR < 0 dB). Consequently, inner products with past recovered spectra lose selective discrimination, flattening attention weights across the historical buffer into an unweighted uniform moving average (αi [t] ≈ 1/Np = 1/6). Rather than retrieving coherent gait patterns, attention simply averages multiple corrupted spectrogram windows, diffusing Doppler energy and blurring subtle limb trajectories into indistinct veils. Moreover, its Np = 6 buffer spans only ≈ 52 ms (6 × 8.64 ms)—far short of a full gait cycle (≈ 1.0 s). B. Contributions: Continuous-Time State-Space Dynamics To address this challenge, we introduce StarBOA (Fig. 1): • Mamba-Based Recurrent Context Modeling: We replace STAR’s non-parametric attention mechanism with a selective continuous-time state space [8]. Rather than re-scanning a fixed buffer, StarBOA carries a recurrent state with constant (O(1)) per-step update cost regardless of history length, integrating context across 128 windows in our experiments—extendable further—versus STAR’s fixed 6-window buffer. • Root-Cause Entropy Collapse Analysis: We reveal that under 90% chirp missingness, STAR’s dot-product attention collapses to maximum uniform entropy (H = 2.5840 bits ≈ log2 (6)), explaining why attention degrades to moving-average blur while recurrent state-space models remain structurally resilient against instantaneous noise. • Decoupled Front-End Benchmark Superiority: On the public DISC 60 GHz benchmark, in decoupled evaluation with a frozen LIHT front-end, StarBOA outperforms STAR-Attention across all sparsity tiers (50%, 75%, 90%), lifting kinematic structure correlation above s > 0.888 even at extreme sparsity. • Full Pipeline Verification Across All Sparsity Tiers: Against STAR’s officially reported results, StarBOA surpasses it at every tested sparsity level, with the SSIM gain widening from +0.0379 at 50% to +0.2472 at 90% sparsity. • Incremental Streaming Without History Reprocessing: Unlike attention, which re-compares each new window against a stored frame buffer, StarBOA’s recurrent state

PREPRINT. UNDER REVIEW.

2

updates in constant time with no history reprocessing. Combined with a 10% sensing duty cycle, this sustains real-time ISAC streaming (1.53 ms/window, zero lookahead) while freeing 90% of time-frequency resources for communications.

H[t-1]

State H[t]

Memory

O(1) Recurrence

Because StarBOA directly builds upon the unrolled architecture introduced in [4], we briefly summarize STAR’s three cascaded blocks. A. Channel Model and Incomplete Sampling Under ISAC slot sharing, Channel Impulse Response (CIR) samples are gathered in slow-time windows of K chirps with step shift δ. Due to packet transmission patterns, only a subset Mt ≪ K of samples is gathered at window t. Let h[t] ∈ CMt denote the incomplete measurement vector:

s[t]

Identity Skip Connection

Output

Block B: Mamba Conv1D

II. P RELIMINARIES : T HE STAR F RAMEWORK

to [t+1]

SSM Core

u[t] ≡ ỹ[t]

Block C

90% Missing

MLP

SiLU(·)

Fig. 1. Causal StarBOA architecture and recurrent state memory propagation across slow-time radar windows. Rather than matching against a noisy historical buffer, the compact O(1) recurrent state Ht continuously accumulates sequential context across severe 90% missing chirp bursts, enabling realtime streaming reconstruction with strictly 0.00 ms lookahead.

(1)

where U, V ∈ RK×K and b ∈ RK are learnable weights.

where Mt ∈ RMt ×K selects the sampled indices, FK is the inverse Fourier dictionary, z[t] ∈ CK is the unknown sparse Doppler spectrum, and n[t] is noise [3], [4].

Preserved Components: We retain Blocks A and C unchanged, replacing only the vulnerable non-parametric attention (Block B) with continuous-time selective state-space modeling.

h[t] = Mt FK z[t] + n[t],

B. The Three-Block STAR Pipeline

III. P ROPOSED S TAR BOA A RCHITECTURE

The Single Thresholding with Attention Refinement (STAR) framework [4] recovers sequential Doppler spectra through three cascaded processing blocks: 1) Block A (Single-Layer LIHT Module): To avoid multiiteration solver delay [9], STAR unrolls a single Iterative Hard Thresholding step into a learnable physical layer:   1 T (0) W h[t] , (2) z [t] = HΩ µ    1 1 z[t] = HΩ I − WT W z(0) [t] + WT h[t] , µ µ (3)

A. Sequential State-Space Context Modeling (Block B) As illustrated in Fig. 1, StarBOA replaces STAR’s dotproduct attention mechanism (Block B) with a continuoustime selective state-space model [8]. Instead of computing non-parametric affinity across a limited buffer of noisy frames, Mamba assimilates continuous historical information to predict and reconstruct the missing components in newly arriving slow-time windows. For each feature channel d ∈ {1, . . . , D} (D = 64), the hidden state hd (t) ∈ RN (N = 16) evolves continuously according to:

where HΩ preserves the Ω largest components, µ = 20 is ḣd (t) = Ad hd (t) + Bd,t xconv,d (t), the inverse step size, and W ∈ R2Mt ×2K is a dictionary (6) yd (t) = Cd,t hd (t) + Dd xconv,d (t), initialized as W = R(Mt FK ) via real-complex mapping R(·) [4]. The coarse spectrum is ỹ[t] = [IK IK ]z[t]2 ∈ where A ∈ RD×N is the diagonal state transition matrix with RK . N 2) Block B (Attention Mechanism): To capture tempo- Ad,n < 0 ensuring numerical stability, and Bd,t , Cd,t ∈ R ral correlation, STAR compares candidate ỹ[t] against are input-dependent selective control vectors. Discretization and Selective Recurrence: At each slow-time Np = 6 past recovered spectra Y[t] = [y[t−1], . . . , y[t− D Np ]]T ∈ RNp ×K using non-parametric scaled dot- window t, the coarse normalized spectrum ut ≡ ỹ[t] ∈ R product attention (without learnable projection matrices) (D = K = 64 Doppler bins) produced by Block A is projected into two parallel branches: an SSM branch xt and to extract a context feature vector a[t]: a gating branch zt . Input xt is processed by a depthwise   1 a[t] = Y[t]T Softmax √ Y[t]ỹ[t] ∈ RK . (4) causal 1D convolution (dconv = 4) with SiLU activation: K xconv,t = SiLU(Conv1D(xt )), buffering the prior 3 steps in D 3) Block C (Solution Refinement Module): The coarse streaming mode. The input-dependent timescale ∆t ∈ R and N spectrum ỹ[t] is refined based on the context feature selective vectors Bt , Ct ∈ R are dynamically generated from −4 vector a[t] via additive transformation and multiplicative xconv,t , with ∆t = clip(Softplus(W∆ xconv,t + b∆ ), 10 , 0.1). Discretizing via first-order Euler approximation yields: gating: y[t] = (ỹ[t] + ReLU(Ua[t] + b)) ⊙ σ(Va[t]),

(5)

Āt = exp (∆t A) ,

B̄t = ∆t Bt .

(7)

PREPRINT. UNDER REVIEW.

3

TABLE I D ECOUPLED F RONT-E ND B ENCHMARK ON F ROZEN LIHT U NDER 50%, 75%, AND 90% S PARSITY ON THE DISC DATASET. SSIM ↑

Spar. Model

s

l

c

RMSE ↓

Horizon H

Duration

SSIM ↑ Struct (s) RMSE ↓

H=2 H = 32 H = 128

17.3 ms 276.5 ms 1105.9 ms

0.6861 0.7560 0.7654

50%

STAR-Attn [4] 0.8595 0.9245 0.9461 0.9400 StarBOA (Ours) 0.8822 0.9394 0.9565 0.9531

0.0849 0.0746

75%

STAR-Attn [4] 0.8204 0.8997 0.9289 0.9240 StarBOA (Ours) 0.8536 0.9200 0.9427 0.9433

0.0950 0.0868

STAR-Attn [4] 0.7331 0.8661 0.8538 0.8504 0.1340 StarBOA (Ours) 0.8047 0.8889 0.9137 0.9210 0.1107 s: Structure, l: Luminance, c: Contrast components of SSIM [10]. 90%

mt = Wout

N X

! Ht,·,n Ct,n + D ⊙ xconv,t

0.8238 0.8557 0.8601

0.1457 0.1263 0.1233

B. Decoupled Front-End Benchmark Comparison

In online edge streaming, the recurrent state update is strictly causal (0.00 ms lookahead): Ht = Āt ⊙ Ht−1 + xconv,t ⊗ B̄t ,

TABLE II R ECONSTRUCTION F IDELITY VS . R ECURRENT M EMORY H ORIZON (H ) AT 90% C HIRP S PARSITY.

(8) ! ⊙ SiLU(zt ) .

n=1

(9) B. Residual Solution Refinement (Block C) In STAR [4], Block C required a separate multiplicative Sigmoid gate because attention provided no internal gating. In contrast, since Mamba already incorporates continuous SiLU gating within Block B, StarBOA implements Block C directly on top of the gated context mt via a zero-initialized residual Multi-Layer Perceptron (MLP):   ŝ[t] = clamp ỹ[t] + W2 GELU W1 LayerNorm(mt ) + b1  + b2 , 0, 1 , (10) where W2 , b2 are initialized to zero. This ensures that StarBOA begins training as an exact identity mapping over the physical LIHT reconstruction, incrementally learning kinematic trajectory corrections. C. Multi-Objective Loss Function The network is trained end-to-end using a composite reconstruction objective:

To measure the exact block-level contribution of the temporal refinement mechanism, we conduct a decoupled ceteris-paribus ablation where the physical LIHT front-end (Block A) is frozen, feeding identical coarse inputs to both refinement architectures. As documented in Table I, fidelity is quantified via RMSE and SSIM, decomposed as SSIM(x, y) = [l(x, y)] · [c(x, y)] · [s(x, y)] [10], where l(·), c(·), s(·) denote luminance, contrast, and structure (kinematic trajectory correlation). This decomposition ensures high fidelity reflects genuine gait-curve recovery rather than trivial energy scaling. Across every sparsity step, StarBOA delivers superior performance, improving SSIM by +2.64% at 50%, +4.05% at 75%, and +9.77% at 90% sparsity, while locking kinematic structure correlation above s > 0.888. C. Real-Time Streaming and Recurrent Memory Horizon Because StarBOA operates causally in online streaming mode, we evaluate its recovery fidelity as a function of the recurrent memory horizon H ∈ {2, . . . , 128} past windows. As shown in Table II, recovery precision ascends monotonically as the memory horizon expands, smoothly approaching steady-state performance as the context encompasses a full human gait cycle (H = 128, spanning 1.11 s = 128×8.64 ms). Crucially, total end-to-end inference latency per radar window is only 1.5255 ms (1525.5 µs) in pure software (730.6 µs LIHT, 670.0 µs Mamba step, and 124.9 µs refinement). Operating > 11× faster than the 17.28 ms physical window duration (and > 5× faster than the 8.64 ms arrival stride) with strictly 0.00 ms lookahead, StarBOA provides ample computational headroom for real-time streaming on edge ISAC processors.

L = ∥Ŝ − SGT ∥1 + 0.5∥Ŝ − SGT ∥2F + 0.5LSSIM (Ŝ, SGT ), (11)

D. Root-Cause Analysis: Attention Entropy Collapse vs. Recurrent State Memory

where LSSIM utilizes an accelerated 7 × 7 pool during backpropagation, while testing evaluates exact canonical SSIM over an 11 × 11 Gaussian kernel (σ = 1.5).

To uncover the structural mechanism governing performance under extreme subsampling, Table III(b) analyzes the historical context distribution under 90% chirp missingness. In STAR, candidate ỹ[t] is matched against Np = 6 past windows via dot products. Under 90% missingness, severe noise flattens attention weights across the 6 historical windows to near-identical values (α ≈ 0.172–0.175). Empirical evaluation reveals that attention entropy reaches H = 2.5840 bits, matching the theoretical maximum Huniform = log2 (6) = 2.5850 bits. This entropy collapse compounds STAR’s short 52 ms buffer limitation, explaining why attention degrades to movingaverage blur while recurrent state-space models remain robust across all sparsity tiers (Table III(a)). Why Recurrent State Spaces Prevent Collapse: Unlike attention—which re-evaluates a noisy instantaneous query ỹ[t]

IV. E XPERIMENTAL R ESULTS AND D ISCUSSION A. Experimental Setup We evaluate our framework on the public DISC mmWave radar dataset [4], [11], recorded using 60 GHz IEEE 802.11ay CIR transceivers across four human activities: WALKING, RUNNING, SITTING, and HANDS. Frames consist of K = 64 slow-time samples per window shifted by δ = 32 (Tc = 0.27 ms, physical window duration Tw = KTc = 17.28 ms, window stride Tstride = δTc = 8.64 ms). Compressive sensing is tested at 50%, 75%, and 90% missing chirp ratios under identical data splits, with all models trained for 20 epochs.

PREPRINT. UNDER REVIEW.

4

Fig. 2. Visual reconstruction under 90% extreme sparsity on the DISC mmWave benchmark (v ∈ [−4.5, 4.5] m/s, t ∈ [0, 4] s). (a) Ground Truth reference mmWave spectrogram. (b) Raw LIHT input under 90% pulse dropouts (severe noise and fragmented velocity profile). (c) STAR-Attention [4] (SSIM = 0.536, diffuse moving-average smearing and blurred limb arcs). (d) Proposed StarBOA (SSIM = 0.7832, sharp torso line, continuous limb oscillation arcs, and high-contrast gait trajectory).

TABLE III F ULL E ND - TO -E ND P IPELINE C OMPARISON AND STAR ATTENTION D ISTRIBUTION U NDER 90% S PARSITY. (a) End-to-End Pipeline Reconstruction Across Sparsity Tiers Sparsity Architecture

RMSE ↓ SSIM ↑

∆SSIM

50%

STAR [4] (reported) StarBOA (Ours)

0.0545 0.0506

0.8840 0.9219

Reference +0.0379

75%

STAR [4] (reported) StarBOA (Ours)

0.0779 0.0772

0.7450 0.8679

Reference +0.1229

90%

STAR [4] (reported) StarBOA (Ours)

0.1213 0.1211

0.5360 0.7832

Reference +0.2472

(b) STAR Attention Weights (90% Sparsity, Np = 6) Weights: αt−1 = 0.174, αt−2 = 0.173, αt−3 = 0.175, αt−4 = 0.174, αt−5 = 0.172, αt−6 = 0.174 Uniform: 1/6 ≈ 0.167 | Entropy: H = 2.5840 bits ≈ log2 (6)

against an external buffer of corrupted past estimates—StarBOA maintains an internal recurrent hidden state Ht ∈ RD×N . Because state transitions are governed by continuous recurrence and selective input projections (∆t , Bt ), the state does not depend on static window averaging. Instead, instantaneous noise bursts are naturally suppressed by the selective state transition, while coherent temporal trajectories are integrated across hundreds of slow-time windows (H ≥ 128). Consequently, StarBOA remains robust by construction under severe pulse missingness. E. Visual Spectrogram Inspection Fig. 2 visually corroborates Table I: StarBOA reconstructs a continuous torso trajectory and limb-oscillation arcs where STAR-Attention exhibits moving-average smearing and fragmented footfalls (per-panel SSIM in caption). F. Full End-to-End Pipeline Evaluation Table III(a) compares our full end-to-end StarBOA against STAR’s officially reported numbers [4]: the SSIM gain widens from +0.0379 at 50% to +0.1229 at 75% and +0.2472

at 90% sparsity, consistent with the compounding bufferspan and entropy-collapse mechanisms above (the controlled, same-environment ablation isolating this effect is given in Table I). This benefits ISAC design directly: replacing fragile attention with state-space unrolling lets base stations cut radar transmission to a 10% duty cycle without compromising HARcritical sensing fidelity. V. C ONCLUSION In this letter, we showed that non-parametric attention in unrolled radar recovery suffers from a short 52 ms buffer and collapses to maximum uniform entropy (H=2.584 bits≈ log2 6) under severe missingness. We proposed StarBOA, replacing attention with a causal continuous-time Mamba state space that captures context across a full 1.1 s gait cycle with O(1) memory. On the DISC mmWave benchmark, StarBOA consistently outperforms STAR across all sparsity tiers, with SSIM gains widening to +0.2472 at 90% missing chirps while executing in 1.53 ms with zero lookahead. This allows base stations to allocate 90% of resources to communications while sustaining high-fidelity micro-Doppler sensing. R EFERENCES [1] H. Wymeersch et al., “Integration of communication and sensing in 6G: A joint industrial and academic perspective,” in Proc. IEEE PIMRC, 2021, pp. 1–7. [2] F. Liu, C. Masouros, A. P. Petropulu, H. Griffiths, and L. Hanzo, “Joint radar and communication design: Applications, state-of-the-art, and the road ahead,” IEEE Trans. Commun., vol. 68, no. 6, pp. 3834–3862, 2020. [3] J. Pegoraro, J. O. Lacruz, M. Rossi, and J. Widmer, “SPARCS: A sparse recovery approach for integrated communication and human sensing in mmWave systems,” in Proc. ACM/IEEE IPSN, 2022, pp. 79–91. [4] R. Mazzieri, J. Pegoraro, and M. Rossi, “Attention-refined unrolling for sparse sequential micro-doppler reconstruction,” IEEE J. Sel. Topics Signal Process., vol. 18, no. 5, pp. 812–827, July 2024. [5] V. C. Chen, F. Li, S.-S. Ho, and H. Wechsler, “Micro-Doppler effect in radar: Phenomenon, model, and simulation study,” IEEE Trans. Aerosp. Electron. Syst., vol. 42, no. 1, pp. 2–21, 2006. [6] K. Gregor and Y. LeCun, “Learning fast approximations of sparse coding,” in Proc. ICML, 2010, pp. 399–406. [7] V. Monga, Y. Li, and Y. C. Eldar, “Algorithm unrolling: Interpretable, efficient deep learning for signal and image processing,” IEEE Signal Process. Mag., vol. 38, no. 2, pp. 18–44, 2021. [8] A. Gu and T. Dao, “Mamba: Linear-time sequence modeling with selective state spaces,” arXiv preprint arXiv:2312.00752, 2023.

PREPRINT. UNDER REVIEW.

[9] Y. C. Eldar and G. Kutyniok, Compressed Sensing: Theory and Applications. Cambridge Univ. Press, 2012. [10] Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: From error visibility to structural similarity,” IEEE Trans. Image Process., vol. 13, no. 4, pp. 600–612, 2004. [11] J. Pegoraro, J. O. Lacruz, M. Rossi, and J. Widmer, “DISC: A dataset for integrated sensing and communication in mmWave systems,” IEEE Dataport, 2022.

5

Record · ID 1108660 · SHA-256 3e6de870967a4caf
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.