C O RF: C ROSS -S CENE RF S YNTHESIS BY L EARNING P ROPAGATION AND P RESERVING A RRAY P HYSICS Kang Yang1 1 2
Duaa Nakshbandi2
Mani Srivastava1∗
Department of Electrical and Computer Engineering, University of California, Los Angeles Department of Computer Science and Engineering, University of California, Merced
arXiv:2609.34399v1 [cs.NI] 28 Sep 2026
Wan Du2
{duaanakshbandi,wdu3}@ucmerced.edu
A BSTRACT Existing radio-frequency (RF) neural fields fit each scene separately, making newscene deployment measurement- and optimization-intensive. This work studies amortized cross-scene spatial spectrum synthesis, where a shared model learns propagation across scenes and instantiates an unseen scene from sparse target-scene measurements without scene-specific training. To achieve this, CoRF separates learned scene-dependent propagation from the known receiver-array observation model. An unordered set of spectrum-only references conditions a canonical anchor field, producing arrival directions and query-dependent component powers for each query. Analytic array physics maps these components to the receiver covariance and then to the spatial spectrum. This factorization keeps the pretrained propagation model frozen, enabling synthesis at arbitrary query locations after a single reference-conditioning pass. A necessary local reference-capacity bound and an error decomposition further characterize the formulation. Across 35 simulated scenes spanning seven categories, CoRF outperforms the strongest baseline by 7.09 dB PSNR on unseen variants of represented scene categories and by 6.97 dB on entirely unseen scene categories.
1
I NTRODUCTION
Spatial spectrum synthesis predicts the angular distribution of received signal power at a receiver antenna array for transmitters at unmeasured locations (Krim & Viberg, 1996; Zhao et al., 2023). This directional information supports wireless communication and sensing tasks including beam management, direction-of-arrival estimation, localization, and environment-aware communication (Krim & Viberg, 1996; Schmidt, 1986; Guo et al., 2026; Wang et al., 2026; Zeng et al., 2024). Recent radio-frequency (RF) neural fields and Gaussian representations synthesize spectra at unmeasured locations (Zhao et al., 2023; Lu et al., 2024; Wen et al., 2025; Zhang et al., 2026a; Yang et al., 2025), but fit each scene separately. Deploying these methods in a new scene therefore requires a large target-scene measurement set and a new scene-specific optimization. In contrast, we study amortized cross-scene synthesis, where a shared model learns a propagation prior across scenes and instantiates an unseen scene from sparse target-scene spectra without scene-specific training. GRaF (Yang et al., 2026) is the closest prior work under this input setting, but synthesizes each query from geographically nearby measurements and therefore depends on local reference coverage. Cross-scene RF synthesis raises two challenges. I. Reference correspondence. Generalizable rendering in vision establishes query-reference correspondences by projecting query rays into reference views using scene geometry (Yu et al., 2021; Wang et al., 2021; Charatan et al., 2024; Chen et al., 2024b). RF spatial spectra lack such direct correspondence because each spectrum superposes propagation paths whose directions and powers change with the transmitter location (Zhao et al., 2023; Lu et al., 2024; Orekondy et al., 2023). II. Propagation and array entanglement. Each measured spectrum combines scene-dependent propagation with the receiver-array response (Krim & Viberg, 1996; Schmidt, 1986; Stoica et al., 2011). Propagation varies with scene geometry and material properties (Orekondy et al., 2023; Hoydis et al., 2023; Lu et al., 2024), whereas the observation ∗ Mani Srivastava holds concurrent appointments as a Professor of ECE and CS (joint) at the University of California, Los Angeles.
1
operator is determined by the receiver configuration. Jointly learning both factors expends model capacity on reproducing known array physics rather than modeling cross-scene propagation. We introduce CoRF for cross-scene RF spatial spectrum synthesis from sparse measurements. CoRF is trained across source scenes to learn scene-dependent propagation while retaining the receiver-array response as an analytic operator. Scene propagation is represented by canonical anchors, which are persistent representation slots at fixed coordinates in a normalized volume. At deployment, an unordered set of target-scene spectra, termed references, conditions these anchors once to instantiate the propagation representation of an unseen scene. The resulting anchor field is reused across query locations to predict arrival directions, mean component powers, and power variances. An analytic decoder constructs the receiver covariance from these components and then renders the spatial spectrum through the Bartlett response (Krim & Viberg, 1996), with the power variances providing a query-level reliability estimate. Thus, feed-forward CoRF operates in an unseen scene through reference conditioning rather than scene-specific parameter fitting, keeping the pretrained propagation model frozen. The formulation is characterized by a necessary local reference-capacity bound and a propagation-to-spectrum error bound linking component and renderer errors to spectrum error. Beyond feed-forward CoRF, we develop two optional deployment extensions. CoRFPR refines only the mean component powers using the target-scene reference set, while CoRFCAL calibrates only the analytic receiver response and keeps the learned propagation model fixed. Cross-scene generalization is evaluated on 35 simulated scenes generated with Sionna RT (Hoydis et al., 2023), spanning seven categories with five variants each. Using M = 32 references, feed-forward CoRF outperforms the strongest baseline by 7.09 dB PSNR on unseen variants of categories represented during training and by 6.97 dB on categories excluded entirely from training. Our contributions are as follows: • We formulate cross-scene spatial spectrum synthesis as amortized propagation inference from an unordered set of spectrum-only target-scene references. • We introduce CoRF, which conditions canonical anchors on sparse references to infer scenedependent propagation and renders these components using analytic receiver-array physics. • We demonstrate cross-scene generalization to unseen scene variants and entirely unseen scene categories using sparse target-scene references.
2
R ELATED W ORK
Spatial Spectrum Synthesis. Existing RF neural fields and Gaussian representations reconstruct propagation from scene-specific measurements and typically require separate optimization for each scene (Zhao et al., 2023; Lu et al., 2024; Orekondy et al., 2023; Wen et al., 2025; 2026; Zhang et al., 2026a; Yang et al., 2025; Liu et al., 2026; Zhang et al., 2025; Nukapotula et al., 2025; Huang et al., 2024). Related work explores receiver generalization, sparse differentiable fitting, physics-informed generalizable architectures, and transfer using explicit scene geometry (Chen et al., 2024a; Bian et al., 2025; Shen et al., 2026; Zhang et al., 2026b). Geometry-conditioned methods such as RadTwin require a point cloud or mesh (Zhang et al., 2026b), whereas our setting uses measured spectra when such a scene model is unavailable. GRaF (Yang et al., 2026) is the closest prior work with the same input modality because it performs feed-forward synthesis from reference measurements. However, each query relies on geographically nearby spectra and therefore requires local reference coverage. In contrast, CoRF constructs a reusable scene representation from a sparse reference set collected once for the scene and applies it across query locations. Feed-Forward Reconstruction in Vision. Generalizable neural rendering predicts new views by conditioning on reference observations (Yu et al., 2021; Chen et al., 2021; Wang et al., 2021; Mukund Varma T et al., 2023; Sajjadi et al., 2022; Charatan et al., 2024; Chen et al., 2024b; Wang et al., 2025; Jin et al., 2025). Many of these methods exploit geometric structure and cross-view correspondence between reference and query views. RF spatial spectra lack comparable correspondence because their multipath directions and powers vary across transmitter locations (Zhao et al., 2023; Lu et al., 2024; Orekondy et al., 2023). Permutation-invariant set models and conditional neural processes provide mechanisms for aggregating unordered reference observations (Zaheer et al., 2017; Lee et al., 2019; Garnelo et al., 2018; Kim et al., 2019; Nguyen & Grover, 2022). Accordingly, CoRF conditions a canonical anchor field on unordered RF reference spectra to form a shared scene representation. Analytic Array Processing. Classical array processing recovers angular structure from observed measurements through beamforming, subspace, sparse, and Bayesian methods (Krim & Viberg, 1996; 2
Cross-Scene Propagation Field Reference representation
𝑺𝝅 𝟏
Analytic Array Physics Reference-toanchor projection
⋮ 𝑺𝝅 𝑴
𝑧̃$ 𝑧̃# ⋯ 𝑧̃" ⋯ 𝑧̃!
M spectra
Reference tokens
𝛺& , 𝑚& , 𝜈&
Covariance synthesis 𝓡'
Propagation estimation
Bartlett rendering 𝑆,( Ω = 𝔤 Ω a 𝛺 ) 𝓡( a Ω
Canonical anchor field
Query configuration 𝜋⋆
Spectrum
Uncertainty
Figure 1: Framework of CoRF. Unordered spectrum-only references condition canonical anchors to predict scene-dependent propagation, which analytic array physics renders as the spatial spectrum. Stoica et al., 2011; Yang et al., 2013). Neural methods further support direction estimation and spectral reconstruction (Shmuel et al., 2025; Ji et al., 2024; Merkofer et al., 2024; Berman et al., 2025). CoRF instead predicts propagation for unmeasured transmitter locations and renders the spatial spectrum using the analytic Bartlett response. Keeping the receiver response explicit allows CoRF to support calibration without retraining propagation (Pan et al., 2023). Learned compensation can handle broader hardware mismatch but requires receiver-specific data (Liu et al., 2018).
3
M ETHOD
Problem Setting. We consider a fixed receiver array at position trx ∈ R3 with orientation Q ∈ SO(3). 3 tx rx For a transmitter at ttx i ∈ R , the measurement configuration is πi = (ti ; t , Q). At each H×W configuration, the receiver reports a spatial spectrum Sπi ∈ R+ on an H × W angular grid. Let Itr denote the set of training scenes, where each scene provides measurements πj , Sπj j . For M
each training instance, we sample a scene s ∈ Itr , an unordered reference set Cs = {Sπi }i=1 , and a query configuration π⋆ outside the reference set. A single model fθ is trained across scenes: h i Sbπ⋆ = fθ (Cs , π⋆ ) , min Es∈Itr ECs ,π⋆ ℓspec Sbπ⋆ , Sπ⋆ . (1) θ
At test time, M references from an unseen scene are processed once to construct a reusable scene representation for synthesizing spectra at arbitrary query transmitter locations. Overview. As shown in Figure 1, CoRF first conditions a shared canonical anchor field on the reference spectra to construct the scene representation. Given a query configuration, the conditioned K field predicts K propagation components {(Ωn , mn , vn )}n=1 in §3.1. Analytic array physics then converts these components into the spatial spectrum and uncertainty in §3.2. The formulation is analyzed in §3.3, and training and inference are described in §3.4. 3.1 C ROSS -S CENE P ROPAGATION F IELD The cross-scene propagation field maps the reference spectra and query configuration to queryK specific propagation parameters {(Ωn , mn , vn )}n=1 . It first constructs a scene representation from the unordered references and then conditions this representation on the query to estimate propagation. (I) Reference Representation. Because the reference measurements form an unordered set, their representation should be invariant to input ordering. Each reference spectrum is encoded independently, and the resulting tokens are then contextualized across the full reference set: M M zi = Eϕ (Sπi ) , {e zi }i=1 = Tψ {zi }i=1 . (2) Here, Eϕ is a 2D CNN, while Tψ is a Transformer encoder without positional encoding and is therefore permutation-equivariant over the reference set (Lee et al., 2019). The contextualized tokens capture each reference in relation to the full set. K
(II) Reference-to-Anchor Projection. We define K anchors {xn }n=1 ⊂ R3 at fixed coordinates in a normalized canonical volume, providing shared spatial support across scenes. After reference conditioning, each anchor predicts a position offset, an occupancy weight, and an emission feature that parameterize a candidate propagation component. The conditioning combines global context from the reference set with direction-specific evidence aligned with each anchor. 3
Reference-token conditioning. Let η (xn ) denote the positional embedding of anchor n. Each anchor cross-attends to the contextualized reference tokens, and the attended feature modulates its position-dependent representation: M (γn , βn ) = MLP CrossAttn η (xn ) , {e zi }i=1 , hn = (1 + γn ) ⊙ h0n + βn . (3) Here, h0n = ReLU (W0 η (xn )) is the anchor representation before reference conditioning. The resulting hn captures global reference context associated with anchor n. Bearing-aligned sampling. Because each reference token compresses a full spatial spectrum, we additionally preserve direction-specific evidence from its angular feature map Fi . For anchor n, we sample each feature map along its bearing Ω̄n in the receiver array frame: "M # rx X T xn − t Ω̄n = ∠ Q , bn = wn,i Wv Fi Ω̄n ; max Fi Ω̄n . (4) i ∥xn − trx ∥ i=1 Here, wn,i are attention weights over the references and Wv is a value projection. The attention pool aggregates evidence across references, while the element-wise maximum retains the strongest K response along the anchor bearing. Together, {(hn , bn )}n=1 form the scene-conditioned canonical anchor field used for query-specific propagation estimation. (III) Propagation Estimation. The canonical anchor field is query independent and captures reusable scene structure. For each query, we decode the anchors into refined geometry and query-conditioned propagation parameters. This separation allows the anchor geometry to represent scene structure, while the propagation strength adapts to the query transmitter location. Anchor decoding. We first fuse the global and bearing-aligned features: ∆xn , αn , e0n = Dω (hn + Wb bn ) , x′n = xn + ∆xn .
(5)
Here, Dω is a shared anchor decoder, ∆xn is a learned position offset that refines the canonical anchor location, αn is its occupancy logit, and e0n is its query-independent base emission. Propagation parameters. The refined anchor position x′n determines the final arrival direction Ωn using the same receiver-frame bearing transformation as Equation 4, with xn replaced by x′n . In contrast, the component power depends on the query-anchor geometry, allowing the same anchor rx rx T ′ rx field to support different transmitter locations. Let q⋆ = QT (ttx ⋆ − t ) and xn = Q (xn − t ) denote the query and anchor positions in the receiver array frame. A query-conditioned FiLM head (Perez et al., 2018) modulates the base emission: (γn⋆ , βn⋆ ) = Fq (q⋆ , xrx n ),
en = (1 + γn⋆ ) ⊙ e0n + βn⋆ .
(6)
Using the zeroth-order, direction-independent component of en , denoted by en ∈ C, we compute the mean power and power variance as m2n , (7) κn where κn > 0 is predicted by an uncertainty head. The resulting propagation parameK ters {(Ωn , mn , vn )}n=1 are passed to the analytic array model. 2
mn = sigmoid (αn ) |en | ,
vn =
3.2 A NALYTIC A RRAY P HYSICS K Given the estimated propagation parameters {(Ωn , mn , vn )}n=1 , the analytic array model first synthesizes the receiver covariance, then renders the spatial spectrum through the Bartlett response, and finally propagates component power uncertainty to the spectrum. (I) Covariance Synthesis. For an N -element receiver array, let a (Ω) ∈ CN denote the steering vector at arrival direction Ω. Under the narrowband and far-field model (Schmidt, 1986), a component H with direction Ωn and mean power mn contributes the rank-one covariance mn a (Ωn ) a (Ωn ) . Assuming mutually incoherent components, their covariance contributions add without cross terms. With an explicit line-of-sight (LOS) component, the query covariance is Rπ⋆ =
K X
H
H
mn a (Ωn ) a (Ωn ) + mlos a (Ωlos ) a (Ωlos ) .
n=1
4
(8)
Here, Ωlos is the arrival direction of the direct LOS path. The term mlos provides a shared baseline for its normalized power. Because each target spectrum is normalized independently, this term does not represent absolute received power. The conditioned non-LOS components model the remaining variation across scenes and query locations. (II) Bartlett Rendering. The covariance is converted to a spatial spectrum by evaluating the Bartlett response over the angular grid (Krim & Viberg, 1996): q H Seπ⋆ (Ω) = g (Ω) a (Ω) Rπ⋆ a (Ω). (9) Here, g (Ω) denotes the receiver element power pattern, with g ≡ 1 for the simulated arrays. Simulated targets are generated from coherent single snapshots, whereas the decoder in Equation 8 sums mutually incoherent covariance components without path cross terms. Accordingly, the renderer is a structured approximation under the narrowband and far-field assumptions of classical array processing (Krim & Viberg, 1996; Schmidt, 1986). The representation-mismatch term in Proposition 2 captures the residual error under perfect propagation estimates. (III) Component-Power Uncertainty. The analytic renderer also propagates uncertainty in the component powers to the spatial spectrum. Because Bartlett power is linear in the component powers, their variances propagate through the fixed array response in closed form without stochastic sampling. Let pn ≥ 0 be the random power of component n, with E [pn ] = mn and Var [pn ] = vn . H
2
For an = a (Ωn ), define the unit-power array response χn (Ω) = a (Ω) an . With uncorrelated component powers and fixed arrival directions, the random Bartlett power and variance are Pπrand (Ω) = g (Ω) ⋆
K X
pn χn (Ω) ,
K X 2 2 Var Pπrand (Ω) = g (Ω) vn χn (Ω) . ⋆
n=0
(10)
n=0
The deterministic LOS component is included as n = 0, with p0 = mlos and v0 = 0, so the mean 2 Bartlett power equals Seπ⋆ (Ω) . Because the synthesized spectrum is the square root of the Bartlett power, we propagate its variance through the square root using a first-order approximation: rand Var Pπrand (Ω) ⋆ Var Sπ⋆ (Ω) ≈ . (11) 2 4Seπ (Ω) ⋆
This quantity conditions on fixed arrival directions and analytically propagates component-power uncertainty into a query-level reliability signal. The full derivation is provided in Appendix A. 3.3 T HEORETICAL A NALYSIS Two bounds characterize the information supplied by reference spectra and the effect of propagation errors on rendering. For reference i, Ji measures scene sensitivity after projecting out pose effects. Proposition 1 (Reference Capacity). Let r⋆ be the locally observable scene dimension, M ⋆ the minimum number of references to span it, and N the number of array elements. Then, r⋆ ⋆ rank Ji ≤ N − 1, M ≥ . (12) N −1 This necessary local bound shows why complementary references matter, although it does not determine the reference count needed in practice. The derivation is provided in Appendix B. Proposition 2 (Propagation-to-Spectrum Error). Let P and Pb be the target and rendered power spectra, εphys the best representable spectrum error, and Em and EΩ the matched power and power-weighted direction errors. If the array response is bounded and Lipschitz, then Pb − P
∞
≤ εphys + Cm Em + CΩ EΩ ,
(13)
where Cm and CΩ depend on the receiver array. This bound separates renderer mismatch from errors in predicted component powers and arrival directions. The derivation is provided in Appendix C. 5
3.4
T RAINING AND I NFERENCE
Training. We pretrain the spectrum encoder Eϕ and reference contextualizer Tψ with supervised contrastive learning over scene labels, encouraging consistent representations for references from the same scene. With both modules frozen, we train the reference-to-anchor projection and propagation estimation modules to predict spectra at query locations outside the reference set. The loss combines spectrum reconstruction, reference-use regularization, and heteroscedastic uncertainty learning, as detailed in Appendix D.1. Model architecture is described in Appendix D.2. Inference. For an unseen scene, CoRF conditions the canonical anchors on its reference set once and reuses the resulting field across queries. For each query configuration π⋆ , it preK dicts {(Ωn , mn , vn )}n=1 and analytically renders the spatial spectrum and uncertainty. Unless stated otherwise, CoRF denotes this feed-forward setting. Power Refinement. We also study CoRFPR , a lightweight, power-only target-scene adaptation that K fits a correction to the mean component powers {mn }n=1 . It uses the same M references without additional measurements and leaves the pretrained encoder and propagation field frozen. Arrival directions remain fixed during refinement. The correction is fitted once per target scene and applied to the predicted powers at each query, as detailed in Appendix D.3. Receiver Calibration. Array geometry may come from hardware specifications, and antenna metadata from AISG Antenna Location and Orientation Sensor (ALS) (AISG, 2024; 3GPP, 2024). If the resulting analytic response is accurate, calibration is unnecessary. Otherwise, CoRFCAL fits a correction to the nominal response using NC additional spectra, disjoint from the M references and test queries, while keeping the propagation model and scene representation fixed. The fitted response is reused across queries; details appear in Appendix D.4.
4
E VALUATION
Dataset. We use 35 simulated scenes generated with Sionna RT (Hoydis et al., 2023) and one measured scene from NeRF2 (Zhao et al., 2023). The simulated scenes span seven categories with five variants each: conference room, classroom, bedroom, corridor, open office, laboratory, and lounge. Each variant contains thousands of transmitter measurements. All scenes use a fixed 4 × 4 receiver array with λ/2 spacing and produce 90 × 360 spatial spectra at 920 MHz. Further dataset details are provided in Appendix D.5. Protocol. Each scene uses a 70/10/20 train, validation, and test split. Cross-scene evaluations use M = 32 target-scene references from the training split and queries from the test split. (I) For per-scene evaluation, each model is trained and tested within the same scene variant. (II) For generalization within represented scene categories, five folds exclude one variant per category. (III) For generalization to excluded scene categories, seven folds exclude an entire category. Baselines. We compare with GRaF (Yang et al., 2026), the closest generalizable RF field. GRaF is trained on the same source scenes as CoRF and receives the same M references from the target scene. For each query, its L nearest neighbors are selected from these M references. We also compare with three per-scene RF fields: NeRF2 (Zhao et al., 2023), GSRF (Yang et al., 2025), and WRF-GS (Wen et al., 2025) using its deformable WRF-GS+ variant (Wen et al., 2026). For each baseline X, X-FS trains from scratch on the M target-scene references. X-TR adapts a model trained on one source scene using the same references, with results averaged over all eligible source scenes. CoRF Variants. CoRF denotes the feed-forward version in the results. Power refinement, receiver calibration, and their combination are labeled CoRFPR , CoRFCAL , and CoRFPR,CAL , respectively. Metrics. We report peak signal-to-noise ratio (PSNR)↑, structural similarity index (SSIM)↑, normalized mean squared error (NMSE)↓, and angle-of-arrival (AoA) error↓. AoA error measures the angular distance between the peak directions of the predicted and target spectra. 4.1
P ER -S CENE P ERFORMANCE
We first evaluate the conventional per-scene setting, where Table 1: Per-scene performance. each method is trained and tested within the same scene Method PSNR↑ SSIM↑ NMSE↓ AoA↓ variant. For this comparison, CoRFPR uses the power- NeRF2 21.57 ±3.57 0.698 ±0.104 0.130 ±0.086 11.02 ±7.81 20.45 ±2.43 0.677 ±0.066 0.151 ±0.070 11.58 ±6.90 refinement protocol. Across 36 scenes, CoRFPR achieves GSRF WRF-GS 21.93 ±3.06 0.705 ±0.081 0.117 ±0.064 9.89 ±6.67 GRaF
21.96 ±3.11 0.683 ±0.083 0.113 ±0.060 8.91 ±6.48
CoRFPR 24.27 ±3.38 0.791 ±0.071 0.083 ±0.052 7.68 ±5.82
6
Table 2: Generalization to excluded scene vari- Table 3: Generalization to excluded scene cateants. Mean ± standard deviation over 5 folds. gories. Mean ± standard deviation over 7 folds. Method
PSNR↑
SSIM↑
NMSE↓
AoA↓
Method
PSNR↑
SSIM↑
NMSE↓
AoA↓
NeRF2 -FS NeRF2 -TR
13.74 ±0.72 13.00 ±0.61
0.414 ±0.032 0.382 ±0.025
0.636 ±0.091 0.734 ±0.101
24.65 ±7.10 31.20 ±7.27
NeRF2 -FS NeRF2 -TR
13.75 ±0.63 12.98 ±1.15
0.414 ±0.027 0.379 ±0.056
0.635 ±0.078 0.745 ±0.139
24.65 ±7.04 32.87 ±10.89
GSRF-FS GSRF-TR
12.37 ±0.24 12.71 ±0.71
0.335 ±0.004 0.413 ±0.026
0.717 ±0.023 0.766 ±0.127
51.40 ±3.96 43.68 ±8.73
GSRF-FS GSRF-TR
12.39 ±0.63 12.77 ±0.47
0.335 ±0.020 0.413 ±0.015
0.715 ±0.094 0.772 ±0.100
51.29 ±6.91 39.73 ±3.82
WRF-GS-FS WRF-GS-TR
15.86 ±0.83 13.49 ±1.03
0.495 ±0.034 0.448 ±0.027
0.398 ±0.053 0.624 ±0.138
22.23 ±5.52 34.08 ±9.23
WRF-GS-FS WRF-GS-TR
15.86 ±0.81 12.75 ±1.12
0.495 ±0.030 0.426 ±0.031
0.398 ±0.043 0.740 ±0.202
22.25 ±5.19 35.04 ±5.13
GRaF
13.86 ±0.38
0.441 ±0.020
0.578 ±0.044
28.32 ±5.63
GRaF
13.67 ±0.47
0.435 ±0.025
0.607 ±0.093
29.14 ±4.86
CoRFPR
23.02 ±2.18
0.749 ±0.055
0.119 ±0.054
8.81 ±4.15
CoRFPR
22.89 ±2.72
0.743 ±0.071
0.122 ±0.062
8.89 ±5.43
the best average performance on all four metrics in Table 1. Relative to the strongest baseline for each metric, PSNR improves from 21.96 to 24.27 dB and SSIM from 0.705 to 0.791. NMSE decreases from 0.113 to 0.083, and AoA error from 8.91◦ to 7.68◦ . Together, these gains show improved spectral fidelity and angular accuracy with component-based propagation estimation and analytic array rendering. Qualitative examples and per-scene breakdowns appear in Appendices E.1.1 and E.2.1. 4.2 G ENERALIZATION TO U NSEEN VARIANTS OF K NOWN S CENE C ATEGORIES We evaluate five folds, each excluding one scene variant per category; all methods use the same M = 32 references from each test scene. As shown in Table 2, CoRFPR achieves 23.02 dB PSNR, 0.749 SSIM, 0.119 NMSE, and 8.81◦ AoA error. Compared with WRF-GS-FS, the strongest baseline on all four metrics, it improves PSNR by 7.16 dB and SSIM by 0.254, while reducing NMSE by 0.279 and AoA error by 13.42◦ . Transfer adaptation is inconsistent: NeRF2 -TR and WRF-GS-TR perform worse than training from scratch, while GSRF-TR improves only some metrics. Sparse target-scene references may be insufficient to adapt a representation fitted to another scene. GRaF selects nearby references for each query, so sparse random sampling may leave some queries without local coverage. This comparison therefore measures performance under a matched sparse-reference budget, not under denser local sampling. In contrast, CoRFPR conditions a reusable anchor field on the full reference set and refines component powers using those same measurements. Its gains in spectral and AoA metrics show improved spectrum synthesis and dominant-direction accuracy on unseen variants. Qualitative examples and scene-level breakdowns appear in Appendices E.1.2 and E.2.2, respectively. 4.3 G ENERALIZATION TO U NSEEN S CENE C ATEGORIES We further evaluate category-level generalization in seven folds, each excluding all five variants of one category from training. All methods use the same M = 32 references from each test scene. As shown in Table 3, CoRFPR achieves 22.89 dB PSNR, 0.743 SSIM, 0.122 NMSE, and 8.89◦ AoA error. Compared with WRF-GS-FS, the strongest baseline on all four metrics, it improves PSNR by 7.03 dB and SSIM by 0.248, while reducing NMSE by 0.276 and AoA error by 13.36◦ . Under this sparse reference budget, GRaF selects nearby measurements for each query, whereas CoRFPR conditions a reusable anchor field on the full set and applies a target-scene power correction at each query. Performance remains close to the unseen-variant setting: PSNR changes from 23.02 to 22.89 dB, SSIM from 0.749 to 0.743, NMSE from 0.119 to 0.122, and AoA error from 8.81◦ to 8.89◦ . These results support transfer to excluded categories within the same simulation pipeline, with target-scene references supplying scene-specific information. Qualitative examples and scene-level breakdowns appear in Appendices E.1.3 and E.2.3, respectively.
7
SSIM
NMSE
4.4 C OMPONENT-P OWER R ELIABILITY CoRF analytically propagates component-power variance 0.9 NMSE SSIM to the spatial spectrum while holding arrival directions fixed. 0.3 For each query, the reliability score Uπ is the mean prop- 0.2 0.7 agated variance over the H × W angular grid. Because variance scales differ across folds, we group test queries 0.1 into within-fold quintiles of Uπ before aggregating results 0.5 over the 12 generalization folds. As shown in Figure 2, 0.0 Q1 Q2 Q3 Q4 Q5 NMSE increases and SSIM decreases monotonically across Predicted uncertainty these quintiles. Across 68,232 test queries, mean NMSE Figure 2: Reliability analysis. rises from 0.050 in the lowest-uncertainty quintile to 0.223 in the highest, a 4.5× increase, while mean SSIM falls from 0.828 to 0.630. Mean within-fold Spearman correlations are ρ = 0.53 for NMSE and ρ = −0.55 for SSIM; the NMSE correlation is
positive in every fold, ranging from 0.42 to 0.65. Therefore, propagated power variance provides a useful query-level ranking of synthesis reliability, though it is not calibrated pixel-level predictive uncertainty. These results use feed-forward CoRF; the corresponding power-refinement analysis, qualitative examples, and further details appear in Appendix E.3. 4.5
R ECEIVER A RRAY A DAPTATION
Separating propagation from array physics allows CoRF Table 4: Adaptation to unseen arrays. to update the analytic response for a new receiver without Array Array model PSNR↑ SSIM↑ NMSE↓ retraining the propagation model. When the receiver con- A (3 × 3) Training array 15.64 0.528 0.190 Calibrated 20.20 0.751 0.095 figuration, including its geometry and pose, is accurately Training array 14.96 0.458 0.327 known, this update requires no calibration fit. We test three A (8 × 2) Calibrated 22.09 0.697 0.097 arrays unseen during training on the seven held-out scene A (4 × 4 narrow) Training array 16.47 0.553 0.179 Calibrated 22.60 0.795 0.073 variants of one fold: A1 , a 3 × 3 array at the training spacing; A2 , an 8 × 2 array at the same spacing; and A3 , a 4 × 4 array with spacing reduced from 0.16 m to 0.12 m. Transmitter locations, data splits, and reference sets remain unchanged. The analytic response uses array geometry, wavelength, position, and orientation from receiver specifications or deployment metadata (AISG, 2024; 3GPP, 2024; 2026; 2019). 1
2
3
To evaluate imperfect specifications, we start from a perturbed array response and fit an in-plane rotation and two axis scalings using measured spectra. The pretrained neural modules remain frozen. This optional calibration uses NC = 64 measurements per scene, separate from the M = 32 references and test queries. In Table 4, retaining the training-array response yields 0.458–0.553 SSIM, whereas calibration reaches 0.697–0.795 SSIM and improves PSNR by 4.6–7.1 dB. Because the imposed errors match the three-parameter calibration model, these results demonstrate correction of the tested specification errors, not robustness to arbitrary receiver mismatch. Calibration procedures and parameter-error sensitivity are detailed in Appendix E.4. 4.6
A BLATION S TUDY
Table 5 examines three design choices of feed-forward CoRF under both generalization settings. Each variant is retrained on the same data with the same objective. Table 5: Ablation study. Each variant changes one design choice of CoRF. Unseen Scene Variants
Unseen Scene Categories
Variant
PSNR↑ SSIM↑ NMSE↓ AoA↓ PSNR↑ SSIM↑ NMSE↓ AoA↓
Learned renderer − bearing-aligned sampling − query conditioning
16.09 21.07 20.75
0.544 0.697 0.683
0.369 0.159 0.224
16.27 9.95 13.57
16.00 20.79 18.32
0.542 0.709 0.620
0.396 0.149 0.291
15.42 9.67 17.21
CoRF
22.95
0.740
0.121
8.40
22.83
0.733
0.124
8.41
Analytic Array Physics. Replacing the analytic covariance and Bartlett renderer with a learned CNN produces the largest degradation. For unseen variants, PSNR falls from 22.95 to 16.09 dB, SSIM from 0.740 to 0.544, and AoA error rises from 8.40◦ to 16.27◦ ; the same pattern holds for unseen categories. The analytic renderer applies the specified array response to each predicted component, whereas the CNN must learn this mapping from training examples. The comparison therefore supports retaining the known array mapping while learning scene-dependent propagation. Bearing-Aligned Sampling. Removing bearing-aligned sampling reduces PSNR by 1.88 and 2.04 dB in the two settings and increases AoA error by 1.55◦ and 1.26◦ , respectively. SSIM and NMSE also worsen in both settings. These results suggest that set-level reference aggregation captures broad scene context, while bearing-aligned sampling supplies complementary directional evidence to the anchors and improves query-specific propagation estimates. Query Conditioning. Removing query conditioning reduces PSNR by 2.20 dB for unseen variants and 4.51 dB for unseen categories. For unseen categories, SSIM also falls from 0.733 to 0.620, NMSE rises from 0.124 to 0.291, and AoA error increases from 8.41◦ to 17.21◦ . These losses exceed those from removing bearing-aligned sampling, particularly when the target category is excluded from training. The result highlights the importance of specializing the shared scene representation to each query transmitter location when estimating propagation components. 8
4.7
C ASE S TUDY: B EAM M ANAGEMENT WITH S YNTHESIZED S PATIAL S PECTRA
Beam management aims to find a high-power beam from a codebook, but probing every candidate incurs measurement overhead (Xue et al., 2024; Alkhateeb et al., 2017). For each query location, CoRF uses its synthesized spatial spectrum to rank the candidate beams. We first evaluate the loss from probing only the highest-ranked beams. We then test whether probing more beams for less reliable predictions reduces loss at the same average search budget. B
Spectrum-Based Beam Search. For a B = 64-beam codebook B = {Ωb }b=1 , spectrum-based search selects the Kb beams with the highest synthesized values Sbπ (Ωb ), rather than searching the b π,K denote the selected set. Using target spatial power Pπ (Ω) = Sπ (Ω)2 as full codebook. Let B b the beam-power proxy, we measure loss relative to exhaustive search: ! maxΩb ∈B Pπ (Ωb ) Lπ,Kb = 10 log10 . (14) maxΩb ∈B b π,K Pπ (Ωb ) b
A loss of 0 dB means the selected set contains a beam maximizing the target spatial power. We evaluate Kb ∈ {1, 2, 4, 8, 16} with the same codebook, queries, and reference budget for all methods. Reliability-Aware Beam Search. A fixed Kb assigns every query the same search budget despite differences in synthesis reliability. We rank queries within each fold by Uπ from §4.4 and divide them into five groups. Groups with higher propagated variance receive more beams while preserving a target average budget K b . The allocation is chosen on validation data and fixed for testing, so fixed and adaptive policies use the same average number of beam probes.
Beam loss (dB)
In Figure 3, CoRF achieves 0.93 dB beam loss with one GRaF GSRF Ours selected beam, compared with 5.02 dB for the strongest WRFGS Ours + UQ NeRF2 baseline. It selects a beam maximizing target spatial 10 power in 63.7% of queries, versus 21.8% for that base3 line. With Kb = 4, CoRF reaches 0.46 dB, comparable to the baseline’s 0.44 dB at Kb = 16, using one quarter 1 as many beams. For the unseen-variant generalization 0.3 setting, reliability-aware allocation further reduces beam 0.1 loss from 0.271 to 0.205 dB at K b = 8 and from 0.111 to 0.072 dB at K b = 16. For unseen scene categories, val1 2 4 8 16 Searched beams Kb idation selects a nearly uniform allocation and the adaptive gain disappears. Therefore, spectrum synthesis reduces the Figure 3: Beam management study. beam budget needed for low loss under this proxy, while reliability-aware allocation helps when Uπ meaningfully ranks query difficulty. A localization case study using the same synthesized spectra appears in Appendix E.5. 4.8
G ENERALIZATION TO R EAL M EASUREMENTS
We compare CoRF trained on 35 simulated scenes with Table 6: Simulation-to-real transfer on baselines on measured scene S36 under a matched budget the measured-scene dataset (S36). of 96 target-domain spectra. CoRFCAL and CoRFPR,CAL Method PSNR↑ SSIM↑ NMSE↓ AoA↓ use M = 32 references and NC = 64 disjoint calibraNeRF2 12.94 0.457 0.287 16.24 tion spectra; CoRFPR,CAL reuses both for power refine- GSRF 16.85 0.587 0.145 12.09 16.96 0.592 0.140 9.26 ment. Calibration improves PSNR by 1.95 dB over uncali- WRF-GS+ 14.24 0.510 0.219 7.87 brated CoRF, and refinement adds 0.71 dB without further GRaF 16.74 0.620 0.125 8.12 measurements. In Table 6, CoRFPR,CAL leads in PSNR, CoRFCAL SSIM, and NMSE, while GRaF has a marginally lower CoRFPR,CAL 17.45 0.653 0.118 7.88 AoA error. The smaller PSNR margin partly reflects WRF-GS+ fitting S36 directly, whereas CoRF transfers a simulation-trained propagation prior. Held-out simulations share the material library and propagation pipeline; S36 may have unfamiliar reflected and diffuse paths and receiver mismatch. Sparse references may recover dominant directions but miss weaker paths, reducing spectral fidelity while leaving AoA nearly unchanged. Calibration partly corrects receiver mismatch and refinement adjusts powers, but neither recovers missing directions. Further analysis appears in Appendix E.6. Additional experiments examine reference evidence (§E.7), representation and renderer choices (§E.8), deployment efficiency (§E.9), and material-property shifts (§E.10). 9
5
D ISCUSSION AND C ONCLUSION
We introduce CoRF for cross-scene RF spatial spectrum synthesis from sparse target-scene measurements. Unordered references condition a canonical anchor field whose predicted propagation components are rendered through analytic receiver-array physics. Feed-forward CoRF requires no scene-specific optimization, while CoRFPR optionally fits a correction limited to component powers using the same reference set. Across 35 simulated scenes, CoRFPR exceeds the strongest baseline by about 7 dB PSNR on both unseen scene variants and excluded categories. Three limitations remain. First, the performance gain narrows under simulation-to-real transfer, where physical receiver responses can differ from the nominal array model. Extending analytic calibration beyond geometry to account for element-wise gain and phase errors may reduce this gap. Second, the method relies on spatial-spectrum references; extending it to other measurement modalities, such as received signal strength indicator (RSSI) or channel state information (CSI), may require modality-specific representations and reference-acquisition strategies. Third, CoRF assumes static scenes; temporal representations and online reference updates could support dynamic scenes.
ACKNOWLEDGMENTS The research reported in this paper was sponsored in part by the DEVCOM Army Research Laboratory (award # W911NF1720196) and the National Science Foundation (awards # 2325956 and 2525614). Wan Du and Duaa Nakshbandi were funded in part by the National Science Foundation through grant CCSS-2525613. The views and conclusions contained in this document are those of the authors and should not be interpreted as representing the official policies, either expressed or implied, of the funding agencies.
R EFERENCES 3GPP. Study on Channel Model for Frequencies from 0.5 to 100 GHz. Technical Report 38.901, V16.1.0 (Release 16), 3rd Generation Partnership Project, December 2019. URL https://www. 3gpp.org/ftp/Specs/archive/38_series/38.901/38901-g10.zip. 3GPP. NG-RAN; NR Positioning Protocol a (NRPPa). Technical Specification 38.455, V18.4.0 (Release 18), 3rd Generation Partnership Project, December 2024. URL https://www.3gpp. org/ftp/Specs/archive/38_series/38.455/38455-i40.zip. 3GPP. LTE Positioning Protocol (LPP). Technical Specification 37.355, V19.3.0 (Release 19), 3rd Generation Partnership Project, June 2026. URL https://www.3gpp.org/ftp/Specs/ archive/37_series/37.355/37355-j30.zip. AISG. Antenna Location and Orientation Sensor. Subunit Type Standard AISG-ST-ALS, vALS3.0.2.2, Antenna Interface Standards Group, June 2024. URL https://aisg.org.uk/file/ AISG-ALS-vALS3.0.2.pdf. Ahmed Alkhateeb, Young-Han Nam, Md Saifur Rahman, Jianzhong Zhang, and Robert W. Heath. Initial Beam Association in Millimeter Wave Cellular Systems: Analysis and Design Insights. IEEE Transactions on Wireless Communications, 16(5):2807–2821, 2017. Lioz Berman, Sharon Gannot, and Tom Tirer. (SP)2 -Net: A Neural Spatial Spectrum Method for DOA Estimation. arXiv preprint arXiv:2509.15475, 2025. Kejia Bian, Meixia Tao, Shu Sun, Tongjia Zhang, and Jun Yu. GeNeRT: A Physics-Informed Approach to Intelligent Wireless Channel Modeling via Generalizable Neural Ray Tracing. arXiv preprint arXiv:2506.18295, 2025. David Charatan, Sizhe Lester Li, Andrea Tagliasacchi, and Vincent Sitzmann. pixelSplat: 3D Gaussian Splats from Image Pairs for Scalable Generalizable 3D Reconstruction. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. Anpei Chen, Zexiang Xu, Fuqiang Zhao, Xiaoshuai Zhang, Fanbo Xiang, Jingyi Yu, and Hao Su. MVSNeRF: Fast Generalizable Radiance Field Reconstruction from Multi-View Stereo. In IEEE/CVF International Conference on Computer Vision (ICCV), 2021. 10
Xingyu Chen, Zihao Feng, Ke Sun, Kun Qian, and Xinyu Zhang. RFCanvas: Modeling RF Channel by Fusing Visual Priors and Few-Shot RF Measurements. In ACM International Conference on Embedded Networked Sensor Systems (SenSys), 2024a. Yuedong Chen, Haofei Xu, Chuanxia Zheng, Bohan Zhuang, Marc Pollefeys, Andreas Geiger, Tat-Jen Cham, and Jianfei Cai. MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View Images. In European Conference on Computer Vision (ECCV), 2024b. Marta Garnelo, Dan Rosenbaum, Chris J. Maddison, Tiago Ramalho, David Saxton, Murray Shanahan, Yee Whye Teh, Danilo J. Rezende, and S. M. Ali Eslami. Conditional Neural Processes. In International Conference on Machine Learning (ICML), 2018. Keqiang Guo, Yuheng Zhong, Xin Tong, Jiangbin Lyu, and Rui Zhang. Neural Beam Field for Spatial Beam RSRP Prediction. In IEEE Wireless Communications and Networking Conference (WCNC), 2026. Jakob Hoydis, Fayçal Aït Aoudia, Sebastian Cammerer, Merlin Nimier-David, Nikolaus Binder, Guillermo Marcus, and Alexander Keller. Sionna RT: Differentiable Ray Tracing for Radio Propagation Modeling. In IEEE Globecom Workshops (GC Wkshps), 2023. Tianshu Huang, John Miller, Akarsh Prabhakara, Tao Jin, Tarana Laroia, Zico Kolter, and Anthony Rowe. DART: Implicit Doppler Tomography for Radar Novel View Synthesis. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. ITU-R. Effects of Building Materials and Structures on Radiowave Propagation above about 100 MHz. Recommendation ITU-R P.2040-1, International Telecommunication Union, July 2015. URL https://www.itu.int/rec/R-REC-P.2040-1-201507-S/en. ITU-R. Effects of Building Materials and Structures on Radiowave Propagation above about 100 MHz. Recommendation ITU-R P.2040-3, International Telecommunication Union, August 2023. URL https://www.itu.int/rec/R-REC-P.2040-3-202308-S/en. Junkai Ji, Wei Mao, Feng Xi, and Shengyao Chen. TransMUSIC: A Transformer-Aided Subspace Method for DOA Estimation with Low-Resolution ADCs. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2024. Haian Jin, Hanwen Jiang, Hao Tan, Kai Zhang, Sai Bi, Tianyuan Zhang, Fujun Luan, Noah Snavely, and Zexiang Xu. LVSM: A Large View Synthesis Model with Minimal 3D Inductive Bias. In International Conference on Learning Representations (ICLR), 2025. Hyunjik Kim, Andriy Mnih, Jonathan Schwarz, Marta Garnelo, Ali Eslami, Dan Rosenbaum, Oriol Vinyals, and Yee Whye Teh. Attentive Neural Processes. In International Conference on Learning Representations (ICLR), 2019. Hamid Krim and Mats Viberg. Two Decades of Array Signal Processing Research: The Parametric Approach. IEEE Signal Processing Magazine, 13(4):67–94, 1996. Juho Lee, Yoonho Lee, Jungtaek Kim, Adam R. Kosiorek, Seungjin Choi, and Yee Whye Teh. Set Transformer: A Framework for Attention-Based Permutation-Invariant Neural Networks. In International Conference on Machine Learning (ICML), 2019. Chenhe Li, Qiang Xu, Zhe Gong, and Rong Zheng. TuRF: Fast Data Collection for Fingerprint-Based Indoor Localization. In International Conference on Indoor Positioning and Indoor Navigation (IPIN), 2017. Mufan Liu, Cixiao Zhang, Qi Yang, Yujie Cao, Yiling Xu, Yin Xu, Shu Sun, Mingzeng Dai, and Yunfeng Guan. Deformable 2D Gaussian Splatting for Efficient Wireless Radiance Field Rendering. IEEE Transactions on Visualization and Computer Graphics, 32(7):6363–6378, 2026. Zhang-Meng Liu, Chenwei Zhang, and Philip S. Yu. Direction-Of-Arrival Estimation Based on Deep Neural Networks with Robustness to Array Imperfections. IEEE Transactions on Antennas and Propagation, 66(12):7315–7327, 2018. 11
Haofan Lu, Christopher Vattheuer, Baharan Mirzasoleiman, and Omid Abari. NeWRF: A Deep Learning Framework for Wireless Radiation Field Reconstruction and Channel Prediction. In International Conference on Machine Learning (ICML), 2024. Julian P. Merkofer, Guy Revach, Nir Shlezinger, Tirza Routtenberg, and Ruud J. G. van Sloun. DA-MUSIC: Data-Driven DoA Estimation via Deep Augmented MUSIC Algorithm. IEEE Transactions on Vehicular Technology, 73(2):2771–2785, 2024. Mukund Varma T, Peihao Wang, Xuxi Chen, Tianlong Chen, Subhashini Venugopalan, and Zhangyang Wang. Is Attention All That NeRF Needs? In International Conference on Learning Representations (ICLR), 2023. Tung Nguyen and Aditya Grover. Transformer Neural Processes: Uncertainty-Aware Meta Learning via Sequence Modeling. In International Conference on Machine Learning (ICML), 2022. Bhavya Sai Nukapotula, Rishabh Tripathi, Seth Pregler, Dileep Kalathil, Srinivas Shakkottai, and Theodore S. Rappaport. GSpaRC: Gaussian Splatting for Real-Time Reconstruction of RF Channels. arXiv preprint arXiv:2511.22793, 2025. Tribhuvanesh Orekondy, Kumar Pratik, Shreya Kadambi, Hao Ye, Joseph Soriaga, and Arash Behboodi. WiNeRT: Towards Neural Ray Tracing for Wireless Channel Modelling and Differentiable Simulations. In International Conference on Learning Representations (ICLR), 2023. Mengguan Pan, Shengheng Liu, Peng Liu, Wangdong Qi, Yongming Huang, Wang Zheng, Qihui Wu, and Markus Gardill. In Situ Calibration of Antenna Arrays for Positioning with 5G Networks. IEEE Transactions on Microwave Theory and Techniques, 71(10):4600–4613, 2023. Ethan Perez, Florian Strub, Harm de Vries, Vincent Dumoulin, and Aaron Courville. FiLM: Visual Reasoning with a General Conditioning Layer. In AAAI Conference on Artificial Intelligence (AAAI), 2018. Mehdi S. M. Sajjadi, Henning Meyer, Etienne Pot, Urs Bergmann, Klaus Greff, Noha Radwan, Suhani Vora, Mario Lučić, Daniel Duckworth, Alexey Dosovitskiy, Jakob Uszkoreit, Thomas Funkhouser, and Andrea Tagliasacchi. Scene Representation Transformer: Geometry-Free Novel View Synthesis through Set-Latent Scene Representations. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022. Ralph O. Schmidt. Multiple Emitter Location and Signal Parameter Estimation. IEEE Transactions on Antennas and Propagation, 34(3):276–280, 1986. Jingzhou Shen, Luis Lago Enamorado, Shiwen Mao, and Xuyu Wang. A Geometric AlgebraInformed NeRF Framework for Generalizable Wireless Channel Prediction. In IEEE Conference on Computer Communications (INFOCOM), 2026. Dor H. Shmuel, Julian P. Merkofer, Guy Revach, Ruud J. G. van Sloun, and Nir Shlezinger. SubspaceNet: Deep Learning-Aided Subspace Methods for DoA Estimation. IEEE Transactions on Vehicular Technology, 74(3):4962–4976, 2025. Petre Stoica, Prabhu Babu, and Jian Li. SPICE: A Sparse Covariance-Based Estimation Method for Array Processing. IEEE Transactions on Signal Processing, 59(2):629–638, 2011. Jianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi, Christian Rupprecht, and David Novotny. VGGT: Visual Geometry Grounded Transformer. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025. Qianqian Wang, Zhicheng Wang, Kyle Genova, Pratul Srinivasan, Howard Zhou, Jonathan T. Barron, Ricardo Martin-Brualla, Noah Snavely, and Thomas Funkhouser. IBRNet: Learning Multi-View Image-Based Rendering. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021. Xiucheng Wang, Junxi Huang, and Nan Cheng. RadioDiff-v2: Generative Angular Radio Maps for Multi-Beam Selection and Localization. arXiv preprint arXiv:2607.08045, 2026. 12
Chaozheng Wen, Jingwen Tong, Yingdong Hu, Zehong Lin, and Jun Zhang. WRF-GS: Wireless Radiation Field Reconstruction with 3D Gaussian Splatting. In IEEE Conference on Computer Communications (INFOCOM), 2025. Chaozheng Wen, Jingwen Tong, Yingdong Hu, Zehong Lin, and Jun Zhang. Neural Representation for Wireless Radiation Field Reconstruction: A 3D Gaussian Splatting Approach. IEEE Transactions on Wireless Communications, 25:7490–7504, 2026. Qing Xue, Chengwang Ji, Shaodan Ma, Jiajia Guo, Yongjun Xu, Qianbin Chen, and Wei Zhang. A Survey of Beam Management for mmWave and THz Communications towards 6G. IEEE Communications Surveys and Tutorials, 26(3):1520–1559, 2024. Kang Yang, Gaofeng Dong, Sijie Ji, Wan Du, and Mani Srivastava. GSRF: Complex-Valued 3D Gaussian Splatting for Efficient Radio-Frequency Data Synthesis. In Conference on Neural Information Processing Systems (NeurIPS), 2025. Kang Yang, Yuning Chen, and Wan Du. Generalizable Radio-Frequency Radiance Fields for Spatial Spectrum Synthesis. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026. Zai Yang, Lihua Xie, and Cishen Zhang. Off-Grid Direction of Arrival Estimation Using Sparse Bayesian Inference. IEEE Transactions on Signal Processing, 61(1):38–43, 2013. Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa. pixelNeRF: Neural Radiance Fields from One or Few Images. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021. Manzil Zaheer, Satwik Kottur, Siamak Ravanbakhsh, Barnabás Póczos, Ruslan Salakhutdinov, and Alexander Smola. Deep Sets. In Conference on Neural Information Processing Systems (NeurIPS), 2017. Yong Zeng, Junting Chen, Jie Xu, Di Wu, Xiaoli Xu, Shi Jin, Xiqi Gao, David Gesbert, Shuguang Cui, and Rui Zhang. A Tutorial on Environment-Aware Communications via Channel Knowledge Map for 6G. IEEE Communications Surveys and Tutorials, 26(3):1478–1519, 2024. Lihao Zhang, Zongtan Li, and Haijian Sun. RF-PGS: Fully-Structured Spatial Wireless Channel Representation with Planar Gaussian Splatting. arXiv preprint arXiv:2508.16849, 2025. Lihao Zhang, Haijian Sun, Samuel Berweger, Camillo Gentile, and Rose Qingyang Hu. RF-3DGS: Wireless Channel Modeling with Radio Radiance Field and 3D Gaussian Splatting. IEEE Transactions on Wireless Communications, 25:10419–10433, 2026a. Yuru Zhang, Ming Zhao, Qiang Liu, Ahmed Alkhateeb, Abhishek K. Agrawal, and Qi Qu. RadTwin: Generalizable Wireless Digital Twin for Dynamic Environments. In International Conference on Computer Communications and Networks (ICCCN), 2026b. Xiaopeng Zhao, Zhenlin An, Qingrui Pan, and Lei Yang. NeRF2 : Neural Radio-Frequency Radiance Fields. In Annual International Conference on Mobile Computing and Networking (MobiCom), 2023.
13
A PPENDIX C ONTENTS A Analytic Component Power Uncertainty
15
B Observability from Reference Spectra
17
C Propagation-to-Spectrum Error Bound
19
D Implementation Details
22
D.1 Training Objectives and Hyperparameters . . . . . . . . . . . . . . . . . . . . . . . . . . . .
22
D.2 Model Architecture . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
24
D.3 Target-Scene Power Refinement . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
25
D.4 Receiver Array Calibration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
26
D.5 Dataset and Preprocessing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
28
E Additional Experimental Results and Analyses
29
E.1 Qualitative Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
29
E.2 Scene-Level Spatial Spectrum Synthesis Results . . . . . . . . . . . . . . . . . . . . . . . .
31
E.3 Reliability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
35
E.4 Tolerance to Array-Parameter Error . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
37
E.5 Case Study: Angular Localization with Synthesized Spatial Spectra . . . . . . . . . . . . . .
38
E.6 Simulation-to-Real Transfer Study . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
39
E.7 Role of Target-Scene References . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
40
E.8 Representation and Renderer Diagnostics . . . . . . . . . . . . . . . . . . . . . . . . . . . .
44
E.9 Deployment Efficiency . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
46
E.10 Physical Distribution Shift . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
47
14
A
A NALYTIC C OMPONENT P OWER U NCERTAINTY
We derive how uncertainty in the predicted component powers propagates through the analytic decoder to the reported spatial spectrum. We first exploit the linearity of Bartlett power in the component powers to obtain its mean and variance exactly. We then analyze the square root transformation, provide a rigorous bound on the resulting magnitude uncertainty, and derive the first order delta approximation used by CoRF. Finally, we propagate this approximation through the fixed normalization used by the model and obtain the reported variance proxy after global calibration. Exact Power Domain Moments. Let an = a (Ωn ) ,
n = 1, . . . , K.
The predicted arrival directions are treated as deterministic. We include the LOS component as n = 0, with a0 = a (Ωlos ) , m0 = mlos , v0 = 0. Let pn ≥ 0 almost surely denote the random power of component n, with E [pn ] = mn ,
Var [pn ] = vn < ∞.
We assume that the component powers are pairwise uncorrelated. The deterministic LOS contribution is represented by p0 = mlos almost surely. Define
2
H
χn (Ω) = a (Ω) an . The corresponding random Bartlett power is Pπrand (Ω) = g (Ω)
K X
pn χn (Ω) .
(15)
n=0
Since g (Ω) ≥ 0, pn ≥ 0, and χn (Ω) ≥ 0, we have Pπrand (Ω) ≥ 0 almost surely. Proposition 3 (Exact Power Domain Moments). The mean and variance of equation 15 are K X µP (Ω) := E Pπrand (Ω) = g (Ω) mn χn (Ω) , n=0 K X rand 2 2 2 := σP (Ω) Var Pπ (Ω) = g (Ω) vn χn (Ω) .
(16)
n=0
Proof. The first identity in equation 16 follows directly from linearity of expectation and E [pn ] = mn . Since the component powers are pairwise uncorrelated, "K # K X X 2 Var pn χn (Ω) = vn χn (Ω) , n=0
n=0
which gives the second identity in equation 16. The deterministic decoder substitutes the mean powers mn into the same analytic renderer. Therefore, p Seπ (Ω) = µP (Ω). Thus, introducing the component power uncertainty leaves the deterministic point prediction unchanged. Magnitude Domain Uncertainty. Define Sπrand (Ω) =
q
Pπrand (Ω). √ The deterministic prediction µP is generally not equal to E Sπrand . 15
Lemma 1 (Magnitude Uncertainty). At any angular location with µP > 0, √ 2 E Sπrand ≤ µP , Var Sπrand = µP − E Sπrand , and
σ2 Var Sπrand ≤ P . µP
(17)
Proof. The first inequality follows from the concavity of the square root and Jensen’s inequality. 2 Since Sπrand = Pπrand , 2 Var Sπrand = E Pπrand − E Sπrand . For the variance bound, Var Sπrand ≤ E
"q
Pπrand −
2 # µP
2 Pπrand − µP
= E p ≤
√
Pπrand +
√
µP
2
σP2 , µP
which gives equation 17. If µP (Ω) = 0, nonnegativity and zero expectation imply Pπrand (Ω) = 0 almost surely. Hence the corresponding component induced magnitude variance is also zero. Delta Approximation. The exact magnitude variance cannot in general be determined from only µP 2 and √ σP . For a lightweight uncertainty proxy, CoRF applies the first order delta method to φ (P ) = P . For µP (Ω) > 0, σ 2 (Ω) 2 . (18) σS,∆ (Ω) := P 4µP (Ω) This is the approximation used in equation 11. It is expected to be accurate when σP2 (Ω) µP (Ω)
2 ≪ 1,
but it is not a distribution free upper or lower bound on the true magnitude variance. In particular, equation 18 equals one quarter of the upper bound expression in equation 17, which does not imply that the true variance differs from the approximation by a constant factor. For µP (Ω) = 0, we 2 set σS,∆ (Ω) = 0. Normalization and Reported Uncertainty. The deterministic prediction is compared with measurements after per spectrum min max normalization. Define rπ = max Seπ (Ω) − ℓπ ,
ℓπ = min Seπ (Ω) , Ω
Ω
rπ > 0.
The reported deterministic spectrum is Seπ (Ω) − ℓπ Sbπ (Ω) = . rπ Exactly normalizing every realization of Sπrand would make its minimum and maximum random and would couple all angular locations. Instead, the uncertainty propagation holds ℓπ and rπ fixed at the deterministic prediction. Under this fixed affine approximation, 2 σnorm,∆ (Ω) =
σP2 (Ω) , 4µP (Ω) rπ2 16
µP (Ω) > 0.
(19)
2 For µP (Ω) = 0, we set σnorm,∆ (Ω) = 0.
The final reported variance proxy is 2 σ bπ2 (Ω) = τu2 σnorm,∆ (Ω) + ε0 .
(20)
The temperature τu rescales the uncertainty proxy, while ε0 provides a residual variance floor. Both are fitted by the heteroscedastic objective in §3.4, with their gradients isolated from the deterministic mean branch. The reported quantity σ bπ2 is a component power uncertainty proxy rather than the complete predictive variance. It is conditional on the predicted component means and arrival directions. It does not account for errors in those means or directions, scene state estimation error, renderer mismatch, or distribution shift.
B
O BSERVABILITY FROM R EFERENCE S PECTRA
We study how many reference spectra are needed to constrain the scene-dependent propagation state. Because each reference depends on both the scene and its unknown transmitter pose, only scene variations that cannot be explained by pose variation provide effective scene information. We use this observation to bound the information provided by each reference and characterize how overlap among references affects the required reference budget. Observation Model and Effective Scene Information. Let p ∈ Rdp denote a local parameterization of the scene-dependent, query-independent propagation state, and ξi ∈ Ξ ⊂ Rdπ the unknown physical pose that generated reference i. The vectorized reference spectrum is yi = F (p, ξi ) ∈ RHW . We assume that F is differentiable at the operating points considered and that the reference poses can vary independently. For the per spectrum min max normalization, this excludes zero range spectra and points where the active extrema change under an infinitesimal perturbation. Around (p0 , ξi ), F (p0 + ∆p, ξi + ∆ξi ) − F (p0 , ξi ) ∆p , = Jp,i ∆p + Jξ,i ∆ξi + o ∆ξi 2
(21)
where Jp,i = Dp F (p0 , ξi ) ,
Jξ,i = Dξ F (p0 , ξi ) .
A scene perturbation is informative only through the part of its spectral effect that cannot be reproduced by changing the unknown reference pose. Define † Π⊥ i = I − Jξ,i Jξ,i ,
Ji = Π⊥ i Jp,i ,
(22)
†
where (·) denotes the Moore–Penrose pseudoinverse. The matrix Π⊥ i projects onto the orthogonal complement of col Jξ,i , so Ji retains only scene variations that cannot be absorbed by an infinitesimal pose change. Lemma 2 (Effective Scene Information). A scene perturbation ∆p is indistinguishable to first order from a perturbation of the unknown reference pose if and only if Ji ∆p = 0. Proof. By equation 21, the effect of ∆p can be canceled by some ∆ξi if and only if Jp,i ∆p ∈ col Jξ,i . This is equivalent to Π⊥ i Jp,i ∆p = 0, which gives the result by equation 22. 17
M
For a reference configuration C = {ξi }i=1 , define the stacked effective Jacobian J1 .. JC = . . JM Permuting the references only permutes its block rows, so the analysis is invariant to their ordering. Observable Dimension and Reference Sufficiency. At p0 , define the total locally observable scene subspace and its dimension as O⋆ (p0 ) = span row J (p0 , ξ) : ξ ∈ Ξ , r⋆ (p0 ) = dim O⋆ (p0 ) . (23) Thus, r⋆ is the number of independent scene directions that can be constrained locally by the admissible reference measurements after accounting for unknown reference pose. A finite reference configuration C is locally sufficient when rank JC = r⋆ (p0 ) . The minimum sufficient reference budget is therefore M ⋆ (p0 ) = min |C| : rank JC = r⋆ (p0 ) . C⊂Ξ |C|<∞
(24)
Since O⋆ (p0 ) is finite dimensional, a finite collection of admissible reference row spaces can span it whenever r⋆ > 0. Reference Capacity and Minimum Budget. For the spatial spectra used by CoRF, each reference is formed from the per element phases of an N element array. Let T e (ϕ) = ejϕ1 · · · ejϕN , x ϕ ∈ RN . Its Bartlett power is, up to a fixed scale, 2
H e (ϕ) . Pϕ (Ω) = g (Ω) a (Ω) x
(25)
Let H (ϕ) denote the differentiable map from the phase vector to the vectorized spatial spectrum, including the square root and per spectrum normalization. The reference map therefore factors locally as F (p, ξ) = H (ϕ (p, ξ)) . Proof of Proposition 1. For any scalar phase α, e (ϕ + α1) = ejα x e (ϕ) . x The common phase factor disappears from the magnitude in equation 25, so H (ϕ + α1) = H (ϕ) . Differentiating with respect to α at α = 0 gives DH (ϕ) 1 = 0. Hence DH (ϕ) has a nontrivial null direction and rank DH (ϕ) ≤ N − 1. By the chain rule, Jp,i = DH (ϕi ) Dp ϕ (p0 , ξi ) , and therefore rank Jp,i ≤ N − 1. Since Ji = Π⊥ i Jp,i and left multiplication cannot increase rank, rank Ji ≤ N − 1. 18
(26)
For any configuration C containing M references, rank JC ≤
M X
rank Ji ≤ M (N − 1) .
(27)
i=1
If C is locally sufficient, then r⋆ (p0 ) = rank JC ≤ M (N − 1) . Taking the minimum over all locally sufficient configurations gives r⋆ (p0 ) ⋆ M (p0 ) ≥ , N −1
(28)
which proves Proposition 1. Reference Complementarity. The bound in equation 28 is necessary but not sufficient because different references may constrain overlapping scene directions. For an existing configuration C and a candidate reference at ξ, define its marginal observable dimension as JC (29) ∆r (ξ | C) = rank − rank JC . J (p0 , ξ) By the dimension formula for sums of subspaces, ∆r (ξ | C) = rank J (p0 , ξ) − dim row J (p0 , ξ) ∩ row JC . Hence
0 ≤ ∆r (ξ | C) ≤ N − 1. A reference with ∆r = 0 adds no new first order observable direction, while a larger ∆r indicates greater complementarity with the existing references. Thus, reference utility depends on overlap between effective observable subspaces rather than reference count alone. Scope. The analysis is local around p0 and characterizes first order observability rather than global uniqueness of the nonlinear propagation state. It is also noiseless, so a reference with zero marginal rank may still improve statistical precision under measurement noise. The nuisance projection accounts for ambiguity caused by unknown reference poses and is independent of the estimator used to infer p. If the reference poses are observed, the known pose case follows by setting Jξ,i = 0, giving Ji = Jp,i . Finally, M ⋆ is a local identifiability budget for the reference measurements, not a prediction of the number of references required by an amortized learned estimator such as CoRF to achieve a particular reconstruction accuracy.
C
P ROPAGATION - TO -S PECTRUM E RROR B OUND
We prove Proposition 2, which characterizes how propagation estimation errors affect the rendered spatial spectrum. The proof separates representation mismatch from component power and direction errors, whose sensitivities are determined by the fixed receive array. Setup and Representation Error. Let P denote the target power spectrum and Pb = B (b z) the spectrum rendered from the estimated propagation parameters. For K
z = {(Ωn , mn )}n=0 , where n = 0 denotes the LOS component, the analytic renderer is B (z) (Ω) = g (Ω)
K X
H
2
mn a (Ω) a (Ωn ) .
(30)
n=0
The LOS direction is fixed by the query geometry, while its power is allowed to vary in the admissible renderer class. 19
Let ZK be a compact set of admissible renderer states with bounded nonnegative powers and admissible component directions. Choose z⋆ ∈ arg min ∥P − B (z)∥∞ ,
P ⋆ = B (z⋆ ) ,
z∈ZK
εphys = ∥P − P ⋆ ∥∞ .
(31)
Such a minimizer exists because ZK is compact and B is continuous on the finite angular grid. The term εphys captures spectrum structure that cannot be represented by the analytic renderer, including finite component approximation and measurement effects outside the renderer model. Array Response Sensitivity. Define the unit power Bartlett response 2
H
χ (Ω, Ω′ ) = a (Ω) a (Ω′ ) . Assume 0 ≤ g (Ω) ≤ gmax ,
∥a (Ω)∥2 ≤ A,
∥a (Ω) − a (Ω′ )∥2 ≤ La da (Ω, Ω′ ) ,
(32)
where da is a response relevant direction discrepancy. Lemma 3 (Array Response Sensitivity). Under equation 32, χ (Ω, Ω′ ) ≤ A4 ,
|χ (Ω, Ω1 ) − χ (Ω, Ω2 )| ≤ 2A3 La da (Ω1 , Ω2 ) .
(33)
Proof. By Cauchy–Schwarz, H
a (Ω) a (Ω′ ) ≤ ∥a (Ω)∥2 ∥a (Ω′ )∥2 ≤ A2 , which gives the first inequality. For
H
H
x = a (Ω) a (Ω1 ) , 2 we have |x| , |y| ≤ A , and therefore 2
2
|x| − |y|
y = a (Ω) a (Ω2 ) ,
≤ (|x| + |y|) |x − y| ≤ 2A3 ∥a (Ω1 ) − a (Ω2 )∥2 ≤ 2A3 La da (Ω1 , Ω2 ) .
For a planar unit modulus array with element locations du ∈ R2 , 2π , λ where u∥ (Ω) is the projection of the arrival direction onto the array plane. Define T
au (Ω) = e−jκdu u∥ (Ω) ,
κ=
d∥ (Ω, Ω′ ) = u∥ (Ω) − u∥ (Ω′ ) 2 . If D contains the centered element locations as its rows, then ejx − ejy ≤ |x − y| gives ∥a (Ω) − a (Ω′ )∥2 ≤ κ ∥D∥2 d∥ (Ω, Ω′ ) . Hence one may choose
√ La = κ ∥D∥2 . A = N, Moreover, projection is nonexpansive, so if dgeo is the geodesic angle between two arrival directions, dgeo (Ω, Ω′ ) ′ d∥ (Ω, Ω ) ≤ 2 sin ≤ dgeo (Ω, Ω′ ) . 2 Thus, the same sensitivity bound also holds with geodesic angular error. Propagation to Spectrum Error. Write n oK b n, m b z= Ω bn
K
z⋆ = {(Ω⋆n , m⋆n )}n=0 .
,
n=0
20
b 0 and Ω⋆ equal the known LOS direction. Because the learned components are unordered, Both Ω 0 let υ denote a matching permutation over {1, . . . , K}. Define the matched component power error and power weighted direction error as Em = |m b 0 − m⋆0 | +
K X
m b υ(n) − m⋆n ,
n=1
EΩ =
K X
b υ(n) , Ω⋆ . m⋆n da Ω n
n=1
Proof of Proposition 2. By the triangle inequality, Pb − P
∞
≤ εphys + Pb − P ⋆
∞
.
For each learned component, b υ(n) − m⋆n χ (Ω, Ω⋆n ) m b υ(n) χ Ω, Ω b υ(n) ≤ m b υ(n) − m⋆n χ Ω, Ω b υ(n) − χ (Ω, Ω⋆n ) + m⋆n χ Ω, Ω b υ(n) , Ω⋆n , ≤ A4 m b υ(n) − m⋆n + 2A3 La m⋆n da Ω where the last inequality follows from Lemma 3. For the LOS component, the direction is fixed, so its contribution is bounded by A4 |m b 0 − m⋆0 |. Summing over the components, using g (Ω) ≤ gmax , and taking the maximum over the angular grid gives Pb − P ⋆ ≤ gmax A4 Em + 2gmax A3 La EΩ . ∞
Therefore, Pb − P
∞
Cm = gmax A4 ,
≤ εphys + Cm Em + CΩ EΩ ,
CΩ = 2gmax A3 La .
(34)
This proves Proposition 2. For the planar unit modulus array above, the sensitivity constants become Cm = gmax N 2 ,
CΩ = 2gmax N 3/2 κ ∥D∥2 .
(35)
Thus, for a fixed receive array, the conversion from propagation errors to spectrum error is determined by fixed array physics rather than the scene. Magnitude and Normalized Spectrum. The preceding result is stated in the power domain because the analytic renderer is additive in the component powers. Let p √ S = P, Sbu = Pb √ √ 2 denote the target and rendered unnormalized magnitude spectra. Since x − y ≤ |x − y| for x, y ≥ 0, r Sbu − S
∞
≤
Pb − P
∞
.
For a nonconstant spectrum u, define N (u) =
u − min u , max u − min u 21
ru = max u − min u.
(36)
Lemma 4 (Normalization Stability). For any two nonconstant spectra u and v on the finite angular grid, 2 ∥N (u) − N (v)∥∞ ≤ ∥u − v∥∞ . (37) max {ru , rv } Proof. Set ε = ∥u − v∥∞ , a = min u, b = max u, a′ = min v, and b′ = max v. Then |a − a′ | ≤ ε and |b − b′ | ≤ ε. At any grid point, write v = a′ + trv with t ∈ [0, 1] and u = v + δ with |δ| ≤ ε. It follows that |(1 − t) (a′ − a) + t (b′ − b) + δ| 2ε |N (u) − N (v)| = ≤ . ru ru Interchanging u and v gives the same bound with rv , and taking the smaller of the two bounds proves equation 37. Applying Lemma 4 to Sbu and S, and then using equation 36, gives r 2 n o ≤ . N Sbu − N (S) Pb − P ∞ ∞ max rSbu , rS
(38)
Thus, the power domain bound transfers directly to the normalized magnitude spectrum whenever both spectra have nonzero range. Scope. The result is a deterministic worst case stability bound and does not assume that εphys is small. It characterizes how specified propagation parameter errors are converted into spectrum error, rather than how accurately the learned model estimates those parameters. For a planar array, d∥ reflects the intrinsic response ambiguity between directions with identical in plane direction cosines. The result therefore separates scene-dependent propagation inference from the fixed propagation-to-spectrum sensitivity of the analytic renderer.
D
I MPLEMENTATION D ETAILS
This appendix describes the training objectives and hyperparameters (§D.1), model architecture (§D.2), optional power refinement (§D.3), receiver array calibration (§D.4), and dataset (§D.5). D.1
T RAINING O BJECTIVES AND H YPERPARAMETERS
Training Protocol. Training proceeds in two stages. Stage I pretrains the spectrum encoder Eϕ and reference contextualizer Tψ , after which these modules are frozen. Stage II trains the reference-toanchor projection, propagation estimation, query conditioning, and uncertainty modules. Unless stated otherwise, each training instance uses M = 32 references, with the queries excluded from the reference set. Unless stated otherwise, the canonical main-paper checkpoints use random seed 0. Training-randomness estimates repeat these checkpoints with seeds {0, 1, 2} on all 12 folds, and the coordinate and pretraining ablations also use three seeds per condition; these exceptions are identified with their results. Stage I: Reference Representation Pretraining. We pretrain Eϕ and Tψ using supervised contrastive learning over scene identity. Reference sets from the same scene are treated as positives, while those from different scenes are treated as negatives. Each training step draws Bs = 4 reference sets from every training scene. Each set b is encoded by Eϕ and Tψ , mean-pooled over its M contextualized tokens, passed through a projection head, and ℓ2 -normalized to obtain zb . Let yb denote the scene label of set b, let P (b) = {b′ ̸= b : yb′ = yb } denote its positive sets, and let B be the total number of reference sets in the training step. The supervised contrastive objective is B X exp z⊤ 1 X 1 b zb′ /τc . (39) Lsup = − log P ⊤ B |P (b)| ′ ′′ ̸=b exp zb zb′′ /τc b b=1 b ∈P(b)
We use the contrastive temperature τc = 0.1. Stage I is trained for 1500 steps using Adam with learning rate 3 × 10−4 . 22
Stage II: Propagation Field Training. We freeze the Stage I modules and train the remaining model using query spectra outside the reference set. The full objective is L = Lrec + λcu Lcu + λmar Lmar + λunc Lunc .
(40)
Here, Lrec is the spectrum reconstruction loss, Lcu is the content-use loss, Lmar is the referenceseparation margin loss, and Lunc is the uncertainty loss. The coefficients λcu , λmar , and λunc weight the corresponding auxiliary objectives. 1) Reconstruction Loss Lrec . The reconstruction objective combines amplitude-weighted pointwise error, summed-amplitude consistency, structural similarity, and normalized energy error: Lrec = Lamp + λe Lenergy + λs LSSIM + λn Lnorm .
(41)
All four terms compare the predicted spectrum Sb and target spectrum S on the H × W angular grid G, using the per-measurement normalization described in §D.5. For the amplitude weighting, we use γa = 0.7 and εa = 0.02. The reconstruction-loss weights are λe = 0.2, λs = 0.6, and λn = 0.5. 2) Amplitude-Weighted Loss Lamp . The amplitude-weighted pointwise loss emphasizes high-response regions of the target spectrum: P b Ω∈G w (Ω) S (Ω) − S (Ω) γ P (42) , w (Ω) = S (Ω) a + εa . Lamp = w (Ω) Ω∈G 3) Summed-Amplitude Loss Lenergy . This loss matches the summed spectrum amplitude and normalizes the error by the number of angular cells: Lenergy =
X 1 X b S (Ω) − S (Ω) . HW Ω∈G
(43)
Ω∈G
4) Structural Loss LSSIM . The structural loss is one minus the mean SSIM map, computed using a uniform 7 × 7 window and the standard constants C1 = 0.012 and C2 = 0.032 : 2µSbµS + C1 2σSS 1 X b + C2 . LSSIM = 1 − (44) HW 2 + µ2 + C σS2b + σS2 + C2 1 Ω∈G µS S b Here, µSb and µS denote the local means, σS2b and σS2 denote the local variances, and σSS b denotes the local covariance at Ω. 5) Normalized Energy Loss Lnorm . The normalized energy loss measures squared reconstruction error relative to the target energy and corresponds to the NMSE metric: 2 P b Ω∈G S (Ω) − S (Ω) Lnorm = , ε = 10−8 . (45) P 2 S (Ω) + ε Ω∈G 6) Reference-Use Regularization Lcu , Lmar . To encourage the conditioned field to use the targetscene references, we introduce a content-use loss and a reference-separation margin loss. The content-use loss reconstructs the spectra contained in the reference set: M
Lcu =
1 X Lrec Sbπi , Sπi . M i=1
(46)
We additionally sample a mismatched reference set C′ from another scene and encourage the matched reference set C to produce a lower query reconstruction loss: h ′ i Lmar = Lrec SbπC⋆ , Sπ⋆ − Lrec SbπC⋆ , Sπ⋆ + δm . (47) +
We use λcu = 0.5, λmar = 1.0, and δm = 0.02. The content-use loss is introduced with a 400-step warm-in. Both reference-use losses are enabled only for cross-scene training and are disabled for the per-scene setting. 23
Table 7: Model architecture of CoRF used in reported experiments. Module
Setting
Value
Canonical anchor field
Anchor lattice Canonical extent Physical extent Backbone width Anchor outputs
20 × 20 × 10 = 4,000 [−1, 1]3 [−5, 5] × [−4, 4] × [0, 3] m 512 ∆xn , αn , e0n
Reference encoder
Spectrum CNN Pooling Contextualizer Attention FFN width Positional encoding
Conv 32-64-128-128 AdaptiveAvgPool8 Transformer, dim 256 4 heads, 2 layers 512 none
Anchor conditioning
Cross-attention Bearing-aligned CNN
dim 512, 4 heads 32 channels, angular grid preserved
Query conditioning
Model width Fourier frequencies Modulation
64 6 affine FiLM
Uncertainty
Parameterization Calibration
vn = m2n /κn global τu , variance floor ε0
Renderer
Receiver array Rendering LOS power Uncertainty
4 × 4 at 920 MHz analytic Bartlett with explicit LOS global, softplus (·)2 analytic moment propagation
Model size
Total Reference encoder Field and conditioning
12.7M 3.8M 8.9M
7) Uncertainty Loss Lunc . The uncertainty branch is trained using a heteroscedastic Gaussian negative log-likelihood based on the analytically propagated spectrum variance. Gradients from this objective are detached from the mean-prediction branch. Let σ 2 (Ω) denote the per-cell predictive variance of the normalized spectrum. The uncertainty objective is h i2 b S (Ω) − sg S (Ω) 1 X Lunc = + log σ 2 (Ω) + εu . (48) 2 2HW σ (Ω) + εu Ω∈G
Here, sg [·] denotes the stop-gradient operator. The calibrated predictive variance is 2 σ 2 (Ω) = τu2 σphys (Ω) + ε0 ,
(49)
2 where σphys (Ω) is the analytically propagated variance from Equation 11, expressed in the units of
the normalized spectrum. For unconstrained optimization, we parameterize the global uncertainty temperature and variance floor as τu = eϑτ and ε0 = softplus (ϑε ), with ϑτ and ϑε initialized to 0 and −4.6, respectively. We additionally use a fixed numerical floor εu = 10−4 . The uncertainty objective updates only the concentration head that produces the component variances and the global calibration parameters τu and ε0 through their unconstrained parameters ϑτ and ϑε . The mean prediction, canonical anchor field, and analytic array response enter this objective with stopped gradients. We use λunc = 1.0. D.2
M ODEL A RCHITECTURE
We use the same architecture across all reported experiments unless stated otherwise. Table 7 summarizes the main architectural configuration. The implementation follows the cross-scene propagation field in §3.1 and the analytic array physics in §3.2. Reference Encoder and Contextualizer. Each reference spectrum is processed independently by a four-stage 2D CNN with channel widths 32, 64, 128, and 128, followed by adaptive average pooling. The resulting reference token is contextualized by a two-layer Transformer with embedding 24
dimension 256, four attention heads, and feed-forward width 512. No positional encoding is applied to the reference sequence, so the contextualizer remains permutation equivariant over the unordered reference set. In parallel, a shallow bearing-aligned CNN with 32 channels preserves the angular grid and provides direction-specific reference features. Canonical Anchor Field. The shared field contains K = 20 × 20 × 10 = 4,000 anchors over the 3 canonical extent [−1, 1] , corresponding to the normalized physical extent [−5, 5]×[−4, 4]×[0, 3] m. Each anchor uses a 512-dimensional backbone representation and cross-attends to the contextualized reference tokens using four attention heads. The resulting global reference-conditioned feature is fused with the bearing-aligned feature before anchor decoding. The anchor decoder predicts the position refinement ∆xn , occupancy parameter αn , and base emission e0n . The position refinement is constrained within half of one anchor cell as described in §3.1. Query and Uncertainty Heads. Query and anchor geometry in the receiver-array frame are encoded using six Fourier frequencies. A width-64 network produces affine FiLM modulation of the base anchor emission, yielding the query-conditioned emission used to compute the mean component power mn . A separate uncertainty head predicts a positive concentration κn , giving componentpower variance vn = m2n /κn . A learned global temperature τu and variance floor ϵ0 calibrate the propagated component-power variance. Uncertainty gradients are detached from the mean branch, leaving the synthesized mean spectrum unchanged. K
Analytic Renderer. The learned model outputs the propagation parameters {(Ωn , mn , vn )}n=1 , which are rendered through the analytic receiver-array model in §3.2. For the simulated arrays, the receiver elements are isotropic and the element power pattern is g ≡ 1. The explicit LOS power mlos is parameterized as a squared softplus and shared across scenes and queries. Because the targets are normalized independently, this scalar represents a normalized direct-path prior rather than absolute received power. The renderer constructs the receiver covariance, evaluates the Bartlett response, and propagates component-power uncertainty analytically through the same array response. Implementation Compatibility. The released checkpoint format inherits a degree-9 angular emission head and scale and rotation outputs from the Gaussian-field implementation on which the code was built. CoRF reads only the zeroth-order complex emission coefficient; the other emission coefficients and the scale and rotation outputs are disconnected from the covariance, spectrum, uncertainty, and loss. They are retained solely for checkpoint compatibility and can be pruned without changing any prediction. The reported checkpoint parameter totals conservatively include this dormant head, while the mathematical model and analyses use only the active variables defined in §3.1. D.3
TARGET-S CENE P OWER R EFINEMENT
Power refinement is an optional target-scene optimization that provides a lightweight correction while keeping the pretrained propagation model fixed. For a target scene, we freeze the spectrum encoder, reference contextualizer, canonical anchor field, propagation estimation modules, and analytic array model. Only the component powers are refined through lightweight scene-specific and query-dependent residual corrections. Refinement uses only the same M target-scene references already used to condition the canonical anchor field and requires no additional measurements, but it is distinct from feed-forward conditioning. Although the cross-scene model transfers the propagation structure to an unseen scene, the predicted component powers can still exhibit scene-specific mismatch because attenuation and reflection strengths vary across environments. Rather than fitting a new scene-specific propagation field, we use K the available target-scene references to apply a lightweight correction to {mn }n=1 while preserving the learned anchor geometry and arrival directions. After this one-time refinement, the corrected model is fixed and reused for all queries in the target scene. Residual Power Correction. For each target scene, we introduce a lightweight residual model that corrects the pretrained component powers while leaving the propagation geometry fixed. For anchor n and query configuration π⋆ , the residual correction is δn (π⋆ ) = bPR n + rζ (ΓF (π⋆ ))n ,
(50)
where bPR n is a learned per-anchor power bias and rζ is a residual MLP applied to a Fourier encoding ΓF (π⋆ ) of the query position, array-local direction, and range. The bias bPR n captures a 25
scene-specific correction shared across all queries, whereas rζ (ΓF (π⋆ ))n captures an additional K query-dependent correction. Both bPR and rζ are fitted once for the target scene using n n=1 the same M reference measurements used for scene conditioning. We apply this residual in the pre-softplus domain so that the corrected component power remains nonnegative: m e n (π⋆ ) = softplus softplus−1 (mn (π⋆ )) + δn (π⋆ ) . (51) The anchor biases and the output layer of rζ are initialized to zero, so m e n = mn before refinement. After fitting, the residual model is fixed and reused for all queries in the target scene. Because Equation 51 modifies only the non-LOS component powers, the canonical anchor geometry and arrival K directions {Ωn }n=1 remain unchanged. In the per-scene setting, we additionally allow a scalar correction to the explicit LOS power, as described below. Residual-Map Architecture. The residual-map capacity depends on the amount of target-scene data available for refinement. In the few-shot setting, rζ is a two-layer MLP with width 64, applied to a two-octave Fourier encoding of the query geometry with 35 input dimensions, for 0.266M parameters. Together with the 4,000 anchor biases, the full few-shot correction adapts 0.270M parameters per target scene. When the full scene training split is available, we use a three-layer MLP with width 512 and an eight-octave Fourier encoding with 119 input dimensions, for approximately 2.6M parameters. Before refinement, the pretrained model is evaluated once to cache the base component K powers {mn (π)}n=1 for the refinement measurements. The residual model then learns only the correction in Equation 50, while these pretrained powers remain fixed. Protocol. For each target scene, we first use the M = 32 references to construct the n scene-conditioned o K , rζ for canonical anchor field with the pretrained model. We then fit one residual model bPR n n=1 that scene using the same references while keeping the pretrained model frozen. After refinement, the fitted residual model is fixed and reused for all test queries in the scene without further optimization. We optimize the residual map using the reconstruction objective Lrec defined in §D.1. Adam is used K with learning rate 10−3 for rζ and 5 × 10−3 for the anchor biases bPR , together with a cosine n n=1 learning-rate schedule, batch size 8, and gradient clipping at 1.0. In the few-shot setting, the same M = 32 target-scene references constitute the complete refinement budget. All 32 references are used to condition the canonical anchor field, while 24 are used to fit the residual model and the remaining 8 are reserved for step selection. We optimize for up to 3000 steps with weight decay 10−4 , evaluate the selection references every 100 steps, and retain the step with the highest mean PSNR. Thus, optional power refinement requires no target-scene measurements beyond the original M = 32 references. For the per-scene setting, the scene training split is used for fitting and the validation split for step selection. We optimize for up to 16,000 steps without weight decay and evaluate validation performance every 500 steps. In this setting, the residual map additionally includes a linear head on its hidden features that rescales the explicit LOS power by a bounded exponential factor initialized to one. The step with the highest validation PSNR is retained. D.4
R ECEIVER A RRAY C ALIBRATION
Receiver array calibration adapts the analytic receiver response when the deployed array differs from its nominal specification. Unlike power refinement, calibration leaves the predicted propagation K parameters {(Ωn , mn , vn )}n=1 unchanged and modifies only the receiver response used by the analytic renderer. Such calibration is needed because errors in array orientation, element geometry, or element response can distort the mapping from the predicted propagation components to the observed spatial spectrum. The spectrum encoder, reference contextualizer, canonical anchor field, and propagation estimation modules therefore remain frozen throughout calibration. Calibration uses a separate set of NC measurements and is performed once for each receiver, after which the calibrated response is reused for all queries collected with that receiver. Calibration Formulation. For a receiver array with nominal steering response a (Ω), we introduce receiver-specific calibration parameters νarr and replace the nominal response with a (Ω) −→ aνarr (Ω) . 26
(52)
We estimate νarr from the NC calibration measurements while keeping the learned propagation field fixed. The calibration measurements are disjoint from the M references used for scene conditioning and from the test queries. Simulated Receiver Arrays. For a simulated receiver array, we construct the analytic receiver response from its nominal geometry and fit the three-parameter calibration vector νarr = (∆φarr , sarr,x , sarr,y ) ,
(53)
where ∆φarr corrects the in-plane array orientation, and sarr,x and sarr,y correct the element spacing along the two array axes. We estimate νarr from NC calibration measurements while keeping the learned propagation field fixed. After calibration, the corrected response aνarr (Ω) is fixed and reused for all subsequent queries collected with that receiver. Real Measured Receiver. For the real measured receiver, we keep the learned propagation model frozen and estimate only receiver-specific calibration parameters from NC calibration measurements. Because a physical receiver can exhibit element-dependent gain, geometry, and directional-response mismatch, we model the calibrated response using a complex gain per element, an in-plane position offset per element, and an angle-dependent element response. For receiver element u, the complex gain is gu = exp (ℓu + jϕu ) ,
(54)
where ℓu and ϕu parameterize the log-amplitude and phase, respectively. The log-amplitude ℓu is constrained to [−3, 3], and the element gains are normalized to unit mean magnitude. The position correction ∆pu ∈ R2 is added to the nominal element position when computing the steering phase. The angle-dependent element response is ⊤ cu (φ, θ) = exp α⊤ u f (φ, θ) + j βu f (φ, θ) , where φ and θ denote the array-local azimuth and elevation. We use the angular basis h i Kel Kaz f (φ, θ) = 1, {cos (kφ) , sin (kφ)}k=1 , {cos (jθ) , sin (jθ)}j=1 .
(55)
(56)
We set Kaz = 12 and Kel = 4, giving 33 coefficients for the log-amplitude αu and 33 coefficients for the phase βu of each of the 16 receiver elements. Including complex gain and the two-dimensional position correction, the real-array model fits 70 scalar parameters per element, or 1,120 in total. The resulting log-amplitude of the angle-dependent response is constrained to [−4, 4]. The calibrated steering response evaluates the steering phase at the corrected element positions and multiplies the response of element u by gu cu (Ω). The same calibrated response is applied to both the anchor directions and the explicit line-of-sight direction. All calibration parameters are initialized to reproduce the nominal receiver response, with gu = 1, ∆pu = 0, and cu (Ω) = 1. No additional regularization is applied beyond the parameter bounds described above. Protocol. We use NC = 64 calibration measurements by default. These measurements are separate from the M = 32 references used to condition the target-scene field and from the test queries. For the simulated arrays, the three calibration parameters in Equation 53 are fitted by coordinate search because rebuilding the analytic response for a new array geometry is not differentiable in our implementation. We perform three coordinate-search sweeps over the in-plane rotation from −12◦ to 12◦ and the two axis scalings from 0.86 to 1.16, using 13 candidate values for each parameter. Within each sweep, the parameters are optimized sequentially by retaining the candidate that maximizes the mean SSIM over the NC calibration measurements. For the real receiver, we optimize the calibration parameters using the reconstruction objective Lrec defined in §D.1, while keeping the conditioned propagation parameters fixed. We use Adam with learning rate 5 × 10−3 , a cosine learning-rate schedule, and gradient clipping at 1.0, and optimize for up to 8000 steps. We reserve one quarter of the NC measurements for step selection and use the remaining three quarters for fitting. With the default NC = 64, this corresponds to 48 fitting measurements and 16 selection measurements. The calibration checkpoint with the highest mean PSNR on the selection measurements is retained. After fitting, the receiver calibration is fixed and applied unchanged to all test queries. The test queries are not used for calibration or model selection. The default calibration-only protocol above 27
uses disjoint conditioning and calibration sets, for a total of 32 + 64 = 96 target-domain spectra. For the joint measurement-matched real-data result in §E.6, the scene conditioner uses M = 32 references, and a disjoint set of NC = 64 spectra calibrates the element-wise receiver response. The power-refinement map defined in §D.3 reuses both sets without additional measurements; a held-out subset of the same 96-spectrum budget is used for checkpoint selection. The pretrained propagation model remains frozen, and the test split remains untouched. This combined real-data deployment protocol is denoted by CoRFPR,CAL . The real experiment evaluates one initialization and one measured receiver, and therefore does not characterize sensitivity to calibration initialization, measurement noise, or unmodeled mutual coupling. D.5
DATASET AND P REPROCESSING
Simulated Dataset. We generate 35 simulated scenes using the Sionna RT differentiable ray tracer (Hoydis et al., 2023). The corpus contains seven scene categories: conference room, classroom, bedroom, corridor, open office, laboratory, and lounge, with five variants per category. We refer to these scenes as S1–S35 in category order: S1–S5 are conference rooms, S6–S10 classrooms, S11–S15 bedrooms, S16–S20 corridors, S21–S25 open offices, S26–S30 laboratories, and S31–S35 lounges. The room dimensions are 5 × 4 × 2.8 m for bedrooms, 6.5 × 5.5 × 3.2 m for lounges, 7 × 7 × 3 m for laboratories, 8 × 6 × 3 m for conference rooms, 10 × 7 × 3 m for classrooms, 12 × 2.2 × 3 m for corridors, and 13 × 9 × 3 m for open offices. Surfaces use fixed permittivity and conductivity values at 920 MHz adapted from the ITU-R P.2040 building-material models, with concrete and brick from revision 3 (ITU-R, 2023) and plasterboard and glass from revision 1 (ITU-R, 2015). The five variants of each category differ in furniture layout and clutter. Ray tracing uses up to five interactions per path with specular reflection and refraction. Each scene contains 4.1k–5.9k transmitter locations, with approximately 4.9k on average, distributed throughout the room volume. The per-scene median SNR is sampled uniformly from 8 to 18 dB, with a realized range of 8.2 to 17.9 dB and a mean of 12.6 dB. Complex Gaussian noise is added to the per-element channel response using the scene-level median signal to determine the noise level. All simulated scenes share the same generation pipeline, carrier frequency, material-value table, and interaction model. The category-exclusion protocol therefore measures semantic scene transfer inside this distribution rather than transfer across simulators or physical propagation regimes. Real Measured Dataset. We use the publicly released NeRF2 RFID measurements (Zhao et al., 2023), denoted S36. The dataset contains 6,123 measurements collected with a physical 4×4 receiver array operating at 915 MHz. Following the released processing pipeline, we construct Bartlett spectra using the same angular-grid convention as for the simulated data. The real measurements are not used to train the cross-scene model and serve only for simulation-to-real evaluation. Spectrum Construction and Scene Normalization. The simulated receiver is a fixed 4 × 4 planar array with 16 elements and 0.16 m spacing. At 920 MHz, the wavelength is λ = 0.326 m, corresponding to an element spacing of approximately 0.49λ. For each transmitter, we construct the single-snapshot, phase-only Bartlett magnitude N
S (Ω) =
1 X exp j φst , u (Ω) + ∠Hu N u=1
(57)
where Hu is the channel response at receiver element u, and φst u (Ω) is its steering phase. Following NeRF2 (Zhao et al., 2023), element magnitudes are discarded. Each spectrum is sampled on a 90 × 360 angular grid covering elevations from 1◦ to 90◦ and azimuths from 0◦ to 359◦ , and is min-max normalized per measurement. The target spectrum is constructed from a coherent single snapshot, whereas the analytic renderer in §3.2 models propagation components through an incoherent covariance sum. We treat this difference as a modeling approximation. Each scene contains the receiver position and orientation Q, together with all transmitter positions. These coordinates define the query configurations and the canonical normalization, but the reference encoder receives only measured spectra and not transmitter positions associated with individual references. Before constructing the canonical anchor field, we isotropically rescale scene geometry as 5 4 3 e = cs x, x cs = 0.9 min , , , (58) maxj |xj | maxj |yj | maxj zj 28
where j ranges over all transmitter and receiver positions in the scene. This maps each scene into the canonical extent [−5, 5] × [−4, 4] × [0, 3] m with a 10% margin. Because the same isotropic factor is applied to all coordinates, angular bearings and receiver orientation are preserved. Data Splits. We use the same per-scene 70/10/20 train, validation, and test split for all methods. For evaluation, reference sets are sampled from the training split and queries from the test split; the validation split remains disjoint from both. A simulated scene therefore contains approximately 3.4k training measurements and 1.0k test measurements, while the real dataset contains 4,286/612/1,225 train, validation, and test measurements. The default reference budget is M = 32. For generalization to unseen scene variants, we use five folds, each excluding one variant from every scene category. For generalization to unseen scene categories, we use seven folds, each excluding all five variants of one category. For simulation-to-real evaluation, the cross-scene model is trained on all 35 simulated scenes and evaluated on the real measurements. All methods use identical splits and reference sets where applicable.
E
A DDITIONAL E XPERIMENTAL R ESULTS AND A NALYSES
Roadmap. The first six subsections provide qualitative results (§E.1), scene-level comparisons (§E.2), reliability analysis (§E.3), receiver-array adaptation (§E.4), a localization case study (§E.5), and realdata transfer (§E.6). The remaining subsections examine reference evidence (§E.7), representation and renderer choices (§E.8), deployment efficiency (§E.9), and physical distribution shift (§E.10). The main-paper synthesis tables and their scene-level breakdowns report CoRFPR ; other analyses specify the evaluated variant. Shared Protocol. Unless stated otherwise, simulated experiments use the same 35 scenes, 4 × 4 array, 920 MHz carrier, and 90 × 360 normalized-amplitude spectra as the main experiments. Reported cross-fold deviations reflect variation among held-out scenes. Across three seeds on all 12 folds, the within-fold PSNR standard deviation averages 0.21 dB and reaches at most 0.73 dB. E.1
Q UALITATIVE R ESULTS
This subsection presents qualitative spectra and reflection-dominated failures. Figure 4 compares predictions for per-scene fitting, unseen scene variants, and unseen scene categories. Each panel shows ten selected test locations (P1–P10), with at most one per scene. Rows show the ground-truth (GT) spectrum, CoRFPR (“Ours”), and the corresponding baselines. The figures use the saved predictions underlying the quantitative results. For display, each spectrum is normalized by its maximum; azimuth is on the horizontal axis and elevation on the vertical axis. E.1.1
P ER -S CENE R ESULTS
Figure 4(a) shows per-scene synthesis, with each method using the same scene-specific training and test splits; location P1 comes from the measured scene S36. The baselines generally recover the dominant lobe but often blur or misplace weaker multipath structure, and some produce spurious high-response bands near the upper elevation boundary. Across the displayed locations, their SSIM ranges from 0.51 to 0.58, compared with 0.74 to 0.77 for CoRFPR . E.1.2
U NSEEN S CENE VARIANTS
As illustrated in Figure 4(b), the few-shot baselines often misplace lobes or collapse multipath structure into a diffuse region with only M = 32 target-scene references. Transfer from another scene in the same category does not reliably correct these errors. GRaF produces low-contrast, vertically streaked spectra that sometimes approximately localize the strongest lobe. In contrast, CoRFPR better preserves dominant and secondary lobe directions; its remaining errors mainly concern their relative strengths and the weakest components. Across the displayed locations, the strongest baseline achieves 0.29–0.38 SSIM. 29
(a) Per-scene
(b) Unseen scene variants
(c) Unseen scene categories
Figure 4: Predicted spatial spectra across the three evaluation settings. Rows compare ground truth (GT), CoRFPR (Ours), and the corresponding baselines.
30
Figure 5: Worst-case predictions. The ten lowest-SSIM predictions of feed-forward CoRF in the first unseen-variant fold are shown with all baselines on the same queries. E.1.3
U NSEEN S CENE C ATEGORIES
The qualitative trend persists when an entire scene category is excluded from training, as illustrated in Figure 4(c). The baselines, including models transferred from other categories, often misplace or blur multipath lobes and achieve only 0.29–0.37 SSIM across the displayed locations. In contrast, CoRFPR continues to preserve the dominant lobe directions; its remaining errors mainly concern their relative strengths. E.1.4
R EFLECTION -D OMINATED FAILURE C ASES
This diagnostic uses feed-forward CoRF on the first unseen-variant fold; the preceding panels show CoRFPR . Across 6,436 test queries, mean SSIM is 0.804, with 95.0% above 0.5 and only 0.3% below 0.3. We examine the worst 5%: 322 queries with SSIM ≤ 0.50. Figure 5 shows the ten lowestSSIM predictions. All ten targets have multiple narrow lobes and a strongest arrival away from the direct path, while CoRF generally predicts one or two broad lobes. The strongest baseline also falls from 0.60 SSIM overall to 0.46 on the worst 5%. On the ten displayed queries, it achieves 0.23– 0.41 SSIM vs. 0.16–0.26 for CoRF; neither method captures the full reflected structure. The worst-5% queries are predominantly reflection dominated. A spectrum formed from the true line-of-sight direction alone has median SSIM 0.32 on this tail vs. 0.82 on the remaining queries. The strongest ground-truth lobe lies a median 23.6◦ from line of sight vs. 2.1◦ , while the spectral energy within 20◦ of line of sight falls from 62% to 32%. The diffuse floor rises from 0.13 to 0.28. Across all queries, prediction SSIM correlates strongly with line-of-sight explainability, with Spearman ρ = 0.96. The errors also have a consistent structure. Ground-truth spectra in the worst 5% contain a median of three lobes with azimuth width 26◦ , whereas CoRF typically predicts one broader lobe of width 56◦ . Nearest-reference distance is comparable to that of a random control set: 0.76 vs. 0.85 m. The best-matching reference is less similar for failure cases, with SSIM 0.52 vs. 0.69. Thus, the difficulty is associated more with poorly matched reference spectra than with reference distance alone; narrow reflected arrivals tend to be smoothed into broader components. E.2
S CENE -L EVEL S PATIAL S PECTRUM S YNTHESIS R ESULTS
Crosswalk to the Main Results. The per-scene averages in Table 8 reproduce Table 1 after rounding. The unseen-variant and unseen-category breakdowns correspond to the CoRFPR results in Table 2 and Table 3. Those main tables aggregate over folds, whereas the breakdowns pool queries within scenes; unweighted averages of the scene rows may therefore differ. E.2.1
P ER -S CENE R ESULTS
Table 8 reports all four metrics for each of the 36 scenes, grouped by category. Queries are pooled within scenes, and the table averages reproduce Table 1 after rounding. Because every method is 31
Table 8: Per-scene results in the per-scene setting. Results are reported for all 36 scenes and all four metrics. S36 denotes the real measured dataset, and the averages correspond to Table 1. PSNR↑
SSIM↑
AoA err (◦ )↓
NMSE↓
Scene NeRF2 GSRF WRF-GS GRaF CoRFPR NeRF2 GSRF WRF-GS GRaF CoRFPR NeRF2 GSRF WRF-GS GRaF CoRFPR NeRF2 GSRF WRF-GS GRaF CoRFPR Conference S1 27.13 S2 25.11 S3 19.46 S4 19.27 S5 18.06 Classroom S6 23.51 S7 29.27 S8 22.63 S9 24.83 S10 21.36 Bedroom S11 28.40 S12 25.47 S13 23.23 S14 22.86 S15 23.79 Corridor S16 15.14 S17 21.56 S18 15.55 S19 17.89 S20 17.25 Open office S21 23.61 S22 25.60 S23 24.20 S24 19.63 S25 19.85 Lab S26 21.70 S27 19.53 S28 18.70 S29 18.58 S30 15.57 Lounge S31 25.03 S32 23.75 S33 19.79 S34 20.21 S35 18.14
23.97 22.69 19.12 18.73 18.16
26.84 24.65 20.06 19.15 19.95
27.06 25.35 19.55 19.56 19.59
29.17 27.61 22.03 22.00 21.52
0.847 0.801 0.644 0.638 0.590
0.763 0.735 0.642 0.635 0.612
0.832 0.794 0.638 0.626 0.666
0.816 0.781 0.612 0.615 0.626
0.891 0.865 0.742 0.749 0.731
0.030 0.057 0.152 0.146 0.219
0.071 0.097 0.170 0.171 0.224
0.034 0.059 0.150 0.155 0.162
0.030 0.053 0.151 0.139 0.158
0.021 0.040 0.104 0.096 0.130
3.10 5.89 4.69 7.91 11.46 11.30 10.53 9.64 15.53 14.01
3.29 4.91 11.18 10.74 12.85
2.47 3.80 9.12 8.14 11.85
2.00 3.21 7.25 7.51 11.07
21.83 23.86 20.89 21.96 19.73
23.32 29.70 22.45 25.71 21.20
23.21 29.50 22.64 24.90 20.55
26.25 31.83 25.29 27.97 23.78
0.770 0.875 0.735 0.780 0.691
0.718 0.749 0.688 0.702 0.651
0.747 0.877 0.721 0.801 0.666
0.702 0.855 0.704 0.766 0.616
0.841 0.917 0.811 0.854 0.778
0.063 0.023 0.091 0.090 0.128
0.092 0.088 0.126 0.133 0.177
0.067 0.022 0.096 0.069 0.125
0.067 0.021 0.091 0.073 0.140
0.039 0.015 0.068 0.050 0.080
4.08 5.05 3.15 7.01 7.47 7.59 7.87 11.04 11.06 11.05
4.40 2.70 6.85 6.31 9.05
3.08 2.13 6.81 5.95 8.89
2.74 1.77 5.36 4.51 7.42
25.81 23.56 22.30 21.84 22.62
27.01 24.73 23.26 22.50 23.49
28.41 24.86 22.70 22.71 23.56
30.02 27.87 25.63 24.64 26.17
0.863 0.813 0.760 0.749 0.766
0.802 0.753 0.724 0.715 0.733
0.822 0.780 0.746 0.733 0.747
0.836 0.749 0.695 0.712 0.730
0.902 0.868 0.829 0.811 0.832
0.024 0.043 0.073 0.098 0.073
0.045 0.076 0.098 0.117 0.093
0.033 0.050 0.076 0.094 0.077
0.024 0.050 0.085 0.090 0.076
0.017 0.029 0.049 0.066 0.059
2.64 3.54 5.02 9.24 6.10
4.43 4.88 5.77 8.80 6.36
2.92 3.53 4.66 8.87 6.50
2.00 2.58 4.27 8.78 5.50
1.62 2.18 3.61 7.42 4.61
15.02 20.51 15.54 18.41 17.65
16.04 21.91 16.89 19.14 17.91
16.56 21.23 16.96 19.00 18.45
17.44 23.45 17.69 20.39 20.35
0.534 0.690 0.531 0.557 0.537
0.557 0.662 0.562 0.596 0.575
0.581 0.670 0.597 0.597 0.556
0.579 0.629 0.588 0.571 0.554
0.655 0.767 0.653 0.676 0.687
0.284 0.123 0.292 0.244 0.253
0.305 0.147 0.300 0.219 0.235
0.236 0.123 0.230 0.210 0.238
0.206 0.127 0.226 0.193 0.201
0.186 0.085 0.204 0.157 0.146
22.53 19.50 17.35 29.30 29.68
21.23 21.33 16.44 29.87 26.42
20.02 24.39 15.93 22.57 20.72
19.19 15.71 15.48 23.04 21.78
19.11 11.87 14.70 20.48 18.57
20.62 22.70 21.68 18.55 18.90
23.47 24.85 24.28 20.41 21.49
23.20 24.89 24.04 20.23 20.74
25.79 28.12 26.51 23.12 23.73
0.767 0.806 0.784 0.649 0.628
0.689 0.735 0.714 0.635 0.625
0.743 0.781 0.766 0.681 0.682
0.717 0.745 0.742 0.663 0.626
0.829 0.867 0.842 0.773 0.763
0.069 0.049 0.060 0.147 0.174
0.143 0.093 0.118 0.191 0.209
0.077 0.057 0.061 0.137 0.141
0.078 0.056 0.061 0.136 0.145
0.048 0.030 0.042 0.091 0.096
4.44 8.66 4.93 8.12 4.76 6.13 10.21 10.69 18.11 22.11
4.22 4.56 4.75 9.27 12.46
3.68 3.25 3.70 9.40 16.20
3.00 2.43 3.30 7.37 10.19
21.08 19.59 18.64 18.42 16.76
21.85 20.71 19.10 19.46 17.85
22.19 20.10 19.39 19.07 18.11
24.45 22.69 21.68 21.26 19.42
0.719 0.637 0.619 0.615 0.477
0.698 0.659 0.629 0.623 0.566
0.712 0.672 0.612 0.626 0.591
0.683 0.619 0.604 0.603 0.587
0.804 0.766 0.740 0.728 0.666
0.102 0.147 0.181 0.176 0.385
0.115 0.153 0.186 0.183 0.329
0.101 0.123 0.172 0.160 0.266
0.094 0.136 0.156 0.166 0.256
0.067 0.083 0.111 0.115 0.222
6.37 6.31 8.37 8.30 15.10 13.65 14.69 12.72 28.90 27.80
6.68 7.72 13.21 11.46 29.10
4.91 7.10 11.29 10.35 27.38
4.38 5.66 10.37 9.54 24.04
22.96 21.68 19.99 20.25 18.47
25.45 23.37 21.95 21.11 19.04
25.93 23.90 21.76 21.55 18.66
27.83 26.61 24.18 23.60 21.31
0.808 0.769 0.639 0.680 0.599
0.742 0.706 0.663 0.685 0.620
0.806 0.761 0.697 0.716 0.605
0.805 0.728 0.677 0.710 0.583
0.871 0.842 0.794 0.789 0.726
0.052 0.070 0.148 0.153 0.207
0.083 0.107 0.158 0.147 0.198
0.047 0.075 0.109 0.130 0.176
0.043 0.070 0.108 0.116 0.179
0.033 0.047 0.079 0.103 0.130
3.95 5.29 5.15 7.84 8.21 7.73 15.55 14.77 19.10 15.31
3.85 5.92 6.40 14.97 13.48
2.81 3.73 5.71 12.29 12.68
2.35 3.38 5.01 11.95 12.26
Real measured dataset S36 21.00 21.59
19.23
20.31
22.18
0.755 0.803
0.732
0.776
0.806
0.052 0.051
0.079
0.067
0.043
5.17
5.45
5.61
4.12
5.38
trained and evaluated within the same scene, this setting measures spectrum-synthesis fidelity without testing cross-scene generalization. The Advantage Is Consistent Across Scenes. CoRFPR leads on every scene and metric, covering all 144 scene-metric comparisons. Thus, the aggregate improvement is not driven by a few favorable scenes. Relative to the strongest baseline for each scene and metric, the average gains are 2.02 dB in PSNR and 0.073 in SSIM, with a 1.07◦ reduction in AoA error. The largest PSNR gain is 2.74 dB on classroom scene S6, while the largest SSIM gain is 0.112 on corridor scene S20. The margin is smallest on measured scene S36, where CoRFPR reaches 0.806 SSIM against 0.803 for GSRF. Scene Difficulty Is Consistent Across Methods. Per-scene SSIM is strongly correlated between CoRFPR and each baseline, with r = 0.93–0.98. Corridors are among the most difficult scenes, with CoRFPR achieving 0.653–0.767 SSIM, followed by laboratories. Several bedrooms and classrooms are easier, reaching 0.917 SSIM on S7. These patterns suggest that scene properties contribute substantially to performance variation across methods. Scene Variants Produce Substantial Variation. Across categories, the SSIM spread among five variants is 0.09–0.16, comparable to or larger than the average 0.073 gap between CoRFPR and the strongest baseline. Layouts and clutter therefore matter even within a category, motivating the scene-level breakdown rather than relying on category averages alone. E.2.2
C ROSS -S CENE VARIANT R ESULTS
Table 9 and Table 10 break down the CoRFPR results in Table 2 by scene, using the same M = 32 target-scene references. They compare few-shot baselines and baselines transferred from another scene in the same category, respectively. These tables report scene-level summaries, whereas Table 2 aggregates over folds; unweighted averages of their rows need not reproduce the main-table means. Comparison with the per-scene fits in §E.2.1 shows the combined effect of sparse references and unseen-scene inference. Per-Scene Baselines Depend Strongly on Target-Scene Fitting. Across the 35 simulated scenes, moving from full per-scene fitting to 32 target-scene references reduces baseline PSNR by 6.2–8.0 dB. NeRF2 falls from 21.59 to 13.75 dB, GSRF from 20.41 to 12.40 dB, and WRF-GS from 22.01 32
Table 9: Scene-level results for unseen scene variants with few-shot baselines. Each scene is evaluated by the within-category fold that excludes it using M =32 references. PSNR↑
SSIM↑
AoA err (◦ )↓
NMSE↓
Scene NeRF2 -FS GSRF-FS WRF-GS-FS GRaF CoRFPR NeRF2 -FS GSRF-FS WRF-GS-FS GRaF CoRFPR NeRF2 -FS GSRF-FS WRF-GS-FS GRaF CoRFPR NeRF2 -FS GSRF-FS WRF-GS-FS GRaF CoRFPR Conference S1 15.18 S2 15.26 S3 12.80 S4 14.11 S5 13.87 Classroom S6 15.71 S7 14.25 S8 14.80 S9 13.59 S10 10.99 Bedroom S11 14.89 S12 15.03 S13 14.53 S14 14.73 S15 14.23 Corridor S16 12.04 S17 14.16 S18 11.99 S19 11.76 S20 12.93 Open office S21 14.00 S22 13.59 S23 15.05 S24 11.90 S25 13.15 Lab S26 14.17 S27 14.35 S28 13.65 S29 13.56 S30 11.65 Lounge S31 15.11 S32 14.27 S33 13.92 S34 12.89 S35 13.12
12.09 11.66 12.16 11.99 12.15
17.13 17.12 15.68 15.17 14.92
15.53 13.83 13.73 14.10 14.45
28.44 25.61 21.12 20.97 19.10
0.431 0.480 0.371 0.444 0.412
0.302 0.301 0.336 0.336 0.329
0.538 0.549 0.485 0.484 0.473
0.487 0.439 0.441 0.444 0.468
0.876 0.811 0.710 0.712 0.643
0.499 0.427 0.657 0.491 0.538
0.805 0.850 0.687 0.694 0.740
0.354 0.287 0.364 0.380 0.424
0.418 0.558 0.518 0.460 0.452
0.023 0.069 0.128 0.117 0.227
14.40 16.48 25.35 19.19 23.96
51.26 50.03 54.77 52.14 55.71
16.62 14.66 21.14 17.87 23.81
20.27 26.26 25.87 25.50 27.87
2.10 4.40 8.17 7.54 14.92
11.60 11.71 11.45 12.25 11.73
17.67 17.83 16.47 17.13 15.49
15.25 13.61 13.53 14.79 13.47
25.62 30.40 24.42 25.27 22.89
0.456 0.412 0.481 0.426 0.300
0.284 0.281 0.304 0.325 0.323
0.562 0.566 0.532 0.513 0.463
0.500 0.443 0.465 0.449 0.433
0.824 0.897 0.787 0.803 0.746
0.420 0.644 0.506 0.786 1.348
0.798 0.944 0.836 0.853 0.845
0.275 0.313 0.380 0.360 0.445
0.451 0.668 0.659 0.579 0.639
0.042 0.021 0.082 0.072 0.104
12.99 23.38 16.42 19.83 34.94
48.06 67.48 63.04 68.48 57.91
12.33 14.24 19.94 17.20 23.82
15.24 24.85 23.45 26.54 25.16
2.78 1.92 5.33 4.96 7.67
14.72 12.99 13.39 13.62 12.72
16.96 15.81 16.59 16.31 15.51
13.92 13.10 13.97 14.24 13.78
29.28 26.74 23.94 22.29 25.45
0.446 0.458 0.433 0.462 0.428
0.392 0.339 0.375 0.381 0.339
0.514 0.472 0.512 0.510 0.468
0.415 0.408 0.430 0.442 0.421
0.888 0.848 0.785 0.748 0.806
0.577 0.533 0.566 0.504 0.567
0.533 0.690 0.622 0.573 0.713
0.407 0.434 0.384 0.374 0.423
0.641 0.706 0.590 0.507 0.606
0.020 0.034 0.070 0.106 0.067
16.31 14.94 19.17 17.06 16.51
34.36 39.90 35.24 35.44 48.30
20.66 19.30 16.97 18.41 23.63
23.42 27.25 25.72 28.48 25.01
1.65 2.20 3.80 8.42 4.56
11.68 14.16 11.25 13.59 12.94
12.92 16.33 13.86 14.54 14.46
13.19 14.10 12.79 13.53 13.56
15.36 21.90 14.81 19.69 18.21
0.402 0.411 0.349 0.290 0.342
0.349 0.395 0.318 0.350 0.360
0.420 0.487 0.445 0.421 0.417
0.460 0.441 0.421 0.383 0.384
0.573 0.712 0.532 0.643 0.610
0.621 0.633 0.648 0.871 0.651
0.579 0.561 0.681 0.521 0.619
0.494 0.399 0.422 0.454 0.457
0.477 0.594 0.590 0.558 0.559
0.287 0.116 0.348 0.176 0.270
31.91 40.63 24.90 57.47 38.39
48.11 55.37 46.52 48.95 61.16
31.30 33.01 21.72 43.13 34.75
27.52 46.88 24.85 52.41 45.70
22.37 13.23 19.94 21.03 24.02
11.76 13.04 10.79 11.03 12.03
17.30 17.56 17.91 14.31 15.37
14.71 15.03 14.02 12.75 14.26
25.19 27.33 25.77 20.12 20.93
0.405 0.452 0.474 0.374 0.380
0.348 0.398 0.279 0.296 0.349
0.550 0.559 0.576 0.465 0.464
0.502 0.478 0.484 0.436 0.444
0.813 0.850 0.820 0.686 0.686
0.562 0.767 0.486 0.817 0.633
0.767 0.682 0.951 0.830 0.783
0.271 0.343 0.269 0.466 0.438
0.478 0.536 0.598 0.677 0.485
0.055 0.036 0.047 0.182 0.191
16.61 15.52 13.85 27.30 43.08
53.11 42.97 63.81 61.17 62.38
15.08 16.49 11.32 23.52 29.35
17.44 25.00 21.62 27.91 32.67
3.24 2.55 3.35 12.10 15.18
13.07 12.21 12.31 12.02 11.94
15.49 16.35 15.57 15.08 12.86
14.11 13.76 13.94 14.53 12.93
23.97 22.04 20.99 20.31 16.59
0.444 0.447 0.375 0.413 0.321
0.339 0.326 0.339 0.333 0.315
0.496 0.514 0.503 0.478 0.375
0.473 0.435 0.455 0.449 0.394
0.775 0.743 0.716 0.699 0.562
0.626 0.515 0.627 0.576 0.928
0.617 0.697 0.686 0.707 0.758
0.459 0.312 0.437 0.400 0.720
0.607 0.515 0.553 0.430 0.679
0.077 0.094 0.136 0.138 0.403
16.92 18.83 26.91 30.32 49.25
34.46 48.92 50.64 52.15 58.62
18.49 16.06 23.77 23.79 45.54
22.66 26.43 27.24 32.41 46.01
4.40 5.80 11.21 9.70 28.68
12.87 12.10 12.76 12.98 13.23
17.53 16.21 15.87 15.12 14.22
14.45 12.85 12.81 13.09 12.98
27.36 25.84 23.84 22.31 20.48
0.501 0.447 0.449 0.387 0.381
0.332 0.295 0.358 0.347 0.372
0.562 0.497 0.509 0.485 0.437
0.452 0.398 0.420 0.392 0.403
0.858 0.827 0.776 0.733 0.703
0.502 0.615 0.644 0.736 0.618
0.701 0.789 0.666 0.626 0.534
0.313 0.391 0.419 0.422 0.472
0.550 0.758 0.844 0.682 0.634
0.036 0.052 0.086 0.129 0.150
14.18 18.40 17.37 34.83 40.61
42.54 54.48 47.18 46.72 47.04
12.62 17.10 16.07 31.17 36.32
17.87 26.83 26.13 40.58 36.94
2.48 3.41 5.11 11.50 12.18
Table 10: Scene-level results for unseen scene variants with transfer baselines. The three transfer baselines fine-tune source-scene models using the same M =32 references; GRaF and CoRFPR are repeated unchanged from the few-shot comparison. PSNR↑
SSIM↑
AoA err (◦ )↓
NMSE↓
Scene NeRF2 -TR GSRF-TR WRF-GS-TR GRaF CoRFPR NeRF2 -TR GSRF-TR WRF-GS-TR GRaF CoRFPR NeRF2 -TR GSRF-TR WRF-GS-TR GRaF CoRFPR NeRF2 -TR GSRF-TR WRF-GS-TR GRaF CoRFPR Conference S1 17.33 S2 14.50 S3 12.84 S4 12.19 S5 13.07 Classroom S6 12.28 S7 14.90 S8 13.51 S9 12.53 S10 13.23 Bedroom S11 11.52 S12 13.60 S13 13.16 S14 12.89 S15 13.60 Corridor S16 10.67 S17 13.15 S18 12.59 S19 11.61 S20 14.09 Open office S21 11.33 S22 13.88 S23 13.69 S24 11.54 S25 12.45 Lab S26 12.58 S27 13.70 S28 13.67 S29 12.60 S30 11.79 Lounge S31 12.25 S32 13.54 S33 13.05 S34 12.28 S35 12.78
19.22 11.14 12.83 12.45 13.85
20.94 12.47 13.22 12.69 13.71
15.53 13.83 13.73 14.10 14.45
28.44 25.61 21.12 20.97 19.10
0.504 0.442 0.375 0.368 0.397
0.626 0.362 0.428 0.405 0.435
0.653 0.397 0.429 0.449 0.480
0.487 0.439 0.441 0.444 0.468
0.876 0.811 0.710 0.712 0.643
0.319 0.483 0.693 0.766 0.661
0.202 1.004 0.725 0.683 0.591
0.139 0.710 0.590 0.647 0.523
0.418 0.558 0.518 0.460 0.452
0.023 0.069 0.128 0.117 0.227
7.11 19.55 24.05 30.25 30.71
13.57 51.60 30.46 42.80 33.75
11.58 23.46 22.72 28.89 30.66
20.27 26.26 25.87 25.50 27.87
2.10 4.40 8.17 7.54 14.92
11.21 12.25 13.11 11.99 12.41
12.87 13.84 13.92 12.14 13.56
15.25 13.61 13.53 14.79 13.47
25.62 30.40 24.42 25.27 22.89
0.370 0.412 0.391 0.351 0.427
0.360 0.381 0.441 0.353 0.400
0.444 0.460 0.474 0.402 0.453
0.500 0.443 0.465 0.449 0.433
0.824 0.897 0.787 0.803 0.746
0.864 0.533 0.635 1.004 0.712
0.950 0.942 0.717 0.994 0.885
0.735 0.611 0.577 0.844 0.528
0.451 0.668 0.659 0.579 0.639
0.042 0.021 0.082 0.072 0.104
20.48 19.59 35.24 31.34 27.13
52.78 48.85 33.33 47.57 47.20
26.97 42.60 27.19 33.40 43.14
15.24 24.85 23.45 26.54 25.16
2.78 1.92 5.33 4.96 7.67
13.03 11.18 14.14 13.02 13.43
13.61 11.95 13.48 11.69 14.17
13.92 13.10 13.97 14.24 13.78
29.28 26.74 23.94 22.29 25.45
0.294 0.398 0.354 0.387 0.423
0.425 0.353 0.455 0.425 0.437
0.427 0.397 0.448 0.384 0.447
0.415 0.408 0.430 0.442 0.421
0.888 0.848 0.785 0.748 0.806
1.188 0.747 0.698 0.675 0.651
0.876 1.120 0.665 0.628 0.672
0.657 0.887 0.616 0.857 0.526
0.641 0.706 0.590 0.507 0.606
0.020 0.034 0.070 0.106 0.067
37.81 20.39 24.47 28.13 28.23
33.11 50.38 25.72 37.98 38.97
39.07 27.36 24.86 34.26 34.19
23.42 27.25 25.72 28.48 25.01
1.65 2.20 3.80 8.42 4.56
12.15 11.45 13.15 12.22 14.77
13.16 13.70 14.37 8.64 14.65
13.19 14.10 12.79 13.53 13.56
15.36 21.90 14.81 19.69 18.21
0.316 0.376 0.331 0.295 0.433
0.399 0.372 0.442 0.380 0.483
0.460 0.465 0.478 0.315 0.467
0.460 0.441 0.421 0.383 0.384
0.573 0.712 0.532 0.643 0.610
0.848 0.757 0.530 0.880 0.521
0.553 1.070 0.485 0.776 0.466
0.472 0.573 0.364 1.781 0.417
0.477 0.594 0.590 0.558 0.559
0.287 0.116 0.348 0.176 0.270
44.73 46.22 24.68 60.44 40.30
39.26 59.26 27.10 67.37 38.73
31.57 56.93 20.47 80.24 33.15
27.52 46.88 24.85 52.41 45.70
22.37 13.23 19.94 21.03 24.02
12.27 12.12 11.87 12.22 13.19
14.40 16.19 13.83 12.39 12.58
14.71 15.03 14.02 12.75 14.26
25.19 27.33 25.77 20.12 20.93
0.338 0.446 0.417 0.336 0.367
0.400 0.402 0.402 0.399 0.420
0.492 0.512 0.472 0.438 0.412
0.502 0.478 0.484 0.436 0.444
0.813 0.850 0.820 0.686 0.686
1.082 0.711 0.631 0.914 0.716
0.793 0.923 0.886 0.728 0.645
0.523 0.411 0.548 0.675 0.631
0.478 0.536 0.598 0.677 0.485
0.055 0.036 0.047 0.182 0.191
33.82 20.52 17.44 34.45 54.60
40.15 41.86 41.03 40.31 50.05
23.01 21.30 20.72 28.58 43.77
17.44 25.00 21.62 27.91 32.67
3.24 2.55 3.35 12.10 15.18
12.46 11.74 12.97 12.76 12.31
13.90 13.73 13.28 12.83 12.54
14.11 13.76 13.94 14.53 12.93
23.97 22.04 20.99 20.31 16.59
0.381 0.442 0.391 0.358 0.337
0.407 0.389 0.427 0.417 0.394
0.446 0.470 0.450 0.441 0.401
0.473 0.435 0.455 0.449 0.394
0.775 0.743 0.716 0.699 0.562
0.826 0.579 0.612 0.702 0.871
0.765 0.854 0.689 0.675 0.753
0.565 0.476 0.586 0.613 0.673
0.607 0.515 0.553 0.430 0.679
0.077 0.094 0.136 0.138 0.403
22.52 21.06 26.96 39.80 49.67
36.17 60.21 32.39 48.43 60.86
24.81 36.45 29.86 33.79 61.93
22.66 26.43 27.24 32.41 46.01
4.40 5.80 11.21 9.70 28.68
11.81 11.12 13.38 12.42 12.51
13.30 12.59 13.49 12.28 13.54
14.45 12.85 12.81 13.09 12.98
27.36 25.84 23.84 22.31 20.48
0.337 0.388 0.400 0.362 0.394
0.384 0.358 0.433 0.397 0.417
0.419 0.411 0.472 0.413 0.435
0.452 0.398 0.420 0.392 0.403
0.858 0.827 0.776 0.733 0.703
0.893 0.675 0.759 0.869 0.692
0.914 1.034 0.731 0.743 0.692
0.686 0.750 0.660 0.775 0.531
0.550 0.758 0.844 0.682 0.634
0.036 0.052 0.086 0.129 0.150
25.26 30.20 21.84 49.40 50.71
49.74 54.78 31.89 63.11 63.57
26.70 31.82 22.80 82.30 39.64
17.87 26.83 26.13 40.58 36.94
2.48 3.41 5.11 11.50 12.18
to 15.85 dB. Feed-forward GRaF also falls from 22.00 to 13.85 dB under sparse references. By comparison, CoRFPR decreases by 1.34 dB, from 24.33 to 22.99 dB. The Advantage Persists Across Scenes. CoRFPR leads on all 35 scenes against both few-shot and transfer baselines in these breakdowns. Its average PSNR gain over the strongest few-shot competitor on each scene is 7.13 dB, with a minimum of 0.94 dB on corridor scene S18. Against the strongest transfer competitor, the average gain is 8.72 dB, with a minimum of 0.44 dB, again on S18. The feed-forward and power-refinement variants are compared separately in Table 13. Source-Scene Transfer Provides Limited Benefit. Initializing a baseline from another scene in the same category and adapting it with the 32 references does not consistently help. GSRF rises from 12.40 to 12.69 dB and improves on 20 scenes. NeRF2 loses 0.77 dB on average and improves on only 9 scenes; WRF-GS loses 2.43 dB and improves on only 4. A field fitted to one scene is therefore not generally a reliable initialization for another, even within the same category. 33
Table 11: Scene-level results with the entire scene category unseen using few-shot baselines. Each scene is evaluated by the cross-category fold that excludes its category using M =32 references. Few-shot baseline results are unchanged from the unseen-variant setting. PSNR↑
SSIM↑
AoA err (◦ )↓
NMSE↓
Scene NeRF2 -FS GSRF-FS WRF-GS-FS GRaF CoRFPR NeRF2 -FS GSRF-FS WRF-GS-FS GRaF CoRFPR NeRF2 -FS GSRF-FS WRF-GS-FS GRaF CoRFPR NeRF2 -FS GSRF-FS WRF-GS-FS GRaF CoRFPR Conference S1 15.18 S2 15.26 S3 12.80 S4 14.11 S5 13.87 Classroom S6 15.71 S7 14.25 S8 14.80 S9 13.59 S10 10.99 Bedroom S11 14.89 S12 15.03 S13 14.53 S14 14.73 S15 14.23 Corridor S16 12.04 S17 14.16 S18 11.99 S19 11.76 S20 12.93 Open office S21 14.00 S22 13.59 S23 15.05 S24 11.90 S25 13.15 Lab S26 14.17 S27 14.35 S28 13.65 S29 13.56 S30 11.65 Lounge S31 15.11 S32 14.27 S33 13.92 S34 12.89 S35 13.12
12.09 11.66 12.16 11.99 12.15
17.13 17.12 15.68 15.17 14.92
14.55 14.48 14.32 14.09 14.60
28.24 25.72 20.77 20.34 19.22
0.431 0.480 0.371 0.444 0.412
0.302 0.301 0.336 0.336 0.329
0.538 0.549 0.485 0.484 0.473
0.461 0.451 0.451 0.448 0.471
0.871 0.823 0.698 0.698 0.647
0.499 0.427 0.657 0.491 0.538
0.805 0.850 0.687 0.694 0.740
0.354 0.287 0.364 0.380 0.424
0.524 0.478 0.459 0.465 0.437
0.024 0.060 0.141 0.151 0.217
14.40 16.48 25.35 19.19 23.96
51.26 50.03 54.77 52.14 55.71
16.62 14.66 21.14 17.87 23.81
21.53 26.08 24.54 24.14 27.42
2.10 4.14 9.05 8.13 14.98
11.60 11.71 11.45 12.25 11.73
17.67 17.83 16.47 17.13 15.49
14.04 13.47 13.71 14.31 12.86
25.68 30.43 23.79 26.52 22.95
0.456 0.412 0.481 0.426 0.300
0.284 0.281 0.304 0.325 0.323
0.562 0.566 0.532 0.513 0.463
0.471 0.441 0.472 0.450 0.424
0.820 0.897 0.767 0.819 0.742
0.420 0.644 0.506 0.786 1.348
0.798 0.944 0.836 0.853 0.845
0.275 0.313 0.380 0.360 0.445
0.637 0.725 0.652 0.673 0.757
0.044 0.020 0.098 0.068 0.101
12.99 23.38 16.42 19.83 34.94
48.06 67.48 63.04 68.48 57.91
12.33 14.24 19.94 17.20 23.82
20.36 24.44 22.46 27.22 24.91
2.75 1.92 5.84 4.96 7.58
14.72 12.99 13.39 13.62 12.72
16.96 15.81 16.59 16.31 15.51
12.80 13.59 14.20 13.67 13.24
29.49 26.74 25.02 22.67 25.55
0.446 0.458 0.433 0.462 0.428
0.392 0.339 0.375 0.381 0.339
0.514 0.472 0.512 0.510 0.468
0.371 0.410 0.437 0.426 0.400
0.892 0.831 0.811 0.751 0.813
0.577 0.533 0.566 0.504 0.567
0.533 0.690 0.622 0.573 0.713
0.407 0.434 0.384 0.374 0.423
0.796 0.647 0.542 0.560 0.677
0.019 0.035 0.058 0.105 0.065
16.31 14.94 19.17 17.06 16.51
34.36 39.90 35.24 35.44 48.30
20.66 19.30 16.97 18.41 23.63
37.57 27.51 26.03 29.84 27.73
1.65 2.17 3.79 8.39 4.59
11.68 14.16 11.25 13.59 12.94
12.92 16.33 13.86 14.54 14.46
13.16 14.97 13.72 13.55 13.80
15.26 22.14 15.76 19.12 18.06
0.402 0.411 0.349 0.290 0.342
0.349 0.395 0.318 0.350 0.360
0.420 0.487 0.445 0.421 0.417
0.457 0.457 0.450 0.401 0.395
0.569 0.695 0.566 0.624 0.603
0.621 0.633 0.648 0.871 0.651
0.579 0.561 0.681 0.521 0.619
0.494 0.399 0.422 0.454 0.457
0.488 0.485 0.457 0.552 0.530
0.294 0.121 0.286 0.195 0.278
31.91 40.63 24.90 57.47 38.39
48.11 55.37 46.52 48.95 61.16
31.30 33.01 21.72 43.13 34.75
28.04 42.46 23.63 50.13 45.93
22.30 13.30 17.03 21.74 26.50
11.76 13.04 10.79 11.03 12.03
17.30 17.56 17.91 14.31 15.37
13.93 15.04 14.27 12.58 14.08
23.42 26.67 24.01 19.76 19.96
0.405 0.452 0.474 0.374 0.380
0.348 0.398 0.279 0.296 0.349
0.550 0.559 0.576 0.465 0.464
0.478 0.474 0.492 0.439 0.442
0.768 0.839 0.788 0.676 0.665
0.562 0.767 0.486 0.817 0.633
0.767 0.682 0.951 0.830 0.783
0.271 0.343 0.269 0.466 0.438
0.567 0.526 0.544 0.712 0.499
0.086 0.039 0.057 0.193 0.245
16.61 15.52 13.85 27.30 43.08
53.11 42.97 63.81 61.17 62.38
15.08 16.49 11.32 23.52 29.35
20.60 24.14 21.45 27.74 37.30
3.74 2.66 3.41 11.42 16.43
13.07 12.21 12.31 12.02 11.94
15.49 16.35 15.57 15.08 12.86
12.71 13.70 13.93 13.71 12.02
23.69 22.18 20.58 20.47 16.63
0.444 0.447 0.375 0.413 0.321
0.339 0.326 0.339 0.333 0.315
0.496 0.514 0.503 0.478 0.375
0.441 0.432 0.460 0.440 0.382
0.776 0.743 0.682 0.688 0.547
0.626 0.515 0.627 0.576 0.928
0.617 0.697 0.686 0.707 0.758
0.459 0.312 0.437 0.400 0.720
0.823 0.537 0.555 0.542 0.833
0.080 0.096 0.155 0.137 0.404
16.92 18.83 26.91 30.32 49.25
34.46 48.92 50.64 52.15 58.62
18.49 16.06 23.77 23.79 45.54
26.47 26.25 27.83 30.23 45.77
4.50 5.83 11.18 9.60 28.76
12.87 12.10 12.76 12.98 13.23
17.53 16.21 15.87 15.12 14.22
13.53 12.99 13.30 12.57 12.82
27.37 25.84 23.55 21.71 20.24
0.501 0.447 0.449 0.387 0.381
0.332 0.295 0.358 0.347 0.372
0.562 0.497 0.509 0.485 0.437
0.414 0.386 0.426 0.383 0.392
0.858 0.823 0.766 0.718 0.697
0.502 0.615 0.644 0.736 0.618
0.701 0.789 0.666 0.626 0.534
0.313 0.391 0.419 0.422 0.472
0.644 0.748 0.744 0.762 0.661
0.036 0.053 0.091 0.126 0.148
14.18 18.40 17.37 34.83 40.61
42.54 54.48 47.18 46.72 47.04
12.62 17.10 16.07 31.17 36.32
22.92 26.49 25.12 41.10 38.86
2.45 3.39 5.16 11.47 12.24
Table 12: Scene-level results with the entire scene category unseen using transfer baselines. The three transfer baselines fine-tune models from outside the target category using the same M =32 references; GRaF and CoRFPR are unchanged from the few-shot comparison. PSNR↑
SSIM↑
AoA err (◦ )↓
NMSE↓
Scene NeRF2 -TR GSRF-TR WRF-GS-TR GRaF CoRFPR NeRF2 -TR GSRF-TR WRF-GS-TR GRaF CoRFPR NeRF2 -TR GSRF-TR WRF-GS-TR GRaF CoRFPR NeRF2 -TR GSRF-TR WRF-GS-TR GRaF CoRFPR Conference S1 12.77 S2 12.77 S3 12.52 S4 12.22 S5 12.51 Classroom S6 14.29 S7 14.30 S8 15.36 S9 14.10 S10 12.76 Bedroom S11 13.93 S12 13.36 S13 16.70 S14 13.83 S15 13.05 Corridor S16 10.48 S17 11.46 S18 10.94 S19 10.13 S20 10.82 Open office S21 13.80 S22 13.34 S23 13.17 S24 11.73 S25 12.47 Lab S26 13.22 S27 16.17 S28 13.23 S29 13.53 S30 11.36 Lounge S31 13.08 S32 12.98 S33 12.45 S34 12.67 S35 12.19
13.22 13.39 13.11 12.45 13.00
12.21 10.72 10.95 12.40 10.35
14.55 14.48 14.32 14.09 14.60
28.24 25.72 20.77 20.34 19.22
0.348 0.338 0.371 0.369 0.352
0.408 0.415 0.421 0.404 0.408
0.404 0.358 0.406 0.436 0.375
0.461 0.451 0.451 0.448 0.471
0.871 0.823 0.698 0.698 0.647
0.782 0.681 0.791 0.761 0.792
0.729 0.657 0.646 0.682 0.706
0.924 1.102 1.047 0.705 1.150
0.524 0.478 0.459 0.465 0.437
0.024 0.060 0.141 0.151 0.217
30.64 31.41 33.03 29.71 39.05
38.85 35.87 39.49 42.88 37.67
27.99 39.86 36.36 29.38 32.44
21.53 26.08 24.54 24.14 27.42
2.10 4.14 9.05 8.13 14.98
12.06 12.10 13.99 12.07 12.45
9.15 12.53 14.46 11.96 12.75
14.04 13.47 13.71 14.31 12.86
25.68 30.43 23.79 26.52 22.95
0.453 0.422 0.488 0.415 0.396
0.416 0.389 0.472 0.379 0.420
0.317 0.425 0.484 0.400 0.450
0.471 0.441 0.472 0.450 0.424
0.820 0.897 0.767 0.819 0.742
0.593 0.638 0.424 0.683 0.806
0.916 1.040 0.564 1.089 0.890
1.547 0.798 0.450 0.964 0.706
0.637 0.725 0.652 0.673 0.757
0.044 0.020 0.098 0.068 0.101
17.36 20.96 17.56 23.18 32.34
36.61 35.67 20.56 39.75 35.22
32.80 40.02 26.77 39.46 39.22
20.36 24.44 22.46 27.22 24.91
2.75 1.92 5.84 4.96 7.58
13.41 12.15 17.38 12.52 12.05
14.16 14.15 16.77 12.70 12.71
12.80 13.59 14.20 13.67 13.24
29.49 26.74 25.02 22.67 25.55
0.385 0.379 0.515 0.393 0.361
0.416 0.398 0.573 0.411 0.395
0.457 0.457 0.546 0.415 0.421
0.371 0.410 0.437 0.426 0.400
0.892 0.831 0.811 0.751 0.813
0.719 0.728 0.354 0.595 0.728
0.838 0.959 0.315 0.798 0.938
0.588 0.530 0.341 0.696 0.731
0.796 0.647 0.542 0.560 0.677
0.019 0.035 0.058 0.105 0.065
26.56 24.17 12.01 23.39 27.50
32.28 48.11 11.86 49.58 50.49
36.20 25.34 21.86 32.96 31.31
37.57 27.51 26.03 29.84 27.73
1.65 2.17 3.79 8.39 4.59
12.44 13.06 12.05 12.18 11.97
13.63 14.74 13.88 12.57 13.75
13.16 14.97 13.72 13.55 13.80
15.26 22.14 15.76 19.12 18.06
0.274 0.277 0.270 0.248 0.273
0.422 0.425 0.399 0.373 0.381
0.464 0.493 0.459 0.432 0.462
0.457 0.457 0.450 0.401 0.395
0.569 0.695 0.566 0.624 0.603
0.873 1.106 0.799 1.280 1.084
0.503 0.757 0.614 0.861 0.808
0.400 0.482 0.423 0.660 0.503
0.488 0.485 0.457 0.552 0.530
0.294 0.121 0.286 0.195 0.278
48.77 59.78 45.62 67.11 56.01
35.55 43.08 36.06 41.91 54.19
28.54 39.90 23.04 53.63 56.23
28.04 42.46 23.63 50.13 45.93
22.30 13.30 17.03 21.74 26.50
13.33 14.05 13.33 12.54 13.21
12.71 12.99 13.32 12.18 12.73
13.93 15.04 14.27 12.58 14.08
23.42 26.67 24.01 19.76 19.96
0.430 0.406 0.419 0.375 0.370
0.436 0.446 0.430 0.408 0.421
0.452 0.369 0.471 0.427 0.388
0.478 0.474 0.492 0.439 0.442
0.768 0.839 0.788 0.676 0.665
0.568 0.777 0.708 0.917 0.709
0.612 0.679 0.649 0.661 0.642
0.708 0.704 0.622 0.799 0.635
0.567 0.526 0.544 0.712 0.499
0.086 0.039 0.057 0.193 0.245
24.64 29.96 21.18 32.35 54.64
39.25 41.21 32.94 40.73 49.65
29.89 23.68 24.24 31.25 40.45
20.60 24.14 21.45 27.74 37.30
3.74 2.66 3.41 11.42 16.43
11.99 14.19 12.38 12.10 11.25
13.81 15.00 13.63 12.12 13.42
12.71 13.70 13.93 13.71 12.02
23.69 22.18 20.58 20.47 16.63
0.420 0.526 0.406 0.406 0.345
0.395 0.463 0.399 0.400 0.353
0.468 0.478 0.454 0.413 0.418
0.441 0.432 0.460 0.440 0.382
0.776 0.743 0.682 0.688 0.547
0.713 0.414 0.630 0.576 0.941
0.891 0.536 0.752 0.800 0.974
0.591 0.430 0.507 0.840 0.571
0.823 0.537 0.555 0.542 0.833
0.080 0.096 0.155 0.137 0.404
27.90 17.41 34.57 30.77 57.68
40.05 34.88 46.65 52.29 54.25
20.67 24.65 40.68 26.14 56.97
26.47 26.25 27.83 30.23 45.77
4.50 5.83 11.18 9.60 28.76
12.80 12.47 12.49 11.78 11.63
11.64 12.48 10.72 10.77 11.22
13.53 12.99 13.30 12.57 12.82
27.37 25.84 23.55 21.71 20.24
0.373 0.375 0.382 0.351 0.355
0.404 0.399 0.404 0.378 0.383
0.377 0.399 0.365 0.378 0.382
0.414 0.386 0.426 0.383 0.392
0.858 0.823 0.766 0.718 0.697
0.741 0.763 0.867 0.755 0.797
0.884 0.916 0.873 0.985 0.906
0.889 0.816 1.075 1.103 1.002
0.644 0.748 0.744 0.762 0.661
0.036 0.053 0.091 0.126 0.148
23.60 27.00 26.03 38.69 42.09
31.73 39.34 30.32 44.31 48.45
33.40 34.64 49.79 50.75 49.75
22.92 26.49 25.12 41.10 38.86
2.45 3.39 5.16 11.47 12.24
Harder Scenes Incur Larger Generalization Losses. For CoRFPR , the SSIM decrease from perscene fitting to unseen-variant inference is largest for corridors, from 0.687 to 0.614, and smallest for lounges, from 0.805 to 0.779. At the scene level, the gap is negatively correlated with per-scene SSIM, with r = −0.64. The six largest gaps occur on S18, S30, S5, S24, S16, and S20, which are also among the more difficult scenes under per-scene fitting. Thus, scenes that are difficult with full target-scene data tend to lose more when synthesis relies on sparse references in an unseen scene. E.2.3
C ROSS -S CENE C ATEGORY R ESULTS
Table 3 reports CoRFPR results averaged over seven excluded scene categories. Table 11 and Table 12 provide scene-level comparisons with few-shot and transfer baselines. Their unweighted scene 34
averages may differ from the fold-level means in Table 3. The category comparisons below pool results over each category’s five scenes. Excluding an Entire Category Causes a Small Additional Drop. Relative to unseen-variant evaluation, the unweighted scene summaries for CoRFPR decrease by 0.14 dB in PSNR and 0.007 in SSIM; GRaF changes by 0.18 dB and 0.005, respectively. For CoRFPR , SSIM falls on 24 scenes and rises on 11, with the largest scene-level decrease of 0.045 on S21. The largest category-level change is 0.024 for open offices, while the overall category mean moves from 0.750 to 0.743. Thus, category exclusion causes a modest average loss, although its effect varies by scene. The Margin over GRaF Remains Large. CoRFPR exceeds GRaF on all four metrics in every category. Its category-level SSIM margin ranges from 0.18 on corridors to 0.41 on bedrooms. AoA error ranges from 4.1◦ to 20.2◦ for CoRFPR , compared with 23.9◦ to 38.0◦ for GRaF. Across individual scenes, CoRFPR spans 0.547–0.897 SSIM, showing that absolute synthesis quality still varies despite the consistent margin. Source-Scene Transfer Has Mixed Effects. With target references fixed, using a source scene outside the target category changes PSNR by −0.02 dB for NeRF2 , +0.07 dB for GSRF, and −0.70 dB for WRF-GS relative to within-category transfer. A same-category source therefore provides no consistent advantage. Likewise, Table 3 shows that transfer improves some GSRF metrics but degrades NeRF2 and WRF-GS relative to their few-shot fits. The Advantage Persists Across Scenes. CoRFPR leads on all 35 scenes and all four metrics against both baseline families, covering 280 scene-metric comparisons. The aggregate gain is therefore not concentrated in a few favorable scenes. E.2.4
P OWER R EFINEMENT
This ablation repeats the unseen-variant and unseen-category protocols of Table 2 and Table 3 in an independent run. All variants condition the same pretrained model on M = 32 target-scene references. Feed-forward denotes CoRF and performs no target-scene fitting. Anchor offsets fits one scene-specific power correction per anchor, shared across queries. Query network fits a network that adjusts component powers for each query. Both fits these two corrections and denotes CoRFPR . The pretrained model, anchor positions, and arrival directions remain fixed. Each fitted variant uses 24 references for optimization and 8 for step selection. Setup time includes conditioning and fitting to the selected step; query time is per synthesized spectrum. Table 13: Power-refinement ablation. Mean ± standard deviation over folds. Setting
Variant
Unseen variants
Feed-forward 22.95±1.94 0.740±0.054 0.121±0.050 8.40±3.85 Anchor offsets 23.16±1.98 0.751±0.052 0.116±0.048 8.39±3.85 Query network 22.90±2.11 0.743±0.054 0.124±0.052 8.72±3.94 Both 22.92±2.09 0.744±0.054 0.124±0.052 8.70±4.02
PSNR↑
SSIM↑
NMSE↓
AoA↓
Adapted params. Setup Query 0 0.004M 0.266M 0.270M
0.01 s 8.5 ms 202 s 8.5 ms 47 s 8.5 ms 80 s 8.5 ms
Unseen categories
Feed-forward 22.83±2.69 0.733±0.074 0.124±0.069 8.41±5.18 Anchor offsets 23.01±2.70 0.743±0.074 0.120±0.067 8.40±5.17 Query network 22.84±2.73 0.740±0.074 0.123±0.067 8.67±5.25 Both 22.83±2.72 0.739±0.074 0.124±0.068 8.69±5.32
0 0.004M 0.266M 0.270M
0.01 s 8.5 ms 202 s 8.5 ms 47 s 8.5 ms 80 s 8.5 ms
As shown in Table 13, joint refinement changes PSNR relative to feed-forward CoRF by −0.03 dB on unseen variants and +0.01 dB on unseen categories. The respective 95% confidence intervals are [−0.24, 0.16] and [−0.15, 0.16]. Anchor offsets give the largest gains, 0.21 and 0.18 dB, but require 202 seconds of setup. The query network peaks near 50 optimization steps and then overfits, falling to roughly 21.16 dB after 3,000 steps. The selection gate rejects a fitted correction in 28 of 70 scene evaluations. These results show that the cross-scene advantage is largely present in feed-forward CoRF. Power refinement provides limited additional accuracy at a one-time setup cost, while leaving per-query rendering time unchanged. E.3
R ELIABILITY
We first examine whether the variance from feed-forward CoRF ranks query difficulty, has an appropriate overall scale, and localizes error within a spectrum. We then test how optional power refinement affects these properties. 35
Figure 6: Predictions at low and high uncertainty. P1–P5 are low-uncertainty queries and P6–P10 are high-uncertainty queries, paired by scene. Rows show ground truth (GT), predicted mean (Ours), normalized squared error, and predicted standard deviation. Table 14: Uncertainty scale and within-spectrum error localization. Results are mean ± standard deviation across folds. Setting Unseen scene variants Unseen scene categories
E.3.1
rpix
cov@1σ
cov@2σ
cσ
−0.027 ±0.034 −0.034 ±0.088
0.809 ±0.068 0.803 ±0.111
0.931 ±0.033 0.925 ±0.062
1.11 ±0.28 1.12 ±0.44
F EED -F ORWARD R ELIABILITY
The main paper reports the aggregate query-ranking result in §4.4. Figure 6 illustrates the corresponding low- and high-uncertainty predictions. Query-level separation. In the first unseen-variant fold, the 10% of 6,436 test queries with the lowest propagated variance achieve mean SSIM 0.889 and NMSE 0.017. The highest-variance 10% achieve 0.553 SSIM and 0.326 NMSE. The separation persists among queries whose ground-truth spectra have at least three lobes: mean SSIM is 0.862 in the low-uncertainty group and 0.539 in the high-uncertainty group. For the figure, the two groups are matched by scene, and each column shows the median-SSIM query from that scene within its uncertainty group. The low-uncertainty examples achieve 0.86–0.89 SSIM and 0.012–0.034 NMSE. The high-uncertainty examples achieve 0.40– 0.72 SSIM and 0.11–0.34 NMSE, with larger errors around the dominant lobes. Uncertainty scale and spatial localization. We assess the overall scale of predicted uncertainty and whether it identifies angular cells with larger amplitude error ε = Sb − S. For each spectrum, rpix is the correlation between predicted standard deviation σ and |ε| over 400 randomly sampled angular cells; we then average this correlation across queries. Computing the correlation within each spectrum prevents differences in query difficulty from driving this spatial measure. We also report empirical coverage within 1σ and 2σ, whose Gaussian reference values are 0.68 and 0.95, and the scale factor s ε2 cσ = E 2 , (59) σ where cσ ≈ 1 indicates approximately correct average scale. Table 14 shows that the average uncertainty scale is similar in the two simulated generalization settings. Coverage within 2σ is about 0.93, near the Gaussian reference of 0.95, and the scale factors are 1.11 and 1.12. Coverage within 1σ is higher than the Gaussian reference, at about 0.81 vs. 0.68, indicating that the residual distribution is not exactly Gaussian. Spatial localization is much weaker than query-level ranking. The within-spectrum correlations are approximately −0.03 in both settings, so angular cells assigned higher uncertainty are not systematically those with larger errors in the same spectrum. The propagated variance is therefore most informative as a query-level reliability signal. E.3.2
R ELIABILITY AFTER P OWER R EFINEMENT
We test whether power refinement preserves query-level ranking and amplitude-domain calibration on the first 300 test queries per scene across all 12 folds. In Table 15, Spearman correlates mean predicted variance with mean squared amplitude error across queries, while NLL is pixel-level Gaussian negative log-likelihood. Coverage is the fraction of angular cells with absolute error within one or two predicted standard deviations. The scale factor cσ is defined in Equation 59, and sharpness is the mean predicted standard deviation. 36
Table 15: Amplitude-domain reliability before and after power refinement. Condition
Spearman
NLL
Cov. 1σ
Cov. 2σ
cσ
Sharpness
0.603 0.599 0.523 0.498
-0.630 -0.641 -0.650 -0.741
0.809 0.804 0.819 0.762
0.930 0.929 0.938 0.922
1.10 1.09 1.04 1.20
0.1110 0.1110 0.1217 0.1057
CoRF CoRFPR mean, CoRF variance CoRFPR variance at refined powers CoRFPR variance, 8-reference recalibration
Table 16: Tolerance to array-parameter error on A2 (8 × 2). Uncalibrated results use the corrupted array parameters directly, while calibrated results fit νarr using NC = 64 measurements per scene. Uncalibrated Orientation Spacing ◦
0 2◦ 4◦ 6◦ 9◦ 12◦
— +2/ − 2% +4/ − 3% +7/ − 6% +10/ − 9% +14/ − 12%
Calibrated
PSNR↑ SSIM↑ NMSE↓ AoA↓ PSNR↑ SSIM↑ NMSE↓ AoA↓ 22.32 21.22 19.79 18.45 17.05 16.01
0.701 0.671 0.624 0.570 0.514 0.475
0.095 0.107 0.134 0.175 0.235 0.293
6.73 6.74 7.70 9.01 10.44 12.22
22.21 22.17 22.22 22.09 22.14 22.07
0.701 0.700 0.700 0.697 0.697 0.695
0.096 0.096 0.096 0.097 0.096 0.097
6.73 6.73 6.73 6.78 6.77 6.71
Carrying the feed-forward variance unchanged to the refined mean nearly preserves Spearman correlation, from 0.603 to 0.599, and 1σ coverage, from 0.809 to 0.804. Recomputing variance at refined powers lowers the correlation to 0.523. Recalibration using eight selection references improves NLL but lowers 1σ coverage to 0.762 and raises cσ to 1.20, indicating underestimated error scale. These results favor carrying the feed-forward variance unchanged alongside a powerrefined mean. Converting variance to the power domain by the delta method requires scale factors of about 9.7–11.8; the calibration results here concern normalized amplitude. E.4
T OLERANCE TO A RRAY-PARAMETER E RROR
§4.5 evaluates receiver calibration at one level of geometry mismatch. Here we vary the nominal array error and test how much accuracy calibration recovers. Deployment Requirements for a New Array. The analytic decoder requires receiver element positions, carrier wavelength, and array position and orientation. Element geometry and wavelength are typically known from the array specification, while receiver position is obtained at installation. Orientation may be reported by deployed infrastructure or estimated through calibration. For example, AISG supports reporting antenna azimuth, tilt, and roll (AISG, 2024), while 3GPP specifies positioning information and array-orientation conventions (3GPP, 2024; 2026; 2019). When this metadata is unavailable or inaccurate, the receiver response can be calibrated using measurements collected at known locations, as demonstrated on commercial 5G arrays (Pan et al., 2023). A new array therefore requires either an accurate response specification or receiver-specific calibration measurements. Protocol. We use the unseen 8 × 2 receiver array A2 and perturb its in-plane orientation and element spacing over six levels, from exact geometry to 12◦ of orientation error with +14% and −12% axis scaling. At each level, we evaluate the nominal response and a response calibrated with NC = 64 measurements. The propagation model, M = 32 target-scene references, and test queries remain unchanged; calibration measurements are disjoint from the references and queries. Uncalibrated Accuracy Degrades with Array Error. Table 16 reports the results. As the nominal geometry becomes less accurate, each additional 2◦ –3◦ of orientation error with its corresponding spacing error reduces PSNR by approximately 1.0–1.4 dB. At the largest perturbation, PSNR falls to 16.01 dB, approaching the 14.96 dB obtained by retaining the training-array response for A2 . NMSE increases from 0.095 to 0.293, and AoA error from 6.73◦ to 12.22◦ . Thus, receiver-geometry errors distort rendered spectrum even though the predicted propagation components are unchanged. Calibration Recovers In-Family Geometry Errors. After calibration, PSNR remains between 22.07 and 22.22 dB across the perturbation range. NMSE remains between 0.096 and 0.097, and AoA error between 6.71◦ and 6.78◦ . The injected errors match the fitted calibration family: one inplane rotation and two axis scalings. After fitting, recovered element positions differ from their true 37
positions by 0.010–0.028 wavelengths on average, with maximum displacement below 0.058λ. These results establish recovery within the modeled geometry family, not for arbitrary receiver mismatch. Calibration Becomes More Valuable as Mismatch Increases. Calibration improves PSNR by 0.95 dB at the smallest nonzero perturbation and 6.06 dB at the largest. With exact geometry, it changes PSNR by only −0.11 dB. Calibration therefore adds little when the receiver specification is accurate but becomes more valuable as its geometry becomes less reliable. Calibration Robustness. Under the same in-family perturbation, using NC = 8, 16, 32, and 64 calibration measurements yields 21.96, 22.03, 22.05, and 22.08 dB, respectively, compared with 14.96 dB without calibration and 22.32 dB with exact geometry. Ten random initializations at NC = 64 yield 22.06 ± 0.02 dB. Calibration-measurement SNRs of 20, 10, and 0 dB yield 22.07, 22.00, and 21.33 dB. Thus, the geometric fit is stable across these measurement budgets and initializations, although severe measurement noise reduces accuracy. E.5
C ASE S TUDY: A NGULAR L OCALIZATION WITH S YNTHESIZED S PATIAL S PECTRA
We evaluate whether spectra synthesized from sparse references can augment training for downstream localization; beam management is studied in §4.7. Following Zhao et al. (2023), we train an angular artificial neural network (AANN) with a ResNet-50 backbone and an MLP head. Given a spatial spectrum, it predicts the geometric direction from the receiver to the transmitter as a unit vector. The training loss is cosine distance to the direction computed from the known transmitter and receiver positions. Protocol. Only the M = 32 target-scene reference spectra are available to each synthesis method. Each method generates spectra at the remaining training locations; feed-forward methods condition on the references, while per-scene baselines are fitted from scratch using their few-shot configurations. For synthesis method m, the localization model is trained on M b Dm loc = DGT ∪ Dm ,
(60)
b where DM GT contains the 32 reference spectra and Dm contains spectra synthesized at the remaining training locations. Synthesized spectra retain the transmitter locations and geometric localization labels of their corresponding ground-truth spectra. We compare seven training conditions: the 32 references alone; augmentation with NeRF2 -FS, GSRFFS, WRF-GS-FS, GRaF, or CoRF; and the full ground-truth training set. The evaluation covers seven unseen scene variants in one fold, with one scene from each category. All AANNs use the same architecture, training configuration, and seed. Every condition is evaluated on the same 6,436 ground-truth test spectra, never on synthesized test spectra. For test query π, localization error is gt b loc Eπloc = dang Ω , Ω , (61) π π gt b loc where Ω π is the AANN prediction, Ωπ is the geometric transmitter direction, and dang is spherical angular distance.
Results. As shown in Figure 7, training on the 32 references alone yields mean angular error 28.27◦ . Augmentation with CoRF spectra reduces it to 4.11◦ , close to the 3.05◦ obtained with the full groundtruth training set. This closes 96% of the gap between the reference-only and full-data conditions. The strongest baseline, NeRF2 -FS, achieves 11.36◦ ; GSRF-FS, GRaF, and WRF-GS-FS achieve 12.06◦ , 12.75◦ , and 13.27◦ , respectively. With CoRF augmentation, 77% of test queries fall within 5◦ of the ground-truth direction, compared with 28%–35% for the baselines. CoRF also yields the lowest mean localization error in each of the seven evaluated scenes. Synthesis Fidelity and Localization. Localization labels always use the true geometric directions, so synthesis errors affect the training inputs rather than the labels. Misplaced or distorted spectral lobes create a mismatch between synthesized training inputs and ground-truth test spectra. Across method–scene pairs, localization error is negatively correlated with the SSIM of synthesized training spectra, with Pearson r = −0.73. This association is consistent with improved synthesis fidelity contributing to lower localization error. 38
Localization error (°)
100 10 1 0.1 0.01 M real NeRF2 GSRF WRFGS GRaF
Ours All real
Figure 7: Localization with M = 32 reference spectra and synthesized training spectra. Only the 32 references per scene are treated as available observations; each method synthesizes the remaining training spectra. Localization error is evaluated on the same test queries across unseen scene variants. Overall, augmenting 32 reference spectra with CoRF predictions trains a localization model whose accuracy approaches that obtained with the full ground-truth training set. Augmentation with baselinesynthesized spectra leaves substantially larger localization error.
E.6
S IMULATION - TO -R EAL T RANSFER S TUDY
We transfer the propagation model trained on 35 simulated scenes to measured scene S36, which uses a physical 4 × 4 receiver array. Each method has a budget of 96 target-domain spectra, although the methods use these measurements differently. Both CoRFCAL and CoRFPR,CAL condition on M = 32 references and calibrate the receiver with NC = 64 disjoint spectra. CoRFPR,CAL also reuses both sets for power refinement without additional measurements. Table 17: Measurement-matched transfer to measured scene S36. Method
PSNR↑
SSIM↑
NMSE↓
AoA↓
2
NeRF GSRF WRF-GS+ GRaF
12.94 16.85 16.96 14.24
0.457 0.587 0.592 0.510
0.287 0.145 0.140 0.219
16.24 12.09 9.26 7.87
CoRFCAL CoRFPR,CAL
16.74 17.45
0.620 0.653
0.125 0.118
8.12 7.88
In Table 17, CoRFPR,CAL leads in PSNR, SSIM, and NMSE, exceeding WRF-GS+ by 0.49 dB in PSNR and 0.061 in SSIM. Even without power refinement, CoRFCAL achieves higher SSIM and lower NMSE and AoA error than WRF-GS+, despite slightly lower PSNR. GRaF has nearly the same AoA error as CoRFPR,CAL but substantially worse full-spectrum metrics. AoA measures the strongest direction, while SSIM and NMSE also reflect the locations and relative strengths of secondary spectral structure. The calibration controls show that modifying the receiver response removes substantial error while the conditioned propagation estimates remain fixed. Uncalibrated feed-forward CoRF reaches 14.79 dB PSNR. Correcting array rotation and spacing with three geometric parameters raises this to 14.98 dB, whereas the element-wise response model reaches 16.74 dB. The small geometric gain suggests that a global rotation and two spacing corrections cannot capture most of the correctable response discrepancy. The element-wise model additionally represents gain, phase, position, and angularresponse differences, providing a plausible mechanism for its larger gain. Because these parameters are fitted together, the comparison does not identify which element-wise effect contributes most. Power refinement addresses a different part of the synthesis pipeline by correcting predicted component strengths while the pretrained propagation model remains frozen. Under the same measurement budget, CoRFPR,CAL reaches 17.45 dB. This further improvement is consistent with component-power error remaining after receiver calibration. The two adaptations therefore address complementary discrepancies: calibration changes the mapping from predicted propagation to the measured spectrum, while refinement changes the powers supplied to that mapping. 39
The 0.49 dB real-data margin is smaller than the 7.16 dB margin over the strongest few-shot baseline on unseen simulated variants. These comparisons use different measurement budgets and adaptation settings, so their margins should be interpreted within their respective protocols. In the measured-scene comparison, WRF-GS+ fits its field to S36, whereas CoRF retains a simulationtrained propagation model and adapts its receiver response and component powers. The baseline can therefore fit scene-specific multipath directly, while CoRF must infer it from the transferred model and sparse references. Calibration and power refinement can correct the response and strengths of predicted arrivals without retraining the learned propagation model, but neither can add a missing propagation component. Secondary reflected structure may remain unresolved even when the dominant direction is accurate, consistent with the failures in §E.1.4. A separate larger-budget diagnostic uses the 612-spectrum validation split in addition to the M = 32 conditioning references. Fitting the receiver response reaches 17.39 dB PSNR, while fitting both the response and power correction reaches 19.99 dB. This shows that additional target-domain data can improve both corrections, but the resulting 644-spectrum protocol differs from the matched 96spectrum comparison in Table 17. In a deployment with an accurately specified effective receiver response, receiver-specific calibration measurements could be omitted. Hardware specifications and network metadata may provide geometry and pose (AISG, 2024; 3GPP, 2024), but the 19.99 dB experiment also uses validation spectra to fit the scene-specific power correction. It therefore does not establish the same accuracy from metadata and M = 32 references alone. E.7 E.7.1
ROLE OF TARGET-S CENE R EFERENCES R EFERENCE D EPENDENCE
A conditional model could achieve strong accuracy while relying primarily on query geometry and ignoring the reference spectra. We therefore directly test whether CoRF uses the target-scene references by replacing them with references from other scenes while keeping the trained model and test queries fixed. All model parameters remain fixed throughout this experiment, so any performance change is caused only by changing the reference set. Protocol. For each test scene in the unseen-variant and unseen-category evaluations, we replace its M = 32 references with the reference set from each of the other 34 simulated scenes. For each replacement scene, we first evaluate 150 evenly spaced test queries and identify the reference set that produces the largest decrease in PSNR. We then evaluate all test queries using both the correct target-scene references and this selected replacement set. We additionally average over all 34 replacement scenes to determine whether the effect persists beyond the single reference set producing the largest degradation. Correct Scene References Matter. Figure 8 shows that predictions respond to which scene supplies the references. With the correct references, the model achieves 22.95 dB PSNR and 0.740 SSIM on unseen scene variants, and 22.83 dB and 0.733 on unseen scene categories. Replacing them with the reference set that produces the largest degradation reduces PSNR by 3.62 dB and 2.65 dB, reduces SSIM by 0.072 and 0.053, and increases NMSE by 63% and 45%, respectively. These are stress-test losses from the selected replacements, not typical scene swaps. The effect occurs broadly across test scenes rather than being driven by a few extreme cases. For unseen scene variants, every test scene loses at least 0.46 dB, with the largest decrease reaching 10.1 dB. For unseen scene categories, the largest decrease reaches 8.1 dB. The effect also remains when averaging over all 34 replacement scenes: PSNR decreases by 0.96 dB for unseen variants and 0.51 dB for unseen categories. Scenes that are more sensitive to the selected replacement set are also more sensitive on average, with Spearman correlations of ρ = 0.92 and 0.77, respectively. These results show that the dependence on the target-scene references is systematic rather than arising from a small number of unfavorable scene pairs. Reference Sensitivity Depends on Scene Structure. The magnitude of the degradation varies across target scenes. For unseen scene variants, the reference-swap PSNR change is negatively correlated with the PSNR obtained using the correct references, with Spearman correlation ρ = −0.83. Classroom and bedroom scenes can lose approximately 4–6 dB, whereas the most difficult corridor scenes often lose less than 1 dB. This suggests that when the model captures more scene-specific spectral structure, replacing the corresponding reference evidence removes more useful information. 40
6 9
0.06 0.12 0.18
Test scenes (sorted)
0.3 0.2 0.1 0.0
Test scenes (sorted)
AoA error change (°)
3
NMSE change
0.00 SSIM change
PSNR change (dB)
0
Test scenes (sorted)
8 4 0
Test scenes (sorted)
3 6 9
NMSE change
0.00 0.06 0.12 0.18 Test scenes (sorted)
0.3 0.2 0.1 0.0
Test scenes (sorted)
AoA error change (°)
0 SSIM change
PSNR change (dB)
(a) Unseen scene variants
Test scenes (sorted)
8 4 0
Test scenes (sorted)
(b) Unseen scene categories
Figure 8: Effect of cross-scene reference replacement. For each test scene, a replacement set is selected by the largest PSNR loss on 150 queries. Each bar shows its mean metric change over all test queries relative to correct-scene references, with scenes sorted separately for each metric. Table 18: Reference-budget sensitivity. Entries are median PSNR losses in dB across scenes relative to each scene’s mean at M = 32. The average first averages over ten reference draws; the least favorable result uses the worst draw for each scene. Unseen variants
Unseen categories
Budget
Average
Least favorable
Average
Least favorable
M =1 M =8 M = 32
0.62 < 0.05 0
2.40 0.53 0.16
0.41 < 0.05 0
2.10 0.30 0.09
References Primarily Affect Full-Spectrum Structure. Changing the references has a much smaller effect on dominant-bearing accuracy than on full-spectrum synthesis. Mean AoA error increases by only 0.35◦ for unseen variants and 0.91◦ for unseen categories, despite the substantially larger changes in PSNR, SSIM, and NMSE. This is consistent with the analytic decoder preserving geometric information about arrival directions while the reference-conditioned field supplies scenespecific information needed to synthesize the complete spatial spectrum. E.7.2
R EFERENCE B UDGET
The main experiments use a reference budget of M = 32. Here we vary the number of references available at inference while keeping the trained model fixed, allowing us to isolate how reference coverage affects cross-scene prediction. Protocol. For each test scene, we evaluate M ∈ {1, 2, 4, 8, 16, 32, 64, 128} using R = 10 reference sets per budget. The reference sets are nested within each draw, so the references available at a smaller budget are retained when the budget increases. At M = 32, draw 0 corresponds to the reference set used in the main evaluation. All unseen-variant and unseen-category folds are evaluated on their complete test splits. For each scene, we report performance relative to its mean result at M = 32, considering both the average over the 10 draws and the least favorable draw. At M = 32, the average difference is zero by definition, but the least favorable draw can fall below the scene mean. Average Performance Saturates Quickly. Table 18 shows that relatively few references are sufficient to recover most of the average performance. At M = 1, the median PSNR decrease across scenes is only 0.62 dB for unseen scene variants and 0.41 dB for unseen scene categories. By M = 8, the median decrease is below 0.05 dB in both settings. Increasing the budget beyond M = 32 produces negligible additional improvement. Thus, the mean prediction quality saturates after a relatively small number of target-scene references. 41
5.0 7.5
0.05 0.10 0.15
1 2 4 8 16 32 64 128
0.30 0.15 0.00
1 2 4 8 16 32 64 128
Number of references M
AoA error change (°)
2.5
NMSE change
SSIM change
PSNR change (dB)
0.00
0.0
Number of references M
1 2 4 8 16 32 64 128
7.5 5.0 2.5 0.0
Number of references M
1 2 4 8 16 32 64 128
Number of references M
5.0 7.5 1 2 4 8 16 32 64 128
Number of references M
0.05 0.10 0.15 1 2 4 8 16 32 64 128
AoA error change (°)
2.5
NMSE change
0.00
0.0 SSIM change
PSNR change (dB)
(a) Unseen scene variants 0.30 0.15 0.00
Number of references M
1 2 4 8 16 32 64 128
Number of references M
7.5 5.0 2.5 0.0
1 2 4 8 16 32 64 128
Number of references M
(b) Unseen scene categories
Figure 9: Effect of the reference budget across test scenes. For each budget M , each point shows the worst change over ten reference draws relative to the scene’s average performance at M = 32. Results are shown for unseen scene variants and unseen scene categories. More References Improve Selection Robustness. Although the average result saturates quickly, small reference budgets are substantially more sensitive to which measurements are selected. Figure 9 shows the least favorable draw for each scene across reference budgets. At M = 1, the 10 reference draws span a median of 2.4 dB across unseen-variant scenes and 3.0 dB across unseen-category scenes. The least favorable draw decreases PSNR by a median of 2.4 dB and 2.1 dB, respectively, with individual scenes losing as much as 8.4 dB and 7.9 dB. This variability decreases rapidly as more references are provided. At M = 4, the loss under the least favorable draw decreases to 0.85 dB for unseen variants and 1.02 dB for unseen categories. At M = 8, it further decreases to 0.53 dB and 0.30 dB, and at M = 32 to only 0.16 dB and 0.09 dB. Reference Sensitivity Varies Across Scenes. Scenes that achieve higher fidelity with the default reference set generally exhibit greater sensitivity when only one or a few references are available. For unseen scene variants, the signed PSNR change under the least favorable draw at M = 1 is negatively correlated with PSNR at M = 32, with Spearman correlation ρ = −0.60. For example, some classroom scenes lose several dB with an unfavorable single reference, whereas the most difficult corridor scenes change much less. This pattern is consistent with §E.7.1: scenes for which the conditioned field captures more scene-specific structure also depend more strongly on receiving representative target-scene evidence. Overall, increasing the reference budget primarily reduces sensitivity to reference selection rather than substantially increasing average accuracy. A small number of references is often sufficient for a typical prediction, but M = 32 provides substantially more robust performance across reference selections and scenes while additional references beyond this budget yield little further benefit. This nested-draw experiment begins at M = 1 and does not evaluate reference-free inference. E.7.3
R EFERENCE -S ET S AMPLING
The main experiments use a fixed budget of M = 32 target-scene references, but the particular measurements available in practice may vary. We therefore evaluate how sensitive CoRF is to reference selection while keeping the reference budget, trained model, and test queries fixed. Protocol. For each test scene s, let Rs denote its pool of candidate reference measurements. We independently sample R = 100 reference sets, C(r) s ∼ UniformSubset (Rs , M ) ,
r = 0, . . . , R − 1,
(62)
with M = 32. Sampling is performed without replacement within each set, while the test queries remain fixed and disjoint from all sampled references. The reference sets are selected before evaluation, and r = 0 corresponds to the original reference selection; this analysis evaluates the fixed feed-forward model, not the power-refined main-table rows. We evaluate all 12 unseen-variant and 42
1 2
0.00 0.02 0.04
Test scenes S1 S35
AoA error change (°)
0
NMSE change
SSIM change
PSNR change (dB)
0.02
1
0.04 0.02 0.00
Test scenes S1 S35
Test scenes S1 S35
0.15 0.00 0.15
Test scenes S1 S35
1 2
0.00 0.02 0.04
Test scenes S1 S35
Test scenes S1 S35
AoA error change (°)
0
NMSE change
0.02
1 SSIM change
PSNR change (dB)
(a) Unseen scene variants 0.04 0.02 0.00 Test scenes S1 S35
0.15 0.00 0.15
Test scenes S1 S35
(b) Unseen scene categories
Figure 10: Variation across R = 100 reference draws. Each box summarizes the metric change for one test scene relative to its mean across draws at M = 32. unseen-category folds, giving 1200 fold-by-reference-set evaluations in total. Because the model parameters remain fixed, any variation across repetitions is caused only by the selected references. Performance Is Robust to Reference Selection. Averaged over folds and reference sets, CoRF achieves 22.95 dB PSNR, 0.739 SSIM, 0.122 NMSE, and 8.40◦ AoA error on unseen scene variants. For unseen scene categories, the corresponding values are 22.86 dB, 0.733, 0.124, and 8.41◦ . The standard deviation of fold-averaged PSNR across the 100 reference sets is only 0.026 dB and 0.037 dB, respectively, and no reference set changes the fold-averaged PSNR by more than 0.12 dB relative to the original reference selection. Figure 10 examines the variation within individual scenes. For most scenes, reference selection has only a small effect. Across scenes, the median width of the interquartile range is 0.11 dB for unseen variants and 0.07 dB for unseen categories, while the median span between the 5th and 95th percentiles is 0.30 dB and 0.20 dB, respectively. The reference set used in the main evaluation is also representative, with a median difference from the corresponding scene mean of only −0.02 dB and 0.00 dB. A small number of scenes exhibit greater sensitivity. Individual reference sets can shift scene-level PSNR by as much as 1.38 dB for unseen variants and 2.09 dB for unseen categories. However, SSIM, NMSE, and especially AoA remain substantially more stable. Across all scenes, a single reference selection changes AoA error by at most 0.27◦ , indicating that reference selection affects detailed spectrum synthesis much more than the dominant arrival direction. Reference Selection Changes Scene-Level Spectral Structure. We further examine the scenes with the largest variation across reference sets. Changing the references tends to shift performance across many queries in the same scene rather than only at queries near the selected measurements. Between a scene’s lowest- and highest-performing reference sets, 63%–93% of its test queries improve together, while performance is essentially uncorrelated with distance to the nearest reference. The dominant spectral peak is also highly stable. Across these scenes, the predicted peak remains at the same angular cell as in the main reference selection for a median of 96%–100% of test queries. Instead, most of the variation appears in the broader spectral background. Reference sets containing more spectra with rich multipath tend to produce a stronger background response across the scene, whereas reference sets dominated by a strong direct path tend to produce a weaker background. These relationships are moderate rather than deterministic, indicating that CoRF combines information across the complete reference set rather than relying on individual measurements. Together with §E.7.1, these results distinguish two properties of the reference-conditioning mechanism. Changing which measurements are selected from the correct target scene usually produces only small variations, whereas replacing them with references from a different scene causes substantial degradation. Thus, M = 32 references provide a stable scene-level representation without requiring a specially selected reference set.
43
Table 19: Reference-budget sensitivity on one unseen-variant fold. GRaF uses the test-selected L = 4. M = 32
M = 64
M = 128
M = 256
GRaF PSNR GRaF SSIM
16.20 0.525
17.82 0.576
19.36 0.627
20.44 0.662
CoRF PSNR CoRF SSIM
24.70 0.785
24.70 0.784
24.74 0.785
24.74 0.783
Method
Table 20: Encoder-training ablation at M = 32. Results are from seed 0. Encoder training
Var. PSNR
Var. SSIM
Cat. PSNR
Cat. SSIM
24.22 24.08 24.75 24.97
0.778 0.770 0.793 0.794
25.06 25.54 24.84 25.70
0.794 0.809 0.790 0.812
Frozen random Synthesis loss only Masked spectrum Supervised contrastive
E.7.4
L OCAL C OVERAGE FOR GR A F
GRaF selects nearby reference measurements for each query, so its performance may depend more strongly on local reference coverage than that of CoRF. We examine this effect on unseen-variant fold 1 using 150 test queries per scene and identical nested reference sets for both methods. The budgets range from M = 32 to M = 256, with each larger set retaining the references in the smaller sets. We select GRaF’s neighbor count from L ∈ {4, 8, 16, 32} using these test queries; the best choice is L = 4. This test-selected choice favors GRaF, so the experiment diagnoses reference-density sensitivity rather than repeating the main-table protocol. As shown in Table 19, GRaF’s PSNR rises by 4.24 dB and its SSIM by 0.137 as the budget increases from M = 32 to M = 256. Over the same range, CoRF’s PSNR and SSIM change little. The PSNR gap narrows from 8.50 to 4.30 dB, but does not close even with eight times as many references. At M = 32, GRaF’s PSNR falls from 16.79 to 15.96 dB between queries in the nearest- and farthestreference quintiles, whereas CoRF shows no monotone distance trend. These observations suggest that GRaF benefits from denser query-local coverage, while CoRF remains comparatively stable across the tested budgets. The experiment covers one fold and one reference draw per budget; unseen-category evaluation and spatially stratified GRaF references were not evaluated. E.8 E.8.1
R EPRESENTATION AND R ENDERER D IAGNOSTICS R EFERENCE C OORDINATES
We train spectrum-only, absolute-coordinate, and query-relative-coordinate reference-conditioning variants on one unseen-variant fold and one unseen-category fold, reusing the same pretrained encoder. Absolute coordinates preserve once-per-scene conditioning, whereas relative coordinates require rebuilding the conditioned field for each query. Zeroing a trained coordinate branch at evaluation changes accuracy by at most 0.13 dB and 0.003 SSIM, with no consistent direction of change. Across three seeds, the absolute-coordinate model differs from spectrum-only by −0.05 dB on the unseen-variant fold and +0.33 dB on the unseencategory fold, within seed spreads of 0.47 and 0.95 dB. Relative coordinates reduce PSNR by 0.70 dB on one fold and improve it by 0.73 dB on the other, while zeroing their trained coordinate branch again has negligible effect. On these tested folds, neither coordinate variant provides a consistent accuracy gain over spectrum-only conditioning, which also avoids reference-pose metadata. E.8.2
P RETRAINING O BJECTIVE
We compare four encoder-training conditions on one unseen-variant fold and one unseen-category fold, using three seeds and the same 5,500 total parameter updates. The conditions are a frozen random encoder, synthesis-loss-only end-to-end training, masked-spectrum self-supervision, and the supervised contrastive objective used by CoRF. Supervised contrastive pretraining gives the highest seed-0 PSNR and SSIM on both folds in Table 20. Across seeds, its PSNR exceeds the best label-free condition by 0.24 and 0.35 dB, gains comparable 44
0.780
24.0 23.5
0.765 500 1.4k 4k 6.9k 11k Anchors K
AoA error (°)
24.5
5.22 0.096 NMSE
0.795 SSIM
PSNR (dB)
25.0
0.090 0.084
500 1.4k 4k 6.9k 11k Anchors K
500 1.4k 4k 6.9k 11k Anchors K
5.16 5.10 5.04
500 1.4k 4k 6.9k 11k Anchors K
Figure 11: Sensitivity to anchor-lattice resolution. Performance is shown across lattice sizes K with M = 32. The default K = 4,000 corresponds to the reported model. to the seed spreads. Accuracy alone, however, does not establish that a model uses its references. In the seed-0 reference-swap probe, seven of the eight trained models respond to a changed reference set, but the synthesis-loss-only model on the unseen-variant fold is exactly reference-blind despite reaching 24.08 dB. Scene-identification accuracy is also insufficient as a reference-use diagnostic: even a frozen random encoder separates scenes using raw spectral statistics. E.8.3
S ENSITIVITY TO A NCHOR R ESOLUTION
The canonical anchor lattice controls the spatial granularity of the conditioned propagation field. The default configuration uses K = 20 × 20 × 10 = 4000 anchors shared across scenes. We evaluate whether the reported performance depends strongly on this choice. Protocol. We evaluate five lattice resolutions with the same 2:2:1 aspect ratio: K ∈ {500, 1372, 4000, 6912, 10976} . For each resolution, we retrain the propagation model on the first unseen-variant fold while reusing the same pretrained reference encoder. All other training hyperparameters, the M = 32 references, and the evaluation split remain unchanged. Because the lattice spacing also determines the allowable anchor-position refinement and the initial spatial scale of each component, changing K jointly changes the spatial discretization of the propagation field. The Default Resolution Provides the Best Accuracy. As shown in Figure 11, the default K = 4000 lattice achieves the highest PSNR of 24.97 dB and SSIM of 0.794, together with 0.083 NMSE. Coarser lattices reduce PSNR by 0.30–0.45 dB, indicating that insufficient spatial resolution moderately limits synthesis accuracy. Increasing the resolution does not improve performance either: K = 6912 and K = 10976 reduce PSNR by 1.29 dB and 0.27 dB, respectively. AoA error remains nearly unchanged across all resolutions, ranging only from 5.06◦ to 5.20◦ . Higher Resolution Adds Cost Without Consistent Benefit. Training time increases steadily with the number of anchors. On one H100, training requires approximately 50, 51, 68, 92, and 124 minutes for the five resolutions from smallest to largest. Thus, the two finer lattices increase training cost by approximately 1.4× and 1.8× relative to the default without improving synthesis accuracy. Analysis of the predicted anchor powers also indicates that the model uses only a fraction of the available anchors. At the default resolution, the anchor-power participation ratio corresponds to approximately 244 effective anchors per query, substantially fewer than the 4000 available anchors. The finer lattices do not consistently increase this effective capacity: one run concentrates its power on very few anchors, while the other distributes it over substantially more anchors, yet neither improves accuracy. Because each alternative resolution is trained once, these differences should not be interpreted as an inherent failure mode of finer lattices. Rather, they show that increasing anchor density beyond the default does not provide a consistent accuracy benefit. Overall, performance is not highly sensitive to anchor resolution around the default operating point. A substantially coarser lattice modestly reduces synthesis fidelity, while finer lattices add computation without improving accuracy. The default K = 4000 therefore provides a practical balance between spatial resolution, synthesis quality, and training cost. E.8.4
O RACLE R ENDERER C OMPARISON
We ray-trace the first 100 test queries in each simulated scene and render the exact paths coherently, preserving phase cross terms, or incoherently, matching CoRF’s component-power approximation. We also evaluate feed-forward CoRF on the same queries. All predictions are scored against the released noisy targets, so even an exact-path rendering need not match them perfectly. 45
SSIM
PSNR (dB)
0.75
20 16 12
0.60 0.45
10
10
100 1k Number of training measurements
100 1k Number of training measurements
(a) PSNR
(b) SSIM
Figure 12: GSRF data efficiency on S36. GSRF is trained on increasing subsets of the S36 measurements and evaluated on the same test set. Table 21: Oracle rendering on 3,500 simulated queries. Prediction
PSNR↑
SSIM↑
NMSE↓
AoA↓
23.77 22.62 23.39 22.73
0.744 0.720 0.746 0.734
0.111 0.129 0.103 0.127
7.91 8.90 7.63 8.67
Oracle coherent renderer Independent noisy coherent sample Oracle incoherent renderer Feed-forward CoRF
As shown in Table 21, the incoherent oracle is 0.38 dB below the coherent oracle in average PSNR, while feed-forward CoRF is another 0.66 dB below the incoherent oracle. The incoherent oracle nevertheless scores slightly better on SSIM, NMSE, and AoA error, so the coherent renderer is not uniformly better against these noisy targets. The independent noisy coherent sample further illustrates the effect of target noise on the measured scores. The average PSNR difference conceals a larger approximation error in dense multipath. When the Ricean factor is below 0 dB, the incoherent oracle reaches 18.34 dB, compared with 19.98 dB for the coherent oracle. When at least four paths lie within 10 dB of the strongest, the corresponding scores are 16.81 and 18.55 dB. Thus, omitting coherent interference incurs a modest average PSNR cost on this corpus but a substantially larger cost in its dense-multipath subsets.
E.9 E.9.1
D EPLOYMENT E FFICIENCY DATA -C OLLECTION E FFICIENCY
Measurement Cost. Full-data per-scene training uses the target scene’s training split: a median of 3376 measurements per scene and a range of 2823–4286. In the cross-scene setting, CoRF and CoRFPR use M = 32 target-scene references, approximately two orders of magnitude fewer measurements than full-data fitting. The reference-budget analysis in §E.7.2 shows that this budget already lies in the regime where additional references provide little improvement in average accuracy. The calibration-only protocol uses M = 32 scene references and NC = 64 separate calibration measurements, for a total of 96 target-domain spectra. The joint protocol conditions on the same 32 references, calibrates with the same separate 64 spectra, and reuses both sets for power refinement. For an illustrative acquisition time of 10 s to 1 min per location, the upper end reflecting a Wi-Fi site-survey estimate (Li et al., 2017), collecting 3376 measurements would take 9.4–56.3 hours, compared with 5.3–32 minutes for 32 references. The 64 separate calibration measurements would add 10.7–64 minutes under the same assumption; the reported protocol collects them for each scene. Empirical Data Efficiency. We directly examine how much target-scene data a per-scene field requires using the real scene S36. We train GSRF on random subsets of its 4898 non-test measurements, using 15 subset sizes from 10 to 4898 and three random seeds per size, and evaluate every model on the same 1225 test measurements. Figure 12 shows that accuracy improves steadily with the amount of real training data. With 30 measurements, close to the reference budget of CoRF, GSRF achieves 15.14 dB PSNR and 0.506 SSIM. With 2449 measurements on S36, GSRF comes within 0.44 dB PSNR and 0.014 SSIM of its full-data result of 21.98 dB and 0.814. 46
Table 22: Cost of deployment in an unseen scene. Runtime is measured on the same NVIDIA H100. Method
Parameters
Pre-training
Measurements
Setup
Inference
CoRF CoRFPR
12.7M 12.7M + 0.004–0.270M
86 min 86 min
32 32
10 ms 47–202 s
8.5 ms 8.5 ms
GRaF
12.6M
177 min
32
none
2.9 s
WRF-GS GSRF NeRF2
1.3M 4.4M 0.69M
none none none
3,376 3,376 3,376
21–25 min 45–99 min 89–172 min
3.7 ms 12 ms 125 ms
Table 23: Material-property shift for feed-forward CoRF, averaged over the five unseen-variant folds.
E.9.2
Scale (εr , σcond )
PSNR
SSIM
NMSE
AoA
∆PSNR
(1, 1) (0.85, 0.67) (0.7, 0.33) (1.2, 1.5) (1.5, 3) (2, 10)
22.90 22.99 23.34 22.19 20.43 17.76
0.738 0.740 0.747 0.725 0.679 0.598
0.122 0.119 0.113 0.132 0.168 0.254
8.49 8.05 7.77 9.35 12.92 21.66
0.00 +0.09 +0.44 -0.71 -2.46 -5.14
C OMPUTATIONAL E FFICIENCY
Table 22 reports wall-clock measurements on one NVIDIA H100. A single CoRF model contains 12.7M inference-time parameters and is shared across scenes. Training takes approximately 86 minutes per fold, after which feed-forward conditioning on the M = 32 references of a new scene requires 10 ms. This timing describes feed-forward CoRF only. Depending on whether anchor biases, the residual network, or both are fitted, the refinement optimizes 0.004–0.270M parameters and requires 47–202 seconds of measured scene setup. Rendering requires approximately 8.5 ms per query. A complete feed-forward test-scene map, containing a median of 957 queries, takes approximately 16–21 s end to end, including scene conditioning and I/O; power refinement adds its one-time target-scene fit. Once trained, the per-scene baselines have comparable per-query rendering costs for WRF-GS and GSRF, at approximately 3.7 and 12 ms, while NeRF2 requires approximately 125 ms. Their primary deployment cost is scenespecific optimization before rendering: 21–25 minutes for WRF-GS, 45–99 minutes for GSRF, and 89–172 minutes for NeRF2 under their released training configurations. GRaF is also generalizable and therefore does not require per-scene fitting. Its pretraining takes 177 minutes per fold, but its neighbor retrieval and volumetric rendering require approximately 2.9 s per query. A complete test-scene map therefore takes approximately 48 minutes, compared with seconds for feed-forward CoRF or a one-time fit plus seconds of rendering for CoRFPR . E.10
P HYSICAL D ISTRIBUTION S HIFT
We test feed-forward CoRF under material-property shifts by re-simulating all 35 scenes after scaling the relative permittivity and conductivity of every non-metal material. Geometry, poses, array, frequency, ray-tracing settings, and per-scene measurement SNR remain fixed. Each scene uses its unchanged unseen-variant-fold checkpoint, one fixed draw of M = 32 shifted references, and its first 300 test queries, without parameter adaptation. The (1, 1) row in Table 23 is a re-simulated control. Reducing the material parameters yields small PSNR gains in this test, whereas increasing them progressively degrades performance. At the (1.2, 1.5) shift, PSNR falls by 0.71 dB. At the most severe (2, 10) shift, it falls by 5.14 dB, and AoA error rises from 8.49◦ to 21.66◦ . This last setting serves as a stress test of the fixed propagation model, not a typical deployment perturbation.
47