arXiv:2609.11894v1 [cs.CV] 10 Sep 2026
3D Point Splatting for mmWave Radar Novel View Synthesis
Adnan Armouti Cornell Tech New York, NY, USA [email protected]
Yixuan Gao Cornell Tech New York, NY, USA [email protected]
Rajalakshmi Nandakumar Cornell Tech New York, NY, USA [email protected]
Abstract Solving novel view synthesis (NVS) for millimeter-wave (mmWave) radar requires a renderer that is physically faithful, complex-valued, and multi-viewpoint-tractable. No prior method achieves these three properties simultaneously. Differentiable Monte Carlo (MC) ray tracers implement the radar forward model directly with explicit material modeling and complex outputs, but do not scale to the multiview optimization NVS demands. Optical-NVS ports of NeRF, hash grids, and 3D Gaussians train fast but discard phase and replace explicit material modeling with opaque learned features, restricting them to power-only range–azimuth (RA) magnitudes. We propose 3D Point Splatting (3DPS), the first differentiable point renderer for radar, derived directly from the standard solid-angle form of the radar equation. Each oriented 3D point carries an ITU-R P.2040 material model, evaluated in closed form, with the resulting complex phasor splatted into range bins through a precomputed point spread function (PSF). The complex-valued output makes the renderer product-agnostic. The same optimized scene yields analog-todigital converter (ADC), complex range profile (CRP), and RA outputs through standard fast Fourier transform (FFT) pipelines without retraining for each format. On six outdoor ColoRadar scenes, 3DPS reaches 0.587 mean Pearson correlation on held-out RA images. This is between 1.7× and 5.2× the three optical-NVS baselines (RadarSplat, Radar Fields, DART). Training takes approximately 3 minutes per scene on a single RTX 4090.
1
Introduction
Millimeter-wave (mmWave) radar penetrates fog, rain, and airborne dust Guan et al. [2020], Bijelic et al. [2020], returns Doppler velocity directly, and is an essential sensor for robust perception in autonomous driving and robotics Gao et al. [2019], Harlow et al. [2024]. However, radar datasets are substantially smaller than their optical or LiDAR counterparts. Radar sensors are not ubiquitous, the captured data structure depends heavily on array geometry and application-specific waveform configuration, and most large-scale datasets remain proprietary. NVS for radar can address this data scarcity by enabling sparse training sets to generalize across novel poses, but this problem is structurally hard. The supervision signal is typically i) sparse with returns dominated by quasi-mirror specular reflections at 60–77 GHz, ii) constrained to 2D as commercial off-the-shelf (COTS) radar array geometries lack an elevation aperture and integrate out the elevation dimension, and iii) limited in view count as NVS aggregates only a handful of training views per scene. Achieving high-fidelity radar NVS requires a renderer with three properties. First, physically faithful, so that predictions at unseen poses match the underlying sensor physics and the optimized scene parameters remain interpretable, editable, and transferable across sensor views. Second, complexvalued, since the radar forward model emits a complex phasor that is product-agnostic. The same optimized scene yields ADC, CRP, and RA outputs at novel viewpoints without retraining (Fig. 1). Preprint.
Mesh + MC ray tracing
NeRF / 3DGS Primitives
3DPS: Point Primitives
✓ physics-based ✓ product-agnostic, complex £ slow ( » 160 min)
£ approximated £ power-only jRAj ✓ fast
✓ physics-based ✓ product-agnostic, complex ✓ fast ( » 3 min)
Applications Product-Agnostic Rendering CRP
ADC
GT
F ¡1 r
Fµ
3DPS
F ¡1 r
Fµ
RS
F ¡1 r
Fµ
jRAj
Novel View Synthesis
Compression, Material Reconstruction and Normals Estimation
GT
£
£
3DPS
N = 20;000 pts
8 train jRAj
⋯
⋯
RS
Figure 1: The right primitive for radar NVS. Top: three method families on the three required properties. MC ray-tracing renderers (e.g., Sionna RT) are physics-faithful but slow. Optical-NVS adaptations (RadarSplat, Radar Fields, DART) are fast but discard phase. 3DPS satisfies all three. Bottom: 3DPS outputs on S1 F185 (|RA|, raw ADC magnitude, point cloud colored by learned ε′r ).
This also allows the renderer to reproduce antenna patterns, sidelobes, and FFT windowing from the signal model, sharpening RA correctness. Phase-discarding renderers forfeit both. Third, multiviewpoint-tractable, since NVS by definition demands joint optimization across many training poses. No prior method has all three properties. Differentiable Monte Carlo (MC) ray tracers Hoydis et al. [2023], Hofmann et al. [2025], Chen et al. [2025] implement the radar forward model with explicit bidirectional scattering distribution function (BSDF) material modeling, real antenna beam patterns, and complex outputs. The cost is structural, since each surface hit must be evaluated against every transmitter–receiver (TX, RX) pair to maintain multiple-input multiple-output (MIMO) coherence. Per-viewpoint MC ray-tracing fits commonly require tens of minutes (Sionna RT Hoydis et al. [2023] measured at approximately 27 minutes per single-viewpoint fit at our 500-iteration budget; Sec. I), which compounds to multiple hours per scene at the multi-viewpoint optimization NVS demands. Optical-NVS adaptations Huang et al. [2024], Borts et al. [2024], Takawale and Roy [2025], Kung et al. [2025] of neural radiance fields, hash grids, and 3D Gaussians Mildenhall et al. [2021], Kerbl et al. [2023] train fast on real radar captures, because they trade explicit modeling of the radar signal chain for runtime. Supervision is on post-FFT power-only RA images, phase is dropped by construction, and the radar’s signal-model components are absorbed into learned features rather than rendered. 3DPS realizes all three properties because we derive our closed-form, deterministic, and differentiable pipeline directly from the standard solid-angle form of the radar equation Skolnik [2001]. We initialize an oriented 3D point set from a co-registered LiDAR point cloud (Sec. 3.3, see Scope below). For physical fidelity, the antenna geometry, BSDF (per ITU-R P.2040), and antenna gains are evaluated in closed form with no learned features, and phase is integrated from the geometric path length rather than learned. Each oriented 3D point contributes exactly once per (TX, RX) pair, with no Monte Carlo sampling. For complex-valued output, the per-(TX, RX) range profile is preserved end-to-end through PSF splatting, enabling product-agnostic rendering to ADC, CRP, and RA without retraining via standard FFT pipelines. For multi-viewpoint tractability, we use a Hann-windowed PSF splat at constant cost per (point, range bin) and fuse BSDF evaluation, splatting, and the analytical backward into three CUDA kernels. Across six outdoor ColoRadar Kramer et al. [2022] scenes, 3DPS reaches 0.587 mean Pearson correlation on held-out novel-view |RA|, between 1.7× and 5.2× the three baselines, in approximately 3 minutes per scene on a single RTX 4090. 2
Table 1: Comparison of radar rendering methods. Signal types: ADC, CIR, IF, RA, RDA (signaloutput formats; full names in Sec. A). Columns mark scene representation (Repr.), explicit material model (Mat.), per-point or per-vertex storage (Per-pt), real-data validation (Real), and per-scene wall-clock on a single RTX 4090 over 8 viewpoints (Time, Sec. 4). Method
Repr.
Signal
Complex
Diff.
Mat. Per-pt BSDF NVS Real
Sionna RT Hoydis et al. [2023] SFCW Inv. Hofmann et al. [2025] InverTwin Chen et al. [2025]
Mesh Mesh Mesh
CIR IF CIR
✓ ✓ ✓
Partial Partial ✓
– ✓ –
– ✓ –
– – –
– – –
– ✓ ✓
– – –
DART Huang et al. [2024] Implicit Radar Fields Borts et al. [2024] Implicit SpINR Takawale and Roy [2025] Implicit RadarSplat Kung et al. [2025] 3DGS
RDA RA RA RA
– – – –
✓ ✓ ✓ ✓
– – – –
– – – –
– – – –
✓ ✓ ✓ ✓
✓ ✓ ✓ ✓
0.4 min 1.6 min – 20 min
✓
✓
✓
✓
✓
✓
✓
3 min
Ours
Points ADC, CRP, RA
Time
Contributions. (i) The first differentiable point renderer for radar, with a radar-native primitive (oriented 3D points carrying ITU-R P.2040 material parameters) and multi-viewpoint-tractable rendering operation (PSF splatting of complex phasors into range bins), derived from the solid-angle form of the radar equation. (ii) Product-agnostic rendering, since the same optimized scene yields CRP, ADC, and RA outputs through standard FFT pipelines without retraining. (iii) Sparse-view radar NVS at the state of the art on |RA|, |CRP|, and ADC envelope across six ColoRadar scenes, in approximately 3 minutes per scene. Scope. (i) LiDAR initialization. 3DPS initializes the oriented 3D point set from a co-registered LiDAR point cloud, following the LiDAR-derived geometric prior described in Sec. 3.3. This is standard for outdoor radar Kung et al. [2025], Kramer et al. [2022] and 3DGS NVS Xiong et al. [2023], Yan et al. [2024]. A 3D prior is necessary because the supervision signal suffers from cascading sparsity: a sparse set of training views, sparse radar returns per view, and elevation integrated out by 2D RA. Together these compound to leave 3D scene structure unrecoverable from RA-only supervision in the sparse-view setting NVS demands. LiDAR-free 3D recovery from radar (shape-from-radar) is an active open problem outside our scope. (ii) Single-bounce path tracing. 3DPS evaluates a single TX→scatterer→RX path per oriented point, matching the single-bounce assumption in RadarSplat Kung et al. [2025].
2
Related Works
Differentiable Monte Carlo (MC) ray tracers. Building on the differentiable-rendering toolkit pioneered for graphics — edge-sampling gradients Li et al. [2018], reparameterized boundary integrals Loubet et al. [2019], unbiased warped-area sampling Bangaru et al. [2020], and the broader differentiable-rendering survey by Kato et al. Kato et al. [2020] — recent radar-domain ray tracers extend the same machinery to electromagnetic propagation: Sionna RT Hoydis et al. [2023], Hofmann et al. Hofmann et al. [2025], InverTwin Chen et al. [2025], and the deterministic MIMO ray tracer of Schüßler et al. Schüßler et al. [2021] variously suspend gradients through their path solvers, do not model wideband frequency-modulated continuous-wave (FMCW) radar, assume known geometry, or validate only on synthetic data. Earlier mmWave FMCW ray-tracing simulators Dudek et al. [2010] establish the methodology but lack differentiability. SAR-domain inverse rendering Wei et al. [2024] pursues a similar materials-fitting agenda but on synthetic-aperture geometry rather than real-aperture FMCW. This family is physically faithful and complex-valued but slow at the multi-viewpoint optimization NVS demands: per-viewpoint MC fits commonly require tens of minutes (Sec. I), which compounds to multiple hours per scene at the 8-viewpoint optimization 3DPS targets. 3DPS preserves the physical fidelity and complex output of MC ray tracers while replacing per-iteration MC sampling with deterministic PSF splatting at constant cost per (point, range bin). Optical-NVS adaptations. Following neural radiance fields Mildenhall et al. [2021] and 3D Gaussian Splatting Kerbl et al. [2023] (with hash-grid acceleration Müller et al. [2022], point-based light-field variants Ost et al. [2022], and subsequent surfel-style refinements Jiang et al. [2025]), DART Huang et al. [2024] trains a NeRF on range–Doppler magnitude, Radar Fields Borts et al. [2024] synthesizes power-only RA via a hash-grid and multilayer perceptron (MLP), and RadarSplat Kung et al. [2025] ports 3DGS to RA rendering with opaque per-Gaussian features. Other 3
neural-field adaptations Takawale and Roy [2025], Zhang et al. [2025], Rafidashti et al. [2025], Lu et al. [2024], Farrell et al. [2023] extend to dynamic scenes or fuse radar with other modalities, but render no complex output. The same neural-field machinery has been ported to other coherent sensors: SAR Lei et al. [2024], Ehret et al. [2024] and ISAR Deng et al. [2024] for radar-like apertures, LiDAR for point-cloud novel views Tao et al. [2024], Huang et al. [2023], sonar via implicit fields Reed et al. [2023] and Gaussian splatting Qu et al. [2024], Sethuraman et al. [2025], and physically-based scattering NeRFs for adverse-weather media Ramazzina et al. [2023]. This family is multi-viewpoint-tractable but neither physically faithful nor complex-valued. Every method supervises on post-FFT power-only RA images, drops phase by construction, and absorbs the radar signal-model components (antenna patterns, sidelobes, FFT windowing, materials) into opaque learned features. 3DPS preserves multi-view tractability while adding explicit ITU-R P.2040 BSDF material modeling, real antenna beam patterns, and per-(TX, RX) complex range profile output.
3
Method
3DPS renders complex-valued FMCW range profiles from a learned set of oriented 3D points, anchored to a LiDAR-derived geometric scaffold. Each point carries a quaternion-encoded surface normal and a six-parameter ITU-R P.2040 material vector, evaluated by a closed-form bidirectional scattering distribution function (BSDF) and splatted into range-profile bins through a precomputed sensor point spread function (PSF). Sec. 3.1 states the forward model from the radar equation through the hemisphere integral to a deterministic discrete sum and its splatted form, plus the per-point BSDF. Sec. 3.2 walks through the rendering pipeline, justifies the choice of primitive, and ties the design back to the three requirements in Sec. 1. Sec. 3.3 details optimization and implementation. 3.1
Radar forward model
Consider an FMCW radar with a MIMO array of NTX transmitters and NRX receivers, chirp sweep s −1 slope µ, carrier wavenumber k = 2π/λ, and per-chirp ADC sampling {tn = n/fs }N n=0 . A scene is illuminated by transmitter t at pt and observed by receiver r at pr . We let Rt (x) = ∥x − pt ∥, Rr (x) = ∥x − pr ∥, ω̂ i the incident direction (TX→ x), and ω̂ o the scattered direction (x →RX). Radar equation. The complex amplitude received from a point scatterer at x is Skolnik [2001] p Pt Gt (ω̂ i ) Gr (ω̂ o ) p λ Str (x) = σ(ω̂ i , ω̂ o ; x) e−jk (Rt (x)+Rr (x)) , (1) 3/2 Rt (x) Rr (x) (4π) where σ is the bistatic radar cross section (RCS) and Gt , Gr are the (complex) antenna gain patterns. Hemisphere integral form. For a single-bounce path TX → x → RX over a scene surface, the dechirped ADC sample is the hemisphere integral over the receiver-centered upper hemisphere Ω+ (pr ) (full surface-to-hemisphere derivation in Sec. B) Z Rr2 (ω̂) atr [n] = Str x(ω̂) e j2π µ τ (ω̂) tn dω, (2) Ω+ (pr ) |n(ω̂)· ω̂| where x(ω̂) is the first scene intersection along ω̂ from pr and τ (ω̂) = (Rt + Rr )/c. Discretization on an oriented point set. We approximate Ω+ (pr ) by N oriented points {(xi , ni , Ai , θ i )}N i=1 , where θ i is the per-point ITU-R P.2040 material vector and each point covers 2 solid angle ∆ωi = Ai |ni · ω̂ o,i |/Rr,i at the receiver. The Riemann sum approximation to Eq. (2) 2 cancels the Rr /|n· ω̂| Jacobian against ∆ωi , leaving the closed-form, deterministic single-bounce ADC sum âtr [n] =
N X
vi Ai Str (xi ) e j2π µ τi tn ,
vi = ⊮[ni · ω̂ o,i > 0] , τi =
Rt,i +Rr,i . c
(3)
i=1
Each point contributes exactly once per (TX, RX) pair, with no Monte Carlo sampling or path enumeration. CRP via Hann-windowed PSF splat. The complex range profile is the windowed range FFT of the ADC, CRPtr [k] = Fn {w[n] atr [n]}. Linearity of Eq. (3) lets us push the FFT inside the sum, so 4
Forward Model
3DPS Scene Representation
Antenna Gain
BSDF
zi =
p
GTX GRX fr e j'i =Ri2
! bin ki = round(Ri =¢r)
Phase Optimised 3DPS points (positions, normals, materials)
Ground Truth
Splat
GTX (µi ) GRX (µo )
fr (µi ; µo ; n i ; "r0 ; ¾; t)
Rendered
Fµ
Fµ RA Loss
'i = ¡2¼ (RiTX + RiRX )=¸
Figure 2: 3DPS rendering pipeline. A LiDAR cloud is reduced to N =20,000 oriented points. Per-point ITU BSDF and complex amplitude are evaluated per (TX, RX) pair (Eqs. 1, 5), then splatted via a Hann-windowed PSF into the L=15 nearest range bins (Eq. 4). Stacking the per-(TX, RX) range profiles and applying an azimuth FFT yields the final RA map. each point becomes a known kernel centered at its (fractional) range bin CRPtr [k] =
N X
vi Ai Str (xi ) Φ k − ki ; δi ,
ki =
Rt,i +Rr,i , ∆R
δi = ki − ⌊ki ⌋,
(4)
i=1
where ∆R is the range bin width and Φ(·; δ) is the precomputed Hann-window PSF (length L = 15 taps). Splatting replaces the per-sample chirp synthesis of Eq. (3) with a single scatter-add into L adjacent range bins per point, dropping the per-pair cost from O(N ·Ns ) to O(N ·L). Validity bounds (Hann-tail truncation, fractional-bin handling, numerical equivalence to the standard ADC-then-FFT pipeline) are quantified in Sec. C. Per-point ITU-R P.2040 BSDF. Each point’s bistatic RCS factors via the Kirchhoff decomposition of the ITU-R P.2040 ITU-R [2023] air–material interface model into a roughness-attenuated coherent specular term and an incoherent diffuse lobe: σi (ω̂ i , ω̂ o ) = |Γ(θi ; ε′r , ε′′r , d)|2 ρcoh (σh ; θi ) Kspec (ω̂ i , ω̂ o ; ni , τ ) + ρinc (σh , ℓc ; ω̂ i , ω̂ o , ni ) , (5) with all material parameters drawn from the per-point vector θ i = (ε′r , ε′′r , σh , ℓc , τ, d)i (subscript i suppressed inside the equation for compactness). Γ is the multilayer-slab Fresnel coefficient at incidence θi = arccos(−ni · ω̂ i ), parameterized by complex permittivity εr = ε′r − jε′′r and slab 2 thickness d; ρcoh = exp −(2kσh cos θi ) is the Ament coherent attenuation; Kspec is the specular angular lobe (Kirchhoff approximation (KA) / small-perturbation method (SPM) blend with weight τ ); and ρinc is the incoherent diffuse lobe. Each parameter is reparameterized via sigmoid (τ ) or shifted softplus (positive-only quantities). The full closed form is in Sec. D. 3.2
Rendering pipeline
Figure 2 traces a single forward pass. Input: an oriented 3D point set {(xi , ni , Ai , θ i )}N i=1 with N =20,000 (Sec. 3.3). For each (TX, RX) pair the renderer (i) computes per-point geometry Rt,i , Rr,i , ω̂ i,i , ω̂ o,i and visibility vi from ni ; (ii) evaluates the per-point BSDF σi from θ i via Eq. (5); (iii) assembles the per-(TX, RX) complex amplitude Str (xi ) of Eq. (1); and (iv) scatteradds each point’s complex contribution into the L=15 nearest range bins of CRPtr [k] via the Hann-windowed PSF kernel Φ. The output of the splat is a set of NTX × NRX complex range profiles. Stacking these into the virtual-array dimension and applying an NTX -point azimuth discrete Fourier transform (DFT) yields the range–azimuth (RA) map of shape (Naz , K) = (127, 256), with arcsin-spaced azimuth bins θa = arcsin(2(a−Naz /2)/(Naz +1)) and forward-cone half-angle arcsin(126/128) ≈ 79.86◦ . ADC samples are recovered by inverse range FFT of CRP. The same optimized scene therefore yields all three downstream products (ADC, CRP, and RA) through standard FFT pipelines without retraining. Choice of primitive. The radar equation (Eq. 1) requires hit positions, surface normals, and material parameters at every scattering site. The hemisphere integral form (Eq. 2) operates on the receivercentered solid angle, removing any need for attached surface area on the primitives. Oriented 3D points are therefore the radar-native primitive. MC ray tracers (Sionna RT Hoydis et al. [2023]) 5
recompute these hit points every iteration. We make them persistent and learnable, removing the per-iteration ray-tracing cost. The 3DGS literature Kerbl et al. [2023] motivated 3D Gaussians over points using “holes, aliasing, strict discontinuity” arguments at sub-pixel optical resolution. Radar resolution is an order of magnitude coarser at ∆R=5.93 cm range and ∼1.4◦ azimuth, and each oriented point splats a continuous L=15-tap Hann PSF over ∼89 cm of range. Those arguments do not apply here. Realizing the three properties (Sec. 1). Physically faithful: antenna gains, the ITU-R P.2040 BSDF, and the carrier phase e−jk(Rt,i +Rr,i ) are all evaluated in closed form from geometric path lengths, so the optimized θ i and ni remain physically interpretable, editable, and transferable across sensor views. Complex-valued: Eq. (4) preserves the per-(TX, RX) complex range profile as a first-class output. The same optimized scene synthesizes ADC, CRP, and RA at novel viewpoints without retraining. Multi-viewpoint-tractable: each oriented point contributes exactly once per (TX, RX) pair with no Monte Carlo sampling or path enumeration. The PSF splat has fixed cost O(N L) per pair, and the fused CUDA implementation (Sec. 3.3) trains across 8 viewpoints jointly in approximately 3 minutes per scene on a single RTX 4090. 3.3
Optimization and implementation
LiDAR-derived geometric prior. A 3D geometric prior is necessary for the cascading-sparsity reasons in Sec. 1. We initialize the oriented point set from a co-registered LiDAR cloud of between 2.7 and 6.5 million raw points across our scenes, reducing to N =20,000 via four stages: frustum and range culling, occlusion ray-casting against the scene with Mitsuba 3 Jakob et al. [2022], cosineweighted resampling biased toward array boresight, and farthest-point sampling. Each surviving point inherits a quaternion-encoded LiDAR surface normal and a uniform ITU-concrete material prior. Full reduction-stage parameters are in Sec. E. Learned parameters and carrier-phase detach. For each point we learn the 6-parameter ITU material vector θ i , the surface-normal quaternion qi (initialized from the LiDAR normal), and the 3D position pi through the amplitude path only. The carrier-phase term e−jk(Rt,i +Rr,i ) in Str (xi ) is detached from position gradients because the bin index ⌊ki ⌋ creates a piecewise-constant dependence on pi , making position-phase gradients degenerate. The forward-pass carrier phase is still computed (0) from the unrounded Rt,i + Rr,i . An L2 anchor λpos ∥pi − pi ∥2 (λpos =100) restricts position drift to approximately 1 mm per iteration, much smaller than the carrier wavelength (λ ≈ 3.9 mm at 76 GHz), so geometry remains LiDAR-anchored and the carrier phase remains correct in expectation. GT Optimization and CUDA implementation. The loss is per-pose MSE on |RApred F | vs. |RAF | 2 (both L -normalized), optimized with Adam (β1 =0.9, β2 =0.999) at per-parameter learning rates ηmat =10−2 , ηrot =5·10−3 , ηpos =10−5 . Following 3DGS Kerbl et al. [2023] but adapted to radar physics, we split the top 5% andP prune the bottom 5% of points by accumulated amplitude-path position-gradient magnitude gi = F ∥∇pi LF ∥ at iters {100, 200, 300, 400}, preserving the budget N =20,000. Position gradients are tracked but excluded from the optimizer step, so positions stay on the LiDAR scaffold. Eqs. (4)–(5) plus the virtual-array DFT run as three fused CUDA kernels (BSDF, scatter-splat, and analytical backward), with the analytical backward validated against finite differences. The fused forward+backward+optimizer step renders all N =20,000 points across 192 (TX, RX) pairs in approximately 10 ms on an RTX 4090. A 500-iteration 8-viewpoint training pass completes in approximately 3 minutes per scene.
4
Experiments
We evaluate 3DPS on six outdoor ColoRadar Kramer et al. [2022] scenes captured with a TI MMWCAS cascaded mmWave radar (12 TX × 16 RX, 76 GHz carrier, 5.93 cm range resolution) and a co-registered Ouster OS1-64 LiDAR. The scenes span parking lots, vehicle staging areas, and outdoor infrastructure with primarily static structure (walls, parked vehicles, traffic infrastructure). Why ColoRadar. 3DPS requires four inputs: complete radar array geometry, raw I/Q ADC samples, co-registered LiDAR, and dense pose annotations. ColoRadar provides all four. Other automotive radar datasets Caesar et al. [2020], Schumann et al. [2021], Barnes et al. [2020], Burnett et al. [2023], Sheeny et al. [2021], Paek et al. [2022], Rebut et al. [2022] either ship pre-processed radar tensors 6
without per-element array geometry, use mechanically-rotated single-channel sensors that do not expose a virtual array, or lack the co-registered LiDAR and dense poses 3DPS requires. 4.1
Experimental setup
Scene splits. We follow the periodic held-out-frame protocol used by Radar Fields Borts et al. [2024] and RadarSplat Kung et al. [2025], both of which reserve every fifth frame from a contiguous trajectory window as the novel-view test frame. Adapted to the cascade’s 5 Hz frame rate and our short-window (9-frame) experimental setup, this becomes a centred 1-of-9 variant: for each scene we select 9 adjacent cascaded radar frames at 5 Hz (200 ms inter-frame spacing) and use only the first chirp loop of each frame. The middle frame F is held out as the novel-view test frame and the eight bracketing frames {F ±1, . . . ,±4} form the training set, giving 8 training RA maps and 1 held-out test RA map per scene. Baselines. Three published radar NVS methods from the optical-NVS family (full descriptions in Sec. 2): DART Huang et al. [2024], Radar Fields Borts et al. [2024], and RadarSplat Kung et al. [2025]. Only 3DPS renders complex range profiles, so CRP and ADC columns are reported for our method only; the three optical-NVS baselines emit power-only RA by construction. Metrics. Following Kramer et al. Kramer et al. [2022], we report Pearson correlation (Corr), peak signal-to-noise ratio (PSNR), structural similarity (SSIM), and root-mean-square error (RMSE), all computed on min-max-normalized 399×399 Cartesian |RA| images (Sec. A.5). Implementation. 3DPS training takes approximately 3 minutes per scene for joint 8-viewpoint optimization on a single NVIDIA RTX 4090, or 30 to 45 seconds per single-viewpoint fit at the same 500-iteration budget (full details in Sec. 3.3). Measured baseline wall-clocks on identical hardware are 0.4, 1.6, and 19.7 min per scene for DART, Radar Fields, and RadarSplat respectively. Per-scene wall-clock and peak GPU memory are in Sec. I. Sample efficiency. 3DPS is robust at the view axis: dropping from the 8-view default to just 2 bracketing training frames yields 0.573 mean test Pearson correlation, within 0.014 of the 8-view default (0.587); per-view-count breakdown in Sec. L. Sparse-view radar deployments can therefore use as few as two training frames with negligible held-out-view loss. 4.2
Quantitative results
Table 2 reports the cross-method, cross-product results, averaged over the six ColoRadar scenes for training views (8 per scene) and the held-out novel-view test frame. The three optical-NVS baselines emit power-only |RA| and have no CRP, ADC, or complex outputs, so those columns are marked ‘–’. Per-scene breakdowns and CRP/ADC evaluation methodology are in supplement Sec. F and G. Magnitude on |RA|. 3DPS achieves 0.812 mean Pearson correlation on training views and 0.587 on the held-out test view, on Cartesian |RA|. This is 2.2× the next-best baseline on training views and 1.7× on the test view. PSNR, SSIM, and RMSE all rank 3DPS first on the mean and on every per-scene metric (supplement Tables 3 and 4). The next-best method on both splits is RadarSplat at 0.365 train and 0.339 test. Among the three optical-NVS baselines, RadarSplat is best, then Radar Fields, then DART. The ranking matches the structural mismatches identified in Sec. 2. RadarSplat opaque per-Gaussian features cannot share information with the radar antenna pattern, capping at 0.37 train and 0.34 test. Radar Fields combines a hash-grid encoder with a multilayer perceptron decoder, supervised on post-FFT magnitudes. Without a virtual-array model it fails to recover scene support from sparse RA alone, reaching only 0.22 train and 0.13 test. DART is designed for Nd =256 chirps but operates here at the cascade’s Nd =16, leaving insufficient Doppler resolution at 0.03 train and 0.11 test. Train-test gap. 3DPS exhibits a 0.225 absolute gap between train and held-out test Pearson correlation. This reflects the cascading sparsity defined in Sec. 1. Outdoor surfaces at 60 to 77 GHz act as quasi-mirrors, so views show little cross-view consistency to regularize the held-out pose. The cascade’s large physical aperture also limits frame rate, leaving inter-viewpoint baselines at 5 Hz that are not ideal for radar NVS. Per-scene breakdown. Test correlation ranges from 0.508 (S2 F300) to 0.657 (S2 F160) across the six scenes (per-scene values: 0.644, 0.519, 0.642, 0.553, 0.657, 0.508 for S0 F135, S1 F185, S1 F438, S2 F105, S2 F160, S2 F300). The strongest scene (S2 F160) is dominated by extended 7
Table 2: Cross-method, cross-product results on six ColoRadar scenes (mean across scenes). Magnitude : Pearson correlation, PSNR, SSIM, RMSE on min-max-normalized | · |. Rel. phase : magnitude-weighted phase coherence ⟨cos(∆φ)⟩ across range (∆φr ) and across virtual antennas (∆φv ); on ADC, ∆φt is the differential across fast-time samples. All in [−1, 1], random reference 0. These metrics are absolute-phase-invariant by construction so are unaffected by per-VA / per-range calibration drift the magnitude-only training loss does not constrain. Bold: best per column (ties jointly). Underline: second-best. Per-scene breakdowns and standard deviations in supplement Tables 3–6. RA cart Method
CRP
Magnitude Magnitude Rel. phase Corr PSNR SSIM RMSE Corr PSNR SSIM ∆φr ∆φv
ADC Env. ρ|·|
Rel. phase ∆φt
Train Mean DART Huang et al. [2024] 0.033 Radar Fields Borts et al. [2024] 0.218 RadarSplat Kung et al. [2025] 0.365 3DPS (Ours) 0.812
23.9 24.9 24.9 31.9
0.540 0.0673 – – – – – – 0.480 0.0581 – – – – – – 0.526 0.0595 – – – – – – 0.819 0.0272 0.701 27.9 0.721 0.518 0.494 0.689
– – – 0.444
Test Mean DART Huang et al. [2024] 0.112 Radar Fields Borts et al. [2024] 0.133 RadarSplat Kung et al. [2025] 0.339 3DPS (Ours) 0.587
22.5 28.0 24.2 28.8
0.408 0.0817 – – – – – – 0.496 0.0402 – – – – – – 0.511 0.0643 – – – – – – 0.734 0.0369 0.603 24.7 0.650 0.371 0.434 0.613
– – – 0.414
planar walls and a parked vehicle that present consistent specular returns across the 8 training views, regularizing the held-out pose well. The weakest scene (S2 F300) is a tighter staging area with dense pedestrian-scale support structure: many small-feature targets contribute high-frequency RA support that the magnitude-only loss does not constrain at the held-out pose, increasing the train-test gap. Magnitude on CRP and ADC. The 3DPS native complex range profile gives mean training and test CRP Pearson correlation of 0.701 and 0.603. These numbers are directly comparable to the |RA| Pearson correlation in the same table. The mild attenuation reflects that per-virtual-antenna evaluation is inherently stricter than azimuth-integrated RA, since the azimuth FFT in |RA| provides √ approximately 86 of coherent gain at target directions of arrival that the per-(v,r) |CRP| does not. The ADC envelope ρ|·| reaches 0.689 train and 0.613 test. PSNR and SSIM are not reported for ADC because image-domain metrics are inappropriate for raw I/Q time series. Relative-phase fidelity. The renderer’s per-(VA, range) absolute phase is unconstrained by the magnitude-only |RA| loss — approximately 22,000 per-bin phase degrees of freedom are free up to per-VA RF calibration drift and per-range ADC sample-zero offsets. Strict complex correlation |ρ| and per-bin phase RMSE penalize that nuisance phase directly, so we instead report the magnitudeweighted phase coherence ⟨cos(∆φ)⟩ of the differential phase across each axis (denoted ∆φr , ∆φv , ∆φt in Table 2). Differential phase across an axis is invariant to any per-bin rotation along the orthogonal axis by construction, so it isolates the relative-phase content that drives downstream products (range-bin localization, beamforming) while ignoring the absolute-phase nuisance subspace the magnitude loss does not cover. The metric lies in [−1, 1] with random reference 0; perfect agreement is 1. 3DPS reaches mean training and test CRP ∆φr of 0.518 and 0.371, ∆φv of 0.494 and 0.434, and ADC ∆φt of 0.444 and 0.414 — 0.37–0.52 of perfect on the held-out view despite no phase-domain supervision. As above, the three optical-NVS baselines emit power-only RA, so the rel. phase columns are inapplicable to those methods. Only 3DPS contributes complex outputs. Qualitative comparison. Figure 3 shows held-out test |RA| renders for 3DPS and the three baselines on the six scenes, sorted left to right by descending 3DPS test Pearson correlation. 3DPS preserves dominant scene structure (walls, parked vehicles, support columns) with consistent range and azimuth localization. RadarSplat produces blurred opaque blobs from its per-Gaussian features. Radar Fields under-fits with a near-constant background dotted by sparse hot-spots, the failure mode of post-FFT supervision without a virtual-array model. DART produces noisy, low-amplitude renders dominated by Doppler ambiguity at the cascade’s Nd =16 chirps per frame. The visual ranking matches the quantitative order. Beyond magnitude, 3DPS preserves ∼40% of held-out relativephase coherence on CRP (∆φr =0.371, ∆φv =0.434) and on ADC (∆φt =0.414) without any phase supervision — meaningful structure for downstream beamforming and range-localization. Per-scene 8
S0-F135
S1-F438
S2-F105
S1-F185
S2-F300
Ours (v5_v4)
GT
S2-F160
0.64
0.64
0.55
0.52
0.51
0.56
0.41
0.26
0.37
0.25
0.15
0.10
0.11
0.04
0.21
0.09
0.17
0.08
0.03
0.07
0.24
0.04
0.03
DART
RadarFields
RadarSplat
0.66
Figure 3: Held-out novel-view |RA| comparison. Columns: six ColoRadar scenes sorted by descending 3DPS test Pearson correlation, ranging from 0.508 on S2 F300 to 0.657 on S2 F160 (mean 0.587). Rows: GT, 3DPS (Ours), RadarSplat, Radar Fields, DART. Per-frame test Pearson correlation overlaid in white. comparisons across all training frames, CRP and ADC heatmaps, runtime, and extended limitations are in supplement Sec. H.1, H.2, I, and K.
5
Conclusion
We presented 3DPS, a differentiable physically based point renderer for FMCW mmWave radar that synthesizes complex range profiles from novel sensor poses. Oriented 3D points carry ITU-R P.2040 material parameters and are evaluated by closed-form BSDFs and splatted through precomputed sensor PSFs, removing the Monte Carlo sampling and post-FFT supervision that constrain prior radarNVS methods. On six outdoor ColoRadar scenes, 3DPS reaches 0.587 mean Pearson correlation on held-out |RA|, between 1.7× and 5.2× the three baselines. Training takes approximately 3 minutes per scene over 8 viewpoints. Limitations and future work. Single-bounce evaluation suffices outdoors but not for indoor multipath. LiDAR-free radar NVS, a phase-aware loss exploiting the renderer complex output, dynamic scenes, downstream perception tasks, and broader radar geometries are natural next steps. We discuss broader impact and dual-use considerations in Sec. N. Reproducibility. Code, configuration files, trained checkpoints, and the analysis scripts that generated every table and figure in this paper will be released upon publication. The ColoRadar dataset is publicly available at https://arpg.github.io/coloradar/.
9
References Sai Praveen Bangaru, Tzu-Mao Li, and Frédo Durand. Unbiased warped-area sampling for differentiable rendering. In SIGGRAPH Asia, 2020. Dan Barnes, Matthew Gadd, Paul Murcutt, Paul Newman, and Ingmar Posner. The oxford radar robotcar dataset: A radar extension to the oxford robotcar dataset. In ICRA, 2020. Mario Bijelic, Tobias Gruber, Fahim Mannan, Florian Kraus, Werner Ritter, Klaus Dietmayer, and Felix Heide. Seeing through fog without seeing fog: Deep multimodal sensor fusion in unseen adverse weather. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020. David Borts, Erich Liang, Tim Brödermann, Andrea Ramazzina, Stefanie Walz, Edoardo Palladin, Jipeng Sun, David Bruggemann, Christos Sakaridis, Luc Van Gool, Mario Bijelic, and Felix Heide. Radar fields: Frequency-space neural scene representations for FMCW radar. In ACM SIGGRAPH 2024 Conference Papers, 2024. doi: 10.1145/3641519.3657510. Keenan Burnett, David J. Yoon, Yuchen Wu, Andrew Z. Li, Haowei Zhang, Shichen Lu, Jingxing Qian, Wei-Kang Tseng, Andrew Lambert, Keith Y. K. Leung, Angela P. Schoellig, and Timothy D. Barfoot. Boreas: A multi-season autonomous driving dataset. The International Journal of Robotics Research, 2023. Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. nuscenes: A multimodal dataset for autonomous driving. In CVPR, 2020. Xingyu Chen, Jianrong Ding, Kai Zheng, Xinmin Fang, Xinyu Zhang, Chris Xiaoxuan Lu, and Zhengxiong Li. InverTwin: Solving inverse problems via differentiable radio frequency digital twin. arXiv preprint arXiv:2508.14204, 2025. Junyuan Deng, Pengfei Xie, Lei Zhang, and Yulai Cong. ISAR-NeRF: Neural radiance fields for 3-D imaging of space target from multiview ISAR images. IEEE Sensors Journal, 24(7):11705–11722, 2024. M. Dudek, R. Wahl, D. Kissinger, R. Weigel, and G. Fischer. Millimeter wave FMCW radar system simulations including a 3D ray tracing channel simulator. In 2010 Asia-Pacific Microwave Conference Proceedings, pages 1665–1668, 2010. Thibaud Ehret, Roger Marí, Dawa Derksen, Nicolas Gasnier, and Gabriele Facciolo. Radar fields: An extension of radiance fields to SAR. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2024. Sean M. Farrell, Vivek Boominathan, Nathaniel Raymondi, Ashutosh Sabharwal, and Ashok Veeraraghavan. CoIR: Compressive implicit radar. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023. Xiangyu Gao, Guanbin Xing, Sumit Roy, and Hui Liu. Experiments with mmwave automotive radar test-bed. In 2019 53rd Asilomar Conference on Signals, Systems, and Computers, 2019. Junfeng Guan, Sohrab Madani, Suraj Jog, Saurabh Gupta, and Haitham Hassanieh. Through fog high-resolution imaging using millimeter wave radar. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020. Kyle Harlow, Hyesu Jang, Timothy D. Barfoot, Ayoung Kim, and Christopher Heckman. A new wave in robotics: Survey on recent mmWave radar applications in robotics. IEEE Transactions on Robotics, 2024. Nikolai Hofmann, Vanessa Wirth, Johanna Bräunig, Ingrid Ullmann, Martin Vossiek, Tim Weyrich, and Marc Stamminger. Inverse rendering of near-field mmWave MIMO radar for material reconstruction. IEEE Journal of Microwaves, 2025. doi: 10.1109/JMW.2025.3535077. Jakob Hoydis, Fayçal Aït Aoudia, Sebastian Cammerer, Merlin Nimier-David, Nikolaus Binder, Guillermo Marcus, and Alexander Keller. Sionna RT: Differentiable ray tracing for radio propagation modeling. arXiv preprint arXiv:2303.11103, 2023. 10
Shengyu Huang, Zan Gojcic, Zian Wang, Francis Williams, Yoni Kasten, Sanja Fidler, Konrad Schindler, and Or Litany. Neural LiDAR fields for novel view synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023. Tianshu Huang, John Miller, Akarsh Prabhakara, Tao Jin, Tarana Laroia, Zico Kolter, and Anthony Rowe. DART: Implicit doppler tomography for radar novel view synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 24118–24129, 2024. Cesar Iovescu and Sandeep Rao. The fundamentals of millimeter wave sensors. Technical Report SPYY005, Texas Instruments Application Report, 2017. ITU-R. Effects of building materials and structures on radiowave propagation above about 100 mhz. Recommendation ITU-R P.2040-3, 2023. Wenzel Jakob, Sébastien Speierer, Nicolas Roussel, and Delio Vicini. Dr.Jit: A just-in-time compiler for differentiable rendering. ACM Transactions on Graphics (Proceedings of SIGGRAPH), 41(4), 2022. M. Jankiraman. FMCW Radar Design. Artech House, 2018. Kaiwen Jiang, Jia-Mu Sun, Zilu Li, Dan Wang, Tzu-Mao Li, and Ravi Ramamoorthi. Differentiable light transport with gaussian surfels via adapted radiosity for efficient relighting and geometry reconstruction. In SIGGRAPH, 2025. Hiroharu Kato, Deniz Beker, Mihai Morariu, Takahiro Ando, Toru Matsuoka, Wadim Kehl, and Adrien Gaidon. Differentiable rendering: A survey. arXiv preprint arXiv:2006.12057, 2020. Michael Kazhdan, Matthew Bolitho, and Hugues Hoppe. Poisson surface reconstruction. In Eurographics Symposium on Geometry Processing, 2006. Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. In SIGGRAPH, 2023. Andrew Kramer, Kyle Harlow, Christopher Williams, and Christoffer Heckman. ColoRadar: The direct 3D millimeter wave radar dataset. The International Journal of Robotics Research, 41(4), 2022. Pou-Chun Kung, Skanda Harisha, Ram Vasudevan, Aline Eid, and Katherine A. Skinner. RadarSplat: Radar gaussian splatting for high-fidelity data synthesis and 3D reconstruction of autonomous driving scenes. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 27596–27606, 2025. Zhengxin Lei, Feng Xu, Jiangtao Wei, Feng Cai, Feng Wang, and Ya-Qiu Jin. SAR-NeRF: Neural radiance fields for synthetic aperture radar multi-view representation. IEEE Transactions on Geoscience and Remote Sensing, 2024. Tzu-Mao Li, Miika Aittala, Frédo Durand, and Jaakko Lehtinen. Differentiable monte carlo ray tracing through edge sampling. In SIGGRAPH Asia, 2018. Guillaume Loubet, Nicolas Holzschuch, and Wenzel Jakob. Reparameterizing discontinuous integrands for differentiable rendering. In SIGGRAPH Asia, 2019. Chris Xiaoxuan Lu, Stefano Rosa, Peijun Zhao, Bing Wang, Changhao Chen, John A. Stankovic, Niki Trigoni, and Andrew Markham. See through smoke: Robust indoor mapping with lowcost mmWave radar. In Proceedings of the 18th International Conference on Mobile Systems, Applications, and Services (MobiSys), 2020. Haofan Lu, Christopher Vattheuer, Baharan Mirzasoleiman, and Omid Abari. NeWRF: A deep learning framework for wireless radiation field reconstruction and channel prediction. In Proceedings of the 41st International Conference on Machine Learning (ICML), pages 33147–33159, 2024. Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. NeRF: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM, 65(1):99–106, 2021. 11
Meysam Moallem and Kamal Sarabandi. Polarimetric study of MMW imaging radars for indoor navigation and mapping. IEEE Transactions on Antennas and Propagation, 62(1):500–504, 2014. Thomas Müller, Alex Evans, Christoph Schied, and Alexander Keller. Instant neural graphics primitives with a multiresolution hash encoding. ACM Transactions on Graphics (ToG), 41(4), 2022. Julian Ost, Issam Laradji, Alejandro Newell, Yuval Bahat, and Felix Heide. Neural point light fields. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022. Dong-Hee Paek, Seung-Hyun Kong, and Kevin Tirta Wijaya. K-Radar: 4d radar object detection for autonomous driving in various weather conditions. In NeurIPS Datasets and Benchmarks Track, 2022. Ziyuan Qu, Omkar Vengurlekar, Mohamad Qadri, Kevin Zhang, Michael Kaess, Christopher Metzler, Suren Jayasuriya, and Adithya Pediredla. Z-Splat: Z-axis Gaussian splatting for camera–sonar fusion. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024. Mahan Rafidashti, Ji Lan, Maryam Fatemi, Junsheng Fu, Lars Hammarstrand, and Lennart Svensson. NeuRadar: Neural radiance fields for automotive radar point clouds. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2025. Andrea Ramazzina, Mario Bijelic, Stefanie Walz, Alessandro Sanvito, Dominik Scheuble, and Felix Heide. ScatterNeRF: Seeing through fog with physically-based inverse neural rendering. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 17957–17968, 2023. Julien Rebut, Arthur Ouaknine, Waqas Malik, and Patrick Pérez. Raw high-definition radar for multi-task learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022. Albert Reed, Juhyeon Kim, Thomas Blanford, Adithya Pediredla, Daniel Brown, and Suren Jayasuriya. Neural volumetric reconstruction for coherent synthetic aperture sonar. ACM Transactions on Graphics, 42(4), 2023. Mark A. Richards. Fundamentals of Radar Signal Processing. McGraw-Hill Education, 2nd edition, 2014. Ole Schumann, Markus Hahn, Nicolas Scheiner, Fabio Weishaupt, Julius F. Tilly, Jürgen Dickmann, and Christian Wöhler. RadarScenes: A real-world radar point cloud data set for automotive applications. In International Conference on Information Fusion (FUSION), 2021. Christian Schüßler, Marcel Hoffmann, Johanna Bräunig, Ingrid Ullmann, Randolf Ebelt, and Martin Vossiek. A realistic radar ray tracing simulator for large mimo-arrays in automotive environments. IEEE Journal of Microwaves, 1(4):962–974, 2021. Advaith V. Sethuraman, Max Rucker, Onur Bagoren, Pou-Chun Kung, Nibarkavi N. B. Amutha, and Katherine A. Skinner. SonarSplat: Novel view synthesis of imaging sonar via Gaussian splatting. IEEE Robotics and Automation Letters, 2025. Marcel Sheeny, Emanuele De Pellegrin, Saptarshi Mukherjee, Alireza Ahrabian, Sen Wang, and Andrew Wallace. RADIATE: A radar dataset for automotive perception in bad weather. In 2021 IEEE International Conference on Robotics and Automation (ICRA), 2021. Merrill I. Skolnik. Introduction to Radar Systems. McGraw-Hill, New York, 3rd edition, 2001. Harshvardhan Takawale and Nirupam Roy. SpINR: Neural volumetric reconstruction for FMCW radars. arXiv preprint arXiv:2503.23313, 2025. Tang Tao, Longfei Gao, Guangrun Wang, Yixing Lao, Peng Chen, Hengshuang Zhao, Dayang Hao, Xiaodan Liang, Mathieu Salzmann, and Kaicheng Yu. LiDAR-NeRF: Novel LiDAR view synthesis via neural radiance fields. In Proceedings of the 32nd ACM International Conference on Multimedia, 2024. 12
Mehrnoosh Vahidpour and Kamal Sarabandi. Millimeter-wave Doppler spectrum and polarimetric response of walking bodies. IEEE Transactions on Geoscience and Remote Sensing, 50(7): 2866–2879, 2012. Bruce Walter, Stephen R Marschner, Hongsong Li, and Kenneth E Torrance. Microfacet models for refraction through rough surfaces. Eurographics Symposium on Rendering, 2007. Jiangtao Wei, Yixiang Luomei, Xu Zhang, and Feng Xu. Learning surface scattering parameters from SAR images using differentiable ray tracing. IEEE Transactions on Geoscience and Remote Sensing, 2024. Haolin Xiong, Sairisheek Muttukuru, Rishi Upadhyay, Pradyumna Chari, and Achuta Kadambi. SparseGS: Real-time 360 sparse view synthesis using Gaussian splatting. arXiv preprint arXiv:2312.00206, 2023. Yunzhi Yan, Haotong Lin, Chenxu Zhou, Weijie Wang, Haiyang Sun, Kun Zhan, Xianpeng Lang, Xiaowei Zhou, and Sida Peng. Street Gaussians: Modeling dynamic urban scenes with Gaussian splatting. In European Conference on Computer Vision (ECCV), 2024. Jiarui Zhang, Zhihao Li, Chong Wang, and Bihan Wen. RF4D: Neural radar fields for novel view synthesis in outdoor dynamic scenes. arXiv preprint arXiv:2505.20967, 2025.
13
The supplement opens with a self-contained primer on FMCW mmWave radar signal processing for readers more familiar with vision/graphics than radar hardware (Sec. A). It then provides forwardmodel details deferred from Sec. 3.1 (full surface-to-hemisphere derivation, PSF-splat validity bounds, full ITU-R P.2040 BSDF closed form, LiDAR reduction parameters), the per-scene quantitative evidence and qualitative comparisons that anchor the main paper’s claims, the methodology details for the CRP and ADC evaluation introduced in Sec. 4.2, implementation details (wall-clock, hardware), and a more detailed discussion of limitations and future work.
A
FMCW mmWave radar primer
This section summarizes the FMCW mmWave radar signal-processing chain that 3DPS targets, for readers more familiar with vision, graphics, or general machine learning than with radar hardware. Notation is kept consistent with the main paper. We follow the standard treatments in Skolnik Skolnik [2001] and Richards Richards [2014], with cascade-radar parameters as in the TI MMWCAS reference Iovescu and Rao [2017]. A.1
FMCW chirp and dechirped ADC samples
FMCW radar Jankiraman [2018] transmits a linear-frequency-sweep waveform (a “chirp”) of duration Tc that rises from carrier frequency fc at slope µ = B/Tc over total swept bandwidth B: sTX (t) = cos 2π(fc t + 12 µ t2 ) , t ∈ [0, Tc ]. (6) A reflection from a scatterer at total round-trip path length R = Rt + Rr arrives at the receiver after delay τ = R/c. Mixing the received signal with a delayed copy of sTX (“dechirping”) and low-pass filtering produces a beat signal whose instantaneous frequency is constant over the chirp at µτ . After analog-to-digital conversion at sample rate fs , the per-chirp dechirped baseband samples are a[n] = Ar e j2π µ τ tn ,
tn = n/fs ,
n = 0, . . . , Ns − 1,
(7)
where the complex amplitude Ar aggregates antenna gains, scattering cross section, path-length attenuation, and the carrier-phase term e−jk(Rt +Rr ) with k = 2π/λ. Eq. (7) is the single-scatterer complex baseband response, integrated over the visible scene in our Eq. (2). A.2
Complex baseband (I/Q) and the range profile
The dechirp mixer is implemented in I/Q form so that each ADC sample is complex, a[n] = aI [n] + j aQ [n]. The complex form is what makes radar a phase-coherent imaging modality: the magnitude |a[n]| encodes scatterer reflectivity, while the phase argument retains fractional-wavelength path information. At a 76 GHz carrier the wavelength is λ ≈ 3.9 mm, an order of magnitude finer than the range-bin width ∆R = c/(2B) ≈ 5.93 cm produced by the range FFT below. Optical-NVS adaptations (Sec. 2) train on | · | alone and discard this phase content; 3DPS retains it through the splat in Eq. (4). The complex range profile (CRP) is the windowed range FFT of the chirp’s ADC samples, CRP[k] =
NX s −1
w[n] a[n] e−j2πnk/Ns ,
k = 0, . . . , Ns − 1,
(8)
n=0
with w[·] a windowing function (we use Hann, sidelobes ∼−50 dB). A scatterer at delay τ appears as a peak at fractional bin k ⋆ = µτ Ns /fs , smeared by the discrete-Fourier image of w across neighboring bins — the “point spread function” our PSF splat (Eq. 4) inverts. A.3
TDM-MIMO and the virtual array
A MIMO radar array transmits and receives on multiple antenna elements. Time-division multiplexing (TDM) cycles through TX elements one at a time within a chirp loop, so each (TX, RX) pair gets its own dedicated chirp and resulting CRP. Under far-field assumptions and electrically small array baselines, every (TX, RX) pair behaves like a single “virtual element” located at the position pv (t, r) = pt + pr , 14
(9)
the convolution of the TX and RX position sets Iovescu and Rao [2017]. Thus NTX ·NRX physical channels yield up to NTX NRX virtual elements at the locations of {pt + pr }. For the TI MMWCAS cascade (NTX =12, NRX =16), the 192 channels reduce to 86 unique virtual positions after de-duplication, all on a single horizontal axis. The cascade has zero elevation extent in its virtual array, which fixes the 2D-only supervision constraint discussed below. A.4
Range–azimuth (RA) image and the 2D-only constraint
Stacking per-(TX, RX) CRPs along the virtual-element axis gives a virtual-array CRP of shape (NVA , Ns ). Applying a windowed FFT across the virtual-element axis converts inter-element phase progression (which encodes the angle of arrival of each scatterer) into an azimuth-bin index, yielding the range–azimuth image RA[a, k]. Bin widths are ∆R = c/(2B) ≈ 5.93 cm in range and approximately θres ≈ λ/DVA ≈ 1.4◦ in azimuth, where DVA is the virtual-array aperture (∼16 cm for the cascade’s 86-element extent). Crucially, a planar virtual array with zero elevation aperture cannot resolve elevation: scatterers at any height in the same (range, azimuth) cell project to the same RA pixel. RA supervision therefore integrates out the elevation dimension by construction, which is the underlying reason for the cascading-sparsity argument in Sec. 1: 2D RA + sparse views + sparse returns make 3D scene structure information-theoretically unrecoverable from RA alone. A.5
Polar-to-Cartesian conversion for evaluation
The native RA image is intrinsically polar — rows index range bins k, columns index azimuth bins a — and is the form on which all radar imaging metrics (Pearson correlation, PSNR, SSIM, RMSE) are conventionally reported in the ColoRadar benchmark. For human inspection and to keep image-domain metrics meaningful in physical units, we resample the polar RA onto a Cartesian grid (x, y) in the radar frame: √ 2 2 x +y RAcart [x, y] = RApolar k(x, y), a(x, y) , k = ∆R , a = N2az 1 + Naz2+1 sin θ , (10) with θ = arctan(x/y) and bilinear interpolation in (k, a). The arcsin-spaced inverse for a(x, y) matches the azimuth-bin grid in Sec. 3.2. We use a 399×399 Cartesian grid spanning the radar’s forward cone, retain range bins 15–110 to reject TX–RX coupling near the origin and DFT wraparound of near-field energy at the far range, and apply the same interpolation harness to all baselines for fairness. A.6
ITU-R P.2040 mmWave material model
ITU-R Recommendation P.2040 ITU-R [2023] specifies bistatic radio-wave scattering parameters for common surfaces in the 100 MHz–100 GHz band. Each surface is parameterized by complex relative permittivity εr = ε′r − jε′′r , RMS surface roughness σh , and (for layered slabs) thickness d. The standard provides closed-form expressions for the coherent specular reflection coefficient (with the Ament roughness attenuation exp(−(2kσh cos θi )2 )) and an incoherent diffuse lobe parameterized by surface correlation length ℓc . 3DPS attaches a per-point 6-vector θ i = (ε′r , ε′′r , σh , ℓc , τ, d)i to each oriented point, where τ is the Kirchhoff/SPM blend introduced in Eq. (5); the full closed form is in Sec. D.
B
Forward model: surface-to-hemisphere derivation
Building on the FMCW signal model in Sec. A, this section derives Eq. (2) (the hemisphere ADC integral) starting from the surface integral form. The discretization in Eq. (3) then follows from a Riemann sum. Surface integral form. For a single-bounce path TX → x → RX over a scene surface S, the dechirped FMCW ADC sample is the surface integral Z r (x) atr [n] = ⊮[n(x)· ω̂ o (x) > 0] Str (x) e j2π µ τ (x) tn dA(x), τ (x) = Rt (x)+R , (11) c S
15
with Str the complex amplitude of Eq. (1) and the indicator enforcing visibility from the receiver. Solid-angle change of variables. Each visible surface element dA(x) at range Rr (x) from the receiver subtends solid angle |n(x)· ω̂ o | dA(x) dω = . Rr2 (x) Inverting this Jacobian, dA(x) = Rr2 /|n·ω̂ o | dω, and parameterizing the visible scene by the receivercentered upper hemisphere Ω+ (pr ) (with x(ω̂) the first scene intersection along ω̂ from pr ), Eq. (11) becomes the hemisphere integral of Eq. (2). First-hit ray casting absorbs the visibility indicator: every direction ω̂ ∈ Ω+ maps to its first intersection, which is by construction visible from pr . Riemann discretization with Jacobian cancellation. We approximate Ω+ (pr ) by the oriented point set {(xi , ni , Ai , θ i )}N i=1 , with each point covering solid angle ∆ωi =
Ai |ni · ω̂ o,i | . 2 Rr,i
Substituting into Eq. (2), the integrand’s Rr2 /|n· ω̂| Jacobian cancels exactly against ∆ωi , and the Riemann sum collapses to Eq. (3). Visibility re-enters at discretization. The hemisphere integral does not need an explicit visibility indicator (first-hit ray casting handles it implicitly), but a LiDAR-derived oriented point set is not in general a hemisphere-conformal partition of Ω+ : back-facing or self-occluded points violate the implicit first-hit assumption. We therefore reintroduce a strict front-facing indicator vi = ⊮[ni·ω̂ o,i > 0] in Eq. (3), and remove self-occluding points up-front via Mitsuba 3 ray-casting in the LiDAR reduction pipeline (Sec. E).
C
Validity of the PSF splat
Eq. (4) replaces the per-sample chirp synthesis of Eq. (3) with an L-tap scatter-add into range bins. We bound the two approximation sources here. Hann-tail truncation. The Hann-windowed PSF Φ(·; δ) is the discrete Fourier transform of a length-Ns Hann window evaluated at fractional bin offset δ ∈ [0, 1). Its sidelobe envelope falls below −50 dB beyond L=15 taps relative to peak, well under the noise floor of real ColoRadar captures. Truncation to 15 taps therefore introduces no measurable energy loss into the rendered CRP. Fractional-bin handling. The splat support index ⌊ki ⌋ rounds the path’s bistatic range to the nearest bin, but the carrier phase e−jk(Rt,i +Rr,i ) in Str (xi ) (Eq. 4) is computed from the unrounded Rt,i + Rr,i , and the kernel Φ(·; δi ) encodes the fractional offset δi = ki − ⌊ki ⌋ exactly. The bin rounding therefore affects only the splat support (which L adjacent bins receive contributions), not the per-bin amplitude or phase. Numerical equivalence to the standard ADC-then-FFT pipeline. Direct evaluation of Eq. (3) followed by a length-Ns Hann-windowed range FFT, against the splat of Eq. (4), on identical (point, TX, RX) inputs gives a maximum forward error of 1.19 × 10−7 across 90,000 path samples — essentially fp32 machine precision. The splat is therefore numerically equivalent to the standard ADC-then-FFT pipeline on the same path set, not an approximation of it.
D
Full ITU-R P.2040 BSDF closed form
Following the ITU-R P.2040 introduction in Sec. A.6, the BSDF σ(ω̂ i , ω̂ o ; θ i ) in Eq. (5) expands as follows. Multilayer-slab Fresnel. The reflection coefficient Γ for a slab of thickness d and complex permittivity εr = ε′r − jε′′r tracks the coherent superposition of the air–material reflection and the internal round-trip through the slab, evaluated independently on the s and p polarizations of the local Jones basis (mmWave polarimetric scattering studies in Vahidpour and Sarabandi [2012], Moallem and Sarabandi [2014] validate the per-polarization treatment at our 76–77 GHz operating band). This captures internal reflections in thin layered materials (plasterboard, vehicle bodies, painted metal) 16
that single-interface Fresnel cannot model. We use the canonical multilayer-slab form documented in ITU-R P.2040 ITU-R [2023]; the coefficient is differentiable in (ε′r , ε′′r , d) in closed form. Coherent specular kernel. Kspec blends a Kirchhoff-approximation (KA) lobe KKA and a small-perturbation-method (SPM) lobe KSPM with weight τ ∈ [0, 1]: Kspec (ω̂ i , ω̂ o ; n, τ ) = τ KKA (ω̂ i , ω̂ o ; n) + (1 − τ ) KSPM (ω̂ i , ω̂ o ; n). The two kernels follow the standard mmWave forms (KA: GGX-style microfacet specular about the mirror direction Walter et al. [2007]; SPM: von Mises– Fisher lobe scaled by the surface power spectrum). τ is reparameterized via sigmoid so optimization remains in [0, 1]. Coherent attenuation (Ament). ρcoh (σh ; θi ) = exp −(2kσh cos θi )2 with k = 2π/λ, attenuating the specular lobe with surface RMS height σh . Incoherent diffuse lobe. ρinc combines a directional and a Lambertian component, ρinc = γ Ldir (ω̂ i , ω̂ o , n; σh , ℓc ) + (1 − γ) Llam (n · ω̂ o ), where γ is set from the surface roughness slope σh /ℓc . The directional lobe’s angular width grows with σh /ℓc as physical roughness increases. Reparameterizations and bounds. All positive-only material parameters (ε′r , ε′′r , σh , ℓc , d) are reparameterized via shifted softplus to keep optimization in physically valid ranges. The τ ∈ [0, 1] blend uses a sigmoid. Per-point material initialization is uniform ITU-concrete ITU-R [2023], with per-point freedom afterward. The full closed form follows the standard per-lobe ITU-R P.2040 forms ITU-R [2023], applied per-point rather than per-vertex.
E
LiDAR reduction parameters
The LiDAR-to-oriented-point-set reduction in Sec. 3.3 runs in four stages on each scene’s coregistered LiDAR cloud (positions, normals, intensity; ∼2.7–6.5M points across our scenes), reducing it to exactly N =20,000 oriented points. We use the raw LiDAR cloud rather than a Poissonreconstructed surface mesh Kazhdan et al. [2006]: the radar’s beamwidth and range resolution are both coarser than typical reconstruction error, so explicit surface fitting offers no benefit while losing the per-point density signal the renderer leverages. Stage 1 — frustum and range cull. Points with cos(∠(d̂, b̂)) ≤ 0.1761 (boresight half-angle ≈79.8◦ matching the radar’s azimuth FOV) or radial range outside [1.5 m, K∆R=15.2 m] are dropped. This typically removes 80–95% of the raw LiDAR points. Stage 2 — occlusion cull. Surviving points are ray-cast toward the array centroid via Mitsuba 3 Jakob et al. [2022]; any point whose ray intersects another scene element before reaching the centroid is dropped as self-occluded. This is what motivates the explicit vi visibility indicator in Eq. (3) (the residual self-occlusion that survives Stage 2 is handled at evaluation time). Stage 3 — cosine-weighted resampling. We draw 3N candidates with replacement at probability ∝ max(cos(∠(d̂, b̂)), 0.01), biasing selection toward boresight-aligned regions where the antenna pattern is strongest while preserving a long tail of off-boresight returns. Stage 4 — farthest-point sampling. The 3N candidates are reduced to exactly N =20,000 via greedy farthest-point sampling, yielding a near-uniform spatial distribution across the boresight-biased candidate pool. Each surviving point inherits a quaternion encoding its LiDAR surface normal and a uniform ITU(0) concrete material prior θ i ITU-R [2023]. Per-point material drift afterward is unconstrained (Sec. 3.3).
F
Per-scene |RA| breakdowns
The main paper Table 2 reports the 6-scene mean for each method on training views (8 views per scene) and the held-out test view. Tables 3 and 4 below give the per-scene breakdown that anchors those means, retaining the canonical |RA| metric set (Corr, PSNR, SSIM, RMSE on min-maxnormalized 399×399 Cartesian images, computed via a shared evaluation harness applied identically to all methods). The Mean row of each table matches the corresponding 3DPS row in Table 2 by construction.
17
Table 3: Per-scene held-out test |RA| metrics on six ColoRadar scenes. Bold: best per scene per metric (ties bolded jointly). Underline: second-best. The Mean row anchors the corresponding test row of main paper Table 2, and the Std row reports the per-method standard deviation across the six scenes. RA Corr ↑ Scene
S0 F135 0.644 S1 F185 0.519 S1 F438 0.642 S2 F105 0.553 S2 F160 0.657 S2 F300 0.508 Mean Std
RA PSNR ↑
RA SSIM ↑
RA RMSE ↓
Ours RSplat RFields DART Ours RSplat RFields DART Ours RSplat RFields DART Ours RSplat RFields DART 0.407 0.305 0.226 0.369 0.559 0.168
0.058 0.087 0.210 0.226 0.075 0.141
0.114 0.057 0.060 0.241 0.085 0.117
31.3 29.5 27.1 27.6 27.7 29.4
26.4 26.9 22.7 27.0 20.8 21.5
28.8 28.2 27.4 28.2 26.2 28.9
16.6 21.1 21.6 22.4 27.8 25.3
0.839 0.780 0.663 0.683 0.653 0.785
0.545 0.618 0.502 0.518 0.372 0.512
0.616 0.583 0.435 0.439 0.319 0.585
0.177 0.276 0.349 0.516 0.687 0.446
0.0273 0.0478 0.0333 0.0452 0.0442 0.0732 0.0415 0.0446 0.0413 0.0908 0.0340 0.0841
0.587 0.339 0.068 0.140
0.133 0.072
0.112 28.8 0.068 1.58
24.2 2.86
28.0 1.00
22.5 3.83
0.734 0.511 0.077 0.080
0.496 0.117
0.408 0.0369 0.0643 0.0402 0.0817 0.182 0.0064 0.0210 0.0048 0.0372
0.0363 0.0389 0.0426 0.0388 0.0488 0.0360
0.1481 0.0880 0.0827 0.0760 0.0407 0.0545
Table 4: Per-scene training-view |RA| metrics on six ColoRadar scenes. Each cell is the mean over the 8 training frames per scene. Same metric harness as Table 3. Bold: best per scene per metric (ties bolded jointly). Underline: second-best. The Mean row anchors the corresponding train row of main paper Table 2, and the Std row reports the per-method standard deviation across the six scenes. RA Corr ↑ Scene
S0 F135 0.822 S1 F185 0.836 S1 F438 0.848 S2 F105 0.804 S2 F160 0.784 S2 F300 0.774 Mean Std
G
RA PSNR ↑
RA SSIM ↑
RA RMSE ↓
Ours RSplat RFields DART Ours RSplat RFields DART Ours RSplat RFields DART Ours RSplat RFields DART 0.356 0.354 0.268 0.427 0.575 0.211
0.812 0.365 0.029 0.128
0.308 -0.028 32.0 0.196 0.016 32.9 0.091 0.033 32.6 0.229 0.153 31.3 0.306 -0.038 29.7 0.174 0.062 32.9
25.5 28.0 23.2 26.5 22.2 24.0
25.6 26.3 23.4 24.0 23.6 26.7
22.5 26.0 25.3 23.6 20.4 25.8
0.837 0.819 0.852 0.815 0.743 0.847
0.540 0.639 0.512 0.494 0.377 0.596
0.522 0.594 0.419 0.409 0.296 0.639
0.468 0.484 0.578 0.656 0.541 0.511
0.033 31.9 0.070 1.25
24.9 2.16
24.9 1.43
23.9 2.20
0.819 0.526 0.040 0.091
0.480 0.129
0.540 0.0272 0.0595 0.0581 0.0673 0.069 0.0038 0.0146 0.0088 0.0173
0.218 0.083
0.0261 0.0540 0.0266 0.0406 0.0241 0.0707 0.0289 0.0477 0.0338 0.0790 0.0236 0.0649
0.0527 0.0495 0.0684 0.0633 0.0661 0.0487
0.0763 0.0504 0.0578 0.0678 0.0968 0.0546
CRP and ADC evaluation
Because 3DPS produces complex range profiles, the same optimized scene yields range–azimuth and raw ADC outputs by applying the standard azimuth FFT and inverse range FFT to the per-(TX, RX) range-profile output of Eq. (4) — with no retraining. The aggregate fidelity numbers were reported in Table 2 of the main paper. This section provides the eval pipeline conventions and the per-scene CRP/ADC tables that anchor the 3DPS row of Table 2. G.1
Pipeline conventions
CRP and ADC are evaluated in the trainer’s loss domain to enable direct comparison with the published |RA| numbers. The GT side applies range Hann + range FFT to the raw ADC; both GT and predicted CRPs then receive the azimuth Hann window on the virtual-array axis, matching the conventions used during training. Range bins [0:15] are zeroed in both signals to reject TX–RX coupling. ADC is the inverse range FFT of the resulting CRP. Magnitude metrics (ρ, PSNR, SSIM, ρ|·| ) are computed on independently min-max-normalized | · |, identical to the |RA| convention. Relative-phase metrics (∆φr , ∆φv , ∆φt ) are computed on the raw (uncorrected) signals and are invariant to per-bin absolute phase, so they require no calibration removal. A round-trip FFT/IFFT sanity test (no Hann, no zero-pad) is exact to 5×10−16 relative error on every scene. G.2
Per-scene CRP and ADC fidelity
Tables 5 and 6 give the per-scene CRP/ADC fidelity that anchors the 3DPS row of Table 2. For each scene, the value is the 8-train-frame mean (train table) or the held-out test frame value (test table); the Mean row at the bottom of each is the average across the six scenes and matches the corresponding 3DPS Train Mean / Test Mean row of Table 2 by construction. Same metric set and column layout as Table 2.
18
Table 5: Per-scene CRP and ADC fidelity (training-view), 3DPS (Ours) only. Mean row anchors the 3DPS (Ours) train row of main paper Table 2. Std is across the six scenes. Magnitude columns: Pearson correlation, PSNR, SSIM, RMSE on min-max-normalized | · |. Rel. phase columns: magnitude-weighted ⟨cos(∆φ)⟩ across range (∆φr ), across virtual antennas (∆φv ), and across fast-time (∆φt ); all in [−1, 1], random reference 0. CRP
ADC
Scene
Corr
Magnitude PSNR SSIM RMSE
Rel. phase ∆φr ∆φv
Env. ρ|·|
Rel. phase ∆φt
S0 F135 S1 F185 S1 F438 S2 F105 S2 F160 S2 F300
0.677 0.741 0.718 0.690 0.671 0.709
28.1 26.5 29.6 27.3 26.4 29.6
0.697 0.753 0.742 0.707 0.627 0.798
0.598 0.533 0.468 0.494 0.435 0.579
0.528 0.507 0.586 0.395 0.376 0.569
0.665 0.674 0.725 0.666 0.681 0.719
0.422 0.757 0.128 0.422 0.419 0.519
Mean Std
0.701 0.027
27.9 1.44
0.721 0.0422 0.518 0.494 0.058 0.0068 0.064 0.088
0.689 0.027
0.444 0.203
0.0398 0.0483 0.0343 0.0449 0.0505 0.0352
Table 6: Per-scene CRP and ADC fidelity (held-out test), 3DPS (Ours) only. Mean row anchors the 3DPS (Ours) test row of main paper Table 2. Std is across the six scenes. Magnitude columns: Pearson correlation, PSNR, SSIM, RMSE on min-max-normalized | · |. Rel. phase columns: magnitude-weighted ⟨cos(∆φ)⟩ across range (∆φr ), across virtual antennas (∆φv ), and across fast-time (∆φt ); all in [−1, 1], random reference 0. CRP
ADC
Scene
Corr
Magnitude PSNR SSIM RMSE
Rel. phase ∆φr ∆φv
Env. ρ|·|
Rel. phase ∆φt
S0 F135 S1 F185 S1 F438 S2 F105 S2 F160 S2 F300
0.626 0.594 0.554 0.611 0.615 0.619
26.2 26.0 25.0 23.2 23.6 24.2
0.714 0.705 0.632 0.639 0.536 0.671
0.345 0.437 0.352 0.336 0.317 0.439
0.514 0.482 0.590 0.309 0.266 0.445
0.667 0.612 0.586 0.590 0.588 0.636
0.505 0.711 -0.019 0.434 0.460 0.395
Mean Std
0.603 0.026
24.7 1.24
0.650 0.0589 0.371 0.434 0.065 0.0083 0.053 0.124
0.613 0.033
0.414 0.239
0.0492 0.0501 0.0565 0.0694 0.0662 0.0617
H
Qualitative comparisons
H.1
Per-scene |RA| comparison across methods
Figures 4–9 show per-scene side-by-side |RA| renders for all four methods on every training frame, plus the held-out test frame (F , last column with the red header). Rows: GT, 3DPS (Ours), RadarSplat, Radar Fields, DART (cascaded). Per-cell train/test CC overlaid in white. These figures complement the per-scene tables in Sec. F and confirm visually that 3DPS preserves wall and vehicle structure across all eight training frames as well as the held-out test frame, while the three baselines exhibit the failure modes discussed in Sec. 4.2. H.2
CRP and ADC heatmaps and I/Q traces
Figures 10 and 11 show the qualitative comparison of |CRP| and |ADC| between GT and 3DPS across all six scenes (held-out test frames). Only 3DPS is shown alongside GT because it is the only method in our benchmark that emits complex outputs — the three optical-NVS baselines (RadarSplat, Radar Fields, DART) emit magnitude-only RA and have no CRP/ADC outputs (Sec. 2). Each panel 19
F¡3
F¡2
F¡1
F
F+1
F+2
F+3
F+4
Ours
GT
F¡4
0.85
0.81
0.84
0.64
0.84
0.81
0.77
0.76
0.19
0.21
0.38
0.41
0.41
0.46
0.52
0.43
0.16
0.05
0.21
0.13
0.28
0.11
0.13
0.15
0.22
0.25
-0.04
-0.04
-0.04
-0.04
0.03
-0.05
-0.04
-0.03
0.03
DART
Radar Fields
RadarSplat
0.90
Figure 4: Per-frame |RA| comparison, scene S0 F135. Cols: 8 train frames (F ±1. . .±4) plus held-out test frame F (red header). Rows: GT, 3DPS, RadarSplat, Radar Fields, DART. Per-frame Pearson correlation overlaid. F¡3
F¡2
F¡1
F
F+1
F+2
F+3
F+4
Ours
GT
F¡4
0.79
0.74
0.87
0.52
0.85
0.87
0.85
0.90
0.19
0.23
0.47
0.39
0.25
0.46
0.22
0.74
0.35
0.17
0.09
0.06
0.14
0.09
0.21
0.12
0.10
0.12
-0.01
0.00
-0.02
-0.02
0.04
-0.01
0.05
0.01
0.12
DART
Radar Fields
RadarSplat
0.81
Figure 5: Per-frame |RA| comparison, scene S1 F185. is normalized identically to its |RA| counterpart in the main paper, so the visual structure is directly comparable.
20
F¡3
F¡2
F¡1
F
F+1
F+2
F+3
F+4
Ours
GT
F¡4
0.75
0.80
0.86
0.64
0.87
0.85
0.88
0.92
0.57
0.35
0.34
0.27
0.26
0.18
0.21
0.17
0.03
0.02
0.06
-0.01
0.07
0.04
0.09
0.06
0.11
0.05
0.03
0.03
0.04
0.02
0.07
0.02
0.05
0.01
0.06
DART
Radar Fields
RadarSplat
0.85
Figure 6: Per-frame |RA| comparison, scene S1 F438.
F¡3
F¡2
F¡1
F
F+1
F+2
F+3
F+4
Ours
GT
F¡4
0.78
0.80
0.80
0.55
0.80
0.83
0.85
0.83
0.51
0.33
0.48
0.50
0.37
0.39
0.44
0.43
0.36
0.12
0.09
0.20
0.37
0.21
0.10
0.08
0.07
0.28
0.05
0.10
0.18
0.14
0.24
0.24
0.17
0.18
0.18
DART
Radar Fields
RadarSplat
0.74
Figure 7: Per-frame |RA| comparison, scene S2 F105.
21
F¡3
F¡2
F¡1
F
F+1
F+2
F+3
F+4
Ours
GT
F¡4
0.79
0.80
0.72
0.66
0.85
0.83
0.76
0.73
0.61
0.65
0.60
0.57
0.56
0.57
0.68
0.59
0.44
0.08
0.21
0.08
0.35
0.10
0.14
0.12
0.14
0.19
0.00
-0.03
-0.05
-0.09
0.08
-0.08
-0.05
0.00
-0.03
DART
Radar Fields
RadarSplat
0.79
Figure 8: Per-frame |RA| comparison, scene S2 F160.
F¡3
F¡2
F¡1
F
F+1
F+2
F+3
F+4
Ours
GT
F¡4
0.72
0.79
0.86
0.51
0.81
0.75
0.82
0.74
0.26
0.20
0.13
0.17
0.15
0.23
0.27
0.24
0.24
0.02
0.06
0.00
0.39
0.17
-0.00
0.06
0.13
0.03
-0.02
0.04
0.02
0.13
0.03
0.09
0.14
0.05
0.03
DART
Radar Fields
RadarSplat
0.70
Figure 9: Per-frame |RA| comparison, scene S2 F300.
22
jCRPj
Ours
GT
jADCj
Ours
S0-F135
GT
½=0.67 j½j VR =0.31 ¾Á =79°
½=0.59 PSNR=26.0 SSIM=0.71 j½j VR =0.23
½=0.61 j½j VR =0.23 ¾Á =85°
½=0.55 PSNR=25.0 SSIM=0.63 j½j VR =0.28
½=0.59 j½j VR =0.28 ¾Á =81°
½=0.61 PSNR=23.2 SSIM=0.64 j½j VR =0.22
½=0.59 j½j VR =0.22 ¾Á =84°
½=0.62 PSNR=23.6 SSIM=0.54 j½j VR =0.20
½=0.59 j½j VR =0.20 ¾Á =86°
S2-F300
S2-F160
S2-F105
S1-F438
S1-F185
½=0.63 PSNR=26.2 SSIM=0.71 j½j VR =0.31
½=0.64 j½j VR =0.29 ¾Á =80°
½=0.62 PSNR=24.2 SSIM=0.67 j½j VR =0.29
ADC sample
range bin (offset by 15)
Figure 10: |CRP| and |ADC| heatmaps on held-out test frames. Rows: six ColoRadar scenes. Top block: |CRP|, range bins [15:256], columns GT vs 3DPS. Bottom block: |ADC|, the inverse range FFT of the near-field-masked CRP, columns GT vs 3DPS. Panels are min-max normalized to [0, 1]. Per-cell overlay: ρ / PSNR / SSIM (magnitude) and ∆φr / ∆φv (CRP magnitude-weighted phase coherence ⟨cos(∆φ)⟩ in [−1, 1], random reference 0); envelope ρ|·| and ∆φt (ADC).
23
CRP S0-F135
VA 41 j½j VR = 0:31
¾Á = 79°
VA 37 j½j VR = 0:23
¾Á = 85°
VA 45 j½j VR = 0:28
¾Á = 81°
VA 49 j½j VR = 0:22
¾Á = 84°
VA 43 j½j VR = 0:20
¾Á = 86°
VA 48 j½j VR = 0:29
¾Á = 80°
0
−1
−1 VA 37 PSNR VR ℂ = 22:6 dB
1
S1-F185
1
0
0
−1
−1 VA 45 PSNR VR ℂ = 21:8 dB
1 0 −1
VA 49 PSNR VR ℂ = 19:1 dB
1 0 −1
1
amplitude (norm.)
S1-F438
1
0
amplitude (norm.)
S2-F105 S2-F160 S2-F300
ADC VA 41 PSNR VR ℂ = 23:2 dB
1
0 −1 1 0 −1
VA 43 PSNR VR ℂ = 19:3 dB
1 0
1 0
−1
−1 VA 48 PSNR VR ℂ = 20:9 dB
1 0
1 0
−1
−1 16
48
80
112
144
176
208
240
0
32
64
96
range bin
128
160
192
224
256
ADC sample GT Re
GT Im
pred Re
pred Im
Figure 11: CRP and ADC I/Q traces on held-out test frames. Rows: six ColoRadar scenes. Top block: CRP I/Q traces at the highest-energy virtual antenna per scene, range bins [15:256]. Bottom block: ADC I/Q traces at the same VA, across 256 fast-time samples. Solid lines are GT real (blue) and imaginary (red); dashed lines are 3DPS predictions. Per-panel: ∆φr / ∆φv (CRP relative-phase coherence ⟨cos(∆φ)⟩); envelope ρ|·| and ∆φt (ADC). The traces show that 3DPS recovers the dominant amplitude envelope while the per-bin absolute phase varies (per-(VA, range) absolute phase is unconstrained by the magnitude-only training loss; only the relative-phase content is captured).
24
I
Wall-clock and peak GPU memory
Table 7 reports per-scene wall-clock and peak GPU memory measured on a single NVIDIA RTX 4090 (24 GB) for 3DPS and the three optical-NVS baselines, on the same 6-scene ColoRadar benchmark and the same 9-frame split (8 training viewpoints, 1 held-out test frame). All measurements include forward+backward+optimizer+I/O. We also include a single-viewpoint Sionna RT Hoydis et al. [2023] reference figure as a representative MC ray-tracing cost anchor, measured at 8.2 min for 150 iterations on the same hardware and linearly extrapolated to ∼27.3 min at our 500-iteration budget; this is per single-viewpoint fit and would compound across the 8 training viewpoints in a multi-viewpoint NVS regime. Table 7: Wall-clock and peak GPU memory per scene on the 6-scene benchmark, single RTX 4090. Wall-clock is reported as mean across scenes with per-scene min–max range for measured methods. Sionna RT is reported as a single-viewpoint reference figure linearly extrapolated from a 150-iteration measurement to our 500-iteration budget. Method
Mean wall-clock (min) Per-scene range (min) Peak GPU mem (GB)
DART (cascaded) Huang et al. [2024] Radar Fields Borts et al. [2024] RadarSplat Kung et al. [2025] Sionna RT Hoydis et al. [2023]*
0.43 1.63 19.74 ∼27.3
0.4–0.5 1.6–1.7 14.2–21.4 —
not reported 2.4 14.2 —
3DPS (Ours)
3.17
2.3–3.7
—
*
Sionna RT figure is a single-viewpoint reference, linearly extrapolated from 8.2 min at 150 iterations on the same RTX 4090 to our 500-iteration budget. Reported as a representative MC ray-tracing cost anchor only; it is not a measured baseline on our 8-viewpoint benchmark.
J
RadarSplat hyperparameter selection
We benchmark RadarSplat Kung et al. [2025] using the public reference implementation. Hyperparameter selection mirrors the canonical configuration: we adopt the published per-Gaussian intensity head, the public per-scene Gaussian density schedule (initial densification at iteration 500, prune-andclone at 1000-iteration intervals), and the documented spherical-harmonic order. The Gaussian budget per scene is taken from the published default (90,000) and matches the order of magnitude implied by the public training scripts; we did not down-scale to match the 3DPS budget of N =20,000 oriented points because doing so would not be a fair use of RadarSplat’s representation. The training loss is the published L1+SSIM combination on RA magnitude. Adam learning rates and 30,000-iteration schedule match the public configuration. Per-scene wall-clock and test correlation are reported in Table 7 and Sec. F, respectively. We did not perform per-scene hyperparameter tuning on RadarSplat; the numbers reflect the public default applied uniformly across all six ColoRadar scenes.
K
Limitations and future work
The conclusion (Sec. 5) summarizes the main limitations briefly. This section expands on each and identifies the corresponding planned investigations. Single-bounce path tracing. The renderer evaluates a single TX→scatterer→RX path per oriented point. For static outdoor scenes dominated by direct surface returns (the regime of our ColoRadar benchmark) this is sufficient, and the headline 0.587 mean test |RA| Pearson reflects that. For indoor multipath-dominated environments where second- and third-order reflections carry comparable energy to direct returns Lu et al. [2020], multi-bounce path tracing in the style of MC ray tracers (e.g., Sionna RT Hoydis et al. [2023]) would remain necessary; integrating the closed-form BSDF + PSF splatting recipe into a multi-bounce framework is open. Magnitude-only loss leaves per-bin absolute phase unconstrained. 3DPS supervises on |RA| for fair comparison with the three magnitude-only baselines, leaving ∼ 22,000 per-(VA, range) absolute phase DOFs unconstrained by the loss — specifically, free up to per-VA RF calibration drift and per-range ADC sample-zero offsets. We side-step that nuisance phase at evaluation time by 25
reporting differential-phase metrics (∆φr , ∆φv , ∆φt in Table 2) that are absolute-phase-invariant by construction; a tighter absolute phase fit would require adding a complex (Re/Im) loss term during training. A natural future extension is a two-stage training recipe (first stage: magnitude loss to fit geometry and material; second stage: bounded position refinement under a complex loss to recover absolute phase) that locks in the magnitude performance and adds phase fidelity on top. LiDAR-dependent initialization. The point cloud that seeds 3DPS comes from a co-registered LiDAR scan. This is a strong prior that simplifies geometric initialization and accelerates convergence to a reasonable scene fit. Pure radar-driven NVS — removing the LiDAR dependence — is an important direction for radar-only deployments and an open challenge for our framework. Dynamic scenes and downstream perception. Extending range-Doppler-coherent rendering to scenes with moving targets, and validating NVS-augmented training data on radar-trained perception tasks (object detection, semantic segmentation), are natural next steps that the product-agnostic complex-output property of 3DPS enables but does not directly address.
L
Ablations
We ablate every design choice of 3DPS that is non-trivial to defend by appealing to the canonical 3D Gaussian Splatting literature: point count N , the multi-frame training-view budget, our adaptive density-control schedule, each of the four LiDAR-initialisation stages (azimuth-cone cull / occlusion ray-cast / cosine-weighted resample / farthest-point sampling), the quasi-static MIMO factorisation that fuses the per-(TX, RX) BSDF + range-splat into a single CUDA kernel, the Hann-PSF kernel half-width L, the carrier-phase detach used by the position gradient, and the L2 anchor strength λpos that ties the optimised positions to the LiDAR seed. All ablations train for the same 500 Adam iterations on the same 6 ColoRadar scenes used in the main paper, with all other hyper-parameters held at the canonical 3DPS recipe (Table 8, top row, bolded). Point count N . Test |RA| correlation is essentially saturated by N =10k (0.586) and peaks at the N =20k default (0.587); N =50k regresses to 0.538 because the adaptive split/prune schedule begins to displace already-converged points. The implication is that 20k is near-optimal rather than merely past the knee, and that 5k–10k remain viable choices for memory-constrained deployments. Number of training views. 3DPS is sample-efficient at the view axis: 2 / 4 / 6 / 8 (default) views land within 0.014 of one another on test correlation (0.573, 0.575, 0.586, 0.587). Wall-clock scales linearly (0.9, 1.4, 2.6, 3.0 minutes per scene) but quality does not. Sparse-view radar deployments can therefore drop to as few as two bracketing frames with negligible loss in the held-out view. Adaptive density control. Disabling the iter-100/200/300/400 split/prune schedule drops test correlation by 0.034 (0.553 vs. 0.587). The CRP and ADC envelope numbers stay close to default (0.587, 0.586), suggesting that the adaptive density mechanism contributes mostly to the magnitude-domain fit rather than to phase coherence. Validates the choice of importing the 3DGS-style densification into a radar setting. LiDAR initialization stages. Of the four init stages, three contribute: azimuth-cone cull (−0.031 when removed), cosine-weighted resample (−0.016), and farthest-point sampling (−0.045). The Mitsuba-based occlusion ray-cast is the surprising outlier — removing it actually improves test correlation by +0.013 (0.600 vs. 0.587) while saving ∼2 s of init time per scene. The likely cause is that adaptive density combined with the fisher-weighted split criterion already prunes occluded points more accurately than the explicit ray-cast, which over-rejects points that are partially visible across the chirp loop’s pose interpolation. We retain occlusion ray-casting in the canonical recipe for now to keep the LiDAR-prior pipeline aligned with the rendering literature, but flag it as removable in a future revision. Quasi-static MIMO factorisation. Disabling the fused CUDA Step-4 (BSDF) and Step-5 (rangesplat) kernels and falling through to the PyTorch path that materialises the full (M, NT X , NRX ) BSDF and (spread, M NT X NRX ) splat tensors raises per-scene wall-clock from 3.0 to 5.3 minutes (1.8×) and drops test correlation from 0.587 to 0.555. The factorisation is therefore both faster and numerically more stable in the backward pass, since the fused kernel uses a hand-derived gradient 26
Table 8: Ablation sweep on the 6-scene benchmark. Mean held-out test |RA| metrics, test |CRP| correlation, ADC envelope correlation, mean training |RA| correlation, and mean wall-clock per scene on a single RTX 4090. The 3DPS default (gray reference row) is N =20k points, 8 training views, adaptive density on, full 4-stage LiDAR init, MIMO-factored CUDA kernels, L=15, carrierphase detach on, λpos =100. Each subsequent row flips one design choice. Bold: best per column (ties jointly). Underline: second-best. Time is informational, not ranked. Test |RA|
Configuration
Complex / Env.
Train
Time
Corr ↑ PSNR ↑ SSIM ↑ RMSE ↓ |CRP| Corr ↑ ADC env. ↑ Corr ↑ (min) Reference 3DPS default
0.587
28.8
0.734
0.0369
0.603
0.613
0.812
3.0
Point count N N = 2,000 N = 5,000 N =10,000 N =50,000
0.578 0.571 0.586 0.538
28.4 28.0 28.2 27.3
0.696 0.691 0.704 0.696
0.0389 0.0421 0.0394 0.0439
0.586 0.574 0.591 0.565
0.571 0.571 0.615 0.602
0.673 0.731 0.777 0.858
4.0 4.0 2.7 5.1
Training views 2 train views 4 train views 6 train views
0.573 0.575 0.586
27.0 28.4 28.2
0.674 0.717 0.704
0.0448 0.0393 0.0393
0.591 0.581 0.584
0.591 0.599 0.605
0.888 0.859 0.838
0.9 1.4 2.6
Adaptive density no adaptive density
0.553
27.0
0.638
0.0459
0.587
0.586
0.760
3.3
LiDAR init stages no azimuth-cone cull no occlusion ray-cast no cosine resample no FPS (random sub-sample)
0.556 0.600 0.571 0.542
27.0 27.7 27.7 28.1
0.665 0.697 0.692 0.717
0.0458 0.0411 0.0417 0.0403
0.581 0.612 0.585 0.538
0.586 0.587 0.590 0.626
0.812 0.825 0.812 0.808
2.5 3.6 3.5 3.3
MIMO factorization no MIMO factorization (PyTorch fallback)
0.555
27.4
0.685
0.0432
0.576
0.587
0.813
5.3
PSF kernel L PSF L=5 PSF L=9 PSF L=21 PSF L=25
0.569 0.572 0.578 0.569
28.0 28.2 27.8 27.4
0.716 0.708 0.691 0.682
0.0402 0.0399 0.0413 0.0433
0.590 0.590 0.589 0.605
0.610 0.577 0.595 0.574
0.812 0.814 0.813 0.814
4.8 3.4 3.6 3.7
Carrier-phase detach no carrier-phase detach
0.537
27.1
0.676
0.0450
0.562
0.585
0.894
4.5
Position anchor λpos λpos =0 λpos =1 λpos =1000
0.584 0.584 0.566
28.1 27.7 28.2
0.725 0.710 0.707
0.0398 0.0419 0.0394
0.606 0.608 0.590
0.594 0.616 0.584
0.821 0.819 0.822
2.3 3.7 3.3
while the PyTorch fallback relies on autograd’s automatic differentiation through the expanded tensors. Hann-PSF kernel half-width L. The sweep {5, 9, 15, 21, 25} produces a clean unimodal curve with the peak at the L=15 default (0.587). Smaller (0.569, 0.572) and larger (0.578, 0.569) all underperform by 0.009–0.018. A small L underspreads the per-(TX, RX) range bins and aliases nearby scatterers; a large L smears the response across range, washing out fine-scale geometric detail. The unimodal shape supports our derivation that the Hann window’s discrete-time impulse response decays into the noise floor at |n|> ∼ 7 bins, making L=15 (±7 neighbours) the natural truncation. Carrier-phase detach. Allowing position gradients to flow through the carrier phase ϕcarrier = 2πf0 (dTX + dRX )/c rather than detaching it drops test correlation by 0.050 (0.537 vs. 0.587), the largest single ablation effect. The wavelength at 77 GHz is 3.9 mm, so a sub-mm position update produces a O(1) phase rotation: the loss landscape becomes locally periodic and Adam oscillates rather than converging. Detaching the carrier phase preserves only the amplitude-path gradient (geometric path length, antenna pattern, BSDF), which is smooth in position. This validates the design choice and explains why the fused CUDA Step-5 kernel is allowed to assume ϕcarrier is a constant w.r.t. position. L2 position anchor λpos . The sweep {0, 1, 100, 1000} shows that the L2 anchor is a soft regulariser: λpos = 0 and λpos = 1 both yield 0.584, essentially indistinguishable from the λpos =100 default 27
(0.587). Pushing to λpos = 1000 over-constrains positions toward the LiDAR seed and degrades fit to 0.566. The amplitude-path-only gradient combined with the small position learning rate (1e-5) already keeps the optimised geometry sub-millimetre from the seed; the explicit anchor is largely redundant in the canonical recipe but provides a guard against high-N regimes where individual points carry less data and can drift. Summary. Of the nine ablated axes, four are load-bearing for the |RA| result (∆ ≥ 0.03): adaptive density, FPS-stage init, MIMO factorisation, and carrier-phase detach. Three are mild contributors (0.01 ≤ ∆ < 0.03): cull-stage init, cosine-resample init, point count. Two are essentially neutral within MC noise: number of training views (provided ≥ 2) and λpos (anywhere in [0, 100]). The one negative result — occlusion-stage init at −0.013 — is a concrete actionable finding for the next revision of the system.
M
Per-point material visualizations
3DPS optimizes a 6-vector ITU-R P.2040 material per oriented point (ε′r , ε′′r , σh , ℓc , τ , d; Sec. D) jointly with the magnitude-only |RA| loss. Because the parameters are physically meaningful per point, the optimized scene is interpretable as a per-point material map rather than as opaque learned features. We visualize the material content three ways below: (i) a single “overall change” heatmap per scene quantifying per-point deviation from the ITU concrete prior, (ii) a side-by-side initial-vs-optimized comparison showing what the 8-viewpoint training actually moved, and (iii) a per-parameter breakdown across all six ITU axes for each scene. All renders use the canonical paper-results checkpoints, the inferno colormap on per-parameter physics ranges (linear for sigmoid params ε′r and τ ; log-scale for ε′′r , σh , ℓc , d; bounds in Sec. D), and the LiDAR scaffold mesh as faint grey context geometry. Per-scene viewpoints follow the same canonical scheme used for the qualitative supplement figures. S1-F185
S1-F438
S2-F105
S2-F160
S2-F300
Material Deviation
S0-F135
Figure 12: Per-point material deviation from ITU concrete defaults (single row, six ColoRadar scenes). Inferno colormap on the magnitude-of-deviation in physics-space (mean across the 6 ITU parameters; log-space for ε′′r , σh , ℓc , d and linear-space for ε′r , τ ). Brighter = larger deviation. Most points settle close to ITU concrete (the seed prior), with localized hot regions corresponding to vehicles, signage, and other strong specular returns where the optimizer pushed materials away from the prior to fit the radar response.
28
S1-F185
S1-F438
S2-F105
S2-F160
S2-F300
Optimized
Initial (ITU Default)
S0-F135
Figure 13: Initial vs. optimized per-point materials. Top row: every point initialized to ITU concrete (uniform purple = zero deviation). Bottom row: per-point materials after 500 Adam iterations on the 8-viewpoint |RA| loss. The contrast between rows visualizes which regions the optimizer chose to move away from the concrete prior — predominantly the boresight strip and locally on parked vehicles — consistent with the cosine-weighted resample (Sec. E) concentrating points where the antenna pattern weights them most. 0
00
r
h
c
d
Overall
S2-F300
S2-F160
S2-F105
S1-F438
S1-F185
S0-F135
r
Figure 14: Per-parameter optimized materials, all six ITU-R P.2040 axes. Rows: six ColoRadar scenes. Columns: ε′r , ε′′r , σh , ℓc , τ , d, plus an overall-deviation column on the right (matching Fig. 12). Inferno colormap with per-parameter normalization to the ITU-R P.2040 valid bounds (Sec. D); linear scale for sigmoid params (ε′r , τ ) and log scale for the others. Brighter = higher value. The per-parameter columns confirm that the optimizer makes physically interpretable choices — e.g., elevated ε′r and σh on vehicle bodies, near-default values on extended ground returns — rather than treating the 6-vector as an opaque embedding.
29
M.1
Per-point surface normals
3DPS also optimizes a per-point quaternion that orients the local surface frame; the surface normal is the rotation of +ẑ by that quaternion (convention from Sec. E, init seeded from LiDAR). The quaternion is trainable so the optimizer can refine the LiDAR-derived normals against the magnitudeonly |RA| loss; the figures below show what that refinement does. We use the same 6-scene benchmark and viewpoint convention as the material visualizations (Sec. M). S1-F185
S1-F438
S2-F105
S2-F160
S2-F300
Normal Deviation
S0-F135
Figure 15: Per-point normal deviation from the LiDAR seed (single row, six ColoRadar scenes). Inferno colormap on per-point angular deviation in degrees, clipped to [0◦ , 30◦ ] (dark = LiDARaligned, bright yellow = ≥ 30◦ rotation). Initial-state normals come from nearest-vertex lookup into the LiDAR mesh at each optimized point’s position; sub-mm position drift makes this a faithful reconstruction of the iteration-0 seed. Most points stay within ∼10◦ of the seed (consistent with LiDAR normal noise), with localized hot regions where the radar response disagrees with the LiDARderived normal — e.g., on smooth specular surfaces (vehicle bodies, signage) where the LiDAR’s local-plane fit averages over surface fine-structure that the radar return is sensitive to. S1-F185
S1-F438
S2-F105
S2-F160
S2-F300
Optimized
Initial (LiDAR seed)
S0-F135
Figure 16: Initial vs. optimized per-point normals, RGB-encoded. World-space normal-map encoding ((n+1)/2 → RGB) with back-facing normals camera-flip-corrected for visual consistency. Top row: initial state (LiDAR mesh nearest-neighbour at each optimized point’s position). Bottom row: learned state (per-point quaternion → normal after 500 Adam iterations). The visual contrast between rows tracks Fig. 15: the optimizer adds high-frequency normal variation to the locally-flat LiDAR seed, predominantly on the boresight strip where the antenna pattern weights points most heavily. S1-F185
S1-F438
S2-F105
S2-F160
S2-F300
Learned Normals (RGB)
S0-F135
Figure 17: Learned per-point normals only, RGB-encoded as in Fig. 16 (single row, six scenes). Provided alongside Fig. 16 for direct visual reference of the optimized normal field at full per-scene scale.
30
N
Broader impact and dual-use considerations
3DPS targets sparse-view novel view synthesis for mmWave radar in static outdoor scenes. The intended applications are radar simulation, sensor-driven scene understanding, and data augmentation for downstream perception tasks such as object detection and free-space estimation in autonomous driving and robotics. By making radar data more consistent across viewpoints and reducing the data-collection cost of new deployments, 3DPS lowers the barrier to safety-critical evaluation of radar perception systems and supports sim-to-real transfer for robust autonomy. Dual-use considerations. mmWave radar is also used in surveillance and military sensing systems. The methods presented here apply, in principle, to these settings, but the inputs we require (a coregistered LiDAR scan, full radar array geometry, and dense pose annotations) are not generally available in covert deployments, and our contribution does not advance any covert sensing capability beyond what is already achievable with conventional radar simulation tools. We do not provide trained models or reconstructions of any sensitive site, and the released code is intended for open scientific use on public datasets such as ColoRadar Kramer et al. [2022]. Failure modes and deployment caveats. 3DPS is not a substitute for direct radar measurement. The single-bounce assumption fails in indoor multipath-dominated environments, and the LiDARconditioned scaffold means failure of the scaffold (poor calibration, missing scan coverage) propagates to the renderer output. Users deploying 3DPS in safety-critical settings should treat its output as supplementary to, not replacing, real radar acquisitions. We list further limitations in Sec. K. Energy and compute footprint. A complete training run for one scene takes approximately 3 minutes on a single NVIDIA RTX 4090 (Sec. I). The full set of experiments reported in this paper consumed less than 100 GPU hours of compute.
31