Conceptio › Archive › arXiv CS
arXiv CSopen access

PILOT: One Physics-Integrated Generation Framework to Unify 2D and 3D Radio Map Construction

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
distributed-systemsinternetnetworkingprotocols
networking, internet, protocols, distributed systems

PILOT: One Physics-Integrated Generation Framework to Unify 2D and 3D Radio Map Construction Weiming Huang∗ , Hao Sun† , Junting Chen∗ ∗ School of Science and Engineering (SSE) and Shenzhen Future Network of Intelligence Institute (FNii-Shenzhen)

The Chinese University of Hong Kong, Shenzhen, Guangdong 518172, China

arXiv:2604.23533v1 [eess.SP] 26 Apr 2026

† Department of Electrical Engineering, City University of Hong Kong, Hong Kong

Abstract—Unified 2D and 3D radio map construction supports network planning, wireless digital twins, and unmanned aerial vehicle (UAV) applications. In urban environments, blockage, reflection, and diffraction make accurate construction expensive for physics-based solvers. Autoregressive next-token prediction offers a single sequential formulation that can cover both 2D and 3D generation, but standard raster ordering ignores the spatial structure of radio propagation. When generation follows propagation, each token is predicted from propagation-relevant history rather than spatially arbitrary context, which provides more causally informative conditioning and lowers conditional uncertainty. We propose PILOT, a pretrained autoregressive framework that replaces raster scan with a wavefront sequence expanding outward from the transmitter. Each prediction step is guided by an environment-aware instruction that spatially aligns environment features with the queried radio map region. The same framework extends to 3D radio maps through height-slice stacking while a gradient loss enforces vertical continuity. On standard 2D benchmarks, PILOT achieves the lowest NMSE among all baselines. For volumetric generation, it reduces NMSE by 78% relative to the diffusion baseline at roughly 2500× faster inference. It also outperforms methods that rely on 10% sparse measurements and achieves the best zero-shot results in the crossdomain evaluation. Code: https://github.com/Uminan/PILOT Index Terms—radio map construction, pathloss estimation, autoregressive generation, wireless communications

I. I NTRODUCTION As sixth-generation wireless networks extend coverage from the ground plane to aerial and multi-floor users, channel characterization needs to cover 3D regions rather than streetlevel slices [1], [2]. Wireless digital twins, cellular-connected unmanned aerial vehicle (UAV) networking, and volumetric coverage analysis therefore require radio maps in both 2D and 3D over the same scene [2], [3]. Dense measurements over such regions are infeasible: ground surveys are laborintensive, and upper-floor interiors and aerial corridors cannot be sampled as densely [2]. By contrast, environment maps encoding building geometry are available before deployment from urban geographic databases and digital twin assets, and extend from planar to volumetric scenes [4]. In dense urban scenes, pathloss is shaped by blockage, reflection, and diffraction around irregular structures, so local pathloss depends on geometry well beyond the line-of-sight path. Classical ray tracing models these interactions, but its cost grows rapidly

with resolution and multi-bounce order, becoming impractical for 3D volumes [4]. In this paper, we focus on the construction of pathloss radio maps from scene geometry and transmitter configuration, using one formulation for both 2D slices and 3D volumes over the same urban scene. The map characterizes pathloss across ground-level locations and low-altitude airspace, with the 2D slice and the 3D field coinciding at shared coordinates. Wireless digital twins rely on such a map for network planning, where one volumetric field must cover street-level corridors and upper-floor interiors within a single urban block. Cellularconnected unmanned aerial vehicle (UAV) networking relies on it for aerial base-station placement and trajectory design, which depend on altitude-dependent pathloss rather than a single ground-level slice. These applications demand efficient re-evaluation when the carrier frequency or link-budget setting changes, together with source-aware prediction in which the pathloss field over the scene is structured by signal propagation from the transmitter. For this geometry-to-pathloss construction setting, prior work mainly follows physics-based simulation and learningbased prediction. Classical ray tracing is the representative physics-based approach, modeling reflection, diffraction, and scattering over a detailed 3D environment with encoded geometry and material properties [5]. Its fidelity, however, incurs a computational cost that makes repeated 3D evaluation impractical [4]. Learning-based regressors reduce this cost by predicting 2D radio maps from scene geometry and transmitter location in a single forward pass. RadioUNet employs a cascaded two-UNet architecture that maps citymap geometry and transmitter location to the 2D radio map [6]. RadioMamba retains this 2D regression formulation and replaces the convolutional neural network (CNN) backbone with a hybrid Mamba-UNet to capture long-range spatial dependencies and to improve the efficiency-accuracy tradeoff [7]. Other methods instead reconstruct radio maps from sparse on-site measurements, with RME-GAN using a twophase conditional generative adversarial framework and deep completion autoencoders inferring the field directly from measured samples [8], [9]. Across these learning-based pipelines, construction remains posed as direct regression or reconstruction rather than as a source-aware process. Most methods

are restricted to 2D prediction, and measurement-conditioned variants further depend on on-site samples that are difficult to acquire densely over the target region. Among existing generative approaches, conditional diffusion is the representative framework that formulates radio-map construction as iterative denoising rather than direct regression or measurement-conditioned reconstruction. RadioDiff encodes the radio map into a variational autoencoder (VAE) latent space and applies a 2D UNet for denoising conditioned on environment geometry and transmitter location. RadioDiff3D extends the backbone to a 3D UNet for volumetric generation. In both formulations, however, the denoising trajectory follows a fixed noise schedule rather than a source-aware signal-propagation process. Iterative denoising further grows costly as spatial resolution and sampling steps increase, which conflicts with real-time construction requirements. This paper develops a physics-guided pretrained autoregressive framework for source-aware radio map construction. Radio propagation guides the generation process, so each token is predicted along propagation paths from the transmitter rather than in raster order. Most existing methods use separate architectures for different scenarios, each trained from scratch; an autoregressive decoder consolidates these into one tokenized backbone that scales from 2D to 3D and across carrier frequencies without redesign. By drawing on spatial priors from large-scale visual pretraining, limited wireless supervision adapts the decoder to propagation-induced effects, enabling zero-shot cross-domain transfer. Fig. 1 illustrates the progression from next-token prediction in text and vision models to the proposed physics-integrated latent ordered transformer (PILOT) framework. Recent LLM-based work has introduced sequence modeling into radio-map construction, yet a propagation-aligned mechanism is still not embedded in the generation process. LLM4PG [10] uses an LLM backbone as a single-pass regressor without an explicit generation order. Ripple [11] formulates radio map construction as incontext learning with a frozen autoregressive vision model and a geometric spiral-out order, but relies on same-layout demonstrations as prompts and ignores blockage, making it unsuitable for the sampling-free cross-scene setting. In visual autoregressive models, each token depends on spatially adjacent context. In a radio field, however, the informative context for a target location traces the propagation path from the transmitter and may bypass neighboring regions entirely. This propagation dependency raises three questions. How to guide the generation order so that the model follows the physical propagation of radio signals? How to represent the radio field so that it preserves both horizontal boundary fidelity and vertical continuity under real-time constraints? How to keep prediction consistent when carrier frequency and linkbudget settings shift the coarse pathloss scale? We tackle these challenges through a propagation-guided autoregressive framework built on a pretrained visual backbone. To resolve the generation-order mismatch, PILOT adopts a wavefront propagation order that prioritizes nearby lowblockage regions before shadowed areas. This order reflects

"Define AI" Prompt

Next-Token Prediction

"AI is a field..." Answer

Llama

Vision Decoder

"A cute cat" Raster Order

Prompt LlamaGen

Visual Image

Physics-Guided Mechanism Propagation Order

Environment Map PILOT

2D/3D Radio Map

Fig. 1. Overview of the PILOT framework.

blockage-aware propagation costs accumulated along paths from the transmitter to each target region, using the building environment map. At each step, the radio token is predicted with guidance from its spatially corresponding environment token, and spatial coordinates are embedded via 3D rotary position embedding (3D-RoPE) to ensure positional registration and semantic alignment. By adjusting the height coordinate in 3D-RoPE, we further harness the framework to generate at the target height for various height-layer applications. To address field structure complexity, a tokenizer compresses the radio field into a discrete token space to accommodate both planar and volumetric generation. Specifically, height information of the 3D map is encoded along the channel dimension to avoid volumetric processing, assisted by a 3D gradient regularizer that preserves boundary fidelity and vertical continuity. To compensate for pathloss scale variation, PILOT inputs a pathloss anchor map based on free-space path loss (FSPL) with shadowing that encodes carrier-frequency and pathloss-range information to calibrate the coarse field scale. One architecture therefore supports 2D and 3D construction and enables zeroshot transfer across carrier frequencies and scene geometries without retraining. To validate the role of generation order in training convergence and inference accuracy, we construct a true-PL order that ranks tokens from low to high ground-truth pathloss. This order serves as an oracle upper bound at inference, yet training on it alone overfits. Mixing physics-guided orders improves generalization over single-order training and further outperforms random permutations and geometric orders such as raster and Z-curve orders. Among all candidate orders, the wavefront order yields the lowest predictive entropy, with entropy increasing monotonically over successive generation steps. This confirms the design principle of wavefront order: prioritizing low-uncertainty tokens first limits error accumulation. On the standard 2D benchmark, PILOT reduces normalized mean square error (NMSE) by 9.5% relative to RadioMamba [7]. Under zero-shot transfer across carrier fre-

quency, pathloss range, and scene geometry, PILOT reduces NMSE by 7.9% relative to the closest zero-shot baseline, RME-GAN [8]. For volumetric generation, PILOT reduces NMSE by 78% at roughly 2500× faster inference than the diffusion baseline [12] and outperforms a variant with 10% sparse measurements. The main contributions are: • We propose PILOT, a pretrained autoregressive framework that unifies 2D and 3D radio map construction as environment-aware next-token prediction. To align radio and environment tokens in a shared spatial frame, physical coordinates are encoded into the autoregressive sequence through 3D-RoPE. • We develop a wavefront propagation order that mirrors signal propagation from the transmitter along blockageaware paths to guide the autoregressive generation sequence, so that each prediction step conditions on propagation-relevant context. • We design a shared 2D tokenizer for both tasks by encoding height information in the channel dimension, thereby avoiding volumetric processing, while a 3D gradient regularizer preserves boundary fidelity and vertical continuity. The pathloss anchor map further calibrates the coarse field scale, thereby enabling zero-shot transfer across carrier frequencies and link-budget settings. • On three ray-traced benchmarks, PILOT reduces NMSE by 9.5% in standard 2D construction and by 7.9% under zero-shot cross-domain transfer, relative to the respective state-of-the-art baselines. For 3D generation, it reduces NMSE by 78% over the diffusion baseline while achieving approximately 2500× faster inference in a samplingfree setting. The rest of this paper is organized as follows. Section II presents the system model and problem formulation. Section III details the proposed PILOT framework. Section IV reports the experimental results and analysis. Section V concludes the paper. The core variables are summarized in Table I.

TABLE I N OTATION TABLE

Symbol E H Manc R, R̂ e r, r̂ Nz Fθ pθ π P u β(·) I(·) D N f ∆P |C| H(·) L λ, α

Description Environment map Building height map Pathloss anchor map Ground-truth and predicted radio maps Environment token Ground-truth and predicted radio tokens Number of height-slice channels Autoregressive NN with parameters θ Autoregressive conditional distribution Generation-order permutation Set of spatial patches 3D spatial coordinate Blockage ratio along a ray Indicator function Accumulated propagation cost 8-connected neighborhood of a patch Carrier frequency Pathloss range Codebook size Predictive entropy function Block-diagonal direct sum Hyperparameters

zrx and Nz , where zrx specifies the center of the receiverheight range and Nz determines the number of predicted height levels. B. Physics-Aware Input The environment input is defined as E = [H, Mtx , Manc ],

(2)

We consider sampling-free radio map construction in dense urban environments with static geometry and a single transmitter. Let E denote the environment-side input consisting of the building map and transmitter configuration. Let R, R̂ ∈ RS×S×Nz denote the ground-truth and predicted pathloss fields on an S × S spatial grid. The field-level prediction task is R̂ = Fθ (E, zrx , Nz ), (1)

where H is the building height map, Mtx is the transmitter position map, and Manc is the pathloss anchor map. Prior methods are typically developed for a single domain with fixed carrier frequency and link-budget setting, so the coarse pathloss scale is absorbed by the training data distribution and need not be encoded explicitly in the input. Under cross-domain shifts, the free-space attenuation varies with carrier frequency f , while the valid pathloss range is further determined by the link-budget setting of the target domain. The third channel Manc is therefore introduced to provide a frequency-aware anchor for the coarse pathloss scale. The pathloss anchor map is defined as

where Fθ denotes a neural network parameterized by θ. The model learns θ by minimizing the prediction loss L(R̂, R). This formulation unifies 2D and 3D radio map prediction in one field representation. When Nz = 1, the output is a 2D radio map at receiver height zrx . When Nz > 1, the output is a 3D radio map over the receiver-height range determined by

Manc (u; f, Lthr ) = LFSPL (∥u − utx ∥, f ) + Lshd (u; f, Lthr ), (3) where LFSPL is the free-space pathloss and the shadow correction is   Lshd (u; f, Lthr ) = β(utx , u) LFSPL (d0 , f ) − Lthr . (4)

II. S YSTEM M ODEL AND P ROBLEM F ORMULATION A. Scenario and Task Definition

explicit next-token interface for modeling pθ (R | E), which is the formulation adopted in this work. To match the autoregressive generation paradigm, the continuous radio map is split into patches and represented by N discrete radio tokens r = [r1 , . . . , rN ]. The environment input is encoded into an aligned sequence e = [e1 , . . . , eN ] with the same spatial indexing as r. In the proposed pipeline, the vector-quantization (VQ) stage first establishes this discrete target space, so that 2D and 3D radio map construction share a common token interface for autoregressive prediction. Given r and e, the conditional distribution factorizes as

Global Initialize Local Update

pθ (r | e) = Fig. 2. Blockage ratio computation. K sample points are placed along the ground-plane projection from utx to uj ; blue segments indicate blocked portions where zk < H(xk , yk ).

Here Lthr = 10 log10 (W N0 ) + N F − (PTx )dB is a domaindependent threshold determined by the link-budget setting of the target domain, where W is the bandwidth, N F is the noise figure, and PTx is the transmit power; and d0 is a near-field reference distance. The correction Lshd uses the blockage ratio β to scale the shadow range [LFSPL (d0 , f ) − Lthr ], so that Manc encodes the frequency-aware coarse pathloss scale that H and Mtx cannot convey. The blockage ratio is defined as β(utx , u) =

K 1 X

K

I(zk < H(xk , yk )) ,

(5)

k=1

where (xk , yk , zk ) is the k-th sample along the ray from utx to u, H(xk , yk ) is the building height at that sample, K is determined by the ray length and grid resolution, and I(·) is the indicator function. The quantity β(utx , u) measures the blocked fraction of the direct path and scales the shadow correction Lshd , as illustrated in Fig. 2. As a result, the anchor map Manc varies with the blockage level along the direct transmitter-to-target path. C. Order-Aware Sequence Formulation Prior work [13] formulates radio map construction as continuous field prediction conditioned on E. The relevant dependencies in this field are long-range and propagationdependent, since blockage, reflection, and diffraction couple distant regions of the map rather than purely local neighborhoods. The task is reformulated as a sequential generation problem, in which the conditional distribution pθ (R | E)

(6)

is realized by sequential decoding. A single-pass discriminative predictor does not provide such an explicit generation process. A diffusion model is generative, but its trajectory follows a denoising schedule rather than a spatial generation order over the radio map. Decoder-only autoregressive next-token prediction provides an

N Y

pθ (rn | e, r1 , . . . , rn−1 ),

(7)

n=1

which defines the prefix conditioning structure under the default ordering n = 1, . . . , N but leaves the mapping from token index to spatial position unspecified. This mapping is not fixed by decoder-only autoregressive generation itself. The work in [14] treats the generation order as a design choice without changing the decoder. Let π ∈ SN be a permutation of the N token indices, so that π(n) is the index predicted at step n. The order-aware factorization is pθ (r | e; π) =

N Y

 pθ rπ(n) | e, rπ(1) , . . . , rπ(n−1) ,

(8)

n=1

where e provides global context at every step and π determines which radio-token prefix is available when rπ(n) is predicted. The available prefix shapes the conditional uncertainty at each step. For a radio field, earlier tokens are informative only when they are propagation-relevant to the current target, so that the prefix encodes the blockage and detour structure along the propagation path. Raster or arbitrary orders may expose prefixes that are nearby in space yet irrelevant to the propagation state of the current target, thereby increasing the conditional uncertainty in shadowed regions. The remaining design question is how to construct π so that generation follows the propagation process from the transmitter. In particular, propagation-relevant regions should be predicted before the shadowed regions that depend on them. III. PILOT N ETWORK A RCHITECTURE In this section, we develop PILOT, an environment-aware autoregressive architecture for radio map construction. A vanilla decoder-only autoregressive predicts each token from a raster prefix that is often misaligned with the propagation state of the current target, and offers no built-in way for the environment E to shape either the generation order or the per-step prediction. To align the order-aware factorization in (8) with radio propagation, we need to address the following technical challenges: • How to convert the continuous radio field R into discrete tokens that retain sharp shadow boundaries and preserve vertical continuity for 3D maps formed by channelstacked height slices, so that 2D and 3D maps share a single token space?

How to construct the generation order π from the environment E so that the prefix rπ(<n) at each step exposes propagation-relevant context for predicting rπ(n) , rather than spatially nearby but propagation-irrelevant tokens? • How to inject the environment E into the decoder so that, under the order π, each radio-token prediction is conditioned on environment evidence at the same spatial coordinate, in both 2D and 3D? To address these challenges, PILOT combines a blockageaware graph relaxation that constructs π from E, an autoregressive decoder conditioned on environment evidence at every step, and a tokenizer that yields a discrete space shared between 2D and 3D maps. The overall architecture is shown in Fig. 3.

entries remain unused. PILOT introduces two design changes. First, the convolutional code projector is replaced with a shared linear projector, which couples all code vectors through the same trainable mapping and lets them evolve jointly during training. Second, a differentiable vector quantization scheme is adopted to restore gradient flow through the discrete bottleneck without an auxiliary commitment loss. These changes mitigate codebook collapse, as examined in Section IV-F. For 3D inputs, channel-stacked height slices avoid volumetric convolution and preserve a shared codebook with 2D maps, but they can weaken the vertical smoothness of the reconstruction. To compensate, the VQ-stage loss augments the reconstruction term with a multi-scale gradient regularizer,

A. Overall Pipeline

where the gradient term is X   Lgrad3D = E ∥∇xy R̂(s) − ∇xy R(s) ∥

•

PILOT decomposes radio map construction into a tokenizer and an autoregressive generator. The tokenizer is trained first and then frozen, after which the autoregressive generator is trained to predict indices into the resulting fixed codebook. The tokenizer maps the continuous radio map R ∈ R256×256×Nz to a 16×16 grid of N = 256 discrete tokens drawn from a shared codebook, and reconstructs R from these tokens at inference. Stacking Nz height slices along the channel dimension allows the same 2D tokenizer to handle both 2D and 3D maps and produce N = 256 radio tokens in either case. The wavefront ordering module takes the environment E and outputs a permutation π that ranks the N patches by accumulated propagation cost from the transmitter. The cost combines Euclidean distance with a blockage penalty derived from H and utx , so π visits propagation-relevant patches before the shadowed patches that depend on them. At inference, this order is fixed. During training, it is sampled together with other candidate orders. The autoregressive generator encodes E into a sequence of environment tokens through a DINOv3-initialized vision transformer (ViT), interleaves them with the radio tokens under the order π, and predicts the next radio token at each step. A shared 3D rotary position embedding registers each environment-radio token pair to its spatial coordinate, so a single decoder handles both 2D and 3D maps. The frozen tokenizer then decodes the predicted token sequence to R̂. B. Radio Map Tokenizer

Lvq = R̂ − R + λgrad Lgrad3D ,

The tokenizer compresses the radio map R ∈ R into a 16 × 16 latent grid through a 2D convolutional encoder, and reconstructs R̂ from the latent grid through a 2D convolutional decoder. Each spatial position of the latent grid is then quantized to one entry of a shared codebook, yielding N = 256 discrete tokens that form the target space for the autoregressive stage and that are common to 2D and 3D inputs. A 512-dimensional latent space is used to cover the wide dynamic range of pathloss. We adopt the LlamaGen [15] codebook size, |C| = 16384. However, raising the latent dimension under standard VQVAE training causes severe codebook collapse, in which most

(10)

s∈S

  + λz E ∥∇z R̂ − ∇z R∥ . Here, ∇xy and ∇z denote finite in-plane and vertical differences, S is the set of downsampling scales used in multi-scale evaluation, λgrad controls the weight of the gradient term relative to reconstruction, and λz sets the relative importance of vertical versus in-plane gradients. For 2D maps with Nz = 1, the vertical term is dropped. C. Wavefront Propagation Order Raster or fixed spatial orders do not reflect the physical nature of radio propagation. The prefix available when predicting a shadowed patch may then contain spatially nearby tokens that are weakly related to its propagation state, while omitting the lower-cost regions that physically shape it. PILOT therefore adopts a wavefront propagation order, denoted by πwavefront , which expands outward from the transmitter along blockage-aware low-cost paths. 1) Order Construction: The wavefront order is defined by ranking patches according to their accumulated blockageaware propagation cost from the transmitter, so that the sequence expands outward along lower-cost paths rather than following a fixed spatial scan. At the initial step t = 0, each (0) patch i receives a cost Di that combines Euclidean distance with a direct-path blockage penalty, (0)

256×256×Nz

(9)

Di

=

utx − ui 2 αLoS , 1 − β(utx , ui )

(11)

where β(·, ·) ∈ [0, 1] denotes the blockage between two locations computed from H, and αLoS controls how strongly directpath blockage inflates the cost. A patch blocked along the direct path may still be reachable at lower total cost through a less obstructed intermediate region. PILOT therefore performs Dijkstra-style relaxation over the 8-connected neighborhood N (i) of each settled patch i, updating each neighbor j ∈ N (i) from step t to t + 1 as ) ( ui − uj 2 (t+1) (t) (t) αNLoS , (12) Dj = min Dj , Di + 1 − β(ui , uj )

Loss Codebook

Quantize

(ⅰ) Scene Geometry

Radio Token

B

... (16 × 16 × 512)

Tx

61×↑

(256 × 256 × Nz)

61×↓

VQ Encoder

Radio Feature

VQ Decoder

2D/3D Radio Map

B

B

(16 × 16 × 1)

(256 × 256 × 1) (ⅱ) Propagation Cost

(a) Radio Map Tokenizer Environment Map Pathloss Anchor Map Tx Position Height Map

ViT Encoder

Projector

cls

...

1.4 0.6

∞

4.8

1.0

∞

∞

0

1.2 0.8 1.6 3.2 2.4 2.0 2.8 3.6

Propagation-Guided Permutation (ⅲ) Sort by Increasing

Shared 3D-RoPE Predict

cls

(256 × 256 × 3) cls

Global Token

... Autoregressive Transformers

Environment Token ...

Null Token

#6

#2 #15 #13

#4

#1 #14 #16

#5

#3

#9

#8 #10 #12

#7 #11

Radio Token

(b) Autoregressive Generator

(c) Wavefront Ordering

Fig. 3. Overall framework of the proposed PILOT.

where αNLoS governs the blockage sensitivity of local hops. The full procedure is summarized in Algorithm 1, and example propagation paths are shown in Fig. 4(a). Let Di denote the final relaxed cost of patch i at convergence. The deployed wavefront order is then defined as πwavefront = argsorti∈P Di ,

(13)

where P = {1, . . . , N } indexes the N patches and πwavefront (n) returns the index of the patch with the n-th smallest relaxed cost. The relaxation runs in O(N log N ) on the 8-connected patch graph with N = 256, so the order is computed once per scene at negligible cost relative to inference. 2) Information-Entropy Analysis: The wavefront order πwavefront is further analyzed through an information-theoretic lens that links the choice of order to the predictive uncertainty realized by a finite-capacity autoregressive model. The analysis characterizes the prefix property induced by πwavefront . Under the joint distribution p(r | e) defined in Section II, the chain rule of entropy gives N X

  H rπ(n) | e, rπ(<n) = H r | e ,

(14)

n=1

for every permutation π, so the aggregate conditional entropy is invariant to ordering. Reordering cannot reduce the total uncertainty when the model matches the true distribution. This invariance does not constrain the predictive entropy of a finite-capacity model with pθ ̸= p. The predictive entropy

Algorithm 1 Wavefront Propagation Order Input: Patch index set P = {1, . . . , N }; transmitter position utx ; blockage function β(·, ·) derived from the height map H; exponents αLoS , αNLoS . Output: Wavefront propagation order πwavefront . 1: for each i ∈ P do  αLoS 2: Di ← ∥utx − ui ∥2 1 − β(utx , ui ) 3: end for 4: V ← P 5: while V ̸= ∅ do 6: i ← arg mink∈V Dk 7: V ← V \ {i} 8: for each j ∈ N (i) ∩ V do αNLoS 9: wij ← ∥ui − uj ∥2 1 − β(ui , uj ) 10: Dj ← min{Dj , Di + wij } 11: end for 12: end while 13: πwavefront ← argsorti∈P Di 14: return πwavefront

of pθ at step n under order π is Hπθ (n) = −

|C| X

  pθ c | e, rπ(<n) log pθ c | e, rπ(<n) ,

(15)

c=1

which equals the entropy of the codebook-index softmax produced by the decoder at step n. Different orders π induce different prefixes rπ(<n) , hence different per-step enP tropies Hπθ (n) and different averages H̄ θ (π) = N1 n Hπθ (n). Whether π helps the model depends on whether the prefix

Tx Low Cost Zone High Cost Zone Obstacles Path (Near Tx) Path (Far Tx)

Wavefront Context Raster Context Target Patch Tx

(a)

(b)

Fig. 4. (a) Wavefront propagation paths in an urban scene, routing around blocked regions through lower-cost detours. (b) Visible prefix comparison for the same target patch (red): the raster-order prefix (green) consists of spatially preceding patches largely irrelevant to the target propagation state, whereas the wavefront-order prefix (yellow) concentrates along the propagation route to the target.

surfaces the propagation-relevant predecessors of patch π(n), a property observed empirically in autoregressive language and image generation. For the wavefront order, the predecessor containment property can be stated explicitly. For each patch i, let P ∗ (i) = {j0 , j1 , . . . , jK−1 } denote any shortest-cost predecessor chain induced by the final cost graph from utx to i. Along this chain, the recorded costs satisfy Dj0 < Dj1 < · · · < DjK−1 < Di , since every relaxation step adds a positive edge weight. When πwavefront (n) = i, sorting patches by ascending D to form πwavefront guarantees  P ∗ (i) ⊆ πwavefront (1), . . . , πwavefront (n − 1) . (16) The wavefront prefix at the step that predicts patch i thus contains an entire shortest-cost predecessor chain to i. Under raster ordering, (16) does not hold in general, since the predecessors on the propagation path to a shadowed patch may lie outside the raster prefix block. By ranking patches according to D, the wavefront prefix exposes a lower-cost surrogate propagation route that is more aligned with blockage-aware reachability than a raster prefix, and this difference is expected to manifest as a lower realized H̄ θ (πwavefront ) on patches that depend on long propagation paths. Fig. 4(b) illustrates the prefix difference on a representative target patch: under raster order the visible prefix is dominated by spatially adjacent patches with no propagation link to the target, while under wavefront order the visible prefix concentrates along the lower-cost propagation route to the target. Section IV-E2 evaluates this prediction through inference-time measurements of the realized predictive entropy. D. Environment-Aware Autoregressive Generation Wavefront ordering specifies which patch is predicted next, but not what environment evidence should be exposed at that step. Using the full environment sequence as a one-shot prefix would introduce many regions unrelated to the current propagation state, dilute attention, and increase decoding

cost. PILOT instead uses one global CLS token and stepaligned environment tokens. The CLS token provides scenelevel context, and the patch-aligned token eπ(n) supplies the environment cue for each target rπ(n) . 1) Environment-Guided Context: The environment input E ∈ R256×256×3 is encoded by a ViT initialized from DINOv3 into 1 + N = 257 tokens of dimension 768, comprising one global CLS token and N = 256 patch-aligned environment tokens {ei }. An multilayer perceptron (MLP) projector then maps each token to the decoder hidden dimension of 1024. The ViT and the projector are fine-tuned end-to-end with the decoder. The CLS token is prepended once, and the N environment tokens are interleaved with the radio tokens under the order π to form the input sequence   CLS, eπ(1) , rπ(1) , eπ(2) , rπ(2) , . . . , eπ(N ) , rπ(N ) , (17) with causal masking applied throughout. At step n, the visible prefix consists of the CLS token, the previously decoded environment-radio pairs, and the current environment token eπ(n) . Environment tokens enter the prefix incrementally under the wavefront order, rather than as a fixed scene-wide prefix. The corresponding π-conditioned likelihood factorizes as N  Y pθ rπ(n) CLS, eπ(1) , rπ(1) , pθ r | e; π =

(18)

n=1



. . . , eπ(n−1) , rπ(n−1) , eπ(n) . 2) 3D Spatial Registration: The ViT token carries environment semantics for the queried patch, and 3D-RoPE assigns each matched pair (ei , ri ) to the same 3D index (xi , yi , zi ). The 1D rotary position embedding is defined as    d/2  M  cos(mϕj ) − sin(mϕj )  q, (19) RoPE q, m =  sin(mϕj ) cos(mϕj ) j=1

L

where denotes the block-diagonal direct sum, d is the perhead rotary dimension, m is the 1D position index, and ϕj = 10000−2(j−1)/d is the j-th rotary frequency. For 3D position indexing, each per-head query vector is partitioned into three axis-specific subvectors as  (x)  q q = q(y)  , (20) q(z) where the per-head rotary dimension is split approximately evenly across the three axes. The 3D rotary embedding then applies the 1D rotation to each subvector with its corresponding axis index,   RoPE q(x) , x  3D-RoPE q, x, y, z =  RoPE q(y) , y  . (21) (z) RoPE q , z The same coordinate rule applies in 2D and 3D, with zi instantiated as the fixed receiver height in the 2D case and as the per-slice height across the Nz channels in the 3D case. The decoder weights are shared across both settings.

E. Training and Inference During teacher forcing, the cross-entropy loss is evaluated only at the steps whose target is a radio token. For each training sample, the order is sampled uniformly from three candidates: πwavefront from (13); πpriorPL , which ranks patches by ascending value of the anchor map Manc derived from frequency-aware free-space and blockage information; and πtruePL , which ranks patches by ascending value of the groundtruth radio map R and serves as an oracle reference. At inference, πwavefront is computed from E via Algorithm 1 and held fixed throughout decoding. The decoder generates the N = 256 radio tokens autoregressively by greedy decoding, and the frozen tokenizer reconstructs R̂ ∈ R256×256×Nz from the predicted token grid. The number of autoregressive steps is fixed at 256 for both 2D and 3D settings. IV. E VALUATION A. Experimental Setup We evaluate the proposed method on three tasks. Standard 2D radio map construction measures in-domain reconstruction accuracy on held-out samples. Zero-shot cross-domain transfer evaluates robustness to simultaneous shifts in carrier frequency, transmitter geometry, and pathloss range without retraining. 3D radio map generation evaluates volumetric pathloss prediction across receiver heights. Table II summarizes the datasets. RadioMapSeer [6] contains 56,000 ray-traced 256 × 256 pathloss maps at 5.9 GHz, with both transmitter and receiver fixed at 1.5 m and building height fixed at 25 m. RadioMap3DSeer [6] uses the same horizontal resolution but operates at 3.5 GHz, places the transmitter 3 m above the rooftop, and has a narrower pathloss range, making it the target domain for zero-shot transfer. UrbanRadio3D [12] provides a 256 × 256 × 20 volumetric benchmark at 5.9 GHz with 1 m spatial resolution, realistic building heights from 6.6 to 19.8 m, and receiver heights from 1 to 20 m. For pathloss prediction, we use its 2.84 million pathloss maps. We compare with five baselines from three paradigms: • RadioUNet [6]: a cascaded two-UNet architecture that predicts the radio map in a single forward pass from environment-conditioned inputs and serves as the standard discriminative baseline for 2D construction. • RadioMamba [7]: a hybrid Mamba-U-Net model that combines convolutional feature extraction with linearcomplexity global context modeling and serves as a stronger single-pass discriminative baseline. • RME-GAN [8]: a conditional adversarial reconstruction framework originally designed for sparse-measurementconditioned radio map estimation. In this work, it is evaluated as an adapted variant without sparse measurements so that its conditioning setting matches the common 2D baseline setting. • RadioDiff [13]: a conditional diffusion model for sampling-free 2D radio map construction that performs iterative denoising in a VAE latent space.

RadioDiff-3D [12]: a 3D diffusion baseline with volumetric convolutional operators for pathloss generation on UrbanRadio3D. RadioMapSeer is split into training, validation, and test sets of 500, 100, and 100 maps, respectively. For the 3D task, we use the 1–4 m subset of UrbanRadio3D, which is divided at the file level into 90% training files and 10% test files, to match the evaluation setting of RadioDiff-3D. For the 2D tasks, all models receive the building height map, transmitter position, and pathloss anchor map as inputs. Since the original RME-GAN conditions on sparse measurements, we remove its sparse-measurement branch and evaluate an adapted version under the same 2D input setting. For the 3D task, the pathloss anchor map is withheld, and each model receives only the volumetric environment representation and transmitter encoding. Pathloss maps are min-max normalized to [0, 1] over [−47, −169] dB for cross-dataset comparability. NMSE, structural similarity index measure (SSIM), and peak signal-to-noise ratio (PSNR) are computed in the normalized domain, while root mean squared error (RMSE) is reported in dB. All reproducible baselines are retrained with AdamW, a learning rate of 10−4 , batch size 32, and mixed-precision bf16 on two NVIDIA RTX 4090 GPUs. Inference is measured on a single RTX 4090. •

TABLE II DATASET SPECIFICATIONS Parameter RadioMapSeer [6] Task type Standard 2D estimation Dataset size 56k Map size 256 × 256 1m Pixel length Center carrier 5.9 GHz Building height 25 m Tx height 1.5 m 1.5 m Rx height Pathloss range −47 to −147 dB ∗ Reflects pathloss maps exclusively.

RadioMap3DSeer [6] Zero-shot generalization 56k 256 × 256 1m 3.5 GHz 6.6–19.8 m 3 m above rooftop 1.5 m −75 to −111 dB

UrbanRadio3D [12] 3D estimation 2.8M∗ 256 × 256 × 20 1m 5.9 GHz 6.6–19.8 m 1.5 m 1.0–20.0 m −92 to −169 dB

B. Standard 2D Comparison On RadioMapSeer, PILOT obtains the lowest NMSE as reported in Table III. Relative to RadioMamba, NMSE drops from 0.0349 to 0.0316. Qualitatively, PILOT better resolves shadow boundaries and diffraction corridors as shown in Fig. 5(a). TABLE III 2D RADIO MAP CONSTRUCTION ON R ADIO M AP S EER Model RME-GAN [8] RadioUnet [6] RadioDiff [13] RadioMamba [7] PILOT (Ours)

NMSE ↓ 0.1911 0.1041 0.1111 0.0349 0.0316

RMSE (dB) ↓ 7.730 5.636 5.785 3.286 3.076

SSIM ↑ 0.8707 0.8991 0.9059 0.9322 0.9462

PSNR ↑ 22.40 25.25 24.99 29.93 30.58

Infer Time (s) ↓ 0.0025 0.0024 0.2960 0.0393 0.0252

C. Zero-Shot Cross-Domain Transfer Table V evaluates zero-shot transfer from RadioMapSeer to RadioMap3DSeer under joint frequency, pathloss-range and geometry shift. PILOT reduces NMSE by approximately 7.9 % over the strongest zero-shot baseline RME-GAN, while

-75

-111

Pathloss (dB)

-47

-147

-75

-111

Pathloss (dB)

-47

-147

RME-GAN

RadioUnet

RadioDiff RadioMamba (a) 2D Prediction Visualization

PILOT (Ours)

Ground Truth

-75

-111

Pathloss (dB)

-47

-147

-75

-111

Pathloss (dB)

-47

-147

RME-GAN

RadioUnet

RadioDiff RadioMamba PILOT (Ours) (b) Zero-Shot Transfer Visualization

Ground Truth

Fig. 5. Predicted radio maps. (a) 2D construction on RadioMapSeer. (b) Zero-shot transfer to RadioMap3DSeer under frequency and geometry shift.

TABLE IV D ISTRIBUTIONAL STATISTICS OF PIXEL - SPACE AND VQ- SPACE REPRESENTATIONS ACROSS R ADIO M AP S EER AND R ADIO M AP 3DS EER Representation Pixel Space VQ Space

Norm. Entropy ↑

Gini Coeff. ↓

Seer

3DSeer

Seer

3DSeer

0.658 0.968

0.519 0.973

0.859 0.390

0.941 0.365

DJS ↓

ρ↑

0.705 0.030

0.577 0.824

RadioMamba degrades under domain shift. Fig. 5(b) shows that PILOT and RME-GAN output pathloss ranges that match the target domain, whereas the other baselines remain locked to the training range. The shared pathloss anchor map anchors the pathloss range across frequencies, and the VQ latent space reduces the Jensen–Shannon divergence DJS from 0.705 in pixel space to 0.030, as shown in Fig. 6 and Table IV. D. 3D Radio Map Generation Table VI and Fig. 7 report 3D results on UrbanRadio3D at receiver heights 1–4 m. PILOT in volumetric mode generates multi-layer radio map tokens jointly and reduces NMSE from 0.0534 to 0.0120 relative to the 1000-step RadioDiff-3D

TABLE V Z ERO - SHOT GENERALIZATION COMPARISON ON R ADIO M AP 3DS EER

Model RME-GAN [8] RadioUnet [6] RadioDiff [13] RadioMamba [7] PILOT (Ours)

NMSE ↓ 0.3115 0.5844 0.5081 0.6188 0.2869

RMSE (dB) ↓ 3.873 5.116 4.871 5.277 3.582

Pixel Space DJS = 0.705

RadioMapSeer

SSIM ↑ 0.7035 0.6699 0.6327 0.6255 0.7451

PSNR ↑ 19.43 17.18 17.50 16.92 20.24

VQ Latent Space DJS = 0.030

RadioMap3DSeer

Fig. 6. t-SNE of RadioMapSeer and RadioMap3DSeer distributions in pixel space and VQ latent space.

diffusion baseline, at 0.049 s per map versus 121.76 s. PILOT in height-conditioned mode produces a single-height map by setting the z coordinate of the 3D-RoPE embedding. The 200step RadioDiff-3D variant with 10 % sparse measurements still falls behind both PILOT modes on NMSE and SSIM. Fig. 9(c) shows that the 3D gradient loss preserves inter-slice continuity across the four height levels, with the 90th-percentile vertical error reduced by 50 %.

PILOT-3D

Pathloss (dB)

-92

-169

H1

H2

H3

H4

Ground Truth

Pathloss (dB)

-92

-169

H1

H2

H3

H4

Fig. 7. 3D radio maps on UrbanRadio3D at receiver heights 1–4 m.

TABLE VI 3D RADIO MAP CONSTRUCTION ON U RBAN R ADIO 3D Model RadioDiff-3D [12] (T = 200) RadioDiff-3D (T = 1000) RadioDiff-3D (T = 200 + 10%Samp.) PILOT volumetric PILOT height-cond.

NMSE ↓

RMSE (dB) ↓

SSIM ↑

PSNR ↑

Infer Time (s) ↓

0.3472

10.20

0.6453

19.57

24.3526

0.0534

5.03

0.8309

24.00

121.7630

0.0550

3.70

0.8187

29.23

24.3526

0.0120 0.0071

2.60 2.00

0.9290 0.9498

29.92 32.49

0.0486 0.1010

E. Propagation Mechanism Analysis Section IV-B established end-task accuracy. This subsection investigates why propagation-aligned ordering improves generation by analyzing training convergence, comparing sequence variants at inference, and measuring predictive entropy per generation step. 1) Generation Order Comparison: Training stage. Fig. 8 plots validation cross-entropy for eight individual orders and three hybrid strategies. Among the geometric orders, raster scanning reaches the lowest loss at 5.12, while Hilbert, Zcurve, subsample, and alternative patterns settle between 6.19 and 7.51. The physics-guided sequences, wavefront and priorPL, converge to 5.89 and 5.95. The true-PL order stalls at 12.87 because the ground-truth pathloss ranking varies abruptly across samples and provides no consistent spatial structure for the network to exploit. The physical hybrid strategy reaches 3.33, well below the geometric hybrid at 4.53 and random ordering at 4.54. Its constituent orders preserve physical structure through blockageweighted costs and pathloss rankings, so each permutation supplies spatial context correlated with the propagation state. Geometric mixtures, by contrast, merely rearrange scan patterns without physical alignment. Mixing a bounded set of physics-guided orders also provides the sequence diversity

needed for regularization while avoiding the intractable N ! permutation space, where model capacity would be diluted across uninformative arbitrary sequences. Inference stage. Tables VII and VIII evaluate a single physical hybrid checkpoint across eight generation orders. On the in-domain RadioMapSeer benchmark in Table VII, all nonoracle orders achieve NMSE between 0.0313 and 0.0326, indicating that inference-time ordering has minor effect when the test distribution matches the training data. The true-PL oracle reaches 0.0255. Although true-PL diverges during single-order training, the physical hybrid checkpoint encodes the correct mapping from propagation structure to tokens, and the oracle order then provides the sequence of lowest uncertainty at test time. Under zero-shot transfer to RadioMap3DSeer in Table VIII, the gap between sequences widens. The five geometric orders [14] cluster near NMSE = 0.2900 because scene-agnostic scan patterns cannot adapt to unseen layouts and frequencies. Both wavefront and true-PL achieve lower errors at 0.2869 and 0.2871. The blockage-weighted Dijkstra cost thus serves as an accurate, training-free surrogate for the oracle order on unseen spatial distributions. TABLE VII A BLATION OF GENERATION ORDERS ON R ADIO M AP S EER

Order Hilbert Z-curve Prior PL Subsample Raster Alternative Wavefront True PL

NMSE ↓ 0.0326 0.0316 0.0315 0.0315 0.0314 0.0313 0.0316 0.0255

RMSE (dB) ↓ 3.136 3.091 3.089 3.088 3.085 3.079 3.076 2.799

SSIM ↑ 0.9448 0.9455 0.9455 0.9455 0.9456 0.9457 0.9462 0.9492

PSNR ↑ 30.42 30.54 30.54 30.55 30.56 30.57 30.58 31.39

TABLE VIII A BLATION OF GENERATION ORDERS ON R ADIO M AP 3DS EER

Order Subsample Raster Hilbert Alternative Z-curve Prior PL Wavefront True PL

NMSE ↓ 0.2899 0.2899 0.2900 0.2900 0.2899 0.2878 0.2869 0.2871

RMSE (dB) ↓ 3.601 3.601 3.601 3.601 3.600 3.586 3.582 3.582

SSIM ↑ 0.7447 0.7447 0.7448 0.7448 0.7448 0.7451 0.7451 0.7452

PSNR ↑ 20.19 20.19 20.19 20.19 20.19 20.23 20.24 20.24

2) Conditional Entropy Verification: We measure predictive entropy at each autoregressive step to quantify how the generation sequence affects model confidence. At step n, the model outputs a logit vector zn ∈ R|C| , and the corresponding entropy is

H(n) = −

|C| X c=1

softmax(zn )c log softmax(zn )c .

(22)

Hilbert-Curve Z-Curve

Subsample Alternative

Raster

Wavefront

Random

Validation Loss

12

Validation Loss

Validation Loss

Prior PL

10

8 6 1

2

3

4

5

6

Training Epochs

Geometric Hybrid

Physical Hybrid

10

10

0

True PL

14

7

8

8 6 0

9

8 6 4

1

2

3

4

5

6

Training Epochs

(a)

7

8

9

0

1

2

3

4

5

6

Training Epochs

(b)

7

8

9

(c)

Fig. 8. Validation cross-entropy by generation order. (a) Geometric orders. (b) Physics-guided orders. (c) Hybrid strategies. Z-curve Subsample

Hilbert Alternative

4

Wavefront True PL

3 2

Prior PL Raster

Z-curve Subsample

Hilbert Alternative

0

50

100

150

Step n

200

250

0.8

1

0.6

0.95

grad3D

4

3

0.4

0.9

50% ↓

0.85

0.2

2 1

grad2D

1

CDF

Prior PL Raster

̄ Entropy H(n)

̄ Entropy H(n)

Wavefront True PL

0.8 0

0

50

100

(a)

150

Step n

200

250

(b)

0

0.2

1

0.4

2

0.6

3

0.8

4

Vertical Gradient Error (dB)

(c)

Fig. 9. Predictive entropy and vertical gradient error measurements. (a) Sample-averaged predictive entropy H̄(n) over 256 steps on RadioMapSeer. (b) H̄(n) on RadioMap3DSeer under zero-shot transfer. (c) CDF of vertical gradient error on UrbanRadio3D.

The physical hybrid checkpoint generates 256 tokens under each evaluated order. We denote the sample-averaged predictive entropy at step n as H̄(n). Fig. 9(a) shows H̄(n) on RadioMapSeer. The true-PL and prior-PL sequences follow concave profiles that peak near the midpoint, where complex multipath effects concentrate. The wavefront order increases monotonically from H̄(0) = 1.2 to H̄(255) = 4.2, confirming that generation begins with low-uncertainty near-field patches and defers complex shadow regions to later steps. The five geometric orders oscillate without a clear trend and share an identical mean entropy of H̄ = 2.87; rearranging a purely spatial scan pattern does not alter the aggregate conditional uncertainty. The wavefront order yields H̄ = 2.72, a 5.3 % reduction in mean predictive uncertainty. Fig. 9(b) repeats this measurement on RadioMap3DSeer under zero-shot transfer. The three curve profiles persist across the domain shift: the wavefront order averages H̄ = 2.96, compared with 3.01 for the geometric cluster. To localize this reduction geographically, we define ∆H = Hraster − Hwavefront per patch and evaluate three spatial regimes in Fig. 10. In the edge transmitter scenario, signals route around multiple obstacles before reaching deep non-lineof-sight regions, producing µ∆H = 0.731 and σ 2 = 0.890;

the positive ∆H values concentrate directly behind buildings. In the urban canyon layout, the mean remains stable at µ∆H = 0.707 but the variance rises to 2.101, with the uncertainty reduction localizing at canyon intersections and transition boundaries while straight corridors show negligible difference. In the sparse obstacle environment, the metrics drop to µ∆H = 0.001 and σ 2 = 0.346, confirming that the wavefront constraint introduces no detrimental bias when the geometry is structurally simple. These results corroborate the formulation in Section II: propagation-aligned ordering reduces predictive entropy because each token conditions on the correct physical history, and the magnitude of this reduction scales with local propagation complexity. F. Ablation Studies 1) Training Strategy: Without pretraining, the autoregressive model reaches only NMSE 0.0828 on RadioMapSeer. Fine-tuning from ImageNet-pretrained LlamaGen-L [15] reduces this to 0.0316, a 62 % drop. LoRA [16] stops at 0.0559 because its low-rank subspace cannot bridge the full domain gap from natural images to radio maps, and full weight updates are needed as confirmed in Table IX. Fig. 9(c) shows that the

8

Edge Tx

8

0

-8

8

8

Urban Canyon

0

Sparse Obstacles

0

0

-8

8

8

0

0

Radio Map

-8

Fig. 10. Spatial entropy maps for three propagation regimes.

3D gradient loss preserves inter-slice continuity, with the 90thpercentile vertical error reduced by 50 %. TABLE IX A BLATION OF TRAINING STRATEGIES ON R ADIO M AP S EER Strategy Training from Scratch LoRA [16] Full Fine-tuning

NMSE ↓ 0.0828 0.0559 0.0316

RMSE (dB) ↓ 4.852 4.123 3.076

SSIM ↑ 0.9063 0.9195 0.9462

PSNR ↑ 26.57 28.02 30.58

2) Tokenizer Design: The VQ-VAE tokenizer determines the upper bound of autoregressive generation, because details lost during tokenization cannot be recovered later. Table X quantifies the contribution of each component to the tokenizer performance. The baseline VQ-VAE from LlamaGen [15] suffers from severe codebook collapse, with only 0.25 % utilization and poor reconstruction. Replacing the convolutional projector with the linear projector of DiVeQ [17] raises utilization to 99.7 % and reduces NMSE from 0.1507 to 0.0605. Differentiable quantization [17] further reduces NMSE to 0.0253 by restoring gradient flow through the discrete bottleneck. Adding the multi-scale gradient loss [18] further improves boundary fidelity and achieves the best results, with NMSE 0.0189, SSIM 0.9626, and PSNR 33.46 dB. TABLE X A BLATION OF RECONSTRUCTION COMPONENTS IN THE VQ-VAE TOKENIZER ON R ADIO M AP S EER Method

Codebook Util. ↑

NMSE ↓

RMSE (dB) ↓

SSIM ↑

PSNR ↑

Baseline [15] + Linear Projector [19] + Diff. Quantization [17] + Gradient Loss [18]

0.25% 99.7% 99.6% 99.7%

0.1507 0.0605 0.0253 0.0189

5.661 3.264 2.682 2.305

0.8365 0.9213 0.9534 0.9626

25.15 30.11 32.10 33.46

V. C ONCLUSION This paper presented PILOT, a physics-integrated autoregressive framework for unified 2D and 3D radio map construction. By coupling propagation-guided generation with a

frequency-aware pathloss anchor map, PILOT achieves strong performance in standard 2D radio map estimation, enables zero-shot cross-domain transfer without sparse measurements, and extends effectively to 3D radio map generation. On the evaluated benchmarks, PILOT achieved up to 2500× faster inference than the diffusion baseline while outperforming a 3D baseline supported by 10% measurements. These results show that physics-guided autoregressive generation can produce accurate radio maps without sparse measurements, at inference speeds suitable for real-time planning. R EFERENCES [1] Z. Zhang, Y. Liu, C.-X. Wang, H. Chang, J. Bian, and J. Zhang, “Machine learning based clustering and modeling for 6g uav-to-ground communication channels,” IEEE Transactions on Vehicular Technology, vol. 73, no. 10, pp. 14 113–14 126, 2024. [2] C. Li, Z. Dou, and Y. Lin, “Fast 3-d radio map reconstruction via cross tensor approximation,” IEEE Internet of Things Journal, vol. 11, no. 24, pp. 40 619–40 633, 2024. [3] H. Wang, J. Zhang, G. Nie, L. Yu, Z. Yuan, T. Li, J. Wang, and G. Liu, “Digital twin channel for 6g: Concepts, architectures and potential applications,” IEEE Communications Magazine, vol. 63, no. 3, pp. 24– 30, 2025. [4] J.-H. Lee and A. F. Molisch, “A scalable and generalizable pathloss map prediction,” IEEE Transactions on Wireless Communications, vol. 23, no. 11, pp. 17 793–17 806, 2024. [5] Z. Wu, D. Wu, S. Fu, Y. Qiu, and Y. Zeng, “Ckmimagenet: A dataset for ai-based channel knowledge map toward environment-aware communication and sensing,” IEEE Transactions on Communications, vol. 73, no. 12, pp. 14 430–14 443, 2025. [6] R. Levie, Ç. Yapar, G. Kutyniok, and G. Caire, “Radiounet: Fast radio map estimation with convolutional neural networks,” IEEE Transactions on Wireless Communications, vol. 20, no. 6, pp. 4001–4015, 2021. [7] H. Jia, N. Cheng, X. Wang, C. Zhou, R. Sun, and X. Shen, “Radiomamba: Breaking the accuracy-efficiency trade-off in radio map construction via a hybrid mamba-unet,” IEEE Transactions on Network Science and Engineering, vol. 13, pp. 2454–2468, 2026. [8] S. Zhang, A. Wijesinghe, and Z. Ding, “Rme-gan: A learning framework for radio map estimation based on conditional generative adversarial network,” IEEE Internet of Things Journal, vol. 10, no. 20, pp. 18 016– 18 027, 2023. [9] Y. Teganya and D. Romero, “Deep completion autoencoders for radio map estimation,” IEEE Transactions on Wireless Communications, vol. 21, no. 3, pp. 1710–1724, 2022. [10] M. Sun, L. Bai, X. Cheng, and J. Wu, “Llm4pg: Adapting large language model for pathloss map generation via synesthesia of machines,” arXiv preprint arXiv:2511.02423, 2025. [11] Y. Peng and J. Xu, “In-context radio map estimation via ripple autoregressive modeling,” in NeurIPS 2025 Workshop: AI and ML for NextGeneration Wireless Communications and Networking. [12] X. Wang, Q. Zhang, N. Cheng, J. Chen, Z. Zhang, Z. Li, S. Cui, and X. Shen, “Radiodiff-3d: A 3d× 3d radio map dataset and generative diffusion based benchmark for 6g environment-aware communication,” IEEE Transactions on Network Science and Engineering, vol. 13, pp. 3773–3789, 2026. [13] X. Wang, K. Tao, N. Cheng, Z. Yin, Z. Li, Y. Zhang, and X. Shen, “Radiodiff: An effective generative diffusion model for sampling-free dynamic radio map construction,” IEEE Transactions on Cognitive Communications and Networking, vol. 11, no. 2, pp. 738–750, 2025. [14] Z. Pang, T. Zhang, F. Luan, Y. Man, H. Tan, K. Zhang, W. T. Freeman, and Y.-X. Wang, “Randar: Decoder-only autoregressive visual generation in random orders,” in Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR), 2025, pp. 45–55. [15] P. Sun, Y. Jiang, S. Chen, S. Zhang, B. Peng, P. Luo, and Z. Yuan, “Autoregressive model beats diffusion: Llama for scalable image generation,” ArXiv, vol. abs/2406.06525, 2024. [16] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “Lora: Low-rank adaptation of large language models,” in International Conference on Learning Representations, ICLR, 2022.

[17] M. H. Vali, T. Bäckström, and A. Solin, “Diveq: Differentiable vector quantization using the reparameterization trick,” in International Conference on Learning Representations, ICLR, 2026. [18] M. Mathieu, C. Couprie, and Y. LeCun, “Deep multi-scale video prediction beyond mean square error,” in International Conference on Learning Representations (ICLR), 2016. [19] Y. Zhu, B. Li, Y. Xin, Z. Xia, and L. Xu, “Addressing representation collapse in vector quantized models with one linear layer,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025, pp. 22 968–22 977.

Record · ID 138879 · SHA-256 2e0dab27c7732f3a
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.