Conceptio › Archive › arXiv CS
arXiv CSopen access

Task-Oriented Communications for Edge-Assisted Multi-View Localization

Zhengru Fang et al. · arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
distributed-systemsinternetnetworkingprotocols
networking, internet, protocols, distributed systems

1

Task-Oriented Communications for Edge-Assisted Multi-View Localization

arXiv:2609.35173v1 [cs.NI] 28 Sep 2026

Zhengru Fang† , Huanhuan Lou† , Senkang Hu, Yihang Tao, Zongdian Li, Member, IEEE, Yiqin Deng, Member, IEEE, Jingjing Wang, Senior Member, IEEE, and Yuguang Fang, Fellow, IEEE

Abstract—Unmanned aerial vehicles (UAVs) and unmanned ground vehicles (UGVs) tend to lose satellite positioning in urban canyons, indoor facilities, and jammed or spoofed environments, and vision-based matching with a geo-tagged database remains one of the few sources for absolute positioning. However, onboard computation and energy budgets usually cannot host that database and its matching pipeline, so localization is often offloaded to edge servers or roadside units over wireless links whose throughput varies along the route. Such edge-assisted multi-view localization must therefore jointly decide when to offload, which views and semantic rate to transmit, and which client to serve under changing wireless and edge resources. In this paper, we present a network-adaptive task-oriented communication framework that combines scalable orthogonalityregularized variational information bottleneck (O-VIB) encoding, value-of-information (VOI)-guided request control, and VOI-weighted Lyapunov scheduling. We design an O-VIB model that spans prefix lengths from 8 to 128 latent dimensions through nested importance-ordered prefixes and covers the available view subsets through view masks. Edge assistance is requested only when the predicted reduction in localization risk exceeds the communication and service cost. We have demonstrated that under a matched per-route traffic budget, VOI-guided control can lower the mean route error by 24.8% and the p95 route error by 31.0% relative to budgeted periodic offloading on CARLA multi-view UAV data. In indoor UAV and UGV realworld experiments, the deployed pipeline can lower the mean position error by 28.0% and 14.4% relative to uncompressed all-view CLIP retrieval while cutting descriptor traffic by 98.6% and 98.2%, at 0.145 KB per request. Under high congestion, value-aware shaping can lower the edge-side p95 latency for the top-10% high-value requests by 76.2%, from 137.7 ms to 32.8 ms. Index Terms—Task-oriented communication, information bottleneck (IB), value of information, visual localization, unmanned aerial vehicles (UAVs).

† Z. Fang and H. Lou contributed equally to this work (co-first authors).

Z. Fang is with the Department of Electronic and Computer Engineering, The Hong Kong University of Science and Technology, Hong Kong (e-mail: [email protected]). H. Lou, S. Hu, Y. Tao, and Y. Fang are with the Hong Kong JC STEM Lab of Smart City and the Department of Computer Science, City University of Hong Kong, Hong Kong (e-mail: {huanhulou2, senkanghu2-c, yihangtao2c}@my.cityu.edu.hk; [email protected]). Z. Li is with Zhejiang University, Hangzhou 310027, China (e-mail: [email protected]). Y. Deng is with the School of Data Science, Lingnan University, Tuen Mun, Hong Kong, China (e-mail: [email protected]). J. Wang is with the School of Cyber Science and Technology, Beihang University, Beijing, China (e-mail: [email protected]). A preliminary version of this work was presented in part at the IEEE Global Communications Conference (GLOBECOM), 2025 [1].

I. Introduction Unmanned aerial vehicles (UAVs) and unmanned ground vehicles (UGVs) are frequently deployed in indoor warehouses, large industrial facilities, campuses, urban canyons, and outdoor inspection routes, where satellite positioning is often weak, intermittent, or actively spoofed. Under these severe conditions, vision-based localization may have become the only remaining source for absolute positioning, since camera views carry geometric and appearance cues that complement radio, inertial, and marker-based positioning. Recent visual-inertial simultaneous localization and mapping (SLAM), digital-twin positioning, and visual place recognition systems confirm this capability in environments where satellite positioning is unreliable [2]–[7]. Unlabeled adaptation of vision models on movable agents offers a complementary way to handle changing observations [8], while multicamera configurations retain a usable view under occlusion, repetitive layouts, or rapid viewpoint changes. Running this pipeline onboard is nevertheless expensive, because the geotagged database and matching stage often exceed the computation and energy budget on a small platform. Such platforms therefore offload localization inference to an edge server or roadside unit, where task-oriented communication can reduce the transmitted visual payload [9]–[11]. Take the scenario in Fig. 1 as an example. When GNSS is unavailable or spoofed, local motion estimates drift from the true trajectory and navigation risk grows. The proposed system therefore keeps the fast local loop on the mobile platform, while an edge request carries only a selected viewrate semantic representation. The edge decodes that representation, searches the geo-tagged database, estimates the pose, and returns a correction. Because the uplink and the edge server are shared among multiple clients, this separation ties semantic payload selection to request timing and to multiclient service ordering. This pipeline therefore creates four coupled challenges. First, what to transmit: multi-view visual embeddings may contain substantial cross-view redundancy, while localization depends mainly on location-discriminative and task-relevant features. Second, how much to transmit: wireless throughput and endpoint energy vary over time, so a fixed high-rate representation becomes infeasible under bandwidth bottlenecks, whereas a fixed low-rate representation underutilizes available bandwidth when channel conditions improve. Third, when to transmit: local odometry can maintain short-term pose estimates, and an edge request is valuable mainly when

2

(a) GPS-Denied Localization Failure Reference Trajectory Estimated Trajectory

GNSS Unavailable

Spo of Sig ing na l

GNSS Disruption Collision Risk

Wrong Localization

(b) Edge-Assisted Multi-View Localization Front

Multi-View Observations �� Back

Left

Right

Edge Server Down

Semantic Decoder

Retrieval Database

Pose Prediction

Task-Oriented (�) Semantic Features �� UAV-Edge Links

Accurate Localization

Corrected Pose (ŷ� , �� )

Fig. 1. Motivating scenario and edge-assisted multi-view localization pipeline. Panel (a) illustrates localization drift and collision risk when GNSS is unavailable or unreliable. Panel (b) shows the proposed pipeline: the mobile client observes multiple views, selects a task-oriented semantic representation, transmits it over the client-edge link, and receives a corrected pose after edge decoding and database-assisted localization.

uncertainty growth justifies the cost. Fourth, which request to serve: many mobile clients may contend for the same uplink and edge GPU, and first-come-first-served or rate-only scheduling serves them without considering the localization error of each request. These decisions are tightly coupled, since changing the request time also changes the useful view subset, the required semantic rate, and the urgency of edge service. Existing approaches address only fragments of this problem. Conventional image and video codecs such as JPEG, H.264, H.265, and WebP optimize perceptual fidelity rather than localization utility, so aggressive compression may discard task-critical geometric cues [12]–[15]. In parallel, recent task-oriented and information-bottleneck based semantic communication systems reduce task-irrelevant payloads, yet many of them train a single-rate encoder or assume a periodic upload model, without adapting view mode, semantic rate, and request timing together [10], [11], [16]–[18]. A third line of work on information-aware status updating and edge scheduling highlights the importance of freshness, correctness, and service ordering, but it usually assumes fixed-size state packets or leaves scheduling decoupled from semantic compression [19]–[25]. To the best of our knowledge, no existing study combines state-triggered localization requests, view- and rate-adaptive semantic encoding, and value-aware

scheduling under shared wireless and edge-computing resources. Motivated by these observations, we propose a unified task-oriented communication framework for edge-assisted multi-view localization under GPS-denied environments. An orthogonality-regularized variational information bottleneck (O-VIB) encoder addresses what and how much to transmit, and the task-level value of information (VOI) addresses when to transmit and which request to serve. Both components serve the same goal of reducing the localization error under limited communication and edge resources. The main contributions are summarized as follows: • We develop a scalable multi-view O-VIB encoder. Through nested rate dropout and view masks, one model supports latent prefixes from 8 to 128 dimensions and all four UAV view modes. A variational rate term suppresses task-irrelevant latent coordinates, and an orthogonality penalty keeps the posterior-mean projection wellconditioned. At 8 KB/s, encoding and sending a 32dimensional O-VIB code takes 50.4 ms on a Jetson Orin NX, whereas the fastest image codec takes 2.92 s. • We design a VOI-guided online request-and-mode selection policy. A client requests edge assistance only when the predicted reduction in localization risk is worth the costs for the communication, delay requirement, and energy consumption, and it selects the view-rate action with the largest net VOI. Under a matched traffic budget in CARLA, this policy can lower the mean and p95 route errors by 24.8% and 31.0% relative to budgeted periodic offloading. • We formulate multi-client edge scheduling as a VOIweighted drift-plus-penalty problem. The queue-stability and 𝑂 (1/𝑉) utility-gap result is conditional on all assumptions of Theorem 2. We separately evaluate a shaped scheduler with waiting-age and deadline terms, which can lower the edge-side p95 latency of the top10% high-value requests by 76.2% under high congestion, from 137.7 ms to 32.8 ms. • We open-source our code and a multi-view UAV dataset of 357,690 five-view frames from eight CARLA towns, with RGB, depth, semantic segmentation, and sixdegree-of-freedom pose labels. In GPS-denied indoor experiments with a UAV and a UGV under motion-capture ground truth, our pipeline can lower the mean position error by 28.0% and 14.4% relative to uncompressed allview CLIP retrieval while cutting the descriptor traffic by 98.6% and 98.2%. II. Related Work A. Task-Oriented and Semantic Communication Task-oriented, or semantic, communication transmits only the information needed by the receiver’s task. Shao et al. [26] provide an information-theoretic framework, while Zhang et al. [27] survey open problems. Information-bottleneck features have been used for edge video analytics, federated learning, and timely modeling [16]–[18], [28], [29]. Yuan et al. [10] and Furutanpey et al. [11] study scalable

3

or lightweight split encoders, while CASVA [30] and ILCAS [31] adapt video configurations to network conditions. Nested dropout and Matryoshka representation learning train one embedding whose leading dimensions form shorter representations [32], [33]. These works do not jointly adapt view mode, semantic rate, and request time; our VOI controller makes these decisions online under localization loss. B. Edge-Assisted Visual Localization When satellite signals are unavailable, visual localization provides another source for absolute positioning. Collaborative visual-inertial SLAM, digital twins, and coded pavement references support multi-agent, indoor, and road positioning [2]–[4]. Visual place recognition and learned matching improve retrieval under viewpoint and appearance changes [5]–[7], [34]–[36]. Recent collaborative-perception work also studies adaptive mutual-view information against adversarial agents and feed-forward reconstruction from uncalibrated driving views [37], [38]. These methods target perception security or scene completion, whereas our problem is task-oriented localization over a constrained link. Cao et al. [9] offload collaborative visual SLAM to reduce onboard cost, but edge-assisted localization generally assumes complete observations and periodic fixed-rate uploads. We instead select views, latent rate, and request time from predicted localization value. C. Information-Aware Status Updating and VOI Status-update systems decide whether a new update is worth sending. Age of information and age of incorrect information have guided transmission design for massive access, UAV, and platooning networks [19]–[22], [39]. Wang et al. [40] use deep reinforcement learning to decide when to upload for edge video analytics. Online robust control has likewise addressed system and channel uncertainty in lowaltitude UAV swarms [41]. These works mainly update fixedsize state or control information. Our VOI estimate instead decides whether to request edge assistance, which views and rate to send, and how the edge orders requests according to localization loss. D. Multi-Client Edge Scheduling for Inference A shared edge server decides which request to serve first. Prior work schedules video frames, cached models, and concurrent DNN jobs [23]–[25], [42]–[47]. Surveys of mobile edge intelligence and distributed learning for autonomous swarms further identify placement, communication overhead, and unreliable links as common deployment constraints [48], [49]; lightweight federated learning combines pruning, quantization, and power control under delay and energy budgets [50]. These schedulers and learning systems do not rank requests by their expected localization improvement. We therefore combine Lyapunov backlog control with predicted task value under wireless and computing budgets. Table I compares the proposed framework with representative works.

III. System Overview A. System Architecture We consider an edge-assisted multi-view localization system in a GPS-denied environment, where satellite positioning is unavailable, intermittent, or spoofed, as shown in Fig. 2. Each client carries 𝑀 onboard cameras, one for each view. A UAV uses 𝑀 = 5 front, back, left, right, and downward views, and a UGV uses 𝑀 = 4 horizontal views. At slot 𝑡, the client extracts the feature X𝑡(𝑚) = Φ(𝑉𝑡(𝑚) ) of the image 𝑉𝑡(𝑚) from view 𝑚 with the backbone Φ in Section IV-A, and the features 𝑀 form the multi-view feature X . The edge of V𝑡 = {𝑉𝑡(𝑚) } 𝑚=1 𝑡 𝑁D server stores a geo-tagged database D = {( 𝑓𝑛 , 𝑙 𝑛 )} 𝑛=1 , where 𝑓𝑛 is a reference descriptor, 𝑙 𝑛 is its location, and 𝑁 D is the number of entries. As shown in Fig. 3, our design encodes X𝑡 into the latent representation Z𝑡 = E (X𝑡 ; Θ𝐸 ) at the client and transmits Z𝑡 to the edge server. The edge localizer then estimates the position as Ŷ𝑡 = F (Z𝑡 ; Θ𝐹 , D). It combines decoder regression with nearest-neighbor retrieval by cosine similarity reg as Ŷ𝑡 = 𝜂Ŷ𝑡 +(1−𝜂) Ŷret 𝑡 , where the mixing weight 𝜂 ∈ [0, 1] is selected on the validation set. We write Θ𝐸 = 𝜙 and Θ𝐹 = 𝜃 in the rest of the paper. Because exhaustive retrieval over a large database is costly at the edge, we summarize the database with prototypes. A scene prototype is the mean of the descriptors 𝑓𝑛 from one scene, and a tile prototype is the mean of the descriptors whose locations 𝑙 𝑛 fall in one spatial tile of that scene. The retrieval branch first matches scene prototypes and then tile prototypes, and it searches only the descriptors in the selected tiles. This coarse-to-fine search reduces the cost of fine retrieval from 𝑂 (𝑁 D 𝑑r ) to 𝑂 (𝑁tile 𝑑r ), where 𝑑r is the descriptor dimension and 𝑁tile is the number of descriptors in the selected tiles. B. Communication Model We model the uplink between a client and the edge server as a block-fading channel, whose gain stays constant during the transmission of one request. The achievable rate is 𝑅 = 𝐵 log2 (1 + 𝑃𝑔/(𝑁0 𝐵)), where 𝐵 is the channel bandwidth, 𝑃 is the transmit power of the client, and 𝑁0 is the noise power spectral density. The channel gain 𝑔 = 𝑔0 ( 𝜚0 /𝜚) 𝜅 10 𝜉 /10 |ℎ| 2 captures large-scale and small-scale effects, where 𝑔0 is the path gain at the reference distance 𝜚 0 , 𝜚 is the client-edge distance, 𝜅 is the path-loss exponent, 𝜉 is the shadowing in dB, and ℎ is the small-scale fading coefficient. A payload of |Z𝑡 | bits therefore incurs the transmission delay 𝜏 = |Z𝑡 |/𝑅. In our evaluation, 𝑅 follows the controlled throughput profiles described in Section VII. C. Problem Formulation Our objective is to minimize the localization error subject to a per-request payload budget:   min E ∥ Ŷ𝑡 − Y𝑡 ∥ 22 s.t. C(Z𝑡 ) ≤ 𝐶max , (1) Θ

where Ŷ𝑡 and Y𝑡 are the predicted and true positions, Θ = {Θ𝐸 , Θ𝐹 } collects the encoder and localizer parameters, C(Z𝑡 ) = |Z𝑡 | is the payload size of Z𝑡 in bits, and 𝐶max is the

4

TABLE I Feature coverage across representative works. Feature

[16]

[10]

[30], [31]

[5]–[7]

[40]

[9]

[24], [25], [47]

[1]

Proposed work

Task-oriented compression Variable-rate representation Request control View selection Coarse-to-fine retrieval Uncertainty feedback Multi-client scheduling Edge-testbed validation

! △ × × × × × ×

! ! × × × × × !

△ ! △ △ × × △ !

× × × × △ × × ×

△ △ ! × × ! × ×

× × △ × △ △ △ !

× × × × × × ! !

! × × × × × × △

! ! ! ! ! ! ! !

Note: Columns are identified by citation. ! denotes fully addressed, △ denotes partially addressed or conditionally evaluated, and × denotes not addressed. Semantic payload / processing

Client (Mobile Device) Front

Back

Left

Multi-view inputs

Right

Down

Request Queue

Wireless Uplink

(�)

�� = {�� }� �=1

C1

Feature Extractor �(·) Local Odometry & Uncertainty

C2

Semantic Payload

∗

VOI Controller ��,�

C3

Light Metadata

Scalable Semantic Encoder

Prod

Payload

Enqueue

Deadline

Prod

Payload

Enqueue

Deadline

Prod

Payload

Enqueue

Deadline

Control & feedback signals Edge Localization Engine

VOI-aware Edge Scheduler Semantic Decoder

C2

High

C1

Medium

C3

Low

Retrieval Database

Localization Head

...

Corrected position

... Compact Semantic Payload

Pose Correction

Fig. 2. System architecture of network-adaptive edge-assisted multi-view localization. The client maintains a local fast loop, requests edge localization only when the predicted value is high, and receives correction feedback after shared wireless and edge-scheduling stages. View mask � 1 0 1 0

RGB inputs (multi-view)

...

Back

...

Left Right

Shared Feature Extractor �(∙)

... ... ...

Down

...

Selected multi-view features

Front

KL rate regularization

Orthogonal regularization

�� : ��

O-VIB encoder Feature fusion (per-view-selected views)

��

(a) View-masked features

Compact latent bottleneck (compression)

Ordered latent

...

(importance-ordered latent dimensions)

...

� ∈ ℝ�×��

...

⋮ ...

� ≪ ��

��

(b) O-VIB encoder

...

� ∈ ℝ�

(high → low importance)

Low rate

�� : ��

⋮

� ∈ ℝ�

�� : ��

Medium rate ...

High rate

(c) Nested latent prefixes

Fig. 3. Scalable O-VIB encoder for view-rate adaptive semantic localization. View masks select the available camera subset, while nested latent prefixes provide low-, medium-, and high-rate semantic payloads for one deployed encoder.

payload budget of one request. Lowercase x, z, y denote realizations of the flattened feature X𝑡 , the latent representation Z𝑡 , and the position Y𝑡 , respectively. We address (1) in three parts. Section IV trains one encoder that serves a family of payload budgets, Section V decides when to send a request and with which budget, and Section VI schedules the requests of multiple clients at the edge server. IV. Scalable Multi-View O-VIB Encoder A. Task-Oriented Feature Extraction We use a CLIP-based vision backbone for multi-view feature extraction. Each image 𝑉𝑡(𝑚) is processed by a

shared CLIP ViT-B/32 encoder [51]. After the preprocessing 𝜋(·), the embedding of view 𝑚 is X𝑡(𝑚) = Φ(𝑉𝑡(𝑚) ) = 𝑓CLIP (𝜋(𝑉𝑡(𝑚) ); 𝜃 Φ ) ∈ R𝑑 , where 𝜃 Φ denotes the pretrained CLIP parameters and 𝑑 = 512 is the embedding dimension. We use this pretrained backbone because it provides a reusable visual basis across scenes and does not add a taskspecific backbone to the resource-constrained client. We then normalize each embedding as X̃𝑡(𝑚) = X𝑡(𝑚) /∥X𝑡(𝑚) ∥ 2 and concatenate the normalized embeddings into X𝑡 ∈ R 𝑀 ×𝑑 .

5

B. Task-Oriented Feature Compression The multi-view feature X𝑡 ∈ R 𝑀 ×𝑑 is flattened into x ∈ R𝑑 𝑓 for joint encoding, where 𝑑 𝑓 := 𝑀 𝑑. Let 𝐾 denote the full latent dimension and z ∈ R𝐾 the full encoder output. Since the uplink bandwidth is limited, the client needs a compact and task-relevant representation [18]. The information bottleneck (IB) principle [52] and its variational form [53] provide the theoretical framework for learning such a representation. The IB principle seeks a stochastic encoder 𝑞 𝜙 (z|x) that minimizes 𝐼 (x; z) for compactness while maximizing 𝐼 (z; y) for localization relevance. Its objective is min 𝛽 𝐼 (x; z) − 𝐼 (z; y) , 𝜙 | {z } | {z } Rate

(2)

Accuracy

where the non-negative hyperparameter 𝛽 controls the tradeoff between compression and localization accuracy. This trade-off is important in our design, because the same latent representation serves both the edge localizer and the variablerate wireless link. 1) Variational Rate Term: The rate term 𝐼 (x; z) is controlled through a prior 𝑝(z) on the latent representation. The encoder is modeled as a diagonal Gaussian distribution  𝑞 𝜙 (z | x) = N 𝝁 𝜙 (x), diag 𝝈 2𝜙 (x) , and the latent representation is sampled as z = 𝝁 + 𝝈 ⊙ 𝝐 with 𝝐 ∼ N (0, I) through the reparameterization trick [54]. With the standard Gaussian prior 𝑝(z) = N (0, I), the rate term becomes   R (𝜙) := Ex KL 𝑞 𝜙 (z | x) ∥ 𝑝(z) 𝐾 ∑︁ (3)  𝜇𝑖2 + 𝜎𝑖2 − log 𝜎𝑖2 − 1 , = 21 Ex

the induced marginal. The mutual information under the joint density 𝑝(x)𝑞 𝜙 (z | x) satisfies h i  𝐼 (x; z) = Ex KL 𝑞 𝜙 (z | x) ∥ 𝑝(z) − KL 𝑞 𝜙 (z) ∥ 𝑝(z) (5) h i ≤ Ex KL 𝑞 𝜙 (z | x) ∥ 𝑝(z) . (6) Proof: Expanding the mutual information gives 𝐼 (x; z) = Ex,z [log 𝑞 𝜙 (z | x) − log 𝑞 𝜙 (z)]. Adding and subtracting Ez [log 𝑝(z)] yields (5). The second term in (5) is nonnegative for a proper density, and dropping it yields (6). Lemma 2: ´ For any decoder 𝑝 𝜃 (y|z) and the joint density 𝑝(z, y) = 𝑝(x, y)𝑞 𝜙 (z | x) dx induced by the Markov chain y → x → z, the mutual information between the latent representation z and the task variable y is lower-bounded by 𝐼 (z; y) ≥ Ez,y [log 𝑝 𝜃 (y|z)] + h(y), (7) where h(y) = −Ey [log 𝑝(y)] is the differential entropy of the continuous position y and is constant with respect to (𝜙, 𝜃). Proof: By definition, 𝐼 (z; y) = h(y) − h(y|z).

(8)

h i E𝑞 𝜙 (z) KL 𝑝(y|z) ∥ 𝑝 𝜃 (y|z) ≥ 0,

(9)

Moreover,

which implies Ez,y [log 𝑝(y|z)] ≥ Ez,y [log 𝑝 𝜃 (y|z)] .

(10)

Ez,y [log 𝑝(y|z)] = −h(y|z),

(11)

Since

𝑖=1

where KL(·∥·) denotes the Kullback-Leibler divergence. This proper prior gives a tractable rate surrogate shared by every latent prefix in Section IV-B3, and Lemma 1 shows that this surrogate upper-bounds 𝐼 (x; z). For the fixed-rate model of our conference version [1], we instead use a log-uniform prior 𝑝(𝑧 𝑖 ) ∝ |𝑧 𝑖 | −1 , which yields an automatic relevance determination (ARD) penalty that prunes uninformative coordinates [55]. Its KL divergence admits the closed-form approximation Dard (x) :=

𝐾 h ∑︁

−1 

1 2 log 1 + 𝛼𝑖



− 𝑐 1 sigm 𝑐 2 + 𝑐 3 log 𝛼𝑖 + 𝑐 1

i

𝑖=1

 ≈ KL 𝑞 𝜙 (z | x) 𝑝 LU (z) , (4) where 𝑝 LU (𝑧 𝑖 ) ∝ |𝑧 𝑖 | −1 denotes the log-uniform prior, 𝛼𝑖 := 𝜎𝑖2 /𝜇𝑖2 is the noise-to-signal ratio of coordinate 𝑖, sigm(𝑢) = 1/(1 + 𝑒 −𝑢 ) is the sigmoid function, and (𝑐 1 , 𝑐 2 , 𝑐 3 ) = (0.63576, 1.87320, 1.48695) are the constants fitted in [55]. The additive constant 𝑐 1 makes each summand vanish as 𝛼𝑖 → ∞, which corresponds to the pruning regime. This approximation deviates from the exact KL divergence by less than 0.009 over the full range of log 𝛼𝑖 [55]. We write R ard (𝜙) := Ex [Dard (x)] for this variant. Both expectations over x are estimated by minibatch averaging. Lemma 1: Let 𝑞 𝜙 (z | x) be any encoder, let 𝑝(z) be any ´ proper prior density, and let 𝑞 𝜙 (z) = 𝑞 𝜙 (z | x) 𝑝(x) dx be

substitution into 𝐼 (z; y) = h(y) − h(y|z) yields (7), which is the variational lower bound of [56]. Theorem 1: For the proper Gaussian prior 𝑝(z) = N (0, I) used by the scalable encoder and any decoder 𝑝 𝜃 (y | z), the IB objective LIB := 𝛽𝐼 (x; z) − 𝐼 (z; y) in (2) satisfies   LIB ≤ 𝛽 R (𝜙) − Ez,y log 𝑝 𝜃 (y | z) − h(y). (12) Since −h(y) is constant with respect to (𝜙, 𝜃), minimizing the right-hand side of (12) is equivalent to solving   min 𝛽 R (𝜙) − Ez,y log 𝑝 𝜃 (y | z) . (13) 𝜙, 𝜃

Proof: Lemma 1 gives 𝛽𝐼 (x; z) ≤ 𝛽R (𝜙) for 𝛽 ≥ 0, and Lemma 2 gives −𝐼 (z; y) ≤ −Ez,y [log 𝑝 𝜃 (y | z)] − h(y). Adding the two inequalities yields (12), and removing the constant −h(y) yields (13). In practice, we optimize (13) with the Gaussian prior for the scalable model. The fixed-rate model uses R ard in place of R. Since the log-uniform prior is improper, Theorem 1 applies only to the Gaussian-prior formulation. The second term in (13) is induced by the variational decoder 𝑝 𝜃 (y | z). With a Gaussian position decoder of fixed isotropic variance, the negative log-likelihood equals a squared-error regression loss up to additive and multiplicative constants, and the multiplicative constant is absorbed into the localization weight

6

𝛼 in (20). We additionally train a reconstruction decoder that maps the latent representation back to the feature space. Its output is used by the retrieval branch at the edge, and the reconstruction loss also regularizes the training. 2) Orthogonality Under the IB Objective: Let h 𝜙 (x) ∈ R𝑑ℎ be the final hidden feature of the encoder, where 𝑑 ℎ = 256, and let W ∈ R𝐾 ×𝑑ℎ be the weight matrix of the linear projection that maps h 𝜙 (x) to the posterior mean 𝝁 𝜙 (x). The rate penalty pulls posterior coordinates toward the prior, while the nested-prefix training in Section IV-B3 favors early coordinates. Several rows of W may then become aligned, causing redundant latent directions. We therefore impose approximate row-orthogonality on W. Proposition 1 shows that if WW⊤ is close to the identity, all singular values remain bounded away from zero and the mean projection avoids rank collapse across its 𝐾 directions. Proposition 1: Let W ∈ R𝐾 ×𝑑ℎ denote the weight matrix of the posterior-mean projection. Assume the approximate orthogonality condition WW⊤ = I𝐾 + 𝚫,

(14)

where 𝚫 is a symmetric perturbation matrix satisfying ∥𝚫∥ 2 ≤ 𝜀 for some 0 < 𝜀 < 1. Then all singular values 𝜎𝑖 (W) of W are bounded as √ √ 1 − 𝜀 ≤ 𝜎𝑖 (W) ≤ 1 + 𝜀, 𝑖 = 1, . . . , 𝐾. (15) In particular,

√ 𝜎min (W) ≥ 1 − 𝜀,

(16)

so W has full row rank and the posterior-mean projection avoids rank collapse. Proof: The singular values of W are the square roots of the eigenvalues of the 𝐾 × 𝐾 matrix WW⊤ . Since WW⊤ = I𝐾 + 𝚫

with

∥𝚫∥ 2 ≤ 𝜀,

(17)

every eigenvalue of WW⊤ lies in the interval [1 − 𝜀, 1 + 𝜀]. Therefore, for each singular value 𝜎𝑖 (W), 1 − 𝜀 ≤ 𝜎𝑖2 (W) ≤ 1 + 𝜀.

(18)

Taking square roots gives the claimed bound. Since ∥𝚫∥ 2 ≤ ∥𝚫∥ 𝐹 , the normalized Frobenius penalty 𝐾 −2 ∥WW⊤ − I𝐾 ∥ 2𝐹 in (20) controls the perturbation in the proposition. Section VII-B evaluates the empirical effect of the orthogonality penalty. 3) Scalable View-Rate Encoding: Rather than compressing each camera stream in isolation, we concatenate the 𝑀 viewwise embeddings and pass them through one O-VIB encoder so that the joint code captures complementary information across views. To support rate adaptation without a separate encoder for each bitrate, let K ⊆ {1, . . . , 𝐾 }, with 𝐾 ∈ K, denote the supported prefix lengths. For 𝑘 ∈ K, the client transmits z (𝑘 ) = [𝑧 1 , . . . , 𝑧 𝑘 ] and masks the remaining coordinates. A prefix contains 4𝑘 float32 bytes and a 20-byte representation header. The prefix ordering is learned during training through nested rate dropout, following the ordered-representation principle of nested dropout [32] and Matryoshka representation

learning [33]. For each minibatch, we sample a prefix length 𝑘 ∈ K and apply the binary prefix mask m (𝑘 ) = [1, . . . , 1, 0, . . . , 0], | {z } | {z } 𝑘

z̃ (𝑘 ) = m (𝑘 ) ⊙ z.

(19)

𝐾 −𝑘

The same decoder reconstructs features and supports position inference from z̃ (𝑘 ) for all 𝑘. If 𝑖 < 𝑗, coordinate 𝑖 is active whenever coordinate 𝑗 is active, so earlier coordinates participate in at least as many prefix losses and become importance ordered. We handle view adaptation with a view mask. Let S𝑖 be the set of admissible view modes on the platform of client 𝑖. For the UAV, SUAV = {F, FS, H4, A5}, where F uses the front view, FS adds the left and right views, H4 uses the four horizontal views, and A5 uses all five views. For the four-view UGV, SUGV = {F, FS, H4}, where H4 uses all available views. For mode 𝑆 ∈ S𝑖 , inactive embeddings are set to zero before joint encoding, and the same view mask is applied to the decoded retrieval descriptor. Zeroing keeps one fixed encoder input for all view subsets. Training covers the UAV modes and prefixes, while the UGV evaluation uses its feasible modes. The application-layer request size is 𝑏 𝑖 (𝑆, 𝑘) = 4𝑘 + 20 + 4|𝑆| bytes, where |𝑆| is the number of selected views and each selected view adds a 4-byte identifier. Thus, one encoder supports A𝑖 = S𝑖 × K. The online controller uses the compact and rich prefixes Kctrl = {16, 32} and evaluates A𝑖,ctrl = S𝑖 × Kctrl . 4) Overall Training Objective: Let m𝑆 ∈ {0, 1} 𝑀 be the view mask for mode 𝑆, broadcast over the 𝑑 coordinates of each view, and let x𝑆 ∈ R𝑑 𝑓 be the resulting flattened feature. The encoder maps x𝑆 to 𝝁 𝜙 (x𝑆 ) and 𝝈 𝜙 (x𝑆 ). Reparameterization produces z, and (19) produces z̃ (𝑘 ) . The reconstruction decoder outputs x̂𝑆,𝑘 ∈ R𝑑 𝑓 , and the edge localizer outputs ŷ𝑆,𝑘 ∈ R3 from the reconstructed descriptor. By combining the reconstruction loss, the localization loss, the rate term, and the orthogonality penalty, with the view mode and the prefix length sampled during training, we obtain the overall training objective as h i h i 2 2 L (𝜙, 𝜃) = Ex,𝑆,𝑘 x𝑆 − x̂𝑆,𝑘 2 + 𝛼 Ex,y,𝑆,𝑘 y − ŷ𝑆,𝑘 2   + 𝛽 Ex,𝑆 KL 𝑞 𝜙 (z | x𝑆 ) ∥ 𝑝(z) 𝛾 2 + 2 WW⊤ − I𝐾 𝐹 , 𝐾 (20) where 𝑆 ∈ SUAV is the sampled view mode, 𝑘 ∈ K is the sampled prefix length, and 𝜃 denotes the reconstruction decoder and edge-localizer parameters. The coefficients 𝛼, 𝛽, 𝛾 > 0 balance the four terms. The reconstruction term preserves appearance cues needed by retrieval, while the localization term remains the task objective. Sampling (𝑆, 𝑘) implements nested view-rate training. The rate term penalizes retained information, the prefix length determines the transmitted coordinates, and orthogonality constrains the rows of W. V. VOI-Guided Request and Mode Selection The scalable encoder provides multiple view-rate modes, and the client decides whether an edge correction is worth

7

(a) Trigger timing

(b) Channel state

(c) Sparse semantic uploads

4 2 0

50

10 8

40

6

30

4

20

2

Queue (ms)

6

2.0 Cumulative KB

8

Throughput (Mbps)

Uncertainty (m)

12

All5 Front+Sides

1.5 1.0 0.5 0.0

0

1000

2000

3000

4000

Distance (m)

0

1000

2000

3000

4000

0

1000

Distance (m)

2000

3000

4000

Distance (m)

Fig. 4. VOI-Control request timing on a held-out route. The client delays edge requests while local risk is low, triggers sparse semantic uploads around uncertainty growth, and adapts decisions under time-varying throughput and queue state.

requesting in the current slot. The client does not observe ground-truth localization error online, so it estimates the tasklevel VOI before transmission. Gross VOI is the predicted reduction in localization risk from an edge correction, and net VOI further subtracts the communication and service costs. Fig. 4 illustrates the request timing of our online controller, denoted VOI-Control, on the held-out route of Section VII. A. Local Belief and Predicted Edge Gain For client 𝑖, let ȳ𝑖,𝑡 and P𝑖,𝑡 denote the locally propagated position estimate and its error covariance at slot 𝑡, obtained from onboard odometry or a lightweight local estimator. The client navigates with ȳ𝑖,𝑡 between edge corrections, while P𝑖,𝑡 grows as odometry drift accumulates. We use √︁ loc (21) L𝑖,𝑡 = tr(P𝑖,𝑡 ), as a scalar local-risk proxy in meters. The controller uses this uncertainty statistic for action ranking and does not require ground-truth position error online. For each candidate action 𝑎 = (𝑆, 𝑘) ∈ A𝑖,ctrl , the client predicts the residual risk after edge correction with  bedge (𝑎) = 𝑔 𝜓 𝑐 𝑖,𝑡 , 𝑆, 𝑘 , L (22) 𝑖,𝑡 where 𝑔 𝜓 is a residual predictor with parameters 𝜓, and 𝑐 𝑖,𝑡 is a locally available route-context category, such as corridor or intersection. In the evaluated replay, 𝑔 𝜓 returns the median edge residual observed on training routes for the same context and action (𝑆, 𝑘). The lookup uses no held-out position labels. The median reduces the effect of rare high-residual views and stabilizes action ranking. B. Net VOI and Online Action The predicted gross VOI of action 𝑎 = (𝑆, 𝑘) is h i+ b𝑖,𝑡 (𝑎) = L loc − L bedge (𝑎) . ΔL 𝑖,𝑡 𝑖,𝑡

Clients Client 1

Client 2

VOI-aware Edge Scheduler Prioritized Execution Order

Localization Request VOI

Payload

Deadline

High Priority

1

Client 2

VOI

Payload

Deadline

2

Client 1

VOI

Payload

Deadline

Bandwidth

Localization Request VOI

Payload

Shared Edge Resources

Deadline

Localization Service

GPU

Client 3

Localization Request VOI

Payload

Deadline

Low Priority

3

Client 3

VOI

Payload

Deadline

Fig. 5. VOI-weighted multi-client edge scheduling. Each request carries localization value, payload, estimated service time, and deadline metadata, allowing the edge to prioritize requests that are both urgent and task-relevant.

b𝑖,𝑡 (𝑎) be the predicted queueing, service, and return and 𝐷 delay excluding uplink transmission. The net VOI is b𝑖,𝑡 (𝑎) − 𝜆 𝑏 8𝑏(𝑎) − 𝜆 𝑒 𝐸 (𝑎) − 𝜆 𝑑 𝐷 b𝑖,𝑡 (𝑎), (24) b 𝜈𝑖,𝑡 (𝑎) = ΔL 𝑅𝑖,𝑡 where 𝜆 𝑏 , 𝜆 𝑒 , and 𝜆 𝑑 convert transmission time, energy, and delay into the same utility scale as the localization risk. Our trace replay computes the transmission and delay terms from the payload and latency traces. The energy term applies only when an endpoint energy estimate is available. The client chooses 𝑎 ∗𝑖,𝑡 = arg max b 𝜈𝑖,𝑡 (𝑎), (25) 𝑎∈ A𝑖,ctrl

and submits an edge request only if b 𝜈𝑖,𝑡 (𝑎 ∗𝑖,𝑡 ) > 0 and any configured local energy constraint is satisfied. Besides the compressed latent prefix, the request packet carries the selected view-rate mode, predicted VOI, payload size, and deadline metadata, which form the 20-byte header counted in 𝑏(𝑎). VI. VOI-Weighted Edge Scheduling

(23)

Here, [𝑢] + := max{𝑢, 0} denotes the positive part, which assigns zero gross VOI when the predicted residual is no smaller than the local-risk proxy. For client 𝑖 and action 𝑎 = (𝑆, 𝑘), let 𝑏(𝑎) = 𝑏 𝑖 (𝑆, 𝑘) be the application-layer payload in bytes, 𝑅𝑖,𝑡 be the current uplink rate in bits per second, 𝐸 (𝑎) be the estimated endpoint energy consumption,

When multiple clients request edge localization simultaneously, the edge server allocates wireless and computing resources according to queue urgency and localization value, as illustrated in Fig. 5. Let R 𝑡 be the set of clients with at least one pending request at slot 𝑡, and let 𝑄 𝑖,𝑡 be the pendingrequest count for client 𝑖. The head-of-line request of client 𝑖 carries the net VOI b 𝜈𝑖,𝑡 predicted when it is generated, payload 𝑏 𝑖,𝑡 , and estimated service time 𝐺 𝑖,𝑡 . The edge server chooses 𝑥𝑖,𝑡 ∈ {0, 1} by solving

8

max

{ 𝑥𝑖,𝑡 }

s.t.

∑︁

𝑥 𝑖,𝑡 𝑄 𝑖,𝑡 + 𝑉b 𝜈𝑖,𝑡

denoted VOI-Lyapunov. The Base-DPP priority follows the drift-plus-penalty construction [57].



𝑖 ∈ R𝑡

∑︁

𝑥𝑖,𝑡 𝑏 𝑖,𝑡 ≤ 𝑏 max 𝑡 ,

𝑖 ∈ R𝑡

∑︁

(26)

𝑥𝑖,𝑡 𝐺 𝑖,𝑡 ≤ 𝐺 max 𝑡 ,

𝑖 ∈ R𝑡

where 𝑏 max and 𝐺 max are the slot-level wireless and edge𝑡 𝑡 computing budgets, and 𝑉 > 0 controls the utility-delay tradeoff. In the replay, one work-conserving edge server processes one request at a time. Its per-request service time is the measured mode-dependent compute latency plus 3.1 ms for O-VIB decoding, 5.2 ms for descriptor retrieval, 4.0 ms of runtime overhead, and lognormal jitter with a 2.2-ms median. Reported edge latency includes queueing before service. The baseline scheduler, denoted Base-DPP, serves one pending request at a time using the priority 𝑄 𝑖,𝑡 + 𝑉b 𝜈𝑖,𝑡 , rather than jointly optimizing multiple requests within a slot. Each queue evolves as  + 𝑄 𝑖,𝑡+1 = 𝑄 𝑖,𝑡 − 𝑥 𝑖,𝑡 + 𝐴𝑖,𝑡 , (27) where 𝐴𝑖,𝑡 indicates whether a new request from client 𝑖 is admitted into the edge queue. The rule in (26) raises priority with backlog and predicted localization value. It does not explicitly encode waiting age or impending deadline misses, which motivates the shaped rule below. We therefore design a shaped scheduler. For queued request 𝑗, define the mission-weighted scheduling value n  o b b 𝑣 𝑗,𝑡 = max 0.1, 𝜔 𝑗 b 𝜈 𝑗,𝑡 + 1.8𝑟 drift + 0.22L loc , 𝑗 𝑗 − 0.006 𝐷 𝑗,𝑡 (28) is the fitted drift rate in where 𝜔 𝑗 is mission urgency, 𝑟 drift 𝑗 loc b 𝑗,𝑡 is predicted meters per second, L 𝑗 is local risk, and 𝐷 latency in milliseconds. The coefficients convert the drift, risk, and latency terms to the meter scale of b 𝜈 𝑗,𝑡 , and the floor keeps the value positive. Let nrm𝑡 (𝑢 𝑗 ) denote min-max normalization over the queued requests, with a zero output when the range is zero. The priority score is    b 𝑣 𝑗,𝑡 + 𝑤 𝑟 nrm𝑡 (b 𝑣 𝑗,𝑡 ) + 𝑤 ℓ 1 − nrm𝑡 (ℓ 𝑗,𝑡 ) 𝑠 𝑗,𝑡 = 𝑤 𝑣 nrm𝑡 𝐺 𝑗,𝑡 + 𝑤 𝑎 nrm𝑡 (𝜁 𝑗,𝑡 ) + 𝑤 𝑜 nrm𝑡 (𝑜 𝑗,𝑡 ) + 𝑤 𝑠 nrm𝑡 (−𝐺 𝑗,𝑡 ). (29) Here 𝐺 𝑗,𝑡 is the estimated service time. Let 𝜏𝑡 be the rdy clock time at scheduling decision 𝑡, 𝑡 𝑗 the ready time, and 𝑑 𝑗 the absolute deadline, all expressed in the same time unit as the relative deadline 𝐷 ddl 𝑗 . The replay uses rdy + ℓ 𝑗,𝑡 = [𝑑 𝑗 − 𝜏𝑡 − 𝐺 𝑗,𝑡 ] , 𝜁 𝑗,𝑡 = [𝜏𝑡 − 𝑡 𝑗 ] + /𝐷 ddl 𝑗 , and 𝑜 𝑗,𝑡 = [𝜏𝑡 + 𝐺 𝑗,𝑡 − 𝑑 𝑗 ] + /𝐷 ddl . The laxity ℓ measures the 𝑗,𝑡 𝑗 remaining service margin, 𝜁 𝑗,𝑡 is the normalized waiting age, and 𝑜 𝑗,𝑡 measures the predicted completion overrun, which can be positive before the deadline. The selected weights are (𝑤 𝑣 , 𝑤 𝑟 , 𝑤 ℓ , 𝑤 𝑎 , 𝑤 𝑜 , 𝑤 𝑠 ) = (3.2, 1.2, 1.2, 1.4, 0.6, 0.1), fixed before testing. The waiting-age term increases the priority of older requests. This scoring effect is separate from the stability guarantee of (26). We evaluate the single-request Base-DPP rule described above and the shaped scheduler,

Theorem 2: Suppose that the arrival indicators 𝐴𝑖,𝑡 , payloads 𝑏 𝑖,𝑡 , service times 𝐺 𝑖,𝑡 , and predicted values b 𝜈𝑖,𝑡 are uniformly bounded, are independent and identically distributed across slots, and are independent of the queue state. Suppose that (26) is solved exactly in each slot and that a stationary randomized policy satisfies the per-slot budgets while serving every client at a rate above its mean arrival rate by a slack 𝛿 > 0. Then every queue under (26) is mean-rate stable. Í −1 The Í time-average predicted-VOI utility lim inf 𝑇→∞ 𝑇 −1 𝑇𝑡=0 𝜈𝑖,𝑡 ] is within 𝑂 (1/𝑉) of the 𝑖 E[𝑥 𝑖,𝑡 b largest utility achievable by any stabilizing policy, and the lim sup of the time-average total backlog is 𝑂 (𝑉). Í Proof: Let 𝐿 (Q𝑡 ) = 12 𝑖 𝑄 2𝑖,𝑡 be the quadratic Lyapunov function and let Δ(Q𝑡 ) denote its one-slot conditional drift. Using the queue update and bounded arrivals and services gives ∑︁ Δ(Q𝑡 ) ≤ 𝐶0 + 𝑄 𝑖,𝑡 E[ 𝐴𝑖,𝑡 − 𝑥 𝑖,𝑡 | Q𝑡 ] (30) 𝑖

for a finite constant 𝐶0 determined Í by the bounds on arrivals and services. Subtracting 𝑉 E[ 𝑖 𝑥𝑖,𝑡 b 𝜈𝑖,𝑡 | Q𝑡 ] from both sides yields the drift-plus-penalty expression. The scheduler in (26) minimizes its right-hand side in every slot because the arrival terms do not depend on the decision. Comparing this minimum with the stationary randomized policy, taking expectations, and summing over time give mean-rate stability, an 𝑂 (1/𝑉) utility gap, and an 𝑂 (𝑉) backlog bound by the standard drift-plus-penalty argument [57]. Theorem 2 concerns the exact slotted optimizer under the stated independence assumptions on both arrivals and request attributes. It does not establish guarantees for the singlerequest replay, whose head-of-line request attributes persist until service, or for the shaped rule. In the coupled replay, VOI-Control also uses the predicted queueing delay in (24), so arrivals depend on the queue state. Section VII-F therefore evaluates the shaped policy empirically. VII. Performance Evaluation A. Experimental Setup Table II and Fig. 6 summarize the data and evaluation settings. Per-mode residuals and payloads come from trainedmodel inference. Semantic-path latency is measured on a Jetson Orin NX after feature extraction. The request-control experiments additionally simulate odometry drift, request timing, and network delay. Unless otherwise specified, the localization error is the Euclidean distance to the ground truth, and the application-layer request adds selected-view identifiers to the serialized 4𝑘 + 20-byte semantic representation.1 The 4,994-frame Town05 set and the 6,684-frame fivescene collection are CARLA data with logged poses. The 4,593-frame sampled benchmark comes from the latter. Indoor UAV and UGV data use Qualisys motion-capture ground 1 The multi-view UAV dataset was collected by the authors and is released at https://huggingface.co/datasets/Peter341/Multi-View-UAV-Dataset. The code is available at https://github.com/fangzr/TOC-Edge-Aerial.

9

TABLE II Datasets and evaluation configurations used in the study. Setting

Data path

Scenes

Single-scene split Raw multi-scene data Sampled benchmark Control workload Retrieval queries Indoor UAV Indoor UGV

Conference-version Town05 CARLA UAV CARLA subset Town02 route replay Offline query set Qualisys capture Qualisys capture

1 scene 5 scenes 5 scenes Held-out route 5 scenes 4.1 m × 4.0 m 5.0 m × 4.0 m

Frames/routes

Views

Ground truth

Main purpose

4,994; 3,495/749/750 6,684 frames 4,593 frames 593 slots / 7 policies 400 queries 76 ref. / 58 query 78 ref. / 72 query

5 5 5 Variable 5 5 sequential 4 sequential

Simulator pose Simulator pose Simulator pose Simulator pose Simulator pose 6-DoF mocap 6-DoF mocap

Scalable compression and ablations Raw samples and retrieval database Coarse-to-fine retrieval and plots 𝑔 𝜓 request and mode control Scene coarse + tile-pruned fine search Real-scene compact localization Real-scene compact localization

Back

Right Bottom Deploying downwards

Left

Front (a)

(b)

(c)

(d)

(e)

(f)

Fig. 6. Self-collected multi-scene multi-view UAV data rendered in the CARLA simulator. The synchronized five-view RGB observations and logged simulator poses support retrieval, feature extraction, and model training. TABLE III Deployment overhead of fixed-rate and scalable O-VIB configurations under the Jetson Orin NX latency profile. Method

Models

Storage

Rates

Latency @ 8 KB/s

Fixed-rate O-VIB-32 1 2.8 MB 1 50.4 ms Fixed-rate O-VIB-128 1 2.9 MB 1 100.6 ms Fixed-rate bank {32, 128} 2 5.7 MB 2 50.4–100.6 ms Scalable O-VIB 𝑘 = 35 1 13.9 MB 128 52.0 ms Scalable O-VIB 𝑘 = 16 1 13.9 MB 128 42.0 ms Note: Latency combines the measured Jetson compute profile with transmission of the serialized 4𝑘 +20-byte O-VIB payload at 8 KB/s. The Rates column counts prefix lengths supported at runtime, not independently trained models.

truth. Compression experiments fix the downstream localizer after training, which we call the frozen-localizer protocol. Controller configurations are selected on tuning seeds and evaluated on disjoint held-out seeds, where each seed fixes one realization of the channel, drift, request timing, and queue state. The Town05 scalable encoder uses 𝐾 = 128, K = {4, 8, 16, 32, 64, 96, 128}, 50 AdamW epochs, batch size 512, learning rate 10−3 , and (𝛼, 𝛽, 𝛾) = (1, 10−3 , 0.1). Raw-image retrieval uses a 2 Mbps uplink. Semantic-payload experiments use 4 to 12 KB/s profiles. Request replay uses 0.5-s slots and a shared edge queue. B. Scalable Compression Performance Fig. 7 uses 3,495 training, 749 validation, and 750 test frames from Town05. Table III reports deployment overhead. One model is evaluated at every 𝑘 = 1, . . . , 128, so panel (a) is not interpolated. Validation selects 𝑘 = 35, which gives 9.18 m test error at 0.156 KB/request. Prefix 𝑘 = 16 gives

9.31 m at 0.082 KB and 𝑘 = 128 gives 10.11 m at 0.520 KB. Panel (b) compares serialized payloads, while panel (c) checks tail errors. At 8 KB/s after feature extraction, the fixed-rate O-VIB-32 and O-VIB-128 models in Table III take 50.4 and 100.6 ms, compared with 2.92 s for WebP, 3.61 s for H.265, 5.04 s for H.264, and 4.58 s for JPEG. These values recompute transmission from the serialized payload rather than the preliminary conference-version profiles. One checkpoint supports dense runtime rate adaptation. Fig. 8 reports separately trained ordering and component ablations under the same frozen-localizer protocol. Validation selects 𝛾 = 0.1 at 𝑘 = 35, giving 8.85 m validation and 8.75 m test error. This checkpoint is independent of the 9.18 m model in Fig. 7. At 𝑘 = 35, the full model, denoted Full, obtains 8.75 m, compared with 9.44 m without the rate term, 9.30 m without orthogonality, and 116.19 m without nested-prefix supervision. Random and tail-first coordinate orderings remain worse than the learned prefix ordering until most coordinates are retained. Larger 𝛾 values do not improve validation error, and the model without nested-prefix training collapses at short prefixes. Fig. 9 summarizes the conference-version single-scene module results. It uses the same Town05 split as Fig. 7. Panel (a) shows that increasing the ARD weight reduces the active dimensions from 8 to 2 and raises the mean error from 9.57 m to 16.65 m. In the fixed-𝑘 = 16 sweep, the best reproduced OVIB setting is 𝛾 = 0.1 and 𝛽 = 0.001, with 9.06 m mean error over 10 trials, compared with 10.92 m for the conferenceversion multi-view fusion baseline. This fixed-rate model uses the log-uniform prior, and the experiment isolates the effects of ARD and orthogonality from scalable-prefix training. C. Coarse-to-Fine Retrieval Fig. 10 evaluates 400 sampled multi-scene queries. The mean raw five-view payload is 891.5 KB. On a Jetson Orin NX 16 GB with a 2 Mbps uplink, raw-image offloading takes 3.59 s, with 99.3% spent on upload. Semantic brute-force and coarse-to-fine retrieval take 5.75 and 4.21 ms after upload, with 93.3% of fine-search descriptors pruned by the latter. The coarse stage keeps the two most similar scene prototypes and the twelve most similar tile prototypes, and the scene top2 accuracy and tile top-12 recall are both 98.5%. Conditional on retaining the correct tile, both methods obtain 2.60 m mean error. Over all queries, coarse-to-fine rises to 6.52 m because six coarse-stage misses select the wrong scene, whereas brute force remains at 2.60 m. The result quantifies the accuracy cost of pruning at this database size.

10

(b) Payload-error Pareto Per-frame error (m)

H265 JPEG

k=128

10.5 10.0 k=16

k=35

9.5 9.0

WEBP Scalable O-VIB

101 16 35

64

96

90th pct.

20

100

Encode Decode Transmit

15 10 5

5000

k=16

4.58s 3.61s

2.92s

3000 2000 1000 0

102

5.04s

4000

0

128

Active prefix dimensions k

(d) Jetson latency breakdown

mean

Latency at 8 KB/s (ms)

102

11.0

(c) Error distribution

H264

Mean error (m)

Mean localization error (m)

(a) Prefix-rate curve

k=35 k=128

50ms101ms

-32 128 ebP .265 .264 JPEG W H H VIB IBO- O-V

Serialized payload (KB/request)

Fig. 7. Scalable O-VIB compression on Town05. Panel (a) evaluates one trained model over prefix dimension 𝑘. Panel (b) compares serialized payload-error frontiers with image-codec experiments. Panel (c) reports per-frame errors. Diamonds mark means and red ticks mark p90 values. Panel (d) reports postfeature-extraction Jetson latency at 8 KB/s.

(c) Fixed-rate, k=35 10.5

random mask

150

9.4 9.2 9.0

tail first

100

50 caps: 95% test CI teal band: random-order std

8.8

116 m

9.75 Mean error (m)

9.6

(d) Rate compatibility

10.00

ordered prefix

Mean error (m)

Error at k=35 (m)

(b) Prefix-order

val test selected

9.44

9.50

w/o nested: 150 m at k=16

Mean error (m)

(a) Gamma sweep 9.8

9.30

9.25 9.00

10.0

9.5

Full scalable w/o KL rate w/o orth. w/o nested

9.0

8.75

8.75

0 22

0 10−310−210−1 1 3

25

4

l . Ful No KL orth ested No No n

Active dimensions

γ

8

16 32 64 128

Active dimensions

Fig. 8. Scalable encoder ablations on Town05 with the frozen-localizer protocol. Panel (a) sweeps 𝛾 at 𝑘 = 35, with the star marking the validation-selected positive-𝛾 Full model. Panel (b) reports 95% confidence intervals over 750 test frames. The teal band additionally shows standard deviation across eight random coordinate permutations. Panel (c) compares components at 𝑘 = 35. Panel (d) zooms into 8.6 to 10.6 m. (c) Error map

Error

12

4 γ = 0.04

2 100

101

ARD weight β

10

10.5 10.0 9.5

0.02

9.62

9.43

9.50

35.2

0.06

9.23

9.08

9.46

30

9.5 9.4 9.3 9.2

9.0

32.1

0.1

9.45

9.06

Latency (ms)

14

Active dims

(d) Jetson compute cost 9.6

Orth. γ

6

Mean error (m)

16 Mean error (m)

Active latent dims

(b) Trial dispersion 11.0

m

(a) ARD rate-error 8

20 10

9.41

3.1

9.1

8.5

0 0.0004 0.001

0.005

β at γ = 0.1

4e-4

1e-3

5e-3

Enc.

Dec.

Total

Bottleneck β

Fig. 9. Reproduction of the conference-version single-scene O-VIB experiments. Panel (a) shows ARD-controlled active dimensions and error. Panels (b) and (c) report the fixed-𝑘 = 16 sweep, and panel (d) reports Jetson-profile compute cost.

D. View-Rate Selection We compare VOI-Control with five causal request policies under the 2.1 KB/route budget and include an oracle action bound. Periodic sends an FS compact request every 16 s. Budgeted periodic places the same number of FS compact requests as VOI-Control at uniformly spaced slots. Fixedfront sends F compact requests when the local risk exceeds 8 m, with a minimum gap of 1 s. Random uses the same request count and FS compact mode at uniformly sampled slots. Uncertainty sends FS compact requests when the local risk exceeds 6 m, with a minimum gap of 6 s. Oracle action uses the request times selected by VOI-Control and chooses

the action with the lowest realized residual after applying the same latency and payload penalties. Fig. 11 analyzes VOI-Control on Town02 routes containing corridor, intersection, and repetitive scenes. The prefix checkpoint is trained on frames 000000 to 000355 with masked reconstruction, four UAV view modes, and prefix set {16, 32, 35, 64, 128}. It uses neither position labels nor held-out route frames. Candidate residuals use leave-one-out retrieval with the same-scene database, and the median prior is fitted on training routes. Tuning seeds select the lowestrisk feasible configuration, which uses compact prefixes for admitted requests. Across five held-out seeds with 95% con-

11

(b) Jetson E2E @2 Mbps

(c) Pruning reliability

Rate

103

103

102 Semantic search

101

(d) Retained-tile accuracy

1.0

5

0.8

4 Error (m)

Raw-upload share: 99.3%

E2E latency (ms)

Fine candidates/query

(a) Candidate pruning

0.6 0.4 0.2

2 e@

en Sc

Tile

2 @1

ge ima Raw

F

ic B

ant Sem

tic man

C2F

2

0

2 12 e@ e@ en Til Sc

Se

3

1

0.0 te Bru

tile recall=98.5%

Brute mean

ne Pru

C2F mean

Brute p90

C2F p90

Fig. 10. Coarse-to-fine retrieval on the sampled multi-scene benchmark. Scene and tile prototypes reduce the fine-search set. Panel (c) reports the pruned fraction, and panel (d) reports accuracy conditional on retaining the correct tile. All-query accuracy is given in the text.

(a) Train-only action risk

(b) Budget tuning

2.9

2.9

2.9

2.9

2.8

3.0

inter.

3.1

3.2

2.2

2.2

2.2

2.2

2.3

2.2

12

6

10

5

occl.

4.1

4.3

3.0

3.0

3.0

3.0

3.0

3.0

4

open

7.1

6.0

4.0

3.7

3.6

3.5

3.5

3.5

budget=2.1 KB

8 6

3

repet.

4.0

C

F-

4.1

R

F-

2.0

-C FS

2.0

R

FS

2.0

-C

H4

2.0

-R

H4

2.0

-C

A5

over budget

4

2.0

(c) Held-out route 6.5 Mean route error (m)

4.8

median m

4.9

Mean error (m)

corr.

7

feasible grid

0

-R

6.0 Fixed front

5.5

Budgeted periodic Periodic

5.0 4.5

Uncertainty VOI-Control

4.0

selected

A5

Random

Oracle action

1

2

1.8

Payload on tuning seeds (KB)

1.9

2.0

2.1

Traffic (KB/route)

Fig. 11. Calibration and evaluation of VOI-Control. Panel (a) shows the train-only context/action median residual. F, FS, H4, and A5 denote front, frontplus-sides, four-horizontal, and all-five views. C and R denote 𝑘 = 16 and 𝑘 = 32. Panel (b) shows tuning-seed budget search, and panel (c) shows held-out policies with the oracle action.

(a) Scene-selective uploads

(b) Request value by scene

F+S compact

Local before Edge residual

0.175

0.125 0.100 0.075

5.7%

0.050 0.025

8

Train median residual (m)

Upload fraction

0.150

14.0%

Mean risk / residual (m)

All-5 compact

(c) Compact-mode risk prior

6 4 2

5

* = selected mode

Front

H4

F+S

All-5

4 *

3

* 2

*

1

0.5%

0.000

0 Cor.

Int.

Rep.

0 Cor.

Int.

Rep.

Cor.

Int.

Rep.

Fig. 12. View-mode behavior under the budget-constrained held-out route. Panel (a) reports upload fractions by compact view mode, panel (b) compares pre-request risk with returned residuals, and panel (c) shows the train-only compact-mode prior.

fidence intervals, VOI-Control obtains 4.13 ± 0.03 m mean and 6.71 ± 0.08 m p95 route error, compared with 4.21 ± 0.03 m and 7.01 ± 0.10 m for uncertainty triggering. Budgeted periodic offloading obtains 5.49±0.03 m mean and 9.73±0.09 m p95 route error, so VOI-Control reduces these two errors by 24.8% and 31.0%. Relative to uncertainty triggering, the reductions are 1.9% and 4.3%. Budgeted periodic matches the request count and compact prefix but fixes the view mode to FS, so this comparison evaluates request timing and view selection jointly. The oracle action reaches 3.82 m mean route error.

Fig. 12 shows how the selected VOI-Control configuration spends its traffic on route segments with nonzero uploads. On the held-out route, VOI-Control sends rare all-five compact requests in corridor scenes and front-plus-sides compact requests in intersection and repetitive scenes. It uploads on 0.5% of corridor steps, 5.7% of intersection steps, and 14.0% of repetitive steps. These requests reduce local risk from roughly 6 to 9 m down to 2 to 4 m, depending on the scene. The compact-mode prior in panel (c) explains the mode choices. Front-only compact has much larger residuals, while frontplus-sides and all-five compact are close to the best compact

12

(c) Representative route

15

10.0

10

7.5

4 2

Error (m)

6

5

VOI-Control Periodic

(d) VOI requests Uncertainty (m)

(b) Tail accuracy

p95 error (m)

Mean error (m)

(a) Held-out accuracy

5.0 2.5

8 6 4 2 0

0

0

l t c c y n om fron iodi iodi int ntro tio r r nd ta o ac Ra ixed d pe Pe cer I-C cle n VO ra F te U e O dg Bu

l t c c y n om fron iodi iodi int ntro tio r r nd ta o ac Ra ixed d pe Pe cer I-C cle n VO ra F te U e O dg Bu

0

2000

4000

0

2000

Distance (m)

4000

Distance (m)

Fig. 13. Budget-constrained request-control dynamics on the held-out route. Panels report held-out mean error, held-out p95 error, a representative error curve, and the corresponding VOI-Control request timing. VOI-Control concentrates uploads around uncertainty growth and improves both mean and tail localization error over budgeted periodic, uncertainty-triggered, fixed-front, and random baselines.

2 1 0

4

0.05 150 100 50

0.04 0.03 0.02

3

FIFO Round-robin EDF MaxWeight Base-DPP VOI-Lyapunov

2

0.01 0

O in FIF -rob d un o R

(d) Edge tail-cost tradeoff

Weighted cost

3

(c) Stale feedback

Mean stale (m)

4

Cost

(b) Edge tail latency Edge p95 latency (ms)

(a) Value-weighted cost

v F P ht ED eig DP uno ep xW Bas Lya I Ma VO

1

0.00 O in FIF -rob d un o R

t v P gh -DP uno ei e p xW Bas Lya I Ma VO F

O in FIF -rob d un o R

ED

t v P gh -DP uno ei e p xW Bas Lya I Ma VO F

ED

50

100

150

Edge p95 latency (ms)

Fig. 14. Value-aware scheduling under high congestion (𝑁 = 30, 4 KB/s). Panels report cost and latency metrics for the top-10% high-value requests. TABLE IV High-congestion scheduler comparison for top-10% high-value requests. High congestion: 𝑁 = 30 4 KB/s weak link Scheduler

E2E p95

Edge p95

Stale

Cost

MaxWeight Base-DPP VOI-Lyap.

198 ms 193 ms 88.5 ms

142.4 ms 137.7 ms 32.8 ms

0.022 m 0.022 m 0.012 m

1.87 1.82 1.12

Note: All table metrics, including stale-feedback penalty, are computed over the top-10% high-value requests.

residual in scenes where uploads are selected. E. Request-Control Dynamics Fig. 13 shows the same budget-constrained behavior over time. VOI-Control issues 22 requests and transmits 2.07 KB per route. Under the same budget, fixed-front obtains 5.64 m mean and 8.41 m p95 error, while random obtains 6.43 m and 13.43 m. The other policies and confidence intervals appear in Fig. 11(c). F. Multi-Client Scheduling We obtain the multi-client results in Fig. 14 and Table IV by replaying VOI-Control request events from the held-out route on a single shared edge server. The replay covers low, medium, and high congestion, corresponding to 𝑁 = 10, 𝑁 = 20, and 𝑁 = 30 clients with weak-link profiles of 12, 8,

and 4 KB/s. The figure and table focus on high congestion, where service order matters most. Each client count uses five scheduler seeds that perturb route phase, mission urgency, deadlines, and service-time jitter while preserving trainedmodel payloads and localization residuals. The weighted cost is b 𝑣 𝑗 𝐷 𝑗 /𝐷 ddl 𝑗 , where 𝐷 𝑗 is the response latency and ddl 𝐷 𝑗 is the request deadline. The stale-feedback penalty is stale is the observation-tostale 𝑟 drift 𝑗 𝐷 𝑗 /1000 meters, where 𝐷 𝑗 correction delay in milliseconds. The top-10% weighted cost is computed over the largest decile of b 𝑣 𝑗. Each application-layer request is 0.094 to 0.102 KB and contains the latent prefix, a 20-byte representation header, and one 4-byte identifier per selected view. The corresponding p95 transmission times are 8.3, 12.7, and 25.4 ms. For each request, the replay adds the Jetson O-VIB-32 encoding profile and weak-link transmission time to edge queueing and service latency before computing the end-to-end p95. Base-DPP uses 𝑉 = 8 in (26), while VOI-Lyapunov applies (29). We also compare first-in-first-out (FIFO), earliestdeadline-first (EDF), and MaxWeight, which sets 𝑉 = 0 in (26). For the top-10% high-value requests, VOI-Lyapunov reduces the weighted cost to 1.12, compared with 1.82 for Base-DPP and 1.87 for MaxWeight. It lowers edge-side p95 latency from 137.7 ms for Base-DPP to 32.8 ms and endto-end p95 latency to 88.5 ms; the corresponding BaseDPP and MaxWeight end-to-end values are 193.0 ms and 197.8 ms. Across all high-congestion requests, VOI-Lyapunov gives 185.7 ms edge and 241.1 ms end-to-end p95, com-

13

Real Experimental Platform for Cooperative Aerial–Ground Localization Aerial Platform with Edge Computation (UAV)

TABLE V Held-out real-scene localization for representative O-VIB prefixes.

Motor

Commands

RGB Frames Onboard Computer

Front / Downward RGB Cameras

Raspberry Pi 4B

Wi-Fi

Flight Control Board

Unmanned Aerial Vehicle (UAV)

Localization

Image Streams

Results

(RGB Frames)

Shared Infrastructure Real Experiments

Edge Server Position Feedback

Plat.

Representation

UAV UAV UAV UAV UGV UGV UGV UGV

CLIP O-VIB-8 O-VIB-32 O-VIB-128 CLIP O-VIB-8 O-VIB-32 O-VIB-128

KB

Mean (m)

p90 (m)

[email protected]

10.000 0.051 0.145 0.520 8.000 0.051 0.145 0.520

0.407 0.286 0.293 0.274 0.346 0.305 0.296 0.293

0.758 0.458 0.447 0.459 0.396 0.519 0.501 0.481

86.2% 89.7% 91.4% 93.1% 98.6% 88.9% 88.9% 90.3%

Note: Reference samples train O-VIB, form the localization database, and calibrate the frozen-CLIP view-fusion weights; metrics use the disjoint query split. [email protected] denotes recall at 0.5 m position error. O-VIB payload includes a 20-byte representation header.

Motion Capture System Jetson Orin NX Visual Localization for UAV Status / Logs

Ground Station

Position feedback

Visual Localization for UGV

Ground Platform with Onboard Computation (UGV) Motor

RGB-D

Frames

Onboard Computer

Jetson Orin NX

Image stream

Commands

Front RGB-D Camera

STM32 Motion Control Board

USB: Commands / State

Localization

Unmanned Ground Vehicle (UGV)

Position

Status / logs

(Wi-Fi) (bidirectional) localization Fig. 15. Real-world UAV/UGV testbed feedback and edge-assisted results (from motion capture) pipeline. The platforms transmit visual inputs over Wi-Fi, while Qualisys provides pose ground truth.

pared with 232.5/288.0 ms for Base-DPP and 217.8/273.2 ms for MaxWeight. FIFO and EDF give 93.2/148.8 ms and 106.7/162.5 ms for edge/end-to-end p95, respectively, because they prioritize latency rather than request value. The comparison therefore separates value-aware service from latency-first scheduling.

G. Real-Scene Multi-Platform Evaluation We further evaluate O-VIB on two indoor datasets collected with the UAV and UGV testbed in Fig. 15. The UAV set contains 76 reference and 58 query locations with five posealigned views. The UGV set contains 78 reference and 72 query locations with four views. The horizontal directions were captured sequentially near each target pose rather than by a synchronized camera array, while the UAV front and downward views were captured together. Qualisys supplies six-degree-of-freedom pose ground truth while the platforms send observations to the edge server for localization. Reference samples form the localization database and are split spatially into encoder-training and validation subsets. The query samples remain disjoint and are used once for reporting. The mixing weight 𝜂 between O-VIB regression and retrieval is selected on the reference validation subset for each prefix. The CLIP baseline uses pure nearest-neighbor retrieval with nonnegative view-fusion weights calibrated on five spatial

folds of the reference split. The weights are renormalized over the views kept by each view mode. Fig. 16(b) first isolates the contribution of additional views without learned compression. Relative to front-only retrieval, the calibrated H4 mode reduces the mean error from 0.626 to 0.415 m for the UAV, and the calibrated All mode further reduces it to 0.407 m. For the UGV, the corresponding reduction is from 0.622 m to 0.346 m with all four views. The mean distance from each query to its nearest reference location is 0.322 m for the UAV and 0.316 m for the UGV. This spacing lower-bounds the mean error of pure retrieval, which always returns a reference location, while direct regression can interpolate below it. Table V and Figs. 16(c) and (d) report the held-out inference results. At the rich prefix 𝑘 = 32, the deployed OVIB pipeline obtains 0.293 m UAV and 0.296 m UGV mean error using 0.145 KB/request, compared with 0.407 m at 10 KB and 0.346 m at 8 KB for validation-calibrated all-view CLIP retrieval. Relative to the CLIP retrieval baseline, the deployed O-VIB pipeline lowers mean error by 28.0% and 14.4% while reducing descriptor traffic by 98.6% and 98.2%, respectively. The largest prefix further lowers the mean errors to 0.274 and 0.293 m. For the UGV, coordinate regression removes grid quantization for most queries but introduces several larger residuals, which explains why O-VIB has a higher p90 error and a lower recall at 0.5 m, denoted [email protected], than CLIP in Table V. Latency is characterized by the separate Jetson measurements rather than by this test. VIII. Conclusion In this paper, we have presented a task-oriented communication framework that combines scalable multi-view O-VIB encoding, VOI-guided request and mode control, and valueaware edge scheduling. One nested encoder supports runtime prefix adaptation from 8 to 128 dimensions, and its 𝑘 = 35 Town05 operating point reaches 9.18 m mean error with a 0.156 KB semantic representation per request. Under the same traffic budget, VOI-Control reaches 4.13 m mean route error, compared with 5.49 m for budgeted periodic offloading and 4.21 m for uncertainty triggering, and reduces the p95 error of budgeted periodic offloading by 31.0%. Under high congestion, bounded waiting-age and deadline shaping can lower the edge-side p95 latency for the top-10% high-value

14

(b) View contribution

0 −1 −2

UAV ref. UAV query

−2

0

UGV ref. UGV query

2

Ground-truth x (m)

UGV

0.6 0.4 0.2

F

F+S

H4

All

UAV O-VIB

0.45

0.8 0.6 0.4

UAV CLIP-f32 UAV k=32

0.2

UGV CLIP-f32 UGV k=32

0.0

0.0

(d) Real-scene rate-error

0.25

0.50

0.75

Position error (m)

Mean error (m)

1

UAV

(c) Held-out error CDF 1.0

Query fraction

2

0.8

Mean error (m)

Ground-truth y (m)

(a) Metrical sampling

UGV O-VIB

0.40 0.35 0.30

10−1

100

101

Payload (KB/request)

Fig. 16. Real-scene localization with motion-capture ground truth. (a) Reference and disjoint query locations: circles denote UAV, squares denote UGV, and open/filled markers denote reference/query samples. (b) Mean validation-calibrated frozen-CLIP retrieval error for front only (F), front plus sides (F+S), all four horizontal directions (H4), and all available views (All). Caps are 95% confidence intervals over queries. (c) Held-out query-error CDFs for all-view CLIP retrieval and O-VIB at 𝑘 = 32. (d) Mean query error versus the actual serialized representation size. Crosses denote the uncompressed float32 CLIP descriptors.

requests from 137.7 ms for Base-DPP to 32.8 ms. The indoor motion-capture evaluation further shows that a 0.145 KB semantic representation reaches about 0.29 m mean error on both UAV and UGV observations. These results show that semantic rate, request timing, and edge service order can be coordinated through task value. Future work will extend the framework to online context adaptation and live multi-robot closed-loop deployment. References [1] Z. Fang, Z. Liu, J. Wang, S. Hu, Y. Guo, Y. Deng, and Y. Fang, “Taskoriented communications for visual navigation with edge-aerial collaboration in low altitude economy,” in Proc. IEEE Global Communications Conference (GLOBECOM), Taipei, Taiwan, Dec. 2025, pp. 1059–1064. [2] T. Zhang, L. Zhang, Y. Chen, and Y. Zhou, “CVIDS: A collaborative localization and dense mapping framework for multi-agent based visualinertial SLAM,” IEEE Trans. Image Process., vol. 31, pp. 6562–6576, Oct. 2022. [3] K. Gao, H. Wang, H. Lv, and W. Liu, “Localization-oriented digital twinning in 6G: A new indoor-positioning paradigm and proof-ofconcept,” IEEE Trans. Wireless Commun., vol. 23, no. 8, pp. 10 473– 10 486, Aug. 2024. [4] Y. Wang, Y. Wang, I. W.-H. Ho, W. Sheng, and L. Chen, “Pavement marking incorporated with binary code for accurate localization of autonomous vehicles,” IEEE Trans. Intell. Transp. Syst., vol. 23, no. 11, pp. 22 290–22 300, Nov. 2022. [5] N. Keetha, A. Mishra, J. Karhade, K. M. Jatavallabhula, S. Scherer, M. Krishna, and S. Garg, “AnyLoc: Towards universal visual place recognition,” IEEE Robot. Autom. Lett., vol. 9, no. 2, pp. 1286–1293, Feb. 2024. [6] A. Ali-Bey, B. Chaib-Draa, and P. Giguère, “MixVPR: Feature mixing for visual place recognition,” in Proc. IEEE/CVF Winter Conf. Appl. Comput. Vis. (WACV). IEEE, Jan. 2023, pp. 2997–3006. [7] G. Berton, G. Trivigno, B. Caputo, and C. Masone, “Eigenplaces: Training viewpoint robust models for visual place recognition,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV). IEEE, Oct. 2023, pp. 11 046–11 056. [8] Y. Ma, Y. Tao, Z. Fang, S. Hu, X. Chen, and Y. Fang, “Learning by moving closer: Adapting vision models on movable agents without manual labeling,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), 2026. [9] H. Cao, J. Xu, Z. Yang, L. Shangguan, J. Zhang, X. He, and Y. Liu, “Scaling up edge-assisted real-time collaborative visual SLAM applications,” IEEE/ACM Trans. Netw., vol. 32, no. 2, pp. 1823–1838, Apr. 2024. [10] Z. Yuan, S. Rawlekar, S. Garg, E. Erkip, and Y. Wang, “Split computing with scalable feature compression for visual analytics on the edge,” IEEE Trans. Multimedia, vol. 26, pp. 10 121–10 133, May 2024.

[11] A. Furutanpey, P. Raith, and S. Dustdar, “Frankensplit: Efficient neural feature compression with shallow variational bottleneck injection for mobile edge computing,” IEEE Trans. Mobile Comput., vol. 23, no. 12, pp. 10 770–10 786, 2024. [12] G. K. Wallace, “The JPEG still picture compression standard,” IEEE Trans. Consum. Electron., vol. 38, no. 1, pp. xviii–xxxiv, Feb. 1992. [13] Advanced Video Coding for Generic Audiovisual Services, ITU-T Std. Recommendation H.264 and ISO/IEC 14 496-10, 2003. [14] F. Bossen, B. Bross, K. Suhring, and D. Flynn, “HEVC complexity and implementation analysis,” IEEE Trans. Circuits Syst. Video Technol., vol. 22, no. 12, pp. 1685–1696, Oct. 2012. [15] B. Li, J. Shi, W. Li, and H. Li, “WebP-JPEG transcoding detection by spotting re-compression artifacts with CNN-ViT for processing dualdomain features,” IEEE Trans. Circuits Syst. Video Technol., vol. 34, no. 12, pp. 12 535–12 549, Dec. 2024. [16] J. Shao, X. Zhang, and J. Zhang, “Task-oriented communication for edge video analytics,” IEEE Trans. Wireless Commun., vol. 23, no. 5, pp. 4141–4154, May 2024. [17] H. Wei, W. Ni, W. Xu, F. Wang, D. Niyato, and P. Zhang, “Federated semantic learning driven by information bottleneck for task-oriented communications,” IEEE Commun. Lett., vol. 27, no. 10, pp. 2652–2656, Oct. 2023. [18] Z. Fang, S. Hu, J. Wang, Y. Deng, X. Chen, and Y. Fang, “Prioritized information bottleneck theoretic framework with distributed online learning for edge video analytics,” IEEE/ACM Trans. Netw., Jan. 2025. [19] Z. Fang, J. Wang, Y. Ren, Z. Han, H. V. Poor, and L. Hanzo, “Age of information in energy harvesting aided massive multiple access networks,” IEEE J. Sel. Areas Commun., vol. 40, no. 5, pp. 1441–1456, May 2022. [20] J. Chen, J. Wang, C. Jiang, and J. Wang, “Age of incorrect information in semantic communications for NOMA aided XR applications,” IEEE J. Sel. Topics Signal Process., vol. 17, no. 5, pp. 1093–1105, Sep. 2023. [21] J. Wang, L. Bai, Z. Fang, R. Han, J. Wang, and J. Choi, “Age of information based URLLC transmission for UAVs on pylon turn,” IEEE Trans. Veh. Technol., vol. 73, no. 6, pp. 8797–8809, Jun. 2024. [22] H. Feng, J. Wang, Z. Fang, J. Chen, and D.-T. Do, “Evaluating AoIcentric HARQ protocols for UAV networks,” IEEE Trans. Commun., vol. 72, no. 1, pp. 288–301, Jan. 2024. [23] T. Li, J. Sun, Y. Liu, X. Zhang, D. Zhu, Z. Guo, and L. Geng, “ESMO: Joint frame scheduling and model caching for edge video analytics,” IEEE Trans. Parallel Distrib. Syst., vol. 34, no. 8, pp. 2295–2310, Aug. 2023. [24] H. Zhang, Y. Tang, A. Khandelwal, and I. Stoica, “SHEPHERD: Serving DNNs in the wild,” in Proc. 20th USENIX Symp. Networked Syst. Design Implement. (NSDI). Boston, MA, USA: USENIX Association, 2023, pp. 787–808. [25] Z. Li, L. Zheng, Y. Zhong, V. Liu, Y. Sheng, X. Jin, Y. Huang, Z. Chen, H. Zhang, J. E. Gonzalez, and I. Stoica, “Alpaserve: Statistical multiplexing with model parallelism for deep learning serving,” in Proc. 17th USENIX Symp. Operating Syst. Design Implement. (OSDI). Boston, MA, USA: USENIX Association, 2023, pp. 663–679.

15

[26] Y. Shao, Q. Cao, and D. Gündüz, “A theory of semantic communication,” IEEE Trans. Mobile Comput., vol. 23, no. 12, pp. 12 211–12 228, Dec. 2024. [27] P. Zhang, W. Xu, Y. Liu, X. Qin, K. Niu, S. Cui, G. Shi, Z. Qin, X. Xu, F. Wang, Y. Meng, C. Dong, J. Dai, Q. Yang, Y. Sun, D. Gao, H. Gao, S. Han, and X. Song, “Intellicise wireless networks from semantic communications: A survey, research issues, and challenges,” IEEE Commun. Surv. Tutor., vol. 27, no. 3, pp. 2051–2084, Jun. 2025. [28] Z. Meng, K. Chen, Y. Diao, C. She, G. Zhao, M. A. Imran, and B. Vucetic, “Task-oriented cross-system design for timely and accurate modeling in the metaverse,” IEEE J. Sel. Areas Commun., vol. 42, no. 3, pp. 752–766, Mar. 2024. [29] J. Wang, H. Du, Z. Tian, D. Niyato, J. Kang, and X. Shen, “Semanticaware sensing information transmission for metaverse: A contest theoretic approach,” IEEE Trans. Wireless Commun., vol. 22, no. 8, pp. 5214–5228, Aug. 2023. [30] M. Zhang, F. Wang, and J. Liu, “CASVA: Configuration-adaptive streaming for live video analytics,” in Proc. IEEE INFOCOM. IEEE, May 2022, pp. 2168–2177. [31] D. Wu, D. Zhang, M. Zhang, R. Zhang, F. Wang, and S. Cui, “ILCAS: Imitation learning-based configuration-adaptive streaming for live video analytics with cross-camera collaboration,” IEEE Trans. Mobile Comput., vol. 23, no. 6, pp. 6743–6757, Jun. 2024. [32] O. Rippel, M. Gelbart, and R. P. Adams, “Learning ordered representations with nested dropout,” in Proc. 31st Int. Conf. Mach. Learn. (ICML), vol. 32. PMLR, 2014, pp. 1746–1754. [33] A. Kusupati, G. Bhatt, A. Rege, M. Wallingford, A. Sinha, V. Ramanujan, W. Howard-Snyder, K. Chen, S. Kakade, P. Jain, and A. Farhadi, “Matryoshka representation learning,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 35, 2022, pp. 30 233–30 249. [34] S. Izquierdo and J. Civera, “Optimal transport aggregation for visual place recognition,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR). IEEE, Jun. 2024, pp. 17 658–17 668. [35] P. Lindenberger, P.-E. Sarlin, and M. Pollefeys, “Lightglue: Local feature matching at light speed,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV). IEEE, Oct. 2023, pp. 17 581–17 592. [36] R. Pautrat, I. Suárez, Y. Yu, M. Pollefeys, and V. Larsson, “Gluestick: Robust image matching by sticking points and lines together,” in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV). IEEE, Oct. 2023, pp. 9672–9682. [37] Y. Tao, S. Hu, H. An, Z. Fang, H. Cao, and Y. Fang, “Learning mutual view information graph for adaptive adversarial collaborative perception,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2026. [38] Y. Tao, Y. Guo, Z. Fang, H. An, and Y. Fang, “FRUC: Feedforward dynamic scene reconstruction from uncalibrated collaborative driving views,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), 2026. [39] M. R. Abedi, N. Mokari, M. R. Javan, H. Saeedi, E. A. Jorswieck, and H. Yanikomeroglu, “Safety-aware age of information (S-AoI) for collision risk minimization in cell-free mMIMO platooning networks,” IEEE Trans. Netw. Service Manag., vol. 21, no. 3, pp. 3035–3053, Jun. 2024. [40] S. Wang, S. Bi, and Y.-J. A. Zhang, “Edge video analytics with adaptive information gathering: A deep reinforcement learning approach,” IEEE Trans. Wireless Commun., vol. 22, no. 9, pp. 5800–5813, Sep. 2023. [41] M. Tang, C. Feng, G. Min, K. Yang, Y. Huang, and T. Q. S. Quek, “Toward robust UAV swarm coordination in low-altitude networks: Online learning control under system and channel uncertainty,” IEEE Trans. Cogn. Commun. Netw., vol. 12, pp. 7017–7033, 2026. [42] R. Bhardwaj, Z. Xia, G. Ananthanarayanan, J. Jiang, Y. Shu, N. Karianakis, K. Hsieh, P. Bahl, and I. Stoica, “Ekya: Continuous learning of video analytics models on edge compute servers,” in Proc. 19th USENIX Symp. Networked Syst. Design Implement. (NSDI). Renton, WA, USA: USENIX Association, 2022, pp. 119–135. [43] A. Padmanabhan, N. Agarwal, A. Iyer, G. Ananthanarayanan, Y. Shu, N. Karianakis, G. H. Xu, and R. Netravali, “Gemel: Model merging for memory-efficient, real-time video analytics at the edge,” in Proc. 20th USENIX Symp. Networked Syst. Design Implement. (NSDI). Boston, MA, USA: USENIX Association, 2023, pp. 973–994. [44] M. Khani, G. Ananthanarayanan, K. Hsieh, J. Jiang, R. Netravali, Y. Shu, M. Alizadeh, and V. Bahl, “RECL: Responsive resourceefficient continuous learning for video analytics,” in Proc. 20th USENIX Symp. Networked Syst. Design Implement. (NSDI). Boston, MA, USA: USENIX Association, 2023, pp. 917–932. [45] Y. Nan, S. Jiang, and M. Li, “Large-scale video analytics with cloudedge collaborative continuous learning,” ACM Trans. Sensor Netw., vol. 20, no. 1, pp. 1–23, Oct. 2023.

[46] T. Liu, B. Li, W. Wang, Y. Cheng, S. Wang, T. He, and Y. Liu, “FaaSLearner: Resource-efficient edge video analytics via correlationaware multi-model continual learning,” IEEE Trans. Mobile Comput., pp. 1–15, 2026. [47] M. Han, H. Zhang, R. Chen, and H. Chen, “Microsecond-scale preemption for concurrent GPU-accelerated DNN inferences,” in Proc. 16th USENIX Symp. Operating Syst. Design Implement. (OSDI). Carlsbad, CA, USA: USENIX Association, 2022, pp. 539–558. [48] G. Qu, Q. Chen, W. Wei, Z. Lin, X. Chen, and K. Huang, “Mobile edge intelligence for large language models: A contemporary survey,” IEEE Commun. Surv. Tutor., Jan. 2025. [49] X. Hou, J. Wang, J. Du, C. Jiang, and Y. Ren, “Distributed machine learning for autonomous agent swarm: A survey,” IEEE Commun. Surv. Tutor., vol. 28, pp. 1597–1636, 2026. [50] X. Hou, J. Wang, J. Du, C. Jiang, Y. Ren, and D. Niyato, “Lightweight federated learning over wireless edge networks,” IEEE Trans. Mobile Comput., vol. 25, no. 1, pp. 300–312, 2026. [51] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” in Proc. 38th Int. Conf. Mach. Learn. (ICML), vol. 139. PMLR, 2021, pp. 8748–8763. [52] N. Tishby, F. C. Pereira, and W. Bialek, “The information bottleneck method,” in Proc. 37th Annu. Allerton Conf. Commun., Control, Comput., 1999, pp. 368–377. [53] A. Alemi, I. Fischer, J. Dillon, and K. Murphy, “Deep variational information bottleneck,” in Int. Conf. Learn. Represent. (ICLR), 2017. [54] D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” 2013. [55] D. Molchanov, A. Ashukha, and D. Vetrov, “Variational dropout sparsifies deep neural networks,” in Proc. 34th Int. Conf. Mach. Learn. (ICML), Sydney, Australia, July 2017, pp. 2498–2507. [56] D. Barber and F. Agakov, “The IM algorithm: A variational approach to information maximization,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 16, 2003. [57] M. J. Neely, Stochastic Network Optimization with Application to Communication and Queueing Systems. Morgan & Claypool, 2010.

Record · ID 1108651 · SHA-256 5ebb68a521e96110
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.