1
TWIST: Closed-Loop token Synchronization for Application-Aware Wireless Digital Twins
arXiv:2605.27205v1 [eess.IV] 26 May 2026
Sige Liu, Member, IEEE, and Kezhi Wang, Senior Member, IEEE Abstract—Wireless digital twins require repeated synchronization between a time-evolving physical scene and its digital counterpart under limited and time-varying communication resources. For perception-centric twins, pixel-domain transmission or uniformly protected bitstreams can be mismatched to the semantic state consumed by twin-side applications. This paper proposes TWIST, a closed-loop token synchronization framework for application-aware wireless digital twins. TWIST represents each physical observation as a token and synchronizes this state over a wireless link, rather than optimizing visual reconstruction. Token positions are grouped by task relevance and protected through mode-conditioned unequal error protection under low, medium-, and high-synchronization modes. At the twin side, decoding confidence converts unreliable hard token decisions into erasures, which are restored by a completion model before updating the semantic twin state. The recovered state supports traffic-state inference and generates compact feedback statistics, including channel quality, receiver uncertainty, semantic drift, and application priority, for subsequent mode adaptation. Experiments on a dynamic road-scene digital-twin scenario show that TWIST improves traffic-state inference and semantic twin-state synchronization compared with fixed-mode and channel-only adaptation strategies, while reducing the average synchronization cost relative to always-high transmission. Index Terms—Digital twin, semantic communication, token communication, closed-loop synchronization.
I. I NTRODUCTION Digital twins are evolving from static digital replicas into closed-loop, application-aware systems that continuously synchronize with the physical world and support monitoring, prediction, and control. In wireless systems, this evolution makes communication a central bottleneck: the digital counterpart must be updated under fading, bandwidth limitations, latency constraints, and time-varying application requirements. Digital twin networks have been surveyed as a promising paradigm for coupling physical and virtual systems through sensing, communication, computation, and feedback [1]. Wireless digital twins further emphasize how communication systems can both enable and benefit from such physical–virtual coupling [2]. Recent studies have investigated real-time wireless digital twins, digital-twin channels, and DT-assisted beam or CSI prediction [3], [4]. However, they mainly focus on channel and network states, or control decisions, rather than on how a perception-centric twin should synchronize the semantic representation consumed by its downstream application. This work is supported in part by Eureka CELTIC-NEXT 5G4PHealth/Innovate UK project (10093679), UKRI under the Horizon Europe funding guarantee (EP/Y03743X/1), as part of the European Commission MSCA HarmonicAI project (101131117) and Royal Society project (IEC-NSFC-211264). K. Wang would like to acknowledge the support in part by the Royal Society Industry Fellow scheme (IF\R2\23200104). (Corresponding author: Kezhi Wang). S. Liu and K. Wang are with the Department of Computer Science, Brunel University London, Uxbridge UB8 3PH, U.K. (e-mail:[email protected]; [email protected]).
This distinction becomes important as modern perception and generative-model pipelines increasingly operate on tokenized representations. Learned discrete representations, such as VQstyle tokenizers, map visual observations into compact codebook indices [5]. Tokenized visual representations have also been widely used in generative and masked-token modeling, where downstream models operate on token grids rather than raw pixels [6], [7]. For a perception-centric digital twin, the object that needs to be synchronized is therefore not necessarily a pixel-level reconstruction or a uniformly protected bitstream. It can instead be a token: a compact representation that carries the information needed by the twin-side application. This observation raises a central question: how should a wireless digital twin synchronize tokens under limited and time-varying communication resources? The problem differs from conventional image delivery in several aspects. First, token positions are not equally important for the twin-side application, so uniformly protecting all tokens can waste scarce channel uses. Second, in a token-state system, an erroneous hard token update can be more harmful than an explicit erasure, since a wrong token may contaminate the twin state while an erasure can be recovered by a completion prior. Third, synchronization is inherently dynamic: the appropriate resource level depends not only on the wireless channel, but also on the uncertainty of the recovered twin state, the semantic drift of the scene, and possible application-priority signals. Finally, a practical wireless-DT architecture should remain modular and interpretable, rather than relying entirely on a monolithic endto-end learned transceiver. Semantic and task-oriented communication has established that communication objectives should be aligned with meaning or downstream utility rather than exact bit recovery [8], [9]. Task-oriented edge communication further formalizes this principle by optimizing communication with respect to inference performance [10]. In parallel, deep joint source–channel coding has shown that learned image transmission can be effective under bandwidth and channel constraints [11], [12]. More recently, token-centric communication has emerged as a natural direction for large-model-era communication systems. UniToCom treats tokens as communication units from an informationbottleneck perspective [13], while token-domain multiple access considers how tokenized source and modulation codebooks can be shared in multiuser semantic communication [14]. Token-aware semantic-channel coding and modulation further demonstrate that token-domain reliability and modulation can be jointly designed [15]. Despite these advances, existing works do not jointly address token-state synchronization, receiver decisions, and closed-loop mode adaptation in a wireless digital twin. To address the gap, we propose TWIST (Tokenized Wireless Intelligent Synchronization for digital Twins), a closed-loop to-
2
ken synchronization framework for application-aware wireless digital twins. TWIST does not optimize pixel-level reconstruction. Instead, it maintains a synchronized token at the twin side and uses this state both for application-state inference and for subsequent synchronization control. At the physical side, each frame is tokenized and its token positions are grouped according to task relevance. A mode-conditioned unequal protection mechanism then selects group-wise channel protection under low-, medium-, and high-synchronization modes. At the twin side, soft-decoding confidence is used to convert unreliable hard token decisions into erasures, which are restored by a completion model before updating the semantic twin state. The twin further computes compact feedback statistics, including channel quality, erasure-ratio uncertainty, semantic drift, and application priority, to select the next synchronization mode. The main contributions of this paper are summarized as follows: • We formulate wireless digital-twin updating as a closedloop token-state synchronization problem. The synchronized object is a time-evolving token consumed by the twin-side application, rather than a pixel-level reconstruction or a conventional bitstream. • We develop a twin-in-the-loop synchronization architecture and a practical digital realization based on tokenutility grouping, mode-conditioned unequal protection, confidence-aware gating, and completion-assisted recovery. The physical side transmits tokenized scene states, while the twin side updates its semantic state and feeds back the next synchronization mode. • We introduce a closed-loop mode controller that adapts the synchronization mode according to channel quality, twinside uncertainty, semantic drift, and application-priority information. The controller operates on compact feedback statistics and does not require online retraining or heavy control optimization. • We instantiate TWIST on a dynamic road-scene digital twin scenario using sequence data and derived traffic-state labels. The evaluation demonstrates improved traffic-state inference and semantic twin-state synchronization under time-varying wireless conditions, with lower average cost than always-high synchronization. The remainder of this paper is organized as follows. Section II reviews related work on wireless digital twins, semantic and task-oriented communication, token-centric communication, and completion-assisted semantic recovery. Section III introduces the system model and problem formulation. Section IV presents the proposed TWIST framework. Section V provides the design rationale and theoretical characterization. Section VI presents the experimental setup and performance evaluation. Section VII concludes this paper. II. R ELATED W ORK This section reviews wireless digital twins, semantic and task-oriented communication, token-centric communication, and completion-assisted semantic recovery. A. Wireless Digital Twins and Twin-State Synchronization Digital twins maintain virtual representations of physical systems for monitoring, prediction, optimization, and control.
Digital twin networks connect physical systems and virtual replicas through sensing, modeling, communication, computation, and feedback [1]. In wireless systems, DTs can represent wireless entities and support application-aware communication decisions [2]. The 6G DT-network vision highlights real-time simulation and AI-driven network optimization [16], while digital-twin-empowered communications frame wireless networks as both enablers and beneficiaries of DT-based decisions [17]. Recent wireless-DT studies have developed more specific communication-layer realizations. Real-time wireless twins fuse environmental, sensing, and communication data for communication and sensing decisions [3]. Digital twin channel models map radio propagation behavior into virtual channel representations [4]. Digital replicas and ray-tracing environments have been used for beam prediction and massiveMIMO CSI feedback [18]. Generative AI and autonomousvehicle testbeds have also been explored for wireless network twins and mobility-oriented evaluation [19], [20]. These works establish the value of physical–virtual coupling, but they mainly focus on channel replicas, network states, beam management, CSI feedback, or testbed validation. They do not address how a perception-centric twin should synchronize the semantic representation consumed by its application. B. Semantic and Task-Oriented Communication Semantic and task-oriented communication optimize communication for meaning or downstream utility rather than bitlevel recovery. Broad perspectives have emphasized contextand task-aware objectives [8]. DeepSC demonstrated semantic feature transmission for text communication [9], while taskoriented edge communication formulated inference-driven communication through an information-bottleneck principle [10]. Related works extended semantic communication to speech, visual question answering, and explainable task-oriented systems [21]–[23]. Deep joint source–channel coding provides another foundation for application-aware wireless transmission. DeepJSCC showed robust neural image transmission under channel impairments [11]. Bandwidth-agile, constellationconstrained, Transformer-based, and semantic-oriented variants further improve adaptability under bandwidth, modulation, or semantic constraints [12], [24]–[26]. Most of these methods, however, consider one-shot transmission or fixed transmitter– receiver mappings. Wireless digital twins require repeated state synchronization and online feedback-driven mode adaptation. C. Token-Centric Communication Foundation-model pipelines increasingly operate on tokenized representations, making token sequences or token grids natural communication objects. Recent work has argued for token-domain semantic information theory [27]. UniToCom studies token communication from an information-bottleneck perspective [13]. ToDMA extends token-domain design to semantic multiple access [14], while token-aware semanticchannel coding and modulation maps token representations into channel-compatible symbols [15]. Selective-token and largemodel-assisted mechanisms have also been studied for multimodal and communication-efficient token transmission [28], [29]. These works show that token-level representations are
3
(a) Conventional bit-centric DT system
(b) Proposed TWIST: closed-loop token-centric DT framework Fig. 1. Conventional bit-centric DT versus the proposed token-centric TWIST.
important communication objects, but they mainly target token transmission, token-domain access, or large-model-assisted communication efficiency. They do not formulate the digital twin as a time-evolving token, nor do they study how token synchronization should be controlled over time by twin-side feedback.
D. Completion-Assisted Semantic Recovery Completion and generative priors provide natural mechanisms for recovering missing semantic content. VQ-VAE introduced learned discrete latent representations [5]. Taming Transformers and MaskGIT showed that tokenized image representations can support high-resolution generation and masked-token recovery [6], [7]. Latent diffusion and generative semantic communication further demonstrate the value of generative priors in compressed latent or bandwidth-constrained settings [30], [31]. For token-state synchronization, completion priors change the receiver-side decision problem: a low-confidence token can be converted into an erasure and recovered from context instead of being forced into a hard decision. Existing completion and generative communication studies typically do not treat this erasure decision as part of closed-loop digital-twin synchronization. TWIST integrates confidence-aware gating with completionassisted recovery and uses the erasure ratio as an online uncertainty statistic for feedback-controlled synchronization. TWIST is positioned as a wireless-DT token-state synchronization framework with feedback-controlled synchronization modes. The synchronized object is not a pixel-domain image, channel replica, or generic feature vector, but a token maintained by the digital twin. Token utility guides modeconditioned unequal protection, receiver confidence determines whether hard token updates are accepted or converted into erasures, and the recovered twin state drives both application-state inference and subsequent mode adaptation. This distinguishes TWIST from one-shot token transmission, reconstruction-
oriented semantic communication, and wireless-DT works focused primarily on channel or network-state replication. III. S YSTEM M ODEL AND P ROBLEM F ORMULATION We consider a perception-centric wireless digital twin in which a physical-side sensing device observes a time-evolving scene and synchronizes its semantic state to a twin-side server through a wireless link. Time is indexed by t = 1, 2, . . . , T . At the beginning of slot t, the physical side uses the synchronization mode at selected from previous twin-side feedback. It then observes an image frame xt ∈ RH×W ×C , converts it into a token, and transmits this state to the twin side. The twin side updates its semantic state, performs application-state inference, computes compact feedback statistics, and selects the next synchronization mode at+1 . This causal timing is central to the proposed setting. The mode at is fixed before frame xt is transmitted and recovered. The next mode at+1 may depend on statistics observed after processing frame t, but it must not depend on unavailable future information or on the ground-truth outcome of frame t + 1. Fig. 1 contrasts this token-centric closed-loop synchronization architecture with a conventional bit-centric digital-twin communication pipeline. Throughout this paper, bold lowercase letters denote vectors or stacked token/signal representations, bold uppercase letters denote matrices, and calligraphic letters denote sets. Scalars are written in standard italic form. A. Physical Observation and token A discrete tokenizer T (·) maps the physical observation xt to a token sequence tt = [tt,1 , tt,2 , . . . , tt,L ]T ,
tt,i ∈ K ≜ {1, 2, . . . , K}, (1)
where L is the token sequence length and K is the tokenizer codebook size. The sequence tt is the physical token at time t.
4
Fig. 2. Offline preparation and online deployment workflow of TWIST.
Let E ∈ RK×D denote the token embedding table. The embedding of token position i is et,i = E[tt,i , :] ∈ RD ,
C. Wireless Channel Model We consider a flat block-fading complex baseband channel rt = ht st + wt ,
(2)
(6)
(at )
and the corresponding embedding sequence is Zt = [et,1 , et,2 , . . . , et,L ]T ∈ RL×D .
(3)
In TWIST, the object to be synchronized is tt , rather than the raw image xt or a pixel-domain reconstruction. B. Synchronization Modes and Mode-Conditioned Transmission To support closed-loop adaptation, TWIST uses a finite synchronization-mode set A ≜ {low, med, high},
(4)
where “med” denotes the medium synchronization mode. Each mode a ∈ A is associated with three quantities: a per-frame communication budget N (a) in channel uses; (a) (a) (a) • a group-wise protection profile π = [π1 , . . . , πG ]; (a) • a group-wise confidence-threshold profile τ = (a) (a) [τ1 , . . . , τG ]. •
Here G is the number of token utility groups, and g(i) ∈ {1, . . . , G} maps token position i to its group. The construction of g(i), π (a) , and τ (a) is described in Section IV. Given tt and the selected mode at , the physical side generates a transmitted block (at ) st = ftx tt ; π (at ) , st ∈ CN , (5) where ftx (·) includes token-to-bit mapping, channel coding, digital modulation, and mode-conditioned group-wise unequal protection. The budget N (at ) determines the number of channel uses allocated to frame t.
where rt ∈ CN is the received block, ht ∈ C is the channel coefficient, and wt ∼ CN (0, σt2 I) is circularly symmetric complex Gaussian noise. The channel coefficient is constant over one frame-level transmission block. The receiver is assumed to have the channel-state information required for coherent demodulation and decoding. In the experiments, (6) is instantiated as AWGN or Rayleigh block fading. D. Twin-Side Token Recovery and State Update From rt , the twin side performs coherent demodulation and soft decoding to obtain token-level posterior information. For token position i, let pt,i (k) ≜ Pr(Θt,i = k | rt , ht ),
k ∈ K,
(7)
where Θt,i denotes the random source token at position i. The hard token estimate and confidence score are t̂t,i = arg max pt,i (k), k∈K
ct,i = max pt,i (k). k∈K
(8)
To avoid directly accepting unreliable hard token updates, the receiver may output an erasure symbol ⊥. Let K⊥ ≜ K ∪ {⊥}. Under mode at , the gated token at position i is ( (a ) t̂t,i , ct,i ≥ τg(i)t , t̃t,i = (9) (a ) ⊥, ct,i < τg(i)t . The gated token sequence is L t̃t = [t̃t,1 , . . . , t̃t,L ]T ∈ K⊥ .
(10)
A completion model restores the erased positions: t̄t = fcomp (t̃t ),
t̄t ∈ KL .
(11)
The synchronized semantic state maintained by the twin is then tDT ≜ t̄t . t
(12)
5
Fig. 3. Relation between intra-frame utility-aware token grouping and inter-frame synchronization-mode adaptation.
Thus, the digital twin is updated in the token domain after confidence-aware gating and completion.
E. Twin-Side Application-State Inference The twin-side application head operates on the embedding sequence induced by tDT t . Let DT T L×D Z̄t = [E[tDT . t,1 , :], . . . , E[tt,L , :]] ∈ R
(13)
The twin predicts the current application state as
G. Problem Formulation The goal is not to reconstruct a visually faithful image, but to maintain a semantic twin state that is useful for downstream application inference while respecting communicationcost constraints. This leads to three coupled design objectives: application-state inference quality, semantic twin-state synchronization quality, and communication efficiency. Let Lapp (ŷt , yt ) denote the application-state loss, and let Lsync (tDT t , tt ) denote a token-state mismatch loss. We write the system-level closed-loop design objective as " T X min E ω(qt )Lapp (ŷt , yt ) Ω
ŷt = fstate (Z̄t ),
(14)
t=1
where fstate (·) is the twin-side application-state inference head. In the road-scene instantiation, yt contains derived traffic-state labels, such as vehicle presence, pedestrian/cyclist presence, and traffic-density level.
F. Closed-Loop Synchronization Feedback After processing frame t, the twin computes compact feedback statistics for the next synchronization decision. Let γt denote a channel-quality indicator, ρt a receiver-side uncertainty statistic, dt a semantic drift statistic, and qt an applicationpriority input. A general closed-loop synchronization policy is at+1 = Πsync (γt , ρt , dt , qt ),
(15)
where Πsync (·) maps communication and twin-state statistics to a mode in A. The variable qt is modeled as an exogenous priority input. Depending on the deployment setting, it may be provided by an upper-layer event detector, a safety policy, or a prior service requirement. When qt is derived from labels in evaluation, the corresponding result should be interpreted as an oracle-priority or exogenous-priority analysis rather than as a claim that the current ground-truth urgency is available online.
+ αLsync (tDT t , tt ) + βC(at )
(16)
# ,
where C(at ) is the communication cost of mode at , ω(qt ) is a priority-dependent application weight, and α, β ≥ 0 balance semantic synchronization quality and communication cost. The expectation is over the physical scene sequence, wireless channel realizations, and possible stochastic components in the receiver pipeline. The objective in (16) is a system-level formulation. It is not intended to be solved as a single monolithic end-to-end neural optimization problem. In TWIST, the tokenizer, completion model, and application-state head are trained offline; the communication-specific design is realized through tokenutility estimation, utility grouping, mode-conditioned protection, confidence-threshold calibration, and lightweight closedloop mode selection. The detailed realization is presented next. IV. C LOSED -L OOP TOKEN S YNCHRONIZATION F RAMEWORK This section presents the proposed TWIST framework. As illustrated in Fig. 2, TWIST separates offline preparation from online deployment. Offline preparation constructs reusable artifacts, including the token utility profile, group map, modeconditioned protection profiles, and mode-conditioned confidence thresholds. Online deployment then performs causal
6
physical-to-twin synchronization, updates the semantic twin state, and feeds back the next synchronization mode. Fig. 3 illustrates the relation between intra-frame utility grouping and inter-frame mode adaptation: the group map is shared across frames, whereas the current synchronization mode selects the protection and threshold profiles used in the current slot. A. Offline Artifact Preparation The offline stage prepares the components required for low-complexity online synchronization. These components are computed from training and calibration sets and reused during deployment. 1) Token utility estimation: Token positions are not equally important for twin-side application-state inference. For a training sample (x, y), let t = T (x) and ei = E[ti , :]. TWIST uses the gradient-based proxy wigrad (x, y) =
∂Lapp (fstate (Z), y) , ∂ei 2
(17)
where Z is the clean embedding sequence induced by t. The deployment-time utility profile is the calibration-set average X 1 w̄i = wigrad (x, y). (18) |Dcal |
Algorithm 1 Offline Preparation of TWIST Artifacts Require: Training set Dtr , calibration set Dcal , tokenizer T (·), embedding table E, mode set A, budgets {N (a) }a∈A , PHY policy set P Ensure: Application head fstate , completion model fcomp , utility profile {w̄i }L i=1 , group map g(i), mode-conditioned profiles {π (a) }a∈A , mode-conditioned thresholds {τ (a) }a∈A 1: Train or load the twin-side application head fstate on Dtr . 2: Train or load the completion model fcomp using randomly masked token sequences from Dtr . 3: Tokenize calibration samples in Dcal . 4: Compute gradient-based token utilities on Dcal using (17). 5: Average token utilities to obtain {w̄i }L i=1 . 6: Quantize token positions into G utility groups and obtain g(i), Lg , and Wg . 7: for each synchronization mode a ∈ A do 8: Select the group-wise protection profile π (a) under budget N (a) . 9: Calibrate the group-wise confidence thresholds τ (a) on Dcal after mode-a transmission, gating, and completion. 10: end for 11: Store all offline artifacts for online deployment.
(x,y)∈Dcal
2) Utility grouping: To avoid per-token protection and threshold design, TWIST quantizes token positions into G utility groups: g(i) ∈ {1, . . . , G}. (19) For group g, define Lg = |{i : g(i) = g}|,
X
Wg =
w̄i .
(20)
performs soft demodulation and decoding, applies the thresh(a ) old τg(i)t to each token confidence, converts unreliable hard decisions into erasures, and invokes the completion model. The current mode at therefore determines both the transmitter-side protection profile and the receiver-side confidence thresholds. After completion, the recovered token sequence is assigned as tDT t , which is used for traffic-state inference and feedback-statistic computation.
i:g(i)=g
Here Lg is the group size and Wg is its aggregate task relevance. 3) Mode-conditioned protection and threshold profiles: For each mode a ∈ A, TWIST prepares a group-wise protection profile (a) (a) π (a) = [π1 , . . . , πG ], (21)
C. Twin-Side Feedback Statistics TWIST uses compact statistics observable at the twin side. The erasure-ratio uncertainty is L
ρt =
(a)
where πg is selected from a finite digital-PHY policy set. Higher synchronization modes have larger budgets and can therefore use stronger group-wise protection. The receiver also stores a mode-conditioned threshold profile (a)
(a)
τ (a) = [τ1 , . . . , τG ],
(22)
1X 1{t̃t,i = ⊥}. L i=1
(23)
It measures the fraction of positions that the receiver does not trust enough to accept as hard updates. Unlike token mismatch metrics, ρt does not require ground-truth tokens and is available online. The semantic drift statistic is
(a)
where τg is the confidence threshold for group g under mode a. These thresholds are calibrated offline after mode-a transmission, gating, and completion. B. Online Physical-to-Twin Synchronization During deployment, TWIST processes the sequence causally. At the beginning of slot t, the mode at has already been selected from feedback generated after slot t − 1. The physical side tokenizes xt and transmits the resulting token using π (at ) . The twin side follows the state-update interface in Section III: it
L
dt =
1X DT 1{tDT t,i ̸= tt−1,i }, L i=1
t ≥ 2,
(24)
with d1 = 0. It measures the frame-to-frame change of the recovered twin state. The channel-quality statistic γt is obtained from the receiver-side channel estimate or nominal SNR state. We also use γ̄t ∈ [0, 1], where a larger value indicates a better channel condition. The priority input qt represents an external or upper-layer indication that the current state should be synchronized more reliably.
7
D. Closed-Loop Mode Selection The controller maps (γt , ρt , dt , qt ) to the next synchronization mode. TWIST uses a validation-calibrated synchronization-risk score ψt = ηρ ρt + ηd dt + ηq qt − ηγ γ̄t ,
(25)
where ηρ , ηd , ηq , ηγ ≥ 0. A larger ψt indicates a higher need for reliable synchronization due to receiver uncertainty, semantic drift, application priority, or poor channel quality. The next mode is selected as high, ψt ≥ θhigh , at+1 = low, ψt ≤ θlow and qt = 0, (26) med, otherwise, where θlow and θhigh are calibrated on the validation split. The condition qt = 0 in the low-mode branch prevents priority frames from being assigned to the weakest synchronization mode. The priority input therefore biases the controller away from low-cost synchronization when the application context indicates higher risk, but it does not unconditionally force highmode transmission. The controller is not intended to be an optimal control law. Its role is to provide a transparent and low-overhead realization of twin-in-the-loop synchronization. TWIST thereby separates two adaptations: intra-frame utility-aware protection through g(i) and π (a) , and inter-frame mode adaptation through twin-side feedback. The expensive operations are performed offline, while deployment only switches among a small set of precomputed profiles. TWIST combines two forms of adaptation. The first is intraframe adaptation: token positions are grouped according to task relevance, and more important groups receive stronger protection under a given synchronization mode. The second is inter-frame adaptation: the twin selects the next synchronization mode according to channel quality, uncertainty, semantic drift, and application priority. This separation keeps the online controller lightweight. The expensive components, including utility profiling, group-wise protection design, completion training, and threshold calibration, are performed offline, while online deployment only switches among a small set of precomputed synchronization profiles. V. D ESIGN R ATIONALE AND T HEORETICAL C HARACTERIZATION This section provides local design characterizations for the main TWIST components. The goal is not to solve the full closed-loop objective in (16) optimally, nor to claim a tight endto-end performance bound for the complete trajectory. Instead, we analyze two local mechanisms that make the modular realization interpretable: utility-aware group-wise protection at the transmitter and confidence-aware acceptance-or-erasure decisions at the twin receiver. Proposition 1 connects token errors to application-loss sensitivity and motivates the utilityweighted protection surrogate in (27). Theorem 1 characterizes the receiver-side accept-or-erase decision under a local Bayesrisk model and motivates the group-wise threshold structure calibrated in (42). The closed-loop controller then selects among these precomputed mode profiles online, while its endto-end behavior is evaluated empirically in Section VI.
A. Mode-Conditioned Utility-Aware Protection For a fixed synchronization mode a ∈ A, TWIST assigns one protection policy to each token group. Let P denote a finite set of digital-PHY protection policies, such as channel-coding rates or equivalent protection levels. For π ∈ P, let c(π) denote (a) the channel-use cost per token. Let ϵg (π) denote the offlineprofiled token error rate of group g under mode a and policy π, measured after channel decoding but before receiver-side gating and completion. Given the group size Lg and aggregate utility Wg in (20), a practical mode-conditioned protection design is min (a) {πg }G g=1
s.t.
G X
(a) Wg ϵ(a) g (πg )
g=1 G X
Lg c(πg(a) ) ≤ N (a) ,
(27)
g=1
πg(a) ∈ P,
g = 1, . . . , G.
The objective in (27) allocates stronger protection to groups whose token errors are expected to have larger impact on the twin-side application loss. The constraint enforces the modespecific communication budget. Since both G and |P| are small in deployment, (27) can be solved offline by exhaustive search, dynamic programming, or a simple discrete search over feasible group-wise profiles. The design in (27) is a tractable surrogate for the full closedloop objective. It does not model receiver-side completion or future controller decisions explicitly. Instead, it provides a deployable intra-frame protection profile for each mode, while inter-frame adaptation is handled by the controller in (26). B. Utility-Weighted Loss Degradation Bound The following bound motivates the utility weighting used in (27). It is not a tight characterization of the full recovery pipeline, which also includes confidence gating, completion, and closed-loop mode selection. Rather, it shows that, before receiver-side recovery, token errors at application-sensitive positions can induce larger application-loss perturbations. This justifies using task-relevance weights in the group-wise protection surrogate. Proposition 1: Utility-weighted upper bound. Let Zt = [et,1 , . . . , et,L ]T denote the clean embedding sequence induced by tt , and let Ẑt = [êt,1 , . . . , êt,L ]T denote the embedding sequence induced by the hard-decoded token sequence t̂t , where êt,i = E[t̂t,i , :]. Define Zt (α) = Zt + α(Ẑt − Zt ),
α ∈ [0, 1].
(28)
Assume that Lapp (fstate (Z), yt ) is differentiable with respect to Z, and that the token-embedding diameter is bounded as ∥E[u, :] − E[v, :]∥2 ≤ ∆max ,
∀u, v ∈ K.
(29)
Let zt,i (α) denote the i-th row of Zt (α), and define the pathdependent sensitivity sup wt,i = sup
α∈[0,1]
∂Lapp (fstate (Zt (α)), yt ) . ∂zt,i (α) 2
(30)
8
(a) AWGN: macro-F1 versus SNR.
(b) Rayleigh: macro-F1 versus SNR.
(c) AWGN: normalized cost.
(d) Rayleigh: normalized cost.
Fig. 4. Traffic-state inference performance and normalized communication cost under AWGN and Rayleigh channels.
Then Lapp (fstate (Ẑt ), yt ) − Lapp (fstate (Zt ), yt ) ≤ ∆max
L X
sup wt,i 1{t̂t,i ̸= tt,i }.
(31)
i=1
Proof: Define ϕ(α) = Lapp (fstate (Zt (α)), yt ). By the fundamental theorem of calculus, Z 1 dϕ(α) dα. (32) ϕ(1) − ϕ(0) = dα 0 Using the chain rule along Zt (α), L dϕ(α) X ∂Lapp (fstate (Zt (α)), yt ) = , êt,i − et,i . (33) dα ∂zt,i (α) i=1 For notational compactness, define gt,i (α) ≜
∂Lapp (fstate (Zt (α)), yt ) . ∂zt,i (α)
(34)
The triangle inequality and Cauchy–Schwarz inequality give Z 1X L |ϕ(1) − ϕ(0)| ≤ ∥gt,i (α)∥2 ∥êt,i − et,i ∥2 dα 0
≤
L X
i=1
(35)
sup wt,i ∥êt,i − et,i ∥2 .
i=1
If t̂t,i = tt,i , then êt,i = et,i . If t̂t,i ̸= tt,i , then (29) gives ∥êt,i −et,i ∥2 ≤ ∆max . Substituting this relation into (35) yields (31). ■
Proposition 1 shows that token errors on application-sensitive positions can contribute more strongly to application-loss degradation. The gradient-based utility in (17) is a practical local approximation of this sensitivity, while the group aggregate Wg in (20) provides a compact criterion for group-wise protection design. C. Confidence-Aware Acceptance-or-Erasure Rule We next characterize the receiver-side gating rule. The key question is whether the twin should accept a hard-decoded token or convert it into an erasure for completion. This decision is important because a wrong token update can directly contaminate the semantic twin state, whereas an erased token can be recovered from context by the completion model. For token position i at time t, let the receiver choose an output ut,i ∈ K⊥ . We assign zero cost to accepting the correct token, cost wi > 0 to accepting an incorrect token, and cost (a) λi ≥ 0 to erasing the token under synchronization mode a. Under the posterior distribution pt,i (k), the Bayes risk of outputting a token u ∈ K is Rt,i (u) = wi (1 − pt,i (u)),
u ∈ K,
(36)
while the risk of erasure is (a)
Rt,i (⊥) = λi .
(37)
The parameter wi reflects the cost of a wrong accepted update (a) at token position i, whereas λi reflects the effective penalty of sending this position to completion under mode a. In deployment, wi is approximated by the offline token utility w̄i ,
9
Fig. 5. Temporal behavior of TWIST on a representative KITTI sequence under blockwise time-varying Rayleigh fading.
or by a group-level representative value for group g(i). The (a) erasure penalty λi is not modeled analytically; it is absorbed (a) into the calibrated group-wise threshold τg(i) . Theorem 1: Mode-conditioned confidence threshold. Assume (a) wi > 0 and 0 ≤ λi ≤ wi . Let t̂t,i = arg max pt,i (k),
ct,i = max pt,i (k).
k∈K
k∈K
Then, under synchronization mode a, the Bayes-optimal receiver output is ( (a) t̂t,i , wi (1 − ct,i ) ≤ λi , u⋆t,i = (38) (a) ⊥, wi (1 − ct,i ) > λi . Equivalently, the rule can be written as a confidence-threshold test: (a)
ct,i ≥ τi
⇒ u⋆t,i = t̂t,i ,
(a)
ct,i < τi
where
⇒ u⋆t,i = ⊥, (39)
(a)
λi . (40) wi Proof: For any accepted output u ∈ K, the Bayes risk is given by (36). Since wi > 0, minimizing Rt,i (u) over u ∈ K is equivalent to maximizing pt,i (u). Therefore, the minimum-risk accepted token is the MAP estimate t̂t,i , and the corresponding accepted-token risk is (a)
τi
=1−
min Rt,i (u) = wi (1 − ct,i ). u∈K
(a)
The receiver then compares this risk with the erasure risk λi . (a) Accepting t̂t,i is optimal when wi (1 − ct,i ) ≤ λi , and erasure (a) is optimal otherwise. Rearranging gives ct,i ≥ 1 − λi /wi , which yields (39) and (40). ■ Theorem 1 explains why confidence gating is a natural receiver operation for token-state synchronization. When confidence is low, accepting a hard token creates a high risk
of semantic-state contamination. Erasing the token can be preferable because the completion model can exploit contextual structure to recover the missing position. D. Group-Wise Threshold Calibration The token-level threshold in (40) depends on the unknown (a) effective erasure penalty λi . This penalty is difficult to model analytically because it depends on the completion model, neighboring tokens, the downstream application head, and the synchronization mode. TWIST therefore uses group-wise validation calibration. For each mode a ∈ A, the implementation sets (a)
τi
(a)
= τg(i) ,
(41)
so that all token positions in the same utility group share the same threshold. Given M calibration samples, the modeconditioned thresholds are selected by M 1 X (m) min Lapp fstate Z̄(m) ({τg(a) }G ) , y , g=1 (a) M m=1 {τg ∈[0,1]}G g=1 (42) (a) where Z̄(m) ({τg }G ) is the recovered embedding sequence g=1 after mode-a transmission, confidence gating, and completion under the candidate threshold set. The optimization in (42) is low-dimensional because it searches over only G thresholds for each mode. In practice, it can be performed offline by grid search or coordinate search on a validation set. Equations (16), (27), (26), and (42) therefore play different roles. Equation (16) defines the system-level goal; (27) gives a tractable transmitterside surrogate for each mode; (42) calibrates the receiverside accept-or-erase behavior after the full gating–completion pipeline; and (26) selects among the resulting mode profiles online. The theoretical results justify the structure of the modular design, while the closed-loop performance is evaluated empirically.
10
VI. P ERFORMANCE E VALUATION This section evaluates TWIST in a dynamic road-scene digital-twin scenario. The experiments examine four aspects: traffic-state inference under wireless impairments, the performance–cost tradeoff of closed-loop mode adaptation, the quality of semantic twin-state recovery, and the contribution of feedback and recovery modules. A. Dynamic Road-Scene Digital Twin Scenario We instantiate TWIST on a road-scene digital twin constructed from KITTI sequence data. Each frame is treated as the physical observation xt , tokenized into tt , transmitted over the wireless link, and recovered as the twin-side semantic state tDT after soft decoding, confidence gating, and completion. t The twin-side task is derived traffic-state inference. For frame t, the label vector is (43) yt = ytcar , ytped , ytden , where ytcar indicates car-like object presence, ytped indicates pedestrian/cyclist/person-like presence, and ytden denotes the traffic-density class. The density label is obtained by quantizing the number of relevant traffic objects into low-, medium-, and high-density classes. This derived task evaluates whether the synchronized semantic twin state supports application-level scene understanding; it is not intended to replace the official KITTI detection or tracking benchmarks. The physical-layer realization follows the digital tokentransmission chain in Sections III–V. Tokens are mapped to binary representations, protected by group-wise LDPC coding, modulated using 16QAM, transmitted through AWGN or Rayleigh block-fading channels, and recovered by coherent demodulation and soft decoding. In the Rayleigh case, each frame experiences one block-fading realization, and the nominal SNR specifies the average operating condition. The synchronization-mode set is A = {low, med, high}. The nominal budget is B0 = N (med) , with N (low) = 0.5B0 ,
N (high) = 2B0 .
(44)
(a) Twin-state mismatch rate.
(b) Accepted-update error ratio.
(c) Erasure ratio. Fig. 6. Token-level diagnostics for semantic twin-state synchronization.
B. Baselines, Variants, and Metrics
(a)
For each mode a ∈ A, the protection profile π and threshold profile τ (a) are prepared offline. Static baselines keep a fixed mode for all frames, whereas adaptive schemes select at frame by frame. In the experiments, the tokenizer has codebook size K = 1024 and token length L = 576, arranged as a 24 × 24 token grid. Token positions are partitioned into G = 4 utility groups. The nominal budget is B0 = 4096 channel uses per frame, so the low, medium, and high modes use 2048, 4096, and 8192 channel uses, respectively. The LDPC rates are selected from {1/2, 2/3, 3/4, 5/6}. The traffic-state head, completion model, utility profile, group map, mode-conditioned protection profiles, and group-wise thresholds are trained or calibrated on the training/validation split. All results are evaluated on KITTI sequences using the same sequence split and channel random seeds across compared methods. For bar-chart results, error bars denote one standard deviation over repeated channel realizations.
We compare TWIST with fixed-mode, channel-adaptive, pixel-domain, and ablated variants. Static-Low, Static-Med, and Static-High use at = low, at = med, and at = high for all frames, respectively. Channel-Adaptive selects the mode using only γt or γ̄t , without twin-side uncertainty, drift, or priority information. The JPEG reference provides a conventional pixeldomain source-coding comparison under the considered setting. Unless otherwise specified, qt is treated as an exogenous priority signal; if derived from labels, the result is interpreted as oracle- or exogenous-priority analysis. We also evaluate ablations. The variants w/o γ, w/o ρ, w/o d, and w/o q remove the corresponding controller input. Uniform removes utility-aware unequal protection. No Gating accepts all hard-decoded tokens. No Completion applies gating but does not recover erased positions. Hard-only directly uses harddecoded tokens without gating or completion.
11
Semantic twin-state synchronization is measured by TSMRt =
(a) Controller-input ablation: macro-F1.
L 1 X DT 1 tt,i ̸= tt,i . L i=1
(46)
TSMR measures the mismatch between the physical token tt and recovered twin state tDT t . We further define PL i=1 1 t̃t,i ̸= ⊥, t̂t,i ̸= tt,i , (47) AUERt = PL i=1 1 t̃t,i ̸= ⊥ where t̂t,i is the hard-decoded token before gating and t̃t,i is the gated token. AUER measures the fraction of wrong hard-token updates among accepted updates; frames with zero accepted tokens are omitted for this metric. The erasure ratio is ρt =
(b) Controller-input ablation: normalized cost.
L 1X 1 t̃t,i = ⊥ . L i=1
(48)
Unlike TSMR and AUER, which require the physical token, ρt is observable at the twin side and can be used by the controller. The normalized communication cost is T
C̄ =
1 X N (at ) . T t=1 B0
(49)
The low, medium, and high modes therefore have normalized per-frame costs 0.5, 1, and 2, respectively. C. Main Performance and Communication Cost
(c) Protection/recovery ablation: macro-F1.
(d) Protection/recovery ablation: TSMR. Fig. 7. Ablation study of closed-loop synchronization inputs and token protection/recovery modules.
For application performance, we report
MacroF1traffic =
1 F1car + F1ped/cyc + F1density , (45) 3
where F1car and F1ped/cyc are binary F1 scores, and F1density is the macro-F1 over the three density classes. Macro-F1 is used because the derived labels can be imbalanced.
Fig. 4 reports traffic-state macro-F1 and normalized cost under AWGN and Rayleigh channels. The static baselines show the expected performance–cost tradeoff: Static-Low uses the smallest budget but gives limited reliability, Static-High provides the strongest fixed-mode reference at the highest cost, and Static-Med lies between these two operating points. Channel-Adaptive reacts to the wireless condition but does not observe the recovered twin-state quality. TWIST uses both channel and twin-side feedback statistics, and approaches the high-mode performance over a wide SNR range while reducing average cost relative to always-high synchronization. The stacked bars in Figs. 4(c) and 4(d) show that TWIST distributes synchronization resources across low, medium, and high modes rather than using a fixed budget. Relative to Static-High, it lowers the average cost by avoiding high-mode transmission when the channel and twin-state statistics do not indicate a high-risk update. Relative to Static-Med, it can allocate additional resources to difficult or priority-sensitive frames, contributing to its higher macro-F1 at the reported operating points. The JPEG reference provides a conventional pixeldomain comparison under the considered source-coding setting; it is not an upper bound for token-domain synchronization. D. Temporal Closed-Loop Behavior Fig. 5 illustrates TWIST on a representative sequence under blockwise time-varying Rayleigh fading. The panels show the nominal SNR state, twin-side uncertainty ρt , semantic drift dt , selected mode at , and rolling traffic-state correctness. The trace
12
Fig. 8. Qualitative example of token-level synchronization and twin-side traffic-state inference on a KITTI frame.
supporting the motivation for confidence-aware gating. Fig. 6(c) shows the erasure ratio. No-Gating has zero erasures but accepts more erroneous updates, whereas No-Completion can gate unreliable tokens but cannot restore the missing positions. These results support the proposed receiver design: gating reduces wrong-but-accepted updates, while completion converts erasures into recoverable semantic estimates.
F. Ablation Study Fig. 9. Macro-F1 on urgent frames.
shows that TWIST operates as a closed-loop synchronization policy rather than as a static token-transmission pipeline: the controller tends to increase the mode when the channel degrades or uncertainty rises, and can reduce the mode when the channel and twin-state statistics become more favorable. The rolling correctness curves show a more stable trend for TWIST than for Static-Med and Channel-Adaptive in this representative sequence. Static-Med cannot react to difficult channel or twinstate periods, while Channel-Adaptive responds to the channel but not directly to the recovered twin-state quality. TWIST combines channel quality, uncertainty, drift, and priority information to allocate stronger updates to more demanding intervals without using always-high synchronization. E. Semantic Twin-State Synchronization Diagnostics Fig. 6 reports token-level diagnostics for semantic twin-state synchronization. The TSMR curve in Fig. 6(a) measures the mismatch between the physical token and the recovered twin state. Static-High provides the strongest fixed-mode reference because it uses the largest budget for every frame, while TWIST approaches this high-reliability behavior with lower average cost. Static-Low and Static-Med exhibit larger mismatch at low and medium SNRs, indicating that insufficient synchronization resources can degrade the maintained twin state. Fig. 6(b) shows AUER. The No-Gating variant has a high acceptedupdate error ratio because it accepts unreliable hard decisions,
Fig. 7 studies the roles of closed-loop inputs and token protection/recovery modules. Figs. 7(a) and 7(b) compare TWIST with variants that remove one controller input at a time. The full controller achieves the highest macro-F1 among the compared variants at the reported operating point. Removing the channel-quality input weakens adaptation to link variation. Removing the uncertainty input changes resource allocation and lowers the cost, but also reduces application-state inference performance. Removing the drift or priority input leads to different performance–cost operating points. These results should be interpreted as a performance–cost analysis rather than as evidence that every signal improves every metric by the same margin. The channel-quality statistic captures the wireless condition; the erasure ratio reflects receiver uncertainty; the drift statistic captures changes in the recovered twin state; and the priority input biases the controller toward application-critical frames. Their joint use helps TWIST select a favorable operating point compared with Static-Med and Channel-Adaptive. Figs. 7(c) and 7(d) evaluate transmitter protection and receiver recovery. Uniform protection weakens utility-aware allocation. No-Completion degrades traffic-state macro-F1 and increases TSMR, showing the role of completion in converting erasures into usable semantic estimates. Hard-only synchronization performs worst among the recovery variants, indicating vulnerability to token substitutions. NoGating can preserve task performance in some cases, but results in higher TSMR, suggesting that confidence gating is more directly reflected in twin-state quality than in task-level performance alone.
13
G. Qualitative and Priority-Aware Analysis Fig. 8 provides a qualitative example of token-level synchronization on a KITTI frame. The original frame is shown with the receiver confidence map, erasure mask, selected synchronization mode, and traffic-state prediction. The example illustrates that TWIST operates on the token rather than on pixel reconstruction: low-confidence positions are masked, completion restores missing tokens, and the recovered semantic state is used for twin-side traffic-state inference. Fig. 9 shows performance on urgent frames. TWIST attains higher urgentframe macro-F1 than Static-Med and Channel-Adaptive over most evaluated SNRs. The variant without the priority input remains competitive but does not consistently match the full TWIST policy, indicating that the priority signal can bias synchronization resources toward application-critical frames. When qt is derived from labels, this result should be interpreted as exogenous- or oracle-priority analysis. In practical systems, such a signal may be provided by an upper-layer event detector, a safety policy, or a prior service requirement. VII. C ONCLUSION This paper studies digital twin synchronization from a tokencentric perspective. Instead of synchronizing pixel-level observations or uniformly protected bitstreams, TWIST maintains a recovered token at the twin side and uses it for both traffic-state inference and subsequent synchronization control. The framework combines token-utility-aware grouping, unequal protection, confidence-aware gating, completion-assisted recovery, and feedback-based mode adaptation. A utility-weighted loss bound motivates task-relevance-aware protection, while a Bayes-risk interpretation explains why low-confidence hard token decisions can be converted into erasures before completion. Experiments on a dynamic road-scene scenario show that TWIST provides a favorable balance between traffic-state inference, semantic twin-state synchronization, and communication cost compared with fixed-mode and channel-only adaptation strategies. These results show that tokens can provide a practical synchronization interface for digital twins. R EFERENCES [1] Y. Wu, K. Zhang, and Y. Zhang, “Digital twin networks: A survey,” IEEE Internet of Things Journal, vol. 8, no. 18, pp. 13 789–13 804, Sep. 2021. [2] L. U. Khan, Z. Han, W. Saad, E. Hossain, M. Guizani, and C. S. Hong, “Digital twin of wireless systems: Overview, taxonomy, challenges, and opportunities,” IEEE Communications Surveys & Tutorials, vol. 24, no. 4, pp. 2230–2254, 2022. [3] A. Alkhateeb, S. Jiang, and G. Charan, “Real-time digital twins: Vision and research directions for 6g and beyond,” IEEE Communications Magazine, vol. 61, no. 11, pp. 128–134, Nov. 2023. [4] H. Wang, J. Zhang, G. Nie, L. Yu, Z. Yuan, T. Li, J. Wang, and G. Liu, “Digital twin channel for 6g: Concepts, architectures and potential applications,” IEEE Communications Magazine, vol. 63, no. 3, pp. 24–30, Mar 2025. [5] A. van den Oord, O. Vinyals, and K. Kavukcuoglu, “Neural discrete representation learning,” in Adv. Neural Inf. Process. Syst., 2017, pp. 6306–6315. [6] P. Esser, R. Rombach, and B. Ommer, “Taming transformers for highresolution image synthesis,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2021, pp. 12 873–12 883. [7] H. Chang, H. Zhang, L. Jiang, C. Liu, and W. T. Freeman, “Maskgit: Masked generative image transformer,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022, pp. 11 315–11 325. [8] D. Gündüz, Z. Qin, I. E. Aguerri, H. S. Dhillon, Z. Yang, A. Yener, K.K. Wong, and C.-B. Chae, “Beyond transmitting bits: Context, semantics, and task-oriented communications,” IEEE J. Sel. Areas Commun., vol. 41, no. 1, pp. 5–41, Jan. 2023.
[9] H. Xie, Z. Qin, G. Y. Li, and B.-H. Juang, “Deep learning enabled semantic communication systems,” IEEE Trans. Signal Process., vol. 69, pp. 2663–2675, Apr. 2021. [10] J. Shao, Y. Mao, and J. Zhang, “Learning task-oriented communication for edge inference: An information bottleneck approach,” IEEE J. Sel. Areas Commun., vol. 40, no. 1, pp. 197–211, Jan. 2022. [11] E. Bourtsoulatze, D. B. Kurka, and D. Gündüz, “Deep joint sourcechannel coding for wireless image transmission,” IEEE Trans. Cogn. Commun. Netw., vol. 5, no. 3, pp. 567–579, 2019. [12] D. B. Kurka and D. Gündüz, “Bandwidth-agile image transmission with deep joint source-channel coding,” IEEE Trans. Wireless Commun., vol. 20, no. 12, pp. 8081–8095, 2021. [13] H. Wei, W. Ni, W. Wang, W. Xu, D. Niyato, and P. Zhang, “Token communication in the era of large models: An information bottleneckbased approach,” IEEE Wireless Commun. Lett., 2025, early access. [14] L. Qiao, M. B. Mashhadi, Z. Gao, R. Schober, and D. Gündüz, “ToDMA: Large model-driven token-domain multiple access for semantic communications,” May 2025. [15] J. Ying, Z. Qin, Y. Feng, L. Wang, and X. Tao, “Joint semantic-channel coding and modulation for token communications,” IEEE Trans. Wireless Commun., vol. 25, pp. 8179–8193, 2026. [16] X. Lin, L. Kundu, C. Dick, E. Obiodu, T. Mostak, and M. Flaxman, “6g digital twin networks: From theory to practice,” IEEE Communications Magazine, vol. 61, no. 11, pp. 72–78, Nov. 2023. [17] L. Bariah, H. Sari, and M. Debbah, “Digital twin-empowered communications: A new frontier of wireless networks,” IEEE Communications Magazine, vol. 61, no. 12, pp. 24–36, Dec. 2023. [18] S. Jiang and A. Alkhateeb, “Digital twin based beam prediction: Can we train in the digital world and deploy in reality?” in Proc. IEEE International Conference on Communications Workshops (ICC Workshops), May 2023, pp. 36–41. [19] Z. Tao, W. Xu, Y. Huang, X. Wang, and X. You, “Wireless network digital twin for 6g: Generative ai as a key enabler,” IEEE Wireless Communications, vol. 31, no. 4, pp. 24–31, Aug. 2024. [20] A. Gürses, G. Reddy, S. Masrur, Ö. Özdemir, İ. Güvenç, M. L. Sichitiu, A. Şahin, A. Alkhateeb, M. Mushi, and R. Dutta, “Digital twins and testbeds for supporting ai research with autonomous vehicle networks,” IEEE Communications Magazine, vol. 63, no. 4, pp. 56–62, Apr. 2025. [21] Z. Weng and Z. Qin, “Semantic communication systems for speech transmission,” IEEE J. Sel. Areas Commun., vol. 39, no. 8, pp. 2434– 2444, Aug. 2021. [22] H. Xie, Z. Qin, and G. Y. Li, “Task-oriented multi-user semantic communications for VQA task,” IEEE Wireless Commun. Lett., vol. 11, no. 3, pp. 553–557, 2022. [23] S. Ma, W. Qiao, Y. Wu, H. Li, G. Shi, D. Gao, Y. Shi, S. Li, and N. Al-Dhahir, “Task-oriented explainable semantic communications,” IEEE Trans. Wireless Commun., vol. 22, no. 12, pp. 9248–9262, 2023. [24] T.-Y. Tung, D. B. Kurka, M. Jankowski, and D. Gündüz, “Deepjscc-q: Constellation constrained deep joint source-channel coding,” IEEE J. Sel. Areas Inf. Theory, vol. 3, no. 4, pp. 720–731, 2022. [25] K. Yang, S. Wang, J. Dai, X. Qin, K. Niu, and P. Zhang, “Swinjscc: Taming Swin transformer for deep joint source-channel coding,” IEEE Trans. Cogn. Commun. Netw., vol. 11, no. 1, pp. 90–104, 2025. [26] J. Xu, T.-Y. Tung, B. Ai, W. Chen, Y. Sun, and D. Gündüz, “Deep joint source-channel coding for semantic communications,” IEEE Commun. Mag., vol. 61, no. 11, pp. 42–48, 2023. [27] B. Bai, “Forget bit, it is all about token: Towards semantic information theory for llms,” 2025, technical report. [28] J. Peng, H. Xing, Z. Xiao, L. Xu, and X. Lei, “Large model empowered multi-modal semantic communication with selective tokens for training,” IEEE Signal Process. Lett., vol. 32, pp. 2967–2971, 2025. [29] F. Solat, J. Lee, M. Seif, D. Niyato, and H. V. Poor, “Federated learning-enabled hybrid language models for communication-efficient token transmission,” IEEE Internet Things J., vol. 12, no. 24, pp. 53 574– 53 592, 2025. [30] R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2022, pp. 10 684–10 695. [31] L. Guo, W. Chen, Y. Sun, B. Ai, N. Pappas, and T. Q. S. Quek, “Diffusiondriven semantic communication for generative models with bandwidth constraints,” IEEE Trans. Wireless Commun., vol. 24, no. 8, pp. 6490– 6503, Aug. 2025.