1
Adapting Diffusion Language Models for Lossless Pixel-Level Image Transmission
arXiv:2606.06273v1 [cs.IT] 4 Jun 2026
Tianqi Ren, Rongpeng Li, Xianfu Chen, Yingyu Li, and Zhifeng Zhao
Abstract—Lossless pixel-level image transmission is a fundamental regime beyond semantic communications, because exact recovery requires both accurate symbol probability modeling and reliable delivery over noisy channels. This paper proposes DDM-SSCC, a discrete-diffusion-model-based separate sourcechannel coding framework for lossless image transmission. Different from raster-order autoregressive coding, the proposed source codec adapts a diffusion language model to pixel-token restoration and performs synchronized reverse arithmetic coding under bidirectional attention, allowing multiple masked tokens to be coded within one reverse denoising step. This progressive restoration process also yields a more favorable source representation for noisy transmission, since newly restored tokens can serve as bidirectional context in subsequent denoising steps. To bridge the gap between generation-oriented masked denoising and lossless arithmetic coding, we further introduce a Haltonguided denoising order, a mask-ratio-aware cosine schedule, and a lightweight temperature calibration module. These designs respectively improve spatial coverage, adapt the denoising pace to context reliability, and calibrate the probability tables used by arithmetic coding. Experiments on CIFAR10, DIV2K-LR-X4, and Kodak over additive white Gaussian noise and Rayleigh fading channels show that DDM-SSCC achieves better exact-recovery performance than representative lossless and semantic communication baselines, while ablation studies verify the effectiveness of the proposed denoising order, schedule, and calibration modules. Index Terms—Lossless image transmission, pixel-level communication, separate source-channel coding, discrete diffusion models
I. I NTRODUCTION
R
ecently, semantic communications have manifested strong effectiveness for distortion-oriented delivery. However, the underlying deep joint source-channel coding (JSCC) inevitably incurs pixel-wise reconstruction errors [1]– [4]. This limitation becomes critical in fidelity-sensitive scenarios such as remote medical imaging, industrial inspection, and scientific image delivery, where the receiver must recover the original image exactly. By contrast, lossless pixel-level transmission provides a necessary probabilistic interface to the raw image signal, and serves as a foundation for future patch-level or higher-level semantic representations. Recent progress in large-model-based compression and pixel-level sequence modeling suggests that powerful probabilistic models T. Ren and R. Li are with College of Information Science and Electronic Engineering, Zhejiang University, Hangzhou 310027, China (email: {rentianqi, lirongpeng}@zju.edu.cn). X. Chen is with Shenzhen CyberAray Network Technology Co., Ltd, Shenzhen 518000, China (email: [email protected]). Y. Li is with School of Mechanical Engineering and Electronic Information, China University of Geosciences, Wuhan 430074, China (email: [email protected]). Z. Zhao is with Zhejiang Lab, Hangzhou 311121, China (email: [email protected]).
can serve as effective source coders when their predictive distributions are coupled with entropy coding [5]–[9]. This trend motivates the investigation of lossless pixel-level transmission as a complementary direction beyond semantic reconstruction. Correspondingly, a separate source-channel coding (SSCC) architecture provides a natural and principled solution, particularly when combined with modern neural source models and reliability-oriented channel decoding techniques [10]– [13]. Under noisy channels, however, the source representation should be evaluated not only by compression efficiency, but also by how well it supports reliable exact recovery after imperfect channel decoding. A. Related Works Existing studies on semantic communications have mainly focused on task-oriented or distortion-oriented signal reconstruction. Deep JSCC directly maps source signals to channel symbols and has shown strong robustness for wireless image transmission [1]. Subsequent works further improve semantic visual transmission by introducing sparse transmission, adaptive learning, generative priors, and foundation models [2], [14]–[18]. More recently, diffusion priors have also been introduced into semantic communication systems. For example, DiT-JSCC employs generative diffusion modeling to improve perceptual reconstruction quality under bandwidth- or SNRlimited channels [19]. However, these diffusion-aided semantic communication methods are still designed for semantic or perceptual fidelity rather than bit-exact pixel recovery. Therefore, digital or hybrid solutions based on SSCC remain indispensable when lossless reconstruction is required. In such systems, source coding determines the number of bits to be protected, while channel coding determines whether these bits can be reliably delivered over a noisy link. Classical channel codes, such as LDPC and Polar codes, provide strong digital protection mechanisms [20], [21]. Recent neural decoders further enhance reliability-oriented channel decoding [11]– [13], [22]. In this work, we follow the SSCC principle and focus on improving the source-coding component for lossless pixel-level image transmission. From the source-coding viewpoint, lossless image transmission depends on accurate probability estimation before entropy coding. Classical lossless image codecs improve compression by exploiting local context and adaptive prediction [23]– [25]. Learned lossless codecs further strengthen conditional modeling with neural networks, flows, hierarchical latent variables, or bit-plane representations [26]–[32]. Meanwhile, arithmetic coding provides a principled interface between predictive distributions and binary lossless representations
2
[33]–[35]. Recent large-model approaches further show that powerful sequence models can serve as effective compressors when their token probabilities are coupled with arithmetic coding [5]–[8], [36]. Representative image-oriented sequence models and coders, including iGPT and P2 -LLM, confirm that autoregressive next-token modeling can be adapted to pixellevel compression [9], [37]. Nevertheless, strict autoregression still suffers from a fixed coding order and inherently sequential decoding over long pixel streams. This limits scalability and offers little flexibility in balancing compression performance against inference latency. These deficiencies motivate nonautoregressive sequence modeling mechanisms that can exploit broader contexts and support more flexible token restoration orders. Discrete diffusion models provide an appealing alternative by restoring masked tokens under bidirectional attention [38]– [42]. Concretely, diffusion models adopt a reverse process where coding starts from a fully masked state and progressively removes masks until completely recovering the clean pixel-token sequence. Compared to predicting the next pixel in a rigid causal order, this non-causal mechanism is especially attractive for images with spatially heterogeneous predictability. It also provides a progressively recoverable representation, since restored tokens can be reused as bidirectional context in later denoising steps. Recent masked diffusion language models and block diffusion models further demonstrate that non-autoregressive or semi-autoregressive token restoration can provide flexible generation and decoding behaviors beyond strict left-to-right prediction [40]–[43]. Consequently, diffusion models could alleviate the latency induced by strict onetoken-per-step autoregression while offering a source representation better suited to noisy transmission. However, existing diffusion priors in visual communication are mainly used for generative reconstruction or perceptual enhancement [17], [19], rather than for entropy-coded lossless compression and transmission. Transforming a diffusion model into a lossless source coder is non-trivial. Firstly, the deterministic left-to-right decoding order underlying autoregressive arithmetic coding no longer holds. Therefore, to support arithmetic coding, the encoder and decoder must share the same denoising path together with aligned predictive probabilities at each denoising stage. Secondly, during the iterative restoration in DDM [38], [40]– [42], confidence-based token selection is a common and effective sampling strategy [44]. However, confidence-driven restoration may repeatedly select locally smooth and highly correlated pixels [34], [35], so the tokens encoded within the same jump step are predicted from the same pre-update context and may introduce redundant coding cost. A similar clustering issue has also been observed in masked image generation, motivating low-discrepancy alternatives such as Halton scheduling [45]. Furthermore, the step-wise denoising size is usually not matched to the reliability of the current context: early denoising states contain little information and require conservative updates, whereas late states can safely process more tokens. Lastly, the predictive distributions of diffusion language models are optimized for denoising rather than directly for arithmetic coding, while the latter is sensitive to the
exact probability table used for interval subdivision [34], [35]. Therefore, overconfident probabilities at highly masked states may increase the code length even when the top prediction is plausible. B. Contributions In light of the above observations, this paper develops a Discrete Diffusion Model-based Separate Source-Channel Coding (DDM-SSCC) framework for pixel-level lossless image transmission. The proposed source codec turns diffusion denoising into a synchronized arithmetic-coding process and further improves it with schedule-aware denoising designs. To tackle the practical mismatch between generation-oriented diffusion denoising and lossless arithmetic coding, we further develop the Halton-guided denoising rule, so as to improve the spatial coverage of denoised tokens. Besides, we introduce a cosine denoising schedule and mask-ratio-aware calibration module, thus adapting the denoising pace to context reliability and calibrating the probability tables consumed by arithmetic coding. While providing a summary of the key differences between our algorithm and relevant literature in Table I, we summarize the main contributions as follows: • We propose a synchronized discrete-diffusion source coding protocol for lossless pixel-token compression. Starting from a fully masked sequence, the encoder and decoder follow the same reverse denoising trajectory, use identical target positions and predictive distributions, and therefore make non-autoregressive diffusion compatible with arithmetic coding. To the best of our knowledge, this work is the first to make a DDM compatible with arithmetic coding and integrate it into SSCC for exact pixel-level image transmission. • We analyze the denoising-position strategy and introduce a Halton-guided low-discrepancy denoising order that improves spatial coverage while remaining exactly reproducible at the decoder. Meanwhile, we design two maskratio-aware improvements for the diffusion source coder: a cosine denoising schedule that controls the number of processed tokens at each step, and a lightweight temperature calibration module that adapts the probability distribution used by arithmetic coding. • We embed the proposed source coder into an SSCC transmission system with a fixed digital channel-protection backend and evaluate it under AWGN and Rayleigh fading channels. Experimental results show that the proposed source representation yields a more favorable exactrecovery operating point, while ablation studies verify the effectiveness of the proposed denoising order, schedule, and calibration modules. C. Paper Structure The remainder of this paper is organized as follows. Section II introduces the preliminaries, system model, and problem formulation. Section III presents the proposed discrete diffusion-based lossless source coding framework. Section IV reports the simulation settings and experimental results. Section V concludes this paper.
3
TABLE I S UMMARY AND COMPARISON OF RELATED PAPERS . Lossless Pixel-Level Recovery
LLM-Based Source Coding
Bidirectional Diffusion Modeling
Bourtsoulatze et al. [1], Tong et al. [2]
#
#
#
Tan et al. [19], Liang et al. [17]
#
#
References
Alakuijala et al. [25], Bai et al. [31]
#
#
H #
Ren et al. [10]
# #
Chen et al. [9] Chang et al. [44], Besnier et al. [45]
#
# H
This paper
Brief Description Performing distortion-oriented or task-oriented visual transmission with deep JSCC and semantic feature delivery. Employing diffusion or diffusion-transformer priors for semantic or perceptual JSCC reconstruction without bit-exact pixel recovery. Improving lossless image compression through conventional codecs or learned residual probability modeling. Combining LLM-based lossless source decoding with reliable channel protection under the SSCC architecture. Utilizing autoregressive large language models for next-pixel probability estimation and arithmetic coding. Enabling bidirectional masked token restoration with iterative generation and low-discrepancy scheduling. First applying discrete diffusion to lossless compression and integrating synchronized diffusion source coding into SSCC for exact pixel-level transmission.
indicates fully included; # H means partially included; # denotes not included.
Notations:
II. P RELIMINARIES & S YSTEM M ODEL Beforehand, mainly used notations are summarized in Table II. A. Preliminaries 1) Lossless Source Coding with Probabilistic Sequence Models: Let x1:N = (x1 , · · · , xN ) ∈ DN denote an N length discrete source sequence with true distribution ρ. Under ideal entropy coding, the code length is determined by the negative log-likelihood [46], namely − log2 ρ(x1:N ). Thus, a probabilistic model becomes a compressor by assigning high probability to the observed sequence. For an autoregressive model, N Y ρ(x1:N ) = ρ(xi | x<i ), (1)
Along with encoding all symbols, the bitstream is chosen as a finite binary fraction in the final interval. The decoder uses the same contexts and probability tables to identify the unique ith symbol whose subinterval contains this fraction [Li , Ui ), so losslessness requires synchronized symbol order, probability tables, and context updates. Finally, the resulting model-based codelength is approximately X ℓ(x1:n ) ≃ − log2 ρ̃(yi | ci ), (4) i
with expected codelength characterized by the cross-entropy "
#
H(ρ, ρ̃) = Ex1:n ∼ρ −
X
log2 ρ̃(yi | ci ) .
(5)
i
i=1
and the corresponding ideal codelength becomes − log2 ρ(x1:N ) = −
N X
log2 ρ(xi | x<i ).
(2)
i=1
In practice, the true distribution ρ is approximated by a predictive model ρ̃, whose output probabilities can be coupled with arithmetic coding to produce a lossless bitstream m [33]– [35]. At coding step i, let yi ∈ D be the target symbol and ci be the shared encoder-decoder context. Assume that the values in a dictionary D are arranged according to a fixed order shared by the encoder and decoder. For a symbol value d ∈ D, P define the cumulative probability preceding d as ′ Fi (d) = d′ ≺d ρ̃(d | ci ). Starting from I0 = [L0 , U0 ) = [0, 1), encoding yi updates the interval as Li = Li−1 + (Ui−1 − Li−1 ) Fi (yi ), Ui = Li−1 + (Ui−1 − Li−1 ) [Fi (yi ) + ρ̃(yi | ci )] .
(3)
Thus, better probability estimation directly improves lossless compression, while non-standard coding orders additionally require deterministic synchronization. 2) Discrete Diffusion-Based Sequence Modeling: Discrete diffusion models define a forward corruption process and a learned reverse denoising process over discrete sequences [38], [40]. Let x(0) ∈ DN be a clean sequence. For consistency with the source-coding formulation in Section III, we use the superscript (t) to index the diffusion step and the subscript ℓ to index the sequence position. A categorical forward process can be written as q(x(t) | x(t−1) ) =
N Y
(t)
(t−1)
q(xℓ | xℓ
).
(6)
ℓ=1
This work focuses on the absorbing-state formulation, where a special mask token m is used as the absorbing state [38],
4
TABLE II L IST OF KEY NOTATIONS USED IN THIS PAPER . Notation
Definition
X, X̂ H, W, C
Source image and reconstructed image. Numbers of image or patch rows, columns, and color channels. X(p) , s(p) The p-th image patch and its flattened pixel-value sequence. N Number of maskable image tokens in one patch-level coding unit. Φp (·), Φ−1 Patch flattening and inverse patch reconstruction operap (·) tions. T (·), T −1 (·) One-to-one pixel-to-token mapping and inverse token-topixel mapping. V, D Full tokenizer vocabulary and valid 256-symbol pixeltoken dictionary. x Patch-level token sequence, where x0 = <bos> is fixed and xi ∈ D for i ≥ 1. m Absorbing mask token. ρ, ρ̃ True source distribution and model-based predictive distribution used for arithmetic coding. m, B(x), Rs Source bitstream, its length, and source rate in bits per pixel. Cs (·), Cs−1 (·) Source encoder and decoder. Ce (·), Ce−1 (·) Channel encoder/modulator and channel decoder. J, K, Nc , Rc Number of channel blocks, information bits, codeword length, and code rate. 2 h, n, σn Channel coefficient, additive noise vector, and noise variance. y, m̂ Received signal and decoded source bitstream. θ, fθ (·) Parameters and neural backbone of the diffusion source coder. T, t Number of diffusion steps and step index. x(t) Partially denoised token state before reverse step t. M(t) , C (t) Masked and visible token-position sets at step t. βt , ᾱt Forward masking probability and cumulative survival probability in the absorbing diffusion process. (t) Z(t) , zi,d Shifted pixel-token logits and the logit of token value d at position i during reverse step t. (t) ρ̂θ (d, i) Model-predicted clean-token probability at position i. (t) pi (d) Probability of token d at position i used for coding. π (t) , kt Ordered denoising list and number of tokens processed at reverse step t. ri Halton-based priority score assigned to position i. εt , εmin , εmax , γ Mask-ratio-aware calibration temperature and its shared hyperparameters. u(t) Synchronized encoder/decoder runtime state used in Algorithm 2. τ, µτ Sampled forward diffusion step and corresponding training mask ratio. x̃(τ ) , Lft Corrupted training sequence and fine-tuning loss. st , α t Reverse progress variable and target cumulative denoising ratio in the cosine schedule. Bcomm , B0 Communication-resource consumption and prescribed communication budget. Pimg , η Exact image-recovery probability and communication variables.
[40]. At each forward diffusion step, an unmasked symbol is either preserved or replaced by m: (t) (t−1) xℓ = m, xℓ ̸= m; βt , (t) (t−1) (t) (t−1) (t−1) q(xℓ | xℓ ) = 1 − βt , xℓ = xℓ , xℓ ̸= m; (t) (t−1) 1, xℓ = m, xℓ = m, (7) where βt is the masking probability. Equivalently, conditioned
on the clean sequence, the marginal corruption distribution is (t) (0) xℓ = xℓ , ᾱt , (t) (0) q(xℓ | xℓ ) = 1 − ᾱt , x(t) (8) ℓ = m, 0, otherwise, Qt where ᾱt = τ =1 (1 − βτ ). The reverse process uses a neural model to estimate clean symbols from partially masked states. For a state x(t) , the model outputs categorical distributions over the discrete al(t) phabet. Let M(t) = {ℓ ∈ {1, · · · , N } | xℓ = m} denote the masked position set until denoising step t and let C (t) = {1, · · · , N } \ M(t) denote the visible position set, corresponding to a context ct . The current partially observed state x(t) therefore contains the denoised tokens at positions in C (t) and mask tokens at positions in M(t) . For any masked position ℓ ∈ M(t) , the model provides a conditional distribution as (t) (0) ρ̂θ (d, ℓ) ≜ pθ xℓ = d | x(t) , ℓ , d ∈ D. (9) Here, conditioned on the whole partially observed state C (t) , (t) ρ̂θ denotes the model-based estimate of the true source distribution ρ, and gives the probability table supplied to arithmetic coding. Accordingly, a discrete diffusion model can be viewed as a probabilistic sequence model that supplies context-dependent symbol probabilities for masked positions. B. System Model As illustrated in Fig. 1, we consider pixel-level lossless image transmission over a noisy digital link under an SSCC architecture. Consistent with patch-wise source processing in learned lossless image compression and large-model image coders [9], the source coder operates patch by patch, and the final image bitstream is obtained by concatenating the compressed outputs of all patches. Without loss of generality, for a source image X ∈ {0, 1, · · · , 255}H×W ×C , where H, W , and C denote the numbers of pixel rows, columns, and color channels, the image is first partitioned into nonoverlapping patches. The p-th patch is serialized into a pixel sequence s(p) = Φp (X(p) ) ∈ {0, 1, · · · , 255}Np ,
(10)
where Φp (·) denotes patch flattening with RGB-ordered serialization and Np is the number of pixel values in the patch1 . Each pixel value is then mapped to a unique token in a 256dimensional digital dictionary D, and a beginning-of-sequence token <bos> is prepended: x = [x0 , T (s)] ∈ V N +1 ,
(11)
where V denotes the tokenizer vocabulary, x0 = <bos> is fixed and not entropy-coded, and T (·) is a one-to-one pixelto-token mapping from valid pixel values to the digital token dictionary D ⊂ V. Only the image-token positions are entropycoded, while the <bos> token is shared by the codec as 1 For convenience of representation, we ignore the subscript p in the subsequent description.
5
Language Model-Based Lossless Source Encoding
Image-to-Token Processing
Original Image 𝐗
Context sequence 𝐜𝑖
Patch Flattening (RGB Sequence) 6 24 183 3 23 168 0 25 163 …
Patch
one pixel
R G B R G B
Digital Token Dict 𝒟
Token Sequence 𝐱 50256, 21, 1731, 24839, 18, 1954, 14656, …
𝑝
𝕀𝑖−1 :
Softmax
Initialize: Start with [0,1)
Token Sequence
get prob interval 𝕀𝑖 = 𝕀𝑖−1 (𝐷𝑗 )
Digital Tokenizer 𝒯 “0”: 15, “1”: 16, “2”: 17, … , “255”: 13381
Language Model 𝜌( 𝑥𝑖 |𝐜𝑖 )
Previous Input
𝑗
Token-to-Image Processing
Entropy Encoder
Output
Patch
Digital Token Dict 𝒟
Pixel Sequence 𝐬Ƹ
“0”: 15, “1”: 16, “2”: 17, … , “255”: 13381
<bos>, “6”, “24”, “183”, “3”, “23”, “168”, “0”, …
Context Sequence 𝐜𝑖
Patch Reconstruction (RGB Ordered)
Modulator Channel
Demodulator
𝜌(𝑥 𝑖 |𝐜𝑖 )
Previous Output
𝑝
𝕀𝑖−1 :
Softmax find 𝐷𝑗 satisfied: 𝕀𝑖 ⊆ 𝕀𝑖−1 (𝐷𝑗 )
logits
Information Bits
Output
𝑗
𝐷∈𝒟
Channel Decoder Neural Decoders, e.g. ECCT
Decoded Token 𝑥𝑖 = 𝐷𝑗
𝕀𝑖
…
AWGN/Fading
…Language Model
Initialize: Start with <bos>
Token Sequence
LDPC/Polar
[Step 𝒊]
Language Model-Based Lossless Source Decoding
Digital Detokenizer 𝒯 −1
Channel Encoder
Encoded Token 𝑥𝑖 = 𝐷𝑗
find 𝑥𝑖 in 𝒟
prob intervals → bits
Reconstructed Image 𝐗
Information Bits
𝐷∈𝒟
𝕀𝑖 Prob Interval 𝕀𝑖
logits
Input
Prob Interval 𝕀𝑖
Entropy Decoder bits → prob intervals
Input
[Step 𝒊]
Fig. 1. Overall DDM-SSCC pipeline for lossless pixel-level image transmission.
a deterministic anchor. Therefore, lossless recovery in token space is equivalent to lossless recovery in the original image space. By a compressor Cs (·), x is transformed into a binary bitstream m = Cs (x), whose length is denoted by B(x) = |m|. The source rate of one coding unit is measured by B(x) Rs = bpp. (12) N where bpp stands for the number of bits per pixel. For channel transmission, the compressed bitstream m is segmented into J = ⌈B(x)/K⌉ information blocks, where K denotes the number of information bits per channel-coding block. Each K-bit block is then encoded into an Nc -bit codeword by a channel code of rate Rc = K/Nc and mapped to channel symbols through modulation. The encoded source bitstream is then transmitted over a channel. By denoting channel coding and modulation as Ce (·), the received signal can be written as ee (m) + n, y = hC
(13)
where h is the channel coefficient and n ∼ N (0, σn2 I). At the receiver, a channel decoder recovers an estimate m̂. The source decoder then reconstructs the image as: T −1 Cs−1 (m̂) . (14) X̂ = Φ−1 p
of the diffusion-based source coder, and let η collect the communication-side design variables such as channel rate and transmission energy. The system-level objective can be written as max Pimg = Pr(X̂ = X) θ,η
s.t.
m = Cs (x; θ), m̂ = Ce−1 (y; η),
(16)
Bcomm (η, m) ≤ B0 , where Bcomm denotes the prescribed communication budget. In this paper, we focus on the source-coding component and treat the channel decoder as a fixed external module (e.g., error correction code transformer (ECCT) [11]). Even if the channel receiver is fixed, reducing the source burden and improving the stage-wise probability quality of arithmetic coding can still shift the full system toward a better reliability regime. Moreover, the source representation itself affects how reliably exact recovery can be maintained under the same communication budget. A diffusion-based source coder is therefore appealing not only for non-autoregressive modeling, but also for its progressive restoration process, where restored tokens can enhance later predictions as bidirectional context. Consequently, the central question becomes whether a diffusionbased source coder can achieve competitive compression while yielding a more favorable exact-recovery operating point under the same communication budget.
C. Problem Formulation For pixel-level lossless transmission, it aims to accomplish exact end-to-end image recovery, namely
III. P ROPOSED D ISCRETE D IFFUSION - BASED L OSSLESS S OURCE C ODING F RAMEWORK
(15)
A. Discrete Diffusion Model-based Source Codec with a Synchronized denoising schedule
Unlike distortion-oriented communication, any pixel mismatch is regarded as a failure. Let θ denote the parameters
As introduced in Section II, the source image is first partitioned into patches, flattened into RGB-ordered pixel
X̂ = X.
6
Autoregressive Casual Attention 𝑥1
··· 𝑥𝑡−1
𝑥𝑡
Given the ordered denoising list π (t) and the shifted logits (t) Z = [zi,d ] ∈ RN ×|D| in (18), for the j-th selected position (t) ij , Eq. (9) can be formally written as
Discrete Diffusion
(t)
Bidirectional Attention
𝑥𝑡+1
𝑥𝑁
···
𝑥1
[M]
[M]
[M]
[M]
𝑥5
𝑥𝑁
(t)
Coding State (partially revealed)
Coding State (raster-scan)
[M] known
[M]
[M]
Step 3/16
Inputs
<bos>
𝑥1
𝑥2
masked
ℎ0
ℎ1
ℎ2
ℎ3
ℎ0
ℎ1
ℎ2
ℎ3
predict 𝑥𝑖+1
Shifted Logits
∅
ℎ0
ℎ1
ℎ2
<bos>
𝑥1
𝑥2
𝑥3
use ℎ𝑖 for 𝑥𝑖+1
CE Loss
𝑥1
𝑥2
𝑥3
𝑥4
Labels
Fig. 2. Comparison between raster-order autoregressive coding and discrete diffusion restoration.
sequences, and converted into a digital token sequence by the one-to-one tokenizer T (·). For a patch-level coding unit x = [x0 , x1 , · · · , xN ], instead of processing tokens in the raster-scan order of autoregressive coding, the discrete diffusion source coder starts from a fully masked state x(T ) = [x0 , m, · · · , m],
(17)
where m denotes the absorbing mask token, and progressively restores the masked image tokens until x̂ = x(0) . At reverse denoising step t, the current state x(t) contains both restored and masked positions. The diffusion backbone fθ (·) takes this partially denoised sequence as input and produces shifted pixel-token logits: Z(t) = Shift fθ (x(t) ) . (18) D
The shift operation follows the DiffuGPT adaptation recipe from autoregressive language models and is shared by finetuning and inference [43]. Compared to autoregressive source coding, wherein the next coding context and the next probability table are naturally synchronized due to the left-to-right factorization, i.e., p(x) = QN p(x | x0 , x1 , · · · , xi−1 ), a discrete diffusion model i i=1 predicts masked tokens from a shared partially denoised state under bidirectional attention [40], [44]. Thus, as illustrated in Fig. 2, the coding context for position i is the pair (x(t) , i) rather than a causal prefix, requiring an explicit denoisingpath synchronization rule. In other words, compatibility with arithmetic coding requires the encoder and decoder to share the coding state, denoising positions, and probability distributions at each denoising step. Specifically, as elaborated later in Section III-B, a synchronized scheduler selects an ordered denoising list π
(t)
(t) (t) (t) = (i1 , i2 , · · · , ikt ),
ij ,d
(t) d′ ∈D exp(zi(t) ,d′ )
,
d ∈ D. (20)
π
(t)
(t)
⊆M ,
The resulting discrete distribution is then used by arithmetic coding for interval subdivision. Following (3), the encoder (t) t , then narrows the arithmetic-coding interval using {pj }kj=1 kt fills the selected kt positions to obtain {xi(t) }j=1 and then j
Shift
CE Loss
Labels
[M]
[M]
Alignment (shift operation) Raw Logits
𝑥3
Model Forward
Logits
[M] Step 3/5
Alignment (natural next-token)
exp(z (t) ) j
current step (multi-tokens)
[M] [M]
j
known
current step (one token) unknown
(t) (t) pj (d) = ρ̂ (t) (d) = P θ,i
(19)
where kt is the number of tokens processed at this denoising step (i.e., decoded tokens). Since π (t) is deterministically generated from a shared scheduling rule and the current masked set, no side information about the denoising order is required.
x(t−1) . Eventually, a reconstructed image X̂ can be obtained through inverse tokenization and patch reconstruction. Notably, the number of diffusion steps T and the step-wise denoising size kt jointly control the trade-off between context refresh frequency and coding complexity. B. Halton-Guided Denoising Position Selection Generally, the synchronized codec can in principle use any deterministic denoising order, such as the confidence-based method in MaskGIT [44]. However, as discussed in masked image generation [45], a well-spread denoising set π(t) can reduce same-step redundancy. Therefore, we introduce Halton scheduling to guide denoising-position selection for improved spatial coverage. For a patch with spatial size H × W and channel dimension C, let N = HW C be the number of maskable image tokens. As shown in Fig. 3(a), each position i ∈ {1, · · · , N } is assigned according to a Halton-based priority map constructed by quantizing a Halton sequence to the patch grid and recording the first-visit order of discrete token positions. For lossless diffusion coding, the Halton priority map depends only on the patch geometry and shared deterministic rules, not on the source content, the logits, or any decoder-unknown information. Hence, this property also reduces same-step dependence while remaining exactly reproducible at the decoder. Specifically, the Halton-based priority map is built from the radical-inverse sequence [47]. For an integer g ≥ 1, let g = PL ℓ a b be its base-b expansion, where L is a non-negative ℓ=0 ℓ integer and aℓ ∈ {0, 1, · · · , b−1}. The radical-inverse function is defined in terms of a and b as L X ϕb (g) = aℓ b−(ℓ+1) . (21) ℓ=0
That is, ϕb (g) maps the base-b digits of g to the fractional part in reverse order. Given a set of pairwise coprime bases b = (b1 , · · · , bs ), the g-th Halton point is obtained by ug = [ϕb1 (g), · · · , ϕbs (g)] ∈ [0, 1)s .
(22)
s The points {ug }∞ g=1 form a Halton sequence in [0, 1) . In our patch-level codec, we use b = (2, 3) for grayscale patches and b = (2, 3, 5) for RGB patches. For an RGB patch, i.e., C = 3, a Halton point uq = (ug,1 , ug,2 , ug,3 ) is quantized to a discrete patch coordinate by
hg = 1+⌊Hug,1 ⌋ ,
wg = 1+⌊W ug,2 ⌋ ,
cg = 1+⌊Cug,3 ⌋ , (23)
7
(a) Halton-guided denoising position selection
(b) Cosine reverse denoising schedule
decides where to denoise
decides how many tokens to denoise at each step 1
0.74 0.16 0.86 0.23 0.50 0.98 0.30 0.61 0.56 0.92 0.03 0.34 0.67 0.45 0.80 0.19 0.88 0.52 1.00 0.08 0.31 0.63 0.41 0.75
(fraction of denoised tokens)
token index
cosine
0.36 0.49 0.13 0.81 0.20 0.58 0.94 0.27
0.97 0.06 0.28 0.72 0.39 0.84 0.22 0.48 0.17 0.44 0.78 0.55 0.91 0.02 0.66 0.11
denoised
Deterministic low-discrepancy priority map
high priority
low priority
Halton order 𝑖𝑔 ⇒
𝑟𝑖𝑔 = 𝑁 − 𝑔 + 1
logit
0.5
0
time (reverse steps) early conservative (few tokens early)
Cumulative denoising ratio after reverse step 𝑡
t=1
late aggressive (more tokens later)
calibrated probability
raw logits
token index
𝜋 𝛼𝑡 = 1 − cos(𝑠𝑡 ⋅ ) 2
token index
Late stage (low mask ratio)
linear t=T
- high uncertainty - avoid overconfidence
logit
0.47 0.83 0.60 0.95 0.05 0.70 0.14 0.38
calibrated probability
raw logits
unmask ratio 𝜶𝒕
logit
0.09 0.33 0.64 0.42 0.77 0.25 0.53 0.89
masked
decides which probability distribution to use Early stage (high mask ratio)
Halton priority map
logit
Masked State
(c) Mask-ratio-aware probability calibration
Temperature
- more context - slightly more confident token index
𝜀𝑡 = 𝜀min + (𝜀max − 𝜀min )
ℳ (𝑡) 𝑁
𝛾
Fig. 3. Overview of the improved denoising strategy: (a) Halton-guided denoising position selection; (b) Cosine reverse denoising schedule; (c) Mask-ratioaware probability calibration.
Original Image
where hg ∈ {1, · · · , H}, wg ∈ {1, · · · , W }, and cg ∈ {1, · · · , C}. The corresponding flattened token position is ig = (hg −1)W +(wg −1) C +cg , ig ∈ {1, · · · , N }. (24)
π (t) = TopKi∈M(t) (ri , kt ) ,
3.0
Entropy
2.5
Halton
2.0 1.5 1.0 4.5 4.0 3.5 3.0
Random
(25)
2.5 2.0
(t)
where π (t) = (i1 , · · · , ikt ). In implementation, already denoised and non-maskable <bos> token are assigned −∞ before sorting, so only positions in M(t) can be selected. Fig. 4 visualizes the remaining-mask entropy at the middle denoising stage. Compared with confidence-based MaskGITstyle selection [44], the Halton priority map is independent of the model logits and therefore avoids repeatedly selecting only the easiest local regions. Compared with independent pseudorandom priorities, the Halton priority map provides a quasirandom but more evenly distributed ordering over the patch. In this sense, Halton-guided denoising is a low-discrepancy derandomization of randomized denoising: it keeps the nonconfidence-driven property while producing a more uniformly covered denoising trajectory. C. Mask-Ratio-Aware Denoising Schedule After determining where to denoise by the Halton rule (i.e., π (t) ), the codec still needs to decide how many tokens to process (kt ) and which probability table to use. We address these issues with a cosine schedule and mask-ratio-aware calibration, inspired by masked token generation and discrete diffusion models [40], [44]. 1) Cosine Schedule: Recalling that M(t) denotes the set of positions that remain masked before denoising step t, while N is the number of maskable image tokens. A simple linear (t) schedule ktlin = |Mt | processes an approximately uniform fraction of the remaining masked tokens over the remaining denoising steps. This rule guarantees completion at t = 1,
1.5
5.0 4.5 4.0
Entropy
(t)
3.5
Entropy
Because multiple Halton points may be quantized to the same discrete position, we scan g = 1, 2, · · · but keep only the first visit to each position. For the position ig , which is first visited by the g-th retained Halton point, the priority is assigned as rig = N −g +1. Then a deterministic ordering of all maskable positions can be obtained, where earlier visited positions receive larger priorities. At denoising step t, the denoising list is obtained by selecting the top-kt active positions:
Entropy on Remaining Masks (t=T/2)
Confidence
3.5 3.0 2.5 2.0
Fig. 4. Visualization of remaining-mask entropy at the middle denoising stage for different denoising-position selection rules.
but it does not distinguish sparse early contexts from richer late contexts [44]. We therefore resort to a mask-ratio-aware cosine schedule. Without loss of generality, let st = T −t+1 , T ∀t = T, T −1, · · · , 1, and the target cumulative denoising ratio after denoising step t is defined as ( 1 − cos st · π2 , t = T, T − 1, · · · , 1 αt = . (26) 0, t = T + 1, where αT +1 = 0 for notational convenience. Accordingly, the number of newly processed tokens at denoising step t is kt = N (αt − αt+1 ) ,
(27)
except that the final step uses k1 = |M(1) | to guarantee complete restoration.
8
Algorithm 1 Domain-Adapted Fine-Tuning of DiffuGPT Input: Image mini-batches, pretrained DiffuGPT fθ , pixeltoken dictionary D, mask token m, maximum forward step T. Output: Fine-tuned model fθ 1: for r = 1 to epochs do 2: Tokenize each image patch as x = [x0 , x1 , · · · , xN ], where x0 = <bos> and xi ∈ D for i ≥ 1. 3: Sample a forward diffusion step τ ∈ {1, · · · , T } and compute the corresponding mask ratio µτ = 1 − ᾱτ . 4: Form a corrupted sequence x̃(τ ) by replacing each image-token position with m independently with probability µτ ; keep x0 unchanged. (τ ) 5: Let M(τ ) = {ℓ ∈ {1, · · · , N } : x̃ℓ = m} be the masked image-token set. 6: if M(τ ) = ∅ then 7: Continue. 8: end if 9: Compute shifted pixel-token logits Z(τ ) ← Shift(fθ (x̃(τ ) ))D by Eq. (18). 10: Compute Lft (θ) by Eq. (30) over M(τ ) and update θ. 11: end for 12: return fθ As shown in Fig. 3(b), this cosine schedule is conservative at the beginning because αt grows slowly when t is close to T . Only a few anchor tokens are denoised when the available context is limited. As t decreases, more tokens have been restored and the context becomes richer; the schedule then increases the step size and processes more tokens per denoising step. Thus, the cosine schedule improves early-stage probability reliability while retaining the jump-step efficiency of diffusion coding. 2) Mask Ratio-Aware Calibration: Since the accuracy of diffusion predictions changes markedly with the remaining mask ratio, using the raw logits at every denoising step is suboptimal for arithmetic coding, and temperature scaling provides a simple way to correct this kind of probability miscalibration [48]. We therefore re-write the distribution for (t) the j-th denoising position ij in Eq. (20) as (t) exp(z (t) /εt ) ij ,d (t) pj (d) = P , (t) d′ ∈D exp(zi(t) ,d′ /εt )
(28)
j
where the mask-ratio-aware temperature is defined as γ |M(t) | εt = εmin + (εmax − εmin ) , (29) N where εmin , εmax , and γ are shared hyperparameters. As shown in Fig. 3(c), when many tokens remain masked, εt > 1 softens overconfident predictions and reduces high-probability assignments in arithmetic coding. When only a few tokens remain masked, εt < 1 sharpens the distribution and improves compression efficiency. D. Fine-Tuning and Inference Procedures 1) Domain-Adapted Fine-Tuning: During fine-tuning, the pretrained DiffuGPT backbone is adapted to image-token
Algorithm 2 Synchronized Diffusion Arithmetic Coding and Decoding Input: Fine-tuned model fθ , pixel dictionary D, mask token m, Halton map H, steps T , calibration parameters (εmin , εmax , γ), arithmetic coder Input: Encoder input: clean sequence x = [x0 , x1 , · · · , xN ]; decoder input: bitstream Output: Encoder output: bitstream; decoder output: reconstructed sequence x̂ 1: Initialize the synchronized state u(T ) ← [x0 , m, · · · , m] and masked set M(T ) ← {1, · · · , N }. 2: for t = T, T − 1, · · · , 1 while M(t) ̸= ∅ do 3: Compute shifted pixel-token logits Z(t) ← Shift(fθ (u(t) ))D by Eq. (18). 4: Determine kt by Eq. (27), clipped to 1 ≤ kt ≤ |M(t) |, with k1 = |M(1) |. 5: Select the ordered denoising list π (t) = TopKℓ∈M(t) (rℓ , kt ) by Eq. (25). 6: Compute the remaining mask ratio |M(t) |/N and the temperature εt by Eq. (29). 7: for each position ℓ ∈ π (t) do (t) (t) 8: Form pℓ (d) = softmax(zℓ,d /εt ) over d ∈ D by Eq. (28). (t) 9: Compute the distribution function Fℓ (d) = P (t) ′ d′ ≺d pℓ (d ) under the fixed ordering of D. 10: if encoder mode then 11: Arithmetic-encode the ground-truth token xi us(t) (t) ing Fℓ (xℓ ) and pℓ (xℓ ) as in Eq. (3), and set (t−1) uℓ ← xℓ . 12: else 13: Arithmetic-decode x̂ℓ from the current interval (t) (t) using Fℓ (·) and pℓ (·) as in Eq. (3), and set (t−1) uℓ ← x̂ℓ . 14: end if 15: end for 16: Copy all unprocessed positions from u(t) to u(t−1) and set M(t−1) ← M(t) \ π (t) . 17: end for 18: Output the terminated bitstream in encoder mode, or x̂ ← u(0) in decoder mode. denoising by masking pixel-token positions and keeping the <bos> token fixed, applying the same logit-shift operation that will be used during inference, and computing the loss only over the 256 valid pixel tokens. Since DiffuGPT is pretrained with a next-token objective, the hidden state at position ℓ−1 predicts token xℓ . Thus, for a corrupted sequence x̃(τ ) generated at sampled forward step τ , we shift the raw logits by one position and restrict the output vocabulary to the pixel-token dictionary D as in Eq. (18). The sampled step τ determines the training mask ratio µτ = 1 − ᾱτ , which is consistent with the absorbing forward marginal in Eq. (8). Only image-token positions are randomly replaced by the mask token m, while <bos> is kept unchanged. Let (τ ) M(τ ) = {ℓ ∈ {1, · · · , N } | x̃ℓ = m} denote the masked image-token positions. For each ℓ ∈ M(τ ) , the training (τ ) distribution pθ (d | x̃(τ ) , ℓ) is given by Eq. (28). The fine-
9
TABLE III K EY PARAMETER S ETTINGS FOR THE S IMULATION AND E XPERIMENTS
Parameter Description
Value
Source Coder Training (DiffuGPT) Base architecture Training dataset Vocabulary size (original → valid pixel tokens) Patch size (spatial resolution) Learning rate Effective batch size Number of training epochs
GPT-2 DIV2K-HR 50,257 → 256 16 × 16 3 × 10−4 256 3
Synchronized Diffusion Inference Inference diffusion steps T Arithmetic coder precision Temperature scaling bounds [εmin , εmax ] Temperature scaling exponent γ
20 (default) 32 bits [0.9, 1.2] 1.5
Channel Protection (ECCT) 10−3 256 6 128 8 1/2
Learning rate Batch size Number of encoder layers Dimension of embedding Number of attention heads Coding rate Evaluation Protocol Test datasets
CIFAR10, DIV2K-LRX4, Kodak AWGN, Rayleigh PSNR, SSIM
Channel models Quality metrics
tuning objective is then written as X w(µτ ) (τ ) Lft (θ) = Eτ,x − log pθ (xℓ | x̃(τ ) , ℓ) , |M(τ ) | (τ ) ℓ∈M
(30) where w(µτ ) is the diffusion-rate weight. Mini-batches with no valid masked pixel tokens are ignored. The complete finetuning procedure is summarized in Algorithm 1. 2) Synchronized Inference: At inference time, compression and decompression use the same masked state u(t) , masked set M(t) , denoising list π (t) , schedule kt , temperature εt , shifted (t) logits Z(t) , and normalized probability tables pi . For each (t) selected position i ∈ π , the encoder arithmetically encodes (t) the ground-truth pixel value xi ∈ D according to pi (d), whereas the decoder decodes x̂i ∈ D from the bitstream using (t) the same pi (d). Since π (t) is generated by the shared Haltonguided denoising rule in Eq. (19), all shared parameters, including the Halton priority map, denoising rule, denoising schedule, and calibration parameters, must be identical at both sides. Algorithm 2 summarizes the inference procedures. IV. E XPERIMENTS A. Experimental Settings We evaluate the proposed DDM-SSCC on CIFAR10, DIV2K-LR-X4 (validation) and Kodak [49]–[51], which re-
spectively represent low-resolution and higher-resolution image transmission scenarios. For both datasets, the source coder follows the patch-wise processing in Section II-B, and the final image is obtained by concatenating all patch-level bitstreams and reassembling the decoded patches. To focus on the source-coding effect, all digital schemes use the same channel-protection backend, namely a rate-1/2 channel code together with an ECCT-enhanced receiver [11]. All schemes are tested under both additive white Gaussian noise (AWGN) and Rayleigh channels. We primarily compare DDM-SSCC with four representative baselines: deep lossy plus residual coding (DLPR)+ECCT [31], JPEG-XL+ECCT [25], SparseSBC [2], and DeepJSCC [1]. To highlight the difference between one-token causal decoding and the proposed multi-token reverse restoration, we further include an iGPT-based autoregressive source coder, denoted by LVMSSCC [37]. Mainly used hyperparameters are summarized in Table III. The reconstruction quality is measured by peak signal-tonoise ratio (PSNR) and structural similarity (SSIM). Since exact reconstruction gives zero MSE and hence an infinite PSNR under the standard definition, we report such cases as PSNR = 100 dB only for finite numerical display [52]. To compare heterogeneous transmission schemes under the same communication budget, the horizontal axis is the unified signal-to-noise ratio (SNR) [46], denoted by SNRunified . In other words, SNRunified aligns the same total communication budget. Meanwhile, SNRunified serves as the experimental instantiation of the prescribed budget constraint Bcomm (η, m) ≤ B0 in Eq.(16): under a fixed total energy and reference channel-use budget, a source coder that produces a shorter protected bitstream can operate at a more favorable physical SNR. B. Experimental Results 1) Main Results: Figs. 5 and 6 compare the end-toend transmission performance of different schemes. Overall, DDM-SSCC achieves the most favorable performance. On three datasets under AWGN, DDM-SSCC reaches perfect reconstruction when SNRunified is larger than 2 dB, a smaller threshold than the requirement for DLPR+ECCT and JPEGXL+ECCT. Comparatively, the semantic communication baselines show a different behavior and remain far from the exactrecovery region. In other words, compared with semanticoriented transmission, DDM-SSCC is much better aligned with fidelity-sensitive pixel-level delivery. A similar phenomenon is also observed for the Rayleigh fading channel. The advantage becomes even more pronounced on DIV2K-LR-X4 and Kodak, where high-resolution textures impose stronger requirements on probability modeling. Notably, compared with the iGPT-based autoregressive source coder LVM-SSCC , the proposed DDM-SSCC consistently moves the recovery threshold to a lower SNR region, validating the advantage of bidirectional masked restoration over one-token causal decoding for pixel-level lossless transmission. The qualitative results in Figs. 7 and 8 are consistent with the quantitative curves. At SNRunified = 0 dB, DDMSSCC is already close to perfect reconstruction: the image
10
PSNR (dB)
CIFAR10
DIV2K-LR-X4
Kodak
100
100
100
80
80
80
60
60
60
40
40
40
20
20
20
0
2
4
6
8
10
1.0
0
2
4
6
8
10
1.0 0.9
SSIM
0.8
0
2
0
2
4
6
8
10
4
6
8
10
1.0 0.9
0.8
0.8
0.7
0.6
0.7
0.6 0.4
0.2
0.5
0.6
0.4
0.5
0.3 0
2
4
6
SNRunified (dB)
8
10
DDM-SSCC
0.4 0
2
LVM-SSCC
4
6
SNRunified (dB) DLPR+ECCT
8
JPEG-XL+ECCT
10
SparseSBC
SNRunified (dB)
DeepJSCC
Fig. 5. Main performance comparison on CIFAR10, DIV2K-LR-X4 and Kodak under AWGN channels. The upper row reports PSNR and the lower row reports SSIM. Due to the awfully high running-time, the result of LVM-SSCC for DIV2K-LR-X4 (validation) and Kodak can not be obtained at a reasonable amount time (several hours on NVIDIA GeForce RTX 4090 for a single image, consistent with the results in [7], [9]), and omitted here.
PSNR (dB)
CIFAR10
DIV2K-LR-X4
Kodak
100
100
100
80
80
80
60
60
60
40
40
40
20
20
20
2
4
6
8
10
12
2
4
6
8
10
12
1.0
1.0
2
4
6
8
10
12
6
8
10
12
0.9
0.8
0.8
0.8
SSIM
4
1.0
0.9
0.9
2
0.7 0.7
0.6
0.6
0.5
0.7 0.6 0.5
0.4
0.5
0.3
0.4
0.4 2
4
6
8
SNRunified (dB)
10
12
DDM-SSCC
2
LVM-SSCC
4
6
8
SNRunified (dB) DLPR+ECCT
JPEG-XL+ECCT
10
12
SparseSBC
SNRunified (dB)
DeepJSCC
Fig. 6. Main performance comparison on CIFAR10, DIV2K-LR-X4 and Kodak under Rayleigh fading channels. The upper row reports PSNR and the lower row reports SSIM. Due to the awfully high running-time, the result of LVM-SSCC for DIV2K-LR-X4 (validation) and Kodak can not be obtained at a reasonable amount time (several hours on NVIDIA GeForce RTX 4090 for a single image, consistent with the results in [7], [9]), and omitted here.
is visually almost lossless. By contrast, DLPR+ECCT reconstructs the early part of the image well, but once bit errors appear, its lossy-plus-residual decoding mechanism causes the subsequent patches to fail, which leads to the severe corruption visible on the right side of the reconstructed image. SparseSBC and DeepJSCC still provide visually acceptable reconstructions, with SparseSBC being clearly better; however,
both remain far from lossless recovery in fine textures and local details. 2) Impact of Denoising Steps: Fig. 9 shows the performance–complexity tradeoff controlled by tuning the number of diffusion steps. More denoising steps (T ) generally improve source modeling and move the exact-recovery threshold to lower SNRunified values under both AWGN and Rayleigh fading channels. The T = 50 setting performs best,
11
DIV2K-LR-X4 (Validation) Avg. Resolution: 256×512 DDM-SSCC 41.7 dB / 0.99
DLPR+ECCT 17.5 dB / 0.68
SparseSBC 33.5 dB / 0.91
DeepJSCC 25.3 dB / 0.76
Ground Truth PSNR / SSIM
Fig. 7. Representative reconstruction example at SNRunified = 0 dB on DIV2K-LR-X4. Kodak24 Avg. Resolution: 768×512
Ground Truth PSNR / SSIM
DDM-SSCC 49.4 dB / 0.99
DLPR+ECCT 18.7 dB / 0.68
SparseSBC 35.2 dB / 0.96
DeepJSCC 26.7 dB / 0.78
Fig. 8. Representative reconstruction example at SNRunified = 0 dB on Kodak.
while the aggressive T = 5 setting still maintains competitive quality with much lower decoding complexity. 3) Contribution of Halton-Guided Denoising: Fig. 10 compares Halton-guided denoising with random and confidencedriven position selection. Halton-guided denoising reaches the exact-recovery region earlier and provides better pre-threshold PSNR/SSIM, especially under Rayleigh fading. This verifies that the denoising order affects arithmetic-coding probability quality, and that the deterministic low-discrepancy order can provide more spatially dispersed context anchors. 4) Effectiveness of Cosine denoising schedule: Fig. 11 validates the cosine denoising schedule. Compared with the linear schedule, it achieves a lower exact-recovery threshold under both channel models, with a clearer advantage under Rayleigh fading. This confirms that conservative early updates and more aggressive late updates improve probability reliability while preserving jump-step efficiency. 5) Ablation Studies: Table IV further validates the component-wise effectiveness of Halton-guided denoising, cosine scheduling, and mask-ratio-aware calibration. The three modules respectively improve spatial coverage, adapt the denoising pace to context reliability, and mitigate stage-
TABLE IV A BLATION STUDY OF H ALTON SAMPLING , COSINE SCHEDULING , AND CALIBRATION . AWGN, SNRunified = 0 dB
Components Halton
Cosine
Calib.
PSNR (dB) ↑
SSIM ↑
× × × × ✓ ✓ ✓ ✓
× × ✓ ✓ × × ✓ ✓
× ✓ × ✓ × ✓ × ✓
19.35 22.33 24.42 24.00 22.28 21.15 25.30 26.85
0.8739 0.9104 0.9359 0.9283 0.9202 0.9067 0.9462 0.9581
dependent probability mismatch. All variants share the same high-SNR lossless ceiling, while their pre-threshold differences reflect the reliability of source-bitstream delivery under the fixed channel-protection backend. V. C ONCLUSION This paper has studied lossless pixel-level image transmission from a beyond-semantics perspective. We have pro-
12
80
80
PSNR (dB)
100
PSNR (dB)
100
60
40
60 40 20
20 0
2
4
6
SNRunified (dB)
8
10
0
2
4
6
8
10
0
2
4
6
8
10
SNRunified (dB)
1.0
1.00
0.8
0.95
SSIM
SSIM
0.90 0.85 0.80 0.75
0.6 0.4 0.2
0.70
0.0
0.65 0
2
4
6
SNRunified (dB)
DDM-SSCC (T=50, AWGN) DDM-SSCC (T=20, AWGN) DDM-SSCC (T=10, AWGN) DDM-SSCC (T=5, AWGN)
8
10
SNRunified (dB)
Cosine (Ours), AWGN Linear, AWGN
DDM-SSCC (T=50, Rayleigh) DDM-SSCC (T=20, Rayleigh) DDM-SSCC (T=10, Rayleigh) DDM-SSCC (T=5, Rayleigh)
Cosine (Ours), Rayleigh Linear, Rayleigh
Fig. 11. Ablation study of the reverse denoising schedule on CIFAR10. Fig. 9. Impact of the diffusion step number on CIFAR10.
0.6
mask-ratio-aware cosine schedule, and temperature calibration for diffusion source coding. These modules respectively reduce confidence-induced spatial clustering, match the denoising pace to context reliability, and alleviate stage-dependent probability miscalibration in the reverse process. Experiments on CIFAR10, DIV2K-LR-X4, and Kodak over AWGN and Rayleigh fading channels show that DDM-SSCC consistently provides a more favorable source-channel operating point than representative lossless, autoregressive, and semantic baselines, with earlier entry into the exact-recovery regime and stronger pre-threshold fidelity. Ablation studies further confirm the component-wise benefits of the proposed denoising order, schedule, and calibration strategy. These results suggest that discrete diffusion is a promising direction for future pixel-level communication systems and can serve as a stronger foundation for subsequent patch-level or higher-level extensions.
0.4
R EFERENCES
100
PSNR (dB)
80 60 40 20 0
2
4
6
SNRunified (dB)
8
10
1.0
SSIM
0.8
0.2 0
2
4
6
SNRunified (dB)
Halton (Ours), AWGN Random, AWGN Confidence, AWGN
8
10
Halton (Ours), Rayleigh Random, Rayleigh Confidence, Rayleigh
Fig. 10. Ablation study of the denoising-position selection rule on CIFAR10.
posed DDM-SSCC, a discrete-diffusion-model-based SSCC framework that turns bidirectional masked restoration into an arithmetic-coding-compatible source codec. Beyond the basic synchronized diffusion codec, we further clarify the codelength and mutual-information perspective of jump-step denoising, and develop Halton-guided position selection, a
[1] E. Bourtsoulatze, D. B. Kurka, and D. Gündüz, “Deep joint sourcechannel coding for wireless image transmission,” IEEE Trans. Cogn. Commun. Netw., vol. 5, no. 3, pp. 567–579, Sep. 2019. [2] S. Tong, X. Yu, R. Li, K. Lu, Z. Zhao, and H. Zhang, “Alternate learningbased SNR-adaptive sparse semantic visual transmission,” IEEE Trans. Wireless Commun., vol. 24, no. 2, pp. 1737–1752, Feb. 2025. [3] W. Zhang, K. Bai, S. Zeadally, H. Zhang, H. Shao, H. Ma, and V. C. M. Leung, “DeepMA: End-to-end deep multiple access for wireless image transmission in semantic communication,” IEEE Trans. Cogn. Commun. Netw., vol. 10, no. 2, pp. 387–402, Apr. 2024. [4] Y. Jia, Z. Huang, K. Luo, and W. Wen, “Lightweight joint source-channel coding for semantic communications,” IEEE Commun. Lett., vol. 27, no. 12, pp. 3161–3165, Dec. 2023. [5] G. Delétang, A. Ruoss, P.-A. Duquenne, E. Catt, T. Genewein, C. Mattern, J. Grau-Moya, L. K. Wenliang, M. Aitchison, L. Orseau, M. Hutter, and J. Veness, “Language modeling is compression,” in Proc. Int. Conf. Learn. Represent. (ICLR), Vienna, Austria, May 2024. [6] Y. Huang, J. Zhang, Z. Shan, and J. He, “Compression represents intelligence linearly,” Apr. 2024, arXiv preprint arXiv:2404.09937.
13
[7] Z. Li, C. Huang, X. Wang, H. Hu, C. Wyeth, D. Bu, Q. Yu, W. Gao, X. Liu, and M. Li, “Lossless data compression by large models,” Nat. Mach. Intell., vol. 7, no. 5, pp. 794–799, May 2025. [8] J. Du, C. Zhou, N. Cao, G. Chen, Y. Chen, Z. Cheng, L. Song, G. Lu, and W. Zhang, “Large language model for lossless image compression with visual prompts,” Feb. 2025, arXiv preprint arXiv:2502.16163. [9] K. Chen, P. Zhang, H. Liu, J. Liu, Y. Liu, J. Huang, S. Wang, H. Yan, and H. Li, “Large language models for lossless image compression: Next-pixel prediction in language space is all you need,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), San Diego, USA, Dec. 2025. [10] T. Ren, R. Li, M. min Zhao, X. Chen, G. Liu, Y. Yang, Z. Zhao, and H. Zhang, “Separate source channel coding is still what you need: An LLM-based rethinking,” ZTE Commun., vol. 23, no. 1, pp. 30–44, Mar. 2025. [11] Y. Choukroun and L. Wolf, “Error correction code transformer,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), New Orleans, USA, Nov. 2022, pp. 38 695–38 705. [12] ——, “A foundation model for error correction codes,” in Proc. Int. Conf. Learn. Represent. (ICLR), Vienna, Austria, May 2024. [13] D.-T. Nguyen and S. Kim, “U-shaped error correction code transformers,” IEEE Trans. Cogn. Commun. Netw., vol. 11, no. 2, pp. 975–986, Apr. 2025. [14] Z. Lu, R. Li, M. Lei, C. Wang, Z. Zhao, and H. Zhang, “Selfcritical alternate learning-based semantic broadcast communication,” IEEE Trans. Commun., vol. 73, no. 5, pp. 3347–3361, May 2025. [15] P. Jiang, C.-K. Wen, X. Yi, X. Li, S. Jin, and J. Zhang, “Semantic communications using foundation models: Design approaches and open issues,” IEEE Wireless Commun., vol. 31, no. 3, pp. 76–84, Jun. 2024. [16] F. Jiang, Y. Peng, L. Dong, K. Wang, K. Yang, C. Pan, and X. You, “Large AI model-based semantic communications,” IEEE Wireless Commun., vol. 31, no. 3, pp. 68–75, Jun. 2024. [17] C. Liang, H. Du, Y. Sun, D. Niyato, J. Kang, D. Zhao, and M. A. Imran, “Generative AI-driven semantic communication networks: Architecture, technologies and applications,” IEEE Trans. Cogn. Commun. Netw., vol. 11, no. 1, pp. 27–47, Feb. 2025. [18] F. Jiang, L. Dong, Y. Peng, K. Wang, K. Yang, C. Pan, and X. You, “Large AI model empowered multimodal semantic communications,” IEEE Commun. Mag., vol. 63, no. 1, pp. 76–82, Jan. 2025. [19] K. Tan, J. Dai, S. Wang, G. Lu, S. Shao, K. Niu, W. Zhang, and P. Zhang, “DiT-JSCC: Rethinking deep JSCC with diffusion transformers and semantic representations,” 2026, arXiv preprint arXiv:2601.03112. [20] R. G. Gallager, “Low-density parity-check codes,” IRE Trans. Inf. Theory, vol. 8, no. 1, pp. 21–28, Jan. 1962. [21] E. Arikan, “Channel polarization: A method for constructing capacityachieving codes for symmetric binary-input memoryless channels,” IEEE Trans. Inf. Theory, vol. 55, no. 7, pp. 3051–3073, Jul. 2009. [22] E. Nachmani, E. Marciano, L. Lugosch, W. J. Gross, D. Burshtein, and Y. Be’ery, “Deep learning methods for improved decoding of linear codes,” IEEE J. Sel. Top. Signal Process., vol. 12, no. 1, pp. 119–131, Feb. 2018. [23] X. Wu and N. Memon, “Context-based, adaptive, lossless image coding,” IEEE Trans. Commun., vol. 45, no. 4, pp. 437–444, Apr. 1997. [24] M. J. Weinberger, G. Seroussi, and G. Sapiro, “The LOCO-I lossless image compression algorithm: Principles and standardization into JPEGLS,” IEEE Trans. Image Process., vol. 9, no. 8, pp. 1309–1324, Aug. 2000. [25] J. Alakuijala, R. van Asseldonk, S. Boukortt, M. Bruse, I.-M. Comsa, M. Firsching, T. Fischbacher, S. Gomez, E. Kliuchnikov, R. Obryk, K. Potempa, A. Rhatushnyak, J. Sneyers, Z. Szabadka, L. Vandevenne, L. Versari, and J. Wassenberg, “JPEG XL next-generation image compression architecture and coding tools,” in Proc. SPIE Appl. Digit. Image Process. XLII, vol. 11137, San Diego, USA, Sep. 2019, p. 111370K. [26] F. Mentzer, E. Agustsson, M. Tschannen, R. Timofte, and L. V. Gool, “Practical full resolution learned lossless image compression,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Long Beach, USA, Jun. 2019, pp. 10 629–10 638. [27] E. Hoogeboom, J. W. T. Peters, R. van den Berg, and M. Welling, “Integer discrete flows and lossless compression,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 32, Vancouver, Canada, Dec. 2019, pp. 12 134–12 144. [28] N. Kang, S. Qiu, S. Zhang, Z. Li, and S.-T. Xia, “PILC: Practical image lossless compression with an end-to-end GPU oriented neural framework,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), New Orleans, USA, Jun. 2022, pp. 3739–3748. [29] Z. Zhang, H. Wang, Z. Chen, and S. Liu, “Learned lossless image compression based on bit plane slicing,” in Proc. IEEE/CVF Conf.
Comput. Vis. Pattern Recognit. (CVPR), Seattle, USA, Jun. 2024, pp. 27 579–27 588. [30] D. Li, Y. Bai, K. Wang, J. Jiang, X. Liu, and W. Gao, “CALLIC: Content adaptive learning for lossless image compression,” in Proc. AAAI Conf. Artif. Intell., vol. 39, no. 5, Philadelphia, USA, Feb. 2025, pp. 4679– 4688. [31] Y. Bai, X. Liu, K. Wang, X. Ji, X. Wu, and W. Gao, “Deep lossy plus residual coding for lossless and near-lossless image compression,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 46, no. 5, pp. 3577–3594, May 2024. [32] Z. Zhang, Z. Chen, and S. Liu, “Fitted neural lossless image compression,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), Nashville, USA, Jun. 2025, pp. 23 249–23 258. [33] J. J. Rissanen, “Generalized Kraft inequality and arithmetic coding,” IBM J. Res. Dev., vol. 20, no. 3, pp. 198–203, May 1976. [34] I. H. Witten, R. M. Neal, and J. G. Cleary, “Arithmetic coding for data compression,” Commun. ACM, vol. 30, no. 6, pp. 520–540, Jun. 1987. [35] P. G. Howard and J. S. Vitter, “Arithmetic coding for data compression,” Proc. IEEE, vol. 82, no. 6, pp. 857–865, Jun. 1994. [36] C. S. K. Valmeekam, K. Narayanan, D. Kalathil, J.-F. Chamberland, and S. Shakkottai, “LLMZip: Lossless text compression using large language models,” Jun. 2023, arXiv preprint arXiv:2306.04050. [37] M. Chen, A. Radford, R. Child, J. Wu, H. Jun, D. Luan, and I. Sutskever, “Generative pretraining from pixels,” in Proc. Int. Conf. Mach. Learn. (ICML), ser. Proc. Mach. Learn. Res., vol. 119, Virtual Edition, Jul. 2020, pp. 1691–1703. [38] J. Austin, D. D. Johnson, J. Ho, D. Tarlow, and R. van den Berg, “Structured denoising diffusion models in discrete state-spaces,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 34, Virtual Edition, Dec. 2021, pp. 17 981–17 993. [39] X. L. Li, J. Thickstun, I. Gulrajani, P. Liang, and T. B. Hashimoto, “Diffusion-LM improves controllable text generation,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 35, New Orleans, USA, Dec. 2022, pp. 4328–4343. [40] S. S. Sahoo, M. Arriola, Y. Schiff, A. Gokaslan, E. Marroquin, J. T. Chiu, A. M. Rush, and V. Kuleshov, “Simple and effective masked diffusion language models,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 37, Vancouver, Canada, Dec. 2024. [41] M. Arriola, A. Gokaslan, J. T. Chiu, Z. Yang, Z. Qi, J. Han, S. S. Sahoo, and V. Kuleshov, “Block diffusion: Interpolating between autoregressive and diffusion language models,” in Proc. Int. Conf. Learn. Represent. (ICLR), Singapore, Apr. 2025. [42] S. Nie, F. Zhu, Z. You, X. Zhang, J. Ou, J. Hu, J. Zhou, Y. Lin, J.R. Wen, and C. Li, “Large language diffusion models,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), San Diego, USA, Dec. 2025. [43] S. Gong, S. Agarwal, Y. Zhang, J. Ye, L. Zheng, M. Li, C. An, P. Zhao, W. Bi, H. Peng, J. Han, and L. Kong, “Scaling diffusion language models via adaptation from autoregressive models,” in Proc. Int. Conf. Learn. Represent. (ICLR), Singapore, Apr. 2025. [44] H. Chang, H. Zhang, L. Jiang, C. Liu, and W. T. Freeman, “MaskGIT: Masked generative image transformer,” in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), New Orleans, USA, Jun. 2022, pp. 11 315–11 325. [45] V. Besnier, M. Chen, D. Hurych, E. Valle, and M. Cord, “Halton scheduler for masked generative image transformer,” in Proc. Int. Conf. Learn. Represent. (ICLR), Singapore, Apr. 2025. [46] T. M. Cover and J. A. Thomas, Elements of Information Theory, 2nd ed. Hoboken, NJ, USA: John Wiley & Sons, 2006. [47] J. H. Halton and G. B. Smith, “Algorithm 247: Radical-inverse quasirandom point sequence,” Commun. ACM, vol. 7, no. 12, pp. 701–702, Dec. 1964. [48] C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, “On calibration of modern neural networks,” in Proc. Int. Conf. Mach. Learn. (ICML), ser. Proc. Mach. Learn. Res., vol. 70, Sydney, Australia, Aug. 2017, pp. 1321–1330. [49] A. Krizhevsky, “Learning multiple layers of features from tiny images,” University of Toronto, Toronto, ON, Canada, Tech. Rep., 2009. [50] E. Agustsson and R. Timofte, “NTIRE 2017 challenge on single image super-resolution: Dataset and study,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. Workshops (CVPRW), Honolulu, USA, Jul. 2017, pp. 1122–1131. [51] R. Franzen, “Kodak lossless true color image suite,” Available online: http://r0k.us/graphics/kodak/, 1999. [52] A. Benazza-Benyahia, J.-C. Pesquet, and H. Masmoudi, “Blockbased adaptive vector lifting schemes for multichannel image coding,” EURASIP J. Image Video Process., vol. 2007, pp. 1–12, 2007.