1
Storage-Scalable Progressive Semantic Communication via Knowledge-Base Reuse
arXiv:2609.10112v1 [cs.LG] 9 Sep 2026
Heng Zhu, Ye Liu, Kun Zhu, Member, IEEE, Feifei Song
Abstract—Existing knowledge-base-assisted semantic communication schemes commonly adopt either single knowledge-base quantization (SKBQ) or multi-knowledge-base residual quantization (MKBQ). SKBQ incurs limited storage overhead but has restricted quantization capacity, whereas MKBQ supports progressive refinement by assigning an independent knowledge base (KB) to each stage, causing the KB storage to grow linearly with the transmission depth. To address this problem, we propose storage-scalable knowledge-base reuse quantization (SSKBQ), which reuses a compact set of KBs across multiple residual refinement stages and thereby decouples the number of transmission stages from the number of maintained KBs. A stageaware residual supervision mechanism is further introduced to regularize intermediate quantized representations and encourage progressive refinement. Experimental results demonstrate that KB reuse provides an effective solution to the storage scalability problem while maintaining competitive progressive reconstruction performance. Index Terms—Semantic communication, progressive image reconstruction, vector quantization, knowledge-base reuse, storage scalability.
I. I NTRODUCTION EMANTIC communication has recently emerged as a promising paradigm and has attracted extensive research interest [1]–[3]. Unlike conventional communication systems that pursue bit-level fidelity, semantic communication focuses on delivering task-relevant information. By removing taskirrelevant redundancy, it can significantly reduce transmission rates while preserving task performance. Despite these advantages, early studies reveal a fundamental challenge: without carefully designed encoding and decoding mechanisms, semantic communication may even require higher transmission rates than conventional schemes. This issue largely stems from the use of deep neural networks for semantic feature extraction, where integer-valued inputs are transformed into high-dimensional floating-point representations. According to the IEEE 754 standard [4], a doubleprecision floating-point number typically occupies 64 bits, whereas an integer often requires only 8 bits. Consequently, directly transmitting semantic features requires substantial compression to maintain the same transmission cost, which is often impractical. To alleviate this issue, quantization has been introduced into semantic encoding. A representative approach is SKBQ [5], where semantic features are mapped to the nearest codewords in a predefined KB. Instead of transmitting floatingpoint features, the corresponding integer indices are conveyed,
S
Heng Zhu, Ye Liu, Kun Zhu and Feifei Song are with the College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics, Nanjing 210016, China (e-mail: {zhuheng, sz2516087, zhukun, songff}@nuaa.edu.cn).
Quantization Module Transmitter
Receiver Semantic Encoder
Semantic Decoder
Fig. 1: End-to-End semantic communication system for image reconstruction. thereby reducing transmission rates. However, the use of a single KB limits the representation capacity and may lead to considerable quantization distortion. Moreover, it typically supports only single-stage transmission and does not naturally support progressive refinement. To enable progressive refinement, MKBQ schemes have been proposed [6]. These methods progressively quantize residual errors across multiple KBs, enabling progressive reconstruction through multi-stage transmission. However, the storage overhead grows linearly with the number of stages, since each stage requires an independent KB. When each KB is large, the cumulative storage cost becomes prohibitive. To resolve this trade-off, we propose an SSKBQ scheme for progressive semantic communication. The key idea is to decouple transmission stages from dedicated KBs by reusing a compact set of KBs across multiple residual refinement stages, thereby enabling progressive reconstruction without linear KB storage growth. In addition, a stage-aware residual supervision mechanism is designed to regularize intermediate quantized representations and encourage progressive refinement across stages. Experimental results show that the proposed scheme provides a favorable trade-off between progressive reconstruction performance and KB storage overhead compared with SKBQ and MKBQ.
II. S EMANTIC C OMMUNICATION S YSTEM In this work, we consider an end-to-end (E2E) semantic communication system for image reconstruction, as illustrated in Fig. 1. We focus on the semantic-layer design of storagescalable progressive transmission rather than physical-layer transmission or channel-robust modulation and coding. Let I ∈ R3×H×W denote an input image sampled from dataset D, where the three dimensions represent the RGB channels, height, and width, respectively. At the transmitter, the image is processed by a semantic encoder ϕ(·) to extract task-relevant semantic features: ze = ϕ(I),
ze ∈ Rc×h×w ,
(1)
2
where c, h, and w denote the number of channels, height, and width of the feature map, respectively. The extracted semantic feature ze is then transmitted to the receiver. However, direct transmission of ze in floating-point format incurs substantial transmission overhead. To reduce this overhead, a quantization module ν(·) is introduced to quantize ze as zq = ν(ze ),
zq ∈ Rc×h×w ,
(2)
which is then transmitted to the receiver. We assume perfect physical-layer transmission. Accordingly, channel coding, modulation, equalization, and error-control mechanisms are abstracted as a reliable delivery interface. At the receiver, the received zq is fed into the semantic decoder ϕ−1 (·) to reconstruct the image: Î = ϕ−1 (zq ),
Î ∈ R3×H×W .
(3)
KB
(a) Single Knowledge-Base Quantization
KB
......
KB
...
...
(b) Storage-Scalable Knowledge-Base Reuse Quantization
KB
KB
......
KB
KB
III. S TORAGE -S CALABLE K NOWLEDGE -BASE R EUSE S CHEME A. Knowledge-Base Reuse Scheme
(c) Multi-Knowledge-Base Residual Quantization
A straightforward implementation of the quantization module in Fig. 1 is based on vector quantization. Specifically, the quantization module ν(·) is parameterized by a learnable KB φ ∈ RN ×c , which contains N codewords of dimension c. The semantic feature ze is reshaped into hw vectors {zie }hw i=1 , each of dimension c. For each vector, the Euclidean distances to all codewords in the KB are computed, and the nearest codeword is selected as
Fig. 2: Knowledge-base quantization schemes.
ziq = φk ,
2
k = arg min zie − φj 2 , j
(4)
where φj denotes the j-th codeword. By concatenating all ziq and reshaping them into size c×h×w, the quantized semantic feature zq is obtained. As illustrated in Fig. 2(a), SKBQ requires only one KB and thus incurs limited storage overhead. However, it provides limited reconstruction performance and does not support progressive transmission. Therefore, its reconstruction quality is often insufficient for high-resolution images. To enable progressive transmission, MKBQ extends SKBQ by employing multiple homogeneous KBs. As illustrated in Fig. 2(c), the transmission process is divided into T stages. Accordingly, the quantization module ν(·) includes T KBs, denoted by {φj }Tj=1 , where φj ∈ RN ×c . The first KB generates an initial approximation z̃q = φ1 (ze ), while the remaining KBs iteratively quantize the residuals. Specifically, the first residual is r1 = ze − z̃q , and each subsequent KB Pi−1 produces r̃i = φi+1 (ri ), where ri = ze − z̃q − j=1 r̃j . After T stages, the final quantized semantic feature is zq = z̃q +
T −1 X
r̃i .
construction, resulting in unstable progressive refinement. Second, the KB storage grows linearly with the number of transmission stages because each stage requires an independent KB containing numerous codewords. Since high-resolution image reconstruction generally requires more transmission stages, the resulting storage overhead can become prohibitively high. To overcome these limitations, we propose the SSKBQ scheme, illustrated in Fig. 2(b). Unlike MKBQ, which assigns an independent KB to each transmission stage, SSKBQ reuses a compact set of KBs across multiple residual refinement steps, thereby decoupling the number of maintained KBs from the transmission depth. Let τj denote the number of residual refinement steps assigned to the j-th KB. The first KB quantizes the semantic feature ze to obtain the initial approximation z̃q and is then reused to quantize the following τ1 − 1 residuals. The second KB is subsequently reused to quantize the residuals from step τ1 + 1 to τ2 , and the remaining KBs follow the same strategy. Consequently, using only K reusable KBs (K ≤ T ), the proposed scheme supports T progressive transmission stages through KB reuse. The resulting quantized semantic feature is expressed as τj K X X zq = z̃q + r̃j,i , (6) j=1 i=1
where r̃j,i denotes the quantization result produced by the j-th KB at its i-th residual quantization step.
(5)
i=1
By adaptively selecting the number of participating KBs, MKBQ supports progressive transmission. However, two limitations remain. First, MKBQ does not guarantee monotonic refinement across transmission stages. Consequently, a laterstage reconstruction may not outperform an earlier-stage re-
B. Training Storage-Scalable Knowledge Base The semantic codec and the proposed storage-scalable KB are jointly optimized in an E2E manner. The overall objective consists of three components: the reconstruction loss Lrec , the quantization loss Lvq , and the residual supervision loss Lres . The reconstruction loss measures the distortion between the
3
reconstructed image Î and the original image I using mean square error: 2
EI∼D
Î − I
. 2
(7)
The quantization loss enforces consistency between the semantic feature ze and quantized feature zq : 2 2 (8) EI∼D α ∥ze − sg[zq ]∥2 + ∥sg[ze ] − zq ∥2 , where sg(·) denotes the stop-gradient operator. The first term updates the semantic encoder by encouraging ze to approach zq , whereas the second term updates the KBs by aligning zq with ze . The hyperparameter α balances these two effects. Although the quantization loss explicitly aligns the final quantized feature zq with the semantic feature ze , it imposes no constraint on the intermediate quantized representations generated during progressive transmission. Therefore, we introduce a residual supervision loss to explicitly regularize the stage-wise refinement process. Let the intermediate quantized feature after t refinement stages be defined as: zq (t) = z̃q +
t X
r̃i .
(9)
i=1
The residual supervision loss is formulated as # " T X Λ 2 2 EI∼D β ze − sg zq (i) 2 + sg ze − zq (i) 2 . i i=1 Similar to the quantization loss, the first term updates the semantic encoder, while the second term updates the KBs to progressively reduce the residual error. The constant Λ controls the overall strength of progressive supervision, and the weighting factor 1/i assigns larger penalties to earlier refinement stages. By progressively minimizing the stage-wise residual error, the proposed supervision explicitly regularizes intermediate quantized representations under KB reuse, encouraging successive refinement stages to gradually reduce the residual error. The overall training objective is defined as: Ltotal = Lrec + Lvq + Lres .
(10)
IV. S IMULATION R ESULTS A. Experiment Settings 1) Datasets: The Cityscapes and COCO datasets are used to evaluate the proposed SSKBQ scheme. Cityscapes contains 2,975 training images and 500 validation images of urban street scenes, whereas COCO contains more diverse objects and visual scenes. For a consistent evaluation, images from both datasets are resized to 128 × 64. 2) Evaluation Metrics: Reconstruction quality is evaluated using PSNR [7], SSIM [8], FID [9], and KID [10]. 3) Baselines: We compare SSKBQ with representative non-progressive and progressive schemes. JSCC [11] directly transmits continuous semantic features and serves as an unquantized reference. For SKBQ, the VQVAE [5] system is implemented using one-hot and Gumbel-softmax quantization.
MKBQ employs an independent KB at each transmission stage. 4) Parameter Settings: For both datasets, the input image and semantic feature dimensions are 3 × 128 × 64 and 256 × 32 × 16, respectively. Each KB contains N = 512 codewords of dimension c = 256, and hw = 512 spatial feature vectors are quantized at each stage. We consider 16stage progressive transmission and set α = β = 0.25 and Λ = 8. During testing, only the codeword indices associated with the quantized feature vectors are transmitted. The receiver retrieves the corresponding codewords from the shared KBs according to the received indices and reconstructs the quantized semantic feature. B. Complexity and Scalability SKBQ, MKBQ, and SSKBQ require O(N c), O(T N c), and O(KN c) KB storage, respectively. By reusing K KBs across T stages, SSKBQ requires only K/T of the KB storage of MKBQ. At each stage, nearest-neighbor assignment compares hw feature vectors of dimension c with N codewords, resulting in O(chwN ) transmitter-side complexity. The cumulative complexity after t stages is therefore O(tchwN ) for both MKBQ and SSKBQ. At the receiver, the transmitted indices directly identify the codewords. Index lookup, feature assembly, and residual accumulation require O(chw) operations per stage and O(tchw) after t stages, independent of N . Let Cdec denote the cost of one decoder execution. Reconstruction after all t stages requires O(tchw + Cdec ), whereas reconstruction after every stage requires O(tchw + tCdec ). Hence, progressive inference latency is mainly determined by repeated decoder executions, while KB processing grows only linearly with t. Actual wall-clock latency also depends on the hardware platform and implementation. C. Comparison with Existing Schemes To compare the proposed SSKBQ with existing semantic communication schemes, experiments are conducted on the Cityscapes and COCO datasets. The progressive reconstruction results are presented in Figs. 3 and 4, where different numbers of reusable KBs are evaluated over 16 transmission stages. Across both datasets, SSKBQ exhibits consistent progressive reconstruction trends. As more transmission stages are received, the reconstruction quality gradually improves in terms of PSNR, SSIM, FID, and KID, validating the effectiveness of KB reuse for progressive semantic refinement. Moreover, compared with MKBQ under comparable settings, SSKBQ achieves better reconstruction performance, demonstrating that the proposed stage-aware residual supervision facilitates the optimization of reused KBs across multiple refinement stages. Compared with single-stage quantization schemes, including VQVAE(OH) and VQVAE(GS), SSKBQ achieves consistent improvements on both datasets, demonstrating the benefit of residual refinement. The comparison with JSCC shows dataset-dependent behavior. On Cityscapes, SSKBQ achieves
4
PSNR vs Transmission Stage
SSIM vs Transmission Stage
KB=1 KB=2 KB=4 KB=8 KB=16
15
1
3
5
7
9
11
ResVQVAE JSCC One-hot Softmax
13
0.6 KB=1 KB=2 KB=4 KB=8 KB=16
0.4
15
1
Transmission Stage
KB=1 KB=2 KB=4 KB=8 KB=16
3
5
7
9
ResVQVAE JSCC One-hot Softmax
11
13
0.4
KB=1 KB=2 KB=4 KB=8 KB=16
400
ResVQVAE JSCC One-hot Softmax
300 200
0.2
100
15
1
3
Transmission Stage
(a) PSNR
ResVQVAE JSCC One-hot Softmax
FID
20
KID
0.8
SSIM
PSNR (dB)
25
FID vs Transmission Stage
KID vs Transmission Stage 0.6
5
7
9
11
13
1
15
3
(b) SSIM
5
7
9
11
13
15
Transmission Stage
Transmission Stage
(c) KID
(d) FID
Fig. 3: Reconstruction performance comparison on the Cityscapes dataset. PSNR vs Transmission Stage 0.8
20
15
KB=1 KB=2 KB=4 KB=8 KB=16
0.4
1
3
5
7
9
11
13
15
KB=1 KB=2 KB=4 KB=8 KB=16
0.3
0.6
1
3
5
7
9
ResVQVAE JSCC One-hot Softmax
FID vs Transmission Stage
KID vs Transmission Stage
0.4
400
ResVQVAE JSCC One-hot Softmax
0.2
KB=1 KB=2 KB=4 KB=8 KB=16
300
FID
25
SSIM vs Transmission Stage
ResVQVAE JSCC One-hot Softmax
KID
KB=1 KB=2 KB=4 KB=8 KB=16
SSIM
PSNR (dB)
30
0.1
ResVQVAE JSCC One-hot Softmax
200 100
0.0
11
13
Transmission Stage
Transmission Stage
(a) PSNR
(b) SSIM
15
1
3
5
7
9
11
13
15
Transmission Stage
(c) KID
1
3
5
7
9
11
13
15
Transmission Stage
(d) FID
Fig. 4: Reconstruction performance comparison on the COCO dataset.
(a) Ground Truth
(b) Stage1(SSKBQ)
(c) Stage4(SSKBQ)
(d) Stage8(SSKBQ)
(e) Stage16(SSKBQ)
(f) JSCC
(g) VQVAE(Gumbel-Softmax)
(h) VQVAE(One-Hot)
(i) Stage8(MKBQ)
(j) Stage16(MKBQ)
Fig. 5: Image reconstruction performance under different baselines.
higher PSNR after sufficient progressive stages, whereas JSCC maintains better performance on COCO. This difference results from the interaction among dataset complexity, semantic representation capability, and codec capacity. JSCC avoids quantization distortion by transmitting continuous semantic features, which is advantageous for complex image distributions. In contrast, SSKBQ provides a more efficient progressive transmission mechanism when semantic information can be effectively organized through reusable KBs. The influence of KB number is consistent with the observations in the previous subsection. In early transmission stages, smaller KB configurations provide better performance because semantic information is more concentrated within each reusable KB. With increasing transmission stages, larger KB configurations gradually benefit from their higher representation capacity and achieve better final reconstruction quality. This result highlights the necessity of balancing representation capacity and KB scalability, which motivates the proposed KB reuse strategy. A representative reconstruction example is
shown in Fig. 5. For quantitative comparison, Tables I and II summarize the final reconstruction performance of different schemes on the Cityscapes and COCO datasets, respectively. The JSCC results are highlighted in bold as the continuous-feature transmission baseline. VQVAE(GS) and VQVAE(OH) denote the Gumbel– Softmax and one-hot quantization implementations, respectively. In MKBQ(t) and SSKBQ(t), t represents the number of received progressive transmission stages. To evaluate communication efficiency, the number of transmitted bits for one image sample is adopted as the transmission rate. Specifically, JSCC directly transmits the continuous semantic feature ze ∈ R256×32×16 , resulting in 8,388,608 transmitted bits under 64-bit floating-point representation. For MKBQ and SSKBQ, only the indices of selected codewords are transmitted. Since each KB contains N = 512 codewords, each index requires 9 bits. Therefore, the transmission rates for 1, 8, and 16 progressive stages are 4,608, 36,864, and 73,728 bits, respectively. For VQVAE(GS), transmitting the
5
TABLE I: Image Reconstruction Performance Comparison (Cityscapes). Method JSCC MKBQ(1) MKBQ(8) MKBQ(16) SSKBQ(1) SSKBQ(8) SSKBQ(16) VQVAE(GS) VQVAE(OH)
PSNR (dB) 24.24 15.69 21.62 24.58 25.28 25.15 27.36 23.16 21.22
SSIM 0.9123 0.3739 0.6912 0.7972 0.8288 0.8137 0.8739 0.6463 0.6297
FID 71.38 426.77 299.22 183.94 155.60 167.59 111.77 317.84 287.42
KID 0.0398 0.5728 0.3569 0.1831 0.1418 0.1602 0.0860 0.3793 0.3336
512-dimensional soft assignment weights for each spatial feature vector requires 16,777,216 transmitted bits. These results demonstrate that SSKBQ enables discrete-index semantic transmission with substantially reduced communication overhead compared with continuous-feature and soft-assignmentbased schemes. As shown in Table I, SSKBQ(16) achieves a PSNR of 27.36 dB on the Cityscapes dataset, improving by 11.31%, 28.93%, and 18.13% over MKBQ(16), VQVAE(OH), and VQVAE(GS), respectively. Meanwhile, SSKBQ(16) requires only 73,728 transmitted bits, corresponding to less than 1% of the transmission overhead of JSCC. Despite this substantial reduction in communication cost, SSKBQ achieves a 12.87% PSNR improvement over JSCC. However, JSCC obtains better SSIM, FID, and KID values, which can be attributed to its direct transmission of continuous semantic features and the resulting avoidance of quantization distortion. The results on the COCO dataset exhibit a different trend. As shown in Table II, SSKBQ(16) achieves a PSNR of 23.53 dB, outperforming MKBQ(16), VQVAE(OH), and VQVAE(GS) by 10.16%, 24.17%, and 27.60%, respectively. However, JSCC still achieves higher reconstruction quality on COCO due to its stronger representation capability for diverse object categories and complex visual distributions. This observation indicates that SSKBQ is not intended to universally replace continuous-feature JSCC, but rather provides a storage-efficient and progressive semantic transmission framework that achieves a favorable trade-off between communication overhead and reconstruction performance. V. C ONCLUSION This letter investigates the storage-performance trade-off in progressive semantic communication with knowledge-base quantization. Conventional single knowledge-base schemes achieve low storage overhead but suffer from limited refinement capability, while multi-knowledge-base residual schemes support progressive transmission at the cost of storage complexity that scales with the number of refinement stages. To address this issue, we propose a storage-scalable knowledgebase reuse quantization scheme that decouples the number of maintained knowledge bases from the number of progressive transmission stages. By reusing a limited number of knowledge bases across multiple refinement stages, SSKBQ reduces storage requirements while preserving progressive reconstruction capability.
TABLE II: Image Reconstruction Performance Comparison (COCO) Method JSCC MKBQ(1) MKBQ(8) MKBQ(16) SSKBQ(1) SSKBQ(8) SSKBQ(16) VQVAE(GS) VQVAE(OH)
PSNR (dB) 29.20 13.22 18.77 21.36 13.02 21.60 23.53 18.44 18.95
SSIM 0.9379 0.3238 0.5886 0.7116 0.3721 0.7240 0.7939 0.5341 0.5734
FID 30.96 400.74 238.82 168.82 344.97 171.49 122.80 221.64 215.05
KID 0.0027 0.3876 0.1763 0.0936 0.3111 0.0974 0.0512 0.1627 0.1450
Experimental results on different image datasets demonstrate that SSKBQ achieves competitive reconstruction performance with substantially reduced storage overhead compared with existing knowledge-base quantization schemes. Moreover, the results reveal that increasing the number of knowledge bases does not always guarantee better performance, since semantic fragmentation and optimization difficulty may limit their effective utilization. Developing more effective training strategies for multi-KB architectures is therefore an important direction for future research. Furthermore, the current study focuses on semantic-layer KB reuse under reliable index delivery. Extending SSKBQ to practical communication environments, including noisy index transmission, fading, packet loss, and channel-aware KB adaptation, will be investigated in future work. R EFERENCES [1] Y. Shao, Q. Cao, and D. Gündüz, “A theory of semantic communication,” IEEE Transactions on Mobile Computing, vol. 23, no. 12, pp. 12 211– 12 228, 2024. [2] X. Luo, H.-H. Chen, and Q. Guo, “Semantic communications: Overview, open issues, and future research directions,” IEEE Wireless communications, vol. 29, no. 1, pp. 210–219, 2022. [3] W. Yang, H. Du, Z. Q. Liew, W. Y. B. Lim, Z. Xiong, D. Niyato, X. Chi, X. Shen, and C. Miao, “Semantic communications for future internet: Fundamentals, applications, and challenges,” IEEE Communications Surveys & Tutorials, vol. 25, no. 1, pp. 213–250, 2022. [4] P. Markstein, “The new ieee-754 standard for floating point arithmetic.” Schloss Dagstuhl–Leibniz-Zentrum für Informatik, 2008. [5] A. Van Den Oord, O. Vinyals et al., “Neural discrete representation learning,” Advances in neural information processing systems, vol. 30, 2017. [6] M. Adiban, K. Stefanov, S. M. Siniscalchi, and G. Salvi, “S-hr-vqvae: Sequential hierarchical residual learning vector quantized variational autoencoder for video prediction,” IEEE Transactions on Multimedia, 2025. [7] B. Jähne, Digital image processing. Springer Science & Business Media, 2005. [8] Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: from error visibility to structural similarity,” IEEE transactions on image processing, vol. 13, no. 4, pp. 600–612, 2004. [9] M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “Gans trained by a two time-scale update rule converge to a local nash equilibrium,” in Advances in Neural Information Processing Systems, vol. 30. Curran Associates, Inc., 2017. [10] M. Bińkowski, D. J. Sutherland, M. Arbel, and A. Gretton, “Demystifying mmd gans,” 2021. [11] E. Bourtsoulatze, D. B. Kurka, and D. Gündüz, “Deep joint sourcechannel coding for wireless image transmission,” IEEE Transactions on Cognitive Communications and Networking, vol. 5, no. 3, pp. 567–579, 2019.