arXiv:2609.04970v1 [cs.CR] 4 Sep 2026
Robust Coverless Linguistic Steganography via Sentence Embedding Space with Global Resynchronization Lizhi Xiong✉
Yuping Lu
Jun Li
Nanjing University of Information Science and Technology Nanjing, China [email protected]
Nanjing University of Information Science and Technology Nanjing, China [email protected]
Nanjing University of Information Science and Technology Nanjing, China [email protected]
Ziqiang Li
Zhangjie Fu
Nanjing University of Information Science and Technology Nanjing, China [email protected]
Nanjing University of Information Science and Technology Nanjing, China [email protected]
Abstract
1
Linguistic steganography enables covert communication through natural language. Existing methods heavily rely on token-level operations and struggle to maintain reliability under word- and sentencelevel textual perturbations. Moreover, variable-length coding-based schemes are highly susceptible to bit-slippage under minor disturbances, as perturbations cause desynchronization between embedded and extracted bit sequences. To address these issues, we propose a robust coverless steganographic framework that operates in the sentence embedding space rather than the token space. Specifically, secret messages are encoded as hierarchical clustering paths in the sentence embedding space, which enhances decoding stability against word- and sentence-level textual perturbations. To tackle the bit-slippage problem, we introduce a Global Resynchronization Mechanism (GRM) that reframes variable-length bitstreams as discrete symbols anchored to semantic subspaces, decoupling local embedding failures from global message recovery. Experimental results demonstrate that under word- and sentence-level perturbations, our approach achieves substantial improvements in robustness, while maintaining effective embedding capacity and exhibiting strong resistance to statistical analysis.
With the advancement and development of social networks, today’s social media platforms increasingly rely on automated systems for large-scale and strict information regulation to monitor, filter, and shape public opinion. Under the circumstances, users are turning to steganography as a covert means of conveying sensitive information, aiming to protect personal privacy. Steganography is a technique that embeds secret information into seemingly innocuous carriers to enable covert communication without arousing suspicion. In the age of digital steganography, people use texts, images [28], videos [2], and audios [29] as media to transmit secret messages. As the most fundamental medium of human communication, text is generated, transmitted, and consumed at massive scale across social media platforms every day, providing a natural and abundant cover for information hiding. Consequently, embedding secret information in text has become an important and active branch of steganography research. Existing steganography methods are classified into modificationbased, retrieval-based and generative steganography. In the first two methods, secret messages are conveyed by altering or selecting words [36] or sentences [27] according to predefined coding rules. Such methods are inherently fragile due to the sensitivity of the characters. Even slight deletions or replacements may break the shared dictionary/codebook correspondence, leading to incorrect decoding or complete extraction failure once the edited content falls outside the shared dictionary. Subsequently, generative linguistic steganography [34, 42] based on Autoregressive Models (ARMs) [24] gained popularity due to fluent and natural generation. In these methods, secret information is embedded by sampling from context-conditioned token probability distributions; however, their strong sequential dependency makes them extremely fragile, as tampering with a single token can propagate errors to subsequent tokens and cause irreversible information loss, especially when attacks introduce error contents or segmentation ambiguity [18]. As shown in Figure 1, for methods that use characters or tokens as steganographic channels, even minor perturbations can cause decoding to fail. Therefore, existing text steganography methods generally suffer from insufficient robustness when facing realistic text attacks, fundamentally limiting their practical applicability.
CCS Concepts • Security and privacy → Human and societal aspects of security and privacy; • Computing methodologies → Lexical semantics.
Keywords Robust Linguistic Steganography; Sentence Embedding Space; Error Correction Code
Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. This paper has been accepted by ACM Multimedia 2026.
Introduction
• We propose a Hierarchical Clustering Mechanism (HCM) that cluster the semantic space into codable subspaces efficiently. By recursive pruning and hierarchical mapping, HCM achieves an optimal balance between embedding capacity and computational efficiency, reducing the complexity from exponential to linear. • To resolve the critical issue of bit-slippage in variable-length provably secure encoding, we introduce a Global Resynchronization Mechanism (GRM). By reframing local desynchronization as symbol erasures and integrating LT codes for global recovery, GRM decouples local embedding failures from global message integrity. • Extensive experiments demonstrate that our method achieves strong robustness and effective embedding capacity under word- and sentence-level textual attacks.
Figure 1: An example of decoding error due to attacked token not in the vocabulary.
To address the problem, later robust steganographic methods [5, 16, 19] introduced some defense mechanisms at the token-level to enable partial recovery. Nevertheless, these defense mechanisms focus only on token-level perturbations while ignoring word- and sentence-level attacks. In practice, real-world adversaries are unaware of the steganographic system’s segmentation rules and tend to operate at the word- or sentence-level through deletion, insertion anywhere or paraphrasing. Consequently, while token-level defenses can improve robustness against token-boundary-aligned perturbations, robustness under practical word- and sentence-level rewriting remains a fundamental and unresolved challenge for current text steganography frameworks. To address these challenges, we propose a robust coverless linguistic steganography framework in the sentence embedding space. Specifically, a shared sentence encoder maps sentences into a unified semantic space, which is partitioned into multiple codable semantic subspaces for information hiding. We cluster the semantic space in a hierarchical manner to reduce the computational overhead associated with the clustering required for high-capacity. Secret information is conveyed through subspace selection and sentence sampling within the selected subspace. To ensure the security of subspace selection, a provably secure coding-and-sampling strategy is adopted to selectively encode only a subset of semantic subspaces, so that the resulting sampling distribution remains consistent with the underlying natural distribution. However, this strategy is inherently variable-length, and thus a subspace prediction error at the receiver may lead to a local mismatch between the extracted and embedded message lengths, causing bit-slippage across subsequent bits. To address this issue, we organize the hidden information as discrete symbols aligned with the partitioning structure of the semantic space and integrate LT codes for global recovery, thereby localizing decoding errors and improving robustness against text perturbations. Our contributions can be summarized as follows:
2 Related Work 2.1 Retrieval Linguistic Steganography Retrieval steganography requires the sender and receiver to share a large dataset and a set of coding rules, whose advantage is that the transmitted carrier remains natural without any modification. The sender selects a group of elements that can be mapped to the secret messages and sends them to the receiver to realize the transmission of secret information. For example, [37] established the mapping relationship through the word-frequency characteristics in the corpus; [9] encoded text according to the distance in semantic space; [25] built a search tree according to the character structure; [6] jointly realized information hiding through a database combining graphics and text. Retrieval linguistic steganography is a type of coverless steganography. Current retrieval-based text steganography methods have not systematically investigated steganographic robustness, which needs to be further explored. These methods rely on ARMs and hardly change the probability distribution of the generated words or semantics to ensure security when embedding secret information.
2.2
Provably Secure Linguistic Steganography
In recent years, researchers have proposed a variety of provably secure linguistic steganography methods based on generative language models. These methods are dedicated to designing messageembedding algorithms that are indistinguishable from the conventional text-generation process. ADG [40] adaptively and dynamically grouped tokens based on the probability generated by the pre-trained language model and recursively embeds secret information. Meteor [7] used the random prefix shared by the generation model to embed the message into the sampling random number and masked it at once, dynamically adapted to entropy changes, and takes both security and efficiency into account. Discop [4] proposed the “distributed copies” technique, which uses the index value of these “copies” to express secret information during the generation process and realizes practical steganography with provable security and high capacity. [15] made two major improvements to the prefix-encoding method—quantitative fine-tuning and distributed coupling—to repair security and capacity defects.
• We propose a novel coverless linguistic steganography framework that shifts the embedding operation from fragile token sequences to a robust, high-dimensional sentence embedding space. By leveraging the approximate semantic invariance of sentence embeddings, our method inherently absorbs wordand sentence-level textual perturbations. 2
2.3
Robust Linguistic Steganography
without altering meaning. Consider a Probabilistic PolynomialTime (PPT) adversary operating in a highly regulated environment whose duty is to detect and eliminate covert communication that may be hidden in public texts. The adversary intercepts any text transmitted over public channels and possesses two capabilities: Detection capability. The adversary can apply steganalysis techniques to determine whether the suspicious texts may contain secret information. This corresponds to the security requirement. Tampering capability. The adversary can tamper a text 𝑠 into 𝑠 ′ , but within limits: the semantic content of 𝑠 remain unchanged. Let textual attack be a transformation operator: A : 𝑠 → 𝑠 ′ . This transformation satisfies approximate semantic invariance:
In recent years, researchers began to pay attention to the problem of message loss caused by segmentation ambiguity, and designed corresponding countermeasures. SECC [5] overlaid an ECC layer similar to LT codes [10] on the conventional generative steganography pipeline to improve character-level erasure resistance; however, it didn’t mitigate the original extraction-error problem. To avoid altering the probability distribution of candidate words, Syncpool [18] merged tokens that share a prefix into an “ambiguity pool” and used a shared encrypted pseudo-random generator to ensure consistent selection, yet the secret information remained sensitive to vocabulary shifts. Active textual attacks—such as deletion, insertion, and tampering—have also been studied. GTSD [30] employed a diffusion model to generate candidate texts and embeds information via prompt-andbatch mapping, achieving local robustness. Winstega [16] sacrificed capacity for anti-editing capability by using a sliding window and entropy-threshold-based discontinuous embedding during inference. STEAD [19] embedded duplicate codes in parallel at multiple locations within a single denoising step of a diffusion model and coordinated with neighborhood-search alignment to realize provably secure steganography resistant to insertion and deletion. These approaches achieved stable extraction under their respective tokenlevel active attacks.
2.4
∥v(𝑠 ′ ) − v(𝑠)∥ 2 ≤ 𝜖,
(2)
where v(𝑠) denotes the semantic vector of 𝑠 and 𝜖 is the tampering threshold. This corresponds to the robustness requirement of steganography.
3.2
Overview
In this section, we introduce the proposed steganography framework. First of all, we introduce how to construct and partition semantic space and how to improve the efficiency. Then, we will describe the security application of SparSamp sampling method in this framework, and point out its limitations in robustness. On this basis, we employ a Global Resynchronization Mechanism (GRM) to improve it. Figure 2 shows the proposed stegosystem. In the embedding phase, we cluster the sentence embedding spaces mapped by the dataset, and code and sample the divided semantic subspaces at multiple levels. At the same time, we encode the secret messages to be sent into symbol blocks, and embed them into each subspace layer in units of symbol blocks. After the iteration, we randomly select sentence samples in the selected subspaces and send them as steganography carriers. Text attacks may occur during transmission. In the extraction phase, we use the same sentence encoder as in the embedding phase to reconstruct the semantic space of the shared dataset, predict the semantic subspace of each sentence embedding in the steganography texts, reproduce the same hierarchical clustering results as the embedding phase, and decode the symbol blocks according to the sampling process. Finally, we use Belief Propagation (BP) decoder to accurately recover the original secret messages according to the collected symbol blocks.
Sentence Embedding Encoder
Sentence embeddings are obtained by encoding text sentences with a sentence embedding encoder. After sufficient training, embeddings of semantically similar sentences exhibit higher similarity and smaller distance. Semantic similarity is computed by: Í𝑛 𝐴·𝐵 𝑖=1 𝐴𝑖 𝐵𝑖 Sim(𝐴, 𝐵) = = √︃ , (1) √︃ Í𝑛 ∥𝐴∥ ∥𝐵∥ 2 Í𝑛 𝐵 2 𝐴 𝑖=1 𝑖 𝑖=1 𝑖 where 𝐴 and 𝐵 denote the vector representations of the sentences, and 𝐴𝑖 and 𝐵𝑖 are their 𝑖-th components. Owing to different training parameters and embedding dimensions, each encoder exhibits distinct characteristics. For instance, the static encoder [12], trained with modest parameters, yields a relatively flat and uniform semantic space suitable for small models, but its advantage is speed. In contrast, encoders based on large models [1] produce a more comprehensive and high-dimensional semantic space at the cost of increased time and memory, yet offer stronger text-processing capabilities. Due to the strong similaritymatching ability, sentence embeddings are widely employed in downstream tasks such as retrieval [8], clustering [23], etc.
3.3
Construction of Semantic Space
Robust Subspace Construction. The core of our robust framework lies in the structural stability inherent in the high-dimensional sentence manifold. High-quality sentence embeddings, generated by pre-trained encoders, exhibit a crucial property of semantic invariance: sentences with minor textual perturbations tend to cluster within a constrained neighborhood in the semantic space 𝑉𝑆 . By applying clustering algorithms to partition 𝑉𝑆 into a set of disjoint subspaces—rather than arbitrary partitioning—we capitalize on the natural distribution of the data to minimize intra-cluster variance. Due to the sparsity of high-dimensional spaces, most samples are distributed near the centroid. Consequently, we can establish a stable mapping between secret information and regional semantic
3 Proposed Method 3.1 Deployment Assumption and Threat Model The proposed framework focuses on a symmetric key steganography system: the sender and receiver need to share the same dataset, sentence embedding encoder, LT Encoding Symbol Identifiers (ESI) and pseudo-random number key to ensure the construction of the same sentence-level semantic space and the success of extraction. The adversary can apply semantic-preserving perturbations but has no access to shared secrets, aiming to disrupt communication 3
Figure 2: An overview of the proposed stegosystem. Embedding and extraction are mutually inverse processes. Hierarchical Clustering Mechanism. If a high steganography capacity is required, a large dataset must be used to construct a richer semantic space, thereby enabling the partitioning of more semantic subspaces. As shown in Eq. 3, the time complexity of the k-means algorithm for 𝑁 samples is 𝑂 (𝑁 · 𝑘). Assuming that a single sentence is to embed 𝑏 secret bits, Shannon’s source-coding principle [22] dictates that 𝑏 secret bits requires at least 2𝑏 distinguishable states. Consequently, the number of sub-clusters k must scale as:
characteristics. In this schema, we select the semantic subspaces as the fundamental steganographic units rather than individual tokens or specific sentence instances. Since the secret bits are tied to the subspace identity rather than the precise coordinates of a single vector, the system can tolerate significant word-level or sentence-level attacks. As long as the perturbed stego sentence 𝑠 ′ satisfies the semantic proximity constraint Eq. 2 and remains within the original cluster boundary, the receiver can accurately identify the corresponding subspace and retrieve the secret message. This structure-level design effectively decouples the hidden information from fragile textual surface forms, thereby providing intrinsic robustness against various practical attacks. We use k-means algorithm to cluster a semantic space 𝑉𝑆 . Given the number of clusters 𝑘 and pseudo-random seed 𝑟𝑐 , k-means clustering can be formalized as the following optimization problem: {𝝁 1, . . . , 𝝁 𝑘 } = arg min
𝑁 ∑︁
min v𝑖 − 𝝁 𝑗
{𝝁 𝑗 }𝑘𝑗=1 𝑖=1 1≤ 𝑗 ≤𝑘
s.t.
𝑘 ∝ 2𝑏 .
Then the time complexity of the algorithm is 𝑂 (𝑁 · 2𝑏 ). Therefore, when clustering large semantic spaces, increases in the number of samples 𝑁 and in the number of secret bits 𝑏 incur significant computational overhead. To address this problem, we design a Hierarchical Clustering Mechanism (HCM) that embeds 𝑏 secret bits across 𝑙 layers, with each layer carrying 𝑏¯ bits. In each layer 𝑖, the encoding process (𝑖 ) selects a specific cluster 𝑉ℎ𝑖𝑡 as the “hit” subspace based on the secret bits and designates it as the semantic space 𝑉 (𝑖+1) for the next layer. This cycle repeats 𝑙 times. ¯ The time complexity of this algorithm is 𝑙 ·𝑂 (𝑁 ·2𝑏 ). Based on the actual clustering efficiency, we set 𝑏¯ to 2 or 3. Then the complexity of each layer of clustering can be controlled and stable, and the overall time complexity can be regarded as linear complexity 𝑂 (𝑁 · 𝑙). By concentrating computational overhead solely on the steganographic path, this method significantly improves the overall efficiency of the framework and effectively performs recursive pruning of irrelevant branches in the semantic space.
2
, 2
(3)
𝝁 (0) 𝑗 ∼ P (𝑟𝑐 ),
where 𝜇 𝑗 is the 𝑗-th cluster center, 𝜇 𝑗(0) is the initial centroid initialized by the pseudo-random number 𝑟𝑐 , and P (𝑟𝑐 ) is the initialization distribution determined by the random seed 𝑟𝑐 . After clustering, the semantic space 𝑉𝑆 is divided into 𝑘 disjoint semantic subspaces. To ensure that each semantic subspace has sufficient stability, we test different values of 𝑘 to determine the appropriate threshold. When the optimization convergence time is too long or the subspaces are too small to be robust after clustering, we consider 𝑘 to be too large. When the value of 𝑘 is set too small, although there is no concern about the stability of the subspace, the embedding capacity becomes too limited. Since both sides of the communication share the same 𝑟𝑐 , the clustering function is determined: C𝑟𝑐 : 𝑉𝑆 → {𝑉1, 𝑉2, . . . , 𝑉𝑘 },
(5)
3.4
Global Resynchronization
Subspace Sampling. During clustering, the discrete distribution of samples makes it impossible to partition the semantic space perfectly evenly, so the number of samples in each subspace is usually different. If the subspaces {𝑉1, 𝑉2, ..., 𝑉𝑘 } are encoded in simple sequential order, the secret-driven selection will bias the semantic
(4)
thereby ensuring that the sender and receiver obtain consistent semantic subspace partition results. 4
Figure 4: An example of GRM. The green arrow indicates that the clustering is correct, and the red arrow indicates that the clustering is wrong. Figure 3: An example of bit-slippage. Due to a local decoding error, the subsequent extracted bit sequence is out of sync with the embedded bits.
because the internal spatial structures of different clusters vary, the frequency of message conflicts also differs. This may cause the length of the bits extracted from that sentence to differ from the original bits, leading to misalignment in the subsequent bit sequence and significantly increasing the error rate. Figure 3 illustrates an example of bit-slippage. To address this issue, we propose a Global Resynchronization Mechanism (GRM) based on LT codes to dynamically align the extracted bit positions and decouple local extraction failures from global message recovery. During embedding, we let the length of each symbol block of the LT codes be fixed to the single-layer embedding capacity 𝑏 and standardize the number of layers for each sentence clustering to 𝑙. The LT encoder can generate an unbounded checksum stream 𝐶 = {𝑐 1, 𝑐 2, . . .}, where each symbol 𝑐𝑖 ∈ {0, 1}𝑏 represents a complete LT-encoded symbol block. Each layer carries exactly at most one symbol block. If the subspace sampling at layer 𝑗 of sentence 𝑖 does not trigger a conflict, the corresponding symbol block 𝐵𝑖𝑏+𝑗 is embedded into that layer. Otherwise, the layer is skipped and the symbol block 𝐵 (𝑖 −1)𝑙+𝑗 is not embedded. For the 𝑖-th sentence, the symbol blocks are preassigned to the interval [(𝑖−1)𝑙+1, 𝑖𝑙] globally, where missing symbol blocks are treated as natural erasures. During extracting, the same decoding procedure is applied to the received sentence 𝑠 ′ : If the subspace sampling at layer 𝑗 of sentence ′ 𝑖 does not trigger a conflict, the corresponding symbol block 𝐵𝑖𝑏+𝑗 ′ is extracted; Otherwise, the layer is skipped and 𝐵𝑖𝑏+𝑗 is none. Due to clustering errors or message conflicts, symbol blocks may not be successfully recovered from sentence 𝑠 ′ that matches the one from the embedding stage. All extracted symbol blocks are subsequently reorganized into a global verification stream 𝐶 ′ according to their predefined ESIs. Once the total number of collected symbols exceeds the LT decoding threshold, the original secret message can be reconstructed in a single decoding step. In this manner, local layer-level deviations are isolated at the sentence level, preventing error propagation across bits and significantly enhancing the robustness of the steganographic system. Figure 4 illustrates an example of GRM when 𝑏 = 3 and 𝑙 = 3. After being attacked, the sentence 𝑠 2′ is unfortunately divided into
distribution of the output sentences away from the natural distribution, introducing security risks. To eliminate this bias, we adopt the SparSamp [26] approach—message-driven pseudo-random number—for subspace sampling. The construction of the semantic space and the procedure of HCM remain unchanged. Let the subspaces obtained after the 𝑡-th layer partition driven by the pseudo-random number 𝑟𝑐 be Í C𝑟𝑐 (𝑉 (𝑡 ) ) = {𝑉1(𝑡 ) , 𝑉2(𝑡 ) , ..., 𝑉𝑘(𝑡 ) }, where 𝑉 (𝑡 ) = 𝑘𝑖=1 𝑉𝑖 (𝑡 ) . The cumulative probability distribution of the samples is then: 𝐹 (𝑡 ) ( 𝑗) =
𝑗 ∑︁ |𝑉 (𝑡 ) | 𝑖
|𝑉 (𝑡 ) | 𝑖=1
,
(6)
where 𝑗 ∈ {0, 1, . . . , 𝑘 }. If 𝑏 secret bits 𝑚(𝑡) ∈ {0, 1}𝑏 are to be embedded in this layer, we partition the probability interval [0, 1) uniformly with the minimum granularity 𝛿 = 1/2𝑏 , yielding 2𝑏 equally wide intervals whose starting positions are 𝑥𝑖 = 𝑖 · 𝛿, where 𝑖 ∈ {0, 1, ..., 2𝑏 − 1}. To simulate real random sampling and ensure reproducibility, we introduce a pseudo-random number 𝑟 Δ ∈ [0, 1) derived from the shared key and construct the pseudo-random position: 𝑟 (𝑖) = (𝑥𝑖 + 𝑟 Δ ) 𝑚𝑜𝑑 1, (7) where 𝑖 ∈ {0, 1, . . . , 2𝑏 − 1}. Let 𝑖 ∗ = 𝑑𝑒𝑐 (𝑚(𝑡)) be the decimal value of the secret bits. The subspace selected for the next layer is given by the inverse CDF: 𝑉 (𝑡 +1) = 𝑉𝑗(𝑡∗ ) , (8) 𝑗 ∗ = argmin 𝐹 (𝑡 ) ( 𝑗) ≥ 𝑟 (𝑖 ∗ ) . (9) 𝑗
If there are two or more pseudo-random positions corresponding to the target subspace 𝑉 (𝑡 +1) , no message is embedded in this layer. This phenomenon is known as a message conflict. Global Resynchronization Mechanism. After being subjected to a perturbation, there is a small probability that the receiver will incorrectly predict the cluster to which a sentence belongs, thereby extracting an incorrect secret message. More seriously, however, 5
Table 1: Robustness against paraphrase attack.
the wrong subspace, and the involved symbol blocks is wrongly extracted. Because LT codes with a certain degree of redundancy possess inherent robustness, the correct message can still be recovered even if some symbol blocks are missing.
4 Experiment 4.1 Implementation Details
Method
AC
ADG
SECC
Discop
STEAD
FSBTS
Ours
BER BLR 𝑃𝐸
0.6145 0 0.6145
0.3933 0 0.3933
0.5500 0 0.5500
1 0.5000
1 0.5000
0.1090 0.2681 0.2138
0.0038 0 0.0038
Table 2: Comparison of PPL.
Setup. We select the general text-embedding model Sentence-T5Large [13] optimized by T5 model as the sentence-vector encoder. We retrieve and generate texts on PersonaChat [39] corpus. For HCM, we set the number of layers 𝑙 to 3 and the number of clusters per layer 𝑘 to 8. For decoding, we use the BP decoding algorithm. Baselines. We select AC [43], ADG [40], Discop [4], SECC [5], STEAD [19] and FSBTS [32] as the baseline methods. STEAD adopts Dream-7B [35] and the rest adopt GPT-2 [20]. FSBTS uses openConcepts [38] as its dataset, while other methods randomly select texts from the PersonaChat corpus as inputs for text generation tasks.
4.2
AC ADG ADG+SECC Discop STEAD Ours
Metrics
PersonaChat
C4
IMDB
112.70 127.24 125.16 109.34 128.55 53.85
157.78 398.31 390.18 193.18 94.07 43.21
213.95 394.51 431.72 176.15 291.76 41.52
Table 3: Embedding capacity when 𝛼 = 5. The “Ours (P/C/I)” column represents the result of the stegos of our method on PersonaChat/C4/IMDB.
Robustness. In order to demonstrate the robustness, we test the extraction error rate 𝑃𝐸 of stegos after five types of textual attacks. In the actual extraction process, the damage of secret information is mainly manifested in two physical forms: fragment loss and extraction error. According to information theory, the mutual information provided by the completely lost bits is 0, which is equivalent to the random guess of equal probability (i.e. the error rate is 0.5) at the receiver in the binary channel. Therefore, we use the full probability formula to convert the bit loss rate (BLR) of the extracted message and the bit error rate (BER) of the extracted message into the equivalent comprehensive bit error rate 𝑃𝐸 : 𝑃𝐸 = (1 − 𝐵𝐿𝑅) ∗ 𝐵𝐸𝑅 + 0.5 ∗ 𝐵𝐿𝑅.
Dataset
Method
Method
AC
ADG
ADG+SECC
Discop
STEAD
FSBTS
Ours (P/C/I)
𝑃𝐸 𝐻2 ER (bit/token) EER (bit/token)
0.4001 0.971 2.47 0.07
0.3852 0.958 1.36 0.06
0.4938 1 0.45 0.00
0.2856 0.863 0.43 0.06
0.4204 0.977 0.08 0.00
0.0438 0.0259 0.48 0.36
0.0038 0.036 0.14/0.08/0.07 0.13/0.08/0.07
as EER = (1 − 𝐻 2 (𝑃𝐸 )) ∗ ER, where ER denotes the raw embedding capacity. As 𝑃𝐸 approaches 50% and [1 − 𝐻 2 (𝑃𝐸 )] approaches 0, more redundant codes need to be introduced, resulting in less EER. We conducted tests with 𝛼 = 5, as a representative value for moderate intensity. Under this condition, all methods functioned properly, and there were significant differences in terms of robustness. Through numerous experiments, we have found that the relevant conclusions hold true for other values of 𝛼 as well. Security. We use five steganalysis strategies—FCN [33], R-BiLSTMC [14], LSTMATT [44], SANet [31] and HiDuNet [17]—to test security of stegos. The closer the accuracy of steganalysis to 50%, the better the concealment and the securer the stegos. Since there is no comparable cover text for FSBTS, it is not included in the evaluation.
(10)
The types of text attacks we use include: deletion, replacement, insertion, swap and paraphrase. We use Parrot_Paraphraser [3] model to implement the paraphrase operation. Among them, paraphrase belongs to sentence-level attacks, and other attack types belong to word-level attacks. For word-level attacks, we set the ratio of the attacked words to the total number of words in a sample as 𝛼, which can be regarded as the attack intensity. Quality of Stegos. We use Perplexity (PPL) of the model DialoGPTmedium [41] as the index to evaluate the quality of steganography. PPL measures the instability of an LLM in predicting text sequences and represents the fluency of texts. The smaller the PPL, the smoother the stegos and the better the concealment. Since the dataset of FSBTS is fixed, it is not included in the evaluation. Embedding Capacity. The ultimate transmission rate for reliable, error-free communication over a memoryless binary channel is bounded by the channel capacity [22]. The channel uncertainty induced by the composite error rate 𝑃𝐸 is formulated by the binary entropy function 𝐻 2 (𝑃𝐸 ) = −𝑃𝐸 log2 (𝑃𝐸 ) − (1 − 𝑃𝐸 ) log2 (1 − 𝑃𝐸 ). At the theoretical limit, 𝐻 2 (𝑃𝐸 ) represents the minimum proportion of redundant parity bits inherently required to correct channel noise. Conversely, the term [1 − 𝐻 2 (𝑃𝐸 )] represents the capacity factor, which is the fraction genuinely available for carrying the useful payload after deducting the error-correction overhead. Consequently, the actual Effective Embedding Rate (EER) is computed
4.3
Results
Robustness. Figure 6 shows the robust performance of different methods under different word-level attack types with different attack intensities. Our method sets the redundancy of LT code to three times. It can be seen that the baseline methods are quite vulnerable to word-level attacks, and their 𝑃𝐸 increase as the attacks become stronger, tending to be 50% of random guess. In contrast, the 𝑃𝐸 of our proposed method remains below 2.7%, achieved A significant improvement compared to the baseline methods. Figure 5 illustrates the performance of various methods in terms of BER, BLR, and 𝑃𝐸 under different types of attacks when 𝛼 = 5. Based on the results, we can thoroughly analyze the primary factors contributing to the extraction errors for each method. First, 6
Figure 5: The specific robustness performance of different methods under different attack types when 𝛼=5. Table 4: Anti-steganalysis comparison. The ADG results are the same whether or not SECC error correction is added. Method AC ADG(+SECC) Discop STEAD Ours
FCN 65.10% 58.30% 52.70% 55.56% 50.54%
R-BiLSTM-C 59.56% 56.93% 56.90% 54.44% 49.60%
LSTMATT 59.75% 53.82% 53.92% 52.22% 50.63%
SANet 77.58% 90.67% 90.67% 52.54% 53.43%
Table 5: The effect of the number of clusters 𝑘 on robustness and capacity (𝛼 = 5).
HiDuNet 74.00% 96.33% 92.71% 93.33% 56.49%
𝑘 2 4 8 16
𝑃𝐸𝐷 0.0038 0.0132 0.0188 0.2876
𝑃𝐸𝑅 0.0132 0.0225 0.0263 0.1147
𝑃𝐸𝐼 0 0 0.0038 0.3214
𝑃𝐸𝑆 0 0.0019 0.0038 0.3195
𝑃𝐸𝑃 0 0 0.0038 0.3214
ER 0.0299 0.0667 0.1295 0.1338
Table 6: Ablation study of HCM. We test the efficiency under different numbers of clustering layers.
the 𝑃𝐸 of AC and ADG is dominated by BER, as assigning default values to undecodable tokens converts potential loss into extraction errors. Although ADG extracts sufficient blocks for SECC to achieve a 0 BLR, its high inherent symbol error rate causes SECC’s error correction to fail, yielding a BER near 50%. Conversely, our method shares SECC’s BLR mechanism but ensures highly accurate symbol extraction via resynchronization and semantic space stability. This successfully averts SECC’s decoding failure, maintaining exceptionally low BER and 𝑃𝐸 under various attacks. Second, the 𝑃𝐸 of Discop and STEAD is predominantly driven by BLR, because they halt decoding at undecodable tokens, converting potential errors into message loss. Notably, Discop’s apparent error rate is artificially lowered by its zero-padding strategy, as these padded sequences are excluded from actual BLR calculations. Table 1 shows the robustness performance of different methods under paraphrase attack. It can be clearly seen that the baseline methods are also vulnerable to sentence-level attack, with Discop and STEAD completely unable to decode the attacked stegos, while our method still has strong robustness and significant advantages. Among all attacks, the ER is slightly higher for replacement and deletion attacks on our system, as these operations may alter semantic keywords, causing the prediction to fail to be assigned to the correct cluster. Quality of Stegos. Table 2 shows the PPL metrics of different methods on three datasets: PersonaChat [39], C4 [21], and IMDB [11]. By selecting natural sentences directly from the corpus, our method preserves the original linguistic distribution and fluency, achieving significantly lower PPL compared to generative methods. Embedding Capacity. As shown in Table 3, we test the embedding capacity of different methods under mixed attacks (including all word-level attacks) when 𝛼=5. Our method sets the redundancy of
Number of Layers 1 2 3 4
Encoding Rate(s/bit) 20.96 1.37 1.52 2.18
Encoding Rate(s/layer) 167.71 5.48 4.61 1.09
Decoding Rate(s/bit) 10.75 1.17 1.49 2.11
Decoding Rate(s/layer) 85.96 4.67 4.47 1.06
Table 7: Capacity under different redundancy levels for deletion attack (𝛼 = 5). 𝑅𝐿𝑇 ER(bit/token)
1.5× 0.2552
2× 0.1925
2.5× 0.1598
3× 0.1338
LT code to three times. Although our method offers no advantage in ER under lossless transmission, its EER in the attack scenario outperforms that of other methods except FSBTS, which employs a larger dataset than ours. In addition to the comparison with the baseline methods, we test the embedding capacity of the proposed method on the same three datasets as mentioned above. It is clear that the ER of the stego texts in PersonaChat is higher, as the average length of its sentences is significantly shorter than that of the other two datasets. Security. Table 4 shows that the detection accuracy of our proposed method is close to 50%, which shows that the performance of steganalysis methods is not better than random guess when detecting the stegos generated by our method, thus verifying the security of our method. 7
Figure 6: Robustness against word-level attacks. 𝛼 represents the ratio of the attacked words to the total number of words in a sample.
Figure 7: Ablation study of GRM. We test the BER of systems without GRM and systems with GRM combined with LT codes at different redundancy levels under various types of attacks of varying intensity.
4.4
Ablation Study
robustness of semantic-space steganography and confirms that bitslippage, rather than semantic distortion, is the predominant factor behind decoding failure in the baseline system. In addition, the BER decreases consistently as the redundancy of LT codes increases. While 2× redundancy keeps the BER below 10%, it still exhibits an error floor. In contrast, the 3× configuration achieves near-perfect reconstruction across all attack types and intensities. The system maintains consistent robustness across four types of attacks. Furthermore, as shown in Table 7, as the redundancy of the LT code increases, the amount of actual secret messages embedded decreases, and the ER shows a downward trend. This illustrates the trade-off between robustness and embedding capacity.
Number of Clusters. We test how the system’s robustness and embedding capacity vary with different cluster sizes 𝑘. Columns 2 through 6 represent the 𝑃𝐸 values under deletion, replacement, insertion, swap, and paraphrase attacks, respectively. Table 5 shows that as 𝑘 increases, 𝑃𝐸 exhibits an upward trend, indicating that this reduces the size of the subspace, thereby making extraction errors more likely to occur. In addition, an increase in 𝑘 also increases the codable entropy and improves the embedding rate (ER). Hierarchical Clustering Mechanism (HCM). We test the time consumption with different number of levels. Table 6 shows that, compared with the single-layer flat structure, the introduction of HCM greatly reduces the single clustering time and achieves an order of huge improvement in efficiency. When the number is 2, the system reaches the efficiency inflection point; When the number continue to increase, time spent on each level falls, but the crosslayer scheduling overhead caused by multi-layer begins to exceed the benefit, resulting in a rebound in the overall time per bit. Global Resynchronization Mechanism (GRM). As illustrated in Figure 7, the performance drops significantly without GRM module. The system with GRM achieves an extremely low BER, but the BER increases slightly when LT codes are added. This is because GRM converts error bits into lost bits, whereas LT’s redundancy coding converts them back into fewer error bits. This verifies the inherent
5
Conclusion
This paper proposes a robust coverless linguistic steganography framework that operates in the sentence embedding space rather than the fragile token space. By leveraging the semantic invariance of sentence representations, our method inherently absorbs wordand sentence-level textual perturbations. We introduce HCM to efficiently partition the semantic space and GRM to resolve the critical bit-slippage problem through LT codes. Experimental results demonstrate that our approach achieves substantial improvements in robustness, while maintaining effective embedding capacity, superior linguistic quality, and strong security against steganalysis. 8
References
[23] Jianlin Su, Jiarun Cao, Weijie Liu, and Yangyiwen Ou. 2021. Whitening sentence representations for better semantics and faster retrieval. arXiv preprint arXiv:2103.15316 (2021). [24] Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971 (2023). [25] Kaixi Wang and Quansheng Gao. 2019. A coverless plain text steganography based on character features. IEEE Access 7 (2019), 95665–95676. [26] Yaofei Wang, Gang Pei, Kejiang Chen, Jinyang Ding, Chao Pan, Weilong Pang, Donghui Hu, and Weiming Zhang. 2025. { SparSamp } : Efficient Provably Secure Steganography Based on Sparse Sampling. In 34th USENIX Security Symposium (USENIX Security 25). 6817–6835. [27] Alex Wilson and Andrew D Ker. 2016. Avoiding detection on twitter: embedding strategies for linguistic steganography. Electronic Imaging 28 (2016), 1–9. [28] Han-Zhou Wu, Yun-Qing Shi, Hong-Xia Wang, and Lin-Na Zhou. 2016. Separable reversible data hiding for encrypted palette images with color partitioning and flipping verification. IEEE transactions on circuits and systems for video technology 27, 8 (2016), 1620–1631. [29] Junqi Wu, Bolin Chen, Weiqi Luo, and Yanmei Fang. 2020. Audio steganography based on iterative adversarial attacks against convolutional neural networks. IEEE transactions on information forensics and security 15 (2020), 2282–2294. [30] Zhengxian Wu, Juan Wen, Yiming Xue, Ziwei Zhang, and Yinghan Zhou. 2024. GTSD: Generative Text Steganography Based on Diffusion Model. In International Conference on Neural Information Processing. Springer, 168–183. [31] Yiming Xue, Jiaxuan Wu, Ronghua Ji, Ping Zhong, Juan Wen, and Wanli Peng. 2023. Adaptive domain-invariant feature extraction for cross-domain linguistic steganalysis. IEEE Transactions on Information Forensics and Security 19 (2023), 920–933. [32] Jinshuai Yang, Minghao Bai, Kaiyi Pang, Yue Gao, Jiajun Zou, and Yongfeng Huang. 2025. A Novel Framework of Semantic-Based Text Steganography. IEEE Transactions on Dependable and Secure Computing (2025). [33] Zhongliang Yang, Yongfeng Huang, and Yu-Jin Zhang. 2019. A fast and efficient text steganalysis method. IEEE Signal Processing Letters 26, 4 (2019), 627–631. [34] Zhong-Liang Yang, Xiao-Qing Guo, Zi-Ming Chen, Yong-Feng Huang, and Yu-Jin Zhang. 2018. RNN-stega: Linguistic steganography based on recurrent neural networks. IEEE Transactions on Information Forensics and Security 14, 5 (2018), 1280–1295. [35] Jiacheng Ye, Zhihui Xie, Lin Zheng, Jiahui Gao, Zirui Wu, Xin Jiang, Zhenguo Li, and Lingpeng Kong. 2025. Dream 7b: Diffusion large language models. arXiv preprint arXiv:2508.15487 (2025). [36] Biao Yi, Hanzhou Wu, Guorui Feng, and Xinpeng Zhang. 2022. ALiSa: Acrostic linguistic steganography based on BERT and Gibbs sampling. IEEE Signal Processing Letters 29 (2022), 687–691. [37] Jianjun Zhang, Yicheng Xie, Lucai Wang, and Haijun Lin. 2017. Coverless text information hiding method using the frequent words distance. In International Conference on Cloud Computing and Security. Springer, 121–132. [38] Ningyu Zhang, Qianghuai Jia, Shumin Deng, Xiang Chen, Hongbin Ye, Hui Chen, Huaixiao Tou, Gang Huang, Zhao Wang, Nengwei Hua, et al. 2021. Alicg: Fine-grained and evolvable conceptual graph construction for semantic search at alibaba. In Proceedings of the 27th ACM SIGKDD conference on knowledge discovery & data mining. 3895–3905. [39] Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston. 2018. Personalizing dialogue agents: I have a dog, do you have pets too? arXiv preprint arXiv:1801.07243 (2018). [40] Siyu Zhang, Zhongliang Yang, Jinshuai Yang, and Yongfeng Huang. 2021. Provably secure generative linguistic steganography. arXiv preprint arXiv:2106.02011 (2021). [41] Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and William B Dolan. 2020. Dialogpt: Large-scale generative pre-training for conversational response generation. In Proceedings of the 58th annual meeting of the association for computational linguistics: system demonstrations. 270–278. [42] Xuejing Zhou, Wanli Peng, Boya Yang, Juan Wen, Yiming Xue, and Ping Zhong. 2021. Linguistic steganography based on adaptive probability distribution. IEEE Transactions on Dependable and Secure Computing 19, 5 (2021), 2982–2997. [43] Zachary Ziegler, Yuntian Deng, and Alexander M Rush. 2019. Neural linguistic steganography. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 1210–1215. [44] Jiajun Zou, Zhongliang Yang, Siyu Zhang, Sadaqat ur Rehman, and Yongfeng Huang. 2020. High-performance linguistic steganalysis, capacity estimation and steganographic positioning. In International Workshop on Digital Watermarking. Springer, 80–93.
[1] Jianlv Chen, Shitao Xiao, Peitian Zhang, Kun Luo, Defu Lian, and Zheng Liu. 2024. Bge m3-embedding: Multi-lingual, multi-functionality, multi-granularity text embeddings through self-knowledge distillation. arXiv preprint arXiv:2402.03216 4, 5 (2024). [2] Yanli Chen, Hongxia Wang, Hanzhou Wu, Zhiqiang Wu, Tao Li, and Asad Malik. 2019. Adaptive video data hiding through cost assignment and STCs. IEEE Transactions on Dependable and Secure Computing 18, 3 (2019), 1320–1335. [3] Prithiviraj Damodaran. 2021. Parrot: Paraphrase generation for NLU. [4] Jinyang Ding, Kejiang Chen, Yaofei Wang, Na Zhao, Weiming Zhang, and Nenghai Yu. 2023. Discop: Provably secure steganography in practice based on" distribution copies". In 2023 IEEE Symposium on Security and Privacy (SP). IEEE, 2238–2255. [5] Yuzhe Guo, Zhongliang Yang, Zhuang Wang, Zhili Zhou, and Linna Zhou. 2025. SECC-Stega: Generative Linguistic Steganographic Framework Based on Error Correcting Codes. In ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1–5. [6] Yuting Hu, Haoyun Li, Jianni Song, and Yongfeng Huang. 2020. MM-stega: multimodal steganography based on text-image matching. In International Conference on Artificial Intelligence and Security. Springer, 313–325. [7] Gabriel Kaptchuk, Tushar M Jois, Matthew Green, and Aviel D Rubin. 2021. Meteor: Cryptographically secure steganography for realistic distributions. In Proceedings of the 2021 ACM SIGSAC Conference on Computer and Communications Security. 1529–1548. [8] Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick SH Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense Passage Retrieval for Open-Domain Question Answering.. In EMNLP (1). 6769–6781. [9] Yi Long and Yuling Liu. 2018. Text coverless information hiding based on word2vec. In International Conference on Cloud Computing and Security. Springer, 463–472. [10] Michael Luby. 2002. LT codes. In The 43rd Annual IEEE Symposium on Foundations of Computer Science, 2002. Proceedings. IEEE Computer Society, 271–271. [11] Andrew Maas, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts. 2011. Learning word vectors for sentiment analysis. In Proceedings of the 49th annual meeting of the association for computational linguistics: Human language technologies. 142–150. [12] Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781 (2013). [13] Jianmo Ni, Gustavo Hernandez Abrego, Noah Constant, Ji Ma, Keith Hall, Daniel Cer, and Yinfei Yang. 2022. Sentence-t5: Scalable sentence encoders from pretrained text-to-text models. In Findings of the association for computational linguistics: ACL 2022. 1864–1874. [14] Yan Niu, Juan Wen, Ping Zhong, and Yiming Xue. 2019. A hybrid R-BILSTM-C neural network based text steganalysis. IEEE Signal Processing Letters 26, 12 (2019), 1907–1911. [15] Chao Pan, Donghui Hu, Yaofei Wang, Kejiang Chen, Yinyin Peng, Xianjin Rong, Chen Gu, and Meng Li. 2025. Rethinking Prefix-Based Steganography for Enhanced Security and Efficiency. IEEE Transactions on Information Forensics and Security (2025). [16] Kaiyi Pang, Minhao Bai, Jinshuai Yang, Wei-Qiang Zhang, Minghu Jiang, and Yongfeng Huang. 2025. WinStega: An Adaptive Robust Enhancement Framework for Generative Linguistic Steganography. In ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1–5. [17] Wanli Peng, Sheng Li, Zhenxing Qian, and Xinpeng Zhang. 2023. Text steganalysis based on hierarchical supervised learning and dual attention mechanism. IEEE/ACM Transactions on Audio, Speech, and Language Processing 31 (2023), 3513–3526. [18] Yuang Qi, Kejiang Chen, Kai Zeng, Weiming Zhang, and Nenghai Yu. 2024. Provably secure disambiguating neural linguistic steganography. IEEE Transactions on Dependable and Secure Computing (2024). [19] Yuang Qi, Na Zhao, Qiyi Yao, Benlong Wu, Weiming Zhang, Nenghai Yu, and Kejiang Chen. 2025. STEAD: Robust Provably Secure Linguistic Steganography with Diffusion Language Model. In The Thirty-ninth Annual Conference on Neural Information Processing Systems. [20] Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog 1, 8 (2019), 9. [21] Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research 21, 140 (2020), 1–67. [22] Claude E Shannon. 1948. A mathematical theory of communication. The Bell system technical journal 27, 3 (1948), 379–423.
9