Conceptio › Archive › arXiv CS
arXiv CSopen access

MarkSec: Capability-Aware Evaluation of Adversarial Attacks Against LLM Watermarks

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

MarkSec: Capability-Aware Evaluation of Adversarial Attacks Against LLM Watermarks Kairong Li

Zhikun Zhang Xiao Ren

Yunjun Gao∗

Zhejiang University, Hangzhou, China {kairong.li,zhikun,renx,gaoyj}@zju.edu.cn

arXiv:2609.16681v1 [cs.CR] 15 Sep 2026

Abstract LLM watermarking has emerged as a promising solution for tracing the origin of LLM-generated content to help limit accidental or deliberate misuse. It faces various adversarial attacks: stealing, which recovers watermark-related information; scrubbing, which removes watermark signals from marked text; and spoofing, which forges text that a detector accepts as watermarked without actual watermark embedding. However, these attacks are often studied in isolation, leaving their connections unclear. Current evaluations also often lack shared detector calibration, metric definitions, and reporting semantics, making it difficult to compare the risks faced by different watermarks. In addition, attack success and text quality are measured separately, making it difficult to identify attacks that are both effective and quality-preserving. To this end, we propose MarkSec, a general framework that unifies analyses of stealing, scrubbing, and spoofing. Building on this framework, we evaluate different attacks under a unified reporting protocol. We introduce a quality-constrained attack success metric to assess attack effectiveness and text quality jointly and support comparisons across watermark methods. Evaluations across representative watermark families, attack methods, LLMs, and datasets reveal three key findings. First, attacks that appear strongest by watermark removal alone can fall behind general rewriting under a criterion that jointly measures success and text quality. Second, general rewriting remains a strong baseline across watermark families, while its advantage over other scrubbing attacks varies by family. Third, in a case study of one watermark family, stealingbased scrubbers often underperform the best general-scrubbing baselines when text quality is required. These results show that apparent attack winners depend on text-quality constraints, attack generality, and capability assumptions.

Keywords LLM Watermarks, Watermark Security, LLM Content Provenance

1

Introduction

Large language models (LLMs) are increasingly applied in software development [13, 26], education [33, 37], and online content production [20, 34]. As model outputs become more fluent and human-like, distinguishing LLM-generated text from humanwritten text becomes more difficult [1, 35], which creates provenance and misuse concerns. For instance, schools and universities have debated or restricted ChatGPT use because of concerns about plagiarism and academic dishonesty [7, 9]. AI-generated spam and low-quality synthetic content are increasingly polluting online platforms and search results [28]. These trends make tracing ∗Corresponding author.

LLM-generated content increasingly important in real-world LLM deployments [6]. LLM watermarking has therefore emerged as a promising solution because it embeds machine-verifiable signals during generation [16] and does not depend only on post-hoc detection, which is fragile under paraphrasing [17, 23]. LLM watermarking is useful only if it remains reliable and secure under adversarial manipulation. In practice, it faces three main threat types. Stealing attacks recover watermark-related information, such as key-dependent token preferences or detector scoring rules, and can use it to support stronger downstream attacks [14]. Scrubbing attacks remove or weaken watermark signals in marked text through paraphrasing, rewriting, or targeted editing [4, 31]. Spoofing attacks forge watermark-like outputs without actual watermark embedding, potentially creating false tracing signals [3]. Despite the broad attack surfaces, existing studies investigate these attacks mostly in isolation. Furthermore, these attacks are rarely evaluated with shared detector calibration, metric definitions, and reporting semantics. Prior studies use different models, watermark families, datasets, attack budgets, and reporting protocols, making it difficult to compare attack results across studies [19, 32]. In addition, the attack success rate and text quality are often measured separately, making it hard to identify attacks that are both effective and quality-preserving. As a result, it remains unclear how these attacks relate to one another, which attacks act as broad baselines versus narrow specialists, and which vulnerabilities matter most in practice. Unified Framework. To address these issues, we study stealing, scrubbing, and spoofing as a connected attack space. Our key observation is that the three attack types are connected by how watermark-related information is obtained and reused. In particular, stealing can recover reusable artifacts, such as token preferences, detector statistics, or proxy distributions. The same artifacts can then support downstream scrubbing, which removes watermark evidence from existing text, or spoofing, which creates false attribution by generating watermark-like text. Motivated by this observation, we introduce a unified taxonomy that organizes these attacks by objective, capability, recovered artifact, and downstream use. Experimental Evaluation. Building on this taxonomy, we evaluate adversarial attacks against LLM watermarks under a unified benchmark protocol. Our evaluation covers representative watermark families, attack methods, datasets, and LLMs. We report quality-constrained attack success, QSR, which jointly counts recorded removal and acceptance by a declared automatic quality gate. The main general-scrubbing benchmark focuses on three questions: how robust watermarks are overall, whether attack strength is broad or family-conditioned, and how rankings change

Li et al.

under stricter Prometheus quality gates. Two further analyses examine Base–Instruct KGW differences across prompting protocols and the use of recovered watermark information for targeted scrubbing or spoofing. We also analyze quality-gate sensitivity across all six watermark families on common Llama/Qwen support. Main Findings. Our evaluation shows that the reported attack ordering depends on the success definition and the automatic quality gate. On C4, SIRA leads raw removal, but LLMP leads quality-aware scrubbing; DIPPER remains the closest qualityaware alternative in several cells. Watermark removal alone can therefore overstate practical attack strength. Family-conditioned analysis further shows that LLMP is the strongest quality-aware general-scrubbing baseline on C4, with the narrowest margins under UW and KGW. LLMP achieves the highest mean QSR on C4 at all three Prometheus gates. The Instruct–Base QSR gap is smaller under wrapper-disabled prompting than under native prompting in the evaluated KGW settings. Among the stealing-based scrubbers, B4 leads QSR@3 in all four KGW settings; in spoofing, DE-MARK-Sp reaches 98–100% detector acceptance and 47–64% QSSR@3. The results call for a capability-aware reading: raw removal, quality-aware success, and recovered-artifact assumptions must be reported together. These findings show that apparent attack winners depend on text-quality constraints, watermark family, deployment format, and attacker capability assumptions. MarkSec. We implement MarkSec as a modular and reusable tool for evaluating adversarial attacks against LLM watermarks. MarkSec provides common interfaces for watermark generation, attack execution, detector calibration, text-quality evaluation, and result reporting. It supports the joint study of stealing, scrubbing, and spoofing while keeping their capability assumptions explicit. The reporting pipeline groups general scrubbing, stealing-based scrubbing, and spoofing results by their access assumptions and downstream objectives.

2 Preliminaries 2.1 LLM Text Generation Most modern text-generation LLMs are Transformer-based autoregressive models. Given an input prompt, the LLM first predicts a probability distribution over its vocabulary, which consists of all possible tokens. It then selects one token from this distribution and appends it to the current text. This process repeats until the model generates an end-of-sequence token or reaches a maximum length. Formally, we view an LLM as a probabilistic generator parameterized by 𝜃 . Let 𝒱 denote the token vocabulary, where |𝒱| is the vocabulary size. Given a prompt sequence x(𝑝) = {𝑥1 , … , 𝑥𝑀 }, the model generates a response sequence x(𝑟) = {𝑥𝑀+1 , … , 𝑥𝑀+𝑇 } token by token. At each generation step 𝑡 , the model takes the prompt and all previously generated tokens x<𝑡 = {𝑥1 , … , 𝑥𝑡−1 } as input. It computes a logit vector l𝑡 ∈ ℝ|𝒱| , where each entry gives an unnormalized score for one token in the vocabulary. The logits are then normalized by a softmax function to produce a probability distribution 𝑃𝑡 over the vocabulary:

𝑃𝑡 (𝑣 ∣ x<𝑡 ) =

exp(l𝑡 [𝑣]) , ∑𝑣 ′ ∈𝒱 exp(l𝑡 [𝑣 ′ ])

𝑣 ∈ 𝒱.

(1)

The next token 𝑥𝑡 is selected from this distribution using a decoding strategy such as multinomial sampling, greedy decoding, or beam search.

2.2

LLM Watermarks

An LLM watermark is a machine-verifiable signal embedded into LLM-generated text for later provenance verification. A watermarked generation process should produce text that remains natural to human readers while carrying detectable statistical evidence. In general, LLM watermarks can be categorized by when and where the watermark signal is introduced. Model-parameter watermarks embed the signal into model-side parameters or learned watermark components, for example, by modifying parameters, finetuning behavior, or trigger-response patterns [21, 22]. Inferencetime watermarks embed the signal during decoding by modifying logits, token probabilities, or the sampling rule without changing model parameters [6, 16]. Post-hoc watermarks introduce the signal after ordinary text generation, typically by editing, selecting, or filtering outputs according to a watermark rule [18]. In this paper, we focus on inference-time LLM watermarks. As described in Section 2.1, an LLM first produces token logits, then converts them into a probability distribution (i.e., posterior), and finally samples the next token. Therefore, inference-time LLM watermarks can be further classified into three types according to where they intervene in this pipeline: logits-based, posterior-based, and sampling-based. Logits-Based Watermarks. The general idea is to modify the model logits so that some key-dependent tokens become more likely to be sampled. At generation step 𝑡 , a key-dependent rule assigns a bias to each candidate token, and the watermarked logits can be written as l𝑊 𝑡 [𝑣] = l𝑡 [𝑣] + 𝑏𝜉 ,𝑡 (𝑣),

(2)

where l𝑡 [𝑣] is the original logit of token 𝑣 , 𝜉 is the watermark key, and 𝑏𝜉 ,𝑡 (𝑣) is a key-dependent bias. After softmax, tokens with larger positive bias receive higher sampling probability. The detector then tests whether the generated sequence contains more favored tokens than expected under unwatermarked generation. KGW-style watermarks are representative examples of this class. They use the watermark key and previous context to partition candidate tokens into green and red sets, add a positive bias to green-token logits during decoding, and detect the watermark by testing whether the green-token count is statistically significant [16]. Posterior-Based Watermarks. Posterior-based watermarks operate on the normalized next-token distribution rather than on the pre-softmax logits. At generation step 𝑡 , they apply a keydependent transformation to the original distribution:

𝑃𝑡𝑊 = T𝜉 ,𝑡 (𝑃𝑡 ),

(3)

where 𝑃𝑡 is the original next-token distribution and T𝜉 ,𝑡 is a keydependent transformation. This transformation reallocates probability mass across tokens according to the watermark key while aiming to preserve the overall generation quality. The detector applies the same key-dependent rule to the observed sequence and tests whether the accumulated token statistics are more consistent with the watermarked distribution than with ordinary generation.

MarkSec : Capability-Aware Evaluation of Adversarial Attacks Against LLM Watermarks

UW is a representative example of this class. It constructs a distribution-preserving transformation so that the marginal output distribution remains close to the original model distribution while still leaving a detectable keyed signal [10]. Sampling-Based Watermarks. These methods keep the probability distribution largely unchanged, but use a key-dependent sampling rule to choose the next token. Formally, a sampling-stage watermark can be written as

𝑥𝑡 = Sample𝜉 ,𝑡 (𝑃𝑡 ),

(4)

where Sample𝜉 ,𝑡 denotes a key-dependent sampling rule. The detector applies the same key-dependent rule to the generated sequence and accumulates a statistic that measures whether the observed tokens are more consistent with the watermarked sampling process than with ordinary sampling. SynthID [6] is a representative example of this class. It uses tournament sampling during decoding. A key-dependent random function scores candidate tokens, and the sampler selects winners through pairwise token competitions, leaving a recoverable statistical signal in the generated sequence [6]. Watermark Detection. For all the above classes, watermark detection can be abstracted as a key-dependent hypothesis test. Given a text sequence x and a watermark key 𝜉 , a detector computes a score 𝑆𝜉 (x) that measures how strongly the text matches the expected watermark signal. The detector then compares this score with a threshold 𝜏 :

𝐷𝜉 (x) = 1[𝑆𝜉 (x) ≥ 𝜏 ].

(5)

Different watermark methods instantiate 𝑆𝜉 differently, such as a green-token count for logits-based schemes, a distributional consistency score for posterior-based schemes, or a key-dependent sampling statistic for sampling-based schemes. This unified detection view allows us to define attacks and evaluation metrics without depending on one specific watermark design.

3 Attacks Against LLM Watermarks 3.1 Threat Model Attack Objectives. Stealing aims to recover reusable watermarkrelated information, such as token preferences, counting rules, green-list structure, or a surrogate distribution. Scrubbing aims to transform a watermarked text 𝑦𝑤 into an attacked text 𝑦𝑎 that is no longer detected as watermarked while preserving text quality. Spoofing aims to generate a forged text 𝑦𝑠 that is accepted as watermarked without being produced by the genuine watermarking process. Attacker Access. We assume the attacker does not know the model parameters 𝜃 or the secret watermark key 𝜉 . The attacker may have black-box access to the watermarked generation API. Depending on the attack setting, the attacker may also receive the original prompt 𝑝 , the generated watermarked text 𝑦𝑤 , or victimside token information such as probabilities, logits, or top-𝑘 scores. These access differences define the attack capabilities and determine which attacks are directly comparable in our evaluation. We do not assume access to hidden system prompts or proprietary serving templates. When prompt context is available, it refers to the user-facing original prompt.

3.2

Unified Framework

We organize attacks against LLM watermarks by their operational role. This avoids treating watermark attacks as a flat list of unrelated methods. Our key observation is that stealing, scrubbing, and spoofing are connected by how watermark-related information is obtained and used. Stealing is an information-recovery objective. It recovers information that can later support stronger attacks. For example, a recovered green list, token-color map, or count table can be reused for targeted scrubbing or spoofing, while a distilled proxy distribution is used in this paper only for stealing-based scrubbing. Scrubbing is a removal objective. It takes watermarked text as input and tries to erase the watermark signal while preserving text quality. We distinguish between direct scrubbing and stealing-based scrubbing. Direct scrubbing attacks rewrite, edit, or search over the given watermarked text without first recovering a reusable watermark-related artifact. Stealing-based scrubbing attacks first recover watermark-related information, such as token preferences or green-list structure, and then use this information to guide removal. Spoofing is a forgery objective. It tries to create new text that is accepted as watermarked by the detector, even though the text was not produced by the genuine watermark process. Spoofing therefore attacks the attribution guarantee of the watermark rather than its removal robustness. In the methods we study, spoofing is usually downstream of stealing or inference because the attacker must recover enough watermark-related information to reproduce the watermark signal and trigger detection. This relationship is illustrated in Figure 1. Direct scrubbing attacks form the main comparable removal setting because they operate on the same input type and do not require a prior recovery step. Stealing-based scrubbing and spoofing are analyzed separately because they rely on additional recovery steps or reusable artifacts. This separation keeps the main comparison focused while still capturing stronger attack paths in the full threat surface. The taxonomy also defines the comparison boundary used in our experiments. Table 1 summarizes stealing-based methods by their stealing source, recovered artifact, and downstream use. Table 2 summarizes the direct scrubbing attacks used in the main comparison. This split keeps general scrubbing results comparable and still covers stronger attack paths that rely on additional capabilities. Role Notation. Some attack papers support multiple downstream objectives. To avoid mixing different roles of the same method, we attach an objective suffix when a method appears in more than one part of the taxonomy. We use St for the stealing role, Sc for the scrubbing role, and Sp for the spoofing role. For example, JSVSt denotes the information-recovery stage of JSV, while JSV-Sp denotes its use for spoofing generation.

3.3

Stealing Attacks

Stealing attacks recover reusable watermark-related information before downstream use. The recovered artifact can be explicit, such as a green list or token-color map, or implicit, such as a context count table or proxy distribution. Its role is to expose information that later supports scrubbing or spoofing.

Li et al.

Scrubbing Attacks

Stealing Attacks MIP-St

JSV-St

Stealing-based Scrubbing

B4-St

MIP-Sc DE-MARK-St

SCTS-St

JSV-Sc

DE-MARK-Sc

B4-Sc SCTS-Sc

General Scrubbing

Spoofing Attacks JSV-Sp

LLMP

DIPPER

DE-MARK-Sp

RW

SA

SIRA

Figure 1: Attack taxonomy organized by adversarial objective: Stealing, Scrubbing, and Spoofing. Table 1: Required capability of stealing-based attacks. Method MIP [40] SCTS [36] DE-MARK [3] JSV [14] B4 [11]

Online

Prompt

Scores

Spoof

✗ ✓ ✓ ✗ ✗

✗ ✓ ✓ ✗ ✓

✗ ✗ ✓ ✗ ✗

✗ ✗ ✓ ✓ ✗

SCTS-St [36]. Wu and Chandrasekaran propose SCTS for KGWstyle watermarks, where token choices are biased by a hidden redgreen partition. In its stealing role, SCTS estimates token-color information through black-box querying. The attacker queries the watermarked generator with adaptive query prompts and observes which candidate tokens are more likely to appear. These observations are used to infer a token-color map that approximates the hidden green and red sets. This recovered color map is the artifact later used by SCTS-Sc for targeted substitution. JSV-St [14]. Jovanović et al. study whether hidden watermark behavior can be approximated from generated samples. In the stealing role, JSV builds context-dependent statistics from base and watermarked continuations. These statistics capture how the watermark shifts token choices under different contexts. The recovered count table is reusable because it summarizes watermark behavior beyond one attacked sample. This artifact can later guide JSV-Sc for removal or JSV-Sp for false attribution. B4-St [11]. Huang et al. formulate B4 around a black-box view of the victim watermarked generator. Its stealing role is to learn a proxy distribution from victim-generated data. This proxy approximates how the watermarked system allocates probability mass over candidate continuations. The learned distribution becomes a reusable artifact for downstream rewriting. B4-Sc then uses this proxy to search for attacked text that reduces the watermark signal while preserving fidelity.

Table 2: Capability questions for general scrubbing attacks. Method DIPPER [17] LLMP SIRA [4] RW [38] SA [2]

Prompt

Scores

Locate

Local

Search

✗ ✗ ✗ ✓ ✓

✗ ✗ ✗ ✗ ✓

✗ ✗ ✓ ✗ ✓

✗ ✗ ✓ ✓ ✓

✗ ✗ ✗ ✓ ✗

MIP-St [40]. Zhang et al. formulate watermark stealing for greenlist-style schemes as a mixed-integer optimization problem. The attacker uses the collected watermarked and reference text to infer which tokens are likely to belong to the hidden green list. The optimization stage produces a stolen green-list artifact. This artifact is reusable because it can guide future token replacement without repeating the full recovery process. MIP-Sc uses the recovered green list for targeted watermark removal. DE-MARK-St [3]. Chen et al. design DE-MARK to infer watermark rules through token-level querying. In its stealing role, the attacker sends prompt-guided queries to the victim system and observes how candidate tokens behave under the hidden watermark. The query process estimates which tokens are favored by the watermark and may also recover parameter-level information. The resulting query evidence becomes a reusable guide for downstream attack.

3.4

Scrubbing Attacks

Scrubbing attacks form the main general-scrubbing comparison in MarkSec. These attacks take already watermarked text as input and try to remove the watermark signal while preserving text quality. They differ mainly in what information they use during rewriting, such as the text alone, the original prompt, proxy-model signals, or victim-side scores. DIPPER [17]. As a generic paraphrasing baseline, DIPPER rewrites watermarked text into a semantically similar form. Its

MarkSec : Capability-Aware Evaluation of Adversarial Attacks Against LLM Watermarks

intuition is that paraphrasing changes surface tokens and local token patterns, which can disrupt watermark signals embedded in the original wording. The attack does not use the watermark key, detector feedback, victim-side scores, or a recovered artifact. In MarkSec, DIPPER represents text-only direct scrubbing. LLMP. Our LLMP baseline uses an LLM to rewrite the full watermarked continuation. It treats scrubbing as sequence-level rewriting: the model is prompted to preserve the meaning of 𝑦𝑤 while producing a new expression. This attack can preserve high-level semantics because it rewrites the whole continuation at once. It does not use detector access, victim-side token scores, or a prior stealing stage. In MarkSec, LLMP represents general LLM-based direct scrubbing. SIRA [4]. Cheng et al. propose SIRA to avoid rewriting the entire sequence. The attack uses self-information from a proxy model to locate spans that appear suspicious under the watermarked text distribution. It then rewrites only these high-suspicion spans. This targeted strategy aims to remove watermark evidence while changing fewer parts of the text. In MarkSec, SIRA represents proxyguided direct scrubbing. RW [38]. RW performs prompt-aware rewriting through a quality-guided search process. The attack generates candidate rewrites and keeps candidates that better balance watermark removal and text quality. The original prompt provides reference context during the search, helping the attack preserve the intended continuation. This makes RW different from one-shot paraphrasing baselines. In MarkSec, RW represents direct scrubbing with prompt-aware search and explicit quality control. SA [2]. Chang et al. propose SA as a score-aware editing attack. The attack uses victim-side token information, such as probabilities, logits, or top-𝑘 scores, to identify tokens that contribute strongly to the watermark signal. It edits the text at the token level and searches for alternatives that preserve the local context. This gives SA finer control than full-sequence paraphrasing or spanlevel rewriting. In MarkSec, SA is a direct scrubber with stronger attack-time access because it uses victim-side token scores. SCTS-Sc [36]. The scrubbing role of SCTS uses the token-color map recovered by SCTS-St. Given a watermarked continuation, the attack identifies tokens that are likely to be green under the hidden KGW-style partition. It then replaces these tokens with semantically similar alternatives that are less likely to be green. This directly targets the green-token count used by the detector. We classify SCTS-Sc as stealing-based scrubbing because its substitutions depend on the prior color-recovery stage. JSV-Sc [14]. JSV-Sc uses the context-dependent statistics recovered by JSV-St to guide removal. The attack identifies token patterns that are likely to reflect watermarked generation under the recovered statistics. During rewriting or token selection, it avoids patterns that make the text look watermarked. This role uses stolen watermark behavior to reduce detector evidence in an existing continuation. We therefore treat JSV-Sc as stealing-based scrubbing. B4-Sc [11]. B4-Sc uses the proxy distribution learned by B4-St to guide rewriting. The attack searches for continuations that reduce the watermark signal while staying close to the original text under the learned proxy. Its core trade-off is removal versus fidelity: aggressive changes may evade detection but damage text quality. The proxy distribution helps the attack navigate this trade-off in

a black-box setting. We classify B4-Sc as stealing-based scrubbing because its removal strategy depends on a reusable learned distribution. MIP-Sc [40]. MIP-Sc uses the green list recovered by MIP-St. The attacker searches for edits that reduce the number of green-list tokens in the watermarked continuation. Because KGW-style detection depends on green-token overrepresentation, replacing green tokens with non-green alternatives can reduce the detection score. This scrubbing role is targeted and watermark-specific. We classify MIP-Sc as stealing-based scrubbing because it depends on the stolen green-list artifact. DE-MARK-Sc [3]. DE-MARK-Sc uses query-derived watermark evidence to remove signals from an existing continuation. The attack identifies token choices likely favored by the hidden watermark rule. It then edits the text to avoid these choices while preserving meaning. This approach uses the recovered rule in the opposite direction of spoofing. We classify DE-MARK-Sc as stealingbased scrubbing because its removal strategy is guided by prior stealing. Direct scrubbing attacks share a common goal: transforming watermarked text into attacked text without recovering a reusable watermark artifact. They differ in the side information and resources used for rewriting. All direct scrubbers in our study are detectorfree; Table 2 asks whether an attack uses prompt context, victimside scores, suspicious span/token localization, local edits, or iterative search.

3.5

Spoofing Attacks

Spoofing targets false attribution rather than watermark removal. A spoofing adversary generates text that is accepted by the victim detector as watermarked, even though the text was not produced by the actual watermark embedding process. In our taxonomy, spoofing is usually downstream of stealing because the attacker must reproduce enough of the watermark signal to trigger detection. We therefore analyze spoofing-related methods separately from direct scrubbing attacks. JSV-Sp [14]. JSV-Sp uses context-dependent statistics recovered from JSV-St for detector-targeted generation. Instead of avoiding watermarked patterns, the attack biases generation toward them. The goal is to make unwatermarked text resemble outputs from the watermarked generator. This creates a false-attribution risk because the detector may accept forged text as authentic. We classify JSV-Sp as spoofing because it uses recovered watermark behavior to trigger detection. DE-MARK-Sp [3]. DE-MARK-Sp uses the watermark rules recovered by DE-MARK-St during generation. The attacker favors tokens that the hidden watermark mechanism is expected to favor. This pushes the generated continuation toward the detector’s expected watermark pattern. Unlike DE-MARK-Sc, which avoids watermark-favored tokens, DE-MARK-Sp uses the same recovered information to imitate the watermark signal. We classify DE-MARK-Sp as spoofing because it forges text that can be falsely accepted as watermarked.

Li et al.

1 Generation

2 Targeted Attack

3 Evaluation

Direct Scrubbing

Datasets Watermarks Models

Stealing

Scrubbing

Stealing

Spoofing

Shared Infrastructure

Configuration

Logging

4

Reporting

Detect Score

Raw Records

Threshold

Summaries

Quality Score

Tables/Plots

Job Manager

Registry factory

Figure 2: System overview of MarkSec.

4

MarkSec

In this section, we introduce MarkSec, a modular toolkit for evaluating adversarial attacks against LLM watermarks. MarkSec operationalizes the unified framework in Section 3 and turns it into an executable evaluation pipeline. MarkSec runs attack implementations and makes their results comparable under explicit capability assumptions. In particular, MarkSec jointly supports stealing, scrubbing, and spoofing, while keeping direct scrubbing attacks separate from stronger attacks that first recover reusable watermarkrelated artifacts. Figure 2 presents the system overview. Existing studies on LLM watermark evaluation mainly focus on implementing watermark methods or reporting benchmark results [19, 24, 32]. Attack evaluation requires more than collecting attack scripts. Different watermark methods expose different generation interfaces, detector APIs, score directions, and threshold choices. Different attacks also require different inputs, access assumptions, recovered artifacts, and downstream goals. Without a shared toolkit, these differences can easily lead to results that are difficult to reproduce or compare across papers. MarkSec addresses this gap by providing a configuration-driven pipeline that loads common components, executes attacks under explicit capability assumptions, and reports attack success together with text quality. Modules. MarkSec consists of four main modules: input and generation, attack orchestration, detection and metrics, analysis and reporting. Input and Generation. This module loads prompt datasets, target LLMs, watermark methods, and generation configurations. It builds on existing watermark implementations and wraps them into a common generation pipeline. For each run, it records the prompt, the watermarked text, the watermark configuration, and the decoding metadata needed by later modules. This standardized input format decouples downstream attacks and metrics from watermark-specific implementation details. Attack Orchestration. This module executes attacks according to their capability assumptions. Direct scrubbing attacks take watermarked text as input and produce attacked text under a shared removal protocol. Stealing-based methods first run a recovery stage,

such as querying, estimation, or distillation, and then pass the recovered artifact to downstream scrubbing or spoofing. The module records whether an attack uses the original prompt, victim-side scores, helper models, or reusable recovered artifacts. This makes the comparison boundary explicit at execution time rather than leaving it implicit in separate attack scripts. Detection and Metrics. This module applies the detector associated with each watermark method to natural, watermarked, and attacked outputs. Because detectors may produce scores with different scales or directions, MarkSec normalizes score direction when needed and supports shared threshold calibration. The module computes detection metrics, attack success metrics, text-quality metrics, and quality-constrained attack success. This module lets MarkSec compare attack success together with text quality. Analysis and Reporting. This module stores per-sample records, per-run summaries, metric details, and aggregate exports. The raw records preserve prompts, watermarked text, attacked text, detector outputs, quality metrics, and attack-specific metadata. The summary reports aggregate detection results, quality results, sample statistics, and the full run configuration. MarkSec also maintains a registry-backed reporting layer that supports large-scale tracking, analysis scripts, and paper-table exports. This reporting design keeps raw experimental artifacts available while allowing higherlevel benchmark views to be regenerated. Modular Design. MarkSec is organized around configurationloaded wrappers for attacks, watermark generators and detectors, datasets, metrics, and experiment runners. New attacks are integrated through a shared attack interface and registered in an attack factory. This interface supports preparation, optional recovery or learning, editing, and saving or loading learned artifacts. Watermark generation and detection are loaded through wrapper interfaces, while datasets and metrics are managed through registries and metric managers. This design does not require every component to share the same internal implementation. Instead, it provides a common boundary where heterogeneous components can be loaded, executed, and reported through the same pipeline. Supported Workflows. MarkSec supports three main workflows. First, it supports unified general-scrubbing evaluation. In

MarkSec : Capability-Aware Evaluation of Adversarial Attacks Against LLM Watermarks

this workflow, comparable removal attacks are run under shared settings and summarized in a common report. Second, it supports capability-aware analysis of stronger attacks. In this workflow, attacks that first recover watermark-related information are evaluated together with their recovery assumptions, recovered artifacts, and downstream use. Third, it supports plug-and-play extension. Researchers can add new watermark methods, attack methods, models, datasets, or metrics while reusing the same generation, detection, and reporting pipeline. Using MarkSec. Watermark designers can use MarkSec to assess whether a watermark is vulnerable to direct scrubbing or to stronger attacks that first recover watermark-related information. Attack researchers can use MarkSec to compare new attacks against existing baselines under a unified reporting protocol. Practitioners can use MarkSec to interpret watermark scrubbing success together with output utility, so that watermark removal is not mistaken for practical attack success when the resulting text is no longer useful. Because MarkSec separates direct scrubbing from stronger stealing- and spoofing-related methods, its reports make clear which results depend on additional attacker capabilities. This makes MarkSec both the implementation vehicle of our evaluation and a reusable toolkit for future studies of LLM watermark robustness.

5 Experimental Setup 5.1 Evaluation Settings LLM Watermarks. We conduct experiments on six representative watermark schemes, including three logits-based schemes, KGW, Unigram, and SIR; two posterior-based schemes, UW and DIP; and one sampling-based scheme, SynthID. DIP is a watermark family and is unrelated to the DIPPER paraphraser attack. For implementation, we build on the MarkLLM repository [24] where available and adapt their input-output format to our protocol. Evaluated SynthID Implementation. We evaluate MarkLLM’s SynthID-Text implementation in non-distortionary mode with the mean detector. The detector averages keyed 𝑔 -values over channels and eligible completion n-grams, excluding repeated contexts and positions from EOS onward. Appendix F specifies its configuration and the cross-watermark sensitivity analysis. Attacks. We include five general scrubbing attacks in the main experiments: DIPPER, LLMP, RW, SA, and SIRA. These attacks cover paraphrasing, LLM rewriting, quality-guided search, score-aware editing, and targeted proxy-guided rewriting. We further include five stealing-based methods in the stronger-capability analysis: MIP, SCTS, DE-MARK, JSV, and B4. Their objectives and capability assumptions are summarized in Section 3. For SCTS, we implement the attack in PyTorch [25]; for the remaining attacks, we adapt the authors’ official implementations to the benchmark settings. Detailed attack defaults are reported in Appendix D. Target LLMs. The main benchmark evaluates three widely used open-source LLM families: Llama 3.1 8B Instruct [8], Qwen2.5 7B Instruct [29], and Mistral 7B Instruct v0.3 [12]. These targets cover different tokenizer and model families while keeping the matrix computationally tractable. All local models are loaded from fixed

checkpoints to avoid inconsistencies caused by remote model updates. For the instruction-tuning sensitivity study, we additionally compare Qwen2.5-7B Base and Instruct checkpoints under the same named KGW settings. Prompt Datasets. We conduct experiments on three prompt datasets: C4 [30], Dolly-15K [5], and MMW BookReport, the bookreport prompt set from Mark My Words (MMW) [27]. We select 100 prompts per dataset. To better match instruction-tuned models, we prepend an instruction-style continuation prompt to each C4 sample. Sample Universe. The main matrix uses 26,096 recorded outputs from 261 runs conducted in April 2026. Each attack block is computed on its own completed-output population; execution completion includes empty or otherwise invalid outputs. Generated source texts and detector thresholds can differ across attack runs, so cross-method comparisons aggregate configuration-level rates. Within a run, paired detection refers to the same sample before and after attack. The records include 35 samples with insufficient context for SIR detection and two empty SA outputs with unavailable after-detection. Structural exclusions are marked “N/A”; undefined measured rates are marked “–”.

5.2

Evaluation Metrics

Robustness Metrics. True positive rate (TPR) is the fraction of watermarked texts detected as watermarked. False positive rate (FPR) is the fraction of unwatermarked texts incorrectly detected as watermarked. Detector calibration targets an FPR of 1%; results use the decision and threshold recorded for each run. This is a calibration target, not a separately measured test-set FPR. When a detector’s raw score uses the opposite direction, we invert the decision direction so that all reported results share the same interpretation. Let 𝐶 contain the completed-output sample IDs of a run. Define 𝑆𝑖 = 1 only when both detector states are present and sample 𝑖 moves from detected to undetected; otherwise it earns no removal credit. Let 𝐶+,pair contain samples with both states observed and a positive before-state, and let 𝑃𝑖 be the recorded Prometheus score. We compute ASR =

∑𝑖∈𝐶 𝑆𝑖 , |𝐶+,pair |

(6)

∑𝑖∈𝐶 1[𝑆𝑖 = 1, 𝑃𝑖 ≥ 𝑢] . (7) |𝐶| Unavailable quality or detector evidence cannot pass the jointsuccess gate. ASR uses only before-positive samples with both detector states observed; it is undefined when this set is empty. TPR𝑏 is the positive fraction among all available before-detection results, regardless of after-detection availability. ASR macro-averages and ranks use only cells with defined ASR; QSR macro-averages use all available cells, including those with zero detected inputs. Each included cell receives equal weight, and figure data report the contributing cell counts. The ASR–QSR difference therefore combines eligibility, measurement coverage, and the quality gate. Text-Quality Metrics. We use Prometheus (Prom) [15] to assess output quality at thresholds 2, 3, and 4. BERTScore with robertalarge and PSP provide complementary similarity measures. QSR denotes removal jointly satisfying the specified automatic quality QSR@𝑢 =

Li et al.

gate; QSR@3 is the default. BERTScore and PSP compare full-text fields that include a shared prompt when present (Appendix E.2). Quality evaluation uses automatic scores; supplementary content screening uses deterministic rules and Codex-assisted source– output inspection, with no independent human annotation. Source adequacy and preservation are assessed separately: a faithful edit can retain an inadequate source.

5.3

Evaluation Matrix And Reporting Scope

Each benchmark run is defined by a tuple (𝐷, 𝑀, 𝑊 , 𝐴), where 𝐷 represents a prompt dataset, 𝑀 a target LLM, 𝑊 a watermark scheme, and 𝐴 an attack method. To ensure readability and protocol control, the main text focuses on the canonical C4 slice. Additional Dolly-15K and MMW BookReport slices are included in the appendix to examine dataset dependence. We report general scrubbing attacks as the main comparison surface because they share the same downstream objective: given watermarked text, produce an attacked text with weakened watermark evidence while preserving utility. Stealing- and spoofingrelated methods are reported separately because they require additional capabilities, such as query access, query budgets, prebuilt corpora, recovered artifacts, or attack-side generation control. This shared-protocol comparison answers how robust watermark families are under general scrubbing. The stronger-capability analysis asks when recovered watermark-related information changes the risk picture through targeted scrubbing or spoofing.

6

Experimental Results

In this section, we answer five empirical questions. RQ1. How robust are LLM watermarks under the shared generalscrubbing protocol? RQ2. Are general scrubbing attacks broad quality-aware baselines or watermark-family specialists? RQ3. How does Prometheus-based utility filtering change attack rankings? RQ4. How do the Base–Instruct observations vary with prompting protocol? RQ5. How does recovered watermark-related information change scrubbing and spoofing risk? The first three questions use the shared general-scrubbing benchmark, where each attack receives watermarked text and outputs attacked text under the same reporting protocol. RQ1 measures robustness under this shared protocol. RQ2 separates broad qualityaware baselines from family-aligned specialists. RQ3 asks whether rankings survive Prometheus-gated utility filtering. RQ4 compares Base–Instruct outcomes under two prompting protocols. RQ5 studies whether recovered watermark-related information changes downstream scrubbing and spoofing risk. We report RQ5 separately from the general-scrubbing leaderboard because its access assumptions are different.

6.1

RQ1: Overall General-Scrubbing Robustness

We first evaluate general scrubbing on the canonical C4 slice. Table 3 reports pre-attack detectability, raw attack success, and

quality-constrained success for each watermark–victim–attack combination. We abbreviate pre-attack TPR as TPR𝑏 . Higher ASR and QSR mean stronger attacks. Bold marks the row-wise best QSR. We use “N/A” for combinations excluded by the benchmark protocol. In this table, “N/A” appears for SA under selected Mistral settings because SA requires a tokenizer-aligned helper model, and our Mistral setup does not include a matching small helper model. The relevant distinctions are baseline detectability, paired detector evasion, and automatic-quality-gated success. Two patterns stand out. First, attack success has to be interpreted relative to the pre-attack detector strength. Some watermark families start from much stronger detectability than others, so a high ASR on a weakly detectable setting is not equivalent to breaking a strongly embedded watermark. Second, no general scrubbing attack dominates all axes. SIRA leads the recorded C4 removal rate, while LLMP leads joint success under the automatic Prometheus gate. DIPPER remains the closest quality-aware alternative on C4. Figure 3 visualizes the difference between the two reported metrics. On C4, SA illustrates the danger of using ASR alone. It often drives detector scores down, but fewer records receive joint credit under the automatic gate and joint-rate denominator. LLMP and DIPPER lose less of their apparent success when QSR is applied. Across dataset slices, attack rankings differ between ASR and QSR@3. Finding 1. On C4, SIRA achieves the highest mean ASR, whereas LLMP achieves the highest mean QSR@3.

6.2

RQ2: Family-Specific Robustness Patterns

RQ2 asks whether general scrubbing attacks are broadly effective across watermark families or mainly strong when their mechanism aligns with a specific watermark. This question matters because overall averages can hide two different behaviors: a stable broad baseline and a narrow family specialist. To answer RQ2, we read the C4 matrix by watermark family and summarize the interaction in Figure 4. The main result is that quality-aware attack rankings are less fragmented than raw removal suggests. LLMP has the highest mean QSR@3 for every C4 watermark family. The family-conditioned view still matters because the margin changes by family. UW and KGW leave DIPPER close to LLMP, while SynthID shows the widest LLMP lead. Unigram, SIR, and DIP fall between these two cases. Figure 5 adds a supporting view over dataset–watermark cells. Unlike Figure 4, which summarizes family-level averages, this figure shows how stable each attack is across individual (𝐷, 𝑊 ) conditions. Each point is one cell, and each box summarizes the QSR@3 distribution for one attack. LLMP has the highest qualityaware success distribution across these cells. Other attacks either sit lower overall or show wider variation. This interaction view changes how we should interpret attack rankings. Overall averages identify LLMP as the strongest qualityaware attack on C4. Family-conditioned analysis explains where that lead is narrow and where it is robust. SIRA leads mean ASR on C4, while LLMP leads mean QSR@3 in every C4 watermark

MarkSec : Capability-Aware Evaluation of Adversarial Attacks Against LLM Watermarks

Table 3: C4 general-scrubbing results across watermark families, victim models, and attacks. DIPPER

Watermark Model

LLMP

SIRA

RW

SA

TPR𝑏 ↑ ASR↑ QSR↑ TPR𝑏 ↑ ASR↑ QSR↑ TPR𝑏 ↑ ASR↑ QSR↑ TPR𝑏 ↑ ASR↑ QSR↑ TPR𝑏 ↑ ASR↑ QSR↑

KGW

Llama Qwen Mistral

0.93 0.96 0.94

0.78 0.75 0.88

0.63 0.61 0.71

0.92 1.00 0.97

0.67 0.88 0.82

0.57 0.72 0.78

0.93 0.96 0.94

0.82 0.91 0.89

0.42 0.74 0.44

0.96 0.99 0.95

0.61 0.62 0.81

0.35 0.26 0.39

1.00 0.95 N/A

0.46 0.68 N/A

0.26 0.33 N/A

Unigram

Llama Qwen Mistral

0.72 0.81 0.90

0.75 0.85 0.60

0.46 0.57 0.50

0.77 0.86 0.81

0.75 0.98 0.78

0.53 0.78 0.58

0.72 0.81 0.90

0.96 0.95 0.88

0.43 0.65 0.47

0.73 0.80 0.91

0.88 0.64 0.59

0.30 0.24 0.31

0.73 0.75 N/A

0.49 0.89 N/A

0.19 0.22 N/A

SIR

Llama Qwen Mistral

0.62 0.69 0.91

0.62 0.79 0.55

0.33 0.41 0.43

0.55 0.70 0.91

0.78 0.96 0.69

0.40 0.63 0.61

0.51 0.77 0.91

0.84 0.95 0.73

0.27 0.62 0.38

0.49 0.71 0.93

0.71 0.76 0.80

0.23 0.32 0.32

0.59 0.70 N/A

0.76 0.91 N/A

0.26 0.24 N/A

UW

Llama Qwen Mistral

0.90 0.84 0.95

0.93 0.99 0.96

0.71 0.68 0.84

0.85 0.83 0.94

0.92 1.00 0.93

0.75 0.76 0.82

0.90 0.84 0.95

0.96 1.00 0.96

0.47 0.71 0.43

0.87 0.75 0.98

0.78 0.93 0.84

0.48 0.37 0.50

0.90 0.82 0.95

0.73 1.00 0.99

0.24 0.26 0.46

DIP

Llama Qwen Mistral

0.57 0.82 0.85

1.00 0.94 0.94

0.49 0.59 0.68

0.72 0.81 0.84

0.90 0.96 0.95

0.61 0.72 0.76

0.57 0.82 0.85

0.93 0.96 0.92

0.37 0.69 0.35

0.70 0.78 0.85

0.89 0.88 0.85

0.40 0.34 0.37

0.75 0.86 0.78

0.97 0.99 0.99

0.19 0.32 0.46

SynthID

Llama Qwen Mistral

0.95 0.98 0.97

0.80 0.85 0.69

0.67 0.70 0.58

0.99 0.93 0.99

0.89 0.97 0.81

0.85 0.85 0.76

0.95 0.98 0.97

0.87 0.98 0.74

0.43 0.87 0.40

0.96 0.94 0.96

0.86 0.88 0.82

0.47 0.42 0.42

0.96 0.95 0.97

0.95 1.00 0.91

0.22 0.27 0.38

DIPPER LLMP

SIRA RW

ASR QSR@3 1

DIPPER Attack rank

LLMP SIRA RW SA 0.00

SA

2 3 4 5

0.25

0.50 Mean rate

0.75

1.00

ASR

QSR C4

ASR

QSR

Dolly

ASR

QSR

Book

Figure 3: Mean ASR and QSR@3 on C4 and attack ranks across datasets. ASR averages exclude undefined cells; QSR averages include all available cells. family. Family conditioning mainly changes the margin, not the top QSR attack on C4. Finding 2. LLMP is the strongest quality-aware generalscrubbing baseline on C4, with the narrowest margins under UW and KGW. SynthID: Shared Trend and Model-Level Exceptions. SynthID shares the C4 family-average trend, but its Qwen cell gives SIRA QSR@3 = 0.87 versus 0.85 for LLMP; the Llama values are 0.43 and 0.85. Across these two models, SA removes 186 of 191 recorded pre-positive passages, but only 49 removals pass

Prometheus 3. For LLMP, 178 of 192 recorded pre-positive passages are removed and 170 pass the gate. For SynthID–LLMP, C4 has 192/200 pre-positive records, 178 removals, and 170 gated successes, whereas Dolly has 113/200, 112, and 104. Dolly’s lower joint rate, 0.52 versus 0.85, accompanies near-complete removal of its smaller eligible set.

6.3

RQ3: Utility-Aware Robustness Trade-Offs

RQ3 asks how the recorded joint-success rates change as the automatic acceptance threshold is tightened.

Li et al.

0.65

0.69

0.53

0.33

0.30

Unigram

0.51

0.63

0.52

0.28

0.21

SIR

0.39

0.55

0.42

0.29

0.25

UW

0.74

0.78

0.54

0.45

0.32

DIP

0.59

0.70

0.47

0.37

0.32

SynthID

0.65

0.82

0.57

0.44

0.29

DIPPER LLMP

SIRA

RW

SA

0.00

0.25

0.50 0.75 Mean QSR@3

1.00

Figure 4: Mean QSR@3 by watermark on C4, averaging available victim-model cells. UW denotes Unbiased.

Group mean QSR@3

1.00

0.75

DIPPER LLMP

SIRA RW

SA

1.0 0.8 Mean QSR

KGW

0.6 0.4 0.2 0.0

2

3 Prometheus gate

4

Figure 6: Mean QSR on C4 at Prometheus gates 2, 3, and 4. Each method averages its available victim-model cells. also remains the strongest attack at each reported threshold. SA degrades much faster. SIRA has the strongest raw ASR, but its qualityaware success drops once the utility gate tightens. DIPPER remains the closest quality-aware alternative to LLMP on C4. The counterintuitive point is that a stronger detector-side attack can become weaker under a stricter utility gate. Figure 7 adds a cell-level view. Each point is one watermark– victim-model unit. The plot shows how automatic quality scores and joint success vary across watermark and victim-model conditions. It complements the threshold view by exposing the variation hidden by aggregate attack rankings. Finding 3. LLMP achieves the highest mean QSR on C4 at all three Prometheus gates.

0.50

0.25

0.00 DIPPER

LLMP

SIRA

RW

SA

Figure 5: QSR@3 over 18 dataset–watermark groups per attack, each averaged over available models. Points show every group; boxes indicate quartiles and whiskers span the observed range. For readability, the main text uses two C4 views: threshold sensitivity and attack-conditioned cell variation. Appendix E.2 reports the full per-dataset matrices for detection, QSR, BERTScore, PSP, and Prometheus. Figure 6 makes the threshold effect explicit. LLMP retains the largest fraction of its QSR under stricter quality thresholds. LLMP

Cross-Watermark Sensitivity on Common Support. The supplementary analysis fixes Prometheus gates 2, 3, and 4 and uses 180 cells with the same two-model support for all six watermarks and five attacks, keeping the three datasets separate. For SynthID on C4, LLMP’s joint rates are 0.875, 0.850, and 0.770; its margins over the best other attack are 12.11, 16.65, and 24.50 percentage points. LLMP leads each C4 watermark family at all three gates, and each family’s margin is larger at 4 than at 2. SynthID has the largest C4 margin at gate 4. The Qwen exception favors SIRA at gates 2 and 3 but LLMP at gate 4, 0.78 versus 0.76. Table 14 reports the corresponding results for all six watermark families.

6.4

RQ4: Instruction-Tuning Sensitivity

RQ4 asks how the Base–Instruct difference in watermark robustness varies with prompting protocol. We compare Qwen2.5-7B and Qwen2.5-7B-Instruct on C4 at two KGW settings, ℎ=1, 𝛾 =0.25 and ℎ=1, 𝛾 =0.5, using LLMP and SIRA. The native condition uses automatic template selection; the wrapper-disabled condition omits the chat wrapper. The comparison contains 16 checkpoint– prompting–watermark–attack combinations. Each run uses its

MarkSec : Capability-Aware Evaluation of Adversarial Attacks Against LLM Watermarks

1

1.0

3 Mean Prometheus

QSR@3 1

1

1.0

RW

0.5

0.0

0.5

0.0

5

3 Mean Prometheus

5

1.0

LLMP

QSR@3

0.5

0.0

QSR@3

1.0

DIPPER

QSR@3

QSR@3

1.0

3 Mean Prometheus

5

0.5

0.0

1

3 Mean Prometheus

5

SA KGW Unigram SIR UW DIP SynthID Llama Mistral Qwen

0.5

0.0

SIRA

1

3 Mean Prometheus

5

Figure 7: Prometheus scores and QSR@3 on C4. Each point is one victim-model–watermark cell. Color identifies the watermark; shape identifies the model. generated source texts and recorded detector threshold, so the table reports configuration-level Instruct-minus-Base differences. Under native prompting, the Instruct-minus-Base pre-detection differences are −0.045 for 𝛾 = 0.25 and −0.015 for 𝛾 = 0.5. For 𝛾 = 0.25, LLMP and SIRA have native QSR differences of +0.170 and +0.160; disabling the wrapper reduces these descriptive differences to +0.020 and +0.090. For 𝛾 = 0.5, the native QSR differences are +0.180 and +0.210, compared with +0.060 and +0.120 without the wrapper. The corresponding ASR differences change sign in three of four wrapper-disabled entries. Finding 4. In the evaluated Qwen2.5-7B/KGW runs, the Instruct–Base QSR gap is smaller under wrapper-disabled prompting than under native prompting.

6.5

RQ5: Stealing-Based Attacks

RQ5 asks whether recovered watermark information changes the risk picture beyond the general-scrubbing benchmark. Unlike

RQ1–RQ3, these attacks assume that the attacker first steals or infers reusable watermark information, such as green-list structure, token preferences, or decision statistics. We compare them with strong general-scrubbing baselines, but we do not merge them into the same leaderboard. The five stealing-based methods share KGW as a target, enabling a comparison across four settings with ℎ ∈ {1, 3}, 𝛾 ∈ {0.25, 0.5}, and 𝛿 = 2. Mechanistically, ℎ defines how much prior context enters the favored-token rule, while 𝛾 defines the greentoken base rate. Recovered artifacts guide scrubbing by avoiding or replacing watermark-favored tokens; spoofing-facing use would reverse this direction and favor such tokens during generation. The runs use method-specific helper models, source texts, and recovery budgets. Execution time and recovery resources are reported separately in Table 6. Table 5 compares representative stealing-based scrubbing methods against the best general-scrubbing baselines on C4. It reports paired-removal ASR, QSR@3, sample counts, Prometheus scores, and differences from the highest general-scrubbing QSR at each setting.

Li et al.

Table 4: Instruct-minus-Base differences on Qwen2.5-7B/KGW under native and wrapper-disabled prompting. Setting

Protocol

ΔTPR𝑏

ΔASRLLMP

ΔASRSIRA

ΔQSRLLMP

ΔQSRSIRA

ℎ=1, 𝛾 =0.25 ℎ=1, 𝛾 =0.25 ℎ=1, 𝛾 =0.5 ℎ=1, 𝛾 =0.5

Native Wrapper-disabled Native Wrapper-disabled

-0.045 +0.000 -0.015 +0.020

+0.106 -0.070 +0.061 -0.028

+0.046 +0.070 +0.059 -0.017

+0.170 +0.020 +0.180 +0.060

+0.160 +0.090 +0.210 +0.120

Table 5: Stealing-based scrubbing on C4/Llama under four KGW settings. General comparators are the highest QSR@3 values among general scrubbers at each setting. KGW

Method

𝑛

ASR

QSR@3

Best general QSR

ΔQSR

Prom

ℎ = 1, 𝛾 = 0.25

B4 MIP DE-MARK SCTS JSV

100 100 100 100 100

0.78 0.17 0.61 0.22 0.75

0.72 0.10 0.42 0.20 0.42

DIPPER 0.59 DIPPER 0.59 DIPPER 0.59 DIPPER 0.59 DIPPER 0.59

+0.13 -0.49 -0.17 -0.39 -0.17

4.35 3.21 2.91 4.11 3.01

ℎ = 1, 𝛾 = 0.5

B4 MIP DE-MARK SCTS JSV B4 MIP DE-MARK SCTS JSV B4 MIP DE-MARK SCTS JSV

100 100 100 100 99 100 100 100 100 100 100 100 100 100 99

0.71 0.33 0.38 0.26 0.84 0.89 0.72 0.00 0.49 0.88 0.89 0.78 0.00 0.39 0.83

0.60 0.23 0.27 0.23 0.41 0.78 0.44 0.00 0.46 0.53 0.83 0.55 0.00 0.37 0.48

DIPPER 0.63 DIPPER 0.63 DIPPER 0.63 DIPPER 0.63 DIPPER 0.63 DIPPER 0.80 DIPPER 0.80 DIPPER 0.80 DIPPER 0.80 DIPPER 0.80 DIPPER 0.80 DIPPER 0.80 DIPPER 0.80 DIPPER 0.80 DIPPER 0.80

-0.03 -0.40 -0.36 -0.40 -0.22 -0.02 -0.36 -0.80 -0.34 -0.27 +0.03 -0.25 -0.80 -0.43 -0.32

4.32 3.04 2.97 4.38 2.89 4.22 3.16 2.79 4.46 3.05 4.31 3.09 2.64 4.48 3.13

ℎ = 3, 𝛾 = 0.25

ℎ = 3, 𝛾 = 0.5

B4 achieves the highest QSR@3 among the stealing-based scrubbers in all four settings, with rates from 0.60 to 0.83. Its QSR@3 exceeds the best general-scrubbing comparator by 0.13 and 0.03 at two settings and is lower by 0.03 and 0.02 at the other two. The ASR and QSR comparisons show different outcomes for stealing-based attacks. JSV often approaches the best generalscrubbing baseline in ASR, while its QSR@3 remains lower. SCTS shows the opposite failure mode: it preserves high Prometheus quality, yet its QSR remains below the strongest general-scrubbing baseline because removal is too weak. MIP and DE-MARK are more setting-sensitive and do not provide a consistent gain across the four KGW settings. Spoofing is the separate case where recovered information is used for false attribution rather than removal. We run a small KGW spoofing feasibility check for JSV-Sp and DE-MARK-Sp, reported in Table 7. QSSR@3 counts forged outputs that trigger the watermark detector and receive Prometheus ≥ 3 against their paired watermarked continuation. The eight spoofing runs each contain 100 outputs; natural-generation detector acceptance is zero in each run. DE-MARK-Sp reaches 98–100% detector acceptance and 47– 64% QSSR@3 across the four settings. JSV-Sp reaches 0–1% detector acceptance and zero QSSR@3.

Finding 5. B4 leads stealing-based QSR@3 in all four KGW settings; DE-MARK-Sp attains 98–100% detector acceptance and 47–64% QSSR@3.

6.6

Supplementary Output Screening

We supplement the automatic metrics with output screening against the task and source text. Deterministic rules identify empty and punctuation-only outputs; length reduction and repetition flag texts for Codex-assisted inspection. Across 26,096 outputs, we inspect 251 source–output pairs and identify 254 invalid outputs, including 19 deterministic failures. An invalid output contributes zero to screened QSR while remaining in its original denominator. This excludes 20 QSR@3 successes across 20 cells, and LLMP retains the highest dataset-level mean QSR@3 on C4, Dolly-15K, and BookReport. Appendix E.3 gives the screening criteria, review coverage, and per-method changes.

6.7

Scope And Limitations

The general-scrubbing matrix covers six watermark families, three instruction-tuned model families, and three prompt datasets. The Base–Instruct and recovered-information studies cover KGW under their specified prompting and access conditions. These

MarkSec : Capability-Aware Evaluation of Adversarial Attacks Against LLM Watermarks

configuration-level comparisons use the run-specific sources, detector thresholds, and resource scopes defined in Section 5.2 and Appendix D. Interpretation and Decision Value. A low joint-success rate can reflect few detectable inputs, low removal among eligible inputs, a strict automatic gate, or missing measurements. These cases lead to different security decisions and should not be collapsed into a watermark ranking. Evaluators should inspect baseline eligibility and paired transitions before choosing attack tests, and verify source identity before pairing methods for uncertainty estimation.

7

Related Work

Recent LLM watermark benchmarks and toolkits have improved the reproducibility of watermark implementation and evaluation. MarkLLM is a representative implementation-oriented toolkit, providing unified interfaces for implementing watermark methods and evaluating generation, detection, quality, and robustness behavior [24]. Complementing this toolkit perspective, WaterBench studies fair comparison of watermark methods under aligned watermarking strength, with particular attention to detection– generation trade-offs and instruction-following quality [32]. From a robustness perspective, WaterPark is the closest prior platform: it integrates a broad suite of removal attacks to evaluate the resilience of LLM watermarkers under adversarial perturbations [19]. Other work broadens the evaluation criteria. CEFW proposes a multi-dimensional scoring framework that covers detectability, text quality, usability, robustness, and imperceptibility [39]. MarkMyWords further benchmarks watermark schemes across quality, size, and tamper-resistance under practical attacks [27]. These frameworks are complementary to MarkSec, but they are mostly organized around watermark methods rather than attack methods. In these systems, attacks typically serve as robustness tests, quality–robustness trade-off factors, or components of a composite watermark score. WaterPark is the closest exception because it substantially expands the robustness evaluation. WaterPark also examines semantic preservation and watermark design factors. Its main evaluation target remains the resilience of watermarkers under attacks, not a capability-stratified benchmark of attack methods themselves. Prior frameworks also provide limited support for organizing stealing, spoofing, recovered artifacts, and general scrubbing as separate evaluation tracks. MarkSec addresses this attack-centric gap by asking how attack methods behave under explicit capability assumptions. It treats attack methods as a first-class benchmark dimension. It separates the general-scrubbing comparison from stronger stealing and spoofing analyses. It also reports attack effectiveness together with text quality. This makes MarkSec complementary to watermark implementation and watermark-scoring frameworks: it stress-tests LLM watermarks through modern attack methods and reports the assumptions behind those attacks.

8

Conclusion

We presented MarkSec, a general framework for evaluating adversarial attacks against LLM watermarks under a unified reporting

protocol. The framework is built around an explicit separation between comparable general scrubbing attacks and stronger stealingbased attacks that require additional access, recovered artifacts, or attack-side control. This design allows shared-protocol robustness comparisons while still exposing escalation risks that would be hidden by a text-only evaluation. Across the evaluated matrix, three findings stand out. First, there is no universal winner across all evaluation axes. Second, family-specific behavior matters: overall averages are useful, but they can hide whether an attack is a broad baseline or a narrow specialist. Third, SIRA leads mean removal on C4, while LLMP leads mean Prometheus-gated joint success across watermark families and retains that lead under stricter gates. Supplementary output screening preserves LLMP’s lead in dataset-level mean QSR@3. The Base–Instruct comparison shows smaller QSR gaps with the prompt wrapper disabled. Under stronger recovered-information access, B4 achieves the highest QSR@3 among the evaluated stealing-based scrubbers in all four KGW settings, and DE-MARKSp achieves high detector acceptance with lower quality-gated spoofing success. More broadly, our results argue for capability-aware watermark evaluation. Stronger attacks should not be ignored, but they should be reported together with the assumptions, query budgets, runtime costs, and artifact provenance that make them possible. We hope MarkSec helps move LLM watermark evaluation toward more reproducible, interpretable, and practically meaningful security assessment.

References [1] Sahar Abdelnabi, Kai Greshake, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz. 2023. Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. In AISec@CCS. 79–90. [2] Hongyan Chang, Hamed Hassani, and Reza Shokri. 2025. Watermark Smoothing Attacks against Language Models. In Findings of EMNLP. 4915–4941. [3] Ruibo Chen, Yihan Wu, Junfeng Guo, and Heng Huang. 2024. De-Mark: Watermark Removal in Large Language Models. arXiv preprint (2024). [4] Yixin Cheng, Hongcheng Guo, Yangming Li, and Leonid Sigal. 2025. Revealing Weaknesses in Text Watermarking Through Self-Information Rewrite Attacks. In ICML. 9982–10009. [5] Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin. 2023. Free Dolly: Introducing the World’s First Truly Open Instruction-Tuned LLM. https://www.databricks.com/blog/2023/04/12/dolly-first-open-commerci ally-viable-instruction-tuned-llm. Introduces the Databricks Dolly 15K dataset and Dolly 2.0. [6] Sumanth Dathathri, Abigail See, Shruti Ghaisas, et al. 2024. Scalable Watermarking for Identifying Large Language Model Outputs. Nature (2024), 818–823. [7] Ashleigh Davis. 2023. ChatGPT Sparks Cheating, Ethical Concerns as Students Try Realistic Essay Writing Technology. https://www.abc.net.au/news/202301-26/chatgpt-sparks-cheating-ethical-concerns-in-schools-universities/101 888440. ABC News, January 25, 2023; accessed 2026-04-26. [8] Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, et al. 2024. The Llama 3 Herd of Models. arXiv preprint arXiv:2407.21783 (2024). [9] Jocelyn Gecker. 2023. Amid ChatGPT Outcry, Some Teachers Are Inviting AI to Class. https://apnews.com/article/chatgpt-ai-use-school-essay-7bc171932ff9b9 94e04f6eaefc09319f. AP News, February 14, 2023; accessed 2026-04-26. [10] Zhengmian Hu, Lichang Chen, Xidong Wu, Yihan Wu, Hongyang Zhang, and Heng Huang. 2023. Unbiased Watermark for Large Language Models. arXiv preprint (2023). [11] Baizhou Huang, Xiao Pu, and Xiaojun Wan. 2025. B4: A Black-Box Scrubbing Attack on LLM Watermarks. In NAACL. 9113–9126. [12] Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, et al. 2023. Mistral 7B. arXiv preprint arXiv:2310.06825 (2023). [13] Juyong Jiang, Fan Wang, Jiasi Shen, Sungju Kim, and Sunghun Kim. 2026. A Survey on Large Language Models for Code Generation. ACM Transactions on Software Engineering and Methodology 35, 2 (Jan. 2026), 1–72. https://doi.org/

Li et al.

10.1145/3747588 [14] Nikola Jovanović, Robin Staab, and Martin Vechev. 2024. Watermark Stealing in Large Language Models. arXiv preprint (2024). [15] Seungone Kim, Juyoung Suk, Shayne Longpre, Bill Yuchen Lin, Jamin Shin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, and Minjoon Seo. 2024. Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models. arXiv preprint arXiv:2405.01535 (2024). [16] John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein. 2023. A Watermark for Large Language Models. In ICML. 17061– 17084. [17] Kalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting, and Mohit Iyyer. 2023. Paraphrasing Evades Detectors of AI-Generated Text, but Retrieval Is an Effective Defense. In NeurIPS. [18] Gregory Kang Ruey Lau, Xinyuan Niu, Hieu Dao, Jiangwei Chen, Chuan-Sheng Foo, and Bryan Kian Hsiang Low. 2024. Waterfall: Scalable Framework for Robust Text Watermarking and Provenance for LLMs. In EMNLP. 20432–20466. [19] Jiacheng Liang, Zian Wang, Spencer Hong, Shouling Ji, and Ting Wang. 2025. Watermark under Fire: A Robustness Evaluation of LLM Watermarking. In Findings of EMNLP. 21050–21074. [20] Weixin Liang, Yaohui Zhang, Mihai Codreanu, Jiayu Wang, Hancheng Cao, and James Zou. 2025. The Widespread Adoption of Large Language Model-Assisted Writing Across Society. arXiv:2502.09747 [cs.CL] https://arxiv.org/abs/2502.0 9747 [21] Aiwei Liu, Leyi Pan, Xuming Hu, Shu’ang Li, Lijie Wen, Irwin King, and Philip S Yu. 2023. An Unforgeable Publicly Verifiable Watermark for Large Language Models. arXiv preprint (2023). [22] Aiwei Liu, Leyi Pan, Xuming Hu, Shiao Meng, and Lijie Wen. 2023. A Semantic Invariant Robust Watermark for Large Language Models. arXiv preprint (2023). [23] Eric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning, and Chelsea Finn. 2023. DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature. arXiv:2301.11305 [cs.CL] https://arxiv.org/abs/23 01.11305 [24] Leyi Pan, Aiwei Liu, Zhiwei He, Zitian Gao, Xuandong Zhao, Yijian Lu, Binglin Zhou, Shuliang Liu, Xuming Hu, Lijie Wen, Irwin King, and Philip S. Yu. 2024. MarkLLM: An Open-Source Toolkit for LLM Watermarking. In EMNLP. 61–71. [25] Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In NeurIPS. 8024–8035. [26] Sida Peng, Eirini Kalliamvakou, Peter Cihon, and Mert Demirer. 2023. The Impact of AI on Developer Productivity: Evidence from GitHub Copilot. arXiv:2302.06590 [cs.SE] https://arxiv.org/abs/2302.06590 [27] Julien Piet, Chawin Sitawarin, Vivian Fang, Norman Mu, and David Wagner. 2023. Mark My Words: Analyzing and Evaluating Language Model Watermarks. arXiv preprint (2023). [28] James Purtill. 2024. Twitter Is Becoming a ’Ghost Town’ of Bots as AI-Generated Spam Content Floods the Internet. https://www.abc.net.au/news/science/202402-28/twitter-x-fighting-bot-problem-as-ai-spam-floods-the-internet/10349 8070. ABC News, February 27, 2024; accessed 2026-04-26. [29] Qwen Team, An Yang, Baosong Yang, Beichen Zhang, et al. 2024. Qwen2.5 Technical Report. arXiv preprint arXiv:2412.15115 (2024). [30] Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. Journal of Machine Learning Research (2020), 1–67. [31] Vinu Sankar Sadasivan, Aounon Kumar, Sriram Balasubramanian, Wenxiao Wang, and Soheil Feizi. 2023. Can AI-Generated Text Be Reliably Detected? Transactions on Machine Learning Research (2023). [32] Shangqing Tu, Yuliang Sun, Yushi Bai, Jifan Yu, Lei Hou, and Juanzi Li. 2024. WaterBench: Towards Holistic Evaluation of Watermarks for Large Language Models. In ACL. 1517–1542. [33] Shen Wang, Tianlong Xu, Hang Li, Chaoli Zhang, Joleen Liang, Jiliang Tang, Philip S. Yu, and Qingsong Wen. 2024. Large Language Models for Education: A Survey and Outlook. arXiv:2403.18105 [cs.CL] https://arxiv.org/abs/2403.18105 [34] Tiannan Wang, Jiamin Chen, Qingrui Jia, Shuai Wang, Ruoyu Fang, Huilin Wang, Zhaowei Gao, Chunzhao Xie, Chuou Xu, Jihong Dai, Yibin Liu, Jialong Wu, Shengwei Ding, Long Li, Zhiwei Huang, Xinle Deng, Teng Yu, Gangan Ma, Han Xiao, Zixin Chen, Danjun Xiang, Yunxia Wang, Yuanyuan Zhu, Yi Xiao, Jing Wang, Yiru Wang, Siran Ding, Jiayang Huang, Jiayi Xu, Yilihamu Tayier, Zhenyu Hu, Yuan Gao, Chengfeng Zheng, Yueshu Ye, Yihang Li, Lei Wan, Xinyue Jiang, Yujie Wang, Siyu Cheng, Zhule Song, Xiangru Tang, Xiaohua Xu, Ningyu Zhang, Huajun Chen, Yuchen Eleanor Jiang, and Wangchunshu Zhou. 2024. Weaver: Foundation Models for Creative Writing. arXiv:2401.17268 [cs.CL] https://arxiv.org/abs/2401.17268 [35] Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa

Kasirzadeh, Courtney Biles, Sasha Brown, Zac Kenton, Will Hawkins, Tom Stepleton, Abeba Birhane, Lisa Anne Hendricks, Laura Rimell, William Isaac, Julia Haas, Sean Legassick, Geoffrey Irving, and Iason Gabriel. 2022. Taxonomy of Risks Posed by Language Models. In FAccT. 214–229. [36] Qilong Wu and Varun Chandrasekaran. 2024. Bypassing LLM Watermarks with Color-Aware Substitutions. In ACL. 8549–8581. [37] Hanyi Xu, Wensheng Gan, Zhenlian Qi, Jiayang Wu, and Philip S. Yu. 2024. Large Language Models for Education: A Survey. arXiv:2405.13001 [cs.CL] https://arxiv.org/abs/2405.13001 [38] Hanlin Zhang, Benjamin L. Edelman, Danilo Francati, Daniele Venturi, Giuseppe Ateniese, and Boaz Barak. 2024. Watermarks in the Sand: Impossibility of Strong Watermarking for Language Models. In ICML. 58851–58880. [39] Shuhao Zhang, Bo Cheng, Jiale Han, Yuli Chen, Zhixuan Wu, Changbao Li, and Pingli Gu. 2025. CEFW: A Comprehensive Evaluation Framework for Watermark in Large Language Models. In ICME. 1–6. [40] Zhaoxi Zhang, Xiaomei Zhang, Yanjun Zhang, Leo Yu Zhang, Chao Chen, Shengshan Hu, Asif Gill, and Shirui Pan. 2024. Stealing Watermarks of Large Language Models via Mixed Integer Programming. In ACSAC.

A Open Science Reproducibility Scope. The reproducibility materials distinguish the manuscript source from the benchmark code, experimental records, and third-party assets. We view reproducibility as a core requirement for benchmark papers. The main contribution of MarkSec is a reusable evaluation protocol that others can inspect, rerun, and extend. To support that goal, the project is organized around explicit configuration files, versioned benchmark defaults, and shared interfaces for watermark generation, attack execution, detection, quality evaluation, and result persistence. Artifact Scope. The benchmark repository contains the framework code needed to instantiate the reported evaluation pipeline, including attack adapters, watermark wrappers, detector calibration logic, experiment orchestration, quality-analysis modules, and paper-facing plotting/reporting scripts. The configuration-driven design makes benchmark assumptions auditable: attack settings, watermark parameters, datasets, victim models, and reporting thresholds are defined in named configs rather than buried in ad hoc scripts. Reproducibility Assets. The evidence manifest identifies TeX and figure dependencies, per-sample records, configurations, and analysis outputs by cryptographic hashes. The experiment records identify the April 2026 runs, September supplementary scoring, and output-screening decisions as separate measurements. The source archive accompanying the PDF is a manuscript archive, not a claim that all third-party models, datasets, or recovered artifacts are redistributable. This is especially important for stealingbased attacks, where recovered artifacts and corpus provenance materially affect interpretation. Dependencies And Non-Redistributable Components. Some components of the benchmark depend on external model checkpoints, third-party repositories, or datasets that may have their own licenses or access controls. Redistribution of these components is governed by their respective licenses and access conditions; users may need to obtain third-party assets separately under their original terms. The framework is designed to resolve local model paths through repo configuration so that experiments can be rerun on machines with different storage layouts without changing benchmark logic.

MarkSec : Capability-Aware Evaluation of Adversarial Attacks Against LLM Watermarks

Planned Benchmark Release. We intend to release the framework source, configuration files, experiment manifests, and figuregeneration scripts needed to reproduce the reported benchmark protocol and extend it with new attacks, watermark families, datasets, and models. This planned release does not include thirdparty model checkpoints, datasets, or recovered artifacts unless their redistribution terms permit it.

B

Ethical Considerations

This paper studies attacks on LLM watermarking systems, which creates an unavoidable dual-use tension. Stronger attack benchmarks are necessary for understanding whether watermarking methods provide meaningful robustness in practice. The same evaluation infrastructure can also lower the barrier to studying or operationalizing watermark removal, stealing, or spoofing strategies. We therefore frame MarkSec as a stress-testing benchmark for defensive evaluation, not as a claim that watermark removal should be optimized without qualification. Why This Evaluation Is Necessary. Watermarks are often discussed as mechanisms for provenance, platform governance, and misuse mitigation. If their robustness is evaluated only against weak or incomparable baselines, practitioners may overestimate the security they provide. A realistic attack benchmark is therefore important for preventing false confidence, identifying failure modes before deployment, and making defender-side trade-offs visible. Dual-Use Risk And Mitigation. The highest-risk part of this space is capability-escalated attacks that recover reusable watermark artifacts or support downstream spoofing. For that reason, the paper does not collapse all attacks into a single leaderboard. It reports stronger stealing-based threats separately and emphasizes the assumptions that make them possible, including access patterns, query budgets, runtime costs, and watermark-family alignment. This reporting choice is itself a mitigation: it keeps the analysis focused on security interpretation rather than on decontextualized “best attack” claims. Release And Use Considerations. Any public release of the benchmark should prioritize reproducibility for defenders and evaluators while avoiding unnecessary concentration of turnkey high-risk assets. In practice, this means documenting assumptions, preserving artifact provenance, and distinguishing benchmark infrastructure from third-party models or recovered attack-side artifacts that may require additional controls or independent reconstruction. Users of the benchmark should interpret the reported attacks as stress tests for robustness claims and should avoid deploying them against systems they do not own or have permission to evaluate. Broader Impact. Our goal is to improve the reliability of claims about LLM watermark robustness. We believe transparent evaluation with explicit capability assumptions is more responsible than relying on optimistic or incomparable robustness reports, especially for techniques that may inform provenance or policy decisions. At the same time, benchmark results should not be read as evidence that watermarking is either universally sufficient or universally futile. The more responsible conclusion is narrower:

robustness depends on attacker capability, watermark family, deployment format, and the extent to which attacked outputs remain useful after manipulation.

C Generative AI Usage Generative AI assistance was used for manuscript revision, code inspection and changes, analysis-script development, experimentartifact analysis, and source–output content inspection.

D Implementation Details This appendix summarizes the concrete implementation choices behind MarkSec.

D.1 Default Benchmark Configuration Benchmark Matrix. Unless otherwise noted, the main removal benchmark uses three instruction-tuned victim families: Llama 3.1 8B Instruct, Qwen2.5 7B Instruct, and Mistral 7B Instruct v0.3. We evaluate on C4, Dolly-15K, and MMW BookReport, and randomly sample 100 prompts per dataset. For C4, we prepend an instruction-style continuation prefix so that continuation-style data better matches the chat-oriented victims used in the benchmark. Watermark Families. The reported benchmark covers six watermark families: KGW, Unigram, SIR, UW, DIP, and SynthID. For KGW stealing-based analysis, we additionally expose an explicit four-setting grid through repo configs with ℎ ∈ {1, 3}, 𝛾 ∈ {0.25, 0.5}, and 𝛿 = 2.0. Other watermark families are run through the unified framework interface and compared under the same downstream reporting protocol. Detection And Quality Reporting. Detector calibration targets an FPR of 1%, with decisions taken at the recorded run-specific thresholds. We then report TPRbefore , TPRafter , and ASR. For utilityaware reporting, we use Prometheus as the default gating metric and report 𝑄𝑆𝑅@2, 𝑄𝑆𝑅@3, and 𝑄𝑆𝑅@4, with 𝑄𝑆𝑅 in the main text referring to 𝑄𝑆𝑅@3 unless otherwise noted. The paper-facing quality matrices also report BERTScore and PSP as supporting semantic diagnostics. Sample Universe. Each attack block uses its completed-output population and generated source texts. TPR𝑏 uses available beforedetection results; ASR uses before-positive samples with both detector states, and QSR uses all completed outputs (Section 5.2). The result data record the numerator and denominator of each rate and the count of missing predictions. Attack Defaults. The attack defaults below are repo settings used to instantiate benchmark runs. For the general-scrubbing benchmark, DIPPER uses lexical diversity 60, order diversity 0, sentence interval 1, and a bf16 paraphraser backend; LLMP uses the rewrite prompt template with max new tokens 256, temperature 0.7, and sampling enabled; SIRA uses self-information threshold 30, max new tokens 256, and deterministic decoding, following the official implementation’s percentile-style blanking semantics; RW uses 200 total steps, 40 target valid steps, span length 6, backtrack patience 40, reward-model tie threshold 0.02, and T5 sampling with top-𝑝 = 0.8, top-𝑘 = 40, and temperature 1.0; and SA uses 𝛼 = 1.0, top-𝑘 = 10, 200 confidence-normalization samples, 100 confidence

Li et al.

bins, and max new tokens 200. For protocol-separated stealingbased analysis, MIP uses Gurobi with a 1800s time limit, MIP gap 0.05, AS2 attack mode, 2000 stealing samples, max edit ratio 0.3, and the greedy removal backend by default. SCTS uses context size 1, at most five candidates, query top-𝑘 = 10, step size 2, max substitution ratio 0.1, and chi-square threshold 0.01. DE-MARK is run against a KGW-style target with context width 1, query batch size 32, sampled token budget 20, predicted/actual 𝛿 = 2.0, and query caching enabled. JSV uses a query budget of 30000, batch size 10, previous-context width 3, prebuilt corpus reuse, and downstream scrub/spoof deltas of ±7.5. B4 runs in prebuilt-corpus mode with automatic corpus discovery, OPT-IML-1.3B as the default proxyreference model, contrastive top-𝑘 = 10, max new tokens 256, and fixed 𝜆 = 0.5. Capability And Cost Transparency. Table 6 reports attack execution time summed from per-sample durations. Preparation, corpus acquisition, and offline recovery are excluded as specified for each method. Cost Drivers. MIP’s short online editing time does not include its offline recovery stage. DE-MARK’s spoof rows record about 20,000 token queries, while JSV’s 30,000 figure describes a prebuilt corpus; these quantities are not interchangeable.

D.2 Implementation Notes Unified Framework Integration. We integrate watermarking methods through the MarkLLM-based interface layer and adapt attack implementations to a shared pipeline for data loading, generation, attack execution, detection, quality analysis, and result persistence. Direct Removal Versus Stealing-Based Attacks. We intentionally separate general scrubbers from stealing-based attacks in the paper’s reporting. General scrubbers take watermarked text and optional side information as input. Stealing-based attacks first recover a reusable artifact such as a greenlist, token-color map, count table, or proxy distribution. We therefore do not collapse them into a single headline leaderboard. Local Models And Fixed Checkpoints. All experiments use local checkpoints rather than floating remote model identifiers. This avoids silent version drift during a long-running benchmark campaign and keeps attack-side helper models reproducible under the same local environment. Prebuilt Stealing Corpora. For stealing-based attacks that depend on large collections of watermarked and natural continuations, the benchmark defaults to canonical prebuilt corpora. These corpora are generated with vLLM V0 rather than V1 because, at the time of writing, the V1 path did not support the custom logitsprocessor interface required by our watermark injection stack. Attack Adaptations. SA uses the unified edit pipeline, RW includes prompt-aware cleanup and local quality control, and SIRA’s proxy and rewriting roles share the framework interface. These adaptations connect the attack implementations to the same generation, detection, and reporting components.

E Additional Results E.1 Family-Level Aggregation for RQ2 RQ2 in the main text summarizes family-conditioned QSR patterns with compact figures. The detailed evidence is the set of appendix tables below. Table 7 reports the limited KGW spoofing feasibility slice used in RQ5. Table 8 and Table 9 report the canonical C4 detection/QSR and attacked-text quality matrices. Table 10 and Table 11 provide the same matrices for Dolly-15K. Table 12 and Table 13 provide the same matrices for MMW BookReport. These tables provide the cell-level values behind the main RQ figures.

E.2

Full General-Scrubbing Matrices

The main text reports compact C4 views for readability and protocol control. This appendix reports the full watermark–victimmodel–attack matrices for the canonical C4 slice and two externalvalidity slices. Detection tables report pre-attack TPR, raw ASR, and QSR@3. The quality tables score attacked text with BERTScore, PSP, and Prometheus. The quality tables report the original runlevel automatic measurements. Original per-record BERTScore and PSP measurements are available for 36 cells. For the remaining 225 cells, supplementary scoring in September 2026 evaluates the same 22,497 text pairs; all 450 cell-mean differences from the original scores are below 0.001. The supplementary scores use full-text fields, including a shared prompt when present, and are reported separately in the result data. Appendix E.3 gives the output-screened QSR analysis. Thirteen Dolly–SIR cells have no observable removal opportunities; their ASR is marked “–” and excluded from ASR averages. Unless otherwise noted, QSR uses Prometheus threshold 3. “N/A” denotes combinations excluded by the benchmark protocol. Bold in detection tables marks the rowwise best QSR@3. Bold in quality tables marks the row-wise best Prometheus score.

E.3

Content Validity of Attack Outputs

We screen all 26,096 outputs across the 261 general-scrubbing cells and five attack implementations, with individual source–output inspection of 251 pairs. Empty outputs and outputs containing only punctuation are invalid. Repetition, unrelated responses, and loss of core content require inspection against the source and task; length alone does not establish failure. An invalid output contributes zero to quality-gated success at every gate and remains in the original denominator. The separate screening gate leaves the measured similarity scores and detector states unchanged. Deterministic checks identify 2 empty and 17 punctuation-only outputs. Source–output inspection assisted by Codex identifies 235 further invalid observations, giving 254 in total. Of the 251 inspected pairs, 16 have no established categorical failure; 25,826 other outputs receive only the mechanical screen. Short verse retaining its core imagery is not rejected for length alone. Some sources are themselves degraded, so output invalidity does not by itself establish attack-caused degradation.

MarkSec : Capability-Aware Evaluation of Adversarial Attacks Against LLM Watermarks

Table 6: Attack execution time across four KGW settings per method. Ranges sum per-sample durations within each run; excluded preparation stages are listed separately. Method

Objective

B4 DE-MARK JSV MIP SCTS DE-MARK-Sp JSV-Sp

Scrub Scrub Scrub Scrub Scrub Spoof Spoof

Runs

Execution time (min)

4 4 4 4 4 4 4

16.6–18.9 666.0–690.2 6.0–13.6 0.09–0.15 2384.7–2731.0 673.7–701.6 9.3–32.3

Excluded stages Preparation and corpus acquisition Preparation and prior cache construction Preparation and corpus acquisition Offline recovery and corpus acquisition Preparation outside per-sample calls Preparation and prior cache construction Preparation and corpus acquisition

Each run records 100 outputs except two JSV scrub runs with 99. Scoring performed after attack execution is excluded.

Let 𝐼𝑖 indicate a confirmed-invalid output, 𝑅𝑖 an observed detector removal, and 𝑃𝑖 the Prometheus score. The additional descriptive statistic is 1 ∑ 1[𝑅𝑖 ∧ 𝑃𝑖 ≥ 𝜏 ∧ 𝐼𝑖 = 0]. = QSRscreen 𝜏 |𝐶| 𝑖∈𝐶 Here 𝐼𝑖 = 0 means that invalidity has not been established, rather than a positive content-validity judgment. This check removes 20 successes across 20 cells at gate 3, 32 at gate 2, and 5 at gate 4. The table summarizes all dataset–attack groups whose gate-3 values change. The denominator and treatment of missing detector states follow Section 5.2. Dataset

Attack

Removed

QSR@3

Screened

BookReport BookReport C4 C4 Dolly Dolly

DIPPER RW DIPPER RW DIPPER RW

3 2 7 2 1 5

30.22 17.33 58.82 36.06 37.72 23.17

30.06 17.22 58.43 35.94 37.67 22.89

Removed counts observations; the last two columns are cell-macro QSR@3 percentages, each over 18 cells. Unlisted groups retain their gate-3 values. LLMP remains the leader of the dataset-level cell macro-averages under this check.

E.4

Spoofing Feasibility

Table 7 reports the limited KGW spoofing feasibility slice used in RQ5. Natural AR is the detector accept rate on ordinary unwatermarked generations. QSSR@3 counts spoofing success only when the forged text is detected as watermarked and reaches Prometheus ≥ 3 against the paired watermarked output. All eight rows have 100 outputs and complete spoof-specific Prometheus scores. The resource columns distinguish token probes for DEMARK from prebuilt-corpus size for JSV; execution time sums per-sample attack durations.

Li et al.

Table 7: Spoofing feasibility under KGW settings on C4 with Llama-3.1-8B-Instruct. Genuine TPR↑ Natural AR↓

KGW

ℎ=1, 𝛾 =0.25 ℎ=1, 𝛾 =0.25 ℎ=1, 𝛾 =0.5 ℎ=1, 𝛾 =0.5 ℎ=3, 𝛾 =0.25 ℎ=3, 𝛾 =0.25 ℎ=3, 𝛾 =0.5 ℎ=3, 𝛾 =0.5

0.80 0.74 0.56 0.64 0.71 0.75 0.59 0.67

0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00

Method DE-MARK-Sp JSV-Sp DE-MARK-Sp JSV-Sp DE-MARK-Sp JSV-Sp DE-MARK-Sp JSV-Sp

SSR↑ QSSR@3↑ Prom↑ Probes / corpus

Exec. time

1.00 0.00 1.00 0.00 1.00 0.01 0.98 0.01

701.6m 9.3m 688.2m 32.1m 674.0m 32.3m 673.7m 9.4m

0.47 0.00 0.64 0.00 0.55 0.00 0.56 0.00

2.51 2.02 2.91 1.95 2.80 2.05 2.98 1.85

20k token 30k corpus 20k token 30k corpus 20k token 30k corpus 20k token 30k corpus

Table 8: Full C4 general-scrubbing detection and QSR matrix. DIPPER

Watermark Model

LLMP

SIRA

RW

SA

TPR𝑏 ↑ ASR↑ QSR@3↑ TPR𝑏 ↑ ASR↑ QSR@3↑ TPR𝑏 ↑ ASR↑ QSR@3↑ TPR𝑏 ↑ ASR↑ QSR@3↑ TPR𝑏 ↑ ASR↑ QSR@3↑ KGW

Llama Qwen Mistral

0.93 0.96 0.94

0.78 0.75 0.88

0.63 0.61 0.71

0.92 1.00 0.97

0.67 0.88 0.82

0.57 0.72 0.78

0.93 0.96 0.94

0.82 0.91 0.89

0.42 0.74 0.44

0.96 0.99 0.95

0.61 0.62 0.81

0.35 0.26 0.39

1.00 0.95 N/A

0.46 0.68 N/A

0.26 0.33 N/A

Unigram

Llama Qwen Mistral

0.72 0.81 0.90

0.75 0.85 0.60

0.46 0.57 0.50

0.77 0.86 0.81

0.75 0.98 0.78

0.53 0.78 0.58

0.72 0.81 0.90

0.96 0.95 0.88

0.43 0.65 0.47

0.73 0.80 0.91

0.88 0.64 0.59

0.30 0.24 0.31

0.73 0.75 N/A

0.49 0.89 N/A

0.19 0.22 N/A

SIR

Llama Qwen Mistral

0.62 0.69 0.91

0.62 0.79 0.55

0.33 0.41 0.43

0.55 0.70 0.91

0.78 0.96 0.69

0.40 0.63 0.61

0.51 0.77 0.91

0.84 0.95 0.73

0.27 0.62 0.38

0.49 0.71 0.93

0.71 0.76 0.80

0.23 0.32 0.32

0.59 0.70 N/A

0.76 0.91 N/A

0.26 0.24 N/A

UW

Llama Qwen Mistral

0.90 0.84 0.95

0.93 0.99 0.96

0.71 0.68 0.84

0.85 0.83 0.94

0.92 1.00 0.93

0.75 0.76 0.82

0.90 0.84 0.95

0.96 1.00 0.96

0.47 0.71 0.43

0.87 0.75 0.98

0.78 0.93 0.84

0.48 0.37 0.50

0.90 0.82 0.95

0.73 1.00 0.99

0.24 0.26 0.46

DIP

Llama Qwen Mistral

0.57 0.82 0.85

1.00 0.94 0.94

0.49 0.59 0.68

0.72 0.81 0.84

0.90 0.96 0.95

0.61 0.72 0.76

0.57 0.82 0.85

0.93 0.96 0.92

0.37 0.69 0.35

0.70 0.78 0.85

0.89 0.88 0.85

0.40 0.34 0.37

0.75 0.86 0.78

0.97 0.99 0.99

0.19 0.32 0.46

SynthID

Llama Qwen Mistral

0.95 0.98 0.97

0.80 0.85 0.69

0.67 0.70 0.58

0.99 0.93 0.99

0.89 0.97 0.81

0.85 0.85 0.76

0.95 0.98 0.97

0.87 0.98 0.74

0.43 0.87 0.40

0.96 0.94 0.96

0.86 0.88 0.82

0.47 0.42 0.42

0.96 0.95 0.97

0.95 1.00 0.91

0.22 0.27 0.38

Table 9: Full C4 general-scrubbing attacked-text quality matrix. DIPPER

Watermark Model

LLMP

SIRA

RW

SA

BERT↑

PSP↑

Prom↑

BERT↑

PSP↑

Prom↑

BERT↑

PSP↑

Prom↑

BERT↑

PSP↑

Prom↑

BERT↑

PSP↑

Prom↑

KGW

Llama Qwen Mistral

0.929 0.894 0.921

0.835 0.807 0.821

3.70 3.55 3.57

0.932 0.918 0.938

0.844 0.838 0.873

4.18 3.67 4.45

0.910 0.901 0.897

0.781 0.819 0.749

3.09 3.93 2.79

0.917 0.904 0.899

0.819 0.834 0.797

3.06 2.78 2.70

0.877 0.850 N/A

0.655 0.614 N/A

2.61 2.56 N/A

Unigram

Llama Qwen Mistral

0.921 0.911 0.932

0.814 0.837 0.854

3.57 3.63 3.89

0.941 0.922 0.934

0.879 0.863 0.863

4.26 4.19 4.42

0.900 0.911 0.906

0.764 0.836 0.775

2.83 4.09 3.02

0.907 0.908 0.909

0.811 0.824 0.811

2.88 2.71 2.92

0.877 0.855 N/A

0.643 0.604 N/A

2.65 2.18 N/A

SIR

Llama Qwen Mistral

0.929 0.900 0.924

0.822 0.822 0.833

3.58 3.46 3.67

0.936 0.918 0.935

0.846 0.858 0.870

4.11 4.47 4.35

0.910 0.905 0.897

0.765 0.832 0.768

3.16 4.17 2.92

0.915 0.908 0.902

0.799 0.849 0.803

3.06 3.03 2.68

0.880 0.851 N/A

0.643 0.593 N/A

2.65 2.45 N/A

UW

Llama Qwen Mistral

0.928 0.909 0.928

0.831 0.836 0.844

3.68 3.61 3.90

0.941 0.923 0.937

0.871 0.872 0.875

4.53 4.17 4.13

0.904 0.909 0.901

0.770 0.839 0.765

2.80 4.01 2.72

0.915 0.906 0.906

0.834 0.836 0.815

3.17 2.90 2.86

0.866 0.863 0.864

0.754 0.789 0.761

2.05 2.09 2.56

DIP

Llama Qwen Mistral

0.928 0.905 0.925

0.832 0.833 0.831

3.54 3.45 3.59

0.940 0.920 0.940

0.872 0.868 0.881

4.24 4.27 4.41

0.906 0.910 0.898

0.771 0.840 0.749

3.21 4.18 2.58

0.919 0.909 0.908

0.834 0.849 0.813

3.07 2.78 2.89

0.864 0.865 0.862

0.719 0.787 0.770

1.83 2.22 2.75

SynthID

Llama Qwen Mistral

0.929 0.903 0.929

0.838 0.826 0.850

3.63 3.69 3.85

0.940 0.917 0.939

0.872 0.862 0.872

4.38 4.25 4.27

0.904 0.905 0.896

0.771 0.833 0.751

2.72 4.05 2.80

0.914 0.906 0.908

0.812 0.840 0.819

2.94 2.71 2.89

0.867 0.862 0.862

0.735 0.791 0.770

1.82 1.97 2.29

MarkSec : Capability-Aware Evaluation of Adversarial Attacks Against LLM Watermarks

Table 10: Full Dolly-15K general-scrubbing detection and QSR matrix. DIPPER

Watermark Model

LLMP

SIRA

RW

SA

TPR𝑏 ↑ ASR↑ QSR@3↑ TPR𝑏 ↑ ASR↑ QSR@3↑ TPR𝑏 ↑ ASR↑ QSR@3↑ TPR𝑏 ↑ ASR↑ QSR@3↑ TPR𝑏 ↑ ASR↑ QSR@3↑ KGW

Llama Qwen Mistral

0.77 0.85 0.85

0.87 0.88 0.89

0.44 0.58 0.56

0.81 0.90 0.87

0.83 0.99 0.86

0.63 0.83 0.70

0.77 0.85 0.85

0.83 0.98 0.88

0.45 0.67 0.48

0.82 0.85 0.80

0.67 0.68 0.82

0.27 0.30 0.28

0.77 0.87 N/A

0.58 0.86 N/A

0.21 0.38 N/A

Unigram

Llama Qwen Mistral

0.39 0.92 0.56

0.82 0.50 0.77

0.25 0.31 0.31

0.42 0.93 0.48

0.90 0.83 0.88

0.35 0.73 0.41

0.39 0.92 0.56

0.95 0.82 0.82

0.18 0.61 0.30

0.42 0.91 0.51

0.83 0.35 0.61

0.19 0.19 0.13

0.36 0.93 N/A

0.56 0.54 N/A

0.12 0.31 N/A

SIR

Llama Qwen Mistral

0.00 0.00 0.00

– – –

0.00 0.00 0.00

0.00 0.00 0.00

– – –

0.00 0.00 0.00

0.00 0.00 0.00

– – –

0.00 0.00 0.00

0.01 0.00 0.00

1.00 – –

0.00 0.00 0.00

0.00 0.00 N/A

– – N/A

0.00 0.00 N/A

UW

Llama Qwen Mistral

0.71 0.90 0.59

1.00 1.00 1.00

0.57 0.74 0.47

0.70 0.89 0.70

1.00 1.00 1.00

0.67 0.82 0.65

0.71 0.90 0.59

0.97 1.00 1.00

0.49 0.75 0.41

0.67 0.94 0.70

0.97 0.84 0.91

0.40 0.48 0.30

0.72 0.91 0.64

1.00 1.00 1.00

0.23 0.55 0.38

DIP

Llama Qwen Mistral

0.50 0.67 0.85

0.98 0.96 0.93

0.36 0.44 0.59

0.48 0.60 0.78

1.00 1.00 0.94

0.47 0.58 0.64

0.50 0.67 0.72

0.80 0.97 0.92

0.21 0.47 0.37

0.51 0.63 0.71

0.86 0.87 0.90

0.29 0.26 0.34

0.47 0.65 0.78

0.91 1.00 0.96

0.10 0.37 0.48

SynthID

Llama Qwen Mistral

0.49 0.61 0.51

0.96 0.93 0.98

0.39 0.40 0.38

0.46 0.67 0.47

0.98 1.00 0.98

0.44 0.60 0.44

0.49 0.61 0.51

0.94 1.00 0.98

0.28 0.48 0.35

0.51 0.63 0.42

0.86 0.87 0.88

0.22 0.37 0.15

0.50 0.58 0.53

1.00 1.00 1.00

0.17 0.28 0.35

Table 11: Full Dolly-15K general-scrubbing attacked-text quality matrix. DIPPER

Watermark Model

LLMP

SIRA

RW

SA

BERT↑

PSP↑

Prom↑

BERT↑

PSP↑

Prom↑

BERT↑

PSP↑

Prom↑

BERT↑

PSP↑

Prom↑

BERT↑

PSP↑

Prom↑

KGW

Llama Qwen Mistral

0.910 0.905 0.914

0.818 0.801 0.826

3.10 3.18 3.19

0.939 0.928 0.932

0.882 0.856 0.861

4.47 4.32 4.43

0.897 0.894 0.895

0.758 0.778 0.755

3.49 3.80 3.32

0.915 0.910 0.894

0.850 0.854 0.772

3.16 2.99 2.74

0.872 0.865 N/A

0.677 0.651 N/A

2.68 2.54 N/A

Unigram

Llama Qwen Mistral

0.910 0.905 0.919

0.810 0.810 0.840

3.24 3.25 3.39

0.938 0.926 0.928

0.869 0.850 0.856

4.41 4.40 4.27

0.893 0.899 0.897

0.756 0.786 0.755

3.15 3.71 3.12

0.921 0.911 0.901

0.868 0.844 0.792

3.30 3.09 2.87

0.871 0.863 N/A

0.683 0.652 N/A

2.99 2.70 N/A

SIR

Llama Qwen Mistral

0.910 0.898 0.915

0.813 0.783 0.832

3.10 3.15 3.21

0.938 0.924 0.929

0.875 0.853 0.857

4.48 4.44 4.41

0.898 0.897 0.890

0.766 0.782 0.734

3.34 3.76 3.36

0.912 0.919 0.902

0.856 0.870 0.802

2.82 3.37 2.81

0.870 0.862 N/A

0.666 0.641 N/A

2.99 2.83 N/A

UW

Llama Qwen Mistral

0.916 0.909 0.919

0.836 0.817 0.844

3.22 3.35 3.39

0.941 0.925 0.931

0.890 0.851 0.860

4.49 4.42 4.33

0.904 0.895 0.894

0.790 0.771 0.751

3.23 3.73 3.39

0.920 0.916 0.906

0.868 0.867 0.820

2.96 3.07 2.68

0.857 0.889 0.884

0.756 0.845 0.833

1.86 2.86 2.99

DIP

Llama Qwen Mistral

0.914 0.903 0.915

0.833 0.784 0.825

3.19 3.19 3.13

0.936 0.928 0.931

0.877 0.856 0.862

4.44 4.38 4.12

0.903 0.897 0.890

0.787 0.772 0.737

3.29 3.80 3.03

0.924 0.906 0.905

0.875 0.831 0.811

3.07 2.98 2.91

0.858 0.889 0.881

0.759 0.848 0.828

1.80 2.63 2.99

SynthID

Llama Qwen Mistral

0.911 0.902 0.920

0.808 0.779 0.837

3.41 3.11 3.33

0.938 0.927 0.931

0.878 0.848 0.859

4.36 4.44 4.34

0.901 0.898 0.896

0.781 0.781 0.753

3.33 3.79 3.25

0.917 0.913 0.905

0.859 0.840 0.808

3.00 3.16 2.95

0.858 0.884 0.882

0.765 0.837 0.827

2.03 2.63 2.90

Li et al.

Table 12: Full MMW BookReport general-scrubbing detection and QSR matrix. DIPPER

Watermark Model

LLMP

SIRA

RW

SA

TPR𝑏 ↑ ASR↑ QSR@3↑ TPR𝑏 ↑ ASR↑ QSR@3↑ TPR𝑏 ↑ ASR↑ QSR@3↑ TPR𝑏 ↑ ASR↑ QSR@3↑ TPR𝑏 ↑ ASR↑ QSR@3↑ KGW

Llama Qwen Mistral

0.90 0.99 0.90

0.97 0.89 0.99

0.40 0.24 0.36

0.93 0.98 0.89

0.69 0.97 0.83

0.55 0.56 0.66

0.90 0.99 0.90

0.90 0.94 0.98

0.38 0.68 0.34

0.93 0.97 0.95

0.76 0.74 0.89

0.30 0.13 0.10

0.94 1.00 N/A

0.59 0.60 N/A

0.24 0.18 N/A

Unigram

Llama Qwen Mistral

0.74 0.98 0.76

0.69 0.69 0.88

0.27 0.28 0.30

0.71 0.98 0.57

0.65 0.81 0.89

0.44 0.69 0.45

0.74 0.98 0.76

0.96 0.72 0.92

0.33 0.53 0.23

0.67 0.93 0.71

0.85 0.59 0.55

0.19 0.16 0.04

0.76 0.96 N/A

0.37 0.48 N/A

0.13 0.22 N/A

SIR

Llama Qwen Mistral

0.68 0.90 0.77

0.71 0.72 0.66

0.25 0.26 0.17

0.65 0.88 0.82

0.77 0.98 0.68

0.45 0.52 0.51

0.68 0.90 0.77

0.87 0.87 0.94

0.38 0.53 0.30

0.54 0.67 0.84

0.74 0.76 0.70

0.12 0.17 0.08

0.59 0.59 N/A

0.80 0.90 N/A

0.16 0.22 N/A

UW

Llama Qwen Mistral

0.99 1.00 1.00

0.98 1.00 0.95

0.40 0.34 0.36

1.00 1.00 1.00

0.91 1.00 0.97

0.83 0.86 0.85

0.99 1.00 1.00

0.96 1.00 0.94

0.58 0.81 0.34

0.99 1.00 1.00

0.72 0.83 0.88

0.28 0.39 0.09

0.98 1.00 1.00

0.32 0.90 0.80

0.07 0.35 0.45

DIP

Llama Qwen Mistral

0.50 0.82 0.68

0.98 0.99 1.00

0.27 0.45 0.26

0.44 0.89 0.83

1.00 0.93 0.95

0.39 0.66 0.71

0.50 0.82 0.69

0.96 0.99 1.00

0.27 0.75 0.24

0.34 0.87 0.70

1.00 0.87 0.96

0.08 0.32 0.08

0.37 0.89 0.80

1.00 0.99 1.00

0.03 0.46 0.27

SynthID

Llama Qwen Mistral

0.54 0.98 0.89

0.93 0.83 0.84

0.20 0.31 0.32

0.82 0.99 0.85

0.94 0.94 0.92

0.71 0.74 0.70

0.54 0.98 0.89

0.93 0.92 0.91

0.26 0.71 0.36

0.75 0.98 0.87

0.93 0.78 0.90

0.25 0.26 0.08

0.77 1.00 0.91

0.97 0.87 0.90

0.15 0.40 0.35

Table 13: Full MMW BookReport general-scrubbing attacked-text quality matrix. DIPPER

Watermark Model

LLMP

SIRA

RW

SA

BERT↑

PSP↑

Prom↑

BERT↑

PSP↑

Prom↑

BERT↑

PSP↑

Prom↑

BERT↑

PSP↑

Prom↑

BERT↑

PSP↑

Prom↑

KGW

Llama Qwen Mistral

0.900 0.888 0.898

0.814 0.763 0.786

2.26 1.79 2.27

0.945 0.928 0.932

0.913 0.878 0.885

4.05 3.19 4.15

0.874 0.885 0.866

0.681 0.754 0.650

2.56 3.33 2.28

0.899 0.893 0.856

0.796 0.784 0.636

2.26 1.71 1.35

0.881 0.872 N/A

0.741 0.724 N/A

2.43 1.94 N/A

Unigram

Llama Qwen Mistral

0.892 0.886 0.898

0.787 0.750 0.795

2.39 2.16 2.21

0.944 0.929 0.936

0.903 0.876 0.892

4.22 3.94 4.17

0.869 0.882 0.862

0.660 0.743 0.633

2.27 3.51 2.07

0.894 0.903 0.873

0.777 0.822 0.702

2.14 2.10 1.59

0.880 0.871 N/A

0.753 0.723 N/A

2.37 2.26 N/A

SIR

Llama Qwen Mistral

0.900 0.892 0.897

0.814 0.792 0.785

2.40 2.09 2.20

0.947 0.924 0.932

0.909 0.863 0.891

4.12 3.10 4.15

0.878 0.884 0.872

0.694 0.751 0.676

2.96 3.20 2.18

0.900 0.895 0.866

0.818 0.789 0.673

2.04 2.01 1.53

0.881 0.875 N/A

0.743 0.735 N/A

2.47 2.15 N/A

UW

Llama Qwen Mistral

0.902 0.888 0.900

0.830 0.767 0.791

2.28 2.00 2.16

0.947 0.926 0.935

0.910 0.862 0.889

4.10 3.89 4.11

0.876 0.881 0.867

0.702 0.746 0.648

2.87 3.73 2.11

0.899 0.905 0.861

0.797 0.829 0.662

2.26 2.43 1.36

0.859 0.900 0.890

0.802 0.882 0.864

1.56 2.18 2.62

DIP

Llama Qwen Mistral

0.900 0.893 0.900

0.812 0.783 0.789

2.44 2.43 2.19

0.944 0.928 0.934

0.905 0.868 0.895

4.12 3.78 4.01

0.874 0.885 0.869

0.688 0.754 0.663

2.75 4.00 2.08

0.898 0.901 0.864

0.799 0.821 0.656

2.10 2.33 1.46

0.860 0.897 0.882

0.782 0.877 0.856

1.43 2.53 2.18

SynthID

Llama Qwen Mistral

0.899 0.887 0.903

0.819 0.771 0.805

2.26 2.09 2.32

0.942 0.925 0.939

0.903 0.868 0.898

4.30 3.75 3.99

0.873 0.882 0.877

0.677 0.748 0.686

2.60 3.53 2.51

0.901 0.898 0.865

0.818 0.814 0.686

2.29 2.23 1.51

0.860 0.899 0.891

0.793 0.881 0.870

1.58 2.54 2.45

MarkSec : Capability-Aware Evaluation of Adversarial Attacks Against LLM Watermarks

F

Cross-Watermark Quality-Gate Sensitivity

SynthID configuration. The evaluated MarkLLM SynthID-Text implementation uses iterative probability reweighting with 30 keyed binary 𝑔 -value channels over five-token n-grams, including four preceding context tokens. Its configuration uses a 65,536-entry sampling table, table seed 0, and a 1,024-context repetition history. The mean detector uses the victim tokenizer without added special tokens and excludes repeated contexts and positions from EOS onward. The stored z_score field contains the mean 𝑔 -value. All 45 SynthID runs use dynamic detection with increasing-score decisions and target calibration FPR 0.01. The run-specific threshold is selected from natural and watermarked score populations and applied as score ≥ 𝜏 ; thresholds range from 0.5099 to 0.5228. Configuration and implementation identifiers accompany the result data. Analysis population. The analysis uses 180 cells: three datasets, six watermarks, Llama and Qwen, and five attacks. These two victim models provide all five attacks for every watermark; Mistral has structural attack exclusions and remains in the full benchmark tables. Each entry averages the Llama and Qwen cell rates with equal weight at Prometheus gates 2, 3, and 4, using the QSR denominator in Section 5.2. The 17,996 outputs include 33 with at least one missing detector prediction; these outputs remain in the QSR denominator and receive no success credit. Table 14: Quality-gate sensitivity on common Llama/Qwen support. LLMP QSR@𝑢 (%)

LLMP minus best other (pp)

Dataset

Watermark

𝑢=2

𝑢=3

𝑢=4

𝑢=2

𝑢=3

𝑢=4

C4

KGW Unigram SIR UW DIP SynthID

68.00 68.00 53.00 77.00 70.00 87.50

64.50 65.50 51.50 75.50 66.50 85.00

56.00 59.50 48.50 70.00 60.50 77.00

+2.00 (S) +8.00 (S) +6.50 (S) +0.50 (D) +9.00 (D) +12.11 (D)

+2.50 (D) +11.50 (S) +7.00 (S) +6.00 (D) +12.50 (D) +16.65 (D)

+6.00 (S) +12.50 (S) +12.00 (S) +16.50 (D) +15.00 (S) +24.50 (S)

Dolly-15K

KGW Unigram SIR UW DIP SynthID

75.50 55.50 0.00 75.77 53.00 52.50

73.00 54.00 0.00 74.74 52.50 52.00

69.50 49.50 0.00 71.66 48.00 50.00

+14.50 (S) +12.00 (S) +0.00 (T) +2.77 (D) +7.00 (D) +9.50 (D)

+17.00 (S) +14.50 (S) +0.00 (T) +9.24 (D) +12.50 (D) +12.50 (D)

+21.00 (S) +16.00 (S) +0.00 (T) +23.16 (S) +18.50 (S) +18.50 (S)

BookReport

KGW Unigram SIR UW DIP SynthID

62.50 57.00 52.50 87.50 56.50 76.50

55.50 56.50 48.50 84.50 52.50 72.50

49.00 50.00 42.50 74.50 44.50 62.50

+5.00 (S) +10.00 (S) +3.50 (S) +11.00 (S) +3.50 (S) +24.50 (S)

+2.50 (S) +13.50 (S) +3.00 (S) +15.00 (S) +1.50 (S) +24.00 (S)

+9.50 (S) +18.00 (S) +7.50 (S) +22.00 (S) +3.00 (S) +27.00 (S)

Reading the table. Rates are percentages; differences are percentage points relative to the highest-rate non-LLMP attack at that gate. D denotes DIPPER and S denotes SIRA; T denotes a tie among all four non-LLMP attacks. The comparator is selected separately at each gate. The result data provide per-cell denominators, model-specific rates, and all five attack rankings.

Record · ID 919254 · SHA-256 b0ed57d3bf3b3aa6
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.