Watermarking Should Be Treated as a Monitoring Primitive
arXiv:2605.13095v1 [cs.CR] 13 May 2026
Toluwani Aremu1 Nils Lukas1 Jie Zhang2 1 MBZUAI, UAE 2 A*STAR, Singapore
Abstract Watermarking is widely proposed for provenance, attribution, and safety monitoring in generative models, yet is typically evaluated only under adversaries who attempt to evade detection or induce false positives at the level of individual samples. We argue that watermarking should be treated as a monitoring primitive, and that internal monitoring is unavoidable given per-entity attribution keys and messages, as well as detector access. We introduce an observer-based threat model in which observers can aggregate watermark signals across outputs to infer entitylevel information, showing that even zero-bit watermarking enables attribution under multi-key settings. We further show that external monitoring can emerge over time from persistent, key-dependent statistical structure, although this depends on watermark design and may be mitigated by distribution-preserving or undetectable schemes. Our findings reveal a fundamental dual-use tension between attribution and monitoring, motivating evaluation of watermarking beyond per-sample robustness to account for aggregation and observer-based capabilities.
1
Introduction
Watermarking has emerged as a promising mechanism for establishing provenance [Zhao et al., 2024b, Pang et al., 2024a, Zhou et al., 2024], enabling attribution [Aaronson, 2023, Kirchenbauer et al., 2023a, Liu et al., 2024, Hou et al., 2023, Dathathri et al., 2024], and supporting safety monitoring [Aremu et al., 2026] in generative models. By embedding detectable signals into model outputs, watermarking allows downstream systems to distinguish AI-generated content, enforce usage policies, and provide accountability in increasingly automated pipelines. As generative models become widely deployed, watermarking is increasingly positioned as a key building block for trustworthy and responsible AI systems [Bartz et al., 2023, EU AI Act, 2024, California Legislature, 2024]. Existing work on watermarking primarily evaluates security under adversaries who attempt to evade detection [Diaa et al., 2024, Lukas et al., 2024, Pang et al., 2024b, Wu and Chandrasekaran, 2024, Krishna et al., 2023] or induce false positives [Jovanović et al., 2024, Gloaguen et al., 2024, Aremu et al., 2025, Müller et al., 2025]. This has led to a focus on robustness at the level of individual samples, measuring whether watermark signals persist under paraphrasing, rewriting, or other transformations [Kirchenbauer et al., 2023b, Pan et al., 2024, Piet et al., 2023, Zhao et al., 2024a, Christ et al., 2024a]. While these threat models are important, they capture only one side of the security landscape: adversaries who seek to remove, spoof, or manipulate watermark signals. Position. Watermarking should be treated as a monitoring primitive, as it enables entity-level inference when signals are aggregated across outputs. Rather than treating watermark signals solely as targets for removal or forgery, we consider the capabilities they enable when observed over time. Even weak, per-sample signals can accumulate across multiple outputs to reveal usage patterns, link related content, and support attribution. Watermarking, as increasingly mandated by emerging regulatory and standardization efforts (e.g., [EU AI Act, 2024]), effectively introduces a persistent monitoring capability that cannot be cleanly Preprint.
Key / Message
Prompt
Key
Key
Prompt
Was this generated, using this key and model?
Which entity (key) most likely generated this output?
Decode (Optional) Is the decoded message accurate? Generation
Detection
Generation
Attribution
Prompt
Do these outputs come from the same source?
Generation
Reidentification
Figure 1: Comparison of watermarking usage under different observer models. Left: Standard watermarking, where a detector determines whether an output is watermarked and optionally decodes an embedded message. Middle: Internal observer, who has access to watermark keys and performs attribution by identifying which entity generated an output. Right: External observer, who does not have access to keys and instead learns to identify which entity generated an output from publicly observed data by exploiting watermark-induced statistical patterns. This illustrates a shift from per-sample detection to entity-level inference, showing that watermarking can act as a monitoring primitive by enabling user attribution and re-identification over time. separated from its intended roles in attribution and safety. To formalize this perspective, we introduce an observer-based threat model in which observers passively aggregate watermark signals across outputs. In this setting, monitoring arises directly from watermarking design. This is immediate in multi-bit watermarking schemes that explicitly encode information [Wang et al., 2024]. More importantly, we show that it is inherent in zero-bit watermarking under multi-key deployments, where distinct keys induce persistent statistical structure that enables entity-level attribution even without explicit identity encoding. We also show that such structure may support linkability across outputs, creating a pathway toward re-identification by any actor without knowledge of the watermark, depending on how closely the watermark preserves the underlying data distribution. Hence, the same properties that enable detection, attribution, and safety monitoring also enable tracking and inference over time, which are not captured by current evaluation protocols. As a result, robustness-focused evaluations may significantly understate the monitoring capabilities of watermarking systems. We therefore argue that watermarking evaluation must extend beyond per-sample robustness to account for aggregation and observer-based capabilities. In particular, evaluation should distinguish between inherent monitoring (internal observers with key access) and emergent monitoring (external observers without keys). This reframing introduces a new dimension in watermark design: balancing robustness and detectability with the potential for monitoring and privacy leakage. Contributions. We (i) introduce an observer-based threat model for watermarking, where observers aggregate signals across outputs to perform entity-level inference, (ii) show that monitoring is inherent under multi-key deployments, even in zero-bit watermarking without explicit identity encoding, (iii) demonstrate that persistent statistical structure can support attribution and re-identification over time, and (iv) highlight a fundamental dual-use tension, arguing that watermarking should be evaluated beyond per-sample robustness to account for aggregation and observer-based capabilities.
2
Background
Generative Models. Modern generative models map an input prompt p ∈ P to an output x ∈ X , where x may represent text, images, or other modalities. Formally, a model samples x ∼ M(· | p) from a conditional distribution over outputs given the prompt [Achiam et al., 2023, Bubeck et al., 2023]. These models are widely deployed across applications, where outputs may be consumed, transformed, or redistributed in downstream pipelines. Watermarking. Watermarking embeds a detectable signal into generated content to enable downstream verification [Zhao et al., 2024b]. A watermarking scheme typically consists of a secret key k, an embedding procedure that modifies generation, and a detector that determines whether a given output contains the watermark. Watermarks may be zero-bit, indicating only presence or absence, or 2
multi-bit, encoding additional information such as identifiers [Zhao et al., 2024b, Wang et al., 2024]. A key design goal is robustness, i.e., the watermark should remain detectable under transformations [Kirchenbauer et al., 2023b, Christ et al., 2024a]. Threat Models for Watermarking. Prior work primarily studies watermarking under adversaries who aim to disrupt or exploit the watermark signal. Common threats include evasion, where outputs are modified to remove the watermark [Krishna et al., 2023, Diaa et al., 2024, Pang et al., 2024b], forgery, where non-watermarked content is attributed to a watermarked source [Aremu et al., 2025, Jovanović et al., 2024, Gloaguen et al., 2024, Müller et al., 2025], and secret extraction, where the watermarking key or decision boundary is inferred [Zhang et al., 2024, Gu et al., 2024]. These threat models focus on adversaries who manipulate individual outputs. Watermarking for Monitoring. Recent work has explored using watermarking for safety monitoring, embedding signals into model behavior to detect policy-violating outputs [Aremu et al., 2026]. This expands watermarking beyond provenance into detection of unsafe behavior. Our work is also related to classical and modern approaches to attribution and fingerprinting [Kumarage et al., 2024], including traitor tracing [Kumarage et al., 2023], stylometry [Przystalski et al., 2025], and fingerprinting [Kumarage and Liu, 2023]. These methods show that aggregation across samples can reveal source identity even when individual observations are weak, typically relying on intrinsic properties of the data. In contrast, we show that similar linkability arises as a consequence of watermarking design and deployment, particularly under multi-key settings. This reframes watermarking from a purely defensive mechanism into a system that can enable monitoring under realistic deployment conditions. From Detection to Monitoring. Building on these perspectives, we consider a broader notion of monitoring, where watermark signals are aggregated across outputs to support entity-level inference. This shifts the focus from per-sample detection to cross-sample inference and motivates the observerbased threat model introduced in the next section.
3
Threat Model
We formalize an observer-based threat model for watermarking, focusing on entities that exploit watermark signals to perform monitoring over time. Unlike prior work that considers adversaries manipulating individual outputs, we study observers that passively aggregate signals across multiple outputs. Setting. Let M denote a generative model that maps a prompt p ∈ P to an output x ∈ X , where x ∼ M(· | p). A watermarking scheme modifies generation using a secret key k ∈ K to produce watermarked outputs. Given an output x, a detector Dk (x) produces a score or decision indicating the presence of a watermark under key k. We consider a set of entities E = {e1 , . . . , en } interacting (e) with the model over time, each generating a sequence of outputs {xt }Tt=1 . Observer. An observer O passively observes outputs over time and aggregates signals to infer information about the generating entities. The observer does not modify outputs or interact with the generation process. We distinguish two types of observers: (Internal observer.) The observer has access to watermark detectors and keys. Under multi-key deployments, each entity may be associated with a distinct key ke , allowing the observer to evaluate Dke (x) and directly attribute outputs to entities. (External observer.) The observer does not have access to watermark keys. Instead, it relies on observable outputs and applies statistical or learned methods to extract signals, aggregating weak evidence across samples to infer relationships between outputs. Capabilities. The observer is assumed to have: (i) access to a stream of outputs over time, (ii) the ability to evaluate watermark detectors (internal) or compute surrogate signals (external), and (iii) the ability to aggregate observations across multiple samples. The observer does not control generation or modify outputs. Goals. The observer aims to perform: (Monitoring.) Determine whether and how frequently an entity uses the model. (Linkability.) Determine whether multiple outputs originate from the same entity. (Re-identification.) Associate outputs with specific entities using watermark-based signals. Key Distinction. Prior watermarking threat models focus on adversaries acting on individual samples. In contrast, our model considers observers that exploit the persistence of watermark signals across multiple samples. This shift from per-sample robustness to cross-sample aggregation changes what 3
Example Monitoring Scenarios Enabled by Watermarking Provider-side attribution
Platform-level aggregation
A model provider assigns distinct watermarking keys to users or accounts. With detector access, the provider can attribute outputs to specific entities and monitor their usage over time for auditing, abuse detection, or enforcement.
A platform hosting AI-generated posts, messages, or media may aggregate outputs across users. Even without prompts, repeated watermark detections can reveal which entities are active and how frequently they use the system.
External source identification
Coordinated-activity surveillance
A third party collects public outputs believed to be watermarked and trains a classifier on observable patterns. Over time, the observer may identify which entity generated unseen outputs without access to watermark keys or detectors.
A regulator or law-enforcement agency may use provider-accessible watermark detectors to monitor a set of entities suspected of coordinated harmful activity. Attribution over time can reveal behavioral patterns or collaboration signals without direct account access.
Figure 2: Example scenarios in which watermarking can enable monitoring. The first two arise naturally for internal observers with detector access, while the latter two illustrate how monitoring may also extend to external observers or institutional surveillance settings.
constitutes a successful attack or use of watermarking systems. We illustrate representative scenarios in Figure 2. Importantly, these scenarios do not require malicious intent; monitoring can arise naturally from watermarking design and deployment choices. We formalize this in the next section.
4
Watermarking as a Monitoring Primitive
We now show that watermarking inherently enables monitoring under the observer-based threat model introduced in Section 3 and Figure 1. Our key observation is that watermarking introduces a persistent, detectable signal into generated content, which can be aggregated across outputs to support entity-level inference over time. Importantly, we do not assume a specific notion of behavior, task, or intent for the observer. Our goal is to characterize this previously underexplored capability induced by watermarking, that persistent signals, when observed across multiple outputs, can enable monitoring, irrespective of whether the use is benign or adversarial. 4.1
Conceptual Description
Watermarking is typically evaluated as a per-sample detection problem, i.e., given an output x, a detector Dk (x) determines whether the watermark is present. Formally, a watermarking scheme consists of an embedding function Ek and a detector Dk , parameterized by a secret key k ∈ K. Given a prompt p, the model generates x ∼ Ek (M(· | p)), (1) where the embedding process biases generation according to k. The detector computes a statistic s = Dk (x),
(2)
and determines whether x is watermarked via a hypothesis test s ≷ τ . In the observer setting, detection is applied across a sequence of outputs {xt }Tt=1 . Rather than making per-sample decisions, the observer aggregates signals across samples, allowing even weak per-output signals to induce stable, entity-level structure. In multi-bit watermarking schemes, the embedding function encodes a message m ∈ M, x ∼ Ek,m (M(· | p)), 4
(3)
and the detector recovers m̂ = Dk (x) [Wang et al., 2024]. In this case, monitoring is immediate, as the observer can directly decode entity-level information. We now consider the more subtle case of zero-bit watermarking, where no explicit identity is encoded. Under multi-key deployments [Aremu et al., 2025], each entity e ∈ E is associated with a distinct key ke , inducing a key-dependent distribution x ∼ Eke (M(· | p)).
(4)
An internal observer with access to {ke } can directly attribute outputs by evaluating Dke (x) across keys. More importantly, even without access to keys, an external observer may exploit these persistent statistical differences. Let ϕ(x) denote observable features derived from x (e.g., lexical or embedding-based features). Given public outputs from known entities, the observer can train a sourceidentification model to predict which entity generated an unseen output. Thus, watermark-induced statistical structure alone can support entity-level inference. 4.2
Entity Re-identification
We formalize entity identification as determining whether two outputs xi and xj originate from the same entity. Let e(x) denote the (unknown) source of x. The observer aims to infer whether e(xi ) = e(xj ).
(5)
Internal Observer. An observer with access to watermark keys evaluates Dke (x) for each candidate key and selects ê(x) = arg max Dke (x), (6) e∈E
enabling direct attribution and tracking. External Observer. An observer without access to keys relies on observable structure. Given labeled outputs {(xi , ei )}N i=1 , it trains a classifier f : ϕ(x) 7→ ê and predicts ê(x) = f (ϕ(x)).
(7)
If watermarking induces persistent key-dependent structure, the observer can identify the most likely source without access to watermark mechanisms.
5
Experiments
Watermarking Methods. We evaluate multiple zero-bit watermarking methods for both text and image generation. For text, we instantiate several methods [Kirchenbauer et al., 2023a, Christ et al., 2024b, Zhao et al., 2024a, Lu et al., 2024, Wang et al., 2025, Hou et al., 2023, Dathathri et al., 2024, Gu et al., 2025, Lee et al., 2024] from MarkLLM [Pan et al., 2024]. For images, we instantiate several methods [Wen et al., 2023, Yang et al., 2024, Arabi et al., 2024] from MarkDiffusion [Pan et al., 2025]. Our goal is not to benchmark watermarking methods, but to test whether zero-bit watermarking under multi-key deployment enables monitoring. Models. We use Qwen2.5-14B [Team, 2024] for text and Stable Diffusion v2.1 [Rombach et al., 2022] for image generation. For each modality, we study watermarking under standard deployment and under multi-key deployment, where each entity is assigned a distinct watermarking key. Datasets. We use shared prompt pools across entities to ensure that attribution and linkability are not trivially explained by prompt differences. For text, we use C4 [Raffel et al., 2019], a set of Common Crawl’s corpus spanning multiple content categories. For images, we use the Stable Diffusion Prompt dataset1 . For our experiments, we use a prompt-matched setting, where all entities generate outputs from the same prompt under different keys. We ensure that training and evaluation are performed on disjoint sets of generated outputs, with no overlap in prompts or samples across splits, and that prompts are partitioned such that training and test sets are prompt-disjoint. This prevents leakage and ensures that external observer performance reflects generalization rather than memorization. Metrics. For internal observers, we report top-1 attribution accuracy (TPR@1%FPR). Concretely, for each key k, we calibrate a detection threshold τk to achieve a 1% false positive rate on nonmatching samples (i.e., outputs generated under other keys). Given an output x, we compute detector 1 https://huggingface.co/datasets/Gustavosta/Stable-Diffusion-Prompts
5
scores Dk (x) for all candidate keys and attribute the output to the entity whose key yields the highest score. The reported TPR is the fraction of correctly attributed samples under this argmax decision rule. We note that the 1% FPR is controlled per key, and does not directly correspond to a global false positive rate under multi-key selection, as correlations between detector scores may affect attribution at larger scales. For external observers, we report standard top-1 and top-3 classification accuracy, where random guessing corresponds to 1/n and 3/n, respectively, for n entities. External observer evaluation is performed on a held-out set of 100 samples per entity.
5.1
Internal Observer: Attribution under Zero-Bit Multi-Key Watermarking
KGW
Unigram
EWD
MorphMark
EXP
IE
SWEET
SEMSTAMP
SynthID
TreeRing
Wind
Guassian Shading
1.0
TPR@1%FPR
0.9 0.8 0.7 0.6 0.5 1.0
TPR@1%FPR
0.9 0.8 0.7 0.6 0.5 1.0
TPR@1%FPR
0.9 0.8 0.7 0.6 0.5
4
8
12
Number of Entities
16
4
8
12
Number of Entities
16
4
8
12
Number of Entities
16
4
8
12
Number of Entities
16
Figure 3: Internal attribution performance under zero-bit multi-key watermarking. We report the top-1 attribution accuracy (TPR@1%FPR) as the number of entities increases across watermarking methods.
We evaluate the internal observer setting under zero-bit watermarking with multi-key deployment. In this setting, each entity is assigned a distinct watermarking key, and the observer has access to the corresponding detectors. The observer’s goal is to identify which entity generated a given output. We measure top-1 attribution accuracy (TPR@1%FPR) as the number of candidate entities increases (from 1 to 16), using multiple watermarking methods across text and image generation. Each entity contributes 100 samples for evaluation, and attribution is performed by selecting the key that yields the highest detector score among all candidates. Figure 3 shows that attribution remains consistently high across all watermarking methods, with only mild degradation as the number of entities increases. Most methods achieve near-perfect attribution for small numbers of entities, and maintain strong performance even at larger scales. While some methods exhibit slight drops at higher entity counts (e.g., Unigram, SEMSTAMP, and TreeRing), overall attribution accuracy remains well above chance. These results demonstrate that zero-bit watermarking under multi-key deployment enables reliable attribution. Despite the absence of explicit identity encoding, the watermarking process introduces consistent statistical structure that allows an internal observer to monitor entities over time. 6
5.2
External Observer: Emergent Re-Identification through Aggregation over Time
We evaluate the external observer setting, where the observer does not have access to watermark keys or detectors, but can collect public outputs over time and learn to infer their source from observable structure. We consider n ∈ {2, 4, 8, 16} entities, each assigned a distinct watermarking key under a zero-bit watermarking scheme. For text, we evaluate KGW watermarking, and for images, we evaluate Tree-Ring watermarking. The observer trains a classifier (BERT-Base [Devlin et al., 2018] for text, CLIP-RN50 [Radford et al., 2021] for images) on observable features to predict the generating entity. Training uses between 100 and 4000 samples per entity (batch size 16, 10 epochs, AdamW/Adam, learning rate 3 × 10−5 ), with evaluation on a held-out set of 100 samples per entity. We report top-1 and top-3 identification accuracy, where random guessing corresponds to 1/n. 2 Entities (KGW)
1.0
4 Entities (KGW)
8 Entities (KGW)
16 Entities (KGW)
Accuracy
0.8 0.6
Top-3 Accuracy Top-1 Accuracy Random Baseline
0.4 0.2 0.0
2 Entities (Tree-Ring)
1.0
4 Entities (Tree-Ring)
8 Entities (Tree-Ring)
16 Entities (Tree-Ring)
Accuracy
0.8 0.6
Top-3 Accuracy Top-1 Accuracy Random Baseline
0.4 0.2 0.0
100
1000
2000
3000
Samples per Entity
4000
100
1000
2000
3000
Samples per Entity
4000
100
1000
2000
3000
Samples per Entity
4000
100
1000
2000
3000
Samples per Entity
4000
Number of Samples Observed (per entity) Over Time
Figure 4: External observer identification under zero-bit multi-key watermarking across text (KGW) and image (Tree-Ring) models. We report top-1 and top-3 accuracy as a function of the number of samples observed per entity for n ∈ {2, 4, 8, 16} entities. Random guessing corresponds to 1/n. Identification accuracy is initially near random, but improves substantially as more samples are observed. These results demonstrate that external monitoring emerges over time through aggregation, even without access to watermark keys or detectors. Figure 4 shows that identification is initially near random, but improves substantially as more samples are observed. For example, under KGW with 16 entities, top-1 accuracy increases from 11.3% (near the 6.25% random baseline) to 73.0%, while top-3 accuracy reaches 90.0%. Tree-Ring exhibits similar but faster convergence, achieving higher accuracy with fewer samples (e.g., 91.0% top-1 for n = 16 at 4000 samples per entity). These results show that external monitoring emerges over time through aggregation, even without access to watermark mechanisms. For tractability, we evaluate up to n = 16 entities, which is sufficient to demonstrate the emergence of this effect. Importantly, many practical monitoring scenarios are targeted i.e., where the observer seeks to identify a specific entity rather than perform full multi-class attribution. This reduces the problem to a one-vs-all task, which is substantially easier than full multi-class attribution and may require fewer samples. We discuss implications for larger-scale and targeted monitoring in Section 6. Controls. To isolate the role of watermarking in enabling external identification, we evaluate two additional settings (see Figure 5). First, we consider a no-watermark baseline, where outputs are generated without watermarking but labeled during evaluation by entity. In this case, identification accuracy remains near random, with slight deviations at small n due to finite-sample effects and residual classifier bias, but quickly collapses toward chance as the number of entities increases. Second, we evaluate a shared-key setting, where all entities use the same watermarking key. Here, identification accuracy again collapses toward chance. These controls confirm that the observed performance arises from key-dependent watermarking structure rather than spurious correlations. This behavior arises because, under a shared-key deployment, all entities induce identical watermarkconditioned distributions, eliminating the key-dependent structure required for attribution. 7
2 Entities (KGW)
4 Entities (KGW)
8 Entities (KGW)
16 Entities (KGW) Random Baseline
1.0
Accuracy
0.8 0.6 0.4 0.2 0.0
2 Entities (Tree-Ring)
4 Entities (Tree-Ring)
8 Entities (Tree-Ring)
16 Entities (Tree-Ring)
1.0
Accuracy
0.8 0.6 0.4 0.2 0.0
Internal
External
No-WM Shared-Key
Internal
External
No-WM Shared-Key
Internal
External
No-WM Shared-Key
Internal
External
No-WM Shared-Key
Observer Scenario per Number of Observed Entities
Figure 5: Control experiments isolating the role of watermarking in enabling monitoring. We compare four settings: internal observer (with key access), external observer (learned classifier), no watermark, and shared-key deployment. Results are shown for both text (KGW) and image (Tree-Ring) watermarking across n ∈ {2, 4, 8, 16} entities, with random guessing indicated by the dashed baseline (1/n). Internal attribution remains near-perfect under multi-key deployment, while external identification remains strong but lower. In contrast, both no-watermark and shared-key settings collapse toward random performance, confirming that monitoring arises from key-dependent watermark structure rather than content or prompt artifacts.
6
Discussion
Implications for Monitoring. Our results establish a clear distinction between internal and external monitoring capabilities in watermarking systems. For internal observers, monitoring is inherent under multi-key deployments. When each entity is assigned a distinct watermarking key, attribution follows directly from detector access, even for zero-bit watermarking schemes. This implies that monitoring is not a side effect, but a direct consequence of the system design. Importantly, there is no technical mechanism to prevent such monitoring once per-entity keys are deployed. The only viable constraint is at the level of system design or governance, for example by enforcing a shared global key or message across users. However, such approaches weaken attribution guarantees, complicate auditing, and introduce operational challenges such as key rotation and reduced robustness. While alternative deployment strategies such as shared keys or rotating group keys may reduce monitoring granularity, they come at the cost of reduced utility. In contrast, external monitoring is an emergent capability. Our results show that an external observer can learn to identify entities from public outputs over time, even without access to watermark keys or detectors. However, this capability depends on the presence of persistent, key-dependent statistical structure in generated outputs, and is therefore not guaranteed across all watermarking designs. Mitigations and Design Considerations. For internal observers, mitigation is fundamentally limited. Since monitoring follows directly from key- or bit-based attribution, reducing it requires restricting key assignment, for example through shared or group-level keys. However, this comes at the cost of weaker attribution and reduced utility. For external observers, mitigation depends on watermark design. Schemes that aim to satisfy distribution-preserving and undetectability properties [Christ and Gunn, 2024, Christ et al., 2024b, Gunn et al., 2024] are expected to reduce the statistical signals exploited in our experiments. In a preliminary experiment with EXP [Aaronson, 2023] and EXP-Edit [Kuditipudi et al., 2023], external identification remains near chance even as training data increases, suggesting that reducing key-dependent distortion can weaken monitoring signals. We include this result only as preliminary evidence for the mitigation hypothesis, not as a comprehensive evaluation of undetectable watermarking. Even if such designs eliminate external monitoring, they do not affect internal monitoring under multi-key or entity-focused multi-bit deployments. On Scaling and Partial Supervision. A common concern is whether external identification scales to large numbers of entities. While multi-class attribution becomes more challenging as the number 8
of entities grows, many practical monitoring scenarios are inherently targeted. In such cases, the observer seeks to determine whether a specific entity is responsible for a given output, reducing the problem to a one-vs-all task that is substantially easier. Similarly, while external identification may appear to require labeled data, many realistic scenarios are semi-supervised. For example, an observer may have access to outputs from a known entity and seek to identify additional outputs generated by the same entity. In this setting, binary classification or clustering can be used to separate outputs corresponding to the target entity from others. This suggests that monitoring may remain feasible in practice even when scaling or labeling assumptions are relaxed. Alternative Views and Limitations. A natural counterargument is that watermarking schemes can be designed to avoid the risks identified in this work, particularly through distribution-preserving or undetectable constructions [Zhao et al., 2024b]. We agree with this perspective in part. Our results apply to watermarking schemes that introduce persistent, key-dependent structure, which includes many practical methods used in current deployments. Whether monitoring remains possible under strictly distribution-preserving watermarking remains an open question [Zhao et al., 2024b]. Prior work has shown that watermark signals can be inferred [Gloaguen et al., 2025] or attacked [Jovanović et al., 2024, Pang et al., 2024b] in black-box settings, and that even robust watermarking schemes may exhibit residual structure under practical conditions [Liu et al., 2025]. This suggests that whether such designs fully eliminate external monitoring capabilities remains to be empirically validated. However, even under ideal watermark designs, our results for internal observers remain unaffected. As long as distinct keys or messages are assigned to different entities, monitoring is unavoidable for any observer with access to the corresponding detectors or decoders. This highlights a fundamental asymmetry between internal and external monitoring. Our experiments also do not evaluate robustness of external identification under variations in decoding strategies (e.g., temperature, sampling) or post-processing (e.g., paraphrasing or summarization), which are known to affect watermark detectability [Pan et al., 2024, Piet et al., 2023]. Such transformations may weaken per-sample signals and reduce external monitoring effectiveness, although the extent to which aggregation compensates for this effect remains an open question. In contrast, internal observers with detector access may remain more robust to such transformations. Finally, our external observer experiments are limited to specific models and watermarking schemes. While we observe consistent trends across modalities, further work is needed to evaluate generalization across models, watermark designs, observer capabilities, and real-world conditions. Accordingly, our results demonstrate feasibility for a broad class of practical watermarking schemes, rather than a universal property of all possible designs. Takeaway. Watermarking evaluation should explicitly account for monitoring risk. At a minimum, we recommend reporting (i) attribution or identification accuracy as a function of the number of entities, and (ii) performance as a function of samples observed per entity over time. This reframing introduces monitoring as a first-class evaluation dimension alongside robustness and detectability.
7
Conclusion
We argue that watermarking should be treated as a monitoring primitive. We show that internal monitoring is unavoidable under multi-key deployments, even for zero-bit watermarking, while external monitoring can emerge over time through aggregation depending on watermark design. These findings suggest that existing regulatory and standardization efforts may be incomplete if they treat watermarking solely as a mechanism for provenance and attribution. We argue that watermarking systems should be evaluated and governed as monitoring technologies, with explicit consideration of how deployment choices (e.g., per-entity keying) enable tracking and inference over time and under realistic observer models. We therefore highlight the need for greater transparency and call for broader discussion on how watermarking systems should be designed, evaluated, and governed in light of their inherent monitoring capabilities.
Ethical Considerations This work highlights a dual-use property of watermarking systems. While watermarking is intended to support provenance, attribution, and safety monitoring, our results show that it can also enable 9
tracking and inference over time. We do not advocate for the use of watermarking for surveillance, but aim to inform the design and evaluation of such systems by identifying monitoring as an inherent or emergent capability. We encourage careful consideration of transparency, user awareness, and governance in the deployment of watermarking technologies.
References S. Aaronson. Watermarking of large language models. https://www.youtube.com/watch?v=2Kx9jbSMZqA, 2023. Simons Institute, YouTube video. J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023. K. Arabi, B. Feuer, R. T. Witter, C. Hegde, and N. Cohen. Hidden in the noise: Two-stage robust watermarking for images. arXiv preprint arXiv:2412.04653, 2024. T. Aremu, N. Hussein, M. Nwadike, S. Poppi, J. Zhang, K. Nandakumar, N. Gong, and N. Lukas. Mitigating watermark forgery in generative models via randomized key selection. arXiv preprint arXiv:2507.07871, 2025. T. Aremu, D. Ognev, S. Poppi, and N. Lukas. Robust safety monitoring of language models via activation watermarking. arXiv preprint arXiv:2603.23171, 2026. D. Bartz, K. Hu, D. Bartz, and K. Hu. Openai, google, others pledge to watermark ai content for safety, white house says. Reuters, Jul 2023. URL https://www.reuters.com/technology/ openai-google-others-pledge-watermark-ai-content-safety-white-house-2023-07-21/. S. Bubeck, V. Chadrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, et al. Sparks of artificial general intelligence: Early experiments with gpt-4, 2023. California Legislature. California ai transparency act (sb 942). California Legislative Information, Sept. 2024. URL https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_ id=202320240SB942. Chapter 291, Statutes of 2024; operative Jan 1, 2026. M. Christ and S. Gunn. Pseudorandom error-correcting codes. In Annual International Cryptology Conference, pages 325–347. Springer, 2024. M. Christ, S. Gunn, T. Malkin, and M. Raykova. Provably robust watermarks for open-source language models. arXiv preprint arXiv:2410.18861, 2024a. M. Christ, S. Gunn, and O. Zamir. Undetectable watermarks for language models. In The Thirty Seventh Annual Conference on Learning Theory, pages 1125–1139. PMLR, 2024b. S. Dathathri, A. See, S. Ghaisas, P.-S. Huang, R. McAdam, J. Welbl, V. Bachani, A. Kaskasoli, R. Stanforth, T. Matejovicova, J. Hayes, N. Vyas, M. A. Merey, J. Brown-Cohen, R. Bunel, B. Balle, A. T. Cemgil, Z. Ahmed, K. Stacpoole, I. Shumailov, C. Baetu, S. Gowal, D. Hassabis, and P. Kohli. Scalable watermarking for identifying large language model outputs. Nat., 634(8035):818–823, October 2024. URL https: //doi.org/10.1038/s41586-024-08025-4. J. Devlin, M. Chang, K. Lee, and K. Toutanova. BERT: pre-training of deep bidirectional transformers for language understanding. CoRR, abs/1810.04805, 2018. URL http://arxiv.org/abs/1810.04805. A. Diaa, T. Aremu, and N. Lukas. Optimizing adaptive attacks against watermarks for language models. arXiv preprint arXiv:2410.02440, 2024. EU AI Act. Artificial intelligence act. Official Journal of the European Union, July 2024. URL http: //data.europa.eu/eli/reg/2024/1689/oj. Adopted 13 June 2024; OJ L, 12 July 2024. T. Gloaguen, N. Jovanović, R. Staab, and M. Vechev. Discovering clues of spoofed lm watermarks. arXiv preprint arXiv:2410.02693, 2024. T. Gloaguen, N. Jovanović, R. Staab, and M. Vechev. Black-box detection of language model watermarks. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview. net/forum?id=E4LAVLXAHW. C. Gu, X. L. Li, P. Liang, and T. Hashimoto. On the learnability of watermarks for language models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/ forum?id=9k0krNzvlV.
10
T. Gu, Z. Wang, K. Huang, Y. Yao, X. Zhang, Y. Yang, and X. Chen. Invisible entropy: Towards safe and efficient low-entropy llm watermarking. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 6727–6744, 2025. S. Gunn, X. Zhao, and D. Song. An undetectable watermark for generative image models. arXiv preprint arXiv:2410.07369, 2024. A. B. Hou, J. Zhang, T. He, Y. Wang, Y.-S. Chuang, H. Wang, L. Shen, B. Van Durme, D. Khashabi, and Y. Tsvetkov. Semstamp: A semantic watermark with paraphrastic robustness for text generation. arXiv preprint arXiv:2310.03991, 2023. N. Jovanović, R. Staab, and M. Vechev. Watermark stealing in large language models. ICML, 2024. J. Kirchenbauer, J. Geiping, Y. Wen, J. Katz, I. Miers, and T. Goldstein. A watermark for large language models. In International Conference on Machine Learning, pages 17061–17084. PMLR, 2023a. J. Kirchenbauer, J. Geiping, Y. Wen, M. Shu, K. Saifullah, K. Kong, K. Fernando, A. Saha, M. Goldblum, and T. Goldstein. On the reliability of watermarks for large language models. arXiv preprint arXiv:2306.04634, 2023b. K. Krishna, Y. Song, M. Karpinska, J. Wieting, and M. Iyyer. Paraphrasing evades detectors of ai-generated text, but retrieval is an effective defense. Advances in Neural Information Processing Systems, 36:27469–27500, 2023. R. Kuditipudi, J. Thickstun, T. Hashimoto, and P. Liang. Robust distortion-free watermarks for language models. Trans. Mach. Learn. Res., 2024, 2023. URL https://api.semanticscholar.org/CorpusID: 260315804. T. Kumarage and H. Liu. Neural authorship attribution: Stylometric analysis on large language models. In 2023 International conference on cyber-enabled distributed computing and knowledge discovery (cyberc), pages 51–54. IEEE, 2023. T. Kumarage, J. Garland, A. Bhattacharjee, K. Trapeznikov, S. Ruston, and H. Liu. Stylometric detection of ai-generated text in twitter timelines. arXiv preprint arXiv:2303.03697, 2023. T. Kumarage, G. Agrawal, P. Sheth, R. Moraffah, A. Chadha, J. Garland, and H. Liu. A survey of ai-generated text forensic systems: Detection, attribution, and characterization. arXiv preprint arXiv:2403.01152, 2024. T. Lee, S. Hong, J. Ahn, I. Hong, H. Lee, S. Yun, J. Shin, and G. Kim. Who wrote this code? watermarking for code generation. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 4890–4911, 2024. A. Liu, L. Pan, X. Hu, S. Li, L. Wen, I. King, and P. S. Yu. An unforgeable publicly verifiable watermark for large language models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=gMLQwKDY3N. Y. Liu, X. Zhao, D. X. Song, G. W. Wornell, and Y. Bu. Position: Llm watermarking should align stakeholders’ incentives for practical adoption. ArXiv, abs/2510.18333, 2025. URL https://api.semanticscholar. org/CorpusID:282246443. Y. Lu, A. Liu, D. Yu, J. Li, and I. King. An entropy-based text watermarking detection method. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 11724–11735, 2024. N. Lukas, A. Diaa, L. Fenaux, and F. Kerschbaum. Leveraging optimization for adaptive attacks on image watermarks. In The Twelfth International Conference on Learning Representations, 2024. URL https: //openreview.net/forum?id=O9PArxKLe1. A. Müller, D. Lukovnikov, J. Thietke, A. Fischer, and E. Quiring. Black-box forgery attacks on semantic watermarks for diffusion models. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 20937–20946, 2025. L. Pan, A. Liu, Z. He, Z. Gao, X. Zhao, Y. Lu, B. Zhou, S. Liu, X. Hu, L. Wen, et al. Markllm: An open-source toolkit for llm watermarking. arXiv preprint arXiv:2405.10051, 2024. L. Pan, S. Guan, Z. Fu, L. Si, Z. Wang, X. Hu, I. King, P. S. Yu, A. Liu, and L. Wen. Markdiffusion: An open-source toolkit for generative watermarking of latent diffusion models. arXiv preprint arXiv:2509.10569, 2025.
11
Q. Pang, S. Hu, W. Zheng, and V. Smith. No free lunch in llm watermarking: Trade-offs in watermarking design choices. In Neural Information Processing Systems, 2024a. URL https://api.semanticscholar.org/ CorpusID:267938448. Q. Pang, S. Hu, W. Zheng, and V. Smith. Attacking LLM watermarks by exploiting their strengths. In ICLR 2024 Workshop on Secure and Trustworthy Large Language Models, 2024b. URL https://openreview. net/forum?id=P2FFPRxr3Q. J. Piet, C. Sitawarin, V. Fang, N. Mu, and D. Wagner. Mark my words: Analyzing and evaluating language model watermarks. ArXiv, abs/2312.00273, 2023. URL https://api.semanticscholar.org/CorpusID: 265552122. K. Przystalski, J. K. Argasiński, I. Grabska-Gradzińska, and J. Ochab. Stylometry recognizes human and llm-generated texts in short samples. Expert Systems with Applications, page 129001, 2025. A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PmLR, 2021. C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv e-prints, 2019. R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. Q. Team. Qwen2.5: A party of foundation models, September 2024. URL https://qwenlm.github.io/ blog/qwen2.5/. L. Wang, W. Yang, D. Chen, H. Zhou, Y. Lin, F. Meng, J. Zhou, and X. Sun. Towards codable watermarking for injecting multi-bits information to LLMs. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=JYu5Flqm9D. Z. Wang, T. Gu, B. Wu, and Y. Yang. Morphmark: Flexible adaptive watermarking for large language models. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 4842–4860, 2025. Y. Wen, J. Kirchenbauer, J. Geiping, and T. Goldstein. Tree-ring watermarks: Fingerprints for diffusion images that are invisible and robust. Advances in Neural Information Processing Systems, 37, 2023. Q. Wu and V. Chandrasekaran. Bypassing llm watermarks with color-aware substitutions. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8549–8581, 2024. Z. Yang, K. Zeng, K. Chen, H. Fang, W. Zhang, and N. Yu. Gaussian shading: Provable performance-lossless image watermarking for diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12162–12171, 2024. Z. Zhang, X. Zhang, Y. Zhang, L. Y. Zhang, C. Chen, S. Hu, A. Gill, and S. Pan. Large language model watermark stealing with mixed integer programming. arXiv preprint arXiv:2405.19677, 2024. X. Zhao, P. V. Ananth, L. Li, and Y.-X. Wang. Provable robust watermarking for AI-generated text. In The Twelfth International Conference on Learning Representations, 2024a. URL https://openreview.net/ forum?id=SsmT8aO45L. X. Zhao, S. Gunn, M. Christ, J. Fairoze, A. Fabrega, N. Carlini, S. Garg, S. Hong, M. Nasr, F. Tramer, et al. Sok: Watermarking for ai-generated content. arXiv preprint arXiv:2411.18479, 2024b. T. Zhou, X. Zhao, X. Xu, and S. Ren. Bileve: Securing text provenance in large language models against spoofing with bi-level signature. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. URL https://openreview.net/forum?id=vjCFnYTg67.
12