Do Modern Post-Hoc Watermarking Methods Beat Broken-Arrows? Enoal Gesny
Eva Giboulot
Inria Rennes, France
Inria Rennes, France
arXiv:2605.27135v1 [cs.CR] 26 May 2026
Abstract With the rapid proliferation of generative models, such as diffusion models, digital watermarking has emerged as a crucial solution for identifying AI-generated images. Modern post-hoc watermarking schemes use neural networks to achieve an extremely low false-alarm rate while remaining robust to common image transformations. However, there is a lack of comparison between these modern methods and classic ones, particularly in real-world scenarios where robustness and security take precedence over achieving an extremely low false-alarm probability. In this paper, we propose a fair comparison of robustness and security between modern and classic post-hoc watermarking across various types of classic augmentations and recent sophisticated attacks. Our experiments show that, in a realistic scenario, classic watermarking outperforms modern techniques in terms of security while maintaining robustness.
Keywords Watermarking, Adversarial Machine Learning, Information Security ACM Reference Format: Enoal Gesny and Eva Giboulot. 2026. Do Modern Post-Hoc Watermarking Methods Beat Broken-Arrows?. In ACM Workshop on Information Hiding and Multimedia Security (IH&MMSec ’26), June 17–19, 2026, Firenze, Italy. ACM, New York, NY, USA, 6 pages. https://doi.org/10.1145/3785353.3815081
1
Introduction
With recent advancements in diffusion models [33, 35], generative AI now produces high-quality, diverse, and realistic visuals that are indistinguishable from real images. The spread of platforms and services makes this technology accessible to a wide audience. This rise in generated images has prompted regulatory entities to react. The EU AI Act [11], the White House executive order [37], and Chinese AI governance [6] require that AI-generated content be easily identifiable and traceable. Among existing methods, such as metadata [3] and forensics [8], digital watermarking stands out as a key approach to address the issue. This leads to a revival of interest in the design of new watermarking techniques. Classical methods modified Fourier or wavelet coefficients [10, 16] to leverage the theoretical literature on watermarking designs [9, 26]. This approach combined knowledge from signal processing [29], information theory [25], and statistical theory [27] in order to inform their design. The modern approach, on
This work is licensed under a Creative Commons Attribution 4.0 International License. IH&MMSec ’26, Firenze, Italy © 2026 Copyright held by the owner/author(s). ACM ISBN 979-8-4007-2376-6/2026/06 https://doi.org/10.1145/3785353.3815081
the other hand, is now fundamentally based on machine learning, leveraging the flexibility of deep neural networks [2, 13, 42]. The need to trace AI-generated content spurred a line of research specifically dedicated to take advantage of the generation process itself [12, 18, 39]. However, the difficulty of integrating these methods into existing infrastructure has led the industry to favor post-hoc methods [2, 13, 20]. Post-hoc methods embed the watermark in content after it has been generated, meaning they cannot reliably be used for open-source/open-weights workflows. This is of no consequence for proprietary models: the generation pipeline is kept behind an API. In this setting, modern post-hoc methods offer a plug-and-play approach with detection performance that can easily be tuned. What we term the modern approach in this paper can be characterized by three main ingredients: 1) a multi-bit watermarking framework, 2) an embedder and a decoder designed as deep-neural networks, and 3) a training pipeline that explicitly includes robustness and quality considerations by introducing image-processing augmentations as well as psycho-visual masking [7]. The simplicity and flexibility of this recipe allowed the rapid refinement of embedders and decoders since the seminal HiDDeN architecture [42]. Nevertheless, the modern approach vastly increases the attack surface on the watermarking channel. First of all, they inherit the vulnerability of deep-neural networks to adversarial examples [19], making detectors extremely brittle against white and black-box attacks [22]. Secondly, modern detectors lack the concept of the watermarking key. Combined with the first vulnerability, this means that the watermarking channel is fully compromised as soon as the detector is made public, since erasure and copy attacks are then trivial to perform. This is fundamentally a security problem, which is vastly more understood and studied in the classical approach [1, 4, 15]. Another arresting decision with the modern approach is its focus on multi-bit watermarking. In practice, content traceability does not necessitate decoding a message. Indeed, re-framing watermarking as a detection problem – i.e. zero-bit watermarking – is a proven way to tackle the traceability problem [5, 32]. We further discuss the question of multi-bit and zero-bit watermarking in Section 2. For now, it suffices to observe that modern multi-bit decoders can be converted to zero-bit detectors easily – see [18], which demonstrates the excellent performance of such conversion. Stemming from these observations, this work aims to challenge the assumed superiority of the modern approach by asking the following question: Does modern post-hoc watermarking consistently outperform classical approaches in a zero-bit setting? We answer this question through the following contributions: We establish a framework to compare methods and attacks under zero-bit assumption, we provide a fair comparison of modern and
IH&MMSec ’26, June 17–19, 2026, Firenze, Italy
Gesny et al.
traditional watermarking, and we compare recent attacks within the same realistic framework.
2
Zero-Bit Watermarking
Zero-bit watermarking is fundamentally different from multi-bit watermarking: the former is a detection problem, whereas the latter is a communication problem. We frame it as a hypothesis test between two hypotheses: H0 the image is not watermarked, H1 the image is watermarked. Zero-bit watermarking aims to embed a signal in the host such that the power PD of the detector is maximized under a guaranteed, fixed, probability of false-alarm PFA . The watermarked detector is a function 𝜙 : R𝐷 × K → R that takes an image x ∈ R𝐷 and a key 𝑘 ∈ K as input. The test 𝛿𝑘 (x) then decides on either hypothesis by comparing the detector’s output against a fixed threshold 𝜏 ∈ R: H0 𝛿𝑘 (x) := 𝜙 (x, 𝑘) ≶ H 𝜏. 1
(1)
In practice, the distribution of the detector’s output is often known under H0 . It is then simpler to work directly with the 𝑝values 𝑝𝜙 (x, 𝑘) := 𝐹𝜙−1 (𝜙 (x, 𝑘)) ,where 𝐹𝜙−1 is the quantile function associated with the distribution of 𝜙 (x, 𝑘). When 𝐹𝜙 is absolutely continuous, the distribution of the 𝑝values is uniform. Replacing 𝜙 by 𝑝𝜙 in Eq.(1), one guarantees a level 𝛼 ∈ [0, 1] for the test simply by setting 𝜏 = 𝛼.
2.1
The hypercone detector
Let k be a normalized secret vector in a (secret or not) 𝑀-dimensional subspace Sk ⊂ R𝑀 . We can build a detector 𝜙 in two steps. Let x ∈ R𝐷 be an input image and 𝑃k : R𝐷 → Sk the projection to the key subspace. We can compare the extracted vector r := 𝑃k (x) to the secret vector k by computing the angle 𝜃 between the two, 𝑇 𝑐 (r, k) := |r| |r|k|| = cos 𝜃 . Now, if we construct 𝑃k such that the distribution of 𝑃 k (x) is isotropic under H0 , the probability that r falls inside the hypercone of axis k and half-angle 𝜃 is given by: PFA = 1 − 𝐼 cos2 𝜃 (1/2, (𝑀 − 1)/2) .
(2)
This is indeed the probability of false alarm of the test based on the detector 𝜙 hypercone defined as 𝜙 hypercone (x, k) := 𝑐 (𝑃k (x), k) .
2.2
Applications
Broken-Arrows. In the original work [16], the authors use 𝜙 hypercone as the detector. Both the secret vector k as well as the projection 𝑃k are secret. The projection is composed of two steps. The first step performs a (public) wavelet transform on the input and conserves the first 𝑁 𝑓 coefficients. These coefficients are then projected into the secret subspace using 𝑀 secret (pseudo)-orthogonal basis vectors. Within this subspace, 𝑁𝑐 hypercones are defined, using the closest one to the host is selected for embedding: this is the secret vector k. The host vector is then pushed as far as possible from the detection border in order to maximize robustness. Note that due to the use of several hypercones, the PFA is slightly higher than reported in Eq. (2). Using a simple union bound, the PFA for (BA) Broken-Arrows is such that PFA ≤ 𝑁𝑐 PFA .
Modern multi-bit approach. The hypercone detector can be trivially applied to modern multi-bit watermarking by observing that their decoder are all based on projecting an input into a 𝑀 dimensional subspace, 𝑀 being the number of bits in the message. The message is then decoded by thresholding the value of each vector element, usually assigning 0 and 1 to negative and positive elements respectively. By skipping the thresholding step, we obtain the projection 𝑃 k for free. It remains to define the secret vector k. To do so, notice that the original multi-bit schemes embed a message m ∈ {0, 1}𝑀 . We can construct the secret vector by modulating m with antipodal modulation: we map 0s to -1s and 1s to 1s. Finally, we normalize√the modulated message to obtain the final secret vector √ k ∈ {−1/ 𝑀, 1/ 𝑀 }𝑀 . The isotropy assumption must be checked carefully for this approach – see [18][Appendix C] for a generic methodology to enforce this assumption.
3
Evaluation methodology
The recent watermarking reference [1] defines a watermarking system as: “[...] the embedding of a robust, imperceptible and secure information.”. These properties were given thorough definitions in [23]: “Robust watermarking is a mechanism to create a communication channel that is multiplexed into an original content [where] the perceptual degradation of the marked content [...] with respect to the original content is minimal and [where] the capacity of the watermark channel degrades as a smooth function of the degradation of the marked content.” and “Watermark security refers to the inability by unauthorized users to have access to the raw watermarking channel.”. We herein propose a methodology to compare and rank the performance of each watermarking based on these three axes. We motivate our decision in each case, with the goal to design an evaluation which is fair and relevant to a realistic implementation of watermarking systems. A first important choice is the nature of images to use. Experiments are conducted on 1024×1024 natural images to align with the scale-dependent capacity of watermarking and the high-resolution standards of generative models and digital media.
3.1
Detectability/Robustness
We can rank the robustness of watermarking schemes according to two main methods. The first is by computing some statistics on the 𝑝-values of each scheme across different attack scenarios. The second is by setting a detection level 𝛼 guaranteeing a chosen PFA . In practical scenarios, a detection threshold is always fixed, we thus focus on the second method. The hypothesis test is thus simply: H0 𝑝 (x) ≶ H 𝛼. 1 How low should we set the level 𝛼? We can observe a race towards higher capacities in modern watermarking schemes, with TrustMark [2] starting at 100 bits, followed by Videoseal [13] at 256 bits, and most recently ChunkySeal [30] at 1024 bits. When no attack is performed, more capacity usually translates to higher detectability after converting into a zero-bit detector. In a realworld scenario, we argue that guaranteeing extremely low PFA , say 10−100 , under a few image processing attacks is far less valuable than guaranteeing a modest PFA , say 10−6 , across all possible attack scenarios.
Do Modern Post-Hoc Watermarking Methods Beat Broken-Arrows?
IH&MMSec ’26, June 17–19, 2026, Firenze, Italy
Decision. We follow the BOWS-2 competition [16, Section 6] and set 𝛼 = 10−6 so that in expectation, one image out of one million is a false positive.
3.2
Imperceptibility
In this work, we compare the quality between two using the images 1 Í𝐷 standard choice of the PSNR(x, y) = −10·log10 𝐷 𝑖=1 (𝑥𝑖 − 𝑦𝑖 ) 2 . We are aware of the use of perceptual metrics such as the LPIPS [41] and psycho-visual masks (JND) [7] during the training of modern watermarking encoders. However, in order to be fair with respect to detectability, we prefer working at a fixed watermark signal power which makes the PSNR a natural choice in this case. To ensure fairness, we thus fix the PSNR of watermarking signal to a fixed level 𝜂 by scaling the signal after embedding such that: s := 𝜓 (x) − x;
√ 𝜂 xwm = x + s/||s|| 𝐷10 20 .
(3)
Decision. We measure the impact of watermarking on image quality using the PSNR between the original and the watermarked images. We fix 𝜂 =42dB. We measure the impact of an attack using the PSNR between the watermarked and the attacked images.
3.3
Security
We now propose four scenarios where the attacker’s goal is to erase the watermark of the image. More precisely, an attack is deemed successful if the attacked image 𝑝 (xatk ) < 𝛼, that is, if its p-value is below the level 𝛼 of the detector. We present each scenario in decreasing order of knowledge of the attacker about the detection process. White-box scenario. We provide the attacker with all the knowledge pertaining to the detection process. For the modern approach, this corresponds to access to the detector as a white-box. This most closely reflects the standard Kerchoff’s principle, where everything should be public but the key. For the modern approach, we formulate the problem as the optimization of an adversarial perturbation: ( min𝝐 ∈R𝐷 ||𝝐 || 22 (4) s.t. 𝑝 (x + 𝝐) > 𝛼 This is usually solved iteratively, with a fixed model query budget of 𝑄. For this paper, we settled on the boundary projection DDN Attack [34] with gradient. For Broken-Arrows, we also disclose the key to the attacker. Though unfair to Broken-Arrows, this allows us to use closed-form formulas to compute the optimal attack using Eq.(11) and Eq.(12) in [16]. Black-box scenario. The attacker only has access to the detector as a black box: they can query it and obtain a yes/no answer, but the gradient cannot be computed exactly through backpropagation. This attacker still aims to solve the optimization problem in Eq. 4. We settled on CGBA [31], which is currently the most efficient black-box attack for locally linear classifiers. Oracle scenario. Contrary to previous scenarios, the attacker does not have access to the detector. They instead use an oracle 𝑂 : R𝐷 → R which serves as a proxy to the detector’s output. We
selected two attacks as representative of this scenario, with two different approaches. The first, WmForger [36], uses a preference model as an oracle and performs a gradient ascent on the image. In other words at the 𝑖-th iteration, we have: (𝑖 ) (𝑖 −1) (𝑖 −1) xatk = xatk − ∇x (𝑖 −1) 𝑂 xatk . (WmForger) atk
The attack stops after a fixed number of iterations 𝑄. The second attack, Watermark In the Sand (WIS) [40], uses the oracle as a stopping condition. It locally modifies the image at each step and continues until the condition is reached. We found the design in the original paper – based on inpainting and GPT-as-a-judge oracle for quality – computationally intensive and inefficient. We simplified the design for this paper, purifying local patches using a VAE and using the PSNR as an oracle. This resulted in far better attack performance at a fraction of the computational cost – see Appendix A for details about our implementation. Blind scenario. The attacker has no access to the detector and uses no feedback from an oracle. This is the classic robustness scenario in watermarking that tests detectability against sets of classic image processing transformations – e.g. JPEG compression, sharpening, gamma transform, etc. However, recent attacks based on the use of diffusion models and VAEs, such as Purification [28], also fall in this category since the attacker must set the number of steps blindly. For this paper, we made the decision to focus exclusively on valuemetric operations. The reason is that the robustness against geometric operations such as crops, rotations, perspective shifts, etc, can be generically improved by decomposing the image into patches – this is the approach recently taken by SynthID [20] – and synchronization techniques [14][10][Section 9.3]. Since we want to focus on the performance of the baseline watermarking systems at a fixed image size, we forego the comparisons to these operations.
4 Results 4.1 Experimental Settings Dataset. We perform watermark embedding and attacks on 200 natural images from the MFlickr [21] dataset. All images were resized to 1024 × 1024 using bilinear interpolation, with further cropping applied to preserve the aspect ratio. Modern watermarking. We chose Videoseal [13] and TrustMark (without error correcting codes) [2] as modern representatives of post-hoc methods. They were converted to a zero-bit watermarking scheme using the methodology in Section 2. The detectors were whitened using the methodology presented in Appendix C of [18] in order to match the isotropy assumption of the hypercone detector. The dimensions of the watermark subspace for Videoseal and TrustMark are 𝑀 = 256 and 𝑀 = 100 respectively. Classical watermarking. We chose Broken-Arrows [16] as the state-of-the-art of the classical approach. We perform a three levels wavelet decomposition on the GPU using the Pytorch Wavelet Toolbox library1 with Daubechies-9 wavelets. We embed only within the first 𝑁 𝑓 = 60492 coefficients. We set the dimension of the subspace as 𝑀 = 128 and use 𝑁𝑐 = 50 secret hypercones. 1 see https://github.com/v0lta/PyTorch-Wavelet-Toolbox
IH&MMSec ’26, June 17–19, 2026, Firenze, Italy
For all methods, the key is randomized for each image, and the watermarking signal is scaled to achieve a fixed PSNR of 42dB with respect to the original image. White-box scenario. We performed the DDN Attack [34] on modern watermarking techniques with a fixed query budget 𝑄 = 250. For Broken-Arrows, we perform the optimal attack as described in Section 3.3. Black-box scenario. We perform the state-of-the-art CGBA attack [31] for both modern and classic watermarking with a fixed budget of 𝑄 = 2000 queries. We restrict the search space by keeping only the lowest 1.25% DCT frequencies (including DC) – see Section 5.2 in [24]. Preliminary experiments showed an average search overhead of 10 queries to find the adversarial point at each iteration, irrespective of the method. Following the theoretical recommendations in [17], we thus set the number of queries for estimating the gradient to 2. Oracle scenario. We use two novel methods: WmForger [36] and our version of the Watermark In the Sand (WIS) attack as defined in Appendix A. For the latter, the image is downsampled to 512 × 512 before being passed through the VAE. We use the SANA [38]2 . Blind scenario. We use the SANA and Stable-Diffusion-2.13 diffusion models for purification. The Flow-Matching Euler Discrete scheduler and Euler Discrete schedulers are used respectively. For basic image processing valuemetric operations, we only succeeded to attack the detectors with JPEG compression at a quality factor of 5. We reported examples of all used attacks with their residue compare to the watermarked image in figure 1. Metrics. We compute the attack success rate of each attack as Í𝑁
1(𝑝 (atk(x𝑖 ),𝑘𝑖 ) >𝛼 )
𝐴𝑆𝑅 = 𝑖=1 𝜙 𝑁 , where 𝑁 is the number of images 𝑥𝑖 and atk is the evaluated attack. The threshold for attack success is set at 10−6 . We measure the distortion due to the attack using the PSNR with respect to the watermarked image.
4.2
Experimental Results
We report the ASR at a threshold of 𝛼 = 10−6 as a function of PSNR for each scenario and watermarking technique in figure 2. White-box scenario. Broken-Arrows clearly outperforms TrustMark and Videoseal. For the latter, the attack achieves a distortion nearing the sub-quantization limit around 60dB, whereas the true optimal attack for Broken-Arrows achieves 100% only around 40dB. Due to their finite training set, the detection region of DNN-based detectors in pixel space contains many "blind spot" making them vulnerable to small perturbations [19]. The possibility of computing gradients makes these blind spots easy to identify with a gradient projection attack. The way Broken-Arrows detection is built precludes the existence of such adversarial signals Black-box scenario. Both Videoseal and TrustMark are once again far more vulnerable than Broken-Arrows. The gap between the modern and classical methods is surprisingly large: nearly 15dB 2 Huggingface ID: Efficient-Large-Model/Sana_600M_512px_diffusers 3 Huggingface ID: stabilityai/stable-diffusion-2-1-base
Gesny et al.
separates the 100% ASR threshold between the two approaches. The reason for the lack of success against Broken-Arrows is surprising as it fully meets the assumptions of the attack, notably the linearity of the detector. From the results in the white-box scenario, we have an upper-bound on the smallest distortion for a successful attack. Leveraging the theoretical study in [17], we can estimate the expected distortion at each attack iteration depending on the distortion after one iteration. The main insight from this analysis is that if finding a good adversarial example is difficult early on, convergence to the optimum will be slow. Figure 3 presents the PSNR distribution for the initial adversarial points and their state after 2000 queries across both watermarking techniques. The results indicate that initial boundary searches against Broken-Arrows incur a higher distortion cost compared to Videoseal. We hypothesize that Videoseal susceptibility to find a better initial boundary point in the pixel space arises from the inherent geometry of DNN detectors. Oracle scenario. For lower distortion attacks (𝑃𝑆𝑁 𝑅 ≥ 35 dB), Broken-Arrows is the most secure method. If we accept higher distortion, Videoseal becomes better. The figure 4 shows the details of each attack depending on the watermarking methods. It shows that WmForger is the most efficient oracle attack against TrustMark, making it the least secure attack for 𝑃𝑆𝑁 𝑅 ≥ 28 dB. This attack is less efficient against Videoseal and has no effect on Broken-Arrows watermarked images. Our WIS attack has a lot more impact on Broken-Arrows. The VAE purification + downsampling operation results in systematic removal of the mid-frequency wavelet coefficients (𝐿𝐻, 𝐻𝐿, 𝐻𝐻 ), effectively filtering out the watermark signal embedded in those sub-bands. A possible defense is to concentrate the signal in the lowest frequency coefficients – this would further improve robustness, but the resulting wavelet artifacts Blind Scenario. The results in Figure 2 show the best performance of Regeneration, VAE Purification, and JPEG with a quality factor of 5, as it was the only classic value-metric augmentation that had an impact on the watermarking methods. In this scenario, the three methods are robust, with at least 𝑃𝑆𝑁 𝑅 ≤ 30 dB required to erase the watermark. Notably, the Regeneration attack has the same efficiency across all methods. A first difference between the methods lies in the VAE purification, which has a greater impact on BrokenArrows than on modern methods for the same reasons as for WIS. Secondly, TrustMark is not robust against strong JPEG compression, unlike Videoseal and Broken-Arrows. Even if Videoseal stands out as the most robust method in this scenario, the gain compared to Broken-Arrows is marginal. TrustMark is the overall worst choice of the three.
5
Discussion and Conclusion
Our findings reveal a critical trade-off between security and robustness when comparing modern and traditional watermarking. While modern DNN-based methods offer no significant gain in robustness under the tested conditions, they exhibit a significant decline in security. However, modern techniques are primarily optimized for high-capacity multi-bit payloads rather than the zero-bit identification scenario addressed here. Furthermore, this study does not account for geometric distortions, typically handled by modern methods.
Do Modern Post-Hoc Watermarking Methods Beat Broken-Arrows?
IH&MMSec ’26, June 17–19, 2026, Firenze, Italy
Figure 1: Examples of attacks on Videoseal.
0.6 0.4
0.4
0.2
0.2
0.2
0.0 10
0.0 10
0.0 10
20
30
40
50
60
70
20
30
PSNR (dB)
50
60
70
0.4
0.8 0.6
0.2 0.0 10
40
50
60
70
60
0.6 0.4 0.2 0.0 10
70
30
40
30
50
60
70
Figure 2: Comparison of watermarking schemes across scenarios. Each curve is the worst-case attack envelope computed from ASRs of attacks corresponding to the same scenario as a function of the PSNR. A smaller area is better for the watermark.
Figure 3: Distribution of the PSNR of the images attacked after 10 queries and 2000 queries of CGBA against Videoseal and Broken-Arrows.
References [1] Patrick Bas, Teddy Furon, François Cayre, Gwenaël Doërr, and Benjamin Mathon. 2016. Watermarking Security. Springer Singapore.
0.0 10
50
60
70
JPEG WmForger CGBA Optimal Purification VAE WMInTheSand
0.2
PSNR (dB)
40
Broken Arrows
0.4 VideoSeal TrustMark Broken Arrows 20
20
PSNR (dB)
1.0
0.6
0.0 10
PSNR (dB)
50
0.8
0.2
30
40
JPEG WmForger CGBA DDNAttack Purification VAE WMInTheSand
0.8
PSNR (dB)
0.4
20
30
ASR
ASR
0.6
20
White-box
1.0
VideoSeal TrustMark Broken Arrows
0.8
40
PSNR (dB)
Black-box
1.0
ASR
0.6
0.4
TrustMark
1.0 JPEG WmForger CGBA DDNAttack Purification VAE WMInTheSand
0.8
ASR
0.6
Videoseal
1.0
VideoSeal TrustMark Broken Arrows
0.8
ASR
ASR
0.8
Oracle
1.0
VideoSeal TrustMark Broken Arrows
ASR
Blind
1.0
20
30
40
50
60
70
PSNR (dB)
Figure 4: Comparison of attacks on Videoseal, TrustMark, and Broken-Arrows. Each curve represents the ASR as a function of the PSNR for a given attack. A smaller area is better for the watermark.
[2] Tu Bui, Shruti Agarwal, and John Collomosse. 2025. TrustMark: Robust Watermarking and Watermark Removal for Arbitrary Resolution Images. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 18629–18639. [3] C2PA. 2024. C2PA: The Coalition for Content Provenance and Authenticity. https://c2pa.org. [4] François Cayre, Caroline Fontaine, and Teddy Furon. 2005. Watermarking security part I: theory, Vol. 5681. SPIE, 746. https://inria.hal.science/inria-00083329 [5] Vivien Chappelier, Mathieu Desoubeaux, and Jonathan Delhumeau. 2018. Procede d’enregistrement d’un contenu multimedia, procede de detection d’une marque au sein d’un contenu multimedia, dispositifs et programme d’ordinateurs correspondants. [6] China. 2023. Chinese AI Governance Rules. http://www.cac.gov.cn/2023-07/13/ c_1690898327029107.htm. [7] Chun-Hsien Chou and Yun-Chin Li. 1995. A perceptually tuned subband image coder based on the measure of just-noticeable-distortion profile. 5, 6 (1995), 467–476. https://ieeexplore.ieee.org/document/475889 [8] Riccardo Corvi, Davide Cozzolino, Giada Zingarini, Giovanni Poggi, Koki Nagano, and Luisa Verdoliva. 2023. On the detection of synthetic images generated by diffusion models. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1–5. [9] M. Costa. 1983. Writing on dirty paper (Corresp.). 29, 3 (1983), 439–441. https: //ieeexplore.ieee.org/document/1056659 [10] Ingemar J. Cox. 2008. Digital watermarking and steganography (2nd ed ed.). Morgan Kaufmann Publishers, Amsterdam Boston. [11] Europe. 2023. European AI Act. https://artificialintelligenceact.eu/.
IH&MMSec ’26, June 17–19, 2026, Firenze, Italy
[12] Pierre Fernandez, Guillaume Couairon, Hervé Jégou, Matthijs Douze, and Teddy Furon. 2023. The Stable Signature: Rooting Watermarks in Latent Diffusion Models. ICCV (2023). [13] Pierre Fernandez, Hady Elsahar, I. Zeki Yalniz, and Alexandre Mourachko. 2024. Video Seal: Open and Efficient Video Watermarking. [14] Pierre Fernandez, Tomáš Souček, Nikola Jovanović, Hady Elsahar, SylvestreAlvise Rebuffi, Valeriu Lacatusu, Tuan Tran, and Alexandre Mourachko. 2025. Geometric Image Synchronization with Deep Watermarking. [15] Teddy Furon and Patrick Bas. 2012. A New Measure of Watermarking Security Applied on DC-DM QIM. In Information Hiding (Berkeley, United States, 2012-05). TBA. https://hal.science/hal-00702689 [16] Teddy Furon and Patrick Bas. 2008. Broken Arrows. EURASIP Journal on Information Security 2008 (Oct. 2008), ID 597040. https://hal.science/hal-00335311 [17] Enoal Gesny, Eva Giboulot, and Teddy Furon. 2024. When does gradient estimation improve black-box adversarial attacks?. In 2024 IEEE International Workshop on Information Forensics and Security (WIFS) (2024-12). https://ieeexplore.ieee. org/document/10810691/ ISSN: 2157-4774. [18] Enoal Gesny, Eva Giboulot, Teddy Furon, and Vivien Chappelier. 2026. Guidance Watermarking for Diffusion Models. In The Fourteenth International Conference on Learning Representations. https://openreview.net/forum?id=5ifzhjMCKq [19] Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy. 2015. Explaining and Harnessing Adversarial Examples. arXiv:1412.6572 [stat] http://arxiv.org/abs/ 1412.6572 [20] Sven Gowal, Rudy Bunel, Florian Stimberg, David Stutz, Guillermo Ortiz-Jimenez, Christina Kouridi, Mel Vecerik, Jamie Hayes, Sylvestre-Alvise Rebuffi, Paul Bernard, Chris Gamble, Miklós Z. Horváth, Fabian Kaczmarczyck, Alex Kaskasoli, Aleksandar Petrov, Ilia Shumailov, Meghana Thotakuri, Olivia Wiles, Jessica Yung, Zahra Ahmed, Victor Martin, Simon Rosen, Christopher Savčak, Armin Senoner, Nidhi Vyas, and Pushmeet Kohli. 2025. SynthID-Image: Image watermarking at internet scale. arXiv:2510.09263 [cs.CR] https://arxiv.org/abs/2510.09263 [21] Mark J. Huiskes and Michael S. Lew. 2008. The MIR flickr retrieval evaluation. In Proceedings of the 1st ACM International Conference on Multimedia Information Retrieval (Vancouver, British Columbia, Canada) (MIR ’08). Association for Computing Machinery, New York, NY, USA, 39–43. https://doi.org/10.1145/1460096. 1460104 [22] Chloé Imadache, Eva Giboulot, and Teddy Furon. 2025. Evaluating the security of public surrogate watermark detectors. In ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (2025-04). https: //ieeexplore.ieee.org/document/10889821/ ISSN: 2379-190X. [23] T. Kalker. 2001. Considerations on watermarking security. In 2001 IEEE Fourth Workshop on Multimedia Signal Processing (Cat. No.01TH8564) (2001-10). 201–206. https://ieeexplore.ieee.org/document/962734 [24] Thibault Maho, Teddy Furon, and Erwan Le Merrer. 2021. SurFree: a fast surrogatefree black-box attack. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (Nashville, TN, USA, 2021-06). IEEE, 10425–10434. https: //ieeexplore.ieee.org/document/9578850/ [25] Neri Merhav and Erez Sabbag. 2006. Optimal Watermark Embedding and Detection Strategies Under Limited Detection Resources. In 2006 IEEE International Symposium on Information Theory (2006-07). 173–177. arXiv:0705.1919 [cs] http://arxiv.org/abs/0705.1919 [26] M.L. Miller, I.J. Cox, and J.A. Bloom. 2000. Informed embedding: exploiting image and detector information during watermark insertion. In Proceedings 2000 International Conference on Image Processing (Cat. No.00CH37101) (Vancouver, BC, Canada, 2000). IEEE. http://ieeexplore.ieee.org/document/899260/ [27] Matt L. Miller and Jeffrey A. Bloom. 2000. Computing the Probability of False Watermark Detection. In Information Hiding (Berlin, Heidelberg, 2000), Andreas Pfitzmann (Ed.). Springer, 146–158. [28] Weili Nie, Brandon Guo, Yujia Huang, Chaowei Xiao, Arash Vahdat, and Anima Anandkumar. 2022. Diffusion Models for Adversarial Purification. In International Conference on Machine Learning (ICML). [29] Stéphane Pateux and Gaëtan Le Guelvouit. 2003. Practical watermarking scheme based on wide spread spectrum and game theory. 18, 4 (2003), 283–296. https: //www.sciencedirect.com/science/article/pii/S0923596502001455 [30] Aleksandar Petrov, Pierre Fernandez, Tomas Soucek, and Hady Elsahar. 2026. We Can Hide More Bits: The Unused Watermarking Capacity in Theory and in Practice. https://openreview.net/forum?id=Ry8jLSYIUG [31] Md Farhamdur Reza, Ali Rahmati, Tianfu Wu, and Huaiyu Dai. 2023. CGBA: Curvature-aware Geometric Black-box Attack. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 124–133. [32] Geoffrey B. Rhoads. 2010. Detecting embedded signals in media content using coincidence metrics. [33] Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-Resolution Image Synthesis With Latent Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 10684–10695. [34] Jérôme Rony, Luiz Gustavo Hafemann, Luiz Soares de Oliveira, Ismail Ben Ayed, and Eric Granger. 2019. Decoupling Direction and Norm for Efficient GradientBased L2 Adversarial Attacks and Defenses. 4317–4325.
Gesny et al.
Algorithm 1 Watermark In The Sand (WIS) Attack Require: Image x, VAE (E, D), Oracle O, Threshold 𝛽 Ensure: Adversarial Image x𝐴 1: x0 ← x 2: 𝑖 ← 0 // Phase 1: Iterative VAE Purification 3: while O (x𝑖 ) > 𝛽 do 4: x𝑑𝑜𝑤𝑛 ← Downsample(x𝑖 , 𝑠) 5: z ← E (x𝑑𝑜𝑤𝑛 ) ⊲ Latent encoding 6: x̂𝑑𝑜𝑤𝑛 ← D (z) ⊲ Reconstruction 7: x𝑖+1 ← Upsample( x̂𝑑𝑜𝑤𝑛 , shape(x)) 8: 𝑖 ←𝑖 +1 9: end while // Phase 2: Patch-based refinement 10: x𝐴 , x𝐵 ← x𝑖 −1 , x𝑖 11: while O (x𝐴 ) > 𝛽 do 12: 𝑚 ← SelectPatchMask() 13: x𝐴 ← (1 − 𝑚) ⊙ x𝐴 + 𝑚 ⊙ x𝐵 ⊲ Patch replacement 14: end while 15: return x𝐴
[35] Jiaming Song, Chenlin Meng, and Stefano Ermon. 2022. Denoising Diffusion Implicit Models. arXiv:2010.02502 [cs.LG] https://arxiv.org/abs/2010.02502 [36] Tomas Soucek, Sylvestre-Alvise Rebuffi, Pierre Fernandez, Nikola Jovanović, Hady Elsahar, Valeriu Lacatusu, Tuan A. Tran, and Alexandre Mourachko. 2025. Transferable Black-Box One-Shot Forging of Watermarks via Image Preference Models. In The Thirty-ninth Annual Conference on Neural Information Processing Systems. https://openreview.net/forum?id=yb5JOOmfxA Ensuring Safe, Secure, and Trustworthy AI. https: [37] USA. 2023. //www.whitehouse.gov/wp-content/uploads/2023/07/Ensuring-Safe-Secureand-Trustworthy-AI.pdf. [38] Enze Xie, Junsong Chen, Junyu Chen, Han Cai, Haotian Tang, Yujun Lin, Zhekai Zhang, Muyang Li, Ligeng Zhu, Yao Lu, and Song Han. 2024. Sana: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer. arXiv:2410.10629 [cs.CV] https://arxiv.org/abs/2410.10629 [39] Zijin Yang, Kai Zeng, Kejiang Chen, Han Fang, Weiming Zhang, and Nenghai Yu. 2024. Gaussian Shading: Provable Performance-Lossless Image Watermarking for Diffusion Models. arXiv preprint arXiv:2404.04956 (2024). [40] Hanlin Zhang, Benjamin L. Edelman, Danilo Francati, Daniele Venturi, Giuseppe Ateniese, and Boaz Barak. 2024. Watermarks in the sand: impossibility of strong watermarking for language models. In Proceedings of the 41st International Conference on Machine Learning (Vienna, Austria) (ICML’24). JMLR.org, Article 2429. [41] Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shechtman, and Oliver Wang. 2018. The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. arXiv:1801.03924 [cs] http://arxiv.org/abs/1801.03924 [42] Jiren Zhu, Russell Kaplan, Justin Johnson, and Li Fei-Fei. 2018. HiDDeN: Hiding Data With Deep Networks. In Computer Vision – ECCV 2018, Vittorio Ferrari, Martial Hebert, Cristian Sminchisescu, and Yair Weiss (Eds.). Vol. 11219. Springer International Publishing, Cham, 682–697. https://link.springer.com/10.1007/9783-030-01267-0_40 Series Title: Lecture Notes in Computer Science.
A
Our Watermark in The Sand
We developed a custom implementation of the WIS attack, adapting the original concept to improve visual fidelity and efficiency. The original WIS method relies on an iterative patch-inpainting process governed by a GPT-as-a-Judge oracle. This approach often introduces significant semantic distortions due to the inpainting and the stochastic nature of this type of oracle. To mitigate these issues, our implementation introduces two key modifications. First, we substitute the oracle with a distortion-based metric relative to the original image. Second, we employ a VAE-purification process at the patch level. The complete procedure is formalized in Algorithm 1.