ConceptioArchivearXiv CS
arXiv CSopen access

On the Robustness of Watermarking for Autoregressive Image Generation

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
cryptography, security, privacy, cybersecurity

On the Robustness of Watermarking for Autoregressive Image Generation Andreas Müller1 , Denis Lukovnikov1 , Shingo Kodama2 , Minh Pham3 , Anubhav Jain4⋆ , Jonathan Petit5 , Niv Cohen3 , and Asja Fischer1

arXiv:2604.11720v1 [cs.CV] 13 Apr 2026

1

Ruhr University Bochum 2 Middlebury College 3 New York University 4 Independent Researcher 5 Qualcomm

Abstract. The proliferation of autoregressive (AR) image generators demands reliable detection and attribution of their outputs to mitigate misinformation, and to filter synthetic images from training data to prevent model collapse. To address this need, watermarking techniques, specifically designed for AR models, embed a subtle signal at generation time, enabling downstream verification through a corresponding watermark detector. In this work, we study these schemes and demonstrate their vulnerability to both watermark removal and forgery attacks. We assess existing attacks and further introduce three new attacks: (i) a vector-quantized regeneration removal attack, (ii) adversarial optimization–based attack, and (iii) a frequency injection attack. Our evaluation reveals that removal and forgery attacks can be effective with access to a single watermarked reference image and without access to original model parameters or watermarking secrets. Our findings indicate that existing watermarking schemes for AR image generation do not reliably support synthetic content detection for dataset filtering. Moreover, they enable Watermark Mimicry, whereby authentic images can be manipulated to imitate a generator’s watermark and trigger false detection to prevent their inclusion in future model training.

1

Introduction

Autoregressive (AR) image generation has emerged as a powerful paradigm for visual synthesis, with recent models such as VAR [36], HMAR [21], Infinity [13], and other proprietary closed source models demonstrating remarkable capabilities. Unlike diffusion-based approaches, these models generate images by predicting discrete visual tokens in an autoregressive manner, enabling efficient sampling and fine-grained control over the generation process. With the growing adoption of autoregressive image generators, provenance and misuse have become pressing concerns, leading to several in-generation watermarking methods tailored to these models [16, 18, 26, 37]. These watermarking techniques embed ⋆

Now at Apple

2

A. Müller et al. BitMark for Infinity Authentic Image

Spoofed Signal

+ <latexit sha1_base64="7CDz+hFii/hnzm/SPcG6JVj1JjA=">AAAB6HicbVDLSgNBEOyNrxhfUY9eBoMgCGFXJHoMevGYgHlAsoTZSW8yZnZ2mZkVQsgXePGgiFc/yZt/4yTZgyYWNBRV3XR3BYng2rjut5NbW9/Y3MpvF3Z29/YPiodHTR2nimGDxSJW7YBqFFxiw3AjsJ0opFEgsBWM7mZ+6wmV5rF8MOME/YgOJA85o8ZK9YteseSW3TnIKvEyUoIMtV7xq9uPWRqhNExQrTuemxh/QpXhTOC00E01JpSN6AA7lkoaofYn80On5MwqfRLGypY0ZK7+npjQSOtxFNjOiJqhXvZm4n9eJzXhjT/hMkkNSrZYFKaCmJjMviZ9rpAZMbaEMsXtrYQNqaLM2GwKNgRv+eVV0rwse5VypX5Vqt5mceThBE7hHDy4hircQw0awADhGV7hzXl0Xpx352PRmnOymWP4A+fzB3UXjLo=</latexit>

BitMark Forgery

= <latexit sha1_base64="dfa12l/r87/D93uS51ywGerJGEM=">AAAB6HicbVDLSgNBEOyNrxhfUY9eBoPgKeyKRC9C0IvHBMwDkiXMTnqTMbOzy8ysEEK+wIsHRbz6Sd78GyfJHjSxoKGo6qa7K0gE18Z1v53c2vrG5lZ+u7Czu7d/UDw8auo4VQwbLBaxagdUo+ASG4Ybge1EIY0Cga1gdDfzW0+oNI/lgxkn6Ed0IHnIGTVWqt/0iiW37M5BVomXkRJkqPWKX91+zNIIpWGCat3x3MT4E6oMZwKnhW6qMaFsRAfYsVTSCLU/mR86JWdW6ZMwVrakIXP198SERlqPo8B2RtQM9bI3E//zOqkJr/0Jl0lqULLFojAVxMRk9jXpc4XMiLEllClubyVsSBVlxmZTsCF4yy+vkuZF2auUK/XLUvU2iyMPJ3AK5+DBFVThHmrQAAYIz/AKb86j8+K8Ox+L1pyTzRzDHzifP5BfjMw=</latexit>

high definition photo of a poisonous frog

Is generated! Do not consume!

Service provider recognizes AI content online by their watermark and Þlters training data

Is generated! Do not consume!

Authentic images are protected from being consumed by mimicking the watermark

Fig. 1: Watermark Mimicry subverts synthetic content filtering. An image generated with the radioactive BitMark watermark (left) carries an embedded signal that allows service providers to identify synthetic content and exclude it from future training to prevent model collapse. However, other parties can transfer this watermark onto unrelated images, causing the detector to mistake authentic images (here, Ranitomeya imitator - the mimic poison frog) as AI-generated (right). This adversarial strategy inverts the protective intent of radioactive watermarking: rather than preventing real data from being mislabeled, one may exploit the detector’s own decision boundary to shield genuine content from being harvested for training. Photo by andes2amazonexpeditions is licensed under CC BY-NC 4.0. Sourced from iNaturalist.

a signal during the generation process, making them inherently more resistant to post-hoc attacks compared to post-processing watermarking schemes [40, 46]. Typically, they embed a signal by slightly shifting the probabilities of the tokens selected during generation. The resulting watermarks exhibit high robustness against common transformation such as compression, additive noise, color jitter, as well as existing regeneration attacks using diffusion models [47]. However, targeted attacks specifically designed to exploit the structure of autoregressive image generators and their watermarking mechanisms remain unexplored. This gap in understanding poses significant risks, as sophisticated adversaries could potentially circumvent watermark detection or falsely claim a non-watermarked image to be watermarked while maintaining perceptual image quality. The latter point is of particular interest in the context of recent radioactive watermarking techniques, specifically BitMark for bitwise autoregressive image generation, which are intended as a means of filtering synthetic content from training data to prevent performance degradation, i.e. model collapse [1, 31]: By triggering false detection via forgery attacks, third parties are able to perform Watermark Mimicry, where authentic images can be protected from being harvested for model training (see Fig. 1). Additionally, as pointed out by earlier work on forgery attacks for diffusion model watermarks [15, 27], watermark forgery can be used to discredit evidence. These risks motivate our investigation of forgery attacks on watermarks for AR image generators.

On the Robustness of Watermarking for Autoregressive Image Generation

3

Contributions. (i) We demonstrate that recent in-generation watermarks for autoregressive image generation can, in most cases, be successfully removed with limited impact on image quality, requiring neither access to the original generator model nor knowledge of the secret watermarking parameters. Moreover, results show that forgery of watermarks from just one reference image onto unrelated images is also possible. (ii) We propose three new watermark attacks. (iii) We show that attacks cannot be mitigated solely by adjusting decision boundaries.

2

Background

2.1

Autoregressive Image Generation

Recent works [3,4,8,12,13,13,21,22,34,36,38,39,44] have developed AR methods for generating images, using different sampling approaches. We first look at the common approach of decoding discrete tokens from a latent codebook. Since generating high-resolution images at pixel-level is prohibitively costly, modern AR methods instead decode images in the latent space, where a high-resolution image x ∈ R3×H×W is represented by a low-resolution latent z ∈ Rd×h×w , and subsequently map z back to the pixel space using a decoder D. In order to produce vectors in latent space, earlier AR image generators vector-quantize the latent space, building a vocabulary V of visual tokens q, each corresponding to a d-dimensional vector in the latent space. The mapping C : V → Rd is called the codebook C. An AR image generator decomposes the joint Qh·w distribution of visual tokens p(q1 , q2 , . . . , qh·w ) as the product i pθ (qi |q<i , c). Here, the distribution pθ over the vocabulary V is obtained from the logits computed by a neural network fθ using a softmax layer: pθ (·) = softmax ◦ fθ (·). Generation then proceeds by sampling a token from this categorical distribution, conditioned on previously produced tokens [33, 35, 44]: q̂t ∼ softmax(fθ (q̂<t )). In this work, we also study attacks on the recent watermark scheme BitMark [18] that was developed for the multi-scale bitwise autoregressive Infinity generator [13]. Rather than generating latent tokens one-by-one, Infinity samples the latent z by generating a sequence of residuals at different scales, using the next scale prediction approach (VAR [36]). In contrast to regular left-toright token-level autoregressive decoding, all elements in one scale are sampled independently from each other, conditioned on all previous scales. Where VAR predicts tokens from a fixed vocabulary, Infinity predicts residuals for every channel of the latent independently, and the residuals are quantized to two levels (one bit) per channel. Alternatively, this can be thought of as the prediction of bits that point to a token from a very large codebook of 2d tokens. 2.2

KGW Watermarking

Most of the watermarking schemes for AR image generators are derived from KGW watermarking originally proposed for large language models [19,20]. KGW takes the following approach: First, at every generation step, the previously

4

A. Müller et al. Verify

Generate Determine green token set

Count green tokens & test

T6 É <latexit sha1_base64="cwx7JxB20iz98SKvoa2XDs34Pfw=">AAAB6nicbVBNS8NAEJ3Ur1q/qh69LBbBU0lEqseiF48VW1toQ9lsJ+3SzSbsboQS+hO8eFDEq7/Im//GbZuDtj4YeLw3w8y8IBFcG9f9dgpr6xubW8Xt0s7u3v5B+fDoUcepYthisYhVJ6AaBZfYMtwI7CQKaRQIbAfj25nffkKleSybZpKgH9Gh5CFn1Fjpodmv9csVt+rOQVaJl5MK5Gj0y1+9QczSCKVhgmrd9dzE+BlVhjOB01Iv1ZhQNqZD7FoqaYTaz+anTsmZVQYkjJUtachc/T2R0UjrSRTYzoiakV72ZuJ/Xjc14bWfcZmkBiVbLApTQUxMZn+TAVfIjJhYQpni9lbCRlRRZmw6JRuCt/zyKnm8qHq1au3+slK/yeMowgmcwjl4cAV1uIMGtIDBEJ7hFd4c4bw4787HorXg5DPH8AfO5w/f742M</latexit>

É É É É

T6 T1 T19 T144 <latexit sha1_base64="ZVFuL/frNJY8X2Q2y99GZc4Lq+Y=">AAAB6nicbVBNS8NAEJ3Ur1q/qh69LBbBU0lEqseiF48VW1toQ9lsJ+3SzSbsboQS+hO8eFDEq7/Im//GbZuDtj4YeLw3w8y8IBFcG9f9dgpr6xubW8Xt0s7u3v5B+fDoUcepYthisYhVJ6AaBZfYMtwI7CQKaRQIbAfj25nffkKleSybZpKgH9Gh5CFn1Fjpodn3+uWKW3XnIKvEy0kFcjT65a/eIGZphNIwQbXuem5i/Iwqw5nAaamXakwoG9Mhdi2VNELtZ/NTp+TMKgMSxsqWNGSu/p7IaKT1JApsZ0TNSC97M/E/r5ua8NrPuExSg5ItFoWpICYms7/JgCtkRkwsoUxxeythI6ooMzadkg3BW355lTxeVL1atXZ/Wanf5HEU4QRO4Rw8uII63EEDWsBgCM/wCm+OcF6cd+dj0Vpw8plj+APn8wfYW42H</latexit>

<latexit sha1_base64="esrYUA0zhResIRRfLnHuJcX0uMU=">AAAB7XicbVBNSwMxEJ2tX7V+VT16CRbBU9kVqXorevFYoV/QLiWbZtvYbLIkWaEs/Q9ePCji1f/jzX9jtt2Dtj4YeLw3w8y8IOZMG9f9dgpr6xubW8Xt0s7u3v5B+fCorWWiCG0RyaXqBlhTzgRtGWY47caK4ijgtBNM7jK/80SVZlI0zTSmfoRHgoWMYGOldnOQejezQbniVt050CrxclKBHI1B+as/lCSJqDCEY617nhsbP8XKMMLprNRPNI0xmeAR7VkqcES1n86vnaEzqwxRKJUtYdBc/T2R4kjraRTYzgibsV72MvE/r5eY8NpPmYgTQwVZLAoTjoxE2etoyBQlhk8twUQxeysiY6wwMTagkg3BW355lbQvql6tWnu4rNRv8ziKcAKncA4eXEEd7qEBLSDwCM/wCm+OdF6cd+dj0Vpw8plj+APn8wcZr47W</latexit>

<latexit sha1_base64="cwx7JxB20iz98SKvoa2XDs34Pfw=">AAAB6nicbVBNS8NAEJ3Ur1q/qh69LBbBU0lEqseiF48VW1toQ9lsJ+3SzSbsboQS+hO8eFDEq7/Im//GbZuDtj4YeLw3w8y8IBFcG9f9dgpr6xubW8Xt0s7u3v5B+fDoUcepYthisYhVJ6AaBZfYMtwI7CQKaRQIbAfj25nffkKleSybZpKgH9Gh5CFn1Fjpodmv9csVt+rOQVaJl5MK5Gj0y1+9QczSCKVhgmrd9dzE+BlVhjOB01Iv1ZhQNqZD7FoqaYTaz+anTsmZVQYkjJUtachc/T2R0UjrSRTYzoiakV72ZuJ/Xjc14bWfcZmkBiVbLApTQUxMZn+TAVfIjJhYQpni9lbCRlRRZmw6JRuCt/zyKnm8qHq1au3+slK/yeMowgmcwjl4cAV1uIMGtIDBEJ7hFd4c4bw4787HorXg5DPH8AfO5w/f742M</latexit>

<latexit sha1_base64="VbREsMKlF8oXO9blCj4TDV1yPTM=">AAAB7nicbVBNS8NAEJ34WetX1aOXxSJ4KomU6rHoxWOFfkEbyma7aZduNmF3IpTQH+HFgyJe/T3e/Ddu2xy09cHA470ZZuYFiRQGXffb2djc2t7ZLewV9w8Oj45LJ6dtE6ea8RaLZay7ATVcCsVbKFDybqI5jQLJO8Hkfu53nrg2IlZNnCbcj+hIiVAwilbqNAeZV63OBqWyW3EXIOvEy0kZcjQGpa/+MGZpxBUySY3peW6CfkY1Cib5rNhPDU8om9AR71mqaMSNny3OnZFLqwxJGGtbCslC/T2R0ciYaRTYzoji2Kx6c/E/r5dieOtnQiUpcsWWi8JUEozJ/HcyFJozlFNLKNPC3krYmGrK0CZUtCF4qy+vk/Z1xatVao/Vcv0uj6MA53ABV+DBDdThARrQAgYTeIZXeHMS58V5dz6WrRtOPnMGf+B8/gCHm48P</latexit>

Encode & Quantize

É É É É

<latexit sha1_base64="H4YV6LvZLtPqryGzlcFL699u7Ts=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSQi1WPRi8cK9gPaUDabTbt2sxt2J0Ip/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemApu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVpyppUCaU7ITFMcMmayFGwTqoZSULB2uHodua3n5g2XMkHHKcsSMhA8phTglZq9SImkPTLFa/qzeGuEj8nFcjR6Je/epGiWcIkUkGM6fpeisGEaORUsGmplxmWEjoiA9a1VJKEmWAyv3bqnlklcmOlbUl05+rviQlJjBknoe1MCA7NsjcT//O6GcbXwYTLNEMm6WJRnAkXlTt73Y24ZhTF2BJCNbe3unRINKFoAyrZEPzll1dJ66Lq16q1+8tK/SaPowgncArn4MMV1OEOGtAECo/wDK/w5ijnxXl3PhatBSefOYY/cD5/AJUejyc=</latexit>

Boost green token logits by

Embed & Decode

G(T6 ) = {T1 , T3 , ...} <latexit sha1_base64="L2a8Q269DHeufxemiL/YHuEkyng=">AAACB3icbVDLSsNAFJ3UV62vqEtBBotQQUKiUt0IRRe6rNAXNCFMptN26GQSZiZCCdm58VfcuFDErb/gzr9x2mah1QPDPZxzL3fuCWJGpbLtL6OwsLi0vFJcLa2tb2xumds7LRklApMmjlgkOgGShFFOmooqRjqxICgMGGkHo+uJ374nQtKIN9Q4Jl6IBpz2KUZKS765f1Np+NUjeAndtOGnTnYMdTnVxbIsN/PNsm3ZU8C/xMlJGeSo++an24twEhKuMENSdh07Vl6KhKKYkazkJpLECI/QgHQ15Sgk0kund2TwUCs92I+EflzBqfpzIkWhlOMw0J0hUkM5703E/7xuovoXXkp5nCjC8WxRP2FQRXASCuxRQbBiY00QFlT/FeIhEggrHV1Jh+DMn/yXtE4sp2pV787Ktas8jiLYAwegAhxwDmrgFtRBE2DwAJ7AC3g1Ho1n4814n7UWjHxmF/yC8fENSWOWbA==</latexit>

T0 T67 <latexit sha1_base64="4P5UQRgOYdoVZjyzN25v+T5nAgQ=">AAAB6nicbVBNS8NAEJ3Ur1q/qh69LBbBU0lEqseiF48VW1toQ9lsJ+3SzSbsboQS+hO8eFDEq7/Im//GbZuDtj4YeLw3w8y8IBFcG9f9dgpr6xubW8Xt0s7u3v5B+fDoUcepYthisYhVJ6AaBZfYMtwI7CQKaRQIbAfj25nffkKleSybZpKgH9Gh5CFn1Fjpodl3++WKW3XnIKvEy0kFcjT65a/eIGZphNIwQbXuem5i/Iwqw5nAaamXakwoG9Mhdi2VNELtZ/NTp+TMKgMSxsqWNGSu/p7IaKT1JApsZ0TNSC97M/E/r5ua8NrPuExSg5ItFoWpICYms7/JgCtkRkwsoUxxeythI6ooMzadkg3BW355lTxeVL1atXZ/Wanf5HEU4QRO4Rw8uII63EEDWsBgCM/wCm+OcF6cd+dj0Vpw8plj+APn8wfW142G</latexit>

<latexit sha1_base64="xrwcvfYbSX2XM7XShT88/ZChphg=">AAAB7XicbVBNSwMxEJ2tX7V+VT16CRbBU9kV2XosevFYoV/QLiWbZtvYbLIkWaEs/Q9ePCji1f/jzX9j2u5BWx8MPN6bYWZemHCmjet+O4WNza3tneJuaW//4PCofHzS1jJVhLaI5FJ1Q6wpZ4K2DDOcdhNFcRxy2gknd3O/80SVZlI0zTShQYxHgkWMYGOldnOQ+bXZoFxxq+4CaJ14OalAjsag/NUfSpLGVBjCsdY9z01MkGFlGOF0VuqnmiaYTPCI9iwVOKY6yBbXztCFVYYoksqWMGih/p7IcKz1NA5tZ4zNWK96c/E/r5ea6CbImEhSQwVZLopSjoxE89fRkClKDJ9agoli9lZExlhhYmxAJRuCt/ryOmlfVT2/6j9cV+q3eRxFOINzuAQPalCHe2hACwg8wjO8wpsjnRfn3flYthacfOYU/sD5/AEeQ47Z</latexit>

G(T6 ) 3 T1 +1 G(T1 ) 3 T19 +1 G(T19 ) 3 T144+1 <latexit sha1_base64="8hD/pR6Bgmu++73pgHjb7Xz7IZk=">AAAB73icbVBNSwMxEJ3Ur1q/qh69BItQL2VXpHosetBjhX5Bu5Rsmm1Ds9k1yQpl6Z/w4kERr/4db/4b03YP2vpg4PHeDDPz/FhwbRznG+XW1jc2t/LbhZ3dvf2D4uFRS0eJoqxJIxGpjk80E1yypuFGsE6sGAl9wdr++Hbmt5+Y0jySDTOJmReSoeQBp8RYqXNXbvTT6vS8Xyw5FWcOvErcjJQgQ71f/OoNIpqETBoqiNZd14mNlxJlOBVsWuglmsWEjsmQdS2VJGTaS+f3TvGZVQY4iJQtafBc/T2RklDrSejbzpCYkV72ZuJ/XjcxwbWXchknhkm6WBQkApsIz57HA64YNWJiCaGK21sxHRFFqLERFWwI7vLLq6R1UXGrlerDZal2k8WRhxM4hTK4cAU1uIc6NIGCgGd4hTf0iF7QO/pYtOZQNnMMf4A+fwD+TI9O</latexit>

<latexit sha1_base64="qWfOeH2KzF+zAKYdV+SM7Kn0k6g=">AAAB7HicbVBNS8NAEJ3Ur1q/qh69LBbBU0lEqseiF48VmrbQhrLZbtulm03YnQgl9Dd48aCIV3+QN/+N2zYHbX0w8Hhvhpl5YSKFQdf9dgobm1vbO8Xd0t7+weFR+fikZeJUM+6zWMa6E1LDpVDcR4GSdxLNaRRK3g4n93O//cS1EbFq4jThQURHSgwFo2glv9nPvFm/XHGr7gJknXg5qUCORr/81RvELI24QiapMV3PTTDIqEbBJJ+VeqnhCWUTOuJdSxWNuAmyxbEzcmGVARnG2pZCslB/T2Q0MmYahbYzojg2q95c/M/rpji8DTKhkhS5YstFw1QSjMn8czIQmjOUU0so08LeStiYasrQ5lOyIXirL6+T1lXVq1Vrj9eV+l0eRxHO4BwuwYMbqMMDNMAHBgKe4RXeHOW8OO/Ox7K14OQzp/AHzucPnLiOkw==</latexit>

<latexit sha1_base64="Mx4dJlBZd0WdhYo0BwxmB7UE4GY=">AAAB6nicbVBNS8NAEJ3Ur1q/qh69LBbBU0lEqseiF48V7Qe0oWy2k3bpZhN2N0IJ/QlePCji1V/kzX/jts1BWx8MPN6bYWZekAiujet+O4W19Y3NreJ2aWd3b/+gfHjU0nGqGDZZLGLVCahGwSU2DTcCO4lCGgUC28H4dua3n1BpHstHM0nQj+hQ8pAzaqz00JO8X664VXcOskq8nFQgR6Nf/uoNYpZGKA0TVOuu5ybGz6gynAmclnqpxoSyMR1i11JJI9R+Nj91Ss6sMiBhrGxJQ+bq74mMRlpPosB2RtSM9LI3E//zuqkJr/2MyyQ1KNliUZgKYmIy+5sMuEJmxMQSyhS3txI2oooyY9Mp2RC85ZdXSeui6tWqtfvLSv0mj6MIJ3AK5+DBFdThDhrQBAZDeIZXeHOE8+K8Ox+L1oKTzxzDHzifP1BFjdY=</latexit>

<latexit sha1_base64="k+c/6ceNrnh9MiSpT5TH840RKZM=">AAAB73icbVBNSwMxEJ3Ur1q/qh69BItQL2VXpHosetBjhX5Bu5Rsmm1Ds9k1yQpl6Z/w4kERr/4db/4b03YP2vpg4PHeDDPz/FhwbRznG+XW1jc2t/LbhZ3dvf2D4uFRS0eJoqxJIxGpjk80E1yypuFGsE6sGAl9wdr++Hbmt5+Y0jySDTOJmReSoeQBp8RYqXNXbvRTd3reL5acijMHXiVuRkqQod4vfvUGEU1CJg0VROuu68TGS4kynAo2LfQSzWJCx2TIupZKEjLtpfN7p/jMKgMcRMqWNHiu/p5ISaj1JPRtZ0jMSC97M/E/r5uY4NpLuYwTwyRdLAoSgU2EZ8/jAVeMGjGxhFDF7a2Yjogi1NiICjYEd/nlVdK6qLjVSvXhslS7yeLIwwmcQhlcuIIa3EMdmkBBwDO8wht6RC/oHX0sWnMomzmGP0CfP/auj0k=</latexit>

<latexit sha1_base64="esrYUA0zhResIRRfLnHuJcX0uMU=">AAAB7XicbVBNSwMxEJ2tX7V+VT16CRbBU9kVqXorevFYoV/QLiWbZtvYbLIkWaEs/Q9ePCji1f/jzX9jtt2Dtj4YeLw3w8y8IOZMG9f9dgpr6xubW8Xt0s7u3v5B+fCorWWiCG0RyaXqBlhTzgRtGWY47caK4ijgtBNM7jK/80SVZlI0zTSmfoRHgoWMYGOldnOQejezQbniVt050CrxclKBHI1B+as/lCSJqDCEY617nhsbP8XKMMLprNRPNI0xmeAR7VkqcES1n86vnaEzqwxRKJUtYdBc/T2R4kjraRTYzgibsV72MvE/r5eY8NpPmYgTQwVZLAoTjoxE2etoyBQlhk8twUQxeysiY6wwMTagkg3BW355lbQvql6tWnu4rNRv8ziKcAKncA4eXEEd7qEBLSDwCM/wCm+OdF6cd+dj0Vpw8plj+APn8wcZr47W</latexit>

T3 T19 T42 T6 <latexit sha1_base64="MoFCcybhDUclq3mWXoQLkga09pc=">AAAB6nicbVDLSgNBEOyNrxhfUY9eBoPgKeyqRI9BLx4j5gXJEmYnvcmQ2dllZlYIIZ/gxYMiXv0ib/6Nk2QPmljQUFR1090VJIJr47rfTm5tfWNzK79d2Nnd2z8oHh41dZwqhg0Wi1i1A6pRcIkNw43AdqKQRoHAVjC6m/mtJ1Sax7Juxgn6ER1IHnJGjZUe673LXrHklt05yCrxMlKCDLVe8avbj1kaoTRMUK07npsYf0KV4UzgtNBNNSaUjegAO5ZKGqH2J/NTp+TMKn0SxsqWNGSu/p6Y0EjrcRTYzoiaoV72ZuJ/Xic14Y0/4TJJDUq2WBSmgpiYzP4mfa6QGTG2hDLF7a2EDamizNh0CjYEb/nlVdK8KHuVcuXhqlS9zeLIwwmcwjl4cA1VuIcaNIDBAJ7hFd4c4bw4787HojXnZDPH8AfO5w/bY42J</latexit>

<latexit sha1_base64="esrYUA0zhResIRRfLnHuJcX0uMU=">AAAB7XicbVBNSwMxEJ2tX7V+VT16CRbBU9kVqXorevFYoV/QLiWbZtvYbLIkWaEs/Q9ePCji1f/jzX9jtt2Dtj4YeLw3w8y8IOZMG9f9dgpr6xubW8Xt0s7u3v5B+fCorWWiCG0RyaXqBlhTzgRtGWY47caK4ijgtBNM7jK/80SVZlI0zTSmfoRHgoWMYGOldnOQejezQbniVt050CrxclKBHI1B+as/lCSJqDCEY617nhsbP8XKMMLprNRPNI0xmeAR7VkqcES1n86vnaEzqwxRKJUtYdBc/T2R4kjraRTYzgibsV72MvE/r5eY8NpPmYgTQwVZLAoTjoxE2etoyBQlhk8twUQxeysiY6wwMTagkg3BW355lbQvql6tWnu4rNRv8ziKcAKncA4eXEEd7qEBLSDwCM/wCm+OdF6cd+dj0Vpw8plj+APn8wcZr47W</latexit>

<latexit sha1_base64="YBypGNs0ww79T4IikLMeQlU46QY=">AAAB7XicbVBNSwMxEJ2tX7V+VT16CRbBU9ktpfVY9OKxQr+gXUo2zbax2WRJskJZ+h+8eFDEq//Hm//GtN2Dtj4YeLw3w8y8IOZMG9f9dnJb2zu7e/n9wsHh0fFJ8fSso2WiCG0TyaXqBVhTzgRtG2Y47cWK4ijgtBtM7xZ+94kqzaRomVlM/QiPBQsZwcZKndYwrVbmw2LJLbtLoE3iZaQEGZrD4tdgJEkSUWEIx1r3PTc2foqVYYTTeWGQaBpjMsVj2rdU4IhqP11eO0dXVhmhUCpbwqCl+nsixZHWsyiwnRE2E73uLcT/vH5iwhs/ZSJODBVktShMODISLV5HI6YoMXxmCSaK2VsRmWCFibEBFWwI3vrLm6RTKXu1cu2hWmrcZnHk4QIu4Ro8qEMD7qEJbSDwCM/wCm+OdF6cd+dj1Zpzsplz+APn8wcTno7S</latexit>

<latexit sha1_base64="cwx7JxB20iz98SKvoa2XDs34Pfw=">AAAB6nicbVBNS8NAEJ3Ur1q/qh69LBbBU0lEqseiF48VW1toQ9lsJ+3SzSbsboQS+hO8eFDEq7/Im//GbZuDtj4YeLw3w8y8IBFcG9f9dgpr6xubW8Xt0s7u3v5B+fDoUcepYthisYhVJ6AaBZfYMtwI7CQKaRQIbAfj25nffkKleSybZpKgH9Gh5CFn1Fjpodmv9csVt+rOQVaJl5MK5Gj0y1+9QczSCKVhgmrd9dzE+BlVhjOB01Iv1ZhQNqZD7FoqaYTaz+anTsmZVQYkjJUtachc/T2R0UjrSRTYzoiakV72ZuJ/Xjc14bWfcZmkBiVbLApTQUxMZn+TAVfIjJhYQpni9lbCRlRRZmw6JRuCt/zyKnm8qHq1au3+slK/yeMowgmcwjl4cAV1uIMGtIDBEJ7hFd4c4bw4787HorXg5DPH8AfO5w/f742M</latexit>

<latexit sha1_base64="Mx4dJlBZd0WdhYo0BwxmB7UE4GY=">AAAB6nicbVBNS8NAEJ3Ur1q/qh69LBbBU0lEqseiF48V7Qe0oWy2k3bpZhN2N0IJ/QlePCji1V/kzX/jts1BWx8MPN6bYWZekAiujet+O4W19Y3NreJ2aWd3b/+gfHjU0nGqGDZZLGLVCahGwSU2DTcCO4lCGgUC28H4dua3n1BpHstHM0nQj+hQ8pAzaqz00JO8X664VXcOskq8nFQgR6Nf/uoNYpZGKA0TVOuu5ybGz6gynAmclnqpxoSyMR1i11JJI9R+Nj91Ss6sMiBhrGxJQ+bq74mMRlpPosB2RtSM9LI3E//zuqkJr/2MyyQ1KNliUZgKYmIy+5sMuEJmxMQSyhS3txI2oooyY9Mp2RC85ZdXSeui6tWqtfvLSv0mj6MIJ3AK5+DBFdThDhrQBAZDeIZXeHOE8+K8Ox+L1oKTzxzDHzifP1BFjdY=</latexit>

<latexit sha1_base64="JD2FpJmF28nM/k99oPaBSP8phL0=">AAAB8HicbVBNSwMxEJ31s9avqkcvwSLUS9kVqXoretBjhX5Ju5Rsmm1Dk+ySZIWy9Fd48aCIV3+ON/+NabsHbX0w8Hhvhpl5QcyZNq777aysrq1vbOa28ts7u3v7hYPDpo4SRWiDRDxS7QBrypmkDcMMp+1YUSwCTlvB6Hbqt56o0iySdTOOqS/wQLKQEWys9HhXqvdS73py1isU3bI7A1omXkaKkKHWK3x1+xFJBJWGcKx1x3Nj46dYGUY4neS7iaYxJiM8oB1LJRZU++ns4Ak6tUofhZGyJQ2aqb8nUiy0HovAdgpshnrRm4r/eZ3EhFd+ymScGCrJfFGYcGQiNP0e9ZmixPCxJZgoZm9FZIgVJsZmlLcheIsvL5PmedmrlCsPF8XqTRZHDo7hBErgwSVU4R5q0AACAp7hFd4c5bw4787HvHXFyWaO4A+czx90U4+M</latexit>

<latexit sha1_base64="VbREsMKlF8oXO9blCj4TDV1yPTM=">AAAB7nicbVBNS8NAEJ34WetX1aOXxSJ4KomU6rHoxWOFfkEbyma7aZduNmF3IpTQH+HFgyJe/T3e/Ddu2xy09cHA470ZZuYFiRQGXffb2djc2t7ZLewV9w8Oj45LJ6dtE6ea8RaLZay7ATVcCsVbKFDybqI5jQLJO8Hkfu53nrg2IlZNnCbcj+hIiVAwilbqNAeZV63OBqWyW3EXIOvEy0kZcjQGpa/+MGZpxBUySY3peW6CfkY1Cib5rNhPDU8om9AR71mqaMSNny3OnZFLqwxJGGtbCslC/T2R0ciYaRTYzoji2Kx6c/E/r5dieOtnQiUpcsWWi8JUEozJ/HcyFJozlFNLKNPC3krYmGrK0CZUtCF4qy+vk/Z1xatVao/Vcv0uj6MA53ABV+DBDdThARrQAgYTeIZXeHMS58V5dz6WrRtOPnMGf+B8/gCHm48P</latexit>

<latexit sha1_base64="Mx4dJlBZd0WdhYo0BwxmB7UE4GY=">AAAB6nicbVBNS8NAEJ3Ur1q/qh69LBbBU0lEqseiF48V7Qe0oWy2k3bpZhN2N0IJ/QlePCji1V/kzX/jts1BWx8MPN6bYWZekAiujet+O4W19Y3NreJ2aWd3b/+gfHjU0nGqGDZZLGLVCahGwSU2DTcCO4lCGgUC28H4dua3n1BpHstHM0nQj+hQ8pAzaqz00JO8X664VXcOskq8nFQgR6Nf/uoNYpZGKA0TVOuu5ybGz6gynAmclnqpxoSyMR1i11JJI9R+Nj91Ss6sMiBhrGxJQ+bq74mMRlpPosB2RtSM9LI3E//zuqkJr/2MyyQ1KNliUZgKYmIy+5sMuEJmxMQSyhS3txI2oooyY9Mp2RC85ZdXSeui6tWqtfvLSv0mj6MIJ3AK5+DBFdThDhrQBAZDeIZXeHOE8+K8Ox+L1oKTzxzDHzifP1BFjdY=</latexit>

É É É É

T4 T5 T71 T229 <latexit sha1_base64="djxGdettxtZKsOrpLUXqqEyKfnE=">AAAB6nicbVDLSgNBEOyNrxhfUY9eBoPgKexKiB6DXjxGzAuSJcxOOsmQ2dllZlYISz7BiwdFvPpF3vwbJ8keNLGgoajqprsriAXXxnW/ndzG5tb2Tn63sLd/cHhUPD5p6ShRDJssEpHqBFSj4BKbhhuBnVghDQOB7WByN/fbT6g0j2TDTGP0QzqSfMgZNVZ6bPQr/WLJLbsLkHXiZaQEGer94ldvELEkRGmYoFp3PTc2fkqV4UzgrNBLNMaUTegIu5ZKGqL208WpM3JhlQEZRsqWNGSh/p5Iaaj1NAxsZ0jNWK96c/E/r5uY4Y2fchknBiVbLhomgpiIzP8mA66QGTG1hDLF7a2EjamizNh0CjYEb/XlddK6KnvVcvWhUqrdZnHk4QzO4RI8uIYa3EMdmsBgBM/wCm+OcF6cd+dj2ZpzsplT+APn8wfc542K</latexit>

<latexit sha1_base64="XaDArZEwEKH2gq7lVG0y8JsXPLM=">AAAB6nicbVDLSgNBEOyNrxhfUY9eBoPgKeyKRo9BLx4j5gXJEmYnvcmQ2dllZlYIIZ/gxYMiXv0ib/6Nk2QPmljQUFR1090VJIJr47rfTm5tfWNzK79d2Nnd2z8oHh41dZwqhg0Wi1i1A6pRcIkNw43AdqKQRoHAVjC6m/mtJ1Sax7Juxgn6ER1IHnJGjZUe672rXrHklt05yCrxMlKCDLVe8avbj1kaoTRMUK07npsYf0KV4UzgtNBNNSaUjegAO5ZKGqH2J/NTp+TMKn0SxsqWNGSu/p6Y0EjrcRTYzoiaoV72ZuJ/Xic14Y0/4TJJDUq2WBSmgpiYzP4mfa6QGTG2hDLF7a2EDamizNh0CjYEb/nlVdK8KHuVcuXhslS9zeLIwwmcwjl4cA1VuIcaNIDBAJ7hFd4c4bw4787HojXnZDPH8AfO5w/ea42L</latexit>

<latexit sha1_base64="6aHL3odxG2uksK3tdbWtSIgsEr8=">AAAB7XicbVBNSwMxEJ2tX7V+VT16CRbBU9kVaT0WvXis0C9ol5JNs21sNlmSrFCW/gcvHhTx6v/x5r8x2+5BWx8MPN6bYWZeEHOmjet+O4WNza3tneJuaW//4PCofHzS0TJRhLaJ5FL1AqwpZ4K2DTOc9mJFcRRw2g2md5nffaJKMylaZhZTP8JjwUJGsLFSpzVM6958WK64VXcBtE68nFQgR3NY/hqMJEkiKgzhWOu+58bGT7EyjHA6Lw0STWNMpnhM+5YKHFHtp4tr5+jCKiMUSmVLGLRQf0+kONJ6FgW2M8Jmole9TPzP6ycmvPFTJuLEUEGWi8KEIyNR9joaMUWJ4TNLMFHM3orIBCtMjA2oZEPwVl9eJ52rqler1h6uK43bPI4inME5XIIHdWjAPTShDQQe4Rle4c2Rzovz7nwsWwtOPnMKf+B8/gAWq47U</latexit>

<latexit sha1_base64="moDQStNZrQ/d4cqeZxXeqZOJM4Y=">AAAB7nicbVDLSgNBEOyNrxhfUY9eBoPgKewGiXoLevEYIS9IljA7mSRDZmeXmV4hLPkILx4U8er3ePNvnCR70MSChqKqm+6uIJbCoOt+O7mNza3tnfxuYW//4PCoeHzSMlGiGW+ySEa6E1DDpVC8iQIl78Sa0zCQvB1M7ud++4lrIyLVwGnM/ZCOlBgKRtFK7UY/rVRuZ/1iyS27C5B14mWkBBnq/eJXbxCxJOQKmaTGdD03Rj+lGgWTfFboJYbHlE3oiHctVTTkxk8X587IhVUGZBhpWwrJQv09kdLQmGkY2M6Q4tisenPxP6+b4PDGT4WKE+SKLRcNE0kwIvPfyUBozlBOLaFMC3srYWOqKUObUMGG4K2+vE5albJXLVcfr0q1uyyOPJzBOVyCB9dQgweoQxMYTOAZXuHNiZ0X5935WLbmnGzmFP7A+fwBja+PEw==</latexit>

É

T0 T1 T2 É <latexit sha1_base64="4P5UQRgOYdoVZjyzN25v+T5nAgQ=">AAAB6nicbVBNS8NAEJ3Ur1q/qh69LBbBU0lEqseiF48VW1toQ9lsJ+3SzSbsboQS+hO8eFDEq7/Im//GbZuDtj4YeLw3w8y8IBFcG9f9dgpr6xubW8Xt0s7u3v5B+fDoUcepYthisYhVJ6AaBZfYMtwI7CQKaRQIbAfj25nffkKleSybZpKgH9Gh5CFn1Fjpodl3++WKW3XnIKvEy0kFcjT65a/eIGZphNIwQbXuem5i/Iwqw5nAaamXakwoG9Mhdi2VNELtZ/NTp+TMKgMSxsqWNGSu/p7IaKT1JApsZ0TNSC97M/E/r5ua8NrPuExSg5ItFoWpICYms7/JgCtkRkwsoUxxeythI6ooMzadkg3BW355lTxeVL1atXZ/Wanf5HEU4QRO4Rw8uII63EEDWsBgCM/wCm+OcF6cd+dj0Vpw8plj+APn8wfW142G</latexit>

<latexit sha1_base64="ZVFuL/frNJY8X2Q2y99GZc4Lq+Y=">AAAB6nicbVBNS8NAEJ3Ur1q/qh69LBbBU0lEqseiF48VW1toQ9lsJ+3SzSbsboQS+hO8eFDEq7/Im//GbZuDtj4YeLw3w8y8IBFcG9f9dgpr6xubW8Xt0s7u3v5B+fDoUcepYthisYhVJ6AaBZfYMtwI7CQKaRQIbAfj25nffkKleSybZpKgH9Gh5CFn1Fjpodn3+uWKW3XnIKvEy0kFcjT65a/eIGZphNIwQbXuem5i/Iwqw5nAaamXakwoG9Mhdi2VNELtZ/NTp+TMKgMSxsqWNGSu/p7IaKT1JApsZ0TNSC97M/E/r5ua8NrPuExSg5ItFoWpICYms7/JgCtkRkwsoUxxeythI6ooMzadkg3BW355lTxeVL1atXZ/Wanf5HEU4QRO4Rw8uII63EEDWsBgCM/wCm+OcF6cd+dj0Vpw8plj+APn8wfYW42H</latexit>

<latexit sha1_base64="P5uTVhnDiNW6LSz/zb4pEApYqlI=">AAAB6nicbVDLSgNBEOyNrxhfUY9eBoPgKewGiR6DXjxGzAuSJcxOOsmQ2dllZlYISz7BiwdFvPpF3vwbJ8keNLGgoajqprsriAXXxnW/ndzG5tb2Tn63sLd/cHhUPD5p6ShRDJssEpHqBFSj4BKbhhuBnVghDQOB7WByN/fbT6g0j2TDTGP0QzqSfMgZNVZ6bPQr/WLJLbsLkHXiZaQEGer94ldvELEkRGmYoFp3PTc2fkqV4UzgrNBLNMaUTegIu5ZKGqL208WpM3JhlQEZRsqWNGSh/p5Iaaj1NAxsZ0jNWK96c/E/r5uY4Y2fchknBiVbLhomgpiIzP8mA66QGTG1hDLF7a2EjamizNh0CjYEb/XlddKqlL1qufpwVardZnHk4QzO4RI8uIYa3EMdmsBgBM/wCm+OcF6cd+dj2ZpzsplT+APn8wfZ342I</latexit>

Green Fraction

Fig. 2: Token-based semantic watermarking. During generation (left), green tokens (e.g., T1 ) are boosted. During verification (right), a statistical test is performed based on the fraction of present green tokens to determine the presence of a watermark. Generate

Verify

Red Set 00, 11

É

Boost next bit prediction

0

1

1

1

0

1 Decode

Green Set 01, 10

0

Encode

Static red/green n-gram split

É

0

1

1

1

1

0

1

Count green n-grams & test

+1

G 3 01 +1 G 3 10 +1 G 3 01 <latexit sha1_base64="8DsWlPjpfee24a6GpeHlh2caBls=">AAAB6HicbVDLSgNBEOyNrxhfUY9eBoPgKeyKRI9BD3pMwDwgWcLspDcZMzu7zMwKIeQLvHhQxKuf5M2/cZLsQRMLGoqqbrq7gkRwbVz328mtrW9sbuW3Czu7e/sHxcOjpo5TxbDBYhGrdkA1Ci6xYbgR2E4U0igQ2ApGtzO/9YRK81g+mHGCfkQHkoecUWOl+l2vWHLL7hxklXgZKUGGWq/41e3HLI1QGiao1h3PTYw/ocpwJnBa6KYaE8pGdIAdSyWNUPuT+aFTcmaVPgljZUsaMld/T0xopPU4CmxnRM1QL3sz8T+vk5rw2p9wmaQGJVssClNBTExmX5M+V8iMGFtCmeL2VsKGVFFmbDYFG4K3/PIqaV6UvUq5Ur8sVW+yOPJwAqdwDh5cQRXuoQYNYIDwDK/w5jw6L86787FozTnZzDH8gfP5A5+HjNY=</latexit>

<latexit sha1_base64="Mx4dJlBZd0WdhYo0BwxmB7UE4GY=">AAAB6nicbVBNS8NAEJ3Ur1q/qh69LBbBU0lEqseiF48V7Qe0oWy2k3bpZhN2N0IJ/QlePCji1V/kzX/jts1BWx8MPN6bYWZekAiujet+O4W19Y3NreJ2aWd3b/+gfHjU0nGqGDZZLGLVCahGwSU2DTcCO4lCGgUC28H4dua3n1BpHstHM0nQj+hQ8pAzaqz00JO8X664VXcOskq8nFQgR6Nf/uoNYpZGKA0TVOuu5ybGz6gynAmclnqpxoSyMR1i11JJI9R+Nj91Ss6sMiBhrGxJQ+bq74mMRlpPosB2RtSM9LI3E//zuqkJr/2MyyQ1KNliUZgKYmIy+5sMuEJmxMQSyhS3txI2oooyY9Mp2RC85ZdXSeui6tWqtfvLSv0mj6MIJ3AK5+DBFdThDhrQBAZDeIZXeHOE8+K8Ox+L1oKTzxzDHzifP1BFjdY=</latexit>

<latexit sha1_base64="8DsWlPjpfee24a6GpeHlh2caBls=">AAAB6HicbVDLSgNBEOyNrxhfUY9eBoPgKeyKRI9BD3pMwDwgWcLspDcZMzu7zMwKIeQLvHhQxKuf5M2/cZLsQRMLGoqqbrq7gkRwbVz328mtrW9sbuW3Czu7e/sHxcOjpo5TxbDBYhGrdkA1Ci6xYbgR2E4U0igQ2ApGtzO/9YRK81g+mHGCfkQHkoecUWOl+l2vWHLL7hxklXgZKUGGWq/41e3HLI1QGiao1h3PTYw/ocpwJnBa6KYaE8pGdIAdSyWNUPuT+aFTcmaVPgljZUsaMld/T0xopPU4CmxnRM1QL3sz8T+vk5rw2p9wmaQGJVssClNBTExmX5M+V8iMGFtCmeL2VsKGVFFmbDYFG4K3/PIqaV6UvUq5Ur8sVW+yOPJwAqdwDh5cQRXuoQYNYIDwDK/w5jw6L86787FozTnZzDH8gfP5A5+HjNY=</latexit>

<latexit sha1_base64="Mx4dJlBZd0WdhYo0BwxmB7UE4GY=">AAAB6nicbVBNS8NAEJ3Ur1q/qh69LBbBU0lEqseiF48V7Qe0oWy2k3bpZhN2N0IJ/QlePCji1V/kzX/jts1BWx8MPN6bYWZekAiujet+O4W19Y3NreJ2aWd3b/+gfHjU0nGqGDZZLGLVCahGwSU2DTcCO4lCGgUC28H4dua3n1BpHstHM0nQj+hQ8pAzaqz00JO8X664VXcOskq8nFQgR6Nf/uoNYpZGKA0TVOuu5ybGz6gynAmclnqpxoSyMR1i11JJI9R+Nj91Ss6sMiBhrGxJQ+bq74mMRlpPosB2RtSM9LI3E//zuqkJr/2MyyQ1KNliUZgKYmIy+5sMuEJmxMQSyhS3txI2oooyY9Mp2RC85ZdXSeui6tWqtfvLSv0mj6MIJ3AK5+DBFdThDhrQBAZDeIZXeHOE8+K8Ox+L1oKTzxzDHzifP1BFjdY=</latexit>

0

0

<latexit sha1_base64="8DsWlPjpfee24a6GpeHlh2caBls=">AAAB6HicbVDLSgNBEOyNrxhfUY9eBoPgKeyKRI9BD3pMwDwgWcLspDcZMzu7zMwKIeQLvHhQxKuf5M2/cZLsQRMLGoqqbrq7gkRwbVz328mtrW9sbuW3Czu7e/sHxcOjpo5TxbDBYhGrdkA1Ci6xYbgR2E4U0igQ2ApGtzO/9YRK81g+mHGCfkQHkoecUWOl+l2vWHLL7hxklXgZKUGGWq/41e3HLI1QGiao1h3PTYw/ocpwJnBa6KYaE8pGdIAdSyWNUPuT+aFTcmaVPgljZUsaMld/T0xopPU4CmxnRM1QL3sz8T+vk5rw2p9wmaQGJVssClNBTExmX5M+V8iMGFtCmeL2VsKGVFFmbDYFG4K3/PIqaV6UvUq5Ur8sVW+yOPJwAqdwDh5cQRXuoQYNYIDwDK/w5jw6L86787FozTnZzDH8gfP5A5+HjNY=</latexit>

+

<latexit sha1_base64="Mx4dJlBZd0WdhYo0BwxmB7UE4GY=">AAAB6nicbVBNS8NAEJ3Ur1q/qh69LBbBU0lEqseiF48V7Qe0oWy2k3bpZhN2N0IJ/QlePCji1V/kzX/jts1BWx8MPN6bYWZekAiujet+O4W19Y3NreJ2aWd3b/+gfHjU0nGqGDZZLGLVCahGwSU2DTcCO4lCGgUC28H4dua3n1BpHstHM0nQj+hQ8pAzaqz00JO8X664VXcOskq8nFQgR6Nf/uoNYpZGKA0TVOuu5ybGz6gynAmclnqpxoSyMR1i11JJI9R+Nj91Ss6sMiBhrGxJQ+bq74mMRlpPosB2RtSM9LI3E//zuqkJr/2MyyQ1KNliUZgKYmIy+5sMuEJmxMQSyhS3txI2oooyY9Mp2RC85ZdXSeui6tWqtfvLSv0mj6MIJ3AK5+DBFdThDhrQBAZDeIZXeHOE8+K8Ox+L1oKTzxzDHzifP1BFjdY=</latexit>

<latexit sha1_base64="MQil7zbQIavt9hqFmKnrSNGXeqE=">AAAB7nicbVBNS8NAEJ3Ur1q/qh69LBZBEEoiUj0WvXisYD+gDWWz2bRLN5uwOxFK6Y/w4kERr/4eb/4bt20O2vpg4PHeDDPzglQKg6777RTW1jc2t4rbpZ3dvf2D8uFRyySZZrzJEpnoTkANl0LxJgqUvJNqTuNA8nYwupv57SeujUjUI45T7sd0oEQkGEUrtS96IZdI++WKW3XnIKvEy0kFcjT65a9emLAs5gqZpMZ0PTdFf0I1Cib5tNTLDE8pG9EB71qqaMyNP5mfOyVnVglJlGhbCslc/T0xobEx4ziwnTHFoVn2ZuJ/XjfD6MafCJVmyBVbLIoySTAhs99JKDRnKMeWUKaFvZWwIdWUoU2oZEPwll9eJa3Lqler1h6uKvXbPI4inMApnIMH11CHe2hAExiM4Ble4c1JnRfn3flYtBacfOYY/sD5/AH8Oo9c</latexit>

Prev. Bit

1

Next Bit

1

Embed, scale up, and accumulate

Retrieve residuals and get bits

É

Green Fraction

Fig. 3: Watermarking for bitwise autoregressive models. Here, shown for the multi-scale generation model Infinity. Generation (left): A fixed green set of n-grams (here, bigrams G = {01, 10}) is used to bias sampling of bits on a given scale by δ. Verification (right): Bits are recovered, the fraction of green n-grams is determined, and a statistical hypothesis test is performed to check if it is significantly above chance level, indicating the presence of a watermark.

decoded tokens are used to partition the vocabulary of the generator into two sets, a “green” set Gi , and a “red” set Ri . Since this partitioning is dependent on the previous tokens, the green and red sets can be different between different timesteps. It is done by first computing a hash oi based on the previous l tokens and a secret key κ: oi = hash(yi−1 , . . . , yi−l , κ). The hash oi is used to seed a pseudorandom number generator (PRNG), which is then used to randomly sample a fraction γ of tokens from the vocabulary V to form the set of green tokens Gi of size ⌊γ ∗ |V|⌋. Finally, the model’s distribution over V is biased to favor tokens from the green set by adding δ to their logits: m_i[v] &= \mathbf {1}_{G_i}(v) \enspace , \\ p'_{\theta }(y_i | y_{<i}) &= \text {softmax} (f_{\theta }(y_{<i}) + m_i * \wmstrength ) \enspace . (2) A smaller δ minimizes the impact on the generator’s distribution but weakens the watermark signal. In order to verify the watermark for the given sequence (y1 , . . . , yT ), the green sets Gi are recomputed for every token given the preceding context and the number of green tokens is counted. Then, a one-sided right-tailed statistical test is used to assess how unlikely it is to observe the given number of green tokens, Ng = #green , in the entire sequence by chance. The null hypothesis is that Ng is distributed according to an unbiased binomial distribution: H0 : Ng ∼ Binomial(T, γ). The right-tailed p-value p = Pr(X ≥ Ng ) is computed and H0 is rejected for a sequence if p is lower than a threshold chosen for a specified false positive rate (FPR).

On the Robustness of Watermarking for Autoregressive Image Generation

2.3

5

Watermarking for Autoregressive Image Generation

In this work, we study removal and forgery attacks on a representative selection of recently proposed watermarking schemes for autoregressive image generation models for which we were able to obtain source code6 : IndexMark [37] divides the vocabulary into a set of pairs of tokens, maximizing token similarity within every pair. One of the tokens in the pair is assigned to the green set and the other to the red. In order to embed the watermark, in the generated token sequence, every red token is replaced with its green counterpart. This results minimizes image degradation and introduces a detectable watermark signal. The encoder is fine-tuned to improve token reconstruction accuracy. WMAR(NeurIPS′ 25) [16], following KGW, uses the preceding tokens to partition the vocabulary into green and red sets, where green token’s logits are boosted during sampling. See Fig. 2 for an illustration. In order to improve robustness, both the pixel-latent encoder and decoder are fine-tuned with perturbations. Specifically, it aims to improve reverse cycle consistency, i.e. the reconstruction accuracy of originally generated tokens from watermarked images. Lastly, WMAR proposes to use an image synchronization layer watermark like SyncSeal [10] to recover original image orientations, improving robustness against removal attempts by geometric transformations such as rotation. ClusterMark(CVPR′ 26) [26] also uses the preceding token to partition the vocabulary and studies token clustering as a possible means to improve the robustness against removal attempts (since similar tokens are clustered and assigned the same color, switching between similar tokens does not destroy the watermark). Furthermore, the encoder is augmented with a classification head and trained with a perturbation-augmented dataset to further boost the reconstruction accuracy of original tokens or clusters from watermarked images. BitMark(NeurIPS′ 25) [18] has been developed specifically for the Infinity [13] generator. Recall that during image generation, Infinity [13] computes logits over bits at scales of increasing resolution, conditioning the parallel prediction of all bits in one scale on bits in all previously predicted scales. With this in mind, BitMark follows an approach similar to KGW, where first the distributions over each bit in one scale is viewed as a sequence (unfolded over the channel dimension first). Given the sequence of binary distributions, the watermark is embedded by biasing the sampling of a bit based on the l previously sampled bits. Given that the vocabulary size is only 2, there are only a few possible conditional partitionings of the vocabulary (e.g. G2 = {01, 10} for l = 1), where each one is empirically evaluated and the best ones are used as the final method. See Fig. 3 for an illustation. BitMark was specifically proposed to enable service providers to re-identify their generated content and exclude it from subsequent training, thereby mitigating model collapse [1, 18, 31]. This goal is reinforced 6

C-Reweight [41] is not included in our evaluation due to their code being unavailable.

6

A. Müller et al.

by its radioactive property: the watermark signal propagates to models trained on watermarked data, even transferring across architectures (e.g., to diffusion models).

3

Attacks on Autoregressive Image Watermarks

In this work, we evaluate existing removal and forgery attacks on AR model watermarking and propose three new attack methods that specifically target AR generative models. The first is Vector-Quantized Regeneration (VQ-Regen), a removal attack that aims at disrupting the precise token index recovery by regenerating a watermarked image from close-by tokens. The second is Latent Encoder Optimization (LatentOpt), which can be applied as both a removal and a forgery attack. The third is Frequency Injection forgery, which performs watermark forgery by injecting localized peaks into a cover image’s Fourier representation. 3.1

Vector-Quantized Regeneration Attack (VQ-Regen)

In the Vector-Quantized Regeneration Attack (VQ-Regen) for watermark removal, we use a VQ-VAE codebook to generate a perturbed reconstruction x′ as follows. First, we encode the image: z = E(x), where z ∈ Rd×h×w . For each spatial location (i, j) ∈ {1, . . . h} × {1, . . . w}, we sort all codebook vectors in C by their Euclidean distance from z·,i,j . This gives us a ranking si,j of V, such that si,j [1] is the closest and si,j [k] is the k-th closest token index. Next, we construct the attacked token map t′ ∈ V h×w by replacing the standard quantization step with the selection of the k-th nearest token index: t′i,j = si,j [k]. Finally, the attacked image is produced by mapping t′ back to latent space via codebook lookup z ′ = C[t′ ], and decoding the result: x′ = D(z ′ ). This is explained in detail in Algorithm 3 in Supplementary Material. Note that setting k = 1 recovers the (unattacked) standard nearest-neighbor reconstruction, while k > 1 induces a controlled structural deviation in the generated image. 3.2

Latent Encoder Optimization (LatentOpt)

This attack adds a subtle pixel-level perturbation ∆x to an image x, which propagates through the autoencoder to shift the latent representation in a targeted, gradient-guided manner, while minimally affecting the image appearance. For removal, the goal is to move the latents away from the current latents, such that different tokens are decoded, and the watermark signal is lost. For forgery, the latents of the target image are moved towards those of a watermarked reference image. LatentOpt-Removal. We use a VQ-VAE encoder E to obtain the latent z from the watermarked image x. We optimize a pixel-delta ∆x such that the E(x +

On the Robustness of Watermarking for Autoregressive Image Generation

7

∆x) moves away from the initial E(x), while keeping its p-norm (we use p=∞) constrained by a budget c: \pert ^* = & ~\argmax _{\pert } |\encoder (x + \pert ) - \encoder (x)|^2 \\ & ~ s.t. ~~ |\pert |^p < c \enspace . (4) Here, E can be any encoder that maps pixels to a latent space, either from the model’s own VAE with the same weights (leading to a white-box setting), from another, unrelated VAE (black-box), or a VAE of similar architecture (grey-box). LatentOpt-Forgery. For forgery, the approach is similar, but we optimize the latent of a cover image xc towards the latent xw of a watermarked image: \pert ^* = & ~\argmin _{\pert } |\encoder (x_c+ \pert ) - \encoder (x_w)|^2 \\ & ~ s.t. ~~ |\pert |^p < c \enspace . (6) BitOpt-Removal for BitMark. To stress test BitMark against the strongest possible adversary, we developed a white-box+ adversarial optimization attack specifically tailored for BitMark. This attack assumes access to the latent encoder E, the scale resolutions used, and the knowledge of the green set G. Note that in the original work [18], the BitFlipper attack was proposed with identical assumptions and was found unsuccessful due to severe image quality degradation. In our white-box BitOpt-Removal attack, the target image x is first encoded using E to obtain its latent z. Then, given the scale resolutions, z is converted to a sequence of residuals, similarly to the regular encoding procedure, where both the unquantized and the single-bit-quantized residuals are retained. In the next step, the positions of bits that can be flipped to reduce the green token count are identified. For example, BitMark [18] by default uses 0 → 1 and 1 → 0 as the green bigrams. With this knowledge, we identify the positions of 010 or 101 trigrams as flipping targets, since flipping the 1 in 010 would decrease the green token count by 2. In the final step, we perform adversarial optimization on the identified target positions, with the objective to flip the signs of the unquantized latent residuals using an L1 loss. For reference, this method is formally elaborated in Algorithm 1 in the Supplementary Material. 3.3

Frequency Injection

We conducted a steganographic analysis of BitMark-generated images by averaging 5000 samples and found that the residual signal verifies against the watermark detector with p-values in the order of 10−33 (see Fig. 4). Yet, adding the averaged signal to cover images (for forgery) or subtracting it from watermarked images (for removal), similarly to the averaging attack [42], resulted in poor image quality while not sufficiently affecting detection metrics. Inspired by this finding, we develop a simple forgery attack, in which a pattern of magnitude peaks is injected into the frequency representation of a cover image in regular intervals along the diagonals to introduce a regular pixel pattern. We set

8

A. Müller et al. Average of 5000 BitMark Images (p=3e-33)

Real Cover Image (p=0.02, PSNR= 1) <latexit sha1_base64="bF/NP/OHOrpLPjZtD9cibHjZtzk=">AAAB7XicbVBNS8NAEJ3Ur1q/qh69BIvgqSQi1WPRi8cK9gPaUDbbTbt2sxt2J0Io/Q9ePCji1f/jzX/jts1BWx8MPN6bYWZemAhu0PO+ncLa+sbmVnG7tLO7t39QPjxqGZVqyppUCaU7ITFMcMmayFGwTqIZiUPB2uH4dua3n5g2XMkHzBIWxGQoecQpQSu1elxGmPXLFa/qzeGuEj8nFcjR6Je/egNF05hJpIIY0/W9BIMJ0cipYNNSLzUsIXRMhqxrqSQxM8Fkfu3UPbPKwI2UtiXRnau/JyYkNiaLQ9sZExyZZW8m/ud1U4yugwmXSYpM0sWiKBUuKnf2ujvgmlEUmSWEam5vdemIaELRBlSyIfjLL6+S1kXVr1Vr95eV+k0eRxFO4BTOwYcrqMMdNKAJFB7hGV7hzVHOi/PufCxaC04+cwx/4Hz+AMXBj0c=</latexit>

Forgery Attack (p=1e-19 | PSNR=35.9)

Fig. 4: BitMark Forgery via Frequency Injection. (Left) Averaging 5000 BitMark-watermarked images reveals a structured spatial artifact detectable by the watermarking scheme (p = 3 × 10−33 ), and its Fourier transform on the right. (Center) A real cover image and its Fourier transform are not detected as watermarked (p = 0.02, PSNR = ∞). (Right) Injecting the frequency patterns marked by red circles into the cover image achieves high visual fidelity (PSNR = 35.9 dB) while successfully triggering BitMark’s detector (p = 1 × 10−19 ).

random phases for each color channel to reduce the saliency of the injected pattern and reduce visibility. We found three settings that yielded three desirable tradeoffs ranging from median p-values between 10−7 to 10−31 , and mean PSNR between 37.5 to 29.75 over a set of 100 runs. The exact settings, as well as the formal description of the frequency injection procedure are described in Sec. A and Algorithm 4 in the Supplementary Material, respectively.

4

Evaluation

4.1

Experimental Setup

Attacked Watermarking Schemes. We investigate attacks on three discrete token watermarking schemes (IndexMark, WMAR, ClusterMark) and one for bitwise autoregressive models (BitMark). Each scheme is tested with multiple generative models and different settings, as listed below. Unless otherwise specified, we use the default settings in the code or paper provided with each watermarking scheme. – IndexMark [37] is deployed with the LlamaGen [33] model in both the GPT-B (generating images of size 256×256) and GPT-L (384×384) variants. Both models are class-conditional, using 1k ImageNet [29] classes. – WMAR [16] is tested with different models: RAR-XL [44] (class-conditional, 256×256), Taming Transformers [7] (class-conditional, 256×256) and Anole [5] (text-conditional, 512 × 512). WMAR is used with the fine-tuned encoder and decoder, but without the synchronization layer to focus the evaluation on non-geometric transformation robustness. – ClusterMark [26] is tested with LlamaGen both with GPT-B and GPT-L, as well as RAR-XL in the default settings with 64 clusters. We also use the finetuned cluster classifier. – BitMark [18] is tested with Infinity-2B [13] (text-conditional generation, image size of 1024 × 1024) and default settings (green set G = {01, 10} active on all scales, and δ = 2).

On the Robustness of Watermarking for Autoregressive Image Generation

9

Deployed Attacks. We evaluate two categories of attacks. Full experimental details are provided in Sec. A of the Supplementary Material. Our attacks. We evaluate LatentOpt-Removal (LatentOpt-R) and LatentOptForgery (LatentOpt-F), our optimization-based adversarial attacks, as well as VQ-Regen, a token substitution attack. Specifically for BitMark, we also evaluate BitOpt-Removal (BitOpt-R), a white-box+ variant with additional knowledge of the watermarking settings, as well as Frequency Injection forgery under multiple quality/success trade-off settings. Existing attacks. We evaluate the watermarking schemes against strong Perturbations (JPEG compression, additive noise, color jitter), geometric transformations, as well as existing diffusion-based regeneration attacks, Regen. and Rinse [47], and CtrlRegen+ [24]. For BitMark, we also evaluate their BitFlipper attack [18]. Visual examples of perturbations are provided in Sec. A of the Supp. Material. Datasets and Metrics. For both removal and forgery, each attack is performed against 100 watermarked target images generated with each of the multiple deployed target models per watermarking scheme. For forgery attacks, we use cover images from the MS-COCO dataset [23]. In line with previous work [16,18,26,37], we report average TPR@FPR=1% as accuracy metric (i.e., can one verify the watermark correctly?), as well as median p-values as a threshold-agnostic alternative to TPR. We also report quality degradation between attacked images and watermarked target images (removal), as well as cover images (forgery), in terms of PSNR↑ and LPIPS↓ [45] scores respectively. For BitMark, we additionally include accuracy metrics on radioactive data in different settings: We finetune Infinity-2B and Stable Diffusion v2.1 [28] on 1,000 images generated using watermarked Infinity-2B (using captions from MS-COCO). During fine-tuning, different ratios of watermarked to unwatermarked images were used. Each watermarking scheme is evaluated across multiple target models (e.g., WMAR is tested on RAR, Anole and Taming), with results aggregated over all settings. White-box, Grey-box, and Black-box Settings. For white-box attacks, we aggregate results from attacks that use the exact same fine-tuned encoder as the one deployed by the verifier. For grey-box settings, we use a closely related attacker model to the deployed verifier model. Specifically, for token-based watermarks, the grey-box attacker model is the non-finetuned version of the verifier model. For bit-autoregressive watermarks, the grey-box attacker model is an identically trained model of the same architecture but different vocabularies (Vd = 216 , Vd = 224 , Vd = 264 ). For black-box settings, we aggregate attacks launched with an entirely different encoder than the one deployed by the verifier. 4.2

Results for Token-based Watermarks

The results for both removal and forgery attacks on token-based watermarks are shown in Tab. 1. Fig. 6 shows qualitative examples for a subset of evaluated

10

A. Müller et al.

Table 1: Results for token-based methods, aggregated over multiple target models deployed with each watermarking scheme. ■ is black-box, ■ is grey-box, □ is whitebox. P-values are reported as median, and all others as mean. For removal, attackers target lower TPR (higher p-values); for forgery, higher TPR (lower p-values).

TPR WM

IndexMark P-value PSNR LPIPS

1.00 4.3×10−78

-

-

TPR

WMAR P-value PSNR LPIPS

1.00 4.9×10−45

-

-

TPR

ClusterMark P-value PSNR LPIPS

1.00 4.6×10−80

-

-

Removal 0.58 5.1×10−3 18.38 0.53 4.3×10−3 17.12

0.12 0.44

0.55 7.3×10−3 18.80 0.43 3.0×10−2 17.14

0.12 0.43

0.78 9.9×10−6 18.51 0.80 2.3×10−23 17.15

0.12 0.45

Regen. ■ 0.52 1.0×10−2 22.91 Rinse ■ 0.03 3.5×10−1 20.14 CtrlRegen ■ 0.36 4.0×10−2 23.40

0.13 0.29 0.14

0.36 5.3×10−2 24.09 0.02 3.7×10−1 21.11 0.22 1.0×10−1 24.29

0.11 0.29 0.11

0.85 1.6×10−5 22.95 0.29 6.2×10−2 19.95 0.80 5.5×10−5 23.73

0.13 0.31 0.14

VQ-Regen ■ 0.03 4.2×10−1 20.32 ■ 0.01 7.9×10−1 22.30 □ 0.00 1.0 21.81

0.15 0.10 0.12

0.12 2.3×10−1 21.59 0.19 2.4×10−1 22.13 0.08 3.3×10−1 21.76

0.12 0.11 0.13

0.39 2.5×10−2 20.49 0.79 3.2×10−5 21.12 0.82 1.8×10−5 17.54

0.15 0.12 0.22

LatentOpt ■ 0.63 1.8×10−4 31.94 (Removal) ■ 0.10 6.2×10−1 31.67 □ 0.00 1.0 31.81

0.12 0.15 0.04

0.42 5.2×10−2 32.13 0.11 8.4×10−1 32.21 0.00 9.4×10−1 32.28

0.13 0.13 0.13

0.78 4.1×10−7 31.97 0.42 3.0×10−2 31.82 0.00 9.8×10−1 32.79

0.12 0.14 0.10

LatentOpt ■ 0.10 1.1×10−1 34.34 (Forgery) ■ 0.14 9.5×10−2 34.65 □ 1.00 8.6×10−78 32.85

0.07 0.06 0.03

0.06 1.7×10−1 32.26 0.10 1.6×10−1 32.26 0.15 1.0×10−1 31.92

0.10 0.10 0.10

0.10 1.3×10−1 34.49 0.12 1.1×10−1 34.65 0.79 3.6×10−11 33.00

0.06 0.06 0.06

Geom. Perturb.

Forgery

attacks. A model-wise breakdown of all reported values is shown in Sec. B in the Supplementary Material. Overall, the results indicate that the three discrete token watermarks tested are removable, even in black-box settings, but are not easily forgeable (unless in a white-box setting). Firstly, we observe that strong image perturbations and geometric transformations can disturb the watermark detection for all schemes. Note that, while discrete token watermark schemes have already been shown to be vulnerable to geometric transformations [16], this can be addressed by adding an additional synchronization layer like SyncSeal [10]. Secondly, we analyze diffusion regeneration attacks. We observe that Regen. and CtrlRegen+, while generally effective with a success rate of 78%-48% for IndexMark and WMAR, display high pixel distortion with ∼23 dB PSNR, while maintaining relatively high perceptual quality (low LPIPS of ∼0.13). Rinse, while effective, shows considerable quality loss in terms of both PSNR and LPIPS. ClusterMark generally exhibits more robustness against these attacks. Similarly, VQ-Regen causes severe pixel-wise distortion (low PSNR), while maintaining relatively high perceptual quality (low LPIPS). Generally, VQRegen in the black-box setting achieves slightly higher removal success with comparable quality degradation. Interestingly, for ClusterMark, using similar or even the exact same proxy VQ-VAE as the deployed target model renders the attack less effective: The close alignment of the proxy VQ-VAE to the target in the grey- and white-box setting actually hinders the attacker from removing the watermark, as the substituted tokens are still very likely to fall into correct token clusters, maintaining the integrity of green set assignments with consecu-

On the Robustness of Watermarking for Autoregressive Image Generation

0.2 45

40

35

30

25

PSNR (dB)

20

15

0.4 0.2 45

40

35

30

25

PSNR (dB)

20

15

10

0.6 0.4 0.2 0.0

0.6 0.4 0.2 45

40

35

30

PSNR (dB)

25

20

0.8 0.6 0.4 0.2 0.0

40

35

30

25

PSNR (dB)

20

15

10

45

40

35

30

PSNR (dB)

25

20

0.6

Perturbations

0.4

JPEG Gaussian Noise Salt & Pepper Gaussian Blur

0.2 0.0

45

40

35

30

25

PSNR (dB)

20

15

10

1.0

0.8 0.6 0.4 0.2 0.0

Watermarked =1.00 LatentOpt-Removal Black-box Grey-box White-Box

0.8

1.0

TPR @ FPR=1%

0.8

45

BitMark

1.0

0.8

1.0

TPR @ FPR=1%

TPR @ FPR=1%

0.6

0.0

10

1.0

0.0

0.8

TPR @ FPR=1%

0.4

ClusterMark

1.0

TPR @ FPR=1%

TPR @ FPR=1%

TPR @ FPR=1%

0.6

0.0

WMAR

1.0

0.8

TPR @ FPR=1%

IndexMark

1.0

11

45

40

35

30

PSNR (dB)

25

20

0.8 0.6

Watermarked =1.00 LatentOpt-Forgery Black-box Grey-box White-Box

0.4 0.2 0.0

45

40

35

30

PSNR (dB)

25

20

Fig. 5: LatentOpt-Removal (top) and LatentOpt-Forgery (bottom) results 2 4 8 16 32 for different budgets c ∈ { 255 , 255 , 255 , 255 , 255 }. The top row also shows different perturbation baselines in varying strengths.

tive tokens. With less alignment in the black-box setting, disturbing the latent sufficiently to escape clusters is more feasible. Our LatentOpt-Removal attack is also effective and is able to maintain low pixel-wise distortion (30-33 dB PSNR), resulting in different and potentially more favourable trade-off between quality and attack success. Finally, we see that forgery is significantly more difficult, failing to achieve >10% TPR@FPR=1% in black-box settings. In the white-box setting, forgery is most effective when LlamaGen was used by both the verifier and attacker, as is the case for part of the setup for IndexMark and ClusterMark. A possible explanation is the relatively small embedding dimension (8) used in LlamaGen’s VQ-VAE, in contrast to the 256 dimensions in all other models. Entering correct Voronoi cells for each token is easier with fewer embedding dimensions because the decision boundaries depend on fewer orthogonal directions, making the target region geometrically less fragmented and reducing the number of independent constraints that the forgery attack instance must simultaneously satisfy. Effect of Perturbation Budget. We study the tradeoff between attack success (TPR@1%FPR) and quality degradation (PSNR) of our LatentOpt attacks in different box-settings and with varying perturbation budgets c, and compare them to different naive image perturbations (noise, blur, etc.) of varying strength (see Fig. 5). We observe the following: (1) using gradients informed by a VQ-VAE encoder in order to move an image’s latents enables more efficient removal attacks than naive pertubrations, (2) both removal and forgery efficiency improves with closer alignment between the proxy encoder and the encoder deployed on the verifier’s side, i.e., using an unrelated proxy encoder (black-box) is outperformed by using the same architecture (grey-box), which is in turn outperformed by using identical weights (white-box). This indicates a transferability property of adversarial optimizations for attacking autoregressive watermarks, similar to the properties shown in regard to attacks on semantic watermarking for diffusion models [27].

12

A. Müller et al.

Table 2: Results for removal attack (left), as well as watermarked, radioactive data baseline, and forgery (right) attacks for BitMark. ■ is black-box, ■ is grey-box, □ is white-box. □+ indicates having access to the green set G. LatentOpt and VQ-Regen (■ & ■) were performed with several attacker models, or specifically with LlamaGen (indicated by ♣ ). P-values are reported as median, and all others as mean. For removal, attackers target lower TPR (higher p-values); for forgery, higher TPR (lower p-values). TPR

P-value

1.00 1.2×10 16.20 0.98 1.3×10−253 17.56

0.24 0.47

■ ■ ■

1.00 1.5×10−233 28.04 1.00 1.1×10−68 25.20 1.00 8.1×10−22 23.60

0.06 0.14 0.22

1.00 3.3×10−83 24.21

0.10

LatentOpt-R LatentOpt-R

■ ■ ■ □

0.88 2.4×10−32 32.34 0.97 5.2×10−100 34.32 0.93 2.0×10−47 32.45 0.93 7.6×10−18 31.67

0.28 0.22 0.26 0.26

BitFlipper

□+ 0.33 9.5×10−1

18.64

0.27

BitOpt-R

−1

46.17

0.01

Geom. Perturb. Regen. Rinse CtrlRegen+ VQ-Regen ♣

4.3

□+ 0.00 3.0×10

TPR

PSNR LPIPS

−71

Watermarked

P-value

1.00 0.0

Radio. ∞-2B, 10% 0.74 2.4×10−4 ∞-2B, 50% 1.00 2.3×10−55 ∞-2B, 100% 1.00 8.4×10−212 SD2.1, 50% 0.81 2.9×10−5 SD2.1, 100% 0.98 6.5×10−14

PSNR LPIPS -

-

-

-

LatentOpt-F♣ ■ LatentOpt-F ■ ■ □

0.88 3.7×10−5 31.59 0.59 3.5×10−3 31.51 0.99 1.7×10−23 32.51 1.00 3.4×10−232 31.96

0.18 0.19 0.15 0.16

F.Inj. Setting A Setting B Setting C

0.73 2.7×10−7 37.53 0.81 3.0×10−17 35.51 0.93 7.5×10−31 29.75

0.06 0.10 0.19

Results for Bitwise-Autoregressive Watermarks

Tab. 2 reports the results of the studied attacks on BitMark. Fig. 6 shows qualitative examples for a subset of evaluated attacks. We observe that BitMark is extremely robust to removal attempts, in part due to the large number of bits, which results in higher statistical significance. Thus, removal attacks are less effective than for token-level watermarks. Previous regeneration attacks are ineffective but they do lower the p-values. Our VQ-Regen is ineffective as well. For LatentOpt-Removal attacks, we notice a difference between LlamaGen’s VAE as attacker model and the others. With LlamaGen’s VAE, watermark removal is more effective for a fraction of examples, lowering TPR below 90% while retaining > 32 dB PSNR. While the vanilla white-box LatentOpt-Removal fails to remove the watermark within the perturbation budget, our BitMark-specific BitOpt-R attack is able to completely erase the watermark while maintaining a PSNR of over 45 dB. Conversely, forgery of BitMark is more effective than for token-based watermarks. In the black-box setting, LatentOpt-Forgery is able to push TPR above 50% while maintaining PSNR greater than 31. Interestingly, LlamaGen’s encoder is significantly more successful in faking a BitMark than the average proxy encoder in black-box settings with a TPR of 88%, outperforming even the average grey-box encoder. This indicates that an attacker can benefit from evaluating multiple unrelated proxy models, as they can differ significantly in terms of attack success. In the white-box setting, the forgery attack is almost indistinguishable from originally generated watermarked images with very low p-values. Frequency injection is able to achieve low p-values without access to any model.

On the Robustness of Watermarking for Autoregressive Image Generation Watermarked Original LatentOpt-R (Black-B.)

BitOpt (White-B.)

BitFlipper

p < 2e-308

p = 0.11

p > 0.99

p > 0.99

VQ-Regen.

Regen.

Rinse

CtrlRegen+

p = 5e-115

p = 1e-147

p = 3e-34

p = 2e-9

Authentic Cover Image

Frequency Injection

p = 0.87

p = 2e-67

13

LatentOpt-F (White-B.) LatentOpt-F (Black-B.)

p < 2e-308

p = 3e-7

Fig. 6: Qualitative Eaxmples for BitMark deployed with Infinity-2B for removal attacks (left) and forgery attacks (right). LatentOpt-R and LatentOpt-F attacks are 8 . Black-box settings and the VQ-Regen attack use capped by a budget of ∥∆x∥∞ = 255 LlamaGen’s VQ-VAE and Infinity-2B’s VQ-VAE as proxy models, respectively. Watermarked Radioactive ( -2B, 50%) Radioactive (SD2.1, 50%) Frequency Injection Forgery (A,B,C) LatentOpt-F LlamaGen LatentOpt-F LatentOpt-F LatentOpt-F VQ-Regen Diffusion Regen. (all) LatentOpt-R LlamaGen LatentOpt-R LatentOpt-R LatentOpt-R BitOpt-R +

0

10

20

30

40

Z-Score

Watermarked: z=79.02, p<2e-308 Radioactive Inf-2B: z=15.63, p=2e-55 Radioactive SD2.1: z=4.02, p=3e-05 50 60 70 80

Fig. 7: Box plot showing the distributions of z-scores for different attacks on BitMark (forgery is orange, removal is blue) as well as original watermarked and radioactive images (grey).

Can attacks be mitigated by adjusting p-value threshold? The box-plot in Fig. 7 shows the spread of z-scores for the different forgery and removal attacks from Tab. 2, compared to the z-scores generators (Infinity-2B, SD2.1) affected by radioactivity, i.e. they have been finetuned on 50% watermarked images. The median z-scores for the radioactive model data, as well as the original watermarked images are shown as dashed lines. We observe that removal attacks in certain black-box settings cannot be clearly separated from Frequency Injection forgery - a significant portion of their distribution mass overlaps. Hence, perfectly protecting against both removal and forgery by adjusting the detection threshold is not possible, even for a black-box attacker. Furthermore, in order to enable the detection of radioactively generated images, the detection threshold has to be set below the bulk of the corresponding distributions. However, in doing so, the verifier also enables false positive detection of forgery attack instances. For example, frequency injection forgery well surpasses any reasonable threshold that includes detection of by the radioactive SD2.1 model trained on 50% watermarked training data. Note that assuming that a model would have been finetuned on 50% watermarked data is already a strong assumption.

14

A. Müller et al.

Enabling detection of even lower fractions increases the opportunity for forgery attempts. For example, in a more realistic scenario of an Infinity-2B model being finetuned on 10% watermarked data (z = 3.48, p = 2.4 × 10−4 ), completely rules out being able to both detect radioactive images and prevent forgery attacks. Overall, this suggests that a service provider that prioritizes reliable identification of radioactively emitted data or robustness to removal attacks inevitably becomes susceptible to forgery attacks, even in black-box setting.

5

Related Work

Prior to the recent watermarking schemes for autoregressive image (AR) generation models we studied in this work, several in-generation watermarking schemes have been proposed for latent diffusion models (LDM) [6,9,11,40,43], which rely on DDIM inversion [14, 32] for watermark verification. This was followed by works studying removal [2, 15, 17, 24, 25, 30, 42, 47] and forgery [15, 27, 42] attacks against these watermarks. Of these attacks, we included regeneration-based [24, 47] removal attacks in our comparison. As discussed in Sec. 3.3, we also tried the averaging attack [42]. We further tried UnMarker [17] for removal, but were unable to perform it on images bigger than 384 × 384 due to prohibitevely expensive memory requirements, ruling out large parts of our evaluation setup. We were unable to obtain source code for the very recent RAVEN [30] attack. Note that while [25,27] could potentially be effective on in-generation watermarking for AR models, these attacks are specifically tailored to the mechanics of in-generation watermarks for LDMs, since they include DDIM inversion to mimic. Finally, the optimizationbased attacks in [2,15] on LDM watermarks are similar to our LatentOpt attack, but use LDM-specific VAE encoders and different optimization objectives.

6

Conclusion

In this work, we examine attacks on autoregressive (AR) image generators. We find that watermarks for token-based generators [16, 26, 37] can be successfully removed even at low detection thresholds and when the service provider keeps both the model and watermarking parameters secret, although they are generally more difficult to forge. BitMark [18], in contrast, is extremely robust to removal attacks except under the strongest attacker model, but remains broadly vulnerable to forgery attacks by uninformed attackers. Furthermore, this weakness cannot be easily avoided if the detection of images generated through BitMark’s radioactivity feature is required. Since BitMark is designed to identify and exclude synthetic images during data collection, forging this watermark enables Watermark Mimicry (see Fig. 1), where authentic images can be protected from being harvested for model training. This robustness analysis provides a foundation for future research on watermarking for autoregressive image generation. The source code will be released to support future work.

On the Robustness of Watermarking for Autoregressive Image Generation

15

Acknowledgements This work was funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under Germany’s Excellence Strategy – EXC 2092 CASA – 390781972 and by the Ministry of Culture and Science of North RhineWestphalia as part of the Lamarr Fellow Network. Minh Pham would like to acknowledge support from the AI Research Institutes program supported by NSF and USDA-NIFA under AI Institute: for Resilient Agriculture, Award No. 2021-67021-35329.

References 1. Alemohammad, S., Casco-Rodriguez, J., Luzi, L., Humayun, A.I., Babaei, H., LeJeune, D., Siahkoohi, A., Baraniuk, R.: Self-consuming generative models go MAD. In: ICLR (2024) 2. An, B., Ding, M., Rabbani, T., Agrawal, A., Xu, Y., Deng, C., Zhu, S., Mohamed, A., Wen, Y., Goldstein, T., Huang, F.: WAVES: benchmarking the robustness of image watermarks. In: ICML (2024) 3. Chang, H., Zhang, H., Barber, J., Maschinot, A., Lezama, J., Jiang, L., Yang, M.H., Murphy, K.P., Freeman, W.T., Rubinstein, M., Li, Y., Krishnan, D.: Muse: Textto-image generation via masked generative transformers. In: Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., Scarlett, J. (eds.) ICML. Proceedings of Machine Learning Research, vol. 202, pp. 4055–4075. PMLR (23–29 Jul 2023) 4. Chang, H., Zhang, H., Jiang, L., Liu, C., Freeman, W.T.: Maskgit: Masked generative image transformer. In: CVPR. pp. 11315–11325 (June 2022) 5. Chern, E., Su, J., Ma, Y., Liu, P.: ANOLE: An open, autoregressive, native large multimodal models for interleaved image-text generation. arXiv:2405.06135 (2024) 6. Ci, H., Yang, P., Song, Y., Shou, M.Z.: RingID: Rethinking tree-ring watermarking for enhanced multi-key identification. In: ECCV (2024) 7. Esser, P., Rombach, R., Ommer, B.: Taming transformers for high-resolution image synthesis. In: CVPR (2021) 8. Fan, L., Li, T., Qin, S., Li, Y., Sun, C., Rubinstein, M., Sun, D., He, K., Tian, Y.: Fluid: Scaling autoregressive text-to-image generative models with continuous tokens. In: ICLR (2025) 9. Fernandez, P., Couairon, G., Jégou, H., Douze, M., Furon, T.: The stable signature: Rooting watermarks in latent diffusion models. In: CVPR (2023) 10. Fernandez, P., Souček, T., Jovanović, N., Elsahar, H., Rebuffi, S.A., Lacatusu, V., Tran, T., Mourachko, A.: Geometric image synchronization with deep watermarking. arXiv:2509.15208 (2025) 11. Gunn, S., Zhao, X., Song, D.: An undetectable watermark for generative image models. In: ICLR (2025) 12. Han, C., Li, G., Wu, J., Sun, Q., Cai, Y., Peng, Y., Ge, Z., Zhou, D., Tang, H., Zhou, H., Liu, K., Xia, S.T., Jiao, B., Jiang, D., Zhang, X., Zhu, Y.: Nextstep-1: Toward autoregressive image generation with continuous tokens at scale. In: ICLR (2026) 13. Han, J., Liu, J., Jiang, Y., Yan, B., Zhang, Y., Yuan, Z., Peng, B., Liu, X.: Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis. In: CVPR. pp. 15733–15744 (June 2025)

16

A. Müller et al.

14. Hong, S., Lee, K., Jeon, S.Y., Bae, H., Chun, S.Y.: On exact inversion of dpmsolvers. In: CVPR (June 2024) 15. Jain, A., Kobayashi, Y., Murata, N., Takida, Y., Shibuya, T., Mitsufuji, Y., Cohen, N., Memon, N., Togelius, J.: Forging and removing latent-noise diffusion watermarks using a single image. arXiv:2504.20111 (2025) 16. Jovanović, N., Labiad, I., Soucek, T., Vechev, M., Fernandez, P.: Watermarking autoregressive image generation. In: NeurIPS (2025) 17. Kassis, A., Hengartner, U.: Unmarker: a universal attack on defensive image watermarking. In: IEEE Symposium on Security and Privacy (SP). pp. 2602–2620. IEEE (2025) 18. Kerner, L., Meintz, M., Zhao, B., Boenisch, F., Dziedzic, A.: Bitmark for infinity: Watermarking bitwise autoregressive image generative models. In: NeurIPS (2025) 19. Kirchenbauer, J., Geiping, J., Wen, Y., Katz, J., Miers, I., Goldstein, T.: A watermark for large language models. In: ICML (2023) 20. Kirchenbauer, J., Geiping, J., Wen, Y., Shu, M., Saifullah, K., Kong, K., Fernando, K., Saha, A., Goldblum, M., Goldstein, T.: On the reliability of watermarks for large language models. In: ICLR (2024) 21. Kumbong, H., Liu, X., Lin, T.Y., Liu, M.Y., Liu, X., Liu, Z., Fu, D.Y., Re, C., Romero, D.W.: Hmar: Efficient hierarchical masked auto-regressive image generation. In: CVPR. pp. 2535–2544 (June 2025) 22. Li, T., Tian, Y., Li, H., Deng, M., He, K.: Autoregressive image generation without vector quantization. In: NeurIPS (2024) 23. Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft COCO: Common objects in context. In: ECCV (2014) 24. Liu, Y., Song, Y., Ci, H., Zhang, Y., Wang, H., Shou, M.Z., Bu, Y.: Image watermarks are removable using controllable regeneration from clean noise. In: ICLR (2025) 25. Lukas, N., Diaa, A., Fenaux, L., Kerschbaum, F.: Leveraging optimization for adaptive attacks on image watermarks. In: ICLR (2024) 26. Lukovnikov, D., Müller, A., Quiring, E., Fischer, A.: Clustermark: Towards robust watermarking for autoregressive image generators with visual token clustering. In: CVPR (2026) 27. Müller, A., Lukovnikov, D., Thietke, J., Fischer, A., Quiring, E.: Black-box forgery attacks on semantic watermarks for diffusion models. In: CVPR (2025) 28. Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-resolution image synthesis with latent diffusion models. In: CVPR. pp. 10684–10695 (June 2022) 29. Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A.C., Fei-Fei, L.: ImageNet large scale visual recognition challenge. IJCV 115(3), 211–252 (2015). https://doi. org/10.1007/s11263-015-0816-y 30. Shamshad, F., Lukas, N., Nandakumar, K.: Raven: Erasing invisible watermarks via novel view synthesis. In: CVPR (2026) 31. Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., Gal, Y.: Ai models collapse when trained on recursively generated data. Nature 631(8022), 755–759 (2024). https://doi.org/10.1038/s41586-024-07566-y 32. Song, J., Meng, C., Ermon, S.: Denoising diffusion implicit models. In: ICLR (2021) 33. Sun, P., Jiang, Y., Chen, S., Zhang, S., Peng, B., Luo, P., Yuan, Z.: Autoregressive model beats diffusion: Llama for scalable image generation. arXiv:2406.06525 (2024)

On the Robustness of Watermarking for Autoregressive Image Generation

17

34. Tang, H., Wu, Y., Yang, S., Xie, E., Chen, J., Chen, J., Zhang, Z., Cai, H., Lu, Y., Han, S.: HART: Efficient visual generation with hybrid autoregressive transformer. In: ICLR (2025) 35. Team, C.: Chameleon: Mixed-modal early-fusion foundation models. arXiv:2405.09818 (2025) 36. Tian, K., Jiang, Y., Yuan, Z., PENG, B., Wang, L.: Visual autoregressive modeling: Scalable image generation via next-scale prediction. In: NeurIPS (2024) 37. Tong, Y., Pan, Z., Yang, S., Zhou, K.: Training-free watermarking for autoregressive image generation. arXiv:2505.14673 (2025) 38. Wang, X., Zhang, X., Luo, Z., Sun, Q., Cui, Y., Wang, J., Zhang, F., Wang, Y., Li, Z., Yu, Q., Zhao, Y., Ao, Y., Min, X., Li, T., Wu, B., Zhao, B., Zhang, B., Wang, L., Liu, G., He, Z., Yang, X., Liu, J., Lin, Y., Huang, T., Wang, Z.: Emu3: Next-token prediction is all you need. arXiv:2409.18869 (2024) 39. Wang, Y., Lin, Z., Teng, Y., Zhu, Y., Ren, S., Feng, J., Liu, X.: Bridging continuous and discrete tokens for autoregressive visual generation. In: ICCV (2025) 40. Wen, Y., Kirchenbauer, J., Geiping, J., Goldstein, T.: Tree-Ring watermarks: Invisible fingerprints for diffusion images. In: NeurIPS (2023) 41. Wu, Y., Cui, X., Chen, R., Milis, G., Huang, H.: A watermark for auto-regressive image generation models. arXiv:2506.11371 (2025) 42. Yang, P., Ci, H., Song, Y., Shou, M.Z.: Can simple averaging defeat modern watermarks? In: Globerson, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J., Zhang, C. (eds.) NeurIPS (2024) 43. Yang, Z., Zeng, K., Chen, K., Fang, H., Zhang, W., Yu, N.: Gaussian Shading: Provable performance-lossless image watermarking for diffusion models. In: CVPR (2024) 44. Yu, Q., He, J., Deng, X., Shen, X., Chen, L.C.: Randomized autoregressive visual generation. In: ICCV (October 2025) 45. Zhang, R., Isola, P., Efros, A.A., Shechtman, E., Wang, O.: The unreasonable effectiveness of deep features as a perceptual metric. In: CVPR (2018) 46. Zhao, X., Zhang, K., Su, Z., Vasan, S., Grishchenko, I., Kruegel, C., Vigna, G., Wang, Y.X., Li, L.: Invisible image watermarks are provably removable using generative ai. NeurIPS (2024) 47. Zhao, X., Zhang, K., Su, Z., Vasan, S., Grishchenko, I., Kruegel, C., Vigna, G., Wang, Y.X., Li, L.: Invisible image watermarks are provably removable using generative AI. In: NeurIPS (2024)

Supplementary Material for On the Robustness of Watermarking for Autoregressive Image Generation

A

Full Experimental Settings

We provide full details on our experimental setup and implementation details, namely details on the watermarking schemes and generators (Sec. A.1) and deployed attacks (Sec. A.2), For reference, Table 1 summarizes details of the VQ-VAEs used in our evaluation, while Table 2 lists the models used for each watermarking scheme, proxy model setting, and corresponding box setting, i.e. the level of access and knowledge needed for launching different attacks. A.1

Details on the Watermarking Schemes and Generators

We clarify the details of each targeted watermarking scheme and model: – IndexMark [37]. We used the pre-constructed token pairs, and experimented with the GPT-B (at 256 × 256 resolution) and GPT-L (384 × 384). 100% of the red tokens were replaced, which is the strongest watermarking setting. The rest of the settings are set to default, as specified in their repository1 . With each model, we generate 1,000 watermarked images using one ImageNet [29] class each. – WMAR [16]. We used the default settings available in the official respository2 . For each of the deployed models (Anole [5] 512 × 512, Taming [7] 256 × 256, and RAR-XL [44] 256 × 256), the watermarking parameters are set to γ = 0.25 (green token fraction) and δ = 2.0 (bias strength). Every model has the stratification strategy enabled, which excludes dead tokens from green token assignment. For Taming, this leads to 971 alive tokens out of 16,384 tokens in total. For Anole, there are 57,344 alive tokens out of a total of 65,536 tokens. The RAR-XL VQ-VAE does not have dead tokens, so all of the 1,024 tokens in the vocabulary are alive. Furthermore, we use WMAR’s pretrained finetunes for each model’s VQ-VAE decoder (necessary for watermarked generation) and encoder (necessary for watermark verification) for improved reverse cycle consistency. For RAR-XL experiments with WMAR, we use the optimal generator settings reported in RAR [44] (rather than those in WMAR’s codebase) since these led to better FIDs matching the RAR paper. Finally, we do not enable WMAR’s synchronization layer for improved robustness against geometric transformations, to study 1 2

https://github.com/maifoundations/IndexMark https://github.com/facebookresearch/wmar

On the Robustness of Watermarking for Autoregressive Image Generation

19

WMAR’s robustness to non-geometric attacks in isolation. For each model, we generate 1,000 watermarked images, using 1,000 ImageNet classes (Taming, RAR-XL), or 1,000 random captions from the MS-COCO dataset [23] (Anole). – ClusterMark [26]. We obtained the code from the authors and used the default settings for the image generators LlamaGen [33] GPT-B (256 × 256) and GPT-L 384 × 384, as well as RAR-XL [44]. We used the default watermarking settings, with green fraction γ = 0.25, watermark strength δ = 5 and 64 clusters, while using their pretrained cluster predictor for watermark verification. For all three models, we generate 1,000 watermarked images each, using 1,000 ImageNet classes. – BitMark [18]. For image generation, we used the ∞-2B model [13] in the 1 million pixel (1024 × 1024) setting. We further used a watermarking strength of δ = 2, which is the moderate setting between the three settings δ ∈ {1, 2, 3} which have been reported yielding desirable tradeoffs between image quality degradation and watermark robustness. Otherwise, we used the default settings from the official respository3 , including enabling watermark embedding and detection at every scale. Watermarked images are generated using 1,000 random captions from the MS COCO dataset. A.2

Attack Dataset and Deployed Attacks.

Removal attacks are performed on 100 watermarked images, while forgery attacks use 100 cover images sampled from the MS COCO dataset. For LatentOptForgery, a single watermarked image from the corresponding watermarking scheme is used as a reference image. Frequency Injection Forgery does not require reference images. Since the evaluated watermarking schemes operate on images of different resolutions, the COCO dataset is filtered to include only images whose shortest edge is at least as large as the target resolution. These images are then resized and center-cropped accordingly to match the required input size. The evaluated attacks and their exact settings are listed below. – VQ-Regen: A detailed description is provided in Sec. A.2. – LatentOpt-Removal and -Forgery: we run optimization through the encoder of a proxy model’s VQ-VAE to obtain pre-quantization latents, i.e. the representation right before nearest neighbor matching to the closest tokens in the vocabulary. This is done for at most 300 steps with a budget of 8 and verify the resulting attack instances every 10 optimization |∆x|∞ < 255 steps against the target verifier. The Adam optimizer is used, with constant learning rate of 0.001 across all setups, with some higher learning rates for LatentOpt-Forgery attacks, as reported in Tab. 2. During experimentation, we observed that individual proxy encoders used by the attacker display different progressions in terms of attack success and quality degradation, most likely due to differences in latent sizes and value ranges of latent embeddings. 3

https://github.com/sprintml/BitMark

20

A. Müller et al.

Furthermore, individual attack instances may oscillate between detectable and non-detectable states with successive optimization steps, especially in black-box settings. This behaviour could also be addressed by tuning learning rates with a dedicated schedule for each setup. However, due to the large amount of attacker and verifier model combinations, we opt to set constant learning rates of 0.001 (except for some LatentOpt-Forgery setups, see Tab. 2 for details) and normalize for the reported values for the LatentOpt attacks by reporting the accuracy (TPR, z-score, p-value) and quality metrics (PSNR, LPIPS) achieved at the step with the highest (removal) or lowest (forgery) per attack instance within a run of 300 optimization steps. – BitOpt-Removal(BitMark-only) is tested with 100 steps, targeting the finest scales (9-12). Accuracy and quality metrics are reported for the first occurence of a p > 0.01, i.e. the first successful removal at the F P R = 1% threshold. Algorithm 1 and 2 formally describe the BitOpt attack and the watermark verification procedure in BitMark for reference, respectively. – Forgery via Frequency Injection(BitMark-only) is reported with three best settings with different trade-offs between image quality and attack success. A detailed description is provided in Sec. A.2 (see below). – Perturbations and Geometric Transformations: The different perturbations are applied individually to a watermarked target image and the average is reported. Visual examples and the parameters for each perturbation are provided in Fig. 1. – Regen4 and Rinse from [47] are tested with the default setting of 60 steps and 3 rounds for Rinse, using Stable Diffusion v2.1 [28]. – CtrlRegen+ [24] is tested with moderate settings (guidance of 2, strength 0.3) using their official repository5 , again using SD-v2.1. – BitFlipper [18](BitMark-only) assumes full access to the exact VQ-VAE deployed with the watermark verifier, as well as knowledge of watermarking settings. We used the implementation from BitMark’s official repository with minor modifications, which were necessary to bring the code in working condition. We used a strength of ϕ = 2.2, which was reported as the most effective setting in the original work.

4 5

https://github.com/XuandongZhao/WatermarkAttacker https://github.com/yepengliu/CtrlRegen

Table 1: Overview of the visual tokenizer in each model. |V| and d refer to the covabulary size and the embedding dimension, respectively. LlamaGen Anole Taming RAR |V| 16384 d 8

65536 256

16384 256

1024 256

∞-2B 216 224 232 264 16 24 32 64

On the Robustness of Watermarking for Autoregressive Image Generation IndexMark [37]

WMAR [16]

LlamaGen [33] GPT-B GPT-L

Anole [5]

LlamaGen [33] ∞-2B [13] Taming [7] RAR-XL [44] GPT-B GPT-L RAR-XL [44] (Vd = 232 )

ClusterMark [26]

256 × 256 384 × 384

512 × 512

256 × 256

256 × 256

256 × 256 384 × 384

256 × 256

21

BitMark [18]

1024 × 1024

Regen. [47] (60 steps) SD-v2.1 VAE

■□ ■ ■ ■

■ ■ ■ ■□

■ ■ ■ ■

■ ■ ■ ■□ -

■ ■ ■ ■ ■ ■ □ ■

Rinse (3× Regen. w. 60 steps) SD-v2.1 VAE

CtrlRegen+ (Guidance=2, Strength=0.3) SD-v2.1 VAE

VQ-Regen (Substitution Rank k = 2) LlamaGen Anole Taming RAR-XL

■□ ■ ■ ■

■□ ■ ■ ■

■ ■□ ■ ■

■ ■ ■□ ■

■ ■ ■ ■□

■□ ■ ■ ■

8 LatentOpt-Removal (300 steps, c = 255 , LR=0.001)

LlamaGen Anole Taming RAR-XL ∞-2B (d = 16) ∞-2B (d = 24) ∞-2B (d = 32) ∞-2B (d = 64)

■□ ■ ■ ■ -

■□ ■ ■ ■ -

■ ■□ ■ ■ -

■ ■ ■□ ■ -

LlamaGen Anole Taming RAR-XL ∞ (Vd = 216 ) ∞ (Vd = 224 ) ∞ (Vd = 232 ) ∞ (Vd = 264 )

■□ ■ ■ ■ -

■□ ■ ■ ■ -

■ LR=0.01 ■ □ LR=0.01 ■ LR=0.01 ■ LR=0.01 -

■ LR=0.01 ■ LR=0.01 ■ □ LR=0.01 ■ LR=0.01 -

∞ (Vd = 232 )

-

-

-

-

■ ■ ■ ■□ -

■□ ■ ■ ■ -

■□ ■ ■ ■ -

8 LatentOpt-Forgery (300 steps, c = 255 , LR=0.001)

■ ■ ■ ■□ -

■□ ■ ■ ■ -

■□ ■ ■ ■ -

■ ■ ■ ■□ -

■ LR=0.01 ■ LR=0.01 ■ LR=0.01 ■ LR=0.01 ■ ■ □ ■

-

-

-

□+

-

-

-

□+

-

-

BitFlipper (ϕ = 2.2) -

BitOpt (100 steps, LR=0.0005) ∞ (Vd = 232 )

-

-

-

-

-

Frequency Injection Forgery (Setting A,B,C) -

-

-

-

-

-

-

Table 2: Detailed breakdown of settings used in each attack, excluding perturbations and geometric transformations. For each watermarking scheme and deployed verifier model, the attacker model is shown on the left side. Using an unrelated model on the attacker side (or no model, as in the Frequency Injection Forgery attack) is considered a black-box setting and is indicated by ■. Using a model of similar architecture is considered a grey-box setting and is indicated by ■. For token-based watermarking schemes (IndexMark, WMAR, ClusterMark), this corresponds to using the default version of the deployed verifier model’s encoder, without acces the finetuned weights. White-box settings require full access to the exact weights of the verifier’s encoder and is indicated by □. In the case of BitMark, having additional knowledge of the green set G is indicated by □+. For the LatentOpt-Forgery attack on the Anole and Taming models deployed with WMAR, a higher learning rate (LR=0.01) was chosen. This was done in order to saturate the attack success for experiments with higher perturbation 32 (see Sec. 4.2), since the default LR=0.001 would not show budgets such as c = 255 forgery success within 300 steps. This was done to show that forgery is possible for all verifier models, albeit at the cost of extreme quality degradtion.

22

A. Müller et al.

Algorithm 1 White-Box+ BitOpt-Removal Attack Require: target image x, encoder E, quantizer Q, green list G, scale resolutions (hi , wi )K i=1 , margin γ, perturbation budget ϵ, step size α, number of steps T 1: δ ← 0 2: for t = 1, . . . , T do 3: z ← E(x + δ) 4: ẑ ← z ▷ ẑ is overall residual 5: for i = 1, . . . , K do 6: ẽi ← Interpolate(ẑ, hi , wi ) ▷ Unquantized residual 7: ui ← Q(ẽi ) ▷ Quantized residual 8: z̃i ← ẑ − Interpolate(ui , hK , wK ) ▷ z̃i is unquantized residual at scale i 9: ẑ ← ẑ − ui (i) (i) ▷ Extract bit sequence 10: (b1 , . . . , bhi ·wi ·m ) ← B(ui > 0) 11: T ←∅ 12: for i = 1, . . . , K do 13: for j ∈ {2, . . . , hi · wi · m − 1} do (i) (i) (i) 14: if (bj−1 , bj , bj+1 ) ∈ {(0, 1, 0), (1, 0, 1)} then 15: XT ← T ∪ {(i, j)}  ▷ Flipping bj reduces green count by 2 16: L← |ẽi [j] + sign ẽi [j] · γ| ▷ L1 Loss (i, j) ∈ T

17: δ ← δ − α · ∇δ L 18: δ ← clip(δ, −ϵ, ϵ) 19: return x + δ

▷ Project back to ℓ∞ ball

Algorithm 2 Watermark Detection (Alg. 2 from BitMark [18]) Inputs: raw image x, green list G, red list R, encoder E, quantizer Q Hyperparameters: steps K (number of resolutions), resolutions (hi , wi )K i=1 , the number of tokens for resolution i is ri , number of latent channels n. 1: e = E(im) 2: C = 0 3: for i = 1, . . . , K do 4: ui = Q(Interpolate(e, hi , wi )) 5: ui = (b1 , . . . , bri ·m ) 6: C = Count((b1 , . . . , bri ·m ), G) 7: zi = Lookup(ui ) 8: zi = Interpolate(zi , hK , wK ) 9: e = e − ϕi (zi ) 10: Return: StatisticalTest(H0 , (C))

On the Robustness of Watermarking for Autoregressive Image Generation Original

JPEG 10

Gauss. std 0.35

Salt&Pepper 0.15

Brightness 6

Contrast 4

Hue 0.5

Saturation 5

23

(a) Visual examples of the perturbations.

Original

Rotation 5°

Rotation 7°

Random Crop 5%

Random Crop 10%

(b) Visual examples for the geometric transformations.

Fig. 1: Visual examples of perturbations and geometric transformations.

VQ-Regen Algorithm and Effect of Different VQ-VAEs and Substitution Ranks. Algorithm 3 provides a formal description of the VQ-Regen attack. Examples of VQ-Regen applied to watermarked images (generated by BitMark) are provided in Fig. 2. From the figure, we observe that the choice of VAE and quantizer plays a large role in the VQ-Regen attack. A VQ-VAE with a larger vocabulary and a smaller embedding dimension (d = 8 for the first row) appears to have a much better reconstruction while larger embedding dimensions (d = 256 for the bottom three rows) results in lower image quality. With a lower vocabulary size (|V| = 1024 in the last row), the reconstruction quality further suffers. The effect of vocabulary size is straightforward: with fewer available tokens, the second closest token used in the VQ-Regen attack will be further away from the original token in latent space, resulting in more change. The effect of latent dimension d is less obvious. We speculate that the observed results are caused by (1) Euclidean distances being larger with a growing number of dimensions and (2) the distances between tokens becoming less spread out with growing space dimension, and thus the ranks becoming more sensitive to randomness. A possible improvement to the VQ-Regen attack is to also involve the attacker backbone in order to use not just the n-th closest token, but also take into account which of the n closest tokens would be more consistent with the preceding tokens and would result in better image quality.

<latexit sha1_base64="tOYzeM52sSuDNJqYcUdBdb1jaC8=">AAAB/nicbVDLSgMxFL1TX7W+RsWVm2ARXJWZUqsIhaIblxVsK7RDyWQybWjmQZIRyrTgr7hxoYhbv8Odf2PazkJbz+XC4Zx7yc1xY86ksqxvI7eyura+kd8sbG3v7O6Z+wctGSWC0CaJeCQeXCwpZyFtKqY4fYgFxYHLadsd3kz99iMVkkXhvRrF1AlwP2Q+I1hpqWcejVvjmm2VK92reXm18nm1ZxatkjUDWiZ2RoqQodEzv7peRJKAhopwLGXHtmLlpFgoRjidFLqJpDEmQ9ynHU1DHFDppLPzJ+hUKx7yI6E7VGim/t5IcSDlKHD1ZIDVQC56U/E/r5Mo/9JJWRgnioZk/pCfcKQiNM0CeUxQovhIE0wE07ciMsACE6UTK+gQ7MUvL5NWuWRXS9W7SrF+ncWRh2M4gTOw4QLqcAsNaAKBFJ7hFd6MJ+PFeDc+5qM5I9s5hD8wPn8AwyGTcA==</latexit>

RAR-XL VQ-VAE |V | = 1024 d = 256 <latexit sha1_base64="V/FztQjgoZoyy+pzYEBaupBuspQ=">AAAB/3icbVDLSsNAFJ3UV62vqODGzWARXJWk1liEQtGNywq2FtpQJpNJO3QyCTMToaRd+CtuXCji1t9w5984bbPQ1nO5cDjnXubO8WJGpbKsbyO3srq2vpHfLGxt7+zumfsHLRklApMmjlgk2h6ShFFOmooqRtqxICj0GHnwhjdT/+GRCEkjfq9GMXFD1Oc0oBgpLfXMo3FrXLOd82qlezUvv1a+cHpm0SpZM8BlYmekCDI0euZX149wEhKuMENSdmwrVm6KhKKYkUmhm0gSIzxEfdLRlKOQSDed3T+Bp1rxYRAJ3VzBmfp7I0WhlKPQ05MhUgO56E3F/7xOooKqm1IeJ4pwPH8oSBhUEZyGAX0qCFZspAnCgupbIR4ggbDSkRV0CPbil5dJq1yynZJzVynWr7M48uAYnIAzYINLUAe3oAGaAIMxeAav4M14Ml6Md+NjPpozsp1D8AfG5w9NdJO5</latexit>

Taming VQ-VAE |V | = 16384 d = 256 <latexit sha1_base64="fYyOEeAHrdYJbkBPXjHPCoI1EQc=">AAAB/3icbVDLSsNAFJ3UV62vqODGzWARXJWk2ihCoejGZQX7gDaUyWTSDp1MwsxEKGkX/oobF4q49Tfc+TdO2yy09VwuHM65l7lzvJhRqSzr28itrK6tb+Q3C1vbO7t75v5BU0aJwKSBIxaJtockYZSThqKKkXYsCAo9Rlre8Hbqtx6JkDTiD2oUEzdEfU4DipHSUs88GjfHVadSOXe61/Pyq+WK0zOLVsmaAS4TOyNFkKHeM7+6foSTkHCFGZKyY1uxclMkFMWMTArdRJIY4SHqk46mHIVEuuns/gk81YoPg0jo5grO1N8bKQqlHIWengyRGshFbyr+53USFVy5KeVxogjH84eChEEVwWkY0KeCYMVGmiAsqL4V4gESCCsdWUGHYC9+eZk0yyXbKTn3F8XaTRZHHhyDE3AGbHAJauAO1EEDYDAGz+AVvBlPxovxbnzMR3NGtnMI/sD4/AFSN5O8</latexit>

Anole VQ-VAE |V | = 65536 d = 256 <latexit sha1_base64="WiYmvSloX2sUb/pvo+hUyD6xShs=">AAAB/XicbVDLSsNAFL2pr1pf8bFzEyyCq5JoqUUoFN24rGAf0IYymUzaoZNJmJkINS3+ihsXirj1P9z5N04fC209lwuHc+5l7hwvZlQq2/42Miura+sb2c3c1vbO7p65f9CQUSIwqeOIRaLlIUkY5aSuqGKkFQuCQo+Rpje4mfjNByIkjfi9GsbEDVGP04BipLTUNY9GjVHFKV2Ui52rWfmVctfM2wV7CmuZOHOShzlqXfOr40c4CQlXmCEp244dKzdFQlHMyDjXSSSJER6gHmlrylFIpJtOrx9bp1rxrSASurmypurvjRSFUg5DT0+GSPXlojcR//PaiQrKbkp5nCjC8eyhIGGWiqxJFJZPBcGKDTVBWFB9q4X7SCCsdGA5HYKz+OVl0jgvOKVC6a6Yr17P48jCMZzAGThwCVW4hRrUAcMjPMMrvBlPxovxbnzMRjPGfOcQ/sD4/AFeUJNA</latexit>

LlamaGen VQ-VAE |V | = 16384 d = 8

24 A. Müller et al.

Original Watermarked Subst. Rank k=2 Subst. Rank k=3 Subst. Rank k=4 Subst. Rank k=5 Subst. Rank k=6

Fig. 2: Visual examples of the VQ-Regen attack. The original watermarked image is generated using BitMark deployed with ∞-2B.

Algorithm 3 Vector-Quantized Regeneration Attack (VQ-Regen)

Require: Image x; encoder E; decoder D; codebook C ∈ R|V|×d ; substitution rank k with 1 ≤ k ≤ |V| 1: z ← E(x) ▷ z ∈ Rd×h×w ′ h×w 2: Initialize t ∈ V 3: for i ← 1 to h do 4: for j ← 1 to w do 5: u ← z:,i,j ∈ Rd  6: s ← Argsort( ∥ u − C[v, :] ∥2 v∈V ) ▷ s[1] is nearest, s[k] is k-th nearest 7: t′i,j ← s[k] ′ 8: z ′ ← C[t′ ] ▷ lookup: z:,i,j = C[t′i,j , :] ′ ′ 9: x ← D(z ) 10: return x′

On the Robustness of Watermarking for Autoregressive Image Generation

25

Forgery via Frequency Injection. As explained in Sec. 3.3, averaging many BitMarked images reveals a pixel pattern, appearing as a regular grid of high magnitude peaks in the frequency domain, which verifies against the BitMark verifier with a p-value of 3e − 33. Based on this observation, we tried to recreate this effect by introducing a similar pattern via frequency injection into authentic cover images. For this, we ran a set of trials where magnitude peak locations are controlled by different step sizes s ∈ {16, 32, 64, 128} by either (i) arranging them in a full lattice with these spacings like in the FFT representation of the averaged image, or (ii) restricting the peaks to the main diagonals to minimize quality degradation. We also varied the magnitude across ln(α) ∈ {7.5, 7.75, 8.0, 8.25, 8.5} and restricted the peaks to the first n occurrences closest to the center. We then evaluated the accuracy metrics (p-value produced by the BitMark verifier) and visual quality metrics (PSNR) on 100 images for each setting, filtering out any result above p = 1e − 5 and sorting by PSNR in descending order. Finally, we chose three desirable tradeoffs (settings A, B and C). Note that this procedure is not guaranteed to yield optimal results and there may be better tradeoffs. Fig. 3 shows the settings for our forgery attack via frequency injection. The attack is performed by injecting magnitude peaks of strength α along the main diagonals in regular intervals (32 frequency bins in both axes) with random phases for each color channel’s FFT. Setting A and B are limited to the first four lower frequency bins, while setting C includes all frequency bins along the main diagonals. Setting A uses a smaller magnitude of ln(α) = 7.75, while settings B and C use a magnitude of ln(α) = 8.0. With this, we achieve different tradeoffs between the shift towards positive BitMark detection (lower p-values as determined by the BitMark verifier), and the quality degradation in terms of PSNR. Algorithm 4 provides a formal description of the injection procedure. Fig. 4 shows the effect of frequency injection on the BitMark verifier: the periodic patterns trigger the ∞-2B multi-scale VAE to observe a high count of green bigrams G = {01, 10} in the last (finest) scale. Fig. 4 also shows a breakdown of the number of bits on each scale during generation: approximately 40% of all bits are located in the last scale. While one could counter our attack by disabling watermark embedding and verification at the finest scale entirely, this would forfeit 40% of the available embedding capacity, substantially reducing BitMark’s robustness. Alternative strategies for detecting scale imbalances may exist, but evaluating their effectiveness in the context of other attacks like removal attempts is beyond the scope of this work.

26

A. Müller et al. Original Cover Image

Setting A

Setting B

ln(↵) = 7.75

ln(↵) = 8.0

<latexit sha1_base64="16H4td35kgdg8t5+DR55CRPVSJE=">AAAB+nicbVBNS8NAEN34WetXqkcvi0Wol5CItl6EohePFewHtKFMttt26WYTdjdKif0pXjwo4tVf4s1/47bNQVsfDDzem2FmXhBzprTrflsrq2vrG5u5rfz2zu7evl04aKgokYTWScQj2QpAUc4ErWumOW3FkkIYcNoMRjdTv/lApWKRuNfjmPohDATrMwLaSF27wEWpAzwewim+whWnctG1i67jzoCXiZeRIspQ69pfnV5EkpAKTTgo1fbcWPspSM0Ip5N8J1E0BjKCAW0bKiCkyk9np0/wiVF6uB9JU0Ljmfp7IoVQqXEYmM4Q9FAtelPxP6+d6P6lnzIRJ5oKMl/UTzjWEZ7mgHtMUqL52BAgkplbMRmCBKJNWnkTgrf48jJpnDle2SnfnRer11kcOXSEjlEJeaiCqugW1VAdEfSIntErerOerBfr3fqYt65Y2cwh+gPr8weEppI7</latexit>

Setting C

<latexit sha1_base64="AjQqYHLtPLIpYpbQt6gPUXDfSJo=">AAAB+XicbVBNS8NAEN3Ur1q/oh69LBahXkIiUnsRil48VrAf0IYy2W7apZtN2N0USug/8eJBEa/+E2/+G7dtDtr6YODx3gwz84KEM6Vd99sqbGxube8Ud0t7+weHR/bxSUvFqSS0SWIey04AinImaFMzzWknkRSigNN2ML6f++0JlYrF4klPE+pHMBQsZAS0kfq2zUWlBzwZwSW+xTXH7dtl13EXwOvEy0kZ5Wj07a/eICZpRIUmHJTqem6i/QykZoTTWamXKpoAGcOQdg0VEFHlZ4vLZ/jCKAMcxtKU0Hih/p7IIFJqGgWmMwI9UqveXPzP66Y6rPkZE0mqqSDLRWHKsY7xPAY8YJISzaeGAJHM3IrJCCQQbcIqmRC81ZfXSevK8apO9fG6XL/L4yiiM3SOKshDN6iOHlADNRFBE/SMXtGblVkv1rv1sWwtWPnMKfoD6/MHAXWR9g==</latexit>

<latexit sha1_base64="AjQqYHLtPLIpYpbQt6gPUXDfSJo=">AAAB+XicbVBNS8NAEN3Ur1q/oh69LBahXkIiUnsRil48VrAf0IYy2W7apZtN2N0USug/8eJBEa/+E2/+G7dtDtr6YODx3gwz84KEM6Vd99sqbGxube8Ud0t7+weHR/bxSUvFqSS0SWIey04AinImaFMzzWknkRSigNN2ML6f++0JlYrF4klPE+pHMBQsZAS0kfq2zUWlBzwZwSW+xTXH7dtl13EXwOvEy0kZ5Wj07a/eICZpRIUmHJTqem6i/QykZoTTWamXKpoAGcOQdg0VEFHlZ4vLZ/jCKAMcxtKU0Hih/p7IIFJqGgWmMwI9UqveXPzP66Y6rPkZE0mqqSDLRWHKsY7xPAY8YJISzaeGAJHM3IrJCCQQbcIqmRC81ZfXSevK8apO9fG6XL/L4yiiM3SOKshDN6iOHlADNRFBE/SMXtGblVkv1rv1sWwtWPnMKfoD6/MHAXWR9g==</latexit>

ln(↵) = 8.0

Frequency component injection with random phase and magnitude 𝛼 In regular intervals

32 32 32

p = 0.645 | PSNR = inf dB

p = 4e-05 | PSNR = 38.7 dB

p = 3e-27 | PSNR = 36.3 dB

p = 1e-191 | PSNR = 30.0 dB

3k 2k 1k 0

Green Bigrams per Generation Scale Real cover images Frequency injected images 0 1 2 3 4 5 6 7 8 9 10 11 12 Generation Scale (a) Difference

Number of Bits

Green

Fig. 3: BitMark Forgery via Frequency Injection Settings. We take an authentic cover image, and then compute its FFT. We inject spectral components with magnitude α and a channel-wise random phase along the diagonals and spaced 32 pixels apart in both axes, directly overwriting the original coefficients. Finally, we apply inverse FFT to reconstruct the attacked image. Settings A and B are limited to the first four frequency bins along the diagonals (lower frequencies). Setting C modifies all frequency bins along the diagonals. Setting A uses a low magnitude of ln(α) = 7.75, while settings B and C use a higher magnitude ln(α) = 8.0.

40000k

Number of Bits per Scale Generation Scale 0-11: ~60% of Bits Generation Scale 12: ~40% of Bits

20000k 0k

0 1 2 3 4 5 6 7 8 9 10 11 12 Generation Scale (b) Bits per scale

Fig. 4: Total difference ∆ between green (G = {01, 10}) and red ((R = {00, 11})) bigrams counts across different scales for 100 real cover images (blue) and 100 frequency injection forgery attack instances (setting A, orange) as verified by the BitMark watermarking schemes deployed with ∞-2B (left). At the finest generation scale (scale 12), the attack introduces a strong surplus of green bigrams G = 01, 10, causing the verifier to detect the presence of the watermark. The amount of bits across each generation scale of ∞-2B is shown to the right. The scale most affected by the attack (scale 12) accounts for roughly 40% of all embedded bits.

On the Robustness of Watermarking for Autoregressive Image Generation

27

Algorithm 4 Frequency Injection in the FFT Domain Require: image x ∈ R3×H×W , magnitude α, spacing s = 32 Ensure: forged image x′ 1: X ← fftshift(F(x)) ▷ FFT per channel 2: M ← diagonal mask with spacing d in the right half of the spectrum ▷ points (cy ± kd, cx + kd) only (right-side quadrants) 3: for channel c ∈ {1, 2, 3} do 4: for (y, x) with M [y, x] = 1 do 5: sample ϕ ∼ U(0, 2π) 6: A ← αeiϕ ▷ fixed magnitude, random phase 7: Xc [y, x] ← Xc [y, x] + A 8: (y ′ , x′ ) ← ((H − y) mod H, (W − x) mod W ) 9: Xc [y ′ , x′ ] ← Xc [y ′ , x′ ] + A ▷ write conjugate to opposite quadrant to enforce Hermitian symmetry 10: x′ ← ℜ F −1 (ifftshift(X)) ▷ real output guaranteed by Hermitian symmetry 11: return x′

B

Full Experimental Results

Figs. 5 to 7 show more visual examples for every attack. Fig. 8 shows the full results for LatentOpt-Removal and -Forgery attacks using different perturbation budgets c, broken down for each deployed model for every watermarking scheme. Tables 3 to 11 show a detailed breakdown of all experimental results for all verifiers, their deployed models, and every attack setting, including individual proxy models used by the attacker in terms of mean TPR@FPR=1%, median p-value, as well as mean with standard deviation for both PSNR and LPIPS. For BitMark, we also show the mean z-score with standard deviation. Again, removal attack aim for lower TPR, z-score, and higher p-value, while forgery attacks aim for the opposite. All attacks aim for higher PSNR and lower LPIPS.

28

A. Müller et al.

WMAR (Anole)

WMAR (Taming)

ClusterMark (RAR-XL)

ClusterMark (LlamaGen GPT-L)

BitMark (∞-2B)

CtrlRegen+ (SD-v2.1)

Rinse (SD-v2.1)

VQ-Regen VQ-Regen Regen. (SD-v2.1) (RAR-XL VQ-VAE) Taming VQ-VAE

VQ-Regen Anole VQ-VAE

VQ-Regen LlamaGen VQ-VAE

Original Watermarked

IndexMark (LlamaGen GPT-B)

Fig. 5: Visual examples of VQ-Regen and diffusion regeneration-based attacks.

On the Robustness of Watermarking for Autoregressive Image Generation WMAR (Anole)

WMAR (Taming)

ClusterMark (RAR-XL)

ClusterMark (LlamaGen GPT-L)

BitMark (∞-2B)

LatentOpt RAR-XL VAE

LatentOpt Taming VAE

LatentOpt Anole VAE

LatentOpt LlamaGen VAE

Original Watermarked

IndexMark (LlamaGen GPT-B)

29

Original Watermarked

LatentOpt ∞-2B VAE (d=16)

LatentOpt ∞-2B VAE (d=24)

LatentOpt ∞-2B VAE (d=32)

LatentOpt ∞-2B VAE (d=64)

BitOpt ∞-2B VAE (d=32)

BitFlipper ∞-2B VAE (d=32)

8 Fig. 6: Visual examples of the LatentOpt-Removal attack (c = 255 at step 300).

Authentic Cover Image

RAR-XL VAE

LatentOpt LlamaGen VAE

LatentOpt Anole VAE

∞-2B VAE (Vd=216) ∞-2B VAE (Vd=224) ∞-2B VAE (Vd=232) ∞-2B VAE (Vd=264)

LatentOpt

Taming VAE

Frequency Injection Frequency Injection Frequency Injection Setting A Setting B Setting C

Fig. 7: Visual examples of the Frequency Injection Forgery and LatentOpt8 Forgery (c = 255 at step 300) attacks on BitMark.

A. Müller et al.

0.4

Perturbations JPEG Gaussian Noise Salt & Pepper Gaussian Blur

0.2 0.0

45

40

35

25

PSNR (dB)

15

0.6 0.4

JPEG Gaussian Noise Salt & Pepper Gaussian Blur 45

40

35

30

25

PSNR (dB)

20

15

10

0.6 0.4

Perturbations JPEG Gaussian Noise Salt & Pepper Gaussian Blur

0.2 0.0

45

40

35

30

25

PSNR (dB)

20

15

JPEG Gaussian Noise Salt & Pepper Gaussian Blur 45

40

35

30

25

20

PSNR (dB)

15

0.4

JPEG Gaussian Noise Salt & Pepper Gaussian Blur 45

40

35

30

25

20

PSNR (dB)

15

10

0.4

Perturbations JPEG Gaussian Noise Salt & Pepper Gaussian Blur

0.2 0.0

10

Watermarked =1.00 LatentOpt-Removal Chameleon RAR-XL Taming LlamaGen LlamaGen

0.6

45

40

35

30

25

PSNR (dB)

20

15

0.4

Perturbations

0.2 45

40

35

25

PSNR (dB)

10

Watermarked =1.00 LatentOpt-Removal LlamaGen Chameleon Taming RAR-XL RAR-XL

0.6 0.4

Perturbations JPEG Gaussian Noise Salt & Pepper Gaussian Blur

0.2 45

40

35

30

25

PSNR (dB)

20

15

10

ClusterMark (RAR-XL)

Watermarked =1.00 LatentOpt-Removal LlamaGen Chameleon Taming RAR-XL RAR-XL

0.8 0.6 0.4

Perturbations JPEG Gaussian Noise Salt & Pepper Gaussian Blur

0.2 0.0

JPEG Gaussian Noise Salt & Pepper Gaussian Blur 20 15

WMAR (RAR-XL)

1.0

10

30

0.8

0.0

ClusterMark (LlamaGen GPT-L)

0.8

0.6

1.0

Perturbations

0.2

Watermarked =1.00 LatentOpt-Removal LlamaGen Chameleon Taming RAR-XL -2B -2B

0.8

0.0

10

Watermarked =1.00 LatentOpt-Removal LlamaGen Chameleon RAR-XL Taming Taming

0.6

BitMark ( -2B)

1.0

WMAR (Taming)

0.8

1.0

TPR @ FPR=1%

TPR @ FPR=1%

Watermarked =1.00 LatentOpt-Removal Chameleon RAR-XL Taming LlamaGen LlamaGen

0.8

Perturbations

0.2

0.0

ClusterMark (LlamaGen GPT-B)

1.0

0.4

1.0

Perturbations

0.2

Watermarked =1.00 LatentOpt-Removal Chameleon RAR-XL Taming LlamaGen LlamaGen

0.6

0.0

10

Watermarked =0.99 LatentOpt-Removal LlamaGen RAR-XL Taming Anole Anole

0.8

0.0

20

WMAR (Anole)

1.0

TPR @ FPR=1%

30

IndexMark (LlamaGen GPT-L)

0.8

TPR @ FPR=1%

0.6

1.0

TPR @ FPR=1%

Watermarked =1.00 LatentOpt-Removal Chameleon RAR-XL Taming LlamaGen LlamaGen

TPR @ FPR=1%

IndexMark (LlamaGen GPT-B)

0.8

TPR @ FPR=1%

TPR @ FPR=1%

1.0

TPR @ FPR=1%

30

45

40

35

30

25

PSNR (dB)

20

15

10

(a) Removal

45

40

35

30

25

PSNR (dB)

Watermarked =0.99 LatentOpt-Forgery LlamaGen RAR-XL Taming Anole Anole

0.4 0.2 0.0

45

40

35

30

PSNR (dB)

25

0.0

Watermarked =1.00 LatentOpt-Forgery Chameleon RAR-XL Taming LlamaGen LlamaGen 45

40

35

30

PSNR (dB)

40

25

20

35

30

25

PSNR (dB)

WMAR (Taming) Watermarked =1.00 LatentOpt-Forgery LlamaGen Chameleon RAR-XL Taming Taming

0.4 0.2 45

40

35

30

PSNR (dB)

25

Watermarked =1.00 LatentOpt-Forgery Chameleon RAR-XL Taming LlamaGen LlamaGen

0.4 0.2 0.0

45

40

35

30

PSNR (dB)

0.4 0.2 45

40

25

20

35

30

25

PSNR (dB)

20

WMAR (RAR-XL)

0.8 Watermarked =1.00 LatentOpt-Forgery LlamaGen Chameleon Taming RAR-XL RAR-XL

0.6 0.4 0.2 0.0

20

0.8 0.6

Watermarked =1.00 LatentOpt-Forgery LlamaGen Chameleon Taming RAR-XL -2B -2B

0.6

1.0

0.8 0.6

0.8

0.0

20

45

40

35

30

PSNR (dB)

25

20

ClusterMark (RAR-XL)

1.0

1.0

TPR @ FPR=1%

TPR @ FPR=1%

0.8

0.2

45

ClusterMark (LlamaGen GPT-L)

ClusterMark (LlamaGen GPT-B)

0.4

0.2

0.0

20

1.0

0.6

0.4

1.0

0.8 0.6

Watermarked =1.00 LatentOpt-Forgery Chameleon RAR-XL Taming LlamaGen LlamaGen

0.6

0.0

20

WMAR (Anole)

1.0

0.8

BitMark ( -2B)

1.0

TPR @ FPR=1%

0.2

IndexMark (LlamaGen GPT-L)

TPR @ FPR=1%

Watermarked =1.00 LatentOpt-Forgery Chameleon RAR-XL Taming LlamaGen LlamaGen

0.4

0.0

TPR @ FPR=1%

TPR @ FPR=1%

0.8 0.6

1.0

TPR @ FPR=1%

IndexMark (LlamaGen GPT-B)

TPR @ FPR=1%

TPR @ FPR=1%

1.0

0.8 Watermarked =1.00 LatentOpt-Forgery LlamaGen Chameleon Taming RAR-XL RAR-XL

0.6 0.4 0.2 0.0

45

40

35

30

PSNR (dB)

25

(b) Forgery 2 4 8 16 32 Fig. 8: LatentOpt attacks for perturbation budgets c ∈ { 255 , 255 , 255 , 255 , 255 }.

20

On the Robustness of Watermarking for Autoregressive Image Generation

31

Table 3: Full experimental results (IndexMark deployed w. LlamaGen GPT-B) TPR P-Value

PSNR↑

1.00 8.6×10−78

Watermarked

LPIPS↓ -

-

Removal 0.57 4.3×10−3 18.358±3.186 0.107±0.043 0.50 1.4×10−2 16.882±7.753 0.430±0.371

Geometric Transf. Perturbations ■ ■ ■

Regen. Rinse CtrlRegen+ VQ-Regen

(aggregated) ■ RAR-XL ■ Anole ■ Taming ■ LlamaGen LlamaGen

LatentOpt-R (aggregated) RAR-XL Anole Taming LlamaGen LlamaGen

0.35 5.9×10−2 21.828±2.984 0.138±0.047 0.02 4.3×10−1 18.890±2.252 0.336±0.078 0.43 1.9×10−2 23.622±2.658 0.130±0.046 0.02 0.01 0.02 0.03

4.3×10−1 4.8×10−1 3.8×10−1 4.3×10−1

19.484±3.240 18.088±2.677 20.374±3.362 19.989±3.170

0.159±0.051 0.179±0.050 0.147±0.050 0.152±0.046

■ □

0.02 7.1×10−1 21.496±2.792 0.106±0.033 0.01 1.0 20.997±2.688 0.122±0.037

■ ■ ■ ■

0.62 0.78 0.58 0.50

■ □

0.10 5.7×10−1 31.373±0.947 0.124±0.060 0.00 1.0 31.474±0.917 0.032±0.020

■ ■ ■ ■

0.10 0.10 0.07 0.14

■ □

0.14 7.5×10−2 34.359±3.235 0.057±0.034 1.00 1.2×10−64 32.696±1.065 0.028±0.013

4.4×10−4 1.5×10−8 5.7×10−4 1.2×10−2

32.125±3.350 31.940±1.418 32.919±5.496 31.515±0.790

0.098±0.047 0.101±0.038 0.091±0.053 0.101±0.051

Forgery LatentOpt-F (aggregated) RAR-XL Anole Taming LlamaGen LlamaGen

9.5×10−2 9.5×10−2 9.5×10−2 9.5×10−2

34.687±3.614 34.518±3.625 35.175±3.776 34.370±3.417

0.062±0.038 0.067±0.041 0.059±0.039 0.061±0.036

Table 4: Full experimental results (IndexMark deployed w. LlamaGen GPT-L) TPR P-Value 1.00 4.0×10−174

Watermarked

PSNR↑

LPIPS↓ -

-

Removal 0.59 5.5×10−3 18.408±3.574 0.133±0.056 0.55 1.2×10−3 17.366±8.446 0.459±0.399

Geometric Transf. Perturbations ■ ■ ■

Regen. Rinse CtrlRegen+ VQ-Regen

(aggregated) ■ RAR-XL ■ Anole ■ Taming ■ LlamaGen LlamaGen

LatentOpt-R (aggregated) RAR-XL Anole Taming LlamaGen LlamaGen

0.68 1.2×10−3 23.994±3.639 0.114±0.043 0.04 3.4×10−1 21.395±2.857 0.245±0.070 0.29 7.8×10−2 23.174±2.801 0.154±0.054 0.04 0.02 0.07 0.04

3.9×10−1 5.2×10−1 2.7×10−1 3.5×10−1

21.160±3.673 19.555±3.057 22.187±3.809 21.736±3.563

0.147±0.051 0.168±0.051 0.134±0.047 0.139±0.046

■ □

0.01 8.9×10−1 23.108±3.102 0.098±0.034 0.00 1.0 22.616±2.922 0.110±0.036

■ ■ ■ ■

0.64 0.83 0.56 0.53

■ □

0.11 7.1×10−1 31.966±1.288 0.178±0.079 0.00 1.0 32.150±1.126 0.054±0.034

■ ■ ■ ■

0.09 0.08 0.07 0.13

■ □

0.14 1.1×10−1 34.951±3.514 0.061±0.037 1.00 4.2×10−153 33.000±1.127 0.033±0.014

3.1×10−5 31.753±1.137 6.4×10−11 32.329±1.407 1.1×10−3 31.426±0.982 3.5×10−3 31.505±0.676

0.148±0.071 0.138±0.052 0.152±0.079 0.154±0.079

Forgery LatentOpt-F (aggregated) RAR-XL Anole Taming LlamaGen LlamaGen

1.1×10−1 1.1×10−1 1.2×10−1 1.1×10−1

34.823±3.577 33.618±2.673 36.118±3.872 34.735±3.652

0.074±0.048 0.088±0.044 0.062±0.052 0.072±0.046

32

A. Müller et al. Table 5: Full experimental results (WMAR deployed w. Anole) TPR P-Value

PSNR↑

0.99 9.1×10−50

Watermarked

LPIPS↓ -

-

Removal 0.30 5.1×10−2 17.426±2.889 0.146±0.067 0.41 5.2×10−2 17.053±8.170 0.426±0.391

Geometric Transf. Perturbations ■ ■ ■

Regen. Rinse CtrlRegen+ VQ-Regen

(aggregated) ■ LlamaGen ■ RAR-XL ■ Taming ■ Anole Anole

LatentOpt-R (aggregated) LlamaGen RAR-XL Taming Anole Anole

0.66 1.9×10−3 26.904±3.022 0.058±0.026 0.04 2.5×10−1 24.144±2.503 0.141±0.056 0.17 9.4×10−2 25.274±3.051 0.077±0.032 0.21 0.53 0.01 0.09

1.4×10−1 8.0×10−3 4.6×10−1 1.4×10−1

22.678±3.409 24.132±3.117 20.698±2.930 23.203±3.202

0.094±0.043 0.071±0.030 0.122±0.045 0.090±0.036

■ □

0.52 1.1×10−2 25.553±3.456 0.060±0.028 0.20 6.9×10−2 24.358±3.738 0.068±0.029

■ ■ ■ ■

0.22 0.03 0.49 0.13

■ □

0.01 9.3×10−1 32.435±2.413 0.158±0.057 0.01 9.1×10−1 34.044±3.513 0.117±0.058

■ ■ ■ ■

0.00 0.00 0.00 0.00

■ □

0.00 4.5×10−1 31.271±0.233 0.135±0.064 0.07 3.0×10−1 31.088±0.249 0.109±0.050

3.9×10−1 8.0×10−1 1.9×10−2 3.9×10−1

32.636±1.445 32.350±1.399 33.126±1.463 32.431±1.357

0.155±0.055 0.178±0.057 0.130±0.035 0.156±0.060

Forgery LatentOpt-F (aggregated) LlamaGen RAR-XL Taming Anole Anole

3.5×10−1 3.4×10−1 3.4×10−1 3.6×10−1

31.341±0.338 31.430±0.363 31.251±0.378 31.343±0.235

0.128±0.058 0.120±0.055 0.136±0.058 0.128±0.061

Table 6: Full experimental results (WMAR deployed w. Taming) TPR P-Value 1.00 7.4×10−58

Watermarked

PSNR↑

LPIPS↓ -

-

Removal 0.47 1.9×10−2 20.123±4.256 0.100±0.045 0.39 4.7×10−2 17.229±8.143 0.415±0.346

Geometric Transf. Perturbations ■ ■ ■

Regen. Rinse CtrlRegen+ VQ-Regen

(aggregated) ■ LlamaGen ■ RAR-XL ■ Anole ■ Taming Taming

LatentOpt-R (aggregated) LlamaGen RAR-XL Anole Taming Taming

0.06 2.4×10−1 22.850±3.959 0.139±0.068 0.00 4.7×10−1 19.914±2.895 0.361±0.109 0.13 2.2×10−1 23.671±3.530 0.124±0.063 0.06 0.11 0.02 0.04

3.3×10−1 1.6×10−1 4.8×10−1 3.4×10−1

21.418±4.182 22.330±4.003 19.842±3.662 22.083±4.392

0.114±0.051 0.089±0.037 0.141±0.051 0.111±0.049

■ □

0.02 4.4×10−1 22.267±3.714 0.104±0.035 0.02 5.2×10−1 22.684±3.494 0.108±0.038

■ ■ ■ ■

0.33 0.18 0.52 0.30

■ □

0.01 9.1×10−1 32.289±1.773 0.121±0.066 0.00 9.7×10−1 31.347±0.996 0.127±0.061

■ ■ ■ ■

0.12 0.13 0.09 0.14

■ □

0.17 5.7×10−2 31.191±0.233 0.102±0.045 0.24 3.5×10−2 31.052±0.285 0.099±0.041

1.9×10−1 5.3×10−1 4.5×10−3 3.2×10−1

32.227±1.859 31.961±1.360 32.765±1.741 31.956±2.263

0.122±0.060 0.141±0.069 0.109±0.041 0.117±0.063

Forgery LatentOpt-F (aggregated) LlamaGen RAR-XL Anole Taming Taming

7.5×10−2 7.8×10−2 8.1×10−2 6.2×10−2

31.181±0.301 31.264±0.309 31.101±0.333 31.178±0.233

0.106±0.047 0.098±0.044 0.109±0.045 0.111±0.050

On the Robustness of Watermarking for Autoregressive Image Generation

33

Table 7: Full experimental results (WMAR deployed w. RAR-XL) TPR P-Value

PSNR↑

1.00 3.8×10−36

Watermarked

LPIPS↓ -

-

Removal 0.86 2.0×10−7 18.850±3.139 0.106±0.040 0.50 1.3×10−2 17.148±8.353 0.442±0.379

Geometric Transf. Perturbations ■ ■ ■

Regen. Rinse CtrlRegen+ VQ-Regen

(aggregated) ■ LlamaGen ■ Anole ■ Taming ■ RAR-XL RAR-XL

LatentOpt-R (aggregated) LlamaGen Anole Taming RAR-XL RAR-XL

0.36 4.7×10−2 22.507±2.985 0.147±0.044 0.01 3.6×10−1 19.281±2.193 0.364±0.070 0.36 2.6×10−2 23.917±2.427 0.139±0.045 0.09 0.20 0.03 0.03

2.3×10−1 7.0×10−2 3.3×10−1 3.5×10−1

20.672±2.769 21.013±2.616 20.687±2.888 20.317±2.755

0.147±0.044 0.127±0.036 0.152±0.045 0.162±0.043

■ □

0.01 4.5×10−1 18.579±2.371 0.167±0.041 0.01 5.1×10−1 18.249±2.269 0.204±0.046

■ ■ ■ ■

0.71 0.60 0.78 0.76

■ □

0.31 1.3×10−1 31.910±1.123 0.108±0.036 0.00 9.3×10−1 31.450±1.378 0.131±0.044

■ ■ ■ ■

0.07 0.06 0.08 0.08

■ □

0.13 1.1×10−1 34.313±3.318 0.069±0.042 0.15 8.3×10−2 33.627±3.360 0.079±0.047

2.0×10−5 1.1×10−3 1.7×10−6 6.7×10−7

31.527±0.891 31.508±0.934 31.701±0.973 31.373±0.724

0.111±0.048 0.123±0.052 0.104±0.044 0.107±0.046

Forgery LatentOpt-F (aggregated) LlamaGen Anole Taming RAR-XL RAR-XL

1.3×10−1 1.4×10−1 1.2×10−1 1.3×10−1

34.795±3.705 35.101±3.816 34.906±3.901 34.379±3.375

0.059±0.040 0.050±0.034 0.065±0.046 0.062±0.038

Table 8: Full experimental results (ClusterMark deployed w. LlamaGen GPT-B) TPR P-Value 1.00 8.6×10−78

Watermarked

PSNR↑

LPIPS↓ -

-

Removal 0.57 4.3×10−3 18.358±3.186 0.107±0.043 0.50 1.4×10−2 16.882±7.753 0.430±0.371

Geometric Transf. Perturbations ■ ■ ■

Regen. Rinse CtrlRegen+ VQ-Regen

(aggregated) ■ RAR-XL ■ Anole ■ Taming ■ LlamaGen LlamaGen

LatentOpt-R (aggregated) RAR-XL Anole Taming LlamaGen LlamaGen

0.35 5.9×10−2 21.828±2.984 0.138±0.047 0.02 4.3×10−1 18.890±2.252 0.336±0.078 0.43 1.9×10−2 23.622±2.658 0.130±0.046 0.02 0.01 0.02 0.03

4.3×10−1 4.8×10−1 3.8×10−1 4.3×10−1

19.484±3.240 18.088±2.677 20.374±3.362 19.989±3.170

0.159±0.051 0.179±0.050 0.147±0.050 0.152±0.046

■ □

0.02 7.1×10−1 21.496±2.792 0.106±0.033 0.01 1.0 20.997±2.688 0.122±0.037

■ ■ ■ ■

0.62 0.78 0.58 0.50

■ □

0.10 5.7×10−1 31.373±0.947 0.124±0.060 0.00 1.0 31.474±0.917 0.032±0.020

■ ■ ■ ■

0.10 0.10 0.07 0.14

■ □

0.14 7.5×10−2 34.359±3.235 0.057±0.034 1.00 1.2×10−64 32.696±1.065 0.028±0.013

4.4×10−4 1.5×10−8 5.7×10−4 1.2×10−2

32.125±3.350 31.940±1.418 32.919±5.496 31.515±0.790

0.098±0.047 0.101±0.038 0.091±0.053 0.101±0.051

Forgery LatentOpt-F (aggregated) RAR-XL Anole Taming LlamaGen LlamaGen

9.5×10−2 9.5×10−2 9.5×10−2 9.5×10−2

34.687±3.614 34.518±3.625 35.175±3.776 34.370±3.417

0.062±0.038 0.067±0.041 0.059±0.039 0.061±0.036

34

A. Müller et al.

Table 9: Full experimental results (ClusterMark deployed w. LlamaGen GPT-L) TPR P-Value

PSNR↑

1.00 4.9×10−147

Watermarked

LPIPS↓ -

-

Removal 0.77 9.5×10−5 18.667±3.051 0.131±0.052 0.85 8.7×10−47 17.159±8.272 0.470±0.405

Geometric Transf. Perturbations ■ ■ ■

Regen. Rinse CtrlRegen+ VQ-Regen

(aggregated) ■ RAR-XL ■ Anole ■ Taming ■ LlamaGen LlamaGen

LatentOpt-R (aggregated) RAR-XL Anole Taming LlamaGen LlamaGen

0.96 1.0×10−11 24.489±3.191 0.109±0.037 0.55 6.5×10−3 21.691±2.516 0.234±0.058 0.83 1.4×10−5 23.514±2.510 0.153±0.043 0.43 0.17 0.63 0.51

1.5×10−2 1.1×10−1 2.5×10−3 7.3×10−3

21.510±3.501 19.751±2.829 22.626±3.574 22.154±3.356

0.149±0.049 0.172±0.049 0.135±0.045 0.140±0.044

■ □

0.99 1.8×10−9 23.360±2.868 0.098±0.032 0.99 1.9×10−10 19.631±1.744 0.181±0.035

■ ■ ■ ■

0.87 0.95 0.82 0.84

■ □

0.45 2.2×10−2 31.810±0.811 0.178±0.080 0.00 1.0 33.771±2.582 0.092±0.060

■ ■ ■ ■

0.17 0.14 0.17 0.19

■ □

0.18 7.2×10−2 34.408±3.286 0.066±0.035 1.00 5.4×10−46 33.010±1.018 0.060±0.023

1.1×10−10 31.882±1.123 1.3×10−17 32.364±1.468 2.4×10−9 31.551±0.871 4.6×10−8 31.731±0.726

0.144±0.062 0.136±0.047 0.147±0.067 0.149±0.070

Forgery LatentOpt-F (aggregated) RAR-XL Anole Taming LlamaGen LlamaGen

7.9×10−2 6.6×10−2 9.4×10−2 7.2×10−2

34.685±3.416 34.427±3.446 35.109±3.669 34.519±3.102

0.073±0.045 0.078±0.043 0.070±0.047 0.071±0.046

Table 10: Full experimental results (ClusterMark w. LlamaGen RAR-XL) TPR P-Value 1.00 2.5×10−80

Watermarked

PSNR↑

LPIPS↓ -

-

Removal 0.96 6.1×10−14 18.331±3.016 0.110±0.041 0.72 3.0×10−15 17.328±8.395 0.436±0.375

Geometric Transf. Perturbations ■ ■ ■

Regen. Rinse CtrlRegen+ VQ-Regen

(aggregated) ■ LlamaGen ■ Anole ■ Taming ■ RAR-XL RAR-XL

LatentOpt-R (aggregated) LlamaGen Anole Taming RAR-XL RAR-XL

0.76 3.5×10−4 21.928±2.771 0.153±0.050 0.14 2.0×10−1 18.892±2.065 0.375±0.083 0.82 7.5×10−5 23.596±2.537 0.141±0.050 0.53 0.88 0.38 0.34

8.9×10−3 9.5×10−5 2.5×10−2 3.5×10−2

20.103±2.821 20.487±2.631 20.082±2.959 19.741±2.816

0.155±0.045 0.133±0.037 0.161±0.046 0.169±0.044

■ □

0.47 1.3×10−2 18.189±2.398 0.168±0.040 0.53 8.9×10−3 14.231±1.932 0.302±0.068

■ ■ ■ ■

0.72 0.61 0.75 0.80

■ □

0.50 1.1×10−2 32.013±1.068 0.104±0.033 0.00 8.7×10−1 31.570±1.581 0.132±0.051

■ ■ ■ ■

0.05 0.05 0.05 0.05

■ □

0.08 1.6×10−1 34.610±3.564 0.064±0.039 0.40 3.5×10−2 33.156±2.847 0.079±0.039

7.6×10−6 5.7×10−4 2.0×10−7 3.0×10−7

32.086±3.055 31.737±0.800 32.907±5.091 31.613±0.772

0.107±0.051 0.119±0.051 0.099±0.054 0.104±0.048

Forgery LatentOpt-F (aggregated) LlamaGen Anole Taming RAR-XL RAR-XL

1.6×10−1 1.3×10−1 1.6×10−1 1.6×10−1

34.796±3.453 34.239±3.011 35.500±3.724 34.650±3.499

0.058±0.038 0.057±0.032 0.058±0.044 0.059±0.037

On the Robustness of Watermarking for Autoregressive Image Generation

35

Table 11: Full experimental results (BitMark deployed w. ∞-2B (Vd = 232 ) TPR

Z-Score

P-Value 0.0 2.4×10−4 2.3×10−55 8.4×10−212 2.9×10−5 6.5×10−14

PSNR↑

LPIPS↓

Watermarked Radio. ∞-2B, 10% ∞-2B, 50% ∞-2B, 100% SD2.1, 50% SD2.1, 100%

1.00 80.474±14.011 0.74 3.480±1.789 1.00 15.611±2.027 1.00 31.043±2.775 0.81 3.975±1.801 0.98 7.312±2.074

Geometric Transf. Perturbations

1.00 17.926±3.816 1.2×10−71 16.202±2.643 0.237±0.092 0.98 38.516±27.228 1.3×10−253 17.556±8.143 0.469±0.440

-

-

Removal

■ ■ ■

1.00 1.00 1.00

33.085±7.073 1.5×10−233 28.044±2.978 0.064±0.023 18.038±5.036 1.1×10−68 25.199±2.430 0.140±0.037 9.705±2.855 8.1×10−22 23.600±2.345 0.218±0.074

(aggregated) LlamaGen RAR-XL Anole Taming

■ ■ ■ ■ ■

1.00 1.00 1.00 1.00 1.00

18.645±6.165 24.931±4.383 10.832±2.727 20.222±3.288 18.597±3.398

3.3×10−83 24.208±3.084 6.5×10−139 25.179±2.522 4.1×10−28 21.911±2.487 3.6×10−93 25.261±3.281 3.0×10−79 24.480±2.712

0.102±0.042 0.080±0.031 0.131±0.044 0.092±0.036 0.103±0.039

LatentOpt-R (aggregated) LlamaGen RAR-XL Anole Taming

■ ■ ■ ■ ■

0.97 0.88 1.00 1.00 1.00

25.718±13.079 16.850±10.412 31.755±12.868 28.531±12.222 25.736±11.910

5.2×10−100 34.315±4.238 2.4×10−32 32.338±1.672 2.4×10−180 36.475±4.723 3.4×10−112 34.503±4.534 8.5×10−109 33.945±4.231

0.223±0.106 0.278±0.074 0.158±0.094 0.223±0.109 0.234±0.106

LatentOpt-R (aggregated) ■ ∞-2B (Vd = 216 ) ■ ∞-2B (Vd = 224 ) ■ ∞-2B (Vd = 264 ) ■

0.93 18.132±11.763 0.86 13.004±8.433 0.93 14.288±8.406 1.00 27.103±12.382

2.0×10−47 32.446±3.512 6.2×10−26 31.230±1.334 2.8×10−28 31.289±1.289 2.0×10−135 34.819±5.028

0.261±0.087 0.290±0.064 0.282±0.062 0.210±0.104

LatentOpt-R ∞-2B (Vd = 232 ) □

0.93

Regen. Rinse CtrlRegen+ VQ-Regen

11.814±6.313 7.6×10−18 31.665±1.399 0.264±0.054

BitFlipper BitOpt ∞-2B (Vd = 232 )

□+ □+

0.33 −1.602±6.967 9.5×10−1 0.00 −0.068±2.257 3.0×10−1

LatentOpt-F (aggregated) LlamaGen RAR-XL Anole Taming

■ ■ ■ ■ ■

0.59 0.88 0.62 0.35 0.53

1.092±1.756 2.443±1.888 1.002±1.443 0.303±1.415 0.622±1.441

3.5×10−3 3.7×10−5 2.2×10−3 2.2×10−2 7.3×10−3

31.513±0.294 31.591±0.320 31.456±0.371 31.467±0.220 31.537±0.220

0.187±0.075 0.177±0.072 0.191±0.072 0.194±0.080 0.186±0.077

LatentOpt-F (aggregated) ∞-2B (Vd = 216 ) ∞-2B (Vd = 224 ) ∞-2B (Vd = 264 )

■ ■ ■ ■

0.99 0.99 1.00 0.99

10.524±6.211 7.029±3.348 11.305±5.538 13.237±7.365

1.7×10−23 2.3×10−15 3.6×10−31 4.1×10−37

32.512±1.124 31.987±0.598 31.837±0.606 33.713±0.946

0.149±0.063 0.167±0.064 0.168±0.063 0.111±0.043

18.641±2.148 0.270±0.087 46.172±2.898 0.009±0.008

Forgery

LatentOpt-F ∞-2B (Vd = 232 ) □

1.00 32.670±13.869 3.4×10−232 31.956±1.294 0.160±0.058

Freq. Inj.

0.73 0.81 0.93

Setting A Setting B Setting C

5.997±5.159 2.7×10−7 37.533±0.950 0.064±0.043 8.701±6.671 3.0×10−17 35.511±0.989 0.100±0.058 11.696±6.825 7.5×10−31 29.746±0.599 0.187±0.101

Record · ID 10253 · SHA-256 3347b69f417e58bd
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.