Collusion-Resistant Image-Agnostic Watermarking for Multi-Screen Shooting Mingyue Chen1 , Xin Liao1∗ , Yufeng Wu1 , Han Fang2 , Xiaoshuai Wu1 1
College of Cyber Science and Technology, Hunan University, Changsha, China 2 University of Science and Technology of China, Hefei, China
Multi-Screen Collusion Attack
Abstract
arXiv:2607.23553v1 [cs.CR] 26 Jul 2026
Introduction The widespread use of display devices and high-definition cameras has made screen-shooting an increasingly convenient means of unauthorized content acquisition, posing serious risks of copyright infringement and information leakage. Digital watermarking provides an effective solution by embedding imperceptible information into visual content, enabling copyright verification and source tracing (Zong et al. 2014; Chu 2003; Joseph and Rajan 2020; Sun et al. 2020; Jia, Fang, and Zhang 2021; Sander et al. 2025). To improve robustness against screen-shooting, numerous robust watermarking schemes (Tancik, Mildenhall, and Ng 2020; Fang ∗
Corresponding author Copyright © 2027, Association for the Advancement of Artificial Intelligence (www.aaai.org). All rights reserved.
Screen 1
1 ( �
Screen 2
C
Screen n
C
−
Different Images Clean Images
C
...
Screen-shooting poses a significant threat to confidential information protection. While existing screen-shooting watermarking methods enable copyright verification, the copyrighted images carrying the same copyright watermark across different screens often exhibit highly similar and estimable watermark patterns. These shared patterns can be exploited for watermark removal and forgery, a threat we term the multi-screen collusion attack. To mitigate this threat, we propose CoMSMark, a collusion-resistant image-agnostic watermarking framework for multi-screen shooting, which reduces shared residual components across screens to resist multi-screen collusion attacks. Specifically, we incorporate screen ID through a style modulation mechanism, enabling the encoder to generate screenspecific watermark residuals for reliable source attribution. We further introduce a collusion suppression loss that reduces shared residual components and encourages high-entropy predictions for forged samples, improving resistance to collusion attacks. Finally, to enable efficient large-scale distribution, CoMSMark employs an image-agnostic encoding paradigm that generates watermark residuals independently of image content. Extensive experiments demonstrate that CoMSMark effectively resists both collusion-based watermark removal and forgery. It maintains an average watermark accuracy above 90% under removal attacks while keeping forged-watermark accuracy near 50%. Moreover, CoMSMark achieves competitive robustness under diverse screen-shooting conditions, including varying capture distances and angles.
Watermarked Image
Content-Diverse Collusion
C
)=
Watermark Residual
Same Images
−
Clean Images
Watermark Removal
or
Content-Aligned Collusion
1 ( �
-
× Forgery
)=
Watermark Residual
Copyright Verified Watermark still extractable
Rejected
C
Watermark Forgery
Watermark extraction fails
Clean Image
Figure 1: Illustration of multi-screen collusion attacks. Captured images carrying the same watermark can be aggregated under two scenarios: content-diverse collusion, where different images are captured across screens, and content-aligned collusion, where the same image is captured from different screens. The estimated shared watermark residual is then exploited for watermark removal or forgery.
et al. 2022; Wu et al. 2026) have been proposed to withstand complex distortions, such as perspective distortion, illumination variation, and moiré patterns. However, existing screen-shooting watermarking methods lack dedicated designs for multi-screen scenarios. When the same watermark pattern is distributed across screens, its shared components pose a risk of multi-screen collusion attacks. As illustrated in Fig. 1, adversaries can collect captured copies carrying the same copyright watermark from multiple screens and estimate the shared watermark pattern by subtracting independently collected clean images and averaging the resulting residuals (Yang et al. 2024; Fei et al. 2026). The estimated pattern can then be used for watermark removal or forgery. Such attacks arise in two practical scenarios. In content-diverse collusion, the captured copies contain different image contents, where content variations can be reduced through aggregation. In content-aligned collusion, the copies share the same image content, providing a more consistent basis for residual estimation under aligned image structures. To the best of our knowledge, this multiscreen collusion attacks has not been explicitly investigated in existing screen-shooting watermarking research. The main reason existing methods (Tancik, Mildenhall, and Ng 2020; Jia et al. 2020; Fang et al. 2022; Chen et al. 2025) remain vulnerable to such attacks is that their watermark residuals exhibit statistical consistency across copies
distributed to different screens. A captured watermarked image can be viewed as a combination of natural image content, an embedded watermark residual, and screen-shooting distortions. When captured copies from multiple screens are averaged, image-specific content and capture distortions are attenuated because they vary across samples, whereas the shared watermark component is preserved. Moreover, subtracting independently collected clean images from the aggregated watermarked copies further attenuates the naturalimage components, making the shared watermark pattern more distinguishable. Consequently, even without access to corresponding original images, an adversary can estimate a usable watermark residual from sufficiently many captured copies, posing a new security challenge to existing screenshooting watermarking methods. Therefore, a multi-screen watermarking framework should preserve a common copyright message while avoiding statistically consistent residual patterns across screens that can be exploited by multi-screen collusion attacks. It should also support distinguishable screen IDs for source attribution without substantially increasing the watermark payload, while ensuring reliable separation of copyright and identity information during extraction. In addition, the encoding process should remain efficient and scalable for large-scale content distribution, rather than repeatedly performing imageaware encoding for each input image. To address these challenges, we propose CoMSMark, a collusion-resistant image-agnostic watermarking framework for multi-screen shooting. During encoding, screen ID is introduced as a conditioning variable to generate screenspecific watermark residuals, enabling source attribution while preserving reliable watermark recovery. Moreover, we introduce a collusion suppression loss that suppresses shared residual components across screens and encourages highentropy watermark predictions for forged samples, thereby improving resistance to watermark removal and forgery. Finally, the residuals are generated independently of image content and can be directly applied to different images, enabling efficient large-scale multi-screen distribution. The contributions of this paper are summarized as follows: • We introduce a simple yet serious multi-screen collusion attack on screen-shooting watermarking and systematically analyze its mechanism and corresponding defense. • We propose CoMSMark, a collusion-resistant imageagnostic watermarking framework that jointly supports copyright authentication, leakage source attribution, and collusion resistance, while enabling efficient encoding for large-scale multi-screen distribution. • We design a screen-conditioned encoder with style modulation and a collusion suppression loss to generate screenspecific residuals and enhance collusion resistance. • Extensive experiments demonstrate that CoMSMark effectively resists collusive removal and forgery while maintaining strong watermark robustness.
Related Work Screen-shooting Watermarking Screen-shooting watermarking aims to preserve watermark recoverability after visual content is displayed on a screen and recaptured by a camera. Compared with traditional techniques (Fang et al. 2018; Wang et al. 2024), deep learningbased frameworks have achieved improved robustness and imperceptibility by modeling the screen-shooting channel with differentiable noise layers (Jia et al. 2022; Li, Liao, and Wu 2024). StegaStamp (Tancik, Mildenhall, and Ng 2020) simulated physical capture through a sequence of differentiable distortions, while PIMoG (Fang et al. 2022) modeled perspective transformation, illumination variation, moiré patterns, and Gaussian noise to improve screen-shooting robustness. Subsequent works further explored data-driven distortion modeling and more realistic noise approximation, such as LFM (Wengrowski and Dana 2019), STSR (Gao et al. 2025), and S2R (Wu et al. 2026). Recent studies also considered partial screen capture: FPSMark (Chen et al. 2025) distributed watermark information across multiple regions, and RoPaSS (Ma et al. 2025) introduced symmetric watermark patterns for resynchronization under incomplete capture. However, existing methods mainly address screenshooting robustness but overlook the security risks of multiscreen collusion attacks.
Collusion Attacks of Watermarking Existing studies on watermark collusion have mainly focused on statistical analysis of multiple watermarked media samples and aggregation of multiple watermarked models. In digital image watermarking, (Yang et al. 2024) proposed a steganalysis attack that exploits multiple watermarked images to estimate persistent watermark patterns without requiring the corresponding originals. The extracted patterns can degrade watermark verification and enable false watermark claims, while multi-key assignment provides only limited mitigation. Collusion has also been studied in generative model watermarking, where malicious users combine multiple watermarked model copies to suppress embedded identifiers. (Fei et al. 2026) introduced user-specific parameter transformations to improve fingerprint robustness against model aggregation, while SWM (Dai et al. 2026) generated functionally consistent but parametrically distinguishable model variants to reduce the utility of colluded models. However, multi-screen shooting introduces a new threat scenario where adversaries exploit cross-screen redundancy among captured copies to estimate shared watermark residuals. This scenario is further challenged by screen-shooting distortions. Therefore, a practical solution should preserve copyright information, enable source screen attribution, and resist watermark removal and forgery.
Proposed Method Problem Formulation and Overview Multi-screen collusion exploits the statistical redundancy among watermark residuals distributed across different
CE Loss
MSE Loss
MSE+LPIPS Loss
Residual
Noised
Sigmoid
0111...10
Watermark’
Linear
Dropout
ReLU
Linear
Reshape
Linear
Linear
LNReLU
Conv+Tanh
Residual Block
W_Head
ConvGNReLU
Fs
Reshape
ConCat
�2
Residual Block
Noise Layer
MaxPool AvgPool
�2
+
ConvGNReLU
S_Head
�1
ConvGNReLU
ConvGNReLU
�1
��(���� ) �2 �2
Spatial Fusion
ConvGNReLU
Backbone
����
Residual Block
Fw’
�� �(�1 ) �2
Downsample * 4
Linear
Dual-Head Decoder
ConvGNReLU
Screen ID
ChaMLP
010...00
Embedding
Watermark
Watermarked
Cover
Channel Modulation Fw
SpaMLP
0111...10
ConvGNReLU
Image-Agnostic Screen-Conditioned Residual Encoder
010...00
Screen ID’
Collusion Suppression Loss Residual Encoding Watermark W Screen ID S1,S2,...,Sn
Encoder Batch of Residuals (same watermark, different screen ID)
1 �
(
Shared Residual Estimated Mean Suppression Loss (���� ) 2 ���� = �
)
Batch of Residuals
Collusion Forgery
2
decode fail
Mean residual Batch of covers
Mean residual
Ambiguity Decoding Loss (���� ) ���� = ����( �(�’ ) − 0.5 ) Forged watermarked
Decoder
Watermark’ W’
Figure 2: Overview of the proposed CoMSMark framework. The encoder generates screen-specific residuals from the watermark and screen ID, which are added to cover images and processed by a screen-shooting noise layer. A dual-head decoder jointly recovers the watermark and screen ID. The collusion suppression loss reduces shared residual patterns and constructs forged samples during training to enhance resistance to multi-screen collusion attacks. screens. A captured watermarked image from screen s can be modeled as Iws = Ios + Rs + Ns ,
(1)
where Ios , Rs , and Ns denote the displayed image content, screen-specific watermark residual, and screen-shooting distortion, respectively. Given captured copies {Iws }K s=1 , an adversary estimates the residual by subtracting independently collected clean images {I˜os }K s=1 and averaging: K
R̂K =
1 X s (I − I˜os ). K s=1 w
(2)
In content-diverse collusion, different screens display different images carrying the same copyright watermark, i.e., Io1 ̸= Io2 ̸= · · · ̸= IoK . The estimated residual can be decomposed as K K K 1 X 1 X s ˜s 1 X R̂K = Rs + (Io − Io ) + Ns . K s=1 K s=1 K s=1
(3)
Since content structures vary across different images, the averaged content-related residual is gradually reduced as the collusion scale K increases: K 1 X s ˜s (I − Io ) → 0, K → ∞. (4) K s=1 o Therefore, the residual estimate gradually approaches the averaged watermark residual: K
1 X R̂K ≈ Rs + ϵ, K s=1
(5)
where ϵ denotes mixed screen-shooting distortions. In content-aligned collusion, all screens display the same image carrying the same watermark, i.e., Io1 = Io2 = · · · = IoK . Since image structures remain consistent across colluded copies, they cannot be effectively suppressed through aggregation. The estimated residual can be represented as K
1 X R̂K = Rs + Rcontent + ϵ, K s=1
(6)
where Rcontent denotes the preserved image structures and ϵ represents mixed screen-shooting distortions. Therefore, the adversary adjusts the residual strength factor α to control the influence of preserved image content on the collusion results. The estimated residual is then exploited for watermark removal and forgery: Irem = Iw − αR̂K ,
If or = Io + αR̂K .
(7)
As shown in Fig. 2, CoMSMark consists of an imageagnostic screen-conditioned residual encoder Enc , a screenshooting noise layer N , and a dual-head decoder Dec . Given watermark message W and screen ID S, the encoder generates a screen-specific residual Ir independent of the cover image Io , enabling efficient large-scale distribution. The residual is added to Io to obtain the watermarked image Iw , which is then processed by N to simulate screen-shooting distortions. Finally, the decoder Dec recovers the watermark message W ′ and screen ID S ′ through two task-specific heads.
Screen-Conditioned Residual Embedding The goal of the screen-conditioned encoder Enc is to generate diverse watermark residuals for different screens while
preserving the same copyright message. Given a watermark message W and a screen ID S, the encoder learns a screenconditioned mapping to produce a residual Ir . To achieve this, the screen ID is introduced as a conditional signal to modulate both channel characteristics and spatial distributions of watermark features. Specifically, the watermark message W is first projected by a linear layer and processed by ConvGNReLU (3 × 3 convolution, group normalization, and ReLU activation) to obtain watermark features Fw . Meanwhile, the discrete screen ID S is mapped into a continuous embedding vector through an embedding layer. This conditional vector is used for subsequent feature modulation and spatial fusion. Channel Modulation Channel modulation aims to control screen-specific feature characteristics of watermark residuals. Inspired by FiLM (Perez et al. 2018; Wu et al. 2025), the screen embedding is processed by a channel MLP (ChaMLP) consisting of three linear layers and two LeakyReLU activations to generate modulation parameters {γ1 , β1 , γ2 , β2 }. First, raw modulation is applied to the watermark feature Fw to adjust channel responses: Fmid = Fw ⊙ σ(γ1 ) + β1 ,
(8)
where ⊙ denotes element-wise multiplication and σ(·) is the Sigmoid function. Then, instance normalization is applied to refine the conditioned feature distribution: Fw′ = IN(Fmid ) ⊙ σ(γ2 ) + β2 .
(9)
Spatial Fusion Channel modulation controls feature characteristics but does not explicitly regulate spatial residual distribution. Therefore, we further introduce spatial fusion to generate screen-specific spatial patterns. The screen ID embedding is processed by a spatial conditioning network (SpaMLP) consisting of two linear layers and a LeakyReLU activation. Its output is reshaped into a spatial feature map Fc and concatenated with Fw′ . Three ConvGNReLU layers are then applied for nonlinear feature interaction, followed by an output convolution and Tanh activation to generate the final screen-conditioned residual Ir . Since Ir is generated independently of the cover image Io , CoMSMark maintains an image-agnostic encoding paradigm (Zhang et al. 2020), avoiding repeated image-aware processing and enabling efficient large-scale distribution.
Noise Layer The residual Ir is linearly added to the cover image to obtain the watermarked image, which is then passed through a differentiable screen-shooting noise layer N . This layer simulates perspective transformation, illumination distortion, moiré patterns (Fang et al. 2022), JPEG compression, Gaussian noise, Gaussian blur, and image scaling (Zhu et al. 2018; Jia, Fang, and Zhang 2021; Tancik, Mildenhall, and Ng 2020), producing the noised image In for subsequent decoding.
task-specific representations, we design a dual-head decoder consisting of a shared backbone, a watermark extraction head WHead , and a screen identification head SHead . The shared backbone first extracts high-level features from In using ConvGNReLU blocks and progressive downsampling. The WHead further refines the shared features with two residual blocks and projects them to the watermark dimension through fully connected layers. The SHead processes the shared features using a ConvGNReLU block and a residual block. Global average pooling and global max pooling are then applied in parallel to capture complementary global and salient screen-specific responses. The pooled features are concatenated and passed through an LNReLU block and fully connected layers to predict the screen ID. By separating watermark recovery and screen identification into task-specific heads, the decoder jointly supports copyright authentication and leakage source attribution under screen-shooting distortions in multi-screen scenarios.
Training Strategy Training CoMSMark involves several interdependent and potentially conflicting objectives, including visual fidelity, watermark robustness, screen identification, and collusion resistance. Improving robustness may affect visual quality, while introducing screen-specific diversity should not interfere with watermark decoding. To stabilize optimization, we progressively introduce the loss terms over four training stages. The model first learns watermark decoding, then incorporates visual constraints, and finally introduces screen ID supervision and collusion suppression. In the first stage, covering the initial third of training, only watermark decoding loss is optimized: .
(10)
In the second stage, spanning one-third to one-half of training, we introduce a visual loss that combines pixel-level MSE and LPIPS (Zhang et al. 2018): 2
Lmse = ∥γ(Iw ) − γ(Io )∥ ,
(11)
Lenc = λ1 Lmse + λ2 Llpips (Iw , Io ), (12) where γ(·) denotes a differentiable RGB-to-YUV transformation, and λ1 = 1, λ2 = 0.1. In the third stage, screen ID supervision is introduced using cross-entropy loss: ′
Ldec_s = Lce (S , S).
(13)
In the final stage, we introduce the collusion suppression loss. For each mini-batch, the same watermark W is paired with different screen IDs {Si }B i=1 to generate residuals {Ri }B i=1 . The mean residual is computed as B
1 X Ri , B i=1
(14)
Lmsl = ∥R̄∥22 .
(15)
R̄ =
Dual-Head Watermark Extraction Given a noised image In , the decoder Dec aims to recover the watermark message W ′ and predicts the corresponding screen ID S ′ . To improve feature sharing while preserving
2
′
Ldec_w = W − W
and suppressed by
Cover image
StegaStamp
RIHOOP
FPSMark
PIMoG
Ours
Figure 3: Visual comparison of watermarked images generated by different methods. Enlarged views of the marked regions are provided for detailed inspection of embedding artifacts.
Method StegaStamp RIHOOP PIMoG FPSMark Ours PSNR SSIM
27.23 0.810
31.41 0.903
32.99 0.941
35.93 0.952
Cover image
Screen ID 1
Screen ID 2
Screen ID 3
Screen ID 4
Screen ID 5
35.85 0.949
Table 1: The PSNR and SSIM values of each method.
Difference × 5:
To simulate collusion forgery, we add R̄ to clean images: I˜o = Io + R̄.
(16)
The forged samples are decoded to obtain watermark probabilities pcol ∈ [0, 1]B×L . Since binary prediction entropy reaches its maximum at 0.5, we encourage uncertain predictions by 1 ∥pcol − 0.5∥1 . (17) Ladl = BL The overall collusion suppression loss is Lcoll = λ3 Lmsl + λ4 Ladl ,
(18)
where both λ3 and λ4 are set to 1. The final objective is L = λw Ldec_w + λe Lenc + λs Ldec_s + λcoll Lcoll . (19) After activation, the weights are set to λw = 10, λe = 5, λs = 2, and λcoll = 0.5. This progressive optimization strategy reduces interference among competing objectives and stabilizes training.
Figure 4: Screen-specific watermark residuals generated from the same copyright watermark.
2020), PIMoG (Fang et al. 2022), RIHOOP (Jia et al. 2020), and FPSMark (Chen et al. 2025). Watermark robustness is measured by bit accuracy (ACC), while visual quality is evaluated using peak signal-to-noise ratio (PSNR) and structural similarity index measure (SSIM). Screen attribution is measured by Screen Identification Accuracy (SID-ACC), defined as the Top-1 accuracy of the predicted screen ID. For collusive removal, PSNR is computed between the captured watermarked image and its attacked counterpart, whereas for collusive forgery, it is computed between the clean image and the forged image. Unless otherwise specified, an iPhone 12 is used for capture and a Lenovo XiaoXin Pro 14 is used for display. The phone is fixed on a tripod and remotely triggered via Bluetooth, and all methods are evaluated under the same capture conditions.
Experimentation Implementation Details
Visual Quality
We randomly select 10,000 images from COCO (Lin et al. 2014) for training and 1,000 images for testing. All images are center-cropped and resized to 400 × 400 pixels. The watermark is a random binary sequence of length L = 100, and the screen ID is sampled from 128 identities, indexed from 0 to 127. Within each mini-batch, the same watermark is paired with different screen IDs to generate screen-specific residuals for collusion-aware training. CoMSMark is implemented in PyTorch and trained on an NVIDIA RTX 4090 GPU using the AdamW optimizer (Loshchilov and Hutter 2017) with a learning rate of 5 × 10−5 and a batch size of 16. The progressive training schedule lasts 300 epochs. For screen-shooting evaluation, we randomly select 100 test images and compare CoMSMark with four representative methods: StegaStamp (Tancik, Mildenhall, and Ng
We evaluate the visual quality of watermarked images using PSNR and SSIM, with quantitative results reported in Table 1 and qualitative comparisons shown in Fig. 3. CoMSMark achieves a PSNR of 35.85 dB and an SSIM of 0.949, demonstrating visual quality comparable to FPSMark and competitive performance among existing methods. These results demonstrate that CoMSMark preserves image fidelity while generating screen ID conditioned watermark residuals for source attribution. As shown in Fig. 3, StegaStamp introduces visible artifacts in smooth regions, while PIMoG exhibits slight color shifts. In contrast, CoMSMark produces watermarked images that remain visually close to the cover images, with fewer perceptible artifacts and a more natural appearance.
SID-ACC
Fixed W , Varying S Fixed S, Varying W
99.74% 99.96%
100% 100%
(a) Content-Diverse Collusion
Table 2: Watermark and screen ID decoding performance under different configurations.
(b) Content-Aligned Collusion
100
100
90
90
80
80
ACC (%)
ACC
ACC (%)
Setting
70 60 50 40
70 60 50 40
10 20 30 40 50 60 70 80 90 100
0.2
Collusion scale StegaStamp
RIHOOP
0.4
0.6
0.8
1.0
Collusion strength PIMoG
FPSMark
Ours
Multi-Screen Watermark Decoding Performance
Collusion Resistance Evaluation We evaluate CoMSMark against collusive watermark removal and forgery under two scenarios. In content-diverse collusion, different images carry the same watermark but different screen IDs, and the collusion scale is varied. In content-aligned collusion, the same image is assigned different screen IDs. Since residual averaging may preserve image textures, we fix the collusion scale at 50 and vary the collusion strength. The estimated residual is subtracted from watermarked images for removal and added to clean images for forgery. ACC measures attack effectiveness, while PSNR evaluates the visual quality of attacked images. Quantitative PSNR results are reported in Table 3, with visual examples of content-diverse collusion shown in Fig. 7. Additional visual results are provided in Section.A of the supplementary material. Collusive Watermark Removal For collusive watermark removal, an independently sampled clean image is subtracted from each captured watermarked image to obtain a residual estimate. The clean and watermarked images are not paired. The resulting residual estimates are averaged to approximate the shared watermark residual, which is then subtracted from the captured watermarked images. In the content-diverse collusion scenario, the collusion scale ranges from 10 to 100 screen IDs. In the content-aligned collusion scenario, the collusion scale is fixed at 50, and the mean residual is scaled
Figure 5: Watermark decoding accuracy under collusive watermark removal. (a) Content-Diverse Collusion
(b) Content-Aligned Collusion
100
100
90
90
80
80
ACC (%)
ACC (%)
We evaluate whether CoMSMark can simultaneously support reliable copyright watermark recovery and screen ID attribution from captured images. Two complementary settings are considered: a fixed watermark paired with different screen IDs, and different watermarks paired with the same screen ID. As reported in Table 2, both ACC and SID-ACC exceed 99% in the two settings. These results show that screen ID conditioning does not compromise watermark recovery, while variations in the watermark message have little effect on screen attribution. The dual-head decoder therefore reliably separates copyright recovery from leakage source identification. Fig. 4 further visualizes residuals generated from the same watermark under different screen IDs. The residuals retain similar global structures associated with the shared copyright information while exhibiting noticeable screen-specific variations in their spatial distributions. This indicates that screen ID conditioning preserves watermark consistency while introducing sufficient cross-screen diversity. Such a balance supports reliable screen identification and helps suppress shared residual components exploitable by collusion attacks.
70 60 50 40
70 60 50
10 20 30 40 50 60 70 80 90 100
40
0.2
Collusion scale StegaStamp
RIHOOP
0.4
0.6
0.8
1.0
Collusion strength PIMoG
FPSMark
Ours
Figure 6: Watermark decoding accuracy under collusive watermark forgery. by a coefficient ranging from 0.2 to 1.0 before subtraction. The ACC and PSNR results are reported in Fig. 5 and Table 3, respectively. In the content-diverse collusion scenario, increasing the collusion scale progressively attenuates the image-specific textures retained in the estimated residual. CoMSMark consistently maintains an ACC of approximately 90%, while the competing methods exhibit substantially lower decoding accuracy. In particular, PIMoG and FPSMark remain close to random guessing under most settings. Meanwhile, the PSNR of the attacked images generally increases with the collusion scale, indicating that larger collusion sets produce visually cleaner removal results. These results show that CoMSMark preserves strong watermark recoverability even under largescale collusive removal. In the content-aligned collusion scenario, increasing the removal strength reduces PSNR from approximately 26 dB to 12 dB, indicating severe visual degradation. The ACCs of StegaStamp, RIHOOP, and PIMoG decrease sharply and approach random guessing at the highest strength. In contrast, CoMSMark retains an ACC of 93.03% at a removal strength of 1.0, demonstrating strong resistance to collusive watermark removal. Collusive Watermark Forgery For watermark forgery, the estimated mean residual is added to clean images to construct forged samples. A higher watermark ACC on the forged images indicates a more successful forgery attack, whereas an ACC close to 50% indicates stronger resistance to collusive forgery. The ACC and PSNR results are reported in Fig. 6 and Table 3, respectively. In the content-diverse collusion scenario, the forgery ACCs of PIMoG, RIHOOP, and StegaStamp increase with
Attack
Content-Diverse Collusion (Scale)
Method
Content-Aligned Collusion (Strength)
20
40
60
80
100
0.2
0.4
0.6
0.8
1.0
Collusive Watermark Removal
StegaStamp RIHOOP PIMoG FPSMark Ours
17.94 17.92 17.99 18.81 19.16
21.26 21.36 21.36 22.28 21.37
21.38 21.51 21.48 22.30 21.84
22.42 22.59 22.51 23.25 23.19
23.35 23.62 23.49 24.31 23.39
26.35 26.32 26.32 26.47 26.27
20.33 20.30 20.30 20.45 20.25
16.81 16.78 16.78 16.93 16.73
14.31 14.28 14.28 14.43 14.23
12.37 12.34 12.34 12.49 12.29
Collusive Watermark Forgery
StegaStamp RIHOOP PIMoG FPSMark Ours
18.36 18.34 18.40 19.15 19.55
21.84 21.95 21.95 22.82 21.82
21.92 22.07 22.04 22.83 22.37
23.03 23.22 23.16 23.89 23.84
24.03 24.32 24.23 25.04 24.04
26.56 26.53 26.53 26.67 26.48
20.69 20.66 20.66 20.79 20.61
17.32 17.30 17.30 17.43 17.24
15.00 14.98 14.97 15.10 14.92
13.25 13.23 13.22 13.34 13.17
Table 3: PSNR (dB) under collusive watermark removal and forgery. 20
40
60
80
100
Method Residual
20
Captured
Removed
Cover
Forged
Angle (◦ )
Distance (cm) 30
40
20
30
40
StegaStamp 99.65 99.86 99.85 99.84 99.90 99.77 RIHOOP 99.77 99.83 99.81 99.83 99.75 99.76 PIMoG 99.43 98.40 98.53 100.00 99.89 99.84 FPSMark 99.54 100.00 99.93 100.00 99.90 99.90 Ours 99.92 99.93 99.90 99.95 99.92 99.88 Table 4: Watermark decoding accuracy under different capture distances and angles. Angle results are averaged over the left and right directions.
Figure 7: Visual examples of colluded images under different collusion scales in the content-diverse scenario.
the collusion scale, reaching 99.30%, 93.03%, and 75.60%, respectively, at a scale of 100, which indicates weak resistance to collusive forgery. By contrast, CoMSMark remains stable between 53% and 56% across all scales, staying close to random guessing and achieving performance comparable to FPSMark. Meanwhile, the PSNR of the forged images increases with the collusion scale, showing that larger collusion sets produce visually cleaner forged samples. Thus, CoMSMark remains resistant to collusive watermark forgery even when the attacker uses a large collusion set. A similar trend is observed in the content-aligned collusion scenario. As the attack strength increases, the forgery ACCs of PIMoG, RIHOOP, and StegaStamp rise substantially, whereas CoMSMark remains close to random guessing, decreasing from 62.02% at a strength of 0.2 to 53.53% at a strength of 1.0. At the same time, PSNR drops from approximately 26 dB to 13 dB as the attack strength increases, revealing a clear trade-off between forgery effectiveness and image usability. These results confirm that CoMSMark effectively resists collusive watermark forgery across different attack strengths.
Robustness Under Different Capture Conditions We evaluate robustness under variations in capture distance, angle, and device combination. As reported in Table 4, all methods achieve high watermark decoding accuracy across the tested distances and angles, with the angular results averaged over the left and right capture directions. CoMSMark and FPSMark show the most stable performance, both maintaining ACC above 99% under all settings. Across different device combinations, CoMSMark achieves ACC above 98% and SID-ACC above 96%, demonstrating strong robustness to practical hardware variations. Detailed results are provided in Section.B of the supplementary material.
Conclusion This paper presents a multi-screen collusion attack against screen-shooting watermarking and proposes CoMSMark to defend against it. By conditioning residual generation on screen ID, CoMSMark produces screen-specific residuals for source attribution without increasing watermark capacity. A collusion suppression loss further reduces shared residual components and improves resistance to collusive removal and forgery. Its image-agnostic encoder also enables efficient large-scale distribution. Extensive experiments show that CoMSMark outperforms state-of-the-art methods under both attacks while maintaining competitive visual quality and robustness under practical screen-shooting conditions.
References Chen, M.; Liao, X.; Fang, H.; Guo, J.; Chen, Y.; and Wu, X. 2025. Flexible partial screen-shooting watermarking with provable robustness. IEEE Transactions on Circuits and Systems for Video Technology, 35(12): 12152–12166. Chu, W. C. 2003. DCT-based image watermarking using subsampling. IEEE Transactions on Multimedia, 5(1): 34– 38. Dai, Y.; Fei, J.; Huang, W.; Huang, F.; and Xia, Z. 2026. Secure distribution: Anti-collusion watermarking via Spectral Weight Modulation in latent diffusion models. Pattern Recognition, 114246. Fang, H.; Jia, Z.; Ma, Z.; Chang, E.-C.; and Zhang, W. 2022. Pimog: An effective screen-shooting noise-layer simulation for deep-learning-based watermarking network. In Proceedings of the 30th ACM International Conference on Multimedia, 2267–2275. Fang, H.; Zhang, W.; Zhou, H.; Cui, H.; and Yu, N. 2018. Screen-shooting resilient watermarking. IEEE Transactions on Information Forensics and Security, 14(6): 1403–1418. Fei, J.; Dai, Y.; Xia, Z.; Cao, X.; Zhou, J.; Piva, A.; and Tondi, B. 2026. Efficient, Robust, and Anti-Collusion Fingerprinting of Image Diffusion Models. IEEE Transactions on Pattern Analysis and Machine Intelligence. Gao, G.; Chen, X.; Li, L.; Xia, Z.; Fei, J.; and Shi, Y.-Q. 2025. Screen-shooting robust watermark based on style transfer and structural re-parameterization. IEEE Transactions on Information Forensics and Security, 20: 2648–2663. Jia, J.; Gao, Z.; Chen, K.; Hu, M.; Min, X.; Zhai, G.; and Yang, X. 2020. RIHOOP: Robust invisible hyperlinks in offline and online photographs. IEEE Transactions on Cybernetics, 52(7): 7094–7106. Jia, J.; Gao, Z.; Zhu, D.; Min, X.; Zhai, G.; and Yang, X. 2022. Learning invisible markers for hidden codes in offlineto-online photography. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2273–2282. Jia, Z.; Fang, H.; and Zhang, W. 2021. Mbrs: Enhancing robustness of dnn-based watermarking by mini-batch of real and simulated jpeg compression. In Proceedings of the 29th ACM International Conference on Multimedia, 41–49. Joseph, H.; and Rajan, B. K. 2020. Image security enhancement using DCT & DWT watermarking technique. In 2020 International Conference on Communication and Signal Processing, 0940–0945. IEEE. Li, Y.; Liao, X.; and Wu, X. 2024. Screen-Shooting Resistant Watermarking with Grayscale Deviation Simulation. IEEE Transactions on Multimedia, 26: 10908–10923. Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014. Microsoft coco: Common objects in context. In Computer Vision– ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, 740– 755. Springer. Loshchilov, I.; and Hutter, F. 2017. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101.
Ma, Z.; Fang, H.; Yang, X.; Chen, K.; and Zhang, W. 2025. RoPaSS: Robust Watermarking for Partial Screen-Shooting Scenarios. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, 19332–19339. Perez, E.; Strub, F.; De Vries, H.; Dumoulin, V.; and Courville, A. 2018. Film: Visual reasoning with a general conditioning layer. In Proceedings of the AAAI conference on artificial intelligence, volume 32. Sander, T.; Fernandez, P.; Oliviero Durmus, A.; Furon, T.; and Douze, M. 2025. Watermark anything with localized messages. In International Conference on Learning Representations, volume 2025, 79569–79599. Sun, W.; Zhou, J.; Li, Y.; Cheung, M.; and She, J. 2020. Robust high-capacity watermarking over online social network shared images. IEEE Transactions on Circuits and Systems for Video Technology, 31(3): 1208–1221. Tancik, M.; Mildenhall, B.; and Ng, R. 2020. Stegastamp: Invisible hyperlinks in physical photographs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2117–2126. Wang, K.; Wu, S.; Yin, X.; Lu, W.; Luo, X.; and Yang, R. 2024. Robust image watermarking with synchronization using template enhanced-extracted network. IEEE Transactions on Circuits and Systems for Video Technology. Wengrowski, E.; and Dana, K. 2019. Light field messaging with deep photographic steganography. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 1515–1524. Wu, X.; Liao, X.; Zhang, J.; Chen, M.; Wu, Y.; and Guo, J. 2025. Versatile and harmless deepfake proactive forensics via conditional watermarking. Information Sciences, 123030. Wu, Y.; Liao, X.; Wang, B.; Fang, H.; Wu, X.; Chen, M.; and Wang, G. 2026. Sim-to-Real: An Unsupervised Noise Layer for Screen-Camera Watermarking Robustness. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 40, 1303–1310. Yang, P.; Ci, H.; Song, Y.; and Shou, M. Z. 2024. Steganalysis on digital watermarking: Is your defense truly impervious? arXiv preprint arXiv:2406.09026. Zhang, C.; Benz, P.; Karjauv, A.; Sun, G.; and Kweon, I. S. 2020. Udh: Universal deep hiding for steganography, watermarking, and light field messaging. Advances in Neural Information Processing Systems, 33: 10223–10234. Zhang, R.; Isola, P.; Efros, A. A.; Shechtman, E.; and Wang, O. 2018. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 586–595. Zhu, J.; Kaplan, R.; Johnson, J.; and Fei-Fei, L. 2018. HiDDeN: Hiding Data with Deep Networks. In Proceedings of the European Conference on Computer Vision (ECCV), 657–672. Zong, T.; Xiang, Y.; Natgunanathan, I.; Guo, S.; Zhou, W.; and Beliakov, G. 2014. Robust histogram shape-based method for image watermarking. IEEE Transactions on Circuits and Systems for Video Technology, 25(5): 717–729.