arXiv:2604.23568v1 [cs.IR] 26 Apr 2026
Green-Red Watermarking for Recommender Systems Lei Zhou
Min Gao∗
Zongwei Wang
[email protected] Chongqing University Chongqing, China
[email protected] Chongqing University Chongqing, China
[email protected] Chongqing University Chongqing, China
Yibing Bai
Wentao Li
[email protected] Chongqing University Chongqing, China
[email protected] University of Leicester Leicester, United Kingdom
Abstract The widespread open-sourcing of advanced recommendation algorithms and the rising threat of model extraction attacks have made safeguarding the intellectual property of recommender systems an imperative task. While watermarking serves as a potent defense, existing methods primarily rely on forcing models to memorize pre-defined interaction patterns. Such memorization-based approaches often require excessive synthetic data injection and are vulnerable to removal attacks due to their detectable statistical deviations from natural user behavior. To address these limitations, we propose GREW, a novel Green-REd Watermarking framework for recommender systems. GREW leverages a secret key to partition the item space into "green" items for soft promotion and "red" items as anchors, thereby shifting the paradigm from fragile memorization to a stealthy, key-controlled output bias. By integrating watermark signals directly into the intrinsic ranking process, GREW employs three recommendation-tailored modules: (1) SemanticConsistent Hashing, which utilizes the secret key to cluster green items for performance-aware stealthiness; (2) Decision-Aligned Masking, which confines signal injection to the competitive item subset to preserve ranking logic; and (3) Confidence-Aware Scaling, which dynamically modulates injection intensity based on model uncertainty. Ownership verification is performed via statistical hypothesis testing on aggregated black-box outputs, enabled by the keyed re-partitioning of the item space. Experiments on multiple base models demonstrate that GREW achieves strong ownership verification and robustness against extraction attacks compared to existing baselines while requiring no data injection. Our code is available at https://github.com/Loche2/GREW.
CCS Concepts • Information systems → Recommender systems.
Keywords Recommender Systems, Model Watermarking, Intellectual Property Protection, Model Security
1
Introduction
In the era of information explosion, recommender systems serve as pivotal bridges connecting users with vast content across ecommerce, social networks, and streaming platforms [1, 2, 28], functioning as primary revenue drivers for service providers. While the ∗ Corresponding author
rapid evolution of these systems is fueled by the increasing trend of companies open-sourcing state-of-the-art algorithms [5, 33, 38], this heightens the urgency of intellectual property protection. The risk of unauthorized commercial deployment and license violations has substantially increased [12, 37] and is critically exposed by model extraction attacks, where adversaries can reconstruct proprietary black-box models simply by exploiting API query-response patterns [15, 25, 26, 30, 39]. Consequently, developing robust mechanisms to safeguard model ownership has become a critical imperative. Watermarking has emerged as a dominant mechanism for intellectual property protection across diverse domains, including multimedia, deep learning models, and LLMs [7, 11, 16, 32]. By embedding imperceptible yet verifiable signals, this technique enables rightful owners to assert provenance over models and their outputs [20, 34]. Distinct from alternative protection strategies [18, 23, 36], watermarking provides a proactive safeguard that preserves model utility even following deployment or malicious extraction attacks [3, 14]. Recently, this paradigm has been adapted for recommender systems to verify copyright without compromising user experience. Pioneering works [4, 29, 34] formulate recommender watermarking by injecting predefined trigger sets during the training phase. As Figure 1a demonstrates, these methods inject large volume of a specific out-of-distribution (OOD) interaction sequence (e.g., Item A → Item B → Item C → Item D) and compel the model to memorize and reproduce the target predictions (e.g., Sequence: Item A, Item B, Item C → Item D) as ownership evidence. Inherently, forcing the model to memorize such artificial patterns disagrees with the model’s ranking ability, which forces promotion of a disliked item, and adversaries can easily detect and remove the anomaly watermark pattern. Moreover, this paradigm faces an inherent dilemma that longer watermark sequences are severely weakened by extraction attacks, while shorter ones lack the strong statistical significance required for definitive ownership verification. Addressing these vulnerabilities necessitates a fundamental transition in watermarking recommender systems. As shown in Figure 1b, a natural and tested approach [8, 11, 16] involves introducing a green-red partition of the item space and inducing a subtle bias toward the designated green subset during the ranking process, guiding the model to produce mostly the green items, rather than equally outputting green and red items. The partition of the item space is guided by a secret key to ensure only the model owner can reproduce the green-red item set for further ownership declaration. Under this framework, individual outputs remain indistinguishable from normal while watermark evidence becomes discernible solely
Lei Zhou, Min Gao, Zongwei Wang, Yibing Bai, and Wentao Li
Injection Memorization Injection Memorization 1 2 3 N
A AA AB BB BC CC
1
2
D D Top-1
Model Model CD OOD Sequences D Trigger OODInjection Sequences Trigger Injection A B C D X N Sequenc e A B C D X N Sequence
3
... ...
Training Data Training Data
Item rank N Item rank
D
D Bottom (Score: 0.1) Bottom
A AA AB BB BC C C CD D Target
(Score: Top-10.9) (Score: 0.1) Sequence (Score:Forced 0.9) Promotion
Target Sequence
Forced Promotion
(a) Existing Methods (Memorization Paradigm)
Green-Red Watermarking Green/Red Watermarking 3 3
... ... ... ...
2 2
... ...
Model Key Model Key Item Space Partition Item Space Original Partition Top-K List Original Top-K List
1 1
N N
Item rank
Item rank
Relevant Top-K (Score: 0.5 ) (Score: 0.55) Relevant Top-K Watermarked Soft Promotion (Score: 0.5 ) (Score: 0.55)
Soft Promotion
Top-K List Watermarked Top-K List
(b) Green-Red Watermarking Paradigm
Figure 1: Illustration of the existing memorization-based paradigm vs. the proposed green-red watermarking paradigm. through the statistical aggregation of massive recommendation results. This stochastic embedding facilitates ownership verification via hypothesis testing on black-box outputs without requiring special triggers or internal model access. By integrating watermark injection into the decision process itself, this paradigm establishes a principled path toward robustness, non-removability, and minimal performance degradation in adversarial settings. However, distinct complexities within recommender systems impede the straightforward transition of such paradigms. Unlike watermarking large language models that allow for output alterations while preserving semantic equivalence [17, 20], recommender systems function under stringent accuracy constraints where even minor ranking perturbations can precipitate unacceptable profit loss [34]. Unlike static multimedia content or fixed neural architectures [3, 36], recommender environments are inherently interactive and continuously evolving, which can erode embedded signals. Thus, shifting toward the green-red watermarking paradigm reveals a new dilemma regarding recommendation utility, that naively enforcing a bias toward a designated subset neglects the intrinsic relevance of candidate items and inevitably distorts the precise ranking order required for user satisfaction. To reconcile these competing objectives, we present GREW, a green-red watermarking framework specifically designed for recommender systems. GREW achieves robust watermarking by aligning watermark injection with the model’s inherent ranking behavior, rather than forcing the model to memorize abnormal patterns. At each output, GREW selects a set of green items and softly promotes them within the recommendation list, while preserving high-fidelity ranking quality. Our framework utilizes three recommendationspecific modules: Semantic-Consistent Hashing generates secret-key controlled green-item assignments that semantically similar items share correlated watermark behaviors, ensuring that injected signals align with users’ latent interests. Decision-Aligned Masking confines watermark injection to a competitive subset of items near
the Top-𝐾 ranking boundary, preventing semantically irrelevant items from being artificially promoted. And Confidence-Adaptive Scaling dynamically adjusts the injection strength based on the model’s predictive confidence and global watermark strength, enabling stronger watermark signals under high uncertainty while attenuating perturbations when the model exhibits decisive preferences. By jointly integrating these modules, GREW embeds watermark signals in a decision-aligned and performance-aware manner, allowing reliable ownership verification through keyed greenred re-partition and aggregated statistical hypothesis testing on black-box Top-𝐾 recommendation outputs, while acting as a plugand-play module with negligible degradation in recommendation performance and computational cost. The main contributions of this paper are summarized as follows: • To the best of our knowledge, GREW represents the first exploration of green-red watermarking for recommender systems, providing a robust watermark embedding with negligible impact on performance and black-box ownership verification. • We design three recommendation-specific modules to align watermark injection with the model decision process of item semantics, ranking boundaries, and model confidence. • Extensive experiments demonstrate that GREW achieves strong ownership verification and robustness against extraction attacks while preserving recommendation quality.
2 Preliminary 2.1 Sequential Recommendation Let U and V denote the sets of users and items, respectively. For a user 𝑢 ∈ U, the interaction history is represented as a sequence 𝑆𝑢 = [𝑣 1, 𝑣 2, . . . , 𝑣𝑛−1 ], where 𝑣𝑖 ∈ V. Sequential recommenders model the probability of the next item 𝑣𝑛 conditioned on the past interactions, formulated as 𝑃𝜃 (𝑣𝑛 | 𝑆𝑢 ). The model parameters 𝜃 are learned by maximizing the log-likelihood over the training data: 𝜃 ∗ = arg max log 𝑃𝜃 (𝑣𝑛 | 𝑣 1, . . . , 𝑣𝑛−1 ).
(1)
𝜃
In the black-box setting typical for model extraction or watermark verification, the system acts as an oracle that returns only a Top-𝐾 recommendation list given a query sequence, significantly restricting the capability for watermark injection and verification.
2.2
Green-Red Watermarking
To watermarking generative systems like LLMs, introducing watermarking signals deep into the generation process of the model allows the model to generate watermarked outputs directly while preserving generation performance unbiased. One possible approach is to alter the logits generation process, which does not require modifying the model parameters, and can be formulated as: 𝑙˜(𝑖 ) = A (𝑀 (t0:(𝑖 −1) ), 𝑤), (2) where the watermarking algorithm A alters the logits from the target model 𝑀 with the input of previous token t0:(𝑖 −1) to incorporate the watermark signal 𝑤 and returns the watermarked logits 𝑙˜(𝑖 ) at the 𝑖-th auto-regressive generation step. The green-red watermarking scheme [11] is based on logits modification. At each generation step, the token space is partitioned into a green set (G) and a red set (R) using a hash function conditioned
Green-Red Watermarking for Recommender Systems
Ownership Verification
Watermark Injection Semantic-Alignment Unwatermarked Recommender
Green-Red Partition
Green-Red Reproduction
Semantic-Consistent Hashing
Hash / cluster by semantics
Hash(E, key)
Hash(E’, key)
Mgreen
Item Space
Statistical Testing
Decision-Alignment
Boundary Mask
Semantic Mask
Original Top-K List
Decision-Aligned Masking
Extract top-K boundary
Mbound
Altered Item Logits 𝐳’𝐢
1
Candidate Item Logits 𝐳𝐢
Watermarked Top-K List
⨀
Green Item Hit Rate Z-test / P-value
Confidence-Alignment
Original Top-K List
Confidence-Adaptive Scaling
Scaling Weight
Reject / Accept
Entropy of logits to watermark strength
𝛼#$'&#
Watermarked Top-K List
𝛼"#$%&#
H0
Not watermarked
H1
Watermarked
Figure 2: Overview of our GREW framework for watermark injection and ownership verification. Mgreen ← Hash(E’, key)
on the previous token. For the 𝑖-th generation by the watermarked model 𝑀𝑤 , a small perturbation 𝛿 is applied to the logits of green items, leading to a higher proportion in watermarked outputs. The adjusted logit 𝑙˜(𝑖 ) for item 1 𝑣2𝑗 at3step4𝑖 is defined as: N ( (𝑖 ) 𝑙 + 𝛿, 𝑣 ∈ G, 𝑗 𝑙˜𝑗(𝑖 ) = 𝑗(𝑖 ) (3) 𝑙 , 𝑣 𝑗 ∈ R. Relevant Item 𝑗 Top-K Green Item
...
... ...
As a result, the watermark manifests as a statistically significant Soft Promotion increase in the occurrence of green items in the generated outputs, which can be detected by re-partitioning items via the hash function and evaluating the green-item ratio with a one proportion 𝑍 -test. The 𝑍 -statistic for this test is: √︁ 𝑍 = (|𝑠 |𝐺 − 𝛾𝑇 )/ 𝑇𝛾 (1 − 𝛾), (4) where |𝑠 |𝐺 is the green item count in the output, 𝛾 is the ratio of Final Injection: 𝐳’𝐢 = 𝐳 𝐢 + 𝛿"#$%&' ⋅ Mgreen⨀Mbound the green set and 𝑇 is the whole output count. We reject the null hypothesis (𝐻 0 : The output is generated with no knowledge of the green set rule) and detect the watermark if Z exceeds a certain green token threshold (with a false positive rate 𝑃 < 10−5 , when 𝑍 > 4). However, transitioning to the green-red watermarking paradigm in recommender systems is non-trivial. Naively enforcing a bias toward the green subset neglects the intrinsic decision process of recommender systems and inevitably distorts the precise ranking order, thereby compromising user satisfaction. Thus, we provide our green-red watermarking framework specifically developed for recommender systems below.
3
GREW: Green-Red Watermarking for Recommendation
In this section, we present GREW, a novel green-red watermarking framework designed to secure the intellectual property of recommender models while maintaining good recommendation utility. As illustrated in Figure 2, our approach integrates three recommendation-tailored and decision-aligned synergistic mechanisms: Semantic-Consistent Hashing ensures that the promotion
of the secret-key controlled items form coherent semantic clusters, making them indistinguishable from natural recommendations and Logits preserving normalNext-Item recommendation performance. Decision-Aligned Masking to ensure that the watermark signal is decision-aligned both semantically and stealthily to preserve the fundamental ranking logic of recommendation. And Confidence-Adaptive Scaling to dynamically adapt injection intensity based on prediction uncertainty and watermark detectability. In the following section, we detail the constituent modules of the injector and the statistical ownership verification process. A complete procedure of watermark injection is illustrated in Appendix C.
3.1
Semantic-Consistent Hashing
The primary setup of green-red watermarking is to partition the candidate item pool into green items and red items, GREW employs Semantic-Consistent Hashing. Unlike traditional schemes that randomly scatter "green" labels based on discrete IDs, our approach ensures that items with high semantic similarity in the embedding space E ∈ R | V | ×𝑑 are mapped to close hash values, thereby sharing the same green-red labels. This design promotes a collective semantic cluster of items with similar hash values. If this category aligns with user interests, the items jointly leverage this signal to easily occupy the Top-𝐾. Conversely, if the category is irrelevant, the collective boost proves futile against their intrinsically low scores, leaving them submerged at the bottom of the ranking. Semantic-Aligned Projection. The first step involves mapping the high-dimensional embedding e𝑖 of item space into a onedimensional semantic coordinate. We utilize a projection vector v𝑝𝑟𝑜 𝑗 ∈ R𝑑 , which is randomized by a specific secret key 𝐾𝑤 , ensuring that the resulting green–red partition remains confidential and resistant to external inference. This key serves as the private credential of the watermark owner and is never disclosed during deployment. Consequently, any subsequent watermark verification requires access to 𝐾𝑤 , preventing unauthorized parties from reconstructing the partition or forging ownership claims. The semantic
Lei Zhou, Min Gao, Zongwei Wang, Yibing Bai, and Wentao Li
coordinate 𝑐𝑖 for the item 𝑖 is calculated as: e𝑖 · vproj 𝑐𝑖 = √ . (5) 𝑑 According to the Johnson-Lindenstrauss lemma [9], the relative distances between items in the original embedding space are approximately preserved in this lower-dimensional projection, ensuring that items within the same semantic cluster remain proximal on the coordinate axis. Continuous Hashing. At each recommendation step 𝑡, the green-red partition of item space is randomized by the hash value of the previous sequence, following the existing green-red watermark paradigm [11]. This dynamic re-partitioning avoids persistent favoritism toward a fixed subset of items, reducing the risk of structural artifacts that could otherwise be detected or exploited. We first compute a step-dependent seed 𝑠𝑡 by combining the secret key 𝐾𝑤 with the previous interaction sequences S𝑢𝑡 −1 : 𝑠𝑡 = (𝑎 ∗ S𝑢𝑡 −1 + 𝐾𝑤 ) mod 232,
(6)
where 𝑎 is a multiplicative hashing constant. For numerical stability, the seed is normalized to the unit interval as 𝑠ˆ𝑡 = 𝑠𝑡 /232 , and evaluated in double precision to avoid inconsistencies when interacting with floating-point coordinates. Then, to generate a pseudo-random yet continuous hash value, we transform the semantic coordinate 𝑐𝑖 using a randomized sinusoidal mapping of Random Fourier Features mappings [21]. The continuous hash ℎ𝑖 is computed as: ℎ𝑖𝑡 = sin (𝑐𝑖 + 𝑠ˆ𝑡 ) · 𝜔 , (7) where 𝜔 is a frequency scaling factor governing the semantic granularity; specifically, a smaller 𝜔 yields larger continuous "green" regions in the embedding space. Finally, the green item mask Mgreen ∈ {0, 1} | V | is generated by thresholding the hash value against a target green item density 𝛾: ( 1, if ℎ𝑖𝑡 ≤ 𝛾, Mgreen,𝑖 = (8) 0, otherwise. This mechanism aligns the watermark signal with the intrinsic geometry of the embedding space, ensuring that green items form semantically consistent clusters while preserving stealthiness by secret-key hashing.
3.2
Decision-Aligned Masking
The goal of our watermarking strategy is to act as a selective catalyst: it should reinforce the model’s recommendation logic when aligned with user interests, while remaining imperceptible as a subtle "style shift" rather than a disruption. However, relying solely on the semantic-consistent mask (Mgreen ) is insufficient to achieve this balance. A naive approach that applies perturbations to all green items inevitably leads to the mis-promotion of irrelevant items. If an item is semantically "green" but intrinsically disliked by the user, artificially elevating it to the Top-𝐾 list severely compromises ranking fidelity and user experience. To mitigate this, we confine the watermark injection to a competitive subset of items that reside near the model’s decision boundary. Specifically, for a predicted logit vector z, we determine a candidate threshold 𝑧𝑡ℎ𝑟𝑒𝑠ℎ = z𝑘𝑐𝑎𝑛𝑑 , which represents the score of the
𝑘𝑐𝑎𝑛𝑑 -th ranked item (e.g., 𝑘𝑐𝑎𝑛𝑑 = 100). A boundary-aware mask Mbound ∈ {0, 1} | I | is then constructed: ( 1, if z𝑖 ≥ 𝑧 thresh, Mbound,𝑖 = (9) 0, otherwise. By applying this mask, the injector ignores irrelevant tail items (where z𝑖 ≪ 𝑧 thresh ), ensuring that only items already in the "neighborhood" of the Top-𝐾 list are considered for promotion. To combine both the boundary-aware mask and the semanticconsistent mask, the final set of promoted items for watermarking is determined by the intersection of boundary constraints and greenitem assignments. Formally, we define the integrated injection mask Minject,𝑖 as the Hadamard product of the two masks: Minject,𝑖 = Mbound,𝑖 ⊙ Mgreen,𝑖 .
(10)
This construction ensures that the watermark signal is applied if and only if the items are both semantically consistent and userfavored, therefore allowing reliable watermark signal embedding while eliminating the risk of mis-promotion and maintaining highfidelity recommendation ranking quality.
3.3
Confidence-Adaptive Scaling
Using the above two modules, the watermark can already be effectively embedded. However, we further aim to enable GREW to dynamically modulate the injection strength to balance stealthiness and verifiability. Instead of applying a uniform perturbation, which may disrupt the model’s decisive preferences, we introduce a mechanism that scales the watermark intensity based on the model’s local uncertainty and global detection performance. Confidence-Aware Local Scaling. The core intuition of our adaptive strategy is to act as a "tie-breaker" when the model is uncertain, while attenuating the signal when the model exhibits high confidence. At every single injection, we quantify this uncertainty using the Shannon entropy of the Top-𝐾 prediction. Let K be the set of indices for the Top-𝐾 items and 𝑝 𝑗 denote the 𝑗-th item. We first compute the localized softmax probabilities: 𝑝𝑗 = Í
exp(𝑧 𝑗 )
𝑘 ∈ K exp(𝑧𝑘 )
,
∀𝑗 ∈ K.
(11)
The normalized entropy is then calculated to represent the model’s confidence and used as a local watermark strength coefficient: Í − 𝑗 ∈ K 𝑝 𝑗 log 𝑝 𝑗 𝛼 local = . (12) log 𝐾 Under this formulation, when the model is highly confident (i.e., low entropy), 𝛼 local is correspondingly low, thereby shielding the high-confidence recommendation from perturbation. Global Feedback Control and Final Injection. To maintain a stable detection rate across varying data distributions, GREW incorporates a global feedback factor 𝛼 global . This factor is dynamically adjusted during training to ensure the aggregated hit rate meets a predefined target 𝜏: 𝛼 global ← 𝛼 global + 𝜂 (𝜏 − 𝑟¯),
(13)
where 𝜂 represents the adjustment step size and 𝑟¯ denotes the smoothed moving average of the validation batch hit rate. This feedback mechanism serves as a self-correcting controller: if the
Green-Red Watermarking for Recommender Systems
current hit rate 𝑟¯ falls below the target 𝜏, 𝛼 global automatically increases to strengthen the watermark signal. Combining the local uncertainty and global stability, the final watermarked logits z𝑖′ are computed as: z𝑖′ = z𝑖 + (𝛿 · 𝛼 global · 𝛼 local ) · Minject,𝑖 ,
(14)
where 𝛿 is the base injection magnitude and Minject,𝑖 is the dualmask defined in Eq. 8. This decision-aligned injection ensures that the watermark signal is "softly" embedded: it is strongest when the model deems multiple items nearly equally plausible, allowing the watermark to guide the final ranking without compromising the model’s fundamental predictive fidelity.
4
Ownership Verification Scheme of GREW
This section describes how ownership can be verified once a watermark has been embedded and a model is suspected of unauthorized use. Specifically, we establish a statistical verification framework to examine whether the recommendation outputs contain the predefined watermark signal. The presence of the watermark within the generated recommendation lists is verified by formulating the verification procedure as a hypothesis testing problem. The core objective is to determine whether the observed frequency of green items in the target model’s output exhibits a statistically significant deviation from the unwatermarked baseline, thereby ruling out coincidental occurrences. A detailed procedure of ownership verification is presented in Appendix C. Hypothesis Testing Formulation. Let R𝑢 be the Top-𝐾 recommendation list provided to user 𝑢. We define the null hypothesis 𝐻 0 as the case where the model does not contain our watermark, meaning the probability of any item in the list being selected solely based on user preferences and model parameters. Consequently, the probability of a green item appearing in the list follows the predefined target density 𝛾 with the secret-keyed partition. Conversely, the alternative hypothesis 𝐻 1 states that the model has been watermarked, leading to a higher-than-expected occurrence of green items in the Top-𝐾 results. Green Set Reproduction To perform statistical hypothesis testing, the first step is to reproduce the green-red partition of the item space. The 𝑀green is reproduced by following the exact procedure in Section 3.1, with the same secret key 𝐾𝑤 . This reproduction can only be done by the model owner; therefore, to declare strong model ownership. Let 𝑁 = |U| × 𝐾 be the total number of items inspected across the recommendation lists for all users in a test set. We count the total number of green items observed, denoted by |𝑠 |𝐺 , where I(·) is the indicator function: ∑︁ ∑︁ |𝑠 |𝐺 = I(Mgreen,𝑖 = 1). (15) 𝑢 ∈ U 𝑖 ∈ R𝑢
Aggregated Z-test Analysis. The verification is framed as an aggregated statistical test rather than a single-query verification. This design stems from the need to preserve the ranking order; naively enforcing the presence of green items would inevitably compromise the user experience. By aggregating observations across multiple user sequences, the proposed aggregated 𝑍 -test mitigates the risk of watermark detectability in cases where all green items are filtered in one sequence. This ensures robust ownership verification without sacrificing the system’s core utility.
The empirical hit rate is defined as 𝑝ˆ = |𝑠 |𝐺 /𝑁 . To determine if 𝑝ˆ significantly deviates from the expected hit rate 𝛾 (the predestinated green item density of the secret key-controlled partition of item space) of an unwatermarked recommender, we calculate the 𝑍 -score using the normal approximation of the binomial distribution: 𝑝ˆ − 𝛾 𝑍 = √︁ . 𝛾 (1 − 𝛾)/𝑁
(16)
Ownership Verification. The 𝑍 -score represents the number of standard deviations the observed hit rate 𝑝ˆ lies away from the unwatermarked base model under 𝐻 0 . The corresponding P-value, which represents the probability of a false positive (verifying the watermark where none exists), is derived from the standard normal cumulative distribution function: 𝑃 = 1 − Φ(𝑍 ). In our framework, the cumulative evidence aggregated over thousands of items drives the 𝑍 -score to high values. We adopt a rigorous detection threshold (e.g., 𝑍 > 4, where 𝑃 < 10−5 ) to claim ownership. This confirms that the overall statistical distinction between the watermarked model and the clean model is undeniable.
5
Experiments
In this section, we design empirical evaluations across different datasets and recommender models to verify the watermark injection and ownership verification performance of GREW by answering the following research questions (RQs). • RQ1: How does the proposed GREW perform compared to the existing watermarking method for sequential recommenders on both watermark validity and recommendation performance? • RQ2: Can GREW distinguish between watermarked and clean models with a significant statistical margin? • RQ3: How does GREW achieve better robustness against model extraction attacks compared to the existing baseline? • RQ4: How do the hyperparameters influence the trade-off between watermark validity and model utility? • RQ5: What are the contributions of each core module?
5.1
Experimental Setup
5.1.1 Datasets. To evaluate the performance of the proposed GREW framework, we conducted experiments on three recommendation datasets: MovieLens-1M (ML-1M) [6], Amazon Beauty [19], and Steam Review [10]. Detailed statistics for these datasets are presented in Appendix A. Following established evaluation protocols [10, 24], the leave-one-out strategy is employed for training. 5.1.2 Base Models. We evaluate our framework using three representative sequential recommendation models: • NARM [13] is an RNN-based recommender that combines an attention-based GRU unit with a bilinear scoring function to capture both global and local user intent within sessions. • SASRec (SAS) [10] is a Transformer-based model that utilizes unidirectional self-attention blocks to model long-term and shortterm dependencies in user behavior sequences. • BERT4Rec (BERT) [24] employs a bidirectional self-attention mechanism to capture complex transitions in user sequences, optimized through the Cloze task objective.
Lei Zhou, Min Gao, Zongwei Wang, Yibing Bai, and Wentao Li
Table 1: Comparison of recommendation performance (Recall/NDCG) and watermark verification on three datasets. Statistics are presented in percentages. The best results, except the Base Model, are highlighted in bold.
Model Method
MovieLens-1M R@10 R@5 N@10 N@5 𝑉 @1
Steam Review
Amazon Beauty
𝑉˜ @1 R@10 R@5 N@10 N@5 𝑉 @1
𝑉˜ @1 R@10 R@5 N@10 N@5 𝑉 @1
𝑉˜ @1
BERT
Base 20.30 12.90 10.70 8.30 19.00 15.80 14.90 13.90 2.90 1.70 1.50 AOW 19.62 11.80 10.25 7.74 100.00 26.09 18.99 15.84 14.94 13.93 100.00 24.83 2.90 1.81 1.46 GREW 20.25 12.78 10.70 8.29 100.00 95.90 19.10 15.90 15.00 13.90 100.00 100.00 3.24 1.98 1.61
1.10 1.11 100.00 25.02 1.21 100.00 100.00
SAS
Base 27.23 19.23 15.92 13.35 20.29 16.77 15.68 14.50 3.63 1.86 1.58 AOW 27.18 18.78 15.73 13.02 100.00 0.00 20.10 16.72 15.63 14.51 100.00 0.00 3.94 2.08 1.77 GREW 27.19 18.77 15.56 12.86 100.00 100.00 20.09 16.69 15.59 14.50 100.00 100.00 4.01 2.30 1.85
1.01 1.17 100.00 0.00 1.30 100.00 100.00
Base 27.56 19.55 16.23 13.65 20.31 16.75 15.62 14.74 3.49 2.15 1.81 NARM AOW 23.15 15.06 12.41 9.79 100.00 0.00 19.95 16.43 15.35 14.22 100.00 0.00 3.38 2.16 1.79 GREW 26.29 18.36 15.21 12.66 100.00 100.00 20.11 16.57 15.48 14.35 100.00 100.00 3.48 2.20 1.83
1.38 1.39 100.00 23.90 1.42 100.00 100.00
5.1.3 Evaluation Metrics. We comprehensively evaluate performance from two perspectives: recommendation utility and watermark verifiability. For utility, we employ the standard Recall@𝐾 and Normalized Discounted Cumulative Gain (NDCG@𝐾) (𝐾 ∈ {5, 10}) to ensure the watermarked model maintains high-quality ranking performance. For verifiability, we conduct statistical testing on the overall top-𝐾 list as demonstrated in Section 4, and introduce three statistical metrics to quantify the significance of the watermark signal within the model’s output: 𝑍 @𝐾 represents the 𝑍 -score measuring the deviation of the observed hit rate of green items from the theoretical target density of a clean model. Additionally, 𝑃@𝐾 denotes the corresponding probability value (𝑃-value) and 𝑉 @𝐾 = 1 − 𝑃@𝐾 denotes the confidence to claim ownership.
5.2
successfully achieves a utility-preservation level competitive with the memorization-based paradigms without any data injection.
5.3
Table 2: Statistical analysis between the unwatermarked base model and our GREW watermarked model.
Watermarking Performance
Table 1 presents the experimental results of recommendation utility and watermark verifiability. And for watermark robustness, we performed the model extraction attack, DFME[30] on both watermark methods, 𝑉˜ @1 denotes the watermark verification result of the extracted model. Firstly, in terms of watermark validity, GREW achieves 100% (𝑃 < 10−4 ) detection confidence across all datasets and architectures, demonstrating an efficacy in ownership verification that is fully on par with the state-of-the-art memorizationbased watermarking baseline (AOW [34]). However, when verifying the watermark signal from an extracted model, AOW yields a significant drop from 100% to 0% of watermark detection in most cases, while GREW maintains 100% verification, demonstrating high watermark robustness against extraction. Secondly, regarding recommendation utility, GREW exhibits better performance preservation compared to AOW in most cases. For example, on the Movielens-1M dataset with the base model NARM, GREW performs a better recommendation utility maintenance (R@10: 26.29%) than AOW (R@10: 23.15%), and more closely aligns with the unwatermarked model (R@10: 27.56%). This similarity in performance can be attributed to our decision-aligned design: unlike AOW, which relies on memorizing rigid out-of-distribution patterns, GREW subtly modulates the ranking process within the contestable item subset. By constraining watermark injection to the decision boundary rather than forcing memorization, GREW
Watermarking Verification Analysis
Model
Method
H@20
HR@20
𝑍 @20
𝑃 @20
BERT
Base GREW
36,137 54,308
29.91 44.96
114.21
< 10−6
SAS
Base GREW
40,989 70,942
33.93 58.73
182.02
< 10−6
NARM
Base GREW
10,172 17,314
32.23 54.86
86.02
< 10−6
To provide rigorous statistical proof of our watermark injection, we analyze the hit count 𝐻 @20 and hit rate 𝐻𝑅@20 of green items within the Top-20 recommendation outputs (target green item density 𝛾 = 0.33), demonstrated in Table 2. We first observe that unwatermarked Base models exhibit a natural, floating green item hit rate (𝐻𝑅@20 ≈ 30% − 34%), an expected outcome of the random semantic-consistent hashing mapping with the predestinated green density of 33%. Crucially, GREW induces a substantial distribution shift, amplifying the 𝐻𝑅@20 significantly across all architectures—most notably on SASRec, which surges by 33.93% to reach 58.73%. This deviation is statistically irrefutable, as evidenced by the exceptionally high 𝑍 -metrics (e.g., 182.02 for SASRec ) and near-zero false positive 𝑃-values (< 10−6 ), thereby confirming that GREW provides definitive ownership verification by effectively watermarking the output distribution.
5.4
The Watermark Stealthiness and Robustness Against Model Extraction
Table 3 presents a rigorous test on the MovieLens-1M dataset, designed to expose the fundamental fragility of memorization-based watermarking. To evaluate robustness, we subjected GREW to a challenging heterogeneous extraction setting (NARM → BERT),
Green-Red Watermarking for Recommender Systems
Table 3: Comprehensive comparison of stealthiness and robustness. Injection: Ratio of modified user sequences. Detectability: Recall of the watermark by an adversary. Retention: Retained watermark sequences within the extraction data.
AOW (𝑛 = 5) AOW (𝑛 = 20) GREW (Ours)
Injection
Detectability
Teacher 𝑉 @1
10.0% 10.0% 0%
100.0% 100.0% 0.00
100.0% 100.0% 100.0%
Table 4: Performance after model extraction attack. Model
Method
R@10 N@10 Agr@10 Agr@1 Z@20
BERT
Base 0.722 Base+DFME 0.754 GREW+DFME 0.771
0.531 0.514 0.534
0.498 0.419
0.220 0.211
5.056
SAS
Base 0.823 Base+DFME 0.767 GREW+DFME 0.710
0.607 0.530 0.447
0.424 0.326
0.212 0.233
6.470
26.09% 0.00% 95.90%
100.00% 6.51% 99.65%
100.00% 11.21% 99.94%
100.00% 26.13% 100.00%
0% 0% 100.00%
Table 4 (Metrics here were calculated with 100 negative samples, following the model extraction attack method, DFME [30]) highlights the dual capability of GREW as both an ownership verification tool and an active defense mechanism against model extraction. We focus specifically on the Agreement metrics (Agr@10), which measure the functional similarity between the extracted model output and the victim teacher output: |Ttop−𝐾 ∩ Stop−𝐾 | , (17) 𝐾 where Ttop−𝐾 denotes the top-k recommendation list predicted by the target black-box model, and Stop−𝐾 denotes the top-k recommendation list predicted by the extracted surrogate model. The results show a consistent decline in model extraction fidelity when the adversary targets a watermarked model compared to attacking an unwatermarked model; for instance, in the SASRec setting, Agr@10 drops significantly from 0.424 to 0.326. This indicates that the subtle perturbations introduced by the watermark effectively shift the output distribution, making it harder for the adversary to perfectly clone the teacher’s recommendation ability. Crucially, despite this degradation in extraction quality, the ownership signal remains robustly embedded: the Z-scores in the extracted models reach 5.056 (BERT) and 6.470 (SAS), far exceeding the statistical significance threshold (𝑍 > 4.0), ensuring that any stolen copy, however imperfect, remains undeniably traceable. Agr@K =
5.5
Hyperparameter Analysis R10
R5
N10
Candidate Size (kcand )
25
N5
Z@20
Base Magnitude ( )
100
20
80
15
60
10
40
5
20
0
30
50
100
200
1
2
5
10
Z@20 Score
whereas the baseline AOW was tested under the standard homogeneous setting (BERT → BERT). Despite this disadvantage, the results reveal a striking contrast in robustness and stealthiness. First, using a simple anomaly detector based on item popularity and transition patterns, we observe that AOW—driven by a brute-force 10.0% injection strategy—is completely exposed (100.0% Detectability), rendering it distinguishable from natural user behaviors. It is worth noting that while AOW can achieve high recall with shorter triggers (e.g., lengths 𝑛 = 2 or 𝑛 = 5 ), such brief sequences lack the statistical significance required for ownership verification; specifically, a length-2 fixed trigger pattern is indistinguishable from natural user interactions with niche items, creating high falsepositive risks that undermine its legal validity. Crucially, AOW fails to survive model extraction in the easier homogeneous setting: despite effective verification on the watermarked Teacher model, it suffers a catastrophic collapse on the extracted Student model, with V@1 plummeting to 0.00%, and no watermark sequences are retained in the extraction data. We attribute this to the "denoising effect" of the model extraction attack, where the student filters out the non-robust OOD noise forced by AOW. In contrast, GREW embeds the watermark signal deep into the semantic space, which the distillation process inherently aims to replicate by learning the teacher’s semantic distribution. Because the watermark is encoded as part of the underlying semantics, where green items remain semantically consistent with user preferences, it is preserved during knowledge transfer rather than treated as noise. Consequently, the student model internalizes the watermark while approximating the teacher’s ranking ability. Even under the harder heterogeneous extraction, GREW requires zero data injection and transforms the watermark into a robust semantic signal deep into the model, maintaining a 𝑉 @1 of 95.90% and 𝑉 @20 of 100.00% in the extracted student model. This proves that watermarking and aligning with the model’s intrinsic ranking logic is the viable path for persistent ownership verification.
Robustness against Model Extraction Student 𝑉˜ @1 𝑉˜ @5 𝑉˜ @10 𝑉˜ @20 Retention
R & N Score
Stealthiness Method
0
Figure 3: Watermarking performance on different hyperparameter settings under ML-1M and BERT4Rec as base model. We further investigate the sensitivity of GREW to two key hyperparameters: the size of the candidate item pool 𝑘𝑐𝑎𝑛𝑑 for bias injection and the base watermark injection magnitude 𝛿. First, regarding the candidate pool size 𝑘𝑐𝑎𝑛𝑑 , we observe a fundamental
Lei Zhou, Min Gao, Zongwei Wang, Yibing Bai, and Wentao Li
trade-off between utility and verifiability. As shown in Figure 3, an overly restrictive pool (𝑘𝑐𝑎𝑛𝑑 = 30) causes model collapse, reducing R@10 to 9.72% and failing to inject a valid watermark signal. Expanding the pool from 50 to 200 restores model utility (19.35%), but it gradually dilutes the watermark signal (yet verifiable with 𝑃-value < 10−6 ). We identify 𝐾𝑐𝑎𝑛𝑑 = 50 as the optimal setting, achieving a high 𝑍 -score of 74.69 while maintaining recommendation utility. Second, regarding the base injection magnitude 𝛿, increasing 𝛿 from 1 to 10 all provides 𝑍 -scores high enough to claim model ownership. Notably, setting 𝛿 = 2 offers the most balanced performance, achieving a superior R@10 of 20.25% while maintaining a robust verification confidence (𝑍 =66.52), demonstrating that a moderate but precise bias is sufficient for effective watermarking.
5.6
large language models, to synthesize informative surrogate data that closely approximates the target distribution [25, 26, 39]. These attacks are further amplified by optimized sampling constraints and distillation objectives, allowing adversaries to replicate complex ranking behaviors even under strict query budgets [15]. Once a high-quality surrogate is acquired, the adversaries are able to gain a surrogate white-box control, develop target attacks, or unauthorized replication. Since the memorization-based watermarks manifest as statistical outliers rather than intrinsic preferences, they are susceptible to being pruned during the surrogate learning process or explicitly removed. This vulnerability highlights the critical need for watermarking mechanisms that are intrinsically aligned with the recommendation process, ensuring persistence even when the model is subjected to high-fidelity extraction.
Ablation Study
In Table 5, we further conducted an ablation study to validate the contribution of each designed component. First, replacing the Semantic-Consistent Hash module with naive ID-based hashing ("w/o Semantic") reduces detectability (Z@20 drops from 87.39 to 68.93), confirming that, unlike simple ID-based hashing, our semantic-aligned approach enables the model to internalize the watermark as a consistent preference pattern. Second, regarding the Confidence-Adaptive Scaling mechanisms, removing the entropybased local scaling constraint ("w/o 𝛼 local ") results in the lowest recommendation utility preservation (R@10 drops to 19.33%), indicating that a uniform bias is detrimental for hard samples; similarly, the absence of global feedback control ("w/o 𝛼 global ") disrupts the overall equilibrium, leading to suboptimal performance in both utility and detection. Consequently, the full GREW outperforms all variants, demonstrating that semantic-aligned hashing coupled with multi-granularity (global and local) strength scaling is essential for achieving the optimal balance between verification robustness and recommendation quality. Table 5: Watermarking performance with different variants under ML-1M and BERT4Rec as base model. Method
R@10
R@5
N@10
N@5
𝑍 @20
GREW w/o Semantic w/o 𝛼 local w/o 𝛼 global
20.45 19.72 19.33 19.85
12.98 12.30 12.34 12.33
10.68 10.28 10.19 10.33
8.28 7.90 7.96 7.92
87.39 68.93 74.75 85.00
6 Related Work 6.1 Model Extraction in Recommender Systems Intellectual property theft in recommender systems increasingly manifests through model extraction attacks, where adversaries reconstruct proprietary algorithms by systematically exploiting public API interfaces to train high-fidelity surrogate models [30]. Although deployed systems typically restrict outputs to Top-K rankings, attackers have developed sophisticated mechanisms to overcome this information bottleneck. Contemporary extraction frameworks leverage auto-regressive querying strategies and advanced generative backbones, including graph neural networks and
6.2
Watermarking in Recommender Systems
Model watermarking has been widely studied as a means of protecting the intellectual property of machine learning models, by embedding property-related signals or triggers into the model itself [32]. Existing approaches can be broadly categorized into whitebox watermarking, which embeds signals into model parameters or internal representations, and black-box watermarking, which verifies ownership solely through model outputs [4]. In particular, white-box watermarking often needs modification of the model’s inner weights and structures, which impractically requires full access to the model infrastructure during the detection process [31]. On the other hand, black-box watermarking methods often rely on trigger-based inputs or predefined signals that the model is trained to memorize, with ownership verified through strong watermarked outputs [22, 27]. Recent progress in generative models has motivated green-red watermarking techniques that embed predefined statistical patterns into the generation process. These methods enable GenAI models to maintain normal generation quality while producing outputs that conform to owner-specified rules, allowing ownership to be verified via statistical hypothesis testing of the generated results [11, 16, 17, 20]. While these techniques have shown effectiveness in classification and generation tasks, they typically assume independent prediction targets and do not consider the structured decision mechanisms present in recommender systems. Compared to general-purpose model watermarking, watermarking for recommender systems remains relatively underexplored. Existing studies often inject watermarks by forcing the model to memorize abnormal or rarely occurring interaction sequences [34], special trigger users [29, 35], or user-item pairs [4]. However, such designs are not naturally aligned with recommendation behaviors and may introduce noticeable distribution shifts, making the watermark susceptible to detection, removal, or performance degradation. Moreover, these approaches overlook the ranking-based nature of recommendation, treating item predictions independently rather than as competitive decisions deep in the model. These limitations motivate the need for watermarking mechanisms that are inherently compatible with recommendation decision processes.
7
Conclusion
We presented GREW, the first Green-Red watermarking framework designed specifically for recommender systems. Departing from
Green-Red Watermarking for Recommender Systems
the conventional strategy of injecting detectable synthetic data triggers, GREW leverages a secret key to stealthily partition the item space into green-red sets and subtly modulate the ranking probability of green items. Strong ownership verification can be carried out via key-controlled re-partitioning and statistical hypothesis testing on black-box outputs. Through the synergistic design of Semantic-Consistent Hashing, Decision-Aligned Masking, and Confidence-Adaptive Scaling, our framework successfully embeds robust watermark signals directly into the ranking behavior of a model while preserving the accuracy of Top-𝐾 recommendations. Our evaluation confirms that GREW offers strong verifiability and robustness against model extraction attack. Future work may explore extending this paradigm to graph or generative recommendation architectures to further broaden the scope of intellectual property protection.
References [1] J. Bobadilla, F. Ortega, A. Hernando, and A. Gutiérrez. 2013. Recommender systems survey. Knowledge-Based Systems 46 (July 2013), 109–132. doi:10.1016/j. knosys.2013.03.012 [2] Jiawei Chen, Hande Dong, Xiang Wang, Fuli Feng, Meng Wang, and Xiangnan He. 2023. Bias and Debias in Recommender System: A Survey and Future Directions. ACM Trans. Inf. Syst. 41, 3 (Feb. 2023), 67:1–67:39. doi:10.1145/3564284 [3] Ingemar Cox, Matthew Miller, Jeffrey Bloom, and Chris Honsinger. 2002. Digital watermarking. Journal of Electronic Imaging 11, 3 (2002), 414–414. [4] Xiaocui Dang, Priyadarsi Nanda, Heng Xu, Haiyu Deng, and Manoranjan Mohanty. 2024. Recommendation System Model Ownership Verification via NonInfluential Watermarking. In 2024 17th International Conference on Security of Information and Networks (SIN). 1–8. doi:10.1109/SIN63213.2024.10871674 [5] Jiaxin Deng, Shiyao Wang, Kuo Cai, Lejian Ren, Qigen Hu, Weifeng Ding, Qiang Luo, and Guorui Zhou. 2025. Onerec: Unifying retrieve and rank with generative recommender and iterative preference alignment. arXiv preprint arXiv:2502.18965 (2025). [6] F. Maxwell Harper and Joseph A. Konstan. 2015. The MovieLens Datasets: History and Context. ACM Trans. Interact. Intell. Syst. 5, 4 (Dec. 2015), 19:1–19:19. doi:10.1145/2827872 [7] Frank Hartung and Martin Kutter. 2002. Multimedia watermarking techniques. Proc. IEEE 87, 7 (2002), 1079–1107. [8] Jiahao Huo, Shuliang Liu, Bin Wang, Junyan Zhang, Yibo Yan, Aiwei Liu, Xuming Hu, and Mingxun Zhou. 2025. PMark: Towards Robust and Distortionfree Semantic-level Watermarking with Channel Constraints. arXiv preprint arXiv:2509.21057 (2025). [9] William B Johnson, Joram Lindenstrauss, et al. 1984. Extensions of Lipschitz mappings into a Hilbert space. Contemporary mathematics 26, 189-206 (1984), 1. [10] Wang-Cheng Kang and Julian McAuley. 2018. Self-Attentive Sequential Recommendation. In 2018 IEEE International Conference on Data Mining (ICDM). 197–206. doi:10.1109/ICDM.2018.00035 ISSN: 2374-8486. [11] John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein. 2023. A Watermark for Large Language Models. In Proceedings of the 40th International Conference on Machine Learning. PMLR, 17061–17084. https://proceedings.mlr.press/v202/kirchenbauer23a.html ISSN: 2640-3498. [12] Kalpesh Krishna, Gaurav Singh Tomar, Ankur P. Parikh, Nicolas Papernot, and Mohit Iyyer. 2019. Thieves on Sesame Street! Model Extraction of BERT-based APIs. https://openreview.net/forum?id=Byl5NREFDr [13] Jing Li, Pengjie Ren, Zhumin Chen, Zhaochun Ren, Tao Lian, and Jun Ma. 2017. Neural attentive session-based recommendation. In Proceedings of the 2017 ACM on Conference on Information and Knowledge Management. 1419–1428. https: //dl.acm.org/doi/abs/10.1145/3132847.3132926 [14] Yuqing Liang, Jiancheng Xiao, Wensheng Gan, and Philip S Yu. 2024. Watermarking techniques for large language models: A survey. arXiv preprint arXiv:2409.00089 (2024). [15] Fu Liu, Hui Zhang, Yuqin Lan, and Min Li. 2025. FewMEA: Few-shot Model Extraction Attack against Sequential Recommenders. In Proceedings of the 2025 International Conference on Multimedia Retrieval (ICMR ’25). Association for Computing Machinery, New York, NY, USA, 917–925. doi:10.1145/3731715.3733340 [16] Yijian Lu, Aiwei Liu, Dianzhi Yu, Jingjing Li, and Irwin King. 2024. An Entropybased Text Watermarking Detection Method. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 11724–11735. [17] Minjia Mao, Dongjun Wei, Zeyu Chen, Xiao Fang, and Michael Chau. 2025. Watermarking Large Language Models: An Unbiased and Low-risk Method.
In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (Eds.). Association for Computational Linguistics, Vienna, Austria, 7939–7960. doi:10.18653/v1/2025.acl-long.391 [18] Alex Martinez, Mihnea Tufis, and Ludovico Boratto. 2024. Unmasking Privacy: A Reproduction and Evaluation Study of Obfuscation-based Perturbation Techniques for Collaborative Filtering. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’24). Association for Computing Machinery, New York, NY, USA, 1753–1762. doi:10.1145/3626772.3657858 [19] Jianmo Ni, Jiacheng Li, and Julian McAuley. 2019. Justifying recommendations using distantly-labeled reviews and fine-grained aspects. In Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP). 188–197. [20] Michael-Andrei Panaitescu-Liess, Zora Che, Bang An, Yuancheng Xu, Pankayaraj Pathmanathan, Souradip Chakraborty, Sicheng Zhu, Tom Goldstein, and Furong Huang. 2025. Can Watermarking Large Language Models Prevent Copyrighted Text Generation and Hide Training Data? Proceedings of the AAAI Conference on Artificial Intelligence 39, 23 (April 2025), 25002–25009. doi:10.1609/aaai.v39i23. 34684 [21] Ali Rahimi and Benjamin Recht. 2007. Random features for large-scale kernel machines. Advances in neural information processing systems 20 (2007). [22] Preston K Robinette, Thuy Dung Nguyen, Samuel Sasaki, and Taylor T Johnson. 2025. Trigger-Based Fragile Model Watermarking for Image Transformation Networks. In European Symposium on Research in Computer Security. Springer, 346–365. [23] Zijie Song, Jiawei Chen, Sheng Zhou, Qihao Shi, Yan Feng, Chun Chen, and Can Wang. 2023. CDR: Conservative doubly robust learning for debiased recommendation. In Proceedings of the 32nd ACM International Conference on Information and Knowledge Management. 2321–2330. [24] Fei Sun, Jun Liu, Jian Wu, Changhua Pei, Xiao Lin, Wenwu Ou, and Peng Jiang. 2019. BERT4Rec: Sequential Recommendation with Bidirectional Encoder Representations from Transformer. In Proceedings of the 28th ACM International Conference on Information and Knowledge Management (CIKM ’19). Association for Computing Machinery, New York, NY, USA, 1441–1450. doi:10.1145/3357384.3357895 [25] Yihao Wang, Jiajie Su, Chaochao Chen, Meng Han, Chi Zhang, and Jun Wang. 2025. Sim4Rec: Data-Free Model Extraction Attack on Sequential Recommendation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 12766–12774. https://ojs.aaai.org/index.php/AAAI/article/view/33392 Issue: 12. [26] Zeyu Wang, Yidan Song, Shihao Qin, Yu Shanqing, Yujin Huang, Qi Xuan, and Xin Zheng. 2025. Data-Free Model Extraction for Black-box Recommender Systems via Graph Convolutions. https://openreview.net/forum?id=eS3xJTjSgm [27] Zhenyi Wang, Yihan Wu, and Heng Huang. 2024. Defense against model extraction attack by bayesian active watermarking. In Forty-first International Conference on Machine Learning. [28] Zongwei Wang, Junliang Yu, Min Gao, Wei Yuan, Guanhua Ye, Shazia Sadiq, and Hongzhi Yin. 2024. Poisoning Attacks and Defenses in Recommender Systems: A Survey. CoRR abs/2406.01022 (2024). doi:10.48550/ARXIV.2406.01022 arXiv: 2406.01022. [29] Enyue Yang, Weike Pan, Lixin Fan, Hanlin Gu, Zhitao Li, Qiang Yang, and Zhong Ming. 2025. Ownership Verification for Federated Recommendation. ACM Trans. Inf. Syst. 43, 3 (March 2025), 69:1–69:27. doi:10.1145/3715320 [30] Zhenrui Yue, Zhankui He, Huimin Zeng, and Julian McAuley. 2021. BlackBox Attacks on Sequential Recommenders via Data-Free Model Extraction. In Proceedings of the 15th ACM Conference on Recommender Systems (RecSys ’21). Association for Computing Machinery, New York, NY, USA, 44–54. doi:10.1145/ 3460231.3474275 [31] Jie Zhang, Dongdong Chen, Jing Liao, Weiming Zhang, Huamin Feng, Gang Hua, and Nenghai Yu. 2021. Deep model intellectual property protection via deep watermarking. IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 8 (2021), 4005–4020. [32] Jialong Zhang, Zhongshu Gu, Jiyong Jang, Hui Wu, Marc Ph Stoecklin, Heqing Huang, and Ian Molloy. 2018. Protecting intellectual property of deep neural networks with watermarking. In Proceedings of the 2018 on Asia conference on computer and communications security. 159–172. [33] Jun Zhang, Yi Li, Yue Liu, Changping Wang, Yuan Wang, Yuling Xiong, Xun Liu, Haiyang Wu, Qian Li, Enming Zhang, et al. 2025. GPR: Towards a Generative Pre-trained One-Model Paradigm for Large-Scale Advertising Recommendation. arXiv preprint arXiv:2511.10138 (2025). [34] Sixiao Zhang, Cheng Long, Wei Yuan, Hongxu Chen, and Hongzhi Yin. 2024. Watermarking Recommender Systems. In Proceedings of the 33rd ACM International Conference on Information and Knowledge Management (CIKM ’24). Association for Computing Machinery, New York, NY, USA, 3217–3226. doi:10.1145/3627673. 3679617 [35] Sixiao Zhang, Cheng Long, Wei Yuan, Hongxu Chen, and Hongzhi Yin. 2025. Data Watermarking for Sequential Recommender Systems. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 (KDD
Lei Zhou, Min Gao, Zongwei Wang, Yibing Bai, and Wentao Li
’25). Association for Computing Machinery, New York, NY, USA, 3819–3830. doi:10.1145/3711896.3736903 [36] Bin Zhao, Chuangbai Xiao, Yu Zhang, Peng Zhai, and Zhi Wang. 2019. Assessment of recommendation trust for access control in open networks. Cluster Computing 22, 1 (Jan. 2019), 565–571. doi:10.1007/s10586-017-1338-x [37] Kaixiang Zhao, Lincan Li, Kaize Ding, Neil Zhenqiang Gong, Yue Zhao, and Yushun Dong. 2025. A Survey on Model Extraction Attacks and Defenses for Large Language Models. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 (KDD ’25). Association for Computing Machinery, New York, NY, USA, 6227–6236. doi:10.1145/3711896.3736573 [38] Xu Zhao, Ruibo Ma, Jiaqi Chen, Weiqi Zhao, Ping Yang, and Yao Hu. 2025. Multi-Granularity Distribution Modeling for Video Watch Time Prediction via Exponential-Gaussian Mixture Network. In Proceedings of the Nineteenth ACM Conference on Recommender Systems. 309–318. [39] Lei Zhou, Min Gao, Zongwei Wang, and Yibing Bai. 2025. Budget and Frequency Controlled Cost-Aware Model Extraction Attack on Sequential Recommenders. In Proceedings of the 34th ACM International Conference on Information and Knowledge Management (CIKM ’25). Association for Computing Machinery, New York, NY, USA, 4477–4486. doi:10.1145/3746252.3761032
A
Dataset Statistics
To provide a comprehensive overview of the data scale and structural characteristics, the detailed statistics of the datasets utilized in our empirical evaluations are summarized in Table 6. Table 6: Statistics of the datasets.
B
Dataset
#Users
#Items
#Interacts
Avg.len
Density
ML-1M Beauty Steam
6,040 40,226 334,542
3,416 54,542 13,046
1,000,209 353,989 3,546,145
163.5 8.8 10.6
4.84% 0.02% 0.10%
C
Pseudo-codes
Algorithm 1 GREW Watermark Injection Procedure 1: Input: Logits z, item embeddings E, key 𝐾𝑤 , target density 𝛾,
target hit rate 𝜏, step size 𝜂, momentum 𝑚. 2: Buffers: Global strength 𝛼 global (init: 𝛿𝑏𝑎𝑠𝑒 ), running hit rate 𝑟¯. 3: Output: Watermarked logits z′ . 4: // Step 1: Semantic-Consistent Hashing 5: vproj ← PRNG(𝐾𝑤 );
𝜔 ← 2𝜋
𝑡 −1 + 𝐾 ) mod 232 , 6: 𝑠𝑡 ← (𝑎 · S𝑢 𝑤
𝑠ˆ𝑡 ← 𝑠𝑡 /232
7: for each item 𝑖 ∈ V √ do
𝑐𝑖 ← (e𝑖 · vproj )/ 𝑑 ℎ𝑖𝑡 ← | sin((𝑐𝑖 + 𝑠ˆ𝑡 ) · 𝜔)| ⊲ Double precision mapping 10: Mgreen,𝑖 ← I(ℎ𝑖𝑡 < 𝛾) 11: end for 12: // Step 2: Decision-Aligned Masking 13: 𝑧 thresh ← Top-Value(z, 𝑘𝑐𝑎𝑛𝑑 ) 14: Minject ← I(z ≥ 𝑧 thresh ) ⊙ Mgreen 15: // Step 3: Confidence-Adaptive Scaling 16: 𝑝 𝑗 ← Softmax(z𝑡𝑜𝑝𝐾 ) 𝛽 Í 17: 𝛼 local ← − 𝑝 𝑗 log 𝑝 𝑗 /log 𝐾 ′ 18: z ← z + (𝛼 global · 𝛼 local ) · Minject 19: if is_training then 20: 𝑟𝑐𝑢𝑟𝑟 ← Mean(Mgreen in Top-𝐾 of z′ ) 21: 𝑟¯ ← 𝑚 · 𝑟¯ + (1 − 𝑚) · 𝑟𝑐𝑢𝑟𝑟 ⊲ Running average update 22: 𝛼 global ← Clip(𝛼 global + 𝜂 · (𝜏 − 𝑟¯), 𝛿𝑚𝑖𝑛 , 𝛿𝑚𝑎𝑥 ) 23: end if 24: return z′ 8:
9:
Implementation Details
The detailed training configurations for each base model are summarized in Table 7. We employ the Adam optimizer across all architectures, with a weight decay of 0.01, a learning rate of 0.001, and a batch size of 512 (on a single NVIDIA GeForce RTX 4090 GPU with 24GB of memory). To prevent overfitting, the dropout rates for the ML-1M, Beauty, and Steam datasets are meticulously tuned to 0.1, 0.5, and 0.2, respectively. Following the protocols established in the existing methods [30, 34], the maximum sequence lengths are set to 200 for ML-1M and 50 for both Beauty and Steam. Regarding our GREW framework, the green-item density 𝛾 is set to 0.5, and the base injection magnitude 𝛿 is initialized at 0.1. The candidate pool size 𝑘𝑐𝑎𝑛𝑑 is defined as top-100, and the global feedback scaling is configured to maintain a target aggregated hit rate 𝜏 = 0.65. Table 7: Model configurations. ly: layer, h: attention head, mp: masking probability. Phase
Model
Config. on {ML-1M, Steam, Beauty}
Training
NARM SAS BERT
GRU ly:1; TRM ly:2; h:2; TRM ly:2; h:2; mp={0.2, 0.2, 0.6}.
Algorithm 2 GREW Ownership Verification Procedure 1: Input: Suspected model M, test users 𝑈 , interaction histories
{S𝑢𝑡 −1 }𝑢 ∈𝑈 , secret key 𝐾𝑤 , item embeddings E, target density 𝛾. 2: Output: 𝑍 -score, 𝑃-value, 3: Step 1: Green Set Reproduction 4: Initialize |𝑠 |𝐺 ← 0, 𝑁 ← |U| × 𝐾 5: vproj ← PRNG(𝐾𝑤 ) 6: for each user 𝑢 ∈ 𝑈 do ⊲ Reproduce partition seed 7: Obtain Top-𝐾 recommendation list R𝑢 from M 8: Compute seed 𝑠ˆ𝑡 using 𝐾𝑤 and history S𝑢𝑡 −1 9: for each item 𝑖 ∈ R𝑢√do 10: 𝑐𝑖 ← (e𝑖 · vproj )/ 𝑑 11: ℎ𝑖𝑡 ← | sin((𝑐𝑖 + 𝑠ˆ𝑡 ) · 𝜔)| 12: if ℎ𝑖𝑡 ≤ 𝛾 then 13: |𝑠 |𝐺 ← |𝑠 |𝐺 + 1 ⊲ Count observed green items 14: end if 15: end for 16: end for 17: Step 2: Statistical Hypothesis Testing 18: 𝑝ˆ ← |𝑠 |𝐺 /𝑁 ⊲ Empirical hit rate 𝑝ˆ −𝛾 19: 𝑍 ← √ ⊲ Deviation from null hypothesis 𝐻 0 𝛾 (1−𝛾 )/𝑁
20: 𝑃 ← 1 − Φ(𝑍 )
⊲ Calculate 𝑃-value via standard normal CDF