PEMark:Watermarking API Responses Based on Proxy Gateways and Position Encoding Yifei Zhou1 , Xianjun Gu1,2,∗ , Xinyu Dai1 , Ming Liu1,3 , Lansheng Han∗1,3
arXiv:2605.21865v1 [cs.CR] 21 May 2026
1
Huazhong University of Science and Technology, Wuhan, China 2 State Grid Hubei Electric Power Co., Ltd., Wuhan, China 3 Wuhan Jinyinhu Laboratory, Wuhan, China [email protected]
Abstract—Data leakage from API responses has drawn wide attention. APIs are often not fully regulated, making them easy to abuse. One common solution is to embed watermarks into API responses for traceability. However, existing watermarking methods often require modifying database content or API response data. This forces changes to business system code, and may even disrupt normal business operations because data values are altered. In this paper, we propose an original pluggable watermarking scheme based on a watermark proxy gateway and PEMark (Position Encoding-based Watermarking). The key novelty of our approach is exploiting the inherent permutation redundancy in the ordering of JSON/XML key-value pairs—an overlooked dimension that carries no semantic information yet provides abundant encoding capacity. First, we forward server responses to the watermark proxy gateway, a design that requires zero modification to existing business systems. Then, we embed a watermark into each API response using position encoding, which reorders keys without altering any data values. To the best of our knowledge, this is the first work to achieve distortion-free API response watermarking via position encoding over a proxy gateway. Our method does not modify any data values, so normal business operations continue seamlessly after watermark embedding. Experimental results show that our framework maintains business usability while ensuring that returned API data is traceable. Compared with current mainstream schemes, our method is robust against tampering and insertion attacks (100% similarity), and can withstand certain levels of deletion attacks. Index Terms—API response watermarking, proxy gateway, position encoding, traceability, robustness.
I. Introduction The API (Application Programming Interface) is widely used for service provision, data acquisition, and AI model responses. However, when an API is not properly authorized or returns sensitive data, malicious calls may cause serious data leakage. A common solution is to embed watermark data into the database, and then use that watermarked data during API interactions [1]–[6]. Another approach is to modify semi-structured JSON data in API responses to embed watermarks [7]–[10]. As shown in Fig. 1, these approaches enable tracing the leakage source by embedding identifiable information either in backend data or API responses. Sadly, adding watermarks to a database can interfere with normal business operations. It often forces developers to change the business code [11]–[14]. Modifying the
semi-structured data [7], [15] returned by an API also brings uncertainty to the business. Even a small change, such as a shift in decimal places, may cause business errors [16]. In short, both types of solutions face big challenges in practice because they require changes to existing business code [14]. Business systems are often unwilling to make such changes [17]. In recent years, some zero-watermarking methods have been proposed to avoid data distortion. For example, Zhang et al. [18] proposed PhiMark, a zero-distortion robust database watermarking scheme using a partitionand-mask approach. Tang et al. [19] proposed PKMark, a zero-distortion blind reversible scheme based on primary key marking. Han et al. [20] proposed a distortionfree scheme using auxiliary data encoding. Zhao et al. [10] used a robust solution distribution-based zerowatermarking method to trace semi-structured power data. Other zero-watermarking schemes [21]–[24] also rely on feature libraries. However, the main drawback of these methods is that they need to maintain a trusted thirdparty feature library for watermark extraction. We surveyed watermarking methods for API interfaces in recent years, as shown in Table I. There are three types of methods that can protect API responses: modifying structured data [4]–[6], [13], [25]–[32], modifying semistructured data [8]–[10], and zero-watermark [18], [21]– [23]. Besides these, frequency-based watermarking [33] and format-independent frameworks [34] have also been proposed. Most studies focus on structured data such as databases and CSV files. Few studies focus on protecting semi-structured data like JSON and XML. Although existing schemes [8], [18] can be applied to semi-structured data, they have several weaknesses: they need a thirdparty feature library, they require changes to existing business code, and they affect normal business operations. These issues hinder the adoption of such technologies. Our proposed method based on position encoding avoids the above issues. The key originality of this work lies in observing a little-known redundancy — the order of JSON/XML elements returned by an API call. This is a subtle property that has been largely overlooked in the watermarking literature. This order redundancy gives us a novel insight: we can embed data
Fig. 1. API data leakage risks and watermarking-based solutions. TABLE I Comparison of Different Watermarking Methods Method Classification Modify structured data [4], [5], [25]–[27] Modify semi-structured data [8], [9] Zero-Watermark [10], [18], [21] Position-Encoding (Ours)
Requires Business Code Change ✓ ✓ × ×
into the order using position encoding, thereby achieving distortion-free watermarking that requires neither data value modification nor a third-party feature library. To the best of our knowledge, we are the first to exploit keyordering redundancy for API response watermarking via a proxy gateway architecture. The main contributions of this paper are summarized as follows. We are the first to identify and exploit key-ordering redundancy in semi-structured API responses as a watermarking channel, and propose a novel position encoding algorithm that embeds watermarks without changing any data values, thus leaving the original business completely unaffected. • We design an original pluggable watermarking framework based on a proxy gateway for on-the-fly protection of API response data [35], requiring zero
•
Requires Third-Party Library × × ✓ ×
Affects Business Usability ✓ ✓ × ×
modification to existing server or client code. Compared with current mainstream schemes, our method is highly robust. It fully resists numerical tampering attacks and insertion attacks (100% similarity), and can withstand certain levels of deletion attacks — all while maintaining negligible time overhead (below 0.65 ms).
•
The rest of this paper is organized as follows. Section II describes our proposed watermark proxy gateway, which proxies API responses from the server for watermark embedding. Section III presents the position encoding method used to embed watermarks. Section IV presents and discusses our experimental results. Section V discusses the limitations of our scheme and future work. Finally, Section VI concludes the paper.
II. Watermarking Proxy Gateway We use a proxy gateway to embed watermarks into API responses. As illustrated in Fig. 2, the gateway is deployed between the API server and the client, intercepting each response before it reaches the client. This design is inspired by dynamic watermarking solutions for web content [35] and distortion‑free watermarking for relational databases [23].
In each group Ji , the total number of possible permutations is n!. This means we have n! different orders that can be used in Ji . Each group Ji is used to embed one complete watermarks. A capacity threshold T is also needed for the watermark, because each group Ji needs enough keys to carry a watermark of a certain length. The values of g and T balance watermark capacity against robustness. We discuss this trade-off in Section IV-D. 2) Factorial Decomposition: Given a binary watermark sequence M = m1 , m2 , ..., mx , The value of M is decomposed into a sequence of factorial coefficients through successive division with remainder as formula (2). For the capacity threshold T , we compute coefficients a1 , a2 , . . . , aT −1 satisfying: M=
T −1 X
aj × (T − j)!
(2)
j=1
Fig. 2. Proxy gateway positioned between API server and client.
Upon receiving a response from the server, the gateway extracts the JSON data and applies the position encoding method described in Section III to embed the watermark. Interestingly, the gateway operates transparently: neither the server nor the client is modified. The server continues to function as usual, while the client receives the watermarked response without any awareness of the change. This design preserves the original API logic and requires no modifications to business code. III. Watermarking Scheme This section introduces PEMark (Position Encodingbased Watermarking), a scheme that leverages the permutation characteristics of keys in semi-structured data to embed watermarks. We take JSON as an example to illustrate how position encoding works on semi-structured data. The entire scheme consists of two core components: embedding and extraction, as illustrated in Fig. 3.
where each factorial coefficient aj ranges in [0, T − j − 1] and corresponds to the index of the selected key in the remaining key set at each step. 3) Group-based Reordering: Fig. 4 illustrates how the coefficient sequence from factorial decomposition determines the key order within a group. The coefficients (a1 , a2 , . . . , aT −1 ) constitute a Lehmer code [36] — a bijection between integers [0, T ! − 1] and all T ! permutations. A Lehmer code (L1 , . . . , LT ) satisfies 0 ≤ Lk ≤ T − k, where Lk counts how many remaining elements are smaller than the current pick; our final coefficient aT is implicitly zero. For a group Ji with keys {k1 , . . . , kT }, let the sorted key sequence be: K(0) = k(1) , k(2) , . . . , k(T ) ,
k(1) < k(2) < · · · < k(T ) , (3) where the ordering is lexicographic. The watermarked permutation π = (π1 , π2 , . . . , πT ) is then constructed recursively: πj = K(j−1) [aj ] K(j) = K(j−1) \ {πj } πT = K(T −1) [0]
, j = 1, 2, ..., T − 1
(4)
A. Watermark Embedding The essence of watermark embedding is to convert a watermark value into a specific order of keys in JSON data. The embedding process is shown in Fig. 3. 1) Data Grouping: For a JSON object J that contains N keys, denoted as {j1 , j2 , ..., jN }, we group J into g groups: J1 , J2 , ..., Jg . This process can be expressed as:
where K(j−1) [aj ] denotes the element at index aj (zerobased) in the ordered sequence K(j−1) , and \ denotes set removal. The constraint 0 ≤ aj ≤ T − j − 1 guarantees that the index is always valid. The resulting watermarked group is:
J = {J1 , J2 , ..., Jg }
After all g groups are processed, any remaining keys (those not forming a complete group) are appended in their original order. These leftover keys carry no watermark information.
(1)
where g is the number of groups, and each group Ji contains g = ⌊N/T ⌋ keys.
Jw(i) = (π1 , vπ1 ), (π2 , vπ2 ), . . . , (πT , vπT ) .
(5)
Fig. 3. Overview of PEMark watermark embedding and extraction process.
B. Watermark Extraction Extraction recovers the watermark by inverting the permutation process. Given the watermarked data Jw , the group count is ⌊N ′ /T ⌋ where N ′ is the total key count. Only complete groups are used. For a watermarked group with observed key order (q1 , q2 , . . . , qT ), the coefficients are reconstructed by reversing (4): (0) b K = lex_sort {q1 , ..., qT } (j−1) b b a = index_of q , K j j (j) (j−1) b b \ {qj } K= K
wfinal [j] = mode w1 [j], w2 [j], . . . , wG [j] , , j = 1, 2, ..., T −1 (6)
Each b aj is the zero-based position of key qj in the current sorted remainder, which exactly matches the original coefficient aj when no keys have been deleted or reordered by an attacker. The watermark integer is then recovered from the factorial series: c= M
T −1 X
where bin(x, L) returns the L-bit binary representation of x, padding with leading zeros as needed. Because embedding and extraction share the same lexicographic baseline and factorial mapping, the recovered coefficients satisfy b aj = aj in the absence of attacks, guaranteeing zero-bit-error extraction. Finally, majority voting is applied across all G groups. For each bit position j:
b aj × (T − j)!,
(7)
j=1
and is converted to a binary sequence of length L: c = bin M c, L , W
(8)
(9)
where wg [j] is the j-th bit from group g. This mechanism ensures that the correct watermark is recoverable as long as a majority of groups remain intact. IV. Experiment This section introduces evaluation metrics (IV-A), experimental setup (IV-B), capacity experiment (IV-C), time overhead analysis (IV-D), robustness comparison with baseline methods (IV-E), and testing in real-world systems (IV-F). A. Experimental Evaluation Three metrics are adopted to evaluate the performance of the proposed method: time overhead, attack intensity, and watermark similarity.
Fig. 4. Key position determination inside a group during reordering.
1) Time Overhead: Time overhead refers to the execution time required to embed the watermark into semistructured data and extract it from the watermarked data. Specifically, it includes: Embedding Time: The total time required to generate the watermarked version from the original semistructured data. • Extraction Time: The total time required to recover the watermark information from the watermarked semi-structured data. •
Time overhead is measured in milliseconds (ms). A larger time overhead indicates a greater negative impact on normal business processes for the given dataset, i.e., lower operational efficiency of the method.
2) Attack Intensity: Attack intensity quantifies the degree of malicious modification imposed on semistructured data. Given a semi-structured data sample containing N key-value pairs, the attacker can perform one of the following three operations: • Deletion: Randomly remove M key-value pairs. • Tampering: Randomly modify the values of M keyvalue pairs. • Insertion: Randomly add M new key-value pairs. Attack intensity is defined as follows: Attack Intensity =
M × 100% N
(10)
where M denotes the number of key-value pairs that are deleted, tampered, or inserted, and N denotes the total
number of key-value pairs in the original semi-structured data. As defined in (1), a higher attack intensity indicates more severe damage to the semi-structured data, making correct watermark extraction more challenging. In this experiment, we set the attack intensity range from 0% to 50% to evaluate the robustness of the method under different levels of destruction. 3) Watermark Similarity: The watermark similarity is calculated as: 1X Watermark Similarity = I(wi = wi′ ) × 100% (11) L i=1
TABLE II Configuration of constructed datasets ID 1 2 3 4 5 6 7 8 9
Number of Keys 5–25 5–25 5–25 50–250 50–250 50–250 50 100 200
Max Nesting Depth 1 3 3 1 3 3 3 3 3
Nesting Probability 0 0.3 0.7 0 0.3 0.7 0.3 0.3 0.3
L
B. Experimental Setup 1) Testbed: The watermarking system was built upon the OpenResty1 Proxy framework for HTTP request/response interception. Each experiment was repeated ten times with random seeds, and the average results are reported. 2) Baseline Methods: We compared our method with three state-of-the-art semi-structured data watermarking schemes: He et al. [4], which embeds watermarks by modifying the values of key-value pairs; Phimark [18], a zero-distortion watermarking scheme for relational data; and Zhao et al. [10], a zero-watermarking scheme based on soliton distribution for semi-structured power data. C. Datasets Due to data sensitivity, we used desensitized data from the power industry to construct the API server-side datasets for our experiments. Table II summarizes the configuration of the nine datasets used in our evaluation. • Number of Keys: The count of key-value pairs in each JSON object. • Max Nesting Depth: The deepest nested level within the JSON structure. Depth 1 means no nesting (flat key-value pairs). • Nesting Probability: The chance that a given key contains a nested object instead of a primitive value. These settings cover a wide range of real-world semistructured data and help us test PEMark under different structural conditions. D. Capacity experiment The capacity of PEMark is determined by the number of keys in each group and the factorial decomposition method. This section analyzes the relationship between watermark length, capacity threshold T , and the number of groups g. 1 https://github.com/openresty/openresty
60
50
Threshold T
where I(·) is the indicator function that takes the value 1 when the condition holds and 0 otherwise. According to (2), a higher watermark similarity indicates less damage to the embedded watermark, implying a greater probability of successfully recovering the correct original watermark under error correction coding.
Threshold vs Length
40
30
20
10
0 0
50
100
150
200
250
Length L (bits)
Fig. 5. Watermark length L vs. capacity threshold T
1) Watermark Length vs. Capacity Threshold: A longer watermark requires more permutations to encode all possible values. The capacity threshold T is the minimum number of keys needed in a group to embed a watermark of length L bits. The condition is given by: 2L ≤ T !
(12)
Fig. 5 shows the relationship between L and T . For instance, a 64-bit watermark requires T = 21 keys, while 128-bit watermarks require T = 34 keys. 2) Robustness vs. Number of Groups: The number of groups g affects robustness. A larger g means the watermark is repeated in more groups. This gives more votes during extraction. As a result, the method can tolerate more group damage. 3) Handling Insufficient Keys: When a group has fewer than T keys, we cannot directly embed the full watermark. Our solution is to add fake key-value pairs to that group. E. Time Overhead Analysis We measured embedding and extraction time on six datasets with different sizes. Fig. 6 shows the results. As the number of keys increases, both embedding and extraction time increase. For datasets with 5–25 keys, the
Dataset 1
0.5
Dataset 2
0.5
0.3
0.2
0.4
0.3
0.2 10
15
20
25
10
Dataset 4
15
20
25
0.6
200
250
15
20
25
Dataset 6
Embedding Time Extraction Time
0.4
0.2 150
10
Number of JSON Keys
Time Overhead (ms)
Time Overhead (ms)
0.2 100
5
Dataset 5
0.4
50
0.3
Number of JSON Keys
Embedding Time Extraction Time
0.6
0.4
0.2 5
Number of JSON Keys
Time Overhead (ms)
Embedding Time Extraction Time
Time Overhead (ms)
0.4
5
Dataset 3
0.5
Embedding Time Extraction Time
Time Overhead (ms)
Time Overhead (ms)
Embedding Time Extraction Time
Embedding Time Extraction Time
0.6
0.4
0.2 50
Number of JSON Keys
100
150
200
250
50
100
Number of JSON Keys
150
200
250
Number of JSON Keys
Fig. 6. Time overhead comparison of embedding and extraction on six datasets.
time ranges from 0.22 ms to 0.47 ms. For datasets with 50–250 keys, the time ranges from 0.27 ms to 0.65 ms. PEMark keeps time overhead low. Even on datasets with 250 keys and high nesting complexity, the time stays between 0.3 ms and 0.6 ms. F. Robustness Comparison with Baseline Methods We compared PEMark with three methods: He’s method [4], Phimark [18], and Zhao’s method [10]. We tested three attacks: deletion, tampering, and insertion. We used three datasets: Dataset 7 (50 keys), Dataset 8 (100 keys), and Dataset 9 (200 keys). Fig. 7 shows the results. 1) Delete Attack: PEMark keeps similarity above 94% when attack intensity is below 15%. At 50% attack intensity, PEMark gives 49.23%-60.45% similarity. He’s method drops to 47.40%-61.98% with large fluctuations. Phimark stays at 75.10%-96.75%. Zhao’s method keeps above 91%. Delete attacks break the order of keys. So PEMark does not beat zero-watermark methods like Zhao’s method. But PEMark works better than He’s method with more stable results. 2) Tamper Attack: PEMark keeps 100% similarity under all attacks on all datasets. He’s method drops to 79.17%-82.81% at 50% attack. Phimark stays at 71.32%96.88%. Zhao’s method keeps above 91%. PEMark is perfect against tamper attacks. Our method does not change key values. It uses position encoding. So value changes do not affect the watermark. No other method achieves this.
3) Insert Attack: PEMark keeps 100% similarity under all attacks on all datasets. He’s method drops to 69.27%76.56% at 50% attack. Phimark stays at 72.42%-98.04%. Zhao’s method drops to 81.56%-84.53%. PEMark is also perfect against insert attacks. Two reasons explain this. First, insert attacks do not change existing keys that hold the watermark. Second, our voting mechanism rejects fake watermarks from new keys. Together, they make PEMark fully robust to insert attacks. G. Testing in Real-World System We tested PEMark on three real public APIs: DeepSeek2 , OpenAI3 , and GitHub4 . For each service, we measured the response time in two scenarios: (1) calling the API normally, and (2) applying PEMark to reorder the keys in the JSON response before forwarding it to the client. Fig. 8 illustrates the effect of PEMark on the JSON response from the GitHub API. As shown, the semantic content remains unchanged, while the order of keys is rearranged according to the embedded watermark. Table III shows the results. The overhead of PEMark is consistent and low. All APIs remain fully usable after processing. PEMark does not change the API server or the logical data. Clients get the same data and work as before. The added latency is about 20 ms per request, which is acceptable for most real-world applications. 2 https://api.deepseek.com/models 3 https://api.openai.com/v1/models 4 https://api.github.com/users/google
Dataset 7 - Delete Attack
Dataset 8 - Delete Attack
80 70 PEMark He's method Phimark Zhao's method
10
80 70 PEMark He's method Phimark Zhao's method
60 50
20
30
40
50
0
10
Attack Intensity (%)
Watermark Similarity (%)
90
80
70
10
Phimark Zhao's method
20
30
40
20
30
40
50
0
10
80
70
0
10
Phimark Zhao's method
20
30
30
40
PEMark He's method
50
0
10
20
30
40
50
Attack Intensity (%)
100
90
80
70
0
10
Phimark Zhao's method
20
30
50
Phimark Zhao's method
20
30
Dataset 9 - Insert Attack
PEMark He's method
40
70
Dataset 8 - Insert Attack
Phimark Zhao's method
50
80
Dataset 7 - Insert Attack
PEMark He's method
40
90
Attack Intensity (%)
70
50
100
Attack Intensity (%)
80
40
Dataset 9 - Tamper Attack
90
PEMark He's method
20
Attack Intensity (%)
100
50
90
10
PEMark He's method Phimark Zhao's method
60
Attack Intensity (%)
100
0
70
Dataset 8 - Tamper Attack
Watermark Similarity (%)
Watermark Similarity (%)
Dataset 7 - Tamper Attack
0
80
Attack Intensity (%)
100
PEMark He's method
90
50
Watermark Similarity (%)
0
90
Watermark Similarity (%)
60
100
Watermark Similarity (%)
90
50
Watermark Similarity (%)
Dataset 9 - Delete Attack
100
Watermark Similarity (%)
Watermark Similarity (%)
100
40
100
90
80
70
50
Attack Intensity (%)
PEMark He's method
0
10
Phimark Zhao's method
20
30
Attack Intensity (%)
Fig. 7. Robustness comparison of PEMark with baseline methods under different attack types.
TABLE III API response time before and after PEMark processing. API Service DeepSeek OpenAI GitHub
Fig. 8. Comparison of GitHub API JSON response before (left) and after PEMark processing (right).
V. Limitation Our method has several limitations.
Original (ms) 543.24 933.92 702.74
With PEMark (ms) 563.19 952.47 723.32
Overhead (ms) 19.95 18.55 20.58
First, it does not work well under strong delete attacks. At 50% attack strength, the similarity drops to 49%-60%. Delete attacks change the order of keys. This order is the base of our method. The topology of the watermark source significantly affects robustness [37]. Future work can add error correction codes to fix this. Second, our method uses a fixed key order. The keys are sorted from a to z. If an attacker knows this, they can change specific keys to break the watermark. In untrusted environments, watermark resilience to brute force attacks must be carefully considered [38]. One way to fix this is to use a secret key. The secret key randomizes the order. Only the sender and receiver know it.
Third, our method does not use deep nesting. It only looks at the top level of the JSON data. Nested objects have more space for watermarks. Future work can use this space to store more data or improve robustness. We will address these limitations in future work. VI. Conclusions We introduced PEMark, a novel position encodingbased watermarking scheme for API responses. The originality of this work is threefold: (1) we are the first to exploit key-ordering redundancy as a watermarking channel in semi-structured API responses; (2) we design a pluggable proxy-gateway architecture that requires zero modification to business code; and (3) we achieve distortion-free watermarking without any thirdparty feature library. Compared to existing methods, our scheme needs no third-party library and no business code changes. PEMark achieves perfect robustness against tamper and insert attacks (100.00% similarity under all attacks). It also keeps 71%-83% similarity under 30% delete attacks. Time overhead stays below 0.65 ms even for datasets with 250 keys. Future work will explore adaptive encoding for strong delete attacks and keydependent ordering to stop targeted attacks. Acknowledgment This work is supported by the Research on Key Technologies for Marking-Driven Detection and Traceability of Sensitive Electric Power Data Leakage. (No.52150025001H-146-ZN). References [1] R. Halder, S. Pal, and A. Cortesi, “Watermarking techniques for relational databases: Survey, classification and comparison,” Journal of Universal Computer Science, vol. 16, no. 21, pp. 3164–3190, 2010. [2] R. Agrawal and J. Kiernan, “Watermarking relational databases,” in VLDB’02: Proceedings of the 28th International Conference on Very Large Databases. Elsevier, 2002, pp. 155– 166. [3] M. Shehab, E. Bertino, and A. Ghafoor, “Watermarking relational databases using optimization-based techniques,” IEEE transactions on Knowledge and Data Engineering, vol. 20, no. 1, pp. 116–129, 2008. [4] D. Hu, D. Zhao, and S. Zheng, “A new robust approach for reversible database watermarking with distortion control,” IEEE Transactions on Knowledge and Data Engineering, vol. 31, no. 6, pp. 1024–1037, 2018. [5] Z. Ren, H. Fang, J. Zhang, Z. Ma, R. Lin, W. Zhang, and N. Yu, “A robust database watermarking scheme that preserves statistical characteristics,” IEEE Transactions on Knowledge and Data Engineering, vol. 36, no. 6, pp. 2329–2342, 2023. [6] S. Iftikhar, M. Kamran, and Z. Anwar, “Rrw—a robust and reversible watermarking technique for relational data,” IEEE Transactions on Knowledge and Data Engineering, vol. 27, no. 4, pp. 1132–1145, 2015. [7] D. Florescu and D. Kossmann, “A performance evaluation of alternative mapping schemes for storing xml data in a relational database,” Ph.D. dissertation, INRIA, 1999. [8] J. He, Q. Ying, Z. Qian, G. Feng, and X. Zhang, “Semistructured data protection scheme based on robust watermarking,” EURASIP Journal on Image and Video Processing, vol. 2020, no. 1, p. 12, 2020.
[9] I. Constantin, C. Dobre, R. Ciobanu, O. Dochia, and A. Barbu, “Watermark decoding solution for semi-structured data,” in ICERI2024 Proceedings. IATED, 2024, pp. 6927–6935. [10] L. Zhao, Y. Zou, C. Xu, Y. Ma, W. Shen, Q. Shan, S. Jiang, Y. Yu, Y. Cai, Y. Song et al., “Robust soliton distributionbased zero-watermarking for semi-structured power data,” Electronics, vol. 13, no. 3, p. 655, 2024. [11] S. Kumar, B. K. Singh, and M. Yadav, “A recent survey on multimedia and database watermarking,” Multimedia Tools and Applications, vol. 79, no. 27, pp. 20 149–20 197, 2020. [12] M. L. P. Gort, M. Olliaro, A. Cortesi, and C. F. Uribe, “Semantic-driven watermarking of relational textual databases,” Expert Systems with Applications, vol. 167, p. 114013, 2021. [13] C.-C. Chen, Y. He, X. Peng et al., “A reversible database watermark scheme for textual and numerical datasets,” IEEE Transactions on Knowledge and Data Engineering, vol. 34, no. 8, pp. 1–14, 2022, dOI: 10.1109/TKDE.2022.3144763. [14] M. L. P. Gort and A. Cortesi, “A robust scheme for securing relational data incremental watermarking,” International Journal of Information Management Data Insights, vol. 5, no. 1, p. 100320, 2025. [15] S. Batra and S. Tyagi, “Comparative analysis of relational and graph databases for social networks,” in 2018 5th International Conference on Signal Processing and Integrated Networks (SPIN). IEEE, 2018, pp. 210–215. [16] A. C. Bart, E. Tilevich, S. Hall, T. Allevato, and C. A. Shaffer, “Transforming introductory computer science projects via realtime web data,” in Proceedings of the 45th ACM technical symposium on Computer science education, 2014, pp. 289–294. [17] J. Fritsch and A. Bales, “Market guide for data masking and synthetic data,” Gartner, Inc., Market Guide G00787177, 8 2024. [18] J. Ji, Y. Peng, W. Ma, H. Li, J. Cui, and X. Gao, “Phimark: watermarking relational data robustly with zero distortion,” Information Processing & Management, vol. 63, no. 7, p. 104782, 2026. [19] X. Tang, Z. Cao, X. Dong, and J. Shen, “Pkmark: A robust zero-distortion blind reversible scheme for watermarking relational databases,” in 2021 IEEE 15th International Conference on Big Data Science and Engineering (BigDataSE). IEEE, 2021, pp. 72–79. [20] J. Han, X. Peng, H. Xian, and D. Yang, “A distortion free watermark scheme for relational databases,” in 2024 IEEE 33rd International Conference on Computer Communications and Networks (ICCCN). IEEE, 2024, pp. 1–6. [21] N. Ren, S. Guo, C. Zhu, and Y. Hu, “A zero-watermarking scheme based on spatial topological relations for vector dataset,” Expert Systems with Applications, vol. 226, p. 120217, 2023. [22] J. Xu, Z. Guo, Y. Tang, and B. Han, “A zero-watermarking scheme for medical images based on a stacked sparse autoencoder network,” Expert Systems with Applications, vol. 303, p. 130651, 2025. [23] S. Bhattacharya and A. Cortesi, “A distortion free watermark scheme for relational databases,” in 2024 33rd International Conference on Computer Communications and Networks (ICCCN), 2024, pp. 1–6. [24] S. Jiao, Y. Qiu, Q. Su, C. Shi, and Z. Liu, “Enhancing watermarking robustness and invisibility with growth optimizer and improved lu decomposition,” Optik, vol. 329, p. 172353, 2025. [25] W. Li, N. Li, J. Yan, Z. Zhang, P. Yu, and G. Long, “Secure and high-quality watermarking algorithms for relational database based on semantic,” IEEE Transactions on Knowledge and Data Engineering, vol. 35, no. 7, pp. 7440–7456, 2022. [26] Q. Li, X. Wang, Q. Pei, X. Chen, and K.-Y. Lam, “Consistency preserving database watermarking algorithm for decision trees,” Digital Communications and Networks, vol. 10, no. 6, pp. 1851–1863, 2024. [27] Q. Li, X. Wang, Q. Pei, K.-Y. Lam, N. Zhang, M. Dong, and V. C. Leung, “Database watermarking algorithm based on decision tree shift correction,” IEEE Internet of Things Journal, vol. 9, no. 23, pp. 24 373–24 387, 2022.
[28] K. Jawad and A. Khan, “Genetic algorithm and difference expansion based reversible watermarking for relational databases,” Journal of Systems and Software, vol. 86, no. 10, pp. 2742–2753, 2013. [29] A. S. Alghamdi, S. Naz, A. Saeed, E. Al Solami, M. Kamran, and M. S. Alkatheiri, “A novel database watermarking technique using blockchain as trusted third party,” Computers, Materials & Continua, vol. 70, no. 1, pp. 1585–1601, 2021. [30] S. Rani and R. Halder, “Comparative analysis of relational database watermarking techniques: An empirical study,” IEEE Access, vol. 10, pp. 27 970–27 989, 2022. [31] C. Cai, C. Peng, J. Niu, W. Tan, and H. Tang, “Low distortion reversible database watermarking based on hybrid intelligent algorithm,” Mathematical Biosciences and Engineering, vol. 20, no. 12, pp. 21 315–21 336, 2023. [32] W. Qi, C. Li, and X. Han, “Research on blind reversible database watermarking algorithm based on dual embedding strategy,” The Computer Journal, p. bxae080, 2024. [33] D. İşler, E. Cabana, A. García-Recuero, G. Koutrika, and N. Laoutaris, “Freqywm: Frequency watermarking for the new data economy,” in 2024 IEEE 40th International Conference on Data Engineering (ICDE). IEEE, 2024, pp. 4993–5007. [34] S. Rani and R. Halder, “An efficient format-independent watermarking framework for large-scale data sets,” Expert Systems with Applications, vol. 210, p. 118085, 2022. [35] N. Fränkel, “Dynamic watermarking with imgproxy and apache apisix,” DZone, 2024. [36] D. H. Lehmer, “Teaching combinatorial tricks to a computer,” Proceedings of Symposia in Applied Mathematics, vol. 10, pp. 179–193, 1960. [37] M. L. Pérez Gort, M. Olliaro, and A. Cortesi, “Study of the watermark source’s topology role on relational data watermarking robustness,” IEEE Access, vol. 12, pp. 25 857–25 875, 2024. [38] ——, “Relational data watermarking resilience to brute force attacks in untrusted environments,” Expert Systems with Applications, vol. 212, p. 118713, 2023.