ConceptioArchivearXiv CS
arXiv CSopen access

AI Watermark Evidence Fails Forensic Readiness: An Empirical Evaluation

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
cryptography, security, privacy, cybersecurity

AI Watermark Evidence Fails Forensic Readiness: An Empirical Evaluation Saifur Rahman Tamim, Amir Labib Khan Department of Computer Science and Engineering Northern University Bangladesh tamim [email protected], [email protected]

Preprint. A version of this paper was submitted to the AAAI/ACM Conference on AI, Ethics, and Society (AIES) 2026.

arXiv:2607.16010v1 [cs.CR] 17 Jul 2026

Abstract Governments are increasingly mandating that LLMgenerated content carry watermarks. The EU AI Act calls for markings that are “sufficiently reliable and robust.” California’s SB 942 requires disclosure that is “permanent or extraordinarily difficult to remove.” Both mandates rest on an untested assumption: that watermark detection yields evidence reliable enough for courts. This paper tests that assumption directly. We evaluate three representative LLM watermarking methods—KGW, Unigram, and the MarkLLM implementation of SynthID-Text—against the Daubert admissibility criteria and the NIST SP 800-86 digital forensic process. To structure this evaluation, we propose a Forensic Readiness Score (FRS) framework with 12 criteria, three mandatory gates, and a 60-point scoring system. We focus on meaning-preserving paraphrase as the attack vector, since it is both legally realistic and difficult to dismiss as evidence tampering. The results raise serious evidentiary concerns. Out of 846 valid paraphrase runs across 15 diverse prompts per method, every single initially-detected KGW and Unigram text lost its watermark after paraphrasing—100% conditional removal. SynthID fared only slightly better at 98.3%. Even before any attack, false-negative rates were already high: 70% for KGW, 83% for Unigram, 80% for SynthID. The SynthID configuration also flagged 5.4% of paraphrased human-written controls as AI-generated and showed an 18.6% paradox rate, with 80% of its own pristine watermarked output landing in the uncertainty deadband. None of the three methods satisfy more than two of five Daubert factors. We also find that the FRS point-based scoring system, despite working as designed, cannot fully capture forensic uselessness—a limitation worth noting for future framework design. These configurations, as tested, do not meet the evidentiary bar that courts require.

Introduction Consider a courtroom scenario. A prosecutor offers a text as AI-generated, citing a positive watermark detection result. The defense attorney takes that same text, runs it through a paraphrasing tool anyone can access online, and hands back a version that means the same thing—but no longer

triggers the watermark detector. The prosecution’s evidence just evaporated. The meaning survived. What is the judge supposed to do with that? This is not a contrived hypothetical. The EU AI Act requires AI-generated content markings to be “sufficiently reliable, interoperable, effective and robust as far as this is technically feasible” (European Parliament and Council of the European Union 2024). California’s SB 942 insists that disclosure be “permanent or extraordinarily difficult to remove” (California State Legislature 2024). Executive Order 14110 went further, mandating “state-of-the-art” provenance tools and explicitly naming watermarking (The White House 2023)—though the order was later rescinded (The White House 2025), which itself says something about how quickly the policy ground can shift. What these mandates have in common is an assumption: that the underlying watermarking technology actually works well enough to produce courtroom-grade evidence. But this assumption has not been directly evaluated against forensic admissibility standards. Robustness studies measure how well detectors hold up against attacks using ML metrics like TPR and AUC. Governance analyses ask whether the policy infrastructure around watermarking is adequate. Neither tradition provides the thing courts actually need: an end-toend assessment of whether watermark evidence meets the Daubert admissibility standard and follows the NIST SP 800-86 forensic process (Liang et al. 2025; Nemecek, Jiang, and Ayday 2025). Some background is useful here. The Daubert criteria (Daubert) are how U.S. federal courts decide whether scientific evidence is admissible—they ask about testability, peer review, known error rates, standards, and general acceptance. NIST SP 800-86 (National Institute of Standards and Technology 2006) lays out how digital forensic evidence should be collected, examined, analyzed, and reported. Recent work has warned that watermarking risks becoming “symbolic compliance” without real enforceable standards (Nemecek, Jiang, and Ayday 2025). We wanted to see if the empirical data backs up that concern. It does. This paper contributes four things: 1. An empirical evaluation of representative LLM watermarks against forensic admissibility standards (both NIST SP 800-86 and Daubert). 2. A Forensic Readiness Score (FRS) framework: 12

scorable criteria grounded in Daubert factors and NIST phases, plus 3 mandatory gates and a 60-point scale. 3. Paraphrase experiments across KGW, Unigram, and SynthID-Text, with reproducibility verified through timestamped same-seed computational reruns. 4. Evidence that meaning-preserving paraphrase achieves 100% conditional watermark removal for two of three methods (98.3% for the third), while keeping semantic similarity high enough that a court would struggle to call it evidence tampering. As watermarked AI systems reach wider deployment, courts will encounter this kind of evidence without any established framework for evaluating it. We provide that framework—and show that, at least for these configurations, the evidence does not hold up.

Background and Related Work LLM Watermarking Methods We test three watermarking methods that represent different design philosophies, all available through the open-source MarkLLM toolkit (Pan et al. 2024). KGW (Kirchenbauer et al. 2023) works by splitting the token vocabulary into “green” and “red” lists using a hash of the previous token. During text generation, green-list tokens get a logit boost controlled by a parameter δ, pushing the model toward those tokens. Detection then checks whether the output has an unusually high proportion of green tokens via a z-score. One important consequence of this design: because the partition depends on each token’s predecessor, changing even one token can ripple through the rest of the partitions. Unigram (Zhao et al. 2024) takes a different approach—it uses a fixed, context-independent vocabulary partition. The green/red split stays the same regardless of what came before, which in theory makes it more resilient to local edits. Detection still uses a z-score over green-token frequency, but the fixed partition means the method should be more robust to bounded edit-distance attacks. SynthID-Text (Dathathri et al. 2024) is the highestprofile method in our evaluation—published in Nature, deployed by Google in production. It uses tournamentbased scoring with probabilistic detection. We evaluate the open-source MarkLLM implementation, not Google’s proprietary system (an important distinction we return to in the limitations). Unlike the other two methods, SynthID uses three detection states—WATERMARKED, NOT WATERMARKED, and UNCERTAIN—rather than a simple binary threshold. In our experiments we use the MarkLLM weighted-mean detector with a threshold of 0.5 and an UNCERTAIN deadband of ±0.03. These three methods together cover the main design space: context-dependent partitioning (KGW), context-free partitioning (Unigram), and tournament-based probabilistic scoring (SynthID). That range matters because we want to test whether forensic failure is a property of the general approach, not just one particular design.

Forensic Evidence Standards Two frameworks are relevant to how courts handle scientific and digital evidence. Daubert criteria. Under Daubert v. Merrell Dow Pharmaceuticals (Daubert) and Federal Rule of Evidence 702 (Federal Rules), judges evaluate scientific evidence on five factors: (1) testability and falsifiability, (2) peer review and publication, (3) known or potential error rate, (4) existence of controlling standards, and (5) general acceptance. For watermarking, Factor 3—the known error rate—turns out to be the critical weakness. If the error rate is unstable or unknown, the evidence is hard to admit. We note that Daubert is a US federal evidentiary standard; state courts following Frye and civil-law jurisdictions (e.g., under EU procedural rules) use different admissibility tests. We use Daubert as our primary lens because of its close conceptual fit with NIST’s error-rate language, but the forensic weaknesses we document—unstable error rates, no controlling standards— are relevant to any admissibility framework that asks similar questions of scientific evidence. NIST SP 800-86. NIST’s guide for digital forensics (National Institute of Standards and Technology 2006) defines four phases: Collection, Examination, Analysis, and Reporting. None of the watermark evaluations we are aware of document an end-to-end workflow aligned to these phases. This matters because watermark detection alone is not a forensic process; it becomes forensic evidence only when embedded in a documented collection, examination, analysis, and reporting workflow. NIST has separately acknowledged the broader challenge of digital content transparency (National Institute of Standards and Technology 2024), but no existing guidance bridges watermark detection to forensic evidence handling. Historical precedent. There is a pattern here worth noting. The 2009 NAS report (National Research Council 2009) found that several forensic disciplines—fingerprint analysis, bite mark comparison, hair microscopy—had been used in courtrooms for decades without adequate scientific validation. The 2016 PCAST report (President’s Council of Advisors on Science and Technology 2016) made similar points. The consequences were real: wrongful convictions. AI watermarking looks like it could be heading down the same path—deployed in practice, written into regulation, but not actually validated against the standards courts use to decide whether evidence is reliable enough to hear.

Related Work Watermark robustness evaluation. The most thorough robustness study to date is WaterPark by Liang et al. (2025), which tested 10 watermarking methods against 12 attack types across three language models (OPT-1.3B, LLaMA37B, Qwen2.5-14B) and five datasets. Their findings are directly relevant to ours: SynthID’s true positive rate dropped from 0.998 on clean text to 0.498 under moderate paraphrasing. A single ChatGPT paraphrase pass brought every tested method below 30% detection. They also found enormous cross-model variability—KGW’s TPR went from 0.858 on OPT to 0.334 on LLaMA3 under the same attack conditions.

Table 1: How this paper fits alongside prior watermark evaluations. We go narrower on attacks but deeper on forensic admissibility—to our knowledge, the first study to jointly apply Daubert and NIST SP 800-86 forensic standards to LLM watermark evaluation, building on prior robustness (Liang et al.) and governance (Nemecek et al.) work. Work

Scale

Attacks

Forensic framework

Liang et al. Nemecek et al. Pan et al.; Yi et al. This work

10 methods 1 method 3 methods 3 methods

12 types 1 type Distillation 1 type

ML metrics (TPR, FPR) Governance scorecard Not addressed FRS + Daubert + NIST

But WaterPark evaluates entirely through ML metrics. It does not ask whether these numbers would satisfy a Daubert hearing, and it does not map results to NIST forensic phases. Other work has established theoretical impossibility results for strong watermarking (Zhang et al. 2024), shown that reliable AI text detection may be impossible in adversarial settings (Sadasivan et al. 2023), and demonstrated attacks that exploit watermark design properties (Pang et al. 2024). These all point toward fragility, but none address admissibility. Notably, the WaterPark finding that one round of ChatGPT paraphrasing drops all methods below 30% TPR suggests our same-model paraphraser actually represents a conservative lower bound on how bad things can get. Distillation-based attacks and attribution ambiguity. Pan et al. (2025) demonstrated that watermark traces can be stripped away during unauthorized knowledge distillation. Yi et al. (2025) extended this to unified spoofing and scrubbing attacks across KGW, Unigram, and SynthID-Text. The forensic implication goes beyond robustness: even when a watermark is detected, the signal might not mean what you think it means. It could reflect direct generation, inherited traces from distillation, or deliberate spoofing. That kind of attribution ambiguity makes it very difficult to establish a stable error-rate interpretation in court. Governance and policy. Nemecek, Jiang, and Ayday (2025) argued that watermarking without enforceable standards amounts to “symbolic compliance” and laid out a three-layer governance framework covering technical requirements, audit infrastructure, and enforcement. Their paper included a prototype evaluation scorecard with 0–5 scoring—a design that ended up converging with our FRS. They also ran SynthID on Gemma-2-9b-it independently and found paraphrasing eliminated detection in 4 of 5 cases. But their contribution is fundamentally a governance argument, not a systematic empirical test: five prompts, no forensic pipeline, no FRS-style scoring. Our position. There are three threads in the existing literature. Liang et al. (2025) ask: how robust are watermarks? Nemecek, Jiang, and Ayday (2025) ask: why is governance failing? Distillation-attack studies ask whether watermark traces survive model transfer (Pan et al. 2025; Yi et al. 2025). We ask a different question: would watermark evidence actually survive a Daubert hearing? For the three methods and configurations we tested, the answer is no.

Table 2: FRS criteria mapped to forensic standards. Cat.

#

Criterion

Standard

Tech.

T1 T2 T3 T4

Repeatability Robustness Detectability Quantifiability

NIST reprod. NIST exam. Daubert F3 Daubert F1

Legal

L1 L2 L3 L4

Reliability Peer Review Known Error Rate Testability

Daubert F3 Daubert F2 Daubert F3 Daubert F1

Oper.

O1 O2 O3 O4

Accessibility Efficiency Documentation Independence

NIST cross. NIST cross. NIST cross. NIST cross.

The Forensic Readiness Score Framework ML benchmarks ask: “does watermark detection work on average?” Courts need to know something different: “can an expert witness testify, for this specific piece of evidence, to a known and stable error rate?” The FRS framework is our attempt to bridge that gap.

Design Rationale The FRS maps NIST SP 800-86’s four phases onto five framework components (Evidence Collection, Verification Protocol, FRS Evaluation, Evidence Report Template, and Implementation Guidelines) and turns Daubert’s five factors into scorable criteria. We were not the first to think along these lines—Nemecek, Jiang, and Ayday (2025) independently proposed a scorecard with 0–5 scoring across robustness, detection quality, and auditability. Our FRS shares that philosophy but takes it further: each criterion is grounded in a specific Daubert factor or NIST phase, we add mandatory gates that can override point scores entirely, and we validate the whole thing empirically.

The 12 Criteria Each criterion gets a score from 0 to 5, for a maximum of 60 points. They fall into three categories (Table 2).

The Three Mandatory Gates Points alone are not enough. Even if a method scores ≥40/60, it still has to clear three gates. Failing any one of them results in a NOT FORENSIC READY verdict, no matter how high the score. Gate G1: FPR and FNR must be documented and independently computable. Without known error rates, courts cannot evaluate reliability (Daubert Factor 3). Gate G2: The paradox rate must stay below 20%. If attacking a watermark increases the detection score more than a fifth of the time, something is fundamentally wrong with the method’s logic. Evidence corruption should not help verification. Gate G3: Results must be repeatable across sessions within documented tolerance. This is a basic NIST reproducibility requirement.

Table 3: Experimental configuration. A single attack model (Qwen2.5-1.5B) is used across all methods, controlling the attacker variable so that cross-method differences reflect watermark properties rather than paraphraser capability. Method

Source Model

Attack Model

Venue

Detector

KGW Unigram SynthID

Qwen2.5-1.5B Qwen2.5-1.5B Gemma-2-9b-it

Qwen2.5-1.5B Qwen2.5-1.5B Qwen2.5-1.5B

ICML ’23 ICLR ’24 Nature ’24

z-score (τ =4.0) z-score (τ =4.0) weighted mean (τ =0.5±0.03)

Table 4: Session provenance. Replicate sessions used identical seeds and produced byte-identical results, confirming deterministic reproducibility. We report canonical session results; replicate sessions are archived for provenance. Session

Notebook

Seed

Status

qwen v4.1 rep2 qwen v4.1 rep3 synthid v4 synthid v4 rep2

KGW+Unigram KGW+Unigram SynthID SynthID

123 123 123 123

Canonical Reproducibility Canonical Reproducibility

Scoring Limits as a Finding One of the more interesting outcomes of this work is what happens with the scoring system itself. As Section 5 will show, KGW ends up at 37/60—NOT FORENSIC READY, and fairly classified. But Unigram lands at exactly 40/60, right on the CONDITIONALLY FORENSIC READY threshold, despite having the worst FNR of any method (83%) and 100% conditional removal. It passes every gate and clears the point threshold while being, in practice, forensically useless. We think this is worth highlighting not as a flaw in the framework but as a genuine finding: point-based scoring has inherent limits when it comes to capturing forensic uselessness. It is a bit like how a good credit score can mask individual risk. Courts will need both the quantitative score and a qualitative assessment of how the method actually behaves under adversarial conditions.

Experimental Setup Methods and Models We run all experiments through the MarkLLM toolkit (Pan et al. 2024). Table 3 summarizes the configuration. For KGW and Unigram, Qwen2.5-1.5B serves as both the watermark generator and the paraphrase attacker, which keeps the comparison clean across methods. For SynthID, we use Gemma-2-9b-it as the generator—following the setup in Nemecek, Jiang, and Ayday (2025)—and Qwen2.51.5B as a cross-model attacker. This is actually a more realistic threat model: in practice, someone trying to strip a watermark does not need access to the original model. The fact that we use different source models for different methods is deliberate. We are testing whether forensic failure is a property of statistical watermarking as a category, not something specific to one model.

Attack Design Meaning-preserving paraphrase. Qwen2.5-1.5B rewrites each watermarked text for all three methods. For SynthID, this creates a cross-model attack (Gemma-generated text attacked by Qwen); for KGW and Unigram, it is a same-model attack. We filter outputs through a triple validity gate: (1) cosine similarity ≥ 0.75 via all-MiniLM-L6-v2 (Reimers and Gurevych 2019), (2) normalized Levenshtein distance ≥ 0.15, and (3) length ratio between 0.5 and 2.0. Anything failing any gate gets thrown out. The retained paraphrases ended up with high semantic similarity—medians of 0.83–0.84 across methods, with every single retained sample above 0.75. We chose paraphras-

ing as the sole attack because it is the one that matters most in a legal setting. A defense attorney can point to a paraphrased text and say: “the meaning is identical—you cannot call this evidence destruction.” A court would have a very hard time excluding that argument.

Experimental Protocol We designed 15 prompts per method, covering conversational, technical, news, professional, and creative domains. Each prompt gets 2 generation seeds, giving 30 base watermarked texts per method. We then attack each base text at three temperatures (0.7, 1.0, 1.3) with five template variants per temperature, which means up to 450 attack attempts per method before filtering. After the validity gate, we end up with 304 valid runs for KGW, 306 for Unigram, and 236 for SynthID. We use prompt-level statistics as the primary unit of analysis and compute SHA-256 hashes at every stage. For baseline error rates, the 30 pristine texts per method give us FNR estimates. Paraphrased controls—237 for KGW and Unigram, 184 for SynthID—provide FPR baselines. The Qwen experiments ran on Colab free tier (T4 GPU); SynthID needed Colab Pro for the A100 to handle Gemma-2-9b-it. Every result we report comes from archived, versioncontrolled runs verified with SHA-256 hashes (Table 4). We ran each experiment twice under identical seeds and configuration; the reruns produced byte-identical outputs and serve as deterministic reproducibility evidence, not additional independent samples. Code and archived artifacts will be released upon acceptance.

Results The results tell a story in three layers: watermarks fail before any attack, they collapse under paraphrase, and the paraphrases preserve enough meaning that a court cannot dismiss them as evidence destruction.

Pre-Attack Baseline Failure Before anyone even tries to attack the watermarks, they are already unreliable. The fraction of watermarked texts that the detectors correctly identify—with no adversarial intervention at all—is surprisingly low. KGW picks up 9 out of 30 pristine texts (FNR = 70%, Wilson 95% CI: 52.1–83.3%). Unigram does worse: 5 out of 30 (FNR = 83%, CI: 66.4– 92.7%). SynthID manages 6 out of 30 (FNR = 80%, CI: 62.7–90.5%).

Step 4 Step 1

Watermarked Text Evidence

Step 3

Step 2

Verification Protocol

Score 12 Criteria

Compute Metrics

T1–T4, L1– L4, O1–O4 (max 60 pts)

FPR, FNR, Paradox, Repeatability

3 Mandatory Gates G1: Error rates documented? G2: Paradox <20%? G3: Cross-session repeatable?

FORENSIC READY

COND. READY

Score ≥40 & all gates pass

Score ≥40, gates pass, caveats noted

NOT FORENSIC READY Any gate fails or score <40

Figure 1: How the FRS evaluation works. Watermarked text moves through four stages: verification testing, metric computation, scoring against 12 criteria (0–5 each, 60 max), and three mandatory gates. Any single gate failure overrides the point score—a method that fails a gate is NOT FORENSIC READY no matter how many points it earned. In our results, Unigram passes all gates and lands at exactly 40/60 (the minimum COND. READY threshold) despite 100% conditional removal and 83% FNR—a clear scoring-limit case. Table 5: Paraphrase attack results across three methods. Conditional removal is the headline finding: among initially detectable texts, what fraction lost detection after paraphrase. Wilson 95% CIs reported for FPR, FNR, and conditional removal. Metric Valid runs Yield Raw removal Cond. removal FPR (controls) FPR 95% CI FNR (pristine) FNR 95% CI Paradox rate Sem. sim. (median)

KGW

Unigram

SynthID

304 34% 32.9% 100% (100/100) 0/237 (0%) [0, 1.6]% 21/30 (70%) [52.1, 83.3]% 2.3% 0.84

306 34.5% 13.4% 100% (41/41) 0/237 (0%) [0, 1.6]% 25/30 (83%) [66.4, 92.7]% 4.9% 0.83

236 23% 24.6% 98.3% (58/59) 10/184 (5.4%) [3.0, 9.7]% 24/30 (80%) [62.7, 90.5]% 18.6% 0.83

For SynthID the picture is particularly bleak: 24 of those 30 pristine texts scored inside the ±0.03 uncertainty deadband. The detector fails to issue a confident verdict on 80% of its own watermarked output. This matters for Daubert. If a detector cannot reliably find its own watermarks in clean, unmodified text, an expert witness cannot credibly testify to a “known error rate.” The evidence is already compromised before any adversary gets involved.

Paraphrase Attack: Primary Analysis Table 5 has the core numbers. Among texts where the watermark was initially detected, paraphrasing wiped it out in every single case for KGW and Unigram. For SynthID, 58 out of 59 initially-detected texts lost their watermark—98.3%. That 100% conditional removal for KGW and Unigram is the central result of this paper (Figure 4). What it means concretely: for any watermarked text that a forensic examiner could actually detect, a straightforward paraphrase will erase the watermark entirely. The paraphrasing tools needed are freely available. The meaning is preserved.

We should note that we only discarded paraphrases when they failed the validity gate (similarity below 0.75 or too little lexical change), so Table 5 represents the conservative, high-quality subset that would be hardest for a prosecutor to challenge. The compounding problem. These failures stack. KGW detects 30% of its own watermarked texts to begin with, and then loses all of those to paraphrase: 0.30 × 0.00 = 0.00 effective detection. Unigram is even worse—17% baseline detection, then 0% survival. SynthID detects 20% and retains 1.7% of those: 0.20 × 0.017 ≈ 0.003. A court relying on watermark evidence is looking at effective detection rates somewhere between 0% and 0.3% after a trivial, meaningpreserving transformation. SynthID-specific findings. SynthID has five forensically noteworthy behaviors. First, phantom watermarks: the detector flagged 10 out of 184 paraphrased control texts—clean, unwatermarked human text—as watermarked (5.4%, Wilson CI: 3.0–9.7%). That means human-authored text can be falsely attributed to an AI system. One of the flagged samples was a standard typing test sentence, which scored 0.5656 and cleared the detection threshold. If a standard typing test sentence can exceed the detection threshold, the false-attribution risk is non-trivial. Second, 24 out of 30 pristine watermarked texts (80%) scored inside the ±0.03 uncertainty deadband. The detector’s usable operating range is about 0.07 points wide (0.487 to 0.558). There is almost no room between “watermarked” and “not watermarked.” Third, threshold sensitivity: if you shift the threshold by just ±10%, 93.6% of verdicts flip. That earns a T4 (quantifiability) score of 0/5 in our FRS. There is no stable quantitative signal to testify about. Fourth, after paraphrasing, 69.5% of attacked texts stay stuck in the UNCERTAIN zone while 23.7% transition from DETECTED to UNCERTAIN. The three-state system, which was supposed to reduce false certainty, instead creates a forensic dead zone.

Table 6: FRS audit results under paraphrase attack. The “Empirical reality” column contextualizes the FRS score— methods that pass the point threshold can still be forensically useless. Method

FRS

Gates

Verdict

KGW Unigram SynthID

37/60 40/60 35/60

All pass All pass All pass

Not ready Cond. ready Not ready

Empirical reality 100% cond. removal, 70% FNR 100% cond. removal, 83% FNR 5.4% FPR, 80% UNCERTAIN, T4=0

Fifth, despite being published in Nature and deployed by Google, the MarkLLM SynthID configuration scores lowest of all three methods at 35/60. Liang et al. (2025) independently found SynthID’s TPR dropping to 0.498 under DP-40 paraphrase and 0.232 under translation, confirming that the method works on margins too thin for forensic use.

Semantic Preservation The retained paraphrases stay semantically close to the originals: median cosine similarities of 0.84 (KGW), 0.83 (Unigram), and 0.83 (SynthID), with everything above 0.75. This is what makes the attack so damaging from a legal perspective. You cannot call something “evidence destruction” when the meaning is substantially intact. The transformation is just rewording—the kind of thing people do naturally every day. We also ran preliminary word deletion experiments on earlier pilot datasets and saw the same general failure patterns, but we chose to focus the paper entirely on paraphrase since it is the legally realistic scenario. Our results line up with Liang et al. (2025): their WaterPark evaluation tested 12 attack types across three LLMs and found fragility everywhere. Their SynthID TPR dropped to 0.498 under moderate paraphrasing, 0.232 under translation. And a single ChatGPT paraphrase pass brought every method below 30% TPR—meaning our same-model attacker is actually a conservative test.

FRS Audit and Daubert Scorecard Table 6 gives the FRS scores as an interpretive lens over the empirical results. The experiments in Sections 5.1–5.3 are the primary evidence; FRS adds structure. KGW at 37/60 is correctly flagged as NOT FORENSIC READY. So is SynthID at 35/60—it fails on score alone, and the empirical picture (false positives, near-total uncertainty) is even worse than the number suggests. The interesting case is Unigram. It lands at exactly 40/60—the minimum passing score—and clears all three gates. By the FRS rubric, it is CONDITIONALLY FORENSIC READY. But it has the highest FNR (83%), 100% conditional removal, and the worst baseline detection of any method. One point less and it would be NOT READY. This is the clearest demonstration we have that point-based forensic scoring can mask underlying uselessness. Table 7 maps each method against the five Daubert factors. All three pass Factors 1 and 2 (they are testable and peer-reviewed). All three fail Factor 3 (no stable error rate) and Factor 4 (no forensic standards existed before this

Table 7: Daubert factor assessment by method. “Part.” denotes partial satisfaction: accepted in research literature, but not in forensic practice. Method

F1

F2

F3

F4

F5

KGW Unigram SynthID

Pass Pass Pass

Pass Pass Pass

Fail Fail Fail

Fail Fail Fail

Part. Part. Part.

work). Factor 5 gets a “partial”—accepted in the research community but not in forensic practice. To be clear: we do not mean that error rates cannot be measured in a controlled experimental notebook. Rather, the measured rates are high, attack-sensitive, and configuration-dependent, preventing a stable operational error rate from being generalized to caselevel forensic use under Rule 702. Passing only two out of five Daubert factors makes admissibility a steep climb.

Discussion Implications for Courts We want to be precise about what these experiments show. They do not show that watermark detectors are stochastically unreliable in the sense of producing random outputs. The detectors are deterministically repeatable—run the same input twice, get the same score. What collapses is the evidentiary signal: once a text undergoes meaningpreserving transformation, the detection result changes. This is reproducible forensic failure, not noise. Judges are going to encounter watermark evidence. The FRS framework gives them a structured way to evaluate it using the forensic language they already work with. But the bottom line is simpler than the framework: until watermarking methods can demonstrate stable error rates under realistic adversarial conditions, watermark evidence deserves serious skepticism. There is a broader issue, too. Even a positive watermark signal may not settle the authorship question. Distillation and spoofing research shows that a detected watermark might reflect direct generation, inherited traces from model distillation, or deliberate forgery (Pan et al. 2025; Yi et al. 2025). Watermarking sits inside a larger, unsolved problem of digital evidence authentication in the AI era (Bellovin et al. 2024). Courts have been down this road before. The NAS report (National Research Council 2009) documented how forensic methods were admitted and relied upon for years before anyone checked whether the science held up. We think AI watermarking is at risk of repeating that mistake.

Implications for Policymakers The EU AI Act (European Parliament and Council of the European Union 2024) and California’s SB 942 (California State Legislature 2024) both mandate watermarking, but every method we tested fails forensic admissibility standards in our evaluated configurations. This is essentially what Nemecek, Jiang, and Ayday (2025) called “symbolic compliance”—the mandate exists, implementations exist,

Not removed

KGW

Paradox/increase

No score change

Unigram

2

1

Attacked score - pristine score

Removed

Similarity threshold (0.75)

1

0

0.02

1

0

2

1

3

2

4

3

0.04

5

4

0.06

6

0.75

0.80

0.00 0.02

5

median sim = 0.84 removed = 100/304

0.85

0.90

Semantic similarity

0.95

SynthID

0.04

median sim = 0.83 removed = 41/306

0.75

0.80

0.08 0.85

0.90

Semantic similarity

0.95

1.00

median sim = 0.83 removed = 58/236

0.750 0.775 0.800 0.825 0.850 0.875 0.900 0.925 0.950

Semantic similarity

Figure 2: Within-method score change under paraphrase attack for KGW, Unigram, and SynthID. Each point is a valid paraphrase run; the x-axis gives semantic similarity and the y-axis gives attacked score minus pristine score. Orange points denote removal events, green points denote non-removal, and purple points denote paradoxical score increases. Across all three methods, score drops persist even when semantic similarity remains high (≥ 0.75), supporting the forensic claim that watermark disruption does not require meaning destruction. Score scales differ by method, so comparisons are interpreted within each panel rather than across panels. n = 304 (KGW), 306 (Unigram), 236 (SynthID). but the connection between the two has not been validated. Liang et al. (2025) showed that watermark designers face a basic trade-off: methods that keep text quality high (like KGW and Unigram) tend to be fragile, while robust methods degrade quality. And the cross-model variability they documented (TPR ranging from 0.858 to 0.334 for the same method on different models) makes it essentially impossible to establish a stable “known error rate” across deployments—a direct Daubert Factor 3 violation.

Implications for AI Companies The MarkLLM SynthID configuration scored lowest of the three methods (35/60), with a detection margin of roughly 0.03, a 5.4% false positive rate, 80% UNCERTAIN verdicts on its own pristine output, and T4 = 0 (93.6% verdict instability under threshold perturbation). A standard human typing test sentence scored above the detection threshold. The pattern of findings across our work, Nemecek, Jiang, and Ayday (2025), and Liang et al. (2025) all point in the same direction: deployment interest has gotten ahead of forensic validation.

The Scoring Limit Finding Unigram’s exact-threshold pass (40/60) while exhibiting the worst empirical performance of any method—100% conditional removal, 83% FNR—is a finding in its own right. It shows that criteria-based scoring frameworks, no matter how carefully designed, have inherent limits. Forensic validation for AI systems needs adversarial testing alongside the scorecard, not folded into it.

Limitations and Future Work Model scale and families. We tested two model families (Qwen2.5 and Gemma-2) at 1.5B–9B scale. Liang et al. (2025) found cross-model fragility persists at 14B, and

the statistical properties underlying these failures are scaleindependent. Testing additional families like LLaMA or Mistral would widen coverage, but given the category-wide nature of the failure, we would not expect different forensic conclusions. Attack scope. We tested one attack type. Liang et al. (2025) tested twelve and found consistent fragility everywhere. Their ChatGPT paraphrase result—all methods dropping below 30% TPR—suggests stronger attackers would produce worse outcomes than what we report. We also ran earlier word deletion pilot experiments that informed our anomaly taxonomy but used a different experimental protocol, so we do not report those as primary results. Text length and sample size. Our texts are 50 tokens long, and we use 30 pristine texts (15 prompts × 2 seeds) per method. Longer texts and larger sample sets are needed in future work. Reproducibility scope. All experiments use a single global seed (123). Our replicate sessions confirm deterministic reproducibility, but they are not cross-seed replications. We do have cross-seed data from earlier pilot experiments (seed 42 vs. 123 for KGW/Unigram), which showed identical failure patterns, but this was under a slightly different protocol. Implementation vs. production. We test the open-source MarkLLM implementation of SynthID, not Google’s proprietary production system. We cannot speak to how the production version would perform. The FRS is also one possible operationalization of forensic readiness—alternative designs might score differently. Framework validation breadth. The FRS criteria and gates draw on Daubert and NIST, but we have not yet conducted inter-rater reliability testing or a systematic study of how different gate thresholds would affect outcomes. Those are next steps for maturing the framework.

120

1.00

0.840

0.832

0.829

Min. threshold (0.75) KGW (n=304)

Unigram (n=306)

Watermarking Method

SynthID (n=236)

Figure 3: Semantic similarity distribution across paraphrase attacks (all three methods). Box plots show median, quartiles, and range of cosine similarity for retained paraphrases. The green dashed line indicates the minimum acceptable semantic threshold (0.75). All methods show medians between 0.83–0.84 and retain samples exclusively above 0.75, confirming that watermark removal occurs without meaning destruction. n = 304 (KGW), 306 (Unigram), 236 (SynthID).

98.3%

60 40

0

0.75

100.0%

80

32.9%

24.6%

20

0.80

0.70

Removal rate (%)

Cosine Similarity

0.90

0.85

100.0%

100

0.95

13.4%

KGW

Unigram

Watermarking Method

Raw removal rate

SynthID

Conditional removal rate

Figure 4: Raw vs. conditional watermark removal rates. Raw removal (blue) counts all texts losing detection; conditional removal (red) counts only texts that were initially detectable. The gap exposes the forensic problem: raw rates appear modest (13–33%), but among the small fraction of texts a court could actually rely on, removal is near-total (98– 100%).

new capabilities; we are documenting the forensic consequences of capabilities that already exist, which we believe serves the public interest.

Conclusion All three watermarking methods we tested—KGW, Unigram, and the MarkLLM SynthID configuration—fall short of forensic evidence admissibility standards. Meaningpreserving paraphrase eliminates watermark detection in 100% of initially-detected texts for KGW and Unigram, and 98.3% for SynthID, while also revealing a 5.4% false positive rate on clean text for SynthID. All three methods fail at least two of five Daubert factors. The FRS framework we propose gives courts and policymakers a structured tool for evaluating watermark evidence—while also revealing, through the Unigram boundary case, that point-based scoring has inherent limits. The gap between what “technical robustness” means in an ML paper and what “forensic admissibility” means in a courtroom has to be closed before this evidence shows up at scale. Our framework is a first step. The alternative—letting unvalidated evidence into legal proceedings without anyone having checked it against the standards courts actually use—risks repeating the failures documented by the NAS (National Research Council 2009) and PCAST (President’s Council of Advisors on Science and Technology 2016) reports. Those failures were measured in wrongful convictions.

Ethical Considerations Everything in this paper evaluates publicly available systems. We use open-source models (Qwen2.5, Gemma-2) and a public toolkit (MarkLLM). No human subjects are involved. The attack we test—paraphrasing—is something anyone with internet access can do. We are not introducing

Researcher Positionality Our background is in digital forensics and adversarial security, not ML optimization. That shaped how we approached this work: we evaluated watermarks against the standards courts actually use to decide whether evidence is admissible, rather than measuring aggregate detection accuracy the way most ML robustness studies do. We think that framing is important because policymakers are already writing watermarking into law, and someone needs to ask whether the evidence holds up in the room where it will actually be used. That said, this lens has blind spots. Watermarks may serve purposes beyond courtroom evidence — content provenance tracking, platform-level filtering, or internal audit — where the forensic bar we apply is not the right yardstick. Our evaluation does not speak to those use cases, and readers should keep that scope in mind.

Adverse Impact Statement Showing that watermarks are fragile could undermine trust in watermarking prematurely, or be used to argue against mandates. But we think the greater risk is the alternative: deploying unreliable evidence technology in courts without anyone having validated it first. History supports this view. Delayed validation of forensic methods—bite mark analysis, hair microscopy, algorithmic risk scores like COMPAS (State v Loomis)—has led to documented wrongful convictions. Proactive validation, even when the results are negative, is the responsible path.

Acknowledgments We thank Alexander Nemecek for reviewing an earlier version of this work and for helpful discussion on watermark governance.

References Bellovin, S. M.; et al. 2024. Seeking Reliable Digital Evidence in the Age of AI. Columbia University Working Paper. California State Legislature. 2024. California AI Transparency Act, SB 942, Chapter 291. Approved by Governor, September 19, 2024; effective January 1, 2026. Dathathri, S.; See, A.; Ghaisas, S.; Huang, P.-S.; McAdam, R.; Welbl, J.; Bachani, V.; Kaskasoli, A.; Stanforth, R.; Matejovicova, T.; et al. 2024. Scalable Watermarking for Identifying Large Language Model Outputs. Nature, 634(8035): 818–823. Daubert. 1993. Daubert v. Merrell Dow Pharmaceuticals, Inc. 509 U.S. 579. Supreme Court of the United States. European Parliament and Council of the European Union. 2024. Regulation (EU) 2024/1689: Artificial Intelligence Act. Official Journal of the European Union, L 2024/1689. Article 50. Federal Rules. 2023. Federal Rules of Evidence, Rule 702: Testimony by Expert Witnesses. As amended Dec. 1, 2023. Kirchenbauer, J.; Geiping, J.; Wen, Y.; Katz, J.; Miers, I.; and Goldstein, T. 2023. A Watermark for Large Language Models. In Proceedings of the 40th International Conference on Machine Learning (ICML), 17061–17084. PMLR. Liang, J.; Wang, Z.; Hong, S.; Ji, S.; and Wang, T. 2025. Watermark under Fire: A Robustness Evaluation of LLM Watermarking. In Findings of the Association for Computational Linguistics: EMNLP 2025, 21050–21074. National Institute of Standards and Technology. 2006. Guide to Integrating Forensic Techniques into Incident Response. Technical Report SP 800-86, NIST. National Institute of Standards and Technology. 2024. Reducing Risks Posed by Synthetic Content: An Overview of Technical Approaches to Digital Content Transparency. Technical Report AI 100-4, NIST. National Research Council. 2009. Strengthening Forensic Science in the United States: A Path Forward. Technical report, The National Academies Press, Washington, DC. Nemecek, A.; Jiang, Y.; and Ayday, E. 2025. Watermarking Without Standards Is Not AI Governance. In Workshop on Technical AI Governance (TAIG), ICML 2025. Pan, L.; Liu, A.; He, Z.; Gao, Z.; Zhao, X.; Lu, Y.; Zhou, B.; Liu, S.; Hu, X.; Wen, L.; et al. 2024. MarkLLM: An OpenSource Toolkit for LLM Watermarking. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: System Demonstrations (EMNLP Demo). Pan, L.; Liu, A.; Huang, S.; Lu, Y.; Hu, X.; Wen, L.; King, I.; and Yu, P. S. 2025. Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation? In Proceedings

of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL), Volume 1: Long Papers, 13228– 13251. Vienna, Austria: Association for Computational Linguistics. Pang, Q.; Hu, S.; Zheng, W.; and Smith, V. 2024. Attacking LLM Watermarks by Exploiting Their Strengths. arXiv:2402.16187. President’s Council of Advisors on Science and Technology. 2016. Forensic Science in Criminal Courts: Ensuring Scientific Validity of Feature-Comparison Methods. Technical report, Executive Office of the President. Reimers, N.; and Gurevych, I. 2019. Sentence-BERT: Sentence Embeddings Using Siamese BERT-Networks. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP). Sadasivan, V. S.; Kumar, A.; Balasubramanian, S.; Wang, W.; and Feizi, S. 2023. Can AI-Generated Text Be Reliably Detected? arXiv:2303.11156. State v Loomis. 2016. State v. Loomis. 881 N.W.2d 749 (Wis. 2016). Supreme Court of Wisconsin. The White House. 2023. Executive Order 14110: Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence. Federal Register, Vol. 88, No. 210. Rescinded by Executive Order 14179, January 2025. The White House. 2025. Executive Order 14179: Removing Barriers to American Leadership in Artificial Intelligence. Yi, X.; Li, Y.; Zheng, S.; Wang, L.; Wang, X.; and He, L. 2025. Unified Attacks to Large Language Model Watermarks: Spoofing and Scrubbing in Unauthorized Knowledge Distillation. Knowledge-Based Systems, 329: 114295. Zhang, H.; Edelman, B. L.; Francati, D.; Venturi, D.; Ateniese, G.; and Barak, B. 2024. Watermarks in the Sand: Impossibility of Strong Watermarking for Generative Models. In Proceedings of the 41st International Conference on Machine Learning (ICML). Zhao, X.; Ananth, P.; Li, L.; and Wang, Y.-X. 2024. Provable Robust Watermarking for AI-Generated Text. In Proceedings of the 12th International Conference on Learning Representations (ICLR).

Record · ID 381682 · SHA-256 0321c3042bbfc7a8
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.