Conceptio › Archive › arXiv CS
arXiv CSopen access

ZK-Trace: Certified Collusion Tracing with Zero-Knowledge Credentials for Federated GNSS Interference Monitoring

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

Index Terms— Federated Learning, Model Attribution, Traitor Tracing, Watermarking, Zero-Knowledge Proof, Few-Shot Learning, Global Navigation Satellite System, Interference Classification

arXiv:2609.08763v1 [cs.CR] 8 Sep 2026

ZK-Trace: Certified Collusion Tracing with Zero-Knowledge Credentials for Federated GNSS Interference Monitoring

Redwanul Karim Nisha L. Raichur Lucas Heublein Tobias Feigl Christopher Mutschler Felix Ott Fraunhofer Institute for Integrated Circuits IIS, 90411 Nürnberg University of Technology Nürnberg (UTN), 90461 Nürnberg Abstract— Federated global navigation satellite system (GNSS) monitoring distributes a proprietary classifier to partly trusted stations, any of which may leak its copy. ZK-Trace combines public identity marks, recipient-specific Tardos fingerprints, and zero-knowledge credential verification. The registry supports offline tracing without the leaker’s cooperation. We establish conditional false-accusation bounds for arbitrary recovered bit patterns, a finite completeness bound under a hidden-bias residual channel, and a deterministic tracing-score bound for correlated feature-distillation errors. An interval-arithmetic checker makes the conditional bound executable and allocates a common budget across accusation and tamper decisions. Under innocent-row independence, the certificatebased evaluation uses a false-naming budget of 10−3 per investigation. It isolates all 160 single-owner copies and traces 712 of 720 two-owner mixtures without naming an innocent. Experiments use a simulated GNSS federation and CIFAR-10. Feature matching preserves the feature mark in 20/20 runs and cross-architecture transfer in 19/20, at copy-accuracy costs of 4.8 and 6.1 percentage points on GNSS and CIFAR-10. Function-only distillation erases the feature mark, and distillation also removes weight-space marks. These results support verifiable tracing under explicit statistical and cryptographic assumptions. Credential knowledge and recipient evidence serve distinct roles.

Manuscript received September 08, 2026; revised XXXXX 00, 0000; accepted XXXXX 00, 0000. This work has been carried out within the DARCII project, funding code 50NA2401, sponsored by the German Federal Ministry for Economic Affairs and Climate Action (BMWK), and supported by the German Space Agency at DLR, the Bundesnetzagentur (BNetzA), and the Federal Agency for Cartography and Geodesy (BKG). We used ChatGPT 5.4 solely for writing improvements. (Corresponding author: F. Ott). All authors are with the Fraunhofer Institute for Integrated Circuits IIS, 90411 Nürnberg, Germany (e-mail: {redwanul.karim, nisha.lakshmana.raichur, lucas.heublein, tobias.feigl, christopher.mutschler, felix.ott}@iis.fraunhofer.de).

0018-9251 © 2026 IEEE UNDER REVIEW

VOL. XX, No. XX

September 2026

I. INTRODUCTION

Federated learning (FL) [1] allows sensor stations to train a shared model while keeping their measurements local. This is a natural fit for global navigation satellite system (GNSS) interference monitoring [2], but it also distributes a proprietary classifier to partly trusted participants. A station that leaks its issued controlled-receptionpattern antenna model exposes the operator’s detection capabilities. An adversary can then inspect which jamming and spoofing patterns evade detection, with consequences for positioning, navigation, and timing (PNT). Tracing such a leak requires evidence linking the model to an enrolled station. Existing federated watermarks provide only part of that evidence. Zero-knowledge ownership schemes such as FedZKP [3] certify group membership without identifying a particular client. Perclient fingerprints identify recipients, but are assigned by the server and read in plaintext, without binding to a client secret [4], [5]. Collusion adds a further difficulty: stations can combine their copies to weaken individual marks. Collusion-secure fingerprinting codes address this attack [6], [7], but their guarantees do not directly cover the extraction errors introduced by federated averaging (FedAvg). ZK-Trace separates two questions: which credentials are represented in the shared model, and which enrolled recipient is associated with a leaked copy (Fig. 1). Public identity codewords address the first question, while a recipient-specific Tardos fingerprint [7], [8] addresses the second question. The registry supports offline comparison, and a zero-knowledge proof authenticates the claimant during a dispute. The central analysis turns tracing into an auditable decision: it bounds an innocent score conditional on the recovered bits, then separately bounds missed colluders under a specified channel. The construction supports a BN-scale carrier [9] and a feature carrier whose score can be stable under feature matching. We evaluate ZK-Trace for federated few-shot GNSS interference monitoring, where labeled events are scarce [10]. Prototypical networks [11] provide the classifier. Experiments use ten seeds, ten simulated stations partitioned from real GNSS recordings, and CIFAR-10 as a transfer check. This setting tests whether tracing remains useful when data are imbalanced, participation varies, and attackers modify their issued copies. Contributions. (C1) Finite and executable tracing guarantees. We give a conditional false-accusation bound that does not assume independent extraction errors (Theorem 3), together with a finite hidden-bias completeness bound allowing coalition decisions across positions (Theorem 4). An interval checker enforces the bound and a shared error budget. Exhaustive evaluation of two1

owner mixtures traces 712/720 mixtures and isolates all 160 single-owner copies without naming an innocent (Sec. VII-B). (C2) Credential verification and score stability. We extend FedZKP’s credential mechanism [3] with perclient identity blocks, recipient fingerprints, and a registry. The artifact-bound verifier authenticates a credential statement about a complete model state. It does not claim authorship. A deterministic weighted-score bound establishes when feature perturbations preserve tracing even with correlated or targeted errors (Theorem 5). (C3) A GNSS evaluation with explicit robustness limits. A common attack suite evaluates five baselines on their native mark metrics. Distillation erases all tested weight-space marks. Our feature carrier survives feature-matching distillation and nineteen of twenty crossarchitecture runs, but function-only distillation erases it. Its dispatched-copy accuracy cost is 4.8 percentage points on GNSS and 6.1 on CIFAR-10. No innocent is accused in the coalition sweep, whereas a DeepMarks-family comparator [12] produces false accusations (Sec. VII-C). II. RELATED WORK

Watermarking and fingerprinting codes. Model ownership, recipient identification, and survival under model modification are distinct capabilities. White-box watermarks encode information in model parameters, whereas black-box watermarks use queryable triggers. Collusion-secure fingerprinting identifies a source when recipients combine their copies [6]. Tardos codes achieve length O(c2 log(N/ε1 )) for coalition size c and falseaccusation target ε1 [7]. Symmetric scoring [8] and later refinements improve the score and length analysis [13]– [16]. Our analysis complements noisy-Tardos work on additive white Gaussian noise [17] with conditional score certification and a finite residual-channel bound for model fingerprints. Federated model ownership. FedIPR [9] and WAFFLE [18] embed ownership marks in models redistributed through FedAvg [1]. Their verification establishes mark presence but may expose the secret or depend on its holder. Proofs of knowledge address credential disclosure. Our verification uses the Σ-protocol [19] for exact-weight Learning Parity with Noise (xLPN) given by Jain et al. [20], following Véron’s identification formulation [21]. This cryptographic layer authenticates a credential. Recipient tracing additionally requires a per-copy identifier. Per-client attribution. The closest federated methods provide different parts of the required evidence. FedZKP [3] derives a group watermark from all clients’ xLPN public inputs and proves ownership in zero knowledge, but does not identify a recipient. FedTracker [4] and DUW [5] provide per-client fingerprints assigned by a trusted party, without a collusion-secure code. DeepMarks [12] embeds anti-collusion codes centrally, but assumes exact extraction and lacks a false-accusation bound at arbitrary coalition sizes. 2

Cryptographic participation evidence serves a different purpose. FedPoP [22] proves participation anonymously and leaves no in-model tracing artifact. FedAaT [23] adds per-client output-space sequences, but reads them in plaintext and confines zero knowledge to the credential layer. BlackCATT [24] uses per-client Tardos labels in a federated trigger channel. It does not establish their composition through the aggregation channel considered here. Distillation and the remaining gap. Removal attacks include extraction-based erasure [25] and knowledge distillation [26], within the taxonomy of [27]. DAWN [28] marks prediction-API responses and evaluates coalition resistance empirically. Entangled watermarks [29] can survive distillation but do not identify individual leakers. ZK-Trace combines recipient tracing with credential authentication and certificate-gated adjudication (Table IX). ZK-Trace builds on FedZKP’s credential-based watermarking framework [3], using BN scaling parameters as the embedding carrier, as in FedIPR [9]. It analyzes Tardos tracing through FedAvg and adds a feature carrier whose survival depends on what the distiller reproduces. Federated few-shot GNSS monitoring. Gaikwad et al. [10] combine episodic prototypical learning with FedAvg for GNSS interference classification under noni.i.d. data. We build on this learning setup to address model ownership and tracing. The evaluation studies watermarking on an episodic prototypical model. III. PRELIMINARIES

Notation. For binary vectors x, y ∈ {0, 1}m , ∥x∥1 denotes Hamming weight and x⊕y denotes bitwise XOR. Their Hamming distance is ∥x ⊕ y∥1 . We write x || y R for concatenation and x ← − X for a uniform draw. The watermark bit subset B ⊆ {1, . . . , n} differs from the Bernoulli noise law Bτ . Table I summarizes the symbols. Section IV-A defines the threat model. Federated learning. A server coordinates Nc clients, each with private data Di and local parameters Wi . PN c FedAvg forms W = k=1 λk Wk , where λk ≥ 0 and P k λk = 1 [1]. The same weighted average applies to the BN scales that carry the identity marks. The survival analysis asks whether those marks remain distinguishable after averaging. Prototypical networks. An N -way K -shot episode samples N classes, with K labeled support examples and Q query examples per class. Let Sk be class k ’s support set. The embedding fθ : P X → Rd forms its prototype −1 as the mean ck = |Sk | (xj ,yj )∈Sk fθ (xj ). A query is classified by a softmax over negative squared Euclidean distances to these prototypes [11]. Training minimizes the query negative log-likelihood over episodes. We use the federated few-shot setup of Gaikwad et al. [10]. Credentials. In search LPN, the public pair (A, y) satisfies y = As ⊕ e, with noise rate 0 < τ < 21 . Recovering the secret is an average-case random-code decoding problem [30]. The xLPN variant fixes the error UNDER REVIEW

VOL. XX, No. XX

September 2026

TABLE I: Principal notation and deployed dimensions. Identity length n and tracing length nt describe separate watermark layers. Symbol

Meaning (space / value)

Nc ci , Di W, Wi λk γ, γ agg ω n wi , ĥi E, Ei ρ Ai , yi si , ei Bτ , τ B errn , pr fθ , d N, K, Q ck Xi , nt c, k Si , Z, ε1 q

number of clients (= 10) client i and its private dataset global / local model P parameters FedAvg weights ( k λk =1) BN scale carrier (∈ Rω ) carrier dimension (= 4800) per-client codeword length (= 128) codeword / extracted component (∈ {0, 1}n ) shared / client projection (Rω×Nc n , Rω×n ) identity-column load Nc n/ω xLPN public input ({0, 1}m×l , {0, 1}m ) xLPN witness ({0, 1}l , {0, 1}m ; ∥ei ∥1 = wτ ) i.i.d. Bernoulli noise, rate τ per-episode WM bit subset (⊆ {1, . . . , n}) near-collision radius, detection threshold embedding net, feature dim (Rd , d=512) episodic way / shot / query class prototype (∈ Rd ) tracing row / length (512 weight, 2048 feature) collusion design target / tested coalition size tracing score / threshold / false-accusation target residual flip probability relative to intended output

weight to wτ = ⌊mτ + 0.5⌋ [20]. Client i publishes (Ai , yi ) and retains the witness (si , ei ). Verification establishes knowledge of this witness using a three-move commitment–challenge–response Σprotocol [19]. Its non-interactive form binds the challenges to the extracted model component and the registered credential (Sec. IV-C). Appendix C-C specifies the commitment and random-oracle assumptions. IV. ZK-TRACE CONSTRUCTION

Pipeline. ZK-Trace has two watermark layers (Fig. 1). The identity layer extends FedZKP [3]: each client’s xLPN public input determines a codeword embedded in the shared model’s BN scales. The tracing layer adds a secret Tardos fingerprint only to the copy dispatched to that client. It supports attribution after modifications invalidate an exact dispatch-log hash, providing evidence for trace-and-revoke [31]. Fig. 1 follows one station through four stages. At (A), station ci registers (Ai , yi ) and its derived codeword wi , keeping the witness (si , ei ) private. At (B), federated training aggregates the identity codewords into W while leaving the tracing block ET empty. At (C), the tracer embeds the secret row Xi into the copy issued to station i. At (D), the tracer extracts the leaked copy’s overlay, scores each registered row, and accuses a station when its score exceeds Z . Binding intuition. Each client has an assigned set of projection directions at which to read its public codeword.

Agreement with that codeword is statistical evidence of a mark. Possession of its private xLPN witness is a separate credential fact. A valid proof authenticates the witness holder. The verifier also includes a digest of the entire model state in its context, so equal extracted bits do not permit replay onto different model bytes. Neither presence nor authentication establishes who trained the model: public marks can be copied. Recipient tracing therefore uses the separate secret row and registry. A. Setting and Threat Model

The server acts as tracer and issues each authorized client a copy bearing its secret Tardos row. The registry allows a recovered copy to be traced without contacting its holder. Clients keep their witnesses private, and the tracer holds the tracing rows. Threat model. We consider authorized clients that hold valid credentials but may misuse their dispatched copies. Access control for unauthorized parties is a separate concern. • T1, weight-space post-processor. A leaker applies utility-preserving edits such as fine-tuning, pruning, quantization, or noise before redistributing its copy. We evaluate both identity attribution and Tardos tracing under these edits. Their guarantees require the extraction conditions stated in Sec. V. • T2, bounded coalition. Clients may average their copies to attenuate the marks [6]. The design targets are two colluders on the weight carrier and three on the feature carrier. Soundness requires innocentrow independence, A5(a). Finite channel completeness additionally requires A5(b). A coalition may base its intended symbols on its entire row matrix under that theorem. The experiments distinguish model averaging from decoder-space adaptive steering. • T3, framing adversary. An adversary may copy a victim’s public identity mark. It lacks the victim’s witness and secret tracing row. Certified false naming is bounded under A5(a), while credential verification authenticates the claimant. Authentication alone does not exculpate a recipient: the judge must examine the accusation evidence. • T4, distiller. A leaker may train a fresh student whose BN parameters contain no identity codeword. This erases the weight carrier in our benchmark [27]. The feature carrier can survive when the student copies the teacher’s feature geometry closely enough (Proposition 2). Function-only distillation falls outside Assumption 6 and erases this carrier too. Out of scope. ZK-Trace does not detect or prevent poisoned updates. It can be combined with robust aggregation to address this threat (Sec. VII-C). Credential redistribution is also outside its scope. The fingerprint identifies the enrolled recipient of the issued copy, and subsequent revocation is handled through the registry [31].

KARIM ET AL.: ZK-Trace: Collusion-Secure Traitor Tracing with Zero-Knowledge Exculpation

3

(A) ENROL

(B) FEDERATED TRAINING

Station ci witness (si, ei) secret

E EETTT local models each station writes only its own block

ET copyi — dispatched to station i

SHAKE-128 ET

wi — public codeword

global model W — all codewords superposed ET empty — no tracing mark in the shared model Theorem 3: the accusation composes through this channel

Station

(D) LEAK, TRACE, ADJUDICATE

Tracer writes station i's secret Tardos row Xi into the dispatched copy

FedAvg: γagg = Σk λk γ(k)

public (Ai, yi)

Ci

(C) DISPATCH

Server / Tracer

Adversary

Leaked copy: fine-tune, prune, quantise, distil decoded from the leaked copy flip rate q E T

Z ET = secret row Xi — a framer cannot stamp it

accuse iff max Si > Z

Tardos scores Si weight carrier: nt = 512, collusion c = 2 feature carrier: probes → margins, c ≈ 6, survives distillation Judge

certified-attribute tamper forensic-lead

Exculpate: zero-knowledge proof of (si, ei) if framed

clean

Nc = 10 · ω = 4800 · n = 128 · ρ = 0.27 · r = 331

Enrolment registry (append-only): Ci = Com(wi ‖ Xi ‖ Ai ‖ yi) opening escrowed with the judge — tracing needs no cooperation from the leaker

Fig. 1: ZK-Trace life cycle and carrier layout. The bar is the projection codebook read from the batch-normalization scales γ , partitioned into per-client identity blocks E1 , . . . , ENc and the tracing overlay block ET . Solid cells carry the public codeword wi , hatched cells the tracer-secret Tardos row Xi . Si is the Škorić-symmetric score, Z the accusation threshold, and q the per-bit flip rate induced by aggregation.

Server as leaker. The server and tracer are honest-butcurious. The server holds the shared model and issued copies, so the scheme cannot distinguish a server leak from a recipient leak of the same copy. Registry commitments prevent post-hoc row substitution. Authenticated delivery and server accountability require additional evidence (App. A-D). B. Codeword Embedding

b∈B

Before training, initialization fixes the shared Gaussian matrix E ∈ Rω×Nc n , with ω=4,800 and n=128, and the detection threshold pr and radius errn . Following FedZKP [3], we select an xLPN instance with a syndrome-decoding work estimate reported near 2152 [30] as a separate computational-hardness estimate (Sec. IVC; dimensions in App. A). Each client’s public input determines its identity codeword, wi = SHAKE-128(Ai | yi ) ∈ {0, 1}n .

(1)

Client i uses block Ei ∈ Rω×n , comprising columns [(i−1)n, in) of E . The server retains the concatenated public inputs (Aagg , yagg ) so a verifier can recompute any station’s codeword. Carrier. The vector γ ∈ Rω concatenates the scales of all L=20 BN layers. Each projection spreads a bit across these scales, so editing one layer affects only part of its carrier. At Nc =10, the identity-codebook load is Nc n/ω ≈ 0.27. BN scales rescale activations after 4

normalization. Section VII-C evaluates how this affects feature geometry and classification. Projection and objective. For bit b, the projection zi,b = γ ⊤ Ei,b gives the extracted bit ĥi,b = ⊮[zi,b > 0], with signed target ti,b = 2wi,b − 1. The hinge loss penalizes a wrong sign or a signed projection smaller than the target margin µ > 0,  1 X Lwm = max 0, µ − ti,b · zi,b , (2) |B| with the active subset B ⊂ {1, . . . , n} drawn each step by bit-dropout (pdrop = 0.5, µ = 1.0). The local objective adds this to the prototypical task loss, L = Ltask + λwm Lwm , with λwm = 0.1 trading accuracy against watermark strength. Federated training. Cross-entropy pre-training on base classes produces an unwatermarked initial model [10]. Federated episodic training then follows Sec. III, with λk proportional to local sample count. Each client fixes its public input and codeword in the first round and trains under L in subsequent rounds. Projection onto Ei recovers client i’s component from the averaged BN scales, subject to cross-talk from the other clients (Sec. V). C. Zero-Knowledge Verification

Algorithm 1 verifies credential knowledge using a Stern-type non-interactive proof [3]. All commitments UNDER REVIEW

VOL. XX, No. XX

September 2026

Algorithm 1 Non-interactive credential and modelpresence verification, from FedZKP’s Stern-type xLPN Σ-protocol [3]. Require: witness (si , ei ), yi =Ai si ⊕ei ; block (Ai , yi ); model W; rounds r=331 Ensure: verify credential knowledge and mark presence, else reject 1: ĥi ← ⊮[γ(W)⊤ Ei > 0], ctx ← tag∥H256 (W)∥ĥi ∥Ai ∥yi 2: Prover (offline), for k=1, . . . , r with fresh π, v, f : 3: t0 ←Ai v⊕f , t1 ←π(f ), t2 ←π(f ⊕ei ) 4: C0 ←com(π, t0 ), C1 ←com(t1 ), C2 ←com(t2 )  (1) (r) 5: (c(k) )rk=1 ← SHAKE-128 C0 ∥ · · · ∥C2 ∥ ctx mod 3 6: open the two commitments per c(k) and send transcript τ 7: Verifier (offline). Recompute (c(k) ) from τ, ctx and check per round: 8: c(k) =0: t0 ⊕ π −1 (t1 ) ∈ img(Ai ) 9: c(k) =1: t0 ⊕ π −1 (t2 ) ⊕ yi ∈ img(Ai ) 10: c(k) =2: wt(t1 ⊕ t2 ) = wτ 11: accept iff all r checks pass and HD(ĥi , wi ) ≤ errn , else reject

and the public context determine the challenge vector. In the verifier, ctx = tag∥H256 (W)∥ĥi ∥Ai ∥yi , where H256 hashes a canonical full state including buffers. We use r = 331 repetitions. Appendix C-C separates the grinding calibration from end-to-end witness-recovery security. Attribution and acceptance. From the public block Ei and codeword  wi alone, a verifier extracts ĥi = ⊮ γ(W)⊤ Ei > 0 and tests codeword presence,  HD ĥi , wi ≤ errn . (3) Against an independent codeword at n=128, calibration to 2−128 requires errn =0. Exact presence is restrictive for the measured noisy extraction, so the deployed identity metric uses maximum codeword agreement (Prop. 4, App. C). A match is statistical evidence of a codeword, but it does not authenticate its holder because the public mark can be copied. Algorithm 1 adds proof of the private witness and retains the explicit presence condition in its acceptance rule. Recipient tracing instead uses the secret overlay below.

The proof, with extractor, simulator, and grinding accounting, is in Appendix C. Two soundness layers. Credential verification and recipient tracing answer different questions. The former proves knowledge of a private witness and binds the statement to an artifact. The latter bounds naming an innocent recipient under a conditional code model. The 331-round grinding calibration, commitment failures, computational witness recovery, and statistical tracing budget must be accounted for separately (Table XIII). E. Tracing Overlay and Two Carriers

Dispatch and readout. The identity layer cannot isolate a recipient because the aggregate contains every contributor’s mark. At dispatch, the tracer fine-tunes a copy on server-held proxy data, combining the task loss with a hinge that embeds the recipient’s secret Tardos row Xi . The weight carrier reads each bit from a BNscale projection. The feature carrier reads the sign of a feature projection on a fixed probe input, with co-training used to align those signs with Xi . Its design target is c=3, compared with c=2 for the weight carrier. Featurematching distillation can preserve these probe responses under Assumption 6. Tracing and adjudication. The tracer scores the recovered bits against registered rows. The adjudication rule certifies each positive or negative tail before issuing a certified decision, allocating one budget across both tails and any jointly used carriers (Prop. 3). A threshold exceedance without a passing certificate remains an uncertified lead. If neither threshold is exceeded, the outcome is no certified evidence. Threshold-only decisions are evaluated separately to measure score separation. An independent judge checks the enrollment opening and reconstructs the score evidence. A credential proof authenticates the claimant but does not decide whether that claimant leaked. V. THEORETICAL ANALYSIS

D. Security Guarantees

Here Q denotes the number of random-oracle queries, distinct from query examples per episode. P ROPOSITION 1 (Credential knowledge and artifact binding) Under the random-oracle and commitment assumptions in Appendix C-C, the ideal Stern/Fiat–Shamir protocol is complete for a witness holder whose codeword passes the requested presence test, and admits a zeroknowledge simulation. With ideal binding commitments and uniform challenges, its knowledge error is at most (Q+1)(2/3)r [32]. The artifact-bound context ties a transcript to the registered credential, extracted component, and full-state digest. Transfer to a different state requires a digest collision or a new-context challenge coincidence. These are knowledge and binding properties, not a proof of authorship or a numerical bound on all xLPN-recovery attacks.

The analysis separates three questions: when identity bits survive averaging, when a tracing score supports a bounded false-accusation probability, and when feature matching preserves probe bits. The identity probability model is idealized. The deterministic decoding and probemargin results do not require that model. Appendix B-A states the assumptions and Appendix C gives the proofs. Identity Write γ (i) = γ0 + ∆i , G = ∥γ0 ∥, P survival. (k) and ui = k̸=i λk γ . A local signed margin of at least µ gives ti,b ⟨γagg , Ei,b ⟩ ≥ λi µ + ti,b ⟨ui , Ei,b ⟩.

(4)

A probabilistic recovery bound follows from this inequality when the signed directions are independent of the cross-vector. T HEOREM 1 (Conditional per-bit recovery) Assume A1– A3, ω > 4, and λi > 0. Conditional on (ui , ti ), let Fω be

KARIM ET AL.: ZK-Trace: Collusion-Secure Traitor Tracing with Zero-Knowledge Exculpation

5

√ the CDF of ω times the first coordinate of a uniform √ unit vector. For ui ̸= 0, put si = λi µ ω/∥ui ∥. Then

P[ĥi,b = wi,b | ui , ti ] ≥ Fω (si ), 8 . ω−4 (5) The proxy events in (4) are conditionally independent across bits. If ui = 0, every bit is correct. Fω (si ) ≥ Φ(si ) − bω ,

bω =

A deterministic bound ∥∆k ∥ ≤ Dk gives ∥ui ∥ ≤ P (1−λi )G+ k̸=i λk Dk by the triangle inequality. Random projection columns alone do not establish the smaller quadrature norm. For interpreting the measurements we retain the approximation √ µ ω SNRquad := p . (6) (Nc − 1)[(Nc − 1)G2 + nµ2 ] √ It assumes uniform weights, update norms near µ n, and negligible cross terms. It is a diagnostic model, not a proved floor for trained federated networks. Attribution by codeword separation. Maximum agreement is nearest-neighbor decoding in Hamming distance. A rival participates in training, so its codeword need not be independent of the decoded aggregate. The following criterion avoids that independence assumption. T HEOREM 2 (Attribution from a decoding radius) Let di = HD(ĥi , wi ) and dmin = mini̸=j HD(wi , wj ). If 2di < dmin for every client, all maximum-agreement decisions are uniquely correct. Under A4, if P[maxi di > rn] ≤ η for some 0 ≤ r < 1/4, then   Nc −2n(1/2−2r)2 P[any attribution error] ≤ η + e . (7) 2 The radius event may depend on the entire codebook. C OROLLARY 1 (A sufficient operating condition) Under A1–A4, suppose Fω (si ) ≥ p uniformly over clients and conditioning values. Choose ξ > 0 with r = 1 − p + ξ < 1/4. All clients are attributed correctly with probability at least 1 − δ if   Nc −2n(1/2−2r)2 −2nξ 2 Nc e + e ≤ δ. (8) 2 These are sufficient conditions, not necessary thresholds. A measured mean bit-accuracy cannot substitute for the uniform conditional p. The load ρ = Nc n/ω describes the number of identity constraints relative to scale dimension. ρ = 1 is a rank boundary, not a universal attribution-failure point (Lemma 5). Partial participation changes the current averaging weights. Marks from earlier rounds can remain in a skipped client’s absence. Presence additionally requires a specified absolute Hamming radius, calibrated in Proposition 4. Tracing soundness. For P a recovered tracing word y , biases pb , and score Si = b U (Xi,b , yb , pb ), an innocent row has independent Bernoulli entries conditional on (y, p) under A5(a). Its first two moments are unchanged by the output, although its full distribution depends on the output. 6

T HEOREM 3 (Conditional tracing bounds) Assume A5(a) and Definition 1. For any threshold z > 0 and α > 0, define √ Mb (α; y, p) = (1 − pb )e−α(2yb −1) pb /(1−pb ) √ + pb eα(2yb −1) (1−pb )/pb , (9) X Eα (y, p; z) = αz − log Mb (α; y, p). b

Then P[∃ innocent i : Si > z | y, p] ≤ min{1, N e−Eα (y,p;z) }. (10) A separate a-priori bound holds uniformly over (y, p) at s 2 BL BL + 2nt L, (11) + zB = 3 3 p (1 − δc )/δc , δc = 1/(300c), and L = where B = ln(N/ε1 ): the false-accusation probability is at most ε1 . √ The candidate threshold Z = 2nt L used in part of the evaluation is smaller than zB . It supports a certificate only when an evaluated exponent passes the required budget. The coalition sweep measures threshold exceedances, while the certificate-based evaluation checks the probability bound for each decoded word (Table VI). Appendix A-E specifies a certificate-gated rule, including separate budgets for positive and negative decisions. Finite completeness. Theorem 4 supplies a lowertail guarantee for the actual cutoff, without assuming independent intended coalition symbols. For a coalition of size s ≤ c, its exponential-moment bound is P[max Si ≤ z] ≤ etsz Js (t, q)nt , i∈C

t > 0.

(12)

The function Js integrates over the posterior secret bias given each coalition column, maximizing over permitted output bits. Conditioning on the entire coalition matrix allows the attacker to choose its intended word jointly across positions. Only the residual channel must flip these bits independently with probability q . Appendix CB proves the bound and gives a compatible a-priori soundness threshold. Interval integration establishes both guarantees at the two deployed lengths: for N = 10 and false-accusation budget 10−3 , the worst missed-coalition bounds over s ≤ c are 0.0096 at (c, nt , q) = (2, 512, 0.04) and 0.00201 at (3, 2048, 0.15) (Table XI). These are finite channel guarantees, not estimates of the networks’ residual channel. Per-colluder disagreement after averaging cannot identify q. For planning, the asymptotic length estimate is   2 π 2 −2 c (1 − 2q) ln(N/ε ) . (13) ndesign = 1 t 2 Its inverse-square channel cost [8], [15] describes a design trend. Equation (12) establishes the finite guarantee for the specified threshold and cutoff. Figure 2 distinguishes that trend from threshold certification. UNDER REVIEW

VOL. XX, No. XX

September 2026

(a) design cost

(b) thresholds 350 one-sided threshold

design length multiplier

6 5 4 3

300 250 200 150 100

2 1

candidate Bernstein

50 0.0 0.1 0.2 0.3 modeled residual rate q

512

2048 tracing length

4096

Fig. 2: Tracing design and certification are separate. (a) Inverse-square length multiplier for a hypothetical independent residual flip channel. (b) Candidate and Bernstein-certified one-sided thresholds versus tracing length, at c = 2, N = 10, and ε1 = 10−3 . The certified threshold does not assert completeness.

Feature stability under distillation. Let fT , fS use P the same feature coordinates and let ε2KD = −1 2 nt b ∥fS (pb ) − fT (pb )∥2 on the watermark probes. P ROPOSITION 2 (Empirical probe-margin stability) For the unit projections in Definition 2, suppose all but a fraction qbad of teacher probes have the correct signed margin at least µ > 0. The student’s disagreement with the embedded binary row satisfies   ε2 qKD ≤ min 1, qbad + KD . (14) µ2 This bound counts wrong and insufficient-margin teacher probes in qbad . A stronger score-level result handles arbitrary correlated errors: if at most k decoded bits change, then SiS ≥ SiT − Ai (k), p where Ai (k) sums the k largest weights 2|Xi,b − pb |/ pb (1 − pb ). Thus SiT − Ai (k) > z certifies threshold survival without a binary-symmetric channel (Theorem 5). Feature matching can preserve probe margins. Function-only matching, however, can change feature coordinates and violate the margin condition. Weight projections are not fixed by a feature-matching objective. Section VII-C reports the observed outcomes for both carriers. VI. EXPERIMENTAL SETUP

Datasets. We re-render the GNSS recordings of Heublein et al. [2] as 100,000 four-channel 4×32×32 antenna-array tensors. The GNSS benchmark contains six interference classes and excludes the interferencefree class. We use a sample-level 80/10/10 split (not a held-out-class split) into training, validation, and test sets. Class imbalance reaches approximately 178×. A Dirichlet partition with α=2.0 gives approximately 8,000 samples per client and moderate label skew, without guaranteeing

class coverage. CIFAR-10 (3×32×32) provides the transfer check. All ten stations use partitions of one measurement campaign. These partitions vary class availability and sample count, but do not reproduce separately sited receivers with distinct calibration, multipath, and local interference. The receiver-shift experiment is a proxy for these effects. Appendix D-A discusses the unmodeled variation and the need to recheck tracing certificates on decodes from new sites. Implementation. A ResNet-18 backbone [33], with its final fully connected layer removed, feeds a prototypical head [11]. We pre-train each dataset for 200 crossentropy epochs, then run 100 FedAvg rounds [1] with ten clients. Each client trains on 25 local five-way five-shot episodes per round, with 15 query examples per class. Both phases use stochastic gradient descent (SGD), with optimizer settings and seeds listed in Appendix D-A. Watermark configuration. Each client has n=128 identity bits. We evaluate the deployed Nc =10 setting (ρ≈0.27) and a higher-load setting with Nc =40 (ρ=1.07). The presence calibration uses pr =2−128 , giving errn =0 (Prop. 4). The identity and presence criteria are distinguished below. Credential proofs use r=331 parallel repetitions. On dispatched copies, the weight overlay uses nt =512 BN-γ projections with design target c=2. The feature carrier uses nt =2048 in-distribution probes with design target c=3. Metrics. We measure classification accuracy over 200 five-way five-shot test episodes. Each GNSS episode samples five of the six classes. On the aggregate, attribution accuracy is the fraction of clients identified by maximum codeword agreement. Self bit-accuracy measures agreement with the client’s own codeword, while crosstalk measures mean agreement with other codewords. Presence is the stricter Hamming-distance test in (3). On dispatched copies, single-leaker isolation requires naming the owner without accusing an innocent. Testing every recipient over ten seeds gives 100 trials at Nc =10 and 400 at Nc =40, per dataset. Collusion traceability requires accusing at least one colluder and no innocent, for sampled coalition sizes k ∈ {2, 3, 5, 8}. We also test offline tracing and resistance to codeword copying and evidence substitution (Sec. IV-E). Distillation counts pool ten runs per dataset, giving twenty runs per distiller class. Coalition sampling counts are in Appendix D-A. The certificate-based evaluation uses Nc = 10 weight-carrier models from ten CIFAR-10 and six GNSS seeds. For each seed, it tests all ten single copies and all 45 pair averages with the adjudication rule in Proposition 3. Feature geometry is measured by silhouette score, intra- and inter-class distance, and leave-one-out 1-NN accuracy. We report ten-seed means and standard deviations or pooled counts, as indicated. Matched watermark-off/on models differ only in λwm . Two one-sided tests (TOST) assess global-model equivalence at the stated margins (0.5 pp on CIFAR-10 and 1.0 pp on GNSS, App. D-F).

KARIM ET AL.: ZK-Trace: Collusion-Secure Traitor Tracing with Zero-Knowledge Exculpation

7

Attr. (%)

Self

Cross

Acc (%)

GNSS

98.0±4.0 10.0†

0.96±.02 –

0.50±.01 –

93.4±1.0 93.9±0.4

CIFAR-10

Ours FedZKP (group)

100.0±0.0 10.0†

0.94±.01 –

0.50±.01 –

84.9±0.3 84.7±0.2

Dataset Setting

p̂

Self

Cross

Attr.

Acc

CIFAR-10 (G=7.32)

Method Ours FedZKP (group)

Nc =5 0.13 0.97 Nc =10 0.27 0.83 Nc =20 0.53 0.68 Nc =40 1.07 0.59 Nc =80 2.13 0.55 Nc =120 3.20 0.53

1.00 0.94 0.86 0.76 0.66 0.63

0.50 0.50 0.50 0.50 0.50 0.50

1.00 1.00 1.00 1.00 0.88±0.04 0.61±0.05

0.85 0.85 0.85 0.86 0.86 0.86

GNSS (G=4.73)

Dataset

TABLE III: Predicted survival vs. ten-seed measurements under FedAvg (R=100, n=128, µ=1, ω=4800; means, with std shown where >0.01). ρ=Nc n/ω : identity load. p̂=Φ(SNRquad ): quadrature diagnostic (6). Self, Cross, Attr.: measured per-bit self-agreement, cross-talk, and attribution. Acc: few-shot accuracy. G=∥γ0 ∥: backbone baseline scale.

Nc =5 Nc =10 Nc =20 Nc =40

1.00 0.96 0.88 0.77

0.50 0.50 0.50 0.50

1.00 0.98±0.04 0.96±0.03 0.87±0.04

0.92 0.93 0.92 0.86

† Chance level 1/N (group mark, no per-client component). c

VII. RESULTS

We evaluate identity attribution on the aggregate, recipient tracing on dispatched copies, and robustness after attacks. The results distinguish the two carriers’ tracing performance and copy-accuracy costs from the utility of the deployed global model. A. Attribution and Survival

The maximum-agreement rule identifies 98% of contributors on GNSS and 100% on CIFAR-10 (Table II). Mean cross-talk is 0.50, consistent with the analysis of non-matching codewords. FedZKP’s group mark contains no per-client identifier, so its reference attribution rate is chance (1/Nc =10%). These results identify contributors 8

0.13 0.27 0.53 1.07

0.99 0.90 0.75 0.64

(b) per-bit accuracy vs. p̂

(a) attribution survival attribution accuracy

Attacks. Each attack is evaluated on 100 few-shot test episodes. We report task accuracy together with presence and attribution, so successful erasure can be distinguished from destruction of the classifier. Model-modification attacks include BN-scale pruning, Gaussian noise, quantization, and combined pruning and quantization. Targeted attacks use projected gradient descent (PGD) against a bit-flip objective or reset selected BN layers. Trainingbased attacks use fine-tuning without the watermark loss or knowledge distillation (KD) [26]. The latter includes feature matching [34], cross-architecture transfer, logitonly distillation, and feature isometry. Additional tests cover insider own-row erasure (negation, fresh re-randomization, and their per-carrier hybrid), robust aggregation [35] using the coordinate-wise median [36], and GNSS receiver covariate shift. Appendix DH gives the grids and budgets. Baselines. Five federated-watermarking methods use the same learning pipeline: FedZKP [3], FedTracker [4], DUW [5], FedIPR [9], and WAFFLE [18]. FedZKP supplies the group-ownership reference, while FedTracker and DUW supply per-client fingerprints. FedIPR uses BN scales, while WAFFLE uses its native trigger carrier. Each method is evaluated under a shared seven-attack suite using its own mark metric. We additionally test a DeepMarks-style balanced incomplete block design (BIBD) comparator [12] under the coalition sweep, because the five federated baselines lack collusion-secure codes.

ρ

1.0

1.0 per-bit accuracy

TABLE II: Per-client attribution on the aggregated model (Nc =10, n=128, ρ=0.27; 10-seed mean±std). Attr.: clients named correctly. Self, Cross: bit-agreement with own and with other codewords. Acc: few-shot accuracy.

0.8 0.6

ρ=1 0

0.8 0.6

cross-talk 0.5

2 0 identity load ρ = Nc n/ω CIFAR-10

GNSS

measured

2

diagnostic p̂

Fig. 3: Attribution versus identity load from Table III (n=128, ω=4800, R=100, 10-seed). (a) Attribution accuracy against ρ=Nc n/ω , error bars one std. Dotted: rank reference ρ=1. (b) Measured per-bit self-agreement (filled, solid) and the diagnostic p̂=Φ(SNRquad ) of (6) (open, dashed). Dotted: chance agreement 0.5.

to the shared aggregate. Recipient tracing is evaluated separately in Sec. VII-B. Survival under aggregation. Self-agreement decreases as the load ratio ρ grows, but remains above the quadrature diagnostic p̂ in every cell (Table III, Fig. 3). CIFAR-10 attribution remains exact through ρ=1.07 (Nc =40). GNSS attribution falls to 0.87 there despite similar self-agreement (0.77 versus 0.76). Thus the per-bit floor alone does not explain the dataset difference. In particular, GNSS’s smaller baseline scale G raises the quadrature diagnostic and cannot explain its earlier attribution decline. Mean cross-talk remains near 0.50, but this average does not determine the largest competing score. Appendix D-G discusses the bound’s limits near capacity. UNDER REVIEW

VOL. XX, No. XX

September 2026

Dataset

Nc

ρ

Isolated

qowner

Aggr. attr.

GNSS

10 40

0.27 100/100 1.07 400/400

0.12 0.05

0.98 0.87

CIFAR-10

10 40

0.27 100/100 1.07 400/400

0.15 0.06

1.00 1.00

Innocent baseline q≈0.29–0.31 at Nc =10 and 0.25–0.26 at Nc =40 (the minimum over the Nc −1 innocents is an order statistic that falls with Nc ). The owner sits far below it.

Checked attribution margins. On the Nc = 10 models used in the certificate-based evaluation, we also check the deterministic radius condition of Theorem 2 on the shared aggregates before dispatch. It holds for 100/100 CIFAR-10 and 59/60 GNSS client components, exactly the correctly attributed components in those sets. Thus each correct decision has a verified Hamming-separation margin, without invoking the idealized conditional sphere model. Heterogeneity and partial participation. Attribution is 100% when evaluated only over stations that participate in the non-IID and partial-participation sweeps. Across the full roster, sampling 30% of clients per round gives 98% attribution on GNSS and 100% on CIFAR-10. Strong label skew has a different effect: some clients lack the classes needed for a five-way episode and therefore cannot train a mark. At α=0.5, roster-wide attribution is 100% on CIFAR-10 and 79% on GNSS. GNSS falls to 16% at α=0.1. Appendix D-B reports both denominators.

B. Tracing, Collusion, and the Two Carriers

Isolating the leaked copy. The tracer decodes the dispatched copy’s overlay and compares it with the registered secret rows, without the holder’s cooperation. Every single-leaker trial identifies exactly the owner: 100/100 per dataset at Nc =10 and 400/400 at Nc =40 (Table IV). Isolation therefore persists beyond the identity-layer rank reference N∗ =37.5. Registry adjudication succeeds in 10/10 trials. The credential proof accepts the legitimate station and rejects every tested forgery, allowing a third party to check the evidence. Collusion tracing and false accusations. The weight overlay traces every tested two-client coalition, at its design target c=2. Success declines for larger coalitions (Table V, Fig. 4). The feature carrier is designed for c=3 and traces every sampled coalition through k=5 on GNSS and k=8 on CIFAR-10. These outcomes are empirical. The finite channel guarantees are given in Table XI. Neither carrier accuses an innocent station at any tested coalition size, including the Nc =40 experiments.

TABLE V: Collusion traceability under Boneh–Shaw copy averaging, both carriers, Nc =10, 10-seed pooled (GNSS first). Decisions use score thresholds without a certificate requirement. Traceability: fraction of sampled coalitions with ≥1 colluder accused and no innocent accused. Framed: fraction of coalitions in which any innocent is accused, at every k . Carrier

Dataset

k=2

k=3

k=5

k=8

Framed

Weight (c=2)§

GNSS CIFAR-10

1.00 1.00

0.95 0.75

0.40 0.35

0.15 0.00

0.000 0.000

Feature (c=3)

GNSS 1.00 1.00 1.00 0.76 CIFAR-10 1.00 1.00 1.00 1.00

0.000 0.000

§ Weight-carrier traceabilities use the implemented central-limit threshold. At the larger candidate threshold Z=97.1, the mid-coalition cells fall to 0.65/0.50 at k=3 and 0.10/0.05 at k=5 (GNSS/CIFAR-10). The c=2 design point (1.00) and the zero-framing column are invariant to the threshold.

(a) collusion

(b) margin

1.0 fraction of coalitions

TABLE IV: Single-leaker isolation of the dispatched copy (weight-space overlay), 10-seed. Isolated: trials whose owner the Tardos accusation names exactly with no innocent accused (Nc ×10 seeds). Aggr. attr.: the shared-model argmax attribution from Table III. qowner : the owner’s perbit flip rate.

Z

0.8

innocent (max)

0.6

86.8

DeepMarks framed

0.4 owner (min)

0.2 framed 0.000

0.0 2

3

5 coalition size k weight feature

8

191.4

0

CIFAR-10 GNSS

100 200 accusation score framed

Fig. 4: Collusion tracing and the accusation margin. (a) Fraction of sampled coalitions traced, and the fraction in which any innocent is framed, against coalition size k for both carriers and datasets (Table V). Color keys the carrier, marker and line style the dataset. Shaded band: the DeepMarks-BIBD framing rate, whose traceability is not plotted. (b) Largest measured innocent score and smallest owner score over the realized p decodes. Shaded band: the accusation threshold Z= 2nt ln(N/ε1 ) over the deployed federation sizes.

Certificate-based tracing. We evaluate Nc = 10 weight-carrier models over ten CIFAR-10 seeds and  six GNSS seeds, testing all ten single copies and all 10 2 = 45 pair averages per seed. The certificate-gated rule allocates 10−3 per investigation across accusation and tamper tails. All 880 positive-tail certificates pass. The adjudication rule isolates 160/160 single copies and traces 712/720 pair mixtures without naming an innocent (Table VI). The eight missed mixtures are CIFAR-10 pairs. This evaluation enumerates every pair within each seed. The

KARIM ET AL.: ZK-Trace: Collusion-Secure Traitor Tracing with Zero-Knowledge Exculpation

9

TABLE VI: Certificate-based weight-carrier tracing. Each entry counts decisions naming a colluder and no innocent. Positive-tail certificates pass for every tested decode. No innocent is named on either tail. Dataset

Seeds

Single copies

All pair averages

GNSS CIFAR-10

6 10

60/60 100/100

270/270 442/450

TABLE VII: Two-carrier comparison. Where two figures appear they are GNSS / CIFAR-10. Tracing lengths: 512 weight and 2048 feature. Max k : largest coalition traced at 100%. Copy cost: accuracy drop on the dispatched copy. Global cost: accuracy change on the deployed model. Property Collusion design target c Max k traced at 100% KD survival (feature-match) Copy-accuracy cost (pp) Global-accuracy cost (pp)

Weight-space Feature-space 2 2/2 0/10 ≈4.5 0

3 5/8 10/10 4.8 / 6.1 0

coalition sweep in Table V samples coalitions and uses threshold-only decisions. Each certificate establishes its numerical bound under A5(a), without requiring a binary-symmetric residual channel. Decoded arrays and interval enclosures accompany the results so that each bound can be checked independently. It does not empirically establish the independence premise. Unanimous-position disagreement averages 0.112 on GNSS pairs and 0.140 on CIFAR10 pairs. These are measured channel diagnostics, not estimates validating the modeled q values in Table XI. Carrier tradeoff. Both carriers encode the secret tracing row and use the same registry and credentialauthentication procedure (Table VII). The weight carrier targets two colluders and has a dispatched-copy accuracy cost of approximately 4.5 points in the reported 300episode, λt = 6 configuration. Distillation erases this carrier. The feature carrier targets three colluders and survives the tested feature-matching distillation runs, with a sufficient margin certificate given by Theorem 5. Its copy-accuracy cost is 4.8 points on GNSS and 6.1 on CIFAR-10. Function-only distillation erases this carrier too. Both tracing marks are added at dispatch, so neither changes the global model. C. Comparison, Utility, and Robustness

Table IX compares the five baselines using their native mark metrics. The capability columns denote cooperationfree recipient identification (Attribution), client-secret proof of knowledge (Credential), verification without disclosure of reusable secrets (ZK), distillation survival (KD), and coded tracing with a conditional falseaccusation bound (Coll.). ZK-Trace combines these capabilities subject to the carrier, channel, and per-decode qualifications above. 10

TABLE VIII: Feature-space metrics, unwatermarked (Base) vs. watermarked (WM) backbone (10-seed mean±std, GNSS first). 1-NN LOO: leave-one-out nearest-neighbor accuracy. GNSS Metric

Base

CIFAR-10 WM

Base

WM

Silhouette 0.605±.027 0.556±.040 0.340±.002 0.318±.004 1-NN LOO acc. 0.905±.010 0.911±.013 0.794±.002 0.797±.003 Intra-class dist. 3.36±.25 5.01±.39 7.31±.11 10.16±.18 Inter-class dist. 12.60±.53 16.35±.50 14.53±.17 18.50±.31

FedZKP and FedIPR verify ownership of a shared model rather than identify its recipient. WAFFLE’s groupmark metric is weak on both datasets (0.22 GNSS, 0.29 CIFAR-10). FedTracker and DUW achieve per-client traceability 1.0, but their fingerprints are server-assigned and lack binding to a client secret or zero-knowledge verification. Their distillation outcomes depend on method and dataset (App. D-D). DeepMarks-BIBD [12] addresses collusion under exact extraction, an assumption disrupted by FedAvg bit errors. It frames innocents in 50% of plain coalitions and 58% within its design resilience. ZK-Trace frames none in the same sweep (Fig. 4). Global-model utility. Dispatch fingerprinting does not modify the global model. For the identity watermark, paired ten-seed watermark-off/on comparisons establish equivalence at the stated margins on both datasets: ∆ = −0.04 pp on GNSS (margin 1.0 pp, TOST p = 0.0052) and +0.24 pp on CIFAR-10 (margin 0.5 pp, p = 0.0489). Appendix D-F gives confidence intervals and the matched-artifact protocol. Feature-space geometry. Table VIII evaluates the representation learned with the identity watermark. Leaveone-out 1-NN accuracy, a diagnostic of local class separation, shows no significant change on either dataset. Silhouette scores decrease by 0.048 on GNSS (p=0.009) and 0.023 on CIFAR-10 (p<10−4 ). Both intra- and inter-class distances increase. The observed pattern indicates altered cluster compactness alongside similar nearest-neighbor classification. It does not imply that watermarking leaves feature geometry unchanged. Removal attacks. Pruning, noise, quantization, PGD, and fine-tuning preserve weight-carrier attribution at its clean level in Table X. Layer reset reduces attribution, but also severely degrades task accuracy. Distillation is the tested attack that erases the weight mark while retaining approximately 80% task accuracy (Fig. 5). All tested weight-space baselines also lose their marks under the same 80-epoch distillation. Feature-carrier distillation. Feature-matching distillation preserves the feature mark in 20/20 runs, ten per dataset. Cross-architecture transfer from ResNet-18 to ResNet-34 preserves it in 19/20 runs. The single loss is GNSS seed 271, whose flip rate rises to 0.4976. Across seeds, the architecture change produces larger and more variable flip-rate increases on GNSS than on CIFAR10. Sparse four-channel inputs or subtle class differences UNDER REVIEW

VOL. XX, No. XX

September 2026

TABLE IX: Federated-watermarking baselines (Nc =10, 10-seed mean±std). Capability columns are defined in the text. ✓/× denote presence/absence, partial a qualified guarantee, and a parenthesis names the party the guarantee rests on. Cost: one-time verification payload. Acc: few-shot accuracy. Mark: each method’s native mark metric, DeepMarksBIBD being scored on its native Boolean tracing only, hence the empty cells. GNSS Method

Attribution Credential ZK

KD

Coll.

Cost

Ours FedZKP FedTracker DUW FedIPR WAFFLE

✓ × (group) ✓ ✓ partial ×

✓ ✓ × (server) × (server) partial ×

✓ ✓ × × × ×

✓∗ × × partial × partial

✓ × × × × ×

0.6 MB interactive 2.3 MB 1.5 MB 2.3 MB 0.3 MB

✓

× (server)

×

×

partial

–

DeepMarks-BIBD

Acc

CIFAR-10 Mark

0.934±.010 0.98±.04 0.939±.004 0.999±.001 0.926±.009 1.00±.00 0.922±.019 1.00±.00 0.925±.016 0.997±.002 0.933±.009 0.218±.126

–

–

Acc

Mark

0.849±.003 1.00±.00 0.847±.002 0.986±.003 0.838±.005 1.00±.00 0.825±.018 1.00±.00 0.849±.003 0.980±.003 0.846±.003 0.289±.035

–

–

∗ Feature carrier survives feature-matching KD (10/10 per dataset), but not function-only KD. All tested weight-space marks fail.

GNSS Attack None (clean) Structured prune 50% Noise σ=2.0 Quantize 2-bit PGD (ϵ=0.1) Reset layer3+4 Fine-tune 100 ep KD 80 ep

Self

Acc

0.96±.02 93.4±1.0 0.95±.02 31.2±2.8 0.78±.01 26.7±1.9 0.93±.02 29.4±5.0 0.95±.02 82.4±9.1 0.59±.02 41.9±4.6 0.96±.02 82.2±8.4 0.50±.01 80.4±10.8

CIFAR-10 Attr.

Self

Acc

Attr.

0.98±.04 0.98±.04 0.98±.04 0.98±.04 0.98±.04 0.63±.23 0.98±.04 0.15±.11

0.94±.01 0.94±.01 0.76±.01 0.92±.01 0.94±.01 0.56±.02 0.94±.01 0.51±.01

84.9±.3 52.0±3.6 26.0±.8 29.4±5.4 80.4±.8 26.3±2.2 79.1±.7 80.6±1.9

1.00±.00 1.00±.00 1.00±.00 1.00±.00 1.00±.00 0.48±.17 1.00±.00 0.15±.08

100 few-shot accuracy (%)

TABLE X: Robustness of attribution for the weight-space carrier under the parameter-space and training attacks (Nc =10, 100 test episodes, 10-seed mean±std, GNSS first). Self: self bit-accuracy (≈0.5 = codeword erased). Acc: few-shot accuracy (%). Attr.: attribution accuracy. Green/red: attribution preserved / broken.

80

(a) CIFAR-10 Clean

danger zone KD

PGD

danger zone KD

Feature

Fine-tune Prune

60

PGD Fine-tune Reset

40 20 0

Noise

Noise Quant

Reset

0.6

0.8

1.0

Prune Quant

0.6

0.8

1.0

self bit-accuracy (mark integrity) weight intact weight erased

may contribute, but these experiments do not isolate either cause. The baseline scale G governs the weightcarrier bound, not this feature-carrier failure (App. D-D). Function-only distillation erases the feature mark in all twenty runs. Survival therefore depends on the featurematching condition in Assumption 6. Own-row erasure and hybrid escape. The thresholdonly tests detect complementary single-strategy erasures (Prop. 3, Lem. 4). A hybrid attack defeats that complementarity by negating the weight overlay while independently re-randomizing the feature mark. At matched compute, it escapes the threshold-only tests on both carriers in 5/20 runs. The hybrid attack moves the weight score along a line with two decision thresholds. Scores above Z trigger accusation, while sufficiently negative scores trigger the reflected tamper test. Moderate negation moves the score into the interval between these regions, while stronger negation can push it through the interval and into tamper detection. At the matched 300-episode split, 15/20 copies remain above the positive threshold. All twenty copies escape at twice that budget, but all twenty trigger tamper detection at four times the budget. Under the stated custody model, the attacker cannot directly monitor the tracer-secret score, so the tested attack selects compute without observing its stopping point. These results characterize an unresolved evasion route. They do not establish

Feature

(b) GNSS Clean

feature copy integrity floor

feature after KD

Fig. 5: Accuracy–integrity tradeoff under the weightcarrier attacks of Table X: (a) CIFAR-10, (b) GNSS. xaxis: self bit-accuracy (≈0.5 = codeword erased). y -axis: few-shot accuracy. Shaded: high accuracy with the codeword at chance. Square: clean feature-carrier copy. Dashdot: its bit-accuracy after the same 80-epoch distillation.

robustness to an adaptively chosen budget. No innocent is flagged at any tested budget (App. A-E). Deployment stressors. Coordinate-wise median aggregation preserves attribution (1.00 CIFAR-10, 0.98 GNSS), although task accuracy declines. Under GNSS receiver covariate shift, tracing succeeds in all 100 trials and aggregate attribution remains 0.98, while exact singleleaker isolation degrades toward chance. Per-class copy cost. On the dispatched copies, mean per-class recall falls by 4.4 points on CIFAR-10 and 10.9 on GNSS, with larger losses on the harder interference classes. These values use the original label space rather than five-way episodes and therefore differ from the episode-accuracy costs in Table VII (App. D-E).

KARIM ET AL.: ZK-Trace: Collusion-Secure Traitor Tracing with Zero-Knowledge Exculpation

11

VIII. CONCLUSION

ZK-Trace combines recipient tracing, client-secret credentials, and independently checkable false-accusation certificates for federated GNSS models. The analysis provides finite completeness bounds for whole-codeword coalition strategies under a specified residual channel, together with deterministic score-stability bounds for correlated feature errors. At the two tracing lengths, verified finite examples give missed-coalition bounds below 0.0096 and 0.00201 with false-accusation budget 10−3 . These modeled-channel results complement the certificate-based evaluation: all 880 positive-tail certificates pass, all 160 single copies are isolated, and 712/720 pair mixtures are traced without naming an innocent. Carrier choice controls the robustness and utility tradeoff. Feature-space tracing survives the tested featurematching distillation runs, while function-only distillation and hybrid erasure remain evasion routes. Dispatch fingerprints leave the shared model unchanged. Matched controls establish global-model utility equivalence on both datasets at the stated margins. The remaining deployment work is to validate the conditional independence and probe-margin premises across independently sited GNSS stations and to address the demonstrated evasion routes.

[7]

[8]

[9]

[10]

[11]

[12]

[13]

Data and Code Availability

Code, per-seed artifacts, and configuration files covering both carriers, the Tardos code, the registry and verifier, the attack suite, and the table-reproduction scripts will be released upon publication. The GNSS tensors are re-rendered from the interference recordings released by their originators [2].

[14]

[15]

[16]

REFERENCES [1]

[2]

[3]

[4]

[5]

[6]

12

H. B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y Arcas, “Communication-Efficient Learning of Deep Networks from Decentralized Data,” in International Conference on Artificial Intelligence and Statistics, 2016. [Online]. Available: https://api.semanticscholar.org/CorpusID:14955348 L. Heublein, T. Feigl, T. Nowak, A. Rügamer, C. Mutschler, and F. Ott, “Evaluating ML Robustness in GNSS Interference Classification, Characterization & Localization,” in 2025 International Conference on Localization and GNSS (ICL-GNSS), 2025, pp. 1–7. W. Yang, Y. Yin, G. Zhu, H. Gu, L. Fan, X. Cao, and Q. Yang, “FedZKP: Federated Model Ownership Verification with Zero-knowledge Proof,” 2023. [Online]. Available: https://arxiv.org/abs/2305.04507 S. Shao, W. Yang, H. Gu, Z. Qin, L. Fan, Q. Yang, and K. Ren, “FedTracker: Furnishing Ownership Verification and Traceability for Federated Learning Model,” IEEE Transactions on Dependable and Secure Computing, vol. 22, no. 1, pp. 114–131, 2025. S. Yu, J. Hong, Y. Zeng, F. Wang, R. Jia, and J. Zhou, “Who Leaked the Model? Tracking IP Infringers in Accountable Federated Learning,” 2023. [Online]. Available: https://arxiv. org/abs/2312.03205 D. Boneh and J. Shaw, “Collusion-secure fingerprinting for digital data,” IEEE Transactions on Information Theory, vol. 44, no. 5, pp. 1897–1905, 1998.

[17]

[18]

[19] [20]

[21]

[22]

[23]

G. Tardos, “Optimal probabilistic fingerprint codes,” J. ACM, vol. 55, no. 2, May 2008. [Online]. Available: https: //doi.org/10.1145/1346330.1346335 B. Škorić, S. Katzenbeisser, and M. U. Celik, “Symmetric Tardos fingerprinting codes for arbitrary alphabet sizes,” Designs, Codes and Cryptography, vol. 46, no. 2, pp. 137– 166, Feb. 2008. [Online]. Available: https://doi.org/10.1007/ s10623-007-9142-x B. Li, L. Fan, H. Gu, J. Li, and Q. Yang, “FedIPR: Ownership Verification for Federated Deep Neural Network Models,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 4, pp. 4521–4536, 2023. N. S. Gaikwad, L. Heublein, N. L. Raichur, T. Feigl, C. Mutschler, and F. Ott, “Federated Learning with MMD-based Early Stopping for Adaptive GNSS Interference Classification,” in NOMS 2025-2025 IEEE Network Operations and Management Symposium, 2025, pp. 01–10. J. Snell, K. Swersky, and R. S. Zemel, “Prototypical Networks for Few-shot Learning,” 2017. [Online]. Available: https: //arxiv.org/abs/1703.05175 H. Chen, B. D. Rouhani, C. Fu, J. Zhao, and F. Koushanfar, “DeepMarks: A Secure Fingerprinting Framework for Digital Rights Management of Deep Learning Models,” in Proceedings of the 2019 on International Conference on Multimedia Retrieval, ser. ICMR ’19. New York, NY, USA: Association for Computing Machinery, 2019, p. 105–113. [Online]. Available: https://doi.org/10.1145/3323873.3325042 J.-J. Oosterwijk, B. Škoric, and J. Doumen, “Optimal suspicion functions for Tardos traitor tracing schemes,” in Proceedings of the First ACM Workshop on Information Hiding and Multimedia Security, ser. IH&MMSec ’13. New York, NY, USA: Association for Computing Machinery, 2013, p. 19–28. [Online]. Available: https://doi.org/10.1145/2482513.2482527 B. Skoric, T. U. Vladimirova, M. Celik, and J. C. Talstra, “Tardos Fingerprinting is Better Than We Thought,” IEEE Transactions on Information Theory, vol. 54, no. 8, pp. 3663–3676, 2008. T. Laarhoven and B. de Weger, “Optimal symmetric Tardos traitor tracing schemes,” Designs, Codes and Cryptography, vol. 71, no. 1, pp. 83–103, Apr. 2014. [Online]. Available: https://doi.org/10.1007/s10623-012-9718-y B. Škorić and J.-J. Oosterwijk, “Binary and q-ary Tardos codes, revisited,” Designs, Codes and Cryptography, vol. 74, no. 1, pp. 75–111, Jan. 2015. [Online]. Available: https: //doi.org/10.1007/s10623-013-9842-3 M. Kuribayashi, “Tardos’s Fingerprinting Code over AWGN Channel,” in Information Hiding, R. Böhme, P. W. L. Fong, and R. Safavi-Naini, Eds. Berlin, Heidelberg: Springer Berlin Heidelberg, 2010, pp. 103–117. B. G. A. Tekgul, Y. Xia, S. Marchal, and N. Asokan, “WAFFLE: Watermarking in Federated Learning,” in 2021 40th International Symposium on Reliable Distributed Systems (SRDS), 2021, pp. 310–320. I. Damgård, “On Σ-Protocols,” Lecture Notes, Department of Computer Science, Aarhus University, 2010. A. Jain, S. Krenn, K. Pietrzak, and A. Tentes, “Commitments and Efficient Zero-Knowledge Proofs from Learning Parity with Noise,” in Advances in Cryptology – ASIACRYPT 2012, X. Wang and K. Sako, Eds. Berlin, Heidelberg: Springer Berlin Heidelberg, 2012, pp. 663–680. P. Véron, “Improved identification schemes based on errorcorrecting codes,” Applicable Algebra in Engineering, Communication and Computing, vol. 8, no. 1, pp. 57–69, Jan. 1997. [Online]. Available: https://doi.org/10.1007/s002000050053 D. İşler, E. van Kempen, S. Hwang, and N. Laoutaris, “FedPoP: Federated Learning Meets Proof of Participation,” 2025. [Online]. Available: https://arxiv.org/abs/2511.08207 H. Liu, J. Wei, Z. Xu et al., “Authentication and Traceability for Federated Learning Models via Group Signatures,” Research

UNDER REVIEW

VOL. XX, No. XX

September 2026

[24]

[25]

[26]

[27]

[28]

[29]

[30]

[31]

[32]

[33]

[34]

[35]

[36]

[37]

[38]

Square, Sep. 2024, preprint, Version 1. [Online]. Available: https://doi.org/10.21203/rs.3.rs-4867383/v1 E. Rodrı́guez-Lois, F. Brau, M. Pintor, B. Biggio, and F. PérezGonzález, “BlackCATT: Black-box Collusion Aware Traitor Tracing in Federated Learning,” 2026. [Online]. Available: https://arxiv.org/abs/2602.12138 M. Shafieinejad, N. Lukas, J. Wang, X. Li, and F. Kerschbaum, “On the Robustness of Backdoor-based Watermarking in Deep Neural Networks,” in Proceedings of the 2021 ACM Workshop on Information Hiding and Multimedia Security, ser. IH&MMSec ’21. New York, NY, USA: Association for Computing Machinery, 2021, p. 177–188. [Online]. Available: https://doi.org/10.1145/3437880.3460401 G. Hinton, O. Vinyals, and J. Dean, “Distilling the Knowledge in a Neural Network,” 2015. [Online]. Available: https: //arxiv.org/abs/1503.02531 N. Lukas, E. Jiang, X. Li, and F. Kerschbaum, “SoK: How Robust is Image Classification Deep Neural Network Watermarking?” in 2022 IEEE Symposium on Security and Privacy (SP), 2022, pp. 787–804. S. Szyller, B. G. Atli, S. Marchal, and N. Asokan, “DAWN: Dynamic Adversarial Watermarking of Neural Networks,” in Proceedings of the 29th ACM International Conference on Multimedia, ser. MM ’21. New York, NY, USA: Association for Computing Machinery, 2021, p. 4417–4425. [Online]. Available: https://doi.org/10.1145/3474085.3475591 H. Jia, C. A. Choquette-Choo, V. Chandrasekaran, and N. Papernot, “Entangled Watermarks as a Defense against Model Extraction,” 2021. [Online]. Available: https://arxiv.org/abs/2002.12200 A. Esser and E. Bellini, “Syndrome Decoding Estimator,” in Public-Key Cryptography – PKC 2022, G. Hanaoka, J. Shikata, and Y. Watanabe, Eds. Cham: Springer International Publishing, 2022, pp. 112–141. D. Naor, M. Naor, and J. Lotspiech, “Revocation and Tracing Schemes for Stateless Receivers,” in Advances in Cryptology — CRYPTO 2001, J. Kilian, Ed. Berlin, Heidelberg: Springer Berlin Heidelberg, 2001, pp. 41–62. T. Attema, S. Fehr, and M. Klooß, “Fiat-Shamir Transformation of Multi-round Interactive Proofs,” in Theory of Cryptography, E. Kiltz and V. Vaikuntanathan, Eds. Cham: Springer Nature Switzerland, 2022, pp. 113–142. K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 770–778. A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, and Y. Bengio, “FitNets: Hints for Thin Deep Nets,” 2015. [Online]. Available: https://arxiv.org/abs/1412.6550 P. Blanchard, E. M. El Mhamdi, R. Guerraoui, and J. Stainer, “Machine Learning with Adversaries: Byzantine Tolerant Gradient Descent,” in Advances in Neural Information Processing Systems, I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, Eds., vol. 30. Curran Associates, Inc., 2017. [Online]. Available: https://proceedings.neurips.cc/paper files/ paper/2017/file/f4b9ec30ad9f68f89b29639786cb62ef-Paper.pdf D. Yin, Y. Chen, R. Kannan, and P. Bartlett, “Byzantine-Robust Distributed Learning: Towards Optimal Statistical Rates,” in Proceedings of the 35th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, J. Dy and A. Krause, Eds., vol. 80. PMLR, 10–15 Jul 2018, pp. 5650–5659. [Online]. Available: https: //proceedings.mlr.press/v80/yin18a.html K. Nuida, S. Fujitsu, M. Hagiwara, T. Kitagawa, H. Watanabe, K. Ogawa, and H. Imai, “An improvement of discrete Tardos fingerprinting codes,” Designs, Codes and Cryptography, vol. 52, no. 3, pp. 339–362, Sep. 2009. [Online]. Available: https://doi.org/10.1007/s10623-009-9285-z P. Meerwald and T. Furon, “Toward Practical Joint Decoding of Binary Tardos Fingerprinting Codes,” IEEE Transactions on

[39]

[40]

[41]

[42]

Information Forensics and Security, vol. 7, no. 4, pp. 1168– 1180, 2012. H. D. Hollmann, J. H. van Lint, J.-P. Linnartz, and L. M. Tolhuizen, “On Codes with the Identifiable Parent Property,” Journal of Combinatorial Theory, Series A, vol. 82, no. 2, pp. 121–133, 1998. [Online]. Available: https://www.sciencedirect. com/science/article/pii/S009731659792851X R. Vershynin, High-Dimensional Probability: An Introduction with Applications in Data Science, ser. Cambridge Series in Statistical and Probabilistic Mathematics. Cambridge University Press, 2018. P. Diaconis and D. Freedman, “A dozen de Finetti-style results in search of a theory,” in Annales de l’IHP Probabilités et statistiques, vol. 23, no. S2, 1987, pp. 397–423. T. Attema, S. Fehr, M. Klooß, and N. Resch, “The fiat–shamir transformation of (γ1 , . . . , γµ )-special-sound interactive proofs,” Cryptology ePrint Archive, Paper 2023/1945, 2023. [Online]. Available: https://eprint.iacr.org/2023/1945

KARIM ET AL.: ZK-Trace: Collusion-Secure Traitor Tracing with Zero-Knowledge Exculpation

13

Appendix A Tracing Construction and Analysis

The construction separates credential authentication, identity decoding, and recipient tracing. The bounds below specify their different probability spaces: identity code generation, an innocent row conditional on the recovered artifact, and a coalition channel with hidden biases. Experiments use fixed reproducible code seeds. Their success counts are distinct from these probability bounds.

separately from the shared model. The tracing operation does not change that shared model. Table VII reports copy costs for the evaluated configurations, not a universal cost of embedding. The codebook and tracing overlay have different purposes. The global model contains multiple identity marks, so those marks cannot isolate a recipient. The dispatched copy contains the recipient’s tracing row, allowing offline comparison with the enrollment registry. Neither layer detects poisoned updates or prevents voluntary credential sharing. These threats require separate controls, such as robust aggregation and revocation [31], [36].

A. Identity Codebook and Tardos Overlay

The identity codewords wi occupy separate column blocks Ei . Every block is a set of dense directions in the same ω -dimensional scale vector. They do not occupy disjoint physical coordinates. An additional shared block ET carries the dispatched recipient’s tracing row. The identity layer detects agreement with public credentials. It does not establish who trained the model. D EFINITION 1 (Tardos overlay on a shared block) Fix a design target c ≥ 2, tracing length nt , and δc = 1/(300c). Draw biases independently with density 1 fδc (p) = √ p (π − 4 arcsin δc ) p(1 − p) on [δc , 1 − δc ]. Conditional on the biases, draw all entries Xi,b ∼ Bernoulli(pb ) independently. At dispatch, embed row Xi in recipient i’s copy using the shared unit columns ET,b . Decode yb = 1{⟨γpir , ET,b ⟩ > 0} and score nt X (2y − 1)(x − p) U (Xi,b , yb , pb ), U (x, y, p) = p . Si = p(1 − p) b=1 (15) The score is the symmetric Tardos score [7], [8]. It satisfies U (x, 1 − y, p) = −U (x, y, p), p (16) |U | ≤ B := (1 − δc )/δc . The carrier uses nt shared columns rather than N nt clientspecific columns. The full code matrix and biases are held by the tracer. Clients can estimate their own rows from their copies and may receive their rows in the insiderattack model. No secrecy from one’s own recipient is needed for the innocent-row analysis.

C. Finite Tracing Bounds and the Residual Channel

L EMMA 1 (Innocent-score moments) Under A5(a), for an innocent i and any realized (y, p), E[Si | y, p] = 0,

Var(Si | y, p) = nt .

(18)

The individual increments are independent and bounded by B . Their conditional distribution and upper tails may depend on y . L EMMA 2 (Flip-scaled moments) Let Tb = P ∗ U (X , y , p ) and let F ∼ Bernoulli(q) be i,b b b b i∈C independent of Tb . Then Tb′ = (1 − 2Fb )Tb obeys E[Tb′ ] = (1 − 2q)E[Tb ], E[(Tb′ )2 ] = E[Tb2 ], Var(Tb′ ) = Var(Tb ) + 4q(1 − q)(E[Tb ])2 .

(19)

Here q denotes disagreement with the intended coalition output y ∗ , not with each colluder’s row. On positions where colluders disagree, a noiseless marking-consistent output already differs from some rows. Consequently, the reported per-row disagreement rates do not identify a residual BSC parameter. L EMMA 3 (A deterministic weight-carrier P perturbation bound) For a fixed noiseless mixture γ ∗ = i∈C λi γ (i) and perturbation η , let ab = ⟨γ ∗ , ET,b ⟩. Any position with |ab | > ∥η∥2 keeps its sign after adding η .

B. Dispatch Fingerprinting

The unit-column Cauchy–Schwarz bound proves this statement without a Gaussian or independence assumption. Equal and opposite colluder margins may cancel, so local margins do not guarantee a positive pooled margin on every bit. Because averaging neural-network weights does not imply averaging feature responses, the featurecarrier analysis uses probe margins directly.

Federated training embeds the public identity layer. The tracer subsequently fine-tunes a separate copy for each recipient on a server-held proxy pool, with λt X Li = Ltask + max{0, µ − (2Xi,b − 1)⟨γ, ET,b ⟩}. nt b (17) The reported weight configuration uses roughly 300 episodes and λt = 6. The task term preserves classifier utility during embedding. The copy is then evaluated

R EMARK 1 (A certificate is an evaluated inequality) Theorem 3 gives a conditional bound for each (y, p) at any positive α. The checker selects an α numerically and encloses the resulting exponent using interval arithmetic. A passing certificate requires its lower endpoint to exceed the upper endpoint of ln(N/ε) for the allocated tail budget ε. This avoids assuming that an optimizer has found the exact supremum. An uncertified threshold exceedance remains an investigative lead. A valid certificate does not

14

UNDER REVIEW

VOL. XX, No. XX

September 2026

TABLE XI: Finite guarantees for the specified residual channels, N = 10, ε1 = 10−3 , δc = 1/(300c). The completeness bound covers every coalition size 1 ≤ s ≤ c. Displayed thresholds are rounded. The calculation uses the full-precision values in the accompanying artifact. c

nt

q

z

Upper bound on ε2

2 3

512 2048

0.04 0.15

111.0004 214.9809

0.0096 0.00201

prove that the innocent-row independence premise holds in a deployment. The finite completeness bound integrates over the hidden biases before maximizing a coalition’s allowed output. This order is essential: a pirate knows its rows but not the biases. T HEOREM 4 (Finite completeness against row-dependent marking strategies) Fix a coalition of s ≤ c rows from Definition 1. It may choose its full intended word as any randomized function of those rows and side information independent of the biases conditional on the rows. At unanimous positions it must output the common bit. Apply independent residual BSC flips with common rate q . For k ∈ {0, . . . , s} set k − sp ak (p) = p , Y0 = {0}, Ys = {1}, p(1 − p) Yk = {0, 1} (0 < k < s). For t > 0 define Z 1−δc s   X s Js (t, q) = max fδc (p)pk (1 − p)s−k k v∈Yk δc k=0  · (1 − q)e−t(2v−1)ak (p) + qet(2v−1)ak (p) dp. (20) At a fixed threshold z , the probability of missing every colluder satisfies P[max Si ≤ z] ≤ min{1, etsz Js (t, q)nt }. i∈C

(21)

This holds even when the intended symbols are dependent across positions. The bound can be checked for every s ≤ c and optimized over t. It is finite at the implemented cutoff and specifies its own completeness level. For a compatible fixed threshold, define Z 1−δc I(a) = fδc (p) max M (a; v, p) dp. (22) v∈{0,1}

δc −az

nt

Under A5(a), N e I(a) ≤ ε1 suffices for a-priori soundness. This bound averages over independently generated biases. Theorem 3, by comparison, conditions on the realized biases. Neither bound requires the innocent score distribution to be invariant under changes of output. Verified finite examples. Table XI evaluates both inequalities using 4096 interval panels in the arcsine angle coordinate and 30-decimal-digit interval arithmetic. Numerical optimization selects candidate values of (a, t).

Interval integration then encloses the integrals from above to establish the displayed bounds. All s = 1, . . . , c are checked. The residual rates are stipulated model parameters, not fitted neural extraction rates. The script and fullprecision enclosures are supplied with the reproducibility artifacts. Relation to the design equation. The common planning rule   = dc2 (1 − 2q)−2 ln(N/ε1 ) , d = π 2 /2, ndesign t (23) is an asymptotic estimate [8]. Laarhoven–de Weger [15] obtain a finite constant 23.79 jointly with threshold coefficient 8.06 and cutoff coefficient 28.31. Those constants jointly specify a different code configuration. Equations (20)–(22) give soundness and completeness bounds for cutoff coefficient 300 and the thresholds used here. Alternative codes and decoders offer other tradeoffs [13], [37]–[39]. D. Registry, Authentication, and Disputes

The registry binds the identity and tracing assignments before a dispute: Ci = Com(wi ∥Xi ∥Ai ∥yi ; ρi ).

(24)

The tracer creates the enrollment tuple, and the judge holds its opening. Tracing needs no response from the suspect. A judge checks the opening, registered decoder setup, and score evidence. A commitment prevents substitution of a different row, but a tracer that knows the original row can embed it again. As in the evaluated protocol, provenance therefore assumes an honest tracer. A signed delivery receipt and authenticated setup would be needed to prove issuance independently. A client nonce alone does not prove delivery. The credential proof authenticates a claimant in a dispute. Both an innocent recipient and a genuine leaker can possess a valid witness, so witness knowledge alone is not exculpatory evidence. Exculpation requires rejection of an invalid accusation, such as a mismatched enrollment opening or a failed tracing certificate. A credential-only authentication proof can omit model presence. Algorithm 1 checks both credential knowledge and presence. The implementation’s registry helper checks the opening. A deployment’s judge must independently reconstruct the score and setup from the evidence package. E. Certified Decisions and Residual Erasure

C OROLLARY 2 (Budgeted two-tail decisions) For K declared carriers and per-investigation budget ε, allocate ε/(2K) to each tail. For each carrier, certify the positive tail using y and the negative tail using 1 − y . Under A5(a), a union bound gives probability at most ε that any certified decision names an innocent, across all K carriers and both tails. Independence between carriers is unnecessary.

KARIM ET AL.: ZK-Trace: Collusion-Secure Traitor Tracing with Zero-Knowledge Exculpation

15

TABLE XII: Custody and scope of the guarantees. Artifact

Custody

Consequence of disclosure

xLPN witness

client

Recipient row

tracer; recipient can estimate it tracer

credential impersonation own-row removal is possible innocent-row independence may fail targeted carrier edits are possible integrity still depends on binding

Full code and biases Projection directions Registry opening

public judge/tracer

P ROPOSITION 3 (Certificate-gated adjudication) Use a p candidate threshold z = 2nt ln(2KN/ε) on each tail. Return CERTIFIED - ATTRIBUTE if a positive score exceeds z and its tail certificate passes. Otherwise, return CERTIFIED - TAMPER for a negative score below −z with a passing reflected certificate. A threshold exceedance without its certificate is an UNCERTIFIED - LEAD. If neither threshold is exceeded, return NO - CERTIFIED - EVIDENCE. Positive decisions take precedence, with the largest positive score selected. Negative decisions select the smallest score. The probability bound is that of Corollary 2. The lead carries no calibrated accusation guarantee. No-certified-evidence does not establish that a copy was never leaked. Repeated investigations need a separately allocated total budget. Adaptive feedback also requires maintaining the conditional independence premise. The attack sweeps measure threshold exceedances. The certificate-based evaluation additionally requires the corresponding tail bound to pass. L EMMA 4 (Perfect p the score) For Y = P negation reverses 1 − Xi , Si = − b |Xi,b − pb |/ pb (1 − pb )p P . Its expectation over row i, conditional on p, is −2 b pb (1 − pb ). A negative verdict still requires its threshold and certificate. The experiments show different erasure responses on the two carriers. Negating the weight mark can leave a positive score below its threshold, whereas negating feature marks produces large negative scores. Rerandomizing feature signs can bring their scores near zero. These are measured attack outcomes, not a guarantee that every negative-training objective reaches perfect inversion. The escape window, geometrically. In the hybrid sweep, the weight score determines the observed decision because the feature score remains inside its non-triggering interval. At the matched 300-episode split, 15/20 weight scores remain above the accusation threshold. At 600 episodes, all twenty lie between the positive and negative thresholds. At 1200, all twenty cross the negative threshold. Moderate negation moves the score into this interval, while stronger negation can move it through the interval. The corresponding score ranges are 82–155, −6–73, and 16

−236–−163, respectively. No innocent is flagged in the sweep. The attacker does not observe the biases or exact score in this experiment, but may estimate useful attack budgets by other means. The sweep does not establish resistance to adaptive budget selection. F. The Feature-Space Carrier

D EFINITION 2 (Feature-space margin carrier) For a backbone fθ : X → RD , fixed probes pb , and unit directions vb ∈ SD−1 , define mb (θ) = ⟨vb , fθ (pb )⟩,

yb = 1{mb (θ) > 0}.

(25)

Embed row Xi by minimizing λt X Lfi = Ltask + max{0, µ−(2Xi,b −1)mb (θ)}. (26) nt b

Here D = 512, smaller than ω = 4800. The ability to fit many probe constraints comes from training a nonlinear function on different inputs, not from feature width exceeding the scale dimension. The experiments demonstrate that the carrier can fit 2048 or 4096 probes. R EMARK 2 (Carrier-independent scoring) The conditional soundness bound uses only the decoded bits, biases, and innocent-row independence, so it applies to both readouts. Completeness additionally depends on the attack and channel. Inverting the asymptotic design equation gives the planning index p (27) cplan = (1 − 2q) nt /(d ln(N/ε1 )). It is not a finite certified coalition size. Theorem 4 provides a finite channel analysis instead. Feature matching optimizes min Ex∼D ∥fS (x) − fT (x)∥22 .

(28)

θS

Its relation to a finite probe set requires control on that set, not merely a small training loss. Proposition 2 is stated directly for the probe residual and counts all incorrect or insufficient-margin teacher probes. A stronger result connects feature stability directly to the weighted tracing score, without independent flips. T HEOREM 5 (Tracing-score stability under arbitrary probe p errors) For client i, define ai,b = 2|Xi,b − pb |/ pb (1 − pb ). If teacher and student decodes differ on a set J , then X |SiS − SiT | ≤ ai,b . (29) b∈J

If at most k bits differ, the right side is bounded by the sum Ai (k) of the k largest ai,b . Consequently SiT − Ai (k) > z guarantees that client i still exceeds threshold z , regardless of error dependence. For feature readouts, let rb = ∥fS (pb ) − fT (pb )∥2 and Jµ = {b : |mb (θT )| < µ}. One may take k = min{nt , |Jµ | + ⌊nt ε2KD /µ2 ⌋}.

(30)

Alternatively the directly measured set {b : rb ≥ |mb (θT )|} supplies a sharper bound. UNDER REVIEW

VOL. XX, No. XX

September 2026

This yields a sufficient tracing-survival certificate for correlated or targeted errors. It certifies threshold survival, while Theorem 3 separately certifies the falseaccusation risk of the resulting decode. When the bound is too loose, direct score evaluation is still possible. Function-only distillation can rotate feature coordinates, so it need not obey a small probe-residual condition. The observed 20/20 feature-matching and 19/20 crossarchitecture outcomes, and 0/20 function-only outcomes, remain empirical evidence rather than substitutes for a margin check. Carrier tradeoff. The weight carrier is inexpensive to read and has a reported episodic copy cost of about 4.5 points in the 300-episode, λt = 6 configuration, but distillation removes it. The feature carrier costs 4.8 and 6.1 accuracy points per copy on GNSS and CIFAR10, and supports the longer tracing rows used in the experiments. The design targets are two and three colluders. Measured tracing extends beyond those targets. The score-stability theorem explains a sufficient mechanism for survival without claiming that feature matching always preserves the mark. Appendix B Assumptions and Supporting Lemmas A. Assumptions

A SSUMPTION 1 (Unconditioned codebook geometry) Identity columns are independent uniform directions on Sω−1 . The implementation normalizes one Gaussian draw without a global rejection test. A SSUMPTION 2 (Achieved local margins) At the round analyzed, each participating client satisfies ti,b ⟨γ (i) , Ei,b ⟩ ≥ µ for every bit. A hinge objective alone does not establish this premise, so the achieved margins must be checked or treated as idealized. A SSUMPTION 3 (Conditional signed-direction model) Conditional on (ui , ti ), the signed directions {ti,b Ei,b }b remain independent uniform sphere directions. This is an idealized single-round model. Repeated federated training is not asserted to satisfy it. A SSUMPTION 4 (Identity codeword model) Identity codewords are independent uniform binary vectors. They need not be independent of the trained decoded aggregate. A SSUMPTION 5 (Tracing probability models) (a) Conditional on the recovered word and biases, each innocent row retains independent Bernoulli(pb ) entries. This holds when the artifact and its selection expose no information about that row beyond the biases. (b) For the finite completeness theorem only, the coalition’s side information is independent of biases conditional on its own rows, its intended word satisfies marking, and residual flips are independent of all rows and biases and mutually independent with common rate q . Part (a) does not require part (b).

A SSUMPTION 6 (Comparable feature coordinates) Teacher and student features use the same coordinates and unit projection vectors. The residual is evaluated on the actual fixed probes. Inferring that residual from a population or training objective needs a separate generalization argument. A3 is not derived from the fact that clients train on separate blocks: earlier aggregates carry other clients’ marks. Similarly, uniform model averaging does not prove the BSC premise in A5(b). Innocent-score soundness, finite BSC completeness, and deterministic score stability have different assumptions and should be applied separately. B. Projection and Geometry

The match score is M (i, j) = 1 − HD(ĥi , wj )/n.

(31)

The projection decomposition is exact: ⟨γagg , Ei,b ⟩ = λi ⟨γ (i) , Ei,b ⟩ + ⟨ui , Ei,b ⟩.

For γ (k) = γ0 + ∆k , the norm inequality is X ∥ui ∥ ≤ (1 − λi )G + λk ∥∆k ∥.

(32)

(33)

k̸=i

The quadrature approximation (6) drops baseline-update and update-update cross terms. It is not implied by A1 or √ by a Gram norm below 2 ρ + ρ. L EMMA 5 (Rank and random-code separation) If Nc n > ω , the identity Gram matrix is singular and ∥E ⊤ E − I∥op ≥ 1. This does not preclude useful sign decoding. Under A4, for 0 < ζ < 1/2,   Nc −2ζ 2 n P[∃i < j : HD(wi , wj ) ≤ (1/2 − ζ)n] ≤ e . 2 (34) For instance, E = [I I] has load two and Gram deviation one, and passes the inequality with right side √ 2 ρ + ρ. Thus N ∗ = ω/n = 37.5 is a dimensional reference, not an impossibility theorem. For unit-normalized Gaussian columns the limiting nonzero singular-value √ √ edges use 1 ± ρ, without another division by ω [40]. The empirical failure point also depends on codeword separation and decoding errors. Appendix C Proofs A. Identity Recovery and Attribution

L EMMA 6 (Sphere projection) For X uniform on Sω−1 √ and fixed v ̸= 0, the normalized projection ω⟨X, v⟩/∥v∥ is symmetric with variance one and, for ω > 4, sup |Fω (a) − Φ(a)| ≤ 8/(ω − 4).

(35)

a

Proof:

KARIM ET AL.: ZK-Trace: Collusion-Secure Traitor Tracing with Zero-Knowledge Exculpation

17

Rotation reduces the projection to the first coordinate. Symmetry follows by reflection. The variance is 1/ω before normalization because the squared coordinates sum to one and have equal expectations. The first-coordinate normal approximation follows from the finite-sphere bound of Diaconis–Freedman [41]. For v = 0 the unnormalized projection is identically zero and normalization is unnecessary. Proof of Theorem 1: A2 and (32) imply (4). Conditional on (ui , ti ), A3 makes the signed noise projection a fixed vector projected onto an independent random unit direction. Its symmetry gives proxy success probability Fω (si ). The true success event contains the proxy event, so it has at least that probability. Different proxy bits use independent signed directions. Lemma 6 supplies the normal approximation. If ui = 0, the local positive margin survives multiplication by λi > 0.

For completeness at least .95, exact enumeration gives minimum lengths 202 at p = .95 (t = 15) and 233 at p = .93 (t = 23). Proof: Count the binary vectors in the Hamming ball for the false-accept law. Under the separate i.i.d. error model the true Hamming distance is binomial. At n = 128, radius zero contains one vector and radius one contains 129, proving the exact calibration. Evaluating the two binomial tails jointly gives the stated example lengths. At n = 128, perfect extraction has completeness one, while p = .9999 gives .9873. Dependent decoded bits require their own completeness analysis. The false-accept calculation concerns an independent codeword, not deliberate copying of a public mark. B. Tracing and Feature-Stability Proofs

Proof of Theorem 2: For any rival j , the triangle inequality gives HD(ĥi , wj ) ≥ HD(wi , wj ) − di . Thus 2di < dmin makes every rival farther away than the true codeword. For the probabilistic statement, an error requires either maxi di > rn or a pair of codewords at distance at most 2rn. Under A4, each pair distance is Binomial(n, 1/2), so Hoeffding’s inequality bounds the latter event by the second term of (7). The union bound requires no independence between decoding errors and codewords.

Proof of Lemma 1 and Theorem 3: For X ∼ Bernoulli(p) independent of the decoded bit, E[X − p] = 0 and E[(X − p)2 ] = p(1 − p). Substitution into (15) gives zero mean and unit second moment per increment. A5(a) gives independence across positions, so the score variance is nt . The two-point MGF is exactly (9). Markov’s inequality applied to eαSi and a union over innocent rows prove (10). For the uniform alternative, Bernstein’s inequality gives

Proof of Corollary 1: The true error count is bounded above by the number of failed proxy bits. Conditional on (ui , ti ), these bits are independent with mean at most 1 − p. Hoeffding 2 gives P[di > (1 − p + ξ)n] ≤ e−2nξ after removing the conditioning. Union over clients and apply Theorem 2. This gives (8).

Solving z 2 /(2nt + 2Bz/3) = L yields (11). Equal moments do not imply equal MGFs: at p = .1, the laws for y = 0 and y = 1 are reflected asymmetric two-point distributions.

Proof of Lemma 5: More columns than rows imply a zero eigenvalue of E ⊤ E , hence an eigenvalue −1 of E ⊤ E − I . This proves only the rank assertion. Each independent codeword-pair distance is binomial. Applying the lower Hoeffding tail and taking a union over pairs proves (34). P ROPOSITION 4 (Exact presence calibration) Against an independent uniform binary codeword, a Hammingradius-t test has false-accept probability t   X n PFA (n, t) = 2−n . (36) j j=0

At n = 128, the largest radius satisfying PFA ≤ 2−128 is zero. Under an additional i.i.d. true-bit model with accuracy p, completeness is t   X n Paccept = (1 − p)j pn−j . (37) j j=0

18

P[Si > z | y, p] ≤ exp{−z 2 /(2nt + 2Bz/3)}.

Proof of Theorem 4: Condition on all coalition rows. Biases remain independent across positions under this conditioning. A position containing k ones has posterior bias density proportional to fδc (p)pk (1 − p)s−k . Conditional on the rows and an intended output word, the independent BSC flips give the exponential factor in (20). A randomized strategy is a mixture of such words. Its conditional product of factors is at most the product of the largest allowed factor at each position, even when it chooses its symbols jointly. Now average over the independent coalition columns. The posterior normalization cancels the column probability,  and summing over the ks columns with k ones gives P Js (t, q) per position. Hence E[e−t i∈C Si ] ≤ Js (t, q)nt . If all colluder scores are at most z , their sum is at most sz . Exponential Markov bounds that event by (21). For soundness with (22), first condition on biases and the output. Each innocent MGF is bounded by the maximum over the two output symbols. The resulting product depends only on the biases, whose independent draws give I(a)nt after averaging. Markov and a union over at most N innocents give N e−az I(a)nt . This is UNDER REVIEW

VOL. XX, No. XX

September 2026

an a-priori guarantee over code generation. It does not condition on selecting favorable realized codebooks.

D EFINITION 3 (Credential knowledge relation) The relation is R = {((A, y), e) : wt(e) = wτ , y ⊕ e ∈ Im(A)}. A witness (s, e) with y = As ⊕ e satisfies it.

Proof of Lemmas 2 and 3: For an independent sign flip, E[1 − 2F ] = 1 − 2q and (1 − 2F )2 = 1. Subtracting the squared new mean from the unchanged second moment gives the variance increase 4q(1−q)(E[T ])2 . For a unit carrier direction, |⟨η, ET,b ⟩| ≤ ∥η∥2 , so a larger absolute noiseless margin cannot change sign.

L EMMA 7 (Three-transcript extraction) Three accepting Stern transcripts with the same commitments and all three challenges reveal a witness for R, provided the commitments bind and the encoded permutation is valid.

Proof of Corollary 2, Proposition 3, and Lemma 4: Antisymmetry makes the score of the complemented decode equal −Si , allowing the same upper-tail certificate to treat negative scores. For each fixed decode, a passing check bounds the allocated tail event. A failed check prevents a certified decision. Union over the 2K allocated events proves the total budget without requiring carrier independence. Precedence can only remove decisions. For perfect negation, direct substitution into the score gives thepnegative absolute increment, whose expectation is −2 pb (1 − pb ). Proof of Proposition 2 and Theorem 5: Cauchy–Schwarz bounds the probe-margin change by rb = ∥fS (pb ) − fT (pb )∥2 . A correctly embedded teacher margin at least µ is preserved if rb < µ. Charge the qbad nt remaining probes in full and use #{b : rb ≥ µ}µ2 ≤ P 2 b rb to prove (14). This is a deterministic finite-sample inequality. Changing one binary output reverses its score increment, changing Si by absolute amount ai,b . Summing over changed positions proves (29). Maximizing a sum of k nonnegative weights selects the largest k . Only probes with small teacher margins or rb ≥ µ can change, giving (30). The direct residual-to-margin comparison is valid without in-distribution sampling or independence of the errors. It does not imply that the upper bound is tight.

C. Credential Proof and Security Accounting

A SSUMPTION 7 (Random-oracle and commitment model) Fiat–Shamir is modeled with a classical random oracle. Commitments use a fresh random salt and SHAKE-256 with a 256-bit output, modeled as hiding and  binding. The ideal query-limited collision bound is Q2 2−256 . An extraction reduction must account for all of its oracle queries when applying this bound. A SSUMPTION 8 (Computational witness recovery) Recovering a weight-wτ error e with y ⊕ e ∈ Im(A) from a random registered instance is assumed computationally hard. The experimental parameters are m = 1024, l = 512, τ = .125, wτ = 128. Information-set-decoding estimates quantify computational work [30], not forgery probabilities.

Proof: The openings jointly determine π, t0 , t1 , t2 . The first two checks imply t0 ⊕ π −1 (t1 ) ∈ Im(A) and t0 ⊕ π −1 (t2 ) ⊕ y ∈ Im(A). XOR gives y ⊕ π −1 (t1 ⊕ t2 ) ∈ Im(A). The third check and permutation invariance give weight wτ , so e′ = π −1 (t1 ⊕ t2 ) is a witness. This is the Stern-type extraction used in the xLPN protocol [3], [20], [21]. L EMMA 8 (Ideal Fiat–Shamir knowledge error) With ideal binding commitments and uniformly sampled ternary challenges, the r-fold parallel protocol has knowledge error (2/3)r . Its single-challenge-phase Fiat–Shamir transform has knowledge error at most (Q + 1)(2/3)r under generalized special-soundness extraction [42]. Proof: Let Γ contain challenge sets exhibiting all three values in some coordinate. Such a set extracts by Lemma 7. A non-extracting set has at most 2r vectors, giving κΓ = (2/3)r . Each useful challenge adds an unseen coordinate value, so tΓ ≤ 2r + 1. Useful challenges can be sampled by rejection outside the Cartesian product of previously seen values. Until extraction its probability is at most (2/3)r . Theorem 5 of [42] therefore applies with polynomial TΓ ≤ 2r + 2. Its extractor uses at most (Q + 1)(2r + 2)/(1 − κΓ ) expected prover calls and succeeds with probability at least (ϵ−(Q+1)κΓ )/(1−κΓ ). The linear bound on useful challenges establishes efficient knowledge extraction. Commitment failures must be added at the reduction’s actual query budget. C OROLLARY 3 (Round-count calibration) For the ideal knowledge-error term to be at most 2−129 , it suffices that r≥

129 + log2 (Q + 1) . log2 (3/2)

(38)

At Q = 264 , 330 rounds suffice and the implementation uses 331. The ideal term is approximately 2−129.623 . Proof: Take logarithms of (Q + 1)(2/3)r ≤ 2−129 and round upward. Here log2 (264 + 1) > 64. Keeping that term does not change the integer result. The experimental mapping of 16-bit words modulo three has maximum two-answer probability 43691/65536, giving approximately 2−129.619 instead. The verifier uses this mapping, with its bias included in the calibration. Rejection sampling provides an exactly uniform alternative. Proof of Proposition 1:

KARIM ET AL.: ZK-Trace: Collusion-Secure Traitor Tracing with Zero-Knowledge Exculpation

19

TABLE XIII: Security quantities and their separate meanings. Quantity

Interpretation

Ideal FS knowledge error Experimental ternary mapping Commitment failure

(Q + 1)(2/3)r ; extraction statement replace 2/3 by 43691/65536 for grinding calibration bounded at the reduction’s total query budget estimated attack cost; not 2−128 forgery probability certified tail budget under A5(a) finite bound (21) under A5(b)

xLPN recovery work Tracing false accusation Tracing completeness

Completeness. Honest openings satisfy the three algebraic checks for their respective challenges. If the required presence condition also holds, verification accepts. This is conditional completeness. Noisy extraction can fail presence even for an honest credential holder. Knowledge and zero knowledge. The extraction statement follows from Lemmas 7 and 8 in the ideal model. A simulator chooses challenges first, samples the corresponding accepting masked openings, commits arbitrary hidden values in unopened slots, and programs the challenge oracle at the complete commitment/context input. Hiding and fresh commitment entropy bound the distinguishing effects of unopened values and prior oracle queries. This yields the usual computational randomoracle simulation, subject to the commitment assumptions [19], [32]. Binding. The challenge context contains the extracted component, registered credential, and a canonical digest of the complete model state. An unchanged transcript then transfers to different model bytes only through a digest collision or a new-context challenge coincidence. Binding only the extracted component would not distinguish models with equal extracted bits. Model-state binding authenticates the credential statement rather than authorship: a public watermark can be copied, and a legitimate witness holder can authenticate after that copying. An end-to-end credential-forgery reduction must additionally bound witness-recovery advantage and extraction cost. The implementation validates encodings, uses operating-system randomness for proof masks and permutations, and binds transcripts to the full model state. Tracing-code generation uses a domain-separated SHAKE-256 stream with a 32-byte secret key. The experiments use fixed codebooks generated from a fixed experimental key and 64-bit NumPy seeds for reproducibility. These reproducible codebooks support experimental comparisons. Operational secrecy requires secret cryptographic generation, authenticated decoder setup, and the row-independence premise. 20

Appendix D Extended Results A. Configuration Details

The ten fixed seeds behind every reported mean ± standard deviation are {42, 137, 271, 314, 1729, 2718, 3141, 5772, 6561, 9999}. Collusion traceability averages 20 random coalitions’ copies per size for the weight carrier and 50 for the feature carrier, pooled over the ten seeds. Both training phases of Sec. VI use SGD with momentum 0.9, weight decay 5×10−4 , and cosine annealing, at learning rate 0.01 for the 200-epoch cross-entropy pre-training and η=0.001 annealed over the R federated rounds. Deployment scope of the simulated federation. All stations use partitions of one GNSS recording campaign [2]. The Dirichlet partition varies class coverage and sample count, but does not measure physical differences between sites. We sweep label skew to α=0.1 and participation to 30% per round (Table XIV). Extreme skew leaves some clients unable to form a five-way episode. Partial participation models absence for an entire round. The receiver-shift experiment adds a +3 dB gain offset and per-station SNRs of 5–20 dB (App. D-E). These tests do not cover distinct antenna and front-end calibrations, independent multipath, or local interference at separately sited receivers. The shift sweep is a proxy for receiver variation. Compute heterogeneity is also untested: all clients use the same local-episode budget, so the evidence does not cover stragglers, mid-round dropout, or unequal training progress. Physical deployment may change extraction errors, the required length in (23), and tracing completeness. Each deployment decode requires its own interval certificate (Remark 1). Under A5(a), innocent scores have mean zero and variance nt , but their tail bounds depend on the recovered bits and biases. In the proxy experiment, tracing persists in all 100 trials while exact single-leaker isolation degrades toward chance.

B. Attribution under Non-IID Data and Partial Participation

Table XIV separates attribution over the full roster from attribution over clients that can train. With 30% participation per round, roster-wide attribution remains 98% on GNSS and 100% on CIFAR-10. Under strong label skew, some clients have too few classes for a fiveway episode. They do not embed a mark, even though no shard is empty. At α=0.5, GNSS roster-wide attribution is 79%. Participating-only attribution is 100% in every reported cell on both datasets. The sweep therefore identifies episode formation as the source of the roster-wide losses in these configurations. UNDER REVIEW

VOL. XX, No. XX

September 2026

TABLE XIV: Attribution and few-shot accuracy under label non-IID (Dirichlet concentration α) and partial participation (per-round client fraction q ), Nc =10, 10-seed mean±std, GNSS first. Raw denotes attribution over all ten clients. Participated-only attribution is 100% in every cell (footnote). ∗ The deployed configuration is evaluated in an independent run of the sweep. Its GNSS accuracy sits within one standard deviation of Table II, whose wider-spread GNSS partition draws are the source of the difference. GNSS

CIFAR-10

Setting

Attr. (%)

Acc (%)

Attr. (%)

Acc (%)

α=0.1, q=1.0 α=0.25, q=1.0 α=0.5, q=1.0 α=2.0, q=1.0∗ α=2.0, q=0.5 α=2.0, q=0.3

16.0±12.0 32.0±16.6 79.0±13.0 98.0±4.0 98.0±4.0 98.0±4.0

58.2±9.4 84.1±8.3 92.0±0.9 92.3±2.6 91.4±2.2 90.7±2.2

51.0±14.5 91.0±7.0 100.0±0.0 100.0±0.0 100.0±0.0 100.0±0.0

74.6±4.5 81.7±0.6 83.3±0.4 84.9±0.2 84.5±0.3 84.0±0.4

Participated-only attribution (restricted to clients that form at least one episode) is 100% in every cell on both datasets, and no shard is ever empty. Under α=0.1 a mean of 4.1 (CIFAR-10) and 0.5 (GNSS) of ten clients can assemble a 5-way 5-shot episode, so the raw rate tracks episode-formability rather than attribution loss. On GNSS no client forms one in six of the ten seeds, leaving that cell’s participated-only figure resting on the remaining four.

C. Tracing, Collusion, and Anti-Framing

Collusion for removal, extended settings. Section VII-B reports traceability against coalition size k at the operating point. Two extensions complete the picture. The weight-space overlay repeats its profile at Nc =40, tracing all k=2 and k=3 coalitions (10/10) and 0.70 of k=5 on GNSS (0.20 on CIFAR-10), again with zero framing. In the adaptive-steering experiment, the simulator edits the decoded word toward the coalition majority. It does not optimize model parameters toward a chosen innocent. At k=2, every tested steered coalition is traced. Stronger steering lowers the colluders’ own flip rate from 0.28 to 0.11, increasing traceability in these configurations. No innocent is framed across 400 steered trials. These observations characterize the tested attack. The formal soundness guarantee remains subject to Theorem 3 and Assumption 5. The mean disagreement with individual colluder rows q̄ on the weight carrier rises with k , measuring 0.22/0.24/0.26/0.26 (GNSS) and 0.25/0.26/0.27/0.28 (CIFAR-10) at k=2/3/5/8. Cooperation-free tracing and registry adjudication. Tracing an unknown-origin model needs no help from the leaker. Rebuilding the carriers from the public seeds and decoding the copy offline names the exact leaker in all ten seeds (10/10),1 and the registry adjudicates the accusation against the enrolled commitments in all ten (10/10), so a disputed model is resolved without the suspect ever participating. This closes the gap left by proof-of-ownership schemes that confirm only a cooperating owner. 1 The released per-seed logs record these two outcomes under the metric

identifiers trace_exact and frame_tardos_hit.

Anti-framing and credential verification. Copying a victim’s public identity codeword raises its self bitaccuracy to 1.00, showing that public-codeword presence alone can be forged. The secret overlay accuses the framed victim in 0/10 trials. The credential proof accepts the legitimate victim in 10/10 trials and accepts the tested forgeries in 0/10. Registry adjudication also rejects evidence substitution in all ten trials. The credential acceptance and forgery benchmarks use transcripts bound to the extracted component and public credential. To evaluate full-state binding, we change the model state while holding the extracted component fixed. The verifier rejects the transferred transcript. These empirical rejection counts are distinct from the proof knowledge-error bound below 2−128 at r = 331 under the stated query budget (Proposition 1). Wrong-client, serverside, and replay attempts are included in the forgery tests. A credential proof costs approximately 0.6 MB and 37 s per client. The repetitions can be verified in parallel. This is an offline dispute procedure, not part of each training round. Dispatch without client data. The two carriers share every downstream step, the code, the registry, credential authentication, and the conditional score bound (Remark 2, Theorem 3). Section VII-B sets their operational trade side by side. Neither requires the operator to hold client data. Writing the feature carrier from a server-held proxy set reproduces the same profile, 100/100 isolation and 10/10 distillation survival at a 4–5-point copy cost, with tracing unaffected by adaptive steering. The deployed global model is untouched by either carrier, since both are written only into dispatched copies (App. A-B). The global-accuracy cost is zero. D. Distillation Leveling

The single cross-architecture failure. ResNet-18to-ResNet-34 distillation preserves the feature mark in 19/20 runs. The failed run is GNSS seed 271: its postdistillation flip rate is 0.4976 and its score is 54.8. Its same-architecture flip rate was 0.147, below the GNSS median. The increase of 0.351 is the largest among all twenty runs, so the failure reflects a large architectureinduced change rather than a mark already close to chance. The sweep holds the probes, code, and 80-epoch budget fixed. ResNet-34 is chosen because its penultimate layer has the same width, 512, as the teacher’s. On GNSS, the mean flip rate increases from 0.195 under samearchitecture distillation to 0.291 under cross-architecture distillation, a rise of 0.096. On CIFAR-10 it increases from 0.119 to 0.165, a rise of 0.046. Cross-architecture rates also span a wider range on GNSS (0.165–0.498) than on CIFAR-10 (0.132–0.211). The feature carrier’s clean flip rates are below 0.09 and 0.002, respectively. Equation (14) provides an interpretation: featurematching errors that exceed the teacher’s probe margins can flip decoded bits. It bounds the fraction of potentially

KARIM ET AL.: ZK-Trace: Collusion-Secure Traitor Tracing with Zero-Knowledge Exculpation

21

TABLE XV: Feature-carrier survival by distiller class, 20 seeds (10 GNSS + 10 CIFAR-10). Survival counts are shown for each tested class. Weight-space marks fail where extraction is defined. The incompatible crossarchitecture DeepMarks-BIBD runs return no framing verdict. Distiller

Student copies

Survival

Method

Carrier

GNSS

CIFAR-10

Feature-matching (FitNets) Cross-architecture (R18→R34) Logit-only KD Feature isometry

feature geometry feature geometry output function output function

20/20 19/20 0/20 0/20

Ours (feature) Ours (weight) FedIPR FedZKP FedTracker DUW WAFFLE

feature-space weight-space weight-space weight-space weight-space output-space output-space

10/10 0/10‡ 0/10 0/10 0.10† 0.30† 0.30

10/10 0/10‡ 0/10 0/10 0.10† 1.00 0.00

changed bits. Theorem 5 translates those changes into a score bound by weighting each position by its score impact. The measured flip rates are outcomes, however, and do not directly measure εKD or establish that this upper bound is tight. GNSS’s sparse four-channel inputs and subtler inter-class differences are plausible sources of transfer difficulty, but the experiment does not isolate them. The larger per-class recall cost on GNSS (10.9 versus 4.4 points) is consistent with a more demanding representation problem. The baseline scale G appears only in the weight-carrier bound and does not explain this feature-carrier failure. The failed run loses the mark without a false accusation. Across the distillation evaluation, no innocent is framed in the 180 runs returning a framing verdict. The remaining twenty rows are the DeepMarks-BIBD weightcarrier comparator under cross-architecture transfer, for which the BN-γ extraction layout is incompatible with ResNet-34 and no framing verdict is returned. Distillation across the benchmark. Table XVI evaluates each native mark after the same 80-epoch distillation. FedIPR, FedZKP, and our weight carrier lose their marks on both datasets (0/10 per dataset). FedTracker’s attribution is at chance (0.10). Output-space methods vary by dataset: WAFFLE retains its mark in 3/10 GNSS runs and 0/10 CIFAR-10 runs, while DUW retains its perclient key in 3/10 and 10/10, respectively. WAFFLE’s clean GNSS mark is already near its chance level. DUW’s CIFAR-10 survival shows that a key in the teacher’s function can transfer to the student, but DUW lacks a collusion-secure code and certified falseaccusation bound. Our feature carrier combines that transfer mechanism with coded tracing. It survives feature matching in 20/20 runs and cross-architecture transfer in 19/20. Logit-only distillation and feature isometry erase it (0/20 each; Table XV). Thus the observed advantage is specific to distillers that reproduce the feature geometry closely enough, as required by Assumption 6. E. Insider Attacks and Deployment Stressors

Own-row erasure: scores and budgets. Negated erasure moves the feature score from approximately +1317 to −1265. The two-sided tamper test (Ztwo =201.4) detects 22

TABLE XVI: Distillation leveling: native mark retained after 80-epoch KD, as a fraction of 10 seeds surviving (GNSS first). Under this feature-matching protocol, weight-space marks fail, output-space results depend on the dataset, and the feature carrier survives on both datasets.

† Consistent with per-seed Bernoulli outcomes at the 1/N =0.1 c chance level (upper-tail p=0.65 for FedTracker, p=0.07 for DUW on GNSS). DUW on CIFAR-10 is genuine survival, not chance (10/10, p≈10−10 under the chance model), and is discussed in the text. ‡ Retention criterion for our weight carrier: self bit-accuracy, which collapses to its 0.5 chance level in every seed. Under a twice-chance attribution criterion (attribution >0.2 against chance 0.1) the cells read 0/10 on GNSS and 2/10 on CIFAR-10.

it in all twenty seeds, at a task-accuracy cost of three to four points. On the weight overlay, the same strategy leaves only an uncertified forensic lead. Fresh erasure has the complementary effect: it drives the feature score near zero, while the weight overlay still traces the true row in all twenty runs. Neither strategy frames an innocent. The hybrid combines weight negation with fresh feature erasure. It escapes the threshold-only tests on both carriers in 5/20 runs at matched compute, 20/20 at twice that budget, and 0/20 at four times the budget, where the weight score crosses the negative tamper threshold. The budgets are 300, 600, and 1200 episodes, with twenty runs per level. Appendix A-E gives the score ranges and geometric explanation. These observations characterize the tested budgets rather than establish an upper bound on an adaptive hybrid’s success. Framing remains 0/20 at each level. Robust aggregation. With coordinate-wise median aggregation, attribution is 1.00 on CIFAR-10 and 0.98 on GNSS, matching the reported FedAvg rates. Self bitaccuracies are 0.92 and 0.95, respectively. Task accuracy falls to 80.2% on CIFAR-10 and 63.4% on GNSS, with high GNSS variance. These results show that the tested median configuration retains identity attribution. They do not establish invariance to arbitrary robust aggregators. Receiver covariate shift (GNSS). We apply a +3 dB gain offset and per-station SNRs from 5 to 20 dB to model receiver variation. Tracing succeeds in all 100 shifted trials, and aggregate attribution remains 0.98. Exact singleleaker isolation degrades toward chance under the same perturbations. Thus retaining a tracing accusation does not imply retaining the stronger isolation outcome. Per-class cost on dispatched copies. We evaluate class recall against a clean control over ten seeds in the original label space. Mean recall falls by 4.4 points UNDER REVIEW

VOL. XX, No. XX

September 2026

TABLE XVII: Paired global-model equivalence at Nc = 10, ten seeds. ∆ is watermark-on minus watermark-off accuracy in percentage points. Confidence intervals are paired 90% intervals. Both TOST tests pass at their stated margins. Dataset GNSS CIFAR-10

∆

90% CI

Margin

TOST p

−0.04 +0.24

[−0.585, 0.505] [−0.018, 0.498]

1.0 0.5

0.0052 0.0489

on CIFAR-10 (0.844 → 0.799), approximately uniformly across classes, and by 10.9 points on GNSS (0.932 → 0.823). GNSS losses range from 0.9 to 18.0 points, with the largest losses on the harder interference types. These class-level differences provide information that episodeaveraged accuracy does not capture. The benchmark excludes the interference-free class, so it cannot measure false alarms on clean signals. The reported far proxy is instead inter-class confusion, P (pred ̸= true | true = c). Open-set abstention is not evaluated. These costs concern the dispatched copies, not the global model tested in App. D-F. Embedding channel. The weight-space carrier is held almost entirely by the BN γ scales, whose standard deviation broadens by 6.7–7.1× while β is nearly untouched, consistent with a hinge loss that steers each projection to the correct sign with margin. The representation changes are consistent with, but do not isolate, the effect of the identity hinge on the feature-space geometry reported in Section VII-C. F. Global-Model Utility Equivalence

Matched global-model comparison. We pair the clean rows of the ten per-seed attribution-attack tables with the corresponding watermark-off controls, using the same dataset and seed and differing only in λwm . The test uses the sample standard deviation of paired differences and Student’s t distribution with nine degrees of freedom. Table XVII reports the mean, 90% interval, and the larger one-sided TOST p-value. Both datasets establish equivalence at the specified margins. This result concerns episodic accuracy at Nc = 10 and does not establish equality of feature geometry or dispatched-copy utility. The reproduction artifact records every input path, file hash, and paired difference. G. Identity-Survival Diagnostics

The quadrature approximation uses the measured pretrained scale norm G (7.32 on CIFAR-10, 4.73 on GNSS), with no fitted coefficient. Its per-bit prediction and the observed self-agreement both decrease with identity load in Table III. The observations exceed the approximation in every cell. Repeated embedding may contribute to that gap, but this comparison does not establish the geometry or conditional independence needed by Theorem 1.

At Nc = 40, GNSS and CIFAR-10 have similar selfagreement (0.77 and 0.76), but attribution rates of 0.87 and 1.00. Mean agreement therefore does not determine attribution: the decoding radius and minimum codeword separation in Theorem 2 control the competing decisions. The smaller GNSS scale norm improves the quadrature prediction and cannot explain its earlier attribution loss. Above ρ = 1, the projection system is overcomplete, but rank alone does not force decoding failure. These data establish the dataset difference without identifying its cause.

H. Adversarial Attack Protocol

Model-modification attacks. Attack (1) prunes 5– 90% of the smallest-magnitude BN γ values. Its structured variant removes entire channels, including their γ , β , and running statistics. Attack (2) adds Gaussian noise N (0, σ std(γ)) to all BN γ , with σ ∈ {0.01, 0.05, 0.1, 0.5, 1.0, 2.0}. Attack (3) uniformly quantizes BN γ to b bits, where b ∈ {16, 8, 6, 4, 3, 2}. Attack (4) combines pruning and quantization in six configurations. Targeted attacks. Attack (5) uses norm-constrained PGD to maximize a bit-flip objective. Each step is projected into an ϵ-ball around the watermarked γ , using an ℓ∞ clamp or ℓ2 rescaling with ϵ ∈ {0.05, 0.1, 0.2}. Attack (6) resets the BN γ values in selected ResNet-18 stages to their default 1.0. We test resets from a single stage through all BN layers. Training-based attacks. Attack (7) fine-tunes the model for 10–500 episodic ProtoNet episodes without the watermark term. Attack (8) distills a freshly initialized ResNet-18 student for 5–100 epochs using the same data and optimizer. The feature-matching objective combines mean-squared embedding error with a Kullback– Leibler term on temperature-softened embedding coordinates (T = 4), following the hint-matching approach of [34]. Both terms align feature coordinates without matching task logits. The distillation-class sweep distinguishes this objective from soft-label distillation [26]. It tests feature matching, cross-architecture transfer to ResNet-34, logitonly distillation, and feature isometry. The isometry composes a frozen backbone with a random orthogonal map, preserving the output function while changing feature coordinates (Sec. VII-C). Insider and deployment stressor budgets. Ownrow erasure targets the recipient’s tracing row on each carrier. The negation strategy writes the opposite bits, while fresh erasure writes an independent random row. The hybrid combines weight-overlay negation with fresh feature-carrier erasure. We test 300, 600, and 1200 attack episodes, corresponding to 1–4× the matched split budget. For GNSS receiver covariate shift, we perturb the GNSS front-end with a +3 dB gain offset and sweep per-station SNR over 5–20 dB.

KARIM ET AL.: ZK-Trace: Collusion-Secure Traitor Tracing with Zero-Knowledge Exculpation

23

Record · ID 667893 · SHA-256 e64104188d31f338
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.