Conformal Privacy Auditing: Calibrated Re-identification Attacks with Statistical Guarantees Shuo Huang1
Gholamreza Haffari1 Xingliang Yuan2 Ting Yu3 Lizhen Qu1 * 1 Monash University 2 The University of Melbourne 3 Mohamed bin Zayed University of Artificial Intelligence {shuo.huang1,gholamreza.haffari,lizhen.qu}@monash.edu [email protected] [email protected] Abstract
arXiv:2609.21340v1 [cs.CR] 18 Sep 2026
Empirical identity leakage from released text is increasingly driven by attackers that combine large language models (LLMs) with auxiliary knowledge to link documents to individuals. Existing audits typically report success rates for specific attack pipelines but lack finitesample statistical guarantees, while trainingtime protections such as differential privacy are difficult to translate into release-time decisions for individual natural-language documents. We introduce Conformal Privacy Auditing (CPA), a distribution-free calibration framework that provides a statistical certificate of reidentification risk for each released document against LLM-empowered adversaries. CPA outputs a conformal ambiguity set of candidate identities that is guaranteed to contain the true identity with user-chosen confidence under exchangeability, together with an interpretable leakage proxy derived from set size. CPA supports both logit-access and sampling-only attackers, enabling audits of open-source models and proprietary API models in a unified framework. Across multiple release benchmarks and attacker configurations, CPA achieves calibrated coverage and reveals sharp shifts in certified identifiability as auxiliary knowledge, LLM augmentation, and release mechanisms vary, providing a statistically grounded basis for reporting and comparing release-time linkage risk across attacker configurations, datasets, and release mechanisms alike.
1
Introduction
Privacy concerns have intensified alongside the rapid growth of large language models (LLMs) and their widespread applications (Xin et al., 2025). These applications are inherently data-intensive. In many domains, including healthcare and law, organizations face a strong need to release textual data, such as clinical notes (Lison et al., 2021), legal documents (Deuber et al., 2023), and online posts (Dou et al., 2024), to support research and product development. At the same time, they must ensure that releasing such data does not compromise the privacy of the individuals it describes. * Corresponding author.
A common practice for releasing sensitive textual data is redaction (Pilán et al., 2022), which replaces private information belonging to predefined categories with symbolic placeholders. However, as illustrated in Fig. 1, combinations of seemingly innocuous attributes, such as “37-year-old”, “female”, “nephrologist”, and “Brisbane”, can remain in redacted texts and collectively serve as linkable signals (Xin et al., 2025), enabling the re-identification of individuals who were intended to be anonymized. Recent studies (Manzanares-Salor et al., 2024; Staab et al., 2024) further show that the widespread use of LLMs has effectively turned them into practical privacy adversaries, capable of extracting salient cues from text and integrating auxiliary information to substantially amplify re-identification risks, even for carefully redacted text. Existing approaches for auditing privacy risks in released textual data remain limited, particularly in the presence of LLMs with strong reasoning capabilities. Differential Privacy (DP) provides formal worst-case guarantees by injecting carefully calibrated noise into the data (Dwork, 2006). However, these guarantees are often overly conservative in practice, leading to significant utility degradation, while failing to capture the nuanced and context-dependent privacy risks observed in real-world deployments. In contrast, empirical privacy auditing methods quantify privacy leakage using attack-based metrics, such as re-identification success rates; such evaluations provide practical evidence of observed leakage but lack theoretical guarantees, leaving it unclear how well the measured risks generalize to unseen data or stronger attackers. To address these challenges, we propose Conformal Privacy Auditing (CPA), a theoretically grounded black-box framework for auditing re-identification risks of individuals in released texts under a specified attacker model. CPA builds on the theory of Conformal Prediction (Vovk et al., 2005; Bates et al., 2021), a distributionand model-agnostic framework for uncertainty quantification with formal statistical guarantees. As illustrated in Fig. 1, given a released document and attacker background knowledge, CPA employs a score function to identify a certified ambiguity set that covers the true target individual with a high probability, e.g., 90%, under the chosen attack model. The size of this ambiguity set serves as an interpretable measure of re-identification uncertainty—a large set (e.g., 106 candidates) indicates high uncertainty for the attacker to correctly identify
the individual. Importantly, CPA provides explicit statistical conditions under which these guarantees hold, enabling principled interpretation and comparison of auditing outcomes across threat models. The contributions of this work are three-fold. • We formulate re-identification auditing of released texts as constructing conformal ambiguity sets that serve as privacy certificates, and develop a theoretical framework specifying the statistical conditions under which these certificates provably hold. • We design nonconformity scores for LLM and retrieval-based attackers requiring only candidate log-probabilities or similarity scores, enabling fully black-box auditing of open-source models and proprietary API-based models without any internal or logit-level access. • We demonstrate that CPA certifies privacy risks against a declared family of attackers, with formal statistical guarantees across diverse release benchmarks, attacker families, and adversarial settings.
2
Preliminaries: Conformal Prediction
Conformal prediction (CP) turns arbitrary model scores into set-valued predictions with distribution-free, finitesample guarantees (Vovk et al., 2005). It requires no correctly specified probabilistic model; it relies only on exchangeability: the joint distribution of (Z1 , . . . , Zn+1 ) with Zi = (Xi , Yi ) is invariant under permutations of indices (i.i.d. sampling is sufficient), so that a test example is statistically indistinguishable from the n examples of a held-out calibration sample and ranks of calibration scores form valid quantiles. Concretely, fix any trained (possibly black-box) predictive system and a nonconformity score s : X × Y → R, where larger values indicate that label y is less compatible with input x. Given calibration scores Si = s(Xi , Yi ), split conformal sets the cutoff q̂α to the ⌈(n + 1)(1 − α)⌉-th smallest calibration score and returns the prediction set Cα (x) = {y ∈ Y : s(x, y) ≤ q̂α }. A standard rank argument yields the finite-sample marginal coverage guarantee Pr Yn+1 ∈ Cα (Xn+1 ) ≥ 1 − α, (1)
and inherits any miscalibration of the underlying score. Split CP instead specifies the exact order statistic whose use guarantees marginal coverage at level 1 − α for the next exchangeable example—for any score function, without parametric or calibration assumptions. In our setting, this is what turns an arbitrary attacker score into a statistically calibrated ambiguity certificate with finite-sample validity guarantees.
3
Auditing Identity Linkage Risk
3.1
Running example: Release-time identity linkage auditing
Consider a healthcare provider releasing de-identified clinical case summaries (x) who must ensure that an attacker cannot link them to specific patients using auxiliary background knowledge (R), such as demographics or public registries. The provider uses CPA to simulate a realistic linkage attacker that extracts clues with LLMs and ranks potential matches against a candidate profile database. For each document, the audit produces a certified ambiguity set Cα (x): a statistically guaranteed “short-list” of suspects containing the true patient with confidence 1 − α. A small set (e.g., |Cα (x)| = 3) immediately signals that the attacker can narrow the identity to a tiny group, so the provider can withhold or further redact documents whose certified ambiguity falls below a minimum threshold k—directly mirroring government guidance that ties release decisions to numeric re-identification risk (e.g., the 0.09 threshold used by the Information and Privacy Commissioner of Ontario, 2016). 3.2
Released documents and identities. We study identity linkage risk for released text. Let X ∈ X denote a released document (possibly anonymized or rewritten) and let Y ∈ U denote the true identity of the subject. Here U is a reference population (e.g., a candidate profile database). The goal of auditing is not to protect against an unbounded attacker, but to provide a quantitative, finite-sample certificate of linkage risk under a declared attacker capability and budget. 3.3
which is distribution-free and holds for any score function s computed consistently on calibration and test examples (Vovk et al., 2005). In multiclass settings, adaptive prediction sets (APS) construct s(x, y) from the model’s ranked probabilities so that sets shrink on easy examples and expand on hard ones while preserving marginal coverage (Romano et al., 2020). Relation to threshold tuning on a hold-out set. Split CP resembles the familiar practice of tuning a classifier’s decision threshold on held-out data, but differs in what is guaranteed. Ad-hoc threshold tuning targets an empirical operating point (e.g., a precision/recall tradeoff), makes no finite-sample statement about future data,
Task Formulation
Threat Model
Every CPA certificate is issued relative to a declared threat configuration. Following standard practice in security research, we make each component of this configuration explicit below; Table 7 summarizes the knowledge ladder instantiated in our experiments, and Table 8 in Appendix D maps every experimental suite to its full configuration. A notation table consolidating the symbols used in this and the following sections is provided in Appendix C (Table 6) for reference. Attacker goal. Given a released document x, the attacker attempts to link x to the true subject identity Y within the reference population U.
Figure 1: CPA auditing vs. top-1 prediction. A conventional attack pipeline (top) outputs a single point prediction; CPA (bottom) calibrates the same attacker into a conformal ambiguity set that contains the true identity with probability at least 1 − α under the declared threat configuration. Side information and tooling as a knowledge level. Privacy leakage is attacker-relative: an attacker with access to a detailed candidate profile database and strong search tooling can link text that would be safe under a weaker attacker. We represent attacker side information and tooling by a variable W ∈ W, which encodes (i) which fields are accessible in the candidate profile database (metadata-only vs. enriched profiles), (ii) which auxiliary corpora or snapshots are available, (iii) whether the attacker may use LLM-based clue extraction or web-search augmentation, and (iv) decoding/sampling settings. We interpret different values of W as knowledge levels (threat configurations) and report certificates separately for each knowledge level. Candidate-pool access (closed vs. open world). Linkage is meaningful only relative to a finite candidate pool. We assume the attacker operates on a set of candidates Y(x, w) ⊆ U ,
|Y(x, w)| < ∞,
Auditing goal and outputs. For each document (x, w), the auditor outputs a certified ambiguity set Cα (x, w) ⊆ Y(x, w) such that the true identity is included with user-chosen confidence 1 − α. The set size |Cα (x, w)| is the primary certificate: it counts how many candidate identities remain statistically plausible after the attack. For reporting convenience we also summarize leakage by the inverse ambiguity proxy
(2)
derived from an attacker-accessible candidate profile database, demographic filters, or a fixed snapshot corpus. In the closed-world setting the true identity is guaranteed to lie in Y(x, w); in the open-world setting it may be absent. When we restrict scoring to a top-K retrieved subset for computational reasons, this truncation is an engineering choice that can re-introduce open-world behavior; we quantify the resulting effect on coverage empirically in the stress tests of Section 7.1. Query interface and budget. Given (x, w) and candidate pool Y(x, w), the attacker pipeline produces a ranking or distribution over candidates: pA (· | x, w) ∈ ∆(Y(x, w)).
Scope of the certificate. The candidate-pool construction, release mechanism, query mode, knowledge level w, sampling budget m, and the attacker’s randomness model are all part of the declared configuration, and CPA certifies coverage for that configuration only: any change to the release distribution or attacker pipeline (e.g., a new prompt, model, or web snapshot) requires recalibration. We stress-test this assumption boundary empirically in the robustness analysis of Section 7.1 and the full tables of Appendix E.
(3)
Our framework treats the attacker as a black box. In the weakest interface we assume only sampling-only access via repeated forced-choice queries over Y(x, w), with a declared per-document sampling budget m; when candidate log-probabilities are exposed we use them directly. This covers retrieval-based linkers (e.g., similarity search over profiles) and LLM-based pipelines that infer attributes/clues and match them to profiles, including proprietary API models.
L(x, w) :=
1 . max{1, |Cα (x, w)|}
(4)
We emphasize that L(x, w) is not a probability of reidentification: it is a monotone transform of the certified set size, and because it decays rapidly with |Cα |, seemingly small values can still represent substantial risk (e.g., L = 0.1 means the attacker has certifiably narrowed the subject to 10 candidates, which may be unacceptable under many release policies). L(x, w) should be interpreted as a certificate of attacker ambiguity rather than a cryptographic privacy guarantee. Candidate-miss (open-world) caveat. In open-world settings the true identity may be absent from the candidate pool. We measure this via the candidate-miss probability ρ(w) := Pr Y ∈ / Y(X, w) ,
(5)
reported empirically as candidate_miss_rate. Our coverage guarantees apply conditional on Y ∈ Y(X, w) and degrade gracefully with ρ(w), as quantified by the open-world bound in Corollary 5.3.
4
Methodology: Conformal Privacy Auditing (CPA)
4.1
Nonconformity for Identity Linkage: APS Mass Score
We require a scalar score that measures how much probability mass must be accumulated to include the true identity under attacker A. Let π1 (x, w), π2 (x, w), . . . be candidates sorted by decreasing pA (· | x, w). We use an APS-style mass score, restricted to the candidate pool: s(x, y; w) :=
X
pA (πj | x, w) ∈ (0, 1], (6)
j: pA (πj |x,w) ≥pA (y|x,w)
where we abbreviate πj ≡ πj (x, w). Smaller s(x, y; w) indicates that the attacker concentrates probability on the truth (high linkage risk). Larger s(x, y; w) indicates diffuse belief (lower linkage risk). CPA is not tied to a single score. The calibration layer of CPA requires only an exchangeable scalar score; the APS mass score is one instantiation. In particular, the weights in (6) need not be calibrated probabilities: for retrieval attackers they are pseudo-posteriors obtained by softmax-normalizing similarity scores, and for ranking-only attackers a purely rank-based nonconformity score is equally valid. Our ablations (Section 7.1, Table 11) show that APS, rank-based, and raw-probability constructors all achieve valid coverage; we adopt APS as a robust default rather than claiming it is uniformly dominant. Sampling-only interface. When logits are unavailable, we query the attacker m times with forced-choice decoding over Y(x, w) and form empirical frequencies p̂m (· | x, w). We then compute the plug-in score ŝm (x, y; w) by replacing pA with p̂m in (6). This yields a fully black-box conformal pipeline applicable to proprietary LLM APIs that expose no logits. 4.2
Split Conformal Ambiguity Sets
Given calibration data {(Xi , Yi , Wi )}ni=1 and a fixed knowledge level w (or conditioning on W = w), we compute scalar scores ( s(Xi , Yi ; w), probability/logits available, Ti := ŝm (Xi , Yi ; w), sampling-only. (7) Let T(1) ≤ · · · ≤ T(n) be the sorted scores and define the split-conformal threshold k := ⌈(n + 1)(1 − α)⌉ ,
q̂α := T(k) .
(8)
For a new document (x, w), we output the calibrated ambiguity set, i.e., the prediction set that contains the true identity with high confidence: Cα (x, w) := {y ∈ Y(x, w) : s(x, y; w) ≤ q̂α }, (9)
using s or ŝm depending on access. Note that the inequality direction “≤” in (9) follows the nonconformity convention: top-ranked candidates have smaller cumulative mass scores, so the set retains all candidates that the attacker cannot statistically rule out. In particular, q̂α is a calibrated quantile of nonconformity scores, not a probability-of-linkage threshold. Section 5 gives the finite-sample coverage guarantee under exchangeability, together with its open-world extension. 4.3
Instantiations of the Attacker Pipeline
Our framework is attacker-agnostic: any pipeline that induces a distribution over candidates can be conformalized without any modification. Retrieval-CPA (profile matching). We construct a profile text representation for each candidate profile (using a controlled schema) and define pA (· | x, w) from normalized similarity scores (e.g., TF–IDF cosine similarity). Knowledge level w determines which profile fields and which query transformations are permitted (e.g., metadata-only vs. keyword profiles; with or without LLM clue extraction at query time). Profile-CPA (LLM extraction → matching). An LLM attacker extracts attributes from x (e.g., occupation, location, age range), producing distributions over attribute values via sampling. These distributions induce probabilistic match scores against candidate profiles, yielding pA (· | x, w) over identities. We then apply CPA exactly as in (8)–(9). This instantiation matches a realistic threat model in which LLMs infer quasi-identifiers and link to a candidate profile. The full procedure is listed in Algorithm 1 in Appendix A.1.
5
Privacy guarantee for CPA
This section formalizes the statistical guarantees provided by Conformal Privacy Auditing (CPA). Our results are distribution-free (no parametric assumptions on data) and apply to black-box attackers, provided an exchangeability assumption holds for the calibration and evaluation distributions under the declared attacker knowledge level w (Section 3.3). 5.1
Assumptions
Exchangeability (within a declared knowledge level). Fix a knowledge level w ∈ W that specifies the attacker’s side information and tooling (Section 3.3). We assume that the sequence {(Xi , Yi )}n+1 i=1 used for calibration/evaluation is exchangeable under this fixed w (i.i.d. sampling is sufficient). Equivalently, in the most general form we assume {(Xi , Yi , Wi )}n+1 i=1 is exchangeable; when experiments are run stratified by category and knowledge level (as in our suites), Wi ≡ w is constant and exchangeability reduces to exchangeability of (Xi , Yi ) within that stratum. Attacker pipeline and randomness. For each example i, the attacker pipeline (retrieval or LLM-based) may involve internal randomness—e.g., sampling-only
forced-choice decoding, stochastic clue extraction, or stochastic rewriting. We denote this randomness by Ui and assume: (i) Ui are i.i.d. across examples, and (ii) Ui are independent of the data (Xi , Yi , Wi ). This reflects operational practice: each audit query uses an independent random seed / sampling stream, and the auditor does not adversarially choose randomness based on the content of any individual example. Candidate-pool miss probability. In open-world linkage, the true identity may be absent from the candidate pool Y(X, w). We measure this using the miss probability Pr Y ∈ / Y(X, w) ≤ ρ(w), (10) reported empirically as candidate_miss_rate. Our strongest guarantees apply conditional on Y ∈ Y(X, w) and degrade gracefully with the miss probability ρ(w), as shown in Corollary 5.3. 5.2
CPA Validity for a Fixed Attacker Configuration
Recall the APS-style nonconformity score (restricted to Y(x, w)): X s(x, y; w) = pA (y ′ | x, w) ∈ (0, 1]. y ′ : pA (y ′ |x,w) ≥pA (y|x,w)
In the sampling-only interface, pA is replaced by empirical frequencies p̂m , yielding ŝm (x, y; w). CPA uses split conformal calibration: compute scores on a calibration set, take q̂α as the ⌈(n + 1)(1 − α)⌉-th order statistic, and return Cα (x, w) = {y : s(x, y; w) ≤ q̂α }. Lemma 5.1 (Score exchangeability under sampling-only access). Let {(Xi , Yi , Wi )}n+1 i=1 be exchangeable and let Ui be i.i.d. independent randomness used by the attacker pipeline to compute ŝm (Xi , Yi ; Wi ). Define scalar scores Ti := ŝm (Xi , Yi ; Wi ) = f (Xi , Yi , Wi , Ui ) for a deterministic measurable function f (deterministic given the example and the sampled outputs). Then {Ti }n+1 i=1 is an exchangeable sequence of scalar scores. Proof sketch. Augment each example to (Xi , Yi , Wi , Ui ). Exchangeability of (Xi , Yi , Wi ) and i.i.d. independence of Ui implies the augmented sequence is exchangeable. Since Ti is a deterministic function of the augmented example, exchangeability is preserved under deterministic measurable mappings of the entire augmented example sequence. Theorem 5.2 (Finite-sample CPA validity for identity linkage). Fix a knowledge level w and attacker pipeline A. Assume: (i) exchangeability holds for calibration and test examples under w, (ii) Y ∈ Y(X, w) almost surely (closed-world or ρ(w) = 0), and (iii) sampling randomness (if used) satisfies Lemma 5.1. Let q̂α be the
split-conformal threshold computed from n calibration scores at level α, and let Cα (·, w) be defined by the rule s(x, y; w) ≤ q̂α (or ŝm in the sampling-only case). Then for a new test example (Xn+1 , Yn+1 ), Pr Yn+1 ∈ Cα (Xn+1 , w) ≥ 1 − α. (11) Proof sketch. Let Ti denote the scalar scores used for conformalization. By exchangeability (Lemma 5.1 if needed), the rank of Tn+1 among {T1 , . . . , Tn , Tn+1 } is uniform. With q̂α chosen as the k-th order statistic for k = ⌈(n + 1)(1 − α)⌉, we obtain Pr(Tn+1 ≤ q̂α ) ≥ 1 − α. Membership Y ∈ Cα (X, w) is equivalent to T ≤ q̂α . Corollary 5.3 (Open-world validity with candidate-miss probability). Fix w and suppose Pr(Y ∈ Y(X, w)) ≥ 1 − ρ(w). If conditional on the event E = {Y ∈ Y(X, w)} the CPA set satisfies Pr(Y ∈ Cα (X, w) | E) ≥ 1 − α, then Pr Y ∈ Cα (X, w) ≥ (1 − ρ(w))(1 − α) (12) ≥ 1 − (ρ(w) + α). Proof. By the law of total probability, Pr(Y ∈ Cα ) = Pr(E) Pr(Y ∈ Cα | E), and apply the two bounds. Theorem 5.2 states that CPA outputs a valid ambiguity set containing the truth with probability at least 1 − α under the declared attacker knowledge level w. The ambiguity proxy 1/|Cα | is therefore a calibrated certificate of attacker ambiguity, not a claim about universal privacy against unbounded attackers.
6
Experiments
Goal. We evaluate conformal privacy auditing as a reporting layer for document release. Given a released text x and an attacker-side candidate profile database R that encodes auxiliary knowledge, our auditor outputs a certified ambiguity set Cα (x) ⊆ R. Intuitively, |Cα (x)| quantifies how many plausible identities remain after the attack pipeline is applied, while the finite-sample conformal guarantee ensures calibrated uncertainty under exchangeability. 6.1
Datasets and Release Mechanisms
We use datasets where each released document is associated with a ground-truth identity/profile, enabling identity linkage evaluation. For each dataset we construct (i) released documents x (e.g., anonymized text) and (ii) a candidate profile database R of candidate profiles built from non-released sources (e.g., original text, metadata, or historical documents), reflecting attacker side information. Because certified ambiguity is only meaningful relative to the candidate pool, we describe the per-dataset candidate-pool construction explicitly below; Table 8 in Appendix D maps every experimental suite to its full threat configuration.
TextWash. TextWash (Kleinberg et al., 2022) provides paired orig and anon person descriptions. We treat anon as the released document and build the candidate profiles from orig. The benchmark includes three categories (famous, semi-famous, fictional), which we interpret primarily as differing availability of external context: external web knowledge is abundant for famous individuals, limited for semi-famous, and absent for fictional entities. Candidate pool: one profile per identity in the same category, constructed from the orig description (with identity tokens removed); linkage is evaluated within-category, so the pool size equals the size of the category (≈ 400 candidate profiles). TAB (text anonymization benchmark). TAB (Pilán et al., 2022) consists of text records where anonymization may preserve substantial lexical content. We use TAB as an easy leakage regime to verify that the auditor reports near-singleton ambiguity sets when linkage is straightforward. Candidate pool: all 1,268 case records, each represented by a profile built from metadata only (K1), extracted keywords (K2), or case text (K3), with no verbatim copying of the released document into the profiles. WikiBio. WikiBio (Stranisci et al., 2023) provides biography-style text with structured fields. We use WikiBio to control profile completeness by varying which fields are exposed to the attacker. Candidate pool: 1,000 candidate profiles built from infobox-style fields or extracted keywords, audited against 500 released biographies drawn from the prepared slice. Blog Authorship Corpus. We use the Blog Authorship corpus (Schler et al., 2006) as a large-scale linkage setting. To keep experiments tractable, we subsample to a fixed number of authors and released posts. Candidate pool: one profile per author, built by concatenating W held-out posts (W ∈ {1, 3, 10}) that are disjoint from the released posts, so that W directly controls the strength of attacker side information in this suite. Candidate-pool truncation and miss rate. For expensive attacker pipelines we optionally restrict conformal scoring to the top-K retrieved candidates—an engineering choice that can violate the closed-world condition of Theorem 5.2 if the true identity is dropped. We therefore report the candidate-miss rate with both conditional coverage (given the truth is retained) and unconditional coverage, and quantify the effect of the truncation level K in the stress tests of Section 7.1. 6.2
Attack Pipelines and Knowledge Levels
Our threat model is a declared attacker pipeline that maps a released document to a ranked list or distribution over candidates, given side information such as locations, gender, and occupations (Section 3.3). We study two complementary families of attackers. Retrieval attacker. We build a text index over profiles. Given a released document x, the attacker forms a query q(x) and retrieves a ranked list of candidates.
We instantiate a knowledge ladder by varying the information available in profiles and the query construction: (i) a minimal schema subset (weak), (ii) a richer profile description including auxiliary facts (medium), and (iii) query augmentation using LLM-extracted clues from x (strong). For TextWash famous/semi-famous, we additionally consider a web-assisted attacker that can enrich the query q(x) with web search results. LLM attribute/clue attacker (sampling-based). We also evaluate an LLM that extracts attributes or salient clues from x via repeated sampling. These predictions are matched against the profiles to produce candidate scores. We treat the number of LLM samples m as a controllable attacker resource and study its impact on the resulting audit outcomes. Models. We use two open-source instructionfollowing LLMs as local attackers (Llama-3.1-8BInstruct (Grattafiori et al., 2024) and Qwen3-8BInstruct (Yang et al., 2025)), and one stronger cloud model (GPT-5). Local models are used for clue extraction and attribute prediction in the sampling-based attacker. The retrieval attacker does not require an LLM unless explicitly augmented. Metrics. We report both non-conformal ranking metrics and conformal auditing metrics. Non-conformal baselines. We report Top-1/Top-k accuracy, mean reciprocal rank (MRR), and mean rank of the true identity in the retrieved list. Conformal privacy auditing. For each α ∈ {0.01, 0.05, 0.1} we report empirical coverPN age N1 i=1 1{yi ∈ Cα (xi )}, the distribution of set sizes |Cα (x)| (median and tail percentiles), and a leakage proxy 1/ max{1, |Cα (x)|}. When the candidate pool may not contain the truth, we also report the empirical candidate-miss rate for each suite. 6.3
Protocol and Reproducibility
For each dataset and attacker configuration we use a split conformal protocol with disjoint calibration and test partitions. Because the conformal threshold q̂α is an order statistic of the calibration scores, it depends on the calibration size; we therefore treat the split sizes as part of the declared configuration rather than as fixed parameters, state them explicitly wherever attackers are compared (e.g., ncal =ntest =100 per TextWash category), and ablate the effect of ncal in Section 7.1. When the attack pipeline includes sampling (LLM attribute/clue extraction), sampling randomness is generated independently per example, with a declared per-document sampling budget m. We calibrate a separate conformal threshold per dataset and knowledge level, enabling statistically grounded comparison across threat models without re-tuning attack heuristics. Attacker-to-attacker comparisons (e.g., direct retrieval vs. LLM-assisted clues) are made only on matched calibration/test budgets, since certified set sizes are not comparable across different split sizes. Whenever the release distribution or the attacker pipeline changes—including prompt, model,
Attacker setting
Top-1
Cov.
Med. |C|
Mean |C|
Mean 1/|C|
0.052 0.973 0.994 0.715
0.929 0.973 0.994 0.970
1268 1 1 20
1179.5 1 1 20.1
9.57e-04 1.0000 1.0000 0.0499
0.003 0.872 0.964 0.615
0.932 0.953 0.964 0.945
1268 4 1 27
1173.6 3.7 1 27.1
0.0017 0.2717 1.0000 0.0369
TAB-Direct K1 meta / direct K2 keywords / direct K3 case-text / direct K2 keywords / LLM-clues TAB-Direct (Quasi-ID) K1 meta / direct K2 keywords / direct K3 case-text / direct K2 keywords / LLM-clues
Table 1: TAB linkage auditing (α = 0.05). Smaller ambiguity sets indicate higher certified identifiability. direct uses the released document as the retrieval query; LLM-clues uses an LLM to extract a short clue string from the released document before retrieval. meta/keywords/case-text describe which fields are available in the candidate profile database. All suites use split conformal calibration with disjoint calibration/test partitions; the full declared threat configuration per suite is summarized in Table 8 in Appendix D. or web-snapshot changes—we treat the result as a new attacker configuration and recalibrate before issuing certificates.
7
Experimental Results
CPA yields comparable, threat-model-specific certificates. Across tables and datasets, the empirical coverage is generally close to the target 1 − α (e.g., 0.95 when α = 0.05), supporting that CPA produces a statistically meaningful certificate that is comparable across attacker configurations rather than a single, uncalibrated, attack-specific success rate reported in isolation from the declared threat model. Release mechanism × auxiliary knowledge jointly determine certified identifiability. Table 1 shows a sharp transition in certified identifiability on TAB as the attacker’s auxiliary knowledge increases. With weak side information (K1: metadata-only profiles), CPA returns near-full ambiguity sets (median |Cα | close to the full candidate pool), certifying substantial residual uncertainty even if the attacker occasionally ranks the true identity highly. When the attacker is granted richer profile signals (K2: keyword profiles; K3: case-text profiles), the ambiguity sets collapse toward singletons, certifying that the same released documents become highly linkable under stronger auxiliary knowledge. The quasiidentifier release variant exhibits the expected mitigation effect: compared to direct release, CPA sets expand (e.g., median |Cα | ≈ 4 under K2 keywords), indicating reduced—but still non-trivial—linkability. Overall, TAB illustrates the central message of CPA: privacy risk is not an intrinsic property of the released text alone, but depends critically on the attacker’s declared side information and tooling. The miscoverage level α is a practical audit “dial” that trades certificate strength for ambiguity. Figure 2 (Appendix E) visualizes the conformal trade-off between statistical confidence and the size of the cer-
tified ambiguity set. On TAB (direct retrieval under the quasi-identifier release), tightening the guarantee (smaller α) yields larger ambiguity sets: at α=0.01, empirical coverage is 0.998 and the median set size is 19 (low leakage proxy E[1/|Cα |] ≈ 0.052). Relaxing the guarantee shrinks sets and increases the leakage proxy: at α=0.05, the median set size drops to 4 (E[1/|Cα |] ≈ 0.233), and at α=0.1 the median reaches 1 (E[1/|Cα |] ≈ 0.825). Once α becomes sufficiently large, CPA enters a near-singleton regime: ambiguity sets cannot shrink below size 1, so the conformal certificate effectively coincides with the attacker’s point prediction and empirical coverage approaches the attacker’s Top-1 accuracy. This provides a clean interpretation for auditors: smaller α certifies residual ambiguity more conservatively; larger α yields smaller certified sets at correspondingly weaker confidence. Blog Authorship shows the same qualitative “dial” behavior: with only coarse demographic-style profile fields (“W_demo”), calibration saturates (q̂α = 1) and CPA returns near-full candidate sets across α, whereas with text-enriched profiles (“W_text”, W =10 posts/author) the median |Cα | drops from 9 (α=0.01) to 7 (α=0.05) to 5 (α=0.1), with coverage decreasing accordingly as the guarantee weakens. LLM-assisted clue extraction can amplify linkage— but the effect is model-dependent and visible only with set-valued certificates. Table 2 shows that augmenting retrieval with LLM-generated clues can either reduce ambiguity sets (higher certified identifiability) or fail to help, depending on the attacker model. For WikiBio, Llama-based clue extraction increases Top-1 accuracy and shrinks the median ambiguity set relative to direct retrieval, whereas Qwen-based clues behave closer to the direct baseline. This highlights a key methodological point for LLM-era privacy auditing: different plausible attacker models can induce meaningfully different linkage risks, and CPA offers a principled way to compare them under a common calibrated out-
Table 2: Retrieval-CPA on WikiBio (500 released bios; 1k candidates) under direct query vs LLM-assisted clue extraction (α = 0.05), with split conformal calibration over disjoint calibration/test partitions of the released bios; the full declared threat configuration is summarized in Table 8 in Appendix D. Attacker setting
Top-1
Coverage
Median |Cα |
Mean |Cα |
E[1/|Cα |]
q̂α
Direct query LLM clues (Llama-3.1-8B) LLM clues (Qwen3-8B)
0.362 0.420 0.380
0.956 0.950 0.958
264 189 265
254.98 186.87 255.69
0.00466 0.00562 0.00465
0.277 0.2 0.277
put. A similar, but more nuanced, pattern appears on TextWash under matched evaluation budgets (Table 5): LLM clues reduce certified ambiguity only where external context exists, making shifts in certified residual uncertainty visible where fixed Top-k summaries would show little or no measurable change. Scale and cohort effects: narrowing the candidate pool can dominate the certificate. When the candidate pool is large, certified ambiguity sets remain large even with enriched side information (Table 4), indicating substantial residual uncertainty at the chosen confidence level. In contrast, in “motivated intruder” scenarios where the pool is implicitly narrowed (organizational, demographic, or geographic constraints), CPA sets can collapse dramatically, certifying much higher linkability. Realistic auditing should therefore report results under multiple candidate-pool assumptions, not only under a single global pool. TextWash Retrieval-CPA under matched budgets. Table 5 reports conformal privacy auditing results on TextWash at α = 0.05, with direct retrieval and LLM-clue attackers evaluated on the same capped split (ncal = ntest = 100 per category), so that certified set sizes are directly comparable across attackers. Direct retrieval yields large certified ambiguity sets (median |Cα | of 120–169.5 depending on category), indicating that even when Top-1 accuracy is moderate, the attacker still cannot be certifiably confident about a small identity set. The effect of LLM-assisted clue extraction is category- and model-dependent rather than uniform: for famous subjects, where external context exists, Llama clues shrink the certified median from 120 to 90 (Top-1 0.470 → 0.510), whereas for fictional and semi-famous subjects—where little external knowledge is available— clue extraction leaves the certified sets essentially unchanged relative to direct retrieval. We therefore do not claim that LLM clues uniformly dominate direct retrieval; rather, CPA makes visible where LLM augmentation genuinely increases certified identifiability. Empirical coverage remains close to the nominal target (≈ 0.95) across all matched runs. 7.1
Robustness under Stressed Assumptions
We now stress each assumption behind the CPA guarantee empirically, making the boundary of the certificate visible rather than only theoretical (full tables and persuite breakdowns in Appendix E).
Exchangeability and attacker/release drift. We calibrate under one configuration and test under another (Table 9). Matched configurations meet the 1−α = 0.95 target (TAB direct→direct 0.965; Blog W10→W10 0.940), while shifted configurations fall below it (TAB direct→quasi-ID 0.923; Blog W10→W1 0.835). A CPA certificate is thus calibrated for a declared threat configuration and does not automatically transfer under attacker or release drift—recalibration is the operational protocol. More robust variants (e.g., Mondrian or weighted conformal prediction) are possible when the drift model is explicit, but do not provide distributionfree transfer under arbitrary prompt, model, or snapshot drift, which remains an open problem for conformal methods. Candidate-pool miss and top-K truncation. Sweeping the truncation level K (Table 10) separates conformal miscoverage from candidate-pool failure. For TextWash fiction with direct retrieval at K = 50, the candidate-miss rate is 0.160: coverage conditional on candidate inclusion remains 1.000, while unconditional coverage drops to 0.840; at K = 200 the miss rate vanishes and unconditional coverage recovers to 0.970 (WikiBio+Llama behaves analogously). Closed-world validity (Theorem 5.2) applies conditional on candidate inclusion, and the observed degradation matches the multiplicative bound in Corollary 5.3. Alternative score constructors at matched coverage. Replacing the APS mass score with rank-based, raw-probability, or calibrated top-K constructors over cached LLM-clue queries (Table 11) shows that multiple constructors achieve valid coverage: on WikiBio+Llama at α = 0.05, APS attains coverage 0.950 (median |Cα | = 189), rank-based 0.950 (median 179), and raw probability 0.960 (median 212). CPA’s calibration layer is therefore not tied to a single score; APS is a robust default rather than uniformly dominant, and rank-based scores extend CPA to attackers that expose nothing but a ranked list of candidate identities. Calibration-set size. Varying ncal ∈ {50, 100, 150} on matched TextWash direct retrieval at α = 0.05 (Table 12) yields average coverage 0.962/0.957/0.927 with average median set sizes 151.2/148.2/140.3: coverage stays near target while sets tighten with more calibration data. Since q̂α is an order statistic of the calibration scores, certified set sizes are comparable only
at matched splits; we state the split sizes explicitly for all matched comparisons and ablations. Sampling budget and split stability. Increasing the sampling-only budget m from 1 to 10 (Llama-3.1-8BInstruct clue extraction on TextWash famous; Table 13) keeps coverage above target (0.970–0.980) while the median certified set size drops from 120 to 78: the budget affects efficiency, not validity, once recalibrated. Across three random calibration/test splits, coverage is stable at 0.951 ± 0.037 (median |Cα | 143.3 ± 33.6). Prompt, model, or decoding changes are treated as attacker-configuration drift requiring recalibration. Operational interpretation. The key advantage of CPA over fixed Top-k reporting is that it turns a ranked list into a certificate: for a chosen α, the auditor can interpret |Cα (x)| as the (calibrated) residual uncertainty remaining after the attack. This supports concrete release policies such as “withhold or further redact documents with median |Cα | < k under attacker configuration w,” and encourages tail-focused auditing of the smallest, highest-risk certified sets first.
8
Related Work
Differential privacy for training and fine-tuning language models. Differential privacy (DP) provides a formal stability-based guarantee for randomized algorithms (Dwork et al., 2006), with DP-SGD (Abadi et al., 2016) as the standard training approach, adapted to NLP fine-tuning and generation (Yu et al., 2022; Li et al., 2022). These works enforce a mechanism-level guarantee; in contrast, we audit task-level re-identification risk under a declared threat model, with finite-sample guarantees derived from split conformal calibration rather than from mechanism design or calibrated noise injection. Unintended memorization and data extraction attacks. Language models can memorize and regurgitate training data or sensitive spans, enabling extractionstyle attacks (Carlini et al., 2019, 2021, 2023), which motivates attack-based evaluation protocols such as membership inference and canary tests (Shokri et al., 2017; Yeom et al., 2018). Our approach is compatible with these attack definitions: any attack score can serve as the underlying nonconformity signal. Privacy auditing and post hoc evaluation. A parallel literature audits trained models to estimate or falsify privacy claims, including auditing of DP pipelines (Jagielski et al., 2020; Steinke et al., 2023). Hu et al. (2025) suggest that recovering a single scalar “privacy number” may be brittle across settings. Our method instead outputs a set-valued certificate with a coverage guarantee under a declared threat family. Text anonymization benchmarks and reidentification. The Text Anonymization Benchmark (TAB) provides a corpus and evaluation framework
for reducing disclosure risk in released text (Pilán et al., 2022), and recent work argues anonymization should be assessed via de-anonymization experiments with realistic adversaries (Deuber et al., 2023). We complement these efforts with a distribution-free risk-control layer: given a threat model and attacker pipeline, the reported ambiguity sets have guaranteed finite-sample coverage on held-out data. Statistical disclosure control and record linkage. Release-time disclosure risk has a long history in statistical disclosure control (SDC) and record linkage: probabilistic record linkage scores matches against a candidate list (Fellegi and Sunter, 1969); disclosure-risk estimation quantifies per-record re-identification probability (Lambert, 1993; Reiter, 2005; Hundepool et al., 2012); and k-anonymity bounds risk via a minimum cohort of indistinguishable records (Sweeney, 2002). CPA shares this candidate-list view, and its certified ambiguity set can be read as a calibrated, per-document analogue of a k-anonymity cohort under a modern, declared attacker. Unlike classical SDC, which targets structured microdata and model-based populationfrequency assumptions, CPA operates on unstructured text attacked by retrieval- and LLM-based pipelines and replaces model-based risk estimates with a distributionfree, finite-sample set-valued coverage guarantee. CPA does not replace classical SDC analysis; it wraps a declared modern attacker pipeline with a statistical certificate that classical disclosure-risk metrics by themselves do not provide for released text. Conformal prediction and risk control. Conformal prediction provides finite-sample, distribution-free uncertainty quantification under exchangeability (Vovk et al., 2005), with adaptive prediction sets controlling coverage while adjusting set size to difficulty (Romano et al., 2020) and conformal risk control extending to generic risks (Bates et al., 2021). CPA adapts this principle to privacy: the risk is re-identification under a declared threat model and the output is certified attacker ambiguity instead of a point estimate.
9
Conclusion
We studied release-time privacy risk for released text against black-box attackers that combine LLMs with auxiliary knowledge for identity linkage. Our framework, Conformal Privacy Auditing (CPA), wraps arbitrary attacker pipelines into a finite-sample privacy certificate, valid in both logit-access and sampling-only settings and thus applicable to open-source and proprietary API models alike. Empirically, CPA delivers calibrated coverage across benchmarks and shows that privacy risk is heterogeneous—a release that looks safe under limited side information can become highly identifiable under richer registries or stronger reasoning—so evaluation should favor calibrated, threat-model-specific certificates over fragile point estimates.
Limitations CPA is an auditing framework, not a cryptographic privacy definition. Its guarantees hold under the declared threat model: a finite candidate pool and an attacker pipeline family specified by the auditor, together with an exchangeability assumption between calibration and evaluation examples within each configuration. In particular, CPA does not assume the auditor can enumerate all future attackers, and it provides no universal protection against unknown or unanticipated attacks—the history of linkage attacks shows that even expert practitioners underestimate attacker ingenuity. Each certificate is valid only for the declared candidate pool, attacker pipeline, knowledge level, and calibration/test distribution; we therefore recommend running CPA across multiple plausible attacker families and recalibrating whenever attacker tooling, external knowledge, or release mechanisms change. Our drift stress tests (Section 7.1) quantify how coverage degrades when this protocol is not followed. If the candidate pool omits the true identity (open-world miss) or if deployment data drift violates exchangeability, coverage can degrade; we therefore report candidate-miss rates together with conditional and unconditional coverage, and recommend recalibration when the release distribution changes (e.g., after rewriting). More robust calibration variants (e.g., Mondrian or weighted conformal prediction under covariate shift) are possible when the drift model is explicit, but they require additional assumptions and do not provide distribution-free transfer under arbitrary prompt or model drift. Finally, this work focuses empirically on identity linkage; the same calibration layer applies to any attack that yields a scalar score or ranked candidate set, and extending CPA to membership inference is a natural direction for future work. These limitations are not unique to CPA but reflect the fundamental challenge of giving operationally meaningful, release-time guarantees for natural language under evolving adversaries.
uating and testing unintended memorization in neural networks. In 28th USENIX Security Symposium (USENIX Security). Nicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Úlfar Erlingsson, Alina Oprea, and Colin Raffel. 2021. Extracting training data from large language models. In 30th USENIX Security Symposium (USENIX Security). Dominic Deuber, Michael Keuchen, and Nicolas Christin. 2023. Assessing anonymity techniques employed in german court decisions: A deanonymization experiment. In 32nd USENIX Security Symposium (USENIX Security 23). Yao Dou, Isadora Krsek, Tarek Naous, Anubha Kabra, Sauvik Das, Alan Ritter, and Wei Xu. 2024. Reducing privacy risks in online self-disclosures with language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 13732– 13754, Bangkok, Thailand. Association for Computational Linguistics. Cynthia Dwork. 2006. Differential privacy. In International Colloquium on Automata, Languages, and Programming (ICALP), pages 1–12. Springer. Cynthia Dwork, Frank McSherry, Kobbi Nissim, and Adam Smith. 2006. Calibrating noise to sensitivity in private data analysis. In Theory of cryptography conference, pages 265–284. Springer. Ivan P Fellegi and Alan B Sunter. 1969. A theory for record linkage. Journal of the American Statistical Association, 64(328):1183–1210. Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, and 1 others. 2024. The Llama 3 herd of models. arXiv preprint arXiv:2407.21783.
Martin Abadi, Andy Chu, Ian Goodfellow, H. Brendan McMahan, Ilya Mironov, Kunal Talwar, and Li Zhang. 2016. Deep learning with differential privacy. In Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security (CCS).
Yuzheng Hu, Fan Wu, Ruicheng Xian, Yuhang Liu, Lydia Zakynthinou, Pritish Kamath, Chiyuan Zhang, and David Forsyth. 2025. Empirical privacy variance. In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, Vancouver, Canada. PMLR.
Stephen Bates, Anastasios Angelopoulos, Lihua Lei, Jitendra Malik, and Michael I. Jordan. 2021. Distribution-free, risk-controlling prediction sets. Journal of the ACM, 68(6):1–34.
Anco Hundepool, Josep Domingo-Ferrer, Luisa Franconi, Sarah Giessing, Eric Schulte Nordholt, Keith Spicer, and Peter-Paul De Wolf. 2012. Statistical Disclosure Control. John Wiley & Sons.
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramèr, and Chiyuan Zhang. 2023. Quantifying memorization across neural language models. In International Conference on Learning Representations (ICLR).
Information and Privacy Commissioner of Ontario. 2016. De-identification guidelines for structured data. Toronto, ON, Canada.
References
Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, and Dawn Song. 2019. The secret sharer: Eval-
Matthew Jagielski, Jonathan Ullman, and Alina Oprea. 2020. Auditing differentially private machine learning: How private is private SGD? In Advances in Neural Information Processing Systems (NeurIPS).
Bennett Kleinberg, Toby Davies, and Maximilian Mozes. 2022. Textwash—an automated opensource text anonymisation tool. arXiv preprint arXiv:2208.13081. Diane Lambert. 1993. Measures of disclosure risk and harm. Journal of Official Statistics, 9(2):313–331. Xuechen Li, Florian Tramèr, Percy Liang, and Tatsunori Hashimoto. 2022. Large language models can be strong differentially private learners. In International Conference on Learning Representations (ICLR). Pierre Lison, Ildikó Pilán, David Sánchez, Montserrat Batet, and Lilja Øvrelid. 2021. Anonymisation models for text data: State of the art, challenges and future directions. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4188–4203. Benet Manzanares-Salor, David Sánchez, and Pierre Lison. 2024. Evaluating the disclosure risk of anonymized documents via a machine learning-based re-identification attack. Data Mining and Knowledge Discovery, 38(6):4040–4075. Ildikó Pilán, Pierre Lison, Lilja Øvrelid, Anthi Papadopoulou, David Sánchez, and Montserrat Batet. 2022. The Text Anonymization Benchmark (TAB): A dedicated corpus and evaluation framework for text anonymization. Computational Linguistics, 48(4):1053–1101. Jerome P Reiter. 2005. Estimating risks of identification disclosure in microdata. Journal of the American Statistical Association, 100(472):1103–1112. Yaniv Romano, Matteo Sesia, and Emmanuel Candès. 2020. Classification with valid and adaptive coverage. In Advances in Neural Information Processing Systems (NeurIPS). Jonathan Schler, Moshe Koppel, Shlomo Argamon, and James W. Pennebaker. 2006. Effects of age and gender on blogging. In AAAI Spring Symposium: Computational Approaches to Analyzing Weblogs, pages 199–205. Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov. 2017. Membership inference attacks against machine learning models. In IEEE Symposium on Security and Privacy. Robin Staab, Mark Vero, Mislav Balunović, and Martin Vechev. 2024. Beyond memorization: Violating privacy via inference with large language models. In International Conference on Learning Representations (ICLR). Thomas Steinke, Milad Nasr, and Matthew Jagielski. 2023. Privacy auditing with one (1) training run. In Advances in Neural Information Processing Systems (NeurIPS).
Marco Antonio Stranisci, Rossana Damiano, Enrico Mensa, Viviana Patti, Daniele Radicioni, and Tommaso Caselli. 2023. WikiBio: A semantic resource for the intersectional analysis of biographical events. arXiv preprint arXiv:2306.09505. Latanya Sweeney. 2002. k-anonymity: A model for protecting privacy. International Journal of Uncertainty, Fuzziness and Knowledge-Based Systems, 10(5):557– 570. Vladimir Vovk, Alex Gammerman, and Glenn Shafer. 2005. Algorithmic Learning in a Random World. Springer. Rui Xin, Niloofar Mireshghallah, Shuyue Stella Li, Michael Duan, Hyunwoo Kim, Yejin Choi, Yulia Tsvetkov, Sewoong Oh, and Pang Wei Koh. 2025. A false sense of privacy: Evaluating textual data sanitization beyond surface-level privacy leakage. arXiv preprint arXiv:2504.21035. An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, and 1 others. 2025. Qwen3 technical report. arXiv preprint arXiv:2505.09388. Samuel Yeom, Irene Giacomelli, Matt Fredrikson, and Somesh Jha. 2018. Privacy risk in machine learning: Analyzing the connection to overfitting. In IEEE 31st Computer Security Foundations Symposium (CSF). Da Yu, Saurabh Naik, Arturs Backurs, Sivakanth Gopi, Huseyin A. Inan, Gautam Kamath, Janardhan Kulkarni, Yin Tat Lee, Andre Manoel, Lukas Wutschitz, Sergey Yekhanin, and Huishuai Zhang. 2022. Differentially private fine-tuning of language models. In International Conference on Learning Representations (ICLR).
A
Appendix
A.1
CPA Algorithm
The full CPA procedure is shown in Algorithm 1. Algorithm 1 Split Conformal Privacy Auditing (CPA) Require: Calibration set Dcal = {(Xi , Yi )}ni=1 under fixed knowledge level w; candidate-pool generator Y(·, w); attacker pipeline A; miscoverage level α ∈ (0, 1); sampling budget m (used only in samplingonly access). Ensure: Conformal threshold q̂α and ambiguity-set function Cα (·, w). 1: Score function (APS mass). For any (x, y) define X s(x, y; w) = pA (y ′ | x, w). y ′ ∈Y(x,w): pA (y ′ |x,w)≥pA (y|x,w)
2: Calibration scores. 3: for i = 1 to n do 4: Construct candidate pool Yi ← Y(Xi , w). 5: Obtain a distribution p̃i (·) ≈ pA (· | Xi , w) over
Yi : if logits/logprobs available then p̃i ← pA (· | Xi , w) restricted/normalized to Yi . 8: else {sampling-only access} (1) (m) 9: Draw Ỹi , . . . , Ỹi ∼ A(· | Xi , w) over Yi . Pm (j) 1 10: Set p̃i (y) ← m = y} for y ∈ j=1 1{Ỹi Yi . 11: end if 12: Compute calibration score Ti ← P ′ p̃ (y ). ′ ′ i y ∈Yi : p̃i (y )≥p̃i (Yi ) 13: end for 14: Sort T(1) ≤ · · · ≤ T(n) and set 6: 7:
k ← ⌈(n + 1)(1 − α)⌉ ,
q̂α ← T(k) .
15: Define ambiguity set for a new document (x, w). 16: Construct Y(x, w) and obtain p̃(·) ≈ pA (· | x, w)
as above. 17: Compute plug-in scores s̃(y) for all y ∈ Y(x, w) as
in Step 1 (with p̃ in place of pA ), and output Cα (x, w) ← y ∈ Y(x, w) : s̃(y) ≤ q̂α . 18: leakage
proxy: 1/ max{1, |Cα (x, w)|}.
B
L(x, w)
←
Dataset Statistics
Table 3 summarizes the datasets used in our privacyauditing experiments, together with their split statistics. The TAB benchmark contains 1014 training documents, 127 development documents, and 127 test documents. These records contain long judicial narratives,
which is reflected in the substantially larger average document length relative to the other datasets. The TextWash benchmark contains 1202 original and 1202 released descriptions, with category counts of 401 famous, 400 semi-famous, and 401 fiction/other examples. Average document length is approximately 130 words before release and 128 words after release. The Blog authorship corpus is stored as 19320 authorlevel XML files. Parsing those XML files yields an average of 9.74 posts per author and a median of 6 posts per author. This corpus is substantially larger in author count than the other datasets and is used to study how CPA behaves when the attacker registry is constructed from varying numbers of posts per author. The WikiBio experiments use the prepared audit slice emitted by the dataset-preparation script. It contains 1000 released biographies with 1000 unique identities, and the matching profile database contains 1000 profile rows. The average original and released biography lengths are 102.7 and 101.9 words respectively, reflecting the lighter editing typical of the prepared WikiBio audit slice used in our experiments. B.1
Implementation Details
Retrieval attacker (non-LLM and LLM-assisted). In Retrieval-CPA, the attacker maps a released document x to a probability distribution over a candidate pool by: (i) building candidate profile texts for each identity in the auxiliary database, (ii) embedding them with a bag-of-words TF–IDF vectorizer, and (iii) ranking candidates by cosine similarity to a query string derived from x. We convert similarity scores to a pseudoposterior distribution via a temperature-scaled softmax (parameter –softmax_temperature). Direct query (–query_mode direct) uses the released text itself as the query string. LLM clues (–query_mode direct_plus_clues) uses an LLM to extract a short list of missing/high-value clues from the released text (optionally augmented by web search), and appends them to the query string. This is implemented in adversary/clue_extract.py and parsed via linkage/clues.py within privacy_audit_cp. In both cases, we can optionally restrict to a top-K candidate pool (–top_k_pool) to reduce runtime while preserving a well-defined candidate set for the split conformal calibration performed at audit time. Profile attacker (LLM attribute extraction + candidate filtering). In Profile-CPA, the LLM adversary does not see (and is not prompted with) the full candidate list. Instead, it predicts attributes from the released text (e.g., gender, age-range, occupation, location, or dataset-specific fields) by sampling m times per attribute. The sampled attribute values induce empirical distributions p̂m (· | x). For each attribute, we conformalize these distributions to produce a certified set of plausible attribute values; the final identity ambiguity set is formed by filtering the auxiliary profile database to candidates consistent with all certified attribute sets.
Dataset TAB (ECHR) TextWash Blog Authorship WikiBio
Unit
Total
Split / subset
Avg. words
documents biographies authors biographies
1,268 1,202 orig + 1,202 anon 19,320 1,000
train/dev/test = 1014/127/127 famous 401, semifamous 400, fiction/other 401 balanced male/female author roster one category in the prepared audit slice
train 1343.4; dev 874.6; test 843.4 orig 130.5; anon 127.7 9.74 posts/author orig 102.7; anon 101.9
Table 3: Dataset statistics for the release-audit experiments. For TAB we report split sizes from the benchmark JSON files; for TextWash we report file counts from the original and anonymized person-description trees; for Blog we report author-level XML statistics; and for WikiBio we report the prepared audit slice used by the canonical experiment runner in all reported WikiBio suites and ablations. Conformal calibration and outputs. Both RetrievalCPA and Profile-CPA use split conformal calibration: we compute a scalar nonconformity score for the true identity (or true attribute value) on a calibration split and take an order statistic threshold q̂α . At test time we output an ambiguity set Cα (x) consisting of candidates whose score is below q̂α . Per run we report: (i) empirical coverage, (ii) distribution of |Cα (x)| (median/mean/min/max), (iii) leakage proxy E[1/ max{1, |Cα (x)|}], and (iv) diagnostic rates such as empty-set and candidate-miss (when applicable). For reproducibility and debugging, the runner scripts write jsonl files for both calibration and test splits containing the query string, top-K ranked candidates, conformal scores, and the final ambiguity set for every audited document in both conformal splits. Randomness and the Ui variables. Whenever an attacker uses sampling (e.g., m LLM generations for clue extraction or attribute extraction), the randomness is represented abstractly as a per-example seed/noise variable Ui . In code, Ui corresponds to the randomness consumed by stochastic decoding (and any parsing postprocessing). Conditional on (Xi , Wi ) and the fixed prompt/decoding settings, the resulting score Ti is a deterministic function of (Xi , Yi , Wi , Ui ), which is exactly the condition needed by the exchangeability-based conformal validity argument in Section 5. Default hyperparameters. Unless otherwise stated, we use: TF–IDF with max_features=50,000, ngram_max=1, and dataset-dependent min_df; softmax temperature in [0.5, 2.0]; LLM decoding with temperature=0.8, top_p=0.95; and m ∈ {5, 10, 30} samples depending on the ablation. We use fixed random seeds for dataset splits and sampling (whenever seeding is supported by the inference backend). Blog Authorship: attacker side-information via heldout posts. Table 4 reports Retrieval-CPA on Blog Authorship under a controlled increase in attacker side information. Each released document is a post whose true identity is its author. To model auxiliary knowledge, we build a candidate database of author profiles by concatenating W held-out posts per author (W ∈ {1, 3, 10}), and run a retrieval attacker that ranks authors by text similarity between the released post and each author profile. We then apply split conformal calibration at α = 0.05 to convert the attacker ranking into a certified ambigu-
ity set Cα (x). As W increases, the attacker becomes stronger: Top-1 accuracy rises (0.087 → 0.209) and the conformal threshold q̂α decreases (0.948 → 0.763), yielding smaller certified ambiguity sets (median |Cα | drops from 222 to 176). Importantly, empirical coverage remains near the nominal 1 − α ≈ 0.95, indicating that the certificates remain valid while exposing how additional auxiliary text reduces certified uncertainty. Despite this trend, the certified sets remain large in absolute terms, suggesting substantial residual ambiguity in a large candidate pool even when the attacker has multiple posts per author at its disposal.
C
Notation
Table 6 consolidates the symbols used in Sections 3–5.
D
Threat-Model Configurations per Experimental Suite
Table 7 summarizes the knowledge ladder used to instantiate attacker side information, and Table 8 maps each experimental suite to its declared threat configuration (Section 3.3): candidate-pool construction, query mode, LLM/web augmentation, and closed- vs. open-world assumption. The conformal split sizes and sampling budgets of the matched and ablation suites are stated in the corresponding table captions (Tables 5, 12, and 13). Each CPA certificate is valid for exactly one row of this table; changing any entry defines a new attacker configuration that requires recalibration.
E
Robustness and Ablation Studies
This section reports the stress tests and ablations summarized in Section 7.1, together with the α-sweep visualization for the TAB quasi-identifier suite (Figure 2). E.1
Exchangeability and Attacker/Release Drift
Table 9 calibrates CPA under one configuration and evaluates it under another. For Blog, W =k denotes an attacker registry built from k known posts per author, so W10 is a stronger author-profile registry than W1. Matched configurations meet the 1 − α = 0.95 target, while shifted configurations fall below target, confirming that certificates do not transfer across attacker or release drift and must be recalibrated.
Table 4: Retrieval-CPA on Blog Authorship under increasing attacker side information (number of posts per author; α = 0.05), with split conformal calibration over disjoint calibration/test partitions of the released posts; the full declared threat configuration is summarized in Table 8 in Appendix D. Attacker setting
Top-1
Coverage
Median |Cα |
Mean |Cα |
E[1/|Cα |]
q̂α
W=1 posts/author W=3 posts/author W=10 posts/author
0.087 0.143 0.209
0.946 0.956 0.946
222 210 176
200.50 191.82 147.14
0.00566 0.00604 0.00677
0.948 0.903 0.763
Figure 2: Effect of the miscoverage level α on TAB (direct retrieval, quasi-identifier release): mean leakage proxy E[1/|Cα |] (left), empirical coverage (center), and median certified set size |Cα | (right). Smaller α certifies residual ambiguity more conservatively; larger α yields smaller sets with weaker statistical protection. Table 5: TextWash Retrieval-CPA under matched budgets. Retrieval-CPA on TextWash under direct retrieval vs. LLM-assisted clue extraction (α = 0.05). All attackers are evaluated on the same capped split with ncal = ntest = 100 per category, so certified set sizes are directly comparable across attackers; LLM rows use cached clue outputs throughout. Cov. Med. |Cα | Mean |Cα |
Attacker
Top-1
fiction Direct query GPT-5 clues Llama clues
0.500 0.970 0.500 0.970 0.600 0.970
169.5 169.5 171.0
172.44 172.44 174.26
semifamous Direct query 0.380 0.910 GPT-5 clues 0.380 0.910
155.0 155.0
167.12 167.12
famous Direct query Llama clues
120.0 90.0
119.88 89.70
0.470 0.990 0.510 0.960
Smaller |Cα | implies a more identifying attack (less ambiguity). Coverage should be near 1 − α = 0.95 under exchangeability, up to finite-sample fluctuation.
E.2
Candidate-Pool Miss and Top-K Truncation
Table 10 sweeps the top-K truncation level, reporting the candidate-miss rate, coverage conditional on the true identity being retained in the pool, and unconditional coverage. Conditional coverage stays at or above target whenever the true identity is retained, while unconditional coverage degrades with the miss rate, exactly as predicted by Corollary 5.3. This separates conformal miscoverage from candidate-pool construction failure:
top-K truncation is an engineering choice, and closedworld validity applies conditional on candidate inclusion in the truncated candidate pool. E.3
Alternative Score Constructors
Table 11 compares nonconformity-score constructors at matched coverage over cached LLM-clue queries (no new API calls): the APS mass score, a rank-based score, and the raw attacker probability. Multiple constructors achieve valid coverage, confirming that the CPA calibration layer is not tied to a single score; APS is a robust default rather than uniformly dominant, and the rankbased variant extends CPA to attackers that return only ranked candidate lists without scores. E.4
Calibration-Set Size
Table 12 varies the calibration size ncal on matched TextWash direct retrieval (α = 0.05, ntest = 100 per category). Coverage remains near target while the certified sets tighten with more calibration data; per-category behavior (e.g., famous coverage 0.980 → 0.940 with median 124 → 97 as ncal grows from 50 to 150, while fiction remains at 0.980/0.970/0.980) shows the expected finite-sample variability of the order-statistic threshold. Because q̂α depends on ncal through the split-conformal order statistic, the split sizes are part of the declared threat configuration and are stated with every matched comparison and ablation. E.5
Sampling Budget and Split Stability
Table 13 reports a local Llama-3.1-8B-Instruct clueextraction ablation on TextWash famous (direct-plusclues, ncal = ntest = 100, candidate-miss rate 0) over the per-document sampling budget m. Increasing m
Symbol
Meaning
x, X
Released document (possibly anonymized/rewritten) y, Y True identity of the document subject U Reference population of identities Attacker knowledge level / threat conw, W figuration (side information and tooling) Y(x, w) Finite candidate pool available to the attacker K Top-K truncation level of the candidate pool (engineering choice) pA (· | x, w) Attacker’s distribution (or normalized ranking) over Y(x, w) m Per-document sampling budget in the sampling-only interface p̂m Empirical attacker distribution from m forced-choice samples s(x, y; w), ŝm APS mass nonconformity score (exact / plug-in) Ti Scalar calibration score of example i α Target miscoverage level Split-conformal threshold (⌈(n + q̂α 1)(1 − α)⌉-th order statistic) ncal , ntest Calibration / test split sizes Cα (x, w) Certified ambiguity set (primary certificate) L(x, w) Inverse ambiguity proxy 1/ max{1, |Cα (x, w)|} (not a reidentification probability) ρ(w) Candidate-miss probability Pr(Y ∈ / Y(X, w)) Attacker-internal randomness (samUi pling, decoding) for example i
Table 6: Notation used throughout the paper. from 1 to 10 preserves above-target coverage while reducing the median certified set size from 120 to 78: the sampling budget affects efficiency after recalibration, not validity. Across three random calibration/test splits of TextWash direct retrieval (nine category-byseed rows), coverage is 0.951 ± 0.037 with median |Cα | 143.3 ± 33.6 (mean 146.5 ± 35.2), indicating stability to the split randomness. We treat prompt, model, or decoding changes as attacker-configuration drift requiring recalibration (Table 9); we do not compare or interpret certificates across such configuration changes.
Table 7: Knowledge ladder (side information W ) used to instantiate attacker pipelines in our experiments. Level K1 (weak)
Attacker access / capability
Profiles contain only coarse structured fields (e.g., occupation bucket, coarse location, age range). Retrieval query uses released text directly. K2 (medium) Profiles add additional non-identifying evidence facts (e.g., coarsened events/keywords). Retrieval uses released text; no LLM assistance. K3 (strong) LLM-assisted clue extraction from released text produces a query expansion (optional), and/or profile is augmented (RAG/web search).
Suite
Candidate pool (size) Query mode LLM/web aug.
World
TAB K1–K3 TAB K2/K3 + clues WikiBio direct WikiBio + clues Blog W∈ {1, 3, 10} TextWash direct (matched) TextWash + clues (matched) TextWash m-ablation
case records (1268) case records (1268) profile rows (1000) top-K=200 of 1000 author registry category (≈400) category (≈400) category (≈400)
closed closed closed top-K truncated closed closed closed closed
direct direct+clues direct direct+clues direct direct direct+clues direct+clues
none LLM clues none LLM clues none none LLM clues LLM clues (local HF)
Table 8: Declared threat configuration for each experimental suite. “direct” uses the released text as the retrieval query; “direct+clues” appends LLM-extracted clues. The matched TextWash suites use ncal =ntest =100 per category (Table 5), and the m-ablation sweeps the per-document sampling budget m ∈ {1, 3, 5, 10} (Table 13).
Cal. → test
Cov. Med. |Cα | Mean |Cα |
TAB direct → direct 0.965 TAB direct → quasi-ID 0.923 TAB quasi-ID → quasi-ID 0.965 Blog W10 → W10 0.940 Blog W10 → W1 0.835
1 1 3 176 176
1.000 1.000 2.658 175.691 175.224
Table 9: Exchangeability stress tests (α = 0.05). Matched calibration/test configurations achieve target coverage; shifted configurations under-cover, so recalibration is the operational protocol under suspected attacker-side or release-side drift.
Suite
K
Avg. coverage
Avg. median |Cα |
50 100 150
0.962 0.957 0.927
151.2 148.2 140.3
Table 12: Calibration-set-size ablation on matched TextWash direct retrieval (α = 0.05, ntest = 100), averaged over the three TextWash categories.
Miss Cond. Uncond. Med.
TW fiction direct 50 0.160 TW fiction direct 100 0.080 TW fiction direct 200 0.000 TW famous direct 50 0.110 TW famous direct 100 0.030 WB + Llama clues 50 0.118 WB + Llama clues 200 0.044
1.000 1.000 0.970 1.000 1.000 1.000 0.994
0.840 50 0.920 100 0.970 169.5 0.890 50 0.970 100 0.882 50 0.950 182
Table 10: Candidate-pool / top-K truncation stress tests (α = 0.05). TW = TextWash, WB = WikiBio. “Miss” is the empirical candidate-miss rate; “Cond.”/“Uncond.” are coverage conditional on the true identity being retained in the truncated pool and unconditional coverage; “Med.” is the median certified set size |Cα | over the evaluated test split.
Attacker
ncal
Score
WikiBio + Llama APS WikiBio + Llama rank WikiBio + Llama prob. Blog W10 + Llama APS Blog W10 + Llama rank TAB K3 + clues APS TAB K3 + clues prob.
Cov. Med. 0.950 0.950 0.960 0.950 0.962 0.967 0.951
Mean
189 186.870 179 179.000 212 209.434 8 8.341 9 9.000 31 30.803 19 24.465
Table 11: Score-constructor ablation for cached LLMclue attackers (α = 0.05). “Med.”/“Mean” refer to |Cα |. All constructors achieve valid coverage; efficiency (certified set size) varies moderately across constructors and across the different attacker pipelines.
m
Top-1
Cov.
Med. |Cα |
Mean |Cα |
1 3 5 10
0.490 0.510 0.510 0.490
0.980 0.970 0.970 0.970
120 83 82 78
120.020 82.590 82.110 78.060
Table 13: Sampling-budget ablation with Llama-3.18B-Instruct on TextWash famous (α = 0.05, ncal = ntest = 100, candidate-miss rate 0). Coverage stays above target while larger budgets yield tighter certified sets at the same nominal coverage level.