ConceptioArchivearXiv CS
arXiv CSopen access

PromptGraph: Graph-Guided Prompt Sanitization for Balancing Privacy and Utility in LLM Inference

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
cryptography, security, privacy, cybersecurity

1

PromptGraph: Graph-Guided Prompt Sanitization for Balancing Privacy and Utility in LLM Inference

arXiv:2607.10709v1 [cs.CR] 12 Jul 2026

Chen Gu, Hui Wan, Donghui Hu, Hui Wang, and Zhuoer Gu

Abstract—Large Language Model (LLM) services introduce a fundamental privacy challenge. Sensitive information may be inferred not only from explicit identifiers, such as names or phone numbers, but also from contextual associations among otherwise innocuous spans. Existing sanitizers typically assign privacy or utility signals to individual spans without explicitly modeling pairwise relationships among them. In this paper, we propose PromptGraph, a graph-guided prompt-sanitization approach for privacy-preserving LLM inference. PromptGraph estimates privacy leakage at the span level and utility-relevant contextual dependencies between pairs of spans. It represents each prompt as an attributed graph, in which nodes carry spanlevel privacy scores and edges encode contextual dependencies needed to preserve utility. The sanitization objective selects a protected span set that maximizes privacy gain while penalizing the loss of contextual dependencies. This formulation explicitly balances privacy and utility when contextual evidence is hidden. Protected spans are sanitized locally, and returned placeholders are restored only after passing local consistency checks. We conduct extensive experiments showing that PromptGraph achieves a more favorable balance between privacy and utility than prompt-privacy baselines.

Local User Prompt Alice asks whether fasting log and dosage adjustment should be included in the treatment plan.

Cloud LLM Masking

dosage treatment fasting adjustment plan log Privacy Cue Attribute Task Relation

Alice

Observed issue Independent span scores overlook how context couples privacy and utility.

Under-sanitization [MASK] asks whether fasting log and dosage adjustment should be included in the treatment plan. Health status remains inferable

Over-sanitization [MASK] asks whether [MASK] and [MASK] should be included in [MASK]. Intent Underspecified

Fig. 1: Context-coupled privacy risks of independent span sanitization. Under-sanitization may mask an explicit value while leaving contextual cues that support sensitive-attribute inference, whereas over-sanitization may remove task-relevant relations and leave the prompt underspecified for downstream LLM inference.

setting under differential privacy. Closer to our setting, ProSan [8] balances privacy leakage risk with word importance, Index Terms—Prompt sanitization, privacy-preserving inference, while ALSA [9] considers privacy leakage risk, contextual large language models, graph-guided selection. information importance, and task relevance when assigning anonymization actions. These studies suggest that lightweight I. I NTRODUCTION local sanitization can be integrated into online LLM workflows. Cloud Large Language Models (LLMs) are widely deployed Despite this progress, existing prompt sanitization methods through a prompt–response workflow, in which users submit still do not explicitly model how privacy leakage and task prompts to remote servers and receive generated outputs [1]. utility are coupled through contextual relations. As illustrated Although this paradigm makes foundation models broadly in Figure 1, independent span decisions may hide an explicit accessible, it requires prompt content to be disclosed during sensitive value while leaving nearby cues that support recovery each remote inference request [2]. Such prompts may contain or attribute inference. They may also remove contextual sensitive information, including personal attributes and account relations that downstream inference needs. Existing methods credentials. Consequently, prompts transmitted to remote LLM such as ProSan [8] and ALSA [9] mainly attach importance or servers may expose users’ private information. relevance signals to individual words or spans, rather than to Prior work has explored several strategies for reducing the relations among them. In ALSA, for example, contextual privacy leakage in prompts. Differentially private text sani- information importance is summarized as a word-level feature tization [3], [4] perturbs tokens or representations to limit and combined with privacy and task relevance in a clusteringdisclosure. However, these classical perturbation methods are based action rule. Such designs do not identify which visible not primarily designed for prompt workflows that require spans provide evidence for recovering a sensitive or hidden local restoration after generation. Prompt sanitization methods span, nor do they score the utility loss caused by removing a instead protect the input before cloud inference. HaS [5] hides particular relation. The key problem is therefore how to select private entities locally and restores anonymized responses after prompt content for protection while reducing contextual generation. InferDPT [6] studies privacy-preserving inference privacy risk and preserving relations needed for downstream for black-box LLMs through perturbation and extraction. DP- inference. OPT [7] investigates a related but distinct prompt optimization In this paper, we propose PromptGraph, a graph-guided prompt sanitization approach that hides evidence of private Chen Gu, Hui Wan, Donghui Hu, and Hui Wang are with the School of Computer Science and Information Engineering, Hefei University of Tech- information while preserving contextual relations. It first nology, Hefei, China (e-mail: [email protected]; [email protected]; estimates span privacy from local patterns and counterfactual [email protected]; [email protected]). masking, so contextual privacy cues are handled as node-level Zhuoer Gu is with International College Beijing, China Agricultural University, China (e-mail: [email protected]). privacy evidence. It then estimates contextual dependency for

2

utility preservation by checking how much one span supports information. In-context learning research demonstrates that local reconstruction of another. This edge-level dependency demonstration retrieval [10], prompt format, example selection, goes beyond span relevance by making utility-relevant pairwise and example ordering can all substantially affect downstream relations explicit in the selection objective. The scored spans performance [11]. Studies on long-context models [12] further and retained pairwise relations are organized as a prompt graph reveal that LLMs do not utilize contextual evidence uniformly, for selection. The approach selects protected spans through with performance often depending on the position of relevant a principled trade-off between privacy gain and contextual information within the prompt. While these findings highlight utility loss. Protected spans are locally sanitized, and returned the importance of contextual evidence for task performance, placeholders are restored only when they pass local consistency they do not address how such information should be protected checks. This graph-guided selection protects high-risk spans or preserved during prompt sanitization. Structured prompting while avoiding unnecessary removal of contextual relations provides another perspective on how contextual dependencies needed for downstream tasks. The novelty of PromptGraph can be represented. Prior prompting methods elicit intermediate lies in reframing prompt sanitization from isolated span scoring reasoning steps in explicit chains or organize reasoning states as relation-aware graph selection, enabling a more explicit as graphs [13], [14], showing that structure can support more balance between privacy protection and task utility. flexible LLM reasoning. However, these structures are designed Our contributions are summarized as follows. for reasoning rather than for selecting which prompt content should be hidden before cloud inference. • We propose PromptGraph, a graph-guided prompt sanitization approach that protects sensitive prompt content while preserving contextual relations needed for downIII. P RELIMINARIES stream inference. A. Counterfactual Masking • The approach first segments a prompt into textual spans, Counterfactual analysis [15] estimates the effect of an estimates span privacy risk and contextual dependency input component by comparing a score before and after for utility preservation, and organizes the scored spans a controlled intervention. Let x = [u1 , u2 , . . . , un ] denote and relations as an attributed prompt graph. It then selects an input represented as a sequence of components, and let protected spans with an edge-aware objective and restores T ⊆ {1, . . . , n} be an intervention set. A masked counterfactual returned placeholders only after local consistency checks. input is defined as: • We conduct extensive experiments on three task datasets, ( three mainstream LLMs, and two server-side attack eval[MASK], k ∈ T, x\T = [ū1 , . . . , ūn ], ūk = (1) uations, demonstrating a more favorable balance between uk , k∈ / T. privacy and utility than baselines. For a scoring function H(·), the positive counterfactual II. R ELATED W ORK contribution of component ui can be written as: A. Privacy-Preserving Prompt Solutions ∆H (2) i = [H(x) − H(x\{i} )]+ , Privacy-preserving prompt solutions aim to protect sensitive information during cloud LLM inference while preserving the where [·]+ = max(·, 0) retains only positive score reductions utility of sanitized prompts for downstream tasks. Practical caused by masking. This notation quantifies the marginal effect approaches typically anonymize prompts locally and restore of an intervened component on the output of an estimator, protected content after generation. HaS [5] replaces sensitive without specifying how components are selected or how the entities before inference and restores them in the returned estimator is implemented. response. Another line of work reduces disclosure through calibrated perturbations, often under differential privacy guar- B. Threat Model antees. Early text sanitization methods [3], [4] perturb tokens We consider black-box LLM inference with a local sanitizer. or representations to limit information leakage. InferDPT [6] Given an original prompt q, the client sends only the sanitized investigates privacy-preserving inference for black-box LLMs prompt q̃ to the cloud model, receives a sanitized response r̃, through perturbation and extraction, while DP-OPT [7] applies and locally restores the placeholders using a private restoration differential privacy to prompt optimization rather than direct table to obtain the final response r. prompt sanitization. More recent methods aim to balance The cloud service is honest but curious and can observe privacy and utility. ProSan [8] balances privacy leakage risk only artifacts visible to the server: the sanitized prompt q̃, against word importance, and ALSA [9] additionally incorpo- the raw response r̃, and the public sanitization procedure. It rates contextual importance and task relevance when selecting does not observe the original prompt q, the restoration table, anonymization actions. Nevertheless, existing approaches still local detector scores, or restoration decisions on the client. make protection decisions primarily at the span level, even Under this view, we consider two risks on the server side. A when contextual or task-aware signals are considered. recovery attack attempts to reconstruct hidden sensitive values B. Contextual Evidence and Structured Prompting Recent studies have shown that LLM behavior is strongly influenced by both the content and organization of contextual

from visible context. An attribute inference attack attempts to infer sensitive attribute types from residual cues even when the original values remain hidden. We exclude client compromise and attacks using records specific to the user.

3

C. Problem Formulation Let q = [u1 , u2 , . . . , un ] be a prompt segmented into textual spans. A span may reveal sensitive information either directly through its surface form or indirectly through its contextual relations with other spans. Prompt sanitization aims to construct a sanitized prompt q̃ that reduces such evidence while preserving the information required for downstream inference. We formulate this problem over a prompt graph: G(q) = (V, E, P, I),

(3)

attribute exposure from contextual cues. In practice, G can be implemented as an ensemble of lightweight attribute classifiers and cue detectors on the client. This local estimator enables counterfactual assessment of contextual evidence. Beyond explicitly sensitive strings, otherwise benign spans may become privacy sensitive through their contextual associations. For each span ui , we construct a counterfactual prompt q\i by replacing ui with a designated placeholder. The placeholder preserves the span position while withholding its original content. We then estimate the contribution of ui by comparing attribute evidence before and after masking:   Ci = max G(a | q) − G(a | q\i ) + , (4)

where V = {vi }ni=1 and each node vi corresponds to span ui . The edge set E ⊆ {(vi , vj ) : 1 ≤ i < j ≤ n} contains retained a∈A contextual relations, P stores span privacy scores Pi , and I where [·] retains only positive reductions in attribute evidence. + stores utility-oriented dependency scores Iij . A protected set Intuitively, a span receives a high contextual score when maskS ⊆ V induces a sanitized prompt q̃. Because replacing either ing it substantially reduces evidence for inferring a sensitive endpoint makes the original pairwise context unavailable to attribute. Consequently, spans without explicit sensitive strings the cloud, every edge incident to S is affected. We therefore may still receive high privacy scores if they make sensitive select S to maximize privacy gain while minimizing the total attributes easier to infer. Because both Di and Ci are on the weight of affected dependencies. same normalized scale, the final span privacy score combines both sources conservatively: IV. M ETHODOLOGY Pi = max(Di , Ci ).

A. Overview

(5)

The overview of PromptGraph is shown in Figure 2. The maximum operator protects explicitly sensitive spans It operates on the client side for prompt sanitization before even when the contextual estimator is uncertain, while still inference and response restoration after inference. Given an identifying ordinary spans whose removal substantially reduces original prompt, the local sanitizer first segments it into sensitive attribute evidence. textual spans, estimates span privacy leakage risk, and scores contextual dependency for utility preservation for retained pairs. C. Pairwise Dependency Scoring These scored spans and pairwise relations are then organized as For each retained candidate pair (ui , uj ), we assign a a prompt graph that guides protection selection. The selection contextual dependency score I . Unlike the span privacy score ij algorithm then identifies protected spans and replaces each of P , I estimates the potential utility loss caused by removing i ij them with an opaque placeholder carrying only a local identifier a relation. A pair is contextually dependent when either span generated by the client. Only the sanitized prompt is sent to the supports local reconstruction of the other. cloud LLM. After the cloud model returns a response, the local To reduce computation, we score contextual dependencies postprocessor checks placeholder consistency before restoration. only for a sparse candidate set C rather than for all span pairs. Thus the cloud service observes only the sanitized prompt, We construct C with the top k nearest neighbors in the span the model response, and intentionally opaque placeholder embedding space. Each span is encoded, compared with other identifiers. spans by cosine similarity, and connected to its k most similar neighbors. We treat the resulting candidate pairs as undirected. This keeps the relation graph compact while retaining pairs B. Span Privacy Scoring Span privacy leakage risk is estimated from two comple- that are likely to share useful context. Let ℓi (x) denote a local reconstruction score that measures mentary sources. The first is the direct evidence score Di , how strongly the visible context x supports recovery of target obtained from local rules. As a surface-form score, spans that span ui . It is a utility signal, rather than a privacy-risk or attack exactly match or are confidently identified by privacy detectors score. The pairwise construction below applies to any local receive high risk values, while unmatched spans receive low scorer that produces ℓi (x). For example, a masked language or zero scores. It captures explicit sensitive patterns such as model with parameters θ provides the following token-level email addresses, physical addresses, and database credentials. instantiation [16]. Let ui = (wi,1 , . . . , wi,mi ) contain mi However, direct evidence covers only surface patterns that local tokens. For a prompt x that retains ui but may mask other detectors can explicitly match and therefore cannot capture spans, let x \(i,t) mask only wi,t : attribute evidence implied by otherwise ordinary contextual spans. The second source relies on a local privacy estimator G(a|x), which returns an exposure score for sensitive attribute a ∈ A in text x, where A denotes the set of sensitive attributes considered locally. Unlike Di , which scores a single span by direct surface matches, G scores the whole text and estimates

ℓi (x) =

mi  1 X log Pθ wi,t | x\(i,t) . mi t=1

(6)

Length normalization makes scores comparable across spans of different lengths. Larger ℓi (x) values indicate stronger contextual support for recovering ui .

4

Span Privacy Scoring

Sanitization

User with Prompt Selecting

Segmentation Pairwise Dependency Scoring Masking & Normalizing

[P1]

Spans

[P2]

[p3]

[p4]

Spans after Sanitization Original Spans

Sensitive Spans

Sanitized Spans

Response Spans

Sanitized Prompt

Local Restoration [P1]

Restoring

Cloud LLM

[P2] [P3]

Response

[p4]

Response

Fig. 2: Overview of PromptGraph. The local client segments the input prompt into textual spans, estimates span-level privacy risk and pairwise contextual dependency for utility preservation, and forms an attributed graph whose nodes correspond to spans and whose edges encode contextual dependencies. Graph-guided selection replaces protected spans with locally generated placeholders ([P1]–[P4]) before cloud inference. Only the sanitized prompt is sent to the cloud. The private mapping and restoration remain on the client after the response is returned.

For a retained candidate pair (ui , uj ), we measure dependency by the reduction in one span’s reconstruction score when the other is masked: 1 aij = [ℓi (q) − ℓi (q\j )]+ 2 (7)  + [ℓj (q) − ℓj (q\i )]+ . The two terms capture support in opposite directions, and the positive part retains only utility-relevant dependencies. Contextual dependency scores are then normalized within the retained candidate set C: Iij = aij /

max (up ,uq )∈C

apq .

(8)

If all raw scores are zero, we set Iij = 0 for every candidate pair. After normalization, larger Iij values indicate stronger contextual dependency. An edge is affected once either endpoint is protected and is penalized only once.

D. Dependency-Aware Sanitization After scoring, PromptGraph seeks a protected set that maximizes the following objective:

max F (S), F (S) = R(S) − λL(S), S⊆V X   Pi ,  R(S) =    v ∈S i  Eaff (S) = {(vi , vj ) ∈ E : {vi , vj } ∩ S ̸= ∅}, where  X    L(S) = Iij .   (vi ,vj )∈Eaff (S)

(9) Here, λ controls the privacy–utility trade-off. Each edge is charged when its first endpoint enters S and is not charged again if the other endpoint is later selected. When λ = 0, the objective performs span-only sanitization. Larger λ increasingly prioritizes contextual-dependency preservation, trading privacy gain for utility. We approximately optimize this discrete objective with a greedy marginal heuristic. For each unprotected node, its marginal gain equals its privacy score minus the λ-weighted sum of newly affected edges to other unprotected nodes. Starting from S = ∅, PromptGraph repeatedly adds the node with the largest positive marginal gain and stops when none remains. After selection, each span in S is replaced with an opaque placeholder carrying a local identifier, producing the sanitized prompt q̃. Only q̃ is sent to the cloud LLM for inference,

5

while the original spans, placeholder mapping, and private type metadata remain local.

Model

PHR↑

Method RegexMasking HaS InferDPT DP-OPT ProSan ALSA Ours

ASR↓

Med

SAM

Code

Avg.

Rec

Attr

0.652 0.859 0.759 0.211 0.653 0.656 0.879

0.662 0.884 0.780 0.203 0.749 0.552 0.847

0.547 0.458 0.898 0.524 0.547 0.472 0.894

0.620 0.734 0.812 0.313 0.650 0.560 0.873

0.087 0.283 0.078 0.320 0.094 0.202 0.045

0.700 0.705 0.472 0.676 0.612 0.654 0.347

E. Local Restoration After the cloud LLM returns a response r̃, the local Llama postprocessor scans it for placeholder strings. For each observed placeholder, it first checks whether its local identifier matches RegexMasking 0.652 0.662 0.547 0.620 0.089 0.700 an entry in the private restoration table. It then applies the local HaS 0.869 0.887 0.391 0.716 0.298 0.717 consistency checks using the private metadata stored with that InferDPT 0.755 0.784 0.816 0.785 0.087 0.477 0.205 0.213 0.489 0.302 0.296 0.659 entry. A placeholder is restored only when both checks succeed, Mistral DP-OPT ProSan 0.650 0.750 0.551 0.650 0.099 0.616 replacing it with the corresponding original protected span in ALSA 0.649 0.558 0.468 0.558 0.210 0.669 Ours 0.886 0.850 0.876 0.871 0.050 0.352 the final response r. The restoration table and its metadata remain on the client throughout this process. RegexMasking 0.652 0.662 0.547 0.620 0.086 0.700 HaS 0.851 0.892 0.438 0.727 0.288 0.710 If the model rewrites a placeholder, invents a new identiInferDPT 0.736 0.777 0.884 0.799 0.082 0.483 fier, or omits a placeholder, the corresponding item remains Qwen3 DP-OPT 0.216 0.221 0.505 0.314 0.330 0.654 ProSan 0.647 0.748 0.528 0.641 0.093 0.615 unresolved rather than being mapped to a protected span. ALSA 0.648 0.565 0.470 0.561 0.201 0.645 Thus, the postprocessor does not restore a placeholder merely Ours 0.898 0.848 0.890 0.879 0.047 0.345 because it resembles a valid one. This validation confines restoration to client-side entries that pass local checks and TABLE I: Privacy evaluation across downstream LLMs, prevents malformed or fabricated placeholders from triggering datasets, and attacks on the server side. unintended restoration of sensitive content. 4) Parameter Settings: The main experiments set the tradeV. E XPERIMENTS off coefficient to λ = 0.2 and the embedding neighbor A. Experimental Setup sparsification parameter to k = 4. The direct evidence term Di 1) Datasets and Models: We evaluate PromptGraph on is computed from the outputs of rule-based entity recognizers three task datasets: MedQA [17] is used for medical question and custom regular expression recognizers [28]. Mapped answering, SAMSum [18] for dialogue summarization, and detector scores are normalized on the validation split of CodeAlpaca-20k for instruction-style code generation [19]. To AI4Privacy PII-Masking-300k [29], a PII masking dataset used evaluate response quality from LLMs, we use three mainstream in recent evaluations of masking models. downstream LLMs: Llama-3.1-8B [20], Mistral-7B [21], and Qwen3-8B [22]. 2) Baselines: We compare six baselines spanning rule- B. Experimental Results 1) Privacy Evaluation: As shown in Table I, based masking and representative prompt privacy methods. PromptGraph achieved the highest average PHR for RegexMasking removes structured sensitive strings using predefined patterns. HaS [5] performs local hide-and-seek each of the three downstream LLMs. PHR is computed anonymization with restoration after generation. InferDPT [6] over the server-visible sanitized prompt and raw response and DP-OPT [7] represent private inference by perturbation before local restoration. The identical RegexMasking values and differentially private prompt optimization, respectively. across model blocks reflect its deterministic prompt-side ProSan [8] balances privacy risk with word importance, while masking behavior in these runs. The consistently high PHR ALSA [9] further incorporates contextual importance and task of PromptGraph suggests that its selected spans protect both explicit sensitive values and contextually related attribute relevance when assigning anonymization actions. 3) Evaluation Metrics: Privacy is measured with Privacy cues. Although some baselines performed particularly well Hiding Rate (PHR) [8] and two attack metrics on the server side: in individual settings, including HaS on SAMSum and Recovery Attack Success Rate (Rec-ASR) [23] and Attribute InferDPT on CodeAlpaca with Llama-3.1-8B, their PHR Inference Attack Success Rate (Attr-ASR) [24], [25]. Rec- performance varied more substantially across datasets. Overall, ASR measures whether the attacker can recover the exact PromptGraph provided more stable privacy coverage across sensitive value, while Attr-ASR measures whether the attacker the three task formats. Attack results showed the same trend. PromptGraph can infer the sensitive attribute type or meaning even without recovering the exact value. We utilize task utility to measure consistently achieved the lowest Rec-ASR and Attr-ASR, whether the cloud LLM’s response remains correct and useful indicating reduced exposure of both exact sensitive values and after prompt privacy protection. Efficiency is evaluated by the attribute evidence. By comparison, RegexMasking remained local preprocessing overhead, including Sanitization Time Cost susceptible to attribute inference, HaS still leaked contextual information despite its high SAMSum PHR, and InferDPT (STC) and Peak Memory Overhead (PMO). Please note that, because the three datasets represent distinct retained more attribute evidence. These results indicate that task types, we use task-specific utility metrics: answer accuracy contextual span privacy scoring covered both explicit sensitive for MedQA, ROUGE-L [26] for SAMSum, and equal-weight content and attribute cues, while the dependency-aware penalty CodeBLEU [27] for CodeAlpaca. We collectively refer to these helped preserve utility-relevant context during this privacymeasures as Task Utility. driven selection.

6

SAMSum

CodeAlpaca

w/o contextual term Ci

w/o pairwise weights Iij

RegexMasking HaS InferDPT DP-OPT ProSan ALSA Ours

0.171 0.230 0.241 0.241 0.246 0.234 0.248

0.104 0.157 0.042 0.088 0.081 0.103 0.149

0.510 0.403 0.263 0.439 0.459 0.321 0.536

MedQA SAMSum CodeAlpaca

0.879 0.847 0.894

0.875 0.808 0.750

0.940 0.853 0.985

ASR↓

Rec-ASR Attr-ASR

0.045 0.347

0.233 0.555

0.021 0.238

Mistral

RegexMasking HaS InferDPT DP-OPT ProSan ALSA Ours

0.332 0.320 0.154 0.327 0.304 0.318 0.335

0.202 0.211 0.045 0.194 0.168 0.154 0.234

0.489 0.442 0.292 0.457 0.444 0.380 0.472

Task Utility↑

MedQA SAMSum CodeAlpaca

0.248 0.149 0.536

0.252 0.197 0.650

0.117 0.057 0.514

Qwen3

RegexMasking HaS InferDPT DP-OPT ProSan ALSA Ours

0.333 0.320 0.113 0.331 0.326 0.289 0.336

0.141 0.210 0.018 0.145 0.163 0.170 0.240

0.619 0.540 0.251 0.570 0.562 0.381 0.592

TABLE II: Utility evaluation across various LLM task models and datasets.

Metric

Setting

PHR↑

TABLE III: Ablation of PromptGraph’s contextual privacy term and pairwise dependency weights using Llama-3.1-8B.

1.00

0.38

0.75

0.37

Task Utility

MedQA

Llama

Full PromptGraph

Task Utility↑

Method

PHR

Model

0.50 0.25

k=1

k=2

k=4

k=8

0.1

0.2

λ value 0.6

(a) PHR

0.35 0.34

0.00 0

0.36

0.5

1.0

0

0.1

0.2

0.5

1.0

λ value

(b) Task Utility

ASR

PHR

0.8 0.6 0.4 0.2

RegexMasking InferDPT ProSan PromptGraph

HaS DP-OPT ALSA

Fig. 4: Effects of the trade-off coefficient λ and graph sparsity k on PromptGraph’s privacy and task utility.

0.4

0.2

privacy protection while preserving task utility. 3) Ablation Studies: To disentangle the contributions of (a) PHR–utility trade-off (b) ASR–utility trade-off the two graph-aware signals, we ablate each component and Fig. 3: Privacy–utility trade-offs. Faint points indicate individual observe complementary effects on privacy and utility. In downstream task models, and larger points indicate method Table III, removing the contextual term Ci lowered PHR and averages. ASR is computed as the average of Rec-ASR and raised both attack success rates, while task utility increased. This setting still protects direct detector hits but no longer Attr-ASR. assigns privacy value to ordinary spans through counterfactual attribute evidence. The result is therefore consistent with Ci 2) Utility Evaluation: The utility evaluation tests whether reducing inferential leakage at the cost of hiding some context the privacy gains above came at the cost of downstream that supports task execution. By contrast, eliminating the pairwise weights Iij strengthtask performance. As shown in Table II, PromptGraph achieved the best score in six of the nine model-dataset ened measured privacy but reduced utility, with the clearest settings and remained close to the best baseline in the other drops on MedQA and SAMSum. Because the edge penalty three. This stable utility is especially clear on MedQA, where then vanishes from the greedy marginal gain, selection is driven PromptGraph was best across all models, and on SAMSum, only by span privacy scores and becomes more aggressive. It where preserving high-dependency speaker and event relations consequently removes additional sensitive evidence, but can kept summaries competitive. On CodeAlpaca, RegexMasking also break relationships that the downstream model uses. The remained strong because many programming instructions full method retains both signals, thereby avoiding the two survived simple masking, but PromptGraph still achieved one-sided behaviors: high contextual exposure without Ci and the best score and stayed close on the other two models. These excessive context removal without Iij . results indicate that the privacy gains of PromptGraph were 4) Sensitivity Analysis of Hyperparameters: We assessed not obtained by excessive removal of task-relevant content. the sensitivity of PromptGraph to the trade-off coefficient Based on the results in Tables I and II, Figure 3 summarizes λ and graph sparsity k. Figure 4 reveals a broad stable region the trade-off between privacy and utility. Figure 3a shows as λ increases from 0 to 0.2, with nearly unchanged PHR and that PromptGraph occupies the desirable high-privacy, high- task utility across sparsity settings. The main configuration, utility region, achieving the highest average PHR while with a coefficient of 0.2 and k set to 4, lies within this plateau. maintaining the highest average task utility among all compared Beyond this region, larger coefficients shift selection toward methods. The results in Figure 3b show that PromptGraph retaining dependency-weighted context. Whenλ is 1.0, the also achieves the lowest average ASR without sacrificing utility. edge-dependency penalty causes settings to select almost no These results demonstrate that PromptGraph strengthens content for protection. Privacy declined first for denser graphs, 0.1

0.2

0.3

Task Utility

0.4

0.1

0.2

0.3

Task Utility

0.4

7

Span Privacy Scoring Pairwise Dependency Scoring

PromptGraph

0.59

SAMSum

MedQA

(a) STC Results InferDPT DP-OPT

ProSan ALSA

1.0 0.5 0.0

STC

MedQA

SAMSum

CodeAlpaca

<0.1

<0.1

PMO

1.2

6

0.8

4

0.4

2

0.0

64

128

192

256

0

Textual spans

(c) PMO Results

single-turn prompts to multi-turn interactions, where privacy evidence and useful contextual dependencies evolve across dialogue turns. R EFERENCES

CodeAlpaca

(b) Stage STC Breakdown PromptGraph

STC (s)

PMO (MB)

1.5

RegexMasking HaS

SAMSum

<0.1

0.0

CodeAlpaca

PMO (MB)

MedQA

0.2 <0.1

0

0.54

<0.1

1

Dependency-Aware Sanitization Local Restoration

0.4

<0.1

2

STC (s)

STC (s)

0.6

<0.1

ProSan ALSA

<0.1

InferDPT DP-OPT

<0.1

RegexMasking HaS

<0.1

3

(d) Impact of Span Count

Fig. 5: Overhead Results. PMO measures the additional memory used during sanitization, excluding resident localmodel parameters.

whereas utility improved only marginally until the largest coefficient, where privacy degradation became pronounced. Thus, the sensitivity analysis identifies a moderate trade-off coefficient as the most robust operating regime. 5) Overhead: Finally, we assessed the local preprocessing overhead of PromptGraph. As shown in Figure 5(a), its sanitization time cost was comparable to that of DP-OPT and ProSan, while remaining consistently below ALSA, particularly on MedQA and SAMSum. The stage breakdown in Figure 5(b) identifies pairwise dependency scoring as the dominant preprocessing cost across all three datasets. Span privacy scoring, dependency-aware sanitization, and local restoration each contributed comparatively little. Figure 5(c) further shows that PMO, excluding resident local-model memory, remained below 0.5 MB across datasets, matching DP-OPT and ProSan and staying well below ALSA. Moreover, Figure 5(d) tracks the effect of prompt length. As the number of textual spans increased from 64 to 256, both STC and PMO rose monotonically, reaching approximately 1.1 s and 6 MB at the largest setting. This trend is consistent with the growth in candidate span pairs required for dependency scoring. Overall, the overhead is concentrated in contextual pair construction rather than in the subsequent selection or restoration steps. VI. C ONCLUSION This paper introduced PromptGraph, a graph-guided prompt sanitizer for privacy-preserving LLM inference. PromptGraph selects protected spans over a graph whose nodes capture privacy risks and whose edges capture contextual dependencies needed for utility, reducing privacy exposure through contextual span scoring while preserving task-relevant relations. Across diverse datasets and task models, it improves the balance between privacy and utility by increasing PHR and reducing attack success rates while preserving competitive utility. Future work will extend this graph formulation from

[1] L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe, “Training language models to follow instructions with human feedback,” in Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 27 730–27 744. [2] E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell, “On the dangers of stochastic parrots: Can language models be too big?” in Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, ser. FAccT ’21. ACM, 2021, pp. 610–623. [3] O. Feyisetan, B. Balle, T. Drake, and T. Diethe, “Privacy- and utilitypreserving textual analysis via calibrated multivariate perturbations,” in Proceedings of the Thirteenth ACM International Conference on Web Search and Data Mining. ACM, 2020, pp. 178–186. [4] S. Chen, F. Mo, Y. Wang, C. Chen, J.-Y. Nie, C. Wang, and J. Cui, “A customized text sanitization mechanism with differential privacy,” in Findings of the Association for Computational Linguistics: ACL 2023. Toronto, Canada: Association for Computational Linguistics, Jul. 2023, pp. 5747–5758. [5] Y. Chen, T. Li, H. Liu, and Y. Yu, “Hide and seek (HaS): A lightweight framework for prompt privacy protection,” 2023. [6] M. Tong, K. Chen, J. Zhang, Y. Qi, W. Zhang, N. Yu, T. Zhang, and Z. Zhang, “InferDPT: Privacy-preserving inference for closed-box large language models,” IEEE Transactions on Dependable and Secure Computing, vol. 22, no. 5, pp. 4625–4640, 2025. [7] J. Hong, J. T. Wang, C. Zhang, Z. Li, B. Li, and Z. Wang, “DP-OPT: Make large language model your privacy-preserving prompt engineer,” in International Conference on Learning Representations, 2024, pp. 58 003–58 026. [8] Z. Shen, Z. Xi, Y. He, W. Tong, J. Hua, and S. Zhong, “The fire thief is also the keeper: Balancing usability and privacy in prompts,” 2024. [9] H. Ma, W. Lu, Y. Liang, T. Wang, Q. Zhang, Y. Zhu, and J. Si, “ALSA: Context-sensitive prompt privacy preservation in large language models,” in Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2, ser. KDD ’25. New York, NY, USA: ACM, 2025, pp. 2042–2053. [10] O. Rubin, J. Herzig, and J. Berant, “Learning to retrieve prompts for in-context learning,” in Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 2655–2671. [11] Z. Zhao, E. Wallace, S. Feng, D. Klein, and S. Singh, “Calibrate before use: Improving few-shot performance of language models,” in Proceedings of the 38th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, vol. 139. PMLR, 2021, pp. 12 697–12 706. [12] N. F. Liu, K. Lin, J. Hewitt, A. Paranjape, M. Bevilacqua, F. Petroni, and P. Liang, “Lost in the middle: How language models use long contexts,” Transactions of the Association for Computational Linguistics, vol. 12, pp. 157–173, 2024. [13] J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. V. Le, and D. Zhou, “Chain-of-thought prompting elicits reasoning in large language models,” in Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 24 824–24 837. [14] M. Besta, N. Blach, A. Kubicek, R. Gerstenberger, M. Podstawski, L. Gianinazzi, J. Gajda, T. Lehmann, H. Niewiadomski, P. Nyczyk, and T. Hoefler, “Graph of thoughts: Solving elaborate problems with large language models,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, 2024, pp. 17 682–17 690. [15] S. Wachter, B. Mittelstadt, and C. Russell, “Counterfactual explanations without opening the black box: Automated decisions and the GDPR,” Harvard Journal of Law & Technology, vol. 31, no. 2, pp. 841–887, 2018. [16] J. Salazar, D. Liang, T. Q. Nguyen, and K. Kirchhoff, “Masked language model scoring,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, 2020, pp. 2699–2712. [17] D. Jin, E. Pan, N. Oufattole, W.-H. Weng, H. Fang, and P. Szolovits, “What disease does this patient have? a large-scale open domain question answering dataset from medical exams,” Applied Sciences, vol. 11, no. 14, p. 6421, 2021.

8

[18] B. Gliwa, I. Mochol, M. Biesek, and A. Wawer, “SAMSum corpus: A human-annotated dialogue dataset for abstractive summarization,” in Proceedings of the 2nd Workshop on New Frontiers in Summarization, 2019, pp. 70–79. [19] S. Chaudhary, “Code alpaca: An instruction-following LLaMA model for code generation,” GitHub repository, 2023, gitHub repository. [20] A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, A. Yang, A. Fan, A. Goyal, A. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru, B. Roziere, B. Biron et al., “The Llama 3 herd of models,” 2024. [21] A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de Las Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, L. R. Lavaud, M.-A. Lachaux, P. Stock, T. L. Scao, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed, “Mistral 7B,” 2023. [22] A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, C. Zheng, D. Liu, F. Zhou, F. Huang, F. Hu, H. Ge, H. Wei, H. Lin, J. Tang, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Zhou, J. Lin, K. Dang, K. Bao, K. Yang, L. Yu, L. Deng, M. Li, M. Xue, M. Li, P. Zhang, P. Wang, Q. Zhu, R. Men, R. Gao, S. Liu, S. Luo, T. Li, T. Tang, W. Yin, X. Ren, X. Wang, X. Zhang, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Zhang, Y. Wan, Y. Liu, Z. Wang, Z. Cui, Z. Zhang, Z. Zhou, and Z. Qiu, “Qwen3 technical report,” 2025. [23] N. Carlini, F. Tramer, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. Brown, D. Song, U. Erlingsson, A. Oprea, and C. Raffel, “Extracting training data from large language models,” in 30th USENIX Security Symposium, 2021, pp. 2633–2650. [24] L. Melis, C. Song, E. D. Cristofaro, and V. Shmatikov, “Exploiting unintended feature leakage in collaborative learning,” in 2019 IEEE Symposium on Security and Privacy. IEEE, 2019, pp. 691–706. [25] R. Staab, M. Vero, M. Balunovic, and M. T. Vechev, “Beyond memorization: Violating privacy via inference with large language models,” in The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7–11, 2024. OpenReview.net, 2024. [26] C.-Y. Lin, “ROUGE: A package for automatic evaluation of summaries,” in Text Summarization Branches Out, 2004, pp. 74–81. [27] S. Ren, D. Guo, S. Lu, L. Zhou, S. Liu, D. Tang, N. Sundaresan, M. Zhou, A. Blanco, and S. Ma, “CodeBLEU: A method for automatic evaluation of code synthesis,” 2020. [28] I. Neamatullah, M. M. Douglass, L. wei H. Lehman, A. Reisner, M. Villarroel, W. J. Long, P. Szolovits, G. B. Moody, R. G. Mark, and G. D. Clifford, “Automated de-identification of free-text medical records,” BMC Medical Informatics and Decision Making, vol. 8, p. 32, 2008. [29] D. Singh and S. Narayanan, “Unmasking the reality of PII masking models: Performance gaps and the call for accountability,” 2025.

Record · ID 363182 · SHA-256 394600a583a4f7e9
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.