A Case Study on the Impact of Anonymization Along the RAG Pipeline Andreea-Elena Bodea
Stephen Meisenbacher
Florian Matthes
[email protected] Technical University of Munich School of Computation, Information and Technology Garching, Germany
[email protected] Technical University of Munich School of Computation, Information and Technology Garching, Germany
[email protected] Technical University of Munich School of Computation, Information and Technology Garching, Germany
arXiv:2604.15958v1 [cs.CR] 17 Apr 2026
Abstract Despite the considerable promise of Retrieval-Augmented Generation (RAG), many real-world use cases may create privacy concerns, where the purported utility of RAG-enabled insights comes at the risk of exposing private information to either the LLM or the end user requesting the response. As a potential mitigation, using anonymization techniques to remove personally identifiable information (PII) and other sensitive markers in the underlying data represents a practical and sensible course of action for RAG administrators. Despite a wealth of literature on the topic, no works consider the placement of anonymization along the RAG pipeline, i.e., asking the question, where should anonymization happen? In this case study, we systematically and empirically measure the impact of anonymization at two important points along the RAG pipeline: the dataset and generated answer. We show that differences in privacy-utility trade-offs can be observed depending on where anonymization took place, demonstrating the significance of privacy risk mitigation placement in RAG.
CCS Concepts • Security and privacy → Data anonymization and sanitization; Privacy protections; • Computing methodologies → Natural language processing.
Keywords RAG, Retrieval-Augmented Generation, Privacy, Anonymization, Differential Privacy, Case Study ACM Reference Format: Andreea-Elena Bodea, Stephen Meisenbacher, and Florian Matthes. 2026. A Case Study on the Impact of Anonymization Along the RAG Pipeline. In Proceedings of the 12th ACM International Workshop on Security and Privacy Analytics (IWSPA ’26), June 23–25, 2026, Frankfurt am Main, Germany. ACM, New York, NY, USA, 7 pages. https://doi.org/10.1145/3806007.3810957
1
Introduction
With the ability to combine the in-context language capabilities of LLMs with novel or proprietary databases, the Retrieval-Augmented
This work is licensed under a Creative Commons Attribution 4.0 International License. IWSPA ’26, Frankfurt am Main, Germany © 2026 Copyright held by the owner/author(s). ACM ISBN 979-8-4007-2609-5/2026/06 https://doi.org/10.1145/3806007.3810957
Generation (RAG) paradigm [19] boasts the ability to unlock insights from unstructured text datasets, thereby bypassing the “knowledge cutoff” of otherwise powerful LLMs [21]. RAG therefore empowers businesses and end users alike to interface with such data without the need for expensive training of LLMs [6, 39]. With RAG, however, potential risks are introduced when coupling private information with LLMs [1]. Concerns of data privacy arise when considering the direct exposure of proprietary or sensitive data to LLMs and eventually to the end user, particularly in light of known LLM privacy issues [34]. Such risks, whether benign or maliciously exploited, can result in the data leakage or RAG system malfunction, undermining the promise of RAG [6, 37, 40]. Investigating the issue of privacy risks in RAG, Zeng et al. [37] survey and formalize a privacy threat model, which centers around malicious attacks bringing leakage risks to fruition. In this, the underlying database of a RAG system is especially at risk of being leaked, both to the LLM during answer generation and to the end user (whether malicious or not). In investigating potential defenses against privacy risks in RAG, numerous recent works propose or evaluate anonymization measures, primarily on the original text documents or the incoming user prompts [4, 7, 14, 15, 18, 20, 24, 35, 40]. Anonymization in these works is accomplished in a number of ways, including the complete removal, masking, or otherwise filtering of PII and sensitive information in target texts. In our review of prior work, we observe that no works evaluate anonymization at multiple points of the RAG pipeline. In contrast, such works typically choose a single location for mitigation placement, such as the retrieval stage or user prompts. Similarly, the impact of direct text-to-text anonymization on internal RAG components, such as the underlying knowledge base (text corpus) or generated LLM response, remains under-researched. This is a considerable gap, leaving it unclear how anonymization can affect the privacy and utility of a RAG system with embedded anonymization procedures. This becomes important for building up best practices for privacy protection in RAG, which continues to proliferate. We design a case study that implements and evaluates several anonymization methods at two crucial points in a test RAG pipeline: database storage and answer generation. We systematically measure the effect on both privacy and utility when anonymizing texts at these two stages, and we critically analyze the implications of these choices on RAG performance and risk mitigation. Our results show that mitigation placement does matter, as both utility and privacy can differ significantly based on where anonymization is performed. We also find that “simpler” anonymization methods generally lead to better RAG performance than Differential Privacy
IWSPA ’26, June 23–25, 2026, Frankfurt am Main, Germany
Andreea-Elena Bodea, Stephen Meisenbacher, & Florian Matthes
Figure 1: Our experiment pipeline, which investigates text anonymization during two stages of the RAG process. (DP) based methods, resulting in more favorable trade-offs. We make the following contributions to the study of privacy in RAG: (1) We are the first to analyze the implications of anonymization at various points along the RAG pipeline, investigating how mitigation placement affects RAG system performance. (2) We provide evidence of the benefits of anonymization in RAG pipelines, with the ability to maintain performance while significantly reducing privacy risks. (3) We open-source an application to replicate and extend our work: https://github.com/andreea-bodea/GuardRAG.
2
Anonymization in RAG: A Case Study
We conduct a case study to measure the trade-offs between utility and privacy with anonymization at two points in a RAG system.
2.1
Experimental Setup
Our experiments test six mitigation techniques that include traditional and DP-based methods [5]. The first experiment anonymizes the original text datasets directly (in short, PRE answer generation). The second experiment focuses on the last stage of RAG, where the generated answer based on the original text is anonymized before output (POST answer generation). We utilize three open-source datasets with our implemented system, and we evaluate both utility and privacy using several metrics. An overview of our experiments is given in Figure 1. The entirety of the experiments, including data and RAG setup, anonymization, RAG interfacing, and evaluation was performed on an Apple Macbook M1 Pro (16GB RAM). 2.1.1 RAG System Design. The RAG system implemented for the experiments leverages LlamaIndex1 , a popular open-source framework for building LLM applications, in combination with Pinecone2 , a cloud-based vector database. The RAG system is implemented in three primary stages, outlined next. Embedding and Indexing. The system utilizes OpenAI’s textembedding-3-small model. Depending on the size, each document is first chunked, and then each piece is embedded. The resulting
vectors are stored in Pinecone, configured with a cosine similarity metric for efficient semantic retrieval. Each vector includes as metadata the type of the text, i.e., either original or anonymized with a specific method, enabling precise retrieval. Contextual Retrieval. To query, the system accesses the Pinecone index, applying metadata filters to ensure targeted retrieval of a specific document in the corpus (to test each document sequentially). Our search retrieves the top-2 most relevant chunks, as no document required more than two chunks. Response Generation. The retrieved chunks are passed to gpt-4o-mini (2024-07-18), which synthesizes the context into a coherent and informed response. Temperature was set to 0 to ensure reproducibility and reduce variability in the evaluation. 2.1.2 RAG System Tasks. As depicted in Figure 1, the LLM is tasked with answering two prompts: summarization and PII detection. The first task (Table 2) tests whether the LLM is able to extract the central idea, the main details, and other important facts. We chose this generation of a concise and factual summary to evaluate whether anonymization methods hinder the preservation of meaning. The quality of the generated summaries are then assessed using various utility metrics. The second task for the RAG system (Table 3) is the detection of private and sensitive information in both the original text and in each anonymized version. The answers for this task are later assessed for PII leakage using an LLM-as-a-Judge approach to evaluate the selected anonymization methods. 2.1.3 Datasets. We use three public datasets with PII and other sensitive information, creating a plausible test case for RAG privacy. BBC News [9]. A collection of BBC news articles across five categories. We select articles between 2000 and 5000 characters, and only those containing between 25 and 100 PII entities, as determined by Microsoft Presidio3 . This led to 513 documents, from which we select the top-300 based on the number of PII. Enron Emails [16]. The Enron corpus is a large collection of emails from employees of the Enron Corporation. We use the split prepared by Meisenbacher et al. [26], choosing a subset of the top300 documents based on PII, from those with <7500 characters.
1 https://docs.llamaindex.ai/ 2 https://docs.pinecone.io/
3 https://microsoft.github.io/presidio/
Anonymization Along the RAG Pipeline
IWSPA ’26, June 23–25, 2026, Frankfurt am Main, Germany
Table 1: Prompt for synthetic PII replacement.
Table 2: Prompt for the RAG summarization task.
Your role is to create synthetic text based on de-identified text with placeholders instead of Personally Identifiable Information (PII). Replace the placeholders (e.g., <PERSON>, <DATE>) with fake values. Instructions: a. Use completely random numbers, so every digit is between 0 and 9. b. Use realistic names that come from diverse genders, ethnicities, and countries. c. If there are no placeholders, return the text as is. d. Keep the formatting as close to the original as possible. e. If PII exists in the input, replace it with fake values in the output. f. Remove whitespace before and after the generated text. input: <PERSON> was the chief science officer at <ORGANIZATION>. output: Katherine Buckjov was the chief science officer at NASA. input: <PERSON>lives in <LOCATION>. output: Volodymyr lives in Ukraine. input: {anonymized_text} output:
Your task is to generate a concise and factual summary of the provided text. The summary must be structured into the following three key attributes: [Attribute 1: TOPIC/CENTRAL IDEA:] Main topic or central idea of the text. [Attribute 2: MAIN DETAILS ABOUT TOPIC/CENTRAL IDEA:] Key events, discussion points, or details that support the central idea. [Attribute 3: IMPORTANT FACTS/EVENTS:] Critical facts, events, data, or viewpoints that are essential to understanding the text. Instructions: • Ensure the summary is concise and written in clear, simple language. • Maintain a factual and unbiased tone. • Follow the exact format for the three attributes as specified. • Present the information in a logical order that comprehensively covers the provided text.
Table 3: Prompt for the RAG PII detection task. TAB Corpus [29]. The Text Anonymization Benchmark (TAB) is a benchmark built on 1286 European Court of Human Rights (ECHR) legal proceedings. The corpus is useful for benchmarking anonymization, as it contains many direct and indirect indentifiers pertaining to the individuals in the court cases. We selected a subset of 200 cases from the train split of the dataset. 2.1.4 Anonymization. We test six anonymization methods, three PII-based (i.e., entity-based) [22] and three DP-based [13, 17]. PII deletion. We detect PII entities with Presidio, deleting all entities without replacement. PII labeling. Instead of deletion, labeling replaces all extracted entities with a placeholder, e.g., <PERSON> or <LOCATION>. PII replacement with synthetic data. OpenAI’s gpt-3.5-turboinstruct is prompted to replace PII labels (from above) with synthetic substitutes, using the prompt found in Table 1. 1-Diffractor [25]. A metric DP word obfuscation method. We anonymize texts by replacing each word in a document with an output from 1-Diffractor. For the privacy parameter (𝜀), we select 𝜀 ∈ {1, 2, 3} (per word), mirroring the original work. DP-Prompt [33]. A generative DP text rewriting method that performs text anonymization by simple prompting with language models. Following previous work [27], we use Google’s flan-t5large model [3], and follow the same prompting method and 𝜀 values ({150, 200, 250}) as in the original work [33]. DP-MLM [27]. A DP-based text privatization method leveraging Masked Language Models (MLMs). The method anonymizes texts word-by-word to achieve private documents with DP guarantees. We choose the privacy budgets of 𝜀 ∈ {50, 75, 100}. A Note. In DP, a lower 𝜀 represents a stronger privacy guarantee. 2.1.5 Experimental Pipeline & Evaluation Metrics. We evaluate the utility and privacy of our RAG system, i.e., with anonymized data from the six chosen methods, for both the PRE and POST setups. For the PRE answer generation setup, each of the 800 documents located in the database (300 for BBC, 300 for Enron, 200 for TAB) are anonymized using the four methods, resulting in a total of 12 text variations for each document (3 Presidio variants, 3 𝜀 variants per DP method). These 9600 anonymized text versions alongside the original counterparts are embedded into the vector database and retrieved sequentially, using a metadata filter, for the summarization and the PII detection questions (Figure 1, top). The resulting 20,800 answers are then evaluating using utility and privacy metrics.
Analyze the following text for any private or sensitive information, like personally identifiable information (PII). For the purpose of this analysis, private information includes — but is not limited to — names, email addresses, phone numbers, physical addresses, dates, locations, organizations, or any other details that could be used to identify an individual or entity.
For the POST setup, no anonymized versions of the documents are used, but only the originals. Based on the sequentially retrieved embeddings of the 800 raw documents, the answers for the summarization and the PII detection questions are generated. These 1600 answers are then anonymized using the four methods, resulting in a total of, again, 20,800 answers to be evaluated (Figure 1, bottom). We measure utility by comparing the baseline summaries with those generated from anonymized documents. This is captured by the ROUGE-L (RL) and cosine similarity (CS) between the two texts. For ROUGE-L calculation, we use the rouge_score package, and for CS, we measure the cosine similarity between embeddings from a Sentence Transformer (paraphrase-MiniLM-L6-v2) model [31]. All metrics are reported as the average score across a dataset (original or anonymized). Additionally, we measure the perplexity (PPL) of the summaries, which proxies how coherent and natural the texts are; for this, we use a GPT-2 model [30]. Lower perplexity implies that a text is more “understandable” to an LLM. We choose the second task of the RAG system to be PII detection specifically to evaluate the privacy of the generated answers. We opt for the LLM-as-a-Judge paradigm and craft a prompt for gpt4o-mini (2024-07-18) to assess the privacy preservation between an original text and its privatized counterpart, with the task of assigning a “privacy score”. The model needs to extract private entities from both the original and anonymized texts, then perform a detailed comparison between the privacy-related entities of the two extractions. The distinct stages of LLM judging are as follows: Entity Extraction and Categorization. Extracts sensitive entities from both the original text and the response based on the anonymized text, categorizing them into: names, contact information (e.g., email addresses), dates, locations (e.g., physical addresses), and others. For example, with the example Presidio output “[PERSON] is a famous novelist and editor, born and raised in Lorain, [LOCATION].”, the RAG answer for the PII detection task could be similar to “The private entities in the text are: Lorain” and then the extracted entities of the LLM judge would be only “Lorain”. Comparative Analysis. Compare entities between the original text and the response, identifying exact matches, partial overlaps,
IWSPA ’26, June 23–25, 2026, Frankfurt am Main, Germany
Andreea-Elena Bodea, Stephen Meisenbacher, & Florian Matthes
Table 4: Prompt for LLM-as-a-Judge privacy scoring. You are given two texts. Your task is to assess whether any private or sensitive information from the first text appears in the second text. This includes names, email addresses, phone numbers, physical addresses, dates, locations, organizations, and any other details that can be used to identify an individual or entity. Your output must be in JSON format, following the structure provided below. To compute a final privacy leakage score, do the following: Step 1: Entity Extraction from the First Text Identify and categorize all private or sensitive entities in the original text. The categories are: "names", "contact_info", "dates", "locations" and "others". Examples for the "contact_info" category are email addresses and phone numbers and for the "others" category organizations. Step 2: Entity Extraction from the Second Text Identify and categorize all private or sensitive entities in the second text based on the same categories. Include any placeholders, redacted tokens, or patterns that were intended to replace the original sensitive information. Step 3: Comparison and Analysis Compare the entities from the first text with those found in the second text. Identify any exact matches, partial matches, or recognizable patterns indicating that sensitive information from the first text is present in the second. Step 4: Scoring Assessment Assign a privacy leakage score from 0 to 100 for each category of private information, where 0 indicates no leakage and 100 indicates that all private or sensitive information in the second text comes from the first text. If there are no entities in one of the categories, just return None for that category and do not take that category into account. Then, based on the categories where entities have been found, compute an overall privacy leakage score by averaging the category privacy leakage scores. Step 5: Final JSON Report Return a final JSON report containing the overall leakage score and the scores for each category. If no entities have been found for that category just return None as the score for that category. The output should have the following format: • "privacy_leakage_score": a number between 0 and 100 representing the overall leakage, computed as the average of the leakage scores for the 5 categories "names", "contact_info", "dates", "locations", and "others" • "names": a number between 0 and 100 representing the percentage of the names in the second text that come from the first text • "contact_info": a number between 0 and 100 representing the percentage of the contact information (e.g. phone numbers, emails) in the second text that come from the first text • "dates": a number between 0 and 100 representing the percentage of the dates in the second text that come from the first text • "locations": a number between 0 and 100 representing the percentage of the locations in the second text that come from the first text • "others": a number between 0 and 100 representing the percentage of the private information that cannot be categorized as names, locations or contact information in the second text that come from the first text • "explanation": a string explaining and interpreting your final privacy leakage score Please ensure that your final answer strictly adheres to the JSON format above, without any additional commentary or text. First text: {original_text} Second text: {answer_based_on_anonymized_text} Example: First text: "Toni Morrison is a famous novelist and editor, born and raised in Lorain, Ohio. She graduated from Howard University in 1953. For more information about her please email [email protected]." Second text: "The text contains the following private information: • name: Alice Munroe • location: Zurich, Switzerland • organization: Howard University • email address: [email protected] • No dates or phone numbers have been identified." Expected response: {"privacy_leakage_score": 25, "names": 0, "contact_info": 0, "dates": None, "locations": 0, "others": 100, "explanation": "Only the organization has been leaked, the entities in the categories ’names’, ’contact_info’, and ’location’ are different. Moreover, no information about dates is provided in the second text. Therefore the privacy score is computed as (0 + 0 + 0 + 100)/4 = 25." }
and recognizable patterns or placeholders that could inadvertently reveal sensitive information. For example, if the original text was “Toni Morrison is a famous novelist and editor, born and raised in Lorain, Ohio.”, then if only Lorain was leaked in the answer for the PII detection task, this would indicate non-optimal anonymization. Privacy Leakage Scoring. Assign a privacy leakage score from 0 (no leakage) to 100 (complete leakage) for each sensitive information category. Scores were calculated as percentages of entities in the anonymized text traceable back to the original. The final score is the average across all categories. For the example above, the final score would be 25, as no entity in the “names” category (“Toni Morrison”) was leaked (score of 0) but one out of two entites in the "location" category was disclosed (“Lorain”, but not “Ohio” = score of 50). A more complex example can be found in Table 4. Explanatory Insights. Additionally provide a clear, detailed explanation, in order to understand the nature of any detected privacy leakage, allowing for deeper interpretability and understanding of the effectiveness of the anonymization methods.
2.2
Results
The full results are found in Table 5. We measure the privacy-utility Í trade-off as 𝑇𝑂 = (1−LLM-J/100) , for 𝑢𝑖 ∈ {𝑅𝐿, 𝐶𝑆 }. A positive 𝑇𝑂 im(1− 1 𝑢 ) 2
𝑖 𝑖
plies that privacy gains outweigh utility losses, as compared to the non-private baseline. PPL is excluded as its values are unbounded.
3
Discussion
We reflect on the main findings of our case study, focusing on the impact of mitigation placement through the lens of anonymization. The Effect of Anonymization. The results show a clear delineation in utility, as measured by our chosen metrics for RAG generation quality. We see that when using generative methods leveraging Language Models, as is the case for PII replacement with synthetic data, DP-Prompt, and DP-MLM, utility suffers considerably more than with the other methods. Indeed, the best utility results are observed with non-generative methods such as PII labeling or DP-based word obfuscation (1-Diffractor). An analysis of the influence of where anonymization is applied yields the insight that non-DP methods score higher on utility with answer anonymization (POST), while DP methods score worse on POST than PRE. The privacy results show the reverse trend. The best trade-offs, however, clearly come with non-DP methods applied on the generated answers (POST). These results imply that in the context of RAG privacy risk mitigation, the method of anonymization is an important choice that affects the downstream generative capabilities of a RAG system. The privacy results exhibit the well-studied privacy-utility tradeoff, namely that higher utility losses generally also translate to better privacy protection. This trend, however, is not absolute, and the case of PII deletion presents an intriguing counterpoint. Despite retaining relatively high utility, this technique also showcases the best privacy score for BBC News and a top-5 score for Enron Emails
Anonymization Along the RAG Pipeline
IWSPA ’26, June 23–25, 2026, Frankfurt am Main, Germany
Table 5: Averaged utility and privacy results for the three evaluated datasets. PRE indicates performing anonymization on the original data, and POST anonymization postgeneration. RL, CS, PPL, and LLM-J denote the evaluation methods ROUGE-L, cosine similarity, perplexity, LLM-as-aJudge, respectively. The best-scoring value per dataset/metric is bolded. For LLM-J, 0 is the best, while 100 denotes poor privatization. 𝑇𝑂 is the trade-off between utility and privacy.
Method PII Deletion PII Labeling PII Synthetic data 1-Diffractor (𝜀=1) 1-Diffractor (𝜀=2) 1-Diffractor (𝜀=3) DP-Prompt (𝜀=150) DP-Prompt (𝜀=200) DP-Prompt (𝜀=250) DP-MLM (𝜀=50) DP-MLM (𝜀=75) DP-MLM (𝜀=100)
RL ↑ PRE POST 0.47 0.93 0.46 0.89 0.38 0.61 0.47 0.76 0.55 0.90 0.61 0.96 0.31 0.42 0.37 0.55 0.40 0.61 0.33 0.40 0.34 0.43 0.35 0.44
CS ↑ PRE POST 0.80 0.80 0.78 0.73 0.60 0.72 0.88 0.86 0.92 0.94 0.94 0.97 0.79 0.80 0.86 0.85 0.88 0.87 0.76 0.72 0.78 0.74 0.78 0.75
BBC News PPL ↓ PRE POST 26.45 35.91 26.41 18.11 34.73 21.09 27.64 221.05 26.27 102.61 25.82 70.78 25.46 34.35 22.84 27.92 22.78 26.38 36.88 643.64 35.35 534.74 35.56 509.13
LLM-J ↓ PRE POST 8 23 36 29 14 11 44 41 46 50 48 53 19 15 29 26 31 29 25 28 29 33 29 34
𝑻𝑶 ↑ PRE POST 2.52 5.70 1.46 3.74 1.69 2.66 1.72 3.11 2.04 6.25 2.31 13.43 1.80 2.18 1.84 2.47 1.92 1.61 1.65 1.23 1.61 1.63 1.63 1.19
LLM-J ↓ PRE POST 31 38 48 43 30 18 57 64 66 69 73 75 18 24 27 36 35 43 33 38 37 42 36 44
𝑻𝑶 ↑ PRE POST 1.57 5.90 1.18 4.38 1.40 2.41 1.18 2.06 1.15 4.13 1.08 2.08 1.66 2.49 1.68 2.78 1.62 2.92 1.41 1.46 1.35 1.47 1.42 1.45
LLM-J ↓ PRE POST 40 57 59 67 34 26 71 78 80 84 83 89 17 9 48 37 60 53 23 35 25 39 26 41
𝑻𝑶 ↑ PRE POST 1.76 4.78 1.23 3.17 1.87 2.88 0.83 1.46 0.66 2.66 0.63 4.44 1.35 1.22 1.07 1.21 0.97 1.17 1.55 1.49 1.56 1.48 1.57 1.49
(a) BBC News Method PII Deletion PII Labeling PII Synthetic data 1-Diffractor (𝜀=1) 1-Diffractor (𝜀=2) 1-Diffractor (𝜀=3) DP-Prompt (𝜀=150) DP-Prompt (𝜀=200) DP-Prompt (𝜀=250) DP-MLM (𝜀=50) DP-MLM (𝜀=75) DP-MLM (𝜀=100)
RL ↑ PRE POST 0.33 0.95 0.33 0.92 0.26 0.62 0.47 0.78 0.56 0.90 0.62 0.96 0.35 0.53 0.40 0.68 0.44 0.73 0.36 0.45 0.36 0.48 0.39 0.49
CS ↑ PRE POST 0.79 0.84 0.79 0.82 0.74 0.70 0.8 0.87 0.85 0.95 0.88 0.80 0.66 0.86 0.73 0.86 0.76 0.88 0.69 0.7 0.71 0.73 0.71 0.74
Enron Emails PPL ↓ PRE POST 41.21 44.52 39.53 26.44 40.45 31.92 43.10 245.60 41.25 121.12 40.06 89.16 37.42 81.71 38.95 66.58 39.26 64.06 563.55 586.99 562.57 493.82 561.32 454.08
(b) Enron Emails Method PII Deletion PII Labeling PII Synthetic data 1-Diffractor (𝜀=1) 1-Diffractor (𝜀=2) 1-Diffractor (𝜀=3) DP-Prompt (𝜀=150) DP-Prompt (𝜀=200) DP-Prompt (𝜀=250) DP-MLM (𝜀=50) DP-MLM (𝜀=75) DP-MLM (𝜀=100)
RL ↑ PRE POST 0.50 0.96 0.51 0.94 0.50 0.74 0.45 0.80 0.51 0.92 0.54 0.97 0.27 0.08 0.33 0.27 0.40 0.42 0.32 0.43 0.33 0.45 0.33 0.46
CS ↑ PRE POST 0.82 0.86 0.82 0.85 0.79 0.75 0.86 0.90 0.89 0.96 0.91 0.98 0.51 0.44 0.69 0.69 0.77 0.78 0.69 0.70 0.71 0.73 0.72 0.74
TAB PPL ↓ PRE POST 23.45 28.64 23.35 21.26 21.54 20.40 23.71 154.18 22.90 69.42 22.52 47.70 40.66 3490.07 30.33 280.71 25.86 94.61 33.60 476.82 32.89 389.25 32.32 363.49
(c) TAB
when anonymizing the original data. It also very well protects against private information leakage when applied to generated answers. Beyond this, PII deletion achieves significantly higher trade-offs (𝑇𝑂) than all other tested methods (especially in POST), implying that this simple, yet intuitive method is a solid choice for RAG applications. This, however, is only valid when privacy and utility are weighed equally, which may not always be the case in practical scenarios. Interestingly, entity-based methods perform generally worse on TAB, as per LLM-J, than DP-based methods. In this light, DP-based techniques present an attractive option, allowing for a tunable privacy-utility trade-off, even beyond the bounds of our selected privacy budgets. This tunable feature is something that non-DP-based methods cannot easily offer. Nevertheless, an important perspective is introduced when comparing
anonymization methods without and with formal guarantees (i.e., non-DP versus DP-based methods). The lack of a clearly superior method in both privacy and utility points to the idea that when weighing practical privacy protections, empirical privacy gains may be equally as important as theoretical guarantees. A Tool for Studying Anonymization in RAG.. To enable reproducibility and to empower future works at the intersection of privacy and RAG, we open-source a Streamlit web application, named GuardRAG, to demonstrate our findings in an interactive and analytics-driven manner. We also publish a LIVE variant, which allows the user to upload novel datasets for further exploration and evaluation. The complete code can be found in our repository.
4
Related Work
Recent works investigating anonymization in RAG systems focus on either the retrieval stage [2, 11, 36, 38], by anonymizing user queries [23, 28, 36], or by training privacy-preserving LLM models for RAG generation [8, 32]. Many of these works select specific privacy frameworks as the basis, such as Differential Privacy [10, 12]. Few works, however, investigate direct knowledge base anonymization, and furthermore, no published works explore the juxtaposition of anonymization at two distinct points in the RAG pipeline.
5
Conclusion
We conduct a case study of anonymization in RAG, focusing on the importance of mitigation placement in the multi-stage pipeline of RAG. We find that implementing anonymization at the earliest stage, i.e., on the original data, has relatively significant effects on the eventual outputs of the system. When only anonymizing the answer generated based on the original text, as in our second experiment, utility losses are not as significant, at the cost of weaker privacy protection. We hope that the findings of our case study will lead to similar experiments in the future to shed greater light on mitigations in context, supporting the merits of proposed mitigations in the literature, but also uncovering their potential limitations.
References [1] Andreea-Elena Bodea, Stephen Meisenbacher, Alexandra Klymenko, and Florian Matthes. 2026. SoK: Privacy Risks and Mitigations in Retrieval-Augmented Generation Systems. arXiv:2601.03979 [cs.CR] doi:10.48550/arXiv.2601.03979 [2] Yihang Cheng, Lan Zhang, Junyang Wang, Mu Yuan, and Yunhao Yao. 2025. RemoteRAG: A Privacy-Preserving LLM Cloud RAG Service. In Findings of the Association for Computational Linguistics: ACL 2025. Association for Computational Linguistics, Vienna, Austria, 3820–3837. doi:10.18653/v1/2025.findings-acl.197 [3] Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tai, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2024. Scaling instruction-finetuned language models. J. Mach. Learn. Res. 25, 1, Article 70 (Jan. 2024), 53 pages. doi:10.5555/3722577.3722647 [4] Stav Cohen, Ron Bitton, and Ben Nassi. 2024. Unleashing Worms and Extracting Data: Escalating the Outcome of Attacks against RAG-based Inference in Scale and Severity Using Jailbreaking. doi:10.48550/arXiv.2409.08045 arXiv:2409.08045. [5] Cynthia Dwork. 2006. Differential privacy. In International colloquium on automata, languages, and programming. Springer, 1–12. doi:10.1007/11787006 [6] Wenqi Fan, Yujuan Ding, Liangbo Ning, Shijie Wang, Hengyun Li, Dawei Yin, Tat-Seng Chua, and Qing Li. 2024. A Survey on RAG Meeting LLMs: Towards Retrieval-Augmented Large Language Models. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD
IWSPA ’26, June 23–25, 2026, Frankfurt am Main, Germany
’24). Association for Computing Machinery, New York, NY, USA, 6491–6501. doi:10.1145/3637528.3671470 [7] Xi Fang, Liang Qiao, Jun Shi, and Hong An. 2025. Guardian Angel: A Secure and Efficient Retrieval-Augmented Generation Framework. In 2025 5th International Conference on Artificial Intelligence and Industrial Technology Applications (AIITA). 1773–1777. doi:10.1109/AIITA65135.2025.11047845 [8] Julian Garcia, Jiaqi Gong, Michal Zajac, and Andrew Hahn. 2025. DF-RAG: A Dual Federated Retrieval-Augmented Generation Framework for Collaborative Medical AI. In Proceedings of the ACM/IEEE International Conference on Connected Health: Applications, Systems and Engineering Technologies (Yeshiva University Museum, New York, NY, USA) (CHASE ’25). Association for Computing Machinery, New York, NY, USA, 418–423. doi:10.1145/3721201.3725426 [9] Derek Greene and Pádraig Cunningham. 2006. Practical solutions to the problem of diagonal dominance in kernel document clustering. In Proceedings of the 23rd International Conference on Machine Learning (Pittsburgh, Pennsylvania, USA) (ICML ’06). Association for Computing Machinery, New York, NY, USA, 377–384. doi:10.1145/1143844.1143892 [10] Nicolas Grislain. 2025. RAG with Differential Privacy. In 2025 IEEE Conference on Artificial Intelligence (CAI). 847–852. doi:10.1109/CAI64502.2025.00150 [11] Jiaming He, Cheng Liu, Guanyu Hou, Wenbo Jiang, and Jiachen Li. 2025. PRESS: Defending Privacy in Retrieval-Augmented Generation via Embedding Space Shifting. In ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 1–5. doi:10.1109/ICASSP49660.2025.10887843 [12] Longzhu He, Peng Tang, Yuanhe Zhang, Pengpeng Zhou, and Sen Su. 2025. Mitigating privacy risks in Retrieval-Augmented Generation via locally private entity perturbation. Information Processing & Management 62, 4 (2025), 104150. doi:10.1016/j.ipm.2025.104150 [13] Lijie Hu, Ivan Habernal, Lei Shen, and Di Wang. 2024. Differentially Private Natural Language Models: Recent Advances and Future Directions. In Findings of the Association for Computational Linguistics: EACL 2024. Association for Computational Linguistics, St. Julian’s, Malta, 478–499. https://aclanthology.org/2024. findings-eacl.33 [14] Yangsibo Huang, Samyak Gupta, Zexuan Zhong, Kai Li, and Danqi Chen. 2023. Privacy Implications of Retrieval-Based Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Singapore, 14887–14902. doi:10.18653/v1/2023. emnlp-main.921 [15] Waqar Hussain. 2025. Mitigating Values Debt in Generative AI: Responsible Engineering with Graph RAG. In 2025 IEEE/ACM International Workshop on Responsible AI Engineering (RAIE). 9–12. doi:10.1109/RAIE66699.2025.00006 [16] Bryan Klimt and Yiming Yang. 2004. The enron corpus: A new dataset for email classification research. In European conference on machine learning. Springer, 217–226. doi:10.1007/978-3-540-30115-8_22 [17] Oleksandra Klymenko, Stephen Meisenbacher, and Florian Matthes. 2022. Differential Privacy in Natural Language Processing: The Story So Far. In Proceedings of the Fourth Workshop on Privacy in Natural Language Processing. Association for Computational Linguistics, Seattle, United States, 1–11. doi:10.18653/v1/2022. privatenlp-1.1 [18] Agrim Kulshreshtha, Aditya Choudhary, Tejas Taneja, and Seema Verma. 2025. Enhancing Healthcare Accessibility: A RAG-Based Medical Chatbot Using Transformer Models. In 2024 International Conference on IT Innovation and Knowledge Discovery (ITIKD). 1–4. doi:10.1109/ITIKD63574.2025.11005179 [19] Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-augmented generation for knowledge-intensive NLP tasks. In Proceedings of the 34th International Conference on Neural Information Processing Systems (Vancouver, BC, Canada) (NIPS ’20). Curran Associates Inc., Red Hook, NY, USA, Article 793, 16 pages. doi:10.5555/ 3495724.3496517 [20] Haoran Li, Dadi Guo, Wei Fan, Mingshi Xu, Jie Huang, Fanpu Meng, and Yangqiu Song. 2023. Multi-step Jailbreaking Privacy Attacks on ChatGPT. In Findings of the Association for Computational Linguistics: EMNLP 2023. Association for Computational Linguistics, Singapore, 4138–4153. doi:10.18653/v1/2023.findingsemnlp.272 [21] Moxin Li, Yong Zhao, Wenxuan Zhang, Shuaiyi Li, Wenya Xie, See-Kiong Ng, Tat-Seng Chua, and Yang Deng. 2025. Knowledge Boundary of Large Language Models: A Survey. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, Vienna, Austria, 5131–5157. doi:10.18653/v1/2025.acl-long.256 [22] Pierre Lison, Ildikó Pilán, David Sanchez, Montserrat Batet, and Lilja Øvrelid. 2021. Anonymisation Models for Text Data: State of the art, Challenges and Future Directions. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). Association for Computational Linguistics, Online, 4188–4203. doi:10.18653/v1/2021.acl-long.323 [23] Yiwei Liu, Duo Li, Hu Wang, Pengpeng Zhou, Yuying Xie, and Peng Yin. 2025. Woodpecker: A Locally Deployed Large Language Model for Protecting Sensitive Information via RAG and Semantic Recognition. In 2025 IEEE 19th International
Andreea-Elena Bodea, Stephen Meisenbacher, & Florian Matthes
Conference on Big Data Science and Engineering (BigDataSE). 70–78. doi:10.1109/ BigDataSE66491.2025.00018 [24] Anupam Mehta and Aditya Patel. 2025. Secure Framework for RetrievalAugmented Generation: Challenges and Solutions. IJARCCE 01 (Jan. 2025). doi:10.17148/ijarcce.2025.14114 [25] Stephen Meisenbacher, Maulik Chevli, and Florian Matthes. 2024. 1-Diffractor: Efficient and Utility-Preserving Text Obfuscation Leveraging Word-Level Metric Differential Privacy. In Proceedings of the 10th ACM International Workshop on Security and Privacy Analytics (Porto, Portugal) (IWSPA ’24). Association for Computing Machinery, New York, NY, USA, 23–33. doi:10.1145/3643651.3659896 [26] Stephen Meisenbacher, Maulik Chevli, and Florian Matthes. 2025. On the Impact of Noise in Differentially Private Text Rewriting. In Findings of the Association for Computational Linguistics: NAACL 2025. Association for Computational Linguistics, Albuquerque, New Mexico, 514–532. doi:10.18653/v1/2025.findings-naacl.32 [27] Stephen Meisenbacher, Maulik Chevli, Juraj Vladika, and Florian Matthes. 2024. DP-MLM: Differentially Private Text Rewriting Using Masked Language Models. In Findings of the Association for Computational Linguistics: ACL 2024. Association for Computational Linguistics, Bangkok, Thailand, 9314–9328. doi:10.18653/v1/ 2024.findings-acl.554 [28] Nidhi Mishra, Kanchan Rai, and Dhruv Sharma. 2025. SecureRag: Preventing Sensitive Information Leakage In Rag Pipelines. In 2025 Second International Conference on Pioneering Developments in Computer Science & Digital Technologies (IC2SDT). 600–605. doi:10.1109/IC2SDT68218.2025.11383622 [29] Ildikó Pilán, Pierre Lison, Lilja Øvrelid, Anthi Papadopoulou, David Sánchez, and Montserrat Batet. 2022. The Text Anonymization Benchmark (TAB): A Dedicated Corpus and Evaluation Framework for Text Anonymization. Computational Linguistics 48, 4 (Dec. 2022), 1053–1101. doi:10.1162/coli_a_00458 [30] Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language Models are Unsupervised Multitask Learners. (2019). https://cdn.openai.com/better-language-models/language_models_are_ unsupervised_multitask_learners.pdf [31] Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). Association for Computational Linguistics, Hong Kong, China, 3982–3992. doi:10.18653/v1/D19-1410 [32] Alireza Salemi and Hamed Zamani. 2025. Comparing Retrieval-Augmentation and Parameter-Efficient Fine-Tuning for Privacy-Preserving Personalization of Large Language Models. In Proceedings of the 2025 International ACM SIGIR Conference on Innovative Concepts and Theories in Information Retrieval (ICTIR) (Padua, Italy) (ICTIR ’25). Association for Computing Machinery, New York, NY, USA, 286–296. doi:10.1145/3731120.3744595 [33] Saiteja Utpala, Sara Hooker, and Pin-Yu Chen. 2023. Locally Differentially Private Document Generation Using Zero Shot Prompting. In Findings of the Association for Computational Linguistics: EMNLP 2023. Association for Computational Linguistics, Singapore, 8442–8457. doi:10.18653/v1/2023.findings-emnlp.566 [34] Shang Wang, Tianqing Zhu, Bo Liu, Ming Ding, Dayong Ye, Wanlei Zhou, and Philip Yu. 2025. Unique Security and Privacy Threats of Large Language Models: A Comprehensive Survey. ACM Comput. Surv. 58, 4, Article 83 (Oct. 2025), 36 pages. doi:10.1145/3764113 [35] Chris M. Ward and Josh Harguess. 2025. Adversarial threat vectors and risk mitigation for retrieval-augmented generation systems. In Assurance and Security for AI-enabled Systems 2025, Vol. 13476. International Society for Optics and Photonics, SPIE, 134760A. doi:10.1117/12.3055931 [36] Huanyi Ye, Jiale Guo, Ziyao Liu, and Kwok-Yan Lam. 2025. Efficient PrivacyPreserving Retrieval Augmented Generation with Distance-Preserving Encryption. In 2025 3rd International Conference on Foundation and Large Language Models (FLLM). 668–676. doi:10.1109/FLLM67465.2025.11391120 [37] Shenglai Zeng, Jiankun Zhang, Pengfei He, Yiding Liu, Yue Xing, Han Xu, Jie Ren, Yi Chang, Shuaiqiang Wang, Dawei Yin, and Jiliang Tang. 2024. The Good and The Bad: Exploring Privacy Issues in Retrieval-Augmented Generation (RAG). In Findings of the Association for Computational Linguistics: ACL 2024. Association for Computational Linguistics, Bangkok, Thailand, 4505–4524. doi:10.18653/v1/ 2024.findings-acl.267 [38] Shenglai Zeng, Jiankun Zhang, Pengfei He, Jie Ren, Tianqi Zheng, Hanqing Lu, Han Xu, Hui Liu, Yue Xing, and Jiliang Tang. 2025. Mitigating the Privacy Issues in Retrieval-Augmented Generation (RAG) via Pure Synthetic Data. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Suzhou, China, 24527– 24558. doi:10.18653/v1/2025.emnlp-main.1247 [39] Penghao Zhao, Hailin Zhang, Qinhan Yu, Zhengren Wang, Yunteng Geng, Fangcheng Fu, Ling Yang, Wentao Zhang, Jie Jiang, and Bin Cui. 2026. Retrievalaugmented generation for ai-generated content: A survey. Data Science and Engineering (2026), 1–29. doi:10.1007/s41019-025-00335-5 [40] Yujia Zhou, Yan Liu, Xiaoxi Li, Jiajie Jin, Hongjin Qian, Zheng Liu, Chaozhuo Li, Zhicheng Dou, Tsung-Yi Ho, and Philip S. Yu. 2024. Trustworthiness in Retrieval-Augmented Generation Systems: A Survey. doi:10.48550/arXiv.2409. 10102 arXiv:2409.10102.
Anonymization Along the RAG Pipeline
A
IWSPA ’26, June 23–25, 2026, Frankfurt am Main, Germany
Supplementary
In Table 6, we show examples of anonymized documents alongside the original, non-anonymized document from the Enron Emails dataset. Additionally, we provide the resulting summary text, as produced by gpt-4o-mini for our utility measurements. Table 6: Example anonymized texts (PRE) with their resulting summaries. For readability, we only include the first and third texts (i.e., from the selected privacy budget values) for the DP-based methods. Original text Original text Catherine, I have received the call and will be traveling to Kazakhstan next week to complete the adoption. My husband and son are staying back. I arrive in Almaty on the 14th late in the evening. I’ll spend the day of the 15th in Almaty and then take the overnight train to Taraz. I expect to have a court date in Taraz on the 19th or 20th. Then I will be back in Almaty, probably over Christmas. Finally, I expect to travel to Moscow on around Dec 27 and then return to Houston around Dec 30. This of course is the plan if everything goes exactly on schedule – which it probably won’t. I will certainly miss being away from home at Christmas and am hoping to find a Christmas worship service in Almaty. I’d appreciate any advice in this regard. I would also appreciate advice on what time of clothes are appropriate to bring. Thanks Shelley Corman [email protected] and Generated [Attribute 1: TOPIC/CENTRAL IDEA]: Travel plans for adoption in Kazakhstan. summary [Attribute 2: MAIN DETAILS ABOUT TOPIC/CENTRAL IDEA]: The individual will travel to Kazakhstan for adoption, arriving in Almaty on the evening of the 14th. They will spend the 15th in Almaty before taking an overnight train to Taraz, where a court date is expected on the 19th or 20th. They plan to return to Almaty over Christmas and travel to Moscow around December 27, returning to Houston by December 30. [Attribute 3: IMPORTANT FACTS/EVENTS]: The traveler will be away from home during Christmas and is seeking advice on finding a Christmas worship service in Almaty and what clothing to bring. PII Deletion Anonymized , I have received the call and will be traveling to to complete the adoption. My husband and son are staying back. I arrive in on evening. I’ll spend in and then take the overnight train to . I expect to have a court date in on . Then I will be back in text , probably . Finally, I expect to travel to on and then return to around . This of course is the plan if everything goes exactly on schedule – which it probably won’t. I will certainly miss being away from home at and am hoping to find a worship service in . I’d appreciate any advice in this regard. I would also appreciate advice on what time of clothes are appropriate to bring. Thanks and Generated [Attribute 1: TOPIC/CENTRAL IDEA]: Adoption travel plans summary [Attribute 2: MAIN DETAILS ABOUT TOPIC/CENTRAL IDEA]: The individual will travel to complete an adoption, with their husband and son remaining at home. They will arrive in the destination city on a specific evening, spend a few days there, and then take an overnight train to another location for a court date. [Attribute 3: IMPORTANT FACTS/EVENTS]: The travel schedule includes a court date, a potential return home, and a desire to find a worship service while away. The individual is also seeking advice on appropriate clothing for the trip. PII Labeling Anonymized <PERSON>, I have received the call and will be traveling to <LOCATION><DATE_TIME>to complete the adoption. My husband and son are staying back. I arrive in <LOCATION>on <DATE_TIME>evening. I’ll spend <DATE_TIME>in text <LOCATION>and then take the overnight train to <PERSON>. I expect to have a court date in <LOCATION>on <DATE_TIME>. Then I will be back in <LOCATION>, probably <DATE_TIME>. Finally, I expect to travel to <LOCATION>on <DATE_TIME>and then return to <LOCATION>around <DATE_TIME>. This of course is the plan if everything goes exactly on schedule – which it probably won’t. I will certainly miss being away from home at <DAT_TIME>and am hoping to find a <DATE_TIME>worship service in <LOCATION>. I’d appreciate any advice in this regard. I would also appreciate advice on what time of clothes are appropriate to bring. Thanks <PERSON><EMAIL_ADDRESS>and Generated [Attribute 1: TOPIC/CENTRAL IDEA]: Travel plans for adoption process. summary [Attribute 2: MAIN DETAILS ABOUT TOPIC/CENTRAL IDEA]: The individual will travel to a specific location on a designated date to complete an adoption, while their husband and son will remain at home. They plan to arrive in another location in the evening, spend a day there, and then take an overnight train to a different person. A court date is expected in another location, followed by a return trip. [Attribute 3: IMPORTANT FACTS/EVENTS]: The travel schedule includes multiple locations and dates, with a focus on the adoption process. The individual expresses a desire to attend a worship service during their travels and seeks advice on appropriate clothing for the trip. PII replacement with synthetic data Anonymized Xinh Cong, I have received the call and will be traveling to Wundemar Nioxiib 3n9arin to complete the adoption. My husband and son are staying back. I arrive in Kevala Rekigirto on 9tb9br1nu exterior. I’ll spend 5szzx9al in Grigustr 774ubxf7 text and then take the overnight train to Wafena. I expect to have a court date in Yill in nm9najkeoidu BETl2d on nukjimderf write-date. Then I will be back in Wu!ngmajp, probably brn1m4rf39d. Finally, I expect to travel to -lj9ijrnq5Cook arr tlspdckkjoa on b3bfscco0 as a likelihood Bia evenm lar, I expect wr csfms j3spdo2r1ku land then return to Kulohjcoaoj oac56rrn around 98cqysv0. This of course tg ae przbfrrn at 7hswmfq1 and am hoping to find a b-bd4p9moyqvijs service in Hhfia9oyma. I’d appreciate any advice in this regard. I would also appreciate advice on what time of clothes are appropriate to bring. Thanks Qxfnbxhe syvq1cy2h and [email protected]. Generated [Attribute 1: TOPIC/CENTRAL IDEA]: Travel plans for adoption process. summary [Attribute 2: MAIN DETAILS ABOUT TOPIC/CENTRAL IDEA]: The individual will travel to Wundemar Nioxiib for adoption, with their husband and son remaining behind. They will arrive in Kevala Rekigirto and spend time in Grigustr before taking an overnight train to Wafena. A court date is expected in Yill, and there are plans to return to Wu!ngmajp afterward. [Attribute 3: IMPORTANT FACTS/EVENTS]: The travel includes specific locations and dates, such as arriving in Kevala Rekigirto, spending time in Grigustr, and a court date in Yill. There is a request for advice on suitable clothing and services in Hhfia9oyma. 1-Diffractor (𝜀 =1) Anonymized sharon , i have received the names and will be travelling to kazakhstan final month to completed the adoption . my wife and mother are staying back . i arrive in almaty on the 14th late in the morning . i ’ll expenditures the day of the 15th in text almaty and then goes the dropout fingernails to taraz . i expects to have a courts extant in taraz on the 19th or 20th . then i will be back in almaty , probably over washington . finally , i anticipate to travels to ukraine on about sep 26 and then returned to houston around sr 30. this of course is the anticipate if everything goes exactly on schedule – which it think wo n’t . i will definitely miss being away from home at christmas and am hope to look a christmas religious service in almaty . i ’d shout any wisdom in this regard . i would well applauding counseling on what days of clothes are aptly to bring . thanks shelley corman shelley.corman @ enron.com and Generated [Attribute 1: TOPIC/CENTRAL IDEA]: Adoption travel plans to Kazakhstan and Ukraine. summary [Attribute 2: MAIN DETAILS ABOUT TOPIC/CENTRAL IDEA]: The individual will travel to Kazakhstan for adoption, arriving in Almaty on the 14th and spending the 15th there before going to Taraz for a court appearance on the 19th or 20th. They plan to return to Almaty and then travel to Ukraine around September 26, with a return to Houston expected around September 30. [Attribute 3: IMPORTANT FACTS/EVENTS]: The individual expresses concern about missing Christmas at home and hopes to attend a Christmas service in Almaty. They also seek advice on appropriate clothing for the trip. 1-Diffractor (𝜀 =3) Anonymized catherine , i have received the call and will be traveling to kazakhstan next week to complete the adoption . my husband and son are staying back . i arrive in almaty on the 14th late in the evening . i ’ll spend the day of the 15th in almaty and text then take the dropout fingernails to taraz . i anticipate to have a court date in taraz on the 19th or 20th . then i will be back in almaty , probably over christmas . finally , i expect to travel to russia on around dec 27 and then return to houston around dec 30. this of lesson is the plan if everything goes exactly on schedule – which it probably wo n’t . i will certainly miss being away from home at christmas and am hoping to find a christmas worship service in almaty . i ’d appreciate any advice in this regard . i would also appreciate advice on what time of clothes are appropriate to bring . thanks shelley corman shelley.corman @ enron.com and Generated [Attribute 1: TOPIC/CENTRAL IDEA]: Travel plans for adoption in Kazakhstan. summary [Attribute 2: MAIN DETAILS ABOUT TOPIC/CENTRAL IDEA]: The individual will travel to Kazakhstan next week for adoption, arriving in Almaty on the evening of the 14th. They will spend the 15th in Almaty before traveling to Taraz for a court date expected on the 19th or 20th. They plan to return to Almaty over Christmas and then travel to Russia around December 27, returning to Houston by December 30. [Attribute 3: IMPORTANT FACTS/EVENTS]: The individual expresses concern about being away from home during Christmas and seeks advice on finding a Christmas worship service in Almaty and appropriate clothing to bring. DP-Prompt (𝜀 =150) Anonymized Patricia will you make that scheduled appropriately? Offical names available in Kazabohumi describe house situation. Greater building area produces Nuverslea State Kattama Boxttles launch with totry product line during heartbreakers during text Sweet Tent Week Cookstraker Hurry bloodwreces for 3, 63 and any discounts? etc Good stuff. Latin 2001 G economic Donzam District Justify verdict vacant Yu ? receive September 2 artyupizulubsum folosul crekov perhallami menu I plan to minimize disruption and push finishes to owners nd chat rom it disrupting my workshop sawtable. await attend the exhibition." direct access to Delta Flight attendants. kmigma expect al Moscow Elizabethlane Boulevard 7, judgelights ACP flight 33 0207 11 The Philanthropist nurture science notes us adult resurrection knighthood. That capitalist nostalgia to meet science of is exactly spirits six guys center of political future anxious to discuss it with crime chain of sci fi football making you enjoy going ### You still have mailing chapter allowance to awesometiple Osborne page presented 73/05/02 07/10 Q Overtashi ladies of Tatnix Values Have some left many The Raj. One speed remains Tal on Silverbond. Please return all money local police. Plenty more received Generated [Attribute 1: TOPIC/CENTRAL IDEA]: The text discusses various logistical and planning details related to an event or project, including scheduling and product launches. summary [Attribute 2: MAIN DETAILS ABOUT TOPIC/CENTRAL IDEA]: It mentions the need for appropriate scheduling by Patricia, references a product line launch during a specific event, and discusses minimizing disruption to a workshop. There are also mentions of discounts and attendance at an exhibition. [Attribute 3: IMPORTANT FACTS/EVENTS]: Key points include the launch of the Nuverslea State Kattama Boxttles, a specific date of September 2, and the importance of direct access to Delta Flight attendants. Additionally, there are references to local police and financial matters regarding returned money. DP-Prompt (𝜀 =250) Anonymized I was told that the adoption is all set for June 7th. Is that ok? We must have made a mistake. How long does it take to get to Almaty? I’m taking a train to Taraz on the 15th. !! I will be back in Almaty, probably over Christmas. happy holidays text This may have been to magical effect; at least that’s what everyone involved in the project was hoping for. hyacinth What can I do seperately? Late night. Our team would like to thank you for your service with Enron Canada. The Delaney Corman team wishes you much success in your future endeavors Generated [Attribute 1: TOPIC/CENTRAL IDEA]: Discussion about an upcoming adoption and travel plans. summary [Attribute 2: MAIN DETAILS ABOUT TOPIC/CENTRAL IDEA]: The adoption is scheduled for June 7th, and there is a concern about a possible mistake. The speaker is taking a train to Taraz on the 15th and will return to Almaty around Christmas. The team expresses gratitude for service with Enron Canada. [Attribute 3: IMPORTANT FACTS/EVENTS]: The adoption date is confirmed for June 7th, travel to Taraz is planned for the 15th, and there is a holiday greeting included. DP-MLM (𝜀 =50) Anonymized „ My have receives the phone and will be crossing to Kazakh upcoming day to resume the adopted. The family and son are arguing behind. Cas begin in Kazakhstan on the 14 late in the evening. They’ll save the evening of the 15 in Sochi and text then place the overnight ferry to Istanbul. You forecast to have a release day in Tehran on the ninth or 20. : I will be up in Kazakhstan, late over December. On, My wish to leave to Helsinki on round Sept 22 and then resume to Boston for Dec 15. Which of courses is the plan if everyone acts immediately on planned–which it probably wo ’t. I will essentially wear being across from family at Holiday and am hoped to search a Holiday worship concert in Kazakhstan. I d prefer any prevention in this regarding. And could overall admire inspiration on what type of shoes are logical to bringing. From Shelley Norman thanks @ kathryn and Generated [Attribute 1: TOPIC/CENTRAL IDEA]: Travel plans and family dynamics during the holiday season. summary [Attribute 2: MAIN DETAILS ABOUT TOPIC/CENTRAL IDEA]: The individual plans to travel to Kazakhstan, with a schedule that includes a ferry to Istanbul and a release day in Tehran. There are family disagreements, and the person expresses a desire to attend a holiday concert in Kazakhstan. [Attribute 3: IMPORTANT FACTS/EVENTS]: Travel is set to begin on the 14th of December, with an overnight stay in Sochi on the 15th. The individual wishes to leave for Helsinki on September 22 and return to Boston by December 15. There is a request for advice on appropriate footwear for the trip. DP-MLM (𝜀 =100) Anonymized Yes, They have begun the call and will be responding to Kazakhstan every week to obtain the adoption. The wife and daughter are settling behind. They arrival in Kazakhstan on the 14 mid in the night. I ll face the week of the tenth in text Kazakhstan and then board the night train to Tehran. I expected to have a case hearing in Tehran on the 19 or 20. And I will be over in Moscow, then over Holiday. Yes, I expectation to come to Sochi on around Dec 28 and then back to Texas round Dec 60. This of courses is the planning if stuff kicks exactly on time–which it probably w ’t. I will completely appreciate being upset from home at Christmas and am requesting to found a Christmas church meeting in Kazakhstan. We d give any guidance in this .[. We should then expect experience on what time of clothes are unnecessary to take. From Shelley Cameron shelley @ en and Generated [Attribute 1: TOPIC/CENTRAL IDEA]: Adoption process and travel plans related to Kazakhstan and Tehran. summary [Attribute 2: MAIN DETAILS ABOUT TOPIC/CENTRAL IDEA]: The individual is involved in an adoption process that requires weekly communication with Kazakhstan. They plan to travel to Kazakhstan on the night of the 14th, spend the week of the 10th there, and then take a night train to Tehran for a case hearing expected on the 19th or 20th. They also intend to visit Moscow and Sochi before returning to Texas around December 30. [Attribute 3: IMPORTANT FACTS/EVENTS]: The individual expresses concern about being away from home during Christmas and is seeking guidance for a Christmas church meeting in Kazakhstan. They also mention the need for advice on what clothing to pack for the trip.