CI-Work: Benchmarking Contextual Integrity in Enterprise LLM Agents Wenjie Fu1 * Xiaoting Qin2† Jue Zhang2† Qingwei Lin2 Lukas Wutschitz2 Robert Sim2 Saravan Rajmohan2 Dongmei Zhang2 1 Huazhong University of Science and Technology, China 2 Microsoft [email protected], {xiaotingqin, juezhang}@microsoft.com Disclaimer: This work evaluates privacy risks in simulated enterprise agent environments using synthetic data. Findings should not be interpreted as validation for deployment without additional safeguards.
arXiv:2604.21308v1 [cs.CR] 23 Apr 2026
Abstract Enterprise LLM agents can dramatically improve workplace productivity, but their core capability, retrieving and using internal context to act on a user’s behalf, also creates new risks for sensitive information leakage. We introduce CI-Work, a Contextual Integrity (CI)grounded benchmark that simulates enterprise workflows across five information-flow directions and evaluates whether agents can convey essential content while withholding sensitive context in dense retrieval settings. Our evaluation of frontier models reveals that privacy failures are prevalent (violation rates range from 15.8%–50.9%, with leakage reaching up to 26.7%) and uncovers a counterintuitive tradeoff critical for industrial deployment: higher task utility often correlates with increased privacy violations. Moreover, the massive scale of enterprise data and potential user behavior further amplify this vulnerability. Simply increasing model size or reasoning depth fails to address the problem. We conclude that safeguarding enterprise workflows requires a paradigm shift, moving beyond model-centric scaling toward context-centric architectures.
1
Introduction
Large Language Models (LLMs) have evolved from static text generators to dynamic agents capable of leveraging external tools to navigate complex environments (Yao et al., 2023; Schick et al., 2023). Integration of these agents into enterprise workflows represents a paradigm shift in productivity, transforming them into active assistants with direct access to proprietary data stores, ranging from emails to meeting transcripts, to execute sophisticated tasks (Microsoft, 2025; Anthropic, 2026). However, this utility introduces a critical security paradox: the very mechanism that empowers * Work is done during an internship at Microsoft. †
Corresponding authors. The data and source code are available at https://aka. ms/ci-work.
Figure 1: Risk model of enterprise LLM agents.
agents, the ability to retrieve and manipulate vast amounts of internal data, simultaneously positions them as potential vectors for sensitive information leakage. As illustrated in Figure 1, enterprise workflows require agents to disentangle essential information from sensitive one. The failure to distinguish between the two results in violations of Contextual Integrity (CI) (Nissenbaum, 2004), where information flows breach privacy norms regarding who receives what data and in which context. Although recent studies have increasingly utilized CI theory to evaluate agent privacy, they remain focused on daily assistant tasks and fail to capture the complexities of professional environments. First, current evaluations typically isolate single information flows, overlooking the parallel flows inherent in enterprise settings where agents retrieve multiple entangled entries simultaneously (Cheng et al., 2024; Li et al., 2025). Second, while prior studies incorporate utility metrics, their evaluations are constrained by simplistic, isolated contexts (Mireshghallah et al., 2024a; Shao et al., 2024). In these settings, utility metrics primarily measure the extent of task completion rather than the contextual precision required to disentangle essential data from sensitive information. Finally, prior studies often rely on simplified contexts or short attributes that fail to replicate the scale and density of real enterprise data (Mireshghallah et al., 2025), where agents must discern sensitive needles in a corporate haystack. To bridge these gaps, we introduce CI-Work, a benchmark designed to evaluate the CI of LLM
agents within high-fidelity enterprise workflows. CI-Work simulates complex corporate dynamics across five distinct organizational directions (e.g., Upward, Lateral). Moving beyond the isolated scenarios of prior benchmarks, each instance in CI-Work requires the agent to navigate a dense retrieval context partitioned into an Essential Set and a Sensitive Set, explicitly quantifying the trade-off between task utility and privacy adherence. To ensure high realism, we employ a rigorous construction pipeline combining human-in-the-loop seed generation with an automated self-iterative refinement mechanism, ensuring scenarios strictly align with business logic and hierarchical constraints. Using CI-Work, we find that frontier LLM agents exhibit a persistent privacy-utility trade-off that is amplified by the high scale and density characteristic of realistic enterprise data. We observe that while models grasp high-level organizational boundaries, they struggle to adjudicate fine-grained information flows. This fragility is further exposed under potential user behavior, where even unintentional instruction precipitate a dual collapse of privacy and utility, causing agents to simultaneously leak more sensitive information while failing to convey essential data. Crucially, such vulnerability cannot be simply addressed by increasing model size and reasoning effort, even leading to an “inverse scaling” phenomenon, where larger models paradoxically exacerbate leakage rather than mitigating it. These findings underscore the inadequacy of current safety alignment for professional domains and highlight the urgent need for contextaware privacy mechanisms in enterprise agents.
2
Related Works
Contextual Privacy Benchmarks. Prior work has increasingly leveraged Nissenbaum’s Contextual Integrity theory (Nissenbaum, 2004) to evaluate privacy reasoning capabilities in LLMs. Early benchmarks focused on static reasoning: ConfAide (Mireshghallah et al., 2024b) assesses tier-based information disclosure, while CI-Bench (Cheng et al., 2024) develops large-scale synthetic datasets to test privacy norms across diverse domains. Recent research has shifted toward agentic dynamics and persistent states: PrivacyLens (Shao et al., 2024) and PrivacyLensLive (Wang et al., 2025) evaluate leakage in evolving agent trajectories and realistic multiagent workflows respectively, while CIMemo-
ries (Mireshghallah et al., 2025) measures how violations accumulate in persistent memory. However, these studies predominantly focus on general daily life, building from real-world court cases (Fan et al., 2024), IoT device (Shvartzshnaider and Duddu, 2025) or web agent assistant perspective (Ghalebikesabi et al., 2025; Zharmagambetov et al., 2025). In contrast, our work specifically targets the enterprise domain, where privacy norms are governed by complex, implicit organizational hierarchies and proprietary workflows. Enterprise Agents Evaluation. Various benchmarks have been established to evaluate the task execution capabilities of LLM agents in enterprise scenarios. Workbench (Styles et al., 2024) assesses the performance of agents in accomplishing complex tasks within enterprise contexts. OfficeBench (Wang et al., 2024) further extends the evaluation to office tasks across multiple applications. WorkArena (Drouin et al., 2024) focuses on testing web agents on typical daily knowledge work tasks. Beyond general enterprise tasks, subsequent benchmarks are designed for more specific professional domains. TheAgentCompany (Xu et al., 2024) simulates a small software company environment, while CRMArena (Huang et al., 2025) evaluates LLMs’ performance on customer service workflows. HERB (Choubey et al., 2025) evaluates deep search capabilities over heterogeneous enterprise data. However, these benchmarks primarily focus on agent utility and task success, while largely overlooking the critical risks of privacy leakage and sensitive information exposure. In contrast, our work specifically addresses this gap by investigating the challenges and solutions related to safeguarding contextual privacy when deploying LLM agents in enterprise environments.
3
Risk Model of Enterprise LLM Agents
We consider an enterprise environment as shown in Figure 1 where a user u possesses a private, unstructured data store D. An LLM-based agent A is tasked with an instruction I that requires retrieving information from D and sharing it to a recipient. The agent interacts with D via a toolkit T , generating an execution trajectory H = {(at , ot )}Tt=1 . At each step t, the action at invokes a retrieval tool τ ∈ T with query parameters qt , yielding a set of data entries ot = τ (D, qt ) as the observation. Let E = ⋃Tt=1 ot denote the cumulative set of retrieved entries. To evaluate CI, we partition E into two disjoint subsets based on the task context: the es-
sential set Eess , containing entries indispensable for fulfilling I, and the sensitive set Esens , containing entries that violate privacy norms if disclosed to the recipient. Since the contexts retrieved by the retrieval tool are typically semantically related, we do not explicitly discuss a set of irrelevant contexts. The privacy risk arises when the agent’s final response af in discloses any entries from Esens while conveying entries from Eess to complete the task.
4
CI-Work Benchmark
Drawing on the theory of CI, an information flow can be described by five key parameters: data subject, sender, recipient, data type and transmission principle (Nissenbaum, 2004). Unlike existing benchmarks that isolate single information flow (Shao et al., 2024; Cheng et al., 2024), CIWork simulates the entangled nature of enterprise workflows where multiple parallel flows coexist. To capture this complexity, all test cases in CIWork are constructed and evaluated in four stages (depicted in Figure 2): 1) Task-oriented Seed Generation (§ 4.1), 2) Contextual Entries Generation (§ 4.2), 3) Case Episode Generation (§ 4.3), and 4) Trajectory Simulation and Evaluation (§ 4.4). 4.1
Task-oriented Seed Generation
A task-oriented seed S is composed of a sender, a recipient, and a task assigned by the sender (transmission principle), that yields an information flow direction from the sender to the recipient. Leveraging standard organizational communication taxonomy (Robbins and Judge, 2009; Slack, 2025), we categorize these flows into five distinct directions: 1) Downward (management to staff), 2) Upward (reporting to superiors), 3) Lateral (peer collaboration), 4) Diagonal (cross-organizational), and 5) External (stakeholder engagement). To construct a high-fidelity benchmark, we employed a continuous human-in-the-loop generation paradigm. We began by manually crafting a set of high-quality seed exemplars across diverse industries (e.g., technology, healthcare, finance). These served as fewshot demonstrations for Gemini-3-Pro (DeepMind, 2025), where human experts interactively monitored the generation stream, filtering out implausible scenarios and refining prompts in real-time to ensure structural diversity. Subsequent validation by enterprise practitioners confirmed that the finalized seeds exhibit high realism and strictly adhere to actual business practices (see Appendix C).
Algorithm 1: Self-Iterative Refinement Input :Task Seed S, Max Iterations Tmax Output :Refined Entry Set E 1 E ← GenInitialEntries(S) t ← 0; 2 while t < Tmax do 3 is_consistent ← true 4 foreach entry e ∈ E do 5 ℓ⋆ ← GetIntendedLabel(e) 6 ℓ̂ ← Evaluate(e) 7 if ℓ̂ ≠ ℓ⋆ then 8 feedback ← Critique(e, ℓ̂, ℓ⋆ ) 9 e ← Refine(e, feedback) 10 is_consistent ← false
12
if is_consistent then break;
13
t ← t + 1;
11
14
return E
4.2
Contextual Entries Generation
Initial Entries Generation. The task-oriented seeds only demonstrate the direction of information flow, the data type is still missing from the key parameters of CI. To fill this gap, a series of contextual entries are required to instantiate a set of parallel information flows for each seed, which yields a key challenge. To address this challenge, we employ an LLM-driven, template-based approach to automate the generation process. We prompt the LLM to synthesize all essential and sensitive entries for each seed in a single pass, explicitly instructing it to craft every item as a distinct, atomic summary of a discrete information unit. This ensures that entries are ready for independent instantiation during trajectory simulation, while strictly aligning with both the seed’s transmission principle and the definitions of essential versus sensitive information. Keeping Eess and Esens strictly disjoint at this stage establishes a clean ground truth for evaluation; the realistic entanglement of sensitive and essential content is re-introduced later during trajectory simulation (§4.4). The sensitive entries are categorized into nine distinct types based on their content and associated privacy risks (see Appendix A). Self-iterative Refinement. However, prior work has observed that LLM-synthesized data deviates from privacy norms (Li et al., 2025), and this phenomenon becomes more severe when multiple con-
(1) Task-oriented Seed Generation
Crafted Seeds
Filter
w-
sh
ot
LLM
Fe
Curate
Human
Synthetic Seeds
Data Sender: Alice, a data scientist Task: Request approval for additional GPU budget due to model training failures Data Recipient: Tom, the department head
(3) Case Episode Generation
Seed
Essential Sets
Sensitive Sets
(2) Contextual Enteries Generation
Instruction: Send an email to Tom requesting approval for additional GPU budget so I can address the failures LLM Outline: Alice is a data scientist responsible for ... . Tom is the department head ... . Alice asks the Agent to ... Retrieval Trajectory (Tool-based): Episode (1) From Email: retrieve E1, S2, S3 (2) From Chat: retrieve E2, S1 (3) From Calendar: retrieve E3, ...
✓
Evalu ate
Seed
ize
LLM
Critic
Essential Set
k dbac
Fee
Refin e
Sensitive Set
Essential Entry Set (Should share) E1. previous budget approval emails E2. cost-benefit analysis of GPU resources E3. current GPU usage statistics E4. ...
✗
Sensitive Entry Set (Should not share) S1. time spent on video games during failures S2. complaints regarding unfair allocation S3. records of GPU usage for personal matters S4. ...
(4) Trajectory Simulation and Evaluation SearchEmail ("GPU")
Agent
{"To": "...", "Subject": "...", "Body": "..."} Simulator
Revealed Entries: E1, E2, E4, S1 Dropped Entries: S2, S3, S4, E3
Leakage: Final Action: SeedEmail Input: {"To": "Tom", "From": "Alice", "Subject": "Request approval for Violation: additional GPU", "Body": " ... current GPU usage is 95%. During downtime Conveyance: we pass the time by playing games"}
Figure 2: Overview of the CI-Work construction and evaluation pipeline.
text entries are generated in a single pass. We draw insight from recent CI benchmarks that highlight LLMs are highly aligned with humans when labeling or evaluating sensitive information (Shao et al., 2024; Mireshghallah et al., 2025). To bridge this generation-evaluation gap, we introduce a selfcorrection mechanism inspired by iterative refinement techniques (Madaan et al., 2023; Wang et al., 2023) that enhances the quality of generated responses without human intervention. For each case, after generating the initial entries, LLM performs a blind classification of all entries into Essential, Sensitive, and Ambiguous. Any discrepancy between the intended category ℓ⋆ and the model’s perceived label ℓ̂ triggers an automated revision loop, where the entry is revised based on the classification reason. The pseudocode of self-correction is provided in Algorithm 1. Subsequent human verification shows that the Essential/Sensitive labels align with annotators’ privacy-norm judgments at 82.5%–95.0% agreement (Appendix C). 4.3
Case Episode Generation
Although our theory-guided schema enables the privacy sensitive seeds to be contextual, these seeds have limited details for guiding trajectroy simulation. For instance, the data sources of contextual entries are not specified under the agent interaction trajectory. As shown in Figure 2, we further generate a detailed case episode P for each seed, including a brief description of the sender, the recipient,
and the current environment, the instruction that the sender provides to the agent, as well as the entries that the agent can retrieve under each available tool. Overall, the case episode depicts a coherent and realistic enterprise scenario, which acts as a semantic blueprint to be concretely instantiated into full textual observations during trajectory simulation. 4.4
Trajectory Simulation and Evaluation
Trajectory Simulation. To evaluate LLMs over the constructed CI-Work benchmark in an agent setting, we developed a tool-centric simulation environment based on ToolEmu (Ruan et al., 2024) and PrivacyLens (Shao et al., 2024) that employs an LLM to mimic the observation yielded by various enterprise tools (e.g., email, chat, calendar, meeting, etc.) during task execution. The sandbox environment was further adapted to support the simultaneous instantiation of multiple essential and sensitive entries, enabling the simulation of enterprise scenarios in which an agent retrieves multiple semantically related pieces of content at once. Crucially, when instantiating a sensitive entry into a concrete artifact (e.g., meeting transcript), the simulator is allowed to weave in contextually relevant non-sensitive content alongside the sensitive atom, mirroring the mixed nature of real enterprise data. At each step, when the agent invokes a specific tool to retrieve information, the observation is generated by an LLM-powered simulator: ot = sim(qt , P, dτ ), which synthesizes the tool’s
4.5
Benchmark Instantiation and Statistics
We curate 25 task-oriented seeds spanning five information-flow directions and synthesize an additional 100 seeds via interactions with Gemini3-Pro, yielding 125 seeds in total. We employ GPT-5.2 for both benchmark generation (entries, episodes, trajectories) and evaluation. Unless otherwise specified, each seed is instantiated with 4 sensitive and 4 essential entries, resulting in 1,000 contextual entries overall. LLM agents are deployed with ReAct (Yao et al., 2023), which require the LLM to reason before taking actions. Detailed benchmark statistics regarding data types and domains are provided in Appendix A.
5
Evaluating Frontier LLMs on CI-Work
We evaluate a wide range of frontier LLMs, including four open-source LLMs and five close-source LLMs (refer Appendix B for the exact models). 5.1
Overall Performance
We summarize the leakage, violation, and conveyance performance of all LLMs across five enterprise information flows in Table 1. Overall, we find that current frontier LLMs fail to adequately protect contextual privacy in the enterprise scenario, with violation rates ranging from 15.80% (DeepSeek-R1) to 50.87% (Grok-3) and leakage rates remaining non-trivial up to 26.66% (mostly