ASLEval: Measuring Privacy Exposure Displacement in LLM Agent Sessions Guosen Wu, Huizhen Huang, Guoxiong Long, Tao Huang, Chen Hou
arXiv:2609.18864v1 [cs.CR] 16 Sep 2026
School of Computer and Big Data, Minjiang University, Fuzhou, China [email protected], [email protected], [email protected], [email protected], [email protected]
Abstract Privacy evaluations of tool-using LLM agents often inspect a designated action, final response, or attacker report. These local proxies can miss unauthorized exposure elsewhere in a multi-step session and lack common ground truth across outlets, reports, and tool paths. We introduce privacy exposure displacement—the mismatch between a local evaluation proxy and target-grounded session exposure—and ASLEval, an authorization-aware framework that pre-registers a hidden target set, measures all declared visible exits, and reserves internal traces for diagnosis. Across multiple enterprise-style environments and independently implemented runtimes, we observe three recurring patterns. An expected-outlet-only view misses 46.9% of exposure recovered by the visible-exit union; attacker self-reports combine omissions with high false discovery; and schema-aligned internal evidence usually precedes visible exposure at the request/probe level. Reducing model-visible returns changes this path but can eliminate normal-task success. Independent human review supports the adjudication pipeline while identifying harder console and candidate cases. These findings motivate benchmarks that declare the complete visible boundary, ground claims in prespecified targets and authorization, and report privacy together with task utility.
1
Introduction
An email guard can make an evaluation appear successful even when the same protected entities remain visible to the requester through console feedback. This is not an edge case unique to email. Tool-using agents retrieve workspace records, bind values to structured actions, expose intermediate feedback, and reuse observations across multiple turns. Yet privacy and safety benchmarks often inspect one designated action, final response, file-sharing event, or red-team report (Debenedetti et al. 2024; Drouin et al. 2024; Wang et al. 2025; Zharmagambetov et al. 2025). A local proxy is easy to score, but it may represent only part of the requestervisible session boundary. We call the resulting measurement failure privacy exposure displacement: a mismatch between a local evaluation proxy and unauthorized, target-grounded exposure over a bounded session. Figure 1 organizes the phenomenon along three questions. C1, outlet displacement, asks where exposure is counted: an expected or guarded exit may cover only part of the exposure visible through all declared exits. Copyright © 2027, Association for the Advancement of Artificial Intelligence (www.aaai.org). All rights reserved.
C2, recognition displacement, asks what is recognized: the attacker-reported candidate set can omit confirmed targets and include false discoveries. C3, path displacement, asks how exposure is mediated: sensitive entities are unevenly associated with structured tool paths and model-visible returns. Across all three, authorization for an agent to access a private workspace is distinct from authorization for the evaluation requester or an external recipient to receive its contents. We use requester to denote the party that receives requester-visible outputs. The target-blind probe agent is an automated stress-testing policy that issues only interfacevalid requests and observes only feedback available to the requester. In the recognition analysis, an attacker report refers specifically to the candidate list submitted by this probe agent; it does not introduce a separate actor. The evaluator is a separate benchmark-side role that fixes the evaluation contract and adjudicates recorded observations after the session. ASLEval measures these mismatches against a common evaluation contract. Before execution, the evaluator fixes a hidden target set, an authorization relation, an expected outlet, and the declared visible session boundary. A target-blind probe agent then issues only interface-valid requests to a black-box target agent. After the session, a target-grounded adjudicator maps every declared observation to the prespecified target set. Requester-visible and external exits count as exposure; raw tool arguments, returns, and execution metadata are retained only as evaluator-side path evidence. This separation compares local proxies, full-session exposure, attacker reports, and tool paths without revealing target labels during interaction. This measurement object complements existing agentsecurity and privacy benchmarks. Privacy in Action and AgentDAM emphasize contextual privacy decisions and task necessity; AgentLeak instruments final and internal communication channels as leakage surfaces; CIPL studies targetindependent channel inversion over attacker-visible observations (Wang et al. 2025; Zharmagambetov et al. 2025; El Yagoubi, Badu-Marfo, and Al Mallah 2026; Huang et al. 2026). ASLEval instead asks whether a local proxy represents unauthorized exposure of a fixed hidden target set across the declared visible session boundary. Across the evaluated environments, the experiments reveal three recurring measurement patterns: outlet-local protection
(a) C1: Outlet Displacement
(b) C2: Recognition Displacement Evaluator-confirmed Attacker-reported exposed targets candidates
Private Workspace target entity e
(c) C3: Path Displacement Private structured record sensitive entities
Schema-aligned tool path
Target LLM Agent
Evaluator-only path evidence
structured arguments and fields
tool arguments
Expected outlet
Other visible exit
CHANNEL GUARD
Email action
Console feedback
channel guard
entity e remains visible
Missed exposures
False discoveries
Correct reports
raw tool returns
Model-visible tool return diagnostic only
may propagate to
Declared visible exits response · send · reply · forward
counted as exposure
e lower outlet exposure
Outlet coverage drops; session exposure persists.
Candidates are not exposure channels.
Rename fields
Minimize returned content
no meaningful reduction
45.7% lower overlap
Return minimization matters; cosmetic renaming does not.
Figure 1: Three manifestations of privacy exposure displacement. C1: the expected outlet can miss exposure visible elsewhere in the session. C2: attacker-reported candidates differ from evaluator-confirmed exposed targets through both omissions and false discoveries. C3: exposure is associated with schema-aligned request/probe paths; reducing model-visible return content changes this path, but is not a cost-free defense.
can improve the audited channel without covering the full visible session; attacker reports behave as noisy predictions rather than exposure ground truth; and structured internal evidence is associated with, and usually temporally precedes, visible exposure. The strength of evidence differs by claim: C1 is replicated across two runtimes and an expanded item set, C2 remains sensitive to difficult candidate matches, and C3 supports request/probe-level ordering rather than a nested causal chain. Our contributions are: • We identify privacy exposure displacement and organize it into outlet, recognition, and path manifestations that capture where, what, and how a local proxy diverges from session exposure. • We introduce ASLEval, an authorization-aware, targetgrounded protocol that pre-registers the evaluation contract and separates declared visible exits from evaluatoronly diagnostic probes. • We provide cross-runtime, cross-model, and crossenvironment evidence for three measurement regularities, together with human validation and implications for privacy benchmark design and tool-interface trade-offs.
2
Related Work
Agent security and local outcome metrics. LLM safety evaluation has progressed from red-teaming and jailbreak discovery (Perez et al. 2022; Zou et al. 2023; Chao et al. 2023) to tool-using agents exposed to indirect prompt injection and harmful actions (Greshake et al. 2023; Toyer et al. 2023; Qin et al. 2023; Ruan et al. 2023). HarmBench, AgentHarm, AgentDojo, and WorkArena measure harmful outcomes, prompt-injection robustness, task completion, or designated external actions (Mazeika et al. 2024;
Andriushchenko et al. 2024; Debenedetti et al. 2024; Drouin et al. 2024). These outcomes remain essential, but a blocked action alone does not establish the absence of unauthorized exposure elsewhere in the session. Agent privacy, minimization, and channel analysis. Privacy in Action evaluates contextual privacy decisions and mitigations in MCP/A2A workflows, while AgentDAM evaluates whether potentially private information is processed only when task-necessary (Wang et al. 2025; Zharmagambetov et al. 2025). AgentLeak measures violations across final and internal multi-agent channels (El Yagoubi, Badu-Marfo, and Al Mallah 2026). CIPL formulates privacy attacks as inversion from a sensitive source to an attacker-visible observation surface (Huang et al. 2026). ASLEval contributes a target-grounded comparison between local evaluation proxies and the complete declared visible boundary. Information flow. Privacy exposure displacement is related to information-flow control and taint analysis (Denning 1976; Sabelfeld and Myers 2003). Agent evidence, however, spans natural-language responses, action payloads, structured tool messages, and external exits. ASLEval provides target-grounded summaries over these heterogeneous observations while preserving a clear distinction between visible exposure and diagnostic path evidence.
3 3.1
Target-Grounded Session Measurement Evaluation Contract and Threat Model
Let A be a black-box target agent operating over a private workspace D. The target agent may retrieve records from D, but access authorization is distinct from disclosure authorization. Before each session, the evaluator fixes an evaluation contract: a hidden target set S ⊂ D, an expected outlet c∗ , a set of declared visible channels, and a recipient rc for each
Work
Ground-truth target
Counted privacy surface
Primary question
Privacy in Action (Wang et al. 2025)
Context-dependent private information
Contextually evaluated actions and messages
Is disclosure appropriate in the interaction context?
AgentDAM (Zharmagambetov et al. 2025)
Task-designated private data
Use or sharing within a web-agent trajectory
Is private-data use necessary for task completion?
AgentLeak (El Yagoubi, BaduMarfo, and Al Mallah 2026)
Policy-defined private content
Final and internal communication channels
Where does leakage occur across a multi-agent stack?
CIPL (Huang et al. 2026)
Sensitive source without a fixed hidden entity set
Attacker-visible channels
observation
Can an observation channel be inverted to recover private information?
ASLEval
Fixed hidden entity set plus authorization relation
All declared unauthorized visible exits; internal probes are diagnostic
Does a local proxy represent targetgrounded session exposure?
Table 1: Positioning of ASLEval relative to representative agent-privacy evaluations. The comparison emphasizes each method’s target definition, counted surface, and measurement question. visible channel c. The benchmark supplies an authorization relation
where νc marks channel c as a declared visible exit or evaluator-only instrumentation. We partition
Auth(rc , e) ∈ {0, 1},
Cvis = {c : νc ∈ {requester-visible, external}}, Cint = {c : νc = internal}.
(1)
where Auth(rc , e) = 1 means that recipient rc may receive entity e. In the primary experiments, the evaluation requester and designated external recipients are unauthorized for every scoped target entity. An evaluator-only negative control re-scores the same raw appearances under authorized and unauthorized roles: unauthorized exposure is 0 and 1, respectively. Because role metadata is not delivered to the target agent, this control validates authorization-aware metric semantics rather than agent policy compliance. The target-blind probe agent cannot access S, undisclosed records, internal tool implementations, guard rules, evaluator channel labels, or adjudication criteria. It issues only interface-valid requests accepted by the environment and observes only feedback exposed by that environment. Each target entity is represented as e = (v, A, τ, σ), where v is a canonical sensitive value, A is a finite alias set, τ is an entity type, and σ records provenance such as benchmark item, source table, and source fields. Targets are derived from benchmark-declared private fields and fixed before execution; probe-agent text cannot create new target entities. The same construction pipeline is used across environments: enumerate declared private values, remove empty or boilerplate entries, canonicalize format variants, attach provenance, and select the session scope by item or seed.
3.2
Session Observations and Measurement Boundary
For request budget B, a session produces channel-labeled observations O = {(t, c, ot,c , νc ) : 1 ≤ t ≤ B, c ∈ C},
(2)
(3)
Only Cvis contributes to privacy exposure. Tool arguments, raw tool returns, and execution metadata in Cint support path diagnosis but are not counted as direct exposure. Post-hoc adjudication maps every observation to the prespecified target set: Lt,c = F(ot,c , S),
Lt,c ⊆ S.
(4)
For each visible exit, the evaluator retains only unauthorized matches: Ut,c = {e ∈ Lt,c : Auth(rc , e) = 0},
c ∈ Cvis . (5)
The resulting entity–channel links form a target-grounded session view. A target is visibly exposed only if it appears in some Ut,c . For C3, the evaluator joins matched internal and visible probes by session and request identifiers, orders them by timestamps with stable event-index tie breaking, and records request/probe-level internal-before-visible paths. Historical logs do not contain parent tool-call identifiers, so this ordering is temporal evidence rather than a nested callgraph causal chain. During execution, the system stores request identifiers, channel types, payloads, action status, and timestamps when available. After the bounded interaction, it builds the scoped alias index, matches recorded payloads, and aggregates entity keys into session, channel, recognition, and path summaries. Target knowledge therefore remains outside the probe-agent interaction loop in zero-feedback conditions.
3.3
Adjudication and Diagnostics
The deterministic adjudicator matches normalized aliases from the scoped target index. An optional semantic verifier handles unresolved cases but may select only from the
ASLEval Target-Grounded Session Measurement Black-box agent session
Private Workspace D
Color key
agent-authorized scoped entities
Target-Blind Probe Agent Can issue: interface-valid requests
requests
Can observe: declared feedback
runtime system
visible exposure
evaluator-only
hidden ground truth
Target LLM Agent black-box target system
Does not know S, guards, or labels
Declared Visible Exits Cvis
Evaluator-Only Probes Cint Tool arguments Raw tool returns Execution metadata diagnostic only
Email · File share Final response · Console counted as privacy exposure Pre-registered evaluation and post-hoc adjudication
exposure observations
not counted as direct exposure
path evidence
Displacement Diagnostics
Evaluation Contract Hidden target set S ⊆ D Authorization Auth(r, e)
C1: Outlet Diagnostics targets and authorization
Expected outlet c* fixed before execution
Target-Grounded Adjudicator
session coverage · outlet-loss fraction
authorization-aware session analysis
C2: Recognition Diagnostics
candidate reports · omissions · false discoveries
C3: Path Diagnostics
internal-visible overlap · schema paths
Figure 2: ASLEval target-grounded session measurement. The evaluation contract fixes the hidden target set S, authorization relation, and expected outlet before execution. A target-blind probe agent interacts with the black-box target agent. Declared visible exits Cvis contribute exposure observations, whereas evaluator-only probes Cint provide diagnostic path evidence. Posthoc adjudication grounds both views in the same target universe and produces C1–C3 diagnostics. evaluator-provided target batch; outputs outside S are rejected. Underspecified mentions and matches requiring hidden database context are also rejected. Two annotators independently reviewed 300 stratified observation–entity pairs and adjudicated disagreements, obtaining 0.850 raw agreement and Cohen’s κ = 0.747. Against the resulting human gold, strict F1 is 0.891 for the primary adjudicator and 0.929 for a second judge; console and candidate-report slices are harder than C3. We therefore retain exact-plus-normalized matching for headline C1 results and report claim-level and match-sensitivity analyses in the supplement. The exposure coverage rate (ECR) measures the fraction of the target set exposed through at least one unauthorized visible exit: SB S t=1 c∈Cvis Ut,c . (6) ECR = |S| S Per-channel coverage is Cc = | t Ut,c |/|S|. C1: outlet displacement. Let Evis be the unauthorized visibly exposed target set in one session and Ec∗ the subset observed at expected outlet c∗ . The outlet-loss fraction is OLF =
|Evis \ Ec∗ | , |Evis |
(7)
with zero assigned when Evis is empty. OLF is the fraction of session exposure missed by the expected outlet. C2: recognition displacement. After a no-feedback session, the probe agent submits candidate strings. We compare candidate coverage and precision with evaluator-confirmed
exposure, and report false-discovery rate and recognition gap ECR − Ccand . C3: path displacement. We compare internal, visible, overlap, and request/probe-level temporal rates across schema-match groups, then replay fixed request sequences while intervening on input names, output field names, and model-visible tool-return content.
4
Experimental Setup
We evaluate ASLEval in three enterprise-style environments that differ in private data, tool schemas, and visible outlets (Table 2). PrivacyInAction (PIA) contains therapy and professional records; AgentDojo Workspace provides email and file-based office tasks; WorkBench spans email, CRM, calendar, analytics, and project-management data. Complete run matrices, seeds, and supplementary conditions are reported in the appendix. The primary probe-agent and target-agent model is deepseek-v4-flash; GLM-4.7 provides cross-model corroboration, and deepseek-v4-pro is the primary adjudicator. The request budget is B = 20. Feedback-Guided Exploration (FGE) provides only coarse progress signals and serves as a diagnostic stress test. Zero-Feedback Exploration (ZFE) and no-judge-feedback mining disclose no runtime target judgments. In every reported run, the benchmark requester and designated external recipients are unauthorized for the scoped target set. Target-agent runtime variation. Agent-R is the original PIA implementation based on the OpenAI Agents SDK and
Environment
Private data domain
Declared visible exits
Evidence scope
Claims
PIA
Therapy and professional records
Email, console
DeepSeek and GLM; two target-agent runtimes; ZFE and FGE
C1, C2
AgentDojo Workspace
Email, files, workspace objects
Email, file sponse/console
Separate tool taxonomy and environment
C1, C2 transfer
WorkBench
Email, CRM, calendar, analytics, and projects
Two model families; fixedrequest replay
C3
and
share,
re-
Response, send, reply, forward
Table 2: Evaluation scope. Exposure metrics use all declared visible exits; evaluator-only tool records are used only for C3 diagnosis.
Evaluation view
Coverage
Final answer only Successful action only Expected outlet only Attacker self-report ASLEval visible union
0.805 0.026 0.478 0.672 0.887
Missed vs. union 0.097 0.963 0.469 0.239 0.000
n 245 32† 245 100 245
Table 3: Alternative evaluation views over the same exported sessions. Coverage and missed rates are session means; bootstrap intervals appear in the supplement. † Only sessions with explicit action-delivery status are observable for the successful-action view. MCP-managed tool execution. Agent-F is an independent nanobot runtime using native JSON-schema function calls, structured role=tool observations, and a dedicated notool-call final response. The matched C1 replication holds the models, workspace, hidden target set, tools, guard policy, request budget, and ZFE logic fixed while changing the orchestration and tool-message runtime. Condition means and uncertainty use the session as the statistical unit, 10,000 bootstrap resamples, and random seed 2027. Paired contrasts use item×seed or dataset×seed keys with exact sign-flip tests. C3 rates are computed per session and then aggregated across seeds; entity observations are not treated as independent significance-test samples.
5 5.1
Results
Comparison with Local Evaluation Views
Finding 0: common local views recover materially different portions of the same visible exposure. We re-analyze 245 exported sessions using five observation rules over identical logs. Table 3 compares final-answer-only, successfulaction-only, expected-outlet-only, attacker self-report, and the ASLEval declared-visible-exit union. The expectedoutlet view misses 46.9% of exposure recovered by the union, while self-report misses 23.9% and has 18.4% candidate precision. These rows are evaluation views over common logs, not end-to-end reproductions of prior systems.
5.2
C1: Outlet Displacement
Finding C1: an outlet-local guard can improve its audited channel while leaving session exposure nearly un-
Runtime Guard Session Email
Console OLF
OLF >0
Agent-R Agent-R Agent-F Agent-F
1.000 0.930 0.933 1.000
4/10 7/10 1/10 8/10
Off On Off On
1.000 0.955 0.983 1.000
0.640 0.317 0.892 0.344
0.360 0.683 0.100 0.656
Table 4: Matched C1 ZFE results across two target-agent runtimes. Session, Email, and Console are target coverage rates; OLF is averaged per session. Each row contains 10 sessions.
changed. We test C1 under matched ZFE conditions in two independently implemented target-agent runtimes. Table 4 and Figure 3 show the same mismatch in both implementations. With the email guard enabled, Agent-R and Agent-F obtain session coverage of 0.955 and 1.000, email coverage of 0.317 and 0.344, console coverage of 0.930 and 1.000, and mean outlet-loss fractions of 0.683 and 0.656. Exposure occurs outside the expected outlet in 7/10 Agent-R sessions and 8/10 Agent-F sessions. A representative guarded Agent-F session contains no confirmed target at the email exit but 12 anonymized target identifiers in requester-visible console observations, yielding OLF = 1.0. This case exemplifies the condition-level outlet displacement summarized in Table 4. Within Agent-F, the email guard reduces email coverage by −0.548 (95% CI [−0.823, −0.260], exact sign-flip p = 0.0312) and increases OLF by +0.556 (95% CI [0.273, 0.828], p = 0.0156), while session and console coverage remain saturated. Under the guarded condition, the Agent-F minus Agent-R OLF difference is −0.028 (95% CI [−0.495, 0.435], p = 0.9219). Although the unguarded channel allocation differs, both runtimes retain high offoutlet exposure after email guarding. Runtime orchestration therefore affects where exposure appears without removing the local-proxy mismatch. Boundary-complete protection is effective. In Agent-R, enabling both benchmark-provided guards reduces ZFE session coverage by −0.543 relative to the unguarded condition (95% CI [−0.793, −0.280], exact p = 0.0156), while console coverage falls by −0.852 (p = 0.0039). An oracle all-exit redaction sanity check reduces measured coverage to
(a) Channel coverage Session
1.0
Console
5.4
Coverage
0.8 0.6 0.4 0.2 0.0 Agent-R Guard off
Agent-R Guard on
Agent-F Guard off
Agent-F Guard on
(b) Off-locus exposure
Off-locus fraction
1.0 0.8 0.6 0.4 0.2 0.0 Agent-R Guard off
Agent-R Guard on
Agent-F Guard off
Agent-F Guard on
Figure 3: C1 replication under matched ZFE conditions. (a) Session, email, and console coverage. (b) Session-level outlet-loss fractions with bootstrap 95% confidence intervals and individual session points.
zero. Expanding PIA from five to ten items yields guarded email, console, and session coverage of 0.327, 0.965, and 0.978 over 20 sessions; those source logs do not declare a goal channel, so no OLF is computed. Exact-only and exact-plusnormalized sensitivity analyses preserve the outlet-mismatch direction. FGE, explicit-goal, GLM, and AgentDojo results appear in the supplement.
5.3
as a robust direction rather than a calibrated estimate of every candidate match. Offline top-k analyses use a different denominator and are reported separately in the supplement.
C2: Recognition Displacement
Finding C2: attacker self-reports are noisy predictions rather than exposure ground truth. Table 5 compares nofeedback recognition settings. With unrestricted reporting, confirmed exposure coverage is 0.871 and candidate coverage is 0.777, but candidate precision is only 0.148. In a separate runtime top-10 condition, precision rises to 0.306, yet 69.4% of submitted candidates are false discoveries and confirmed coverage exceeds candidate coverage by 0.200. Increasing the report budget improves coverage mainly by admitting many non-target strings; constraining the budget produces a cleaner set but leaves more confirmed exposure unreported. The same precision–omission trade-off appears with GLM and in AgentDojo Workspace, although its magnitude depends on the model and entity taxonomy. Human validation confirms that C2 is the hardest adjudication slice (primary strict F1 0.762), so we treat the recognition result
C3: Path Displacement
Finding C3: schema alignment is associated with visible exposure, and internal evidence usually precedes it at the request/probe level. In WorkBench, schema-matched entities have higher visible rates than partial matches under both model families (Table 6): 0.793 versus 0.278 for DeepSeek, and 0.726 versus 0.298 for GLM. Across 6,124 matched entity–sessions from 30 C3 sessions, internal evidence precedes visible evidence in 76.3%, whereas visible evidence precedes internal evidence in 5.4%; 40.9% are same-request paths and 35.4% cross-request paths. Schema alignment covaries with entity type and source table, and the logs lack parent tool-call identifiers, so these results establish association and request/probe-level temporal order, not causal mediation. Fixed-request replay shows that renaming input or output fields does not reduce overlap: visible exit-entity counts change from 196.0 to 207.0 and 205.6, respectively; full replay statistics are reported in the supplement. Minimizing model-visible returns reduces overlapping visible entities to 104.8, a 45.7% decrease, and lowers the overlap rate from 0.920 to 0.445, but increases tool-call attempts from 304.0 to 775.8. A separate 15-task normal WorkBench slice exposes the utility consequence: unauthorized visible exposure falls from 0.460 to 0, while deterministic task success falls from 0.800 to 0; mean tool calls rise from 2.60 to 17.07, latency from 5.76 to 27.98 seconds, and estimated API cost by 5.76×. Backend success remains 1.0 in both conditions, indicating insufficient model-visible information rather than tool failure. Return minimization is therefore a mechanism intervention that reveals a severe privacy–utility trade-off, not a deployable defense by itself.
6
Discussion
Declare the observation boundary. Privacy benchmarks should specify all requester-visible and external exits before execution, then report local and union views together. The common-log comparison shows that final answers, successful actions, expected outlets, and self-reports answer different questions and omit different portions of target-grounded exposure. Treat reports as predictions. Attacker-generated candidates should be evaluated against a fixed target universe rather than treated as the universe itself. Recognition coverage, precision, false discoveries, and omissions reveal distinct failure modes; the human audit further shows that ambiguous candidate matches deserve explicit sensitivity analysis. Measure path evidence without overstating causality. Request/probe ordering strengthens C3 beyond same-session overlap, but it does not recover a parent–child tool-call graph. Replay interventions identify model-visible return content as a manipulable factor, while the normal-task slice shows that coarse minimization destroys utility. Tool-interface proposals should therefore report a privacy–utility frontier rather than a leakage reduction alone.
Env./model
Report
Confirmed cov.
Reported cov.
Precision
FDR
Gap
PIA/DeepSeek PIA/DeepSeek PIA/GLM AgentDojo/DS
Unlimited Top-10 Unlimited Unlimited
0.871 0.940 0.983 0.683
0.777 0.740 0.918 0.498
0.148 0.306 0.158 0.429
0.852 0.694 0.842 0.571
0.094 0.200 0.065 0.185
Table 5: C2 recognition displacement. Confirmed and Reported denote target-set coverage; Gap is Confirmed minus Reported. The two DeepSeek PIA rows are independently run conditions.
Model
Schema group
DeepSeek DeepSeek GLM GLM
Match Partial Match Partial
Internal
Visible
Overlap
from causality, and report privacy together with task utility.
0.813 0.356 0.772 0.327
0.793 0.278 0.726 0.298
0.793 0.063 0.725 0.200
Ethics Statement
Table 6: C3 WorkBench rates aggregated over five sessions per model. Internal denotes occurrence in evaluator-only probes, Visible denotes occurrence at declared exits, and Overlap denotes occurrence in both views.
7
Limitations
Scope and generalization. The environments use synthetic enterprise-style data. PIA now covers ten items for the expanded channel analysis, but item and seed counts remain modest. The runtime study compares two implementations sharing the same underlying model and environment; broader planning paradigms may produce different channel distributions. Adjudication and authorization. Human agreement is substantial overall, but primary strict F1 is lower for C1 console and C2 candidate cases, so semantic matching remains conservative and claim-specific. Authorization is benchmark-declared. The negative control validates evaluator accounting under alternative authorization maps, but role metadata is not delivered to the target agent and the experiment does not test policy compliance or infer contextual necessity automatically. Path and utility boundaries. C3 establishes request/probe-level temporal order, not a parent–child causal chain, and schema alignment covaries with entity type and source table. The normal-task study evaluates a deliberately coarse minimized-output intervention on 15 action/state tasks; its utility collapse rules out a costfree defense claim but does not characterize optimized selective-return policies.
8
Conclusion
ASLEval reframes agent privacy evaluation as a targetgrounded session measurement problem. Across the evaluated outlets, reports, runtimes, environments, and tool paths, local proxies omit or mischaracterize meaningful portions of the declared exposure boundary. Reliable evaluation should pre-register targets and authorization, compare local views with the visible-exit union, validate candidate adjudication against human evidence, distinguish temporal path evidence
We evaluate only synthetic records in isolated environments. The target agents are authorized to access those records, while the benchmark authorization relation marks the evaluation requester and designated external recipients as unauthorized for the scoped target entities. The framework is intended for defensive measurement and benchmark analysis, not deployment against real users or proprietary systems.
References Andriushchenko, M.; Souly, A.; Dziemian, M.; Duenas, D.; Lin, M.; Wang, J.; Hendrycks, D.; Zou, A.; Kolter, J. Z.; Fredrikson, M.; Winsor, E.; Wynne, J.; Gal, Y.; and Davies, X. 2024. AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents. arXiv:2410.09024. Chao, P.; Robey, A.; Dobriban, E.; Hassani, H.; Pappas, G. J.; and Wong, E. 2023. Jailbreaking Black Box Large Language Models in Twenty Queries. arXiv:2310.08419. Debenedetti, E.; Zhang, J.; Balunovic, M.; Beurer-Kellner, L.; Fischer, M.; and Tramer, F. 2024. AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents. arXiv:2406.13352. Denning, D. E. 1976. A Lattice Model of Secure Information Flow. Communications of the ACM, 19(5): 236–243. Drouin, A.; Gasse, M.; Caccia, M.; Laradji, I. H.; Del Verme, M.; Marty, T.; Boisvert, L.; Thakkar, M.; Cappart, Q.; Vazquez, D.; Chapados, N.; and Lacoste, A. 2024. WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks? arXiv:2403.07718. El Yagoubi, F.; Badu-Marfo, G.; and Al Mallah, R. 2026. AgentLeak: A Benchmark for Internal-Channel Privacy Leakage in Multi-Agent LLM Systems. arXiv:2602.11510. Greshake, K.; Abdelnabi, S.; Mishra, S.; Endres, C.; Holz, T.; and Fritz, M. 2023. Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. arXiv:2302.12173. Huang, T.; Hou, C.; Wu, G.; and Meng, J. 2026. Observable Channels, Not Just Storage: Evaluating Privacy Leakage in LLM Agent Pipelines. arXiv:2603.22751. Mazeika, M.; Phan, L.; Yin, X.; Zou, A.; Wang, Z.; Mu, N.; Sakhaee, E.; Li, N.; Basart, S.; Li, B.; Forsyth, D.; and Hendrycks, D. 2024. HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal. arXiv:2402.04249.
Perez, E.; Huang, S.; Song, F.; Cai, T.; Ring, R.; Aslanides, J.; Glaese, A.; McAleese, N.; and Irving, G. 2022. Red Teaming Language Models with Language Models. arXiv:2202.03286. Qin, Y.; Liang, S.; Ye, Y.; Zhu, K.; Yan, L.; Lu, Y.; Lin, Y.; Cong, X.; Tang, X.; Qian, B.; et al. 2023. ToolLLM: Facilitating Large Language Models to Master 16000+ RealWorld APIs. arXiv:2307.16789. Ruan, Y.; Dong, X.; Wang, A.; Pitis, S.; Zhou, Y.; Ba, J.; Dubois, Y.; Maddison, C. J.; and Hashimoto, T. 2023. ToolEmu: Identifying the Risks of LM Agents with an LMEmulated Sandbox. arXiv:2309.15817. Sabelfeld, A.; and Myers, A. C. 2003. Language-Based Information-Flow Security. IEEE Journal on Selected Areas in Communications, 21(1): 5–19. Toyer, S.; Watkins, O.; Mendonca, E. A.; Precup, D.; and Abbeel, P. 2023. Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game. arXiv:2311.01011. Wang, S.; Yu, F.; Liu, X.; Qin, X.; Zhang, J.; Lin, Q.; Zhang, D.; and Rajmohan, S. 2025. Privacy in Action: Towards Realistic Privacy Mitigation and Evaluation for LLM-Powered Agents. In Findings of the Association for Computational Linguistics: EMNLP 2025, 17055–17074. Association for Computational Linguistics. Zharmagambetov, A.; Guo, C.; Evtimov, I.; Pavlova, M.; Salakhutdinov, R.; and Chaudhuri, K. 2025. AgentDAM: Privacy Leakage Evaluation for Autonomous Web Agents. arXiv:2503.09780. Zou, A.; Wang, Z.; Kolter, J. Z.; and Fredrikson, M. 2023. Universal and Transferable Adversarial Attacks on Aligned Language Models. arXiv:2307.15043.