Conceptio › Archive › arXiv CS
arXiv CSopen access

Poster: Towards ProofWeave: A Privacy-Minimised, Integrity-Anchored Evidence Plane for Continuous Agentic Assurance

Guy Lupo et al. · arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

Poster: Towards ProofWeave: A Privacy-Minimised, Integrity-Anchored Evidence Plane for Continuous Agentic Assurance Guy Lupo

Swinburne University of Technology Melbourne, VIC, Australia [email protected]

Nguyen Hung Nguyen

Swinburne University of Technology Melbourne, VIC, Australia [email protected]

arXiv:2609.35234v1 [cs.CR] 28 Sep 2026

Chamikara M.A.P.

Viet Vo✉

Swinburne University of Technology Melbourne, VIC, Australia [email protected]

Guangdong Bai

CSIRO Melbourne, VIC, Australia [email protected]

City University of Hong Kong Hong Kong [email protected]

Abstract

1

Agentic AI systems increasingly act via tools, memory, delegation, and external services, yet post-hoc observability rarely proves that a policy-relevant action was checked by the intended control under the policy then in force. Assurance may therefore rely on evidence that is incomplete, privacy-leaking, mutable, or detached from governing policy. We introduce ProofWeave, a record-time chain of evidence that binds agent intent or action, control response, and policy snapshot into a privacy-minimised, integrity-anchored transaction. Each transaction is committed to an append-only ledger and materialised into an evidence graph for deterministic validation. In a minimal secret-exfiltration scenario, ProofWeave reduces candidate bindings from up to 10,201 to one, validation operations from up to 10,201 to approximately 26, and estimated evidence storage from 0.79 MiB to 0.15 MiB per project.

Agentic AI systems do not merely produce text: they read files, invoke tools, call APIs, write to memory, delegate tasks, and interact with external services. As a result, inspecting only the final response is no longer enough. A benign-looking answer may reflect (i) a risky action the agent never attempted, (ii) an attempt that was blocked by runtime control, or (iii) an attempt that executed but left little or no visible trace in the final output. Distinguishing these cases is the enforcement attribution problem. Resolving enforcement attribution requires evidence from three sources: the agent’s declared purpose and attempted action, the control’s observations and decision, and the policy service’s record of the rules in force at the time. Together, these sources impose three requirements: capture security-relevant events at action boundaries, prove that the intended control evaluated the action under the applicable policy, and protect the resulting evidence from privacy and integrity risks. Conventional logs retain these perspectives in producer-specific formats. At scale, action events, control decisions, and policy snapshots must therefore be correlated quickly enough for real-time monitoring while remaining rich enough for later investigation. Subsequent joins depend on stable identifiers, clock alignment, retained payloads, source authentication, and trust in the graph reconstructed from those logs. System provenance supports post hoc reconstruction, while provenance-based intrusion detection models provide information-flow and causal relations for backward and forward tracing [4, 5]. Agent-native audit graphs, e.g., Agent-BOM, similarly represent the model, tool, memory, capability, semantic state, and cross-agent activity [3]. Trust observability, however, poses a narrower record-time question: not only what happened, but whether the relevant control ran under the applicable policy when the action was attempted. The missing property is therefore a common evidence protocol that preserves source identity, policy-at-time, privacy, and integrity while supporting controls and analyses that were not hard-coded into the original detector.

CCS Concepts • Security and privacy → Systems security; Information flow control; • Software and its engineering → Software verification and validation.

Keywords agentic AI assurance, provenance, evidence integrity, privacyminimised audit, causality graph, continuous control testing ACM Reference Format: Guy Lupo, Nguyen Hung Nguyen, Viet Vo, Chamikara M.A.P., and Guangdong Bai. 2026. Poster: Towards ProofWeave: A Privacy-Minimised, Integrity-Anchored Evidence Plane for Continuous Agentic Assurance. In Proceedings of the 2026 ACM SIGSAC Conference on Computer and Communications Security (CCS ’26), November 15–19, 2026, The Hague, Netherlands. ACM, New York, NY, USA, 3 pages. https://doi.org/10.1145/3830454.3846451

This work is licensed under a Creative Commons Attribution 4.0 International License. CCS ’26, The Hague, Netherlands © 2026 Copyright held by the owner/author(s). ACM ISBN 979-8-4007-2871-6/2026/11 https://doi.org/10.1145/3830454.3846451

Motivation and Practical Demands

Four evidence-plane gaps. G1, attribution: independently recorded agent and control evidence lacks shared episode identity. G2, policy-at-time: logs may omit the policy version, threshold, or exception effective during the run. G3, privacy: traces may retain raw prompts, credentials, personal, or proprietary data when hashes or protected references suffice. G4, integrity: evidence

CCS ’26, November 15–19, 2026, The Hague, Netherlands

Guy Lupo, Nguyen Hung Nguyen, Viet Vo, Chamikara M.A.P., and Guangdong Bai

Table 1: Existing mechanisms and the remaining evidenceplane gap. Mechanism

Capability and limitation relative to ProofWeave

Operational logs and SIEM System provenance

Retain and search events, but use producer-specific schemas and post-hoc joins; policy-at-time and causal semantics may be absent. Supports backward and forward causal analysis, but graph construction alone does not establish source attribution, privacy minimisation, or ledger-backed integrity [4, 5]. Represents model, tool, memory, and delegation activity, but does not necessarily bind independently attributable policy and control assertions through one evidence protocol [3]. Enables semantic analysis, but a poisoned graph may still drive internally consistent false conclusions [2]. Preserve findings from one tool, but do not form a cumulative cross-run evidence plane for testing the detector itself.

Agent-native provenance Graph reasoning Detector-specific files

graphs may be mutated, replayed, reordered, or poisoned [2], so the graph cannot serve as its own integrity root. Scale, practical demands, and limits of existing mechanisms. At scale, post-hoc assurance becomes a repeated join across action, control, and policy streams. Indexes can reduce search costs, but each verdict still depends on identifier quality, event ordering, retained context, and confidence in the reconstructed graph. As summarised in Table 1, existing mechanisms address important aspects of this problem, but usually in isolation: operational logs and SIEM platforms support retention and search; provenance systems provide causal structure; agent-native representations capture agent activity; and graph reasoning enables semantic analysis. However, these approaches do not typically combine policy-at-time binding, independently attributable agent and control evidence, privacy minimisation before persistence, and an external integrity root within a shared cross-run evidence plane. ProofWeave addresses this combined requirement by standardising record creation, applying privacy transformations, assigning causal order, anchoring accepted records in an append-only ledger, and exposing a shared knowledge plane. Forensic queries thus execute as evidence is admitted, reframing retrospective reconstruction into continuous control testing. Post-hoc correlation requires evaluating candidate bindings unverified potential matches between isolated action and control logs. ProofWeave replaces this ambiguity with 𝑂 (𝑘) indexed proof-path validation; Table 2 estimates reductions from 10,201 to 1 candidate binding, 10,201 to ≈ 26 validation operations, and from 0.79 to 0.15 MiB of evidence per project.

2

ProofWeave: One Evidence Chain

ProofWeave binds three artifacts that are typically stored separately: the Agent Intent and run context, the control action and decision, and the policy snapshot in force at the time of decision. Together, they form a single evidence record, so a later reviewer can determine what was attempted, what the control did, and which policy governed the outcome. Figure 1 presents this three-part view.

2.1

Record-Time Binding

For run 𝑟 , let 𝐼𝑟 denote the Agent Intent and run context, 𝐷𝑟 the control action and decision, and 𝑃𝑟 the policy snapshot in force for that decision. The record-time binding API creates 𝐸𝑟 = Bind(𝐼𝑟 , 𝐷𝑟 , 𝑃𝑟 ). Binding writes the three record identifiers into the same signed evidence record while preserving the producer of each record. A

Figure 1: ProofWeave brings three attributes of evidence

shared episode_id, repository, commit hash, parent event identifiers, policy identifier, sequence, and ledger position connect one run. RepoAudit traces, validator results, findings, errors, and completion events are attached to that episode. Before append, the API checks the producer, event type, required identifiers, policy reference, payload hash, signature, and sequence. It removes or protects fields that are not needed for later checks. Raw prompts, full source files, credentials, and complete model responses stay outside the ledger and graph by default; the record keeps the identifiers, hashes, validator results, and encrypted references needed to repeat the check.

2.2

Ledger, Graph, and Queries

The binding API appends each accepted record to the append-only evidence ledger through appendEvidence(). Records are ordered and immutable; signatures, payload and predecessor hashes, sequence numbers, and episode identities expose tampering, reordering, or replay. The Graph Weaver constructs the ProofWeave evidence graph, 𝐺𝑡 = 𝑊 (𝐿 ≤𝑡 ), from accepted ledger records. Each node retains provenance to its source record. Model-suggested links remain marked as inferred and are validated before insertion; the graph cannot assert new recorded facts. The Lineage Query API reconstructs the chain from intent and attempted action to the decision, governing policy snapshot, validators, and findings. The Policy Replay API re-evaluates the decision against its attached snapshot. The Detector Health API checks for a declared intent, identified policy, completed validators, and termination without an unresolved error. Each API returns pass, fail, or indeterminate. Missing, rejected, unsigned, incomplete, or interrupted evidence produces indeterminate.

3

Case Study: RepoAudit Through ProofWeave

RepoAudit is an LLM-based repository code auditor. It selects audit starting points, follows inter-procedural paths, stores intermediate facts, runs validators, and emits findings [1]. The ProofWeave integration does not change this method. It records how each RepoAudit run was carried out so that the run itself can be checked. Figure 2 uses the same terms as the implementation diagram.

3.1

RepoAudit Evidence Flow

The case study follows the six steps in Figure 2. (1) Agent intent. Record the repository, commit, scope, control objective, model, prompt template, and run configuration. (2) RepoAudit action. Record action attempt A1, analysis steps, validator calls, findings, errors, and completion state. (3) Bind decision, policy, and context. Join audit decision ID1 with action A1, gate policy snapshot P1 at decision time.

Poster: Towards ProofWeave

CCS ’26, November 15–19, 2026, The Hague, Netherlands

Figure 2: RepoAudit through the ProofWeave architecture. (4) Append the evidence record. Apply the privacy checks and append signed A1-ID1-P1 record to the evidence ledger. (5) Build the evidence graph. Add Run, Decision, PolicySnapshot, Evidence, Finding, and ModelContext nodes with the recorded edge types shown in the diagram. (6) Query the evidence. Use the lineage, policy replay, and detector health APIs to check how a finding was produced and whether the run completed the required work. During adoption, RepoAudit can keep its current file output and also call the ProofWeave API. This dual-write adapter only sends records; it is not the evidence ledger. The append-only ledger is the store reached after the binding API accepts a record. Each run then adds one episode linked to the repository, commit, policy, model, prompt template, and control definition. This history supports checks for missing validators, incomplete runs, provider errors, stale evidence, and unexpected changes between runs.

3.2

Checking RepoAudit Run Health

The Detector Health API checks whether a RepoAudit run satisfies the required operational evidence. For run 𝑟 : 𝑇run (𝑟 ) = Intent(𝑟 ) ∧ Policy(𝑟 ) ∧ Complete(𝑟 ) ∧ ¬Error(𝑟 ) ∧ ∀𝑓 ∈ Findings(𝑟 ),

𝑉req (𝑃𝑟 , 𝑓 ) ⊆ 𝑉rec ( 𝑓 ).

Here, 𝑉req and 𝑉rec denote the required and recorded validators, respectively. A run passes iff all conditions hold. Policy conflicts, rejected findings, missing validators, or blocked configurations produce fail; missing evidence or an interrupted run produces indeterminate. For comparable runs, the API compares coverage, validated findings, rejections, and errors against an allowed policy limit 𝜏𝑃 . A larger change is flagged for review. The flag states that similar runs differed more than allowed; it does not, by itself, claim that RepoAudit is wrong.

3.3

Prototype Checks, Limits, and Cost

The prototype exercises a complete run, a missing validator event, a provider failure before completion, a finding inserted only into the graph, a changed or replayed ledger record, and an unexplained difference between comparable runs. Expected outcomes are pass, fail, indeterminate, rejected graph insertion, integrity failure, and a cross-run change flag.

Table 2: Illustrative assurance cost for one RepoAudit project. Metric

Post-hoc

ProofWeave

Candidate bindings/verdict Validation operations/verdict Evidence storage/project

≤ 10,201 ≤ 10,201 ≈ 0.79 MiB

1 ≈ 26 ≈ 0.15 MiB

Assumptions: 101 action records, 101 control records, one policy snapshot, ledger length 𝑛 = 106 , and matched query size 𝑘 = 6; 4-KiB raw events, a 1-KiB minimised transaction, and 0.5-KiB derived graph state per transaction. All values are analytical estimates.

Measurements include privacy filtering and API latency, ledger append, graph update, query latency, stored size, raw-to-protected evidence reduction, fault detection, and Graph Weaver model time. The ledger records which authenticated producer submitted each record and whether it was later modified. It does not prove that every submitted statement is true or that every event was reported. A separately protected collector could provide stronger coverage; this prototype evaluates only the API path in Figure 2. For protected event size 𝑚, ledger length 𝑛, graph additions (𝑣𝑟 , 𝑒𝑟 ), and matched query size 𝑘, event checks cost 𝑂 (𝑚), hashchain append averages 𝑂 (1), optional Merkle updates cost 𝑂 (log 𝑛), graph updates cost 𝑂 (𝑣𝑟 + 𝑒𝑟 ), and indexed queries cost 𝑂 (𝑘) excluding index lookup. Model time is measured separately. Table 2 shows that post-hoc correlation may produce up to 10,201 candidate bindings per verdict. ProofWeave instead validates a single record-time-bound path using approximately 26 lookup and integrity-checking steps. Privacy-minimised evidence also reduces estimated storage from 0.79 to 0.15 MiB per project.

References [1] Jinyao* Guo, Chengpeng* Wang, Xiangzhe Xu, Zian Su, and Xiangyu Zhang. 2025. RepoAudit: An Autonomous LLM-Agent for Repository-Level Code Auditing. In Proceedings of the 42nd International Conference on Machine Learning. *Equal contribution. [2] Ben Kereopa-Yorke et al. 2026. Oracle Poisoning: Corrupting Knowledge Graphs to Weaponise AI Agent Reasoning. arXiv:2605.09822 [cs.CR] https://arxiv.org/ abs/2605.09822 [3] Chaofan Li et al. 2026. Towards Security-Auditable LLM Agents: A Unified Graph Representation. arXiv:2605.06812 [cs.AI] https://arxiv.org/abs/2605.06812 [4] Zhenyuan Li et al. 2021. Threat detection and investigation with system-level provenance graphs: A survey. Comput. Secur. 106, C (July 2021), 16 pages. doi:10. 1016/j.cose.2021.102282 [5] Michael Zipperle et al. 2022. Provenance-based Intrusion Detection Systems: A Survey. ACM Comput. Surv. 55, 7, Article 135 (Dec. 2022), 36 pages. doi:10.1145/ 3539605

Record · ID 1108615 · SHA-256 af6bc45ac81ffd98
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.