An Evidence Model for Agentic Processes: Evidence Claims, Trust Assumptions, and Policy Assessment Arslan Brömme CISSP, CISM, CISA, CAISE Independent Researcher [email protected] Draft v0.9.0.6 – 8 September 2026
arXiv:2609.08481v1 [cs.CR] 8 Sep 2026
Preprint / working paper. This version is a work in progress and may be updated. Comments are welcome.
Abstract
its human operators, or the responsible legal entity may need to reconstruct not only what was stored, but also which evidentiary claims can legitimately be made about the stored records. The author’s earlier architecture paper proposed a product- and vendor-neutral black-box architecture for agentic processes, based on the pipeline capture, canonicalize, hash, anchor, certify, and verify [1]. That paper introduced the architectural pattern and anchoring pipeline. The present paper substantially extends it by specifying the evidentiary semantics and limits of the resulting records, including capture-completeness, decision-session, audit, risk-treatment, mitigation-implementation, and management-response claims. The earlier architecture is treated as prior work and motivation, not as independent external validation. The motivating question is deliberately narrow: which evidentiary claims can an agentic black box support? The answer is not binary. A record can be integrity-verifiable but semantically false. It can be signed by an agent identity but still violate policy. It can be anchored externally and yet omit a suppressed event. For that reason, agentic evidence systems need explicit evidence claims and explicit trust assumptions. The novelty is not a new hash construction, ledger protocol, or agent runtime, but a claim-oriented decomposition of agentic evidence with explicit limitations and trust assumptions. The paper makes three contributions: an evidence claim structure for agentic process records, a mapping from mechanisms such as hashes, signatures, sequence references, historian classifications, audit-agent findings, risk and mitigation records, and anchors to supported evidentiary properties, and an agent-organization model that segregates operational execution, evidence production, audit assessment, and management response. This position and architecture paper does not present an empirical performance, security, or compliance evaluation. It proposes a conceptual model and uses a simplified governance example to illustrate evidentiary claims and their limits.
Agentic AI systems increasingly exchange messages, invoke tools, request approvals, hold structured decision sessions, and modify shared artifacts. Logs and anchors can make selected records tamper-evident, but they can also mislead if their evidentiary meaning is implicit: a hash does not establish semantic truth, a signature does not establish authorization, and an external anchor does not establish capture completeness. This paper proposes an evidence claim model for agentic processes. It distinguishes artifact integrity, temporal existence, provenance, approval evidence, declared ordering, capture claim, relevance claim, deliberation traceability, monitoring claim, anchoring authorization claim, policy assessment claim, risk treatment claim, mitigation implementation claim, and management response claim. Semantic validity is treated as a recurring limitation. The model maps these claims to mechanisms, assumptions, limitations, and threats, and situates them in an agent organization with functional CEO agent, executive, operational, evidence, and audit roles, plus a plan–do–check–act-inspired management response loop. The contribution is conceptual: it does not validate a particular implementation, prevent all failures, or automate legal compliance. It provides a vocabulary for stating which claims an agentic black box can support, which claims it cannot establish, and which controls are required around it.
1
Introduction
Autonomous and semi-autonomous AI agents increasingly operate as participants in organizational processes and may also be arranged as agent organizations with internal management, specialist, operational, evidence, and audit roles. They exchange instructions, draft recommendations, invoke tools, update artifacts, and request or receive human approvals. They may also participate in structured discussions in which specialized agents compare recommendations, risks, policy constraints, and dissenting views before an action is executed. These activities create an auditability problem: after an incident, compliance review, or management escalation, an agent organization,
1
2
Related Work
process improvements as an internal management and evidence role. In real deployments, binding decisions, legal accountability, and final risk acceptance remain with the responsible human officers or legal entity. Agentic workflows are owned by an operational executive agent, such as a CTO agent, CIO agent, or Headof-AI-Operations agent. The evidence function is placed under a separate executive agent responsible for security, risk, compliance, or AI governance. The audit function is either an internal audit-agent function or a separate assurance function with reporting access to the CEO agent and, where required, to human oversight. This separation matters because an evidence system should not be controlled only by the same agents and runtime that it is supposed to observe. If operational agents, logging, evidence classification, anchoring keys, audit assessment, and management reporting are all inside the same trust domain, the resulting evidentiary value is weaker. In smaller teams, open-source projects, or decentralized multi-organization settings, the same separation can be approximated through independent maintainers, separate keys, external audit services, read-only observers, or technically isolated evidence components rather than formal agent-executive roles. The following roles are used throughout the paper.
The model builds on long-standing work on digital timestamping, secure audit logs, canonicalization, transparency logs, remote attestation, and threshold signatures. Haber and Stornetta introduced cryptographic time-stamping for digital documents [2]. RFC 3161 defines a time-stamp protocol that can provide evidence that a datum existed before a particular time [4]. Secure audit logging addresses tamper evidence and post-compromise reconstruction for logs on untrusted machines [3]. JSON canonicalization is relevant because evidence records must be represented in a repeatable form before hashing or signing. The JSON Canonicalization Scheme is an informational example of such a method, not the only possible approach [5]. Transparency-log systems such as Certificate Transparency demonstrate externally auditable append-only logs and Merkle-tree inclusion evidence in a certificatespecific context [6]. Remote attestation provides a vocabulary for evidence, claims, verifiers, and relying parties about the state of a computing environment [7]. Threshold signatures such as FROST can distribute trust for sensitive signing operations, but their value depends on key generation, participant independence, threshold assumptions, and operational governance [8]. NIST policy states that federal agencies may use SHA-2 family hash functions, including SHA-256 and SHA-512, for applications that employ secure hash algorithms [10]. These mechanisms support different assurance claims. None of them, by itself, establishes semantic truth, capture completeness, or organizational compliance. The threat model is also related to work on indirect prompt injection, where malicious instructions embedded in external content can influence LLM-integrated applications and agent behavior [9]. AI governance and cybersecurity regulation provide context for risk management, record-keeping, human oversight, incident handling, and management accountability. For high-risk AI systems within the relevant scope, the EU AI Act contains provisions on risk management (Art. 9), record-keeping (Art. 12), human oversight (Art. 14), and deployer obligations (Art. 26), but these provisions do not apply to every agentic system. For entities within its scope, NIS2 covers management governance (Art. 20), cybersecurity risk-management measures (Art. 21), and reporting obligations (Art. 23). NIS2 does not establish a general evidentiary model for AI agents. Applicability depends on system classification, intended purpose, operator role, entity, sector, service, transitional rules, and national transposition [11, 12].
• Operational agents execute tasks, communicate with users or other agents, invoke tools, and update artifacts. Each operational agent has evidence duties, for example to emit events for approvals, privileged tool calls, and policy exceptions. • Agentic decision session is a structured deliberation among specialized agents and, where required, human approvers. It produces an agentic decision record containing topic, participants, policy context, arguments, dissenting views, uncertainty, and outcome. • Logger captures the event stream and stores raw or normalized records in an append-only local evidence store. • Historian classifies selected events as decisionrelevant, policy-relevant, incident-relevant, or auditrelevant. The historian is not a truth oracle. Instead, it signs relevance claims. Its classification output is itself recorded as an evidence record. • Evidence supervisor monitors whether evidence capture policies are applied, checks for missing expected events, and escalates capture gaps. • Evidence controller canonicalizes records, checks signatures and policy metadata, computes hashes or Merkle roots, and approves anchoring.
• Anchoring service publishes commitments to an external evidence layer, such as a public blockchain, timestamping service, transparency log, or permisThe model assumes a deliberately simple agent organizasioned ledger. The anchoring wallet is a governed techtion: a CEO agent at the top, executive agents beneath nical component of this service, not a free-standing the CEO agent, and employee or operational agents bedecision authority. neath them. The terms CEO agent and executive agent denote functional roles inside the illustrative agent or- • Audit-agent evaluates evidence records against security policies, control requirements, and defined workganization, not legal persons, corporate organs, or indeflow rules. pendent legal authorities. The CEO agent receives audit reports, reviews risk-treatment proposals, and initiates
3
System and Governance Model
2
• CEO agent and executive agents receive audit processed it, which policy checks were applied, and which reports, prepare management responses, propose im- actions followed. provements, accept or reject risks within delegated authority, and update policies or processes. Binding 4.3 Threats Against Evidence Infrastruclegal authority remains with the responsible human ture officer or legal entity. The organizational pattern can be summarized as fol- A third threat group targets the evidence infrastructure itself. Examples include bypassing the logger, causing the lows: historian to misclassify events, preventing the evidence CEO agent supervisor from detecting missing records, compromising |- Executive AI Operations agent the evidence controller, misusing anchoring keys, delaying | ‘- Employee / operational agents anchoring, or selectively anchoring only benign records. | and decision sessions A related concern is the evaluative component itself. |- Exec. Security / Risk / AI Governance agent The audit-agent occupies a privileged position: it inter| ‘- Logger, Historian, Evidence Supervisor, prets policies and issues findings that management relies | Evidence Controller, Anchoring Service on, yet it is itself an agent subject to misconfiguration, ‘- Internal Audit / Assurance agent compromise, or manipulation of its inputs. An audit‘- Audit-agent and reports to CEO agent agent finding is therefore a claim, not a ground truth, The exact reporting lines may vary. The central require- and its own evidentiary chain requires the same scrutiny ment is that evidence production and audit assessment as the records it evaluates. remain distinguishable from operational execution, both The model does not eliminate these threats. It requires inside the agent organization and in the surrounding them to be made explicit. Its value depends on segrehuman governance structure. gation of duties, key governance, authenticated capture, monitoring, append-only buffers, independent verification, organizational escalation, and human review of selected 4 Threat Model high-impact findings. In particular, capture completeness The threat model distinguishes three groups: internal can only be assessed relative to an expectation model. threats, external threats, and threats against the evidence Expected-event definitions, mandatory capture policies, sequence counters, tool-gateway logs, independent obinfrastructure itself. servers, and remote attestation can improve confidence, but they cannot establish that no off-system event oc4.1 Internal Threats curred. A simple operational indicator for mandatory-event Internal threats originate from deployed agents, tools, or organizational users inside the agentic system. They coverage can be expressed as a matching rate. Let M include agents that intentionally or unintentionally vio- be the number of expected mandatory events with at late policies, suppress events, manipulate communication least one validated observation, and let Nexpected be the records, bypass logging mechanisms, or execute unau- number of mandatory events expected under an indethorized tool calls. Such behavior may result from ma- pendently defined and versioned evidence-capture policy, licious objectives, misconfiguration, hallucination, over- workflow specification, or tool-gateway rule. Then optimization, defective tool integration, or inadequate M C= . human oversight. Nexpected Internal threats are especially relevant because the agent whose action is being evaluated may also be the Here, C denotes mandatory-event coverage. It is not a component asked to report the action. Therefore, self- completeness proof. It is the quantitative signal underlyreporting by operational agents is useful, but insufficient. ing the evidence supervisor’s monitoring claim in Table 1. It creates a self-reporting claim, not a completeness guar- C = 1 means that every expected mandatory record has a validated observation, while C < 1 indicates a capantee. ture gap. Multiple validated observations of the same expected event count once toward M , with consistency 4.2 External Threats assessed separately. The critical value remains Nexpected : External threats include malicious users, external agents, it must not be inferred only from the observed evidence compromised systems, adversarial documents, manipu- stream, because missing records would then disappear lated tool outputs, identity spoofing, replay attempts, from the denominator. Without such an expectation and prompt injections [9]. In these cases, a malicious model, capture coverage cannot be meaningfully comexternal input may influence an internal agent to follow puted. The ratio is not evidence of global completeness. unauthorized instructions, ignore a policy, misuse a tool, Its usefulness depends on the independence and quality of the expectation model. or rely on corrupted context. For evidence purposes, prompt injection is not only a prevention problem. It is also a reconstruction problem. 5 Evidence Claim Model In the proposed model, an organization may later need to show which external content was accessed, which in- An evidence claim is a statement about what a retained structions were embedded in that content, which agent record, artifact, process step, or audit result can support. 3
A claim is not merely a fact stored in a log. It is a struc- For the illustrative approval event ae-0042, the tured assertion whose strength depends on mechanism record would bind human_approval, human-analyst-01, and assumptions. We model an evidence claim as: remediation-agent-02, AI-GOV-07 version 1.3, a created-at timestamp, a SHA-512 artifact hash, a prior event hash, and the actor, historian, and evidenceclaim = (event, property, mechanism, assumptions, controller signatures. limitation, threat scope). record_id: ae-0042 type: human_approval actor: human-analyst-01 subject: remediation-agent-02 policy: [email protected] created_at: 2026-09-06T15:55:00Z related_records: [decision-session-009, tool-call-017] artifact_hash: sha512:<digest> previous_record_hash: sha512:<previous-digest> classifications: [decision-relevant] signatures: [actor, historian, evidence-controller] anchor_reference: batch-2026-09-06-15
The event may be an agent message, human approval, tool call, policy exception, decision session record, audit finding, risk treatment record, mitigation implementation record, or management decision. The property is the evidentiary property being asserted. The mechanism is the technical or organizational mechanism that supports the claim. The assumptions state what must be true for the claim to hold. The limitation states what the claim does not establish. The threat scope states which threat the claim is meant to address. The tuple is a conceptual notation. It is not yet a full formal semantics for claim composition, conflicting claims, confidence levels, or support states such as supported, partially supported, and not supported. For example, consider a human approval event. A canonicalized record, a SHA-512 digest, and an external anchor can support the claim that the retained approval record still matches the committed representation and that the commitment existed no later than the anchor time. They do not establish that the approval was legally valid, that the human understood all consequences, or that no other unrecorded discussion occurred. Table 1 illustrates the central design principle: the evidentiary question is not whether a system is verifiable in general, but which property is supported, by which mechanism, under which assumptions, and against which threat.
6
The created_at field represents locally asserted event time, whereas anchor_reference points to later external inclusion or finality evidence. The same envelope can represent approval records, decision-session records, audit findings, risk treatment records, mitigation implementation records, and management responses by varying type and adding type-specific payload fields. Interoperability therefore does not require one universal schema for all agentic evidence. It requires a small common envelope, stable identifiers, policy-version binding, canonicalization rules, and clearly defined extensions for different record classes. Typical extensions remain compact. A decision_session record may add participants, inputs considered, arguments, dissent, uncertainty, and outcome. An audit_finding record may add checked records, policy version, result, severity, and finding rationale. A risk_treatment record may add risk identifier, likelihood, impact, owner, treatment option, due date, and residual risk. A mitigation_implementation record may add implementation evidence, affected agents or tools, changed policy or workflow, verification status, and follow-up audit reference. Historian classifications should be captured as separate evidence records that state which event was classified, which policy or heuristic was applied, what rationale was given, when the classification occurred, and which historian identity signed it. This prevents the historian from becoming an invisible relevance gatekeeper. Negative or low-relevance classifications may not all be anchored individually, but sampling, monitoring, or batch commitments can make selective suppression harder to hide. A practical implementation must balance assurance against cost, privacy, and operability. Immediate anchoring gives stronger temporal granularity but increases cost and metadata exposure. Merkle batching reduces cost and leakage but weakens immediacy. Broader capture improves reconstruction but may create privacy, retention, and deletion constraints that require deployment-specific legal and policy analysis. More signatures and review steps improve accountability but add operational overhead. These trade-offs should be policy decisions rather than hidden implementation side effects. The audit-agent consumes evidence records and evaluates them against security policies and control objectives. Its output is not compliance itself, but a signed policyassessment claim over selected records. For example, it
Evidence Capture and Policy Assessment
A practical agentic evidence system needs an evidence capture policy. Otherwise, evidence selection becomes arbitrary and the historian becomes too powerful. The policy should define mandatory capture, risk-based capture, and historian-nominated capture. Mandatory capture should cover human approvals, rejected approvals, privileged tool calls, external system access, policy exceptions, security-relevant decisions, changes to important artifacts, incident-related communication, audit findings, risk treatment records, mitigation implementation records, and responses by the CEO agent or executive agents. Risk-based capture can apply to unusual agent-agent communication, conflicting recommendations, repeated failed tool calls, high-impact decisions, deviations from standard workflows, or structured decision sessions. Historian-nominated capture can preserve contextual discussions and rationale that are not mandatory but may later be relevant. A minimal event record should include an event identifier, event type, actor, related agent or human role, policy identifier and policy version, locally asserted event time, prior-record reference, artifact hash, relevance classifications, and signatures. 4
Mechanism
Supported evidence claim
Trust assumption
Limitation
Canonicalization + hash
Artifact integrity: retained artifact matches committed digest.
External anchor
Temporal existence: commitment existed no later than inclusion or finality time. Provenance: record was signed by a specific key or identity.
Canonicalization is deterministic and the retained artifact is available. Anchor system is independently verifiable and available.
Does not establish truth, intent, correctness, or completeness. Does not establish exact event creation time or causal order.
Key binding, authentication, and key protection are reliable.
Does not establish signer intent, honesty, or legal responsibility. Does not by itself establish wisdom, legal sufficiency, or that all consequences were understood. Does not establish real-world causality or absence of parallel events. Does not establish that all relevant events were captured. Does not establish objective importance or semantic validity. Does not create organizational authority by itself.
Agent or human signature Recorded approval
Sequence references
Logger capture Historian classification
Decision session record
Evidence supervisor
Evidence controller
Audit-agent finding
Risk treatment record
Mitigation implementation record
Management response record
Approval evidence: a specific approval or rejection decision was recorded and linked to policy context. Declared ordering: records reference previous events or a sequence. Capture claim: event passed through the logging path. Relevance claim: event was classified as policy-, decision-, audit-, or incident-relevant. Deliberation traceability: participants, rationale, dissent, and outcome were recorded. Monitoring claim: capture policy was checked against observed event stream. Anchoring authorization claim: record was validated and released for anchoring. Policy assessment claim: selected records were evaluated against policies. Risk treatment claim: risk, owner, treatment option, due date, residual risk, and management decision were documented. Mitigation implementation claim: a specific mitigation action was documented as implemented or partially implemented. Management response claim: report receipt, decision, or risk acceptance was recorded.
Approver identity, role binding, authorization scope, validity interval, delegation state, and policy version are separately verifiable. Sequence generation is controlled and gaps are detectable. Logger is correctly integrated and cannot be silently bypassed. Historian identity, policy basis, and classification time are recorded. Required participants and decision rules are defined and enforced. Expected events can be defined independently of agent self-report. Controller rules, roles, and keys are governed independently.
Unknown off-system events may remain invisible.
Policies are machine-readable, current, and correctly interpreted.
Does not establish content correctness or upstream capture completeness. Does not create compliance or establish legal sufficiency.
Risk workflow, owner assignment, and management authority are defined.
Does not establish that the risk assessment was complete or objectively correct.
Implementation evidence is captured from a reliable source and linked to the correct risk and mitigation. Executive action was captured and bound to the report.
Does not establish mitigation effectiveness or residual-risk reduction. Does not establish that the chosen measure was adequate.
Table 1: Evidence mechanisms, claims, assumptions, and limitations. may state that a privileged tool call was preceded by a recorded human approval under policy AI-GOV-07, or that a required escalation event is missing. The finding should bind the checked record set, input commitments, policy version, audit-agent model or prompt configuration, and output commitment. This allows later reviewers to distinguish conformance under the rule in force at event time from later policy changes or audit-agent configuration changes.
reject, or escalate the recommendation where human oversight is required. The session should produce a structured decision record that captures topic, participants, inputs considered, policy context, arguments, dissenting views, confidence levels, required approvals, and outcome. An agentic decision session does not create organizational authority by itself. It creates a documented decision process inside the agent organization. Legal, managerial, or policy authority remains with the responsible human role, organizational function, or legal entity. Its evidentiary value lies in preserving how the decision 7 Decision Sessions and the Man- was discussed, what warnings were raised, and which rationale was available when approval or escalation ocagement Cycle curred. Audit-agent findings should feed a risk and mitigation Agentic governance requires not only evidence of actions, but also evidence of deliberation. Before high-impact or cycle rather than remain isolated reports. Where an policy-sensitive actions, an agent organization may con- audit finding identifies a policy deviation, missing apvene agentic decision sessions in which specialized agents proval, insufficient evidence capture, unclear agent-agent discuss a decision. A planning agent may propose an communication, or exposure to prompt-injected content, action, a risk agent may evaluate impact, a security agent the agent organization should create a risk treatment may identify abuse paths, a compliance agent may assess record. Such a record links the finding to a risk descrippolicy requirements, and a human approver may accept, tion, affected policy, proposed mitigation, responsible 5
owner, due date, residual risk, and management decision. The implementation of mitigation measures should itself become part of the evidence stream. A mitigation implementation record can document whether the measure was implemented, which policy or workflow changed, which agents or tools were affected, and which follow-up verification is required. This management response loop is plan–do–check–actinspired. In the Plan phase, the agent organization defines security policies, evidence capture policies, risk treatment rules, and control objectives. In the Do phase, operational agents execute workflows and mitigation measures while evidence duties are applied. Agentic decision-session records document structured deliberation before or during operational action. In the Check phase, the evidence supervisor and audit-agent assess whether required records exist, whether policies were followed, and whether mitigation implementation can be verified. In the Act phase, the CEO agent or executive agents update policies, workflows, agent communication rules, tool permissions, or propose residual-risk acceptance for human ratification where required. Audit findings, risk treatment records, mitigation implementation records, and follow-up verification records therefore support the Check and Act phases, while decision-session records primarily support the Do phase. The resulting management decisions and policy changes become new evidence records and define the baseline for the next cycle.
Instantiating the claim tuple above for approval event ae-0042 yields the following bounded evidence claim: Event: approval event ae-0042 Property: evidence of a recorded approval Mechanism: signed approval record + sequence reference + anchored hash Assumptions: approver key valid, capture path active, role record and AI-GOV-07 v1.3 available Limitation: does not establish actual authorization, wisdom, or legal sufficiency by itself Threat scope: later denial or alteration of approval record
A linked mitigation record can instantiate the same structure: Event: mitigation record mitigation-031 Property: mitigation implementation claim Mechanism: owner-signed implementation record + linked risk treatment record + follow-up verification reference Assumptions: owner binding valid, implementation evidence captured, affected agents/tools correctly linked Limitation: does not establish mitigation effectiveness or residual-risk reduction by itself Threat scope: denial or alteration of mitigation status
The CEO agent receives the audit report and prepares a management response within the agent organization. If the audit-agent identifies that external content was processed too close to a privileged execution path, the management layer creates a risk treatment record. A mitigation may require external content to be isolated, summarized by a separate agent, and prevented from directly influencing privileged tool calls. Implementation is tracked through a mitigation implementation record. Evidence records → audit-agent finding → risk treatment A later audit-agent run verifies whether the updated → management decision → mitigation implementation workflow was applied in subsequent cases. Where legal → follow-up verification → policy improvement. authority, binding organizational change, or formal risk acceptance is required, the response must be reviewed or 8 Example: Agent Organization ratified by the responsible human officer or legal entity. The resulting evidence is useful but bounded: it can show and Decision Session that records, findings, risks, mitigations, and management responses existed in a committed form. It cannot Consider an agent organization coordinating a remediaestablish that no unobserved communication occurred, tion workflow. A planning agent proposes a remediation that the mitigation was effective in all cases, or that the step, a risk agent evaluates impact, a security agent idenorganization was legally compliant. tifies a prompt-injection risk, and a compliance agent checks whether policy AI-GOV-07 requires human approval. Where human oversight is required, a human 9 Limitations and Future Work approver accepts, rejects, or escalates the proposed action. An execution agent invokes the tool, and a verification The proposed model improves evidentiary precision, but agent checks the result. The workflow also processes an it does not solve all assurance problems. It does not guarexternal document containing adversarial instructions antee truthful agent behavior, complete capture, correct that attempt to cause a policy bypass. policy interpretation, legal compliance, effective mitigaBefore execution, the agents enter an agentic decision tions, or adequate management decisions. Events that session. The decision record preserves the proposed ac- are never captured cannot be recovered by later hashtion, external input, policy version, risk concern, prompt- ing. A compromised capture component can still record injection warning, approval requirement, dissenting views, a sanitized representation. A historian can misclassify and final recommendation. The historian classifies the relevance. A valid signature can belong to a compromised record as decision-relevant and, because adversarial con- key. tent was involved, incident-relevant. The evidence superMore fundamentally, the model does not resolve who visor checks that mandatory capture rules were triggered assesses the assessor. If the audit-agent’s findings are genfor external content, human approval, privileged tool use, erated, transmitted, and stored through the same infrasand the decision session. The evidence controller canoni- tructure it is meant to evaluate, an unaddressed regress calizes records, verifies signatures, computes hashes, and remains: a compromised or misconfigured audit-agent submits a Merkle root to the anchoring service, which pub- could produce a well-formed, signed, anchored finding lishes the commitment. The audit-agent later evaluates that nonetheless misrepresents the underlying evidence. whether human approval was recorded before execution Mitigating this requires the audit function to sit in a and whether the mitigation workflow was followed. 6
distinct trust domain from the evidence controller and anchoring service. At minimum, audit-agent output should use separately governed signing keys, bind immutable policy versions, record input and output commitments, document the model and prompt configuration used for the assessment, follow an independent capture path, and be subject to periodic human sampling against the underlying records. Future work should examine whether second-order evidence about the audit process is a proportionate control, or whether it merely relocates the same trust problem one level up. The model also depends on practical governance choices. Key management, segregation of duties, retention rules, privacy constraints, access control, monitoring quality, and the independence of the anchoring layer affect the evidentiary strength of the system. Capture completeness remains relative: it can be assessed against defined expectations and independent signals, but cannot establish the non-existence of every unobserved communication or side channel. Blockchain anchoring is one possible externalization mechanism, but timestamping services, transparency logs, and permissioned ledgers may also be appropriate depending on threat model and regulatory environment [4, 6]. Future work should formalize an evidence claim language, including claim composition, conflict handling, confidence levels, and support states. It should define interoperable event schemas, evaluate canonicalization and Merkle batching strategies, test capture completeness under adversarial workflows, integrate remote attestation for evidence controllers, and study how audit-agent findings can be reviewed by human and external auditors. Schema registries, evidence profiles, and conformance tests for canonicalization and claim validation should also be investigated. Another open question is how external auditors can independently verify selected evidence without exposing sensitive agent communications, risk records, mitigation details, or business data. Future work should also evaluate the model in realistic multi-agent deployment environments, including chatoperated gateways, tool-executing agent runtimes, and managed multi-agent settings. Experiments should instrument communication, tool invocation, approvals, auditagent findings, and management responses, test whether expectation models for mandatory events can be defined, assess whether the matching-based coverage indicator C produces useful audit signals, and examine whether adversarial workflows, missing approvals, or privileged tool misuse can be reconstructed. The objective is model usability under realistic communication, execution, and governance conditions, not benchmarking a specific product.
10
traceability, monitoring claim, anchoring authorization claim, policy assessment claim, risk treatment claim, mitigation implementation claim, and management response claim. Semantic validity remains a recurring limitation rather than something established by anchoring or signing alone. The model’s main conclusion is simple: the relevant question is not whether agentic evidence is verifiable in general, but which evidentiary property is supported, by which mechanism, under which trust assumptions, and against which threat. A well-designed agentic black box should therefore not only preserve records. It should make the evidentiary meaning of those records explicit and support a closed management cycle from evidence capture and audit findings to risk treatment, mitigation implementation, follow-up verification, and policy improvement.
References [1] A. Brömme, “A Black Box for Agentic Processes: BlockchainAnchored Evidence for AI Agent Communication, Human Oversight, and GRC Audits,” arXiv:2609.04017 [cs.CR], 2026. https://arxiv.org/abs/2609.04017 [2] S. Haber and W. S. Stornetta, “How to time-stamp a digital document,” Journal of Cryptology, vol. 3, no. 2, pp. 99–111, 1991. doi:10.1007/BF00196791. https : / / doi . org / 10 . 1007 / BF00196791 [3] B. Schneier and J. Kelsey, “Secure audit logs to support computer forensics,” ACM Transactions on Information and System Security, vol. 2, no. 2, pp. 159–176, 1999. doi:10.1145/317087.317089. https://doi.org/10.1145/317087. 317089 [4] C. Adams, P. Cain, D. Pinkas, and R. Zuccherato, “Internet X.509 Public Key Infrastructure Time-Stamp Protocol (TSP),” RFC 3161, IETF, Aug. 2001, updated by RFC 5816. doi:10.17487/RFC3161. https://www.rfc-editor.org/info/rfc3161 [5] A. Rundgren, B. Jordan, and S. Erdtman, “JSON Canonicalization Scheme (JCS),” RFC 8785, Informational, June 2020. doi:10.17487/RFC8785. https://www.rfc-editor.org/info/rfc8785 [6] B. Laurie, E. Messeri, and R. Stradling, “Certificate Transparency Version 2.0,” RFC 9162, Experimental, Dec. 2021. doi:10.17487/RFC9162. https://www.rfc-editor.org/info/rfc9162 [7] H. Birkholz, D. Thaler, M. Richardson, N. Smith, and W. Pan, “Remote ATtestation procedureS (RATS) Architecture,” RFC 9334, Informational, Jan. 2023. doi:10.17487/RFC9334. https: //www.rfc-editor.org/info/rfc9334 [8] D. Connolly, C. Komlo, I. Goldberg, and C. A. Wood, “The Flexible Round-Optimized Schnorr Threshold (FROST) Protocol for Two-Round Schnorr Signatures,” RFC 9591, IRTF/CFRG, Informational, June 2024. doi:10.17487/RFC9591. https://www. rfc-editor.org/info/rfc9591 [9] K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz, “Not what you’ve signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection,” arXiv:2302.12173 [cs.CR], v2, 2023. doi:10.48550/arXiv.2302.12173. https://arxiv.org/abs/2302. 12173 [10] National Institute of Standards and Technology, “NIST Policy on Hash Functions,” NIST Computer Security Resource Center, policy update dated Dec. 15, 2022. Webpage updated Sept. 9, 2024. Accessed: Sept. 7, 2026. https://csrc.nist.gov/projects/ hash-functions/nist-policy-on-hash-functions [11] European Union, “Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence,” Official Journal of the European Union, 2024. https://eur-lex.europa.eu/eli/reg/ 2024/1689/oj [12] European Union, “Directive (EU) 2022/2555 on measures for a high common level of cybersecurity across the Union (NIS2 Directive),” Official Journal of the European Union, 2022. https://eur-lex.europa.eu/eli/dir/2022/2555/oj
Conclusion
Agent organizations need more than logs. They need a precise understanding of which claims their evidence can support. This paper proposed a compact evidence model for agentic processes that separates artifact integrity, temporal existence, provenance, approval evidence, declared ordering, capture claim, relevance claim, deliberation
Declaration on the Use of AI Tools. AI language models, in particular OpenAI GPT-5.6 and Anthropic Claude Sonnet 5, were used as tools during the preparation of this paper for language drafting, critical review, source checking, discussion of examples, and LaTeX/PDF artifact generation. The author remains solely responsible for the content, conceptual decisions, source selection, and final version.
7