Conceptio › Archive › arXiv CS
arXiv CSopen access

Observational Indistinguishability and Integrity Blind Regions in Hybrid Quantum-Classical Workflows

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

1

Observational Indistinguishability and Integrity Blind Regions in Hybrid Quantum-Classical Workflows Roberto Fernández-Barrios, Iker Pastor-López, Amaia Pikatza-Huerga, and Pablo Garcı́a Bringas

arXiv:2609.17150v1 [cs.CR] 15 Sep 2026

✦

Abstract—We present a claim-relative evidence/reference framework for hybrid quantum-classical workflow integrity. Observational indistinguishability yields structural blind regions, distinct from finite-batch statistical misses. Within the declared lattice, a trusted same-batch scalar R0 suffices for conclusion integrity, aggregate M0 for aggregate plus conclusion integrity, and item-aligned binding for item identity. In 3,600 label interventions, feature/prediction views realize exact label-path invariance; all 764 geometry-aligned aggregate-blind rows equal their paired-clean responses, giving zero attack-only increment. For statistical response, the geometry-aligned construction detects 343/2,700 conclusion-changing (τ → 0+ ) label interventions with the conformal rule and 1,183/2,700 with the uncorrected union; the original frozen same-item geometry yields 11/2,617 and 43/2,617, respectively. The executed conformal clean falseaction rates are 0.048–0.059 descriptively; its finite-sample guarantee requires exchangeability, which the overlapping-draw design violates. The cluster-preserving adaptive stress test (Gate A) reduces response versus matched controls in 25–40 of 40 environment/split cells while retaining conclusion changes. A bounded 165-design-cell ideal-statevector and finite-shot-emulation branch directly instantiates semantic, estimated and observed kernel transitions. The fixed equal-weight design estimates neither deployment prevalence nor QPU, provider or deployed-service assurance. Index Terms—hybrid quantum-classical systems, integrity auditing, observability, quantum kernels, workflow integrity

1

I NTRODUCTION

H

YBRID quantum-classical software distributes one reported result across heterogeneous computational and evidential boundaries. Classical records are acquired, cleaned, projected and encoded; a circuit and execution stack produce quantum evidence; a learner turns a kernel into a decision; and an evaluation layer aligns that decision with reference outcomes and reports a metric. A plausible final number may therefore coexist with a failure in the data, circuit, kernel, prediction, label or reporting path. Security analyses of hybrid systems consequently call for lifecycle-aware controls rather than trust in an isolated algorithm [1], [2].

•

•

The authors are with the Faculty of Engineering, University of Deusto, Avda. de las Universidades 24, 48007 Bilbao, Spain. E-mail: [email protected] (corresponding author), [email protected], [email protected], [email protected]. This work is part of grant PID2024-155693NB-C43, ATHENA-AEGIS (Advanced Secure Technologies for Hybrid Quantum-Classical Environments and Applications), funded by MICIU/AEI/10.13039/501100011033 and by ERDF/EU.

Robustness benchmarking collapses this chain into a performance change. That endpoint cannot identify what changed or whether the auditor could have seen it. Consider a hybrid intrusion-detection service that scores each batch with a fixed predictor, then joins the predictions to a groundtruth feed from another subsystem (analyst verdicts, incident tickets or a delayed label store). The predictor is intact, but the join or store is stale, corrupted or rewritten, so balanced accuracy moves although the model has not. Feature and prediction monitors cannot see the event because their inputs are unchanged. Label counts expose an unbalanced corruption but not a balanced exchange of class identities. A signature authenticates the label file, not its item-level alignment with the predictions, so a correctly signed stale or misaligned file still passes. Calling the event stealthy therefore requires a qualifier: stealthy to which auditor, holding which evidence and trusting which reference? The constructive result is claimrelative. A trusted scalar reference R0 for the same batch certifies whether the reported conclusion changed. A trusted aggregate M0 additionally certifies aggregate integrity and, by Proposition 3, every material conclusion change in the declared setting; only item-identity integrity requires item alignment. If one party controls both the label store and its checking reference, no local verifier can distinguish a substituted honest pair from the original (Proposition 6). The paper identifies a necessary external root and claim-relative sufficient granularity within the declared reference model; it does not remove the need for that root. The scenario is a reading aid, not an incidence estimate. Prior art establishes observation-relative detectability [3]– [5], authenticated evaluation artifacts [6], detector-aware drift [7] and quantum-stage contracts [8], [9]. We claim none of these building blocks individually. Their composition here yields a claim-relative integrity lattice with structural blind regions, an explicit evaluation-label path, reference sufficiency/necessity for three claim levels, and a separation between structural blindness and finite-sample miss. The experiments instantiate those distinctions, policy consequences, one declared adaptive fingerprint and a bounded quantum branch. The contribution is one package: •

a claim-relative workflow view/reference lattice, with

2

•

•

•

•

monotonicity, closure, materiality and minimal counterexamples, including distinct semantic, estimated and observed kernels; an explicit adversary/failure model, including an adaptive attacker that preserves the declared batch fingerprint; multi-dataset fixed-OOD validation with duplicatecontrolled result aggregation, numerical identity checks, cluster-level uncertainty and prespecified null calibration; a conformal family rule valid under exchangeability and an offline allow/hold/block analysis over the frozen grid; and a bounded quantum-path instantiation and fourcontract prototype.

The lattice and its claim-relative derivation are the contribution; Max-Rank, contracts, hash chains and cryptographic roots remain prior-art components. Results have three layers: structural propositions, finite-design statistical responses and constructive exact checks. Hardware, calibrated backend noise, provider security, deployed services and the ATHENA fleet-management case study (WP5) are outside scope (subsequent operational work of the ATHENAAEGIS subproject); the supplemental model profile is neither quantum advantage nor vulnerability ordering.

2

R ELATED W ORK AND P OSITIONING

2.1 Stealth, Detectability and Observability in Secure Systems Observation-relative detectability is established: cyberphysical attacks can be unobservable, residual-stealthy or confined to an estimator’s range space [3]–[5]; later work quantifies the detectability–impact trade-off and the residuals a monitor can form [10]–[12]. Hinder et al. construct drifts designed to evade drift detectors [7]. Our clusterpreserving adaptive stress test (Gate A, Section V-E) instantiates this against the declared batch fingerprint; it measures conclusion-changing fractions, policy consequences and evidence regimes without claiming priority for adaptive drift evasion. Simplex supplies the canonical decision/fallback pattern [13], [14], runtime verification checks executions [15], and integrity checkers compare digests with trusted baselines [16]. Our claim is not generic information-relative detectability, but blind regions and calibrated decisions for workflow views whose protected boundaries include evaluation ground truth. 2.2

Monitoring, Labels and Evaluation Integrity

Monitoring work treats shifted priors, selective labels and proxy sufficiency under delayed ground truth [17]–[20]. Benchmark-quality work treats label error, evaluation pitfalls, spatial/temporal sampling bias and measurement blindness [21]–[24]. These lines primarily model labels as shifted, delayed, noisy, erroneous or measurement-limited. VAMP [6], by contrast, explicitly treats evaluation artifacts as authenticated assets that may be poisoned or substituted. Our distinction is therefore not that evaluation data can be attacked, but that auditability and trusted-reference

granularity are derived relative to the protected claim and the auditor’s view. VAMP protects named datasets, software, models and evaluation sets through authenticated provenance, providing a concrete external root of the kind required by Proposition 6. Its focus is artifact authentication; our complementary question is which view and granularity suffice for conclusion, aggregate or item-identity integrity, and what remains indistinguishable with weaker evidence. A VAMP-style manifest can instantiate the references assumed here. In-toto, SLSA and Sigstore bind supply-chain steps, build provenance, identity and transparency [25]–[27]. Remote attestation supplies fresh evidence to an appraiser [28]; RATS names Attesters, Verifiers, Evidence, Reference Values and appraisal policy [29]. These mechanisms establish provenance or appraisal, not whether aggregate or item-aligned evidence suffices for a claim; provider/runtime authentication is subsequent operational work of the ATHENA-AEGIS subproject. 2.3

Calibration of Multi-Sensor Decisions

The family rule of Section III-D is a conformal p-value [30], [31]. We use the Max-Rank statistic of Timans et al. within a full-conformal family construction [32]: the maximum of the per-sensor ranks is equivalently a Tippett minimum-p ordering, closely related to Westfall–Young resampling corrections [33], [34]. The augmented leave-one-out construction and ties-against-rejection convention supply the stated finitesample argument under exchangeability [30], [35]. Conformal anomaly detection and shared-calibration p-values provide the surrounding statistical context [36], [37]. We claim no novelty for the statistic or rule: the family score calibrates an information-regime decision under its stated premise and exposes the executed design’s premise violation. Beyondexchangeability conformal methods exist [38], but this paper does not implement one. The asymmetric comparison rule had no such guarantee (Proposition 5(c)). 2.4

Hybrid and Quantum Software Integrity

Hybrid-system analyses decompose control, compilation, execution and measurement surfaces and cover malicious compilation, hardware faults, cloud exposure and side channels [1], [2], [39]. Quantum Vulnerability Factor scores fault sensitivity, while Quantum Leak demonstrates cloud timing leakage [40], [41]. QCIVET combines stage contracts, hash-chained traces, behavioural subtyping, calibrated observable tests and QPU validation [8]. QML-PipeGuard uses a structured observable family and calibrated drift tolerance to detect channel substitution on IBM hardware [9]; multilevel work evaluates circuit-integrity metrics [42]. Bensoussan et al. derive a taxonomy from real hybridsoftware faults [43]; our synthetic interventions are controlled distinguishability probes, not a prevalence model. Quantum Squeeziness formalizes quantum-program testability and fault masking [44]; we instead study workflow-wide observational indistinguishability across a claim-relative evidence lattice, including evaluation labels and trusted-reference granularity. QProv records provider-independent provenance across circuit, compilation, execution and hardware context [45];

3

recent cross-provider work extends this line toward unified quantum-job provenance [46]. Differential, metamorphic and equivalence-modulo-input tests address the oracle problem [47]–[49]; compiler verification, program logics and runtime assertions supply other formal/runtime mechanisms [50]– [54]. We claim no priority for these mechanisms. QCIVET and QML-PipeGuard reason directly about quantum-stage contracts and behavioural evidence, including calibrated observables, drift and hardware validation; QProv and crossprovider work address provenance. They are stronger here on QPU/hardware and runtime/provider evidence. Our orthogonal contribution is a workflow-wide, claim-relative information-set lattice that protects the evaluation-label path and distinguishes conclusion, aggregate and item-identity integrity. Its quantum evidence is simulator-only, its hash chain unsigned, and its policy offline. 2.5

Relation to Companion Work

item-wise, so only R is recomputed (its prior-preserving subclass Tyπ also preserves the class histogram); feature-side mechanisms overwrite X̃ or X and recompute ŷ and R; circuit-side interventions overwrite C and recompute Ksem and everything below it; estimation variation overwrites ξ or E with Ksem fixed; post-processing overwrites g or Kobs with Ksem and K̂ fixed. “Fixes everything else” thus means every non-descendant. The signed conclusion change is

∆R (a) = R(s0 ) − R(sa ).

(2)

An intervention is material at level τ if |∆R (a)| ≥ τ . The limit τ → 0+ is a structural-sensitivity endpoint—whether the reported conclusion changes at all—not an operational risk threshold or an assertion that every epsilon has the same consequence. The fixed τ = 0.02 and 0.05 analyses show larger effect sizes; service-level risk is outside this study. The signed endpoint, unlike the positive part max(∆R , 0), counts an apparent improvement caused by an integrity failure as a conclusion change.

Three companion studies by the authors are disclosed to delimit the claim. Sharp target-domain certificates [55] identifies where a fixed candidate advantage is supported under 3.2 Views, Refinement and Trusted References shift. Conditional validity [56] studies conditional evidence An information set I is a view map VI : S → OI ; the auditor and abstention under collider and estimation uncertainty. observes VI (s) and nothing else. The regimes studied are Candidate comparability [57] governs challenger promotion IX (the multiset of rows of X̃ ), IXF (the multiset of pairs after drift. This paper instead concerns adversarial workflow (x̃i , ŷi )), IYm (the class-count vector of y ), IXF Y (the multiset integrity: observational blindness, trusted roots, label-store of triples (x̃i , ŷi , yi )) and IQ (circuit, kernel and execution substitution and reference granularity. Shared data sources, evidence). I ⊑ I ′ (refinement) iff VI = π ◦ VI ′ for some π ; fidelity kernels and selected monitors are disclosed; the hence IX ⊑ IXF ⊑ IXF Y and IYm ⊑ IXF Y , while IYm is formal objects, interventions, endpoints, tables and pri- incomparable with IX and IQ is a separate branch. The join mary experiments are distinct. Operational QPU/context I ∨ W is the pair of views. A reference profile separates three dimensions: proveassurance belongs to subsequent operational work of the nance (historical, benchmark-protected, or deploymentATHENA-AEGIS subproject. authenticated), granularity (aggregate or item-aligned), and decision rule (statistical or exact). A reference is trusted 3 O BSERVATION M ODEL AND B LIND R EGIONS if no intervention in the class under study can alter it. The supplement gives complete proofs. The formal core The A–C labels are shorthand, not a one-dimensional trust derives which evidence separates each intervention class and scale. Class A is a statistically thresholded aggregate comparintegrity claim. ison; its reference may be historical or, as executed here, a protected clean same-item-set oracle used without item pairing. The oracle is outside the benchmark attack API, 3.1 States, Interventions and Materiality but no deployed authentication mechanism is demonstrated. A workflow state separates stored artifacts from those the Class B is an exact invariant against a trusted aggregate sameworkflow derives. The primitive artifacts of one evaluation batch reference, such as M (s ), hist(y ) or R(s ). Class C 0 0 0 run are is an exact invariant against a trusted item-aligned sames = (X, P, C, E, ξ, g, f, y) ∈ S, (1) batch reference such as (y ) . Thus statistical versus exact 0,i i acquired features X , preprocessing P , circuit or feature and aggregate versus item-aligned are not synonyms for map C , execution context E , estimation randomness ξ , post- untrusted versus trusted. For a label-path intervention, itemprocessing map g , fitted predictor f and evaluation labels identity integrity holds if ya = y0 item-wise, aggregate integrity y . The derived artifacts are recomputed from them: the if M (sa ) = M (s0 ) and conclusion integrity if R(sa )⋆ = R(s0 ); representation X̃ = P (X); the semantic kernel Ksem = Φ(C) each implies the next and no converse holds. I denotes induced by the circuit on a fixed probe set; the finite- a ⋆regime augmented with trusted same-batch references shot estimate K̂ = est(Ksem , E, ξ); the observed kernel (IXF Y holds the aggregate confusion profile and itemKobs = g(K̂) delivered downstream after estimation, trans- aligned labels and predictions; feature coverage remains mission and post-processing; the predictions ŷ = f (X̃, Kobs ), statistical). A label available in IXF Y is not thereby trusted: indexed by item identity i = 1, . . . , n; and the result if the label store is the asset under attack, the auditor sees ya , R = r(ŷ, y), balanced accuracy of ŷ against y . An inter- not y0 . vention a overwrites a declared set of artifacts, primitive or derived, keeps every non-descendant fixed and recomputes the descendants through the workflow with ξ retained; the baseline is s0 and sa = a(s0 ). Label-only Ty overwrites y

3.3

Equivalence, Sensors and Blind Regions

Definition 1. s ∼I s′ iff VI (s) = VI (s′ ). A sensor admissible in I is a measurable S : OI → Rk , written S(s) = S(VI (s)).

4

Lemma 1 (structural blindness). If sa ∼I s0 then S(sa ) = S(s0 ) for every sensor admissible in I (path-wise for a seeded randomized sensor; in distribution for a fresh seed). Definition 2. For a class A, the structural blind region of I is BI (A) = {a ∈ A : a(s0 ) ∼I s0 } and the sensor blind region of a family S is BI,S (A) = {a : S(a(s0 )) = S(s0 ) ∀S ∈ S}. By Lemma 1, BI ⊆ BI,S , with equality iff S is baseline-separating on the orbit of A (a different value at a(s0 ) than at s0 whenever the views differ); detection needs baseline separation, identification would need pairwise separation, which is not claimed. Sensor blindness thus decomposes into information-level blindness and sensor insufficiency; the adaptive attacker of Section VI-D exploits the second. The structural and sensor auditability gaps at level τ are

GI (τ ) = {a : |∆R (a)| ≥ τ, a ∈ BI },

GI,S (τ ) ⊇ GI (τ ). (3) Proposition 1 (monotonicity). If I ⊑ I ′ then BI ′ (A) ⊆ BI (A) and GI ′ (τ ) ⊆ GI (τ ) for every τ . Proof: VI ′ (sa ) = VI ′ (s0 ) implies VI (sa ) = π(VI ′ (sa )) = π(VI ′ (s0 )) = VI (s0 ). Corollary 1 (label-path boundaries). Let f be fixed and deterministic. (a) Ty ⊆ BIXF ⊆ BIX : every sensor admissible in IX or IXF is invariant under a label-only change. (b) Tyπ ⊆ BIYm : every sensor admissible in IYm is invariant under a prior-preserving change. Proposition 2 (closure). BI∨W (A) = BI (A) ∩ BW (A); a class A ⊆ BI becomes completely separable after adding W iff VW (a(s0 )) ̸= VW (s0 ) for every a ∈ A. Corollary 2 (label-path closure). No view factoring through (X̃, ŷ, hist(y)) separates any a ∈ Tyπ . The multiset IXF Y separates it unless labels move only among items with identical (x̃, ŷ); an item-aligned view separates every non-identity relabeling. Proposition 3 (materiality forces aggregate separability). Let R = g(M ) be a function of the confusion matrix M (s), as balanced accuracy is. If a ∈ Ty and ∆R (a) ̸= 0 then M (sa ) ̸= M (s0 ); hence a is separable in the confusion view and in IXF Y , and GIXF Y (τ ) ∩ Ty = ∅ for every τ > 0. Proof: Contrapositive of M (sa ) = M (s0 ) ⇒ R(sa ) = R(s0 ); the confusion view is a coarsening of the multiset of triples. The proposition holds for every metric that is a function of the binarised confusion matrix; for a ranking metric such as ROC-AUC the same argument holds with the multiset of (score, label) pairs as the aggregate (supplement). No AUC experiment is run. 3.4 Three Auditor Classes and Decision-Level Calibration A reference-anchored auditor holds a trusted same-batch reference ρ (class B or C) and computes D(s) = d(VJ (s), ρ) with a metric d; then D(s0 ) = 0 exactly, the rule “fire iff D > 0” has zero false-action probability, and separability in J is sufficient for detection. A batch-level statistical auditor computes aggregate scores and thresholds them against a clean calibration population (class A). Its scoring reference can be historical or same-item-set; either way, separability

need not yield power beyond clean-score variability. A posthoc certifier recomputes a view from inputs it trusts (the evaluation contract recomputes R from (ŷ, y)) and needs trusted inputs rather than a stored value. Proposition 4. (i) For any auditor whose decision is a function of VI (s), a ∈ BI implies fire(sa ) = fire(s0 ) pathwise under the same state and reference construction. This equality does not in general imply equality with a falseaction rate estimated under a different clean-resampling or reference construction. (ii) A reference-anchored auditor detects every a ∈ / BJ , deterministically. Corollary 3 (which reference certifies which integrity). Let a ∈ Ty with f fixed. (a) A trusted exact same-batch claim reference R0 = R(s0 ) detects exactly every R(sa ) ̸= R0 and suffices for a conclusion-only claim. A trusted M0 = M (s0 ) detects every violation of aggregate integrity and, by Proposition 3, every material conclusion change. (b) References that factor through M , hist(y) or R miss C1 relabelings; item identity needs level C, whose aligned y0 detects every nonidentity relabeling. Thus R0 , M0 and item alignment protect conclusion, aggregate-plus-conclusion and identity claims, respectively. Their sufficiency and the counterexample necessity are claim-relative within the declared reference lattice; the minimality statement is not over every encoding or audit architecture. Proposition 5 (union versus conformal family calibration). (a) If m sensors fire underSthe null with probabilities P pj ≤ α, then maxj pj ≤ Pr0 [ j Ej ] ≤ min(1, j pj ) ≤ min(1, mα), the bounds being attained by nested and disjoint events. (b) Let v1 , . . . , vn be the sensor vectors of n calibration draws and vn+1 that of the audited batch; give every member j of the augmented set the leave-one-out family score Ũj = maxs #{i ̸= j : vs,i < vs,j }/n, let p = (1 + #{i ≤ n : Ũi ≥ Ũn+1 })/(n + 1), and fire iff p ≤ α (ties against firing). At most ⌊α(n + 1)⌋ members of any augmented set would fire if audited; hence, if the n+1 draws are exchangeable, Pr0 [p ≤ α] ≤ ⌊α(n + 1)⌋/(n + 1) ≤ α with no continuity or tie assumption [30], [33], [34]; 10/201 = 0.0498 for n = 200, α = 0.05. For one sensor without ties the rule is the per-sensor null-calibration rule (Gate N, Section V-C). The level is marginal over the joint draw of calibration set and audited batch; nothing is claimed when the premise fails. (c) An asymmetric construction that scored calibration draws against the other n − 1 draws only, has no such guarantee: with three sensors and cyclic orderings over five vectors it fires with probability 0.60 at α = 0.2 under exchangeability (supplement), and on the frozen draws it exceeds α under exchangeable re-splits where the conformal rule does not (Section VI-B). Proposition 6 (no local authentication without an uncontrolled root). Let a local verifier be any map acc(VI (s), ρ) of the observed view and a reference bundle, with honest references ρH (s) = VJ (s). If it accepts every honest pair and the class can produce (VI (s′ ), ρH (s′ )) for some honest s′ ̸∼I s0 (joint control), the substitution is accepted as an honest run of s′ : no function of that input establishes authenticity relative to s0 , and if R(s′ ) ̸= R(s0 ) a material change is served. A bundle component VK (s0 ) the class cannot write rejects every substitution with VK (s′ ) ̸= VK (s0 ) (Proposition 4(ii)): authenticity requires a root outside the declared class. The result concerns authenticity, not consistency; by Proposition 3

5

a material substitution does change the joint view. 3.5

Counterexamples

Eight minimal constructions separate the events a robustness number conflates; the supplement tabulates them with the frozen expansion observations that realise each. C1 swaps the labels of two items with equal predictions: M , the histogram and R are invariant while item identity changes (808 of 3,600 label rows), an integrity violation with no conclusion impact, visible only item-wise against a trusted reference. C2–C4 separate marginal invariance, confusion-matrix invariance and conclusion invariance, C5–C6 do the same on the feature side, C7 shows that the clipped “harm” endpoint of the earlier evidence hid 1,534 conclusion changes (433 on the label path), and C8 is the empirical face of Proposition 3 (2,617 of 2,617 material label rows change M ). Of the 3,600 label rows, a trusted aggregate reference of the same batch detects 2,792 (every material one, Corollary 3a) and the remaining 808 are C1 relabelings that only an item-aligned reference exposes (Corollary 3b). A ninth construction, the cluster-preserving drift of Section VI-D, is not a structural blind region but sensor insufficiency against an adaptive attacker, measured rather than proved. 3.6

The Quantum Branch as an Instance of the Lattice

Let C be the circuit representations (canonical OpenQASM 3 with bound parameters), Φ : C → K the semantic map to the ideal fidelity kernel on a fixed probe set, Ksem = Φ(C), K̂ and Kobs the estimated and observed kernels of Section III-A, h(C) the provenance hash, A(·) the algebraic invariants and fK (X̃) the decision computed from a kernel. Three intervention classes are kept apart: (a) circuit-side, overwriting C ; (b) estimation variation, with C and Ksem fixed and K̂ changing; (c) post-processing or kernel substitution, with C , Ksem and K̂ fixed and Kobs changing. Each inclusion below names its class; none is claimed across classes. Proposition 7 (quantum lattice). Assume h collision-free on the circuits under study. (i) [class (a)] [C] = Φ−1 (Φ(C)) is the semantic class of C ; approved transpilation and common-unitary rewrites map C into [C] without fixing h(C), so provenance refines the semantic view, Bh ⊆ BKsem , strictly whenever the class contains an approved rewrite: separability in provenance is not harm, and the auditor needs the approved class, implemented as semantic equality of probe kernels against the trusted Ksem,0 = Φ(C0 ) (level B). (ii) [class (c)] A and fK coarsen the observed kernel, so BKobs ⊆ BA and BKobs ⊆ Bf ; a PSD-preserving substitution K ′ with A(K ′ ) = A(Kobs,0 ) lies in BA \ BKobs and, since h(C) and the semantic probe of the unchanged circuit see nothing, is closed only by an anchored comparison of Kobs itself: the trusted semantic probe and the trusted observedkernel reference are different anchors. A kernel change that crosses no decision boundary lies in Bf \BKobs . (iii) [class (b)] For an honest finite-shot estimate, exact equality K̂ = Ksem,0 has false-alarm probability 1 − Pr[K̂ = Ksem,0 | honest], which depends on the discrete support of the estimator and may be large or equal to one but is not universally one (fidelity 1/2 at two shots gives 1/2). Exact equality is therefore not an appropriate acceptance criterion; d(K̂, Ksem,0 ) must

be treated as a level-A statistic with a null from repeated honest estimation, to which Proposition 5(b) applies with one sensor when the repeated and audited estimates are exchangeable, and where no such null is calibrated no statistical integrity claim is made. The proposition adds no quantum mechanics: it names the two features the classical branches lack, an approved equivalence class coarser than provenance and an estimation step that turns an exact anchor into a statistical one. Hash chaining of the audit envelope is a reference-anchored comparison over the record itself; its root is not assumed uncontrolled, so Proposition 6 applies.

4

A DVERSARY AND FAILURE M ODEL

Table 1 assigns every executed intervention to one of four classes with the roots it cannot write and the least separating evaluated regime (complete per-class record in the artifact). Evaluation-label classes stand for corruption of the groundtruth store or label join; feature-side classes for corruption or drift of the acquisition path feeding both branches; the cluster-preserving class for an attacker who knows the batchlevel fingerprint; circuit and kernel classes for corruption of the circuit/parameter store, the compiler output or the postestimation kernel; shot emulation for legitimate estimation uncertainty. Equal weighting is a stress profile, not a threat distribution; no prevalence is estimated. The evidence trusts the local runtime and host, the training data and fitted model, the reference input and preprocessing declaration, the reference circuit and probe kernels, SHA-256 collision resistance, the item-level label reference ⋆ only where IXF Y is declared, and the code that builds and verifies the envelope. Compromise of the operating system, runtime or verifier; provider identity, forged job or calibration metadata and provider attestation; malicious schedulers, physical mapping and unapproved passes on hardware; device-calibrated noise, crosstalk, multi-tenant interference and QPU faults; side channels and circuit confidentiality; denial of service; and the ATHENA fleet-management case study (WP5) are outside this study (subsequent operational work of the subproject).

5

M ETHODOLOGY

5.1

Data, Pipelines and Interventions

The frozen first gate (Gate 1) uses a balanced binary subset of CICIDS2017 [58]: 3,000 rows, 77 numeric columns and 128/128 capped training/evaluation samples. The expansion adds eight fixed environments E1–E8 (supplement): five from CICIDS2017 (a 256/256 scale gate and four temporal source-target pairs), two from UNSW-NB15 [59] (a balanced ID subset and a temporal pair) and one from ToN-IoT [60] (a balanced ID subset), all at 128/128 unless stated and projected to dimensions 8, 10 and 12. CICIDS2017 therefore carries six of the nine environments including Gate 1; the environments are not a probability sample of deployments. Train-defined preprocessing is applied without evaluation labels: TruncatedSVD per branch, training-fitted standardization for the RBF support-vector classifier (SVC) and frozen angular scaling to [0, 2π] for the fidelity-kernel branches (ZZ, PauliXYZ and a Z/PauliXZ-equivalent profile, one repetition). Gate 1 uses the Qiskit Machine Learning

6

TABLE 1 Adversary and failure model of the executed interventions. Roots: assets the class cannot write. Least separating evidence in the frozen design: the evaluated regime that first separates the class under the conformal family rule; “exact” means an item-aligned trusted reference; (a)–(c) are the intervention classes of Proposition 7. Mechanism and capability (bound- Class ary)

Roots it cannot write

Least separating evidence in frozen design

Supported claim

Random label flip: rewrite r of y at Integrity cor- X, P, C, E, ξ, g, f ; random (evaluation/report) ruption item-aligned y ⋆ if declared

exact only; IYm /IXF Y batch 0.00–0.01

Prior-preserving label flip: pairwise swaps preserving class counts (evaluation/report) Confusion-preserving relabeling: swaps among equal predictions (realised, not executed) Feature sign flip: negate p of X̃ entries (acquisition/preprocessing) Mean shift, scaling drift: shift or scale all features (acquisition/preprocessing) Feature dropout: impute p of entries (preprocessing)

Integrity cor- training data, fitted model, ruption reference features Fault robust- as above ness

invisible to IX , IXF (Cor. 1a); undetected at budget by marginal or batch joint evidence at these rates exact only; IXF Y batch invisible to IX , IXF , IYm (Cor. 1); 0.00–0.01 every material instance changes M (Prop. 3) item-aligned reference only item-identity violation without aggregate or conclusion impact (C1; 808 rows) IX 0.09–0.72; IXF partially separable; strength0.13–0.84; exact 0.79–0.91 dependent residual IX 0.99–1.00 by cluster detected in this design by a fragile fingerprint fingerprint

Fault robust- as above ness

batch 0.00–0.03; exact 0.49–0.76

Adaptive at- as above tacker Adaptive at- as above tacker

Cluster-preserving mean shift and Adaptive at- as above scaling: same mechanisms on the tacker 32% of entries outside tight clusters (executed, Gate A)

IX 0.01–0.57, IXF 0.05–0.65 at matched strengths; exact for every material row

Circuit-parameter mutation; repetition change: alter C before execution (design; (a)) Benign transpilation; commonunitary rewrite: compiler pass ((a)) Asymmetric, diagonal-eroded kernel edit: edit Kobs after estimation ((c)) PSD-preserving kernel substitution: valid-looking Kobs ((c))

circuit hash and semantic probe 15/15

Adversarial attack

reference circuit hash, probe kernel Ksem,0

Approved – equivalence Integrity cor- reference kernel, algebraic ruption invariants

hash changes 15/15, Ksem unchanged algebra 15/15

Adaptive at- reference kernel tacker

algebra 0/15; circuit hash 0/15; anchored Kobs comparison 15/15 repeated estimation 15/15 in these cells

Binomial shot emulation: finite-shot Fault robust- – estimate (execution; (b)) ness Post-hoc envelope edit: rewrite Adversarial stored contracts (record) attack

none assumed

reference evaluator [61]; the expansion uses a tagged exactstatevector evaluator validated against it at 10−10 on unit matrices and by an end-to-end replay (supplement); neither is shot-noise or QPU evidence. The intervention suite has six mechanisms at strengths 0.02, 0.05 and 0.10: feature sign flip, feature-wise mean shift, scaling drift, feature dropout with median imputation, random evaluation-label flip and prior-preserving evaluation-label flip. Feature-only sensors are feature Jensen–Shannon divergence (JSD), maximum mean discrepancy (MMD) and the Kolmogorov–Smirnov (KS) rejection rate against the clean version of the same item set; prediction, label-marginal and joint-outcome evidence likewise compare aggregate score, class and confusion-profile distributions with their clean same-item-set values. None uses item pairing. The runner protects these clean arrays as an experimental oracle; the policy API receives only sensor and family flags, but producing those flags still requires the upstream oracle. This is statistical class A, not evidence that ⋆ a deployed system authenticates the reference. Under IXF Y, confusion-profile deltas against a trusted aggregate are exact class B invariants, while prediction disagreement and label mismatch use trusted item alignment (class C). The complete

hash chain

separable in IXF yet missed by the executed finite-batch level-A rule reduces response against the declared fingerprint while keeping 83–91% of the conclusion changes; served by P2 in 27–39% of material rows identified by anchored provenance and semantics; outputs miss subdecision changes separability in h(C) is not harm (Prop. 7(i)) detected by structural invariants of Kobs sensor-level blindness closed only by a reference on the observed kernel (Prop. 7(ii)) held for adjudication; exact equality is not the criterion (Prop. 7(iii)); not hardware evidence tamper-evident only; no identity or non-repudiation

sensor-to-reference audit is in the supplement. The duplicate, clustered and near-constant structure on which this benchmark fingerprint depends is consistent with documented CICIDS2017 data-quality limitations [62]; the frozen staging is described rather than re-engineered here. 5.2

Experimental Units and Uncertainty

Gate 1 crosses five split seeds with four nested model seeds; its 7,980 raw rows contain 3,420 exact repeated SVC rows, leaving 4,560 unique observations. The expansion crosses the same five split seeds with two model seeds; all 360 job pairs completed and 13,680 raw rows contain 2,280 verified repeats, leaving 11,400 unique observations. Coverage endpoints are exact counts in the prespecified design; for the secondary model-profile endpoint (supplement) the five split seeds are the inferential units with two-sided 95% Student-t intervals (four degrees of freedom), describing within-environment resampling only. 5.3

Null Calibration and Decision-Level Calibration

The prespecified null-calibration gate (Gate N) calibrates the ten distributional sensors per (environment, dimension,

7

model) cell at a nominal per-sensor α = 0.05: the threshold is the 191st of 200 clean calibration draws from one half of the held-out evaluation pool, and false alarms are measured on 200 draws from the disjoint other half. Gate F (decisionlevel calibration) calibrates each regime as a family with Proposition 5(b)’s conformal rule at α = 0.05 per decision, using only the calibration draws. The same draws also evaluate the union and asymmetric comparison rules (Proposition 5(c)). The asymmetric rule’s pooled excess over nominal is decomposed into rule bias (its rate minus the conformal rate on identical draws) and design effect (the conformal rate minus nominal). Both rules are also evaluated under 30 exchangeable re-splits of the 400 pooled draws per cell. These re-splits diagnose the rule construction without changing the nonexchangeable primary design. Because a run’s 20 draws overlap within one pool half, false-alarm inference uses (environment, split-seed) clusters. Pooled rates are descriptive; intervals are five-cluster t intervals per environment. E1 (256-row draws from 322-row halves) remains in the preregistered primary aggregate and is also reported separately. 5.4 Decision Layer and Offline End-to-End Evaluation Gate D (the decision layer) composes sensors, the conformalrule or union decision, the information regime, the trustedreference status, a materiality threshold and the declared residual blind region into one allow/hold/block decision. block is reserved for violations of an exact invariant against a trusted reference; hold answers statistical evidence or, under P3, missing coverage; the decision composes with the four frozen contracts by the maximum in allow < hold < block. Four policies are fixed and named by what they are: P0 serve-always, the baseline; P1, the uncalibrated union rule of Gate N (risk-tolerant); P2, the conformal family rule, a conformal-rule risk-tolerant policy, nominally calibrated under exchangeability, that may serve under an explicitly declared residual blind region; and P3, the sensor-coverage-complete policy relative to the declared evidence dimensions, which abstains (holds) whenever a mandatory protected boundary (feature, prediction, label) has neither exact nor declared statistical coverage under the regime. P3 fails closed on missing coverage, not on insufficient power: a boundary covered by a low-power statistical sensor is served as under P2, so P3 guarantees no minimum detection power. The prototype’s contracts fail closed on their invariants; P2 does not and is never called fail-closed. Five regimes are evaluated: IX , IXF , IYm , IXF Y (class-A statistical aggregate ⋆ comparisons) and IXF Y (class-B aggregate and class-C itemaligned exact invariants; feature coverage remains statistical). The policy consumes union/family/exact flags, not clean arrays; the upstream scorer has the reference dependency described above. The observation set is frozen: 12,000 clean evaluation draws, 1,200 near-null synthetic controls applied in place (Gaussian noise σ = 0.001, scaling α = 0.001; not operational benign traffic), and the 10,800 intervened observations of the eight environments, 7,008 of which change the reported balanced accuracy (structural-sensitivity endpoint; τ = 0.02 and 0.05 are larger-effect sensitivities, not operational risk thresholds). The primary endpoint is a materially altered audit result served: a conclusionchanging (τ → 0+ ) observation decided allow. The CSV

retains unsafe_allow as a compatibility identifier; unsafe here means serving a materially corrupted audit/report conclusion, not allowing a malicious network event and not necessarily changing the classifier’s live prediction. The interruption endpoints are the clean false-action rate (clean ⋆ draws not allowed; for IXF Y the denominator is the 1,200 exact-zero rows) and the interruption rate on the near-null controls (the same 1,200 rows for every regime), decomposed into statistical holds and exact-reference blocks. The evaluation is offline, on frozen outputs, not a deployed runtime service; the twelve policy-evidence consistency checks fail closed, and every table and figure is generated from the manifested outputs. The 24,000 policy rows derived from the frozen intervention grid are not independent episodes; inference resides at environment/split clusters rather than at row level. 5.5

Adversarial Gate and Geometry Sensitivity

Gate A executes the preregistered cluster-preserving attacker. The standardized projected features hold tight clusters of rows (in every environment some feature keeps 69–87% of the rows within 0.01 standard deviations) that a fresh clean batch reproduces and any in-place displacement smears, which is why mean shift and scaling drift were detected in every cell. The attacker holds the batch, knows the sensor definitions and the fingerprint, and applies the executed mechanisms unchanged to every unclustered entry (an entry is clustered when at least ⌈0.05n⌉ rows, never fewer than two, lie within 0.01 batch standard deviations of it); clustered entries and labels are untouched. Strengths 0.02, 0.05, 0.10 match the frozen suite and 0.25, 0.50 probe whether materiality can be bought on the unclustered entries; the executed mechanisms run in the same jobs as matched controls and must reproduce the frozen expansion exactly. The gate runs in the eight environments with the frozen maps, caps and seeds (240 exact-statevector jobs) and is scored with the calibration draws, rules and policy layer of Gates F and D; window, minimum mass and strength grid were fixed before execution. The geometry-aligned sensitivity leaves frozen evidence intact and replays the same 240 configurations. For fixed E , calibration is s(E, Ck ) and aligned clean versus intervention scoring is s(E, B) versus s(E, T (B)), with fresh B fixed before results. Exact identity s(B, B) is separate from clean resampling. There is one row per frozen model cell and intervention; no label attack, training or selection is rerun (supplement). 5.6

Quantum Gate and Executable Prototype

A separate simulator gate isolates the quantum boundary with CICIDS, 64/64 samples, five split seeds, dimensions 4, 6 and 8 and 11 conditions (165 cells): clean, benign level1 transpilation and a common post-feature-map rewrite as controls; data-dependent RZ mutations at 0.02 and 0.10; a feature-map repetition change; asymmetric, diagonal-eroded and PSD-preserving kernel edits; and binomial fidelityestimation emulators at 256 and 1,024 shots. Each cell records canonical OpenQASM 3 provenance, kernel hashes, the anchored comparison against the trusted reference kernel, algebraic checks, repeated-estimation discrepancy and prediction

8

changes. The prototype instantiates four ordered contracts (input/preprocessing, circuit/kernel, execution/result, evaluation/report) in an SHA-256 chain over one compact CICIDS cell and six prespecified scenarios; it evaluates composition and decision logic, not throughput. The explicit Ksem , K̂ and Kobs transitions, semantic-equivalence checks and finite-shot estimator randomness make this a quantum-stage contract path rather than a generic ML-only pipeline. In ATHENAAEGIS terms, the four-contract prototype instantiates an integrity contract for the hybrid service composition exposed as HSaaS (Hybrid Software as a Service, a composition of classical and quantum services that the customer perceives as one service); it does not constitute deployed HSaaS validation, and provider assurance, confidentiality, QPU protection and the ATHENA fleet-management case study (WP5) are subsequent operational work of the subproject.

6

R ESULTS

6.1

Exact Label-Path Blind Regions

The zero-response cells are validation checks of the structural prediction, not findings: Corollary 1 says that a sensor whose input does not change cannot respond, and the checks confirm that the implementation realises the declared views. All 3,600 expansion and 1,440 Gate-1 evaluationlabel observations (of 11,400 and 4,560 unique observations) have zero prediction disagreement and exactly unchanged feature JSD, MMD, KS rejection and score JSD; all 1,800 and 720 prior-preserving observations also have zero prior shift and label JSD; the verifier recomputes every count. The informative quantity is how often these blind regions coincide with material conclusion changes and, in Section VIC, with serving decisions. The signed conclusion change is not one-sided: in the expansion 2,184 label observations lower the reported balanced accuracy, 983 leave it unchanged and 433 raise it (Gate 1: 1,276, 105, 59). Every one of the 2,617 (Gate 1: 1,335) label observations with a changed conclusion has non-zero item-aligned confusion evidence, and every one of the 2,184 with a lowered conclusion, the count used by the one-sided diagnostic, is among them. 6.2

Decision-Level Calibration

Over the 12,000 evaluation draws the per-sensor false-alarm rates lie between 0.029 and 0.069, but the batch-level rule “fire if any sensor fires” is not calibrated at the level of the decision: its rates are 0.123, 0.201, 0.048 and 0.257 for IX , IXF , IYm and IXF Y (three, six, two perfectly dependent and ten sensors). The conformal family rule brings them to 0.056, 0.059, 0.048 and 0.055 (Fig. 1(a)), observed rates in a design that violates the exchangeability premise. For comparison, the asymmetric construction gives 0.063, 0.074, 0.048 and 0.080 on the same draws. Its excess over 0.05 contains both rule bias and design effect. Under the same exchangeable resplits the conformal rule stays at or below level 0.0498 in all four regimes, while the asymmetric rule exceeds it in some. The remaining executed-design excess is concentrated in E1; E1 stays in the primary aggregate. Per-regime decomposition, E1/excluded-E1 values and the unadopted split construction appear in the supplement.

The executed rules realize Proposition 4(i): across 3,600 label observations, the IX /IXF sensors and rules fire zero times; the IYm rules likewise fire zero times on 1,800 prior-preserving rows. The primary empirical interpretation of finite-batch label response uses the geometry-aligned construction s(E, Ty (B)), calibration s(E, C) and clean s(E, B). It detects 343/2,700 material interventions with the conformal family rule and 1,183/2,700 with the uncorrected union. Among 2,836 aggregate-separable rows, responses are 352 conformal and 1,204 union; all 764 aggregate-blind aligned rows exactly equal their paired-clean sensor and rule responses, so attack-only increment is zero; this is the structural result. The original frozen benchmark geometry s(E, Ty (E)), retained for release reproduction and comparison and structurally valid for invariance, yields 11/2,617 conformal and 43/2,617 union, but is not the preferred estimate of statistical response under the clean-resample geometry. Nonresponse on separable rows is a finite-batch miss, not structural blindness. A trusted same-batch aggregate detects all 2,617 exactly (Corollary 3a) and 2,792 of the 3,600 label rows in total; the 808 relabelings that preserve every aggregate (C1) are exposed only by the item-aligned label reference (Corollary 3b). Feature-side coverage is mechanismdependent: mean shift and scaling are highly responsive because in-place perturbation smears tight projected-feature clusters that clean resampling reproduces; the same fragile fingerprint responds on 41–62% of near-null controls. Feature dropout with imputation is nearly invisible at the batch level (0–3%) yet changes predictions in 49–76% of cells. Within this fixed intervention grid, the conformal rule reduces detection only slightly: containment in IXF is 0.40 under the union and 0.38 under the conformal rule at the primary threshold, 0.70 versus 0.66 at τ = 0.05. Sensor dependence is substantial and explicitly quantified by the existing frozen-only ablation. After removing KS and recalibrating the remaining family, the core retains 3,521/4,361 IX , 4,008/4,462 IXF and 3,959/4,419 IXF Y primary fires, whereas the near-null retained counts fall to 10/650, 62/603 and 62/576. Thus KS dominates much of the near-null response associated with point masses and near-duplicate clusters, while much of the core feature-side response remains. KS stays because it was prespecified, and this descriptive ablation does not replace the primary family. 6.3

Offline End-to-End Decisions

Table 2 and Fig. 1(b,c) report the offline end-to-end evaluation. All aggregate served fractions summarize the prespecified equal-weight intervention grid; they are not deployment prevalence or expected operational unsafe rates. Under P0 every one of the 7,008 material observations is served. Under the batch-level regimes the conformal-rule policy P2 serves 4,496 (IX ), 4,368 (IXF ) and 4,394 (IXF Y ) of them with clean false-action rates 0.056, 0.059 and 0.055, whereas the uncalibrated union serves 4,429, 4,233 and 4,182 at 0.123, 0.201 and 0.257: the union buys 0.40 versus 0.37 containment in IXF Y at more than four times the clean false-action rate and 0.75 versus 0.48 interruption of the near-null controls. The asymmetric comparison rule served 4,490, 4,329 and 4,325; the adopted rule moves the count by less than 1% of the material rows. The served material corruptions are

9

0.0

X

XF

XFY

P0 serve-always

0.4

P1 union P2 conformal-rule P3 sensor-coverage-complete

0.2

P3 in X , XF , Ym : nothing served

0.0

0.6

0.25

0.50

0.75

Bars: pooled rate; whiskers: range of the eight per-environment cluster means. Dashed: nominal 0.05; the asymmetric rule is superseded; conformal exact level 10/201 = 0.0498.

0.77

0.2

X

1.00

Clean false-action rate (12,000 batch draws)

1.00

0.4

0.0 0.00

1.00

1.00

0.8

0.51 0.51

XF, XFY

0.75

0.6

1.0

0.48 0.48

X

0.75

0.8

Trusted check (separate estimand): 0/1,200 exact-zero false actions; 0/7,008 material reports served

0.50

0.080 0.055

Ym

1.0

0.75

0.257

0.1

0.048 0.048 0.048

0.074 0.059

0.2

0.123

0.3

0.201

0.4

0.063 0.056

Clean false-action rate (12,000 draws)

0.5

⋆XFY: 532 statistical holds + 85 gross exact blocks = 617; net +41 vs. batch P2

0.54

P0 (all regimes); P1, P2 in Ym

(c) Near-null control interruption Interruption of 1,200 near-null controls

(b) Batch policy trade-off

Union (uncalibrated) Asymmetric (superseded) Conformal (adopted)

Materially altered audit result served

(a) Executed false actions 0.6

XF

P1 union

 Ym

XFY ⋆XFY

P2 conformal-rule

P3 coverage

Same 1,200 rows for every regime; synthetic near-null variation, not operational traffic.

Fig. 1. Gates F and D (prespecified frozen evidence). (a) Decision-level false-action rate on 12,000 clean draws from evaluation pools disjoint from the calibration pools; resampled draws within each pool may overlap. Union, asymmetric-comparison and conformal rates are shown per regime. (b) Materially altered audit results served versus clean false-action rate at the structural-sensitivity endpoint (12,000-draw batch denominator); the trusted exact-zero check uses 1,200 rows and is a different estimand. (c) Interruption of 1,200 near-null controls; trusted P2 interrupts 617: 532 statistical holds and 85 gross exact blocks.

dominated by structure, not by thresholds: all 2,617 material label observations are served under IX and IXF (residual blind region), 2,606 of them under batch-level IXF Y (missed by the executed finite-batch rule), and 1,058 of 1,079 material dropout observations under IXF . The IYm view provides zero containment of material interventions in this suite; 5,450 material cases remain in its residual blind region at τ → 0+ , while all 7,008 material cases are served once both blind and statistically missed cases are counted. Containment 1.00 follows constructively from Corollary 3 under the declared trusted references; the experiment quantifies interruption within the prespecified near-null stress controls. Consistently, the trusted-reference regime serves none of the 7,008 material observations with zero clean false actions on its 1,200 exact-zero rows and blocks every nonidentity relabeling and every prediction change through its item-aligned references. Its stress-control interruption has two parts. Of the 1,200 near-null synthetic controls the trusted conformal-rule policy interrupts 617: 532 are statistical holds by its batch component, and 85 are gross exact-reference blocks of controls whose predictions changed under a 0.001 perturbation. Of those exact blocks, 44 already belong to the 576 interruptions produced by batch IXF Y /P2. The net increment is therefore 41/1,200, or 3.42 percentage points, not the gross 85/1,200. The controls are synthetic near-null variation, not operational traffic, so neither number is a production false-positive rate, and the two denominators of Table 2 are reported side by side. The sensor-coveragecomplete P3 (relative to the declared evidence dimensions) makes information-set conditionality literal: in IX , IXF and IYm the label boundary has no coverage, so nothing is served and every clean and near-null row is interrupted; ⋆ in IXF Y and IXF Y every boundary is covered and P3 coincides with P2, including P2’s 4,394 materially altered

TABLE 2 Offline decisions on 10,800 unchanged frozen interventions (7,008 conclusion-changing at the structural-sensitivity endpoint τ → 0+ ; τ = 0.05 is in the supplement). P1 is union, P2 conformal and P3 abstains on missing sensor coverage. Batch FPR uses 12,000 clean draws from evaluation pools disjoint from calibration pools; draws within a pool may overlap. The starred exact-zero check uses 1,200 rows and is a different estimand. ‡ P3 holds every row of a regime whose mandatory label or prediction boundary has no declared coverage; that cell is abstention by construction, not a detector false-action rate. Benign is interruption of 1,200 near-null controls; served counts are altered audit results, not operational events. Regime Policy IX IX IX IXF IXF IXF IYm IYm IYm IXF Y IXF Y ⋆ IXF Y

P1 union P2 conf. P3 abst. P1 union P2 conf. P3 abst. P1 union P2 conf. P3 abst. P1 union P2 conf. P2/P3 (exact)

FPR/check Benign Served Contain. Blind 0.123 0.056 n/a‡ 0.201 0.059 n/a‡ 0.048 0.048 n/a‡ 0.257 0.055 0.000

0.75 0.54 1.00 0.75 0.50 1.00 0.00 0.00 1.00 0.75 0.48 0.51

4,429 4,496 0 4,233 4,368 0 7,008 7,008 0 4,182 4,394 0

0.37 0.36 1.00 0.40 0.38 1.00 0.00 0.00 1.00 0.40 0.37 1.00

2,617 2,617 2,617 2,617 2,617 2,617 5,450 5,450 5,450 0 0 0

audit results served in IXF Y : abstention on missing coverage is not a guarantee of power, and the choice between P2 and P3 is a declared risk decision, not a detector property. The composition with the six frozen contract envelopes reproduces their actions (supplement). 6.4

Adaptive Cluster-Preserving Attacker

The frozen Gate A stress test used clean-resample calibration but scored attacked batches as s(E, T (E)). Its matched response falls from 0.94–1.00 to 0.01–0.65 while retaining 83–91% materiality; this is design-specific and against the

10 (b) Conclusion change kept by the attacker 1.0

0.8

0.8

Material fraction (|ΔR| > 0)

Detection rate, conformal family rule

(a) Executed drift (solid) vs. cluster-preserving (dashed) 1.0

0.6

0.4

0.2

0.0

0.6

0.4

0.2

0.0 0.02

0.05

0.10

0.25

0.50

0.02

Strength (δ or α, SD of the scaled feature) Mean shift, executed, X

Scaling, executed, X

Mean shift, cluster-preserving, X

Scaling, cluster-preserving, X

Mean shift, executed, XF

Mean shift, cluster-preserving, XF

Scaling, executed, XF

0.05

0.10

0.25

0.50

Strength Mean shift, executed Mean shift, cluster-preserving

Scaling, executed Scaling, cluster-preserving

Scaling, cluster-preserving, XF

Fig. 2. Corrected-JSD response on the original frozen Gate A rows. (a) Detection under the conformal family rule of the executed drift mechanisms (solid) and of their cluster-preserving variants (dashed) in IX and IXF , 600 observations per point. (b) Material fraction of the same rows: the construction reduces response against the declared batch-level fingerprint at matched strengths while keeping most of the conclusion changes.

declared cluster-sensitive fingerprint, not evidence of general evasion of drift detectors. Dataset cluster structure contributes to the response, and generalization beyond this fingerprint and these staged datasets remains open. Figure 2 shows the matched original-geometry response and materiality profiles used for this declared-fingerprint stress test. The aligned sensitivity confirms the qualitative ordering. Exact paired identity is 0, distinct from the clean-resample false-action rate (family rates 0.056–0.059). Mean shift and scaling remain almost fully responsive; dropout is 0.063– 0.303 and near-null response spans 0.540–0.723. Adaptive response is lower in 25–40 of 40 environment/split cells, never higher, with materiality of 0.262–0.667. Full strength/regime results are in the supplement. Same-item scoring amplified part of the original contrast, but aligned Gate A still shows that calibration controls clean-resample false actions without guaranteeing attack coverage. 6.5

Quantum-Workflow Coverage

All 165 dependent design cells (11 conditions × five splits × three dimensions) and all nine quantum-gate acceptance checks pass. In this ideal-statevector gate one anchored comparison against Φ(C0 ) is a semantic probe on circuit-side rows and an observed-kernel reference on post-processing rows; finite-shot estimation separates those anchors. Approved rewrites change provenance but preserve semantics (class (a)); PSD-preserving substitution escapes algebra and circuit provenance but not the anchored Kobs comparison (class (c)); both shot emulators yield non-zero repeatedestimation discrepancy in their 30 cells without establishing a universal rate (class (b)). Full per-condition responses, contract decisions and the secondary model profile are in the supplement. This is a bounded simulator instantiation of Proposition 7, not QPU/provider validation.

7

I NTEGRITY -AUDIT C ONTRACT

Each contract names its boundary, intervention class, view/reference, premise, residual region and response. For

labels, R0 , M0 and item alignment bind conclusion, aggregate and identity claims. The quantum contract combines provenance, approved equivalence, anchored kernels, algebra and repeated estimation. Exact-invariant violations block; missing mandatory evidence holds; P2 may serve within a residual region, while P3 promises no minimum power.

8

D ISCUSSION AND L IMITATIONS

8.1

Security Interpretation

A metric can change while predictor output is unchanged: input robustness cannot repair corrupted evaluation labels. Zero response establishes blindness only when the evidence could not distinguish the states; otherwise power depends on the null. The 27–39% served fractions summarize the prespecified Gate A adaptive cluster-preserving intervention grid, not deployment prevalence. Gate A is a stress test of one cluster-sensitive fingerprint; its aligned sensitivity preserves the control/adaptive ordering while exposing a geometry effect. Evidence sufficiency and necessity are claim-relative within the declared reference lattice, where granularity means authenticated or bound information, not digest length. A Merkle root or position-binding vector commitment can compactly bind level-C item identity [63], [64]. Level B instead lets aggregate- or conclusion-only audits verify M0 or R0 without retaining or exposing the complete item-wise label store, which can reduce disclosure and support separation of duties over sensitive labels. This is evidence minimization, not a new cryptographic primitive or legal guarantee. An external root remains necessary (Proposition 6); this is neither a deployed protocol nor a minimum over every possible audit architecture. 8.2

Validity Boundaries

The eight fixed settings are not sampled deployments, and balanced staging does not reproduce prevalence. The quantum branch uses ideal-statevector and finite-shot emulation, not QPU/provider validation. The secondary SVC/QSVC profile does not isolate kernel geometry. Structural blind regions are independent of the number of rows: a viewpreserving intervention remains invisible as n grows. For interventions that do change an observed statistic, however, power can increase with batch size as sampling variability often decreases, approximately as n−1/2 for proportion-like statistics under regular conditions. The primary geometryaligned outcomes— 343/2,700 conformal and 1,183/2,700 union—and the original frozen outcomes—11/2,617 conformal and 43/2,617 union—are finite-design results, not power bounds. E1 is the only 256-row environment and exhibits nonzero label response, but because batch size is confounded with environment and null behavior, this cannot identify a batch-size effect. The conformal level requires exchangeability; overlapping draws, different pool rows and E1 violate it, so executed rates are descriptive. Gate A fixes one fingerprint and two mechanisms; generalization beyond its dataset structure remains open. The algebraic quantum check concerns the raw estimated/observed kernel before downstream PSD repair; a repaired-kernel consumer defines a different, unevaluated observable path.

11

The trusted item-aligned reference is an assumption under Proposition 6. P3 abstains on missing coverage, not low power. All policies are offline; runtime non-stationarity, recovery, QPU execution and service enforcement remain open.

9

C ONCLUSION

Hybrid workflow integrity is evidence-relative. Interventions can be structurally invisible, statistically missed in a finite batch, or show reduced response against the declared monitored fingerprint. The geometry-aligned label construction is the primary statistical interpretation; its 764 aggregate-blind rows retain zero attack-only increment. The substantial KS dependence is quantified without replacing the prespecified family. The conformal guarantee remains conditional on exchangeability and the quantum instantiation simulatorbounded. Within the declared lattice, authenticated R0 , M0 and item alignment certify conclusion, aggregate-plusconclusion and item-identity claims, respectively. Assurance must state its view, trust root, premise, residual region and response.

DATA , C ODE , AND E THICS S TATEMENT The version 1.3.9 artifact contains code, tests, frozen and corrective derived evidence and eleven manifests; raw benchmarks are not redistributed. Sources, hashes and staging are documented. Its version DOI is 10.5281/zenodo.22750616 under concept DOI 10.5281/zenodo.22550852; immutable predecessor v1.3.8 has DOI 10.5281/zenodo.22706167. Code is Apache-2.0 and derived evidence/documentation CC BY 4.0. No new human-participant or personal-data collection was performed.

ACKNOWLEDGMENT OpenAI Codex and Anthropic Claude Code assisted with drafting and language review in the Abstract, Introduction, Related Work, Methods, Results, Discussion, data/code statement and supplement; targeted literature-search support; code generation and inspection for analysis, tests and release tools; and table, figure, artifact and submission-package preparation. Human authors independently validated every formal statement and proof, reference, code path, experiment, result and final passage and retain full responsibility. Neither system is an author or proof authority; detailed disclosure accompanies the submission.

R EFERENCES [1]

D. Volya, T. Zhang, N. Alam, M. Tehranipoor, and P. Mishra, “Towards secure classical-quantum systems,” in IEEE Int. Symp. Hardware Oriented Security and Trust (HOST). IEEE, 2023, pp. 283–292. [2] S. Ghosh, S. Upadhyay, and A. A. Saki, “A primer on security of quantum computing hardware,” Proceedings of the IEEE, vol. 113, no. 7, pp. 640–667, 2025. [3] F. Pasqualetti, F. Dörfler, and F. Bullo, “Attack detection and identification in cyber-physical systems,” IEEE Transactions on Automatic Control, vol. 58, no. 11, pp. 2715–2729, 2013. [4] A. Teixeira, I. Shames, H. Sandberg, and K. H. Johansson, “A secure control framework for resource-limited adversaries,” Automatica, vol. 51, pp. 135–148, 2015.

Y. Liu, P. Ning, and M. K. Reiter, “False data injection attacks against state estimation in electric power grids,” ACM Transactions on Information and System Security, vol. 14, no. 1, 2011. [6] J. W. Stokes, P. England, and K. Kane, “Preventing machine learning poisoning attacks using authentication and provenance,” in Proc. IEEE Military Communications Conference (MILCOM), 2021, pp. 181– 188. [7] F. Hinder, V. Vaquet, and B. Hammer, “Adversarial attacks for drift detection,” in ESANN 2025 Proceedings. Ciaco–i6doc.com, 2025, pp. 555–560. [8] E. Yeniaras and M. A. Karimov, “QCIVET: A quantum–classical pipeline integrity framework with contract-based subtype verification and hash-chained audit traces,” arXiv preprint arXiv:2605.13109, 2026. [9] E. Yeniaras, “QML-PipeGuard: Drift-aware behavioral fingerprinting for quantum machine learning pipeline integrity,” arXiv preprint arXiv:2605.25066, 2026, submitted 24 May 2026. [10] C.-Z. Bai, F. Pasqualetti, and V. Gupta, “Data-injection attacks in stochastic control systems: Detectability and performance tradeoffs,” Automatica, vol. 82, pp. 251–260, 2017. [11] D. I. Urbina, J. A. Giraldo, A. A. Cárdenas, N. O. Tippenhauer, J. Valente, M. Faisal, J. Ruths, R. Candell, and H. Sandberg, “Limiting the impact of stealthy attacks on industrial control systems,” in Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security (CCS), 2016, pp. 1092–1105. [12] J. Giraldo, D. Urbina, A. Cardenas, J. Valente, M. Faisal, J. Ruths, N. O. Tippenhauer, H. Sandberg, and R. Candell, “A survey of physics-based attack detection in cyber-physical systems,” ACM Computing Surveys, vol. 51, no. 4, 2018. [13] D. Seto, B. Krogh, L. Sha, and A. Chutinan, “The simplex architecture for safe online control system upgrades,” in Proceedings of the 1998 American Control Conference, vol. 6. IEEE, 1998, pp. 3504–3508. [14] L. Sha, “Using simplicity to control complexity,” IEEE Software, vol. 18, no. 4, pp. 20–28, 2001. [15] M. Leucker and C. Schallhart, “A brief account of runtime verification,” Journal of Logic and Algebraic Programming, vol. 78, no. 5, pp. 293–303, 2009. [16] G. H. Kim and E. H. Spafford, “The design and implementation of Tripwire: A file system integrity checker,” in Proceedings of the 2nd ACM Conference on Computer and Communications Security (CCS), 1994, pp. 18–29. [17] Z. C. Lipton, Y.-X. Wang, and A. J. Smola, “Detecting and correcting for label shift with black box predictors,” in Proc. 35th Int. Conf. Machine Learning, ser. Proc. Machine Learning Research, vol. 80, 2018, pp. 3122–3130. [18] A. A. Ginart, M. J. Zhang, and J. Zou, “MLDemon: Deployment monitoring for machine learning systems,” in Proc. 25th Int. Conf. Artificial Intelligence and Statistics (AISTATS), ser. Proc. Machine Learning Research, vol. 151. PMLR, 2022, pp. 3962–3997. [Online]. Available: https://proceedings.mlr.press/v151/ginart22a.html [19] A. Koebler, T. Decker, I. Thon, V. Tresp, and F. Buettner, “Incremental uncertainty-aware performance monitoring with active labeling intervention,” in Proc. 28th Int. Conf. Artificial Intelligence and Statistics (AISTATS), ser. Proc. Machine Learning Research, vol. 258. PMLR, 2025, pp. 2188–2196. [Online]. Available: https://proceedings.mlr.press/v258/koebler25a.html [20] O. Solozobov, “Evidence sufficiency under delayed ground truth: Proxy monitoring for risk decision systems,” arXiv preprint arXiv:2604.15740, 2026. [21] C. G. Northcutt, A. Athalye, and J. Mueller, “Pervasive label errors in test sets destabilize machine learning benchmarks,” in Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, 2021, arXiv:2103.14749. [22] D. Arp, E. Quiring, F. Pendlebury, A. Warnecke, F. Pierazzi, C. Wressnegger, L. Cavallaro, and K. Rieck, “Dos and don’ts of machine learning in computer security,” in 31st USENIX Security Symposium (USENIX Security 22). Boston, MA: USENIX Association, Aug. 2022, pp. 3971–3988. [Online]. Available: https:// www.usenix.org/conference/usenixsecurity22/presentation/arp [23] F. Pendlebury, F. Pierazzi, R. Jordaney, J. Kinder, and L. Cavallaro, “TESSERACT: Eliminating experimental bias in malware classification across space and time,” in 28th USENIX Security Symposium (USENIX Security 19). Santa Clara, CA: USENIX Association, Aug. 2019, pp. 729–746. [Online]. Available: https://www.usenix.org/conference/usenixsecurity19/ presentation/pendlebury [5]

12

[24] P. Bajaj, “Evaluation blindness: How silent measurement failures corrupt AI systems from training to deployment,” arXiv preprint arXiv:2608.02786, 2026, submitted 3 August 2026. [25] S. Torres-Arias, H. Afzali, T. K. Kuppusamy, R. Curtmola, and J. Cappos, “in-toto: Providing farm-to-table guarantees for bits and bytes,” in 28th USENIX Security Symposium (USENIX Security 19). Santa Clara, CA: USENIX Association, 2019, pp. 1393–1410. [Online]. Available: https://www.usenix.org/ conference/usenixsecurity19/presentation/torres-arias [26] SLSA Community, “Supply-chain levels for software artifacts specification, version 1.2,” https://slsa.dev/spec/v1.2/, 2025, released 24 November 2025; accessed 8 September 2026. [27] Z. Newman, J. S. Meyers, and S. Torres-Arias, “Sigstore: Software signing for everybody,” in Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security. ACM, 2022, pp. 2353–2367. [28] G. Coker, J. D. Guttman, P. A. Loscocco, A. L. Herzog, J. K. Millen, B. O’Hanlon, J. D. Ramsdell, A. Segall, J. Sheehy, and B. T. Sniffen, “Principles of remote attestation,” International Journal of Information Security, vol. 10, no. 2, pp. 63–81, 2011. [29] H. Birkholz, D. Thaler, M. Richardson, N. Smith, and W. Pan, “Remote ATtestation procedureS (RATS) architecture,” RFC Editor, RFC 9334, Jan. 2023. [30] V. Vovk, A. Gammerman, and G. Shafer, Algorithmic Learning in a Random World. New York: Springer, 2005. [31] A. N. Angelopoulos and S. Bates, “Conformal prediction: A gentle introduction,” Foundations and Trends in Machine Learning, vol. 16, no. 4, pp. 494–591, 2023. [32] A. Timans, C.-N. Straehle, K. Sakmann, C. A. Naesseth, and E. Nalisnick, “Max-rank: Efficient multiple testing for conformal prediction,” in Proceedings of the 28th International Conference on Artificial Intelligence and Statistics, ser. Proceedings of Machine Learning Research, vol. 258. PMLR, 2025, pp. 3898–3906. [Online]. Available: https://proceedings.mlr.press/v258/timans25a.html [33] L. H. C. Tippett, The Methods of Statistics: An Introduction Mainly for Workers in the Biological Sciences. London: Williams and Norgate, 1931. [34] P. H. Westfall and S. S. Young, Resampling-Based Multiple Testing: Examples and Methods for p-Value Adjustment. New York: Wiley, 1993. [35] J. Lei, M. G’Sell, A. Rinaldo, R. J. Tibshirani, and L. Wasserman, “Distribution-free predictive inference for regression,” Journal of the American Statistical Association, vol. 113, no. 523, pp. 1094–1111, 2018. [36] R. Laxhammar and G. Falkman, “Inductive conformal anomaly detection for sequential detection of anomalous sub-trajectories,” Annals of Mathematics and Artificial Intelligence, vol. 74, no. 1–2, pp. 67–94, 2015. [37] S. Bates, E. Candès, L. Lei, Y. Romano, and M. Sesia, “Testing for outliers with conformal p-values,” Annals of Statistics, vol. 51, no. 1, pp. 149–178, 2023. [38] R. F. Barber, E. J. Candès, A. Ramdas, and R. J. Tibshirani, “Conformal prediction beyond exchangeability,” Annals of Statistics, vol. 51, no. 2, pp. 816–845, 2023. [39] S. Kundu and S. Ghosh, “Security aspects of quantum machine learning: Opportunities, threats and defenses,” in Proc. Great Lakes Symp. VLSI (GLSVLSI). ACM, 2022, pp. 463–468. [40] D. Oliveira, E. Giusto, B. Baheri, Q. Guan, B. Montrucchio, and P. Rech, “A systematic methodology to compute the quantum vulnerability factors for quantum circuits,” IEEE Transactions on Dependable and Secure Computing, vol. 21, no. 4, pp. 2631–2644, 2024. [41] C. Lu, E. Telang, A. Aysu, and K. Basu, “Quantum leak: Timing sidechannel attacks on cloud-based quantum services,” in Proceedings of the Great Lakes Symposium on VLSI 2025. ACM, 2025, pp. 252–257. [42] E. Ahmed, B. Ye, S. H. Shah, M. A. Akbar, and A. A. Khan, “A multilevel integrity evaluation framework for quantum circuits under controlled anomaly injection,” arXiv preprint arXiv:2604.26430, 2026. [43] A. Bensoussan, G. Jahangirova, and M. R. Mousavi, “A taxonomy of real faults for hybrid quantum-classical software architectures,” ACM Transactions on Software Engineering and Methodology, vol. 35, no. 9, pp. 1–30, 2026. [44] A. Bensoussan, H. D. Menendez, and M. R. Mousavi, “Quantum squeeziness: An information theoretical metric for quantum software testability,” in Proceedings of the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE 2026). Munich, Germany: ACM, 2026, to appear.

[45] B. Weder, J. Barzen, F. Leymann, M. Salm, and K. Wild, “QProv: A provenance system for quantum computing,” IET Quantum Communication, vol. 2, no. 4, pp. 171–181, 2021. [46] J. Peltonen, V. Stirbu, T. Mikkonen, and C. Pautasso, “Toward standardized quantum provenance: A cross-provider analysis, unified API, and reference prototype,” arXiv preprint arXiv:2608.08272, 2026, submitted 8 August 2026. [47] J. Wang, Q. Zhang, G. H. Xu, and M. Kim, “QDiff: Differential testing of quantum software stacks,” in 36th IEEE/ACM Int. Conf. Automated Software Engineering (ASE). IEEE, 2021, pp. 692–704. [48] M. Paltenghi and M. Pradel, “MorphQ: Metamorphic testing of the Qiskit quantum computing platform,” in 45th IEEE/ACM Int. Conf. Software Engineering (ICSE). IEEE, 2023, pp. 2413–2424. [49] J. Luo, S. Xia, F. Zhang, and J. Zhao, “QEMI: A quantum software stacks testing framework via equivalence modulo inputs,” in Fundamental Approaches to Software Engineering, ser. Lecture Notes in Computer Science, vol. 16504. Springer, 2026, pp. 149–169. [50] Y. Shi, R. Tao, X. Li, A. Javadi-Abhari, A. W. Cross, F. T. Chong, and R. Gu, “CertiQ: A mostly-automated verification of a realistic quantum compiler,” arXiv preprint arXiv:1908.08963, 2019. [51] M. Ying, “Floyd–hoare logic for quantum programs,” ACM Transactions on Programming Languages and Systems, vol. 33, no. 6, 2011. [52] L. Zhou, N. Yu, and M. Ying, “An applied quantum hoare logic,” in 40th ACM SIGPLAN Conf. Programming Language Design and Implementation, 2019, pp. 1149–1162. [53] G. Li, L. Zhou, N. Yu, Y. Ding, M. Ying, and Y. Xie, “Projectionbased runtime assertions for testing and debugging quantum programs,” Proceedings of the ACM on Programming Languages, vol. 4, no. OOPSLA, 2020. [54] M. Yamaguchi and N. Yoshioka, “Design by contract framework for quantum software,” in 2023 IEEE/ACM 4th International Workshop on Quantum Software Engineering (Q-SE). IEEE, 2023, pp. 24–25. [55] R. Fernández-Barrios, I. Pastor-López, A. González-Santocildes, and P. Garcı́a Bringas, “Sharp target-domain certificates for quantumkernel advantage under distribution shift,” Submitted manuscript, University of Deusto; public manuscript and reproducibility repository, 2026, submitted to EPJ Quantum Technology, 6 September 2026; the related DOI 10.5281/zenodo.21776862 identifies the software artifact, not the article. [Online]. Available: https://github. com/roberto-fernandez-barrios/target-domain-certificates [56] ——, “Conditional validity of quantum event classifiers under collider systematics and quantum estimation uncertainty,” arXiv preprint arXiv:2609.02781, 2026, submitted 2 September 2026. [57] R. Fernández-Barrios, I. Pastor-López, A. Pikatza-Huerga, and P. Garcı́a Bringas, “Candidate comparability before promotion: Conditional validation in adaptive network intrusion detection,” arXiv preprint arXiv:2609.04388, 2026, submitted 3 September 2026. [58] I. Sharafaldin, A. H. Lashkari, and A. A. Ghorbani, “Toward generating a new intrusion detection dataset and intrusion traffic characterization,” in Proc. 4th Int. Conf. Information Systems Security and Privacy, 2018, pp. 108–116. [59] N. Moustafa and J. Slay, “UNSW-NB15: A comprehensive data set for network intrusion detection systems,” in 2015 Military Communications and Information Systems Conference. IEEE, 2015, pp. 1–6. [60] A. Alsaedi, N. Moustafa, Z. Tari, A. Mahmood, and A. Anwar, “TON IoT telemetry dataset: A new generation dataset of IoT and IIoT for data-driven intrusion detection systems,” IEEE Access, vol. 8, pp. 165 130–165 150, 2020. [61] A. Javadi-Abhari, M. Treinish, K. Krsulich, C. J. Wood, J. Lishman, J. Gacon, S. Martiel, P. D. Nation, L. S. Bishop, A. W. Cross, B. R. Johnson, and J. M. Gambetta, “Quantum computing with Qiskit,” arXiv preprint arXiv:2405.08810, 2024. [62] G. Engelen, V. Rimmer, and W. Joosen, “Troubleshooting an intrusion detection dataset: The CICIDS2017 case study,” in Proc. IEEE Security and Privacy Workshops (SPW). IEEE, 2021, pp. 7–12. [63] R. C. Merkle, “A digital signature based on a conventional encryption function,” in Advances in Cryptology—CRYPTO ’87, ser. Lecture Notes in Computer Science, vol. 293. Springer, 1988, pp. 369–378. [64] D. Catalano and D. Fiore, “Vector commitments and their applications,” in Public-Key Cryptography—PKC 2013, ser. Lecture Notes in Computer Science, vol. 7778. Springer, 2013, pp. 55–72.

Record · ID 919245 · SHA-256 52ff80a18e26edbf
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.