ConceptioArchivearXiv CS
arXiv CSopen access

Decoherence as Defence and the Magnitude of Noise Regularisation: A Rigorous N -Qubit Theory of Stochastic Quantum Neural Networks for Adversarially Robust Network Intrusion Detection

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
cryptography, security, privacy, cybersecurity

arXiv:2606.24219v1 [cs.CL] 23 Jun 2026

Decoherence as Defence and the Magnitude of Noise Regularisation: A Rigorous N -Qubit Theory of Stochastic Quantum Neural Networks for Adversarially Robust Network Intrusion Detection Gautier-Edouard Filardo Efrei Research Lab, Efrei Paris Panthéon-Assas Université, Villejuif, France

Abstract Stochastic quantum neural networks (SQNNs) encode neuronal activations as qubits, synaptic topology as entanglement, and neural noise through a Lindblad master equation. A recent conference study applied a ringentangled SQNN to collaborative intrusion detection and reached three conclusions: ring entanglement is essential for non-local anomaly detection; an adversarial-resilience bound holds but is conservative; and the depolarising channel fails to act as a dropout-style regulariser, behaving instead as output noise. It left open whether a per-gate stochastic deactivation (“true quantum dropout”) could regularise where the depolarising channel could not, and whether the loose robustness bound could be replaced by a predictive theory. This paper resolves both and extends the framework to real data and to neutral-atom hardware. We give an N -qubit formulation through the stochastic master equation and its vectorised Liouvillian, and prove a decoherence-contraction theorem: a depolarising channel of strength γ over L entangling layers contracts every weight-w Pauli read-out by a factor (1 − 4γ/3)wL (for the weight-1 read-out used here, (1 − 4γ/3)L ); building on the general noise-as-defence result of Du et al., we make this quantitative and operational for intrusion detection. On the real NSL-KDD dataset under white-box FGSM and PGD attacks, a depolarising SQNN trained with the channel is, over seven seeds under strong ℓ∞ /ℓ2 attacks, significantly more robust than the noiseless circuit (ℓ∞ PGD-20, p = 0.04, large effect) and, Email address: [email protected] (Gautier-Edouard Filardo)

Preprint submitted to Neurocomputing

June 24, 2026

critically, never suffers the catastrophic robustness collapse that the noiseless model and gradient-trained classical detectors (which fall from 95% to 47%) do, cutting robustness variance roughly twofold; we show this robustness arises from a noise-reshaped training boundary rather than from attacktime gradient contraction. For generalisation, we derive an adaptive-penalty formula showing that per-gate dropout implements a curvature-weighted L2 p(1−p) P 2 2 θ ∂θ L in weight space, maximised at p = 1/2, whereas depopenalty 2 larising noise implements an output-space penalty. A 30-seed study confirms the formula’s quantitative prediction: both mechanisms reduce the train– test gap by a small but statistically significant margin (≈ 0.01; p < 10−4 and p = 0.004), are statistically indistinguishable from each other, and the effect is concentrated where overfitting is largest; increasing the dropout rate past 1/2 does not help, as the formula predicts. The single-seed dichotomy of prior work does not survive replication. We close with a neutral-atom realisation and a feasibility-by-N analysis. Keywords: Stochastic quantum neural networks, Adversarial robustness, Network intrusion detection, Quantum decoherence, Lindblad formalism, Quantum dropout, Adaptive regularisation, Neutral-atom quantum computing, NSL-KDD 1. Introduction Collaborative enterprise infrastructures — industrial IoT, distributed supply chains, and multi-organisational processes — generate data streams whose continuous monitoring is essential for detecting coordinated cyber threats before they propagate [3, 4]. Classical machine-learning detectors are effective but face three structural limitations in this setting: sequential hypothesis exploration, difficulty capturing non-local correlations between distant nodes without explicit feature engineering, and acute vulnerability to adversarial perturbations that nudge a malicious connection across a learned decision boundary while remaining operationally valid [5, 6]. The stakes of detector robustness are rising with the deployment context. Machine-learning detectors increasingly sit inside high-evasion-cost pipelines — automated response systems and, more recently, tool-using software agents — where a single fooled detector is no longer a missed alert but a system-level compromise. Hardening the detector against adversarial evasion is therefore a security primitive, not merely an accuracy concern. 2

Quantum machine learning offers, in principle, superposition for parallel hypothesis exploration and entanglement for native modelling of non-local correlations [7, 8]. Within the noisy intermediate-scale quantum (NISQ) regime [9], variational quantum circuits (VQCs) [10] are the dominant practical model. Our prior work introduced stochastic quantum neural networks (SQNNs) [1], in which qubit-neurons evolve under a stochastic master equation with Wiener-process fluctuations, treating controlled noise as a computational resource, and a subsequent study specialised the SQNN to collaborative anomaly detection in distributed enterprise systems [2]. That study [2] established, on synthetic tasks, that (i) ring-topology entanglement is essential for non-local anomaly detection (84.5% versus 49.9% on an XOR-structured dataset, ring closure +6.0 points); (ii) the adversarial bound |∆A| ≤ ∥ϵ∥/(γ + ∥ϵ∥) holds but overestimates vulnerability by 3–100×; and (iii) the depolarising channel does not act as a dropout-style regulariser, because it perturbs the measured expectation (analogous to label noise) rather than masking subsystems. It proposed, but did not test, replacing each variational rotation by the identity with some probability during training — a per-gate stochastic deactivation analogous to classical dropout — and called for a predictive replacement of its conservative robustness bound. This paper takes up those threads. Its contributions are: 1. an N -qubit formulation of the SQNN through its stochastic master equation and vectorised Liouvillian, with an explicit feasibility-by-N account (Section 3); 2. a decoherence-contraction theorem giving the weight-resolved law (1 − 4γ/3)wL for how the channel scales Pauli read-outs (reducing to (1 − 4γ/3)L for the weight-1 read-out used here); building on the general noise-as-defence result of Du et al. [25], this turns the conservative bound of [2] into a predictive, operational law for intrusion detection (Section 4); 3. a multi-seed demonstration on the real NSL-KDD dataset that a depolarising SQNN trained with the channel is significantly more robust under strong, exact-gradient ℓ∞ /ℓ2 attacks and, unlike the noiseless circuit and gradient-trained classical detectors, free of catastrophic robustness collapse — an effect we trace to a training-time reshaping of the decision boundary, not attack-time gradient contraction, and which generalises to a second real dataset (UNSW-NB15) where the 3

SQNN is benchmarked white-box against defended classical baselines (Sections 4, 8.2, 8.3); 4. an adaptive-penalty formula for per-gate quantum dropout, and a rigorous 30-seed evaluation showing that both depolarising noise and pergate dropout produce a small but significant, and mutually indistinguishable, reduction of the generalisation gap — with the formula predicting both the magnitude and the optimal dropout rate, and the single-seed dichotomy of prior work failing to replicate (Sections 5, 8); 5. the identification of the ring Heisenberg–Ising Hamiltonian with the neutral-atom Rydberg Hamiltonian, with a measured feasibility characterisation and a reservoir realisation (Section 6). We are explicit about what is not claimed. On the low-dimensional tabular tasks studied here, no quantum speed-up or accuracy advantage over strong classical baselines is claimed or observed; the contributions are a rigorous theory of noise-induced robustness and regularisation, a predictive formula, and a hardware-faithful realisation. All reported numbers are produced by simulation; where an experiment was not run we say so. 2. Background and related work SQNNs and VQCs.. The SQNN formalism [1] represents each neuron as a qubit evolving under an open-system stochastic equation with Lindblad dissipators [39, 40, 41, 42]. Variational quantum models [10, 11, 12, 13] parameterise a circuit and optimise an observable by hybrid gradient descent; expressivity and trainability are limited by depth and by barren plateaus [14], which noise can exacerbate [15]. Deep quantum neural networks [16], quantum convolutional networks [17], and capacity analyses [18] situate the SQNN within a broad family. Adversarial robustness.. For a differentiable classifier f , FGSM perturbs an input by ε sign(∇x L) [5], and PGD iterates projected FGSM steps within an ℓ∞ ball [6]; both exploit the model’s input gradient [19, 20, 21]. Randomised smoothing certifies robustness by adding noise at inference [22], the classical analogue of the quantum-channel mechanism we analyse. Quantum adversarial robustness.. Quantum classifiers are themselves vulnerable to adversarial perturbations [23, 24, 26]. Most directly related, Du et al. [25] proved that depolarising noise can protect quantum classifiers against 4

adversaries in general, and follow-up work explored quantum-enhanced robustness [27]. Our contribution is complementary and more specific: we derive the explicit contraction law η wL relating read-out weight to robustness, specialise it to intrusion detection on real data, use it to tighten the conservative bound of [2], and pair it with a regularisation formula and a neutral-atom realisation. We do not claim the noise-as-defence idea as novel in general; we make it quantitative and operational for intrusion detection. Detection in high-evasion-cost settings.. Behavioural anomaly detection — monitoring for unusual action sequences, unexpected resource accesses, and exfiltration-like activity — is a recognised defence family for systems whose misuse is costly, including the tool-using-agent pipelines now entering production. A useful conceptual bridge for the present work is that an instructionlevel manipulation such as prompt injection is, to first approximation, an adversarial attack: a small, crafted perturbation that flips a model’s behaviour. The adversarial robustness of a feature-space classifier and the injection robustness of an ML-integrated system are thus cousins. We stress this is an analogy, not an equivalence: the threat models differ (a bounded ℓ∞ perturbation in feature space versus a natural-language instruction), and we make no agent-level experimental claim here. We return to it only as motivation and as a direction for future work (Sections 1, 9). Datasets and hardware.. We evaluate on NSL-KDD [28] and discuss UNSWNB15 [29] and CIC-DDoS2019 [30]. The Ising-type Hamiltonian of the SQNN maps naturally to neutral-atom Rydberg processors [31, 32, 33, 34], whose pulse-level control is exposed by open-source tooling [35]. Classical regularisation by dropout [36, 37] and its interpretation as adaptive penalisation underpin Section 5. 3. An N -qubit theory of stochastic quantum neural networks Let H = (C2 )⊗N , dim H = 2N , and ρ ∈ B(H) a density operator. Write (i) (i) σa for the Pauli operator on qubit i, σ ± = 12 (σx ± iσy ), and ni = 21 (I − σz ) for the excitation number. Any operator expands in the Pauli-string frame X O 1 ra σa , σa = σai , ra = Tr(ρ σa ), (1) ρ= N 2 N i a∈{0,x,y,z} P generalising the single-qubit Bloch vector. The read-out, ⟨O⟩ = i ⟨ni ⟩, is a weight-one σz observable; pairwise correlations ⟨ni nj ⟩ are weight-two. 5

3.1. The stochastic master equation For a weakly, continuously measured open network, the conditional state obeys the N -qubit stochastic master equation X X (k) dρt = − ℏi [H, ρt ] dt + D[Lk ]ρt dt + H[Lk ]ρt dWt , (2) k

k

with dissipator D[L]ρ = LρL† − 12 {L† L, ρ}, measurement superoperator H[L]ρ = Lρ + ρL† − Tr[(L + L† )ρ]ρ, and independent Wiener increments dW (k) . Averaging over the record yields the Gorini–Kossakowski–Sudarshan–Lindblad (GKSL) generator X D[Lk ]ρ. (3) ρ̇ = Lρ, Lρ = − ℏi [H, ρ] + k

3.2. Ring Hamiltonian and qubit-neuron channels With biases hi and synaptic graph G = (V, E) we take the anisotropic Heisenberg–Ising Hamiltonian X X  H= hi σz(i) + Jij σx(i) σx(j) + σy(i) σy(j) + λ σz(i) σz(j) , (4) i

(i,j)∈E

with ring topology Ering = {(i, i+1 mod N )}. The qubit-neuron channels are p √ √ (i) (i) = γ2 /2 σz , and Lmeas = γm σz . Lrelax = γ1 σi− , Ldeph i i i 3.3. Vectorised generator and the cost of scale N Stacking ρ into |ρ⟩⟩ ∈ C4 (so vec(AρB) = (B T⊗A)|ρ⟩⟩), the generator (3) becomes the 4N × 4N Liouvillian   X ∗ L = − ℏi I ⊗ H − H T ⊗ I + Lk ⊗ Lk − 12 I ⊗ L†k Lk − 12 (L†k Lk )T ⊗ I , (5) k

and ρt = eLt ρ0 is the completely-positive trace-preserving semigroup. A pureN state trajectory of (2) lives in C2 , but any exact density-matrix treatment N lives in C4 ; for N = 20 this is 420 ≈ 1.1×1012 amplitudes (Table 1). Beyond the dense wall, pure-state trajectory sampling (cost 2N per noise realisation, parallel over seeds) and tensor-network surrogates preserve large N .

6

Table 1: Realisability by network size, measured on one CPU core (≈ 4 GB). Noiseless cost is per driven evolution; the exact-noisy regime is bounded by the 4N generator.

N

state vector 2N

density matrix 2N ×2N

noiseless cost/evol.

exact noisy (Lindblad)

8 10 12 20

256 1024 4096 1.0 × 106

6.5 × 104 1.0 × 106 1.7 × 107 1.1 × 1012

0.12 s ≈ 1s 24 s hours (dense)

feasible (≈ 4 GB) needs ≳ 16 GB workstation-scale infeasible (any hardware)

4. Decoherence as an adversarial defence Model decoherence as a single-qubit depolarising channel applied to every γ P qubit after each of the L entangling layers, Λγ (ρ) = (1 − γ)ρ + 3 a σa ρσa . Theorem 1 (Decoherence contraction). Let ra = Tr(ρ σa ) be the expectation of a Pauli string of weight w. After L layers each followed by Λγ on every qubit, , (6) ra(L) = η wL ra(0) , η ≡ 1 − 4γ 3 while the identity component (w = 0) is preserved. Proof. On the single-qubit Pauli as ⟨σa ⟩ 7→ η⟨σa ⟩ for a ∈ P frame, Λγ acts 1 1 {x, y, z} and ⟨I⟩ 7→ ⟨I⟩, since 3 b σb σa σb = − 3 σa gives the multiplier (1 − γ) − γ3 = η. A weight-w string is acted on by w independent channels per layer, contributing η w per layer; over L layers, η wL . Entangling unitaries are trace-preserving frame permutations. The data-dependent part of ⟨O⟩(x) is thus scaled by η L while a constant offset survives. Proposition 1 (Adversarial gradient shrinkage). Let f (x) = σ s(⟨O⟩(x) −  θ) . Writing ⟨O⟩γ (x) = η L ⟨O⟩0 (x) + c with c independent of x, the firstorder change of the decision logit under an input perturbation δ contracts as ∆logitγ (δ) = η L ∆logit0 (δ), η = 1 − 4γ/3. Proof. The logit is s(η L ⟨O⟩0 (x) + c − θ), with differential s η L ∇x ⟨O⟩0 (x); the constant c does not contribute. With L = 3, η L = 0.512 at γ = 0.15 and 0.216 at γ = 0.30: the adversarial leverage at γ = 0.30 is roughly one fifth of the noiseless value. The same 7

contraction erodes the clean margin, so an optimal γ balances separability against gradient shrinkage. This replaces the 3–100× conservative bound of [2] with a predictive scaling law. A caveat must be stated, and it shapes the empirical analysis below. Proposition 1 concerns the magnitude of the logit response, but sign-based ℓ∞ attacks (FGSM/PGD) move along sign(∇x L), and the positive factor η L cancels under the sign: the perturbation direction is unchanged by γ. Inference-time gradient contraction therefore cannot, on its own, explain robustness to sign-based attacks; it is operative for ℓ2 attacks, where magnitude matters. As Section 8.2 establishes across seeds and under strong attacks, the robustness measured here arises predominantly from the training-time effect of the channel — a noise-reshaped decision boundary — rather than from attack-time gradient shrinkage. Crucially this means the stochasticity must be present at training; a channel applied only at inference does not, by itself, harden a model trained without it. 5. Quantum dropout and the magnitude of noise regularisation The study [2] found that the depolarising channel does not regularise like dropout: it perturbs the measured expectation. It proposed per-gate deactivation as the fix. We formalise that mechanism, derive the penalty it implements, and show that the penalty is small at NISQ scale — which, as Section 8 confirms across 30 seeds, is exactly what is observed. Definition 1 (Per-gate quantum dropout). For each variational rotation RY (θl,i ), draw an independent Bernoulli mask ml,i ∼ Bern(1 − p) at every training forward pass and replace the gate by RY (ml,i θl,i ) (the identity when ml,i = 0); data-encoding rotations are never masked. Inference uses the full circuit (m ≡ 1). Proposition 2 (Adaptive-penalty form of per-gate dropout). Let θ̃l,i = ml,i θl,i with ml,i ∼ Bern(1−p) independent, so E[θ̃] = (1−p)θ and Var(θ̃l,i ) = 2 p(1 − p)θl,i . To second order in the mask fluctuation,    p(1 − p) X 2 ∂ 2 L θl,i 2 . Em L(θ̃) ≈ L (1 − p)θ + 2 ∂θl,i l,i

(7)

Per-gate dropout therefore augments the loss with a curvature-weighted (adaptive L2 ) penalty in weight space, with coefficient p(1 − p) maximised at p = 1/2. By contrast the depolarising channel contracts the read-out by 8

η L (Theorem 1) and adds variance to the measured output, augmenting the loss with an output-space (label-smoothing-type) penalty. Proof. Expand L(θ̃) around θ̄ = E[θ̃] and take the expectation; the firstorder term vanishes and the Hessian’s diagonal survives because the masks 2 are independent, E[(θ̃l,i − θ̄l,i )(θ̃l′ ,i′ − θ̄l′ ,i′ )] = δ p(1 − p)θl,i , giving (7). Equation (7) is the resolution of the puzzle. It makes two falsifiable predictions. First, magnitude: the penalty scales as p(1 − p) θ2 H; at NISQ scale the variational angles are small (initialisation std 0.3) and the curvature H modest, so the induced gap reduction is of order 10−2 — a real but small effect, well below the seed-to-seed variance (≈ 0.03), which is why single-seed estimates are unstable and can appear to favour either noise type. Second, optimal rate: because p(1−p) ≤ 1/4 peaks at p = 1/2 and decreases for larger p, increasing the dropout rate past one half cannot strengthen regularisation and only deepens the underfitting induced by L((1 − p)θ). Both predictions are tested in Section 8; both hold. The depolarising channel’s output-space penalty is of comparable, equally small magnitude at this scale, so the two mechanisms are expected to regularise by similar small amounts — which the 30-seed study confirms. 6. Neutral-atom realisation The Hamiltonian (4) is the effective Hamiltonian of a Rydberg atom array. For atoms at {xi } under a global drive of Rabi frequency Ω and detuning δ, HRyd =

X X C6 ℏΩ X (i) σx − ℏδ ni nj . ni + 2 i ∥xi − xj ∥6 i i<j

(8)

The van-der-Waals term is an Ising σz σz coupling whose range is set by geometry, so a ring of atoms instantiates Ering , and the Rydberg blockade is the physical origin of the entanglement modelled by Jij [31, 32, 33]. On neutralatom hardware the dominant decoherence is σz -dephasing, which — being aligned with the Z read-out — lies in the kernel of the contraction of Theorem 1 and preserves the clean signal, shifting the trade-off relative to isotropic depolarising noise. We realise an SQNN reservoir on this model: a ring register with per-seed geometric jitter, features encoded into a global pulse schedule, and read-out from single-atom excitations and blockade-induced

9

Figure 1: Neutral-atom SQNN reservoir on the XOR task: accuracy distribution at N = 8 (100 seeds) and N = 10. The reservoir improves with register size, as predicted by the higher-dimensional quantum feature map at larger N .

pair correlations with a trained linear head [35]. On the XOR task the noiseless reservoir reaches 97.4 ± 2.7% at N = 8 (100 seeds) and 98.3 ± 3.0% at N = 10 (57 seeds; Fig. 1); exact channel simulation is feasible to N ≈ 12 (Table 1), with larger registers requiring tensor-network or trajectory surrogates subject to a bond-dimension caveat. 7. Methodology Dataset.. NSL-KDD records carry 41 features (three categorical) and a label, cast as binary (normal vs. any attack). Categorical features are integerencoded, all features standardised, and the representation reduced by PCA to d = 4 components so quantum and classical models share an identical low-dimensional space in which ℓ∞ perturbations are comparable. The robustness study uses 490 training / 210 test records (attack rate 0.46) over three seeds; the regularisation study uses subsampled training sets N ∈ {40, 80, 160} over thirty seeds. The pipeline is applied to a second real dataset, UNSW-NB15 [29], in Section 8.3; scaling to the full corpora and to CIC-DDoS2019 [30] with mutual-information feature selection is left to a hardware-scale evaluation. Models.. The SQNN has N = d = 4 qubit-neurons, a CNOT ring, and a σz read-out with a trained logistic head, on a density-matrix backend so the depolarising channel of Section 4 is applied exactly at training and inference; 10

γ = 0 is the noiseless VQC. Per-gate dropout (Definition 1) uses p = 0.3 at depth 3, and p ∈ {0.5, 0.7} at depth 5 for the rate sweep. Classical baselines (random forest, MLP, RBF-SVM) use the same features. Training is by Adam [38] on the binary cross-entropy with analytic gradients. Attacks.. White-box FGSM and 5-step PGD (step ε/4) in the shared feature space use the model’s exact input gradient for the quantum classifiers. Tree ensembles and the MLP expose no clean white-box gradient on tree splits; following standard practice we attack them by transfer from a logistic surrogate, a conservative (weaker) attack that understates their vulnerability. 8. Results 8.1. Adversarial robustness Table 2 and Fig. 2 report clean and adversarial accuracy on NSL-KDD. Classical neural detectors are accurate but brittle: the MLP falls from 94.8% to 47.1% and the random forest from 92.9% to 39.5% under FGSM at ε = 0.3, despite the weaker transfer attack. The noiseless VQC degrades from 89.5% to 79.0% under PGD. The SQNN with γ = 0.30 retains 85.2% under PGD — the best robust accuracy at the strongest budget — at the cost of about one clean point. This single-seed table illustrates the per-budget behaviour; because few-seed estimates are unreliable (a point this paper makes elsewhere), the rigorous evaluation — seven seeds, a stronger PGD attack, and an ℓ2 attack — is given in Section 8.2, where we also show that the operative mechanism is not the inference-time gradient contraction of Proposition 1 but a training-time reshaping of the decision boundary. 8.2. Robustness under strong and ℓ2 attacks (seven seeds) Because Table 2 is single-seed, we re-evaluate the γ = 0 versus γ = 0.30 contrast over seven seeds, under a stronger attack (PGD with 20 steps and exact gradients — the appropriate adaptive test for a density-matrix model, whose gradients are not obfuscated by shot noise) and under an ℓ2 attack, where the gradient magnitude contracted by Theorem 1 actually matters. Table 3 and Fig. 3 report the result. Under the strong ℓ∞ attack the γ = 0.30 model is significantly more robust (paired t-test p = 0.044, Cohen’s d = 0.96, favoured on 6/7 seeds); under the ℓ2 attack the advantage has comparable effect size but does not reach significance at this sample (d = 0.82, p = 0.07). Robustness therefore survives a strong, exact-gradient attack, which rules 11

Table 2: Clean and adversarial accuracy (%) on NSL-KDD (d = 4, seed 0). FGSM/PGD at ℓ∞ budget ε; classical models attacked by transfer (weaker). The noisy SQNN is the most robust under the strongest attack; classical neural detectors collapse.

Model

clean

FGSM ε=0.1

FGSM ε=0.3

PGD ε=0.3

SQNN γ=0 (VQC) SQNN γ=0.05 SQNN γ=0.15 SQNN γ=0.30

89.5 90.5 91.9 89.0

87.6 85.7 91.4 85.7

80.0 73.3 81.0 84.3

79.0 78.1 81.4 85.2

Random forest MLP RBF-SVM

92.9 94.8 91.9

71.0 82.4 91.0

39.5 47.1 89.0

— — —

out gradient masking as its source. The most reliable effect, however, is stability: the noiseless model suffers catastrophic robustness collapse (below 50% under PGD-20) on 2/7 seeds, falling as low as 20%, whereas the γ = 0.30 model never collapses and its across-seed standard deviation is 2.3× smaller (0.12 versus 0.27). These observations fix the mechanism. Since sign(η L ∇) = sign(∇), the inference-time gradient contraction of Proposition 1 cannot explain the ℓ∞ robustness; the advantage comes from training with the channel, which reshapes the learned boundary into a flatter, less attackable one — a property of the trained model, which is why it persists under the strongest attack and why a channel applied only at inference would not reproduce it. For a deployment that samples the channel at inference (shot-based or on hardware), a defence-aware attacker should additionally be evaluated with Expectationover-Transformation; we leave this to a hardware study (Section 10). 8.3. A second dataset and defended classical baselines (UNSW-NB15) Two questions remain: does the training-time robustness effect generalise beyond NSL-KDD, and how does the SQNN compare to defended classical models under a white-box attack (removing the transfer-attack handicap of Table 2)? We repeat the pipeline on UNSW-NB15 [29] (identical d = 4 PCA preprocessing, a 700-record stratified subsample, three seeds) and compare the SQNN (γ = 0, 0.30) against three classical baselines, each attacked whitebox with exact input gradients: a vanilla MLP, an adversarially-trained MLP (Madry PGD [6]), and a Gaussian-noise-injection MLP. Table 4 and the 12

Figure 2: Accuracy versus FGSM budget ε on NSL-KDD. Increasing the SQNN decoherence γ flattens the degradation (Proposition 1); gradient-trained classical models collapse. Table 3: Robustness over seven seeds: noiseless VQC (γ = 0) versus γ = 0.30, under strong attacks with exact gradients. Mean ± 95% CI; paired t-test and Cohen’s d for the γ=0.30 minus γ=0 difference. The strong ℓ∞ advantage is significant; the noiseless model collapses catastrophically on 2/7 seeds, γ = 0.30 on 0/7.

Attack

γ = 0 (VQC)

γ = 0.30

Clean 0.941 ± 0.013 0.918 ± 0.012 PGD-5 ℓ∞ (ε=0.3) 0.690 ± 0.162 0.799 ± 0.086 PGD-20 ℓ∞ (ε=0.3) 0.624 ± 0.197 0.790 ± 0.086 PGD-20 ℓ2 (ε2 =0.5) 0.598 ± 0.223 0.688 ± 0.177

paired p

Cohen d

0.067 0.052 0.044 0.073

−0.84 0.91 0.96 0.82

accuracy–robustness frontier of Fig. 4 report the result. Two findings stand out, and we state the second plainly. First, the training-time effect generalises: on UNSW-NB15 the γ = 0.30 model is again more robust than the noiseless circuit (PGD-20 ℓ∞ 0.452 versus 0.409; ℓ2 0.387 versus 0.318), confirming that the boundary-reshaping mechanism of Section 8.2 is not specific to one dataset. Second, the SQNN does not beat strong classical defences: adversarial training is the most robust model (0.706 under ℓ∞ ), and even a white-box-attacked vanilla MLP (0.553) exceeds the SQNN here. This is consistent with our stated position that no quantum advantage is claimed (Section 1); the SQNN’s contribution is the contraction theory and the training-time mechanism, not robustness supremacy over clas13

Figure 3: Robustness over seven seeds under strong, exact-gradient attacks (left: ℓ∞ PGD20; right: ℓ2 PGD-20). Grey lines link paired seeds; markers are means with 95% CI. The γ = 0.30 model is more robust on 6/7 seeds and, unlike the noiseless VQC, never collapses catastrophically.

sical defences. The frontier (Fig. 4) places adversarial training at the robust extreme and the SQNN inside it, with γ = 0.30 strictly improving on γ = 0. Robustness is moreover dataset-dependent: the SQNN γ = 0.30 that was competitive on NSL-KDD (0.79 under ℓ∞ , Section 8.2) is weaker on UNSWNB15 (0.45), a caveat that any operational claim must respect. 8.4. Noise regularisation: a small but significant effect (30 seeds) Table 5 and Fig. 5 report the train–test gap over thirty seeds. Both noise mechanisms reduce the gap relative to the noiseless VQC by a small but statistically significant margin: pooled across sizes (n = 90), the depolarising channel lowers the gap by 0.009 (paired t-test p < 10−4 ) and per-gate dropout by 0.010 (p = 0.004). The two are statistically indistinguishable from each other (p = 0.74, Fig. 5). The effect is concentrated where overfitting is largest — significant at N = 40 (Depol p = 0.001, dropout p = 0.013), partial at N = 80, and vanishing at N = 160 where there is little gap to close. The magnitude (≈ 0.01) and its smallness are exactly the prediction of Proposition 2: at this scale p(1 − p)θ2 H is of order 10−2 , below the seed variance (≈ 0.03). This is why the single-seed dichotomy reported in preliminary work does not survive: with one seed the sign of the difference is essentially noise. Thirty seeds are required to resolve a real effect of this size.

14

Table 4: Second dataset (UNSW-NB15, d = 4 PCA, three seeds): SQNN versus defended classical baselines, all attacked white-box with exact gradients (PGD-20). Mean ± std. The training-time effect generalises (γ=0.30 > γ=0), but classical adversarial training is the most robust model; no quantum advantage is claimed.

Model SQNN γ = 0 SQNN γ = 0.30 MLP (white-box) MLP adv-trained MLP+Gauss

clean

PGD-20 ℓ∞

0.872 ± 0.027 0.409 ± 0.163 0.793 ± 0.048 0.452 ± 0.195 0.878 ± 0.031 0.553 ± 0.120 0.820 ± 0.060 0.706 ± 0.091 0.869 ± 0.032 0.573 ± 0.107

PGD-20 ℓ2 0.318 ± 0.123 0.387 ± 0.179 0.583 ± 0.130 0.670 ± 0.070 0.605 ± 0.101

Table 5: Train–test accuracy gap (mean ± std over 30 seeds; positive = overfitting) on NSL-KDD. Both noise mechanisms reduce the gap; they are statistically equivalent (pooled p = 0.74).

Training size N

VQC (no noise)

Depolarising

Quantum dropout

40 80 160

+0.087 ± 0.030 +0.049 ± 0.028 +0.019 ± 0.018

+0.072 ± 0.026 +0.039 ± 0.028 +0.016 ± 0.018

+0.070 ± 0.034 +0.041 ± 0.030 +0.013 ± 0.034

0.0515

0.0424 (p<10−4 )

0.0414 (p=0.004)

pooled mean (n=90)

8.5. Increasing the dropout rate does not help (depth 5) Proposition 2 predicts that the regularisation strength peaks at p = 1/2. We test this at depth 5 (more variational parameters, hence more overfitting), N = 40, 10 seeds, tracking train and test accuracy (Fig. 6). Raising the dropout rate from 0.5 to 0.7 does not reduce the gap (it stays at ≈ 0.08, indistinguishable from the VQC’s 0.088); instead it lowers both train and test accuracy together (test 0.907 → 0.872 → 0.848). Strong dropout thus degrades the model rather than regularising it, precisely because p(1 − p) decreases past p = 1/2 while the mean contraction (1 − p)θ deepens. The depolarising channel retains the lowest gap (0.067) while keeping test accuracy high (0.908), consistent with its output-space mechanism. 8.6. Decoherence-resolved robustness Fig. 7 isolates the channel dependence predicted by Theorem 1: under pure dephasing — the D[σz ] dissipator aligned with the read-out — accuracy 15

Figure 4: Accuracy–robustness frontier on UNSW-NB15 (three seeds; left ℓ∞ , right ℓ2 ), all models attacked white-box with exact gradients. Adversarial training occupies the robust extreme; the SQNN lies inside the frontier, with γ = 0.30 improving on γ = 0. No quantum advantage over classical defences is claimed.

is preserved up to large noise, whereas under isotropic depolarising noise it collapses past a critical strength. The separation is the channel-resolved content of the contraction law: σz -aligned decoherence sets no threshold, isotropic decoherence a depth-dependent one. 9. Discussion The results cohere around the contraction law and the penalty formula, with an important correction to the naive reading of the former. Theorem 1 contracts the data-dependent read-out by η L , yet for sign-based ℓ∞ attacks this contraction cancels under the sign (Section 4), so it is not the source of the measured robustness. The seven-seed strong-attack study (Section 8.2) shows instead that training with the channel reshapes the decision boundary into a flatter, less attackable one: the γ = 0.30 model is significantly more robust under exact-gradient PGD-20 and, most reliably, never suffers the catastrophic collapse that befalls the noiseless model on a sizeable fraction of seeds. This reframes the noise term of [2] from an empirical knob into a training-time regulariser of the boundary, and clarifies that hardening requires the stochasticity to be present at training, not merely at inference. For generalisation, Proposition 2 shows that per-gate dropout and depolarising noise regularise in different spaces (weight versus output), but that 16

Figure 5: Generalisation gap over 30 seeds. Left: gap by training size and mode; effects are small and shrink as N grows. Right: pooled mean gap with 95% CI — both depolarising noise and per-gate dropout reduce the gap significantly and are indistinguishable from each other.

both penalties are small at NISQ scale; the 30-seed study confirms a real, significant, and mutually equivalent gap reduction of ≈ 0.01, and explains why prior single-seed claims were unstable. The honest reading is not that one noise type “wins”, but that noise regularisation of SQNNs is a small, quantifiable effect whose magnitude and optimal rate the formula predicts. For collaborative intrusion detection the implications are concrete: the ring closure carries the non-local signal at minimal circuit cost [2]; a depolarising channel present at training buys stable robustness with a sub-twopoint clean cost; and the variance reduction it induces makes a defended detector behave predictably across deployments. Looking ahead, the natural and highest-stakes extension is to behavioural detection in tool-using-agent pipelines, where evasion is a system-level compromise and where, by the analogy of Section 2, instruction-level manipulations such as prompt injection are an adversarial problem cousin to the one studied here. Extending SQNN anomaly detection to agent execution traces — with its own dataset of traces, protocol, and baselines — is left as future work; we make no agentlevel claim from the present experiments. 10. Limitations and threats to validity (i) A single simulation environment and a d = 4 PCA feature space; multi-seed sweeps at larger d and N are the next step, and the ℓ2 robustness advantage (Table 3), though of large effect size, is not yet significant at seven seeds. (ii) Non-differentiable classical models are attacked by transfer, 17

Figure 6: Cranking per-gate dropout (depth 5, N = 40, 10 seeds): increasing p lowers train and test accuracy together, leaving the gap unchanged — the regularisation does not strengthen past p = 1/2, as Proposition 2 predicts.

understating their vulnerability; white-box attacks would widen the gap to the SQNN. Because the simulated channel is deterministic (density matrix), the model’s gradients are exact and not obfuscated, so the strong PGD-20 attack is a valid adaptive test here; a shot-based or hardware deployment that samples the channel at inference should additionally be stress-tested with Expectation-over-Transformation. (iii) The depolarising channel is an idealisation; on neutral-atom hardware dephasing dominates, preserving the clean signal and shifting the trade-off favourably but requiring on-device validation. (iv) Exact density-matrix simulation is feasible only to N ≈ 12 (Table 1); larger networks need trajectory or tensor-network surrogates with a bond-dimension caveat. (v) We evaluate on two real datasets (NSLKDD and UNSW-NB15); validation on the full corpora, on CIC-DDoS2019, and at larger d and N remains for future work. No quantum speed-up or accuracy advantage over classical baselines is claimed — indeed, on UNSWNB15 classical adversarial training is the most robust model (Section 8.3). 11. Conclusion We have given an N -qubit theory of stochastic quantum neural networks and used it to resolve two questions left open by the preceding collaborativedetection study. A decoherence-contraction theorem gives the law (1 − 4γ/3)wL for how the channel scales Pauli read-outs; we show this contraction does not, by itself, explain ℓ∞ robustness (the attack’s sign cancels it), and a 18

Figure 7: Channel-resolved robustness: dephasing (D[σz ]) preserves the σz -basis signal, while isotropic depolarising noise contracts all Pauli components and collapses accuracy past a threshold, consistent with Theorem 1.

seven-seed strong-attack study locates the effect instead in a training-time reshaping of the boundary: the depolarising SQNN is significantly more robust under exact-gradient PGD-20 and never suffers the catastrophic collapse that the noiseless model and classical detectors do. An adaptive-penalty formula shows that per-gate dropout implements a curvature-weighted weight-space penalty maximised at p = 1/2, and a thirty-seed study confirms its predictions: both depolarising noise and per-gate dropout produce a small (≈ 0.01) but significant and mutually equivalent reduction of the generalisation gap, increasing the dropout rate past one half does not help, and the single-seed dichotomy of prior work does not replicate. Together with a neutral-atom realisation and a feasibility-by-N characterisation, these results turn the SQNN into a quantitatively grounded, honestly bounded architecture for adversarially robust intrusion detection. CRediT authorship contribution statement Gautier-Edouard Filardo: Conceptualization; Methodology; Software; Formal analysis; Investigation; Writing – original draft; Writing – review & editing.

19

Declaration of competing interest The author declares no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. Data availability NSL-KDD is publicly available [28]. The source code reproducing all experiments—density-matrix SQNN, depolarising channel, per-gate quantum dropout, FGSM/PGD attacks, the depth-5 dropout-rate sweep, and the neutral-atom reservoir—is publicly available at https://github.com/ gautierfilardo-efrei/sqnn-adversarial and archived at https://doi. org/10.5281/zenodo.20786870. References [1] G.-E. Filardo, T. Heckmann, Stochastic quantum neural networks for neuroinspired intelligence: mathematical foundations, comparative benchmarks, and prospects, Neurocomputing 663 (2026) 132031. [2] G.-E. Filardo, Stochastic quantum neural networks for collaborative anomaly detection in distributed enterprise systems, in: Proc. IEEE WETICE, 2026. [3] D.C. Nguyen, M. Ding, P.N. Pathirana, A. Seneviratne, J. Li, H.V. Poor, Federated learning for internet of things: a comprehensive survey, IEEE Commun. Surv. Tutor. 23 (3) (2021) 1622–1658. [4] Z. Ahmad, A. Shahid Khan, C. Wai Shiang, J. Abdullah, F. Ahmad, Network intrusion detection system: a systematic study of machine learning and deep learning approaches, Trans. Emerg. Telecommun. Technol. 32 (1) (2021) e4150. [5] I.J. Goodfellow, J. Shlens, C. Szegedy, Explaining and harnessing adversarial examples, in: Proc. ICLR, 2015. [6] A. Madry, A. Makelov, L. Schmidt, D. Tsipras, A. Vladu, Towards deep learning models resistant to adversarial attacks, in: Proc. ICLR, 2018. [7] J. Biamonte, P. Wittek, N. Pancotti, P. Rebentrost, N. Wiebe, S. Lloyd, Quantum machine learning, Nature 549 (7671) (2017) 195–202.

20

[8] V. Havlı́ček, A.D. Córcoles, K. Temme, A.W. Harrow, A. Kandala, J.M. Chow, J.M. Gambetta, Supervised learning with quantum-enhanced feature spaces, Nature 567 (7747) (2019) 209–212. [9] J. Preskill, Quantum computing in the NISQ era and beyond, Quantum 2 (2018) 79. [10] M. Cerezo, A. Arrasmith, R. Babbush, S.C. Benjamin, S. Endo, K. Fujii, J.R. McClean, K. Mitarai, X. Yuan, L. Cincio, P.J. Coles, Variational quantum algorithms, Nat. Rev. Phys. 3 (9) (2021) 625–644. [11] K. Mitarai, M. Negoro, M. Kitagawa, K. Fujii, Quantum circuit learning, Phys. Rev. A 98 (3) (2018) 032309. [12] M. Benedetti, E. Lloyd, S. Sack, M. Fiorentini, Parameterized quantum circuits as machine learning models, Quantum Sci. Technol. 4 (4) (2019) 043001. [13] M. Schuld, V. Bergholm, C. Gogolin, J. Izaac, N. Killoran, Evaluating analytic gradients on quantum hardware, Phys. Rev. A 99 (3) (2019) 032331. [14] J.R. McClean, S. Boixo, V.N. Smelyanskiy, R. Babbush, H. Neven, Barren plateaus in quantum neural network training landscapes, Nat. Commun. 9 (1) (2018) 4812. [15] S. Wang, E. Fontana, M. Cerezo, K. Sharma, A. Sone, L. Cincio, P.J. Coles, Noise-induced barren plateaus in variational quantum algorithms, Nat. Commun. 12 (1) (2021) 6961. [16] K. Beer, D. Bondarenko, T. Farrelly, T.J. Osborne, R. Salzmann, D. Scheiermann, R. Wolf, Training deep quantum neural networks, Nat. Commun. 11 (1) (2020) 808. [17] I. Cong, S. Choi, M.D. Lukin, Quantum convolutional neural networks, Nat. Phys. 15 (12) (2019) 1273–1278. [18] A. Abbas, D. Sutter, C. Zoufal, A. Lucchi, A. Figalli, S. Woerner, The power of quantum neural networks, Nat. Comput. Sci. 1 (6) (2021) 403–409. [19] C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, R. Fergus, Intriguing properties of neural networks, in: Proc. ICLR, 2014. [20] N. Papernot, P. McDaniel, S. Jha, M. Fredrikson, Z.B. Celik, A. Swami, The limitations of deep learning in adversarial settings, in: Proc. IEEE EuroS&P, 2016, pp. 372–387.

21

[21] A. Kurakin, I.J. Goodfellow, S. Bengio, Adversarial examples in the physical world, in: Proc. ICLR Workshop, 2017. [22] J. Cohen, E. Rosenfeld, J.Z. Kolter, Certified adversarial robustness via randomized smoothing, in: Proc. ICML, 2019, pp. 1310–1320. [23] S. Lu, L.-M. Duan, D.-L. Deng, Quantum adversarial machine learning, Phys. Rev. Res. 2 (3) (2020) 033212. [24] N. Liu, P. Wittek, Vulnerability of quantum classifiers to adversarial perturbations, Phys. Rev. A 101 (6) (2020) 062331. [25] Y. Du, M.-H. Hsieh, T. Liu, D. Tao, N. Liu, Quantum noise protects quantum classifiers against adversaries, Phys. Rev. Res. 3 (2) (2021) 023153. [26] W. Gong, D.-L. Deng, Universal adversarial examples and perturbations for quantum classifiers, Natl. Sci. Rev. 9 (6) (2022) nwab130. [27] M.T. West, S.M. Erfani, C. Leckie, M. Sevior, L.C.L. Hollenberg, M. Usman, Towards quantum enhanced adversarial robustness in machine learning, Nat. Mach. Intell. 5 (6) (2023) 581–589. [28] M. Tavallaee, E. Bagheri, W. Lu, A.A. Ghorbani, A detailed analysis of the KDD CUP 99 data set, in: Proc. IEEE CISDA, 2009, pp. 1–6. [29] N. Moustafa, J. Slay, UNSW-NB15: a comprehensive data set for network intrusion detection systems, in: Proc. MilCIS, 2015, pp. 1–6. [30] I. Sharafaldin, A.H. Lashkari, S. Hakak, A.A. Ghorbani, Developing realistic distributed denial of service (DDoS) attack dataset and taxonomy, in: Proc. IEEE ICCST, 2019, pp. 1–8. [31] L. Henriet, L. Beguin, A. Signoles, T. Lahaye, A. Browaeys, G.-O. Reymond, C. Jurczak, Quantum computing with neutral atoms, Quantum 4 (2020) 327. [32] A. Browaeys, T. Lahaye, Many-body physics with individually controlled Rydberg atoms, Nat. Phys. 16 (2) (2020) 132–142. [33] H. Bernien, S. Schwartz, A. Keesling, H. Levine, A. Omran, H. Pichler, et al., Probing many-body dynamics on a 51-atom quantum simulator, Nature 551 (7682) (2017) 579–584. [34] D. Bluvstein, H. Levine, G. Semeghini, T.T. Wang, S. Ebadi, M. Kalinowski, et al., A quantum processor based on coherent transport of entangled atom arrays, Nature 604 (2022) 451–456.

22

[35] H. Silvério, S. Grijalva, C. Dalyac, L. Leclerc, P.J. Karalekas, N. Shammah, M. Beji, L.-P. Henry, L. Henriet, Pulser: an open-source package for the design of pulse sequences in programmable neutral-atom arrays, Quantum 6 (2022) 629. [36] N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, R. Salakhutdinov, Dropout: a simple way to prevent neural networks from overfitting, J. Mach. Learn. Res. 15 (1) (2014) 1929–1958. [37] Y. Gal, Z. Ghahramani, Dropout as a Bayesian approximation: representing model uncertainty in deep learning, in: Proc. ICML, 2016, pp. 1050–1059. [38] D.P. Kingma, J. Ba, Adam: a method for stochastic optimization, in: Proc. ICLR, 2015. [39] G. Lindblad, On the generators of quantum dynamical semigroups, Commun. Math. Phys. 48 (2) (1976) 119–130. [40] V. Gorini, A. Kossakowski, E.C.G. Sudarshan, Completely positive dynamical semigroups of N-level systems, J. Math. Phys. 17 (5) (1976) 821–825. [41] H.-P. Breuer, F. Petruccione, The Theory of Open Quantum Systems, Oxford University Press, 2002. [42] C.W. Gardiner, P. Zoller, Quantum Noise, 3rd ed., Springer, 2004.

23

Record · ID 303155 · SHA-256 af41d2a838ea509e
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.