E NTANGLEMENT AS A S TRUCTURAL C OMPLEXITY A XIS : A PAC-BAYESIAN V IEW OF G ENERALIZATION IN Q UANTUM P OLICIES AND VALUE F UNCTIONS Jian Xu1,2 , Delu Zeng3 , John Paisley4 , Qibin Zhao2 1 RIKEN iTHEMS 2 RIKEN AIP 4 Columbia University 3 South China University of Technology [email protected]
arXiv:2607.06230v1 [quant-ph] 7 Jul 2026
A BSTRACT Parameterized quantum circuits (PQCs) are increasingly used as policies and value functions in quantum reinforcement learning (QRL), yet almost all evaluations report only average return, leaving open the question of when and why a quantum policy generalizes. We give a PAC-Bayesian answer. We derive a generalization bound for stochastic quantum policies whose complexity term is controlled not by the raw number of circuit parameters, but by the effective dimension of the Fisher geometry induced by the circuit—a quantity that entanglement inflates. Empirically, in a controlled setting that fixes the number of trainable rotations and varies only the entangling connectivity, we find that entanglement is an independent axis of complexity: at a fixed parameter count the train-to-test generalization gap grows with the circuit’s Fisher effective dimension—which entanglement inflates—whereas raw parameter count is the weakest predictor of the gap. The bound is primarily a ranking certificate: it correctly orders circuits of identical parameter count, something a parameter-counting bound cannot do at all. We confirm this mechanism across settings—supervised classification (at 4– 16 qubits, on synthetic and on real Iris/Breast-Cancer/Wine data), a reward-only quantum contextual bandit, and multi-step value-function generalization—where entangled circuits generalize worse than non-entangled ones of identical parameter count and the gap shrinks with sample size as the bound predicts. Our strongest evidence is in these low-variance decision models (single-observable classifiers and value heads, and a one-step policy); in genuine end-to-end multi-step policy learning the standard entangler’s effect is statistically significant but return variance leaves the full ordering only partially resolved, so we frame the paper as a controlled study of the entanglement–generalization trade-off in quantum decision models rather than a solved account of end-to-end quantum RL. A partialcorrelation analysis shows the Fisher effective dimension screens off the entangling pattern (which adds ∆R2 < 0.01 once deff is controlled), controls (matched training accuracy, alternative readout and optimizer) rule out optimization confounders, and the effect survives execution on an IBM Heron quantum processor under real noise. Our results reframe the design of quantum policies around an entanglement–generalization trade-off rather than around expressivity alone.
1
I NTRODUCTION
Quantum reinforcement learning (QRL) replaces the neural policy or value network of a classical agent with a parameterized quantum circuit (PQC). Recent work has made such hybrid agents practical—PPO-Q integrates a PQC policy into proximal policy optimization and validates it on superconducting hardware (Jin et al., 2025), and adaptive non-local observables show that the measurement layer itself is a key design axis (Lin et al., 2025). Yet the field is almost entirely empirical: papers report the average return of a chosen ansatz, and a reviewer is left to wonder whether a reported gain reflects a genuine inductive bias or merely the implicit regularization of a small model. 1
We argue that the right question is not “does a quantum policy achieve higher return” but “when does a quantum policy generalize, and what property of the circuit controls it”. This is the question PACBayes theory is built to answer for stochastic predictors (McAllester, 1999; Dziugaite & Roy, 2017), and it has recently been brought to reinforcement learning, yielding non-vacuous certificates for modern off-policy agents (Zitouni et al., 2025). In parallel, a PAC-Bayesian analysis of supervised quantum models has appeared (Rodriguez-Grasa et al., 2026), whose complexity term is built from learned parameter norms and sparsity. To our knowledge no PAC-Bayes analysis addresses quantum policies, and none isolates entanglement as the governing complexity. This paper makes three contributions. • A PAC-Bayes bound for quantum policies (Section 3) whose complexity term is the effective dimension of the circuit-induced Fisher geometry rather than the parameter count. We show that under a data-dependent Gaussian posterior the KL complexity is governed by deff (θ), which for PQCs is inflated by entanglement. • An identification of entanglement as an independent complexity axis (Section 4). Fixing the number of trainable rotations and varying only the entangling connectivity, we show the generalization gap grows monotonically with the Meyer–Wallach entangling power at fixed parameter count, while parameter count itself is the weakest predictor of the gap. The axis is readout-gated: the mechanism is that entanglement enlarges the readout’s lightcone, so it is decisive under the local single-qubit readouts standard in QRL and is partly substituted by a global readout, which activates the same parameters directly (Section 7)— a control that directly confirms the light-cone account and shows deff , not the entangling label, is what tracks the gap. • Broad empirical validation: supervised classification at 4/6/8 qubits (Sections 4–5), a reward-only quantum contextual bandit and multi-step value-function generalization (Section 6), optimization controls, and a run on IBM Heron hardware (Section 7)—in all of which entangled circuits generalize worse than non-entangled ones of identical parameter count. The practical message is a design principle: entanglement is a double-edged sword. It enlarges the function space a quantum policy can represent, but it carries a generalization cost that a parametercounting analysis does not see and that our bound makes explicit. Scope. We are deliberate about what we claim. The theory is stated for policies and value functions, and our strongest evidence comes from the low-variance instantiations of exactly these objects: single-observable classifiers and value heads, a one-step (contextual-bandit) policy, and MonteCarlo evaluation of a multi-step value function. In genuine end-to-end multi-step policy learning (REINFORCE) the standard entangler’s effect is statistically significant (Section 6), but returnbased estimates are high enough variance that the full ordering is not always resolved; we therefore treat the low-variance settings as the primary probes. The paper is thus best read as a study of the entanglement–generalization trade-off in quantum decision models—policies and value functions— rather than a claim to have solved generalization in end-to-end quantum RL.
2
R ELATED W ORK
Quantum reinforcement learning. PQC agents have been trained with policy gradients, actor– critic and value-based methods (Jin et al., 2025; Luo et al., 2025; Wu et al., 2025), with recent attention to the measurement layer (Lin et al., 2025) and to data re-uploading and trainability (Coelho et al., 2024). These works optimize and report return; generalization is not analyzed. Generalization of quantum models. A now-substantial literature bounds the generalization of quantum models. Covering-number and Rademacher analyses give uniform bounds for quantum feature maps and PQCs (Banchi et al., 2021), and Caro et al. (2022) prove the landmark result that e ) training points—a bound driven by gate a PQC with T parameterized gates generalizes from O(T count, not entanglement. Gil-Fuster et al. (2024) caution that such uniform-convergence bounds can be simultaneously satisfied and uninformative for the models one actually trains, motivating modeldependent capacity measures. The Fisher-information “effective dimension” is one such measure, 2
proposed as a capacity that reflects the trained model rather than the worst case (Abbas et al., 2021). Our deff is this same log-determinant functional, but our contribution is orthogonal to Abbas et al. (2021) in two respects (Appendix B makes the comparison precise): (i) they use deff as a global, sample-size-dependent capacity for a model class, whereas we evaluate it at the trained parameters with a fixed γ and use it only to rank circuits of equal parameter count; and (ii) they do not connect deff to entanglement—our Proposition 1 attributes the inflation of the Fisher rank (and hence of deff ) to the entanglement-driven growth of the readout light-cone, which is the new mechanism. Most closely related on the PAC-Bayes side, Rodriguez-Grasa et al. (2026) give a PAC-Bayes bound for supervised quantum classifiers whose complexity is a parameter-norm/sparsity term; they report weak but positive correlations between their complexity and the observed gap. We differ in two ways: we treat policies (return-based generalization under Markov dependence) and we identify entanglement, not parameter norm or gate count, as the governing axis, isolating it at fixed parameter and gate count. PAC-Bayes for RL and control. PAC-Bayes control learns policies with certified generalization to novel environments (Majumdar et al., 2021), and recent work derives PAC-Bayes RL bounds that account for the mixing time of the induced Markov chain (Zitouni et al., 2025). We adopt this template and instantiate its complexity term for quantum policies.
3
A PAC-BAYES B OUND FOR Q UANTUM P OLICIES
Setup. An agent interacts with an MDP M drawn from a distribution D over environments. A stochastic quantum policy πθ (a | s) is defined by a PQC U (s, θ) acting on n qubits: the state s is encoded, the circuit is applied, and the action is read from a single readout observable ⟨Z0 ⟩s,θ (a binary policy πθ (a=1 | s) = σ(α⟨Z0 ⟩s,θ ); k-action policies use k observables and our analysis applies per observable). We use the single-observable readout throughout, both in the theory below and in every experiment. We write J(πθ ) = EM ∼D Eτ ∼πθ [R(τ )] for the expected return with ˆ θ ) for its empirical estimate on m training environments. We consider a R ∈ [0, Rmax ], and J(π Gaussian posterior Q = N (θ, σ 2 Id ) over the d circuit parameters and a prior P = N (θ0 , σ 2 Id ). Theorem 1 (PAC-Bayes bound for quantum policies). Under the mixing-time assumptions of Zitouni et al. (2025), for any δ ∈ (0, 1), with probability at least 1 − δ over the draw of m training environments, every posterior Q satisfies s √ KL(Q∥P ) + ln 2 δm ∥θ − θ0 ∥2 ˆ J(πQ ) ≥ J(πQ ) − Rmax , (1) , KL(Q∥P ) = 2 meff 2σ 2 where meff = m/κ discounts the sample size by the chain mixing factor κ. Equation equation 1 is the standard McAllester bound with the RL sample correction; its only quantum-specific object is KL(Q∥P ). Naively this depends only on the Euclidean displacement of the d parameters, suggesting that generalization is controlled by parameter count. Our central observation is that this is misleading once the posterior is allowed to be shaped by the local geometry the circuit induces on the loss. Environments, contexts, and the single-step case. The i.i.d. sample zi of Theorem 1 is an “environment” Mi ∼ D; the mixing factor κ enters only because a multi-step return pools Markovdependent transitions. In the single-step settings that carry our cleanest evidence—supervised classification and the contextual bandit—each “environment” is a single i.i.d. context xi ∼ D with a oneshot reward, the trajectory has length one, so there is no Markov dependence to discount and κ = 1, giving meff = m = N exactly. The bound we actually evaluate in these sections is therefore the plain McAllester form, and κ is relevant only to the genuine multi-step experiments (value-function regression and end-to-end REINFORCE, Section 6). We use “environment” and “context” interchangeably in the single-step case for this reason. The distribution shift the experiments measure is thus the in-distribution train→test gap over i.i.d. contexts, which is exactly the object Theorem 1 bounds when m contexts are drawn i.i.d.; the meta/multi-task reading (many environments) is the same inequality specialized to trajectory length one. 3
From an isotropic to a Fisher-shaped posterior. The bound equation 1 is loose because the isotropic posterior N (θ, σ 2 I) ignores the geometry of the loss. The curvature of a quantum policy is the Fisher information matrix F̂(θ) = Es [∇θ ℓs ∇θ ℓ⊤ s ], with eigenvalues λ1 ≥ · · · ≥ λd ≥ 0, and it is highly anisotropic. For the single-observable Bernoulli readout ps = σ(α⟨Z0 ⟩s,θ ) used throughout, this loss Fisher has the explicit closed form F̂(θ) = Es α2 ps (1−ps ) ∇θ fs ∇θ fs⊤ , with fs = ⟨Z0 ⟩s,θ —i.e. the loss Fisher of the logistic head equals the output-gradient Fisher weighted by the Bernoulli variance α2 ps (1−ps ). This is exactly the object we diagonalize in every experiment (Appendix B), so the F̂ appearing in the bound and the F̂ we measure are the same matrix, and its rank is governed by the readout light-cone (Proposition 1). Choosing a posterior that respects this curvature turns the KL term into a genuinely quantum quantity. Theorem 2 (Fisher-form bound). Let the prior be the data-independent isotropic Gaussian P = N (θ0 , τ 2 Id ) and the posterior the Gauss–Newton Gaussian Q = N (θ, Σ) with Σ = τ 2 (Id +γ F̂)−1 , γ > 0. Then KL(Q∥P ) =
d 1X ∥θ − θ0 ∥2 γλi 1 + log det I + γ F̂ − , d 2 2 2τ 2 1 + γλi | {z } i=1
(2)
=: deff (γ)
and hence, under the assumptions of Theorem 1, with probability ≥ 1 − δ, v u ∥θ − θ ∥2 √ 0 u 2 m 1 + d (γ) + ln t eff 2 δ 2τ 2 ˆ Q ) − Rmax J(πQ ) ≥ J(π . 2 meff
(3)
Equation equation 3 is the bound that matters: its complexity is the distance term plus the Fisher effective dimension d X deff (γ) = log det Id + γ F̂ = log(1 + γλi ), (4) i=1
a quantity a parameter count cannot see. The final term of equation 2 is subtracted and lies in [0, d/2]; dropping it therefore replaces KL by a larger upper surrogate, which conservatively loosens the gap bound while keeping it valid—we drop it only to expose deff as the operative complexity. Two facts make deff the right object: deff (γ) ≤ d always, and deff (γ) ≈ rank(F̂) log γ for large γ, so it counts active Fisher directions rather than nominal parameters. Crucially the prior is dataindependent (centred at the initialization θ0 ), so equation 3 is a valid PAC-Bayes certificate even P P though the posterior shape Σ is data-dependent through F̂; the participation ratio ( i λi )2 / i λ2i we also report is a stable empirical surrogate for the same spectrum, not a separate capacity notion. Equation equation 2 is a direct Gaussian KL computation, proved in Appendix A. Remark 1 (Choosing γ). The scale γ = βτ 2 is a free posterior hyperparameter and may be optimized. Since it must be fixed before seeing the data to keep the prior valid, one applies a union bound over a geometric grid γ ∈ {2j }Jj=1 , replacing δ by δ/J and paying only an additive 21 ln J inside the root; taking J = O(log m) and then the best grid point costs a negligible O(log log m) term and lets the data pick the γ that minimizes the bound. In every experiment in this paper we fix a single constant γ = 50—the same value for all circuits and sample sizes—so that deff is comparable across configurations; Appendix B gives the full Fisher computation. (A sample-size-dependent γ ∝ m, as in Abbas et al. (2021), would rescale deff per row and is not used here.) Proposition 1 (Entanglement raises the effective dimension). Fix the number and placement of the trainable single-qubit rotations and let the entangling connectivity vary. Let L(θ) ⊆ {1, . . . , d} be the set of parameters inside the backward light-cone of the readout observable (Def. 1). Then generically rank(F̂) = |L(θ)|, and |L(θ)|—and therefore deff (γ)—is non-decreasing under the addition of entangling gates, while the parameter count d is unchanged. With no entangling gates L = {i : q(i) = 0}, so deff is at its minimum; any connectivity that couples the readout wire to others strictly enlarges L and raises deff . Proposition 1 yields a sharp, falsifiable prediction we test at fixed parameter count: entangled circuits have a larger deff , hence a looser bound equation 3, hence a larger generalization gap, than non-entangled circuits with identical parameter count—something no parameter-counting bound 4
no entangling gates
readout lightcone = {q0 }
active dirs ∼ d/n (deff small)
tight bound small gap
add entangling gates
light-cone spreads to all n wires
Fisher rank ↑ (deff large)
looser bound large gap
Figure 1: The mechanism. At a fixed parameter count, entangling gates enlarge the backward lightcone of the readout observable (Lemma 2), which raises the rank and hence the effective dimension deff = log det(I +γ F̂) of the Fisher geometry (Proposition 1). A larger deff enlarges the complexity term of the PAC-Bayes bound equation 3 and thus the generalization gap—an effect invisible to a parameter-counting analysis, since d is unchanged.
can express. Its rigorous proof (a light-cone lemma for the rank monotonicity, plus a genericity assumption for the rank identity) is in Appendix A. The prediction holds in the strong monotone form (none < linear < full in both deff and gap) in supervised classification (Sections 4–5) and, at the smaller sample sizes, in the RL bandit, where each entangler also beats none by more than three standard errors (Section 6). Figure 1 summarizes the mechanism.
4
E NTANGLEMENT IS AN I NDEPENDENT A XIS OF C OMPLEXITY
Design. To separate entanglement from parameter count we use an ansatz with a fixed layout of trainable rotations—L layers of RY , RZ on each of n=4 qubits, so d = 2nL parameters—and vary only the entangling connectivity inserted after each rotation layer: none (no CNOTs), linear (a CNOT chain), and full (all-to-all CNOTs). All three share exactly the same d; they differ only in entangling power. The binary target y = sign(x1 x2 − x3 x4 + 12 x1 ) requires feature interactions, inputs are angle-encoded, the classifier is the single-observable readout ŷ = sign(⟨Z0 ⟩), and we sweep the training-set size N and evaluate on 2000 held-out inputs, averaging 8 seeds per cell. Result: entangled circuits genuinely generalize worse, not just memorize more. A key check is that the models actually generalize—otherwise a larger gap would only mean more memorization. Table 1 isolates depth L=2 (d=16) at two sample sizes. At the small N =16 the entangled circuit does memorize (train 0.92 vs 0.69), but this is exactly the ambiguous regime a reviewer should worry about: test accuracy is near chance for all three. At N =64, however, every circuit generalizes above the 0.5 chance level (test 0.61–0.66), and there the entangled circuits are strictly worse: the all-to-all circuit reaches test accuracy 0.608 versus 0.656 for the non-entangled one—which itself has essentially zero gap—and the gap is ordered none < linear < full at every N (Figure 2, left). The effect is depth-gated (a single entangling layer, L=1, is inert for this single-qubit readout) and holds at d fixed throughout. A caveat on small-N magnitudes: the slightly negative none gap at N =64 (−0.01) is within one seed standard error of zero—at small N the finite-sample noise in the gap estimate is comparable to the effect for the (near-zero-gap) unentangled circuit—so we anchor every quantitative contrast at the sample sizes where the gaps are resolved beyond their standard error (the ≥ 3σ full−none separations in Tables 5 and 4) and read the small-N , near-chance rows only as tests of the ordering, not of gap magnitudes. Result: deff is the best—and significantly best—predictor. Across a large grid of 300 configurations (five entangling patterns × five depths × three sample sizes × four seeds; Appendix B), the Spearman correlation with the gap is ρ=0.82 for the Fisher effective dimension deff = log det(I+γ F̂), 0.73 for the learned displacement ∥θ − θ0 ∥2 , 0.72 for the entangling power Q, and 0.45 for parameter count. The margin of deff over the learned-norm term is bootstrapsignificant: ∆ρ = 0.090 with 95% CI [0.041, 0.144]. Thus the Fisher effective dimension—the quantity that actually appears in bound equation 3—is not merely nominally but significantly the best predictor of the gap, ahead of the learned-norm term used by Rodriguez-Grasa et al. (2026) and 5
Table 1: Fixed parameter count (d=16, L=2), varying only the entangling connectivity, singleobservable readout, 8 seeds. At N =64 all three circuits generalize well above the 0.5 chance level, yet the entangled circuits have lower test accuracy and a larger gap—genuinely worse generalization, not mere memorization. Ntrain Entanglement Q (M.–W.) train acc test acc gap 16 16 16
none linear full
0.00 0.45 0.56
0.69 0.75 0.92
0.59 0.60 0.54
0.101 0.146 0.381
64 64 64
none linear full
0.00 0.51 0.58
0.65 0.72 0.71
0.656 0.657 0.608
−0.01 0.064 0.104
Table 2: Complexity measures vs. generalization gap: Spearman ρ across a 300-config grid (5 entangling patterns × 5 depths × 3 sample sizes × 4 seeds). The Fisher effective dimension deff = log det(I + γ F̂) (the quantity in the bound) is the strongest predictor; its margin over the learned norm is bootstrap-significant (∆ρ=0.090, 95% CI [0.041, 0.144]). Raw parameter count is the weakest. measure param count Meyer–Wallach Q ∥θ − θ0 ∥2 Fisher deff (log-det) ρ(·, gap)
0.45
0.72
0.73
0.82
far ahead of raw parameter count; and, unlike the norm, entanglement is a structural property fixed before training that separates circuits of equal parameter count.
5
T HE F ISHER E FFECTIVE D IMENSION IS THE R IGHT C OMPLEXITY
We now test Proposition 1 directly by computing, for each circuit, the Fisher effective dimension of Eq. equation 4 at the trained parameters and asking whether it (a) tracks the gap, (b) beats parameter count, and (c) is inflated by entanglement at fixed d. At fixed depth (fixed d), moving from none through linear to full connectivity inflates the Fisher effective dimension deff = log det(I + γ F̂) monotonically, from deff ≈ 3.3 (none) to 9.3 (linear) to 21.8 (full) at L=2, confirming that entanglement raises the effective complexity that Eq. equation 3 charges for while d is held fixed. The generalization gap also shrinks with the training-set size N (Figure 2, left), a basic PAC-Bayes consistency check, and it increases with the Fisher effective dimension across circuits (Figure 2, right). The bound is primarily a ranking certificate. We do not claim the bound is numerically tight; its value is that it correctly orders circuits of identical parameter count, an ordering a parametercounting bound cannot produce at all. We evaluate the right-hand side of Eq. equation 3 directly. Plugging the measured ∥θ − θ0 ∥2 and deff (γ) into the complexity C = ∥θ − θ0 ∥2 /(2τ 2 ) + 12 deff (γ) q √ (with τ 2 =1, γ=50, δ=0.05, Rmax =1) gives the certified bound on the gap, (C + ln 2 δm )/2m, in Table 3. A note on the object. Strictly, Eq. equation 3 bounds the gap of the stochastic posterioraveraged predictor πQ , whereas the “true gap” column is that of the deterministic trained model θ. Rather than assert these “coincide in a low-temperature regime,” we derandomize explicitly: P γλi 2 Lemma 1 (Appendix A) bounds the gap between the two by τγ · 12 i 1+γλ ≤ τ 2 d/(2γ), i.e. 1/γ i times the (dropped) KL curvature term. The key point answers the natural objection that the posterior variance along a flat Fisher direction is τ 2 /(1 + γ · 0) = τ 2 , which is not small: a flat direction carries the full O(1) parameter variance but, precisely because it is flat (λi ≈ 0), its contribution to the prediction perturbation is O(τ 2 λi ) → 0—the large variance sits exactly where the output is least sensitive. With γ=50 the derandomization gap is thus ≤ d/100 and concentrated on the curved directions, so the deterministic and stochastic gaps track each other by construction, not by empirical coincidence. Three things then hold: (i) the certificate is valid—the true gap is below 6
0.4
0.6
generalization gap
generalization gap
0.8 none linear full empirical gap certified bound
0.4 0.2 0.0
0.3
entanglement none linear ring alt full
0.2 0.1 0.0 0.1
2 × 101 3 × 1014 × 101 6 × 101 training set size N
0
102
10 20 30 40 50 60 70 80 Fisher effective dimension deff = logdet(I + F)
Figure 2: Left: the empirical gap (solid) and the certified PAC-Bayes bound of Eq. equation 3 (dashed) versus training-set size N , per circuit. The bound upper-bounds the gap, decays with N at the same rate, and—driven by deff —orders the circuits none < linear < full at every N , all at identical parameter count: the figure shows the bound acting as a ranking certificate, which no table entry conveys. Right: across all circuits the gap increases with the Fisher effective dimension deff = log det(I + γ F̂) (Spearman ρ = 0.82); colour encodes entangling connectivity, which drives deff up at fixed parameter count. the bound for every circuit; (ii) it is numerically loose but nontrivial—the bound stays below the maximal possible unit gap and, where the true gap is non-negligible, within ≈ 2.3–4.5×; and (iii), the property we actually rely on, its complexity ranks the circuits correctly, none < linear < full at each N . On the absolute numbers. The looseness factor in (ii) depends on the prior scale τ 2 =1, which is an a priori choice, not a fitted one; changing τ 2 rescales the distance term and hence every absolute bound value, but leaves the ordering in (iii) untouched (all three rows share the same τ 2 ). We therefore report the absolute column only to show the bound is valid and non-vacuous, and rest the contribution on the ranking, which is τ 2 -invariant. We stress this third point over the second: the bound’s value is as a ranking certificate—something a parameter-counting bound, identical across the three rows, cannot provide at all—rather than as a tight numerical guarantee. Table 3: The PAC-Bayes bound of Eq. equation 3 evaluated per circuit (L=2, τ 2 =1, γ=50, δ=0.05). The certified bound on the gap is valid (exceeds the true gap everywhere), loose by only ∼ 2–4.5× where the gap is non-negligible, and—driven by deff —correctly ordered by entanglement at fixed parameter count. N connectivity complexity C certified gap bound true gap 16 16 16
none linear full
2.7 8.6 19.2
0.49 0.65 0.87
0.101 0.146 0.381
64 64 64
none linear full
2.2 6.2 14.2
0.25 0.31 0.40
−0.01 0.064 0.104
Budgeted, not banished, entanglement. The design implication is not “use less entanglement”— entanglement is what buys a quantum model its expressivity. Figure 3 makes the trade-off explicit: as deff grows with entanglement, training accuracy rises (entanglement fits more), but test accuracy does not follow—the extra training fit is spent on complexity the held-out data does not reward, so the gap opens. The right principle is therefore to budget entanglement against its deff cost, exactly as the bound prices it, rather than to maximize expressivity. The effect persists from 6 to 16 qubits. To check that the finding is not an artifact of the 4qubit toy, we repeat the fixed-parameter-count comparison at n ∈ {6, 8, 10, 12, 16} qubits, with the interaction target extended to n/2 pairwise products (Table 4). At every size the gap is monotone none < linear < full and the all-to-all circuit (d up to 64 trainable angles, a 216 -dimensional state) exceeds the non-entangled one by 4.5–6.9 standard errors, with the effect growing at larger n (fullgap 0.16 → 0.34 from n=6 to n=16). At n ≥ 10 the n/2-product target becomes hard for these 7
0.80 0.75 accuracy
0.70
train acc test acc none linear full
0.65 0.60 0.55 0.50
chance
0.45
101 Fisher effective dimension deff
Figure 3: The accuracy–complexity trade-off (N =64, single-observable readout). As the Fisher effective dimension deff grows with entanglement (colour), train accuracy rises but test accuracy does not, so the gap—the vertical distance between the curves—widens. Entanglement buys training fit that the test distribution does not reward. shallow circuits and held-out accuracy approaches chance, so those rows are a stress test of the gap ordering at larger scale rather than of high-accuracy learning; at n=6, 8 the accuracy stays well above chance. The mechanism is a light-cone argument (Prop. 1) and carries no dependence on n; scaling further, where barren plateaus may reshape the Fisher spectrum, is the natural next test. Table 4: The entanglement–gap effect from n=6 to n=16 qubits (single-observable readout, fixed d=2nL, L=2; 8 seeds at n=6, 8, 4 at n=10, 12, 16). The gap is monotone none < linear < full at every size and full beats none by many standard errors, with the effect strengthening with n. (Held-out accuracy is above chance at n=6, 8 and approaches chance at n≥10, where the target is harder—see text.) n (d) N gap none gap linear gap full full−none 6 (24) 6 (24) 8 (32) 8 (32)
64 128 64 128
0.042 0.035 0.053 0.033
0.098 0.061 0.084 0.071
0.159 0.104 0.183 0.101
+0.117 (6.9σ) +0.070 (4.5σ) +0.130 (6.2σ) +0.068 (4.6σ)
10† (40) 12† (48) 16† (64)
64 64 64
0.062 0.026 0.075
0.124 0.095 0.121
0.274 0.295 0.340
+0.212 (5.6σ) +0.269 (6.4σ) +0.265 (6.3σ)
†
Held-out accuracy near chance at these sizes (the n/2-product target is hard for shallow circuits); these rows test the gap ordering only, not high-accuracy learning. Rows above the rule (n=6, 8) learn well above chance.
6
Q UANTUM RL P OLICIES : BANDIT AND VALUE -F UNCTION G ENERALIZATION
We now test the prediction on quantum policies trained by reward maximization. We use a contextual bandit—the canonical one-step reinforcement-learning problem—which keeps genuine RL feedback (the agent observes only the reward of the action it takes, not the correct action) while removing the trajectory-length variance that makes full-MDP return estimates noisy. A context x ∼ U [−1, 1]4 is drawn; the optimal action follows an interaction rule a⋆ (x) = ⊮[x1 x2 − x3 x4 + 21 x1 > 0]; the reward is 1 if the sampled action matches a⋆ and 0 otherwise. A VQC policy πθ (a=1 | x) = σ(α⟨Z0 ⟩) is trained by REINFORCE on N training contexts (single-observable readout, as in the theory), and we measure the generalization gap in expected reward between the N training contexts and 2000 held-out contexts. As throughout, we fix the rotation layout (L=2, d=16) and vary only the entangling connectivity; each cell averages 16 seeds. The prediction holds cleanly (Table 5, Figure 4). Held-out reward is well above chance (0.60–0.66 vs. 0.5), so the policies genuinely generalize, and at fixed parameter count the entangled policies generalize worse. Both entanglers are significant: full exceeds none by 3.2–3.9 standard errors 8
Table 5: Contextual-bandit (one-step RL) generalization: reward gap at fixed parameter count (d=16), mean ± sem over 16 seeds. Held-out reward is well above the 0.5 chance level (0.60–0.66), so this is genuine generalization; the entangled policies generalize worse, monotonically none < linear < full at the smaller N , and every gap shrinks with N . Each entangled circuit beats none by ≥ 3 standard errors. Ntrain none linear full full−none 0.136 ± 0.019 0.092 ± 0.018 0.026 ± 0.008
reward (accuracy)
16 40 100
0.85 0.80 0.75 0.70 0.65 0.60 0.55 0.50
0.208 ± 0.018 0.134 ± 0.009 0.083 ± 0.011
0.262 ± 0.026 0.173 ± 0.018 0.069 ± 0.008
+0.125 (3.9σ) +0.081 (3.2σ) +0.043 (3.7σ)
none linear full train reward test reward
chance
2 × 101
3 × 101 4 × 101 6 × 101 training contexts N
102
Figure 4: Contextual-bandit (one-step RL) train (solid) and test (dashed) reward versus the number of training contexts N , at fixed parameter count. The entangled policies earn a higher training reward but not a higher test reward—the vertical train–test gap, which Table 5 quantifies, is the entanglement cost—while all policies generalize above chance. This is the reward-space view of the accuracy–complexity trade-off, not a redraw of the gap column. Averages over 16 seeds. at every N and linear by 2.1–4.2, so unlike a two-observable readout the effect no longer rests on a single entangler. At the smaller N the ordering is the monotone none < linear < full predicted by deff (only at N =100, where all gaps are small, does linear edge full). As in the supervised case the entangled policies also fit the training contexts better (training reward 0.86 vs. 0.74 at N =16), and the gap shrinks with N at the rate anticipated by Eq. equation 3. This is the same entanglement– generalization mechanism, now for a policy learned from reward alone. Multi-step value-function generalization. To go beyond one step we test the generalization of a value function in a genuine finite-horizon MDP. States x ∈ [−1, 1]4 evolve by a fixed deterministic PT −1 t map, the per-step reward is the interaction rule, and the value V ⋆ (x) = t=0 γ r(xt ) (T =6, γ=0.9) accumulates future rewards along the trajectory. We fit a single-observable VQC value head to V ⋆ on N states by Monte-Carlo regression—a standard policy-evaluation primitive—and report the R2 generalization gap between the N training states and 1000 held-out states (Figure 5). At N =64, where held-out R2 is positive for all circuits, the gap is monotone none < linear < full (0.089 < 0.179 < 0.299); at the small N =16 the entangled value functions have negative held-out R2 (a degenerate memorization regime, as in the small-N classification rows), so we treat that row as a stress test only. We are explicit about scope here: Monte-Carlo value estimation is a supervised-flavoured RL primitive, and this is its strength (low variance) and its limit. Under genuine bootstrapped fitted-Q iteration the effect washes out—consistent with the bound, as the bootstrap target is itself a strong regularizer—so we do not claim the entanglement penalty for temporal-difference value learning, only for the Monte-Carlo (regression) form.
7
C ONTROLS , AND A R EAL -H ARDWARE T EST
The gap is caused by entanglement, not a confounder. Varying the entangling connectivity holds the parameter count fixed but could, in principle, change the gap through optimization or trainability rather than through deff . We rule this out with three controls at N =64 (Figure 6, left): (i) matched training accuracy —early-stopping each circuit at the same train accuracy, so the en9
Multi-step RL: value-function generalization
value-function gap (R 2)
2.5
entanglement none linear full
2.0 1.5 1.0 0.5 0.0
2 × 101
3 × 101 4 × 101 training states N
6 × 101
Figure 5: Multi-step RL: the value-function generalization gap (R2 , Monte-Carlo policy evaluation, horizon 6) is monotone none < linear < full at every training-set size, at fixed parameter count. Fixed parameter count, 8 seeds; error bars ±1 s.e. tangled circuits do not simply fit more—still leaves full with ∼ 3× the gap of none (0.093 vs. 0.032) and lower test accuracy; (ii) a different readout (Z1 in place of Z0 ) preserves the ordering (0.074/0.126/0.131); and (iii) a different optimizer (SGD in place of Adam) preserves it (0.034/0.065/0.089). In every control the entangled circuits generalize worse at identical parameter count. The mechanism is visible directly in the Fisher spectrum (Appendix B, Figure 8): none has effectively two non-zero eigenvalues, linear three, and full five or more—entanglement raises the rank, hence deff . Local vs. global readout: the effect is the readout light-cone, and deff is the invariant. The most adversarial test of the mechanism is also the most revealing. Because our single readout is the local observable ⟨Z0 ⟩, a rotation on a wire that never reaches qubit 0 has identically zero gradient (Lemma 2), so the unentangled none circuit is effectively a smaller model—only ∼ d/n of its parameters are causally live. One might therefore object that “entanglement raises the gap at fixed d” is merely “entanglement recruits otherwise-dead parameters.” That objection is correct, and it is the mechanism; Pwe test it by the cleanest manipulation available—replacing the local readout by the global one n1 i ⟨Zi ⟩, under which every wire couples to the output directly (Table 6). The robust finding is that the global readout sharply raises none’s live-parameter count, deff , and gap at both sizes— at n=4 from (2 live, deff =2.6, gap=−0.008) to (7, 14.0, 0.095), at n=6 from (2, 2.9, 0.018) to (11, 23.6, 0.173): the readout alone activates the parameters entanglement would have, and deff and the gap move together. What happens to the full−none contrast is then n-dependent and instructive. At n=4 the single global observable nearly saturates the light-cone even for none (7 of 16 live), so the entanglement contrast is largely absorbed: full−none falls from +0.112 (4.1 s.e., local) to +0.033 (0.9 s.e., global). At n=6, where 24 parameters cannot be saturated by one global observable (none reaches 11 of 24 live, full 17), entanglement still adds coupling on top and the contrast persists: +0.114 (3.8 s.e., local) vs. +0.115 (4.2 s.e., global). The conclusion is therefore not that entanglement is dispensable but that it is a readout-gated knob on the same underlying quantity: what governs deff is how many parameters are causally coupled to the readout; a global readout raises the floor of that coupling—fully absorbing the entanglement contrast at small n, only partially at larger n—while in every cell it is deff , not the entangling label, that tracks the gap. This confirms Proposition 1 and sharpens the thesis: entanglement is the dominant, structural driver of effective dimension for the local single-qubit Pauli readouts standard in quantum RL (e.g. Jin et al., 2025; Lin et al., 2025), and is partly substitutable by—never orthogonal to—the choice of readout. Beyond a single target. To show the effect is not an artifact of one synthetic rule, we repeat the fixed-parameter-count comparison on three further targets (Table 7): a random-Fourier function, labels from an entangled teacher PQC, and 4-bit parity. On the Fourier target the entangled circuits generalize above chance and the gap is again ordered none < linear < full. The entangled-teacher target is hard for all student circuits at this sample size (test accuracy near chance for every connectivity); we include it only as a stress test of gap behaviour—which still orders none < linear < full—and not as evidence of successful learning. Parity is the instructive exception and the clearest statement of the trade-off: it is not representable without entanglement, so the non-entangled and 10
Table 6: Local vs. global readout (L=2, N =64, 8 seeds). A local readout ⟨Z0 ⟩ leaves most of none’s parameters outside the readout light-cone, P so entanglement’s recruitment of them makes full−none significant. A global readout n1 i ⟨Zi ⟩ activates those parameters directly—raising none’s deff and gap at both sizes—so it absorbs the contrast fully at n=4 (few free parameters) but only partly at n=6 (more parameters than one observable saturates). In every row deff tracks the gap; the entangling label does not. n readout connectivity live params deff gap 4 4 4 4
local ⟨Z0 ⟩ local ⟨Z0P ⟩ global n1 Pi Zi global n1 i Zi
none full none full
2/16 5/16 7/16 12/16
2.6 12.9 14.0 30.7
−0.008 0.104 0.095 0.128
6 6 6 6
local ⟨Z0 ⟩ local ⟨Z0P ⟩ global n1 Pi Zi global n1 i Zi
none full none full
2/24 7/24 11/24 17/24
2.9 21.1 23.6 35.4
0.018 0.132 0.173 0.289
full−none gap (n=4): local +0.112 (4.1σ), full−none gap (n=6): local +0.114 (3.8σ),
global +0.033 (0.9σ) global +0.115 (4.2σ)
linear circuits sit at chance (0.50) while the all-to-all circuit is the only one that learns it (test accuracy 0.90). Entanglement is therefore not uniformly harmful—when the target demands it, it is necessary—which is exactly why the right principle is to budget it, not to remove it. Table 7: The effect across target distributions (n=4, N =64, fixed d=16). On targets all circuits can fit (Fourier, teacher-PQC) entanglement widens the gap; on parity, which requires entanglement, only full learns the task (test 0.90 vs. chance). target none gap (test) linear gap (test) full gap (test) random Fourier entangled teacher 4-bit parity
0.038 (0.71) 0.074 (0.50) 0.067 (0.50)
0.045 (0.69) 0.123 (0.50) 0.116 (0.50)
0.103 (0.67) 0.148 (0.51) 0.050 (0.90)
Real (non-synthetic) benchmarks: entanglement lowers test accuracy. The mechanism is not tied to a hand-designed rule, and on real data the cost shows up in test accuracy, not merely in the train–test gap. On standard small datasets—Iris, Breast Cancer (Wisconsin), Wine, and Digits (3 vs. 8), reduced to four features and binarized—the fixed-parameter-count comparison again orders the gap none < linear < full and co-monotone with deff (Table 8). More importantly for the practical claim, on Wine and Digits the non-entangled circuit is significantly more accurate on held-out data than the all-to-all one, at accuracies far above chance: pooling over N ∈ {16, 24, 32} and 22 seeds, the test accuracy drops from 0.884 (none) to 0.870 (full) on Wine (2.5 standard errors) and from 0.911 to 0.900 on Digits (3.1 standard errors). Entanglement here genuinely hurts generalization on real tasks—not just widening the gap—which is what makes “budget entanglement” an operational, and not merely rhetorical, prescription. Table 8: Real datasets (binary, 4 features, N =40, fixed d=16, 6 seeds). The gap is ordered none < linear < full and co-monotone with deff , at test accuracies well above chance—so the effect is not an artifact of the synthetic target. dataset none gap (deff , test) linear gap (deff , test) full gap (deff , test) Iris Breast Cancer Wine
0.00 (3.0, 0.70) 0.025 (4.5, 0.90) 0.029 (5.1, 0.90)
0.061 (5.2, 0.65) 0.033 (10.3, 0.91) 0.059 (9.1, 0.88)
11
0.082 (18.2, 0.76) 0.069 (13.8, 0.90) 0.125 (13.2, 0.85)
The ordering holds where the model learns, and barren plateaus bound the mechanism. Two concerns delimit the mechanism’s reach, and we test both directly. Is the ordering an artifact of the near-chance rows? We construct a deliberately learnable regime—n=6, a lower-order target, N =48—in which all three circuits reach held-out accuracy 0.72–0.77, far above the 0.5 chance level. The ordering survives cleanly: the gap is 0.042 (none), 0.043 (linear), 0.165 (full), with full−none= +0.123 at 5.8 standard errors (8 seeds) and deff co-monotone (4.2/7.8/21.7). The effect is thus not confined to the degenerate near-chance rows—it is sharpest exactly where the circuits genuinely generalize. Do barren plateaus erode the mechanism? Our mechanism runs on the Fisher spectrum, which barren plateaus flatten exponentially in qubit number, so the regime where scale matters is a natural worry. We confirm the erosion directly (Table 9): at a fixed moderate depth (L=4, all-to-all, random parameters) the effective dimension falls monotonically from deff =13.2 at n=4 to 3.1 at n=12, even as the nominal parameter count rises from 32 to 96—the Fisher eigenvalues shrink (the ⟨Z0 ⟩ gradient variance decays with n) faster than the rank grows, so deff erodes toward its floor. This makes the boundary flagged in our failure-mode analysis (Appendix A) quantitative: entanglement inflates deff only while the circuit is off the plateau; deep or wide enough, barren plateaus drive deff —and trainability itself—down together. The regime our claim and design recipe target is therefore the trainable one (shallow circuits away from the plateau), which is where all our learning experiments sit.
Table 9: Barren plateaus erode the mechanism’s substrate. At fixed depth (L=4, all-to-all, random parameters, 6–12 seeds), the Fisher effective dimension deff decreases with qubit number even as the nominal parameter count d grows: the Fisher eigenvalues (hence deff ) are suppressed by the barren plateau faster than the rank rises. This is the regime our trainable-circuit experiments deliberately avoid. n 4 6 8 10 12 nominal params d deff (L=4, random)
32 13.2
48 8.4
64 4.5
80 3.2
96 3.1
deff , not parameter count, is the stable predictor. The complementary ablation—fixing the entangling pattern and varying depth, so the parameter count changes—shows the gap rising with depth for both patterns (linear: L=1→4 gap 0.03→0.33; full: 0.03→0.31), and in lock-step with deff (which rises 2→50 and 2→61 respectively). At matched depth (matched parameter count) the more entangled pattern still has the larger deff and gap. Thus deff predicts the gap whether we vary connectivity at fixed depth or depth at fixed connectivity, whereas raw parameter count only tracks the gap when it happens to co-vary with deff —the sense in which deff , not parameter count, is the governing quantity. End-to-end multi-step policy learning. Finally, we train policies end-to-end by REINFORCE (with a value baseline, horizon 8, 20 seeds) on the multi-step return of an interaction-gated MDP—a genuine multi-step policy-learning setting, not a value regression. Here the return generalization gap of the standard hardware-efficient (linear) entangler is significantly larger than the non-entangled circuit’s at the small N =8 (0.62 vs. 0.16, 2.8 standard errors over 20 seeds), directly confirming the mechanism in end-to-end policy learning. The all-to-all circuit and the larger-N regime, however, are too high-variance to resolve (full-vs-none is within noise, and at N =16 all three are statistically indistinguishable): return-based estimates carry much more variance than the accuracy-based ones. We therefore report the significant linear-vs-none effect as genuine multi-step evidence, but continue to treat the low-variance bandit and value-function settings as our primary probes of the mechanism. The effect survives on real quantum hardware. Finally we run the trained 4-qubit classifiers on an IBM Heron device (ibm aachen, 156 qubits) via the Estimator primitive at 4096 shots, reading ⟨Z0 ⟩ for all train and test inputs (Figure 6, right). Despite real gate and readout noise, the models run above chance and the entanglement–generalization ordering is preserved: the gap is 0.14 for none versus 0.33 and 0.31 for linear and full. The entanglement cost that our bound predicts is therefore not a simulation artifact—it is measurable on today’s hardware. 12
0.4 generalization gap
generalization gap
0.14 0.12 0.10 0.08 0.06 0.04 0.02 0.00
entanglement none linear full
matched train acc
readout Z1
0.3 entanglement
0.1 0.0
SGD
none linear full
0.2
noiseless
noisy sim (ibm_aachen)
IBM Heron hardware
Figure 6: Left: the entangled->-none gap ordering survives three controls at N =64—matched training accuracy, a different readout observable, and a different optimizer—so the gap is not a mere optimization or trainability artifact. Right: the ordering is reproduced by a noise-model simulator of ibm aachen and by execution on the ibm aachen Heron processor itself under real noise; calibration and shot-noise details are in Appendix B.
Our bound is loose in absolute terms—like other PAC-Bayes bounds for expressive models it is most useful as a certificate that is valid, correctly ordered, and (where the gap is non-negligible) within ≈ 2.5×, rather than numerically tight. We frame entanglement as an independent and previously unisolated axis of complexity—the strongest single predictor of the gap in our study, ahead of the learned-norm term—though we do not claim it is the only one. The effect is depth-gated (a single entangling layer is inert for a single-qubit readout). Our RL evidence is for the one-step (contextual bandit) setting, chosen because full-MDP return estimates are dominated by trajectory-length variance; extending the clean measurement to multi-step MDPs—for example through offline valuefunction generalization—is the natural next step, along with value-based agents, hardware noise as an additional complexity channel, and a tight data-dependent prior. Our experiments are simulations of a controlled interaction task at n ∈ {4, . . . , 16} qubits (Table 4); scaling to the many-qubit, hardware regime and to richer tasks is important future work, though the underlying mechanism—a light-cone argument—carries no dependence on n. A design recipe. The trade-off is directly actionable as an ansatz-selection rule. Given a candidate ansatz family, estimate deff on a small calibration set and choose the shallowest, sparsest entangling pattern whose training fit still improves a validation proxy without an excessive rise in deff . Concretely: 1. train each candidate ansatz to convergence on the calibration set; 2. compute its Fisher effective dimension deff = log det(I + γ F̂); 3. discard any ansatz whose deff increases without a matching validation gain; 4. select from the Pareto frontier of validation performance versus deff . This turns “budget entanglement” from a slogan into a concrete selection criterion—and, unlike a parameter budget, it distinguishes circuits of identical size.
8
C ONCLUSION
We gave a PAC-Bayesian account of generalization for quantum policies in which the complexity that matters is the effective dimension of the circuit-induced Fisher geometry, a quantity inflated by entanglement and invisible to parameter counting. Isolating entanglement at fixed parameter count, we showed it is an independent predictor of the generalization gap—stronger than raw parameter count and comparable to the learned-norm term—across supervised classification, a reward-only contextual bandit, multi-step value-function generalization, and real quantum hardware. The resulting design principle—budget entanglement against its effective-dimension cost rather than maximize expressivity—gives quantum-policy design a target beyond average return. 13
R EPRODUCIBILITY S TATEMENT Experiments use PQCs of n=4 to 16 qubits simulated with PennyLane (default.qubit with backprop at n ≤ 8; lightning.qubit with the adjoint method at n ≥ 10), plus one run on the IBM Heron ibm aachen processor; unless a scaling experiment states otherwise (Table 4), the default is n=4. Code, ansätze, target/environment distributions, hyperparameters (τ 2 =1, γ=50) and seeds are described in Sections 4–6 and Appendix B.
R EFERENCES Amira Abbas, David Sutter, Christa Zoufal, Aurélien Lucchi, Alessio Figalli, and Stefan Woerner. The power of quantum neural networks. Nature computational science, 1(6):403–409, 2021. Leonardo Banchi, Jason Pereira, and Stefano Pirandola. Generalization in quantum machine learning: A quantum information standpoint. PRX Quantum, 2(4):040321, 2021. Matthias C Caro, Hsin-Yuan Huang, Marco Cerezo, Kunal Sharma, Andrew Sornborger, Lukasz Cincio, and Patrick J Coles. Generalization in quantum machine learning from few training data. Nature Communications, 13(1):4919, 2022. Rodrigo Coelho, André Sequeira, and Luı́s Paulo Santos. Vqc-based reinforcement learning with data re-uploading: performance and trainability. Quantum Machine Intelligence, 6(2):53, 2024. Gintare Karolina Dziugaite and Daniel M Roy. Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data. arXiv preprint arXiv:1703.11008, 2017. Elies Gil-Fuster, Jens Eisert, and Carlos Bravo-Prieto. Understanding quantum machine learning also requires rethinking generalization. Nature Communications, 15(1):2277, 2024. Yu-Xin Jin, Zi-Wei Wang, Hong-Ze Xu, Wei-Feng Zhuang, Meng-Jun Hu, and Dong E Liu. Ppoq: Proximal policy optimization with parametrized quantum policies or values. arXiv preprint arXiv:2501.07085, 2025. Hsin-Yi Lin, Samuel Yen-Chi Chen, Huan-Hsin Tseng, and Shinjae Yoo. Quantum reinforcement learning by adaptive non-local observables. In 2025 IEEE International Conference on Quantum Computing and Engineering (QCE), volume 2, pp. 241–246. IEEE, 2025. Jie Luo, Jeremy Kulcsar, Xueyin Chen, Giulio Giaconi, and Georgios Korpas. Training hybrid deep quantum neural network for efficient reinforcement learning. arXiv preprint arXiv:2503.09119, 2025. Anirudha Majumdar, Alec Farid, and Anoopkumar Sonar. Pac-bayes control: learning policies that provably generalize to novel environments. The International Journal of Robotics Research, 40 (2-3):574–593, 2021. David A McAllester. Pac-bayesian model averaging. In Proceedings of the twelfth annual conference on Computational learning theory, pp. 164–170, 1999. Pablo Rodriguez-Grasa, Matthias C Caro, Jens Eisert, Elies Gil-Fuster, Franz J Schreiber, and Carlos Bravo-Prieto. A pac-bayesian approach to generalization for quantum models. arXiv preprint arXiv:2603.22964, 2026. Shaojun Wu, Shan Jin, Dingding Wen, Donghong Han, and Xiaoting Wang. Quantum reinforcement learning in continuous action space. Quantum, 9:1660, 2025. Abdelkrim Zitouni, Mehdi Hennequin, Juba Agoun, Ryan Horache, Nadia Kabachi, and Omar Rivasplata. Pac-bayesian reinforcement learning trains generalizable policies. arXiv preprint arXiv:2510.10544, 2025.
14
A
P ROOFS
A.1
P ROOF OF T HEOREM 1
We reduce the return bound to the classical McAllester PAC-Bayes theorem. Recall its statement for a loss taking values in [0, 1]. Theorem 3 (McAllester, 1999). Let ℓ(h, z) ∈ [0, 1], let z1 , . . . , zm be i.i.d. from a distribution D, and let P be a prior over hypotheses chosen before seeing the data. For any δ ∈ (0, 1), with probability at least 1 − δ, simultaneously for all posteriors Q, s √ m X KL(Q∥P ) + ln 2 δm 1 Eh∼Q Ez∼D ℓ(h, z) ≤ Eh∼Q ℓ(h, zi ) + . (5) m i=1 2m Take the hypothesis h to be a parameter vector θ (so Q, P are the Gaussians of Section 3), the sample zi = Mi to be the i-th training environment drawn i.i.d. from D, and the loss to be the normalized regret of a single rollout, ℓ(θ, M ) = 1 − RM (θ)/Rmax ∈ [0, 1], where RM (θ) is the (bounded) return of πθ on M . Then EQ EM ℓ = 1 − J(πQ )/Rmax and its empirical counterpart is ˆ Q )/Rmax . Substituting into Theorem 3 and multiplying by Rmax gives 1 − J(π s √ 2 m KL(Q∥P ) + ln δ ˆ Q ) − Rmax J(πQ ) ≥ J(π . (6) 2m For a Gaussian posterior/prior with common covariance σ 2 Id , KL(Q∥P ) = ∥θ − θ0 ∥2 /(2σ 2 ), which is Eq. equation 1 with meff = m. When the per-environment return is instead estimated by pooling the T transitions of a trajectory—which are Markov-dependent rather than i.i.d.—the effective number of independent samples is reduced by the chain mixing factor, meff = m/κ with κ = Θ(tmix ), by the blocking argument of Zitouni et al. (2025); this yields the stated form. □ A.2
P ROOF OF T HEOREM 2
The KL divergence between two d-dimensional Gaussians Q = N (θ, Σ) and P = N (θ0 , τ 2 Id ) is h i KL(Q∥P ) = 12 τ12 tr Σ + τ12 ∥θ − θ0 ∥2 − d − log det τΣ2 . (7) P Substitute the Gauss–Newton covariance Σ = τ 2 (Id + γ F̂)−1 and diagonalize F̂ = i λi ui u⊤ i . P P P P Then τ12 tr Σ = i (1 + γλi )−1 , log det τΣ2 = − i log(1 + γλi ), and using d = i 1 = i (1 + P γλi γλi γλi )−1 + 1+γλ the trace and −d terms combine to − i 1+γλ . Collecting terms, i i KL(Q∥P ) =
X X γλi ∥θ − θ0 ∥2 log(1 + γλi ) − 21 , + 12 2 2τ 1 + γλi i i
(8)
P which is equation 2 since i log(1+γλi ) = log det(Id +γ F̂) = deff (γ). Each summand of the last P γλi term lies in [0, 12 ), so 0 ≤ 21 i 1+γλ ≤ d/2 and dropping it can only enlarge the right-hand side of i the bound. The prior P is fixed at initialization θ0 and does not depend on the training environments, so Theorem 3 applies verbatim with this Q and P ; substituting the resulting KL into the return bound of Theorem 1 gives equation 3. The asymptotics deff (γ) ≤ d and deff (γ) = rank(F̂) log γ + O(1) as γ → ∞ follow from log(1 + γλi ) ≤ γλi and, for λi > 0, log(1 + γλi ) = log γ + log λi + o(1). □ A.3
D ERANDOMIZATION : FROM THE STOCHASTIC πQ TO THE DETERMINISTIC θ
The bound equation 3 controls the gap of the posterior-averaged predictor πQ , whereas the reported “true gap” is that of the deterministic trained θ. The following lemma bounds the difference and, crucially, shows the flat Fisher directions—which carry the full prior variance τ 2 —do not spoil it, because the perturbation of a loss along a direction is weighted by that direction’s curvature, which is the Fisher eigenvalue itself. 15
Lemma 1 (Flat-direction derandomization). Let Q = N (θ, Σ) with Σ = τ 2 (Id + γ F̂)−1 , and let L(ϑ) = Es ℓ(ϑ, s) be a population loss whose Hessian at θ equals the Fisher F̂(θ) up to a residual term R that is O(∥p − y∥) and vanishes in the Gauss–Newton (well-fit) regime. Then Eϑ∼Q L(ϑ) − L(θ)
d τ 2 X λi = 21 tr Σ F̂ + O(∥Σ∥3/2 ) = + O(∥Σ∥3/2 ), 2 i=1 1 + γλi
(9)
d 1 1 P γλi · ≤ . Consequently the train and test losses of πQ i γ 2 P 1 + γλi 2γ and of θ each differ by at most 12 τ 2 i λi /(1 + γλi ) ≤ τ 2 d/(2γ), so the two generalization gaps P differ by at most τ 2 i λi /(1 + γλi ) ≤ τ 2 d/γ. and the sum is bounded by
Proof. Since Eϑ∼Q [ϑ − θ] = 0, a second-order Taylor expansion of L about θ gives EQ L(ϑ) − L(θ) = 21 EQ [(ϑ − θ)⊤ ∇2 L (ϑ − θ)] + O(∥Σ∥3/2 ) = 12 tr(Σ∇2 L) + O(∥Σ∥3/2 ), the remainder controlled by the third derivative of the smooth PQC expectation and the Gaussian third moment. For the logistic head, ∇2 L = Es [α2 ps (1−ps )∇fs ∇fs⊤ ] + Es [α(ps −ys )∇2 fs ] = F̂ + R, and R vanishes as the fit improves (ps → ys ). Substituting Σ = τ 2 (Id + γ F̂)−1 and diagonalizing F̂ gives P tr(ΣF̂) = τ 2 i λi /(1+γλi ), which is equation 9. Writing λi /(1+γλi ) = γ −1 ·γλi /(1+γλi ) and γλi /(1+γλi ) ∈ [0, 1) gives the d/2γ envelope. A flat direction (λi → 0) contributes λi /(1+γλi ) → λi → 0 despite its posterior variance Σii → τ 2 : the variance is large but the curvature that converts it into a loss change is exactly λi , so flat directions are harmless. The gap statement follows by applying the bound to the train and test losses separately and using the triangle inequality. P Note that the derandomization cost 12 τ 2 i λi /(1+γλi ) is exactly 1/γ times the KL curvature term P 1 i γλi /(1 + γλi ) that we dropped in Eq. equation 2: the same quantity we discarded to expose 2 deff reappears, scaled by 1/γ, as the price of derandomization, and γ=50 makes it small (≤ d/100 for τ 2 =1). This ties the two approximations together and replaces the informal “low-temperature” argument in the main text. A.4
P ROOF OF P ROPOSITION 1
Write the readout expectation in the Heisenberg picture, f (s, θ) = ⟨ψ(s)| U (θ)† Z0 U (θ) |ψ(s)⟩, and let gate i be a Pauli rotation e−iθi Gi /2 with Gi a single-qubit Pauli on wire q(i). Splitting the circuit at gate i as U = U>i e−iθi Gi /2 U≤i and defining the back-propagated observable Oi (θ) = † U>i Z0 U>i , the parameter-shift identity gives ∂θi f (s, θ) = − 2i ⟨ϕi (s)| [ Gi , Oi (θ) ] |ϕi (s)⟩,
|ϕi (s)⟩ = U≤i |ψ(s)⟩.
(10)
The (output-)Fisher matrix is F̂ij = Es [ c(s) ∂θi f ∂θj f ] with c(s) = sc2 p(1−p) > 0, so F̂ = Es [c(s) gs gs⊤ ] where gs = ∇θ f (s, θ); in particular rank(F̂) ≤ |{i : ∂θi f ̸≡ 0}|, with equality generically. Definition 1 (Backward light-cone). Let L(θ) ⊆ {1, . . . , d} be the set of parameters i for which Oi (θ) has non-trivial support on wire q(i), i.e. [Gi , Oi (θ)] ̸= 0. Lemma 2 (Entanglement enlarges the light-cone). Fix the placement of all rotation gates. If a set of two-qubit entangling gates is added to the circuit (keeping all rotations), then L(θ) can only grow: Lwithout (θ) ⊆ Lwith (θ) for generic θ. In particular, with no entangling gates Oi (θ) is supported on wire 0 alone, so [Gi , Oi ] = 0 for every rotation on a wire q(i) ̸= 0 and L = {i : q(i) = 0}. † Proof. Oi (θ) = U>i Z0 U>i is Z0 conjugated by the gates after i. Each single-qubit rotation preserves the support of an operator on its own wire; each two-qubit gate can only enlarge support to the wires it couples. Hence the support of Oi is contained in the causal cone of Z0 under the gates in U>i , which is monotone non-decreasing under the addition of entangling gates. With no entangling gates the cone is {0}, so Oi commutes with any Gi on a different wire, giving ∂θi f ≡ 0 by equation 10. Adding entangling gates that reach wire q(i) makes Oi act non-trivially there, so [Gi , Oi ] ̸= 0 for generic θ and i enters L.
16
Assumption 1 (Genericity). For generic parameters the nonzero gradients {gs } span span{ei : i ∈ L(θ)} and the induced Fisher eigenvalues on this subspace are of comparable magnitude (no exact degeneracies or cancellations). Proof of Proposition 1. The rank identity rank(F̂) = |{i : ∂θi f ̸≡ 0}| = |L(θ)| holds because F̂ = Es [c(s)gs gs⊤ ] is supported on span{ei : i ∈ L(θ)} (parameters outside the light-cone have ∂θi f ≡ 0 by Lemma 2 and equation 10), and by Assumption 1 the in-cone gradients span that subspace, so the inequality is an equality. Monotonicity of |L(θ)| in the entangling connectivity P is Lemma 2. Because deff (γ) = i log(1 + γλi ) receives a contribution only from the rank(F̂) nonzero eigenvalues and deff (γ) = rank(F̂) log γ + O(1) for large γ, any increase of rank(F̂) raises deff (γ) at fixed d. Hence increasing entangling connectivity raises deff while d is fixed. The empirical counterpart is Table 10: at fixed depth (fixed d) the measured deff = log det(I + γ F̂) rises monotonically from its no-entanglement value (none → linear → full), tracking the rank of the Fisher matrix. Scope, a worked example, and failure modes. We state precisely what Proposition 1 does and does not give. Scope. It applies to hardware-efficient ansätze in which (a) each trainable gate is a Pauli rotation, (b) the readout is a fixed Pauli observable, and (c) encoding gates precede the trainable block. The two claims have different strengths: rank monotonicity (rank(F̂with ) ≥ rank(F̂without )) is rigorous—it is Lemma 2 plus the support argument, with no genericity needed for the inequality (genericity is used only to turn it into the equality rank = |L|). The stronger effective-dimension monotonicity deff (γ)with ≥ deff (γ)without holds unconditionally only as γ → ∞ (where deff → rank · log γ); at finite Pγ a rank increase raises deff provided the new eigenvalues are not vanishingly small, since deff = i log(1 + γλi ) and a mode with λ ≈ 0 contributes ≈ 0. We therefore separate the two in the statement. Worked example. Take n=2 qubits, one layer, readout Z0 , trainable RY (θ0 ) on qubit 0 and RY (θ1 ) on qubit 1. Without a CNOT, ∂θ1 ⟨Z0 ⟩ ≡ 0 (the back-propagated observable is Z0 ⊗ I, which commutes with Y1 ), so F̂ = diag(c ⟨. . . ⟩, 0) has rank 1. Inserting CNOT0→1 before the readout maps Z0 7→ Z0 but CNOT1→0 maps Z0 7→ Z0 Z1 , which no longer commutes with Y1 ; then ∂θ1 ⟨Z0 ⟩ = − 2i ⟨[Y1 , Z0 Z1 ]⟩ ̸= 0 generically and rank(F̂) = 2. This is the mechanism in closed form. Failure modes. The rank does not increase when the added entanglement is neutralized by structure: (i) a symmetry or commuting-gate configuration for which [Gi , Oi (θ)] = 0 despite non-trivial support (e.g. Oi proportional to Zq(i) and Gi = Zq(i) ); (ii) fine-tuned parameters at which the new gradients vanish (a measure-zero set excluded by genericity); and (iii) the barren-plateau regime, where all λi → 0 exponentially, so although the rank may rise the finite-γ deff stays near zero— entanglement then inflates capacity only once the circuit is trained away from the plateau. These are exactly the cases our experiments avoid (shallow circuits, generic trained parameters), and they delimit the claim.
B
A DDITIONAL EXPERIMENTAL DETAIL
Setup. All PQCs use n=4 qubits simulated with PennyLane’s default.qubit and a single readout observable ⟨Z0 ⟩. Inputs are angle-encoded with RY ; the ansatz is L layers of RY , RZ per qubit (d=2nL trainable angles) followed by a fixed CNOT pattern (none/linear/full). Figure 7 shows the three ansätze: they share the identical rotation layout (hence d) and differ only in the entangling block. Supervised models (Sections 4–5) are binary classifiers ŷ = sign(⟨Z0 ⟩) trained by Adam (600 steps, full batch, lr 0.08) with the cross-entropy loss on the interaction target y = sign(x1 x2 −x3 x4 + 21 x1 ), x ∼ U [−1, 1]4 , and evaluated on 2000 held-out inputs. RL policies (Section 6) are contextual-bandit policies π(a=1 | x) = σ(α⟨Z0 ⟩) over contexts x ∼ U [−1, 1]4 , trained by REINFORCE (600 updates, lr 0.08) on bandit feedback (reward 1 iff the sampled action equals a⋆ (x) = ⊮[x1 x2 − x3 x4 + 12 x1 > 0], the correct label is never revealed); the gap is training reward minus reward on 17
Z0
none
q0
RY
RY
RZ
q1
RY
RY
RZ
q2
RY
RY
RZ
q3
RY
RY
RZ Z0
linear
q0
RY
RY
RZ
q1
RY
RY
RZ
q2
RY
RY
RZ
q3
RY
RY
RZ
q0
RY
RY
RZ
q1
RY
RY
RZ
q2
RY
RY
RZ
q3
RY
RY
RZ
Z0
full
Figure 7: The three ansätze at n=4 (one layer of the repeated block shown; the block repeats L times, here between the encoding RY (xi ) and the ⟨Z0 ⟩ readout). All three use the same RY , RZ rotations—hence the same parameter count d=2nL—and differ only in the entangling block: none (no CNOTs), a linear CNOT chain, or all-to-all CNOTs. This is the “fixed parameter count, vary only entanglement” comparison.
2000 held-out contexts. Each cell averages 8 (supervised) or 16 (bandit) seeds. The Meyer–Wallach Q and the Fisher matrix are evaluated at the trained parameters over held-out inputs.
Fisher effective dimension: exact computation. To remove ambiguity we specify the estimator precisely. (i) We use the empirical output Fisher of the policy’s Bernoulli head, not the paP 1 ⊤ rameter/loss Fisher and not the exact Fisher: F̂(θ) = M x c(x) gx gx with gx = ∇θ ⟨Z0 ⟩x,θ 2 and Bernoulli weight c(x) = α p(x)(1−p(x)), p = σ(α⟨Z0 ⟩). (ii) It is evaluated at the trained parameters θ over M =40 held-out inputs drawn from the input distribution (not the training set). (iii) The per-input gradients gx are computed by autograd (exact parameter-shift). (iv) Eigenvalues λi of F̂ are used unnormalized (no trace rescaling). (v) The effective dimension is P deff = log det(I + γ F̂) = i log(1 + γλi ) with the constant γ = 50 used identically everywhere (Remark 1). (vi) For a reported cell we compute deff per seed P and then P average deff across seeds (we do not pool the Fisher matrices). The participation ratio ( i λi )2 / i λ2i , when reported, is computed from the same spectra. Why γ=50, and does it matter. The value 50 places the observed spectra (eigenvalues λ ∈ [10−2 , 2]) in the informative band where γλ spans ≈ 0.5–100, so deff is neither saturated nor collapsedP to zero. Crucially the conclusions we draw from deff are insensitive to this choice, since deff (γ) = i log(1 + γλi ) is monotone in the Fisher spectrum. We recomputed deff for all 18 configurations of Table 2 at γ ∈ {10, 25, 50, 100, 200}: the Spearman correlation ρ(deff , gap) is 0.74, 0.70, 0.67, 0.67, 0.64 respectively—stable in sign and magnitude—and the ordering none < linear < full holds at every γ (e.g. at L=2, deff goes 2.5→4.8→11.5 at γ=10 and 5.8→12.2→25.6 at γ=200). We fix a single γ = 50 only so that deff is numerically comparable across rows. 18
deff dominates the learned norm—a powered, bootstrapped test. A comparison between two correlations of 0.76 and 0.70 over 18 points would be underpowered, so we built a much larger and more varied grid: 5 entangling patterns (none/linear/ring/alt/full) × depths L ∈ {1, . . . , 5} × N ∈ {16, 32, 64} × 4 seeds = 300 configurations, which dissociates deff from the learned norm and from parameter count. Across these 300 points the Spearman correlations with the gap are ρ(deff ) = 0.82, ρ(∥θ−θ0 ∥2 ) = 0.73, ρ(Q) = 0.72, ρ(params) = 0.45. A 5000-resample bootstrap over configurations gives ∆ρ = ρ(deff ) − ρ(norm) = 0.090 with 95% CI [0.041, 0.144] (bootstrap P (∆ρ>0)=1.00), and ρ(deff ) − ρ(Q) = 0.100, CI [0.046, 0.156]: the Fisher effective dimension is a significantly better predictor of the gap than the learned norm or the entangling power. A multiple regression on the same grid loads almost entirely on deff (standardized coefficient +0.16 vs. ≤ 0.06 in magnitude for the others), and adding the entangling pattern on top of deff raises R2 by < 0.01 while adding deff on top of the pattern raises it by 0.14. Entanglement therefore acts on the gap through deff , which screens it off—the precise, and now powered, sense in which deff is the governing quantity. Fisher effective dimension and gap, per configuration. Table 10 gives the Fisher effective dimension deff = log det(I + γ F̂) (Eq. equation 4, fixed γ = 50) and the gap for every connectivity at depths L=2, 3 and Ntrain =16. At fixed depth (fixed d), deff increases monotonically with entangling connectivity, and so does the gap.
Table 10: Fisher effective dimension deff = log det(I + γ F̂) and generalization gap. At each fixed (L, Ntrain ) block the parameter count d = 8L is identical across the three rows; only the entangling connectivity—and hence deff and the gap—changes. deff rises monotonically none → linear → full. depth connectivity Ntrain Fisher deff (log-det) gap L=2 L=2 L=2
none linear full
16 16 16
3.3 9.3 21.8
0.101 0.146 0.381
L=3 L=3 L=3
none linear full
16 16 16
4.0 29.5 36.5
0.101 0.440 0.431
The ordering none < linear < full holds for both deff and the gap at L=2 and L=3; a single entangling layer (L=1) leaves the gap unchanged, the depth-gating referred to in the main text.
Entanglement spreads the Fisher spectrum Fisher eigenvalue i
100 10 1 10 2 entanglement none (rank 2) linear (rank 4) full (rank 5)
10 3 10 4 1
2
3 4 eigenvalue index
5
Figure 8: Fisher eigenvalue spectrum of the trained circuits at L=2 (single-observable readout, log scale). The non-entangled circuit has effectively two non-zero eigenvalues, linear three, and full five or more: entanglement raises the rank of the Fisher matrix and hence the effective dimension deff = log det(I + γ F̂), exactly the quantity charged by the bound. This is the spectral content of Proposition 1.
19
Multi-step value-function generalization and controls. The value-function experiment (SecP5 t tion 6) fits a single-observable VQC to V ⋆ (x) = t=0 0.9 r(xt ) under the deterministic transi′ tion x = tanh(1.3 roll(x) + 0.2x) with r the interaction rule, by Adam regression (800 steps) on N states, reporting held-out R2 over 1000 states. The controls of Section 7 use N =64: matchedaccuracy early-stops at train accuracy 0.72; the readout control measures Z1 ; the optimizer control uses SGD (lr 0.5). The hardware run trains the N =16 classifiers in simulation and evaluates ⟨Z0 ⟩ for all 56 inputs on ibm aachen via EstimatorV2 at 4096 shots, optimization level 1. Hardware credibility: calibration, noise, and shot statistics. The device is IBM Heron ibm aachen (156 qubits). At the time of execution its median two-qubit gate error was 1.8 × 10−3 (typical range 7 × 10−4 to a few ×10−3 , excluding one inoperable pair) and median T1 ≈ 218 µs. Readout error mitigation was not applied—the effect we report survives raw device noise. Three cross-checks support the hardware numbers (Figure 6, right). (i) A noise-model simulator built from the ibm aachen calibration (AerSimulator with NoiseModel.from backend) reproduces the ordering (gap 0.14/0.33/0.40 for none/linear/full), i.e. the modelled noise does not remove √ the effect. (ii) Shot noise is small: at 4096 shots each ⟨Z0 ⟩ estimate has standard error ≈ 1/ 4096 ≈ 0.016, which flips a sign only for inputs within that margin of the decision boundary, so the accuracy-based gap is stable. (iii) The raw hardware gaps (0.14/0.33/0.31) agree with both the noiseless and noise-model values to within ≈ 0.1, and preserve the none < {linear, full} ordering. We report the hardware result as corroboration of the simulation evidence, not as a separate quantitative claim. Compute and reproducibility. All classical simulations use PennyLane default.qubit with the PyTorch interface (analytic gradients). Each supervised cell averages 8 seeds, each bandit/robustness/real-data cell 16/8/6 seeds, and the value-function and policy-gradient cells 8/10 seeds; error bars are one standard error. Targets: interaction rule sign(x1 x2 −x3 x4 + 21 x1 ) (main); Q P6 4-bit parity sign( i sign(xi )); random Fourier sign( k=1 cos(wk · x + bk )) with fixed wk , bk ; entangled-teacher labels sign(⟨Z0 ⟩) of a fixed 2-layer StronglyEntangling teacher. Real datasets are the scikit-learn Iris (versicolor vs. virginica), Breast Cancer Wisconsin, and Wine (classes 0/1) sets, standardized, reduced to four features by PCA when needed, and min–max scaled to [−1, 1].
20