Phase-cycled randomized benchmarking of quantum processors: recovering hidden classical noise correlations Mirza Samad Ahmed Baig,1, ∗ Syeda Anshrah Gillani,2, † Abdul Akbar Khan,3, ‡ and Muhammad Omer Khan4, § 1
arXiv:2609.06448v1 [cs.ET] 6 Sep 2026
2
Fandaqah, Al Khobar, Saudi Arabia Heidelberg University, Heidelberg, Germany 3 Argaam, Riyadh, Saudi Arabia 4 Fortanixor, United Arab Emirates
Randomized benchmarking can hide classical temporal correlations because its Clifford-twirled response is even in the noise phase. For a stationary symmetric telegraph fluctuator, we show that continuous evolution and independent stationary resets at slot boundaries yield identical mean responses for arbitrary fixed idle modulations. We construct an eight-setting phase-cycle measurement of the connected sine-phase covariance under ideal Clifford twirling and classical idle dephasing. This observable vanishes for independent slot noise and fixed detuning without a weak-phase or Gaussian approximation. A closed telegraph response, independent circuit calculations and 800 simulation trials validate the construction and quantify the empirical coverage of a paired bootstrap estimator. A separate conservative confidence set states its finite-sample assumptions. Two acquisitions on an IBM processor compare engineered shared-sign and independently reset phases with identical marginals. Their primary contrasts are 0.254 and 0.211, with empirical 95% intervals [0.177, 0.331] and [0.136, 0.285], respectively; all negative-control intervals include zero. At equal shot and sensingwindow budgets, an ideal Ramsey/echo estimator is more precise in every tested class. The result supplies an explicit connection between a benchmarking identifiability limitation and a controlled correlation measurement. No native or quantum-memory detection is claimed.
I.
INTRODUCTION
Randomized benchmarking (RB) estimates a compact description of gate performance by averaging random circuits [1–3]. The single exponential associated with gateindependent memoryless noise is attractive experimentally, but observing that shape does not establish the noise assumptions. Classical environments can preserve information across gates while producing the same averaged decay [4–6]. This distinction matters because a fitted average error is not itself a worst-case error certificate. Idle-duration scans are an established way to probe this problem. RB Ramsey and RB echo already measured dephasing versus idle duration, resolved telegraph noise, and separated linear and quadratic error contributions on superconducting qubits [7]. Engineered-noise RB experiments also measured and suppressed error correlations, with theoretical predictions tested on trapped ions [8]. Accordingly, neither insertion of an idle window nor calibration against injected correlated noise is claimed here as a new principle. Our question is narrower: which feature of the Clifford average hides a symmetric fluctuator, and can a small change of the measured observable recover the information it removes? We answer by separating two tasks that an idle scan can otherwise conflate. Measuring the duration dependence of the phase distribution within a slot is
∗ [email protected] † [email protected] ‡ [email protected] § [email protected]
a spectroscopy task. Distinguishing continuous environmental evolution from independent resets between slots is an identifiability task. Even an ideal, arbitrarily dense gap scan need not solve the second task. The contribution is a constructive pair of results. First, we establish an exact reset equivalence for a symmetric telegraph environment, including unequal idle durations. Second, controlled phase cycling converts the even Clifford response into a connected odd-phase correlator. The latter responds to inter-slot dependence while rejecting a fixed detuning and a reset model with identical marginal distributions. Both statements concern an explicitly specified dephasing model; neither is a universal classifier of classical versus quantum non-Markovianity. Existing analyses explain non-exponential RB under correlated dephasing [5, 9], temporal-noise effects with finite-duration gates [10], and learning correlated dynamics from RB data [11, 12]. Multi-exponential estimation has a general framework [13], and finite-sample RB statistics have substantial prior development [14]. Sine-phase correlations already appear in Ramsey noise spectroscopy [15]; phase cycling and correlation spectroscopy are established beyond RB [16], including the recent RESOLUTE protocol [17]. Our proposed contribution is the explicit Clifford-twirled implementation of a connected sine covariance, its constructive resetequivalence motivation, and its matched-marginal controlled test. Neither the existence of RB blindness nor sine-correlation spectroscopy itself is new. Section V B compares an ideal Ramsey/echo construction at equal circuit-shot and sensing-window budgets.
2 II.
WHAT A GAP SCAN CAN IDENTIFY
Consider ideal independent single-qubit Cliffords separated by dephasing slots. Only the Cliffords enter the final inverse. A slot accumulates a classical phase ϕk , independent of the chosen Cliffords. Averaging the Clifford sequence gives "m # Y 1 + 2 cos ϕ p(ϕ) = , Zm = E p(ϕk ) , (1) 3 k=1
where ASF(m) is the average sequence fidelity (survival probability) and Zm = 2ASF(m) − 1 for ideal preparation and measurement. In particular, Z0 = 1; Zm is a normalized contrast, not the survival probability itself. For an exponential field autocorrelation ⟨b(t)b(t′ )⟩ = ′ σ 2 e−|t−t |/τc , integration over an idle window δ yields v(δ) = 2σ 2 τc2 f (δ/τc ),
f (x) = x + e−x − 1.
(2)
This expression is shared by Gaussian exponential noise and symmetric telegraph noise with the same autocorrelation, although their complete phase distributions differ. The weak-phase single-slot error per Clifford (EPC) is v/6. Its local exponent crosses from 2 at δ ≪ τc to 1 at δ ≫ τc . For Gaussian phases specifically, EPC = (1 − e−v/2 )/3 is exact for the averaged single-slot channel. That formula does not make a multi-slot correlated Gaussian decay a single exponential, and finite phase saturation can move the observed exponent outside its weakphase limits. In hardware analysis we fit the phenomenological local model EPC(δ) ≃ ε0 + Aδ ν ,
ν=
d ln(EPC − ε0 ) . d ln δ
(3)
The offset must be fitted and the abscissa is δ, not gate duration plus δ. The fit is an effective description on a finite interval, rather than an exact power law through the full crossover. A quadratic response can arise from a deterministic frequency error as well as slowly fluctuating noise. A.
An exact reset equivalence
Proposition 1 Let b(t) = ±σ be a stationary symmetric random telegraph noise (RTN) process with flip rate γ = 1/(2τc ). Each slot has an arbitrary fixed real modulation of its coupling to this field, and ideal independent Clifford twirling produces the even factor p(ϕk ). Then Zm =
m Y
λk ,
λk = E[p(ϕk )].
(4)
To prove the statement, define the weighted transition matrix [Mk ]s′ s = E[p(ϕk )⊮send =s′ | sstart = s]. The telegraph dynamics are invariant under global sign reversal, and the weight is even. Therefore SMk S = Mk , where S swaps the two telegraph states. The one-dimensional symmetric subspace is invariant, so Mk π = λk π for π = (1, 1)T /2. Applying the matrices in order gives Eq. (4), including slots with different durations or modulation. Uncoupled environmental evolution between slots also preserves π. Independent stationary resets give the same product directly. ■ For identical slots the result reduces to Zm = λm . The stronger statement is that a scan over any set of durations remains identical to a slot-reset model. This reset model retains correlations within each slot; it is memoryless across slot boundaries, not necessarily a continuous-time Lindblad semigroup. Thus a gap scan can reject a whitenoise duration law without certifying inter-slot memory. No procedure using only these ideal mean contrasts distinguishes this particular pair. Sequence-resolved distributions may contain further information and are not covered by this no-go statement. The proof relies on a single symmetric two-state fluctuator and gate-independent coupling. Multiple fluctuators, unequal switching rates, or gate errors that depend on the environment need not share this equivalence. An even modulation of the same scalar phase cannot break it; changing the parity of the measured response can. The quasi-static equal-branch case is already within the blindness mechanism described in Ref. [6]. The construction above makes the continuous telegraph/reset pair explicit for arbitrary fixed slot modulations.
III.
PHASE CYCLING RECOVERS THE HIDDEN CORRELATOR
Apply a known virtual rotation Rz (sβ) in a slot, with s = ±1, keeping this rotation out of the Clifford inverse. The twirled factor becomes ps (ϕ) = a(ϕ) − sb(ϕ), 1 + 2 cos β cos ϕ a(ϕ) = , 3
2 sin β sin ϕ . 3
(5)
For two slots, let Zst = E[ps (ϕ1 )pt (ϕ2 )], and obtain the matched marginal responses zk,s = E[ps (ϕk )] from singleslot experiments. Stationarity or engineered replay must ensure that the single-slot law agrees with the corresponding marginal of the two-slot experiment. Proposition 2 For any classical joint distribution of (ϕ1 , ϕ2 ) obeying the dephasing and Clifford-independence assumptions above, define
k=1
The same contrasts result if the environment is independently reset to its stationary distribution before every slot while preserving each slot’s phase law.
b(ϕ) =
Jβ =
1 X st Zst , 4 s,t=±1
Cβ = Jβ − u1 u2 .
uk =
zk,+ − zk,− , 2 (6)
3 Then
IV.
Cβ =
4 sin2 β Cov(sin ϕ1 , sin ϕ2 ). 9
(7)
Substitution of Eq. (5) shows that the sum weighted by st removes the even and mixed-parity terms, leaving E[b(ϕ1 )b(ϕ2 )]. The single-slot differences are uk = −E[b(ϕk )]. Subtracting their product proves Eq. (7). ■ No Gaussian or small-angle approximation was used. A nonzero Cβ excludes independence of the two phases within this measurement model. A fixed detuning has zero covariance even when both marginal sine responses are nonzero; omitting the connected subtraction would fail this essential control. Independent resets also give zero, whereas a slowly fluctuating sign shared by both slots generally does not. The choice β = π/2 maximizes the coefficient, but also introduces a substantial known coherent rotation and is not an ordinary low-error EPC measurement. In particular, a revival under this controlled bias is not interpreted as evidence of quantum memory. The witness is sufficient but incomplete. Correlated distributions can have zero sine covariance, including phase aliases near integer multiples of π. Multiple values of the idle duration or coupling strength can address particular zeros but do not turn this second-moment witness into a universal dependence test. In the weak-phase limit, Cβ ≃ (4/9) sin2 β Cov(ϕ1 , ϕ2 ). A.
Closed response for continuous telegraph noise
For two equal windows of duration δ separated by an uncoupled interval h, define κ2 = γ 2 − σ 2 . Exact telegraph evolution gives 2 4 2 −h/τc −γδ sinh(κδ) Cβ (h) = sin β e σe . 9 κ
(8)
For imaginary κ, the bracket is evaluated with a real sine; at κ = 0 its continuous limit is σδe−γδ . The transfer derivation is given in Appendix A. In the quasistatic limit, Eq. (8) becomes (4/9) sin2 β sin2 (σδ). At fixed window duration and nonzero response, its dependence on an ideal uncoupled separation directly gives τc−1 = −d ln Cβ /dh. This last relation is a model prediction, not a claimed hardware correlation-time measurement. A bare delay on a physical qubit usually remains coupled to its environment; engineering an effectively uncoupled separation requires a separate control and its error characterization. The injected hardware experiment below tests the shared-sign limit and an independent reset, while finite switching times and the h dependence are validated numerically.
ESTIMATION AND FINITE-SAMPLE SCOPE
One block contains the four two-slot settings ++, +−, −+, −−, followed by marginal settings 1+, 1−, 2+, 2−. These share a Clifford pair and a programmed phase pair. Let Pi,j be the observed survival fraction in block i and setting j. The block quantities are Ji = (Pi,++ − Pi,+− − Pi,−+ + Pi,−− )/2, Ui = Pi,1+ − Pi,1− , Vi = Pi,2+ − Pi,2− .
(9)
Using distinct blocks for the product of means gives the unbiased estimator P P P ( i Ui )( i Vi ) − i Ui Vi b C=J− . (10) n(n − 1) Independence across blocks makes each i ̸= j product unbiased for EU EV , while arbitrary dependence within a block is allowed. This also removes the finite-sample covariance bias of simply multiplying two paired sample means. The practical interval resamples whole blocks, keeping phase settings and matched control arms together. A percentile bootstrap is an empirical approximation, not a finite-sample coverage theorem. We assess it on independently seeded simulation trials using a fixed decision rule: a two-sided nominal 95% interval must exclude zero. No threshold is optimized on these validation trials. A separate conservative confidence set is available because Ji , Ui , Vi ∈ [−1, 1]. For independent identically distributed blocks, Hoeffding’s inequality and a union bound imply that all three population means lie within p ϵn = 2 ln(6/α)/n (11) of their sample means with probability at least 1 − α. Intersect each mean interval with [−1, 1], then evaluate the minimum and maximum of j − uv over that rectangular set. The resulting interval covers Cβ at least 1 − α and need not be centered on Eq. (10). Its conservatism is reported separately from bootstrap performance. Independent blocks are a substantive assumption. Repeated observations of a slowly drifting native environment can violate it. Such observations require time ordering, an appropriate block length or independent acquisitions; the independent injected trajectories used in simulation do not establish coverage for arbitrary hardware drift. For an affine readout response Pobs = aP + b, phase differences cancel b. The joint term scales as a and the product of marginal differences as a2 . We therefore calibrate a from prepared zero and one states and divide each of J, U, V by the measured gain before Eq. (10). Hardware bootstrap replicates resample those calibration counts as well. The finite bound above applies to the ideal bounded observations; it is not asserted unchanged after estimating and dividing by readout gain. Stable,
4 TABLE I. Held-out simulation calibration. Entries are percentages over 200 trials per class. Detection means the nominal two-sided bootstrap interval excludes zero. The conservative bound uses Eq. (11). Model
Detection Bootstrap Finite bound coverage coverage Continuous RTN 100.0 92.5 100.0 Independent reset 4.5 95.5 100.0 Fixed detuning 6.5 93.5 100.0 Independent Gaussian 5.5 94.5 100.0
TABLE II. Root mean squared error for the same sine covariance in ideal simulation. Neither method has native drift or imperfect gates in this comparison. Model Continuous RTN Independent reset Fixed detuning Independent Gaussian
B.
setting-independent state preparation and measurement (SPAM) and gate errors remain experimental assumptions tested in part by the controls.
V.
A.
Ratio 2.89 1.85 4.22 2.59
A coherent Ramsey/echo pair measures the mean cosines of the sum and difference of two phases. Two single-slot Y -quadrature settings measure their sine means. Thus a four-setting ideal construction gives CR = 21 E[cos(ϕ1 − ϕ2 ) − cos(ϕ1 + ϕ2 )] − E[sin ϕ1 ]E[sin ϕ2 ] = Cov(sin ϕ1 , sin ϕ2 ).
(12)
Independent state-amplitude propagation verifies all four ideal responses. This comparator is an ideal Ramsey/echo construction motivated by established correlation spectroscopy [15, 17]; it is not a reproduction of the RESOLUTE experiment or its full control sequence. We simulate 200 trials per class, with 128 independent phase-pair blocks. RB uses eight settings and 8 shots per setting; Ramsey uses four settings and 16 shots. Both therefore use 8,192 circuit shots and 12,288 sensingwindow exposures per trial. They share the same sampled phase pairs and use the distinct-block connected estimator. RB estimates are multiplied by 9/4 so both target CR . Noise parameters match the preceding validation; shots and blocks differ. This compares shot and sensing-window resources, not gate duration or wall time. The Ramsey/echo estimate is more precise in all tested classes (Table II). The supplementary data retain all trial estimates and paired Monte Carlo uncertainty for each ratio. These results support the RB embedding as an interpretive construction, not an efficiency claim.
Additional assumption checks
An additional 200 randomly generated schedules use unequal slots with piecewise real modulation of either sign. Their continuous/reset mean discrepancy is below 10−14 . Testing 100 discrete, nonsymmetric joint phase laws and varying the cycle angle verifies Eq. (7) below the same tolerance. A deliberately asymmetric telegraph example gives a reset discrepancy 0.000852; the symmetry assumption cannot simply be dropped.
Ramsey/echo 0.0269 0.0432 0.0188 0.0318
An ideal resource-matched Ramsey comparator
NUMERICAL VALIDATION
The exact transfer calculation, closed response, and explicit random-Clifford Bloch propagation are implemented separately. A scan over 19 logarithmically spaced correlation times and three idle durations checks reset equivalence for lengths m = 0, . . . , 50. The maximum discrepancy is 4.66 × 10−15 . The closed form agrees with the eight transfer-matrix measurement means to 1.39×10−16 , including a separate critical-damping check. Enumeration of all 242 Clifford pairs agrees with the product-twirl identity to 3.33 × 10−16 for both weak and strong test angles. Figure 1 displays the exact response and sampling distributions. The sampling study uses 200 independent trials per class, 512 independent blocks per trial, and 64 binomial shots per setting. The fixed continuous-RTN parameters are σ = 1.6 rad/µs, τc = 2 µs, δ = 0.5 µs and h = 0.1 µs. Controls comprise independent resets with the same complete single-slot phase law, a deterministic phase, and independent Gaussian slot phases. The exact RTN covariance is 0.1704. Results are given in Table I; coverage variation makes the distinction between empirical and guaranteed intervals material.
RB 0.0777 0.0798 0.0793 0.0824
VI. A.
HARDWARE CONTROLS
A test that holds the slot marginals fixed
The phase-cycle design is saved before submission and uses a shared Clifford pair across four arms. With θ = 0.8 rad, the correlated arm draws a random sign and injects (ϕ1 , ϕ2 ) = (sθ, sθ). The reset arm draws two independent signs, yielding (sθ, tθ) with identical single-slot distributions. The deterministic arm injects (θ, θ), and the last arm has no injected phase. The ideal connected response is (4/9) sin2 θ for the correlated arm and zero for each control. In particular, ordinary unbiased RB means coincide for the first three arms because cos(θ) = cos(−θ).
5
(a) Identical gap scans connected contrast
10−2 10−3
.02
.05
.1
.2
idle duration (us)
.5
0.175 0.150 0.125 0.100 0.075 0.050 0.025 0.000
(c) Simulated estimates 0.20
continuous RTN reset / fixed detuning
connected contrast
continuous RTN independent slot resets
10−1
unbiased idle EPC
(b) Exact response
0
2
4
6
8
uncoupled separation h (us)
0.15 0.10 0.05 0.00 rtn
reset
static
white
FIG. 1. Simulation and exact theory. (a) Continuous telegraph noise and independent slot resets have identical mean gap scans. (b) The phase-cycle covariance distinguishes them; the separation dependence is the ideal uncoupled-interval prediction of Eq. (8). (c) Median and central 95% sampling ranges over independent simulated experiments, with crosses marking population values. These are sampling-distribution ranges, not confidence intervals from one experiment.
Each block contains eight settings. Each slot contains a 320 ns idle; the injected phase and the ±π/2 bias are virtual rotations protected by circuit barriers. Circuits use Qiskit [18]. The programmed rotations remain outside the Clifford inverse. Readout calibration adds 64 circuits for each prepared state. The processor is ibm fez q46. Initial submissions used individually serialized circuits. A subsequent version groups phase bindings into shared Clifford templates without changing their unitaries; the template order and the order of settings within each template are randomized. Sample counts and acquisition outcomes are reported below. Randomization reduces order confounding but does not itself prove stationarity. The primary comparison, chosen before collection, is the paired difference between the correlated and reset contrasts. Per-arm intervals additionally check the fixed detuning and no-injection controls. Injected phases are held fixed over each circuit’s shot group; native fluctuations are not replayed between distinct circuits. Thus the experiment calibrates recovery of engineered interslot dependence rather than detecting a native hardware memory process.
B.
Injected phase-cycle results
The archive contains 2 completed acquisitions. Acquisition 1 uses 128 programmed blocks per arm and 8 shots per setting (4,224 circuit evaluations, 33,792 circuit shots), with random seed 2026090501 and readout gain 0.9980. Acquisition 2 uses 128 programmed blocks per arm and 8 shots per setting (4,224 circuit evaluations, 33,792 circuit shots), with random seed 2026090502 and readout gain 0.9883. The second acquisition was specified after inspecting the first result, with a fresh seed and the same block count, shots, phase amplitude, ob-
TABLE III. Measured injected phase-cycle contrasts and empirical 95% intervals. The ideal population contrast is 0.2287 for shared signs and zero for the other arms. Arm Shared sign Independent reset Fixed detuning No injection Shared sign Independent reset Fixed detuning No injection
Estimate Acquisition 1 0.2580 0.0040 0.0456 0.0416 Acquisition 2 0.1704 -0.0409 -0.0119 -0.0154
95% interval [0.1935, 0.3182] [−0.0704, 0.0779] [−0.0304, 0.1195] [−0.0410, 0.1265] [0.1091, 0.2272] [−0.1053, 0.0249] [−0.0769, 0.0495] [−0.0874, 0.0568]
servable and controls. Both acquisitions used the same qubit on the same day; they test within-session repeatability, not robustness across devices or calibration cycles. Intervals use 4,000 paired block bootstrap replicates and resample readout calibration counts. They are empirical intervals conditional on the stability assumptions in Sec. IV. The conservative coverage guarantee for bounded uncorrected observations is not transferred to these calibration-corrected hardware intervals. The primary correlated-minus-reset contrast in acquisition 1 is 0.2540, with nominal 95% interval [0.1775, 0.3310]. It excludes zero in the predicted direction. The primary correlated-minus-reset contrast in acquisition 2 is 0.2114, with nominal 95% interval [0.1359, 0.2851]. It excludes zero in the predicted direction. All individual negative-control intervals contain zero. In acquisition 2, the shared-sign interval does not contain the ideal population prediction 0.2287. The positive separation therefore does not imply exact quantita-
6 D. Ideal Acquisition 1 Acquisition 2
connected contrast
0.3
0.2
0.1
0.0
−0.1 shared sign
reset
fixed
none
FIG. 2. Hardware response to engineered phase correlations. Points and error bars are connected estimates and paired block bootstrap intervals for each acquisition; crosses are ideal population predictions.
tive agreement with an ideal device. All arms appear in Table III and Fig. 2. The archive also retains 3 unsuccessful jobs with no usable counts and 45 seconds of recorded charges. They are excluded from the estimator and not counted as replications. Reductions in sample count and Runtime packaging preceded the first successful measurement. The completed results are engineered-noise proofs of concept; they establish neither native memory nor a physical correlation-time scan.
C.
Complementary idle-gap calibration
The earlier campaign used synthetic dephasing trajectories with known correlation times in idle windows. Effective exponents on ibm fez q46 were ν = 1.00 ± 0.10, 1.51 ± 0.08, and 1.72 ± 0.15 for fast, crossover and quasistatic arms, respectively. The largest discrepancy from the stored response prediction was 1.61 standard errors. These are calibration observations for Eq. (3), not independent evidence establishing priority or universal classifier performance. Several limitations affect that legacy analysis. The baseline-subtracted EPC is a leading-order approximation: even independent depolarizing contractions compose multiplicatively, so rcombined = r0 + ri − 2r0 ri and subtraction does not cancel background exactly. The stored analysis uses data-dependent signal-to-noise cuts and separately estimated EPC uncertainties rather than a complete paired joint fit. Its interval coverage under those choices has not been established. The quasistatic Hahn control demonstrates refocusing of the programmed phase; because a slot-integrated phase is distributed proportionally between echo segments, it does not validate the full within-slot switching response of finite-time telegraph noise. These observations motivate the more direct matched-marginal phase-cycle test. Repeated-sequence variance removes a separate confound: different Clifford sequences can have different intrinsic fidelities even with stationary noise. The injected rep = 4.02 and native controls quasi-static arm gave Wvar near unity. Nevertheless, this statistic probes changes between repeated circuit executions, a different timescale from memory between gates. Shot groups that average many independent trajectories may suppress it. It is a diagnostic, not a necessary condition for all inter-gate memory.
Readout calibration sensitivity
A zero-error calibration count is not evidence of a zero population error rate. To assess this boundary, we construct two exact binomial intervals at 97.5% per acquisition and combine them into a simultaneous 95% interval for the readout gain by a union bound. With experimental survivals held fixed, we recompute the primary contrast over a dense gain grid. For acquisition 1, this calibration-only range is [0.2536, 0.2590]. For acquisition 2, this calibration-only range is [0.2093, 0.2171]. These ranges do not include experimental sampling error and are not confidence intervals for the primary contrast. They indicate that the observed positive point estimates do not depend on setting unobserved calibration errors exactly to zero. They do not test drift or setting-dependent SPAM.
VII.
NATIVE SURVEY AND LIMITS OF A QUANTITATIVE NULL
The earlier screen covered 16 qubits across ibm fez and ibm marrakesh. Two gap-exponent candidates did not reproduce on follow-up. On ibm fez, q53 changed from 1.87 ± 0.22 to 1.26 ± 0.14. On ibm marrakesh, q154 changed from 2.63 ± 0.54 to 1.56 ± 0.28. These measurements support the statement that there was no replicated detection under the screening rule. Regression after selecting a maximum is compatible with selection effects; temporal drift is another explanation, and the data do not uniquely attribute the changes to either cause. A stored six-gap EPC curve on q46 has the effective exponent 1.05 ± 0.05. To ask whether it supports an exclusion, consider δ g(δ/τc ) r(δ) = ε0 + a (1 − f ) +f , (13) δmax g(δmax /τc ) where g(x) = x + e−x − 1, a ≥ 0, ε0 ≥ 0, and 0 ≤ f ≤ 1.
7 The fraction is defined at the largest gap. At every proposed (f, τc ), we refit both nuisance parameters by weighted nonnegative least squares. Holding them at their previously fitted values understates their uncertainty. With the stored diagonal EPC error model, the minimum residual χ2 ranges from 15.81 to 18.97 over the 13 tested correlation times 20–200,000 ns. All exceed 12.59, the 95th percentile of χ26 . If the six-dimensional Gaussian covariance were known, the ellipsoid (y−µ)T Σ−1 (y−µ) ≤ χ26,0.95 would cover the true mean with probability 95%. An empty intersection with the proposed model family on this grid signals model incompatibility, rather than exclusion of every possible correlated contribution. Here the covariance is itself estimated and omits unavailable cross-gap covariance, so this is an adequacy diagnostic. Consequently, we do not quote an unconditional 95% correlated-noise exclusion from these summaries. A profile-likelihood threshold of ∆χ2 = 2.71 does not repair a poor absolute fit, establish finite-sample coverage, or provide a simultaneous bound over correlation times. A quantitative exclusion requires reanalysis of raw blocks with covariance, checks of model adequacy, and coverage calibration for the resulting procedure. Figure 3 displays this distinction explicitly.
VIII.
DISCUSSION
The reset equivalence explains why a useful durationdependent noise measurement can still leave inter-slot dependence unidentified. Phase cycling changes the observable rather than merely improving the precision of the same even response. The connected subtraction is essential: it separates shared random phase from a deterministic frequency shift, while the reset control tests specificity at fixed single-slot noise statistics. The current scope is classical idle dephasing with Clifford-independent phases. Finite gate durations, gatedependent error, leakage, nonstationary SPAM and unmatched single-slot marginals can all invalidate the simple interpretation. The fixed-detuning and no-injection controls test some of these effects but cannot establish their absence under all operating conditions. The exact telegraph formula applies to ideal uncoupled separation; demonstrating a physical lag scan requires additional control validation. Only one- and two-slot circuits are needed, but multiple settings and independent blocks are required. The ideal resource-matched comparison in Sec. V B favors Ramsey/echo estimation in every tested noise class. The present method therefore offers an explicit connection to the Clifford RB observable, rather than a demonstrated precision advantage. It does not inherit ordinary RB’s fitted separation of gate decay and arbitrary SPAM coefficients. Repeated acquisitions across calibration epochs and more general gate-error models are needed to assess its experimental generality.
FIG. 3. Audit of the stored native EPC summaries. (a) Offset-aware phenomenological fit. (b) Best residual after profiling nonnegative nuisance parameters, versus the Gaussianball cutoff. The tested model curves fail this adequacy diagnostic, so no exclusion region is shaded or claimed.
The finite-sample construction also separates what is proved from what is measured. The conservative meanrectangle confidence set has a coverage statement under independent bounded blocks. The bootstrap is more precise but its coverage is empirical. Neither implies quantum-memory certification or a worst-case gate-error bound. Similarly, failure to replicate a native screening candidate is a useful negative result without implying an accurately calibrated exclusion fraction.
IX.
CONCLUSION
For a symmetric telegraph environment, mean gapscan RB remains exactly unchanged when memory across slot boundaries is erased. A controlled phase cycle recovers the sine-phase covariance that the ordinary even Clifford response hides. The exact identity supplies explicit independent-noise and deterministic-detuning nulls, a telegraph response function, and a connected estimator with separately stated empirical and conservative uncertainty assessments. The hardware controls and the
8 native-data audit distinguish an operationally testable memory observable from stronger claims that the present data do not support.
Feynman–Kac propagation of a phase characteristic function over an idle window is Ek = exp[(Q + ikD)δ], for k = −1, 0, 1. The biased one-slot weighted matrix is
DATA AND CODE AVAILABILITY
The accompanying repository contains the data, analysis code and reproduction instructions. Its inventory distinguishes observed hardware counts, simulated data, failed acquisitions and legacy summaries; four older datasets remain summary-only. Code uses a custom source-available license requiring attribution in public work substantially using it, and original data and figures use CC BY 4.0. See the repository attribution and citation files for the authors and license scope. The companion project repository is quantum-phase-cycled-rb.
Mβ = eQh
E0 + eiβ E1 + e−iβ E−1 . 3
(A2)
For equal biases, Zm = 1T Mβm π; distinct phase signs are obtained by ordered multiplication of the corresponding matrices. At zero bias the symmetric stationary subspace is invariant. The sine component Sδ = (E1 − E−1 )/(2i) instead exchanges the symmetric and antisymmetric subspaces. In their normalized basis its off-diagonal entry is
ACKNOWLEDGMENTS
We acknowledge the use of IBM Quantum services. The views expressed are those of the authors and do not reflect the official policy or position of IBM or the IBM Quantum team.
R = σe−γδ
sinh(κδ) . κ
(A3)
Appendix A: Telegraph response derivation
(A1)
The antisymmetric component decays by e−2γh between windows. Stationary reflection symmetry gives E sin ϕk = 0, while E[sin ϕ1 sin ϕ2 ] = R2 e−2γh . Substitution into Eq. (7) proves Eq. (8).
[1] J. Emerson, R. Alicki, and K. Życzkowski, J. Opt. B: Quantum Semiclass. Opt. 7, S347 (2005). [2] E. Knill, D. Leibfried, R. Reichle, J. Britton, R. B. Blakestad, J. D. Jost, C. Langer, R. Ozeri, S. Seidelin, and D. J. Wineland, Phys. Rev. A 77, 012307 (2008). [3] E. Magesan, J. M. Gambetta, and J. Emerson, Phys. Rev. Lett. 106, 180504 (2011). [4] H. Ball, T. M. Stace, S. T. Flammia, and M. J. Biercuk, Phys. Rev. A 93, 022303 (2016), arXiv:1504.05307. [5] J. Qi and H. K. Ng, Phys. Rev. A 103, 022607 (2021), arXiv:2010.11498. [6] V. Srivastava, A. K. Roy, S. Mahanti, J. Kaur, S. Karuvade, and A. Gilchrist, Phys. Rev. Research 8, 023258 (2026), arXiv:2510.13051. [7] P. J. J. O’Malley, J. Kelly, R. Barends, B. Campbell, Y. Chen, Z. Chen, B. Chiaro, A. Dunsworth, A. G. Fowler, I.-C. Hoi, E. Jeffrey, A. Megrant, J. Mutus, C. Neill, C. Quintana, P. Roushan, D. Sank, A. Vainsencher, J. Wenner, T. C. White, A. N. Korotkov, A. N. Cleland, and J. M. Martinis, Phys. Rev. Applied 3, 044009 (2015), arXiv:1411.2613. [8] C. L. Edmunds, C. Hempel, R. J. Harris, V. Frey, T. M. Stace, and M. J. Biercuk, Phys. Rev. Research 2, 013156 (2020).
[9] M. A. Fogarty, M. Veldhorst, R. Harper, C. H. Yang, S. D. Bartlett, S. T. Flammia, and A. S. Dzurak, Phys. Rev. A 92, 022326 (2015), arXiv:1502.05119. [10] A. Brillant, P. Groszkowski, A. Seif, J. Koch, and A. A. Clerk, Phys. Rev. Lett. 135, 070601 (2025), arXiv:2501.06172 [quant-ph]. [11] S.-X. Yang, P. Figueroa-Romero, and M.-H. Hsieh, Machine learning of average non-Markovianity from randomized benchmarking (2022), arXiv:2207.01542 [quant-ph]. [12] X. Zhang, Z. Wu, G. A. L. White, Z. Xiang, S. Hu, Z. Peng, Y. Liu, D. Zheng, X. Fu, A. Huang, D. Poletti, K. Modi, J. Wu, M. Deng, and C. Guo, Commun. Phys. 8, 29 (2025), arXiv:2312.06062. [13] J. Helsen, I. Roth, E. Onorati, A. H. Werner, and J. Eisert, PRX Quantum 3, 020357 (2022), arXiv:2010.07974. [14] J. J. Wallman and S. T. Flammia, New J. Phys. 16, 103032 (2014), arXiv:1404.6025. [15] F. Yan, J. Bylander, S. Gustavsson, F. Yoshihara, K. Harrabi, D. G. Cory, T. P. Orlando, Y. Nakamura, J.-S. Tsai, and W. D. Oliver, Phys. Rev. B 85, 174521 (2012). [16] J. Meinel, V. Vorobyov, P. Wang, B. Yavkin, M. Pfender, H. Sumiya, S. Onoda, J. Isoya, R.-B. Liu, and J. Wrachtrup, Nature Communications 13, 5318 (2022).
Let Q=
−γ γ , γ −γ
D = diag(σ, −σ).
9 [17] I. Zohar, S. Oviedo-Casado, A. Denisenko, R. Stöhr, and A. Finkler, Ramsey correlation spectroscopy with phase cycling using a single quantum sensor (2026), arXiv:2603.05650 [quant-ph].
[18] A. Javadi-Abhari, M. Treinish, K. Krsulich, C. J. Wood, J. Lishman, J. Gacon, S. Martiel, P. D. Nation, L. S. Bishop, A. W. Cross, B. R. Johnson, and J. M. Gambetta, Quantum computing with Qiskit (2024), arXiv:2405.08810 [quant-ph].