ConceptioArchivearXiv CS
arXiv CSopen access

Claim against Measurement: Statistical Artefacts in Quantum Error Mitigation Benchmarks

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
softwarearchitecturesoftwareengineeringtesting
software engineering, software architecture, testing

arXiv:2605.29872v1 [quant-ph] 28 May 2026

Claim against Measurement: Statistical Artefacts in Quantum Error Mitigation Benchmarks Dominik Koester

Wolfgang Mauerer

Technical University of Applied Science Regensburg Regensburg, Germany [email protected]

Technical University of Applied Science Regensburg Siemens AG, Foundational Technology Regensburg/Munich, Germany [email protected]

Abstract—Quantum Error Mitigation (QEM) is widely regarded as a plausible bridge from Noisy Intermediate Scale Quantum (NISQ) devices to Fault Tolerant Quantum Computers (FTQC). Yet the empirical studies used to assess the effectiveness of QEM techniques on concrete problems have received comparatively little scrutiny with respect to the validity of their conclusions. We systematically review 81 recent QEM papers using an eightcriterion framework covering statistical rigour, reproducibility, and reporting quality. Among the 59 papers for which statistical evidence is applicable, only 15 (25%) use inferential methods, while 25 (42%) report uncertainty only descriptively, without testing whether the claimed effects are statistically supported. To demonstrate the consequences of these omissions, we use Zero-Noise Extrapolation (ZNE) as a representative and widely used case study and identify two compounding sources of artefacts in current QEM benchmarks. First, we observe parameter sensitivity: in a 132-configuration sweep, implicitly assumed choices such as scale factors, extrapolation method, and hardware calibration are not merely incidental but active, with variations changing conclusions from statistically significant improvement to statistically significant degradation. Second, we identify a driftinduced effectiveness illusion: in a 72-hour longitudinal study on real hardware, temporal drift alone can make the same ZNE configuration exhibit an effect size more than three times as large, depending solely on when it is executed, and also drastically reduces the effective number of independent observations. These findings do not imply that QEM methods are intrinsically unsound; rather, they show that current evaluation practice can make mitigation performance appear more robust than the evidence warrants. We therefore propose minimum reporting standards for QEM evaluations, including explicit parameter documentation, robustness checks, longitudinal drift assessment, and inferential statistical testing with effect-size reporting. Index Terms—quantum error mitigation, benchmarking, statistical artefacts, zero-noise extrapolation, hypothesis testing, reproducibility, NISQ

I. I NTRODUCTION We remain in the NISQ era of quantum computing [1]–[4], still some distance from FTQC architectures with sufficiently many qubits, low error rates, and operational Quantum Error Correction (QEC) [5], [6]. Although QEC is clearly the long-term route to reliable quantum computation, present implementations still incur substantial qubit overheads [7], [8]. In this intermediate regime, QEM has emerged (among hardware-efficient problem formulations [9]–[11] or target design automation techniques [12], [13]) as a pragmatic

approach for reducing the impact of noise [14] on current NISQ devices without implementing full error correction [15]– [18]. Techniques such as ZNE [15], [16], Probabilistic Error Cancellation (PEC) [15], [19], and Clifford data regression [20] have shown promise in improving the accuracy of quantum computations. Beyond dedicated benchmarking studies, QEM techniques are increasingly adopted as standard components in broader quantum computing research – from variational quantum eigensolvers and quantum structure simulations [21], [22] via optimisation [23]–[27] and machine learning [28]– [31] to quantum simulation of many-body systems [32] and demonstrations of evidence for quantum utility [33] – where framework-default parameters are silently inherited and their influence is not examined. Artefacts arising from such implicitly assumed parameter choices therefore propagate beyond QEM evaluation into any higher-level result that relies on mitigation as a building block. However, the empirical evaluation of QEM techniques themselves frequently rests on experiments whose statistical foundations are not always commensurate with the strength of the conclusions drawn from them. Previous work has already identified important statistical challenges in the evaluation of QEM techniques, including the need for explicit hypothesis testing [34]. In this work, however, our systematic review of 81 recent QEM papers (Section IV) shows that such concerns remain largely unaddressed in practice: most papers rely on descriptive reporting rather than inferential statistical evidence. This is not merely a matter of presentation, but one with direct experimental consequences. Reported QEM improvements are often small, since current NISQ hardware constrains experiments to shallow circuits and shot budgets limited by cost, placing many studies in regimes where the true effect is modest at best [35], [36]. In such settings, shot noise, hardware drift, and implicitly assumed parameter choices can each be sufficient to alter the experimental conclusion. Without the use of careful inferential testing, a comprehensive awareness of influence factors and boundary conditions, as well as the adequate availability of all ingredients required to perform faithful reproductions or replications, neither the original study nor any subsequent replication can reliably distinguish genuine mitigation effects from statistical or experimental artefacts. In this paper, we identify two compounding sources of

artefacts in QEM evaluation, each independently capable of III. BACKGROUND producing misleading results. First, a systematic replication A. Zero-Noise Extrapolation (ZNE) study shows that documented parameters usually represent only ZNE, first introduced by Li and Benjamin [16] and Temme et a small portion of the full space of possible parameters for an al. [15], is a widely used QEM technique that estimates the experiment: the choices experiments do not explicitly specify – noise-free result of a quantum computation by extrapolating such as scale factors, folding strategy, or calibration snapshot – results obtained at amplified noise levels. Given the noise are active, meaning their variation can shift the interpretation scale λ, the result of the smallest error rate on a circuit is of the outcome of an experiment from significant improvement given by the expectation value E(λ1 ) [18], [45]. By artificially to significant worsening compared to an unmitigated baseline. increasing the noise to levels λ1 < λ2 < · · · < λK – called Second, a longitudinal hardware study reveals that temporal scale factors – we can fit a model to extrapolate the noisedrift alone can produce large variation in apparent ZNE free expectation value E(0). Common amplification strategies effectiveness on the same device at different access times; include pulse stretching [15], [21], unitary folding [46], and given the dominant use of cloud services with indeterministic gate-level folding [47]. To compute the extrapolated value, batch access for many quantum experiments, this presents a besides standard approaches like linear [16], polynomial substantial challenge for both reproducibility and interpretation of empirical results. We focus on ZNE with Richardson and exponential extrapolation [17], [45], a technique called extrapolation as a representative case, as it is the most widely Richardson extrapolation is a commonly employed method [15], reported QEM technique in our corpus, and propose a minimum [17], [39], [46] that fits a polynomial of degree K−1 through K data points. is in this case estimated [18] as reporting checklist for QEM evaluations. PKThe zero-noise limit P K Ê(0) = c E(λ ), where k k=1 k k=1 ck = 1, and coefficients Concretely, our contributions are as follows: ck are determined by Lagrange interpolation. PK • We provide a systematic eight-criterion review of 81 Since Var(ÊZNE ) = k=1 c2k Var(Ê(λk )), a convenient preQEM papers, revealing that the majority of papers lack PK experiment bound on variance amplification is k=1 |ck | [18], inferential statistics and drift control. [39] – computable from the scale factors alone, without running • We introduce a replication pipeline and case study demonstrating that implicitly assumed ZNE parameters are active: any PKcircuits. For the widely used default set {1, 3, 5} [45], [46], k=1 |ck | = 3.5, whereas {1, 1.1, 1.25, 1.5} – motivated by their variation shifts experimental outcomes across a large arguments fraction of tested configurations. PKthat finer spacing reduces extrapolation error [39] – yields k=1 |ck | = 681: a 194× difference in the variance • A longitudinal drift study shows that temporal hardware drift produces a drift-induced effectiveness illusion, where bound. Scale factor choice alone can therefore dominate the apparent ZNE effectiveness varies substantially across variance budget of a ZNE experiment – motivating its treatment as an active parameter in Section VI. identical experiments at different times. • We derive minimum reporting standards for QEM benchB. Statistical Methods marks to avoid incorrect claims and misinterpretations of Our approach, implemented in a reproducible pipeline QEM experiments. (link in PDF) [48], relies on standard tools from frequentist II. R ELATED W ORK statistics to quantify how much a QEM technique improves Cai et al. [18] provide a comprehensive review of QEM performance. Unlike domains such as clinical trials or psytechniques, including their theoretical foundations, practical chology, QEM evaluation has no established domain-specific implementations, and performance analysis. Takagi et al. [37] statistical standards: there is no agreed effect-size threshold derive fundamental sampling-overhead bounds under global for claiming practical mitigation, no canonical hypothesis test, depolarising noise, generalised by Quek et al. [38] for local and no prescribed power requirement. In the absence of such noise. Krebsbacher et al. [39] derive variance bounds for conventions, we adopt the well-validated general-purpose tools Richardson extrapolation, guiding scale factor selection to from empirical science [49]–[51], which are commonly used in minimise amplification. On the statistical side, Saki et al. [34] related empirical fields. Despite frequentist hypothesis testing propose a hypothesis-testing framework for QEM evaluation, (e.g., paired t-tests and p-values) and effect-size estimation (e.g., while Li et al. [40] survey the use of statistical methods Cohen’s d) being standard, well-validated tools in empirical in quantum software testing, finding that formal statistical science, these methods remain rare in quantum computing. As methods are similarly rare in that domain. Moguel et al. [41] they are textbook knowledge [50], [51], we only briefly review review quantum benchmarking methods, proposing a quantum their essential properties below, especially to fix notation of experiment guideline extending established best practices the approaches in the context of quantum error mitigation. from classical software benchmarking. On reproducibility, a) Paired t-test: A paired t-test compares differences Senapati et al. [42], [43] highlight the reproducibility challenges between two groups: Given n independent repetitions producing in quantum machine learning, being device variability and raw errors |ϵraw,i | and mitigated errors |ϵmit,i |, the paired differtemporal drift. Hirasaki et al. [44] demonstrate temporal ence δi = |ϵraw,i | − |ϵmit,i | captures per-repetition improvement. fluctuations in superconducting qubits, yielding different mea- The paired √ t-test evaluates the null hypothesis H0 : µδ = 0 via surements over time points. t = δ̄/(sδ / n). We classify each configuration as significantly

better (p < 0.05, d > 0), not significant (p ≥ 0.05), or significantly worse (p < 0.05, d < 0) based on a 5% significance level (while using a prescribed explicit significance level is known to be exhibit a number of issues [52], we stick to this approach as it is common practice in the considered references, and conclusion stability can only be assessed when the same approach as in the original work is employed). As a non-parametric alternative that does not assume normality of the paired differences, we compute the Wilcoxon signed-rank test for every configuration. b) Effect-Size Measures: We use two complementary effect-size quantities. Cohen’s d = δ̄/sδ [49] is the standard effect-size measure in empirical sciences such as clinical trials or psychology, with well-established conventions |d| = 0.2/0.5/0.8 (small/medium/large) [49] and 1.2/2.0 (very large/huge) [53]; positive d means QEM reduces, negative d means it increases error. These thresholds were, however, calibrated for domains where effect sizes are bounded by natural variability rather than a controllable experimental parameter. In √ shot-based quantum experiments, sδ ∝ 1/ nshots , so d scales with shot count: at nshots = 4096, a modest improvement of ∆ = 0.18 already yields d ≈ 6, far above the established thresholds. Absolute d values are not comparable to those conventions and we use d primarily as a relative metric for cross-configuration comparison. Because no single effect-size measure is universally agreed on for QEM evaluation, we additionally report Cliff’s δ = (N> − N< )/n – where N> and N< count paired differences δi > 0 and δi < 0 – as a non-parametric, distribution-free alternative in [−1, +1] that is insensitive to this shot-count inflation and directly reflects the probability of a genuine directional improvement across repetitions. IV. S YSTEMATIC R EVIEW

TABLE I E IGHT- CRITERION ANALYSIS FRAMEWORK . E ACH PAPER IS RATED ADEQUATE , PARTIAL , MISSING , OR N / A .

ID

Criterion

Key Question(s)

C1 C2 C3

Sample Size Variance Stat. Evidence

C4 C5 C6 C7 C8

Drift Control Overhead Noise Model Reproducibility Neg. Results

How many shots/circuits? Size justified? Error bars, CIs, or variance reported? Inferential or descriptive evidence for claims? Temporal hardware drift accounted for? Classical/quantum overhead quantified? Noise model validated or discussed? Code, data, or sufficient detail? Failure cases or limitations reported?

human effort and comprehensiveness/completeness of coverage. Two raters jointly reviewed all 81 papers against the criteria. Two additional raters each independently rated a (different) subset of 15 papers. Disagreements among the raters were then resolved by consensus discussion. Additionally, we established a comparison baseline using automatic text processing: A regular expression scanner matched textual evidence against patterns targeting key statistical terms – for example, p-value reporting phrases, mentions of confidence intervals, or software repository URLs – and an LLM (Claude Opus 4.6) filtered candidate ratings for known false positives such as physical noise probabilities misidentified as statistical p-values, or framework citations misidentified as reproduction packages. Automated ratings agreed with the human consensus in 77% of applicable paper-criterion pairs. Full scan logs, rating sets, and a per-paper LLM evidence report are provided in the reproduction package. Overall, we believe the approach provides a sound and sufficiently large-scale basis that leads to generalisable conclusions.

To evaluate the current state of statistical reporting in QEM related papers and to identify common pitfalls and areas Missing Partial Adequate Rating for improvement, we systematically reviewed 81 publications published over half a decade (2022–2026), collected via Google Scholar, arXiv, IEEE digital libraries, and forward/backward 1. Sample Size citation tracking from the review by Cai et al. [18]. We 2. Variance used queries like “quantum error mitigation”, “zero-noise extrapolation”, “probabilistic error cancellation” in combination 3. Stat. Evidence with terms like “algorithm”, “drift”, or “experiment” to identify 4. Drift Control relevant papers (the full list is provided in the reproduction 5. Overhead package). Each paper was evaluated based on eight criteria 6. Noise Model detailed in Table I that assesses the presence and quality of statistical reporting in the experimental evaluation of QEM 7. Reproducibility techniques. Each criterion is rated as adequate (clear, specific 8. Neg. Results evidence), partial (mentioned but incomplete), missing, or 0% 25% 50% 75% 100% not applicable. The criteria themselves are derived from a Percentage of applicable papers priori expert knowledge, typical desiderata in the literature, and requirements documented in reviews [54]–[56] or existing Fig. 1. Summary of the systematic review results across the eight criteria. work on quantum reproducibility [48], [57], [58]. a) Review process: All raters are among the paper authors; Figure 1 shows the observed compliance of the 81 papers the process follows established patterns for lightweight semi- with each of the eight criteria. Five criteria are well addressed, formal reviews [59], and aims at balancing required manual with over 60% adequate reporting: Sample Size (C1, 77%),

Variance (C2, 77%), Noise Model (C6, 78%), Overhead (C5, 66%), and Negative Results (C8, 66%). Reproducibility (C7, 55%) is in the middle, but plays a crucial role as it is often the only way to confirm and build upon reported results. The two criteria with the lowest compliance are Drift Control (C4, 30%) and Statistical Evidence (C3, 25%). Of the 59 applicable papers for statistical evidence, only 15 employ inferential statistical methods: hypothesis tests, Bayesian inference, bootstrap-based comparisons, or scaling analysis with uncertainty quantification. The remaining 42% report uncertainty only descriptively (error bars, standard deviations, improvement factors), and 19 papers (32%) do not provide statistical evidence. The “partial” rating on C3 deserves explanation and credit to the respective studies. These papers report variance, quantitative improvement metrics, accuracy thresholds, or bootstrap error bars, all forms of descriptive statistical evidence common in quantum experiments, but do not use them for inferential comparison. QEM improvements are often small [35], [36]: in regimes where noise, drift, and parameter choice can each be sufficient to shift the outcome of an experiment, simple descriptive evidence alone is insufficient to confirm genuine effects. The two empirical analyses below demonstrate what practical consequences arise from a failure – as is omnipresent in the literature – to provide more complete descriptions and a more rigorous statistical analysis. Both artefact sources are independently capable of producing misleading evaluations, and the two factors compound when present simultaneously. V. T HE R EPRODUCTION PARAMETER S PACE & E XTRACTION P IPELINE sig. better

sig. worse

not sig.

Full parameter space P Replication’s search region {θdoc }× Range(θundoc )

Original experiment’s configs

Reproducibility risk

Fig. 2. The full parameter space P of a QEM experiment, with regions of improvement (green), no improvement (grey), and worsening (red). Studies typically test a small subset of parameter configurations (orange).

Our analysis shows C3 and C4 have the smallest compliance, yet matter most when parameter choices and timing can

influence or even determine the outcome of an experiment. Before we demonstrate this experimentally, let us formalise the structure of the problem. Figure 2 illustrates the relation between a quantum (software) experiment and an attempt to reproduce or replicate the findings. Given the overall variability and currently fast-paced change in quantum hardware, rapidly changing software environments [55], [58], [60], and other factors, an exact reproduction is rarely possible even if the original software artefacts are available, which leads to an effectively different set of empirical parameters. The space of relevant parameters for quantum experiments can be divided into three categories of outcomes for a given configuration: regions of significant improvement, regions of significant worsening, and regions with no statistically significant change. In typical experiments, only a specific subset of possible values is tested per parameter (highlighted as the orange circle), where some are explicitly specified (θdoc ) and others are left unspecified (θundoc ). A parameter is inert if varying it does not change the experimental conclusion, and active if it does. A. Parameter Space P A QEM experiment requires many choices beyond the circuit and observable. We formalise the reproduction parameter space for a ZNE experiment as: P = H × C × Q × F × E × S As |P| > 104 , and the size of parameter space for reproductions can easily grow to hundreds of reasonable configurations. TABLE II R EPRODUCTION PARAMETER SPACE .

Axis

Name

Examples

H

Hardware/Backend

C Q F E

Circuit Shots & reps Folding Extrapolation

S

Scale factors

QPU vendor/model, noise model, calibration snapshot Qubit count, depth, transp. level Shot count, number of repetitions Local-left, local-right, global Linear, polynomial, exponential, Richardson {1, 3, 5}, {1, 2, 3}, {1, 1.5, . . . , 3}

B. Reproduction Pipeline Following established terminology, we distinguish reproduction, where a different team re-runs the same experimental artefacts from replication, where a different team uses a different experimental setup, yet addresses the same question as an existing piece of research. Our case study addresses both aspects: we use the original circuit specification but systematically vary the unspecified parameters across a wider range than explored in the original study. To systematically explore the parameter space, we apply a four-stage pipeline to a target paper: (1) Parameter Extraction: extract all documented parameters θdoc and identify unspecified ones θundoc . Also categorise any parameters θamb that are either ambiguously mentioned or very likely given the context, but not explicitly stated. (2) Parameter-Space Sampling: sample θundoc and θamb across reasonable values. Reasonable is contextdependent, but in most cases, refers to the most commonly

used values in literature. The result is a set of configurations {p1 , p2 , . . . , pn } ⊂ P. (3) Statistical Analysis: run nreps repetitions for each configuration and perform statistical tests e.g. conduct paired t-tests, and compute Cohen’s d effect size. (4) Artefact Classification: Classify each configuration based on the statistical analysis. If the claimed improvement only holds for a subset of P, it is not generalisable and relies on active parameters. This pipeline is designed to be extensible to any QEM by replacing the ZNE specific parameters (F, E, S). VI. T HE PARAMETER S PACE IS ACTIVE To test whether implicitly assumed parameters θundoc and the statistical test gap (C3) have practical consequences – that is, whether an experiment with seemingly the same prerequisites can yield different outcomes – we demonstrate how varying different axes of P can shift the experimental outcome. A. Case Study: Khan et al. (2024) 1) Parameter Extraction: To evaluate the effectiveness of different QEM techniques for NISQ devices, Khan et al. [61] apply dynamic decoupling, twirled readout error extraction and ZNE to four-qubit Quantum Trotter Circuits (QTC) whose gate structure ( [61], Algorithm 1) corresponds to the Trotterisation of a transverse-field Ising chain (C). The paper executes these circuits on IBM Kyoto and Osaka machines (both now retired [62]), and compares them with an ideal QASM Simulation (H) [63]. Error rates, qubits properties and architecture (at the time of execution) are documented, as well as pseudo-code for the QEM procedures. Unfortunately, parameters are neither explicitly nor implicitly (via a reproduction package) available. This leaves us with the unknown F, E, S, and Q. 2) Parameter-Space Sampling: For the missing parameters, we applied common default values: 4096 measurement shots and transpilation level 1 [60]. Additionally, for ZNE, foldingfrom-left strategy, Richardson extrapolation method, and scale factors {1, 3, 5} [18], [45], [46] are employed. From this baseline, we sweep values for reasonable alternatives and execute the configuration on calibration snapshots for IBM Kyoto and Osaka. Inspecting the Kyoto snapshot, we find ECR gate errors of all 144 qubits are near hundred percent, producing a noise-floor output. Given this is not the exact environment the work conducted its simulations, we added a depolarising noise model matching the error rate of the original work. Additionally, we added the reported QASM simulator for ideal results (negative control, as QEM cannot improve ideal results). Ten configurations arise for a one-at-a-time sweep of the five parameters plus the baseline. We vary each axis independently while keeping the others at their default values and tests them against three Trotter depths (TC1, TC3 , TC5) specified in the original paper. In Summary 11 configurations × 4 backends × 3 Trotter depths = 132 configurations are tested in total. 3) Statistical Analysis: For each configuration, we run nreps = 200 independent repetitions, then conduct paired ttests on the per-repetition absolute error |ϵ| relative to the ideal expectation value, and compute Cohen’s d as an effect-size measure. We apply a significance threshold of α = 0.05.

4) Artefact Classification: Figure 3 shows the results of the parameter-space sampling. As expected, the ideal simulator shows a significant worsening with a median Cohen’s d of −0.95: this serves as a negative control, confirming that ZNE cannot improve results that are already ideal, since any introduced noise amplification only degrades an already noiseless expectation value. In the last three columns we see neutral and mostly non-significant results for the fake IBM Kyoto (median Cohen’s d of −0.03), which is also expected given the fully faulty ECR gates – their errors drive the output to the totally mixed state, yielding near-zero expectation values regardless of the circuit. The other two backends show a more positive pattern. The depolarising simulation (Kyoto error rates from the paper) shows significant improvement in 29/33 configurations with a median Cohen’s d of +5.95. Meanwhile the Osaka snapshot has a less positive improvement of median +2.23. As expected, the improvement is more pronounced for the depolarising simulation given its more idealised noise model. In contrast, the Osaka snapshot is closer to real hardware and therefore shows a more mixed pattern of results with other error patterns. The non-improving results and smaller positive improvements are largely driven by the scale factors S, due to variance amplification, and the exponential extrapolation method E. We motivated the amplification of variance by the scale factors in Section III. Different extrapolation methods assume different functional forms of noise decay and have different sensitivities to the noise model [17], [46]. The exponential model assumes an exponential decay of the noise, motivated by global depolarising noise, where E(λ) ∝ (1 − p)λn decays exponentially [17]. However, real hardware noise is often more complex and may not follow this idealised model, as in our case. The extrapolation method E is therefore an active parameter that can significantly affect the results, and its sensitivity to the noise model can lead to different conclusions about the effectiveness of our ZNE. Cliff’s δ values (right half of each cell) closely mirror Cohen’s d patterns, confirming directional results are not an artefact of the shot-count inflation: Wherever d is strongly positive, δ is near +1, and wherever d is negative, δ is near -1. a) Multiple-comparisons perspective: Given that we apply repeated identical tests that lead to a near-certain probability for false positives, we applied the standard Bonferroni and Benjamini-Hochberg corrections to all paired t-test p-values. Of the 107 out of 132 uncorrected significant results (58 better, 49 worse), 103 survive the strict Bonferroni threshold (57 better, 46 worse) and 106 survive Benjamini-Hochberg. Only four results are dropped by Bonferroni, all corner cases near the noise floor. As improvements and degradations, are robust to multiplicity correction, this reinforces that experimental outcomes depend on implicitly assumed parameter values, and degradations observed on the noiseless and FakeKyoto backends are genuine rather than statistical artefacts. b) Hardware calibration and effect size: Khan et al. report expectation values and variances (Table III in [61]), but it is not specified whether these variances represent perrun estimator uncertainty, an average of shot-variance across

Cohen’s d (left) 0

10

15

-1.0

-0.5

0.0

0.5

1.0

-1.0

-0.66

-1.0

-0.62

-0.8

-0.56

+11

+1.00 +9.6 +1.00 +4.9 +1.00 +1.2 +0.80 +2.3 +0.97 +2.6 +0.96

-0.07

+0.0

-0.01

+0.0 +0.04

-1.0

-0.70

-1.0

-0.70

-1.1

-0.74

+11

+1.00

+10

+1.00 +5.5 +1.00 +1.1 +0.69 +2.3 +0.97 +2.6 +0.99 +0.0 +0.01

-0.1

-0.10

-0.1

-0.05

-1.0

-0.69

-0.9

-0.64

-0.9

-0.59

+12

+1.00

+10

+1.00 +4.7 +1.00 +1.5 +0.83 +2.4 +0.99 +2.5 +0.99 +0.0 +0.00 +0.0 +0.04 +0.0

-0.01

-0.4

-0.28

-0.4

-0.35

-0.4

-0.27

+12

+1.00 +5.7 +1.00 +1.8 +0.92 +3.4 +1.00 +5.2 +1.00 +5.4 +1.00 +0.1 +0.10 +0.1 +0.08

-0.1

-1.0

-0.70

-0.9

-0.53

-1.0

-0.63

+12

+1.00 +9.7 +1.00 +5.1 +1.00 +1.2 +0.79 +2.7 +1.00 +2.7 +1.00

-0.5

-0.78

-0.5

-0.67

-0.5

-0.71

+7.1 +1.00

+0.16

-0.1

-1.1

-0.80

-1.0

-0.69

-1.0

-0.76

+2.6 +0.98 +5.1 +1.00 +3.3 +0.99

-0.9

-1.2

-0.99

-1.3

-0.98

-1.4

-0.98

-3.2

-1.00

-0.8

-0.09

-0.5

-0.26

-0.7

-0.52

-1.0

-0.64

-1.0

-0.69

+12

+1.00 +8.9 +1.00 +4.9 +1.00 +1.3 +0.80 +2.6 +0.99 +2.8 +0.99 +0.1 +0.08 +0.0 +0.06

-0.1

-0.10

-0.9

-0.59

-1.0

-0.70

-1.1

-0.72

+6.7 +1.00 +4.5 +1.00 +2.4 +0.98 +0.6 +0.45 +1.2 +0.78 +1.3 +0.83

-1.0

-0.68

-1.0

-0.71

-0.9

-0.64

+15

+1.00

-0.1 -1.3

+14

+0.82 -0.77

-0.7 -0.5

-0.20

-0.1

-0.1

-0.06

+0.99 +2.6 +1.00 +2.5 +0.99

-0.4

-0.62

+0.1 +0.15 +0.6 +0.40

-0.1

-0.48

-0.9

-0.3

-0.56

-1.0

-0.72

-0.1

-0.12

-0.0

-0.06

+0.0 +0.09

-0.23

-0.5

-0.23

-0.5

-0.01

+0.1 +0.01 +0.0 +0.00

-0.19

-0.3

-0.23

-0.07

-0.0

+0.01 +0.0 +0.05

+1.00 +7.0 +1.00 +1.8 +0.91 +3.3 +1.00 +3.3 +1.00 +0.0 +0.05

-0.1

+0.01

-0.2

-0.10

(n e

g. No co ise nt les ro s (n TC l) eg N 1 . c ois on el tr ess o (n TC l) eg N 3 . c ois on el t es T rol) s C 5 D K epo y T oto l. C 1 D K epo y T oto l. C 3 D K epo y T oto l. C 5 O Fa s k T aka e C 1 O Fa s k T aka e C 3 O Fa s k T aka e C 5 K Fak y T oto e C 1 K Fak y T oto e C 3 K Fak y T oto e C 5

Default (Richardson) Fold: from-right Fold: global Extrap: linear Extrap: polynomial Extrap: exponential Scales: 1,2,3 Scales: 1,1.5,2,2.5,3 Transpiler: 3 Shots: 1024 Shots: 8192

5

Cliff’s δ (right)

Fig. 3. Parameter space heatmap for 4 backends × 3 Trotter depths (separated by vertical lines), one-at-a-time sweep from defaults. Each cell is split: the left half shows Cohen’s d (colour scale left), the right half shows Cliff’s δ (colour scale right). Green = ZNE significantly reduces error, yellow = ZNE significantly increases error (α=0.05), white = not significant.

runs, or a different measure. The number of independent repetitions per configuration is also not reported. This means we cannot directly verify the statistical significance of the reported improvements, nor recover paired differences needed to compute Cohen’s d. We estimate dˆ for the original paper by computing ∆ = EZNE − Eraw from the reported values and divide by the simulated σimprovement as proxy. Since Khan et al. used real IBM hardware, the IBM Osaka calibration snapshot is the closest available approximation for comparison. Our replication yields d = +1.16 at TC1 – a moderate but statistically significant improvement. For the original paper [61], Eraw = 0.728, EZNE = 0.814, Eideal = 0.828; dividing the improvement ∆ = 0.086 by our σimprovement = 0.016 gives dˆ ≈ +5.3, more than four times our replication value. Under idealised depolarising noise (same circuit, same configuration), the same calculation yields d = +11.3. The original effect size therefore depends critically on the specific hardware calibration (H) at the time of the experiment, which is no longer accessible. We therefore cannot distinguish whether d ≈ 1.16, d ≈ 5.3, or d ≈ 11.3 is the better characterisation of the true effect, let alone assess its robustness across noise models [42], [64].

outweighs any potential correction, pushing the mitigated value further from ideal. At TC3 (18 ECR gates), sufficient error accumulates for ZNE to provide genuine improvement. This not only confirms hardware dependence of the outcome, but also reveals a regime where low-error hardware makes ZNE counterproductive for shallow circuits. c) Shot-count variance and statistical power: Figure 4(a) √ confirms the expected d ∝ nshots scaling: on the depolarising model and FakeOsaka, d grows monotonically with shot count, while FakeKyoto stays near zero (no genuine signal) and the ideal simulator is fixed at d ≈ −0.95. A bootstrap power analysis (see Figure 4(b)) shows that nreps ≥ 20 suffices for 80% power at moderate effects, while large depolarising effects need as few as nreps = 5; the near-zero FakeKyoto effect never reaches 80% power even at nreps = 200. Low shot counts (Q) and few repetitions therefore risk both false negatives for genuine effects and an inability to confirm the absence of improvement where none exists. B. Case Study: Desdentado et al. (2025)

So far, we have assumed that multiple runs of an identical configuration yield the (essentially) identical result. DesdenTo check if simulation-based findings transfer to quantum tado et al. [66] provide a case where this assumption breaks: hardware, we ran the default configuration on the 156-qubit they observe a temporal confound whose structure is consistent IBM Marrakesh processor [65] for TC1 and TC3. The Heron- with calibration drift, as studied below in Section VII. Their based machine has a median two-qubit error rate of 0.241%, work proposes an algorithm to estimate the ideal shot count for roughly four times lower than the 0.947% reported for IBM a given quantum circuit to achieve the best possible result under Kyoto. Our experiment yields d = −0.75 for TC1, but shot noise – the statistical sampling variance that decreases as √ d = +0.55 for TC3. The negative effect at TC1 arises as low 1/ nshots when more measurement shots are taken. Hardware error rates leaves the shallow circuit (six ECR gates) nearly noise (e.g., gate errors, calibration drift, . . . ) is, in contrast, not ideal (|ϵraw | = 0.013), so ZNE’s variance amplification (2.7×) reduced by increasing shot count.

(a) Effect size vs. shot count

(b) Power vs. repetition count

Noise backends (genuine effect)

15

1.0

Statistical power

10

Cohen’s d

5 0 Neg. control / broken calibration

0.00 -0.25 -0.50 -0.75 -1.00

0.8

80% power

0.6 0.4 0.2 0.0

128

256

512

1K

2K

4K

8K

5

nshots Backend

Noiseless (neg. ctrl.)

10

20

30

50

100

TC1

TC3

200

nreps Depol. Kyoto

Fig. 4. Sensitivity analysis. (a) Cohen’s d vs. shot count; d scales with resamples); 80% power requires nreps ≥ 20 for moderate effects.

FakeOsaka

FakeKyoto

Trotter depth

TC5

nshots for genuine effects. (b) Statistical power vs. repetition count (1,000 bootstrap

as shot noise decreases. Job identifiers reveal that shot count groups were executed in sequence over a 20 minute time span, 0.07 except for a gap of 13 minutes between parts of group 8458 and the entire group 10936. Because shot count is confounded 0.06 with execution time and the group size is small (n = 10), inter0.05 group differences may also reflect calibration drift. Particularly 1024 3502 5980 8458 10936 given the non-monotonic pattern (the last complete executed (Default) (Mid-Low) (Estimated) (Mid-High) (High) shot group, ideal shot estimation, coincides with lowest error), Shot count (configuration) we cannot distinguish hardware drift, shot count or pure chance as cause. Fig. 5. Shot estimation results from Desdentado et al. [66] for a five qubit Grover circuit on the IBM Brisbane. The proposed shot estimation yields the This illustrates how temporal drift can masquerade as parambest target state probability, while higher shot counts drift to the noise floor. eter effect with un-randomised execution order. To isolate drift as an independent factor, we conduct a controlled longitudinal The paper proposes an algorithm to estimate the ideal shot experiment measuring its impact on ZNE effectiveness. count and tests five configurations for a five-qubit Grover circuit Cross-vendor fragility To test whether these findings transfer on IBM Brisbane. If finds an improvement of 0.37% for two, beyond IBM hardware, we executed the same circuit on the and 0.53% for four target states. Their proposed shot estimation 54-qubit IQM Euro-Q-Exa [70] machine. At TC1 (6 CZ gates yields the best target state probability for all their circuits while at λ1 ), the circuit retains 27.7% of the ideal expectation value lower and higher shot counts yield worse results. Uniquely, the (Ē(λ1 ) = 0.290 vs. Eideal = 0.980). At λ3 and λ5 , the paper provides a complete reproduction package that allows expectation value is near the noise floor (E(λ3 ) = 0.024) or us to analyse exact hardware results from published data (see negative (E(λ5 ) = −0.105). TC3 (18 CZ) retains only 10.9% Figure 5). Circuit depth and hardware error rate are dominant of the ideal expectation value on Euro-Q-Exa, compared to challenges: IBM Brisbane (Eagle r3) reports a median two- 113% on IBM Marrakesh (coherent over-rotation). The circuit qubit gate error rate of 0.77% error rate through the last IBM depth at which ZNE “works” is therefore a property of the calibration snapshot. Alternate reported median ECR error rates circuit–hardware pair: cross-vendor studies are necessary for ranging from 0.762% [67], 0.79% [68] to 0.832% [69]. The any generalisable claim. As the signal is near a fully mixed state transpiled circuit uses 904 ECR gates, which leads to a total at TC3 on the Euro-Q-Exa, TC1 is used for the longitudinal circuit fidelity of F = (1 − 0.00762)904 ≈ 0.09% with the drift study in the next section. most optimistic error rate. The low circuit fidelity explains a near-random target state VII. T HE C ONTINGENCY OF T IME probability. Desdentado et al. (see Figure 5) report probability degradation after exceeding the estimated ideal shot count, with Even with fixed parameters, replications can depend on when a non-monotonic pattern. We only show results for one circuit, an experiment is run, given time-varying HW properties. We but the pattern applies to all. This contradicts the shrinking find only 28% of papers address hardware drift (C4). Given variance of shot noise with increasing shot count, which should that superconducting processors exhibit calibration fluctuations result in a monotonic improvement in target state probability for minutes to hours [42], [44], this raises concerns.

Success rate

0.08

non-stationary phenomenon that can yield different outcomes A longitudinal experiment on the 54-qubit IQM Euro-Q- for ZNE effectiveness at different times. We extended this Exa system available at LRZ [70] tests the effect of temporal experiment to a full seven-day window (163 scheduled hours, drift using the Khan et al. [61] QTC at depth TC1 (6 CZ at Figure 7): the trace shows qualitatively different overnight λ1 ). We chose the smallest circuit, because the expectation behaviour, and the baseline level after a 43-hour outage (red value Ē(λ1 ) of TC3 is already dominated by noise on this gap) differs visibly from before, confirming that drift persists hardware. To minimise parameter confounds, we use default over days and is not resolved by recalibration. b) Drift is pervasive and severe (RQ2): In the 48-hour configurations from Section VI-A2 and run an experiment weekend study, more than half of the total variance in the raw every 30 minutes for different time periods. This allows us signal is attributable to the time of measurement (η 2 = 0.55 to observe any temporal variability attributable to hardware drift. Four days with two 12 hour periods and one 48 hour (see Table III)), and consecutive time points are strongly weekend period with a total of 147 time points with nreps = 30 correlated (r1 = 0.83). Each measurement carries information independent runs allow for observing short-term and long-term about the next, violating independence assumptions of standard paired tests. Additionally the autocorrelation is asymmetric: drift patterns. We address three questions: Is drift reproducible across r1 = 0.17 in the first 24 hours vs. 0.91 in the second, sessions (RQ1); do raw expectation values at λ1 exhibit coinciding with a visible upward shift during the second night temporal autocorrelation inconsistent with the i.i.d. assumption (t = 39−43 h) which results in a higher E(λ1 ) up to 0.346 of paired tests (RQ2); does drift cause per-time-point ZNE compared to most of the other time points with expectation 0.300. The 12-hour sessions confirm drift at lower effect size to vary substantially such that identical experiments values below 2 severity (η = 0.20−0.35, r1 = 0.21−0.55). at different times yield different effectiveness conclusions c) The drift-induced effectiveness illusion (RQ3): Figure (RQ3). 6 (b) shows per-time-point Cohen’s d of ZNE vs. raw across all B. Results sessions. The d values vary substantially over time (3.3–12.9) (although inflated by the high measurement precision at nshots = Across all sessions, Ē(λ1 ) ≈ 0.28−0.30, corresponding 4096). What matters is not effect size, but relative variation to approximately 29% of the ideal expectation value Eideal , over time. For instance, in the weekend session, the experiment E(λ3 ) ≈ 0.02 (near noise floor), and E(λ5 ) is consistently yields d = 3.3 at one time point and d = 11.3 twelve hours negative (mean −0.10), which is inconsistent with depolarising later: a 3.4× difference in apparent ZNE effectiveness. For noise decay and indicating coherent over-rotation past the comparison, switching from the Osaka calibration snapshot to zero-crossing. Figure 6 shows the time series across all three a fundamentally different depolarising noise model, produces sessions alongside the associated ZNE effect size per time only a ratio of 2.7. Temporal drift on a single back-end can point. Table III summarises the drift severity metrics. produce larger variation in apparent ZNE effectiveness than changing the entire noise model. TABLE III In our experiment, the effect is significantly positive at every D RIFT SEVERITY ON E URO -Q-E XA . η 2 : FRACTION OF TOTAL VARIANCE BETWEEN TIME POINTS . r1 : LAG -1 AUTOCORRELATION OF Ē(λ1 ). nEFF : time point (d > 3 throughout), as expected: Ē(λ1 ) is well EFFECTIVE INDEPENDENT REPETITIONS ( NOMINAL nREPS =30). d RANGE : above the noise floor and the high shot count inflates d (SecC OHEN ’ S d; ALL PER TIME - POINT. tion III-B). The illusion manifests not as a sign reversal, but as uncontrolled magnitude variation: for moderate effects (d ≈ 1– Session TPs η2 r1 neff d range 2) commonly reported in QEM hardware evaluations [33], this Day 1 (12 h) 25 0.35 0.55 3.5 6.1–10.9 temporal variation alone is sufficient to determine the binary Day 2 (12 h) 25 0.20 0.21 4.9 7.2–12.9 Weekend (48 h) 97 0.55 0.83 1.8 3.3–11.3 conclusion of statistical significance. Additionally, the withintime-point Intraclass Correlation Coefficient (ICC) reduces a) Three sessions, three drift patterns (RQ1): The top the nreps =30 nominal repetitions to as few as neff = 1.8 part of Fig. 6 reveals qualitatively different dynamics: Day effective independent observations (Table III), following Kish’s 1 shows a slight upward drift with a discrete step-change design-effect formula neff = n/(1 + (n−1) · ICC) [71]. Note at t ≈ 9.5 h (that is, asymmetric across scale factors: λ1 that this ICC is distinct from the between-time-point r1 : it affected, λ3,5 insensitive). Contrary to the first session, day is computed via one-way ANOVA within each time point. A 2 exhibits a gradual downward trend to the lowest measured genuinely moderate improvement that appears highly significant expectation value. Session 3 was conducted over the weekend, at nreps =30 becomes non-significant once this effective sample and reveals what appears to be an overnight recalibration size is accounted for. Two tests of the same ZNE configuration – shift in the second night (between a Saturday and Sunday), one in the morning, one the next day – can give contradictory which is absent in the first night. Between days 1 and 2, conclusions. Neither would be wrong given measured results, E(λ3 ) crosses zero (+0.024 → −0.011), which the negative but neither would paint an accurate picture. This limits Richardson coefficients convert into a large change in the ZNE generalisability of typical quantum (software) experiments that estimate at around t = 18 h. Different patterns across sessions are often based on immediately consecutive hardware runs indicate that drift is not a stable, reproducible process, but a and limited repetition counts, especially given the financial A. Experimental Design

(a) Session 1 (12 h)

(b) Session 2 (12 h)

(c) Weekend (48 h) Night 1

Ē(λ1 )

0.350

Night 2

0.325 0.300 0.275 0h

3h

6h

9h

12h 0h

3h

6h

0h

3h

6h

9h

12h 0h

3h

6h

9h

12h

0h

12h

24h

36h

48h

9h

12h

0h

12h

24h

36h

48h

Cohen’s d

12.5 10.0 7.5 5.0

Hours since start

Fig. 6. Longitudinal drift study on IQM Euro-Q-Exa (147 time points across 72 hours in three independent sessions). Top: Ē(λ1 ) averaged values with 95% CI vs. elapsed time; each session exhibits a qualitatively different drift pattern (step-change, gradual decline, overnight shift). Bottom: per time point Cohen’s d (ZNE vs. raw): d varies around 3× across the weekend, this is a drift-induced effectiveness illusion.

commitment for such experiments. IBM ibm_brussels (TC3, 12-hour session, Figure 8) shows Ē(λ1 ) drifts by approximately 0.14 within a single session, confirming that drift-induced effectiveness illusion is not specific to IQM. VIII. D ISCUSSION A. The compound nature of QEM artefacts Our analysis reveals that a favourable ZNE outcome requires both conditions to hold: the chosen parameter configuration must fall in the improving region of P (Section VI), and the hardware calibration at the time of measurement must produce a stable, noise-discernible expectation value (Section VII). Critically, these two artefact sources interact and amplify each other. A paper that reports a single configuration at one point in time risks confounding both: the observed outcome may be an artefact of a fortunate parameter choice or a favourable calibration window. The Khan et al. reproduction illustrates this compounding: even within the moderate Osaka noise model, d ranges from −0.75 on IBM Marrakesh to +11.3 under idealised depolarising noise (Section VI). Adding temporal drift introduces an additional 3.4× variation in apparent ZNE effectiveness within 48 hours on the same device. This compound structure explains why few papers survive both challenges: a result must be robust against parameter choices and stable under drift. The rarity of meeting all review criteria suggests that the current literature overestimates ZNE reliability not because of method flaws, but because evaluation does not control for confounders. B. Scope and limitations We focus on ZNE with Richardson extrapolation, the most widely used QEM method in our corpus. Artefact sources – parameter sensitivity and temporal drift – are method-agnostic and apply to any quantum experiment and hardware platform,

but patterns may differ. Our design does not capture interaction effects; a full factorial design could reveal additional effects. Internal validity: Paired t-tests assume approximate normality of differences. While robust to moderate non-normality, smaller configurations may violate this assumption. We therefore computed the Wilcoxon signed-rank test for every configuration, where the tests agree on 129 of 132 outcomes, (all on the FakeKyoto backend at marginal p-values). Bonferroni and Benjamini–Hochberg corrections on 132 parameter-space configurations leave qualitative conclusions unchanged. External validity: Drift patterns may differ on other devices, architectures, or time scales. The Khan et al. reproduction uses noise-model snapshots rather than live hardware for the parameter sweep, which may not capture all device effects. Construct validity: Our eight-criterion review framework is a pragmatic operationalisation of statistical rigour. Consequently, other framings could yield different compliance rates. C. Recommendations Based on our analysis, we propose a reporting checklist, ordered by implementation effort: 1) Document all Active Parameters, including calibration snapshot (H), shot count and repetition count (Q), transpilation seed, or method-specific hyper-parameters. 2) Report Inferential Statistics: At least pair claimed improvements with a hypothesis test and effect-size measure. 3) Provide a Reproduction Package containing all code, data, transpiled circuits, calibration snapshots, etc. as explicit baseline for independent verification. 4) Ensure Result Robustness by evaluating at least a small grid of configurations (e.g., multiple scale factors, backends, . . . ) to avoid misleading results.

Night 1

Night 2

Night 3

Night 4

Night 5

Night 6

Night 7

Ē(λ1 )

0.35 0.30 0.25 ideal

0.20 0h

24h

48h

QPU outage

72h

96h

Scheduled hours since start

120h

144h

Fig. 7. Seven-day longitudinal drift study on IQM Euro-Q-Exa (163 scheduled hours, 2026-04-08 to 2026-04-15). Ē(λ1 ) per time point with 95% CI; dashed lines show piecewise-linear interpolation through the mean of each night-bounded daytime cluster (pre- and post-outage separately); grey bands mark local night (21:00–06:00 CEST) and the red gap a 43-hour QPU maintenance outage. Ē(λ1 ) post-outage shows HW recalibration does not restore a stable baseline.

effective independent observations, substantially weakening the evidential basis of nominally repeated measurements. Parameter sensitivity and temporal drift compound on real 0.60 hardware. Their interaction challenges the validity of QEM 0.55 benchmarks that do not include inferential testing, robustness analysis, and drift control. An apparent QEM improvement 0.50 may reflect a favourable point in parameter space, a favourable calibration window, or both. This is not an issue of the methods, 0h 3h 6h 9h 12h but relates to their use. We believe this is partly because Hours since session start especially use-case-centric empirical measurements are often Fig. 8. Twelve-hour drift session on IBM ibm_brussels (TC3, 25 time carried out by domain specialists who may lack deeper training points). Ē(λ1 ) with 95% CI ribbon and linear trend (dashed); the dotted in quantum computing. It is reasonable for them to rely on line marks the ideal value (Eideal = 0.846). Despite a fixed hardware and standard settings and avoid involvement with low-level details. parameter configuration, Ē(λ1 ) drifts across the session. We believe the core of the problem is a software challenge: Providing good mechanisms, abstractions and reference patterns 5) Quantify Temporal Stability by distributing HW experi- would alleviate “users” from having to deal with such details. ments over in time, or randomise execution order to de- We hope our reproduction pipeline, together with the proposed confound drift from parameter effects. reporting standards, will support more robust QEM evaluation (and results with improved practical credibility, as well as IX. C ONCLUSION scientific soundness) as the field progresses towards practical Our systematic review of 81 QEM papers reveals pre- quantum advantage. dominantly descriptive reporting: only 25% of the papers employ inferential statistics, and only 30% address hardware DATA AVAILABILITY drift. A two-stage analysis shows this methodological gap has Code, data, and plotting scripts are available in our reprosubstantial consequences: Prevailing practices yield different interpretations of the same technique when uncontrolled factors duction package that can build the paper including analysis result. HW calibration snapshots and logs allow for analysing (e.g., parameter choice, execution time) vary. First, we show implicitly assumed ZNE parameters, including parameters and drift without machine access. scale factors, extrapolation method, and hardware calibration, Acknowledgments The authors gratefully acknowledge the use of the are active: on two physics-based noise models, variations shift quantum system Euro-Q-Exa, co-funded by the EuroHPC JU, BMFTR the conclusion from significant improvement to significant (grant 13N16690), and the Bavarian State Ministry of Science and degradation in 8 of 66 configurations (12%), while 55 (83%) the Arts, operated by the Leibniz Supercomputing Centre (LRZ) in show significant improvement – depending solely on which Garching, Germany, for providing the computational resources for parameters were implicitly assumed. We found that on IBM this work. We acknowledge partial support by the German Research Marrakesh QPUs, ZNE can even be counterproductive for Foundation, grant MA 9739/1-1, and the High-Tech Agenda of the shallow circuits whose unmitigated output is close to ideal. Free State of Bavaria. We also acknowledge partial support by the Second, our 72-hour longitudinal study on IQM Euro-Q-Exa European Union (Project Reference 101083427), the European Funds shows temporal drift alone induces a 3.4-fold variation in for Regional Development (EFRE) (Project Reference 20-3092.10apparent ZNE effectiveness, exceeding the 2.7-fold variation THD-105), by the European Regional Development Fund (ERDF) observed when changing the entire noise model. The same drift and by the Free State of Bavaria as part of the project AIM-SMEs reduces the nreps =30 nominal repetitions to as few as neff =1.8 (Grant No. 2506-014-3.2), co-funded by the European Union.

Ē(λ1 )

0.65

R EFERENCES [1] J. Preskill, “Quantum Computing in the NISQ era and beyond,” Quantum, vol. 2, p. 79, Aug. 2018, arXiv:1801.00862 [quant-ph]. [Online]. Available: http://arxiv.org/abs/1801.00862 [2] F. Greiwe, T. Krüger, and W. Mauerer, “Effects of imperfections on quantum algorithms: A software engineering perspective,” in IEEE International Conference on Quantum Software (QSW). IEEE, 2023, pp. 31–42. [Online]. Available: https://doi.org/10.1109/QSW59989.2023. 00014 [3] S. Thelen, H. Safi, and W. Mauerer, “Approximating under the influence of quantum noise and compute power,” in IEEE International Conference on Quantum Computing and Engineering (QCE). IEEE, 2024, pp. 274– 279. [Online]. Available: https://doi.org/10.1109/QCE60285.2024.10291 [4] K. Bharti, A. Cervera-Lierta, T. H. Kyaw, T. Haug, S. AlperinLea et al., “Noisy intermediate-scale quantum algorithms,” Rev. Mod. Phys., vol. 94, p. 015004, Feb 2022. [Online]. Available: https://link.aps.org/doi/10.1103/RevModPhys.94.015004 [5] J. Preskill, “Fault-tolerant quantum computation,” Dec. 1997, arXiv:quantph/9712048. [Online]. Available: http://arxiv.org/abs/quant-ph/9712048 [6] ——, “Beyond NISQ: The Megaquop Machine,” ACM Transactions on Quantum Computing, vol. 6, no. 3, pp. 18:1–18:7, Apr. 2025. [Online]. Available: https://dl.acm.org/doi/10.1145/3723153 [7] M. Beverland, V. Kliuchnikov, and E. Schoute, “Surface code compilation via edge-disjoint paths,” PRX Quantum, vol. 3, no. 2, p. 020342, May 2022, arXiv:2110.11493 [quant-ph]. [Online]. Available: http://arxiv.org/abs/2110.11493 [8] C. Gidney, M. Newman, P. Brooks, and C. Jones, “Yoked surface codes,” Dec. 2023, arXiv:2312.04522 [quant-ph]. [Online]. Available: http://arxiv.org/abs/2312.04522 [9] L. Schmidbauer and W. Mauerer, “SAT strikes back: Parameter and path relations in quantum toolchains,” in Proceedings of the IEEE International Conference on Quantum Software (QSW). IEEE, 2025, pp. 1–12. [Online]. Available: https://doi.org/10.1109/QSW67625.2025.00021 [10] L. Schmidbauer, K. Wintersperger, E. Lobe, and W. Mauerer, “Polynomial reduction methods and their impact on QAOA circuits,” in IEEE International Conference on Quantum Software (QSW), 2024, pp. 35–45. [Online]. Available: https://doi.org/10.1109/QSW62656.2024.00018 [11] L. Schmidbauer, E. Lobe, I. Schaefer, and W. Mauerer, “It’s quick to be square: Fast quadratisation for quantum toolchains,” ACM Transactions on Quantum Computing, Mar. 2026, just Accepted. [Online]. Available: https://doi.org/10.1145/3800943 [12] S. Thelen and W. Mauerer, “Predict and conquer: Navigating algorithm trade-offs with quantum design automation,” in IEEE International Conference on Quantum Computing and Engineering (QCE). Los Alamitos, CA, USA: IEEE Computer Society, 2025, pp. 591–602. [Online]. Available: https://doi.ieeecomputersociety.org/10. 1109/QCE65121.2025.00071 [13] R. Wille, L. Berent, T. Forster, J. Kunasaikaran, K. Mato et al., “The mqt handbook : A summary of design automation tools and software for quantum computing,” in 2024 IEEE International Conference on Quantum Software (QSW), 2024, pp. 1–8. [14] S. R. Maschek, J. Schwittalla, M. Franz, and W. Mauerer, “Make some noise! measuring noise model quality in real-world quantum software,” in Proceedings of the IEEE International Conference on Quantum Software (QSW). IEEE, 2025, pp. 1–11. [Online]. Available: https://doi.org/10.1109/QSW67625.2025.00010 [15] K. Temme, S. Bravyi, and J. M. Gambetta, “Error mitigation for short-depth quantum circuits,” Physical Review Letters, vol. 119, no. 18, p. 180509, Nov. 2017, arXiv:1612.02058 [quant-ph]. [Online]. Available: http://arxiv.org/abs/1612.02058 [16] Y. Li and S. C. Benjamin, “Efficient variational quantum simulator incorporating active error minimization,” Physical Review X, vol. 7, p. 021050, 2017. [17] S. Endo, S. C. Benjamin, and Y. Li, “Practical quantum error mitigation for near-future applications,” Physical Review X, vol. 8, p. 031027, 2018. [18] Z. Cai, “Quantum error mitigation,” Reviews of Modern Physics, vol. 95, no. 4, 2023. [19] E. v. d. Berg, Z. K. Minev, A. Kandala, and K. Temme, “Probabilistic error cancellation with sparse Pauli-Lindblad models on noisy quantum processors,” Nature Physics, vol. 19, no. 8, pp. 1116– 1121, Aug. 2023, arXiv:2201.09866 [quant-ph]. [Online]. Available: http://arxiv.org/abs/2201.09866

[20] P. Czarnik, A. Arrasmith, P. J. Coles, and L. Cincio, “Error mitigation with Clifford quantum-circuit data,” Quantum, vol. 5, p. 592, Nov. 2021, arXiv:2005.10189 [quant-ph]. [Online]. Available: http://arxiv.org/abs/2005.10189 [21] A. Kandala, K. Temme, A. D. Córcoles, A. Mezzacapo, J. M. Chow et al., “Error mitigation extends the computational reach of a noisy quantum processor,” Nature, vol. 567, pp. 491–495, 2019. [22] E. F. Dumitrescu, A. J. McCaskey, G. Hagen, G. R. Jansen, T. D. Morris et al., “Cloud quantum computing of an atomic nucleus,” Physical review letters, vol. 120, no. 21, p. 210501, 2018. [23] L. Schmidbauer, E. Lobe, I. Schaefer, and W. Mauerer, “It’s quick to be square: Fast quadratisation for quantum toolchains,” ACM Transactions on Quantum Computing, vol. 7, no. 2, p. 46, 2026. [Online]. Available: https://doi.org/10.1145/3800943 [24] A. Lucas, “Ising formulations of many np problems,” Frontiers in Physics, vol. 2, 2014. [Online]. Available: http://dx.doi.org/10.3389/fphy. 2014.00005 [25] T. Krüger and W. Mauerer, “Out of the Loop: Structural Approximation of Optimisation Landscapes and non-Iterative Quantum Optimisation,” Quantum, vol. 9, p. 1903, Nov. 2025. [Online]. Available: https: //doi.org/10.22331/q-2025-11-06-1903 [26] L. Schmidbauer, C. A. Riofrío, F. Heinrich, V. Junk, U. Schwenk et al., “Path matters: Industrial data meet quantum optimization,” in IEEE International Conference on Quantum Computing and Engineering (QCE). IEEE, 2025, pp. 2101–2111. [Online]. Available: https://doi.org/10.1109/QCE65121.2025.00230 [27] M. Schönberger, I. Trummer, and W. Mauerer, “Quantum-inspired digital annealing for join ordering,” Proc. VLDB Endow., vol. 17, no. 3, p. 511–524, Nov. 2023. [Online]. Available: https://doi.org/10.14778/ 3632093.3632112 [28] M. Schuld, I. Sinayskiy, and F. Petruccione, “An introduction to quantum machine learning,” Contemporary Physics, vol. 56, no. 2, pp. 172–185, 2015. [29] M. Franz, T. Winker, S. Groppe, and W. Mauerer, “Hype or heuristic? quantum reinforcement learning for join order optimisation,” in IEEE International Conference on Quantum Computing and Engineering (QCE), vol. 01, 2024, pp. 409–420. [30] M. Schuld and F. Petruccione, Machine Learning with Quantum Computers, ser. Quantum Science and Technology. Springer Cham, 2021. [31] P. Wittek, Quantum Machine Learning: What Quantum Computing Means to Data Mining. Boston: Academic Press, 2014. [32] I.-C. Chen, B. Burdick, Y. Yao, P. P. Orth, and T. Iadecola, “Errormitigated simulation of quantum many-body scars on quantum computers with pulse-level control,” Physical Review Research, vol. 4, no. 4, p. 043027, 2022. [33] Y. Kim, A. Eddins, S. Anand, K. X. Wei, E. Van Den Berg et al., “Evidence for the utility of quantum computing before fault tolerance,” Nature, vol. 618, no. 7965, pp. 500–505, Jun. 2023. [Online]. Available: https://www.nature.com/articles/s41586-023-06096-3 [34] A. A. Saki, A. Katabarwa, S. Resch, and G. Umbrarescu, “Hypothesis Testing for Error Mitigation: How to Evaluate Error Mitigation,” Jan. 2023, arXiv:2301.02690 [quant-ph]. [Online]. Available: http://arxiv.org/abs/2301.02690 [35] V. Russo, A. Mari, N. Shammah, R. LaRose, and W. J. Zeng, “Testing Platform-Independent Quantum Error Mitigation on Noisy Quantum Computers,” IEEE Transactions on Quantum Engineering, vol. 4, pp. 1–18, 2023. [Online]. Available: https://ieeexplore.ieee.org/document/ 10219054/ [36] O. G. Maupin, A. D. Burch, B. Ruzic, C. G. Yale, A. Russo et al., “Error mitigation, optimization, and extrapolation on a trapped ion testbed,” Physical Review A, vol. 110, no. 3, p. 032416, Sep. 2024, arXiv:2307.07027 [quant-ph]. [Online]. Available: http://arxiv.org/abs/2307.07027 [37] R. Takagi, S. Endo, S. Minagawa, and M. Gu, “Fundamental limits of quantum error mitigation,” npj Quantum Information, vol. 8, no. 1, p. 114, Sep. 2022. [Online]. Available: https: //www.nature.com/articles/s41534-022-00618-z [38] Y. Quek, D. Stilck França, S. Khatri, J. J. Meyer, and J. Eisert, “Exponentially tighter bounds on limitations of quantum error mitigation,” Nature Physics, vol. 20, no. 10, pp. 1648–1658, Oct. 2024. [Online]. Available: https://www.nature.com/articles/s41567-024-02536-7 [39] M. Krebsbach, B. Trauzettel, and A. Calzona, “Optimization of Richardson extrapolation for quantum error mitigation,” Physical

Review A, vol. 106, no. 6, p. 062436, Dec. 2022. [Online]. Available: https://link.aps.org/doi/10.1103/PhysRevA.106.062436 [40] Y. Li, M. Shao, J. Zhao, and Q. Wang, “A methodological analysis of empirical studies in quantum software testing,” 2026, arXiv:2601.08367 [quant-ph]. [41] E. Moguel, J. A. Parejo, A. Ruiz-Cortés, J. Garcia-Alonso, and J. M. Murillo, “Quantum software experiments: A reporting and laboratory package structure guidelines,” May 2024, arXiv:2405.04192 [cs]. [Online]. Available: http://arxiv.org/abs/2405.04192 [42] P. Senapati, Z. Wang, W. Jiang, T. S. Humble, B. Fang et al., “Towards Redefining the Reproducibility in Quantum Computing: A Data Analysis Approach on NISQ Devices,” in 2023 IEEE International Conference on Quantum Computing and Engineering (QCE), vol. 01, Sep. 2023, pp. 468–474. [Online]. Available: https://ieeexplore.ieee.org/document/10313593/ [43] P. Senapati, S. Y.-C. Chen, B. Fang, T. M. Athawale, A. Li et al., “PQML: Enabling the Predictive Reproducibility on NISQ Machines for Quantum ML Applications,” in 2024 IEEE International Conference on Quantum Computing and Engineering (QCE), vol. 01, Sep. 2024, pp. 1413–1424. [Online]. Available: https://ieeexplore.ieee.org/document/10821454/ [44] Y. Hirasaki, S. Daimon, T. Itoko, N. Kanazawa, and E. Saitoh, “Detection of temporal fluctuation in superconducting qubits for quantum error mitigation,” Applied Physics Letters, vol. 123, no. 18, p. 184002, Nov. 2023. [Online]. Available: https://doi.org/10.1063/5.0166739 [45] R. Majumdar, P. Rivero, F. Metz, A. Hasan, and D. S. Wang, “Best practices for quantum error mitigation with digital zero-noise extrapolation,” Jul. 2023, arXiv:2307.05203 [quant-ph]. [Online]. Available: http://arxiv.org/abs/2307.05203 [46] T. Giurgica-Tiron, Y. Hindy, R. LaRose, A. Mari, and W. J. Zeng, “Digital zero noise extrapolation for quantum error mitigation,” 2020 IEEE International Conference on Quantum Computing and Engineering (QCE), pp. 306–316, 2020. [47] L. Hour, M. Go, and Y. Han, “Improving Zero-noise Extrapolation for Quantum-gate Error Mitigation using a Noise-aware Folding Method,” May 2024, arXiv:2401.12495 [quant-ph]. [Online]. Available: http://arxiv.org/abs/2401.12495 [48] W. Mauerer and S. Scherzinger, “1-2-3 reproducibility for quantum software experiments,” in IEEE International Conference on Software Analysis, Evolution and Reengineering, 2022, pp. 1247–1248. [49] J. Cohen, Statistical Power Analysis for the Behavioral Sciences, 2nd ed. Hillsdale, NJ: Lawrence Erlbaum Associates, 1988. [50] L. Fahrmeir, C. Heumann, R. Künstler, I. Pigeot, and G. Tutz, Statistik: Der Weg zur Datenanalyse. Berlin, Heidelberg: Springer, 2023. [Online]. Available: https://link.springer.com/10.1007/978-3-662-67526-7 [51] J. L. Devore, “Probability and statistics for engineering and the sciences,” 2008. [52] R. L. Wasserstein and N. A. Lazar, “The asa statement on p-values: context, process, and purpose,” pp. 129–133, 2016. [53] S. S. Sawilowsky, “New effect size rules of thumb,” Journal of Modern Applied Statistical Methods, vol. 8, no. 2, pp. 597–599, 2009. [54] J. M. Murillo, J. Garcia-Alonso, E. Moguel, J. Barzen, F. Leymann et al., “Quantum software engineering: Roadmap and challenges ahead,” ACM Trans. Softw. Eng. Methodol., vol. 34, no. 5, May 2025. [Online]. Available: https://doi.org/10.1145/3712002 [55] C. Carbonelli, M. Felderer, M. Jung, E. Lobe, M. Lochau et al., Challenges for Quantum Software Engineering: An Industrial Application Scenario Perspective. Springer Nature Switzerland, 2024, p. 311–335. [Online]. Available: http://dx.doi.org/10.1007/978-3-031-64136-7_12 [56] T. Yue, W. Mauerer, S. Ali, and D. Taibi, Challenges and Opportunities in Quantum Software Architecture. Springer Nature Switzerland, 2023, p. 1– 23. [Online]. Available: http://dx.doi.org/10.1007/978-3-031-36847-9_1 [57] I. M. Veiga and E. Hänggi, “Reproducible builds for quantum computing,” 2025. [Online]. Available: https://arxiv.org/abs/2510.02251 [58] V. Gierisch and W. Mauerer, “Qef: Reproducible and exploratory quantum software experiments,” 1 2026. [Online]. Available: https: //arxiv.org/pdf/2511.04563 [59] B. Kitchenham, S. Charters et al., “Guidelines for performing systematic literature reviews in software engineering,” 2007. [60] A. Javadi-Abhari, M. Treinish, K. Krsulich, C. J. Wood, J. Lishman et al., “Quantum computing with Qiskit,” 2024. [61] M. U. Khan, M. A. Kamran, W. R. Khan, M. M. Ibrahim, M. U. Ali et al., “Error Mitigation in the NISQ Era: Applying Measurement Error Mitigation Techniques to Enhance Quantum Circuit Performance,”

Mathematics, vol. 12, no. 14, p. 2235, Jan. 2024. [Online]. Available: https://www.mdpi.com/2227-7390/12/14/2235 [62] IBM Quantum, “Retired QPUs.” [Online]. Available: https://quantum.cloud.ibm.com/docs/en/guides/quantum.cloud.ibm. com/docs/en/guides/processor-types [63] A. W. Cross, A. Javadi-Abhari, T. Alexander, N. d. Beaudrap, L. S. Bishop et al., “OpenQASM 3: A broader and deeper quantum assembly language,” ACM Transactions on Quantum Computing, vol. 3, no. 3, pp. 1–50, Sep. 2022, arXiv:2104.14722 [quant-ph]. [Online]. Available: http://arxiv.org/abs/2104.14722 [64] M. Zheng, A. Li, T. Terlaky, and X. Yang, “A Bayesian Approach for Characterizing and Mitigating Gate and Measurement Errors,” ACM Transactions on Quantum Computing, vol. 4, no. 2, pp. 11:1–11:21, Feb. 2023. [Online]. Available: https://dl.acm.org/doi/10.1145/3563397 [65] IBM Quantum, “Processor types.” [Online]. Available: https://eu-de.quantum.cloud.ibm.com/docs/en/guides/eu-de. quantum.cloud.ibm.com/docs/en/guides/processor-types [66] E. Desdentado, M. Polo, and C. Calero, “Estimating the number of shots to improve results accuracy,” 2025, preprint. [Online]. Available: https://github.com/GreenTeamAlarcos/ Estimating-The-Number-Of-Shots-To-Improve-Results-Accuracy [67] R. Robertson, E. Doucet, E. Spicer, and S. Deffner, “Simon’s algorithm in the NISQ cloud,” Entropy, vol. 27, no. 7, p. 658, Jun. 2025, arXiv:2406.11771 [quant-ph]. [Online]. Available: http://arxiv.org/abs/2406.11771 [68] C. Benito, E. López, B. Peropadre, and A. Bermudez, “Comparative study of quantum error correction strategies for the heavy-hexagonal lattice,” Quantum, vol. 9, p. 1623, Feb. 2025, arXiv:2402.02185 [quant-ph]. [Online]. Available: http://arxiv.org/abs/2402.02185 [69] M. AbuGhanem, “Practical Fidelity Limits of Toffoli Gates in Superconducting Quantum Processors,” Sep. 2025, arXiv:2509.05395 [quant-ph] version: 1. [Online]. Available: http://arxiv.org/abs/2509.05395 [70] Leibniz Supercomputing Centre, “First European quantum computer for Germany: Euro-Q-Exa starts operation at LRZ - LeibnizRechenzentrum.” [Online]. Available: https://www.lrz.de/en/news/detail/ first-european-quantum-computer-for-germany-euro-q-exa-starts-operation-at-lrz [71] L. Kish, Survey Sampling. New York: John Wiley & Sons, 1965.

Related documents

Record · ID 241561 · SHA-256 7444e1f019e735d0
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.