Impact of Data Loss in Postprocessing on Training and Inference of Quantum Neural Networks
arXiv:2609.05060v1 [quant-ph] 4 Sep 2026
Soraya V Panambalom1,∗ , Edoardo Altamura2,3 , Nick Chancellor1 , and Jonte R Hance1,†
Abstract—As quantum hardware scales to larger devices, the classical software layers that interface with it must evolve in step. Postprocessing routines developed and tested primarily in simulator settings can encode assumptions that no longer hold on utility-scale devices, leading to data loss that can be difficult to detect from high-level model outputs alone. We present a case study of SamplerQNN, the sampling-based quantum neural network class in the Qiskit Machine Learning library. Here, the postprocessing method applies a filter that assumes measurement bit-strings are in virtual qubit space. On our quantum hardware runs, where bit-strings span over 100 physical qubits, this filter led to the loss of 85 to 99.6% of valid measurement shots, depending on the transpiler’s qubit placement. The resulting probability vector is unnormalised, allowing distorted prediction and loss values to propagate through the model without an APIlevel warning. We demonstrate the impact across five experiments on two IBM backends: for inference, accuracy drops from 0.94 to 0.39 on the same raw measurements; for training, the loss signal is compressed by 22 to 27×, substantially reducing the sensitivity of the optimiser to the objective landscape. The behaviour arises in all released versions of the library (0.8.4 to 0.9.0). We implemented a layout-based marginalisation fix, merged into the GitHub codebase as Pull Request #1041, that makes SamplerQNN postprocessing forward-compatible with current and upcoming hardware. Index Terms—Quantum computing, Neural networks, Software reliability, Data postprocessing, Machine learning
I. I NTRODUCTION Reliable software behaviour is key for reproducible scientific workflows, particularly when deploying algorithms to specialised hardware. When software assumptions fail silently, the resulting effects can be difficult to distinguish from physical noise, and can propagate into the final results and their interpretation [1]. Reproducibility studies have shown that scientific code can be difficult to execute and validate across computational environments [2]; and a study of Harvard replication datasets found that approximately half of the code provided successfully executed [3], even after automatic codecleaning corrected simpler errors (e.g., use of absolute rather than relative paths). This has obvious impacts for reproducibility of data based on this code. In a quantum setting, it is crucial that the classical software layers that interface with the machines and perform supporting computations are correct. This is especially important in hybrid 1 Quantum Group, School of Computing, Newcastle University, 1 Science Square, Newcastle upon Tyne, NE4 5TG, UK 2 National Quantum Computing Centre, Didcot, OX11 0QX, UK 3 Yusuf Hamied Department of Chemistry, University of Cambridge, Lensfield Road, Cambridge CB2 1EW, UK ∗ [email protected] † [email protected]
quantum-classical settings, where the classical components are integral pieces of the computational model [4]. Most current quantum machine learning algorithms fall into this category as they are often variational, with training loops based on sampling distributions [5] to compute the objective value from raw counts or observables. Even outside of a hybrid setting, some classical processing is needed (e.g., collection and processing of sample data, and deciding QPU inputs). In this manuscript, we present a case study of data loss in a sampling-based postprocessing routine within the Qiskit Machine Learning (ML) library [6], arising when code paths originally consistent with simulator-style outputs were used with utility-scale hardware outputs. This case study highlights how classical software must evolve alongside the platforms it supports, and illustrates the impact that such issues can have on reported data. The behaviour is traced to SamplerQNN, a Samplerderived primitive implemented in Qiskit ML that interfaces parameterised quantum circuits to classical optimisers. It executes quantum circuits via a Sampler backend, collects the measurement outcomes, and postprocesses them into a probability vector that can be used for classification, regression, or other downstream tasks. The postprocessing logic was designed for simulators, which return measurement bitstrings over the virtual qubits only. However, current IBM quantum hardware behaves differently: the Qiskit transpiler maps n virtual qubits onto N physical qubits (where N > 100 for Eagle and Heron devices), and the backend returns bitstrings over all N measured qubits. This discrepancy between the expected and actual measurement formats causes the postprocessing step to produce biased results in hardware runs. The postprocessing filter in SamplerQNN._postprocess discards bit-strings whose decimal integer value exceeds 2n . On real hardware, this eliminates 85 to 99.6% of valid shots depending on the transpiler’s qubit placement, producing a probability vector that does not sum to 1. Thus, the ML model receives an incorrect signal, but without warning the user. This behaviour manifests itself in both training and inference. When the optimiser runs on hardware, the unnormalised probabilities propagate into the loss function at every iteration, and the resulting weight updates cannot be corrected after the fact. When the model is trained classically via a simulator and only inference is performed on hardware, the predictions are incorrect, but the trained weights are intact and the correct results can be recovered by reprocessing the raw measurements. We identified this issue while running quantum machine
learning experiments on IBM quantum hardware, traced it to the source code, and submitted a fix covering layout-based marginalisation. The issue was reported on the Qiskit ML GitHub repository (#1040 [7]), and our fix was submitted as Pull Request (PR) #1041 [8]. The PR was reviewed by the library maintainers and merged into the official codebase on 7 May 2026. The fix will be included in the upcoming version release (installable from PyPI via pip); in the meantime, users can access the updated SamplerQNN by manually installing Qiskit ML from source1 . We verified this behaviour in recent versions of the library, including 0.8.4 and 0.9.0. However, it was not triggered when using classical simulators because the default Sampler returns measurements in virtual qubit space, where the filter condition is never met. It only manifests when using real IBM quantum hardware or SamplerV2, with all qubits measured by default. This paper is organised as follows. Section II describes the issue in detail: how the postprocessing filter works, and the two distinct failure modes we identified: one driven by qubit placement, the other by ancilla noise. Section III presents the experimental evidence across five experiments on two real IBM quantum devices, showing that the issue affects both inference results (accuracy dropping from 0.94 to 0.39 on identical measurements) and the training loop (loss signal compressed by 22 to 27×). Section IV describes our fix and explains why the issue was not triggered earlier. Finally, Sections V and VI discuss the implications for Qiskit ML users, concluding with an outlook on forward-compatibility of open source software for quantum machine learning applications. II. T HE P OSTPROCESSING DATA L OSS M ECHANISM To understand the issue, it helps to recall what happens when a quantum circuit runs on real hardware. In Qiskit, the qubits in the user’s circuit are called virtual qubits, while the qubits on the chip are called physical qubits. A user defines a circuit with n virtual qubits (for example, n = 4 for a 4-feature classifier with linear entanglement). Before execution, the circuit is transpiled: the transpiler maps each virtual qubit onto a physical qubit on the chip. On a 156-qubit backend like ibm_kingston, the 4 virtual qubits might be placed at positions [0, 1, 2, 3], or equally at positions [108, 109, 110, 118], depending on the chip’s connectivity graph and qubit calibration data at runtime. The remaining 152 physical qubits, called ancillas, play no role in the computation. After execution, the backend returns one measurement bitstring per shot, covering all 156 physical qubits, not just the 4 virtual ones. Each bit-string is 156 bits long: the 4 bits at the virtual qubit positions carry the actual computation result, and the 152 ancilla bits are irrelevant. We note that SamplerQNN measures all qubits in the transpiled circuit by default. To evaluate the cost function, the postprocessing step must extract the virtual qubit bits and discard the rest. 1 main branch at github.com/qiskit-community/qiskit-machine-learning.
A. The Filter SamplerQNN._postprocess extracts the virtual qubit results using the following filter: # keys -> ints, filter to valid range for k, v in counts_i.items(): ki = _key_to_int(k) if ki < 2**self.num_virtual_qubits: probs_i[ki] = v / total_shots It converts each N -bit measurement string to an integer and keeps it only if that integer is less than 2n , where n is the number of virtual qubits. For n = 4, the threshold is 24 = 16: any bit-string whose base-10 integer value is 16 or above is discarded. This works when the virtual qubits happen to be at positions [0, 1, 2, 3], the lowest bits. In that case, a measurement where only virtual qubits are non-zero produces an integer below 16, and the filter keeps it. However, the transpiler does not guarantee low positions by default. B. Case 1: Data Loss from High Qubit Placement Consider a 4-qubit circuit where the transpiler places the four virtual qubits at physical positions [108, 109, 110, 118] on a 156-qubit backend. Suppose the circuit produces the virtual outcome |0110⟩, meaning virtual qubit 0 measures 0, virtual qubit 1 measures 1, virtual qubit 2 measures 1, and virtual qubit 3 measures 0. On the physical chip, this sets bit 109 to 1 (virtual qubit 1) and bit 110 to 1 (virtual qubit 2), while the remaining 154 bits are ideally 0. The resulting 156-bit integer is 2109 + 2110 ≈ 1.9 × 1033 , far larger than the filter threshold of 24 = 16. Therefore, the filter discards this perfectly valid measurement. In this example, the same happens for any virtual outcome except |0000⟩: as soon as any virtual qubit measures |1⟩, the corresponding physical bit (at position 108, 109, 110, or 118) produces an overall decimal integer far above 24 , and the shot is discarded. For instance ⊗152
→ integer 0 < 24 → Accepted
⊗152
→ integer 2108 ≫ 24 → Rejected
|0000⟩ |0⟩
|0001⟩ |0⟩
The only virtual outcome that can pass the filter is |0000⟩, where all four virtual qubits measure zero, and even then, only if all 152 ancilla qubits also measure zero. On real hardware, this almost never happens. In one of our experiments with this qubit placement, only 16 out of 4096 shots (0.4%) passed the filter (see Section III). C. Case 2: Data Loss from Ancilla-Bit Contributions The filter can also discard valid measurements even when virtual qubits are in low positions. Consider a circuit where the virtual qubits are at positions [0, 1, 2, 3]. Suppose a shot produces the virtual outcome |1010⟩, a perfectly valid measurement. The integer contribution from the virtual qubits alone is 21 + 23 = 10, which is below the threshold of
16. However, the 152 ancilla qubits are also measured, and hardware noise (thermal excitations, readout errors) causes some of them to flip. If a single ancilla at position 20 flips to |1⟩, the full 156-bit integer becomes at least 220 ≈ 106 , far above the threshold. Because of this flip, the shot is discarded even though the active (i.e. non-ancilla) qubit measurement was acceptable. Even at the lowest possible qubit positions, only 15.2% of shots survive for the 4-qubit model. The remaining 85 to 96% of shots carried valid virtual qubit information but were discarded because of noise on qubits that play no role in the computation. D. Consequence: Unnormalised Probabilities Since total_shots is computed before filtering (= 4096) and only surviving shots contribute toP the numerator, the output probabilities pi are not normalised: pi < 1. With a 15.2% survival rate they, sum to ∼0.15; with 0.4% survival they sum to ∼0.004. Therefore, the cost function is evaluated with a probability vector that is not a valid distribution. We verified that the same filter code is present in all released versions of the library (0.8.4, 0.9.0, and the main development branch at the time of our report). Two prior issues on the repository (#674 and #819) had reported problems with the output shape of transpiled circuits, but neither identified this effect on the postprocessing step. III. E XPERIMENTAL E VIDENCE We discovered the issue while running five quantum machine learning experiments for a binary breast cancer classification task [9] on IBM Heron devices. Each experiment used a variational classifier (either a Sampler-based Classifier or a VQC) trained on a noisy classical simulator and transferred to hardware for inference. The five experiments ran on two backends, ibm_kingston (156 qubits) and ibm_torino (133 qubits), depending on device availability. The models tested are the following: a Sampler-based Classifier (CS), the same classifier with small-angle initialisation [10] and full (all-to-all) entanglement (CS SI+FE), a Variational Quantum Classifier [11] with ZZFeatureMap (VQC ZZ), and a Variational Quantum Classifier with ZFeatureMap (VQC Z). Small-angle initialisation refers to initialising all trainable parameters near zero, so the circuit starts close to the identity. Full entanglement means all-toall qubit connectivity, which increases the number of CNOT gates. All models were trained on a noisy classical simulator using Qiskit Aer [12], [13] with a noise model extracted from FakeJakartaV2, which reproduces the noise characteristics of real IBM hardware (gate errors, readout errors, and decoherence). Training used shot-based sampling with 1024 shots per circuit evaluation. The trained weights were then transferred to real IBM hardware for inference. Table I shows the physical qubit positions assigned by the transpiler for each experiment. Two experiments returned near-random accuracy (0.39), despite having trained well on the simulator. The other three
TABLE I P HYSICAL QUBIT POSITIONS ASSIGNED BY THE TRANSPILER AND PERCENTAGE OF MEASUREMENT SHOTS SURVIVING THE S A M P L E R QNN FILTER FOR EACH EXPERIMENT.
Model CS CS CS (SI+FE) VQC (ZZ) VQC (Z)
Q 4 7 4 4 7
Backend ibm_kingston ibm_torino ibm_kingston ibm_torino ibm_torino
Positions [0,1,2,3] [0,1,2,3,4,5,6] [108,109,110,118] [61,62,54,60] [0,1,2,3,4,5,6]
Surv. 15.2% 4.1% 0.4% 0.2% 5.2%
transferred reasonably. The variable that cleanly separated correct results from incorrect ones was the physical qubit positions: experiments where the transpiler placed virtual qubits at positions 0 to 6 gave correct results, while those mapped to positions in the 50s or above 100 gave nearrandom accuracy. Since the quantum circuit performs the same computation regardless of which physical qubits it uses, this pointed to a position-dependent error in the postprocessing step. A. Impact on Inference The experiments in this section use models trained classically on a simulator. The trained weights are then used to run inference on quantum hardware, with no further optimisation on the device. To confirm the impact described in Section II, we retrieved the raw measurement counts from all five hardware jobs and computed predictions using three different methods from the same raw data: Method A (all qubits): compute parity over all physical qubits. This is a naı̈ve baseline: the ancilla qubits are in a noisy state, so including them introduces substantial noise. • Method B (marginalisation): use the circuit layout to identify the physical positions of the virtual qubits, extract only those bits from each measurement, and compute parity over them. The ancilla qubits are marginalised out. All shots contribute to the prediction. • Method C (SamplerQNN filter): convert each bit-string to an integer and keep only entries below 2n . This reproduces exactly what SamplerQNN._postprocess does internally. •
Tables II and III compare the three methods across all five experiments. Method C reproduced the original SamplerQNN results exactly across all five experiments: not just accuracy, but all 35 individual classification metrics (precision, recall, and F1 per class, for both classes, across all five experiments). This confirmed that, for these experiments, the discrepancies were caused by the postprocessing filter, not by the hardware execution or our circuit design. To illustrate the magnitude of the effect, Table IV shows the full class-wise metrics for the most affected experiment, a 4qubit classifier whose virtual qubits were mapped to positions
TABLE II ACCURACY COMPARISON OF THREE POSTPROCESSING METHODS ON THE SAME RAW HARDWARE MEASUREMENTS . CS: C LASSIFIER S AMPLER . SI+FE: S MALL -A NGLE I NITIALISATION WITH F ULL E NTANGLEMENT. M ETHODS A, B, AND C ARE DEFINED IN S ECTION III-A.
Model CS CS CS (SI+FE) VQC (ZZ) VQC (Z)
Qubits 4 7 4 4 7
Method A 0.57 0.63 0.56 0.44 0.45
Method B 0.94 0.92 0.94 0.54 0.59
Method C 0.94 0.89 0.39 0.39 0.61
TABLE III F1 SCORE ( MACRO ) COMPARISON OF THREE POSTPROCESSING METHODS ON THE SAME RAW HARDWARE MEASUREMENTS . CS: C LASSIFIER S AMPLER . SI+FE: S MALL -A NGLE I NITIALISATION WITH F ULL E NTANGLEMENT. M ETHODS A, B, AND C ARE DEFINED IN S ECTION III-A.
Model CS CS CS (SI+FE) VQC (ZZ) VQC (Z)
Qubits 4 7 4 4 7
Method A 0.56 0.60 0.56 0.43 0.45
Method B 0.93 0.91 0.93 0.50 0.53
Method C 0.93 0.88 0.37 0.31 0.57
[108, 109, 110, 118]. Both columns are computed from the same 4096 raw measurement shots. TABLE IV P ER - CLASS CLASSIFICATION METRICS COMPUTED FROM THE SAME RAW HARDWARE MEASUREMENTS . M ETHOD C RETAINS ONLY 16 OF 4096 SHOTS (0.4%); M ETHOD B USES ALL 4096.
Class Malignant Benign Accuracy
Met. C (filter) Prec. Rec. F1 0.34 0.74 0.47 0.54 0.18 0.27 0.39
Met. B (marginal.) Prec. Rec. F1 0.95 0.88 0.91 0.93 0.97 0.95 0.94
The classically trained model had learned well, achieving an accuracy of 0.94 when all shots were used (Method B). For this experiment, the dominant effect is postprocessing data loss: Method C discards 99.6% of valid shots, reducing accuracy to 0.39. The correction changes the results for two experiments significantly, and the scientific conclusions along with them. Without the correction, the Classifier Sampler with smallangle initialisation and full (all-to-all) entanglement appeared to perform poorly on hardware (accuracy=0.39, F1=0.37), and the VQC with ZZFeatureMap showed the same behaviour (accuracy=0.39, F1=0.31). The common factor between these two models is dense qubit connectivity: both involve many CN OT gates due to full entanglement or pairwise feature interactions. Without inspecting the raw counts, this pattern could reasonably be interpreted as hardware-noise sensitivity in deeper or more entangling circuits, an explanation commonly invoked in quantum machine learning studies using near-term quantum devices. [14] After correction, the SI+FE model achieves 0.94 accuracy on ibm_kingston (Table II), identical to the standard Clas-
sifier Sampler. Despite these architectural differences, both models achieve the same accuracy after correction, showing that neither caused the poor performance. The drop to 0.39 was predominantly due to the postprocessing issue: these were simply the models where the transpiler placed qubits at high physical positions. The VQC with ZZFeatureMap shows lower accuracy on hardware after correction (0.54 with Method B) compared to the simulator (0.70), but the F1 score improves from 0.31 to 0.50, enough to change the interpretation from “complete failure” to “moderate hardware degradation.” What emerges after correction is a clearer picture: the real performance gap is between the VQC architecture and the Classifier Sampler, not between entanglement strategies. This conclusion was hidden by the postprocessing artefact. For quantum ML applications, such as cancer classification and other clinical contexts, the difference between 0.39 and 0.94 accuracy would lead to entirely different assessments of whether a model is viable for deployment. The uncorrected results would have led to a substantially different assessment of the SI+FE model, whereas the corrected postprocessing identifies it as the best-performing configuration within this study. B. Impact on Training So far, the experiments have explored the effects of the postprocessing filter at inference time, i.e. the models were trained on a classical simulator and evaluated on hardware, so the issue only affected the final prediction step (sampling). The model weights, learned in this way, were unaffected, and we could recover the correct results simply by saving and reprocessing the raw measurements with marginalisation. Here, we show that the issue has a greater impact if the training is performed on hardware. To test this, we progressed the pre-trained models in Section III-A for 5 additional COBYLA iterations directly on IBM quantum hardware (ibm_torino for 4 qubits, ibm_fez for 7 qubits), using SamplerQNN to evaluate the loss at every iteration. This means the postprocessing data loss affects both the loss values used during the five fine-tuning evaluations and the subsequent inference step. The loss function computes a probability-weighted expected loss as follows: L(θ) = N
1 X P (0 | xi , θ) · (0 − yi )2 + P (1 | xi , θ) · (1 − yi )2 N i=1 (1) where P (k | xi , θ) are the class probabilities returned by SamplerQNN. With the issue present and a survival rate of ∼4%, the probabilities are compressed by a factor of ∼25, i.e. P pi ∼ 0.04 instead of 1. The loss is compressed by the same factor. At iteration 1 of the on-device fine-tuning, the 4-qubit model has a correct loss of 0.33 but the optimiser received 0.012, a 27× reduction in the loss signal (Table V). Similarly,
TABLE V L OSS RECOMPUTED FROM RAW MEASUREMENTS AT EACH HARDWARE ITERATION . M ETHOD C IS WHAT THE OPTIMISER RECEIVED ; M ETHOD B IS THE CORRECT LOSS . I TERATION 1 USES CLEAN SIMULATOR - TRAINED WEIGHTS ; ITERATIONS 2 TO 5 USE WEIGHTS AFFECTED BY PRIOR UNCORRECTED UPDATES .
Model 4Q
7Q
Iter. 1 (clean) 2 (affected) 3 (affected) 4 (affected) 5 (affected) 1 (clean) 2 (affected) 3 (affected) 4 (affected) 5 (affected)
A 0.50 0.50 0.50 0.50 0.50 0.50 0.50 0.50 0.50 0.50
B 0.33 0.42 0.33 0.41 0.34 0.41 0.42 0.43 0.44 0.44
C 0.012 0.016 0.012 0.016 0.014 0.019 0.019 0.018 0.015 0.015
Surv. 3.7% 3.7% 3.7% 4.0% 4.0% 4.6% 4.4% 4.2% 3.4% 3.4%
for the 7-qubit model we expect a loss of 0.41, but observed a raw loss 0.019, giving a 22× reduction. Our results show that the optimiser receives a near-constant loss signal (∼0.01 to 0.02) regardless of the model’s actual performance. Because of the very small gradients, it is effectively unable to suggest reliable optimisation pathways. This behaviour looks like a barren plateau [15], where vanishing gradients prevent the optimiser from learning. However, the cause here is not the circuit structure or expressibility, but the postprocessing step compressing the probability vector. This distinction matters because the symptom (near-constant loss, small gradients) is identical, and could easily be misdiagnosed as a barren plateau in practice. Crucially, sampling loss during training is irreversible unless all sampled bit-strings are stored (which is typically impractical due to the storage overhead across many optimisation iterations). The first iteration uses the clean simulator-trained weights, but from iteration 2 onwards, COBYLA updates the weights based on the uncorrected loss values. Each subsequent evaluation may therefore use weights modified according to compressed and noisy loss estimates, allowing the effect to accumulate across the hardware fine-tuning sequence. The model parameters after 5 hardware iterations are the product of an optimiser that could not see the landscape it was navigating. 1) Loss Recomputation: To measure the gap between what the optimiser received and what a correct implementation would have provided, we retrieved the raw measurement data from all 10 optimisation jobs (5 iterations × 2 models) and recomputed the loss using the three methods. Table V shows the results. Figure 1 shows the comparison visually. The loss in Method C remains approximately constant between 0.01 and 0.02 throughout the fine-tuning, while Method B shows the fiducial loss values. The loss gap between Methods B and C persists and is large: the optimiser was navigating a landscape it could not see. 2) COBYLA and the Simplex Constraint: Not only was the optimiser blind, it never even began optimising. COBYLA [16] is a trust-region-based (gradient-free) optimiser that works
by maintaining a simplex: a set of n + 1 points in the n-dimensional parameter space, where n is the number of trainable weights. It uses these points to build a local linear model of the loss function. But before it can take a single optimisation step, it must first construct this simplex by evaluating the loss at n + 1 different parameter settings: the starting point, plus n perturbations where one weight at a time is shifted by +1.0 rad. The 4-qubit model has 12 trainable weights, so the simplex requires 13 evaluations. The 7-qubit model has 21 weights, requiring 22 evaluations. With only 5 hardware iterations available, COBYLA never completes its initial simplex for either model: it builds 5 out of 13 vertices for the 4-qubit model and 5 out of 22 for the 7-qubit model. Zero actual optimisation steps are taken. Every iteration we observed was part of the simplex construction phase [16]. This means the experiment was structurally underdimensioned regardless of the issue: even if SamplerQNN had worked correctly and COBYLA had received accurate loss values, 5 iterations would still have been insufficient for a single optimisation step. This imposed a practical runtime constraint: each iteration takes ∼8 minutes of hardware time, and completing the simplex alone would require 13 × 8 = 104 minutes (4 qubits) or 22 × 8 = 176 minutes (7 qubits) out of our total quantum computing hardware allocation of 180 minutes. 3) Weight Drift Analysis: To verify that the iterations correspond to simplex construction rather than optimisation, we extracted the trainable weights at each iteration and compared them to the clean simulator-trained values. Table VI shows these results. Every perturbation is exactly +1.0 radian along a single parameter axis, consistent with simplex construction, not optimisation. Whether COBYLA keeps or reverts a perturbation depends on whether the perturbed loss appears better or worse than the reference. Without the issue, these decisions would have been informed by the real loss values (Method B), which show clear differences between iterations. With the issue, COBYLA received compressed values of ∼0.01 to 0.02 where the differences (∼0.001 to 0.003) are dominated by shot noise. In this regime, the keep-or-revert decisions are effectively random. For the 4-qubit model, only 1 of 4 perturbations survives to inference, and the model stays near the simulator optimum. For the 7-qubit model, all 4 perturbations are kept, and the correct loss rises monotonically from 0.41 to 0.44, and each accumulated perturbation pushes the model further from the optimum. IV. L AYOUT- BASED MARGINALISATION A hardware-compatible postprocessing implementation requires layout-aware marginalisation, outlined as follows. For each measurement bit-string, extract only the bits at the physical positions corresponding to virtual qubits, reconstruct an integer in the virtual qubit space, and accumulate probabilities. This is equivalent to summing over all possible ancilla states,
Fig. 1. Loss for 5 on-device fine-tuning iterations, recomputed from raw measurements. Method B (correct marginalisation) reveals the actual loss landscape; Method C (original filter) shows what COBYLA received. The star marks iteration 1 (clean simulator-trained weights); the shaded region marks iterations where the weights had already been affected by prior uncorrected updates. TABLE VI W EIGHT DRIFT RELATIVE TO CLEAN SIMULATOR - TRAINED VALUES AT EACH ITERATION . A LL NON - ZERO CHANGES ARE EXACTLY ±1.0 RAD (= R H O B E G ), CONFIRMING SIMPLEX PROBING . L OSS C IS THE VALUE THE OPTIMISER RECEIVED AND USED TO DECIDE WHETHER TO KEEP OR REVERT EACH PERTURBATION .
Model
4Q (12 weights)
7Q (21 weights)
Iter. 1 (clean) 2 3 4 5 inference 1 (clean) 2 3 4 5 inference
∆w0 0 +1.0 0 0 0 0 0 +1.0 +1.0 +1.0 +1.0 +1.0
∆w1 0 0 +1.0 +1.0 +1.0 +1.0 0 0 +1.0 +1.0 +1.0 +1.0
the standard mathematical operation in quantum mechanics for projecting a joint distribution onto a subspace. It uses every shot, is unaffected by ancilla noise, and produces a probability distribution that sums to 1.0 by construction. We reported the issue on the Qiskit ML repository (issue #1040 [7]), together with the full diagnostic: the qubit position pattern across five experiments, the three postprocessing methods with accuracy and F1 comparisons, the per-class classification reports confirming exact reproduction of the original results, and a proposed fix based on marginalisation. The maintainers flagged the report as high priority and invited a PR implementing the proposed fix. Our PR [8] replaces the integer value filter in _postprocess with layout-based marginalisation as follows: pos = layout.final_index_layout()
∆w2 0 0 0 +1.0 0 0 0 0 0 +1.0 +1.0 +1.0
∆w3 0 0 0 0 +1.0 0 0 0 0 0 +1.0 +1.0
# changed 0 1 1 2 2 1 0 1 2 3 4 4
Loss C 0.012 0.016 0.012 0.016 0.014 − 0.019 0.019 0.018 0.015 0.015 −
for key, count in counts.items(): key_int = to_integer(key) # extract bit at each virtual position bits = [(key_int >> q) & 1 for q in pos] # reconstruct virtual-space integer virtual = sum( b << i for i, b in enumerate(bits) ) probs[virtual] += count / total_shots When a circuit layout is present (i.e., the circuit has been transpiled), the fix reads the physical positions of the virtual qubits from the layout, extracts only those bits from each measurement bit-string via bit-shifting, and accumulates probabilities in the virtual qubit space. Shots that differ only in their ancilla bits are mapped to the same virtual key, effectively marginalising the ancillas out. When no layout is present (local
simulation with the default sampler), the original behaviour is preserved, as no ancilla qubits exist in that case and the filter works correctly. The PR was reviewed and merged into the official codebase on 7 May 2026, and will be distributed on PyPI in the next version release. The issue had not been reported before because the default sampler used by SamplerQNN (the internal QMLSampler, based on state vector simulation) returns bit-strings in virtual qubit space: n-bit strings rather than Nphysical -bit strings. In that case, the integer value of any measurement is at most 2n − 1, so the filter threshold is never exceeded. The problem only manifests when using real hardware or SamplerV2, which return measurements over all physical qubits. Since most tutorials and examples in the Qiskit ecosystem use simulators, and real hardware access requires specific allocation programmes, the issue went undetected despite being present in all released versions of the library (including 0.8.4 and 0.9.0 at the time of our report). Moreover, the filter may not have caused issues on earlier IBM quantum devices, which had approximately 30 qubits or less [11], [17], making it more likely for virtual qubits to be mapped to low-numbered physical positions that would pass the threshold. As hardware evolved toward larger physical qubit registers, like the Eagle and Heron IBM architectures with over 100 qubits, the same assumption became more likely to induce substantial data loss. V. B ROADER I MPLICATIONS To assess the potential scope of the behaviour, we searched for published studies whose workflows match the conditions under which the postprocessing data loss can occur: SamplerQNN usage, the Qiskit ML versions we tested, and execution on real quantum hardware or SamplerV2-like outputs. We discuss two examples of such works below. Martin-Perez et al. [18] proposed a hybrid classical-quantum transfer learning pipeline in which a pre-trained convolutional neural network (CNN) backbone (ResNet18, EfficientNet-B0, or MobileNetV2) extracts image features, and a SamplerQNN with 4 qubits replaces the final classifier. They used Qiskit ML 0.8.0 and ran hardware experiments on ibm_torino (133 qubits) using SamplerV2. Their results showed that one backbone (ResNet18) slightly improves on hardware compared to noisy simulation (+1.9%), while another (EfficientNet-B0) drops by 19 percentage points (from 91.03% to 71.79%). The authors attributed this gap to “transpilation overhead and stochastic gradients” [18]. Because the same quantum circuit was used for all three backbones, the selective drop is compatible with several explanations, including hardware noise, input-dependent sampling effects, and the postprocessing dataloss mechanism described in our work. We note that we cannot confirm whether the behaviour we document caused the gap in Ref. [18] without inspecting their raw measurement data, but the setup (4 virtual qubits on a 133-qubit backend via SamplerQNN and SamplerV2) matches the conditions under which the behaviour arises.
Chaudhary et al. [19] proposed a quantum-enhanced graph neural network for intrusion detection. Their methodology describes expectation values of Pauli observables (consistent with EstimatorQNN), and their simulation experiments use StatevectorEstimator and BackendEstimatorV2. However, for their hardware validation on ibm_fez (156 qubits), they switch to SamplerQNN with 128 shots and Qiskit ML 0.9.0. With 4 virtual qubits on a 156-qubit backend, this workflow satisfies the conditions under which the original range filter can discard a large fraction of valid measured shots, depending on layout and measurement-register details. In this case, the impact of data loss is not directly measurable: the hardware experiment uses a very small dataset (8 training nodes, 5 testing nodes), and both the hardware and the noisy simulator produce flat loss curves over 30 epochs. As with Ref. [18], we cannot confirm the issue affected their results without access to the raw measurement data, however, the setup matches the conditions for potential data loss in postprocessing. More broadly, performance degradation when moving from simulation to real hardware is common in quantum machine learning and is typically attributed to hardware noise. Our findings show that software-level errors in the postprocessing pipeline can produce a similar effect, and the two are difficult to distinguish without determining whether raw bitstring counts are normalised. For previously published results obtained with affected versions (0.8.4 to 0.9.0), correct inference results can still be recovered if the raw measurement counts were saved, by applying marginalisation as described in Section IV. Training results cannot be corrected after the fact, since the uncorrected loss signal has already been applied to the model weights at each optimisation step. It is worth noting that only SamplerQNN-based workflows may trigger data loss under the conditions above. Other Qiskit ML tools, like quantum kernels or Bayesian inference are not expected to suffer data loss of the same nature, as their internal workflows never call SamplerQNN. VI. C ONCLUSION We analysed and corrected a postprocessing data-loss mechanism in SamplerQNN that, in the hardware experiments we performed, discarded 85 to 99.6% of valid measurement shots. We verified the behaviour in the affected released versions examined here, including Qiskit ML 0.8.4 and 0.9.0, and found the root cause to a filter that assumed bit-strings are in virtual qubit space. This assumption holds for classical simulators but not for real hardware backends where bit-strings span large (over 100) physical qubit registers in the device. Our experiments showed that the issue affects both inference and training. For inference, it can reduce accuracy from 0.94 to 0.39 on identical raw measurements. For training, it degrades the cost function evaluation by 22 to 27×, preventing the classical optimiser from making meaningful updates. Our fix, based on layout marginalisation, was merged into the official repository (PR #1041) and will be available on PyPI from version 0.9.1. Inference results from previous Qiskit
ML versions can be corrected if the raw measurement counts were saved; however, models trained on hardware cannot be corrected post hoc and require re-running with the fix. While researchers routinely validate their own workflows, this case study shows that assumptions at the boundary between user code, library code, transpilation, and hardware outputs also require validation as platforms evolve. Softwarelevel explanations should be considered alongside hardware noise, model expressibility, and optimisation artifacts when interpreting unexpected experimental results. This is especially crucial as software scope and complexity grows with algorithmic and hardware advances. Open-source quantum computing frameworks have been instrumental in making quantum hardware accessible to researchers; maintaining their reliability as hardware scales requires close feedback between users, maintainers, and hardware providers, and benefits the entire community. ACKNOWLEDGMENT The authors thank Laura Martin for her help with data recovery. They also thank the Qiskit ML maintainers who reviewed and merged the fix. SP acknowledges support from the UKRI National Edge AI Hub for Real Data: Edge Intelligence for Cyber-disturbances and Data Quality (EPSRC EP/Y028813/1). NC and JRH acknowledge support from their EPSRC Mathematical Sciences Small Grant (UKRI3647). NC acknowledges support from QCI3 - the Hub for Quantum Computing via Integrated and Interconnected Implementations (EP/Z53318X/1). JRH acknowledges support from a Royal Society Research Grant (RG/R1/251590), and from their EPSRC Quantum Technologies Career Acceleration Fellowship (UKRI1217). This project was funded and supported by the UK National Quantum Computer Centre (NQCC) [NQCC200921], which is a UKRI Centre and part of the UK National Quantum Technologies Programme (NQTP). Access to quantum processing units was enabled through the NQCC SparQ programme, with additional compute time provided through the NQCC’s in-kind contribution to JRH’s EPSRC Quantum Technologies Career Acceleration Fellowship. DATA AND C ODE ACCESSIBILITY All data and code supporting this work can be accessed from doi:10.5281/zenodo.22304590. R EFERENCES [1] B. Ferman and L. Finamor, “There must be an error here! experimental evidence on coding errors’ biases,” 2025, arXiv:2508.20069. [Online]. Available: https://arxiv.org/abs/2508.20069 [2] A. Brodeur et al., “Mass reproducibility and replicability: A new hope,” s.l., I4R Discussion Paper Series 107, 2024. [Online]. Available: https://hdl.handle.net/10419/289437 [3] A. Trisovic, M. K. Lau, T. Pasquier, and M. Crosas, “A largescale study on research code quality and execution,” Scientific Data, vol. 9, no. 1, p. 60, Feb 2022. [Online]. Available: https://doi.org/10.1038/s41597-022-01143-6 [4] A. Callison and N. Chancellor, “Hybrid quantum-classical algorithms in the noisy intermediate-scale quantum era and beyond,” Phys. Rev. A, vol. 106, p. 010101, Jul 2022. [Online]. Available: https://link.aps.org/doi/10.1103/PhysRevA.106.010101
[5] S. S. Gokhale, R. R. Dhote, and R. Delhibabu, “A review of quantum machine learning algorithms, applications, and emerging advantages,” Discover Computing, vol. 29, no. 1, p. 226, Apr 2026. [Online]. Available: https://doi.org/10.1007/s10791-026-10085-1 [6] M. E. Sahin, E. Altamura, O. Wallis, S. P. Wood, A. Dekusar, D. A. Millar, T. Imamichi, A. Matsuo, and S. Mensa, “Qiskit machine learning: an open-source library for quantum machine learning tasks at scale on quantum hardware and classical simulators,” 2025. [7] S. V. Panambalom, “Unexpected results from SamplerQNN when using pre-transpiled circuits with non-trivial qubit layouts,” GitHub Issue #1040, qiskit-community/qiskit-machine-learning, 2026, accessed: 2026-04-27. [Online]. Available: https://github.com/qiskit-community/ qiskit-machine-learning/issues/1040 [8] ——, “Fix SamplerQNN. postprocess layout-based marginalization,” Pull Request #1041, qiskit-community/qiskit-machine-learning, 2026, merged 2026-05-07. [Online]. Available: https://github.com/ qiskit-community/qiskit-machine-learning/pull/1041 [9] S. V. Panambalom, N. Chancellor, and J. R. Hance, “Using quantum neural networks for cancer identification,” In Preparation, 2026. [10] E. Grant, L. Wossnig, M. Ostaszewski, and M. Benedetti, “An initialization strategy for addressing barren plateaus in parametrized quantum circuits,” Quantum, vol. 3, p. 214, 2019. [11] V. Havlı́ček, A. D. Córcoles, K. Temme, A. W. Harrow, A. Kandala, J. M. Chow, and J. M. Gambetta, “Supervised learning with quantumenhanced feature spaces,” Nature, vol. 567, pp. 209–212, 2019. [12] A. Javadi-Abhari, M. Treinish, K. Krsulich, C. J. Wood, J. Lishman, J. Gacon, S. Martiel, P. D. Nation, L. S. Bishop, A. W. Cross, B. R. Johnson, and J. M. Gambetta, “Quantum computing with Qiskit,” 2024. [13] Qiskit Development Team, “Qiskit Aer,” 2024. [Online]. Available: https://github.com/Qiskit/qiskit-aer [14] A. Kumar, N. Sharma, N. K. Marriwala, S. Panda, M. Aruna, and J. Kumar, “Quantum machine learning with qiskit: Evaluating regression accuracy and noise impact,” IET Quantum Communication, vol. 5, no. 4, pp. 310–321, 2024. [15] J. R. McClean, S. Boixo, V. N. Smelyanskiy, R. Babbush, and H. Neven, “Barren plateaus in quantum neural network training landscapes,” Nature Communications, vol. 9, p. 4812, 2018. [16] M. J. D. Powell, “A direct search optimization method that models the objective and constraint functions by linear interpolation,” in Advances in Optimization and Numerical Analysis. Springer, 1994, pp. 51–67. [17] S. Mensa, E. Sahin, F. Tacchino, P. Kl Barkoutsos, and I. Tavernelli, “Quantum machine learning framework for virtual screening in drug discovery: a prospective quantum advantage,” Machine Learning: Science and Technology, vol. 4, no. 1, p. 015023, 2023. [18] A. Martin-Perez, D. Fernández-de-las Heras, M. Terrón-Cuadrado, and M. A. Serrano, “Hybrid classical-quantum transfer learning with noisy quantum circuits,” 2026. [19] P. Chaudhary et al., “Q-AGNN: Quantum-enhanced attentive graph neural network for intrusion detection,” 2026.