Conceptio › Archive › arXiv CS
arXiv CSopen access

Noise-Robust Quantum State Characterization for Remote State Preparation with Deep Learning

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Noise-Robust Quantum State Characterization for Remote State Preparation with Deep Learning Bo Tang,1, ∗ Zixuan Liao,1, ∗ Hao Li,1, ∗ Yilin Yang,1 Jiani Lei,1 Zengya Li,1 Jing Qiu,1 Zhaohui Dong,1 Zhengyang Mao,1 Yuanhua Li,2, † Yuanlin Zheng,1, 3, 4, ‡ and Xianfeng Chen1, 3, 4, 5, §

arXiv:2609.20523v1 [quant-ph] 17 Sep 2026

1

State Key Laboratory of Photonics and Communications, School of Physics and Astronomy, Shanghai Jiao Tong University, Shanghai 200240, China 2 Department of Physics, Shanghai Key Laboratory of Materials Protection and Advanced Materials in Electric Power, Shanghai University of Electric Power, Shanghai 200090, China 3 Hefei National Laboratory, Hefei 230088, China 4 Shanghai Research Center for Quantum Sciences, Shanghai 201315, China 5 Collaborative Innovation Center of Light Manipulations and Applications, Shandong Normal University, Jinan 250358, China (Dated: September 18, 2026) Quantum communication underpins secure information processing and scalable quantum networks. In particular, remote state preparation (RSP) enables efficient quantum state transfer, but accurately estimating target states under complex noise remains challenging. Here, we propose a Transformer-based Quantum State Characterizer (TQSC) model for noisy RSP experiments. Our model reconstructs experimentally prepared pure and mixed photonic polarization states from noisy measurements in complex scattering environments, while its attention patterns provide physically grounded insights into correlations among the measured observables. The method achieves a mean estimator-target fidelity exceeding 99.999% under complex scattering and dynamic Gaussian noise, while its robustness and generalization are further examined using Qiskit-simulated Bloch-ball states. Furthermore, in a practical MNIST image transmission task with held-out states, the decoded bit error rate is reduced from 50.34% to zero after TQSC post-processing. The TQSC model enables accurate tomographic characterization under dynamic noise and provides physically grounded posthoc insights, holding promise for intelligent quantum information processing applications.

I.

INTRODUCTION

Quantum information science drives advancements in communication, computing, and sensing by harnessing quantum mechanics [1–3]. Quantum communication is the cornerstone of the future quantum internet, enabling secure information transfer [4, 5]. Reliable distribution of quantum states between distant nodes is a fundamental task for practical quantum networks [6]. Among protocols, remote state preparation (RSP) is a promising approach [7, 8]. RSP enables Alice to remotely prepare a known quantum state at Bob’s site using shared entanglement and classical communication, offering reduced classical resources than teleportation [9] and improved security over direct transmission [10]. Firstly demonstrated in liquid-state nuclear magnetic resonance [11], RSP is extended to single- and multi-qubit states [12– 14]. Recent advances include photonic states with nonclassical features [15], high-dimensional states [16] using hybrid entanglement [17], orbital-angular-momentum lattices [18], and metasurfaces [19], and implementations in hybrid platforms [20]. Reliable quantum state preparation and characterization is central to practical quantum information processing. In RSP protocols, pure states are essential for computation, communication, and metrology, while mixed states offer noise resilience and resource efficiency, enabling advantages in realistic quantum communication and computation [21, 22]. Accurate reconstruction

through quantum state tomography (QST) and fidelity estimation is therefore critical [23, 24]. Nevertheless, hardware imperfections—including decoherence [25], dynamic scattering [26], control errors [27], and state preparation and measurement (SPAM) errors [28]—invariably constrain the achievable fidelity. To combat complex noise environments, machine learning has emerged as a powerful tool for quantum estimation and control [29, 30]. Specifically, learningbased approaches have been applied to quantum error mitigation [31–34], dynamic state preparation [35, 36], and QST, where neural networks map noisy measurement statistics to density matrices [37–41]. The Transformer [42], featuring self-attention to model sequential dependencies and multi-head attention to capture diverse contextual relationships, has demonstrated strong multimodal performance [43, 44]. In quantum physics, this architecture has been employed across a wide range of problems [45], including electronic structures [46], quantum correlations [47], noise-robust quantum communication [48], automated circuit generation [49], scalable quantum error correction [50]. Previous attention-based QST methods have learned measurement distributions [51], denoised LI/MLE estimates through Cholesky representations under simulated noise [52], and reconstructed states from structured measurements with IBM-device validation [53]. However, these studies do not address quantum-state reconstruction in real physical environments involving com-

2 plex scattering and dynamically varying noise. Moreover, conventional ‘black-box’ neural networks provide limited physical insight, making it difficult to fully trust the reconstructed outputs. High model performance does not necessarily indicate that the model has learned meaningful and task-relevant features [54–56]. Here, we propose a Transformer-based Quantum State Characterizer (TQSC) model with built-in physical constraints for quantum state reconstruction, designed to improve reconstruction fidelity and robustness in noisy RSP while enabling physically grounded post-hoc interpretation. Using photonic polarization as the experimental platform, TQSC reconstructs density matrices from noisy measurements under both MMF transmission alone and MMF transmission with additional dynamic perturbations. Beyond reconstruction fidelity, we further analyze the learned internal representation using statistical, intervention-based, and end-to-end attribution methods rather than relying on attention maps alone. Experimentally, TQSC substantially improves reconstruction performance, achieving an average fidelity above 99.999%. These results demonstrate a measurement-level and experimentally validated framework for robust quantumstate reconstruction in dynamically perturbed optical RSP systems. More broadly, reliable state reconstruction under realistic experimental noise and perturbations may support state monitoring, validation, and diagnostics in practical quantum networks and distributed quantum information processing. II.

THE SCHEME

The positive operator-valued measure (POVM) provides a fundamental framework for quantum measurement theory [57]. In our RSP protocol, the probability of obtaining measurement outcome m is given by † pm = Tr(Mm Mm ρ),

(1)

where the measurement {Mm } satisfy the comP operators † pleteness condition m Mm Mm = I. Following the theoretical framework in Ref. [58], but with a distinct entanglement source, we implement the RSP protocol for mixed states. Taking a mixed state as an example, Bob’s target mixed state to be transmitted is expressed as ρB = α2 |H⟩⟨H| + β 2 |V ⟩⟨V |,

(2)

(where α2 +β 2 = 1). The maximally entangled Bell state generated by our entanglement source takes the form 1 |Φ ⟩AB = √ (|H⟩A |H⟩B + |V ⟩A |V ⟩B ) . 2 +

(3)

After passing through the POVM-based pre-processing module, projective measurements on specific polarization states are implemented using analyzers. When Alice’s photon is projected onto the |H⟩ state, the remote preparation of Bob’s photon collapses to the desired state ρ̂IB

(see eq. (2)). For other measurement outcomes, Bob applies appropriate local unitary operations {σ̂x , σ̂z , σ̂y } to recover the target state. The experimental implementation of RSP is inevitably affected by various physical imperfections. Quantum noise originates from uncontrollable interactions between a quantum system and its environment, leading to deviations from ideal unitary evolution. We describe this noise using the Kraus operator formalism of quantum channels. In our analysis of RSP, we consider four primary sources of noise: photon loss, decoherence, fiber-induced noise, and detector imperfections. The final state of Bob, incorporating these noise effects, is given by   X (k′ ) † ′ ′ ′ρ ρfinal P (k = )E U U . (4) ′ k k B B,ideal k k′

Here, P (k ′ ) denotes the experimentally observed proba(k′ ) bility of Alice obtaining measurement outcome k ′ , ρB,ideal is Bob’s ideal conditional state associated with outcome k ′ before the correction operation, Uk′ is the corresponding correction unitary applied by Bob, and Ek′ represents an outcome-dependent effective trace-preserving noise map for the entire noisy RSP branch, not only for postcorrection noise. The final state is obtained by averaging the noise-affected corrected states over all possible measurement outcomes. Our goal is to suppress this noise. In the context of neural networks, the entire process can be viewed as a complex quantum-to-classical channel, denoted by C, which maps an ideal target quantum state ρtarget onto a set of classical measurement outcomes. The data vector thus constitutes a noise-corrupted quantum state: M = C(ρtarget ).

(5)

The map C formally integrates the four previously discussed noise sources, representing them as a unified framework of coherent and incoherent dynamic disturbances. To counteract these effects, we propose a TQSC framework that employs a neural network to parameterize an inverse channel, denoted as ΦTQSC (·; θ). The predicted quantum state ρpred is obtained from the measurement data M via the inverse map: ρpred = ΦTQSC (M; θ).

(6)

Further details on the POVM implementation, RSP communication noise, and the model are provided in the Supplemental Material (SM). III.

EXPERIMENTAL SETUP

The experimental setup (Fig. 1) utilizes a polarization Sagnac interferometric loop with an integrated spontaneous parametric down-conversion (SPDC) source to generate phase-stable polarization-entangled photon

3

(a)

(b)

HWP

PC

FBS

WDM

PBS

Coupler

DHWP

QWP

FPBS

DWDM

DPBS

SNSPD

PP L

N

Sagnac Loop

MMF

CH31

DM

×

TCSPC

PM AWG

CH33 ODL

Polarization Entanglement

Bob

H Model

Laser

V

Coincidence Counting

Alice

VOA

Source

User Distribution

(c)

Density-Matrix Reconstruction Matrix Reconstruction

Add&Norm Linear

Feed Forward

Unsqueeze

Add&Norm

Input Tensor

Multi-Head Attention

Lower-Triangular Factor C

Diagonal

… Tokens

Data and Pre-processing

Model

Lower Re

Lower Im

Parameter Splitting

Linear

Positional Encoding

+

Linear

Transformer Encoders

Post-processing

Figure 1: Schematic diagram of the experimental setup for RSP. (a) Polarization entanglement source based on Sagnac loop. (b) RSP of user distribution device. PPLN, periodically poled lithium niobate waveguide; DM, dichroic mirror; HWP, halfwave plate; DHWP, dual-wavelength half-wave plate; QWP, quarter-wave plate; PC, polarizaton controller; FBS, fiber beam splitter; FPBS, fiber polarization beam splitter; WDM, wavelength division multiplexing; DWDM, dense wavelength division multiplexing; PM, phase modulator; AWG, arbitrary waveform generator; PBS, polarizing beam splitter; DPBS, dual-polarizing beam splitter; MMF, multimode fiber; ODL, optical delay line; VOA, variable optical attenuator; SNSPD, Superconducting nanowire single-photon detector; TCSPC, timecorrelated single-photon counting. (c) Multidimensional measurement data are processed through a preprocessing module, a three-layer Transformer Encoders module for feature extraction, and a postprocessing module that reconstructs the predicted density matrix. The loss function quantifies the difference between predicted and actual labels, minimized via the Adam optimizer.

pairs. A femtosecond laser pumps a periodically poled lithium niobate (PPLN) waveguide within the loop, producing degenerate photon pairs at the telecom wavelength. Waveplates optimize the source for maximal Bell state generation (Eq. 3). Subsequent wavelength filtering routes photons to separate receivers for Alice and Bob. Alice’s detection module employs a beam splitter to divide the incoming photon. One path incorporates a tunable delay for synchronization, while the other includes a variable attenuator to adjust the intensity ratio η = α2 /β 2 between measurement bases. Polarization controllers (PC) compensate fiber effects before polarization projection. A hybrid detection scheme combines specific polarization outputs at a second beam splitter for

state analysis, followed by final projective measurement. Bob’s receiver includes a PC for fiber-effect compensation, followed by a segment of MMF mimicking spatially scrambling in scattering media [59, 60]. Dynamic noise is generated by driving the phase modulator (PM) with signals from the arbitrary waveform generator (AWG). QST is performed with a polarization analyzer composed of a quarter-wave plate (QWP), a half-wave plate (HWP), and a polarizing beam splitter (PBS) arranged in sequence. For preparing pure states, the PBS in Alice’s polarization analyzer is retained and the QWP at the entanglement source is adjusted to control the phase. For mixed states, the PBS in Alice’s analyzer is removed. Both users employ superconducting nanowire single-photon detectors (SNSPDs). Detection events

4 (a)

(b)

IV.

0.8

0.6

0.4

0.2

0

-0.2

Initially, we prepare a set of mixed states with η ranging from 1 to 7, given by ρη =

Real Part

Imaginary Part

(c) 1.0 0.9

Fidelity

RESULT

0.8 0.7 0.6 0.5

Figure 2: Characterization and verification of RSP. (a) Remote quantum states are prepared on the Bloch ball, with the pure state shown as a red dot and mixed states shown as blue dots [61]. (b) The real and imaginary parts of the density matrix with η = 1. (c) Reconstruction fidelity of the pure state and mixed states prepared via RSP through an MMF.

1 (|H⟩⟨H| + η|V ⟩⟨V |). η+1

Furthermore, we prepare a pure state |ψp ⟩ as  1  |ψp ⟩ = √ 3|H⟩ + eiπ/3 |V ⟩ . 10

The structure of the TQSC is illustrated in Fig. 1(c). The model takes a 13-dimensional input vector representing measurement data for quantum states and outputs the corresponding density matrix. Input data points are treated as tokens. After preprocessing with linear transformation and dimension adjustment, the data passes through a stack of Transformer encoders. These incorporate positional encoding and multihead self-attention to capture dependencies between sequence elements, with residual connections and layer normalization stabilizing the learning process. A feedforward network then extracts deeper nonlinear features. In the post-processing stage, the extracted features are linearly transformed into real factor parameters and split into diagonal, lower-real, and lower-imaginary components. These components form a lower-triangular complex factor C, and the final density matrix is reconstructed as ρ = CC † /Tr(CC † ). Model training minimizes the Mean Squared Error (MSE) between predicted and true density matrices using the Adam optimizer. Details of the experimental setup, model architecture, and training parameters are provided in the SM.

(8)

The corresponding data points are shown in Fig. 2(a). QST uses the projection bases H, V , D, and R, corresponding to horizontal, vertical, diagonal, and rightcircular polarizations, respectively. We prepare eight quantum states, averaging five measurements per basis without the MMF to benchmark state-preparation fidelity, and utilize single-shot measurements for all subsequent MMF data. To verify the prepared states, we reconstruct the density matrix ρB using the maximum likelihood method [62] and calculate the fidelity ⟨F ⟩ with the target prepared state ρp [63]. The fidelity is computed using the following formula:  q√ √ 2 ρp ρB ρp F (ρp , ρB ) = Tr .

are recorded with time-correlated single-photon counting (TCSPC).

(7)

(9)

Without the MMF, the average fidelity of these states reaches (99.25 ± 0.20)%, demonstrating the high statepreparation capability of our experimental setup. Upon introducing the MMF, however, the fidelities of the states decrease to varying degrees, as illustrated in Fig. 2(c). As an example, the real and imaginary parts of the density matrix for the case η = 1 are reconstructed and shown in Fig. 2(b). The impact of dynamic disturbances on the system coherence is then evaluated. In the absence of noise, the prepared quantum states exhibit a purity of 1 and a fidelity of (99.23 ± 0.20)%. When zero-mean Gaussian dynamic noise (σ = 1 V) generated by the AWG is introduced by driving the PM (Vπ = 5 V) in the MMF setup, the purity decreases to 0.6724 ± 0.0698, while the fidelity drops to (38.31 ± 12.30)%, indicating significant decoherence and fidelity degradation under the applied noise conditions. To overcome these limitations, our model is comprehensively trained on pure states, mixed states, and pure states under dynamic noise, achieving an overall average estimation fidelity of 0.99999497 ± 8.78 × 10−6 when evaluated on both training and test sets. As illustrated in Fig. 3(a), the estimation fidelity distributions for both the training and test sets are tightly clustered around values approaching 1, indicating precise agreement between predicted results and theoretical values. To benchmark the reconstruction performance, we compare the TQSC

5 (a) 4

(b)

10

Train Data Test Data 3

(a)

Coinc. H/V/D/R Alice singles

Send

Bob singles Alice D/R

102

Image Reconstruction over Noisy Channel Receive MMF Transmission

TQSC Inference

··· 0 1 0 0 1 1 ··· ··· ···

··· 0 1 0 0 1 1 ··· ··· ···

Bob D/R Alice H/V Bob H/V

100 0.99990

0.99995

1.00000

Fidelity

(c)

BTC Compressed −10−7 0

10−7 10−6 10−5

Excess fidelity drop

(d)

(b)

Noisy Channel TQSC Model

3

0.20

0.06

12

12

0.05

0

3

6

9

12

0

3

6

9

12

0.00

Figure 3: Performance and attention analysis of the TQSC model. (a) Fidelity distributions for the training and test datasets. (b) Excess fidelity drop from targeted token-group interventions relative to size-matched random ablations. (c) Mean attention map of Layer 3 Head 4 (L3H4). (d) Attention variance map of L3H4. In (c) and (d), the query (y) and key (x) axes index the 13 input tokens (0–12), comprising H/V/D/R coincidence counts, the corresponding Alice and Bob single-photon counts, and measurement time.

model with representative baselines; it achieves the highest average test fidelity, with a slight advantage over the parameter-matched MLP baseline. Furthermore, the model demonstrates robust generalization to new parameter values, maintaining estimation fidelities above 99% in tests involving both interpolation (η = 4.5) and leave-one-out validation (η = 6). In addition, we test the model on Qiskit-simulated Bloch-ball states beyond the experimentally measured state family, where it also achieves estimation fidelities exceeding 99%. This result demonstrates that the model maintains high estimation fidelity on unseen test data with a slight performance degradation, indicating strong generalization ability and robustness without evident signs of overfitting. Using only single-shot measurement data as input, the TQSC model reliably reconstructs quantum states with high estimation fidelity. Recent MPO-based methods target large-scale mixedstate learning under noise [64], rather than direct purestate optimization under dynamic noise. Admittedly, the train set scattered high-fidelity outliers suggest potential sensitivity of the model to edgecase data. Additionally, the limited sample size of the test set (n < 400) may necessitate expansion of the validation set to enhance statistical inference reliability. L3H4-targeted token-group ablations in Fig. 3(b) re-

Bit 0

122

475

597

0

500 400 300

Bit 1

117

462

0

579

Bit 0

Bit 1

Bit 0

Bit 1

200 100 0

0.02

9

6 9

0.10

6

0.04 0.15

Overlap

Actual Bits

0 3

0.25

Model Denoised Result BER = 0%

(c)

Attention variance

0

Mean attention

Reconstructed with Noise BER = 50.34%

Received Bits

Number of Samples

101

Probability

Count (log)

10

Predicted Bits

Number of outcomes (N=100)

Figure 4: TQSC-based image transmission and reconstruction. (a) Degradation and recovery of BTC-compressed MNIST images through the noisy channel. (b) Probability distributions of |V ⟩ measurement outcomes (N = 100) for mixed states ρ2 and ρ3 . (c) Confusion matrices for the noisy channel (left) and TQSC predictions (right).

veal a nonuniform local feature dependence, dominated by Bob-side single counts, with smaller positive contributions from coincidence and Bob polarization features. To gain insight into how the TQSC model integrates experimental information when optimizing quantumstate fidelity, we analyze the attention maps of Layer 3 Head 4 (L3H4), a deeper Transformer layer that may capture higher-order, task-relevant dependencies among input features [65]. As shown in Fig. 3(c), the mean attention map exhibits localized high-attention regions, indicating preferential interactions among selected query–key pairs. The attention variance in Fig. 3(d) is concentrated on a smaller subset of pairs, showing that only selected interactions vary appreciably across samples. Together, these features indicate selective rather than global information integration. Quantitative analysis of these localized attention patterns further reveals enhanced Bob H/V–Bob D/R and Coinc. H/V–Coinc. D/R interactions, consistent with the model integrating complementary population- and coherence/phase-sensitive information. In addition, interactions between the coincidence features and acquisition time suggest that the model interprets the coincidence statistics in conjunction with the measurement duration, which determines the degree of statistical averaging and hence the reliability of the measured counts. Together, these observations suggest that the deeper-layer attention does not merely emphasize individual observables, but captures structured correlations among multi-

6 ple experimentally relevant features. Quantitative attention and intervention analyses are provided in the SM. Ultimately, to further validate the practicality of the proposed scheme, we conduct an image transmission experiment under the same system configuration as described above. In the complex scattering environment of MMF, mixed states ρ2 and ρ3 are respectively employed to encode bits 0 and 1, and fidelity-based criteria are used for bit decision. In Fig. 4(b), with the state fidelity of F (ρ2 , ρ3 ) ≈ 99.16%, measurement-induced Gaussian noise causes the distributions to overlap, making the states difficult to resolve even without the MMF. The N = 100 measurements are used only to characterize the projection-outcome distributions and are not used for decoding or BER evaluation. A grayscale image of the handwritten digit “0” from the MNIST dataset is selected. After BTC compression [66], the 28 × 28 pixel image is encoded into 1176 bits, followed by quantum state preparation and transmission using the experimental setup. The transmitted quantum-state data are all held-out states unseen by the model, with each state measured only once. As shown in Fig. 4(a), the original bit error rate (BER) is 50.34%; after applying the TQSC model, the BER is reduced to 0%. As depicted in Fig. 4(c), the confusion matrices show bit errors in the noisy channel and accurate recovery by the TQSC model. These results demonstrate the strong practicality and generalization capability of the model. Additional results on baseline comparisons, generalization validation, the explanation of “key” and “query”, and further details of the image transmission experiments are provided in the SM. V.

CONCLUSION

In summary, we have proposed and implemented a TQSC-assisted tomographic characterization scheme for mixed- and pure-state reconstruction under complex scattering, with additional dynamic Gaussian noise for pure states. Through this framework, our approach achieves an average estimator-target fidelity exceeding 99.999% between the reconstructed and target states in the RSP experiments, while simultaneously providing interpretable data-driven insights into measurement correlations. Although the training phase requires a dataset of paired (M, ρtarget ) examples, during the inference stage, the scheme does not require additional quantum operations and performs classical tomographic denoising based on a single acquired measurement vector. Moreover, it demonstrates high single-acquisition efficiency, generalization across the tested conditions, and interpretability, providing a data-driven approach to quantum state estimation in complex noisy environments. While our current implementation demonstrates remarkable performance in a controlled environment, we acknowledge that its scalability and resilience against more complex noise models warrant further investigation.

Future work will focus on integrating this scheme as a post-processing and diagnostic tool in larger, multi-node quantum networks and exploring its compatibility with different quantum information protocols. These continued efforts are expected to further unlock the potential of our approach, paving the way for its broader application in intelligent quantum communication and quantum state characterization.

VI.

ACKNOWLEDGMENTS

This work is supported in part by the National Natural Science Foundation of China (Grants No. 62375164, No. 12574360 and No. 12192252), the Foundation for Shanghai Municipal Science Technology Major Project (Grant No. 2019SHZDZX01-ZX06), Quantum Science and Technology-National Science and Technology Major Project (Grant No. 2021ZD0300802), the Shuguang Program of Shanghai Education Development Foundation and Shanghai Municipal Education Commission (Grant No. 24SG53), and Guangdong Provincial Quantum Science Strategic Initiative (Grant No. GDZX2403003) and Shanghai Oriental Talent Plan Youth Project (Grant no. QNKJ2025013).

∗

These authors contributed equally to this work. [email protected] ‡ [email protected] § [email protected] [1] D. Main, P. Drmota, D. P. Nadlinger, E. M. Ainley, A. Agrawal, B. C. Nichol, R. Srinivas, G. Araneda, and D. M. Lucas, Distributed quantum computing across an optical network link, Nature 638, 383 (2025). [2] M. Pittaluga, Y. S. Lo, A. Brzosko, R. I. Woodward, D. Scalcon, M. S. Winnel, T. Roger, J. F. Dynes, K. A. Owen, S. Juárez, P. Rydlichowski, D. Vicinanza, G. Roberts, and A. J. Shields, Long-distance coherent quantum communications in deployed telecom networks, Nature 640, 911 (2025). [3] P. Bhattacharyya, W. Chen, X. Huang, S. Chatterjee, B. Huang, B. Kobrin, Y. Lyu, T. J. Smart, M. Block, E. Wang, et al., Imaging the meissner effect in hydride superconductors using quantum sensors, Nature 627, 73 (2024). [4] H. J. Kimble, The quantum internet, Nature 453, 1023 (2008). [5] Y. Yang, Y. Li, H. Li, C. Wu, Y. Zheng, and X. Chen, A 300-km fully-connected quantum secure direct communication network, Science Bulletin (2025). [6] S. Wehner, D. Elkouss, and R. Hanson, Quantum internet: A vision for the road ahead, Science 362, eaam9288 (2018). [7] A. K. Pati, Minimum classical bit for remote preparation and measurement of a qubit, Physical Review A 63, 014302 (2000). [8] H.-K. Lo, Classical-communication cost in distributed quantum-information processing: a generalization of †

7 quantum-communication complexity, Physical Review A 62, 012313 (2000). [9] C. H. Bennett, G. Brassard, C. Crépeau, R. Jozsa, A. Peres, and W. K. Wootters, Teleporting an unknown quantum state via dual classical and einstein-podolskyrosen channels, Phys. Rev. Lett. 70, 1895 (1993). [10] S. Pogorzalek, K. Fedorov, M. Xu, A. Parra-Rodriguez, M. Sanz, M. Fischer, E. Xie, K. Inomata, Y. Nakamura, E. Solano, et al., Secure quantum remote state preparation of squeezed microwave states, Nat. Commun. 10, 2604 (2019). [11] X. Peng, X. Zhu, X. Fang, M. Feng, M. Liu, and K. Gao, Experimental implementation of remote state preparation by nuclear magnetic resonance, Phys. Lett. A 306, 271 (2003). [12] N. A. Peters, J. T. Barreiro, M. E. Goggin, T.-C. Wei, and P. G. Kwiat, Remote state preparation: Arbitrary remote control of photon polarization, Phys. Rev. Lett. 94, 150502 (2005). [13] G.-Y. Xiang, J. Li, B. Yu, and G.-C. Guo, Remote preparation of mixed states via noisy entanglement, Phys. Rev. A 72, 012315 (2005). [14] M. Rådmark, M. Wieśniak, M. Żukowski, and M. Bourennane, Experimental multilocation remote state preparation, Phys. Rev. A 88, 032304 (2013). [15] S. Liu, D. Han, N. Wang, Y. Xiang, F. Sun, M. Wang, Z. Qin, Q. Gong, X. Su, and Q. He, Experimental demonstration of remotely creating wigner negativity via quantum steering, Phys. Rev. Lett. 128, 200401 (2022). [16] M. Erhard, M. Krenn, and A. Zeilinger, Advances in high-dimensional quantum entanglement, Nat. Rev. Phys. 2, 365 (2020). [17] M. Erhard, H. Qassim, H. Mand, E. Karimi, and R. W. Boyd, Real-time imaging of spin-to-orbital angular momentum hybrid remote state preparation, Phys. Rev. A 92, 022321 (2015). [18] A. R. Cameron, S. W. Cheng, S. Schwarz, C. Kapahi, D. Sarenac, M. Grabowecky, D. G. Cory, T. Jennewein, D. A. Pushin, and K. J. Resch, Remote state preparation of single-photon orbital-angular-momentum lattices, Phys. Rev. A 104, L051701 (2021). [19] M. Ning, N. Li, H. Zhong, Q. Yuan, L.-e. Zhang, S. Wang, M.-K. Chen, X. Qiu, X. Ren, X. Hu, et al., Highdimensional remote state preparation with metasurfacebased hyper-entangled photon source, Laser Photonics Rev. , e01968 (2025). [20] F.-X. Sun, S.-S. Zheng, Y. Xiao, Q. Gong, Q. He, and K. Xia, Remote generation of magnon schrödinger cat state via magnon-photon entanglement, Phys. Rev. Lett. 127, 087203 (2021). [21] E. Chitambar and G. Gour, Quantum resource theories, Reviews of modern physics 91, 025001 (2019). [22] K. Modi, A. Brodutch, H. Cable, T. Paterek, and V. Vedral, The classical-quantum boundary for correlations: Discord and related measures, Rev. Mod. Phys. 84, 1655 (2012). [23] X. Zhang, M. Luo, Z. Wen, Q. Feng, S. Pang, W. Luo, and X. Zhou, Direct fidelity estimation of quantum states using machine learning, Physical Review Letters 127, 130503 (2021). [24] H. Qin, L. Che, C. Wei, F. Xu, Y. Huang, and T. Xin, Experimental direct quantum fidelity learning via a datadriven approach, Physical Review Letters 132, 190801

(2024). [25] W. H. Zurek, Decoherence, einselection, and the quantum origins of the classical, Rev. Mod. Phys. 75, 715 (2003). [26] H. Defienne, M. Reichert, and J. W. Fleischer, Adaptive quantum optics with spatially entangled photon pairs, Physical review letters 121, 233601 (2018). [27] E. Magesan, J. M. Gambetta, and J. Emerson, Scalable and robust randomized benchmarking of quantum processes, Phys. Rev. Lett. 106, 180504 (2011). [28] R. Blume-Kohout, J. K. Gamble, E. Nielsen, K. Rudinger, J. Mizrahi, K. Fortier, and P. Maunz, Demonstration of qubit operations below a rigorous fault tolerance threshold with gate set tomography, Nat. Commun. 8, 14485 (2017). [29] H. Ma, B. Qi, I. R. Petersen, R.-B. Wu, H. Rabitz, and D. Dong, Machine learning for estimation and control of quantum systems, National Science Review 12, nwaf269 (2025). [30] G. Torlai and R. G. Melko, Machine-learning quantum states in the nisq era, Annual Review of Condensed Matter Physics 11, 325 (2020). [31] A. Strikis, D. Qin, Y. Chen, S. C. Benjamin, and Y. Li, Learning-based quantum error mitigation, PRX Quantum 2, 040330 (2021). [32] Z. Cai, R. Babbush, S. C. Benjamin, S. Endo, W. J. Huggins, Y. Li, J. R. McClean, and T. E. O’Brien, Quantum error mitigation, Reviews of Modern Physics 95, 045005 (2023). [33] H. Liao, D. S. Wang, I. Sitdikov, C. Salcedo, A. Seif, and Z. K. Minev, Machine learning for practical quantum error mitigation, Nature Machine Intelligence 6, 1478 (2024). [34] M. Liao, Y. Zhu, G. Chiribella, and Y. Yang, Noiseagnostic quantum error mitigation with data augmented neural models, npj Quantum Information 11, 8 (2025). [35] Z.-M. Wang and T.-Z. Chen, Adaptive denoising quantum state preparation in a dynamic environment, Physical Review Research 6, 043195 (2024). [36] C.-C. Li, R.-H. He, and Z.-M. Wang, Enhanced quantum state preparation via stochastic predictions of neural networks, Physical Review A 108, 052418 (2023). [37] G. Torlai, G. Mazzola, J. Carrasquilla, M. Troyer, R. Melko, and G. Carleo, Neural-network quantum state tomography, Nat. Phys. 14, 447 (2018). [38] J. Carrasquilla, G. Torlai, R. G. Melko, and L. Aolita, Reconstructing quantum states with generative models, Nature Machine Intelligence 1, 155 (2019). [39] A. M. Palmieri, E. Kovlakov, F. Bianchi, D. Yudin, S. Straupe, J. D. Biamonte, and S. Kulik, Experimental neural network enhanced quantum tomography, npj Quantum Information 6, 20 (2020). [40] S. Ahmed, C. Sanchez Munoz, F. Nori, and A. F. Kockum, Quantum state tomography with conditional generative adversarial networks, Phys. Rev. Lett. 127, 140502 (2021). [41] Y. Hu, M. Ma, and J. Shang, Error-mitigated quantum state tomography using neural networks, arXiv:2602.09733 [quant-ph] (2026). [42] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. K. Kaiser, and I. Polosukhin, Attention is all you need, Adv. Neural Inf. Process. Syst. 30 (2017).

8 [43] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, An image is worth 16x16 words: Transformers for image recognition at scale, in International Conference on Learning Representations (2021). [44] L. Dong, S. Xu, and B. Xu, Speech-transformer: A norecurrence sequence-to-sequence model for speech recognition, Proc. IEEE Int. Conf. Acoust., Speech Signal Process. (ICASSP) 2018, 5884 (2018). [45] H. Zhang, Q. Zhao, M. Zhou, L. Feng, D. Niyato, S. Zheng, and L. Chen, A survey of quantum transformers: Architectures, challenges and outlooks, arXiv preprint arXiv:2504.03192 (2025). [46] J. A. Sobral, M. Perle, and M. S. Scheurer, Physicsinformed transformers for electronic quantum states, Nature Communications (2025). [47] Y.-H. Zhang and M. Di Ventra, Transformer quantum state: A multipurpose model for quantum many-body problems, Physical Review B 107, 075147 (2023). [48] Y. Li, Z. Shi, H. Ma, L. Shen, J. Bao, and Y. Xiao, Language model for large-text transmission in noisy quantum communications, arXiv preprint arXiv:2504.20842 (2025). [49] S. Daimon and Y.-i. Matsushita, Quantum circuit generation for amplitude encoding using a transformer decoder, Physical Review Applied 22, L041001 (2024). [50] J. Bausch, A. W. Senior, F. J. H. Heras, T. Edlich, A. Davies, M. Newman, C. Jones, K. Satzinger, M. Y. Niu, S. Blackwell, G. Holland, D. Kafri, J. Atalaya, C. Gidney, D. Hassabis, S. Boixo, H. Neven, and P. Kohli, Learning high-accuracy error decoding for quantum processors, Nature 635, 834 (2024). [51] P. Cha, P. Ginsparg, F. Wu, J. Carrasquilla, P. L. McMahon, and E.-A. Kim, Attention-based quantum tomography, Machine Learning: Science and Technology 3, 01LT01 (2022). [52] A. M. Palmieri, G. Müller-Rigat, A. K. Srivastava, M. Lewenstein, G. Rajchel-Mieldzioć, and M. Plodzień, Enhancing quantum state tomography via resourceefficient attention-based neural networks, Phys. Rev. Res. 6, 033248 (2024). [53] H. Ma, Z. Sun, D. Dong, C. Chen, and H. Rabitz, Tomography of quantum states from structured measurements via quantum-aware transformer, IEEE Transactions on Cybernetics (2025). [54] M. T. Ribeiro, S. Singh, and C. Guestrin, ” why should i trust you?” explaining the predictions of any classi-

fier, in Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining (2016) pp. 1135–1144. [55] S. Lapuschkin, S. Wäldchen, A. Binder, G. Montavon, W. Samek, and K.-R. Müller, Unmasking clever hans predictors and assessing what machines really learn, Nature communications 10, 1096 (2019). [56] A. J. DeGrave, J. D. Janizek, and S.-I. Lee, Ai for radiographic covid-19 detection selects shortcuts over signal, Nature Machine Intelligence 3, 610 (2021). [57] K. Kraus, A. Böhm, J. D. Dollard, and W. Wootters, States, effects, and operations fundamental notions of quantum theory: Lectures in mathematical physics at the university of Texas at Austin (Springer, 1983). [58] W. Wu, W.-T. Liu, P.-X. Chen, and C.-Z. Li, Deterministic remote preparation of pure and mixed polarization states, Phys. Rev. A 81, 042301 (2010). [59] L.-Y. Yu and S. You, High-fidelity and high-speed wavefront shaping by leveraging complex media, Science Advances 10, eadn2846 (2024). [60] M. W. Matthès, Y. Bromberg, J. De Rosny, and S. M. Popoff, Learning and avoiding disorder in multimode fibers, Physical Review X 11, 021060 (2021). [61] N. Lambert, E. Giguère, P. Menczel, B. Li, P. Hopf, G. Suárez, M. Gali, J. Lishman, R. Gadhvi, R. Agarwal, A. Galicia, N. Shammah, P. Nation, J. Johansson, S. Ahmed, S. Cross, A. Pitchford, and F. Nori, Qutip 5: The quantum toolbox in python, Physics Reports 1153, 1 (2026). [62] D. F. V. James, P. G. Kwiat, W. J. Munro, and A. G. White, Measurement of qubits, Phys. Rev. A 64, 052312 (2001). [63] J. B. Altepeter, E. R. Jeffrey, and P. G. Kwiat, Photonic state tomography, Adv. At. Mol. Opt. Phys. 52, 105 (2005). [64] M. Votto, M. Ljubotina, C. Lancien, J. I. Cirac, P. Zoller, M. Serbyn, L. Piroli, and B. Vermersch, Learning mixed quantum states in large-scale experiments, Physical Review Letters 136, 090801 (2026). [65] G. Jawahar, B. Sagot, and D. Seddah, What does bert learn about the structure of language?, in Proceedings of the 57th annual meeting of the association for computational linguistics (2019) pp. 3651–3657. [66] E. Delp and O. Mitchell, Image compression using block truncation coding, IEEE transactions on Communications 27, 1335 (2003).

Supplementary Material for: Noise-Robust Quantum State Characterization for Remote State Preparation with Deep Learning Bo Tang,1, ∗ Zixuan Liao,1, ∗ Hao Li,1, ∗ Yilin Yang,1 Jiani Lei,1 Zengya Li,1 Jing Qiu,1 Zhaohui Dong,1 Zhengyang Mao,1 Yuanhua Li,2, † Yuanlin Zheng,1, 3, 4, ‡ and Xianfeng Chen1, 3, 4, 5, § 1

State Key Laboratory of Photonics and Communications, School of Physics and Astronomy, Shanghai Jiao Tong University, Shanghai 200240, China 2 Department of Physics, Shanghai Key Laboratory of Materials Protection and Advanced Materials in Electric Power, Shanghai University of Electric Power, Shanghai 200090, China 3 Hefei National Laboratory, Hefei 230088, China 4 Shanghai Research Center for Quantum Sciences, Shanghai 201315, China 5 Collaborative Innovation Center of Light Manipulations and Applications, Shandong Normal University, Jinan 250358, China (Dated: September 18, 2026)

I.

PREPARATION OF MIXED STATE THEORY SCHEME

Bob’s photon onto one of the four mixed states: ⊥ ρ̂IB = p2 |φB ⟩⟨φB | + q 2 |φ⊥ B ⟩⟨φB |,

Consider that the desired mixed state is [1] ⊥ ρB = p2 |φB ⟩⟨φB | + q 2 |φ⊥ B ⟩⟨φB |

(S1)

with iϕ

|φB ⟩ = α|HB ⟩ + βe |VB ⟩,

(S2)

−iϕ |φ⊥ |HB ⟩ − α|VB ⟩, B ⟩ = βe

(S3)

where α, β ∈ C with |α|2 + |β|2 = 1, and p, q are probability amplitudes satisfying p2 + q 2 = 1. Without loss of generality, we assume that p, q are real numbers, p2 + q 2 = 1, and α, β, ϕ are the same as before. To prepare arbitrary mixed states we need to achieve complete control over all five parameters. In Alice’s setup, two polarization controllers, PC1 and PC2, are inserted in the two distinct optical paths to rotate the polarization states as |H⟩ → p|H⟩ + q|V ⟩,

(S4)

|V ⟩ → p|H⟩ + q|V ⟩.

(S5)

To achieve independent polarization control in both paths, two additional polarization controllers, PC1′ and PC2′ , are employed. These are adjusted to implement the transformation

(S8b) ⊥ ρ̂YB = p2 (σ̂y |φB ⟩)(⟨φB |σ̂y ) + q 2 (σ̂y |φ⊥ B ⟩)(⟨φB |σ̂y ), (S8c) 2 2 ⊥ ⊥ ρ̂Z B = p (σ̂z |φB ⟩)(⟨φB |σ̂z ) + q (σ̂z |φB ⟩)(⟨φB |σ̂z ). (S8d)

Bob can obtain the desired mixed state by applying local ˆ σ̂x , σ̂y or σ̂z according to Alice’s unitary operations I, measurement results. The required classical communication is two bits. For simplicity, we do not use PC1′ and PC2′ to adjust the transformation. The final mixed states are as follows:

ρ̂IB = α2 |H⟩⟨H| + β 2 |V ⟩⟨V |,

(S9a)

2 2 ρ̂X B = α (σ̂x |H⟩)(⟨H|σ̂x ) + β (σ̂x |V ⟩)(⟨V |σ̂x ), ρ̂YB = α2 (σ̂y |H⟩)(⟨H|σ̂y ) + β 2 (σ̂y |V ⟩)(⟨V |σ̂y ),

(S9b)

2 2 ρ̂Z B = α (σ̂z |H⟩)(⟨H|σ̂z ) + β (σ̂z |V ⟩)(⟨V |σ̂z ).

(S9d)

|H⟩ → α|H⟩ + βeiϕ |V ⟩,

(S6)

−iϕ

(S7)

|H⟩ − α|V ⟩.

Then POVM measurement described by M1 and M2 are performed on the H-polarized and V-polarized components. Then the Alice’s side HWP (∼22.5◦ ) and the detectors perform the projection measurement , which projects

(S9c)

For pure states, the derivation follows analogously to the mixed-state case (or refer to Ref.[2]), similarly leading to: |ϕ⟩ =

|V ⟩ → βe

(S8a)

2 2 ⊥ ⊥ ρ̂X B = p (σ̂x |φB ⟩)(⟨φB |σ̂x ) + q (σ̂x |φB ⟩)(⟨φB |σ̂x ),

1 {|H2A′ ⟩(α|H B ⟩ + βeiφ |V B ⟩) 2 − |V2A′ ⟩[σ̂z (α|H B ⟩ + βeiφ |V B ⟩)]

(S10)

+ |H1A′ ⟩[σ̂x (α|H B ⟩ + βeiφ |V B ⟩)] − |V1A′ ⟩[iσ̂y (α|H B ⟩ + βeiφ |V B ⟩)]}, In a similar manner, Bob can obtain the desired pure ˆ σ̂x , σ̂y , or σ̂z state by applying local unitary operations I, according to Alice’s measurement results. The required

2 classical communication is two bits. 1 (α|H B ⟩ + βeiφ |V B ⟩), 2 1 X |ψ⟩B = σ̂x (α|H B ⟩ + βeiφ |V B ⟩), 2 1 Y |ψ⟩B = iσ̂y (α|H B ⟩ + βeiφ |V B ⟩), 2 1 Z |ψ⟩B = σ̂z (α|H B ⟩ + βeiφ |V B ⟩). 2 I

|ψ⟩B =

II.

(S11a) (S11b) (S11c) (S11d)

QUANTUM NOISE MODELS

Quantum noise arises from the uncontrollable interaction between a quantum system and its environment, causing the system’s evolution to deviate from ideal unitary dynamics. We employ the framework of quantum channels to mathematically describe such noise. A quantum channel is a completely positive (CP) and trace-nonincreasing (TNI) map that transforms an input density matrix ρ into an output state E(ρ). According to Kraus’ theorem [3], any completely positive trace-preserving (CPTP) map can be expressed as E(ρ) =

X

Ek ρEk† ,

(S12)

k

where {Ek } are the Kraus operators, satisfying P † For tracek Ek Ek = I (trace-preserving condition). P non-increasing maps (lossy channels), k Ek† Ek ≤ I. When multiple noise channels E1 , E2 , . . . , EN act sequentially, the overall effect is described by the composite channel ETotal = EN ◦ EN −1 ◦ · · · ◦ E1 ,

For a single qubit with loss probability p, the Kraus operators are   √   0 p 1 √ 0 , E1 = . (S15) E0 = 0 0 0 1−p Here E0 corresponds to no photon loss, while E1 describes decay from |1⟩ to |0⟩. If Alice and Bob’s photons experience independent losses with probabilities pA and pB , the overall channel is   (A) (B) ELoss (ρAB ) = EAD (pA ) ⊗ EAD (pB ) (ρAB ), (S16) (Loss)

with Kraus operators Eij

(A)

= Ei

(B)

⊗Ej

, i, j ∈ {0, 1}.

1.2. Phase Damping Channel (PDC)

Phase damping accounts for decoherence without energy exchange, e.g., random refractive index fluctuations or scattering-induced phase noise. For a single qubit with decoherence probability q, the Kraus operators are     0 0 1 √ 0 √ . , F1 = (S17) F0 = 0 q 0 1−q With decoherence probabilities qA and qB for Alice and Bob, respectively:   (A) (B) EDecoherence (ρAB ) = EP D (qA ) ⊗ EP D (qB ) (ρAB ), (Decoherence)

with Kraus operators Fkl

(A)

= Fk

(B)

⊗ Fl

(S18) .

(S13)

whose Kraus operators are given by chained products:

1.3. MMF Induced Noise

Kj = EN,kN EN −1,kN −1 · · · E1,k1 ,

When a single-mode input qubit is transmitted through a segment of MMF, the different spatial modes supported by the MMF experience varying propagation constants and mode-mixing. This can lead to mode dispersion, differential modal attenuation, and effectively, a loss of coherence or entanglement if the output modes are not perfectly demultiplexed or if the measurement setup cannot distinguish them. For a qubit encoded in polarization or time bins, the interaction with multiple spatial modes can scramble the encoded information, acting as a depolarizing-like channel or a more complex unitary scrambling. A simplified model treats the MMF as inducing random unitary transformations or causing effective dephasing and amplitude damping through mode coupling and differential loss, ultimately leading to a mixed state for the output single mode. We model the MMF noise on Bob’s qubit as a generalized depolarizing channel or a channel that effectively

(S14)

where j indexes all possible choices of {k1 , . . . , kN }.

1. Representative Noise Models

In our analysis of remote state preparation (RSP), we focus on four dominant sources of noise: photon loss, decoherence, multimode fiber (MMF) induced noise, and detector imperfections.

1.1. Amplitude Damping Channel (ADC)

Photon loss during transmission (e.g., absorption or scattering in optical fibers or free-space channels) can be modeled by the amplitude damping channel.

3 mixes the input state due to coupling to higher-order modes that are subsequently lost or unmeasured. For simplicity, we model the MMF as a depolarizing channel on Bob’s qubit. The Kraus operators for a depolarizing channel on Bob’s qubit are: q 1 − 3dMMF I (B) , 4 q (B) , H1 = dMMF 4 X q (B) H2 = dMMF , 4 Y q (B) , H3 = dMMF 4 Z H0 =

(S19)

2. Physical Mechanism of Fidelity Degradation in Multi-Mode Fibers and Noise Modeling Analysis

This section provides a physical interpretation of the fidelity degradation observed when introducing MultiMode Fiber (MMF) into the quantum channel and elucidates the role of the noise model employed in this study within this context.

(S20) 1.

Physical Principles of MMF-Induced Degradation

(S21) (S22)

where dMMF is the depolarization probability due to the MMF, and {I, X, Y, Z} are Pauli matrices. Since the MMF noise is specific to Bob’s transmission path, the overall channel for the two-qubit system is:  (B) EMMF (ρAB ) = I (A) ⊗ EDepolarizing (dMMF ) (ρAB ),

Multi-mode fibers support the simultaneous propagation of multiple spatial modes when the core radius or the refractive index contrast between the core and cladding is sufficiently large. The approximate number of supported modes, D, is given by:  D≈

2π λ

2 Z R

 2 1/2 n (r) − n2c r dr,

(S29)

0



(S23) (B) (MMF) = I (A) ⊗ Hm , for m ∈ with Kraus operators Hm {0, 1, 2, 3}.

1.4. Detector Imperfections

Detector efficiency (q) can be absorbed into the loss model, with effective loss p = 1 − q. Spurious detector clicks (dark counts) introduce random errors in measurement outcomes. We model this effect as an effective depolarizing channel acting on Bob’s qubit: q G0 = 1 − 3d4det I, q G1 = ddet 4 X, q G2 = ddet 4 Y, q G3 = ddet 4 Z,

(S24) (S25) (S26) (S27)

where ddet is the depolarization probability representing the impact of dark counts and other detector-related errors, and {I, X, Y, Z} are Pauli matrices. For the two-qubit system, we assume detector imperfections manifest primarily on Bob’s side:   (B) EDetector (ρAB ) = I (A) ⊗ EDepolarizing (ddet ) (ρAB ), (S28) (Detector) (B) with Kraus operators Gm = I (A) ⊗ Gm .

where n(r) represents the radial refractive index profile of the fiber core, nc is the cladding refractive index, R is the core radius, and λ is the optical wavelength [4]. The number of propagating modes is roughly proportional to the refractive index contrast and (R/λ)2 . The introduction of MMF into a quantum transmission channel induces several physical effects that degrade the fidelity of the transmitted quantum states, including mode coupling, polarization mode dispersion (PMD), random phase fluctuations, and mode-dependent loss (MDL). Specifically, unintended mode coupling can arise from manufacturing imperfections (e.g., non-circular core geometry, rough core-cladding interfaces, or refractive index variations) and external mechanical perturbations (e.g., stress, micro-bending, or twisting). In our experiments, controlled twisting was applied to the MMF to induce perturbations, thereby enhancing inter-mode coupling. Such perturbations facilitate energy exchange between different spatial modes and lead to the random evolution of the transmitted optical field. Consequently, the initially prepared pure quantum states undergo decoherence during propagation, resulting in a reduction in measured fidelity.

2.

Theoretical Modeling Based on Field Coupling

The inter-mode coupling effects described above can be quantitatively characterized using the coupled-mode theory. In this framework, the total optical field propagating along the longitudinal z-direction is expanded as a superposition of orthogonal eigenmodes. The evolution of the modal amplitudes is governed by the following cou-

4 pled differential equation: dAµ = −jβµ Aµ + dz

X

Cµν (z)Aµ ,

(S30)

ν̸=µ

where Aµ denotes the complex amplitude of the µ-th mode, βµ is its propagation constant, and Cµν (z) represents the coupling coefficient induced by structural perturbations in the fiber [5]. The first term describes the independent phase evolution in the absence of coupling, while the second term accounts for energy transfer between modes. These coupling-induced effects ultimately manifest as crosstalk and decoherence in the transmitted quantum states. 3.

the correct sequence. The entangled state is distributed. Bob’s qubit then goes through the MMF, and finally, both Alice and Bob perform detection. The Kraus operators for Alice’s total noise channel are (k,i) (A) (A) KA = Fk Ei . The Kraus operators for Bob’s total (m,n,l,j) (B) (B) (B) (B) noise channel are KB = Gm Hn Fl Ej . Then, the noisy state is given by:  X  (k,i) (m,n,l,j) ρnoisy = K ⊗ K ρin AB AB A B k,i m,n,l,j

3. Composite Noise Model

In RSP, the shared entangled state first undergoes transmission, which includes loss, decoherence, and MMF noise on Bob’s side, followed by Alice’s measurement and Bob’s conditional unitary. We model the total noise channel as a sequential application of the individual noise channels. The composite Kraus operator for the total noise will be a product of the individual Kraus operators, applied in

(k,i)† (m,n,l,j)† KA ⊗ KB



.

+ + Here, ρin AB = |Φ ⟩ ⟨Φ | is the initial Bell state.

Discussion on Noise Suppression and Model Validity

Experimentally, in the absence of MMF (i.e., a direct quantum channel), the average fidelity of the prepared states reaches 99.25% ± 0.22%. As shown in the main text, the introduction of MMF causes a significant drop in fidelity when no noise suppression model is applied. This degradation reflects the raw impact of multi-mode perturbations prior to correction. The optimized TQSC (Transformer-based Quantum State Characterizer) noise model used in this work is trained on measurement data from the quantum channel incorporating the MMF. Consequently, the TQSC model inherently accounts for the effective noise introduced by the MMF, including contributions from mode coupling, phase perturbations, and mode-dependent loss. Combined with our Transformer-based post-processing reconstruction, the model recovers the experimental density matrices with high fidelity, aligning them closely with the target states. While the current framework achieves high reconstruction accuracy by treating the channel as a complex noise process, future work could benefit from incorporating more explicit physical constraints. Specifically, independently modeling the distinct noise contributions-such as the statistical characterization of coupling coefficients Cµν (z), phase noise spectra, and polarization effects-and integrating them into the noise prior or loss function could further refine the post-processing inversion.

(S31) 

4. Remote State Preparation under Noise

Alice performs a Bell-state measurement (BSM) on her (A) subsystem. For outcome k ′ , with projector Πk′ , the probability is h i (A) P (k ′ ) = Tr (Πk′ ⊗ I (B) )ρnoisy , (S32) AB and the post-measurement state of Bob’s qubit is i h (A) (A) TrA (Πk′ ⊗ I (B) )ρnoisy (Πk′ ⊗ I (B) ) ′ AB (k ) . (S33) ρB = P (k ′ ) Depending on Alice’s reported result k ′ , Bob applies a correction unitary Uk′ ∈ {I, X, Y, Z}, leading to the final ensemble state: X (k′ ) ρfinal = P (k ′ ) Uk′ ρB Uk†′ . (S34) B k′

5. Performance Metric: Fidelity

The performance of RSP is quantified by the fidelity between Bob’s final state and the target state |ψ⟩: F = ⟨ψ|ρfinal B |ψ⟩ .

(S35)

This allows us to systematically evaluate the impact of loss, decoherence, MMF induced noise, and detector imperfections on RSP. III.

THEORETICAL FRAMEWORK OF THE TQSC MODEL

A.

Problem Formulation: Learning an Inverse Quantum-to-Classical Map

The experimental realization of RSP is inevitably influenced by a multitude of physical imperfections. The entire process can be conceptualized as a complex quantumto-classical channel, denoted by C, which maps an ideal

5 target quantum state ρtarget to a set of classical measurement outcomes. In our work, each data record of these outcomes is tokenized into a 13-element feature vector M ∈ R13 . This vector acts as a comprehensive, noisecorrupted classical “snapshot” of the quantum state: M = C(ρtarget ).

(S36)

To intuitively understand this formulation, M aggregates the multidimensional profile of the state observed from 4 distinct perspectives. Specifically, it consists of direct physical observables: the coincidence counts between Alice and Bob, the corresponding single-photon counts for each party, all measured across four different measurement bases, alongside the total coincidence measurement time. This specific design ensures that the vector is informationally complete, encapsulating the full spectrum of data required for quantum state tomography. Through this numerical encoding, we map raw physical variables into a structured feature space, enabling our TQSC Transformer model to process quantum data with the same architectural logic used for word embeddings in natural language processing. The map C encapsulates both coherent errors, such as misalignments in optical components, and incoherent noise processes, including channel decoherence, detector dark counts, and Poissonian shot noise. Conventional quantum state tomography (QST) techniques attempt to reconstruct the state by inverting this map from the data M. However, they often falter when faced with complex, non-Gaussian, or correlated noise structures inherent in real-world systems. To overcome this limitation, we propose a data-driven TQSC model. The TQSC is a deep neural network, parameterized by a set of weights θ, that learns an effective inverse map ΦTQSC (·; θ). This learned function directly transforms the noisy classical measurement data M into a highfidelity estimate of the density matrix, ρpred : ρpred = ΦTQSC (M; θ).

(S37)

The core objective is to train the TQSC such that the learned map ΦTQSC not only inverts the deterministic dynamics of the channel but also actively suppresses the stochastic noise components.

B.

formation to the Transformer, we apply a Sinusoidal Positional Encoding Ppos ∈ RL×dmodel . This encoding is generated by fixed sine and cosine functions of different frequencies, where each position pos in the sequence is mapped to a unique vector. The initial representation H(0) ∈ RL×dmodel is formed as:

Model Architecture and Mathematical Formalism

The architecture of the TQSC is meticulously designed to leverage the powerful sequence-processing capabilities of the Transformer model [6], adapting it to the structured nature of our physics-based input vector. Input Embedding with Encoding The input vector M ∈ R13 is interpreted as a sequence of L = 13 tokens, S = {m1 , m2 , . . . , mL }. To provide positional in-

H(0) = Linear(S) + Ppos .

(S38)

This injects absolute positional information but does not incorporate explicit knowledge of physical feature types. Multi-Head Self-Attention Core The heart of the TQSC is a stack of N = 3 identical encoder layers. The central mechanism in each layer is Multi-Head SelfAttention (MHSA) with h = 4 heads, which allows the model to weigh the importance of all input tokens relative to each other. Each head independently learns different subspace features of the input, such as photon counting patterns, noise correlations, and temporal correlations in quantum signals. For an input representation X ≜ H(l−1) ∈ Rn×d at layer l (where n is the sequence length, d the feature dimension), the MHSA output is computed as: MHSA(X) = Concat(head1 , . . . , headh )WO ,

(S39)

where WO ∈ Rhdv ×d is the output projection matrix, and each head is defined by: headi = Attention(Qi , Ki , Vi ).

(S40)

Here the query, key, and value matrices for head i are computed as: Qi = XWiQ ,

WiQ ∈ Rd×dk

(S41)

Ki = XWiK , Vi = XWiV ,

WiK ∈ Rd×dk WiV ∈ Rd×dv

(S42) (S43)

with dk = dv = d/h. The attention function is then given by:  Attention(Qi , Ki , Vi ) = softmax

Qi K⊤ √ i dk

 Vi . (S44)

This attention mechanism enables the model to identify and exploit complex, non-linear correlations within the measurement data—for instance, the relationship between single-photon counts and coincidence events, which is a key indicator of channel loss and noise. The output of the attention block is then processed through a positionwise Feed-Forward Network (FFN), with residual connections and layer normalization applied after each sub-layer to ensure stable training. State Vector Prediction Head Following the final encoder layer, we obtain the output representation H(N ) ∈ RL×dmodel . To aggregate the information from the entire

6 sequence into a single vector, we utilize the fixed positional encoding scheme inherent in the Transformer architecture. Specifically, the sinusoidal positional encoding Ppos provides absolute position information for each element in the sequence. The aggregated representation hagg ∈ Rdmodel is computed as: L X

1 (N ) h hagg = L i=1 i

ρout =

(S46)

For the single-qubit case considered in this work, d = 2 and the network outputs four real parameters r = (r0 , r1 , r2 , r3 ). We define   s0 0 C= , r2 + ir3 s1 (S47) sj = softplus(rj ) + ϵ,

A brief non-technical explanation of Key and Query.

Keys and queries are fundamental components of the attention mechanism widely used in modern neural networks. Intuitively, a query represents the question or probe used to search for relevant information, while a key is a descriptor associated with each candidate item. The value stores the actual information to be retrieved. In simple terms, the query asks “what am I looking for?”, the keys label the available items, and the values contain the corresponding content. Items whose keys better match the query receive higher weights and thus contribute more strongly to the output representation. A useful analogy is searching for a book in a library. The query corresponds to what one is searching for, for example, “an introductory book on machine learning.” Each book in the library is associated with a key that summarizes its main characteristics, such as subject, difficulty level, or keywords. The value represents the actual content of the book. The attention mechanism compares the query with the keys of all books and assigns higher weights to those whose keys better match the query. As a result, books that are more relevant to the search request contribute more to the final output, while irrelevant books receive negligible weights. From this perspective, the attention mechanism can be viewed as a soft, learnable retrieval process, rather than a hard selection of a single item.

D.

CC † . Tr(CC † )

(S45)

where L is the sequence length. This mean-pooling strategy leverages the position-aware representations created by the positional encoding module.

C.

in maximum-likelihood quantum-state tomography. Instead of directly predicting the entries of ρ, the network 2 outputs an unconstrained real vector r ∈ Rd , which is used to construct a lower-triangular complex matrix C. The output density matrix is then reconstructed as

Physically-Constrained Output Layer

A critical challenge is ensuring that the neural-network output corresponds to a physically valid density matrix. A density matrix ρ must satisfy three constraints: (i) Hermiticity, ρ = ρ† , (ii) unit trace, Tr(ρ) = 1, and (iii) positive semidefiniteness, ρ ≥ 0. To enforce these constraints by construction, we adopt a Cholesky-style parametrization, as commonly used

j = 0, 1,

where ϵ > 0 is a small constant used for numerical stability. This choice ensures that Tr(CC † ) > 0 for every finite network output. This parametrization guarantees Hermiticity because (CC † )† = CC † .

(S48)

It also guarantees unit trace by construction: Tr(ρout ) =

Tr(CC † ) = 1. Tr(CC † )

(S49)

Finally, for any complex vector v, we have v † ρout v =

v † CC † v ∥C † v∥2 = ≥ 0. † Tr(CC ) Tr(CC † )

(S50)

Therefore, ρout is positive semidefinite for any real-valued network output r. For d = 2, the reconstructed density matrix can be written explicitly as ρout =

1 s20 + s21 + r22 + r32 ! s0 (r2 − ir3 ) s20

·

s0 (r2 + ir3 ) s21 + r22 + r32

(S51) .

Thus, the physically constrained output layer maps every unconstrained real-valued output of the Transformer backbone to a bona fide density matrix, while preserving end-to-end differentiability.

E.

Optimization and Theoretical Merit

The network parameters θ are optimized by minimizing the Mean Squared Error (MSE) loss function, defined as the squared Frobenius norm between the predicted state ρpred and the known ideal target state ρtarget . Over a

7 training dataset of K samples {(Mk , ρtarget,k )}K k=1 , the loss is: K

L(θ) =

1 X 2 ∥ΦTQSC (Mk ; θ) − ρtarget,k ∥F . K

(S52)

k=1

Optimization is performed using the Adam optimizer, which implements a variant of stochastic gradient descent. From a theoretical standpoint, the TQSC framework offers a significant advantage. The Transformer architecture is a universal function approximator, guaranteeing its capacity to model the highly complex inverse map ΦTQSC ≈ C −1 . More pointedly, the self-attention mechanism functions as a learned, context-aware adaptive filter. It can dynamically identify and down-weight measurement values that are inconsistent with the context provided by the rest of the data, effectively learning to suppress noise-induced anomalies. Therefore, the minimization of the MSE loss in Eq. (S52) does not merely fit a curve; it drives the model to learn the underlying statistical structure of the noise and systematically correct for it. This establishes a robust framework where deep learning directly enhances quantum state fidelity by learning noise-resilient inversion of physical measurements. The TQSC framework derives its theoretical strength from two fundamental properties of the Transformer architecture. First, as a universal function approximator, the Transformer possesses the inherent capacity to model arbitrarily complex inverse mappings with any desired precision. This mathematical guarantee ensures that TQSC can in principle learn the exact inverse transformation from measurement data to quantum states. Second, the self-attention mechanism provides a sophisticated context-aware processing capability. By dynamically evaluating the consistency of each measurement value against the full context of the dataset, it learns to identify and suppress noise-induced anomalies while preserving genuine quantum signatures. This adaptive filtering operates as a learned denoising function that transcends traditional signal processing approaches. Consequently, minimizing the mean squared error loss accomplishes more than simple data fitting. It drives the model to internalize the statistical patterns of measurement noise and implement physically meaningful corrections. This creates a principled framework where deep learning directly enhances the accuracy of quantum state reconstruction through intelligent noise-resilient inversion of experimental data. IV.

EXPERIMENTAL SETUP

The core of the system is a Sagnac interferometric loop housing a spontaneous parametric down-conversion (SPDC) source, designed for generating polarizationentangled photon pairs with inherent phase stability. A

femtosecond laser operating at a repetition rate of 60 MHz undergoes second-harmonic generation to produce 775 nm pump light. This pump beam is injected into the polarization-entangled Sagnac loop, where SPDC occurs in a periodically poled lithium niobate (PPLN) crystal, generating degenerate photon pairs at 1550 nm wavelength. The half-wave plate (HWP) and quarterwave plate (QWP) within the Sagnac loop were initially aligned with their fast axes at 0° relative to the horizontal polarization axis to ensure maximal generation of the target Bell state. Following generation, wavelength-division multiplexing (WDM) filters select photons within the approximately 1550 nm band. These photons are then routed through specific dense wavelength-division multiplexing (DWDM) channels: channel 31 (1552.54 nm) directs photons to Alice, and channel 33 (1550.92 nm) directs photons to Bob. In Alice’s detection module, the incoming photon is first split by a 9:1 beam splitter (BS1). One output arm incorporates a tunable optical delay line for precise temporal synchronization before recombination at the second beam splitter (BS2). The alternate output path from BS1 contains a variable optical attenuator used to adjust the intensity ratio ‘η = α2 /β 2 ’ between the two measurement bases. Both arms subsequently pass through fiber-based polarizing beam splitters (FPBS) for projection onto the horizontal (H) or vertical (V) polarization basis. For the critical state analysis step, a hybrid detection scheme is implemented by combining the H-polarized output from the upper FPBS and the V-polarized output from the lower FPBS through BS2. Polarization controllers (PC) in both arms compensate for polarization rotation induced by the optical fibers. Following BS2, projective measurements are performed using a 22.5° half-wave plate (HWP) and polarizing beam splitters (PBS) before coupling into fiber-coupled singlemode collection paths. Bob’s receiver module employs a polarization controller (PC) for initial compensation of fiber-induced birefringence. A commercially available OM2 gradedindex MMF (core diameter 50 µm, cladding diameter 125 µm, length 1.5 m) is used in the experiment. To simulate the spatial mode scrambling encountered in complex scattering environments, the photon is transmitted through this fiber section. Full quantum state tomography is implemented using a polarization analyzer consisting of a quarter-wave plate (QWP), a half-wave plate (HWP), and a polarizing beam splitter (PBS) in sequence, with the outputs coupled via single-mode fibers to detectors. Both Alice and Bob employ superconducting nanowire single-photon detectors (SNSPDs) for photon detection, characterized by a quantum efficiency of approximately 80% at the operating wavelength of 1550 nm. Detection events are recorded and time-tagged using time-

8 correlated single-photon counting (TCSPC) electronics offering a timing resolution of 81 ps. The coincidence window was set to 500 ps, optimized relative to the 16.7 ns repetition period of the femtosecond laser source.

V.

CHARACTERIZATION OF THE SPDC SOURCE

To systematically quantify the performance of the Spontaneous Parametric Down-Conversion (SPDC) source, we provide a comprehensive experimental characterization covering its efficiency and other key experimental parameters. The TQSC model, trained on experimental data, implicitly captures SPDC failure events such as vacuum and multi-pair contributions, learning to suppress noise and restore fidelity without requiring an analytic noise model. The characterization results and analysis are organized as follows: Firstly, the spectral properties and normalized conversion efficiency of the PPLN waveguide are investigated to clarify the phase-matching bandwidth and the generation efficiency of the source. Secondly, the signal-to-noise performance is evaluated by measuring the coincidenceto-accidental ratio (CAR) as a function of pump power. Thirdly, the single-photon purity and suppression of multi-pair emissions are verified using the second-order correlation function g (2) (0). Fourthly, the high quality of the generated states is demonstrated through polarization entanglement characterization. Finally, regarding SPDC failure events, the TQSC model implicitly learns to suppress the resulting noise and restore fidelity, without the need for an explicit analytic model of the imperfections. Detailed discussions and measurement results for each of these parameters are presented below.

A.

Spectral Properties and Normalized Conversion Efficiency

To characterize the phase-matching bandwidth and generation efficiency of the PPLN waveguide, we investigated the spectral distribution of the SPDC photons and compared it with the theoretical model. Fig. S1 presents the measured normalized photon counts for 10 symmetric signal-idler channel pairs centered at ITU Channel 32 (CH32), ranging from the outermost pair (CH22 & CH42) to the innermost pair (CH31 & CH33). In the experiment, the pump wavelength was tuned to generate degenerate photon pairs centered at CH32. The experimental data (colored bars) show excellent agreement with the theoretical phase-matching curve (red solid line), confirming the reliability of the waveguide design. Based on the measured spectral distribution, we calculated a Full-Width at Half-Maximum

(FWHM) bandwidth of 60.74 nm. This broad phasematching bandwidth ensures a relatively flat and high normalized conversion efficiency across a wide range of channels. This characteristic verifies the uniformity of photon generation and confirms that the source is wellsuited for multi-channel quantum communication protocols requiring high spectral purity and broad wavelength tunability.

B. Source Noise Performance and Coincidence-to-Accidental Ratio (CAR)

To evaluate the signal-to-noise ratio performance of the source, we measured the Coincidence-to-Accidental Ratio (CAR) as a function of pump power for the 10 symmetric channel pairs centered at CH32 (spanning from CH22/CH42 to CH31/CH33), as summarized in Fig. S16. As shown, the CAR values for all channel pairs exhibit a consistent dependence on pump power, strictly following the theoretical inverse relationship CAR ∝ 1/Ppump . This behavior indicates that accidental coincidence counts are dominated by multi-pair emissions rather than background noise or dark counts, confirming the low-noise nature of the experimental setup. Quantitatively, at low pump powers (in the µW regime), the maximum CAR values for all measured channel pairs exceed 20,000, with the central channel pair (CH30 and CH34) reaching a peak of 30,446. Even at higher pump powers, where the statistical significance of multi-photon events increases, the fitted curves maintain excellent agreement with the experimental data across the entire spectral range. These results demonstrate that the device maintains high entanglement purity and a high signal-to-noise ratio over a broad bandwidth, validating the robust performance of the source for multi-channel quantum communication applications.

C.

Single-Photon Purity and Second-Order Correlation Function g (2) (0)

To verify the single-photon purity and the suppression of multi-pair emissions, we measured the zero-delay second-order correlation function, g (2) (0). The results for the 10 symmetric channel pairs (from CH22/CH42 to CH31/CH33) are summarized in Fig. S2. To ensure statistical reliability, each value represents the mean of three independent repeated measurements, with error bars indicating the standard deviation. As shown in Fig. S2, all channel pairs exhibit consistently low g (2) (0) values, ranging from approximately 0.012 to 0.019. These values are well below the classical limit (g (2) (0) ≥ 1) and significantly lower than the typical benchmark for high-quality single-photon sources (g (2) (0) < 0.1), clearly demonstrating strong

1RUPDOL]HG3KRWRQ1XPEHU'LVWULEXWLRQ

9

7KHRUHWLFDO0RGHO

      







:DYHOHQJWK QP





Figure S1: Characterization of SPDC spectral properties. The red solid line represents the theoretical normalized conversion efficiency (phase-matching curve). The colored bars indicate the measured normalized photon counts for 10 symmetric channel pairs centered at CH32 (from CH22 & CH42 to CH31 & CH33). The measured Full-Width at Half-Maximum (FWHM) bandwidth is 60.74 nm.

non-classical statistical properties and effective suppression of multi-channel noise. Furthermore, the small standard deviations and the narrow distribution of means across different channel pairs confirm the uniformity and stability of the multi-channel system, indicating that high-quality quantum properties are achieved simultaneously across the entire spectral range without significant channel-dependent systematic errors.

D.

Polarization Entanglement Characterization

To demonstrate the high quality of the generated states, we performed polarization entanglement characterization across all 10 symmetric channel pairs. Fig. S17 presents the polarization-dependent fringe measurements. To ensure experimental rigor, entanglement fringes were measured in two non-orthogonal polarization bases (0◦ and 22.5◦ ). Each data point represents the average of three independent trials. The entanglement fringes exhibit robust and consistent sinusoidal behavior across all channels, in excellent agreement with theoretical predictions. Quantitatively, we achieved an average visibility of 98.74%±0.69% across all channels and settings. Specifically, the average visibility was 98.91%±0.37% in the 0◦ basis and 98.56%±0.89% in the 22.5◦ basis. The close agreement between these two bases, combined with low statistical uncertainty, confirms that the SPDC source maintains high-fidelity entanglement and uniform performance across all channel pairs with no observable systematic bias.

E.

Accounting for SPDC Failure Events in the Noise Model

In a coincidence-based post-selection scheme, coincidence measurements can filter out the most fundamental photon losses and the vast majority of |0, 0⟩ vacuum noise. However, the SPDC source still faces inherent failure issues. Specifically, high-order multi-photon emissions intrinsic to the SPDC Hamiltonian, such as |2, 2⟩ states, can masquerade as valid single-photon pairs when coupled with detectors lacking photon-number resolution and partial channel losses. Furthermore, detector dark counts within the coincidence window also lead to accidental coincidences. Mathematically, these phenomena manifest as a finite CAR, and a non-zero g (2) (0), which is consistent with the characterization results discussed earlier. Instead of relying on idealized theoretical assumptions, our model learns from the actual photon statistics where these imperfections are naturally present. In the resulting coincidence data, these physical failures manifest as specific noise patterns. For instance, reduced total counts from a finite heralding efficiency introduce Poissonian shot noise, while accidental coincidences mix the pure state with classical white noise to create depolarizing noise. The source imperfections and intrinsic noise effects of the SPDC process are naturally reflected in the measurement data. By using these noisy measurements as the input to train the TQSC model, we ensure that the model’s predictions faithfully reflect the performance of the physical hardware. The core advantage of our Transformer model is that it

10

$YHUDJHg(2)(0)DFURVVFKDQQHOV 0HDQg(2)(0)

g(2)(0)

   



 &+

&

+



 &+

&

+



 &+

&

+



 &+

&

+



 &+

&

+



 &+

&

+

 &+

&

+



 &+

&

+

 +

 &+

&

 &+

&

+





Figure S2: Measurement of the second-order correlation function g (2) (0) for 10 symmetric channel pairs. The bar chart shows the average g (2) (0) values from the outer pairs (CH22 & CH42) to the inner pairs (CH31 & CH33). Error bars represent the standard deviation calculated from three independent measurements.

does not require an explicit analytical inversion of these physical failures. Instead, through the self-attention mechanism trained on large datasets generated by this noise model, the Transformer automatically learns the complex mapping from the noise-corrupted projection data back to the high-fidelity, physically valid (positive semi-definite) density matrix ρideal . It inherently learns to suppress the statistical noise caused by SPDC low efficiency and correct the fidelity drop caused by accidental coincidences.

VI.

QUANTUM STATES PREPARED AND MEASUREMENT STATISTICS

First, it is clarified that the reported five repetitions do not correspond to five single-shot projective measurements. Instead, they refer to five independent experimental batches. In each batch, coincidence counts and the corresponding single-photon counts are accumulated across all projection bases within a specific time. Given the source brightness, approximately 102 coincidence counts per second are collected, while singlephoton counts per user range from 102 to 103 depending on the basis. Conducting multiple independent batches allows for evaluating the long-term stability of the setup and for averaging out random fluctuations or system drifts. Second, to bound the margin of error within E = 0.05 at a 95% confidence level (Z ≈ 1.96), the minimum required sample size is given by n = (Zσ/E)2 .

(S53)

Taking the mixed state η = 1 as an example, preliminary tests (n = 3) indicated that five independent replicates would be sufficient. Subsequent formal measurements at n = 5 yielded σ ≈ 0.0539, confirming this design. As validated in Fig. S3(a) and Fig. S3(b), both the 95% confidence interval and the margin of error converge securely within the predefined tolerance at this exact sample size. The observed fidelity decrease reflects the physical impact of the MMF, rather than an omission in the noise model. Experimentally, in the absence of the MMF (i.e., when the quantum channel is directly connected without multimode transmission), the average fidelity of the seven prepared quantum states reaches (99.25 ± 0.22)%. As illustrated in Fig. S4, the baseline fidelity of the calibrated setup without the MMF (blue bars) is slightly below unity due to inherent experimental imperfections, such as Poissonian statistical noise, detector dark counts, and imperfect optical components. Upon introducing the MMF, the uncorrected fidelity (red bars) decreases to an average of (87.04 ± 3.77)% due to multimode perturbations. The optimized TQSC model is trained directly on measurement data from the quantum channel with the MMF already incorporated. Consequently, it inherently accounts for both inherent experimental imperfections and the MMF-induced effective noise (e.g., mode coupling, phase perturbations, and mode-dependent loss). When paired with the Transformer-based post-processing, this framework reliably recovers the experimentally measured density matrices with high fidelity toward the target states.

11 (a)

(b)

Figure S3: Dependence of measurement precision on replicate number n. (a) Measured values with 95% confidence intervals. (b) Margin of error.

Without MMF

With MMF

1.0

Fidelity

0.8 0.6 0.4 0.2 0.0 ρ1

ρ2

ρ3

ρ4

ρ5

ρ6

ρ7

Figure S4: State fidelities without and with the MMF.

A.

Measurement statistics and BER evaluation in the MNIST demonstration

In the MNIST transmission demonstration, the two logical bit values are encoded into the mixed states ρη =

1 (|H⟩⟨H| + η|V ⟩⟨V |) , 1+η

(S54)

with η = 2 and η = 3. Using the squared Uhlmann fidelity convention, their fidelity is F (ρ2 , ρ3 ) = 0.9916.

(S55)

Thus, the two ideal encoded states are intrinsically close to each other. For equal prior probabilities, the Helstrom minimum error probability for direct single-copy discrimination is   1 1 Hel Pe = 1 − ∥ρ2 − ρ3 ∥1 = 45.83%, (S56) 2 2

where ∥ρ2 − ρ3 ∥1 = 1/6. This value quantifies the minimum possible error probability for directly distinguishing one ideal copy of ρ2 from one ideal copy of ρ3 using the optimal quantum measurement. It therefore reflects the intrinsic closeness of the two encoded ideal states, but it is not the operational BER bound for the experimental MNIST transmission task, where each sample is represented by a finite-time acquisition record. We emphasize that the N = 100 measurements shown in Fig. 4(b) are used only to characterize and visualize the marginal projection-outcome distributions of the two encoded states. These measurements are not concatenated, averaged, or majority-voted for the bit-error-rate evaluation. In the bit error rate (BER) evaluation, each transmitted bit is represented by one independent experimental measurement record, x = (C1 , . . . , C4 , S1 , . . . , S8 , τ ) ∈ R13 ,

(S57)

where C1 , . . . , C4 and S1 , . . . , S8 denote four coincidence and eight single-photon counts, respectively, and τ is the timing feature. The TQSC model reconstructs the output state and assigns the corresponding label from this single 13-dimensional experimental record. Therefore, the BER result is not obtained by statistically averaging the N = 100 repeated measurements shown in Fig. 4(b) of the main text, but by evaluating independent experimental records after MMF transmission. For comparison, the conventional maximum-likelihood reconstruction baseline uses the four coincidence-count measurements. For the ith experimental record, we denote the state reconstructed by method m ∈ {ML, TQSC} as ρ̂m,i , and evaluate its fidelity to the target state ρη as F (ρ̂m,i , ρη ). Averaging over 700 experi-

12 mental realizations for each target state gives (η=3)

= 0.8732 ± 0.0360,

(S58)

(η=3) F̄TQSC = 0.9625 ± 0.0020,

(S59)

F̄ML

(η=2)

= 0.9353 ± 0.0930,

(S60)

(η=2) F̄TQSC = 0.9843 ± 0.0005.

(S61)

F̄ML

Here, the values after ± denote the standard deviations over different experimental realizations. The decoded label is assigned according to the reconstructed state’s fidelity to the two target states, η̂ = arg max F (ρ̂, ρη ).

(S62)

η∈{2,3}

On the 1400 test records, including 700 realizations of ρ2 and 700 realizations of ρ3 , the TQSC reconstruction gives a BER of 0%, whereas the conventional ML reconstruction gives a BER of 49.57%. The experiment evaluates the robustness of different reconstruction pipelines when applied to noisy experimental records after multimode-fiber transmission and detection. The performance improvement of TQSC originates from its use of the complete 13-dimensional experimental measurement record, rather than from repeatedmeasurement averaging of the N = 100 statistics shown in Fig. 4(b).

VII.

DETAILED TQSC ARCHITECTURE AND TRAINING

The TQSC processes a 13-dimensional input vector composed of three key components for each measured mixed state: the coincidence counts of relevant projection bases, the corresponding photon counts recorded at Alice’s and Bob’s measurement stations, and the measurement coincidence time. The target output for the fidelity optimization task is the density matrix representing the quantum state. Unlike conventional quantum state tomography approaches that rely on repeated measurements to estimate expectation values, the proposed TQSC model infers the quantum state from a single set of measurement outcomes. Specifically, for each quantum state, only one measurement shot per measurement setting is performed, and no averaging over multiple measurement rounds is required. The Transformer model learns to reconstruct the underlying quantum state directly from this single-shot measurement data. Each element in the input vector is treated as a token. The input tensor first undergoes preprocessing where it passes through a linear transformation layer that projects it into a higher-dimensional feature space. This is followed by an unsqueeze operation to adjust the tensor

dimensions for compatibility with the subsequent Transformer architecture. The preprocessed tensor is then fed into a stack of three identical Transformer encoder layers. Each encoder layer begins by incorporating positional encoding to preserve the sequence order information. The core component is a multi-head self-attention mechanism that operates with four parallel attention heads. Each head learns distinct types of correlations within different representation subspaces, significantly enhancing the model’s capacity to capture complex dependencies between the input tokens. The outputs from the attention heads are integrated and passed through a residual connection followed by layer normalization (Add & Norm block), which stabilizes training and facilitates convergence. The normalized output is then processed by a position-wise feedforward neural network with nonlinear activation, enabling deeper extraction of complex features from the attention outputs. After processing through the three encoder layers, the enriched feature representation enters the postprocessing stage. The feature dimension is first adjusted through a final linear transformation. The resulting tensor is then decomposed into two components: one corresponding to the diagonal elements of the density matrix and the other to the off-diagonal elements. The diagonal elements undergo softmax normalization to ensure they sum to unity and represent valid probabilities. The off-diagonal elements are processed through a hyperbolic tangent (tanh) activation function, constraining their values to the range [-1, 1] to reflect the properties of coherence terms. Finally, the processed diagonal and off-diagonal components are recombined appropriately to reconstruct the full Hermitian density matrix, which constitutes the final output of the model. For model development and evaluation, the experimental dataset was partitioned using stratified random sampling, allocating 80% of the data to the training set and 20% to the independent test set. This approach preserves the underlying distribution of quantum states in both subsets. During training, the model was optimized using the Adam optimizer, an adaptive variant of stochastic gradient descent. The MSE loss function was employed to quantify the discrepancy between the predicted density matrix (ρpred ) and the true label density matrix (ρtrue ): N

LMSE =

1 X (i) (i) ∥ρ − ρtrue ∥2F , N i=1 pred

(S63)

where N is the batch size and ∥ · ∥F denotes the Frobenius norm. The initial learning rate was set to 0.001 and training proceeded with a batch size of 64. Over successive training epochs, the loss value steadily decreased as the predicted density matrices converged toward their ground-truth counterparts.

13 VIII.

BASELINE COMPARISON

To evaluate reconstruction performance under the same experimental protocol, we compared the Transformer–Cholesky model with representative neural, tomographic, Bayesian, and maximum-likelihood baselines. All learning-based models used the same experimental input vectors, normalization protocol, train/validation/test split, and fidelity metric. The neural baselines were also constrained to output physical density matrices through the Cholesky parametrization, enabling a direct comparison between attention-based and non-attention learning-based reconstructions. The conventional reconstruction baselines include HVDR linear inversion, HVDR least-squares tomography, and HVDR maximum-likelihood estimation (MLE), which are standard density-matrix reconstruction approaches in quantum-state tomography. For the HVDRMLE baseline, the maximum-likelihood procedure was applied to the experimental remote-state-preparation projective measurements obtained under dynamic Gaussian noise superimposed with MMF noise. For each target state, five repeated projective measurements were performed, the density matrix was reconstructed for each repetition, and the reported fidelity is the average fidelity between the reconstructed density matrices and the corresponding target states. The baseline set covers several complementary comparisons: (i) a parameter-matched non-attention MLP with the same Cholesky physicality constraint as the Transformer; (ii) a smaller Cholesky-constrained MLP; (iii) closed-form HVDR linear-inversion and least-squares tomography from the measured HVDR coincidence counts, followed by positive-semidefinite (PSD) projection; (iv) a Bayesian known-family estimator over the experimentally used discrete mixed-state family; and (v) an empirical training-calibrated HVDR-MLE estimator. Randomized benchmarking and randomized-compiling-based noise tailoring are gate-sequence characterization and mitigation protocols [7, 8]; they are therefore outside the direct scope of the present projection-measurement data set. The quantitative results are summarized in Table S1. While the results presented in the main text are obtained using a fixed random seed for consistency, here we evaluate five different random seeds to assess the robustness of the model architectures against variations in random initialization. The Transformer–Cholesky model achieved the highest average test fidelity among the compared methods, reaching 0.999982 ± 0.000046. The parametermatched MLP–Cholesky baseline also achieved high fidelity, reaching 0.999978 ± 0.000166 with a comparable number of trainable parameters. This comparison shows that physically constrained learning-based reconstruction is highly effective for the present dynamic Gaussian plus

MMF noise condition, and that the Transformer formulation retains a slight average advantage under an intentionally strong non-attention neural baseline. The smaller MLP reached 0.999933 ± 0.000429, further confirming the robustness of the Cholesky-constrained neural reconstruction strategy. The Bayesian known-family estimator reached 0.996200 ± 0.007437, confirming that the discrete state-family information provides a useful prior for this task. The Transformer–Cholesky and MLP–Cholesky models further improve the average fidelity by learning directly from the experimental count vectors and auxiliary channels. The HVDR-LI-PSD, HVDR-LS-PSD, and training-calibrated HVDR-MLE baselines provide conventional reconstruction references under the same dynamic-noise data. Overall, the comparison supports the use of physically constrained neural reconstruction for high-fidelity single-qubit state estimation in this experiment. Fig. S5 visualizes the baseline comparison from three complementary perspectives. Fig. S5(a) compares the test infidelity, 1 − F , of different reconstruction methods on a logarithmic scale, showing that the Transformer– Cholesky model attains the lowest mean infidelity among the evaluated methods. Fig. S5(b) presents the per-sample infidelity-tail distribution, where the Transformer–Cholesky results remain concentrated in the high-fidelity regime across the test set. Fig. S5(c) shows the sorted paired fidelity differences relative to the Transformer–Cholesky model; positive values correspond to test samples for which the Transformer–Cholesky reconstruction gives higher fidelity than the corresponding baseline, and the dashed horizontal line denotes parity with the Transformer. The Transformer formulation is well matched to the structure of the present tomographic input. The HVDR counts and auxiliary channels are basis-specific observations of the same density matrix under a shared noise environment. Tokenizing these components allows selfattention to model inter-channel interactions and provides a natural route to larger measurement sets, additional monitoring channels, time-dependent records, and potentially multi-qubit tomography [6, 10, 11]. The attention maps are used as post-hoc diagnostic tools to inspect how the model weights different measurement channels [12, 13].

IX.

GENERALIZATION CAPABILITY OF THE TQSC MODEL A.

Experimental generalization within the measured state family

To further evaluate the generalization capability of the proposed TQSC model beyond the initially selected eight

14 Table S1: Quantitative comparison of the strengthened baselines. For neural models, the mean is averaged over five random seeds. The reported standard deviation is the mean per-sample fidelity standard deviation on the test set, not the seed-to-seed standard deviation. The parameter column reports trainable parameters; ‘–’ denotes non-neural baselines without a comparable trainable parameter count. Method

Parameters Test Fidelity (Mean ± Std)

Purpose

Transformer-Cholesky Proposed model MLP-Cholesky-ParamMatched Non-attention MLP, comparable size MLP-Cholesky-Small Small MLP Bayesian-Known-Family Oracle-like known-family posterior Training-Calibrated HVDR-MLE Noise-aware reconstruction HVDR-LS-PSD Least-squares tomography HVDR-LI-PSD Linear-inversion tomography Legacy HVDR-MLE Ref. [9] Legacy MLE reference

(a)

(b)

Fraction with infidelity >= x

MLP small Bayesian TC HVDR-MLE HVDR-LS

100

1.0

seed mean

MLP matched

(c)

10−1 0.8

10−2 10−3

F = 0.9999

0.6

Delta F

F = 0.9999

Transformer

0.999982 ± 0.000046 0.999978 ± 0.000166 0.999933 ± 0.000429 0.996200 ± 0.007437 0.871839 ± 0.235055 0.834470 ± 0.179945 0.830808 ± 0.193338 0.383100 ± 0.123000

152,292 152,306 7,268 0 – – – –

0.4

10−4 10−5 0

0.2

−10−5 −10−4

HVDR-LI

0.0

10−6 10−5 10−4 10−3 10−2 10−1

Test infidelity (1 - F) Transformer

MLP matched

10−6

10−4

10−2

100

Per-sample infidelity (1 - F) MLP small

Bayesian

TC HVDR-MLE

0

100

200

300

Test samples sorted by Delta F HVDR-LS

HVDR-LI

Figure S5: Comparison with strengthened baselines. (a) Test infidelity 1 − F on a logarithmic scale; circles denote individual seeds and diamonds denote means. (b) Per-sample infidelity 1−F tail distribution. (c) Sorted paired fidelity differences relative to the Transformer; the dashed horizontal line marks parity. Sorted paired fidelity differences are computed by subtracting the Transformer fidelity from each baseline fidelity for the same input state.

training states, we conducted additional experiments on previously unseen quantum states. Interpolation Test at an Unseen Intermediate Parameter (η = 4.5). A model trained on the original dataset with parameter values η ∈ {1, 2, 3, 4, 5, 6, 7} was tested on quantum states generated using an intermediate value η = 4.5, which was not included during training. The model successfully reconstructed the corresponding quantum states with an average fidelity of 0.993712 (standard deviation: 4.28×10−6 ), demonstrating its ability to interpolate within the learned parameter manifold. Leave-One-Out Validation (η = 6). To further examine predictive capability, we retrained the model on a reduced dataset that explicitly excluded η = 6, i.e., η ∈ {1, 2, 3, 4, 5, 7}. The retrained model was then eval-

uated on quantum states corresponding to η = 6. The resulting reconstruction achieved an average fidelity of 0.999989 (standard deviation: 0.32 × 10−6 ), indicating strong predictive performance on unseen data. Discussion. As shown in Fig. S6, these results confirm that the TQSC model exhibits strong generalization capability to quantum states not explicitly present in the training set. This behavior can be attributed to the regression-based formulation of TQSC, which learns a continuous mapping between measurement data and quantum state parameters rather than memorizing discrete categories. Nevertheless, the predictive performance is inherently dependent on the statistical proximity between unseen data and the training distribution. The model performs

15

0HDQ)LGHOLW\

   

7UDLQLQJ6WDWHV 8QVHHQ6WDWH /HDYH2QH2XW

 





  

6WDWH3DUDPHWHU





Figure S6: Mean fidelity vs η. Circles: training. Square: leave-one-out. Diamond: unseen. Error bars: ±1σ.

best in interpolation regimes (e.g., η = 4.5) or when the test states share underlying physical characteristics with the training data. As is typical in regression-based learning tasks, prediction accuracy may degrade for states that deviate significantly from the training manifold [14, 15]. Therefore, enlarging and diversifying the training dataset can improve the sampling density of the state space, leading to enhanced predictive accuracy and robustness against distributional shifts [16, 17]. B.

Numerical generalization over Bloch-ball states

The experimentally acquired photonic data used in the main TQSC demonstration do not constitute an exhaustive sampling of the full single-qubit Bloch ball. In particular, the mixed states in Eq. (2) of the main text form a one-parameter family that is diagonal in the {|H⟩, |V ⟩} basis, and the experimentally prepared pure-state example in Eq. (8) represents one selected target state. Therefore, the experimental leave-one-out and interpolation tests should be interpreted as proof-of-principle tests within this restricted measured state family, rather than as a demonstration over arbitrary single-qubit states. To examine the behavior of the same TQSC pipeline beyond this diagonal family, we performed an additional Qiskit-based numerical benchmark using pure and mixed single-qubit target states distributed on and inside the Bloch ball. The target states are parameterized as ρ(r, θ, ϕ) =

1 [I + r n̂(θ, ϕ) · σ] , 2

0 ≤ r ≤ 1,

(S64)

where n̂(θ, ϕ) = (sin θ cos ϕ, sin θ sin ϕ, cos θ).

(S65)

Here, r = 1 corresponds to pure states on the Blochsphere surface, while 0 ≤ r < 1 corresponds to mixed

Figure S7: Qiskit-generated Bloch-ball target states used in the supplementary numerical benchmark. Pure states are sampled on the Bloch-sphere surface, while mixed states are distributed inside the Bloch ball. The dataset contains 480 pure states and 5041 mixed states, covering both diagonal and off-diagonal single-qubit density matrices.

states inside the Bloch ball. This construction includes both diagonal and off-diagonal density matrices in the {|H⟩, |V ⟩} basis. Fig. S7 shows the state-space coverage used in this supplementary benchmark. We used 480 pure-state directions on the Bloch-sphere surface and 360 angular directions for each of 14 nonzero mixed-state radial shells, together with the maximally mixed state at the origin. Therefore, the total number of target states before repeated noisy measurement sampling is Nstate = 480 + 14 × 360 + 1 = 5521.

(S66)

The noisy measurement samples were generated using an MMF-inspired Bob-channel model. This model is not intended to serve as a device-specific calibration of a particular fiber. Instead, it defines a physically motivated noisy-channel benchmark including transmission loss, polarization-dependent loss, polarization rotation, depolarization, phase damping, amplitude damping, temporal jitter, detector dark counts, and photoncounting shot noise. The parameters used in this simulation are summarized in Table S2. To avoid overlap between training and testing target states, we evaluated four strict train–test splits: leavestate-out, off-diagonal mixed-state holdout, radial-shell holdout, and Bloch-octant holdout. These splits are illustrated in Fig. S8, where the training and held-out states are separated by target-state identity. The corresponding held-out fidelities are shown in

16

Figure S8: Train and holdout splits in the Bloch ball. Gray points denote training states and red points denote held-out test states. The target-state-ID overlap between the training and held-out test sets is zero for all splits.

Fig. S9. Across the four strict generalization splits, the mean fidelities are 0.998911, 0.997943, 0.999145, and 0.998286, respectively. In total, these tests include 4831 held-out target states and 483100 repeated noisy measurement samples.

These results show that, in the supplementary numerical benchmark, the TQSC model can generalize beyond the original one-parameter diagonal mixed-state family. We emphasize, however, that this benchmark is numerical and should not be interpreted as a replacement for experimentally measured Bloch-ball data. The experimental dataset in the main text remains a cost-limited proof-of-principle demonstration under realistic optical measurement noise.

X. ATTENTION-ASSISTED POST-HOC ATTRIBUTION AND FUNCTIONAL ANALYSIS OF L3H4

This section provides additional analyses of the attention structure learned by the Transformer, with particular emphasis on Head 4 in Layer 3 (L3H4). A self-attention map characterizes how each token representation weights and integrates information from the other measurement tokens. Whereas earlier layers operate more directly on input-level measurement statistics, later layers act on progressively contextualized representations formed through repeated attention and nonlinear transformations. We therefore examine deeperlayer attention because it provides a natural setting in which task-relevant relations among complementary ob-

17 Table S2: Parameters used in the MMF-inspired noisy Bob-channel simulation. Parameter

Value

Role in the simulation 5

−1

Pair-generation rate 2.5 × 10 s Sets the photon-pair flux before detection and post-selection. Acquisition time per basis 0.25 s Integration time used to generate finite HVDR coincidence counts. Coincidence window 5.0 ns Defines the temporal acceptance for coincidence counting. Alice detection efficiency 0.55 Models finite Alice-side photon detection efficiency. Bob detection efficiency 0.50 Models finite Bob-side photon detection efficiency. RSP success probability 0.50 Accounts for successful remote-state-preparation events. Bob dark-count rate 80 s−1 Adds detector dark photons and accidental coincidences. Baseline Bob-channel transmission 0.88 Baseline transmission before MMF attenuation and timing acceptance. Depolarizing probability 0.06 Mixes the transmitted state toward I/2. Amplitude-damping probability γ = 0.04 Models polarization-dependent amplitude relaxation. Phase-damping parameter λ = 0.03 Models loss of phase coherence. MMF length 2.0 m Propagation length of the multimode-fiber segment. MMF attenuation 0.015 dB m−1 Propagation loss in the MMF channel. Polarization-dependent loss 0.25 dB Models unequal attenuation of orthogonal polarization components. MMF rotation angle 0.18 rad Birefringence-like polarization rotation angle. MMF rotation azimuth 0.70 rad Azimuth of the polarization-rotation axis. MMF modal depolarization 0.04 Additional modal depolarizing contribution. MMF modal dephasing 0.025 Additional modal coherence damping. MMF temporal jitter 1.0 ns Temporal broadening that reduces coincidence-window acceptance.

Mean

Median

1.000

Fidelity

0.999 0.998 0.997 0.996 0.995 Unseen states

Off-diagonal mixed

Radial shell

Bloch octant

Figure S9: Generalization fidelity on held-out Bloch-ball states. Bars show the mean fidelity, error bars denote one standard deviation, and diamonds indicate the median fidelity.

The input sequence consists of 13 tokens whose ordering is fixed by the measurement design. These tokens comprise four coincidence-count channels, Aliceand Bob-side single-count channels in the H/V/D/R bases, and acquisition time. The individual tokens and composite groups used throughout this analysis are summarized in Table S3 and Table S4. See Sec. X F for further details. Because the token order is identical for all samples, the relevant question is not whether attention changes with token ordering, but whether different physical input states produce systematically different attention structures under the same attention head.

A.

State Dependence of Sample-Specific L3H4 Attention

For each test sample i, the sample-specific L3H4 attention matrix is denoted by servables may emerge, including the joint use of H/V population-sensitive statistics, D/R coherence/phasesensitive statistics, Alice–Bob coincidence information, and acquisition-time context. This does not imply that deeper layers necessarily encode “deeper physical laws”; rather, they allow us to test whether lower-level measurement statistics have been integrated into more structured relational representations. The purpose of these analyses is not to treat an averaged attention map as direct evidence of model interpretability. Instead, attention is used as a post-hoc indicator of information routing, whose structure is further examined through state-dependence analysis, pathwayspecific intervention, whole-model perturbation, Integrated Gradients (IG), and relational analyses.

Ai ∈ R13×13 .

(S67)

State dependence was analyzed using a matched acquisition-condition subset containing N = 229 samples from five ground-truth state groups. Fig. S10(a,b) shows representative raw L3H4 attention maps for a pure-state sample (sample 143) and an η = 4 sample (sample 77). These examples were not selected randomly or chosen to maximize their visual difference. Within each state group, the representative was defined as the sample having the minimum Frobenius distance from the corresponding group-mean attention map. The two representatives exhibit different query–key structures under the same attention head. These representative

18

0.15

0.10

0.05

0.00

Pure medoid (143)

0.05

0.00

η = 4 medoid (77)

0.2 0.1 0.0 −0.1 −0.2

Query token

(d) Mean-centered L3H4 Mean-centered weight

Query token

0.10

Key token

(c) Mean-centered L3H4

0 1 2 3 4 5 6 7 8 9 10 11 12

0 1 2 3 4 5 6 7 8 9 101112

0.2 0.1 0.0 −0.1 −0.2 0 1 2 3 4 5 6 7 8 9 101112

Key token

Key token

(e)

State-explained variance

(f)

Residual-distance separation

0.14

Null zoom

0.14

Null zoom

0.12 0.10 0.08

0.02

0.04

0.06

0.06

Obs. = 0.674 Perm. p = 1.00 × 10−4

0.04 0.02 0.00 0.0

Permutation fraction

Permutation fraction

0.15

0 1 2 3 4 5 6 7 8 9 101112

Key token

0 1 2 3 4 5 6 7 8 9 10 11 12

0.20

0 1 2 3 4 5 6 7 8 9 10 11 12

Mean-centered weight

0 1 2 3 4 5 6 7 8 9 101112

η = 4 medoid (77) Raw L3H4

Attention weight

0.20

Query token

Query token

0 1 2 3 4 5 6 7 8 9 10 11 12

(b)

Attention weight

Pure medoid (143) Raw L3H4

(a)

0.12 0.10 0.000

0.08

0.025

0.050

Obs. = 0.965 Perm. p = 1.00 × 10−4

0.06 0.04 0.02

0.2

0.4

0.6

2 State-explained variance, Rstate

0.00

0.0

0.2

0.4

0.6

0.8

1.0

ΔD = Dbetween − Dwithin

Figure S10: State-dependent L3H4 attention. (a,b) Raw attention maps for the pure-state medoid (sample 143) and the η = 4 medoid (sample 77). (c,d) Corresponding mean-centered maps after subtraction of the test-subset mean. (e) State2 explained variance Rstate . (f) Between-state minus within-state residual distance ∆D. Null distributions were obtained from 104 state-label permutations; the insets show enlarged views of the corresponding null distributions.

19 samples illustrate the variation in L3H4 attention patterns across physical states. The statistical association between attention structure and state identity is evaluated subsequently using the full matched subset. The raw attention maps contain a pronounced verticalstripe component shared across the analyzed test subset. To separate this common component from samplespecific attention variation, the mean-centered attention matrix is defined as

0.0176, a central 95% interval of [0.0088, 0.0300], and a maximum of 0.0585. None of the 10,000 permutations reached the observed value. Using the finite-permutation correction X 1+ I(Sb ≥ Sobs )

Ri = Ai − A,

pperm = 9.999 × 10−5 .

SSbetween , SStotal

(S69)

where SSbetween =

X

2

ng R g − R F ,

(S70)

B+1

,

(S73)

(S74)

The observed statistic and permutation distribution are shown in Fig. S10(e). As a complementary analysis, we tested whether attention maps from samples belonging to the same physical state are more similar than maps obtained from different states. For samples i and j, the correlation distance between their vectorized mean-centered attention maps is defined as dij = 1 − corr [vec(Ri ), vec(Rj )] .

(S75)

The distance ranges from 0 to 2: dij = 0 corresponds to perfect positive correlation, dij = 1 to zero Pearson correlation, and dij = 2 to perfect negative correlation. The average within-state and between-state distances are

g

and

b

the resulting permutation probability is

(S68)

where A denotes the mean L3H4 attention matrix over the analyzed test subset. The resulting Ri therefore contains deviations from the shared attention pattern. Positive and negative entries indicate attention weights above and below the corresponding subset mean, respectively; they do not represent positive or negative contributions to the final model output. Representative mean-centered maps are shown in Fig. S10(c,d). To quantify the association between the mean-centered attention structure and the ground-truth physical states, we calculate the state-explained variance 2 Rstate =

pperm =

Dwithin = mean (dij | gi = gj ) ,

(S76)

Dbetween = mean (dij | gi ̸= gj ) ,

(S77)

and SStotal =

X

Ri − R

2 . F

(S71)

i

Here, ng is the number of samples in state group g, Rg is the mean-centered attention centroid of group g, and R is the mean of the mean-centered attention matrices 2 over the analyzed subset. A value of Rstate = 0 indicates that the state labels explain none of the residual attention variation, as expected for a shared fixed attention pattern with state-independent fluctuations. In contrast, 2 Rstate = 1 corresponds to the limiting case in which all residual variation is attributable to differences between the state-group centroids, with no within-state variation. For L3H4, the observed value is 2 Rstate = 0.674.

respectively. Their separation is defined as ∆D = Dbetween − Dwithin .

(S78)

If sample-specific attention variation is unrelated to physical state, Dwithin and Dbetween should be comparable and ∆D should be close to zero. Conversely, ∆D > 0 indicates that attention maps are more similar within a physical state than between different states. For L3H4, the measured distances are Dwithin = 0.229, Dbetween = 1.195, ∆D = 0.965. (S79)

(S72)

Thus, 67.4% of the mean-centered L3H4 attention variance is associated with differences among the five groundtruth state groups. Statistical significance was evaluated by randomly permuting the state labels among the 229 samples while keeping the attention matrices and group sizes unchanged. A total of B = 10,000 permutations were performed. The resulting null distribution had a mean of

The corresponding average Pearson correlations are approximately 0.771 within states and −0.195 between states. The average different-state distance is therefore approximately 5.2 times the average same-state distance. Sample-label permutation was again used for statistical inference, avoiding the incorrect treatment of the large number of pairwise distances as independent observations. The null distribution of ∆D had a mean of −5.95×10−5 , a central 95% interval of [−0.0121, 0.0193],

20 and a maximum of 0.0565. The observed ∆D = 0.965 exceeded all 10,000 permuted values, yielding pperm = 9.999 × 10−5 .

(S80)

The corresponding permutation analysis is shown in Fig. S10(f). Together, these results distinguish the shared keyposition preference from sample-specific attention routing. After removal of the shared attention component, the residual attention structure retains a robust and reproducible association with the physical input state.

B.

L3H4 Attention Structure and Pathway-Specific Intervention

L3H4 exhibits a clearly structured attention pattern. Fig. S11(a) shows the L3H4 attention weights averaged over the entire training set. Rather than being uniformly distributed, the attention weights form distinct block- and stripe-like motifs across specific query–key token pairs. Because the Alice- and Bob-side single-count tokens are interleaved in the input sequence, structured organization within the single-count sector appears as an alternating stripe-like pattern rather than as a single contiguous block. A particularly pronounced substructure is visually apparent among the Bob-side single-count tokens, consistent with coordinated processing of complementary measurements of the noise-affected Bob output. The coincidence tokens and the acquisition-time token also show partially overlapping attention profiles, which motivates examining whether temporal context participates in the organization of coincidence-related information. At this stage, these patterns are treated only as descriptive features of the attention map rather than as evidence of functional or physical causality. The corresponding cross-sample variance is shown in Fig. S11(b). The variance remains low over most query– key pairs, with only localized regions showing appreciably larger fluctuations. The most visible variations occur within the single-count sector, particularly around several Bob-side relations. Thus, the mean attention structure is not accompanied by broad sample-to-sample variability, but instead consists of a largely reproducible backbone together with a limited number of more sampledependent relations. To characterize the consistency of individual query–key relations, the stability score is defined as Sij = µij − σij ,

(S81)

where µij is the mean attention weight of query–key pair (i, j) and σij is its standard deviation across samples. A larger positive stability score identifies a relation that combines relatively strong mean attention with relatively

small cross-sample variability, whereas smaller or negative values indicate either weak mean attention or larger sample-dependent variation. As shown in Fig. S11(c), the stability map preserves several of the structured motifs observed in the mean attention map. In particular, stable relations are visually apparent within the Bob-side single-count sector and among coincidence-related entries, whereas several cross-domain entries are comparatively weaker or more variable. These descriptive patterns motivate the quantitative relational analysis presented in Sec. X D. Attention structure alone does not establish functional relevance. To examine whether selected token groups contribute functionally through the L3H4 pathway, targeted L3H4 interventions were performed. These interventions do not remove the complete attention head. Instead, only the key columns corresponding to the selected token group are masked in the L3H4 attention logits. The intervened attention is calculated as   QK ⊤ √ +M , (S82) A = softmax d where the entries of M associated with the ablated keys are assigned a large negative value. This prevents all queries in L3H4 from attending to the selected tokens and renormalizes the attention probabilities over the remaining keys. All other attention heads, residual connections, feed-forward layers, and subsequent Transformer layers are left unchanged. The fidelity of the intervened forward pass is then compared with that of the corresponding unmasked forward pass. The resulting fidelity reduction therefore measures the functional contribution of the selected token group specifically through L3H4, rather than its contribution through the complete network. To control for the number of tokens included in each predefined group, each targeted intervention is compared with a size-matched random key ablation. The excess fidelity drop is defined as ∆Fexcess = ∆Ftarget − ∆Frandom ,

(S83)

where ∆Ftarget denotes the fidelity drop induced by the predefined token group and ∆Frandom denotes the corresponding size-matched random-ablation baseline. Positive values indicate that the predefined token group causes a larger fidelity degradation than an equally sized random group, whereas negative values indicate an effect weaker than the random baseline. As shown in Fig. S11(d), the intervention effects are strongly group dependent. Bob-related groups exhibit positive excess fidelity drops, with the Bob-singles intervention producing the strongest effect among the analyzed groups. In contrast, the corresponding Alicerelated groups exhibit negative excess values. The Coinc. H/V/D/R group, where “Coinc.” denotes coincidence, also produces a positive excess fidelity drop.

21

0

(b) Attention variance

0

(a) Mean attention

0.06

3

3

0.25

0.20

0.02

9

6 9

0.10

6

0.04 0.15

12

12

0.05

0

3

6

9

12

0

(c) Stability score

3

6

9

12

0.00

(d) Excess fidelity drop

0

0.10

Coinc. H/V/D/R 3

0.05

Alice singles Bob singles Alice D/R

6

0.00

Bob D/R Alice H/V

9

-0.05

Bob H/V −10−7 0

12

-0.10

0

3

6

9

10−7 10−6 10−5

12

Figure S11: Attention characteristics of Layer 3 Head 4 (L3H4) and intervention effects beyond random ablation. (a) Mean attention weights averaged across all samples. (b) Corresponding cross-sample variance, highlighting localized input-dependent variability. (c) Stability score Sij , defined as the mean attention weight minus its cross-sample standard deviation. (d) Excess fidelity drop for targeted token-group interventions relative to size-matched random key ablations.

These results cannot be explained solely by differences in the number of masked tokens. Instead, different token groups make different functional contributions specifically through the L3H4 pathway. The intervention results therefore complement the attention-weight analysis by showing that the structured attention patterns in Fig. S11(a–c) are associated with group-dependent functional effects. It is important to distinguish this L3H4-specific intervention from the whole-model input perturbation analyzed below. The L3H4 intervention is local and pathway specific: selected tokens are blocked only as keys within L3H4, while the remainder of the network is unchanged. Whole-model input perturbation instead acts directly on the model input, making the perturbed in-

formation unavailable to all attention heads and all subsequent computational pathways. These two interventions therefore address complementary questions: L3H4 intervention evaluates whether selected information contributes through this particular attention head, whereas whole-model input perturbation evaluates its overall contribution to the network.

C.

Comparison with Whole-Model Perturbation and Integrated Gradients

To clarify the relationship between the pathwayspecific L3H4 intervention and global input-importance measures, three analyses were compared using the same

22 Ablations beyond Random

(a)

Absolute IG beyond Random

(b)

Coinc. H/V/D/R

Coinc. H/V/D/R

Alice singles

Alice singles

Bob singles

Bob singles

Alice D/R

Alice D/R

Bob D/R

Bob D/R

Alice H/V

Alice H/V

Bob H/V

Bob H/V

−10−1

−10−3

0

10−1

10−3

2

Alice singles

3

Bob H/V

4

Bob singles

5

Alice H/V

6

Bob D/R

10−4

10−2

L3H4

L3H4

Whole Model

Absolute IG

1.0

1.00

0.5

7

Whole Model

-0.54

1.00

Absolute IG

-0.68

0.89

0.0

Spearman ρ

Within-method corrected-effect rank

(d)

1

0

Excess mean |IG| vs random

(c) Alice D/R

−10−4

−10−2

Extra fidelity drop vs random

−0.5 1.00

Coinc. H/V/D/R

L3H4

Whole Model

Absolute IG

−1.0

Figure S12: Cross-method comparison of token-group importance. (a) Size-matched-random-corrected whole-model input permutation effects. (b) Corrected absolute Integrated Gradients (IG) attribution. (c) Corresponding within-method rankings across L3H4 intervention, whole-model perturbation, and IG. (d) Spearman rank correlations between methods. Whole-model perturbation and IG yield similar end-to-end rankings, whereas L3H4 exhibits a distinct head-specific pattern.

346 held-out test samples: (i) L3H4-specific intervention, (ii) whole-model input permutation/ablation, and (iii) Integrated Gradients (IG). For an input x and a reference baseline x′ , standard IG attributes the final model output to each input feature by integrating the gradient along the straight-line path between x′ and x [18]. For input feature i, the attribution is Z 1 ∂F (x′ + α(x − x′ )) dα, (S84) IGi (x) = (xi − x′i ) ∂xi 0 where F denotes the final model output. Because the gradient is taken with respect to the final output, IG aggregates the influence of an input token through all computational paths connecting that token to

the prediction, including multiple attention heads, residual connections, feed-forward blocks, and subsequent Transformer layers. IG therefore provides an end-to-end attribution of the model output to the input rather than an attribution restricted to a particular internal attention pathway. Whole-model input perturbation is conceptually closer to IG than to the L3H4-specific intervention because it also evaluates end-to-end dependence on the original input features. Whole-model perturbation changes a token group at the model input and measures the resulting change in final prediction or fidelity after the information has been disrupted throughout the network. L3H4 intervention, by contrast, masks only selected key columns

23 within one attention head while leaving other heads and computational pathways unchanged. Nevertheless, IG and whole-model perturbation are not mathematically equivalent. IG is a path-integrated gradient attribution, whereas explicit ablation or permutation measures a finite output change caused by an input perturbation [19]. Perturbation-based analyses consequently provide a complementary means of assessing the functional relevance of features identified by attribution methods [20]. Nonlinearity, feature interactions, and redundancy can lead to quantitative differences between gradient-based attribution and explicit perturbation. For both the whole-model and IG analyses, tokengroup size was controlled by subtracting the mean effect obtained from 20 size-matched random groups. Consequently, the corrected quantities in Fig. S12(a,b) indicate whether a predefined token group produces a stronger effect than would be expected from an equally sized random group. Uncertainty was estimated using 10,000 percentile-bootstrap resamples of the 346 held-out samples, and 95% confidence intervals are reported in the figure. As shown in Fig. S12(a), the whole-model analysis assigns the strongest corrected fidelity effect to Alice D/R, followed by Alice singles and Bob H/V. Bob D/R and Alice H/V fall below their respective size-matched random baselines, whereas the confidence intervals for Coinc. H/V/D/R and Bob singles include zero. These values characterize the end-to-end sensitivity of the network to disruption of the corresponding input information. The absolute IG analysis in Fig. S12(b) yields a broadly similar global pattern. Alice D/R and Alice singles again receive the largest corrected importance values. Several groups exhibit corrected IG values below zero. Because absolute IG magnitude is used here, a negative corrected value does not indicate a negative attribution direction. It indicates only that the absolute attribution magnitude of the predefined group is smaller than the mean absolute attribution magnitude of size-matched random groups. The ranking comparison in Fig. S12(c) further distinguishes global and head-specific importance. Wholemodel perturbation and absolute IG both rank Alice D/R and Alice singles first and second, respectively. L3H4 instead assigns relatively greater importance to Bob singles, Coinc. H/V/D/R, and Bob D/R. The L3H4 ranking is therefore not simply a noisier version of the global ranking; rather, it reflects a different head-specific pattern of information dependence. The relationships among the three rankings are quantified by Spearman correlation in Fig. S12(d). Wholemodel perturbation and absolute IG are strongly positively correlated, ρWhole,IG = 0.893, 95% CI = [0.750, 0.964].

(S85)

By contrast, the L3H4 ranking is negatively correlated with the whole-model analysis, ρL3H4,Whole = −0.536, 95% CI = [−0.643, −0.214],

(S86)

and with absolute IG, ρL3H4,IG = −0.679, 95% CI = [−0.714, −0.571].

(S87)

The confidence intervals were obtained by paired bootstrap over the test samples. These correlations are treated as descriptive cross-method comparisons rather than conventional significance tests across groups, because the seven predefined token groups partially overlap and are therefore not statistically independent. The strong agreement between whole-model perturbation and absolute IG is consistent with their shared end-to-end scope, despite their different mathematical mechanisms. The weaker and opposite relationship between L3H4 and the two global measures is consistent with L3H4 probing a more localized internal informationrouting pathway. More generally, attention-level importance and prediction-level or gradient-based feature importance need not coincide [21]. Accordingly, the three analyses should be regarded as complementary rather than interchangeable. Wholemodel perturbation provides an explicit end-to-end perturbation test, absolute IG provides gradient-based endto-end attribution, and L3H4 intervention isolates the contribution of selected tokens through a specific internal attention pathway. The agreement between whole-model perturbation and IG, together with the distinct L3H4 ranking, supports the interpretation that L3H4 exhibits a specialized information-routing pattern rather than simply reproducing global token importance. D.

Selective Relational Organization in L3H4

To further characterize the role of L3H4, we analyzed pairwise interactions, attention-scope localization, group-level attention reorganization, and standalone intervention importance. The 13 input tokens encode physically distinct information. Alice-side single-count channels provide preparation-side reference information, whereas Bobside single-count channels describe measurements after dynamic Gaussian phase modulation and MMF noise. Coincidence-count channels retain conditional correlation information, while acquisition time specifies the measurement duration and provides information needed to interpret the associated counting statistics and their statistical reliability. In addition, H/V and D/R measurements provide complementary population-sensitive and coherence/phase-sensitive information.

24 The relational analysis indicates that L3H4 exhibits selective rather than uniform information integration. Here, A ↔ B denotes a bidirectional relational analysis between predefined token groups A and B. The interaction excess is defined relative to a size-matched randomtoken baseline, such that positive values indicate enrichment relative to the random baseline and negative values indicate depletion. Fig. S13 summarizes the pairwise L3H4 interaction results. The most pronounced enriched interaction is between Bob H/V and Bob D/R. The observed bidirectional interaction at L3H4 is 0.1527, compared with a size-matched random baseline of 0.0719, giving an interaction excess of Iexcess = 0.1527 − 0.0719 = 0.0808.

(S88)

This corresponds to an approximately 112.4% enhancement relative to the random expectation. The interaction remains significant after Benjamini–Hochberg correction across the 23 unique pairwise tests, p = 0.00498,

q ≃ 0.038.

(S89)

The strong Bob H/V ↔ Bob D/R relation (↔ denotes bidirectional relational analysis) is consistent with localized joint use of complementary output-state information. H/V measurements primarily provide populationsensitive statistics, while D/R measurements provide complementary coherence- or phase-sensitive information. The observed coupling is therefore consistent with an L3H4 representation that jointly organizes these complementary measurements. Coincidence-related relations also exhibit selective enrichment. In particular, the Coinc. D/R ↔ Coinc. H/V interaction and the Coinc. H/V/D/R ↔ Time interaction are enhanced when the analysis is localized to L3H4. These patterns are consistent with a localized organization of complementary-basis coincidence information and with the joint interpretation of coincidence statistics and their acquisition duration. In contrast, several interactions between distinct information domains are depleted relative to their corresponding size-matched random baselines. Alice singles ↔ Bob singles, Alice H/V ↔ Bob H/V, and Alice D/R ↔ Bob D/R show reductions of approximately 29.3%, 27.5%, and 27.6%, respectively. The Coinc. H/V/D/R ↔ all-singles relation is reduced by approximately 18.9%. The coexistence of enriched and depleted interactions argues against interpreting L3H4 as a generic global mixing head. 1.

Localization of L3H4 Relational Structure

To determine whether the interaction structure described above reflects model-wide relations or is preferentially localized to L3H4, the analysis was evaluated at

three nested scopes: all layers and heads, all heads in Layer 3, and L3H4 alone. As shown in Fig. S14(a), the Bob H/V ↔ Bob D/R relation becomes progressively enriched as the analysis is localized. Its interaction excess increases from +0.0397 across all layers and heads to +0.0492 across Layer 3 and finally to +0.0808 at L3H4. The prominent L3H4 interaction therefore cannot be attributed solely to a uniform model-wide background relation; instead, it becomes increasingly concentrated toward Layer 3 and most prominently toward L3H4. Selected coincidence-related relations show an even more localized pattern. The Coinc. D/R ↔ Coinc. H/V interaction excess changes from −0.0247 across all layers and −0.0092 in Layer 3 to +0.0325 at L3H4. The L3H4 value corresponds to an approximately 45.7% enhancement relative to its size-matched random baseline. Similarly, the Coinc. H/V/D/R ↔ Time interaction changes from −0.0074 across all layers and −0.0095 in Layer 3 to +0.0312 at L3H4, corresponding to an approximately 42.9% enhancement relative to random expectation. The sign reversals show that these relations are not uniformly elevated throughout the Transformer, but emerge specifically as the analysis is restricted to L3H4. The group-level statistics in Fig. S14(c,d) provide complementary evidence for this redistribution. Relative to the all-layer baseline, the attention key mass of Coinc. H/V, Coinc. D/R, and Coinc. H/V/D/R increases by approximately 42.4%, 61.3%, and 51.5%, respectively. Their corresponding within-group attention densities increase by approximately 78.7%, 106.4%, and 94.6%. The L3H4 coincidence-related pattern therefore involves both increased recruitment of coincidence information and stronger internal organization of that information. The opposite localization trend is observed for several cross-domain relations, as shown in Fig. S14(b). Alice singles ↔ Bob singles decreases from +0.0115 across all layers to +0.0021 within Layer 3 and −0.0209 at L3H4. Alice H/V ↔ Bob H/V changes from +0.0214 to +0.0127 and then to −0.0197, while Alice D/R ↔ Bob D/R changes from +0.0033 to −0.0091 and finally to −0.0199. Localization to L3H4 therefore produces opposite effects for different classes of relations: selected withindomain interactions become enriched, whereas several cross-domain interactions become depleted. This pattern is more consistent with selective within-domain integration accompanied by cross-domain segregation than with indiscriminate aggregation of all input features.

2.

Relational Coupling and Standalone Intervention Importance

The L3H4 interaction structure is also distinct from the standalone intervention importance of individual token

25 L3H4 pairwise attention interaction Bob HV - Bob DR Bob singles - Bob HV Bob singles - Bob DR Time - Coin D Coinc. DR - Coinc. HV Coinc. HVDR - Time Alice HV - Alice DR D-related - R-related Coinc. D - Alice D Time - Alice D DR-related - HV-related Coinc. D - Bob H H-related - HV-related D-related - DR-related R-related - DR-related Time - Bob H H-related - V-related V-related - HV-related Coinc. HVDR - all singles Alice HV - Bob HV Alice DR - Bob DR BH-FDR Nominal

Alice singles - Bob singles Alice D - Bob H −0.04

−0.02

0.00

0.02

0.04

0.06

0.08

Interaction excess vs. size-matched random baseline

Figure S13: Pairwise attention interactions at L3H4. Interaction excess denotes the observed bidirectional interaction relative to a size-matched random-token baseline. Error bars show 95% confidence intervals, and the dashed line marks zero excess. Filled diamonds indicate interactions that remain significant after Benjamini–Hochberg false-discovery-rate correction (q < 0.05); open circles indicate nominal significance only (p < 0.05, q ≥ 0.05); filled circles indicate all remaining interactions.

groups, as illustrated by Fig. S14(e,f). Bob H/V produces a mean-mask fidelity drop of approximately 12.11%, whereas Bob D/R produces a substantially smaller standalone effect of approximately 0.62%. Nevertheless, Bob H/V ↔ Bob D/R forms the strongest pairwise L3H4 interaction, with an interaction excess of +0.0808. A similar dissociation is observed for coincidence information. Coinc. H/V and Coinc. D/R produce standalone fidelity drops of only approximately 0.19% and 0.16%, respectively, whereas their L3H4 pairwise interaction excess reaches +0.0325. Likewise, Coinc. H/V/D/R and Time have standalone fidelity drops of approximately 0.67% and 0.32%, while their interaction excess is +0.0312. The group-level attention redistribution further indicates that these pairwise effects are not explained by a nonspecific increase in total attention toward the participating groups. Relative to the all-layer baseline, the L3H4 key mass of Bob H/V and Bob singles decreases by approximately 32.3% and 18.9%, respectively, while the

Bob D/R key mass changes only modestly by approximately +3.1%. Nevertheless, their within-group attention densities increase by approximately 28.3%, 38.6%, and 66.3%, respectively.

The strong Bob H/V ↔ Bob D/R interaction therefore does not arise from a nonspecific increase in overall attention toward Bob-related tokens. Instead, the observations are consistent with redistribution of Bob-related attention into a more concentrated internal interaction structure within L3H4.

More generally, these results demonstrate that standalone intervention importance and localized relational coupling quantify different properties of the learned computation. A token group can exhibit a comparatively weak standalone effect while nevertheless participating strongly in a localized relational structure.

26

(a) 0.08

Interaction excess vs. random

(b)

Within-domain integration

Cross-domain segregation

Bob HV - Bob DR Coinc. DR - Coinc. HV Coinc. HVDR - Time Alice HV - Alice DR

0.06

Alice singles - Bob singles Alice HV - Bob HV Alice DR - Bob DR Coinc. HVDR - all singles

0.04

0.02

0.00

−0.02 All layers

Layer 3

L3H4

All layers

Layer 3

L3H4

Attention scope (c)

(d)

Key-mass change at L3H4

Alice singles

-6%

Time

-6%

+25% +28% +95%

+52%

Coinc. HVDR

+106%

+61%

Coinc. DR

+79%

+42%

Coinc. HV

+39%

-19%

Bob singles

+66%

+3%

Bob DR

+28%

-32%

Bob HV −40

(e)

Within-group density change at L3H4

−20

0

20

40

0

60

20

(f)

Standalone causal contribution Alice HV

40

60

80

100

120

Change vs. all-layer baseline (%)

Change vs. all-layer baseline (%)

Local pairwise coupling

Alice HV - DR

Alice DR Bob HV

Bob HV - DR

Bob DR Coinc. HV Coinc. HV - DR

Coinc. DR Coinc. HVDR

Coinc. HVDR - Time

Time 0.0

2.5

5.0

7.5

10.0

12.5

15.0

Mean-mask fidelity drop (%)

17.5

0.00

0.02

0.04

0.06

0.08

L3H4 interaction excess vs. random

Figure S14: Scope localization and functional organization of L3H4 attention interactions. (a,b) Interaction excess relative to size-matched random-token baselines across all layers and heads, all heads in Layer 3, and L3H4 for selected (a) within-domain and (b) cross-domain token-group pairs. (c,d) Percentage change at L3H4 relative to the all-layer baseline in (c) attention key mass and (d) within-group attention density. (e) Mean-mask fidelity drop for selected token groups. (f) L3H4 interaction excess for selected within-domain and contextual pairs. Dashed lines denote zero excess where applicable, and horizontal bars in panels (e) and (f) indicate 95% confidence intervals.

27 E.

Interpretation and Scope of the Attention-Assisted Analysis

The analyses above provide complementary information about the organization and functional relevance of L3H4. First, after the attention component shared across samples is removed, sample-specific L3H4 attention variation remains strongly associated with the ground-truth physical state. This state dependence is supported both by the state-explained variance and by the separation between within-state and between-state residual attention distances. Second, targeted key interventions demonstrate groupdependent functional effects specifically through the L3H4 pathway. The use of size-matched random ablations indicates that these differences cannot be explained simply by the number of tokens included in each group. Third, comparison with whole-model input perturbation and IG establishes a clear distinction between pathway-specific and end-to-end feature importance. Whole-model perturbation and absolute IG show strongly concordant rankings, whereas L3H4 exhibits a distinct ranking. This observation is consistent with L3H4 probing a specialized internal pathway rather than reproducing the global importance of the original input features. Finally, the relational analysis shows that L3H4 contains a selective and domain-structured organization of token-group interactions. The most prominent features include strong Bob H/V ↔ Bob D/R coupling, localized Coinc. D/R ↔ Coinc. H/V organization, localized Coinc. H/V/D/R ↔ Time organization, and reduced interaction among several Alice–Bob and coincidence–singlecount relations. Within the present experimental setting, these observations are consistent with a role for L3H4 in selectively organizing population-sensitive, coherence/phasesensitive, coincidence, and acquisition-duration information during reconstruction of the remotely prepared state under dynamic noise. However, these analyses should not be interpreted as establishing attention weights themselves as direct explanations of the final model prediction. Rather, they provide an attention-assisted, post-hoc attribution analysis supported by complementary statistical and intervention-based validation.

F.

Token and Group Nomenclature

The individual input tokens and composite token groups used in the analyses are summarized below. The abbreviation “Coinc.” is used for coincidence in the text and figures. H/V and D/R denote paired measurementbasis subsets. Composite groups are analytical groupings constructed from the listed input tokens and do not

constitute additional model-input tokens. The notation A ↔ B denotes a bidirectional relational analysis between two predefined token groups.

G.

Interpretability Analysis

High fidelity alone does not guarantee that a machinelearning model has learned physically relevant features. A model may instead exploit spurious correlations or dataset-specific artifacts that are unrelated to the underlying physics. Such shortcut learning has been identified in a variety of contexts, including image classification based on background snow rather than the animal itself [22], source-specific watermarks rather than object features [23], and textual or acquisition-related markers rather than pathological signatures in radiographic COVID-19 detection [24]. This issue is particularly relevant in quantum experiments, where technical noise, experimental imperfections, and correlations introduced during data acquisition may provide unintended predictive cues. Interpretability therefore offers an important complement to fidelitybased benchmarks by testing whether the model response is associated with physically meaningful variables. Here, we use interpretability analysis to examine the dependence of the denoising model on experimentally relevant variables and its response to noise. The above results suggest that the improvement in fidelity does not arise solely from dataset-specific shortcuts. While interpretability does not by itself prove that the model has learned the complete underlying physics, it provides an additional physical consistency check and strengthens the reliability of the denoising results.

XI.

COMPARISON WITH ATTENTION-BASED QUANTUM TOMOGRAPHY METHODS

The present work is closely related to recent attentionbased machine-learning approaches to quantum state tomography (QST) in Refs. [25–27]. Since density-matrix reconstruction and denoising are also integral components of the present framework, the distinction should not be understood simply as one between QST and RSP. Rather, the relevant differences concern the physical task and experimental platform, the data and noise regime, the representation supplied to the learning model, and the level at which the learned model is validated. A quantitative comparison is summarized in Table S5. Because the studies involve different state families, Hilbert-space dimensions, physical platforms, and noise conditions, the reported fidelities and resource requirements should not be interpreted as direct head-to-head benchmarks. Ref. [25] introduced attention-based quantum tomography (AQT), in which a transformer-based generative

28 Table S3: Individual input-token nomenclature used in the attention and intervention analyses. Recommended name

Physical meaning

Coincidence H Coincidence V Coincidence D Coincidence R Alice H Bob H Alice V Bob V Alice D Bob D Alice R Bob R Time

Coincidence count in the H basis. Coincidence count in the V basis. Coincidence count in the D basis. Coincidence count in the R basis. Alice-side single count in the H basis. Bob-side single count in the H basis. Alice-side single count in the V basis. Bob-side single count in the V basis. Alice-side single count in the D basis. Bob-side single count in the D basis. Alice-side single count in the R basis. Bob-side single count in the R basis. Acquisition time.

Token 0 1 2 3 4 5 6 7 8 9 10 11 12

Table S4: Composite token groups used in the attention and intervention analyses. The listed groups are constructed analytically from the original 13 input tokens and are not additional model-input tokens. Recommended name

Physical meaning

Alice singles Bob singles Alice H/V Alice D/R Bob H/V Bob D/R Coinc. H/V Coinc. D/R Coinc. H/V/D/R All singles H/V-related D/R-related H-related V-related D-related R-related

All Alice-side single-count channels. All Bob-side single-count channels. Alice-side H/V single-count channels. Alice-side D/R single-count channels. Bob-side H/V single-count channels. Bob-side D/R single-count channels. H/V coincidence-count channels. D/R coincidence-count channels. All coincidence-count channels. All Alice- and Bob-side single-count channels. All H/V-associated coincidence and single-count channels. All D/R-associated coincidence and single-count channels. H-basis coincidence and Alice/Bob single-count channels. V-basis coincidence and Alice/Bob single-count channels. D-basis coincidence and Alice/Bob single-count channels. R-basis coincidence and Alice/Bob single-count channels.

model learns tomographic measurement statistics and is subsequently used for density-matrix reconstruction. The method was demonstrated numerically and using experimental data acquired from an IBM superconducting quantum processor. The use of self-attention was motivated by its ability to capture correlations among different subsystems. Ref. [27] further incorporated quantum structure into the learning architecture through a quantum-aware transformer, in which measured frequencies and measurement-operator information are jointly encoded and the Bures distance is incorporated into the learning objective. These works therefore establish important attention-based and physics-aware strategies for QST reconstruction. Among Refs. [25–27], Ref. [26] is most closely related to the present work because both approaches use neural networks to improve density-matrix reconstruc-

Token(s) [4, 6, 8, 10] [5, 7, 9, 11] [4, 6] [8, 10] [5, 7] [9, 11] [0, 1] [2, 3] [0, 1, 2, 3] [4, 5, 6, 7, 8, 9, 10, 11] [0, 1, 4, 5, 6, 7] [2, 3, 8, 9, 10, 11] [0, 4, 5] [1, 6, 7] [2, 8, 9] [3, 10, 11]

tion in the presence of noise. In Ref. [26], measurement frequencies are first processed using a conventional QST estimator, such as linear inversion (LI) or MLE. The resulting density-matrix estimate is converted into a vectorized Cholesky representation and supplied to an attention-based neural network, which performs matrix-to-matrix post-processing to obtain an improved density-matrix estimate. The training and evaluation are based on numerically generated tomography data, including finite-statistics noise, depolarization, and measurement/calibration errors. A first major distinction concerns the experimental task and physical platform. Refs. [25, 27] combine numerical QST studies with experimental demonstrations on IBM superconducting quantum processors. In particular, Ref. [25] applies AQT to measurement data acquired from the IBMQ OURENSE device and bench-

29 Table S5: Technical comparison between the present work and Refs. [25–27] The numbers of trainable parameters for Refs. [25, 26] are calculated from the network architectures specified in the corresponding works. Work

Task / platform

Model, data scale, and NN input

States and noise models

Reported performance

[51] Cha et al., MLST 2022

Attention-based generative QST; reconstruction of density matrices from tomographic measurements, including experimental data acquired on an IBM quantum processor.

101,513 trainable parameters; two transformer layers with a 64-dimensional embedding; 2,700 measurements for the 3-qubit IBMQ benchmark and 20,000 classically generated measurements for the 6-qubit demonstration.

GHZ pure-state reconstruction; simulated mixed GHZ states with single-qubit or distributed bit-flip errors; IBMQ hardware noise.

For the 3-qubit IBMQ GHZ benchmark, AQT reports F = 0.917, compared with F = 0.897 for MLE.

[52] Palmieri et al., PRR 2024

NN-enhanced QST; LI/MLE density-matrix estimates are denoised through matrix-to-matrix neural post-processing using a Cholesky representation.

549,503 trainable parameters; 10,000 training and 1,500 validation samples; vectorized Cholesky-matrix input of length d2 , corresponding to 256 real components for the d = 16 four-qubit OAT benchmark.

Simulated finite-shot noise, depolarizing noise, and measurement/calibration bias; Haar/OAT and related tomography benchmarks.

For four-qubit OAT states, the NN-enhanced reconstruction reaches (99.3 ± 0.2)% at 106 trials and (87.6 ± 4.1)% at 103 trials. OOD calibration-noise tests report (88.7 ± 2.3)% and (91.0 ± 1.9)%.

[53] Ma et al., IEEE Trans. Cybernetics 2025

Quantum-aware transformer for QST from structured measurements; simulations and experiments on IBM quantum computers.

Approximately 144K and 152K trainable parameters for the 2- and 4-qubit models, respectively; 95,000 training states and 5,000 evaluation states; measured frequencies and measurement-operator information are incorporated into the transformer.

Pure and mixed 2-, 3-, and 4-qubit states; finite-copy/statistical noise; IBM hardware noise.

On ibmq manila, the mean QAT fidelities are 0.975340 and 0.994796 for 100 and 1000 shots, respectively; on ibmq belem, the corresponding values are 0.976820 and 0.993845.

This work

Interpretable machine-learning-assisted state reconstruction and denoising in an optical RSP experiment through a MMF channel.

152,292 trainable parameters; 1,727 real experimental training samples; 13-dimensional experimental input consisting of four coincidence counts, eight Alice/Bob single counts in the H/V/D/R bases, and acquisition time; density-matrix output.

Real optical experimental data; mixed states, static pure states, and pure states under a time-dependent Gaussian perturbation physically introduced through an AWG-driven phase modulator; polarization perturbations and preparation or measurement imperfections.

Static-state MLE fidelity: 0.8684 ± 0.0093. Under the dynamic Gaussian-noise benchmark, the MLE fidelity decreases to (38.31 ± 12.30)%, whereas the proposed model reconstructs the RSP output with an average fidelity above 99.999%.

marks the reconstruction against MLE tomography, whereas Ref. [27] evaluates the quantum-aware transformer using experimental measurements from the IBM ibmq manila and ibmq belem processors. By contrast, Ref. [26] presents a simulation-based QST-denoising study in which experimentally relevant noise mechanisms are introduced numerically rather than through an operating physical quantum experiment. The present work is implemented directly in an optical remote-state-preparation experiment, in which the remotely prepared state propagates through a MMF channel and both training and test data are acquired from the physical optical system. The measured observables there-

fore contain the combined effects of state-preparation and measurement imperfections, MMF-induced modal mixing and scattering, and associated channel perturbations. In addition, a controlled time-dependent disturbance is physically introduced during acquisition by driving a phase modulator with an arbitrary waveform generator (AWG) generated voltage waveform. The model is consequently required to reconstruct the output density matrix from experimental observables while the optical system is simultaneously affected by intrinsic channel imperfections and externally imposed dynamic perturbations. The learning task is therefore embedded in an actual state-preparation, transmission, and measure-

30 ment process, providing a direct test of machine-learningassisted reconstruction and denoising in a dynamically varying optical RSP environment. A second major distinction lies in the level at which the learned model is validated. The previous works already incorporate physical knowledge or model-level validation in different forms: Ref. [25] motivates selfattention through its ability to represent quantum correlations, Ref. [26] provides a theoretical interpretation of the network as a conditional “debiaser” and examines the role of the transformer architecture, and Ref. [27] explicitly incorporates quantum-measurement information into the network design and learning objective. These analyses, however, address physical motivation, architectural design, or the performance contribution of model components rather than the functional role of a specific learned internal pathway. As discussed in the preceding interpretability analysis, the present work additionally probes a specific internal attention pathway using complementary statistical, intervention-based, and end-to-end attribution analyses. The results provide mutually consistent evidence that the identified internal representation contains statedependent and experimentally structured information and contributes functionally to the reconstruction process. Importantly, this interpretation is not inferred from attention weights alone, but is supported by independent perturbation and attribution tests. To the best of our knowledge, Refs. [25–27] do not report a comparable pathway-level post-hoc analysis combining these complementary forms of validation. The interpretability analysis therefore provides an additional physical-consistency check on the learned reconstruction beyond fidelity-based evaluation alone. A further, although more implementation-specific, difference concerns the input representation and data regime. As summarized in Table S5, the present model uses 1,727 experimentally acquired training samples and a 13-dimensional input consisting of four coincidence counts, eight Alice/Bob single counts in the H/V/D/R bases, and acquisition time, from which the density matrix is reconstructed. In comparison, Ref. [26] uses 10,000 training and 1,500 validation states and applies the neural network to a vectorized Cholesky representation of an already reconstructed density matrix; for the d = 16 four-qubit OAT benchmark, this representation contains 256 real components. The corresponding models contain 152,292 and 549,503 trainable parameters, respectively. These differences illustrate the compact measurementlevel representation and relatively small experimental training set used in the present implementation. They should not, however, be interpreted as establishing universal computational superiority, since the Hilbert-space dimensions, training distributions, and reconstruction tasks are different. The distinction in representation is particularly rel-

evant to Ref. [26]. Whereas its neural network performs post-processing on an LI/MLE-reconstructed density matrix, the present model maps experimentally identifiable observables directly to the density matrix. Moreover, retaining acquisition time as an explicit input provides temporal context associated with the dynamically driven optical system. This measurement-level representation also provides a natural interface for the interpretability analysis, since the learned internal structure can be related directly to physically identifiable quantities such as polarization-resolved coincidence counts, single-count channels, and temporal information. The performance values in Table S5 should be interpreted within these distinct experimental settings. For example, under the dynamic-noise benchmark considered here, the MLE fidelity decreases to (38.31 ± 12.30)%, whereas the proposed model reconstructs the remotely prepared states with an average fidelity above 99.999%. This result does not imply numerical superiority over Refs. [25–27], whose reported fidelities correspond to different states, noise models, and platforms. Rather, it demonstrates that the learned reconstruction remains effective when the measured data are acquired from an operating optical RSP system subject to continuously varying physical perturbations. Taken together, the comparison in Table S5 shows that the contribution of the present work does not lie in the use of attention or machine learning for QST per se. The main distinction is instead the integration of machine-learning-assisted density-matrix reconstruction and denoising into a real optical RSP experiment with an MMF channel and physically imposed dynamic perturbations, together with a mechanism-level analysis of the learned internal representation. Relative in particular to Ref. [26], the present framework therefore extends neural QST denoising from density-matrix post-processing under numerically generated noise toward measurementlevel reconstruction in a dynamically perturbed physical optical system. The additional interpretability analysis further tests whether this learned correction is organized in a manner consistent with physically meaningful experimental information.

XII.

SCALABILITY ANALYSIS AND MBQC PERSPECTIVES

To better contextualize the practical utility of our proposed Transformer-based approach, it is valuable to assess the TQSC model’s scalability and explore its broader implications. In the following subsections, we first examine the classical processing latency and platform-specific applicability of our model. We then analyze the intrinsic scalability of the Transformer architecture itself. Finally, we discuss the potential perspectives of our data-driven noise mitigation strategy for measurement-based quan-

31 tum computing (MBQC) architectures.

TQSC Inference (GPU, this work)

A. Evaluation of TQSC Classical Processing Latency and Platform-Specific Scalability

TQSC Inference (FPGA, projected)

Photonics cycle

The TQSC model requires no additional quantum hardware resources; all TQSC model processing is performed entirely in classical post-processing. Here, the primary scope of this work is quantum state preparation for quantum communication tasks. Nevertheless, classical latency is a critical factor in real-time resource state generation and quantum error correction (QEC). We therefore evaluated the training and inference times of the TQSC model, as well as its feasibility across different quantum platforms. The analysis is organized as follows: 1. Training and Processing Time on Current Hardware. The TQSC inference time is benchmarked on a standard GPU (NVIDIA RTX 3060), yielding about 5.15 µs per state. Training is performed offline and incurs no runtime overhead. While this latency exceeds the ∼ 1 µs throughput target for superconducting qubits, recent Transformer-based surface-code decoders have outperformed conventional algorithmic decoders on physical hardware, confirming the potential of Transformers for practical QEC [28]. Moreover, established hardware-software co-design techniques—lower-precision arithmetic, weight pruning, and deployment on FPGAs or ASICs exploiting fixed-point and sparsity—are expected to bring Transformer decoding into the submicrosecond regime. For instance, a recent FPGA-based QEC decoder has demonstrated decoding latencies of 6590 ns and a total feedback latency of 446 ns [29]. 2. Real-time feasibility across platforms. - Photonic: Resource state generation operates at a cycle time of approximately 1 µs [30]. A 5 µs delay requires roughly 1 km of coiled fiber, a common approach in photonics that introduces low loss (< 5%) [31, 32]. - Superconducting: The syndrome cycle time is approximately 1.1 µs, and the 5.15 µs absolute latency is competitive with recent 63 µs decoding demonstrations [33]. Sustained throughput at the 1.1 µs cycle rate will require future FPGA deployment to eliminate GPU scheduling overhead [29]. - Trapped ions: QEC cycles in trapped-ion platforms extend from tens to hundreds of milliseconds [34]. The 5.15 µs TQSC latency is orders of magnitude below these timescales and is therefore negligible. These comparisons are summarized in Fig. S15, which plots the TQSC inference latency against characteristic timescales of the relevant quantum computing platforms. 3. Extension to QEC Protocols.

(Bartolucci et al., 2023)

Superconducting cycle (Acharya et al., 2025)

Superconducting decoder (Acharya et al., 2025)

Trapped-ion cycle (Ryan-Anderson et al., 2021)

≈ 5.15 µs (Unoptimized) < 1 µs

≈ 1.00 µs ≈ 1.10 µs ≈ 63.00 µs < 200 ms

101 103 105 107 Latency / Cycle Time (µs)

Figure S15: Classical processing latency of the TQSC model compared to characteristic timescales of quantum computing platforms.

The TQSC model currently acts as a data-driven state reconstruction and error mitigation tool at the physical level; it is not a logical QEC decoder. However, the underlying mechanism—learning complex noise distributions to map corrupted measurements back to ideal states—shares a fundamental mathematical isomorphism with QEC decoding. Extension of this Transformer architecture to real-time QEC is feasible and has been validated by recent work such as Google DeepMind’s AlphaQubit, which successfully decodes surface codes using Transformers [28]. To adapt the TQSC architecture into a real-time QEC scheme, particularly for photonic quantum computing, several structural modifications would be considered: - Syndrome decoding adaptation: Rather than reconstructing the full density matrix from projection measurements, the model’s input could be modified to process streams of noisy syndrome measurements. The output layer would be restructured to perform a classification task, predicting probable logical errors or required correction operators. The inherent sequence-modeling capabilities of Transformers suggest an advantage in processing the history of stabilizer measurements to extract encoded logical information. - Mitigation of complex, realistic noise: In physical systems, syndrome information extracted from redundancy checks often suffers from non-local noise, including correlated crosstalk and measurement errors. While traditional algorithms, such as Minimum-Weight Perfect Matching, can face challenges with these asymmetric noise profiles, the Transformer architecture might offer an alternative. By directly processing raw, soft measurement data—such as unthresholded coincidence counts— the model has the potential to adaptively learn complex

32 underlying error distributions, helping to relax the reliance on simplified theoretical noise assumptions. - Continuous decoding for resource states: For protocols that require continuous resource state generation, common in measurement-based quantum computing, the decoder must handle an ongoing stream of data. The Transformer architecture could be augmented with sliding-window attention or recurrent mechanisms. These additions allow the model to maintain an internal decoder state, process incoming continuous syndrome data, and potentially generalize to deeper circuits without incurring memory overflow.

B.

orchestration of fault-tolerant logical gates [39]. Consequently, the Transformer model would target the logical qubit layer rather than the physical one. Since logical qubits are encoded using thousands of physical qubits, the effective dimensionality of the problem is reduced by orders of magnitude, ensuring that the computational overhead remains tractable. Architectural Pathways for Scalable Inference. To further reduce inference time as N grows, several algorithmic enhancements independent of physical hardware accelerations (such as FPGAs) can be employed. • Linear Attention and Efficient Transformers. Integrating mechanisms such as Linformer or Performer can provably reduce the theoretical inference and memory complexity from O(N 2 ) to O(N log N ) or strictly O(N ) [35].

Scalability of the Transformer Architecture

A potential concern regarding scalability arises from the fact that the standard Transformer’s self-attention scales as O(N 2 ) with the system size N [6, 35]. Nevertheless, our approach remains practical and promising for the following reasons. Asymmetry of Training and Inference Costs. As previously noted, the training and inference costs are highly asymmetric: pre-training is a one-time offline cost, whereas inference—generating the remote state preparation protocol—is highly efficient. For our application, inference requires only a single forward pass (approximately 5.15 µs), which is orders of magnitude faster than training [35]. Focus on the NISQ Era. Our current framework is targeted at the Noisy Intermediate-Scale Quantum (NISQ) regime [36]. In this regime, as defined by Preskill, the number of available qubits is not yet sufficient for fault tolerance (typically ranging from tens to a few hundreds), and the primary bottleneck is not asymptotic scalability, but rather maximizing the fidelity of operations under constrained circuit depths. The Transformer’s ability to capture complex, non-local correlations enables the discovery of highly optimized, hardware-specific state preparation protocols that are difficult to obtain with analytical methods. Although the FLOPs of self-attention scale quadratically as O(N 2 ), the wall-clock latency on modern GPUs and TPUs does not follow this trend for NISQ-relevant sequence lengths of up to a few hundred. Matrix multiplications in self-attention are highly parallelized and, for N below 1000, are memory-bandwidth bound, leaving many cores under-utilized [37, 38]. Consequently, increasing N within this regime typically adds little latency, as the extra operations can be largely absorbed by the parallel hardware. Applicability in the Future FTQC Era. Looking forward to the Fault-Tolerant Quantum Computing (FTQC) era, where systems will scale to millions of physical qubits, the classical processing bottleneck shifts from physical-level calibration to syndrome decoding and the

• State Space Models. State-space models such as Mamba substitute attention with linear-time sequence models, achieving O(N ) inference with constant memory while matching Transformer performance [40]. • Knowledge Distillation and Quantization. The heavy pre-trained Transformer can be distilled into a lightweight student model, or post-training quantization (e.g., INT8) can be applied [41]. This drastically reduces the number of active parameters during the forward pass, accelerating inference speed without sacrificing the fidelity of the generated quantum protocols. • Transfer Learning. As previously noted [42], models can be pre-trained on smaller quantum subsystems and efficiently fine-tuned or generalized to larger ones, effectively managing the scaling of classical computation.

C.

Potential Perspectives for MBQC Architectures

Standard MBQC relies heavily on the deterministic generation of pristine, large-scale cluster states [43]. However, maintaining such states is experimentally demanding: hardware noise and decoherence inevitably degrade state fidelity, and errors in the initial cluster state can propagate detrimentally during computation [44, 45]. Our machine-learning approach is inherently data-driven and has the potential to address this challenge directly at the data-processing level. Similar to the hardwareaware noise resilience observed in variational quantum algorithms [46], our model can incorporate specific hardware noise profiles into its training process. By treating noisy projective measurement outcomes as sequence data, the Transformer shows the capability to adaptively learn and mitigate the underlying hardware noise. This

33 demonstrates the potential for efficient reconstruction of high-fidelity density matrices, offering a characterization strategy with inherent robustness to specific decoherence and gate errors, without the need for physical errorcorrection overhead.

∗

These authors contributed equally to this work. [email protected] ‡ [email protected] § [email protected] [1] W. Wu, W.-T. Liu, P.-X. Chen, and C.-Z. Li, Deterministic remote preparation of pure and mixed polarization states, Phys. Rev. A 81, 042301 (2010). [2] W.-T. Liu, W. Wu, B.-Q. Ou, P.-X. Chen, C.-Z. Li, and J.-M. Yuan, Experimental remote preparation of arbitrary photon polarization states, Physical Review A—Atomic, Molecular, and Optical Physics 76, 022308 (2007). [3] K. Kraus, A. Böhm, J. D. Dollard, and W. Wootters, States, effects, and operations fundamental notions of quantum theory: Lectures in mathematical physics at the university of Texas at Austin (Springer, 1983). [4] D. GLOGE, Multimode theory of graded-core fibers, BSTL 52, 935 (1975). [5] D. Marcuse, Theory of dielectric optical waveguides (Elsevier, 2013). [6] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. K. Kaiser, and I. Polosukhin, Attention is all you need, Adv. Neural Inf. Process. Syst. 30 (2017). [7] E. Magesan, J. M. Gambetta, and J. Emerson, Scalable and robust randomized benchmarking of quantum processes, Phys. Rev. Lett. 106, 180504 (2011). [8] J. J. Wallman and J. Emerson, Noise tailoring for scalable quantum computation via randomized compiling, Physical Review A 94, 052325 (2016). [9] D. F. V. James, P. G. Kwiat, W. J. Munro, and A. G. White, Measurement of qubits, Phys. Rev. A 64, 052312 (2001). [10] W. Song, C. Shi, Z. Xiao, Z. Duan, Y. Xu, M. Zhang, and J. Tang, AutoInt: Automatic feature interaction learning via self-attentive neural networks, in Proceedings of the 28th ACM International Conference on Information and Knowledge Management, CIKM ’19 (Association for Computing Machinery, New York, NY, USA, 2019) pp. 1161–1170. [11] Y. Gorishniy, I. Rubachev, V. Khrulkov, and A. Babenko, Revisiting deep learning models for tabular data, in Advances in Neural Information Processing Systems, Vol. 34 (Curran Associates, Inc., 2021) pp. 18932–18943. [12] J. Vig, A multiscale visualization of attention in the transformer model, in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, edited by M. R. Costajussà and E. Alfonseca (Association for Computational Linguistics, Florence, Italy, 2019) pp. 37–42. [13] S. Abnar and W. Zuidema, Quantifying attention flow in transformers, in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, edited by D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault (As†

sociation for Computational Linguistics, Online, 2020) pp. 4190–4197. [14] J. Carrasquilla and R. G. Melko, Machine learning phases of matter, Nature Physics 13, 431 (2017). [15] G. Carleo, I. Cirac, K. Cranmer, L. Daudet, M. Schuld, N. Tishby, L. Vogt-Maranto, and L. Zdeborová, Machine learning and the physical sciences, Reviews of Modern Physics 91, 045002 (2019). [16] J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei, Scaling laws for neural language models, arXiv preprint arXiv:2001.08361 (2020). [17] M. C. Caro, H.-Y. Huang, M. Cerezo, K. Sharma, A. Sornborger, L. Cincio, and P. J. Coles, Generalization in quantum machine learning from few training data, Nature communications 13, 4919 (2022). [18] M. Sundararajan, A. Taly, and Q. Yan, Axiomatic attribution for deep networks, in International conference on machine learning (PMLR, 2017) pp. 3319–3328. [19] M. Ancona, E. Ceolini, C. Öztireli, and M. Gross, Towards better understanding of gradient-based attribution methods for deep neural networks, in International Conference on Learning Representations (2018). [20] S. Hooker, D. Erhan, P.-J. Kindermans, and B. Kim, A benchmark for interpretability methods in deep neural networks, Advances in neural information processing systems 32 (2019). [21] S. Jain and B. C. Wallace, Attention is not explanation, in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) (2019) pp. 3543–3556. [22] M. T. Ribeiro, S. Singh, and C. Guestrin, ” why should i trust you?” explaining the predictions of any classifier, in Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining (2016) pp. 1135–1144. [23] S. Lapuschkin, S. Wäldchen, A. Binder, G. Montavon, W. Samek, and K.-R. Müller, Unmasking clever hans predictors and assessing what machines really learn, Nature communications 10, 1096 (2019). [24] A. J. DeGrave, J. D. Janizek, and S.-I. Lee, Ai for radiographic covid-19 detection selects shortcuts over signal, Nature Machine Intelligence 3, 610 (2021). [25] P. Cha, P. Ginsparg, F. Wu, J. Carrasquilla, P. L. McMahon, and E.-A. Kim, Attention-based quantum tomography, Machine Learning: Science and Technology 3, 01LT01 (2022). [26] A. M. Palmieri, G. Müller-Rigat, A. K. Srivastava, M. Lewenstein, G. Rajchel-Mieldzioć, and M. Plodzień, Enhancing quantum state tomography via resourceefficient attention-based neural networks, Phys. Rev. Res. 6, 033248 (2024). [27] H. Ma, Z. Sun, D. Dong, C. Chen, and H. Rabitz, Tomography of quantum states from structured measurements via quantum-aware transformer, IEEE Transactions on Cybernetics (2025). [28] J. Bausch, A. W. Senior, F. J. H. Heras, T. Edlich, A. Davies, M. Newman, C. Jones, K. Satzinger, M. Y. Niu, S. Blackwell, G. Holland, D. Kafri, J. Atalaya, C. Gidney, D. Hassabis, S. Boixo, H. Neven, and P. Kohli, Learning high-accuracy error decoding for quantum processors, Nature 635, 834 (2024).

34 [29] J. Liu, Y. Lee, Y. Xu, G. Huang, and X. Wu, A scalable open-source qec system with sub-microsecond decodingfeedback latency, arXiv preprint arXiv:2603.16203 (2026). [30] H. Aghaee Rad, T. Ainsworth, R. N. Alexander, B. Altieri, M. F. Askarani, R. Baby, L. Banchi, B. Q. Baragiola, J. E. Bourassa, R. Chadwick, et al., Scaling and networking a modular photonic quantum computer, Nature 638, 912 (2025). [31] S. Bartolucci, P. Birchall, H. Bombin, H. Cable, C. Dawson, M. Gimeno-Segovia, E. Johnston, K. Kieling, N. Nickerson, M. Pant, et al., Fusion-based quantum computation, Nature Communications 14, 912 (2023). [32] M.-J. Li and T. Hayashi, Advances in low-loss, largearea, and multicore fibers, in Optical Fiber Telecommunications VII (Elsevier, 2020) pp. 3–50. [33] Quantum error correction below the surface code threshold, Nature 638, 920 (2025). [34] C. Ryan-Anderson, J. G. Bohnet, K. Lee, D. Gresh, A. Hankin, J. P. Gaebler, D. Francois, A. Chernoguzov, D. Lucchetti, N. C. Brown, et al., Realization of real-time fault-tolerant quantum error correction, Physical Review X 11, 041058 (2021). [35] Y. Tay, M. Dehghani, D. Bahri, and D. Metzler, Efficient transformers: A survey, ACM Computing Surveys 55, 1 (2022). [36] J. Preskill, Quantum computing in the nisq era and beyond, Quantum 2, 79 (2018). [37] W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica, Efficient memory management for large language model serving with

pagedattention, in Proceedings of the 29th symposium on operating systems principles (2023) pp. 611–626. [38] R. Pope, S. Douglas, A. Chowdhery, J. Devlin, J. Bradbury, J. Heek, K. Xiao, S. Agrawal, and J. Dean, Efficiently scaling transformer inference, Proceedings of machine learning and systems 5, 606 (2023). [39] E. T. Campbell, B. M. Terhal, and C. Vuillot, Roads towards fault-tolerant universal quantum computation, Nature 549, 172 (2017). [40] A. Gu and T. Dao, Mamba: Linear-time sequence modeling with selective state spaces, arXiv preprint arXiv:2312.00752 (2023). [41] A. Gholami, S. Kim, Z. Dong, Z. Yao, M. W. Mahoney, and K. Keutzer, A survey of quantization methods for efficient neural network inference, in Low-power computer vision (Chapman and Hall/CRC, 2022) pp. 291–326. [42] A. Mari, T. R. Bromley, J. Izaac, M. Schuld, and N. Killoran, Transfer learning in hybrid classical-quantum neural networks, Quantum 4, 340 (2020). [43] R. Raussendorf and H. J. Briegel, A one-way quantum computer, Physical review letters 86, 5188 (2001). [44] M. A. Nielsen, Cluster-state quantum computation, Reports on Mathematical Physics 57, 147 (2006). [45] C. M. Dawson, H. L. Haselgrove, and M. A. Nielsen, Noise thresholds for optical quantum computers, Physical review letters 96, 020501 (2006). [46] M. Cerezo, A. Arrasmith, R. Babbush, S. C. Benjamin, S. Endo, K. Fujii, J. R. McClean, K. Mitarai, X. Yuan, L. Cincio, et al., Variational quantum algorithms, Nature Reviews Physics 3, 625 (2021).

105

Fit Curve Measured CAR

10

5

Fit Curve Measured CAR

104

Channel CH22&CH42

104

Channel CH23&CH41

CAR

CAR

35

103

103

102

102 0

10

20

30

40

0

10

20

30

Pump power (µW)

Pump power (µW)

(a)

(b)

106

Fit Curve Measured CAR

105

40

Fit Curve Measured CAR

Channel CH24&CH40

104 10

CAR

CAR

105

Channel CH25&CH39 10

4

3

103 102 0

10

20

40

0

5

10

15

20

Pump power (µW)

(c)

(d)

Fit Curve Measured CAR

Channel CH26&CH38

104 103

25

Fit Curve Measured CAR

105

CAR

105

CAR

30

Pump power (µW)

Channel CH27&CH37

104

103

102 102 0

10

20

30

40

0

10

20

30

Pump power (µW)

Pump power (µW)

(e)

(f)

40

Figure S16: Characterization of full-spectrum signal-to-noise ratio performance. The Coincidence-to-Accidental Ratio (CAR) as a function of pump power is shown for 10 symmetric channel pairs. (a)-(f) show the outer pairs from CH22 & CH42 to CH27 & CH37.

105

Fit Curve Measured CAR

104

Channel CH28&CH36

10

CAR

CAR

36

103

Fit Curve Measured CAR

5

Channel CH29&CH35

104

103

102 102 10

20

0

10

20

30

Pump power (µW)

(g)

(h)

40

Fit Curve Measured CAR

105

Fit Curve Measured CAR

Channel CH30&CH34

104

Channel CH31&CH33

105

CAR

40

Pump power (µW)

106

10

30

4

CAR

0

103 103 102 0

10

20

30

40

0

10

20

30

Pump power (µW)

Pump power (µW)

(i)

(j)

40

Figure S16: (Continued) Characterization of full-spectrum signal-to-noise ratio performance. (g)-(j) show the inner pairs from CH28 & CH36 to CH31 & CH33. Blue triangles represent experimental data, and dashed lines indicate theoretical fits (CAR ∝ 1/Ppump ). All channels exhibit CAR values exceeding 200,000 at low pump powers, indicating consistently high signal quality across the entire generated bandwidth.

37

CH23&CH41 1200

1250

Coincidence count / s

Coincidence count / s

CH22&CH42

1000 750 500 250

Fit 0.0° Fit 22.5° Data 0.0° Data 22.5°

0 0

20

1000 800 600 400 Fit 0.0° Fit 22.5° Data 0.0° Data 22.5°

200 0

40

60

80

0

∘

20

(b)

CH24&CH40

CH25&CH39 Coincidence count / s

Coincidence count / s

1200 1000 800 600 400 Fit 0.0° Fit 22.5° Data 0.0° Data 22.5°

0

20

1250 1000 750 500 250

Fit 0.0° Fit 22.5° Data 0.0° Data 22.5°

0 40

60

80

0

Rotation angle ( ∘ )

20

Coincidence count / s

Coincidence count / s

800 600 400 Fit 0.0° Fit 22.5° Data 0.0° Data 22.5°

20

1000 800 600 400 Fit 0.0° Fit 22.5° Data 0.0° Data 22.5°

200 0

40

60

Rotation angle ( ∘ ) (e)

80

CH27&CH37

1000

0

60

(d)

CH26&CH38

0

40

Rotation angle ( ∘ )

(c)

200

80

Rotation angle ( )

(a)

0

60 ∘

Rotation angle ( )

200

40

80

0

20

40

60

80

Rotation angle ( ∘ ) (f)

Figure S17: Polarization entanglement correlation fringes for 10 symmetric channel pairs. (a)-(f) show the outer pairs from CH22 & CH42 to CH27 & CH37.

38

CH29&CH35

1000

Coincidence count / s

Coincidence count / s

CH28&CH36

800 600 400 200

Fit 0.0° Fit 22.5° Data 0.0° Data 22.5°

0 0

20

1200 1000 800 600 400 Fit 0.0° Fit 22.5° Data 0.0° Data 22.5°

200 0

40

60

80

0

∘

20

(h)

CH30&CH34

CH31&CH33 Coincidence count / s

Coincidence count / s

1500 1250 1000 750 500 Fit 0.0° Fit 22.5° Data 0.0° Data 22.5°

0

20

40

60

Rotation angle ( ∘ ) (i)

80

Rotation angle ( )

(g)

0

60 ∘

Rotation angle ( )

250

40

80

1250 1000 750 500 250

Fit 0.0° Fit 22.5° Data 0.0° Data 22.5°

0 0

20

40

60

80

Rotation angle ( ∘ ) (j)

Figure S17: (Continued) Polarization entanglement correlation fringes for 10 symmetric channel pairs. (g)-(j) show the inner pairs from CH28 & CH36 to CH31 & CH33. The plots show coincidence counts as a function of the polarization analyzer angle. Blue and red curves correspond to measurements in the 0◦ and 22.5◦ bases, respectively (points represent experimental data; solid lines represent sinusoidal fits). High interference visibility (average > 98%) is observed in both bases for all channels.

Record · ID 978434 · SHA-256 cbd336616c23e0d7
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.