Noise-Robust Quantum State Characterization for Remote State Preparation with Deep Learning Bo Tang,1, ∗ Zixuan Liao,1, ∗ Hao Li,1, ∗ Yilin Yang,1 Jiani Lei,1 Zengya Li,1 Jing Qiu,1 Zhaohui Dong,1 Zhengyang Mao,1 Yuanhua Li,2, † Yuanlin Zheng,1, 3, 4, ‡ and Xianfeng Chen1, 3, 4, 5, §
arXiv:2609.20523v1 [quant-ph] 17 Sep 2026
1
State Key Laboratory of Photonics and Communications, School of Physics and Astronomy, Shanghai Jiao Tong University, Shanghai 200240, China 2 Department of Physics, Shanghai Key Laboratory of Materials Protection and Advanced Materials in Electric Power, Shanghai University of Electric Power, Shanghai 200090, China 3 Hefei National Laboratory, Hefei 230088, China 4 Shanghai Research Center for Quantum Sciences, Shanghai 201315, China 5 Collaborative Innovation Center of Light Manipulations and Applications, Shandong Normal University, Jinan 250358, China (Dated: September 18, 2026) Quantum communication underpins secure information processing and scalable quantum networks. In particular, remote state preparation (RSP) enables efficient quantum state transfer, but accurately estimating target states under complex noise remains challenging. Here, we propose a Transformer-based Quantum State Characterizer (TQSC) model for noisy RSP experiments. Our model reconstructs experimentally prepared pure and mixed photonic polarization states from noisy measurements in complex scattering environments, while its attention patterns provide physically grounded insights into correlations among the measured observables. The method achieves a mean estimator-target fidelity exceeding 99.999% under complex scattering and dynamic Gaussian noise, while its robustness and generalization are further examined using Qiskit-simulated Bloch-ball states. Furthermore, in a practical MNIST image transmission task with held-out states, the decoded bit error rate is reduced from 50.34% to zero after TQSC post-processing. The TQSC model enables accurate tomographic characterization under dynamic noise and provides physically grounded posthoc insights, holding promise for intelligent quantum information processing applications.
I.
INTRODUCTION
Quantum information science drives advancements in communication, computing, and sensing by harnessing quantum mechanics [1–3]. Quantum communication is the cornerstone of the future quantum internet, enabling secure information transfer [4, 5]. Reliable distribution of quantum states between distant nodes is a fundamental task for practical quantum networks [6]. Among protocols, remote state preparation (RSP) is a promising approach [7, 8]. RSP enables Alice to remotely prepare a known quantum state at Bob’s site using shared entanglement and classical communication, offering reduced classical resources than teleportation [9] and improved security over direct transmission [10]. Firstly demonstrated in liquid-state nuclear magnetic resonance [11], RSP is extended to single- and multi-qubit states [12– 14]. Recent advances include photonic states with nonclassical features [15], high-dimensional states [16] using hybrid entanglement [17], orbital-angular-momentum lattices [18], and metasurfaces [19], and implementations in hybrid platforms [20]. Reliable quantum state preparation and characterization is central to practical quantum information processing. In RSP protocols, pure states are essential for computation, communication, and metrology, while mixed states offer noise resilience and resource efficiency, enabling advantages in realistic quantum communication and computation [21, 22]. Accurate reconstruction
through quantum state tomography (QST) and fidelity estimation is therefore critical [23, 24]. Nevertheless, hardware imperfections—including decoherence [25], dynamic scattering [26], control errors [27], and state preparation and measurement (SPAM) errors [28]—invariably constrain the achievable fidelity. To combat complex noise environments, machine learning has emerged as a powerful tool for quantum estimation and control [29, 30]. Specifically, learningbased approaches have been applied to quantum error mitigation [31–34], dynamic state preparation [35, 36], and QST, where neural networks map noisy measurement statistics to density matrices [37–41]. The Transformer [42], featuring self-attention to model sequential dependencies and multi-head attention to capture diverse contextual relationships, has demonstrated strong multimodal performance [43, 44]. In quantum physics, this architecture has been employed across a wide range of problems [45], including electronic structures [46], quantum correlations [47], noise-robust quantum communication [48], automated circuit generation [49], scalable quantum error correction [50]. Previous attention-based QST methods have learned measurement distributions [51], denoised LI/MLE estimates through Cholesky representations under simulated noise [52], and reconstructed states from structured measurements with IBM-device validation [53]. However, these studies do not address quantum-state reconstruction in real physical environments involving com-
2 plex scattering and dynamically varying noise. Moreover, conventional ‘black-box’ neural networks provide limited physical insight, making it difficult to fully trust the reconstructed outputs. High model performance does not necessarily indicate that the model has learned meaningful and task-relevant features [54–56]. Here, we propose a Transformer-based Quantum State Characterizer (TQSC) model with built-in physical constraints for quantum state reconstruction, designed to improve reconstruction fidelity and robustness in noisy RSP while enabling physically grounded post-hoc interpretation. Using photonic polarization as the experimental platform, TQSC reconstructs density matrices from noisy measurements under both MMF transmission alone and MMF transmission with additional dynamic perturbations. Beyond reconstruction fidelity, we further analyze the learned internal representation using statistical, intervention-based, and end-to-end attribution methods rather than relying on attention maps alone. Experimentally, TQSC substantially improves reconstruction performance, achieving an average fidelity above 99.999%. These results demonstrate a measurement-level and experimentally validated framework for robust quantumstate reconstruction in dynamically perturbed optical RSP systems. More broadly, reliable state reconstruction under realistic experimental noise and perturbations may support state monitoring, validation, and diagnostics in practical quantum networks and distributed quantum information processing. II.
THE SCHEME
The positive operator-valued measure (POVM) provides a fundamental framework for quantum measurement theory [57]. In our RSP protocol, the probability of obtaining measurement outcome m is given by † pm = Tr(Mm Mm ρ),
(1)
where the measurement {Mm } satisfy the comP operators † pleteness condition m Mm Mm = I. Following the theoretical framework in Ref. [58], but with a distinct entanglement source, we implement the RSP protocol for mixed states. Taking a mixed state as an example, Bob’s target mixed state to be transmitted is expressed as ρB = α2 |H⟩⟨H| + β 2 |V ⟩⟨V |,
(2)
(where α2 +β 2 = 1). The maximally entangled Bell state generated by our entanglement source takes the form 1 |Φ ⟩AB = √ (|H⟩A |H⟩B + |V ⟩A |V ⟩B ) . 2 +
(3)
After passing through the POVM-based pre-processing module, projective measurements on specific polarization states are implemented using analyzers. When Alice’s photon is projected onto the |H⟩ state, the remote preparation of Bob’s photon collapses to the desired state ρ̂IB
(see eq. (2)). For other measurement outcomes, Bob applies appropriate local unitary operations {σ̂x , σ̂z , σ̂y } to recover the target state. The experimental implementation of RSP is inevitably affected by various physical imperfections. Quantum noise originates from uncontrollable interactions between a quantum system and its environment, leading to deviations from ideal unitary evolution. We describe this noise using the Kraus operator formalism of quantum channels. In our analysis of RSP, we consider four primary sources of noise: photon loss, decoherence, fiber-induced noise, and detector imperfections. The final state of Bob, incorporating these noise effects, is given by X (k′ ) † ′ ′ ′ρ ρfinal P (k = )E U U . (4) ′ k k B B,ideal k k′
Here, P (k ′ ) denotes the experimentally observed proba(k′ ) bility of Alice obtaining measurement outcome k ′ , ρB,ideal is Bob’s ideal conditional state associated with outcome k ′ before the correction operation, Uk′ is the corresponding correction unitary applied by Bob, and Ek′ represents an outcome-dependent effective trace-preserving noise map for the entire noisy RSP branch, not only for postcorrection noise. The final state is obtained by averaging the noise-affected corrected states over all possible measurement outcomes. Our goal is to suppress this noise. In the context of neural networks, the entire process can be viewed as a complex quantum-to-classical channel, denoted by C, which maps an ideal target quantum state ρtarget onto a set of classical measurement outcomes. The data vector thus constitutes a noise-corrupted quantum state: M = C(ρtarget ).
(5)
The map C formally integrates the four previously discussed noise sources, representing them as a unified framework of coherent and incoherent dynamic disturbances. To counteract these effects, we propose a TQSC framework that employs a neural network to parameterize an inverse channel, denoted as ΦTQSC (·; θ). The predicted quantum state ρpred is obtained from the measurement data M via the inverse map: ρpred = ΦTQSC (M; θ).
(6)
Further details on the POVM implementation, RSP communication noise, and the model are provided in the Supplemental Material (SM). III.
EXPERIMENTAL SETUP
The experimental setup (Fig. 1) utilizes a polarization Sagnac interferometric loop with an integrated spontaneous parametric down-conversion (SPDC) source to generate phase-stable polarization-entangled photon
3
(a)
(b)
HWP
PC
FBS
WDM
PBS
Coupler
DHWP
QWP
FPBS
DWDM
DPBS
SNSPD
PP L
N
Sagnac Loop
MMF
CH31
DM
×
TCSPC
PM AWG
CH33 ODL
Polarization Entanglement
Bob
H Model
Laser
V
Coincidence Counting
Alice
VOA
Source
User Distribution
(c)
Density-Matrix Reconstruction Matrix Reconstruction
Add&Norm Linear
Feed Forward
Unsqueeze
Add&Norm
Input Tensor
Multi-Head Attention
Lower-Triangular Factor C
Diagonal
… Tokens
Data and Pre-processing
Model
Lower Re
Lower Im
Parameter Splitting
Linear
Positional Encoding
+
Linear
Transformer Encoders
Post-processing
Figure 1: Schematic diagram of the experimental setup for RSP. (a) Polarization entanglement source based on Sagnac loop. (b) RSP of user distribution device. PPLN, periodically poled lithium niobate waveguide; DM, dichroic mirror; HWP, halfwave plate; DHWP, dual-wavelength half-wave plate; QWP, quarter-wave plate; PC, polarizaton controller; FBS, fiber beam splitter; FPBS, fiber polarization beam splitter; WDM, wavelength division multiplexing; DWDM, dense wavelength division multiplexing; PM, phase modulator; AWG, arbitrary waveform generator; PBS, polarizing beam splitter; DPBS, dual-polarizing beam splitter; MMF, multimode fiber; ODL, optical delay line; VOA, variable optical attenuator; SNSPD, Superconducting nanowire single-photon detector; TCSPC, timecorrelated single-photon counting. (c) Multidimensional measurement data are processed through a preprocessing module, a three-layer Transformer Encoders module for feature extraction, and a postprocessing module that reconstructs the predicted density matrix. The loss function quantifies the difference between predicted and actual labels, minimized via the Adam optimizer.
pairs. A femtosecond laser pumps a periodically poled lithium niobate (PPLN) waveguide within the loop, producing degenerate photon pairs at the telecom wavelength. Waveplates optimize the source for maximal Bell state generation (Eq. 3). Subsequent wavelength filtering routes photons to separate receivers for Alice and Bob. Alice’s detection module employs a beam splitter to divide the incoming photon. One path incorporates a tunable delay for synchronization, while the other includes a variable attenuator to adjust the intensity ratio η = α2 /β 2 between measurement bases. Polarization controllers (PC) compensate fiber effects before polarization projection. A hybrid detection scheme combines specific polarization outputs at a second beam splitter for
state analysis, followed by final projective measurement. Bob’s receiver includes a PC for fiber-effect compensation, followed by a segment of MMF mimicking spatially scrambling in scattering media [59, 60]. Dynamic noise is generated by driving the phase modulator (PM) with signals from the arbitrary waveform generator (AWG). QST is performed with a polarization analyzer composed of a quarter-wave plate (QWP), a half-wave plate (HWP), and a polarizing beam splitter (PBS) arranged in sequence. For preparing pure states, the PBS in Alice’s polarization analyzer is retained and the QWP at the entanglement source is adjusted to control the phase. For mixed states, the PBS in Alice’s analyzer is removed. Both users employ superconducting nanowire single-photon detectors (SNSPDs). Detection events
4 (a)
(b)
IV.
0.8
0.6
0.4
0.2
0
-0.2
Initially, we prepare a set of mixed states with η ranging from 1 to 7, given by ρη =
Real Part
Imaginary Part
(c) 1.0 0.9
Fidelity
RESULT
0.8 0.7 0.6 0.5
Figure 2: Characterization and verification of RSP. (a) Remote quantum states are prepared on the Bloch ball, with the pure state shown as a red dot and mixed states shown as blue dots [61]. (b) The real and imaginary parts of the density matrix with η = 1. (c) Reconstruction fidelity of the pure state and mixed states prepared via RSP through an MMF.
1 (|H⟩⟨H| + η|V ⟩⟨V |). η+1
Furthermore, we prepare a pure state |ψp ⟩ as 1 |ψp ⟩ = √ 3|H⟩ + eiπ/3 |V ⟩ . 10
The structure of the TQSC is illustrated in Fig. 1(c). The model takes a 13-dimensional input vector representing measurement data for quantum states and outputs the corresponding density matrix. Input data points are treated as tokens. After preprocessing with linear transformation and dimension adjustment, the data passes through a stack of Transformer encoders. These incorporate positional encoding and multihead self-attention to capture dependencies between sequence elements, with residual connections and layer normalization stabilizing the learning process. A feedforward network then extracts deeper nonlinear features. In the post-processing stage, the extracted features are linearly transformed into real factor parameters and split into diagonal, lower-real, and lower-imaginary components. These components form a lower-triangular complex factor C, and the final density matrix is reconstructed as ρ = CC † /Tr(CC † ). Model training minimizes the Mean Squared Error (MSE) between predicted and true density matrices using the Adam optimizer. Details of the experimental setup, model architecture, and training parameters are provided in the SM.
(8)
The corresponding data points are shown in Fig. 2(a). QST uses the projection bases H, V , D, and R, corresponding to horizontal, vertical, diagonal, and rightcircular polarizations, respectively. We prepare eight quantum states, averaging five measurements per basis without the MMF to benchmark state-preparation fidelity, and utilize single-shot measurements for all subsequent MMF data. To verify the prepared states, we reconstruct the density matrix ρB using the maximum likelihood method [62] and calculate the fidelity ⟨F ⟩ with the target prepared state ρp [63]. The fidelity is computed using the following formula: q√ √ 2 ρp ρB ρp F (ρp , ρB ) = Tr .
are recorded with time-correlated single-photon counting (TCSPC).
(7)
(9)
Without the MMF, the average fidelity of these states reaches (99.25 ± 0.20)%, demonstrating the high statepreparation capability of our experimental setup. Upon introducing the MMF, however, the fidelities of the states decrease to varying degrees, as illustrated in Fig. 2(c). As an example, the real and imaginary parts of the density matrix for the case η = 1 are reconstructed and shown in Fig. 2(b). The impact of dynamic disturbances on the system coherence is then evaluated. In the absence of noise, the prepared quantum states exhibit a purity of 1 and a fidelity of (99.23 ± 0.20)%. When zero-mean Gaussian dynamic noise (σ = 1 V) generated by the AWG is introduced by driving the PM (Vπ = 5 V) in the MMF setup, the purity decreases to 0.6724 ± 0.0698, while the fidelity drops to (38.31 ± 12.30)%, indicating significant decoherence and fidelity degradation under the applied noise conditions. To overcome these limitations, our model is comprehensively trained on pure states, mixed states, and pure states under dynamic noise, achieving an overall average estimation fidelity of 0.99999497 ± 8.78 × 10−6 when evaluated on both training and test sets. As illustrated in Fig. 3(a), the estimation fidelity distributions for both the training and test sets are tightly clustered around values approaching 1, indicating precise agreement between predicted results and theoretical values. To benchmark the reconstruction performance, we compare the TQSC
5 (a) 4
(b)
10
Train Data Test Data 3
(a)
Coinc. H/V/D/R Alice singles
Send
Bob singles Alice D/R
102
Image Reconstruction over Noisy Channel Receive MMF Transmission
TQSC Inference
··· 0 1 0 0 1 1 ··· ··· ···
··· 0 1 0 0 1 1 ··· ··· ···
Bob D/R Alice H/V Bob H/V
100 0.99990
0.99995
1.00000
Fidelity
(c)
BTC Compressed −10−7 0
10−7 10−6 10−5
Excess fidelity drop
(d)
(b)
Noisy Channel TQSC Model
3
0.20
0.06
12
12
0.05
0
3
6
9
12
0
3
6
9
12
0.00
Figure 3: Performance and attention analysis of the TQSC model. (a) Fidelity distributions for the training and test datasets. (b) Excess fidelity drop from targeted token-group interventions relative to size-matched random ablations. (c) Mean attention map of Layer 3 Head 4 (L3H4). (d) Attention variance map of L3H4. In (c) and (d), the query (y) and key (x) axes index the 13 input tokens (0–12), comprising H/V/D/R coincidence counts, the corresponding Alice and Bob single-photon counts, and measurement time.
model with representative baselines; it achieves the highest average test fidelity, with a slight advantage over the parameter-matched MLP baseline. Furthermore, the model demonstrates robust generalization to new parameter values, maintaining estimation fidelities above 99% in tests involving both interpolation (η = 4.5) and leave-one-out validation (η = 6). In addition, we test the model on Qiskit-simulated Bloch-ball states beyond the experimentally measured state family, where it also achieves estimation fidelities exceeding 99%. This result demonstrates that the model maintains high estimation fidelity on unseen test data with a slight performance degradation, indicating strong generalization ability and robustness without evident signs of overfitting. Using only single-shot measurement data as input, the TQSC model reliably reconstructs quantum states with high estimation fidelity. Recent MPO-based methods target large-scale mixedstate learning under noise [64], rather than direct purestate optimization under dynamic noise. Admittedly, the train set scattered high-fidelity outliers suggest potential sensitivity of the model to edgecase data. Additionally, the limited sample size of the test set (n < 400) may necessitate expansion of the validation set to enhance statistical inference reliability. L3H4-targeted token-group ablations in Fig. 3(b) re-
Bit 0
122
475
597
0
500 400 300
Bit 1
117
462
0
579
Bit 0
Bit 1
Bit 0
Bit 1
200 100 0
0.02
9
6 9
0.10
6
0.04 0.15
Overlap
Actual Bits
0 3
0.25
Model Denoised Result BER = 0%
(c)
Attention variance
0
Mean attention
Reconstructed with Noise BER = 50.34%
Received Bits
Number of Samples
101
Probability
Count (log)
10
Predicted Bits
Number of outcomes (N=100)
Figure 4: TQSC-based image transmission and reconstruction. (a) Degradation and recovery of BTC-compressed MNIST images through the noisy channel. (b) Probability distributions of |V ⟩ measurement outcomes (N = 100) for mixed states ρ2 and ρ3 . (c) Confusion matrices for the noisy channel (left) and TQSC predictions (right).
veal a nonuniform local feature dependence, dominated by Bob-side single counts, with smaller positive contributions from coincidence and Bob polarization features. To gain insight into how the TQSC model integrates experimental information when optimizing quantumstate fidelity, we analyze the attention maps of Layer 3 Head 4 (L3H4), a deeper Transformer layer that may capture higher-order, task-relevant dependencies among input features [65]. As shown in Fig. 3(c), the mean attention map exhibits localized high-attention regions, indicating preferential interactions among selected query–key pairs. The attention variance in Fig. 3(d) is concentrated on a smaller subset of pairs, showing that only selected interactions vary appreciably across samples. Together, these features indicate selective rather than global information integration. Quantitative analysis of these localized attention patterns further reveals enhanced Bob H/V–Bob D/R and Coinc. H/V–Coinc. D/R interactions, consistent with the model integrating complementary population- and coherence/phase-sensitive information. In addition, interactions between the coincidence features and acquisition time suggest that the model interprets the coincidence statistics in conjunction with the measurement duration, which determines the degree of statistical averaging and hence the reliability of the measured counts. Together, these observations suggest that the deeper-layer attention does not merely emphasize individual observables, but captures structured correlations among multi-
6 ple experimentally relevant features. Quantitative attention and intervention analyses are provided in the SM. Ultimately, to further validate the practicality of the proposed scheme, we conduct an image transmission experiment under the same system configuration as described above. In the complex scattering environment of MMF, mixed states ρ2 and ρ3 are respectively employed to encode bits 0 and 1, and fidelity-based criteria are used for bit decision. In Fig. 4(b), with the state fidelity of F (ρ2 , ρ3 ) ≈ 99.16%, measurement-induced Gaussian noise causes the distributions to overlap, making the states difficult to resolve even without the MMF. The N = 100 measurements are used only to characterize the projection-outcome distributions and are not used for decoding or BER evaluation. A grayscale image of the handwritten digit “0” from the MNIST dataset is selected. After BTC compression [66], the 28 × 28 pixel image is encoded into 1176 bits, followed by quantum state preparation and transmission using the experimental setup. The transmitted quantum-state data are all held-out states unseen by the model, with each state measured only once. As shown in Fig. 4(a), the original bit error rate (BER) is 50.34%; after applying the TQSC model, the BER is reduced to 0%. As depicted in Fig. 4(c), the confusion matrices show bit errors in the noisy channel and accurate recovery by the TQSC model. These results demonstrate the strong practicality and generalization capability of the model. Additional results on baseline comparisons, generalization validation, the explanation of “key” and “query”, and further details of the image transmission experiments are provided in the SM. V.
CONCLUSION
In summary, we have proposed and implemented a TQSC-assisted tomographic characterization scheme for mixed- and pure-state reconstruction under complex scattering, with additional dynamic Gaussian noise for pure states. Through this framework, our approach achieves an average estimator-target fidelity exceeding 99.999% between the reconstructed and target states in the RSP experiments, while simultaneously providing interpretable data-driven insights into measurement correlations. Although the training phase requires a dataset of paired (M, ρtarget ) examples, during the inference stage, the scheme does not require additional quantum operations and performs classical tomographic denoising based on a single acquired measurement vector. Moreover, it demonstrates high single-acquisition efficiency, generalization across the tested conditions, and interpretability, providing a data-driven approach to quantum state estimation in complex noisy environments. While our current implementation demonstrates remarkable performance in a controlled environment, we acknowledge that its scalability and resilience against more complex noise models warrant further investigation.
Future work will focus on integrating this scheme as a post-processing and diagnostic tool in larger, multi-node quantum networks and exploring its compatibility with different quantum information protocols. These continued efforts are expected to further unlock the potential of our approach, paving the way for its broader application in intelligent quantum communication and quantum state characterization.
VI.
ACKNOWLEDGMENTS
This work is supported in part by the National Natural Science Foundation of China (Grants No. 62375164, No. 12574360 and No. 12192252), the Foundation for Shanghai Municipal Science Technology Major Project (Grant No. 2019SHZDZX01-ZX06), Quantum Science and Technology-National Science and Technology Major Project (Grant No. 2021ZD0300802), the Shuguang Program of Shanghai Education Development Foundation and Shanghai Municipal Education Commission (Grant No. 24SG53), and Guangdong Provincial Quantum Science Strategic Initiative (Grant No. GDZX2403003) and Shanghai Oriental Talent Plan Youth Project (Grant no. QNKJ2025013).
∗
These authors contributed equally to this work. [email protected] ‡ [email protected] § [email protected] [1] D. Main, P. Drmota, D. P. Nadlinger, E. M. Ainley, A. Agrawal, B. C. Nichol, R. Srinivas, G. Araneda, and D. M. Lucas, Distributed quantum computing across an optical network link, Nature 638, 383 (2025). [2] M. Pittaluga, Y. S. Lo, A. Brzosko, R. I. Woodward, D. Scalcon, M. S. Winnel, T. Roger, J. F. Dynes, K. A. Owen, S. Juárez, P. Rydlichowski, D. Vicinanza, G. Roberts, and A. J. Shields, Long-distance coherent quantum communications in deployed telecom networks, Nature 640, 911 (2025). [3] P. Bhattacharyya, W. Chen, X. Huang, S. Chatterjee, B. Huang, B. Kobrin, Y. Lyu, T. J. Smart, M. Block, E. Wang, et al., Imaging the meissner effect in hydride superconductors using quantum sensors, Nature 627, 73 (2024). [4] H. J. Kimble, The quantum internet, Nature 453, 1023 (2008). [5] Y. Yang, Y. Li, H. Li, C. Wu, Y. Zheng, and X. Chen, A 300-km fully-connected quantum secure direct communication network, Science Bulletin (2025). [6] S. Wehner, D. Elkouss, and R. Hanson, Quantum internet: A vision for the road ahead, Science 362, eaam9288 (2018). [7] A. K. Pati, Minimum classical bit for remote preparation and measurement of a qubit, Physical Review A 63, 014302 (2000). [8] H.-K. Lo, Classical-communication cost in distributed quantum-information processing: a generalization of †
7 quantum-communication complexity, Physical Review A 62, 012313 (2000). [9] C. H. Bennett, G. Brassard, C. Crépeau, R. Jozsa, A. Peres, and W. K. Wootters, Teleporting an unknown quantum state via dual classical and einstein-podolskyrosen channels, Phys. Rev. Lett. 70, 1895 (1993). [10] S. Pogorzalek, K. Fedorov, M. Xu, A. Parra-Rodriguez, M. Sanz, M. Fischer, E. Xie, K. Inomata, Y. Nakamura, E. Solano, et al., Secure quantum remote state preparation of squeezed microwave states, Nat. Commun. 10, 2604 (2019). [11] X. Peng, X. Zhu, X. Fang, M. Feng, M. Liu, and K. Gao, Experimental implementation of remote state preparation by nuclear magnetic resonance, Phys. Lett. A 306, 271 (2003). [12] N. A. Peters, J. T. Barreiro, M. E. Goggin, T.-C. Wei, and P. G. Kwiat, Remote state preparation: Arbitrary remote control of photon polarization, Phys. Rev. Lett. 94, 150502 (2005). [13] G.-Y. Xiang, J. Li, B. Yu, and G.-C. Guo, Remote preparation of mixed states via noisy entanglement, Phys. Rev. A 72, 012315 (2005). [14] M. Rådmark, M. Wieśniak, M. Żukowski, and M. Bourennane, Experimental multilocation remote state preparation, Phys. Rev. A 88, 032304 (2013). [15] S. Liu, D. Han, N. Wang, Y. Xiang, F. Sun, M. Wang, Z. Qin, Q. Gong, X. Su, and Q. He, Experimental demonstration of remotely creating wigner negativity via quantum steering, Phys. Rev. Lett. 128, 200401 (2022). [16] M. Erhard, M. Krenn, and A. Zeilinger, Advances in high-dimensional quantum entanglement, Nat. Rev. Phys. 2, 365 (2020). [17] M. Erhard, H. Qassim, H. Mand, E. Karimi, and R. W. Boyd, Real-time imaging of spin-to-orbital angular momentum hybrid remote state preparation, Phys. Rev. A 92, 022321 (2015). [18] A. R. Cameron, S. W. Cheng, S. Schwarz, C. Kapahi, D. Sarenac, M. Grabowecky, D. G. Cory, T. Jennewein, D. A. Pushin, and K. J. Resch, Remote state preparation of single-photon orbital-angular-momentum lattices, Phys. Rev. A 104, L051701 (2021). [19] M. Ning, N. Li, H. Zhong, Q. Yuan, L.-e. Zhang, S. Wang, M.-K. Chen, X. Qiu, X. Ren, X. Hu, et al., Highdimensional remote state preparation with metasurfacebased hyper-entangled photon source, Laser Photonics Rev. , e01968 (2025). [20] F.-X. Sun, S.-S. Zheng, Y. Xiao, Q. Gong, Q. He, and K. Xia, Remote generation of magnon schrödinger cat state via magnon-photon entanglement, Phys. Rev. Lett. 127, 087203 (2021). [21] E. Chitambar and G. Gour, Quantum resource theories, Reviews of modern physics 91, 025001 (2019). [22] K. Modi, A. Brodutch, H. Cable, T. Paterek, and V. Vedral, The classical-quantum boundary for correlations: Discord and related measures, Rev. Mod. Phys. 84, 1655 (2012). [23] X. Zhang, M. Luo, Z. Wen, Q. Feng, S. Pang, W. Luo, and X. Zhou, Direct fidelity estimation of quantum states using machine learning, Physical Review Letters 127, 130503 (2021). [24] H. Qin, L. Che, C. Wei, F. Xu, Y. Huang, and T. Xin, Experimental direct quantum fidelity learning via a datadriven approach, Physical Review Letters 132, 190801
(2024). [25] W. H. Zurek, Decoherence, einselection, and the quantum origins of the classical, Rev. Mod. Phys. 75, 715 (2003). [26] H. Defienne, M. Reichert, and J. W. Fleischer, Adaptive quantum optics with spatially entangled photon pairs, Physical review letters 121, 233601 (2018). [27] E. Magesan, J. M. Gambetta, and J. Emerson, Scalable and robust randomized benchmarking of quantum processes, Phys. Rev. Lett. 106, 180504 (2011). [28] R. Blume-Kohout, J. K. Gamble, E. Nielsen, K. Rudinger, J. Mizrahi, K. Fortier, and P. Maunz, Demonstration of qubit operations below a rigorous fault tolerance threshold with gate set tomography, Nat. Commun. 8, 14485 (2017). [29] H. Ma, B. Qi, I. R. Petersen, R.-B. Wu, H. Rabitz, and D. Dong, Machine learning for estimation and control of quantum systems, National Science Review 12, nwaf269 (2025). [30] G. Torlai and R. G. Melko, Machine-learning quantum states in the nisq era, Annual Review of Condensed Matter Physics 11, 325 (2020). [31] A. Strikis, D. Qin, Y. Chen, S. C. Benjamin, and Y. Li, Learning-based quantum error mitigation, PRX Quantum 2, 040330 (2021). [32] Z. Cai, R. Babbush, S. C. Benjamin, S. Endo, W. J. Huggins, Y. Li, J. R. McClean, and T. E. O’Brien, Quantum error mitigation, Reviews of Modern Physics 95, 045005 (2023). [33] H. Liao, D. S. Wang, I. Sitdikov, C. Salcedo, A. Seif, and Z. K. Minev, Machine learning for practical quantum error mitigation, Nature Machine Intelligence 6, 1478 (2024). [34] M. Liao, Y. Zhu, G. Chiribella, and Y. Yang, Noiseagnostic quantum error mitigation with data augmented neural models, npj Quantum Information 11, 8 (2025). [35] Z.-M. Wang and T.-Z. Chen, Adaptive denoising quantum state preparation in a dynamic environment, Physical Review Research 6, 043195 (2024). [36] C.-C. Li, R.-H. He, and Z.-M. Wang, Enhanced quantum state preparation via stochastic predictions of neural networks, Physical Review A 108, 052418 (2023). [37] G. Torlai, G. Mazzola, J. Carrasquilla, M. Troyer, R. Melko, and G. Carleo, Neural-network quantum state tomography, Nat. Phys. 14, 447 (2018). [38] J. Carrasquilla, G. Torlai, R. G. Melko, and L. Aolita, Reconstructing quantum states with generative models, Nature Machine Intelligence 1, 155 (2019). [39] A. M. Palmieri, E. Kovlakov, F. Bianchi, D. Yudin, S. Straupe, J. D. Biamonte, and S. Kulik, Experimental neural network enhanced quantum tomography, npj Quantum Information 6, 20 (2020). [40] S. Ahmed, C. Sanchez Munoz, F. Nori, and A. F. Kockum, Quantum state tomography with conditional generative adversarial networks, Phys. Rev. Lett. 127, 140502 (2021). [41] Y. Hu, M. Ma, and J. Shang, Error-mitigated quantum state tomography using neural networks, arXiv:2602.09733 [quant-ph] (2026). [42] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. K. Kaiser, and I. Polosukhin, Attention is all you need, Adv. Neural Inf. Process. Syst. 30 (2017).
8 [43] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, An image is worth 16x16 words: Transformers for image recognition at scale, in International Conference on Learning Representations (2021). [44] L. Dong, S. Xu, and B. Xu, Speech-transformer: A norecurrence sequence-to-sequence model for speech recognition, Proc. IEEE Int. Conf. Acoust., Speech Signal Process. (ICASSP) 2018, 5884 (2018). [45] H. Zhang, Q. Zhao, M. Zhou, L. Feng, D. Niyato, S. Zheng, and L. Chen, A survey of quantum transformers: Architectures, challenges and outlooks, arXiv preprint arXiv:2504.03192 (2025). [46] J. A. Sobral, M. Perle, and M. S. Scheurer, Physicsinformed transformers for electronic quantum states, Nature Communications (2025). [47] Y.-H. Zhang and M. Di Ventra, Transformer quantum state: A multipurpose model for quantum many-body problems, Physical Review B 107, 075147 (2023). [48] Y. Li, Z. Shi, H. Ma, L. Shen, J. Bao, and Y. Xiao, Language model for large-text transmission in noisy quantum communications, arXiv preprint arXiv:2504.20842 (2025). [49] S. Daimon and Y.-i. Matsushita, Quantum circuit generation for amplitude encoding using a transformer decoder, Physical Review Applied 22, L041001 (2024). [50] J. Bausch, A. W. Senior, F. J. H. Heras, T. Edlich, A. Davies, M. Newman, C. Jones, K. Satzinger, M. Y. Niu, S. Blackwell, G. Holland, D. Kafri, J. Atalaya, C. Gidney, D. Hassabis, S. Boixo, H. Neven, and P. Kohli, Learning high-accuracy error decoding for quantum processors, Nature 635, 834 (2024). [51] P. Cha, P. Ginsparg, F. Wu, J. Carrasquilla, P. L. McMahon, and E.-A. Kim, Attention-based quantum tomography, Machine Learning: Science and Technology 3, 01LT01 (2022). [52] A. M. Palmieri, G. Müller-Rigat, A. K. Srivastava, M. Lewenstein, G. Rajchel-Mieldzioć, and M. Plodzień, Enhancing quantum state tomography via resourceefficient attention-based neural networks, Phys. Rev. Res. 6, 033248 (2024). [53] H. Ma, Z. Sun, D. Dong, C. Chen, and H. Rabitz, Tomography of quantum states from structured measurements via quantum-aware transformer, IEEE Transactions on Cybernetics (2025). [54] M. T. Ribeiro, S. Singh, and C. Guestrin, ” why should i trust you?” explaining the predictions of any classi-
fier, in Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining (2016) pp. 1135–1144. [55] S. Lapuschkin, S. Wäldchen, A. Binder, G. Montavon, W. Samek, and K.-R. Müller, Unmasking clever hans predictors and assessing what machines really learn, Nature communications 10, 1096 (2019). [56] A. J. DeGrave, J. D. Janizek, and S.-I. Lee, Ai for radiographic covid-19 detection selects shortcuts over signal, Nature Machine Intelligence 3, 610 (2021). [57] K. Kraus, A. Böhm, J. D. Dollard, and W. Wootters, States, effects, and operations fundamental notions of quantum theory: Lectures in mathematical physics at the university of Texas at Austin (Springer, 1983). [58] W. Wu, W.-T. Liu, P.-X. Chen, and C.-Z. Li, Deterministic remote preparation of pure and mixed polarization states, Phys. Rev. A 81, 042301 (2010). [59] L.-Y. Yu and S. You, High-fidelity and high-speed wavefront shaping by leveraging complex media, Science Advances 10, eadn2846 (2024). [60] M. W. Matthès, Y. Bromberg, J. De Rosny, and S. M. Popoff, Learning and avoiding disorder in multimode fibers, Physical Review X 11, 021060 (2021). [61] N. Lambert, E. Giguère, P. Menczel, B. Li, P. Hopf, G. Suárez, M. Gali, J. Lishman, R. Gadhvi, R. Agarwal, A. Galicia, N. Shammah, P. Nation, J. Johansson, S. Ahmed, S. Cross, A. Pitchford, and F. Nori, Qutip 5: The quantum toolbox in python, Physics Reports 1153, 1 (2026). [62] D. F. V. James, P. G. Kwiat, W. J. Munro, and A. G. White, Measurement of qubits, Phys. Rev. A 64, 052312 (2001). [63] J. B. Altepeter, E. R. Jeffrey, and P. G. Kwiat, Photonic state tomography, Adv. At. Mol. Opt. Phys. 52, 105 (2005). [64] M. Votto, M. Ljubotina, C. Lancien, J. I. Cirac, P. Zoller, M. Serbyn, L. Piroli, and B. Vermersch, Learning mixed quantum states in large-scale experiments, Physical Review Letters 136, 090801 (2026). [65] G. Jawahar, B. Sagot, and D. Seddah, What does bert learn about the structure of language?, in Proceedings of the 57th annual meeting of the association for computational linguistics (2019) pp. 3651–3657. [66] E. Delp and O. Mitchell, Image compression using block truncation coding, IEEE transactions on Communications 27, 1335 (2003).
Supplementary Material for: Noise-Robust Quantum State Characterization for Remote State Preparation with Deep Learning Bo Tang,1, ∗ Zixuan Liao,1, ∗ Hao Li,1, ∗ Yilin Yang,1 Jiani Lei,1 Zengya Li,1 Jing Qiu,1 Zhaohui Dong,1 Zhengyang Mao,1 Yuanhua Li,2, † Yuanlin Zheng,1, 3, 4, ‡ and Xianfeng Chen1, 3, 4, 5, § 1
State Key Laboratory of Photonics and Communications, School of Physics and Astronomy, Shanghai Jiao Tong University, Shanghai 200240, China 2 Department of Physics, Shanghai Key Laboratory of Materials Protection and Advanced Materials in Electric Power, Shanghai University of Electric Power, Shanghai 200090, China 3 Hefei National Laboratory, Hefei 230088, China 4 Shanghai Research Center for Quantum Sciences, Shanghai 201315, China 5 Collaborative Innovation Center of Light Manipulations and Applications, Shandong Normal University, Jinan 250358, China (Dated: September 18, 2026)
I.
PREPARATION OF MIXED STATE THEORY SCHEME
Bob’s photon onto one of the four mixed states: ⊥ ρ̂IB = p2 |φB ⟩⟨φB | + q 2 |φ⊥ B ⟩⟨φB |,
Consider that the desired mixed state is [1] ⊥ ρB = p2 |φB ⟩⟨φB | + q 2 |φ⊥ B ⟩⟨φB |
(S1)
with iϕ
|φB ⟩ = α|HB ⟩ + βe |VB ⟩,
(S2)
−iϕ |φ⊥ |HB ⟩ − α|VB ⟩, B ⟩ = βe
(S3)
where α, β ∈ C with |α|2 + |β|2 = 1, and p, q are probability amplitudes satisfying p2 + q 2 = 1. Without loss of generality, we assume that p, q are real numbers, p2 + q 2 = 1, and α, β, ϕ are the same as before. To prepare arbitrary mixed states we need to achieve complete control over all five parameters. In Alice’s setup, two polarization controllers, PC1 and PC2, are inserted in the two distinct optical paths to rotate the polarization states as |H⟩ → p|H⟩ + q|V ⟩,
(S4)
|V ⟩ → p|H⟩ + q|V ⟩.
(S5)
To achieve independent polarization control in both paths, two additional polarization controllers, PC1′ and PC2′ , are employed. These are adjusted to implement the transformation
(S8b) ⊥ ρ̂YB = p2 (σ̂y |φB ⟩)(⟨φB |σ̂y ) + q 2 (σ̂y |φ⊥ B ⟩)(⟨φB |σ̂y ), (S8c) 2 2 ⊥ ⊥ ρ̂Z B = p (σ̂z |φB ⟩)(⟨φB |σ̂z ) + q (σ̂z |φB ⟩)(⟨φB |σ̂z ). (S8d)
Bob can obtain the desired mixed state by applying local ˆ σ̂x , σ̂y or σ̂z according to Alice’s unitary operations I, measurement results. The required classical communication is two bits. For simplicity, we do not use PC1′ and PC2′ to adjust the transformation. The final mixed states are as follows:
ρ̂IB = α2 |H⟩⟨H| + β 2 |V ⟩⟨V |,
(S9a)
2 2 ρ̂X B = α (σ̂x |H⟩)(⟨H|σ̂x ) + β (σ̂x |V ⟩)(⟨V |σ̂x ), ρ̂YB = α2 (σ̂y |H⟩)(⟨H|σ̂y ) + β 2 (σ̂y |V ⟩)(⟨V |σ̂y ),
(S9b)
2 2 ρ̂Z B = α (σ̂z |H⟩)(⟨H|σ̂z ) + β (σ̂z |V ⟩)(⟨V |σ̂z ).
(S9d)
|H⟩ → α|H⟩ + βeiϕ |V ⟩,
(S6)
−iϕ
(S7)
|H⟩ − α|V ⟩.
Then POVM measurement described by M1 and M2 are performed on the H-polarized and V-polarized components. Then the Alice’s side HWP (∼22.5◦ ) and the detectors perform the projection measurement , which projects
(S9c)
For pure states, the derivation follows analogously to the mixed-state case (or refer to Ref.[2]), similarly leading to: |ϕ⟩ =
|V ⟩ → βe
(S8a)
2 2 ⊥ ⊥ ρ̂X B = p (σ̂x |φB ⟩)(⟨φB |σ̂x ) + q (σ̂x |φB ⟩)(⟨φB |σ̂x ),
1 {|H2A′ ⟩(α|H B ⟩ + βeiφ |V B ⟩) 2 − |V2A′ ⟩[σ̂z (α|H B ⟩ + βeiφ |V B ⟩)]
(S10)
+ |H1A′ ⟩[σ̂x (α|H B ⟩ + βeiφ |V B ⟩)] − |V1A′ ⟩[iσ̂y (α|H B ⟩ + βeiφ |V B ⟩)]}, In a similar manner, Bob can obtain the desired pure ˆ σ̂x , σ̂y , or σ̂z state by applying local unitary operations I, according to Alice’s measurement results. The required
2 classical communication is two bits. 1 (α|H B ⟩ + βeiφ |V B ⟩), 2 1 X |ψ⟩B = σ̂x (α|H B ⟩ + βeiφ |V B ⟩), 2 1 Y |ψ⟩B = iσ̂y (α|H B ⟩ + βeiφ |V B ⟩), 2 1 Z |ψ⟩B = σ̂z (α|H B ⟩ + βeiφ |V B ⟩). 2 I
|ψ⟩B =
II.
(S11a) (S11b) (S11c) (S11d)
QUANTUM NOISE MODELS
Quantum noise arises from the uncontrollable interaction between a quantum system and its environment, causing the system’s evolution to deviate from ideal unitary dynamics. We employ the framework of quantum channels to mathematically describe such noise. A quantum channel is a completely positive (CP) and trace-nonincreasing (TNI) map that transforms an input density matrix ρ into an output state E(ρ). According to Kraus’ theorem [3], any completely positive trace-preserving (CPTP) map can be expressed as E(ρ) =
X
Ek ρEk† ,
(S12)
k
where {Ek } are the Kraus operators, satisfying P † For tracek Ek Ek = I (trace-preserving condition). P non-increasing maps (lossy channels), k Ek† Ek ≤ I. When multiple noise channels E1 , E2 , . . . , EN act sequentially, the overall effect is described by the composite channel ETotal = EN ◦ EN −1 ◦ · · · ◦ E1 ,
For a single qubit with loss probability p, the Kraus operators are √ 0 p 1 √ 0 , E1 = . (S15) E0 = 0 0 0 1−p Here E0 corresponds to no photon loss, while E1 describes decay from |1⟩ to |0⟩. If Alice and Bob’s photons experience independent losses with probabilities pA and pB , the overall channel is (A) (B) ELoss (ρAB ) = EAD (pA ) ⊗ EAD (pB ) (ρAB ), (S16) (Loss)
with Kraus operators Eij
(A)
= Ei
(B)
⊗Ej
, i, j ∈ {0, 1}.
1.2. Phase Damping Channel (PDC)
Phase damping accounts for decoherence without energy exchange, e.g., random refractive index fluctuations or scattering-induced phase noise. For a single qubit with decoherence probability q, the Kraus operators are 0 0 1 √ 0 √ . , F1 = (S17) F0 = 0 q 0 1−q With decoherence probabilities qA and qB for Alice and Bob, respectively: (A) (B) EDecoherence (ρAB ) = EP D (qA ) ⊗ EP D (qB ) (ρAB ), (Decoherence)
with Kraus operators Fkl
(A)
= Fk
(B)
⊗ Fl
(S18) .
(S13)
whose Kraus operators are given by chained products:
1.3. MMF Induced Noise
Kj = EN,kN EN −1,kN −1 · · · E1,k1 ,
When a single-mode input qubit is transmitted through a segment of MMF, the different spatial modes supported by the MMF experience varying propagation constants and mode-mixing. This can lead to mode dispersion, differential modal attenuation, and effectively, a loss of coherence or entanglement if the output modes are not perfectly demultiplexed or if the measurement setup cannot distinguish them. For a qubit encoded in polarization or time bins, the interaction with multiple spatial modes can scramble the encoded information, acting as a depolarizing-like channel or a more complex unitary scrambling. A simplified model treats the MMF as inducing random unitary transformations or causing effective dephasing and amplitude damping through mode coupling and differential loss, ultimately leading to a mixed state for the output single mode. We model the MMF noise on Bob’s qubit as a generalized depolarizing channel or a channel that effectively
(S14)
where j indexes all possible choices of {k1 , . . . , kN }.
1. Representative Noise Models
In our analysis of remote state preparation (RSP), we focus on four dominant sources of noise: photon loss, decoherence, multimode fiber (MMF) induced noise, and detector imperfections.
1.1. Amplitude Damping Channel (ADC)
Photon loss during transmission (e.g., absorption or scattering in optical fibers or free-space channels) can be modeled by the amplitude damping channel.
3 mixes the input state due to coupling to higher-order modes that are subsequently lost or unmeasured. For simplicity, we model the MMF as a depolarizing channel on Bob’s qubit. The Kraus operators for a depolarizing channel on Bob’s qubit are: q 1 − 3dMMF I (B) , 4 q (B) , H1 = dMMF 4 X q (B) H2 = dMMF , 4 Y q (B) , H3 = dMMF 4 Z H0 =
(S19)
2. Physical Mechanism of Fidelity Degradation in Multi-Mode Fibers and Noise Modeling Analysis
This section provides a physical interpretation of the fidelity degradation observed when introducing MultiMode Fiber (MMF) into the quantum channel and elucidates the role of the noise model employed in this study within this context.
(S20) 1.
Physical Principles of MMF-Induced Degradation
(S21) (S22)
where dMMF is the depolarization probability due to the MMF, and {I, X, Y, Z} are Pauli matrices. Since the MMF noise is specific to Bob’s transmission path, the overall channel for the two-qubit system is: (B) EMMF (ρAB ) = I (A) ⊗ EDepolarizing (dMMF ) (ρAB ),
Multi-mode fibers support the simultaneous propagation of multiple spatial modes when the core radius or the refractive index contrast between the core and cladding is sufficiently large. The approximate number of supported modes, D, is given by: D≈
2π λ
2 Z R
2 1/2 n (r) − n2c r dr,
(S29)
0
(S23) (B) (MMF) = I (A) ⊗ Hm , for m ∈ with Kraus operators Hm {0, 1, 2, 3}.
1.4. Detector Imperfections
Detector efficiency (q) can be absorbed into the loss model, with effective loss p = 1 − q. Spurious detector clicks (dark counts) introduce random errors in measurement outcomes. We model this effect as an effective depolarizing channel acting on Bob’s qubit: q G0 = 1 − 3d4det I, q G1 = ddet 4 X, q G2 = ddet 4 Y, q G3 = ddet 4 Z,
(S24) (S25) (S26) (S27)
where ddet is the depolarization probability representing the impact of dark counts and other detector-related errors, and {I, X, Y, Z} are Pauli matrices. For the two-qubit system, we assume detector imperfections manifest primarily on Bob’s side: (B) EDetector (ρAB ) = I (A) ⊗ EDepolarizing (ddet ) (ρAB ), (S28) (Detector) (B) with Kraus operators Gm = I (A) ⊗ Gm .
where n(r) represents the radial refractive index profile of the fiber core, nc is the cladding refractive index, R is the core radius, and λ is the optical wavelength [4]. The number of propagating modes is roughly proportional to the refractive index contrast and (R/λ)2 . The introduction of MMF into a quantum transmission channel induces several physical effects that degrade the fidelity of the transmitted quantum states, including mode coupling, polarization mode dispersion (PMD), random phase fluctuations, and mode-dependent loss (MDL). Specifically, unintended mode coupling can arise from manufacturing imperfections (e.g., non-circular core geometry, rough core-cladding interfaces, or refractive index variations) and external mechanical perturbations (e.g., stress, micro-bending, or twisting). In our experiments, controlled twisting was applied to the MMF to induce perturbations, thereby enhancing inter-mode coupling. Such perturbations facilitate energy exchange between different spatial modes and lead to the random evolution of the transmitted optical field. Consequently, the initially prepared pure quantum states undergo decoherence during propagation, resulting in a reduction in measured fidelity.
2.
Theoretical Modeling Based on Field Coupling
The inter-mode coupling effects described above can be quantitatively characterized using the coupled-mode theory. In this framework, the total optical field propagating along the longitudinal z-direction is expanded as a superposition of orthogonal eigenmodes. The evolution of the modal amplitudes is governed by the following cou-
4 pled differential equation: dAµ = −jβµ Aµ + dz
X
Cµν (z)Aµ ,
(S30)
ν̸=µ
where Aµ denotes the complex amplitude of the µ-th mode, βµ is its propagation constant, and Cµν (z) represents the coupling coefficient induced by structural perturbations in the fiber [5]. The first term describes the independent phase evolution in the absence of coupling, while the second term accounts for energy transfer between modes. These coupling-induced effects ultimately manifest as crosstalk and decoherence in the transmitted quantum states. 3.
the correct sequence. The entangled state is distributed. Bob’s qubit then goes through the MMF, and finally, both Alice and Bob perform detection. The Kraus operators for Alice’s total noise channel are (k,i) (A) (A) KA = Fk Ei . The Kraus operators for Bob’s total (m,n,l,j) (B) (B) (B) (B) noise channel are KB = Gm Hn Fl Ej . Then, the noisy state is given by: X (k,i) (m,n,l,j) ρnoisy = K ⊗ K ρin AB AB A B k,i m,n,l,j
3. Composite Noise Model
In RSP, the shared entangled state first undergoes transmission, which includes loss, decoherence, and MMF noise on Bob’s side, followed by Alice’s measurement and Bob’s conditional unitary. We model the total noise channel as a sequential application of the individual noise channels. The composite Kraus operator for the total noise will be a product of the individual Kraus operators, applied in
(k,i)† (m,n,l,j)† KA ⊗ KB
.
+ + Here, ρin AB = |Φ ⟩ ⟨Φ | is the initial Bell state.
Discussion on Noise Suppression and Model Validity
Experimentally, in the absence of MMF (i.e., a direct quantum channel), the average fidelity of the prepared states reaches 99.25% ± 0.22%. As shown in the main text, the introduction of MMF causes a significant drop in fidelity when no noise suppression model is applied. This degradation reflects the raw impact of multi-mode perturbations prior to correction. The optimized TQSC (Transformer-based Quantum State Characterizer) noise model used in this work is trained on measurement data from the quantum channel incorporating the MMF. Consequently, the TQSC model inherently accounts for the effective noise introduced by the MMF, including contributions from mode coupling, phase perturbations, and mode-dependent loss. Combined with our Transformer-based post-processing reconstruction, the model recovers the experimental density matrices with high fidelity, aligning them closely with the target states. While the current framework achieves high reconstruction accuracy by treating the channel as a complex noise process, future work could benefit from incorporating more explicit physical constraints. Specifically, independently modeling the distinct noise contributions-such as the statistical characterization of coupling coefficients Cµν (z), phase noise spectra, and polarization effects-and integrating them into the noise prior or loss function could further refine the post-processing inversion.
(S31)
4. Remote State Preparation under Noise
Alice performs a Bell-state measurement (BSM) on her (A) subsystem. For outcome k ′ , with projector Πk′ , the probability is h i (A) P (k ′ ) = Tr (Πk′ ⊗ I (B) )ρnoisy , (S32) AB and the post-measurement state of Bob’s qubit is i h (A) (A) TrA (Πk′ ⊗ I (B) )ρnoisy (Πk′ ⊗ I (B) ) ′ AB (k ) . (S33) ρB = P (k ′ ) Depending on Alice’s reported result k ′ , Bob applies a correction unitary Uk′ ∈ {I, X, Y, Z}, leading to the final ensemble state: X (k′ ) ρfinal = P (k ′ ) Uk′ ρB Uk†′ . (S34) B k′
5. Performance Metric: Fidelity
The performance of RSP is quantified by the fidelity between Bob’s final state and the target state |ψ⟩: F = ⟨ψ|ρfinal B |ψ⟩ .
(S35)
This allows us to systematically evaluate the impact of loss, decoherence, MMF induced noise, and detector imperfections on RSP. III.
THEORETICAL FRAMEWORK OF THE TQSC MODEL
A.
Problem Formulation: Learning an Inverse Quantum-to-Classical Map
The experimental realization of RSP is inevitably influenced by a multitude of physical imperfections. The entire process can be conceptualized as a complex quantumto-classical channel, denoted by C, which maps an ideal
5 target quantum state ρtarget to a set of classical measurement outcomes. In our work, each data record of these outcomes is tokenized into a 13-element feature vector M ∈ R13 . This vector acts as a comprehensive, noisecorrupted classical “snapshot” of the quantum state: M = C(ρtarget ).
(S36)
To intuitively understand this formulation, M aggregates the multidimensional profile of the state observed from 4 distinct perspectives. Specifically, it consists of direct physical observables: the coincidence counts between Alice and Bob, the corresponding single-photon counts for each party, all measured across four different measurement bases, alongside the total coincidence measurement time. This specific design ensures that the vector is informationally complete, encapsulating the full spectrum of data required for quantum state tomography. Through this numerical encoding, we map raw physical variables into a structured feature space, enabling our TQSC Transformer model to process quantum data with the same architectural logic used for word embeddings in natural language processing. The map C encapsulates both coherent errors, such as misalignments in optical components, and incoherent noise processes, including channel decoherence, detector dark counts, and Poissonian shot noise. Conventional quantum state tomography (QST) techniques attempt to reconstruct the state by inverting this map from the data M. However, they often falter when faced with complex, non-Gaussian, or correlated noise structures inherent in real-world systems. To overcome this limitation, we propose a data-driven TQSC model. The TQSC is a deep neural network, parameterized by a set of weights θ, that learns an effective inverse map ΦTQSC (·; θ). This learned function directly transforms the noisy classical measurement data M into a highfidelity estimate of the density matrix, ρpred : ρpred = ΦTQSC (M; θ).
(S37)
The core objective is to train the TQSC such that the learned map ΦTQSC not only inverts the deterministic dynamics of the channel but also actively suppresses the stochastic noise components.
B.
formation to the Transformer, we apply a Sinusoidal Positional Encoding Ppos ∈ RL×dmodel . This encoding is generated by fixed sine and cosine functions of different frequencies, where each position pos in the sequence is mapped to a unique vector. The initial representation H(0) ∈ RL×dmodel is formed as:
Model Architecture and Mathematical Formalism
The architecture of the TQSC is meticulously designed to leverage the powerful sequence-processing capabilities of the Transformer model [6], adapting it to the structured nature of our physics-based input vector. Input Embedding with Encoding The input vector M ∈ R13 is interpreted as a sequence of L = 13 tokens, S = {m1 , m2 , . . . , mL }. To provide positional in-
H(0) = Linear(S) + Ppos .
(S38)
This injects absolute positional information but does not incorporate explicit knowledge of physical feature types. Multi-Head Self-Attention Core The heart of the TQSC is a stack of N = 3 identical encoder layers. The central mechanism in each layer is Multi-Head SelfAttention (MHSA) with h = 4 heads, which allows the model to weigh the importance of all input tokens relative to each other. Each head independently learns different subspace features of the input, such as photon counting patterns, noise correlations, and temporal correlations in quantum signals. For an input representation X ≜ H(l−1) ∈ Rn×d at layer l (where n is the sequence length, d the feature dimension), the MHSA output is computed as: MHSA(X) = Concat(head1 , . . . , headh )WO ,
(S39)
where WO ∈ Rhdv ×d is the output projection matrix, and each head is defined by: headi = Attention(Qi , Ki , Vi ).
(S40)
Here the query, key, and value matrices for head i are computed as: Qi = XWiQ ,
WiQ ∈ Rd×dk
(S41)
Ki = XWiK , Vi = XWiV ,
WiK ∈ Rd×dk WiV ∈ Rd×dv
(S42) (S43)
with dk = dv = d/h. The attention function is then given by: Attention(Qi , Ki , Vi ) = softmax
Qi K⊤ √ i dk
Vi . (S44)
This attention mechanism enables the model to identify and exploit complex, non-linear correlations within the measurement data—for instance, the relationship between single-photon counts and coincidence events, which is a key indicator of channel loss and noise. The output of the attention block is then processed through a positionwise Feed-Forward Network (FFN), with residual connections and layer normalization applied after each sub-layer to ensure stable training. State Vector Prediction Head Following the final encoder layer, we obtain the output representation H(N ) ∈ RL×dmodel . To aggregate the information from the entire
6 sequence into a single vector, we utilize the fixed positional encoding scheme inherent in the Transformer architecture. Specifically, the sinusoidal positional encoding Ppos provides absolute position information for each element in the sequence. The aggregated representation hagg ∈ Rdmodel is computed as: L X
1 (N ) h hagg = L i=1 i
ρout =
(S46)
For the single-qubit case considered in this work, d = 2 and the network outputs four real parameters r = (r0 , r1 , r2 , r3 ). We define s0 0 C= , r2 + ir3 s1 (S47) sj = softplus(rj ) + ϵ,
A brief non-technical explanation of Key and Query.
Keys and queries are fundamental components of the attention mechanism widely used in modern neural networks. Intuitively, a query represents the question or probe used to search for relevant information, while a key is a descriptor associated with each candidate item. The value stores the actual information to be retrieved. In simple terms, the query asks “what am I looking for?”, the keys label the available items, and the values contain the corresponding content. Items whose keys better match the query receive higher weights and thus contribute more strongly to the output representation. A useful analogy is searching for a book in a library. The query corresponds to what one is searching for, for example, “an introductory book on machine learning.” Each book in the library is associated with a key that summarizes its main characteristics, such as subject, difficulty level, or keywords. The value represents the actual content of the book. The attention mechanism compares the query with the keys of all books and assigns higher weights to those whose keys better match the query. As a result, books that are more relevant to the search request contribute more to the final output, while irrelevant books receive negligible weights. From this perspective, the attention mechanism can be viewed as a soft, learnable retrieval process, rather than a hard selection of a single item.
D.
CC † . Tr(CC † )
(S45)
where L is the sequence length. This mean-pooling strategy leverages the position-aware representations created by the positional encoding module.
C.
in maximum-likelihood quantum-state tomography. Instead of directly predicting the entries of ρ, the network 2 outputs an unconstrained real vector r ∈ Rd , which is used to construct a lower-triangular complex matrix C. The output density matrix is then reconstructed as
Physically-Constrained Output Layer
A critical challenge is ensuring that the neural-network output corresponds to a physically valid density matrix. A density matrix ρ must satisfy three constraints: (i) Hermiticity, ρ = ρ† , (ii) unit trace, Tr(ρ) = 1, and (iii) positive semidefiniteness, ρ ≥ 0. To enforce these constraints by construction, we adopt a Cholesky-style parametrization, as commonly used
j = 0, 1,
where ϵ > 0 is a small constant used for numerical stability. This choice ensures that Tr(CC † ) > 0 for every finite network output. This parametrization guarantees Hermiticity because (CC † )† = CC † .
(S48)
It also guarantees unit trace by construction: Tr(ρout ) =
Tr(CC † ) = 1. Tr(CC † )
(S49)
Finally, for any complex vector v, we have v † ρout v =
v † CC † v ∥C † v∥2 = ≥ 0. † Tr(CC ) Tr(CC † )
(S50)
Therefore, ρout is positive semidefinite for any real-valued network output r. For d = 2, the reconstructed density matrix can be written explicitly as ρout =
1 s20 + s21 + r22 + r32 ! s0 (r2 − ir3 ) s20
·
s0 (r2 + ir3 ) s21 + r22 + r32
(S51) .
Thus, the physically constrained output layer maps every unconstrained real-valued output of the Transformer backbone to a bona fide density matrix, while preserving end-to-end differentiability.
E.
Optimization and Theoretical Merit
The network parameters θ are optimized by minimizing the Mean Squared Error (MSE) loss function, defined as the squared Frobenius norm between the predicted state ρpred and the known ideal target state ρtarget . Over a
7 training dataset of K samples {(Mk , ρtarget,k )}K k=1 , the loss is: K
L(θ) =
1 X 2 ∥ΦTQSC (Mk ; θ) − ρtarget,k ∥F . K
(S52)
k=1
Optimization is performed using the Adam optimizer, which implements a variant of stochastic gradient descent. From a theoretical standpoint, the TQSC framework offers a significant advantage. The Transformer architecture is a universal function approximator, guaranteeing its capacity to model the highly complex inverse map ΦTQSC ≈ C −1 . More pointedly, the self-attention mechanism functions as a learned, context-aware adaptive filter. It can dynamically identify and down-weight measurement values that are inconsistent with the context provided by the rest of the data, effectively learning to suppress noise-induced anomalies. Therefore, the minimization of the MSE loss in Eq. (S52) does not merely fit a curve; it drives the model to learn the underlying statistical structure of the noise and systematically correct for it. This establishes a robust framework where deep learning directly enhances quantum state fidelity by learning noise-resilient inversion of physical measurements. The TQSC framework derives its theoretical strength from two fundamental properties of the Transformer architecture. First, as a universal function approximator, the Transformer possesses the inherent capacity to model arbitrarily complex inverse mappings with any desired precision. This mathematical guarantee ensures that TQSC can in principle learn the exact inverse transformation from measurement data to quantum states. Second, the self-attention mechanism provides a sophisticated context-aware processing capability. By dynamically evaluating the consistency of each measurement value against the full context of the dataset, it learns to identify and suppress noise-induced anomalies while preserving genuine quantum signatures. This adaptive filtering operates as a learned denoising function that transcends traditional signal processing approaches. Consequently, minimizing the mean squared error loss accomplishes more than simple data fitting. It drives the model to internalize the statistical patterns of measurement noise and implement physically meaningful corrections. This creates a principled framework where deep learning directly enhances the accuracy of quantum state reconstruction through intelligent noise-resilient inversion of experimental data. IV.
EXPERIMENTAL SETUP
The core of the system is a Sagnac interferometric loop housing a spontaneous parametric down-conversion (SPDC) source, designed for generating polarizationentangled photon pairs with inherent phase stability. A
femtosecond laser operating at a repetition rate of 60 MHz undergoes second-harmonic generation to produce 775 nm pump light. This pump beam is injected into the polarization-entangled Sagnac loop, where SPDC occurs in a periodically poled lithium niobate (PPLN) crystal, generating degenerate photon pairs at 1550 nm wavelength. The half-wave plate (HWP) and quarterwave plate (QWP) within the Sagnac loop were initially aligned with their fast axes at 0° relative to the horizontal polarization axis to ensure maximal generation of the target Bell state. Following generation, wavelength-division multiplexing (WDM) filters select photons within the approximately 1550 nm band. These photons are then routed through specific dense wavelength-division multiplexing (DWDM) channels: channel 31 (1552.54 nm) directs photons to Alice, and channel 33 (1550.92 nm) directs photons to Bob. In Alice’s detection module, the incoming photon is first split by a 9:1 beam splitter (BS1). One output arm incorporates a tunable optical delay line for precise temporal synchronization before recombination at the second beam splitter (BS2). The alternate output path from BS1 contains a variable optical attenuator used to adjust the intensity ratio ‘η = α2 /β 2 ’ between the two measurement bases. Both arms subsequently pass through fiber-based polarizing beam splitters (FPBS) for projection onto the horizontal (H) or vertical (V) polarization basis. For the critical state analysis step, a hybrid detection scheme is implemented by combining the H-polarized output from the upper FPBS and the V-polarized output from the lower FPBS through BS2. Polarization controllers (PC) in both arms compensate for polarization rotation induced by the optical fibers. Following BS2, projective measurements are performed using a 22.5° half-wave plate (HWP) and polarizing beam splitters (PBS) before coupling into fiber-coupled singlemode collection paths. Bob’s receiver module employs a polarization controller (PC) for initial compensation of fiber-induced birefringence. A commercially available OM2 gradedindex MMF (core diameter 50 µm, cladding diameter 125 µm, length 1.5 m) is used in the experiment. To simulate the spatial mode scrambling encountered in complex scattering environments, the photon is transmitted through this fiber section. Full quantum state tomography is implemented using a polarization analyzer consisting of a quarter-wave plate (QWP), a half-wave plate (HWP), and a polarizing beam splitter (PBS) in sequence, with the outputs coupled via single-mode fibers to detectors. Both Alice and Bob employ superconducting nanowire single-photon detectors (SNSPDs) for photon detection, characterized by a quantum efficiency of approximately 80% at the operating wavelength of 1550 nm. Detection events are recorded and time-tagged using time-
8 correlated single-photon counting (TCSPC) electronics offering a timing resolution of 81 ps. The coincidence window was set to 500 ps, optimized relative to the 16.7 ns repetition period of the femtosecond laser source.
V.
CHARACTERIZATION OF THE SPDC SOURCE
To systematically quantify the performance of the Spontaneous Parametric Down-Conversion (SPDC) source, we provide a comprehensive experimental characterization covering its efficiency and other key experimental parameters. The TQSC model, trained on experimental data, implicitly captures SPDC failure events such as vacuum and multi-pair contributions, learning to suppress noise and restore fidelity without requiring an analytic noise model. The characterization results and analysis are organized as follows: Firstly, the spectral properties and normalized conversion efficiency of the PPLN waveguide are investigated to clarify the phase-matching bandwidth and the generation efficiency of the source. Secondly, the signal-to-noise performance is evaluated by measuring the coincidenceto-accidental ratio (CAR) as a function of pump power. Thirdly, the single-photon purity and suppression of multi-pair emissions are verified using the second-order correlation function g (2) (0). Fourthly, the high quality of the generated states is demonstrated through polarization entanglement characterization. Finally, regarding SPDC failure events, the TQSC model implicitly learns to suppress the resulting noise and restore fidelity, without the need for an explicit analytic model of the imperfections. Detailed discussions and measurement results for each of these parameters are presented below.
A.
Spectral Properties and Normalized Conversion Efficiency
To characterize the phase-matching bandwidth and generation efficiency of the PPLN waveguide, we investigated the spectral distribution of the SPDC photons and compared it with the theoretical model. Fig. S1 presents the measured normalized photon counts for 10 symmetric signal-idler channel pairs centered at ITU Channel 32 (CH32), ranging from the outermost pair (CH22 & CH42) to the innermost pair (CH31 & CH33). In the experiment, the pump wavelength was tuned to generate degenerate photon pairs centered at CH32. The experimental data (colored bars) show excellent agreement with the theoretical phase-matching curve (red solid line), confirming the reliability of the waveguide design. Based on the measured spectral distribution, we calculated a Full-Width at Half-Maximum
(FWHM) bandwidth of 60.74 nm. This broad phasematching bandwidth ensures a relatively flat and high normalized conversion efficiency across a wide range of channels. This characteristic verifies the uniformity of photon generation and confirms that the source is wellsuited for multi-channel quantum communication protocols requiring high spectral purity and broad wavelength tunability.
B. Source Noise Performance and Coincidence-to-Accidental Ratio (CAR)
To evaluate the signal-to-noise ratio performance of the source, we measured the Coincidence-to-Accidental Ratio (CAR) as a function of pump power for the 10 symmetric channel pairs centered at CH32 (spanning from CH22/CH42 to CH31/CH33), as summarized in Fig. S16. As shown, the CAR values for all channel pairs exhibit a consistent dependence on pump power, strictly following the theoretical inverse relationship CAR ∝ 1/Ppump . This behavior indicates that accidental coincidence counts are dominated by multi-pair emissions rather than background noise or dark counts, confirming the low-noise nature of the experimental setup. Quantitatively, at low pump powers (in the µW regime), the maximum CAR values for all measured channel pairs exceed 20,000, with the central channel pair (CH30 and CH34) reaching a peak of 30,446. Even at higher pump powers, where the statistical significance of multi-photon events increases, the fitted curves maintain excellent agreement with the experimental data across the entire spectral range. These results demonstrate that the device maintains high entanglement purity and a high signal-to-noise ratio over a broad bandwidth, validating the robust performance of the source for multi-channel quantum communication applications.
C.
Single-Photon Purity and Second-Order Correlation Function g (2) (0)
To verify the single-photon purity and the suppression of multi-pair emissions, we measured the zero-delay second-order correlation function, g (2) (0). The results for the 10 symmetric channel pairs (from CH22/CH42 to CH31/CH33) are summarized in Fig. S2. To ensure statistical reliability, each value represents the mean of three independent repeated measurements, with error bars indicating the standard deviation. As shown in Fig. S2, all channel pairs exhibit consistently low g (2) (0) values, ranging from approximately 0.012 to 0.019. These values are well below the classical limit (g (2) (0) ≥ 1) and significantly lower than the typical benchmark for high-quality single-photon sources (g (2) (0) < 0.1), clearly demonstrating strong