ConceptioArchivearXiv CS
arXiv CSopen access

Recursive Quantum Long Short-Term Memory for Stable Short-Horizon Temperature Forecasting

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Recursive Quantum Long Short-Term Memory for Stable Short-Horizon Temperature Forecasting Mu-En Lee∗ , Yen-Ku Liu† , Samuel Yen-Chi Chen‡ , Yun-Cheng Tsai§ ∗ University of Toronto, Toronto, ON, Canada † National Yang Ming Chiao Tung University, Hsinchu, Taiwan ‡ Brookhaven National Laboratory, Upton, NY, USA § PecuLab LLC, Seattle, WA, USA

arXiv:2609.20594v1 [cs.LG] 17 Sep 2026

Yun-Cheng Tsai is the corresponding author: [email protected]

Abstract—Quantum long short-term memory (QLSTM) models extend recurrent sequence learning with variational quantum circuits, but their optimization behavior can vary substantially across random initializations and temporal contexts. This paper evaluates a recursive QLSTM architecture against a standard QLSTM for one-step-ahead prediction of daily minimum and maximum temperature. Using daily weather observations from Toronto and identical training settings, we compare convergence, predictive accuracy, and generalization across input windows of 8, 16, and 32 days over 20 random seeds. The recursive model consistently reaches a near-optimal test loss earlier, reduces mean absolute error and root mean squared error, and exhibits a smaller generalization gap. These results indicate that recursive quantum feature transformations can improve stability and outof-sample performance for compact hybrid quantum–classical temporal models. Index Terms—quantum machine learning, QLSTM, recursive QLSTM, weather forecasting, variational quantum circuits, timeseries prediction

I. I NTRODUCTION Long short-term memory (LSTM) networks were designed to model long-range temporal dependencies while mitigating vanishing-gradient behavior in recurrent learning [1]. Their gated structure has made them a standard baseline for sequential prediction tasks, including meteorological forecasting. Recent quantum machine-learning research has explored hybrid architectures in which trainable variational quantum circuits (VQCs) replace or augment classical transformations [2], [3]. Because VQC optimization can be affected by expressibility, initialization, and barren-plateau behavior [4], [5], empirical comparisons should report both accuracy and training stability. One such model, the quantum long short-term memory (QLSTM), embeds quantum neural-network modules in an LSTMlike recurrent structure to learn temporal data using shallow, NISQ-compatible circuits [6]. Although QLSTM models can be expressive, their behavior is shaped by circuit initialization, circuit depth, and the amount of temporal context presented to the model. These factors can make both optimization and generalization sensitive to the selected configuration. Recent QLSTM variants have also explored distributed, federated, and fast-weightstyle parameter-generation mechanisms for temporal quantum models [7]–[9]. Related quantum sequential-modeling studies have examined shallow quantum temporal embeddings,

financial decision systems using LSTM forecasting signals, batch-size/runtime tradeoffs in QLSTM-style training, and sequence-length sensitivity in urban telecommunication forecasting [10]–[13]. Recursive QLSTM introduces a metacorebased recursive construction intended to improve temporal information propagation while retaining a compact hybrid architecture [14]. This study provides an empirical comparison of standard QLSTM and Recursive QLSTM for short-horizon temperature prediction. Our contributions includes: First, evaluating both architectures under identical data, optimization, and quantum-circuit settings; Second, assessing not only MAE and RMSE but also convergence behavior and the generalization gap; Third, reporting mean and standard deviation over 20 random seeds, separating performance trends from single-run randomness. II. BACKGROUND AND E XPERIMENTAL D ESIGN A. QLSTM and Recursive QLSTM A conventional LSTM updates its cell state and hidden state through gated transformations of the current input xt and prior hidden state ht−1 [1]. In QLSTM, selected affine transformations are replaced by VQCs that encode projected features, apply parameterized rotations and entangling operations, and return expectation values to the classical recurrent computation [6]. This hybrid design can represent nonlinear transformations with a small quantum circuit while leaving training compatible with gradient-based optimization. Recursive QLSTM augments this design by applying a recursive metacore transformation within the recurrent update. Rather than treating the quantum feature map as a single isolated transformation, the recursive formulation reuses intermediate representations to improve information propagation across the sequence [14]. In this paper, we use the MetaCore_Single_NN configuration for the recursive model and hold the remaining architecture settings fixed. See [14] for underlying model and circuit structure. B. Data and Prediction Task We use daily observations from a Toronto weather station for 1 May 2024 through 30 April 2026. Environment and Climate Change Canada provides historical daily weather observations, including temperature variables and degree-day

TABLE I I NPUT AND OUTPUT VARIABLES . Role

Variable

Unit/scale

Input Input Input Input Output

Day-of-year cosine; day-of-year sine Year number Mean, minimum, and maximum temperature Heating and cooling degree days Next-day minimum and maximum temperature

[−1, 1] Count ◦C ◦C ◦C

v u n X 2 u 1 X 2 RMSE = t (yij − ŷij ) . 2n i=1 j=1

For target-wise analysis, the same metrics are computed separately for minimum and maximum temperature. We also report the 95th-percentile absolute error (P95 AE), coefficient of determination (R2 ), and absolute bias. Absolute bias is defined as the absolute value of the mean signed prediction error across the held-out test set. For each seed, predictive metrics are calculated at the epoch with the minimum test loss for that run. For each seed, t95 is measured relative to that run’s initial and best test losses. This definition is preferable to a fixed loss threshold because the two architectures and their random initializations can begin at different loss levels. All reported error bars represent the standard deviation across the 20 independent runs.

TABLE II S HARED EXPERIMENTAL CONFIGURATION . Parameter

Value

QNN depth; hidden size; input projection size; number of qubit Window length L Prediction horizon; input/output dimensions Epochs; batch size; learning rate Random seeds Recursive metacore

1; 3; 3; 6 {8, 16, 32} 1; 8/2 60; 8; 0.0005 0–19 (20 trials) MetaCore_Single_NN

Data Input Toronto City Weather Stations Data (8-dimensional Feature → 3D Feature Spac)

measures, through its public climate-data service [15]. The chronological dataset is split into 80% training observations and 20% held-out test observations. The task is one-step-ahead prediction of daily minimum and maximum temperature. Table I summarizes the model variables. Calendar seasonality is represented by sine and cosine encodings of day of year. All input features are continuous and are scaled using trainingset statistics only, preventing test-set information from entering preprocessing.

Select Window Length L ∈ {8, 16, 32}

Figure 1 illustrates the overall experimental workflow. The models receive a temporal window L ∈ {8, 16, 32} and predict the following day. Each configuration is trained for 60 epochs using a batch size of 8 and learning rate 5 × 10−4 . To characterize stochastic training behavior, each setting is repeated for 20 seeds. Core settings are listed in Table II. We report mean absolute error (MAE), root mean squared error (RMSE), and the generalization gap, defined as test loss minus training loss. A smaller generalization gap indicates a smaller train–test discrepancy. We additionally report the epoch of the minimum test loss (Best Epoch) and the first epoch that reaches 95% of each run’s total improvement (t95 ). The latter gives a practical indicator of how quickly a run obtains a near-optimal test result. For a test set of n daily forecasts with two temperature targets, let yij and ŷij denote the observed and predicted values for forecast i and target j, where j ∈ {1, 2} corresponds to next-day minimum and maximum temperature. Aggregate MAE and RMSE are computed over both targets and all heldout test observations as n

2

(1)

Iterate Window Length

Training & Evaluation Iterate across 20 Random Seeds Models: QLSTM, Recursive QLSTM

Analysis I: Convergence Epoch to Best Loss & PR 95 Test Loss

C. Protocol and Metrics

1 XX MAE = |yij − ŷij | . 2n i=1 j=1

(2)

Analysis II: Performance Predictive Metrics (RMSE & MAE)

Analysis III: Robustness Model Overfitting (Generalization Gap)

Fig. 1. Experimental workflow for comparative study across QLSTM and Recursive QLSTM.

III. R ESULTS A. Convergence Behavior Figure 2 compares the epoch of the minimum test loss and t95 . The two models reached their minimum test loss at broadly similar epochs across the three window lengths. Recursive QLSTM reached its best test loss slightly earlier at L = 8, slightly later at L = 16, and at a comparable epoch at L = 32. Thus, the best-epoch statistic does not indicate a consistent timing advantage for either architecture. In contrast, Recursive QLSTM reaches 95% of its eventual test-loss improvement in fewer epochs across all three window lengths. This pattern suggests that the recursive architecture reaches a useful near-optimal performance range earlier, even when the final epoch of minimum test loss occurs at a similar point in training. At L = 16 and L = 32, Recursive QLSTM also shows narrower error bars, indicating more consistent convergence across seeds.

Fig. 2. Best-epoch and t95 convergence statistics for QLSTM and Recursive QLSTM across window lengths of 8, 16, and 32. Error bars denote the standard deviation across 20 random seeds.

B. Predictive Accuracy Figure 3 reports aggregate MAE and RMSE on the heldout test set, calculated at the minimum-test-loss epoch of each seed. Recursive QLSTM achieves MAE below 3.9◦ C for all window lengths, compared with approximately 4.1–4.2◦ C for QLSTM. The RMSE results show the same direction of effect, indicating that the recursive architecture reduces both average prediction error and the influence of larger residuals. Recursive QLSTM remains favorable at every tested context length, with smaller cross-seed deviations in Figure 3 and Table III.

Fig. 4. Target-wise held-out test-error comparison for next-day maximum and minimum temperature. Metrics are evaluated at the minimum-test-loss epoch of each seed. TABLE III B EST- EPOCH RESIDUAL ROBUSTNESS ON THE HELD - OUT TEST SET. VALUES ARE MEAN ± STANDARD DEVIATION ACROSS 20 SEEDS . L

Model

MAE

RMSE

P95 AE

R2

8 QLSTM 4.08 ± 0.24 5.15 ± 0.31 10.25 ± 0.88 0.625 ± 0.046 8 Rec. QLSTM 3.73 ± 0.18 4.66 ± 0.21 8.97 ± 0.67 0.69 ± 0.028 16 QLSTM 4.14 ± 0.31 5.22 ± 0.38 10.38 ± 1.04 0.616 ± 0.059 16 Rec. QLSTM 3.79 ± 0.20 4.74 ± 0.24 9.07 ± 0.71 0.685 ± 0.032 32 QLSTM 4.17 ± 0.31 5.25 ± 0.39 10.38 ± 1.04 0.617 ± 0.060 32 Rec. QLSTM 3.84 ± 0.21 4.80 ± 0.26 9.22 ± 0.74 0.681 ± 0.034 All error metrics are in ◦ C.

Fig. 3. Aggregate held-out test MAE and RMSE for QLSTM and Recursive QLSTM across window lengths of 8, 16, and 32. Metrics are evaluated at the minimum-test-loss epoch of each seed and averaged across 20 random seeds.

Figure 4 shows that the improvement holds for both maximum and minimum temperature. Recursive QLSTM has lower MAE, RMSE, P95 absolute error, and absolute bias at every window length. Table III confirms lower mean and tail errors, higher R2 , and smaller cross-seed variation for Recursive QLSTM.

QLSTM reaches a near-optimal solution earlier, as reflected by a smaller t95 across the tested window lengths. However, the epoch of the minimum test loss is broadly comparable between the two models and does not show a consistent advantage for either architecture. Reporting both measures therefore helps distinguish rapid early improvement from the point at which the best observed test loss occurs. The contribution of this study is not only lower average forecasting error. Recursive QLSTM also shows a more con-

C. Generalization Gap Figure 5 shows a consistently smaller train–test loss gap for Recursive QLSTM at both the best and final epochs, with reductions of roughly one-quarter to one-third. Its lower standard deviations further indicate more repeatable out-ofsample behavior across seeds. IV. D ISCUSSION The results distinguish between early practical improvement and the timing of the single lowest test-loss epoch. Recursive

Fig. 5. Train–test generalization-gap comparison for QLSTM and Recursive QLSTM across window lengths of 8, 16, and 32. Results are shown at the minimum-test-loss epoch and at the final training epoch. Error bars denote the standard deviation across 20 random seeds.

sistent training pattern. Under the same set of random seeds, it reaches a useful performance range earlier, reduces both average and larger residual errors, and varies less across runs. This matters in small VQC-based experiments, where results can be affected by initialization. Better mean performance together with lower variation gives more confidence that the recursive design is contributing beyond a favorable single run. The weather-station task also clarifies when the model is most appropriate. Daily minimum and maximum temperature are not arbitrary time-series signals: they contain strong annual seasonality, short-term persistence, and local departures caused by changing weather conditions. The input variables used here combine calendar encodings with recent temperature and degree-day information, so the model is being asked to refine a short-horizon forecast from a compact but physically meaningful context. Recursive QLSTM is most suitable for this kind of setting: one-step or short-horizon prediction, limited input dimensionality, and data where recent temporal context is informative but longer windows may introduce redundant or noisy seasonal information. The target-wise results further suggest a useful operational characteristic. Recursive QLSTM improves both maximum and minimum temperature forecasts, including the 95thpercentile absolute error. For weather applications, this matters because a model with similar mean error but fewer large residuals is preferable for decision support: missed cold nights or unusually warm days can matter more than small average deviations. The lower generalization gap indicates that the recursive mechanism is not merely fitting the training portion more aggressively; in this dataset it transfers the learned shortterm structure more consistently to held-out dates. These findings should be interpreted cautiously. Our experiments use a single urban weather station, one forecasting horizon, shallow simulated quantum circuits, and fixed hyperparameters. Recursive QLSTM appears stable for this compact forecasting setting, but the results do not demonstrate a general quantum advantage. This is in line with recent studies showing that quantum sequential models remain sensitive to task characteristics and architectural choices [12], [13]. Future work should test more stations, longer horizons, noisy and missing-data settings, broader hyperparameter ranges, and stronger classical baselines. NARMA-style benchmarks may also help assess whether the observed benefits extend beyond weather forecasting [16]. V. C ONCLUSION We compared QLSTM and Recursive QLSTM for onestep temperature forecasting across three window lengths and 20 random seeds. Recursive QLSTM reached a useful nearoptimal test-performance range earlier, achieved lower MAE and RMSE, reduced larger residual errors, and showed less variation across runs. It also exhibited a consistently smaller train–test generalization gap, indicating more repeatable outof-sample behavior within this controlled setting. The epoch of the minimum test loss was broadly comparable between the two architectures, suggesting that the principal

convergence advantage of Recursive QLSTM lies in faster early improvement rather than a consistently earlier final optimum. These findings suggest that recursive quantum feature transformations may improve training stability and shorthorizon forecasting performance in compact hybrid recurrent models. However, the study is limited to simulated circuits, a single weather station, one prediction horizon, and fixed hyperparameters. The results therefore do not establish a broader quantum advantage or guarantee performance gains across other datasets and tasks. Future work should evaluate additional stations, longer forecasting horizons, noisy and missing-data settings, stronger classical baselines, broader hyperparameter ranges, and noisy quantum hardware conditions. R EFERENCES [1] S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation, vol. 9, no. 8, pp. 1735–1780, 1997. [2] J. Biamonte, P. Wittek, N. Pancotti, P. Rebentrost, N. Wiebe, and S. Lloyd, “Quantum machine learning,” Nature, vol. 549, no. 7671, pp. 195–202, 2017. [3] M. Schuld and N. Killoran, “Quantum machine learning in feature hilbert spaces,” Physical Review Letters, vol. 122, no. 4, p. 040504, 2019. [4] J. R. McClean, S. Boixo, V. N. Smelyanskiy, R. Babbush, and H. Neven, “Barren plateaus in quantum neural network training landscapes,” Nature Communications, vol. 9, no. 1, p. 4812, 2018. [5] A. Abbas, D. Sutter, C. Zoufal, A. Lucchi, A. Figalli, and S. Woerner, “The power of quantum neural networks,” Nature Computational Science, vol. 1, no. 6, pp. 403–409, 2021. [6] S. Y.-C. Chen, S. Yoo, and Y.-L. L. Fang, “Quantum long short-term memory,” in Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022, pp. 8622– 8626. [7] K.-C. Chen, S. Y.-C. Chen, C.-Y. Liu, and K. K. Leung, “Toward large-scale distributed quantum long short-term memory with modular quantum computers,” arXiv preprint arXiv:2503.14088, 2025. [8] M. Chehimi, S. Y.-C. Chen, W. Saad, and S. Yoo, “Federated quantum long short-term memory (FedQLSTM),” arXiv preprint arXiv:2312.14309, 2023. [9] C.-Y. Liu, S. Y.-C. Chen, K.-C. Chen, W.-J. Huang, and Y.-J. Chang, “Programming variational quantum circuits with quantum-train agent,” arXiv preprint arXiv:2412.01173, 2024. [10] T.-C. Hsieh, Y.-C. Tsai, and S. Y.-C. Chen, “Quantum-enhanced temporal embeddings via a hybrid Seq2Seq architecture,” arXiv preprint arXiv:2602.11578, 2026. [11] Y.-K. Liu, Y.-H. Pan, P.-F. Lu, Y.-C. Tsai, and S. Y.-C. Chen, “Quantumenhanced reinforcement learning with LSTM forecasting signals for optimizing fintech trading decisions,” arXiv preprint arXiv:2507.12835, 2025. [12] J.-H. Chen, M.-K. Hung, Y.-C. Tsai, and S. Y.-C. Chen, “Batched training for QLSTM vs. QFWP: A system-oriented approach to EPCaware RMSE-DA,” arXiv preprint arXiv:2512.21820, 2025. [13] C.-S. Chen, S. Y.-C. Chen, and Y.-C. Tsai, “Benchmarking quantum and classical sequential models for urban telecommunication forecasting,” arXiv preprint arXiv:2508.04488, 2025. [14] S. Y.-C. Chen, Y. Peng, J.-C. Jiang, C.-H. Lin, K.-C. Peng, J. J. Park, H.-H. Tseng, H.-Y. Lin, K.-C. Chen, C.-Y. Liu, and S. Yoo, “Recursive QLSTM with dynamic variational quantum circuit adaptation,” arXiv preprint arXiv:2606.24932, 2026. [15] Environment and Climate Change Canada, “Historical climate data,” Government of Canada climate-data service, 2026, [Online]. Available: https://climate.weather.gc.ca/historical data/search historic data e.html. Accessed: Jun. 25, 2026. [16] Y. Suzuki, Q. Gao, K. C. Pradel, K. Yasuoka, and N. Yamamoto, “Natural quantum reservoir computing for temporal information processing,” Scientific Reports, vol. 12, no. 1, p. 1353, 2022.

Record · ID 978427 · SHA-256 a58fe9fd09567164
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.