Conceptio › Archive › arXiv CS
arXiv CSopen access

Learning to Program Adaptive Non-Local Observables for Machine Learning

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

LEARNING TO PROGRAM ADAPTIVE NON-LOCAL OBSERVABLES FOR MACHINE LEARNING Yu-Ting Lee1

arXiv:2609.18655v1 [cs.LG] 16 Sep 2026

1

Samuel Yen-Chi Chen2

Huan-Hsin Tseng3

Graduate Institute of Communication Engineering, National Taiwan University, Taipei, Taiwan 2 Wells Fargo, New York, NY, USA 3 Brookhaven National Laboratory, AI & ML Department, Upton, NY, USA [email protected], [email protected], [email protected] ABSTRACT

Quantum neural networks (QNNs) are typically built from variational quantum circuits (VQCs), which are limited by local measurements. Adaptive non-local observables (ANO) address this by jointly optimizing circuit parameters and multiqubit measurements. However, existing ANO-based VQCs learn only a single static observable that remains invariant across all inputs. We propose QFWP-ANO, a novel architecture which employs a classical hypernetwork to dynamically program VQC parameters and/or non-local observables conditioned on each input. On multivariate time-series forecasting across four ETT datasets, QFWP-ANO achieves the lowest MSE in 16 of 20 settings and second-lowest in the remaining four, surpassing ANO-based and other strong baselines. On reinforcement learning tasks, QFWP-ANO consistently surpasses ANO-VQCs. Our results establish input-conditioned ANO as an effective approach for enhancing QNNs. Index Terms— Quantum machine learning, Variational quantum circuits, Quantum neural networks, Non-local observables 1. INTRODUCTION Quantum machine learning (QML) and quantum neural networks (QNNs) have emerged as a promising paradigm that leverages the representational expressivity of quantum phenomena such as superposition, entanglement, and quantum interference [1, 2]. In particular, QNNs are increasingly being applied to a wide range of complex machine learning tasks, including reinforcement learning [3–6], classification [7, 8], anomaly detection [9], and time-series prediction [10–14]. However, QNNs are typically built from variational quantum circuits (VQCs) and depend on local measurements, which limit the network’s ability to learn complex data distributions. The views expressed in this article are those of the authors and do not represent the views of Wells Fargo. This article is for informational purposes only. Nothing contained in this article should be construed as investment advice. Wells Fargo makes no express or implied warranties and expressly disclaims all legal, tax, and accounting implications related to this article.

To address this bottleneck, recent work proposes to jointly optimize circuit parameters with trainable observables [15]. Specifically, the adaptive non-local observables (ANO) framework [16] leverages trainable multi-qubit Hermitian observables and shows strong potential across super-resolution [17], reinforcement learning [18], classification [16], and time-series forecasting [14]. Yet, a key limitation of existing ANO-based QNNs is their reliance on a static observable per task, rendering the measurement invariant to the input data stream. To overcome this limitation, we propose QFWP-ANO, which renders non-local observables adaptive to the input. Building on quantum fast weight programmers (QFWP) [12], our approach employs a classical hypernetwork to program the circuit parameters, the non-local observables, or both, yielding a family of input-conditioned models. Experiments on multivariate time-series forecasting (MTSF) and reinforcement learning (RL) demonstrate the superior performance of QFWP-ANO over standard ANO-VQCs. Our contributions: • We introduce QFWP-ANO, a novel QNN architecture that utilizes a classical hypernetwork to dynamically program VQCs along with non-local observables. • On multivariate time-series forecasting across four ETT datasets, QFWP-ANO achieves the lowest MSE in 16 of 20 settings and second-lowest in the remaining four, outperforming standard ANO-VQCs in 17 of 20 cases. • In standard RL environments, QFWP-ANO outperforms ANO-VQCs, while programming the non-local observables can accelerate learning. 2. METHODOLOGY 2.1. Variational Quantum Circuits Variational quantum circuits (VQCs), also known as parameterized quantum circuits (PQCs), are trainable quantum models that process classical data in three stages. Initially, a classical input x is mapped into an n-qubit system via a data encoding unitary circuit U (x), yielding the encoded states U (x)|0⟩⊗n ,

where |0⟩⊗n is the ground state. Subsequently, a parameterized unitary circuit V (θ) evolves this state into V (θ)U (x)|0⟩⊗n . This variational circuit V (θ) is typically structured with alternating layers of trainable single-qubit rotations and multi-qubit entangling gates. Lastly, a measurement layer extracts classical information by calculating the expectation values of a fixed Hermitian observable H. The computation of a VQC can be summarized as a quantum function fVQC (x; θ): ⊗n

fVQC (x; θ) = ⟨0|

†

†

U (x)V (θ)HV (θ)U (x)|0⟩

⊗n

.

D layers Variational V (θ)

Encoding U (x)

|0⟩0

H

Rz (θ0,1 )

Ry (θ0,2 )

Ry (x0 )

Rz (x0 )

|0⟩1

H

Rz (θ1,1 )

Ry (θ1,2 )

Ry (x1 )

Rz (x1 )

|0⟩2

H

Rz (θ2,1 )

Ry (θ2,2 )

Ry (x2 )

Rz (x2 )

|0⟩3

H

Rz (θ3,1 )

Ry (θ3,2 )

Ry (x3 )

Rz (x3 )

H(ϕ)

(1)

2.2. Adaptive Non-Local Observables Motivated by the Heisenberg picture, where quantum evolution is characterized by dynamical observables, adaptive non-local observables (ANO) [16] replace the fixed observable of a standard VQC with a trainable Hermitian operator H(ϕ) parameterized by ϕ. A k-local observable takes the form:   c11 a12 + ib12 a13 + ib13 · · · a1K + ib1K ∗ c22 a23 + ib23 · · · a2K + ib2K    ∗ ∗ c33 · · · a3K + ib3K  H(ϕ) =    ..  .. .. .. ..  .  . . . . ∗ ∗ ∗ ··· cKK (2) where k ≤ n, K = 2k , and the parameter set ϕ = 2 (aij , bij , cii )K i,j=1 consists of K real parameters. Utilizing a trainable observable strictly expands a QNN’s function class: a conventional VQC with a fixed local observable is a special case of ANO [16]. Moreover, a k-local observable H(ϕ) operates jointly across k qubits, coupling features between distant qubits to facilitate an information mixture that common Pauli-Z measurements cannot capture.

Fig. 1: VQC architecture of QFWP-ANO. Each layer consists parameterized Rz , Ry rotations, circular CNOT gates, and encoding U (x). Trainable k-local observables H(ϕ) are employed for measurement. Table 1: Statistics of the datasets and environments. (a) The four ETT datasets.

Datasets

ETTh1 & ETTh2

ETTm1 & ETTm2

Variates Timesteps Sample rate

7 17,420 1 hour

7 69,680 5 min

(b) RL environments.

Environments

Acrobot

SimpleCrossingS9N1

State Space Action Space State dim Actions Reward

Continuous Discrete 6 3 −1/step

Discrete Discrete 147 7 +1 on goal, else 0

2.3. QFWP-ANO QFWP-ANO employs a classical neural network as a hypernetwork to program the rotation angles of the variational circuit and/or the ANO parameters, conditioned on each input. Specifically, this hypernetwork utilizes an encoder-decoder architecture. A shared MLP encoder first processes the input into a latent representation z. Taking z as input, specific decoders then generate the necessary parameters: a linear layer for the rotation angles θ, and a separate two-layer MLP for the ANO parameters ϕ. In contrast to QFWP, which stores temporal memory by recurrently accumulating updates, QFWP-ANO maps each input independently, carrying no accumulated state. We use a data re-uploading VQC structure [8, 19] combined with ANO. Initialized by a layer of Hadamard gates, each of the D layers consists of 3 parts: parameterized Rz and Ry rotation gates, a circular topology of CNOT gates, and Ry and Rz encoding gates for data re-uploading (Fig. 1). Finally, a combinatorial measurement scheme with k-local observables is applied. We compute expected values across  all nk combinations of k qubits from the n available; this

generates one output value per combination to account for multi-qubit correlations.

3. EXPERIMENTAL SETTINGS We evaluate three QFWP-ANO variants in which the hypernetwork programs only the rotation angles (QFWP-ANO-R), only the ANO parameters (QFWP-ANO-O), or both (QFWPANO-RO). We test QFWP-ANO on two common QML tasks: multivariate time-series forecasting (MTSF) and RL. 3.1. Multivariate Time Series Forecasting For a multivariate series with C variates (channels), let L be the lookback window and H the forecasting horizon. Given historical data Xt ∈ RC×L , MTSF aims to predict future valb t ∈ RC×H . The ground truth is denoted Yt ∈ RC×H . ues Y Following Lee et al. [14], we apply instance normaliza-

tion [20] to each channel of Xt against distribution shift: X′t = (Xt − µ) ⊙ (σ 2 + ϵ)−1/2 ,

(3)

where µ, σ 2 ∈ RC are the per-channel mean and variance over the lookback window and ϵ ensures numerical stability. The normalized input is flattened and projected to a latent ht , transformed by QFWP-ANO, and mapped to forecasts via a linear layer and de-normalization:  p Ŷt = Wout f VQC (ht ; θ, ϕ) + bout ⊙ σ 2 + ϵ + µ, (4) n where f VQC (ht ; θ, ϕ) ∈ R(k ) . We conduct experiments on four real-world datasets from the Electricity Transformer Temperature (ETT) benchmark [21] (Table 1a). We adopt a standard 6:2:2 train/validation/test split and report MSE and MAE as our main metrics [10, 14, 21]. We compare against several state-of-theart baselines: MTSF-ANO [14], an ANO-based MTSF model; QuLTSF [10], a quantum method for long-term MTSF; and the strong linear baselines DLinear/NLinear/Linear [22], which are known to outperform many Transformer-based methods in long-term MTSF. A QFWP [12] baseline is also included. We set n = C qubits and sweep VQC depth D = {1, 3} and non-locality k ∈ {1, . . . , 7}, reporting the best result over k and D. We train up to 100 epochs with Adam using MSE loss.

3.2. Reinforcement Learning Following Lin et al. [18], we embed QFWP-ANO as the function approximator in Asynchronous Advantage Actor-Critic (A3C) [23]: a linear layer encodes the state st into a latent ht , which is the circuit input for two independent QFWP-ANO instances (policy and value) that produce πθ (a | st ) and Vψ (st ). Each QFWP-ANO instance’s hypernetwork reads st directly to generate the parameters. Following prior work [3, 6, 18], we evaluate on Acrobot, a swing-up control problem, and MiniGrid-SimpleCrossingS9N1, a sparse-reward navigation task (Table 1b). We compare our method against an ANO-VQC baseline. In all configurations, we use n = 4 qubits, D = 4 VQC layers, and non-locality k = 3. Further, we introduce a trainable input scaling w (e.g., Ry (wi xi ) rather than Ry (xi )). Since trainable input scaling w has proven beneficial in standard QRL agents [3, 5], we evaluate models both with and without w to investigate how it affects ANO-based agents. All models are trained using A3C for 5,000 (Acrobot) or 10,000 (SimpleCrossingS9N1) episodes, averaged over 10 random seeds. 4. EXPERIMENTS 4.1. MTSF Results Table 2 summarizes the MTSF performance at a fixed lookback L = 16. QFWP-ANO-O achieves the lowest MSE in 13 of the 20 settings; combined with QFWP-ANO-RO, they rank

Fig. 2: Reward over episodes. Band is ± standard deviation.

first in 16 settings and secure the second-lowest MSE in the remaining four. This advantage is largest and most consistent on ETTm1, where the improvement over the best baseline by up to 6.6% and the improvement grows steadily with the horizon. Further, QFWP-ANO outperforms standard ANO-VQCs (MTSF-ANO) in 17 of 20 cases. Across all ETT datasets, programming the observable (QFWP-ANO-O, QFWP-ANO-RO) is more effective than programming only the rotation angles (QFWP-ANO-R). Our finding indicate that input-conditioned non-local measurement provides significant gains. 4.2. RL Results Fig. 2 shows the learning curves, with (dark) and without (light) the trainable input scaling w. In Acrobot, all three QFWP-ANO variants substantially outperform the ANO-VQC baseline from episode 1000 to 4000. Although the models ultimately converge to similar final returns, QFWP-ANO-O and QFWP-ANO-RO reach their peak performance as early as around episode 1000. This early convergence demonstrates that programming the observable effectively improves sample efficiency. In SimpleCrossingS9N1, QFWP-ANO-RO reaches

Table 2: Multivariate forecasting results. Results are averaged over 10 different random seeds. Lookback L = 16 and prediction horizon H ∈ {1, 5, 16, 32, 48}. Lower MSE and MAE are better. The best result is highlighted in bold and the second best is highlighted with underline. IMP. is the improvement between the best QFWP-ANO method and the best baseline, where a larger value indicates better improvement. QFWP-ANO models implemented by us; other results from [14].

ETTm2

ETTm1

ETTh2

ETTh1

Methods Metric 1 5 16 32 48 1 5 16 32 48 1 5 16 32 48 1 5 16 32 48

IMP. MSE 3.6% 4.4% 2.3% 0.7% -2.3% -1.4% 2.5% 2.3% 0.5% 0.0% 0.0% 3.6% 5.2% 5.7% 6.6% -3.1% -1.7% 1.0% 0.7% 0.6%

QFWP-ANO-RO QFWP-ANO-R QFWP-ANO-O MSE MAE MSE MAE MSE MAE 0.133 0.239 0.137 0.243 0.132 0.237 0.324 0.364 0.337 0.371 0.323 0.362 0.380 0.400 0.397 0.410 0.379 0.398 0.425 0.426 0.437 0.432 0.423 0.425 0.439 0.431 0.446 0.435 0.437 0.429 0.075 0.170 0.082 0.179 0.075 0.169 0.118 0.215 0.129 0.229 0.118 0.215 0.174 0.262 0.177 0.265 0.173 0.261 0.222 0.293 0.223 0.294 0.221 0.292 0.261 0.315 0.261 0.316 0.259 0.314 0.048 0.136 0.050 0.138 0.048 0.136 0.106 0.200 0.111 0.205 0.107 0.200 0.311 0.332 0.324 0.341 0.311 0.331 0.560 0.448 0.575 0.459 0.564 0.451 0.660 0.501 0.668 0.507 0.668 0.505 0.033 0.103 0.035 0.107 0.033 0.103 0.059 0.140 0.060 0.143 0.059 0.140 0.099 0.190 0.101 0.193 0.099 0.190 0.146 0.237 0.149 0.241 0.145 0.236 0.179 0.266 0.183 0.271 0.178 0.265

MTSF-ANO MSE MAE 0.137 0.243 0.338 0.371 0.388 0.405 0.426 0.427 0.427 0.425 0.079 0.173 0.121 0.218 0.180 0.268 0.225 0.295 0.264 0.318 0.048 0.135 0.110 0.204 0.328 0.342 0.594 0.462 0.720 0.522 0.034 0.106 0.059 0.141 0.100 0.191 0.146 0.238 0.179 0.267

Table 3: Impact of ANO non-locality and VQC depth. Color shows the best variant among R/O/RO. Bold indicates lowest MSE across all k = 1 to k = 7 within that column, i.e., the best variant and non-locality k for a given (H, D) pair. (a) VQC depth D = 1.

Settings H = 1 H = 5 H = 16 H = 32 H = 48 k=1 k=2 k=3 k=4 k=5 k=6 k=7

0.073 0.072 0.073 0.074 0.074 0.077 0.179

0.161 0.152 0.152 0.152 0.155 0.168 0.268

0.260 0.242 0.240 0.242 0.248 0.263 0.379

0.359 0.340 0.338 0.340 0.347 0.360 0.475

0.404 0.388 0.385 0.388 0.391 0.404 0.519

(b) VQC depth D = 3.

Settings H = 1 H = 5 H = 16 H = 32 H = 48 k=1 k=2 k=3 k=4 k=5 k=6 k=7

0.073 0.072 0.073 0.074 0.076 0.079 0.167

0.160 0.152 0.153 0.153 0.155 0.168 0.277

0.260 0.243 0.242 0.243 0.251 0.267 0.374

0.359 0.341 0.340 0.342 0.350 0.363 0.475

0.403 0.388 0.386 0.388 0.395 0.408 0.520

the highest final reward and remains the best throughout training, ahead of all other models. Employing the trainable input scaling w only significantly boosts ANO-VQC in Acrobot.

QuLTSF MSE MAE 0.162 0.254 0.431 0.409 0.437 0.422 0.481 0.447 0.474 0.441 0.074 0.167 0.122 0.222 0.177 0.271 0.222 0.297 0.259 0.318 0.051 0.136 0.128 0.210 0.432 0.369 0.788 0.517 0.956 0.588 0.032 0.095 0.058 0.138 0.104 0.196 0.158 0.250 0.195 0.284

QFWP MSE MAE 0.397 0.406 0.700 0.522 0.662 0.514 0.670 0.523 0.645 0.514 0.117 0.225 0.155 0.261 0.193 0.284 0.243 0.314 0.277 0.331 0.105 0.201 0.198 0.271 0.520 0.423 0.859 0.552 0.986 0.608 0.051 0.136 0.073 0.166 0.118 0.216 0.168 0.263 0.205 0.295

DLinear MSE MAE 0.164 0.258 0.478 0.429 0.459 0.428 0.494 0.447 0.481 0.439 0.074 0.166 0.126 0.225 0.177 0.269 0.222 0.294 0.259 0.315 0.052 0.138 0.133 0.214 0.451 0.378 0.835 0.534 1.018 0.610 0.032 0.095 0.060 0.142 0.108 0.201 0.162 0.255 0.198 0.288

NLinear MSE MAE 0.162 0.256 0.489 0.432 0.487 0.441 0.518 0.458 0.502 0.450 0.074 0.166 0.126 0.226 0.179 0.271 0.223 0.296 0.261 0.317 0.052 0.137 0.132 0.214 0.451 0.378 0.836 0.534 1.019 0.610 0.032 0.096 0.060 0.142 0.108 0.201 0.162 0.255 0.198 0.288

Linear MSE MAE 0.199 0.284 0.510 0.443 0.474 0.435 0.507 0.453 0.490 0.443 0.076 0.170 0.127 0.228 0.178 0.270 0.222 0.295 0.259 0.316 0.052 0.138 0.133 0.214 0.452 0.378 0.836 0.534 1.019 0.610 0.032 0.096 0.060 0.142 0.108 0.201 0.162 0.255 0.198 0.288

4.3. Impact of ANO Non-Locality and VQC Depth We evaluate how VQC depth and ANO non-locality influence MTSF performance of QFWP-ANO variants. Table 3 reports the lowest average MSE across the four ETT datasets for each (k, H) configuration and circuit depth D ∈ {1, 3}, colored by the best-performing QFWP-ANO variant. We make four key observations: (1) MSE initially decreases from k = 1, reaches a minimum around k ∈ {2, 3}, and rises sharply by k = 7, which is likely due to combinatorial ANO scheme leaving only a single expectation value at k = 7; (2) QFWP-ANO-RO and QFWP-ANO-O achieve the lowest MSE for k ≤ 4, while QFWP-ANO-R starting to overtake them from k = 5 and dominates across all horizons at k = 7; (3) moderate nonlocality consistently yields the best performance, regardless of the specific variant; and (4) shallower circuits (D = 1) match or outperform deeper ones (D = 3) in most settings. 5. CONCLUSION We introduced QFWP-ANO, a novel QNN architecture that utilizes a classical hypernetwork to dynamically program VQCs along with non-local observables. Across several MTSF and RL benchmarks, programming the non-local observables yields the most substantial improvements. Specifically, in MTSF, QFWP-ANO ranks first in MSE in 16 of 20 settings and ranks second in the remaining four. In RL environments, QFWP-ANO consistently outperforms ANO-VQCs, and programming the non-local observables can effectively accelerate learning. These findings establish input-conditioned ANO as a promising route to more capable quantum models.

6. REFERENCES [1] M. Cerezo, Andrew Arrasmith, Ryan Babbush, Simon C. Benjamin, Suguru Endo, Keisuke Fujii, Jarrod R. McClean, Kosuke Mitarai, Xiao Yuan, Lukasz Cincio, and Patrick J. Coles, “Variational quantum algorithms,” Nature Reviews Physics, vol. 3, no. 9, pp. 625–644, Sep 2021. [2] Jarrod R McClean, Jonathan Romero, Ryan Babbush, and Alán Aspuru-Guzik, “The theory of variational hybrid quantumclassical algorithms,” New Journal of Physics, vol. 18, no. 2, pp. 023023, feb 2016. [3] Sofiene Jerbi, Casper Gyurik, Simon Marshall, Hans Briegel, and Vedran Dunjko, “Parametrized quantum policies for reinforcement learning,” in Advances in Neural Information Processing Systems, M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan, Eds. 2021, vol. 34, pp. 28362–28375, Curran Associates, Inc.

[14] Yu-Ting Lee, Huan-Hsin Tseng, and Samuel Yen-Chi Chen, “Multivariate time series forecasting with adaptive non-local observables,” 2026. [15] Samuel Yen-Chi Chen, Huan-Hsin Tseng, Hsin-Yi Lin, and Shinjae Yoo, “Learning to measure quantum neural networks,” in 2025 IEEE International Conference on Acoustics, Speech, and Signal Processing Workshops (ICASSPW), 2025, pp. 1–5. [16] Hsin-Yi Lin, Huan-Hsin Tseng, Samuel Yen-Chi Chen, and Shinjae Yoo, “Adaptive non-local observable on quantum neural networks,” in 2025 IEEE International Conference on Quantum Computing and Engineering (QCE), 2025, vol. 01, pp. 1884– 1893. [17] Hsin-Yi Lin, Huan-Hsin Tseng, Samuel Yen-Chi Chen, and Shinjae Yoo, “Quantum super-resolution by adaptive non-local observables,” in ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026, pp. 22027–22031.

[4] Gyu Seon Kim, Samuel Yen-Chi Chen, Soohyun Park, and Joongheon Kim, “Quantum reinforcement learning for coordinated satellite systems,” in ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2025, pp. 1–5.

[18] Hsin-Yi Lin, Samuel Yen-Chi Chen, Huan-Hsin Tseng, and Shinjae Yoo, “Quantum reinforcement learning by adaptive non-local observables,” in 2025 IEEE International Conference on Quantum Computing and Engineering (QCE), 2025, vol. 02, pp. 241–246.

[5] Yu-Ting Lee, Samuel Yen-Chi Chen, and Fu-Chieh Chang, “Quantum hierarchical reinforcement learning via variational quantum circuits,” 2026.

[19] Maria Schuld, Ryan Sweke, and Johannes Jakob Meyer, “Effect of data encoding on the expressive power of variational quantum-machine-learning models,” Phys. Rev. A, vol. 103, pp. 032430, Mar 2021.

[6] Samuel Yen-Chi Chen, Chao-Han Huck Yang, Jun Qi, Pin-Yu Chen, Xiaoli Ma, and Hsi-Sheng Goan, “Variational Quantum Circuits for Deep Reinforcement Learning,” IEEE Access, vol. 8, pp. 141007–141024, 2020. [7] Edward Farhi and Hartmut Neven, “Classification with quantum neural networks on near term processors,” 2018. [8] Adrián Pérez-Salinas, Alba Cervera-Lierta, Elies Gil-Fuster, and José I. Latorre, “Data re-uploading for a universal quantum classifier,” Quantum, vol. 4, pp. 226, Feb. 2020.

[20] Taesung Kim, Jinhee Kim, Yunwon Tae, Cheonbok Park, JangHo Choi, and Jaegul Choo, “Reversible instance normalization for accurate time-series forecasting against distribution shift,” in International Conference on Learning Representations, 2022. [21] Haoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang, Jianxin Li, Hui Xiong, and Wancai Zhang, “Informer: Beyond efficient transformer for long sequence time-series forecasting,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, no. 12, pp. 11106–11115, May 2021.

[9] Marco Casalbore, Leonardo Lavagna, Antonello Rosato, and Massimo Panella, “Time series anomaly detection with quantum variational methods and set covering,” in ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026, pp. 1846–1850.

[22] Ailing Zeng, Muxi Chen, Lei Zhang, and Qiang Xu, “Are transformers effective for time series forecasting?,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, no. 9, pp. 11121–11128, Jun. 2023.

[10] Hari Hara Suthan Chittoor, Paul Robert Griffin, Ariel Neufeld, Jayne Thompson, and Mile Gu, “Qultsf: Long-term time series forecasting with quantum machine learning,” in Proceedings of the 17th International Conference on Agents and Artificial Intelligence - Volume 1: QAIO. INSTICC, 2025, pp. 824–829, SciTePress.

[23] Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in Proceedings of The 33rd International Conference on Machine Learning, Maria Florina Balcan and Kilian Q. Weinberger, Eds., New York, New York, USA, 20–22 Jun 2016, vol. 48 of Proceedings of Machine Learning Research, pp. 1928–1937, PMLR.

[11] Samuel Yen-Chi Chen, Shinjae Yoo, and Yao-Lung L. Fang, “Quantum long short-term memory,” in ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022, pp. 8622–8626. [12] Samuel Yen-Chi Chen, “Learning to program variational quantum circuits with fast weights,” in 2024 International Joint Conference on Neural Networks (IJCNN), 2024, pp. 1–9. [13] Andrea Ceschini, Antonello Rosato, Massimo Panella, and Samuel Yen-Chi Chen, “Quantum fast weight programming for time series prediction,” in ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026, pp. 22032–22036.

Record · ID 965428 · SHA-256 18887636eea0267e
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.