Hybrid Variational Quantum Circuits for Multivariate Regression and High-Dimensional Data Reconstruction
arXiv:2609.17358v1 [cs.LG] 15 Sep 2026
Koffi O. AYENA
Frédéric HOLWECK
Serge IOVLEFF
Amah S. D’ALMEIDA
SINERGIES (UR 4662), UMLP LAMMA, Universite de Lomé ICB, UTBM ICB, UTBM 01 BP 1515 Lomé-Togo F-90000 Belfort, France F-90000 Belfort, France F-90000 Belfort, France [email protected] [email protected] LAMMA, Universite de Lomé [email protected] 01 BP 1515 Lomé-Togo ORCID: 0009-0009-0858-258X
Abstract—Variational quantum circuits (VQCs) are parameterized quantum circuits optimized classically. We propose a hybrid variational quantum circuit (HVQC) extending VQCs with a classical affine post-measurement layer, enabling vector-valued regression without the linear overhead of independent scalar circuits. Theoretically, we show that elementary one- and two-qubit circuits can approximate quadratic functions and products via data re-uploading and entanglement, providing the foundations of the full architecture. Experimentally, on two synthetic image reconstruction datasets and the Friedman1 benchmark (40,568 test samples), our HVQC matches Gaussian Process Regression and outperforms XGBoost and Random Forest. An ablation study confirms that both quantum and classical components are essential, and results highlight the central role of the feature map in hybrid quantum-classical models. Index Terms—Variational quantum circuits, quantum machine learning, multivariate regression, feature maps
I. I NTRODUCTION Quantum computing has experienced significant advances in the era of noisy intermediate-scale quantum (NISQ) processors [1]. These processors, although limited in the number of qubits and subject to noise, have paved the way for new computing paradigms, including variational quantum algorithms (VQAs). Popularized by the foundational work of Peruzzo [2] on the variational solver for quantum chemistry, variational quantum circuits (VQCs) have become central models for exploring the potential advantages of quantum computing in various fields, notably quantum chemistry and, more recently, quantum machine learning (QML) [3], [4]. The enthusiasm for variational quantum circuit (HVQC) stems from their hybrid architecture: a parameterized quantum circuit, whose parameters are optimized by a classical computer. This approach partially circumvents the limitations of current quantum hardware. Formally, a VQC can be described as a parameterized unitary U (z; θ) applied to an initial state This work was supported by the Agence Française de Développement (AFD) through the Agence Nationale de la Recherche (ANR-21-PEA20007), under the Partenariats with "l’Enseignement supérieur Africain" (PEA) program.
m0
|0⟩ = |0⟩⊗ , producing a final state whose expectation value of an observable Ô provides the model’s output [5]: f (z; θ) = ⟨0| U † (z; θ)ÔU (z; θ) |0⟩ d
(1)
where z ∈ R represents the input data, typically encoded via parameterized rotation gates [6]. In the QML landscape, while classification tasks [7]–[9], and "unpublished" [10] have been extensively studied, quantum regression has remained a theoretically less explored topic until recently, as highlighted by authors in [11] on applying QML to practical regression on NISQ hardware. Yet, the ability to predict continuous values is crucial for numerous scientific and industrial applications, including financial time series forecasting [12], physical data modeling [13], or climate prediction [14]. Research on quantum regression circuits has experienced a remarkable acceleration since 2024, marked by several important milestones which we organize chronologically. The year 2024 constitutes a turning point with the first significant experimental demonstrations. A pioneering study [11] applied variational quantum regression to the Auto-MPG dataset on NISQ hardware with error mitigation. Their results demonstrate that VQAs can outperform classical models like XGBoost, and that error mitigation techniques are effective in bringing the performance of noisy simulators closer to that of ideal simulators. Concurrently, the PennyLane platform published a tutorial demonstrating multidimensional regression with a two-qubit variational circuit to approximate the function f (x1 , x2 ) = 1 2 2 2 2 (x1 + x2 ), achieving an R score of 0.983 in "unpublished" [15]. This work illustrates the ability of VQCs to build partial Fourier series for function approximation, consistent with the VQC expressivity theory established in [6]. Reference [16] introduced the PQML (Predictive Quantum Machine Learning) tool to predict the reproducibility of results across different quantum machines, a crucial advance for the reliability of quantum regression applications in a heterogeneous NISQ context.
The year 2025 saw the emergence of sophisticated optimization techniques and applications to complex problems. In [17], the authors proposed a novel state preparation method for variational quantum regression, using optimization techniques based on the ZX-calculus (Pauli pushing, phase folding, Hadamard pushing). Their results demonstrate that these optimizations enable the successful execution of quantum regression algorithms on current hardware, significantly reducing circuit depth. This approach was extended to multivariate time series in "unpublished" [18] through the MTS-QRC (Multivariate Time Series Quantum Reservoir Computing) framework. Applied to Lorenz and ENSO (El Niño-Southern Oscillation) systems, this method achieved a mean squared error (MSE) of 0.0087 and 0.0036, respectively. Interestingly, their work revealed that hardware noise can sometimes act as an implicit regularizer, improving performance compared to ideal simulators a counter-intuitive yet promising phenomenon for NISQ applications. Parallel advances in error mitigation directly benefited regression applications. The authors of [19] proposed significant improvements to Clifford Data Regression (CDR) with Energy Sampling (ES) and Non-Clifford Extrapolation (NCE), enhancing the fidelity of computations on noisy hardware without additional quantum overhead. Automatic design of quantum architectures for regression using genetic algorithms was explored in "unpublished" [20]. Their Reduced Regressor QNN framework explores circuit depth, configuration of parameterized gates, and data reuploading patterns, demonstrating that these evolved circuits, although compact, can achieve competitive performance against 17 classical regression models on 22 non-linear benchmark functions. The year 2026 marks a consolidation of the field. Joo’s editorial [21] in Frontiers in Physics reviews the progress in algorithm optimization and error mitigation, confirming that these areas are now mature enough to support practical applications such as quantum regression. The challenges are progressively shifting from fundamental feasibility towards comparative efficiency and demonstrable quantum advantage. Although variational quantum circuits (VQCs) have shown promise for scalar regression, they suffer from the barren plateau phenomenon [22], [23] that makes training difficult, and their naive extension to vector-valued regression incurs a prohibitive linear complexity in the output dimension. To address these limitations, we introduce a hybrid variational quantum circuit (HVQC) that handles vector-valued regression in a single unified circuit, thereby avoiding the linear overhead of independent approaches [11], [15]. Unlike quantum kernel methods [24], [25] where the circuit is fixed, our architecture jointly optimizes all quantum and classical parameters end-toend, with the final affine layer enabling the model to reach any target in the output vector space. This paper is organized as follows. Section II presents the mathematical foundations of HVQCs. In Section III, we demonstrate the approximation capabilities of elementary
one- and two-qubit circuits, thereby providing the theoretical foundations for our architecture. Complete HVQC architecture is established in Section IV, where we define the key components: data encoding, and the variational parameterization that enables learning. Section V presents experimental results comparing the performance of our HVQC, evaluated using the R2 score and mean squared error (MSE), on two simulated datasets with different feature maps, against several classical regression models, including Gaussian process regression (GPR), random forest regression (RFR), multioutput XGBoost Regression (XGB), and two classical neural networks. The section concludes with a global analysis of the quantum states after measurement, revealing an implicit clustering behavior. II. T HEORETICAL FRAMEWORK OF A HVQC We consider the following supervised learning problem: (i) given a training set D = {(z(i) , x(i) )}N ∈ Rd i=1 , where z (i) m and x ∈ R , the objective is to learn a function fΘ : Rd → Rm , parameterized by Θ, minimizing the mean squared error: N
L(Θ) =
1 X ∥fΘ (z(i) ) − x(i) ∥22 . N i=1
(2)
a) Baseline architecture of the quantum regressor: The conventional architecture of the variational quantum regressor proceeds in three steps. First, a quantum feature map embeds the input vector z into the Hilbert space of a system of m0 ≥ ⌈log2 m⌉ qubits, producing the state |ϕ(z)⟩ ∈ H ∼ = (C2 )⊗m0 . In second step, this encoding, whose dimensionality is chosen to match the target dimension m, is then evolved by a parameterized quantum circuit U (θ), generating the variational state: ⊗m0
|ψ(z; θ)⟩ = ϕ(z)U (θ L ) · · · ϕ(z)U (θ 1 )ϕ(z) |0⟩
.
(3)
The model’s prediction emerges from measuring the m0 m0 qubits in the computational basis {|k⟩}2k=0−1 . The theoretical probability distribution of outcomes is given by the projectors Mk = |k⟩⟨k|: pk (z; θ) = ⟨ψ(z; θ)|Mk |ψ(z; θ)⟩.
(4)
This distribution p(z; θ) = (p0 , . . . , p2m0 −1 ) resides in the m0 probability simplex ∆2 −1 , defined by: m0 −1
∆2 (
= 2m0
(q0 , q1 , · · · , q2m0 −1 ) ∈ R
| qk ≥ 0,
0 −1 2m X
) qk = 1 .
k=0
In practice, access to this distribution is obtained through sampling (shots). For S measurement repetitions, one obtains a frequentist estimate p̂k (θ, z) = ck /S, where ck is the count of outcome k. Each projector Mk = |k⟩⟨k| defines an observable whose measurement yields the probability of observing the computational basis state |k⟩. It can be expressed as a tensor product of single-qubit observables
Mk =
m0 O
1
2 i=1
I + (−1)ki σz ,
with
1 σz = 0
0 −1
where ki ∈ {0, 1} denotes the i-th bit of the integer k on m0 bits, σz is the Pauli Z matrix and I denotes the identity matrix. In the third step, we perform post-processing to handle model limitations and ensure that the output dimensions are consistent. This final step is discussed in the following two paragraphs. b) Geometric limitation of the baseline model: A fundamental limitation of the architecture defined by equations (3) and (4) lies in the confinement of its output to the m0 probabilistic simplex ∆2 −1 . Although a trivial linear postprocessing could theoretically project this output onto Rm , the internal geometry of the learned representations remains that of a convex polytope with extremal properties — its points being convex combinations of computational basis states. This geometric constraint intrinsically limits the model’s capacity to capture data structures exhibiting a different underlying geometry (unbounded, non-trivial topology), necessitating an appropriate preprocessing of input data. c) Extension via post-variational affine transformation: To overcome this limitation while preserving the model’s differentiability, we propose incorporating a learnable affine transformation of the probability distribution p(θ, z), like "unpublished" [26]. The complete regression function is then written as: fΘ (z) = W p(θ, z) + b, (5) where Θ = {θ, W, b} encompasses all parameters. The m0 matrix W ∈ Rm×2 and the bias vector b ∈ Rm perform a fundamental geometric transformation: they allow the model to reach any point in Rm by learning a linear map from the simplex to the target space. This formulation breaks the simplexial geometry of internal representations and offers a clear geometric interpretation, conferring upon the model increased flexibility to adapt to complex data patterns. To better understand the capabilities of our model, we now turn to the expressivity of HVQCs, focusing on their ability to approximate non-linear functions. III. E XPRESSIVITY OF VQC S We demonstrate here the capacity of elementary one- and two-qubit quantum circuits to approximate fundamental nonlinearities, thereby establishing the theoretical basis for our architecture. We validate Propositions 1 and 2 via statevector simulation using PennyLane, neglecting shot noise in the probability estimation. Proposition 1 (Approximation of the square function). Let K = [−2, 2]. For any ε ∈ (0, 1), there exist β < 1 and k ≥ 1, such that the measurement of the quantum state |ψ(z)⟩ = RY (z) |0⟩ in the computational basis satisfies sup z 2 − 4β −2k P ψ(β k z) = |1⟩ ≤ ε (6) z∈K
Proof. We have the quantum state |ψ(z)⟩ = Ry (z) |0⟩. Computing explicitly, we obtain: cos z2 |ψ(z)⟩ = sin z2 The probability of measuring |1⟩ is: P (|ψ(z)⟩ = |1⟩) = sin
z 2 2
(7)
Hence, βkz ψ(β k z) = |1⟩ = 4β −2k sin2 2
4β −2k P
2
For small u, we have the Taylor expansion sin2 u2 = u4 − u 6 k 48 + O(u ). Here u = β z is small for large k. Thus: 2k 2 βkz β z β 4k z 4 4β −2k sin2 = 4β −2k − + O(β 6k z 6 ) 2 4 48 4
β 2k z 4 + O(β 4k z 6 ) 12 By Taylor’s theorem with the Lagrange remainder, for any u ∈ R, we have = z2 −
sin2
|u|4 u u2 − ≤ 2 4 48
Applied to u = β k z: 4β −2k sin2
β 2k |z|4 βkz |β k z|4 − z 2 ≤ 4β −2k = 2 48 12
Then for all z ∈ K: 4β −2k P
4β 2k ψ(β k z) = |1⟩ − z 2 ≤ 3
For sufficiently large k, β 2k becomes arbitrarily small (since β < 1), so this difference tends to 0 uniformly on K. Thus, for any ε ∈ (0, 1), it suffices to choose β ≤ 1 and k 2k large enough such that 4β3 ≤ ε, which yields: sup z 2 − 4β −2k P ψ(β k z) = |1⟩ ≤ ε z∈K
The approximation of the function x2 by the HVQC defined in Proposition 1 achieves an R2 score of 1 on the training set and 1 on the test set, using 900 points split in a 70%–30% proportion, as illustrated in Fig. 1. Using the identity 1 (x + y)2 − (x − y)2 x×y = 4 combined with the square function approximation obtained in Proposition 1, we can then approximate the product x × y using two RY rotation gates applied to two qubits. Proposition 2 (Quantum approximation of the product). For any ε ∈ (0, 1), there exist parameters β < 1, k ≥ 1, and a quantum circuit acting on a 2-qubit system such that measuring a quantum state yields an approximation x fy satisfying:
two distinct qubits, which allows obtaining both probabilities in execution of the circuit. On four qubits, Proposition 2 with CNOT gates can be rewritten as: RY (+π/2) ⊗ I ⊗ RY (−π/2) ⊗ I ◦ CNOT0,1 (14) ⊗ CNOT2,3 ◦ RY (β k x)⊗2 ⊗ RY (β k y)⊗2 |0000⟩,
4.0 3.5 3.0
y
2.5 2.0
with CN OTi,j |qi qj ⟩ → |qi , (qj ⊕ qi )⟩ and ⊕ means addition modulo 2 (classical XOR). And denoting by P01 and P23 the probabilities of measuring |11⟩ on qubits (0, 1) and (2, 3) respectively, β22k P01 − P23 approximates xy. We further observe the following
1.5 1.0 0.5 0.0
test prediction train prediction Real 2.0
1.5
1.0
0.5
0.0
x
0.5
1.0
1.5
2.0
Observation 1 (Approximations realized by the two-qubit circuit). Consider the quantum circuit C(α) acting on two qubits defined by:
Fig. 1. Square function approximation.
sup (x,y)∈[−1,1]2
|x × y − x fy| ≤ ε
(8)
More precisely, the approximation is given by: x fy = β −2k × P ψ(β k (x + y)) = |1⟩ − P
CN OT0,1 · (I ⊗ RY (α)) · (RY (x) ⊗ RY (y)) |00⟩
(15)
For small values of x and y (in the neighborhood of 0), the measurement probabilities in the computational basis realize the following approximations:
ψ(β k (x − y)) = |1⟩ (9)
where |ψ(z)⟩ = RY (z) |0⟩ is the quantum state defined in Proposition 1. Proof. From Proposition 1, for any z ∈ K, we have: 4β 2k ψ(β k z) = |1⟩ ≤ (10) 3 Since x + y, x − y ∈ K, and using the remarkable identity: z 2 − 4β −2k P
1 (x + y)2 − (x − y)2 (11) 4 The approximation error follows from the triangle inequality: x×y =
|x × y − x fy| 1h (x + y)2 − 4β −2k P ψ(β k (x + y)) = |1⟩ ≤ 4 i + (x − y)2 − 4β −2k P ψ(β k (x − y)) = |1⟩ 4β 2k 2β 2k 1 4β 2k + = (12) ≤ 4 3 3 3 For any ε ∈ (0, 1), it therefore suffices to choose β < 1 and k sufficiently large such that: 2β 2k ≤ε 3
C(α) : |00⟩ 7→ |ψ(x, y, α)⟩ =
(13)
The corresponding quantum circuit can be implemented by preparing the states ψ(β k (x + y)) and ψ(β k (x − y)) on
y2 x2 y 2 x2 − + (16) 4 4 16 y2 x2 y 2 P01 (x, y, 0) = P00 (x, y, π) ≈ − (17) 4 16 2 2 x y (18) P10 (x, y, 0) = P11 (x, y, π) ≈ 16 2 2 2 x x y P11 (x, y, 0) = P10 (x, y, π) ≈ − (19) 4 16 where Pij (x, y, α) is the probability of measuring the state |ij⟩. P00 (x, y, 0) = P01 (x, y, π) ≈ 1 −
In Observation 1, entanglement is preserved only when RY (α) precedes the CNOT, yielding a non-zero determinant 1 2 sin x cos(y + α). The only exceptions are when sin x = 0 or cos(y + α) = 0 (first qubit in a computational basis state). IV. P ROPOSED HVQC T OOLS AND A RCHITECTURE Three design principles follow from Propositions 1 and 2: (i) data re-uploading — re-encoding z′ at each layer (22) builds complex dependencies like β k z in (6); (ii) entanglement — the cascade CNOT (26) generalizes (14) to all feature dimensions; (iii) affine post-processing — (29) generalizes the rescaling 4β −2k in (6) to arbitrary linear combinations, lifting the simplex constraint. This section details the complete architecture of our hybrid variational quantum circuit. A. Hilbert space and data encoding Let H = (C2 )⊗m0 be the Hilbert space associated with the m0 -qubit system, of dimension dim(H) = 2m0 ≥ m. A quantum state |ψ⟩ ∈ H can be decomposed in the computational m0 basis {|i⟩}2i=0 −1 as:
1
y 2 + x 2y 2 4 16
x2 4
1.0 0.5 0.0 0.5
y2 4
1.0 0.9 0.8 0.7 0.6 1.0 0.5 0.0 1.0 0.5 0.5 0.0 0.5 1.0 1.0
1.0
P00(x, y, 0)
1.0 0.5 0.0 0.5
1.0 0.9 0.8 0.7 0.6 1.0 0.5 0.0 1.0 0.5 0.5 0.0 0.5 1.0
1.0
0.25 0.20 0.15 0.10 0.05 0.00 1.0 0.5 0.0 1.0 0.5 0.5 0.0 0.5 1.0
1.0
x 2y 2 16
1.0
x 2y 2 16
0.25 0.20 0.15 0.10 0.05 0.00 1.0 0.5 0.0 1.0 0.5 0.5 0.0 0.5 1.0
P10(x, y, 0)
0.20 0.15 0.10 0.05 0.00 1.0 0.5 0.0 1.0 0.5 0.5 0.0 0.5 1.0
P01(x, y, )
0.20 0.15 0.10 0.05 0.00 1.0 0.5 0.0 1.0 0.5 0.5 0.0 0.5 1.0 1.0
1.0
x2 4
P01(x, y, 0)
P00(x, y, )
1.0 0.5 0.0 0.5
x 2y 2 16
1.0
1.0
U (z′ ; θ) = V (θL ) ◦ Φ(z′ ) ◦ V (θL−1 ) ◦ Φ(z′ ) ◦ · · ·
0.06 0.04 0.02 0.00 1.0 0.5 0.0 0.5 1.0
◦V (θ1 ) ◦ Φ(z′ ) ◦ V (θ0 )
0.04 0.02 0.00 1.0 0.5 0.0 1.0 0.5 0.5 0.0 0.5 1.0
1.0
0.20 0.15 0.10 0.05 0.00 1.0 0.5 0.0 0.5 1.0
P11(x, y, )
0.20 0.15 0.10 0.05 0.00 1.0 0.5 0.0 1.0 0.5 0.5 0.0 0.5 1.0 1.0
(24)
1) Strongly entangling layer: For the ℓ-th layer, V (θℓ ) = Uent · R(θℓ ), with the order suggested by Observation 1 decomposes into two parts: • Local rotations on each qubit:
P11(x, y, 0)
P10(x, y, )
1.0 0.9 0.8 0.7 0.6 1.0 0.5 0.0 1.0 0.5 0.5 0.0 0.5 1.0 1.0
layers. The architecture alternates encoding layers Φ(z′ ) and strongly entangling layers V (θℓ )
R(θℓ ) =
m0 O
(1)
(2)
(3)
RX (θℓ,k )RY (θℓ,k )RZ (θℓ,k )
(25)
k=1
0.04 0.02 0.00 1.0 0.5 0.0 0.5 1.0
•
Entangling CNOT gates in a cascade pattern: Uent =
mY 0 −1
CN OTi,(i+r) mod m0 .
(26)
i=0
Fig. 2. The measurement probabilities of |ψ(x, y, α)⟩ in the computational basis states |ij⟩ , i, j ∈ {0, 1}.
with CN OTi,j |q0 . . . qi . . . qj . . . qm0 −1 ⟩ → |q0 . . . qi . . . (qj ⊕ qi ) . . . qm0 −1 ⟩
|ψ⟩ =
0 −1 2m X
0 −1 2m X
αi |i⟩ ,
i=0
|αi |2 = 1
(20)
m0
F :R →R
F (z) = z
′
(21)
which augments the d components of z to m0 components of z′ (padding encoding) with ′
m0
z = (z1 , z2 , · · · , zd , c0 , . . . , cm0 −d−1 ) ∈ R and m0
Φ:R
→ H,
and r is a hyperparameter called the range which defines the distance between the control qubit i and the target qubit j.
i=0
The encoding of classical data is achieved by the mapping: d
(27)
′
Φ(z ) =
m0 O
RY (z′k ) |ψ⟩
(22)
k=1
0
RY
Rot
1
RY
Rot
2
RY
Rot
3
RY
Rot
4
RY
Rot
5
RY
Rot
6
RY
Rot
7
RY
Rot
Fig. 3. Architecture of a single layer of the VQC. Rot denotes (1) (2) (3) RX (θℓ,k )RY (θℓ,k )RZ (θℓ,k ).
−i θ2 Y
is the rotation operator about the Ywhere RY (θ) = e axis, whose matrix representation in the computational basis is: RY (θ) =
cos(θ/2) sin(θ/2)
− sin(θ/2) cos(θ/2)
(23)
It can be noted that the coefficients c0 , . . . , cm0 −d−1 may be constants, or alternatively that z′ = z⊗m0 for a suitably chosen integer m0 , as detailed in reference [9] (tensor product encoding). Furthermore, the components of z′ may also depend on the zi (non-linear feature encoding in Fig. 3); in this case, it is shown that this can improve learning (Table I). B. Architecture of the parameterized circuit The quantum circuit implements a parameterized unitary transformation U (z′ ; θ) : H → H. Let L be the number of
C. Measurement and probability distribution Measurement in the computational basis yields a probability m0 m0 vector p(z′ ; θ) ∈ R2 belonging to the simplex ∆2 −1 Each component is expressed as: pj (z′ ; θ) = Tr (|j⟩ ⟨j| ρ(z′ ; θ))
(28)
where ρ(z′ ; θ) = U (z′ ; θ) |0⊗m0 ⟩ ⟨0⊗m0 | U † (z′ ; θ) is the density matrix of the final state. Thus, the circuit U (z′ ; θ) constructs a family of functions of the form (28), which, according to the universal approximation theorem [27], can be uniformly approximated on compact sets by single-qubit gates and a CNOT gate, provided the depth L is sufficient. A classical post-processing (an affine transformation) m0 projects the probability space R2 onto the target space Rm fΘ (z) = Wp(F (z); θ) + b
(29)
m0
with Θ = {θ = (θ0 , · · · , θL ), W ∈ Rm×2 , b ∈ Rm }. The cost function (2) admits a variational interpretation as the expectation of the reconstruction error under the empirical data distribution: (30) L(Θ) = E(z,x)∼D ∥x − (W p(F (z); θ) + b)∥22 D. Gradient computation The gradient with respect to the classical parameters (W, b) is computed via standard automatic differentiation. For the quantum parameters θ, we use the parameter-shift rule ∂pj 1 pj (θi + π2 ) − pj (θi − π2 ) = ∂θi 2
(31)
This property follows from the structure of rotation gates and enables an exact analytical computation of the gradients. Optimization is performed using the Adam (Adaptive Moment Estimation) algorithm. V. R ESULTS ON SIMULATION A. Synthetic dataset To evaluate our HVQC to reconstruct images from input variables, We consider two synthetic dataset (Fig. 4 and 5) composed of N = 900 samples (z(i) , I˜(i) ). Each input variable (i) (i) z(i) = (z1 , z2 ) ∈ R2 parametrizes an image I˜(i) defined over a discrete spatial grid. The output image has resolution 16 × 16. a) Dataset 1: The spatial grid is given by two uniform subdivisions xk and yl of [−2, 2] into 16 points each. Latent variables are sampled from the uniform distribution on [−1, 1]2 . For each z = (z1 , z2 ), two vectors are defined on the grid v1 (k) = sin(z1 gk ) + cos(z2 gk ), v2 (l) = cos(z1 gl ) + sin(z2 gl ). (32) The image is constructed as |Iz (xk , yl )| I˜z (xk , yl ) = P , Iz (xk , yl ) = v1 (k)v2 (l). k,l |Iz (xk , yl )| (33) b) Dataset 2: Input variables are sampled from [−2, 2]2 . For a given z = (z1 , z2 ), three spatial patterns are defined: (x − z1 )2 + (y − z2 )2 h1 (x, y) = exp − , (34) 2 h2 (x, y) = sin(z1 x + z2 y), (35) ! p (r − z12 + z22 /2)2 h3 (x, y) = exp − , (36) 0.5 p where r = x2 + y 2 . The resulting image is given by the mixture
c) Dataset 3: The Friedman1 dataset is a standard synthetic regression benchmark defined by y = 10 sin(πx1 x2 ) + 20(x3 −0.5)2 +10x4 +5x5 +ε, with 5 input features uniformly drawn from [0, 1] and additive gaussian noise ε ∼ N (0, 1). B. Experimental setup All experiments were conducted on a simulated quantum environment using PennyLane, with classical components implemented in Tensorflow. The training procedure followed a supervised learning framework as defined, with the Adam optimizer and a learning rate of 10−2 . For the experiment, we aim to evaluate the impact 3 of the mappings z 7→ F1 (z) = z⊗ , z 7→ F2 (z) where F2 (z) = z′(i) ∈ R8 is defined by F2 (z1 , z2 ) = (z1 , z2 , z1 z2 , z2 z1 , z12 , z22 , z12 z2 , z22 z1 ), and z 7→ F3 (z) where F3 (z) = z′(i) ∈ R8 is defined by F3 (z1 , z2 ) = (z1 , z2 , 0, 0, 0, 0, 0, 0). To benchmark the HVQC with 66,032 trainable parameters (240 quantum, 65,792 classical), we selected four classical regression models: GPR with an RBF kernel, alpha 0.1, 10 optimizer restarts, and target normalization ; RFR with 100 estimators and default unlimited depth ; and XGB with 100 estimators, maximum depth 6, and learning rate 0.1 ; and two feedforward neural networks with identical architectures (5 hidden layers, 67,280 parameters) but different activation functions (ReLU and SiLU), trained with the Adam optimizer for up to 1000 epochs. As noted in Section III, all experiments are conducted using statevector simulation, without shot noise. The impact of shot noise and hardware noise on regression performance is left for future work.
15.0 12.5 10.0 7.5 5.0 2.5 0.0
15.0 12.5 10.0 7.5 5.0 2.5 0.0
0
0
5
5
10
10
15
15
15.0 12.5 10.0 7.5 5.0 2.5 0.0
15.0 12.5 10.0 7.5 5.0 2.5 0.0
0
0
5
5
10
10
15
15
15.0 12.5 10.0 7.5 5.0 2.5 0.0
15.0 12.5 10.0 7.5 5.0 2.5 0.0
0
0
5
5
10
10
15
15
15.0 12.5 10.0 7.5 5.0 2.5 0.0
15.0 12.5 10.0 7.5 5.0 2.5 0.0
0
0
5
5
10
10
15
15
15.0 12.5 10.0 7.5 5.0 2.5 0.0
15.0 12.5 10.0 7.5 5.0 2.5 0.0
0
0
5
5
10
10
15
15
15.0 12.5 10.0 7.5 5.0 2.5 0.0
15.0 12.5 10.0 7.5 5.0 2.5 0.0
0
5
10
15
0
5
10
15
Fig. 4. Dataset 1: Reconstruction of 16×16 maps by the variational quantum circuit whose architecture is described in (24) with number of layers L = 9.
15.0 12.5 10.0 7.5 5.0 2.5 0.0
15.0 12.5 10.0 7.5 5.0 2.5 0.0
0
0
5
5
10
10
15
15
15.0 12.5 10.0 7.5 5.0 2.5 0.0
15.0 12.5 10.0 7.5 5.0 2.5 0.0
0
0
5
5
10
10
15
15
15.0 12.5 10.0 7.5 5.0 2.5 0.0
15.0 12.5 10.0 7.5 5.0 2.5 0.0
0
0
5
5
10
10
15
15
15.0 12.5 10.0 7.5 5.0 2.5 0.0
15.0 12.5 10.0 7.5 5.0 2.5 0.0
0
0
5
5
10
10
15
15
15.0 12.5 10.0 7.5 5.0 2.5 0.0
15.0 12.5 10.0 7.5 5.0 2.5 0.0
0
0
5
5
10
10
15
15
15.0 12.5 10.0 7.5 5.0 2.5 0.0
15.0 12.5 10.0 7.5 5.0 2.5 0.0
0
5
10
15
0
5
10
15
Fig. 5. Dataset 2:Reconstruction of 16 × 16 maps by the variational quantum circuit whose architecture is described in (24) with number of layers L = 9.
Iz (x, y) = αh1 (x, y)+βh2 (x, y)+(1−α−β)h3 (x, y), (37)
C. Impact of feature map, ablation study on hybrid architecture, and benchmark against classical models
with α = 0.5 + 0.5 tanh(z1 ) and β = 0.5 + 0.5 tanh(z2 ). As before, the image is normalized.
The results in Table I reveal that the choice of feature map critically influences the predictive performance of the
0.000200
F1 map F3 map F2 map
0.000175 0.000150 6
loss
0.000125
1e 7
5 4
0.000100
3 2
0.000075
1 0
0.000050
200
400
600
800
1000
0.000025 0.000000 0
200
400
600
800
iterations
1000
1200
1400
1600
Fig. 6. Learning curve for map reconstruction by the variational quantum 3 circuit (24) with z′ = z⊗ in blue and z′ = F1 (z) in red (L = 9) for dataset 1. F1 map F3 map F2 map
0.00020
TABLE II A BLATION STUDY: R2 SCORES ON DATASET 1 ( LEFT ) AND DATASET 2 ( RIGHT ) FOR THREE CONFIGURATIONS : VQC- ONLY ( QUANTUM CIRCUIT WITHOUT AFFINE LAYER ), A FFINE - ONLY ( CLASSICAL AFFINE LAYER APPLIED DIRECTLY TO THE ENCODED INPUT, WITHOUT QUANTUM PROCESSING ), AND F ULL HVQC. FM DENOTES FEATURE MAP.
1e 6
0.00015
2.5
Configuration
2.0
loss
to 0.325 (F3, Dataset 2). Conversely, removing the affine layer (VQC-only) yields negative R2 values, confirming that the final linear projection is equally indispensable to lift the probability simplex constraint. Compared to classical baselines, the HVQC with F3 performs on par with Gaussian Process Regression. Random Forest achieves good performance on Dataset 1 but degrades slightly on Dataset 2, while XGBoost consistently lags behind the other methods. HVQC achieves the best performance (0.938) on the Friedman1 benchmark compared to XGB, RFR, and GPR (Table III).
1.5
0.00010
1.0 0.5 0.0
200
400
600
800
VQC-only
1000
0.00005
Affine-only
0.00000 0
200
400
600
800
iterations
1000
1200
1400
1600
Fig. 7. Learning curve for map reconstruction by the variational quantum 3 circuit (24) with z′ = z⊗ in blue and z′ = F1 (z) in red (L = 9) for dataset 2.
HVQC. Feature map F1 contains redundant components, which degrades its performance. Feature map F2 shows mixed results, performing well on some datasets but poorly on others. In contrast, feature map F3 which retains only the original components and pads the remaining entries with zeros consistently achieves strong predictive performance. Although the affine layer dominates the parameter count (65,792 vs. 240 quantum parameters), the ablation study (Table II) shows that removing the quantum circuit reduces the test R2 from 0.978
TABLE I C OMPARISON OF PREDICTION PERFORMANCE FOR HVQC MODELS WITH DIFFERENT FEATURE MAPS AND CLASSICAL (GPR, RFR AND XGB) ACROSS TWO DATASETS , USING MSE AND R2 METRICS ON TRAIN AND TEST SPLITS ( m0 = 2 QUBITS AND L = 9 LAYERS ). Models HVQC with F 1 HVQC with F 2 HVQC with F 3 GPR RFR XGB
FM
Metrics mse(×10−7 ) R2 mse(×10−7 ) R2 mse(×10−7 ) R2 mse(10−7 ) R2 mse(10−7 ) R2 mse(10−7 ) R2
Dataset 1 train test 2.55 3.06 0.970 0.965 0.218 0.336 0.997 0.996 0.609 0.745 0.991 0.990 0.605 0.788 0.993 0.991 0.102 0.782 0.998 0.995 0.740 1.797 0.990 0.979
Dataset 2 train test 5.151 14.398 0.909 0.762 1.737 10.262 0.970 0.843 0.701 1.082 0.986 0.978 0.683 1.075 0.986 0.978 0.238 1.654 0.991 0.968 1.944 4.751 0.960 0.905
Full HVQC
F1 F2 F3 F1 F2 F3 F1 F2 F3
R2 (dataset 1) Train Test −0.747 −0.692 −1.373 −1.387 −0.372 −0.348 0.520 0.532 0.858 0.857 0.638 0.664 0.970 0.965 0.997 0.996 0.991 0.990
R2 (dataset 2) Train Test −1.881 −2.053 −2.244 −2.409 −1.383 −1.512 0.075 0.151 0.703 0.651 0.364 0.325 0.909 0.762 0.970 0.843 0.986 0.978
Two classical neural networks (ReLU and SiLU), with 67,280 parameters, were trained for comparison. At 100 epochs, they achieve slightly lower performance: for dataset 1, R2 ≈ 0.994 (both); for dataset 2, R2 ≈ 0.975 (SiLU) and 0.972 (ReLU). After 1000 epochs, both reach R2 ≈ 0.999 (dataset 1) and 0.993 (dataset 2), matching HVQC performance. TABLE III XGB, RFR AND GPR ARE EVALUATED ON THE F RIEDMAN 1 DATASET USING 20 INDEPENDENT SPLITS (200 TRAIN / 40 568 TEST ) WITH 10- FOLD CROSS - VALIDATION FOR HYPERPARAMETER SELECTION . F OR HVQC ( m0 = 5 QUBITS AND L = 3 LAYERS , WITHOUT FEATURE MAP ). Method HVQC XGB RFR GPR
R2 on test set 0.938 0.858 0.790 0.934
D. Global analysis of quantum state after measure and implicit clustering behavior Although the significant states (Fig. 8) obtained immediately after quantum measurements on all VQC output qubits, representing the hottest pixels, do not form sharply separated clusters according to classical metrics, they reveal a clear overall tendency to group images by category (Fig. 9). Thus, the VQC implicitly reflects the distribution of thresholdactivated elements, facilitating a structured interpretation of the identification of significant pixels across the entire training set of dataset 1.
250
1.00 0.75
200 Significant quantum states
0.50 0.25
150
0.00
100
0.25 0.50
50
0.75 1.00 1.00
0.75
0.50
0.25
0.00
0.25
0.50
0.75
1.00
0
Fig. 8. Emergent activation patterns from HVQC output probabilities before post-processing, on training set of dataset 1. 0
2
4
6
8
14
20
23
24
25
32
38
40
42
72
87
91
97
98
107
119
120
128
129
136
137
139
145
151
158
166
171
181
182
183
187
192
194
201
231
234
236
237
239
245
247
253
254
Fig. 9. Sample images from each of 48 clusters of significant quantum state, on training set of dataset 1.
VI. C ONCLUSION We proposed a hybrid variational quantum circuit (HVQC) for multivariate regression, combining a parameterized quantum circuit with a classical affine post-measurement layer that lifts the probability simplex constraint. An ablation study confirms that both the quantum and classical components are essential, and that the architecture follows naturally from elementary circuits via data re-uploading, entanglement, and affine post-processing. Experimentally, the HVQC matches Gaussian Process Regression on synthetic image datasets and achieves R2 = 0.938 on the larger Friedman1 benchmark (40,568 test samples), outperforming XGB and RFR. Results highlight the central role of the feature map, whose expressivity proves more influential. Future work includes validation on real-world datasets, analysis of shot noise and hardware effects, and principled feature map selection strategies. These results suggest that HVQC architectures are a viable alternative to classical methods for structured regression tasks, particularly when the output space is high-dimensional. R EFERENCES [1] J. Preskill, “Quantum computing in the nisq era and beyond,” Quantum, vol. 2, p. 79, 2018. [2] A. Peruzzo, J. McClean, P. Shadbolt, M.-H. Yung, X.-Q. Zhou, P. J. Love, A. Aspuru-Guzik, and J. L. O’brien, “A variational eigenvalue solver on a photonic quantum processor,” Nature communications, vol. 5, no. 1, p. 4213, 2014. [3] J. Biamonte, P. Wittek, N. Pancotti, P. Rebentrost, N. Wiebe, and S. Lloyd, “Quantum machine learning,” Nature, vol. 549, no. 7671, pp. 195–202, 2017.
[4] M. Cerezo, A. Arrasmith, R. Babbush, S. C. Benjamin, S. Endo, K. Fujii, J. R. McClean, K. Mitarai, X. Yuan, L. Cincio et al., “Variational quantum algorithms,” Nature Reviews Physics, vol. 3, no. 9, pp. 625– 644, 2021. [5] M. Benedetti, E. Lloyd, S. Sack, and M. Fiorentini, “Parameterized quantum circuits as machine learning models,” Quantum science and technology, vol. 4, no. 4, p. 043001, 2019. [6] M. Schuld, R. Sweke, and J. J. Meyer, “Effect of data encoding on the expressive power of variational quantum-machine-learning models,” Physical Review A, vol. 103, no. 3, p. 032430, 2021. [7] M. Schuld and N. Killoran, “Quantum machine learning in feature hilbert spaces,” Physical review letters, vol. 122, no. 4, p. 040504, 2019. [8] A. Pérez-Salinas, A. Cervera-Lierta, E. Gil-Fuster, and J. I. Latorre, “Data re-uploading for a universal quantum classifier,” Quantum, vol. 4, p. 226, 2020. [9] M. Schuld, A. Bocharov, K. M. Svore, and N. Wiebe, “Circuit-centric quantum classifiers,” Physical Review A, vol. 101, no. 3, p. 032308, 2020. [10] S. Lloyd, M. Schuld, A. Ijaz, J. Izaac, and N. Killoran, “Quantum embeddings for machine learning,” 2020, unpublished. [11] E. Garate, P. San Sebastian, G. Valverde, A. Ruiz, and M. Gómez, “Variational quantum regression on nisq hardware with error mitigation,” in 2024 Artificial Intelligence Revolutions (AIR). IEEE, 2024, pp. 32– 39. [12] R. Orús, S. Mugel, and E. Lizaso, “Quantum computing for finance: Overview and prospects,” Reviews in Physics, vol. 4, p. 100028, 2019. [13] C. Ciliberto, M. Herbster, A. D. Ialongo, M. Pontil, A. Rocchetto, S. Severini, and L. Wossnig, “Quantum machine learning: a classical perspective,” Proceedings of the Royal Society A: Mathematical, Physical and Engineering Sciences, vol. 474, no. 2209, p. 20170551, 2018. [14] M. H. F. da Silva, G. F. de Jesus, C. Nascimento, V. L. da Silva, and C. S. Cruz, “Exploring quantum machine learning for weather forecasting,” Brazilian Journal of Physics, vol. 56, no. 1, p. 22, 2026. [15] J. G. M. de Lejarza and S. Shum, “Multidimensional regression with a variational quantum circuit,” 2024, unpublished. [16] P. Senapati, S. Y.-C. Chen, B. Fang, T. M. Athawale, A. Li, W. Jiang, C. C. Lu, and Q. Guan, “Pqml: Enabling the predictive reproducibility on nisq machines for quantum ml applications,” in 2024 IEEE International Conference on Quantum Computing and Engineering (QCE), vol. 1. IEEE, 2024, pp. 1413–1424. [17] F. Perkkola, I. Salmeperä, A. Meijer-van de Griend, C.-C. J. Wang, R. S. Bennink, and J. K. Nurminen, “Optimizing state preparation for variational quantum regression on nisq hardware,” in 2025 IEEE International Conference on Quantum Computing and Engineering (QCE), vol. 1. IEEE, 2025, pp. 302–311. [18] W. Hamhoum, S. Cherkaoui, J.-F. Laprade, O. Ahmed, and S. Wang, “Multivariate time series forecasting with gate-based quantum reservoir computing on nisq hardware,” 2025, unpublished. [19] P. Czarnik, A. Arrasmith, P. J. Coles, and L. Cincio, “Error mitigation with clifford quantum-circuit data,” Quantum, vol. 5, p. 592, 2021. [20] F. M. Neto, L. d. R. Silva, P. S. Neto, and F. F. Fanchini, “Regression of functions by quantum neural networks circuits,” 2025, unpublished. [21] J. Joo, “Advancing quantum computation: optimizing algorithms and error mitigation in nisq devices,” Frontiers in Physics, vol. 14, p. 1788075, 2026. [22] J. R. McClean, S. Boixo, V. N. Smelyanskiy, R. Babbush, and H. Neven, “Barren plateaus in quantum neural network training landscapes,” Nature communications, vol. 9, no. 1, p. 4812, 2018. [23] J. Cunningham and J. Zhuang, “Investigating and mitigating barren plateaus in variational quantum circuits: a survey: J. cunningham, j. zhuang,” Quantum Information Processing, vol. 24, no. 2, p. 48, 2025. [24] V. Havlíček, A. D. Córcoles, K. Temme, A. W. Harrow, A. Kandala, J. M. Chow, and J. M. Gambetta, “Supervised learning with quantumenhanced feature spaces,” Nature, vol. 567, no. 7747, pp. 209–212, 2019. [25] M. Schuld, “Supervised quantum machine learning models are kernel methods,” arXiv preprint arXiv:2101.11020, 2021. [26] C. Wilson, J. Otterbach, N. Tezak, R. S. Smith, A. Polloreno, P. J. Karalekas, S. Heidel, M. S. Alam, G. Crooks, and M. Da Silva, “Quantum kitchen sinks: An algorithm for machine learning on nearterm quantum computers,” 2018, unpublished. [27] J.-L. Brylinski and R. Brylinski, “Universal quantum gates,” in Mathematics of quantum computation. Chapman and Hall/CRC, 2002, pp. 117–134.