ConceptioArchivearXiv CS
arXiv CSopen access

Kernel-based Operator Learning: Error Analysis, Budget Allocation, and a Physics-Informed Extension

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

Kernel-based Operator Learning

Kernel-based Operator Learning: Error Analysis, Budget Allocation, and a Physics-Informed Extension Rüdiger Kempf

[email protected]

Applied and Numerical Analysis University of Bayreuth 95440 Bayreuth, Germany

Editor:

arXiv:2607.06287v1 [math.NA] 7 Jul 2026

Abstract We study kernel-based operator learning in a two-stage sampling framework, where an offline kernel regression operator learns a discretized representation of the target operator from input-output pairs and an online kernel reconstruction operator recovers the output function from predicted observations. Our main theoretical contribution is an explicit budget allocation condition relating the number N of training pairs, the number n of input observations, and the output resolution m. The condition is derived from a coupled error analysis that interprets the surrogate as a reconstruction from approximate data. This yields a decomposition of the total error into reconstruction and learning contributions that can be analyzed independently. As a consequence, we obtain quantitative scaling laws describing how N , n, and m must be coupled to guarantee convergence and to balance offline learning and online reconstruction errors. The resulting estimates extend previous analyses of kernel-based operator learning. We further introduce a physics-informed extension that incorporates knowledge of the underlying PDE at evaluation time. Rather than encoding constraints directly into the kernel, we augment the online reconstruction step by penalizing PDE residuals at collocation points. The method requires no retraining for new inputs. Numerical experiments illustrate the theoretical findings and demonstrate the effectiveness of the proposed physics-informed reconstruction strategy. Keywords: kernel methods, operator learning, budget allocation, physics-informed learning

1 Introduction Operator learning has emerged as a central problem in scientific machine learning: given a finite collection of input-output pairs generated by an unknown operator G : U → V between function spaces, the goal is to construct a surrogate A ≈ G that generalizes to unseen inputs. The canonical motivation is the solution operator of a PDE, where G maps a coefficient or forcing function to the corresponding solution, and one wishes to replace expensive repeated PDE solves with a fast surrogate at deployment time. The dominant paradigm for operator learning is neural-network-based. Architectures such as DeepONet (Lu et al., 2021, 2022; Wang et al., 2021, 2022) and the Fourier Neural Operator (Li et al., 2021) have achieved impressive empirical performance on a range of benchmarks. However, neural-network approaches offer limited theoretical guarantees: 1

R. Kempf

convergence rates, sample complexity, and the influence of design parameters such as training set size and output resolution are difficult to characterize rigorously. Kernel methods, by contrast, come with a well-developed approximation theory (Wendland, 2004; Schaback and Wendland, 2006; Owhadi and Scovel, 2019; Schölkopf and Smola, 2001), deterministic error bounds (Narcowich et al., 2005; Arcangeli et al., 2012; Duchon, 1978), and transparent dependence on all design parameters, such as the correlation length of the kernel (Wendland and Rieger, 2005; Sun and Wang, 2026). While classical kernel methods are well understood in the finite-dimensional function approximation setting, their extension to operator learning is more recent and mainly of numerical nature (Kadri et al., 2016; Owhadi, 2023; Batlle et al., 2024; Sharma et al., 2026; Nelsen and Stuart, 2024). This paper contributes to the rigorous, approximation-theoretic foundation of kernel-based operator learning. The framework we study builds on the two-stage approach of Batlle et al. (2024); Sharma et al. (2026), in which an offline kernel regression operator Aoff learns a discretized version of G from training pairs, and an online kernel reconstruction operator Aon recovers the output function from the predicted observations. Related kernel-based approaches include Long et al. (2024), who use a three-step framework to first learn the PDE form from data and then solve it with a kernel solver, and Mora et al. (2025), who propose a hybrid GP/NN framework approximating the operator’s associated bilinear form. Our approach differs from both: we work directly within the two-stage sampling structure of Batlle et al. (2024) and focus on a self-contained, interpretable error analysis. Our main theoretical contribution is an explicit budget allocation condition relating the number N of training pairs, the number n of input observations, and the output resolution m. The condition is derived from a new coupled error analysis of kernel-based operator learning. By interpreting the surrogate as a reconstruction from approximate data, we decompose the total error into reconstruction and learning components and analyze them separately. This yields quantitative scaling laws that determine how training effort must increase as the online discretization is refined, extending the analyses of Batlle et al. (2024); Sharma et al. (2026). Finally, we augment the framework to incorporate knowledge of the underlying PDE at evaluation time. In the kernel setting, one classical approach encodes physical constraints directly into the kernel KV used for Aon : constructing KV whose native space is contained in a constrained function space, such as divergence-free or curl-free vector fields (Narcowich and Ward, 1994; Fuselier, 2008; Narcowich et al., 2007), ensures every surrogate output satisfies the constraint by construction. This hard-constraint approach was recently used for operator learning in Sharma et al. (2026). However, it requires the physical constraint to be independent of the parameter u and analytically identifiable at kernel construction time, ruling out parameter-dependent operators such as Lu = −div(eu ∇·). We therefore follow a soft-constraint approach in the spirit of kernel collocation (Kansa, 1990; Schaback, 2009; Fasshauer and McCourt, 2014): we penalize the PDE residual at a finite set of collocation points by augmenting the Tikhonov functional with a weighted residual term. Unlike classical kernel collocation, which solves a PDE from scratch, our method augments an already-trained operator learning surrogate at evaluation time, using the collocation penalty only in the online reconstruction step Aon . No PDE solver is invoked, and the method adapts to new parameter inputs u without retraining, in contrast to the training-time physics-informed approach of Wang et al. (2021). The resulting reconstruction 2

Kernel-based Operator Learning

operator APDE admits a closed-form representer theorem via the framework of Micchelli and on Pontil (2005). We also analyse the numerical costs for the standard and the physics-informed learning method. The rest of the paper is structured as follows. Section 2 introduces the operator learning framework and its generalizations over Batlle et al. (2024), with particular attention to the existence of the discretized operator g. Section 3 recalls the relevant kernel approximation theory, including sampling inequalities. The error analysis is developed in Section 4, covering both interpolation and regularized least-squares data generation, and yields the budget allocation condition. The physics-informed extension is treated in Section 5. Numerical experiments validating the theory are presented in Section 6, and Section 7 gives a brief conclusion.

2 The Two-Stage Operator Learning Framework We start with giving the precise setting and fixing notation used throughout. Compared to Batlle et al. (2024) we study operator learning where input and output functions are vector fields, similar to Sharma et al. (2026). Furthermore, we discuss the connection of the fill distance of the sampled input functions and the fill distance of the data we actually use for learning the surrogate. This also yields first results on the existence of the discretized operator g. 2.1 The Setting Let U and V be Banach spaces and let G : U → V be an operator. The goal of operator learning is to learn G from given N -many input/output pairs, i.e., to find a surrogate N A : U → V for G using only the training data {(ui , vi )}N i=1 := {(ui , G(ui ))}i=1 ⊆ U × V. In our framework, the input/output functions are also only known by discrete observations, modeled by sampling operators: Let Φ = {ϕ1 , . . . , ϕn } and Ψ := {ψ1 , . . . , ψm } be sets of bounded linear observational mappings on U and V, respectively. Define the sampling operators SΦ : U → Rn and SΨ : V → Rm by SΦ (u) = [ϕ1 (u), . . . , ϕn (u)]T

and

SΨ (v) = [ψ1 (v), . . . , ψm (v)]T .

With this, we can formalize the operator learning task the following way: Let {(ui , vi )}N i=1 ∈ U × V be such that vi = G(ui ),

1 ≤ i ≤ N.

and let SΦ : U → Rn and SΨ : V → Rm be sampling operators. Then the operator learning task is to find a surrogate A for G using only the data N {(SΦ (ui ), SΨ (vi ))}N i=1 := {(ui , v i )}i=1 .

Of particular interest, cf. Batlle et al. (2024), are operators G that are connected to (non-linear) partial differential operators, e.g., coefficient-to-solution operators. In this case, U and V are Banach spaces of functions u : Ω → Rp and v : D → Rℓ respectively, where 3

R. Kempf

G

U

Ale

SΦ Rnp

V

g Aoff

Aon Rmℓ

Figure 1: The fully symmetric learning approach of Batlle et al. (2024): To learn the operator G, employ the two sampling operators SΦ and SΨ and three reconstructions Ale , Aoff and Aon . The discretized operator g is give as SΨ ◦ G ◦ Ale . The learned surrogate is then given as A = Aon ◦ Aoff ◦ SΦ . Ω ⊆ Rk and D ⊆ Rd are non-empty domains. Ω and D and k and d do not necessarily need to coincide. The sampling operators SΦ and SΨ are assumed to be point evaluation operators, i.e., we assume there are sets of collocation points X = {x1 , . . . , xn } ⊆ Ω and Y := {y 1 , . . . , y m } ⊆ D and define SX := SΦ and SY := SΨ by 

 u(x1 )  u(x2 )    SX (u) := u :=  .  ∈ Rnp , .  . 

and

 v(y 1 )  v(y )  2   SY (v) := v :=  .  ∈ Rmℓ , .  . 

u(xn )

v(y m )

as point-evaluation mappings, cf. Sharma et al. (2026). Our methodology follows a two-stage approach. We assume the existence of a discretized operator g : Rnp → Rmℓ satisfying g(SX (u)) = SY (G(u)) for all u ∈ U. In practice, we will restrict to a compact model calss M ⊆ U , see Theorem 23. In an offline phase, we learn a np surrogate Aoff ≈ g from the training data {(ui , v i )}N i=1 by kernel regression on R . In the mℓ online phase, for a new input u, a kernel reconstruction operator Aon : R → V recovers the output function from the predicted observations Aoff (SX (u)) ≈ SY (G(u)), giving the surrogate A = Aon ◦ Aoff ◦ SX : U → V. The key insight is to view Aon as a reconstruction from perturbed data: the exact discretized output v † = SY (G(u)) is unknown, and Aoff (SX (u)) is only an approximation to it. This yields a clean decomposition of the total error into a reconstruction error, measuring how well Aon recovers G(u) from exact data v † , and a learning error, measuring how well Aoff approximates g. The two contributions can be analyzed independently, which is the basis of the error analysis in Section 4. The existence and properties of g are discussed in Sections 2.4 and 4.3. This differs from Batlle et al. (2024), where an additional reconstruction operator Ale : np R → U is assumed to exist and g is expressed as SY ◦ G ◦ Ale . Further differences are discussed in Sections 2.2 and 2.3. 4

Kernel-based Operator Learning

2.2 Observations about the Established Framework The operator learning framework described above is general, but several assumptions implicit in the existing literature (Batlle et al., 2024; Sharma et al., 2026) can be relaxed without sacrificing the error analysis. We identify three such generalizations, each of which broadens the applicability of the framework. Observation 1 (Symmetry of the formulation) The diagram in Fig. 1 is symmetric in U and V: it assumes the simultaneous existence of reconstruction operators Ale : Rnp → U and Aon : Rmℓ → V. This is unnecessarily restrictive. In our framework, Ale plays no role in either the surrogate construction or the error analysis: the surrogate A = Aon ◦ Aoff ◦ SX requires only a reconstruction on the V-side. Consequently, U need not carry any additional structure, it suffices for U to be a Banach space supporting bounded point evaluations. This asymmetry also opens the door to treating inverse problems, where the roles of U and V are exchanged. Observation 2 (Condition 3.2 in Batlle et al. (2024)) The framework of Batlle et al. (2024) assumes that the training inputs satisfy ui = Ale (SΦ (ui )) for all 1 ≤ i ≤ N , i.e., that the training inputs are exactly reconstructible from their observations. While natural in their setting, this assumption ties the training procedure to a specific reconstruction operator Ale that must be known prior to data collection. Our framework avoids this assumption entirely, since Ale does not appear in our surrogate or error analysis. Observation 3 (Fill distance in U) For the error analysis, it is necessary to assume np are sufficiently rich, typically expressed that the discrete training inputs {SX (ui )}N i=1 ⊆ R via a small fill distance. However, this condition on {SX (ui )} carries limited information about the original functions {ui } ⊆ U. The sampling operator SX is in general not injective: distinct functions u, u′ ∈ U may satisfy SX (u) = SX (u′ ), so fill distance in Rnp does not imply coverage of U . This persists even as the number of observations grows, as the classical example of Runge Runge (1905) shows. 2.3 Proposed Framework G

U

V SY

SX Rnp

Aoff

G −1

U Aon

SX

(a) Forward problem.

SY

Aon Rnp

Rmℓ

V

Aoff

Rmℓ

(b) Inverse problem.

Figure 2: Asymmetric diagrams addressing Observation 1 and Observation 2: only one reconstruction operator is required in each case. It turns out that Observation 1 and Observation 2 can be resolved simultaneously by augmenting the general setting. Instead of assuming the symmetric diagram in Fig. 1, we 5

R. Kempf

propose to separate the forward and inverse problems, as summarized in Fig. 2. In both cases, only a single online reconstruction operator is required, either on the V side for the forward problem, or on the U side for the inverse problem. This asymmetric formulation resolves both observations at once: since Ale no longer appears, the assumption ui = Ale (SΦ (ui )) of Batlle et al. (2024) is not needed, and U need not carry any additional structure. In particular, U may be any Banach space of functions supporting bounded point evaluations. 2.4 Fill Distance and Existence of the Discretized Operator Addressing Observation 3 is more involved. We begin by formally defining the fill distance. Definition 1 (Fill distance) Let (V, ∥ · ∥V ) be a normed space and D ⊆ V a bounded domain. Let Ξ := {ξ 1 , . . . , ξ M } ⊆ D be a discrete set of points. Then the fill distance hΞ,D of Ξ in D is defined as hΞ,D := sup min ∥ξ − ξi ∥V . ξ∈D 1≤i≤M

The following theorem links the fill distance of the training inputs {ui } in U to the fill distance of their discretizations {ui } in Rnp . Theorem 2 Let M ⊆ U be a compact model class and let {u1 , . . . , uN } ⊆ M. Assume that SΦ consists of bounded linear observational mappings and that there exists h0 > 0 such that h{ui },SΦ (M) ≤ h0 . Define ιΦ : M → R≥0 by ιΦ (u) :=

inf

w∈ker SΦ

∥u − w∥U ,

and assume there exists a constant C > 0 such that for all u, u e ∈ M, ∥u − u e∥U ≤ C∥SΦ (u) − SΦ (e u)∥2 + ιΦ (u) + ιΦ (e u).

(1)

Then sup min ∥u − ui ∥U ≤ Ch{ui },SΦ (M) + 2 sup ιΦ (u)

u∈M 1≤i≤N

u∈M

holds for all u ∈ M. Proof Let u ∈ M be fixed. Since M is compact and SΦ consists of continuous mappings, there exists an index 1 ≤ i ≤ N such that ∥SΦ (u) − SΦ (ui )∥2 ≤ h{ui },SΦ (M) . Then (1) yields ∥u − ui ∥U ≤ Ch{ui },SΦ (M) + ιΦ (u) + ιΦ (ui ) ≤ Ch{ui },SΦ (M) + 2 sup ιΦ (u). u∈M

Taking the minimum over i and then the supremum over u ∈ M yields the claim.

6

Kernel-based Operator Learning

The term ιΦ in Theorem 2 quantifies the intrinsic non-identifiability of the sampling operator: functions u and u e satisfying SΦ (u) = SΦ (e u), i.e., u − u e ∈ ker SΦ , are indistinguishable from the observations alone. The kernel ker SΦ can therefore be interpreted as the space of invisible perturbations. In particular, if U is infinite-dimensional, ιΦ will in general not vanish identically on M. In the case where SΦ = SX is a point evaluation operator, ιΦ (u) measures the part of u that is not observed at the points X. Inequality (1) is a Lipschitz stability condition for the inverse sampling map. In the −1 injective case ιΦ ≡ 0, the assumption reduces to SΦ : SΦ (M) → M is C-Lipschitz, i.e., that the inverse problem of recovering u from its observations is well-posed on M. A concrete setting in which (1) is satisfied is discussed in Section 4.3. Remark 3 (Solution for Observation 3) Theorem 2 shows that the fill-distance of {ui } in M decomposes into two contributions: the fill distance of {ui } in SΦ (M), which is directly controllable by the choice of training data, and the intrinsic non-identifiability term supu∈M ιΦ (u), which is an irreducible consequence of the non-injectivity of SΦ . In our error analysis, we therefore work directly with the fill distance of {ui } in SΦ (M) ⊆ Rnp , treating the non-identifiability term as a separate contribution, which we ignore here. We further note that the assumption ιΦ ≡ 0 on M, i.e., that SΦ is effectively injective on the model class, is sufficient for the existence of a well-defined discretized operator g : Rnp → Rmℓ satisfying g(SX (u)) = SY (G(u)) for all u ∈ M.

3 Kernel-based Learning We give a brief introduction to kernel-based learning of vector-valued functions, which forms the theoretical backbone of our two-stage surrogate construction. Recall that the surrogate A = Aon ◦ Aoff ◦ SX involves two distinct kernel-based learning problems: a regression problem for Aoff on Rnp , approximating the discretized operator g, and a reconstruction problem for Aon in V, recovering an output function from disturbed data. The present section develops the tools needed for both. Throughout, we assume d, r ∈ N and D ⊆ Rd is an arbitrary domain. For background on kernel methods we refer to Wendland (2004) for the approximation theory perspective and, for the statistical learning perspective, to Steinwart and Christmann (2008) . 3.1 Matrix-Valued Positive Definite Kernels and RKHSs Definition 4 We call a continuous kernel K : D × D → Rr×r positive (semi-)definite, if 1. K is symmetric, i.e., K(ξ, ζ) = K(ζ, ξ)T , for all ξ, ζ ∈ D, and 2. the quadratic form M X

⟨wi , K(ξ i , ξj )wj ⟩2

(2)

i,j=1

is non-negative for all M ∈ N, all pair-wise distinct {ξ 1 , . . . , ξ M } ⊆ D and all w1 , . . . wM ∈ Rr , not all being the zero vector. 7

R. Kempf

The kernel is called positive definite if the quadratic form (2) is equal to zero only if wi = 0 for indices 1 ≤ i ≤ N where ξ i ̸= ξj , i ̸= j. A key feature of positive definite kernels is that they implicitly define a feature map φ : D → H such that K(ξ, ζ) = ⟨φ(ξ), φ(ζ)⟩H , allowing learning algorithms to operate in a potentially infinite-dimensional feature space H using only kernel evaluations, this is known as the kernel trick (Steinwart and Christmann, 2008). The associated feature space is a Hilbert space of functions with a special structure, called native space of the kernel K or a reproducing kernel Hilbert space (RKHS) with kernel K. Definition 5 A Hilbert space (H, ⟨·, ·⟩H ) of vector-valued functions f : D → Rr is a reproducing kernel Hilbert space if there is a kernel K : D × D → Rr×r that satisfies 1. K(·, ξ)w ∈ H for all ξ ∈ D and w ∈ Rr , 2. ⟨v, K(·, ξ)w⟩H = ⟨v(ξ), w⟩2 for all v ∈ H, w ∈ Rr and ξ ∈ D. The latter is called reproducing property of K. A particularly important class of matrix-valued kernels are separable kernels, which decouple two distinct modeling aspects: the spatial regularity of the function, encoded by scalar-valued positive definite kernels ki : D × D → R, and the interaction between output components, encoded by symmetric positive semidefinite matrices Bi ∈ Rr×r . This separation makes both the theoretical analysis and the computational implementation tractable. For fixed b ∈ N, we define a matrix-valued kernel K : D × D → Rr×r by K(ξ, ζ) :=

b X

ki (ξ, ζ)Bi ,

ξ, ζ ∈ D.

(3)

i=1

In the special case b = 1 and B1 = I, the kernel acts componentwise and the associated RKHS is the tensor product of identical scalar RKHSs. Of particular interest are translation-invariant kernels on Rd . A kernel K : Rd × Rd → r×r R is called translation invariant if there exists a function Φ : Rd → Rr×r such that K(ξ, ζ) := Φ(ξ − ζ),

ξ, ζ ∈ Rd .

In the separable setting (3), this corresponds to choosing translation-invariant kernels ki = Φi . For the error analysis later, we want to identify RKHSs with known function spaces. To this end, we recall the definition of Sobolev spaces of vector-valued functions: For 1 ≤ q ≤ ∞, let Lq (D)r denote the standard Lebesgue space of q-integrable vector-valued functions f : D → Rr , equipped with the norm ∥ · ∥Lq (D)r . For s ∈ N0 , the Sobolev space Wqs (D)r is given by Wqs (D)r := {f ∈ Lq (D)r : ∥f ∥Wqs (D)r < ∞}, 8

Kernel-based Operator Learning

where ∥f ∥Wqs (D)r :=

s X

!1

1

q

|f |qW l (D)r q

, with |f |W l (D)r := 

q

X

q

∥D

α

f ∥qLq (D)r 

.

|α|=l

l=0

Fractional orders s ≥ 0 can be defined by, e.g., interpolation (Brenner and Scott, 2008). In the Hilbert space case q = 2, Sobolev spaces admit a convenient Fourier characterization. For s ≥ 0, we define H s (Rd )r := {f ∈ L2 (Rd )r : (1 + ∥ · ∥2 )s/2 fb ∈ L2 (Rd )r }, equipped with the inner product Z ⟨f, g⟩H s (Rd )r :=

Rd

T

(1 + ∥ω∥2 )s fb(ω) gb(ω) dω,

(4)

where the Fourier transform is applied componentwise. The connection between RKHSs of translation-invariant kernels and Sobolev spaces is well-understood in the scalar-valued case (Wendland, 2004). Corollary 6 Let k : Rd × Rd → R be a RBF associated to Φ ∈ L1 (Rd ). Let s > d/2 and assume that there are constants c1 , c2 > 0 such that b c1 (1 + ∥ω∥22 )−s ≤ Φ(ω) ≤ c2 (1 + ∥ω∥22 )−s ,

ω ∈ Rd .

Then the RKHS H with kernel k coincides with H s (Rd ), with equivalent norms. Prominent classes of kernels satisfying the assumptions of Theorem 6 are Matérn kernels (Matérn, 1986) and compactly supported Wendland kernels (Wendland, 1995, 2004). Using the separable construction (3), this result extends directly to the vector-valued setting. Corollary 7 Let (ki )1≤i≤b and (Bi )1≤i≤b be as in (3) and assume that each ki satisfies the assumptions of Theorem 6 with the same smoothness parameter s > d/2. Then the RKHS associated with the matrix-valued kernel K defined as in (3) coincides with H s (Rd )r . Moreover, the induced norm is equivalent to the Sobolev norm induced by the inner product in (4), with constants depending on c1 , c2 and the spectral bounds of the matrices Bi . This construction shows that separable kernels provide a systematic way to lift scalar Sobolev kernels to vector-valued function spaces, while allowing for flexible modeling of correlations between output components via the matrices Bi . Remark 8 (Restriction to bounded domains) In many applications, functions are defined on a bounded domain D ⊆ Rd . Kernels for Sobolev spaces on D can be obtained by restriction. Let k : Rd × Rd → R be a translation-invariant kernel whose RKHS is H s (Rd ), and define k|D (ξ, ζ) := k(ξ, ζ), ξ, ζ ∈ D. Then the RKHS associated with k|D is norm-equivalent to H s (D), provided that D is, e.g., a bounded Lipschitz domain. An analogous construction applies to matrix-valued kernels via (3). 9

R. Kempf

3.2 Learning Methods We consider two standard learning approaches within the kernel-based setting. Throughout this section, we assume that H is a RKHS with kernel K. Given data at a finite set of training points, we seek an approximation in a finite-dimensional hypothesis space induced by the kernel. Definition 9 For a fixed set of training points Ξ := {ξ 1 , . . . , ξ M } ⊆ D define the hypothesis space or parametric model class VM ⊆ H as (M ) X VM := K(·, ξ i )wi : wi ∈ Rr , 1 ≤ i ≤ M . i=1

If the data is exact, interpolation is the natural choice: it fits the data perfectly and, as we show below, minimizes the H-norm among all interpolants. If the data is disturbed, as is the case for Aon , which receives the inexact output Aoff (SX (u)) ≈ v † , interpolation is undesirable as it overfits the noise. In this case, regularized least-squares introduces a bias controlled by λ > 0 in exchange for improved stability, embodying the classical bias-variance tradeoff familiar from statistical learning theory (Steinwart and Christmann, 2008). 3.2.1 Norm-minimal Interpolation Definition 10 For given training points Ξ := {ξ 1 , . . . , ξ M } ⊆ D and corresponding labels F := {f1 , . . . fM } ⊆ Rr the interpolation problem is given as Find s0 ∈ VM such that s0 (ξ i ) = fi for all 1 ≤ i ≤ M . The solution of the interpolation problem is called the (kernel-) interpolant to F in VM . Kernel-interpolants in RKHSs have many advantageous properties that follow directly from the reproducing property in Theorem 5. Theorem 11 Assume that the labels F are generated by f ∈ H, i.e., fi = f (ξ i ), and let s0 ∈ VM be the interpolant to F . Then sF is the unique interpolant in VM and the orthogonal projection of f onto VM . In particular, ∥s0 ∥H ≤ ∥f ∥H holds. The interpolant s0 defines a linear interpolation operator IΞ : RrM → VM , IΞ ([f ]) = s0 , where [f ] = [f1 , . . . , fM ]T ∈ RrM . If the labels are generated by a function f ∈ C(D)r , this can be extended to an operator IΞ : C(D)r → VM by setting IΞ (f ) = IΞ ([f ]). 3.2.2 Regularized Least-Squares Regression If the data is not exact interpolation is typically undesirable, as it leads to overfitting. We try to learn the surrogate via another kernel method, regularized least-squares or ridge regression. 10

Kernel-based Operator Learning

Definition 12 Let λ > 0 be a penalization parameter. Let Ξ := {ξ 1 , . . . , ξ M } ⊆ D be a set of training points and F = {f1 , . . . , fM } ⊆ Rr be the labels. Define the functional Jλ : H → R by Jλ (s) :=

M X

∥s(ξ j ) − fj ∥22 + λ∥s∥2H .

(5)

j=1

Then the regularized least-squares problem is given as Find sλ ∈ H such that sλ = argmins∈H Jλ (s). It turns out that, if H is a RKHS, the solution of the regularized least-squares problem, a minimization problem in an infinite dimensional space, is unique and an element of the finite dimensional space VM , see, e.g., Steinwart and Christmann (2008). Theorem 13 For any λ > 0, there is a unique minimizer of the regularized least-squares problem of Theorem 12 in H. In addition, there exists a coefficient vector w ∈ RrM such that sλ =

M X

K(·, ξ i )wi ,

i=1

i.e., sλ ∈ VM . The following estimate is important for the error analysis. Proposition 14 Let sλ ∈ VM be the unique minimizer of the regularized least-squares problem and assume that the labels are generated by f ∈ H, i.e., fi = f (ξi ), 1 ≤ i ≤ M . Then ∥sλ ∥H ≤ ∥f ∥H . Remark 15 (Computational form) Both the interpolant and the regularized least-squares solution admit explicit representations in terms of the kernel matrix K Ξ,Ξ ∈ RrM ×rM with block entries K(ξ i , ξj ) ∈ Rr×r : s0 = K Ξ (·) K −1 Ξ,Ξ [f ], sλ = K Ξ (·) (K Ξ,Ξ + λI)−1 [f ], where [f ] = [f1 , . . . , fM ]T ∈ RrM is the vector containing the labels. In particular, IΞ and Qλ,Ξ are both linear maps that can be evaluated via a single matrix-vector product once the respective system matrix has been factored offline. Remark 16 It is also possible to augment the functional Jλ in (5) by additional penalization terms. Of particular interest for our framework is a PDE residual penalty µ∥Ls − f ∥2L2 enforcing a soft PDE constraint on the reconstruction. We revisit this extension in Section 5. 11

R. Kempf

Again, the solution of the regularized least-squares problem sλ defines a linear operator Qλ,Ξ : RrM → VM by Qλ,Ξ ([f ]) = sλ . If the labels are generated by a function f ∈ C(D)r , this can be extended to an operator Qλ,Ξ : C(D)r → VM by setting Qλ,Ξ (f ) = Qλ,Ξ ([f ]). Indeed, the label generating function does not need to be in the RKHS. In the error analysis for kernel-based operator learning, we will make use of the following estimate. Proposition 17 With the notation and assumptions of Theorem 12, let Qλ,Ξ : RrM → VM be the regularized least-squares operator. Then the bound 1 ∥Qλ,Ξ ([f ])∥H ≤ √ ∥[f ]∥2 λ holds for all [f ] ∈ RrM . Proof Since Qλ,Ξ ([f ]) is the unique minimizer of Jλ we have λ∥Qλ,Ξ ([f ])∥2H ≤ Jλ (Qλ,Ξ ([f ])) ≤ Jλ (0) = ∥[f ]∥22 .

In particular, Theorem 17 means that 1 ∥Qλ,Ξ ∥RrM →H ≤ √ . λ This bound is central to the error analysis of Section 4: it shows that the reconstruction operator Qλ,Ξ is bounded as a map from data space to H, with stability constant √1λ that is independent of the point set Ξ and the kernel K. This uniform bound is what allows us to control the propagation of the learning error of Aoff through the reconstruction Aon . 3.3 Sampling Inequalities Sampling inequalities are the key tool connecting discrete training errors to continuous function space norms. Concretely, they bound a weak Sobolev norm ∥ · ∥Wqt of a function by a combination of a strong norm ∥ · ∥H s and a discrete norm at the training points, with the fill distance hΞ,D controlling the relative weight of the two terms. In this sense, they play an analogous role to covering number or Rademacher complexity bounds in statistical learning theory (Steinwart and Christmann, 2008): they quantify how well discrete information at a finite point set represents the continuous function. The version we give here can be found in Gia et al. (2025, Theorem 2.13). We use the notation (·)+ := max(0, ·). Theorem 18 Let D ⊆ Rd be a bounded Lipschitz domain, p, q ∈ [1, ∞], s > d/2, and γ := max(2, p, q). Then there exist constants h0 , C > 0 such that for all Ξ ⊆ D with hΞ,D < h0 , all f ∈ H s (D)r , and all admissible 0 ≤ t ≤ ℓ,   ! 1 1 s−t−d

∥f ∥Wqt (D)r ≤ C

hΞ,D

2

−q

d

+

−t

γ ∥f ∥H s (D)r + hΞ,D ∥f ∥ℓp (Ξ)r

12

.

(6)

Kernel-based Operator Learning

Remark  19 The admissible range of t depends on s, q, and d. Specifically, ℓ = ℓ0 := when s ∈ N and either q > 2 with ℓ0 ∈ N, or q = 2. Otherwise ℓ = ⌈ℓ0 ⌉ − 1. s − d 21 − 1q +

For q = ∞, one additionally requires t ∈ N0 . We obtain the following estimates for the two learning method introduced above. Proposition 20 Let D ⊆ Rd be a bounded Lipschitz domain and let s > d/2. Assume that the reproducing kernel Hilbert space H is norm-equivalent to H s (D)r . Let Ξ ⊆ D be a set of training points with sufficiently small fill distance hΞ,D . Then the following error estimates hold for all admissible t and q as in Theorem 18 and all f ∈ H s (D)r : 1. For interpolation,   s−t−d 12 − 1q

+

∥f − IΞ (f )∥Wqt (D)r ≤ ChΞ,D

∥f ∥H s (D)r .

2. For regularized least-squares learning, ∥f − Qλ,Ξ (f )∥Wqt (D)r ≤ C

  s−t−d 12 − 1q

hΞ,D

+

√ +

d −t γ

λhΞ,D

! ∥f ∥H s (D)r ,

where γ = max(2, q). Proof The interpolation case follows immediately from Theorem 18 and Theorem 11, using the norm equivalence of H and H s (D)r . To see the statement for the penalized least-squares case, we use (6) with u = f − sλ and p = 2. First, by Theorem 14 and norm equivalence of H and H s (D)r , we obtain ∥f − sλ ∥H s (D)r ≤ ∥f ∥H s (D)r + ∥sλ,F ∥H s (D)r ≤ C∥f ∥H s (D)r . and second, since sλ is the unique minimizer of Jλ ∥f − sλ ∥2ℓ2 (Ξ)r =

M X

∥f (ξj ) − sλ (ξ j )∥22

j=1

≤ Jλ (sλ ) ≤ Jλ (f ) = λ∥f ∥2H ≤ Cλ∥f ∥2H s (D)r .

In the error estimates of Theorem 20, we assumed that the target function f lies in H, i.e., f ∈ H s (D)r . In practice, this assumption may be too restrictive: the true operator output G(u) needs not to have the same smoothness as the kernel used for reconstruction. This situation, where the target function has smoothness b strictly less than the kernel smoothness s is known as the mismatch case or escaping the native space (Narcowich et al., 2006). From a machine learning perspective, this corresponds to a mild form of model misspecification: the hypothesis space VM is richer than necessary for the target function, 13

R. Kempf

yet useful approximation rates can still be recovered, albeit at a reduced rate reflecting the true smoothness b rather than the kernel smoothness s. To give the corresponding error estimate, we define the separation radius qΞ of Ξ to be qΞ :=

1 min ∥ξ − ξ j ∥2 , 2 i̸=j i

and the mesh-ratio ρΞ,D of Ξ in D to be ρΞ,D :=

hΞ,D . qΞ

With this, we have the following error estimates (Gia et al., 2025, Theorem 4.5 and Theorem 4.10). Proposition 21 Let D ⊆ Rd be a bounded Lipschitz domain and let s > d/2. Assume that the reproducing kernel Hilbert space H is norm-equivalent to H s (D)r . Let Ξ ⊆ D be a set of training points with sufficiently small fill distance hΞ,D , and assume that Ξ is quasi-uniform, i.e., ρΞ,D ≤ c. Let f ∈ H b (D)r with d/2 < b < s. Then the following error estimates hold for all admissible t and q as in Theorem 18: 1. For interpolation b−t−d

∥f − IΞ (f )∥Wqt (D)r ≤ ChΞ,D



1 − 1q 2

 +

ρs−b Ξ,D ∥f ∥H b (D)r ,

with a constant C > 0. 2. For regularized least-squares learning ∥f − Qλ,Ξ (f )∥Wqt (D)r ≤ b−t−d

≤C

hΞ,D



1 − 1q 2

 +

ρs−b Ξ,D +

b−s+ d −t λhΞ,D γ ρs−b Ξ,D

! ∥f ∥H b (D)r .

Remark 22 For a quasi-uniform point set Ξ = {ξ 1 , . . . , ξ M } in a bounded domain D ⊆ Rd , i.e., ρΞ,D ≤ c, the fill distance satisfies hΞ,D ∼ M −1/d , where the implicit constants depend only on D and the mesh-ratio bound c. This relation allows all fill distance conditions in this paper to be restated directly in terms of sample sizes.

4 Error Analysis and Budget Allocation for Kernel-based Operator Learning The purpose of the following analysis is to establish both convergence of the surrogate A and an explicit budget allocation condition quantifying how the three design parameters must be coupled. The resulting balance condition will provide an explicit resource-allocation rule for kernel-based operator learning. We recall that A = Aon ◦ Aoff ◦ SX 14

Kernel-based Operator Learning

consists of three components: SX discretizes the input function, Aoff learns a finite-dimensional map between discretized inputs and outputs, and Aon reconstructs an output function from the predicted discretized observations. Assumption 23 We make the following assumptions: 1. Let U be a Banach space and M ⊆ U a compact model class. 2. The output space V = HKV is a reproducing kernel Hilbert space norm-equivalent to H σ (D)ℓ , σ > d/2. 3. The bottom RKHS HKb is norm-equivalent to H α (B)mℓ , α > np/2, where B = SX (M) ⊆ Rnp is the image of the model class under SX . Let B be a bounded Lipschitz domain. 4. There exists a function g ∈ HKb satisfying g(SX (u)) = SY (G(u)) for all u ∈ M. 5. The output observation points Y ⊆ D have fill distance hY,D . 6. The discretized training inputs U = {SX (ui )}N i=1 ⊆ B have fill distance hU,B . Theorem 23(4) formalizes the requirement that the operator G can be represented consistently at the discretized level. In particular, it guarantees that the continuous operator learning problem induces a well-defined finite-dimensional learning problem. The key idea underlying our error analysis, which distinguishes our approach from Batlle et al. (2024), is the following: rather than tracing the error through the full composition Aon ◦ Aoff ◦ SX directly, we interpret A(u) = Aon (Aoff (SX (u))) as a reconstruction in V from perturbed data. Specifically, the exact discretized output v † = SY (G(u)) ∈ Rmℓ is unknown, and Aon instead receives the disturbed data Aoff (SX (u)) ≈ v † . This justifies the natural choice Aon = Qλ,Y . Since Qλ,Y is linear, we can insert the exact data and decompose the total error as ∥G(u) − A(u)∥Wqτ (D)ℓ ≤ ∥G(u) − Aon (v † )∥Wqτ (D)ℓ + ∥Aon ∥Rmℓ →V · ∥Aoff (SX (u)) − v † ∥2 , {z } | {z } | Term II

Term I

(7) where Term I is the reconstruction error incurring when reconstructing G(u) from v † . Its decay is governed by the output discretization density hY,D and Term II is the learning error measuring the approximation error of the learned discretized operator. It is amplified by the reconstruction stability factor ∥Aon ∥ ≤ λ−1/2 . This error is governed by the training discretization density hU,B . The two terms are coupled only through the regularization parameter λ: smaller λ reduces the reconstruction error in Term I but amplifies the learning error in Term II via the stability constant, and vice versa. This explicit coupling allows an optimal choice of λ, which we discuss after deriving general error estimates. 15

R. Kempf

4.1 Generating the Data by Interpolation We first consider the noiseless setting in which the discretized operator is learned via kernel interpolation. This corresponds to deterministic training outputs generated without additional noise. The following theorem quantifies how reconstruction accuracy and discretized learning accuracy interact through the regularization parameter λ. Theorem 24 With the notation and assumptions of Theorem 23, let Aoff = IU be the interpolation operator of Section 3.2.1. Then there exist constants h1 , h2 , C1 , C2 > 0 such that for all hY,D < h1 , all hU,B < h2 , all admissible τ , q as in Theorem 18, and all u ∈ M, the estimate ∥G(u) − A(u)∥Wqτ (D)ℓ ≤ C1 f1 ∥G(u)∥H σ (D)ℓ + C2 f2 ∥g∥H α (B)mℓ holds, where σ−τ −d

f1 := hY,D



1 − 1q 2

 +

√ +

d

−τ

γ λ hY,D

and

1 α− 1 np f2 := √ hU,B2 , λ

with γ = max(2, q). Proof The error decomposition (7) gives Term I and Term II. Term I is bounded by applying Theorem 20 to Qλ,Y with f = G(u) ∈ V. Term II is bounded using Theorem 17 for ∥Aon ∥ and Theorem 20 applied to IU with f = g ∈ HKb . Combining the two bounds yields the result.

Remark 25 (Mismatch cases) Theorem 24 assumes that G(u) ∈ V = HKV and g ∈ HKb , i.e., both the target output function and the discretized operator have smoothness matching their respective native spaces. In practice, either or both of these assumptions may fail, leading to two possible mismatch scenarios. In both cases, we additionally assume that the respective point sets are quasi-uniform, i.e., ρY,D := hY,D /qY ≤ c and ρU,B := hU,B /qU ≤ c for a constant c > 0, where qY and qU denote the separation radii of Y and U , respectively. In the output mismatch case, G(u) ∈ H µ (D)ℓ with d/2 < µ < σ. The factor f1 is replaced by µ−τ −d(1/2−1/q)+ σ−µ ρY,D ,

fmiss := hY,D 1

reflecting the reduced smoothness of the target output at the cost of an additional mesh-ratio factor ρσ−µ Y,D . Additionally, ∥G(u)∥H σ (D)ℓ has to be replaced by ∥G(u)∥H µ (D)ℓ In the bottom mismatch case, g ∈ H β (B)mℓ with np/2 < β < α. The factor f2 is replaced by β− 1 np α−β ρU,B ,

fmiss := hU,B2 2

together with the respective norm of g. Both mismatches can occur simultaneously. Then both replacements apply. 16

Kernel-based Operator Learning

The regularization parameter λ controls an explicit bias-variance tradeoff: smaller values improve reconstruction accuracy, but amplify errors propagated from the learned discretized operator. For readability, we only consider q = 2. Similar results can be obtained for all admissible q. Choosing the regularization parameter such that  λ∗ = cλ 

hσ−τ Y,D d −τ 2

2

2  d  = cλ hσ− 2 , Y,D

(8)

hY,D with a constant cλ > 0, this yields that

α− 1 np

f1 = cλ hσ−τ Y,D

and

f2 = cλ

hU,B2

σ− d

.

hY,D2

This choice of λ∗ yields an explicit scaling law relating the output discretization complexity, the training sample complexity, and the regularity of the discretized operator. Corollary 26 (Budget Allocation Rule) With the notation and assumptions of Theorem 24, set q = 2, assume that Y and U are quasi-uniform with mesh-ratios ρY,D ≤ c and ρU,B ≤ c, so that hY,D ∼ m−1/d and hU,B ∼ N −1/(np) , and choose λ = λ∗ as in (8). Then the error estimate becomes σ−τ

∥G(u) − A(u)∥H τ (D)ℓ ≤ C1 m− d ∥G(u)∥H σ (D)ℓ + C2 N

− 2α−np 2np

2σ−d

m 2d ∥g∥H α (B)mℓ ,

(9)

and the total error converges to zero as m, N → ∞ provided that log N np(2σ − d) ≥ . log m d(2α − np)

(10)

The condition (10) expresses a budget allocation rule between three free design parameters: the number of training pairs N , the number of input observation points n and the output resolution m. It prescribes how N must scale relative to m and n in order to prevent the learning error from dominating the reconstruction error. The interplay between these parameters is discussed further in Section 4.3. np(2σ−d)

Remark 27 (Unified Convergence rate) If N = m d(2α−np) , i.e., (10) holds with equality, then both terms in (9) decay at the same rate and the total error satisfies σ−τ

∥G(u) − A(u)∥H τ (D)ℓ ≤ C m− d



 ∥G(u)∥H σ (D)ℓ + ∥g∥H α (B)mℓ .

In other words, if N grows sufficiently fast relative to m, the learning error of Aoff does not pollute the overall convergence rate, and the surrogate A converges at the same rate m−(σ−τ )/d as the pure reconstruction error of Aon from exact data. 17

R. Kempf

4.2 Generating the Data by Regularized Least-Squares We now replace interpolation at the offline phase by regularized least-squares (RLS) approximation. This setting is particularly relevant in operator learning applications, where training outputs are often generated numerically and therefore contain discretization or solver error. In this case, exact interpolation may lead to instability or overfitting of the discretized operator. The following theorem shows that we obtain a similar convergence estimate as in the interpolation setting, up to an additional bias term. Theorem 28 With the notation and assumptions of Theorem 23, let Aoff = Qµ,U be the regularized least-squares operator of Section 3.2.2 with regularization parameter µ > 0. Then there exist constants h1 , h2 , C1 , C2 > 0 such that for all hY,D < h1 , all hU,B < h2 , all admissible τ , q as in Theorem 18, and all u ∈ M, the estimate ∥G(u) − A(u)∥Wqτ (D)ℓ ≤ C1 f1 ∥G(u)∥H σ (D)ℓ + C2ef2 ∥g∥H α (B)mℓ holds, where f1 is as in Theorem 24 and ef2 is given by   1 ef2 := √1 hα− 2 np + √µ . U,B λ Remark 29 (Mismatch cases) The mismatch scenarios discussed in Theorem 25 extend directly to the this setting and yield analogous modified rates. The regularization parameter µ introduces a second bias-variance tradeoff at the bottom learning stage. Choosing µ too small reduces regularization bias but amplifies instability, whereas large µ improves stability at the cost of approximation accuracy. Choosing λ as in (8) and µ as  2 α− 1 np µ∗ := cµ hU,B2 ,

(11) α− 1 np

with a free constant cµ > 0, exactly balances the fill-distance term hU,B2 and the regu√ larization bias µ in ef2 , yielding the following result. Remarkably, after optimal parameter balancing, the RLS and interpolation regimes yield identical asymptotic scaling laws. Corollary 30 With the notation and assumptions of Theorem 28, set q = 2, assume that Y and U are quasi-uniform with mesh-ratios ρY,D ≤ c and ρU,B ≤ c, and choose λ = λ∗ as in (8) and µ = µ∗ as in (11). Then the error estimate and budget allocation condition (9) and (10) of Theorem 26 hold verbatim. Theorem 30 shows that, with the optimal choice of µ, regularized least-squares at the bottom achieves the same asymptotic convergence rates and budget allocation condition as interpolation. Although RLS does not improve the asymptotic rate, it improves robustness in practice by better conditioning of the kernel system, particularly when training outputs contain numerical noise or the training set is large. 18

Kernel-based Operator Learning

Remark 31 (Comparison with Batlle et al. (2024)) The error results presented above are complementary to the main quantitative result of Batlle et al. (2024) (Theorem 3.3), reflecting the different goals of the two frameworks. We highlight three structural differences. ′ (i) Target norm. We extend the range of norms the error is measured from H t -norm results to more general Wqτ -norm estimates. This includes in particular L1 - and L∞ -error results. (ii) Input discretization term. We can drop the assumption that the input space U has RKHS structure and do not need a reconstruction Ale . Moreover, Condition 3.2 of Batlle et al. (2024), namely ui = Ale ◦ SX (ui ) for all training inputs is dropped. The input discretization error is therefore entirely absent in Theorem 24. (iii) Explicit vs implicit bottom error rates. In Batlle et al. (2024, Equation 3.8), the learning error is expressed in terms of maxj ∥fj† ∥S , the RKHS norm of the components of f † , which characterizes the smoothness of the target but does not directly give a rate in terms of α− 1 np

the number of training pairs N . In Theorem 24, the corresponding factor is f2 = √1λ hU,B2 , which provides an explicit convergence rate in the fill distance hU,B of the training inputs and directly yields the budget allocation condition (10) of Theorem 26. 4.3 Existence of the Discretized Operator We expand the discussion of the existence of the discretized operator g begun in Section 2.4, focusing here on the competing demands that existence and learnability of g place on the discretization parameters n and m. A practically important special case is M = BR [0] ∩ Un0 , where Un0 ⊆ U is a subspace of dimension n0 and BR [0] denotes the closed ball of radius R in U about 0. Then M is compact and SX : Un0 → Rnp is a linear map between finite-dimensional spaces. Whenever the point evaluations {δx1 , . . . , δxn } separate the elements of Un0 , satisfied in particular when np ≥ n0 , the map SX is injective on Un0 , so ιΦ ≡ 0 on M and g is well-defined. This setting closely resembles the framework of Batlle et al. (2024), where training inputs satisfy the reconstruction condition ui = Ale ◦ ϕ(ui ), effectively restricting to a finite-dimensional subspace of the input RKHS. In practice, PCA preprocessing of the training inputs, as used in Section 6, realizes precisely this setting: Un0 is the span of the leading n0 PCA components, and injectivity of SX on Un0 holds whenever the PCA basis functions are separated by the sampling points X. Remark 32 (Verification of (1)) In this finite-dimensional setting, condition (1) is automatically satisfied. Since SX : Un0 → Rnp is a linear injective map between finite−1 dimensional spaces, its left inverse SX : SX (Un0 ) → Un0 is bounded. Hence, (1) holds −1 with ιΦ ≡ 0 and C = ∥SX ∥. In the PCA setting, C depends on the condition number of the matrix formed by the leading n0 PCA basis functions at the sampling points X, and is controlled by choosing X with sufficiently small fill distance relative to the spatial frequency of the basis functions. Furthermore, the existence and learnability of g impose competing demands on n and m. Increasing n reduces information loss under discretization and improves identifiability on M: in the degenerate case n = 0, all inputs become indistinguishable. At the same 19

R. Kempf

time, the budget allocation condition (10) shows that the required number of training pairs grows as np(2σ−d)

N ≳ m d(2α−np) , so larger n simultaneously increases the statistical complexity of learning g. The output parameter m plays a different role: larger m improves reconstruction accuracy of Aon through the decay of hY,D , but increases the output dimension mℓ of g and hence the complexity of the regression problem. Thus, n and m together control a fundamental representationlearnability tradeoff in the two-stage operator learning framework.

5 Physics-Informed Reconstruction In many operator learning applications, G maps a parameter or coefficient function u to the solution v of a PDE. Concretely, we assume that for each u ∈ M, the output G(u) ∈ V is the unique solution of Lu v = f (u)

in D,

(12)

where Lu : V → L2 (D)ℓ is a linear differential operator of order ν depending on u, and f : M → L2 (D)ℓ is a known forcing term. We emphasize that (12) covers both the coefficient-to-solution map (e.g., Lu = −div(eu ∇·), f (u) = 1) and the forcing-to-solution map (e.g., Lu = −∆, f (u) = u). We wish the surrogate APDE to approximately satisfy (12), i.e., we want Lu APDE (u) ≈ f (u) in D for each u ∈ M. As recalled in Section 1, one classical approach encodes physical constraints directly into the kernel KV . However, this requires the constraint to be identifiable from the analytical structure of Lu and fails when Lu depends on u. We therefore adopt a soft-constraint approach: we penalize the PDE residual at a finite set of collocation points Z = {z 1 , . . . , z mL } ⊆ D, requiring only the ability to evaluate Lu and f (u) pointwise for each u ∈ M. Assumption 33 We make the following assumptions: 1. The operator Lu is a linear differential operator of order ν for each fixed u ∈ M, and the native space V = HKV = H σ (D)ℓ satisfies V ,→ C 2ν (D), which holds whenever σ > 2ν + d/2. 2. Lu satisfies a uniform a priori estimate over M, i.e., there exists a constant cmin > 0 such that for all u ∈ M and all v ∈ V, ∥v∥H ν (D)ℓ ≤

1  cmin

 ∥Lu v∥L2 (D)ℓ + ∥v∥L2 (D)ℓ .

3. The PDE collocation points Z = {z 1 , . . . , z mL } ⊆ D have fill distance hZ,D . The linearity of Lu is essential: it ensures that v 7→ Lu v(z k ) is a bounded linear functional on V, which is required for the representer theorem of Micchelli and Pontil 20

Kernel-based Operator Learning

(2005). Particular examples of operators satisfying Theorem 33 (2) are Darcy-type operators Lu = −div(a(u, ·)∇·) whenever a(u, x) ≥ amin > 0 uniformly over M and D, which holds for example when u is bounded on M and a = eu . For second-order uniformly elliptic operators with suitable boundary conditions, the estimate follows from classical elliptic regularity theory (Evans, 2010). 5.1 The Physics-Informed Reconstruction Operator and Representer Theorem We define the collocation sampling operator SZLu : V → RmL ℓ by [SZLu v]k := Lu v(z k ),

k = 1, . . . , mL ,

which is the natural analogue of SY for differential operator evaluations. For labels v ∈ Rmℓ PDE : V → R by and u ∈ M, define the physics-informed Tikhonov functional Jλ,µ,u PDE (s) := ∥SY s − v∥22 + λ∥s∥2V + Jλ,µ,u

µ ∥S Lu s − [f (u)]∥22 , mL Z

(13)

where [f (u)] := SZ (f (u)) ∈ RmL ℓ collects the values of the forcing term at the collocation points. The three terms in (13) penalize the data misfit at Y , the RKHS norm of v, and the PDE residual at Z, respectively. Definition 34 Let Theorem 23 and Theorem 33 hold and let λ, µ > 0. The physicsinformed reconstruction operator APDE : Rmℓ × M → V is defined by on PDE APDE on (v, u) := argmin Jλ,µ,u (s).

(14)

s∈V

The corresponding physics-informed surrogate operator APDE : U → V is APDE (u) := APDE on (Aoff (SX (u)), u),

u ∈ M.

Remark 35 Note that APDE (u) is no longer of the form Aon ◦ Aoff ◦ SX with a fixed Aon , since APDE on (·, u) depends on u through Lu and f (u). However, for each fixed u it is still a linear map from Rmℓ to V, and the error decomposition (7) applies verbatim with APDE on (·, u) in place of Aon . Since both SY and SZLu consist of bounded linear functionals on V, the generalized representer theorem of Micchelli and Pontil (2005) applies to (14). Theorem 36 Let Theorem 23 and Theorem 33 hold. Then for each v ∈ Rmℓ and u ∈ M there exists a unique minimizer APDE on (v, u) of (14), which lies in the collocation-augmented hypothesis space  n omL  m VY,Z (u) := span KV (·, y j ) j=1 ∪ (L(2) , u KV )(·, ζ) ζ=z k

k=1

(2) where Lu denotes Lu acting on the second argument of KV . Writing

APDE on (v, u) =

m X

KV (·, y j )αj +

j=1

mL X ((L(2) u KV )(·, ζ) ζ=z )β k , k

k=1

21

R. Kempf

the coefficients (α, β) ∈ Rmℓ × RmL ℓ are the unique solution of the block linear system   ! ! K Y Y + λI KL YZ α v  = , (15) λmL  LL T β f (u) (K L ) K + I YZ ZZ µ where the kernel matrices are defined entry-wise by [K Y Y ]ij := KV (y i , y j ), (2) [K L Y Z ]jk := (Lu KV )(y j , ζ) ζ=z , k

(1) (2) [K LL ZZ ]kl := (Lu Lu KV )(ξ, ζ) ζ=z , ξ=z . k

l

Proof Let u ∈ M be fixed. Because of the strict convexity of ∥ · ∥V there is a unique PDE in V. minimizer s∗ of Jλ,µ,u To show that s∗ ∈ VY,Z (u), we split s∗ = s∥ + s⊥ , where s∥ ∈ VY,Z and s⊥ ∈ VY,Z (u)⊥ , the orthogonal complement of VY,Z in V. Because VY,Z (u) is finite dimensional, this splitting is direct. We have, with the reproducing property of KV , that s⊥ (y i ) = ⟨s⊥ , KV (·, y i )⟩V = 0,

yi ∈ Y

and ∥s∗ ∥2V = ∥s∥ ∥2V + ∥s⊥ ∥2V + 2⟨s∥ , s⊥ ⟩V = ∥s∥ ∥2V + ∥s⊥ ∥2V . Furthermore, by Theorem 33, the mapping h 7→ (Lu h)(z k ), h ∈ V, is bounded and linear. This means that there is a rk ∈ V such that (Lu h)(z k ) = ⟨h, rk ⟩V by the Riesz representation theorem. Using the reproducing property of KV , we obtain that (2) rk = Lu KV (·, ζ)|ζ=zk ∈ VY,Z (u). This also yields that ⟨s⊥ , rk ⟩ = (Lu s⊥ )(z k ) = 0 for all z k ∈ Z. Together, we see that µ PDE ∗ (s ) = ∥SY (s∗ ) − v∥22 + λ∥s∗ ∥2V + Jλ,µ,u ∥S Lu (s∗ ) − [f (u)]∥22 mL Z m X µ = ∥v j − s∥ (y j )∥22 + λ∥s∥ ∥2V + λ∥s⊥ ∥2V + ∥S Lu (s∥ ) − [f (u)]∥22 . mL Z j=1

PDE , this means that ∥s⊥ ∥2 = 0, i.e., s⊥ = 0. However, since s∗ is the unique minimizer of Jλ,µ,u V In turn, this yields s∗ ∈ VY,Z (u). Hence, there are coefficient vectors α ∈ Rmℓ and β ∈ RmL ℓ such that   mL m X X α ∗ (2) L s = KV (·, y j )αj + (Lu KV (·, ζ)|ζ=zk )β k = (K Y (·), K Z (·)) . β j=1

k=1

PDE yields a finite-dimensional quadratic functional in Inserting the ansatz for s∗ into Jλ,µ,u mℓ m ℓ (α, β) ∈ R × R L . Setting the gradient to zero yields (15).

22

Kernel-based Operator Learning

Remark 37 (Smoothness requirement) The kernel matrices in (15) reveal why we require σ > 2ν + d/2 rather than σ > ν + d/2. The matrix K LL ZZ involves applying Lu to both arguments of KV simultaneously, requiring 2ν derivatives of the kernel to exist and be bounded. This is the standard smoothness-doubling phenomenon in symmetric kernel collocation (Schaback, 2009; Chen et al., 2021). For Matérn kernels, σ > 2ν + d/2 is satisfied for example by choosing σ = 2ν + d/2 + 1/2; for second-order operators (ν = 2) in two dimensions (d = 2), this requires σ > 5. Remark 38 (Connection to kernel collocation) In the limit µ → ∞, the diagonal regL ularization λm µ I in the lower-right block of (15) tends to zero, and the system enforces SZLu v = f (u) exactly, recovering symmetric kernel collocation (Schaback, 2009) augmented L T with data fit at Y . In the limit µ → 0, the off-diagonal blocks K L Y Z and (K Y Z ) become negligible relative to the diagonal regularization, and the system reduces to the standard Tikhonov system of Section 3.2.2 with β = 0. The parameter µ therefore continuously interpolates between pure data fitting (µ = 0) and exact PDE enforcement (µ → ∞). 5.2 Computational Cost We now briefly discuss the computational cost of the physics informed operator learning method. The costs for the non-physics informed method can be recovered by setting mL = 0. is assembling and solving the block system (15) of The dominant online cost of APDE on size (m + mL ) × (m + mL ). The matrix K Y Y is the standard output kernel matrix and is independent of u. It is assembled and factored once per output grid Y . The matrices K L YZ and K LL ZZ require evaluating derivatives of KV under Lu : for a second-order operator this means second derivatives of KV , available analytically for Matérn kernels or via automatic 2 differentiation. The assembly of K LL ZZ costs O(mL ) kernel evaluations and is the dominant assembly cost when Lu depends on u. The Cholesky factorization of the full block system costs O((m + mL )3 ), the same order as the standard Tikhonov solve but with a larger constant. When Lu = L is independent of u (e.g., the forcing-to-solution map Lu = −∆, f (u) = u), all three kernel matrices are independent of u and can be assembled and factored once offline. The online cost per test input then reduces to a single triangular solve of size (m + mL ) plus one evaluation of f (u) at Z. When Lu depends on u (e.g., the coefficient-to-solution map LL Lu = −div(eu ∇·)), the matrices K L Y Z and K ZZ must be recomputed for each test input u, and the full O((m + mL )3 ) solve is required online. A rigorous error analysis of APDE , including the optimal choice of µ and the resulting convergence rate, is left for future work.

6 Numerical Experiments We validate Section 4 on the Darcy flow equation and test the physics-informed reconstruction of Section 5 on the Poisson equation. Shared setup. Inputs u are drawn from a Gaussian random field obtained by applying a Gaussian filter with bandwidth σinput = 8 to white noise and clipping to [−4, 4], giving C ∞ -like realizations. Solutions are computed by a finite difference scheme on a uniform 23

R. Kempf

64 × 64 interior grid with zero Dirichlet boundary conditions. We generate Ntotal = 60000 input-output pairs, split into Ntrain = 50000 training and Ntest = 10000 test pairs. Input functions are reduced via PCA, retaining 99% of variance. This is achieved with nPCA = 28 √ √ components. Output observation points Y form a uniform m × m grid of interior points on (0, 1)2 . For the PI experiments, error averages are computed over a subset of 500 test samples due to the additional per-sample cost of solving the augmented system. 6.1 Kernel Operator Learning: Darcy Flow We consider the parametric Darcy flow problem −div(eu(x) ∇v(x)) = 1

on (0, 1)2 ,

v = 0 on ∂(0, 1)2 ,

Mean relative L2 error

where u is the log-permeability and G : u 7→ v the parameter-to-solution map. Both the output kernel KV and the bottom kernel Kb are matrix valued kernels with the Matérn-3/2 function on the diagonal. The native space of Matérn-3/2 on R2 is H 2 (R2 ), giving σ = 2. The regularization parameter is λ(m) = m−5/2 , which lies well within the flat plateau of the error curve identified in Fig. 5.

10−1

10−2

9

25

49 121 225 Output observation points m

Oracle Aon

Full surrogate A

441

784

Ref. slope −1

Figure 3: Convergence of the oracle and full surrogate error as a function of output resolution m with N = 50,000 fixed. The oracle error transitions from a pre-asymptotic regime (local slope ≈ −0.87 for m ≤ 121) toward the theoretical rate m−1 for Matérn- 32 (σ = 2, d = 2), with local slope −1.03 for m ≥ 121. The full surrogate tracks the oracle throughout. Convergence in m. Fig. 3 shows the mean relative L2 error as a function of m with N = Ntrain fixed. The oracle error exhibits a clear pre-asymptotic regime for small m, with local slopes around −0.87 for m ≤ 121, steepening toward the theoretical prediction m−σ/d = m−1 of Theorem 24 (σ = 2, d = 2) as m increases. The local slope over m ∈ {121, . . . , 784} is −1.03, consistent with the asymptotic rate. The full surrogate error tracks 24

Kernel-based Operator Learning

the oracle throughout the entire range, confirming that the bottom-level learning error is negligible when N is large.

Mean relative L2 error

Budget allocation. Fig. 4 shows the full surrogate error for N = mκ with varying κ, using nPCA = 20. For κ = 0.5 the error stagnates, indicating that the training set is too small to keep pace with the growing output resolution. For κ ≥ 1.0 the curves converge at a rate comparable to the oracle, and the three curves with κ ∈ {1.0, 1.5, 2.0} are essentially indistinguishable from one another. This suggests that a moderate superlinear growth N ∼ mκ with κ somewhere between 0.5 and 1.0 is sufficient to avoid stagnation, and that investing further in training data beyond that threshold yields no measurable benefit.

10−1

10−2

9

25

49 121 225 Output observation points m

κ = 0.5 κ = 2.0

κ = 1.0 Ref. slope −1

441

784

κ = 1.5

Figure 4: Budget allocation experiment with N = mκ and nPCA = 20. The curve with κ = 0.5 stagnates, while all curves with κ ≥ 1.0 converge at the oracle rate and are essentially indistinguishable from each other. Regularization sensitivity. Fig. 5 shows the oracle error as a function of λ at m = 225. The curve is flat over many orders of magnitude around the theoretical optimum λ∗ = 225−1 ≈ 4.44 × 10−3 , with a sharp increase only for λ ≫ λ∗ . The schedule used in practice, λ(m) = m−5/2 , gives λ(225) ≈ 1.32 × 10−6 , which is well within the flat region. 6.2 Physics-Informed Reconstruction: Poisson Equation We consider the Poisson problem −∆v(x) = u(x)

on (0, 1)2 ,

v = 0 on ∂(0, 1)2 .

The choice of −∆ as the differential operator is deliberate: its action on the Matérn kernel has a closed-form expression, so the representer-theorem system of Section 5 can be assembled exactly without approximating kernel derivatives. Both matrix valued kernels 25

R. Kempf

Mean relative L2 error

0.6

0.4

0.2

0 10−14

10−12

10−10 10−8 10−6 10−4 10−2 Regularization parameter λ

100

102

λ∗ = 4.44 × 10−3 (theory)

Oracle Ar

Figure 5: Oracle error as a function of λ for fixed m = 225. The error is flat over more than ten orders of magnitude to the left of the theoretical optimum λ∗ = 225−1 ≈ 4.44 × 10−3 , confirming robustness of the regularization schedule λ(m) = m−5/2 .

used are again diagonal kernels, where the the output kernel is chosen to be Matérn-9/2 (ν = 4.5, σ = 5.5), because the system matrix requires evaluating (−∆)2 KV , which demands KV ∈ C 4 . Matérn-9/2 provides C 8 smoothness, giving a comfortable margin. The lengthscale is selected per test sample by leave-one-out cross-validation on the Y -block. The PDE constraint weight is µ = λ(m), tying the PI regularization to the data regularization so that neither block dominates across the full range of m. The bottom kernel is Matérn3/2. We compare three reconstructors: the plain oracle Aon (kernel interpolant using exact SY v ∗ ), the PI oracle API on (same plus PDE enforced at mL collocation points Z), and the PI surrogate (same as PI oracle but with SY v ∗ replaced by Aoff (u∗ )). Results are averaged over 500 test samples. Convergence in m. The PI oracle error is smaller than the plain oracle error at every tested value of m, demonstrating that enforcing the PDE at collocation points reduces the reconstruction error. Both the PI oracle and plain oracle saturate for m > 441 at an error of approximately 1.3 × 10−3 , which coincides with the O(h2 ) discretization error of the 64 × 64 FD solver; this is a solver precision ceiling rather than a limitation of the reconstruction method. The PI surrogate tracks the PI oracle for small m but saturates at ≈ 0.023 from m = 121 onward. This saturation reflects the irreducible learning error of Aoff in nPCA = 28 input dimensions: as the PI reconstruction becomes more accurate, it faithfully propagates the fixed prediction bias of Aoff into the output, consistent with the budget allocation analysis of Theorem 27. Sensitivity to mL . Fig. 7 shows the PI oracle error at m = 49 as a function of the number of collocation points. The error decreases up to mL = m = 49 and then stagnates. Adding 26

Kernel-based Operator Learning

Mean relative L2 error

100

10−1

10−2

10−3 9

25

49 121 225 Output observation points m PI oracle API r

Oracle Ar Ref. slope −1

441

784

PI surrogate

Figure 6: Oracle, PI oracle, and PI surrogate errors as a function of output resolution m (mL = m, µ = λ(m), lengthscale by LOO-CV). The PI oracle is smaller than the plain oracle at every tested m; both saturate for m > 441 at the O(h2 ) FD discretization floor. The PI surrogate saturates at ≈ 0.023 from m = 121, reflecting the irreducible learning error of Aoff .

collocation points beyond the output resolution gives no further benefit. This confirms that mL = m is a sufficient and essentially optimal choice.

7 Conclusion We studied kernel-based operator learning in the two-stage sampling framework of Batlle et al. (2024) and Sharma et al. (2026) and derived an explicit budget allocation rule relating the number N of training pairs, the number of input observations n, and the output resolution m. The rule quantifies how offline learning effort and online discretization must be balanced in order to achieve optimal approximation performance. It is obtained from a coupled error analysis based on interpreting the online reconstruction as recovery from perturbed data, which yields a decomposition into learning and reconstruction errors that can be analyzed independently. As a consequence, we recover the oracle convergence rate m−(σ−τ )/d whenever N scales appropriately with m and n. The theoretical predictions are validated on the Darcy flow benchmark. As a second contribution, we introduced a physics-informed extension of the online reconstruction stage. By augmenting the reconstruction functional with a soft PDE collocation penalty, the governing equation can be enforced at evaluation time for each new test input without retraining and without additional PDE solves. The resulting physics-informed Tikhonov functional admits a closed-form representer theorem. Numerical experiments for 27

R. Kempf

Mean relative L2 error

0.15

mL = m

0.1

0.05

0 9

25 49 Collocation points mL PI oracle API r

Plain oracle (0.132)

121

225

PI surrogate

Figure 7: Effect of the number of collocation points mL at fixed output resolution m = 49 (µ = λ(m), lengthscale by LOO-CV). The PI oracle error decreases up to mL = m = 49 and then stagnates; adding further collocation points yields no benefit. The choice mL = m is therefore both sufficient and essentially optimal.

the Poisson equation demonstrate consistent reductions in reconstruction error and indicate that mL = m collocation points provide an effective accuracy-cost tradeoff. Several theoretical questions remain open. Most notably, a convergence-rate analysis for the physics-informed surrogate would provide a rigorous characterization of the interplay between the training budget N , the number of input observations n, the reconstruction resolution m, and the number of collocation points mL . Such a result could lead to a physics-informed budget allocation rule and extend the present analysis beyond the purely data-driven setting

References R. Arcangeli, M. Cruz Lopez de Silanes, and J. J. Torrens. Extension of sampling inequalities to Sobolev semi-norms of fractional order and derivative data. Numer. Math., 121(3): 587–608, 2012. Pau Batlle, Matthieu Darcy, Bamdad Hosseini, and Houman Owhadi. Kernel methods are competitive for operator learning. Journal of Computational Physics, 496:112549, 2024. ISSN 0021-9991. doi: https://doi.org/10.1016/j.jcp.2023.112549. URL https: //www.sciencedirect.com/science/article/pii/S0021999123006447. Susanne C. Brenner and L. Ridgway Scott. The Mathematical Theory of Finite Element Methods. Springer New York, 2008. ISBN 9780387759340. doi: 10.1007/ 978-0-387-75934-0. URL http://dx.doi.org/10.1007/978-0-387-75934-0. 28

Kernel-based Operator Learning

Yifan Chen, Bamdad Hosseini, Houman Owhadi, and Andrew M. Stuart. Solving and learning nonlinear pdes with gaussian processes. Journal of Computational Physics, 447: 110668, 2021. ISSN 0021-9991. doi: https://doi.org/10.1016/j.jcp.2021.110668. URL https://www.sciencedirect.com/science/article/pii/S0021999121005635. Jean Duchon. Sur l’erreur d’interpolation des fonctions de plusieurs variables par les dm splines. RAIRO. Analyse numérique, 12(4):325–334, 1978. URL https://www.numdam. org/item/M2AN_1978__12_4_325_0/. Lawrence C Evans. Partial Differential Equations. Graduate Studies in Mathematics. American Mathematical Society, Providence, RI, 2 edition, March 2010. Gregory Fasshauer and Michael McCourt. Kernel-based Approximation Methods using MATLAB. WORLD SCIENTIFIC, June 2014. ISBN 9789814630146. doi: 10.1142/9335. URL http://dx.doi.org/10.1142/9335. Edward J. Fuselier. Sobolev-type approximation rates for divergence-free and curl-free rbf interpolants. Mathematics of Computation, 77(263):1407–1423, 2008. ISSN 00255718, 10886842. URL http://www.jstor.org/stable/40234564. Quoc Thong Le Gia, Ian Hugh Sloan, and Holger Wendland. Vector-valued gaussian processes for approximating divergence- or rotation-free vector fields, 2025. URL https: //arxiv.org/abs/2511.12535. Hachem Kadri, Emmanuel Duflos, Philippe Preux, Stéphane Canu, Alain Rakotomamonjy, and Julien Audiffren. Operator-valued kernels for learning from functional response data. Journal of Machine Learning Research, 17(20):1–54, 2016. URL http: //jmlr.org/papers/v17/11-315.html. E.J. Kansa. Multiquadrics—a scattered data approximation scheme with applications to computational fluid-dynamics—i surface approximations and partial derivative estimates. Computers and Mathematics with Applications, 19(8-9):127–145, 1990. ISSN 0898-1221. doi: 10.1016/0898-1221(90)90270-t. URL http://dx.doi.org/10.1016/0898-1221(90) 90270-T. Zongyi Li, Nikola Borislavov Kovachki, Kamyar Azizzadenesheli, Burigede Liu, Kaushik Bhattacharya, Andrew M. Stuart, and Anima Anandkumar. Fourier neural operator for parametric partial differential equations. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021. URL https://openreview.net/forum?id=c8P9NQVtmnO. Da Long, Nicole Mrvaljević, Shandian Zhe, and Bamdad Hosseini. A kernel framework for learning differential equations and their solution operators. Physica D: Nonlinear Phenomena, 460:134095, 2024. ISSN 0167-2789. doi: https://doi.org/10.1016/ j.physd.2024.134095. URL https://www.sciencedirect.com/science/article/pii/ S0167278924000460. 29

R. Kempf

Lu Lu, Pengzhan Jin, Guofei Pang, Zhongqiang Zhang, and George Em Karniadakis. Learning nonlinear operators via deeponet based on the universal approximation theorem of operators. Nature Machine Intelligence, 3(3):218–229, March 2021. ISSN 2522-5839. doi: 10. 1038/s42256-021-00302-5. URL http://dx.doi.org/10.1038/s42256-021-00302-5. Lu Lu, Xuhui Meng, Shengze Cai, Zhiping Mao, Somdatta Goswami, Zhongqiang Zhang, and George Em Karniadakis. A comprehensive and fair comparison of two neural operators (with practical extensions) based on fair data. Computer Methods in Applied Mechanics and Engineering, 393:114778, 2022. ISSN 0045-7825. doi: https://doi.org/10. 1016/j.cma.2022.114778. URL https://www.sciencedirect.com/science/article/ pii/S0045782522001207. B. Matérn. Spatial variation, volume 36 of Lecture Notes in Statistics. Springer, Berlin, second edition, 1986. ISBN 3-540-96365-0. doi: 10.1007/978-1-4615-7892-5. Charles A. Micchelli and Massimiliano Pontil. On learning vector-valued functions. Neural Computation, 17(1):177–204, January 2005. ISSN 1530-888X. doi: 10.1162/ 0899766052530802. URL http://dx.doi.org/10.1162/0899766052530802. Carlos Mora, Amin Yousefpour, Shirin Hosseinmardi, Houman Owhadi, and Ramin Bostanabad. Operator learning with gaussian processes. Computer Methods in Applied Mechanics and Engineering, 434:117581, 2025. ISSN 0045-7825. doi: https://doi.org/10. 1016/j.cma.2024.117581. URL https://www.sciencedirect.com/science/article/ pii/S0045782524008351. Francis J. Narcowich and Joseph D. Ward. Generalized hermite interpolation via matrixvalued conditionally positive definite functions. Mathematics of Computation, 63(208): 661–661, 1994. ISSN 0025-5718. doi: 10.1090/s0025-5718-1994-1254147-6. URL http: //dx.doi.org/10.1090/S0025-5718-1994-1254147-6. Francis J. Narcowich, Joseph D. Ward, and Holger Wendland. Sobolev bounds on functions with scattered zeros, with applications to radial basis function surface fitting. Mathematics of Computation, 74(250):743–763, 2005. Francis J Narcowich, Joseph D Ward, and Holger Wendland. Sobolev error estimates and a bernstein inequality for scattered data interpolation via radial basis functions. Constructive Approximation, 24(2):175–186, 9 2006. Francis J. Narcowich, Joseph D. Ward, and Grady B. Wright. Divergence-free rbfs on surfaces. Journal of Fourier Analysis and Applications, 13(6):643–663, October 2007. ISSN 1531-5851. doi: 10.1007/s00041-006-6903-2. URL http://dx.doi.org/10.1007/ s00041-006-6903-2. Nicholas H. Nelsen and Andrew M. Stuart. Operator learning using random features: A tool for scientific computing. SIAM Review, 66(3):535–571, 2024. doi: 10.1137/24M1648703. URL https://doi.org/10.1137/24M1648703. Houman Owhadi. Do ideas have shape? idea registration as the continuous limit of artificial neural networks. Physica D: Nonlinear Phenomena, 444:133592, 2023. ISSN 0167-2789. 30

Kernel-based Operator Learning

doi: https://doi.org/10.1016/j.physd.2022.133592. URL https://www.sciencedirect. com/science/article/pii/S0167278922002962. Houman Owhadi and Clint Scovel. Operator-Adapted Wavelets, Fast Solvers, and Numerical Homogenization: From a Game Theoretic Approach to Numerical Approximation and Algorithm Design. Cambridge University Press, October 2019. ISBN 9781108484367. doi: 10.1017/9781108594967. URL http://dx.doi.org/10.1017/9781108594967. C. Runge. Über die Zerlegung einer empirischen Funktion in Sinuswellen. Schlömilch Z., 52:117–123, 1905. Robert Schaback. Unsymmetric meshless methods for operator equations. Numerische Mathematik, 114(4):629–651, October 2009. ISSN 0945-3245. doi: 10.1007/ s00211-009-0265-z. URL http://dx.doi.org/10.1007/s00211-009-0265-z. Robert Schaback and Holger Wendland. Kernel techniques: From machine learning to meshless methods. Acta Numerica, 15:543–639, 2006. doi: 10.1017/S0962492906270016. Bernhard Schölkopf and Alexander J. Smola. Learning with Kernels: Support Vector Machines, Regularization, Optimization, and Beyond. The MIT Press, December 2001. ISBN 9780262256933. doi: 10.7551/mitpress/4175.001.0001. URL http://dx.doi.org/ 10.7551/mitpress/4175.001.0001. Ramansh Sharma, Matthew Lowery, Houman Owhadi, and Varun Shankar. Fluids you can trust: Property-preserving operator learning for incompressible flows, 2026. URL https://arxiv.org/abs/2602.15472. Ingo Steinwart and Andreas Christmann. Support Vector Machines. Springer New York, 2008. ISBN 9780387772424. doi: 10.1007/978-0-387-77242-4. URL http://dx.doi.org/ 10.1007/978-0-387-77242-4. Jian Sun and Wenshuai Wang. Optimizing shape parameters in rbf methods: A systematic review of techniques, applications, and computational challenges. Computer Science Review, 59:100842, 2026. ISSN 1574-0137. doi: https://doi.org/10.1016/j. cosrev.2025.100842. URL https://www.sciencedirect.com/science/article/pii/ S1574013725001182. Sifan Wang, Hanwen Wang, and Paris Perdikaris. Learning the solution operator of parametric partial differential equations with physics-informed DeepONets. Science Advances, 7(40):eabi8605, 2021. doi: 10.1126/sciadv.abi8605. Sifan Wang, Hanwen Wang, and Paris Perdikaris. Improved architectures and training algorithms for deep operator networks. Journal of Scientific Computing, 92(2), 2022. ISSN 1573-7691. doi: 10.1007/s10915-022-01881-0. URL http://dx.doi.org/10.1007/ s10915-022-01881-0. Holger Wendland. Piecewise polynomial, positive definite and compactly supported radial functions of minimal degree. Advances in Computational Mathematics, 4(1):389–396, December 1995. ISSN 1572-9044. doi: 10.1007/bf02123482. URL http://dx.doi.org/ 10.1007/BF02123482. 31

R. Kempf

Holger Wendland. Scattered Data Approximation. Cambridge Monographs on Applied and Computational Mathematics. Cambridge University Press, 2004. doi: 10.1017/ CBO9780511617539. Holger Wendland and Christian Rieger. Approximate interpolation with applications to selecting smoothing parameters. Numerische Mathematik, 101(4):729–748, August 2005. ISSN 0945-3245. doi: 10.1007/s00211-005-0637-y. URL http://dx.doi.org/10.1007/ s00211-005-0637-y.

32

Record · ID 346504 · SHA-256 fe1a77daf98899b3
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.