Conceptio › Archive › arXiv CS
arXiv CSopen access

Optimization Geometry of Equivalent Brownian RKHS Representations

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Optimization Geometry of Equivalent Brownian RKHS Representations Gustau Camps-Valls Image Processing Laboratory (IPL) Universitat de València Paterna, València 46980, Spain [email protected]

arXiv:2609.21693v1 [cs.LG] 18 Sep 2026

Mahdi Mohammadigohari Faculty of Engineering Free University of Bozen-Bolzano Bruno Buozzi 1, Bolzano 39100, Italy [email protected]

Abstract Equivalent finite parameterizations can represent the same functions and intrinsic norm yet induce different optimization algorithms. We study this effect in a controlled finite Brownian RKHS with nodal, increment, and spectral coordinates. Classical finite-element, RKHS-interpolation, Brownian-covariance, and mixed-boundary DCT identities make the shared hypothesis class, Brownian energy, approximation operator, and coordinate maps explicit. Our main results concern the optimization geometry of this fixed model. With mapped initialization, identical scalar steps, and identical minibatches, nodal and spectral GD/SGD have exactly the same mapped trajectories. Increment GD is an explicit Euler step for the constant Brownian/Sobolev metric, with factor 1/h. For Brownian-regularized least squares, κ2 (Hinc ) ≤ 1 + A/ρ, independently of grid resolution G for fixed A, ρ > 0, and the stated normalization. Under the stated standard-Adam convention, the universal orthogonal equivariance group is exactly the signed permutations; the block DCT-VIII transform is not one. Float64 tests over five grids numerically verify the finite identities, mapped one-layer and recursive trajectories, conditioning predictions, and theorem-matched Adam separation. Thus coordinate effects are isolated without changing the represented functions, intrinsic regularizer, or approximation space.

1

Introduction

Reproducing kernel Hilbert spaces describe a model through the functions it represents and the intrinsic norm controlling them (Aronszajn, 1950; Schölkopf and Smola, 2002; Berlinet and Thomas-Agnan, 2004; Steinwart and Christmann, 2008). Splines, Gaussian processes, random features, and neural tangent kernels all exploit this viewpoint (Wahba, 1990; Rasmussen and Williams, 2006; Rahimi and Recht, 2007; Jacot et al., 2018; Lee et al., 2019). A complementary line studies the geometry of optimization: natural gradient, mirror descent, implicit-bias analysis, and reparameterization theory show that equivalent descriptions of one model can induce different algorithms (Amari, 1998; Beck and Teboulle, 2003; Gunasekar et al., 2018; Amid and Warmuth, 2020; Kristiadi et al., 2023; Li et al., 2022; Martens, 2020). Coordinatewise adaptive methods add a basis dependence that ordinary Euclidean gradient descent does not have under orthogonal changes (Kingma and Ba, 2015; Ling et al., 2022; Zhang et al., 2025). These viewpoints are difficult to separate because a neural reparameterization often also changes initialization, implicit regularization, architecture, or the represented class. A controlled comparison instead needs several coordinate systems for one fixed Hilbert space. Brownian profile layers provide this setting: BKLs and VBKLs recursively use one-dimensional Brownian RKHS profiles (Mohammadigohari et al., 2026; Mohammadigohari, 2026), and the VBKL companion already employs anchored piecewise-linear profiles and their increment energy. The required ingredients are largely classical: piecewise-linear stiffness energies come from finite elements (Strang and Fix, 1973; Brenner and Scott, 2008); nodal interpolation and power functions from spline and kernel approximation (Wahba, 1990; Fasshauer, 2007; Wendland, 2005); Brownian covariance identities from Gaussian processes (Rasmussen and Williams, 2006); and the mixed-boundary spectrum from DCT theory (Martucci, 1994; Strang, 1999; Masera et al., 2017). We assemble them in one two-sided anchored normalization so that the represented space, norm, approximation error, and coordinate maps are simultaneously fixed and explicit, rather than claiming these ingredients separately as new. 1

In this controlled model, the spectral map is orthogonal whereas D⊤ 0 D0 = hK0 . We derive the resulting mapped optimizer trajectories, the constant-metric Brownian interpretation of increment descent, and its blockwise recursive consequence. We also obtain the fixed-A, ρ grid-resolution-independent least-squares bound and the maximal universal orthogonal equivariance group of the stated standard-Adam update. Our contributions are: 1. one controlled finite Brownian RKHS with exactly equivalent nodal, increment, and block DCT-VIII coordinates, preserving the hypothesis class, intrinsic norm, and approximation space; 2. exact mapped GD/SGD identities: nodal and spectral trajectories coincide, while increment descent is a constant-metric Brownian/Sobolev Euler step with the precise mesh scaling; 3. the blockwise recursive consequence and κ2 (Hinc ) ≤ 1 + A/ρ, independent of G under the stated fixed quantities; and 4. the signed-permutation maximality theorem for standard Adam, with source-traced float64 verification. EuroSAT and Salinas remain secondary external-validity studies, not coordinate-equivalence tests. Proofs and supplementary material are in the appendix; Table 3 summarizes the results.

2

Notation and Preliminaries

We write [n] = {1, . . . , n}, ⟨·, ·⟩ and ∥ · ∥2 for the Euclidean inner product and norm, In for the identity, ej for a canonical basis vector, δij for the Kronecker delta, and blkdiag for block-diagonal assembly. Vectors are bold lowercase, matrices bold uppercase, scalars italic; grid quantities are indexed from 0 and spectral quantities from 1. Three symbols are disambiguated relative to the underlying frameworks: increment coordinates are w, reserving δ for the Kronecker delta and the evaluation functional δt ; eigenvalues of the block Tm are νk , reserving µ for measures; and eigenvalues of the anchored stiffness matrix are λk , k ∈ [m], each of multiplicity two. Notation used only inside proofs is collected in Appendix A. BKLs construct recursive reproducing-kernel representations from the one-dimensional Brownian kernel (Mohammadigohari et al., 2026), |x| + |x′ | − |x − x′ | = min(|x|, |x′ |) ⊮{xx′ ≥ 0}, x, x′ ∈ R, (1) 2 the covariance of two-sided Brownian motion anchored at the origin. RIts RKHS on I, denoted HB , is the Cameron–Martin space HB = {f ∈ H 1 (I) : f (0) = 0} with ⟨f, g⟩HB = I f ′ g ′ (Berlinet and Thomas-Agnan, 2004). Given support functions S on an inputR space X , each u : X → R, and a probability measure µ on S, the Brownian integral kernel k[S, µ](x, x′ ) = S kB (u(x), u(x′ )) dµ(u) has an RKHS defining the next layer; iterating from linear projections yields a hierarchy of Brownian RKHSs. Variation Brownian Kernel Ladders (VBKLs) instead use an atomic representation (Mohammadigohari, 2026): composing support functions  with Brownian profiles yields a dictionary UL = g ◦ u : u ∈ UL−1 , g ∈ HB , ∥g∥HB ≤ 1 of depth-L atoms, R (L) with variation space V (L) = {F = UL u dµ(u) : ∥µ∥TV < ∞} and complexity Cbvar (F ) = inf{∥µ∥TV : F = R u dµ(u)}. In both, profile functions lie in infinite-dimensional Brownian RKHSs, so implementations UL must discretize them. The VBKL construction already employs anchored piecewise-linear profiles and their increment-energy constraint; the present paper makes the corresponding finite RKHS and coordinate metrics explicit for a controlled optimization comparison. kB (x, x′ ) =

3

Finite Brownian RKHS

We collect the classical finite-element and RKHS machinery in the paper’s two-sided anchored normalization (Strang and Fix, 1973; Brenner and Scott, 2008), ensuring that later optimizer comparisons use the same function space, norm, and approximation operator. Definition 1 (Finite Brownian profile space). Let I = [−A, A] for some A > 0, let G ∈ N, and let ti = −A+ih for i = 0, . . . , G be the uniform grid with mesh size h = 2A G , so that −A = t0 < t1 < · · · < tG = A. For each i = 0, . . . , G, let ϕi ∈ C(I) be the hat function that is affine on each subinterval [tj , tj+1 ] and satisfies ϕi (tj ) = δij for every j = 0, . . . , G (Brenner and Scott, 2008). Define the finite-element reconstruction operator 2

PG R : RG+1 → C(I) by R(v) = i=0 vi ϕi , where v = (v0 , . . . , vG )⊤ ∈ RG+1 . The finite Brownian profile space is Hh = R(RG+1 ) = span(ϕ0 , . . . , ϕG ). Since ϕi (tj ) = δij , the family {ϕi }G i=0 is linearly independent, so every f ∈ Hh has a unique coefficient vector v ∈ RG+1 with f = R(v). Definition 2 (Brownian energy form). On Hh we define the symmetric bilinear form ⟨f, g⟩B,h = RA ′ f (t)g ′ (t) dt, with derivatives taken elementwise on each [ti , ti+1 ] and hence defined almost every−A p where, and the Brownian seminorm ∥f ∥B,h = ⟨f, f ⟩B,h . Proposition 1 (Finite Brownian energy representation). Let D ∈ RG×(G+1) be the first-difference matrix, (Dv)i = vi+1 − vi for i = 0, . . . , G − 1, and let K = (Krs )G r,s=0 with Krs = ⟨ϕr , ϕs ⟩B,h be the stiffness matrix. For f = R(v) ∈ Hh , 2

∥f ∥B,h = v⊤ Kv =

G−1 1 1 X 2 2 ∥Dv∥2 = (vi+1 − vi ) . h h i=0

(2)

Since both matrices are symmetric, this is equivalent to the matrix identity K = h1 D⊤ D. This recovers the classical stiffness-matrix form of the Dirichlet energy (Strang and Fix, 1973; Brenner and Scott, 2008); we record it because it identifies the regularizer that finite realizations implement. PG−1 Corollary 1 (Implemented Brownian regularizer). For v ∈ RG+1 let Ωh (v) := i=0 (vi+1 − vi )2 . Then Ωh (v) = h∥R(v)∥2B,h : the implemented regularizer is the Brownian energy times the fixed mesh factor h > 0. 3.1

Anchored Finite Brownian RKHS

The Brownian energy is constant-invariant and hence only a seminorm on Hh . As in the Cameron–Martin space (Berlinet and Thomas-Agnan, 2004), uniqueness is recovered by fixing the value at a reference point. Assumption 1. G = 2m with m ∈ N, and i0 := m, so ti0 = 0. In force for the remainder of the paper. Definition 3 (Anchored finite Brownian profile space). Hh0 = {f ∈ Hh : f (0) = 0}; equivalently f = R(v) ∈ Hh0 iff vi0 = 0, and dim Hh0 = G. Since the anchor removes one coordinate, we work with a free parameter in RG . The ordering below traverses each half-grid from its exterior endpoint toward the anchor, which is what makes the two diagonal blocks of Theorem 3 the same matrix. Definition 4 (Reduced anchored coordinates). For an anchored nodal vector v ∈ RG+1 with vi0 = 0, the e := (v0 , . . . , vm−1 , v2m , v2m−1 , . . . , vm+1 )⊤ ∈ RG , and the anchoring reconstruction reduced anchored vector is v matrix is R := (e0 , . . . , em−1 , e2m , e2m−1 , . . . , em+1 ) ∈ R(G+1)×G , so that Re v = v. We write K0 := R⊤ KR G×G for the anchored stiffness matrix and D0 := DR ∈ R . The anchored energy is positive definite, which is what makes the anchored space an inner-product space rather than merely a seminormed one. Lemma 3.1 (Positive definiteness). ker K = span(1) and K0 ≻ 0. Consequently ⟨·, ·⟩B,h is an inner product on Hh0 . We can now record the Brownian-specific kernel formula and its covariance interpretation. The value of the result here is its exact normalization and its use in the subsequent controlled comparison, rather than finite-dimensional RKHS existence itself. Theorem 1 (Finite Brownian RKHS and its reproducing kernel). The space (Hh0 , ⟨·, ·⟩B,h ) is a G-dimensional RKHS. Let J := {0, . . . , G} \ {i0 } and let π : J → [G] give the position of a node in the reduced ordering of e := (ϕπ−1 (1) (t), . . . , ϕπ−1 (G) (t))⊤ , its reproducing kernel satisfies: Definition 4. With ϕ(t) e ⊤ K−1 ϕ(t) e for all s, t ∈ I. (i) kh (s, t) = ϕ(s) 0

3

(ii) (K−1 0 )π(r),π(q) = kB (tr , tq ) for r, q ∈ J , and kh (·, tr ) = kB (·, tr ) for every grid node. More generally, kh is the tensor-product piecewise-bilinear interpolant of kB on the grid, and therefore agrees with kB whenever at least one argument is a node. Corollary 2 (Conforming Brownian interpolation space). Hh0 = span {kB (·, tr ) : r ∈ J } ⊂ HB ,

∥g∥B,h = ∥g∥HB

(g ∈ Hh0 ).

(3)

Thus the inclusion is isometric, and the HB -orthogonal projection Πh : HB → Hh0 is exactly nodal interpolation. Corollary 3 (Exact residual kernel and sharp approximation). Let rh := kB − kh and let Ti = [ai , bi ] = [ti , ti+1 ]. Then   (min{s, t} − ai ) (bi − max{s, t}) , s, t ∈ T , i h rh (s, t) = (4)  0, s, t lie in different grid elements. Consequently, for t ∈ Ti the power function is ph (t)2 := rh (t, t) =

(t − ai )(bi − t) , h

|f (t) − Πh f (t)| ≤ ph (t)∥f ∥HB .

(5)

For every f ∈ HB , ∥f − Πh f ∥2HB = ∥f ∥2HB − ∥Πh f ∥2HB , √ h ∥f − Πh f ∥∞ ≤ ∥f ∥HB , 2

(6) ∥f − Πh f ∥L2 (I) ≤

h ∥f ∥HB . π

(7)

Both constants are sharp over the unit ball. These quantities depend only on the subspace Hh0 , so they are identical under every coordinate realization of that space. The residual has the Brownian-bridge covariance form, and the bounds are standard power-function and sharp interval-inequality consequences (Rasmussen and Williams, 2006; Fasshauer, 2007; Brenner and Scott, 2008); their role is to certify identical approximation across coordinates.

4

Equivalent Coordinate Representations

The finite Brownian RKHS admits nodal, increment, and spectral realizations. Their invertible linear relation is elementary; the substantive control is that they realize exactly the same space, Brownian norm, and approximation operator. Consequently, any mapped optimization difference must arise from the optimizer or the coordinate metric rather than representational capacity. 4.1

Nodal and Increment Coordinate Representations

We first consider the nodal finite-element realization and its associated increment representation. Definition 5 (Nodal coordinates). Every function f ∈ Hh0 admits the unique nodal representation f = R(Re v) 2 e ∈ RG , and ∥f ∥B,h = v e⊤ K0 v e = v⊤ Kv for v = Re e ∈ RG is with reduced anchored vector v v. Throughout, v the free parameter; the anchor is not a trainable coordinate. The nodal realization is the parameterization induced directly by the finite-element discretization: each trainable parameter is the value of the profile at a grid point. Definition 6 (Increment coordinates). The increment coordinates associated with the reduced nodal vector e are defined by w = D0 v e, or equivalently, wi = vi+1 − vi for i = 0, . . . , G − 1 with v = Re v v. The increment realization stores local differences rather than grid values. Because the profile is anchored, D0 is invertible (Theorem 2ii) and the increments determine f uniquely. In these coordinates the Brownian energy is simply ∥f ∥2B,h = h1 ∥w∥22 .

4

4.2

Spectral Coordinate Representation

A complementary realization is obtained by diagonalizing the anchored Brownian energy, yielding an orthogonal spectral parameterization of the finite Brownian RKHS. Definition 7 (Spectral coordinates). By Lemma 3.1, K0 ≻ 0. Let K0 = QΛQ⊤ be a fixed orthogonal eigendecomposition, with Q ∈ RG×G orthogonal and Λ = diag(λ1 , . . . , λG ) having strictly positive diagonal. e ∈ RG , define c = Q⊤ v e, or equivalently, v e = Qc. For the reduced anchored nodal vector v These coordinates diagonalize the Brownian norm, so each spectral coefficient contributes independently to the total energy. The reconstruction and norm identities are proved in Appendix B.4; Section 5 gives the explicit DCT-VIII block-structured choice of Q. The three coordinate systems differ in how they store a Brownian profile: the nodal realization stores profile values, the increment realization stores local differences, and the spectral realization stores coefficients in an orthogonal Brownian basis. The following theorem shows that they are nevertheless equivalent at the level of the function space they represent. e ∈ RG denote its reduced Theorem 2 (Equivalence of implementations). Let f ∈ Hh0 be arbitrary, let v G+1 G×G anchored nodal coefficient vector, and set v = Re v∈R and D0 := DR ∈ R . Define the increment e and c = Q⊤ v e, respectively. Then the following statements hold. and spectral coordinates by w = D0 v (i) The nodal, increment, and spectral representations determine exactly the same function f ∈ Hh0 . e 7→ w and v e 7→ c are invertible linear transformations. Consequently, each coordinate system (ii) The maps v uniquely determines the other two. (iii) The Brownian norm admits the equivalent representations 1 2 2 ∥f ∥B,h = v⊤ Kv = ∥w∥2 = c⊤ Λc. h

(8)

(iv) The three realizations therefore represent the same hypothesis class and, by Corollary 3, share the same approximation properties. (v) The nodal-to-spectral transformation is an isometry for the Euclidean metric on the reduced anchored coordinates, whereas the nodal-to-increment transformation satisfies D⊤ 0 D0 = hK0 . Consequently, for G = 2m ≥ 4, the increment coordinate map is nonorthogonal, whereas the spectral coordinate map is orthogonal. The restriction G ≥ 4 excludes only G = 2, where hK0 = I2 and the increment map is itself orthogonal. Theorem 2 gives both the function-space equivalence and the exact relation between the induced Euclidean metrics: the spectral map is orthogonal, the increment map induces the metric hK0 . Section 6 shows that this one matrix identity determines which optimizers see the three realizations as the same problem.

5

Explicit Spectral Representation

Using classical DCT nomenclature (Martucci, 1994; Strang, 1999; Masera et al., 2017), we give all eigenpairs of the fixed mixed-boundary stiffness matrix in closed block DCT-VIII form; “complete” does not mean a new transform family. 5.1

Block Decomposition

After anchoring, the stiffness matrix decouples into two independent tridiagonal blocks. The reason is that the anchor node ti0 = 0 is the unique node coupling the two half-grids; removing it severs them. In the reduced ordering of Definition 4, which traverses both half-grids from the exterior endpoint inward, the two blocks are literally the same matrix. m×m Theorem 3 (Complete DCT-VIII eigendecomposition). Let Tm = (τrs )m−1 be the tridiagonal r,s=0 ∈ R matrix with τ00 = 1, τrr = 2 for 1 ≤ r ≤ m − 1, τrs = −1 for |r − s| = 1, and τrs = 0 otherwise. Then K0 =  (2k−1)π 1 − 2 cos(θk ) = 4 sin2 θ2k , h blkdiag(Tm , Tm ) (Lemma D3). For each k ∈ [m], define θk = 2m+1 and νk = 2   2 and let the vector qk ∈ Rm be defined componentwise by qk (j) = √2m+1 cos j + 12 θk for j = 0, . . . , m − 1. Then the following statements hold.

5

(i) The vectors q1 , . . . , qm form an orthonormal basis of Rm . (ii) With Qm = [q1 , . . . , qm ], the matrix Tm admits the eigendecomposition Tm = Qm diag(ν1 , . . . , νm )Q⊤ m. (iii) The anchored stiffness matrix admits the orthogonal eigendecomposition K0 = QΛQ⊤ , where Q = blkdiag(Qm , Qm ) and Λ = h1 diag(ν1 , . . . , νm , ν1 , . . . , νm ). (2k−1)π  (iv) The eigenvalues of K0 are λk = h4 sin2 2(2m+1) for k ∈ [m], each of multiplicity two. The eigendecomposition induces the basis ψi := R(RQei ) of Hh0 , satisfying ⟨ψi , ψj ⟩B,h = λi δij (Lemma D6). Hence the finite RKHS kernel has the spectral basis expansion kh (s, t) =

G X

λ−1 i ψi (s)ψi (t).

(9)

i=1

This is a finite RKHS expansion; no input measure or integral-operator eigenproblem is required. Remark 1. Every eigenvalue has multiplicity two, so the eigenbasis is not canonical inside each left–right eigenspace. We fix Q = blkdiag(Qm , Qm ), the explicit block-structured choice matching the reduced ordering. Any such orthogonal choice gives the same gradient descent trajectory after mapping, but coordinatewise adaptive methods may distinguish the choices. Corollary 4 (Exact stiffness and coordinate conditioning). Under Assumption 1,   π cos2 2m+1 4   = 2 (2m + 1)2 + O(1) = Θ(G2 ), κ2 (K0 ) = π π sin2 4m+2     π π σmin (D0 ) = 2 sin , σmax (D0 ) = 2 cos , 4m + 2 2m + 1   π cos 2m+1 p   = κ2 (K0 ) = Θ(G). κ2 (D0 ) = π sin 4m+2

(10)

(11)

(12)

The condition numbers are independent of A. Moreover, hλmax (K0 ) → 4 and λmin (K0 ) ∼ π 2 /[h(2m + 1)2 ] as m → ∞. The first identity gives the classical O(h−2 ) stiffness conditioning; the second and third quantify the anisotropy of the actual nodal-to-increment coordinate map.

6

Coordinates and Optimizers

We begin with the standard chain-rule reparameterization identity (Gunasekar et al., 2018; Amid and Warmuth, 2020; Kristiadi et al., 2023). Let L : RG → R be differentiable in reduced nodal coordinates and, for an invertible matrix A, define LA (z) := L(A−1 z). Proposition 2 (Linear reparameterization is exact preconditioning). If zk+1 = zk − ηk ∇LA (zk ) and ek = A−1 zk , then v ek+1 = v ek − ηk (A⊤ A)−1 ∇L(e v vk ).

(13)

The same identity holds for stochastic gradients computed from the same minibatches and transformed by the chain rule. Proposition 3 (Coordinate–optimizer correspondence). For corresponding initializations, identical scalar step sizes, identical minibatches, and exact arithmetic: e and Q is orthogonal, nodal and spectral gradient descent produce identical mapped (i) Since c = Q⊤ v iterates and hence identical functions at every step.

6

e and D⊤ (ii) Since w = D0 v 0 D0 = hK0 , increment gradient descent maps to ηk ek+1 = v ek − K−1 ∇L(e vk ). (14) v h 0 Thus it is an explicit Euler step generated by the constant Brownian/Sobolev metric with Gram matrix K0 , with time step ηk /h. We use “Riemannian-gradient step” only in this discrete, constant-metric sense; the statement is not equality with the continuous gradient flow. Corollary 5 (Blockwise extension to recursive architectures). Let an arbitrary differentiable come(1) , . . . , v e(B) and possibly other parameters, and apply the posed objective depend on profile blocks v block transformation A = blkdiag(A1 , . . . , AB , I). Then Proposition 2 holds with block preconditioner −1 −1 , . . . , (A⊤ , I). Hence simultaneous spectral changes leave the full GD/SGD trajecblkdiag((A⊤ 1 A1 ) B AB ) tory invariant, while simultaneous increment coordinates implement blockwise Brownian Riemannian-gradient steps with the corresponding per-block factors 1/hb . No quadratic or layer-separability assumption is required. This is the block-diagonal application of Proposition 2; cross-block coupling does not invalidate the identity. Here “mesh-independent” means independent of G for fixed A, ρ > 0, sample domain, and normalization, not uniform in A or ρ. The previous identities describe trajectories. The next result gives the corresponding quantitative conditioning statement for the regularized regression problem generated by the finite RKHS. Proposition 4 (Mesh-independent increment conditioning). For samples x1 , . . . , xn ∈ I, targets y1 , . . . , yn ∈ R, and ρ > 0, consider n 2 ρ 1 Xe e − yi + v e ⊤ K0 v e. L(e v) = ϕ(xi )⊤ v (15) 2n i=1 2 e, the Hessian Hinc Nodal and spectral Hessians are orthogonally similar. In increment coordinates w = D0 v satisfies ρ+A A ρ IG ⪯ Hinc ⪯ IG , κ2 (Hinc ) ≤ 1 + , (16) h h ρ independently of G. For the pure Brownian quadratic, the increment condition number is exactly 1, whereas the nodal and spectral condition numbers equal κ2 (K0 ) = Θ(G2 ). Standard Adam requires a separate statement because momentum prevents its update from being represented as a diagonal matrix times only the current gradient. Proposition 5 (Exact orthogonal equivariance group of standard Adam). Consider standard Adam with scalar β1 , β2 ∈ [0, 1), scalar ε > 0, zero initial moment states, bias correction, a positive first step size, and no weight decay, clipping, or other modification. Among orthogonal coordinate changes, Adam is equivariant for every gradient sequence and corresponding initialization if and only if the coordinate change is a signed permutation. For m ≥ 2, the DCT-VIII matrix Q is not a signed permutation; therefore there exists even a linear objective for which corresponding nodal and spectral Adam runs differ at the first step. This maximality statement strengthens basis dependence to an if-and-only-if universal group under the exact convention; it does not cover AdamW, clipping, nonzero moments, or matrix-valued adaptivity. Thus GD gives exact nodal–spectral agreement; increment GD gives exact Brownian preconditioning; and Adam is permitted, but not forced, to separate the orthogonally related realizations. A separation theorem does not impose a universal ordering, and independently sampled initializations do not test trajectory equivalence. These predictions require mapped initial states, the same data order and scalar step sequence, and the same intrinsic Brownian penalty.

7

Experiments

We separate direct theorem-verification tests from predictive comparisons. The verification suite comprises E0, the one-layer optimization tests E1A–E1C, and the recursive extension E2. All five tests use frozen configurations, float64 arithmetic, one common represented initialization, exact coordinate maps, the same scalar step sequences, and, for SGD, identical minibatches; no identity-test arm is tuned separately. Secondary EuroSAT and spatially separated Salinas studies compare different model classes and are reported only in the appendix. 7

1-layer N/S 1-layer Inc./Brown. Recursive N/S Recursive Inc./Brown. Tolerance 10 8

10 11 10 13 10 15

(b) Conditioning versus mesh Pure N/S Pure Inc. LS N/S LS Inc. Bound 1 + A/

103 102 101 100

8

16

32

64

Maximum grid size G

128

8

16

32

Grid size G

64

128

Mapped nodal--spectral 2 discrepancy

10 9

104

Condition number

Maximum diagnostic residual

(a) Mapped GD/SGD identities

10 1 10 3 10 5 10 7 10 9 10 11 10 13 10 15

(c) First-step basis dependence

Linear GD Linear Adam Learning Adam

8

16

32

Grid size G

64

128

Figure 1: Direct numerical verification of the coordinate–optimizer theory. (a) Each marker is the maximum over the recorded parameter, function, objective, gradient, and preconditioner diagnostics and over full-batch GD and shared-minibatch SGD. The recursive markers correspond to grid triples (8, 16, 32) and (32, 64, 128), plotted at their maximum G; dashed segments guide the eye. All measured residuals remain far below the predeclared 10−8 tolerance. (b) For the pure Brownian quadratic, nodal/spectral conditioning grows with the mesh whereas increment conditioning is exactly one. Least-squares markers are five-seed means; increment conditioning remains independent of G under the stated fixed quantities and below 1 + A/ρ = 21. (c) Ordinary GD remains equivariant, whereas standard Adam has a nonzero mapped nodal–spectral first-step discrepancy. Learning-objective markers are five-seed means and bars span the minimum-to-maximum range. The ordinate is a coordinate-mapping discrepancy, not a predictive loss or error. These results establish permitted basis dependence, not a universal performance ordering. Table 1: Float64 verification of the exact finite-Brownian and mapped-optimizer identities. The last four rows report maxima over the listed diagnostics; these are identity checks, not predictive-performance comparisons.

Check

Scope

Maximum residual Type

D⊤ 0 D0 = hK0 K0 = QΛQ⊤ Closed-form spectrum K−1 versus Brownian Gram 0 Coordinate reconstruction Intrinsic Brownian energy One-layer Nodal–Spectral GD/SGD One-layer Increment–Brownian GD/SGD Recursive Nodal–Spectral GD/SGD Recursive Increment–Brownian GD/SGD

5 grids 5 grids 5 grids 5 grids 125 random trials 125 random trials 50 mapped runs 50 mapped runs 20 mapped runs 20 mapped runs

0 9.99 × 10−15 1.41 × 10−13 0 3.12 × 10−14 3.99 × 10−15 5.02 × 10−15 1.77 × 10−14 2.57 × 10−15 8.10 × 10−15

absolute relative relative absolute absolute relative diagnostic envelope diagnostic envelope diagnostic envelope diagnostic envelope

Finite and mapped identities. Across G ∈ {8, 16, 32, 64, 128}, the matrix, DCT-VIII spectrum, BrownianGram inverse, reconstruction, and intrinsic-energy residuals remain at floating-point precision. Mapped nodal–spectral full-batch GD and shared-minibatch SGD agree in parameters, functions, objectives, and gradients. Increment GD agrees with the correctly scaled Brownian-preconditioned nodal update. The same conclusions hold for a recursively coupled three-profile model with all non-profile parameters trained and shared across arms. Conditioning and standard Adam. For the pure Brownian quadratic, nodal/spectral conditioning grows from 29.3 at G = 8 to 6740.7 at G = 128, whereas increment conditioning is exactly one. For Brownianregularized least squares with ρ = 0.1, the largest measured increment condition number is 5.64, below the stated G-independent bound 1 + A/ρ = 21, while nodal/spectral conditioning and fixed-step iteration counts increase with G. A functional standard-Adam implementation matches torch.optim.Adam, but mapped nodal and spectral first steps differ on both the deterministic linear counterexample and a nontrivial regularized learning objective. This demonstrates permitted basis dependence, not a universal performance ordering.

8

Table 2: Conditioning, fixed-step convergence, and standard-Adam basis dependence across meshes. “N/S” denotes identical nodal/spectral quantities and “Inc.” the increment realization. Iteration counts use the analytic optimal scalar step and relative objective-gap tolerance 10−8 ; least squares uses ρ = 0.1, with increment bound 1 + A/ρ = 21. The last columns report mapped nodal–spectral first-step discrepancies. G

κ(K0 )

Pure iters N/S / Inc. LS κ N/S / Inc. LS iters N/S / Inc. Adam ∆1 linear Adam ∆1 learning

8 29.3 16 113.5 32 437.7 64 1.7 × 103 128 6.7 × 103

123.2 / 1.0 467.0 / 1.0 1795.2 / 1.0 7004.2 / 1.0 27629.0 / 1.0

6.94 / 4.86 24.45 / 5.11 92.15 / 5.16 357.64 / 5.17 1408.71 / 5.17

31.2 / 21.6 109.6 / 22.6 413.0 / 23.4 1601.8 / 23.4 6310.2 / 23.4

0.02 0.04 0.06 0.09 0.13

7.09 × 10−3 0.01 0.02 0.03 0.04

External validity. Appendix G reports independently audited EuroSAT and spatially separated Salinas BKL-versus-MLP robustness studies. They provide secondary predictive evidence, not tests of coordinate equivariance.

8

Conclusion

Equivalent coordinates of one finite function space need not define equivalent optimizers. Classical finiteelement, interpolation, covariance, and DCT identities provide our fixed Brownian comparison model. Within it, nodal and spectral GD/SGD coincide; increment descent is a 1/h-scaled constant-metric Brownian/Sobolev Euler step, also blockwise; the least-squares condition number is at most 1 + A/ρ independently of G under the stated fixed quantities; and signed permutations are the maximal universal orthogonal equivariances of standard Adam. Float64 tests numerically verify these predictions. Limitations. Closed forms assume a one-dimensional uniform grid and fixed anchor. The least-squares bound is not uniform in A or ρ or a nonlinear convergence rate; the Adam theorem covers only its stated update and gives no coordinate ranking. EuroSAT and Salinas compare model classes rather than coordinates, and eigenvalue multiplicity makes the displayed DCT-VIII basis noncanonical within each two-dimensional eigenspace.

9

References Shun-ichi Amari. Natural gradient works efficiently in learning. Neural Computation, 10(2):251–276, 1998. doi: 10.1162/089976698300017746. Ehsan Amid and Manfred K. Warmuth. Reparameterizing mirror descent as gradient descent. In Advances in Neural Information Processing Systems, volume 33, pages 8430–8439, 2020. Nachman Aronszajn. Theory of reproducing kernels. Transactions of the American Mathematical Society, 68 (3):337–404, 1950. Amir Beck and Marc Teboulle. Mirror descent and nonlinear projected subgradient methods for convex optimization. Operations Research Letters, 31(3):167–175, 2003. Alain Berlinet and Christine Thomas-Agnan. Reproducing Kernel Hilbert Spaces in Probability and Statistics. Kluwer Academic Publishers, 2004. Susanne C. Brenner and L. Ridgway Scott. The Mathematical Theory of Finite Element Methods. Springer, 3 edition, 2008. Gregory E. Fasshauer. Meshfree Approximation Methods with MATLAB. World Scientific, 2007. Suriya Gunasekar, Jason D. Lee, Daniel Soudry, and Nathan Srebro. Characterizing implicit bias in terms of optimization geometry. In Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pages 1832–1841, 2018. Arthur Jacot, Franck Gabriel, and Clément Hongler. Neural tangent kernel: Convergence and generalization in neural networks. In Advances in Neural Information Processing Systems, volume 31, 2018. Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In International Conference on Learning Representations, 2015. Agustinus Kristiadi, Felix Dangel, and Philipp Hennig. The geometry of neural nets’ parameter spaces under reparametrization. In Advances in Neural Information Processing Systems, volume 36, 2023. Jaehoon Lee, Lechao Xiao, Samuel S. Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington. Wide neural networks of any depth evolve as linear models under gradient descent. In Advances in Neural Information Processing Systems, volume 32, 2019. Zhiyuan Li, Tianhao Wang, Jason D. Lee, and Sanjeev Arora. Implicit bias of gradient descent on reparametrized models: On equivalence to mirror descent. In Advances in Neural Information Processing Systems, volume 35, 2022. Selena Ling, Nicholas Sharp, and Alec Jacobson. VectorAdam for rotation equivariant geometry optimization, 2022. James Martens. New insights and perspectives on the natural gradient method. Journal of Machine Learning Research, 21(146):1–76, 2020. Stephen A. Martucci. Symmetric convolution and the discrete sine and cosine transforms. IEEE Transactions on Signal Processing, 42(5):1038–1051, 1994. Maurizio Masera, Maurizio Martina, and Guido Masera. Odd type DCT/DST for video coding: Relationships and low-complexity implementations. In 2017 IEEE International Workshop on Signal Processing Systems (SiPS), pages 1–6, 2017. doi: 10.1109/SiPS.2017.8110009. Mahdi Mohammadigohari. Variation brownian kernel ladders, 2026. URL https://arxiv.org/abs/2608. 13882. Mahdi Mohammadigohari, Giuseppe Di Fatta, Giuseppe Nicosia, and Panos M. Pardalos. Brownian kernel ladders, 2026. URL https://arxiv.org/abs/2606.15812. 10

Ali Rahimi and Benjamin Recht. Random features for large-scale kernel machines. In Advances in Neural Information Processing Systems, volume 20, 2007. Carl Edward Rasmussen and Christopher K. I. Williams. Gaussian Processes for Machine Learning. MIT Press, 2006. Bernhard Schölkopf and Alexander J. Smola. Learning with Kernels: Support Vector Machines, Regularization, Optimization, and Beyond. MIT Press, 2002. Ingo Steinwart and Andreas Christmann. Support Vector Machines. Springer, 2008. Gilbert Strang. The discrete cosine transform. S0036144598336745.

SIAM Review, 41(1):135–147, 1999.

doi: 10.1137/

Gilbert Strang and George J. Fix. An Analysis of the Finite Element Method. Prentice-Hall, 1973. Grace Wahba. Spline Models for Observational Data. Society for Industrial and Applied Mathematics, 1990. Holger Wendland. Scattered Data Approximation. Cambridge University Press, 2005. Tianyue H. Zhang, Lucas Maes, Alan Milligan, Alexia Jolicoeur-Martineau, Ioannis Mitliagkas, Damien Scieur, Simon Lacoste-Julien, and Charles Guille-Escuret. Understanding Adam requires better rotation-dependent assumptions. In Advances in Neural Information Processing Systems, 2025. AI use statement Generative AI tools were used to assist with the writing and revision of mathematical proofs, feedback on experimental design, editing of research code, language editing, LaTeX formatting, and literature search. All AI-assisted material incorporated into the paper was reviewed and checked for correctness. The authors take responsibility for the final content of this work, including its mathematical claims, proofs, code, experiments, citations, and other artifacts.

11

A

Additional Notation Table 3: Summary of the main theoretical results.

Result

Content

Proposition 1 Theorem 1 Corollaries 2 and 3 Theorem 2 Theorem 3 and Corollary 4 Propositions 2 and 3 Corollary 5 Proposition 4 Proposition 5

Classical stiffness identity in the Brownian normalization Brownian-specific kernel and covariance–precision formula Classical interpolation consequences with exact constants Exact equivalence of the three realizations Closed-form block DCT-VIII specialization Standard reparameterization and Brownian metric consequence Exact blockwise recursive extension G-independent conditioning for fixed A and ρ Maximal universal orthogonal group of standard Adam

The reduced anchored ordering, the anchoring reconstruction matrix R, the anchored stiffness matrix K0 = R⊤ KR, and D0 = DR are now defined in the main text (Definition 4), together with the standing assumption G = 2m, i0 = m (Assumption 1); they are not repeated here. For reference, the reduced ordering is ⊤

e := (v0 , . . . , vm−1 , v2m , v2m−1 , . . . , vm+1 ) ∈ R2m , v

Re v = v,

(17)

so that both half-grids run from their exterior endpoint toward the anchor, and R := (e0 , . . . , em−1 , e2m , e2m−1 , . . . , em+1 ) ∈ R(2m+1)×2m .

(18)

That the Brownian energy has a one-dimensional constant nullspace, so that anchoring yields K0 ≻ 0, is established by Lemmas 3.1 and D2; it is proved rather than assumed. The first-difference operator D : RG+1 → RG is defined by (Dv)i = vi+1 − vi for i = 0, . . . , G − 1, and the RA ′ ′ finite-element stiffness matrix K = (Krs )G r,s=0 entrywise by Krs = −A ϕr (t)ϕs (t) dt.

B

Proofs

This section proves the foundational finite Brownian and coordinate results. We first establish the finite Brownian energy and stiffness representation in Proposition 1 and its implemented-regularizer consequence in Corollary 1. We then prove the G-dimensional Hilbert-space and reproducing-kernel structure asserted in Theorem 1; the explicit kernel formula and Brownian Gram identification in items i–ii are isolated in Section C. Next, we prove the exact equivalence of the nodal, increment, and spectral realizations in Theorem 2, including the coordinate maps, Brownian-energy identities, and induced metric relations. Finally, we derive the complete DCT-VIII eigendecomposition in Theorem 3. The auxiliary stiffness, nullspace, block-decomposition, and coordinate-energy results used throughout are collected in Section D. B.1

Proof of Proposition 1

Since A > 0 and the grid is strictly increasing, its mesh size satisfies h = 2A G > 0. Because f = R(v), the definition of the reconstruction operator gives, for every t ∈ I, (a)

G (b) X

f (t) = R(v)(t) =

vi ϕi (t).

(19)

i=0

Here (a) uses the assumed identity f = R(v), and (b) is the definition of R. Each basis function ϕi is continuous and piecewise affine on the finite grid, and is therefore absolutely continuous on I and differentiable away from the grid points. Consequently, f is absolutely continuous and, for almost every t ∈ I, !′ G G X (a) (b) X ′ f (t) = vi ϕi (t) = vi ϕ′i (t). (20) i=0

i=0

12

A. FINITE BROWNIAN RKHS Lemma C1 Structure of the stiffness matrix Lemma C2 Kernel of the stiffness matrix

Proposition 1 Finite Brownian energy representation

Lemma 3.1 Positive definiteness of K0

Corollary 1 Implemented Brownian regularizer

Theorem 1 Finite RKHS and reproducing kernel

Corollary 2 Conforming Brownian interpolation space

Corollary 3 Exact residual kernel and sharp approximation

B. EQUIVALENT COORDINATE REALIZATIONS Lemma C5 Brownian energy in spectral coordinates

Lemma C3 Block decomposition

Lemma C4 Brownian energy in increment coordinates

Theorem 2 Equivalence of implementations

C. EXPLICIT SPECTRUM AND CONDITIONING Theorem 3 Complete DCT-VIII eigendecomposition Corollary 4 Exact stiffness and coordinate conditioning

Lemma C6 Explicit spectral Brownian basis

D. OPTIMIZATION CONSEQUENCES Proposition 2 Linear reparameterization is exact preconditioning

Proposition 5 Exact orthogonal equivariance group of standard Adam

Proposition 3 Coordinate–optimizer correspondence

Corollary 5 Recursive blockwise extension

Proposition 4 Mesh-independent increment conditioning

Figure 2: Proof-dependency structure of the theoretical results. The displayed artwork is the supplied dependency graph without redrawing or re-typesetting; panel headings and result nodes are clickable and jump to the corresponding section or statement.

Here (a) differentiates (19) at a point at which all basis derivatives exist, and (b) uses the linearity of differentiation and the fact that each coefficient vi is independent of t. The possible failure of differentiability at the finitely many grid points does not affect any of the integrals below.

13

We first identify the stiffness-matrix representation of the Brownian energy. By Definition 2, Z A (a) (b) 2 ∥f ∥B,h = ⟨f, f ⟩B,h = f ′ (t)f ′ (t) dt −A

(c)

Z A

G X

−A

r=0

=

(d)

=

! vr ϕ′r (t)

G X

! vs ϕ′s (t)

dt

s=0

Z AX G X G

vr vs ϕ′r (t)ϕ′s (t) dt

−A r=0 s=0 G G (e) X X

=

Z A vr vs

r=0 s=0 G G (f ) X X

=

ϕ′r (t)ϕ′s (t) dt

−A G G (g) X X

vr ⟨ϕr , ϕs ⟩B,h vs =

r=0 s=0

vr Krs vs

r=0 s=0

(h)

= v⊤ Kv.

(21)

Here (a) is the definition of the Brownian seminorm associated with the symmetric bilinear energy form, (b) is the definition of that energy form, (c) substitutes (20), (d) expands the product of the two finite sums, (e) interchanges the integral with finite sums by linearity, (f) applies the definition of ⟨ϕr , ϕs ⟩B,h , (g) uses the definition Krs = ⟨ϕr , ϕs ⟩B,h , and (h) is the componentwise expansion of the quadratic form v⊤ Kv. By Lemma D1, (a) 1

D⊤ D. (22) h Here (a) invokes the stiffness-matrix factorization proved in Lemma D1. In particular, the dimensions D ∈ RG×(G+1) and v ∈ RG+1 ensure that Dv ∈ RG and that every product below is well defined. Using (22),   1 ⊤ (a) (b) 1 v⊤ Kv = v⊤ D D v = v⊤ D⊤ Dv h h 1 (c) 1 (d) ⊤ 2 = (Dv) (Dv) = ∥Dv∥2 . (23) h h K =

Here (a) substitutes (22), (b) moves the scalar h1 outside the quadratic form, (c) uses the transpose identity (Dv)⊤ = v⊤ D⊤ and associativity of matrix multiplication, and (d) applies w⊤ w = ∥w∥22 with w = Dv. By the definition of the Euclidean norm and the componentwise definition of D, 2 (a) ∥Dv∥2 =

G−1 X

2 (b)

((Dv)i ) =

i=0

G−1 X

2

(vi+1 − vi ) .

(24)

i=0

Here (a) expands the squared Euclidean norm of the vector Dv ∈ RG , and (b) uses (Dv)i = vi+1 − vi for every i = 0, . . . , G − 1. Combining (21), (23), and (24) yields G−1 X (a) ⊤ (b) 1 2 2 (c) 1 2 ∥f ∥B,h = v Kv = ∥Dv∥2 = (vi+1 − vi ) . h h i=0

(25)

Here (a) is (21), (b) is (23), and (c) substitutes (24). Finally, (22) is exactly the equivalent matrix identity asserted in the proposition. Therefore all claimed representations follow. □ B.2

Proof of Corollary 1

Fix an arbitrary vector v = (v0 , . . . , vG )⊤ ∈ RG+1 and define f := R (v) .

14

(26)

By Definition 1, R maps RG+1 into Hh . Consequently, (b)

(a)

f = R (v) ∈ Hh .

(27)

Here (a) is the definition (26), and (b) follows from R(RG+1 ) = Hh and v ∈ RG+1 . Therefore Proposition 1 applies to f and its nodal vector v. By the definition of the implemented finite Brownian regularizer, Ωh (v) :=

G−1 X

2

(vi+1 − vi ) .

(28)

i=0

On the other hand, Proposition 1, applied using (27), gives G−1 (a) 1 X

2

∥R (v)∥B,h =

h i=0

2

(vi+1 − vi ) .

(29)

Here (a) is the final identity in Proposition 1, applied to the function f = R(v) established in (27). Since h = 2A/G > 0, multiplying (29) by h gives 2

Ωh (v) = h ∥R (v)∥B,h .

(30)

As v ∈ RG+1 was arbitrary and h depends only on the fixed grid, the implemented regularizer is the discrete Brownian energy scaled by the fixed positive factor h. □ B.3

Proof of Theorem 1: RKHS existence and dimension

We first identify the linear structure and dimension of the anchored space. Set I0 := {0, . . . , G} \ {i0 } .

(31)

Since ti0 = 0 and the nodal basis satisfies ϕj (tr ) = δjr , we have, for every j = 0, . . . , G, (a)

(b)

ϕj (0) = ϕj (ti0 ) = δji0 .

(32)

Here (a) uses ti0 = 0, and (b) is the nodal interpolation property of the finite-element basis. Let f ∈ Hh be arbitrary, and let v = (v0 , . . . , vG )⊤ be its unique nodal vector, so that f = R(v). Evaluating this representation at the anchor and using (32) gives (a)

G (b) X

f (0) = R(v)(0) =

G (c) X

vj ϕj (0) =

j=0

(d)

vj δji0 = vi0 .

(33)

j=0

Here (a) uses f = R(v); (b) is the definition of the reconstruction operator; (c) substitutes (32); and (d) uses the defining property of the Kronecker delta, which removes every term except the term with index j = i0 . Consequently, X (a) (b) (c) f ∈ Hh0 ⇐⇒ f (0) = 0 ⇐⇒ vi0 = 0 ⇐⇒ f = v j ϕj . (34) j∈I0

Here (a) is the definition of Hh0 ; (b) uses (33); and (c) uses the unique nodal representation and the fact that the coefficient of ϕi0 is zero exactly when vi0 = 0. It follows that (a)

Hh0 = span {ϕj : j ∈ I0 } .

(35)

Here (a) is precisely the set equality expressed by (34). In particular, Hh0 is a vector subspace of Hh . We next verify that the spanning family in (35) is linearly independent. Suppose that real coefficients (aj )j∈I0 satisfy X aj ϕj = 0. (36) j∈I0

15

Fix an arbitrary r ∈ I0 and evaluate (36) at the grid point tr . Then (c) (b) X (a) X aj δjr = ar . aj ϕj (tr ) = 0 =

(37)

j∈I0

j∈I0

Here (a) evaluates the zero-function identity (36) at tr ; (b) uses ϕj (tr ) = δjr ; and (c) again uses the defining property of the Kronecker delta. Since r ∈ I0 was arbitrary, every coefficient ar vanishes. Therefore {ϕj : j ∈ I0 } is a basis of Hh0 . Because I0 is obtained by removing one index from the G + 1 indices 0, . . . , G,  (a) (b) (c) dim Hh0 = |I0 | = (G + 1) − 1 = G. (38) Here (a) uses the basis just established; (b) uses the definition (31); and (c) simplifies the integer expression. Thus the anchored space is finite-dimensional. We now prove that the Brownian energy form is an inner product on this space. Every function in Hh0 ⊂ Hh is continuous and piecewise affine on a finite partition of I. Hence it is absolutely continuous on I, its derivative exists almost everywhere, and its derivative is piecewise constant and therefore belongs to L2 (I). For arbitrary f, g ∈ Hh0 , the Cauchy–Schwarz inequality gives !1/2 Z !1/2 Z Z A

A

(a)

−A

A

2

|f ′ (t)g ′ (t)| dt ≤

2

|f ′ (t)| dt

|g ′ (t)| dt

−A

(b)

< ∞.

(39)

−A

Here (a) is the Cauchy–Schwarz inequality on L2 (I), and (b) follows from f ′ , g ′ ∈ L2 (I). Thus the integral defining ⟨f, g⟩B,h is finite and the form is well defined. Let f1 , f2 , g ∈ Hh0 and α, β ∈ R. Because Hh0 is a vector space, the function αf1 + βf2 also belongs to Hh0 . Using linearity of the weak derivative and of the integral, Z A (a) ′ ⟨αf1 + βf2 , g⟩B,h = (αf1 + βf2 ) (t)g ′ (t) dt −A

(b)

Z A

(αf1′ (t) + βf2′ (t)) g ′ (t) dt

=

−A (c)

Z A

=α

f1′ (t)g ′ (t) dt + β

Z A

−A

f2′ (t)g ′ (t) dt

−A

(d)

= α ⟨f1 , g⟩B,h + β ⟨f2 , g⟩B,h .

(40) ′

= αf1′ + βf2′ almost everywhere;

Here (a) is the definition of the Brownian energy form; (b) uses (αf1 + βf2 ) (c) distributes the product and uses linearity of the integral; and (d) applies the definition of the Brownian energy form to the two resulting integrals. For arbitrary f, g ∈ Hh0 , commutativity of multiplication in R yields Z A Z A (a) (b) (c) ⟨f, g⟩B,h = f ′ (t)g ′ (t) dt = g ′ (t)f ′ (t) dt = ⟨g, f ⟩B,h . (41) −A

−A

Here (a) and (c) are the definition of the Brownian energy form, and (b) uses f ′ (t)g ′ (t) = g ′ (t)f ′ (t) for almost every t ∈ I. Linearity in the second argument now follows explicitly from (40) and (41): for f, g1 , g2 ∈ Hh0 and α, β ∈ R, (a)

⟨f, αg1 + βg2 ⟩B,h = ⟨αg1 + βg2 , f ⟩B,h (b)

= α ⟨g1 , f ⟩B,h + β ⟨g2 , f ⟩B,h

(c)

= α ⟨f, g1 ⟩B,h + β ⟨f, g2 ⟩B,h .

(42)

Here (a) uses symmetry; (b) applies (40); and (c) uses symmetry separately in each term. Therefore the form is bilinear and symmetric. Moreover, for every f ∈ Hh0 , Z A Z A (c) (a) (b) 2 ⟨f, f ⟩B,h = f ′ (t)f ′ (t) dt = |f ′ (t)| dt ≥ 0. (43) −A

−A

16

Here (a) is the definition of the energy form; (b) uses that the functions are real valued; and (c) follows because |f ′ (t)|2 ≥ 0 almost everywhere. It remains to prove definiteness. Fix f ∈ Hh0 and let v ∈ RG+1 be its unique nodal vector, so that (a)

(b)

f = R(v),

vi0 = 0.

(44)

Here (a) follows from f ∈ Hh and Definition 1, and (b) follows from f ∈ Hh0 and (33). Assume that ⟨f, f ⟩B,h = 0.

(45)

By the definition of the Brownian seminorm and Proposition 1, (a)

(b)

(d) 1

(c)

2

2

0 = ⟨f, f ⟩B,h = ∥f ∥B,h = v⊤ Kv =

∥Dv∥2 . (46) h Here (a) is the assumption (45); (b) is the definition of the Brownian seminorm; and (c)–(d) are the stiffness and first-difference representations in Proposition 1. The mesh size is strictly positive: (a) 2A (b)

h =

> 0. (47) G Here (a) is the mesh-size definition, and (b) follows because A > 0 and the strictly increasing grid contains G ≥ 1 subintervals. Multiplying the first and last quantities in (46) by h and using (47), we obtain   1 (b) (c) 2 2 (a) ∥Dv∥2 = h · 0 = 0. (48) ∥Dv∥2 = h h Here (a) uses h(1/h) = 1; (b) uses (46); and (c) uses h · 0 = 0. For each i = 0, . . . , G − 1, (a)

2

X (b) G−1

0 ≤ ((Dv)i ) ≤

2 (c)

2 (d)

((Dv)r ) = ∥Dv∥2 = 0.

(49)

r=0

Here (a) uses nonnegativity of a square; (b) uses that the sum on the right contains the nonnegative term on the left; (c) is the componentwise definition of the Euclidean norm; and (d) uses (48). Therefore every component (Dv)i is zero and hence (a)

Dv = 0.

(50)

Here (a) follows from (49) for all component indices. Using the factorization in Proposition 1,   1 ⊤ (b) 1 (c) 1 (d) (a) D D v = D⊤ (Dv) = D⊤ 0 = 0. Kv = h h h

(51)

Here (a) uses K = h1 D⊤ D; (b) uses associativity of matrix multiplication; (c) substitutes (50); and (d) uses that every matrix maps the zero vector to the zero vector. Thus v ∈ ker(K). Lemma D2 now yields (a)

(b)

(c)

v ∈ ker(K) = span(1) =⇒ v = c1

for some c ∈ R.

(52)

Here (a) follows from (51); (b) is Lemma D2; and (c) is the definition of the one-dimensional span of 1 = (1, . . . , 1)⊤ ∈ RG+1 . The anchoring constraint in (44) then gives (a)

(b)

(c)

0 = vi0 = (c1)i0 = c.

(53)

Here (a) is the anchoring condition; (b) substitutes v = c1; and (c) uses 1i0 = 1. Therefore c = 0 and v = 0. Substituting this vector into the reconstruction formula gives (a)

G (c) X

(b)

f = R(v) = R(0) =

(d)

0ϕj = 0.

(54)

j=0

Here (a) is (44); (b) uses v = 0; (c) is the definition of the reconstruction operator; and (d) simplifies the zero linear combination. Together with nonnegativity in (43), this proves that the Brownian energy form is positive definite on Hh0 . Hence ⟨·, ·⟩B,h is an inner product on Hh0 , and ∥ · ∥B,h is its induced norm.

17

Completeness and boundedness of evaluation are automatic in finite dimensions, but the two quantitative bounds behind them are used later and are recorded here. Pi0 −1 Let f = R(v) ∈ Hh0 , so vi0 = 0. Telescoping from the anchor, vj = − r=j (vr+1 − vr ) for j < i0 Pj−1 and vj = r=i0 (vr+1 − vr ) for j > i0 ; in either case Cauchy–Schwarz over at most G terms gives vj2 ≤ PG−1 G i=0 (vi+1 − vi )2 = Gh∥f ∥2B,h by Proposition 1. Summing over the G + 1 nodes, p 2 ∥v∥2 ≤ G(G + 1)h ∥f ∥B,h , and conversely ∥f ∥B,h ≤ √ ∥v∥2 , (55) h 2 the second bound from (vi+1 − vi )2 ≤ 2(vi+1 + vi2 ). Thus ∥ · ∥B,h and the Euclidean norm on the coefficient 0 vector are equivalent, so (Hh , ⟨·, ·⟩B,h ) is a finite-dimensional inner-product space and hence complete: a ∥ · ∥B,h -Cauchy sequence has Cauchy coefficient vectors by (55), and the limit vector reconstructs the limit function, which lies in Hh0 because vi0 = 0 is preserved. Rt For t ∈ I the evaluation functional δt (f ) = f (t) is linear, and writing f (t) = f (t) − f (0) = 0 f ′ and applying Cauchy–Schwarz on an interval of length at most A, Z 1/2 Z t p √ ′ ′ 2 |δt (f )| = f ≤ |t| |f | ≤ A ∥f ∥B,h . (56) I

0

√

Hence ∥δt ∥ ≤ A, uniformly in h: the bound does not degrade under mesh refinement, which is the quantitative form of the isometric embedding of Corollary 2. Since Hh0 is a real Hilbert space and δt is a bounded linear functional, the Riesz representation theorem provides a unique function kt ∈ Hh0 such that (a)

for every f ∈ Hh0 .

δt (f ) = ⟨f, kt ⟩B,h

(57)

Here (a) is the Riesz representation theorem applied to the bounded linear functional δt . Define kh0 : I × I −→ R,

kh0 (s, t) := kt (s).

(58)

For every f ∈ Hh0 and t ∈ I, (a)

(b)

(c)

f (t) = δt (f ) = ⟨f, kt ⟩B,h = f, kh0 (·, t) B,h .

(59)

Here (a) is the definition of point evaluation; (b) is (57); and (c) uses kh0 (·, t) = kt from (58). Thus kh0 reproduces every point evaluation. For completeness, the kernel is symmetric because, for arbitrary s, t ∈ I, (a)

(b)

(c)

(d)

(e)

kh0 (s, t) = kt (s) = ⟨kt , ks ⟩B,h = ⟨ks , kt ⟩B,h = ks (t) = kh0 (t, s).

(60)

Here (a) and (e) use the kernel definition; (b) applies (57) with the function f = kt and evaluation point s; (c) uses symmetry of the inner product; and (d) applies (57) with the function f = ks and evaluation point t. Moreover, for arbitrary N ≥ 1, points t1 , . . . , tN ∈ I, and coefficients a1 , . . . , aN ∈ R, *N + N X N N N N X X X (a) X X (b) 0 ar as kh (tr , ts ) = ar as ⟨kts , ktr ⟩B,h = as kts , ar ktr r=1 s=1

r=1 s=1 (c)

=

N X r=1

s=1

r=1

B,h

2 (d)

≥ 0.

ar ktr

(61)

B,h

Here (a) uses kh0 (tr , ts ) = kts (tr ) = ⟨kts , ktr ⟩B,h , which follows from (58) and (57); (b) uses bilinearity of the inner product and the finiteness of the two sums; (c) observes that the two displayed sums represent the same function after renaming the dummy index; and (d) uses nonnegativity of the squared norm. Therefore kh0 is a positive-semidefinite reproducing kernel for Hh0 . Combining this reproducing property with (38) and the completeness proved above shows that   Hh0 , ⟨·, ·⟩B,h (62) is a finite-dimensional reproducing kernel Hilbert space, which completes the proof. Items i–ii are established in Section C.2. □ 18

B.4

Proof of Theorem 2

The increment-energy identity used in part iii is Lemma D4, and the positivity of Λ used in the spectral part is Lemma 3.1 with Lemma D5. The derivation below is self-contained, but those lemmas may be substituted for the corresponding steps. Let f ∈ Hh0 be arbitrary. By Definition 3, G is even. Write G = 2m and i0 = m. Since f ∈ Hh0 ⊆ Hh , Definition 1 yields a unique vector v = (v0 , . . . , vG )⊤ ∈ RG+1 such that G (b) X

(a)

f = R(v) =

vr ϕ r .

(63)

r=0

Here (a) is the unique finite-element representation of f , and (b) is the definition of the reconstruction operator R. For every grid index j = 0, . . . , G, the nodal property of the basis gives G (a) X

f (tj ) =

G (b) X

vr ϕr (tj ) =

r=0

(c)

vr δrj = vj .

(64)

r=0

Here (a) evaluates (63) at tj , (b) uses ϕr (tj ) = δrj , and (c) uses the defining property of the Kronecker delta. Since ti0 = 0 and f (0) = 0, it follows that (a)

(b)

(c)

vi0 = f (ti0 ) = f (0) = 0.

(65)

Here (a) applies (64), (b) uses ti0 = 0, and (c) uses f ∈ Hh0 . G

e ∈ R be the reduced anchored vector ordered as in (17). By the definition of the anchoring reconstruction Let v matrix in (18), (a)

v = Re v.

(66)

Equality (a) inserts the zero component at index i0 and places every remaining reduced coordinate at its corresponding nodal index. Define D0 := DR ∈ RG×G .

(67)

Then the increment coordinates satisfy (a)

(b)

(c)

e. w = Dv = DRe v = D0 v

(68)

Here (a) is Definition 6, (b) substitutes (66), and (c) uses (67). The spectral coordinates satisfy (a)

e. c = Q⊤ v

(69)

Equality (a) is Definition 7. We first establish ii, because the inverse coordinate maps will then make the assertion in i explicit. For arbitrary x, y ∈ RG and α, β ∈ R, (a)

D0 (αx + βy) = DR (αx + βy) (b)

= D (αRx + βRy)

(c)

(d)

= αDRx + βDRy = αD0 x + βD0 y.

(70)

Here (a) substitutes D0 = DR, (b) uses the linearity of multiplication by R, (c) uses the linearity of multiplication by D, and (d) again uses the definition of D0 . Thus x 7→ D0 x is linear. We next prove injectivity. Let w ∈ RG satisfy D0 w = 0.

(71)

Set z := Rw ∈ RG+1 . Then (a)

(b)

(c)

Dz = DRw = D0 w = 0.

19

(72)

Here (a) substitutes z = Rw, (b) uses D0 = DR, and (c) applies (71). By the componentwise definition of D, for every i = 0, . . . , G − 1, (a)

(b)

0 = (Dz)i = zi+1 − zi ,

(73)

where (a) follows from (72), and (b) is the definition of the first-difference operator. Therefore zi+1 = zi for every i = 0, . . . , G − 1. Repeated application of these equalities gives (a)

zj = z0 ,

j = 0, . . . , G.

(74)

Equality (a) follows by induction on j: it is immediate for j = 0, and if zj = z0 , then zj+1 = zj = z0 . The row of R corresponding to the anchor index i0 is zero because none of the columns of R equals the omitted canonical vector ei0 . Hence (a)

(b)

zi0 = (Rw)i0 = 0.

(75)

Here (a) uses z = Rw, and (b) uses the zero anchor row of R. Combining (74) and (75) gives, for every j, (a)

(b)

zj = zi0 = 0,

(76)

where (a) uses the constancy of z, and (b) uses the anchored value. Thus z = 0. The columns of R are pairwise distinct canonical basis vectors of RG+1 . Consequently, (a)

R ⊤ R = IG .

(77)

⊤

Indeed, the (r, s) entry of R R is the Euclidean inner product of columns r and s of R, which is 1 when r = s and 0 otherwise. It follows that (a)

(b)

(c)

(d)

w = IG w = R⊤ Rw = R⊤ z = 0.

(78)

Here (a) uses the identity matrix, (b) applies (77), (c) uses z = Rw, and (d) uses z = 0. Therefore ker(D0 ) = {0}, so the increment map is injective. We now prove surjectivity directly. Let d = (d0 , . . . , dG−1 )⊤ ∈ RG be arbitrary, and define u(d) = (u0 , . . . , uG )⊤ ∈ RG+1 by  i0 −1 X    − dr , 0 ≤ j ≤ i0 ,   r=j uj := j−1 (79)  X    dr , i0 ≤ j ≤ G.  r=i0

The two formulas agree at j = i0 because both sums are empty, and every empty sum is understood to equal zero. In particular, (a)

ui0 = 0.

(80)

Equality (a) follows from the empty-sum convention. For i = 0, . . . , i0 − 1, (a)

ui+1 − ui = −

iX 0 −1

dr +

r=i+1

iX 0 −1

(b)

dr = d i .

(81)

r=i

Here (a) applies the first branch of (79), including its empty-sum value when i = i0 − 1, and (b) cancels the common terms di+1 , . . . , di0 −1 , leaving only di . For i = i0 , . . . , G − 1, i (a) X

ui+1 − ui =

dr −

r=i0

i−1 X

(b)

dr = d i .

(82)

r=i0

Here (a) applies the second branch of (79), including the empty second sum when i = i0 , and (b) cancels the common terms di0 , . . . , di−1 , leaving only di . Hence, for every i = 0, . . . , G − 1, (a)

(b)

(Du(d))i = ui+1 − ui = di , 20

(83)

where (a) is the definition of D, while (b) applies (81) when i < i0 and (82) when i ≥ i0 . Therefore (a)

Du(d) = d.

(84)

Equality (a) follows because the two vectors have equal components at every index. Because the columns of R contain every canonical basis vector except ei0 , one also has (b) (a) X ⊤ ej e⊤ RR⊤ = j = IG+1 − ei0 ei0 .

(85)

j̸=i0

Here (a) expands the product as the sum of the outer products of the columns of R, and (b) removes the single omitted canonical projector from the canonical decomposition of the identity. Define x(d) := R⊤ u(d) ∈ RG .

(86)

Then (a)

Rx(d) = RR⊤ u(d) (b)

 = IG+1 − ei0 e⊤ i0 u(d)

(c)

(d)

= u(d) − ei0 ui0 = u(d).

(87)

Here (a) substitutes (86), (b) uses (85), (c) evaluates e⊤ i0 u(d) = ui0 , and (d) uses (80). Consequently, (a)

(b)

(c)

D0 x(d) = DRx(d) = Du(d) = d.

(88)

Here (a) uses D0 = DR, (b) applies (87), and (c) applies (84). Since d ∈ RG was arbitrary, D0 is surjective. Together with injectivity, this proves that D0 is invertible. Moreover, the preceding construction gives the explicit inverse (a)

⊤ D−1 0 d = R u(d),

(89)

where (a) follows from uniqueness of the preimage under the injective map D0 and (88). We next treat the spectral map. For arbitrary x, y ∈ RG and α, β ∈ R, (a)

Q⊤ (αx + βy) = αQ⊤ x + βQ⊤ y.

(90)

Equality (a) is the distributivity and homogeneity of matrix multiplication, so the spectral map is linear. Since Q is orthogonal, (a)

(b)

Q⊤ Q = I G ,

QQ⊤ = IG .

(91)

Equalities (a) and (b) are the two defining inverse identities for an orthogonal square matrix. Thus, for every x, a ∈ RG ,  (a)  (b) (c) Q Q⊤ x = QQ⊤ x = IG x = x, (92)  (a) (b) (c) Q⊤ (Qa) = Q⊤ Q a = IG a = a. (93) In each line, (a) uses associativity, (b) applies (91), and (c) uses the identity matrix. Therefore x 7→ Q⊤ x is invertible, with inverse a 7→ Qa. The complete coordinate conversion formulas are now (a)

(b)

e, c = Q⊤ v

e, w = D0 v (c)

(d)

e = D−1 v 0 w,

e = Qc, v

(e)

(f )

c = Q⊤ D−1 0 w,

w = D0 Qc.

(94)

Here (a) and (b) are the coordinate definitions, (c) uses the inverse of D0 , (d) uses the inverse of Q⊤ , (e) substitutes (c) into (b), and (f) substitutes (d) into (a). Hence each coordinate system uniquely determines the other two, proving ii. 21

We now prove i. Define the three reconstruction maps by Φnod (x) := R(Rx),

(95)

RD−1 0 d



Φinc (d) := R

,

(96)

Φspec (a) := R (RQa) .

(97)

Each map takes values in Hh0 . Indeed, for every x ∈ RG , G (a) X

Φnod (x)(0) =

G (b) X

(Rx)j ϕj (0) =

j=0 (c)

(Rx)j δji0

j=0 (d)

= (Rx)i0 = 0.

(98)

Here (a) expands R(Rx), (b) uses 0 = ti0 and ϕj (ti0 ) = δji0 , (c) applies the Kronecker delta, and (d) uses the zero anchor row of R. Thus Φnod (x) ∈ Hh0 . The increment and spectral maps are compositions of Φnod with vectors in RG , so they also take values in Hh0 . For the coordinates associated with the fixed function f , (a)

(b)

(c)

Φnod (e v) = R (Re v) = R(v) = f. Here (a) uses (95), (b) uses (66), and (c) uses (63). Likewise,  (b)  (c) (a) (d) (e) −1 e = R (RIG v e) = R (Re Φinc (w) = R RD−1 v) = f. 0 w = R RD0 D0 v

(99)

(100)

e, (c) uses D−1 Here (a) uses (96), (b) substitutes w = D0 v 0 D0 = IG , (d) uses the identity matrix, and (e) applies (99). Finally,  (c) (a) (b) (d) (e) e = R (RIG v e) = R (Re Φspec (c) = R (RQc) = R RQQ⊤ v v) = f. (101) e, (c) uses QQ⊤ = IG , (d) uses the identity matrix, and (e) applies Here (a) uses (97), (b) substitutes c = Q⊤ v (99). Therefore the three coordinate representations reconstruct exactly the same function f , proving i. We next prove iii. By Proposition 1, 2

(a)

(b) 1

∥f ∥B,h = v⊤ Kv =

⊤

(Dv) (Dv) h (c) 1 ⊤ (d) 1 2 = w w = ∥w∥2 . h h

(102)

Here (a) is the stiffness representation of the Brownian energy, (b) uses K = h−1 D⊤ D and associativity, (c) uses w = Dv, and (d) is the definition of the Euclidean norm. For the spectral representation, first 2

(a)

⊤

(b)

(c)

e⊤ R⊤ KRe e ⊤ K0 v e. ∥f ∥B,h = (Re v) K (Re v) = v v=v

(103)

e⊤ R⊤ and associativity, and Here (a) substitutes v = Re v into the first equality of (102), (b) uses (Re v )⊤ = v ⊤ (c) uses K0 = R KR. By Definition 7, (a)

e = Qc, v

(b)

K0 = QΛQ⊤ .

(104)

Here (a) is the inverse spectral transformation proved above, and (b) is the eigendecomposition of the anchored stiffness matrix, whose existence with strictly positive spectrum is established in Lemma 3.1 and Lemma D5, rather than assumed in Definition 7. Consequently,  (a) (b) (c) (d) 2 ⊤ ∥f ∥B,h = (Qc) QΛQ⊤ (Qc) = c⊤ Q⊤ QΛQ⊤ Qc = c⊤ IG ΛIG c = c⊤ Λc. (105) Here (a) substitutes both identities in (104) into (103), (b) uses (Qc)⊤ = c⊤ Q⊤ and associativity, (c) uses Q⊤ Q = IG , and (d) uses the identity matrix. Combining (102) and (105) gives 2

(a)

(b) 1

∥f ∥B,h = v⊤ Kv =

22

h

2 (c)

∥w∥2 = c⊤ Λc.

(106)

Here (a) and (b) are the nodal and increment equalities in (102), while (c) uses (105). This proves iii. We now establish iv. Define the ranges of the three reconstruction maps by Fnod := Φnod (RG ),

Finc := Φinc (RG ),

Fspec := Φspec (RG ).

(107)

Equation (98) shows that Fnod ⊆ Hh0 . Conversely, let g ∈ Hh0 be arbitrary and let a ∈ RG+1 be its unique nodal vector. As in (65), ai0 = 0. Set x := R⊤ a ∈ RG . Then  (a) (b) Rx = RR⊤ a = IG+1 − ei0 e⊤ i0 a (c)

(d)

= a − ei0 ai0 = a.

(108)

Here (a) uses x = R⊤ a, (b) applies (85), (c) evaluates the rank-one projector, and (d) uses ai0 = 0. Therefore (a)

(b)

(c)

g = R(a) = R(Rx) = Φnod (x),

(109)

where (a) is the nodal representation of g, (b) applies (108), and (c) uses (95). Hence g ∈ Fnod , so Hh0 ⊆ Fnod . We have proved (a)

Fnod = Hh0 . Equality (a) combines the two set inclusions. Using the invertibility of D0 ,  (a)  G Finc = Φnod D−1 0 d :d∈R

(110)

(b) 

= Φnod (x) : x ∈ RG

(c)

= Fnod .

(111)

Here (a) uses (96), (b) uses the bijective change of variable x = D−1 0 d, and (c) uses the definition of Fnod . Similarly, using the orthogonality and hence surjectivity of Q, (a)  (b)  (c) Fspec = Φnod (Qa) : a ∈ RG = Φnod (x) : x ∈ RG = Fnod . (112) Here (a) uses (97), (b) uses the bijective change of variable x = Qa, and (c) again uses the definition of Fnod . Combining (110)–(112) gives (a)

(b)

(c)

Fnod = Finc = Fspec = Hh0 .

(113)

Here (a) is (111), (b) is (112), and (c) is (110). For completeness, we also verify that the coordinate realizations pull back the same Brownian inner product, not merely the same norm. Let x, y ∈ RG , define g := Φnod (x),

h := Φnod (y),

α := D0 x,

β := D0 y,

⊤

q := Q⊤ y.

p := Q x,

(114)

The bilinear stiffness identity gives (a)

⊤

(b)

⟨g, h⟩B,h = (Rx) K (Ry) = x⊤ K0 y.

(115)

Here (a) follows by expanding both reconstructed functions in the hat basis and using Krs = ⟨ϕr , ϕs ⟩B,h , while (b) uses K0 = R⊤ KR. Moreover, (a)

(b) 1

x⊤ K0 y = x⊤ R⊤ KRy =

h

(c) 1

x⊤ R⊤ D⊤ DRy =

h

⊤

(d) 1

(D0 x) (D0 y) =

h

α⊤ β.

(116)

Here (a) uses the definition of K0 , (b) substitutes K = h−1 D⊤ D, (c) uses D0 = DR and the transpose identity (D0 x)⊤ = x⊤ D⊤ 0 , and (d) uses the definitions of α and β. Since x = Qp and y = Qq,  (a) ⊤ x⊤ K0 y = (Qp) QΛQ⊤ (Qq) (b)

(c)

= p⊤ Q⊤ QΛQ⊤ Qq = p⊤ Λq.

23

(117)

Here (a) substitutes the inverse spectral transformations and the eigendecomposition of K0 , (b) uses the transpose identity and associativity, and (c) applies Q⊤ Q = IG . Thus (b) 1

(a)

(c)

α⊤ β = p⊤ Λq. (118) h Here (a), (b), and (c) apply (115), (116), and (117), respectively. Therefore, when the nodal, increment, and spectral parameter spaces are respectively equipped with the pulled-back inner products displayed in (118), their reconstruction maps are linear isometric isomorphisms onto the common space Hh0 equipped with its Brownian inner product. By Theorem 1, this common inner-product space is a finite-dimensional RKHS. Hence all three realizations generate the same finite Brownian RKHS and the same hypothesis class, proving iv. e1 , v e2 ∈ RG be arbitrary, and define It remains to establish v. Let v ⟨g, h⟩B,h = x⊤ K0 y =

e1 − v e2 , w := v er , cr := Q⊤ v

er , wr := D0 v

r ∈ {1, 2}.

(119)

For the spectral coordinates, 2 (a)

2 (b)

e 2 ) 2 = Q⊤ w v1 − v ∥c1 − c2 ∥2 = Q⊤ (e (c)

(d)

(e)

⊤

Q⊤ w



2

= w⊤ QQ⊤ w = w⊤ IG w = ∥w∥2 .

(120)

Here (a) uses linearity of Q⊤ and the definitions of c1 and c2 , (b) is the definition of the squared Euclidean norm, (c) uses (Q⊤ w)⊤ = w⊤ Q, (d) uses QQ⊤ = IG , and (e) again uses the definition of the Euclidean norm. Thus the nodal-to-spectral map is a Euclidean isometry. For the increment map, (a)

⊤

(b)

(c)

(d)

(e)

⊤ ⊤ ⊤ ⊤ D⊤ 0 D0 = (DR) (DR) = R D DR = R (hK) R = hR KR = hK0 . ⊤

⊤

(121)

⊤

Here (a) substitutes D0 = DR, (b) uses (DR) = R D and associativity, (c) rearranges the stiffness identity K = h−1 D⊤ D as D⊤ D = hK, (d) extracts the scalar h, and (e) uses K0 = R⊤ KR. Consequently, 2 (a)

2 (b)

(c)

⊤

(d)

⊤ e2 )∥2 = (D0 w) (D0 w) = w⊤ D⊤ ∥w1 − w2 ∥2 = ∥D0 (e v1 − v 0 D0 w = hw K0 w.

(122)

Here (a) uses linearity of D0 and the definitions of w1 and w2 , (b) is the definition of the squared Euclidean norm, (c) uses the transpose identity and associativity, and (d) applies (121). Finally, Lemma D3 gives   (a) Tm 0 hK0 = . (123) 0 Tm Equality (a) is exactly the block decomposition after multiplication by h. If m ≥ 2, then the entrywise definition of Tm gives (a)

(b)

(c)

(Tm )0,1 = −1 ̸= 0 = (Im )0,1 .

(124)

Here (a) uses |0 − 1| = 1, (b) is the elementary inequality −1 ̸= 0, and (c) uses the zero off-diagonal entries of the identity matrix. Therefore Tm ̸= Im and, by (123), (a)

D⊤ 0 D0 ̸= IG ,

m ≥ 2.

(125)

Inequality (a) follows from (121), (123), and Tm = ̸ Im . Hence the increment map is nonorthogonal for G = 2m ≥ 4. In contrast, (120) shows that the spectral map is an orthogonal change of coordinates of the reduced nodal coordinates. This proves v and completes the proof. □

24

B.5

Proof of Theorem 3

By Lemma D3, (a) 1

K0 =



 0 , Tm

(126)

r = s = 0, 1 ≤ r = s ≤ m − 1, |r − s| = 1, otherwise.

(127)

h

Tm 0

where Tm = (τrs )m−1 r,s=0 is defined by  1,    2, τrs :=  −1,    0,

Equality (a) is exactly the block decomposition proved in Lemma D3. In particular, T1 = (1). It is therefore sufficient first to construct an orthonormal eigendecomposition of Tm and then to apply it independently to the two diagonal blocks in (126). For every k ∈ [m], set αm := √

2 , 2m + 1

θk :=

(2k − 1) π , 2m + 1

νk := 2 − 2 cos (θk ) ,

(128)

and define the extended cosine sequence  qbk (j) := αm cos

j+

1 2



 θk ,

j ∈ {−1, 0, . . . , m} .

(129)

By the definition of qk in the theorem statement, (a)

j = 0, . . . , m − 1.

qk (j) = qbk (j),

(130)

Equality (a) follows because (129) and the theorem statement use the same normalization factor and the same cosine formula at these indices. To identify the transform type explicitly, set r := k − 1 ∈ {0, . . . , m − 1}. Then     1 r + 21 1 (a) 2π j + 2 θk = , j, r ∈ {0, . . . , m − 1}. (131) j+ 2 2m + 1  Equality (a) uses r = k − 1, and hence 2k − 1 = 2 r + 21 . Therefore the entries of Qm are the type-VIII 2 discrete-cosine kernel with the normalization factor √2m+1 ; the orthonormality of this normalization is proved below rather than assumed. We first verify the two boundary relations encoded by the extended sequence. At the outer boundary,     θk (c) θk (b) (a) = αm cos = qbk (0). (132) qbk (−1) = αm cos − 2 2 Here (a) substitutes j = −1 into (129), (b) uses the evenness of the cosine, and (c) substitutes j = 0 into the same definition. At the anchored boundary,   1 (a) 2m + 1 (2k − 1) π (b) (2k − 1) π m+ θk = = , (133) 2 2 2m + 1 2 where (a) uses m + 12 = 2m+1 and the definition of θk , while (b) cancels the nonzero factor 2m + 1. 2 Consequently,   (2k − 1) π (b) (a) qbk (m) = αm cos = 0. (134) 2 Here (a) combines (129) and (133), and (b) uses that the cosine of every odd multiple of π2 is zero. We now prove the eigenvalue relation Tm qk = νk qk . The entrywise definition (127), together with (132) and (134), gives, for every j ∈ {0, . . . , m − 1}, (a)

(Tm qk )j = −b qk (j − 1) + 2b qk (j) − qbk (j + 1). 25

(135)

To justify (a) without suppressing the endpoint cases, observe the following. If m = 1, then j = 0, T1 = (1), qbk (−1) = qk (0), and qbk (1) = 0, so the right-hand side of (135) equals qk (0). If m ≥ 2 and j = 0, the first row of Tm gives qk (0) − qk (1), which equals the displayed second difference because qbk (−1) = qk (0). If m ≥ 3 and 1 ≤ j ≤ m − 2, the equality is the ordinary tridiagonal row formula. Finally, if m ≥ 2 and j = m − 1, the last row gives −qk (m − 2) + 2qk (m − 1), which equals the displayed second difference because qbk (m) = 0. These cases exhaust all admissible pairs (m, j). For fixed k ∈ [m] and j ∈ {0, . . . , m − 1}, define   1 aj,k := j + θk . (136) 2 Then (a)

(b)

qbk (j − 1) = αm cos (aj,k − θk ) ,

qbk (j) = αm cos (aj,k ) ,

(c)

qbk (j + 1) = αm cos (aj,k + θk ) .

(137)

Equalities (a)–(c) follow by substituting j − 1, j, and j + 1 into (129) and using (136). Therefore, (a)

(Tm qk )j = αm [− cos (aj,k − θk ) + 2 cos (aj,k ) − cos (aj,k + θk )] (b)

(c)

= αm [2 cos (aj,k ) − 2 cos (aj,k ) cos (θk )] = (2 − 2 cos (θk )) αm cos (aj,k )

(d)

(e)

= νk qbk (j) = νk qk (j).

(138)

Here (a) combines (135) and (137); (b) uses (a)

cos (aj,k − θk ) + cos (aj,k + θk ) = 2 cos (aj,k ) cos (θk ) ,

(139)

which is the elementary cosine addition identity; (c) factors the common quantity αm cos (aj,k ); (d) uses the definitions of νk , qbk (j), and aj,k ; and (e) applies (130). Since (138) holds at every component, (a)

k ∈ [m].

Tm qk = νk qk ,

(140) m

Equality (a) follows from equality of the corresponding components in R . We next prove that the eigenvalues ν1 , . . . , νm are positive and pairwise distinct. For every k ∈ [m], (a)

(b)

(c)

1 ≤ 2k − 1 ≤ 2m − 1 < 2m + 1. Here (a) uses k ≥ 1 and (b) uses k ≤ m. Since 2m + 1 > 0, multiplication by inequalities and gives (a)

(141) π 2m+1

> 0 preserves the

(b)

0 < θk < π.

(142)

Here (a) follows from the first inequality in (141), and (b) follows from its last strict inequality. Moreover, if 1 ≤ k < ℓ ≤ m, then (a)

2k − 1 < 2ℓ − 1

(b)

=⇒

θk < θℓ .

(143)

Inequality (a) follows from k < ℓ, and the implication to inequality (b) follows after multiplication by the π d positive factor 2m+1 . Since dϑ (2 − 2 cos ϑ) = 2 sin ϑ > 0 on (0, π), the map ϑ 7→ 2 − 2 cos ϑ is strictly increasing there. Combining (142), (143) with this monotonicity yields (a)

νk > 0, (b)

νk < νℓ ,

k ∈ [m], 1 ≤ k < ℓ ≤ m.

(144)

Here (a) uses θk > 0 and 2 − 2 cos (0) = 0, while (b) uses strict monotonicity and the strict ordering of the angles. In particular, the values ν1 , . . . , νm are positive and pairwise distinct.

26

We now prove pairwise orthogonality directly, without invoking the spectral theorem. The entrywise formula (127) is invariant under interchange of r and s, because the conditions r = s, r = s = 0, and |r − s| = 1 are all symmetric. Thus (a)

τrs = τsr

=⇒

(b)

T⊤ m = Tm .

(145)

Here (a) follows from the cases in (127), and (b) is the entrywise definition of a symmetric matrix. Fix k, ℓ ∈ [m] with k ̸= ℓ. Then (a)

⊤

(b)

⊤

νk qk⊤ qℓ = (νk qk ) qℓ = (Tm qk ) qℓ (c)

(d)

(e)

(f )

⊤ = qk⊤ T⊤ m qℓ = qk Tm qℓ

= qk⊤ (νℓ qℓ ) = νℓ qk⊤ qℓ .

(146)

Here (a) moves the real scalar νk inside the transpose; (b) applies the eigenvalue relation (140) for k; (c) ⊤ uses (Ax) = x⊤ A⊤ ; (d) uses (145); (e) applies (140) for ℓ; and (f) extracts the scalar νℓ . Subtracting the right-hand side from the left-hand side gives (a)

(νk − νℓ ) qk⊤ qℓ = 0.

(147)

Equality (a) is a rearrangement of (146). Since νk ̸= νℓ by (144), division by the nonzero scalar νk − νℓ yields (a)

qk⊤ qℓ = 0,

k ̸= ℓ.

(148)

Equality (a) uses the elementary fact that a product of a nonzero real number and a real number can vanish only if the second factor vanishes. It remains to compute the norm of each qk . By definition,    m−1 m−1 X X 4 1 2 (a) 2 (b) 2 ∥qk ∥2 = qk (j) = cos j+ θk . (149) 2m + 1 j=0 2 j=0 Here (a) is the definition of the squared Euclidean norm, and (b) substitutes the component formula and 4 2 uses αm = 2m+1 . The identity cos2 (x) = 1+cos(2x) gives 2    m−1 m−1 X 1 (a) 1 X cos2 j+ [1 + cos ((2j + 1) θk )] θk = 2 2 j=0 j=0 (b) m

=

2

m−1

+

1 X cos ((2j + 1) θk ) . 2 j=0

Here (a) applies the double-angle identity term by term, and (b) uses finite sums. Define Sk :=

m−1 X

Pm−1 j=0

(150)

1 = m and separates the two

cos ((2j + 1) θk ) .

(151)

j=0

For every j ∈ {0, . . . , m − 1}, the sine addition formula gives (a)

2 sin (θk ) cos ((2j + 1) θk ) = sin ((2j + 2) θk ) − sin (2jθk ) .

(152)

Indeed, (a) is the identity 2 sin (x) cos (y) = sin (x + y) + sin (x − y) with x = θk and y = (2j + 1) θk , together with the oddness of the sine. Summing (152) from j = 0 to j = m − 1 yields m−1 (a) X

2 sin (θk ) Sk =

[sin ((2j + 2) θk ) − sin (2jθk )]

j=0 (b)

(c)

= sin (2mθk ) − sin (0) = sin (2mθk ) . 27

(153)

Here (a) uses the definition of Sk and linearity of finite summation; (b) uses telescopic cancellation, since every intermediate term sin (2rθk ) for r = 1, . . . , m − 1 appears once with sign +1 and once with sign −1; and (c) uses sin (0) = 0. Furthermore, (a)

(b)

2mθk = (2m + 1) θk − θk = (2k − 1) π − θk ,

(154)

where (a) adds and subtracts θk , and (b) uses (2m + 1) θk = (2k − 1) π. Since 2k − 1 is odd, (a)

(b)

sin (2mθk ) = sin ((2k − 1) π − θk ) = sin (θk ) .

(155)

Here (a) applies (154), and (b) uses sin ((2r − 1) π − x) = sin (x) for every integer r. Combining (153) and (155) gives (a)

2 sin (θk ) Sk = sin (θk ) .

(156)

Equality (a) substitutes (155) into (153). By (142), sin (θk ) > 0, and hence it is nonzero. Division by 2 sin (θk ) therefore yields (a) 1

. (157) 2 Equality (a) performs the valid division by the nonzero quantity 2 sin (θk ). Substituting (157) into (150) gives      m−1 X 1 1 1 (b) 2m + 1 (a) m 2 cos j+ θk = + = . (158) 2 2 2 2 4 j=0 Sk =

Here (a) uses (157), and (b) puts the two terms over the common denominator 4. Finally, 2 (a)

∥qk ∥2 =

4 2m + 1 (b) = 1. 2m + 1 4

(159)

Here (a) combines (149) and (158), and (b) cancels the positive factor 2m + 1 and the nonzero factor 4. Equations (148) and (159) imply (a)

qk⊤ qℓ = δkℓ ,

k, ℓ ∈ [m].

(160)

Here (a) uses the zero cross-products when k ̸= ℓ and the unit norm when k = ℓ. To verify explicitly that the family is a basis, suppose that a1 , . . . , am ∈ R satisfy m X

(a)

ak qk = 0.

(161)

k=1

For every fixed ℓ ∈ [m], left multiplication by qℓ⊤ gives m m X (c) X ⊤ (b) ⊤ 0 = qℓ 0 = qℓ ak qk = ak qℓ⊤ qk k=1 k=1 m (e) (d) X = ak δℓk = aℓ . (a)

(162)

k=1

Here (a) uses qℓ⊤ 0 = 0; (b) applies (161); (c) uses linearity of the inner product in a finite sum; (d) uses (160); and (e) uses the defining property of the Kronecker delta. Thus aℓ = 0 for every ℓ ∈ [m], so the family is linearly independent. Since it contains exactly m vectors in the m-dimensional space Rm , it is a basis. This proves i. We next prove ii. Define Qm := [q1 , . . . , qm ] ,

Θ := diag (ν1 , . . . , νm ) .

For k, ℓ ∈ [m], the (k, ℓ) entry of Q⊤ m Qm satisfies  (a) ⊤ (b) (c) Q⊤ m Qm kℓ = qk qℓ = δkℓ = (Im )kℓ .

28

(163)

(164)

Here (a) is the rule for multiplying a matrix by its transpose in terms of its columns; (b) applies (160); and (c) is the entrywise definition of the identity matrix. Equality of all entries gives (a)

Q⊤ m Qm = Im .

(165)

Equality (a) follows because the two m × m matrices have identical entries. Since the columns of Qm form a basis, Qm is invertible. Multiplying (165) from the right by Q−1 m gives (a)

−1 Q⊤ m = Qm ,

(166)

−1 −1 where (a) uses Q⊤ m Qm Qm = Im Qm . Therefore, (a)

(b)

−1 Qm Q⊤ m = Q m Qm = I m .

(167)

−1 Here (a) substitutes Q⊤ m = Qm , and (b) uses the defining property of an inverse. Collecting the eigenvalue identities (140) columnwise gives (a)

(b)

(c)

(d)

Tm Qm = Tm [q1 , . . . , qm ] = [Tm q1 , . . . , Tm qm ] = [ν1 q1 , . . . , νm qm ] = Qm Θ.

(168)

Here (a) substitutes the definition of Qm ; (b) applies matrix multiplication column by column; (c) uses (140); and (d) uses the definition of the diagonal matrix Θ. Consequently, (a)

(b)

(c)

⊤ Tm = Tm Im = Tm Qm Q⊤ m = Qm ΘQm (d)

= Qm diag (ν1 , . . . , νm ) Q⊤ m.

(169)

Here (a) inserts the identity on the right; (b) applies (167); (c) uses (168); and (d) substitutes the definition of Θ. This proves ii. We now prove iii. Because A > 0, G = 2m > 0, and h = 2A G , (a) 2A (b) A (c)

h =

= > 0. (170) 2m m Here (a) substitutes G = 2m, (b) cancels the nonzero factor 2, and (c) uses A > 0 and m ≥ 1. Define     1 Θ 0 Qm 0 Q := , Λ := . (171) 0 Qm h 0 Θ Using (165), (a)



(b)



Q⊤ Q = =

Q⊤ m 0

  Qm 0 0 Qm   (c) Im 0 = 0 Q⊤ m Qm

0 Q⊤ m

Q⊤ m Qm 0

0 Im



(d)

(e)

= I2m = IG .

(172)

Here (a) transposes the block-diagonal matrix in (171); (b) performs the block multiplication; (c) uses (165); (d) identifies the resulting block matrix as the identity in dimension 2m; and (e) uses G = 2m. Similarly, using (167),   ⊤ (b) 0 ⊤ (a) Qm Qm QQ = = IG . (173) 0 Qm Q ⊤ m Here (a) performs the analogous block multiplication, and (b) applies (167) and G = 2m. Thus Q is orthogonal. Substituting (169) into (126) gives   (a) 1 Qm ΘQ⊤ 0 m K0 = 0 Qm ΘQ⊤ h m      ⊤  1 Θ 0 (b) Qm 0 Qm 0 = 0 Qm 0 Q⊤ h 0 Θ m (c)

= QΛQ⊤ .

(174) 29

Here (a) substitutes the eigendecomposition of both copies of Tm ; (b) factors the block-diagonal product and places the scalar factor h1 in the middle block; and (c) applies (171). This proves iii. It remains to prove iv. The elementary half-angle identity gives   θk (a) (b) νk = 2 − 2 cos (θk ) = 4 sin2 , (175) 2  where (a) is the definition of νk , and (b) uses 1 − cos (x) = 2 sin2 x2 . Thus the distinct eigenvalue values appearing on the diagonal of Λ are   θk (a) νk (b) 4 = sin2 λk = h h 2   (2k − 1) π (c) 4 = sin2 , k = 1, . . . , m. (176) h 2 (2m + 1) Here (a) reads the diagonal scaling from (171); (b) applies (175); and (c) substitutes the definition of θk . Because h > 0 by (170) and the νk are strictly ordered by (144), (a)

λk > 0, (b)

λk < λ ℓ ,

k ∈ [m], 1 ≤ k < ℓ ≤ m.

(177)

Here (a) and (b) follow by division of the corresponding inequalities in (144) by the positive scalar h. By (171), (a)

Λ = diag (λ1 , . . . , λm , λ1 , . . . , λm ) .

(178)

Equality (a) uses λk = νhk in each of the two diagonal blocks. Since (174) and (172) imply Q−1 = Q⊤ , the matrices K0 and Λ are similar. Hence, for an indeterminate z,  (b)  (a) det (zIG − K0 ) = det zIG − QΛQ⊤ = det Q (zIG − Λ) Q⊤  (d)  (c) = det (Q) det (zIG − Λ) det Q⊤ = det QQ⊤ det (zIG − Λ) m (e) (f ) Y 2 = det (zIG − Λ) = (z − λk ) . (179) k=1 ⊤ Here (a) applies (174); (b) uses zIG = Q follows from (173); (c) uses multiplicativity of the  (zIG ) Q , which  determinant; (d) uses det (Q) det Q⊤ = det QQ⊤ ; (e) applies (173) and det (IG ) = 1; and (f) uses the repeated diagonal form (178). By (177), the numbers λ1 , . . . , λm are distinct. Therefore each factor z − λk occurs exactly twice in the characteristic polynomial (179), and each λk has algebraic multiplicity two. This proves iv and completes the proof. □

Table 4: Summary of the auxiliary theoretical results. Their logical dependencies are illustrated in Fig. 2. For the main theoretical results, see Table 3.

Result

Content

Page

Lemma D1 Lemma D2 Lemma D3 Lemma D4 Lemma D5 Lemma D6

Structure of the stiffness matrix Kernel of the stiffness matrix Block decomposition of the anchored stiffness matrix Brownian energy in increment coordinates Brownian energy in spectral coordinates Explicit spectral Brownian basis

page 34 page 36 page 38 page 41 page 41 page 43

30

C

Proofs of the Kernel, Conditioning, and Optimization Results

This section completes the explicit kernel, approximation, conditioning, and optimization consequences of the foundational results proved above. We first establish the positive definiteness of the anchored stiffness matrix in Lemma 3.1, identify the reproducing kernel and Brownian Gram inverse in Theorem 1, items i–ii, and derive the conforming-subspace and sharp-approximation results in Corollaries 2 and 3. We then obtain the exact stiffness and coordinate condition numbers in Corollary 4. Finally, we prove the linear-reparameterization identity in Proposition 2, the nodal–spectral and increment–Brownian gradient correspondences in Proposition 3, their recursive blockwise extension in Corollary 5, the mesh-independent increment least-squares bound in Proposition 4, and the signed-permutation characterization of standard Adam equivariance in Proposition 5. e Throughout, J = {0, . . . , G} \ {i0 } indexes the reduced anchored coordinates of Definition 4, and ϕ(s) = 0 e, u e, (ϕj (s))j∈J is understood in that same reduced ordering. Thus, for f, g ∈ Hh with reduced coordinates v e ⊤v e, f (s) = ϕ(s) C.1

e ⊤ K0 u e. ⟨f, g⟩B,h = v

(180)

Proof of Lemma 3.1

By Lemma D2, ker K = span(1), and by Proposition 1, K ⪰ 0. If x⊤ K0 x = 0, then (Rx)⊤ K(Rx) = 0, so Rx = c1. The anchor entry of Rx is zero; hence c = 0. Since R⊤ R = IG , x = 0. Thus K0 ≻ 0, and (180) is an inner product. □ C.2

Proof of Theorem 1, items (i)–(ii)

Finite-dimensional RKHS existence and dimension are proved in Appendix B.3. Fix t ∈ I. If the representer et , then for every v e, kh (·, t) has reduced vector k e ⊤v et = f (t) = ϕ(t) e⊤ K0 k e. v

(181)

e e et = ϕ(t), et = K−1 ϕ(t), Therefore K0 k and positive definiteness gives k proving item (i). 0 For r ∈ J , set gr = kB (·, tr ). Its only possible kinks are at 0 and tr , both grid nodes, so gr ∈ Hh0 . If tr > 0, then gr′ = ⊮(0,tr ) almost everywhere; if tr < 0, then gr′ = −⊮(tr ,0) . Hence, for every f ∈ Hh0 , (R t r f ′ (t) dt, tr > 0, 0R ⟨f, gr ⟩B,h = = f (tr ) − f (0) = f (tr ). (182) 0 ′ − tr f (t) dt, tr < 0, e r ) = eπ(r) , Thus gr = kh (·, tr ). Since ϕ(t (K−1 0 )π(r),π(q) = kh (tr , tq ) = kB (tr , tq ).

(183)

Substitution into item (i) gives kh (s, t) =

X

ϕr (s)kB (tr , tq )ϕq (t),

(184)

r,q∈J

which is the tensor-product bilinear interpolant. If one argument is a node, one sum collapses and the preceding representer identity gives exact agreement. □ C.3

Proof of Corollary 2

 The functions gr = kB (·, tr ) belong to Hh0 , and their Gram matrix is kB (tr , tq ) r,q = K−1 0 in reduced order. 0 They are therefore linearly independent, and their number equals dim Hh = G, so they span the space. Every g ∈R Hh0 is continuous, piecewise affine, anchored, and has derivative in L2 (I); hence g ∈ HB and ∥g∥2HB = I (g ′ )2 = ∥g∥2B,h . Let Ih f be nodal interpolation of f ∈ HB . For every r ∈ J , ⟨f − Ih f, gr ⟩HB = (f − Ih f )(tr ) = 0. Since the gr span Hh0 , the residual is orthogonal to that space, so Ih = Πh .

31

(185) □

C.4

Proof of Corollary 3

Fix a cell Ti = [a, b] and write h = b − a. No cell crosses the anchor because 0 is a grid node. Suppose first 0 ≤ a ≤ s ≤ t ≤ b, and set α = (s − a)/h, β = (t − a)/h. The four Brownian corner values are a, a, a, b, so bilinear interpolation gives kh (s, t) = (1 − α)(1 − β)a + (1 − α)βa + α(1 − β)a + αβb = a + αβh.

(186)

Since kB (s, t) = s = a + αh, (s − a)(b − t) . (187) h For a ≤ s ≤ t ≤ b ≤ 0, the corner values are −a, −b, −b, −b and the same bilinear calculation gives kh (s, t) = −b + (1 − α)(1 − β)h. Since kB (s, t) = −t, the same residual (s − a)(b − t)/h follows. Symmetry replaces s, t by their minimum and maximum. On two different cells the diagonal kink s = t is absent from the corresponding rectangle and kB is affine in each argument there; hence its bilinear interpolant is exact. This proves (4). Because Πh is orthogonal, the error functional f 7→ f (t) − Πh f (t) has representer rh (·, t) and squared norm rh (t, t). Cauchy–Schwarz gives (5); equality holds for a normalized residual section whenever t is not a node. √ Maximizing (t − a)(b − t)/h on a cell gives h/4, proving the sharp h/2 constant. The Pythagorean identity follows from f − Πh f ⊥ Πh f . For the L2 bound, set e = f − Πh f . On every cell, e(a) = e(b) = 0, so the sharp Wirtinger inequality gives kB (s, t) − kh (s, t) = α(1 − β)h =

∥e∥2L2 (Ti ) ≤

h2 ′ 2 ∥e ∥L2 (Ti ) . π2

(188)

The derivative of Πh f on Ti is the cell average of f ′ , so e′ is the residual after the L2 (Ti )-orthogonal projection onto constants. Summing over cells, h2 ′ 2 h2 ′ 2 ∥e ∥ ∥f ∥L2 (I) . (189) 2 (I) ≤ L π2 π2 Sharpness follows by taking a sine arch supported on one cell, f (t) = sin(π(t − a)/h) on that cell and zero elsewhere; it has zero nodal values and attains the Wirtinger constant. □ ∥e∥2L2 (I) ≤

C.5

Proof of Corollary 4

By Theorem 3, the distinct stiffness eigenvalues are   (2k − 1)π 4 λk = sin2 , h 2(2m + 1)

k ∈ [m].

(190)

They increase with k. Thus   π 4 2 , λmin = sin h 4m + 2   4 π λmax = cos2 , h 2m + 1

(191) (192)

which gives (10). Since D⊤ 0 D0 = hK0 , its squared singular values are hλk . Taking square roots yields (11) and (12). Taylor expansion of sine and cosine gives the stated asymptotics. In particular hλmax → 4; writing λmax → 4/h would be meaningless because h varies with m. □ C.6

Proof of Proposition 2

The chain rule gives ∇LA (z) = A−⊤ ∇L(A−1 z). Applying A−1 to the update gives ek+1 = A−1 zk − ηk A−1 A−⊤ ∇L(e v vk )

(193)

⊤

(194)

ek − ηk (A A) =v

−1

∇L(e vk ).

The stochastic statement is identical with a shared stochastic gradient and its chain-rule transform.

32

□

C.7

Proof of Proposition 3

For A = Q⊤ , orthogonality gives A⊤ A = IG , so (13) is exactly the nodal update. Mapped initializations and induction give identical function iterates. For A = D0 , A⊤ A = hK0 , and (13) becomes (14). The Riemannian gradient for the constant metric with Gram matrix K0 is K−1 0 ∇L; hence the discrete update is explicit Euler for that flow with time increment ηk /h. □ C.8

Proof of Corollary 5

Apply Proposition 2 to the block-diagonal matrix A. Its Gram matrix and inverse are block diagonal:  −1 −1 (A⊤ A)−1 = blkdiag (A⊤ , . . . , (A⊤ ,I . (195) 1 A1 ) B AB ) If every Ab is orthogonal, this is the identity. If every Ab = D0,b , the bth block is (hb K0,b )−1 . The proof never separates the objective by blocks, so arbitrary cross-layer coupling is allowed. □ C.9

Proof of Proposition 4

Let n

B :=

1Xe e i )⊤ ⪰ 0. ϕ(xi )ϕ(x n i=1

(196)

The nodal Hessian is Hnod = B + ρK0 ; the spectral Hessian is Q⊤√Hnod Q, so the two have identical spectra. 1/2 Since D⊤ h UK0 for an orthogonal U. Therefore 0 D0 = hK0 , the polar decomposition has the form D0 = −1 Hinc = D−⊤ 0 (B + ρK0 )D0   1 −1/2 −1/2 = U K0 BK0 + ρIG U⊤ . h −1/2

Set C = K0

−1/2

BK0

(197) (198)

⪰ 0. Then n

1Xe e λmax (C) ≤ tr(C) = ϕ(xi )⊤ K−1 0 ϕ(xi ) n i=1 =

(199)

n n n 1X 1X 1X kh (xi , xi ) ≤ kB (xi , xi ) = |xi | ≤ A. n i=1 n i=1 n i=1

(200)

The penultimate inequality follows from the nonnegative residual diagonal in (5). Hence every eigenvalue of C + ρIG lies in [ρ, ρ + A], and (198) proves (16). If B = 0, the increment Hessian is (ρ/h)IG , while nodal and spectral Hessians have condition number κ2 (K0 ). □ C.10

Proof of Proposition 5

Write standard Adam, for gradient gk , as mk = β1 mk−1 + (1 − β1 )gk ,

(201)

sk = β2 sk−1 + (1 − β2 )(gk ⊙ gk ), mk bk = m , 1 − β1k+1 bk m xk+1 = xk − ηk √ , b sk + ε1

(202) sk b sk = , 1 − β2k+1

(203) (204)

where the last two operations are coordinatewise. If S is a signed permutation, gradients and first moments transform by S⊤ , while squared gradients and second moments are permuted by |S|⊤ . Induction through the recursion shows that the mapped update is exactly the original one, so signed permutations are equivariances. Conversely, let U be orthogonal and suppose Adam is equivariant for every gradient sequence. At the first step, bias correction reduces the Adam direction to Fε (g) = g/(|g| + ε1). Equivariance under z = U⊤ x requires UFε (U⊤ g) = Fε (g) 33

for every g.

(205)

Take g = tUej with t > 0 and let t → ∞. The left side tends to Uej , while the right side tends coordinatewise to sgn(Uej ). Thus every nonzero coordinate of every column Uej equals ±1. Since each column has unit norm, it has exactly one nonzero entry; orthogonality then makes U a signed permutation. For m ≥ 2, the first column of each DCT-VIII block has m strictly positive entries, so Q is not a signed permutation. Failure of (205) for some vector gives a linear objective with that constant gradient; because the first step size is positive, the first mapped Adam iterates differ. □

D

Internal Lemmas

This section proves the six auxiliary linear-algebra and coordinate results used in Sections B and C: the stiffness factorization and nullspace, the anchored block decomposition, the increment and spectral Brownian-energy identities, and the explicit spectral Brownian basis. Their logical roles are displayed in Figure 2, and their statements and locations are summarized in Table 4. Lemma D1 (Structure of the stiffness matrix). The finite-element stiffness matrix K = (Krs )G r,s=0 , with Z A Krs = ϕ′r (t)ϕ′s (t) dt, r, s = 0, . . . , G, (206) −A

satisfies K=

1 ⊤ D D. h

(207)

Equivalently, 1

−1

 −1 1  K=  0 h  .  ..

2 −1 .. .

0

···

··· . −1 . . .. . 2 .. .. . . 0 −1 0

 0 ..  .    . 0   −1 1

(208)

Consequently, for every v ∈ RG+1 and f = R(v) ∈ Hh , G−1 1 1 X 2 2 ∥Dv∥2 = (vi+1 − vi ) . h h i=0

2

∥f ∥B,h =

(209)

Proof. Let v = (v0 , . . . , vG )⊤ ∈ RG+1 be arbitrary, and set f = R(v). Fix an interval index i ∈ {0, . . . , G − 1}. For every t ∈ [ti , ti+1 ], only the two basis functions ϕi and ϕi+1 can be nonzero. Moreover, on this interval, ϕi (t) =

ti+1 − t , h

ϕi+1 (t) =

t − ti . h

(210)

Therefore, G (a) X

f (t) =

(b)

vj ϕj (t) = vi ϕi (t) + vi+1 ϕi+1 (t)

j=0

ti+1 − t t − ti + vi+1 . (211) h h Here (a) is the definition of the finite-element reconstruction operator, (b) uses the local support of the nodal basis functions, and (c) substitutes the explicit affine formulas for ϕi and ϕi+1 on [ti , ti+1 ]. Differentiating the preceding affine expression on the open interval (ti , ti+1 ) gives     1 1 (b) vi+1 − vi (a) ′ f (t) = vi − + vi+1 = . (212) h h h (c)

= vi

Here (a) differentiates the two affine terms, and (b) combines them over the common denominator h. The derivative may fail to exist at the finitely many grid points, but (212) holds almost everywhere on I, which is sufficient for every integral below. 34

By the definition of the discrete Brownian seminorm, 2 Z A G−1 Z ti+1 G−1 Z ti+1  vi+1 − vi (b) X (c) X (a) 2 2 2 ′ ′ (f (t)) dt = (f (t)) dt = ∥f ∥B,h = dt h −A i=0 ti i=0 ti 2 Z ti+1 2 G−1  G−1  vi+1 − vi vi+1 − vi (d) X (e) X = 1 dt = (ti+1 − ti ) h h ti i=0 i=0 2 G−1 G−1  vi+1 − vi (f ) X (g) 1 X 2 = h = (vi+1 − vi ) . h h i=0 i=0

(213)

Here (a) expands the squared seminorm, (b) partitions [−A, A] into its G mesh intervals, (c) substitutes (212), (d) moves the intervalwise constant factor outside each integral, (e) evaluates the integral of the constant function 1, (f) uses the uniform-grid identity ti+1 − ti = h, and (g) simplifies h/h2 = 1/h. By the componentwise definition of the first-difference operator, 2 (a)

∥Dv∥2 =

G−1 X

2 (b)

((Dv)i ) =

i=0

G−1 X

2

(vi+1 − vi ) .

(214)

i=0

Here (a) is the definition of the squared Euclidean norm on RG , and (b) uses (Dv)i = vi+1 − vi . Combining (213) and (214) yields 2

∥f ∥B,h =

1 1 2 ∥Dv∥2 = v⊤ D⊤ Dv. h h

(215)

The last equality follows from ∥w∥22 = w⊤ w with w = Dv, together with (Dv)⊤ = v⊤ D⊤ . We next express the same energy through the stiffness matrix. By the entrywise definition of K, Z A G G G G (b) X X (a) X X ⊤ v r vs vr Krs vs = ϕ′r (t)ϕ′s (t) dt v Kv = (c)

=

−A

r=0 s=0

r=0 s=0

Z AX G X G

(d) vr vs ϕ′r (t)ϕ′s (t) dt =

−A r=0 s=0 (e)

Z A

=

−A

2

(f )

Z A

G X

−A

r=0

! vr ϕ′r (t)

G X

! vs ϕ′s (t)

dt

s=0

2

(f ′ (t)) dt = ∥f ∥B,h .

(216)

Here (a) expands the quadratic form componentwise, (b) substitutes the definition of Krs , (c) interchanges PG the integral with two finite sums, (d) factors the double sum, (e) uses f ′ = r=0 vr ϕ′r almost everywhere, and (f) is the definition of the discrete Brownian seminorm. Combining (215) and (216) gives   1 v⊤ K − D⊤ D v = 0, ∀v ∈ RG+1 . (217) h Define M := K −

1 ⊤ D D. h

(218)

The matrix M is symmetric because Krs = Ksr and (D⊤ D)⊤ = D⊤ D. Let x, y ∈ RG+1 be arbitrary. Then (a)

(b)

2x⊤ My = (x + y)⊤ M(x + y) − x⊤ Mx − y⊤ My = 0.

(219)

Here (a) expands the first quadratic form and uses the symmetry identity y⊤ Mx = x⊤ My, while (b) applies (217) to x + y, x, and y. Thus x⊤ My = 0 for all x, y. Taking x = er and y = es gives (a)

(b)

Mrs = e⊤ r Mes = 0,

35

r, s = 0, . . . , G.

(220)

Here (a) extracts the (r, s) entry of M, and (b) uses the preceding bilinear identity. Therefore M = 0, and hence 1 K = D⊤ D. (221) h It remains to expand D⊤ D. The entries of D are   −1, r = i, Dir = 1, r = i + 1,   0, otherwise,

i = 0, . . . , G − 1,

r = 0, . . . , G.

(222)

Consequently,  1,    G−1 2, X  (a) (b) D⊤ D rs = Dir Dis = −1,  i=0   0,

r = s ∈ {0, G}, r = s ∈ {1, . . . , G − 1}, |r − s| = 1, |r − s| ≥ 2.

(223)

Here (a) is the entrywise formula for a matrix product. For (b), an endpoint column of D contains one nonzero entry of magnitude one, an interior column contains two nonzero entries of magnitude one, adjacent columns overlap in exactly one row with product −1, and nonadjacent columns have disjoint supports. Thus   1 −1 0 · · · 0  ..  −1 2 −1 . . . .      D⊤ D =  0 −1 2 . . . 0  . (224)    .  .. .. ..  .. . . . −1 0 · · · 0 −1 1 Combining this identity with (221) proves the matrix formula, while (213) and (214) prove the asserted energy identities. Lemma D2 (Kernel of the stiffness matrix). The stiffness matrix satisfies ker (K) = span (1) ,

(225)

where ⊤

1 := (1, . . . , 1) ∈ RG+1 .

(226)

Proof. By Lemma D1, the stiffness matrix admits the factorization 1 K = D⊤ D. h We first prove that

(227)

ker(K) = ker(D).

(228)

Let v ∈ ker(K) be arbitrary. By the definition of the kernel, Kv = 0. Multiplying this identity from the left by v (a)

(b)

⊤

(229)

and using (227) gives (c)

0 = v⊤ 0 = v⊤ (Kv) = v⊤ Kv (d) 1 ⊤ ⊤ (e) 1 (f ) 1 ⊤ 2 = v D Dv = (Dv) (Dv) = ∥Dv∥2 . h h h

(230)

Here (a) uses the identity v⊤ 0 = 0, (b) substitutes (229), (c) uses the associativity of matrix multiplication, (d) substitutes the factorization (227), (e) uses v⊤ D⊤ = (Dv)⊤ , and (f) applies the Euclidean identity 36

w⊤ w = ∥w∥22 with w = Dv. Since A > 0 and G ≥ 1, the mesh size h = 2A/G satisfies h > 0. Multiplying (230) by h therefore yields 2

∥Dv∥2 = 0.

(231)

By the componentwise definition of the Euclidean norm, (a)

2 (b)

0 = ∥Dv∥2 =

G−1 X

2

((Dv)i ) .

(232)

i=0

Here (a) is (231), and (b) is the definition of the squared Euclidean norm on RG . Fix i ∈ {0, . . . , G − 1}. Since every real square is nonnegative, we have (a)

2

X (b) G−1

0 ≤ ((Dv)i ) ≤

2 (c)

((Dv)r ) = 0.

(233)

r=0

Here (a) uses the nonnegativity of the square of a real number, (b) uses the fact that the sum contains the displayed term and that all of its remaining terms are nonnegative, and (c) applies (232). The two inequalities in (233) force ((Dv)i )2 = 0. A real number has square zero if and only if the number itself is zero, so (Dv)i = 0.

(234)

Since i was arbitrary, every component of Dv is zero. It follows that Dv = 0.

(235)

ker(K) ⊆ ker(D).

(236)

Thus v ∈ ker(D), and consequently

Conversely, let v ∈ ker(D) be arbitrary. Then Dv = 0.

(237)

Using (227), we obtain (a) 1

Kv =

(b) 1

D⊤ Dv =

(c) 1

D⊤ (Dv) =

(d)

D⊤ 0 = 0.

(238) h h h Here (a) substitutes (227), (b) uses the associativity of matrix multiplication, (c) substitutes (237), and (d) uses D⊤ 0 = 0 and (1/h)0 = 0. Therefore v ∈ ker(K), and hence ker(D) ⊆ ker(K).

(239)

Combining (236) and (239) proves (228). It remains to characterize ker(D). Let v = (v0 , . . . , vG )⊤ ∈ ker(D). Then Dv = 0.

(240)

For every i ∈ {0, . . . , G − 1}, we therefore have (a)

(b)

0 = (Dv)i = vi+1 − vi .

(241)

Here (a) takes the ith component of (240), and (b) uses the definition of the first-difference operator. Rearranging (241) gives vi+1 = vi ,

i = 0, . . . , G − 1.

(242)

j = 0, . . . , G.

(243)

We now show by induction that vj = v0 ,

The statement is immediate for j = 0. Assume that it holds for some j ∈ {0, . . . , G − 1}, so that vj = v0 . Applying (242) with i = j gives (a)

(b)

vj+1 = vj = v0 . 37

(244)

Here (a) is (242), and (b) is the induction hypothesis. Thus (243) holds for every j = 0, . . . , G. Setting c := v0 and recalling that 1 = (1, . . . , 1)⊤ ∈ RG+1 , we obtain     v0 c (a)  .  (b)  .  (c) (245) v =  ..  =  ..  = c1. vG

c

Here (a) writes v componentwise, (b) uses (243), and (c) factors out the common scalar c. Therefore every vector in ker(D) belongs to span(1), and hence ker(D) ⊆ span(1).

(246)

For the reverse inclusion, let v ∈ span(1). By the definition of the linear span, there exists c ∈ R such that v = c1.

(247)

For every i ∈ {0, . . . , G − 1}, the corresponding component of Dv satisfies (a)

(b)

(c)

(Dv)i = vi+1 − vi = c − c = 0.

(248)

Here (a) uses the definition of D, (b) uses (247), and (c) simplifies the scalar difference. Thus every component of Dv is zero, so Dv = 0.

(249)

span(1) ⊆ ker(D).

(250)

ker(D) = span(1).

(251)

ker(K) = ker(D) = span(1),

(252)

Therefore v ∈ ker(D), which proves

Combining (246) and (250) yields

Finally, combining (228) and (251) gives

which proves the claim. Lemma D3 (Block decomposition). Assume that G = 2m, so that the anchor is located at ti0 = 0 with i0 = m, and use the reduced anchored ordering in (17). Then the anchored stiffness matrix satisfies   1 Tm 0 K0 = , (253) 0 Tm h m×m where Tm = (τrs )m−1 is defined entrywise by r,s=0 ∈ R  1, r = s = 0,    2, r = s ∈ {1, . . . , m − 1}, τrs :=  −1, |r − s| = 1,    0, otherwise.

(254)

Equivalently, for m ≥ 2, 1

−1

 −1   Tm =  0   .  ..

2 −1 .. .

0

···

whereas T1 = (1). 38

··· . −1 . . .. . 2 .. .. . . 0 −1 0

 0 ..  .    , 0   −1 2

(255)

e ∈ R2m be arbitrary, and partition it as Proof. Let v   a e= v , a = (a0 , . . . , am−1 )⊤ , b

b = (b0 , . . . , bm−1 )⊤ .

(256)

Set v := Re v ∈ R2m+1 .

(257)

By the reduced-coordinate ordering in (17), the components of the three vectors satisfy aj = vej = vj ,

bj = vem+j = v2m−j ,

j = 0, . . . , m − 1,

(258)

and the inserted anchored component satisfies vm = 0.

(259)

m×m Define Em = (Ers )m−1 entrywise by r,s=0 ∈ R   −1, s = r, Ers := 1, s = r + 1 and r ∈ {0, . . . , m − 2},   0, otherwise,

and let Jm = (Jrs )m−1 r,s=0 denote the reversal permutation matrix defined by ( 1, r + s = m − 1, Jrs := 0, otherwise. The definition of Em gives, for every x = (x0 , . . . , xm−1 )⊤ ∈ Rm , ( xr+1 − xr , r = 0, . . . , m − 2, (Em x)r = −xm−1 , r = m − 1.

(260)

(261)

(262)

We first compute the action of DR on the left half-grid. For i = 0, . . . , m − 1, (a)

(b)

(DRe v)i = (Dv)i = vi+1 − vi ( ai+1 − ai , i = 0, . . . , m − 2, (d) (c) = = (Em a)i . −am−1 , i = m − 1,

(263)

Here (a) uses v = Re v, (b) uses the componentwise definition of D, (c) uses (258) and (259), and (d) follows from (262). We next compute the right-half action. By (261), for every s = 0, . . . , m − 1, (a)

(−Jm Em b)s = −(Em b)m−1−s ( bm−1 , (b) = bm−s−1 − bm−s ,

s = 0, s = 1, . . . , m − 1.

(264)

Here (a) uses the fact that Jm reverses the order of the components, and (b) applies (262): when s = 0, the index is m − 1, whereas when s ≥ 1, the index m − 1 − s belongs to {0, . . . , m − 2}. On the other hand, for s = 0, . . . , m − 1, (a)

(b)

(DRe v)m+s = (Dv)m+s = vm+s+1 − vm+s ( bm−1 , s = 0, (c) (d) = = (−Jm Em b)s . bm−s−1 − bm−s , s = 1, . . . , m − 1,

(265)

Here (a) again uses v = Re v, (b) uses the definition of D, (c) uses vm = 0 and the identity vm+r = bm−r for r = 1, . . . , m, which follows from (258), and (d) invokes (264).

39

Combining (263) and (265) yields (a)



(c)



DRe v = =

Em a −Jm Em b

Em 0



(b)



= 

Em 0

  a −Jm Em b 0

0 e. v −Jm Em

(266)

Here (a) concatenates the first m components from (263) and the last m components from (265), (b) is e ∈ R2m was arbitrary, (266) implies the matrix block-matrix multiplication, and (c) uses (256). Because v identity   Em 0 DR = . (267) 0 −Jm Em The reversal matrix is orthogonal. Indeed, for every r, s ∈ {0, . . . , m − 1}, ( m−1 1, r = s, (c) (a) X (b) ⊤ (Jm Jm )rs = Jqr Jqs = = (Im )rs . 0, r ̸= s, q=0

(268)

Here (a) is the componentwise formula for a matrix product. For (b), the factor Jqr is nonzero only for q = m − 1 − r, whereas Jqs is nonzero only for q = m − 1 − s; these two row indices coincide exactly when r = s. Finally, (c) is the entrywise definition of the identity matrix. Since the entries agree for every r, s, J⊤ m Jm = Im . We now apply the stiffness factorization from Lemma D1. The anchored stiffness matrix satisfies (a) (b) 1 (c) 1 K0 = R⊤ KR = R⊤ D⊤ DR = (DR)⊤ (DR) h h    ⊤ (d) 1 Em 0 0 Em = ⊤ 0 −Jm Em h 0 −E⊤ m Jm    ⊤  ⊤ (f ) 1 (e) 1 0 0 Em Em Em Em = . = ⊤ 0 E⊤ 0 E⊤ h h m Jm Jm Em m Em

(269)

(270)

Here (a) is the definition of K0 , (b) substitutes K = h1 D⊤ D from Lemma D1, (c) uses (DR)⊤ = R⊤ D⊤ , (d) ⊤ substitutes (267) and uses (−Jm Em )⊤ = −E⊤ m Jm , (e) performs the block-matrix multiplication, and (f) uses (269). It remains to identify E⊤ m Em . For r, s ∈ {0, . . . , m − 1},  1, r = s = 0,    m−1  X 2, r = s ∈ {1, . . . , m − 1}, (c) (a) (b) (E⊤ Eqr Eqs = = τrs . (271) m Em )rs =  −1, |r − s| = 1,  q=0   0, otherwise, Here (a) is the componentwise formula for matrix multiplication. To justify (b), observe from (260) that column zero has the single nonzero entry E00 = −1, so its squared Euclidean norm is one. Every column r ∈ {1, . . . , m − 1} has exactly the two nonzero entries Er−1,r = 1 and Err = −1, so its squared Euclidean norm is 12 + (−1)2 = 2. Two consecutive columns have exactly one common nonzero row: if s = r + 1, their product in row r is Err Er,r+1 = (−1)(1) = −1, and the symmetric case r = s + 1 is identical. Columns whose indices differ by at least two have no common nonzero row, so their inner product is zero. Finally, (c) uses the definition of τrs in (254). Since the entries agree for all r, s, E⊤ m Em = Tm .

(272)

 Tm 0

(273)

Substituting (272) into (270) gives K0 =

1 h

which is exactly (253) and proves the lemma.

40

 0 , Tm

e ∈ RG be its reduced anchored Lemma D4 (Brownian energy in increment coordinates). Let f ∈ Hh0 , and let v nodal coefficient vector. Define the corresponding full anchored nodal vector and increment vector by v := Re v ∈ RG+1 ,

(274) ⊤

G

w := Dv = DRe v = (w0 , . . . , wG−1 ) ∈ R ,

(275)

so that f = R(v) = R (Re v). Then 1 1 2 2 ∥f ∥B,h = ∥w∥2 = h

G−1 X

h i=0

wi2 .

(276)

Proof. By (274), v ∈ RG+1 , and by (275), w ∈ RG . For every i ∈ {0, . . . , G − 1}, the ith increment satisfies (a)

(b)

wi = (Dv)i = vi+1 − vi .

(277)

Here (a) takes the ith component of the defining identity w = Dv in (275), and (b) uses the definition of the first-difference operator D. Since f = R(v), Proposition 1 gives both the quadratic-form representation of the Brownian energy and the factorization of the stiffness matrix. Therefore,   1 ⊤ (a) (b) (c) 1 (d) 1 2 ⊤ ∥f ∥B,h = v⊤ Kv = v⊤ D D v = v⊤ D⊤ Dv = (Dv) (Dv) h h h (e) 1

=

h

(f ) 1

w⊤ w =

h

2 (g) 1

∥w∥2 =

G−1 X

h i=0

wi2 .

(278)

Here (a) is the Brownian-energy quadratic-form identity established in Proposition 1; (b) substitutes the stiffness factorization K = h1 D⊤ D from the same proposition; (c) moves the scalar factor 1/h outside the matrix product; (d) uses (Dv)⊤ = v⊤ D⊤ and the associativity of matrix multiplication; (e) substitutes w = Dv from (275); (f) uses the definition ∥x∥22 = x⊤ x for vectors in Euclidean space; and (g) expands that norm componentwise using w = (w0 , . . . , wG−1 )⊤ . Thus (278) is exactly (276), which proves the lemma. e ∈ RG be its reduced anchored Lemma D5 (Brownian energy in spectral coordinates). Let f ∈ Hh0 , and let v nodal coefficient vector, so that f = R (Re v) .

(279)

K0 = R⊤ KR

(280)

Then the anchored stiffness matrix

is symmetric positive definite. Consequently, it admits an orthogonal eigendecomposition K0 = QΛQ⊤ , where Q ∈ R

G×G

Λ = diag (λ1 , . . . , λG ) ,

λi > 0,

i = 1, . . . , G.

(281)

is orthogonal. Define ⊤

(282)

λi c2i .

(283)

e = (c1 , . . . , cG ) . c := Q⊤ v Then 2

∥f ∥B,h = c⊤ Λc =

G X i=1

D.1

Proof of Lemma D5

Proof. Set v := Re v ∈ RG+1 .

(284)

f = R(v).

(285)

By (279) and (284),

41

We first verify that K0 is symmetric positive definite. Write the anchoring reconstruction matrix as R = [r1 , . . . , rG ] ,

(286)

where r1 , . . . , rG ∈ RG+1 are its columns. By the explicit construction in (18), these columns are pairwise distinct canonical basis vectors of RG+1 . Therefore, for every r, s ∈ [G],  (a) (b) (c) R⊤ R rs = r⊤ (287) r rs = δrs = (IG )rs . Here (a) is the entrywise rule for the product R⊤ R, (b) uses the orthonormality of pairwise distinct canonical basis vectors, and (c) is the entrywise definition of the identity matrix. Since (287) holds for every r, s ∈ [G], R⊤ R = IG .

(288)

1 ⊤ D D. h

(289)

By Proposition 1, K= Hence (a)

K⊤ =



⊤ ⊤ (c) 1 ⊤ (d) 1 ⊤ (b) 1 ⊤ D D = D D⊤ = D D = K. h h h

(290)

Here (a) substitutes the stiffness factorization, (b) uses (αA)⊤ = αA⊤ for real α and (AB)⊤ = B⊤ A⊤ , (c) uses (D⊤ )⊤ = D, and (d) substitutes the stiffness factorization again. It follows that ⊤ (b) ⊤ ⊤ (c) ⊤ (a) (d) ⊤ K⊤ = R K R = R KR = K0 . (291) 0 = R KR Here (a) uses the definition of K0 , (b) applies the transpose-of-a-product rule, (c) uses (290), and (d) again uses the definition of K0 . Thus K0 is symmetric. To prove positive definiteness, let z ∈ RG \ {0} be arbitrary and define y := Rz ∈ RG+1 .

(292)

Using (288), we obtain 2 (a)

(b)

⊤

∥y∥2 = (Rz) (Rz) = z⊤ R⊤ Rz (c)

2 (e)

(d)

= z⊤ IG z = ∥z∥2 > 0.

(293)

Here (a) uses ∥y∥22 = y⊤ y and (292), (b) uses (Rz)⊤ = z⊤ R⊤ and associativity, (c) substitutes (288), (d) uses IG z = z and the definition of the Euclidean norm, and (e) follows from z ̸= 0. Consequently, y ̸= 0.

(294)

Moreover, the construction of R inserts a zero at the anchor index, and therefore yi0 = 0.

(295)

The quadratic form of K0 at z satisfies (a)

(b)

⊤

z⊤ K0 z = z⊤ R⊤ KRz = (Rz) K (Rz) (d) 1

(c)

= y⊤ Ky =

h

2

(e)

∥Dy∥2 ≥ 0.

(296)

Here (a) substitutes K0 = R⊤ KR, (b) uses z⊤ R⊤ = (Rz)⊤ and associativity, (c) substitutes (292), (d) applies Proposition 1 to y, and (e) uses h > 0 and the nonnegativity of a squared Euclidean norm. Suppose, for contradiction, that z⊤ K0 z = 0.

(297)

Combining (296) and (297) gives (a) 1

0 =

h

2 (b)

2

(c)

∥Dy∥2 =⇒ ∥Dy∥2 = 0 =⇒ Dy = 0.

42

(298)

Here (a) uses (296) with equality as imposed by (297), (b) multiplies by the positive scalar h, and (c) uses the fact that a vector has zero squared Euclidean norm if and only if it is the zero vector. By Lemma D2, (298) implies that there exists a ∈ R such that y = a1.

(299)

Taking the anchor component in (299) and using (295) yields (a)

(b)

0 = yi0 = a,

(300)

where (a) is (295) and (b) takes the i0 th component of (299). Thus a = 0, and (299) gives y = 0, contradicting (294). Therefore z⊤ K0 z > 0

for every z ∈ RG \ {0}.

(301)

Together with (291), this proves that K0 is symmetric positive definite. By the spectral theorem for real symmetric matrices, there exist an orthogonal matrix Q ∈ RG×G and a real diagonal matrix Λ such that K0 = QΛQ⊤ .

(302)

The positive definiteness in (301) implies that every diagonal entry λi of Λ is strictly positive. Indeed, if qi denotes the ith column of Q, then ∥qi ∥2 = 1 and K0 qi = λi qi , so (a)

(c)

(b)

λi = λi qi⊤ qi = qi⊤ K0 qi > 0.

(303)

Here (a) uses qi⊤ qi = 1, (b) uses K0 qi = λi qi , and (c) applies (301) to the nonzero vector qi . We now derive the Brownian energy identity. By (216), (284), and the definition of K0 , 2

(a)

(b)

⊤

∥f ∥B,h = v⊤ Kv = (Re v) K (Re v) (c)

(d)

e⊤ R⊤ KRe e⊤ K0 v e =v v = v

(e)

e⊤ QΛQ⊤ v e =v   (g) (f ) ⊤ e Λ Q⊤ v e = c⊤ Λc. = Q⊤ v

(304)

e⊤ R⊤ and associativity, (d) Here (a) applies (216) to f = R(v), (b) substitutes (284), (c) uses (Re v )⊤ = v ⊤ e⊤ Q = (Q⊤ v e)⊤ , and uses K0 = R KR, (e) substitutes the orthogonal eigendecomposition of K0 , (f) uses v (g) substitutes (282). Because Λ is diagonal, i, j ∈ [G].

Λij = λi δij ,

(305)

Therefore G G (a) X X

c⊤ Λc =

G G (b) X X

ci Λij cj =

i=1 j=1 G (c) X

=

G (d) X

ci λi ci =

i=1

ci λi δij cj

i=1 j=1

λi c2i .

(306)

i=1

Here (a) expands the quadratic form componentwise, (b) substitutes (305), (c) uses δij = 0 for i ̸= j and δii = 1 to collapse the inner sum to its j = i term, and (d) rearranges the scalar factors and uses ci ci = c2i . Combining (304) and (306) yields (283), which proves the lemma. Lemma D6 (Explicit spectral Brownian basis). Assume that G = 2m, and let K0 = QΛQ⊤ ,

Λ = diag (λ1 , . . . , λG )

43

(307)

be the orthogonal eigendecomposition obtained in Theorem 3. Thus Q ∈ RG×G is orthogonal and λi > 0 for every i ∈ [G]. Let ei ∈ RG denote the ith canonical basis vector. For every i ∈ [G], define ψi := R (RQei ) =

G X

(RQei )j ϕj =

G X

(RQ)ji ϕj .

(308)

j=0

j=0

Then the following statements hold. (i) Each function ψi belongs to Hh0 , and the family {ψ1 , . . . , ψG } forms a basis of Hh0 . e ∈ RG be its reduced anchored nodal coefficient vector. Then f admits the unique (ii) Let f ∈ Hh0 , and let v representation f=

G X

ci ψ i ,

(309)

i=1

where ⊤

e ∈ RG c := (c1 , . . . , cG ) = Q⊤ v

(310)

is the spectral coefficient vector of f . (iii) The basis is orthogonal with respect to the Brownian inner product. More precisely, for every i, j ∈ [G], ⟨ψi , ψj ⟩B,h = λi δij .

(311)

Proof. We first verify that every function defined in (308) belongs to the anchored space. Fix i ∈ [G]. Because ψi is a finite linear combination of ϕ0 , . . . , ϕG , we have ψi ∈ Hh . Moreover, the anchor is the grid point ti0 = 0, and the nodal basis satisfies ϕj (0) = δji0 . Therefore, G (a) X

G (b) X

ψi (0) =

(RQ)ji ϕj (0) =

j=0

(RQ)ji δji0

j=0 G (d) X

(c)

= (RQ)i0 i =

(e)

Ri0 r Qri = 0.

(312)

r=1

Here (a) uses the definition of ψi in (308); (b) substitutes ϕj (0) = δji0 ; (c) uses the defining property of the Kronecker delta to retain only the term j = i0 ; (d) expands the (i0 , i) entry of the matrix product RQ; and (e) uses the construction of the anchoring reconstruction matrix R, whose i0 th row is the zero row because the anchor coefficient is fixed at zero. Thus ψi (0) = 0, and hence ψi ∈ Hh0 . Since i ∈ [G] was arbitrary, this conclusion holds for every member of the family. We next prove existence and uniqueness of the spectral expansion. Let f ∈ Hh0 be arbitrary. By the reduced e ∈ RG such that anchored nodal representation, there exists a unique vector v f = R (Re v) .

(313)

Define c by (310). Because Q is orthogonal, Q⊤ Q = IG ,

QQ⊤ = IG .

(314)

Consequently, (a)

(b)

(c)

e = IG v e =v e. Qc = QQ⊤ v

(315)

e; (b) uses QQ⊤ = IG from (314); and (c) uses the defining action Here (a) substitutes the definition c = Q⊤ v of the identity matrix.

44

Writing c =

PG

i=1 ci ei and using (313) and (315), we obtain (a)

(b)

f = R (Re v) = R (RQc) ! ! G G X X (c) (d) = R RQ ci e i = R ci RQei i=1 (e)

=

G X

i=1 (f )

ci R (RQei ) =

i=1

G X

ci ψi .

(316)

i=1

e = Qc from (315); (c) expands c in the canonical basis of RG ; (d) uses the Here (a) is (313); (b) substitutes v linearity of the matrix maps Q and R; (e) uses the linearity of the finite-element reconstruction operator R; and (f) applies the definition of ψi in (308). This proves the existence of the representation (309). To prove uniqueness, suppose that another vector ⊤

(317)

di ψi .

(318)

d := (d1 , . . . , dG ) ∈ RG satisfies f=

G X i=1

Repeating the linearity calculation in (316) with d in place of c gives G X

di ψi = R (RQd) .

(319)

R (Re v) = R (RQd) .

(320)

i=1

Combining (313), (318), and (319) yields

The nodal representation in Hh is unique. Therefore, Re v = RQd.

(321)

By (288), R⊤ R = IG . Left-multiplying (321) by R⊤ gives (a)

(b)

(c)

e = IG v e = R⊤ Re v v = R⊤ RQd (d)

(e)

= IG Qd = Qd.

(322)

Here (a) uses the action of the identity matrix; (b) substitutes R⊤ R = IG ; (c) left-multiplies (321) by R⊤ ; (d) again uses R⊤ R = IG ; and (e) uses IG Qd = Qd. It follows that (a)

(b)

(c)

(d)

e = c. d = IG d = Q⊤ Qd = Q⊤ v

(323)

Here (a) uses the action of the identity matrix; (b) uses Q⊤ Q = IG ; (c) substitutes (322); and (d) uses the definition of c in (310). Thus the coefficients in (309) are unique, proving (ii). The existence of the expansion for every f ∈ Hh0 shows that {ψ1 , . . . , ψG } spans Hh0 . To verify linear PG e=0 independence, suppose i=1 di ψi = 0. The zero function has reduced anchored nodal coefficient vector v and hence, by (310), spectral coefficient vector Q⊤ 0 = 0. The uniqueness established in (323) therefore gives d = 0. Thus the family is linearly independent. Since it both spans Hh0 and is linearly independent, it is a basis, proving (i). It remains to prove the Brownian orthogonality relation. For each i ∈ [G], set ⊤

ai := RQei = (a0i , . . . , aGi ) ∈ RG+1 .

45

(324)

Then ψi = R(ai ). Fix arbitrary i, j ∈ [G]. Using the definition of the Brownian inner product and the entrywise definition of the stiffness matrix in Proposition 1, we obtain ! G ! Z A Z A X G X (a) (b) ′ ′ ′ ′ ⟨ψi , ψj ⟩B,h = ψi (t)ψj (t) dt = ari ϕr (t) asj ϕs (t) dt −A

(c)

=

G X G X

−A

Z A ari asj

(f )

⊤

= (RQei )

s=0 (d)

ϕ′r (t)ϕ′s (t) dt =

−A

r=0 s=0

r=0 G X G X

(e)

ari Krs asj = ai⊤ Kaj

r=0 s=0 (g)

(h) ⊤ ⊤ ⊤ ⊤ K (RQej ) = e⊤ i Q R KRQej = ei Q K0 Qej

(i)

   (k) ⊤ (j) ⊤ ⊤ ⊤ = e⊤ QΛQ⊤ Qej = e⊤ i Q i Q Q Λ Q Q ej = ei Λej

(l)

(m)

= Λij = λi δij .

(325) PG

PG

Here (a) is the definition of the Brownian inner product; (b) uses ψi = r=0 ari ϕr and ψj = s=0 asj ϕs , and differentiates these finite sums on each open mesh interval; (c) expands the product and uses the linearity of RA the integral for finite sums; (d) uses Krs = −A ϕ′r (t)ϕ′s (t) dt; (e) is the componentwise formula for the bilinear ⊤ ⊤ matrix product ai⊤ Kaj ; (f) substitutes (324); (g) uses (RQei )⊤ = e⊤ i Q R and associativity; (h) substitutes ⊤ K0 = R KR; (i) substitutes the eigendecomposition (307); (j) uses associativity of matrix multiplication; (k) uses Q⊤ Q = IG and the action of the identity matrix; (l) uses the defining property of canonical basis vectors, e⊤ i Λej = Λij ; and (m) uses the diagonal structure Λij = λi δij . This proves (311) and hence (iii). In particular, if i ̸= j, then ⟨ψi , ψj ⟩B,h = 0, whereas if i = j, then 2

∥ψi ∥B,h = ⟨ψi , ψi ⟩B,h = λi > 0,

(326)

because λi > 0. This completes the proof.

E

Implications of the Equivalent Coordinate Representations

Function-space invariance. Theorem 2 establishes that the nodal, increment, and spectral parameterizations are mathematically equivalent realizations of the same finite Brownian RKHS. Consequently, the represented function space, the Brownian norm, and the associated approximation properties are intrinsic to the underlying RKHS and remain invariant under changes of coordinate realization. Controlled optimization. Because all three realizations represent exactly the same hypothesis class, differences cannot be attributed to representational capacity, approximation error, or intrinsic Brownian regularization. The sharper distinction is optimizer-specific: Proposition 3 forces mapped nodal and spectral GD/SGD trajectories to agree, Proposition 3 identifies increment GD with Brownian preconditioning, and Proposition 5 permits standard Adam to distinguish the DCT-VIII basis. The direct one-layer and recursive tests in Section 7 numerically verify these predictions under mapped initialization and shared update protocols. Connection to the spectral realization. The equivalence theorem is independent of the particular choice of coordinates. The explicit spectral realization presented in Section 5 constructs a block-structured orthogonal basis of the anchored finite Brownian RKHS through the complete eigendecomposition of the anchored stiffness matrix, providing the spectral implementation used in the optimization study.

F

Implications of the Explicit Spectral Representation

Explicit block-structured spectral coordinates. Unlike the nodal and increment parameterizations, the spectral coordinates diagonalize the Brownian energy. The energy determines the eigenspaces, but the twofold multiplicity of every eigenvalue does not determine a unique eigenbasis inside each left–right eigenspace. The displayed block-diagonal eigendecomposition fixes an explicit orthogonal coordinate system by respecting the reduced-coordinate ordering and the two half-grid blocks. Diagonalization of the Brownian energy. The explicit spectral basis diagonalizes the Brownian inner product, reducing the Brownian norm to a weighted Euclidean norm of the spectral coefficients. Consequently, each spectral mode contributes independently to the total Brownian energy, providing a natural decomposition of the finite Brownian RKHS into orthogonal Brownian spectral modes. 46

Implications for optimization. The spectral transformation is orthogonal, so mapped nodal–spectral GD/SGD must agree after coordinates, initialization, scalar steps, and minibatches are mapped consistently. Standard Adam is different: its universal orthogonal equivariance group is only the signed permutations, and the DCT-VIII matrix is not one. The controlled tests in Section 7 numerically verify both the exact GD/SGD identity and the permitted Adam separation.

G

Supplementary Experimental Evidence

This appendix separates direct theorem-verification protocols from secondary predictive evidence. The old synthetic tables and legacy coordinate curves are omitted because their raw outputs mixed incompatible sample sizes or lacked the mapped-initialization and optimizer metadata required for a valid equivariance test. The retained EuroSAT and spatially separated Salinas studies were regenerated from frozen pipelines and independently audited from saved predictions and probabilities. G.1

Theory-Validation Protocols and Reproducibility Audit

Evidence classes. The direct tests E0, E1A–E1C, and E2 numerically verify algebraic identities, mapped optimizer trajectories, conditioning statements, and the theorem-specific standard-Adam counterexample. They are not tuned predictive comparisons. The EuroSAT and spatially separated Salinas studies instead compare different model classes and are retained only as secondary predictive and robustness evidence; they are not used to establish coordinate equivariance. Table 5: Frozen verification and audit protocols. Identity-test arms share one mapped represented initialization, objective normalization, scalar step sequence, and, for SGD, minibatch order; they are never tuned separately. Stage

Purpose

Data/runs

Mesh

Budget

Precision

E0 E1A E1B E1C E2 Real-data audit

Finite Brownian RKHS, spectrum, and coordinate identities One-layer mapped-trajectory identities for GD and SGD Conditioning and fixed-step convergence under mesh refinement Basis dependence of standard Adam Recursive blockwise mapped-trajectory identities Independent metric and statistical reconstruction

125 random vectors 256 samples, 5 seeds 256 samples, 5 seeds 256 samples, 5 seeds 128 samples, 5 seeds 390 prediction artifacts

G ∈ {8, 16, 32, 64, 128} G ∈ {8, 16, 32, 64, 128} G ∈ {8, 16, 32, 64, 128} G ∈ {8, 16, 32, 64, 128} (8, 16, 32) and (32, 64, 128) EuroSAT and Salinas

4097-point dense grid 40 GD; 80 SGD; batch 32 gap 10−8 ; analytic steps 200 steps; full batch 8 GD; 16 SGD; batch 32 5720 published-value checks

float64 CPU float64 CPU float64 CPU float64 CPU float64 CPU NumPy reconstruction

E0: finite Brownian RKHS identities. We use A = 2, G ∈ {8, 16, 32, 64, 128}, seeds {0, 1, 2, 3, 4}, and five random vectors per seed. All matrix, spectrum, Brownian-Gram, reconstruction, and energy diagnostics are evaluated in float64 on CPU; reconstruction is checked on 4097 points. E1A: one-layer mapped-trajectory identities for GD and SGD. The Brownian-regularized leastsquares objective uses n = 256, ρ = 0.1, and noise standard deviation 0.03. Every coordinate arm begins from one reduced-nodal function and is mapped exactly. Full-batch GD uses 40 updates; SGD uses 80 updates with batch size 32 and an identical minibatch sequence. No coordinate arm is tuned separately. E1B: conditioning and convergence under mesh refinement. The pure Brownian Hessians are constructed directly in nodal, spectral, and increment coordinates. For least squares, n = 256 and ρ ∈ {0.05, 0.1, 0.2, 0.5}. The convergence diagnostic uses ρ = 0.1, the analytic optimal scalar step for each SPD quadratic, a relative objective-gap tolerance of 10−8 , and at most 200,000 iterations. These steps are analytic, not validation-selected. E1C: basis dependence of standard Adam. We use scalar β1 = 0.9, β2 = 0.999, ε = 10−8 , zero initial moment states, bias correction, learning rate 0.003 for the regularized objective, and 0.01 for the deterministic linear counterexample. The experiment uses 200 full-batch updates, no AdamW, weight decay, clipping, or AMSGrad, and no coordinate-specific tuning. E2: recursive blockwise mapped-trajectory identities. The recursive model contains three trainable Brownian profile blocks with grid sets (8, 16, 32) and (32, 64, 128), trainable cross-layer couplings, and identically initialized non-profile parameters. The objective uses n = 128 and ρ = 0.05. We run eight full-batch steps at η = 5 × 10−5 and sixteen shared-minibatch steps at η = 2 × 10−5 with batch size 32. Each increment block is compared with the nodal Brownian-gradient update scaled by its own 1/hb .

47

Independent real-data audit. For the retained EuroSAT and spatially separated Salinas studies, accuracy, multiclass F1 , calibration, and likelihood metrics were reconstructed directly from 390 saved prediction artifacts, including 5,820 per-class records and 88,920 confusion-matrix cells. A separate implementation reproduced 5,720 reported AUC, exact-test, crossed-bootstrap, and variance-component entries across 11 tables with zero mismatch and maximum discrepancy 8.66 × 10−15 . G.2

EuroSAT Under Test-Time Spectral-Band Dropout

Protocol. Patch MLP and Spatial Path-Atomic BKL receive the same Spatial65 representation of all 13 EuroSAT bands. For each of five paired seeds, both models are trained once on the same clean split. At test time only, one to six bands are zeroed before feature extraction; masks are nested within each seed and shared between models. For a metric M , retention is M (d)/M (0), and the normalized area under the retention curve summarizes the complete degradation path. We additionally record ECE, NLL, per-class F1 , inference time, and parameter count. Table 6: EuroSAT robustness under nested test-time spectral-band dropout. Values are means over five paired seeds. Retention AUC is normalized over zero to six zeroed bands; higher is better. Maximum ECE is lower-is-better.

Model

Accuracy Retention AUC (%)

Macro-F1 Retention AUC (%)

Worst Accuracy Retention (%)

Worst Macro-F1 Retention (%)

Maximum ECE

37.78 46.81

32.16 42.10

21.80 23.38

13.88 17.68

0.774 0.629

Patch MLP Spatial Path-Atomic BKL

Spatial Path-Atomic BKL

Patch MLP (b) Macro-F1 retention

100

80

80

Macro-F1 retention (%)

Accuracy retention (%)

(a) Accuracy retention 100

60

40

20

0

60

40

20

0

1

2 3 4 Number of zeroed spectral bands

5

0

6

0

1

2 3 4 Number of zeroed spectral bands

5

6

Figure 3: EuroSAT accuracy- and macro-F1 -retention trajectories under nested test-time spectral-band dropout. Lines show paired-seed means and shading one standard deviation.

Spatial Path-Atomic BKL

Patch MLP

(a) Calibration error

(b) Negative log-likelihood 20.0

0.8 Negative log-likelihood

Expected calibration error

17.5 0.6

0.4

0.2

15.0 12.5 10.0 7.5 5.0 2.5 0.0

0.0 0

1

2 3 4 Number of zeroed spectral bands

5

6

0

1

2 3 4 Number of zeroed spectral bands

5

6

Figure 4: EuroSAT calibration and likelihood degradation under nested spectral-band dropout; lower is better.

48

Results and scope. Spatial Path-Atomic BKL has normalized accuracy-retention AUC 46.81% versus 37.78% for Patch MLP and macro-F1 -retention AUC 42.10% versus 32.16%. All five pairs favor BKL for both AUCs and for maximum-ECE reduction; each exact two-sided sign-flip test gives p = 0.0625, the minimum nonzero value for five pairs. The result concerns the full degradation trajectory rather than uniform superiority at each severity. This image-level multispectral experiment is a BKL-versus-MLP comparison and is not used to establish coordinate equivariance. Table 7: Exact paired sign-flip analysis over five seeds. Improvements are oriented so that positive values favor Spatial Path-Atomic BKL.

Metric Accuracy retention AUC Macro-F1 retention AUC Worst accuracy retention Worst Macro-F1 retention Maximum ECE reduction G.3

Mean improvement

Two-sided p

Paired dz

BKL-favored seeds

0.0903 0.0994 0.0158 0.0379 0.1449

0.0625 0.0625 0.6875 0.4375 0.0625

1.61 2.11 0.17 0.46 1.53

5/5 5/5 3/5 3/5 5/5

Salinas Evaluation Under Spatial Separation and Missing Spectral Bands

Protocol. We evaluate the mathematically constrained finite VBKL implementation on the 204-band Salinas scene using a spatially separated protocol fixed before the final model evaluations. Labeled pixels are grouped into 24 × 24 spatial blocks, and a four-pixel cross-split buffer removes 19,152 boundary pixels. The retained split contains 23,998 training, 5,424 validation, and 5,555 test pixels, with all 16 classes represented in every split. Standardization is fitted on the training split only. The standardized spectra are divided by the training 0.99-quantile of their Euclidean norms and are radially projected, when necessary, to the radius-two ball. This projection affected 2 training spectra and no validation or test spectra (0 and 0, respectively). The constrained VBKL uses 32 paths, depth four under the convention of one normalized linear projection followed by three Brownian profiles, grid size 32, and interval radius A = 2. Every effective direction has unit Euclidean norm; every profile is anchored at zero and radially normalized to Brownian energy at most one. Hyperparameters were chosen using validation data only: seed 0 supplied the initial search, seeds 1–3 supplied independent validation confirmation, and the final clean models use seeds 4, 5, 6, 7, 8. The fixed Patch MLP has 15,728 parameters, whereas constrained VBKL has 10,128. Table 8: Clean Salinas performance (mean ± standard deviation over five final evaluation seeds). Model Patch MLP Constrained VBKL Model Patch MLP Constrained VBKL

Acc. (%)

Macro-F1 (%)

Bal. Acc. (%)

Worst F1 (%)

ECE

NLL

91.87 ± 0.93 92.18 ± 0.20

94.68 ± 1.14 95.49 ± 0.40

94.83 ± 1.28 95.23 ± 0.60

73.93 ± 2.55 73.83 ± 0.63

0.018 ± 0.002 0.015 ± 0.004

0.245 ± 0.013 0.218 ± 0.008

Parameters

Training time (s)

Inference (ms/sample)

15,728 10,128

5.7 ± 0.6 166.2 ± 18.9

0.001 ± 0.000 0.024 ± 0.000

Clean performance. Across the five final seeds, constrained VBKL obtains 92.18% accuracy and 95.49% Macro-F1, compared with 91.87% and 94.68% for Patch MLP. The corresponding balanced accuracies are 95.23% and 94.83%. Constrained VBKL also has lower mean ECE (0.015 versus 0.018) and lower mean NLL. The worst-class F1 values are essentially tied. Test-time spectral masking. Starting from each clean checkpoint, we mask nested random subsets of 16, 31, 47, 63, 78, and 94 bands. A masked raw coordinate is replaced by its training-set band mean, so it becomes exactly zero after training standardization. We cross five final model seeds with five independent mask seeds, giving 25 cells per corrupted level and 300 model evaluations in total. Confidence intervals resample the model-seed and mask-seed axes separately; exact sign-flip analyses of their marginal means are reported as sensitivity analyses.

49

Patch MLP

(b) Macro-F1 retention 100

80

80

60

40

20

0

(c) Balanced-accuracy retention 100 Balanced-accuracy retention (%)

100

Macro-F1 retention (%)

Accuracy retention (%)

(a) Accuracy retention

Constrained VBKL

60

40

20

0

10

20 30 Masked spectral bands (%)

40

0

0

10

20 30 Masked spectral bands (%)

80

60

40

20

0

40

0

10

20 30 Masked spectral bands (%)

40

Figure 5: Salinas accuracy, Macro-F1, and balanced-accuracy retention under nested spectral-band masking. Lines are crossed-design means and shaded regions are 95% crossed-bootstrap intervals. Table 9: Aggregate Salinas robustness under nested spectral-band masking. Entries are means with 95% crossedbootstrap intervals. Retention is higher-is-better; ECE and NLL increases are lower-is-better. Model

Accuracy-retention AUC (%)

Macro-F1-retention AUC (%)

Balanced-retention AUC (%)

76.30 [69.82, 82.13] 75.12 [71.53, 78.39]

67.98 [60.57, 74.78] 65.66 [60.29, 70.70]

72.26 [66.21, 77.99] 69.97 [65.71, 74.09]

Patch MLP Constrained VBKL Model Patch MLP Constrained VBKL

Level-6 Acc. ret. (%)

Level-6 F1 ret. (%)

ECE-increase AUC

NLL-increase AUC

41.92 [31.22, 53.95] 51.66 [44.13, 57.85]

28.22 [19.89, 37.42] 34.67 [29.02, 40.85]

0.184 [0.135, 0.235] 0.162 [0.134, 0.195]

1.568 [0.866, 2.367] 1.098 [0.850, 1.378]

Crossover rather than uniform dominance. Across the complete corruption path, Patch MLP has slightly higher mean retention AUC: the BKL-minus-MLP differences are -1.17, -2.32, and -2.29 percentage points for accuracy, Macro-F1, and balanced accuracy, respectively. All corresponding crossed-bootstrap intervals include zero. At the most severe level, however, constrained VBKL retains 51.66% of its clean accuracy and 34.67% of its clean Macro-F1, compared with 41.92% and 28.22% for Patch MLP. The level-six improvements are +9.74 and +6.45 percentage points. All five model-seed marginal means favor constrained VBKL for level-six accuracy retention, yielding an exact two-sided sign-flip value of 0.0625; the crossedbootstrap interval nevertheless still includes zero because mask realization and model–mask interaction remain substantial. Calibration and interpretation. Constrained VBKL exhibits a smaller average corruption-induced increase in both ECE and NLL. The crossed mean improvements are +0.021 for ECE-increase AUC and +0.469 for NLL-increase AUC. These effects are directionally favorable but their crossed-bootstrap intervals include zero. Thus this experiment supports a qualified conclusion: Patch MLP has slightly higher average retention AUC over the full masking path, whereas constrained VBKL is more competitive on clean data, has higher mean retention at the most severe masking level, and shows smaller likelihood and calibration deterioration. The effect is heterogeneous across both model initialization and mask realization, so the result does not justify a claim of uniform robustness dominance. Table 10: Crossed Salinas robustness comparisons. Improvements are oriented so that positive values favor constrained VBKL. The confidence interval is obtained by crossed bootstrap. The model-seed and mask-seed columns report BKL-favored/MLP-favored marginal means and exact two-sided sign-flip p-values. Mean improvement Metric

Favored model

Accuracy-retention AUC Macro-F1-retention AUC Balanced-accuracy-retention AUC Level-6 accuracy retention Level-6 Macro-F1 retention ECE-increase AUC NLL-increase AUC

Patch MLP Patch MLP Patch MLP Constrained VBKL Constrained VBKL Constrained VBKL Constrained VBKL

[95% CI]

Cells B/M

Model seeds B/M

pmodel

Mask seeds B/M

pmask

−1.17 [−6.86, +4.33] −2.32 [−10.28, +5.39] −2.29 [−8.88, +4.27] +9.74 [−0.54, +18.58] +6.45 [−2.64, +14.72] +0.021 [−0.032, +0.074] +0.469 [−0.207, +1.194]

14/11 12/13 12/13 19/6 18/7 16/9 15/10

2/3 2/3 2/3 5/0 4/1 4/1 4/1

0.6875 0.6875 0.6875 0.0625 0.1875 0.4375 0.1250

3/2 2/3 1/4 3/2 4/1 3/2 3/2

0.6250 0.4375 0.3125 0.2500 0.1875 0.3125 0.3125

50

(a) Accuracy-retention advantage

(b) Macro-F1-retention advantage 15

15

BKL − MLP Macro-F1 retention (pp)

BKL − MLP accuracy retention (pp)

20

10 5 0 −5 −10

10 5 0 −5 −10 −15 −20

−15 10

15

20 25 30 35 Masked spectral bands (%)

40

45

10

(c) Calibration advantage

15

40

45

40

45

(d) NLL advantage

4

0.25

20 25 30 35 Masked spectral bands (%)

Reduction in NLL increase

Reduction in ECE increase

0.20 0.15 0.10 0.05 0.00 −0.05

3

2

1

0

−0.10 10

15

20 25 30 35 Masked spectral bands (%)

40

45

10

Patch MLP

20 25 30 35 Masked spectral bands (%)

Constrained VBKL

(a) Calibration degradation

(b) Predictive-likelihood degradation 6 NLL increase from clean

0.5 ECE increase from clean

15

0.4 0.3 0.2 0.1

5 4 3 2 1

0.0

0 0

10

20 30 Masked spectral bands (%)

40

0

10

20 30 Masked spectral bands (%)

40

Figure 6: Salinas corruption-path comparisons. The upper panels show BKL-minus-MLP differences, with positive values favoring constrained VBKL; for ECE and NLL, positive values denote a smaller corruption-induced increase. The lower panels show the corresponding calibration and negative-log-likelihood degradation.

Table 11: Severity-by-severity Salinas retention. Means are taken over the full 5 × 5 crossed design at corrupted levels; the clean level uses five model seeds. Differences are BKL minus MLP in percentage points. Level

Bands

MLP Acc. ret.

BKL Acc. ret.

∆ Acc.

MLP F1 ret.

BKL F1 ret.

∆ F1

0 1 2 3 4 5 6

0 16 31 47 63 78 94

100.00 95.09 90.83 78.30 66.72 55.97 41.92

100.00 94.34 86.75 73.17 63.43 57.27 51.66

+0.00 -0.75 -4.09 -5.13 -3.29 +1.29 +9.74

100.00 91.68 85.46 68.70 55.36 42.67 28.22

100.00 93.27 81.60 62.62 48.91 40.28 34.67

+0.00 +1.60 -3.86 -6.08 -6.45 -2.39 +6.45

51

Record · ID 1006874 · SHA-256 cf61e728ca74c4af
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.