WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS ON INFINITE-DIMENSIONAL MANIFOLDS
arXiv:2606.09820v1 [math.FA] 8 Jun 2026
PHILIPP SCHMOCKER AND JOSEF TEICHMANN
Abstract. We generalize the universal approximation theorem for functional input neural networks (FNN) in [28] to differentiable maps by including the approximation of the derivatives. A FNN maps the input from a possibly infinite-dimensional weighted manifold to the real-valued hidden layer, on which a non-linear scalar activation function is applied, and then returns the output into a Banach space via some linear readouts. By proving a weighted Nachbin theorem, we establish a universal approximation theorem (UAT) for differentiable maps, which goes beyond the usual formulation on compact sets and also includes the approximation of the derivatives. This leads us to approximation results for non-anticipative functionals including the horizontal and vertical derivatives. As a further application, we show that linear functions of the signature are able to approximate path space functionals including their directional derivatives.
Contents 1.
2.
3.
4.
5.
Introduction 1.1. Notation 1.2. Bastiani calculus on σ-compact spaces 1.3. Manifolds over σ-compact model spaces 1.4. Examples of σ-compact model spaces Weighted spaces and differentiable maps 2.1. Weighted domains k 2.2. BΨ -maps over weighted domains k 2.3. Weighted manifolds and BΨ -maps thereon Weighted Nachbin theorems 3.1. Classical formulation 3.2. Subalgebras of Ψ-moderate growth 3.3. Weighted Nachbin theorems over finite-dimensional manifolds 3.4. Weighted Nachbin theorems over infinite-dimensional manifolds Weighted universal approximation of functional input neural networks 4.1. Functional input neural networks 4.2. Examples of additive families 4.3. Weighted UAT over finite-dimensional manifolds 4.4. Weighted UAT over infinite-dimensional manifolds Weighted universal approximation of non-anticipative functionals 5.1. Non-anticipative path-neural network 5.2. Weighted UAT for differentiable non-anticipative functionals 5.3. Weighted UAT for horizontal and vertical derivatives
2 4 6 9 10 11 12 14 18 19 20 20 22 27 29 29 31 33 35 37 38 38 39
Date: June 9, 2026. Key words: Machine learning, neural operator, Universal approximation, weighted approximation, infinitedimensional manifold, locally convex topological vector space, bounded approximation property, Stone-Weierstrass theorem, Nachbin theorem, Tauberian theorem, non-anticipative functional, rough path, signature. MSC2020 Subject Classification: 26F16, 41A65, 41A81, 46E15, 58C20, 60L10, 68T07. 1
2
P. SCHMOCKER AND J. TEICHMANN
6.
Weighted universal approximation of linear functions of the signature 6.1. Notation related to the signature of (rough) paths 6.2. Manifold of weakly geometric α-Hölder rough paths 6.3. Weighted universal approximation for differentiable functionals of rough paths 7. Numerical experiments Appendix A. α-Hölder Skorokhod space Dα,r ([0, T ]; Z) Appendix B. BAP of (C α (S; Z), τ∞ ) and (Dα,r ([0, T ]; Z), τw∗ ) Appendix C. Proof of results in Section 2 C.1. Proof of Proposition 2.10 Appendix D. Proof of results in Section 3 D.1. Proof of Lemma 3.6 D.2. Proof of Lemma 3.11 D.3. Proof of Lemma 3.9 Appendix E. Proof of results in Section 4 E.1. Auxiliary lemma for the proof of Theorem 4.11 Appendix F. Proof of results in Section 5 F.1. Proof of Lemma 5.4 F.2. Proof of Corollary 5.5 F.3. Proof of Corollary 5.7 References
40 41 43 47 51 53 56 60 60 62 62 64 65 66 66 67 67 70 72 73
1. Introduction In recent years, machine learning has transformed a wide range of scientific domains with major breakthroughs in image classification [68], speech recognition [51], and computer games [105]. Along these advances, one of the oldest branches of mathematical analysis – approximation theory – has again attracted more attention: Given a target function, can a model class approximate it to arbitrary accuracy? In this paper, we consider functional input neural networks (FNNs) introduced in [28], which extend classical neural networks between Euclidean spaces to infinitedimensional spaces. In particular, we are interested whether such neural networks can also include the approximation of the directional derivatives. This contributes to the rigorous mathematical understanding of supervised machine learning methods in artificial intelligence (see [47,83,84,115]). Neural networks between Euclidean spaces were discovered in the seminal work [80] of W. McCulloch and W. Pitts. They mimic the functionality of a human brain consisting of connections between neurons, i.e., the data is fed into the network, sent along various connections, transformed in the neurons, and then finally returned as output. In mathematical terms, a neural network can be described by a composition of affine and non-linear maps, where the affine maps describe the connections between neurons and the non-linear map the transformation of the data inside a neuron. Neural networks enjoy the so-called universal approximation property, meaning that they can approximate any continuous function uniformly on compact subsets of the Euclidean space. This fundamental result goes back to G. Cybenko and K. Hornik (see [30, 52]) who established in so-called universal approximation theorems (UATs) denseness of the set of neural networks in suitable function spaces. These UATs were then also extended to differentiable functions by taking into account the simultaneous approximation of the derivatives (see [52, 53]). Subsequently, other works [4,15,17] related the approximation error to the network complexity by proving quantitative approximation rates under more restrictive assumptions on the target function.
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
3
The main objective of this article is to generalize the universal approximation theorem (UAT) for functional input neural networks in [28] to differentiable maps, in the sense that not only the values of a given map are approximated but also its directional derivatives. To this end, we extend the weighted framework of [28] by introducing additional weight functions on the higherorder tangent spaces of the input manifold, which in turn requires a slight adaptation of Bastiani calculus [5] to our σ-compact setting. The weights control the functions and their derivatives outside of large compact subsets, which allows us to formulate UATs including the derivatives beyond the usual approximation on compacta. This is relevant for the approximation of stochastic processes as their realizations usually do not stay in a compact path space almost surely. In particular, the weighted setting has been applied as a theoretical framework for generalized Feller processes and their semigroups (see, e.g., [13,29,33,97]), which is important for the approximation of solutions of stochastic (partial) differential equations (see, e.g., [25, 42, 59, 100, 101]). In order to establish the universal approximation property of functional input neural networks (FNNs), we first prove a Nachbin theorem in our weighted setting. The original Nachbin theorem, established by L. Nachbin in [86] over finite-dimensional manifolds, generalizes the classical StoneWeierstrass theorem by including the approximation of the derivatives. This result was later extended by J.B. Prolla and C.S. Guerreiro in [95] as well as R.M. Aron and J.B. Prolla in [3] to infinite-dimensional Banach spaces by using either the compact-open topology or topology of compact convergence, both yielding approximation results over compact subsets of the input space. Only [89] proved a weighted approximation result including the derivatives for polynomials on the Euclidean space. In contrast, our weighted Nachbin theorem is able to approximate a given function and its derivatives globally over an entire chart of an infinite-dimensional manifold. By applying the weighted Nachbin theorem, we can lift the universal approximation theorem (UAT) of neural networks over the real line to a UAT for functional input neural networks (FNNs) defined on infinite-dimensional weighted manifolds. However, even on the real line, the weighted setting requires a global UAT, which is fundamentally different from classical UATs over compact subsets (see, e.g., [20, 30, 52]). Indeed, the weighted UATs in [28, Proposition 4.4 (A3)] and [90, Theorem 2.7] rely on J. Korevaar’s distributional extension [62] of N. Wiener’s Tauberian theorem [118] to obtain sufficient conditions on the Fourier transform of the activation function, ensuring that it is discriminatory for the corresponding linear functionals (in the sense of [30]). To include the approximation of the derivatives, [90, Theorem 2.7] followed [52, 53] and mollified the linear functionals, which allows us to apply integration by parts to eliminate the derivatives from the linear functionals. Finally, we assume the bounded approximation property to lift the UAT from finite-dimensional spaces to infinite-dimensional input and output spaces. Let us remark that there is of course an extensive literature on infinite-dimensional generalizations of neural networks. Early contributions [20, 81, 98, 108] studied the approximation of nonlinear functionals. More recent developments address approximation on non-Euclidean domains [41, 64, 65], on Fréchet spaces [8], and approximation rates for nonlinear functionals on Lp spaces [107]. Moreover, in the setting of adapted maps between suitably defined discrete-time path spaces, echo-state network architectures were shown to be universal [46, 48], while so-called metric hypertransformers were introduced in [1]. In the context of learning solution operators for partial differential equations, we further refer to the works on the deep Galerkin method [106], physics-informed neural networks [96], neural operators [63,73], DeepONets [70,76], and generative equilibrium operators [66], along with the references therein. Apart from neural networks, there are many other families serving as universal approximators on function spaces. By using the weighted Nachbin theorem, we show that linear functions of the signature are able to approximate a given path space functional including its directional derivatives. The signature plays a central role in rough path theory, introduced by T. Lyons
4
P. SCHMOCKER AND J. TEICHMANN
in [78] (see also the textbooks [39,40]), and can be interpreted as polynomials on path space. More precisely, we prove that a path space functional can be approximated with linear functions of the signature on the whole path space, which extends the global universal approximation theorem (UAT) in [28, Theorem 5.4] by including the approximation of the derivatives. The remainder of this article is structured as follows. In Section 2, we introduce weighted domains and manifolds, and characterize maps defined thereon. In Section 3, we prove weighted Nachbin theorems, which are used to show universal approximation theorems for functional input neural networks in Section 4. Subsequently, we apply these weighted approximation results to nonanticipative functionals in Section 5 and to linear functions of the signature in Section 6. Finally, we provide two numerical examples in Section 7. Some proofs are given in Appendices A–F. 1.1. Notation. As usual, we denote by N := {1, 2, 3, ...} and N0 := N ∪ {0} the sets of natural numbers. For n ∈ N, we define Sn as the set of permutations σ : {1, ..., n} → {1, ..., n}, whereas Pn denotes the set of partitions π := {π1 , ..., π|π| } of {1, ..., n}, consisting of disjoint subsets π1 , ..., π|π| ⊆ {1, ..., n} with π1 ∪ ... ∪ π|π| = {1, ..., n}. Moreover, we introduce the set of multiindices as Nd0,n := α := (α1 , ..., αd ) ∈ Nd0 : |α| ≤ n with |α| := α1 + ... + αd . In addition, R √ and C (with imaginary unit i = −1 ∈ C) represent the sets of real and complex numbers, respectively. Furthermore, for d, m ∈ N, we denote by Rd the d-dimensional Euclidean space Pd 2 1/2 equipped with the norm ∥x∥ = , while Rd×m denotes the vector space of matrices i=1 xi Pd Pm j=1,...,m d×m 2 1/2 A := (ai,j )i=1,...,d ∈ R equipped with the Frobenius norm ∥A∥ := . i=1 j=1 |ai,j | Moreover, a topological space (X, τX ) is called Hausdorff if for distinct points x, y ∈ X there exist open sets U, V ∈ τX with x ∈ U and y ∈ V such that U ∩ V = ∅. In addition, a topological vector space (X, τX ) is a vector space X equipped with a topology τX such that addition X × X ∋ (x1 , x2 ) 7→ x1 + x2 ∈ X and scalar multiplication R × X ∋ (λ, x) 7→ λx ∈ X are both continuous. Let us remark that only vector spaces over R are considered in this paper. Furthermore, for topological spaces (X, τX ) and (Y, τY ), we denote by FX := σ(τX ) the Borel σ-algebra of (X, τX ), and define C 0 (X; Y ) as the vector space of continuous maps f : X → Y . In addition, a locally convex topological vector space (X, τX ) is a Hausdorff topological vector space such that τX admits a 0-neighborhood basis consisting of balanced and convex sets. In this case, the topology τX is equivalently generated by a fundamental system of seminorms P(X,τX ) , i.e., by sets of the form {x ∈ X : p(x) < ε}, for ε > 0 and p ∈ P(X,τX ) (see [102, Section II.4]). A seminorm is a map p : X → [0, ∞) such that for every λ ∈ R and x1 , x2 ∈ X it holds that p(λx1 ) = |λ|p(x1 ) and p(x1 + x2 ) ≤ p(x1 ) + p(x2 ), while P(X,τX ) is fundamental if for every p1 , p2 ∈ P(X,τX ) there exist C > 0 and p3 ∈ P(X,τX ) such that for every x ∈ X we have max(p1 (x), p2 (x)) ≤ Cp3 (x). If P(X,τX ) is countable (i.e., (X, τX ) is metrizable) and (X, τX ) is complete, then (X, τX ) is called a Fréchet space. Moreover, if P(X,τX ) consists only of one norm ∥ · ∥X and (X, τX ) is complete, then (X, ∥ · ∥X ) is called a Banach space. In this case, X BrX (x) := {y ∈ X : ∥y − x∥X < r} and B r (x) := {y ∈ X : ∥y − x∥X ≤ r} denote the open and X X closed ball of radius r > 0 around x ∈ X. When x = 0, we set BrX := BrX (0) and B r := B r (0). Furthermore, for a family of locally convex topological vector spaces (Xi , τXi )i∈I , we consider Q the Cartesian product i∈I Xi := {(xi )i∈I : xi ∈ Xi }, which is equipped with the product topolQ ogy i∈I τXi defined as the initial topology with respect to the projections (Q → Xi i∈I Xi (1.1) πi : , i ∈ I, (xi )i∈I 7→ xi Q i.e., the weakest topology on i∈I Xi such that the mappings (1.1) are continuous. Then, Q Q ( i∈I Xi , i∈I τXi ) is again a locally convex topological vector space (see [102, p. 52]). For
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
5
example, if (Xn , ∥ · ∥Xn )n=1,...,N are Banach spaces, with finite N ∈ N, then the product topology QN QN PN Q := n=1 ∥xn ∥Xn . n=1 τXn on n=1 Xn is generated by the norm ∥(xn )n=1,...,N ∥ N n=1 Xn Moreover, for two locally convex topological vector spaces (X, τX ) and (Y, τY ), we denote by L(X; Y ) the vector space of continuous linear maps T : X → Y , which is equipped (unless otherwise specified) with the topology of uniform convergence on bounded subsets of X. In particular, if (X, ∥ · ∥X ) and (Y, ∥ · ∥Y ) are normed vector spaces (resp. Banach spaces), then L(X; Y ) is under the norm ∥T ∥L(X;Y ) := supx∈X, ∥x∥X ≤1 ∥T (x)∥Y again a normed vector space (resp. Banach space). On the other hand, if Y = R, the space X ∗ := L(X; R) is the dual space of (X, τX ) consisting of continuous linear functionals l : X → R. In addition, for two locally convex topological vector spaces (X, τX ) and (Y, τY ), a continuous linear map T : X → Y is called compact if there exists a 0-neighborhood U of (X, τX ) such that T (U ) is relatively compact in (Y, τY ) (see, e.g., [102, p. 98]). If (X, τX ) is a normed vector space (X, ∥ · ∥X ), this is equivalent to the condition that for every ∥ · ∥X -bounded subset B ⊆ X the image T (B) is relatively compact in (Y, τY ). Indeed, the latter implies that T (B1X ) is relatively compact in (Y, τY ). Conversely, if there exists a 0-neighborhood U of (X, ∥ · ∥X ) with BrX ⊆ U , X for some r > 0, then for every ∥ · ∥X -bounded subset B ⊆ BR ⊂ X, with some R > 0, it holds R R R X that T (B) ⊆ T (BR ) ⊆ r T (BrX ) ⊆ r T (U ), where r T (BrX ) is relatively compact in (Y, τY ). Furthermore, a Banach space (X, ∥·∥X ) is called a dual Banach space if there exists an isometric isomorphism I : X → E ∗ into the dual of another Banach space (E, ∥·∥E ), called a predual. Then, the dual pairing E × X ∋ (e, x) 7→ ⟨e, x⟩E×X := I(x)(e) ∈ R is continuous. Hence, X can be equipped with a weak-∗-topology generated by sets of the form {x ∈ X : ⟨e, x⟩E×X ∈ U }, for e ∈ E and U ⊆ R open (see also [28, Appendix A]). Moreover, for locally convex topological vector spaces (X, τX ) and (Y, τY ), we denote by X ∗ ⊗ Y := (X, τX )∗ ⊗ Y := span {X ∋ x 7→ ℓ(x)y ∈ Y : ℓ ∈ X ∗ , y ∈ Y } ⊆ L(X; Y ) the vector subspace of finite rank operators, i.e., continuous linear maps T : X → Y with finite-dimensional range. Then, (X, τX ) is said to have the approximation property (AP) if the identity idX : X → X belongs to the closure of X ∗ ⊗ X with respect to uniform convergence on relatively compact subsets of (X, τX ), i.e., there exists a net (Tγ )γ ⊆ X ∗ ⊗ X approximating the identity idX : X → X uniformly on relatively compact subsets of (X, τX ) (see [102, Section III.9]). Furthermore, we say that (X, τX ) has the (QX -) bounded approximation property (BAP) if there exists a set of seminorms QX generating the topology τQX on X with τQX ⊇ τX such that (X, τX ) has AP with finite rank operators (Tγ )γ ⊆ (X, τX )∗ ⊗ X and for every p ∈ P(X,τX ) there exist q ∈ QX and λ ≥ 0 with p(Tγ (x)) ≤ λq(x) for all γ and x ∈ X. If (X, τX ) is a Banach space, (X, ∥ · ∥X ) has AP if and only if for every ε > 0 and relatively compact subset K ⊂ X there exists some T ∈ X ∗ ⊗ X such that supx∈K ∥x − T (x)∥X < ε. Moreover, (X, ∥ · ∥X ) has BAP if and only if there exists a constant λ ≥ 1 such that (X, ∥ · ∥X ) has AP with finite rank operators (Tγ )γ ⊆ X ∗ ⊗ X satisfying ∥Tγ ∥L(X;X) ≤ λ (see also [74, Section 1.e]). In addition, for U ⊆ X, we say that (U, τX ) has (B)AP if (X, τX ) has (B)AP with net of finite rank operators (Tγ )γ ⊆ X ∗ ⊗ X satisfying Tγ (U ) ⊆ U . In addition, for U ⊆ R open, we denote by Cc∞ (U ; C) the vector space of smooth functions g : U → C with compact support supp(g) := {s ∈ U : g(s) ̸= 0} contained in U . Furthermore, S (R; C) represents the Schwartz space consisting of smooth functions g : R → C with finite seminorms maxj=0,...,n sups∈R (1 + |s|2 )n g (j) (s) , for all n ∈ N0 , which generate the topology of S (R; C). Then, its dual space S ′ (R; C) consists of linear functionals T : S (R; C) → C called |ρ(s)| tempered distributions. For example, ρ ∈ C 0 (R; C) with sups∈R (1+|s| 2 )n < ∞, for some n ∈ N0 , R induces the tempered distribution g 7→ Tρ (g) := R ρ(s)g(s)ds ∈ S ′ (R; C). Moreover, the support of any T ∈ S ′ (R; C) is defined as the complement of the largest open set U ⊆ R on which T ∈ S ′ (R; C) vanishes, i.e., T (g) = 0 for all g ∈ Cc∞ (U ; C). In addition, the Fourier transform
6
P. SCHMOCKER AND J. TEICHMANN
R of any g ∈ L1 (R; C) is defined as R ∋ ξ 7→ gb(ξ) = √12π R e−iξs g(s)ds ∈ C, while the Fourier transform of any T ∈ S ′ (R; C) is defined by g 7→ Tb(g) := T (b g ) ∈ S ′ (R; C). For more details, we refer to [37, Chapter 7 and 9]. In addition, if the functions are real-valued, we use the abbreviations C k (U ) := C k (U ; R), ∞ Cc (U ) := Cc∞ (U ; R), Lp (Ω) := Lp (Ω; R), C α (S) := C α (S; R), Dα,r ([0, T ]) := Dα,r ([0, T ]; R), etc. Most of the function spaces are introduced in the following sections. 1.2. Bastiani calculus on σ-compact spaces. In this section, we first recall the notion of Bastiani calculus (also known as Keller’s Cc1 -theory, see [5, 58] and also [9, 43, 103, 116]) and then introduce a slight generalization onto σ-compact spaces. For an open subset U ⊆ X of a locally convex topological vector space (X, τX ) as input space and a locally convex topological vector space (Y, τY ) as output space, we define the directional derivative of a map f : U → Y at the point u ∈ U in direction v ∈ X (if it exists) as f (u + hv) − f (u) . h→0 h For k ≥ 2, we define the k-th order directional derivatives of a map f : U → Y at the point u ∈ U in directions v1 , ..., vk ∈ X (if they exist) as
(1.2)
(1.3)
df (u; v) := Dv f (u) := lim
dj f (u; v1 , ..., vj ) := Dvj · · · Dv1 f (u).
Then, for k ∈ N0 , the C k -space C k (U ; Y ) in the sense of Bastiani is defined as the vector space of maps f : U → Y whose j-th order directional derivatives exist and the mappings U × X j ∋ (u, v1 , ..., vj ) 7→ dj f (u; v1 , ..., vj ) ∈ Y are continuous, for all j = 0, ..., k, with d0 f := f . k (U ; Y ) as the vector space of Moreover, if the input space (X, τX ) is σ-compact, we define Cloc maps f : U → Y whose j-th order directional derivatives exist and the mappings dj f |K∩(U ×X j ) : K ∩ (U × X j ) → Y are continuous, for any compact subset K ⊆ X × X j and j = 0, ..., k. Compared to Bastiani calculus with globally continuous mappings dj f : U × X j → Y , j = 0, ..., k, k we only require them to be continuous on compact subsets, implying that C k (U ; Y ) ⊆ Cloc (U ; Y ). However, if (X, τX ) is locally compact or first countable (ensuring that (X, τX ) is compactly generated, see [85, Lemma 46.3]), every mapping U × X j ∋ (u, v1 , ..., vj ) 7→ dj f (u; v1 , ..., vk ) ∈ Y that is continuous on compacta, is also globally continuous (see [85, Lemma 46.4]), whence k the two notions are equivalent. Therefore, our notion of Cloc -maps is stronger than Gâteaux differentiability (except on finite-dimensional spaces), but weaker than Fréchet differentiability (except on finite-dimensional spaces). k In order to establish some properties of our Cloc -differential calculus that are known for Bastiani calculus (see, e.g., [5, 43, 58, 103]), we first prove the following auxiliary lemmas. For an open Rb interval I ⊆ R, a, b ∈ R, and f ∈ C 0 (I; Y ), we say that the weak integral a f (t)dt exists if there Rb is a point y ∈ Y such that for every ℓ ∈ Y ∗ it holds that ℓ(y) = a ℓ(f (t))dt. Lemma 1.1 (Fundamental theorem of calculus). Let 0 ∈ I ⊆ R be an open interval and let Rh 1 c ∈ Cloc (I; Y ). Then, for every h ∈ I, the weak integral 0 c′ (t)dt exists and satisfies Z h c(h) − c(0) = c′ (t)dt. 0 1 Proof. Since c ∈ Cloc (I; Y ), we have c′ ∈ C 0 (I; Y ). Hence, we can apply the fundamental theorem of calculus for real-valued functions to conclude for every ℓ ∈ Y ∗ and h ∈ I that Z h Z h ℓ(c(h) − c(0)) = ℓ(c(h)) − ℓ(c(0)) = (ℓ ◦ c)′ (t)dt = ℓ(c′ (t))dt. 0
0
Hence, y := c(h) − c(0) ∈ Y satisfies the defining properties of the weak integral.
□
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
7
1 Note that this fundamental theorem of calculus for Cloc -curves with values in a locally convex topological vector space (Y, τY ) holds irrespective of completeness of Y . Moreover, as an application of the bipolar theorem (see, e.g., [102, Theorem IV.1.5]), we obtain the following result. Rb Lemma 1.2. Let a, b ∈ R and f ∈ C 0 ([a, b]; Y ) such that the weak integral a f (t)dt exists. Then, for every pY ∈ P(Y,τY ) , it holds that ! Z b pY f (t)dt ≤ |b − a| sup pY (f (t)). t∈[a,b]
a
k Next, we prove the following properties of our Cloc -differential calculus including the linearity of
the differential and the chain rule, which are known for Bastiani calculus (see, e.g., [5,43,58,103]). 1 Proposition 1.3. Let f ∈ Cloc (U ; Y ). Then, the following holds true: 0 (i) For every u ∈ U the map X ∋ v 7→ df (u; v) ∈ Y is linear and in Cloc (X; Y ). (ii) Let V ⊆ Y be open with f (U ) ⊆ V , let (Y, τY ) be σ-compact, let (Z, τZ ) be another locally 1 1 convex topological vector space, and let g ∈ Cloc (V ; Z). Then, g ◦ f ∈ Cloc (U ; Z) and for every u ∈ U and v ∈ X we have d(g ◦ f )(u; v) = dg(f (u); df (u; v)). k (iii) If f ∈ Cloc (U ; Y ) with k ≥ 2, then for every u ∈ U and v1 , ..., vk ∈ X it holds that k−1 d d f (·; v1 , ..., vk−1 ) (u; vk ) = dk f (u; v1 , ..., vk ). k (iv) If f ∈ Cloc (U ; Y ), then for every u ∈ U , v1 , ..., vk ∈ X, and every permutation σ ∈ Sk we k have d f (u; vσ(1) , ..., vσ(k) ) = dk f (u; v1 , ..., vk ).
Proof. For (i) we fix some u ∈ U . Then, for every v ∈ X and λ ∈ R, the homogeneity df (u; λv) = λdf (u; v) follows from (1.2). For linearity of X ∋ v 7→ df (u; v) ∈ Y , we fix some ε > 0, pY ∈ P(Y,τY ) , u ∈ U , v1 , v2 ∈ X, and δ > 0 such that u + rv1 + sv2 ∈ U for all r, s ∈ [−δ, δ]. Then, by applying Lemma 1.1 twice, it follows for every h ∈ (−δ, δ) that (1.4) Z 1 f (u + h(v1 + v2 )) − f (u) = f (u + hv1 ) − f (u) + df (u + hv1 + shv2 ; hv2 )ds 0
Z 1 =
Z 1 df (u + shv1 ; hv1 )ds +
0
df (u + hv1 + shv2 ; hv2 )ds 0
Z 1 (df (u + shv1 ; hv1 ) − df (u; hv1 )) ds
= h(df (u; v1 ) + df (u; v2 )) + 0
Z 1 (df (u + hv1 + shv2 ; hv2 ) − df (u; hv2 )) ds,
+ 0
1 where all integrals exist as weak integrals. Moreover, by using that f ∈ Cloc (U ; Y ) and the image of [−δ, δ] ∋ s 7→ u + sv1 ∈ U is compact in U , we conclude that [−δ, δ] ∋ s 7→ df (u + sv1 ; v1 ) − df (u; v1 ) ∈ Y is continuous, thus uniformly continuous, whence there exists some δ1 ∈ (0, δ) such that for every h ∈ (−δ1 , δ1 ) it holds that Z 1 1 (df (u + shv1 ; hv1 ) − df (u; hv1 )) ds pY h 0 Z 1 (1.5) = pY (df (u + shv1 ; v1 ) − df (u; v1 )) ds 0
≤ sup pY (df (u + shv1 ; v1 ) − df (u; v1 )) < s∈[0,1]
ε , 2
where we have applied Lemma 1.2 for the first inequality. Similarly, by using that [−δ, δ]2 ∋ (r, s) 7→ df (u + rv1 + sv2 ; v1 ) − df (u; v1 ) ∈ Y is continuous, thus uniformly continuous, there exists
8
P. SCHMOCKER AND J. TEICHMANN
some δ2 ∈ (0, δ) such that for every h ∈ (−δ2 , δ2 ) we have Z 1 1 pY (df (u + hv1 + shv2 ; hv2 ) − df (u; hv2 )) ds h 0 Z 1 (1.6) (df (u + hv1 + shv2 ; v2 ) − df (u; v2 )) ds = pY 0
≤ sup pY (df (u + hv1 + shv2 ; v2 ) − df (u; v2 )) < s∈[0,1]
ε . 2
Hence, by inserting (1.5)–(1.6) into (1.4) and defining δ0 := min(δ1 , δ2 ) > 0, it follows for every h ∈ (−δ0 , δ0 ) that f (u + h(v1 + v2 )) − f (u) pY − (df (u; v1 ) + df (u; v2 )) h Z 1 1 (df (u + shv1 ; hv1 ) − df (u; hv1 )) ds ≤ pY h 0 Z 1 1 + pY (df (u + hv1 + shv2 ; hv2 ) − df (u; hv2 )) ds h 0 ε ε < + = ε. 2 2 Since ε > 0 was chosen arbitrarily, this and the homogeneity show that X ∋ v 7→ df (u; v) ∈ Y is 1 0 linear. Finally, we use that f ∈ Cloc (U ; Y ) to see that (v 7→ df (u; v)) ∈ Cloc (X; Y ). For (ii), we fix some u ∈ U , v ∈ X, and δ > 0 such that u + sv ∈ U for all s ∈ [−δ, δ]. Then, by using that df (u; v) exists and is continuous on compacta, there exists a continuous function r : [−δ, δ] → Y with r(0) = 0 such that for every h ∈ [−δ, δ] we have (g ◦ f )(u + hv) = g (f (u) + h (df (u; v) + r(h))) . Defining w(h) := df (u; v) + r(h) and shrinking δ > 0 if necessary, we can assume that f (u) + shw(h) ∈ V for all h ∈ [−δ, δ] and s ∈ I, where I is an open interval containing [0, 1]. Moreover, 1 1 (U ; Y ) and g ∈ Cloc (V ; Z), the curve I ∋ s 7→ c(s) := g(f (u) + hsw(h)) ∈ Z is since f ∈ Cloc continuously differentiable with c′ (s) = dg(f (u) + hsw(h); hw(h)) = hdg(f (u) + hsw(h); w(h)), whence Lemma 1.1 implies for every h ∈ [−δ, δ] that (1.7) Z 1 (g ◦ f )(u + hv) = c(1) = c(0) + c′ (s)ds 0
= g(f (u)) + h dg(f (u); df (u; v)) Z 1 +h dg(f (u) + hsw(h); df (u; v) + r(h)) − dg(f (u); df (u; v)) ds 0
= g(f (u)) + h dg(f (u); df (u; v)) Z 1 +h dg(f (u) + hsw(h); df (u; v)) + dg(f (u) + hsw(h); r(h)) − dg(f (u); df (u; v)) ds, 0
where all integrals exist as weak integrals. Now, for every fixed ε > 0 and pZ ∈ P(Z,τZ ) , we use 1 1 that f ∈ Cloc (U ; Y ) and g ∈ Cloc (V ; Z) to conclude that [−δ, δ] × [0, 1] ∋ (h, s) 7→ Φ(h, s) := dg(f (u) + hsw(h); r(h)) ∈ Z [−δ, δ] × [0, 1] ∋ (h, s) 7→ Ψ(h, s) := dg(f (u) + hsw(h); df (u; v)) − dg(f (u); df (u; v)) ∈ Z are continuous, thus uniformly continuous, with Φ(0, s) = Ψ(0, s) = 0 for all s ∈ [0, 1]. Hence, there exists some δ0 ∈ (0, δ) such that for every (h, s) ∈ (−δ0 , δ0 )×[0, 1] it holds that pZ (Φ(h, s)) <
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
9
ε/2 and pZ (Ψ(h, s)) < ε/2, which implies by Lemma 1.2 for every h ∈ (−δ0 , δ0 ) that Z 1 (dg(f (u) + hsw(h); df (u; v)) + dg(f (u) + hsw(h); r(h)) − dg(f (u); df (u; v))) ds pZ 0 (1.8) ε ε ≤ sup pZ (Φ(h, s) + Ψ(h, s)) < + = ε. 2 2 (s,h)∈[0,1]×(−δ0 ,δ0 ) Hence, by inserting (1.8) into (1.7), it follows for every h ∈ (−δ0 , δ0 ) that (g ◦ f )(u + hv) − g(f (u)) pZ − dg(f (u); df (u; v)) < ε. h Since ε > 0 was chosen arbitrarily, this shows that d(g ◦ f )(u; v) = dg(f (u); df (u; v)). While (iii) follows from the definition (1.3), we fix for (iv) some δ > 0 such that u + h1 v1 + ... + hk vk ∈ U for all h1 , ..., hk ∈ [−δ, δ]. Then, by using that the mapping [−δ, δ]k ∋ (h1 , ..., hk )
7→
F (h1 , ..., hk ) := f (u + h1 v1 + ... + hk vk ) ∈ Y
is k-times continuously differentiable on (−δ, δ)k , we can apply the finite-dimensional Schwarz theorem on (−δ, δ)k ⊆ Rk to conclude that dk f (u; vσ(1) , ..., vσ(k) ) =
dk F dk F (0, ..., 0) = (0, ..., 0) = dk f (u; v1 , ..., vk ), dhσ(1) · · · dhσ(k) dh1 · · · dhk □
which completes the proof.
1.3. Manifolds over σ-compact model spaces. In this section, we introduce the notion of manifolds that are modelled over σ-compact locally convex topological vector spaces (see also [9, 67, 103] for more details). To this end, we shall fix some k ∈ N ∪ {∞}, a topological space (M, τM ), and a family of σ-compact locally convex topological vector spaces (Xi , τXi )i∈I , where k I is an arbitrary index set. Then, a Cloc -atlas (Ui , ϕi )i∈I for M consists of an open cover (Ui )i∈I S of M , i.e., i∈I Ui = M , and homeomorphisms ϕi : Ui → ϕi (Ui ) ⊆ Xi called charts such that k the transition maps ϕi1 ◦ ϕi2 |−1 ϕi (Ui ∩Ui ) : ϕi1 (Ui1 ∩ Ui2 ) → ϕi2 (Ui1 ∩ Ui2 ) are Cloc -maps, for all 2
1
2
k k i1 , i2 ∈ I. If such a Cloc -atlas (Ui , ϕi )i∈I exists for M , then we call (M, τM ) a Cloc -manifold (with atlas (Ui , ϕi )i∈I over model spaces (Xi , τXi )i∈I ). For example, if U ⊆ X is an open subset of a ∞ σ-compact locally convex topological vector space (X, τX ), then M := U is a Cloc -manifold with global chart given by the smooth inclusion U ,→ X. Moreover, we follow [82, 111] and define for every j = 0, ..., k the tangent space of order j at x ∈ M as set Txj M of equivalence classes [c]jx of C j -curves c : (−ε, ε) → M with c(0) = x whose accelerations agree up to order j, i.e., c ∼j e c if and only if c(ℓ) (0) = e c(ℓ) (0), for all ℓ = 0, ..., j, 0 j where Tx M := {0}. If x ∈ Ui , then Tx M is topologically isomorphic to Xij with isomorphism ( Txj M → Xij j Φi,x : [c]jx 7→ (ϕi ◦ c)(ℓ) (0) ℓ=1,...,j
and inverse Φ−j i,x :
Xij
→
(v1 , ..., vj )
7→
Txj M h ij . j t 7→ ϕ−1 ϕi (x) + 1!t v1 + ... + tj! vj i x
S Furthermore, we define for every j = 0, ..., k the tangent bundle of order j as T j M := x∈M Txj M := (x, [c]jx ) : x ∈ M, [c]jx ∈ Txj M , with T 0 M := M , which we equip with the final topology τT j M
10
P. SCHMOCKER AND J. TEICHMANN
with respect to the family of mappings ϕi (Ui ) × X j → i −j (1.9) Φi : (u, v1 , ..., vj ) 7→
T jM , −j ϕ−1 i (u), Φi,ϕ−1 (u) (v1 , ..., vj )
i ∈ I,
i
i.e., the finest topology on T j M such that the mappings (1.9) are continuous. Then, by following the proof of [111, Theorem 2.1] (with manifolds over locally convex topological vector spaces 0 instead of Banach manifolds), one can show that T j M is as fibre bundle a Cloc -manifold with j j j −1 atlas (πT j M (Ui ), Φi )i∈I over model spaces (Xi × Xi , τXi × τXi )i∈I , where ( −1 j πT j M (Ui ) → ϕ j , i (Ui ) × Xi i ∈ I, (1.10) Φi : (x, [c]jx ) 7→ ϕi (x), Φji,x ([c]jx ) are the charts, and where T j M ∋ (x, [c]jx ) 7→ πT j M (x, [c]jx ) := x ∈ M is the bundle projection. For example, if M := U ⊆ X is an open subset of a locally convex topological vector space (X, τX ), then it holds that T j M ∼ = U × X j , for all j ∈ N0 . k k In addition, for k ∈ N0 and a given Cloc -manifold (M, τM ), we denote by Cloc (M ; Y ) the vector −1 k space of maps f : M → Y such that f ◦ ϕi ∈ Cloc (ϕi (Ui ); Y ) for all i ∈ I. 1.4. Examples of σ-compact model spaces. In this section, we present some examples of σ-compact locally convex topological vector spaces used as model spaces for manifolds. For α ∈ (0, ∞), a compact metric space (S, dS ) with designated origin 0 ∈ S, and a dual Banach space (Z, ∥ · ∥Z ) with predual (E, ∥ · ∥E ), we denote by C α (S; Z) the space of α-Hölder continuous functions x : (S, dS ) → (Z, ∥ · ∥Z ) satisfying ∥x∥α := ∥x(0)∥Z + |x|α < ∞. Here, |x|α denotes the α-Hölder seminorm of x : S → Z defined as |x|α :=
∥x(s) − x(t)∥Z . dS (s, t)α s,t∈S, s̸=t sup
Then, the norm ∥ · ∥α turns C α (S; Z) into a Banach space (see [40, Theorem 5.25] and [117, ′ Proposition 2.3(b)]). Moreover, for α′ ∈ [0, α], we equip C α (S; Z) also with the weaker w∗ -C α topology τα′ generated by seminorms of the form ∥x∥α′ ,e := |⟨x(0), e⟩Z×E | + |x|α′ ,e , for e ∈ E, where |x|α′ ,e denotes the (α′ , e)-Hölder seminorm of x : S → Z defined as |x|α′ ,e := sup s,t∈S s̸=t
|⟨x(s) − x(t), e⟩Z×E | . dS (s, t)α′
Hence, (C α (S; Z), τα′ ) forms a locally convex topological vector space. Note that for α′ = 0 the w∗ -C 0 -topology τ0 is equivalent to the w∗ -uniform topology τ∞ generated by seminorms of the form ∥x∥∞,e := supt∈S |⟨x(t), e⟩Z×E |, for e ∈ E (see [28, Lemma A.1]). Moreover, for α′ ∈ [0, α), the embedding (C α (S; Z), ∥ · ∥α ) ,→ (C α (S; Z), τα′ ) is by [28, Theorem A.4] compact, whence (C α (S; Z), τα′ ) is as the image of countable many ∥ · ∥α -balls σ-compact. In addition, C α (S; Z) is a dual Banach space (see [28, Theorem A.5]), which is by the Banach-Alaoglu theorem also σ-compact with respect to its weak-∗-topology τw∗ . Furthermore, we denote by C0α (S; Z) ⊆ C α (S; Z) the vector subspace of α-Hölder continuous functions x ∈ C α (S; Z) with x(0) = 0 ∈ Z. Moreover, for T > 0 and a dual Banach space (Z, ∥ · ∥Z ), we denote by D0 ([0, T ]; Z) the vector space of càdlàg paths x : [0, T ] → (Z, ∥ · ∥Z ), whose left limits x(t−) := lims→t− x(s) exist, for all t ∈ (0, T ] and the right limits satisfy x(t+) := lims→t+ x(s) = x(t), for all t ∈ [0, T ).
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
11
Then, the norm ∥x∥∞ := supt∈[0,T ] ∥x(t)∥Z turns D0 ([0, T ]; Z) into a Banach space (see, e.g., [12, Section 12], [35, Section 3.5], and [56, p. 1]). In addition, for (α, r) ∈ [0, 1) × [1, ∞), we define Dα,r ([0, T ]; Z) ⊆ D0 ([0, T ]; Z) as the vector subspace of càdlàg paths x ∈ D0 ([0, T ]; Z) satisfying ∥x∥α,ℓr := max (∥xc ∥α , ∥∆x∥ℓr ) < ∞, P where the continuous part [0, T ] ∋ t 7→ xc (t) := x(t) − s∈(0,t] ∆x(s) ∈ Z is α-Hölder continuous and the jump part (0, T ] ∋ t 7→ ∆x(t) := x(t) − x(t−) ∈ Z is r-summable, i.e. r1 X (1.11) ∥∆x∥ℓr := ∥∆x(t)∥rZ < ∞. t∈(0,T ]
Since every càdlàg path has at most countably many jumps (see [35, Lemma 5.1]), the condition (1.11) is only an assumption on the jump sizes. Then, (Dα,r ([0, T ]; Z), ∥ · ∥α,ℓr ) is a Banach space, which is isometrically isomorphic to the direct sum of the Banach spaces (C α ([0, T ]; Z), ∥ · ∥α ) and (ℓr ((0, T ]; Z), ∥ · ∥ℓr ), where the latter consists of Z-valued sequences (zt )t∈(0,T ] with P r 1/r ∥z∥ℓr := < ∞ (see Theorem A.1). Moreover, if (α, r) ∈ (0, 1) × (1, ∞), then t∈(0,T ] ∥zt ∥Z (Dα,r ([0, T ]; Z), ∥ · ∥α,ℓr ) is a dual Banach space (see Theorem A.2), which is by the BanachAlaoglu theorem also σ-compact with respect to its weak-∗-topology τw∗ . Note that τw∗ coincides on ∥ · ∥α,ℓr -bounded subsets of C α ([0, T ]; Z) ⊆ Dα,r ([0, T ]; Z) with the w∗ -uniform topology τ∞ . In addition, for p ∈ [1, ∞], a σ-finite measure space (Ω, F, µ), and a dual Banach space (Z, ∥·∥Z ) with predual (E, ∥·∥E ), we denote by Lp (Ω; Z) := Lp (Ω, F, µ; Z) the Bochner space of (equivalence classes of) strongly µ-measurable maps x : Ω → Z with finite norm ( R 1/p ∥x(ω)∥pZ µ(dω) , p ∈ [1, ∞), Ω ∥x∥Lp (Ω;Z) := inf {c > 0 : µ ({ω ∈ Ω : ∥x(ω)∥Z > c}) = 0} , p = ∞. Then, the norm ∥ · ∥Lp (Ω;Z) turns Lp (Ω; Z) into a Banach space (see [54, Section 1.2b]). In particular, for p ∈ (1, ∞] and p′ ∈ [1, ∞) with 1/p + 1/p′ = 1, and if (Z, ∥ · ∥Z ) has the RadonNikodym property with respect to µ (see [54, Definition 1.3.9]), then the Bochner space Lp (Ω; Z) ∼ = ′ Lp (Ω; E)∗ is a dual Banach space, which is by the Banach-Alaoglu theorem σ-compact with respect to its weak-∗-topology τw∗ . Furthermore, for a weighted space (Ω, ψΩ ) (see [28, Definition 2.1]), we denote by MψΩ (Ω) the R vector space of signed Radon measures x : FΩ → R with Ω ψΩ (ω)|x|(dω) < ∞. Then, Z ∥x∥MψΩ (Ω) := sup f (ω)x(dω) : f ∈ BψΩ (Ω), ∥f ∥BψΩ (Ω) ≤ 1 Ω
turns MψΩ (Ω) into a Banach space, where the weighted function space BψΩ (Ω) is defined in [28, Definition 2.5]. Then, MψΩ (Ω) = BψΩ (Ω)∗ is by the Riesz representation theorem in [33, Theorem 2.8] a dual Banach space, which is by the Banach-Alaoglu theorem σ-compact with respect to its weak-∗-topology τw∗ . 2. Weighted spaces and differentiable maps For the approximation results on infinite-dimensional manifolds, we endow the input space with a weight function and assume that the output space is a Banach space. This weighted setting is in particular inspired by the works on Kolmogorov equations, splitting schemes of (stochastic) partial differential equations, and generalized Feller processes (see, e.g., [29, 33, 97]). In the following, we first introduce our weighted setting on domains given as open subsets of locally convex topological vector spaces, followed by weighted infinite-dimensional manifolds.
12
P. SCHMOCKER AND J. TEICHMANN
k Later on, we introduce the weighted BΨ -function space that was under slightly different conditions also studied in [10, 11, 28, 87, 89, 93, 110, 112–114].
2.1. Weighted domains. In the following, we shall fix some k ∈ N and consider an open subset U ⊆ X of a locally convex topological vector space (X, τX ). We refer to Section 1.1 for the mathematical background of locally convex topological vector spaces. Definition 2.1. A collection Ψ := (ψj )j=0,...,k of weight functions ψj : U × X j → (0, ∞) is called admissible (on U ) if (i) for every j = 0, ..., k and R > 0 the pre-image Kj,R := ψj−1 ((0, R]) := (u, v1 , ..., vj ) ∈ U × X j : ψj (u, v1 , ..., vj ) ≤ R j is compact with respect to τX × τX , and (ii) Ψ := (ψj )j=0,...,k is monotone, i.e., there exists a constant CΨ ≥ 1 such that for every j = 1, ..., k, ℓ = 1, ..., j, σ ∈ Sj , and (u, v1 , ..., vj ) ∈ U × X j it holds that
ψℓ (u, vσ(1) , ..., vσ(ℓ) )ψj−ℓ (u, vσ(ℓ+1) , ..., vσ(j) ) ≤ CΨ ψj (u, v1 , ..., vj ). In this case, we call (U, Ψ) a weighted domain. Remark 2.2. If (U, Ψ) is a weighted domain, then the following holds true: (i) The weight functions ψj : U × X j → (0, ∞) are necessarily lower semicontinuous and bounded from below by a strictly positive constant (see [28, Remark 2.2 (i)]). S (ii) The domain U is σ-compact with respect to τX as U = R∈N K0,R . (iii) The locally convex topological vector space (X, τX ) is also σ-compact because of X = S R∈N π1 (K1,R ), where π1 (K1,R ) is compact as continuous image of the compact set K1,R under the projection U × X ∋ (u, v1 ) 7→ π1 (u, v1 ) := v1 ∈ X. (iv) If (X, τX ) is complete, then X is finite-dimensional. Indeed, this follows from Baire’s category theorem and (iii) (see also [28, Remark 2.2 (iii)]). Hence, for an infinite-dimensional domain U , we need to consider an incomplete locally convex topological vector space (X, τX ) instead of a Banach space or a Fréchet space. (v) If U = X is separable and ψj : X × X j → (0, ∞) is convex, then ψj : X × X j → (0, ∞) is continuous on a convex subset E ⊆ X × X j if and only if E is locally compact (see also [29, Remark 2.2]). In the following, we present various examples of weighted domains (U, Ψ), where U ∈ τX is a subset of a Banach space (X, ∥ · ∥X ) that is equipped with a weaker topology τX than the norm topology (except X is finite-dimensional). Lemma 2.3. Let U ⊆ X be an open subset of one of the following two types of locally convex topological vector spaces (X, τX ): (i) (X, ∥ · ∥X ) is a Banach space equipped with the initial topology τX := τinit of a compact embedding Γ : (X, ∥ · ∥X ) ,→ (X0 , τ0 ) into another locally convex topological vector space X (X0 , τ0 ) such that B r is closed with respect to τinit , for all r > 0. Here, τinit is the weakest locally convex topology on X such that Γ : (X, τinit ) ,→ (X0 , τ0 ) is continuous. (ii) (X, ∥ · ∥X ) is a dual Banach space equipped with the weak-∗-topology τX := τw∗ . Moreover, let Ψ := (ψj )j=0,...,k be a collection of weight functions of the form (2.1) ! j X j −1 U ×X ∋ (u, v1 , ..., vj ) 7→ ψj (u, v1 , ..., vj ) = η max(j, 1) δU c (u) +∥u∥X + ∥vℓ ∥X ∈ (0, ∞), ℓ=1
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
13
−1 where (U, τinit ) ∋ u 7→ δU c (u) := 1/ inf v∈X\U ∥u−v∥X ∈ [0, ∞) is assumed to be lower semicontinuous, and where η : [0, ∞) → (0, ∞) is a continuous and increasing function with limr→∞ η(r) = ∞. Then, (U, Ψ) is a weighted domain.
Proof. For (i), we fix some j = 0, ..., k, R > 0, and consider the pre-image Kj,R := ψj−1 ((0, R]). Then, for every (u, v1 , ..., vj ) ∈ Kj,R , it holds that j X −1 max(j, 1) δU c (u)−1 + ∥u∥X + ∥vℓ ∥X ≤ ηR := sup{r ≥ 0 : η(r) ≤ R} < ∞, ℓ=1 X which ensures that Kj,R ⊆ B ηR−1 × j
X j B ηR−1 .
j Since the product topology τinit × τinit on X × X coincides with the initial topology induced by the mapping X × X j ∋ (u, v1 , ..., vj ) 7→ (Γ(u), Γ(v1 ), ..., Γ(vj )) ∈ (X0 × X0j , τ0 × τ0j ), it follows that Kj,R is a relatively compact subj j , ). In order to show that Kj,R is also closed with respect to τinit × τinit set of (U × X j , τinit × τinit (γ) (γ) we fix a net u(γ) , v1 , ..., vj γ ⊆ Kj,R converging to some (u, v1 , ..., vj ) ∈ U × X j with respect j to τinit × τinit . Then, by using that δU c : (U, τinit ) → [0, ∞) as well as ∥ · ∥X : (X, τinit ) → [0, ∞) are lower semicontinuous and that η : R → R is continuous, we conclude that ! j X −1 ψj (u, v1 , ..., vj ) = η max(j, 1) δU c (u) + ∥u∥X + ∥vℓ ∥X ℓ=1
≤ lim inf η max(j, 1) δU c (u(γ) )−1 + ∥u(γ) ∥X +
γ
! (γ) ∥vℓ ∥X
ℓ=1
(γ)
= lim inf ψj u γ
j X
(γ) (γ) , v1 , ..., vj
≤ R,
j which shows that Kj,R is closed and therefore compact with respect to τinit × τinit . Moreover, For (ii), let (E, ∥ · ∥E ) be a predual for (X, ∥ · ∥X ). Then, for every fixed j = 0, ..., k, we use that (X × X j , ∥ · ∥X×X j ) is a dual Banach space with predual (E × E j , ∥ · ∥E×E j ), whose weak∗-topology coincides with the product topology τw∗ × τwj ∗ on X × X j . Thus, for every R > 0, we use that Kj,R := ψj−1 ((0, R]) is bounded with respect to ∥ · ∥X×X j to conclude that Kj,R is by the Banach-Alaoglu theorem a compact subset of (U × X j , τw∗ × τwj ∗ ). □
Remark 2.4. By [57, Theorem 1] of S. Kaijser (formally generalizing the Dixmier-Ng theorem in [32, 91]), a Banach space (X, ∥ · ∥X ) is a dual Banach space if there exists a point separating X subset L ⊆ X ∗ such that the unit ball B 1 is compact with respect to the weak topology on X induced by L ⊆ X ∗ . Thus, a compactly embedded Banach space (X, ∥ · ∥X ) as in (i) can be turned into a dual Banach space (see also [28, Appendix A]). In the following, we give some examples of weighted domains (U, Ψ). We refer to Section 1.4 for the precise definition of some of the vector spaces that appear below. Example 2.5. The following are examples of weighted domains (U, Ψ), where U ⊆ X is an open subset of a locally convex topological vector space (X, τX ), and where Ψ := (ψj )j=0,...,k is a collection of weight functions of the form (2.1). (i) First, we consider an open subset U ∈ τinit of a Banach space (X, ∥ · ∥X ) equipped with the initial topology τinit of a compact embedding as in Lemma 2.3 (i): (a) Euclidean space X := Rd with τinit generated by the Euclidean norm ∥ · ∥. (b) α-Hölder space X := C α (S; Z) with τinit generated by the compact embedding (C α (S; Z), ∥ · ∥α ) ,→ (C α (S; Z), τα′ ) (see [28, Theorem A.4]), where 0 ≤ α′ < α < 1, (S, dS ) is a compact metric space, and (Z, ∥ · ∥Z ) is a dual Banach space.
14
P. SCHMOCKER AND J. TEICHMANN
(c) Sobolev space X := W 1,p (Ω) with τinit induced by the compact embedding (W 1,p (Ω), ∥· ∥W 1,p (Ω) ) ,→ (W 1,p (Ω), ∥ · ∥Lq (Ω) ) (see [16, Theorem 9.16]), where p ∈ [1, d), q ∈ dp ), and Ω ⊂ Rd is an open bounded Lipschitz domain. [1, d−p s s (d) Besov space X := Bp,q (Ω) with τinit generated by the compact embedding (Bp,q (Ω), ∥· ′ s d s (Ω) ) ,→ (B ′ ′ (Ω), ∥ · ∥ s′ ∥Bp,q ) (see [113, Theorem 1.97]), where Ω ⊂ R is open p ,q B ′ ′ (Ω) p ,q
and bounded, p, q ∈ (1, ∞] (with dual exponents p′ , q ′ ∈ [1, ∞)), and s′ ∈ (−∞, s) with s − dp > s′ − pd′ . (ii) Second, we consider an open subset U ∈ τw∗ of a dual Banach space (X, ∥ · ∥X ) equipped with the weak-∗-topology τw∗ as in Lemma 2.3 (ii): (e) Euclidean space X := Rd ∼ = (Rd )∗ is a dual Banach space. α (f ) α-Hölder space X := C (S; Z) is a dual Banach space (see [28, Theorem A.4]), where α ∈ (0, 1], (S, dS ) is a compact metric space, and (Z, ∥ · ∥Z ) is a dual Banach space. (g) α-Hölder Skorokhod space X := Dα,r ([0, T ]; Z) is a dual Banach space (see Theorem A.2), where (α, r) ∈ (0, 1) × (1, ∞), T > 0, and (Z, ∥ · ∥Z ) is a dual Banach space. ′ (h) Lp -space X := Lp (Ω; Z) ∼ = Lp (Ω; E)∗ is a dual Banach space (see [54, Theorem 1.3.10]), where p ∈ (1, ∞] (with dual exponent p′ ∈ [1, ∞)), (Ω, F, µ) is a σ-finite measure space, (Z, ∥·∥Z ) is a dual Banach space having the Radon-Nikodym property with respect to µ (see [54, Definition 1.3.9]), and (E, ∥ · ∥E ) is a predual for (Z, ∥ · ∥Z ). (i) Space X := BV (Ω) of integrable functions x : Ω → R with bounded variation is a dual Banach space (see [2, Remark 3.12]), where Ω ⊆ Rd is an open subset. s d ∗ (j) Besov space X := Bp,q (Rd ) ∼ is a dual Banach space (see [114, Theo= Bp−s ′ ,q ′ (R ) rem 2.11.2 (i)]), where p, q ∈ (1, ∞] (with dual exponents p′ , q ′ ∈ [1, ∞)) and s ∈ R. (k) Weighted measure space X := MψΩ (Ω) ∼ = BψΩ (Ω)∗ is a dual Banach space (see [33, Theorem 2.4]), where (Ω, ψΩ ) is a weighted space in the sense of [28, Definition 2.1] and BψΩ (Ω) consists of weighted functions defined on (Ω, ψΩ ) (see [28, Definition 2.5] for the precise definition). k 2.2. BΨ -maps over weighted domains. In this section, we introduce differentiable maps on a weighted domain (U, Ψ) taking values in a Banach space (Y, ∥ · ∥Y ). Let U ⊆ X be an open subset of a σ-compact locally convex topological vector space (X, τX ). Moreover, for a given set QX of seminorms on X, we assume that the admissible collection of weight functions Ψ := (ψj )j=0,...,k grows fast enough such that for every j = 0, ..., k and pX ∈ P(X,τX ) ∪ QX it holds that
(2.2)
lim
sup
R→∞ (u,v1 ,...,vj )∈(U ×X j )\Kj,R
pX (v1 ) · · · pX (vj ) = 0. ψj (u, v1 , ..., vj )
Then, we define Cbk (U ; Y ) ⊆ C k (U ; Y ) as the vector subspace of maps f ∈ C k (U ; Y ) such that (dj f (u; ·))u∈U ⊆ L(X j ; Y ) is equicontinuous, for all j = 0, ..., k, i.e., there exists a constant Cf ≥ 0 and a seminorm pX ∈ P(X,τX ) such that for every j = 0, ..., k, u ∈ U , and v1 , ..., vj ∈ X we have (2.3)
∥dj f (u; v1 , ..., vj )∥Y ≤ Cf pX (v1 ) · · · pX (vj ),
with d0 f (u) := f (u). Furthermore, we define the weighted norm (2.4)
∥dj f (u; v1 , ..., vj )∥Y , j=0,...,k (u,v1 ,...,vj )∈U ×X j ψj (u, v1 , ..., vj )
∥f ∥BΨk (U ;Y ) = max
sup
for f ∈ Cbk (U ; Y ), which is well-defined by (2.2)–(2.3) and Remark 2.2 (i). Now, we can introduce k the weighted function space BΨ (U ; Y ).
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
15
k Definition 2.6. Let (U, Ψ) be a weighted domain. Then, we define BΨ (U ; Y ) as the closure of k Cb (U ; Y ) with respect to ∥ · ∥BΨk (U ;Y ) , which is a Banach space under the weighted norm defined k in (2.4). If Y = R, we shall only write BΨ (U ).
Since the weight functions ψj : U × X j → (0, ∞), j = 0, ..., k, grow on the compact pre-images −1 k ψj ((0, R]), the derivatives of a map f ∈ BΨ (U ; Y ) are typically unbounded. However, the growth j j of d f : U × X → Y is controlled by ψj : U × X j → (0, ∞). Remark 2.7. For simplicity, we always assume that the output space is a Banach space (Y, ∥·∥Y ). However, the following results can be generalized to locally convex topological vector spaces (Y, τY ) as output space. k In order to characterize maps in BΨ (U ; Y ) in Proposition 2.10 below, we first provide some examples of weighted domains, which have the (bounded) approximation property ((B)AP). To this end, we assume that the Banach space (X, ∥ · ∥X ) is equipped with a weaker topology τX than the norm topology (see Lemma 2.3). For more background on (B)AP, we refer to Section 1.1.
Lemma 2.8. Let (X, ∥ · ∥X ) be a Banach space equipped with the initial topology τinit of a compact embedding Γ : (X, ∥ · ∥X ) → (X0 , τX0 ) as in Lemma 2.3 (i). Moreover, let U ∈ τinit be an open subset and assume that (X0 , τX0 ) has AP (resp., P(X0 ,τX0 ) -BAP) with finite rank operators (T0,γ )γ ∈ (X0 , τX0 )∗ ⊗ Γ(X) satisfying T0,γ (Γ(U )) ⊆ Γ(U ). In addition, let Ψ = (ψj )j=0,...,k be a collection of weight functions of the form (2.1) with η : [0, ∞) → (0, ∞) additionally satisfying rk limr→∞ η(r) = 0. Then, (U, τinit ) has AP (resp., P(X,τinit ) -BAP) and (2.2) is satisfied. Proof. First, we observe that the image Γ(K) of any relatively compact subset K of (X, τinit ) under the continuous embedding Γ : (X, ∥ · ∥X ) → (X0 , τX0 ) is relatively compact in (X0 , τX0 ). Now, since (X0 , τX0 ) has AP, there exists a net of finite rank operator (T0,γ )γ ∈ (X0 , τX0 )∗ ⊗ Γ(X) such that for every relatively compact subset K of (X0 , τinit ) and pX0 ∈ P(X0 ,τX0 ) we have lim
(2.5)
sup
γ x ∈Γ(K) 0
pX0 (x0 − T0,γ (x0 )) = 0.
Then, for every γ, there exists some Tγ ∈ (X, τinit )∗ ⊗ X with Γ ◦ Tγ = Tγ,0 ◦ Γ and Γ(Tγ (U )) = T0,γ (Γ(U )) ⊆ Γ(U ) implying that Tγ (U ) ⊆ U . Hence, (2.5) ensures for every relatively compact subset K of (X, τinit ) and x 7→ pX (x) := pX0 (Γ(x)) ∈ P(X,τinit ) that lim sup pX (x − Tγ (x)) = lim sup pX0 (Γ(x − Tγ (x))) γ x∈K
γ x∈K
= lim sup pX0 (Γ(x) − Γ(Tγ (x))) γ x∈K
= lim
sup
γ x ∈Γ(K) 0
pX0 (x0 − T0,γ (x0 )) = 0,
which shows that (U, τinit ) has AP. Moreover, if (X0 , τX0 ) additionally has P(X0 ,τX0 ) -BAP, then for every pX0 ∈ P(X0 ,τX0 ) there exists some qX0 ∈ P(X0 ,τX0 ) such that for every γ and x0 ∈ X0 it holds that pX0 (Tγ,0 (x0 )) ≤ qX0 (x0 ). Hence, for every x 7→ pX (x) := pX0 (Γ(x)) ∈ P(X,τinit ) , we use x 7→ qX (x) := qX0 (Γ(x)) ∈ P(X,τinit ) to conclude for every γ and x ∈ X that pX (Tγ (x)) = pX0 (Γ(Tγ (x))) = pX0 (Tγ,0 (Γ(x))) ≤ qX0 (Γ(x)) = qX (x), which shows that (U, τX ) has P(X,τinit ) -BAP. Finally, by using that Γ : (X, ∥ · ∥X ) → (X0 , τX0 ) is continuous, i.e., that for every x 7→ pX (x) := pX0 (Γ(x)) ∈ P(X,τinit ) there exists a constant CΓ,pX ≥ 1 such that for every v ∈ X it
16
P. SCHMOCKER AND J. TEICHMANN k
r holds that pX (v) := pX0 (Γ(v)) ≤ CΓ,pX ∥v∥X , and the assumption limr→∞ η(r) = 0, we obtain for every j = 0, ..., k and pX ∈ P(X,τinit ) that
pX (v1 ) · · · pX (vj ) R→∞ j=0,...,k (u,v1 ,...,vj )∈(U ×X j )\Kj,R ψj (u, v1 , ..., vj ) lim
max
sup
j ≤ CΓ,p lim X
max
sup
∥v1 ∥X · · · ∥vj ∥X
Pj η max(j, 1) (δU c (u)−1 + ∥u∥X ) + ℓ=1 ∥vℓ ∥X j Pj 1 + ∥u∥X + ℓ=1 ∥vℓ ∥X k = 0, ≤ CΓ,p lim max sup Pj X R→∞ j=0,...,k (u,v1 ,...,vj )∈(U ×X j )\Kj,R η max(j, 1)∥u∥X + ℓ=1 ∥vℓ ∥X R→∞ j=0,...,k (u,v1 ,...,vj )∈(U ×X j )\Kj,R
□
which shows that (2.2) is satisfied.
Lemma 2.9. Let (X, ∥ · ∥X ) be a dual Banach space equipped with the weak-∗-topology τw∗ . Moreover, let U ∈ τw∗ be an open subset with π(U ) ⊆ U , for all projections π ∈ X ∗ ⊗ X. In addition, let Ψ = (ψj )j=0,...,k be a collection of weights of the form (2.1) with η : [0, ∞) → rk (0, ∞) additionally satisfying limr→∞ η(r) = 0. Then, (U, τw∗ ) has AP and (2.2) is satisfied. Furthermore, if the predual (E, ∥ · ∥E ) of (X, ∥ · ∥X ) has BAP with finite rank operators (Qγ )γ satisfying Q∗γ (U ) ⊆ U , then (U, τw∗ ) has ∥ · ∥X -BAP. Proof. For fixed linearly independent e1 , ..., eN ∈ E, we consider the seminorm x 7→ pX (x) := maxn=1,...,N |⟨x, en ⟩X×E | ∈ P(X,τw∗ ) . Then, by using that E ∗ ∼ = X is due to the Hahn-Banach theorem point separating on E, there exist some linearly independent x1 , ..., xN ∈ X such that ⟨xm , en ⟩X×E = δm,n
for all m, n = 1, ..., N, and ⊥ for all x ∈ X1:N and n = 1, ..., N,
⟨x, en ⟩X×E = 0
⊥ ⊥ = X (see also [102, Corollary 4.2.2]). satisfy X1:N ⊕X1:N where X1:N := span{x1 , ..., xN } and X1:N PN From this, we define the finite rank operator x 7→ Te1:N (x) := n=1 ⟨x, en ⟩X×E xn ∈ (X, τw∗ )∗ ⊗ X satisfying Te1:N (x) = x for any x ∈ X1:N and therefore Te1:N (U ) ⊆ U (as Te1:N is the projection to X1:N ). Thus, for every relatively compact subset K of (X, τw∗ ), it holds that
sup pX (x − Te1:N (x)) = sup x∈K
max
x∈K n=1,...,N
= sup
max
x∈K n=1,...,N
N D E X x− ⟨x, em ⟩X×E xm , en
X×E
m=1
⟨x, en ⟩X×E −
N X m=1
⟨x, em ⟩X×E ⟨xm , en ⟩X×E = 0, {z } | δm,n
which shows that the net (Te1:N )e1:N ⊆ (X, τw∗ )∗ ⊗ X converges to idX : X → X uniformly on each relatively compact subset of (X, τw∗ ), whence (U, τw∗ ) has AP. Moreover, if (E, ∥ · ∥E ) has BAP, there exists some λ ≥ 1 and a net of finite rank operators (Qγ )γ ⊆ E ∗ ⊗ E with ∥Qγ ∥L(E;E) ≤ λ, for all γ, such that for every relatively compact subset L of (E, ∥ · ∥E ), it holds that lim sup ∥e − Qγ (e)∥E = 0. γ e∈L
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
17
Hence, by defining Tγ := Q∗γ ∈ (X, τw∗ )∗ ⊗ X, we conclude for every seminorm x 7→ pX (x) := maxn=1,...,N |⟨x, en ⟩X×E | ∈ P(X,τw∗ ) and relatively compact subset K of (X, τw∗ ) that lim sup pX (x − Tγ (x)) = lim sup γ x∈K
max |⟨(idX −Tγ )(x), en ⟩X×E |
γ x∈K n=1,...,N
= lim sup max |⟨x, (idE −Qγ )en ⟩X×E | γ x∈K n=1,...,M ≤ sup ∥x∥X lim max ∥en − Qγ (en )∥E = 0. γ n=1,...,M
x∈K
In addition, for every x 7→ peX (x) := maxm=1,...,M |⟨x, eem ⟩X×E | ∈ P(X,τw∗ ) , we have peX (Tγ (x)) =
max
m=1,...,M
|⟨Tγ (x), eem ⟩X×E | =
max
m=1,...,M
≤ ∥x∥X ∥Qγ ∥L(E;E)
max
m=1,...,M
∥e e m ∥E ≤
|⟨x, Qγ (e em )⟩X×E | λ max ∥e em ∥E ∥x∥X , m=1,...,M
which shows that (U, τw∗ ) has ∥ · ∥X -BAP. Finally, by using that for every x 7→ pX (x) := maxn=1,...,N |⟨x, en ⟩X×E | ∈ P(X,τw∗ ) there exists a constant CpX > 0 such that for every x ∈ X it holds that pX (x) ≤ CpX ∥x∥X and that rk = 0, we obtain for every j = 0, ..., k and pX ∈ P(X,τw∗ ) that limr→∞ η(r) lim
max
sup
R→∞ j=0,...,k (u,v1 ,...,vj )∈(U ×X j )\Kj,R
≤ CpjX lim
max
pX (v1 ) · · · pX (vj ) ψj (u, v1 , ..., vj ) ∥v1 ∥X · · · ∥vj ∥X
sup
Pj η max(j, 1) (δU c (u)−1 + ∥u∥X ) + ℓ=1 ∥vℓ ∥X j Pj max(j, 1)∥u∥X + ℓ=1 ∥vℓ ∥X = 0, ≤ CpjX lim max sup Pj R→∞ j=0,...,k (u,v1 ,...,vj )∈(U ×X j )\Kj,R η max(j, 1)∥u∥X + ℓ=1 ∥vℓ ∥X R→∞ j=0,...,k (u,v1 ,...,vj )∈(U ×X j )\Kj,R
□
which shows that (2.2) is satisfied.
k In the following, we characterize maps in BΨ (U ; Y ), which extends [33, Theorem 2.7] and [28, Lemma 2.3] to non-differentiable maps. The proof is given in Appendix C.1.
Proposition 2.10. Let (U, Ψ) be a weighted domain satisfying (2.2), where Kj,R := ψj−1 ((0, R]) denote the compact pre-images of the admissible collection Ψ := (ψj )j=0,...,k of weight functions, j = 0, ..., k and R > 0. Then, the following holds true: k k (i) If f ∈ BΨ (U ; Y ), then f ∈ Cloc (U ; Y ) and it holds that (2.6)
lim
max
sup
dj f (u; v1 , ..., vj ) Y = 0. ψj (u, v1 , ..., vj )
sup
dj f (u; v1 , ..., vj ) Y = 0. ψj (u, v1 , ..., vj )
R→∞ j=0,...,k (u,v1 ,...,vj )∈(U ×X j )\Kj,R
k (ii) Let f ∈ Cloc (U ; Y ) satisfy
(2.7)
lim
max
R→∞ j=0,...,k (u,v1 ,...,vj )∈(U ×X j )\Kj,R
Moreover, if (X, τX ) is not locally compact, we assume additionally that (U, τX ) has AP with net of finite rank operators (Tγ )γ ⊆ X ∗ ⊗ X satisfying (2.8)
lim sup
R→∞ γ
max
sup
j=0,...,k j L⊆{1,...,j} (u,v1 ,...,vj )∈(U ×X )\Kj,R
k Then, f ∈ BΨ (U ; Y ).
∥d|L| f (Tγ (u); (Tγ (vℓ ))ℓ∈L )∥Y = 0. ψL (u, vL )
18
P. SCHMOCKER AND J. TEICHMANN
Note that (i) is a straightforward generalization of [33, Theorem 2.7] to differentiable maps. However, for (ii), we need to assume the approximation property (AP) and (2.8) (if (X, τX ) is not locally compact), which is more restrictive than the original result in [33, Theorem 2.7] for 0 0 BΨ -maps. There, the Tietze extension theorem is applied to extend a Cloc -map beyond compacta. k 2.3. Weighted manifolds and BΨ -maps thereon. For some k ∈ N0 ∪ {∞}, we now consider a k Cloc -manifold (M, τM ), which we endow similarly as in Definition 2.1 with a collection of weight functions Ψ := (ψj )j=0,...,k . For more details on the notion of manifolds, we refer to Section 1.3. k Definition 2.11. Let (M, τM ) be a Cloc -manifold with atlas (Ui , ϕi )i∈I over the model spaces (Xi , τXi )i∈I . Then, a collection Ψ := (ψj )j=0,...,k of weight functions ψj : T j M → (0, ∞) is called admissible (on M ) if for every i ∈ I the collection Ψi := (ψi,j )j=0,...,k of push-forward weight functions defined by
(2.9)
j ψi,j := ψj ◦ Φ−j i : ϕi (Ui ) × Xi → (0, ∞)
is admissible on ϕi (Ui ), i.e., if for every i ∈ I the pair (ϕi (Ui ), Ψi ) is a weighted domain. In this k case, we call (M, Ψ) a weighted Cloc -manifold. Note that the admissibility of Ψ := (ψj )j=0,...,k is atlas-dependent because an intrinsic (global) version is not suitable for our approximation results (see also Remark 2.16 below). k Remark 2.12. If (M, Ψ) is a weighted Cloc -manifold, then it holds for every i ∈ I that: (i) ϕi (Ui ) is σ-compact with respect to τXi (see Remark 2.2 (ii)). Hence, by using the continuous function ϕ−1 : ϕi (Ui ) → Ui , the set Ui is σ-compact with respect to τM . i (ii) (Xi , τXi ) is also σ-compact (see Remark 2.2 (iii)). (iii) If (Xi , τXi ) is complete, then Xi is finite-dimensional (see Remark 2.2 (iv)). k Hence, for a weighted Cloc -manifold (M, Ψ), the following holds true: S (iv) If |I| < ∞, then (M, τM ) is σ-compact as M = i∈I Ui with σ-compact Ui (see (i)). (v) If (M, τM ) is a Banach manifold or a Fréchet manifold, i.e., the model spaces (Xi , τXi )i∈I are complete, then (M, τM ) is finite-dimensional (see (iii)). Hence, for an infinite-dimensional manifold (M, τM ), we necessarily have to consider incomplete locally convex topological vector spaces (Xi , τXi )i∈I as model spaces.
Now, we relate the pre-images of the weight functions in Ψ := (ψj )j=0,...,k to the pre-images of the push-forward weights in Ψi := (ψi,j )j=0,...,k , for i ∈ I (see (2.9)). k Lemma 2.13. Let (M, τM ) be a Cloc -manifold over model spaces (Xi , τXi )i∈I and let Ψ := (ψj )j=0,...,k be a collection of weight functions ψj : T j M → (0, ∞), j = 0, ..., k. Then: (i) If for every j = 0, ..., k and R > 0 the pre-image (2.10) Kj,R := ψj−1 ((0, R]) = (x, [c]jx ) ∈ T j M : ψj (x, [c]jx ) ≤ R k is compact with respect to τT j M , then (M, Ψ) is a weighted Cloc -manifold. k (ii) If (M, Ψ) is a weighted Cloc -manifold with |I| < ∞, then for every j = 0, ..., k and R > 0 the pre-image Kj,R defined in (2.10) is compact with respect to τT j M . −1 Proof. For (i), we fix some i ∈ I, j = 0, ..., k, and R > 0. Then, Ki,j,R := ψi,j ((0, R]) is a compact j j subset of (ϕi (Ui ) × Xi , τXi × τXi ) as image of the compact set Kj,R under the continuous chart Φji : T j M → ϕi (Ui ) × Xij (see (1.10)), whence (ϕi (Ui ), Ψi ) is a weighted domain. Since i ∈ I was k chosen arbitrarily, (M, Ψ) is a weighted Cloc -manifold. For (ii), we assume that |I| < ∞ and fix some i ∈ I, j = 0, ..., k, and R > 0. Then, Ki,j,R := j −1 ψi,j ((0, R]) is by definition compact subset of (ϕi (Ui ) × Xij , τXi × τX ). Hence, by using that i
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
19
−j j j in (T j M, τT j M ) Φ−j i : ϕi (Ui )×Xi → T M is continuous (see (1.9)), the set Φi (Ki,j,R ) is compact S −1 as continuous image of the compact set Ki,j,R . Hence, Kj,R := ψj ((0, R]) = i∈I Φ−j i (Ki,j,R ) is compact with respect to τT j M as finite union of compact sets. □
Let us give an example of a weighted manifold (M, Ψ) in the following. Example 2.14. Let (X, τX ) be a locally convex topological vector space and let Ψ := (ψj )j=0,...,k be a collection of admissible weight functions on X. Moreover, let s ∈ C k (X; Rd ) have constant rank, i.e., dim ({ds(x; v) : v ∈ X}) = r, for all x ∈ X and some r ∈ N. Then, for any y ∈ Rd , the k pre-image M := s−1 ({y}) is by [44, Theorem F] a (split) Cloc -submanifold of (X, τX ). Moreover, the collection Ψ|M := (ψj |M )j=0,...,k of restricted weight functions is admissible on M . For example, the space measures x : FΩ → [0, 1] over a weighted R M := PΨΩ (Ω) of probability ∞ space (Ω, ψΩ ) satisfying Ω ψ(ω)x(dω) < ∞ is a Cloc -manifold over the model space X := MψΩ (Ω) equipped with the weak-∗-topology (see also Example (ii) (k)), where Ψ := (ψj )j=0,...,k is of the ∞ form (2.1). Indeed, M = s−1 ({1}) is the pre-image of the Cloc -map MψΩ (Ω) ∋ x 7→ s(x) := x(Ω) ∈ R having constant rank equal to one. For further examples of weighted manifolds, we refer to Section 3. In order to introduce maps on weighted manifolds, we assume that the input space (M, Ψ) is a k weighted Cloc -manifold and that the output space (Y, ∥ · ∥Y ) is a Banach space. k Definition 2.15. Let (M, Ψ) be a weighted Cloc -manifold with atlas (Ui , ϕi )i∈I over the model spaces (Xi , τXi )i∈I and let Ψi := (ψi,j )j=0,...,k be the collection of push-forward weight functions k (M ; Y ) as the vector space of functions f : M → Y such introduced in (2.9). Then, we define BΨ −1 k k (M ; Y ) with the initial topology τBΨk (M ;Y ) (ϕ (U ); Y ), for all i ∈ I. We equip BΨ that f ◦ϕi ∈ BΨ i i i with respect to the family of mappings
(2.11)
k BΨ (M ; Y ) ∋ f
7→
k f ◦ ϕ−1 ∈ BΨ (ϕi (Ui ); Y ), i i
i ∈ I,
i.e., the weakest topology such that the mappings (2.11) are continuous. k (M ; Y ) introduced in Definition 2.15 depends Remark 2.16. As in Definition 2.11, the space BΨ on the choice of atlas (Ui , ϕi )i∈I for the manifold M . This dependence cannot be avoided for the infinite-dimensional approximation results in Section 4–6 below, since the lack of partitions of unity on infinite-dimensional model spaces excludes the gluing of finite-dimensional local approximations into a global (atlas-independent) construction.
For simplicity, we shall always assume that the output space is a Banach space (Y, ∥ · ∥Y ). However, Definition 2.15 could be also extended to a locally convex topological vector space (Y, τY ) as output space (see also Remark 2.7). 3. Weighted Nachbin theorems In this section, we extend the Nachbin theorem to weighted (possibly infinite-dimensional) manifolds. Originally established by L. Nachbin in [86] over finite-dimensional manifolds, the theorem generalizes the classical Stone-Weierstrass theorem by including the approximation of the derivatives. Subsequently, the Nachbin theorem was extended in [3,95] to infinite-dimensional Banach spaces as input and output spaces, using the compact-open topology (of higher order) or the topology of compact convergence (of higher order), and in [89] to a weighted approximation result for polynomials over the Euclidean space. First, we recall the classical Nachbin theorems.
20
P. SCHMOCKER AND J. TEICHMANN
3.1. Classical formulation. Let us denote by Pol(Rd ) ⊆ C ∞ (Rd ) the vector space of polynomiP Qd i als of the form Rd ∋ x := (x1 , ..., xd )⊤ 7→ α∈Nm cα i=1 xα i ∈ R, with n ∈ N0 and cα ∈ R. 0,n
Theorem 3.1 (Weierstrass, [75, Theorem 1.1.2]). Pol(Rd ) is a dense subset of C k (Rd ) with respect to the compact-open topology1 of order k. Subsequently, the Weierstrass theorem (Theorem 3.1) was generalized by L. Nachbin in [86] to the notion of subalgebras. Hereby, a vector space G of maps g : X → R is called a subalgebra if G is closed under multiplication, i.e., g1 · g2 ∈ G, for all g1 , g2 ∈ G. ∞ Theorem 3.2 (Nachbin on C k (M ), [86, p. 1550]). Let (M, τM ) be a Cloc -manifold over a finitek dimensional vector spaces. Moreover, let G ⊆ C (M ) be a subalgebra such that (i) G is point separating on M , i.e., for any distinct points x1 , x2 ∈ M there exists some g ∈ G such that g(x1 ) ̸= g(x2 ), (ii) G vanishes nowhere on M , i.e., for every x ∈ X there exists g ∈ G with g(x) ̸= 0, (iii) G has nowhere vanishing derivatives on M , i.e., for every (x, [c]1x ) ∈ T 1 M with c′ (0) ̸= 0 there exists g ∈ G such that (g ◦ c)′ (0) ̸= 0. Then, G is a dense subset of C k (M ) with respect to the compact-open topology1 of order k.
Later, the Nachbin theorem was generalized by J.B. Prolla and C.S. Guerreiro in [95] to the following infinite-dimensional setting. For two locally convex topological vector spaces (X, τX ) and (Y, τY ), the space of continuous homogeneous polynomials of finite type is defined as Pf (X; Y ) = span {X ∋ x 7→ ℓ(x)n y ∈ Y : n ∈ N0 , ℓ ∈ X ∗ , y ∈ Y } ⊆ C 0 (X; Y ). Then, a subset G ⊆ C 0 (X; Y ) is called a polynomial algebra if r ◦ g ∈ G for all g ∈ G and r ∈ Pf (Y ; Y ). This is the case if and only if G ′ := {ℓ ◦ g : ℓ ∈ Y ∗ , g ∈ G} is a subalgebra with G ′ ⊗ Y ⊆ G (see [94, Lemma 4.6]). Theorem 3.3 (Nachbin on C k (U ; Y ), [95, Theorem 3.3]). Let (X, ∥·∥X ) be a Banach space having AP and let U ⊆ X be open. Moreover, let G ⊆ C k (U ; Y ) be a polynomial subalgebra such that (i) G ′ := {ℓ ◦ g : ℓ ∈ Y ∗ , g ∈ G} is point separating, (ii) G ′ vanishes nowhere, i.e., for every x ∈ U there exists g ′ ∈ G ′ with g(x) ̸= 0, (iii) G has nowhere vanishing derivatives, i.e., for every x ∈ U and v ∈ X \ {0} there exists g ∈ G with dg(x; v) ̸= 0, and (iv) for every g ∈ G, T ∈ X ∗ ⊗ X, and open subset V ⊆ U with T (V ) ⊆ U the composition g ◦ T |V ∈ C k (V ; Y ) belongs to the closure of G|V in C k (V ; Y ). 1
Then, G is a dense subset of C k (U ; Y ) with respect to the compact-open topology of order k. The Nachbin theorem was later extended to the topology2 of compact convergence of order k by R.M. Aron and J.B. Prolla in [3], and into other directions (see, e.g., [45, 75, 88]). 3.2. Subalgebras of Ψ-moderate growth. For the weighted Nachbin theorems, we impose the k following conditions on a given subalgebra G ⊆ BΨ (M ). This is analogous to the concept of point separating and nowhere vanishing subalgebras of Ψ-moderate growth, which was introduced in [28, Definition 3.4] for the non-differentiable case and is inspired by Nachbin’s definition of localisability (see [87, Definition 4]). For j = 0, ..., k and a partition π ∈ Pj , we use the notation Q|π| dπ a(u; vπ ) = r=1 d|πr | a(u; vπr ) with d|πr | a(u; vπr ) := d|πr | a(u; (vℓ )ℓ∈πr ). 1For Banach spaces (X, ∥·∥ ), (Y, ∥·∥ ), and U ⊆ X open, the compact-open topology of order k on C k (U ; Y ) is X Y generated by seminorms pK,L (f ) := maxj=0,...,k sup(u,v)∈K×L ∥dj f (u; v, ..., v)∥Y , for compact K ⊆ U and L ⊆ X. 2For Banach spaces (X, ∥ · ∥ ), (Y, ∥ · ∥ ), and U ⊆ X open, the topology of compact convergence of order k on X Y k C (U ; Y ) is generated by seminorms pK (f ) := maxj=0,...,k supu∈K ∥dj f (u; ·, ..., ·)∥L(X j ;Y ) , for compact K ⊆ U . Note that the topology of compact convergence of order k is stronger than the compact-open topology of order k (except, if (X, ∥ · ∥X ) is finite-dimensional; then they are equivalent).
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
21
k Definition 3.4. Let (M, Ψ) be a weighted Cloc -manifold with atlas (Ui , ϕi )i∈I over finite-dimenk sional model spaces (Xi , τXi )i∈I . Then, a subalgebra G ⊆ BΨ (M ) is called strongly point separating and nowhere vanishing of Ψ-moderate growth if there exists a vector subspace Ge ⊆ G such that (M1) Ge is point separating on M , i.e., for any distinct points x1 , x2 ∈ M there exists some ge ∈ Ge
such that ge(x1 ) ̸= ge(x2 ), (M2) Ge is nowhere vanishing on M , i.e., for every x ∈ X there exists ge ∈ Ge with ge(x) ̸= 0, (M3) Ge has nowhere vanishing derivatives on M , i.e., for every (x, [c]1x ) ∈ T 1 M with c′ (0) ̸= 0 there exists some ge ∈ Ge such that (e g ◦ c)′ (0) ̸= 0, ⊤ (M4) for every i ∈ I there exist ge1 , ..., gem ∈ Ge such that ηi := (e g1 ◦ϕ−1 em ◦ϕ−1 i , ..., g i ) : ϕi (Ui ) → m ∞ R is an embedding and there exist some cutoff functions (hi,R )R>0 ∈ Cc (ηi (U )) with 0 ≤ hi,R ≤ 1 and hi,R |ηi (Ki,R ) = 1 such that ∥dℓ (hi,R ◦ ηi )(u)∥Lℓ (Xi ;R) ∥v1 ∥ · · · ∥vj ∥ = 0, R→∞ 1≤ℓ≤j≤k ψi,j (u, v1 , ..., vj ) (u,v1 ,...,vj )∈(ϕi (Ui )×X j )\Ki,j,R lim
max
sup
i
j j=0 πi,0 (Ki,j,R ) with U × Xi ∋ (u, v1:j ) 7→ πi,0 (u, v1:j ) := u ∈ U , and for every i ∈ I the vector space Gei := g ◦ ϕ−1 : g ∈ Ge is of Ψi -moderate growth, i.e. for i
where Ki,R := (M5)
Sk
every gei ∈ Gei and i ∈ I there exists some λ > 0 such that exp (λ |gi (u)|) |dπ gi (u; vπ )| = 0, R→∞ j=0,...,k ψi,j (u, v1 , ..., vj ) (u,v1 ,...,vj )∈(ϕi (Ui )×X j )\Ki,j,R π∈P lim
max
sup
j
i
−1 where Ki,j,R := ψi,j ((0, R]) denotes the compact pre-image of the push-forward weight
functions Ψi := (ψi,j )j=0,...,k defined in (2.9). k If (M, Ψ) is a weighted Cloc -manifold with atlas (Ui , ϕi )i∈I over infinite-dimensional model spaces (Xi , τXi )i∈I each having BAP with finite rank operators (Ti,γ )γ , we replace (M4) and (M5) by (M4’) for every i ∈ I and γ the set Gei,γ := ge ◦ ϕ−1 | γ (ϕi (Ui )) : ge ∈ Ge satisfies (M4), and i T−1 (M5’) for every i ∈ I and γ the set Gei,γ := g ◦ ϕi |Tγ (ϕi (Ui )) : g ∈ Ge is of Ψi,γ -moderate growth, i.e. for every gei,γ ∈ Gei,γ there exists some λ > 0 such that
exp (λ |gi (u)|) |dπ gi (u; vπ )| = 0, R→∞ j=0,...,k (u,v1 ,...,vj )∈(Tγ (ϕi (Ui ))×Tγ (Xi )j )\Ki,γ,j,R ψi,γ,j (u, v1 , ..., vj ) π∈P lim
max
sup
j
−1 where Ki,γ,j,R := ψi,γ,j ((0, R]) denotes the weight functions Ψi,γ := (ψi,γ,j )j=0,...,k defined by ψγ,j (e u, ve1 , ..., vej ) := inf (u,v1 ,...,vj )∈U ×X j , Tγj+1 (u,v1 ,...,vj )=(eu,ev1 ,...,evj ) ψj (u, v1 , ..., vj ).
Remark 3.5. The conditions (M1)–(M3) ensure that the classical Nachbin theorem on compacta (Theorem 3.2) can be applied. Moreover, (M4) is needed to localize compact subsets of M , where the corresponding limit is zero if, e.g., ηi has uniformly bounded derivatives, i.e., maxℓ=1,...,k supu∈U ∥dℓ ηi (u)∥Lℓ (Xi ;R) < ∞ (see (2.2)). In addition, if G consists of bounded maps, then the exponential part in (M5) is bounded, whence (M5) is by monotonicity of Ψ satisfied. While (M1), (M2), and (M5) are similar to the non-differentiable case in [28, Definition 3.4], the conditions (M3) and (M4) are needed to include the approximation of the derivatives in our weighted setting. Moreover, (M5) is an analogue of the exponential moment condition for the uniqueness of the moment problem. The proof can be found in Appendix D.1. k Lemma 3.6. For an open subset U ⊆ Rd , let ge ∈ BΨ (U ) satisfy (M5) with Ui := U and ϕi = idUi . k k Then, R ∋ t 7→ cos(te g (·)) ∈ BΨ (U ) and R ∋ t 7→ sin(te g (·)) ∈ BΨ (U ) are real-analytic.
22
P. SCHMOCKER AND J. TEICHMANN
3.3. Weighted Nachbin theorems over finite-dimensional manifolds. Now, we formulate a generalized version of the Nachbin theorem in our weighted setting, which extends the weighted approximation results in [89,121] for polynomials over Rd to the notion of subalgebras over finitedimensional manifolds. k k Theorem 3.7 (Nachbin on BΨ (M )). Let (M, Ψ) be a weighted Cloc -manifold over finite-dimenk sional vector spaces. Moreover, let G ⊆ BΨ (M ) be a subalgebra such that G is strongly point k separating and nowhere vanishing of Ψ-moderate growth. Then, G is dense in BΨ (M ). k Proof. First, we show the conclusion for a subalgebra G ⊆ BΨ (U ) consisting of bounded maps over e a weighted domain (U, Ψ), where we can choose G := G as strongly point and nowhere vanishing k separating vector subspace (see Remark 3.5). Since BΨ (U ) is defined as the closure of Cbk (U ) with respect to ∥·∥BΨk (U ) , it suffices to approximate any given f ∈ Cbk (U ) by an element of G. To this end, we fix some f ∈ Cbk (U ) and ε > 0. Moreover, by defining Cf := 1 + maxj=0,...,k ∥dj f ∥Lj (Rd ;R) > 0, it holds for every j = 0, ..., k and (u, v1 , ..., vj ) ∈ U × (Rd )j that
|dj f (u; v1 , ..., vj )| ≤ Cf ∥v1 ∥ · · · ∥vj ∥. In addition, by (2.2) and (M4), there exists some R2 > R > 0 such that max
sup
j=0,...,k (u,v ,...,v )∈(U ×(Rd )j )\K 1
j
∥v1 ∥ · · · ∥vj ∥ ε ≤ , kC ψ (u, v , ..., v ) 3 · 2 j 1 j f j,R
∥dℓ (h ◦ η)(u)∥Lℓ (Rd ;R) ∥v1 ∥ · · · ∥vj ∥ ε ≤ , j,ℓ=1,...,k (u,v ,...,v )∈(U ×(Rd )j )\K ψj (u, v1 , ..., vj ) 3 · 2k Cf 1 j j,R max
sup
e and h := hR ∈ Cc∞ (η(U )) where η := (e g1 , ..., gem )⊤ : U → Rm is an embedding (with ge1 , ..., gem ∈ G) is a cutoff function with 0 ≤ h ≤ 1, h|η(KR ) = 1, and h|Rm \η(KR2 ) = 0. Then, by using the Leibniz product rule (if j ≥ 1), we conclude that (3.1) |dj (1 − h ◦ η) · f (u; v1 , ..., vj )| max sup j=0,...,k (u,v1 ,...,vj )∈(U ×X j )\Kj,R ψj (u, v1 , ..., vj ) P |L| (1 − h ◦ η)(u; vL )||dj−|L| f (u, vLc )| L⊆{1,...,j} |d ≤ max sup j=0,...,k (u,v1 ,...,vj )∈(U ×X j )\Kj,R ψj (u, v1 , ..., vj ) ≤ 2k
∥d|L| (1 − h ◦ η)(u)∥L|L| (Rd ;R) ∥dj−|L| f (u)∥Lj−|L| (Rd ;R) ∥v1 ∥ · · · ∥vj ∥ j=0,...,k ψj (u, v1 , ..., vj ) j L⊆{1,...,j} (U ×X )\Kj,R max
sup
∥dℓ (h ◦ η)(u)∥Lℓ (Rd ;R) ∥v1 ∥ · · · ∥vj ∥ 1≤ℓ≤j≤k (u,v ,...,v )∈(U ×(Rd )j )\K ψj (u, v1 , ..., vj ) 1 j j,R ε ε ≤ 2k Cf = , 3 · 2 k Cf 3 ≤ 2k Cf
max
sup
Sk where we used that ∥d0 (h ◦ η)(u)∥L0 (Rd ;R) = |h(η(u))| ≤ 1. Now, on the set K := j=0 π0 (Kj,R2 ) (being compact as continuous image of the compact pre-images Kj,R2 := ψj−1 ((0, R2 ]), j = 0, ..., k), we can apply the classical Nachbin theorem (Theorem 3.2) to obtain some g ∈ G satisfying ε (3.2) max sup ∥dj f (u) − dj g(u)∥Lj (Rd ;R) < , j=0,...,k u∈K 3 · 2k k!CΨ Cη Cinf ∥v ∥···∥v ∥
1 j where the constant Cinf := maxj=0,...,k sup(u,v1 ,...,vj ) ψj (u,v > 0 is by (2.2) finite, and the 1 ,...,vj ) j
constant Cη := maxj=0,...,k sup(u,v1 ,...,vj )∈U ×(Rd )j
|d (h◦η)(u;v1 ,...,vj )| ψj (u,v1 ,...,vj )
> 0 is by (M4) finite. Thus,
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
23
by using again the Leibniz product rule (if j ≥ 1), that supp(h ◦ η) ⊆ K, the monotonicity of Ψ := (ψj )j=0,...,k , and (3.2), it follows that dj (h ◦ η) · (f − g) (u; v1 , ..., vj ) sup ∥(h ◦ η) · (f − g)∥BΨk (U ) = max j=0,...,k (u,v ,...,v )∈U ×(Rd )j ψj (u, v1 , ..., vj ) 1 j P |L| (h ◦ η)(u; vL )||dj−|L| (f − g)(u; vLc )| L⊆{1,...,j} |d ≤ CΨ max sup j=0,...,k (u,v ,...,v )∈K×(Rd )j ψ|L| (u, vL )ψj−|L| (u, vLc ) 1 j (3.3)
∥dj f (u) − dj g(u)∥Lj (Rd ;R) ∥v1 ∥ · · · ∥vj ∥ j=0,...,k (u,v ,...,v )∈K×(Rd )j ψj (u, v1 , ..., vj ) 1 j
≤ 2k CΨ Cη max
sup
≤ 2k CΨ Cη Cinf sup max ∥dj f (u) − dj g(u)∥Lj (Rd ;R) u∈K j=0,...,k
k
< 2 k!CΨ Cη Cinf
ε ε = . 3 · 2k k!CΨ Cη Cinf 3
Therefore, by defining the function Rm × R ∋ (y, z) 7→ H(y, z) := h(y)z ∈ R, we conclude from (3.1) and (3.3) that ∥f − H ◦ (η ⊕ g)∥BΨk (U ) ≤ ∥f − (h ◦ η) · f + (h ◦ η) · (f − g)∥BΨk (U ) ≤ ∥f − (h ◦ η) · f ∥BΨk (U ) + ∥(h ◦ η) · (f − g)∥BΨk (U )
(3.4)
≤
2ε ε ε + = . 3 3 3
e := η(U ) × g(U ) ⊆ Rm × R ∼ Next, we define the set K = Rm+1 , which is compact as G consists of bounded maps. Then, by applying the Weierstrass theorem (Theorem 3.1), there exists some pn ∈ Pol(Rm+1 ) ∼ = Pol(Rm × R) satisfying ε (3.5) max sup ∥dj H(y, z) − dj pn (y, z)∥Lj (Rm ×R;R) < . j=0,...,k 3k!C η,g e (y,z)∈K Hence, by using the Faà di Bruno formula (if j ≥ 1), that |Pj | ≤ j! ≤ k!, and (3.5), we have (3.6) ∥H ◦ (η ⊕ g) − pn ◦ (η ⊕ g)∥BΨk (U ) P |π| π π∈Pj d (H − pn ) (η(u), g(u)); d (η ⊕ g)(u; vπ ) = max sup j=0,...,k (u,v ,...,v )∈U ×(Rd )j ψj (u, v1 , ..., vj ) 1 j Q|π| ∥d|π| H(η(u), g(u))−d|π| pn (η(u), g(u))∥L|π| (Rm ×R;R) r=1 ∥d|πr | (η ⊕ g)(u; vπr )∥ ≤ k! max sup j=0,...,k ψj (u, v1 , ..., vj ) U ×(Rd )j π∈P j
≤ k!Cη,g max
j=0,...,k
sup ∥dj H(y, z) − dj pn (y, z)∥Lj (Rm ×R;R) e (y,z)∈K
ε ε < k!Cη,g = . 3k!Cη,g 3 Finally, by combining (3.4) with (3.6), it follows for pn ◦ (η ⊕ g) ∈ G (as G is a subalgebra) that ∥f − pn ◦ (η ⊕ g)∥BΨk (U ) ≤ ∥f − H ◦ (η ⊕ g)∥BΨk (U ) + ∥H ◦ (η ⊕ g) − pn ◦ (η ⊕ g)∥BΨk (U ) <
2ε ε + = ε. 3 3
k Since f ∈ Cbk (U ) and ε > 0 was chosen arbitrarily, this shows that G is dense in BΨ (U ). k Now, we show the conclusion for a general subalgebra G ⊆ BΨ (U ) over a weighted domain (U, Ψ), where Ge ⊆ G denotes the strongly point separating and nowhere vanishing vector subspace
24
P. SCHMOCKER AND J. TEICHMANN
of Ψ-moderate growth. By using R ∋ s 7→ cos0 (s) := cos(s) − 1 ∈ R, we introduce the set n o n o Gtrig := span cos0 ◦e g : ge ∈ Ge ∪ sin ◦e g : ge ∈ Ge , k which consists of bounded maps. In order to show that Gtrig ⊆ BΨ (U ), we fix some ge ∈ Ge and k k ε > 0. Since ge ∈ BΨ (U ), there exists by definition of BΨ (U ) some b ∈ Cbk (U ) such that (3.7) |dj ge(u; v1 , ..., vj ) − dj b(u; v1 , ..., vj )| ε , ∥e g − b∥BΨk (U ) = max sup < k j=0,...,k (u,v ,...,v )∈U ×(Rd )j ψ (u, v , ..., v ) k · k!(C j 1 j Ψ Cg e) 1 j
where Cge := 1 + ∥e g ∥BΨk (U ) . Thus, the Faà di Bruno formula, |Pj | ≤ j! ≤ k!, that | cos(j) (s)| ≤ 1, a |d|πℓ | b(u,vπℓ )| |d|πℓ | g e(u,vπℓ )| g ∥BΨk (U ) = Cge, and (3.7) imply that ψ|πℓ | (u,vπℓ ) ≤ ε + ψ|πℓ | (u,vπℓ ) ≤ 1 + ∥e
telescoping sum, and (3.8)
(|π|) (e g (u)) dπ ge(u; vπ )−dπ b(u, vπ ) π∈Pjcos0
P ∥ cos0 ◦e g − cos0 ◦b∥BΨk (U ) = max
sup
ψj (u, v1 , ..., vj ) Q|πr | π (|π|) | cos0 (e g (u))| r=1 d|πr | ge(u; vπr ) − r=1 d b(u; vπr ) sup ≤ k! max j=0,...,k ψ (u, v , ..., v ) d j j 1 j π∈Pj (u,v1 ,...,vj )∈U ×(R ) P|π| Qr Q|π| |πℓ | ge(u, vπℓ )||d|πr | ge(u, vπr )−d|πr | b(u, vπr )| ℓ=r+1 |d|πℓ | b(u, vπℓ )| r=1 ℓ=1 |d k sup ≤ k!CΨ max Qr Q|π| j=0,...,k d j π∈Pj U ×(R ) ℓ=1 ψ|πℓ | (u, vπℓ )ψ|πr | (u, vπr ) ℓ=r+1 ψ|πℓ | (u, vπℓ ) j=0,...,k (u,v ,...,v )∈U ×(Rd )j 1
j
Q|π|
≤ k · k!(CΨ Cge)k
|d|πr | ge(u, vπr ) − d|πr | b(u, vπr )| j=0,...,k ψ|πr | (u, vπr ) d j π∈P , r=1,...,|π| (u,v1 ,...,vj )∈U ×(R ) max
sup
j
k
≤ k · k!(CΨ Cge)
ε = ε. k · k!(CΨ Cge)k
k Since ε > 0 was chosen arbitrarily and cos0 ◦b ∈ Cbk (U ), this shows that cos0 ◦e g ∈ BΨ (U ), which k k holds analogously for sin ◦e g ∈ BΨ (U ), thus Gtrig ⊆ BΨ (U ). Moreover, the trigonometric identities
cos0 (s) cos0 (t) = cos(s) cos(t) − cos(s) − cos(t) + 1 1 = cos(s − t) + cos(s + t) − cos(s) − cos(t) + 1 2 1 cos0 (s − t) + cos0 (s + t) − cos0 (s) − cos0 (t), = 2 cos0 (s) sin(t) = cos(s) sin(t) − sin(t) (3.9) 1 = sin(s + t) − sin(s − t) − sin(t), 2 1 sin(s) sin(t) = cos(s − t) − cos(s + t) 2 1 = cos0 (s − t) − cos0 (s + t) , 2 ensure that Gtrig is a subalgebra. Now, we check that Gtrig satisfies (M1)–(M5). For (M1), we find for any distinct points u1 , u2 ∈ U some ge ∈ Ge with ge(u1 ) ̸= ge(u2 ). Thus, for t ̸= 0 small enough, te g (u1 ) ̸= te g (u2 ) are distinct points in (− π2 , π2 ) and the map sin(te g (·)) ∈ Gtrig separates them. For (M2), there exist for every u ∈ U some ge ∈ Ge with ge(u) ̸= 0, whence there exists a suitable t ̸= 0 such that the map sin(te g (·)) ∈ Gtrig satisfies sin(te g (u)) ̸= 0. For (M3), we find for any u ∈ U and d e v ∈ R \ {0} some ge ∈ G with de g (u; v) ̸= 0, thus either the map cos0 ◦e g ∈ Gtrig or sin ◦e g ∈ Gtrig satisfies d(cos0 ◦e g )(v) = − sin(e g (v))de g (u; v) ̸= 0 or d(sin ◦e g )(v) = cos(e g (v))de g (u; v) ̸= 0. For (M4),
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
25
there exist some ge1 , ..., gem ∈ Ge such that η := (e g1 , ..., gem )⊤ : U → Rm is an embedding. Hence, by using that R ∋ s 7→ (cos0 (s), sin(s), cos0 (πs), sin(πs))⊤ ∈ R4 is injective, the map ⊤ U ∋ u 7→ ηtrig (u) := cos0 (e gℓ (u)), sin(e gℓ (u)), cos0 (πe gℓ (u)), sin(πe gℓ (u)) ℓ=1,...,m ∈ R4m is an embedding with components from Gtrig satisfying (M4). For (M5), we use that Gtrig consists of bounded maps implying that (M5) is already satisfied (see Remark 3.5). Thus, we can now k apply the previous step to conclude that Gtrig is dense in BΨ (U ). Next, we show that Gtrig is contained in the closure of G with respect to ∥ · ∥BΨk (U ) . To this end, we fix some ge ∈ Ge and ε > 0. Then, by (M5), there exists some λ > 0 and R > 0 such that 1 + exp(λ|e g (u)|) |λ||π| |dπ ge(u; vπ )| ε (3.10) max sup < . j=0,...,k ψ (u, v , ..., v ) 2k! j j 1 j (u,v1 ,...,vj )∈(U ×X )\Kj,R π∈P j
k
Sk
g (u)| j=0 π0 (Kj,R ) and the constants c := λ supu∈K |e Pn (−1)ℓ 2ℓ . Hence, by using the Taylor polynomial pn (s) := ℓ=1 (2ℓ)! s
From this, we define the compact set K :=
as well as Cge := (1 + ∥λe g ∥BΨk (U ) ) of cos0 : R → R, there exists a large enough n ∈ N such that (j)
sup | cos0 (s) − p(j) n (s)| ≤
max
(3.11)
j=0,...,k s∈[−c,c]
ε c2n−k+1 < . kC (2n − k + 1)! 2k!CΨ g e
Thus, by using the Faà di Bruno formula, |Pj | ≤ j! ≤ k!, the monotonicity of Ψ := (ψj )j=0,...,k , (|π|) that | cos(|π|) (s)| ≤ 1, that |pn (s)| ≤ exp(|s|), and (3.10) as well as (3.11), it follows that ∥ cos0 (λe g (·)) − pn (λe g (·))∥BΨk (U ) P = max
(|π|)
π∈Pj
sup
cos0
j
(|π|)
≤ k! max
cos0
sup
j=0,...,k (u,v ,...,v )∈(U ×(Rd )j )\K 1
j
(λe g (u)) dπ (λe g )(u; vπ )
ψj (u, v1 , ..., vj )
j=0,...,k (u,v ,...,v )∈U ×(Rd )j 1
(|π|)
(λe g (u)) − pn
j,R
(λe g (u)) |dπ (λe g )(u; vπ )| ψj (u, v1 , ..., vj )
(|π|)
(λe g (u)) |dπ (λe g )(u; vπ )| j=0,...,k (u,v ,...,v )∈(U ×(Rd )j )\K ψ (u, v , ..., v j 1 j) 1 j j,R Q|π| (|π|) (|π|) g (u)) r=1 |d|πr | (λe g )(u; vπr )| cos0 (λe g (u)) − pn (λe k + k!CΨ max sup Q |π| j=0,...,k (u,v1 ,...,vj )∈Kj,R r=1 ψ|πr | (u, vπr ) |π| π 1 + exp(λ|e g (u)|) |λ| |d ge(u; vπ )| ≤ k! max sup j=0,...,k (u,v ,...,v )∈(U ×(Rd )j )\K ψj (u, v1 , ..., vj ) 1 j j,R + k! max
pn
sup
k + k!CΨ Cge max
(j)
sup | cos0 (s) − p(j) n (s)|
j=0,...,k s∈[−c,c]
≤ k!
ε ε k + k!CΨ Cge = ε. kC 2k! 2k!CΨ g e
Since ε > 0 was chosen arbitrarily, the map cos0 (λe g (·)) belongs to the closure of G with respect to ∥ · ∥BΨk (U ) , which holds analogously true for sin(λe g (·)). Hence, by using that R ∋ t 7→ cos0 (te g (·)) ∈ k k BΨ (U ) and R ∋ t 7→ sin(te g (·)) ∈ BΨ (U ) are real-analytic (see Lemma 3.6), [28, Lemma 3.7] ensures k k that cos0 (te g (·)) ∈ BΨ (U ) and sin(te g (·)) ∈ BΨ (U ), for all t ∈ R, which shows by taking t = 1 that Gtrig is contained in the closure of Ge with respect to ∥ · ∥BΨk (U ) . Combining this with the previous k k step, i.e., that Gtrig is dense in BΨ (U ), it follows that G is also dense in BΨ (U ).
26
P. SCHMOCKER AND J. TEICHMANN
k k Finally, for a general subalgebra G ⊆ BΨ (M ) over a weighted Cloc -manifold (M, Ψ), we observe −1 k that Gi := g ◦ ϕi : g ∈ G ⊆ BΨi (ϕi (Ui )) is a strongly point separating and nowhere vanishing subalgebra of Ψi -moderate growth (as ϕi : Ui → ϕi (Ui ) is a diffeomorphism), whence Gi is by k k the previous step dense in BΨ (ϕi (Ui )). Thus, by using that BΨ (M ) is equipped with the initial i k topology with respect to (2.11), it follows that G is dense in BΨ (M ). □
Remark 3.8. Theorem 3.7 generalizes the weighted approximation results in [89, Proposition 4] k from polynomials over Rd to the notion of subalgebras on more general weighted Cloc -manifolds. Next, we extend the weighted Nachbin theorem to the vector-valued case. To this end, we first assume that the output space (Y, ∥ · ∥Y ) has the bounded approximation property (BAP), which k k ensures that BΨ (U ) ⊗ Y is dense in BΨ (U ; Y ). Compared to the compact-open topology, for which k C (U ) ⊗ Y is by a compactness argument dense in C k (U ; Y ) (see [95, Lemma 2.1]), the BAP is required in this weighted setting. The proof of the following lemma is given in Appendix D.3. Lemma 3.9. For an open subset U ⊆ Rd , let (U, Ψ) be a weighted domain and assume that (Y, ∥ · ∥Y ) is a Banach space having BAP. Then, k k BΨ (U ) ⊗ Y := span U ∋ u 7→ f (u)y ∈ Y : f ∈ BΨ (U ), y ∈ Y k (U ; Y ). is a dense subset of BΨ k (M ; Y ) (see Section 3.1) with Now, we combine the property of polynomial algebras G ⊆ BΨ Lemma 3.9 to obtain the following vector-valued weighted Nachbin theorem. k k (M ; Y )). Let (M, Ψ) be a weighted Cloc -manifold over some Theorem 3.10 (Nachbin on BΨ finite-dimensional vector spaces and assume that (Y, ∥ · ∥Y ) is a Banach space. Moreover, let k (M ; Y ) be a polynomial subalgebra such that G ′ := {ℓ ◦ g : ℓ ∈ Y ∗ , g ∈ G} is strongly point G ⊆ BΨ k (M ; Y ). separating and nowhere vanishing of Ψ-moderate growth. Then, G is dense in BΨ k (U ; Y ) over a finiteProof. First, we show the conclusion for a polynomial subalgebra G ⊆ BΨ k k (U ; Y ), dimensional weighted domain (U, Ψ). Since BΨ (U ) ⊗ Y is by Lemma 3.9 dense in BΨ k it suffices to approximate any map f ∈ BΨ (U ) ⊗ Y by some element in G. To this end, we PN k k (U ), and fix some f = n=1 fn (·)yn ∈ BΨ (U ) ⊗ Y and ε > 0, where N ∈ N, f1 , ..., fN ∈ BΨ ′ k y1 , ..., yN ∈ Y . Then, by applying Theorem 3.7, we conclude that G is dense in BΨ (U ), which ε implies the existence of some g1 , ..., gN ∈ G ′ such that ∥fn − gn ∥BΨk (U ) < N (1+∥y . Hence, by n ∥Y ) PN ′ using that g := n=1 gn (·)yn ∈ G ⊗ Y ⊆ G (see [94, Lemma 4.6]), it follows that
∥f − g∥BΨk (U ;Y ) ≤
N X
∥fn (·)yn − gn (·)yn ∥Bk (U ;Y ) ≤ Ψ
n=1
≤
N X
∥fn − gn ∥BΨk (U ) ∥yn ∥Y
n=1
N X
ε ∥yn ∥Y ≤ ε. N (1 + ∥yn ∥Y ) n=1
k k Since f ∈ BΨ (U ) ⊗ Y and ε > 0 were chosen arbitrarily, and BΨ (U ) ⊗ Y is a dense subset of k k BΨ (U ; Y ), this shows that G is dense in BΨ (U ; Y ). k k Finally, for a general polynomial subalgebra G ⊆ BΨ (M ; Y ) over a weighted Cloc -manifold −1 k (M, Ψ), we observe that Gi := g ◦ ϕi : g ∈ G ⊆ BΨi (ϕi (Ui ); Y ) is a polynomial subalgebra such that Gi′ := {ℓ ◦ gi : ℓ ∈ Y ∗ , gi ∈ Gi } is a strongly point separating and nowhere vanishing subalgebra of Ψi -moderate growth (as ϕi : Ui → ϕi (Ui ) is a diffeomorphism). Hence, Gi is by the k k previous step dense in BΨ (ϕi (Ui ); Y ). Since BΨ (M ; Y ) is equipped with the initial topology with i k respect to (2.11), G is dense in BΨ (M ; Y ). □
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
27
3.4. Weighted Nachbin theorems over infinite-dimensional manifolds. In this section, we generalize the weighted Nachbin theorem to infinite-dimensional input manifolds by assuming the bounded approximation property (BAP). Recall that a domain (U, τX ) is said to have BAP if (X, τX ) has BAP, i.e., (X, τX ) has BAP with finite rank operators (Tγ )γ ⊆ (X, τX )∗ ⊗X satisfying Tγ (U ) ⊆ U (see also Sections 1.1 and 2.2). Lemma 3.11. Let (U, Ψ) be a weighted domain such that (U, τX ) has BAP and assume that (Y, ∥ · ∥Y ) is a Banach space. Then, for every f ∈ Cbk (U ; Y ) there exists a net of finite rank operators (Tγ )γ ⊆ X ∗ ⊗ X with Tγ (U ) ⊆ U such that limγ ∥f − f ◦ Tγ ∥BΨk (U ;Y ) = 0. We now first apply Lemma 3.9 to reduce the approximation problem onto a finite-dimensional domain and then apply the weighted Nachbin theorem (Theorem 3.7). k k Theorem 3.12 (Nachbin on BΨ (M )). Let (M, Ψ) be a weighted Cloc -manifold with atlas (Ui , ϕi )i∈I over model spaces (Xi , τXi )i∈I such that each domain (ϕi (Ui ), τXi ) has BAP with finite rank opk erators (Ti,γ )γ ⊆ (Xi , τXi )∗ ⊗ Xi . Moreover, let G ⊆ BΨ (M ) be a subalgebra such that
(i) G is strongly point separating and nowhere vanishing of Ψ-moderate growth, and (ii) for every g ∈ G, i ∈ I, and γ the composition g ◦ ϕ−1 ◦ Ti : ϕi (Ui ) → R belongs to the i closure of Gi := g ◦ ϕ−1 : g ∈ G with respect to ∥ · ∥ k (ϕ (U )) . BΨ i i i i
k (M ). Then, G is dense in BΨ k Proof. First, we show the conclusion for a subalgebra G ⊆ BΨ (U ) over a weighted domain (U, Ψ). k k Since BΨ (U ) is defined as closure of Cb (U ) with respect to ∥ · ∥BΨk (U ) , it suffices to approximate any given f ∈ Cbk (U ) by an element of G. Moreover, by using that (U, τX ) has BAP with finite rank operators (Tγ )γ ⊆ (X, τX )∗ ⊗ X, there exists by Lemma 3.11 some γ such that ε (3.12) ∥f − f ◦ Tγ ∥BΨk (U ) < . 3 Now, we define the collection Ψγ := (ψγ,j )j=0,...,k of weights ψγ,j : Tγ (U ) × Tγ (X)j → (0, ∞) by
(3.13)
ψγ,j (e u, ve1 , ..., vej ) :=
inf
(u,v1 ,...,vj )∈U ×X j
ψj (u, v1 , ..., vj ),
j+1 Tγ (u,v1 ,...,vj )=(u,e e v1 ,...,e vj )
for (e u, ve1 , ..., vej ) ∈ Tγ (U ) × Tγ (X)j , where Tγj+1 : U × X j → Tγ (U ) × Tγ (X)j is defined by j+1 Tγ (u, v1 , ..., vj ) := (Tγ (u), Tγ (v1 ), ..., Tγ (vj )), for (u, v1 , ..., vj ) ∈ U × X j . Then, for every R > 0, we claim that −1 ψγ,j ((0, R]) = Tγj+1 (Kj,R ). −1 For ψγ,j ((0, R]) ⊇ Tγj+1 (Kj,R ), there exists for every (e u, ve1 , ..., vej ) ∈ Tγj+1 (Kj,R ) some (u, v1 , ..., vj ) ∈ j+1 Kj,R such that Tγ (u, v1 , ..., vj ) = (e u, ve1 , ..., vej ). Hence, by the definition of ψγ,j , it holds that
ψγ,j (e u, ve1 , ..., vej ) ≤ ψj (u, v1 , ..., vj ) ≤ R, −1 −1 which shows that (e u, ve1 , ..., vej ) ∈ ψγ,j ((0, R]). Conversely, for ψγ,j ((0, R]) ⊆ Tγj+1 (Kj,R ), we −1 fix some (e u, ve1 , ..., vej ) ∈ ψγ,j ((0, R]). Then, by definition of ψγ,j , there exists for every n ∈ (n) (n) (n) (n) N some u(n) , v1 , ..., vj ∈ U × X j with Tγj+1 u(n) , v1 , ..., vj = (e u, ve1 , ..., vej ) such that (n) (n) 1 (n) (n) (n) (n) ∈ Kj,R+1 . Since Kj,R+1 is compact, ≤ R + n , whence u , v1 , ..., vj ψj u , v1 , ..., vj (γ) (γ) (γ) there exists a subnet u , v1 , ..., vj γ , converging to some (u, v1 , ..., vj ) ∈ Kj,R+1 , which together with the continuity of Tγj+1 implies that (γ) (γ) Tγj+1 (u, v1 , ..., vj ) = lim Tγj+1 u(γ) , v1 , ..., vj = (e u, ve1 , ..., vej ). γ
28
P. SCHMOCKER AND J. TEICHMANN
Moreover, by using that ψj is lower semicontinuous, it follows that (γ)
(γ)
ψj (u, v1 , ..., vj ) ≤ lim inf ψj u(γ) , v1 , ..., vj γ
≤ R.
Hence, (u, v1 , ..., vj ) ∈ Kj,R and therefore (e u, ve1 , ..., vej ) = Tγj+1 (u, v1 , ..., vj ) ∈ Tγj+1 (Kj,R ). This −1 proves (3.13), which ensures that ψγ,j ((0, R]) = Tγj+1 (Kj,R ) is compact as continuous image of the compact set Kj,R , showing that Ψγ := (ψγ,j )j=0,...,k is admissible on Tγ (U ) ⊆ Tγ (X). Moreover, we claim that Gγ := G|Tγ (U ) is a strongly point separating and nowhere vanishing of Ψ-moderate growth. Indeed, while the conditions of point separation, nowhere vanishing, and nowhere vanishing derivatives in (M1)–(M3) are inherited to the sub-domain Tγ (U ) ⊆ U , the other conditions (M4’)–(M5’) are defined such that Gγ = G|Tγ (U ) satisfies (M4)–(M5) on (Tγ (U ), Ψγ ). Hence, we can apply the weighted Nachbin theorem (Theorem 3.7) on the map k f |Tγ (U ) ∈ BΨ (Tγ (U ); Y ) to obtain some ge ∈ Gγ such that γ ∥f |Tγ (U ) − ge∥BΨk (Tγ (U )) γ
dj f (e u; ve1 , ..., vej ) − dj ge(e u; ve1 , ..., vej ) ε < . j=0,...,k (e ψ (e u , v e , ..., v e ) 3 j γ,j 1 j u,e v1 ,...,e vj )∈Tγ (U )×Tγ (X)
= max
sup
Thus, by using the chain rule, we conclude that ∥f ◦ Tγ − ge ◦ Tγ ∥BΨk (U ;Y ) |dj (f ◦ Tγ )(u; v1 , ..., vj ) − dj (e g ◦ Tγ )(u; v1 , ..., vj )| j=0,...,k (u,v1 ,...,vj )∈U ×X j ψj (u, v1 , ..., vj )
= max (3.14)
sup
|dj f (Tγ (u); Tγ (v1 ), ..., Tγ (vj )) − dj ge(Tγ (u); Tγ (v1 ), ..., Tγ (vj ))| j=0,...,k (u,v1 ,...,vj )∈U ×X j ψj (u, v1 , ..., vj )
= max
sup
ε |dj f (e u; ve1 , ..., vej ) − dj ge(e u; ve1 , ..., vej )| < . j=0,...,k (e ψ (e u , v e , ..., v e ) 3 j γ,j 1 j u,e v1 ,...,e vj )∈Tγ (U )×Tγ (X)
≤ max
sup
Next, we use that ge ◦ Tγ belongs by (ii) to the closure of G with respect to ∥ · ∥BΨk (U ) to obtain some g ∈ G such that ε (3.15) ∥e g ◦ Tγ − g∥BΨk (U ) < . 3 Finally, by combining (3.12), (3.14), and (3.15), it follows that ∥f − g∥BΨk (U ) ≤ ∥f − f ◦ Tγ ∥BΨk (U ) + ∥f ◦ Tγ − ge ◦ Tγ ∥BΨk (U ) + ∥e g ◦ Tγ − g∥BΨk (U ) ε ε ε < + + = ε. 3 3 3 k Since f ∈ Cbk (U ) and ε > 0 were chosen arbitrarily, this shows that G is dense in BΨ (U ). k k Finally, for a general subalgebra G ⊆ BΨ (M ) over a weighted Cloc -manifold (M, Ψ), we observe k that Gi := g ◦ ϕ−1 : g ∈ G ⊆ BΨ (ϕi (Ui )) is a strongly point separating and nowhere vanishing i i subalgebra of Ψi -moderate growth satisfying (ii). Hence, we can apply the previous step to k k conclude that Gi is dense in BΨ (ϕi (Ui )). Thus, by using that BΨ (M ) is equipped with the initial i k topology with respect to (2.11), it follows that G is dense in BΨ (M ). □ Moreover, by following the arguments of Theorem 3.10, we can derive the following vectorvalued weighted Nachbin theorem over infinite-dimensional manifolds. k k Corollary 3.13 (Nachbin on BΨ (M ; Y )). Let (M, Ψ) be a weighted Cloc -manifold with atlas (Ui , ϕi )i∈I over model spaces (Xi , τXi )i∈I such that each domain (ϕi (Ui ), τXi ) has BAP with finite rank operators (Ti,γ )γ ⊆ (Xi , τXi )∗ ⊗ Xi . Moreover, let (Y, ∥ · ∥Y ) be a Banach space and assume k that G ⊆ BΨ (M ; Y ) is a polynomial subalgebra such that
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
29
(i) G ′ := {ℓ ◦ g : ℓ ∈ Y ∗ , g ∈ G} is strongly point separating and nowhere vanishing of Ψmoderate growth, and (ii) for every g ∈ G,i ∈ I, and γ the composition g ◦ ϕ−1 i ◦ Ti,γ : ϕi (Ui ) → Y belongs to the closure of Gi := g ◦ ϕ−1 : g ∈ G with respect to ∥ · ∥BΨk (ϕi (Ui );Y ) . i i
k Then, G is dense in BΨ (M ; Y ).
Proof. The proof follows from the same arguments as in the proof of Theorem 3.10, where we now apply Theorem 3.12 instead of Theorem 3.7. □ 4. Weighted universal approximation of functional input neural networks We now introduce a generalization of neural networks to infinite-dimensional spaces, called functional input neural networks (FNNs), and show different universal approximation theorems k (UATs) for FNNs. To this end, we assume that the input space (M, Ψ) is a weighted Cloc -manifold j with admissible collection Ψ = (ψj )j=0,...,k of weight functions ψj : T M → (0, ∞), j = 0, ..., k. Moreover, the output space (Y, ∥ · ∥Y ) is supposed to be a Banach space. 4.1. Functional input neural networks. In this section, we define neural networks between infinite-dimensional spaces. To this end, we first introduce the infinite-dimensional analogue of weight matrices that connect adjacent layers in classical neural networks. k -manifold over finite-dimensional model spaces (Xi , τXi )i∈I . Definition 4.1. Let (M, τM ) be a Cloc k A subset A ⊆ Cloc (M ) is called an additive family (on M ) if (A1) A is closed under addition, i.e., for every a1 , a2 ∈ A it holds that a1 + a2 ∈ A, (A2) A is point separating on M , i.e., for any distinct points x1 , x2 ∈ M there exists some a ∈ A such that a(x1 ) ̸= a(x2 ), and (A3) A has nowhere vanishing derivatives on M , i.e., for every (x, [c]1x ) ∈ T 1 M with c′ (0) ̸= 0 there exists some a ∈ A such that (a ◦ c)′ (0) ̸= 0. −1 ⊤ (A4) for every i ∈ I there exist a1 , ..., am ∈ A such that ηi := (a1 ◦ϕ−1 i , ..., am ◦ϕi ) : ϕi (Ui ) → m R is an embedding and there exist some cutoff functions (hi,R )R>0 ∈ Cc∞ (ηi (U )) with 0 ≤ hi,R ≤ 1 and hi,R |ηi (Ki,R ) = 1 such that
∥dℓ (hi,R ◦ ηi )(u)∥Lℓ (Xi ;R) ∥v1 ∥ · · · ∥vj ∥ = 0, R→∞ 1≤ℓ≤j≤k ψi,j (u, v1 , ..., vj ) (u,v1 ,...,vj )∈(ϕi (Ui )×Xij )\Ki,j,R Sk where Ki,R := j=0 πi,0 (Ki,j,R ) with U ×Xij ∋ (u, v1 , ..., vj ) 7→ πi,0 (u, v1 , ..., vj ) := u ∈ U . lim
max
sup
k If (M, τM ) is a Cloc -manifold over infinite-dimensional model spaces (Xi , τXi )i∈I each having BAP with finite rank operators (Ti,γ )γ , we replace condition (A4) by (A4’) for every i ∈ I and γ the restriction Ai |Ti,γ (ϕi (Ui )) := a ◦ ϕ−1 i |Ti,γ (ϕi (Ui )) satisfies (A4).
Remark 4.2. In contrast to [28, Definition 4.1], we do not include the constants into the additive family. Under this consideration, the conditions (A1)–(A2) are the same as in the nondifferentiable case of [28, Definition 4.1], whereas (A3)–(A4) are additionally required for the approximation of the derivatives in our weighted setting. In addition, if the embedding ηi in (A4) has uniformly bounded derivatives, then the corresponding limit is zero (see also Remark 3.5). For an open subset U ⊆ X of the Euclidean space X := Rd , we observe that the weight matrices in classical neural networks form an additive family. d Example 4.3. For an M := U ⊆ X of the Euclidean space open subset X := R ,⊤an additive fam⊤ d ily is given by A = M ∋ x 7→ a x ∈ R : a ∈ R . Note that A = M ∋ x 7→ a x ∈ R : a ∈ Nd0 is an even smaller additive family.
30
P. SCHMOCKER AND J. TEICHMANN
k Definition 4.4. For a given additive family A ⊆ Cloc (M ), a function ρ ∈ C k (R), and a subset L ⊆ Y , we define a functional input neural network (FNN) φ : M → Y as
M ∋x
(4.1)
7→
φ(x) =
N X
yn ρ(an (x) + bn ) ∈ Y,
n=1
where N ∈ N denotes the number of neurons, where a1 , ..., aN ∈ A are the hidden layer maps, where b1 , ..., bN ∈ R represent the biases, and where y1 , ..., yN ∈ L are the linear readouts. Moreover, we denote by N N A,ρ,L M,Y the set of FNNs of the form (4.1). Hidden Layer R
Input Layer (M, Ψ)
Output Layer (Y, ∥ · ∥Y ) φ(x) ∈ Y
M ∋x A⊕R
L ρ
Figure 1. A FNN φ : M → Y with additive family A, activation function ρ ∈ C k (R), linear readout L ⊆ Y , and N = 3 number of neurons. Remark 4.5. Definition 4.4 extends the notion of classical neural networks between Euclidean spaces. Indeed, let φ : Rd → Rm be a classical neural network of the form Rd ∋ x
(4.2)
7→
φ(x) = W ρ(Ax + b) =
N X
m yn ρ a⊤ n x + bn ∈ R ,
n=1
for some W = (y1 , ..., yN ) ∈ Rm×N , A = (a1 , ..., aN )⊤ ∈ RN ×d , and b = (b1 , ..., bN )⊤ ∈ RN , where y1 , ..., yN ∈ Rm denote the columns of W ∈ Rm×N , and where a1 , ..., aN ∈ Rd represent the rows of A ∈ RN ×d . Moreover, by a slight abuse of notation, ρ ∈ C k (R) is applied componentwise to Ax + b ∈ RN after the first equality in (4.2). If we choose A as in Example 4.3 and L = Rm , then . φ : Rd → Rm is a functional input neural network in N N A,ρ,L Rd ,Rm Moreover, we can construct deep functional input neural networks by concatenation. For an k additive family A ⊆ Cloc (M ) and two activation functions ρ1 , ρ2 ∈ C k (R), we introduce a deep FNN with two hidden layers. Indeed, by assuming that ρ1 ∈ C k (R) is (strongly) non-polynomial 1 ,R (see below), the set N N A,ρ is another additive family on M . Hence, a functional input neural M,R network with two hidden layers φ : M → Y is of the form M ∋x
7→
φ(x) =
N2 X
yn2 ρ2 (φn2 (x) + bn2 )
n2 =1
=
N2 X n2 =1
yn2 ρ2
N1 X
! wn2 ,n1 ρ1 (an2 ,n1 (x) + bn1 ) + bn2
∈ Y,
n1 =1
where y1 , ..., yN2 ∈ L are linear readouts, b1 , ..., bN2 ∈ R are the biases of the second layer, w1,1 , ..., wN2 ,N1 ∈ R are the connections between the layers, and where a1,1 , ..., aN2 ,N1 ∈ A and b1 , ..., bN1 ∈ R are the weights and biases of the first layer, respectively. Moreover, φ1 , ..., φN2 ∈ PN 1 ,R N N A,ρ are FNNs of the form φn2 (x) := n11=1 wn2 ,n1 ρ1 (an2 ,n1 (x) + bn1 ), for all x ∈ M and M,R n2 = 1, ..., N2 . Hence, by an analogous concatenation, it is possible to construct deep functional input neural networks with finitely many hidden layers.
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
31
4.2. Examples of additive families. In this section, we give some examples of additive families on weighted manifolds having an atlas with one global chart. More precisely, for k ∈ N ∪ {∞}, k we consider a weighted Cloc -manifold (M, Ψ) with global chart ϕi : Ui → ϕi (Ui ) over a Banach space (X, ∥ · ∥X ) that is equipped with a weaker topology τX than the norm topology (except X is finite-dimensional). This applies in particular to every open subset M := Ui of (X, τX ), where the chart ϕi := idUi : Ui → Ui ⊆ X is equal to the identity. k Lemma 4.6. Let (M, τM ) be a Cloc -manifold with global chart ϕi : M → ϕi (M ) over a Banach space (X, ∥ · ∥X ), which is equipped with the initial topology τinit of a compact embedding Γ : (X, ∥ · ∥X ) → (X0 , ∥ · ∥X0 ) into another Banach space (X0 , ∥ · ∥X0 ). Moreover, assume that (X, τX ) has BAP with finite rank operators (Tγ )γ . Then, k A := {M ∋ x 7→ ℓ(ϕi (x)) ∈ R : ℓ ∈ (X, τinit )∗ } ⊆ Cloc (M )
is an additive family on M . Proof. First, for every ℓ ∈ (X, τinit )∗ , we observe that a := ℓ(ϕi (·)) ∈ A satisfies a ◦ ϕ−1 = i k k ℓ|ϕi (M ) ∈ Cloc (ϕi (M )), which ensures that a ∈ Cloc (M ). Now, we verify the conditions (A1)– (A4). For (A1), A is by definition closed under addition. For (A2), we fix some distinct points x1 , x2 ∈ M , which also satisfy ϕi (x1 ) ̸= ϕi (x2 ) as ϕi : M → ϕi (M ) is injective. Since (X, τinit )∗ is by the Hahn-Banach theorem point separating on X, there exists some ℓ ∈ (X, τinit )∗ such that a := ℓ(ϕi (·)) ∈ A satisfies a(x1 ) = ℓ(ϕi (x1 )) ̸= ℓ(ϕi (x2 )) = a(x2 ). For (A3), we fix some (x, [c]1x ) ∈ T 1 M with c′ (0) ̸= 0, which also satisfies (ϕi ◦ c)′ (0) ̸= 0 as ϕi : M → ϕi (M ) is a diffeomorphism. Thus, by using again the Hahn-Banach theorem, there exists some ℓ ∈ (X, τinit )∗ such that ℓ((ϕi ◦ c)′ (0)) ̸= 0, whence a := ℓ(ϕi (·)) ∈ A satisfies −1 ′ ′ ′ (a ◦ c)′ (0) = (a ◦ ϕ−1 i ◦ ϕi ◦ c) (0) = d(a ◦ ϕi )(ϕi (x); (ϕi ◦ c) (0)) = ℓ((ϕi ◦ c) (0)) ̸= 0.
For (A4’), we fix some γ and consider the finite-dimensional vector subspace Tγ (X) ⊆ X. Then, there exists a basis b1 , ..., bm of Tγ (X) and some ℓ1 , ..., ℓm ∈ (Tγ (X), τinit )∗ such that ℓn (bne ) = δn,en , for all n, n e = 1, ...m. Hence, by the Hahn-Banach theorem, we can extend ℓ1 , ..., ℓm ∈ (Tγ (X), τinit )∗ to some L1 , ..., Lm ∈ (X, τinit )∗ with Ln |Tγ (X) = ℓn , for all n = 1, ..., m, which implies that L := (L1 , ..., Lm )⊤ |Tγ (X) : Tγ (X) → Rm is a linear isomorphism. Hence, by defining an := Ln ◦ ϕi ∈ A, we conclude that −1 ⊤ ηi,γ := a1 ◦ ϕ−1 = L|Tγ (ϕi (M )) : Tγ (ϕi (M )) → ηi,γ (Tγ (ϕi (M ))) i , ..., am ◦ ϕi Tγ (ϕi (M )) is the restriction of a linear isomorphism and therefore an embedding.
□
With the additive family A from Lemma 4.6, a corresponding FNN φ : M → Y is of the form M ∋x
7→
φ(x) :=
N X
yn ρ (ℓn (ϕi (x)) + bn ) ∈ Y,
n=1
where N ∈ N, ℓ1 , ..., ℓN ∈ (X, τinit )∗ , b1 , ..., bN ∈ R, y1 , ..., yN ∈ L ⊆ Y , and ρ ∈ C k (R). k Lemma 4.7. Let (M, τM ) be a Cloc -manifold global chart ϕi : M → ϕi (M ) over a dual Banach space (X, ∥ · ∥X ) which is equipped with the weak-∗-topology τw∗ . Moreover, assume that (X, τw∗ ) has BAP with finite rank operators (Tγ )γ . Then, k A := {M ∋ x 7→ ⟨ϕi (x), e⟩X×E ∈ R : e ∈ E} ⊆ Cloc (M )
is an additive family on M .
32
P. SCHMOCKER AND J. TEICHMANN
Proof. First, for every e ∈ E, we observe that a := ⟨ϕi (·), e⟩X×E ∈ A satisfies a◦ϕ−1 = ⟨·, e⟩X×E ∈ i k k Cloc (ϕi (M )), which ensures that a ∈ Cloc (M ). For (A1), A is by definition closed under addition. For (A2), we fix some distinct points x1 , x2 ∈ M , which also satisfy ϕi (x1 ) ̸= ϕi (x2 ) as ϕi : M → ϕi (M ) is injective. Since E ∗ ∼ = X is by the Hahn-Banach theorem point separating on X, there exists some e ∈ E such that a := ⟨ϕi (·), e⟩X×E ∈ A satisfies a(x1 ) = ⟨ϕi (x1 ), e⟩X×E ̸= ⟨ϕi (x2 ), e⟩X×E = a(x2 ). For (A3), we fix some (x, [c]1x ) ∈ T 1 M with c′ (0) ̸= 0, as ϕi : M → ϕi (M ) is a diffeomorphism. Thus, by using again the Hahn-Banach theorem, there exists some e ∈ E such that ⟨(ϕi ◦ c)′ (0), e⟩X×E ̸= 0, whence a := ⟨ϕi (·), e⟩X×E ∈ A satisfies −1 ′ ′ ′ (a ◦ c)′ (0) = (a ◦ ϕ−1 i ◦ ϕi ◦ c) (0) = d(a ◦ ϕi )(ϕi (x); (ϕi ◦ c) (0)) = ⟨(ϕi ◦ c) (0), e⟩X×E ̸= 0.
For (A4’), we fix some γ and consider the finite-dimensional vector subspace Tγ (X) ⊆ X. Since X∼ = E ∗ is point separating on E, there exist some e1 , ..., em ∈ E such that ⊤ ⊤ Tγ (X) ∋ u 7→ L(u) := L1 (u), ..., Lm (u) := ⟨z, e1 ⟩X×E , ..., ⟨z, em ⟩X×E ∈ Rm is a linear isomorphism. Hence, by defining an := Ln ◦ ϕi ∈ A, we conclude that −1 ⊤ ηi,γ := a1 ◦ ϕ−1 = L|Tγ (ϕi (Ui )) : Tγ (ϕi (Ui )) → ηi,γ (Tγ (ϕi (Ui ))) i , ..., am ◦ ϕi T (ϕ (U )) γ
i
i
is the restriction of a linear isomorphism and therefore an embedding.
□
With the additive family A from Lemma 4.7, a corresponding FNN φ : M → Y is of the form M ∋x
7→
φ(x) :=
N X
yn ρ (⟨ϕi (x), en ⟩X×E + bn ) ∈ Y,
n=1
where N ∈ N, e1 , ..., eN ∈ E, b1 , ..., bN ∈ R, y1 , ..., yN ∈ L ⊆ Y , and ρ ∈ C k (R). In the following, we now apply Lemmas 4.6–4.7 to construct different additive families. Example 4.8. The following examples are additive families: (i) Let M := C α (S; Z) be as in Example 2.5 (b) with admissible collection of weight functions Ψ = (ψj )j=0,...,k . Then, an additive family is given by Z α k A := C (S; Z) ∋ x 7→ ⟨x(s), e⟩Z×E ν(ds) ∈ R : ν : FS → R ⊆ BΨ (C α (S; Z)), S
where ν : FS → R is a finite signed regular Borel measure. (ii) Let M := Lp (Ω; Z) with p ∈ (1, ∞] be as in Example 2.5 (h) with (Z, ∥·∥Z ) having predual (E, ∥ · ∥E ) and with admissible collection of weight functions Ψ = (ψj )j=0,...,k . Then, an additive family is given by Z ′ k A := Lp (Ω; Z) ∋ x 7→ ⟨x(ω), g(ω)⟩Z×E dω ∈ R : g ∈ Lp (Ω; E) ⊆ BΨ (Lp (Ω; Z)), Ω ′
where 1/p + 1/p = 1. (iii) Let M := PψΩ (Ω) ⊆ MψΩ (Ω) be the space of probability measures over a weighted space (Ω, ψΩ ) as in Example 2.14 with admissible collection of weight functions Ψ = (ψj )j=0,...,k . Then, an additive family is given by Z k A := PψΩ (Ω) ∋ x 7→ f (ω)x(dω) ∈ R : f ∈ BψΩ (Ω) ⊆ BΨ (PψΩ (Ω)), Ω
where the function f ∈ BψΩ (Ω) could be replaced by a neural network φ : Ω → R.
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
33
Proof. For (i), we first apply [28, Theorem A.3] to obtain that C α (S; Z) ,→ C 0 (S; Z) is a compact embedding, where (C α (S; Z), τ∞ ) has ∥ · ∥α -BAP by Theorem B.1. While (A1) is satisfied, we use Dirac measures for to see that A is point separating on C α (S; Z). Hence, we can follow the proof of Lemma 4.6 to conclude that A is an additive family. For (ii), we use that Lp (Ω; Z) ∼ = Lq (Ω; E)∗ is a dual Banach space (see also Example 2.5 (h)). ′ Moreover, since (E, ∥·∥E ) has BAP, also (Lp (Ω; E), ∥·∥Lp′ (Ω;E) ) admits BAP. Hence, by using the ′ adjoints of the finite rank operators on Lp (Ω; E) as in Lemma 2.9, we conclude that (Lp (Ω; Z), τw∗ ) has ∥ · ∥Lp (Ω;Z) -BAP. Thus, we can apply Lemma 4.7 to obtain the conclusion. For (iii), we use that PψΩ (Ω) ⊆ MψΩ (Ω) ∼ = BψΩ (Ω)∗ is a subset of a dual Banach space (see also Example 2.5 (k)). Moreover, since (BψΩ (Ω), ∥ · ∥BψΩ (Ω) ) has BAP, we can use again the adjoints of the finite rank operators on BψΩ (Ω) as in Lemma 2.9 to conclude that (MψΩ (Ω), τw∗ ) has ∥ · ∥MψΩ (Ω) -BAP. Thus, we can apply Lemma 4.7 to obtain the conclusion. □ 4.3. Weighted UAT over finite-dimensional manifolds. Neural networks between Euclidean spaces enjoy the universal approximation property, meaning that they can approximate any continuous function uniformly on compact subsets. This fundamental result was first proven by G. Cybenko (see [30]) and K. Hornik (see [52]) in so-called universal approximation theorems (UATs), which establish denseness of neural networks in suitable functions spaces. Subsequently, other works [4, 15, 17] related the approximation error to the network complexity by proving quantitative approximation rates under more restrictive assumptions on the target function. In this section, we now prove a UAT for functional input neural networks on finite-dimensional manifolds. To this end, we assume that the activation function ρ : R → R is non-polynomial, cρ ∈ S ′ (R; C) in the sense of distribution has a non-zero point in its i.e., its Fourier transform T support. This is similar to the works with non-polynomial activation function of [20, 71, 92]. k We now introduce a weighted function space that is similar to BΨ (R) with polynomial weights k k Ψ. For c ∈ (0, ∞), we denote by Bc (R) the closure of Cb (R) with respect to the weighted norm (j)
|f (s)| k k ∥f ∥Bck (R) := maxj=0,...,k sups∈R (1+|s|) c . Then, Bc (R) can be related to BΨ (R) by viewing the (j)
(u)| derivatives as differentials. Moreover, ρ ∈ Bck (R) if and only if ρ ∈ C k (R) with lim|s|→∞ |ρ (1+|s|)c = 0, for all j = 0, ..., k (see [90, In addition, any ρ ∈ Bck (R) induces the tempered R Notation (v)]). distribution g 7→ Tρ (g) := R ρ(s)g(s)ds ∈ S ′ (R; C) (see, e.g., [37, Equation 9.26]).
Definition 4.9. For c ∈ (0, ∞), we introduce the following: cρ ∈ S ′ (R; C) has a non-zero (i) ρ ∈ Bck (R) is called non-polynomial if its Fourier transform T point in its support. cρ ∈ S ′ (R; C) has a (ii) ρ ∈ Bck (R) is called strongly non-polynomial if its Fourier transform T support with 0 ∈ R as inner point. cρ ∈ S ′ (R; C), we refer to Section 1.1. For the definition of the support of T First, we combine the weighted UATs of [90, Theorem 2.7] and [28, Proposition 4.4 (A3)] for classical neural networks on the real line. They both rely on Korevaar’s distributional extension [62] of Wiener’s Tauberian theorem [118], which provides sufficient conditions on the Fourier transform of a function such that the linear span of its translations is dense in L1 (R). More precisely, for ρea (s) := ρ(−as) and a finite signed measure µ on R, the condition (e ρa ∗ µ)(b) := R ρ(as − ab)µ(ds) = 0, for all a, b ∈ R, implies by Korevaar’s argument that µ = 0, which means R that the activation function ρ ∈ Bck (R) is discriminatory (cf., [20, 30, 71] for compactly supported measures µ). However, in order to include the approximation of the derivatives, the weighted UAT of [90, Theorem 2.7] followed the proof ideas of [52, 53] and mollified the linear functionals, which allows the application of integration to eliminate the derivatives.
34
P. SCHMOCKER AND J. TEICHMANN
Proposition 4.10. Let c ∈ (0, ∞). Then, the following holds true: (i) If ρ ∈ Bck (R) is strongly non-polynomial, then 0 ,ρ,R NNN := span {R ∋ s 7→ ρ(as + b) ∈ R : a ∈ N0 , b ∈ R} R,R
is a dense subset of Bck (R). (ii) If ρ ∈ Bck (R) is non-polynomial, then N N R,ρ,R R,R := span {R ∋ s 7→ ρ(as + b) ∈ R : a, b ∈ R} is a dense subset of Bck (R). Proof. Part (ii) follows directly from [90, Theorem 2.7]. For (i), we follow the proof of [90, Theorem 2.7] and replace the auxiliary result [90, Proposition 4.3] by the argument of [28, Proposition 4.4 (A3)]. The latter uses the strongly non-polynomial assumption to conclude from R ρ(as + b)µ(ds) = 0, for all a ∈ N0 and b ∈ R, that µ = 0 ∈ M(1+|·|)c (R). Hence, by conR tinuing the proof of [90, Theorem 2.7], we also obtain denseness in (i). □ Next, we lift the UAT from neural networks on the real line (see Proposition 4.10) to FNNs k defined on a weighted Cloc -manifold over finite-dimensional vector spaces. k k (M ; Y )). Let (M, Ψ) be a weighted Cloc -manifold Theorem 4.11 (Universal approximation on BΨ with atlas (Ui , ϕi )i∈I over finite-dimensional vector spaces (Xi , τXi )i∈I and let (Y, ∥ · ∥Y ) be a Banach space having BAP. Moreover, for c ∈ (0, ∞), let ρ ∈ Bck (R) be strongly non-polynomial k (M ) is an additive family such that for every a ∈ A and i ∈ I we have and assume that A ⊆ BΨ c π 1 + a ◦ ϕ−1 (u) d a ◦ ϕ−1 (u; vπ ) i i sup (4.3) Ca,i := max < ∞. j=0,...,k ψi,j (u, v1 , ..., vj ) (u,v1 ,...,vj )∈ϕi (Ui )×X j π∈P j
i
k In addition, let L ⊆ Y be a dense vector subspace. Then, N N A,ρ,L M,Y is a dense subset of BΨ (M ; Y ).
Proof. First, we show the conclusion for functional input neural networks N N A,ρ,L defined on a U,Y A,ρ,L k weighted domain (U, Ψ), which satisfy N N U,Y ⊆ BΨ (U ; Y ) by Lemma E.1. Now, we show that k G := span ({y cos0 (a(·)) : a ∈ A, y ∈ Y } ∪ {y sin(a(·)) : a ∈ A, y ∈ Y }) ⊆ BΨ (U ; Y ) k is dense in BΨ (U ; Y ). To this end, we follow the same arguments as in (3.8) to obtain G ′ := {ℓ ◦ g : ∗ k k ℓ ∈ Y , g ∈ G} ⊆ BΨ (U ) and therefore G ⊆ BΨ (U ; Y ). Moreover, since G ′ is by (3.9) a subalgebra and it holds that G ′ ⊗ Y ⊆ G, [94, Lemma 4.6] ensures that G is a polynomial subalgebra. In addition, we can follow the arguments below (3.9) to deduce that G ′ is point separating and nowhere vanishing of Ψ-moderate growth, where we use the constant map g(·) := cos(0) ∈ G ′ for the nowhere vanishing condition in (M2). Hence, we can now apply the weighted Nachbin theorem k (Theorem 3.7) to conclude that G is dense in BΨ (U ; Y ). Next, we prove that G is contained in the closure of N N A,ρ,L with respect to ∥ · ∥BΨk (U ;Y ) to U,Y k conclude from denseness of G that N N A,ρ,L is also dense in BΨ (U ; Y ). To this end, we fix some U,Y a ∈ A, y ∈ Y , and ε > 0. Then, by using that L is dense in Y , there exists some ye ∈ L such that ε . ∥y − ye∥Y < 2 1 + ∥ cos0 (a(·))∥BΨk (U ) N0 ,ρ,R Moreover, by applying the UAT in Proposition 4.10 (i), there exists φ0 ∈ N N R,R satisfying (j)
(4.4)
(j)
| cos0 (s) − φ0 (s)| ε < . j=0,...,k s∈R (1 + |s|)c 2Ca k!(1 + ∥e y ∥Y )
∥ cos0 −φ0 ∥Bck (R) = max sup
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
35
Hence, by using that A is closed under addition and that L is a vector space, we can define the FNN φ := yφ0 (a(·)) ∈ N N A,ρ,L U,Y . Thus, by using the Faà di Bruno formula, that |Pj | ≤ j! ≤ k!, the constant Ca > 0 defined in (4.3), and (4.4), it follows that ∥y cos0 (a(·)) − yφ0 (a(·))∥BΨk (U ;Y ) = ∥y cos0 (a(·)) − ye cos0 (a(·)) + ye cos0 (a(·)) − yeφ0 (a(·))∥BΨk (U ;Y ) y ∥Y ∥ cos0 (a(·)) − φ0 (a(·))∥BΨk (U ) ≤ ∥y − ye∥Y ∥ cos0 (a(·))∥BΨk (U ) + ∥e ε ∥ cos0 (a(·))∥Bk (U ) ≤ Ψ 2 1 + ∥ cos0 (a(·))∥BΨk (U ) P (|π|) (|π|) (a(u)) − φ0 (a(u)) |dπ a(u, vπ )| π∈Pj cos0 + ∥e y ∥Y max sup j=0,...,k (u,v ,...,v )∈U ×(Rd )j ψj (u, v1 , ..., vj ) 1 j (|π|)
≤
cos0 ε + Ca k!∥e y ∥Y max sup j=0,...,k 2 (u,v1 ,...,vj )∈U ×(Rd )j π∈P
(|π|)
(a(u)) − φ0 (a(u)) (1 + |a(u)|)c
j
(j)
(j)
cos0 (s) − φ0 (s) ε + Ca k!∥e y ∥Y max sup j=0,...,k 2 (1 + |s|)c s∈R ε ε < + Ca k!∥e y ∥Y ≤ ε. 2 2Ca k!(1 + ∥e y ∥Y )
≤
Since ε > 0 was chosen arbitrarily, this shows that y cos0 (a(·)) belongs to the closure of N N A,ρ,L U,Y with respect to ∥ · ∥BΨk (U ;Y ) , which holds analogously true for the map y sin(a(·)). Thus, we
conclude that the entire polynomial algebra G is contained in the closure of N N A,ρ,L with respect U,Y to ∥ · ∥BΨk (U ;Y ) . Therefore, by combining this with the previous step, i.e., that G is dense in
k k (U ; Y ). (U ; Y ), it follows that N N A,ρ,L is also dense in BΨ BΨ U,Y A,ρ,L k k -manifold (M, Ψ), we Finally, for the set of FNNs N N M,Y ⊆ BΨ (M ; Y ) over a weighted Cloc −1 k observe for every i ∈ I that Ai := a ◦ ϕi : a ∈ A ⊆ BΨ (ϕ (U )) is an additive family on ϕi (Ui ). i i i Ai ,ρ,L k k Hence, N N ϕi (Ui ),Y is by the previous step dense in BΨi (ϕi (Ui ); Y ). Thus, by using that BΨ (M ; Y ) k is equipped with the initial topology with respect to (2.11), N N A,ρ,L M,Y is dense in BΨ (M ; Y ).
□
k Remark 4.12. If A ⊆ BΨ (U ) is a vector subspace, then it suffices to assume that the activation k function ρ ∈ Bc (R) is non-polynomial (see Proposition 4.10 (ii)).
4.4. Weighted UAT over infinite-dimensional manifolds. In this section, we lift the UAT from finite-dimensional input spaces to infinite-dimensional input spaces by assuming the bounded approximation property (BAP). For more details on BAP, we refer to Section 1.1. k k Theorem 4.13 (Universal approximation on BΨ (M ; Y )). Let (M, Ψ) be a weighted Cloc -manifold with atlas (Ui , ϕi )i∈I over model spaces (Xi , τXi )i∈I such that each domain (ϕi (Ui ), τXi ) has BAP with finite rank operators (Ti,γ )γ ⊆ (Xi , τXi )∗ ⊗ Xi . Moreover, let (Y, ∥ · ∥Y ) be a Banach space having BAP. In addition, for c ∈ (0, ∞), let ρ ∈ Bck (R) be strongly non-polynomial and assume k that A ⊆ BΨ (M ) is an additive family such that for every a ∈ A and i ∈ I it holds that c π 1 + a ◦ ϕ−1 (u) d a ◦ ϕ−1 (u; vπ ) i i (4.5) Ca,i := max sup <∞ j=0,...,k ψi,j (u, v1 , ..., vj ) (u,v1 ,...,vj )∈ϕi (Ui )×X j π∈P j
i
and for every a ∈ A, b ∈ R, i ∈ I, and γ the composition (4.6)
ϕi (Ui ) ∋ u
7→
ρ
a ◦ ϕ−1 i ◦ Ti,γ (u) + b ∈ R
i ,ρ,R belongs to the closure of N N A k (ϕ (U )) . Furthermore, let L ⊆ Y be a ϕi (Ui ),R with respect to ∥ · ∥BΨ i i i
k dense vector subspace. Then, N N A,ρ,L M,Y is a dense subset of BΨ (M ; Y ).
36
P. SCHMOCKER AND J. TEICHMANN
Proof. First, we show the conclusion for functional input neural networks N N A,ρ,L defined on a U,Y weighted domain (U, Ψ) such that (U, τX ) has BAP with finite rank operators (Tγ )γ ⊆ (X, τX )∗ ⊗ k k X. Note that Lemma E.1 ensures that N N A,ρ,L ⊆ BΨ (U ; Y ). Since BΨ (U ) ⊗ Y is by Lemma 3.9 U,Y k k k dense in BΨ (U ; Y ) and BΨ (U ) is defined as closure of Cb (U ) with respect to ∥ · ∥BΨk (U ) , it suffices
to approximate any given f ∈ Cbk (U ) ⊗ Y by an element of N N A,ρ,L U,Y . To this end, we fix some PN f := n=1 fn (·)yn ∈ Cbk (U ) ⊗ Y and ε > 0, where N ∈ N, f1 , ..., fN ∈ Cbk (U ), and y1 , ..., yN ∈ Y . Then, by using that L is dense in Y , there exists some ye1 , ..., yeN ∈ L such that ε . (4.7) ∥yn − yen ∥Y < 4N 1 + ∥fn ∥BΨk (U )
Moreover, by using that (U, τX ) has BAP with finite rank operators (Tγ )γ ⊆ (X, τX )∗ ⊗ X, there exists by Lemma 3.11 some γ such that for every n = 1, ..., N it holds that ε (4.8) ∥fn − fn ◦ Tγ ∥BΨk (U ) < . 4N (1 + ∥e yn ∥Y ) Now, we define the collection Ψγ := (ψγ,j )j=0,...,k of weight functions ψγ,j : Tγ (U ) × Tγ (X)j → k (0, ∞) as in (3.13) and observe that A|Tγ (U ) ⊆ BΨ (Tγ (U )) is an additive family on Tγ (U ) γ satisfying (4.3). Hence, for every fixed n = 1, ..., N , the UAT in Theorem 4.11 applied to A|T (U ) ,ρ,R
k fn |Tγ (U ) ∈ BΨ (Tγ (U )) ensures the existence of some φn ∈ N N Tγ (Uγ ),R γ
∥fn |Tγ (U ) − φn ∥BΨk (Tγ (U )) < γ
satisfying
ε . 4N (1 + ∥e yn ∥Y )
Thus, by applying the chain rule as in (3.14), it follows that ∥fn ◦ Tγ − φn ◦ Tγ ∥BΨk (U ) <
(4.9)
ε . 4N (1 + ∥e yn ∥Y )
Now, we use that φn ◦ Tγ ∈ N N A,ρ,R belongs by assumption to the closure of N N A,ρ,R with U,R U,R A,ρ,R respect to ∥ · ∥BΨk (U ) to obtain some φ en ∈ N N U,R such that ∥φn ◦ Tγ − φ en ∥BΨk (U ) <
(4.10) Hence, for φ e :=
en φ en ∈ N N A,ρ,L U,Y , we use (4.7), (4.8), (4.9), and (4.10) to conclude that n=1 y
PN
∥f − φ∥ e BΨk (U ;Y ) =
N X
yn fn (·) −
n=1
<
N X
ε . 4N (1 + ∥e yn ∥Y )
∥yn − yen ∥Y ∥fn ∥BΨk (U ) +
n=1
N X n=1
N X
yen φ en (·) k (U ;Y ) BΨ
∥e yn ∥Y ∥fn − φ en ∥BΨk (U )
n=1
≤
N ε X + ∥e yn ∥Y ∥fn − fn ◦ Tγ ∥BΨk (U ) + ∥fn ◦ Tγ − φn ◦ Tγ ∥BΨk (U ) + ∥φn ◦ Tγ − φ en ∥BΨk (U ) 4 n=1
≤
ε 3ε + = ε. 4 4
k Since ε > 0 and f ∈ Cbk (U ) ⊗ Y were chosen arbitrarily and Cbk (U ) ⊗ Y is dense in BΨ (U ; Y ), this A,ρ,L k shows that N N U,Y is dense in BΨ (U ; Y ). k Finally, for general functional input neural networks N N A,ρ,L M,Y over a weighted Cloc -manifold k (M, Ψ), we observe for every fixed i ∈ I that Ai := a ◦ ϕ−1 : a ∈ A ⊆ BΨ (ϕi (Ui )) is an additive i i i ,ρ,R family such that every composition of the form (4.6) belongs to the closure of N N A ϕi (Ui ),R with
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
37
i ,ρ,L respect to ∥ · ∥BΨk (ϕi (Ui )) . Hence, we can apply the previous step to conclude that N N A ϕi (Ui ),Y is i
k k dense in BΨ (ϕi (Ui ); Y ). Thus, by using that BΨ (M ; Y ) is equipped with the initial topology with i A,ρ,L k respect to (2.11), it follows that N N M,Y is dense in BΨ (M ; Y ). □ k Remark 4.14. If A ⊆ BΨ (M ) is a vector subspace, then it suffices to assume that ρ ∈ Bck (R) is non-polynomial (see Proposition 4.10 (ii)).
5. Weighted universal approximation of non-anticipative functionals In this section, we apply the weighted universal approximation theorem (UAT) in Theorem 4.13 to non-anticipative functionals, which extends the universal approximation result for continuous non-anticipative functionals in [28, Corollary 4.17] by including the directional derivatives. Nonanticipative functional calculus was originally introduced in [22,23,34] to extend Föllmer’s pathwise Ito-calculus (see [38]) to path-dependent functionals. First, we recall some notions of non-anticipative functional calculus (see also [21, Section 5.1]). For a fixed terminal time T ∈ (0, ∞) and a dual Banach space (Z, ∥·∥Z ), we define the stopped path of x ∈ D0 ([0, T ]; Z) at time t ∈ [0, T ] as (s 7→ xts := xs∧t ) ∈ D0 ([0, T ]; Z), where s ∧ t := min(s, t). Note that we adopt in this section the notation x := (xs )s∈[0,T ] , where time s ∈ [0, T ] is now indicated as a subscript. Then, the space of stopped Z-valued càdlàg paths is defined as Λ0 = (t, xt ) : (t, x) ∈ [0, T ] × D0 ([0, T ]; Z) ∼ = [0, T ] × D0 ([0, T ]; Z) / ∼, T,Z
with (t, x) ∼ (s, y) if and only if t = s and xt = y s . Moreover, (Λ0T,Z , d∞ ) with metric d∞ ((t, x), (s, y)) = |t − s| + sup u∈[0,T ]
xtu − yus Z
is a complete metric space (see [21, p. 131]). In addition, for (α, r) ∈ (0, 1) ∈ (0, ∞), we denote 0 by Λα,r T,Z ⊆ ΛT,Z the subspace of stopped α-Hölder càdlàg paths with r-summable jumps. Following the original definitions in [22, 23, 34], the space of stopped paths Λ0T,Z can be also seen as vector bundle. More precisely, we define the space of stopped Z-valued α-Hölder càdlàg paths with r-summable jumps as the vector bundle [ (5.1) Λα,r Dα,r ([0, t], Z), T,Z := t∈(0,T ) α,r over the base space (0, T ) with bundle projection Λα,r T,Z ∋ (t, x) 7→ πΛT ,Z (t, x) := t ∈ (0, T ) and −1 α,r fibers π ({t}) = D ([0, t], Z). For technical reasons, we restrict ourselves here to an open interval (0, T ) instead of [0, T ], which therefore does not include any jump at terminal time ∞ T . In this case, (Λα,r T,Z , d∞ ) is a Cloc -manifold over the locally convex topological vector space (R × Dα,r ([0, T ]; Z), τR × τw∗ ) equipped with the product topology of the topology τR on R and the weak-∗-topology τw∗ on Dα,r ([0, T ]; Z). In addition, the global chart is given by
(5.2)
Λα,r T,Z ∋ (t, x)
7→
ϕi (t, x) := (t, xt ) ∈ R × Dα,r ([0, T ]; Z),
with extension xt ∈ Dα,r ([0, T ]; Z) defined as xts = xs if s ∈ [0, t], and xts = xt if s ∈ (t, T ]. The inverse of the global chart (5.2) is equal to ϕi (Λα,r T,Z ) ∋ (t, x)
7→
α,r ϕ−1 i (t, x) := (t, x|[0,t] ) ∈ ΛT,Z .
Furthermore, by using local trivializations of the vector bundle, one can show that the higher j j α,r ∼ α,r order tangent spaces at any point (t, x) ∈ Λα,r ([0, t]; Z) , T,Z are given by T(t,x) ΛT,Z = R × D j ∈ N0 , and the higher order tangent bundles are equal to n o j α,r j j α,r T j Λα,r = (t, x), [c] : (t, x) ∈ Λ , [c] ∈ T Λ j ∈ N0 . T,Z T,Z (t,x) (t,x) (t,x) T,Z ,
38
P. SCHMOCKER AND J. TEICHMANN
We shall use these higher order tangent bundles to define a weighted manifold. 5.1. Non-anticipative path-neural network. In this section, we introduce a special type of functional input neural networks, so-called non-anticipative path-neural networks, to approximate a differentiable non-anticipative functional. To this end, we first introduce non-anticipative functionals as measurable maps from Λα,r T,Z to a Banach space (Y, ∥ · ∥Y ) as output space. α,r Definition 5.1. A map f : Λα,r T,Z → Y is called a non-anticipative functional if f : ΛT,Z → Y is α,r a measurable map from (ΛT,Z , d∞ ) to (Y, ∥ · ∥Y ).
This notion of causality arises in many physical phenomena and in control theory (see, e.g., [36]). α,r Moreover, a non-anticipative functional f : Λα,r T,Z → Y is said to be continuous if f : ΛT,Z → Y is α,r k a continuous map from (ΛT,Z , d∞ ) to (Y, ∥ · ∥Y ). In addition, we recall that f ∈ Cloc (Λα,r T,Z ; Y ) is α,r −1 k k k a Cloc -map if and only if f ◦ ϕi ∈ Cloc (ϕi (ΛT,Z ); Y ) is a Cloc -map on the model space. While we consider in Section 5.2 the approximation of all directional derivatives, we shall restrict ourselves in Section 5.3 to the horizontal and vertical derivatives. Now, we introduce non-anticipative path-neural networks (path-NNs). To this end, we assume that (E, ∥ · ∥E ) is a predual for the dual Banach space (Z, ∥ · ∥Z ). Definition 5.2. For ρ, ρe ∈ C k (R) and a vector subspace L ⊆ Y , we define a non-anticipative path-neural network (path-NN) φ : Λα,r T,Z → Y as ! Z T N X α,r t (5.3) ΛT,Z ∋ (t, x) 7→ φ(t, x) = yn ρ λn t + en (s)⟩Z×E ds + bn ∈ Y, ⟨xs , φ n=1
0
where N ∈ N denotes the number of neurons, where λ1 , ..., λN ∈ R are the weights, where b1 , ..., bN ∈ R are the biases, where y1 , ..., yN ∈ L are the linear readouts, where ρ,E φ e1 , ..., φ eN ∈ N N R,e e(as + b) : a, b ∈ R, e ∈ E} R,E := span {R ∋ s 7→ e ρ
are E-valued neural networks, and where ρ, ρe ∈ C k (R) represent the activation functions. Moreρ e,ρ,L the set of path-NNs of the form (5.3). over, we denote by PN Λ α,r ,Y T ,Z
Remark 5.3. Let us point out the following remarks concerning Definition 5.2: RT Rt (i) In (5.3), we can rewrite the integral as 0 ⟨xts , φ en (s)⟩Z×E ds = 0 ⟨xs , φ en (s)⟩Z×E ds + RT xt , t φ en (s)ds Z×E , which shows the non-anticipative behaviour of a path-NN. ρ,E (ii) The set of linear readouts E used for the E-valued neural networks N N R,e could be R,E also replaced by a dense vector subspace LE ⊆ E.
5.2. Weighted UAT for differentiable non-anticipative functionals. We now apply the weighted universal approximation theorem (UAT) in Theorem 4.13 to establish a universal approximation result for non-anticipative path-neural networks (path-NNs) over the space of stopped α-Hölder càdlàg paths, including the approximation of all possible directional derivatives. On the space of stopped α-Hölder càdlàg paths Λα,r T,Z introduced in (5.1), we define the collection Ψ := (ψj )j=0,...,k of weight functions ψj : T j Λα,r → (0, ∞) by T,Z ! j X j (ℓ) (5.4) ψj (t, x), [c](t,x) := η max(j, 1)∥(t, x)∥R×Dα,r ([0,t];Z) + ∥c (0)∥R×Dα,r ([0,t];Z) , ℓ=1
(t, x), [c]j(t,x) ∈ T j Λα,r T,Z
for and some continuous non-decreasing function η : [0, ∞) → (0, ∞), where ∥(t, x)∥R×Dα,r ([0,t];Z) := |t| + ∥x∥α,ℓr . Hence, the collection Ψi = (ψi,j )j=0,...,k of pushj α,r forward weight functions ψi,j := ψj ◦ Φ−j : ϕi (Λα,r ([0, T ]; Z) → (0, ∞) is given i T,Z ) × R × D
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
39
by ! j X t ψi,j (t,x ), (s1 ,v1 ), ..., (sj ,vj ) = η max(j, 1)∥(t,x )∥R×Dα,r ([0,T ];Z) + ∥(sℓ ,vℓ )∥R×Dα,r ([0,T ];Z) , t
ℓ=1
j α,r for (t, x ), (s1 , v1 ), ..., (sj , vj ) ∈ ϕi (Λα,r ([0, T ]; Z) . Now, we show that (Λα,r T,Z , Ψ) T,Z ) × R × D k is a weighted Cloc -manifold having BAP and that the additive family of the path-NNs introduced in Definition 5.2 indeed satisfies (A1)-(A4). The proof is given in Appendix F.1. t
Lemma 5.4. Let the predual (E, ∥·∥E ) of (Z, ∥·∥Z ) have BAP and assume that η : [0, ∞) → (0, ∞) rk k is a continuous and increasing function with limr→∞ η(r) = 0. Then, (Λα,r T,Z , Ψ) is a weighted Cloc manifold over (R × Dα,r ([0, T ]; Z), τR × τ∞ ) having BAP, where Ψ satisfies (2.2). Moreover, for non-polynomial ρe ∈ C 0 (R), the set ( ) Z T α,r R,e ρ,E t A := ΛT,Z ∋ (t, x) 7→ λt + ⟨xs , φ(s)⟩ e e ∈ N N R,E Z×E ds ∈ R : λ ∈ R, φ 0
α,r k (Λα,r is a vector subspace of BΨ T,Z ) and an additive family on ΛT,Z .
Now, we apply the weighted UAT for FNNs in Theorem 4.13 to obtain the following UAT for non-anticipative path-NNs on Λα,r T,Z . The proof can be found in Appendix F.2. k (Λα,r Corollary 5.5 (Universal Approximation on BΨ T,Z ; Y )). Let the predual (E, ∥·∥E ) of (Z, ∥·∥Z ) have BAP, let (Y, ∥ · ∥Y ) be a Banach space having BAP, and assume that L ⊆ Y is a dense vector subspace. Moreover, for c ∈ (0, ∞), let ρe, ρ ∈ Bck+1 (R) be non-polynomial with bounded derivatives k max(1,c) and assume that η : [0, ∞) → (0, ∞) is continuous and non-decreasing with limr→∞ r η(r) = 0. α,r ρ e,ρ,L k k (Λα,r is a dense subset of BΨ Then, PN Λ α,r T,Z ; Y ), i.e., for every f ∈ BΨ (ΛT,Z ; Y ) and ε > 0 ,Y T ,Z
ρ e,ρ,L there exists some φ ∈ PN Λ such that α,r ,Y T ,Z
:= max ∥f − φ∥BΨk (Λα,r T ,Z ;Y )
j=0,...,k
sup ((t,x),[c]j(t,x) )∈T j Λα,r T ,Z
(f ◦ c)(j) (0) − (φ ◦ c)(j) (0) Y < ε. ψj (t, x), [c]j(t,x)
5.3. Weighted UAT for horizontal and vertical derivatives. In this section, we present a universal approximation theorem (UAT) for non-anticipative functionals, which only includes the approximation of the horizontal and vertical derivatives over Rd . For some fixed k, ℓ ∈ N0 and α, T ∈ (0, ∞) as well as the Euclidean space Z := Rd , we consider again the space of stopped α-Hölder càdlàg paths in Λα,r . However, we equip Λα,r with the T,Rd T,Rd single weight function of the form (5.5)
Λα,r ∋ (t, x) T,Rd
7→
ψ(t, x) := η (∥x∥α,ℓr ) ∈ (0, ∞),
for some continuous non-decreasing function η : [0, ∞) → (0, ∞). Compared to (5.4) this weight function does no longer depend on the derivatives as we only consider some derivatives in particular directions, which have uniformly bounded norms. Definition 5.6. A non-anticipative functional f : Λα,r → Y is called T,Rd t
t
(i) horizontally differentiable if the limit Df (t, x) := limh→0+ f (t+h,x h)−f (t,x ) exists in (Y, ∥ · ∥Y ), for all (t, x) ∈ Λα,r (see [21, Definition 5.7]). T,Rd f (t,xt +h1
e )−f (t,xt )
[t,T ] i (ii) vertically differentiable if the limit Dei f (t, x) := limh→0 h α,r (Y, ∥ · ∥Y ), for all i = 1, ..., d and (t, x) ∈ ΛT,Rd (see [21, Definition 5.8]).
exists in
40
P. SCHMOCKER AND J. TEICHMANN
Moreover, the higher order horizontal derivatives Dj f (t, x), j ∈ N0 , as well as the higher order vertical derivatives Dβ f (t, x), β ∈ Nd0 , are defined by iteration. α,r For k, ℓ ∈ N0 , we denote by Ck,ℓ b (ΛT,Rd ; Y ) the vector space of bounded continuous nonanticipative functionals f : Λα,r → Y that are k-times horizontally differentiable with bounded T,Rd continuous horizontal derivatives Dj f : Λα,r → Y , j = 0, ..., k, and ℓ-times vertically differenT,Rd tiable with bounded continuous vertical derivatives Dβ f : Λα,r → Y , α ∈ Nd0,ℓ . Then, we define T,Rd α,r k,ℓ α,r Bk,ℓ ψ (ΛT,Rd ; Y ) as the closure of Cb (ΛT,Rd ; Y ) with respect to the weighted norm j ∥D f (t, x)∥ ∥D f (t, x)∥ β Y Y . , max sup ∥f ∥Bk,ℓ (Λα,r ;Y ) := max max sup α,r ψ j=0,...,k (t,x)∈Λα,r ψ(t, x) ψ(t, x) T ,Rd β∈Nd (t,x)∈Λ 0,ℓ d d T ,R
T ,R
Now, we show the following universal approximation theorem for non-anticipative functionals in α,r Bk,ℓ ψ (ΛT,Rd ; Y ), where only the approximation of the horizontal and vertical derivatives is included. The proof can be found in Appendix F.3. α,r Corollary 5.7 (Universal Approximation on Bk,ℓ ψ (ΛT,Rd ; Y )). Let (Y, ∥ · ∥Y ) be a Banach space having BAP and let L ⊆ Y be a dense vector subspace. Moreover, for c ∈ (0, ∞), let ρe, ρ ∈ Bck+1 (R) be non-polynomial with bounded derivatives and assume that η : [0, ∞) → (0, ∞) is max(1,k,c) ρ e,ρ,L continuous and non-decreasing with limr→∞ r η(r) = 0. Then, PN Λ is a dense subset of α,r ,Y T ,Rd
α,r k,ℓ α,r ρ e,ρ,L Bk,ℓ ψ (ΛT,Rd ; Y ), i.e., for every f ∈ Bψ (ΛT,Rd ; Y ) and ε > 0 there exists φ ∈ PN Λα,r ,Y such that T ,Rd
∥f − φ∥Bk,ℓ (Λα,r ;Y ) < ε. ψ
T ,Rd
Corollary 5.7 is of particular interest for applications involving the path-dependent Ito-formula (see, e.g., [22, 23, 34]). Indeed, in this case, only the horizontal and vertical derivatives appear, which can be approximated with non-anticipative path-neural networks (PNNs). 6. Weighted universal approximation of linear functions of the signature In this section, we present an application of the weighted Nachbin theorem (Theorem 3.12) to approximate path space functionals, which is similar to Section 5, but using linear functions of the signature instead of non-anticipative functionals. The notion of the signature was introduced by K.-T. Chen in [19] and plays a central role in rough path theory developed by T. Lyons in [78] (see also the textbooks [39, 40]). Let us assume that the input data is sequentially ordered, representing a discretization of a path X : [0, T ] → Z with values in a Banach space (Z, ∥ · ∥Z ), e.g., the motion of a plane in the airspace depending on time, the evolution of temperature or pressure measured by a sensor, or the stock prices in a financial market. Given a continuous path X : [0, T ] → Z of finite variation, we define its signature (at terminal time) as the infinite collection of iterated integrals Z Z S(X)T := 1, dXu1 , · · · dXu1 ⊗ · · · ⊗ dXuN , · · · ∈ T ((Z)), 0<u1 <T
Q∞
⊗n
0<u1 <...<uN <T
where T ((Z)) := n=0 Z denotes the extended tensor algebra (see Section 6.1 below). For paths of lower regularity, e.g., α-Hölder continuous paths X : [0, T ] → Z, one relies on the theory of rough paths to define its signature. In this case, a linear function of the signature (at terminal time T ) are linear combinations of continuous linear functionals of the components of S(X)T . In the following, we show that (non-linear) path space functionals can be approximated by linear functions of the signatures over the whole path space, which extends the global universal approximation theorem (UAT) in [28, Theorem 5.4] by including the approximation of the derivatives.
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
41
This in turn generalizes the UATs (without derivatives) on compact subsets of the path space, e.g., for finite variation paths or for continuous functions of the whole signature (see [72, Theorem 3.1], [61, Theorem 1], and [77, Section 3]) and for càdlàg paths (see [27, Theorem 3.13]). More recently, UATs have been established on the entire path space in an Lp -sense (see [7, 18]), extended to uniform approximation over the whole time interval rather than at a fixed terminal time (see [6, 26]), and further generalized to infinite-dimensional rough path settings (see [24]). To establish the universality of linear functions of the signature, we apply the weighted Nachbin theorem over infinite-dimensional manifolds (Theorem 5.10), which relies on the following key k features of the signature. First, linear functions of the signature are Cloc -maps on the underlying rough path space with suitable growth conditions. Second, the signature (at terminal time) uniquely determines the path (up to so-called tree-like equivalences, see [14, 49]), ensuring point separation. Third, using the shuffle product, any product of linear functions of the signature can be realized as another linear function of the signature, which asserts the algebra property. 6.1. Notation related to the signature of (rough) paths. We now recall the most important notions. For a dual Banach space (Z, ∥ · ∥Z ) with predual (E, ∥ · ∥E ), we assume that ∥ · ∥Z ⊗n is a norm on the n-th algebraic tensor product Z ⊗a n , n ∈ N0 , with Z ⊗a 0 := R, satisfying ∥a ⊗ b∥Z ⊗(m+n) ≤ ∥a∥Z ⊗m ∥b∥Z ⊗n ,
(6.1) for all m, n ∈ N0 , a ∈ Z (6.2)
⊗a m
, and b ∈ Z ⊗a n , and that
∥z1 ⊗ · · · ⊗ zn ∥Z ⊗n = ∥zσ(1) ⊗ · · · ⊗ zσ(n) ∥Z ⊗n
for all n ∈ N0 , σ ∈ Sn , and z1 , ..., zn ∈ Z. Then, for any m, n ∈ N0 , we define Z ⊗m ⊗ Z ⊗n as the completion of the algebraic tensor product Z ⊗m ⊗a Z ⊗n with respect to ∥ · ∥Z ⊗(m+n) , which ensures that Z ⊗m ⊗ Z ⊗n ∼ = Z ⊗(m+n) are isomorphic as Banach spaces. For example, the injective tensor norm satisfies the two properties (6.1)–(6.2) (see, e.g., [99]). Moreover, we assume for every n ∈ N0 that E ⊗n is a predual for Z ⊗n , which is, e.g., satisfied if E ⊗n is equipped with the projective tensor norm and Z ⊗n with the injective tensor norm (see, e.g., [99, Theorem 2.9]). Then, the extended tensor algebra (over Z) is defined as ∞ Y T ((Z)) := Z ⊗n , n=0
which is endowed with the addition, tensor multiplication, and scalar multiplication defined by !∞ ∞ ∞ ∞ X (n) (n) (n−k) (k) a + b := a + b , a ⊗ b := a ⊗b , λ · a := λa(n) , n=0
n=0
k=0
n=0
(n) ∞ for a := (a(n) )∞ )n=0 ∈ T ((Z)), and λ ∈ R. Moreover, for N ∈ N0 , the n=0 ∈ T ((Z)), b := (b truncated tensor algebra is defined as N Y T N ((Z)) := Z ⊗n , n=0
where addition “+”, tensor multiplication “⊗”, and scalar multiplication “·” defined by !N n N N X (n) (n) (n−k) (k) a + b := a + b , a ⊗ b := a ⊗b , λ · a := λa(n) n=0
k=0
,
n=0 n=0
N (n) N for a := (a(n) )N )n=0 ∈ T N (Z), and λ ∈ R. We equip T N (Z) with the n=0 ∈ T (Z), b := (b (n) N norm ∥a∥T N (Z) := maxn=0,...,N ∥a ∥Z ⊗n , for a := (a(n) )N n=0 ∈ T (Z). In addition, we introduce N the subsets T0N (Z) and T1N (Z) of T N (Z) consisting of elements a := (a(n) )∞ n=0 ∈ T (Z) with (0) (0) a = 0 and a = 1, respectively.
42
P. SCHMOCKER AND J. TEICHMANN
In order to adapt the Lie group point of view on weakly geometric rough paths, we observe that T1N (Z) is a Lie group under ⊗, truncated at level N , with unit element 1 := (1, 0, ..., 0) ∈ T1N (Z). LN Moreover, for any N ∈ N, we define the free step-N nilpotent Lie algebra as gN (Z) := n=0 Ln , with homogeneous Lie polynomials Ln ⊆ T N (Z) of degree n recursively defined by L0 := 0,
L1 := Z,
L2 := [Z, L1 ] = [Z, Z],
L3 := [Z, L2 ] = [Z, [Z, Z]],
...,
where [a, b] := a ⊗ b − b ⊗ a is the Lie bracket, with [A, B] := span({[a, b] : a ∈ A, b ∈ B}). Note that Ln ⊆ Z ⊗n is a vector subspace, ensuring that gN (Z) ⊆ T N (Z) is a vector subspace. In addition, we define the exponential map as (6.3)
T0N (Z) ∋ a
N
7→
exp (a) := 1 +
N X 1 n=1
n!
a⊗n ∈ T1N (Z),
whose inverse is given by the logarithm T1N (Z) ∋ 1 + b
7→
logN (1 + b) :=
N X (−1)n+1 n=1
n
b⊗n ∈ T0N (Z).
From this, we define the free step-N nilpotent Lie group GN (Z) := expN (gN (Z)), which we en1/n dow with the homogeneous norm ∥g∥GN (Z) := maxn=1,...,N ∥g(n) ∥Z ⊗n , inducing the homogeneous metric dGN (Z) (g, h) := ∥g−1 ⊗ h∥GN (Z) , for g, h ∈ GN (Z). Then, GN (Z) is a subgroup of T1N (Z) ∞ -manifold with global chart logN : GN (Z) → gN (Z) over the model space gN (Z). and a Cloc Moreover, the truncated signature at level N of a path X ∈ C 0 ([0, T ]; Z) of finite variation is defined by S N (X)s,t := 1, S (1) (X)s,t , ..., S (N ) (X)s,t Z Z = 1, dXu1 , ..., dXu1 ⊗ · · · ⊗ dXuN , s<u1 <t
s<u1 <...<uN <t N
for 0 ≤ s ≤ t ≤ T , which takes values in G (Z). In addition, the (entire) signature of a path X ∈ C 0 ([0, T ]; Z) of finite variation is defined by S(X)s,t := 1, S (1) (X)s,t , S (2) (X)s,t , ... Z Z = 1, dXu1 , dXu1 ⊗ dXu2 , ... , s<u1 <t
s<u1 <u2 <t
which takes values in the set of group-like elements n o G(Z) := g ∈ T ((Z)) : g(0) , ..., g(N ) ∈ GN (Z) for all N ∈ N0 . Furthermore, for any m, n ∈ N, e := e1 ⊗ · · · ⊗ em ∈ E ⊗m , and ee := em+1 ⊗ · · · ⊗ em+n ∈ E ⊗n , we define the shuffle product X (6.4) e ee := eσ−1 (1) ⊗ · · · ⊗ eσ−1 (m+n) ,
σ∈Sh(m,n)
where Sh(m, n) consists of shuffles σ ∈ Sm+n of {1, ..., m} and {m + 1, ..., m + n}, i.e., σ ∈ Sm+n satisfying σ(1) < ... < σ(m) and σ(m + 1) < ... < σ(m + n). In particular, for every e := e1 ⊗ · · · ⊗ em ∈ E ⊗m , ee := em+1 ⊗ · · · ⊗ em+n ∈ E ⊗n , and g ∈ GN (Z) with m + n ≤ N , we have (6.5)
⟨g(m) , e⟩Z ⊗m ×E ⊗m · ⟨g(n) , ee⟩Z ⊗n ×E ⊗n = ⟨g(m+n) , e
ee⟩
Z ⊗(m+n) ×E ⊗(m+n) ,
which is referred to as the shuffle product property (see [78, Theorem 2.15]).
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
43
6.2. Manifold of weakly geometric α-Hölder rough paths. We now introduce weakly geometric α-Hölder rough paths with values in a dual Banach space (Z, ∥·∥Z ) having predual (E, ∥·∥E ), which can be seen as α-Hölder continuous paths with values in G⌊1/α⌋ (Z) (see also [39, 40]). Definition 6.1. For α ∈ (0, 1], a continuous path X : [0, T ] → G⌊1/α⌋ (Z) of the form (2) (⌊1/α⌋) [0, T ] ∋ t 7→ Xt := 1, Xt , Xt , ..., Xt ∈ G⌊1/α⌋ (Z) with X0 := 1 := (1, 0, ..., 0) ∈ G⌊1/α⌋ (Z) is called a weakly geometric α-Hölder rough path if 1
(n) n dG⌊1/α⌋ (Z) (Xs , Xt ) maxn=1,...,⌊1/α⌋ (X−1 s ⊗ Xt ) Z ⊗n ∥X∥α := sup := sup < ∞. α α |s − t| |s − t| 0≤s<t≤T 0≤s<t≤T
We denote by CTα (Z) := C1α ([0, T ]; G⌊1/α⌋ (Z)) the space of weakly geometric α-Hölder rough paths, which we equip with the w∗ -uniform topology generated by the semi-metrics 1
d∞,e (X, Y) := sup dG⌊1/α⌋ (Z),e (Xt , Yt ) := sup
max
t∈[0,T ] n=1,...,⌊1/α⌋
t∈[0,T ]
for e := (e(1) , ..., e(⌊1/α⌋) ) ∈
L⌊1/α⌋ n=1
(n) ⟨(X−1 , en ⟩Z ⊗n ×E ⊗n n , t ⊗ Yt )
E ⊗n .
Next, we define the truncated signature at level N > ⌊1/α⌋ of a weakly geometric α-Hölder rough path X ∈ CTα (Z) as the unique Lyons extension yielding a path S N (X) : [0, T ] → GN (Z) with finite α-Hölder norm ∥ · ∥α whose n-th component agrees with X(n) , for all n = 0, ..., ⌊1/α⌋ (see [78, Theorem 3.7] and [40, Corollary 9.11 (ii)]). By denoting the n-th signature component taking values in Z ⊗n by S (n) (X), the signature of X ∈ CTα (Z) is defined by [0, T ] ∋ t 7→ S(X)t := 1, Xt , S (2) (X)0,t , S (3) (X)0,t , ... ∈ G(Z). Then, a linear function of the signature (at time T ) is given as CTα (Z) ∋ X
7→
ℓ(S(X)T ),
PN
where a 7→ ℓ(a) := n=0 ⟨a(n) , en ⟩Z ⊗n ×E ⊗n , for some N ∈ N and en ∈ E ⊗n , n = 0, ..., N . Moreover, we use the bijection log⌊1/α⌋ : G⌊1/α⌋ (Z) → g⌊1/α⌋ (Z) to observe that ( α α ⌊1/α⌋ (Z)) → Cα CT (Z) T (Z) := C0 ([0, T ]; g (6.6) ϕi : ⌊1/α⌋ ⌊1/α⌋ X 7→ log (X) := t 7→ log (Xt ) is a bijection onto its image, whose inverse is given by ( ϕi (CTα (Z)) → CTα (Z) −1 . (6.7) ϕi : Y 7→ exp⌊1/α⌋ (Y) := t 7→ exp⌊1/α⌋ (Yt ) L⌊1/α⌋ Since (Z, ∥ · ∥Z ) is a dual Banach space and g⌊1/α⌋ (Z) = n=1 Ln with weak-∗-closed Ln ⊆ Z ⊗n , the free step-⌊1/α⌋ nilpotent Lie algebra g⌊1/α⌋ (Z) has also a predual. Hence, we can equip Cα T (Z) with the w∗ -uniform topology τ∞ generated by seminorms of the form pe (Y) = sup
max
t∈[0,T ] n=1,...,⌊1/α⌋
(n)
⟨Yt , e(n) ⟩Z ⊗n ×E ⊗n ,
L⌊1/α⌋ for all e := (e(1) , ..., e(⌊1/α⌋) ) ∈ n=1 E ⊗n . ∞ Now, we observe that (CTα (Z), τ∞ ) is a Cloc -manifold with global chart (6.6) over the model α space (CT (Z), τ∞ ). This is contrast to considering CTα (Z) as submanifold of C α ([0, T ]; T ⌊1/α⌋ (Z)), which requires like G⌊1/α⌋ (Z) ,→ T ⌊1/α⌋ (Z) infinitely many charts. In our case, the higher order
44
P. SCHMOCKER AND J. TEICHMANN
j α j tangent spaces at any point X ∈ CTα (Z) are given by TX CT (Z) ∼ = Cα T (Z) , for all j ∈ N, and the higher order tangent bundles are equal to
T j CTα (Z) =
n o j α X, [c]jX : X ∈ CTα (Z), [c]jX ∈ TX CT (Z) ,
j ∈ N.
Furthermore, we fix some k ∈ N0 and define the collection Ψ := (ψj )j=0,...,k of weight functions ψj : T j CTα (Z) → (0, ∞) by
(6.8)
ψj (X, [c]jX ) := exp
β1 max(j, 1)∥ log
⌊1/α⌋
(X)∥cα1 + β2
j X
! ∥(log
⌊1/α⌋
(ℓ)
◦c)
(0)∥cα2
,
ℓ=1
for (X, [c]jX ) ∈ T j CTα (Z), with β1 , β2 > 0, c1 ≥ ⌊1/α⌋, and c2 > 0. Here, we use the chart ϕi := log⌊1/α⌋ in the term ∥ log⌊1/α⌋ (X)∥α instead of ∥X∥α as in [28, Section 5.2] considering the case without derivatives. This simplifies the collection Ψi := (ψi,j )j=0,...,k of push-forward weight α j α functions ψi,j := ψj ◦ Φ−j i : ϕi (CT (Z)) × CT (Z) → (0, ∞) to
ψi,j (Y, V1 , ..., Vj ) := exp
(6.9)
β1 max(j, 1)∥Y∥cα1 + β2
j X
! ∥Vℓ ∥cα2
,
ℓ=1
j α for (Y, V1 , ..., Vj ) ∈ ϕi (CTα (Z)) × Cα T (Z) . In the following lemma, we show that (CT (Z), Ψ) is a k weighted Cloc -manifold.
Lemma 6.2. Let α ∈ (0, 1] and assume that the predual (E, ∥ · ∥E ) of (Z, ∥ · ∥Z ) has BAP. Then, k (CTα (Z), Ψ) is a weighted Cloc -manifold with global chart (6.6) over the model space (Cα T (Z), τ∞ ). α Proof. First, we show that ϕi (CTα (Z)) ⊆ Cα T (Z). To this end, we observe for every fixed X ∈ CT (Z) and n ∈ N that (6.10) ⌊1/α⌋ X X 1 ⌊1/α⌋ −1 (k1 ) (km ) (n) log (Xt ⊗ Xs ) ≤ (X−1 ⊗ · · · ⊗ (X−1 t ⊗ Xs ) t ⊗ Xs ) m Z ⊗n Z ⊗n m=1 k1 +...+km =n
≤ cn
max
k1 +...+km =n
m Y ℓ=1
(kℓ ) (X−1 t ⊗Xs )
Z
≤ cn ⊗k ℓ
max
k1 +...+km =n
m Y
k
(∥X∥α |s−t|α ) ℓ = cn ∥X∥nα |s−t|nα ,
ℓ=1
where cn > 0 is a constant. Hence, by using the Baker-Campbell-Hausdorff formula (see, e.g., [40, 1 1 [a, [a, b]] − 12 [b, [a, b]] + ... consisting of Lemma 7.24]) with BCH(a, b) := a + b + 21 [a, b] + 12 iterated Lie brackets of at least one a and b, that ∥(BCH(a, b) − a)(n) ∥Z ⊗n can be bounded via ∥[a, b]∥Z ⊗(l+m) = ∥a ⊗ b − b ⊗ a∥Z ⊗(l+m) ≤ 2∥a∥Z ⊗l ∥b∥Z ⊗m into products of ∥a(j) ∥Z ⊗j and
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
45
∥b(k) ∥Z ⊗k , with j, k ≥ 1, and the inequality (6.10), it holds that (6.11) maxn=1,...,⌊1/α⌋ (log⌊1/α⌋ (Xs ) − log⌊1/α⌋ (Xt ))(n)
∥ log⌊1/α⌋ (X)∥α =
sup
(n) ⌊1/α⌋ BCH(log⌊1/α⌋ (Xt ), log⌊1/α⌋ (X−1 (Xt ) t ⊗ Xs )) − log
maxn=1,...,⌊1/α⌋ ≤
sup
≤ C3 ≤ C4
Z ⊗n
|s − t|α
0≤s<t≤T
≤ C2
Z ⊗n
|s − t|α
0≤s<t≤T
p Y
max n=1,...,⌊1/α⌋ |j|+|k|=n
ℓ=1 p Y
max n=1,...,⌊1/α⌋ |j|+|k|=n
max n=1,...,⌊1/α⌋
log⌊1/α⌋ (Xt )(jℓ )
Z ⊗jℓ
! cjℓ ∥X∥jαℓ |s − t|jℓ α
ℓ=1
! q Y ℓ=1 q Y
(kℓ ) log⌊1/α⌋ (X−1 t ⊗ Xs )
sup
Z
⊗kℓ
|s − t|kℓ α
0≤s<t≤T
! ckℓ ∥X∥kαℓ
ℓ=1
∥X∥nα < ∞,
where C1 , ..., C4 > 0 are constants. This together with log⌊1/α⌋ (X0 ) = log⌊1/α⌋ (1) = 0 proves ⌊1/α⌋ : (CTα (Z), τ∞ ) → (Cα that log⌊1/α⌋ (X) ∈ Cα T (Z), τ∞ ) T (Z). In order to show that ϕi := log L ⌊1/α⌋ (1) (⌊1/α⌋) ⊗n is continuous, we fix some e := (e , ..., e ) ∈ . Then, by using the Bakern=1 E Campbell-Hausdorff formula and similar arguments as in (6.10)–(6.11) but now with the w∗ L⌊1/α⌋ seminorms of Z ⊗n , there exists a finite subset E ⊆ n=1 E ⊗n such that for every X, Z ∈ CTα (Z), we conclude that (6.12) pe log⌊1/α⌋ (X) − log⌊1/α⌋ (Z) = sup
max
t∈[0,T ] n=1,...,⌊1/α⌋
≤ sup
⟨(log⌊1/α⌋ (X) − log⌊1/α⌋ (Z))(n) , en ⟩Z ⊗n ×E ⊗n
max
t∈[0,T ] n=1,...,⌊1/α⌋
≤ C5 sup
max
t∈[0,T ] n=1,...,⌊1/α⌋ |j|+|k|=n
≤ C6 sup
max
t∈[0,T ] n=1,...,⌊1/α⌋ |j|+|k|=n
⌊1/α⌋ BCH(log⌊1/α⌋ (Zt ), log⌊1/α⌋ (Z−1 (Zt ) t ⊗ Xt )) − log p Y
! ⟨log
⌊1/α⌋
(jℓ )
(Zt )
(j,k) , ej,ℓ ⟩
ℓ=1 p Y ℓ=1
q Y
(n)
! ⟨log
⌊1/α⌋
(kℓ ) (j,k) (Z−1 ,e ej,ℓ ⟩ t ⊗ Xt )
ℓ=1
max
|r|=jℓ
M Y
! (r ) (j,k,r) ⟨Zt m , el,m ⟩
m=1
q Y ℓ=1
! max
|s|=kℓ
(sm ) (j,k,s) ⟨(Z−1 ,e el,m ⟩ t ⊗ Xt )
! ≤ C7
, en ⟩Z ⊗n ×E ⊗n
sup ∥Zt ∥G⌊1/α⌋ (Z)
max
max
k=1,...⌊1/α⌋ e∈E
t∈[0,T ]
sup
(k) ⟨(Z−1 , ek ⟩Z ⊗k ×E ⊗k t ⊗ Xt )
1 k
!k
t∈[0,T ]
! ≤ C7
sup ∥Zt ∥G⌊1/α⌋ (Z) t∈[0,T ]
max
max d∞,e (Z, X)k ,
k=1,...⌊1/α⌋ e∈E
where C5 , C6 , C7 > 0 are constants. This proves that ϕi := log⌊1/α⌋ : (CTα (Z), τ∞ ) → (Cα T (Z), τ∞ ) ⌊1/α⌋ α is continuous at Z ∈ CTα (Z). Conversely, in order to show that ϕ−1 := exp : (ϕ (C (Z)), τ∞ ) → i T i (CTα (Z), τ∞ ) is continuous, we use again the Baker-Campbell-Hausdorff formula and similar arL⌊1/α⌋ guments as in (6.12) to obtain that for every e := (e(1) , ..., e(⌊1/α⌋) ) ∈ n=1 E ⊗n there exists a
46
P. SCHMOCKER AND J. TEICHMANN
finite subset E ⊆
L⌊1/α⌋ n=1
E ⊗n such that for every Y, Z ∈ ϕi (CTα (Z)) it holds that
d∞,e (exp⌊1/α⌋ (Z), exp⌊1/α⌋ (Y)) = sup
max
t∈[0,T ] n=1,...,⌊1/α⌋
= sup
max
t∈[0,T ] n=1,...,⌊1/α⌋
≤ C8 sup
(exp⌊1/α⌋ (Zt )−1 ⊗ exp⌊1/α⌋ (Yt ))(n) , en Z ⊗n ×E ⊗n (exp⌊1/α⌋ (BCH(−Zt , Yt ))(n) , en Z ⊗n ×E ⊗n
max
t∈[0,T ] n=1,...,⌊1/α⌋
⟨BCH(−Zt , Yt )(n) , en ⟩Z ⊗n ×E ⊗n
1 n
1 n
1 n
! n1
! ≤ C9
sup ∥Zt ∥T ⌊1/α⌋ (Z)
max
max
k=1,...⌊1/α⌋ e∈E
t∈[0,T ]
sup
⟨(Yt − Zt )
(k)
, ek ⟩Z ⊗k ×E ⊗k
t∈[0,T ]
! ≤ C9
sup ∥Zt ∥T ⌊1/α⌋ (Z) t∈[0,T ]
max
1
max pe (Y − Z) k ,
k=1,...⌊1/α⌋ e∈E
where C8 , C9 > 0 are some constants. This proves that ϕ−1 := exp⌊1/α⌋ : (ϕi (CTα (Z)), τ∞ ) → i ⌊1/α⌋ α α : (CTα (Z), τ∞ ) → (Cα (CT (Z), τ∞ ) is continuous at Z ∈ ϕi (CT (Z)). Hence, ϕi := log T (Z), τ∞ ) ∞ α -manifold. is a homeomorphism onto its image, which shows that (CT (Z), τ∞ ) is a Cloc Finally, we show that the collection of push-forward weight functions Ψ = (ψi,j )j=0,...,k deα fined in (6.8) is admissible. Since the identity Γ : (Cα T (Z), ∥ · ∥α ) ,→ (CT (Z), τ∞ ) is a compact embedding (see [28, Theorem A.4]), we can follow the proof of Lemma 2.3 (i) to conclude that k -manifold. □ Ψi := (ψi,j )j=0,...,k is admissible, which shows that (CTα (Z), Ψ) is a weighted Cloc In order to ensure point separation for the application of the weighted Nachbin theorem (Theorem 3.12), we need to ensure that the signature (at terminal time) uniquely determines the path (up to so-called tree-like equivalences, see [14, 49]). To this end, we define the subspace n o b ∈ C α (Z) bTα (Z) := X b :X bt = (t, Xt ) for all t ∈ [0, T ] and some X ∈ CTα (Z) ⊆ CTα (Z), b C T b τ∞ ), where running time is now added. Then, equipped with the subspace topology of (CTα (Z), b := t 7→ exp⌊1/α⌋ (ta0 ) ⊗ ι(Xt ) ∈ C bTα (Z) (6.13) CTα (Z) ∋ X 7→ X ⌊1/α⌋ b ∞ is a Cloc -diffeomorphism, where a0 := (0, (1, 0), 0, ..., 0) ∈ T0 (Z) and where ι : G⌊1/α⌋ (Z) ,→ b is the canonical embedding, where Zb := R ⊕ Z and E b := R ⊕ E. Its inverse is given by G⌊1/α⌋ (Z) b 7→ X := t 7→ π exp⌊1/α⌋ (−ta0 ) ⊗ X b t ∈ C α (Z), bTα (Z) ∋ X (6.14) C T
b → G⌊1/α⌋ (Z) is the canonical projection. Hence, we define the collection where π : G⌊1/α⌋ (Z) b := (ψbj )j=0,...,k of weight functions ψbj : T j C b α (Z) → (0, ∞) by Ψ T ! j X j ⌊1/α⌋ c (ℓ) c b [c] ) := exp β1 max(j, 1)∥ log ψbj (X, (X)∥ 1 + β2 ∥(ϕbi ◦ c) (0)∥ 2 , α
b X
α
ℓ=1
b [c]j ) ∈ T j C b α (Z), with β1 , β2 > 0, c1 ≥ ⌊1/α⌋, and c2 > 0, where ϕbi : C b α (Z) → Cα (Z) for (X, T T T b X b i := (ψbi,j )j=0,...,k is defined in (6.15) below. Note that the corresponding push-forward weights Ψ b −j = ψj ◦ Φj = ψi,j , for all coincide with Ψi := (ψi,j )j=0,...,k defined in (6.9) since ψbi,j = ψbj ◦ Φ i i b α (Z), Ψ) is also a weighted C k -manifold. j = 0, ..., k. Hence, (C T
loc
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
47
Lemma 6.3. Let α ∈ (0, 1] and assume that the predual (E, ∥ · ∥E ) of (Z, ∥ · ∥Z ) has BAP. Then, b α (Z), Ψ) b is a weighted C k -manifold over the model space (Cα (Z), τ∞ ) with global chart (C T T loc b bTα (Z) ∋ X C
(6.15)
7→
b := log⌊1/α⌋ (X) ∈ Cα (Z) ϕbi (X) T
7→
⌊1/α⌋ bTα (Z). ϕb−1 (Y) ∈ C i (Y) := exp
whose inverse is given by V
(6.16)
b α (Z)) ∋ Y ϕbi (C T
Proof. By using that (6.14) is a homeomorphism (with inverse (6.13)) and ϕi = log⌊1/α⌋ : −1 (CTα (Z), τ∞ ) → (Cα = T (Z), τ∞ ) in (6.6) is a homeomorphism onto its image (with inverse ϕi b α (Z), τ∞ ) → exp⌊1/α⌋ : (ϕi (CTα (Z)), τ∞ ) → (CTα (Z), τ∞ ) in (6.7)), we conclude that ϕbi : (C T (Cα T (Z), τ∞ ) in (6.15) is also a homeomorphism onto its image (with inverse (6.16)), which shows b α (Z), Ψ) is a C k -manifold over (Cα (Z), τ∞ ). Moreover, by using that the push-forward that (C T T loc b i = (ψbi,j )j=0,...,k coincide with Ψi = (ψi,j )j=0,...,k , we can follow the proof of Lemma 6.2 weights Ψ b α (Z), Ψ) b is a weighted C k -manifold. (invoking Lemma 2.3 (i)) to obtain that (C □ T loc 6.3. Weighted universal approximation for differentiable functionals of rough paths. We now present the universal approximation theorem (UAT) for linear functions of the signature, k bα which can approximate any path space functional in BΨ b (CT (Z)) introduced in Section 2.2. In order to show their universality, we apply the weighted Nachbin theorem (Theorem 3.12) over k b α (Z), Ψ) b consisting of (time-extended) weakly the infinite-dimensional weighted Cloc -manifold (C T α-Hölder rough paths with values in a dual Banach space (Z, ∥ · ∥Z ) having predual (E, ∥ · ∥E ). The application of the weighted Nachbin theorem (Theorem 3.12) relies on the following properties of the signature. By using the Magnus expansion of the log-signature, we show that linear k b α (Z), τ∞ ) with appropriate growth conditions. functions of the signature are Cloc -maps on (C T Moreover, the time-extension ensures that the signature (at terminal time) uniquely determines the path (up to so-called tree-like equivalences, see [14, 49]), which ensures point separation. Third, the shuffle product can be used to express any product of linear functions of the signature as another linear function of the signature, which asserts the algebra property. k bα Theorem 6.4 (Universal approximation on BΨ b (CT (Z))). Let α ∈ (0, 1] and assume that the predual (E, ∥ · ∥E ) of (Z, ∥ · ∥Z ) has BAP. Then, the linear span of the set n o b 7→ ⟨S (n) (X) b T , ebn ⟩ b⊗n b⊗n ∈ R : ebn ∈ E b α (Z) ∋ X b ⊗n , n ∈ N0 C T Z ×E k bα k bα is a dense subset of BΨ b (CT (Z)), i.e., for every f ∈ BΨ b (CT (Z)) and ε > 0 there exists some N ∈ N PN (n) b ⊗n , such that and a linear function a 7→ ℓ(a) := ⟨a , ebn ⟩ b⊗n b⊗n , with ebn ∈ E n=0
Z
(j)
max
sup
j=0,...,k b bα (Z) (X,[c]jb )∈T j C T X
(f ◦ c)
×E
(0) − (ℓ ◦ S(·)T ◦ c)(j) (0) < ε. b [c]j ) ψbj (X, b X
Proof. We aim to apply the weighted Nachbin theorem (Theorem 3.12) to n o b 7→ ⟨S (n) (X) b T , ebn ⟩ b⊗n b⊗n ∈ R : ebn ∈ E bTα (Z) ∋ X b ⊗n , n ∈ N0 . (6.17) G := span C Z ×E k bα To this end, we need to show that G ⊆ BΨ b (CT (Z)) is a subalgebra satisfying the conditions (i)–(ii) of Theorem 3.12, where (6.18) n b 7→ S (n+k+1) (X) b T , (0, en ) (1, 0)⊗k ⊗ (1, 0) b⊗(n+k+1) b⊗(n+k+1) ∈ R : bTα (Z) ∋ X Ge := span C Z ,E o ⊗n en ∈ E , n ∈ {0, ..., ⌊1/α⌋}, k ∈ N0 ⊆G
48
P. SCHMOCKER AND J. TEICHMANN
is a possible candidate for a strongly point separating and nowhere vanishing vector subspace of b Ψ-moderate growth, with ⟨(t, z), (0, en )⟩Z× b E b := ⟨z, en ⟩Z×E and ⟨(t, z), (1, 0)⟩Z× b E b := t. k bα First, we show that the vector space G is contained in BΨ b (CT (Z)). By following [50, 104], we N b b s,t ))0≤s≤t≤T at observe that the (truncated) log-signature (L (X)s,t )0≤s≤t≤T := (logN (S N (X) α b b level N ∈ N of X ∈ C (Z) satisfies the (backward) controlled rough differential equation (CRDE) T
b s,t = H ad LN (X) b s,t (dX b s ), dLN (X)
s ∈ [0, t],
b t,t = 0, LN (X) b where T N (Z) b ∋ b 7→ (ad a)(b) := [a, b] ∈ T N (Z), b and where H(z) := in the Lie group gN (Z), 0 0 P∞ Bk k 1 1 z = z with Bernoulli numbers (B ) := (1, − , , ...). Hence, by following [55, k k∈N0 k=0 k! ez −1 2 6 N b α b b 79, 109], the log-signature (L (X)s,t )0≤s≤t≤T of X ∈ C (Z) admits the Magnus expansion T
(6.19)
b s,t = LN (X)
∞ X X n=1 σ∈Sn
Z
h
cσ
i b u , ..., [dX bu bu dX , d X ] ... , σ(1) σ(n−1) σ(n)
s<u1 <...<un <t
for some universal coefficients (cσ )σ ⊆ R ensuring that the series converges. Thus, by inserting b t = d exp⌊1/α⌋ (ta0 ) ⊗ ι(Xt ) dX = d exp⌊1/α⌋ (ta0 ) ⊗ exp⌊1/α⌋ (ι(Yt )) + exp⌊1/α⌋ (ta0 ) ⊗ d exp⌊1/α⌋ (ι(Yt )) = exp⌊1/α⌋ (−ta0 ) ⊗ − G(ad(ta0 ))(a0 dt) ⊗ exp⌊1/α⌋ (ι(Yt )) {z } | =−a0 dt
+ exp⌊1/α⌋ (ta0 ) ⊗ exp⌊1/α⌋ (−ι(Yt )) ⊗ − G(ad ι(Yt ))(ι(dYt ))
P∞ z zk into (6.19), where G(z) := e z−1 = k=0 (k+1)! (see, e.g., [40, Lemma 7.23]), it follows that N b−1 LN (ϕb−1 i (Y))T := L (ϕi (Y))0,T is a finite universal linear combination with polynomial vector b 7→ g(X) b := ⟨S (n) (X) b T , ebn ⟩ b⊗n b⊗n ∈ fields on the finite-step Lie algebra. Therefore, for every X Z ×E b ⊗n , we observe that G, with fixed n = N ∈ N and ebn ∈ E b T )(n) , ebn ⟩ b⊗n b⊗n , gi (Y) := g ◦ ϕb−1 (Y) = ⟨S (n) (ϕb−1 bn ⟩Zb⊗n ×Eb⊗n = ⟨expn (Ln (X) i i (Y))T , e Z ×E is a finite universal linear combination of iterated rough integrals (of degree n in Y). Hence, by b α (Z)) × Cα (Z)j ∋ (Y, V1 , ..., Vj ) 7→ induction on j = 1, ..., k, the directional derivatives ϕbi (C T T j b α (Z)) × Cα (Z)j , and d gi (Y; V1 , ..., Vj ) ∈ R exist, are continuous on compact subsets of ϕbi (C T T form again universal linear combinations of iterated rough integrals (of degree n − j in Y and of degree 1 in each tangent direction V1 , ..., Vj ). Thus, there exists some ee1 , ..., eeM ∈ E such that b α (Z)) × Cα (Z)j it holds that for every j = 1, ..., k and (Y, V1 , ..., Vj ) ∈ ϕbi (C T T max(n−j,0)
dj gi (Y; V1 , ..., Vj ) ≤ (1 + ∥Y∥α,ee1:M )
j Y
∥Vℓ ∥α,ee1:M ,
ℓ=1
where ∥Y∥α,ee1:M := maxm=1,...,M ∥Y∥α,eem . Now, we recall from Remark B.4 that (Cα T (Z), τ∞ ) := α ⌊1/α⌋ α ∗ (C0 ([0, T ]; g (Z)), τ∞ ) has AP with finite rank operators (Rγ )γ ⊆ (CT (Z), τ∞ ) ⊗ Cα T (Z) such that ∥Rγ (Y)∥α,ee1:M ≤ Cgi ∥Y∥α , for all Y ∈ Cα e1 , ..., eeM ∈ E T (Z) and some Cgi ≥ 1 (depending on e
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
49
b α (Z)) × Cα (Z)j that and therefore on gi ). This implies for every (Y, V1 , ..., Vj ) ∈ ϕbi (C T T dj gi (Rγ (Y); Rγ (V1 ), ..., Rγ (Vj )) ≤ (1 + ∥Y∥α,ee1:M )
max(n−j,0)
j Y
∥Vℓ ∥α,ee1:M
ℓ=1 max(n−j,0)
≤ (1 + Cgi ∥Y∥α )
(6.20)
j Y
(Cgi ∥Vℓ ∥α )
ℓ=1 max(n−j,0)
≤ Cgni (1 + ∥Y∥α )
j Y
∥Vℓ ∥α .
ℓ=1
Hence, it follows that dj gi (Rγ (Y); Rγ (V1 ), ..., Rγ (Vj )) R→∞ γ j=0,...,k b bα ψi,j (Y, V1 , ..., Vj ) j (ϕi (CT (Z))×Cα T (Z) )\Ki,j,R max(n−j,0) Qj (1 + ∥Y∥α ) n ℓ=1 ∥Vℓ ∥α ≤ Cgi lim max sup Pj c c2 R→∞ j=0,...,k b bα 1 j )\K + β exp β max(j, 1)∥Y∥ ∥V ∥ (ϕi (CT (Z))×Cα (Z) i,j,R α 2 1 ℓ α T ℓ=1 n Pj 1 + ∥Y∥α + ℓ=1 ∥Vℓ ∥α = 0, ≤ Cgni lim max sup Pj c1 c2 R→∞ j=0,...,k b bα j )\K exp β max(j, 1)∥Y∥ + β ∥V ∥ (ϕi (CT (Z))×Cα (Z) i,j,R α α 1 2 ℓ T ℓ=1 lim sup max
sup
b α (Z)) × Cα (Z)j ) \ Ki,j,R . Thus, Propowhere the supremum is taken over (Y, V1 , ..., Vj ) ∈ (ϕbi (C T T k bα k b bα sition 2.10 (ii) implies that gi ∈ BΨ b (CT (Z)). b i (ϕi (CT (Z))), which ensures that G ⊆ BΨ b α (Z)), which Next, we observe that G is by the shuffle property (6.5) a subalgebra of B k (C b Ψ
T
also contains the constants (by choosing n = 0 in (6.17)). In order to show that G is strongly b point separating and nowhere vanishing of Ψ-moderate growth, we claim that the vector subspace e G ⊆ G defined in (6.18) satisfies the conditions (M1)–(M3) and (M4’)–(M5’). For (M1), we fix b Z b ∈ C b α (Z) and first assume (by contradiction) that for every fixed k ∈ N0 , some distinct X, T n ∈ {0, ..., ⌊1/α⌋}, and en ∈ E ⊗n it holds that b T , (0, en ) (1, 0)⊗k ⊗ (1, 0) b⊗(n+k+1) b⊗(n+k+1) S (n+k+1) (X) Z ,E (6.21) (n+k+1) b ⊗k = S (Z)T , (0, en ) (1, 0) ⊗ (1, 0) Zb⊗(n+k+1) ,Eb⊗(n+k+1) .
c ∈C b α (Z) that Then, by using (6.4)–(6.5), we observe for every W T (n+k+1) c ⊗k S (W)T , (0, en ) (1, 0) ⊗ (1, 0) Zb⊗(n+k+1) ,Eb⊗(n+k+1) Z T c t , (0, en ) (1, 0)⊗k ⟩ b⊗(n+k) b⊗(n+k) dt = ⟨S (n+k) (W) Z ,E
0
(6.22)
Z T = 0
Z T = 0
c t , (0, en )⟩ b⊗n b⊗n ⟨S (k) (W) c t , (1, 0)⊗k ⟩ b⊗k b⊗k dt ⟨S (n) (W) Z ×E Z ×E ⟨S (n) (W)t , en ⟩Z ⊗n ×E ⊗n
tk dt. k!
Hence, by combining (6.21) with (6.22), it follows that Z T tk ⟨S (n) (X)t − S (n) (Z)t , en ⟩Z ⊗n ×E ⊗n dt = 0. k! 0
50
P. SCHMOCKER AND J. TEICHMANN
Thus, by using that Pol(R)|[0,T ] is weakly dense in L1 ([0, T ]) with L1 ([0, T ])∗ ∼ = L∞ ([0, T ]) ⊇ 0 (n) (n) C ([0, 1]), we have ⟨S (X)t − S (Z)t , en ⟩Z ⊗n ×E ⊗n = 0, for all t ∈ [0, T ]. Since E ∗ ∼ = Z is by the Hahn-Banach theorem point separating on Z, it follows that Xt = Zt , for all t ∈ [0, T ]. This b Z b∈C b α (Z) are distinct, which shows that Ge is point however contradicts the assumption that X, T α separating on CT (Z). For (M2), we observe that the map ge(·) := S (1) (·)T , (0, 0) (1, 0)⊗0 ⊗ b = T ̸= 0. For (M3), it suffices to show that Gei := {e e has g ◦ ϕb−1 : ge ∈ G} (1, 0) T satisfies ge(X) i α α k b b b b nowhere vanishing derivatives (as ϕi : CT (Z) → ϕi (CT (Z)) is a Cloc -diffeomorphism). To this b α (Z)) and V ∈ Cα (Z) \ {0}, whence there exists some e1 ∈ E and end, we fix some Y ∈ ϕbi (C T T t ∈ [0, T ] such that ⟨Vt , e1 ⟩Z×E ̸= 0. Moreover, by using (6.21) (with n = 1) and the definition of exp⌊1/α⌋ in (6.3), it holds for every k ∈ N0 that ⌊1/α⌋ (Y)) , (0, e ) S (k+2) (exp\ (1, 0)⊗k ⊗ (1, 0) Zb⊗(k+2) ,Eb⊗(k+2) T 1 Z T Z T (6.23) tk tk ⟨Yt , e1 ⟩Z×E dt. = ⟨S (1) (exp⌊1/α⌋ (Y))t , e1 ⟩Z×E dt = k! k! 0 0
Thus, by defining gei ∈ Gei as the left-hand side of (6.23), taking the directional derivative, and using again that polynomials are weakly dense in L1 ([0, T ]), there exists some k ∈ N0 such that Z T tk de gi (Y; V) = ⟨Vt , e1 ⟩Z×E dt ̸= 0, k! 0 e For (M4’), we use for every which shows that Gei has nowhere vanishing derivatives, and so does G. α b b e is γ that Tγ (ϕi (CT (Z))) is m-dimensional (with some m ∈ N), on which Gei := g ◦ ϕb−1 :g∈G i point separating and has nowhere vanishing derivatives, to obtain some ge1 , ..., gem such that ⊤ m b bα ηi,γ := ge1 ◦ ϕb−1 em ◦ ϕb−1 bi (C bα (Z))) : Tγ (ϕi (CT (Z))) → R i , ..., g i Tγ ( ϕ T
is an embedding, where the cutoff functions are obtained by a finite-dimensional smooth exhaustion argument. For (M5’), we fix gei ∈ Gei and define λ := β1 (Cgi Λ)−⌊1/α⌋ /2 > 0 with Cgi ≥ 1 from b α (Z)), τ∞ ) has AP with finite rank operators (Tγ )γ satisfying (B.3) for some above, where (ϕbi (C T constant Λ ≥ 1. Then, by using (6.20) and c1 ≥ ⌊1/α⌋, it follows for every γ that e |dπ gi (Y; e V e π )| exp λ|gi (Y)| lim max sup j=0,...,k e V e 1 , ..., V e j) R→∞ b (C ψi,γ,j (Y, e V e ,...,V e )∈T (ϕ bα (Z))×T (Cα (Z))j )\K (Y, π∈P j
1
j
γ
i
T
γ
T
i,γ,j,R
λ(1+Cgi Λ∥Y∥α )⌊1/α⌋
Qk (1+Cgi Λ∥Y∥α )max(⌊1/α⌋−|π|,0) ℓ=1 (Cgi Λ∥Vℓ ∥α ) ≤ lim max sup Pj R→∞ j=0,...,k K c exp β1 max(j, 1)∥Y∥cα1 + β2 ℓ=1 ∥Vℓ ∥cα2 i,j,R π∈Pj Qk ⌊1/α⌋ eβ1 /2(1+∥Y∥α ) (1 + ∥Y∥α )max(⌊1/α⌋−|π|,0) ℓ=1 ∥Vℓ ∥α ⌊1/α⌋ ≤ (Cgi Λ) lim max sup = 0, Pj R→∞ j=0,...,k K c exp β1 max(j, 1)∥Y∥cα1 + β2 ℓ=1 ∥Vℓ ∥cα2 i,j,R π∈Pj e
b α (Z)) × Cα (Z)j ) \ Ki,j,R . This shows where the supremum is taken over (Y, V1 , ..., Vj ) ∈ (ϕbi (C T T that (M5’) is satisfied. Finally, we show that condition (ii) of Theorem 3.12 is satisfied. By following the proof of ∗ α α Corollary B.3, we may assume that Rγ ∈ (Cα T (Z), τ∞ ) ⊗ CT (Z) is of the form CT (Z) ∋ Y 7→ PJ α b−1 : g ∈ G} and Rγ (Y) := j=1 Y(tj )hj (·) ∈ CT (Z). Hence, for every gi ∈ Gi := {g ◦ ϕi b α (Z)), we observe that (gi ◦ Rγ )(Y) depends only on products of linear functionals of Y ∈ ϕbi (C T
(Y(tj ))j=1,...,J , which can be approximated by elements from Gi with respect to ∥ · ∥Bk (ϕbi (Ui )) . b Ψ i
k bα Now, we can apply Theorem 3.12 to conclude that G is a dense subset of BΨ b (CT (Z)).
□
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
51
Remark 6.5. Let us point out the following remarks concerning Theorem 6.4: (i) Theorem 6.4 extends the universal approximation theorem of [28, Theorem 5.4] for linear functions of the signature by including the approximation of the directional derivatives. (ii) A similar result could be obtained for weakly geometric p-variation rough paths by intersecting Hölder spaces with p-variation spaces (see [28, Section 5.3]). (iii) Theorem 6.4 could be generalized to the space of stopped α-Hölder rough paths Λα T (Z) given as the vector bundle o [ n b [0,t] ∈ C btα (Z) . Λα (t, X[0,t] ) : X T (Z) := t∈(0,T )
Then, similar universal approximation results as in Corollary 5.5 (all directional derivatives) and Corollary 5.7 (only horizontal and vertical derivatives) can be shown, where the approximation holds uniformly in t ∈ (0, T ). 7. Numerical experiments In this section, we illustrate in two examples3 how to learn path space functionals including their horizontal and (an approximation of the) vertical derivatives. More precisely, given a functional f : Λα,r T,R → R, we use non-anticipative path-neural networks (Section 5) and linear functions of the signature (Section 6) to approximate the functional value f (t, x), the horizontal derivative Df (t, x) = lim+ h→0
f (t + h, xt ) − f (t, xt ) , h
and the vertical derivative f (t, xt + h1[t,T ] ) − f (t, xt ) f (t, xt + hgt ) − f (t, xt ) ≈ lim , h→0 h→0 h h
Df (t, x) := De1 f (t, x) = lim
where the latter is applied for linear functions of the signature (allowing only for continuous paths as input). Here, gt ∈ C α ([0, T ]) is an approximation of 1[t,T ] , e.g., given by if s ∈ [0, t], 0 gt (s) := h32 (s − t)2 − h23 (s − t)3 if s ∈ (t, min(t + h, T )], 1 if s ∈ (max(t + h, T ), T ], for some h ∈ (0, T ). As input data we generate M = 50000 sample paths of a one-dimensional Brownian motion x(m) := (x(m)t )t∈[0,T ] , for m = 1, ..., M , with T = 1, which are discretized over K = 101 equidistant time points (tk )k=1,...,K . Since the sample paths of Brownian motion are a.s. α-Hölder continuous, for all α ∈ (0, 1/2), we consider the weighted space of stopped α-Hölder continuous paths Λα,r T,R introduced in Section 5. On the other hand, every sample path of Brownian motion x(m), m = 1, . . . , M , can therefore be lifted to a weakly geometric α-rough path X(m) ∈ CTα (R), \ t , for t ∈ [0, T ]. for α ∈ (0, 1/2), from which we compute the time-extended signature S(X(m)) Since we only consider two directional derivatives, we define for non-anticipative path-neural networks (PNN) the weight function as in (5.5), i.e., Λα,r T,R ∋ (t, x)
7→
ψPNN := exp (β∥x∥cα ) ∈ (0, ∞),
3The experiments have been implemented in Python using the tensorflow package on a HPC (high-performing computing) cluster of ETH Zurich. The code can be found under the following link: https://github.com/psc25/ GlobalUATDerivatives.
52
P. SCHMOCKER AND J. TEICHMANN
for some β, c > 0. Similarly, for linear functions of the signature, we omit the directional derivative terms in (6.8) and define for β > 0 and c ≥ ⌊1/α⌋ the weight function ⌊1/α⌋ 2 c ∈ (0, ∞). Λα,r ∋ (t, x) → 7 ψ := exp β∥ log (S (X))∥ Sig α T,R In the first example, we consider the non-anticipative functional f1 : Λα,r T,R → R, which is together with its horizontal and vertical derivative for every (t, x) ∈ Λα,r given by T,R Z t f1 (t, x) = xt xs ds, 0
Df1 (t, x) = x2t , Z t Df1 (t, x) = xs ds.
(7.1)
0
In the second example, we consider the non-anticipative functional f2 : Λα,r T,R → R, which is α,r together with its horizontal and vertical derivative for every (t, x) ∈ ΛT,R given by Z t x2 f2 (t, x) = max(xs , 0)ds − t , 2 0 (7.2) Df2 (t, x) = max(xt , 0), Df2 (t, x) = −xt . We split up the data into 80%/20% for training and testing, respectively, and then apply the Adam algorithm (see [60]) over 4000 epochs with learning rate 10−5 and batchsize 500 to minimize the weighted mean squared error " 2 2 M K 1 XX |fi (tk , x(m)) − g(tk , x(m))| |Dfi (tk , x(m)) − Dg(tk , x(m))| + M K m=1 ψMethod (tk , x(m)) ψMethod (tk , x(m)) k=1 (7.3) 2 # |Dfi (tk , x(m)) − Dg(tk , x(m))| + ψMethod (tk , x(m)) ρ for i ∈ {1, 2} and Method ∈ {PNN, Sig}, where g := φ ∈ F N ρ,e for path-NN (PNN) or Λα,r T ,R P \ t , eI ⟩ for linear functions of the signature (Sig). In both cases, we g := 0≤|I|≤NSig aI ⟨S(X(m)) k b (m) )∥α used in ψPNN and ψSig , respectively compute an approximation of ∥x(m)∥α and ∥ log⌊1/α⌋ (X
(see the code). Moreover, we choose α = 0.4, β = 0.01, γ = 2.0, and NSig = 6. For the PNNs, ρ we consider φ ∈ FN ρ,e (see Definition 5.2) with NP N N = 30 neurons, activation functions Λα,r T ,R
ρ(s) = ρe(s) = tanh(s), and classical neural networks (ϕn,1 )n=1,...,NP N N with one hidden layer of N1 = 20 neurons, where the time integral inside is approximated with a left Riemann sum. Figure 2–3 empirically demonstrate that the values of the functionals f1 and f2 together with their horizontal and vertical derivatives can be approximated both by non-anticipative path-neural networks (PNN) as well as linear functions of the signature (Sig). The approximations of the PNNs (dotted lines) and linear functions of the signature (dash-dotted lines) are very accurate as they almost overlap the true values (continuous lines). Notice that the weighted mean squared error (7.3) reflects the weighted aspect of our universal approximation theorems (UATs), analogously to classical UATs over compact subsets, for which the unweighted mean squared error is applied. However, unlike classical UATs on compacta, our framework ensures the existence of an approximation beyond compact subsets, including the derivatives. This overcomes the limitation that, for a pre-specified compact training set (e.g., sample paths of Brownian motion), the test data may lie outside the chosen compactum.
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
(a) Learning performance
53
(b) t 7→ f1 (t, x) for three samples x of test set
(c) t 7→ Df1 (t, x) for three samples x of test set (d) t 7→ Df1 (t, x) for three samples x of test set ρ Figure 2. Learning f1 defined in (7.1) by path-NN φ ∈ FN ρ,e α,r (label FNN) and Λ T ,R P b linear function of the signature 0≤|I|≤NSig aI ⟨eI , S(X)t , eI ⟩ (label Sig). In (a), the weighted mean squared error (7.3) is evaluated on the training set in each epoch (continuous line) as well as on the test after every 200-th epoch (dots). In (b)–(d), three samples x(m) of the test set are shown together with f1 (·, x(m)), Df1 (·, x(m)), Df1 (·, x(m)) and their approximations.
Appendix A. α-Hölder Skorokhod space Dα,r ([0, T ]; Z) In this section, we fix some (α, r) ∈ [0, 1) × [1, ∞), T > 0, and a dual Banach space (Z, ∥ · ∥Z ) with predual (E, ∥ · ∥E ). Then, we first show that the α-Hölder Skorokhod space Dα,r ([0, T ]; Z) ⊆ D0 ([0, T ]; Z) introduced in Section 1.4 is a Banach space, which is isometrically isomorphic to the direct sum of the α-Hölder space (C α ([0, T ]; Z), ∥ · ∥α ) and the Banach space (ℓr ((0, T ]; Z), ∥ · ∥ℓr ) P r 1/r consisting of sequences (zt )t∈(0,T ] ⊆ Z with ∥z∥ℓr := < ∞. t∈(0,T ] ∥zt ∥Z Theorem A.1. Let (α, r) ∈ [0, 1) × [1, ∞). Then, (Dα,r ([0, T ]; Z), ∥ · ∥α,ℓr ) is a Banach space, which is isometrically isomorphic to C α ([0, T ]; Z) ⊕∞ ℓr ((0, T ]; Z). Proof. First, we observe that the embedding (Dα,r ([0, T ]; Z), ∥ · ∥α,ℓr ) ∋ x
7→
Ξ(x) := (xc , ∆x) ∈ (C α ([0, T ]; Z) ⊕∞ ℓr ((0, T ]; Z), ∥ · ∥⊕∞ )
54
P. SCHMOCKER AND J. TEICHMANN
(a) Learning performance
(b) t 7→ f2 (t, x) for three samples x of test set
(c) t 7→ Df2 (t, x) for three samples x of test set
(d) t 7→ Df2 (t, x) for three samples x of test set
ρ Figure 3. Learning f2 defined in (7.2) by path-NN φ ∈ FN ρ,e α,r (label FNN) and Λ
T ,R P b linear function of the signature 0≤|I|≤NSig aI ⟨eI , S(X)t , eI ⟩ (label Sig). In (a), the weighted mean squared error (7.3) is evaluated on the training set in each epoch (continuous line) as well as on the test after every 200-th epoch (dots). In (b)–(d), three samples x(m) of the test set are shown together with f2 (·, x(m)), Df2 (·, x(m)), Df2 (·, x(m)), and their approximations.
is continuous, where C α ([0, T ]; Z)⊕∞ ℓr ((0, T ]; Z) is a Banach space under the norm ∥(y, z)∥⊕∞ := max(∥y∥α , ∥z∥ℓr ), see, e.g., [120, Section II.B.20]. Moreover, the linear mapping X (C α ([0, T ]; Z)⊕∞ ℓr ((0, T ]; Z), ∥·∥⊕∞ ) ∋ (y, z) 7→ t 7→ y(t)+ zs ∈ (Dα,r ([0, T ]; Z), ∥·∥α,ℓr ) s∈(0,t]
is well-defined, continuous, and an inverse of Ξ, which shows that Ξ is an isometric isomorphism. Hence, (Dα,r ([0, T ]; Z), ∥·∥α,ℓr ) is a Banach space as isometrically isomorphic image of the Banach space C α ([0, T ]; Z) ⊕∞ ℓr ((0, T ]; Z). □ In addition, we use the preduals of the Banach spaces (C α ([0, T ]; Z), ∥· ∥α ) and (ℓr ((0, T ]; Z), ∥ · ∥ ) to show that (Dα,r ([0, T ]; Z), ∥ · ∥α,ℓr ) is a dual Banach space. ℓr
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
55
Theorem A.2. Let (α, r) ∈ (0, 1)×(1, ∞). Then, (Dα,r ([0, T ]; Z), ∥·∥α,ℓr ) is a dual Banach space. Moreover, its weak-∗-topology τw∗ coincides on every ∥ · ∥α,ℓr -bounded subset of C α ([0, T ]; Z) ⊆ Dα,r ([0, T ]; Z) with the w∗ -uniform topology τ∞ . Proof. By using the linear isomorphism C α ([0, T ]; Z) ∋ x 7→ (x(0), x − x(0)) ∈ Z ⊕∞ C0α ([0, T ]; Z) and that C0α ([0, T ]; Z) ∼ = L(Æ([0, T ], dα ); Z) (see [117, Theorem 3.6], where Æ([0, T ], dα ) denotes the Arens-Eells space defined in [117, Definition 3.2] over the snow-flaked metric space ([0, T ], dα ), with dα (s, t) = |s − t|α ), we observe that C α ([0, T ]; Z) ∼ = Z ⊕∞ C0α ([0, T ]; Z) ∼ = Z ⊕∞ L(Æ([0, T ], dα ); Z). Hence, by combining this with Theorem A.1, it follows that Dα,r ([0, T ]; Z) ∼ = C α ([0, T ]; Z) ⊕∞ ℓr ((0, T ]; Z) ∼ = Z ⊕∞ L(Æ([0, T ], dα ); Z) ⊕∞ ℓr ((0, T ]; Z). b π E)∗ (see [99, Theorem 2.9], where ⊗ bπ Thus, by using that L(Æ([0, T ], dα ); Z) ∼ = (Æ([0, T ], dα )⊗ r′ ∗ ∼ r denotes the completed projective tensor product) and that ℓ ((0, T ]; E) = ℓ ((0, T ]; Z) with 1/r + 1/r′ = 1 (see [54, Proposition 1.3.3]), we can apply [120, Section II.B.21] to conclude that Dα,r ([0, T ]; Z) ∼ = Z ⊕∞ L(Æ([0, T ], dα ); Z) ⊕∞ ℓr ((0, T ]; Z) ∗ ′ ∼ b π E ⊕1 ℓr ((0, T ]; E) = E ⊕1 Æ([0, T ], dα )⊗
(A.1)
′ b π E ⊕1 ℓr ((0, T ]; E) is is a dual Banach space, where the predual G := E ⊕1 Æ([0, T ], dα )⊗ equipped with the norm ∥(e, T, y)∥⊕1 := ∥e∥E + ∥T ∥Æ([0,T ],dα )⊗ b π E + ∥y∥r ′ . Finally, we show for every fixed ∥ · ∥α,ℓr -bounded subset B ⊆ C α ([0, T ]; Z) ⊆ Dα,r ([0, T ]; Z) that τw∗ |B := {U ∩ B : U ∈ τw∗ } = {U ∩ B : U ∈ τ∞ } := τ∞|B . For τw∗ |B ⊆ τ∞|B , we fix some x ∈ B and a set y ∈ B : maxm=1,...,M ⟨x, gm ⟩Dα,r ([0,T ];Z)×G < δ of the x-neighborhood ′ basis of (B, τw∗ |B ), where δ > 0 and g1 , ..., gM ∈ G := F ⊕1 ℓr ((0, T ]; E) with F := E ⊕1 b π E . Then, by using the canonical projection π : G → F and that the weak-∗Æ([0, T ], dα )⊗ α topology of C ([0, T ]; Z) coincides with τ∞ on the ∥·∥α -bounded subset B (see [28, Theorem A.5]), there exist e1 , ..., eN ∈ E and ε > 0 such that o o n n y ∈ B : max ⟨x − y, gm ⟩Dα,r ([0,T ];Z)×G < δ = y ∈ B : max ⟨x − y, π(gm )⟩C α ([0,T ];Z)×F < δ m m ⊆ y ∈ B : max sup |⟨(x − y)(t), en ⟩Z×E | < ε . n
t∈[0,T ]
Since the set on the right-hand side belongs to the x-neighborhood basis of (B, τ∞|B ), we obtain that τw∗ |B ⊆ τ∞|B . Conversely, for τw∗ |B ⊇ τ∞|B , we fix again some x ∈ B and a set y ∈ B : maxm=1,...,M supt∈[0,T ] |⟨(x − y)(t), em ⟩Z×E | < δ of the x-neighborhood basis of (B, τ∞|B ), where e1 , ..., eM ∈ E and δ > 0. Then, by using the canonical embedding ι : F → G and again that τ∞ coincides on the ∥ · ∥α -bounded subset B with the weak-∗-topology of C α ([0, T ]; Z) (see [28, Theorem A.5]), there exist some f1 , ..., fN ∈ F and ε > 0 such that ( ) n o y ∈ B : max sup |⟨(x−y)(t), em ⟩Z×E | < δ ⊆ y ∈ B : max ⟨x−y, fn ⟩C α ([0,T ];Z)×F < ε m
t∈[0,T ]
n
n o = y ∈ B : max ⟨x−y, ι(fn )⟩Dα,r ([0,T ];Z)×G < ε . n
Since the set on the right-hand side belongs to the x-neighborhood basis of (B, τw∗ ), we obtain that τw∗ |B ⊇ τ∞|B , which shows that τw∗ |B = τ∞|B . □
56
P. SCHMOCKER AND J. TEICHMANN
Appendix B. BAP of (C α (S; Z), τ∞ ) and (Dα,r ([0, T ]; Z), τw∗ ) In this section, we first show when (C α (S; Z), τ∞ ) has the bounded approximation property (BAP), where α ∈ [0, 1), (S, dS ) is a compact metric space, and (Z, ∥ · ∥Z ) is a dual Banach space equipped with the weak-∗-topology τw∗ . To this end, we start with the case α = 0. Theorem B.1. Let (S, dS ) be a compact metric space and let (Z, ∥ · ∥Z ) be a dual Banach space with predual (E, ∥ · ∥E ) having BAP. Then, (C 0 (S; Z), τ∞ ) has ∥ · ∥∞ -BAP. Proof. Since (E, ∥ · ∥E ) has BAP, there exists some λ ≥ 1 and (Qγ )γ ⊆ E ∗ ⊗ E with ∥Qγ ∥L(E;E) ≤ e ⊆ E it holds that λ, for all γ, such that for every relatively compact subset K lim sup ∥e − Qγ (e)∥E = 0.
(B.1)
γ
e e∈K
Now, we fix some ε > 0, a relatively compact subset K of (C 0 (S; Z), τ∞ ), and some e1 , ..., eN ∈ E defining the seminorms z 7→ pZ (z) := maxn=1,...,N |⟨en , z⟩Z×E | ∈ P(Z,τw∗ ) and x 7→ pC 0 (S;Z) (x) := sups∈S pZ (x(s)) ∈ P(C 0 (S;Z),τ∞ ) . Then, by applying the vector-valued ArzelàAscoli theorem in [119, Theorem 43.15], the relatively compact set K is equicontinuous (with respect to pZ ). Hence, there exists δ ∈ (0, 1) such that for every s, t ∈ S and x ∈ K it holds that ε (B.2) dS (s, t) < δ =⇒ pZ (x(t) − x(s)) < . 2 For 0 < r < 2δ , we now use that (S, dS ) is compact and thus totally bounded to obtain a finite maximal r-separated set of points (tj )j=1,...,J ⊂ S, i.e., dS (ti , tj ) ≥ r, for all i ̸= j, such SJ that j=1 Br (tj ) = S. In addition, there exists a partition of unity (gj )j=1,...,J subordinate to PJ (B2r (tj ))j=1,...,J , i.e., supp(gj ) ⊆ B2r (tj ), 0 ≤ gj ≤ 1, and j=1 gj = 1. Furthermore, since B := {x(s) : x ∈ K, s ∈ S} is bounded in (Z, τw∗ ), which implies that C := 1 + supz∈B ∥z∥ < ∞ by the uniform boundedness principle, we use (B.1) to obtain some γ satisfying ε . max ∥en − Qγ en ∥E < n=1,...,N 2C PJ 0 ∗ 0 ∗ Next, we define x 7→ Tε,e1:N ,K,γ (x) := j=1 gj (·)Qγ (x(tj )) ∈ (C (S; Z), τ∞ ) ⊗ C (S; Z). Then, by using that s ∈ supp(gj ) implies dS (s, tj ) < r < 2δ ≤ δ and therefore pZ (x(s) − x(tj )) < 2ε by (B.2), and that pZ x(tj ) − Q∗γ (x(tj )) = maxn=1,...,N |⟨en , x(tj ) − Q∗γ (x(tj ))⟩| = ε maxn=1,...,N |⟨en − Qγ (en ), x(tj )⟩| ≤ maxn=1,...,N ∥en − Qγ (en )∥E ∥x(tj )∥Z < 2C C = 2ε , we have sup pC 0 (S;Z) (x − Tε,e1:N ,K,γ (x)) = sup sup pZ (x(s) − Tε,e1:N ,K,γ (x)(s)) x∈K s∈S J X = sup sup pZ gj (s) x(s) − Q∗γ (x(tj ))
x∈K
x∈K s∈S
≤ sup sup
j=1 J X
x∈K s∈S j=1
≤ sup sup
J X
x∈K s∈S j=1
gj (s)pZ x(s) − Q∗γ (x(tj ))
gj (s)pZ (x(s) − x(tj )) + sup sup
J X
x∈K s∈S j=1
gj (s)pZ x(tj ) − Q∗γ (x(tj ))
ε ε < + = ε. 2 2 Hence, the net (Tε,K,e1:N ,γ )ε,K,e1:N ,γ ⊆ (C 0 (S; Z), τ∞ )∗ ⊗ C 0 (S; Z) converges to the identity idC 0 (S;Z) : C 0 (S; Z) → C 0 (S; Z), uniformly on each relatively compact subset of (C 0 (S; Z), τ∞ ),
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
57
showing that (C 0 (S; Z), τ∞ ) has AP. Moreover, for every x ∈ C 0 (S; Z) and ee1 , ..., eeM ∈ E defining the seminorms en , z⟩Z×E | ∈ P(Z,τw∗ ) and x 7→ peC 0 (S;Z) (x) := z 7→ peZ (z) := maxn=1,...,M |⟨e sups∈S peZ (x(s)) ∈ P(C 0 (S;Z),τ∞ ) , we observe that J X peC 0 (S;Z) (Tε,K,e1:N ,γ (x)) = sup peZ gj (s)Q∗γ (x(tj )) s∈S
≤ sup
j=1 J X
gj (s)
s∈S j=1
|⟨x(tj ), Qγ (e en )⟩Z×E |
∥Qγ een ∥E ∥x∥∞ n=1,...,M ≤ λ max ∥e en ∥E ∥x∥∞ ,
≤
max
n=1,...,M
max
n=1,...,M
which proves that (C 0 (S; Z), τ∞ ) has ∥ · ∥∞ -BAP.
□
For the BAP of (C α (S; Z), τ∞ ) with α ∈ (0, 1), we impose the following condition on (S, dS ) to obtain a specific partition of unity (gj )j=1,...,J . Here, a metric space (S, dS ) is called doubling if there exists a doubling constant M > 0 such that for every s ∈ S and r > 0 the open ball Br (s) can be covered with M open balls of radius r/2. Moreover, Lip(S) denotes the vector space of < ∞. Lipschitz continuous functions g : S → R with |g|1 := sups,t∈S, s̸=t |g(s)−g(t)| dS (s,t) Lemma B.2. Let α ∈ (0, 1), let (S, dS ) be a compact doubling metric space. Then, for every ε > 0, relatively compact subset K of (C 0 (S; Z), τ∞ ), and e1 , ..., eN ∈ E, the partition of unity (gj )j=1,...,J in the proof of Theorem B.1 can be chosen to satisfy g1 , ..., gJ ∈ Lip(S) with supp(gj ) ⊆ B2r (tj ), |gj |1 ≤ 20Cr−1 , and #{j : s ∈ supp(gj )} ≤ C, where r := 4δ and C > 0 is a universal constant independent of ε, K, e1 , ..., eN , δ, and r. Proof. Let (tj )j=1,...,J ⊂ S be the finite maximal r-separated set of points from the proof of Theorem B.1, i.e., dS (ti , tj ) ≥ r, for all i ̸= j. Then, for every j = 1, ..., J, we define the function 3 − r−1 dS (s, tj ), 0 ∈ [0, 32 ], S ∋ s 7→ hj (s) := max 2 SJ which satisfies supp(hj ) ⊆ B2r (tj ) and |hj |1 ≤ r−1 , for all j = 1, ..., J. Since j=1 Br (tj ) = S, there exists for every s ∈ S some j = 1, ..., J with dS (s, tj ) < r, which implies that hj (s) ≥ 12 . PJ Hence, by using the function S ∋ s 7→ H(s) := j=1 hj (s) ∈ [0, ∞), which satisfies H(s) ≥ 12 , for all s ∈ S, we can define for every j = 1, ..., J the function S∋s
7→
gj (s) :=
hj (s) ∈ [0, ∞), H(s)
PJ which satisfies supp(gj ) ⊆ supp(hj ) ⊆ B2r (tj ), 0 ≤ gj ≤ 1, and j=1 gj = 1. Thus, by using that (tj )j=1,...,J ⊂ S are r-separated, there exists a constant C > 0 (depending only on the doubling constant M > 0) such that every ball of radius 5r/2 contains at most C of the points tj , which implies that #{j : s ∈ supp(hj )} = #{j : dS (s, tj ) < 2r} ≤ C and therefore 12 ≤ H(s) ≤ 3C 2 . P Thus, if dS (s, t) < r, we use that |H(s) − H(t)| ≤ j=1, hj (s)̸=hj (t) |hj (s) − hj (t)| ≤ Cr−1 dS (s, t) with hj (s) ̸= hj (t) implying hj (s) ̸= 0 or hj (t) ̸= 0 and therefore tj ∈ B5r/2 (s), and if dS (s, t) ≥ r, we insert that |H(s) − H(t)| ≤ |H(s)| + |H(t)| ≤ 2 32 C ≤ 3Cr−1 dS (s, t) to conclude for every
58
P. SCHMOCKER AND J. TEICHMANN
s, t ∈ S that hj (s) hj (t) |hj (s) − hj (t)| |H(s) − H(t)| ≤ − + |hj (t)| H(s) H(t) H(s) H(s)H(t) 3 ≤ 2|hj (s) − hj (t)| + 4|H(s) − H(t)| 2 ≤ 2r−1 dS (s, t) + 18Cr−1 dS (s, t)
|gj (s) − gj (t)| =
≤ 20Cr−1 dS (s, t), which proves that |gj |1 ≤ 20Cr−1 .
□
Corollary B.3. Let α ∈ (0, 1), let (S, dS ) be a compact doubling metric space with designated origin 0 ∈ S, and let (Z, ∥ · ∥Z ) be a dual Banach space with predual (E, ∥ · ∥E ) having BAP. Then, (C α (S; Z), τ∞ ) has AP with finite rank operators (Tϑ )ϑ ⊆ (C α (S; Z), τ∞ )∗ ⊗ C α (S; Z) and for every ee1 , ..., eeM ∈ E there exists some Λ > 0 such that for every ϑ and x ∈ C α (S; Z) it holds that ∥Tϑ (x)∥α,ee1:M ≤ Λ∥x∥α ,
(B.3)
where ∥x∥α,ee1:M := maxm=1,...,M ∥x∥α,eem . Proof. By using the embedding C α (S; Z) ,→ C 0 (S; Z), we adapt the proof of Theorem B.1 with PJ finite rank operators x 7→ Tε,e1:N ,K,γ (x) := j=1 gj (·)Q∗γ (x(tj )) ∈ (C α (S; Z), τ∞ )∗ ⊗ C α (S; Z), where (Qγ )γ ⊆ E ∗ ⊗ E with ∥Qγ ∥L(E;E) ≤ λ. Note that, by Lemma B.2, we may choose the partition of unity (gj )j=1,...,J to satisfy supp(gj ) ⊆ B2r (tj ), |gj |1 ≤ 20Cr−1 , and #{j : s ∈ for every ε > 0, supp(gj )} ≤ C, where r = 4δ . Hence, by the proof of Theorem B.1, it follows e1 , ..., eN ∈ E defining the seminorm x 7→ pC α (S;Z) (x) := sups∈S pZ (x(s)) ∈ P(C α (S;Z),τ∞ ) , and relatively compact subset K of (C α (S; Z), τ∞ ) that there exists some γ such that sup pC α (S;Z) (x − Tε,e1:N ,K,γ (x)) < ε, x∈K
which shows that (C α (S; Z), τ∞ ) has AP. Moreover, for every fixed x ∈ C α (S; Z) and ee1 , ..., eeM ∈ E, we use that 0 ∈ supp(gj ) ⊆ B2r (tj ) implies dS (tj , 0) < 2r ≤ 1 to obtain that |⟨Tε,e1:N ,K,γ (x)(0), eem ⟩Z×E | ≤
J X
gj (0)⟨Q∗γ (x(tj )), eem ⟩Z×E
j=1
(B.4)
≤ |⟨x(0), Qγ (e em )⟩Z×E | +
J X
gj (0)⟨x(tj ) − x(0), Qγ (e em )⟩Z×E
j=1
≤ ∥Qγ ∥L(E;E) ∥e em ∥E ∥x(0)∥Z + ∥Qγ ∥L(E;E) ∥e em ∥E |x|α
J X
gj (0)dS (tj , 0)α
j=1 gj (0)̸=0
≤ λ∥e em ∥E (∥x(0)∥Z + |x|α ) . In addition, for s, t ∈ S with dS (s, t) < r, we add and subtract ⟨x(s), Qγ (e em )⟩Z×E , use that gj (s) ̸= gj (t) implies s ∈ supp(gj ) ⊆ B2r (tj ) or t ∈ supp(gj ) ⊆ B2r (tj ), ensuring that dS (tj , s) ≤ 2r (if s ∈ supp(gj )) or dS (tj , s) ≤ dS (tj , t) + dS (t, s) < 2r + r = 3r (if t ∈ supp(gj )), and PJ that j=1, gj (s)̸=gj (t) dS (tj , s)α |gj (s) − gj (t)| ≤ C · (3r)α |gj |1 dS (s, t) ≤ 40 · 3α rα−1 C 2 dS (s, t)α ≤
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
59
120C 2 dS (s, t)α to deduce that |⟨Tε,e1:N ,K,γ (x)(s) − Tε,e1:N ,K,γ (x)(t), eem ⟩Z×E | ≤
J X (gj (s) − gj (t))⟨Q∗γ (x(tj )), eem ⟩Z×E j=1
J X
≤ (B.5)
|gj (s) − gj (t)||⟨x(tj ) − x(s), Qγ (e em )⟩Z×E |
j=1 gj (s)̸=gj (t)
≤ ∥Qγ ∥L(E;E) ∥e em ∥E |x|α
J X
dS (tj , s)α |gj (s) − gj (t)|
j=1 gj (s)̸=gj (t)
≤ 120C 2 λ∥e em ∥E |x|α dS (s, t)α . Furthermore, for s, t ∈ S with dS (s, t) ≥ r, we use that s ∈ supp(gj ) ⊆ B2r (tj ) implies dS (tj , s)α ≤ (2r)α ≤ 2dS (s, t)α to conclude that (B.6) J X (gj (s) − gj (t))⟨Q∗γ (x(tj )), eem ⟩Z×E |⟨Tε,e1:N ,K,γ (x)(s) − Tε,e1:N ,K,γ (x)(t), eem ⟩Z×E | ≤ j=1
≤ |⟨x(s) − x(t), Qγ (e em )⟩Z×E | +
J X
gj (s)⟨x(tj ) − x(s), Qγ (e em )⟩Z×E
j=1
+
J X
gj (t)⟨x(tj ) − x(t), Qγ (e em )⟩Z×E
j=1
≤ ∥Qγ ∥L(E;E) ∥e em ∥E ∥x(s) − x(t)∥Z + |x|α
J X
gj (s)dS (tj , s)α +
j=1
J X
gj (t)dS (tj , t)α
j=1
≤ λ∥e em ∥E (|x|α dS (s, t)α + 2|x|α dS (s, t)α + 2|x|α dS (s, t)α ) ≤ 5λ∥e em ∥E |x|α dS (s, t)α . Thus, by using (B.4), (B.5), and (B.6), it follows that ∥Tε,e1:N ,K,γ (x)∥α,ee1:M
|⟨T (x)(s) − T (x)(t), e e ⟩ | ε,e1:N ,K,γ ε,e1:N ,K,γ m Z×E = max |⟨Tε,e1:N ,K,γ (x)(0), eem ⟩Z×E | + sup m=1,...,M dS (s, t)α s,t∈S s̸=t ≤ 121C 2 λ max ∥e em ∥E ∥x∥α , m=1,...,M
which proves (B.3).
□
Remark B.4. Let α ∈ (0, 1), let (S, dS ) be a compact doubling metric space with designated origin 0 ∈ S, and let (Z, ∥ · ∥Z ) be a dual Banach space with predual (E, ∥ · ∥E ) having BAP. Then, for the vector space C0α (S; Z) of α-Hölder continuous functions x : S → Z preserving the origin, i.e., x(0) = 0, it is possible to choose the partition of unity (gj )j=1 to satisfy g(0) = 0, which ensures that (C0α (S; Z), τ∞ ) has AP with (B.3).
60
P. SCHMOCKER AND J. TEICHMANN
Moreover, we follow the proof of Lemma 2.9 to give conditions when (Dα,r ([0, T ]; Z), τw∗ ) has the bounded approximation property (BAP), where (α, r) ∈ (0, 1) × (1, ∞), T > 0, and (Z, ∥ · ∥Z ) is a dual Banach space equipped with the weak-∗-topology τw∗ . Theorem B.5. Let (α, r) ∈ (0, 1) × (1, ∞) and let (Z, ∥ · ∥Z ) be a dual Banach space with predual (E, ∥ · ∥E ) having BAP. Then, (Dα,r ([0, T ]; Z), τw∗ ) has ∥ · ∥α,ℓr -BAP. Proof. First, we recall from (A.1) that a predual of Dα,r ([0, T ]; Z) is given by ′ b π E ⊕1 ℓr ((0, T ]; E). G := E ⊕1 Æ([0, T ], dα )⊗ Now, we observe that (E, ∥ · ∥E ) has BAP by assumption. Moreover, (Æ([0, T ], dα ), ∥ · ∥Æ([0,T ],dα ) ) has BAP by [69, Corollary 2.2], which is preserved under the completed projective tensor product ′ b π (see [99, Section 4.1]). In addition, (ℓr ((0, T ]; E), ∥ · ∥r′ ) has BAP with finite rank operators ⊗ ′ ′ y 7→ TF,γ (y) := (1F (s)Qγ (ys ))s∈(0,T ] ∈ ℓr ((0, T ]; E)∗ ⊗ ℓr ((0, T ]; E), where F ⊆ (0, T ] is finite and (Qγ )γ is the BAP-net of (E, ∥ · ∥E ). Hence, by using that finite ℓ1 -sums preserve norm-BAP, the predual (G, ∥ · ∥⊕1 ) has BAP. Finally, we can use adjoints as in the proof of Lemma 2.9 to conclude that (Dα,r ([0, T ]; Z), τw∗ ) has ∥ · ∥α,ℓr -BAP. □
Appendix C. Proof of results in Section 2 C.1. Proof of Proposition 2.10. k Proof of Proposition 2.10. For (i), let f ∈ BΨ (U ; Y ) and fix some ε > 0. Then, by definition of k k BΨ (U ; Y ), there exists some g ∈ Cb (U ; Y ) with ∥dj g(u; v1 , ..., vj )∥Y ≤ Cg pX (v1 ) · · · pX (vj ) for all j = 0, ..., k, u ∈ U , v1 , ..., vj ∈ X, and some Cg ≥ 0 and pX ∈ P(X,τX ) , such that
(C.1)
ε ∥dj f (u; v1 , ..., vj ) − dj g(u; v1 , ..., vj )∥Y < . j=0,...,k (u,v1 ,...,vj )∈U ×X j ψj (u, v1 , ..., vj ) 2
∥f − g∥BΨk (U ;Y ) = max
sup
Moreover, by using (2.2), there exists some R > 0 such that max
(C.2)
sup
j=0,...,k (u,v1 ,...,vj )∈(U ×X j )\Kj,R
pX (v1 ) · · · pX (vj ) ε < . ψj (u, v1 , ..., vj ) 2 max(Cg , 1)
Hence, by using the inequalities (C.1), (2.2), and (C.2), it follows that ∥dj f (u; v1 , ..., vj )∥Y j=0,...,k (u,v1 ,...,vj )∈(U ×X j )\Kj,R ψj (u, v1 , ..., vj ) max
sup
∥dj f (u; v1 , ..., vj ) − dj g(u; v1 , ..., vj )∥Y j=0,...,k (U ×X j )\Kj,R ψj (u, v1 , ..., vj )
≤ max
sup
∥dj g(u; v1 , ..., vj )∥Y j=0,...,k (u,v1 ,...,vj )∈(U ×X j )\Kj,R ψj (u, v1 , ..., vj )
+ max
sup
≤ ∥f − g∥BΨk (U ;Y ) + Cg max
sup
j=0,...,k (u,v1 ,...,vj )∈(U ×X j )\Kj,R
<
ε ε + Cg ≤ ε. 2 2 max(Cg , 1)
pX (v1 ) · · · pX (vj ) ψj (u, v1 , ..., vj )
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
61
Since ε > 0 was chosen arbitrarily small, we obtain (2.6). Moreover, with g ∈ Cbk (U ; Y ) as above, we observe that for every fixed j = 0, ..., k and R > 0 it holds that ∥dj f (u; v1 , ..., vj ) − dj g(u; v1 , ..., vj )∥Y
sup (u,v1 ,...,vj )∈Kj,R
∥dj f (u; v1 , ..., vj ) − dj g(u; v1 , ..., vj )∥Y ψj (u, v1 , ..., vj ) (u,v1 ,...,vj )∈Kj,R ε ≤ R∥f − g∥BΨk (U ;Y ) < R . 2 This implies that Kj,R ∋ (u, v1 , ..., vj ) 7→ dj f (u; v1 , ..., vj ) ∈ Y is continuous as uniform limit of the k continuous mappings Kj,R ∋ (u, v1 , ..., vj ) 7→ dj g(u; v1 , ..., vj ) ∈ Y , showing that f ∈ Cloc (U ; Y ). For (ii), we first consider the case when (X, τX ) is not locally compact and (U, τX ) has AP with net of finite rank operators (Tγ )γ ⊆ X ∗ ⊗ X satisfying Tγ (U ) ⊆ U and (2.8). Moreover, we define −1 > 0 and fix some ε > 0. the constant Cinf := minj=0,...,k inf (u,v1 ,...,vj )∈U ×X j ψj (u, v1 , ..., vj ) Then, by using (2.7)–(2.8), there exists some R > 0 such that ≤R
ε ∥dj f (u; v1 , ..., vj )∥Y < , j=0,...,k (u,v1 ,...,vj )∈(U ×X j )\Kj,R ψj (u, v1 , ..., vj ) 3 max
(C.3) (C.4)
sup
sup γ
sup
ε ∥d|L| f (Tγ (u); (Tγ (vℓ ))ℓ∈L )∥Y < . kC j=0,...,k ψ (u, v ) 3 · 2 j L L Ψ L⊆{1,...,j} (u,v1 ,...,vj )∈(U ×X )\Kj,R max
sup
Now, we use that Kj,R ∋ (u, v1 , ..., vj ) 7→ dj f (u; v1 , ..., vj ) ∈ Y is continuous, thus uniformly continuous on the compact set Kj,R , to conclude that there exists an open 0-neighborhood Vj,0 × j Vj,1 ×...×Vj,j of (X ×X j , τX ×τX ), with Vj,0 , ..., Vj,j ∈ τX , such that for every (u, v1 , ..., vj ) ∈ Kj,R and (e u, ve1 , ..., vej ) ∈ U × X j with (C.5)
(u, v1 , ..., vj ) − (e u, ve1 , ..., vej ) ∈ Vj,0 × Vj,1 × ... × Vj,j ,
it holds that (C.6)
∥dj f (u; v1 , ..., vj ) − dj f (e u; ve1 , ..., vej )∥Y <
ε . 3Cinf
S k Sj From this, we define the set K := j=0 ℓ=0 πℓ (Kj,R ) that is compact as finite union of compact images under the continuous projection U × X j ∋ (v0 , ..., vℓ ) 7→ πℓ (v0 , ..., vj ) := vℓ ∈ X. Then, by using that (U, τX ) has AP, there exists some Tγ ∈ X ∗ ⊗ X such that for every x ∈ K we have Tk Tj x − Tγ (x) ∈ j=0 ℓ=0 Vj,ℓ . Hence, for every (u, v1 , ..., vj ) ∈ Kj,R , it holds that (u, v1 , ..., vj ) − (Tγ (u), Tγ (v1 ), ..., Tγ (vj )) ∈ Vj,0 × Vj,1 × ... × Vj,j , Thus, by combining this with (C.5) as well as using the chain rule and (C.6), it follows for every j = 0, ..., k and (u, v1 , ..., vj ) ∈ Kj,R that ∥dj f (u; v1 , ..., vj ) − dj (f ◦ Tγ )(u; v1 , ..., vj )∥Y (C.7)
= ∥dj f (u; v1 , ..., vj ) − dj f (Tγ (u); Tγ (v1 ), ..., Tγ (vj ))∥Y <
ε . 3Cinf
Qj Now, for every j = 0, ..., k and R > 0, we observe that the map ℓ=0 Tγ (πℓ (Kj,R )) ∋ (u, v1 , ..., vj ) 7→ dj f (u; v1 , ..., vj ) ∈ Y is continuous as restriction of the continuous map Kj,R ∋ (u, v1 , ..., vj ) 7→ Qj dj f (u; v1 , ..., vj ) ∈ Y , where ℓ=0 Tγ (πℓ (Kj,R )) is compact in Tγ (U ) × Tγ (X)j . Hence, by using that the finite-dimensional spaces Tγ (U ) and Tγ (X) are locally compact, it follows from [85, Lemma 46.3+46.4] that Tγ (U ) × Tγ (X)j ∋ (u, v1 , ..., vj ) 7→ dj f (u; v1 , ..., vj ) ∈ Y is globally continuous. Since Tγ ∈ X ∗ ⊗ X is also continuous, we conclude that U × X j ∋ (u, v1 , ..., vj ) 7→ dj (f ◦ Tγ )(u; v1 , ..., vj ) = dj f (Tγ (u); Tγ (v1 ), ..., Tγ (vj )) ∈ Y is continuous, which shows that
62
P. SCHMOCKER AND J. TEICHMANN
f ◦ Tγ ∈ C k (U ; Y ). Moreover, by using again that Tγ (U ) is locally compact, there exists some h ∈ Cc∞ (Tγ (X)) such that Sk h(u) = 1, for all u ∈ j=0 Tγ (π0 (Kj,R )), 0 ≤ h(u) ≤ 1, for all u ∈ Tγ (U ), (C.8) for all j = 1, ..., k and j d (h ◦ Tγ )(u; v1 , ..., vj ) ≤ ψj (u, v1 , ..., vj ), (u, v1 , ..., vj ) ∈ U × X j , From this, we define the map U ∋ u 7→ g(u) := (h ◦ Tγ )(u)(f ◦ Tγ )(u) ∈ Y . Then, by using the Leibniz product rule, the monotonicity of Ψ := (ψj )j=0,...,k , the last property of (C.8), the chain rule, and (C.4), it holds for every j = 0, ..., k and (u, v1 , ..., vj ) ∈ (U × X j ) \ Kj,R that P j−|L| (h ◦ Tγ )(u; vLc )d|L| (f ◦ Tγ )(u; vL ) Y ∥dj g(u; v1 , ..., vj )∥Y L⊆{1,...,l} d = ψj (u, v1 , ..., vj ) ψj (u, v1 , ..., vj ) P j−|L| (h ◦ Tγ )(u; vLc )|∥d|L| (f ◦ Tγ )(u; vL )∥Y L⊆{1,...,j} |d ≤ CΨ ψj−|L| (u, vLc )ψL (u, vL ) (C.9)
≤ 2k CΨ
∥d|L| (f ◦ Tγ )(u; vL )∥Y ψL (u, vL ) L⊆{1,...,j} max
∥d|L| f (Tγ (u); (Tγ (vℓ ))ℓ∈L )∥Y ψL (u, vL ) L⊆{1,...,j} ε ε k < 2 CΨ = . 3 · 2 k CΨ 3
≤ 2k CΨ
max
Hence, by using (C.3), (C.7), and (C.9), we conclude for g ∈ Cbk (U ; Y ) that ∥f − g∥BΨk (U ;Y ) = max
sup
j=0,...,k (u,v1 ,...,vj )∈U ×X j
≤ max
sup
j=0,...,k (u,v1 ,...,vj )∈(U ×X j )\Kj,R
+ Cinf max
sup
+ max
sup
j=0,...,k (u,v1 ,...,vj )∈Kj,R
dj f (u; v1 , ..., vj ) − dj g(u; v1 , ..., vj ) Y ψj (u, v1 , ..., vj )
dj f (u; v1 , ..., vj ) Y ψj (u, v1 , ..., vj ) dj f (u; v1 , ..., vj ) − dj g(u; v1 , ..., vj ) Y
j=0,...,k (u,v1 ,...,vj )∈(U ×X j )\Kj,R
dj g(u; v1 , ..., vj ) Y ψj (u, v1 , ..., vj )
ε ε ε < + Cinf + = ε. 3 3Cinf 3 k Since ε > 0 was chosen arbitrarily and BΨ (U ; Y ) is defined as closure of Cbk (U ; Y ) with respect to k ∥ · ∥BΨk (U ;Y ) , we conclude that f ∈ BΨ (U ; Y ). In the other case, if (X, τX ) is locally compact, we do not need to concatenate with finite rank operators and can directly obtain some h ∈ Cc∞ (X) satisfying (C.8), i.e., we can replace Tγ : X → X with the identity idX : X → X. □
Appendix D. Proof of results in Section 3 D.1. Proof of Lemma 3.6. For more background on Banach space-valued real-analytic functions, we refer to [31, Chapter IX]. Proof of Lemma 3.6. Let V := {z ∈ C : | Im(z)| < λ/2}, fix some distinct z0 , z1 ∈ C, and (j) define the function R ∋ s 7→ hz0 (s) := s cos(z0 s) ∈ C satisfying hz0 (s) = s cos(j) (z0 s)z0j + j cos(j−1) (z0 s)z0j−1 . Then, by splitting into cosine change and monomial change, by applying
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
63
Taylor’s theorem to the holomorphic functions cos : C → C and C ∋ z 7→ z j ∈ C (see also [28, P∞ ( λ |s|)n ( λ |s|)2 Equation 3.1]), and by using that |s|2 = λ82 2 2! ≤ λ82 n=1 2 n! ≤ λ82 exp λ2 |s| together with 1 ≤ exp(λ|s|), it follows for every j = 0, ..., k and s ∈ R that cos(j) (z1 s) z1j − cos(j) (z0 s) z0j − h(j) z0 (s) z1 − z0 Z 1 cos(j) ((z0 + t(z1 − z0 ))s) − cos(j) (z0 s) dt + sz1j cos(j) (z0 s) − sz0j cos(j) (z0 s) ≤ sz1j 0
+ j cos
(j−1)
(z0 s)
Z 1
(z0 + t(z1 − z0 ))j−1 − z0j−1 dt
0
Z 1
≤ |s||z1 |j
0 Z 1
cos(j) ((z0 + t(z1 − z0 ))s) − cos(j) (z0 s) dt + |s| z1j − z0j
(z0 + t(z1 − z0 ))j−1 − z0j−1 dt
+j 0
Z 1
|s(z1 −z0 )|e|s| max(| Im(z0 +t(z1 −z0 ))|,| Im(z0 )|) dt+j|z0 |j−1 |z1 −z0 |+j(j − 1)|z0 |j−2 |z1 −z0 | 0 λ ≤ (1 + |z0 | + |z1 |)k |s|2 e 2 |s| |z1 − z0 | + 2k 2 |z1 − z0 | 8 k 2 ≤ (1 + |z0 | + |z1 |) + 2k |z1 − z0 |eλ|s| . λ2 j
≤ |s||z1 |
Hence, the Faà di Bruno formula implies for every j = 0, ..., k and (u, v1 , ..., vj ) ∈ U × (Rd )j that cos (z1 ge(·)) − cos (z0 ge(·)) dj (u; v1 , ..., vj ) − dj (e g (·) cos′ (z0 ge(·))) (u; v1 , ..., vj ) z1 − z0 ≤
X cos(|π|) (z1 ge(u)) z |π| − cos(|π|) (z0 ge(u)) z |π| 1
z1 − z0
π∈Pj
≤
X
0
|π| |π| cos(|π|) (z1 ge(u)) z1 − cos(|π|) (z0 ge(u)) z0
z1 − z0
π∈Pj
≤ k!k 2 (1 + |z0 | + |z1 |)k
dπ ge(u; vπ ) −
X
h(|π|) g (u))dπ ge(u; vπ ) z0 (e
π∈Pj
− h(|π|) g (u)) |dπ ge(u; vπ )| z0 (e
8|z1 − z0 | λ|g(u)| e max |dπ ge(u; vπ )| . π∈Pj λ2
Thus, by using (M2) together with continuity of U ×(Rd )j ∋ (u, v1 , ..., vj ) 7→ exp (λ|g(u)|) |dπ ge(u; vπ )| on the compact pre-images Kj,R , for all R > 0, we conclude that cos (z1 ge(·)) − cos (z0 ge(·)) − ge(·) cos′ (z0 ge(·)) z1 − z0 k (U ) BΨ 8 exp (λ|g(u)|) |dπ ge(u; vπ )| 2 ≤ k!(1 + |z0 | + |z1 |)k + 2k |z − z | max sup 1 0 j=0,...,k λ2 ψj (u, v1 , ..., vj ) d j π∈Pj (u,v1 ,...,vj )∈U ×(R ) | {z } <∞
z1 →z0
−→
0,
k which shows that V ∋ z 7→ cos(ze g (·)) ∈ BΨ (U ) is holomorphic, implying that the mapping k k R ∋ t 7→ cos(te g (·)) ∈ BΨ (U ) is real-analytic. By a similar argument, V ∋ z 7→ sin(ze g (·)) ∈ BΨ (U ) k is holomorphic, which ensures that R ∋ t 7→ sin(te g (·)) ∈ BΨ (U ) is real-analytic. □
64
P. SCHMOCKER AND J. TEICHMANN
D.2. Proof of Lemma 3.11. Proof of Lemma 3.11. Since f ∈ Cbk (U ; Y ), there exists some Cf ≥ 0 and pX ∈ P(X,τX ) such that for every j = 0, ..., k and (u, v1 , ..., vj ) ∈ U × X j it holds that dj f (u; v1 , ..., vj ) Y ≤ Cf pX (v1 ) · · · pX (vj ).
(D.1)
Moreover, by using that (U, τX ) has QX -BAP, there exists a net of finite rank operators (Tγ )γ ⊆ X ∗ ⊗ X with Tγ (U ) ⊆ U approximating the identity idX : X → X uniformly on each relatively compact subset of (X, τX ) such that for every pX ∈ P(X,τX ) there exists some λ > 0 and qX ∈ QX satisfying for every γ and v ∈ X that pX (Tγ (v)) ≤ λqX (v).
(D.2)
In addition, by using (2.2), there exists some R > 0 such that
(D.3)
max
sup
ε pX (v1 ) · · · pX (vj ) < , ψj (u, v1 , ..., vj ) 3(1 + Cf )
max
sup
ε qX (v1 ) · · · qX (vj ) < k . ψj (u, v1 , ..., vj ) 3λ (1 + Cf )
j=0,...,k (u,v1 ,...,vj )∈(U ×X j )\Kj,R
j=0,...,k (u,v1 ,...,vj )∈(U ×X j )\Kj,R
Furthermore, we observe that Kj,R ∋ (u, v1 , ..., vj ) 7→ dj f (u; v1 , ..., vj ) ∈ Y is continuous, thus uniformly continuous on the compact set Kj,R , to conclude that there exists an open 0-neighborhood j Vj,0 ×Vj,1 ×...×Vj,j of (X ×X j , τX ×τX ), with Vj,0 , ..., Vj,j ∈ τX , such that for every (u, v1 , ..., vj ) ∈ Kj,R and (e u, ve1 , ..., vej ) ∈ U × X j with (D.4)
(u, v1 , ..., vj ) − (e u, ve1 , ..., vej ) ∈ Vj,0 × Vj,1 × ... × Vj,j ,
it holds that (D.5)
dj f (u; v1 , ..., vj ) − dj f (e u; ve1 , ..., vej ) Y <
ε . 3Cinf
Sk Sj From this, we define the set K := j=0 ℓ=0 πℓ (Kj,R ), which is compact as finite union of compact images under the continuous projection U ×X j ∋ (v0 , ..., vℓ ) 7→ πℓ (v0 , v1 , ..., vj ) := vℓ ∈ X. Hence, there exists some γ such that for every (u, v1 , ..., vj ) ∈ Kj,R , we have (u, v1 , ..., vj ) − (Tγ (u), Tγ (v1 ), ..., Tγ (vj )) ∈ Vj,0 × Vj,1 × ... × Vj,j . Thus, by combining this with (D.4), we conclude from the chain rule and (D.5) that
(D.6)
dj f (u; v1 , ..., vj ) − dj (f ◦ Tγ )(u; v1 , ..., vj ) Y = dj f (u; v1 , ..., vj ) − dj f (Tγ (u); Tγ (v1 ), ..., Tγ (vj )) Y <
ε . 3Cinf
Moreover, by using (D.1)–(D.2), we observe for every (u, v1 , ..., vj ) ∈ U × Xij that dj (f ◦ Tγ )(u; v1 , ..., vj ) Y = dj f (Tγ (u); Tγ (v1 ), ..., Tγ (vj )) Y (D.7)
≤ Cf pX (Tγ (v1 )) · · · pX (Tγ (vj )) ≤ Cf λj qX (v1 ) · · · qX (vj ).
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
65
Finally, by combining the inequalities (D.1), (D.3), (D.6), and (D.7), it follows that ∥f − f ◦ Tγ ∥BΨk (U ;Y ) = max
sup
j=0,...,k (u,v1 ,...,vj )∈U ×X j
≤ max
dj f (u; v1 , ..., vj ) − dj (f ◦ Tγ )(u; v1 , ..., vj ) Y ψj (u, v1 , ..., vj )
dj f (u; v1 , ..., vj ) Y ψj (u, v1 , ..., vj )
sup
j=0,...,k (u,v1 ,...,vj )∈(U ×X j )\Kj,R
dj f (u; v1 , ..., vj ) − dj (f ◦ Tγ )(u; v1 , ..., vj ) Y
+ Cinf max
sup
+ max
sup
dj (f ◦ Tγ )(u; v1 , ..., vj ) Y ψj (u, v1 , ..., vj )
< Cf max
sup
pX (v1 ) · · · pX (vj ) ε + Cinf ψj (u, v1 , ..., vj ) 3Cinf
j=0,...,k (u,v1 ,...,vj )∈Kj,R
j=0,...,k (u,v1 ,...,vj )∈(U ×X j )\Kj,R
j=0,...,k (u,v1 ,...,vj )∈(U ×X j )\Kj,R
qX (v1 ) · · · qX (vj ) ψj (u, v1 , ..., vj ) ε ε ε + Cinf ≤ ε. < Cf + Cf λk k 3(1 + Cf ) 3Cinf 3λ (1 + Cf ) + Cf max
sup
j=0,...,k (u,v1 ,...,vj )∈(U ×X j )\Kj,R
Since ε > 0 was chosen arbitrarily, we obtain the conclusion.
□
D.3. Proof of Lemma 3.9. Proof of Lemma 3.9. Let (Y, ∥ · ∥Y ) have BAP (with constant λ ∈ [1, ∞)) and fix some ε > 0. Then, by Lemma 2.10 (i), there exists some R > 0 such that (D.8)
ε ∥dj f (u; v1 , ..., vj )∥Y < . j=0,...,k (u,v1 ,...,vj )∈(U ×X j )\Kj,R ψj (u, v1 , ..., vj ) 2(1 + λ) max
sup
−1 >0 Moreover, we define the constant Cinf := minj=0,...,k inf (u,v1 ,...,vj )∈U ×X j ψj (u, v1 , ..., vj ) Sk j and the set K := j=0 d f (u; v1 , ..., vj ) : (u, v1 , ..., vj ) ∈ Kj,R , which is compact as finite union of continuous images of compact sets. Then, by using that (Y, ∥ · ∥Y ) has BAP (with constant PN λ ∈ [1, ∞)), there exists some y 7→ T (y) := n=1 ℓn (y)yn ∈ Y ∗ ⊗ Y , with ℓ1 , ..., ℓN ∈ Y ∗ and y1 , ..., yN ∈ Y , satisfying ∥T ∥L(Y ;Y ) ≤ λ such that sup ∥y − T (y)∥Y < y∈K
ε . 2Cinf
This together with the chain rule implies that max
sup
j=0,...,k (u,v1 ,...,vj )∈Kj,R
= max
sup
dj f (u; v1 , ..., vj ) − dj (T ◦ f )(u; v1 , ..., vj ) Y
j=0,...,k (u,v1 ,...,vj )∈Kj,R
≤ sup ∥y − T (y)∥Y < y∈K
dj f (u; v1 , ..., vj ) − T dj f (u; v1 , ..., vj ) Y
ε . 2Cinf
In addition, ∥T ∥L(Y ;Y ) ≤ λ ensures for every j = 0, ..., k and (u, v1 , ..., vj ) ∈ U × X j that dj (T ◦ f )(u; v1 , ..., vj ) Y = T dj f (u; v1 , ..., vj ) Y (D.9)
≤ ∥T ∥L(Y ;Y ) dj f (u; v1 , ..., vj ) Y ≤ λ dj f (u; v1 , ..., vj ) Y .
66
P. SCHMOCKER AND J. TEICHMANN
Furthermore, we claim for every fixed n = 1, ..., N that ℓn ◦ f ∈ BkΨ (U ). Indeed, by definition, f ∈ BkΨ (U ; Y ) can be approximated by a sequence (gm )m∈N ⊆ Cbk (U ; Y ) with respect to ∥·∥Ψk (U ;Y ) , whence (ℓn ◦ gm )m∈N ⊆ Cbk (U ) approximates the function ℓn ◦ f : U → R with respect to ∥ · ∥Ψk (U ) , ensuring that ℓn ◦ f ∈ BkΨ (U ). Thus, (D.8)–(D.9) imply that dj f (u; v1 , ..., vj ) − dj (f ◦ T )(u; v1 , ..., vj ) Y ψj (u, v1 , ..., vj )
∥f − T ◦ f ∥BΨk (U ;Y ) = max
sup
≤ max
dj f (u; v1 , ..., vj ) Y ψj (u, v1 , ..., vj )
j=0,...,k (u,v1 ,...,vj )∈U ×X j
sup
j=0,...,k (u,v1 ,...,vj )∈(U ×X j )\Kj,R
+ Cinf max
sup
+ max
sup
j=0,...,k (u,v1 ,...,vj )∈Kj,R
dj f (u; v1 , ..., vj ) − dj (T ◦ f )(u; v1 , ..., vj ) Y dj (T ◦ f )(u; v1 , ..., vj ) Y ψj (u, v1 , ..., vj )
j=0,...,k (u,v1 ,...,vj )∈(U ×X j )\Kj,R
< (1 + λ) max
sup
j=0,...,k (u,v1 ,...,vj )∈(U ×X j )\Kj,R
dj f (u; v1 , ..., vj ) Y ε + Cinf < ε. ψj (u, v1 , ..., vj ) 2Cinf
k Since f ∈ BΨ (U ; Y ) and ε > 0 were chosen arbitrarily, we obtain the conclusion.
□
Appendix E. Proof of results in Section 4 E.1. Auxiliary lemma for the proof of Theorem 4.11. Lemma E.1. Let (U, Ψ) be a weighted domain and let (Y, ∥ · ∥Y ) be a Banach space. Moreover, k (U ) satisfies for every a ∈ A that for c ∈ (0, ∞), let ρ ∈ Bck (R) and assume that A ⊆ BΨ c
(E.1)
(1 + |a(u)|) |dπ a(u; vπ )| < ∞. j=0,...,k ψj (u, v1 , ..., vj ) (u,v1 ,...,vj )∈U ×X j π∈P
Ca := max
sup
j
k In addition, let L ⊆ Y be a vector subspace. Then, N N A,ρ,L ⊆ BΨ (U ; Y ). U,Y
Proof. Since N N A,ρ,L is defined as linear span of maps of the form U ∋ u 7→ yρ(a(u) + b) ∈ Y , U,Y k k with a ∈ A and y ∈ Y , and the mapping BΨ (U ) ∋ g 7→ y · g ∈ BΨ (U ; Y ) is well-defined and k continuous, it suffices to prove that ρ ◦ (a(·) + b) ∈ BΨ (U ). To this end, we fix some a ∈ A, b ∈ R, and ε ∈ (0, 1). Then, by using that Bck (R) is defined as closure of Cbk (R) with respect to ∥ · ∥Bck (R) , there exists some ρe ∈ Cbk (R) such that (E.2)
ρ(j) (z) − ρe(j) (z) ε < , c j=0,...,k z∈R (1 + |z|) 2k!Ca
∥ρ − ρe∥Bck (R) = max sup
where Ca > 0 is defined in (E.1). Moreover, for ∥e ρ∥Cbk (R) := maxj=0,...,k sups∈R |e ρ(j) (s)| < ∞, we k k use that a ∈ A ⊆ BΨ (U ) to obtain from the definition of BΨ (U ) some e a ∈ Cbk (U ) such that dj a(u; v1 , ..., vj ) − dj e a(u; v1 , ..., vj ) j=0,...,k (u,v1 ,...,vj )∈U ×X j ψj (u, v1 , ..., vj ) ε < k , k 1 + ∥a∥ k 2k · k!CΨ 1 + ∥e ρ∥Cbk (R) BΨ (U )
∥a − e a∥BΨk (U ) = max
sup
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
67
which implies that ∥e a∥BΨk (U ) ≤ 1 + ∥a∥BΨk (U ) . Hence, by using a telescoping sum together with the monotonicity of Ψ := (ψj )j=0,...,k , it follows that (E.3) |dπ a(u; vπ ) − dπ e a(u; vπ )| max sup j=0,...,k ψ (u, v , ..., vj ) j j 1 π∈Pj (u,v1 ,...,vj )∈U ×X P|π| Qr−1 πℓ Q|π| πr πr a(u; vπr )) ℓ=r dπℓ e a(u; vπℓ ) r=1 ℓ=1 d a(u; vπℓ )(d a(u; vπr )−d e k ≤ CΨ max sup j=0,...,k ψ (u, v ) · · · ψ (u, v ) j π1 π1 π|π| π|π| (u,v1 ,...,vj )∈U ×X π∈P j
|dπr a(u; vπr ) − dπr e a(u; vπr )| j=0,...,k r=1,...,|π| ψπr (u, vπr ) (u,v1 ,...,vj )∈U ×X j π∈P
k ≤ kCΨ (1 + ∥a∥BΨk (U ) )k max
max
sup
j
k ≤ kCΨ (1 + ∥a∥BΨk (U ) )k
ε k 2k · k!CΨ
1 + ∥a∥BΨk (U )
k
1 + ∥e ρ∥Cbk (R)
=
ε . 2k! 1 + ∥e ρ∥Cbk (R)
Thus, the Faà di Bruno formula and (E.2)–(E.3) imply for ρe ◦ (e a(·) + b) ∈ Cbk (U ) that ∥ρ ◦ (a(·) + b) − ρe ◦ (e a(·) + b)∥BΨk (U ) ≤ ∥ρ ◦ (a(·) + b) − ρe ◦ (a(·) + b)∥BΨk (U ) + ∥e ρ ◦ (a(·) + b) − ρe ◦ (e a(·) + b)∥BΨk (U ) dj (ρ ◦ (a(·) + b))(u; v1 , ..., vj ) − dj (e ρ ◦ (a(·) + b))(u; v1 , ..., vj ) j=0,...,k (u,v1 ,...,vj )∈U ×X j ψj (u, v1 , ..., vj )
= max
sup
dj (e ρ ◦ (a(·) + b))(u; v1 , ..., vj ) − dj (e ρ ◦ (e a(·) + b))(u; v1 , ..., vj ) j=0,...,k (u,v1 ,...,vj )∈U ×X j ψj (u, v1 , ..., vj ) P (|π|) (a(u) + b)dπ a(u; vπ ) − ρe(|π|) (a(u) + b)dπ a(u; vπ ) π∈Pj ρ ≤ max sup j=0,...,k (u,v1 ,...,vj )∈U ×X j ψj (u, v1 , ..., vj ) P e(|π|) (a(u) + b)dπ a(u; vπ ) − ρe(|π|) (e a(u) + b)dπ e a(u; vπ ) π∈Pj ρ + max sup j=0,...,k (u,v1 ,...,vj )∈U ×X j ψj (u, v1 , ..., vj ) + max
sup
ρ(|π|) (a(u) + b) − ρe(|π|) (a(u) + b) j=0,...,k (1 + |a(u) + b|)c (u,v1 ,...,vj )∈U ×X j π∈P
≤ Ca k! max
sup
j
|dπ a(u; vπ ) − dπ e a(u; vπ )| j=0,...,k ψ (u, v , ..., vj ) j 1 π∈P
+ k!∥e ρ∥Cbk (R) max j
ε ε ≤ ε. < Ca k! + k!∥e ρ∥Cbk (R) 2Ca k! 2k! 1 + ∥e ρ∥C k (R) b
k Since ε ∈ (0, 1) was chosen arbitrarily and BΨ (U ) is defined as closure of Cbk (U ) with respect to k ∥ · ∥BΨk (U ) , this shows that ρ ◦ (a(·) + b) ∈ BΨ (U ). □
Appendix F. Proof of results in Section 5 F.1. Proof of Lemma 5.4. k Proof of Lemma 5.4. First, we show that (Λα,r T,Z , Ψ) is a weighted Cloc -manifold. To this end, we α,r conclude from Theorem A.2 that (R × D ([0, T ]; Z), ∥ · ∥R×Dα,r ([0,T ];Z) ) is a dual Banach space with some predual (G, ∥ · ∥G ), where ∥(t, x)∥R×Dα,r ([0,T ];Z) := |t| + ∥x∥α,ℓr . Thus, (ϕi (Λα,r T,Z ), Ψi ) is k by Lemma 2.3 (ii) a weighted domain, whence (Λα,r , Ψ) is by definition a weighted C loc -manifold. T,Z α,r ∗ Now, we prove that (ϕi (ΛT,Z ), τR ×τw ) has ∥·∥R×Dα,r ([0,T ];Z) -BAP. Indeed, since (Dα,r ([0, T ]; Z), ∥· ∥α,ℓr ) is by Theorem A.2 a dual Banach space with some predual (G, ∥ · ∥G ), we can apply
68
P. SCHMOCKER AND J. TEICHMANN
Theorem B.5 to conclude that (Dα,r ([0, T ]; Z), τw∗ ) has ∥ · ∥α,ℓr -BAP with finite rank operators (Q∗γ )γ ⊆ (Dα,r ([0, T ]; Z), τw∗ )∗ ⊗ Dα,r ([0, T ]; Z), where (Qγ )γ ⊆ G∗ ⊗ G is the BAP net of (G, ∥ · ∥G ) with ∥Qγ ∥L(G;G) ≤ CG , for some CG ≥ 1. From this, we define the finite rank operators (Rγ )γ ⊆ (R × Dα,r ([0, T ]; Z), τR × τw∗ )∗ ⊗ (R × Dα,r ([0, T ]; Z)) by
R × Dα,r ([0, T ]; Z) ∋ (t, x)
7→
Rγ (t, x) := t, Q∗γ (x) ∈ R × Dα,r ([0, T ]; Z).
Then, for every (t, x) 7→ pR×Dα,r ([0,T ];Z) (t, x) := |t| + pDα,r ([0,T ];Z) (x) ∈ P(R×Dα,r ([0,T ];Z),τR ×τw∗ ) with pDα,r ([0,T ];Z) ∈ P(Dα,r ([0,T ];Z),τw∗ ) , we conclude for every relatively compact subset K of (R × Dα,r ([0, T ]; Z), τR × τw∗ ) that
lim sup pR×Dα,r ([0,T ];Z) ((t, x) − Rγ (t, x)) = lim sup pR×Dα,r ([0,T ];Z) (t, x) − t, Q∗γ (x) γ (t,x)∈K
γ (t,x)∈K
= lim sup
γ (t,x)∈K
|t − t| + pDα,r ([0,T ];Z) x − Q∗γ (x)
= lim sup pDα,r ([0,T ];Z) x − Q∗γ (x) = 0, γ (t,x)∈K
which shows that (ϕi (Λα,r T,Z ), τR ×τw∗ ) has AP. In addition, it holds for every g1 , ..., gN ∈ G, (t, x)∈ R×Dα,r ([0, T ]; Z), and (t, x) 7→ peR×Dα,r ([0,T ];Z) (t, x) := |t|+maxn=1,...,N |⟨x, gn ⟩Dα,r ([0,T ];Z)×G | ∈ P(R×Dα,r ([0,T ];Z),τR ×τw∗ ) that
peR×Dα,r ([0,T ];Z) (Rγ (t, x)) = |t| + max |⟨Q∗γ (x), gn ⟩Dα,r ([0,T ];Z)×G | n=1,...,N
= |t| + max |⟨x, Qγ (gn )⟩Dα,r ([0,T ];Z)×G | n=1,...,N
≤ |t| + ∥x∥α,ℓr ∥Qγ ∥L(E;E) max ∥gn ∥G n=1,...,N = 1 + CG max ∥gn ∥G ∥(t, x)∥R×Dα,r ([0,T ];Z) , n=1,...,N
which proves that (ϕi (Λα,r T,Z ), τR × τw∗ ) has ∥ · ∥R×D α,r ([0,T ];Z) -BAP. k In order to show that A ⊆ BΨ (Λα,r T,Z ) is well-defined, we aim to apply Proposition 2.10 (ii). To α,r ρ,E this end, we fix some λ ∈ R as well as φ e ∈ N N R,e R,E and consider the map ΛT,Z ∋ (t, x) 7→ a(t, x) := RT t t e λt + 0 ⟨xs , φ(s)⟩ Z×E ds ∈ R. Then, for every j = 0, ..., k and (t, x ), (s1 , v1 ), ..., (sj , vj ) ∈ α,r ϕi (Λα,r ([0, T ]; Z)), it holds that T,Z ) × (R × D
(F.1)
dj
RT λt + 0 ⟨xts , φ(s)⟩ e if j = 0, Z×E ds, RT −1 a ◦ ϕi ((t, x); (s1 , v1 ), ..., (sj , vj )) = λs1 + 0 ⟨v1,s , φ(s)⟩ e Z×E ds, if j = 1, 0, if j ≥ 2.
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
69
α which shows that a ∈ C k (ϕi (Λα,r T,Z )). Moreover, by using that ∥x∥∞ ≤ max(1, T ) ∥x∥α for any R T x ∈ Dα,r ([0, T ]; Z), and the constant C := 1 + |λ| + max(1, T )α CG 0 ∥φ(s)∥ e E ds > 0, we have −1 j t d a ◦ ϕi (Rγ (t, x ); Rγ (s1 , v1 ), ..., Rγ (sj , vj )) Z T Z T ⟨Q∗γ (xt )s , φ(s)⟩ ⟨Q∗γ (v1 )s , φ(s)⟩ ≤ |λ|(|t| + |s1 |) + e ds + e Z×E Z×E ds 0 0 ! Z T α ≤ |λ|(|t| + |s1 |) + ∥φ(s)∥ e ∥Q∗γ (xt )∥α,ℓr + ∥Q∗γ (v1 )∥α,ℓr E ds max(1, T ) 0
(F.2) ≤ |λ|(|t| + |s1 |) + max(1, T )α CG
0 t
≤ C |t| + ∥x ∥α + |s1 | + ∥v1 ∥α ≤C
!
Z T
∥φ(s)∥ e E ds
∥xt ∥α,ℓr + ∥v1 ∥α,ℓr
max(j, 1)∥(t, xt )∥R×Dα,r ([0,T ];Z) +
j X
! ∥(sℓ , vℓ )∥R×Dα,r ([0,T ];Z)
ℓ=1 k
r Therefore, by using this together with the assumption limr→∞ η(r) = 0, we obtain that d|L| a ◦ ϕ−1 (Rγ (t, xt ); (Rγ (sℓ , vℓ ))ℓ∈L ) i lim max sup R→∞ j=0,...,k ψi,|L| (t, xt ), (sℓ , vℓ )ℓ∈L L⊆{1,...,j} (t,xt ),(s1 ,v1 ),...,(sj ,vj ) ∈K c ij,R P max(|L|, 1)∥(t, xt )∥R×Dα,r ([0,T ];Z) + ℓ∈L ∥(sℓ , vℓ )∥R×Dα,r ([0,T ];Z) P ≤ C lim max sup R→∞ j=0,...,k K c η max(|L|, 1)∥(t, xt )∥R×Dα,r ([0,T ];Z) + ℓ∈L ∥(sℓ , vℓ )∥R×Dα,r ([0,T ];Z) i,j,R L⊆{1,...,j} r = C lim = 0, r→∞ η(r) α,r where the supremum is taken over (t, xt ), (s1 , v1 ), ..., (sj , vj ) ∈ (ϕi (Λα,r ([0, T ]; Z))j )\ T,Z )×(R×D α,r −1 k k (Λα,r Ki,j,R . Hence, Proposition 2.10 (ii) implies a ◦ ϕi ∈ BΨi (ϕi (ΛT,Z )), showing that A ⊆ BΨ T,Z ). α,r −1 Finally, we prove that A is an additive family on ΛT,Z by showing that Ai = a ◦ ϕi : a ∈ A is an additive family on ϕi (Λα,r T,Z ). For (A1), we observe that Ai is a vector space and therefore closed under addition. For (A2), let (t1 , xt11 ), (t2 , xt22 ) ∈ ϕi (Λα,r T,Z ) be two distinct points. If −1 t t t1 ̸= t2 , the map (t, x ) 7→ (a ◦ ϕi )(t, x ) := t ∈ Ai satisfies a(t1 , xt11 ) = t1 ̸= t2 = a(t2 , xt22 ). Otherwise, if t1 = t2 =: t and xt1 = xt2 on [0, t) but xt1,t ̸= xt2,t , there exists some e ∈ E such that RT ρ,R ⟨xt1,t , e⟩Z×E ̸= ⟨xt2,t , e⟩Z×E . Thus, there exists some φ e0 ∈ N N R,e e0 (s)ds ̸= 0, whence R,R with t φ RT t −1 t t (t, x ) 7→ (a ◦ ϕi )(t, x ) := 0 ⟨xs , φ e0 (s)e⟩Z×E ds ∈ Ai satisfies (F.3) Z T Z t Z T D E t1 t t t a(t1 , x1 ) = ⟨x1,s , φ e0 (s)e⟩Z×E ds = ⟨x1,s , φ e0 (s)e⟩Z×E ds + x1,t , e φ e0 (s)ds 0
Z t ̸= 0
0
t
Z T D E t t ⟨x2,s , φ e0 (s)e⟩Z×E ds + x2,t , e φ e0 (s)ds t
Z×E
Z T = 0
Z×E
⟨xt2,s , φ e0 (s)e⟩Z×E ds = a(t2 , xt22 ).
Otherwise, if t1 = t2 =: t and xt1,t = xt2,t but xt1 ̸= xt2 on [0, t), there exists some e ∈ E such that ρ,R is weakly s 7→ ⟨xt1,s , e⟩Z×E differs from s 7→ ⟨xt2,s , e⟩Z×E on [0, t). Thus, by using that N N R,e R,R R t t R,e ρ,R 1 dense in L ([0, t]) (see [28, p. 31]), there exists some φ e0 ∈ N N R,R with 0 ⟨x1,s , φ e0 (s)e⟩Z×E ds ̸= RT t Rt t −1 t t ⟨x2,s , φ e0 (s)e⟩Z×E ds, whence (t, x ) 7→ (a ◦ ϕi )(t, x ) := 0 ⟨xs , φ e0 (s)e⟩Z×E ds ∈ Ai also 0 satisfies (F.3), which shows that Ai is point separating on ϕi (Λα,r T,Z ). For (A3), we fix some α,r t α,r (t, x ) ∈ ϕi (ΛT,Z ) and (s, v) ∈ R × D ([0, T ]; Z). Then, by applying the Hahn-Banach theorem,
70
P. SCHMOCKER AND J. TEICHMANN
RT ρ,R there exists some λ ∈ R, φ e0 ∈ N N R,e e0 (s)e⟩Z×E ds ̸= 0. R,R , and e ∈ E such that λs + 0 ⟨vs , φ RT t −1 t t Hence, the map (t, x ) 7→ (a ◦ ϕi )(t, x ) := λt + 0 ⟨xs , φ e0 (s)e⟩Z×E ds ∈ A satisfies Z T Z T −1 t d a ◦ ϕi ((t, x ); (s, v)) = λs + ⟨vt , φ(s)⟩ e ⟨vs , φ e0 (s)e⟩Z×E ds ̸= 0, Z×E dt = λs + 0
0
which shows that Ai has nowhere vanishing derivatives. For (A4’), we use for every γ that Tγ (ϕi (Λα,r T,Z )) is m-dimensional (with some m ∈ N), on which Ai is point separating and has nowhere vanishing derivatives, to obtain some a1 , ..., am ∈ A such that −1 ⊤ m ηi,γ := a1 ◦ ϕ−1 : Tγ (ϕi (Λα,r i , ..., am ◦ ϕi T,Z )) → R Tγ (ϕi (Λα,r )) T ,Z
is an embedding, where a1 , ..., am ∈ A have uniformly bounded derivatives ensuring that the corresponding limit in (A4) is finite (see also Remark 3.5). □ F.2. Proof of Corollary 5.5. ρ e,ρ,L Proof of Corollary 5.5. We aim to apply Theorem 4.13 to obtain that PN Λ = N N A,ρ,L α,r ,Y Λα,r ,Y T ,Z
T ,Z
α,r k is a dense subset of BΨ (Λα,r T,Z ; Y ). To this end, we fix some a ∈ A of the form ΛT,Z ∋ (t, x) 7→ RT t ρ,E e e ∈ N N R,e a(t, x) := λt + 0 ⟨xs , φ(s)⟩ Z×E ds ∈ R, with some λ ∈ R and φ R,E , and first show that the constant Ca,i > 0 defined in (4.5) is finite. Indeed, by using (F.1)–(F.2), we observe that c |π| 1 + a ◦ ϕ−1 (u) d a ◦ ϕ−1 (u; vπ ) i i sup Ca,i := max j=0,...,k α,r ψ (u, v ) α,r j i,j 1:j ([0,T ];Z)) π∈Pj (u,v1 ,...,vj )∈ϕi (ΛT ,Z )×(R×D max(j,c) Pj 1 + max(j, 1)∥u∥R×Dα,r ([0,T ];Z) + ℓ=1 ∥vℓ ∥R×Dα,r ([0,T ];Z) ≤ max sup Pj j=0,...,k (u,v1 ,...,vj ) η max(j, 1)∥u∥R×Dα,r ([0,T ];Z) + ℓ=1 ∥vℓ ∥R×Dα,r ([0,T ];Z)
≤
(1 + r)k max(1,c) < ∞. η(r) r∈(0,∞) sup
Next, we show (4.6), i.e., that for every ai := a ◦ ϕ−1 ∈ Ai := a ◦ ϕ−1 : a ∈ A , b ∈ R, and i i every finite rank operator (Tγ )g1:N ⊆ (R × Dα,r ([0, T ]; Z), τR × τw∗ )∗ ⊗ (R × Dα,r ([0, T ]; Z)) from i ,ρ,R Lemma 5.4 that the composition ρ((ai ◦ Tγ )(·) + b) belongs to the closure of N N A with ϕi (Λα,r T ,Z ),R P N respect to ∥ · ∥BΨk (ϕi (Λα,r , where Dα,r ([0, T ]; Z) ∋ x 7→ Tγ (x) := n=1 ⟨x, gn ⟩Dα,r ([0,T ];Z)×G xn ∈ T ,Z )) i α,r D ([0, T ]; Z), for some g1 , ..., gN ∈ G nd x1 , ..., xN ∈ Dα,r ([0, T ]; Z), where (G, ∥ · ∥G ) is a ρ,E predual of (Dα,r ([0, T ]; Z), ∥ · ∥α,ℓr ). To this end, we fix some λ ∈ R, φ e ∈ N N R,e R,E , b ∈ R, g1 , ..., gN ∈ G, and ε ∈ (0, 1). Then, for every fixed n = 1, ..., N , we apply the weighted UAT 0 in [28, Theorem 4.13] for BΨ -maps without derivatives (onto (Dα,r ([0, T ]; Z), τw∗ ) equipped with ρ,E the weight function (1 + ∥ · ∥α,ℓr )c ) to obtain some φ en ∈ N N R,e satisfying R,E (F.4)
sup x∈D α,r ([0,T ];Z)
RT ⟨x, gn ⟩Dα,r ([0,T ];Z) − 0 ⟨xs , φ en (s)⟩Z×E ds (1 + ∥x∥α
)c
<
ε , C1 Cη Cρ kN max(1, |rn |)
k k max(1,c) |(ai ◦Tγ )(u)| where C1 := 1+supu∈R×Dα,r ([0,T ];Z) (1+∥u∥R×D ≥ 1, Cη := 1+supr∈(0,∞) r η(r) ≥ α,r ([0,T ];Z) )c RT 1, Cρ := 1 + maxj=1,...,k+1 sups∈R |ρ(j) (s)| ≥ 1, and rn := 0 |⟨xn,s , φ(s)⟩ e Z×E |ds ≥ 0. From this, RT PN α,r we define D ([0, T ]; Z) ∋ x 7→ Qγ (x) := n=1 xn 0 ⟨xs , φ en (s)⟩Z×E ds ∈ Dα,r ([0, T ]; Z) and
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
71
R × Dα,r ([0, T ]; Z) ∋ (t, x) 7→ Rγ (t, x) := (t, Qγ (x)) ∈ R × Dα,r ([0, T ]; Z). Hence, for every (t, xt ) ∈ ϕi (Λα,r T,Z ), we observe that
(ai ◦ Rγ )(t, xt ) = λt +
Z T 0
= λt +
⟨Qγ (xt )s , φ(s)⟩ e Z×E ds
Z T N X 0
n=1
! Z T
⟨xts , φ en (s)⟩Z×E ds
0
! ⟨xn,s , φ(s)⟩ e Z×E ds ,
showing that ai ◦ Rγ ∈ Ai . Moreover, (F.4) ensures for every (t, x) ∈ R × Dα,r ([0, T ]; Z) that
|(ai ◦ Tγ )(t, x) − (ai ◦ Rγ )(t, x)| =
N X
λt +
⟨x, gn ⟩Dα,r ([0,T ];Z)×G 0
n=1
−
λt +
Z T N X n=1
≤
N X
0
N X
0
Z T |rn | ⟨x, gn ⟩Dα,r ([0,T ];Z)×G − 0
|rn |
n=1
≤
⟨xn,s , φ(s)⟩ e Z×E ds
! Z !! T ⟨xs , φ en (s)⟩Z×E ds ⟨xn,s , φ(s)⟩ e Z×E ds
n=1
≤
!
Z T
⟨xs , φ en (s)⟩Z×E ds
ε c (1 + ∥x∥α ) C1 Cη Cρ kN max(1, |rn |)
ε c (1 + ∥x∥α ) , C1 Cη Cρ k
c which implies that |(ai ◦ Rγ )(t, x)| ≤ |(ai ◦ Tγ )(t, x)| + 1 + ∥(t, x)∥R×Dα,r ([0,T ];Z) . Thus, by using a telescoping sum, it holds for every j = 1, ..., k and v1 , ..., vj ∈ R × Dα,r ([0, T ]; Z) that
j Y
(ai ◦ Tγ )(vℓ ) −
ℓ=1
≤
j Y
(ai ◦ Rγ )(vℓ )
ℓ=1
j X
l−1 Y
ℓ=1
m=1
|(ai ◦ Tγ )(vm )| |(ai ◦ Tγ )(vℓ ) − (ai ◦ Rγ )(vℓ )|
m=l+1 j
≤j
j Y
!
Y c ε |(ai ◦ Tγ )(vℓ )| + 1 + ∥vℓ ∥R×Dα,r ([0,T ];Z) kC1 Cη Cρ ℓ=1 j
Y c ε C1 1 + ∥vℓ ∥R×Dα,r ([0,T ];Z) C1 Cη Cρ ℓ=1 !ck j X ε = 1+ ∥vℓ ∥R×Dα,r ([0,T ];Z) . Cη Cρ ≤
ℓ=1
! |(ai ◦ Rγ )(vm )|
72
P. SCHMOCKER AND J. TEICHMANN
Therefore, by using the chain rule, it follows for every j = 1, ..., k and (u, v1 , ..., vj ) ∈ ϕi (Λα,r T,Z ) × (R × Dα,r ([0, T ]; Z))j that dj ρ ((ai ◦ Tγ )(·) + b) (u; v1 , ..., vj ) − dj ρ ((ai ◦ Rγ )(·) + b) (u; v1 , ..., vj ) ≤ ρ(|j|) ((ai ◦ Tγ )(u) + b)
j Y
(ai ◦ Tγ )(vℓ ) − ρ(|j|) ((ai ◦ Rγ )(u) + b)
ℓ=1
j Y
(ai ◦ Rγ )(vℓ )
ℓ=1
≤ ρ(|j|) ((ai ◦ Tγ )(u) + b) − ρ(|j|) ((ai ◦ Rγ )(u) + b)
j Y
|(ai ◦ Tγ )(vℓ )|
ℓ=1
+ ρ(|j|) ((ai ◦ Rγ )(u) + b)
j Y
(ai ◦ Tγ )(vℓ ) −
ℓ=1
j Y
(ai ◦ Rγ )(vℓ )
ℓ=1
≤ Cρ |(ai ◦ Tγ )(u) − (ai ◦ Rγ )(u)| C1 + Cρ
j Y
(ai ◦ Tγ )(vℓ ) −
ℓ=1
j Y
(ai ◦ Rγ )(vℓ )
ℓ=1 j
≤ C1 Cρ
≤
ε Cη
c c ε ε Y 1 + ∥u∥R×Dα,r ([0,T ];Z) + Cρ 1 + ∥vℓ ∥R×Dα,r ([0,T ];Z) C1 Cη Cρ k Cη Cρ ℓ=1 !cj j X 1 + max(j, 1)∥u∥R×Dα,r ([0,T ];Z) + ∥vℓ ∥R×Dα,r ([0,T ];Z) . ℓ=1
Hence, we conclude that ∥ρ ((ai ◦ Tγ )(·) + b) − ρ ((ai ◦ Rγ )(·) + b)∥Bk (ϕi (Λα,r )) Ψi
T ,Z
j
d ρ ((ai ◦ Tγ )(·) + b) (u; v1:j ) − dj ρ ((ai ◦ Rγ )(·) + b) (u; v1:j ) j=0,...,k (u,v1:j ) ψi,j (u, v1:j ) cj Pj α,r ([0,T ];Z) + α,r ([0,T ];Z) 1 + max(j, 1)∥u∥ ∥v ∥ ℓ R×D R×D ℓ=1 ε ≤ max sup Pj Cη j=1,...,k (u,v1:j ) η max(j, 1)∥u∥ α,r ([0,T ];Z) ∥v ∥ ℓ R×D α,r ([0,T ];Z) + R×D ℓ=1 = max
≤
sup
ε ε (1 + r)j max(1,c) = max sup Cη = ε, Cη j=0,...,k r∈(0,∞) η(r) Cη
α,r where the supremum is taken over (u, v1:j ) ∈ ϕi (Λα,r ([0, T ]; Z))j . Since ε ∈ (0, 1) T,Z ) × (R × D i ,ρ,R was chosen arbitrarily, this shows that ρ((ai ◦ Tγ )(·) + b) belongs to the closure of N N A ϕi (Λα,r ) T ,Z
ρ e,ρ with respect to ∥ · ∥BΨk (ϕi (Λα,r . Finally, we can apply Theorem 4.13 to obtain that PN Λ = α,r ,Y T ,Z )) T ,Z
i
k N N A,ρ,L is a dense subset of BΨ (Λα,r T,Z ; Y ). Λα,r ,Y T ,Z
□
F.3. Proof of Corollary 5.7. Proof of Corollary 5.7. Compared to Corollary 5.5, we now restrict ourselves to maps that have derivatives only in certain directions, i.e., the horizontal derivatives Dj f : Λα,r T,Z → Y , j = 1, ..., k, d and vertical derivatives Dβ f : Λα,r → Y , β ∈ N , which are represented on the model space 0,ℓ T,Z α,r ϕi (Λα,r ) ⊆ R × D ([0, T ]; Z) by the directional derivatives T,Z dj f ◦ ϕ−1 ((t, xt ); (1, 0), ..., (1, 0)) i
and d|β| f ◦ ϕ−1 ((t, xt ); (0, e1 1[t,T ] ), ..., (0, ed 1[t,T ] )), i
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
73
respectively, where the i-th unit vector ei ∈ Rd appears βi -times. Hence, in order to show that e,ρ α,r PN ρΛ is dense in Bk,ℓ α,r ψ (ΛT,Rd ; Y ), we apply Nachbin’s theorem not to the full jet space, but to T ,Z ,Y the restricted jet space generated by the horizontal and vertical directions. To this end, we only need that the additive family A has nowhere vanishing derivatives in the directions of interest, i.e., (1, 0) ∈ R × Dα,r ([0, T ]; Z) and (0, ei 1[t,T ] ) ∈ R × Dα,r ([0, T ]; Z), to obtain the conclusion. □ Acknowledgments. P. Schmocker gratefully acknowledges financial support by the FinsureTech Hub of ETH Zurich. J. Teichmann gratefully acknowledges financial support by ETH Foundation. References [1] B. Acciaio, A. Kratsios, and G. Pammer. Designing universal causal deep learning models: The geometric (hyper)transformer. Mathematical Finance, 34(2):671–735, 2024. [2] L. Ambrosio, N. Fusco, and D. Pallara. Functions of bounded variation and free discontinuity problems. Oxford science publications. Clarendon Press, Oxford, 2000. [3] R. M. Aron and J. B. Prolla. Polynomial approximation of differentiable functions on Banach spaces. Journal für die reine und angewandte Mathematik, 313:195–216, 1980. [4] A. R. Barron. Universal approximation bounds for superpositions of a sigmoidal function. IEEE Transactions on Information Theory, 39(3):930–945, 1993. [5] A. Bastiani. Applications différentiables et variétés différentiables de dimension infinie. Journal d’Analyse Mathématique, 13:1–114, 1964. [6] C. Bayer, P. P. Hager, S. Riedel, and J. Schoenmakers. Optimal stopping with signatures. The Annals of Applied Probability, 33(1):238–273, 2023. [7] C. Bayer, L. Pelizzari, and J. Schoenmakers. Primal and dual optimal stopping with signatures. Finance and Stochastics, 29:981–1014, 2025. [8] F. E. Benth, N. Detering, and L. Galimberti. Neural networks in Fréchet spaces. Annals of Mathematics and Artificial Intelligence, 91:75–103, 2023. [9] M. S. Berger. Nonlinearity and functional analysis: Lectures on nonlinear problems in mathematical analysis. Pure and applied mathematics, a series of monographs and textbooks; v. 74. Academic Press, New York, 1977. [10] S. Bernstein. Le problème de l’approximation des fonctions continues sur tout l’axe réel et l’une de ses applications. Bulletin de la Société Mathématique de France, 52:399–410, 1924. [11] K.-D. Bierstedt. Gewichtete Räume stetiger vektorwertiger Funktionen und das injektive Tensorprodukt. PhD thesis, Johannes-Gutenberg Universität Mainz, Mainz, 1971. [12] P. Billingsley. Convergence of probability measures. Wiley series in probability and statistics. Probability and statistics section. Wiley, New York, 2nd ed. edition, 1999. [13] J. Blessing, R. Denk, M. Kupper, and M. Nendel. Convex monotone semigroups and their generators with respect to Γ-convergence. Journal of Functional Analysis, 288(8):110841, 2025. [14] H. Boedihardjo, X. Geng, T. Lyons, and D. Yang. The signature of a rough path: Uniqueness. Advances in Mathematics, 293:720–737, 2016. [15] H. Bölcskei, P. Grohs, G. Kutyniok, and P. Petersen. Optimal approximation with sparsely connected deep neural networks. SIAM Journal on Mathematics of Data Science, 1:8–45, 2019. [16] H. Brézis. Functional Analysis, Sobolev Spaces and Partial Differential Equations. Universitext. Springer, New York, 2011. [17] E. J. Candès. Ridgelets: Theory and Applications. PhD thesis, Stanford University, 1998. https://candes. su.domains/publications/downloads/Thesis.pdf. [18] M. Ceylan and D. J. Prömel. Global universal approximation with Brownian signatures. Preprint arXiv:2512.16396, 2025. [19] K.-T. Chen. Integration of paths, geometric invariants and a generalized Baker-Hausdorff formula. Annals of Mathematics, 65(1):163–178, 1957. [20] T. Chen and H. Chen. Approximation capability to functions of several variables, nonlinear functionals, and operators by radial basis function neural networks. IEEE Transactions on Neural Networks, 6(4):904–910, 1995. [21] R. Cont. Functional Ito Calculus and functional Kolmogorov equations, pages 123–208. Advanced Courses in Mathematics. Birkhauser, Basel, 2016. Lecture Notes of the Barcelona Summer School in Stochastic Analysis, July 2012.
74
P. SCHMOCKER AND J. TEICHMANN
[22] R. Cont and D.-A. Fournié. Change of variable formulas for non-anticipative functionals on path space. Journal of Functional Analysis, 259(4):1043–1072, 2010. [23] R. Cont and D.-A. Fournié. Functional Itô calculus and stochastic integral representation of martingales. Annals of Probability, 41(1):109–133, 01 2013. [24] S. Cox, A. Khedher, and T. Maessen. Universal approximation by signatures for infinite-dimensional rough paths. Preprint arXiv:2603.03058, 2026. [25] C. Cuchiero, G. Gazzani, and S. Svaluto-Ferro. Signature-based models: Theory and calibration. SIAM Journal on Financial Mathematics, 14(3):910–957, 2023. [26] C. Cuchiero and J. Möller. Signature methods in stochastic portfolio theory. SIAM Journal on Financial Mathematics, 16(4):1239–1303, 2025. [27] C. Cuchiero, F. Primavera, and S. Svaluto-Ferro. Universal approximation theorems for continuous functions of càdlàg paths and Lévy-type signature models. Finance and Stochastics, 29:289–342, 2025. [28] C. Cuchiero, P. Schmocker, and J. Teichmann. Global universal approximation of functional input maps on weighted spaces. Constructive Approximation, 63:537–612, 2026. [29] C. Cuchiero and J. Teichmann. Generalized Feller processes and Markovian lifts of stochastic Volterra processes: the affine case. Journal of Evolution Equations, 20:1–48, 2020. [30] G. Cybenko. Approximation by superpositions of a sigmoidal function. Mathematics of Control, Signals and Systems, 2(4):303–314, 1989. [31] J. Dieudonné. Foundations of modern analysis. Enlarged and corrected printing. Academic Press, New York, London, 1969. [32] J. Dixmier. Sur un théorème de Banach. Duke Mathematical Journal, 15(4):1057–1071, 1948. [33] P. Dörsek and J. Teichmann. A semigroup point of view on splitting schemes for stochastic (partial) differential equations. Preprint arXiv:1011.2651, 2010. [34] B. Dupire. Functional Itô calculus. Technical report, Bloomberg, 2009. Bloomberg Portfolio Research Paper No. 2009-04-FRONTIERS. [35] S. N. Ethier and T. G. Kurtz. Markov processes: Characterization and convergence. John Wiley & Sons, 2005. [36] M. Fliess. Fonctionnelles causales non linéaires et indéterminées non commutatives. Bulletin de la Société Mathématique de France, 109:3–40, 1981. [37] G. B. Folland. Fourier analysis and its applications. Brooks/Cole Publishing Company, Belmont, California, 1st edition, 1992. [38] H. Föllmer. Calcul d’Ito sans probabilités. Séminaire de probabilités de Strasbourg, 15:143–150, 1981. [39] P. K. Friz and M. Hairer. A Course on Rough Paths: With an Introduction to Regularity Structures. Universitext. Springer International Publishing, Cham, 2nd edition, 2020. [40] P. K. Friz and N. B. Victoir. Multidimensional Stochastic Processes as Rough Paths: Theory and Applications. Cambridge Studies in Advanced Mathematics. Cambridge University Press, 2010. [41] L. Galimberti, A. Kratsios, and G. Livieri. Designing universal causal deep learning models: The case of infinite-dimensional dynamical systems from stochastic analysis. forthcoming in Constructive Approximation, 2026. [42] P. Gierjatowicz, M. Sabate-Vidales, D. Šiška, L. Szpruch, and Ž. Žurič. Robust pricing and hedging via neural SDEs. Journal of Computational Finance, 26:1–32, 2020. [43] H. Glöckner. Infinite-dimensional Lie groups without completeness restrictions. Banach Center Publications, 55:43–59, 2002. [44] H. Glöckner. Fundamentals of submersions and immersions between infinite-dimensional manifolds. Preprint arXiv:1502.05795, 2015. [45] J. Gómez Gil and J. G. Llavona. Polynomial approximation of weakly differentiable functions on Banach spaces. Proceedings of the Royal Irish Academy. Section A: Mathematical and Physical Sciences, 82A(2):141– 150, 1982. [46] L. Gonon, L. Grigoryeva, and J.-P. Ortega. Infinite-dimensional reservoir computing. Neural Networks, 179:106486, 2024. [47] I. Goodfellow, Y. Bengio, and A. Courville. Deep Learning. MIT Press, 2016. [48] L. Grigoryeva and J.-P. Ortega. Echo state networks are universal. Neural Networks, 108:495–508, 2018. [49] B. Hambly and T. Lyons. Uniqueness for the signature of a path of bounded variation and the reduced path group. Annals of Mathematics, 171(1):109–167, 2010. [50] F. Hausdorff. Die symbolische Exponentialformel in der Gruppentheorie. Ber. Verh. Kgl. Sächs. Ges. Wiss., 58:19–48, 1906.
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
75
[51] G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-R. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, and B. Kingsbury. Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups. IEEE Signal Processing Magazine, 29(6):82–97, 2012. [52] K. Hornik. Approximation capabilities of multilayer feedforward networks. Neural Networks, 4(2):251–257, 1991. [53] K. Hornik, M. Stinchcombe, and H. White. Universal approximation of an unknown mapping and its derivatives using multilayer feedforward networks. Neural Networks, 3(5):551–560, 1990. [54] T. Hytönen, J. van Neerven, M. Veraar, and L. Weis. Analysis in Banach Spaces, volume 63 of Ergebnisse der Mathematik und ihrer Grenzgebiete. 3. Folge. Springer, Cham, 2016. [55] A. Iserles and S. P. Nørsett. On the solution of linear differential equations in Lie groups. Philos. Trans. Roy. Soc. A, 357:983–1019, 1999. [56] A. Jakubowski. The Skorokhod space in functional convergence: a short introduction. In International conference: Skorokhod Space, volume 50, pages 11–18, 2007. [57] S. Kaijser. A note on dual Banach spaces. Mathematica Scandinavica, 41(2):325–330, 1977. [58] H. Keller. Differential calculus in locally convex spaces. Lecture Notes in Mathematics; 417. Springer-Verlag, Berlin, Germany, 1st edition, 1974. [59] P. Kidger, J. Foster, X. Li, and T. J. Lyons. Neural SDEs as infinite-dimensional GANs. In International Conference on Machine Learning, pages 5453–5463. PMLR, 2021. [60] D. P. Kingma and J. Ba. Adam: A method for stochastic optimization. In Y. Bengio and Y. LeCun, editors, 3rd International Conference on Learning Representations, 2015, Conference Track Proceedings, May 2015. [61] F. J. Kiraly and H. Oberhauser. Kernels for sequentially ordered data. Journal of Machine Learning Research, 20(31):1–45, 2019. [62] J. Korevaar. Distribution proof of Wiener’s Tauberian theorem. Proceedings of the American Mathematical Society, 16(3):353–355, 1965. [63] N. Kovachki, Z. Li, B. Liu, K. Azizzadenesheli, K. Bhattacharya, A. Stuart, and A. Anandkumar. Neural operator: Learning maps between function spaces with applications to PDEs. Journal of Machine Learning Research, 24(89):1–97, 2023. [64] A. Kratsios and I. Bilokopytov. Non-Euclidean universal approximation. Advances in Neural Information Processing Systems, 33:10635–10646, 2020. [65] A. Kratsios, C. Liu, M. Lassas, M. V. de Hoop, and I. Dokmanić. An approximation theory for metric space-valued functions with a view towards deep learning. Preprint arXiv:2304.12231, 2023. [66] A. Kratsios, A. Neufeld, and P. Schmocker. Generative neural operators of log-complexity can simultaneously solve infinitely many convex programs. Preprint arXiv:2508.14995, 2025. [67] A. Kriegl and P. W. Michor. The convenient setting of global analysis, volume 53 of Mathematical surveys and monographs. American Mathematical Society, Providence, R.I, 1997. [68] A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In F. Pereira, C. Burges, L. Bottou, and K. Weinberger, editors, Advances in Neural Information Processing Systems, volume 25. Curran Associates, Inc., 2012. [69] G. Lancien and E. Pernecká. Approximation properties and Schauder decompositions in Lipschitz-free spaces. Journal of Functional Analysis, 264(10):2323–2334, 2013. [70] S. Lanthaler, S. Mishra, and G. E. Karniadakis. Error estimates for DeepONets: a deep learning framework in infinite dimensions. Transactions of Mathematics and Its Applications, 6(1), 2022. [71] M. Leshno, Vladimir Ya. Lin, A. Pinkus, and S. Schocken. Multilayer feedforward networks with a nonpolynomial activation function can approximate any function. Neural Networks, 6(6):861–867, 1993. [72] D. Levin, T. Lyons, and H. Ni. Learning from the past, predicting the statistics for the future, learning an evolving system. Preprint arXiv:1309.0260, 2013. [73] Z. Li, N. Kovachki, K. Azizzadenesheli, B. Liu, K. Bhattacharya, A. Stuart, and A. Anandkumar. Fourier neural operator for parametric partial differential equations. Preprint arXiv:2010.08895, 2020. [74] J. Lindenstrauss and L. Tzafriri. Classical Banach spaces I and II. Classics in mathematics. Springer, Berlin, 1996. [75] J. G. Llavona. Approximation of Continuously Differentiable Functions, volume 130 of North-Holland Mathematics Studies. North-Holland, 1986. [76] L. Lu, P. Jin, G. Pang, Z. Zhang, and G. E. Karniadakis. Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators. Nature Machine Intelligence, 3:218–229, 2021. [77] T. Lyons, S. Nejad, and I. Perez Arribas. Non-parametric pricing and hedging of exotic derivatives. Applied Mathematical Finance, 27(6):457–494, 2020.
76
P. SCHMOCKER AND J. TEICHMANN
[78] T. J. Lyons, M. Caruana, and T. Lévy. Differential equations driven by rough paths. Lecture notes in mathematics 1908. Springer, Berlin, 2007. [79] W. Magnus. On the exponential solution of differential equations for a linear operator. Communications on Pure and Applied Mathematics, 7(4):649–673, 1954. [80] W. S. McCulloch and W. Pitts. A logical calculus of the ideas immanent in nervous activity. The bulletin of mathematical biophysics, 5:115–133, 1943. [81] H. N. Mhaskar and N. Hahm. Neural networks for functional approximation and system identification. Neural Computation, 9:143–159, 1997. [82] R. Miron. The geometry of higher-order Lagrange spaces: Applications to mechanics and physics. Fundamental theories of physics; 82. Springer Science+Business Media, B.V., Dordrecht, 1st edition, 1997. [83] T. M. Mitchell. Machine Learning. McGraw-Hill series in computer science. WCB McGraw-Hill, Boston MA, 1997. [84] G. Montavon, G. Orr, and K.-R. Müller. Neural Networks: Tricks of the Trade. Theoretical Computer Science and General Issues; 7700. Springer, Berlin, Heidelberg, 2nd edition, 2012. [85] J. R. Munkres. Topology. Pearson, Harlow, Essex, UK, 2nd, Pearson new international edition, 2014. [86] L. Nachbin. Sur les algèbras denses de fonctions différentiables sur une variété. Comptes rendus de l’Académie des Sciences de Paris, 228:1549–1551, 1949. [87] L. Nachbin. Weighted approximation for algebras and modules of continuous functions: Real and self-adjoint complex cases. Annals of Mathematics, 81(2):289–302, 1965. [88] L. Nachbin. On the closure of modules of continuously differentiable mappings. Rendiconti del Seminario Matematico della Università di Padova, 60:33–42, 1978. [89] L. Nachbin. On the weighted approximation of continuously differentiable functions. Proceedings of the American Mathematical Society, 111(2):481–485, 1991. [90] A. Neufeld and P. Schmocker. Universal approximation results for neural networks with non-polynomial activation function over non-compact domains. Analysis and Applications, 24(05):1123–1173, 2026. [91] K.-F. Ng. On a theorem of Dixmier. Mathematica Scandinavica, 29:279–280, 1971. [92] A. Pinkus. Approximation theory of the MLP model in neural networks. Acta Numerica, 8:143–195, 1999. [93] J. B. Prolla. Weighted spaces of vector-valued continuous functions. Annali di Matematica Pura ed Applicata, 89:145–157, 1971. [94] J. B. Prolla. Approximation of Vector Valued Functions. North-Holland Mathematics Studies 25. NorthHolland, Amsterdam, 1977. [95] J. B. Prolla and C. S. Guerreiro. An extension of Nachbin’s theorem to differentiable functions on Banach spaces with the approximation property. Arkiv för Matematik, 14(1-2):251–258, 1976. [96] M. Raissi, P. Perdikaris, and G. Karniadakis. Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations. Journal of Computational Physics, 378:686–707, 2019. [97] M. Röckner and Z. Sobol. Kolmogorov equations in infinite dimensions: Well-posedness and regularity of solutions, with applications to stochastic generalized Burgers equations. Annals of Probability, 34(2):663– 727, 03 2006. [98] F. Rossi and B. Conan-Guez. Functional multi-layer perceptron: a non-linear tool for functional data analysis. Neural Networks, 18(1):45–60, 2005. [99] R. A. Ryan. Introduction to tensor products of Banach spaces. Springer monographs in mathematics. Springer, London, 2002. [100] C. Salvi, M. Lemercier, and A. Gerasimovics. Neural stochastic PDEs: Resolution-invariant learning of continuous spatiotemporal dynamics. In Advances in Neural Information Processing Systems, 2022. [101] C. R. Samuel N. Cohen and S. Wang. Arbitrage-free neural-SDE market models. Applied Mathematical Finance, 30(1):1–46, 2023. [102] H. H. Schaefer and M. P. Wolff. Topological vector spaces, volume 3 of Graduate texts in mathematics. Springer, New York, 2nd edition, 1999. [103] A. Schmeding. Introduction to Infinite-Dimensional Differential Geometry. Cambridge studies in advanced mathematics. Cambridge University Press, Cambridge, United Kingdom, 1st edition, 2023. [104] F. Schur. Zur Theorie der endlichen Transformationsgruppen. Mathematische Annalen, 38:263–286, 1891. [105] D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis. Mastering the game of Go with deep neural networks and tree search. Nature, 529:484–489, 2016.
WEIGHTED UNIVERSAL APPROXIMATION OF DIFFERENTIABLE MAPS
77
[106] J. Sirignano and K. Spiliopoulos. DGM: A deep learning algorithm for solving partial differential equations. Journal of Computational Physics, 375:1339–1364, 2018. [107] L. Song, Y. Liu, J. Fan, and D.-X. Zhou. Approximation of smooth functionals using deep ReLU networks. Neural Networks, 166:424–436, 2023. [108] M. B. Stinchcombe. Neural network approximation of continuous functionals and continuous functions on compactifications. Neural Networks, 12(3):467–477, 1999. [109] R. S. Strichartz. The Campbell-Baker-Hausdorff-Dynkin formula and solutions of differential equations. Journal of Functional Analysis, 72(2):320–345, 1987. [110] W. H. Summers. Weighted Locally Convex Spaces of Continuous Functions. PhD thesis, Louisiana State University, 1968. https://repository.lsu.edu/gradschool_disstheses/1520. [111] A. Suri. Higher order tangent bundles. Mediterranean Journal of Mathematics, 14(5), 2016. [112] H. Triebel. Theory of function spaces II. Modern Birkhäuser Classics. Birkhäuser Verlag, Basel, Switzerland, 1992. [113] H. Triebel. Theory of Function Spaces III. Monographs in Mathematics. Birkhäuser, Basel, Boston, Berlin, 2006. [114] H. Triebel. Theory of function spaces. Modern Birkhäuser Classics. Birkhäuser Verlag, Basel, Switzerland, 2010. [115] A. M. Turing. Computing machinery and intelligence. Mind, LIX(236):433–460, 10 1950. [116] B. Walter. Weighted diffeomorphism groups of Banach spaces and weighted mapping groups. Dissertationes Mathematicae, 484:1–126, 2012. [117] N. Weaver. Lipschitz Algebras. World Scientific, Singapore, 2nd edition, 1999. [118] N. Wiener. Tauberian theorems. Annals of Mathematics, 33(1):1–100, 1932. [119] S. Willard. General topology. Addison-Wesley Publishing Company, Mineola, NY, 2004. [120] P. Wojtaszczyk. Banach spaces for analysts, volume 25 of Cambridge Studies in Advanced Mathematics. Cambridge University Press, Cambridge, 1991. [121] G. Zapata. Bernstein approximation problem for differentiable functions and quasi-analytic weights. Transactions of the American Mathematical Society, 182:503–509, 1973. Department of Mathematics, ETH Zurich, Switzerland Email address: [email protected] Department of Mathematics, ETH Zurich, Switzerland Email address: [email protected]