ON HIGHER-ORDER GEOMETRIC REFINEMENTS OF CLASSICAL COVARIANCE ASYMPTOTICS: AN APPROACH VIA INTRINSIC AND EXTRINSIC INFORMATION GEOMETRY
arXiv:2604.12725v1 [math.ST] 14 Apr 2026
MALIK AMIR AND SOURANGSHU GHOSH
Abstract. Classical Fisher-information asymptotics describe the covariance of regular efficient estimators through the local quadratic approximation of the log-likelihood, and therefore reflect only the first-order geometry of the statistical model. In curved settings, such as mixtures, curved exponential families, latent-variable models, and manifold-constrained parameter spaces, finite-sample behavior can deviate systematically from these first-order predictions. We develop a coordinateinvariant, curvature-aware refinement of classical covariance asymptotics by viewing a regular parametric family as a Riemannian manifold (Θ, g) equipped with the Fisher–Rao metric and immersed into an ambient Hilbert space L2 (µ) through the square-root density map. Using higher-order likelihood expansions together with the intrinsic and extrinsic geometry of this immersion, we derive, under suitable regularity, moment, and replacement assumptions, an n−2 correction to the leading n−1 I(θ)−1 term in the covariance expansion of score-root, first-order efficient estimators. This correction is governed by a second-order tensor Pij admitting a canonical decomposition into an intrinsic Ricci-type contraction of the Fisher–Rao Riemann curvature tensor, an extrinsic Gram-type contraction of the second fundamental form, and a Hellinger discrepancy tensor capturing higherorder probabilistic content not determined solely by the immersion geometry. The extrinsic term is positive semidefinite by construction, the full correction is invariant under smooth reparameterization, and the correction vanishes identically in full exponential families. In one dimension the intrinsic curvature vanishes, and the refinement reduces to a purely extrinsic contribution closely related to Efron’s statistical curvature. We then examine how this perspective extends beyond the regular locus to singular models, where the Fisher information degenerates and quadratic local geometry breaks down. Using resolution of singularities under an additive normal crossing assumption, we describe the induced resolved metric, the role of the real log canonical threshold in determining learning rates, the corresponding posterior mean-squared-error scaling, and a curvature-based covariance expansion on the resolved space, recovering the regular theory as a special case. Beyond the asymptotic expansion itself, the framework suggests geometric diagnostics of weak identifiability and points toward curvature-aware principles for regularization, diagnostics, and optimization in modern probabilistic learning systems.
Contents 1. Introduction 1.1. Literature review and motivation 1.2. Roadmap of the paper 2. Statistical Manifold and Fisher–Rao Geometry 2.1. The square-root density map 2.2. Tangent space and orthogonality 2.3. The induced Fisher–Rao metric 2.4. Geometric significance
2 4 5 6 6 7 7 7
2020 Mathematics Subject Classification. Primary 62F10; Secondary 14P25, 53B20, 53C21, 62B10. Key words and phrases. information geometry, higher-order asymptotics, statistical manifolds, Fisher–Rao metric, statistical curvature, asymptotic inference, point estimation, singular statistical models, algebraic geometry. All authors contributed equally to this work. 1
2
MALIK AMIR AND SOURANGSHU GHOSH
3. Curvature of the Statistical Manifold 3.1. Intrinsic curvature 3.2. Extrinsic curvature through the Hellinger immersion 3.3. The Gauss equation and the relation between intrinsic and extrinsic curvature 3.4. A scalar measure of extrinsic bending 3.5. Relation with Efron’s statistical curvature 3.6. Motivation for curvature-corrected covariance asymptotics 4. Second-Order Geometry of the Log-Likelihood 4.1. Score, Hessian, and Fisher information 4.2. Taylor expansion and departure from quadraticity 4.3. The covariant Hessian 4.4. Third-order likelihood geometry 5. A Curvature-Corrected Second-Order Covariance Expansion 5.1. Asymptotic setting and second-order expansion 5.2. Structure of the second-order correction 5.3. Main theorem 5.4. Interpretation of the correction 5.5. Relation to classical higher-order asymptotic refinements 6. Extension to Singular Models 6.1. Geometric and analytic structure of regular and singular statistical models 6.2. Resolution of singularities and monomialization of the KL divergence 6.3. Differential geometric structure on the resolved manifold 6.4. The Real Log Canonical Threshold (RLCT) 6.5. Asymptotic posterior mean squared error 6.6. Tangent cone geometry and construction of the effective information metric 6.7. Pushforward of curvature tensors under resolution of singularities 6.8. Recovery of the classical theory 7. Conclusion and Outlook 7.1. Applications to deep learning training and regularization 8. Appendix 8.1. Why is an immersion sufficient? 8.2. Detailed standing assumptions for the proof of the main theorem 8.3. Proof of the main theorem References
8 8 8 9 9 10 10 10 10 11 11 11 12 12 13 13 16 16 16 16 18 19 20 22 22 23 23 24 25 26 26 27 28 31
1. Introduction Let X1 , . . . , Xn be independent and identically distributed random variables drawn from a parametric family of probability distributions {pθ : θ ∈ Θ ⊂ Rd }, where Θ is an open subset and the mapping θ 7→ pθ is assumed to be sufficiently smooth1. An estimator of the parameter θ is a measurable function θ̂ = θ̂(X1 , . . . , Xn ) taking values in Θ. Under standard regularity conditions, the score vector for a single observation is given by U (θ; X) = ∇θ log pθ (X), has zero mean and finite second moments, and allows us to define the Fisher information matrix by h i h i I(θ) = Eθ U (θ; X)U (θ; X)⊤ = Eθ ∇θ log pθ (X) ∇θ log pθ (X)⊤ . 1This will be made more precise in Section 2.
ON HIGHER-ORDER GEOMETRIC REFINEMENTS OF CLASSICAL COVARIANCE ASYMPTOTICS
3
The classical first-order asymptotic picture asserts that, in regular models, many efficient estimation procedures satisfy 1 1 −1 Covθ (θ̂n ) = I(θ) + o , n n or equivalently √ d n (θ̂n − θ) − → N 0, I(θ)−1 . This first-order description is sharp in regular exponential families and plays a foundational role in asymptotic statistics, signal processing, and information theory. However, it is fundamentally local in nature. Its derivation relies on a quadratic approximation of the log-likelihood around the true parameter value and therefore depends only on the second-order structure encoded by the Fisher information. Geometrically, this corresponds to endowing the parameter space Θ with the Fisher– Rao metric and approximating the statistical model by its tangent space at θ. All higher-order features of the model, such as curvature, global nonlinearities, and embedding effects, are absent from this approximation. Many statistical models of contemporary interest are intrinsically or extrinsically curved. Finite mixtures, latent-variable models, curved exponential families, and models with parameters constrained to manifolds such as spheres, rotation groups, or spaces of positive-definite matrices exhibit nonzero statistical curvature. In such models, the statistical manifold cannot be globally or even locally flattened beyond first order by any smooth 2 reparameterization, since the Riemann curvature tensor is invariant under smooth reparameterization. As a consequence, the Fisher information alone does not fully characterize the local distinguishability of nearby distributions, especially when the sample size is small to moderate rather than large. Empirically, this manifests as systematic second-order deviations from Fisher-information-only predictions, even for estimators that are first-order asymptotically efficient3. The purpose of this article is to develop a curvature-aware refinement of the classical first-order covariance asymptotics that remains coordinate invariant while incorporating higher-order geometric information. By viewing the parametric family as a Riemannian manifold equipped with the Fisher–Rao metric and immersed into a Hilbert space through the Hellinger map, we identify two geometrically distinct notions of curvature that contribute to the second-order behavior of estimator covariance. The first is intrinsic curvature, captured by the Riemann curvature tensor of the Fisher–Rao metric, which quantifies the failure of the Fisher–Rao geometry to be locally Euclidean beyond first order and controls geodesic deviation4. The second is extrinsic curvature, captured by the second fundamental form of the square-root density immersion, which measures how the model, viewed as a submanifold of square-root densities, deviates from its tangent space within the unit sphere of the ambient Hilbert space. Incorporating these intrinsic and extrinsic curvature effects leads to a second-order correction of order n−2 in the covariance expansion. This correction quantifies how curvature modifies the effective second-order behavior of estimation error beyond the classical Fisher term. Efron’s statistical curvature, in one dimension, quantifies the extent to which a model deviates from an exponential-family representation and measures the resulting second-order information loss. In the nonlinear regression literature, Bates and Watts distinguish between intrinsic curvature 2It is sufficient to assume C ≥2 .
Let {pθ : θ ∈ Θ ⊂ Rd } be a regular parametric model with Fisher information matrix I(θ). An estimator sequence (θ̂n )n≥1 is called (first-order) asymptotically efficient at θ if it is regular at θ (in the usual Hájék–Le Cam sense) and satisfies √ d n (θ̂n − θ) − → N 0, I(θ)−1 . 3
Equivalently, θ̂n attains the Fisher-information asymptotics to first order, in the sense that its asymptotic covariance matrix equals I(θ)−1 . 4 For a one-parameter family of geodesics γu (s) with central geodesic γ(s) = γ0 (s), the variation field J(s) = ∂u γu (s)|u=0 satisfies the Jacobi equation ∇γ̇ ∇γ̇ J + R(J, γ̇)γ̇ = 0.
4
MALIK AMIR AND SOURANGSHU GHOSH
(a reparameterization-invariant property of the model manifold) and parameter-effects curvature (apparent curvature induced by a particular parametrization and removable by reparameterization). Our decomposition is compatible with this philosophy, but differs in that the intrinsic term in our expansion is governed by the Fisher–Rao Riemann curvature tensor and is therefore purely reparameterization invariant, matching the role of Bates–Watts intrinsic curvature. The extrinsic term in our expansion arises from the second fundamental form of the Hellinger immersion. It is likewise invariant under reparameterization and therefore should not be conflated with parametereffects curvature. Rather, it quantifies the bending of the statistical model in the ambient Hilbert space, yielding the Gram-type positive semidefinite contraction that appears in the n−2 term. 1.1. Literature review and motivation. The classical Cramér–Rao lower bound (CRLB) is the canonical local information inequality in regular parametric inference. Under standard assumptions, it bounds the covariance of (locally) unbiased estimators by the inverse of the Fisher information, translating the local quadratic approximation of the log-likelihood into a first-order variance benchmark [11, 15, 17, 19]. A central refinement, essential for geometric generalizations, is that Fisher information is more naturally interpreted as a coordinate-free Riemannian metric tensor (the Fisher– Rao metric) on the statistical manifold, rather than as a matrix tied to a particular parametrization. √ Under the square-root embedding ψθ = pθ , this metric arises as the pullback of the ambient L2 inner product, yielding an intrinsic geometric language for efficiency [1, 2, 4, 3]. The canonicity of this metric is further supported by invariance characterizations under sufficient statistics, called Markov morphisms, originating in Chentsov theory and extended in subsequent work. This viewpoint explains why CRLB-type statements are best regarded as geometric statements rather than coordinate artifacts [9, 10]. Finally, for general (possibly biased) procedures, the appropriate first-order covariance analysis involves the Jacobian of the estimator mean map m(θ) = Eθ [θ̂], highlighting the role of first-order bias structure in refined asymptotic theory [8, 20]. A separate, but tightly connected tradition, asks what remains beyond first order. Higher-order score expansions and refined local inequalities (e.g. Bhattacharyya-type inequalities) incorporate higher derivatives of the log-likelihood and demonstrate that first-order optimality does not fully control finite-sample performance [7, 17]. In likelihood asymptotics, systematic higher-order expansions quantify second-order risk and reveal how finite-sample efficiency losses depend on model curvature and higher-order structure [5, 14]. The geometric content of these second-order effects was crystallized by Efron, who introduced statistical curvature to measure deviation from exponential-family “straightness” and showed that curvature governs second-order efficiency loss in one-dimensional curved families [12]. Related second-order efficiency analyses for curved exponential families further clarified how curvature-type invariants enter second-order risk in [13, 14]. Independently, the nonlinear regression literature developed a decomposition into intrinsic curvature versus parameter-effects curvature (the latter being removable by reparametrization), providing a conceptual template for separating genuinely geometric obstructions from coordinate-induced artifacts [6]. These lines of work motivate the modern information-geometric stance. To understand efficiency beyond Fisher information, one must account for intrinsic curvature and, under the Hellinger immersion perspective, extrinsic bending encoded by second fundamental form-type objects [2, 3, 16]. Such curvature sensitivity is not only a theoretical nuance. In curved settings, including mixtures, latent-variable models, and manifold-constrained parameter spaces, finite-sample behavior can deviate systematically from Fisher-information-only predictions, even when the model is locally regular, because the local geometry is already highly nontrivial [14, 18]. Beyond this regular regime, many practically important models are singular, requiring different asymptotics and tools (e.g. algebraic-geometric singular learning theory), and thereby highlighting where classical first-order Fisher-information asymptotics can fail outright [21]. These observations motivate a coordinate-invariant refinement of the classical first-order covariance expansion that is sensitive to curvature while remaining interpretable on the regular locus. The goal, for sample size n, is to
ON HIGHER-ORDER GEOMETRIC REFINEMENTS OF CLASSICAL COVARIANCE ASYMPTOTICS
5
identify the n−2 -term in the covariance expansion and to decompose its geometric part into an intrinsic contribution governed by contractions of the Fisher–Rao Riemann curvature tensor and an extrinsic contribution governed by Gram-type contractions of the second fundamental form of the Hellinger immersion, with the extrinsic contribution manifestly positive semidefinite. Extending such an expansion beyond the score-root and first-order efficient setting, and understanding how curvature should then interact with derivatives of the estimator mean map, is a natural direction for future work and would connect curvature-aware second-order asymptotics to the broader theory of efficiency and its extensions toward semiparametric tangent-space formalisms [8, 14, 20].
1.2. Roadmap of the paper. After the introduction, Section 2 establishes the geometric framework in the regular setting. We realize the parametric family {pθ : θ ∈ Θ} through the square-root density, or Hellinger, immersion into L2 (µ), identify the tangent spaces of the resulting statistical manifold, and show that the pullback of the ambient Hilbert inner product is exactly the Fisher– Rao metric. This identifies (Θ, g) as the intrinsic statistical manifold governing first-order local distinguishability. Section 3 develops the curvature structures needed for higher-order analysis. On the intrinsic side, we introduce the Levi–Civita connection and the Riemann curvature tensor of the Fisher–Rao metric, emphasizing that nonvanishing curvature is a coordinate-invariant obstruction to flattening the information geometry beyond first order. On the extrinsic side, we study the Hellinger immersion, define its second fundamental form, derive the Gauss equation relating intrinsic curvature to extrinsic bending, and discuss scalar measures of extrinsic curvature together with their relation to Efron’s statistical curvature in one dimension. Section 4 turns to the higher-order differential structure of the log-likelihood. We introduce the score, Hessian, and third-order score tensors, explain why the ordinary observed Hessian is not tensorial under reparameterization, and replace it by the covariant Hessian. We then derive the third-order likelihood expansion and identify the cubic score moments and exponential-connection coefficients that carry information beyond Fisher’s quadratic approximation. Section 5 contains the main second-order asymptotic result. Under the regularity, moment, and stochastic-expansion assumptions stated later in the paper, we derive a coordinate-invariant covariance expansion with explicit n−2 correction for first-order efficient estimators. The correction is organized through the tensor P , which decomposes canonically into an intrinsic Ricci-type term, an extrinsic Gram-type term coming from the second fundamental form, and a Hellinger discrepancy term capturing higher-order probabilistic information not determined solely by the immersion geometry. We also discuss the structural consequences of this decomposition, including the vanishing of the correction in full exponential families and the simplifications that occur in one dimension. Section 6 extends the discussion from regular models to singular ones. There we explain why the regular manifold picture breaks down when the Fisher information degenerates or the immersion loses rank, and we examine how the geometric covariance analysis must be modified in the presence of non-identifiability, singularities, and related failures of classical first-order Fisher-information asymptotics. Section 7 concludes the paper and places the main results in a broader perspective. It summarizes the curvature-based refinement of the covariance expansion, emphasizes the canonical decomposition P = 12 R♯ + S ♯ + D, and explains how these second-order effects clarify weak identifiability and the failure of naive quadratic approximations. Within this final section, Section 7.1 discusses applications to modern learning systems, including curvature-aware regularization, diagnostic uses of the tensors R♯ , S ♯ , and P , possible refinements of natural-gradient and related second-order optimization methods, and the computational issues involved in approximating the relevant contractions in high dimension.
6
MALIK AMIR AND SOURANGSHU GHOSH
Finally, Section 8 contains the technical material omitted from the main text. It begins by explaining why a C 3 immersion of the square-root density map is sufficient for the local differentialgeometric analysis, without requiring a global embedding, and then develops the auxiliary lemmas, stochastic expansions, moment computations, and algebraic reductions used in the proof of the main theorem. 2. Statistical Manifold and Fisher–Rao Geometry This section sets up the geometric framework underlying the paper. We realize the parametric family through the square-root density map as a smooth immersed manifold in L2 (µ), identify its tangent space, and show that the pullback of the ambient Hilbert inner product is exactly the Fisher–Rao metric. This yields a coordinate-invariant geometric interpretation of local statistical distinguishability and provides the basic objects from which curvature corrections will later be constructed. 2.1. The square-root density map. Let (X , A, µ) be a measurable space with a σ-finite dominating measure µ. Let Θ ⊂ Rd be an open set, and let {pθ : θ ∈ Θ} be a parametric family of probability densities with respect to µ. Assume that pθ (x) > 0 for all θ ∈ Θ and for µ-almost every x, and that the map θ 7→ pθ (x), is C 3 for µ-almost every x. Assume further that for every multi-index α with |α| ≤ 3, (1) the partial derivatives ∂θα pθ exist for µ-almost every x, √ (2) the partial derivatives ∂θα pθ exist for µ-almost every x, √ (3) the functions ∂θα pθ belong to L2 (µ), √ (4) the map θ 7→ ∂θα pθ is continuous in the L2 (µ)-norm, (5) these differentiability and continuity assumptions are compatible with differentiation under the integral sign in the arguments below. Define the square-root density map, also called the Hellinger map, by √ Ψ : Θ → L2 (X , µ), Ψ(θ) = ψθ = pθ . Since ∥ψθ ∥2L2 =
Z pθ (x) dµ(x) = 1,
the image of Ψ lies in the unit sphere S = {f ∈ L2 (µ) : ∥f ∥L2 = 1}. To obtain a d-dimensional immersed statistical manifold, we assume that the Fisher information matrix exists, is finite, and is positive definite, Z ⊤ 0≺ ∇θ log pθ (x) ∇θ log pθ (x) pθ (x) dµ(x) for all θ ∈ Θ. Equivalently, the partial derivatives {∂i ψθ }di=1 are linearly independent in L2 (µ). In that case, we restrict attention to locally identifiable parametrizations. If a model is overparameterized, one should first reduce to an identifiable parametrization or treat the singular case separately. Thus the statistical model is realized through a finite-dimensional C 3 immersion Ψ : Θ → S ⊂ L2 (µ), which endows the image with an immersed submanifold structure. This map is called the Hellinger map because the L2 -distance between square-root densities is proportional to the Hellinger distance between the corresponding probability distributions.
ON HIGHER-ORDER GEOMETRIC REFINEMENTS OF CLASSICAL COVARIANCE ASYMPTOTICS
7
2.2. Tangent space and orthogonality. For any fixed θ ∈ Θ, the tangent space is spanned by the partial derivatives of ψθ with respect to the coordinates θi . Using the chain rule, one obtains ∂i ψθ (x) =
∂ p 1 ∂i pθ (x) 1p pθ (x) = p = pθ (x) ∂i log pθ (x). ∂θi 2 pθ (x) 2
Since L2 (µ) is equipped with its canonical inner product Z ⟨f, g⟩L2 = f (x)g(x) dµ(x), every tangent vector to S, and hence every tangent vector to the immersed statistical manifold, is orthogonal to the radial direction ψθ . Indeed, Z Z 1 1 ⟨ψθ , ∂i ψθ ⟩L2 = ∂i pθ (x) dµ(x) = ∂i pθ (x) dµ(x) = 0. 2 2 Therefore Tψθ M ⊂ {v ∈ L2 (µ) : ⟨v, ψθ ⟩L2 = 0}. The inclusion is strict because the orthogonal complement of ψθ in L2 (µ) is infinite-dimensional, whereas Tψθ M has dimension d. 2.3. The induced Fisher–Rao metric. The ambient inner product on L2 (µ) induces a Riemannian metric on the parameter manifold by pullback through the immersion Ψ. For θ ∈ Θ and coordinate directions i, j, define gij (θ) = 4⟨∂i ψθ , ∂j ψθ ⟩L2 . Substituting the expression for ∂i ψθ yields Z gij (θ) = ∂i log pθ (x) ∂j log pθ (x) pθ (x) dµ(x) = Eθ [∂i log pθ (X) ∂j log pθ (X)] . Thus gij (θ) coincides exactly with the entries of the Fisher information matrix Iij (θ). Consequently, the parameter space Θ is naturally endowed with a Riemannian metric g given by the Fisher information, and the pair (Θ, g) becomes the Fisher–Rao statistical manifold. This metric is intrinsic in the sense that it is invariant under smooth reparameterizations of θ, and it quantifies the local distinguishability of nearby probability distributions. For a smooth curve θ(t) in parameter space, the squared norm of its velocity at t = 0 is # " 2 d log pθ(t) (X) , ∥θ̇(0)∥2g = Eθ dt t=0
which measures the sensitivity of the log-likelihood to infinitesimal perturbations of the parameter. 2.4. Geometric significance. The classical Cramér–Rao bound depends only on this first-order geometric structure encoded by the metric g. However, the immersion M = Ψ(Θ) ⊂ S ⊂ L2 (µ) is generally curved. In particular, second derivatives ∂i ∂j ψθ need not lie in the tangent space Tψθ M. This failure of local flatness leads to intrinsic and extrinsic curvature effects, which enter higherorder expansions of the likelihood and ultimately produce curvature-dependent corrections to the classical Cramér–Rao lower bound.
8
MALIK AMIR AND SOURANGSHU GHOSH
3. Curvature of the Statistical Manifold This section introduces the curvature structure of the Fisher–Rao statistical manifold. We describe its intrinsic geometry through the Levi–Civita connection and the Riemann curvature tensor, and its extrinsic geometry through the Hellinger immersion and its second fundamental form. These objects quantify the failure of the statistical model to be locally flat, both as an abstract Riemannian manifold and as an immersed submanifold of the ambient Hilbert space. 3.1. Intrinsic curvature. Let (Θ, g) denote the statistical manifold induced by the Fisher–Rao metric gij (θ) = Iij (θ), where Θ ⊂ Rd is an open set with local coordinates θ = (θ1 , . . . , θd ). Throughout, indices i, j, k, ℓ, m, r range over {1, . . . , d}, and we adopt the Einstein summation convention over repeated upper and lower indices. Since g is a C 2 Riemannian metric, it admits a unique torsion-free and metric-compatible Levi– Civita connection ∇(Θ) , whose Christoffel symbols are 1 Γkij (θ) = g kℓ (θ) ∂i gjℓ + ∂j giℓ − ∂ℓ gij , 2 kℓ where g (θ) denotes the (k, ℓ)-entry of the inverse matrix g(θ)−1 , so that k g kℓ (θ) gℓm (θ) = δm .
The intrinsic curvature of the statistical manifold is encoded by the Riemann curvature tensor. Its (0, 4)-components are Rijkl (θ) = gim (θ) Rm jkl (θ), where the (1, 3)-curvature tensor is given in local coordinates by m r m r m Rm jkl (θ) = ∂k Γm ℓj − ∂ℓ Γkj + Γkr Γℓj − Γℓr Γkj .
If Rijkl vanishes identically on a neighborhood, then (Θ, g) is locally flat. Conversely, if Rijkl is nonzero at some point, then no smooth change of parameters can flatten the Fisher–Rao metric on any neighborhood of that point. This nonvanishing curvature is therefore an intrinsic, coordinateinvariant obstruction to reducing the local information geometry to a purely quadratic Euclidean form beyond first order. In geodesic normal coordinates centered at θ0 , the metric admits the expansion 1 gij (θ0 + δ) = δij − Rikjℓ (θ0 ) δ k δ ℓ + o(∥δ∥2 ), 3 so the quadratic deviation from the Euclidean metric is governed by curvature and cannot be removed when Rijkl (θ0 ) ̸= 0. 3.2. Extrinsic curvature through the Hellinger immersion. In addition to its intrinsic geometry, the statistical manifold carries an extrinsic geometry inherited from the Hellinger immersion √ ψ : Θ → L2 (µ), ψ(θ) = ψθ = pθ . We regard Θ as immersed into the ambient Hilbert space L2 (µ), equipped with its canonical inner product. 2 Let ∇(L ) denote the Levi–Civita connection of L2 (µ). Since the ambient metric is translation2 invariant, ∇(L ) is flat and coincides with the ordinary directional derivative in the ambient linear space. For the coordinate vector fields ∂ ∂i = i , ∂θ
ON HIGHER-ORDER GEOMETRIC REFINEMENTS OF CLASSICAL COVARIANCE ASYMPTOTICS
9
(L2 )
on Θ, the ambient second derivative ∇∂i (∂j ψθ ) admits a unique orthogonal decomposition into tangential and normal components relative to the immersed manifold. This is the Gauss formula (L2 )
∇∂i (∂j ψθ ) = Γkij (θ) ∂k ψθ + IIij (θ), where
(L2 ) IIij (θ) = Π⊥ ∇∂i (∂j ψθ ) ∈ Nψθ M,
is the vector-valued second fundamental form. Here Π⊥ denotes orthogonal projection onto the normal space Nψθ M = (Tψθ M)⊥ , and the tangent space is Tψθ M = span{∂1 ψθ , . . . , ∂d ψθ }. The tensor IIij measures the extrinsic bending of the immersed statistical manifold inside the ambient Hilbert space. It vanishes identically if and only if every ambient second derivative remains tangent to the model, that is, if and only if the immersion is totally geodesic in L2 (µ). Equivalently, ψ(Θ) is then locally contained in the intersection of the unit sphere with a fixed finite-dimensional affine subspace of L2 (µ). Note that since the second fundamental form arises from the second mixed partial derivatives ∂i ∂j ψθ , it is symmetric in its indices. This symmetry is used in the verification of the algebraic Bianchi identity later on. 3.3. The Gauss equation and the relation between intrinsic and extrinsic curvature. Because the ambient space L2 (µ) is flat, its curvature tensor vanishes identically. The intrinsic curvature of (Θ, g) is therefore completely determined by the second fundamental form through the Gauss equation. In coordinates, this yields Rijkl (θ) = 4 ⟨IIik (θ), IIjℓ (θ)⟩L2 − ⟨IIiℓ (θ), IIjk (θ)⟩L2 , where the factor 4 is consistent with our convention gij = 4⟨∂i ψθ , ∂j ψθ ⟩L2 . Thus the intrinsic curvature of the Fisher–Rao metric and the extrinsic bending of the Hellinger immersion are not independent. The former is recovered from the latter by a quadratic identity, reflecting the fact that the statistical manifold sits inside a flat ambient space. 3.4. A scalar measure of extrinsic bending. A natural scalar measure of extrinsic curvature is obtained by taking the squared norm of the second fundamental form. Contracting its two covariant indices with the inverse Fisher–Rao metric gives κ2 (θ) = g ik (θ) g jℓ (θ) ⟨IIij (θ), IIkℓ (θ)⟩L2 . This quantity is invariant under smooth reparameterizations of θ, and it can be interpreted as the Hilbert–Schmidt norm squared of the bilinear form IIθ . Equivalently, if {ea }da=1 is any g-orthonormal basis of Tθ Θ, then d X d X 2 κ2 (θ) = IIθ (ea , eb ) L2 . a=1 b=1
In particular, κ2 (θ) measures the total squared magnitude of the component of (L2 )
∇∂i (∂j ψθ ), that points orthogonally to the tangent space spanned by the vectors ∂k ψθ . It therefore quantifies how far the immersed model deviates from being locally flat in the ambient Hilbert space.
10
MALIK AMIR AND SOURANGSHU GHOSH
3.5. Relation with Efron’s statistical curvature. For a one-dimensional curved family, Efron’s asymptotic curvature theory shows that nonzero statistical curvature reduces the Fisher information retained by the maximum likelihood estimator θ̂n , viewed as a statistic, relative to the full sample. Let I(θ) denote the Fisher information for a single observation, so that the full i.i.d. sample carries Fisher information nI(θ). Under standard regularity conditions, the Fisher information contained in θ̂n admits the expansion κ2 (θ) 1 , Ieff (θ) = Iθ̂n (θ) = n I(θ) 1 − +o n n where κ2 (θ) ≥ 0 is Efron’s one-parameter statistical curvature. Note that this result concerns the information carried by the MLE as a statistic, not directly its variance. The actual variance of the MLE involves additional terms from the third-order likelihood geometry, as made precise by the correction tensor P in Section 5. In particular, the scalar κ2 (θ) appearing in Efron’s information expansion is related to, but not identical to, the extrinsic curvature ♯ scalar g ij Sij or the full correction g ij Pij . 3.6. Motivation for curvature-corrected covariance asymptotics. In higher-dimensional models, both the intrinsic curvature Rijkl and the extrinsic curvature encoded by IIij contribute to higher-order efficiency effects. Under standard regularity conditions, regular estimators may still attain the classical first-order asymptotic covariance n−1 I(θ)−1 , but nonzero curvature typically produces unavoidable second-order corrections of order n−2 to the covariance, and more generally to the bias–variance tradeoff. This motivates the study of matrix-valued curvature corrections to the classical first-order Fisherinformation asymptotics, in which the leading term I(θ)−1 /n is supplemented by explicit curvaturedependent contributions obtained from contractions of Rijkl and the second fundamental form IIij . Such second-order expansions explain geometrically why the usual Fisher-information approximation can become overly optimistic in strongly curved models, and especially in regimes where regularity deteriorates, such as near singular Fisher information or close to non-identifiability, where higher-order effects become pronounced. 4. Second-Order Geometry of the Log-Likelihood This section studies the log-likelihood beyond the Fisher-information level by analyzing its secondand third-order differential structure from a geometric point of view. We introduce the score and higher-order score tensors, explain why the observed Hessian is not tensorial under reparameterization, and replace it by the covariant Hessian, which restores coordinate invariance. These higherorder likelihood objects will later be related to intrinsic and extrinsic curvature and will furnish the geometric and probabilistic ingredients entering the second-order covariance expansion. 4.1. Score, Hessian, and Fisher information. Let X1 , . . . , Xn be i.i.d. samples drawn from a smooth parametric family {pθ : θ ∈ Θ ⊂ Rd }, of densities with respect to the measurable space (X , A, µ) introduced in Section 2. The loglikelihood is n X ℓ(θ) = log pθ (Xk ). k=1
Define the score vector and higher-order score tensors by Ui (θ) = ∂i ℓ(θ),
Uij (θ) = ∂i ∂j ℓ(θ),
Uijk (θ) = ∂i ∂j ∂k ℓ(θ).
ON HIGHER-ORDER GEOMETRIC REFINEMENTS OF CLASSICAL COVARIANCE ASYMPTOTICS
11
The score has zero expectation, and the negative expectation of the Hessian recovers the Fisher information, Eθ [−Uij (θ)] = −n Eθ [∂i ∂j log pθ (X)] = n Eθ [∂i log pθ (X) ∂j log pθ (X)] = n Iij (θ). This second-order structure is exactly what underlies the classical first-order Fisher-information asymptotics. It reflects only the local quadratic approximation of the log-likelihood and therefore captures only first-order statistical geometry. 4.2. Taylor expansion and departure from quadraticity. To go beyond the quadratic approximation, consider the Taylor expansion of ℓ(θ) around a fixed point θ, ℓ(θ + δ) = ℓ(θ) + Ui (θ) δ i + 21 Uij (θ) δ i δ j + 61 Uijk (θ) δ i δ j δ k + o(∥δ∥3 ), where Einstein summation is understood. The cubic term controls the leading deviation of the log-likelihood from a purely quadratic form. It therefore encodes information that is invisible to the Fisher information matrix alone and is the first place where genuinely higher-order geometric effects appear. 4.3. The covariant Hessian. Although the Fisher information Iij (θ) defines an intrinsic (0, 2)tensor, the observed Hessian Uij (θ) = ∂i ∂j ℓ(θ), is not itself tensorial under reparameterization. The intrinsic second derivative of the scalar field ℓ is instead the covariant Hessian (∇2 ℓ)ij (θ) = ∇i ∇j ℓ(θ) = ∂i ∂j ℓ(θ) − Γkij (θ) ∂k ℓ(θ) = Uij (θ) − Γkij (θ) Uk (θ). Equivalently, Uij (θ) = (∇2 ℓ)ij (θ) + Γkij (θ) Uk (θ). Thus the naive second derivatives of the log-likelihood decompose into a tensorial part and a coordinate-dependent correction involving the score. This reflects the fact that the observed Hessian transforms covariantly rather than tensorially. Taking expectations gives Eθ [Uij (θ)] = −n Iij (θ),
Eθ [(∇2 ℓ)ij (θ)] = −n Iij (θ),
so the connection correction disappears in expectation but remains essential at the sample level. This point is fundamental. Even though ℓ(θ) is a scalar, its second partial derivatives ∂i ∂j ℓ do not transform as the components of a (0, 2)-tensor, because a change of coordinates introduces an extra term proportional to the first derivatives ∂k ℓ = Uk . The affine connection term Γkij Uk is precisely what removes this defect and restores coordinate invariance. 4.4. Third-order likelihood geometry. The third-order score tensor contains genuinely new information beyond Fisher’s quadratic approximation. A direct computation gives h i Eθ [Uijk (θ)] = −n ∂k Iij (θ) − n Eθ ∂i ∂j log pθ (X) ∂k log pθ (X) . To organize this information geometrically, introduce the cubic score moment, also called the Amari–Chentsov tensor, Tijk (θ) = Eθ ∂i log pθ (X) ∂j log pθ (X) ∂k log pθ (X) , together with the exponential-connection coefficients h i (e) Γijk (θ) = Eθ ∂i ∂j log pθ (X) ∂k log pθ (X) . In terms of these tensors, the expectation of the third-order score tensor may be written in the symmetric form (e) (e) (e) Eθ [Uijk (θ)] = −n Γijk + Γikj + Γjki + Tijk .
12
MALIK AMIR AND SOURANGSHU GHOSH
This identity makes explicit that the first deviation from the quadratic Fisher approximation is governed by third-order objects. These higher-order likelihood tensors will later be reorganized into curvature-dependent contributions entering the second-order covariance expansion. 5. A Curvature-Corrected Second-Order Covariance Expansion This section derives the main second-order refinement of the classical first-order Fisher-information asymptotics. Starting from the asymptotic expansion of the mean-squared error of a locally unbiased estimator, we show that the n−2 correction admits a coordinate-invariant decomposition into an intrinsic term coming from the Riemann curvature of the Fisher–Rao metric and an extrinsic term coming from the second fundamental form of the Hellinger immersion. This is the point where the geometric structures introduced in the previous sections become an explicit quantitative correction to the usual first-order covariance approximation. 5.1. Asymptotic setting and second-order expansion. Let θ̂n = θ̂n (X1 , . . . , Xn ) be an estimator of the true parameter θ ∈ Θ ⊂ Rd . We restrict attention to estimators that are unbiased or locally unbiased at θ, meaning ∂i Eθ′ [θ̂na ]
Eθ [θ̂n ] = θ,
θ′ =θ
= δi a ,
where δi a is the Kronecker delta. Under the standing regularity assumptions of Section 2, and for asymptotically efficient regular estimators, we work on the regular locus Θreg = {θ ∈ Θ : λmin (I(θ)) > 0}. On any compact subset K ⊂ Θreg , we assume that θ̂n admits a stochastic expansion of the form 1 1 θ̂n = θ + √ A(θ) + B(θ) + op (n−1 ), n n where A(θ) and B(θ) are random vectors with bounded moments, and their distributions depend smoothly on θ. The leading stochastic fluctuations of θ̂n are therefore of order n−1/2 , while bias and higher-order geometric effects enter at order n−1 and smaller. The matrix mean-squared error at θ is i h MSEθ (θ̂n ) = Eθ (θ̂n − θ)(θ̂n − θ)⊤ = Covθ (θ̂n ) + Biasθ (θ̂n ) Biasθ (θ̂n )⊤ , where Biasθ (θ̂n ) = Eθ [θ̂n ] − θ, and h i Covθ (θ̂n ) = Eθ (θ̂n − Eθ [θ̂n ])(θ̂n − Eθ [θ̂n ])⊤ . Using the stochastic expansion, together with Eθ [A(θ)] = 0
and
Covθ (A(θ)) = I(θ)−1 ,
which encode first-order asymptotic efficiency, one obtains the second-order asymptotic expansion 1 1 1 −1 MSEθ (θ̂n ) = I(θ) + 2 C(θ) + o 2 , n n n uniformly for θ ∈ K.
ON HIGHER-ORDER GEOMETRIC REFINEMENTS OF CLASSICAL COVARIANCE ASYMPTOTICS
13
5.2. Structure of the second-order correction. To identify the second-order matrix C(θ), one must account for both geometric and probabilistic higher-order structure of the statistical model (Θ, g). The geometric contributions involve the Riemann curvature tensor Rijkl of the Levi–Civita connection of the Fisher–Rao metric g and the second fundamental form IIij of the Hellinger immersion √ Ψ : Θ → L2 (µ), Ψ(θ) = ψθ = pθ . Define the Ricci-type contraction of the intrinsic curvature and the Gram-type contraction of the second fundamental form: R♯ (θ) ij = g kℓ (θ) Rikjℓ (θ), S ♯ (θ) ij = g kℓ (θ) IIik (θ), IIjℓ (θ) L2 . The algebraic reduction carried out in the proof (Lemma 8.7) shows that the n−2 -coefficient in the covariance expansion is given by a second-order correction tensor Pij (θ) that admits the canonical decomposition 1 ♯ ♯ (θ) + Sij (θ) + Dij (θ), Pij (θ) = Rij 2 where Dij (θ) is the Hellinger discrepancy tensor. It captures the residual deviation between the probabilistic score-moment structure and the ambient L2 -geometry of the Hellinger immersion, and is defined explicitly in Lemma 8.7. Note the following important structural properties • S ♯ (θ) ⪰ 0 by construction. • In a full exponential family (where R♯ = 0, S ♯ ̸= 0 in general), the discrepancy Dij exactly cancels the extrinsic term, yielding Pij = 0. This is consistent with the well-known fact that full exponential families have no second-order correction. • When d = 1, R♯ = 0 identically and P11 = Varθ (∂11 log pθ (X)) − 41 (Eθ [∂111 log pθ (X)])2 in normal coordinates. The full second-order correction to the covariance is ♯ −1 −1 −1 1 ♯ R (θ) + S (θ) + D(θ) I(θ)−1 , C(θ) = I(θ) P (θ) I(θ) = I(θ) 2 and under the hypotheses of the theorem, for each fixed θ ∈ Θreg the covariance admits the expansion 1 1 1 −1 −1 −1 Covθ (θ̂n ) = I(θ) + 2 I(θ) P (θ) I(θ) + o 2 . n n n 5.3. Main theorem. We now state the main theorem. Its proof is given below, with the detailed technical lemmas and computations deferred to Appendix 8. Theorem 5.1 (Geometric decomposition of the second-order covariance correction). Let {pθ : θ ∈ Θ ⊂ Rd } be a C 3 parametric family with Fisher information matrix I(θ) = (gij (θ)) positive definite at θ. Equip Θ with the Fisher–Rao metric g and let Rikjℓ (θ) denote the Riemann curvature tensor of its Levi–Civita connection. Let √ Ψ : Θ → L2 (µ), Ψ(θ) = pθ , be the Hellinger (square-root density) immersion, with second fundamental form IIij (θ). Define the Ricci-type intrinsic contraction and the Gram-type extrinsic contraction ♯ Rij (θ) := g kℓ (θ) Rikjℓ (θ),
♯ Sij (θ) := g kℓ (θ) ⟨IIik (θ), IIjℓ (θ)⟩L2 .
Suppose θ̂n is a first-order efficient estimator for which the second-order covariance expansion Covθ (θ̂n ) =
1 1 I(θ)−1 + 2 C(θ) + o(n−2 ), n n
14
MALIK AMIR AND SOURANGSHU GHOSH
holds, under the regularity, moment, and stochastic-expansion conditions detailed in the Appendix (Section 8). Then the n−2 correction matrix C(θ) is given by C(θ) = I(θ)−1 P (θ) I(θ)−1 , where the second-order correction tensor P (θ) admits the canonical decomposition (5.1)
Pij (θ) =
1 ♯ Rij (θ) |2 {z }
+
intrinsic (Ricci-type)
♯ Sij (θ) | {z }
+
extrinsic (Gram-type)
Dij (θ). | {z }
Hellinger discrepancy
Here Dij (θ) depends on the fourth-order score moments and mixed third-order score–Hessian moments that are not captured by the L2 -geometry of Ψ. The decomposition (5.1) is coordinateinvariant and satisfies (1) S ♯ (θ) ⪰ 0 (positive semidefinite) by construction, (2) Pij (θ) ≡ 0 in any full exponential family, (3) when d = 1, R♯ ≡ 0 identically and P reduces to a purely extrinsic correction. Proof. The proof consists of five components: the identification of C(θ), the geometric decomposition, and the three structural properties. The detailed technical machinery about stochastic expansion analysis, moment bounds, Wick contractions, and the complete computation of Eθ [Bi,n Bj,n ], is given in Appendix (Section 8). Step 1: Identification of C(θ). Fix θ ∈ Θreg . The assumed stochastic expansion 1 1 θ̂n − θ = √ An (θ) + Bn (θ) + op (n−1 ), n n together with the moment and remainder bounds of Appendix 8 (specifically, those of Proposition 8.2 and the uniform integrability hypotheses), yields 1 1 ⊤ −2 Eθ [An A⊤ n ] + 2 Eθ [Bn Bn ] + o(n ). n n The first-order efficiency condition Covθ (An ) → I(θ)−1 recovers the leading n−1 I(θ)−1 -term. The negligible bias assumption yields Covθ (θ̂n ) = MSEθ (θ̂n ) + o(n−2 ). Hence MSEθ (θ̂n ) =
C(θ) = lim Eθ [Bn (θ)Bn (θ)⊤ ]. n→∞
Step 2: Computation of Eθ [Bi,n Bj,n ]. Choose Riemannian normal coordinates for the Fisher–Rao metric at θ, so that gij (θ) = δij and Γkij (θ) = 0 at the base point. The perturbative solution of the score equation U (θ̂n ) = 0 identifies (see Lemma 8.4) Bi,n = Hij,n Ajn + 21 Kijk,n Ajn Akn + op (1), where Hij,n = 1 3 n (∇ ℓ)ijk .
√1 ((∇2 ℓ)ij + ngij ) n
is the centered covariant-Hessian fluctuation and Kijk,n =
By the joint CLT and Wick contractions (Lemma 8.3), together with the replacement hypotheses, the second moment Eθ [Bi,n Bj,n ] converges to (e),m (e),k (e),k (e),m Pij (θ) = g km (θ) Eθ [sik (X)sjm (X)] − gik gjm + Γik Γjm + Γik Γjm + 14 κikℓ κjrs g kℓ g rs + g kr g ℓs + g ks g ℓr (e),k (e),r (e),s + 21 κjrs Γik g rs + Γik g ks + Γik g kr (e),m (e),k (e),ℓ + 12 κikℓ Γjm g kℓ + Γjm g mℓ + Γjm g mk ,
ON HIGHER-ORDER GEOMETRIC REFINEMENTS OF CLASSICAL COVARIANCE ASYMPTOTICS
15
(e)
where si (X) = ∂i log pθ (X), Γijk = Eθ [sij sk ], and κijk = Eθ [sijk ]. These second-moment computations are carried out in full in the appendix proof. In the inverse-metric frame, C(θ) = I(θ)−1 P (θ)I(θ)−1 . √ Step 3: Geometric decomposition. The Hellinger immersion Ψ(θ) = pθ provides the link between the probabilistic tensors above and Riemannian geometry. Setting ei = ∂i ψθ and eij = ∂ij ψθ , one has ei = 12 si ψθ ,
eij = 21 sij ψθ + 14 si sj ψθ .
In normal coordinates at θ, the Gauss decomposition gives eij = IIij (purely normal, since Γ = 0). From the ambient inner products of these vectors, the Gauss equation, and the tangency condition ⟨IIij , ek ⟩L2 = 0, one derives the identities (e)
Γijk = − 21 Tijk ,
κijk = 12 Tijk ,
where Tijk = Eθ [si sj sk ] is the cubic score moment. Substituting these relations into the expression for Pij and comparing with the Ricci-type and Gram-type contractions computed from the immersion, one obtains (by the algebraic reduction of Lemma 8.7) ♯ ♯ (θ) + Sij (θ) + Dij (θ), Pij (θ) = 21 Rij
where the Hellinger discrepancy Dij collects the terms involving fourth-order score moments and mixed third-order score–Hessian moments that are not determined by the immersion geometry. Step 4: Positive semidefiniteness of S ♯ . For any v ∈ Rd , ♯ v i v j Sij (θ) = g kℓ (θ) ⟨v i IIik (θ), v j IIjℓ (θ)⟩L2 .
In normal coordinates at θ, this becomes ♯ v v Sij (θ) = i j
d X
2
v i IIik (θ) L2 ≥ 0.
k=1
Hence S ♯ (θ) ⪰ 0. Step 5: Structural properties. (2) In a full exponential family, the sufficient statistic maps form a globally affine coordinate system in which the log-likelihood is exactly quadratic. All third and higher covariant derivatives of the log-likelihood are deterministic (in fact, zero in exponential coordinates), ♯ so that Dij exactly cancels Sij and the Riemann curvature vanishes. Hence Pij ≡ 0. (3) When d = 1, the Riemann curvature tensor Rijkl has only one independent component R1111 , and the symmetries of the Riemann tensor force R1111 = 0. Therefore R♯ ≡ 0, and the correction reduces to P = S ♯ + D, a purely extrinsic contribution. □ Remark 5.2 (Role of the Hellinger discrepancy). The Hellinger discrepancy Dij arises because the probabilistic content of the score moments is richer than what is captured by the L2 -geometry of the immersion Ψ. In particular, the fourth score moments E[si sj sk sℓ ] and the mixed third moments E[sij sk sℓ ] contribute to Pij but are not fully determined by the inner products ⟨IIik , IIjℓ ⟩L2 . Note that Dij vanishes identically whenever the second derivatives of the log-likelihood are deterministic given the parameter, which occurs in full exponential families. In such models, Pij = 0, consistent with the classical result that full exponential families achieve the Cramér–Rao bound to all orders.
16
MALIK AMIR AND SOURANGSHU GHOSH
5.4. Interpretation of the correction. This expansion makes precise how both curvature and higher-order probabilistic structure affect estimation beyond first order. The correction tensor Pij (θ) receives contributions from three sources (1) The intrinsic term 21 R♯ (θ), a Ricci-type contraction of the Riemann curvature of the Fisher– Rao metric, which measures the failure of local flatness. (2) The extrinsic term S ♯ (θ) ⪰ 0, a Gram-type contraction of the second fundamental form of the Hellinger immersion, which measures bending of the model in the ambient Hilbert space. (3) The Hellinger discrepancy D(θ), which captures higher-order probabilistic content (fourth score moments and mixed cubic moments) not fully determined by the L2 -geometry of the immersion. When R♯ (θ) is positive semidefinite, it increases the second-order covariance term. The extrinsic contribution S ♯ (θ) is always positive semidefinite by construction. The discrepancy D(θ) can be negative (as in full exponential families, where it exactly cancels S ♯ to give P = 0). In full exponential families, Pij (θ) ≡ 0, and the covariance expansion reduces to the classical first-order Fisher-information term up to order n−2 . More generally, the magnitude and sign of Pij quantify the departure of the model from exponential-family behavior at second order. Thus the usual first-order covariance approximation is not universal at finite sample size. In non-exponential models, the correction tensor P is generically nonzero, producing a systematic n−2 correction whose geometric and probabilistic content is made explicit by the decomposition P = 21 R♯ + S ♯ + D. 5.5. Relation to classical higher-order asymptotic refinements. Classical higher-order refinements in parametric inference, such as Bhattacharyya-type inequalities and related likelihood expansions, also incorporate higher derivatives of the likelihood by enlarging the class of estimating functions or moment constraints entering a Cauchy–Schwarz argument. The present approach is compatible with this general principle, since the second-order covariance correction derived here ultimately depends on third-order likelihood geometry. The difference is mainly organizational and invariant-theoretic. First, the relevant higher-order derivative information is packaged into coordinate-invariant geometric tensors, namely the Fisher– Rao Riemann tensor Rikjl and the second fundamental form IIij of the Hellinger immersion, rather than being retained as coordinate-dependent collections of higher derivatives. Second, the extrinsic term is automatically positive semidefinite because it is a Gram-type contraction in the ambient Hilbert space, yielding a transparent monotonicity property in which stronger bending produces a larger second-order covariance contribution. Third, the split into intrinsic and extrinsic contributions clarifies the mechanism of the correction by separating non-flattenability of the Fisher geometry from bending induced by the immersion into the canonical L2 space. In this sense, the curvature terms are not merely a reformulation of higher-order derivatives. They isolate the canonical invariant content of the third-order expansion that survives reparameterization and carries a natural geometric sign structure. 6. Extension to Singular Models Extending the preceding curvature-based second-order covariance analysis from regular to singular statistical models requires a fundamental re-examination of the geometric and analytic structures underlying parametric inference. 6.1. Geometric and analytic structure of regular and singular statistical models. In the regular setting, one begins with a parametric statistical model M = {pθ : θ ∈ Θ ⊂ Rd }
ON HIGHER-ORDER GEOMETRIC REFINEMENTS OF CLASSICAL COVARIANCE ASYMPTOTICS
where Θ is an open subset of Rd and the map Φ : Θ → L2 (X , µ),
Φ(θ) =
√
17
pθ
is assumed to be a smooth embedding. That is, Φ is injective and its differential DΦ(θ) has full rank d for all θ in a neighborhood of θ0 . The tangent map is explicitly given by ∂ √ 1 pθ (x) = pθ (x)−1/2 ∂i pθ (x) ∂θi 2 and hence the induced inner product on Tθ Θ is Z ∂ √ 1 ∂ √ 1 pθ , pθ = gij (θ) = pθ (x)−1 ∂i pθ (x)∂j pθ (x) dµ(x) = Iij (θ). ∂θi ∂θj 4 4 L2 Remark 6.1. The pullback metric of the Hellinger immersion is gij = 14 Iij . In Sections 2–5, we adopted the rescaled convention gij = Iij = 4⟨∂i ψθ , ∂j ψθ ⟩L2 to simplify the relationship between the metric and the Fisher information. Here, in the singular context, it is sometimes more natural to work with the raw pullback 14 Iij . Since a constant rescaling does not affect the Levi–Civita connection or the (1, 3)-curvature tensor, the qualitative conclusions are unaffected, and only the numerical coefficients in the Gauss equation and the scalar curvature are rescaled. We will continue to write gij = 14 Iij in this section, noting that the Gauss equation becomes Rijkl = ⟨IIik , IIjl ⟩ − ⟨IIil , IIjk ⟩ (without the factor of 4 used in the earlier convention). With this convention, up to a constant factor the Fisher information matrix defines the Riemannian metric tensor. The assumption of strict positive definiteness, det I(θ0 ) > 0, implies the existence of λmin > 0 such that the quadratic form Q(v) = v ⊤ I(θ0 )v ≥ λmin |v|2
∀v ∈ Rd ,
ensuring that the metric is non-degenerate and that Θ is locally diffeomorphic to Rd . Consequently, the Levi–Civita connection is well-defined via 1 Γkij (θ) = I kℓ (θ) (∂i Ijℓ (θ) + ∂j Iiℓ (θ) − ∂ℓ Iij (θ)) 2 and curvature tensors may be constructed in the standard manner. Proposition 6.2 (Quadratic approximation of the KL divergence in the regular case). Under the regularity assumptions of Section 2, the Kullback–Leibler divergence admits the expansion 1 K(θ0 + h) = h⊤ I(θ0 )h + o(|h|2 ) 2 as h → 0. Proof. Write Z K(θ0 + h) =
pθ0 (x) log
pθ0 (x) dµ(x). pθ0 +h (x)
Expanding log pθ0 +h (x) in h, 1 log pθ0 +h (x) = log pθ0 (x) + hi ∂i log pθ0 (x) + hi hj ∂i ∂j log pθ0 (x) + o(|h|2 ). 2 Substituting and using the standard identities (from differentiation under the integral sign) Z pθ0 (x)∂i log pθ0 (x) dµ(x) = 0, Z pθ0 (x)∂i ∂j log pθ0 (x) dµ(x) = −Iij (θ0 ),
18
MALIK AMIR AND SOURANGSHU GHOSH
one obtains
1 K(θ0 + h) = h⊤ I(θ0 )h + o(|h|2 ). 2 In contrast, in singular models, one assumes the existence of θ0 such that
□
rank I(θ0 ) = r < d. Proposition 6.3 (Null directions of singular Fisher information). If rank I(θ0 ) = r < d, then the nullspace N = {v ∈ Rd : I(θ0 )v = 0}, has dimension d − r > 0, and for every v ∈ N , X vi ∂i log pθ0 (x) = 0 for µ-almost every x. i d pγ(t) (x) t=0 = 0 for µ-a.e. x. In particular, along the curve γ(t) = θ0 + tv, the density satisfies dt
Proof. For any v ∈ N , ⊤
Z
0 = v I(θ0 )v =
!2 X
vi ∂i log pθ0 (x)
pθ0 (x) dµ(x).
i
P Since the integrand is non-negative and pθ0 > 0 a.e., we conclude i vi ∂i log pθ0 (x) = 0 for µ-a.e. P d pγ(t) (x)|t=0 = i vi ∂i pθ0 (x) = 0 a.e. □ x. Multiplying by pθ0 (x) gives dt This establishes the non-injectivity of the parametrization along null directions. Higher-order derivatives may also vanish, leading to pγ(t) (x) = pθ0 (x) + O(tk ) for k ≥ 2, and the Kullback–Leibler divergence satisfies a higher-order expansion where the quadratic term vanishes along directions in N . Consequently, the local geometry cannot be described by a non-degenerate quadratic form, and the parameter space acquires the structure of a real-analytic variety with singularities. The study of such structures necessitates the use of tools from real algebraic geometry, including resolution of singularities. 6.2. Resolution of singularities and monomialization of the KL divergence. To systematically extract the local analytic structure of the Kullback–Leibler divergence near a singular point θ0 ∈ Θ, one invokes Hironaka’s resolution of singularities theorem in the real-analytic category. This guarantees the existence of a proper real-analytic mapping e → Θ, π:Θ e is a smooth manifold and π is obtained as a finite composition of blow-up maps along such that Θ smooth centers π = π1 ◦ π2 ◦ · · · ◦ πN . Remark 6.4 (Normal crossing forms). In full generality, Hironaka’s theorem produces a multiplicative Q 2k normal crossing form K(π(u)) = rj=1 uj j ·φK (u), where φK is smooth and positive near u = 0. For expository clarity and explicit computability, we work throughout this section under the assumption that the resolved coordinates can be chosen so that the KL divergence admits an additive normal crossing representation r X 2k K(π(u)) = cj uj j , cj > 0, kj ∈ N. j=1
This additive form arises when the singularity decouples along coordinate axes in the resolved space, as occurs in many concrete examples such as certain mixture models and rank-deficient models with independent degenerate directions. The qualitative features of the theory, such as the role of the
ON HIGHER-ORDER GEOMETRIC REFINEMENTS OF CLASSICAL COVARIANCE ASYMPTOTICS
19
RLCT, the modified convergence rates, and the curvature extensions, all hold under the general multiplicative form as well (see [21] for the general theory). Under the additive assumption of Remark 6.4, we proceed with explicit computations. Proposition 6.5 (Jacobian of the resolution map). For a single blow-up of the form x1 = u1 , x2 = u1 u2 , . . . , xs = u1 us , the Jacobian determinant is det Dπ = us−1 1 . After composing N such blow-ups, the full Jacobian factorizes as | det Dπ(u)| =
r Y
|uj |hj · φ(u),
j=1
where hj ∈ Z≥0 are accumulated from successive blow-ups and φ(u) is smooth with φ(0) ̸= 0. Proof. For the single blow-up, the Jacobian matrix is 1 0 ··· u2 u1 · · · Dπ = .. .. . . . . . us
0 0 .. , .
· · · u1
0
which is block-triangular with diagonal entries 1, u1 , . . . , u1 . Hence det Dπ = us−1 1 . The general statement follows by the multiplicativity of determinants under composition and induction on the number of blow-ups. □ Proposition 6.6 (Hessian and metric structure on the resolved manifold). Under the additive assumption, the Hessian of K(π(u)) is diagonal Gij (u) :=
∂2 2k −2 K(π(u)) = δij · 2kj (2kj − 1)cj uj j . ∂ui ∂uj
In particular (1) If kj = 1, then Gjj (u) = 2cj is constant and non-degenerate. (2) If kj ≥ 2, then Gjj (u) → 0 as uj → 0, so the metric degenerates along {uj = 0}. P 2k Proof. Since K(π(u)) = rj=1 cj uj j is a sum of univariate terms, the mixed partial derivatives 2
2
2k
2k −2
j ∂ vanish: ∂u∂i ∂uj K(π(u)) = 0 for i ̸= j. The diagonal terms are ∂u ) = 2kj (2kj − 1)cj uj j 2 (cj uj
.
j
The degeneracy claims follow immediately from the exponents.
□
The resolved space therefore admits a stratification G e= e : uj = 0 for j ∈ I, uj ̸= 0 for j ∈ Θ SI , SI := {u ∈ Θ / I}, I⊂{1,...,r}
where each stratum SI is a smooth manifold with effective metric rank |{j ∈ / I : kj = 1}|. This stratified structure reflects the intrinsic anisotropy of the statistical model. 6.3. Differential geometric structure on the resolved manifold. On the regular part e reg := Θ e\ Θ
r [
{uj = 0},
j=1
the pullback Fisher-type metric tensor geij (u) := Gij (u) is smooth and non-degenerate, hence defines a Riemannian metric.
20
MALIK AMIR AND SOURANGSHU GHOSH
e reg , the Levi–Civita connection of ge has Proposition 6.7 (Curvature of the resolved metric). On Θ Christoffel symbols e m = 1 gemk (∂i gejk + ∂j geik − ∂k geij ) . Γ ij 2 2k −2
Since geij is diagonal with gejj (u) = 2kj (2kj − 1)cj uj j e j = kj − 1 , Γ jj uj
, the Christoffel symbols simplify to
e j = 0 for i ̸= j, Γ ij
e i = 0 for i ̸= j. Γ jj
The Riemann curvature tensor is em = ∂j Γ e m − ∂k Γ em + Γ em Γ eℓ em eℓ R ij ijk ik jℓ ik − Γkℓ Γij . eik = R em and the scalar curvature is R e = geik R eik . The Ricci curvature is R imk Proof. The diagonal structure of ge implies that the only nonzero first derivatives are ∂j gejj = 2k −3 2kj (2kj −1)(2kj −2)cj uj j . Substituting into the Christoffel symbol formula and using gejj = 1/e gjj , j 1 e = gejj ∂j gejj = (kj − 1)/uj . All mixed Christoffel symbols vanish because ge is diagoone obtains Γ jj 2 nal and its off-diagonal derivatives are zero. The Riemann tensor formula follows from the standard definition. □ √ For the extrinsic geometry, the embedding Φ(u) = pπ(u) ∈ L2 (X , µ) provides a second fundamental form e ij = Π⊥ (∂i ∂j Φ(u)) , II e reg and the Gauss and Codazzi equations hold on Θ eijkl = ⟨II e ik , II e jl ⟩ − ⟨II e il , II e jk ⟩, R
e i II e jk = ∇ e j II e ik . ∇
Verification of the Gauss equation. Since the ambient space L2 (µ) is flat, the Gauss equation for an isometric immersion of a Riemannian manifold into a flat ambient space takes the standard form. e ij is the normal The tangent vectors ∂i Φ(u) span the tangent space, the second fundamental form II component of ∂i ∂j Φ, and the equation follows from the decomposition of the ambient curvature (which is zero) into tangential and normal components via the Gauss–Codazzi formalism. □ 6.4. The Real Log Canonical Threshold (RLCT). The asymptotic behavior of statistical estimators in singular models is fundamentally governed by the real log canonical threshold. Definition 6.8 (RLCT via the zeta function). Let K(θ) ≥ 0 be a real-analytic function with K(θ0 ) = 0, and let φ(θ) > 0 be a smooth prior density. The zeta function of the pair (K, φ) is Z ζ(z) = K(θ)z φ(θ) dθ, Θ
defined initially for Re(z) > 0 (where the integral converges) and extended by analytic continuation to a meromorphic function on C. The real log canonical threshold (RLCT) is the smallest positive pole of ζ(z), or equivalently, the unique λ > 0 such that ζ(z) has a pole at z = −λ and is holomorphic on {−λ < Re(z)}. Remark 6.9. The existence of the meromorphic continuation and the rationality of the poles follow from Hironaka’s resolution of singularities combined with Bernstein–Sato theory (see Atiyah [21] and references therein). Under the additive normal crossing assumption of Remark 6.4, the RLCT and the Laplace asymptotics can be computed by direct calculation, as we now show.
ON HIGHER-ORDER GEOMETRIC REFINEMENTS OF CLASSICAL COVARIANCE ASYMPTOTICS
21
Proposition 6.10 (RLCT under the additive assumption). Under the additive normal crossing P Q 2k representation K(π(u)) = rj=1 cj uj j with Jacobian | det Dπ(u)| = rj=1 |uj |hj ·φ(u), the stochastic complexity integral Z exp − nK(θ) φ(θ) dθ, Zn := Θ
satisfies −λ
Zn ∼ C n
,
λ=
r X hj + 1 j=1
2kj
,
as n → ∞, where C > 0 is a computable constant. Proof. Under the change of variables θ = π(u), we have Z r r X Y 2k Zn = exp −n cj uj j |uj |hj ψ(u) du, e Θ
j=1
j=1
where ψ(u) = φ(π(u))φ(u) is smooth and positive with ψ(0) > 0. Since the dominant contribution arises from a neighborhood of u = 0, we approximate ψ(u) ≈ ψ(0) and note that the exponential of the sum factorizes r Z ε Y 2k Zn ∼ ψ(0) exp(−ncj uj j )|uj |hj duj . j=1 | −ε {z } =:Ij (n)
For each factor, the scaling substitution tj = n1/(2kj ) uj gives h +1
j − 2k
Ij (n) = n
Z n1/(2kj ) ε
j
−n1/(2kj ) ε
h +1
2k
exp(−cj tj j )|tj |hj dtj −→ n
j − 2k
j
Aj ,
where Z ∞ Aj :=
exp(−cj t2kj )|t|hj dt < ∞.
−∞
(The finiteness of Aj follows because exp(−cj t2kj ) decays super-polynomially.) Multiplying the factors yields r P hj +1 Y − rj=1 2k j = C n−λ , Zn ∼ ψ(0) Aj n j=1
with λ =
hj +1 j=1 2kj .
Pr
□
Remark 6.11 (Comparison with Watanabe’s general theory). In Watanabe’s general framework Q 2k using the multiplicative normal crossing form K(π(u)) = j uj j · φK (u), the RLCT takes the value h +1
j λ = minj 2k , with possible logarithmic corrections Zn ∼ C n−λ (log n)m−1 when the minimum is j P attained by m > 1 indices. The difference between min and reflects the different local structure of the singularity. In the multiplicative case, K vanishes on coordinate hyperplanes (not just at the origin), leading to qualitatively different asymptotics. The qualitative conclusions that convergence rates are governed by algebraic invariants (kj , hj ) rather than the parameter dimension d hold in both cases.
22
MALIK AMIR AND SOURANGSHU GHOSH
6.5. Asymptotic posterior mean squared error. Proposition 6.12 (Posterior MSE rate). Under the additive assumption with prior φ(θ) > 0, the Bayesian posterior mean squared error satisfies − minj k1 j , Eπn ∥θ − θ0 ∥2 ∼ C n where πn (θ) = exp(−nK(θ))φ(θ) is the posterior distribution. Zn P Proof. Substituting θ = π(u) and using π(u) − θ0 = rj=1 aj uj + O(∥u∥2 ) for vectors aj ∈ Rd , we have r X X 2 ∥π(u) − θ0 ∥ = bj u2j + bjℓ uj uℓ + O(∥u∥3 ), j=1
j̸=ℓ
where bj = ∥aj ∥2 > 0 and bjℓ = ⟨aj , aℓ ⟩. By symmetry of the integrand, the cross terms uj uℓ (with P 2k j ̸= ℓ) vanish upon integration against the even function exp(−n cj uj j ). Hence the numerator Z
2
∥θ − θ0 ∥ exp(−nK(θ))φ(θ) dθ ∼
r X
Z
bj
j=1
Y 2k Iℓ (n). u2j exp(−ncj uj j )|uj |hj duj · {z } ℓ̸=j | =:Ij′ (n)
The scaling tj = n1/(2kj ) uj gives h +3
Z
− j Ij′ (n) = n 2kj Bj ,
Since
t2 exp(−cj t2kj )|t|hj dt < ∞.
Bj :=
hj +3 hj +1 1 2kj = 2kj + kj , the j-th term in the numerator scales as h +3
n
j − 2k
j
·
Y
h +1
n
ℓ − 2k
ℓ
−λ− k1
=n
j
.
ℓ̸=j
The dominant term corresponds to the smallest additional decay, i.e., minj k1j , giving Numerator ∼ C1 n
−λ−minj k1
j
.
Dividing by Zn ∼ C0 n−λ from Proposition 6.10, we obtain − minj k1
Eπn [∥θ − θ0 ∥2 ] ∼ C n
j
.
In the regular case kj = 1 for all j, this gives n−1 . In the singular case with kj ≥ 2 for some j, the rate is n− minj 1/kj ≫ n−1 . □ 6.6. Tangent cone geometry and construction of the effective information metric. To extend the preceding geometric covariance analysis to the singular setting, one must replace the classical notion of a tangent space by the tangent cone at the singular point θ0 ∈ Θ. Definition 6.13 (Tangent cone). Let K(θ) be real-analytic with K(θ0 ) = 0 and let 2k be the order of the leading nonvanishing term in its Taylor expansion. Define the homogeneous polynomial X 1 Φ(v) := ∂ α K(θ0 )v α . α! |α|=2k
The tangent cone at θ0 is n o Tθ0 = v ∈ Rd : K(θ0 + tv) = O(t2k ) as t → 0 .
ON HIGHER-ORDER GEOMETRIC REFINEMENTS OF CLASSICAL COVARIANCE ASYMPTOTICS
23
Proposition 6.14 (Properties of the tangent cone). Tθ0 is a closed cone, i.e., v ∈ Tθ0 and λ ≥ 0 imply λv ∈ Tθ0 . The leading-order asymptotic behavior of K along rays is K(θ0 + tv) ∼ t2k Φ(v) as t → 0. Proof. K(θ0 + t(λv)) = K(θ0 + (λt)v) ∼ (λt)2k Φ(v) = t2k λ2k Φ(v) = O(t2k ), so λv ∈ Tθ0 .
□
To extract a bilinear structure, define the generalized metric 1 1 1 G(v, w) := lim 2k [K(θ0 + t(v + w)) − K(θ0 + tv) − K(θ0 + tw)] = [Φ(v + w) − Φ(v) − Φ(w)] . t→0 t 2 2 When k = 1, G is a symmetric bilinear form recovering the Fisher information inner product. When k ≥ 2, G(v, w) is symmetric but not bilinear. Nevertheless, it captures the leading-order distinguishability and satisfies G(v, v) = Φ(v). 6.7. Pushforward of curvature tensors under resolution of singularities. Curvature opere and II e from Θ e ators on the original parameter space are defined by pushing forward the tensors R to Θ via the resolution map π. Definition 6.15 (Pushed-forward curvature). On the regular part where Dπ has maximal rank, introduce a right-inverse (Dπ)† satisfying (Dπ)αi (Dπ)† iβ = δβα . The pushed-forward curvature tensor and extrinsic contraction are em (u) (Dπ)† i (Dπ)† jγ (Dπ)† k , Rα (θ) = (Dπ)αm R ijk
βγδ
δ
β
e kℓ (u)⟩ gekℓ (u), e ij (u), II Sαβ (θ) = (Dπ)† iα (Dπ)† jβ ⟨II ♯ = Sαβ . where u is any preimage of θ under π. The contracted operators are R♯αβ = Rγαγβ , Sαβ Proposition 6.16 (Generalized asymptotic covariance expansion). Under the additive resolution and the standing assumptions, the asymptotic covariance of the estimator takes the form −2λ −1 −µ 1 ♯ ♯ Covθ0 (θ̂n ) ∼ n G +n R +S , 2 where µ = minj k1j arises from the next-order scaling. Proof. We outline the argument. The estimator in resolved coordinates has the expansion ûjn = n−1/(2kj ) Zj +n−(1/(2kj )+µj ) Bj +· · · , where Zj are the leading-order random variables. Transforming back via θ̂n −θ0 = (Dπ)(0)ûn + 12 ∂ 2 π(0)[ûn , ûn ]+· · · , the leading covariance is (Dπ) Cov(ûn )(Dπ)⊤ ∼ n−2λ G −1 . The next-order correction arises from the nonlinearity of π (through ∂ 2 π) and the thirde and II e via the order log-likelihood terms, which are expressible through the curvature tensors R 1 ♯ ♯ −µ Gauss equation. After pushforward, these yield 2 R + S at the next order n . □ 6.8. Recovery of the classical theory. Proposition 6.17 (Regular case recovery). In the special case kj = 1 and hj = 0 for all j = 1, . . . , d, the singular framework reduces exactly to the classical second-order theory. P Proof. When kj = 1 and hj = 0, the resolved KL divergence is K(π(u)) = dj=1 cj u2j , which is a non-degenerate quadratic form. By Proposition 6.10, λ=
d X 0+1 j=1
2·1
=
d , 2
µ = min j
1 = 1. 1
Since hj = 0, the Jacobian is non-vanishing | det Dπ(u)| = φ(u) with φ(0) ̸= 0, so π is a local diffeomorphism. The resolved metric Gjj = 2cj is constant and non-degenerate, so the geometry is locally Euclidean.
24
MALIK AMIR AND SOURANGSHU GHOSH
The integral Zn evaluates exactly Zn ∼ φ(0)
d r Y π j=1
ncj
= Cn−d/2 ,
recovering λ = d/2. Identifying I(θ0 ) = 2(A−1 )⊤ CA−1 where A = Dπ(0) and C = diag(c1 , . . . , cd ), we recover K(θ) = 21 (θ − θ0 )⊤ I(θ0 )(θ − θ0 ) + O(∥θ − θ0 ∥3 ). The covariance expansion becomes 1 1 1 ♯ −1 ♯ Covθ0 (θ̂n ) = I(θ0 ) + 2 R + S + o(n−2 ), n n 2 which coincides with the general singular expansion when λ = d/2 and µ = 1, and matches Theorem 5.1 in the regular case (up to the discrepancy D, which arises from the more detailed analysis of score moments beyond what the immersion geometry determines). □ 7. Conclusion and Outlook Classical Fisher-information asymptotics are foundational in statistical estimation, but their standard form depends only on the local quadratic structure of the model. In particular, they capture only the first-order geometry of the statistical manifold through the Fisher–Rao metric and are therefore insensitive to higher-order effects such as intrinsic curvature and extrinsic bending. By realizing a regular parametric model as a Riemannian manifold (Θ, g) and immersing it into an am√ bient Hilbert space through the Hellinger map θ 7→ ψθ = pθ , these higher-order geometric features can be incorporated systematically into second-order asymptotic expansions for estimation. Under the regularity assumptions developed in Section 2, and under the additional moment, approximation, and replacement hypotheses used throughout the proof, we obtained for each fixed θ ∈ Θreg a second-order refinement of the covariance expansion whose leading term is the classical Fisher-information contribution and whose next-order term is governed by curvature. More precisely, the covariance admits the expansion 1 1 1 −1 −1 −1 Covθ (θ̂n ) = I(θ) + 2 I(θ) P (θ) I(θ) + o 2 , n n n where P (θ) is the second-order correction tensor, admitting the canonical decomposition 1 P = R♯ + S ♯ + D, 2 ♯ with R a Ricci-type contraction of the Riemann curvature tensor, S ♯ ⪰ 0 a Gram-type contraction of the second fundamental form, and D a Hellinger discrepancy tensor capturing higher-order probabilistic content not determined solely by the immersion geometry. This decomposition clarifies distinct mechanisms contributing to the second-order behavior of the covariance. The intrinsic term records the failure of the Fisher geometry to be flattened beyond first order. The extrinsic term measures the bending of the statistical model inside the ambient space of square-root densities. The discrepancy term captures the residual deviation between the probabilistic score-moment structure and the ambient L2 -geometry. A key structural property is that Pij ≡ 0 in full exponential families, where D exactly cancels the extrinsic contribution S ♯ . In nonlinear models such as mixtures, latent-variable models, or parameter spaces constrained by geometry, the correction tensor P is typically nonzero, providing a geometric explanation for finitesample deviations from Fisher-information-only predictions. Beyond refining first-order asymptotics, the curvature-aware viewpoint suggests a more nuanced interpretation of statistical efficiency. Estimation error is governed not only by local sensitivity encoded in I(θ), but also by higher-order geometric structure, including non-flatness of the Fisher metric, bending of the Hellinger immersion, and near-singular directions in which distinguishability
ON HIGHER-ORDER GEOMETRIC REFINEMENTS OF CLASSICAL COVARIANCE ASYMPTOTICS
25
deteriorates beyond first order. In this sense, curvature provides a principled quantitative language for describing weak identifiability, higher-order inefficiency, and the breakdown of naive quadratic approximations. These conclusions also point toward applications in modern probabilistic learning systems. In many overparameterized settings, many parameter values induce nearly indistinguishable predictive distributions and therefore achieve nearly equivalent empirical performance. The correction tensor P provides a coordinate-invariant way to distinguish between such solutions at second order, thereby suggesting curvature-aware principles for regularization, diagnostics, and optimization. We now briefly describe these directions. 7.1. Applications to deep learning training and regularization. In many deep learning settings, the model qw (· | x) defines a conditional probability distribution parameterized by a highdimensional weight vector w ∈ Rp , and the loss function is the negative log-likelihood or crossentropy n 1X L(w) = − log qw (yi | xi ). n i=1
When the model is overparameterized, the set of near-optimal weights often forms a manifold, or near-manifold, in parameter space. The resulting question is not only how to fit the data, but also which nearly equivalent solution should be preferred. The second-order correction tensor P (w) suggests a principled answer. Small values of P (w), or of suitable contractions of P (w), indicate that the model departs less strongly from exponentialfamily behavior at second order and therefore has a smaller second-order covariance correction. This motivates curvature-aware regularization strategies based on intrinsic curvature, extrinsic curvature, or the full correction tensor. For example, one may penalize the scalar Ricci-type contraction ♯ R(w) = g ij Rij (w),
the total extrinsic curvature ♯ κ2 (w) = g ij Sij (w), or more directly the trace of the covariance correction
tr I(w)−1 P (w)I(w)−1 . Such penalties favor regions of parameter space in which the second-order geometric distortion is weaker. This viewpoint also clarifies the relation with flat-minimum heuristics and sharpness-aware training. The Fisher–Rao metric already gives a coordinate-invariant first-order notion of local sensitivity, while the tensor P refines this by incorporating higher-order effects. In particular, a point where P (w) is small is one where the usual Fisher approximation remains accurate to second order, so that the model is not only locally stable at first order but also weakly curved in the higher-order sense identified by the present theory. Beyond regularization, the geometric quantities R♯ , S ♯ , and P may also serve as diagnostics. Large values of ∥P ∥ relative to ∥I −1 ∥ indicate that second-order effects are substantial and that first-order Fisher-information asymptotics alone may be unreliable. This suggests uses in weakidentifiability detection, in model comparison when first-order performance is similar, and in the analysis of optimization trajectories by monitoring whether training moves toward regions of lower or higher curvature. A further natural direction concerns optimization. The natural gradient method replaces the Euclidean gradient by the Fisher-adjusted gradient I(w)−1 ∇L(w), thereby incorporating first-order information geometry into the training dynamics. The present framework suggests a refinement in which second-order geometric information enters through P (w), at least approximately, so that one replaces the purely Fisher-based preconditioner by a curvature-corrected one. Whether such
26
MALIK AMIR AND SOURANGSHU GHOSH
corrections can be made computationally useful remains an open problem, but the formal structure of the covariance expansion points naturally in this direction. The main practical obstacle is computational. In high-dimensional models, the relevant geometric quantities are rarely available in closed form. The Fisher metric, its derivatives, and the second derivatives of the Hellinger immersion must typically be estimated numerically, and the full Fisher–Rao curvature tensor has O(d4 ) components. The theory should therefore not be interpreted as requiring exact computation of complete curvature tensors. What enters the expansion are only specific contractions of P , and these can in principle be approximated without explicitly forming Rikjl or IIij . Possible strategies include Monte Carlo or minibatch estimates of score moments, structured or low-rank approximations of the Fisher information and its inverse, Fisher-vector products combined with iterative solvers, automatic differentiation for directional second- and third-order derivatives, and stochastic trace estimators for Ricci-type or trace-type contractions. Overall, the geometry-aware framework developed here provides a coordinate-invariant and secondorder accurate refinement of classical Fisher-information asymptotics. It yields a transparent decomposition of the second-order covariance term into intrinsic and extrinsic geometric contributions, and it suggests that many finite-sample phenomena traditionally viewed as analytic complications are in fact manifestations of underlying statistical curvature. A natural continuation of this program is to weaken the technical hypotheses needed for the expansion, to extend the framework beyond the score-root and first-order efficient setting, to clarify the singular counterpart more fully, and to connect curvature-aware second-order asymptotics with practical optimization, stability, and regularization principles in modern nonlinear learning systems. 8. Appendix 8.1. Why is an immersion sufficient? In this appendix we explain why the curvature-aware Cramér–Rao refinements developed in the main text require only that the square-root density map √ Ψ : Θ → L2 (µ), Ψ(θ) = ψθ = pθ , be a C 3 immersion, rather than a global embedding. The key point is that every geometric object entering the second-order correction, namely the Fisher–Rao metric, its Levi–Civita connection, the Riemann curvature tensor, and the second fundamental form, is defined from local derivatives of Ψ up to order three at a fixed parameter value θ. Consequently, global injectivity of the parametrization and global topological regularity of the image are not needed for the local differential-geometric analysis underlying the bound. Let M be a smooth d-manifold and H a possibly infinite-dimensional Hilbert space. A C 1 map F : M → H is an immersion if its differential dFθ : Tθ M → H is injective for every θ ∈ M . A map is an embedding if it is an immersion and a homeomorphism onto its image F (M ) equipped with the subspace topology. All constructions in Sections 2–5 are local on Θ and depend only on derivatives of ψθ in the ambient Hilbert space. The Fisher–Rao metric is defined by pullback, the Levi–Civita connection and Riemann curvature tensor are computed from g and its derivatives, and the extrinsic geometry is defined pointwise by orthogonal projection. For the square-root map, the immersion condition is equivalent to local identifiability in the Fisher–Rao sense: dΨθ injective ⇐⇒ g(θ) ≻ 0. By the local immersion theorem, every immersion is locally an embedding. Hence the model is locally identifiable near θ, even if global identifiability fails elsewhere. If Ψ is not globally injective, the set-theoretic image may have self-intersections. This does not affect the curvature-corrected bounds, which are stated at a fixed parameter value and depend only on local derivatives.
ON HIGHER-ORDER GEOMETRIC REFINEMENTS OF CLASSICAL COVARIANCE ASYMPTOTICS
27
8.2. Detailed standing assumptions for the proof of the main theorem. For reference, we record the full set of technical assumptions used in the proof of Theorem 5.1. Let K ⊂ Θreg be compact. (1) The standing C 3 regularity assumptions of Section 2 hold, including differentiation under the integral sign and smoothness of the Fisher metric on Θreg . (2) Uniform moment assumptions hold on K. For all relevant indices, sup Eϑ (∂ij log pϑ (X))2 < ∞, sup Eϑ |∂ijk log pϑ (X)| < ∞. ϑ∈K
ϑ∈K
(3) For some δ > 0, uniform (4 + δ)-moment assumptions hold on K. In particular, h i sup Eϑ |∂i log pϑ (X)|4+δ < ∞, ϑ∈K
and, with qij (x; ϑ) := ∂ij log pϑ (x) − Γkij (ϑ) ∂k log pϑ (x), h i 4+δ sup Eϑ qij (X; ϑ) − Eϑ [qij (X; ϑ)] < ∞. ϑ∈K
(4) The estimator θ̂n is a score-root estimator, so that U (θ̂n ) = 0. (5) The estimator θ̂n admits a second-order stochastic expansion uniformly for θ ∈ K, 1 1 ∆n := θ̂n − θ = √ An (θ) + Bn (θ) + op (n−1 ), n n with An (θ) = Op (1) and Bn (θ) = Op (1), uniformly on K. (6) The bias is negligible, in the sense that Biasθ (θ̂n ) = o(n−1 ) uniformly for θ ∈ K. (7) The score one-form admits the covariant Taylor expansion uniformly for θ ∈ K, 1 0 = Ui (θ) + (∇2 ℓ)ij (θ)∆jn + (∇3 ℓ)ijk (θ)∆jn ∆kn + ri,n (θ, ∆n ), 2 with ri,n (θ, ∆n ) = op (n−1/2 ). (8) The stronger L4+δ -identification of An holds uniformly on K, Ui (θ) Aan (θ) − g ai (θ) √ −→ 0 in L4+δ (Pθ ), n together with Ui (θ) Aan (θ) = g ai (θ) √ + op (n−1/2 ). n (9) The remainder and moment bounds satisfy Eθ [∥Rn (θ)∥2 ] = o(n−3 ),
sup sup Eθ [∥An ∥4 + ∥Bn ∥4 ] < ∞.
θ∈K n≥1
(10) The mixed second-order term is negligible, namely −1/2 Eθ [An Bn⊤ + Bn A⊤ ). n ] = o(n
(11) The replacement hypotheses for second moments hold. See the proof below for the precise statement involving Bi,n , Hij,n , and Kijk,n .
28
MALIK AMIR AND SOURANGSHU GHOSH
8.3. Proof of the main theorem. We begin with a definition and a proposition. Definition 8.1 (Big-O in probability). Let (Xn )n≥1 be a sequence of real-valued random variables and let (an )n≥1 be a sequence of positive real numbers. We write Xn = Op (an ) if the sequence Xn /an n≥1 is bounded in probability, i.e. |Xn | ∀ε > 0 ∃M < ∞ such that sup P > M ≤ ε. an n≥1 Proposition 8.2. Let K ⊂ Θreg be compact. Assume the regularity assumptions of Section 2, with Op understood in the sense of Definition 8.1. Assume moreover that, for all relevant indices, sup Eϑ (∂ij log pϑ (X))2 < ∞, sup Eϑ |∂ijk log pϑ (X)| < ∞. ϑ∈K
ϑ∈K
Then, for every θ ∈ K, Eθ [(∇2 ℓ)ij (θ)] = Eθ [Uij (θ)] = −n Iij (θ) = −n gij (θ). Moreover, uniformly for θ ∈ K, √ √ (∇2 ℓ)ij (θ) = −n gij (θ) + Op ( n), (∇3 ℓ)ijk (θ) = Op (n). Ui (θ) = Op ( n), Pn 2 Proof. We write ℓ(θ) = t=1 log pθ (Xt ), Ui = ∂i ℓ, Uij = ∂i ∂j ℓ, Uijk = ∂i ∂j ∂k ℓ, and (∇ ℓ)ij = Uij − Γkij Uk . R Expectation identities. Since pθ dµ = 1, differentiation under the integral gives Eθ [∂i log pθ (X)] = 0 and Eθ [∂ij log pθ (X)] = −Iij (θ). Summing over n i.i.d. observations yields Eθ [Uij ] = −nIij , and since Eθ [Uk ] = 0, also Eθ [(∇2 ℓ)ij ] = −ngij . Op -bounds. For the score, Varθ (Ui ) = nIii (θ) ≤ nCi for Ci := supK Iii < ∞. By Chebyshev, √ Ui = Op ( n). For the covariant Hessian, define qij (x; θ) := ∂ij log pθ (x) − Γkij (θ)∂k log pθ (x), so (∇2 ℓ)ij + ngij = Pn 2 p and boundedness of Γ on K give t=1 (qij (Xt ; θ) + gij ). The assumed L -bounds on ∂ij log √ θ 2 supK Varθ (qij ) < ∞. By Chebyshev, (∇ ℓ)ij + ngij = Op ( n). For the third derivative, by Markov’s inequality with the assumed first-moment bound on |∂ijk log pθ |, Uijk = Op (n). The relation (∇3 ℓ)ijk = Uijk + (terms involving Γ, ∂Γ, Ur , Uab ), where all correction terms are Op (n) or smaller, gives (∇3 ℓ)ijk = Op (n). □ Lemma 8.3 (Joint CLT and asymptotic Wick contractions). Fix a compact set K ⊂ Θreg and θ ∈ K. Define the centered covariant-Hessian fluctuation 1 Hab,n (θ) := √ (∇2 ℓ)ab (θ) + n gab (θ) . n Assume the (4 + δ)-moment bounds of Section 8.2 and the L4+δ -identification of An . Then (An (θ), Hn (θ)) converges jointly in distribution to a centered Gaussian vector, and for every polynomial P of total degree at most 4, Eθ [P (An , Hn )] → E[P (A, H)]. In particular, the asymptotic Wick identities hold k m m k Eθ [Hik,n Akn Hjm,n Am n ] = Eθ [Hik,n Hjm,n ] Eθ [An An ] + Eθ [Hik,n An ] Eθ [An Hjm,n ]
+ Eθ [Hik,n Akn ] Eθ [Hjm,n Am n ] + o(1), and similarly for the quartic and mixed cubic terms. Proof. The vector U√i (θ) , H (θ) is a normalized sum of i.i.d. centered random vectors with finite ab,n n (4 + δ)-moments. By the multivariate Lyapunov CLT (applied via Cramér–Wold), it converges jointly to a centered Gaussian. Since An = g −1 √Un + ρn with ∥ρn ∥L4+δ → 0, Slutsky’s theorem gives joint convergence of (An , Hn ).
ON HIGHER-ORDER GEOMETRIC REFINEMENTS OF CLASSICAL COVARIANCE ASYMPTOTICS
29
For moment convergence, Rosenthal’s inequality gives uniform L4+δ -bounds on the normalized sums, hence on (An , Hn ). For any degree-≤ 4 polynomial P , the family {P (An , Hn )} is uniformly integrable (by de la Vallée–Poussin), so distributional convergence implies convergence of expectations. The Wick identities follow from Isserlis’ theorem applied to the limiting Gaussian. □ Lemma 8.4 (Identification of Bn ). Under the perturbative score expansion and the next-order control Aan = g ai √Uni + op (n−1/2 ) 1 Bi,n (θ) := gij (θ)Bnj (θ) = Hij,n (θ)Ajn (θ) + Kijk,n (θ)Ajn (θ)Akn (θ) + op (1). 2 √ √ Proof. Multiply the score expansion by n and substitute Ui − n gij Ajn = op (1) (from the An identification). The result follows by rearrangement. □ p
Lemma 8.5 (From op (1) to o(1) in expectation). If Zn → − 0 and supn≥1 E[|Zn |1+η ] < ∞ for some η > 0, then E[|Zn |] → 0. Proof. The (1 + η)-moment bound implies uniform integrability (de la Vallée–Poussin criterion). Convergence in probability plus uniform integrability implies L1 -convergence (Vitali’s theorem). □ Corollary 8.6 (op (1) remainders inside second moments). Under the hypotheses of Lemma 8.5 applied to the remainders ρkn := Akn − g km U√mn , we have 1 E[Akn Aℓn ] = g km g ℓr E[Um Ur ] + o(1), n
1 E[Hij,n Akn ] = g km √ E[Hij,n Um ] + o(1). n
Ur Proof. Expand Akn Aℓn = (g km U√mn + ρkn )(g ℓr √ + ρℓn ). Each cross term involves ρn multiplied by n an Op (1) factor and Lemma 8.5 eliminates these in expectation. The mixed HA-identity follows similarly. □
Lemma 8.7 (Algebraic reduction in normal coordinates). Fix θ ∈ Θreg , and choose Riemannian normal coordinates at θ. Define (e),m (e),k (e),k (e),m Pij (θ) := g km Eθ [sik sjm ] − gik gjm + Γik Γjm + Γik Γjm 1 + κikℓ κjrs g kℓ g rs + g kr g ℓs + g ks g ℓr 4 1 (e),k (e),r (e),s + κjrs Γik g rs + Γik g ks + Γik g kr 2 1 (e),m (e),k (e),ℓ + κikℓ Γjm g kℓ + Γjm g mℓ + Γjm g mk . 2 ♯ ♯ Then Pij = 12 Rij + Sij + Dij , where Dij is the Hellinger discrepancy tensor.
Proof. In normal coordinates, gij = δij , Γkij = 0. From the Gauss decomposition, eij = IIij is purely (e)
normal. The tangency condition ⟨IIij , ek ⟩ = 0 forces Γijk = − 21 Tijk and κijk = 12 Tijk . Substituting Γ(e) = −κ into Pij and simplifying (using full symmetry of κijk ) gives the reduced form X 1 1 Pij = E[sik sjk ] − δij − κikl κj kl + κir r κjs s . 2 4 k
For the geometric side, the Gauss equation with convention gij = 4⟨ei , ej ⟩ gives Rikjl = 4(⟨IIij , IIkl ⟩− ♯ ♯ ⟨IIil , IIkj ⟩). Computing 12 Rij + Sij from the inner products 4⟨eab , ecd ⟩ = E[sab scd ] + 21 E[sab sc sd ] + 1 1 2 E[scd sa sb ] + 4 E[sa sb sc sd ], one obtains an expression involving second and fourth score moments. ♯ ♯ The Hellinger discrepancy Dij := Pij − ( 21 Rij + Sij ) collects the remaining terms. □
30
MALIK AMIR AND SOURANGSHU GHOSH
Proof of Theorem 5.1. Fix K ⊂ Θreg compact and θ ∈ K. All bounds below are uniform on K unless stated otherwise. Step 1: Score equation expansion. The score-root condition Ui (θ̂n ) = 0 and the covariant Taylor expansion (Assumption 7 of Section 8.2) give 1 0 = Ui (θ) + (∇2 ℓ)ij ∆jn + (∇3 ℓ)ijk ∆jn ∆kn + ri,n , 2 √ 1 −1/2 with ri,n = op (n ). Substituting ∆n = √n An + n1 Bn + op (n−1 ) and (∇2 ℓ)ij = −ngij + nHij,n , √ then dividing by n, yields 1 1 Ui 0 = √ − gij Ajn + √ (Hij,n Ajn − gij Bnj + Kijk,n Ajn Akn ) + op (n−1/2 ). 2 n n Step 2: Identification of An and Bn . At leading order: Aan = g ai √Uni + op (1). The L4+δ strengthening (Assumption 8) and the Lemma on Bn -identification give Bi,n = Hij,n Ajn + 21 Kijk,n Ajn Akn + op (1). Step 3: MSE expansion. From the stochastic expansion and remainder bounds (Assumption 9): 1 1 ⊤ −2 MSEθ (θ̂n ) = E[An A⊤ n ] + 2 E[Bn Bn ] + o(n ). n n Using E[Aan Abn ] = g ab + o(1) (from Corollary 8.6), the negligible bias (Assumption 6), and the negligible cross term (Assumption 10): 1 1 Covθ (θ̂n ) = I(θ)−1 + 2 Cn (θ) + o(n−2 ), Cn (θ) = E[Bn Bn⊤ ]. n n Step 4: Computation of E[Bi,n Bj,n ]. Substituting the Bn identification and applying the replacement hypotheses (Assumption 11) gives ei,n B ej,n ] + o(1), ei,n = Hij,n Aj + 1 κijk Aj Ak . E[Bi,n Bj,n ] = E[B B n n n 2 Expanding and applying the Wick contractions of Lemma 8.3 gives m k k m E[Bi,n Bj,n ] = E[Hik,n Hjm,n ]E[Akn Am n ] + E[Hik,n An ]E[An Hjm,n ] + E[Hik,n An ]E[Hjm,n An ] 1 + κikℓ κjrs Wick contraction of Akn Aℓn Arn Asn 4 1 + κjrs Wick contraction of Hik,n Akn Arn Asn 2 1 k ℓ + κikℓ Wick contraction of Hjm,n Am n An An + o(1). 2 Step 5: Evaluation of second moments. In normal coordinates at θ, the basic second moments are:
E[Akn Aℓn ] = g kl + o(1), (e),ℓ
E[Hik,n Aℓn ] = Γik
+ o(1),
E[Hik,n Hjl,n ] = E[sik sjl ] − gik gjl . These follow from Corollary 8.6 and direct computation of the i.i.d. sums. Step 6: Assembly and geometric reduction. Inserting the evaluated moments into Step 4 yields precisely the expression Pij (θ) defined in Lemma 8.7. By that lemma, 1 ♯ ♯ Pij = Rij + Sij + Dij . 2 In the inverse-metric frame, C(θ) = I(θ)−1 P (θ)I(θ)−1 . P ♯ Step 7: S ♯ ⪰ 0. For any v, v i v j Sij = k ∥v i IIik ∥2 ≥ 0 in normal coordinates. □
ON HIGHER-ORDER GEOMETRIC REFINEMENTS OF CLASSICAL COVARIANCE ASYMPTOTICS
31
References 1. Shun-ichi Amari, Information geometry and its applications, Applied Mathematical Sciences, vol. 194, Springer, Tokyo, 2016. 2. Shun-ichi Amari and Hiroshi Nagaoka, Methods of information geometry, Translations of Mathematical Monographs, vol. 191, American Mathematical Society and Oxford University Press, Providence, RI and Oxford, 2000, Translated by Daishi Harada. 3. Nihat Ay, Jürgen Jost, Hông Vân Lê, and Lorenz Schwachhöfer, The intrinsic geometry of statistical models, Information Geometry, Springer, Cham, 2017, pp. 185–239. 4. , Parametrized measure models, Information Geometry, Springer, Cham, 2017, pp. 121–184. 5. Ole E. Barndorff-Nielsen and David R. Cox, Inference and asymptotics, Chapman and Hall, London, 1994. 6. Douglas M. Bates and Donald G. Watts, Relative curvature measures of nonlinearity, Journal of the Royal Statistical Society. Series B (Methodological) 42 (1980), no. 1, 1–16. 7. A. Bhattacharyya, On some analogues of the amount of information and their use in statistical estimation, Sankhyā 8 (1946), 1–14. 8. Peter J. Bickel, Chris A. J. Klaassen, Ya’acov Ritov, and Jon A. Wellner, Efficient and adaptive estimation for semiparametric models, Johns Hopkins University Press, Baltimore, MD, 1993. 9. L. L. Campbell, An extended čencov characterization of the information metric, Proceedings of the American Mathematical Society 98 (1986), no. 1, 135–141. 10. N. N. Cencov, Statistical decision rules and optimal inference, Translations of Mathematical Monographs, vol. 53, American Mathematical Society, Providence, RI, 1982, Translated by the Israel Program for Scientific Translations. 11. Harald Cramér, Mathematical methods of statistics, Princeton Landmarks in Mathematics and Physics, Princeton University Press, Princeton, NJ, 1999, Originally published in 1946. 12. Bradley Efron, Defining the curvature of a statistical problem (with applications to second order efficiency), The Annals of Statistics 3 (1975), no. 6, 1189–1242. 13. Shinto Eguchi, Second order efficiency of minimum contrast estimators in a curved exponential family, The Annals of Statistics 11 (1983), no. 3, 793–803. 14. Robert E. Kass and Paul W. Vos, Geometrical foundations of asymptotic inference, John Wiley & Sons, New York, 1997. 15. Steven M. Kay, Fundamentals of statistical signal processing. vol. I: Estimation theory, Prentice Hall, Englewood Cliffs, NJ, 1993. 16. Steffen L. Lauritzen, Statistical manifolds, Differential Geometry in Statistical Inference, IMS Lecture Notes– Monograph Series, vol. 10, Institute of Mathematical Statistics, Hayward, CA, 1987, pp. 163–216. 17. Erich L. Lehmann and George Casella, Theory of point estimation, 2nd ed., Springer, New York, 1998. 18. Paul Marriott, On the local geometry of mixture models, Biometrika 89 (2002), no. 1, 77–93. 19. C. Radhakrishna Rao, Information and the accuracy attainable in the estimation of statistical parameters, Bulletin of the Calcutta Mathematical Society 37 (1945), no. 3, 81–91. 20. Aad W. van der Vaart, Asymptotic statistics, Cambridge Series in Statistical and Probabilistic Mathematics, vol. 3, Cambridge University Press, Cambridge, 1998. 21. Sumio Watanabe, Algebraic geometry and statistical learning theory, Cambridge Monographs on Applied and Computational Mathematics, vol. 25, Cambridge University Press, Cambridge, 2009. centre de recherche du chu de l’université de montréal centre de recherche du chu sainte-justine société québécoise de l’intelligence artificielle en médecine Email address: [email protected] department of civil engineering, indian institute of science bangalore Email address: [email protected]