TAP ACCURACY B ELOW THE F LUCTUATION S CALE AND U NIVERSAL P OSTERIOR G EOMETRY IN S PHERICAL L INEAR M ODELS
arXiv:2609.20577v1 [stat.ML] 17 Sep 2026
A P REPRINT Jingbo Liu Department of Statistics University of Illinois Urbana-Champaign Champaign, IL 61820 [email protected]
Zhiyuan Yu Department of Statistics University of Illinois Urbana-Champaign Champaign, IL 61820 [email protected]
September 18, 2026
A BSTRACT We study the Bayes-optimal spherical linear model as the ambient dimension p and sample size grow proportionally, under a quantitative Marchenko–Pastur spectral-regularity condition on the design. This condition is satisfied by normalized i.i.d. designs with standardized entries of finite fourth moment, but does not require entrywise independence or impose conditions on the singular vectors. Under this condition, we prove a quantitative all-temperature TAP approximation and characterize the posterior geometry. For the natural finite-aspect-ratio TAP functional, the normalized spherical free energy and the TAP optimum differ by OP (p−1 ). Each is within OP (p−1/2 ) of its explicit deterministic equivalent, and this fluctuation scale is sharp. Uniformly over all global TAP maximizers, the normalized squared Euclidean distance to the spherical posterior mean is OP (p−1 ). We also prove that the posterior mass outside a data-dependent band determined by the ridge estimator has sharp exponential order. More precisely, uniformly over sufficiently small band widths ε, the logarithm of this mass is at most −cpε2 + OP (1). For every fixed geometrically admissible width, a spherical-cap construction gives a matching exponential-order lower bound on this mass. For every deterministic sequence of widths εp ≫ p−1/2 , the corresponding bands capture asymptotically all posterior mass. 2020 Mathematics Subject Classification. Primary 60B20; Secondary 62C10, 82B44. Keywords Bayesian linear regression · spherical prior · TAP free energy · posterior-mean approximation · posterior concentration · random matrix theory · design universality
1
Introduction
Random linear models form a basic framework for high-dimensional regression, signal recovery, and related inverse problems. In the Bayesian formulation, the posterior mean is the Bayes estimator under squared-error loss, while the posterior covariance quantifies the remaining uncertainty. The corresponding normalizing constant determines the free energy. Obtaining sharp asymptotic characterizations of these quantities when the sample size and ambient dimension grow proportionally is a central problem in statistics, information theory, and statistical physics Banerjee et al. (2026); Barbier et al. (2020); Fan et al. (2021); Reeves and Pfister (2019). Even when the design matrix has i.i.d. entries, the posterior generally exhibits dependence across the signal coordinates. For Bayesian linear models with product priors, write Π(· | X, Y ) for the posterior and let Q be a tractable variational class. A standard variational approach considers the optimization problem inf KL q∥Π(· | X, Y ) , q∈Q
TAP accuracy and universal posterior geometry
A P REPRINT
where KL denotes Kullback-Leibler divergence Blei et al. (2017); Wainwright and Jordan (2008). The naive meanfield approximation restricts Q to the fully factorized family ( ) p Y QMF := q(dβ) = qi (dβi ) . i=1
In proportional random designs, this approximation can yield inaccurate estimates of the log-normalizing constant and posterior mean. It can also underestimate marginal uncertainty Celentano et al. (2023); Qiu (2023). These failures motivate variational functionals that retain the leading effect of the dependence induced by the design matrix. The Bethe and Thouless-Anderson-Palmer (TAP) free energies provide such corrections. The TAP approach originated in the Sherrington-Kirkpatrick model Mézard et al. (1987); Thouless et al. (1977). Plefka derived the TAP equations from a second-order expansion of the Gibbs potential in the interaction strength Plefka (1982). This expansion concerns the construction of the TAP functional, rather than the finite-size approximation error studied here. Rigorous developments include iterative constructions of solutions to these equations under the de Almeida-Thouless stability condition Bolthausen (2014). TAP free energies also provide variational representations of spin-glass free energies Belius (2022); Chen et al. (2023); Fan et al. (2021); Subag (2023). Variational free energies are closely connected to message-passing algorithms. Under the usual positivity and consistency conditions, stationary points of the Bethe free energy correspond to fixed points of belief propagation Wainwright and Jordan (2008); Yedidia et al. (2001, 2003). In dense random models, TAP functionals play an analogous variational role, while approximate message passing provides an algorithmic counterpart Fan et al. (2021); Krzakala et al. (2014). In compressed sensing, Krzakala et al. made this connection explicit by formulating naive mean-field and Bethe free energies and relating the stationary equations of the Bethe functional to approximate message passing Krzakala et al. (2014). More generally, state evolution describes the iterates of approximate message passing in the large-system limit Bayati and Montanari (2011); Krzakala et al. (2014). Rigorous TAP results have also been established for specific high-dimensional inference models. Fan, Mei, and Montanari analyzed Z2 synchronization. In a sufficiently strong-signal regime, the rank-one matrices formed from low-energy TAP critical points approximate the posterior mean of the signal outer product, the sign-invariant estimand in that model Fan et al. (2021). For Bayesian linear models with compactly supported product priors, Celentano et al. proved, under a nondegeneracy condition, the existence of a local TAP minimizer that approximates the posterior marginals. They also established local convexity and algorithmic guarantees in suitable parameter regimes. In highnoise or sufficiently low-aspect-ratio regimes, they further proved global convexity Celentano et al. (2023). For the Gaussian-design spherical linear model, Qiu and Sen proved a TAP representation of the free energy in a highnoise, or high-temperature, regime Qiu and Sen (2023). Their proof analyzes the posterior on spherical caps but does not provide a quantitative finite-size rate. Together, these works suggest that an appropriately defined TAP functional can provide the correct variational description of proportional high-dimensional Bayesian models. However, they do not establish an all-temperature global variational formula for the spherical linear model. At low temperatures, a local TAP optimizer or an approximate-message-passing fixed point is not necessarily globally optimal. Since the TAP functional can be nonconcave, even for Gaussian and spherical priors, the possibility of other extrema attaining larger values cannot be easily ruled out Yu and Liu (2026). Our earlier conference paper resolved this global-optimality problem for the spherical model with a Gaussian design by constructing a concave ridge surrogate that dominates the TAP functional Yu and Liu (2026). It bounded the distance of global and nearglobal TAP maximizers from the ridge estimator in terms of their optimality gap. The same paper also gave an explicit example of TAP nonconcavity for the spherical prior. The present paper adopts substantially improved proofs than Yu and Liu (2026) that further establish the following results: First, we show that the gap between the spherical free energy and the TAP optimum is smaller than the sample-fluctuations of these quantities by a factor of p−1/2 ; previously a similar sharp approximation of the TAP free energy for the Ising perceptron was conjectured by Ding and Sun (Ding and Sun, 2025, (1.12)). Second, we prove a sharp convergence of the global TAP maximizers to the spherical posterior mean. Third, we prove a concentration of the posterior on a spherical band, clarifying the geometric picture. 1.1
Problem setup and notation
For vectors, ∥·∥ := ∥·∥2 denotes the Euclidean norm and ⟨·, ·⟩ denotes the Euclidean inner product. For matrices, ∥·∥op and ∥·∥F denote the operator and Frobenius norms respectively. We use Tr(·) for the trace and Ik for the k ×k identity matrix. For symmetric matrices, A ⪯ B means that B − A is positive semidefinite. We write x+ := max{x, 0}. The symbols c and C denote positive constants whose values may change from line to line. For deterministic sequences up and rp > 0, we write up = O(rp ) if |up |/rp is bounded for all sufficiently large p, and up = o(rp ) if up /rp → 0. For positive sequences up and rp , we write up ≍ rp if both up = O(rp ) and rp = O(up ). 2
TAP accuracy and universal posterior geometry
A P REPRINT
All probabilistic order notation below is with respect to the joint law of the planted model specified next. For random variables Up , we write Up = OP (rp ) if Up /rp is tight. Equivalently, for every η > 0, there exist finite constants Mη and pη such that P |Up | > Mη rp ≤ η for all p ≥ pη . P
d
P
We write Up = oP (rp ) if Up /rp − → 0. The symbols − → and − → denote convergence in probability and convergence in distribution respectively. When auxiliary random variables are introduced, the underlying probability law is enlarged to include their explicitly specified joint distribution. We consider a Bayes-optimal linear model with a uniform spherical prior. This prior models a rotationally invariant signal with fixed energy, ∥β∥2 = p. It also provides a natural non-product setting for studying TAP approximations: the fixed-norm constraint induces dependence across the coordinates, while rotational symmetry permits explicit geometric and random-matrix analysis. Consequently, the fully factorized KL formulation above does not apply directly to the resulting posterior. The appropriate variational object is therefore a spherical TAP functional. For each p, let n := np and set αp := n/p. We observe Y = Xβ ⋆ + W, √ where X = Xp ∈ Rn×p , β ⋆ is uniformly distributed on S p−1 ( p), and W ∼ N (0, ∆In ) with ∆ > 0. The matrix X, ⋆ the signal β , and the noise W are mutually independent. Throughout, we work in the proportional regime and assume that, for some fixed α ∈ (0, ∞), |αp − α| = O p−1/2 .
(1)
This mild quantitative condition is satisfied, for example, when n = ⌊αp⌋, in which case |αp − α| = O(p−1 ). Let µMP,αp denote the Marchenko–Pastur (MP) law on [0, ∞) corresponding to the aspect ratio p/n = 1/αp . We impose the following quantitative spectral condition on the design matrix. Assumption 1 (Quantitative Marchenko–Pastur spectral regularity). There exists a deterministic constant CX < ∞ such that P (∥X∥op ≤ CX ) −→ 1. Moreover, for every fixed δ > 0, Z 1 1 x log det Ip + X ⊤ X = log 1 + dµMP,αp (x) + OP (p−1 ), p δ δ Z δ δ Tr(X ⊤ X + δIp )−1 = dµMP,αp (x) + OP (p−1/2 ). p x+δ Assumption 1 is verified for normalized i.i.d. designs with finite fourth moment in Lemma 7. Its proof in Appendix A uses the Bai–Yin spectral-edge theorem and Najim and Yao’s linear-spectral-statistic estimates. Apart from the independence condition stated above, Assumption 1 concerns only the singular values of X and allows both random and deterministic design sequences. This assumption is stronger than weak convergence of the empirical spectral distribution. Even when combined with a bounded spectral edge, weak convergence gives only an oP (1) normalized error for the two displayed statistics, without a quantitative rate. This is insufficient for the conclusions below the fluctuation scale. The quantitative estimates displayed above are needed in the proofs. √ Let π denote the uniform measure on S p−1 ( p) and set Z 1 2 ZpS := exp − ∥Y − Xβ∥ dπ(β). (2) √ 2∆ S p−1 ( p) The spherical posterior is 1 1 2 ΠS (dβ | X, Y ) := S exp − ∥Y − Xβ∥ dπ(β). Zp 2∆
(3)
The normalized spherical free energy is FpS :=
1 log ZpS . p 3
(4)
TAP accuracy and universal posterior geometry
We study this free energy, the posterior mean, and the geometry of the posterior measure. √ For ∥a∥ < p, define the spherical TAP functional 1 αp 1 − ∥a∥2 /p 1 ∥a∥2 2 fTAP (a) := − ∥Y − Xa∥ − log 1 + + log 1 − , 2∆p 2 αp ∆ 2 p √ and set fTAP (a) := −∞ for ∥a∥ ≥ p. We further define Mp := arg max fTAP (a). √
FTAP := max √ fTAP (a), ∥a∥≤ p
A P REPRINT
(5)
(6)
∥a∥≤ p
√ For every fixed realization of (X, Y ), all terms other than 12 log(1 − ∥a∥2 /p) √ remain bounded as ∥a∥ ↑ p, whereas this logarithmic term tends to −∞. Hence, fTAP (a) → −∞ as ∥a∥ ↑ p. It follows that Mp is nonempty and compact, and the measurable maximum theorem ensures that the suprema over Mp used below are measurable. 1.2
Contributions and main results P
For the Gaussian-design spherical model, Qiu and Sen proved FpS − FTAP − → 0 in a high-noise regime and left its extension to every fixed ∆ > 0 as an open problem Qiu and Sen (2023). Our earlier conference paper extended this convergence to every fixed ∆ > 0 and localized global and near-global TAP maximizers around the ridge estimator Yu and Liu (2026). The present work strengthens these results in four directions. • TAP accuracy below the fluctuation scale. For the finite-aspect-ratio TAP functional, we prove |FpS − FTAP | = OP (p−1 ), whereas the deviations of FpS and FTAP from their common deterministic equivalent occur on the sharp scale p−1/2 . Equivalently, the unnormalized TAP approximation error is OP (1), com√ pared with individual sample fluctuations on the scale p. This gives a samplewise comparison for the global spherical TAP optimum below the scale of its individual fluctuations. Ding and Sun conjectured an O(1)-accurate samplewise TAP description for the Ising perceptron (Ding and Sun, 2025, (1.12)), while prior rigorous TAP representations generally identify the leading extensive term, with error oP (p) rather than OP (1) Belius (2022); Chen et al. (2023); Qiu and Sen (2023); Yu and Liu (2026). Here, we establish an analogous O(1)-accurate TAP description for the spherical regression problem, leveraging a common sample-dependent ridge surrogate that captures the leading fluctuations. • Uniform identification of global TAP maximizers with the posterior mean. We prove that the normalized squared Euclidean distance between the exact spherical posterior mean and any global TAP maximizer is OP (p−1 ), uniformly over all such maximizers. This strengthens the earlier ridge localization result by quantitatively identifying the global TAP maximizers with the Bayesian posterior mean itself. • Sharp exponential-order posterior geometry. Uniformly over sufficiently small band widths ε, the log posterior mass outside a data-dependent ridge band is at most −cpε2 + OP (1). For every fixed geometrically admissible width, a spherical-cap construction gives a matching exponential-order lower bound on this posterior mass. Thus the order-p exponential speed is sharp for every fixed admissible width, while every sequence εp ≫ p−1/2 yields posterior contraction. • Spectral universality beyond i.i.d. designs. The three results above hold for every random or deterministic design sequence satisfying Assumption 1. In particular, the proofs require neither entrywise independence nor conditions on the singular vectors. This extension does not rely on a new invariance principle. Instead, the reduction to the common Gaussian-prior ridge surrogate uses a high-probability spectral-edge bound and two quantitative Marchenko–Pastur linear-spectral-statistic estimates. Normalized i.i.d. designs with standardized entries of finite fourth moment provide a concrete example, as follows from established samplecovariance results Bai and Silverstein (2010); Najim and Yao (2016). The key point is that these spectral inputs suffice for the quantitative TAP and posterior-geometric conclusions obtained here. Weak convergence of the empirical spectrum alone is not sufficient for the stated rates. Related universality results for messagepassing algorithms and regularized regression are obtained by different methods Bayati et al. (2015); Han and Shen (2023). The free-energy proof combines radial conditioning and ridge domination. For the same observed data, exact radial conditioning and local density estimates show that the spherical and Gaussian-prior free energies differ by OP (p−1 ). Building on our earlier conference paper Yu and Liu (2026), we bound the TAP functional by a strictly concave quadratic ridge functional. The resulting nonnegative gap vanishes quadratically at a deterministic reference value of the normalized squared norm. Consequently, OP (p−1/2 ) deviations of this quantity for the ridge estimator produce an 4
TAP accuracy and universal posterior geometry
A P REPRINT
OP (p−1 ) error. Gaussian integration and the log-determinant estimate in Assumption 1 show that the Gaussian free energy also differs from the same ridge optimum by OP (p−1 ). The matching fluctuation lower bounds follow from an anti-concentration estimate for the Gaussian quadratic polynomial appearing in the ridge optimum and the OP (p−1 ) comparisons with Rp . The posterior-mean approximation and ridge-band results do not rely on the TAP free-energy convergence theorem. For the posterior-mean result, Gaussian conditioning first shows that the spherical and Gaussian posterior means are within OP (1) of each other. The entropy-gap comparison and the quadratic curvature of the ridge surrogate then place every global TAP maximizer within OP (1) of the Gaussian posterior mean. For the ridge-band result, the conditionedGaussian representation and a linear exponential tilt yield the uniform upper bound, while a spherical-cap construction gives the matching exponential-order lower bound. Here and throughout, “all-temperature” means that the results hold for every fixed ∆ > 0. Constants may depend on ∆, and no uniformity as ∆ ↓ 0 is claimed. Retaining αp in (5) is essential for the below-fluctuation-scale comparison in Theorem 2. Replacing αp by α changes the functional by at most C|αp − α|, uniformly over the domain on which it is finite. This bound is only of order p−1/2 under (1). To state the deterministic equivalent, for γ > 0 define p 1 − γ − γ∆ + (γ∆ + γ − 1)2 + 4γ∆ E(γ, ∆) := , 2 γ 1 E(γ, ∆) 1 γ + log E(γ, ∆). ϕ(γ, ∆) := − + 1 − E(γ, ∆) − log 1 + 2 2 2 γ∆ 2 These quantities can be shown to equal the fixed point of the replica symmetric free energy and the corresponding free energy (Yu and Liu, 2026, (45)). We abbreviate Ep := E(αp , ∆),
E∆ := E(α, ∆).
(7)
The first main result quantifies the spherical TAP free-energy formula and shows that the p−1/2 fluctuation scale around the finite-aspect-ratio deterministic equivalent is sharp. Theorem 2 (TAP accuracy below the fluctuation scale). Suppose that Assumption 1 holds, αp = n/p satisfies (1), and ∆ > 0 is fixed. Then FpS − FTAP = OP (p−1 ),
(8)
FpS = ϕ(αp , ∆) + OP (p−1/2 ), −1/2
FTAP = ϕ(αp , ∆) + OP (p
).
Moreover, the p−1/2 rate is sharp: there exist constants c, η > 0, depending only on (α, ∆, CX ), such that c S lim inf P Fp − ϕ(αp , ∆) ≥ √ ≥ η, p→∞ p c lim inf P |FTAP − ϕ(αp , ∆)| ≥ √ ≥ η. p→∞ p
(9) (10)
(11) (12)
Consequently, neither centered quantity is oP (p−1/2 ). The bounds (9) and (10) remain valid with ϕ(αp , ∆) replaced by ϕ(α, ∆). No matching lower bound is asserted for the direct comparison |FpS − FTAP |. The second main result connects the TAP variational problem with the posterior mean. Theorem 3 (Uniform approximation of the posterior mean by global TAP maximizers). Under the assumptions of Theorem 2, let Z mS := β ΠS (dβ | X, Y ) be the spherical posterior mean. Then 1 ∗ ∥a − mS ∥22 = OP (p−1 ). a∗ ∈Mp p sup
In particular, the same estimate holds for every measurable choice a∗ ∈ Mp . 5
TAP accuracy and universal posterior geometry
A P REPRINT
The final result concerns the geometry of the spherical posterior. It gives sharp exponential-order concentration on a ridge band determined by the ridge estimator. This provides an all-temperature counterpart to the spherical-cap localization of Qiu and Sen Qiu and Sen (2023). Set S := X ⊤ X,
A∆ := S + ∆Ip ,
⊤ a∆ := A−1 ∆ X Y.
(13)
For ε > 0, define Bε (a∆ ) :=
β∈S
1 ( p) : ⟨β − a∆ , a∆ ⟩ ≤ ε . p
p−1 √
(14)
Theorem 4 (Sharp exponential-order concentration on ridge bands). Under the assumptions of Theorem 2, there exist constants c, ε0 > 0, depending only on (α, ∆, CX ), and a nonnegative random sequence Rband = OP (1) such that, p simultaneously for all 0 < ε ≤ ε0 , log ΠS (Bε (a∆ )c | X, Y ) ≤ −cpε2 + Rband . p √ Consequently, for every deterministic sequence εp ∈ (0, ε0 ] satisfying p εp → ∞, P ΠS Bεp (a∆ ) X, Y − → 1.
(15)
(16)
√ Set q⋆ := 1 − E∆ ∈ (0, 1). For every fixed 0 < ε < q⋆ + q⋆ , there are constants 0 < cε ≤ Cε < ∞, depending only on (ε, α, ∆, CX ), such that 1 P cε ≤ − log ΠS (Bε (a∆ )c | X, Y ) ≤ Cε −→ 1. (17) p The identity ∥β∥2 = p gives the equivalent representation ∥β − a∆ ∥2 ∥a∆ ∥2 p−1 √ ≤ 2ε . − 1− Bε (a∆ ) = β ∈ S ( p) : p p By (38) and (20), the center 1 − ∥a∆ ∥2 /p converges in probability to E∆ > 0. Thus (16) describes posterior concentration in a shell whose normalized squared radius around the ridge estimator converges to E∆ > 0, rather than contraction to the ridge estimator itself. 1.3
Future directions
It remains open whether the direct OP (p−1 ) TAP error is sharp. The limiting behavior of p(FpS − FTAP ) is also unknown. Establishing such a limit would require stronger assumptions than the order bounds in Assumption 1. For the ridge band, Theorem 4 gives contraction when εp ≫ p−1/2 but determines neither the critical shrinking width nor the exact fixed-width exponential rate. Conditional fluctuation limits for the ridge coordinate, together with joint local large-deviation estimates for this coordinate and the Gaussian posterior radius, may address these questions. Extensions to non-Marchenko–Pastur spectra, nonspherical priors, and nonquadratic likelihoods raise further questions. Spectrum-dependent TAP predictions Maillard et al. (2019) and rigorous high-temperature results for rotationally invariant designs with i.i.d. signal priors Li et al. (2024) provide starting points. The challenge is to obtain analogous quantitative all-temperature conclusions using alternatives to the Gaussian-conditioning and ridge-surrogate arguments where these methods no longer apply.
2
TAP Accuracy Below the Fluctuation Scale
This section proves (8)–(10) in Theorem 2; the sharpness bounds (11)–(12) are proved in Appendix B. We first compare the spherical-prior and Gaussian-prior partition functions for the same data through an exact radial identity. We then compare the Gaussian-prior free energy and the TAP maximum with the maximum of a common quadratic ridge surrogate. We work throughout under the model, notation, and assumptions of Theorem 2, including the aspect-ratio condition (1). In particular, FpS , fTAP , FTAP , and Mp are those defined in (4), (5), and (6). For γ > 0, define E(γ, ∆) L(γ, ∆) := − 1 − E(γ, ∆) + γ log 1 + − log E(γ, ∆). γ∆ 6
(18)
TAP accuracy and universal posterior geometry
A P REPRINT
The quantity E(γ, ∆) defined above is also characterized as the unique solution e ∈ (0, 1) of e=
γ∆ + e . γ∆ + e + γ
(19)
For fixed ∆ > 0, the functions E(·, ∆), L(·, ∆), and ϕ(·, ∆) are continuously differentiable on (0, ∞). Since αp → α > 0, both αp and α lie in [α/2, 2α] for all sufficiently large p. Their derivatives are bounded on this interval, so the mean value theorem and (1) give a deterministic constant C = C(α, ∆) < ∞ such that, for all sufficiently large p, |E(αp , ∆) − E(α, ∆)| + |L(αp , ∆) − L(α, ∆)| + |ϕ(αp , ∆) − ϕ(α, ∆)| ≤ C|αp − α| = O p−1/2 . 2.1
(20)
Gaussian-Spherical Free-Energy Comparison
To compare the two free energies, we introduce a Gaussian-prior reference measure while keeping the data (X, Y ) generated by the spherical model. Let γp := N (0, Ip ), and define the corresponding partition function, free energy, and posterior by Z 1 1 ZpG := ∥Y − Xβ∥2 dγp (β), exp − (21) FpG := log ZpG , 2∆ p Rp 1 1 ∥Y − Xβ∥2 dγp (β). (22) PG (dβ | X, Y ) := G exp − Zp 2∆ Lemma 5 (Mean squared radius of the Gaussian-prior posterior). Let −1 ⊤ ΣG := ∆A−1 X X ∆ = Ip + ∆
−1
.
(23)
For every fixed realization of (X, Y ), if Z is drawn from the posterior PG (· | X, Y ), then Z | X, Y ∼ N (a∆ , ΣG ),
(24)
and the Gaussian-prior posterior mean is mG = a∆ . Under the joint law of the planted spherical model, namely √ √ Y = Xβ ⋆ + ∆z, where β ⋆ is uniform on S p−1 ( p) and z ∼ N (0, In ), with X, β ⋆ , and z mutually independent, we furthermore have √ √ ∥a∆ ∥ = OP ( p), E[∥Z∥2 | X, Y ] = ∥a∆ ∥2 + Tr(ΣG ) = p + OP ( p). (25) Proof. Completing the square in the exponent of the Gaussian-prior posterior density establishes (24) for every fixed (X, Y ). For the probabilistic estimates, use the planted representation above and let B := Ip − ΣG . Then √ ⊤ a∆ = Bβ ⋆ + η, η := ∆A−1 ∆ X z. Conditional on X, the vector η is independent of β ⋆ and satisfies η ∼ N (0, ΣG − Σ2G ). Since E[η | X] = 0 and E[β ⋆ (β ⋆ )⊤ | X] = Ip , conditional independence implies E[∥a∆ ∥2 | X] = Tr(B 2 ) + Tr(ΣG − Σ2G ) = Tr(B). The squared norm can be written as ∥a∆ ∥2 = (β ⋆ )⊤ B 2 β ⋆ + 2(β ⋆ )⊤ Bη + ∥η∥2 . Since β ⋆ and η are independent given X and η is conditionally centered Gaussian, the three terms are pairwise uncorrelated given X, so Var(∥a∆ ∥2 | X) = Var (β ⋆ )⊤ B 2 β ⋆ | X + 4 Tr(B 3 ΣG ) + 2 Tr (ΣG − Σ2G )2 . The spherical quadratic-form variance satisfies Var (β ⋆ )⊤ B 2 β ⋆ | X =
2p p+2
7
Tr(B 4 ) −
(Tr(B 2 ))2 p
≤ 2p.
TAP accuracy and universal posterior geometry
A P REPRINT
Since B = Ip − ΣG and 0 ⪯ B, ΣG ⪯ Ip , the remaining two terms are bounded by 4p and 2p, respectively. Thus Var(∥a∆ ∥2 | X) ≤ 8p. By Chebyshev’s inequality, applied conditionally on X and then averaged over X, √ ∥a∆ ∥2 − Tr(B) = OP ( p). Since Tr(B) ≤ p, the norm bound in (25) follows. The mean-radius bound in the same equation follows from E[∥Z∥2 | X, Y ] − p = ∥a∆ ∥2 − Tr(B), using the Gaussian posterior formula and B = Ip − ΣG . The next lemma compares the spherical and Gaussian-prior free energies through an exact radial identity. The proof combines (25) with the local density bounds in Lemma 12, which is proved in Appendix C and is independent of the linear model. Lemma 6. Under the assumptions of Theorem 2, FpS − FpG = OP (p−1 ). √ Proof. For t > 0, let πt be the uniform probability measure on S p−1 ( t) and define Z 1 2 S dπt (β). Zp (t) := √ exp − 2∆ ∥Y − Xβ∥ S p−1 ( t) Let ρX,Y denote the continuous density of ∥Z∥2 conditional on (X, Y ), where Z is drawn from the Gaussian-prior posterior PG (· | X, Y ) = N (a∆ , ΣG ) in (24). Under the Gaussian prior γp = N (0, Ip ), the squared radius ∥β∥2 has a chi-square distribution with p degrees of freedom, whose density we denote by fχ2p . Moreover, conditional on ∥β∥2 = t, the distribution of β is πt . Reweighting by the likelihood and normalizing by ZpG therefore yields, for every t > 0, ρX,Y (t) = fχ2p (t)
ZpS (t) . ZpG
(26)
Since ZpS (p) = ZpS , (26) yields p(FpS − FpG ) = log ρX,Y (p) − log fχ2p (p).
(27)
2 ). It remains to estimate the two densities at p. By Lemma 5, the posterior satisfies (25). Set c− := ∆/(∆ + CX The operator-norm bound in Assumption 1, together with (25), implies that, for every η > 0, there exist deterministic constants K0 , M > 0 such that, for all sufficiently large p, the following bounds hold simultaneously with probability at least 1 − η: √ √ c− Ip ⪯ ΣG ⪯ Ip , ∥a∆ ∥ ≤ K0 p, ∥a∆ ∥2 + Tr(ΣG ) − p ≤ M p.
On this event, the local density bound (44), proved later in Lemma 12, applies with t = p and gives √ cM ≤ p ρX,Y (p) ≤ CM , where cM , CM > 0 depend only on (c− , 1, K0 , M ). Since η is arbitrary, √ log p ρX,Y (p) = OP (1). (28) √ Using Stirling’s formula Γ(x) = 2πxx−1/2 e−x (1 + O(x−1 )) at x = p/2, the chi-square density at its mean satisfies fχ2p (p) =
pp/2−1 e−p/2 1 + O(p−1 ) √ = . 4πp 2p/2 Γ(p/2)
Thus log fχ2p (p) = − 12 log p − 12 log(4π) + O(p−1 ). Together with (28), substitution into (27) gives p(FpS − FpG ) = OP (1), which proves the claim. 8
TAP accuracy and universal posterior geometry
2.2
A P REPRINT
A Common Quadratic Ridge Surrogate
The key idea of this subsection is to control the generally nonconcave TAP functional by the strictly concave quadratic ridge functional Qp defined below. Concavity of the TAP functional cannot be assumed: Appendix A of our earlier conference paper Yu and Liu (2026) gives an explicit example of nonconcavity already for the spherical prior. The pointwise decomposition of fTAP into Qp minus a nonnegative scalar entropy gap establishes Qp as a global upper bound, while the ridge estimator nearly saturates this bound. This provides the quantitative control of the global TAP optimum needed below. The following random-matrix estimates are the only inputs concerning the distribution of the design. The deterministic terms in (30) and (31) below retain the actual aspect ratio αp = n/p, rather than its limit α. In particular, retaining αp in (30) preserves the OP (p−1 ) remainder needed for the comparison below the fluctuation scale. Lemma 7 (Spectral estimates and verification for i.i.d. designs). Suppose that αp → α ∈ (0, ∞) and that ∆ > 0 is fixed. Under Assumption 1, P (∥X∥op ≤ CX ) −→ 1, (29) 1 1 log det Ip + X ⊤ X = L(αp , ∆) + OP (p−1 ), (30) p ∆ ∆ Tr(X ⊤ X + ∆Ip )−1 = Ep + OP (p−1/2 ). (31) p √ Moreover, Assumption 1 holds when Xij = ξij / n, where the ξij are i.i.d. copies of a fixed real random variable ξ satisfying Eξ = 0,
Eξ 2 = 1,
Eξ 4 < ∞,
and X is independent of (β ⋆ , W ). The proof is deferred to Appendix A. For the comparison that follows, define 1 Cp := − L(αp , ∆), 2 1 1 Qp (a) := − ∥Y − Xa∥2 − ∥a∥2 + Cp , 2∆p 2p Rp := sup Qp (a).
(32) (33) (34)
a∈Rp
By strict concavity, Qp is uniquely maximized by the ridge estimator a∆ defined in (13). The next lemma compares the Gaussian-prior free energy with Rp and evaluates Rp to the required accuracy. Lemma 8. Under the assumptions of Theorem 2, FpG − Rp = OP (p−1 ), Rp = ϕ(αp , ∆) + OP (p−1/2 ). Proof. Completing the square in the Gaussian integral and maximizing Qp yield, respectively, 1 ⊤ 1 1 G Fp = − log det Ip + X X − Y ⊤ (∆In + XX ⊤ )−1 Y, 2p ∆ 2p 1 ⊤ ⊤ −1 Rp = Cp − Y (∆In + XX ) Y. 2p Subtracting the two identities and applying (30), we obtain 1 1 FpG − Rp = − log det Ip + X ⊤ X − pL(αp , ∆) = OP (p−1 ). 2p ∆
(35)
(36)
To prove the ridge-centering estimate (37) below, set D∆ := ∆In + XX ⊤ . Since Y = Xβ ⋆ + W , taking conditional expectation given X yields −1 −1 −1 E[Y ⊤ D∆ Y | X] = Tr(X ⊤ D∆ X) + ∆ Tr(D∆ ) = n.
9
TAP accuracy and universal posterior geometry
A P REPRINT
−1 Set A := X ⊤ D∆ X. Expanding the quadratic form and using the independence of β ⋆ and W , the three centered terms are pairwise uncorrelated given X. Hence −1 Var(Y ⊤ D∆ Y | X) = Var((β ⋆ )⊤ Aβ ⋆ | X) −2 −2 + 4∆ Tr(X ⊤ D∆ X) + 2∆2 Tr(D∆ ).
Since the eigenvalues of A lie in [0, 1], the spherical quadratic-form variance formula bounds the first term by 2p. The inequalities 4∆λ/(∆ + λ)2 ≤ 1 and ∆2 /(∆ + λ)2 ≤ 1, for λ ≥ 0, bound the other two terms by p and 2n, respectively. Thus the conditional variance is at most 3p + 2n = O(p), uniformly in X. Conditional Chebyshev’s inequality, followed by averaging over X, therefore yields √ −1 Y ⊤ D∆ Y = n + OP ( p). Consequently, 1 αp Rp = − L(αp , ∆) − + OP (p−1/2 ) 2 2 = ϕ(αp , ∆) + OP (p−1/2 ),
(37)
where the last identity follows directly from the definitions of L and ϕ. We next record two estimates for the ridge estimator that will be used both in the TAP comparison below and in the posterior-geometry arguments. Lemma 9. Under the assumptions of Theorem 2, q∆ :=
1 ∥a∆ ∥2 = 1 − Ep + OP (p−1/2 ), p
1 ∥Y − Xa∆ ∥2 = ∆Ep + ∆(αp − 1) + OP (p−1/2 ). p
(38) (39)
Here Ep is defined in (7). The proof is deferred to Appendix A. To compare Qp with the TAP functional, we isolate their scalar entropy gap. For 0 ≤ q < 1, define q αp 1−q 1 hp (q) := − + Cp + log 1 + − log(1 − q). (40) 2 2 αp ∆ 2 √ For ∥a∥ < p, ∥a∥2 hp = Qp (a) − fTAP (a). (41) p The next lemma uses (41) and (38) to compare their global optima. Lemma 10 (Quantitative ridge approximation). Under the assumptions of Theorem 2, 0 ≤ Rp − FTAP = OP (p−1 ).
(42)
Proof. Let q0,p := 1 − Ep , where Ep is defined in (7). The fixed-point equation (19), together with (32), implies that hp (q0,p ) = 0. Moreover, αp 1 1 + h′p (q) = − − 2 2(αp ∆ + 1 − q) 2(1 − q) (q − q0,p )(1 − q + αp ∆/Ep ) = , 2(1 − q)(αp ∆ + 1 − q) where the second equality follows from (19). The factors other than q − q0,p are positive for 0 ≤ q < 1. Thus h′p (q) < 0 for q < q0,p and h′p (q) > 0 for q > q0,p , so hp (q) ≥ 0,
0 ≤ q < 1. 10
(43)
TAP accuracy and universal posterior geometry
By (41) and (43), fTAP (a) ≤ Qp (a) ≤ Rp for ∥a∥ <
√
A P REPRINT
p, and hence FTAP ≤ Rp .
By (38), q∆ = q0,p + OP (p−1/2 ). Since Ep → E∆ ∈ (0, 1), there exists a fixed δ > 0 such that the intervals [q0,p − δ, q0,p + δ] lie in one fixed compact subset of (0, 1) for all sufficiently large p. Since αp 1 − , h′′p (q) = 2 2(1 − q) 2(αp ∆ + 1 − q)2 the second derivatives are uniformly bounded on these intervals. The event {|q∆ − q0,p | ≤ δ} has probability tending to one. On this event, a∆ lies in the domain of fTAP , and its optimality for Qp gives FTAP ≥ fTAP (a∆ ) = Rp − hp (q∆ ). ′ Since hp (q0,p ) = hp (q0,p ) = 0, Taylor’s theorem therefore yields a deterministic constant C < ∞ such that, on the same event, 0 ≤ Rp − FTAP ≤ hp (q∆ ) ≤ C(q∆ − q0,p )2 . The quantity p(q∆ − q0,p )2 is bounded in probability, and the exceptional event has probability tending to zero. This proves (42). Remark 11. The sharpness bounds (11) and (12) concern the separate fluctuations of FpS and FTAP around ϕ(αp , ∆). They do not provide a matching lower bound for the direct approximation error |FpS − FTAP |, for which we prove only the upper bound (8). Proof of the upper bounds in Theorem 2. By Lemmas 6, 8, and 10, FpS − FTAP = (FpS − FpG ) + (FpG − Rp ) + (Rp − FTAP ) = OP (p−1 ). This proves (8). The ridge-centering estimate (37), together with the OP (p−1 ) comparisons in Lemmas 6, 8, and 10, gives (9) and (10). Finally, (20) gives the corresponding bounds around ϕ(α, ∆).
3
TAP Maximizers and the Spherical Posterior Mean
This section proves Theorem 3 by comparing the spherical posterior mean and all global TAP maximizers with the Gaussian posterior mean mG = a∆ . The argument uses the entropy-gap and ridge-approximation estimates from the preceding section, but not the free-energy results of Theorem 2. The comparison of the two posterior means requires controlling how the mean of a high-dimensional Gaussian distribution changes when its squared norm is fixed. Since conditioning on a fixed value of ∥Z∥22 is a probability-zero event, a precise continuous version of the conditional mean is needed. The following lemma defines this version and bounds its deviation from the original Gaussian mean when the covariance eigenvalues are uniformly bounded above and away from zero. The result is independent of the linear model and will be applied to the Gaussian posterior. Lemma 12 (Gaussian conditioning near the typical radius). Fix 0 < c− < c+ < ∞ and K < ∞, and let Z ∼ N (a, Σp ) in Rp satisfy √ c− Ip ⪯ Σp ⪯ c+ Ip , ∥a∥2 ≤ K p. Set µp := E∥Z∥22 = ∥a∥22 + Tr(Σp ). For rp > 0, the conditional mean is understood in the continuous coarea sense: Z zfZ (z) dHp−1 (z) ∥z∥=r p E Z | ∥Z∥22 = rp2 := Z , fZ (z) dHp−1 (z) ∥z∥=rp p−1
where fZ is the Gaussian density and H is surface measure. There exist ε0 , C > 0, depending only on (c− , c+ , K), such that, for all sufficiently large p and every rp > 0 satisfying |rp2 − µp | ≤ ε0 p, ! |rp2 − µp | 2 2 E Z | ∥Z∥2 = rp − a 2 ≤ C 1 + . √ p
11
TAP accuracy and universal posterior geometry
A P REPRINT
Moreover, if ρZ denotes the continuous density of ∥Z∥22 on (0, ∞), then for every fixed M < ∞ there exist 0 < cM < CM < ∞, depending only on (c− , c+ , K, M ), such that, for all sufficiently large p, cM CM (44) √ ≤ ρZ (t) ≤ √ p p √ whenever |t − µp | ≤ M p. There also exists Cup < ∞, depending only on (c− , c+ , K), such that, for all sufficiently large p, Cup sup ρZ (t) ≤ √ . (45) p t>0 The thresholds on p are uniform over a and Σp satisfying the displayed hypotheses. The proof is deferred to Appendix C. Returning to the linear model under the standing assumptions, including (1), we apply Lemma 12 to the Gaussian-prior posterior with rp2 = p. Conditioning this posterior on ∥Z∥22 = p yields the spherical-prior posterior. Lemma 5 and Assumption 1 verify the conditions required for the following comparison. Lemma 13. Let mS be as in Theorem 3, and let mG be the Gaussian-prior posterior mean for the same realization of (X, Y ). Then ∥mG − mS ∥2 = OP (1). Proof. Lemma 5 identifies the Gaussian-prior posterior as N (a∆ , ΣG ), with mG = a∆ . By polar disintegration, its 2 conditional law given ∥Z∥22 = p is the spherical posterior, since the Gaussian-prior factor e−∥Z∥2 /2 is constant on √ S p−1 ( p). Therefore, mS = E[Z | X, Y, ∥Z∥22 = p],
(46)
where the conditional expectation is understood in the continuous coarea sense defined in Lemma 12. Let µp denote the conditional mean squared radius in (25). This estimate, together with (29) and (23), implies that, for every δ > 0, there are deterministic constants c− > 0 and K < ∞ such that, for all sufficiently large p, the following event has probability at least 1 − δ: √ √ c− Ip ⪯ ΣG ⪯ Ip , ∥a∆ ∥2 ≤ K p, |µp − p| ≤ K p. √ For all sufficiently large p, K p ≤ ε0 p, so Lemma 12 applies with rp2 = p on this event. Together with (46), the lemma gives, on this event, |µp − p| ∥mS − mG ∥2 ≤ C 1 + √ ≤ C(1 + K). p Since δ > 0 is arbitrary, ∥mS − mG ∥2 = OP (1). The preceding lemma shows that replacing the Gaussian prior by the spherical prior changes the posterior mean by only OP (1). Recall from (6) that Mp is the set of all global maximizers of fTAP . It remains to compare these maximizers with the Gaussian posterior mean. The entropy-gap inequality bounds fTAP above by the strictly concave ridge surrogate Qp , whose unique maximizer is a∆ = mG . Comparing the value at a global TAP maximizer with the value at a∆ and using the quadratic curvature of Qp yields the uniform distance estimate below. Lemma 14. Under the assumptions of Theorem 3, 1 sup ∥a − mG ∥22 = OP (p−1 ). a∈Mp p Proof. By (41) and (43), fTAP ≤ Qp on its finite domain. The ridge functional (33) has unique maximizer a∆ , and, for every a ∈ Rp , 1 1 Qp (a∆ ) − Qp (a) = (a − a∆ )⊤ A∆ (a − a∆ ) ≥ ∥a − a∆ ∥22 . 2∆p 2p For every a ∈ Mp , the definition (6) gives Qp (a) ≥ fTAP (a) = FTAP . Consequently, (42) yields 1 sup ∥a − a∆ ∥22 ≤ 2(Rp − FTAP ) = OP (p−1 ). p a∈Mp Since mG = a∆ by (24), this proves the lemma. 12
TAP accuracy and universal posterior geometry
A P REPRINT
Proof of Theorem 3. The squared triangle inequality and Lemmas 14 and 13 yield 1 2 1 sup ∥a − mS ∥22 ≤ 2 sup ∥a − mG ∥22 + ∥mG − mS ∥22 p a∈Mp p a∈Mp p = OP (p−1 ).
4
Sharp Exponential-Order Concentration on Ridge Bands
This section proves Theorem 4 under the model assumptions, including (1). Recall the ridge estimator a∆ and the ridge band Bε (a∆ ), defined in (13) and (14), respectively. The upper bound (15) uses the exact representation of the spherical posterior as a Gaussian posterior conditioned on its squared radius. To prove the upper bound on the logarithmic rate in (17), we construct a spherical cap contained in the complement of the ridge band. Proof of Theorem 4. We condition throughout on (X, Y ) and abbreviate a := a∆ ,
Σ := ΣG .
By (24), the Gaussian-prior posterior is the law of Z ∼ N (a, Σ). The polar-disintegration argument used to derive (46) shows that the spherical posterior is the coarea conditional law of Z given ∥Z∥2 = p. Define L(Z) := ⟨Z − a, a⟩,
R(Z) := ∥Z∥2 ,
vp := a⊤ Σa.
Thus ΠS (Bε (a)c | X, Y ) = P ( |L(Z)| > pε | R(Z) = p, X, Y ) ,
(47)
where the conditional probability is understood in the coarea sense used in Lemma 12. For τ ∈ {−1, 1} and s ≥ 0, introduce the linear exponential tilt dPτ,s s2 vp , (z) := exp τ sL(z) − dP0 2
(48)
where P0 := N (a, Σ) is the Gaussian-prior posterior conditional on the fixed data (X, Y ). Completing the square gives Pτ,s = N (a + τ sΣa, Σ).
(49)
Let ρτ,s denote the density of R(Z) under Pτ,s , and write ρ0,0 for the corresponding density under P0 , always using the continuous versions. Writing f0 for the Lebesgue density of P0 , define, for r > 0, Z 1{τ L(z)>pε} f0 (z) dHp−1 (z). hτ,ε (r) := 2∥z∥ 2 ∥z∥ =r Here Hp−1 denotes (p − 1)-dimensional Hausdorff measure, which on the sphere {z : ∥z∥2 = r} is the usual unnormalized surface measure. The coarea formula identifies this as a density of the subprobability measure P0 (R(Z) ∈ dr, τ L(Z) > pε | X, Y ). The next step is a change-of-measure argument for binary hypothesis testing, conditional on (X, Y ). For s > 0, the event {τ L(Z) > pε} is the rejection region of the likelihood-ratio test of P0 against Pτ,s with threshold exp{spε − s2 vp /2}. We apply the corresponding pointwise density bound on the sphere before normalizing by its radial density. Let fτ,s be the Lebesgue density of Pτ,s . By (48), for every z ∈ Rp , s2 vp 1{τ L(z)>pε} f0 (z) = 1{τ L(z)>pε} exp −τ sL(z) + fτ,s (z) 2 s2 vp fτ,s (z), ≤ exp −spε + 2 where the inequality uses s ≥ 0. Dividing by 2∥z∥ and integrating over the sphere {z : ∥z∥2 = p} gives Z s2 vp fτ,s (z) hτ,ε (p) ≤ exp −spε + dHp−1 (z) 2 2 ∥z∥ =p 2∥z∥ s2 vp = exp −spε + ρτ,s (p). 2 13
TAP accuracy and universal posterior geometry
The last equality is the coarea formula for R(Z) under Pτ,s . Consequently, s2 vp ρτ,s (p) P ( τ L(Z) > pε | R(Z) = p, X, Y ) ≤ exp −spε + . 2 ρ0,0 (p)
A P REPRINT
(50)
We now make this bound uniform for small ε. By the high-probability operator-norm bound in Assumption 1, together with (23) and (38), there are deterministic constants c− , Ca , V > 0, depending only on (α, ∆, CX ), such that the event √ Epupper := {c− Ip ⪯ Σ ⪯ Ip , ∥a∥ ≤ Ca p, vp ≤ V p} 2 has probability tending to one. For example, one may take c− := ∆/[2(∆ + CX )], Ca := 2, and V := 4, since P
∥a∥2 /p − → q⋆ < 1 and Σ ⪯ I√p . Fix ε0 := 1. On this event, for each 0 < ε ≤ ε0 , set s := ε/V . The mean in (49) has norm at most (1 + ε0 /V )Ca p. Therefore, applying (45) uniformly over these choices of ε and τ ∈ {−1, 1} gives a deterministic Cup < ∞ such that Cup ρτ,s (p) ≤ √ , p
τ ∈ {−1, 1}.
Moreover, pε2 s2 vp ≤− . 2 2V Applying (50) for both signs and using (28), we obtain, simultaneously for 0 < ε ≤ ε0 , −spε +
log ΠS (Bε (a∆ )c | X, Y ) ≤ −
pε2 + Dp 2V
(51)
on Epupper , where √ Dp := log(2Cup ) − log ( p ρ0,0 (p)) = OP (1). Set 1 , 4V Rband := (Dp )+ + cpε20 1(Epupper )c . p c :=
The first term is OP (1), while the second is oP (1) because P((Epupper )c ) → 0. On Epupper , (51) implies (15). On the complementary event, the same inequality follows from log ΠS (Bε (a∆ )c | X, Y ) ≤ 0 and ε ≤ ε0 . Thus (15) holds simultaneously for all such ε and all sufficiently large p. Enlarging finitely many terms of Rband makes the bound p √ valid for every p without affecting its order in probability. If p εp → ∞, division by pε2p proves the concentration statement (16). To complete (17), we now prove the posterior-mass lower bound (55) below. Write qp := q∆ , with q∆ defined in (38). By (38), (20), and the definition of q⋆ , P
→ q⋆ ∈ (0, 1). qp − Fix 0 < ε < q⋆ +
√
(52)
q⋆ . Choose r ∈ (0, 1) such that √ q⋆ + r q⋆ > ε.
By continuity, there is a δ ∈ (0, q⋆ /2) such that √ q+r q >ε
whenever
|q − q⋆ | ≤ δ.
On the event {|qp − q⋆ | ≤ δ}, define a∆ up := , ∥a∆ ∥
Cp (r) :=
β∈S
⟨β, up ⟩ ( p) : √ ≤ −r . p
p−1 √
For every β ∈ Cp (r), 1 ∥a∆ ∥ ⟨β − a∆ , a∆ ⟩ = ⟨β, up ⟩ − qp p p √ ≤ −r qp − qp < −ε, 14
(53)
TAP accuracy and universal posterior geometry
A P REPRINT
where the last inequality follows from (53). Hence Cp (r) ⊆ Bε (a∆ )c .
(54)
Choose any fixed r′ ∈ (r, 1). Rotational invariance and the density of the first coordinate of a uniform point on the unit sphere give Γ(p/2) (1 − (r′ )2 )(p−3)/2 π Γ((p − 1)/2) ≥ exp{−Ccap p}
π Cp (r) ≥ (r′ − r) √
for all sufficiently large p, where Ccap < ∞ depends only on r and r′ . The last inequality follows from Γ(p/2)/Γ((p− √ 1)/2) ≍ p. Write HY (β) := (2∆)−1 ∥Y − Xβ∥2 . Let CX be the constant from Assumption 1, and choose a deterministic CW < ∞ such that the event √ Eplower := {∥X∥op ≤ CX , ∥W ∥ ≤ CW p, |qp − q⋆ | ≤ δ} P
has probability tending to one. Such a CW exists because ∥W ∥2 /p − → α∆. On this event, √ ∥Y ∥ ≤ (CX + CW ) p, √ and hence, uniformly over β ∈ S p−1 ( p), HY (β) ≤
(2CX + CW )2 p =: CH p. 2∆
Since HY ≥ 0, we also have ZpS ≤ 1. Therefore, on Eplower , (54) gives Z c ΠS (Bε (a∆ ) | X, Y ) ≥ e−HY (β) dπ(β) Cp (r)
≥ exp{−(CH + Ccap )p}.
(55)
Applying (15) at min{ε, ε0 } gives, with probability tending to one, 1 c − log ΠS (Bε (a∆ )c | X, Y ) ≥ min{ε, ε0 }2 , p 2
(56)
P
/p − → 0. When ε > ε0 , this uses Bε (a∆ )c ⊆ Bε0 (a∆ )c . Combining (56) with (55) proves (17). because Rband p
Acknowledgments The authors would like to thank Kevin Luo, Zhou Fan, and Subhabrata Sen for helpful discussions. J.L.’s research was supported in part by NSF Grant DMS-2515510.
Statement on AI use OpenAI’s GPT-6 Astra, accessed through Codex, assisted with improving the presentation of the proofs written by the authors. The authors independently verified all claims and take full responsibility.
A
Auxiliary Estimates for the TAP Free-Energy Formula
Proof of Lemma 7. The first condition in Assumption 1 directly gives (29). Let λ1 , . . . , λp be the eigenvalues of X ⊤ X, and define p
µ bX,p :=
1X δλ , p i=1 i
x f∆ (x) := log 1 + . ∆
15
TAP accuracy and universal posterior geometry
A P REPRINT
Then Z 1 1 ⊤ log det Ip + X X = f∆ (x)db µX,p (x). p ∆ Let µMP,αp denote the Marchenko-Pastur law corresponding to the aspect ratio p/n = 1/αp . Explicitly, αp q (bαp − x)(x − aαp )1[aαp ,bαp ] (x) dx, µMP,αp = (1 − αp )+ δ0 + 2πx where aγ = (1−γ −1/2 )2 and bγ = (1+γ −1/2 )2 Bai and Silverstein (2010). By the second condition in Assumption 1, Z Z 1 f∆ (x) db µX,p (x) = f∆ (x) dµMP,αp (x) + OP . p For every fixed γ > 0, the deterministic MP integral satisfies Z x dµMP,γ (x) = L(γ, ∆). log 1 + ∆ We verify the identity for a generic parameter, denoted temporarily by α, to simplify notation. Set Z x L(∆) := log 1 + dµMP,α (x). ∆ For ∆ in a compact subinterval of (0, ∞), the derivative of the integrand is uniformly bounded on the MP support. Differentiation under the integral therefore gives Z 1 1 ′ − dµMP,α (x). L (∆) = x+∆ ∆ For the present normalization, the positive MP Stieltjes transform Z 1 mα (−∆) := dµMP,α (x) x+∆ satisfies the quadratic equation ∆mα (−∆)2 + (α∆ + α − 1)mα (−∆) − α = 0 Bai and Silverstein (2010). Its unique positive root is E∆ /∆, by the definition of E(α, ∆). Hence L′ (∆) =
E∆ − 1 . ∆
The fixed-point equation (19), with γ = α and e = E∆ , can be rearranged as α∆ + E∆ =
αE∆ . 1 − E∆
Differentiating (18) with respect to ∆ yields ′ 1 E′ α + E∆ − − ∆ α∆ + E∆ ∆ E∆ α 1 α2 α ′ = E∆ 1+ − + − . α∆ + E∆ E∆ α∆ + E∆ ∆
′ ∂∆ L(α, ∆) = E∆ +α
′ Using α/(α∆ + E∆ ) = (1 − E∆ )/E∆ , the coefficient of E∆ vanishes. Thus
∂∆ L(α, ∆) =
α2 α α(1 − E∆ ) α − = − . α∆ + E∆ ∆ E∆ ∆
The fixed-point equation also gives α∆(1 − E∆ ) = E∆ (α + E∆ − 1), 16
TAP accuracy and universal posterior geometry
A P REPRINT
and therefore α + E∆ − 1 α E∆ − 1 − = = L′ (∆). ∆ ∆ ∆ Finally, as ∆ → ∞, we have E∆ → 1, and both L(∆) and L(α, ∆) tend to zero. Therefore L(∆) = L(α, ∆) for every fixed ∆ > 0. Applying this identity with γ = αp proves (30). ∂∆ L(α, ∆) =
The third condition in Assumption 1 and the Marchenko-Pastur Stieltjes-transform identity give Z ∆ ∆ ⊤ −1 Tr(X X + ∆Ip ) = dµMP,αp (x) + OP (p−1/2 ) p x+∆ = Ep + OP (p−1/2 ). (57) This proves (31). The argument also applies when α = 1: the test functions and their derivatives are bounded near zero for every fixed ∆ > 0. √ It remains to verify Assumption 1 for the i.i.d. designs specified in Lemma 7. Suppose that Xij = ξij / n, where the ξij are standardized i.i.d. random variables with a fixed distribution and finite fourth moment. By the Bai-Yin spectralP
edge theorem (Bai and Silverstein, 2010, Theorem 5.8), ∥X∥op − → 1 + α−1/2 . Thus any fixed CX > 1 + α−1/2 satisfies P (∥X∥op ≤ CX ) −→ 1. Najim and Yao (Najim and Yao, 2016, Theorems 2 and 3) study Gaussian fluctuations and deterministic bias for linear spectral statistics of large sample covariance matrices with general population covariance. Their results allow standardized i.i.d. entries with finite fourth moment, without requiring the fourth moment to match the Gaussian value. The non-Gaussian fourth cumulant contributes to both covariance and bias. They also replace analyticity of the test function by smoothness: Theorem 2 treats centered fluctuations for Cc3 functions, while Theorem 3 gives the bias expansion for Cc18 functions. Here we need only tightness of the unnormalized fluctuations and boundedness of the bias, not their precise Gaussian approximation. These estimates verify Assumption 1 for the present i.i.d. designs; the TAP and posterior comparisons are established separately. In their notation, take N = p, sample size n = np , Rn = Ip , and cn = p/n = 1/αp . Their aspect-ratio condition allows any positive finite limit, including cn → 1/α > 1 when α < 1. Fix δ > 0 and define x δ fδ (x) := log 1 + , gδ (x) := . δ x+δ 2 For either h = fδ or h = gδ , choose a fixed e h ∈ Cc∞ (R) that agrees with h on [0, CX ]. Enlarging CX if necessary, this interval also contains the support of µMP,αp for all sufficiently large p. We decompose Z p X e h(λi ) − p e h(x) dµMP,αp (x) i=1
=
( p X
e h(λi ) − E
i=1
( +
E
p X
) e h(λi )
i=1 p X
Z e h(λi ) − p
) e h(x) dµMP,αp (x) .
i=1
Theorem 2 of Najim and Yao shows that the first term is OP (1). Their Theorem 3, which requires a Cc18 test function, expresses the second term as Zn2 (e h) + o(1). To bound this bias, use their resolvent bias Bn and smooth extension e Φ17 (h) from (6), with a fixed smooth cutoff equal to one for |y| ≤ 1 and compactly supported. Their (101) and the definition of ∂¯ := ∂x + i∂y give |Bn (x + iy)| ≤ K|x + iy|3 y −7 , 17
y > 0,
(iy) e (18) ¯ 17 (e h (x), 0 < y ≤ 1. ∂Φ h)(x + iy) = 17! The extension has support in a fixed bounded set, so the contribution away from the real axis is uniformly bounded. Their representation (53) therefore implies Z Z 1 ∞ ¯ 17 (e |Zn2 (e h)| ≤ |∂Φ h)(x + iy)||Bn (x + iy)| dx dy π 0 R Z 1 ≤ Ceh 1 + y 10 dy = O(1), 0
17
TAP accuracy and universal posterior geometry
A P REPRINT
with constants independent of p. Thus the second term is O(1). The full bias expansion in Theorem 3 is stronger than needed here; we use only its consequence that the deterministic centering error is uniformly bounded. The spectral-edge bound allows e h to be replaced by h with probability tending to one. Consequently, Z p X h(λi ) − p h(x) dµMP,αp (x) = OP (1). i=1
Applying this conclusion to fδ and gδ and dividing by p verifies the two spectral conditions in Assumption 1. The resulting resolvent estimate is in fact stronger than required. Proof of Lemma 9. We first prove the ridge-norm estimate (38). Write W = dent of (X, β ⋆ ). Then √ ⋆ ⊤ a∆ = A−1 ∆A−1 ∆ Sβ + ∆ X z.
√
∆z, where z ∼ N (0, In ) is indepen-
Therefore,
√ −2 ⊤ −2 ⊤ ⋆ ⊤ ⋆ ⊤ ∥a∆ ∥2 = (β ⋆ )⊤ SA−2 ∆ Sβ + ∆z XA∆ X z + 2 ∆(β ) SA∆ X z. √ We first record the quadratic-form identities used below. If u is uniform on S p−1 ( p) and M is a deterministic symmetric matrix, then E[u⊤ M u] = Tr(M ), Tr(M )2 2p 2 ⊤ Tr(M ) − ≤ 2 Tr(M 2 ). Var(u M u) = p+2 p p These formulas follow from rotational invariance and E[ui uj uk uℓ ] = p+2 (δij δkℓ + δik δjℓ + δiℓ δjk ), where δij is the Kronecker delta. For z ∼ N (0, In ) and a deterministic symmetric matrix N ,
E[z ⊤ N z] = Tr(N ),
Var(z ⊤ N z) = 2 Tr(N 2 ).
−2 ⊤ 2 2 2 The eigenvalues of SA−2 ∆ S are λ /(λ + ∆) ≤ 1, whereas the nonzero eigenvalues of XA∆ X are λ/(λ + ∆) ≤ −1 (4∆) . Their squared Hilbert-Schmidt norms are therefore O(p). Applying Chebyshev’s inequality conditionally on X yields
1 ⋆ ⊤ −2 ⋆ 1 −1/2 (β ) SA∆ Sβ = Tr(S 2 A−2 ), ∆ ) + OP (p p p and ∆ ∆ ⊤ ⊤ −1/2 z XA−2 Tr(SA−2 ). ∆ X z = ∆ ) + OP (p p p It remains to control the cross term. Conditional on (X, β ⋆ ), it is centered Gaussian with variance −2 ⋆ 4∆(β ⋆ )⊤ SA−2 ∆ SA∆ Sβ .
The eigenvalues of the matrix in this quadratic form are λ3 /(λ + ∆)4 , which are uniformly bounded for fixed ∆ > 0. Since ∥β ⋆ ∥2 = p, this variance is O(p). Therefore, √ 2 ∆ ⋆ ⊤ −2 ⊤ (β ) SA∆ X z = OP (p−1/2 ). p Combining the three estimates yields 1 1 ∆ −1/2 ∥a∆ ∥2 = Tr(S 2 A−2 Tr(SA−2 ) ∆ )+ ∆ ) + OP (p p p p 1 = Tr S(S + ∆Ip )A−2 + OP (p−1/2 ) ∆ p 1 −1/2 = Tr(SA−1 ) ∆ ) + OP (p p ∆ −1/2 = 1 − Tr(A−1 ). ∆ ) + OP (p p 18
TAP accuracy and universal posterior geometry
A P REPRINT
Substituting (31) yields (38). −1 We next prove the residual estimate (39). Since Ip − A−1 ∆ S = ∆A∆ , we have √ ⊤ ⋆ ∆(In − XA−1 Y − Xa∆ = X(Ip − A−1 ∆ S)β + ∆ X )z √ ⋆ ⊤ = ∆XA−1 ∆(In − XA−1 ∆ β + ∆ X )z.
Thus −1 ⋆ −1 ⊤ 2 ⊤ ∥Y − Xa∆ ∥2 = ∆2 (β ⋆ )⊤ A−1 ∆ SA∆ β + ∆z (In − XA∆ X ) z −1 ⊤ ⊤ + 2∆3/2 (β ⋆ )⊤ A−1 ∆ X (In − XA∆ X )z. −1 2 We use the same variance calculation for the residual terms. The eigenvalues of A−1 ∆ SA∆ are λ/(λ + ∆) , while −1 ⊤ 0 ⪯ In − XA∆ X ⪯ In . Thus the squared Hilbert-Schmidt norms of the two matrices defining the quadratic forms are O(p) and O(n), √ respectively. Conditional applications of Chebyshev’s inequality show that both centered quadratic forms are OP ( p). The cross term is centered Gaussian conditional on (X, β ⋆ ), with variance −1 ⊤ 2 −1 ⋆ ⊤ 4∆3 (β ⋆ )⊤ A−1 ∆ X (In − XA∆ X ) XA∆ β .
In the singular-vector basis, the eigenvalues of the matrix between the two copies of β ⋆ are ∆2 λ/(λ + ∆)4 . They are uniformly bounded, and ∥β ⋆ ∥2 = p, so the variance is O(p). After division by p, all three centered terms are OP (p−1/2 ). Consequently, ∆2 ∆ 1 ⊤ 2 ∥Y − Xa∆ ∥2 = Tr(SA−2 Tr (In − XA−1 + OP (p−1/2 ). ∆ )+ ∆ X ) p p p It remains to simplify the deterministic trace expression. Let r := min(n, p), and let s1 , . . . , sr denote the singular ⊤ values of X, including zeros with multiplicity. The eigenvalues of In − XA−1 ∆ X are ∆ s2i + ∆ for 1 ≤ i ≤ r, together with n − r additional eigenvalues equal to 1. Therefore, −1 ⊤ 2 ∆2 Tr(SA−2 ∆ ) + ∆ Tr (In − XA∆ X ) r X ∆2 s2i ∆3 = + 2 + ∆(n − r) (s2i + ∆)2 (si + ∆)2 i=1 =
r X
∆2
+ ∆(n − r) s2 + ∆ i=1 i
=∆2 Tr(A−1 ∆ ) + ∆(n − p). Consequently, 1 ∆2 ∥Y − Xa∆ ∥2 = Tr(A−1 ∆ )+∆ p p
n − 1 + OP (p−1/2 ) p
= ∆Ep + ∆(αp − 1) + OP (p−1/2 ), where we used (31) and n/p = αp . All conditional variance bounds above are uniform in X, with constants depending only on ∆ and an upper bound for n/p. Averaging the conditional Chebyshev bounds over X therefore gives the stated unconditional OP (p−1/2 ) estimates. This proves the lemma.
B
Sharpness of the Free-Energy and TAP Fluctuation Scale
Proof of the sharpness assertions in Theorem 2. We first prove the ridge-fluctuation lower bound (58) below, then √ transfer it to (11) and (12). Using the exact formula (35), set B∆ := (∆In + XX ⊤ )−1 . Write z := W/ ∆ and µ := Xβ ⋆ . Conditional on (X, β ⋆ ), z ∼ N (0, In ) and √ Y ⊤ B∆ Y = µ⊤ B∆ µ + 2 ∆z ⊤ B∆ µ + ∆z ⊤ B∆ z. 19
TAP accuracy and universal posterior geometry
A P REPRINT
The centered linear and quadratic terms in z are uncorrelated, since Gaussian third moments vanish. Consequently, 2 2 µ + 2∆2 Tr(B∆ ). Var Y ⊤ B∆ Y X, β ⋆ = 4∆µ⊤ B∆ Let Epop := {∥X∥op ≤ CX }, where CX < ∞ is the constant in Assumption 1, so that P(Epop ) → 1. On Epop , all eigenvalues of B∆ are bounded 2 −1 below by (∆ + CX ) . Since n/p ≥ α/2 for all sufficiently large p, on Epop we have n αp 2 Tr(B∆ )≥ ≥ 2 2 )2 . 2 (∆ + CX ) 2(∆ + CX 2 2 Therefore, with c1 := α∆2 /[4(∆ + CX ) ] > 0,
Var (Rp |X, β ⋆ ) =
c1 1 Var Y ⊤ B∆ Y X, β ⋆ ≥ 2 4p p
on Epop . To establish (58), we use the Carbery-Wright inequality (Carbery and Wright, 2001, Theorem 8). Applied to Gaussian measure with q = 2d in their notation, it states that, if G is a standard Gaussian vector and P is a nonzero real polynomial of degree at most d ≥ 1, then for every ε > 0, 1/2 P |P (G)| ≤ ε EP (G)2 ≤ Cd ε1/d , where Cd depends only on d. Conditionally on (X, β ⋆ ), we apply this inequality with d = 2 to the polynomial PX,β ⋆ (z) := Rp − ϕ(αp , ∆). This is a polynomial of degree at most two in z. On Epop , its positive conditional variance ensures that it is nonzero. Moreover, c1 E PX,β ⋆ (z)2 X, β ⋆ ≥ Var (Rp |X, β ⋆ ) ≥ p op op on Ep . Hence, on Ep , for every ε > 0, r c1 P |Rp − ϕ(αp , ∆)| ≤ ε X, β ⋆ ≤ C2 ε1/2 . p Choose a fixed ε > 0 such that C2 ε1/2 ≤ 1/2, set η := 1/2, and define √ c2 := ε c1 . Then, on Epop , P
c2 |Rp − ϕ(αp , ∆)| ≥ √ X, β ⋆ p
≥ η.
Taking expectations gives
c2 P |Rp − ϕ(αp , ∆)| ≥ √ p
≥ ηP(Epop ) = η − o(1).
(58)
Finally, (42) and Lemmas 6 and 8 give FTAP − Rp = OP (p−1 ) = oP (p−1/2 ), FpS − Rp = OP (p−1 ) = oP (p−1/2 ). For either Hp = FTAP or Hp = FpS , the triangle inequality, (58), and (59) give c2 P |Hp − ϕ(αp , ∆)| ≥ √ 2 p c2 c2 ≥ P |Rp − ϕ(αp , ∆)| ≥ √ − P |Hp − Rp | > √ p 2 p ≥ η − o(1). 20
(59)
TAP accuracy and universal posterior geometry
A P REPRINT
Taking c := c2 /2 and η ′ := η/2, we obtain
c lim inf P |FTAP − ϕ(αp , ∆)| ≥ √ ≥ η′ , p→∞ p c lim inf P FpS − ϕ(αp , ∆) ≥ √ ≥ η′ . p→∞ p The constants c and η ′ depend only on (α, ∆, CX ), as required. This proves (11) and (12).
C
Gaussian Conditioning Near the Typical Radius
Proof of Lemma 12. We first prove the result in the centered-radius case rp2 = µp . For the general case, we introduce an exponential tilt under which rp2 is the mean squared radius and then apply the centered-radius result. Centered-radius case. By an orthogonal change of coordinates, we may assume that Σp = diag(c1 , . . . , cp ),
ci ∈ [c− , c+ ].
U := ∥Z∥22 − E∥Z∥22 ,
φ(t) := E[eitU ].
Let Writing Zi = ai +
√
ci Gi , where G1 , . . . , Gp are independent standard Gaussian random variables, and using 2 u2 E[esGi +uGi ] = (1 − 2s)−1/2 exp , 2(1 − 2s)
valid for Re √ s < 1/2 and u ∈ C, with the square-root branch equal to one at s = 0, we substitute s = itci and u = 2itai ci to obtain h i 2 2 2 ita2i E eit(Zi −ai −ci ) = e−it(ai +ci ) (1 − 2itci )−1/2 exp . 1 − 2itci Independence of the coordinates therefore gives φ(t) =
p Y
2
e−it(ai +ci ) (1 − 2itci )−1/2 exp
i=1
ita2i 1 − 2itci
.
Taking absolute values yields |φ(t)| =
p Y 2t2 ci a2i . (1 + 4c2i t2 )−1/4 exp − 1 + 4c2i t2 i=1
In particular, |φ(t)| ≤ (1 + 4c2− t2 )−p/4 .
(60)
For all sufficiently large p, this estimate implies that φ ∈ L1 (R) and tφ(t) ∈ L1 (R). The Fourier inversion formulas used below are therefore valid. Let fU denote the density of U . We claim that c fU (0) ≥ √ , p
(61)
for a constant c > 0 depending only on (c− , c+ , K). Indeed, σp2 := Var(U ) = 2 Tr(Σ2p ) + 4a⊤ Σp a satisfies cp ≤ σp2 ≤ Cp. Choose t0 > 0 sufficiently small. On [−t0 , t0 ], choose the continuous branch of log φ(t) satisfying log φ(0) = 0. This branch is well defined because every factor in the product representation of φ(t) is nonzero on this interval. A Taylor expansion at the origin gives, uniformly for |t| ≤ t0 , log φ(t) = −
σp2 t2 + Ep (t), 2 21
|Ep (t)| ≤ Cp|t|3 ,
TAP accuracy and universal posterior geometry
A P REPRINT
where the remainder bound follows from p X
(c3i + a2i c2i ) ≤ Cp.
i=1
√
For t = s/ p and |s| ≤ M0 , the preceding expansion and the bounds cp ≤ σp2 ≤ Cp imply, for all sufficiently large p, 2 √ Re φ(s/ p) ≥ ce−Cs . √ Consequently, for each fixed M0 ≥ 1, the substitution t = s/ p gives Z Z M0 2 c c0 Re φ(t) dt ≥ e−Cs ds ≥ √ , √ √ p −M0 p |t|≤M0 / p R 1 −Cs2 for all sufficiently large p, where c0 := c −1 e ds > 0 does not depend on M0 . We next bound the complementary −1 integral. Set t1 := (2c− ) . For each fixed M0 ≥ 1 and all sufficiently large p, the inequality log(1 + x) ≥ x/2 for 0 ≤ x ≤ 1 gives Z (1 + 4c2− t2 )−p/4 dt √ M0 / p<|t|≤t1
2 2 2 c− pt Ce−cM0 exp − . dt ≤ √ √ 2 p M0 / p
Z ∞ ≤2
For p > 4, the change of variables u = 2c− |t|, together with 1 + u2 ≥ 2u for u ≥ 1, yields Z Z 2−p/4 ∞ −p/4 (1 + 4c2− t2 )−p/4 dt ≤ u du c− |t|>t1 1 =
2−p/4 . c− (p/4 − 1)
Together with (60), these two estimates imply Z
2
Ce−cM0 |φ(t)| dt ≤ . √ √ p |t|>M0 / p
Fourier inversion gives 1 fU (0) = 2π
Z
1 φ(t) dt = 2π R
Z Re φ(t) dt. R 2
Choose M0 so large that the constant in the complementary bound satisfies Ce−cM0 < c0 /2. The preceding estimates then prove (61). The same characteristic-function bound and Fourier inversion at an arbitrary u ∈ R give the uniform upper estimate Z 1 C sup fU (u) ≤ |φ(t)| dt ≤ √ . (62) 2π R p u∈R Since ρZ (t) = fU (t − µp ) for t > 0, this also proves (45). For each coordinate, set Ai (t) := E (Zi − ai )eitU . Since Zi ∼ N (ai , ci ) and ∂zi U = 2zi , Gaussian integration by parts gives Ai (t) = ci E ∂zi eitU = 2itci E Zi eitU = 2itci (ai φ(t) + Ai (t)) . Solving for Ai (t) yields 2itci ai E (Zi − ai )eitU = φ(t) . 1 − 2itci 22
TAP accuracy and universal posterior geometry
A P REPRINT
Let νi be the finite signed measure νi (A) := E (Zi − ai )1{U ∈A} . The preceding formula and (60) imply Ai ∈ L1 (R). Hence, by Fourier inversion, νi has the continuous density Z 1 gi (u) := e−itu Ai (t) dt. 2π R For u in a neighborhood of zero, the coarea formula applied to z 7→ ∥z∥22 − µp gives Z fZ (z) fU (u) = dHp−1 (z), 2 2∥z∥ 2 ∥z∥2 =µp +u Z (zi − ai )fZ (z) dHp−1 (z). gi (u) = 2∥z∥2 ∥z∥22 =µp +u √ Initially these identities hold for almost every u. Parametrizing each sphere by z = µp + u θ shows that both righthand sides are continuous near zero, so the identities hold there pointwise. Since ∥z∥2 is constant on each level set, the ratio of the two expressions equals the conditional expectation defined in the statement. Thus, gi (u) = fU (u) (E[Zi | U = u] − ai ) , whenever fU (u) > 0. In particular, fU (0) > 0 by (61), and evaluation at u = 0 yields fU (0) (E[Zi | U = 0] − ai ) = gi (0) Z 1 2itci ai = φ(t) dt. 2π R 1 − 2itci Since |1 − 2itci | ≥ 1, (60) implies Z
|t|(1 + 4c2− t2 )−p/4 dt.
|fU (0) (E[Zi | U = 0] − ai )| ≤ C|ai | R
For p > 4, direct integration gives Z
|t|(1 + 4c2− t2 )−p/4 dt =
R
1 . c2− (p − 4)
Combining this identity with (61) gives C|ai | |E[Zi | U = 0] − ai | ≤ √ . p Squaring the preceding bound and summing over i yields E[Z | ∥Z∥22 = µp ] − a 2 ≤
C∥a∥2 ≤ CK. √ p
(63)
General case. Assume that |rp2 − µp | ≤ ε0 p. To reduce this case to the centered one, define 2
ψ(λ) := log E[eλ∥Z∥2 ]. For λ < 1/(2c+ ), define the tilted probability measure Pλ by dPλ (z) := exp λ∥z∥22 − ψ(λ) . dP 2
Multiplying the density of N (a, Σp ) by eλ∥z∥2 and collecting the quadratic and linear terms gives a density proportional to 1 ⊤ −1 exp − z ⊤ (Σ−1 − 2λI )z + z Σ a . p p p 2 Completing the square shows that Pλ is the Gaussian law N (aλ , Cλ ), where −1 Cλ := (Σ−1 , p − 2λIp )
−1 aλ := Cλ Σ−1 a. p a = (Ip − 2λΣp )
23
TAP accuracy and universal posterior geometry
A P REPRINT
Moreover, ψ ′ (λ) = Eλ ∥Z∥22 ,
ψ ′′ (λ) = Varλ (∥Z∥22 ). √ For |λ| ≤ 1/(4c+ ), the eigenvalues of Cλ lie in [2c− /3, 2c+ ], ∥aλ ∥2 ≤ 2K p, and, uniformly on this interval, ψ ′′ (λ) = 2 Tr(Cλ2 ) + 4a⊤ λ Cλ aλ ≍ p. In particular, ψ ′ is strictly increasing on this neighborhood. After decreasing ε0 if necessary, its image contains [µp − ε0 p, µp + ε0 p]. Hence there is a unique λp in this neighborhood such that ψ ′ (λp ) = rp2 . By construction, rp2 is the mean squared radius under Pλp , so the centered-radius estimate applies. The mean-value theorem also gives |λp | ≤ C
|rp2 − µp | . p
Using the formula for aλ , we also obtain ∥aλp − a∥2 ≤ C|λp |∥a∥2 ≤ C
|rp2 − µp | . √ p
The conditional law on {∥Z∥22 = rp2 } is unchanged by the tilt. Indeed, the Radon-Nikodym derivative equals the constant exp{λp rp2 − ψ(λp )} on this level set, and this factor cancels upon normalization. Moreover, throughout the chosen neighborhood of zero, the covariance matrices C√λ have eigenvalues bounded above and below by positive constants depending only on (c− , c+ ), and ∥aλ ∥2 ≤ 2K p. Thus the constant in (63) may be chosen uniformly for these tilted laws. Applying the centered-radius estimate under Pλp gives E[Z | ∥Z∥22 = rp2 ] − a 2 ≤ Eλp [Z | ∥Z∥22 = rp2 ] − aλp 2 + ∥aλp − a∥2 |rp2 − µp | . √ p
≤C +C
√ It remains to prove the local two-sided density bound. If |rp2 − µp | ≤ M p, then the preceding mean-value estimate gives |λp | ≤ CM p−1/2 . Let ρλp denote the density of ∥Z∥22 under Pλp . Since Eλp ∥Z∥22 = rp2 , the centered bounds (61) and (62), applied uniformly to the tilted Gaussian laws, give CM cM √ ≤ ρλp (rp2 ) ≤ √ . p p The original and tilted densities satisfy ρZ (rp2 ) = exp{ψ(λp ) − λp rp2 }ρλp (rp2 ). Furthermore, 0 ≤ λp rp2 − ψ(λp ) =
Z λp
uψ ′′ (u) du ≤ Cpλ2p ≤ CM .
0
Thus the exponential factor is bounded above and below by positive constants, which proves (44) and the lemma. References Bai, Z. and Silverstein, J. W. (2010). Spectral Analysis of Large Dimensional Random Matrices. Springer Series in Statistics. Springer, New York, 2nd edition. Banerjee, S., Castillo, I., and Ghosal, S. (2026). Bayesian inference in high-dimensional models. Stat. Surv., 20:180– 241. Barbier, J., Macris, N., Dia, M., and Krzakala, F. (2020). Mutual information and optimality of approximate messagepassing in random linear estimation. IEEE Trans. Inform. Theory, 66(7):4270–4303. Bayati, M., Lelarge, M., and Montanari, A. (2015). Universality in polytope phase transitions and message passing algorithms. Ann. Appl. Probab., 25(2):753–822. Bayati, M. and Montanari, A. (2011). The dynamics of message passing on dense graphs, with applications to compressed sensing. IEEE Trans. Inform. Theory, 57(2):764–785. 24
TAP accuracy and universal posterior geometry
A P REPRINT
Belius, D. (2022). High temperature TAP upper bound for the free energy of mean field spin glasses. arXiv:2204.00681. Blei, D. M., Kucukelbir, A., and McAuliffe, J. D. (2017). Variational inference: a review for statisticians. J. Amer. Statist. Assoc., 112(518):859–877. Bolthausen, E. (2014). An iterative construction of solutions of the TAP equations for the Sherrington-Kirkpatrick model. Comm. Math. Phys., 325(1):333–366. Carbery, A. and Wright, J. (2001). Distributional and Lq norm inequalities for polynomials over convex bodies in Rn . Math. Res. Lett., 8(3):233–248. Celentano, M., Fan, Z., Lin, L., and Mei, S. (2023). Mean-field variational inference with the TAP free energy: geometric and statistical properties in linear models. arXiv:2311.08442. Chen, W.-K., Panchenko, D., and Subag, E. (2023). Generalized TAP free energy. Comm. Pure Appl. Math., 76(7):1329–1415. Ding, J. and Sun, N. (2025). Capacity lower bound for the Ising perceptron. Probab. Theory Related Fields, 193:627– 715. Fan, Z., Mei, S., and Montanari, A. (2021). TAP free energy, spin glasses and variational inference. Ann. Probab., 49(1):1–45. Han, Q. and Shen, Y. (2023). Universality of regularized regression estimators in high dimensions. Ann. Statist., 51(4):1799–1823. Krzakala, F., Manoel, A., Tramel, E. W., and Zdeborová, L. (2014). Variational free energies for compressed sensing. In Proceedings of the 2014 IEEE International Symposium on Information Theory (ISIT), pages 1499–1503, Honolulu, HI, USA. Li, Y., Fan, Z., Sen, S., and Wu, Y. (2024). Random linear estimation with rotationally-invariant designs: Asymptotics at high temperature. IEEE Trans. Inform. Theory, 70(3):2118–2153. Maillard, A., Foini, L., Lage Castellanos, A., Krzakala, F., Mézard, M., and Zdeborová, L. (2019). High-temperature expansions and message passing algorithms. J. Stat. Mech. Theory Exp., 2019(11):113301. Mézard, M., Parisi, G., and Virasoro, M. A. (1987). Spin Glass Theory and Beyond: An Introduction to the Replica Method and Its Applications, volume 9 of World Scientific Lecture Notes in Physics. World Scientific. Najim, J. and Yao, J. (2016). Gaussian fluctuations for linear spectral statistics of large random covariance matrices. Ann. Appl. Probab., 26(3):1837–1887. Plefka, T. (1982). Convergence condition of the TAP equation for the infinite-ranged Ising spin glass model. J. Phys. A, 15(6):1971–1978. Qiu, J. (2023). Sub-optimality of the naive mean field approximation for proportional high-dimensional linear regression. In Advances in Neural Information Processing Systems, volume 36, pages 63874–63893. Qiu, J. and Sen, S. (2023). The TAP free energy for high-dimensional linear regression. Ann. Appl. Probab., 33(4):2643–2680. Reeves, G. and Pfister, H. D. (2019). The replica-symmetric prediction for random linear estimation with gaussian matrices is exact. IEEE Trans. Inform. Theory, 65(4):2252–2283. Subag, E. (2023). TAP approach for multispecies spherical spin glasses II: the free energy of the pure models. Ann. Probab., 51(3):1004–1024. Thouless, D. J., Anderson, P. W., and Palmer, R. G. (1977). Solution of “solvable model of a spin glass”. Philos. Mag., 35(3):593–601. Wainwright, M. J. and Jordan, M. I. (2008). Graphical models, exponential families, and variational inference. Found. Trends Mach. Learn., 1(1-2):1–305. Yedidia, J. S., Freeman, W. T., and Weiss, Y. (2001). Bethe free energy, Kikuchi approximations, and belief propagation algorithms. Technical Report TR2001-16, Mitsubishi Electric Research Laboratories, Cambridge, MA. Yedidia, J. S., Freeman, W. T., and Weiss, Y. (2003). Understanding belief propagation and its generalizations. In Lakemeyer, G. and Nebel, B., editors, Exploring Artificial Intelligence in the New Millennium, chapter 8, pages 239–269. Morgan Kaufmann, San Francisco, CA. Yu, Z. and Liu, J. (2026). Proof of the TAP free energy for high-dimensional linear regression with spherical priors at all temperatures. In Proceedings of the Twenty-Ninth International Conference on Artificial Intelligence and Statistics, volume 300 of Proceedings of Machine Learning Research. PMLR. arXiv:2506.20768. 25