ConceptioArchivearXiv CS
arXiv CSopen access

How abundant are good interpolators?

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

How abundant are good interpolators? August Y. Chen ∗

Ahmed El Alaoui †

arXiv:2606.06469v1 [math.ST] 4 Jun 2026

Abstract Let S be the set of unit norm linear classifiers θ ∈ Rd which correctly classify every point of a labeled dataset (Xi , yi )ni=1 , Xi ∈ Rd , yi ∈ {−1, +1}, with a possibly negative margin κ fixed in advance. Under two natural data-generating distributions of the (X, y) pairs – a Gaussian mixture model and a logistic model with Gaussian features – and in the proportional regime n/d → α with small enough α, we establish a large deviation principle on the event that a point θ chosen uniformly at random from S achieves a given generalization error, with high probability over the choice of the data. The associated large deviation rate function is deterministic and describes the proportion, at the exponential scale in d, of interpolating classifiers having a given desired performance. As a consequence, we establish the following concentration phenomenon: all but an exponentially small fraction of interpolating classifiers have approximately the same generalization performance given by the unique maximizer of this rate function. We numerically compare this maximizer to the performance of empirical risk minimization by gradient descent and to the performance of a natural linear program, both finding a point in S, and deduce that in the overparametrized regime of small α, these efficient procedures outperform the vast majority of interpolators, pointing to their nontrivial benign overfitting in this setting.

Contents 1 Introduction

2

2 Settings

3

3 Main results 5 3.1 LDP for a uniform interpolator . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 3.2 LDP for the posterior distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8 3.3 Comparisons and numerical simulations . . . . . . . . . . . . . . . . . . . . . . . . . 10 4 Related work and discussion

12

5 Proof ideas and roadmap 5.1 Relaxation to the Gaussian reference measure . . . . . . . . . . . . . . . . . . . . . . 5.2 Turning off the mollification . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5.3 Conversion to the spherical reference measure . . . . . . . . . . . . . . . . . . . . . .

13 14 15 16

∗ †

Cornell University, Department of Computer Science. Ithaca, NY, USA. Email: [email protected] Cornell University, Department of Statistics and Data Science. Ithaca, NY, USA. Email: [email protected].

1

A Computing the expected log-partition function A.1 Interpolating each datapoint via cavity in n . . . . . . . . . . . . . . . . . . . . . . . A.2 Simplifying the expression from cavity in n . . . . . . . . . . . . . . . . . . . . . . . A.3 Proof of Theorem A.1 for interpolators . . . . . . . . . . . . . . . . . . . . . . . . . . A.4 Proof of Theorem A.2 for posterior . . . . . . . . . . . . . . . . . . . . . . . . . . . .

20 23 28 31 36

B Concentration with respect to Gaussian Measure B.1 Proof of Theorem B.1 for interpolators . . . . . . . . . . . . . . . . . . . . . . . . . . B.1.1 Preliminaries . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . B.1.2 Proof of Theorem B.1 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . B.2 Proof of Theorem B.2 for posterior . . . . . . . . . . . . . . . . . . . . . . . . . . . .

37 39 40 51 64

C Conversion from Gaussian to spherical measure: proof of Theorems 3.1, 3.5 C.1 Proof of Theorems 3.1, 3.5 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . C.2 Proof of Theorem C.1 for interpolators: κ < 0 . . . . . . . . . . . . . . . . . . . . . . C.3 Proof of Theorem C.1 for interpolators: κ ≥ 0 . . . . . . . . . . . . . . . . . . . . . . C.4 Proof of Theorem C.2 for posterior . . . . . . . . . . . . . . . . . . . . . . . . . . . .

68 71 74 81 84

D Technical results for proof of Theorems A.1, A.2 88 D.1 Proof of Lemma A.5 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 89 D.2 Proof of Lemma A.8 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 93 D.3 Proof of Proposition A.6 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 96 D.3.1 First part of Proposition A.6 . . . . . . . . . . . . . . . . . . . . . . . . . . . 97 D.3.2 Second part of Proposition A.6 . . . . . . . . . . . . . . . . . . . . . . . . . . 111 E Properties of the RS equations 113 E.1 Proof of Proposition A.10 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 114 E.2 Proof of Proposition A.3 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 119 E.3 Proof of Lemma C.3 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 132 E.4 Proof of Propositions 3.2, 3.3, 3.6, 3.7, and Lemma 3.8 . . . . . . . . . . . . . . . . . 134 F Additional technical results 139 F.1 Truncated Logarithm . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 140

1

Introduction

Overparametrized statistical models have far more parameters than training points. Their empirical success in the context of deep learning has elicited a flurry of theoretical activity on the reasons such models are able to generalize well. With so many parameters, these models can perfectly fit the training data with simple and efficient algorithms minimizing an empirical loss without hurting generalization performance [BHMM19]. With such an overparametrization one expects the existence of a large set of points in the parameter space with zero training error, suggesting that good generalization is a property of the specific algorithm used to minimize the loss function, rather than a consequence of an explicit control on the model’s complexity, as might be done by explicit regularization [ZBH+ 21]. This phenomenon, known as implicit regularization, has been the topic of much research showing that gradient descent and its variants tend to find solutions of some minimum norm, thereby selecting a model in a statistically favorable way; see for instance [NBMS17, SHN+ 18]. 2

In this paper we are interested in the generalization properties of the set of interpolators: points in parameter space with zero training error. We want to understand the full profile of possible generalization performances, beyond the ones algorithmically selected via empirical risk minimization, and investigate whether there is a typical performance achieved by “most” interpolators. We consider the simple setting of linear binary classification with fixed margin κ, in d dimensions and with n samples, under two distinct data-generating distributions of constant signal-to-noise ratio λ: a mixture of two Gaussians, and the logistic model with Gaussian features. In each case we provide a precise asymptotic characterization of the fraction of interpolators having any desired generalization performance x, in the proportional limit n → ∞ and n/d → α, and α small enough. Our result takes the form of a “quenched” large deviation principle (LDP) on the event that a point chosen uniformly at random from the set of interpolators achieves a performance approximately x. The rate function of this LDP is given by a variational problem whose maximum is achieved at a value x⋆ = x⋆ (α, κ, λ), the performance of a ‘typical’ interpolator. As a byproduct, we establish that with high probability over the data, all but an exponentially small fraction of interpolators have a performance near x⋆ . We next remark that this typical performance is lower than the ones algorithmically achieved: we compare the value of x⋆ to the performances xLP and xGD of two interpolators respectively found by a linear programming procedure proposed in [MZZ24], and gradient descent on the logistic loss function analyzed in [DKT22]. We find that x⋆ is smaller than xLP and xGD for a large range of parameters. We deduce that “good” interpolators — in the sense of being at least as good as the two aforementioned efficient baseline algorithms — are exceedingly rare in this high-dimensional, constant SNR, overparametrized regime. For further comparison we also derive an asymptotic characterization of the Bayes optimal performance xBayes – that of a parameter vector drawn from the posterior distribution given the data – under both data-generating distributions. This is similarly done by establishing an LDP on the generalization performance under the corresponding posterior distributions. Our result confirms a conjecture of Theisen, Klusowski and Mahoney [TKM21] that the generalization error of a uniformly random interpolator should concentrate around a typical value. They address in their work the question of ‘abundance of good interpolators’ where they conduct a heuristic analysis of this typical error in the setting of the Gaussian mixture model, leading them to claim that “good classifiers are abundant in the interpolation regime”, seemingly at odds with our findings. But their analysis implicitly assumes a concentration condition about the Gram matrix of the data which can only hold at high SNR: the separation between the Gaussian centers must be growing with the dimension. In this paper we reach a more nuanced conclusion, namely that good interpolators are exceedingly rare in the constant SNR regime – making implicit regularization a genuinely nontrivial phenomenon – and that the typical performance x⋆ increases with the SNR λ (and with the sampling ratio α) leading to a good typical performance in the large SNR limit, in accordance with the results of [TKM21] in that limit.

2

Settings

We consider the setting of binary classification with data (Xi , yi )ni=1 where Xi ∈ Rd is a vector of √ d covariates and yi ∈ {−1, +1} is the class label. Let θ⋆ ∈ Sd−1 ( d) be a distinguished ‘signal’ direction. We consider two data-generating distributions: 1. The Gaussian mixture model Dgmm : independently for i = 1, · · · , n let yi = ±1 with equal 3

probability, and r Xi | yi ∼ N

λ yi θ⋆ , Id d

! (2.1)

for a fixed signal-to-noise ratio (SNR) parameter λ > 0. 2. The logistic model Dlogistic : independently for i = 1, · · · , n let Xi ∼ N (0, Id ) and conditionally on Xi , ( p  +1 with prob. φ λ/d ⟨Xi , θ⋆ ⟩ , p  yi = (2.2) −1 with prob. 1 − φ λ/d ⟨Xi , θ⋆ ⟩ , for a fixed SNR λ > 0. Here φ : R → [0, 1] is an increasing link function which we fix throughout the paper to have the sigmoid form φ(x) =

1 . 1 + e−x

(2.3)

We will denote the Rademacher distribution on {−1, +1} with probability p assigned to +1 by Rad(p), hence we can write  p  yi | Xi ∼ Rad φ λ/d ⟨Xi , θ⋆ ⟩ . Let X ∈ Rn×d be the matrix with the vectors X √i as its rows, and let κ ∈ R be a fixed margin parameter. We consider the set Sκ (X, y) ⊂ Sd−1 ( d) of directions θ that correctly classifies all data points up to margin κ: n o √ √ Sκ (X, y) := θ ∈ Sd−1 ( d) : yi ⟨Xi , θ⟩ ≥ κ d for all 1 ≤ i ≤ n . (2.4) This is the set of linear interpolators with margin κ. d n The √ generalization error of an estimator of θ⋆ , a measurable function θ̂n : (R × {−1, +1}) → d−1 S ( d), is the probability of misclassifying a new data point:  E(θ̂n ) = P y⟨X, θ̂n ⟩ < 0 , where (X, y) ∼ D ∈ {Dgmm , Dlogistic }. We remark that by rotational invariance of the Gaussian distribution, the above error only depends on the inner product ⟨θ̂n , θ⋆ ⟩/d of which it is a monotonically decreasing function. Therefore we can measure performance by amount of overlap instead of generalization error. Large deviations for interpolating linear classifiers Let µX,y denote the uniform measure over Sκ (X, y), Eq. (2.4), whenever this set is non empty (and define µX,y arbitrarily otherwise). We want to establish an asymptotic, high-probability formula for the random quantity  1 log µX,y {θ : |⟨θ, θ⋆ ⟩/d − x| ≤ ε} , d

(2.5)

where n and d are proportionally large and ε is an arbitrary small number, capturing the fraction of interpolators with overlap near x with θ⋆ in the exponential scale. Equivalently we would like to establish a ‘quenched’ large deviation principle for the event that ⟨θ, θ⋆ ⟩/d ≈ x, where θ is drawn 4

uniformly at random from Sκ (X, y). From (2.5) this boils down to estimating the logarithm of the area of Sκ (X, y) (as a measurable subset of the sphere) and that of its slices Sx,ε (X, y) := Sκ (X, y) ∩ {θ : |⟨θ, θ⋆ ⟩/d − x| ≤ ε} ,

x ∈ (−1, 1) , ε > 0 ,

(2.6)

corresponding to all points having overlap approximately x with the signal direction θ⋆ . The quantity in (2.5) is then equal to (1/d) log |Sx,ε (X, y)|/|Sκ (X, y)| , where |A| denotes the d − 1 √ dimensional surface area of any Lebesgue measurable subset of Sd−1 ( d). This will be the content of Theorem 3.1. Large deviations for Bayes-optimal classifiers We also consider a Bayesian setting where √ the signal direction θ⋆ is drawn uniformly at random from the sphere Sd−1 ( d). If the data (X, y) conditional on θ⋆ are drawn from D ∈ {DGMM , DLogistic √ }, the density of the posterior distribution d−1 (relative to the normalized area measure on S ( d)) of θ⋆ given (X, y) takes the form ! n √ yi Xi , θ 1 Y √ , θ ∈ Sd−1 ( d) , (2.7) exp u pX,y (θ) = ZX,y d i=1 √ √ where u(x) = λ x for D = DGMM , and u(x) = log φ( λ x) for D = DLogistic . The denominator ZX,y is the normalizing constant called the partition function of pX,y . We will similarly be interested in quantifying the probability of the large deviation event {|⟨θ, θ⋆ ⟩/d − x| ≤ ε}, when θ is drawn from the posterior distribution (2.7), and this amounts to establishing an asymptotic formula for  1 log pX,y {θ : |⟨θ, θ⋆ ⟩/d − x| ≤ ε} . d

(2.8)

This will be the content of Theorem 3.5.

3

Main results

Fix the parameters α, κ, λ, a function u : (−1, 1) → [−∞, ∞), and a random variable S. Consider the following function defined on (−1, 1) × [0, 1) → R:   p p Φ(x, q) = α E log EW exp u xS + (1 − x2 )qZ + (1 − x2 )(1 − q)W (3.1)  1 q 1 + + log (1 − x2 )(1 − q) , 21−q 2 where S, Z, W are mutually independent, and Z, W ∼ N (0, 1). EW is the expectation with respect to W while E averages over Z and S.

3.1

LDP for a uniform interpolator

In this section we state our main results on the large deviations of the overlap under the uniform measure µX,y on the set of interpolators. We let Z +∞ N (x) = x

5

dz 2 e−z /2 √ 2π

(3.2)

be the complementary c.d.f. of the standard normal distribution, and let ( 0 u(x) = −∞

if x ≥ κ , otherwise .

(3.3)

! p  κ − xS − (1 − x2 )qZ 1 q 1 p + + log (1 − x2 )(1 − q) . 21−q 2 (1 − x2 )(1 − q)

(3.4)

The expression of Φ simplifies to Φ(x, q) = α E log N

The law of the random variable S will depend on the data-generating distribution:  n • If (Xi , yi ) i=1 ∼ DGMM we let √ S=

λ + G,

G ∼ N (0, 1) .

(3.5)

√  Y |G ∼ Rad φ( λG) .

(3.6)

 n • If (Xi , yi ) i=1 ∼ DLogistic we let S = Y G,

G ∼ N (0, 1) ,

We are now in position to state our quenched large deviation principle for the uniform measure µX,y on Sκ (X, y). Let ϕ = sup inf Φ(x, q) , (3.7) x∈(−1,1) q∈[0,1)

and for an interval I ⊆ (−1, 1), ϕ̄(I) = sup inf Φ(x, q) − ϕ .

(3.8)

x∈I q∈[0,1)

Further let κ+ = max{κ, 0}. Theorem 3.1. Fix λ > 0, κ ∈ R and let n = ⌊αd⌋. There exists α0 = α0 (λ, κ+ ) > 0 such that the following holds for all α < α0 in both cases (3.5) and (3.6). We have ϕ > −∞, and for any interval I ⊆ [−1 + δ, 1 − δ] with δ = δ(α, λ, κ+ ) > 0, there exists ε0 > 0 such that for 0 < ε ≤ ε0 there exists d0 , K > 0 such that for d ≥ d0 we have √ 1 log Sκ (X, y) − ϕ − log 2πe ≤ ε , d

(3.9)

  1 log µX,y θ : ⟨θ, θ⋆ ⟩/d ∈ I − ϕ̄(I) ≤ 2ε , d

(3.10)

  1 log µX,y θ : ⟨θ, θ⋆ ⟩/d ≥ 1 − δ < −c , d

(3.11)

and

with probability at least 1 − e−d/K over the data (X, y), where c > 0 is a universal constant. (Here ε0 depends on α, κ, λ, |I|, and d0 , K depend additionally on ε.) 6

The quantity δ = δ(α, λ, κ+ ) is defined in (C.1); for α ≤ α0 , we have δ ≤ 1/10. We furthermore note that as a corollary of our proof, the above results hold for any prescribed δ > 0 if now α ≤ α0 (δ, λ, κ+ ). The above theorem characterizes the area of the set of interpolators and the rate function of the LDP for the overlap with the signal direction under µX,y . This shows that the normalized overlap ⟨θ, θ⋆ ⟩/d is exponentially concentrated around the maximizers of the function x 7→ inf q Φ(x, q) when θ ∼ µX,y , with high probability over the data. (Note the bound (3.11) shows that the set of interpolators with normalized overlap outside [−1 + δ, 1 − δ] makes an exponentially small contribution.) We remark that α0 depends on the margin κ only when the latter is positive; we will show in Appendix A that one can take α0 (λ, κ) = c(λ)/(1 + κ2 ) in this case. For negative margin the result holds as long as α is small enough as a function of the SNR λ only. We further show that for α small enough the ‘sup-inf’ saddle point problem (3.7) is achieved at a unique pair (x, q) ∈ [0, 1)2 which satisfies a system of two nonlinear equations. We state the results for the Gaussian mixture model and the logistic model in two separate propositions. Let A be the inverse of Mills’ ratio 2

1 e−x /2 A(x) = −N (x)/N (x) = √ . (3.12) 2π N (x) We will say that the saddle point problem sup inf Φ is uniquely achieved if there is a unique x⋆ ∈ [0, 1) achieving supx∈(−1,1) inf q∈[0,1) Φ(x, q), and if there is a unique q⋆ ∈ [0, 1) achieving inf q∈[0,1) Φ(x⋆ , q). ′

Proposition 3.2 (GMM). Let κ ∈ R, λ > 0. There exists α0 = α0 (λ, κ+√ ) such that for all α < α0 , the saddle point problem (3.7) where Φ is defined via (3.1) with S = λ + G, G ∼ N (0, 1), is uniquely achieved at a pair (x⋆ , q⋆ ) solving the system of equations " !# p p κ − xS − (1 − x2 )qZ x √ p = α λ(1 − q) E A , (3.13) 1 − x2 (1 − x2 )(1 − q)  !2  p 2 κ − xS − (1 − x )qZ q , p = α E A (3.14) 1−q (1 − x2 )(1 − q) where Z ∼ N (0, 1) independently of S. Proposition 3.3 (Logistic). Let κ ∈ R, λ > 0. There exists α0 = α0 (λ, κ+ ) such that for all α < α0 , the saddle √ point  problem (3.7) where Φ is defined via (3.1) with S = Y G, G ∼ N (0, 1) and Y |G ∼ Rad φ( λG) , is uniquely achieved at a pair (x⋆ , q⋆ ) solving the system of equations " !# p p √  κ − xS − (1 − x2 )qZ x √ p = α λ(1 − q) E φ − λS A , (3.15) 1 − x2 (1 − x2 )(1 − q)  !2  p 2 )qZ κ − xS − (1 − x q , p (3.16) = α E A 1−q (1 − x2 )(1 − q) where Z ∼ N (0, 1) independently of (G, Y ). Furthermore, in both the above cases, (x⋆ , q⋆ ) ∈ [0, Kα]2 where K > 0 depends on λ and κ+ . Corollary 3.4. The uniqueness of the maximizer of the function inf q Φ(·, q), achieved at x⋆ ∈ [−1 + δ, 1 − δ], combined with the large deviation bound (3.10) implies that all but an exponentially small fraction of points θ ∈ Sκ (X, y) have nearly the same overlap x⋆ with the signal direction θ⋆ , with high probability over the data (X, y). 7

Physical interpretation of q: While the parameter x⋆ represents the asymptotic amount of overlap a direction θ ∼ µX,y has with θ⋆ , the parameter q⋆ has an equally concrete interpretation: let θ and θ′ be two independent copies from µX,y conditional on the same realization of the data √ (X, y), and denote by θ̄ and θ̄′ their projections (of norm d) orthogonal to θ⋆ . Then q⋆ represents the asymptotic value of the overlap ⟨θ̄, θ̄′ ⟩/d. A crucial step in the proofs of the above formulas is to show that this overlap indeed concentrates around q⋆ .

3.2

LDP for the posterior distribution

We state a similar result for the posterior distribution (2.7) in the Bayesian setting. Recall the function Φ from (3.1). In this Bayesian case both the function u and the random variable S depend on the data-generating distribution.  n • If (Xi , yi ) i=1 ∼ DGMM we let √ S=

√ λ+G

and

u(x) =

λx.

(3.17)

 n • If (Xi , yi ) i=1 ∼ DLogistic we let √ u(x) = log φ( λ x) , √  where as before, G ∼ N (0, 1) and Y |G ∼ Rad φ( λG) . S =YG

and

(3.18)

We have the following quenched LDP for the posterior distribution pX,y . Let ϕ=

sup

inf Φ(x, q) ,

(3.19)

x∈(−1,1) q∈[0,1)

and for I ⊂ (−1, 1), ϕ̄(I) = sup inf Φ(x, q) − ϕ .

(3.20)

x∈I q∈[0,1)

Theorem 3.5. Fix λ > 0 and let n = ⌊αd⌋. There exists α0 = α0 (λ) > 0 such that the following holds for all α < α0 in both cases (3.17) and (3.18). For any interval I ⊆ [−1 + δ, 1 − δ] with δ = δ(α, λ) > 0, there exist ε0 > 0 such that for all 0 < ε ≤ ε0 there exists d0 , K > 0 such that for all d ≥ d0 we have √ 1 log ZX,y − ϕ − log 2πe ≤ ε , (3.21) d   1 log pX,y θ : ⟨θ, θ⋆ ⟩/d ∈ I − ϕ̄(I) ≤ 2ε , (3.22) d and

  1 log pX,y θ : ⟨θ, θ⋆ ⟩/d ≥ 1 − δ < −c , d

(3.23)

with probability at least 1 − e−d/K over the data (X, y), where c > 0 is a universal constant. (Again, ε0 depends on α, λ, |I|, and d0 , K depend additionally on ε.) Here, δ = δ(α, λ) is defined in (C.6); for α ≤ α0 , we have δ ≤ 1/10. Again, the above results hold for any prescribed δ > 0 if α ≤ α0 (δ, λ). (The bound (3.23) shows that the set of interpolators with normalized overlap outside [−1 + δ, 1 − δ] makes an exponentially small contribution.) Similarly in this case, the sup-inf saddle point problem is uniquely achieved: 8

Proposition 3.6 (GMM). For λ > 0, there exists α0 = α0 (λ)√such that for √ all α < α0 , the saddle point problem (3.19) where Φ is defined via (3.1) with u(x) = λx and S = λ + G, G ∼ N (0, 1), is achieved at a unique pair (xBayes , qBayes ), which solves the system of equations x=

q , 1−q

αλ (1 − x2 ) =

q . (1 − q)2

(3.24)

and

αλ . 1 + 2αλ

(3.25)

and

These equations admit the explicit solutions xBayes =

αλ , 1 + αλ

qBayes =

Proposition 3.7 (Logistic). Define the random variable √ p p EW φ′ ( λV ) √ , where V := xY G + (1 − x2 )qZ + (1 − x2 )(1 − q)W , (3.26) R(x, q) := EW φ( λV ) √  and G, Z, W ∼ N (0, 1) are mutually independent and Y |G ∼ Rad φ( λG) independently of everything else. For λ > 0 there exists α0 = α0 (λ) such that for √ all α < α0 , the saddle point problem (3.19) where Φ is defined via (3.1) with u(x) = log φ( λx) and S = Y G, is uniquely achieved at a pair (xBayes , qBayes ), which solves the system of equations √    x αλ(1 − q) E φ − λY G R(x, q) = , 1 − x2   q and αλ(1 − x2 ) E R(x, q)2 = . (1 − q)2

(3.27)

The uniqueness results combined with (3.22) shows exponential concentration of the normalized overlap under the posterior distribution, and characterizes the Bayes-optimal performance in estimating the signal direction θ⋆ . Again, in both of the above cases, (x⋆ , q⋆ ) ∈ [0, Kα]2 where K > 0 depends on λ. A special relation between x and q: Keeping in mind the interpretation of q mentioned in the previous section, the present Bayesian setting imposes a special relation between x and q: for two independent copies θ, θ′ from pX,y we have by the law of iterated expectations E⟨θ, θ⋆ ⟩ = E⟨θ, θ′ ⟩ .

(3.28)

Therefore by concentration of the random variables inside each of the above expectations (see Appendix D.3) the Pythagorean theorem would impose the relation x = x2 + (1 − x2 )q ,

i.e.,

x=

q . 1−q

(3.29)

This identity is apparent in the Gaussian mixture case, see Eq. (3.24), but is less clear in the logistic case. We show a similar result: Lemma 3.8. For λ > 0, there exists α0 = α0 (λ) > 0 such that for every α ∈ (0, α0 ), any solution (x, q) ∈ [0, 1)2 to the system (3.27) satisfies x=

q . 1−q 9

More precisely, the solution is of the form x=

s , 1+s

q=

s , 1 + 2s

(3.30)

where s > 0 is the unique solution to the scalar fixed-point equation    s s 2 s = αλ E R , . 1 + s 1 + 2s

(3.31)

This allows to search for the pair (xBayes , qBayes ) along a univariate curve, usually referred to as the Nishimori line, see [Nis01, Chapter 4] and [LM17]. No such simplification occurs in general outside the Bayesian setting. In particular the system of equations governing the pair (x⋆ , q⋆ ) concerning the uniform measure µX,y treated in the previous section remains genuinely bivariate. The proofs of Theorems 3.1 and 3.5 can be found in Appendix C.1. The proofs of Propositions 3.2, 3.3, 3.6, 3.7, and Lemma 3.8 can be found in Appendix E.4.

3.3

Comparisons and numerical simulations ell_2 squared error vs sample-to-dimension ratio (gmm data)

2.0 1.8 1.6

err_* err_Bayes err_GD err_LP

1.8

1.4

| - *|^2

| - *|^2

ell_2 squared error vs sample-to-dimension ratio (logistic data)

2.0

err_* err_Bayes err_GD err_LP

1.2 1.0 0.8

1.6 1.4 1.2 1.0

0.6 0.0

0.2

0.4

0.6

=n/d

0.8

0.8

1.0

0.0

(a) Data from DGMM

0.2

0.4

0.6

=n/d

0.8

1.0

(b) Data from Dlogistic

Figure 1: The squared ℓ2 error ∥θ̂ −θ⋆ ∥22 at parameters κ = 0, λ = 1. The black line is the error err⋆ of an interpolator chosen uniformly at random. The red curve is the Bayes-optimal error errBayes . The blue and orange curves are the mean errors of the GD and LP estimators averaged across T = 20 Monte Carlo trials. The error bars represent one standard deviation. In this section we compare the performance of a typical interpolator relative to that of a draw from the posterior and to two efficient algorithms for finding points in Sκ (X, y). (Note that the posterior distribution is not supported on the set of interpolators; a draw from pX,y will√typically not belong to Sκ (X, y) for any κ.) For convenience we rescale the parameter space by d in this section, so that θ⋆ and all its estimators have unit ℓ2 norm. The first algorithm is a second-order cone programming procedure n

LP :

1X max yi Xi , θ θ∈Rd n

subj. to

∥θ∥2 ≤ 1 ,

yi Xi , θ ≥ κ , ∀i ∈ [n] .

(3.32)

i=1

Note that the spherical constraint ∥θ∥2 = 1 is relaxed to a unit ball norm constraint ∥θ∥2 ≤ 1 making the optimization problem convex. By abuse of terminology we refer to this procedure 10

as a linear program (LP), emphasizing the linear nature of the objective Pn and the data-related constraints. This LP maximizes the correlation with the vector v = (1/n) i=1 yi Xi representing a proxy of the signal direction, while staying on the correct side of the hyperplane corresponding to each data point. In [MZZ24, Theorem 5.3] it was shown that when the data is drawn from DLogistic , the norm constraint is saturated, resulting in an estimator θ̂LP lying on the unit sphere, ∥θ̂LP ∥2 = 1 with high probability. Thus the algorithm produces a point in Sκ (X, y) with high probability. The second algorithm is empirical risk minimization of a smooth loss function via gradient descent: n n X o min R(θ) := ℓ yi Xi , θ − κ∥θ∥2 , (3.33) θ∈Rd

i=1

where ℓ : R → R≥0 is a decreasing function with limx→+∞ ℓ(x) = 0; we only consider the logistic loss ℓ(x) = log(1 + e−x ) for concreteness. The gradient descent iteration on R reads GD :

θt+1 = θt − η∇R(θt ) n X   t =θ −η yi Xi − κθt /∥θt ∥ ℓ′ yi Xi , θt − κ∥θ∥2 ,

(3.34)

i=1

where η > 0 is a fixed stepsize and t ∈ {0, 1, · · · }. It was shown in [MZZ24, Theorem 6.1] that under smoothness assumptions on ℓ (satisfied by the logistic loss) and for small enough stepsize η, this iteration produces a margin-κ classifier: lim min yi Xi , θt /∥θt ∥2 ≥ κ ,

t→∞ 1≤i≤n

(3.35)

as long as ∥θt ∥2 → ∞ as t → ∞. For both data distributions Dlogistic and DGMM and both the LP and GD algorithms we perform a Monte Carlo estimate of the squared ℓ2 distance ∥θ̂ − θ⋆ ∥22 = 2(1 − ⟨θ̂, θ⋆ ⟩) over T = 20 trials where the data is freshly resampled. We fix the dimension to d = 3000, the margin to κ = 0, the signal-to-noise ratio to λ = 1 and vary the number of samples-to-dimension ratio α in the range {0.01, · · · , 1} in increments of 0.1. In the case of LP we use a specialized numerical solver to obtain a solution θ̂LP to (3.32) and verify that the normalization condition ∥θ̂LP ∥2 = 1 is satisfied to 10−2 accuracy in all instances. In the case of GD we run the iteration (3.34) with step size η = 1.5 until either the ℓ2 norm of the gradient of R is smaller than 10−8 or a maximum number of iterations of 500 is reached. We verify that the margin condition min1≤i≤n ⟨yi Xi , θ̂GD ⟩ ≥ κ is satisfied in all instances, where θ̂GD = θt /∥θt ∥2 at the terminal time step t. We compare these empirically obtained values to our theoretical predictions of the squared ℓ2 distances achieved by a typical interpolator θ ∼ µX,y : err⋆ := lim E ∥θ − θ⋆ ∥22 = 2(1 − x⋆ ), d→∞

where x⋆ is described in Lemmas 3.2 and 3.3 for the data-generating distributions DGMM and DLogistic respectively, and the error achieved by the posterior mean θ̂Bayes , errBayes := lim E ∥θ̂Bayes − θ⋆ ∥22 = 1 − xBayes , d→∞

where xBayes is described in Lemmas 3.6 and 3.7 for these same distributions. The corresponding fixed point equations are numerically solved via bisection. 11

The numerical results are summarized in Figure 1. We see that both algorithms significantly outperform a typical interpolator under both distributions and across all tested values of α, with the Bayes error being smallest overall. Since this error concentrates exponentially in d under the uniform measure on Sκ (X, y) (Theorem 3.1), only an exponentially small fraction of points in Sκ (X, y) perform at least as well as LP or GD. This numerical finding can supported by theory: for data generated from the logistic model, √ [MZZ24] established that the LP (3.32) achieves correlation xLP = Θ( α) in the limit κ → −∞ with λ fixed; see Lemmas 22, 23, 24 therein. In contrast we have x⋆ = O(α) for κ ≤ 0. This comparison again establishes that for small α, only an exponentially small fraction of points in Sκ (X, y) perform as well as the LP, in line with the results of our numerical simulations. We remark that this effect is purely due to overparametrization: the performance gap between the above estimators vanishes as the sampling ratio α is increased. Heuristically this can be seen from the fact that the area of the set Sκ (X, y) is a decreasing function of α since the map α 7→ ϕ (Eq. (3.7)) is decreasing, leading to consistent estimation for large enough α (or λ → ∞). This is in agreement with the results of [MZZ24] who provided bounds on the maximum and minimum estimation performance of the points in Sκ (X, y). These bounds converge towards a common value as α increases (see Figure 4 therein).

4

Related work and discussion

Statistical literature and benign overfitting There has been an impressive amount of work on overparametrization and implicit regularization over the last decade suggesting that overparametrization not only does not hurt statistical performance, but is essential for algorithmic tractability. We refer to [BMR21, Bel21] for a partial synthesis of the literature. Most existing work analyzes the performance of specific algorithms and statistical procedures in the overparametrized regime and establish their success under suitable data-generating assumptions. A subset of this literature is concerned with obtaining exact high-dimensional asymptotics in (generalized) linear models under idealized data distributions, allowing a precise tracking of the generalization error and other metrics as the parameters of the model are changed. In this line of work we mention the analysis of maximum likelihood and M-estimation in logistic regression [CS20, DKT22], Bayes-optimal estimation in GLMs [BKM+ 19], boosting and minimum ℓ1 norm interpolators [LR23, LW21, LS22], maximum margin interpolators [SHN+ 18, MRSY25], among many others. Closer to our setting, [MZZ24] identify sharp bounds on the interpolation threshold for data generated from the logistic model: the maximum value of α = n/d under which interpolators exist with positive probability, and above which the set of interpolators is empty with high probability. They prove bounds on the generalization performance of any element of Sκ (X, y) and leave open the question of whether there exists a value within those bounds which is typical. They additionally analyze the performance of the linear program (3.32), while [DKT22] analyze the ERM (3.33), which are both compared against in the previous section. Our work is closely related to [TKM21, HHV+ 24] who address the question of abundance of good interpolators and tempered overfitting of randomly selected interpolating neural networks. As mentioned earlier [TKM21] conducts a heuristic analysis in the same setting as ours, under the Gaussian mixture model, and conclude that most interpolators have good generalization performance, in the sense of being close to Bayes-optimal. A careful reading of their argument reveals a seemingly innocuous assumption (Eq. (12) therein) which can √ only hold at high SNR: the separation between the Gaussian centers need to grow faster than d log n, and is otherwise violated at constant separation. [HHV+ 24] considers the setting of fully connected multi-layer neural networks 12

(NN) under binary features x and clean labels h⋆ (x) ∈ {−1, 1} corrupted with bit flip noise; see also [BHN+ 24]. When the clean labels are generated by a “teacher” network h⋆ with N⋆ number of neurons, they prove that the generalization of a randomly picked interpolating “student” neural network (of potentially much larger size) with N neurons is close to Bayes-optimal when the number of samples n is roughly larger than N⋆4 ∨ N . Importantly, the number of weights of the student network is absent from the sample complexity, which can be very large without hurting generalization, leading to tempered overfitting. But the term N⋆4 in the sample complexity is roughly quadratic in the number of weights of the teacher network, and an analogy with our setting where the teacher and the student have the same number of parameters d suggests that this is still a “large α regime” as n needs to grow at least like d2 . High dimensional spin glasses and the perceptron model The set of interpolators is also known as the set of solutions to the spherical perceptron model (or the half-space model when κ = 0): if one draws n uniformly random half-spaces of margin κ in d dimensions, when is their common intersection with the sphere typically non-empty? And what is the typical area of this intersection? Motivated by questions of learning with threshold functions, Cover [Cov65] determined the interpolation threshold in this model, also known as the storage capacity for κ = 0: the set is typically non-empty if n ≤ 2(1 − o(1))d and typically empty if otherwise n ≥ 2(1 + o(1))d, with additional finer understanding of the critical window. Gardner [Gar88, GD88] proposed formulas which would later bear her name in the literature for the storage capacity and for the logarithm of the surface area, based on the non-rigorous replica method, for general κ ≥ 0. These formulas were proved by Shcherbina and Tirozzi [ST03] leveraging convexity and concentration of measure. Stojnic [Sto13a] found an elegant and simple Gaussian comparison argument establishing Gardner’s storage capacity for κ ≥ 0. Talagrand’s books [Tal10, Tal11] present a unified treatment of a larger family of models and incorporate additional important techniques such as Guerra’s interpolation method [GT02, Gue03] into their study. The behavior of the spherical perceptron model when κ < 0 is much more delicate and few results are available. [Sto13b] showed that Gardner’s formula obtained for κ ≥ 0 is an upper bound when κ < 0. [MZZ24] prove upper and lower bounds on the capacity for κ ≤ 0 which match as κ → −∞, and a detailed study at the physics level of rigor can be found in [FPS+ 17]. An algorithm for finding a feasible point was proposed in [EAS22] and was shown to work for κ < 0 under an unproved analytic condition regarding the structure of a certain variational problem. The binary version of the problem where the sphere is replaced by the binary hypercube has also been intensely studied. The storage capacity question has recently been resolved: the existence of a sharp threshold sequence was established in [NS23, Xu21] and the capacity formula predicted by Krauth and Mézard [KM89] was established in [DS19, Hua24]. In addition, peculiar structural phenomena of freezing and preponderance of isolated solutions were established in a symmetric version of this model [APZ19, PX21, ALS22b] but remain open otherwise, and efficient algorithms for finding (a connected set of) solutions were proposed in [ALS22a].

5

Proof ideas and roadmap

The proofs are executed in a unified way by considering the random probability density on the sphere, n Y √ √  √ 1  µ(θ) = 1 |θ1 / d − x| ≤ ε exp u ⟨gi , θ⟩/ d , θ ∈ Sd−1 ( d) , (5.1) Z i=1

13

where x and ε > 0 are fixed, u : R → R is a smooth concave function with derivatives of controlled growth, and the gi ’s are i.i.d. random vectors whose first coordinates are copies of a random variable S drawn from a general strongly log-concave density,√and the remaining coordinates are standard normal. Without loss of generality we can set θ⋆ = de1 , and the density µ corresponds to each case of interest by appropriately choosing u and the law of S; see Eqs. (3.5), (3.6) for the (formal) correspondence with µX,y and Eqs. (3.17), (3.18) for the correspondence with pX,y . The goal is then to estimate the normalized logarithm of the restricted partition function Z, called the free energy. We follow Talagrand’s approach in proving the Gardner formula for the spherical perceptron model [Tal10, Tal11], with additional technical work required for handling the ‘signal part’ of the data: the vectors yi Xi have a non-Gaussian component with law S in the direction of θ⋆ and are otherwise Gaussian on the complement subspace. This effectively turns the model into a spherical perceptron with Gaussian disorder and random inhomogeneous margins. While Talagrand’s treatment is restricted to κ ≥√0, we are able to extend his techniques to our case. By constraining the overlap ⟨θ, θ⋆ ⟩/d = θ1 / d ≃ x and leveraging the log-concavity of the law of S we can control the fluctuations of the first coordinate to any desired accuracy. This allows us to conduct our analysis for κ < 0 as long as α is suitably small. We additionally remark that as a corollary of our proof, we prove the Gardner formula for the spherical perceptron model for all κ ≤ 0 and α ≤ α0 smaller than a universal constant. We mention that our analysis does not operate on the Nishimori line, in contrast to the recent activity in Bayesian settings; see for instance [LM17, BDM+ 16, BKM+ 19]. As already mentioned in Section 3.2, Eq. (3.28), there is no special relation between the one-replica and two-replica overlaps ⟨θ, θ⋆ ⟩, ⟨θ1 , θ2 ⟩ outside the Bayesian setting. Such a relation can simplify interpolation and concentration arguments, and usually lead to a simpler variational formula. The fact that we are outside this special line when studying the uniform measure µX,y on interpolators is responsible for the ‘sup-inf’ structure of the variational formulas we obtain. Finally, we note that although our arguments are lengthy, all known proofs of the Gardner formula are technically involved [Tal10, Tal11, ST03, BNSX22]. In the remaining of this section we present at a high-level the main steps of the proof.

5.1

Relaxation to the Gaussian reference measure

In more detail, the first step is to relax the spherical constraint to full space by switching to a Gaussian reference measure, and we let ! n X √ √  1  2 µβ (θ) = 1 |θ1 / d − x| ≤ ε exp u ⟨gi , θ⟩/ d − β∥θ∥ , θ ∈ Rd , (5.2) Zβ i=1

now defined relative the Lebesgue measure on Rd , and where β > 0 is a parameter. The quadratic term acts as a confining potential. The measure µβ is a log-concave density which has strong concentration properties. In particular the normalized squared norm ∥θ∥2 /d of a draw θ ∼ µβ , and the normalized overlap ⟨θ1 , θ2 ⟩/d of two independent draws from µβ both concentrate around their expectations. We then use the Gaussian interpolation method to prove a Gardner formula for Zβ of the form n o 1 log Zβ = inf Φ̄(x, q, ρ) − β(ρ + x2 ) + O(ε) + oP (1) , (5.3) 0≤q≤ρ d where oP (1) → 0 in probability in the limit d → ∞ and n/d → α, and   1 q  √ 1 √ + log ρ − q , Φ̄(x, q, ρ) := α E log EW exp u xS + qZ + ρ − qW + 2ρ−q 2 14

(5.4)

This is the content of Appendix A. Central to the successful application of this technique are 1) strong concentration of measure statements of the interpolating Gibbs measures and 2) uniqueness of the solution to the system of Replica-Symmetric (RS) equations Eq. (A.11). We demonstrate these in Appendix D.3 and Appendix E respectively. The unique minimizers (q⋆ , ρ⋆ ) of the formula in (5.3) are respectively the asymptotic values of h h  i  i E E(θ1 ,θ2 )∼µβ ×µβ ⟨θ̄1 , θ̄2 ⟩/d , and E Eθ∼µβ ∥θ̄∥2 /d , where θ̄ = (θ2 , · · · , θd ) omits the first coordinate of θ, and similarly for θ̄1 , θ̄2 . √ We remark that in the Bayesian setting with data generated from DGMM we have u(x) = λx which allows for explicit computations of the log-partition function via a contour integral, and one can obtain the desired formula without resorting to these techniques. Nevertheless we do not pursue this and provide a unified treatment here.

5.2

Turning off the mollification

To obtain an asymptotic expression for the area of the set of interpolators with fixed overlap with the signal direction, Sκ (X, y)∩{|⟨θ, θ⋆ ⟩/d−x| ≤ ε}, we turn off the mollification exp u(x) appearing in (5.2) and (5.4) into a hard step function 1{x ≥ κ}, under the Gaussian reference measure, and then convert back to the spherical setting. Regarding the first step, showing that the desired statement (5.3) survives this passage to a hard step function is done √ by adding one constraint at a time. (Note the i-th datapoint imposes the constraint ⟨g , θ⟩ ≥ κ d on the set of possible i √ d−1 interpolators θ ∈ S ( d), and henceforth we refer to the datapoints also as constraints.) Let √ √  (5.5) Ui = θ ∈ Sd−1 ( d) : ⟨gi , θ⟩ ≥ κ d , and 1 am = E log d

√  1 |θ1 / d − x| ≤ ε exp

Z T

i<m Ui

n X

√  u ⟨gi , θ⟩/ d − β∥θ∥2

! m ≤ n . (5.6)

dθ ,

i=m

We want to estimate the difference am+1 − am which can be written as am+1 − am =

√  1 1 E log Gm (Um ) − E log Gm eu(⟨gm ,θ⟩/ d) , d d

(5.7)

where Gm is the measure with density proportional to o n \ √ 1 θ∈ Ui , |θ1 / d − x| ≤ ε exp i<m

X

√ u(⟨gi , θ⟩/ d) − β∥θ∥2

! ,

i>m

which does not involve the m-th constraint. If exp u approximates the step function to accuracy ε′ then we can hope to establish |am+1 −am | ≤ ε′ /d, then sum over m ∈ [n]. The main difficulty is that Gm (Um ) may apriori be too small compared to the second term in (5.7); for instance the additional constraint Um may have an empty intersection with ∩i<m Ui . Talagrand’s approach is to truncate the logarithm in the definition of am , Eq. (5.6) to a minimum floor value e−cd , c > 0, and control the probability that Gm (Um ) falls below this threshold. We will show that Gm (Um ) is very unlikely to be exponentially small by extracting an exponential number of nearly orthogonal directions in ∩i<m Ui , conditional on this set being large (this is provided by the truncation), and show that the fresh constraint Um has a small probability of being violated by all these directions simultaneously. 15

This is a geometric argument executed via the tools of Gaussian processes. A minor adaptation is needed in our case to handle the non-Gaussian, but still strongly log-concave, first coordinate. This argument will also be used to establish concentration of the logarithms of the relevant partition functions, and is executed in Appendix B. We point out a version of this argument in [NS23] which handles arbitrary i.i.d. sub-Gaussian vectors gi , at the price of a suboptimal concentration rate that scales like a fractional power of d.

5.3

Conversion to the spherical reference measure

We convert back from the Gaussian measure to the spherical measure by exploiting the √ √ concentration of θ̄ on a thin shell of width o( d) around a sphere of radius approximately ρ⋆ d when θ ∼ µβ . The parameter β is then chosen to satisfy the unit radius condition x2 + ρ⋆ = 1. A formula for (1/d) log Z, Eq. (5.1) can then be obtained from (5.3) via a volume-to-surface comparison argument. This eliminates the parameters β and ρ from the equation (5.4), leading to a variational formula with q as the only parameter. The overlap constraint on the first coordinate can be eliminated by taking a maximum over x, resulting in the ‘max-min’ structure of the variational formula. This allows us to finish the proof of the large deviation result for the posterior distributions in the Bayesian setting (in Section 3.2) since the function u already satisfies the required smoothness properties. To obtain the LDP for uniform measure on the set of interpolators (in Section 3.1) we apply this argument after the mollification has been turned off. This is the content of Appendix C; see Theorems C.1 and C.2. This completes the proof of Theorems 3.1 and 3.5. Statement on use of AI: The authors used ChatGPT 5.1 to produce the code used for the numerical simulations presented in Section 3.3. The authors instructed the AI to write Julia code for the Monte Carlo experiment and the numerical solvers for the fixed point equations in (x, q) appearing in Section 3. The code was thoroughly checked by the authors. The authors used ChatGPT 5.5 and Gemini 3.1 and 3.5 to establish Part 2 in the proof of Propositions 3.2, 3.3, 3.6, 3.7, in writing the fixed point argument for the map Tα , Eq. (E.49), in the proof of Lemma 3.8, and to prove Proposition A.11 and Lemma C.4. Acknowledgements: AC would like to thank Shuangping Li for stimulating discussions. AE was supported by the National Science Foundation grant DMS-2450867.

References [ALS22a]

Emmanuel Abbe, Shuangping Li, and Allan Sly. Binary perceptron: efficient algorithms can find solutions in a rare well-connected cluster. In Proceedings of the 54th Annual ACM SIGACT Symposium on Theory of Computing, pages 860–873, 2022.

[ALS22b]

Emmanuel Abbe, Shuangping Li, and Allan Sly. Proof of the contiguity conjecture and lognormal limit for the symmetric perceptron. In 2021 IEEE 62nd Annual Symposium on Foundations of Computer Science (FOCS), pages 327–338. IEEE, 2022.

[APZ19]

Benjamin Aubin, Will Perkins, and Lenka Zdeborova. Storage capacity in symmetric binary perceptrons. Journal of Physics A: Mathematical and Theoretical, 52(29):294003, 2019.

[BDM+ 16] Jean Barbier, Mohamad Dia, Nicolas Macris, Florent Krzakala, Thibault Lesieur, and Lenka Zdeborová. Mutual information for symmetric rank-one matrix estimation: A 16

proof of the replica formula. In Advances in Neural Information Processing Systems, pages 424–432, 2016. [Bel21]

Mikhail Belkin. Fit without fear: remarkable mathematical phenomena of deep learning through the prism of interpolation. Acta Numerica, 30:203–248, 2021.

[BGL13]

Dominique Bakry, Ivan Gentil, and Michel Ledoux. Analysis and geometry of Markov diffusion operators, volume 348. Springer Science & Business Media, 2013.

[BHMM19] Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal. Reconciling modern machine-learning practice and the classical bias–variance trade-off. Proceedings of the National Academy of Sciences, 116(32):15849–15854, 2019. [BHN+ 24] Gon Buzaglo, Itamar Harel, Mor Shpigel Nacson, Alon Brutzkus, Nathan Srebro, and Daniel Soudry. How Uniform Random Weights Induce Non-uniform Bias: Typical Interpolating Neural Networks Generalize with Narrow Teachers. In 41st International Conference on Machine Learning, pages 5035–5081. PMLR, 2024. [BKM+ 19] Jean Barbier, Florent Krzakala, Nicolas Macris, Léo Miolane, and Lenka Zdeborová. Optimal errors and phase transitions in high-dimensional generalized linear models. Proceedings of the National Academy of Sciences, 116(12):5451–5460, 2019. [BMR21]

Peter L Bartlett, Andrea Montanari, and Alexander Rakhlin. Deep learning: a statistical viewpoint. Acta numerica, 30:87–201, 2021.

[BNSX22]

Erwin Bolthausen, Shuta Nakajima, Nike Sun, and Changji Xu. Gardner formula for ising perceptron models at small densities. In Conference on Learning Theory, pages 1787–1911. PMLR, 2022.

[Cov65]

Thomas M Cover. Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition. IEEE transactions on electronic computers, (3):326–334, 1965.

[CS20]

Emmanuel J Candès and Pragya Sur. The phase transition for the existence of the maximum likelihood estimate in high-dimensional logistic regression. The Annals of Statistics, 48(1):27–42, 2020.

[DKT22]

Zeyu Deng, Abla Kammoun, and Christos Thrampoulidis. A model of double descent for high-dimensional binary linear classification. Information and Inference: A Journal of the IMA, 11(2):435–495, 2022.

[DS19]

Jian Ding and Nike Sun. Capacity lower bound for the Ising perceptron. In Proceedings of the 51st Annual ACM SIGACT Symposium on Theory of Computing, pages 816–827, 2019.

[EAS22]

Ahmed El Alaoui and Mark Sellke. Algorithmic pure states for the negative spherical perceptron. Journal of Statistical Physics, 189(2):27, 2022.

[FPS+ 17]

Silvio Franz, Giorgio Parisi, Maxime Sevelev, Pierfrancesco Urbani, and Francesco Zamponi. Universality of the sat-unsat (jamming) threshold in non-convex continuous constraint satisfaction problems. SciPost Physics, 2(3):019, 2017. 17

[Gar88]

Elizabeth Gardner. The space of interactions in neural network models. Journal of physics A: Mathematical and general, 21(1):257, 1988.

[GD88]

Elizabeth Gardner and Bernard Derrida. Optimal storage properties of neural network models. Journal of Physics A: Mathematical and general, 21(1):271, 1988.

[GN65]

David Gale and Hukukane Nikaido. The Jacobian matrix and global univalence of mappings. Mathematische Annalen, 159(2):81–93, 1965.

[GT02]

Francesco Guerra and Fabio Lucio Toninelli. The thermodynamic limit in mean field spin glass models. Communications in Mathematical Physics, 230(1):71–79, 2002.

[Gue03]

Francesco Guerra. Broken replica symmetry bounds in the mean field spin glass model. Communications in Mathematical Physics, 233(1):1–12, 2003.

[HHV+ 24] Itamar Harel, William M Hoza, Gal Vardi, Itay Evron, Nathan Srebro, and Daniel Soudry. Provable tempered overfitting of minimal nets and typical nets. Advances in Neural Information Processing Systems, 37:53458–53524, 2024. [Hua24]

Brice Huang. Capacity threshold for the ising perceptron. In 2024 IEEE 65th Annual Symposium on Foundations of Computer Science (FOCS), pages 1126–1136. IEEE, 2024.

[KM89]

Werner Krauth and Marc Mézard. Storage capacity of memory networks with binary couplings. Journal de Physique, 50(20):3057–3066, 1989.

[LM17]

Marc Lelarge and Léo Miolane. Fundamental limits of symmetric low-rank matrix estimation. Probability Theory and Related Fields, pages 1–71, 2017.

[LR23]

Tengyuan Liang and Benjamin Recht. Interpolating classifiers make few mistakes. Journal of Machine Learning Research, 24(20):1–27, 2023.

[LS22]

Tengyuan Liang and Pragya Sur. A precise high-dimensional asymptotic theory for boosting and minimum-ℓ1 -norm interpolated classifiers. The Annals of Statistics, 50(3):1669–1695, 2022.

[LW21]

Yue Li and Yuting Wei. Minimum ℓ1 -norm interpolators: Precise asymptotics and multiple descent. arXiv preprint arXiv:2110.09502, 2021.

[McS34]

Edward James McShane. Extension of range of functions. Bulletin of the American Mathematical Society, 40(12):837–842, 1934.

[MRSY25] Andrea Montanari, Feng Ruan, Youngtak Sohn, and Jun Yan. The generalization error of max-margin linear classifiers: Benign overfitting and high dimensional asymptotics in the overparametrized regime. The Annals of Statistics, 53(2):822–853, 2025. [MZZ24]

Andrea Montanari, Yiqiao Zhong, and Kangjie Zhou. Tractability from overparametrization: The example of the negative perceptron. Probability Theory and Related Fields, 188(3):805–910, 2024.

[NBMS17] Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nati Srebro. Exploring generalization in deep learning. Advances in Neural Information Processing Systems, 30, 2017. 18

[Nis01]

Hidetoshi Nishimori. Statistical Physics of Spin Glasses and Information Processing: An Introduction. Oxford University Press, 2001.

[NS23]

Shuta Nakajima and Nike Sun. Sharp threshold sequence and universality for Ising perceptron models. In Proceedings of the 2023 Annual ACM-SIAM Symposium on Discrete Algorithms (SODA), pages 638–674. SIAM, 2023.

[Pin19]

Iosif Pinelis. Exact bounds on the inverse mills ratio and its derivatives. Complex Analysis and Operator Theory, 13(4):1643–1651, 2019.

[Pré73]

András Prékopa. On logarithmic concave measures and functions. Acta Scientiarum Mathematicarum, 34:335–343, 1973.

[PX21]

Will Perkins and Changji Xu. Frozen 1-rsb structure of the symmetric Ising perceptron. In Proceedings of the 53rd Annual ACM SIGACT Symposium on Theory of Computing, pages 1579–1588, 2021.

[SHN+ 18]

Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro. The implicit bias of gradient descent on separable data. Journal of Machine Learning Research, 19(70):1–57, 2018.

[ST03]

Mariya Shcherbina and Brunello Tirozzi. Rigorous solution of the gardner problem. Communications in Mathematical Physics, 234(3):383–422, 2003.

[Sto13a]

Mihailo Stojnic. Another look at the gardner problem. arXiv preprint arXiv:1306.3979, 2013.

[Sto13b]

Mihailo Stojnic. Negative spherical perceptron. arXiv preprint arXiv:1306.3980, 2013.

[Tal10]

Michel Talagrand. Mean Field Models for Spin Glasses: Volume I: Basic examples, volume 54. Springer Science & Business Media, 2010.

[Tal11]

Michel Talagrand. Mean Field Models for Spin Glasses. Volume II: Advanced ReplicaSymmetry and Low Temperature, volume 55. Springer Science & Business Media, 2011.

[TKM21]

Ryan Theisen, Jason Klusowski, and Michael Mahoney. Good classifiers are abundant in the interpolating regime. In International Conference on Artificial Intelligence and Statistics, pages 3376–3384. PMLR, 2021.

[Ver18]

Roman Vershynin. High-Dimensional Probability: An Introduction with Applications in Data Science, volume 47. Cambridge university press, 2018.

[Whi34]

Hassler Whitney. Analytic extensions of differentiable functions defined in closed sets. Transactions of the American Mathematical Society, 36(1):63–89, 1934.

[Xu21]

Changji Xu. Sharp threshold for the Ising perceptron model. The Annals of Probability, 49(5):2399–2415, 2021.

[ZBH+ 21]

Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals. Understanding deep learning (still) requires rethinking generalization. Communications of the ACM, 64(3):107–115, 2021.

19

A

Computing the expected log-partition function

Recall the definition of Zβ from (5.2). The goal of this section is to compute the free energy – the expected log-partition function (1/d) E log Zβ – in the asymptotic regime d → ∞, nd → α > 0 for α ≤ α0 , where α0 is defined in (A.8). Our expression depends on the solution to an explicit twovariable system of equations known as the Replica-Symmetric (RS) equations, defined in (A.11). Our precise results for the free energy are then given in Theorems A.1, A.2. √ Notation. By rotational symmetry, we assume without loss of generality that θ⋆ = de1 as per Section 5. We then let si = yi Xi1 for 1 ≤ i ≤ n, the component of yi Xi in the span of θ⋆ . Letting D denote the density of the si , we impose the following technical assumption on D which is satisfied when D ∈ {DLogistic , DGMM }. Assumption 1. S ∼ D satisfies the following properties: 1. S has a βsignal strongly log-concave density for some βsignal > 0. 1/2

2. S is Ksignal sub-Gaussian for Ksignal ≥ 1, and the second and fourth moments of S are upper bounded by Ksignal . Note βsignal , Ksignal depend only on λ (and the link function φ when D ≡ DLogistic , which we suppress for notational simplicity). We let φ0 denote the PDF of a standard univariate normal N (0, 1), and let φ1 denote the PDF of S ∼ D. We now formally define the quantities we are interested in understanding. Following the notation in Section 5, for any θ ∈ Rd , θ̄ = (θ2 , . . . , θd ) omits its first coordinate. For all 1 ≤ i ≤ n, we let  1  Si = Si (θ) := √ si θ1 + ⟨gi , θ̄⟩ , d

(A.1)

where si ∼ D, gi ∼ N (0, Id−1 ) are i.i.d., and where gi,j denotes the j − 1-st coordinate of gi . Hence we consider 2 ≤ j ≤ d. This notation is not to be confused with the definition of gi in Section 5, and for the rest of this paper, we define gi ∼ N (0, Id−1 ). Next, for an arbitrary function u : R → R such that the following integrals (A.3), (A.4) exist, we define the Hamiltonian X H(θ) = Hn,d (θ) := u(Si ) . (A.2) 1≤i≤n

Dependence on u(·) is suppressed here in Appendix A and will alwaysP be clear from context. Since √ √  n θ⋆ = de1 , s1 θ1 + ⟨gi , θ̄⟩ = yi ⟨Xi , θ⟩. Thus (A.2) is identical in law to i=1 u ⟨yi Xi ⟩/ d . Finally, for any β > 0 and any interval I ⊆ R, we define the restricted quantities Z   √ Zβ (I) = Zβ,n,d (I) := 1{θ1 / d ∈ I} exp Hn,d (θ) − β∥θ∥2 dθ , (A.3) ϕβ (I) = ϕβ,n,d (I) :=

1 E log Zβ,n,d (I) . d

(A.4)

In the above definitions, we omit dependence on n, d when they are clear from context. We are now in position to state the main results of this section, regarding the free energy of the interpolators and posterior respectively. For the interpolators, our result applies for u(·) such that exp u(x) is as a sufficiently accurate approximation to 1{x ≥ κ}. 20

Theorem A.1 (For interpolators). Let α0 = α0 (λ, κ+ ) be from (A.8) and consider n such that 1 ≤ n ≤ α0 d. Consider any ε > 0. Then there is a ε′ > 0 depending on ε, β0 , β1 , κ, λ such that for any concave u ≤ 0 satisfying 1. u′ ≥ 0 with u′ ̸= 0 on a set of positive Lebesgue measure, 2. u(x) = 0 for x ≥ κ, 3. exp u(x) ≤ ε′ for all x ≤ κ − ε′ , 4. u is four times differentiable and |u(l) | ≤ D for l = 1, 2, 3, 4, there is a εRS0 > 0 depending on ε, u, β0 , β1 , κ, λ such that the following holds (note dependence on ε′ comes in implicitly through u). For any x ∈ [−1, 1], 0 < ε′′ ≤ εRS0 , and β ∈ [β0 , β1 ], letting I = [x − ε′′ , x + ε′′ ], for d large enough in terms of ε, ε′′ , β0 , β1 , κ, λ, we have ϕβ,n,d (I) − RSGI (n/d, β, x) ≤ ε ,

(A.5)

where ϕβ,n,d (I) is from (A.4) and RSGI is from (A.13). For the posterior, u(·) is now fixed based on the assumed data generation process; our result applies for u(·) from (3.17) in the GMM case or (3.18) in the logistic case. Theorem A.2 (For posterior). Let α0 = α0 (λ) be from (A.8) and consider n such that 1 ≤ n ≤ α0 d. Consider u(x) from (3.17), (3.18) in the GMM and logistic cases respectively. Then for any ε > 0, there is εRS0 > 0 depending on ε, α0 , β0 , β1 , λ such that the following holds. For any x ∈ [−1, 1], 0 < ε′′ ≤ εRSG , and β ∈ [β0 , β1 ], letting I = [x − ε′′ , x + ε′′ ], for d large enough in terms of ε, ε′′ , β0 , β1 , λ, we have ϕβ,n,d (I) − RSGI (n/d, β, x) ≤ ε ,

(A.6)

where ϕβ,n,d (I) is defined in (A.4) and RSGI is from (A.13), both with this exp u(x). In both Theorems, RSGI is well-defined thanks to Proposition A.3. Also, note that d does not need to be large enough in terms of α0 . This is because we apply the bound α0 ≤ 2 in the bounds arising in the following proofs, and hence dependence on α0 in d is captured by a universal constant. We now define α0 and RSGI that appear in the above Theorems. Defining α0 : For both the interpolators and the posterior, we let n1 o 1 β0 = , β1 = 10 , Csol = max , 4β1 + 1 . 8 β0

(A.7)

For the interpolators, we now consider ᾱ0 , c0 ∈ (0, 2] small enough in terms of β0 , Csol , λ, and define ( ᾱ0 : κ < 0, α0 = α0 (λ, κ+ ) := (A.8) c0 : κ ≥ 0. κ2 +1 For the posterior, consider u(·) from (3.17) in the GMM case or (3.18) in the logistic case. In both cases, note |u(l) | ≤ D = D(λ) for l = 1, 2, 3, 4. We let Csol be large enough in terms of D, λ. Our results now apply for α0 ∈ (0, 2] small enough in terms of β0 , β1 , D, Csol , and therefore small enough in terms of β0 , β1 , λ.

21

Defining RSG and the RS equations: We now define RSG. To do so, we must define the Replica-Symmetric (RS) equations. Let S ∼ D, Z, W ∼ N (0, 1) be independent. Recall the definition of Φ̄(x, q, ρ) for any 0 ≤ q < ρ and x ∈ R from (5.4). In what follows consider Φ̄(x, q, ρ) with u(x) therein given by exp u(x) = 1{x ≥ κ} for the interpolators, and given by u(·) from (3.17) in the GMM case or (3.18) in the logistic case for the posterior. In turn, for all such (x, q, ρ) and any interval I ⊆ R we define  FI (x, q, ρ) := Φ̄(x, q, ρ) − β ρ + rI (x)  1 q  √ 1 √ = α E log EW exp u xS + qZ + ρ − qW + + log ρ − q (A.9) 2ρ−q 2  − β ρ + rI (x) , where ( 0 rI (x) := x2

:0∈I, : 0 ̸∈ I .

(A.10)

and where dependence on α, β is implicit in the definition of F . Fixing x ∈ [−1, 1] and I, we define the RS equations by ∂FI ∂FI = = 0. ∂q ∂ρ

(A.11)

We will establish in Appendix E.2 the following crucial property of the system (A.11), which applies for both the interpolators and posterior. Proposition A.3. For all (α,β, x) ∈ [0, α0 ] × [β0 , β1 ] ×  [−1, 1], the system (A.11) has a unique solution q0 (α, β, x), ρ0 (α, β, x) = (q0 , ρ0 ) in the domain (q, ρ) : q < ρ, 0 ≤ q ≤ Csol , C1sol ≤ ρ ≤ Csol . Also, we have ρ0 (α, β, x) − q0 (α, β, x) ≥ 2C1sol . Moreover, q0 (α, β, x), ρ0 (α, β, x) are both infinitely differentiable in (α, β, x) for (α, β, x) ∈ [0, α0 ] × [β0 , β1 ] × [−1, 1]. Thus for x ∈ [−1, 1] and I both fixed, RS0,I (α, β, x) and RSGI (α, β, x) are infinitely differentiable in (α, β) for (α, β) ∈ [0, α0 ] × [β0 , β1 ], where RS0,I , RSGI are defined in (A.12), (A.13) below.  Remark 1. Note the system (A.11) and its solution q0 (α, β, x), ρ0 (α, β, x) do not depend on I, since dependence on I in FI comes only through βrI (x), which does not depend on ρ, q. Define in terms of FI the following quantities, which implicitly depend on u(·):  RS0,I (α, β, x) := FI x, q0 (α, β, x), ρ0 (α, β, x) , 1 RSGI (α, β, x) := RS0,I (α, β, x) + log(2πe) . 2

(A.12) (A.13)

Theorems A.1 and A.2 are proven with a unified strategy that uses a Gaussian interpolation argument. This argument proceeds by understanding the expected contribution of adding each datapoint or constraint i on the free energy (1/d) E log Zβ for each 1 ≤ i ≤ n. See Proposition A.4. The free energy then can be written as a sum of such expected contributions. This interpolation argument is carried out in Appendix A.1. We then show in Appendix A.2 that under the conditions on u(x) in Theorem A.1, we can replace u(x) by 1{x ≥ κ} only changes the the free energy by at most ε. This yields an expression that does not depend on u(·) and depends explicitly on the solution to the RS equations. Finally, we sum these expressions across all 22

the datapoints i, 1 ≤ i ≤ n and finish the proof of Theorem A.1 for the interpolators and Theorem A.2 in Appendix A.3 and Appendix A.4 respectively. Crucial to executing this proof strategy is Proposition A.3 – that the RS equations have a unique solution. This is done by showing that the RS equations satisfy a particular convex-concave structure. Establishing this with the signal in the data and the case κ < 0 does not directly follow from the arguments of Talagrand in [Tal10]; we discuss our steps to handle these complications in Appendix E.

A.1

Interpolating each datapoint via cavity in n

Here in Appendix A.1, we fix any d ≥ 1 and n such that 1 ≤ n ≤ α0 d, and define α = nd . To compute the free energy ϕβ,n,d (I), we consider two interpolations. One interpolation is over a parameter v, 0 ≤ v ≤ 1, and aims to elucidate the effect of the n-th datapoint or constraint on the original probability distribution. When v = 1, the n-th datapoint is fully present, and when v = 0, it is eliminated. Hence this interpolation is called ‘cavity in n’. The other interpolation is over a parameter t, 0 ≤ t ≤ 1, and aims to elucidate the effect of the d-th coordinate. When t = 1, the d-th coordinate is fully present, and when t = 0, it is eliminated. Hence this interpolation is called ‘cavity in d’. Specifically, the cavity in n provides us with an approximate expression for ϕβ,n,d (I)−ϕβ,n−1,d (I). This expression depends on the mollifier u(·) and the overlaps qn,d , ρn,d given in (A.18); see Proposition A.4. Establishing this result is the main goal of Appendix A.1. However to define the cavity in n, we will first need to define the cavity in d. We will then use the cavity in d more extensively next in Appendix A.2 to relate qn,d , ρn,d in the expression from Proposition A.4 to q0 , ρ0 from Proposition A.3, and eliminate u(·) in the case of the interpolators. This lets us rewrite the aforementioned expression for ϕβ,n,d (I) − ϕβ,n−1,d (I) in terms of RSGI . Finally, we note that many of our following bounds do not explicitly depend on α0 ; this is because we use the upper bound α0 ≤ 2. Cavity in d: To define the cavity in n in its full generality, we must first set up the cavity in d. For any 1 ≤ i ≤ n and t ∈ [0, 1], define r X si θ1 t 1 Si,t = Si,t (θ) := √ + √ gi,j θj + gi,d θd . d d d 2≤j≤d−1 In terms of Si,t , define Ht,n,d (θ) :=

X 1≤i≤n

p  (1 − t)(r − r̄) 2 u Si,t (θ) + θd (1 − t)rY − θd , 2

(A.14)

where r, r̄ will be chosen next in (A.22), (A.23), and where Y ∼ N (0, 1). Let ⟨ · ⟩t denote the Gibbs average w.r.t. √  1{θ1 / d ∈ I} exp Ht,n,d (θ) − β∥θ∥2 , and let νt (·) := E⟨ · ⟩t .

(A.15)

When we write Ht without explicitly stating n, we will assume this refers to Ht,n,d as written in (A.14). Now consider the overlaps R1,2 =

1 1 1 2 ⟨θ̄ , θ̄ ⟩ , R1,1 = ∥θ̄∥2 , d d 23

(A.16)

where we view R1,1 as a function on Rd−1 and R1,2 as a function on (Rd−1 )2 . Let qt := νt (R1,2 ) ≤ ρt := νt (R1,1 ) .

(A.17)

Note ⟨ · ⟩1 corresponds to exactly the Hamiltonian Hn,d (θ) − β∥θ∥2 . Indeed, specializing to t = 1, we define qn,d := q1 , ρn,d := ρ1 ,

(A.18)

Consequently qn,d ≤ ρn,d . Now consider independent Z, W ∼ N (0, 1), and S ∼ D following the law of each si . For arguments 0 ≤ q < ρ, define √ √ η = η(q, ρ) := xS + qZ + ρ − qW . (A.19) We define the functions  E [u′ (η) exp u(η)] 2 W , EW [exp u(η)] h E [(u′′ (η) + u′ (η)2 ) exp u(η)] i W ψ̄α (q, ρ) := α E . EW [exp u(η)] ψα (q, ρ) := α E

(A.20) (A.21)

In turn, we let "   !2 #   !2 # EW W exp u(η) EW u′ (η) exp u(η) α     E r = ψα q, ρ = α E = , ρ−q EW exp u(η) EW exp u(η)      EW (W 2 − 1) exp u(η) EW (u′′ (η) + u′ (η)2 ) exp u(η) α     = E , r̄ = ψ̄α q, ρ = α E ρ−q EW exp u(η) EW exp u(η) "



(A.22) (A.23)

where the second equality is only valid when ρ > q. Here the equality follows by Gaussian Integration by Parts. As we will justify later in Lemma E.3, we have r ≥ r̄ (irrespective of whether ρ = q or ρ > q). As |u′ |, |u′′ | ≤ D in both the interpolators and GMM case and as α ≤ α0 ≤ 2, we have r, r̄ ≤ K(D), which we use repeatedly in the following. We also note r, r̄ are independent of ε′′ . Cavity in n: We next consider another interpolation parameter v ∈ [0, 1]. To set up the cavity in n, we need to introduce the notion of replicas. Given a probability distribution ν, we consider L i.i.d. draws θℓ , 1 ≤ ℓ ≤ L from ν. These i.i.d. draws are known θℓ are known as replicas. That is, ⊗L . It is often useful to consider functions of L > 1 replicas; one example is the overlap (θℓ )L ℓ=1 ∼ ν R1,2 = d1 ⟨θ̄1 , θ̄2 ⟩. In particular, functions of L > 1 replicas naturally arise when differentiating quantities of interest such as R1,1 and R1,2 with respect to t and v in the cavity in d and n. Proceeding with the setup, we let Z and W ℓ denote i.i.d. N (0, 1) scalars independent of everything else for each replica ℓ, 1 ≤ ℓ ≤ L. For v ∈ [0, 1], define for each replica ℓ, 1 ≤ ℓ ≤ L, r  √ X  p √  1 t √ ℓ ℓ ℓ Sn,t,v = Sv := v √ gn,j θj + gn,d θdℓ + 1 − v qn,d Z + ρn,d − qn,d W ℓ d d 2≤j≤d−1

 θℓ  + sn v √1 + (1 − v)x . d ℓ Note that in the following when we consider Sn,t,v , t will be always held fixed. When it is clear ℓ ℓ ℓ from context we will just write Sv . Note Sn,t,1 = Sn,t .

24

Next, we let ⟨ · ⟩t,∼ denote a Gibbs measure defined by " # Z Y √  1 ℓ ℓ ℓ 2 EW f 1{θ1 / d ∈ I} exp Ht,n−1,d (θ ) − β∥θ ∥ dθ1 · · · dθL (A.24) ⟨f ⟩t,∼ := L Zt,n−1,d 1≤ℓ≤L for f a function of L replicas and the corresponding annealed disorder W ℓ , where we define: Z √  Zt,n−1,d := 1{θ1 / d ∈ I} exp Ht,n−1,d (θ) − β∥θ∥2 dθ , Ht,n−1,d (θ) :=

X

u(Si,t (θ)) + θd

1≤i≤n−1

p (1 − t)(r − r̄) 2 (1 − t)rY − θd . 2

(A.25)

(Note for a fixed value of L, if f is a function of L + 1 replicas, the Gibbs measure (A.24) will use ℓ + 1 replicas ℓ = 1, . . . , ℓ, ℓ + 1.) Now let   P P ℓ ℓ f exp f exp 1≤ℓ≤L u(Sv ) t,∼ 1≤ℓ≤L u(Sn,t,v ) t,∼ νt,v (f ) := E =E . (A.26) L L+1 L exp u(Sv1 ) t,∼ exp u(Sn,t,v ) t,∼ Observe that νt ≡ νt,1 . We will establish that as one datapoint is removed, the resulting difference in the free energy approximately equals the following expression: 1 , we have Proposition A.4. For all d ≥ dA.4 (ε′′ ) ≥ ε′′2

ϕβ,n,d (I) − ϕβ,n−1,d (I) −

i p 1 h KA.4 ε′′ √ E log EW exp u xsn + qn,d Z + ρn,d − qn,d W ≤ , d d (A.27)

where KA.4 depends on D, β0 , β1 , λ, and also on κ for the interpolators. To prove Proposition A.4, we will need to establish the following. First, in Appendix D.1, we will show the following critical Lemma, which lets us transfer expectations from νt,v to νt,1 . Lemma A.5. For any function f : (Rd−1 )L → R with f ≥ 0,    1 νt,v (f ) ≤ K νt,1 (f ) + ε′′ + sup νt,v (f 2 )1/2 , d t,v∈[0,1]2

(A.28)

where K depends on D, β0 , β1 , λ, L, and also on κ for the interpolators. We will apply the abovefor f = |R1,1 − ρn,d  |, |R1,2 − qn,d |, and thus are left with the task of controlling νt,1 |R1,1 − ρn,d | , νt,1 |R1,2 − qn,d | that appear on the right hand side of the above Lemma. To this end, we will also establish the following via concentration of Lipschitz functions of log-concave measures in Appendix D.3. Proposition A.6. For all d ≥ dA.6 (ε′′ ), all t, v ∈ [0, 1] and all k ≤ d/4, for f ≡ R1,1 or R1,2 ,  2k   Kk k . νt,v f − νt,v (f ) ≤ d and   ∥θ̄∥2  log νt,v exp ≤ Kd K

,

  θ2  log νt,v exp d ≤K, K

where K depends on D, β0 , β1 , λ, and also on κ for the interpolators. 25

We also cite a simple fact to bound higher order moments of relevant random variables:   Lemma A.7 (Lemma 3.1.8, [Tal10]). Consider a random variable X ≥ 0 and let C = log E exp X . Then for each k ≥ 1 we have   E X k ≤ 2k (k k + C k ) . Now we have the necessary tools to prove Proposition A.4. Proof of Proposition A.4. We let dA.4 (ε′′ ) = dA.6 (ε′′ ) and consider any d ≥ dA.4 (ε′′ ). Define for the proof of Proposition A.4 the following function of v: h i φ(v) := E log exp u(Sn,1,v ) 1,∼ . (The above function is not to be confused with the link function φ, which does not explicitly appear in the proof of Proposition A.4.) Notice that 1 φ(1) = ϕβ,n,d (I) − ϕβ,n−1,d (I) , d h φ(0) = E log EW exp u xsn +

qn,d Z +

i p ρn,d − qn,d W .

Consequently it remains to prove that  d 1  φ(v) ≤ K(D, β0 , β1 , κ, λ) ε′′ + √ , dv d where here and in the rest of the proof of Proposition A.4, there is no dependence in K on κ for the posterior. To this end, we compute d d dv exp u(Sn,1,v ) 1,∼ φ(v) = E dv exp u(Sn,1,v ) 1,∼

=E

exp u(Sn,1,v )u′ (Sn,1,v ) · sn √θ1d − x

 1,∼

exp u(Sn,1,v ) 1,∼ D

exp u(Sn,1,v )u′ (Sn,1,v ) ·

+E

√1 2 vd

gM , θ̄⟩ − 2√11−v

qn,d Z +

exp u(Sn,1,v ) 1,∼

:= (I) + (II) . Next observe as exp(·) ≥ 0, we have exp u(Sn,1,v )u′ (Sn,1,v ) · sn √θ1d − x

(I) ≤ E

1,∼

exp u(Sn,1,v ) 1,∼ D

≤E



exp u(Sn,1,v )u′ (Sn,1,v ) · sn √θ1d − x exp u(Sn,1,v ) 1,∼ 26

E 1,∼

.

ρn,d − qn,d W

E 1,∼

Next remark, with the quenched disorder fixed, we have D E exp u(Sn,1,v )u′ (Sn,1,v ) · sn √θ1d − x 1,∼ Z h   √ 1 = EW 1{θ1 / d ∈ I} exp H1,n−1 (θ) − β∥θ∥2 Zt,n−1,d i · exp u(Sn,1,v ) · u′ (Sn,1,v ) · √θ1d − x · |sn | dθ ≤ Dε′′ |sn | · exp u(Sn,1,v ) 1,∼ . Here, the last inequality follows as√sn is independent of W and θ, and as the numerator of the last integral is only nonzero when θ1 / d ∈ I, which implies √θ1d − x ≤ ε′′ . Combining with the earlier display implies " ′′ # Dε |sn | · exp u(Sn,1,v ) 1,∼ (I) ≤ E = Dε′′ · E[|sn |] ≤ K(D, λ)ε′′ . exp u(Sn,1,v ) 1,∼ Moreover, applying Gaussian Integration by Parts on gn , Z, and W yields that   (II) ≤ K(D)ν1,v R1,1 − ρn,d + R1,2 − qn,d . The calculation in the above equality is analogous as in the proof of Lemma 3.3.5 of [Tal10]. The reason why the argument still applies is that the Gaussian random variables we integrate by parts with respect to are the exact same as in [Tal10]; there is no integration by parts w.r.t. the disorder from the signal part (I), which we instead bounded directly above. See the proof of Lemma D.1 for the details of a similar computation. We thus have, applying Lemma A.5 to the non-negative function R1,1 − ρn,d + R1,2 − qn,d , d φ(v) dv ≤ (I) + (II)   ≤ K(D, λ)ε′′ + K(D) ν1,v R1,1 − ρn,d + R1,2 − qn,d ≤ K(D, β0 , β1 , κ, λ) ν1,1 R1,1 − ρn,d + R1,2 − qn,d



 1/2 1/2  1 + ε′′ + sup νt,v (R1,1 − ρn,d )2 + νt,v (R1,2 − qn,d )2 d t,v∈[0,1]2 

!

+ K(D, λ)ε′′ . As ν1,1 (R1,1 ) = ρn,d , ν1,1 (R1,2 ) = qn,d , we thus obtain from the first part of Proposition A.6 that K(D, β0 , β1 , κ, λ) √ . d  2 The second part of Proposition A.6 and Lemma A.7 implies νt,v R1,2 ≤ K(D, β0 , β1 , κ, λ). This gives  1/2 1/2  sup νt,v (R1,1 − ρn,d )2 + νt,v (R1,2 − qn,d )2 ≤ K(D, β0 , β1 , κ, λ) . ν1,1 R1,1 − ρn,d + R1,2 − qn,d

t,v∈[0,1]2

27



We thus can bound  1 d 1 φ(v) ≤ K(D, λ)ε′′ + K(D, β0 , β1 , κ, λ) ε′′ + √ + , dv d d 1 . which as remarked earlier is sufficient to conclude Proposition A.4, since d ≥ ε′′2

A.2

Simplifying the expression from cavity in n

Again let α = n/d, and let ψα , ψ̄α be defined in terms of α as per (A.20), (A.21). Our goal here in Appendix A.2 is to show that the expression from Proposition A.4, specifically i p 1 h √ E log EW exp u xsn + qn,d Z + ρn,d − qn,d W , (A.29) d is very close to RSGI . The difference between (A.29) and RSGI is that RSGI instead depends on ρ0 , q0 , and additionally for the interpolators, the mollifier u(·) in (A.29) is replaced by the hard step function 1{x ≥ κ} in RSGI . Here we will do the work in this argument that is common to the proof for the interpolators and posterior, specifically Propositions A.9, A.10, A.11. The additional work for the interpolators that is needed to replace the mollifier u(·) in (A.29) by the hard step function 1{x ≥ κ} in RSGI is done in Propositions A.14 and A.13, presented later in Appendix A.3. Beginning with the argument, consider the following system of four variables and four equations: r = ψα (q, ρ) , r̄ = ψ̄α (q, ρ) , ρ =

1 r r + ,q= . 2 2β + r − r̄ (2β + r − r̄) (2β + r − r̄)2

(A.30)

We will establish that ρn,d , qn,d approximately satisfy (A.30). This is where we use the careful choice of r, r̄ in terms of ψα , ψ̄α from (A.22), (A.23), which lets us complete the cavity in n argument. In particular, in Appendix D.2, we will complete the cavity in n and also execute the cavity in d to establish the following. Lemma A.8. For all t ∈ [0, 1], we have    d d 1  νt θd1 θd2 , νt (θd1 )2 ≤ K ε′′ + √ , dt dt d

(A.31)

where K depends on D, β0 , β1 , λ. With Lemma A.8 in hand, we now establish: Proposition A.9. The following system of equations is satisfied by ρn,d , qn,d : r = ψα (q, ρ) , r̄ = ψ̄α (q, ρ) , ρ =

1 r r + + δ1 , q = + δ2 , (A.32) 2β + r − r̄ (2β + r − r̄)2 (2β + r − r̄)2

where  1  δ1 , δ2 ≤ KA.9 ε′′ + √ , d

(A.33)

for KA.9 depending on D, β0 , β1 , λ. Here the errors |δ1 |, |δ2 | are upper bounded by a quantity independent of α = nd .1 Moreover, as noted earlier, we have r ≥ r̄. 1

The upper bound only depends on α0 and is monotonically increasing, so we can apply the bound α0 ≤ 2.

28

Proof of Proposition A.9. Consider ν1 (R1,1 ), ν1 (R1,2 ). By symmetry between coordinates,   ν1 (R1,1 ) = ν1 (θd1 )2 , ν1 (R1,1 ) = ν1 θd1 θd2 . Next observe that for the measure √ ν0 , the d-th coordinate completely decouples from the others. Moreover, the restriction 1{θ1 / d ∈ I} only concerns the first coordinate θ1 . Thus, letting Zt denote the relevant normalizing constant, we have P  √  R 2 2 + θ √rY − r−r̄ θ 2 dθ θd 1{θ1 / d ∈ I} exp u S (θ) − β∥θ∥ i,0 d  i≤n 2 d ν0 (θd1 )2 = E Zt  √  R 2 2β+r−r̄ 2 θd exp θd rY − θd dθd 2  √  =E R . 2 dθ exp θd rY − 2β+r−r̄ θ d d 2 Similarly, we have R  ν0 θd1 θd2 = E

=E

     1 2 √ 1 2 + θ2 2 θ θd1 θd2 exp (θd1 + θd2 ) rY − 2β+r−r̄ dθd θd d d 2      R √ 2 2 exp (θd1 + θd2 ) rY − 2β+r−r̄ θd1 + θd2 dθd1 θd2 2   √ R 2 dθ !2 θ θd exp θd rY − 2β+r−r̄ d d 2   √ . R 2 dθ exp θd rY − 2β+r−r̄ θ d d 2

The above are univariate Gaussian integrals w.r.t. θd . As detailed on p. 219 of [Tal10], letting Z be a centered univariate Gaussian and s ∈ R, we obtain from Gaussian Integration by Parts that               E ZesZ ] = s E Z 2 E esZ , E Z 2 esZ ] = E Z 2 E esZ + s2 E Z 2 E esZ , where expectation is w.r.t. the law of Z. We obtain !2    2 2 E Z 2 esZ ]    E ZesZ ] 2     = E Z 2 ] + s2 E Z 2 2 . = s E Z , E esZ E esZ  √ 1 We apply the above display with s = rY , Z ∼ N 0, 2β+r−r̄ to compute the quantity inside the   2 1 2 1 expectation in the integrals equaling ν0 θd θd , ν0 (θd ) . This yields  √  R 2 dθ !2 θd exp θd rY − 2β+r−r̄ θ d  d 2 rY 2  √  ν0 θd1 θd2 = E = , R (2β + r − r̄)2 exp θd rY − 2β+r−r̄ θd2 dθd 2  √  R 2 2β+r−r̄ 2 θ exp rY − θ θ d  d d dθd 2 rY 2 1  √  = + . ν0 (θd1 )2 = E R 2β + r − r̄ (2β + r − r̄)2 exp θ rY − 2β+r−r̄ θ2 dθ d

2

d

d

Finally taking expectation over Y ∼ N (0, 1), we establish that     1 r 1 2 1 2 ν (θ ) − ν (θ ) , ρ = ν1 (θd1 )2 = + + 1 0 d d 2β + r − r̄ (2β + r − r̄)2      r q = ν1 θd1 θd2 = + ν1 θd1 θd2 − ν0 θd1 θd2 . 2 (2β + r − r̄)     The result follows by applying Lemma A.8 to control ν1 (θd1 )2 −ν0 (θd1 )2 , ν1 θd1 θd2 −ν0 θd1 θd2 . 29

In Appendix E.1, we will establish: Proposition A.10. For all (α, β, x) ∈ [0, α0 ] × [β0 , β1 ] × [−1, 1], any solution (q, ρ, r, r̄) of (A.30) 1 must satisfy (q, ρ) ∈ (0, Csol ) × C1sol , Csol and ρ−q < Csol . We will use Proposition A.9 and A.10 to establish the following. Proposition A.11. Consider any (α, β, x) ∈ [0, α0 ] × [β0 , β1 ] × [−1, 1]. Then there is a 0 < ε′′A.11 ≤ 1 ε′ depending on u, β0 , β1 , λ satisfying the following. For all ε′′ ≤ ε′′A.11 and all d ≥ dA.11 ≥ ε′′2 , for n any n ≥ 1 such that d = α ∈ [0, α0 ], we have i   h 1 (qn,d , ρn,d ) ∈ 0, Csol × , Csol Csol

and

1 ≤ Csol , ρn,d − qn,d

Furthermore, in the case of the interpolators, supposing the mollifier u satisfies the conditions of Theorem A.1 in terms of some parameter ε′ , we have |δ1 |, |δ2 | ≤ ε′ where δ1 , δ2 are the errors from Proposition A.9 for ρn,d , qn,d . Proof. By definition of (A.32), Proposition A.9, and as α ≤ 2, for suitable ε′′A.11 depending on 1 , we have the bounds ρn,d , qn,d ≤ Kbound (D, β0 , β1 , λ) u, β0 , β1 , λ and ε′′ ≤ ε′′A.11 and d ≥ dA.11 ≥ ε′′2 for some Kbound (D, β0 , β1 , λ) > 0 (dependence on D comes through u). Without loss of generality, suppose Kbound (D, β0 , β1 , λ) > Csol . Consider h i Du := − Kbound (D, β0 , β1 , λ), Kbound (D, β0 , β1 , λ) h i × − Kbound (D, β0 , β1 , λ), Kbound (D, β0 , β1 , λ) , n  1 o   1 D := (q, ρ) ∈ 0, Csol × , Csol , < Csol . Csol ρ−q Let R(q, ρ, α, β) :=

c :=

ρ−

1 ψα (q, ρ) − 2 , 2β + ψα (q, ρ) − ψ̄α (q, ρ) 2β + ψα (q, ρ) − ψ̄α (q, ρ) ! ψα (q, ρ) q− 2 , 2β + ψα (q, ρ) − ψ̄α (q, ρ) inf

(q,ρ)∈Du −D,α∈[0,2],β∈[β0 ,β1 ]

R(q, ρ, α, β) ≤

inf

(q,ρ)∈Du −D,α∈[0,α0 ],β∈[β0 ,β1 ]

R(q, ρ, α, β) .

Note R(q, ρ, α, β) is continuous in q, ρ, α, β. Moreover in Du − D, by Proposition A.10, we have that R(q, ρ, α, β) > 0 pointwise; else we obtain a solution to (A.30) outside D, contradicting Proposition A.10. Thus as (q, ρ, α, β) : (q, ρ) ∈ Du − D, α ∈ [0, 2], β ∈ [β0 , β1 ] is compact, c > 0. Hence in the case of interpolators, we may take ε′′A.11 = ε′′A.11 (u, β0 , β1 , λ) small enough, and 1 dA.11 (ε′′ ) ≥ ε′′2 large enough in terms of KA.9 (D, β0 , β1 , λ) and ε′ (recall that D depends solely on  ′ u) so that |δ1 |, |δ2 | from (A.33) are each strictly less than min 2c , ε2 . In the case of the posterior, we ignore dependence in ε′ in the above definitions of ε′′A.11 = ε′′A.11 (u, β0 , β1 , λ) and dA.11 (ε′′ ), now guaranteeing |δ1 |, |δ2 | < c/2. By Proposition A.9, it follows that ∥R(qn,d , ρn,d , α, β)∥ < c. Since (qn,d , ρn,d ) ∈ Du , by definition of c, it follows that (qn,d , ρn,d ) ∈ D. The conclusion of this Proposition follows. 30

Note this argument is where we use that |δ1 |, |δ2 | from (A.33) are upper bounded independently of α. In particular this lets us take ε′′A.11 = ε′′A.11 (u, β0 , β1 , λ) independently of α, which is crucial as we will use the same ε′′ across n = 1, 2, . . . , d when computing the expected free energy. The compactness supplied from the above Proposition is crucial in the arguments to follow in Propositions A.14 and A.13. The above work lets us interpolate one datapoint at a time and compute the difference in free energy, as we do in Appendix A.3. The last step we need is to compute ϕβ,0,d (I), which is the base case arising when all n datapoints have been accounted for. Lemma A.12. For any ε > 0, for some ε′′A.12 , dA.12 > 0 depending on ε, we have for all ε′′ ≤ ε′′A.12 and d ≥ dA.12 that   1 ϕβ,0,d (I) − RS0,I (0) + log(2πe) ≤ ε . 2 Proof. A direct computation gives that ϕβ,0,d (I) =

  log d 1 π log − βr(x) + O + ε′′ , 2 β d

where we recall the definition of r(x) from (A.10). Here O(·) only hides universal constants. Now ∂F note when α = 0, ∂F ∂q = 0 implies q = 0, hence q0 (0) = 0. Also, considering ∂ρ = 0 when α = 0 1 . Therefore implies that ρ0 (0) = 2β RS0,I (0) +

  1 1 1 π log(2πe) = log 2πeρ0 (0) − β ρ0 (0) + r(x) = log − βr(x) . 2 2 2 β

For suitable ε′′ ≤ ε′′A.12 (ε) and d ≥ dA.12 (ε), this implies the Lemma.

A.3

Proof of Theorem A.1 for interpolators

We now complete the proof of Theorem A.1 by summing Proposition A.4 and the results of Appendix A.2 across all datapoints i, 1 ≤ i ≤ n. We first handle the necessary steps to replace the mollifier u(·) in (A.29) by the hard step function 1{x ≥ κ} in RSGI . Specifically, this is done in the following Propositions A.13 and A.14, which establish that the impact of the mollifier u(·) is insignificant. At a high level, this is because the conditions of Theorem A.1 imply u(·) is an accurate approximation to 1{x ≥ κ}. Proposition A.13. Consider any δ > 0 and let κ+ = max{κ, 0}. Then there exists ε′A.13 > 0 depending on δ, Csol , κ+ , λ such that the following holds. Suppose u ≤ 0 satisfies for some ε′ ≤ ε′A.13 that: 1. u(x) = 0 for x ≥ κ, 2. exp u(x) ≤ ε′ for all x ≤ κ − ε′ . 1 Then for any (q, ρ) such that (q, ρ) ∈ [0, Csol ] × [ C1sol , Csol ] and ρ−q ≤ Csol , and any S ∼ D, we have h h i i √ √ √ √ E log EW exp u xS + qZ + ρ − qW − E log PW xS + qZ + ρ − qW ≥ κ ≤ δ.

31

  √ √ √ √ Proof. Note EW exp u xS + qZ + ρ − qW ≥ PW xS + qZ + ρ − qW ≥ κ . Breaking into √ √ cases on the value of xS + qZ + ρ − qW and using that exp u(x) ≤ ε′ for all x ≤ κ − ε′ and u ≤ 0, we thus may rewrite h h i i √ √ √ √ 0 ≤ E log EW exp u xS + qZ + ρ − qW − E log PW xS + qZ + ρ − qW ≥ κ "  # √ √ √ √ PW xS + qZ + ρ − qW ≥ κ + ε′ + PW xS + qZ + ρ − qW ∈ (κ − ε′ , κ)  = E log √ √ PW xS + qZ + ρ − qW ≥ κ "  !# √ √ ε′ + PW xS + qZ + ρ − qW ∈ (κ − ε′ , κ)  = E log 1 + √ √ PW xS + qZ + ρ − qW ≥ κ " # √ √ ε′ + PW xS + qZ + ρ − qW ∈ (κ − ε′ , κ)  ≤E , √ √ PW xS + qZ + ρ − qW ≥ κ+ where we used the inequality log(1 + x) ≤ x and replaced κ by κ+ in the denominator by monotonicity. We next bound    h κ − ε′ − sx − z √q κ − sx − z √q i √ √ ′ √ √ PW xS + qZ + ρ − qW ∈ (κ − ε , κ) = PW W ∈ , ρ−q ρ−q ′ 1 ε · sup √ exp(−w2 /2) ≤√ ρ − q w∈R 2π 1/2

≤ Csol ε′ . Also note that as ρ − q ≥ C1sol , E

h PW

1  √ √ xS+ qZ+ ρ−qW ≥κ+

i

is a continuous function of q, ρ and

only depends on q, ρ, κ+ , λ. Hence we may upper bound its value in the compact region (q, ρ) ∈ 1 [0, Csol ] × [ C1sol , Csol ], ρ−q ≤ Csol as follows: # " 1  ≤ K(Csol , κ+ , λ) . sup E √ √ + PW xS + qZ + ρ − qW ≥ κ (q,ρ)∈[0,Csol ]×[ 1 ,Csol ], 1 ≤Csol Csol

ρ−q

It follows that h h i i √ √ √ √ 0 ≤ E log EW exp u xS + qZ + ρ − qW − E log PW xS + qZ + ρ − qW ≥ κ " # 1 1/2  ≤ ε′ (Csol + 1) E √ √ PW xS + qZ + ρ − qW ≥ κ+  1/2 ≤ K(Csol , κ+ , λ) Csol + 1 ε′ . Thus we may take ε′A.13 = ε′A.13 (δ, Csol , κ+ , λ) ≤

δ . 1/2 K(Csol ,κ+ ,λ)(Csol +1)

We next use the uniqueness of the solution (q0 , ρ0 ) to (A.11), given by Proposition A.3, to control the difference of (qn,d , ρn,d ) vs. (q0 , ρ0 ). Proposition A.14. Consider any (β, x) ∈ [β0 , β1 ]×[−1, 1]. Given any δ > 0, there exists ε′A.14 > 0 depending on δ such that the following holds. Consider any u ≤ 0 satisfies the conditions of Theorem A.1 with ε′ = ε′A.14 . Then for all ε′′ ≤ ε′′A.11 and all d ≥ dA.11 , and for any n ≥ 1 such that n d = α ∈ [0, α0 ], the following holds. Defining q0 (α, β, x), ρ0 (α, β, x) as per Proposition A.3, qn,d − q0 (α, β, x)

and 32

ρn,d − ρ0 (α, β, x) ≤ δ .

Proof. Suppose for the sake of contradiction that there exists a δ > 0 and a sequence ε′k → 0, a sequence of functions uk such that uk (x) = 0 for x ≥ κ and exp uk (x) ≤ ε′k for all x ≤ κ − ε′k , ε′′k ≤ ε′′A.11 (uk , β0 , β1 , λ), dk ≥ dA.11 (ε′′k ), nk with ndkk = αk ∈ [0, α0 ], such that max

n

qnk ,dk − q0 (αk , β, x) , ρnk ,dk − ρ0 (αk , β, x)

o

≥ δ.

(A.34)

By Proposition A.11, since ε′′k ≤ ε′′A.11 (uk , β0 , β1 , λ), dk ≥ dA.11 (ε′′k ), we have i   h 1 1 (qnk ,dk , ρnk ,dk ) ∈ 0, Csol × , Csol , ≤ Csol . Csol ρnk ,dk − qnk ,dk

(A.35)

Also, letting δ1,k , δ2,k be the corresponding δ1 , δ2 from the system (A.32) for uk , αk , qnk ,dk , ρnk ,dk , we have by the second part of Proposition A.11 that |δ1,k |, |δ2,k | ≤ ε′k . Thus (qnk ,dk , ρnk ,dk ) lie in a set bounded independently of k, so we may extract a bounded subsequence (αk , qnk ,dk , ρnk ,dk ) converging to some limit (α⋆ , q ⋆ , ρ⋆ ). Upon reindexing, as ε′k → 0, we may suppose that (ε′k , αk , qnk ,dk , ρnk ,dk , δ1,k , δ2,k ) → (0, α⋆ , q ⋆ , ρ⋆ , 0, 0) .

(A.36)

i   h 1 1 (q ⋆ , ρ⋆ ) ∈ 0, Csol × , Csol , ⋆ ≤ Csol . Csol ρ − q⋆

(A.37)

Thus, we have

By (A.34), and as limk→∞ q0 (αk , β, x) = q0 (α⋆ , β, x), limk→∞ ρ0 (αk , β, x) = ρ0 (α⋆ , β, x) by Proposition A.3, it follows that n o max q ⋆ − q0 (α⋆ , β, x) , ρ⋆ − ρ0 (α⋆ , β, x) ≥ δ . (A.38) By (A.35), (A.36), (A.37), in particular as ρnk ,dk − qnk ,dk , ρ⋆ − q ⋆ ≥ C1sol , we also have the existence of the following limit and the following equality: h i  p √ lim EW f (W )1 xS + qnk ,dk Z + ρnk ,dk − qnk ,dk W ≥ κ k→∞ (A.39) i h p p  = EW f (W )1 xS + q ⋆ Z + ρ⋆ − q ⋆ W ≥ κ .  Next, letting f (W ) ∈ 1, W, W 2 − 1 and using our conditions on uk , h i p √ EW f (W ) exp uk xS + qnk ,dk Z + ρnk ,dk − qnk ,dk W h i  p √ − EW f (W )1 xS + qnk ,dk Z + ρnk ,dk − qnk ,dk W ≥ κ h     i p √ ≤ ε′k EW f (W ) + EW f (W )1 xS + qnk ,dk Z + ρnk ,dk − qnk ,dk W ∈ κ − ε′k , κ Z κ2   ′ = εk EW f (W ) + |f (W )|φ0 (W ) dW κ1

  ≤ ε′k EW f (W ) + √

Kε′k ρnk ,dk − qnk ,dk

1/2

≤ KCsol ε′k ,

(A.40) 33

where we let κ1 =

√ √ κ−sX− qnk ,dk Z κ−ε′k −xS− qnk ,dk Z ε′k √ √ , κ = = κ 1 + √ρ . 2 ρnk ,dk −qnk ,dk ρnk ,dk −qnk ,dk nk ,dk −qnk ,dk

Here in the last

step we used that |f (y)|φ0 (y) ≤ K uniformly for all y ∈ R. In light of (A.39), (A.40), since Csol is independent of the mollifier and as ε′k → 0, it follows that the following limit exists and can be computed as: h i p √ lim EW f (W ) exp uk xS + qnk ,dk Z + ρnk ,dk − qnk ,dk W k→∞ h i p p  = EW f (W )1 xS + q ⋆ Z + ρ⋆ − q ⋆ W ≥ κ .

(A.41)

(Note this argument does not assume anything a-priori about the limit of the uk ’s; the existence of q ⋆ , ρ⋆ and (A.40) shows that the limit exists.) Now consider the system (A.32), which we know each (uk , αk , qnk ,dk , ρnk ,dk , δ1,k , δ2,k ) solves by Proposition A.9. Since ρ⋆ − q ⋆ > C1sol , we then may rewrite ψα , ψ̄α and hence r, r̄ as per (A.22), (A.23). Since δ1,k , δ2,k → 0, we may now take limits k → ∞ on both sides of (A.32) (where now r, r̄ are given as per (A.22), (A.23)), since the denominators of (A.32) are bounded away from 0 independently of k as r ≥ r̄ by Lemma E.3 and as β ≤√β1 . Using (A.41) to compute the limit of the κ−xS− qZ expressions involving the uk ’s and letting κ′ = √ρ−q , it follows that (q ⋆ , ρ⋆ ) solves the system h φ (κ′ ) 2 i α⋆ 0 E ρ−q N (κ′ ) 1 r ρ= + 2β + r − r̄ (2β + r − r̄)2 r=

, ,

h κ′ φ (κ′ ) i α⋆ 0 E , ρ−q N (κ′ ) r q= . (2β + r − r̄)2 r̄ =

(x) Manipulating the above system and using the identity A′ (x) = A(x)2 − xA(x) where A(x) = φN0(x) denotes the Inverse Mills’ Ratio E.5), it follows that (q ⋆ , ρ⋆ ) solves (A.11).  1 (see Lemma  1 ⋆ ⋆ ⋆ ⋆ Since (q , ρ ) ∈ [0, Csol ]× Csol , Csol and ρ⋆ −q ⋆ ≤ Csol , and as (q , ρ ) solves (A.11), uniqueness given by Proposition A.3 now implies that ρ⋆ = ρ0 (α⋆ , β, x) and q ⋆ = q0 (α⋆ , β, x). This contradicts (A.38).

We now complete the proof of Theorem A.1. As Csol only depends on β0 , β1 , we may let H1 = H1 (β0 , β1 , κ, λ) denote an upperbound on the norm of  the gradient of the C ∞ function √ √ E log PW xS + qZ + ρ − qW ≥ κ over the compact set (x, ρ, q) : −1 ≤ x ≤ 1, 0 ≤ q ≤ 1 Csol , C1sol ≤ ρ ≤ Csol , ρ−q ≤ Csol . Here the function is C ∞ in this compact domain as ρ−q ≥ C1sol . Dependence on λ arises as the law of S depends on λ. Without loss of generality suppose H1 ≥ 1. By Proposition A.3, we may also let H2p= H2 (β0 , β1 , κ, λ) p be an upper bound on the gra  dient of the C ∞ function E log PW xS + q0 (α, β, x)Z + ρ0 (α, β, x) − q0 (α, β, x)W ≥ κ on (α, β, x) ∈ [0, α0 ] × [β0 , β1 ] × [−1, 1]. Here the function is C ∞ in this compact domain as ρ0 (α, β, x) − q0 (α, β, x) ≥ 2C1sol by Proposition A.3. Now, we let ε , 12H1 n o ε′ = ε′ (ε, β0 , β1 , κ, λ) := min ε′A.13 (δ, Csol , κ+ , λ), ε′A.14 (δ) , n o ε εRS0 = εRS0 (ε, u, β0 , β1 , κ, λ) := min ε′′A.11 (u, β0 , β1 , λ), , ε′′A.12 (ε/4) . 12KA.4 (D, β0 , β1 , κ, λ) δ = δ(ε, β0 , β1 , κ, λ) :=

34

Consider any 0 < ε′′ ≤ εRS0 as in the statement of this Theorem. (Note dependence on ε′ in εRS0 comes implicitly through u.) Finally we let n o 12H2 d(ε, ε′′ , β0 , β1 , κ, λ) := max dA.11 (ε′′ ), , dA.12 (ε/4), dA.4 (ε′′ ) . ε Consider any d ≥ d(ε, ε′′ , β0 , β1 , κ, λ) as in the statement of this Theorem. We now consider any n such that nd ∈ [0, α0 ]. Next, we consider any 1 ≤ i ≤ n and let α = di . In what follows, we apply the results from Appendices A.1, A.2 with i in place of n. By Proposition A.4 and as d ≥ dA.4 (ε′′ ) and ε , we have ε′′ ≤ 12KA.4 (D,β 0 ,β1 ,κ,λ) i p 1 h √ E log EW exp u xsi + qi,d Z + ρi,d − qi,d W d KA.4 (D, β0 , β1 , κ, λ) ′′ ε ≤ ε ≤ . d 12d

ϕβ,i,d (I) − ϕβ,i−1,d (I) −

(A.42)

By Proposition A.11, as ε′′ ≤ ε′′A.11 (u, β0 , β1 , λ) and d ≥ dA.11 (ε′′ ), we have h 1 i  1 qi,d , ρi,d ∈ [0, Csol ] × ≤ Csol . , Csol , Csol ρi,d − qi,d Consequently by Proposition A.13, as ε′ ≤ ε′A.13 (δ, Csol , κ+ , λ) and si ∼ D, h i p √ E log EW exp u xsi + qi,d Z + ρi,d − qi,d W h i p ε √ − E log PW xsi + qi,d Z + ρi,d − qi,d W ≥ κ ≤δ≤ . 12

(A.43)

Next, by Proposition A.14, we have qi,d − q0 (α, β, x) , ρi,d − ρ0 (α, β, x) ≤ δ =

ε . 12H1

Consequently h i p √ E log PW xsi + qi,d Z + ρi,d − qi,d W ≥ κ h p p i − E log PW xsi + q0 (α, β, x)Z + ρ0 (α, β, x) − q0 (α, β, x)W ≥ κ √ ε 2 ε ≤ · H1 < . 12H1 8 Finally, observe that for any α′ ∈

 i−1 d

 , di , α − α′ ≤ d1 , so as d ≥ 12H2 /ε we have

h p p i E log PW xsi + q0 (α, β, x)Z + ρ0 (α, β, x) − q0 (α, β, x)W ≥ κ h p p i − E log PW xsi + q0 (α′ , β, x)Z + ρ0 (α′ , β, x) − q0 (α′ , β, x)W ≥ κ ≤ H2 ·

(A.44)

1 ε ≤ . d 12 35

(A.45)

 d RS0,I (α). Since q0 (α, β, x), ρ0 (α, β, x) By Proposition A.3, as β, x, I are fixed, we may consider dα ∂FI I solves ∂F ∂q = ∂ρ = 0, we have  d ∂FI RS0,I (α) = x, q0 (α, β, x), ρ0 (α, β, x) dα ∂α  ∂ ∂FI + x, q0 (α, β, x), ρ0 (α, β, x) · q0 (α, β, x) ∂q ∂α  ∂ ∂FI x, q0 (α, β, x), ρ0 (α, β, x) · ρ0 (α, β, x) + ∂ρ ∂α h p p i = E log PW xsi + q0 (α, β, x)Z + ρ0 (α) − q0 (α, β, x)W ≥ κ , as si  ∼ D and as independent of Z, W . The Mean Value Theorem now implies that for some i α′ ∈ i−1 d ,d , i − 1 − RS0,I d d p p i 1 h = E log PW xsi + q0 (α′ , β, x)Z + ρ0 (α′ ) − q0 (α′ , β, x)W ≥ κ . d

RS0

i

(A.46)

Combining (A.42), (A.43), (A.44), (A.45), (A.46) implies for all 1 ≤ i ≤ n,  i  i − 1  3ε ϕβ,i,d (I) − ϕβ,i−1,d (I) − RS0,I − RS0,I < . d d 8d Summing across all such i and noting there are at most n ≤ α0 d ≤ 2d such i yields   n 3ε − RS0,I (0) < . ϕβ,n,d (I) − ϕβ,0,d (I) − RS0,I d 4 Last, we combine with Lemma A.12, yielding  n 1  ϕβ,n,d (I) − RS0,I + log(2πe) < ε . d 2 This completes the proof of Theorem A.1.

A.4

Proof of Theorem A.2 for posterior

Here we prove Theorem A.2 with the same argument as the proof of Theorem A.1. Note u(x) ≥ −D(λ)(1+|x|) for u(x) from (3.17), (3.18). Next, by the same argument as the proof of Proposition A.14, we have the following for any (β, x) ∈ [β0 , β1 ] × [−1, 1]. For all ε′′ ≤ ε′′A.11 , all d ≥ dA.11 , and all n ≥ 1 with n/d = α, qn,d − q0 (α, β, x) , ρn,d − ρ0 (α, β, x) ≤ δ .

(A.47)

Here the proof follows by the same argument as the proof of Proposition A.14, using Propositions A.3, A.9, and A.11, but now we do not consider the uk or replace uk by 1{x ≥ κ} as k → ∞. Here we use that a solution to (A.30) solves (A.11) where u(·) is identical in both systems; this can be seen via Gaussian integration by parts. Now to prove Theorem A.1, we let H1 = H1 (β0 , β1 , λ) be an upper boundon the gradient of √ √ the  continuously differentiable function E log EW exp u xS + qZ + ρ − qW over the compact 1 set (x, ρ, q) : −1 ≤ x ≤ 1, 0 ≤ q ≤ Csol , C1sol ≤ ρ ≤ Csol , ρ−q ≤ Csol . Here we can verify 36

that the function is continuously differentiable as ρ − q ≥ C1sol , as u is twice differentiable, and as u(x) ≥ −D(λ)(1 + |x|). Without loss of generality suppose H1 ≥ 1. Similarly by Proposition A.3, we bound on the gradient p0 , β1 , λ) be an upper p  may let H2 = H2 (β of the continuously differentiable E log EW exp u xS+ q0 (α, β, x)Z+ ρ0 (α, β, x) − q0 (α, β, x)W on (α, β, x) ∈ [0, α0 ] × [β0 , β1 ] × [−1, 1]. Here we use that ρ0 (α, β, x) − q0 (α, β, x) ≥ 2C1sol by Proposition A.3. Now, we let ε δ = δ(ε, β0 , β1 , λ) := , 12H1 n o ε εRS0 = εRS0 (u, β0 , β1 , λ) := min ε′′A.11 (u, β0 , β1 , λ), , ε′′A.12 (ε/4) . 12KA.4 (D, β0 , β1 , λ) Considering any 0 < ε′′ ≤ εRS0 as in the statement of this Theorem, we let n o 12H2 d(ε, ε′′ , β0 , β1 , λ) := max dA.11 (ε′′ ), , dA.12 (ε/4), dA.4 (ε′′ ) . ε Consider any d ≥ d(ε, ε′′ , β0 , β1 , λ) as in the statement of this Theorem. We now consider any n such that nd ∈ [0, α0 ]. For any 1 ≤ i ≤ n, letting α = di , we obtain by Proposition A.4 that i p ε 1 h √ ≤ ϕβ,i,d (I) − ϕβ,i−1,d (I) − E log EW exp u xsi + qi,d Z + ρi,d − qi,d W . d 12d By Proposition A.11, we have h 1 i  1 , Csol , ≤ Csol . qi,d , ρi,d ∈ [0, Csol ] × Csol ρi,d − qi,d Next by (A.47), we have qi,d − q0 (α, β, x) , ρi,d − ρ0 (α, β, x) ≤ δ =

ε . 12H1

Consequently h i p √ E log EW exp u xsi + qi,d Z + ρi,d − qi,d W h p p i ε < . − E log EW exp u xsi + q0 (α, β, x)Z + ρ0 (α, β, x) − q0 (α, β, x)W 8   i 1 ′ Finally, we observe that for any α′ ∈ i−1 d , d , we have α − α ≤ d , so as d ≥ 12H2 /ε, h p p i E log EW exp u xsi + q0 (α, β, x)Z + ρ0 (α, β, x) − q0 (α, β, x)W h p p i ε − E log EW exp u xsi + q0 (α′ , β, x)Z + ρ0 (α′ , β, x) − q0 (α′ , β, x)W ≤ . 12 With the above bounds in hand, we finish identically to the proof of Theorem A.1, using Proposition A.3 and Lemma A.12. This proves Theorem A.2.

B

Concentration with respect to Gaussian Measure

The goal of this section is to show that the log-partition function concentrates with respect to the Gaussian measure, as stated in Theorem B.1 for interpolators and Theorem B.2 for the posterior below. Recall that in Theorems A.1, A.2, we computed the value of the free energy, that is the expected normalized log-partition function, corresponding to the interpolators and the posterior respectively. Here, we upgrade these statements to exponentially high-probability concentration of the normalized log-partition function about this same value. 37

Theorem B.1 for interpolators: Recall the definition of the set Ui interpolating the i-th datapoint i = 1, . . . , n (equivalently, satisfying each constraint) defined in (5.5), and consider C, the intersection of all the Ui . Recalling the definition of Si in (A.1), these sets are defined by 

d

Ui := θ ∈ R : Si ≥ κ

and

C :=

n \

Ui .

(B.1)

i=1

In Theorem A.1, we showed ϕβ,n,d (I) from (A.4) equals RSGI (n/d, β,P x), where ϕβ,n,d (I) is the nor-  √ malized logarithm of (A.3). We T now show after replacing the mollifier 1≤i≤n u si θ1 +⟨gi , θ̄⟩ / d in (A.3) by the indicator 1{ 1≤i≤n Ui }, the corresponding normalized log-partition-function concentrates exponentially around the same value. Note this hard indicator arises as we want to study the set of interpolators. We now are in a position to state Theorem B.1: Theorem B.1 (For interpolators). Consider n such that 1 ≤ n ≤ α0 d. Then for any ε > 0, there is εRSG > 0 depending on ε, α0 , β0 , β1 , κ, λ such that the following holds. For any β ∈ [β0 , β1 ], recalling the definition of C in (B.1), define for any interval I ⊆ R, Z √   Zβ (I) := 1 {θ1 / d ∈ I} ∩ C exp − β∥θ∥2 dθ , (B.2) Then for any x ∈ [−1, 1] and 0 < ε′′ ≤ εRSG , letting I = [x − ε′′ , x + ε′′ ], we have for d large enough in terms of ε, ε′′ , α0 , β0 , β1 , κ, λ that !  1 P log Zβ (I) − RSGI (n/d, β, x) ≥ ε ≤ exp − d/K , (B.3) d where RSGI is defined as in (A.13) with exp u(x) = 1{x ≥ κ}, and where K is large enough in terms of ε, α0 , β0 , β1 , κ, λ. We prove Theorem B.1 in Appendix B.1. Specifically, we must handle the non-Gaussian signal component to establish several key technical results on stochastic processes and concentration of measure, which establishes the key ‘add one constraint’ style estimate Theorem B.9. This result upper bounds the probability that a Gibbs measure places small mass on one side of a given constraint. This is done in Appendix B.1.1. To handle the non-Gaussianity, it turns out we need the width ε′′ of I to be small enough in terms of λ but independent of d. See the discussion at the start of Appendix B.1.1 for more details. Theorem B.2 for posterior: We have the following similar result for the posterior, which establishes exponential concentration of the log-partition function about RSGI . Note this Theorem does not involve hard indicators as we are interested in studying the posterior pX,y in (2.7), which is defined in terms of u(x) from (3.17), (3.18) in the GMM and logistic cases respectively. Theorem B.2 (For posterior). Consider n such that 1 ≤ n ≤ α0 d. Consider u(x) from (3.17), (3.18) in the GMM and logistic cases respectively. Then for any ε > 0, there is εRSG > 0 depending on ε, α0 , β0 , β1 , λ such that the following holds. For any β ∈ [β0 , β1 ], define for any interval I ⊆ R, Z Zβ (I) :=

n X   √ 1 θ1 / d ∈ I exp u(Si ) − β∥θ∥2 dθ , i=1

38

(B.4)

where Si is defined as in (A.1). Then for any x ∈ [−1, 1] and 0 < ε′′ ≤ εRSG , letting I = [x − ε′′ , x + ε′′ ], we have for d large enough in terms of ε, ε′′ , α0 , β0 , β1 , λ that !  1 log Zβ (I) − RSGI (n/d, β, x) ≥ ε ≤ exp − d/K , (B.5) P d where RSGI is defined as in (A.13) with this exp u(x), and where K is large enough in terms of ε, α0 , β0 , β1 , λ. We prove Theorem B.2 in Appendix B.2. Unlike Theorem B.1, the proof of Theorem B.2 is relatively direct as a consequence of concentration of Lipschitz functions of log-concave measures, as stated below. Note that in contrast, such tools are not available to prove Theorem B.1 due to the hard indicators in the Hamiltonian therein. Concentration of Lipschitz functions log-concave measures. We next introduce concentration of Lipschitz functions of strongly log-concave measures in the form given in Theorem B.3 below, originally credited to B. Maurey. See also the text of Bakry, Gentil, and Ledoux [BGL13], Chapter 5. This fact allows us to establish the high-probability concentration bounds of both Theorem B.1 and Theorem B.2. We will also use this fact crucially in proving Proposition A.6 in Appendix D.3. Note that this Theorem holds when ψ can take on the value −∞, that is when ψ is supported on a convex set that is a strict subset of Rd , as noted in (3.21) of [Tal10]. Technically, this result is written in [Tal10] when f is H-Lipschitz on all of Rd ; the result writte in [Tal10] immediately implies the following by the McShane-Whitney Extension Theorem [McS34, Whi34].   Theorem B.3 (Theorem 3.1.4, [Tal10]). Consider a measure µ(θ) ∝ exp ψ(θ) defined over a convex set C such that for some β > 0, we have  θ + θ  1 θ1 − θ 2 1 2 ψ(θ1 ) + ψ(θ2 ) − ψ ≤ −β 2 2 2

2

.

(Note ψ can take on the value −∞ in the above.) Then for any set C ⊆ Rd , Z β  1 exp d2 (x, C) dµ(x) ≤ , 2 µ(C) where d(x, C) = inf{d(x, y) : y ∈ C} is the distance from x to C. Moreover, if f is a function on Rd with Lipschitz constant H on the support of µ, i.e., for all x, y ∈ supp(µ) we have ∥f (x) − f (y)∥ ≤ H∥x − y∥, then Z Z  β 2  exp f (x) − f dµ dµ(x) ≤ 4 , 8H 2 and for all k ≥ 1 ,

B.1

2k  k Z  Z 8kH 2 f (x) − f dµ dµ(x) ≤ 4 . β

Proof of Theorem B.1 for interpolators

Here throughout Appendix B.1, we let Zβ (I) be as per (B.2). The goal of this section is to prove Theorem B.1, and to this end there are two main steps. 39

1. First, we establish high-probability concentration of the normalized logarithm of (A.3) around its expectation RSGI (n/d, β, x), and more generally, concentration of the normalized free energy around its expectation. In particular, this also applies for (B.2). This is done via a technically involved application of Bernstein’s Inequality in Lemma B.12. 2. Second, we show that the mollifier u(·) in (A.3) can be eliminated without significantly changing the expectation of its logarithm, as discussed in the proof ideas in Section 5. This is a delicate step done in Lemma B.14. With both of these steps in hand, we then prove Theorem B.1. All these steps are done in Appendix B.1.2. Preliminary steps on stochastic processes and concentration of measure necessary to prove these Lemmas – specifically, the key ‘add one constraint’ style estimate, Theorem B.9 – are presented in Appendix B.1.1. These initial steps are important but highly technical. B.1.1

Preliminaries

Central to the subsequent proof in Appendix B.1.2 will be the combination of the following Theorem B.9 and Lemma B.10. These are important results on concentration of measure that let us control hard-to-bound quantities which involve the addition of the i-th datapoint or constraint, which behave at an exponential scale. In turn, these results hinge on the key Lemma B.7 on stochastic processes, which is proved by several comparison inequalities. Although similar in spirit, the statements and proofs differ from those in Section 8.2 of [Tal11], as handling the non-Gaussian ‘signal’ component si ∼ D of the disorder requires more technical ingredients. Notably, these results hinge on being able to choose the width ε′′ of I small enough in terms of λ, but still independently of d, so that for d ≥ d(ε′′ ), various partition functions are lower bounded by exp(−O(d)) where O(·) is independent of ε′′ . In particular, see the assumption Z ≥ exp(−ad) in Theorem B.9; this assumption applies in the proof of Theorem B.1 with exponentially high probability by Lemma B.12. Recall as per Appendix A that Ksignal , βsignal depend on λ. However the following bounds depend delicately on these parameters, so here in Appendix B.1.1 and only here, we will make the dependence on Ksignal , βsignal fully explicit. We begin with several preliminary results on stochastic processes. Lemma B.4. Let ξ1 and (Xℓ′ )1≤ℓ≤L be a collection of real valued random variables where ξ1 has zero mean, and ξ1 and the collection (Xℓ′ )1≤ℓ≤L are independent (but the elements of (Xℓ′ )1≤ℓ≤L are not necessarily independent of each other). Let u1 , . . . , uL ∈ R be fixed such that |uℓ − x| ≤ ε′′ for all 1 ≤ ℓ ≤ L for some x ∈ R. Define X = (Xℓ )ℓ≤L by Xℓ := uℓ ξ1 + Xℓ′ for all 1 ≤ ℓ ≤ L. Then for all t ≥ 0, L L h i h i X X E log exp(tXℓ ) ≥ E log exp(tXℓ′ ) . ℓ=1

ℓ=1

Proof. Fix s ∈ R. Since the function F (x) := log

PL

ℓ=1 exp(txℓ ) for x ∈ R

L is convex, we have

F (X) ≥ F (X ′ ) + ⟨∇F (X ′ ), X − X ′ ⟩ .   ′ Let u = (uℓ )L ℓ=1 . The expectation of the second term on the right-hand side is E ξ1 ⟨∇F (X ), u⟩ = 0 as ξ1 is zero mean and independent of X ′ , proving the Lemma.

40

Lemma B.5 (Slepian’s Inequality, in the form of Proposition 8.2.2, [Tal11]). Consider two jointly Gaussian families (Uℓ )ℓ≤L , (Vℓ )ℓ≤L , and suppose that for all 1 ≤ ℓ ≤ L we have E[Uℓ2 ] ≥ E[Vℓ2 ], and for all 1 ≤ ℓ1 ̸= ℓ2 ≤ L we have E[Uℓ1 Uℓ2 ] ≤ E[Vℓ1 Vℓ2 ]. Then h i h i X X E log exp(tUℓ ) ≥ E log exp(tVℓ ) . 1≤ℓ≤L

1≤ℓ≤L

Lemma B.6 (Proposition 8.2.3, [Tal11]). There exists a KB.6 > 0 with the following property. √ Letting (Wℓ )ℓ≤L ∼ N (0, 1) be i.i.d., for KB.6 ≤ t ≤ log L/KB.6 , we have h i X t2 E log exp(tWℓ ) ≥ log L + . 5 ℓ≤L

We will now prove the crucial Lemma B.7 that combines the above results. Its aim is to upper bound the probability of the following event: that for L replicas, the number of ℓ, 1 ≤ ℓ ≤ L such that √1d θ1ℓ s + ⟨θ̄ℓ , ḡ⟩ ≥ κ is small. The proof is involved and spans the next several pages. We will then leverage Lemma B.7 to prove the important Theorem B.9. Lemma B.7. Consider any x ∈ [−1, 1] and 0 ≤ c1 < c2 < c3 with c3 ≥

1 c2 − c1

and

c1 ≥ x2 E[s2 ] .

(B.6)

 ′ Then there exists KB.7 ≥ 1 depending on c3 , Ksignal , βsignal and KB.7 ≥ max 12c3 Ksignal , 1 depending on c3 , Ksignal satisfying the following. Consider 0 < ε′′ ≤ K1′ and B.7

θ1 , . . . , θL ∈ Rd

such that

√ θ1ℓ / d ∈ [x − ε′′ , x + ε′′ ] ∀ 1 ≤ ℓ ≤ L .

(B.7)

Consider independent s ∼ D, ḡ ∼ N (0, Id−1 ), and a family of random variables (Xℓ )ℓ≤L defined by Xℓ =

θ1ℓ s + ⟨θ̄ℓ , ḡ⟩ √ , d

such that   c2 ≤ E Xℓ2 ≤ c3 for all 1 ≤ ℓ ≤ L Then for KB.7 ≤ t ≤

and

  E Xℓ1 Xℓ2 ≤ c1 for all 1 ≤ ℓ1 ̸= ℓ2 ≤ L .

(B.8)

log L/KB.7 , we have

 n P # ℓ ≤ L : Xℓ ≥

  t o t2  ≤ L exp − KB.7 t2 ≤ KB.7 exp − . KB.7 KB.7

(B.9)

 Consequently, for all L−1/KB.7 ≤ t′ ≤ exp − KB.7 max{1, κ}2 ,  n o  2 P # ℓ ≤ L : Xℓ ≥ κ ≤ Lt′ ≤ KB.7 t′1/KB.7 .

(B.10)

Proof. We will focus on proving (B.9). The proof of (B.10) then directly follows from (B.9). 41

Proof of (B.9). The idea behind proving (B.9) is to apply √ the second moment method to exp(tXℓ ). Specifically, let us consider t such that KB.7 ≤ t ≤ log L/KB.7 and define the following event E: E :=

n X

exp(tXℓ ) ≥ L exp

1≤ℓ≤L

n X

 t2 o

(B.11)

60c3

exp(2tXℓ ) ≤ L exp (1 + K(c3 , Ksignal ))t2

o

.

(B.12)

1≤ℓ≤L

The main claim is that  P(E) ≥ 1 − 2 exp −

  t2 − exp − t2 . K(c3 , βsignal )

(B.13)

We will prove (B.13) at the end of the proof of this Lemma. Assuming (B.13) for now, let us prove (B.9). Let U[L] denote the uniform measure on [L]. Consider any (X1 , . . . , XL ) ∈ E. By definition of E, we have from the Paley-Zygmund Inequality



Pℓ∼U [L] exp(tXℓ ) ≥

1 E 2 ℓ∼U [L]

h

h i2 i Eℓ∼U [L] exp(tXℓ ) h i. exp(tXℓ ) ≥ 4 Eℓ∼U [L] exp(2tXℓ )

(B.14)

h i P Now, we note Eℓ∼U [L] exp(tXℓ ) = L1 1≤ℓ≤L exp(tXℓ ) for any ∈ R. By definition of E, specifically (B.11), we have for t ≥ KB.7 that h i  t2   t2  Eℓ∼U [L] exp(tXℓ ) ≥ exp ≥ 2 exp ≥ 4. 60c3 80c3 Consequently, (X1 , . . . , XL ) ∈ E implies the following by (B.14) and (B.12):   t2  t o 1 n # ℓ ≤ L : Xℓ ≥ = Pℓ∼U [L] exp(tXℓ ) ≥ exp L 80c3 80c3  h i 1 ≥ Pℓ∼U [L] exp(tXℓ ) ≥ Eℓ∼U [L] exp(tXℓ ) 2  ≥ exp − (1 + K(c3 , Ksignal ))t2 . Combining the above display with (B.13), we obtain  n P # ℓ ≤ L : Xℓ ≥

 t o 2 ≥ L exp − (1 + K(c3 , Ksignal ))t 80c3

≥ P(E)   t2 − exp − t2 K(c3 , βsignal )   t2 ≥ 1 − 3 exp − . K(c3 , βsignal )  ≥ 1 − 2 exp −

42

Thus for KB.7 = KB.7 (c3 , Ksignal , βsignal ) large enough, we obtain from the above that  n  t o P # ℓ ≤ L : Xℓ ≥ ≤ L exp − KB.7 t2 KB.7  n  t o ≤ L exp − KB.7 t2 ≤ P # ℓ ≤ L : Xℓ ≥ 80c3  n  t o ≤ P # ℓ ≤ L : Xℓ ≥ ≤ L exp − (1 + K(c3 , Ksignal ))t2 80c3   t2 ≤ 3 exp − K(c3 , βsignal )  t2  ≤ KB.7 exp − , KB.7 proving (B.9). Proof of (B.10). Note by monotonicity in κ, it suffices to prove the result for κ ≥ 0. First √ suppose κ ≥ 1. Consider any t such that KB.7 κ ≤ t ≤ log L/KB.7 . As κ ≥ 1, we may apply (B.9), which yields  n  n o   t o P # ℓ ≤ L : Xℓ ≥ κ ≤ L exp − KB.7 t2 ≤ P # ℓ ≤ L : Xℓ ≥ ≤ L exp − KB.7 t2 KB.7   2 t ≤ KB.7 exp − . KB.7    3 κ2 ≤ exp − K 2 Letting t′ = exp − KB.7 t2 , we obtain for L−1/KB.7 ≤ t′ ≤ exp − KB.7 B.7 κ ,   n o 2 P # ℓ ≤ L : Xℓ ≥ κ ≤ Lt′ ≤ KB.7 t′1/KB.7 √ t Now suppose κ ∈ [0, 1]. Consider any t such that KB.7 ≤ t ≤ log L/KB.7 . Since κ ≤ 1 ≤ KB.7 , we obtain from (B.9) that  n o  n o   P # ℓ ≤ L : Xℓ ≥ κ ≤ L exp − KB.7 t2 ≤ P # ℓ ≤ L : Xℓ ≥ 1 ≤ L exp − KB.7 t2  n  t o ≤ P # ℓ ≤ L : Xℓ ≥ ≤ L exp − KB.7 t2 KB.7   2 t ≤ KB.7 exp − . KB.7  The result now follows from setting t′ = exp − KB.7 t2 as in the κ ≥ 1 case above. As remarked earlier, the result for κ < 0 follows from the κ = 0 case. This proves (B.25). Proof of (B.13). We will lower bound the probabilities of each of the events (B.11), (B.12) defining E separately and then take a Union Bound. For brevity, consider the vector X = (X1 , . . . , XL ) ∈ RL and let  θℓ s + ⟨θ̄ℓ , ḡ⟩  X X F (X) := log exp(tXℓ ) = log exp t · 1 √ . d 1≤ℓ≤L 1≤ℓ≤L First we lower bound the probability of (B.11). Note F (X) can be viewed a function √ of g′ = (s, ḡ)T , which has min{1, βsignal } strongly log-concave law. Letting wℓ (g ′ ) ∝ exp t⟨θℓ , g ′ ⟩/ d be 43

weights summing to 1 for ℓ,√1 ≤ ℓ ≤ L, a direct calculation shows that the gradient ∇g′ F of F P w.r.t. g ′ is t ℓ≤L wℓ (g ′ )θℓ / d. Also note as ḡ ∼ N (0, Id−1 ) independent of s, E[Xℓ2 ] =

(θ1ℓ )2 E[s2 ] + ∥θ̄ℓ ∥2 ∥θℓ ∥2 + (E[s2 ] − 1)(θ1ℓ )2 = . d d

√ Consequently since θ1ℓ / d ∈ I and I ⊆ [−2, 2],  (θℓ )2  ∥∇g′ F ∥2 ≤ t2 max ∥θℓ ∥2 /d = t2 max E[Xℓ2 ] + (1 − E[s2 ]) · 1 ≤ t2 (c3 + 4) . ℓ ℓ d Applying Theorem B.3 on the concentration of Lipschitz functions of strongly log-concave measures now gives for all t′ > 0,  P

    F (X) − E F (X) ≥ t′ ≤ 2 exp −

 t′2 . K(c3 , βsignal )t2

(B.15)

  We now aim to lower bound E F (X) . Define θℓ θℓ E[s] ⟨θ̄ℓ , ḡ⟩ uℓ := √1 , ξ1 := s − E[s] , Xℓ′ := 1√ + √ . d d d Note that Xℓ = uℓ ξ1 + Xℓ′ , that ξ1 has mean 0, and that ξ1 , Xℓ′ are independent as the θℓ here are fixed and as E[s], s are independent. Lemma B.4 now gives i h X   (B.16) E F (X) ≥ E log exp(tXℓ′ ) . 1≤ℓ≤L

√ 1/2 Recall E[s] ≤ E[s2 ]1/2 ≤ Ksignal , and note θ1ℓ / d ≤ 2 for all ℓ as I ⊆ [−2, 2]. Thus √ 1/2 Xℓ′ ≥ X̄ℓ − E[s] · max θ1ℓ / d ≥ X̄ℓ − 2Ksignal 1≤ℓ≤L

where

1 X̄ℓ := √ ⟨θ̄ℓ , ḡ⟩ . d

Note X̄ℓ is a centered Gaussian family. Combining the above display with (B.16) gives h i X   1/2 E F (X) ≥ E log exp(tX̄ℓ ) − 2tKsignal .

(B.17)

1≤ℓ≤L

We now observe that as ḡ ∼ N (0, Id−1 ) independent of s, (θ1ℓ )2 E[s2 ] + ∥θ̄ℓ ∥2 (θℓ )2 = E[X̄ℓ2 ] + 1 E[s2 ] , d d ℓ1 ℓ2 2 ℓ ℓ 1 2 θ θ E[s ] + ⟨θ̄ , θ̄ ⟩ θℓ1 θℓ2 E[Xℓ1 Xℓ2 ] = 1 1 = E[X̄ℓ1 X̄ℓ2 ] + 1 1 E[s2 ] . d d E[Xℓ2 ] =

√ As θ1ℓ / d ∈ [x − ε′′ , x + ε′′ ], it follows that

θ11 θ12 (θ1ℓ )2 2 − x2 d −x , d

≤ 3ε′′ and therefore  E[X̄ℓ2 ] − E[Xℓ2 ] − x2 E[s2 ]) , E[X̄ℓ1 X̄ℓ2 ] − E[Xℓ1 Xℓ2 ] − x2 E[s2 ] ≤ 3ε′′ Ksignal .

Let c̄1 := c1 − x2 E[s2 ] + 3ε′′ Ksignal , c̄2 := c2 − x2 E[s2 ] − 3ε′′ Ksignal . 44

(B.18)

It follows by (B.18) that for ε′′ ≤ K ′

1

B.7 (c3 ,Ksignal )

1 1 ≤ 12Ksignal c3 , because c2 − c1 ≥ c3 , we have

c̄2 − c̄1 ≥ (c2 − x2 E[s2 ]) − (c1 − x2 E[s2 ]) − 6ε′′ Ksignal ≥ c2 − c1 −

1 1 ≥ , 2c3 2c3

c̄1 > c1 − x2 E[s2 ] ≥ 0 ,

(B.19) (B.20)

where the last step uses the condition (B.6). Additionally by (B.18), we have E[X̄ℓ1 X̄ℓ2 ] ≤ E[Xℓ1 Xℓ2 ] − x2 E[s2 ] + 3ε′′ Ksignal = c̄1 , E[X̄ℓ2 ] ≥ E[Xℓ2 ] − x2 E[s2 ] − 3ε′′ Ksignal = c̄2 . We now consider i.i.d. Z, (Wℓ )ℓ≤L ∼ N (0, 1). By (B.19), (B.20), we may define the centered Gaussian family (Vℓ )L ℓ=1 by √ √ Vℓ := Z c̄1 + Wℓ c̄2 − c̄1 . Thus, E[Vℓ2 ] = c̄2 ≤ E[X̄ℓ2 ] and E[Vℓ1 Vℓ2 ] = c̄1 ≥ E[X̄ℓ1 X̄ℓ2 ] for ℓ1 ̸= ℓ2 . Consequently as X̄ℓ is a centered Gaussian family, the conditions of Lemma B.5 apply. Combining it with (B.17) gives h i X   1/2 E F (X) ≥ E log exp(tX̄ℓ ) − 2tKsignal 1≤ℓ≤L

h

X

≥ E log

i 1/2 exp(tVℓ ) − 2tKsignal

1≤ℓ≤L

h

X

= E log

i √ 1/2 exp(t c̄2 − c̄1 Wℓ ) − 2tKsignal .

1≤ℓ≤L

√ B.6 , where the last inequality follows Take KB.7 = KB.7 (c3 , Ksignal , βsignal ) ≥ KB.6 2c3 ≥ √K √ c̄2 −c̄1 from (B.19). Thus we may apply Lemma B.6 with t c̄2 − c̄1 in place of t, which gives that for t ≥ KB.7 = KB.7 (c3 , Ksignal , βsignal ), h i X   √ 1/2 E F (X) ≥ E log exp(t c̄2 − c̄1 Wℓ ) − 2tKsignal 1≤ℓ≤L

(c̄2 − c̄1 )t2 1/2 − 2tKsignal 5 t2 1/2 − 2tKsignal ≥ log L + 20c3 t2 ≥ log L + , 30c3 ≥ log L +

where we use (B.19) to lower bound c̄2 − c̄1 . Combining with (B.15) gives P

 X 1≤ℓ≤L

exp(tXℓ ) ≥ L exp

 t2  60c3

   t2  ≥ 1 − P F (X) ≤ E F (X) − 60c3  t2 ≥ 1 − 2 exp − . K(c3 , βsignal ) 

This lower bounds the probability of (B.11). 45

(B.21)

θℓ s+⟨θ̄ℓ ,ḡ⟩

Next, we lower bound the probability of (B.12) as follows. As each Xℓ = 1 √d has identical law and as s, ḡ are independent, h X i h  2tθ1 i h  2t⟨θ̄1 , ḡ⟩ i   √ . (B.22) E exp(2tXℓ ) = L E exp(2tX1 ) = L E exp √ 1 E exp ds d 1≤ℓ≤L Standard calculations on the MGF of a Gaussian yield  2t2 ∥θ̄1 ∥2  h 2t⟨θ̄1 , ḡ⟩ i √ = exp E exp . d d 2tθ1

Next we have √d1 s ≤

(B.23)

t2 (θ11 )2 Ksignal s2 , therefore as s2 is Ksignal -sub-Exponential, + Ksignal d

 t2 (θ1 )2 K h  2tθ1 s i  h  s2 i   t2 (θ1 )2 K signal signal 1 1 ≤ 2 exp E exp √ 1 E exp . ≤ exp d Ksignal d d √   (θℓ )2 E[s2 ]+∥θ̄ℓ ∥2 and as θ11 / d ∈ I ⊆ [−2, 2], it follows that Since c3 ≥ E Xℓ2 = 1 d

(B.24)

1 1 2 (θ1 )2 ∥θ ∥ ≤ c3 + 1 ≤ c3 + 4 . d d Combining the above bound with (B.22), (B.23), (B.24) gives h X i  t2 K(K 1 2  signal )∥θ ∥ ≤ L exp K(c3 , Ksignal )t2 , E exp(2tXℓ ) ≤ 2L exp d 1≤ℓ≤L

where we used that t ≥ KB.7 = KB.7 (c3 , Ksignal , βsignal ). Thus by Markov’s Inequality,  X  P exp(2tXℓ ) ≤ L exp (1 + K(c3 , Ksignal ))t2 ≥ 1 − exp(−t2 ) .

(B.25)

1≤ℓ≤L

This lower bounds the probability of (B.12). Combining (B.21), (B.25) proves (B.13), thus completing the proof of Lemma B.7. We also cite the following fact from [Tal11]. Technically, in [Tal11], Theorem B.8 is stated with θ in place of θ̄ in the conclusion (B.26), (B.27), however the proof is identical. Theorem B.8 (Corollary of the proof of Theorem 8.2.7, [Tal11]). Consider a concave U ≤ 0 defined on Rd , and a convex set C ⊆ Rd . For β ∈ [β0 , β1 ], define a measure GC on Rd as follows, where Z ′ denotes the appropriate normalization constant: Z   1 for all B ⊆ Rd , GC (B) = ′ 1{B ∩ C} exp U (θ) − β∥θ∥2 dθ . Z Assume that for some a > 0 we have Z ′ ≥ exp(−ad) . 1 ′ Then for some 0 ≤ c′1 < c′2 < c′3 with c′3 ≥ c′ −c ′ , where c3 only depends on a, β0 , β1 , we have that 2 1 GC satisfies:    d (B.26) GC θ : c′2 d ≤ ∥θ̄∥2 ≤ c′3 d ≥ 1 − exp − ′ , c3    d 1 2 1 2 ′ G⊗2 (θ , θ ) : |⟨ θ̄ , θ̄ ⟩| ≤ c d ≥ 1 − exp − ′ . (B.27) 1 C c3

46

Proof. This follows from the exact same proof of Theorem 8.2.7, [Tal11], except we now consider R1,1 = d1 ∥θ̄∥2 , R1,2 = d1 ⟨θ̄1 , θ̄2 ⟩ as we do throughout this paper, as per (A.16). Following the same proof as therein, it is shown that (B.26), (B.27) hold where we can take c′3 = d4′2 where  ′ 2 d = exp − 2(a + a /4β + 3) . Note now rather than following the proof of (3.26) and Theorem 3.1.11 in [Tal10], we follow the proof of Lemma B.19 and Proposition D.5, directly using the bound Z ′ ≥ exp(−ad) rather than Lemma D.9. (There is no dependence on D here as dependence on D in the proof of Lemma B.19, Proposition D.5 only came due to Lemma D.9.) The rest of the proof is identical to the proof of Theorem 8.2.7 in [Tal11]. Note that c′1 ≥ 0 as ⟨R1,2 ⟩ ≥ 0. Now we prove the following very useful result that upper bounds the probability that a Gibbs measure places small mass on one side of a random hyperplane in Rd , where the hyperplane has law identical to the constraints (the probability is over the hyperplane). We will use this result repeatedly next in Appendix B.1.2. Theorem B.9. Consider a concave U ≤ 0 defined on Rd , and a convex set C ⊆ Rd . Consider any x ∈ [−1, 1] and let I = [x − ε′′ , x + ε′′ ] for some ε′′ > 0. For β ∈ [β0 , β1 ], define a measure GC on Rd as follows, where Z ′ denotes the appropriate normalization constant: Z   √  1 for all B ⊆ Rd , GC (B) = ′ 1 B ∩ C ∩ {θ1 / d ∈ I} exp U (θ) − β∥θ∥2 dθ . Z Assume that for some a > 0 we have Z ′ ≥ exp(−ad) . ′ Then for KB.9 ≥ 2 depending on a, β0 , β1 , Ksignal , βsignal , KB.9 ≥ 2 depending on a, β0 , β1 , Ksignal , and KB.7 ≥ 1 depending on a, β0 , β1 , Ksignal , βsignal , we have the following. Consider s ∼ D and ḡ ∼ N (0, Id−1 ) such that s, ḡ are independent. If ε′′ ≤ K1′ , then GC satisfies the following B.9 implication:    d  1 If exp − ≤ t ≤ exp − KB.7 max{1, κ}2 , KB.9 4 ! (B.28) n o θ1 s + ⟨θ̄, ḡ⟩ 2 1/KB.7 √ then P GC θ : . ≥κ ≤ t ≤ 16KB.7 t d

Proof. We define for all θ, X(θ) :=

sθ1 + ⟨θ̄, ḡ⟩ √ , d

and we let Xℓ := X(θℓ ) for all replicas 1 ≤ ℓ ≤ L. First, we prove that the conditions of Lemma B.7√apply with high enough probability over GC . By Theorem B.8 applied with the convex set {θ1 / d ∈ I} ∩ C, for some 0 ≤ c′1 < c′2 < c′3 = 1 ′ c′3 (a, β0 , β1 ) with c′3 ≥ c′ −c ′ , we have (B.26), (B.27). (Crucially c3 does not explicitly depend on I, 2 1 in particular it does on ε′′ , which will be very√important in the following arguments.) √ not depend  Since GC {θ : θ1 / d ̸∈ I} = 0 (recall the constraint θ1 / d ∈ I is present in the definition of GC from this Theorem), it follows that    d √ GC θ : c′2 d ≤ ∥θ̄∥2 ≤ c′3 d , θ1 / d ∈ I ≥ 1 − exp − ′ , (B.29) c3   d  √ 1 2 1 2 ′ G⊗2 (θ , θ ) : |⟨ θ̄ , θ̄ ⟩| ≤ c d , θ / d ∈ I ≥ 1 − exp − ′ . (B.30) 1 1 C c3 47

Now note for any θℓ , θℓ1 , θℓ2 , we have as ḡ ∼ N (0, Id−1 ) independent of s,    1  ℓ1 ℓ2 E Xℓ1 Xℓ2 = θ1 θ1 E[s2 ] + ⟨θ̄ℓ1 , θ̄ℓ2 ⟩ . d √ Consequently for θℓ such that c′2 d ≤ ∥θ̄ℓ ∥2 ≤ c′3 d , (θℓ )1 / d ∈ I, we have    1 ℓ 2 (θ1 ) E[s2 ] + ∥θ̄ℓ ∥2 , E Xℓ2 = d

   2 E Xℓ2 ≤ c′3 + max |x + ε′′ |, |x − ε′′ | Ksignal ) ≤ c′3 + 4Ksignal ,   E Xℓ2 ≥ c′2 + (x2 − 3ε′′ ) E[s2 ] . √ √ Similarly for (θℓ1 , θℓ2 ) such that |⟨θℓ1 , θℓ2 ⟩| ≤ c′1 d , (θℓ1 )1 / d ∈ I , (θℓ2 )1 / d ∈ I, we have   E Xℓ1 Xℓ2 ≤ c′1 + (x2 + 3ε′′ ) E[s2 ] . Define c3 = 2c′3 + 4Ksignal , c1 := c′1 + x2 E[s2 ] +

1 1 , c2 := c′2 + x2 E[s2 ] − ′ . ′ 4c3 4c3

(B.31)

Let ′ ′ KB.9 (a, β0 , β1 , Ksignal ) = KB.7 (c3 , Ksignal ) ≥ 12c3 Ksignal .

Hence since ε′′ ≤ K ′

1

B.9 (a,β0 ,β1 ,Ksignal )

(B.32)

≤ 12c3 K1signal , it follows that for such θ satisfying the conditions

, θ2 satisfying the conditions from (B.30) from (B.29) we have c3 ≥ E[Xℓ2 ] ≥ c2 , and similarly for θ1√  ′ we have E[Xℓ1 Xℓ2 ] ≤ c1 . As c3 ≥ c3 , and as GC {θ : θ1 / d ̸∈ I} = 0, it follows that  d ≥ 1 − exp − , c3    d G⊗2 (θ1 , θ2 ) : E[Xℓ1 Xℓ2 ] ≤ c1 ≥ 1 − exp − . C c3 GC



θ : c2 ≤ E[Xℓ2 ] ≤ c3



Consider L such that 2L2 ≤ exp(d/c3 ) and define o n √ E L := (θ1 , . . . , θL ) : c2 ≤ E[Xℓ2 ] ≤ c3 ∀ ℓ , E[Xℓ1 Xℓ2 ] ≤ c1 ∀ ℓ1 ̸= ℓ2 , θℓ / d ∈ I ∀ ℓ .

(B.33) (B.34)

(B.35)

Now observe that the conditions of Lemma B.7 apply with c1 , c2 , c3 chosen in (B.31) because we have ε′′ ≤ K ′ (c31,Ksignal ) , and because c′1 ≥ 0 and therefore B.7

c2 − c1 = c′2 − c′1 −

1 1 1 ≥ ′ ≥ ′ 2c3 2c3 c3

,

c1 > c′1 + x2 E[s2 ] ≥ x2 E[s2 ] .

Applying Lemma B.7, we obtain the following. For any (θ1 , . . . , θL ) ∈ E L and t satisfying  L−1/KB.7 ≤ t ≤ exp − KB.7 max{1, κ}2 , (B.36) where KB.7 = KB.7 (c3 , Ksignal , βsignal ) ≥ 1 comes from Lemma B.7 with c3 defined in (B.31), we have    2 P # ℓ ≤ L : Xℓ ≥ κ ≤ tL ≤ KB.7 t1/KB.7 , (B.37) where probability in the above is over ḡ, s. 48

   Consider the event G⊗L (θ1 , . . . , θL ) ∈ E L : # ℓ ≥ L : Xℓ ≥ κ ≥ tL . Observe that when C    1 , . . . , θ L ) ∈ E L : # ℓ ≤ L : X ≥ κ ≥ tL ≥ 1 , we have by Linearity of Expectation, G⊗L (θ ℓ C 4 GC



θ : X(θ) ≥ κ



Z  1 = # ℓ ≤ L : X(θℓ ) ≥ κ dG⊗L C (θ) L    1 ≥ · tL · G⊗L (θ1 , . . . , θL ) ∈ E L : # ℓ ≤ L : Xℓ ≥ κ ≥ tL C L t (B.38) ≥ . 4

Next observe that    1 L L ′ G⊗L (θ , . . . , θ ) ∈ E : # ℓ ≤ L : X ≥ κ ≥ tL ℓ C    ⊗L L 1 L L ′ = G⊗L (E ) − G (θ , . . . , θ ) ∈ E : # ℓ ≤ L : X ≥ κ ≤ tL ℓ C C      d 1 L L ′ − G⊗L (θ , . . . , θ ) ∈ E : # ≥ κ ℓ ≤ L : X ≤ tL ≥ 1 − L2 exp − ℓ C c3    1 (B.39) (θ1 , . . . , θL ) ∈ E L : # ℓ ≤ L : Xℓ′ ≥ κ ≤ tL , ≥ − G⊗L C 2 √  where we apply the condition on L and use (B.33), (B.34) and that GC {θ : θ1 / d ̸∈ I} = 0. Combining (B.38), (B.39), it follows that as events we have ( )   1  G⊗L (θ1 , . . . , θL ) ∈ E L : # ℓ ≤ L : Xℓ′ ≥ κ ≤ tL ≤ C 4 ) (  1   (θ1 , . . . , θL ) ∈ E L : # ℓ ≤ L : Xℓ′ ≥ κ ≥ tL ≥ ⊆ G⊗L C 4 n o   t ⊆ GC θ : X(θ) ≥ κ ≥ . (B.40) 4 Now by Markov’s Inequality and (B.37), we obtain !  1 (θ1 , . . . , θL ) ∈ E L : # ℓ ≤ L : Xℓ′ ≥ κ ≤ tL ≥ P G⊗L C 4 " #    ≤ 4 E G⊗L (θ1 , . . . , θL ) ∈ E L : # ℓ ≤ L : Xℓ′ ≥ κ ≤ tL C 

Z =4



 n  o  1 (θ1 , . . . , θL ) ∈ E L P 1 # ℓ ≤ L : Xℓ′ ≥ κ ≤ tL dG⊗L C (θ) 2

≤ 4KB.7 t1/KB.7 . Therefore by the contrapositive of (B.40), combining with the above display yields 

P GC



θ : X(θ) ≥ κ



  1  t 1 L L ′ ≤ ≤ P G⊗L (θ , . . . , θ ) ∈ E : # ℓ ≤ L : X ≥ κ ≤ tL ≥ ℓ C 4 4 2

≤ 4KB.7 t1/KB.7 .

!

(B.41)

Hence, (B.36) implies (B.41). The result follows by taking 4t in place of t in this implication, noting 2 41/KB.7 ≤ 4 as KB.7 ≥ 1, and recalling this holds for any L with 2L2 ≤ exp(d/c3 ). 49

Now we find a generic way to upper bound random variables satisfying the implication in Theorem B.9. This is analogous to Lemma 8.3.8 of [Tal11], but we require some more generality. Lemma B.10. Consider a random variable V ≥ 0 and assume that for certain numbers C0 , C1 , C2 , D0 ≥ 1 and some a ∈ R we have P(V ≥ exp(ad)) = 0 , D0 ≤ t ≤ exp(d/C0 ) =⇒ P(V ≥ t) ≤ C1 t−1/C2 . Then we have E[V γ ] ≤ KB.10 = D0 + 2C1 C2 + C1

n 1 o 1 , γ = max 1, ∈ (0, 1) . 2C2 aC0 C2

for

Proof. As P(V ≥ exp(ad)) = 0, we may write Z ∞ γ tγ−1 P(V ≥ t) dt E[V ] = γ 0

Z D0 =γ

γ−1

t

Z exp(d/C0 ) P(V ≥ t) dt + γ

0

γ−1

t

Z exp(ad) P(V ≥ t) dt + γ

D0

tγ−1 P(V ≥ t) dt .

exp(d/C0 )

We upper bound these three integrals as follows: • First, we have that Z D0 γ

D0

tγ−1 P(V ≥ t) dt ≤ tγ

0

0

= D0γ ≤ D0 .

• Second, as γ ≤ 2C1 2 ≤ 1 we have by the upper bound on P(V ≥ t) that Z exp(d/C0 ) γ

γ−1

t

Z exp(d/C0 ) P(V ≥ t) dt ≤ γC1

D0

tγ−1−1/C2 dt

D0 −1/2C2

≤ C1 · 2C2 D0 ≤ 2C1 C2 , where we used that D0 ≥ 1. • Third, we have by the upper bound on P(V ≥ t) that Z exp(ad) γ

t

γ−1

Z exp(ad)

tγ−1−1/C2 dt

P(V ≥ t) dt ≤ γC1

exp(d/C0 )

exp(d/C0 )

 ≤ γC1 exp − ≤

d  C0 C2  d

γC1 exp − C0 C2

as t ≥ exp(d/C0 ) and γ ≤ aC10 C2 . The Lemma is proved. 50

γ

Z exp(ad)

tγ−1 dt

exp(d/C0 )

· exp(aγd) ≤ C1 ,

We will also need the following useful fact: Lemma B.11. Consider any U ≤ 0 defined on Rd , an interval I ⊆ R, and a convex set C ⊆ Rd . For β ∈ [β0 , β1 ], define a measure GC on Rd as follows, where Z ′ denotes the appropriate normalization constant: Z   √ 1 d 1{θ1 / d ∈ I} exp U (θ) − β∥θ∥2 dθ . for all B ⊆ R , GC (B) = ′ Z B∩C Assume that for some a > 0 we have Z ′ ≥ exp(−ad) . Then for r ≤ rB.11 where rB.11 depends on a, β0 , β1 , we have  √  GC ∥θ̄∥ ≤ r d ≤ exp(−d) . Proof. By the condition on Z ′ and as U ≤ 0, we have Z √ √    2 GC ∥θ̄∥ ≤ r d ≤ exp(ad) √ 1{θ1 / d ∈ I} exp − β∥θ∥ dθ ∥θ̄∥≤r d

Z  β d/2 Z √  2 2 1{θ1 / d ∈ I} exp(−βθ1 ) dθ · = exp(ad) √ exp − β∥θ̄∥ dθ̄ π R ∥θ̄∥≤r d √   ≤ exp d(a + K(β)) PZ ∥Z∥ ≤ r d ,  where Z ∼ N 0, β1 Id−1 . Standard bounds on the concentration of the norm of multivariate Guassians (see e.g. Theorem 3.1.1, [Ver18]) √ now give that for suitable rB.11 (a, β0 , β1 ) and for all r ≤ rB.11 (a, β0 , β1 ), we have PZ ∥Z∥ ≤ r d ≤ exp − d(a + K(β) + 1) . This yields the desired conclusion. B.1.2

Proof of Theorem B.1

We now have established the technical ingredients necessary to prove Theorem B.1. As discussed earlier, the proof proceeds through two key Lemmas. In Lemma B.12 we show concentration of the normalized free energy around its expectation for quite general mollified Hamiltonians. This generality is needed as we apply this result for both u(·) from Theorem A.1 and with exp u(x) = 1{x ≥ κ}. In Lemma B.14 we show the impact of the mollifier on the expectation is small. Since these arguments will involve changing the mollifier u(·), overloading notation, we make the following definitions for only Appendix B.1. For any function u(·) and interval I ⊆ R, let Z   X √ Hu (θ) := u(Si ) and Zu (I) := 1{θ1 / d ∈ I} exp Hu (θ) − β∥θ∥2 dθ . (B.42) 1≤i≤n

We do not vary β in this proof, so there will be no ambiguity in the above definitions. Next, for A ≥ 0, we define the truncated logarithm logA x := max(log x, −A) .

(B.43)

Several useful properties of the truncated logarithm, Lemmas F.2, F.3, F.4, are presented in Appendix F.1. The reason we introduce the truncated logarithm is to truncate various partition functions Z by logad (Z), allowing us to use Theorem B.9. Later in the proof when we conclude 51

Theorem B.1, we will choose a large enough so the truncation is of no impact with high probability using the concentration result Lemma B.12. Our first step is to show strong concentration properties of logad Zu (I). From here on out, following the convention of this paper excluding Appendix B.1.1, we express dependence on the quantities Ksignal , βsignal through the signal-to-noise ratio λ governing them. Lemma B.12. Consider any κ ∈ R, β ∈ [β0 , β1 ], and a > 0. Then whenever u ≤ 0 is concave with u(x) = 0 for all x ≥ κ, we have the following. Consider any interval I = [x − ε′′ , x + ε′′ ] with ′ ε′′ ≤ K1′ where KB.9 depends on a, β0 , β1 , λ. Then for all d large enough in terms of a, β0 , β1 , κ, λ B.9 and all n ≤ 2d, we have for all t > 0,  1 h1 i    d P logad Zu (I) − E logad Zu (I) ≥ t ≤ 2 exp − min{t2 , t} . (B.44) d d K(a, β0 , β1 , κ, λ) Note Lemma B.12 applies for any concave u ≤ 0 with u(x) = 0 for all x ≥ κ, and does not impose further regularity assumptions on u. In particular, it applies for u given by exp u(x) = 1{x ≥ κ}. Proof. Let Fi be the σ-algebra generated by the randomness of the first i samples. Let Ei denote conditional expectation w.r.t. Fi . Let h1 i i h1 logad Zu (I) − Ei−1 logad Zu (I) . (B.45) Xi (I) := Ei d d Note Xi (I) is Fi -measurable and Ei−1 [Xi (I)] = 0. First, we claim it suffices to prove for all i that: h  Ei−1 exp

i d |Xi (I)| ≤ 2, K(a, β0 , β1 , κ, λ)

(B.46)

To this end, note that upon establishing (B.46) for all 1 ≤ i ≤ n, applying Bernstein’s Inequality on the Xi (I) and recalling n ≤ 2d yields ! ! n h1 i  X 1 d min{t2 , t}  P logad Zβ (I) − E logad Zβ (I) ≥ t = P . Xi (I) ≥ t ≤ 2 exp − d d K(a, β0 , β1 , κ, λ) i=1

The rest of the proof of Lemma B.12 is now devoted to showing (B.46). We define Z   X √ Wi (I) := 1{θ1 / d ∈ I} exp u(Si′ ) − β∥θ∥2 dθ ,

(B.47)

i′ ̸=i,1≤i′ ≤n

n o  Z (I)  u Y1 (I) := 1 Wi (I) ≥ exp(−ad) logad , Wi (I) n Z (I) o u Y2 (I) := 1 < exp(−ad), Wi (I) > 1 logad Zu (I) − logad Wi (I) , Wi (I) Y (I) := Y1 (I) + Y2 (I) .

(B.48) (B.49) (B.50)

Now, we begin with some preparatory work to show (B.46). Lemma B.13. We have the pointwise upper bound logad Zu (I) − logad Wi (I) ≤ Y (I) . Proof of Lemma B.13. Notice Zu (I) ≤ Wi (I) as u ≤ 0. We break into the three possible cases: 52

• If Wi (I) ≤ exp(−ad), then the left hand side is 0 so the above holds. • If Zu (I) ≤ exp(−ad) < Wi (I): If Wi (I) ≤ 1, the above follows as a consequence of Lemma F.2 on truncated logarithms. Else if Wi (I) > 1, we must have Zu (I)/Wi (I) ≤ exp(−ad) in this case, and therefore the left hand side above is upper bounded by Y2 (I). • If Zu (I), Wi (I) > exp(−ad): First suppose Wi (I) ≤ 1. Then this is again a consequence Zu (I) of Lemma F.2. Else if Wi (I) > 1 and W ≥ exp(−ad), then Zu (I) ≥ exp(−ad), and the i (I) Zu (I) < exp(−ad), the Lemma follows by definition of logarithms. Finally if Wi (I) > 1 and W i (I) left hand side above is upper bounded by Y2 (I).

This proves Lemma B.13. Notice on the randomness defining the i-th sample. Therefore we have h Wi (I) does i not depend h i i i−1 that E logad Wi (I) = E logad Wi (I) . Lemma B.13 thus gives     d Xi (I) = Ei logad Zu (I) − Ei−1 logad Zβ (I)     = Ei logad Zu (I) − logad Wi (I) − Ei−1 logad Zu (I) − logad Wi (I) i h i h ≤ Ei logad Zu (I) − logad Wi (I) + Ei−1 logad Zu (I) − logad Wi (I)     ≤ Ei |Y (I)| + Ei−1 |Y (I)| . Thus for any γ > 0, we have h h  i i i i−1 i−1 i−1 exp γ E |Y (I)| exp γ E |Y (I)| exp γd Xi (I) ≤E E h    i = exp γ Ei−1 |Y (I)| Ei−1 exp γ Ei |Y (I)| h h i i ≤ Ei−1 exp γ|Y (I)| · Ei−1 Ei exp γ|Y (I)| h i2 = Ei−1 exp γ|Y (I)| .

(B.51)

Let Ei denote expectation in the randomness defining the i-th sample, thus Ei−1 = Ei Ei = Ei Ei . Hence h h i i Ei−1 exp γY (I) = Ei Ei exp γY (I) . Now, note it suffices to show that for γ = γ(a, β0 , β1 , λ) > 0, we have h i Ei exp γY (I) ≤ KB.52 (a, β0 , β1 , κ, λ) .

(B.52)

This is because Jensen’s Inequality then implies h Ei

i h i KB.52 (a,β10 ,β1 ,κ,λ) γ exp Y (I) ≤ Ei exp γY (I) KB.52 (a, β0 , β1 , κ, λ) 

≤ KB.52 (a, β0 , β1 , κ, λ)1/KB.52 (a,β0 ,β1 ,κ,λ) ≤ 2 , hence implying (B.46) when combined with the above. If Wi (I) ≤ exp(−ad) then Y (I) = 0, and (B.52) immediately follows. Thus suppose from now on that Wi (I) ≥ exp(−ad); we will establish (B.52). Note this step is permitted because Ei is 53

only with respect to the randomness defining the i-th sample, while Wi (I) does not depend on this randomness, and hence the randomness defining Wi (I) and Ei are independent. By AM-GM, to show (B.52), it suffices to show for some γ = γ(a, β0 , β1 , κ) > 0 that h h i i Ei exp γY1 (I) , Ei exp γY2 (I) ≤ K(a, β0 , β1 , κ, λ) . (B.53) To this end, define the measure Gi on Rd by  √ Gi (θ) ∝ 1{θ1 / d ∈ I} exp

X

 u(Si′ ) − β∥θ∥2 .

1≤i′ ≤n,i′ ̸=i

Note as the normalizing constant of Gi (θ) is exactly Wi (I) by (B.47), and as u(x) = 0 for all x ≥ κ and u(x) ≤ 0, we have Z Z Zu (I) Zu (I) = exp u(Si ) dGi (θ) , 1 ≥ = exp u(Si ) dGi (θ) ≥ Gi (Si ≥ κ) . (B.54) Wi (I) Wi (I) √  P We apply Theorem B.9 with U (θ) defined by exp(U (θ)) = 1{θ1 / d ∈ I} exp i′ ̸=i u(Si′ ) . Note U ≤ 0 is concave as u ≤ 0 is concave and the Si are affine. Since Gi (Si ≥ κ) ≤ Zu (I)/Wi (I) from (B.54), and as ε′′ ≤ K1′ in the statement of this Lemma and the condition Wi (I) ≥ exp(−ad), we B.9 obtain !   Zu (I) 1 2 2 exp − d/KB.9 ≤ t ≤ exp − KB.7 max{1, κ} =⇒ P ≤ t ≤ 16KB.7 t1/KB.7 , (B.55) 4 Wi (I) ′ ′ (a, β , β , λ) ≥ 2 and K for KB.9 = KB.9 (a, β0 , β1 , λ), KB.9 = KB.9 0 1 B.7 = KB.7 (a, β0 , β1 , λ) ≥ 1.

  Upper bounding Ei exp(γY1 (I)) . We apply Lemma B.10 with   2 V = exp Y1 (I) , C0 = KB.9 , C1 = 16KB.7 , C2 = KB.7 , D0 = 4 exp KB.7 max{1, κ}2 .   n o Zu (I) Wi (I) As Zu (I) ≤ Wi (I) by (B.54), we have Y1 (I) = logad W and exp Y (I) ≤ min , exp(ad) . 1 Zu (I) i (I) Thus by (B.55), it follows that the conditions for Lemma B.10 are satisfied. Thus there is such a γ1 = γ1 (a, β0 , β1 , λ) > 0 for which h i Ei exp γ1 Y1 (I) ≤ K(a, β0 , β1 , κ, λ) . (B.56)   Upper bounding Ei exp(γY2 (I)) .

As u ≤ 0, we have the pointwise bound Z  2π d/2  . exp − β∥θ∥2 dθ = Zu (I) , Wi (I) ≤ β Rd

By definition of truncated logarithm, this implies logad Zu (I) − logad Wi (I) ≤ K(a, β0 )d . Thus for all γ > 0, we have from the definition of Y2 (I) in (B.49) that  Z (I)     u Ei exp(γY2 ) ≤ Pi < exp(−ad) exp γK(a, β0 )d + 1 . Wi (I) 54

(B.57)

 Our strategy now will be to show Pi Zu (I)/Wi (I) < exp(−ad) is exponentially small in d, and then choose γ small enough to upper bound (B.57). First, note as d > K(a, β0 , β1 , κ, λ), we have   1 ′ exp − d/KB.9 , exp(−ad) ≤ exp − KB.7 max{1, κ}2 . (B.58) 4  Suppose exp(−ad) ≥ exp − d/KB.9 . By (B.58), the condition for (B.55) with t = exp(−ad) thus holds, and (B.55) yields  Z (I) Pi

u

Wi (I)

  2 < exp(−ad) ≤ 16KB.7 exp − ad/KB.7 .

  Else, say exp(−ad) < exp −d/KB.9 . By (B.58), the conditions for (B.55) with t = exp −d/KB.9 thus holds, and (B.55) yields  Z (I) Pi

  Z (I)    d u < exp(−ad) ≤ Pi < exp − d/KB.9 ≤ 16KB.7 exp − . 2 Wi (I) Wi (I) KB.9 KB.7 u

Combining with (B.57) implies that for all γ > 0, we have o   n a   1 + γK(a, β )d . , Ei exp(γY2 ) ≤ 16KB.7 exp − d min 0 2 2 KB.7 KB.9 KB.7 Taking γ2 = γ2 (a, β0 , β1 , λ) =

n a o 1 1 min , >0 2 2 2K(a, β0 ) KB.7 KB.9 KB.7

to be small enough, we obtain from the above that   Ei exp(γ2 Y2 ) ≤ 16KB.7 + 1 .

(B.59)

Taking γ = min{γ1 , γ2 } > 0, as Y1 (I), Y2 (I) ≥ 0, we obtain from (B.56), (B.59) that     Ei exp(λY1 (I)) , Ei exp(λY2 (I)) ≤ K(a, β0 , β1 , κ, λ) , which establishes (B.53) and thus finishes the proof of Lemma B.12. Our next goal is to show that when the ε′ defining the mollifiers u(·) are chosen suitably, the mollification has only a minor impact on the free energy. Lemma B.14. For any a > 0 and ε ∈ (0, 1), there is ε′ > 0 depending on ε, a, β0 , β1 , κ, λ such that for any concave u ≤ 0 satisfying  u(x) = 0 for all x ≥ κ and exp u κ − ε′ ≤ ε′ , the following holds. For all β ∈ [β0 , β1 ], any x ∈ [−1, 1], and any interval I = [x − ε′′ , x + ε′′ ] with ′ depends on a, β0 , β1 , λ, we have for all d ≥ d(ε, a, β0 , β1 , κ, λ) and all n ≤ 2d, ε′′ < K1′ where KB.9 B.9

 1   1  E logad Zβ (I) − E logad Zu (I) ≤ ε , d d where we recall the definition of Zβ (I) from (B.2) and Zu (I) from (B.42). 55

(B.60)

 Proof. We replace each 1{Ui } with the corresponding mollified quantity exp u(Si ) and show this leads to only a small difference per i. For i ≤ n, define Z  X  \ √  Ci := Ui′ , Vi (I) := 1 Ci ∩ {θ1 / d ∈ I} exp u(Si′ ) − β∥θ∥2 dθ . (B.61) i′ <i

i≤i′ ≤n

The quantity of interest (B.60) in the statement of the Lemma is thus  1   1  1 X E logad V1 (I) − E logad Vn+1 (I) ≤ d d d

    E logad Vi+1 (I) − E logad Vi (I) .

1≤i≤n

We will bound each of the terms above. To do so, consider any i, 1 ≤ i ≤ n. We use Ei to denote expectation and Pi to denote probability in the randomness defining the i-th sample. Notice h i     E logad Vi+1 (I) − E logad Vi (I) ≤ E Ei logad Vi+1 (I) − logad Vi (I) . We will prove for any i, 1 ≤ i ≤ n that i h Ei logad Vi+1 (I) − logad Vi (I) ≤ Kε ,

(B.62)

where K > 0 is a universal constant independent of everything else. Upon scaling ε by a universal constant, (B.62) proves the desired result. To this end, define the Gibbs measure Gi on Rd as follows: for all subsets B ⊆ Rd , we let Z   X √  1 Gi (B) := 1 Ci ∩ B ∩ {θ1 / d ∈ I} exp u(Si′ ) − β∥θ∥2 dθ , Zi (I) ′ i<i ≤n

where

Z Zi (I) :=

  X √  1 Ci ∩ {θ1 / d ∈ I} exp u(Si′ ) − β∥θ∥2 dθ

Ci

i<i′ ≤n

is the normalizing constant. Recalling the definition of Ci from (B.61), note that Zi (I) does not depend on the randomness defining the i-th sample. Let ⟨·⟩ denote expectation w.r.t. Gi . Recalling the definition of Ci from (B.61) once more, we notice for analogous reasons as above, ⟨·⟩ does not depend on the randomness defining the i-th sample. Next, notice Vi (I) = Zi (I) exp u(Si ) ,

(B.63)

 Vi+1 (I) = Zi (I) 1{Ui } = Zi (I) Gi (Ui ) = Zi (I) Gi {Si ≥ κ} .

(B.64)

As u ≤ 0, we have Vi (I), Vi+1 (I) ≤ Zi (I). Now, we break into cases based on Zi (I). This step is permitted because Ei is only w.r.t. the randomness defining the i-th sample, while as noted above, Zi (I) does not depend on this randomness. First, suppose Zi (I) ≤ exp(−ad). Here we have h i Ei logad Vi+1 (I) − logad Vi (I) = 0 , establishing (B.62). Now, suppose Zi (I) ≥ exp(−ad) from here on out, which is permitted as justified above. Note as u(x) = 0 for x ≥ κ, observe that  0≤Y ≤X≤1 where X := exp u(Si ) , Y := Gi {Si ≥ κ} . (B.65) 56

The idea is that X and Y should be close if u(·) is a mollifier that accurately represents the indicator; formalizing this intuition is how we make quantitative bounds. Before breaking into further cases on Zi (I), we first present two useful Lemmas. Lemma B.15. Let c :=



ε2

2 KB.7

4KB.7 K1

> 0,

(B.66)

where KB.7 > 0 is from Theorem B.9 and depends on a, β0 , β1 , λ, and where K1 > 0 is large enough in terms of a, β0 , β1 , κ, λ. Then if Zi (I) ≥ exp(−ad), for d large enough in terms of ε, a, β0 , β1 , κ, λ, we have h i  Ei logad Y 1 Y ≤ c ≤ ε , where Y is defined as per (B.65). Proof of Lemma B.15. Since Zi (I) ≥ exp(−ad) and as ε′′ ≤ K1′ in the statement of Lemma B.14, B.9 P we may apply Theorem B.9 with U (θ) = i<i′ ≤n u(Si′ ) and C = Ci . Thus GC defined in Theorem B.9 is precisely Gi . Note U ≤ 0 and is concave as u ≤ 0 is concave and the Si are affine. We obtain exp(−d/KB.9 ) ≤ t ≤

  1 2 exp − KB.7 max{1, κ}2 =⇒ P(Y ≤ t) ≤ 16KB.7 t1/KB.7 . 4

(B.67)

′ ′ (a, β , β , λ), and K for KB.9 = KB.9 (a, β0 , β1 , λ), KB.9 = KB.9 0 1 B.7 = KB.7 (a, β0 , β1 , λ). Next, note Y ≤ 1 by (B.65), so

 logad Y = log V where V := min exp(ad), 1/Y

≥ 1.

2 , Thus we may apply Lemma B.10 with V defined as above, and C0 = KB.9 , C1 = 16KB.7 , C2 = KB.7  and D0 = 4 exp KB.7 max{1, κ}2 . We thus obtain that for γ = γ(a, β0 , β1 , λ) > 0,

i h Ei V γ(a,β0 ,β1 ,λ) ≤ K(a, β0 , β1 , κ, λ) . As V ≥ 1, we have log V < K̃(γ)V γ for all γ > 0, where K̃(γ) is an explicit univariate function. Applying this for γ = γ(a, β0 , β1 , λ)/2 and combining with the above display gives, by taking K1 larger if needed, h h i 2 i Ei log V ≤ K̃(γ(a, β0 , β1 , λ)/2)2 · Ei V γ(a,β0 ,β1 ,λ) ≤ K̃(γ(a, β0 , β1 , λ)/2)2 · K(a, β0 , β1 , κ, λ) . Thus, we may take K1 large enough in terms of a, β0 , β1 , κ, λ so that h 2 i Ei log V ≤ K1 , 2   KB.7  1 1 ≤ exp − KB.7 max{1, κ}2 . 4KB.7 K1 4

(B.68) (B.69)

Since ε ≤ 1, for d ≥ d(ε, a, β0 , β1 , κ, λ), our choice of c > 0 implies that the conditions of (B.67) ′2 ′ c1/KB.7 hold with t = c. Now (B.67) gives Pi (Y ≤ c) ≤ 4KB.7 . Cauchy-Schwarz and our choice of c 57

now implies h Ei

i h  i1/2 i h  2 1/2 Ei 1 Y ≤ c logad Y 1 Y ≤ c ≤ Ei logad Y h 2 i1/2 1/2 = Ei log V Pi Y ≤ c ′2 1/2 1/2 ′ ≤ K1 · 4KB.7 c1/KB.7 ≤ ε,

by our choice of c > 0. This proves Lemma B.15. 2 Lemma B.16. Define X, Y as per (B.65). Then if Zi (I) ≥ exp(−ad), for ε′ ≤ min{1, rB.11 } where rB.11 depends on a, β0 , β1 , we have i h √ (B.70) Ei X − Y ≤ exp(−d) + K ε′ ,

where K > 0 is a universal constant. Proof of Lemma B.16. Since exp u(κ − ε′ ) ≤ ε′ , we have the following pointwise: D  E 0 ≤ X − Y ≤ ε′ + 1 κ − ε′ ≤ Si ≤ κ , where we recall that ⟨·⟩ denotes expectation  w.r.t.√ Gi , which does not depend on the randomness defining the i-th sample. Since the event ∥θ̄∥ > ε′ d is independent of the randomness defining the i-th sample, we obtain h i 0 ≤ Ei X − Y iE D h  ≤ ε′ + Ei 1 κ − ε′ ≤ Si ≤ κ E D  E D  √ √  + Pi {κ − ε′ ≤ Si ≤ κ} 1 ∥θ̄∥ ≤ ε′ d = ε′ + Pi {κ − ε′ ≤ Si ≤ κ} 1{ θ̄∥ > ε′ d D  E D  E   √ √ ≤ ε′ + Pi κ − ε′ ≤ Si ≤ κ 1 ∥θ̄∥ > ε′ d + 1 ∥θ̄∥ ≤ ε′ d . (B.71) Recall Zi (I) ≥ exp(−ad). Thus for ε′ ≤ rB.11 (a, β0 , β1 )2 , noting the conditions apply for the Gibbs measure ⟨·⟩ as u ≤ 0, we obtain from Lemma B.11 that D  E √ 1 ∥θ̄∥ ≤ ε′ d ≤ exp(−d) . Now consider any κ1 < κ2 . Letting P̄, Ē denote probability and expectation in gi , observe that Z  ⟨g , θ̄⟩ h  sθ1 sθ1 i i Pi κ1 < Si < κ2 = ϕ(s) P̄ √ ∈ κ1 − √ , κ2 − √ ds , N d d R √ 1 , κ′ := κ2 − sθ √ 1 and where ϕ(·) denotes the density of si . For a given s ∈ R, let κ′1 := κ1 − sθ 2 d d   ⟨gi ,θ̄⟩ ∥θ̄∥2 √ Z = d . Note as Z ∼ N 0, d is Gaussian,

Z κ′  ⟨g , θ̄⟩ h  2 sθ1 sθ1 i 1 t2  i √ √ √ p ∈ κ1 − , κ2 − = dt P̄ exp − 2 E[Z 2 ] d d d 2π E[Z 2 ] κ′1 K ≤p (κ′2 − κ′1 ) 2 E[Z ] √ K d = (κ2 − κ1 ) , ∥θ̄∥ 58

for K > 0 a universal constant. Thus √ √ Z  K d K d ϕ(s) (κ2 − κ1 ) ds = (κ2 − κ1 ) , Pi κ1 < Si < κ2 ≤ ∥θ̄∥ ∥θ̄∥ R yielding that for any θ, √ √ √ √   Kε′ d ′ Pi κ − ε ≤ Si ≤ κ 1 ∥θ̄∥ > ε d ≤ 1{∥θ̄∥ > ε′ d} ≤ K ε′ . ∥θ̄∥ ′

Combining these steps with (B.71) yields i h √ Ei X − Y ≤ ε′ + exp(−d) + K ε′ , and hence (B.70). We now return to the proof of Lemma B.14. We break into cases on the value of Zi (I), recalling that it suffices to prove (B.62) when Zi (I) ≥ exp(−ad). If exp(−ad) ≤ Zi (I) ≤ 1. Applying Lemma F.3 and Lemma F.4 on truncated logarithms in the first and second inequality below respectively, as X ≥ Y and combining (B.63), (B.64) and (B.65), we obtain for any c > 0 that   logad Vi+1 − logad Vi = logad Zi (I)Y − logad Zi (I)X ≤ logad Y − logad X = logad X − logad Y ≤ logad Y 1{Y ≤ c} +

|X − Y | . c

Taking c = c(ε, a, β0 , β1 , κ, λ) > 0 as per (B.66), Lemmas B.15 and B.16 give i 1 h i h i h Ei logad Vi+1 − logad Vi ≤ Ei logad Y 1{Y ≤ c} + Ei X − Y c √  1 ≤ε+ exp(−d) + K ε′ . c Since c = c(ε, a, β0 , β1 , κ, λ) in (B.66) does not depend on ε′ , we have for ε′ ≤ ε′ (ε, a, β0 , β1 , κ, λ) and d ≥ d(ε, a, β0 , β1 , κ, λ) that h i Ei logad Vi+1 − logad Vi ≤ Kε , establishing (B.62) in this case. If Zi (I) > 1. Let c = c(ε, a, β0 , β1 , κ, λ) > 0 be from (B.66). Note for d ≥ d(ε, a, β0 , β1 , κ, λ) that exp(−ad) ≤ c. Thus we may define the following events: n o n o n o E1 := Y ≤ exp(−ad) , E2 := exp(−ad) < Y ≤ c , E3 := Y ≥ c . Recall 1 ≥ X = ⟨exp u(Si )⟩ ≥ Y from (B.65) and that here Zi (I) > 1. Thus as Y ≥ c > 0 on E3 , Zi (I)X1{E3 } , Zi (I)Y 1{E3 } ≥ c1{E3 } ≥ exp(−ad)1{E3 } , 59

and therefore     logad Zi (I)X − logad Zi (I)Y 1{E3 } = log Zi (I)X − log Zi (I)Y 1{E3 } X 1{E3 } = log Y  X − 1 1{E3 } ≤ Y 1 ≤ |X − Y | 1{E3 } . c By (B.63), (B.64), we thus may write i h h  i Ei logad Vi+1 − logad Vi = Ei logad Zi (I)X − logad Zi (I)Y h  i = Ei 1{E1 } logad Zi (I)X − logad Zi (I)Y h  i + Ei 1{E2 } logad Zi (I)X − logad Zi (I)Y h i 1 + Ei 1{E3 } · |X − Y | c := (I) + (II) + (III) .

(B.72)

Now we upper bound each of (I), (II), (III): • Upper bounding (I): Since Zi (I) ≥ exp(−ad) and ε′′ ≤ K1′ , we may apply Theorem B.9 B.9 identically as in the proof of Lemma B.15, which establishes that the implication (B.67)  holds. First, suppose exp(−ad) ≥ exp − d/KB.9 . Taking t = exp(−ad) in (B.67) gives   ad  . Pi Y < exp(−ad) ≤ 16KB.7 exp − 2 KB.7   Else, say exp(−ad) < exp − d/KB.9 . Taking t = exp − d/KB.9 in (B.67) gives      d Pi Y < exp(−ad) ≤ Pi Y < exp − d/KB.9 ≤ 16KB.7 exp − . 2 KB.9 KB.7 Therefore, it follows that   Pi (E1 ) = Pi Y < exp(−ad) ≤ 16KB.7 exp −

 d . 2 max{1/a, K KB.7 B.9 }

Now as u ≤ 0, we have Z Zi (I) ≤

exp(−β∥θ∥2 ) dθ ≤

Rd

 2π d/2 β0

.

Thus as X, Y ≤ 1 by (B.65) and by definition of the truncated logarithm, we have   logad Zi (I)X − logad Zi (I)Y ≤ K(a, β0 )d . Thus for d ≥ d(ε, a, β0 , β1 , λ), we obtain h  i Ei 1{E1 } logad Zi (I)X − logad Zi (I)Y ≤ K(a, β0 )d · Pi (E1 )  ≤ K(a, β0 )d · 16KB.7 exp − ≤ ε.

 d 2 max{1/a, K KB.7 B.9 } (B.73)

60

• Upper bounding (II): On the event E2 , we have c ≥ Y > exp(−ad). Thus by (B.65), on the event E2 , we have 1 ≥ X ≥ Y ≥ exp(−ad) as well. As Zi (I) ≥ 1 in this case, we obtain     1{E2 } logad Zi (I)X − logad Zi (I)Y = 1{E2 } log Zi (I)X − log Zi (I)Y = 1{E2 }| log X − log Y | ≤ 1{E2 }| log Y | = 1{E2 } logad Y ≤ 1{Y ≤ c} logad Y . Therefore as Zi (I) ≥ 1 ≥ exp(−ad), Lemma B.15 implies for d ≥ d(ε, a, β0 , β1 , κ, λ), h h i  i Ei 1{E2 } logad Zi (I)X − logad Zi (I)Y ≤ Ei 1{Y ≤ c} logad Y ≤ ε .

(B.74)

• Upper bounding (III): By Lemma B.16, as Zi (I) ≥ 1 in this case, for ε′ ≤ ε′ (a, β0 , β1 ) we have i h √ 1 Ei 1{E3 } · |X − Y | ≤ exp(−d) + K ε′ . c ′ ′ Hence, for ε ≤ ε (ε, a, β0 , β1 , κ, λ) and for d ≥ d(ε, a, β0 , β1 , κ, λ), we have h i 1 Ei 1{E3 } · |X − Y | ≤ ε . (B.75) c Combining (B.73), (B.74), (B.75) with (B.72), we obtain that h i Ei logad Vi+1 − logad Vi ≤ 3ε , hence establishing (B.62) when Zi (I) ≥ 1. This completes the proof of Lemma B.14. Finally, we put all these steps together to prove Theorem B.1. Proof of Theorem B.1. Without loss of generality, we assume ε ≤ 14 in the following. Consider Φ̄(x, q, ρ) as defined in (5.4) with exp u(x) = 1{x ≥ κ}, and note this function is a continuous function of its three arguments. By continuity supplied by the second part of Proposition A.3 and compactness, we may let a = a(β0 , β1 , κ, λ) > 0 be large enough so that the following holds. For all (α, β, x) ∈ [0, α0 ] × [β0 , β1 ] × [−1, 1], letting q0 = q0 (α, β, x), ρ0 = ρ0 (α, β, x) in the display below for brevity, we have Φ̄(x, q0 , ρ0 ) − βρ +

1 1 log(2πe) , Φ̄(x, q0 , ρ0 ) − β(ρ + x2 ) + log(2πe) ≥ −a + 1 . 2 2

Here we note α0 depends only on β0 , β1 , λ as Csol depends on β0 , β1 . Note a does not depend on ε′′ , ε′ , since Φ̄(x, q, ρ), q0 (α, β, x), ρ0 (α, β, x) do not depend on I as per Remark 1. As RSGI equals Φ̄(x, q0 , ρ0 ) − βρ + 21 log(2πe) or Φ̄(x, q0 , ρ0 ) − β(ρ + x2 ) + 21 log(2πe) pointwise, we obtain for all (α, β, x) ∈ [0, α0 ] × [β0 , β1 ] × [−1, 1] that RSGI (α, β, x) ≥ −a + 1 .

(B.76)

By Lemma B.14, there is ε′1 = ε′1 (ε, a, β0 , β1 , κ, λ) > 0 such that for all d ≥ d(ε, a, β0 , β1 , κ, λ) ′ ′ (a, β , β , λ). For any concave and all n ≤ 2d, the following holds if ε′′ ≤ K1′ , where KB.9 = KB.9 0 1 B.9 u ≤ 0 such that  u(x) = 0 for all x ≥ κ and exp u κ − ε′1 ≤ ε′1 , 61

then  1   1  E logad Zβ (I) − E logad Zu (I) ≤ ε . d d

(B.77)

Likewise by Theorem A.1, there exists ε′2 = ε′2 (ε, β0 , β1 , κ+ ) such that for any concave u ≤ 0 such that 1. u′ ≥ 0 with u′ ̸= 0 on a set of positive Lebesgue measure, 2. u(x) = 0 for x ≥ κ, 3. exp u(κ − ε′2 ) ≤ ε′2 , 4. u is four times differentiable and |u(l) | ≤ D for l = 1, 2, 3, 4, there is a εRS0 = εRS0 (ε, u, β0 , β1 , λ) > 0 such that if 0 < ε′′ ≤ εRS0 , the following holds. For d ≥ d(ε, ε′′ , β0 , β1 , κ, λ), we have  1  E log Zu (I) − RSGI (n/d, β, x) ≤ ε . d

(B.78)

Let ε′ = ε′ (ε, a, β0 , β1 , κ, λ) := min{ε′1 , ε′2 } . Given ε′ , it is straightforward to construct a mollifier û ≤ 0 that is concave, satisfying û′ ≥ 0 and û′ ̸= 0 on a set of positive Lebesgue measure, satisfying |û(l) | ≤ D(ε′ , κ) for l = 1, 2, 3, 4, and satisfying û(x) = 0 for all x ≥ κ , exp û(κ − ε′ ) ≤ ε′ . Hence as û′ ≥ 0 and ε′ ≤ ε′1 , ε′2 , it follows that exp û(κ − ε′1 ) ≤ ε′1 and exp û(κ − ε′2 ) ≤ ε′2 . Since ε′ = ε′ (ε, a, β0 , β1 , κ, λ), we may let   ′ εRSG = εRSG ε, α0 , β0 , β1 , κ, λ := min 1/KB.9 , εRS0 , where now εRS0 = εRS0 (û, β0 , β1 , λ). This is well-defined as û is constructed using κ, ε′ , and as ε′ in turn depends on ε, a = a(α0 , β0 , β1 , κ, λ), β0 , β1 , κ, and λ. Consider any 0 ≤ ε′′ < εRSG and let I = [x − ε′′ , x + ε′′ ]. Note as exp û(κ − ε′1 ) ≤ ε′1 and exp û(κ − ε′2 ) ≤ ε′2 , for d ≥ d(ε, ε′′ , a, β0 , β1 , κ, λ), we have both (B.77) and (B.78) for Zû (I). With these steps in hand, the finish is as follows. We need one last argument about controlling the error of the truncated logarithm: Lemma B.17. For d large enough in terms of ε, ε′′ , a, β0 , β1 , κ, λ, we have  1   1  E logad Zû (I) − E log Zû (I) ≤ ε . d d Proof of Lemma B.17. Let  Ω := Zû (I) ≤ exp(−ad) . Thus by Cauchy-Schwarz, h  i  1   1  1 E logad Zû (I) − E log Zû (I) = E 1{Ω} a + log Zû (I) d d d h 2 i 1 ≤ P(Ω) E a + log Zû (I) d h i 1 ≤ K P(Ω) E a2 + 2 log2 Zû (I) . d 62

(B.79)

h i First we upper bound E a2 + d12 log2 Zû (I) . Notice as û ≤ 0, Z



Zû (I) ≤

2

exp Hû (θ) − β∥θ∥



Z dθ ≤

2

exp(−β∥θ∥ ) dθ ≤

 2π d/2 β0

,

thus

h i 1 E a2 + 2 log2 Zû (I) ≤ K(a, β0 ) . d Now we upper bound P(Ω). As logA x ≥ log x for all x ∈ R and as ε ≤ 41 , (B.78) and (B.76) yields that for d ≥ d(ε, ε′′ , a, β0 , β1 , κ, λ),  1   1  E logad Zû (I) ≥ E log Zû (I) ≥ RSGI (n/d, β, x) − ε ≥ −a + 1 − ε ≥ −a + 3/4 . d d Therefore,  1  3o E logad Zû (I) ≥ . d d 4 ′ , we may apply Lemma B.12 with u = û. Thus for d ≥ d(ε, ε′′ , a, β , β , κ, λ), Now as ε′′ ≤ 1/KB.9 0 1  1     d 1  3 P(Ω) ≤ P ≤ 2 exp − . logad Zû (I) − E logad Zû (I) ≥ d d 4 K(a, β0 , β1 , κ, λ) Ω⊆

n 1

logad Zû (I) −

Combining the above bounds, it follows for such d that    1   1  d ≤ ε. E logad Zû (I) − E log Zû (I) ≤ K(a, β0 ) · 2 exp − d d K(a, β0 , β1 , κ, λ) This proves Lemma B.17. Now summing together (B.77), (B.78), and (B.79), we obtain for d ≥ d(ε, ε′′ , a, β0 , β1 , κ, λ) that   1   1  1  E logad Zβ (I) − RSGI (n/d, β, x) ≤ E logad Zβ (I) − E logad Zû (I) d d d  1   1  + E logad Zû (I) − E log Zû (I) d d  1  + E log Zû (I) − RSGI (n/d, β, x) d ≤ 3ε . By Lemma B.12 with exp u(x) = 1{x ≥ κ} and t = ε, we obtain for such d,  1     1  d P logad Zβ (I) − E logad Zβ (I) ≥ ε ≤ 2 exp − , d d K(ε, a, β0 , β1 , κ, λ) where we obtain dependence on ε in K as we took t = ε in Lemma B.12. Combining the above two displays, we obtain for d ≥ d(ε, ε′′ , a, β0 , β1 , κ, λ) that     1 d logad Zβ (I) − RSGI (n/d, β, x) ≥ 4ε ≤ exp − . (B.80) P d K(ε, a, β0 , β1 , κ, λ) However, by (B.76), we have RSGI (n/d, β, x) ≥ −a + 1 ≥ −a + 4ε. Thus as events, n 1 o n 1 o logad Zβ (I) − RSGI (n/d, β, x) ≤ 4ε ≡ log Zβ (I) − RSGI (n/d, β, x) ≤ 4ε , d d since under this event logad ≡ log as functions. Combining the above display with (B.80) and applying this result for ε/4 and recalling a = a(β0 , β1 , κ, λ) proves Theorem B.1. 63

B.2

Proof of Theorem B.2 for posterior

Our strategy will be to apply Theorem B.3 on concentration of Lipschitz functions of strongly log-concave measures. Specifically, we consider d1 log Zβ (I) as a function of the disorder (si , ḡi )ni=1 , and bound its resulting Lipschitz constant; recall as in Appendix A that the first four derivatives of u are upper bounded by some quantity D, which only depends on λ. Here in this proof we write dependence on D, and in fact we only need the bound on the first derivative of u. Let us begin by considering the space S = Rnd where the first n coordinates correspond to the sj and the next (n − 1)d coordinates correspond to gi,j . We endow S with the product measure γ where the measure on each coordinate is given by law of each corresponding part of the disorder (that is, sj or gi,j ). Define B ⋆ :=

n X

∥ḡi ∥2

1/2

+

i=1

d

n

j=2

i=1

1 X  X 2 gij 16d

n  X 1/2 + n s2i + i=1

(B.81)

n

 X 2 p 1 si + Y 2 + Z 2 + 2π/β0 d , 8dKsignal i=1

C := {B ≤ 3KB.21 (β0 , λ) d} ,

(B.82)

where KB.21 (β0 , λ) is from Lemma B.21. Note B ⋆ is a convex function of the disorder and therefore C is a convex subset of Rd . Also recall Ksignal depends solely on λ. Next, let γ ′ be a probability measure defined on C with density proportional to that of γ. By Assumption 1, γ ′ is min{1, βsignal } strongly log-concave. Now, let us differentiate d1 log Zβ (I) w.r.t. the disorder, letting ⟨·⟩ below and in what follows √  Pn 2 here in Appendix B.2 denote the Gibbs average w.r.t 1{θ1 / d ∈ I} exp i=1 u(Si ) − β∥θ∥ : ! n  1 Z X  √ √ ∂ 1 log Zβ (I) = 1{θ1 / d ∈ I} exp u(Si ) − β∥θ∥2 · u′ (Si ) · θ1 / d dθ /Zβ (I) ∂si d d d R i=1 √ 1 ′ = u (Si )θ1 / d , d ! Z n   X  √ √ ∂ 1 1 log Zβ (I) = 1{θ1 / d ∈ I} exp u(Si ) − β∥θ∥2 · u′ (Si ) · θj / d dθ /Zβ (I) ∂gij d d Rd i=1 √ 1 ′ = u (Si )θj / d . d Therefore, letting ∇ denote the gradient of d1 log Zβ (I) w.r.t. the disorder, and defining R1,1 as per (A.16), we have n n X d √ 2 X √  1 X ′ u′ (Si )θj / d ∥∇∥ = 2 u (Si )θ1 / d + d 2

i=1

≤ ≤

1 d2

n X

i=1 j=1

u′ (Si )2 θ12 /d +

i=1

D2 n  4+ d2 2D2 d

n X d X

u′ (Si )2 θj2 /d

2



i=1 j=1 d X

θj2 /d



j=1

 4 + R1,1 .

(B.83)

64

√ Here we used that n ≤ α0 d ≤ 2d, and that θ12 /d ≤ 4 due to the indicator 1{θ1 / d ∈ I} in the Gibbs average ⟨·⟩. Upper bounding R1,1 . We now upper bound R1,1 in terms of B ⋆ defined in (B.81). To this end we introduce the following Lemmas. Lemma B.18 (Lemma 3.1.5, [Tal10]). Consider a measure µ on Rd defined as follows: for some C ⊆ Rd , we have for all B ⊆ Rd , Z h  i X 1 µ(B) = 1{B ∩ C} EW exp U (θ) − β∥θ∥2 + aj θj + a0 dθ , (B.84) Z 1≤j≤d

where U ≤ 0, Z here denotes the appropriate normalizing constant, and a0 potentially depends on the annealed disorder W . Then Z  1 X   β   1  2π d/2 exp ∥θ∥2 dµ(θ) ≤ a2j EW exp a0 . (B.85) exp 2 Z β 2β 1≤j≤d

Technically the above Lemma is stated in [Tal10] for C = Rd and with a0 = 0, but the proof is identical. Note we will need this Lemma with a0 depending on the annealed disorder W from the interpolation argument in Appendix D.3. We next find a way to upper bound the moments w.r.t. a measure in the form of Lemma B.18, provided that the normalizing constant is not too small. Lemma B.19. Consider a measure µ in the form (B.84) given in Lemma B.18, let aj be defined as in Lemma B.18, and let Z be the corresponding normalizing constant. Suppose Z ≥ exp − K ′ B) for some K ′ depending on D, β0 , β1 , λ and some B ≥ d. Then we have: 1. We have D β E D β E    1 X 2 exp ≤ exp ≤ exp KB.19 B + ∥θ̄∥2 ∥θ∥2 aj + log EW exp a0 , 2 2 2β0 1≤j≤d

where KB.19 ≥ 1 depends on K ′ , D, β0 , β1 , λ. 2. For k ≤ 4d, we have   k 1 X 2 ∥θ̄∥2k ≤ ∥θ∥2k ≤ KB + , aj + log EW exp a0 2β0 1≤j≤d

where K depends on K ′ , D, β0 , β1 , λ. Proof of Lemma B.19. For the first part of this Lemma, we directly combine Lemma B.18 with the given hypothesis Z ′ ≥ exp − K ′ B). For the second part of this Lemma, apply Lemma A.7 with X = β2 ∥θ∥2 and expectation w.r.t. ⟨·⟩. By the first part of this Lemma, we have     1 X 2 log E exp X ≤ KB.19 B + aj + log EW exp a0 , 2β0 1≤j≤d

where expectation here is w.r.t. ⟨·⟩. The proof is complete upon noting k ≤ 4d ≤ 4B and multiplying by (2/β)k on both sides. 65

Now, note that u(x) ≥ −D(1 + |x|) where u(x) is from (3.17), (3.18) in the GMM and logistic cases respectively. Thus Lemma D.9 implies that for d ≥ d(ε′′ ), we have  Zβ (I) ≥ exp − K(D, β0 , β1 , λ)B ⋆ ,

(B.86)

where B ⋆ is given in (B.81). (The proof of Lemma D.9 is deferred to Appendix D.3 for simplicity of notation, since the Lemma is stated for the interpolating Gibbs measures from Appendix A, of which the measure considered here is a special case with t = v = 1.) We now show in both the GMM and logistic cases that R1,1 ≤

K(D, β0 , β1 , λ) B ⋆ . d

(B.87)

• In the logistic case where u(x) is from (3.18), we have φ ≤ 1 and therefore u ≤ 0. Hence by (B.86), we may apply Lemma B.19 with the aj = 0, which gives (B.87). √ • In the GMM case where u(x) is from (3.17), as u(x) = λx, we can write n X

u(Si ) =

i=1

n X i=1

r  d n n d r X  rλ X   X X λ λ si θ1 + gij θj = si θ1 + gij θj . d d d j=2

i=1

(B.88)

i=1

j=2

Thus applying Lemma B.19 with (B.86), we obtain that n

d

n

λ  X 2 1 λ X  X 2 K(D, β0 , β1 , λ)B ⋆ + R1,1 ≤ si + gij d 2β0 d 2β0 d i=1

j=2

!

i=1

K(D, β0 , β1 , λ) B ⋆ , d

where the last step follows from the definition of B ⋆ in (B.81). Again this proves (B.87). Completing the argument. Now we apply Theorem B.3 with f = d1 log Zβ (I) and the strongly √ log-concave measure γ ′ . By (B.83) and (B.87), f is Lipschitz with constant K(D, β0 , β1 , λ)/ d on C. Since γ ′ is supported on C and C is convex, Theorem B.3 yields for d ≥ d(ε′′ ) that  1 Pγ ′

d

log Zβ (I) −

    1 Eγ ′ log Zβ (I) ≥ ε/2 ≤ exp − d/K(ε, D, β0 , β1 , λ) , d

(B.89)

where ε > 0 is as in the statement of Theorem B.2. We next state the following Lemma that essentially lets us reduce the desired tail bound in Theorem B.2 to (B.89). Its proof will be presented later. Lemma B.20. For d ≥ d(ε, ε′′ , D, β0 , β1 , λ), we have  Pγ C c ≤ exp(−d) ,   1   1 Eγ ′ log Zβ (I) − Eγ log Zβ (I) ≤ ε/4 . d d

(B.90) (B.91)

We will now prove Theorem B.2 given Lemma B.20. Let Eε′ :=

n 1 d

log Zβ (I) −

o   1 Eγ ′ log Zβ (I) ≥ ε/4 . d 66

(B.92)

By Lemma B.20, for d ≥ d(ε, ε′′ , D, β0 , β1 , λ),  1

   1 Pγ log Zβ (I) − Eγ log Zβ (I) ≥ ε/2 d d  1    1   1   1 = Pγ log Zβ (I) − Eγ ′ log Zβ (I) + Eγ ′ log Zβ (I) − Eγ log Zβ (I) ≥ ε/2 d d d d   1   1 log Zβ (I) − Eγ ′ log Zβ (I) ≥ ε/4 ≤ Pγ d h d i = Eγ 1{Eε′ } h i h i = Eγ 1{Eε′ }1{C} + Eγ 1{Eε′ }1{C c } h i h i ≤ γ(C) Eγ ′ 1{Eε′ }1{C} + Eγ 1{C c }  ≤ exp − d/K(ε, D, β0 , β1 , λ) .

(B.93)

Here in the last inequality, we applied (B.89). Finally, recall by Theorem A.2 that for d ≥ d(ε, ε′′ , D, β0 , β1 , λ), we have   1 Eγ log Zβ (I) − RSGI (n/d, β, x) ≤ ε/2 . d Therefore combining with (B.93) gives  1 Pγ

d

log Zβ (I) − RSGI (n/d, β, x)



  1    1 ≥ ε ≤ Pγ log Zβ (I) − Eγ log Zβ (I) ≥ ε/2 d d  ≤ exp − d/K(ε, D, β0 , β1 , λ) .

This proves Theorem B.2 given Lemma B.20. Proof of Lemma B.20. Finally, we prove Lemma B.20. We first introduce the following Lemma which gives us a tail bound on B ⋆ . This Lemma follows as a corollary of the proof of Lemma D.8, which is proven later in Appendix D.3. Lemma B.21. For all k ≤ d we have for KB.21 depending on β0 , λ,   k E B ⋆k ≤ KB.21 d . Given Lemma B.21, by the definition of C, we directly obtain (B.90). To prove (B.91), we introduce an additional Lemma: Lemma B.22. For d ≥ d(ε′′ ), we have log Zβ (I) ≤ K(D, β0 , β1 , λ)B ⋆ . Proof of Lemma B.22. The lower bound on log Zβ (I) follows from B.86. For the upper bound, note Z   X Zβ (I) ≤ exp − β∥θ∥2 + aj θj dθ ≤ (2π/β0 )d/2 , Rd

1≤j≤d

for some (aj )dj=1 ∈ R. (Note the integral is invariant to coordinate translation.) Specifically, in the logistic case we use that u ≤ 0, and in the GMM case we use (B.88). This proves Lemma B.22. 67

Now to prove (B.91), we write   1   1   1 Eγ log Zβ (I) = Eγ log Zβ (I)1{C} + Eγ log Zβ (I)1{C c } d d d   1   1 = γ(C) Eγ ′ log Zβ (I) + Eγ log Zβ (I)1{C c } , d d     where we note Eγ ′ log Zβ (I)1{C} = Eγ ′ log Zβ (I) as the support of γ ′ equals C. Therefore   1   1 Eγ ′ log Zβ (I) − Eγ log Zβ (I) d d   1   1   1 Eγ log Zβ (I)1{C} + Eγ log Zβ (I)1{C c } = Eγ ′ log Zβ (I) − d d d   1   1 − γ(C) = Eγ ′ log Zβ (I) − Eγ log Zβ (I)1{C c } . d d

(B.94)

By Lemma B.22, and since γ ′ is supported on C on which B ⋆ ≤ K(D, β0 , β1 , λ)d, we have     1 1 (B.95) Eγ ′ log Zβ (I) ≤ Eγ ′ log Zβ (I) ≤ K(D, β0 , β1 , λ) . d d Next by Lemma B.21, note  d Pγ B ⋆ ≥ t ≤ KB.21 (D, β0 , β1 , λ)/t for all t ≥ 0 . (B.96) Thus by Lemma B.22 and integrating the tail bound (B.96), we have     Eγ log Zβ (I)1{C c } ≤ K(D, β0 , β1 , λ) Eγ B ⋆ 1{C c } Z ∞  ≤ K(D, β0 , β1 , λ) Pγ B ⋆ ≥ t dt 3KB.21 (D,β0 ,β1 ,λ)d

≤ K(D, β0 , β1 , λ) exp(−d) .

(B.97)

Combining (B.90), (B.94), (B.95), (B.97) now proves (B.91) for d ≥ d(ε, ε′′ , D, β0 , β1 , λ), finishing the proof of Lemma B.20.

C

Conversion from Gaussian to spherical measure: proof of Theorems 3.1, 3.5

Here we complete the proof of our main results Theorems 3.1, 3.5. Theorem 3.1 on interpolators and Theorem 3.5 on the posterior follow as a corollary of the following Theorems C.1, C.2 respectively. Theorem C.1 for interpolators: This result states that for a narrow interval I near x ∈ [−1, 1] (whose width is independent of d), the normalized logarithm of the spherical measure of the interpolators constrained to this interval, approximately equals inf q∈[0,1) Φ(x, q) for d large enough: √ Theorem C.1 (For interpolators). Let µd (·) denote the uniform measure on Sd−1 ( d). Consider any α ≤ α0 and κ ∈ R, and let n = ⌊αd⌋. Define δ = δ(α, λ, κ+ ) as follows, where the inequality uses α ≤ α0 : 1/2 √ 1/2 + αKE.8 Csol / 2 1 δ = δ(α, λ, κ+ ) := ≤ . (C.1) 9 10 Recalling the definition of C in (B.1) and letting Φ(x, q) be as in (3.1) with exp u(x) = 1{x ≥ κ}, we have for all x ∈ [−1 + δ, 1 − δ] and ε > 0: 68

1. Upper bound: Let εRSG = εRSG (ε, α0 , β0 , β1 , κ, λ) be from Theorem B.1 and consider any 0 < ε′′ ≤ min{ε, εRSG }. Define the interval   x−ε′′ ′′  : x ≥ 0, κ < 0, ′′ , x + ε  1−ε   x − ε′′ , x+ε′′  : x < 0 , κ < 0 , 1−ε′′  I :=  (C.2) x+ε′′ ′′  x − ε , : x ≥ 0 , κ ≥ 0 , ′′  1+ε    x−ε′′ , x + ε′′  : x < 0 , κ ≥ 0 . ′′ 1+ε If |x| > ε′′ , for d large enough in terms of ε, ε′′ , α, β0 , β1 , κ, λ and with probability at least 1 − exp −d/K where K depends on ε, α, β0 , β1 , κ, λ, we have   √ 1 log µd 1{θ1 / d ∈ I} ∩ C ≤ inf Φ(x, q) + ε . d q∈[0,1)

(C.3)

(Note this bound is non-vacuous; we have I ̸= ∅.) 2. Lower bound: Suppose ε < K1 ′ where K ′ is large enough in terms of α, β0 , β1 , κ, λ. Then there exists εRS > 0 depending on ε, α, β0 , β1 , κ, λ such that for all 0 < ε′′ ≤ εRS , the following holds. Define the interval   x−ε′′  , x + ε′′ : x ≥ 0 , κ < 0 ,  1+ε   x − ε′′ , x+ε′′  : x < 0 , κ < 0 , 1+ε  I :=  (C.4) ′′ , x+ε′′  x − ε : x ≥ 0, κ ≥ 0,  1−ε    x−ε′′ , x + ε′′  : x < 0 , κ ≥ 0 . 1−ε If |x| > ε′′ , for d large enough in terms of ε, ε′′ , α, β0 , β1 , κ, λ and with probability at least 1 − exp −d/K where K depends on ε, α, β0 , β1 , κ, λ, we have   √ 1 log µd 1{θ1 / d ∈ I} ∩ C ≥ inf Φ(x, q) − ε . d q∈[0,1)

(C.5)

Theorem C.2 for posterior: This result is similar to Theorem C.1, but now is for the normalized logarithm of the integral of the posterior constrained to the interval I, rather than the spherical measure of the interpolators: Theorem C.2 (For posterior). Consider any α ≤ α0 and let n = ⌊αd⌋. Consider u(x) from (3.17), (3.18) in the GMM and logistic cases respectively and let Φ(x, q) be as in (3.1) with this u(x). Define δ = δ(α, λ) as follows, where the inequality uses α ≤ α0 : δ = δ(α, λ) :=

1 1 ≤ . 18 − α(2D2 + D) 10

(C.6)

For any interval I ⊆ R, let Z Z(I) :=

√ Sd−1 ( d)

n X   √ 1 θ1 / d ∈ I exp u(Si ) dθ , i=1

where Si = Si (θ) is as in (A.1). Then for all x ∈ [−1 + δ, 1 − δ] and ε > 0, we have: 69

(C.7)

1. Upper bound: Let εRSG = εRSG (ε, α0 , β0 , β1 , λ) be from Theorem B.2, consider any 0 < ε′′ ≤ min{ε/K ′ , εRSG } where K ′ depends on λ. Define the interval I by ( ′′  x − ε′′ , x+ε : x ≥ 0, ′′ 1+ε  I :=  x−ε′′ (C.8) ′′ : x < 0. 1+ε′′ , x + ε ′′ If |x| > ε′′ , for  d large enough in terms of ε, ε , α, β0 , β1 , λ and with probability at least 1 − exp −d/K where K depends on ε, α, β0 , β1 , λ, we have

1 1 log Z(I) ≤ inf Φ(x, q) + log(2πe) + ε . d 2 q∈[0,1)

(C.9)

(As in Theorem C.1, note this bound is non-vacuous; we have I ̸= ∅.) 2. Lower bound: Suppose ε < K1 ′ where K ′ is large enough in terms of α, β0 , β1 , λ. Then there exists εRS > 0 depending on ε, α, β0 , β1 , λ such that for all 0 < ε′′ ≤ εRS , the following holds. Define the interval ( ′′  x − ε′′ , x+ε : x ≥ 0, 1−ε  I :=  x−ε′′ (C.10) ′′ : x < 0. 1−ε , x + ε ′′ If |x| > ε′′ , for  d large enough in terms of ε, ε , α, β0 , β1 , λ and with probability at least 1 − exp −d/K where K depends on ε, α, β0 , β1 , λ, we have

1 1 log Z(I) ≥ inf Φ(x, q) + log(2πe) − ε . d 2 q∈[0,1)

(C.11)

We will prove Theorem C.1 in the κ < 0 case in Appendix C.2, prove Theorem C.1 in the κ ≥ 0 case in Appendix C.3, and prove Theorem C.2 in Appendix C.4. All three proofs follow the same unified strategy that leverages our work in Appendix B, specifically Theorem B.1 for the interpolators and Theorem B.2 for the posterior. We present a full proof for κ < 0 in Appendix C.2 and highlight the necessary modifications to complete the proof when κ > 0 or for the posterior in Appendices C.3, C.4. Remark 2. We emphasize that the interval I from Theorems C.1, C.2 does not equal [x−ε′′ , x+ε′′ ]. This is because our analysis relies on several comparison inequalities where we scale [x − ε′′ , x + ε′′ ] by a factor close to 1 independent of d; see e.g. (C.24), (C.25). I is then a larger interval containing the scaled intervals for the lower bound, or a smaller interval contained within the scaled intervals for the upper bound. For the upper bound, the parameter w governing the range of these scalings equals 1 ± ε′′ (the sign depending on whether κ ≥ 0, κ < 0), and we can check that I ̸= ∅. For the lower bound, ε′′ delicately depends on the parameter w = 1 ± ε governing these scalings, but now I ⊇ [x − ε′′ , x + ε′′ ] and is therefore nonempty. Before we present the proofs of Theorems C.1, C.2, we combine them with an application of the Laplace Method over x ∈ [−1, 1] to prove our main results: Theorem 3.1 on interpolators and Theorem 3.5 on the posterior respectively. Note the Laplace Method gives the ‘supx inf q ’ structure therein. In preparation, we will present the following properties of the RS equations, which follow from the convex-concave structure established in Appendix E. The proof of this Lemma is in Appendix E.3. 70

Lemma C.3. Let δ = δ(α, λ, κ+ ) from (C.1) for the interpolators, or δ = δ(α, λ) from (C.6) for the posterior. We have the following: 1. For all (α, β, x) ∈ [0, α0 ] × [β0 , β1 ] × [−1, 1], we have ∂ρ0 (α,β,x) < 0. ∂β 2. For all (α, x) ∈ [0, α0 ] × [−1 + δ, 1 − δ], there is a unique β = β(α, x) ∈ [β0 , β1 ] such that ρ0 (α, β, x) = 1 − x2 , and moreover β(α, x) ∈ (1/4, 9). 3. For all (α, x) ∈ [0, α0 ] × [−1 + δ, 1 − δ], recalling the definition of Φ̄ in (5.4),   f (β) := β + Φ̄ x, q0 (α, β, x), ρ0 (α, β, x) − β ρ0 (α, β, x) + x2 attains a unique minimum on β ∈ [β0 , β1 ] at β = β(α, x). Thus for any interval I, if 0 ̸∈ I, then β + RSGI (α, β, x) attains a unique minimum on β ∈ [β0 , β1 ] at β = β(α, x). 4. For all (α, x) ∈ [0, α0 ] × [−1 + δ, 1 − δ], we have  inf Φ(x, q) = f β(α, x) . q∈[0,1)

5. For all (α, x) ∈ [0, α0 ] × [−1 + δ, 1 − δ], inf q∈[0,1) Φ(x, q) is continuous in (α, x) (note α is implicitly an argument in Φ(x, q)). We now prove Theorems 3.1, 3.5 using Theorems C.1, C.2 and the above Lemma C.3.

C.1

Proof of Theorems 3.1, 3.5

We provide the proof of Theorem 3.1 using Theorem C.1 when κ < 0; the proof of Theorem 3.1 in the κ ≥ 0 case and the proof of Theorem 3.5 using Theorem C.2 are analogous. Recall√that yi Xi ≡ (si , gi )T in law as random variables and that without loss of generality we can set θ⋆ = de1 . First, we show (C.3), (C.5) with the interval I replaced by any interval I ⊆ [−1 + δ, 1 − δ]. Recall as per (A.7) that β0 , β1 are universal constants, therefore in this proof we will not explicitly write dependence on them. √ √ √  Proof of (C.3): Note for d ≥ d(ε), d1 log µd 1{θ1 / d ∈ [−1/ d, 1/ d]} ≤ ε. Now if 0 ∈ I, we √ √ subdivide I into the two closed intervals I1 = I ∩[−1, −1/ d], I2 = I ∩[1/ d, 1], and the remaining part I − I1 − I2 (if 0 ̸∈ I this step is not necessary, and we can just apply the same argument as in the following upper bound for I1 ). √  We now upper bound d1 log µd 1{θ1 / d ∈ I1 } ∩ C , the upper bound for I2 being analogous. We divide the closed interval I1 into a set of consecutive closed intervals I¯k each of width εRSG = εRSG (ε, α0 , κ, λ) (recall β0 , β1 are universal constants), 1 ≤ k ≤ K(εRSG ), that overlap at exactly their endpoints. Here εRSG comes from Theorem B.1. b−a For each interval I¯k = [a, b], note 0 < a < b ≤ 1, b − a ≤ εRSG . Letting ε′′ = 2−a < b−a   x−ε′′ ′′ ′′ ′′ ′′ ¯ and x = b − ε , we have Ik = [a, b] = 1−ε′′ , x + ε , x > ε , ε ≤ b − a ≤ εRSG . Since ε′′ ≥ b−a = εRSG = εRSG (ε, α0 , κ, λ), Theorem C.1 yields that for d ≥ d(ε, α, κ, λ), with probability at least 1 − exp −d/K(ε, α, κ, λ) ,   √ 1 ¯ log µd 1{θ1 / d ∈ Ik } ∩ C ≤ inf Φ(x, q) + ε ≤ sup inf Φ(x, q) + ε . d q∈[0,1) x∈I q∈[0,1) Here the second inequality uses that I¯k ⊆ I1 . As we have K(εRSG ) such intervals I¯k , combining the above bound over all 1 ≤ k ≤ K(εRSG ) and applying a Union Bound (note we sum the bounds on 71

√  µd 1{θ1 / d ∈ I1 } ∩ C , which behaves at the exponential scale) yields that for d ≥ d(ε, α, κ, λ),  with probability at least 1 − exp − d/K(ε, α, κ, λ) ,   √ 1 log µd 1{θ1 / d ∈ I1 } ∩ C ≤ sup inf Φ(x, q) + 2ε . d x∈I q∈[0,1) An analogous √ bound applies √ for√I2 via  the same proof, and for I − I1 − I2 we recall the bound 1 d ∈ [−1/ d, 1/ d]} ≤ ε. Combining these three upper bounds and applying a log µ 1{θ / 1 d d  Union Bound yields for d ≥ d(ε, α, κ, λ), with probability at least 1 − exp −d/K(ε, α, κ, λ) ,   √ 1 log µd 1{θ1 / d ∈ I} ∩ C ≤ sup inf Φ(x, q) + 4ε . d x∈I q∈[0,1) Upon taking ε ← ε/4, this proves (C.3) for this I. Proof of (C.5): In this proof, we write I = [a, b] for −1 ≤ a < b ≤ 1. Note it suffices to prove 1 the result when ε < K ′ (α,κ,λ) . First, continuity from Lemma C.3 implies supx∈I inf q∈[0,1) Φ(x, q) is ⋆ attained at some x ∈ I. Noting inf q∈[0,1) Φ(x, q) only depends on α, κ, λ, by continuity supplied by Lemma C.3 and compactness, we let δ̄ = δ̄(ε, α, κ, λ) be small enough so that inf Φ(x1 , q) − inf Φ(x2 , q) ≤ ε ,

q∈[0,1)

q∈[0,1)

∀ x1 , x2 ∈ [−1 + δ, 1 − δ] , |x1 − x2 | ≤ δ̄ .

(C.12)

Since ε ≤ ε0 , taking ε0 small enough in terms of α, κ, λ, |I|, we may suppose that |I| ≥ δ̄. We first prove the result when I ⊆ [0, ∞); the proof when I ⊆ (−∞, 0] is analogous, and later we will provide the argument when neither case holds. 1. Suppose x⋆ > 0 is distance at least δ̄/4 from a, b. Let ε̄ = ε̄(ε, δ̄, α, λ) ≤ min{δ̄/8, ε} and ε′′ < min εRS (ε̄, α, κ, λ), δ̄/8 be small enough, where εRS comes from Theorem C.1. Let h x⋆ − ε′′ i I¯ := , x⋆ + ε′′ . 1 + ε̄

(C.13)

As |a| ≤ 1, ε̄ + ε′′ < δ̄/4, |x⋆ | ≤ 1 and x⋆ is distance at least δ̄/4 from a, b, we have I¯ ⊆ I. Also, clearly x⋆ ≥ δ̄/4 > ε′′ . Applying Theorem C.1, recalling that δ̄ and ε̄ depend only ′′ on ε, α, κ, λ, we obtain for  d ≥ d(ε̄, ε , α, κ, λ) = d(ε, α, κ, λ) that with probability at least 1 − exp −d/K(ε, α, κ, λ) ,   1   √ √ 1 ¯ ∩C log µd 1{θ1 / d ∈ I} ∩ C ≥ log µd 1{θ1 / d ∈ I} d d ≥ inf Φ(x⋆ , q) − ε ,

(C.14)

q∈[0,1)

where we use that ε̄ ≤ ε in the last inequality. 2. Suppose x⋆ is within distance δ̄/4 from a, b. Now, we simply apply the above argument with x⋆ replaced by x̄⋆ = a + δ̄/4 if |a − x⋆ | ≤ δ̄/4 or x̄⋆ = b − δ̄/4 if |b − x⋆ | ≤ δ̄/4. Since |I| ≥ δ̄, the resulting x̄⋆ is always distance at least δ̄/4 from a, b, and hence x̄⋆ > ε′′ as well. Define ε̄, ε′′ identically to case 1) above and I¯ now in terms of x̄⋆ as per (C.13). Applying Theorem 72

C.1 analogously, by |x⋆ − x̄⋆ | ≤ δ̄ and (C.12), we obtain for the same d and with the same probability as (C.14) that  1    √ √ 1 ¯ ∩C log µd 1{θ1 / d ∈ I} ∩ C ≥ log µd 1{θ1 / d ∈ I} d d ≥ inf Φ(x̄⋆ , q) − ε̄ q∈[0,1)

≥ inf Φ(x⋆ , q) − 2ε .

(C.15)

q∈[0,1)

Finally suppose I is not contained within [0, ∞) or (−∞, 0]. We suppose x⋆ ≥ 0 in what follows, the argument when x⋆ < 0 is analogous. 1. Suppose I ∩ [0, ∞) ≥ δ̄/2. If x⋆ ≥ δ̄/4 we simply apply the argument in either cases 1) or 2) above, defining ε̄, ε′′ , I¯ the same way as there. In particular, we still have I¯ ⊆ I as I ∩ [0, ∞) ≥ δ̄/2; if we need to define x̄⋆ as in case 2) above (note this only arises when x⋆ is within δ̄/4 of the right endpoint of I), since I ∩ [0, ∞) ≥ δ̄/2, the resulting x̄⋆ is always distance at least δ̄/4 > ε′′ from the endpoints of I and 0. Next suppose 0 ≤ x⋆ ≤ δ̄/4. We apply the argument in case 2) above with x̄⋆ = δ̄/4. Note that defining I¯ analogously, we have I¯ ⊆ I as I ∩ [0, ∞) ≥ δ̄/2. Combining with the fact that |x̄⋆ − x⋆ | ≤ δ̄/4 and (C.12), we obtain (C.15). 2. Suppose I ∩ [0, ∞) < δ̄/2. Therefore as |I| ≥ δ̄, I ∩ (−∞, 0] ≥ δ̄/2. We now apply the argument in either cases 1) or 2)above now with x̄⋆ = −δ̄/4, defining ε̄, ε′′ analogously and ⋆ +ε′′ now defining I¯ := x̄⋆ − ε′′ , x̄1+ε̄ . Since the interval I is not fully contained in [0, ∞) or (−∞, 0] and I ∩ (−∞, 0] ≥ δ̄/2, we have x̄⋆ ∈ I, and furthermore that I¯ ⊆ I ∩ (−∞, 0]. Also, |x̄⋆ | ≥ δ̄/4 > ε′′ . Finally, as x⋆ ≥ 0 and I ∩ [0, ∞) < δ̄/2, we have |x̄⋆ − x⋆ | ≤ 3δ̄/4. Combining with (C.12) gives (C.15). This yields the desired lower bound in all cases. Now, (C.5) follows by taking ε ← ε/2. √  1 To complete the proof, we must upper bound d1 log µd 1{|θ1 / d| ≥ 1 − δ} ∩ C . As δ ≤ 10 , it remains to show that with probability at least 1 − exp − d/K(ε, α, κ, λ) for d ≥ d(ε, α, κ, λ), we have the following for a universal constant c > 0. For the interpolators,    √ 1 log µd 1 |θ1 / d| ≥ 9/10 ∩ C < sup inf Φ(x, q) − c , d x∈[−1+δ,1−δ] q∈[0,1)

(C.16)

and for the posterior, defining Z(I) as per (C.7),    1 1 log Z [−1, −9/10] + Z [9/10, 1] − log(2πe) < sup inf Φ(x, q) − c . d 2 x∈[−1+δ,1−δ] q∈[0,1)

(C.17)

We now upper bound the LHS in (C.16), (C.17). By following the  same proof as that of Lemma C.4, we have with probability at least 1 − exp − d/K(ε, α, κ, λ) , sup√

n X

√ u(Si ) ≤ K(λ) nd .

θ∈Sd−1 ( d) i=1

73

(C.18)

Standard bounds on the surface area of a spherical cap gives   √ 1 log µd 1{|θ1 / d| ≥ 9/10 ≤ −c′ , d where c′ > 0 is a universal constant. Letting LHS denote the LHS in (C.16), (C.17), combined with the bounds u = log φ ≤ 0 for the logistic case for the posterior or (C.18)for the GMM case for the posterior, we obtain with probability at least 1 − exp − d/K(ε, α, κ, λ) and for d ≥ K, √

LHS ≤ −c′ + K(λ) α .

(C.19)

On the other hand, by the proof of 1) in Appendix E.4, we have supx∈[−1,1] inf q∈[0,1) Φ(x, q) = Φ(x⋆ , q ⋆ ) for a unique (x⋆ , q ⋆ ) ∈ [0, K(λ)α]2 . Since α ≤ α0 (λ, κ+ ) and δ ≤ 1/10, the bound (E.39) implies that sup

inf Φ(x, q) = Φ(x⋆ , q ⋆ ) ≥ −αK(λ, κ+ ) .

(C.20)

x∈[−1+δ,1−δ] q∈[0,1)

Since α ≤ α0 (λ, κ+ ) and c′ > 0 is a universal constant, combining (C.19), (C.20) now implies (C.16) for the interpolators or (C.17) for the posterior. This completes the proof of Theorems 3.1, 3.5.

C.2

Proof of Theorem C.1 for interpolators: κ < 0

Throughout this proof, we let I¯ := [x − ε′′ , x + ε′′ ]. Note 0 ̸∈ I¯ as |x| > ε′′ . Consequently recalling ¯ equals x2 . (A.9), we note RSGI¯ does not depend on ε′′ as rI¯(x), the only part depending on I, In the following Appendices √ C.2, C.3, C.4, we let Area(A) refer to the surface area any Lebesgue d−1 measurable subset of vS ( d) for any v > 0.

Preliminaries: Due to the lack of continuity caused by the hard indicator functions, we instead use the monotonicity from the following comparison inequalities. For any v > 0, consider  o n  o n 1  1  vUi = vθ ∈ Rd : √ si θ1 + ⟨gi , θ̄⟩ ≥ κ = θ ∈ Rd : √ si θ1 + ⟨gi , θ̄⟩ ≥ vκ , d d

(C.21)

where the last equality follows from the transformation θ ← vθ. Therefore as κ < 0, we have For v ≥ 1: For v ≤ 1:

 o n 1  vUi ⊇ θ ∈ Rd : √ si θ1 + ⟨gi , θ̄⟩ ≥ κ = Ui , hence vC ⊇ C , d n  o 1  d √ vUi ⊆ θ ∈ R : si θ1 + ⟨gi , θ̄⟩ ≥ κ = Ui , hence vC ⊆ C . d

(C.22) (C.23)

√ We next set up several comparison inequalities letting us to relate the indicator 1{θ / d ∈ I} to 1 √ ¯ (the difference is between I, I). ¯ For any interval [a, b], remark 1{θ1 / d ∈ I} For 1 ≤ v ≤ w :

v[a, b] ⊇ [wa, b] if b ≥ a ≥ 0

,

v[a, b] ⊇ [a, wb] if a ≤ b ≤ 0 ,

(C.24)

For w ≤ v ≤ 1 :

v[a, b] ⊆ [wa, b] if b ≥ a ≥ 0

,

v[a, b] ⊆ [a, wb] if a ≤ b ≤ 0 .

(C.25)

74

Proof of (C.3): Let w = 1 − ε′′ here in the proof of (C.3) and let I be as per (C.2). Consider any β ∈ [β0 , β1 ]. By (C.23) and (C.25), we have the following. If x ≥ ε′′ > 0, Z √  ¯ ∩ C exp(−β∥θ∥2 ) dθ LHS := 1 {θ1 / d ∈ I}   √   ≥ Vol 1 θ1 / d ∈ [x − ε′′ , x + ε′′ ] ∩ C ∩ 1 w2 d ≤ ∥θ∥2 ≤ d exp(−βd)   √ √ Z 1 √  = exp(−βd) d Area 1 θ1 / d ∈ [x − ε′′ , x + ε′′ ] ∩ C ∩ vSd−1 ( d) dv w  n √ h x − ε′′ io √ Z 1 √  ≥ exp(−βd) d Area 1 θ1 / d ∈ v , x + ε′′ ∩ vC ∩ vSd−1 ( d) dv w w  n √ h x − ε′′ io √ Z 1 √  Area 1 θ1 / d ∈ = exp(−βd) d , x + ε′′ ∩ C ∩ Sd−1 ( d) v d−1 dv w w  √ √  1 = exp(−βd) Area 1{θ1 / d ∈ I} ∩ C ∩ Sd−1 ( d) √ (1 − wd ) d  √ √  1 d−1 ≥ exp(−βd) Area 1{θ1 / d ∈ I} ∩ C ∩ S ( d) √ . 2 d Here we used the change of variable θ ← vθ, and that wd = (1 − ε′′ )d ≤ 12 for d ≥ d(ε′′ ). Similarly if x ≤ −ε′′ < 0, by an analogous derivation as above, again using (C.23) and (C.25),   √ √  √ Z 1 Area 1 θ1 / d ∈ [x − ε′′ , x + ε′′ ] ∩ C ∩ vSd−1 ( d) dv LHS ≥ exp(−βd) d w  n √ h √ Z 1 √  x + ε′′ io ≥ exp(−βd) d Area 1 θ1 / d ∈ v x − ε′′ , ∩ vC ∩ vSd−1 ( d) dv w w  1  √ √ ≥ exp(−βd) Area 1{θ1 / d ∈ I} ∩ C ∩ Sd−1 ( d) √ . 2 d For both cases, taking logarithms gives   1  √ √  1  1 1 log LHS ≥ −β + log Area 1{θ1 / d ∈ I} ∩ C ∩ Sd−1 ( d) + log √ . d d d 2 d

(C.26)

¯ since ε′′ ≤ εRSG (ε, α0 , β0 , β1 , κ, λ), we have for d ≥ By Theorem B.1 applied with the interval I,  ′′ d(ε, ε , α0 , β0 , β1 , κ, λ) with probability at least 1 − exp − d/K(ε, α0 , β0 , β1 , κ, λ) , i  h 1 log LHS ∈ RSGI¯(n/d, β, x) − ε, RSGI¯(n/d, β, x) + ε . d

(C.27)

¯ the Lipschitz constant of RSG ¯(α, β, x) over (α, β, x) ∈ [0, α0 ] × By Proposition A.3, as 0 ̸∈ I, I [β0 , β1 ] ∈ [−1, 1] is upper bounded in terms of α0 , β0 , β1 , κ, λ. Therefore, we obtain that for d ≥ d(ε, α0 , β0 , β1 , κ, λ), we have RSGI¯(n/d, β, x) − RSGI¯(α, β, x) ≤ ε. Noting    √ √  √ √  Area 1{θ1 / d ∈ I} ∩ C ∩ Sd−1 ( d) = µd 1{θ1 / d ∈ I} ∩ C Area Sd−1 ( d) , combining with (C.26), (C.27) yields for the same d and the same probability as (C.27) that   √ 1 log µd 1{θ1 / d ∈ I} ∩ C d √ (C.28) √ log(2 d) 1 d−1 ≤ β + RSGI¯(α, β, x) + 2ε + − log Area(S ( d)) . d d 75

¯ inf q∈[0,1) Φ(x, q) = β + RSG ¯(α, β, x) − 1 log(2πe) for By Lemma C.3, as α ≤ α0 and as 0 ̸∈ I, I 2 √ β = β(α, x). Moreover, recall limd→∞ d1 log Area(Sd−1 ( d)) = 12 log(2eπ). Taking β = β(α, x) in (C.28) thusgives for d ≥ d(ε, ε′′ , α, β0 , β1 , κ, λ) and with probability at least 1 − exp − d/K(ε, α, β0 , β1 , κ, λ) ,   √ 1 log µd 1{θ1 / d ∈ I} ∩ C ≤ inf Φ(x, q) + 3ε , d q∈[0,1) √ where we note limd→∞ d1 log Area(Sd−1 ( d)) − 12 log(2eπ) = 0. The result follows taking ε ← ε/3. Proof of (C.5): We proceed with a similar strategy to the proof of (C.3) above, though this direction is more involved. First, by the second part of Proposition A.3, ρ0 (α, β, x) is infinitely differentiable in β and x for (β, x) ∈ [β0 , β1 ] × [−1, 1]. We thus may define ∂ ρ0 (α, β, x) , (β,x)∈[β0 ,β1 ]×[−1,1] ∂β ∂ H = H(α, κ, λ) := sup − ρ0 (α, β, x) . (β,x)∈[β0 ,β1 ]×[−1,1] ∂β H = H(α, κ, λ) :=

inf

(C.29) (C.30)

By Lemma C.3, Proposition A.3, and compactness, H > 0. H > 0 yields strong convexity, which will be crucial  1 in our subsequent arguments. For the rest of the proof of (C.5), we consider ε < min 1, 12 H . Choice of parameters and properties of variational problem. Let w = 1 + ε and define  f¯(α, β, x) := Φ̄ x, q0 (α, β, x), ρ0 (α, β, x) − β ρ0 (α, β, x) + x2 ) , (C.31) ¯ we have where consider Φ̄(x, q, ρ) with u(x) therein such that exp u(x) = 1{x ≥ κ}. (Here as 0 ̸∈ I, 1 ′′ ¯ ¯ f = RSGI¯ − 2 log(2πe), however since I is technically defined in terms of ε , we will define εRS in terms of f¯ and then consider 0 < ε′′ ≤ εRS .) We now define the following positive parameters:  δ̄ > 0 as the unique solution to ρ0 α, β(α, x) − δ̄, x = w2 − x2 , (C.32) δ̄1 , δ̄2 := δ̄/2, n1   o ,ε , ε1 ≤ min f¯ α, β(α, x) − δ̄1 , x − δ̄1 + f¯ α, β(α, x), x 3 n1 o ε2 ≤ min H δ̄22 , ε , 6 1 β := β(α, x) − δ̄1 > β0 = . 8

(C.33) (C.34) (C.35) (C.36)

Therefore, δ̄, δ̄1 , δ̄2 , ε1 , ε2 above depend on x, ε, α, β0 , β1 , κ, λ. Later on we will show ε1 , ε2 can be taken independently of x. First we will justify the statements implicit in the above, including positivity, and derive some useful estimates. Recall by Lemma C.3, β(α, x) ≥ 14 . Now note:  1 H , we have • By definition of H, for small enough ε ≤ min 1, 12    ρ0 α, 1/8, x ≥ ρ0 α, β(α, x), x + H β(α, x) − 1/8 ≥ 1 − x2 + 3ε > w 2 − x2  > 1 − x2 = ρ0 α, β(α, x), x . 76

As ρ0 (α, β, x) is strictly decreasing on [β0 , β1 ] by Lemma C.3, it follows that such a δ̄ exists, is unique, and that δ̄ > 0. This justifies (C.32). • We have q q q   2 2 1 + ε ≥ w = ρ0 α, β(α, x) − δ̄, x + x ≥ ρ0 α, β(α, x), x + x + H δ̄ = 1 + H δ̄ , 1 thus as ε ≤ 12 H, we have δ̄ < 18 . Hence, β(α, x) − δ̄1 ≥ 14 − δ̄1 > 18 , justifying (C.36).

• By Lemma C.3, we have for all β ∈ [β0 , β1 ] that   β + f¯ α, β, x > β(α, x) + f¯ α, β(α, x), x . Taking β = β(α, x) − δ̄1 > 81 = β0 in the above justifies ε1 > 0 and hence (C.34). Moreover by the definition of ε1 , the above display gives   −δ̄1 + f¯ α, β(α, x) − δ̄1 , x > f¯ α, β(α, x), x + 3ε1 .

(C.37)

∂FI • Let F (x, q, ρ) := Φ̄(x, q, ρ)−β(ρ+x2 ) and let ρ0 = ρ0 (α, β, x), q0 = q0 (α, β, x). Note ∂F ∂ρ = ∂ρ , ∂FI ∂F ∂q = ∂q

∂F for any interval I, thus ∂F ∂ρ (x, q0 , ρ0 ) = ∂q (x, q0 , ρ0 ) = 0. Now for all β ∈ [β0 , β1 ],

 ∂F ∂ ¯ ∂ρ0 ∂F ∂q0 f (α, β, x) = − ρ0 (α, β, x) + x2 + (x, q0 , ρ0 ) · + (x, q0 , ρ0 ) · ∂β ∂ρ ∂β ∂q ∂β  2 (C.38) = − ρ0 (α, β, x) + x . Thus by definition of H, we have for all β ∈ [β0 , β1 ], ∂2 ¯ ∂ f (α, β, x) = − ρ0 (α, β, x) ≥ H > 0 . ∂β 2 ∂β

(C.39)

Now in what follows, let β = β(α, x)− δ̄1 as per (C.36), thus β − δ̄2 = β(α, x)− δ̄ as δ̄1 + δ̄2 = δ̄. By our choice of w, δ̄ from (C.32), we obtain that for this choice of β = β(α, x) − δ̄1 , ∂ 1 f¯(α, β, x) ≥ f¯(α, β − δ̄2 , x) + δ̄2 f¯(α, β − δ̄2 , x) + H δ̄22 ∂β 2  1 = f¯(α, β − δ̄2 , x) − δ̄2 ρ0 (α, β − δ̄2 , x) + x2 + H δ̄22 2 1 2 2 ¯ = f (α, β − δ̄2 , x) − δ̄2 w + H δ̄2 2 ≥ f¯(α, β − δ̄2 , x) − δ̄2 w2 + 3ε2 .

(C.40)

As our last preliminary, we claim we can take ε1 , ε2 , and  εRS := min εRSG (ε1 , α0 , β0 , β1 , κ, λ), εRSG (ε2 , α0 , β0 , β1 , κ, λ) > 0 independently of x, where εRSG is from Theorem B.1. To this end first observe   w2 − x2 = ρ0 α, β(α, x) − δ̄, x ≤ ρ0 α, β(α, x), x + H δ̄ = 1 − x2 + H δ̄ . 77

(C.41)

Since w = 1 + ε, it follows that δ̄ > H2 ε, and hence δ̄, δ̄1 , δ̄2 are lower bounded independently of x. Next note   ∂2  ∂ ∂ β + f¯(α, β, x) = 1 − ρ0 (α, β, x) + x2 , β + f¯(α, β, x) = − ρ0 (α, β, x) ≥ H > 0 2 ∂β ∂β ∂β by (C.38), (C.39). Thus as ρ0 (α, β(α, x), x) = 1 − x2 by Lemma C.3, we have  1   ∂ β + f¯(α, β, x) + H δ̄12 β(α, x) − δ̄1 + f¯ α, β(α, x) − δ̄1 , x ≥ β(α, x) + f¯ α, β(α, x), x − δ̄1 ∂β 2  1 ≥ β(α, x) + f¯ α, β(α, x), x + H δ̄12 . 2 Therefore   H f¯ α, β(α, x) − δ̄1 , x ≥ δ̄1 + f¯ α, β(α, x), x + δ̄12 . 2 Since δ̄ > H2 ε, it follows from the above display and (C.34) that ε1 can be taken independently of x. Recalling δ̄2 is lower bounded independently of x and H does not depend on x, it follows that ε2 can be taken independently of x. Hence εRS can also be taken independently of x. We now turn to establishing (C.5). We consider any ε′′ , 0 < ε′′ ≤ εRS , and define I as in (C.4). Note 0 ̸∈ I¯ as |x| > ε′′ . In the following proofs, we take β = β(α, x) − δ̄1 as per (C.36). ′′

′′ Initial Bound. Suppose x ≥ ε′′ > 0. Recall I = [ x−ε w , x + ε ] where w = 1 + ε. By (C.22) and (C.24), Z √   ¯ ∩ C ∩ {d ≤ ∥θ∥2 ≤ w2 d} exp − β∥θ∥2 dθ LHS := 1 {θ1 / d ∈ I}   √   ≤ exp(−βd)Vol 1 θ1 / d ∈ [x − ε′′ , x + ε′′ ] ∩ C ∩ d ≤ ∥θ∥ ≤ w2 d io  n √ h √  √ Z w x − ε′′ Area 1 θ1 / d ∈ w · , x + ε′′ ∩ C ∩ vSd−1 ( d) dv = exp(−βd) d w 1 Z io  n h w √  √ √ x − ε′′ ≤ exp(−βd) d Area 1 θ1 / d ∈ v , x + ε′′ ∩ vC ∩ vSd−1 ( d) dv w 1  n √ io h x − ε′′ √  √ Z w Area 1 θ1 / d ∈ , x + ε′′ ∩ C ∩ Sd−1 ( d) v d−1 dv = exp(−βd) d w 1  √ √  1 ≤ exp(−βd) √ (wd − 1)Area 1{θ1 / d ∈ I} ∩ C ∩ Sd−1 ( d) . (C.42) d

Analogously, when x ≤ −ε′′ < 0, we have by (C.22) and (C.24) that  n √ h √ Z w √  x + ε′′ io LHS ≤ exp(−βd) d Area 1 θ1 / d ∈ x − ε′′ , w · ∩ C ∩ vSd−1 ( d) dv w 1  n √ h √ Z w √  x + ε′′ io ≤ exp(−βd) d Area 1 θ1 / d ∈ v x − ε′′ , ∩ vC ∩ vSd−1 ( d) dv w 1  √ √  1 d d−1 √ ≤ exp(−βd) (w − 1)Area 1{θ1 / d ∈ I} ∩ C ∩ S ( d) . d ′′

′′ This yields the same inequality as (C.42) in the x < 0 case. (Note here that 0 ≥ x+ε w > x + ε as ′′ x ≤ −ε < 0.) Next, note

LHS = (I) − (II) − (III) ,

78

where Z

√  ¯ ∩ C exp(−β∥θ∥2 ) dθ , 1 {θ1 / d ∈ I}

Z

√ √  ¯ ∩ C ∩ {∥θ∥ ≤ d} exp(−β∥θ∥2 ) dθ , 1 {θ1 / d ∈ I}

Z

√ √  ¯ ∩ C ∩ {∥θ∥ > w d} exp(−β∥θ∥2 ) dθ . 1 {θ1 / d ∈ I}

(I) := (II) := (III) :=

We now lower bound (I) and upper bound (II), (III) via Theorem B.1. Lower bounding (I): As ε′′ ≤ εRS , by definition of εRS in (C.41), and as β = β(α, x) ∈ [β0 , β1 ] by (C.36), Theorem B.1 applied with ε = 12 ε1 ∧ ε2 and the interval I¯ implies that  for ′′ d ≥ d(ε1 , ε2 , ε , α0 , β0 , β1 , κ, λ) and probability at least 1 − exp − d/K(ε1 , ε2 , α0 , β0 , β1 , κ, λ) , √  ¯ ∩ C exp(−β∥θ∥2 ) dθ 1 {θ1 / d ∈ I} n      o ≥ max exp RSGI¯(n/d, β, x) − ε1 /2 d , exp RSGI¯(n/d, β, x) − ε2 /2 d n      o ≥ max exp RSGI¯(α, β, x) − ε1 d , exp RSGI¯(α, β, x) − ε2 d . (C.43) Z

(I) =

Here, we used the second part of Proposition A.3 in the last inequality, following identical reasoning as the proof for the upper bound to convert from RSGI¯(n/d, β, x) to RSGI¯(α, β, x). Upper bounding (II): As ε′′ ≤ εRS ≤ εRSG (ε1 , α0 , β0 , β1 , κ, λ) and β ∈ [β0 , β1 ], by Theorem  B.1, for d ≥ d(ε1 , ε′′ , α0 , β0 , β1 , κ, λ) and probability at least 1 − exp − d/K(ε1 , α0 , β0 , β1 , κ, λ) , Z (II) =

√ √  ¯ ∩ C ∩ {∥θ∥ ≤ d} exp(−β∥θ∥2 ) dθ 1 {θ1 / d ∈ I}

√ √    ¯ ∩ C ∩ {∥θ∥ ≤ d} exp − (β + δ̄1 )∥θ∥2 exp δ̄1 ∥θ∥2 dθ 1 {θ1 / d ∈ I} Z √   ¯ ∩ C exp − (β + δ̄1 )∥θ∥2 dθ ≤ exp(δ̄1 d) 1 {θ1 / d ∈ I}    ≤ exp(δ̄1 d) · exp RSGI¯(n/d, β + δ̄1 , x) + ε1 /2 d    ≤ exp(δ̄1 d) · exp RSGI¯(α, β + δ̄1 , x) + ε1 d    = exp RSGI¯(α, β + δ̄1 , x) + δ̄1 + ε1 d . (C.44) Z

=

Here we again use Proposition A.3 to convert from RSGI¯(n/d, β, x) to RSGI¯(α, β, x). Now by ¯ combining (C.43), (C.44) yields that for d ≥ (C.37) and as f¯ = RSGI¯ − 12 log(2πe) as 0 ̸∈ I,  ′′ d(ε1 , ε2 , ε , α, β0 , β1 , κ, λ), with probability at least 1 − exp d/K(ε1 , ε2 , α0 , β0 , β1 , κ, λ) ,   RSGI¯(α, β, x) − ε1 d    ≥ exp(ε1 d) exp RSGI¯(α, β + δ̄1 , x) + δ̄1 + ε1 d ≥ 4(II) .

(I) ≥ exp



79

(C.45)

Upper bounding (III): As ε′′ ≤ εRS ≤ εRSG (ε2 , α0 , β0 , β1 , κ, λ) and β ∈ [β0 , β1 ], by Theorem  B.1, for d ≥ d(ε2 , ε′′ , α0 , β0 , β1 , κ, λ) and probability at least 1 − exp − d/K(ε2 , α0 , β0 , β1 , κ, λ) , Z √ √  ¯ ∩ C ∩ {∥θ∥ > w d} exp(−β∥θ∥2 ) dθ (III) = 1 {θ1 / d ∈ I} Z √ √  ¯ ∩ C ∩ {∥θ∥ > w d} exp(−(β − δ̄2 )∥θ∥2 ) exp(−δ̄2 ∥θ∥2 ) dθ = 1 {θ1 / d ∈ I} Z √  ¯ ∩ C exp(−(β − δ̄2 )∥θ∥2 ) dθ ≤ exp(−δ̄2 w2 d) 1 {θ1 / d ∈ I}    ≤ exp(−δ̄2 w2 d) · exp RSGI¯(n/d, β − δ̄2 , x) + ε2 /2 d    ≤ exp(−δ̄2 w2 d) · exp RSGI¯(α, β − δ̄2 , x) + ε2 d    = exp RSGI¯(α, β − δ̄2 , x) − δ̄2 w2 + ε2 d . (C.46) Here, we again used Proposition A.3. By (C.40) and as f¯ = RSGI¯ − 21 log(2πe), we obtain that  for d ≥ d(ε1 , ε2 , ε′′ , α0 , β0 , β1 , κ, λ) and probability at least 1 − exp − d/K(ε1 , ε2 , α0 , β0 , β1 , κ, λ) ,    (I) ≥ exp RSGI¯(α, β, x) − ε2 d    ≥ exp(ε2 d) exp RSGI¯(α, β − δ̄2 , x) − δ̄2 w2 + ε2 d ≥ 4(III) . (C.47) Finish. Combining (C.45), (C.47), we obtain for d ≥ d(ε1 , ε2 , ε′′ , α0 , β0 , β1 , κ, λ), with probability at least 1 − exp − d/K(ε1 , ε2 , α0 , β0 , β1 , κ, λ) , we have 1 1 1 (I) − (I) = (I) . 4 4 2 Combining the above display with (C.42) and (C.43) givesfor d ≥ d(ε1 , ε2 , ε′′ , α0 , β0 , β1 , κ, λ), with probability at least 1 − exp − d/K(ε1 , ε2 , α0 , β0 , β1 , κ, λ) , √  √ √  1 d d−1 Area 1{θ1 / d ∈ I} ∩ C ∩ S ( d) ≥ d · exp(βd) · (I) 2 w −1 √    d ≥ exp − log w + β + RSGI¯(α, β, x) − ε d , (C.48) 2 ¯ where we recall that ε1 , ε2 ≤ ε. Recall β ∈ [β0 , β1 ], so by Lemma C.3 and as 0 ̸∈ I, LHS = (I) − (II) − (III) ≥ (I) −

β + RSGI¯(α, β, x) ≥ inf Φ(x, q) + q∈[0,1)

1 log(2eπ) . 2

Since log w ≤ w − 1 ≤ ε, we obtain for the same d and the same probability as (C.48), √      √ √ d 1 d−1 µd 1{θ1 / d ∈ I} ∩ C Area(S ( d)) ≥ exp inf Φ(x, q) + log(2eπ) − 2ε d . (C.49) 2 2 q∈[0,1) √ Again recall limd→∞ d1 log Area(Sd−1 ( d)) = 12 log(2eπ), and that ε1 , ε2 depend only on ε, α, β0 , β1 , κ, λ and do not depend on x as justified earlier. Rearranging (C.49)  yields for d ≥ d(ε, ε′′ , α, β0 , β1 , κ, λ) and probability at least 1 − exp − d/K(ε, α, β0 , β1 , κ, λ) ,   √ 1 log µd 1{θ1 / d ∈ I} ∩ C ≥ inf Φ(x, q) − 4ε . d q∈[0,1) Taking ε ← ε/4 completes the proof of (C.5). 80

C.3

Proof of Theorem C.1 for interpolators: κ ≥ 0

Again, we let I¯ := [x − ε′′ , x + ε′′ ], and we still have 0 ̸∈ I¯ as |x| > ε′′ . In this proof many details are analogous to Appendix C.2 above and we only highlight the differences. The notation in the following is also identical to Appendix C.2. Since κ ≥ 0, (C.21) now yields the comparison inequalities vC ⊆ C for v ≥ 1

,

vC ⊇ C for v ≤ 1 .

(C.50)

Now for any interval [a, b], we have For 1 ≤ v ≤ w :

v[a, b] ⊆ [a, bw] if b ≥ a ≥ 0

,

v[a, b] ⊆ [aw, b] if a ≤ b ≤ 0 ,

(C.51)

For w ≤ v ≤ 1 :

v[a, b] ⊇ [a, bw] if b ≥ a ≥ 0

,

v[a, b] ⊇ [aw, b] if a ≤ b ≤ 0 .

(C.52)

Proof of (C.3): Let w = 1 + ε′′ for the proof of this upper bound. For any β ∈ [β0 , β1 ], we obtain the following from (C.50), (C.51). When x > ε′′ we have for d ≥ d(ε′′ ), Z √  ¯ ∩ C exp(−β∥θ∥2 ) dθ LHS := 1 {θ1 / d ∈ I}  n √ o √ Z w √  2 ≥ exp(−βw d) d Area 1 θ1 / d ∈ [x − ε′′ , x + ε′′ ] ∩ C ∩ vSd−1 ( d) dv Z1 w  n √ h √  √ x + ε′′ io ∩ vC ∩ vSd−1 ( d) dv ≥ exp(−βw2 d) d Area 1 θ1 / d ∈ v x − ε′′ , w 1   1 √ √ ≥ exp(−βw2 d) Area 1{θ1 / d ∈ I} ∩ C ∩ Sd−1 ( d) √ . 2 d Similarly when x < −ε′′ , for d ≥ d(ε′′ ) we have  n √ o √ Z w √  2 LHS ≥ exp(−βw d) d Area 1 θ1 / d ∈ [x − ε′′ , x + ε′′ ] ∩ C ∩ vSd−1 ( d) dv 1  n √ h x − ε′′ io √  √ Z w 2 Area 1 θ1 / d ∈ v , x + ε′′ ∩ vC ∩ vSd−1 ( d) dv ≥ exp(−βw d) d w 1  √ √  1 ≥ exp(−βw2 d) Area 1{θ1 / d ∈ I} ∩ C ∩ Sd−1 ( d) √ . 2 d ¯ we have for d ≥ d(ε, ε′′ , α0 , β0 , β1 , κ, λ) In either case, applying Theorem B.1 with the interval I,  and with probability at least 1 − exp − d/K(ε, α0 , β0 , β1 , κ, λ) , i  h 1 log LHS ∈ RSGI¯(n/d, β, x) − ε, RSGI¯(n/d, β, x) + ε . d

(C.53)

We apply the second part of Proposition A.3 identically as the proof in the κ < 0 case to convert the first argument of RSGI¯(n/d, β, x) to RSGI¯(α, β, x) for d ≥ d(ε, α0 , β0 , β1 , κ, λ). We obtain for the same d and the same probability as (C.53),   √ √ √ 1 1 1 log µd 1{θ1 / d ∈ I} ∩ C ≤ βw2 + RSGI¯(α, β, x) + Kε − log Area(Sd−1 ( d)) + log(2 d) . d d d ¯ and finish identically We now take β = β(α, x) in the above, use Lemma C.3 and the fact that 0 ̸∈ I, to the proof of the upper bound in the κ < 0 case. Here we account for the presence of βw2 rather than β by noting β(w2 − 1) ≤ 3ε′′ β1 ≤ 30ε. This proves (C.3) for the κ ≥ 0 case. 81

Proof of (C.5): We let w = 1 − ε. We define H = H(α, κ, λ), H = H(α, κ, λ) identically as (C.29), (C.30) and define f¯ identically as (C.31). We now let δ̄ = δ̄(ε, α, β0 , β1 , κ, λ) := min

1 o ε . 4 2H

n1

,

∂ By definition of H and as ∂β ρ0 (α, β, x) < 0 by Lemma C.3, it follows that uniformly in x ∈ [−1, 1],

ρ0 (α, β(α, x), x) − ρ0 (α, β(α, x) + 2δ̄, x) ≤ ε , and hence by Lemma C.3, ρ0 (α, β(α, x) + 2δ̄, x) ≥ 1 − x2 − ε .

(C.54)

Also note 2δ̄ + β(α, x) < β1 = 10, where we use that β(α, x) ≤ 9. By (C.38), (C.39), and Lemma C.3, we obtain 1 ∂ f¯(α, β(α, x) + δ̄, x) ≥ f¯(α, β(α, x), x) + δ̄ f¯(α, β(α, x), x) + H δ̄ 2 ∂β 2  1 = f¯(α, β(α, x), x) − δ̄ ρ0 (α, β(α, x), x) + x2 + H δ̄ 2 2 1 2 = f¯(α, β(α, x), x) − δ̄ + H δ̄ . 2

(C.55)

Similarly we obtain  1 f¯(α, β(α, x) + δ̄, x) ≥ f¯(α, β(α, x) + 2δ̄, x) − (−δ̄) ρ0 (α, β(α, x) + 2δ̄, x) + x2 + H δ̄ 2 2 1 ≥ f¯(α, β(α, x) + 2δ̄, x) + δ̄(1 − ε) + H δ̄ 2 2 1 2 ≥ f¯(α, β(α, x) + 2δ̄, x) + δ̄w + H δ̄ 2 , (C.56) 2 where we used (C.54) and w = 1 − ε combined with ε < 1. Finally, we let ε1 , ε2 := min

n1

o H δ̄ 2 , ε ,

6 εRS := min εRSG (ε1 , α0 , β0 , β1 , κ, λ), εRSG (ε2 , α0 , β0 , β1 , κ, λ) > 0 .

(C.57) (C.58)

We now turn to establishing (C.5). We consider any ε′′ , 0 < ε′′ ≤ εRS , and define I as in (C.4). Thus 0 ̸∈ I¯ as |x| > ε′′ . We now let β = β(α, x) + δ̄ for the rest of this proof. First suppose x ≥ ε′′ > 0. Here we have by (C.50), (C.52) that Z √   ¯ ∩ C ∩ {w2 d ≤ ∥θ∥2 ≤ d} exp − β∥θ∥2 dθ LHS := 1 {θ1 / d ∈ I}  n √ h  x + ε′′ io ≤ exp(−βw2 d)Vol 1 θ1 / d ∈ x − ε′′ , w · ∩ C ∩ {w2 d ≤ ∥θ∥ ≤ d} w Z 1  n √ h √  √ x + ε′′ io ≤ exp(−βw2 d) d Area 1 θ1 / d ∈ v x − ε′′ , ∩ vC ∩ vSd−1 ( d) dv w w   √ √ 1 ≤ exp(−βw2 d) √ Area 1{θ1 / d ∈ I} ∩ C ∩ Sd−1 ( d) . (C.59) d 82

Analogously, when x ≤ −ε′′ < 0, we have  n √ h io √ Z 1 √  x − ε′′ 2 Area 1 θ1 / d ∈ w · LHS ≤ exp(−βw d) d , x + ε′′ ∩ C ∩ vSd−1 ( d) dv w w Z 1  n √ h x − ε′′ io √ √  ≤ exp(−βw2 d) d Area 1 θ1 / d ∈ v , x + ε′′ ∩ vC ∩ vSd−1 ( d) dv w w  √ √  1 ≤ exp(−βw2 d) √ Area 1{θ1 / d ∈ I} ∩ C ∩ Sd−1 ( d) . d ′′

(Note here that 0 ≥ x − ε′′ > x−ε w .) Now note LHS = (I) − (II) − (III) ,

where Z

√  ¯ ∩ C exp(−β∥θ∥2 ) dθ , 1 {θ1 / d ∈ I}

Z

√ √  ¯ ∩ C ∩ {∥θ∥ ≥ d} exp(−β∥θ∥2 ) dθ , 1 {θ1 / d ∈ I}

Z

√ √  ¯ ∩ C ∩ {∥θ∥ ≤ w d} exp(−β∥θ∥2 ) dθ . 1 {θ1 / d ∈ I}

(I) := (II) := (III) :=

We again lower bound (I) and upper bound (II), (III) via Theorem B.1. First, by Theorem B.1 with interval I¯ and the second part of Proposition A.3, as ε′′ ≤ εRS and β ∈ [β0 , β1 ], we have for d ≥ d(ε1 , ε2 , α0 , β0 , β1 , κ, λ) and probability at least 1 − exp − d/K(ε1 , ε2 , α0 , β0 , β1 , κ, λ) , n      o (I) ≥ max exp RSGI¯(α, β, x) − ε1 d , exp RSGI¯(α, β, x) − ε2 d . By Theorem B.1 and Proposition A.3, for the same d and with the same probability, Z √ √  ¯ ∩ C ∩ {∥θ∥ ≥ d} exp(−(β − δ̄)∥θ∥2 ) exp(−δ̄∥θ∥2 ) dθ (II) = 1 {θ1 / d ∈ I} Z √  ¯ ∩ C exp(−(β − δ̄)∥θ∥2 ) dθ ≤ exp(−δ̄d) 1 {θ1 / d ∈ I}    ≤ exp RSGI¯(α, β − δ̄, x) − δ̄ + ε1 d . Analogously, for the same d and with the same probability, Z √ √  ¯ ∩ C ∩ {∥θ∥ ≤ w d} exp(−(β + δ̄)∥θ∥2 ) exp(δ̄∥θ∥2 ) dθ (III) = 1 {θ1 / d ∈ I}    ≤ exp RSGI¯(α, β + δ̄, x) + δ̄w2 + ε2 d . ¯ Thus applying (C.55), (C.56) to Recall β − δ̄ = β(α, x) and f¯ = RSGI¯ − 21 log(2πe) as 0 ̸∈ I. upper bound (II), (III) respectively gives for d ≥ d(ε1 , ε2 , α0 , β0 , β1 , κ, λ) and probability at least  1 − exp − d/K(ε1 , ε2 , α0 , β0 , β1 , κ, λ) , we have (II), (III) ≤ 41 (I). Therefore as ε1 , ε2 ≤ ε, for the same d and with the same probability, LHS ≥

   1 1 (I) ≥ exp RSGI¯(α, β, x) − ε d . 2 2 83

Via the same steps as the κ < 0 case, combining the above display with (C.59) yields that for the same d and with the same probability,  √d    1 √ √ 1 1 d−1 2 log µd 1{θ1 / d ∈ I} ∩ C + log Area(S ( d)) ≥ βw + RSGI¯(α, β, x) − ε + log . d d d 2 ¯ we recall Note that as β ≤ β1 = 10, we have βw2 ≥ β − 3βε ≥ β − 30ε. Finally, as 0 ̸∈ I, β + RSGI¯(α, β, x) ≥ inf q∈[0,1) Φ(x, q) by Lemma C.3. Combining these steps with the above display and taking ε ← ε/K for universal constant K, (C.5) follows.

C.4

Proof of Theorem C.2 for posterior

Much of this proof is identical to the above and we will only highlight the differences. We again define I¯ := [x−ε′′ , x+ε′′ ] and follow the same C.2, C.3. Similarly as before, √ notation as Appendices √ we parametrize Rd by the scalar v = ∥θ∥/ d and θ ∈ Sd−1 ( d). By Lipschitzness of u(x), √ instead of the comparison inequalities relating vC and C, we directly  study how u (si θ1 + ⟨gi , θ̄⟩)/ d changes in v via the following Lemma. Lemma C.4. With probability at least 1 − exp(−d/KC.4 ), we have for all v ≥ 0, sup√

n X

θ∈Sd−1 ( d)

n  X  u Si (vθ) − u Si (θ) ≤ KC.4 |v − 1|d ,

i=1

i=1

where KC.4 ≥ 1 depends on λ. Here probability is over the si , gi . Proof of Lemma C.4. Note Si (vθ) = vSi (θ). As u(·) is D-Lipschitz, we have sup√

θ∈Sd−1 ( d)

n X

n  X  u vSi (θ) − u Si (θ)

i=1

i=1 n X

≤ D|v − 1|

Si (θ) sup√ θ∈Sd−1 ( d) i=1 n  X

≤ D|v − 1|

|si θ1 / d| +

sup√

θ∈Sd−1 ( d) i=1

sup√

n X

√  ⟨gi , θ̄⟩/ d

θ∈Sd−1 ( d) i=1



:= D|v − 1| (I) + (II) .

(C.60)

We now establish high-probability upper bounds onP(I), (II). Starting √ √ √ Pwith (I), observe for any θ ∈ Sd−1 ( d) that θ1 / d ≤ 1. Hence supθ∈Sd−1 (√d) ni=1 |si θ1 / d| ≤ ni=1 |si |. Recall the s2i are sub-Exponential with parameter Ksignal = K(λ), and hence the |si | are K(λ) sub-Gaussian. As  n ≤ 2d, Hoeffding’s Inequality implies with probability at least 1 − exp − d/K(λ) , we have (I) ≤ K(λ)d . (C.61) √ √ We now upper bound (II). Since ∥θ̄∥ ≤ d, (II) is maximized for ∥θ̄∥ = d. Shifting d by 1, it suffices to consider F (G) :=

sup√

n X

√ ⟨gi , θ⟩/ d

θ∈Sd−1 ( d) i=1

where now gi ∼ N (0, Id ) and where G refers to the n × d Gaussian matrix with rows gi . 84

Consider two Gaussian matrices G0 = {gi0 }ni=1P , G1 = {gi1 }ni=1√ whose rows gi0 , gi1 are i.i.d. from n ′ 1 N (0, Id ). If F (G1 ) ≥ F (G0 ), by continuity of i=1 ⟨gi , θ⟩/ d in θ, we have for some θ ∈ √ d−1 S ( d), F (G1 ) − F (G0 ) =

n X

⟨gi1 , θ′ ⟩/

d − F (G0 ) ≤

i=1

n X

⟨gi1 , θ′ ⟩/

d −

i=1

n X

√ ⟨gi0 , θ′ ⟩/ d

i=1

n X

√ ⟨gi1 − gi0 , θ′ ⟩/ d

i=1

n X

∥gi1 − gi0 ∥

i=1

n X d 1/2 √ X 1 0 2 . n (gij − gij ) i=1 j=1

√ An analogous bound applies if F (G0 ) ≥ F (G1 ). Therefore F (G) is n-Lipschitz as a function of G measured w.r.t. Euclidean norm. Clearly the density of G is 1 strongly log-concave. It follows by concentration of Lipschiz functions of strongly log-concave measures, Theorem B.3, that     t2  P F (G) ≥ E F (G) + t) ≤ exp − , (C.62) Kn   where probability is over G. Next we upper bound E F (G) . Writing u = (u1 , . . . , un ) where each ui ∈ {±1}, observe that n

n

d

X XX 1 1 √ sup u ⟨g , θ⟩ = sup ui θj gij . F (G) = √ i i d θ∈Sd−1 (√d),u∈{±1}n i=1 d θ∈Sd−1 (√d),u∈{±1}n i=1 j=1 P P Define the centered Gaussian processes Xθ,u := ni=1 dj=1 ui θj gij , Yθ,u := ∥u∥2 ⟨g, θ⟩ + ∥θ∥2 ⟨h, u⟩, where√u = (u1 , . . . , un )T , and where g ∼ N (0, Id ), h ∼ N (0, In ) are independent. For any θ, θ′ ∈ Sd−1 ( d), u, u′ ∈ {±1}n , observe   E (Xθ,u − Xθ′ ,u′ )2 = ∥θ∥2 ∥u∥2 + ∥θ′ ∥2 ∥u′ ∥2 − 2⟨u, u′ ⟩⟨θ, θ′ ⟩ ,    E (Yθ,u − Yθ′ ,u′ )2 = 2∥θ∥2 ∥u∥2 + 2∥θ′ ∥2 ∥u′ ∥2 − 2 ∥u∥∥u′ ∥⟨θ, θ′ ⟩ + ∥θ∥∥θ′ ∥⟨u, u′ ⟩ . √   √ As ∥θ∥ = ∥θ′ ∥ = d, ∥u∥ = ∥u′ ∥ = n and n − ⟨u, u′ ⟩ d − ⟨θ, θ′ ⟩ ≥ 0, it follows that     E (Xθ,u − Xθ′ ,u′ )2 ≤ E (Yθ,u − Yθ′ ,u′ )2 . √ √ By the Sudakov-Fernique inequality (see e.g. Theorem 7.2.8, [Ver18]), noting ∥u∥ = n, ∥θ∥ = d √ always holds for u ∈ {±1}n , θ ∈ Sd−1 ( d), we obtain h i √   d E F (G) = E sup X θ,u √ θ∈Sd−1 ( d),u∈{±1}n

≤E ≤

√ √

=

h

sup √

Yθ,u

i

θ∈Sd−1 ( d),u∈{±1}n

nE



 √  sup√ ⟨g, θ⟩ + d E

θ∈Sd−1 ( d)

  nd E ∥g∥2 +

= O(n1/2 d) . 85

  d E ∥h∥1

 sup ⟨h, u⟩ u∈{±1}n

  Using n ≤ 2d, it follows that E F (G) ≤ O(d). Combining with (C.62) gives that with probability at least 1 − exp(−d/K), we have (II) ≤ Kd .

(C.63)

Now combining (C.60), (C.61), (C.63) with a Union Bound, the Lemma follows. We now have the technical ingredients in place to prove Theorem C.2. Proof of (C.9): Let K ′ (λ) := KC.4 (λ), where K ′ is as in the statement of Theorem. We thus consider ε′′ ≤ min{ε/KC.4 (λ), εRSG } and let I be as per (C.8). Let w = 1 + ε′′ . For any β ∈ [β0 , β1 ], by (C.51), we obtain the following, very similarly to the proof of (C.3). When x > ε′′ , we have for d ≥ d(ε′′ ), Z n X  √  ¯ exp LHS := 1{θ1 / d ∈ I} u Si (θ) − β∥θ∥2 dθ i=1

Z ≥

n o X  n √   1 θ1 / d ∈ [x − ε′′ , x + ε′′ ] ∩ d ≤ ∥θ∥2 ≤ w2 d exp u Si (θ) − β∥θ∥2 dθ i=1

√ Z wZ ≥ exp(−βw2 d) d 1

√ Sd−1 ( d)

n X √  1 vθ1 / d ∈ [x − ε′′ , x + ε′′ ] v d−1 exp u Si (vθ) dθdv



i=1

n X n √ h ′′ io  d−1 ′′ x + ε v exp dθdv , 1 θ / d ∈ x − ε , u S (vθ) 1 i √ w Sd−1 ( d)

√ Z wZ ≥ exp(−βw2 d) d 1

i=1

where the last inequality uses (C.51). When x < −ε′′ , we have for d ≥ d(ε′′ ), now using (C.51) in the a ≤ b ≤ 0 case, n io X n √ h x − ε′′  d−1 ′′ v exp dθdv . 1 θ / d ∈ , x + ε u S (vθ) 1 i √ w 1 Sd−1 ( d) i=1 √  By Lemma C.4, with probability at least 1 − exp − d/K(λ) , we have for all θ ∈ Sd−1 ( d) and all v, 1 ≤ v ≤ w that

√ Z wZ LHS ≥ exp(−βw d) d 2

n X

n  X  u Si (vθ) − u Si (θ) ≤ KC.4 |v − 1|d ≤ εd ,

i=1

i=1

where we use that 1 ≤ v ≤ w ≤ 1 + ε/KC.4 . Combining the above displays and recalling the  ′′ definition of I in (C.8) yields that for d ≥ d(ε ), with probability at least 1 − exp d/K(λ) , Z wZ n X √  √  2 d−1 LHS ≥ exp − βw d − εd d 1 θ / d ∈ I v exp u S (θ) dθdv 1 i √ 1

exp

− βw2 d − εd √ 2 d

Sd−1 ( d)

Z √ Sd−1 ( d)

 √ 1 θ1 / d ∈ I exp

i=1 n X

 u Si (θ) dθ .

i=1

That is, √ 1 1 1 log Z(I) ≤ log(LHS) + log(2 d) + βw2 + ε . d d d

(C.64)

With (C.64), the rest of the proof of (C.9) now follows identically as the proof of (C.3), the one difference being that we apply Theorem B.2 rather than Theorem C.2 to control d1 log(LHS). 86

ε . Recall that Lemma C.3, Proposition Proof of (C.11): Let KC.4 = KC.4 (λ) ≥ 1, and w = 1− KC.4 A.3 still hold here, where the RS equations are given in terms of u(x) rather than the hard indicator 1{· ≥ κ}. Thus we may perform the following steps. In particular, define H, H as in (C.29), (C.30). We now define  f¯(α, β, x) := Φ̄ x, q0 (α, β, x), ρ0 (α, β, x) − β ρ0 (α, β, x) + x2 ) , (C.65)

where consider Φ̄(x, q, ρ) given with u(x) (rather than the hard indicator function 1{· ≥ κ} for the interpolators). We now let n1 o 1 , ε . δ̄ = δ̄(ε, α, β0 , β1 , λ) := min 4 2HKC.4 Analogously as the proof of (C.54), we have uniformly in x ∈ [−1, 1]: ε . ρ0 (α, β(α, x) + 2δ̄, x) ≥ 1 − x2 − KC.4

(C.66)

Again analogously the proof of (C.55), (C.56), now using (C.66) and ε < 1, KC.4 ≥ 1, we obtain 1 f¯(α, β(α, x) + δ̄, x) ≥ f¯(α, β(α, x), x) − δ̄ + H δ̄ 2 , 2 1 ¯ ¯ f (α, β(α, x) + δ̄, x) ≥ f (α, β(α, x) + 2δ̄, x) + δ̄w2 + H δ̄ 2 . 2

(C.67) (C.68)

Finally, we again let ε1 = ε2 := min

n1

o H δ̄ , ε , 2

6 εRS := min εRSG (ε1 , α0 , β0 , β1 , λ), εRSG (ε2 , α0 , β0 , β1 , λ) > 0 ,

(C.69) (C.70)

where now εRSG comes from Theorem B.2. We now turn to establishing (C.11). Again, consider any ε′′ , 0 < ε′′ ≤ εRS , and define I as in (C.10). Thus 0 ̸∈ I¯ as |x| > ε′′ . We now let β = β(α, x) + δ̄ for the rest of this proof. First suppose x ≥ ε′′ > 0. Here we have by (C.52) that Z n X  √   2 2 ¯ LHS := 1 {θ1 / d ∈ I} ∩ {w d ≤ ∥θ∥ ≤ d} exp u Si (θ) − β∥θ∥2 dθ i=1

√ Z 1Z 2 ≤ exp(−βw d) d w

√ Sd−1 ( d)

n X √   ′′ ′′ d−1 1 vθ1 / d ∈ [x − ε , x + ε ] v exp u Si (vθ) dθdv i=1

n n √ h X ′′ io  ′′ x + ε d−1 1 θ / d ∈ x − ε , v exp u S (vθ) dθdv . 1 i √ w Sd−1 ( d)

√ Z 1Z 2 ≤ exp(−βw d) d w

i=1

Analogously, when x ≤ −ε′′ < 0, we have n io X h x − ε′′ √ Z 1 n √  2 , x + ε′′ v d−1 exp LHS ≤ exp(−βw d) d 1 θ1 / d ∈ u Si (vθ) dθdv . w w i=1

√ By Lemma C.4, with probability at least 1 − exp − d/K(λ) , we have for all θ ∈ Sd−1 ( d) and all v, w ≤ v ≤ 1 that 

n X i=1

n  X  u Si (vθ) − u Si (θ) ≤ KC.4 |v − 1|d ≤ εd , i=1

87

where we use that 1 − ε/KC.4 = w ≤ v ≤ 1. Note as KC.4 ≥ 1, w ≥ 1 − ε. Therefore by the definition of I in (C.10), h h x − ε′′ i x + ε′′ i x − ε′′ , ⊆ I when x ≥ ε′′ , , x + ε′′ ⊆ I when x ≤ −ε′′ . w w  Combining the above displays now yields that with probability at least 1 − exp d/K(λ) , 2

LHS ≤ exp − βw d + εd

√

Z 1Z d w

exp

− βw2 d + εd √

d

√ Sd−1 ( d)

Z √ Sd−1 ( d)

n X  √  d−1 1 θ1 / d ∈ I v exp u Si (θ) dθdv i=1

n X  √  1 θ1 / d ∈ I exp u Si (θ) dθ . i=1

That is, √ 1 1 1 log Z(I) ≥ log(LHS) + log d + βw2 − ε . d d d

(C.71)

With (C.71), the rest of the proof of (C.11) now follows identically as the proof of (C.5) in the κ ≥ 0 case. We again write LHS = (I) − (II) − (III) ,

 Pn where now 1{C} in the definition of (I), (II), (III) is replaced by exp i=1 u(Si (θ)) . (Observe that in the proof of (C.5), C takes the same form, for u(x) = 1{x ≥ κ}.) We now apply Theorem B.2 to establish high-probability formulas for (I), (II), (III) up to accuracy ε1 , ε2 , where ε1 , ε2 are as defined in (C.69). Combined with (C.67), (C.68), this proves that 1 (I) ≥ (II), (III) . 4 Combining with (C.71) and finishing as in the proof of (C.5) now establishes (C.11).

D

Technical results for proof of Theorems A.1, A.2

Here we prove the intermediate technical results Lemma A.5, Lemma A.8, and Proposition A.6 on the interpolation argument and concentration of measure, used in the proof of Theorems A.1 and A.2. Throughout Appendix D, for the interpolators we let D be as stated in Theorem A.1, and for the posterior we recall that |ul | ≤ D = D(λ) for l = 1, 2, 3, 4. Also here in Appendix D, as per the proof for the interpolators, we will always write dependence on κ in K when it is present in the following proofs; note there is no such dependence for the posterior. We first prove Lemma A.5, which lets us ‘transfer’ expectations of the form νt,v (f ) for 0 ≤ v ≤ 1 to expectations νt,1 (f ) when v = 1, using some of the ideas from cavity in n. We then finish the cavity in n argument and execute thecavity in d argument to prove Lemma A.8, which gives us  d d 1 2 1 2 arising in the interpolation argument. Both control over the derivatives dt νt θd θd , dt νt (θd ) these arguments will crucially use Proposition A.6 on concentration of the interpolating Gibbs measure νt,v (·), which we prove later in Appendix D.3. The proof of Proposition A.6 will crucially use classical results on concentration of Lipschitz functions of strongly log-concave measures; note the quadratic component −β∥θ∥2 we added when relaxing to the Gaussian measure is strongly concave and u(·) in Theorem A.1 is concave. 88

D.1

Proof of Lemma A.5

Our first step will be to establish the following on the derivative of νt,v w.r.t. v. 1 2 Lemma D.1. Consider any t ∈ [0, 1] and a function Bv that is either 1, u′ (Sn,t,v )u′ (Sn,t,v ), or 1 1 u′′ (Sn,t,v ) + u′ (Sn,t,v )2 . Then for any f : (Rd−1 )L → R,

K

X

νt,v f Rℓ,ℓ − ρn,d



X

+

∂ ∂v νt,v (f Bv )

νt,v f Rℓ1 ,ℓ2 − qn,d

1≤ℓ1 ̸=ℓ2 ≤L+2

ℓ≤L+1

is upper bounded by 

1 + ε + νt,v (f 2 )1/2 d 

′′

! ,

where K depends on D, β0 , β1 , λ, L, and also on κ for the interpolators. For proving Lemma A.5 we will only need Lemma D.1 with Bv = 1. However the full generality 1 2 where Bv equals 1, u′ (Sn,t,v )u′ (Sn,t,v ) = u′ (Sv1 )u′ (Sv2 ), or u′′ (Sv1 ) + u′ (Sv1 )2 will be needed in proving Lemma A.8 later, specifically in the proof of Lemma D.4. ℓ . Recall Proof of Lemma D.1. Throughout this proof, we use the shorthand Svℓ = Sn,t,v  P ℓ h g(v) i f Bv exp ℓ≤L u(Sv ) t,∼ νt,v (f Bv ) = E = E ,  L h(v) exp u(SvL+1 ) t,∼

where we let D X E g(v) := f Bv exp u(Svℓ ) ℓ≤L

t,∼

, h(v) :=

D

EL exp u(SvL+1 ) . t,∼

Note the denominator of the Gibbs average defining ⟨·⟩t,∼ does not depend on v. We thus compute D EL−1 D  ∂ ∂ L+1 E h(v) = L exp u(SvL+1 ) S , exp u(SvL+1 ) u′ (SvL+1 ) · ∂v ∂v v t,∼ t,∼ and D X  X E ∂ ∂ ℓ g(v) = f Bv exp u(Svℓ ) u′ (Svℓ ) · Sv + r(v) , ∂v ∂v t,∼ ℓ≤L

ℓ≤L

where   0 ∂ 1 ∂ 2 r(v) := u′′ (Sv1 )u′ (Sv2 ) · ∂v Sv + u′ (Sv1 )u′′ (Sv2 ) · ∂v Sv   ′′′ 1 ∂ 1 ∂ u (Sv ) · ∂v Sv + 2u′ (Sv1 )u′′ (Sv1 ) · ∂v Sv1

: Bv = 1 , : Bv = u′ (Sv1 )u′ (Sv2 ) , : Bv = u′′ (Sv1 ) + u′ (Sv1 )2 .

Note  ∂ ℓ  θ1ℓ 1  Sv = √ − x sn + √ ∂v d 2 vd

X 2≤j≤d−1

gn,j θjℓ +

  p 1 √ tgn,d θdℓ − √ qn,d Z + ρn,d − qn,d W ℓ . 2 1−v

Collecting common terms into the ‘signal part’ (I) corresponding to sn and the rest of the expression ∂ ℓ (II) for ∂v Sv , we obtain ∂ νt,v (f Bv ) = (I) + (II) , ∂v 89

where (I), (II) are defined by (I) := νt,v f Bv

X

u (Svℓ ) ·

ℓ≤L

! !!  θL+1  θℓ   ′ L+1 1 1 √ − x sn + r1 (v) − Lνt,v f Bv · u (Sv ) · √ − x sn , d d

where  0    1   2   θ θ ′′ 1 ′ 2 ′ 1 ′′ 2 r1 (v) := u (Sv )u (Sv ) · √1d − x sn + u (Sv )u (Sv ) · √1d − x sn   1   θ1  u′′′ (S 1 ) + 2u′ (S 1 )u′′ (S 1 ) · √ −x s v

v

v

n

d

: Bv = 1 , : Bv = u′ (Sv1 )u′ (Sv2 ) , : Bv = u′′ (Sv1 ) + u′ (Sv1 )2 ,

and (II) := νt,v f Bv

X

 u′ (Svℓ )S̄vℓ + r2 (v)

!

  − Lνt,v f Bv u′ (SvL+1 )S̄vL+1 ,

ℓ≤L

where  √  p 1  X 1 √ S̄vℓ := √ gn,j θjℓ + tgn,d θdℓ − √ qn,d Z + ρn,d − qn,d W ℓ , 2 1−v 2 vd 2≤j≤d−1   : Bv = 1 , 0 ′′ 1 ′ 2 1 ′ 1 ′′ 2 2 r2 (v) := u (Sv )u (Sv )S̄v + u (Sv )u (Sv )S̄v : Bv = u′ (Sv1 )u′ (Sv2 ) ,    ′′′ 1 u (Sv ) + 2u′ (Sv1 )u′′ (Sv1 ) S̄v1 : Bv = u′′ (Sv1 ) + u′ (Sv1 )2 . First we upper bound (I) , via an analogous derivation as in the proof of Proposition A.4 (specifically, the derivation of the upper bound on (I) therein). Since exp(·) ≥ 0, we have D E  P θ1ℓ ℓ √ f sn d − x hℓ · exp X 1≤ℓ≤L+1 u(Sv ) t,∼ , (I) ≤ K(L) E L+1 L+2 exp u(Sn,t,v ) t,∼ 1≤ℓ≤L+1 ′

where hℓ is a sum of products of u(l ) (s1v ), . . . , u(l ) (sL+1 ) for l′ = 1, 2, 3, given by the above definition v for (I). Thus we √ have |hℓ | ≤ K(D, L) by our bounds on the first 4 derivatives of u(·). Due to the indicator 1{θ1ℓ / d ∈ I} in the definition of the Gibbs measure ⟨·⟩t,∼ from (A.24), which is applied  θℓ for all 1 ≤ ℓ ≤ L + 1 since each f sn √1d − x hℓ is a function of all the L + 1 replicas, we can upper bound the numerator in the above expression as follows: D

 θℓ   X E D  X E f sn √1 − x hℓ · exp u(Svℓ ) ≤ ε′′ K(D, L) · f sn exp u(Svℓ ) . t,∼ t,∼ d 1≤ℓ≤L+1 1≤ℓ≤L+1

Thus by Cauchy-Schwarz,  (I) ≤ ε′′ K(D, L) · νt,v |f ||sn | ≤ ε′′ K(D, L) · νt,v (s2n )1/2 νt,v (f 2 )1/2 . Since sn does not depend on the replicas θ1 , . . . , θL or the annealed disorder W 1 , . . . , W L , we have νt,v (s2n )1/2 = E[s2n ]1/2 ≤ K(λ). Combining with the above yields (I) ≤ ε′′ K(D, λ, L) · νt,v (f 2 )1/2 . 90

(D.1)

Next we upper bound (II) . We simplify (II) via Guassian Integration by Parts w.r.t. the gn,j , gn,d , z, W ℓ . We give the computation in the case where Bv = u′′ (Sv1 )+u′ (Sv1 )2 , the computation when Bv = 1 or Bv = u′ (Sv1 )u′ (Sv2 ) is analogous. This computation is analogous to the proof of Lemma 3.2.8, [Tal10]; this is because the derivatives of Svℓ w.r.t. the gn,j , gn,d , z, W ℓ are the same as when the Svℓ are defined without the signal part. A similar computation is detailed on the proof of Lemma 2.3.2, p. 172 of [Tal10]. Spelling out the computation in full detail when Bv = u′′ (Sv1 ) + u′ (Sv1 )2 , we obtain ! X  u′ (Svℓ )S̄vℓ + r2 (v) νt,v f Bv ℓ≤L

 1 = νt,v f (R1,1 − ρn,d )u′ (Sv1 ) u′′′ (Sv1 ) + 2u′ (Sv1 )u′′ (Sv1 ) 2    1 X + νt,v f (R1,1 − ρn,d ) u′′ (Svℓ ) + u′ (Svℓ )2 Bv 2 1≤ℓ≤L   1 X νt,v f (R1,l − qn,d )u′ (Svℓ ) u′′′ (Sv1 ) + 2u′ (Sv1 )u′′ (Sv1 ) + 2 2≤ℓ≤L   1 X ′ νt,v f (Rl,l′ − qn,d )u′ (Svℓ )u′ (Svℓ )Bv + 2 1≤ℓ̸=ℓ′ ≤L   L X − νt,v f (Rl,L+1 − qn,d )u′ (Svℓ )u′ (SvL+1 )Bv + err1 , 2 

(D.2)

1≤ℓ≤L

and   1  2 νt,v f Bv r2 (v) = νt,v f (R1,1 − ρn,d ) u′′′ (Sv1 ) + 2u′ (Sv1 )u′′ (Sv1 ) 2  + u′′′′ (Sv1 ) + 2u′′ (Sv1 )2 + 2u′ (Sv1 )u′′′ (Sv1 ) Bv !   ′ 1 ′′′ 1 ′ 1 ′′ 1 + u (Sv ) + 2u (Sv )u (Sv ) u (Sv )Bv

(D.3)

   1 X νt,v f (R1,l − qn,d ) u′′′ (Sv1 ) + 2u′ (Sv1 )u′′ (Sv1 ) u′ (Sv1 )Bv 2 2≤ℓ≤L    L − νt,v f (R1,L+1 − qn,d ) u′′′ (Sv1 ) + 2u′ (Sv1 )u′′ (Sv1 ) u′ (SvL+1 ) + err2 , 2

+

and   1    νt,v f Bv u′ (SvL+1 )S̄vL+1 = νt,v f (RL+1,L+1 − ρn,d ) u′′ (SvL+1 + u′ (SvL )2 ) Bv 2   1 + νt,v f (R1,L+1 − qn,d )u′ (SvL+1 ) u′′′ (Sv1 ) + 2u′ (Sv1 )u′′ (Sv1 ) 2   (D.4) 1 X + νt,v f (Rl,L+1 − qn,d )u′ (Svℓ )u′ (SvL+1 )Bv 2 1≤ℓ≤L   L+1 − νt,v f (RL+1,L+2 − qn,d )u′ (SvL+1 )u′ (SvL+2 )Bv + err3 , 2 where err1 , err2 , err3 are such that  1−t |err1 |, |err2 |, |err3 | ≤ · Kνt,v |f | max (θdℓ )2 . 1≤ℓ≤L+2 2d 91

In particular, err1 , err2 , err3 arise because the last coordinate θdℓ is interpolated via the parameter t ∈ [0, 1]. By the second part of Proposition A.6 and Cauchy-Schwarz we have |err1 |, |err2 |, |err3 | ≤

K(D, β0 , β1 , κ, λ, L) · νt,v (f 2 )1/2 . d

(D.5)

Recall that (II) = νt,v f Bv

X



!

u (Svℓ )S̄vℓ + r2 (v)

  − Lνt,v f Bv u′ (SvL+1 )S̄vL+1 .

ℓ≤L

Thus combining (D.2), (D.3), (D.4) and using the bound (D.5) on |err1 |, |err2 |, |err3 |, we obtain  X  (II) ≤ K(D, β0 , β1 , κ, λ, L) νt,v f Rℓ,ℓ − ρn,d ℓ≤L+1

+

X

νt,v f Rℓ1 ,ℓ2 − qn,d

1≤ℓ1 ̸=ℓ2 ≤L+2



+

 1 · νt,v (f 2 )1/2 . d

Again, we obtain an analogous bound as the above when Bv = 1 or Bv = u′ (Sv1 )u′ (Sv2 ), with the same arguments. Combining this bound with the bound (D.1) on |(I)| proves Lemma D.1. We will also use the following variant of Grönwall’s Inequality: Lemma D.2 (Lemma A.13.1 in [Tal10]). Suppose a differentiable function g : [0, 1] → R, g ≥ 0 is d such that dt g(t) ≤ c1 g(t) + c2 for some c1 , c2 ≥ 0. Then g(t) ≤ exp c1 (1 − t)



g(1) +

c2  . c1

Now we return to the proof of Lemma A.5. By the second part of Proposition A.6, n o νt,v Rℓ,ℓ − ρn,d ≥ K(D, β0 , β1 , κ, λ) ≤ K exp(−N ) , o n νt,v Rℓ1 ,ℓ2 − qn,d ≥ K(D, β0 , β1 , κ, λ) ≤ K exp(−N ) for ℓ1 ̸= ℓ2 . Here to bound ρn,d , qn,d we apply Proposition A.6 with t = v = 1. Additionally, the second pat of Proposition A.6 combined with Lemma A.7 implies that νt,v (Rℓ,ℓ − ρn,d )4

1/4

, νt,v (Rℓ1 ,ℓ2 − qn,d )4

1/4

≤ K(D, β0 , β1 , κ, λ) .

Now Hölder’s Inequality implies that νt,v |f | Rℓ,ℓ − ρn,d



 ≤ K(D, β0 , β1 , κ, λ)νt,v |f |1{|Rℓ,ℓ − ρn,d | ≤ K(D, β0 , β1 , κ, λ)} 1/2 1/4 4 1/4 + K(D, β0 , β1 , κ, λ)νt,v f 2 νt,v Rℓ,ℓ − ρn,d νt,v 1{|Rℓ,ℓ − ρn,d | ≤ K(D, β0 , β1 , κ, λ)}  1/2 ≤ K(D, β0 , β1 , κ, λ)νt,v |f | + K(D, β0 , β1 , κ, λ) exp(−N )νt,v f 2 .  The same argument as above establishes the same upper bound on νt,v |f | Rℓ1 ,ℓ2 − qn,d . 92

Now applying Lemma D.1 with Bv = 1, and combining with the above, we obtain that ∂ νt,v (f ) ≤ K(D, β0 , β1 , κ, λ, L) ∂v

X

νt,v |f | Rℓ,ℓ − ρn,d



ℓ≤L+1

X

+

νt,v |f | Rℓ1 ,ℓ2 − ρn,d



1≤ℓ1 ̸=ℓ2 ≤L+2

1 νt,v (f 2 )1/2 + ε′′ + d 

!

   1/2 1 ≤ K(D, β0 , β1 , κ, λ, L) νt,v |f | + ε′′ + + exp(−d) νt,v f 2 d

! .

For f ≥ 0 we thus have   1/2 ∂ 1 νt,v (f ) ≤ K(D, β0 , β1 , κ, λ, L) νt,v f + ε′′ + sup νt,v f 2 ∂v d t,v∈[0,1]2

! .

Recalling νt ≡ νt,1 , we conclude Lemma A.5 via Lemma D.2.

D.2

Proof of Lemma A.8

d νt (f ) using Gaussian Integration by Parts: First we explicitly compute the derivative dt

Lemma D.3. For a function f : (Rd−1 )L → R, we have d νt (f ) = (I)f + (II)f + (III)f + (IV)f + (V)f , dt where we define the quantities (I)f , (II)f , (III)f , (IV)f , (V)f as follows:   α X L+1 L+1 2 ℓ ℓ 2 νt (θdℓ )2 (u′′ (Sn,t ) + u′ (Sn,t ) )f − Lνt (θdL )2 (u′′ (Sn,t ) + u′ (Sn,t ) )f , (I)f := 2 ℓ≤L  X X   ℓ1 ℓ2 L+1 ℓ (II)f := α νt θdℓ1 θdℓ2 u′ (Sn,t )u′ (Sn,t )u′ (Sn,t )f )f − L νt θdℓ θdL+1 u′ (Sn,t 1≤ℓ1 <ℓ2 ≤L

ℓ≤L

 L(L + 1) L+1 ′ L+2 + νt θdL+1 θdL+2 u′ (Sn,t )u (Sn,t )f , 2  X X    L(L + 1) νt θdL+1 θdL+2 f , (III)f := −r νt θdℓ1 θdℓ2 f − L νt θdℓ θdL+1 f + 2 1≤ℓ1 <ℓ2 ≤n ℓ≤L   X   r (IV)f := − νt (θdℓ )2 f − Lνt (θdL+1 )2 f , 2 ℓ≤L   r − r̄  X νt (θdℓ )2 f − Lνt (θdL+1 )2 f . (V)f := 2 ℓ≤L

Proof of Lemma D.3. Note   ∂  X  ∂ d L+1 ℓ νt (f ) = νt f Ht,n − Lνt f Ht,n , dt ∂t ∂s ℓ≤L r X ∂ ℓ 1 r ℓ r − r̄ ℓ 2 ′ ℓ ℓ Ht,n = u (Si,t ) · √ gi,d θd − θd Y + θd . ∂t 1−t 2 2 td i≤n 93

The result now follows by symmetry w.r.t. i, 1 ≤ i ≤ n, and then applying Gaussian Integration by Parts w.r.t. gn,d , z ′′ . (The argument is analogous to the proof of Proposition 3.2.1, [Tal10].) Next collecting like terms from Lemma D.3, we obtain   d 1 X ℓ ℓ 2 νt (f ) ≤ ανt (θdℓ )2 (u′′ (Sn,t ) + u′ (Sn,t ) )f − r̄νt (θdℓ )2 f dt 2 ℓ≤L  X   ℓ1 ℓ2 )u′ (Sn,t )f − rνt θdℓ1 θdℓ2 f + ανt θdℓ1 θdℓ2 u′ (Sn,t 1≤ℓ1 <ℓ2 ≤L

−L

X

(D.6)

L+1 ℓ ανt θdℓ θdL+1 u′ (Sn,t )u′ (Sn,t )f



− rνt θdℓ θdL+1 f



ℓ≤L

  L(L + 1)  L+1 ′ L+2 ανt θdL+1 θdL+2 u′ (Sn,t )u (Sn,t )f − rνt θdL+1 θdL+2 f . 2 We now upper bound the right hand side in the above. This is where we use the particular choice of r, r̄ from (A.22), (A.23). Specifically, we will employ cavity in n over the parameter v to prove: 2 Lemma D.4. Let f = θd1 or θd1 θd2 . For any t ∈ [0, 1] and any ℓ ≤ L + 2, 1 ≤ ℓ1 ̸= ℓ2 ≤ L + 2,   ℓ ℓ 2 ανt (θdℓ )2 (u′′ (Sn,t ) + u′ (Sn,t ) )f − r̄νt (θdℓ )2 f   ℓ1 ℓ2 ∨ ανt θdℓ1 θdℓ2 u′ (Sn,t )u′ (Sn,t )f − rνt θdℓ1 θdℓ2 f !     1 1/2 1/2 , ≤ K νt (R1,2 − qn,d )2 + νt (R1,1 − ρn,d )2 + ε′′ + d +

where K depends on D, β0 , β1 , λ, L, and also on κ for the interpolators. ℓ = Sℓ Proof of Lemma D.4. Recall Sn,t n,t,1 . First, we show

  ℓ ℓ ανt,0 (θdℓ )2 (u′′ (Sn,t,0 ) + u′ (Sn,t,0 )2 )f = r̄νt,0 (θdℓ )2 f (D.7)   ℓ1 ℓ2 ανt,0 θdℓ1 θdℓ2 u′ (Sn,t,0 )u′ (Sn,t,0 )f = rνt,0 θdℓ1 θdℓ2 f . (D.8) √ √ ℓ We prove (D.7), the proof of (D.8) being analogous. Note Sn,t,0 = xsn + qn,d Z + ρn,d − qn,d W ℓ does not depend on θ, Y , gi,j , si for i ≤ n − 1 and that Ht,n−1,d does not depend on W , z, or sn . Therefore D X E ℓ ℓ ℓ ) + u′ (Sn,t,0 )2 )f exp (θdℓ )2 (u′′ (Sn,t,0 u(Sn,t,0 ) t,∼

ℓ≤L

= EW

h

 ℓ ℓ u′′ (Sn,t,0 ) + u′ (Sn,t,0 )2 exp

X

ℓ u(Sn,t,0 )

i

(θdℓ )2 f t,∼ .

ℓ≤L

By the same rationale, νt,0 (θdℓ )2 f ) = E (θdℓ )2 f t,∼ . Consequently, we have    ℓ ℓ νt,0 (θdℓ )2 u′′ (Sn,t,0 ) + u′ (Sn,t,0 )2 f " # ℓ ℓ ℓ EW [(u′′ (Sn,t,0 ) + u′ (Sn,t,0 )2 ) exp u(Sn,t,0 )] =E · (θdℓ )2 f t,∼ L+1 EW [exp u(Sn,t,0 )] " # ℓ ℓ ℓ h i EW [(u′′ (Sn,t,0 ) + u′ (Sn,t,0 )2 ) exp u(Sn,t,0 )] ℓ 2 = Ez,sn · E (θ ) f Y,g ,s :i≤n−1 d i,j i t,∼ ℓ EW [exp u(Sn,t,0 )]   r̄ = νt,0 (θdℓ )2 f . α 94

This proves (D.7). Again, (D.8) follows with an analogous argument. In light of (D.7), (D.8), recalling νt ≡ νt,1 , we next decompose   ℓ ℓ 2 ανt (θdℓ )2 (u′′ (Sn,t ) + u′ (Sn,t ) )f − r̄νt,1 (θdℓ )2 f   ℓ ℓ ℓ ℓ ≤ ανt,1 (θdℓ )2 (u′′ (Sn,t,1 ) + u′ (Sn,t,1 )2 )f − ανt,0 (θdℓ )2 (u′′ (Sn,t,0 ) + u′ (Sn,t,0 )2 )f   + r̄νt,0 (θdℓ )2 f − r̄νt,1 (θdℓ )2 f ,

(D.9)

and similarly we decompose   ℓ1 ℓ2 )u′ (Sn,t )f − rνt θdℓ1 θdℓ2 f ανt θdℓ1 θdℓ2 u′ (Sn,t   ℓ1 ℓ2 ℓ1 ℓ2 )u′ (Sn,t )f )u′ (Sn,t )f − ανt,0 θdℓ1 θdℓ2 u′ (Sn,t ≤ ανt,1 θdℓ1 θdℓ2 u′ (Sn,t   + rνt,0 θdℓ1 θdℓ2 f − rνt,1 θdℓ1 θdℓ2 f .

(D.10)

Note α ≤ α0 and r , r̄ ≤ α0 · K(D) as |u′ |, |u′′ | ≤ D. Now consider any v ∈ [0, 1]. We apply 2 Lemma D.1 with the function (θdℓ )2 f , recalling here that f = θd1 or θd1 θd2 . Applying CauchySchwarz on each of the terms arising from Lemma D.1, and using Proposition A.6 and Lemma A.7, we obtain that νt,v (f 2 )1/2 ≤ K(D, β0 , β1 , κ, λ). Therefore we get   ∂ ∂ ℓ ℓ ) + u′ (Sn,t,1 )2 )f , νt,v (θdℓ )2 (u′′ (Sn,t,1 νt,v (θdℓ )2 f ∂v ∂v !    1 2 1/2 2 1/2 ′′ ≤ K(D, β0 , β1 , κ, λ, L) νt,v (R1,2 − qn,d ) + νt,v (R1,1 − ρn,d ) + ε + , d where we use symmetry between replicas. Next applying Lemma A.5 with f = R1,1 − ρn,d and f = R1,2 − qn,d , and then applying the second part of Proposition A.6 with Lemma A.7, we obtain 2 1/2

νt,v (R1,2 − qn,d )



≤ K(D, β0 , β1 , κ, λ) νt (R1,2 − qn,d )

and analogously for νt,v (R1,1 − ρn,d )2

1/2

2 1/2



! 1 + ε + , d 

′′

. Combining the above two displays, we obtain

  ∂ ∂ ℓ ℓ ) + u′ (Sn,t,1 )2 )f , νt,v (θdℓ )2 (u′′ (Sn,t,1 νt,v (θdℓ )2 f ∂v ∂v !    1 2 1/2 2 1/2 ′′ ≤ K(D, β0 , β1 , κ, λ, L) νt (R1,2 − qn,d ) + νt (R1,1 − ρn,d ) + ε + . d Similarly, we have   ∂ ∂ ℓ1 ℓ2 νt,v θdℓ1 θdℓ2 u′ (Sn,t )u′ (Sn,t )f , νt,v θdℓ1 θdℓ2 f ∂v ∂v !     1 1/2 1/2 ≤ K(D, β0 , β1 , κ, λ, L) νt (R1,2 − qn,d )2 + νt (R1,1 − ρn,d )2 + ε′′ + . d Combining the above with (D.9), (D.10), and noting α ≤ α0 ≤ 2, the Lemma follows. 95

Combining Lemma D.4 and (D.6), it follows that for f = θd1

2

or θd1 θd2 , we have

!    d 1 2 1/2 2 1/2 ′′ . νt (f ) ≤ K(D, β0 , β1 , κ, λ, L) νt (R1,2 − qn,d ) + νt (R1,1 − ρn,d ) + ε + dt d By Proposition A.6, we have 1/2 1/2 νt (R1,1 − ρn,d )2 ≤ Kνt (R1,1 − νt (R1,1 ))2 + 2 νt (R1,1 ) − ν1 (R1,1 ) K(D, β0 , β1 , κ, λ) √ ≤ + 2 νt (R1,1 ) − ν1 (R1,1 ) , d and similarly for R1,2 . Here we use that ρn,d = ν1 (R1,1 ), qn,d = ν1 (R1,2 ). Thus combining the above two displays, to complete the proof of Lemma A.8, it suffices to upper bound d K(D, β0 , β1 , κ, λ) d √ . νt (R1,1 ) , νt (R1,2 ) ≤ dt dt d To this end, considering f = R1,2 and following the notation of Lemma D.3, we note that (I)f = (I)f − νt (R1,2 ) = (I)f −νt (R1,2 ) , because νt (R1,2 ) is just a constant (note that for any constant function c, (I)c = 0). Similarly (II)f = (II)f −νt (R1,2 ) , (III)f = (III)f −νt (R1,2 ) , (IV)f = (IV)f −νt (R1,2 ) , (V)f = (V)f −νt (R1,2 ) . However note that α ≤ α0 and r, r̄ ≤ K(D, α0 ). Thus by Proposition A.6, since |u′ |, |u′′ | ≤ D, α ≤ α0 ≤ 2, and by Hölder’s Inequality used on each individual summand comprising the following expressions, we obtain (I)f −νt (R1,2 ) , (II)f −νt (R1,2 ) , (III)f −νt (R1,2 ) , (IV)f −νt (R1,2 ) , (V)f −νt (R1,2 ) ≤

K(D, β0 , β1 , κ, λ) √ . d

The same derivation applies for f = R1,1 , thus completing the proof of Lemma A.8.

D.3

Proof of Proposition A.6

Here, we prove Proposition A.6. For only Appendix D.3, we define the following Gibbs average ⟨·⟩:  P ℓ f exp 1≤ℓ≤L u(Sn,t,v ) t,∼ ⟨f ⟩ := . L+1 L exp u(Sn,t,v ) t,∼ Consequently by (A.26), we have νt,v (f ) = E⟨f ⟩. Note we can rewrite ⟨f ⟩ as " # Z   Y √  1 l EW f 1{θ1l / d ∈ I} exp u Sn,t,v (θl ) + Ht,n−1,d (θl ) − β∥θl ∥2 dθ1 · · · dθL , Z ′L 1≤l≤L

(D.11) where ′

Z

Z = Rd

 i h √ EW 1{θ1 / d ∈ I} exp u(Sn,t,v ) + Ht,n−1,d (θ) − β∥θ∥2 dθ .

(D.12)

√ Q Note the constraint 1≤l≤L 1{θ1l / d ∈ I} is convex, u is concave, the Sn,t,v and Si,t are affine in θ, and the rest of Ht,n−1,d is log-concave or affine in θ. It follows that the measure ⟨·⟩ is β-strongly log-concave in θ, a fact that will be exploited repeatedly via concentration of Lipschitz functions of strongly log-concave measures (Theorem B.3) in what follows. 96

D.3.1

First part of Proposition A.6

We prove the first part of Proposition A.6 by directly combining the following Propositions D.5, D.6. First, we have the following concentration property of f = R1,1 or R1,2 about their average ⟨f ⟩ with the quenched disorder fixed: Proposition D.5. Consider u(·) as in Theorem A.1 for the interpolators or Theorem A.2 for the posterior. Then for d ≥ dD.5 (ε′′ ) and all k ≤ d, we have D

2k E D 2k E  kKB ⋆ k , R1,1 − ⟨R1,1 ⟩ , R1,2 − ⟨R1,2 ⟩ ≤ d2

where K depends on D, β0 , β1 , λ and also on κ for the interpolators, and where B ⋆ :=

n X

∥ḡi ∥2

1/2

i=1

+

d

n−1

j=2

i=1

o 1 X n X 2 2 gij + gnj 16d

n  X 1/2 + n s2i + i=1

(D.13)

n−1

n X 2 o 1 si + s2n + Y 2 + Z 2 + d . 8dKsignal i=1

Second, we have concentration of ⟨f ⟩ about E⟨f ⟩: Proposition D.6. Consider u(·) as in Theorem A.1 for the interpolators or Theorem A.2 for the posterior. Then for k ≤ d4 , we have h E

R1,1 − E R1,1

2k i

, E

h

R1,2 − E R1,2

2k i

 Kk k d

,

where K depends on D, β0 , β1 , λ, and also on κ for the posterior. Combining Proposition D.5 and Proposition D.6 yields the first part of Proposition A.6. For the rest of the proof of the first part of Proposition A.6, we assume u(·) satisfies the conditions of Theorem A.1. Central in the proofs of both Proposition D.5 and Proposition D.6 is concentration of Lipschitz functions w.r.t strongly log-concave measures, as written in Theorem B.3. We will also use the following fact which follows by a standard symmetrization argument, see (3.22) of [Tal10]. Lemma D.7. For any m ∈ R, any function f such that the following is defined, and any probability measure µ, we have Z  Z Z 2k dµ ≤ 22k (f − m)2k dµ . f − f dµ Proof of Proposition D.5: We now prove Proposition D.5. First, we show that B ⋆ can be controlled on the exponential scale: Lemma D.8. For all k ≤ d, we have   k E B ⋆k ≤ Kd , where K depends on λ. 97

  Proof. Our strategy will be to upper bound E exp B ⋆ and then convert this to a bound on the k-th moment of B ⋆ by Lemma A.7. Since n ≤ 2d, and by applying Cauchy-Schwarz on the (si ), n

d

n−1

n

j=2

i=1

i=1

o  X  1 1 X 1 X n X 2 2 + B ≤ ∥ḡi ∥2 + 2n + gij + gnj s2i + Ksignal n 8n 16d 2Ksignal ⋆

i=1

+

Y2 4

+

Z2 4

+ Kd

= (I) + (II) + (III) +

Y 2 Z2 + + (Ksignal + 2)n + Kd , 4 4

where we define n

d

n−1

n

i=1

j=2

i=1

i=1

o  X  1 X 1 1 X n X 2 2 (I) := , (III) := ∥ḡi ∥2 , (II) := gij + gnj s2i . 8n 16d 2Ksignal By independence of gij , si , Y, Z, and as n ≤ 2d, we have   E exp B ⋆        ≤ E exp (I) + (II) E exp (III) E exp Y 2 /4] E exp Z 2 /4] exp K(λ)d        ≤ E exp 2(I)] E exp 2(II)] E exp (III) E exp Y 2 /4] E exp Z 2 /4] exp K(λ)d , (D.14)     Z2 Y2 where the last step used Cauchy-Schwarz. E exp   Note   4 ,E  exp 4  ≤ K as Y, Z ∼ N (0, 1). It remains to upper bound E exp 2(I) , E exp 2(II) , and E exp (III) .   • Upper bounding E exp 2(I) : By a direct calculation, we have the following, as stated as (A.11) in [Tal10]: h E

2 i  gij 1 −1/2 exp ≤ 1− . 4n 2n

Thus the independence of the (gij ) implies    1 −nd/2 E exp 2(I) ≤ 1 − ≤ K d ≤ exp(Kd) . 2n   P • Upper bounding E exp 2(II) : By independence of the (gij ), we have n−1 i=1 gij ∼ N (0, n − 1) for all 2 ≤ j ≤ d, and thus n−1

1 2 1 1 n n 2 1  X 2 gij + gnj ∼ N (0, n − 1)2 + N (0, 1)2 ∼ N (0, 1)2 =⇒ 2(II) ∼ χ , 8d 8d 8d 8d 8d 8d d−1 i=1

where χ2d−1 denotes a χ2 distribution with d − 1 degrees of freedom. Since n ≤ 2d, standard bounds on the MGF of a χ2 random variable now gives    d−1 n − d−1 2 E exp 2(II) = 1 − 2 · ≤ 2 2 ≤ exp(Kd) . 8d   • Upper bounding E exp (III) : Since the s2i are  Ksignal sub-Exponential, by e.g. Proposition 2 2.7.1 of [Ver18]), we have E exp si /2Ksignal ≤ K. As the (si ) are independent, we obtain   E exp (III) ≤ K n ≤ exp(Kd) . 98

Combining the bounds in the above three cases together with (D.14) and recalling Ksignal depends only on λ implies    E exp B ⋆ ≤ exp K(λ)d ,

(D.15)

and the conclusion now follows from Lemma A.7. We next lower bound the partition function Z ′ in order to apply Lemma B.19 to control the relevant overlaps. Lemma D.9. For d ≥ dD.9 (ε′′ ), letting Z ′ be as per (D.12), we have  Z ′ ≥ exp − KD.9 B ⋆ , where KD.9 depends on D, β0 , β1 , λ, and also on κ for the posterior. Proof. We will prove this Lemma in two stages. First we will prove it when v = t = 1, which lets us upper bound |qn,d |. Then we use this upper bound on |qn,d | to complete the proof for general t, v ∈ [0, 1]. Note the conditions on u from Theorem A.1 give u(x) ≥ −D(κ)(|x| + 1) for the interpolators; similarly, we have u(x) ≥ −D(λ)(|x| + 1) for the posterior. In an abuse of notation, we will write this bound as u(x) ≥ −D(κ, λ)(|x| + 1), where we emphasize that there is no dependence on κ for the posterior. We begin with steps of the proof that are common to both stages. Rewrite ′

Z

Z

h

Z = √1 θ1 ∈I d

Rd−1



X

EW exp u(Sn,t,v ) +

u(Si,t )

1≤i≤n−1

+ θd

i p (1 − t)(r − r̄) 2 (1 − t)rY − θd − β∥θ̄∥2 − βθ12 dθdθ1 . 2

Notice r̄ ≤ r ≤ K(D) as per their definitions in (A.20), (A.21), thus (1−t)(r−r̄) ≤ r−r̄ 2 2 ≤ KD.9 (D). d−1

Letting β ′ := KD.9 (D) + β, we consider the density γ on Rd−1 with density (β ′ /π) 2 exp(−β ′ ∥θ∥2 ). We observe that Z Rd−1

h  EW exp u(Sn,t,v ) +

≥ exp(−βθ12 ) ≥ exp(−βθ12 )

X

u(Si,t ) + θd

1≤i≤n−1

 π  d−1 Z 2

β′

Rd−1

 π  d−1 2

β′

i p (1 − t)(r − r̄) 2 (1 − t)rY − θd − β∥θ̄∥2 − βθ12 dθ 2

h  EW exp u(Sn,t,v ) +

X

u(Si,t ) + θd

i p (1 − t)rY dγ θ̄)

1≤i≤n−1

Z exp Rd−1

h

EW u(Sn,t,v ) +

! i p u(Si,t ) + θd (1 − t)rY dγ θ̄) .

X 1≤i≤n−1

Next by rotational invariance of γ, we have the following bound, as stated in (3.24) in [Tal10]: r

Z ⟨x, θ⟩ dγ θ̄) =

1 1 ∥x∥ ≤ √ ′ ∥x∥ ′ πβ β 99

for all

x ∈ Rd−1 .

(D.16)

  √ p √ Also define g¯i ′ := gi,2 , . . . , gi,d−1 , tgi,d + nd · (1 − t)rY . Using (D.16), it follows that Z h i X p u(Si,t ) + θd (1 − t)rY dγ θ̄) EW u(Sn,t,v ) + Rd−1

1≤i≤n−1

"

r

 √  p t √ gn,d θd + 1 − v qn,d Z + ρn,d − qn,d W d Rd−1 r X X t si θ1 1 √ +√ + gi,j θj + gi,d θd d d d 2≤j≤d−1 1≤i≤n−1 # ! p sn θ1 + v · √ + (1 − v)sn x + n + θd (1 − t)rY dγ θ̄) d r Z Z X 1 v ′ ≥ −D(κ, λ) n + √ ⟨ḡi , θ̄ ⟩dγ(θ̄) + ⟨ḡn′ , θ̄⟩ dγ(θ̄) d d 1≤i≤n−1 ! q q |θ1 | X +√ |si | + |Z| (1 − v)qn,d + (1 − v)|sn | + K (1 − v)(ρn,d − qn,d ) d 1≤i≤n ! q X p 1 |θ1 | X ′ |si | + |Z| (1 − v)qn,d + |sn | + K ρn,d − qn,d . ≥ −D(κ, λ) n + √ ′ ∥ḡ ∥ + √ dβ 1≤i≤n i d 1≤i≤n   Here, in the second inequality we took expectation with respect to W and used the bound EW |W | ≤ K, and in the third inequality we used (D.16). 1/2 P P . Also by Triangle Inequality and as |r| ≤ Next, remark that 1≤i≤n si ≤ n 1≤i≤n s2i K(D), we have p 1/2 d(1 − t)r   n X 1 X 1 X  ′ √ ∥g¯i ∥ ≤ √ ∥ḡi ∥ + |Y | ≤ ∥ḡi ∥2 + K(D)|Y | . n d d 1≤i≤n d 1≤i≤n 1≤i≤n Z

≥ −D(κ, λ)

EW

X √  1 v √ gn,j θj + d 2≤j≤d−1

Now specializing to when t = v = 1, combining the above displays, recalling the definition β ′ = KD.9 (D) + β, and using that √θ1d ∈ I ⊆ [−1, 1], n ≤ α0 d ≤ 2d and the definition of B ⋆ from (D.13), we obtain Z h  i X EW exp u(Sn,1,1 ) + u(Si,1 )β∥θ̄∥2 − βθ12 dθ Rd−1

1≤i≤n−1

 π  d−1 2 ≥ exp(−βθ12 ) ′ exp β

  X 1 |θ1 | X − D(κ, λ) n + √ ′ ∥ḡi′ ∥ + √ |si | + |sn | dβ 1≤i≤n d 1≤i≤n !   X 1/2  X 1/2  − K(D, β0 , β1 , κ) d + ∥ḡi ∥2 + d s2i + |Y |

≥ exp

1≤i≤n



!

1≤i≤n



≥ exp − K(D, β0 , β1 , κ)B ⋆ . √ Using that | dI| ≥ 2ε′′ , that ε′′ ≥ exp(−d) if d ≥ dD.9 (ε′′ ), and that B ⋆ ≥ d, we thus obtain when t = v = 1, Z     ′ Z ≥ exp − K(D, β0 , β1 , κ)B ⋆ dθ1 ≥ exp − K(D, β0 , β1 , κ)B ⋆ . √1 θ1 ∈I d

100

The condition of Lemma B.19 thus applies, and combining it with Lemma D.8 yields that i 1 h E ∥θ̄∥2 d i 1 h ≤ E K(D, β0 , β1 , κ, λ)B ⋆ d 1 ≤ · K(D, β0 , β1 , κ, λ)d = K(D, β0 , β1 , κ, λ) . d

0 ≤ |qn,d | ≤ ρn,d =

(D.17)

We nowPreturn to the proof t, v ∈ [0, 1]2 . Again by the aforementioned upper P for general 1 ′ bounds on 1≤i≤n |si | and √d 1≤i≤n ∥ḡi ∥, recalling the definition β ′ = KD.9 (D) + β, and using that √θ1d ∈ I ⊆ [−1, 1], n ≤ α0 d ≤ 2d and the definition of B ⋆ , we have Z Rd−1

h  EW exp u(Sn,t,v ) +

X

u(Si,t )β∥θ̄∥2 − βθ12

i

1≤i≤n−1

 π  d−1 2 ≥ exp(−βθ12 ) ′ exp β

 X 1 |θ1 | X − D(κ, λ) n + √ ′ ∥ḡi′ ∥ + √ |si | dβ 1≤i≤n d 1≤i≤n q  p + |Z| (1 − v)qn,d + K ρn,d − qn,d + |sn |

≥ exp



− K(D, β0 , β1 , κ, λ) d +

 X

∥ḡi ∥2

1/2



+ d

X

s2i

1/2

+ |Y | + |Z|



!

!

1≤i≤n

1≤i≤n

  ≥ exp − K(D, β0 , β1 , κ, λ)B ⋆ . √ Here, we used (D.17) to upper bound |ρn,d |, |qn,d | and hence K ρn,d − qn,d . √ Using that | dI| ≥ 2ε′′ , that ε′′ ≥ exp(−d) if d ≥ dD.9 (ε′′ ), and that B ⋆ ≥ d, we thus obtain for general t, v ∈ [0, 1]2 that ′

Z

Z ≥ √1 θ1 ∈I d

    exp − K(D, β0 , β1 , κ, λ)B ⋆ dθ1 ≥ exp − K(D, β0 , β1 , κ, λ)B ⋆ .

This proves Lemma D.9. We now have the ingredients needed to prove Proposition D.5. Proof of Proposition D.5. First, we claim that k 1/2 ∥θ̄∥2k ≤ ∥θ̄∥4k ≤ K(D, β0 , β1 , κ, λ) B ⋆ , D β E  ∥θ̄∥2 exp ≤ exp KB.19 (KD.9 , D, β0 , β1 , κ, λ) B ⋆ . 2

(D.18) (D.19)

Note that the functions we take averages with respect to in the Gibbs measure ⟨·⟩ above are a function of L = 1 replica. Next, note we have Z ′ ≥ exp(−KD.9 B ⋆ ) by Lemma D.9. For the interpolators or for the posterior in the logistic case where u ≤ 0, (D.18), (D.19) follow immediately by Lemma B.19, as the corresponding U (·) from (B.84) is such that U ≤ 0 and all the aj = 0. For 101

√ the posterior in the GMM case, since u(x) = λx, we note Lemma B.18 applies where √  p √ a0 = 1 − v qn,d Z + ρn,d − qn,d W + sn · (1 − v)x , r  n−1  λ X a1 = si + sn v , d i=1 r  n−1  √ λ X gij + vgnj for all 2 ≤ j ≤ d − 1 , aj = d i=1 r  n−1  p √ λt X ad = gij + vgnj + (1 − t)rY , d i=1

U (θ) = −

(1 − t)(r − r̄) 2 θd ≤ 0 . 2

By (D.17), using that EW [exp KW ] ≤ K we have    log EW exp a0 ≤ K(D, β0 , β1 , κ, λ) |Z| + |sn | + K(D, β0 , β1 , κ, λ) . It follows that   1 X 2 aj + log EW exp a0 ≤ K(D, β0 , β1 , κ, λ)B ⋆ . 2β0 1≤j≤d

Combining the above display with Lemma B.19 and Lemma D.9 now establishes (D.18), (D.19). 2 We first prove the desired result for R1,2 . Let m = θ̄ , thus ⟨R1,2 ⟩ = ∥m∥ d , and R1,2 − ⟨R1,2 ⟩ ≤

⟨θ̄1 , θ̄2 ⟩ ⟨θ¯1 , m⟩ ⟨θ̄1 , m⟩ ⟨m, m⟩ − + − . d d N N

1

1

For fixed θ1 , the map f (x) = ⟨θ̄ d,x⟩ is Lipschitz with constant ∥θ̄d ∥ . Letting µI (θl ) denote the Gibbs measure w.r.t. only the l-th replica, which is β strongly log-concave, Theorem B.3 gives Z  1 2  k K(β) ∥θ̄1 ∥2 k ⟨θ̄ , θ̄ ⟩ ⟨θ̄1 , m⟩ 2k − dµI (θ2 ) ≤ . d d d2 We now integrate this inequality for θ1 with respect to µI (θ1 ). By (D.18), D ⟨θ̄1 , θ̄2 ⟩ ⟨θ̄1 , m⟩ 2k E Z  ⟨θ̄1 , θ̄2 ⟩ ⟨θ̄1 , m⟩ 2k − − dµI (θ1 )dµI (θ2 ) = d d d d  k K(β) k ≤ ∥θ̄1 ∥2k d2  k K(D, β , β , κ, λ) B ⋆ k 0 1 ≤ . d2 Likewise, the map f (x) = ⟨x,m⟩ is Lipschitz with constant ∥m∥ d d . Thus Theorem B.3 gives D ⟨θ̄1 , m⟩ ⟨m, m⟩ 2k E Z  ⟨θ̄1 , m⟩ ⟨m, m⟩ 2k  k K(β) ∥m∥2 k − = − dµI (θ1 ) ≤ . d d d d d2 Again by Lemma B.19 and Jensen’s Inequality repeatedly, this gives ∥m∥2k = ⟨θ̄⟩

2k

≤ ∥θ̄∥

2k

≤ ∥θ̄∥2k ≤ K(D, β0 , β1 , κ, λ) B ⋆ 102

k

.

(D.20)

Combining with the above display yields D ⟨θ̄1 , m⟩ d

⟨m, m⟩ 2k E  k K(D, β0 , β1 , κ, λ) B ⋆ k . ≤ d d2

(D.21)

As (a + b)2k ≤ 22k−1 (a2k + b2k ), combining (D.20) and (D.21), we obtain the desired upper bound for (R1,2 − ⟨R1,2 ⟩)2k . We now prove the result for R1,1 . The idea is to truncate R1,1 into a part f where we have the requisite bounds, and another part ϕ which is 0 with high probability. Specifically, define the parameter √ KB.19 (KD.9 , D, β0 , β1 , κ, λ) 6B ⋆ √ >0 a := β0 controlling the truncation, and in terms of a define n ∥θ̄∥2 a2 o f (θ) = min , = d d

min{∥θ̄∥, a} d

2 .

is convex, it follows that f is Lipschitz with constant 2a d . Thus

 As the cylinder θ : ∥θ̄∥ ≤ a Theorem B.3 gives

D

2k E  k K(β) a2 k f − ⟨f ⟩ ≤ . d2

(D.22)

Define ϕ(θ) :=

∥θ̄∥2 − f (θ) ≥ 0 . d

2

Thus ϕ(θ) ≤ ∥θ̄∥ d 1{∥θ̄∥ ≥ a}. By Lemma D.7 and Cauchy–Schwarz, D

ϕ − ⟨ϕ⟩

2k E

≤ 22k ϕ2k D ∥θ̄∥2 2k E ≤ 22k 1{∥θ̄∥ ≥ a} d D ∥θ̄∥2 4k E1/2 D E1/2 1{∥θ̄∥ ≥ a}2k ≤ 22k d D ∥θ̄∥2 4k E1/2 D E1/2 = 22k 1{∥θ̄∥ ≥ a} . d

(D.23)

By (D.18), (D.19), Markov’s Inequality and the choice of a gives D E  β   βa2  1{∥θ̄∥ ≥ a} = P exp ∥θ̄∥2 ≥ exp 2 2   βa2  ≤ exp KB.19 (KD.9 , D, β0 , β1 , κ, λ) B ⋆ − ≤ exp − 2B ⋆ . 2 Plugging in (D.18), (D.24) into (D.23), we obtain D

ϕ − ⟨ϕ⟩

2k E

≤ exp − B ⋆

 K(D, β0 , β1 , κ, λ) B ⋆ 2k . d

103

(D.24)

Since R1,1 = ∥θ̄∥2 /d = f + ϕ, using that (a + b)2k ≤ 22k (a2k + b2k ), combining with (D.22) and recalling our choice of a, we obtain D

R1,1 − R1,1

2k E

≤ 22k ≤

D

! 2k E D 2k E f − ⟨f ⟩ + ϕ − ⟨ϕ⟩

 k K(D, β , β , κ, λ) B ⋆ k 0

1

d2

+ exp − B ⋆

 K(D, β0 , β1 , κ, λ) B ⋆ 2k . d

Note for x ≥ 0 that xk ≤ k k exp(x), thus exp(−x) ≤ (k/x)k for x > 0. Clearly B ⋆ > 0. Hence exp − B ⋆

 K(D, β0 , β1 , κ, λ) B ⋆ 2k  k k  K(D, β0 , β1 , κ, λ) B ⋆ 2k ≤ d B⋆ d  k K(D, β , β , κ, λ) B ⋆ k 0 1 . ≤ d2

Combining the above two displays gives the desired upper bound on (R1,1 − R1,1 )2k . Proof of Proposition D.6 We will follow the strategy of [Tal10] of considering ⟨R1,1 ⟩ and ⟨R1,2 ⟩ as functions of the disorder (sj )j≤n × (gi,j )2≤i≤n,1≤j≤d × (W l )1≤l≤L × Z × Y . The density of the disorder is strongly log-concave by Assumption 1, and to apply Theorem B.3 on concentration of Lipschitz functions of strongly log-concave measures, we need to establish control on the Lipschitz constant of these functions. Lemma D.10. Let f = R1,1 or R1,2 and let ∇ denote the derivative of the function ⟨f ⟩ w.r.t. the disorder. Let     T ∈ Rn , u(θ) := u′ S1,t (θ) , . . . , u′ Sn−1,t (θ) , u′ Sn,t,v (θ) and analogously define ul (θl ) for each replica l. Then ∥∇∥2 is upper bounded by

K

1 2 f˙θ˙d1 + d

D

 E 2 1 X 1 θ1 − ⟨θ1 ⟩ f˙u′ Sn,t,v (θ1 ) + d

D

 E 2 1 θ1 − ⟨θ1 ⟩ f˙u′ Si,t (θ1 )

i≤n−1

1 + 1 + ∥θ1 ∥2 d 

 D

 E 2 u1 (θ1 ) − u1 (θ1 ) f˙

! ,

where K depends on D, β0 , β1 , λ and also on κ for the interpolators, and where we use the notation f˙ := f − ⟨f ⟩ , θ˙l := θl − ⟨θl ⟩ , θ˙jl := θjl − ⟨θjl ⟩ , and u̇′ (S l ) := u′ (S l ) − ⟨u′ (S l )⟩

for

l l S l = Si,t (θl ) or Sn,t,v (θl ) .

Proof. First, consider x an element of the quenched disorder, i.e. not one of the W l . For just this 104

 l proof, we let Ht,v,n,d (θl ) := u Sn,t,v (θl ) + Ht,n−1,d (θl ) − β∥θl ∥2 . By (D.11), we remark that ∂ 1 ⟨f ⟩ = ′L ∂x Z

"

Z Rd

EW f

X n √  ∂ 1{θ1l / d ∈ I} exp Ht,v,n,d (θl ) · Ht,v,n,d (θl ) ∂x 1≤l≤L # Y √ o l′ l′ dθ1 · · · dθL · 1{θ1 / d ∈ I} exp Ht,v,n,d (θ ) 1≤l′ ̸=l≤L

L Z ′L+1

"

Z Rd

EW f

Y

1{θ1l /

#  d ∈ I} exp Ht,v,n,d (θ ) dθ1 · · · dθL l

1≤l≤L

h i √  ∂ EW 1{θ1L+1 / d ∈ I} exp Ht,v,n,d (θL+1 ) · · Ht,v,n,d (θL+1 ) dθL+1 ∂x Rd E D∂ E X D ∂ = f Ht,v,n,d (θl ) − ⟨f ⟩ Ht,v,n,d (θ) . ∂x ∂x Z

1≤l≤L

Note that the θ1 , . . . , θL such that there exists some l′ with √1d θl ̸∈ I make no contribution to ⟨f ⟩. √ Q Thus ⟨f ⟩ = ⟨f ′ ⟩ where f ′ is the restriction of f to 1≤l≤L 1{θ1l / d ∈ I}. Consequently we have D  θ1 E ′ S (θ) · √  u  i,t  D E   d vθ   ′ 1  √ u Sn,t,v (θ) · + (1 − v)x   D  θj E d   ′  √  Du Si,t (θ) · dq E     u′ Si,t (θ) · θd t D∂ E  d Ht,v,n,d (θ) = D  pvE ′  ∂x u Sn,t,v (θ) · θj d    q E D    ′   u Sn,t,v (θ) · θd vt  d  D E    u′ Sn,t,v (θ) · p(1 − v)qn,d    D p E    θd (1 − t)r

: x = si , i ≤ n − 1 : x = sn : x = gi,j , i ≤ n − 1, j ≤ d − 1 : x = gi,d , i ≤ n − 1 : x = gn,j , j ≤ d − 1 : x = gn,d :x=Z :x=Y .

D E ∂ An analogous derivation applies also for f ∂x Ht,v,n,d (θl ) . Now consider x = W l for some 1 ≤ l ≤ L. We remark that h √ i ∂ l 1{θ / d ∈ I} exp H (θ) E l t,v,n,d 1 ∂W l W h √ i = − EW l 1{θ1l / d ∈ I} exp Ht,v,n,d (θ) h √  i  ∂ l l l + EW l 1{θ1l / d ∈ I} exp Ht,v,n,d (θ) · u′ Sn,t,v u S (θ ) (θl ) n,t,v ∂W l h i √  = − EW l 1{θ1l / d ∈ I} exp Ht,v,n,d (θ) h i √  q l + EW l 1{θ1l / d ∈ I} exp Ht,v,n,d (θ) · u′ Sn,t,v (θl ) (1 − v)(ρn,d − qn,d ) , √ where the last step follows from the precense of the constraint 1{θ1l / d ∈ I} and analogous logic as above. Since f = R1,1 or R1,2 does not depend on the W l , it follows from an identical derivation 105

as above that q D D X E E ∂ ′ l l ′ l l ⟨f ⟩ = −⟨f ⟩ + ⟨f ⟩ + (1 − v)(ρ − q ) f u S (θ ) − ⟨f ⟩ u S (θ ) n,d n,d n,t,v n,t,v ∂W l 1≤l≤L D D X q E E l l (1 − v)(ρn,d − qn,d ) f u′ Sn,t,v (θl ) − ⟨f ⟩ u′ Sn,t,v (θl ) . = 1≤l≤L

Note ⟨f˙⟩ = 0. Thus for any 1 ≤ l ≤ L, f˙u′ (S l ) = f˙u̇′ (S l ) , f˙θdl = f˙θ˙dl , and for any 1 ≤ l ≤ L and any 1 ≤ j ≤ d, f˙θjl u′ (S l ) = f˙θ˙jl u′ (S l ) + θjl

f˙u′ (S l ) = f˙θ˙jl u′ (S l ) + θjl

f˙u̇′ (S l ) .

Since t, v ∈ [0, 1]2 , L ≤ 2 for f = R1,1 , R1,2 , and as f = R1,1 , R1,2 are symmetric in θ1 , θ2 , combining all of the displays above implies that   ∥∇∥2 ≤ K (I) + (II) + (III) where X  2 1  2 1 X ˙ ˙1 ′ 1 f θj u Sn,t,v (θ1 ) + f˙θ˙j1 u′ Si,t (θ1 ) , d d 1≤j≤d i≤n−1,1≤j≤d  1 X X  2  2 1 ⟨θj1 ⟩2 f˙u̇′ Sn,t,v (θ1 ) + f˙u̇′ Si,t (θ1 ) , (II) := d i≤n−1 1≤j≤d     2 2 1 (III) := max x2 , qn,d , r, ρn,d − qn,d f˙u̇′ Sn,t,v (θ1 ) + f˙θ˙d1 . (I) :=

Note |x| ≤ 1 and that r ≤ K(D). Moreover by (D.17), we have |qn,d |, |ρn,d | ≤ K(D, β0 , β1 , κ, λ). Also remark that 1 1 X 1 2 1 X ⟨θj ⟩ ≤ ⟨(θj1 )2 ⟩ = ∥θ1 ∥2 . d d d 1≤j≤d

1≤j≤d

1 Next, for S 1 either Sn,t,v (θ1 ) or Si,t (θ1 ), we observe that X D E2 f˙θ˙j1 u′ S 1 (θ1 ) 1≤j≤d

=

X D

E   ˙ f˙(θ1 , . . . , θL )f˙(θL+1 , . . . , θ2L )u′ S 1 (θ1 ) u′ S 1 (θL+1 ) θ˙j1 θjL+1

1≤j≤d

D   T L+1 E = f˙(θ1 , . . . , θL )f˙(θL+1 , . . . , θ2L )u′ S 1 (θ1 ) u′ S 1 (θL+1 ) · θ1 − ⟨θ1 ⟩ θ − ⟨θL+1 ⟩ D E 2  = θ1 − ⟨θ1 ⟩ f˙u′ (S 1 ) . Analogously, we have X  2 1 f˙u̇′ Sn,t,v (θ1 ) +

 2 f˙u̇′ Si,t (θ1 )

i≤n−1

   T    E = f˙(θ1 , . . . , θL )f˙(θL+1 , . . . , θ2L ) · u S 1 (θ1 ) − u S 1 (θ1 ) u S 1 (θL+1 ) − u S 1 (θL+1 ) D    E 2 = u S 1 (θ1 ) − u S 1 (θ1 ) f˙ . D

106

Using these observations, the Lemma follows from combining terms in the above expressions (I), (II), (III). We now aim to upper bound the expressions appearing in the bound from Lemma D.10. Lemma D.11. For any function f of θ1 , . . . , θL mapping to R (possibly depending on the disorder), we have D  E 2 ≤ K⟨f 2 ⟩ , θ1 − ⟨θ1 ⟩ f and consequently we also have f θ˙d1

2

≤ K⟨f 2 ⟩, where K depends on β0 , β1 .

Proof. Consider any y ∈ Rd . By Cauchy-Schwarz, 

 D D T E T 2 E1/2 1/2 T θ1 − ⟨θ1 ⟩ f y = f θ1 − ⟨θ1 ⟩ y ≤ f 2 θ1 − ⟨θ1 ⟩ y .

By Theorem B.3 applied with the function (θ1 )T y of θ1 , . . . , θL which is ∥y∥-Lipschitz, D

 T 2 E θ1 − ⟨θ1 ⟩ y ≤ K(β0 , β1 )∥y∥2 .

Hence, we have D

 ET 1/2 θ1 − ⟨θ1 ⟩ f y ≤ K(β0 , β1 ) f 2 ∥y∥ ,

and since y is arbitrary the first part of the Lemma follows. The second part of this Lemma follows upon observing that analogously as in the proof of Lemma D.10, we have f θ˙d1

2

X

f θ˙j1

2

=

D

 E 2 θ1 − ⟨θ1 ⟩ f .

1≤j≤d

Applying the first part of this Lemma then finishes the proof. To bound the other term in Lemma D.10, define the matrix M as follows: for all 1 ≤ i ≤ n and 1 ≤ j ≤ d, we let   si : 1 ≤ i ≤ n − 1, j = 1      gi,j : 1 ≤ i ≤ n − 1, 2 ≤ j ≤ d − 1    √tg : 1 ≤ i ≤ n − 1, j = d i,d (M )i,j := (D.25)  vsn : i = n, j = 1    √   vgn,j : i = n, 2 ≤ j ≤ d − 1    √vtg n,d : i = n, j = d . B ′ := ∥M ∥op .

(D.26)

We now control B ′ as follows: Lemma D.12. We have B ′2 ≤ Kd with probability at least 1 − K exp(−d), where K depends on λ. 107

Proof. Note B ′2 = ∥M M T ∥op . Consider any unit vector y ∈ Rn . Note for all 1 ≤ j ≤ d that Pn−1  :j=1 Pi=1 si yi + vs√n yn  T n−1 (D.27) M y j= :2≤j ≤d−1 i=1 gi,j yi + vgn,j yn  √ √ Pn−1 t i=1 gi,d yi + tvgn,d yn : j = d . Let M̄ be the n × d − 1 matrix comprising of the last d − 1 columns of M ; note this matrix does not involve the si , which we handle in a separate step. Observe that d−1  n−1 X X j=2

gi,j yi +

vgn,j yn

2

+

√ n−1 2 X √ t gi,d yi + tvgn,d yn = ∥M̄ T y∥2 ≤ ∥M̄ ∥2op .

i=1

(D.28)

i=1

By concentration of operator norm for matrices with √ i.i.d. sub-Gaussian entries, e.g. Theorem 4.4.5 of [Ver18], as t, v ∈ [0, 1]2 we have ∥M̄ ∥op ≤ K d with probability at least 1 − K exp(−d). Also, as v ∈ [0, 1],  n−1 X

si yi + vsn yn

2

≤ ∥y∥2 ∥s∥2 =

i=1

n X

s2i .

(D.29)

i=1

As each s2i is Ksignal sub-Exponential and n ≤ 2d, the latter is at most K(Ksignal )d with probability at least 1 − K exp(−d) by Bernstein’s Inequality. Combining (D.27) with the aforementioned highprobability bounds for (D.28) and (D.29), the Lemma follows. Next note for any sequences (xj )j≤d , (yi )i≤n , we have X

Mi,j xj yi ≤ B ′

x2j

1/2  X

Now applying the above inequality for xj = XX

P

Mi,j yi

yi2

1/2

.

i≤n

j≤d

i≤n,j≤d

j≤d

X

i≤n Mi,j yi , we obtain that

2 1/2

i≤n

≤ B′

X

yi2

1/2

.

(D.30)

i≤n

n 1 L Lemma D.13. Consider any sequence y = (y √ i )i≤n ∈ R . Then the function f (θ , . . . , θ ) = 1 1 ′ u (θ ), y is Lipschitz with constant B D∥y∥/ d.

Proof. Clearly f only depends on θ1 among θ1 , . . . , θL . Note for all 1 ≤ j ≤ d that by construction of Mi,j ,    1  X ∂ 1 1 ′′ 1 ′′ 1 √ u (θ ), y = y u S (θ ) M + y u S (θ ) M i i,t i,j n n,t,v n,j . ∂θj1 d 1≤i≤n−1    Applying (D.30) for the sequence y1 u′′ S1,t (θ1 ), . . . , yn−1 u′′ Sn−1,t (θ1 ) , yn u′′ Sn,t,v (θ1 ) and recalling |u′′ | ≤ D now gives that  X  ∂f 2 1/2 j≤d

∂θj1

  1/2 B ′ D∥y∥ B ′  X 2 ′′ 1 2 2 ′′ 1 2 √ ≤ yi u Si,t (θ ) + yn u Sn,t,v (θ ) ≤ √ . d i≤n−1 d

Since f only depends on θ1 , the conclusion follows. 108

Lemma D.14. For any function f of θ1 , . . . , θL mapping to R (possibly depending on the disorder), we have D  E 2 KB ′2 2 ≤ u1 (θ1 ) − u1 (θ1 ) f ⟨f ⟩ , d where K depends on D, β. Proof. Consider any vector y ∈ Rn . By Cauchy-Schwarz, D

D D T E  ET T 2 E1/2 u1 (θ1 ) − u1 (θ1 ) f y = f u1 (θ1 ) − u1 (θ1 ) y ≤ ⟨f 2 ⟩1/2 u1 (θ1 ) − u1 (θ1 ) . y

Consider the function θ → u(θ), y . By Lemma D.13, this function is Lipschitz with constant √ B ′ D∥y∥/ d. Applying Theorem B.3, we obtain D

u1 (θ1 ) − u1 (θ1 )

T 2 E K(β)B ′2 D2 ∥y∥2 . y ≤ d

Since y ∈ Rn was arbitrary, the desired conclusion follows. Corollary D.15. For f = R1,1 , R1,2 , defining ∇ as in Lemma D.10, we have !  ⋆ ⋆  B ⋆ B ′2 B B ∥∇∥2 ≤ K + 1+ , d2 d d3 where K depends on D, β0 , β1 , λ and also on κ for the interpolators. Proof. Applying the bound from Lemma D.10, it remains to upper bound   1 X 1 2 2 2 1 1 ⟨ θ1 − ⟨θ1 ⟩ f˙u′ (Sn,t,v )⟩ + (I) := f˙θ˙d1 + ⟨ θ1 − ⟨θ1 ⟩ f˙u′ (Si,t )⟩ , d d i≤n−1  D  E 2  1 u1 (θ1 ) − u1 (θ1 ) f˙ . (II) := 1 + ∥θ1 ∥2 d First, we upper bound (I). To this end, applying Lemma D.11 with the function f˙ or f˙u′ (S 1 ) 1 or S 1 where S 1 = Si,t n,t,v , we have  1 X ˙2 ′ 1 2  K(D, β0 , β1 , κ, λ)B ⋆ 1 1 )2 ⟩ + (I) ≤ K(β0 , β1 ) ⟨f˙2 ⟩ + ⟨f˙2 u′ (Sn,t,v ⟨f u (Si,t ) ⟩ ≤ , d d d2 i≤n−1

where we use |u′ | ≤ D, Proposition D.5 to control ⟨f˙2 ⟩, and that n/d ≤ α0 = 2. Next, we upper bound (II). First, by (D.18), 

1+

 1 K(D, β0 , β1 , κ, λ)B ⋆ ∥θ1 ∥2 ≤ 1 + . d d

By Lemma D.14 and Proposition D.5, D

 E 2 K(D, β)B ′2 2 K(D, β0 , β1 , κ, λ)B ⋆ B ′2 u1 (θ1 ) − u1 (θ1 ) f˙ ≤ ⟨f˙ ⟩ ≤ . d d3

This yields an upper bound on (II), and combining the upper bounds on (I) and (II) establishes Corollary D.15. 109

We now have done the necessary preparation to prove Proposition D.6. Proof of Proposition D.6. For this proof, we let f = ⟨R1,1 ⟩ or ⟨R1,2 ⟩. Similarly as the proof of Theorem B.2, we consider the space S = Rnd+L+2 where the first n coordinates correspond to the sj , the next (n − 1)d coordinates correspond to gi,j , the next L coordinates correspond to W L , and the last two coordinates correspond to Z and Y . We endow S with the product measure γ where the measure on each coordinate is given by law of each corresponding part of the disorder, that is, sj , gi,j , W L , Z or Y . Thus integration w.r.t. γ corresponds to taking expectation w.r.t. the disorder. Consider the set C := {x ∈ S : B ⋆ ≤ K(λ)d , B ′2 ≤ K(λ)d} , where by (D.15) and Lemma D.12, K(λ) has been chosen large enough so that P(C c ) ≤ K(λ) exp(−d) , where probability here is w.r.t. γ. Moreover, C is a convex subset of S, because B ⋆ is a convex function of the disorder, the operator norm of a matrix is a convex function of its entries, and the entries of M are an affine transformation of the disorder. Next, let γ ′ be a probability measure defined on C with density proportional to that of γ. By Assumption 1, γ ′ is min{1, βsignal } strongly log-concave. By Corollary D.15, f is Lipschitz with R 1 ,κ,λ) on C. Let m = f dγ ′ . Since C is convex and since γ ′ is supported on C, constant K(D,β√0 ,β d Theorem B.3 gives that for all k ≥ 1, Z  K(D, β , β , κ, λ)k k 0 1 f − m)2k dγ ′ ≤ . d Thus by definition of γ ′ , Z Z  K(D, β , β , κ, λ)k k   0 1 2k 2k E (f − m) 1{C} = (f − m) dγ = γ(C) (f − m)2k dγ ′ ≤ . d C C We now combine Lemma D.9 and Lemma B.19 together as in the proof of (D.18), and then take expectations and applying Lemma D.8. Using that f = R1,1 or R1,2 , we obtain   E f 4k ≤ K(D, β0 , β1 , κ, λ)k . ⋆

Next we upper bound m. By Lemma D.9 and Lemma B.19, we have |f | ≤ K(D,β0 ,βd1 ,κ,λ)B . Hence by definition of C ⋆ , Z |m| = f dγ ′ ≤ K(D, β0 , β1 , κ, λ) .   It follows from the above two displays that E (f − m)4k ≤ K(D, β0 , β1 , κ, λ)k . Thus       E (f − m)2k = E (f − m)2k 1{C} + E (f − m)2k 1{C c }  K(D, β , β , κ, λ)k k   0 1 + E (f − m)4k P(C c ) ≤ d  K(D, β , β , κ, λ)k k 0 1 ≤ + K(D, β0 , β1 , κ, λ)k · K(λ) exp(−d) d  K(D, β , β , κ, λ)k k 0 1 ≤ , d where the last inequality uses that exp(−d) ≤ (k/d)k , valid for all k, d > 0. The result now follows from symmetrization, Lemma D.7. 110

As noted before, combining Proposition D.5 and Proposition D.6 proves the first part of Proposition A.6. D.3.2

Second part of Proposition A.6

Recall that we have (D.19) by Lemma D.9 and Lemma B.19. Now by Hölder’s Inequality and (D.15), taking K(D, β0 , β1 , κ, λ) appropriately and combining with (D.19), we obtain   log νt,v exp

 D  E ∥θ̄∥2 ∥θ̄∥2 = log E exp K(D, β0 , β1 , κ, λ K(D, β0 , β1 , κ, λ) 1 D  β∥θ̄∥2 E 1+KB.19 (KD.9 ,D,β0 ,β1 ,κ,λ) ≤ log E exp 2 ⋆ ≤ log E exp(B ) ≤ K(λ) d .

It remains to prove that   log νt,v exp

 θd2 ≤ K(D, β0 , β1 , κ, λ) . K(D, β0 , β1 , κ, λ)

(D.31)

To this end, we cite the following Lemma from [Tal10]. This function is stated for concave T ≤ 0 and β ′ > 0 in [Tal10], however the proof therein applies verbatim under the conditions stated below. Lemma D.16 (Lemma 3.2.5, [Tal10]). Consider a concave function T (θ) defined on Rd such that the following integrals are well-defined. Consider any (aj )j≤d , β > 0, β ′ ≥ 0 and any convex set C ⊆ Rd . Define the measure µ on Rd by defining for all B ⊆ Rd , Z   X aj θj dθ . µ(B) ∝ exp T (θ) − β∥θ∥2 − β ′ θd2 + B∩C

j≤d

Let  C ′ := θd ∈ R : ∃θ′ ∈ Rd−1 s.t. (θ′ , θd ) ∈ C} . Then the following function f on C ′ is concave, where f is defined by Z   X X f (θd ) := log exp T (θ) − β θj2 + aj θj dθ′ . θ=(θ′ ,θd )∈C

j≤d−1

j≤d−1

Moreover, letting w(θ) := f (θ) − (β + β ′ )θ2 + ad θ , the law of θd under µ is the measure on C ′ with density proportional to exp w(θ). We will also need the following corollary of the Prékopa-Leindler Inequality: Lemma D.17. Consider a log-concave distribution θ and and a random variable W with a logconcave density. Then for any function T (θ, W ) that is jointly concave in θ and W , the function  exp T (θ, W ) is concave in θ. θ → log EW 111

Proof. Letting φ(W ) denote the density of RW , it follows that S(θ, W ) = exp T (θ, W )φ(W ) is logconcave. Thus it remains to prove that θ → S(θ, W )dW is log-concave. Considering any λ ∈ [0, 1] and any θ1 , θ2 , we have by log-concavity of S(θ, W ) that for any W , S(λθ1 + (1 − λ)θ2 , W ) ≥ S(θ1 , W )λ S(θ2 , W )1−λ . It follows from the Prékopa-Leindler Inequality (see [Pré73]) that Z Z λ  Z 1−λ . S(λθ1 + (1 − λ)θ2 , W )dW ≥ S(θ1 , W )dW S(θ2 , W )dW This proves that

R

S(θ, W )dW is log-concave, hence the conclusion.

We now complete the proof of the second part of Proposition A.6. By Theorem B.3 applied with the 1-Lipschitz function θ → θd , we have D

exp

 (θ − ⟨θ ⟩)2 E d

d

K(β0 )

≤K.

(D.32)

 Next, note by convexity of the function exp x2 that for f1 , f2 satisfying νt,v exp(f12 /K) ≤ K,   νt,v exp(f22 /K) ≤ K, we have the upper bound νt,v exp (f1 + f2 )2 /K ≤ K (here K can differ). Thus taking expectations of (D.32) and applying this observation, to establish (D.31), it thus remains to show that h  E exp

i ⟨θd ⟩2 ≤ K(D, β0 , β1 , κ, λ) . K(D, β0 , β1 , κ, λ)

(D.33)

√ To this end, we apply Lemma D.16. Define the convex set C = Rd ∩ 1{θ1 / d ∈ I}, and let T0 (θ, W ) :=

X

u(Si,t ) + u(Sn,t,v ) .

1≤i≤n

Since Si,t is affine in θ and as Sn,t,v is jointly affine in θ and W , since u is concave, it follows   that T0 (θ, W ) is concave jointly in θ, W . Thus letting T (θ) := log EW exp T0 (θ) , T is concave by Lemma D.17. Next, let β ′ = (1−t)(r−r̄) ≥ 0, where the inequality follows as r ≥ r̄. Define aj for 1 ≤ j ≤ d 2 p by letting aj = 0 for all j ≤ d − 1 and ad = (1 − t)rY . Defining µ from Lemma D.16 with these choices, as dependence on W in νt,v is only via Sn,t,v (θ), as T (θ) is defined with expectation w.r.t. W , and as the  number of replicas L = 1 here, it follows that µ is exactly the Gibbs average ⟨·⟩. ′ Note C = θ ∈ R : ∃θ′ ∈ Rd−1 s.t. (θ′ , θ) ∈ C} = R. Hence by Lemma D.16, the following function is concave on all of R: Z h  i X 2 f (θd ) := log E exp T (θ, W ) − β θ dθ′ . 0 W j √ θ=(θ′ ,θd )∈Rd ∩1{θ1 / d∈I}

j≤d−1

Moreover, by Lemma D.16, we know that letting  (1 − t)(r − r̄)  2 p w(θ) := f (θ) − β + θ + (1 − t)rY θ , 2 the law of θd under µ is the measure on R with density proportional to exp w(θ). 112

As r ≥ r̄, we know β + (1−t)(r−r̄) > 0, and thus w(θ) is β strongly concave and attains a unique 2 ⋆ maximum θ on R. By Lemma F.1, it follows that (θd − θ⋆ )2 ≤ K(β0 ). By Jensen and Triangle Inequality, we thus have ⟨θd ⟩ ≤ K(β0 ) + |θ⋆ | (for a different K). Note that w′p (θ⋆ ) = 0. Since w is β strongly concave, |w′ (0)| = |w′ (θ⋆ ) − w′ (0)| ≥ 2β|θ⋆ |. Note ′ ′ w (0) = f (0) + (1 − t)rY . Combining the above steps yields   p θd ≤ K(β0 ) 1 + |f ′ (0)| + (1 − t)r|Y | . (D.34) Note since Y is a standard Gaussian, we have exp(Y 2 ) ≤ K. By (D.34), since r ≤ K(D), it suffices to prove the following to establish (D.33) and hence complete the proof: h  E exp

i f ′ (0)2 ≤ K(D, β0 , β1 , κ, λ) . K(D, β0 , β1 , κ, λ)

(D.35)

Let ⟨·⟩f denote the Gibbs average defined as follows: for any test function h, Z h  i X p 1 2 ⟨h⟩f = E h exp T (θ, W ) − β θ + (1 − t)rY θ dθ , 0 W d j Z̄ θ=(θ′ ,θd )∈Rd ∩1{θ1 /√d∈I} j≤d−1

where Z̄ denotes the corresponding normalizing constant. Explicitly calculating yields + * r r X t tv gi,d + u′ (Sn,t,v ) gn,d . f ′ (θd ) = u′ (Si,t ) d d 1≤i≤n−1

(D.36)

f

p Notice that when θd = 0, Si,t , Sn,t,v , (1 − t)rY θd = 0, and therefore T0 (θ, W ) do not depend on (gi,d )1≤i≤n , Y . Hence when θd = 0, Z̄ also does not depend on (gi,d )1≤i≤n , Y . (D.36) thus implies that f ′ (0) can be written as a linear combination of (gi,d )1≤i≤n with coefficients independent of q q (gi,d )1≤i≤n . Specifically, these coefficients are Gibbs averages of u′ (Si,t ) dt , u′ (Sn,t,v ) tv d (where we consider Si,t , Sn,t,v with θd = 0) w.r.t. ⟨·⟩f . Since the (gi,d )1≤i≤n ∼ N (0, 1), it follows that f ′ (0) is a Gaussian. Let Ē √ denote expectation in the gi,d . As |u′ | ≤ D, the corresponding coefficients on the gi,d are most D/ d in magnitude. By the independence of the gi,d , it follows that  n  Ē f ′ (0)2 ≤ · D2 ≤ K(D) , d where we use that n/d ≤ α0 ≤ 2. Since f ′ (0) is Gaussian, it follows from standard upper bounds for the MGF of a Gaussian with a given variance (see (A.11) in [Tal10]) that for suitable K(D, β0 , β1 , κ, λ), h  E exp

i h  i f ′ (0)2 f ′ (0)2 = E Ē exp ≤ 2. K(D, β0 , β1 , κ, λ) K(D, β0 , β1 , κ, λ)

This establishes (D.35) and hence establishes Proposition A.6.

E

Properties of the RS equations

Here, we establish several important properties of the RS equations. Specifically, we prove Proposition A.10 in Appendix E.1, Proposition A.3 in Appendix E.2, and Lemma C.3 in Appendix E.3. 113

Throughout this section, recall that φ0 denotes the PDF of a standard univariate normal N (0, 1) and φ1 denotes the PDF of S ∼ D. All three of these results are critical, and are proved through similar ideas. Discussing the proof of Proposition A.3 for the interpolators for concreteness, we must establish that the RS equations exhibit a particular convex-concave structure. Talagrand’s proof of this convex-concave structure in Section 3.3, [Tal10] proceeds by showing each part of the expression for the relevant second derivatives is of the correct sign. This relies crucially on the fact that κ ≥ 0 and that there is no signal in the data. Here we instead proceed by showing the part of the expression for the relevant second derivatives whose signs are not correct are small in magnitude, and therefore do not eliminate the desired convex-concave structure. We must also bound similar terms h in proving iProposition A.10. The √ κ−xS− qZ  √ magnitude of these terms all can be upper bounded by α E f for a function f (x) ≤ ρ−q O(|x|2 + 1) that depends on the Inverse Mills’ Ratio from (3.12), where expectation is over Z, S. of the expression for each such part can be directly shown to be of order  The 2 magnitude  κ +q+1 O α · ρ−q for all κ ∈ R. However, to establish our result for all α ≤ α0 for all κ < 0, such a √ κ−xS− qZ

bound is not sufficient. The idea is instead as follows. When √ρ−q ≤ 0, the relevant quantities are simple to control by properties of the Inverse Mills’ Ratio. Else, we let κ′ := κ − xS be the ‘effective margin’ and analyze the sign of κ′ . • If |S| < |κ|, as√ |x| ≤ 1, we must have κ′ < 0. In this case, observe that the probability over √ κ−xS− qZ κ′ − qZ √ Z that = √ρ−q > 0 is exponentially small in |κ′ |2 /q. ρ−q • Else, the probability over S that |S| ≥ |κ| is exponentially small in |κ|2 as the law of S 2 is sub-Exponential. h i √ κ−xS− qZ  √ Combining both these bounds enables us to upper bound E f independently of κ for ρ−q κ < 0, which is sufficient for our purposes.

E.1

Proof of Proposition A.10

We consider an arbitrary solution (q, ρ, r, r̄) of (A.30) and show that (q, ρ) ∈ [0, Csol ) × [0, Csol ) 1 and ρ−q < Csol . This is sufficient to establish Proposition A.10, up to showing q > 0. To this end, note q = 0 implies r = 0. Now as λ > 0, we have u′ > 0 on a set of positive Lebesgue measure. It follows that for every Z, S, there exists a set of W with positive Lebesgue measure such that u′ (η) exp u(η) > 0, where η is given in terms of W, Z, S by (A.19). Since u′ exp u ≥ 0 pointwise, it follows that r > 0 by definition of (A.30). Hence q = 0 yields contradiction, so as q ≥ 0, we must have q > 0. Proof for posterior. The proof of Proposition A.10 for the posterior is direct. Recall that |ul | ≤ D = D(λ) for l = 1, 2, 3, 4. Thus by definition of ψα , ψ̄α in (A.22), (A.23), we have |r|, |r̄| ≤ K(D). Also note r ≥ 0 by definition of ψα . Since r ≥ r̄, the definition of the system 1 1 ≥ K(D) . This proves Proposition (A.30) implies 0 ≤ q, ρ ≤ K(D). Next, we note ρ, ρ − q ≥ 2β+r−r̄ A.10 for the posterior. Proof for interpolators. The rest of Appendix E.1 is now devoted to the proof of Proposition A.10 for the interpolators, which poses significantly more difficulties. We first state and prove the following Lemmas which we will need in the following proof. In the following, we recall the definition of η = η(q, ρ) from (A.19). 114

Lemma E.1. Letting s ∼ D, define   √ v(y) = log EW exp u xS + y + ρ − qW . Then we have v ′ ≥ 0 and  √ 2   √  r = α E v ′ qZ , r̄ − r = α E v ′′ qZ . Proof. The fact that v ′ ≥ 0 follows directly as u′ ≥ 0. The equalities written above follow from direct calculation, using that r = ψα (q, ρ), r̄ = ψ̄α (q, ρ). Lemma E.2. Letting Y :=

√ κ−xS− qZ √ ρ−q

and KE.2 = 20, we have   !2 EW W exp u(η)   ≤ KE.2 + Y 2 1{Y ≥ 0} . EW exp u(η)

Proof. If Y ≤ 1, then as u ≤ 0 and u(η) = 0 for η ≥ κ, we have r     2 , EW W exp u(η) ≤ EW |W | = π   EW exp u(η) ≥ PW (η ≥ κ) = PW (W ≥ Y ) ≥ PW (W ≥ 1) . As KE.2 ≥ 20, we obtain an upper bound of KE.2 in this case. Else suppose Y ≥ 1. First as u′ ≥ 0, observe that     √ EW W exp u(η) = EW exp u(η) · u′ (η) · ρ − q ≥ 0 . Thus we may upper bound       EW W exp u(η) = EW W 1{W ≤ Y } exp u(η) + EW W 1{W ≥ Y } exp u(η)     ≤ Y EW 1{W ≤ Y } exp u(η) + EW W 1{W ≥ Y }   1 ≤ Y EW 1{W ≤ Y } exp u(η) + √ exp(−Y 2 /2) . 2π We next lower bound, using that u(η) = 0 for W ≥ Y ,       EW exp u(η) = EW 1{W ≤ Y } exp u(η) + EW 1{W ≥ Y } exp u(η)   = EW 1{W ≤ Y } exp u(η) + P(W ≥ Y )   1 Y exp(−Y 2 /2) . ≥ EW 1{W ≤ Y } exp u(η) + √ · 2 2π Y + 1 Here the lower bound on P(W ≥ Y ) is standard, see e.g. [Ver18]. We thus obtain     Y EW 1{W ≤ Y } exp u(η) + √12π exp(−Y 2 /2) EW W exp u(η)   ≤   0≤ EW exp u(η) EW 1{W ≤ Y } exp u(η) + √12π · Y 2Y+1 exp(−Y 2 /2) ≤Y +

1 , Y

where we use the following inequality stated on p. 238 of [Tal10] that for a, b > 0, a+aY +b ≤ Y + Y1 . Y b 1+Y 2

Hence we have, as Y ≥ 1 in this case,  E W exp u(η) 2 1 W   ≤ Y 2 + 2 + 2 ≤ Y 2 + 3. Y EW exp u(η) This completes the proof of the Lemma. 115

Lemma E.3. For any q ≤ ρ, we have ψα (q, ρ) ≥ ψ̄α (q, ρ). In particular, we have r ≥ r̄. Proof. When q = ρ the desired inequality follows as u′′ ≤ 0. Else when q < ρ, it suffices to show      EW W exp u(η) 2 EW (W 2 − 1) exp u(η)     ≤ EW exp u(η) EW exp u(η)    2  EW W exp u(η) 2     ≤ EW exp u(η) . ⇐⇒ EW W exp u(η) − EW exp u(η) Letting w(y) = u(xS + yields for some real y ′ , 

qZ +



2

ρ − q · y) − y2 which is 2 strongly-concave as u′′ ≤ 0, Lemma F.1

′ 2





EW exp u(η) ≥ EW (W − y ) exp u(η) ≥ EW



   EW W exp u(η) 2   , W exp u(η) − EW exp u(η) 2

completing the proof of the Lemma. Now we return to the proof of Proposition A.10. Bounding q, ρ: First, clearly 0 ≤ q ≤ ρ by definition of the system (A.30). It remains to upper bound q, ρ. Now, note as r ≥ r̄ by Lemma E.3, ρ−q =

1 1 ≤ . 2β + r − r̄ 2β0

(E.1)

It thus remains to upper bound q. If 0 ≤ q ≤ 1 we obtain 0 ≤ q, ρ ≤

1 + 1 < Csol , 2β0

yielding the required bound. We now suppose that q > 1. Now by definition of (A.30), we have r = q(2β + r − r̄)2 = Next by Lemma E.2, letting Y :=  r = ψα q, ρ =

q . (ρ − q)2

(E.2)

√ κ−xS− qZ √ , we have ρ−q

 E W exp u(η) 2   α α  W   E ≤ KE.2 + E Y 2 1{Y ≥ 0} . ρ−q ρ−q EW exp u(η)

(E.3)

  Now we upper bound E Y 2 1{Y ≥ 0} . Upper bound for κ ≥ 0.

Using E[S 2 ] ≤ Ksignal , E[Z 2 ] = 1, x2 ≤ 1, we have

  E Y 2 1{Y ≥ 0} ≤ E[Y 2 ] ≤

  3(κ2 + Ksignal + q) 3 E κ2 + x2 S 2 + qZ 2 ≤ . ρ−q ρ−q 116

(E.4)

Upper bound for κ < 0. Here as we would like to establish this result for all α ≤ α0 at most a universal constant, the proof idea is more complicated. Define the ‘effective margin’ κ′ := κ − xS .

(E.5)

We now analyze the sign of κ′ . Note: • If |S| < |κ|, as |x| ≤ 1, we must have κ′ = κ − xS < 0. Then for Y < 0 to occur, we must ′| √ √ , the probability of which over Z is exponentially unlikely in |κ′ |/ q. have |Z| ≥ |κ q • Else, the probability over S that |S| ≥ |κ| is exponentially unlikely in |κ|2 as the law of S 2 is sub-Exponential.   Thus in both cases we can upper bound E Y 2 1{Y ≥ 0} . We now execute this idea. Write Z √   (κ′ − z q)2 √ E Y 2 1{Y ≥ 0} = 1{κ′ ≥ z q}φ0 (z)φ1 (s) dzds ρ−q s,z Z Z √ (κ′ − z q)2 √ = 1{κ′ ≥ z q}φ0 (z)φ1 (s) dzds ρ − q s:|s|<|κ| z Z Z √ (κ′ − z q)2 √ + 1{κ′ ≥ z q}φ0 (z)φ1 (s) dzds . ρ − q s:|s|≥|κ| z If |s| < |κ|, then |xs| < |κ|, so κ′ = κ − xs < 0 as κ < 0. Recall q ≥ 1, otherwise the proof is ′| √ √ . Hence already completed. Thus κ′ ≥ z q implies that |z| ≥ |κ q Z Z √ (κ′ − z q)2 √ 1{κ′ ≥ z q}φ0 (z)φ1 (s) dzds ρ−q s:|s|<|κ| z Z Z  ′2  q |κ | 1 ≤ 2 + z 2 · exp(−|κ′ |2 /4q) · √ exp(−z 2 /4)φ1 (s) dzds ρ − q s:|s|<|κ| z q 2π ! Z Z  |κ′ |2   2q  ≤ φ1 (s) ds + z 2 · exp(−|κ′ |2 /4q) exp(−z 2 /4) dz ρ−q s q z ≤

10q . ρ−q

When |s| ≥ |κ|, we employ a similar strategy, but now we analyze the probability in S rather than Z. Since q ≥ 1, we may bound Z Z √ (κ′ − z q)2 √ 1{κ′ ≥ z q}φ0 (z)φ1 (s) dzds ρ−q s:|s|≥|κ| z Z Z  2q 2(κ2 + s2 )  ≤ φ0 (z)φ1 (s) dzds z2 + ρ − q s:|s|≥|κ| z q Z Z Z  2q  2 2 2 ≤ z φ0 (z) dz + 2 s φ1 (s) ds + 2κ φ1 (s) ds ρ−q z s s:|s|≥|κ|    2q 1 + 2Ksignal + 2κ2 P S 2 > κ2 ≤ ρ−q  κ2  2q  1 + 2Ksignal + 4κ2 exp − ≤ ρ−q Ksignal  2q  ≤ 1 + 3.5Ksignal . ρ−q 117

Here we used that S 2 is Ksignal sub-Exponential to upper bound P(S 2 > κ2 ), and used Assumption 1 to upper bound E[S 2 ] ≤ Ksignal . Combining the above displays yields    q  E Y 2 1{Y ≥ 0} ≤ 12 + 7Ksignal . (E.6) ρ−q   Having upper bounded E Y 2 1{Y ≥ 0} in both cases, in the κ > 0 case we obtain via (E.3) and (E.4), using that q ≥ 1 by assumption here (else we have already finished the proof) and that Ksignal ≥ 1,   α  q = r(ρ − q)2 ≤ (ρ − q)2 · KE.2 + E Y 2 1{Y ≥ 0} ρ−q  ≤ αKE.2 (ρ − q) + 3αq κ2 + Ksignal + 1 . Similarly in the κ < 0 case we have by (E.3) and (E.6),  q = r(ρ − q)2 ≤ αKE.2 (ρ − q) + αq 12 + 7Ksignal .  1 when κ < 0 and α0 ≤ 3(κ2 +K1signal +1) when κ ≥ 0, combining Since α0 ≤ min 2K1E.2 , 24+14K signal with (E.1) yield that in either case, we have q≤

q ρ−q ρ q 1 + = ≤ + , 2 2 2 2 4β0

thus

q≤

1 . 2β0

Consequently we obtain ρ ≤ β10 from (E.1), and thus 0 ≤ q, ρ < Csol . 1 Upper bounding ρ−q :

From the definition of the system (A.30), we have  √  1 = 2β + r − r̄ = 2β − α E v ′′ qZ . ρ−q

Next by Gaussian Integration by Parts and as v ′ ≥ 0, we have  √   1 √  E v ′′ qZ = √ E Zv ′ qZ q  1 √  ≥ √ E 1{Z ≤ 0} · Zv ′ qZ q  1/2  √ 2 1/2 1 ≥ − √ E Z 2 1{Z ≤ 0} · E v ′ qZ q  √ 2 1/2 1 = − √ E v ′ qZ . 2q Combining the above with Lemma E.1 gives  √ 2 1/2 1 α ≤ 2β + √ E v ′ qZ = 2β + ρ−q 2q

r

αr . 2q

Combining the above with (E.2) and using α ≤ α0 ≤ 1/2 yields r 1 α 1 1 ≤ 2β + · , thus ≤ 4β ≤ 4β1 < Csol . ρ−q 2 ρ−q ρ−q 118

E.2

Proof of Proposition A.3

Throughout Appendix E.2, I is held fixed, and so we denote FI by F for simplicity of notation. As done in Chapter 3 of [Tal10], consider the transformation y=

q . ρ−q

yρ Hence given y, we set q = 1+y , and have ρ > q. Define

G(y, ρ) := F

 yρ  ,ρ . 1+y

Thus we have r r i y 1  yρ ρ 1 Z+ W + + log ρ − log(1 + y) G(y, ρ) = α E log EW exp u xS + 1+y 1+y 2 2 2 (E.7)  − β ρ + rI (x) , h

where u(·) is given by exp u(x) = 1{x ≥ κ} for the interpolators, and by (3.17) in the GMM case or (3.18) in the logistic case for the posterior. Consider the system ∂G ∂G = = 0. ∂ρ ∂y

(E.8)

Note this system does not depend on the −βr(x) term in the definition of G(y, ρ). Lemma E.4. The system (A.11) has a unique solution (q0 , ρ0 ) ∈ [0, Csol ]×[ C1sol , Csol ] with q0 < ρ0 iff (E.8) has a unique solution (y0 , ρ0 ) ∈ [0, ∞) × [ C1sol , Csol ]. q maps any (q, ρ) ∈ [0, Csol ] × [ C1sol , Csol ] with Proof. First, notice the transformation y = ρ−q q < ρ to some (y, ρ) ∈ [0, ∞) × [ C1sol , Csol ]. Also for any (y, ρ) ∈ [0, ∞) × [ C1sol , Csol ], notice the yρ yρ transformation q = 1+y is such that 0 ≤ q < ρ ≤ Csol . Thus the transformation q = 1+y maps any 1 1 (y, ρ) ∈ [0, ∞) × [ Csol , Csol ] to some (q, ρ) ∈ [0, Csol ] × [ Csol , Csol ] with q < ρ.

Thus, it suffices to show that (q0 , ρ0 ) ∈ [0, Csol ]×[ C1sol , Csol ] with q0 < ρ0 is a solution to (A.11) 0 . Establishing that (q0 , ρ0 ), iff (y0 , ρ0 ) ∈ [0, ∞) × [ C1sol , Csol ] is a solution to (E.8), where y0 = ρ0q−q 0 (y0 , ρ0 ) belong in the appropriate domains has already been established above. To show that this transformation preserves solutions, note  ∂F  yρ  ∂q ∂F  yρ ∂G (y, ρ) = ,ρ + ,ρ · (y, ρ) ∂ρ ∂ρ 1 + y ∂q 1 + y ∂ρ

,

 ∂q ∂G ∂F  yρ (y, ρ) = ,ρ · (y, ρ) . ∂y ∂q 1 + y ∂y

∂q ρ0 1 Note ∂y (y0 , ρ0 ) = (1+y Thus if (y0 , ρ0 ) solves (E.8), then we have 2 > 0 as ρ0 > Csol > 0. 0) ∂F ∂F ∂q (q0 , ρ0 ) = 0, and consequently ∂ρ (q0 , ρ0 ) = 0. Moreover if (q0 , ρ0 ) solves (A.11), plugging into the equations above directly implies that (y0 , ρ0 ) solves (E.8). This proves the Lemma.

Thus in what follows, we now work with (E.8). 119

Preliminary calculations for interpolators. We first explicitly write the system (E.8) and establish some of its useful properties. Since we consider ρ0 ≥ C1sol > 0, we may divide by ρ in the following. Now let r r 1+y 1+y ′ √ √ V := (E.9) (κ − xS) − yZ = κ − yZ , ρ ρ where here and in the following, as in the proof of Proposition A.10, we define the ‘effective margin’ κ′ := κ − xS as per (E.5). Hence letting N (·) denote the complementary c.d.f. of the standard normal as per (3.2), we have 

log PW xS +

r

yρ Z+ 1+y

r

  ρ W ≥ κ ≡ log N 1+y

r

1+y √  (κ − xS) − yZ . ρ

(E.10)

Recalling the definition of the Inverse Mills’ Ratio A(x) in (3.12), we record a few of its properties. Lemma E.5 (Lemma 3.3.7 of [Tal10] and Lemma 17, [EAS22]). We have: 1. A(x) ≥ x for all x ∈ R. 2. A′ (x) = A(x)2 − xA(x) ≥ 0 for all x ∈ R. 3. xA(x)A′ (x) ≤ A(x)2 for all x ∈ R. 4. xA(x) ≤ 1 + x2 for all x ∈ R. 5. limx→∞ A(x) x = 1. 6. A′ (x), A′′ (x) ≤ KMILLS for a universal constant KMILLS ≥ 1, 0 ≤ A(x) ≤ 1 for all x < 0, and 0 ≤ A(x) ≤ 2x + 1 for all x ≥ 0. Proof. All the above are proved in Lemma 3.3.7 of [Tal10] and Lemma 17, [EAS22] except the bound A′′ (x) ≤ KMILLS . This follows from noting as in the proof of Lemma 17, [EAS22] the bound A′′ (x) = A(x)P (x, A(x)) for a polynomial P (·, ·). It thus suffices to control A′′ (x) for x sufficiently large. The x → +∞ case follows from Corollary 1.6 of [Pin19], while the x → −∞ case follows from noting that A(x) → 0 exponentially fast as x → −∞. We now explicitly write the system (E.8). By explicit calculation and (E.10), we have i i ∂G α h (1 + y)1/2 1 α h ′ (1 + y)1/2 1 = E (κ − xS) A(V ) + − β = E κ A(V ) + −β, ∂ρ 2 2ρ 2 2ρ ρ3/2 ρ3/2 h i ∂G κ − xS Z  y = α E − 1/2 . + A(V ) + 1/2 1/2 ∂y 2(1 + y) 2ρ (1 + y) 2y Applying Gaussian Integration by Parts and Lemma E.5, h Z i h i h i E √ A(V ) = − E A′ (V ) = E − A(V )2 + V A(V ) y r h i h i 1+y √  2 − yZ A(V ) . = E − A(V ) + E (κ − xS) ρ 120

(E.11) (E.12)

Rearranging yields h Z 1 E − A(V )2 + (κ − xS) √ E[A(V )] = y 1+y

r

i 1+y A(V ) . ρ

We thus obtain h i ∂G α y =− E A(V )2 + . ∂y 2(1 + y) 2(1 + y)

(E.13)

Following the argument of [ST03], [Tal10], we establish that (E.11), (E.13) have a unique solution 2 by considering ∂∂ 2G and the sign of ∂G ∂y , and arguing that the signs work out to yield a saddle point ρ and thus a unique solution to (E.8). We first compute from (E.11), i i α(1 + y)1/2 h ∂2G 3α(1 + y)1/2 h ′ 1 1 1/2 −3/2 ′ ′ ′ ρ κ − 2 = − (1 + y) E κ A(V ) + E κ A (V ) · − 2 5/2 3/2 ∂ρ 2 2ρ 4ρ 2ρ h i h i 1/2 1 3α(1 + y) α(1 + y) E κ′2 A′ (V ) − 2 . (E.14) =− E κ′ A(V ) − 3 5/2 4ρ 2ρ 4ρ Now in preparation on our work on the sign of ∂G ∂y , we define   g(y, ρ) := y − α E A(V )2 ,

thus

∂G g(y, ρ) (y, ρ) = . ∂y 2(1 + y)

(E.15)

∂g We will need to show ∂y > 0 in an appropriate domain, which we do in Lemma E.10. In preparation, we compute

i h  κ′ ∂g 1 = 1 − α E 2A(V )A′ (V ) p − Zy −1/2 ∂y 2 ρ(1 + y) 2 h i i α α h =1− p E κ′ A(V )A′ (V ) + √ E ZA(V )A′ (V ) . y ρ(1 + y) √ Notice ∂V ∂Z = − y, so Gaussian Integration by Parts gives h i h i √ E ZA(V )A′ (V ) = − y E A′2 (V ) + A(V )A′′ (V ) . Thus h i h i ∂g α =1− p E κ′ A(V )A′ (V ) − α E A′2 (V ) + A(V )A′′ (V ) . ∂y ρ(1 + y)

(E.16)

To control the signs of the above, since we do not have non-negative margin – κ ≥ 0 or κ′ = κ − xS ≥ 0 is not uniformly true – the signs of each component of the derivative do not work out directly in the same way as in [ST03]. Rather we show that for α ≤ α0 , the part of the derivatives that do not have the correct sign are small in magnitude. We will carry this argument out in the proofs to follow. 121

Preliminary calculations for posterior. r V̄ := xS +

We now define yρ Z+ 1+y

r

ρ W. 1+y

(E.17)

Here in Appendix E.2, we let ⟨·⟩ denote a Gibbs average w.r.t. EW exp u(V̄ ). In what follows, when the derivatives ul are written, they are taken with argument V̄ unless otherwise stated. We obtain by Gaussian Integration by Parts, α αy ∂G 1 2 = E u′′ + u′2 − E u′ + −β, ∂ρ 2 2(1 + y) 2ρ ∂G αρ y 2 =− E exp u′ + . 2 ∂y 2(1 + y) 2(1 + y)

(E.18) (E.19)

We will follow the same strategy as in the interpolators of establishing the saddle point structure of the RS equations. As such we compute  α 1 ∂2G = (I) + (II) + (III) + (IV) − 2, 2 ∂ρ 4 2ρ

(E.20)

where h i 1 2 E u′′′′ + 4u′′′ u′ + 3u′′2 + 6u′′ u′2 + u′4 − u′′ + u′2 , 1+y h y 2 E u′′′′ + 4u′′′ u′ + 3u′′2 + 6u′′ u′2 + u′4 + 2 u′′ + u′2 u′ (II) := 1+y (I) :=

i − u′′′ + 3u′′ u′ + u′3 u′ − u′′′ + 3u′′ u′ + u′3 u′′ + u′2 , (E.21) h i −y ′′′ ′′ ′ ′3 ′ ′′ ′2 ′ 2 (III) := E u + 3u u + u u − u + u u , 2(1 + y)2 h i −y 2 ′′ ′2 2 ′′′ ′′ ′ ′3 ′ ′ 4 ′′ ′2 ′ 2 E (IV) := u + u + u + 3u u + u u + 3 u − 5 u + u u . 2(1 + y)2 We also define similarly to (E.15), g(y) := y(1 + y) − αρ E exp u′

2

,

thus

∂G g(y) (y, ρ) = . ∂y 2(1 + y)2

(E.22)

∂g Again in Lemma E.10, we show ∂y > 0. In preparation, we compute

 ∂g αρ  = 1 + 2y − (I) + (II) , ∂y (1 + y)2

(E.23)

where h i 2 (I) := E − u′′′ + 3u′′ u′ + u′3 u′ + u′′ + u′2 u′ , h i 2 4 2 (II) := E u′′ + u′2 + u′′′ + 3u′′ u′ + u′3 u′ + 3 u′ − 5 u′′ + u′2 u′ 122

(E.24)

Bounding relevant quantities for the interpolators. To study the system (E.8) for the interpolators, we will need to bound various quantities involving the Inverse Mills’ Ratio. This step is not necessary to study (E.8) for the posterior. To this end, we first control E[A(V )2 ] in terms of y. We then use this to upper bound  y. Finally  we2 use these two upper  bounds to upper bound the quantities of interest, namely E A(V ) , E A(V ) , E |κ − xS|A(V ) . Lemma E.6. Consider (α, β) ∈ [0, α0 ] × [β0 , β1 ]. Then for any y ≥ 0 and ρ ≥ C1sol , h i E A(V )2 ≤ KE.6 y + KE.6 , where ( 12Csol Ksignal + Ksignal + 51 KE.6 = 24Csol (κ2 + Ksignal ) + 12

: κ < 0, : κ ≥ 0.

Proof. We perform a similar argument as in the proof of Proposition A.10. We break into the two cases κ ≥ 0 and κ < 0. When κ ≥ 0. By Lemma E.5, 0 ≤ A(x) ≤ 2|x| + 1 for all x ∈ R. Consequently using x2 ≤ 1 and ρ ≥ C1sol , we have 1+y + 12yZ 2 + 3 ρ ≤ 24Csol (κ2 + S 2 )(y + 1) + 12yZ 2 + 3 .

A(V )2 ≤ (2|V | + 1)2 ≤ 6V 2 + 3 ≤ 12κ′2 ·

As E[S 2 ] ≤ Ksignal , E[Z 2 ] = 1, we obtain i h E A(V )2 ≤ 24Csol (κ2 + Ksignal )(y + 1) + 12y + 3 ≤ KE.6 y + KE.6 . When κ < 0. Recalling that κ′ = κ − xS, write i Z h A(V )2 φ0 (z)φ1 (s) dzds E A(V )2 = s,z Z Z Z 2 = A(V ) φ0 (z)φ1 (s) dzds + s:|s|<|κ|

z

s:|s|≥|κ|

Z

A(V )2 φ0 (z)φ1 (s) dzds .

z

We separately upper bound each of these two terms: • If |s| < |κ|, then as |xs| < |κ| and κ < 0, we must have κ′ = κ − xs < 0. Now, consider the 1 set E := {(s, z) : V ≤ 1}. Hence for (s, z) ∈ E, A(V ) ≤ √2πN . (1) Otherwise consider (s, z) ̸∈ E, so here we have First note that in this case, y > 0; if q V > 1. √ 1+y ′ ′ y = 0, then as κ < 0, we have 1 < V = κ ρ − z y < 0 which is a contradiction. Thus recalling ρ > 0, r r 1+y 1+y 1 √ ′ ′ 1<V =κ − z y =⇒ z ≤ κ −√ . ρ yρ y Hence as κ′ < 0 and as y, ρ > 0, we have z 2 ≥ |κ′ |2 ·

1+y 1 |κ′ |2 + ≥ . yρ y ρ

123

Moreover when V > 1, by Lemma E.5,   1+y A(V )2 ≤ (2V + 1)2 ≤ 8V 2 + 2 ≤ 16 |κ′ |2 · + z2y + 2 . ρ Consequently we can upper bound Z Z A(V )2 φ0 (z)φ1 (s) dzds s:|s|<|κ| z Z Z A(V )2 φ0 (z)φ1 (s) dzds + = (s,z)̸∈E,|s|<|κ|

Z

A(V )2 φ0 (z)φ1 (s) dzds

(s,z)∈E,|s|<|κ|

 16|κ′ |2 (y + 1)

1 + z 2 y + 2 φ0 (z)φ1 (s) dzds + ρ 2πN (1)2 (s,z)̸∈E,|s|<|κ| Z  |κ′ |2  |κ′ |2 1 1 ≤ 16(y + 1) exp − · √ exp(−z 2 /4)φ1 (s) dzds + y + 2 + ρ 4ρ 2πN (1)2 2 2π R ≤ 35y + 43 .   ′2 ′2 Here we used that E[Z 2 ] = 1, and upper bounded κρ exp − κ4ρ by using that for all x ≥ 0, we have x exp(−x/4) ≤ K where K is a universal constant. =



• Else if |s| ≥ |κ|, note by Lemma E.5 that   1+y A(V )2 ≤ 1 + (2V + 1)2 ≤ 8V 2 + 3 ≤ 16 (κ − xs)2 · + z2y + 3 ρ 2 2 ≤ 32Csol (κ + s )(y + 1) + 16z 2 y + 3 . Thus as S 2 is Ksignal sub-Exponential and E[S 2 ] ≤ Ksignal , Z Z A(V )2 φ0 (z)φ1 (s) dzds s:|s|≥|κ| z Z Z Z φ0 (z)φ1 (s) dzds + (y + 1) s2 φ0 (z)φ1 (s) dzds ≤ 32Csol (y + 1)κ2 2 s:|s|≥|κ| z R Z + 16y z 2 φ0 (z)φ1 (s) dzds + 3 2 R   2 = 32Csol (y + 1)κ · P S 2 ≥ κ2 + Ksignal (y + 1) + 16y + 3  κ2  ≤ 32Csol (y + 1)κ2 · exp − + Ksignal (y + 1) + 16y + 3 Ksignal ≤ y(12Csol Ksignal + Ksignal + 16) + (12Csol Ksignal + Ksignal + 3) .   Summing the above bounds yields the desired upper bound on E A(V )2 . Lemma E.7. Suppose (α, β) ∈ [0, α0 ] × [β0 , β1 ]. For any (y, ρ) such that y ≥ 0, ρ ≥ C1sol that satisfies ∂G ∂y (y, ρ) = 0 – in particular, to any solution to (E.8) – we must have y ≤ 1. Proof. By our expression (E.13) for ∂G ∂y and by Lemma E.6, we obtain   1 1 y = α E A(V )2 ≤ α(KE.6 y + KE.6 ) ≤ y + =⇒ y ≤ 1 , 2 2 where we used that α ≤ 2K1E.6 . 124

The last preliminary step we need is the following Lemma, which lets us control various expectations of A(V ). Lemma E.8. Suppose (α, β) ∈ [0, α0 ] × [β0 , β1 ]. Then we have for any (y, ρ) ∈ [0, 1] × [ C1sol , Csol ] that       0 ≤ E A(V ) , E A(V )2 , E κ′ A(V ) ≤ KE.8 , where KE.8 =

( 3/2 42Csol Ksignal + 13Csol + 7.5

: κ < 0,

1/2 8.3Csol (κ2 + Ksignal ) + 2.3

: κ ≥ 0.

Proof. The desired lower bounds are obvious as A ≥ 0, and the first two upper bounds are  immedi ate by combining Lemma E.6 and Lemma E.7. As A ≥ 0, it remains to upper bound E |κ′ |A(V ) , which we do similarly as in the proof of Lemma E.6. We again break into the cases κ ≥ 0, κ < 0. When κ ≥ 0.

By Lemma E.5, we have A(x) ≤ 2|x| + 1 for all x ∈ R. We thus may write |κ′ |A(V ) ≤ 2|κ′ ||V | + |κ′ | r   1+y √ ′ 2 ≤ 2 |κ | + y|Z||κ′ | + |κ′ | ρ p  2|Z| + 1 ≤ 4(κ2 + x2 S 2 ) 2Csol + · 2(κ2 + x2 S 2 ) + 1 , 2 ′2

where we upper bounded |κ′ | by κ 2+1 . Consequently as S, Z are independent,   1/2 E κ′ A(V ) ≤ 5.7(κ2 + Ksignal )Csol + 2.6(κ2 + Ksignal ) + 2.3 ≤ KE.8 . When κ < 0.

We write Z  ′  E |κ |A(V ) = |κ′ |A(V )φ0 (z)φ1 (s) dzds s,z Z Z Z ′ = |κ |A(V )φ0 (z)φ1 (s) dzds + s:|s|<|κ|

z

s:|s|≥|κ|

Z

|κ′ |A(V )φ0 (z)φ1 (s) dzds .

z

Again, we upper bound each of these two integrals separately. • If |s| < |κ|, then as |xs| < |κ| and κ′ < 0, we have κ′ < 0. Define the sets r n κ′ 1 + y o √ E1 := (s, z) : z y ≥ , 2 ρ r r n 1+y κ′ 1 + y o √ ′ E2 := (s, z) : κ −1<z y < , ρ 2 ρ r n o 1+y √ ′ E3 := (s, z) : z y ≤ κ −1 . ρ Note as κ′ < 0 and y, ρ ≥ 0, E1 , E2 , E3 are all disjoint and partition R2 . We thus can write Z Z |κ′ |A(V )φ0 (z)φ1 (s) dzds = (I) + (II) + (III) , s:|s|<|κ|

z

125

where Z

|κ′ |A(V )φ0 (z)φ1 (s) dzds

(I) := (s,z)∈E1 :|s|<|κ|

Z

|κ′ |A(V )φ0 (z)φ1 (s) dzds

(II) := (s,z)∈E2 :|s|<|κ|

Z

|κ′ |A(V )φ0 (z)φ1 (s) dzds .

(III) := (s,z)∈E3 :|s|<|κ|

We now upper bound (I), (II), (III): 1. For (I): by definition of E1 , for (s, z) ∈ E1 we have r r κ′ 1 + y 1+y √ ′ V =κ −z y ≤ < 0. ρ 2 ρ Hence N (V ) ≥ 21 for (s, z) ∈ E1 , and moreover as κ′ < 0, we have for (s, z) ∈ E1 that V2 ≥

|κ′ |2 1 + y |κ′ |2 |κ′ |2 · ≥ ≥ . 4 ρ 4ρ 4Csol

Thus for (s, z) ∈ E1 , 2

 e−V /2 2 |κ′ |2  0 ≤ A(V ) ≤ √ ≤ √ exp − . 8Csol 2πN (V ) 2π Thus,  p 2|κ′ | |κ′ |2  √ exp − φ0 (z)φ1 (s) dzds ≤ Csol . 8Csol 2π (s,z)∈E1 :|s|<|κ|

Z (I) ≤

2. For (II): for (s, z) ∈ E2 , we have V = κ′

q

√ 1+y ρ − z y ≤ 1.

Therefore for (s, z) ∈ E2 , 1 A(V ) ≤ √2πN (1) . Furthermore, it is not possible in this case that y = 0, as then we have q ′ √ ′ 0 = z y < κ2 1+y ρ < 0 as κ < 0. Thus y > 0 and so κ′ z≤ 2

r

1+y |κ′ |2 1 + 1/y |κ′ |2 < 0 =⇒ z 2 ≥ · ≥ . yρ 4 ρ 4Csol

Therefore,  1 |κ′ |2  (II) ≤ |κ | · exp − exp(−z 2 /4)φ1 (s) dzds 2πN (1) 16C sol (s,z)∈E2 :|s|<|κ| Z p p 1 √ exp(−z 2 /4)φ1 (s) dzds ≤ 2.5 Csol . ≤ 1.72 Csol 2π R2 Z

3. For (III): by Lemma E.5 and Lemma E.7, r p 1+y √ ′ 0 ≤ A(V ) ≤ 2|V | + 1 ≤ 2|κ | + 2|z| y + 1 ≤ 3|κ′ | Csol + 2|z| + 1 . ρ 126

Moreover, here we have as κ′ ≤ 0 and ρ ≤ Csol that r 1+y 1 1+y 1 |κ′ |2 ′ . z≤κ − √ < 0 =⇒ z 2 ≥ |κ′ |2 · + ≥ yρ y yρ y Csol Therefore,  |κ′ |2  1 |κ′ |2 exp − · √ exp(−z 2 /4)φ1 (s) dzds 4Csol 2π (s,z)∈E3 :|s|<|κ| Z   ′ 2 |κ | 1 +2 |z||κ′ | exp − · √ exp(−z 2 /4)φ1 (s) dzds 4Csol 2π (s,z)∈E3 :|s|<|κ| Z   ′ 2 |κ | 1 + |κ′ | exp − · √ exp(−z 2 /4)φ1 (s) dzds 4Csol 2π (s,z)∈E3 :|s|<|κ|

Z p (III) ≤ 3 Csol

3/2

1/2

≤ 6.25Csol + 3.15Csol . Putting everything together and using that Csol ≥ 1, we obtain Z Z 3/2 |κ′ |A(V )φ0 (z)φ1 (s) dzds = (I) + (II) + (III) ≤ 13Csol . s:|s|<|κ|

z

• If |s| ≥ |κ|, we obtain by Lemma E.5 and Lemma E.7 that p 0 ≤ A(V ) ≤ 2V + 1 ≤ 2|V | + 1 ≤ 3|κ′ | Csol + 2|z| + 1 , thus by AM-GM, |κ′ |A(V ) ≤

 |κ′ |2 3 + 9|κ′ |2 Csol + 4|z|2 + 1 ≤ |κ′ |2 (13.5Csol + 0.5) + 6z 2 + 1.5 2 2 ≤ (κ2 + s2 )(27Csol + 1) + 6z 2 + 1.5 .

We obtain Z Z

A(V )2 φ0 (z)φ1 (s) dzds s:|s|≥|κ| z Z Z Z Z 2 φ0 (z)φ1 (s) dzds + (27Csol + 1) s2 φ0 (z)φ1 (s) dzds ≤ (27Csol + 1)κ s:|s|≥|κ| z s:|s|≥|κ| z Z + (6z 2 + 1.5)φ0 (z)φ1 (s) dzds s:|s|≥|κ|   2 2 2 ≤ (27Csol + 1)κ P S > κ + (27Csol + 1)Ksignal + 7.5  κ2  + (27Csol + 1)Ksignal + 7.5 ≤ (27Csol + 1)κ2 exp − Ksignal ≤ 1.5(27Csol + 1)Ksignal + 7.5 ,

where we used that S 2 is Ksignal sub-Exponential and E[S 2 ] ≤ Ksignal . Summing the upper bounds from both cases and using that Csol ≥ 1, we obtain   3/2 E |κ′ |A(V ) ≤ 42Csol Ksignal + 13Csol + 7.5 . This proves the Lemma in both cases κ < 0, κ ≥ 0. 127

Finishing the argument. By Lemma E.7, to show (E.8) has a unique solution in the desired domain, it suffices to show (E.8) has a unique solution in (y, ρ) ∈ [0, 1] × [ C1sol , Csol ]. Now, we consider the signs of the derivatives of G in y and ρ. This is done for the interpolators and posterior together in a unified manner. 2

Lemma E.9. Suppose (α, β) ∈ [0, α0 ]×[β0 , β1 ]. For (y, ρ) ∈ [0, 1]×[ C1sol , Csol ], we have ∂∂ 2G (y, ρ) < ρ 0. Proof. We first establish the desired result for the interpolators. Consider (E.14), which states i α(1 + y) h ∂2G 3α(1 + y)1/2 h ′ 1 E κ A(V ) − (y, ρ) = − E κ′2 A′ (V )] − 2 . 3 5/2 ∂2ρ 4ρ 2ρ 4ρ The second and third terms in this expression are non-positive as A′ ≥ 0 by Lemma E.5, and the third term is strictly negative. Finally, note by Lemma E.7 and Lemma E.8 and since α ≤ √ 2 , we have 3 2Csol KE.8 √ i 3α(1 + y)1/2 h ′ 3 2Csol α 1 − E κ A(V ) ≤ · KE.8 ≤ 2 . 2 5/2 4ρ 4ρ 4ρ It follows that

∂2G 1 1 1 ≤ 2 − 2 = − 2 < 0. 2 ∂ρ 4ρ 2ρ 2ρ For the posterior, we follow the same strategy. Recall (E.20) states that  ∂2G α 1 = (I) + (II) + (III) + (IV) − 2, 2 ∂ρ 4 2ρ

where (I), (II), (III), (IV) are as defined in (E.21). As 1/(1 + y), y/(1 + y) ∈ [0, 1] for y ∈ [0, 1], we have |(I)|, |(II)|, |(III)|, |(IV)| ≤ K(D). Since ρ ≤ Csol and α is small enough in terms of D, Csol , 2 it follows that ∂∂ρG2 < 0 for the posterior as well. Lemma E.10. Suppose (α, β) ∈ [0, α0 ] × [β0 , β1 ]. Then for any ρ′ ∈ [ C1sol , Csol ], we have the following: ′ ′ 1. There exists a unique y ′ ∈ [0, 1] such that ∂G ∂y (y , ρ ) = 0.

2. The one variable function y → G(y, ρ′ ) attains its minimum on [0, ∞) at y = y ′ . 2

3. We have ∂∂yG2 (y ′ , ρ′ ) > 0. 4. Defining y ′ := y ′ (ρ′ ) from 1), y ′ (ρ′ ) is continuous in ρ′ for ρ′ ∈ [ C1sol , Csol ]. Proof. We first prove 1), 2) of the Lemma. Recall g(y, ρ) defined in (E.15) for the interpolators or (E.22) for the posterior. In either case, g(y, ρ) is clearly jointly continuous in both arguments. We claim that to prove 1), 2) of the Lemma, it suffices to show that g(y, ρ′ ) = 0 at a unique y = y ′ , and that g(y, ρ′ ) < 0 for y < y ′ , g(y, ρ′ ) > 0 for y such that y ′ < y ≤ 1. ′ ′ To justify why, note by (E.15) or (E.22) that this implies ∂G ∂y (y, ρ ) = 0 at y = y , and that ∂G ′ ′ ∂G ′ ′ ′′ ∂y (y, ρ ) < 0 for y < y , ∂y (y, ρ ) > 0 for y < y ≤ 1. From here, we note that another y ∈ (1, ∞) ′′ ′ with ∂G ∂y (y , ρ ) = 0 would contradict Lemma E.7. This would imply 1). Moreover by continuity of ∂G ∂G ∂G ′′ ′ ′ ′ ′ ′′ ∂y (·, ρ ) and as ∂y (y, ρ ) > 0 for y < y ≤ 1, since there is no y ∈ (1, ∞) with ∂y (y , ρ ) = 0, this ′′ ′ ′′ would mean that ∂G ∂y (y , ρ ) > 0 for all y ∈ (1, ∞). Together with the above, this would imply 2). ∂g Now, we claim that ∂y (y, ρ′ ) > 0 for all y ∈ [0, 1], for both the interpolators and the posterior:

128

• For the interpolators: By (E.16), h i h i ∂g α =1− p E κ′ A(V )A′ (V ) − α E A′2 (V ) + A(V )A′′ (V ) . ∂y ρ(1 + y) By Lemma E.5 and Lemma E.8 we have h i 2 2 E A′2 (V ) + A(V )A′′ (V ) ≤ KMILLS + KMILLS E[A(V )] ≤ KMILLS + KMILLS KE.8 , i h h i 1 1/2 1/2 p E κ′ A(V )A′ (V ) ≤ Csol KMILLS E κ′ A(V ) ≤ Csol KMILLS KE.8 . ρ(1 + y)  1 Thus since α ≤ 14 min K 2 , 1/2 1 , it follows that +K K MILLS

MILLS

E.8

Csol KMILLS KE.8

h i h i 1 ∂g α =1− p E κ′ A(V )A′ (V ) − α E A′2 (V ) + A(V )A′′ (V ) ≥ > 0 . ∂y 2 ρ(1 + y)

(E.25)

• For the posterior: Recall (E.23), that y ≥ 0 and |ul | ≤ D for l = 1, 2, 3, 4, and that ρ′ ≤ Csol , 1 1+y ≤ 0. Thus since α ≤ α0 and α0 is small enough in terms of D, Csol ,  1 αρ  ∂g = 1 + 2y − (I) + (II) ≥ , ∂y (1 + y)2 2

(E.26)

where (I), (II) are as in (E.24). We have |(I)|, |(II)| ≤ K(D) as |ul | ≤ D for l = 1, 2, 3, 4. Note that the bounds (E.25), (E.26) hold uniformly in (y, ρ) ∈ [0, 1] × [ C1sol , Csol ]. This implies there is at most one y ′ ∈ [0, 1] such that g(y ′ , ρ′ ) = 0. Moreover, if this y ′ exists, our above work clearly implies g(y) < 0 for y < y ′ , g(y) > 0 for y > y ′ . Now we show there exists such an y ′ ∈ [0, 1]. When y = 0, evidently g(y) ≤ 0. Now consider when y = 1. For the interpolators, note as α < K1E.8 , we have by Proposition E.8 that g(1) = 1 − α E[A(V )2 ] > 0. For the posterior, since |u′ | ≤ D, ρ′ ≤ Csol , and α is small enough in terms of D, Csol , we have g(1) = 2 − αρ E exp u′

2

> 0.

Thus for both the interpolators and the posterior, there exists y ′ ∈ [0, 1] with g(y ′ , ρ′ ) = 0 by the Intermediate Value Theorem. The above steps now prove that g(y, ρ′ ) = 0 at a unique y = y ′ , and that g(y, ρ′ ) < 0 for y < y ′ , g(y, ρ′ ) > 0 for y such that y ′ < y ≤ 1, which as discussed earlier implies 1), 2) of the Lemma. Next, we prove 3) of the Lemma for both the interpolators and the posterior: • For the interpolators: Differentiating (E.15), ∂g

∂2G ∂y (y, ρ )(1 + y) − g(y, ρ ) (y, ρ′ ) = . 2 ∂y 2(1 + y)2 Observe g(y, ρ′ ) ≤ y. Also from the above steps, we have  2 ∂g 1/2 (y, ρ′ ) ≥ 1 − α max KMILLS + KMILLS KE.8 , Csol KMILLS KE.8 . ∂y 129

It therefore remains to show  2 1/2 1 − α max KMILLS + KMILLS KE.8 , Csol KMILLS KE.8 >

y′ . y′ + 1

y Since y+1 is increasing and as y ′ ≤ 1 by 1), it suffices to show

 2 2 1/2 + KMILLS KE.8 , Csol KMILLS KE.8 > , 1 − α max KMILLS 3 which holds true by our condition on α. • For the posterior: Differentiating (E.22) and using (E.23),  2 αρ ∂g ′ 1 + y + 2αρ E exp u′ − (1+y) 2 (I) + (II) ∂2G ∂y (y, ρ )(1 + y) − 2g(y) ′ (y, ρ ) = = , ∂y 2 2(1 + y)3 2(1 + y)3 where (I), (II) are as in (E.24). Note |(I)|, |(II)| ≤ K(D) as |ul | ≤ D for l = 1, 2, 3, 4. Thus 2 as y ≥ 0, ρ ≤ Csol , and α ≤ α0 is small enough in terms of D, Csol , we have ∂∂yG2 (y ′ , ρ′ ) > 0. This proves 3) of the Lemma in both cases. Finally, we prove 4) of the Lemma. Note that y(ρ′ ) is a solution to the fixed point equation ∂g g(y(ρ′ ), ρ′ ) = 0. As ∂y > 0 holds uniformly on [0, 1] × [ C1sol , Csol ] as argued above, and as y(ρ′ ) ∈ [0, 1] as shown in 1), we conclude 4) by the Implicit Function Theorem. The above Lemmas establish the desired convex-concave structure of the RS equations, and we are now ready to complete the proof of Proposition A.3. Lemma E.11. Consider (α, β) ∈ [0, α0 ] × [β0 , β1 ]. Then there is at most one solution (y, ρ) ∈ [0, 1] × [ C1sol , Csol ] to (E.8). Proof. Suppose that there are two distinct solutions (y1 , ρ1 ) and (y2 , ρ2 ) both in the domain [0, 1] × [ C1sol , Csol ]. Assume without loss of generality that y1 < y2 . Directly from Lemma E.9 and Lemma E.10, we have y1 ̸= y2 and ρ1 ̸= ρ2 . Applying Taylor Expansion to degree 2 about ∂2G (y1 , ρ1 ), since ∂G ∂ρ (y1 , ρ1 ) = 0 by definition of the system of equations (E.8) and since ∂ρ2 (y1 , ·) < 0 from Lemma E.9, we see that G(y1 , ρ1 ) > G(y1 , ρ2 ) (since we Taylor expand about a ρ′ ∈ [min{ρ1 , ρ2 }, max{ρ1 , ρ2 }] ⊆ [ C1sol , Csol ], we may apply the above work). By Lemma E.10, since ∂G ∂y (y2 , ρ2 ) = 0, we have G(y1 , ρ2 ) > G(y2 , ρ2 ). By analogous reasoning again using Lemma E.9, we have G(y2 , ρ2 ) > G(y2 , ρ1 ). Again by Lemma E.10, we have G(y2 , ρ1 ) > G(y1 , ρ1 ). Combining these inequalities, we obtain G(y1 , ρ1 ) > G(y1 , ρ2 ) > G(y2 , ρ2 ) > G(y2 , ρ1 ) > G(y1 , ρ1 ) , a contradiction. Lemma E.12. Consider (α, β) ∈ [0, α0 ] × [β0 , β1 ]. Then there exists a solution (y, ρ) ∈ [0, 1] × [ C1sol , Csol ] to (E.8). Proof. By Lemma E.10 and Lemma E.7, for any ρ ∈ [ C1sol , Csol ], there exists a unique y(ρ) ∈ 1 ′ [0, 1] where ∂G ∂y (y(ρ), ρ) = 0. Thus it is enough to show there exists a ρ ∈ [ Csol , Csol ] such that ∂G ′ ′ ∂ρ (y(ρ ), ρ ) = 0. Let h(ρ) := 2 ∂G ∂ρ (y(ρ), ρ). By 4) of Lemma E.10 and using the explicit form (E.11), (E.18) for ∂G 1 ∂ρ for the interpolators and posterior respectively, h(ρ) varies continuously in ρ ∈ [ Csol , Csol ]. We now prove that h(1.5/β) < 0, h(1/3.5β) > 0 for both the interpolators and the posterior. 130

• For the interpolators: Observe that because y(ρ) ≤ 1 by Lemma E.7, Lemma E.8 yields i (1 + y(ρ))1/2 h ′ 3/2 E κ A(V ) ≤ 21/2 Csol KE.8 . 3/2 ρ Now note by (E.11) and as α <

0.9β0 , we have 3/2 Csol KE.8

i 2  1.5  α(1 + y(ρ))1/2 h 2 3/2 ′ E κ A(V ) + β − 2β < α21/2 Csol KE.8 + β − 2β < 0 . h = 3/2 β 3 3 ρ Similarly we obtain  1  4 h ≥ − β + 3.5β − 2β > 0 . 3.5β 3 • For the posterior: Again y(ρ) ≤ 1 by Lemma E.7. Thus as |u′ |, |u′′ | ≤ D, E u′′ + u′ −

y 2 E u′ ≤ K(D) . 1+y

Thus as α is small enough in terms of β0 , β1 , D, Csol , by (E.18),  1   1.5  2 ≤ αK(D) + β − 2β < 0 , h ≥ −αK(D) + 3.5β − 2β > 0 . h β 3 3.5β 1 Thus by Intermediate Value Theorem and continuity of h(ρ), there is some ρ′ ∈ [ 3.5β , 1.5β] such ∂G 1 1.5 ′ ′ ′ that h(ρ ) = 0, that is, ∂ρ (y(ρ ), ρ ) = 0. By our definition of Csol , we have [ 3.5β , β ] ⊂ [ C1sol , Csol ] for all β ∈ [β0 , β1 ]. Hence (y(ρ′ ), ρ) is a solution in [0, 1] × [ C1sol , Csol ] to (E.8), and this proves the Lemma.

Finally, recall that for (α, β) ∈ [0, α0 ] × [β0 , β1 ], if there is a solution (y, ρ) ∈ [0, ∞) × [ C1sol , Csol ], then this solution must be in [0, 1] × [ C1sol , Csol ] by Lemma E.7. We conclude that (E.8) has a unique solution (y, ρ) ∈ [0, ∞) × [ C1sol , Csol ] (in particular, this solution must be in [0, 1] × [ C1sol , Csol ]). By our initial remarks in this proof, it follows that (A.11) have a unique solution (q, ρ) ∈ [0, Csol ] × [ C1sol , Csol ], as desired. Finally, as this unique solution (y, ρ) ∈ [0, 1] × [ C1sol , Csol ], ρ we have ρ − q = y+1 ≥ 2C1sol . This proves the first part of Proposition A.3.  2 We now prove the second part of Proposition A.3. First, recall ∂∂ 2G y0 (α, β, x), ρ0 (α, β, x) and y  ∂2G y (α, β, x), ρ (α, β, x) are of opposite signs and nonzero by Lemma E.9 and Lemma E.10. 0 0 2 ∂ ρ Thus the determinant of the following Jacobian is strictly negative: " 2 #  ∂2G ∂ G y (α, β, x), ρ (α, β, x) y (α, β, x), ρ (α, β, x) 0 0 0 0 2 y  ∂y∂ρ  <0. det ∂∂2 G ∂2G y (α, β, x), ρ (α, β, x) y0 (α, β, x), ρ0 (α, β, x) 0 0 ∂y∂ρ ∂2ρ  Since G is infinitely differentiable in x, y, ρ, and as y0 (α, β, x), ρ0 (α, β, x) denotes the unique ∂G solution (y, ρ) ∈ [0, ∞) × [ C1sol , Csol ] to the system ∂G ∂y = ∂ρ = 0, it follows by the Implicit Function Theorem that y0 (α, β, x), ρ0 (α, β, x) are infinitely differentiable in α, β and x. Hence as y0 (α, β, x) ≥ 0, it follows that q0 (α, β, x) =

y0 (α, β, x)ρ0 (α, β, x) 1 + y0 (α, β, x)

is also infinitely differentiable in α, β and x. 131

E.3

Proof of Lemma C.3

For this proof, we make the same transformation to G(y, ρ) we considered in Appendix E.2, where G(y, ρ) is defined as in (E.7) (again, here I is held fixed so F ≡ FI ). We let q0 = q0 (α, β, x), 0 . ρ0 = ρ0 (α, β, x), and y0 = y0 (α, β, x) = ρ0q−q 0 ∂G We first prove 1) of the Lemma. Differentiating the relations ∂G ∂ρ (y0 , ρ0 ) ≡ 0, ∂y (y0 , ρ0 ) ≡ 0 with respect to β (recall y0 , ρ0 depend on β), the Chain Rule gives

∂y0 ∂ 2 G ∂ρ0 ∂2G ∂2G (y0 , ρ0 ) · + (y , ρ ) · + (y0 , ρ0 ) = 0 , 0 0 ∂ρ∂y ∂β ∂ρ2 ∂β ∂β∂ρ ∂2G ∂y0 ∂2G ∂ρ0 ∂2G (y , ρ ) · + (y , ρ ) · + (y0 , ρ0 ) = 0 , 0 0 0 0 ∂y 2 ∂β ∂ρ∂y ∂β ∂β∂y 2

∂ G where differentiability is justified by Proposition A.3. Now identically as functions, we have ∂β∂ρ = 2

∂ G 0 −1, ∂β∂y = 0 by the definition of G in (E.7). Multiplying the first relation above by ∂ρ ∂β and the 2

∂ G 0 second relation above by ∂y ∂β and subtracting, the ∂ρ∂y terms cancel and we obtain

 ∂y 2 ∂ 2 G ∂ρ0  ∂ρ0 2 ∂ 2 G 0 = (y , ρ ) − (y0 , ρ0 ) < 0 , · · 0 0 2 ∂β ∂β ∂ρ ∂β ∂y 2 2

2

where Lemma E.9 and Lemma E.10 justify that ∂∂ρG2 (y0 , ρ0 ) < 0, ∂∂yG2 (y0 , ρ0 ) > 0. This proves 1) of the Lemma. 2

2

Remark 3. Note that we only have ∂∂ρG2 (y0 , ρ0 ) < 0, ∂∂yG2 (y0 , ρ0 ) > 0 at the solution y0 , ρ0 . We next prove 2) of the Lemma. To this end, we will show that ρ0 (α, 1/4, x) > 1 − x2 , ρ0 (α, 9, x) < 1 − x2 . The existence and uniqueness of such a β = β(α, x) ∈ [1/4, 9] ⊆ [β0 , β1 ] then follows by part 1) of this Lemma, and moreover this establishes that β(α, x) ∈ (1/4, 9). We first show that ρ0 (α, 1/4, x) > 1 ≥ 1 − x2 . Define  h i E κ′ (1+y0 )1/2 A(V ) : for interpolators , 3/2 ρ0 (I) := (E.27) E u′′ (V̄ ) + u′ (V̄ )2 − y E u′ (V̄ ) 2 : for posterior , 1+y where V is as in (E.9) and V̄ is as in (E.17), taken with arguments ρ0 , y0 . By (E.11) for the interpolators and (E.18) for the posterior, we have 0=

α 1 (I) + −β. 2 2ρ0

(E.28)

Now note (√ |(I)| ≤

1/2 2 ρ0 KE.8 Csol 2D2 + D

: for interpolators , : for posterior .

(E.29)

For the posterior, this simply uses that |u′ |, |u′′ | ≤ D. For the interpolators, we use the proof of Proposition A.3, that ρ0 ≥ C1sol , y0 ≥ 0, and that y0 ≤ 1 by Lemma E.7. These facts combined with Lemma E.8 imply the above bound on (I) . 132

Now as α ≤ α0 and by our conditions on α0 , we have by (E.29) that α|(I)|/2 < 4ρ10 . Combining with (E.28) now establishes  1 3  1 3 β∈ , =⇒ < ρ0 (α, β, x) < =⇒ ρ0 (α, 1/4, x) > 1 ≥ 1 − x2 . (E.30) 4ρ0 4ρ0 4β 4β Next, we claim that ρ0 (α, 9, x) < 1 − x2 for x ∈ [−1 + δ, 1 − δ]. To this end, (E.28), (E.29) imply that α 1 + |(I)| =⇒ ρ0 (α, 9, x) < 1 − x2 . (E.31) 9≤ 2ρ0 (α, 9, x) 2 Here, the last inequality holds for x ∈ [−1 + δ, 1 − δ], using the definition of δ in (C.1) for the interpolators or (C.6) for the posterior. As mentioned earlier, (E.30), (E.31) yields 2) of the Lemma. Now to prove 3) of the Lemma, for fixed (α, x), it follows from Proposition A.3 that f (β) is ∂FI ∂F ∂FI infinitely differentiable in β. Let F (x, q, ρ) = Φ̄(x, q, ρ) − β(ρ + x2 ). Note ∂F ∂ρ = ∂ρ , ∂q = ∂q , ∂F thus ∂F ∂ρ (x, q0 , ρ0 ) = ∂q (x, q0 , ρ0 ) = 0. Therefore,  ∂F  d ∂ρ0 ∂F ∂q0 f (β) = 1 − ρ0 + x2 + (x, q0 , ρ0 ) · + (x, q0 , ρ0 ) · = 1 − ρ0 + x2 . dβ ∂ρ ∂β ∂q ∂β Thus d2 ∂ρ0 (α, β, x) f (β) = − . 2 dβ ∂β  d2 d By 1) of this Lemma, dβ 2 f (β) > 0 on [β0 , β1 ], and by 2) of this Lemma, dβ f β(α, x) = 0. Hence 3) of the Lemma follows. Next, we prove 4) of the Lemma. Since ρ0 α, β(α, x), x) = 1 − x2 , we have    f β(α, x) = Φ̄ x, q0 α, β(α, x), x , 1 − x2 .   1 Note as x ∈ [−1 + δ, 1 − δ] ⊆ [−17/18, 17/18], 1 ≥ 1 − x2 ≥ 10 . Thus 1 − x2 ∈ C1sol , Csol as Csol ≥ 10. Recalling the definition of Φ̄ in (5.4), define the following function of y:   yρ Φ̄′ (x, y, ρ) := Φ̄ x, ,ρ . y+1     Φ̄′ I Note ∂∂y α, β(α, x), x = 0. Now as a direct y0 α, β(α, x), x , 1 − x2 = ∂G y α, β(α, x), x , ρ 0 0 ∂y ′ consequence of the proof of Lemma E.10 and by 2) of this Lemma, Φ̄ (x, y, 1 − x2 ) attains its unique yρ minimum on [0, ∞) at y = y0 α, β(α, x), x . By monotonicity of the transformation q = y+1 , it  2 2 follows that Φ̄(x, q, 1 − x ) attains a unique minimum on [0, 1 − x ) at q = q0 α, β(α, x), x . Thus  f β(α, x) = inf Φ̄(x, q, 1 − x2 ) = inf Φ(x, q) , q∈[0,1−x2 )

q∈[0,1)

q where the last step follows from the transformation q ← 1−x 2 . This proves 4) of the Lemma.

Finally, to prove 5) of the Lemma, we claim that β(α, x) is a continuous function of (α, x). To this end, note by (E.11) and 2), we know that β(α, x) is the unique solution in [β0 , β1 ] to β=

α 1 (I) + , 2 2(1 − x2 ) 133

(E.32)

where (I) is taken with ρ0 = 1 − x2 . Let f¯(α, β, x) denote the right hand side of (E.32); this is continuous in (α, β, x) for (α, β, x) ∈ [0, α0 ] × [β0 , β1 ] × [−1 + δ, 1 − δ] by Proposition A.3. Consider any sequence (αn , xn ) → (α⋆ , x⋆ ) where αn ∈ [0, α0 ], xn ∈ [−1 + δ, 1 − δ]. Consider the corresponding βn := β(αn , xn ). Since βn ∈ [β0 , β1 ], it follows that we can extract a convergent subsequence. Consider any convergent subsequence βnk , with limit β ⋆ . By the above it follows by compactness that β ⋆ ∈ [β0 , β1 ]. As f¯(α, β, x) is continuous, we have β ⋆ = lim βnk = lim f¯(αnk , βnk , xnk ) = f¯(α⋆ , β ⋆ , x⋆ ) . k→∞

k→∞

However as (α⋆ , x⋆ ) ∈ [0, α0 ] × [−1 + δ, 1 − δ], the uniqueness of β(α, x) in [β0 , β1 ] supplied  by 2) ⋆ ⋆ ⋆ ⋆ implies that β = β(α , x ). Therefore, limk→∞ β(αnk , xnk ) = β = β limk→∞ (αnk , xnk ) . Since the convergent subsequence βnk is arbitrary, it follows that βn converges to β(α⋆ , x⋆ ). Continuity of β(α, x) follows. Combining this claim with 4) and Proposition A.3 now proves 5). This concludes the proof of Lemma C.3.

E.4

Proof of Propositions 3.2, 3.3, 3.6, 3.7, and Lemma 3.8

Here we first prove these four Propositions and then prove Lemma 3.8. We will prove these four Propositions together in a unified manner as follows. We show the following: ∂ ∂ 1. There is at most one pair (x, q) satisfying ∂x Φ(x, q) = ∂q Φ(x, q) = 0.

2. This implies that the saddle point problem is uniquely achieved. ∂ ∂ Proof of Part 1. First, we show that any (x, q) satisfying ∂x Φ(x, q) = 0, ∂q Φ(x, q) = 0 must satisfy the relevant bivariate system of equations given in Propositions 3.2, 3.3, 3.6, 3.7 respec∂ ∂ tively. The proof follows from explicitly computing ∂x Φ(x, q) and ∂q Φ(x, q) and applying Gaussian Integration by Parts to simplify the resulting expression. We discuss the proof for Propositions 3.2, 3.3 for the interpolators; the proof for Propositions 3.6, 3.7 for the posterior is identical. Recalling the definition of S in (3.5) in the GMM case and (3.6) in the logistic case, we let p κ − xS − (1 − x2 )qZ p V := . (1 − x2 )(1 − q)

Explicitly differentiating Φ(x, q) w.r.t. q and simplifying with Guassian Integration by Parts and the identity A′ (V ) = A(V )2 − V A(V ) gives h n oi q κ − xS Z = α E A(V ) p −p 1−q q(1 − q) (1 − q)(1 − x2 ) r h n oi 1−q = α E A(V ) V − Z q     = α E A(V )V + A′ (V ) = α E A(V )2 . Therefore, such (x, q) satisfying the stationary point equations must satisfy (3.14) in the GMM case or (3.16) in the logistic case. We will now show the proof that such (x, q) satisfying the stationary point equations must satisfy (3.15) in the logistic case. The proof for (3.13) in the GMM case is similar (in fact simpler). We differentiate Φ(x, q) w.r.t. x to obtain h h p i i x (1 − q)(1 − x2 ) = α E A(V ) S − κx = α E A(V ) Y G − κx , (E.33) 134

√ where we used the explicit form S = Y G from (3.6) (for the GMM case, we use S = G + λ as per (3.5)). We now simplify E A(V )Y G via Gaussian Integration by Parts to obtain h i h i √   √  x E A′ (V ) + λ E φ − λY G A(V ) . E A(V )Y G = − p (E.34) (1 − x2 )(1 − q) Similar calculations, using the identity A′ (V ) = A(V )2 − V A(V ), yield p       E A′ (V ) = (1 − x2 )(1 − q) E A(V )2 − κ (1 − x2 )(1 − q) E A(V ) h i p √ (E.35)  + x λ(1 − x2 )(1 − q) E φ − λY G A(V ) .   q Combining (E.33), (E.34), (E.35) and using the identity α E A(V )2 = 1−q and rearranging yields (3.15). For Proposition 3.6, this establishes (3.24), which directly implies (3.25) and proves 1). For Proposition 3.7, 1) now follows from Proposition 3.8. We now prove 1) for Propositions 3.2, 3.3 for the interpolators. ∂ ∂ To this end, our next step is to bound the value of x, q for solutions to ∂x Φ(x, q) = 0, ∂q Φ(x, q) = 0. Specifically, we will prove the following Lemma. ∂ ∂ Lemma E.13. For any solution (x, q) to ∂x Φ(x, q) = 0, ∂q Φ(x, q) = 0, we have

0 ≤ x, q ≤ αK(λ) .

(E.36)

To prove Lemma E.13, we will need the following useful Lemma. Its proof is analogous to that of Lemma E.6 and we omit it for brevity. Lemma E.14. We have for all 0 ≤ x < 1, 0 ≤ q < 1 that r  p   κ+ + 1 q E A(V )p ≤ K p + +1 1−q (1 − x2 )(1 − q) for p ∈ {1, 2}, where K > 0 depends on λ. Proof of Lemma E.13. We now complete the proof of Lemma E.13. We start with the bound on x. We apply Lemma E.14 with p = 1. First, note x ≥ 0 as when x < 0, the LHS of (3.15), (3.13) is strictly negative while the RHS is non-negative. Noting that φ ∈ [0, 1] and A ≥ 0, we obtain r   p p   κ+ + 1 q x √ ≤ α λ(1 − q) E A(V ) ≤ α λ(1 − q) · K(λ) p + +1 . 1−q 1 − x2 (1 − x2 )(1 − q) √ As α ≤ α0 (λ, κ+ ), we now use 0 ≤ q ≤ 1 and multiply both sides by 1 − x2 to obtain the bound in (E.36) for x. Note as α ≤ α0 (λ, κ+ ), we have x ≤ 21 . We now prove the desired bound in (E.36) for q. By Lemma E.14 and (3.14) for the GMM case or (3.16) for the logistic case, we have    κ2 + 1  κ2+ + 1 q q q q + ≤ αK(λ) + + 1 =⇒ ≤ + αK(λ) + 1 , 1−q (1 − x2 )(1 − q) 1 − q 1−q 2(1 − q) 1−q where we used α ≤ α0 (λ, κ+ ) and x ≤ 12 . The above display now implies (E.36) for q. We now are in a position to finish the proof of 1). Let f1 (x, q) denote the LHS minus the RHS of (3.13) or (3.15) in the GMM case and the logistic case respectively. We similarly let f2 (x, q) denote the LHS minus the RHS of (3.14) or (3.16) in these two respective cases. We now establish the following Lemma E.15, which shows that all principal minors of the Jacobian of f1 (x, q), f2 (x, q) are strictly positive in the domain (x, q) ∈ [0, αK(λ)]2 . Since any solution (x, q) must lie in [0, αK(λ)]2 by Lemma E.13, injectivity of f1 (x, q), f2 (x, q) on (x, q) ∈ [0, αK(λ)]2 and therefore 1) now follows from the Gale-Nikaido Theorem [GN65]. 135

Lemma E.15. We have for all (x, q) ∈ [0, αK(λ)]2 , ∂ ∂ ∂ ∂ ∂ ∂ f1 (x, q) , f2 (x, q) , f1 (x, q) · f2 (x, q) − f1 (x, q) · f2 (x, q) > 0 . ∂x ∂q ∂x ∂q ∂q ∂x Proof of Lemma E.15. Lemma E.13 implies x, q ≤ 1/2, thus we have ∂ x ∂ x ∂ q q ∂ √ , , ≥ c > 0, , ∂x 1 − x2 ∂x 1 − x2 ∂q 1 − q ∂q (1 − q)2 where c > 0 is a universal constant. Since x, q ≤ 1/2, α ≤ α0 (λ, κ+ ), and φ ∈ [0, 1], it suffices to prove h ∂V i h ∂V i h ′ ∂V i h ∂V i E A(V )A′ (V ) , E A(V )A′ (V ) , E A (V ) , E A(V ) + A′ (V ) ≤ K(λ, κ+ ) . ∂x ∂q ∂x ∂q   q We note that the corresponding ρ = 1 − x2 ∈ C1sol , Csol and y = ρ−q = 1−xq2 −q ≤ 1 as (x, q) ∈   )  ≤ K(λ, κ+ ). [0, αK(λ)]2 . Therefore the conditions of Lemma E.8 apply, giving E (κ − Sx)A(V  Upon using Cauchy-Schwarz together with Lemma E.14 to upper bound E SA(V ) , this implies   that E κA(V ) ≤ K(λ, κ+ ). The upper bounds on the first two quantities above now follow as |A′ | ≤ 1, 0 ≤ x, q ≤ 21 , using the explicit forms of ∂V /∂x, ∂V /∂q, and applying the Cauchy-Schwarz Inequality together with Lemma E.14. To prove bound the third and fourth quantities above, as A′ (V ) = A(V )2 − V A(V ), we have |A′ (V )| ≤ 3|V ||A(V )| + |A(V )| as A(V ) ≤ 2|V | +1. Now, following the exact same proof as that  of Lemma E.8 and since 0 ≤ x, q ≤ 12 , we have E κ′2 A(V ) ≤ K(λ, Cauchy-Schwarz  2κ+ ). Using  and Lemma E.14 to upper bound E S 2 |A(V )| , it follows that E |κ| |A(V )| ≤ K(λ, κ+ ). Again  ′   ′  applying Cauchy-Schwarz, this implies E A (V )∂V /∂x , E A (V )∂V /∂q ≤ K(λ, κ+ ). Lemma E.15 now follows. Proof of Part 2. We now complete the proof of Propositions 3.2, 3.3, 3.6, 3.7. Define the function f (x) := inf q∈[0,1) Φ(x, q). Note f (x) is the infimum of an arbitrary family of continuous functions, and therefore is upper semicontinuous. First notice for the posterior, as u(x) ≥ −D(λ)(|x| + 1), we have   p p E log EW exp u xS + (1 − x2 )qZ + (1 − x2 )(1 − q)W h i (E.37) p p ≥ −D(λ) E 1 + xS + (1 − x2 )qZ + (1 − x2 )(1 − q)W ≥ −K(λ) . For the interpolators, using that A(x) ≤ 2x + 1 for all x ≥ 0 by Lemma E.5 and that − log N (V ) ≤ log 2 when V ≤ 0, we obtain − log N (V ) ≤ K(V+2 + 1) for all V ≥ 0 where V+ = max{V, 0}. Since (a + b)+ ≤ b+ , it follows that E log N (V ) ≥ −K ·

κ2+ + x2 S 2 + (1 − x2 )qZ 2 . (1 − x2 )(1 − q)

(E.38)

Thus for α ≤ α0 (λ, κ+ ), we have for both the posterior and the interpolators, Φ(x, q) ≥

αK(λ)(κ2+ + 1) 1 q 1 − + log(1 − x2 ) + log(1 − q) − αK(λ) 2 4(1 − q) (1 − x )(1 − q) 2 2

(E.39)

Taking x = 0 in the above display, bounding αK(λ)(κ2+ + 1) ≤ 1/8, and noting the infimum of the resulting lower bound is attained at q = 1/2 now proves that f (0) ≥ −K(λ). 136

Next, notice f (x) ≤ Φ(x, 0). Note for the interpolators or the posterior in the logistic case, we have u ≤ 0 and therefore Φ(x, 0) ≤ 12 log(1 − x2 ). For the posterior in the GMM case, we have √ u(x) = λx and therefore   p p E log EW exp u xS + (1 − x2 )qZ + (1 − x2 )(1 − q)W  (E.40) p p √  ≤ log E exp λ xS + (1 − x2 )qZ + (1 − x2 )(1 − q)W ≤ K(λ) . Thus in all cases, we have limx→±1 Φ(x, 0) = −∞, and justifying the existence of the following limit and the equality lim f (x) = −∞ .

x→±1

The above derivation also proves that f (0) ≤ Φ(0, 0) ≤ K(λ). Now as f (0) ≥ −K(λ) and limx→±1 f (x) = −∞, by the upper semicontinuity of f and as f (0) ≤ K(λ), the Weierstrass Extreme Value theorem implies that supx∈(−1,1) f (x) is attained by at least one x⋆ ∈ [−1, 1] and that all such x⋆ must lie in [−1 + δ(λ), 1 − δ(λ)]. Consider any such x⋆ . As x⋆ ∈ [−1 + δ(λ), 1 − δ(λ)] and Φ(x⋆ , q) is continuous in q, (E.39) implies that inf q∈[0,1) Φ(x⋆ , q) is attained at q ∈ [0, 1 − δ ′ (λ)], since the RHS of (E.39) goes to +∞ as q → 1 and as α ≤ α0 (λ, κ+ ). Thus supx∈(−1,1) inf q∈[0,1) Φ(x, q) is attained by at least one (x⋆ , q ⋆ ). Now consider any x ∈ (−1, 1). For the interpolators, we explicitly compute using Gaussian Integration by Parts,  ∂ α  Φ(x, 0) = − E A(V )2 < 0 . ∂q 2 For the posterior, a similar calculation gives ∂ α(1 − x2 ) h EW [u′ · exp u] 2 i Φ(x, 0) = − E < 0. ∂q 2 EW [exp u] It follows that all such q ⋆ > 0, and therefore any such (x⋆ , q ⋆ ) ∈ (−1, 1) × (0, 1). Hence any such ∂ ∂ (x⋆ , q ⋆ ) solves the system ∂x Φ(x, q) = 0, ∂q Φ(x, q) = 0. By 1), it follows that (x⋆ , q ⋆ ) is unique and satisfies (x⋆ , q⋆ ) ∈ [0, K(λ)α]2 – for the posterior these bounds follow directly (GMM) or by Lemma 3.8 (logistic). This completes the proof.

Proof of Lemma 3.8. Let G ∼ N (0, 1) ,

√  Y |G ∼ Rad φ( λG) ,

√ η := φ(− λY G) .

(E.41)

Recall the system of equations √    αλ(1 − q) E φ − λY G R(x, q) =

  x q 2 2 , and αλ(1 − x ) E R(x, q) = , (E.42) 1 − x2 (1 − q)2

where √ p p EW φ′ ( λV ) √ R(x, q) := , V := xY G + (1 − x2 )qZ + (1 − x2 )(1 − q)W , EW φ( λV ) 137

(E.43)

and G, Z, W ∼ N (0, 1) are mutually independent. We wish to show that for α small enough this system has a unique solution which takes the form (x, q) = (s/(1 + s), s/(1 + 2s)). Consider the parametrized curve s s xs := , qs := , s ≥ 0, 1+s 1 + 2s which observes the Nishimori identity qs = xs , 1 − qs and 1 − x2s =

1 + 2s , (1 + s)2

√ p s 2 (1 − xs )qs = , 1+s

We consider the auxiliary observation model √ Ḡ = s G + W ,

p 1 (1 − x2s )(1 − qs ) = √ . 1+s

(E.44)

W ∼ N (0, 1) ,

with W independent of (G, Y ). Conditionally on Ḡ, the posterior law of G before observing Y is √ s 1 Gaussian with mean 1+s Ḡ and variance 1+s . Since √ y ∈ {−1, +1} , P(Y = y | G) = φ( λyG) , Bayes’ formula gives √ √  s 1 EW φ′ λY 1+s Ḡ + √1+s W √ √  = R(xs , qs ) . E[η | Ḡ, Y ] = s 1 EW φ λY 1+s Ḡ + √1+s W

(E.45)

Here we used the identity φ(t)φ(−t) = φ′ (t) and symmetry of the Gaussian distribution, together with the equalities displayed in (E.44). Consequently by the tower property,   E[ηR(xs , qs )] = E E[η | Ḡ, Y ]2 = E[R(xs , qs )2 ]. (E.46) Now define C(s) := E[R(xs , qs )2 ] . Using (E.46), the system of equations (E.42) restricted to the curve s 7→ (xs , qs ) become αλ(1 − qs )C(s) =

xs , 1 − x2s

and

αλ(1 − x2s )C(s) =

qs . (1 − qs )2

Both are equivalent to s = αλC(s) .

(E.47)

We show that for α small enough, there is a unique solution which must lie on this curve. Let B(x, q) := E[R(x, q)2 ] .

A(x, q) := E[ηR(x, q)] ,

Since 0 < φ < 1 and 0 < φ′ < φ, A, B, C ∈ (0, 1). Therefore any solution of (E.42) satisfies x ≤ αλ , 1 − x2

q ≤ αλ . (1 − q)2

In particular, x ≤ αλ ,

q ≤ αλ . 138

(E.48)

Thus every solution lies in an arbitrarily small neighborhood of (0, 0) once α is sufficiently small. We now use local uniqueness. Let √ 2x 1 + 4q − 1 √ H(x) := , K(q) := √ , 2 1 + 4q + 1 1 + 1 + 4x the inverses of x 7→ x/(1 − x2 ) on [0, 1) and q 7→ q/(1 − q)2 on [0, 1), respectively. Hence the system is equivalent to the fixed-point equation (x, q) = Tα (x, q) , where  Tα (x, q) := H(αλ(1 − q)A(x, q)) , K αλ(1 − x2 )B(x, q) .

(E.49)

The functions A and B are locally Lipschitz in a neighborhood of (0, 0). Indeed, after conditioning on (G, Y ), the remaining average (with respect to Z, Eq. (E.43)) is a Gaussian heat semigroup applied to the smooth function √ EW φ′ ( λ(u + σW )) √ u 7→ , EW φ( λ(u + σW )) (and its square) with σ bounded away from 0 near (x, q) = (0, 0). Differentiation under the Gaussian integral is justified by the boundedness and exponential decay properties of the derivatives of the logistic function. Hence there exist δ > 0 and L < ∞, depending only on λ, such that A and B are L-Lipschitz on [0, δ]2 . Since H and K have bounded derivatives on bounded intervals, it follows that, for (x, q), (x′ , q ′ ) ∈ [0, δ]2 , ∥Tα (x, q) − Tα (x′ , q ′ )∥∞ ≤ Cλ α∥(x, q) − (x′ , q ′ )∥∞ for some finite constant Cλ . Choosing α0 > 0 sufficiently small, we may ensure that Cλ α0 < 1 and α0 λ < δ. Then Tα is a contraction on [0, δ]2 for every α ∈ (0, α0 ). By (E.48), every solution of the system (E.42) lies in [0, δ]2 for α < α0 . Hence the system has at most one solution. On the other hand, the scalar equation (E.47) has a solution for small α. Indeed, 0 < C(s) < 1 and C is continuous near 0, while s − αλC(s) is negative at s = 0 and positive at s = αλ for α small enough. Moreover, by the same Lipschitz argument, the map s 7→ αλC(s) is a contraction on a small interval, so this solution is unique. Thus the unique small-α solution of the two-dimensional system lies on the curve   s s (x, q) = , , 1 + s 1 + 2s and satisfies s = αλC(s). This proves the claim.

F

Additional technical results

Lemma F.1 (Lemma 3.2.2, [Tal10]). Consider a real c > 0 and a univariate concave function w(y) such that w′′ ≤ −c < 0. Let y ′ be the solution to w′ (y ′ ) = 0. Then we have Z Z ′ 2 c (y − y ) exp w(y) dy ≤ exp w(y) dy . 139

F.1

Truncated Logarithm

Lemma F.2 (Lemma 8.3.7, [Tal11]). For 0 ≤ x, z ≤ 1, we have x . logA x − logA z ≤ logA z Proof. Suppose z ≤ x. The result is clear if x ≤ exp(−A), when the left hand side is 0, or if z ≥ exp(−A), where the desired inequality follows from definition of logarithms. If z ≤ exp(−A) ≤ x, as here xz ≥ 1 and x, z ≤ 1, we obtain | logA x − logA z| ≤ | log x − log z| = log

x x = logA . z z

Else, suppose x ≤ z. By similar arguments as above, the result is clear if z ≤ exp(−A) or if x ≥ exp(−A), where we now note x ≥ exp(−A) ≥ z exp(−A) in the latter case. If x ≤ exp(−A) ≤ z, then noting log z ∈ [−A, 0], log x ≤ −A, and x, z ≤ 1, we obtain | logA x − logA z| = | − A − log z| ≤ A , | logA x − logA z| ≤ | log x − log z| = log Therefore, we obtain n x x o = logA | logA x − logA z| ≤ min A, log z z in this case as well. Lemma F.3 (Lemma 8.3.10, [Tal11]). If 0 ≤ x, y, z ≤ 1, we have | logA xz − logA yz| ≤ | logA x − logA y|1{z ≥ exp(−A)} . Lemma F.4 (Lemma 8.3.11, [Tal11]). If 0 < y ≤ x ≤ 1, then for any c > 0, | logA x − logA y| ≤ | logA y|1{y ≤ c} +

140

|x − y| . c

x . z

Record · ID 259461 · SHA-256 041d0b0165f6de5d
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.