Conceptio › Archive › arXiv CS
arXiv CSopen access

A Note on Scaling in Randomly Rotated Quantization and Its Connection to the CDEF +1 Pythagorean Relation

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

A Note on Scaling in Randomly Rotated Quantization and Its Connection to the CDEF +1 Pythagorean Relation

arXiv:2609.08759v1 [cs.IT] 8 Sep 2026

Uri Erez School of Electrical Engineering Tel Aviv University Tel Aviv, Israel [email protected] Abstract Quantization schemes based on randomized rotations have recently received renewed attention, including the roles of MMSE and unbiased reconstruction scalings. In this note, we point out the connection to classical results in statistical signal processing and communication theory. Specifically, the two reconstruction scales used in the EDEN line of work admit a natural interpretation as finite-dimensional, realization-dependent counterparts of the Wiener and unbiased coefficients in the classical CDEF formulation. At finite blocklength, the CDEF +1 relation holds pointwise for each rotation realization as an exact geometric (Pythagorean) identity, but does not hold after averaging the distortions over the rotation. The classical SNR relation SNRMMSE = SNRMMSE,U +1 is recovered as d → ∞: once the overall scale is handled separately, the empirical coordinate statistics of a randomly rotated vector approach their i.i.d. Gaussian counterparts, and the rotation-dependent quantities concentrate. Importantly, EDEN goes beyond this classical correspondence: for every finite d, its Haar-rotation formulation guarantees exact conditional unbiasedness, a stronger property than the second-order notion of unbiasedness in CDEF. We further comment on two distinct roles random rotations play in quantization: one is approximate Gaussianization of the coordinates; the other is decorrelation of reconstruction errors across quantization branches.

1

Introduction

Quantization schemes based on randomized rotations have recently received renewed attention, primarily due to their central role in LLM quantization; representative examples include [1, 2, 3]. Some of the underlying ideas, and in particular the random-rotate–quantize–rotate-back (RQRB) pipeline, can be traced back to classical areas in information theory and statistical signal processing; see, e.g., the discussions in [3, 4]. The purpose of this note is to shed some further light on these connections. Specifically, we: 1. Trace two distinct roles of the RQRB pipeline in the quantization literature: approximate Gaussianization on the one hand, and decorrelation of quantization errors on the other, with the latter corresponding to its interpretation as a dithering mechanism [9, 10]. 2. Relate the two reconstruction scalings employed in the EDEN line of work [5, 6] to the classical theory of Cioffi, Dudevoir, Eyuboglu, and Forney (CDEF) [7]. Namely, the EDEN scalings are precisely finitedimensional, realization-dependent counterparts of the classical Wiener and unbiased CDEF scalings. This correspondence is exact for every rotation realization; as the dimension grows, concentration connects these realization-dependent quantities to the classical scalar-Gaussian setting and its familiar +1 SNR relation.

2

The Classical Biased/Unbiased Relation

We begin with the classical scalar setting. Let X and Y be correlated zero-mean random variables, and denote PX = E[X 2 ], PY = E[Y 2 ], and rXY = E[XY ]. 1

As we recall next, Cioffi, Dudevoir, Eyuboglu, and Forney (CDEF) distinguish two natural linear scalings of Y , corresponding to biased and unbiased linear MMSE estimation [7]; see also Forney’s later Hilbert-space exposition [8].

2.1

Wiener scaling

The Wiener, or linear-MMSE, coefficient minimizes E[(X − αY )2 ] and is given by αW = αW (rXY , PY ) ≜

rXY . PY

(1)

By the orthogonality principle, the optimal estimation error satisfies E[EY ] = 0,

(2)

where X̂ = αW Y is the linear MMSE estimator and E = X − X̂.1 Thus, Wiener scaling makes the estimation error orthogonal to the observation and yields the backwardchannel relation X = αW Y + E, (3) with E ⊥ Y .

2.2

Unbiased (unit signal gain) scaling

As developed in CDEF [7], the unbiased coefficient is2 αU = αU (rXY , PX ) ≜

PX . rXY

(4)

The unbiased estimator is then Z ≜ αU Y = X + N,

(5)

E[XN ] = αU E[XY ] − E[X 2 ] = 0.

(6)

where Thus αU makes the effective noise orthogonal to the source (N ⊥ X). The relation (5) is called the forwardchannel realization. Remark 1. Here “unbiased” is used in the the weaker, second-order CDEF sense of unit signal gain. The condition X ⊥ N does not by itself imply the (conditional) unbiasedness relation E[Z|X] = X, equivalently E[N |X] = 0. See discussion of EDEN in the sequel. Writing PN = E[N 2 ], the SNR of the unbiased, forward-channel representation is naturally defined as SNRMMSE,U =

PX . PN

(7)

If one starts with the forward channel relation Z = X + N , the Wiener reconstruction is obtained by the additional scaling X̂ = αW|U Z, with reconstruction error Eα = X − αZ = (1 − α)X − αN.

(8)

E[Eα2 ] = (α − 1)2 PX + α2 PN .

(9)

Since X ⊥ N , 1 2

In the sequel MMSE will be used to denote linear MMSE. More precisely, CDEF begin with the forward-channel relation (5) while noting that it can be enforced to hold via scaling.

2

The minimizing coefficient relative to the unit-gain representation is αW|U =

PX . PX + P N

(10)

Thus, we have αW|U = ααW , or equivalently, αW = αW|U αU . The resulting MMSE is U MMSE =

PX PN . PX + P N

(11)

Defining3 PX , MMSE gives the biased/unbiased SNR relation of Lemma 2 of CDEF, SNRMMSE =

SNRMMSE = SNRMMSE,U + 1.

2.3

(12)

(13)

The Pythagorean interpretation

The identity also has the geometric interpretation emphasized by CDEF [7]. The forward representation Z = X + N, with X ⊥ N , forms one right triangle. The Wiener estimate X̂ = αW|U Z is the orthogonal projection of X onto the one-dimensional subspace spanned by Z, so that X = X̂ + E, with E ⊥ Z (and also E ⊥ X̂). This is the second right triangle in CDEF’s Fig. 9; Forney develops the same projection geometry in his Hilbert-space treatment of MMSE estimation [8]. The +1 relation is therefore Pythagoras applied to the forward- and backward-channel decompositions. We identify in the sequel these rules within EDEN [5]. Namely, replacing random variables by fixed vectors and the mean-square inner product by the Euclidean inner product gives the same two scaling rules for every realization.

2.4

Why the forward-channel SNR is natural for averaging

The forward-channel realization (5) is the natural viewpoint when several quantized reconstructions are to be linearly combined as elaborated on in [9, 10, 11].4 To that end, let us now relate the framework to the problem of averaging independently quantized elements of a vector (X1 , . . . , Xm ), where the elements are correlated. To further simplify the exposition assume the extreme case where all entries are identical; see Section II.B in [10]. Namely, we now identify Yi = Qi (Xi ) = Qi (X)

(14)

as the output of the i-th quantizer Qi (·). Ideally, we would like to conclude that Zi = X + Ni ,

i = 1, . . . , m,

3

Indeed, as should be clear from the exposition to follow, the terminology of signal-to-noise ratio to describe this quantity, though ubiquitous, should be interpreted with caution. 4 CDEF studies a channel coding scenario rather than the dual quantization scenario discussed in this note. The exposition here follows that of [11] where the notation SDRopt and uSDRopt was used in place of SNRMMSE and SNRMMSE,U , respectively.

3

where Zi = αU,i Yi . Assume that the Ni are zero mean, have equal power E[Ni2 ] = PN , and satisfy E[XNi ] = 0 and E[Ni Nj ] = 0 for i ̸= j. Under these assumptions, their average is again a forward channel with unit signal gain, m 1 X Z= Zi = X + N , (15) m i=1

where

m

1 X N= Nj , m j=1

and

m h i 1 X 2 E N = 2 E[Ni2 ] = PN /m. m j=1

We therefore obtain SNRMMSE,U,ave = mSNRMMSE,U .

(16)

Thus, the forward-channel representation is natural for averaging, since the estimation errors behave as additive, zero-mean, mutually uncorrelated noises. We next discuss several mechanisms for realizing this model.

2.5

Comments on the Roles of the Random Rotate-Quantize-Rotate Back Pipeline

Classical subtractive dithering realizes this forward-channel model for a uniform quantizer without overload, using uniform dither over one quantization interval: the noise is independent of the input, and independent dithers make the noises in different descriptions independent [12]. In [9, 10], the random rotate–quantize–rotate-back (RQRB) pipeline was proposed as an alternative form of dither for scalar quantization, with the goal of asymptotically eliminating second-order error correlations in distributed quantization of correlated sources. Both Haar rotations and randomized Hadamard transforms are studied, with the latter providing a computationally efficient implementation. The sources in that work are already modeled as temporally i.i.d. Gaussian (e.g., following transform coding), so RQRB is not used for Gaussianization. Rather, different random rotations are applied across quantization branches, making the resulting quantization errors asymptotically uncorrelated (via random rotations) and uncorrelated from the source (via CDEF scaling); hence enabling noncoherent combining. Shortly afterward, Suresh et al. [13] proposed using the RQRB pipeline in the closely related context of distributed mean estimation, with the random rotation playing a different role. A common randomized Hadamard transform is used across clients (branches in the terminology of [10]) to transform a deterministically modeled source vector so as to control the dynamic range of the elements to be fed into stochastic scalar quantizers. With independent stochastic-rounding randomness across clients, the errors are zero mean and pairwise uncorrelated, ensuring noncoherent combining. In turn, EDEN [5] employs RQRB with an independent Haar rotation for each client and deterministic scalar quantization. The rotation makes the relevant empirical coordinate statistics of an arbitrary input vector asymptotically Gaussian, permitting scalar quantization optimized for a Gaussian source, while a realization-dependent scaling ensures unbiasedness in the strong sense (see Remark 1). Independence across client rotations is not needed for Gaussianization or single-client unbiasedness. It is used at the averaging stage: together with conditional unbiasedness, it makes the reconstruction errors conditionally independent and eliminates their cross correlations. Thus, in EDEN, the rotations serve both as a Gaussianizing transform and, in the second-order sense of [10], as dither enabling noncoherent combining.

4

3

EDEN Scaling as a Deterministic Realization-Dependent Counterpart of CDEF

Let X ∈ Rd be a given non-zero vector and consider a realization of a Haar rotation U . The EDEN scheme forms the rotate–quantize–rotate-back observation Y before reconstruction scaling as −1 T Y = ηX U Q(V),

where

√

ηX =

d , ∥X∥

and V = ηX U X. For this deterministic pair, define the correlation 1 ⟨X, Y⟩, d

(d)

rXY = (d)

(d)

and normalized energies PX = d1 ∥X∥2 , PY = d1 ∥Y∥2 . The two CDEF scalings of Section 2 then become

and

  ⟨X, Y⟩ (d) (d) αW = αW rXY , PY = , ∥Y∥2

(17)

  ∥X∥2 (d) (d) αU = αU rXY , PX = . ⟨X, Y⟩

(18)

Indeed, (d) (d) rXY = PX

! d 1X Vi Q(Vi ) , d i=1

and hence, when EDEN is expressed as a scaling of Y, d

αU = Pd

i=1 Vi Q(Vi )

.

(19)

−1 Equivalently, the coefficient multiplying U T Q(V) is ηX αU = ∥X∥2 /⟨U X, Q(V)⟩, which is the scale used in the original EDEN formulation [5]. EDEN assumes that the reconstruction scale is represented without quantization error. It immediately follows that Z = αU Y = X + N, (20)

b = αW Y is the corresponding Euclidean projection. with X ⊥ N, for every rotation realization, while X These are precisely the deterministic counterparts of the two CDEF scalings. Treating U as random, Theorem 2.1 in [5] additionally proves the stronger finite-dimensional (unbiasedness) statement: EU [Z] = X. (21) Equivalently, if X is viewed as random, and independent of U , then E[Z | X] = X. EDEN’s realizationdependent scaling enforces exact source–error orthogonality for every realization; rotational symmetry additionally yields exact conditional unbiasedness.

5

3.1

SNR Interpretation and Relation to the Stochastic Scalar Setting

Define A(U ) =

∥X∥2 ∥Y∥2 . ⟨X, Y⟩2

The normalized errors of the unit-gain and Wiener reconstructions are given, respectively, by DU (U ) =

∥αU Y − X∥2 = A(U ) − 1, ∥X∥2

(22)

DW (U ) =

∥αW Y − X∥2 1 =1− . 2 ∥X∥ A(U )

(23)

and

Note that, for every realization, they indeed satisfy the CDEF Pythagorean relation: 1 1 = + 1. DW (U ) DU (U ) (d)

(d)

Define the averaged distortions over U by DU = EU [DU (U )] and DW = EU [DW (U )]. Define further the corresponding two SNRs by 1 (d) SNRMMSE,U = (d) DU and

1

(d)

SNRMMSE =

(d)

DW

Concavity of t/(1 + t) implies that by Jensen: (d)

(d)

SNRMMSE ≥ SNRMMSE,U + 1.

(24)

The large-d interpretation is as follows. For a Haar rotation, √ dG d V= , ∥G∥ where G ∼ N (0, Id ). √ As d grows, ∥G∥/ d concentrates around one, and the empirical quantities entering A(U ) approach their scalar Gaussian counterparts. Indeed, 1 Pd Q(Vi )2 d A(U ) =  P i=1 2 . d 1 V Q(V ) i i=1 i d Hence, if a = E[GQ(G)], b = E[Q(G)2 ], and G ∼ N (0, 1), then A(U ) approaches b/a2 . Consequently, the realization-dependent EDEN scalings approach the corresponding CDEF scalings for the scalar Gaussian model, and the Jensen gap vanishes. We therefore recover the CDEF relation asymptotically: (d) (d) SNRMMSE − SNRMMSE,U −→ 1, as d → ∞.

Acknowledgment The author thanks Or Ordentlich for drawing attention to the EDEN line of work and, in particular, to its reconstruction scaling. 6

References [1] J. Chee, Y. Cai, V. Kuleshov, and C. M. De Sa, “QuIP: 2-bit quantization of large language models with guarantees,” Advances in neural information processing systems, vol. 36, pp. 4396–4429, 2023. [2] S. Ashkboos, A. Mohtashami, M. L. Croci, B. Li, P. Cameron, M. Jaggi, D. Alistarh, T. Hoefler, and J. Hensman, “QuaRot: Outlier-free 4-bit inference in rotated LLMs,” Advances in Neural Information Processing Systems, vol. 37, pp. 100 213–100 240, 2024. [3] A. Zandieh, M. Daliri, M. Hadian, and V. Mirrokni, “TurboQuant: Online vector quantization with near-optimal distortion rate,” in International Conference on Learning Representations, vol. 2026, 2026, pp. 56 418–56 439. [4] O. Ordentlich and Y. Polyanskiy, “High-rate quantized matrix multiplication I,” IEEE BITS the Information Theory Magazine, 2026. [5] S. Vargaftik, R. B. Basat, A. Portnoy, G. Mendelson, Y. B. Itzhak, and M. Mitzenmacher, “EDEN: Communication-efficient and robust distributed mean estimation for federated learning,” in International Conference on Machine Learning. PMLR, 2022, pp. 21 984–22 014. [6] R. Ben-Basat, Y. Ben-Itzhak, G. Mendelson, M. Mitzenmacher, A. Portnoy, and S. Vargaftik, “A note on TurboQuant and the earlier DRIVE/EDEN line of work,” arXiv preprint arXiv:2604.18555, 2026. [7] J. M. Cioffi, G. P. Dudevoir, M. V. Eyuboglu, and G. D. Forney, “MMSE decision-feedback equalizers and coding. I. equalization results,” IEEE transactions on Communications, vol. 43, no. 10, pp. 2582– 2594, 1995. [8] G. D. Forney Jr, “Shannon meets Wiener II: On MMSE estimation in successive decoding schemes,” arXiv preprint cs/0409011, 2004. [9] R. Hadad, “Dithered quantization via orthogonal transformations and a Cauchy–Schwarz-like inequality,” Master’s thesis, Department of Electrical Engineering, Tel Aviv University, Tel Aviv, Israel, Jan. 2016. [Online]. Available: http://www.eng.tau.ac.il/∼uri/theses/hadad msc.pdf [10] R. Hadad and U. Erez, “Dithered quantization via orthogonal transformations,” IEEE Transactions on Signal Processing, vol. 64, no. 22, pp. 5887–5900, 2016. [11] J. Østergaard, U. Erez, and R. Zamir, “Incremental refinements and multiple descriptions with feedback,” IEEE Transactions on Information Theory, vol. 68, no. 10, pp. 6915–6940, 2022. [12] R. M. Gray and D. L. Neuhoff, “Quantization,” IEEE transactions on information theory, vol. 44, no. 6, pp. 2325–2383, 1998. [13] A. T. Suresh, X. Y. Felix, S. Kumar, and H. B. McMahan, “Distributed mean estimation with limited communication,” in International conference on machine learning. PMLR, 2017, pp. 3329–3337.

7

Record · ID 668052 · SHA-256 fbc14716b3a573ef
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.