Preprint.
W HY D OES P OST-T RAINING Q UANTIZATION W ORK ? Yuxiang Chen1,2 , Michael Beyer3 , Jun Zhu1 , Jianfei Chen1,∗ Dept. of Comp. Sci. and Tech., Institute for AI, BNRist Center, THBI Lab, Tsinghua-Bosch Joint ML Center, Tsinghua University 2 College of AI, Tsinghua University, 3 Bosch AI Research, Renningen, Germany [email protected], [email protected], {dcszj, jianfeic}@tsinghua.edu.cn
arXiv:2609.11716v1 [cs.LG] 10 Sep 2026
1
A BSTRACT Post-training quantization compresses large language models (LLMs) by storing their weights at reduced precision, and each quantized weight introduces an error into the hidden states. Naively, these errors should accumulate with depth and corrupt next-token prediction; randomly initialized models accumulate these discrepancies rapidly, whereas quantized pretrained models accumulate much less hidden-state error and largely maintain downstream task performance, even though they were never trained with quantization noise. This raises the question we address: why does post-training quantization work? Comparing full-precision and quantized forward passes, we identify two mechanisms that characterize pretrained quantization robustness. First, the error a layer newly introduces tends to oppose the error it inherits from the layer’s input. The two cancel partially such that the discrepancy between full-precision and quantized passes grows slowly. This counteracting residual interaction develops during pretraining. Our quantitative analysis identifies it as a major factor slowing hidden-error growth. Second, LM-head geometry preferentially preserves the scores and probabilities of highranked tokens, which typically represent the model’s most confident predictions. Together, these mechanisms explain why quantization error that passes through numerous layers can still produce only small output changes, and we verify the findings across models and quantization settings.
1
I NTRODUCTION
Weight-only post-training quantization (PTQ) compresses large language models by storing their weights at reduced precision, and methods such as GPTQ (Frantar et al., 2023) and AWQ (Lin et al., 2024) are now widely used to serve them. Quantizing a weight makes it slightly different from its full-precision value, so at every layer, the quantized model’s hidden states differ from its full-precision counterpart, and these errors should naively accumulate as they propagate through the layers. Yet in practice, even directly rounding the weights to 4-bit barely hurts: casting Qwen332B (Yang et al., 2025) to NVFP4 (Alvarez et al., 2025) without calibration lowers benchmark accuracy by only 0.43 percentage points on average across six zero-shot benchmarks (Sec. 4.1), even though the model was never trained with quantization noise. This raises our central question: why does post-training quantization work? The usual answer is that quantized weights remain close to their full-precision values (cosine similarity to their NVFP4 reconstructions ∼ 0.996), so errors should propagate slowly. Yet pretrained and randomly initialized weights have nearly identical reconstruction error (App. D.2), while the hidden-state discrepancy is 5.5× larger at random initialization (Fig. 1). Because the weight-level errors are closely matched, this gap in hidden-error growth should reflect properties acquired during pretraining. Prior work reduces or diagnoses quantization damage (Arai & Ichikawa, 2025; Lee et al., 2026; Lotfi et al., 2026), but leaves a mechanistic gap: it does not explain how quantization errors propagate through a pretrained model or why they produce only small output changes. ∗
Corresponding author
1
Preprint.
We trace the discrepancy block by block and find two mechanisms that act in series. Throughout, we compare the quantized and full-precision models on the same inputs; the error at each layer is the difference between their hidden states, and we study how its size changes with depth. First, the new error a block introduces tends to oppose the error it inherits from earlier blocks, so the two partially cancel and the size of the error grows slowly (Sec. 3). Second, the final hidden-state error reaching the LM head is mainly a rotation whose effect is attenuated by the high-dimensional LM head, especially for top-ranked tokens (Sec. 4). Together, they explain why a large internal discrepancy translates into a small output change. Both mechanisms hold across Qwen3 models of different scales, OLMo3 (Team Olmo et al., 2025), Gemma3 (Gemma Team et al., 2025), and the OLMoE mixture-of-experts model (Muennighoff et al., 2025) (Sec. 5). We summarize the work’s contributions as follows: 1. Reframing quantization robustness as a mechanistic question. We ask why post-training quantization works, and use comparisons with randomly initialized models to show that the usual explanation that weights stay close to their full-precision values is not the full answer; slow hidden-error growth largely comes from the pretrained model itself. 2. A quantitative account of slow error growth. The error added by each layer tends to oppose the error it inherits, slowing error growth. We derive an exact decomposition of this growth and confirm experimentally that this cancellation is significant for limiting hidden-error growth. 3. A theoretical account of stable top-ranked token predictions. The scores of the tokens with the highest output probabilities change only slightly under the surviving final hidden-state error; we derive how hidden-state error induces score changes through the LM-head geometry, and connect the resulting score perturbations to log-probability changes.
2
S ETUP : C OMPARING F ULL -P RECISION AND Q UANTIZED M ODELS
Next-token prediction in LLMs For a tokenized sequence x1:n = (x1 , . . . , xn ) in vocabulary V, (0) each token xi is embedded as hi ∈ Rd , where d is the hidden dimension, and then processed by L Transformer blocks (Vaswani et al., 2017). Block ℓ ∈ {1, . . . , L} computes a residual update: (ℓ) (ℓ−1) (ℓ) (ℓ−1) (ℓ) ui := F(ℓ) h1:i , hi = hi + ui . Here F(ℓ) denotes the complete computation performed by block ℓ, including the Attention and MLP contributions. The final hidden state is mapped to scores z by the LM head to obtain the next-token distribution, with T = 1 unless stated otherwise: hLM = Norm h(L) ∈ Rd , z = WLM hLM ∈ R|V| , p = softmax(z/T ) ∈ R|V| . n Quantifying the quantization impact A linear layer maps an input X to Y = XW⊤ using a b = weight matrix W. Weight quantization stores Q (W) as a low-precision approximation of W. Y ⊤ XQ (W) is thus an approximation of Y. Our study mainly uses the NVFP4 format (Alvarez et al., 2025) with round-to-nearest (RTN) (App. D). We quantize standard linear projections in Attention and MLP to NVFP4 and keep other operations, including the LM head, in full precision (BF16). c we refer to the unquantized model as Quantizing these linear layers gives the quantized model M; the original model M. We compare the two models on the same input sequence x1:n . At each position i ∈ {1, . . . , n − 1}, both models receive the prefix x1:i when predicting xi+1 . We analyze each position separately and omit the position index i below. Unless stated otherwise, the expectation E first averages over positions within each input sequence and then averages equally across sequences. c start with the same At each decoding position i, the original model M and quantized model M (0) (0) (ℓ) b b (ℓ) denote the initial hidden state, h = h . For every block ℓ ∈ {1, . . . , L}, let u and u (ℓ) (ℓ−1) (ℓ−1) b b (h c u(ℓ) = F(ℓ) (h b (ℓ) = F respective block updates of M and M: ) and u ). The two 1:i 1:i residual updates and their hidden-state difference are ( b (ℓ) − h(ℓ) , (ℓ) (ℓ−1) ∆h(ℓ) := h (ℓ) (ℓ) (ℓ−1) (ℓ) b b b , h =h +u , h =h +u (1) | {z } | {z } b (ℓ) − u(ℓ) ∆u(ℓ) := u (ℓ−1) (ℓ) b (ℓ−1) b (ℓ) original model; block ℓ: h
7→h
quantized model; block ℓ: h
2
7→h
Preprint.
Qwen3-32B
‖Δh(ℓ)‖
[‖Δh ‖2] Random Init. 2683.0
1
2000
0
16
32
48
Pretrained 0.165
0
64
OLMo3-7B · Stage 1 [‖Δh(ℓ)‖2]
0
16
[ (ℓ) 2 ] ‖h ‖2
32
48
Stage-1 step: 0
Step 0 254.9
1k
8
0.5
16
24
32
0.0
Stage 1 final 12.6
(A) Abs. error norm
, Δu )]
0
16
10k
100k
Pretrained
32
1M
8
64
final
0.00 Stage 1 final
−0.25 0
48
[cos(Δh(ℓ − 1), Δu(ℓ))]
Step 0 0.498 Stage 1 final 0.136
0
Pretrained (ℓ)
−0.25
64
‖Δh(ℓ)‖
200 0
[cos(Δh
(ℓ − 1)
0.00
Random Init. 1.109
Pretrained 484.1
0
Random Init.
[ (ℓ) 2 ] ‖h ‖2
(ℓ)
16
24
32
0
8
16
24
32
Decoder block ℓ (B) Rel. error norm
(C) Input-vs-update error cos.
Figure 1: Pretraining changes quantized model’s hidden-error growth. Rows compare random/pretrained Qwen3-32B and OLMo3-7B Stage-1 checkpoints on C4. (A) Mean absolute hiddenerror norm; (B) mean relative hidden-error norm; (C) mean cosine between the block-input error and block-update error. Error bars show standard deviation. We refer to the quantization-induced hidden-state difference ∆h(ℓ) as the hidden error at layer ℓ. b LM := Norm h b (L) : The LM-head inputs are then hLM := Norm h(L) and h z = WLM hLM , p = softmax(z) | {z }
b LM , p b b = softmax(b z = WLM h z) . | {z }
original model M; hLM7→p
b LM7→p c h b quantized model M;
b LM − hLM induces the score difference ∆z := b The LM-head input difference ∆hLM := h z−z= WLM ∆hLM . We write TopK (z) ⊆ V for the vocabulary tokens with the largest K logits in M, and measure output quality by cross-entropy (CE) and DKL (p∥b p).
3
T HE G ROWTH OF H IDDEN -E RROR N ORM ACROSS L AYERS
b (ℓ) − h(ℓ) changes across The models start from the same hidden state but hidden error ∆h(ℓ) = h ℓ. This section quantifies the growth by ∥∆h(ℓ) ∥2 and its relative size ∥∆h(ℓ) ∥2 /∥h(ℓ) ∥2 . 3.1
H IDDEN -E RROR G ROWTH IN R ANDOMLY I NITIALIZED AND P RETRAINED M ODELS
NVFP4 weight-quantization reduces Qwen3-32B accuracy by only 0.43 percentage points on average across six zero-shot benchmarks (Sec. 4.1). The final relative hidden-error norm is (L) b (L) (L) (L) ∥∆h ∥2 /∥h ∥2 ≈ 0.15 and cosine similarity is cos ∠ h , h ≈ 0.98 on average for pretrained models. A direct intuition is that networks are robust to NVFP4 because each quantized weight matrix Q(W) remains close to its full-precision counterpart W (CosSim(W, Q(W)) ≈ 0.9955 in App. D.2). This suggests that each quantized matrix multiplication introduces only a small error. One might therefore expect the errors introduced across layers to remain limited, leaving a small final hidden error. However, this account is incomplete. Randomly initialized weights have nearly the same reconstruction cosine similarity as the pretrained model under NVFP4 quantization (App. D.2). If this weight-level similarity alone explained the small final hidden error, the randomly initialized and pretrained models should therefore show similar hidden-error growth. However, the quantized randomly initialized Qwen3-32B reaches a 5.5× larger final absolute hidden-error norm and a 6.7× larger relative hidden-error norm than its pretrained checkpoint (Fig. 1). The same pattern appears for OLMo3-7B: the initialized step-0 has 20.2× larger final absolute error and 3.7× larger relative error than the final checkpoint. The growth trends across layers also differ. Over the 64 layers, randomly initialized Qwen3-32B follows the consistently growing trends E∥∆h(ℓ) ∥2 ∝ ℓ0.80 and E ∥∆h(ℓ) ∥2 /∥h(ℓ) ∥2 ∝ ℓ0.30 , while the pretrained relative-error curve even stops increasing in the later layers (details in App. E.1). 3
Preprint.
ℓ
Squared error: (R(ℓ))2
ℓ
(j)
∑ Tadd
j=1
0.2 √ [(R(L) )2 ] = 1.110
1.0
0
(j)
∑ Talign
j=1
√ [(R(L) )2 ] = 0.178
0.1
0.5 0.0
ℓ
(j)
∑ Tinter
j=1
0.0 −0.1
16
32
48
Decoder block ℓ
64
0
(A) Random-initialized
16
32
48
Decoder block ℓ
64
(B) Pretrained
Figure 2: Accumulated contributions of the three terms in Thm. 1. Randomly initialized and pretrained Qwen3-32B on C4. Each colored curve accumulates one term over blocks 1–ℓ and shows its averaged value across inputs. The black curve is their exact sum. Error bars show standard deviation across inputs. Weight-level similarity alone is therefore insufficient to explain robustness across depth; there should be a model-side mechanism that distinguishes randomly initialized models from pretrained models. The next subsection identifies counteraction between block-input error ∆h(ℓ−1) and block-update error ∆u(ℓ) as a major contributor to this slower growth. 3.2
C OUNTERACTION BETWEEN B LOCK -I NPUT AND B LOCK -U PDATE E RRORS
Eq. (1) shows hidden error after block ℓ is ∆h(ℓ) = ∆h(ℓ−1) + ∆u(ℓ) , where ∆h(ℓ−1) is the input error inherited from earlier blocks and ∆u(ℓ) is the block-update error. Here is the recurrence: Proposition 1 (Squared hidden-error recurrence). For every Transformer block ℓ ∈ {1, . . . , L}, ∥∆h(ℓ) ∥22 − ∥∆h(ℓ−1) ∥22 = ∥∆u(ℓ) ∥22 + 2⟨∆h(ℓ−1) , ∆u(ℓ) ⟩
(2)
The second term on the right determines whether the block-update error ∆u(ℓ) increases or cancels the block-input hidden error ∆h(ℓ−1) . We call the latter case counteraction: ⟨∆h(ℓ−1) , ∆u(ℓ) ⟩ < 0. Counteraction offsets part of the newly introduced error, and slows the hidden-error norm growth. P For pretrained Qwen3-32B, over layers 1–48, the interaction term 2 E⟨∆h(ℓ−1) , ∆u(ℓ) ⟩ cumulaP tively cancels 50.2% of the block-update error term E∥∆u(ℓ) ∥22 contribution, explaining the slow growth in Fig. 1(A). At random initialization, 2E⟨∆h(ℓ−1) , ∆u(ℓ) ⟩ is at most 0.16% of E∥∆u(ℓ) ∥22 for each layer, consistent with the near-zero cosines in Fig. 1(C). We next derive the recurrence for the relative hidden error ∥∆h(ℓ) ∥2 /∥h(ℓ) ∥2 which normalizes for changes in hidden-state scale. (ℓ)
∥2 Theorem 1 (Relative hidden-error recurrence). Define R(ℓ) := ∥∆h . For every ℓ ∈ {1, . . . , L} ∥h(ℓ) ∥ 2
such that h(ℓ−1) , h(ℓ) , u(ℓ) are nonzero, 2 2 R(ℓ) − R(ℓ−1) (ℓ) 2 h 2 i 2⟨∆h(ℓ−1) , ∆u(ℓ) ⟩ 2⟨h(ℓ−1) , u(ℓ) ⟩ 2 ∥∆u(ℓ) ∥2 ∥u ∥2 (ℓ−1) 2 − R + − R(ℓ−1) . = (ℓ) (ℓ) (ℓ) (ℓ) 2 2 ∥u ∥2 ∥h ∥2 ∥h ∥2 ∥h ∥2 {z } | | {z }| {z } relative additional error (block-update vs. block-input): Tadd
error interaction: Tinter
(3)
residual-norm contribution: Talign
A detailed derivation is given in App. C.1, and Fig. 2 quantifies the three terms for Qwen3-32B: • Tadd compares the relative error in the current block update with the relative error already present at the block input, scaled by the update-to-hidden norm ratio. • Tinter measures the signed interaction between the block-input error ∆h(ℓ−1) and block-update error ∆u(ℓ) . This term encompasses the counteraction that results in the slower error growth: with Tinter < 0 for most blocks in pretrained Qwen3-32B, the term cumulatively cancels 63.4% 2 of Tadd over layers 1–L, substantially slowing the growth of R(ℓ) . • Talign captures how the change of hidden-state norm affects the relative error. If ⟨h(ℓ−1) , u(ℓ) ⟩ > 0, the update increases ∥h(ℓ) ∥2 and reduces the relative error. In pretrained Qwen3-32B, this term cancels 18.4% of Tadd cumulatively, and most of the cancellation occurs during later blocks. 4
Preprint.
Quantized ̂ Removed : Δu ↦ Δu ⟂ ; same norm R
8
(ℓ)
= [
Total Err. [(
Reversed : Δu ↦ − Δu
‖h(ℓ)‖2 ] Simulated counteraction removal or reversal
‖Δh(ℓ)‖2
10−2
̂ ‖h(ℓ) ‖2
Angular : [2
‖h(ℓ)‖2
(1 − cos ∠)]
Decoder blocks
LM-head input
Total 0.0318
4
6
0.0281
0.0267
4
2
2 0
2 ‖h(ℓ)‖2 ) ]
‖Δh(ℓ)‖2
̂ Length mismatch : [(‖h(ℓ) ‖2/‖h(ℓ)‖2 − 1)2]
Angular 0.0305
R(L) = 1.387 ; KL = 7.856 R(L) = 0.485 ; KL = 0.628 R(L) = 0.165 ; KL = 0.045
0
16
32
48
Decoder block ℓ
0
64
1
16
32
Length 0.0013
0.0013
48
hLM = Norm(h(L))
Decoder block ℓ
64
(B) Length–angle error decomposition
(A) Counteraction intervention
Figure 3: Counteraction interventions and length–angle decomposition of the hidden error. Qwen3-32B on C4. (A) Relative hidden error after counteraction removal or reversal. (B) Terms of Prop. 2; the separated right region shows the LM-head input terms. Error bars show SD. Together, Tinter and Talign cumulatively cancel 81.8% of Tadd in pretrained Qwen3-32B over layers 1–L. At random initialization, the cumulative sum of these two terms is only 0.09% of Tadd over layers 1–L. These measurements identify counteraction as a major factor in slowing down the hidden-error growth in the pretrained model: across blocks, the block-update error ∆u(ℓ) tends to point against the block-input hidden error ∆h(ℓ−1) to slow the error growth. 3.3
H IDDEN -E RROR G ROWTH WITHOUT C OUNTERACTION
After showing the significance of counteraction, we test what happens when counteraction is removed. We construct layer-wise hidden-state trajectories in which each block-update error ∆u keeps its norm, but its direction is modified to be counteraction-free. We compare the hidden-error trajectories and the final KL divergence across intervention policies. Concretely, at every intervened block, we use one of the hidden-state recursions: b (ℓ) = h b (ℓ−1) + u(ℓ) +∆u(ℓ) ; Removal: h rem rem ⊥
b (ℓ) = h b (ℓ−1) + u(ℓ) −∆u(ℓ) . Reversal: h rev rev (ℓ)
In both recursions, ∆u(ℓ) follows Eq. (1) and is recomputed along the modified trajectory; ∆u⊥ is orthogonal to the block-input error ∆h(ℓ−1) . Both interventions are applied in blocks 17–48, where counteraction is strongest in Qwen3-32B; reversal is applied only when the interaction is negative. At the final block, removal raises the relative hidden error by 2.94×, while reversal raises it by 8.41×. Both interventions also sharply increase KL. Moreover, after counteraction removal, Tadd increases by 4.3× cumulatively (App. E.7), which indicates that counteraction not only contributes through the negative term Tinter , but also alters the subsequent trajectory in a way that limits Tadd . The interventions provide causal evidence that counteraction is a major factor limiting hiddenerror growth. Sec. 4 examines how the remaining hidden-state error affects the outputs.
4
T HE E FFECT OF F INAL H IDDEN E RROR ON O UTPUT Q UALITY
The quantized Transformer blocks leave a nonzero error at the final decoder state h(L) : the mean relative error norm ∥∆h(L) ∥2 /∥h(L) ∥2 across the three datasets is ∼0.245 (App. E.2). To see what kind of error survives to the LM head input, we decompose it into its length and angle components. b (ℓ) − h(ℓ) . Proposition 2 (Length–angle decomposition of relative hidden error). Define ∆h(ℓ) := h (ℓ) b (ℓ) ̸= 0, For every block ℓ such that h ̸= 0 and h ∥∆h(ℓ) ∥ 2 2
∥h(ℓ) ∥2
=
∥h b (ℓ) ∥
2
∥h(ℓ) ∥2
−1
2
(ℓ)
+2
b ∥h
∥2
b (ℓ) . 1 − cos ∠ h(ℓ) , h
∥h(ℓ) ∥2 2 (ℓ) b (ℓ) ∥2 = ∥h(ℓ) ∥2 , then ∥∆h(ℓ) ∥2 = 2[1 − cos ∠(h(ℓ) , h b (ℓ) )]. In particular, if ∥h ∥h ∥ 2
5
(4)
Preprint.
Output discrepancy
Metric
C4
WikiText
GSM8K
ΔCE ↓
+0.019
+0.034
+0.040
Rel. ΔCE ↓
+0.7%
+1.5%
+3.3%
KL ↓
0.045
0.090
0.079
Flip@1 ↓
10.7%
12.7%
8.3%
Ret@10 ↑
88.7%
84.9%
83.9%
Ret@20 ↑
88.9%
84.2%
82.5%
Model
ARC-C
ARC-E
Hella.
MMLU
Wino. TQA-MC1
BF16
61.09
83.42
82.66
80.77
72.85
38.80
W4
61.26
83.33
82.45
80.34
70.48
39.17
±1.42
±0.76 ±0.76
±0.38 ±0.38
±0.33 ±0.34
±1.28
±1.30
10−1 10−3 102 101 10−1 −3
Downstream accuracy (%, mean ± SD) ±1.41
102 Probability (%) 101
±1.69 ±1.71
ARC-C: ARC-Challenge · ARC-E: ARC-Easy · Hella.: HellaSwag Wino.: WinoGrande · TQA-MC1: TruthfulQA MC1
10 102 101 10−1 10−3
BF16
9.20% 10.1%
7.63% 6.13%
7.17% 6.13%
5.24% 5.76%
city (1)
northern (2)
central (3)
part (4)
southern (5)
21.2% 15.5%
12.9% 13.7%
11.4% 13.7%
11.4% 8.29%
6.90% 8.29%
July (1)
November (2)
April (3)
February (4)
March (5)
12.6% 7.78%
11.1% 11.3%
6.76% 7.78%
5.26% 3.24%
turns (2)
drove (3)
goes (4)
drives (5)
38.9% 44.8%
turned (1)
(A) Output metrics and task accuracy (W4 vs. BF16)
W4
C4
22.1% 18.9%
0.488% 0.223%
0.458% 0.154%
⋯ ⋯ Autonomous (29) autonomous (30) WikiText-103
⋯ 0.026%6.7e-03% 0.025%2.3e-03% till (29) – (30) ⋯ GSM8K
⋯ ⋯
0.096% 0.052%
0.096% 0.036%
starts (29) ends (30) Token (rank in original model)
(B) One-step token probability distributions
Figure 4: Output metrics and paired next-token probabilities after W4 quantization. Qwen332B. (A) Output-quality metrics and zero-shot benchmark accuracy. (B) Paired BF16/W4 probabilities at one representative next-token position per dataset, showing BF16 ranks 1–5 and 29–30. We refer to the two terms on the right-hand side of Eq. (4) as the length-mismatch contribution and the angular contribution, respectively. Fig. 3(B) shows that the angular contribution accounts for 88.9–98.7% of the squared relative hidden error across depth, while changes in hidden-state norm remain small. After final normalization, this corresponds to an average 12.47◦ rotation of the LM-head input (Fig. 5(A)). We next measure its effect on output quality (Sec. 4.1) and explain theoretically why top-ranked token probabilities are more stable in the outputs (Sec. 4.2). 4.1
O UTPUT D ISTRIBUTIONS AND D OWNSTREAM ACCURACY AFTER Q UANTIZATION
c with only Hereafter, BF16 denotes the original model M evaluated in bfloat16, and W4 denotes M its Transformer block linear weights quantized to NVFP4. We show that the output quality is mostly retained after quantization, especially for top-ranked tokens. First, we compare the output quality of BF16 and W4 models by loss and divergence: ∆CE := ∆CE CEW4 − CEBF16 , its relative value CE , and the forward KL divergence DKL (p∥b p). Fig. 4(A) BF16 reports all three metrics on each dataset (details in App. D.3), and the accuracy drops by only 0.43 percentage points on average across six benchmarks. Across other models, including Qwen3-30BA3B, Qwen3-8B, and OLMo3-32B, quantization also largely preserves output quality (App. E.10). Beyond the metrics, the BF16 model’s highest-scoring tokens largely remain top-ranked in the W4 model. Let πr ∈ V denote the rank-r token of BF16 model’s score z. Define h |Top (z) ∩ Top (b i K K z)| Flip@1 := Pr[Top1 (b z) ̸= Top1 (z)] , Ret@K := E . K Across the three datasets, Flip@1 is 8.3%–12.7%, while Ret@10 and Ret@20 are both about 85% on average (Fig. 4(A)). In the examples shown in Fig. 4(B), the average relative probability change over ranks 29–30 is 2.5–4.4× that over ranks 1–5. Sec. 4.2 tests this pattern across all evaluated positions and explains why top-ranked token scores are less sensitive to quantization. 4.2
P REFERENTIAL P RESERVATION OF T OP -R ANKED S CORES AND P ROBABILITIES
In this subsection, we find that top-ranked token scores have smaller relative errors after quantization and theoretically explain why this trend arises through the geometry of the LM head. This preferential preservation matters because top-ranked tokens govern the high-probability region of b LM and b b is the the output. Here h z are the LM-head input and score vector of the W4 model, and p corresponding probability vector; hLM , z, and p denote their counterparts in the unquantized model. We first examine how the score z = (z1 , z2 , . . . , z|V| )⊤ changes after quantization. The score of token k ∈ V is the projection of the LM-head input onto the corresponding LM-head weight vector: zk = ⟨wk , hLM ⟩. Because the BF16 and W4 models share the same LM head, a token score is changed only by the input norm ∥hLM ∥2 and projection angle ∠(wk , hLM ); App. C.3 gives the b LM ∥2 /∥hLM ∥2 − 1 is only exact decomposition. The average relative input-norm change E ∥h 6
Preprint.
C4
WikiText-103
Pretrained W4
GSM8K
102
∘
16.54° 101 12.60°
Random-init. W4
0.209° 0.158°
10−1
Second-order approx.
[|Δlog pπr |]
[|zπ̂ r − zπr |/zπr ], zπr > 0
Random-init.
101
8.26° 100
First-order approx.
100
100
Random-init.
Pretrained
10−1
Pretrained
0.109° ̂ ∠(hLM, hLM )
̂ k ∈ | ∠(wk, hLM ) −∠(wk, hLM) |
(A) Hidden rotation vs. LM-head angle shift
10−2
1
10
100
1k 10k || BF16-model rank r
(B) Relative score change by rank
10−1
1
10
100
1k 10k || BF16-model rank r
(C) Log-probability change by rank
Figure 5: Top-ranked token scores and probabilities are more stable after quantization in the pretrained model. Qwen3-32B on C4, WikiText-103, and GSM8K text. (A) Across datasets, the LM-head input rotates by 8.26◦ to 16.54◦ , whereas the vocabulary-mean projection-angle change is only 0.109◦ to 0.209◦ . (B) Relative score error increases toward lower-ranked tokens and is much smaller in the pretrained model than at random initialization. The theoretical approximation is under ρ = 1 of Thm. 2. (C) Log-probability error follows the same rank trend, and the approximations follow Thm. 3. Error bars show one log-SD; details are given in App. D.3. 3.7%. We therefore focus on quantifying how changes in the projection angles affect the token scores. Fig. 5 gives two observations: • Observation 1: projection angles change much less than the LM-head input. In Fig. 5(A), the b LM ) ≈ 12.5◦ , whereas the vocabulary-mean projection-angle change is input rotates by ∠(hLM , h b LM )| ≈ 0.159◦ , a 78.6× attenuation. |∠(wk , hLM ) − ∠(wk , h • Observation 2: top-ranked tokens have smaller score and log-probability changes for the quantized pretrained model. For BF16 rank-r token πr , the relative score change grows with rank in Fig. 5(B). The log-probability change E[|∆ log pπr |] follows the same rank trend, while both quantities are much larger at random initialization (Fig. 5(B–C)). b LM ) changes To explain these, Thm. 2 below derives how the hidden-state rotation α = ∠(hLM , h both projection angles and relative scores. Intuitively, in high dimension, only a small component of this rotation affects the projection onto any fixed LM-head weight vector wk , so the angle between wk and the hidden state changes much less than the hidden state itself. Moreover, a smaller angle between wk and hLM makes the relative score less sensitive to the same rotation, and higher-ranked tokens tend to have such smaller angles empirically (App. E.10), producing the rank-dependent trend in Fig. 5(B). Thm. 3 then links these score changes to the log-probability changes in Fig. 5(C). Theorem 2 (Projection-angle and score changes under an isotropic rotation of the LM-head input). b LM , and LM-head weight vector wk be nonzero. Define Let d ≥ 3 and let LM-head input hLM , h b LM ), α := ∠(hLM , h
θk := ∠(wk , hLM ),
b LM ), θbk := ∠(wk , h
∥hLM ∥2 ρ := ∥h LM ∥2 b
b LM is Conditional on the measured α, assume that the direction of the rotation from hLM to h uniformly distributed over all directions orthogonal to hLM . For zk ̸= 0 and α, θk ∈ (0, π), with all expansions taken as α → 0+ for fixed d and θk , let ∆zk := zbk − zk . Then s Γ d−1 2 2 2 b = 1 + O(d−1 ) , (5) E θk − θk = αµd + O(α ), µd := √ d π(d − 1) πΓ 2 ∆zk 1 d−2 2 = (ρ cos α − 1) + ρ tan θk sin α · qk , qk ∼ Beta , , (6) zk 2 2 ∆zk E = α|tan θk |µd + O(α2 ), if ρ = 1, (7) zk Here qk is symmetric about zero. The corresponding second-order estimate is given in Eq. (13). A detailed derivation is given in App. C.4. The first-order and second-order analytic references from Eqs. (7) and (13) set ρ = 1 to remove the modest norm difference and isolate the effect of changing the LM-head input direction, while the empirical W4 score-error curves in Fig. 5(B) retain the measured input-norm changes. We next relate the score changes to the log-probability changes: 7
Preprint.
Theorem 3 (Log-probability change estimates). At a fixed next-token position, let ∆ log pk := log pbk − log pk . The exact identity and expansions for uniformly small centered score changes are b) (exact), ∆ log pk = ∆zk − log Ej∼p e∆zj = ∆zk − Ej∼p ∆zj − KL(p ∥ p ∆ log pk = ∆zk − Ej∼p ∆zj + O(Varj∼p (∆zj )) (1-order), (8) 3 1 ∆ log pk = ∆zk − Ej∼p ∆zj − 2 Varj∼p (∆zj ) + O(Ej∼p |∆zj − Ei∼p ∆zi | ) (2-order). A detailed derivation is given in App. C.5. Thms. 2 and 3 account for the two observations above: • Why do projection angles change so little? Eq. (5) in Thm. 2 gives E|θbk − θk | = αµd + O(α2 ). Thus, in high dimension, the rotation of the hidden state is strongly attenuated when measured as the angle change to any fixed LM-head row vector. For Qwen3-32B, the predicted attenuation µ−1 d ≈ 89.7× is close to the measured 78.6× in Fig. 5(A). • Why are higher-ranked tokens more stable? Eqs. (6) and (7) in Thm. 2 show that, for positive scores, relative score sensitivity is governed by |tan θπr |. A smaller projection angle therefore gives a smaller relative score change under the same hidden-state rotation. Empirically, higherranked tokens tend to have smaller projection angles (App. E.10), matching the rank dependence in Fig. 5(B). Thm. 3 then connects these score changes to log-probability changes, and both approximations reproduce the increasing rank trend in Fig. 5(C). These results also explain why a non-negligible rotation of hLM produces only modest output changes in Fig. 4(A). KL remains small because it weights tokens by their original probabilities, and the highest-probability tokens are best preserved after quantization. CE remains small because the probability assigned to the ground-truth token changes little on average.
5
D ISCUSSION
5.1
C ONSISTENCY ACROSS M ODELS , Q UANTIZATION S ETTINGS , AND PTQ A LGORITHMS
The counteraction phenomenon introduced in Sec. 3 persists across pretrained Qwen, OLMo, and Gemma models, including dense and mixture-of-experts (MoE) architectures (App. E.2). The layerwise curve of E cos ∠(∆h(ℓ−1) , ∆u(ℓ) ) changes little and remains negative in most blocks when quantizing weights, activations, or both (Fig. 6(A)), and across multiple PTQ algorithms (Fig. 6(B)), showing that calibration-based PTQ changes the counteraction geometry little (App. E.9). Randominitialization comparisons and intervention experiments on additional models further support the role of counteraction in slowing hidden-error growth (Apps. E.1 and E.7). Preferential preservation of top-ranked scores and probabilities also transfers across models. Higherranked tokens tend to have smaller projection angles to the LM-head input and smaller relative score errors (App. E.10), and this rank dependence persists across softmax temperatures (App. E.11). 5.2
A NALYSIS OF THE S OURCE OF C OUNTERACTION
What produces this negative error interaction that slows hidden-error growth (Sec. 3)? Prior work found that residual updates across Transformer layers can partially cancel earlier updates. When an earlier contribution is rescaled, the later block adjusts its opposing update, showing that its response depends on the input signal (Patrawala et al., 2025). Does a block’s input hidden error similarly induce a change in its update that points against and partially cancels that error? This self-correcting tendency is much stronger for errors than for the native hidden-state/update b (ℓ−1) , u b (ℓ) ) are positive in 65.6% and 64.1% of pairs. Specifically, cos ∠(h(ℓ−1) , u(ℓ) ) and cos ∠(h the blocks, whereas interactions between the errors cos ∠(∆h(ℓ−1) , ∆u(ℓ) ) are negative in 82.5% of the blocks (App. E.6). Thus, the strong negative relation is not a general tendency for a block update to oppose its input hidden state. It emerges more consistently between the accumulated input hidden error and the new block-update error. b (ℓ−1) ) − F(ℓ) (h(ℓ−1) ) into two parts. The direct b (ℓ) (h To identify its source, we separate ∆u(ℓ) = F b (ℓ) (h(ℓ−1) ) − F(ℓ) (h(ℓ−1) ), which changes the weights at a fixed input. The inputweight effect is F b (ℓ−1) ) − F b (ℓ) (h b (ℓ) (h(ℓ−1) ), which keeps the weight fixed and changes only its error response is F 8
Preprint.
[cos ∠(Δh(ℓ − 1), Δu(ℓ))]
[cos ∠(Δh(ℓ − 1), Δu(ℓ))]
0.0
0.0 −0.2
−0.2 W4
1
A4
16
32
RTN NVFP4 GPTQ NVFP4 AWQ NVFP4
−0.4
W4A4
48
64
1
9
18
27
Decoder block ℓ
Decoder block ℓ
(A) Weight and activation quantization (Qwen3-32B)
(B) PTQ algorithms (Qwen3-4B)
36
Figure 6: Blockwise counteraction across quantization components and PTQ methods. (A) Qwen3-32B on C4 under W4 (weight-only), A4 (activation-only), and W4A4 (both quantized). (B) Qwen3-4B under RTN, GPTQ, and AWQ, all quantizing the same weights with NVFP4 format.
input. Although the two parts have comparable magnitudes, the input-error response contributes PL (ℓ) 99.9% of the negative cumulative interaction ℓ=1 Tinter , compared with less than 0.1% from the direct weight effect (App. E.6). Counteraction therefore comes mainly from the block’s response to its input error, which tends to oppose accumulated hidden error and slow its propagation. 5.3
R ARE FAILURES IN H IDDEN -E RROR P ROPAGATION
Counteraction significantly slows the hidden-error growth, which provides a self-correction effect that enhances the model’s robustness to the quantization noise. However, at rare decoding positions, the perturbation enters a regime where counteraction is no longer sufficient to keep the quantized hidden-state trajectory close to the original model, producing a large hidden-norm mismatch. These rare failures produce extreme values that can dominate the averages and make the typical error dynamics harder to characterize. We therefore exclude these positions from the recurrence-based aggregate statistics (only ∼0.28% of positions are excluded) using the norm-based rule in App. D.3 and analyze them separately in App. E.3. These cases show that large hidden-state norm distortion can cause output failures, but does not always do so (Fig. 12). 5.4
L IMITATIONS
Our analysis has three limitations. First, we analyze the quantization mechanisms for individual next-token predictions, but not how these mechanisms extend to multi-token generation. Second, our mechanism analysis relies on aggregate statistics and does not precisely characterize every hiddenerror trajectory or output change, especially at rare positions with extreme errors. Third, the LMhead theory uses a uniform-direction assumption for the approximation, although actual hidden-state rotation directions need not be uniform. Two questions remain open: whether training stochasticity produces counteraction, as suggested by implicit-bias analyses of SGD (Smith et al., 2020), and whether these findings can improve post-training quantization.
6
C ONCLUSION
This work studies why pretrained models maintain stable next-token predictions after weight quantization. Comparing the original and quantized forward passes reveals two mechanisms: (1) Even though the weights of randomly initialized and pretrained models have nearly identical reconstruction cosine similarities under NVFP4 quantization, pretrained models accumulate much less hidden error. Our exact recurrence isolates a mechanism specific to pretrained models: the block-update error tends to oppose the block-input hidden error. Interventions that remove this counteraction sharply increase the final hidden error, providing causal evidence that counteraction is a major factor limiting hidden-error growth. (2) The quantization-induced change in the final hidden state is primarily a rotation rather than a change in norm. Our theory shows that, in high dimensions, only a small component of this rotation affects the projection onto any fixed LM-head weight vector and that LM-head geometry makes the resulting score and log-probability changes smaller for higherranked tokens. We verify both mechanisms across models and quantization settings. 9
Preprint.
ACKNOWLEDGMENTS The authors sincerely thank Pengle Zhang and Zichen Liang for insightful discussions.
R EFERENCES Eduardo Alvarez, Omri Almog, Eric Chung, Simon Layton, Dusan Stosic, Ronny Krashinsky, and Kyle Aubrey. Introducing NVFP4 for efficient and accurate low-precision inference. NVIDIA Technical Blog, 2025. URL https://developer.nvidia.com/blog/introducin g-nvfp4-for-efficient-and-accurate-low-precision-inference/. Yamato Arai and Yuma Ichikawa. Quantization error propagation: Revisiting layer-wise posttraining quantization. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. URL https://openreview.net/forum?id=a3l3K9khbL. Nora Belrose, Igor Ostrovsky, Lev McKinney, Zach Furman, Logan Smith, Danny Halawi, Stella Biderman, and Jacob Steinhardt. Eliciting latent predictions from transformers with the tuned lens, 2023. URL https://arxiv.org/abs/2303.08112. Ido Ben-Yair, Gil Ben Shalom, Moshe Eliasof, and Eran Treister. Quantized convolutional neural networks through the lens of partial differential equations. Research in the Mathematical Sciences, 9(4):58, 2022. doi: 10.1007/s40687-022-00354-y. Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al. Pythia: A suite for analyzing large language models across training and scaling. In International Conference on Machine Learning, pp. 2397–2430. PMLR, 2023. Nicola Cancedda. Spectral filters, dark signals, and attention sinks. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 4792–4808, Bangkok, Thailand, August 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.acl-long.263. URL https://aclanthology.org/2024.acl-long.263/. Albert Catalan-Tatjer, Niccolò Ajroldi, and Jonas Geiping. Training dynamics impact post-training quantization robustness. In The Fourteenth International Conference on Learning Representations, 2026. URL https://openreview.net/forum?id=ZXr3Xx7Z1O. Ting-Yun Chang, Muru Zhang, Jesse Thomason, and Robin Jia. Why do some inputs break low-bit LLM quantization? In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pp. 3410–3429, 2025. Jianfei Chen, Yu Gai, Zhewei Yao, Michael Mahoney, and Joseph Gonzalez. A statistical framework for low-bitwidth training of deep neural networks. In H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (eds.), Advances in Neural Information Processing Systems, volume 33, pp. 883–894. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/p aper_files/paper/2020/file/099fe6b0b444c23836c4a5d07346082b-Pap er.pdf. Yuxiang Chen, Yifan Liu, Xiaoming Xu, Pengle Zhang, Michael Beyer, Martin Rapp, Jun Zhu, and Jianfei Chen. TetraJet-v2: Accurate NVFP4 training for large language models with oscillation suppression and outlier control. In Forty-third International Conference on Machine Learning, 2026. URL https://openreview.net/forum?id=7ZQhm5HnOA. Hakaze Cho, Yoshihiro Sakai, Kenshiro Tanaka, Mariko Kato, and Naoya Inoue. Understanding token probability encoding in output embeddings. In Owen Rambow, Leo Wanner, Marianna Apidianaki, Hend Al-Khalifa, Barbara Di Eugenio, and Steven Schockaert (eds.), Proceedings of the 31st International Conference on Computational Linguistics, pp. 10618–10633, Abu Dhabi, UAE, January 2025. Association for Computational Linguistics. URL https://aclantho logy.org/2025.coling-main.708/. 10
Preprint.
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try ARC, the AI2 reasoning challenge, 2018. URL https://arxiv.org/abs/1803.05457. Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems, 2021. URL https://arxiv.or g/abs/2110.14168. Xin Ding, Xiaoyu Liu, Zhijun Tu, Yun Zhang, Wei Li, Jie Hu, Hanting Chen, Yehui Tang, Zhiwei Xiong, Baoqun Yin, et al. CBQ: Cross-block quantization for large language models. In International Conference on Learning Representations, volume 2025, pp. 7056–7075, 2025. Ali Edalati, Alireza Ghaffari, Mahsa Ghazvini Nejad, Lu Hou, Boxing Chen, Masoud Asgharian, and Vahid Partovi Nia. OAC: Output-adaptive calibration for accurate post-training quantization. Proceedings of the AAAI Conference on Artificial Intelligence, 39(16):16453–16461, 2025. Matthew Finlayson, Xiang Ren, and Swabha Swayamdipta. Every language model has a forgeryresistant signature. In The Fourteenth International Conference on Learning Representations, 2026. URL https://openreview.net/forum?id=vLFqOoMBol. Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur. Sharpness-aware minimization for efficiently improving generalization, 2021. URL https://arxiv.org/abs/2010 .01412. Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. OPTQ: Accurate quantization for generative pre-trained transformers. In The Eleventh International Conference on Learning Representations, 2023. Gemma Team et al. Gemma 3 technical report, 2025. URL https://arxiv.org/abs/2503 .19786. Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding, 2021. URL https://arxi v.org/abs/2009.03300. Sepp Hochreiter and Jürgen Schmidhuber. Flat minima. Neural Computation, 9(1):1–42, 01 1997. ISSN 0899-7667. doi: 10.1162/neco.1997.9.1.1. URL https://doi.org/10.1162/neco .1997.9.1.1. Yuezhou Hu, Weiyu Huang, Zichen Liang, Chang Chen, Jintao Zhang, Jun Zhu, and Jianfei Chen. Identifying sensitive weights via post-quantization integral, 2025. URL https://arxiv.or g/abs/2503.01901. Jinuk Kim, Marwa El Halabi, Wonpyo Park, Clemens JS Schaefer, Deokjae Lee, Yeonhong Park, Jae W. Lee, and Hyun Oh Song. GuidedQuant: Large language model quantization via exploiting end loss guidance. In Forty-second International Conference on Machine Learning, 2025. URL https://openreview.net/forum?id=ZawsPjlIGu. Jung Hyun Lee, June Yong Yang, Jungwook Choi, and Eunho Yang. LFQ: Logit-aware final-block quantization for boosting the generation quality of low-bit quantized LLMs. In Forty-third International Conference on Machine Learning, 2026. URL https://openreview.net/for um?id=ykdc70h1ND. Yuhang Li, Ruihao Gong, Xu Tan, Yang Yang, Peng Hu, Qi Zhang, Fengwei Yu, Wei Wang, and Shi Gu. BRECQ: Pushing the limit of post-training quantization by block reconstruction. arXiv preprint arXiv:2102.05426, 2021. Yuhang Li, Ruokai Yin, Donghyun Lee, Shiting Xiao, and Priyadarshini Panda. GPTAQ: Efficient finetuning-free quantization for asymmetric calibration. arXiv preprint arXiv:2504.02692, 2025. 11
Preprint.
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han. AWQ: Activation-aware weight quantization for on-device LLM compression and acceleration. Proceedings of Machine Learning and Systems, 6:87–100, 2024. Stephanie Lin, Jacob Hilton, and Owain Evans. TruthfulQA: Measuring how models mimic human falsehoods. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (eds.), Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 3214–3252, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-long.229. URL https://aclanthology.org/2022.acl-l ong.229/. Ruikang Liu, Yuxuan Sun, Manyi Zhang, Haoli Bai, Xianzhi Yu, Tiezheng YU, Chun Yuan, and Lu Hou. Quantization hurts reasoning? an empirical study on quantized reasoning models. In Second Conference on Language Modeling, 2025. URL https://openreview.net/for um?id=BM192Ps5Nv. Sanae Lotfi, Polina Kirichenko, Steven Li, and Zechun Liu. Quantized reasoning models think they need to think longer, but they do not, 2026. URL https://arxiv.org/abs/2606.00206. Haoqian Meng, Yilun Luo, Yafei Zhao, Wenyuan Liu, Peng Zhang, and Xindian Ma. ARCQuant: Boosting NVFP4 quantization with augmented residual channels for LLMs. arXiv preprint arXiv:2601.07475, 2026. Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models. In International Conference on Learning Representations, 2017. URL https://op enreview.net/forum?id=Byj72udxe. Niklas Muennighoff, Luca Soldaini, Dirk Groeneveld, Kyle Lo, Jacob Morrison, Sewon Min, Weijia Shi, Evan Pete Walsh, Oyvind Tafjord, Nathan Lambert, Yuling Gu, Shane Arora, Akshita Bhagia, Dustin Schwenk, David Wadden, Alexander Wettig, Binyuan Hui, Tim Dettmers, Douwe Kiela, Ali Farhadi, Noah A. Smith, Pang Wei Koh, Amanpreet Singh, and Hannaneh Hajishirzi. OLMoE: Open mixture-of-experts language models. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=xXTkbTBmqq. nostalgebraist. Interpreting GPT: The logit lens, 2020. URL https://www.alignmentfor um.org/posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens. NVIDIA, Felix Abecassis, Anjulie Agrusa, Dong Ahn, Jonah Alben, Stefania Alborghetti, Michael Andersch, Sivakumar Arayandi, Alexis Bjorlin, Aaron Blakeman, Evan Briones, Ian Buck, Bryan Catanzaro, Muya Chang, Jinhang Choi, Mike Chrzanowski, Eric Chung, Victor Cui, Steve Dai, Bita Darvish Rouhani, Carlo del Mundo, Deena Donia, Burc Eryilmaz, Henry Estela, Abhinav Goel, Oleg Goncharov, Yugi Guvvala, Robert Hesse, Russell Hewett, Herbert Hum, Ujval Kapasi, Brucek Khailany, Mikail Khona, Nick Knight, Alex Kondratenko, Ronny Krashinsky, Ben Lanir, Simon Layton, Michael Lightstone, Daniel Lo, Paulius Micikevicius, Asit Mishra, Tim Moon, Deepak Narayanan, Chao Ni, Abhijit Paithankar, Satish Pasumarthi, Ankit Patel, Mostofa Patwary, Ashwin Poojary, Gargi Prasad, Sweta Priyadarshi, Yigong Qin, Xiaowei Ren, Oleg Rybakov, Charbel Sakr, Sanjeev Satheesh, Stas Sergienko, Pasha Shamis, Kirthi Shankar, Nishant Sharma, Mohammad Shoeybi, Michael Siu, Misha Smelyanskiy, Darko Stosic, Dusan Stosic, Bor-Yiing Su, Frank Sun, Nima Tajbakhsh, Shelby Thomas, Przemek Tredak, Evgeny Tsykunov, Gandhi Vaithilingam, Aditya Vavre, Rangharajan Venkatesan, Roger Waleffe, Qiyu Wan, Hexin Wang, Mengdi Wang, Lizzie Wei, Hao Wu, Evan Wu, Keith Wyss, Ning Xu, Jinze Xue, Charlene Yang, Yujia Zhai, Ruoxi Zhang, Jingyang Zhu, and Zhongbo Zhu. Pretraining large language models with NVFP4, 2025. URL https://arxiv.org/abs/2509.25149. Andrei Panferov, Erik Schultheis, Soroush Tabesh, and Dan Alistarh. Quartet II: Accurate LLM pre-training in NVFP4 by improved unbiased gradient estimation. In Forty-third International Conference on Machine Learning, 2026. URL https://openreview.net/forum?id= CciWEZZDVb. Arjun Patrawala, Jiahai Feng, Erik Jones, and Jacob Steinhardt. LLM layers immediately correct each other. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. URL https://openreview.net/forum?id=7DY7kB8wyZ. 12
Preprint.
Ofir Press and Lior Wolf. Using the output embedding to improve language models. In Mirella Lapata, Phil Blunsom, and Alexander Koller (eds.), Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers, pp. 157–163, Valencia, Spain, April 2017. Association for Computational Linguistics. URL https: //aclanthology.org/E17-2025/. Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-totext transformer. Journal of Machine Learning Research, 21(140):1–67, 2020. URL http: //jmlr.org/papers/v21/20-074.html. Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. WinoGrande: an adversarial winograd schema challenge at scale. Commun. ACM, 64(9):99–106, August 2021. ISSN 0001-0782. doi: 10.1145/3474381. URL https://doi.org/10.1145/3474381. Wenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu, Lirui Zhao, Zhiqian Li, Kaipeng Zhang, Gao Peng, Yu Qiao, and Ping Luo. OmniQuant: Omnidirectionally calibrated quantization for large language models. In International Conference on Learning Representations, volume 2024, pp. 45472–45496, 2024. Samuel Smith, Erich Elsen, and Soham De. On the generalization benefit of noise in stochastic gradient descent. In Hal Daumé III and Aarti Singh (eds.), Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pp. 9058–9067. PMLR, 13–18 Jul 2020. URL https://proceedings.mlr.press/v1 19/smith20a.html. Team Olmo, Allyson Ettinger, Amanda Bertsch, Bailey Kuehl, David Graham, David Heineman, Dirk Groeneveld, Faeze Brahman, Finbarr Timbers, Hamish Ivison, Jacob Morrison, Jake Poznanski, Kyle Lo, Luca Soldaini, Matt Jordan, Mayee Chen, Michael Noukhovitch, Nathan Lambert, Pete Walsh, Pradeep Dasigi, Robert Berry, Saumya Malik, Saurabh Shah, Scott Geng, Shane Arora, Shashank Gupta, Taira Anderson, Teng Xiao, Tyler Murray, Tyler Romero, Victoria Graf, Akari Asai, Akshita Bhagia, Alexander Wettig, Alisa Liu, Aman Rangapur, Chloe Anastasiades, Costa Huang, Dustin Schwenk, Harsh Trivedi, Ian Magnusson, Jaron Lochner, Jiacheng Liu, Lester James V. Miranda, Maarten Sap, Malia Morgan, Michael Schmitz, Michal Guerquin, Michael Wilson, Regan Huff, Ronan Le Bras, Rui Xin, Rulin Shao, Sam Skjonsberg, Shannon Zejiang Shen, Shuyue Stella Li, Tucker Wilde, Valentina Pyatkin, Will Merrill, Yapei Chang, Yuling Gu, Zhiyuan Zeng, Ashish Sabharwal, Luke Zettlemoyer, Pang Wei Koh, Ali Farhadi, Noah A. Smith, and Hannaneh Hajishirzi. OLMo 3, 2025. URL https: //arxiv.org/abs/2512.13961. Albert Tseng, Zhaofeng Sun, and Christopher De Sa. Model-preserving adaptive rounding. In Fortythird International Conference on Machine Learning, 2026. URL https://openreview.n et/forum?id=PKFilPWjMI. Yusuke Tsuzuku, Issei Sato, and Masashi Sugiyama. Lipschitz-Margin training: Scalable certification of perturbation invariance for deep neural networks. In S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett (eds.), Advances in Neural Information Processing Systems, volume 31. Curran Associates, Inc., 2018. URL https://proceedings.neurip s.cc/paper_files/paper/2018/file/485843481a7edacbfce101ecb1e4d2a 8-Paper.pdf. Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jing Zhou, Jingren Zhou, Junyang Lin, Kai Dang, Keqin Bao, Kexin Yang, Le Yu, Lianghao Deng, Mei Li, Mingfeng Xue, Mingze Li, Pei Zhang, Peng Wang, Qin Zhu, Rui Men, Ruize Gao, Shixuan Liu, Shuang Luo, Tianhao Li, Tianyi Tang, Wenbiao Yin, Xingzhang Ren, Xinyu Wang, Xinyu Zhang, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yinger 13
Preprint.
Zhang, Yu Wan, Yuqiong Liu, Zekun Wang, Zeyu Cui, Zhenru Zhang, Zhipeng Zhou, and Zihan Qiu. Qwen3 technical report, 2025. URL https://arxiv.org/abs/2505.09388. Zhilin Yang, Zihang Dai, Ruslan Salakhutdinov, and William W. Cohen. Breaking the softmax bottleneck: A high-rank RNN language model. In International Conference on Learning Representations, 2018. URL https://openreview.net/forum?id=HkwZSG-CZ. Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. HellaSwag: Can a machine really finish your sentence? In Anna Korhonen, David Traum, and Lluı́s Màrquez (eds.), Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 4791–4800, Florence, Italy, July 2019. Association for Computational Linguistics. doi: 10.18653 /v1/P19-1472. URL https://aclanthology.org/P19-1472/. Manyi Zhang, Ji-Fu Li, Zhongao Sun, Haoli Bai, Hui-Ling Zhen, Zhenhua Dong, and Xianzhi Yu. Benchmarking post-training quantization of large language models under microscaling floating point formats. arXiv preprint arXiv:2601.09555, 2026a. Shihao Zhang, Haoyu Zhang, Ian Colbert, and Rayan Saab. Qronos: Correcting the past by shaping the future... in post-training quantization. In The Fourteenth International Conference on Learning Representations, 2026b. URL https://openreview.net/forum?id=7axclBCY ul.
14
Preprint.
A PPENDIX C ONTENTS A Related Work
15
A.1 PTQ Optimization and Low-Precision Formats . . . . . . . . . . . . . . . . . . .
16
A.2 Transformer Block Responses to Quantization . . . . . . . . . . . . . . . . . . . .
16
A.3 Output Robustness under Quantization . . . . . . . . . . . . . . . . . . . . . . . .
16
B Notation
17
C Proofs and Theoretical Analysis
17
C.1 Hidden-Error Recurrence Proofs . . . . . . . . . . . . . . . . . . . . . . . . . . .
17
C.2 Length–Angle Decomposition . . . . . . . . . . . . . . . . . . . . . . . . . . . .
18
C.3 LM-Head Score Decomposition . . . . . . . . . . . . . . . . . . . . . . . . . . .
18
C.4 High-Dimensional Rotation Analysis . . . . . . . . . . . . . . . . . . . . . . . . .
19
C.5 Log-Probability Approximation Proof . . . . . . . . . . . . . . . . . . . . . . . .
23
C.6 Softmax and Token-Ranking Stability . . . . . . . . . . . . . . . . . . . . . . . .
23
D Experimental Details
25
D.1 NVFP4 Quantization Format . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
25
D.2 NVFP4 Weight Reconstruction Error . . . . . . . . . . . . . . . . . . . . . . . . .
25
D.3 Reproduction Details . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
26
E Additional Experimental Analysis
A
28
E.1 Hidden-Error Growth at Initialization and after Pretraining . . . . . . . . . . . . .
28
E.2 Hidden-Error Recurrence across Models . . . . . . . . . . . . . . . . . . . . . . .
30
E.3 Hidden-Error Analysis: Abnormal Trajectory Filtering . . . . . . . . . . . . . . .
32
E.4 Hidden-Error Growth: Late-Layer Differences across Data . . . . . . . . . . . . .
34
E.5 Counteraction Emergence during Pretraining . . . . . . . . . . . . . . . . . . . .
35
E.6 Counteraction Source: Block-Input-Error Response . . . . . . . . . . . . . . . . .
36
E.7 Counteraction Intervention: Removal and Reversal . . . . . . . . . . . . . . . . .
38
E.8 Counteraction across Weight and Activation Quantization . . . . . . . . . . . . . .
40
E.9 Counteraction across PTQ Algorithms and Weight Formats . . . . . . . . . . . . .
41
E.10 Output Robustness to Quantization across Models . . . . . . . . . . . . . . . . . .
43
E.11 Output Probability Sensitivity to Quantization across Temperatures . . . . . . . . .
45
R ELATED W ORK
We organize related work around three stages of quantization error: how it is introduced, how it propagates through Transformer blocks, and how it affects model outputs. 15
Preprint.
A.1
PTQ O PTIMIZATION AND L OW-P RECISION F ORMATS
Research on weight-only LLM Post-Training Quantization (PTQ) largely frames PTQ as an optimization problem, exemplified by the second-order weight reconstruction of GPTQ (Frantar et al., 2023) and the activation-aware scaling of AWQ (Lin et al., 2024). A broad family of successors extends the calibration objective to learned clipping and equivalent transformations, block-level and cross-block reconstruction, asymmetric calibration against full-precision outputs, alternating error correction, and per-weight loss sensitivity (Shao et al., 2024; Li et al., 2021; Ding et al., 2025; Li et al., 2025; Zhang et al., 2026b; Hu et al., 2025). These methods improve the quantized model by reducing the errors introduced or retained during calibration. On the format side, a broad microscaling benchmark finds that scale construction and compatibility between formats and algorithms are critical at 4-bit precision (Zhang et al., 2026a). NVFP4 has become common in recent large-model inference and training recipes, pairing fine-grained scaling for more accurate 4-bit representation with methods for stable end-to-end low-precision training (Meng et al., 2026; NVIDIA et al., 2025; Chen et al., 2026; Panferov et al., 2026). We use NVFP4 roundto-nearest (RTN) as our main setting (App. D). This lets us study the pretrained model’s response without adding format-specific optimization. A.2
T RANSFORMER B LOCK R ESPONSES TO Q UANTIZATION
Prior work has identified several factors that affect quantization error. Input-dependent failures have been linked to residual magnitudes, late-layer activations, and MLP gates (Chang et al., 2025), while robustness across training trajectories depends strongly on learning-rate dynamics and other training hyperparameters (Catalan-Tatjer et al., 2026). Broader accounts connect parameter-space robustness to training noise and flat minima (Hochreiter & Schmidhuber, 1997; Foret et al., 2021). Related theory studies gradient variance during quantized training (Chen et al., 2020) and forward stability in quantized convolutional and graph residual networks (Ben-Yair et al., 2022). This literature helps explain when and why quantization error becomes large in the forward and backward passes, but not how errors from different Transformer blocks interact. Some PTQ methods address error propagation during optimization. Quantization Error Propagation (QEP) (Arai & Ichikawa, 2025) carries upstream quantization error into layer calibration and adjusts its reconstruction target to compensate for accumulated error. Model-Preserving Adaptive Rounding (Tseng et al., 2026) optimizes rounding against approximate end-to-end output error because local activation error can be a poor proxy for the final distribution. These methods modify the quantization objective to reduce propagated error. Rather than proposing another objective, we study the propagation process itself: how errors introduced by successive Transformer blocks interact and accumulate through depth. Clarifying this mechanism may inspire future PTQ methods. A closer mechanistic parallel comes from full-precision Transformers. Patrawala et al. (2025) show that residual contributions from different layers can oppose one another. Under quantization, we find that the block-update error tends to point against the hidden error already present at the block input (Sec. 3). This negative interaction between errors is much stronger than that between the hidden state and block update. Our decomposition further shows that this counteraction comes from how the quantized block responds to its input error (Sec. 5.2; App. E.6). A.3
O UTPUT ROBUSTNESS UNDER Q UANTIZATION
Output-aware PTQ starts from the observation that hidden-state reconstruction does not necessarily preserve token predictions. These methods incorporate output cross-entropy distortion or end-loss gradients into calibration (Edalati et al., 2025; Kim et al., 2025). Logit-aware Final-block Quantization (LFQ) (Lee et al., 2026) shows that lower final-block hidden-state mean-squared error can worsen token predictions, and calibrates the final decoder block through the full-precision LM head. These methods use output sensitivity to improve quantization, but do not explain how the output probabilities respond to hidden error accumulated through the Transformer blocks. Sequence-level studies examine the downstream effects of quantization on generation. 4-bit weight quantization can retain reasoning performance in many settings (Liu et al., 2025), while more aggressive PTQ lengthens chains of thought and induces failures after correct intermediate answers (Lotfi et al., 2026). These studies measure behavior in the more demanding end-to-end setting, where the 16
Preprint.
original and quantized models follow their own generated histories. To understand the mechanism more clearly, we compare the two models on the same prefix at each next-token prediction and trace how quantization changes the hidden state, token scores, and output probabilities (Sec. 4). At each prediction, the LM head maps the final hidden state to token scores by projecting it onto rows that serve as output word embeddings (Press & Wolf, 2017). Prior work studies the structure of the LM head and the output distributions it can represent (Yang et al., 2018; Cancedda, 2024; Cho et al., 2025; Finlayson et al., 2026). Related work also connects hidden states to outputs through intermediate-state decoding and general perturbation bounds (nostalgebraist, 2020; Belrose et al., 2023; Tsuzuku et al., 2018). These studies characterize the structure of the LM head and general output sensitivity, but do not explain why quantization affects high-ranked token scores less than the rest of the vocabulary. We explain this trend theoretically in Sec. 4.2.
B
N OTATION
Symbol
Meaning
Dimensions and indices d; L; ℓ V πr ∈ V TopK (z) ⊆ V
hidden dimension; number of decoder blocks; block index token vocabulary with |V| tokens token at rank r of the BF16 logits z at one position the K tokens with the largest entries of z; TopK (z) = {π1 , . . . , πK }
Models and passes c M; M Q (W) b (ℓ) F(ℓ) ; F
original (BF16) model; quantized (NVFP4) model NVFP4 quantize–dequantize reconstruction of a weight matrix W BF16/W4 block maps (Attention and MLP of block ℓ)
Residual stream states and errors b (ℓ) ; u(ℓ) , u b (ℓ) h(ℓ) , h (ℓ) (ℓ) ∆h ; ∆u R(ℓ) Tadd , Tinter , Talign
Output layer b LM hLM , h WLM ∈ R|V|×d ; w⊤ k b z, b z; p, p ρ α θk ; θbk b) g ∈ V; ∆CE; KL(p ∥ p Flip@1; Ret@K
BF16/W4 hidden states and block updates; h(ℓ) = h(ℓ−1) + u(ℓ) b (ℓ) − h(ℓ) (the hidden error entering block ℓ+1); blockhidden error h b update error u(ℓ) − u(ℓ) relative hidden-error norm ∥∆h(ℓ) ∥2 /∥h(ℓ) ∥2 the three terms of the recurrence in Thm. 1: relative added-error contribution, error interaction (named counteraction when negative), and the residual-norm contribution from the clean input–update dot product BF16/W4 LM-head inputs, hLM = Norm(h(L) ) shared BF16 LM head; its row for token k BF16/W4 logits z = WLM hLM and next-token distributions b LM ∥2 /∥hLM ∥2 LM-head input-norm ratio ∥h b hidden rotation ∠(hLM , hLM ) b LM ) projection angles ∠(wk , hLM ) and ∠(wk , h ground-truth next token; CE increase CEW4 − CEBF16 ; forward KL greedy flip rate Pr[Top1 (b z) ̸= Top1 (z)] and mean top-K retention 1 E[ K | TopK (z) ∩ TopK (b z)|]
C
P ROOFS AND T HEORETICAL A NALYSIS
C.1
H IDDEN -E RROR R ECURRENCE P ROOFS
Proof of Prop. 1. Subtracting the original residual update from the quantized residual update in Eq. (1) gives ∆h(ℓ) = ∆h(ℓ−1) + ∆u(ℓ) . Therefore, ∥∆h(ℓ) ∥22 = ∥∆h(ℓ−1) + ∆u(ℓ) ∥22 = ∥∆h(ℓ−1) ∥22 + ∥∆u(ℓ) ∥22 + 2⟨∆h(ℓ−1) , ∆u(ℓ) ⟩. Subtracting ∥∆h(ℓ−1) ∥22 from both sides proves Eq. (2). 17
Preprint.
Proof of Thm. 1. By Prop. 1, 2 ∥∆h(ℓ−1) ∥22 + ∥∆u(ℓ) ∥22 + 2⟨∆h(ℓ−1) , ∆u(ℓ) ⟩ R(ℓ) = . ∥h(ℓ) ∥22 The definition of R(ℓ−1) gives 2 ∥∆h(ℓ−1) ∥22 = ∥h(ℓ−1) ∥22 R(ℓ−1) . 2 Substituting this identity and subtracting R(ℓ−1) yields R
(ℓ) 2
− R
(ℓ−1) 2
=
∥∆u(ℓ) ∥22 + 2⟨∆h(ℓ−1) , ∆u(ℓ) ⟩ ∥h(ℓ) ∥22
+
∥h(ℓ−1) ∥22 ∥h(ℓ) ∥22
! −1
2 R(ℓ−1) .
The original residual update h(ℓ) = h(ℓ−1) + u(ℓ) implies ∥h(ℓ) ∥22 = ∥h(ℓ−1) ∥22 + ∥u(ℓ) ∥22 + 2⟨h(ℓ−1) , u(ℓ) ⟩. Therefore, ∥h(ℓ−1) ∥22 ∥h(ℓ) ∥22
−1=−
∥u(ℓ) ∥22 + 2⟨h(ℓ−1) , u(ℓ) ⟩ ∥h(ℓ) ∥22
.
Substitution gives R
(ℓ) 2
− R
(ℓ−1) 2
=
∥∆u(ℓ) ∥22 − ∥u(ℓ) ∥22 R(ℓ−1)
2
∥h(ℓ) ∥22 +
2⟨∆h(ℓ−1) , ∆u(ℓ) ⟩ ∥h(ℓ) ∥22
−
2⟨h(ℓ−1) , u(ℓ) ⟩ ∥h(ℓ) ∥22
2 R(ℓ−1) .
Finally, the first fraction factors as ∥∆u(ℓ) ∥22 − ∥u(ℓ) ∥22 R(ℓ−1)
2
∥u(ℓ) ∥2
!2 "
∥∆u(ℓ) ∥2 ∥u(ℓ) ∥2
2
= − R ∥h(ℓ) ∥22 ∥h(ℓ) ∥2 The three lines are Tadd , Tinter , and Talign , respectively, which proves Eq. (3). C.2
# (ℓ−1) 2
.
L ENGTH –A NGLE D ECOMPOSITION
b (ℓ) ⟩ = ∥h(ℓ) ∥2 ∥h b (ℓ) ∥2 cos ∠ h(ℓ) , h b (ℓ) . Proof of Prop. 2. By the definition of cosine, ⟨h(ℓ) , h Therefore, b (ℓ) ∥2 + ∥h(ℓ) ∥2 − 2⟨h(ℓ) , h b (ℓ) ⟩ ∥∆h(ℓ) ∥22 = ∥h 2 2 2 b (ℓ) ∥2 − ∥h(ℓ) ∥2 + 2∥h(ℓ) ∥2 ∥h b (ℓ) ∥2 1 − cos ∠ h(ℓ) , h b (ℓ) . = ∥h b (ℓ) ∥2 = ∥h(ℓ) ∥2 gives the equal-length identity Dividing by ∥h(ℓ) ∥22 proves Eq. (4). Substituting ∥h stated in the proposition. C.3
LM-H EAD S CORE D ECOMPOSITION
Sec. 4.2 studies how LM-head input rotation changes token scores. The following exact decomposition separates score changes caused by the input norm from those caused by the projection angle and clarifies what the ρ = 1 analytic references omit. With ρ, θk , and θbk as in Thm. 2, each vocabulary score change splits exactly into a length part and an angle part: zbk − zk = (ρ − 1)zk + ρ∥wk ∥2 ∥hLM ∥2 (cos θbk − cos θk ). (9) Proof of Eq. (9). For each nonzero vocabulary row wk , zk = ∥wk ∥2 ∥hLM ∥2 cos θk ,
b LM ∥2 cos θbk . zbk = ∥wk ∥2 ∥h b LM ∥2 cos θk gives Eq. (9). If Subtracting the two equalities and adding and subtracting ∥wk ∥2 ∥h b hLM = chLM with c > 0, then b z = cz, so every pairwise score ordering is unchanged. 18
Preprint.
C.4
H IGH -D IMENSIONAL ROTATION A NALYSIS
Proof of Thm. 2. We first characterize the geometry of the LM-head input rotation relative to a fixed LM-head weight vector. Geometry of the hidden-state rotation. Let h̄ :=
hLM , ∥hLM ∥2
w̄k :=
wk , ∥wk ∥2
T
P⊥ := I − h̄h̄ .
Because θk ∈ (0, π), the component of w̄k orthogonal to h̄ is nonzero. Define its unit direction by ξk :=
P⊥ w k . ∥P⊥ wk ∥2
Then w̄k decomposes into components parallel and orthogonal to h̄: w̄k = cos θk h̄ + sin θk ξk . b LM ) ∈ (0, π), the normalized quantized LM-head input can be Similarly, because α = ∠(hLM , h written as b LM h = cos α h̄ + sin α η, b LM ∥2 ∥h where η :=
b LM P⊥ h b LM ∥2 ∥P⊥ h
is a unit vector orthogonal to h̄. Define the directional coordinate qk := ⟨η, ξk ⟩. Thus, qk measures how much the hidden-state rotation points toward the component of wk orthogonal to the original hidden state. Taking the inner product between the decompositions of w̄k and b LM /∥h b LM ∥2 gives the exact identity h cos θbk = cos θk cos α + sin θk sin α qk .
(10)
Isotropic rotation direction and the distribution of qk . Conditional on the measured rotation angle α, the theorem assumes that the rotation direction η is uniformly distributed over all unit directions orthogonal to hLM . Equivalently, η is uniform on the unit sphere Sd−2 in the (d − 1)-dimensional subspace h⊥ LM . A convenient way to represent such a uniform direction is to normalize an isotropic Gaussian vector. Choose an orthonormal basis of h⊥ LM whose first basis vector is ξk , and let g = (G1 , . . . , Gd−1 )T ,
i.i.d.
Gj ∼ N (0, 1).
Because the standard Gaussian distribution is rotationally invariant, g/∥g∥2 is uniformly distributed on Sd−2 . Hence we may represent η in this basis as g d . η= ∥g∥2 d
where = denotes equality in distribution. This does not mean that a uniform spherical direction is Gaussian: the Gaussian vector g has a random norm, while g/∥g∥2 has unit norm and only its direction is retained. Since the first basis vector is ξk , the directional coordinate qk := ⟨η, ξk ⟩ is distributed as the first coordinate of this normalized Gaussian vector: G1 d qk = q . G21 + · · · + G2d−1 19
Preprint.
Therefore, d
qk2 =
G21 +
G21 Pd−1
2 j=2 Gj
.
Now d−1 X
G21 ∼ χ21 ,
G2j ∼ χ2d−2 ,
j=2
and these two random variables are independent. For independent X ∼ χ2m and Y ∼ χ2n , X/(X +Y ) ∼ Beta(m/2, n/2). Applying this with m = 1 and n = d − 2 gives 1 d−2 2 qk ∼ Beta , . 2 2 Moreover, qk is symmetric about zero, so E[qk ] = 0. Equivalently, qk has density Γ((d − 1)/2) (1 − q 2 )(d−4)/2 , fd (q) = √ π Γ((d − 2)/2)
−1 < q < 1.
Its mean absolute value is Z 1 Γ((d − 1)/2) √ E|qk | = 2 q(1 − q 2 )(d−4)/2 dq π Γ((d − 2)/2) 0 Γ((d − 1)/2) = √ =: µd . π Γ(d/2) The standard gamma-ratio expansion gives s µd =
2 1 + O(d−1 ) . π(d − 1)
Thus, qk is the coordinate of a random unit direction along one fixed direction in a (d − 1)dimensional space, and its typical magnitude is O(d−1/2 ). This is the source of the high-dimensional attenuation in Eq. (5). Projection-angle change. For fixed qk , define gk (α) := cos θk cos α + sin θk sin α qk , At α = 0,
θbk = arccos gk (α).
gk′ (0) = sin θk qk .
gk (0) = cos θk , Therefore, dθbk dα
g ′ (0) sin θk qk = −qk , = −p k = −√ 2 1 − cos2 θk 1 − g (0) k α=0
where the last equality uses θk ∈ (0, π) and hence sin θk > 0. For fixed θk ∈ (0, π), the denominator above remains bounded away from zero for sufficiently small α, uniformly over qk ∈ [−1, 1]. Thus the Taylor remainder can be taken uniformly in qk , and θbk − θk = −αqk + O(α2 ). Using |θbk − θk | − α|qk | = O(α2 ) uniformly in qk , taking expectation gives E θbk − θk = αµd + O(α2 ). Together with the expression for µd above, this proves Eq. (5). 20
Preprint.
Relative score change. The original and quantized scores are zk = ∥wk ∥2 ∥hLM ∥2 cos θk ,
b LM ∥2 cos θbk . zbk = ∥wk ∥2 ∥h
Using ρ :=
b LM ∥2 ∥h , ∥hLM ∥2
and substituting Eq. (10), zbk cos θbk =ρ zk cos θk = ρ cos α + ρ tan θk sin α qk . Subtracting one gives the exact decomposition zbk − zk = (ρ cos α − 1) + ρ tan θk sin α qk , zk which proves Eq. (6). When ρ = 1, zbk − zk = (cos α − 1) + tan θk sin α qk . zk Since cos α − 1 = O(α2 ),
sin α = α + O(α3 ),
we obtain, uniformly over qk ∈ [−1, 1], zbk − zk = α tan θk qk + O(α2 ). zk Taking the expected absolute value therefore gives E
zbk − zk = α|tan θk |µd + O(α2 ), zk
which proves Eq. (7). Finite-angle expectation. The first-order result above keeps only the leading term in α. For ρ = 1, we can also compute the expected absolute relative score change exactly at a finite rotation angle. Let q denote a random variable with density fd above, and for A, B ≥ 0 define Ψd (A, B) := Eq [|−A + Bq|] .
(11)
Thus, Ψd (A, B) is simply the expected absolute value of a random directional term Bq shifted by the deterministic quantity A. Let Ix (a, b) :=
Bx (a, b) , B(a, b)
Z x Bx (a, b) :=
ta−1 (1 − t)b−1 dt,
0
denote the regularized incomplete beta function, with B(a, b) := B1 (a, b). We now derive a closed form for Ψd . If B = 0 or A ≥ B, then for all q ∈ [−1, 1], −A + Bq ≤ −A + B ≤ 0. Hence Ψd (A, B) = Eq [A − Bq] = A, because Eq [q] = 0. Now suppose 0 ≤ A < B, and define t :=
A ∈ [0, 1). B 21
Preprint.
The quantity −A + Bq changes sign at q = t. Splitting the expectation at this point and using the symmetry fd (q) = fd (−q) gives Ψd (A, B) = Eq | − A + Bq| Z 1 = A Pr(|q| ≤ t) + 2B
qfd (q) dq. t
Because
1 d−2 q ∼ Beta , , 2 2 2
the first term is 2
2
Pr(|q| ≤ t) = Pr(q ≤ t ) = It2
1 d−2 , 2 2
.
For the second term, Z 1 Z 1 Γ((d − 1)/2) q(1 − q 2 )(d−4)/2 dq = µd (1 − t2 )(d−2)/2 . 2 qfd (q) dq = 2 √ π Γ((d − 2)/2) t t Substituting t = A/B gives ( A, Ψd (A, B) = (d−2)/2 + Bµd 1 − (A/B)2 AI(A/B)2 21 , d−2 , 2
B = 0 or A ≥ B, (12) 0 ≤ A < B.
Returning to the relative score error with ρ = 1, zbk − zk = −(1 − cos α) + tan θk sin α qk . zk Since qk is symmetric about zero, the sign of the coefficient multiplying qk does not affect the expected absolute value. Therefore, setting A := 1 − cos α,
B := |tan θk sin α|
gives the exact finite-angle expression E
zbk − zk = Ψd (1 − cos α, |tan θk sin α|) . zk
Finally, 1 − cos α =
α2 + O(α4 ), 2
sin α = α + O(α3 ).
Moreover, since |q| ≤ 1, |Ψd (A, B) − Ψd (A′ , B ′ )| ≤ |A − A′ | + |B − B ′ |. Thus replacing the exact arguments by their second-order approximations changes Ψd by at most O(α3 ), yielding 2 zbk − zk α E = Ψd , α|tan θk | + O(α3 ). (13) zk 2
22
Preprint.
C.5
L OG -P ROBABILITY A PPROXIMATION P ROOF P b := P ezj +∆zj . Since pj = ezj /Z, Proof of Thm. 3. Let Z := j ezj and Z j X Zb = Z pj e∆zj = ZEj∼p e∆zj . j
It follows that pbk =
pk e∆zk , Ej∼p e∆zj
∆ log pk = ∆zk − log Ej∼p e∆zj .
Also, b ) = Ek∼p log KL(p ∥ p
pk = log Ej∼p e∆zj − Ej∼p ∆zj . pbk
Combining these two identities proves the exact line of Eq. (8). Set m := Ej∼p ∆zj ,
xj := ∆zj − m,
Ej∼p xj = 0.
Then log Ej∼p e∆zj = m + log Ej∼p exj . For uniformly small xj , Taylor expansion gives 1 Ej∼p exj = 1 + Ej∼p x2j + O Ej∼p |xj |3 , 2 1 log Ej∼p exj = Ej∼p x2j + O Ej∼p |xj |3 2 1 = Varj∼p (∆zj ) + O Ej∼p |xj |3 . 2 Substituting into the exact identity yields ∆ log pk = ∆zk − m −
1 Varj∼p (∆zj ) + O Ej∼p |xj |3 , 2
which is the second-order estimate. Omitting the quadratic term gives ∆ log pk = ∆zk − m + O(Varj∼p (∆zj )) , which is the first-order estimate. C.6
S OFTMAX AND T OKEN -R ANKING S TABILITY
This appendix gives the exact conditions under which a token’s ranking is preserved or changed, complementing the estimates by rank in Sec. 4.2. Theorem 4 (Probability-change identities and token-preservation conditions). Let ∆z := b z − z, b := softmax(b p := softmax(z), and p z), with |V| ≥ 2. Then b ) = log Ek∼p e∆zk − Ek∼p ∆zk . KL(p ∥ p
(14)
For a ground-truth next token g, − log pbg + log pg = log Ek∼p e∆zk − ∆zg .
(15)
Adding the same scalar to every ∆zk leaves these quantities unchanged. If a is the unique top-1 token under z, it remains top-1 exactly when za − zj > ∆zj − ∆za
for every j ̸= a.
(16)
for every a ∈ SK , j ∈ / SK .
(17)
The analogous condition for a unique top-K set SK is za − zj > ∆zj − ∆za
23
Preprint.
Proof of Thm. 4. Let Z :=
P
zj +∆zj
b = Z Ej∼p e∆zj , Z
pbk =
P
je
zj
b := and Z
je
. Then pk e∆zk . Ej∼p e∆zj
Therefore,
pk = log Ej∼p e∆zj − ∆zk . pbk Taking expectation over k ∼ p proves Eq. (14); setting k = g proves Eq. (15). Replacing every ∆zk by ∆zk + c adds c to both terms on the right-hand sides of Eqs. (14) and (15), so the additions cancel. log
The original top-1 token a remains the unique perturbed top-1 token exactly when za + ∆za > zj + ∆zj
for every j ̸= a.
Rearranging gives Eq. (16). Likewise, SK remains exactly the top-K set if and only if every member stays above every nonmember: za + ∆za > zj + ∆zj
for every a ∈ SK , j ∈ / SK .
Rearranging gives Eq. (17). For Qwen3-32B, Flip@1 is 10.7% on C4, 12.7% on WikiText-103, and 8.3% on GSM8K text under NVFP4 (Fig. 4(A)). Under the theorem’s assumption that the original top-1 token a is unique, each flip implies ∆zj − ∆za ≥ za − zj for at least one competitor j ̸= a, where za − zj > 0 is its original score margin to the top-1 token. Thus, Flip@1 measures how often quantization changes relative token scores enough to overturn the original top-1 prediction, rather than how often any score changes.
24
Preprint.
D
E XPERIMENTAL D ETAILS
D.1
NVFP4 Q UANTIZATION F ORMAT
NVFP4 represents each quantized value by a 4-bit E2M1 floating-point code and two levels of shared scales (Alvarez et al., 2025). The signed E2M1 value set is QE2M1 = {0, ±0.5, ±1, ±1.5, ±2, ±3, ±4, ±6}. Thus, the largest E2M1 magnitude is qmax = 6. Each weight matrix W ∈ Rdout ×din is partitioned row by row into consecutive groups of 16 entries along its input dimension. If necessary, the final group is zero-padded for quantization and the padding is removed afterward. We denote these groups by {Gb }b , and each group Gb contains 16 entries with shape 1 × 16. The quantization of one weight matrix proceeds as follows. First, we compute a tensor-level scale sW =
maxi,j |Wij | maxi,j |Wij | = , qmax fmax 6 × 448
(18)
where fmax = 448 is the largest finite magnitude of the E4M3 scale format. For each 16-entry group Gb , we then compute the ideal local scale from the tensor-normalized weights and store it in E4M3: seb =
1 Wij , max qmax (i,j)∈Gb sW
sb = castE4M3 (e sb ) .
The factor 6 × 448 in Eq. (18) ensures that the required local scales lie within the E4M3 range. Finally, each weight is divided by both scales, rounded to its nearest E2M1 value, and dequantized: Wij qij = RTNQE2M1 , Q (W)ij = sW sb qij . (i, j) ∈ Gb . (19) sW sb Here, round-to-nearest (RTN) selects the closest E2M1 value, with values outside its range clipped to [−6, 6]. An all-zero tensor or group is mapped to zero, avoiding division by a zero scale. We apply RTN NVFP4 to the linear weights in every Transformer block, while embeddings, normalization layers, activations, and the LM head remain in their original precision. For MoE layers, this includes every expert’s gate, up, and down projections but not the router. We simulate quantization by casting the reconstructed weights Q (W) in Eq. (19) back to the computation dtype before each linear operation. The experiments therefore measure the numerical effect of quantization, not the runtime or memory cost of a packed 4-bit kernel. D.2
NVFP4 W EIGHT R ECONSTRUCTION E RROR
We compare the NVFP4 reconstruction errors of pretrained and randomly initialized models directly. The Frobenius cosine-similarity between a weight matrix and its reconstruction is defined ⟨W,Q(W)⟩F . Across all models, pretrained and randomly initialized weights both have reconas ∥W∥ F ∥Q(W)∥F struction cosines of about 0.9955 and relative errors of about 9.5%. These nearly identical matrixlevel errors contrast with their different hidden-error growth in Sec. 3.1. Thus, weight reconstruction alone does not explain the slower growth after pretraining. Table 1: Per-matrix reconstruction error under direct-cast NVFP4. Values are means ± SD. Relative error is ∥Q (W) − W∥F /∥W∥F . Random-init columns use one independently initialized instance per model; all reported SDs are computed across weight matrices. Pretrained
Random initialization
Model
cos(W, Q (W))
Relative error
cos(W, Q (W))
Relative error
Qwen3-4B Qwen3-8B Qwen3-14B Qwen3-32B Pythia-1.4B OLMo3-7B
0.99550 ± 0.00001 0.99551 ± 0.00003 0.99551 ± 0.00004 0.99551 ± 0.00003 0.99549 ± 0.00002 0.99549 ± 0.00002
9.49% ± 0.02% 9.47% ± 0.03% 9.47% ± 0.04% 9.48% ± 0.04% 9.50% ± 0.02% 9.50% ± 0.02%
0.99548 ± 0.00001 0.99548 ± 0.00001 0.99548 ± 0.00001 0.99547 ± 0.00001 0.99548 ± 0.00001 0.99548 ± 0.00001
9.51% ± 0.01% 9.51% ± 0.01% 9.51% ± 0.01% 9.51% ± 0.01% 9.51% ± 0.01% 9.51% ± 0.01%
25
Preprint.
D.3
R EPRODUCTION D ETAILS
Datasets. The token-level experiments use C4 (Raffel et al., 2020), WikiText-103 (Merity et al., 2017), and GSM8K text (Cobbe et al., 2021). Unless noted otherwise, each dataset setting contains 64 randomly sampled sequences with 512 next-token prediction positions per sequence. The crossmodel recurrence accounting in Figs. 10 and 11 instead uses 128 sequences per model and dataset. The PTQ comparison uses eight C4 evaluation inputs, while the temperature replay uses eight inputs from each of C4 and GSM8K text. For GSM8K text, we join each official test-set question with its reference solution and use teacher forcing: the model predicts each reference token from the preceding reference text. It does not generate a complete solution, so this setting does not measure generated-solution accuracy. Benchmarks. We use the full test splits of ARC-Challenge and ARC-Easy (Clark et al., 2018) and MMLU (Hendrycks et al., 2021), and the full validation splits of HellaSwag (Zellers et al., 2019), WinoGrande (Sakaguchi et al., 2021), and TruthfulQA (Lin et al., 2022). All six benchmarks are evaluated zero-shot without a chat template. ARC and HellaSwag use character-normalized candidate likelihoods, MMLU scores answer letters, WinoGrande scores the shared suffix conditioned on each candidate, and TruthfulQA reports MC1 accuracy. These results appear in Fig. 4. Models. We evaluate Qwen3-4B, Qwen3-8B, Qwen3-14B, Qwen3-30B-A3B, and Qwen332B (Yang et al., 2025); OLMo3-7B and OLMo3-32B (Team Olmo et al., 2025); OLMoE-1B7B (Muennighoff et al., 2025); Gemma3-4B (Gemma Team et al., 2025); and Pythia-1.4B and Pythia-2.8B (Biderman et al., 2023). Qwen3-32B is the primary model for the hidden-state and output analyses. Fig. 1 compares its pretrained checkpoint with ten independent random initializations and also compares six released OLMo3-7B Stage-1 checkpoints. The random reference in Fig. 5 uses three fixed, seeded random initializations of Qwen3-32B. Post-training weight quantization. Unless stated otherwise, we use the weight-only NVFP4 RTN conversion defined in App. D. Attention and MLP projection weights are quantized, while embeddings, normalization layers, activations, and the LM head remain in their original precision. The counteraction intervention protocol is given in App. E.7. For the calibrated PTQ comparison, GPTQ and AWQ use the same 64 C4 calibration inputs, disjoint from the eight evaluation inputs; RTN uses no calibration data. Activation and activation-weight joint quantization. For A4, the weights remain in their original precision, while the input to each decoder Attention and MLP projection is quantized immediately before matrix multiplication. We use the same E2M1 values and E4M3 block scales as for W4. Each activation tensor is partitioned into contiguous groups of 16 values along its last dimension. The tensor-wide scale is recomputed for every forward pass, followed by one E4M3 scale per group, making activation quantization dynamic and calibration-free. Residual states, normalization layers, embeddings, the KV cache, and the LM head remain in their original precision. W4A4 combines this activation quantization with the W4 conversion defined in App. D. Evaluation. Unless specified below, statistics are first averaged over positions within each input and then equally across inputs. Error bars are one sample standard deviation across complete inputs. For random models, each combination of an initialization and input is one sample. For the relative-error recurrence analysis in Fig. 2, we form each term at every position before aggregation. Position zero is the only prediction conditioned on a one-token prefix, so we exclude this fixed-window boundary and report the recurrence over within-sequence positions. Rare large differences between the BF16 and W4 hidden-state norms produce extreme terms that can dominate the aggregate, so we retain a complete position trajectory only when its maximum relative hiddennorm difference across depth satisfies (ℓ)
(ℓ)
b ∥2 − ∥h ∥2 ∥h i i gimax := max ℓ
. (ℓ) ∥hi ∥2 We retain position i if gimax ≤ 0.5 and use the same retained positions at every block. This removes 0.08% of Qwen3-32B C4 positions and retains 99.20–99.97% across the 15 cross-model and cross26
Preprint.
dataset settings. Each plotted standard deviation is computed separately for one recurrence term and therefore need not satisfy the additive recurrence. App. E.2 reports the cross-model results, and App. E.3 reports the unfiltered and threshold-sensitivity checks. Fig. 3(B) forms the terms in Prop. 2 at each prediction position before averaging. The LM-headinput region applies the same decomposition after final normalization and uses the same retained positions. Fig. 4(A) evaluates all 512 positions in each input. Its benchmark error bars come from 10,000 bootstrap estimates of accuracy. Fig. 5(A) uses the same complete inputs and reports the mean angle between the BF16 and W4 LMhead inputs together with the vocabulary-mean absolute change in their projection angles to the same LM-head weight vectors. For Fig. 5(B) and Fig. 5(C), πr is the token ranked r by the BF16 scores. Fig. 5(B) retains ranks 1–10 individually and groups the rest of the vocabulary into contiguous logspaced intervals. It includes a position–rank pair only when its BF16 score and FP32-recomputed projection ⟨wπr , hLM ⟩ are both positive. The observed absolute relative score changes are averaged first within each complete input and then equally across inputs. The pretrained band shows one log-SD across the input-level means. For the random reference, each initialization–input pair is one sample, and the curve ends when a rank bin contains no positive BF16 scores. The dotted first-order and dashed second-order curves evaluate Eqs. (7) and (13) at each retained pair’s observed hidden rotation α and projection angle θπr under ρ = 1, then apply the same averaging. These analytic references use the conditional expectation over the rotation direction η (App. C.4); neither is fitted to the W4 measurements. Fig. 5(C) retains all vocabulary rows. Its three curves average the absolute values of the exact, firstorder, and second-order expressions in Thm. 3, first within each complete input and then equally across the three datasets. The random reference uses the same aggregation over initialization–input pairs. Bars show one log-SD across the corresponding input-level means.
27
Preprint.
E
A DDITIONAL E XPERIMENTAL A NALYSIS
E.1
H IDDEN -E RROR G ROWTH AT I NITIALIZATION AND AFTER P RETRAINING
Sec. 3.1 describes the growth of absolute and relative hidden error at random initialization and contrasts it with pretrained models. This subsection derives the random-initialization fit from the recurrence, checks it across architectures, and quantifies the slower growth after pretraining. Hidden error in randomly initialized models We derive the randomly initialized Qwen3-32B growth rate from Fig. 1. For random initialization, expectations first average positions within each sequence and then average equally over initialization–sequence pairs; the pretrained curve uses the same sequence-level averaging. Taking expectations in Prop. 1 gives E∥∆h(ℓ) ∥22 − E∥∆h(ℓ−1) ∥22 = E∥∆u(ℓ) ∥22 + 2E⟨∆h(ℓ−1) , ∆u(ℓ) ⟩.
(20)
Across the 64 measured blocks, the ratio of the left-hand side of Eq. (20) to E∥∆u(ℓ) ∥22 lies in [0.9989, 1.0003]. Thus, over this measured depth range, E∥∆h(L) ∥22 ≈
L X
E∥∆u(ℓ) ∥22 .
(21)
ℓ=1
The mean squared block-update error E∥∆u(ℓ) ∥22 does not follow a single power over blocks 1–64. A three-parameter saturation curve instead gives ℓβ E∥∆u(ℓ) ∥22 ≈ A β , τ + ℓβ
A = 1.7154 × 105 ,
β = 1.143,
τ = 13.31,
with log-scale R2 = 0.99931 and original-scale R2 = 0.99956. Substituting this fit into Eq. (21): #1/2 " L q X ℓβ (L) 2 E∥∆h ∥2 ≈ . A β τ + ℓβ ℓ=1
This predicted curve matches the measured root-mean-square (RMS) hidden-error magnitude with R2 = 0.999981 and differs by 0.014% at block 64. Separately, a power fit over the measured depth range gives exponent 0.801 with log–log R2 = 0.997 and a standard deviation of 0.0014 across trajectory fits for each initialization and sequence; the mean E∥∆h(ℓ) ∥2 reported in the main text has the same fitted exponent. Fig. 7(C) directly fits the mean relative error, yielding exponent 0.302 with log–log R2 = 0.984. These fits support the main-text ℓ0.80 and ℓ0.30 descriptions over the measured 64 blocks, without asserting asymptotic growth laws. ×105
×103 Observed ‖Δu(ℓ)‖22
Fit cℓ
2.4
γabs
√j≤ℓ
[‖Δh(ℓ)‖2/‖h(ℓ)‖2]
0.8
1.6
0.8
0.4
Observed R (ℓ)
Fit cℓ γ (γ = 0.302)
(γabs = 0.801)
∑ Aj β/(τ β + j β)
√ ‖Δh(ℓ) ‖2 2
‖Δu(ℓ)‖22
1.2
Observed √ ‖Δh(ℓ)‖22
Fit Aℓ β/(τ β + ℓ β)
1.2
Qwen3-32B
1.0
0.8
0.6
0.4
β = 1.143, τ = 13.3, R 2 = 0.99956
0.0 1
16
32
48
Decoder block ℓ
64
(A) Mean squared block-update error
1
16
32
48
Decoder block ℓ
(B) RMS hidden error
64
1
16
32
48
Decoder block ℓ
64
(C) Relative hidden error
Figure 7: Fits to block-update and hidden-error growth in randomly initialized Qwen3-32B. (A) A three-parameter curve fits E∥∆u(ℓ) ∥22 over the measured blocks. (B) Summing these fitted values closely matches the measured RMS hidden error, whose growth is also fitted by a power law. (C) A separate power law fits the mean relative hidden error. Error bars show standard deviation. 28
Preprint.
The same fit form for E∥∆u(ℓ) ∥22 also describes randomly initialized Qwen3-4B, Qwen3-8B, and OLMo3-7B Stage-1 step 0 (Fig. 8); across the four models the hidden-error exponents fitted over the measured depth range are 0.70–0.83, the corresponding mean-relative-error exponents are 0.20– 0.33, and the recurrence-predicted final RMS errors match the measured values to within 0.02%. Adding an unconstrained intercept to the power fit yields negative intercepts, violating the boundary ∆h(0) = 0; we therefore report these as descriptive fits over the measured intervals, not growth laws.
Qwen3-4B
0.8 0.6 0.4 Observed ‖Δu(ℓ)‖22
0.2
Aℓ β/(τ β + ℓ β)
0.0 1
8
16
24
Decoder block ℓ
32 36
(A) Mean squared block-update error
Qwen3-8B
OLMo3-7B Stage-1 step 0
1.0
1.0
0.8
0.8
γ = 0.33
[‖Δh(ℓ)‖2/‖h(ℓ)‖2]
1.0
Fraction of final RMS hidden error
Normalized mean squared block-update error
Qwen3-4B (Rand. Init.) / Qwen3-8B (Rand. Init.) / OLMo3-7B (Stage-1 step 0)
0.6 0.4
Observed √ ‖Δh(ℓ)‖22
0.2
∑ Aj β/(τ β + j β) √j≤ℓ
0.0
γ = 0.27
0.6 0.4
γ = 0.20
0.2
Observed R (ℓ) Fit cℓ γ
0.0 1
8
16
24
Decoder block ℓ
32 36
1
8
16
24
Decoder block ℓ (C) Relative hidden error
(B) RMS hidden error
32 36
Figure 8: Fits to block-update and hidden-error growth across randomly initialized models. Results in (A) and (B) are normalized by their final values for comparison across Qwen3-4B, Qwen3-8B, and OLMo3-7B Stage-1 step 0. (A) Fitted curves follow the measured E∥∆u(ℓ) ∥22 . (B) Summing the fitted values closely matches the measured RMS hidden error. (C) Power laws fit the unnormalized mean relative hidden error. All fits are restricted to the measured blocks. Hidden-error growth in pretrained models We do not assign the pretrained curve a global power exponent because its growth is not uniform across depth. We quantify its rate over intervals. Define the average per-block relative-error increase over blocks a + 1, . . . , b as h i h i E ∥∆h(b) ∥2 /∥h(b) ∥2 − E ∥∆h(a) ∥2 /∥h(a) ∥2 srel . a:b := b−a Over blocks 33–64, s33:64 is 1.46 × 10−3 for pretrained Qwen3-32B and 4.45 × 10−3 at random initialization. Thus, the measured second-half relative-error growth is 3.1× slower after pretraining. At block 64, the mean hidden-error magnitude and mean relative hidden error are 5.5× and 6.7× smaller, respectively; these final-block ratios describe the final errors rather than their growth rates. Fig. 9 compares randomly initialized and pretrained models across three architecture families. At the final block, random initialization increases relative hidden error by 7.9× in Qwen3-8B, 5.9× in OLMo3-32B, and 4.0× in Gemma3-4B.
Relative error
Random Init.
Qwen3-8B
1.0
0.5
0.0
0
18
Decoder block
Pretrained
OLMo-3-32B
36
Gemma3-4B
0.50
0.4
0.25
0.2
0.00
0
32
Decoder block
64
0.0
0
17
34
Decoder block
Figure 9: Randomly initialized and pretrained models across three architecture families. Relative hidden errors are shown for Qwen3-8B, OLMo3-32B, and Gemma3-4B under the same NVFP4 conversion. Error bars show standard deviation. Qwen3-8B and OLMo3-32B use BF16 for nonquantized computation; Gemma3 uses FP32 for the original pass and all non-quantized operations. 29
Preprint.
E.2
H IDDEN -E RROR R ECURRENCE ACROSS M ODELS
The exact recurrence in Thm. 1 separates each block’s change in squared relative hidden error into the three terms defined in the main text: relative additional error Tadd , error interaction Tinter , and residual-norm contribution Talign . We test whether the same signed balance holds beyond the Qwen3-32B result in the main text. We evaluate five pretrained dense and mixture-of-experts models on C4, WikiText-103, and GSM8K text. Model precision, quantization, aggregation, and position selection follow App. D.3. In every setting, the negative interaction term Tinter offsets part of Tadd . The residual-norm term Talign gives a further reduction, although the size of each contribution varies by model and dataset. The signs Tadd > 0 and Tinter < 0 remain unchanged across the tested filter thresholds. For the primary Qwen3-32B model, the mean final relative hidden error across the three datasets is 0.245, as reported in Sec. 4. Figs. 10 and 11 show the cumulative and per-block results.
ℓ
(R(ℓ))2
Cumulative relative squared-error accounting
(j)
∑ Tinter
j=1 ℓ
ℓ
(j)
∑ Tadd
(j)
∑ Talign
j=1
j=1
Qwen3-30B-A3B 0.2
0.25
0.3
0.1
0.12
0.15
0 -0.05 -0.1
0 -0.06 -0.12
0 -0.1 -0.2
1
13
24
36
48
1
13
24
36
48
1
13
24
36
48
1
10
18
27
36
1
17
32
48
64
5
8
12
Qwen3-8B 0.25
0.3
0.4
0.12
0.15
0.2
0 -0.06 -0.12
0 -0.1 -0.2
0 -0.1 -0.2
1
10
18
27
36
1
10
18
27
36
OLMo-3-32B 0.15
0.2
0.2
0.075
0.1
0.1
0
0 -0.075 -0.15
0 -0.075 -0.15
-0.075 -0.15
1
17
32
48
64
1
17
32
48
64
OLMoE-1B-7B 0.15
0.2
0.15
0.075
0.1
0.075
0 -0.04 -0.08
0 -0.08
0 -0.04 -0.08
1
5
8
12
16
1
5
8
12
16
1
16
Gemma3-4B 0.05
0.05
0.04
0.025
0.025
0.02
0
0
-0.02
1
9
18
26
Decoder block ℓ
(A) C4
34
-0.02
1
9
18
26
Decoder block ℓ
(B) WikiText-103
34
0 -0.01 -0.02
1
9
18
26
Decoder block ℓ
34
(C) GSM8K text
Figure 10: Cumulative relative squared-error accounting. Rows are models and columns are datasets. The black curve is the exact squared relative hidden error; the colored curves are the cumulative three-term contributions. Error bars show one sample standard deviation across complete input sequences. Position selection is described in App. E.3. 30
Preprint.
Per-block relative squared-error accounting
(R(ℓ))2 − (R(ℓ − 1))2
T(ℓ) inter
T(ℓ) add
T(ℓ) align
Qwen3-30B-A3B 5e-3 3e-3 0
8e-3
0.015
4e-3
7e-3
0
-3e-3
-4e-3
-6e-3
-8e-3
0 -6e-3 -0.012
1
13
24
36
48
1
13
24
36
48
1
13
24
36
48
1
10
18
27
36
1
17
32
48
64
5
8
12
Qwen3-8B 0.012
0.025
0.05
6e-3
0.013
0
0.025
0 -0.01 -0.02
0 -0.013 -0.025
-6e-3 -0.012
1
10
18
27
36
1
10
18
27
36
OLMo-3-32B 6e-3 3e-3 0
8e-3 4e-3 0
-5e-3
-5e-3 -0.01
8e-3 4e-3 0
1
17
32
48
64
-0.01
-5e-3 1
17
32
48
64
-0.01
OLMoE-1B-7B 0.015
0.015
0.02
7e-3
7e-3
0.01
0
0
-7e-3
-7e-3
0 -7e-3 -0.015
-0.015
1
5
8
12
16
-0.015
1
5
8
12
16
1
16
Gemma3-4B 0.015
0.01
6e-3
7e-3
5e-3
3e-3
0 -3e-3
0 -3e-3
1
9
18
26
Decoder block ℓ
(A) C4
34
0 1
9
18
26
Decoder block ℓ
(B) WikiText-103
34
-2e-3
1
9
18
26
Decoder block ℓ
34
(C) GSM8K text
Figure 11: Per-block relative squared-error contribution. Rows are models and columns are datasets. The black curve is the exact blockwise change; the colored curves are the three terms in Thm. 1. Error bars show one sample standard deviation across complete input sequences.
31
Preprint.
E.3
H IDDEN -E RROR A NALYSIS : A BNORMAL T RAJECTORY F ILTERING
Counteraction slows hidden-error growth on average, but it does not keep every quantized next-token trajectory stable. At a small fraction of positions, the hidden error becomes unusually large and the original and W4 hidden-state norms separate sharply across depth. These trajectories can dominate averages of the signed recurrence terms, although the recurrence formula is exact at every position. We therefore separate these cases with a trajectory filter, test its threshold sensitivity, and examine the excluded cases. App. D.3 gives the full selection and aggregation protocol. (ℓ) b (ℓ) L Each decoding position i produces a BF16/W4 hidden-state trajectory {hi , h i }ℓ=0 . We compute the largest relative difference between the BF16 and W4 hidden-state norms along that trajectory: (ℓ)
(ℓ)
b ∥2 − ∥h ∥2 ∥h i i gimax :=
max
(ℓ)
∥hi ∥2
ℓ∈{0,...,L}
.
By default, we retain position i when gimax ≤ 0.5, which retains 99.72% of positions across the evaluated settings. The same positions are used at every layer, so all points on a layerwise curve refer to one fixed population. We form the recurrence terms of Thm. 1 at each retained position and then average over positions within each complete sequence. Error bars in the reported figures show one sample standard deviation across complete input sequences. We repeat the analysis with thresholds 1.0, 0.5, 0.25, and 0.1, and without filtering. Tab. 2 reports E[R(L) ], first averaging retained positions within each sequence and then giving equal weight to each b (L) − h(L) ∥2 /∥h(L) ∥2 is the final relative sequence and model-dataset setting, where R(L) := ∥h hidden error defined in Sec. 3.2. Table 2: Final relative hidden error under different hidden-norm-gap thresholds. We exclude position i with gimax > threshold. Values are macro-averaged over the evaluated model–dataset settings. Both sums in the last column run over decoder blocks 1–L. Norm-gap threshold Unfiltered 1.0 0.5 (default) 0.25 0.1
Positions retained
E[R(L) ]
h P P (ℓ) i (ℓ) E − Tinter / Tadd
100.00% 99.92% 99.72% 98.62% 89.34%
0.148 0.147 0.147 0.144 0.136
30.4% 57.7% 50.4% 50.2% 51.8%
The default threshold is a suitable cutoff: it changes the mean final relative hidden error only from 0.148 to 0.147 while retaining 99.72% of positions. It removes rare scale failures while preserving the main statistics. When all positions are included, the fraction of Tadd canceled by Tinter falls from 50.4% to 30.4%. Counteraction is less effective when the W4 norm trajectory departs sharply from the original trajectory. These rare errors appear to exceed the range that the model can regulate through counteraction. For the LM-head norm statistics used in Sec. 4, the default filter gives !2 b LM ∥2 b LM ∥2 ∥h ∥h E − 1 = 0.0372, E −1 = 0.00320, ∥hLM ∥2 ∥hLM ∥2 averaged equally over C4, WikiText-103, and GSM8K text. Without filtering, the values are 0.0373 and 0.00324, only 0.31% and 1.27% higher. Thus, the LM-head-input norm remains statistically b LM ∥2 /∥hLM ∥2 ≈ 1 used in Sec. 4.2. stable even without the filter, supporting the approximation ∥h Fig. 12 shows the excluded regime directly. For each of Qwen3-32B, Qwen3-8B, and OLMo3-7B, we select the position whose gimax is closest to the median among that model’s excluded positions. This deterministic rule does not inspect the recurrence values or output probabilities. The BF16 and W4 hidden-state norms separate sharply in these examples, and the recurrence terms of Thm. 1 become large enough to dominate an unfiltered mean despite the small number of such positions. 32
Preprint.
Abnormal Case 1 (Qwen3-32B) Hidden-norm ratio across depth 1.5 1.0 0.5
0
10
20
30 40 Block
Context ... how many pens does he get? Answer: He can make 5 pens because 25 / 5 =
1.50 h(i ) 2/ hi( ) 2 1.25 1.00 0.75 0.50 0 5 10
Cumulative recurrence terms
1.0 (j) 0.5 j = 1T 0.0 0.5 1.0 60 0 20 Block 40 BF16 probability (%)
h(i ) 2/ hi( ) 2
50
"5" "4" "1" "<space>" " =" " five"
gimax = 0.896
99.97%
0.013% 3.5e-03% 3.1e-03% 1.1e-04% 2.4e-05%
10 4
10 2
(R ( ))2 (j) Tadd
j=1
(j) Tinter
j=1
(j) Talign
j=1
60 W4 probability (%)
13.07%
100
102
10 4
10 2
0.6
(R ( ))2
T (j)
(j) Tadd
j=1
(j) Tinter
j=1
25
30
(j) Talign
35 0 10 BF16 probability (%)
"=" 2.0e-04% ">>" ">=" 2.5e-06% " ="5.6e-07% "=>" 2.5e-06% 10 4 10 2
Block20
100.00%
j=1
30 W4 probability (%)
h(i ) 2/ hi( ) 2
1.0 0.5
Context
10
15 20 Block
25
2.6e-04% 2.6e-04% 6.5e-05%
100
102
" year"
1.0 T (j) 0.5 j = 1 0.0 0.5 1.0 30 0 10 Block 20 BF16 probability (%) 97.81%
... she paid $756 for the " subscription" other half of the year for " streaming" " season" the streaming service. " month" "<space>"
10 4
10 2
99.99%
9.6e-03%
10 4
10 2
Abnormal Case 3 (OLMo3-7B)
5
102
gimax = 0.628
0.0
... How much does Lloyd make on eggs per week? Answer: His farm produces 252 x 7 = 1764
0
100
0.2
Context
1.5
0.593% 0.476%
Abnormal Case 2 (Qwen3-8B) 0.4 j = 1
15Block20
48.58%
0.091% 0.225%
0.555% 0.211% 0.068% 0.211% 0.060%
100
102
100
102
gimax = 0.869 (R ( ))2 (j) Tadd
j=1
(j) Tinter
j=1
(j) Talign
j=1
30 W4 probability (%) 0.017% 0.016%
10 4
10 2
0.259% 0.194% 0.036%
100
98.87%
102
Figure 12: Rare hidden-state trajectory failures occur when BF16 and W4 hidden norms diverge. All three cases are from GSM8K text. Rows show Qwen3-32B, Qwen3-8B, and OLMo3-7B positions selected by proximity to each model’s median excluded gimax ; the value beside each row title is this selection statistic, not the plotted ordinate. Selection does not use recurrence values or (ℓ) output changes. The left column shows ∥hbi (ℓ) ∥2 /∥hi ∥2 , with the retained range [0.5, 1.5] shaded in gray. The right column shows the cumulative recurrence terms. Both columns use decoder block ℓ on the horizontal axis. The lower strip shows the context and BF16/W4 next-token probabilities.
33
Preprint.
E.4
H IDDEN -E RROR G ROWTH : L ATE -L AYER D IFFERENCES ACROSS DATA
Even with the counteraction mechanism, hidden-error growth can also vary across input domains for pretrained models. Qwen3-32B accumulates relative hidden error faster on GSM8K text than on C4 in the later blocks. Here, we identify the responsible recurrence term and test it by applying the same update rescaling to both streams. The relative additional-error term in Thm. 1 is !2 " # 2 ∥u(ℓ) ∥2 ∥∆u(ℓ) ∥2 (ℓ) (ℓ−1) 2 Tadd = − R . ∥u(ℓ) ∥2 ∥h(ℓ) ∥2 Fig. 13(B) shows that the relative hidden-error trajectories remain close in the earlier blocks but 2 separate sharply later. Over blocks 36–64, the coefficient ∥u(ℓ) ∥2 /∥h(ℓ) ∥2 is larger on GSM8K text (Fig. 13(A)). The block-update norms are comparable across the two datasets, but the hiddenstate norm is smaller on average for GSM8K text. Its updates are therefore stronger relative to (ℓ) the hidden state. When the bracketed term is positive, the larger coefficient increases Tadd and accelerates relative hidden-error growth. Fig. 13(A) and Fig. 13(B) show that the coefficient and hidden error change together but do not isolate the coefficient. We keep the GSM8K inputs and internal block computations fixed, then rescale only the residual updates to match the C4 coefficient at each block. We intervene over blocks 36–64 and denote the modified trajectories by the subscript c: b (ℓ) = h b (ℓ−1) + λ(ℓ) u b (ℓ) h c c c .
(ℓ−1) h(ℓ) + λ(ℓ) u(ℓ) c = hc c ,
2 (ℓ) On the BF16 pass, we choose λ(ℓ) so that the position-averaged value E λ(ℓ) ∥uc ∥2 /∥h(ℓ) on c ∥2 GSM8K dataset exactly matches the averaged value on C4 at the same block, then reuse λ(ℓ) on the W4 pass. Only the residual addition is rescaled; the internal attention and MLP computations remain unchanged. This oracle intervention is used only to test the coefficient, so we do not evaluate output quality. After matching the coefficients, the final E[(R(L) )2 ] decreases from 0.121 to 0.062, corresponding to a 48.8% reduction (Fig. 13(C)). This oracle-style intervention shows that the coefficient contributes substantially to the faster late-layer hidden-error growth on GSM8K text. More generally, data change the relative strength of block updates across depth, producing different hidden-state norm trajectories. The ratio ∥u(ℓ) ∥2 /∥h(ℓ) ∥2 measures the signal generated by block ℓ relative to its hidden state. When the bracketed term is positive, a larger ratio gives newly generated block-update error a larger contribution to the residual stream. Later blocks inherit this larger hidden error, so subsequent propagation starts from a larger value. Thus, data-dependent block-update strength can determine where hidden-error growth accelerates. On GSM8K, the larger ratio after block 36 accounts for much of the rapid growth in Fig. 13(B). [(‖u(ℓ)‖2 /‖h(ℓ)‖2 )2 ]
C4 GSM8K text
0.6
0.200
[(R (ℓ))2 ]
C4 GSM8K text
0.150
[(R (ℓ))2 ]
0.200
GSM8K observed Coefficient-matching intervention
0.150
0.4
0.100
0.2
0.050
0.0 32
0.000
36
40
48
56
Decoder block ℓ (A) Coefficient
‖u(ℓ)‖22 ‖h(ℓ)‖22
in Tadd
64
0.100 0.050 1
16
32 36
48
64
0.000
1
16
32 36
48
Decoder block ℓ
Decoder block ℓ
(B) Relative hidden error
(C) Coefficient matching
64
Figure 13: The update-to-hidden norm ratio explains much of the data-dependent late hiddenerror growth in Qwen3-32B. (A) The coefficient ∥u(ℓ) ∥22 /∥h(ℓ) ∥22 in Tadd is larger on GSM8K text over most late blocks. (B) The relative hidden error then grows faster. (C) Matching the coefficient to the C4 values from block 36 reduces this growth. Error bars show the standard error across 64 sequences at selected blocks.
34
Preprint.
E.5
C OUNTERACTION E MERGENCE DURING P RETRAINING
Sec. 3 identifies counteraction as a major difference between randomly initialized and pretrained models. At random initialization, the interaction between the block-input hidden error ∆h(ℓ−1) and block-update error ∆u(ℓ) is nearly zero. In pretrained Qwen3-32B, it is negative over much of the network and offsets part of the block-update error, slowing hidden-error growth. The OLMo3-7B Stage-1 checkpoints in Fig. 1 show this change along pretraining, from weak interaction at step 0 to the negative interaction observed at the final checkpoint. We use Pythia checkpoints to test whether the same change appears outside the OLMo family. Fig. 14 repeats the checkpoint comparison for Pythia-1.4B and Pythia-2.8B. In both models, the cosine between ∆h(ℓ−1) and ∆u(ℓ) changes from a near-zero value at step 0 to broadly negative values at the final checkpoint. The OLMo3 and Pythia interaction curves show the same change during training: counteraction is weak near initialization and becomes more pronounced at later checkpoints. Pythia-1.4B
Step: 0
[‖Δh(ℓ)‖2/‖h(ℓ)‖2]
final (143k)
0.00 −0.25 0
6
12
18
24
0
Pythia-2.8B
[‖Δh(ℓ)‖2/‖h(ℓ)‖2]
6 Step: 0
12 10k
18 50k
24
final (143k)
[cos(Δh(ℓ − 1), Δu(ℓ))]
0.00
0.2 0.0
50k
[cos(Δh(ℓ − 1), Δu(ℓ))]
0.2 0.0
10k
−0.25 0
8
16
24
32
0
8
16
24
32
Decoder block ℓ (A) Relative hidden-error magnitude
(B) Input-update error cosine
Figure 14: Counteraction strengthens during Pythia pretraining. Rows show four checkpoints of Pythia-1.4B and Pythia-2.8B under the same NVFP4 weight conversion. Columns report relative hidden error and the cosine between the block-input hidden error and block-update error. Error bars show standard deviation.
35
Preprint.
E.6
C OUNTERACTION S OURCE : B LOCK -I NPUT-E RROR R ESPONSE
Sec. 5.2 attributes counteraction mainly to the block’s response to the hidden error at its input. Here, we provide the detailed evidence in two steps. We first show that counteraction is specific to the error dynamics rather than a general relation between a hidden state and its residual update. We then decompose the block-update error to determine whether the negative interaction comes from the direct weight perturbation or from the block’s response to the inherited hidden error. Counteraction is specific to the error dynamics. Counteraction refers to the tendency of the block-update error ∆u(ℓ) to point against the hidden error ∆h(ℓ−1) already present at the block input (Sec. 3.2). A related phenomenon has been observed for native residual contributions: later Transformer layers can partially oppose earlier contributions to correct each other (Patrawala et al., 2025). We therefore ask whether the negative interaction between quantization errors simply reflects a general tendency of a block update to oppose its input hidden state. Fig. 15 compares these relations. The native hidden-state/update cosines cos ∠(h(ℓ−1) , u(ℓ) ) and b (ℓ−1) , u b (ℓ) ) are positive in 65.6% and 64.1% of the blocks, respectively. By contrast, cos ∠(h cos ∠(∆h(ℓ−1) , ∆u(ℓ) ) is negative in 82.5% of the blocks. Thus, the strong negative relation is not a general property of the residual update; it appears much more consistently between the inherited input hidden error and the new block-update error.
cos (h(
0.5
[cos ( , )]
1), u( ))
cos (h( 1), u( )) cos ( h( 1), u( )) cos ( h( 1), u( )) BF16/W4 hidden-state/update cosines (cos > 0 for most blocks)
0.0 0.5
1
16
Block-input/block-update error cosine (cos < 0 for most blocks) 32
48
Decoder block
64
Figure 15: Counteraction is specific to the error dynamics in Qwen3-32B. BF16 and W4 hiddenstate/update cosines are compared with the cosine between the block-input and block-update errors on C4 inputs. Error bars show one sample standard deviation. Decomposing the block-update error. We next ask which part of ∆u(ℓ) produces this negative interaction. At block ℓ, the original and quantized updates are b (ℓ−1) ). b (ℓ) (h b (ℓ) = F u(ℓ) = F(ℓ) (h(ℓ−1) ), u We separate these two effects exactly: b (ℓ−1) ) − F b (ℓ) (h b (ℓ) (h(ℓ−1) ) . b (ℓ) (h(ℓ−1) ) − F(ℓ) (h(ℓ−1) ) + F ∆u(ℓ) = F {z } {z } | | (ℓ) ∆uweight (direct weight effect)
(22)
(ℓ) ∆uresponse (response to block-input error)
The first component changes the block weights while keeping its input fixed at h(ℓ−1) . The second b (ℓ−1) . Because T (ℓ) = keeps the quantized block fixed and changes only its input from h(ℓ−1) to h inter 2⟨∆h(ℓ−1) ,∆u(ℓ) ⟩ is linear in ∆u(ℓ) , this decomposition also gives ∥h(ℓ) ∥22 (ℓ)
(ℓ)
(ℓ)
(ℓ)
Tinter = Tinter,weight + Tinter,response =
2⟨∆h(ℓ−1) , ∆uweight ⟩ ∥h(ℓ) ∥22
(ℓ)
+
2⟨∆h(ℓ−1) , ∆uresponse ⟩ ∥h(ℓ) ∥22
.
Thus, the source of counteraction can be identified by asking which component contributes the negative interaction with ∆h(ℓ−1) . We measure both components by replaying every decoder block of Qwen3-32B on the C4 inputs from Sec. 3.2. 36
Preprint.
cos ( h( cos ( h( cos ( h( [cos ( h( 1), )]
1), u( )) 1), F( ) (h( 1), F( ) (h(
1)) 1))
) = 2 h( 1), u( ) / h( ) 2 T(inter 2 ) () ( ) ( 1)) / h( ) 2 T(inter, response = 2 h( 1), F (h( 1)) F (h 2 ) ( ) (h( 1)) F( ) (h( 1)) / h( ) 2 ( 1) T(inter, F = 2 h , weight 2
F( ) (h( 1))) F( ) (h( 1))) 0
0.0 0.2
×10 3
5
0.4 1
16
32
48
Decoder block (A) Cosines with block-input hidden error
64
1
16
32
48
Decoder block (B) Contributions to Tinter
64
Figure 16: The response to block-input error carries counteraction in Qwen3-32B. C4 inputs. (A) Cosines between the block-input error and the block-update error ∆u(ℓ) , split into the direct (ℓ) weight effect and input-error response from Eq. (22). (B) Their contributions Tinter,weight and (ℓ)
(ℓ)
Tinter,response to Tinter . Error bars show one sample standard deviation. The input-error response carries the negative direction. Fig. 16(A) compares the directions of (ℓ) the two components. The full block-update error ∆u(ℓ) closely follows ∆uresponse , whose cosine with the block-input hidden error is negative in 83% of the blocks. By contrast, the cosine curve (ℓ) for ∆uweight stays within 0.02 of zero at every block. The direct weight-quantization perturbation therefore produces a nonzero update error, but it has no consistent direction relative to the hidden error already present at the block input. Fig. 16(B) measures the quantity directly relevant to counteraction: the signed contribution to Tinter . PL PL (ℓ) (ℓ) After summing over blocks, ℓ=1 Tinter,response accounts for 99.9% of the negative ℓ=1 Tinter , PL (ℓ) whereas ℓ=1 Tinter,weight accounts for less than 0.1%. This attribution is directional rather than ∥∆u(ℓ) weight ∥2 is a consequence of the response component being much larger in norm. In fact, E ∥∆u (ℓ) ∥ 2 ∥∆u(ℓ) ∥ 2 response 84% of E ∥∆u . The two components therefore have comparable magnitudes, but only the (ℓ) ∥ 2 input-error response points consistently against the inherited hidden error. Robustness to the decomposition path. The exact decomposition of ∆u(ℓ) is not unique because one may choose either block map for the intermediate evaluation. To check that the attribution above b (ℓ−1) ), giving does not depend on this choice, we instead add and subtract F(ℓ) (h b (ℓ−1) ) − F(ℓ) (h b (ℓ−1) ) + F(ℓ) (h b (ℓ−1) ) − F(ℓ) (h(ℓ−1) ) . b (ℓ) (h ∆u(ℓ) = F Here the first term measures the direct weight effect at the W4 input, whereas the second measures the response of the original block to the same input error. This alternative decomposition gives the same qualitative attribution: across blocks, the first term above contributes +0.0013 to the cumulaPL (ℓ) tive Tinter , whereas the second contributes −0.1118 to the ℓ=1 Tinter = −0.1105. Therefore, under either exact decomposition, counteraction is carried mainly by how the block responds to the hidden error at its input. The resulting change in the block update tends to oppose the accumulated hidden error and slow its propagation.
37
Preprint.
E.7
C OUNTERACTION I NTERVENTION : R EMOVAL AND R EVERSAL
Sec. 3.3 shows that removing or reversing counteraction increases hidden error and output change in Qwen3-32B. Here, we give the online recursions and term-wise accounting, then repeat the intervention in Qwen3-8B to test whether the effect transfers across model scales. Removal and reversal follow separate intervened trajectories initialized from h(0) . In each reb (ℓ−1) is the current modified block input. We recompute the block-input hidcursion below, h b (ℓ−1) − h(ℓ−1) , ∆u(ℓ) := den error and W4 block-update error at every block: ∆h(ℓ−1) := h b (ℓ−1) ) − F(ℓ) (h(ℓ−1) ). Here, ∆u(ℓ) follows the definition in Eq. (1), but is computed on b (ℓ) (h F the modified trajectory. For removal, the equal-norm error orthogonal to the block-input hidden error is ⟨∆h(ℓ−1) , ∆u(ℓ) ⟩
∆u(ℓ) −
∥∆h(ℓ−1) ∥22 (ℓ−1) (ℓ)
(ℓ)
∆u⊥ := ∥∆u(ℓ) ∥2 ∆u
(ℓ)
−
⟨∆h
, ∆u
∥∆h(ℓ−1) ∥22
∆h(ℓ−1) .
⟩
∆h
(ℓ−1) 2
The removal and reversal trajectories are then propagated directly as ( (ℓ) (ℓ−1) (ℓ) ∆u⊥ , 17 ≤ ℓ ≤ 48, (ℓ) b b + u + = h h rem rem ∆u(ℓ) , otherwise, ( (ℓ−1) (ℓ) , ∆u(ℓ) ⟩ < 0, b (ℓ) = h b (ℓ−1) + u(ℓ) + −∆u , 17 ≤ ℓ ≤ 48 and ⟨∆h h rev rev ∆u(ℓ) , otherwise.
(23)
(24)
Removal makes the applied block-update error orthogonal to the block-input hidden error at every position in blocks 17 ≤ ℓ ≤ 48; reversal flips only negative interactions in this interval. Both preserve ∥∆u(ℓ) ∥2 at the intervened block, although later block-update errors may change along the modified trajectory. A deterministic orthogonal fallback handles degenerate projections but was never triggered. Term-wise analysis follows the abnormal case exclusion selection in App. D.3. The same filtering rule is applied to each modified trajectory. Fig. 17 provides the term-wise recurrence accounting for Qwen3-32B. Across the intervened blocks, removal nearly eliminates cumulative Tinter while increasing cumulative Tadd by 4.3×. Together, these changes increase the final E[(R(L) )2 ] by 8.4×. Although the recurrence separates Tadd and Tinter algebraically, changing the trajectory also changes the subsequent values of Tadd . Thus, counteraction limits error growth both through direct negative interaction and by preventing larger values of Tadd later in the forward pass. ℓ
Squared error: (R(ℓ))2
ℓ
(j)
∑ Tadd
ℓ
(j)
∑ Tinter
j=1
j=1
Qwen3-32B
(j)
∑ Talign
j=1
102
102
0
10
0
100
0
0
0
−100
−100
−100
10
−102
1
17
32
48
Decoder block ℓ
(A) No intervention
64
−102
1
17
32
102
48
Decoder block ℓ
64
(B) Counteraction removed
−102
1
17
32
48
Decoder block ℓ
64
(C) Counteraction reversed
Figure 17: Relative squared-error accounting after counteraction intervention in Qwen3-32B. Columns compare unmodified W4 with counteraction removal and reversal over blocks 17–48, where counteraction is strong. The corresponding hidden-error and output changes are reported in Sec. 3.3. In contrast to the other figures, a symmetric logarithmic scale is used here to compare trajectories that span multiple orders of magnitude. 38
Preprint.
We repeat the intervention in Qwen3-8B over its strongest-counteraction interval, blocks 6–23. Fig. 18 combines the hidden-error and output changes with the same term-wise accounting. Removal again nearly eliminates cumulative Tinter , while cumulative Tadd increases by 2.7× and the final E[(R(L) )2 ] increases by 4.8×. The Qwen3-8B result therefore reproduces the same coupling between the negative interaction and the subsequent hidden-state trajectory. Observed W4 Counteraction removed
Counteraction intervention in Qwen3-8B
1250 [‖Δh(ℓ)‖2]
Reversed: 1213.2 [‖Δh
(ℓ)
Counteraction reversed Intervened blocks 6–23
‖2/‖h(ℓ)‖2]
W4
3
1000 750
2
500
Removed: 456.8
250
W4: 197.3
1
0
Reversed: 0.728 Removed: 0.274 W4: 0.121
0 0
9
18
27
36
0
9
Decoder block ℓ (A) Absolute error norm
18
27
+0.04
+0.26
KL ↓
0.053
0.390
4.316
Flip@1 ↓
10.7%
27.5%
79.1%
Ret@20 ↑
89.2%
73.6%
27.4%
ℓ
(C) Output metrics
ℓ
(j)
∑ Tadd
ℓ
(j)
∑ Tinter
j=1
j=1
(j)
∑ Talign
j=1
101
101
101
10−1
10−1
10−1
0
0
0
−10−1
−10−1
−10−1
−101
−101 1
10
18
27
Decoder block ℓ
(D) No intervention
36
+3.93
36
Decoder block ℓ (B) Relative error norm
Squared error: (R(ℓ))2
Removed Rev.
ΔCE ↓
−101 1
10
18
27
Decoder block ℓ
36
(E) Counteraction removed
1
10
18
27
Decoder block ℓ
36
(F) Counteraction reversed
Figure 18: Counteraction intervention in Qwen3-8B. (A,B) Hidden-error trajectories and (C) output changes after removing or reversing counteraction over blocks 6–23. (D)–(F) Relative squarederror accounting along the same three trajectories. Output metrics are defined in Sec. 4.1; error bars show one sample standard deviation across inputs. The lower panels use a symmetric logarithmic scale.
39
Preprint.
E.8
C OUNTERACTION ACROSS W EIGHT AND ACTIVATION Q UANTIZATION
Sec. 5.1 asks whether the negative interaction between block-input hidden error and block-update error is specific to weight quantization. Fig. 6(A) compares its blockwise cosine under W4 (weightonly quantization), A4 (activation-only quantization), and W4A4 (joint weight and activation quantization), and finds a similar depth-wise pattern in all three settings. The cosine shows the direction of the interaction, but not how much it contributes to the relative hidden-error trajectory or how the remaining error changes the next-token distribution. Fig. 19 applies Thm. 1 separately to each setting. The relative-error trajectories and the amount of newly added error differ, but the cumulative Tinter contribution is negative in all three cases. Thus, the block-update error counteracts the block-input hidden error whether the perturbation enters through weights, activations, or both. Fig. 20 compares the corresponding CE, KL, top-token retention, and rank-wise log-probability changes. The three settings do not produce the same output error, even though their blockwise counteraction curves are similar. Counteraction describes how hidden error is limited through depth; output robustness also depends on how much error each block introduces and how the remaining error affects the LM head. ℓ
Cumulative: (R(ℓ))2
ℓ
(j)
∑ Tadd
ℓ
(j)
∑ Tinter
j=1
j=1
0.4
0.4
0.4
0.2
0.2
0.2
0.0
0.0
0.0
−0.2
−0.2
−0.2
0
16
32
48
64
0
16
32
48
(j)
∑ Talign
j=1
64
0
16
32
48
Decoder block ℓ
Decoder block ℓ
Decoder block ℓ
(A) W4
(B) A4
(C) W4A4
64
Figure 19: Relative-error recurrence under weight and activation quantization. Qwen3-32B on C4 under W4, A4, and W4A4. The black curve is the cumulative squared relative error; the colored curves are the cumulative contributions of Tadd , Tinter , and Talign . Error bars show one sample standard deviation across complete inputs.
Setting
CE diff. ↓
Rel. CE (%) ↓
KL ↓
Setting
Flip@1 (%) ↓
W4
0.019
0.78
0.044
W4
10.6
88.8
88.9
A4
0.019
0.76
0.039
A4
10.0
89.8
89.9
W4A4
0.036
1.51
0.080
W4A4
14.4
85.6
85.6
(A) Next-token distribution changes
Ret@10 (%) ↑ Ret@20 (%) ↑
(B) Top-ranked token preservation
[|Δlog pπr |]
100
W4
10−1 1
10
100
1k
A4 10k
W4A4 ||
Original-model vocabulary rank r
(C) Log-probability change by rank
Figure 20: Output changes under weight and activation quantization. Qwen3-32B on the same C4 inputs and retained next-token positions used in Fig. 19. Metrics follow Sec. 4.1. Error bars show one sample standard deviation across complete inputs.
40
Preprint.
E.9
C OUNTERACTION ACROSS PTQ A LGORITHMS AND W EIGHT F ORMATS
Sec. 5.1 also asks whether counteraction depends on how the 4-bit weights are produced. Fig. 6(B) shows that the blockwise counteraction cosine changes little across calibration-free RTN and the calibration-based GPTQ and AWQ algorithms, all using NVFP4 weights. This comparison does not show whether the same conclusion holds for another weight format, or whether similar counteraction curves imply similar relative-error and output changes. We compare RTN, GPTQ, and AWQ on Qwen3-4B under both NVFP4 and asymmetric INT4. All conditions quantize the same Attention and MLP weights, and we report the exact relative-error recurrence together with the resulting next-token output changes. Evaluation follows App. D.3 and PL (ℓ) retains at least 99.93% of positions. Every condition has ℓ=1 Tinter < 0. For NVFP4, the blockwise interaction-cosine curves under GPTQ and AWQ have correlations above 0.998 with the RTN curve. Asymmetric INT4 quantization details. The INT4 experiments use the W4A16 ASYM scheme from llmcompressor. All Attention and MLP linear weights are quantized in groups of 128 consecutive values along the input dimension, while activations remain in BF16 and the LM head remains unquantized. Each group includes zero and uses signed 4-bit integer values qmin = −8 and qmax = 7. Given group wmin and wmax , the static affine parameters are s=
wmax − wmin , 15
wmin zzp = clip round qmin − , qmin , qmax , s
where zzp is stored as INT8. Quantization and reconstruction use w + zzp , qmin , qmax , q = clip round s
w b = s(q − zzp ).
Calibrated PTQ algorithm details. GPTQ and AWQ use 64 disjoint 512-token C4 calibration sequences; RTN uses no calibration data. For NVFP4, NVIDIA Model Optimizer applies GPTQ with block size 128 and Hessian dampening 0.01, or AWQ-lite with alpha step 0.1, directly to the E2M1 weights and E4M3 block scales defined above. For asymmetric INT4, GPTQ uses the same block size and dampening with static activation ordering, while AWQ uses activation-and-weight duo scaling with 20 grid-search points. Figs. 21, 22, and 23 give the full results. Across the tested algorithms and formats, Tinter remains negative and its blockwise geometry changes little, but Tadd , final relative error, and output metrics still vary. A stronger cumulative Tinter does not imply better PTQ quality. Output changes also depend on newly introduced error and on how the remaining hidden error affects the output layer. RTN NVFP4 RTN INT4
GPTQ NVFP4 GPTQ INT4
AWQ NVFP4 AWQ INT4
R (ℓ) = ‖Δh(ℓ)‖2/‖h(ℓ)‖2
0.20
cos ∠(Δh(ℓ − 1), Δu(ℓ))
0.0 0.15 0.10
−0.2
0.05 −0.4
0.00 0
9
18
27
36
1
9
18
27
Decoder block ℓ
Decoder block ℓ
(A) Relative hidden-error magnitude
(B) Blockwise counteraction cosine
36
Figure 21: Counteraction is consistent across PTQ algorithms and weight formats. Qwen3-4B results. Solid and dashed curves denote NVFP4 and asymmetric INT4. Error bars show one sample standard deviation across inputs. 41
Preprint.
ℓ
Cumulative: (R(ℓ))2
ℓ
(j)
∑ Tadd
ℓ
(j)
∑ Tinter
j=1
j=1
RTN
NVFP4
(j)
∑ Talign
j=1
Asymmetric INT4
0.4 0.2 0.0 −0.2 0
9
18
27
36
0
9
GPTQ
NVFP4
18
27
36
27
36
27
36
Asymmetric INT4
0.1 0.0 −0.1 0
9
18
27
36
0
9
AWQ
NVFP4
18
Asymmetric INT4
0.2 0.0 −0.2
0
9
18
27
36
0
9
Decoder block ℓ
18
Decoder block ℓ
Figure 22: Relative-error recurrence across PTQ algorithms and weight formats. Qwen3-4B under RTN, GPTQ, and AWQ with NVFP4 and asymmetric INT4. Each row shows one PTQ algorithm, and the two columns share a vertical scale within that row. The black curve is the cumulative squared relative error; the colored curves are the cumulative contributions of Tadd , Tinter , and Talign . Error bars show one sample standard deviation across inputs.
Recipe
CE diff. ↓
Rel. CE (%) ↓
KL ↓
Recipe
Flip@1 (%) ↓
RTN NVFP4
0.022
0.63
0.061
RTN NVFP4
11.8
Ret@10 (%) ↑ Ret@20 (%) ↑
88.4
88.5
GPTQ NVFP4
-0.006
-0.17
0.036
GPTQ NVFP4
8.6
91.4
91.5
AWQ NVFP4
0.027
0.78
0.047
AWQ NVFP4
10.2
90.0
89.9
RTN INT4
0.122
3.54
0.103
RTN INT4
14.6
85.6
85.4
GPTQ INT4
0.015
0.44
0.042
GPTQ INT4
8.9
90.6
90.6
AWQ INT4
0.008
0.23
0.062
AWQ INT4
11.8
88.3
88.3
(A) Next-token distribution changes
(B) Top-ranked token preservation
[|Δlog pπr |]
100
RTN NVFP4 RTN INT4
10−1 1
10
100
1k
GPTQ NVFP4 GPTQ INT4
10k
AWQ NVFP4 AWQ INT4
||
Original-model vocabulary rank r
(C) Log-probability change by rank
Figure 23: Output changes across PTQ algorithms and weight formats. Qwen3-4B results; metrics follow Sec. 4.1. (C) compares the rank-wise absolute log-probability change under the same inputs and position filter. Solid and dashed curves denote NVFP4 and asymmetric INT4. Error bars show one sample standard deviation across inputs.
42
Preprint.
E.10
O UTPUT ROBUSTNESS TO Q UANTIZATION ACROSS M ODELS
Sec. 4 reports two connected observations for Qwen3-32B. Weight quantization causes only modest changes in CE and KL, and the scores and probabilities of top-ranked tokens change less than those of lower-ranked tokens. Sec. 4.2 explains the rank dependence through the geometry of the shared LM head. Higher-ranked token rows form smaller angles with the LM-head input, and the rotation of that input produces much smaller changes in its angle to each fixed row. Here we test whether the output stability and rank-dependent LM-head geometry extend beyond the main model. We evaluate pretrained models from the Qwen, OLMo, and Gemma families, including dense and mixture-of-experts architectures, on C4, WikiText-103, and GSM8K text. Each model is evaluated on 64 complete inputs from each data view, with the original and quantized models compared at the same next-token positions. The evaluation and aggregation follow App. D.3. We first make the rank-dependent geometry explicit. At each position, let πr be the token with BF16 score rank r, and define θπr := ∠(wπr , hLM ). For Qwen3-32B, Fig. 24 shows that E[cos θπr ] decreases from the top of the vocabulary toward lower ranks after averaging over all three datasets. Thus, the rows of top-ranked tokens have smaller projection angles to the LM-head input. Under the directional model in Thm. 2, this is the geometry that makes their relative scores less sensitive to the same LM-head input rotation. [cos θπr ] C4
GSM8K
WikiText-103
0.10 0.08 0.06 0.04 0.02 0.00 −0.02 0 10
101
102
103
104
||
BF16 vocabulary rank r
Figure 24: Higher-ranked tokens have larger LM-head cosine values. Error bars show one sample standard deviation across complete inputs. The grouped means decrease toward lower ranks on all three datasets, describing the overall rank trend rather than pointwise monotonicity. Table 3: Output changes across pretrained models. Each entry is the macro-mean over C4, WikiText-103, and GSM8K text. ∆CE and forward KL compare W4 with BF16 on the same next-token positions. Model
∆CE ↓
Rel. ∆CE (%)↓
b) ↓ KL(p ∥ p
Qwen3-30B-A3B Qwen3-8B OLMo3-32B
0.015 0.017 0.006
0.57 0.54 0.24
0.048 0.060 0.019
Tab. 3 checks output stability on three representative models. Their macro-mean ∆CE ranges from 0.006 to 0.017, relative ∆CE from 0.24% to 0.57%, and forward KL from 0.019 to 0.060. The modest changes measured for Qwen3-32B in Sec. 4.1 therefore also occur at other model scales and in another model family. Fig. 25 then repeats the geometric and rank-wise analysis for six dense and mixture-of-experts modb LM ) with the vocabulary-mean projection-angle els. Fig. 25(A) compares the rotation ∠(hLM , h change. Across the three data views, the LM-head input rotates from 5.80◦ to 21.71◦ , while the vocabulary-mean projection-angle change ranges from only 0.093◦ to 0.391◦ . The attenuation observed for Qwen3-32B is therefore present in every tested model. Fig. 25(B) measures the relative score change of the BF16 rank-r token, E[|b zπr − zπr |/|zπr |], over positions with zπr > 0. The observed W4 curves generally rise toward lower ranks. The first- and b LM ∥2 /∥hLM ∥2 = 1, so they isolate the effect of the measured second-order references set ρ = ∥h input rotation rather than its norm change. They reproduce the overall rank trend, although the rela43
Preprint.
tive score changes and local fluctuations vary across models. Fig. 25(C) carries the same comparison into probability space. The top-ranked tokens have the smallest absolute log-probability changes, and the approximations from Thm. 3 follow the observed transition from the top ranks to the rest of the vocabulary. C4
WikiText-103
GSM8K
Observed W4
First-order approximation
Second-order approximation
Qwen3-14B ∘
102 [|zπ̂ r − zπr |/zπr ], zπr > 0
13.35°
101 9.27°
100
101
6.91°
5⋅10−1
100
100
0.171° 0.120° 0.093°
10−1
[|Δlog pπr |]
10−1
2⋅10−1
10−2
10−1
Qwen3-30B-A3B ∘
[|zπ̂ r − zπr |/zπr ], zπr > 0
9.94°
10
1
7.84°
100
0.224° 0.166° 0.119°
10−1
100
100
5.80°
10−1 10
[|Δlog pπr |]
5⋅10−1 2⋅10−1 10−1
−2
OLMo3-7B ∘
8.39°
10
101 9.53°
[|zπ̂ r − zπr |/zπr ], zπr > 0
[|Δlog pπr |] 2⋅10−1
101
8.51° 100
10
2
100 0.226° 0.211° 0.189°
−1
10−1
10−1
OLMo-3-32B ∘
11.09°
10
101 9.55°
6.14°
[|zπ̂ r − zπr |/zπr ], zπr > 0
[|Δlog pπr |] 2⋅10−1
101
100
0.234° 0.216° 0.151°
10−1
2
100 10
10−1
−1
OLMoE-1B-7B ∘
102 [|zπ̂ r − zπr |/zπr ], zπr > 0
10.58°
10
1
11.38°
5⋅10−1
10.13°
100
100
0.204° 0.192° 0.181°
10−1
[|Δlog pπr |]
101 2⋅10−1
10−1 10−1
Gemma3-4B (FP32 base) ∘
102 [|zπ̂ r − zπr |/zπr ], zπr > 0
19.52°
101 21.71° 100
10−1
100
101
20.25°
5⋅10
[|Δlog pπr |]
−1
100 ̂ ∠(hLM, hLM )
0.391° 0.364° 0.358°
10
k ∈ |Δ∠k|
(A) Hidden rotation vs. LM-head angle shift
−1
2⋅10−1
1
10
100
1k
10k
||
BF16-model rank r
(B) Relative score change by rank
1
10
100
1k
10k
||
BF16-model rank r
(C) Log-probability change by rank
Figure 25: Top-ranked token scores and probabilities are more stable after quantization across models. Six models are evaluated on C4, WikiText-103, and GSM8K text. Projection-angle changes remain below one degree despite larger LM-head input rotations, and relative score and log-probability errors increase toward lower ranks. Score and log-probability results are averaged over the three datasets. Gemma3 uses FP32 for non-quantized computation. (B) uses the ρ = 1 approximation from Thm. 2; (C) uses the approximations from Thm. 3.
44
Preprint.
E.11
O UTPUT P ROBABILITY S ENSITIVITY TO Q UANTIZATION ACROSS T EMPERATURES
Sec. 5.1 reports that preferential preservation of top-ranked tokens persists when the softmax temperature changes. To separate score geometry from the probability mapping, we keep each paired BF16/W4 score vector fixed and vary only the temperature. For fixed original and quantized score vectors, define pT = softmax(z/T ),
T ∈ {0.5, 0.75, 1, 1.5, 2}.
b T = softmax(b p z/T ),
No hidden state or LM-head projection is recomputed. Every T > 0 preserves the BF16 and W4 score rankings. Their top-K token sets and Flip@1 therefore remain unchanged, as do the LM-head angles. Probability metrics do change. Fig. 26 shows the mean forward KL. Across the tested range, it spans 0.0247–0.0661 on C4 and 0.1514–0.3253 on GSM8K text. The same figure also reports the actual absolute log-probability change at each rank and temperature. Positive temperature rescaling preserves both score rankings, so it does not change the score-side rank protection. It changes the magnitude of probability errors, while the top-to-tail trend remains inherited from the score changes. Thus, temperature affects probability sensitivity without changing the hidden-state propagation or LM-head projections studied in Secs. 3 and 4. Qwen3-32B
0.08
[DKL(pT‖pT̂ )]
0.5
[DKL(pT‖pT̂ )]
[|Δlog pπr, T|]
T = 0.5 T = 0.75 T=1
0.4
0.06
100
T = 1.5
0.3 0.04
T=2
0.2
0.02
0.1 0.5 0.75
1
1.5
Temperature T (A) C4
2
0.5 0.75
1
1.5
2
Temperature T (B) GSM8K text
100
101
102
103
104
105
BF16-model rank r (C) Log-probability change by rank
Figure 26: Next-token probability changes under temperature rescaling. (A) and (B) show the mean forward KL on C4 and GSM8K text. Error bars show standard deviation across inputs. (C) shows the actual mean absolute log-probability change over BF16-model rank for five temperatures averaged over both datasets.
45