ConceptioArchivearXiv CS
arXiv CSopen access

Robust Ultra Low-Bit Post-Training Quantization via Stable Diagonal Curvature Estimate

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

Robust Ultra Low-Bit Post-Training Quantization via Stable Diagonal Curvature Estimate Jaemin Kim

Sungkyun Kim

Junyeol Lee

Jiwon Seo∗

Seoul National University [email protected]

Hanyang University [email protected]

Hanyang University [email protected]

Seoul National University [email protected]

arXiv:2604.13806v1 [cs.LG] 15 Apr 2026

Abstract Large Language Models (LLMs) are widely used across many domains, but their scale makes deployment challenging. PostTraining Quantization (PTQ) reduces memory footprint without retraining by leveraging a small calibration set. Recent Hessian-based PTQ methods compensate quantization error via cross-channel dependencies, but such approaches degrade at low bit-widths due to noisy curvature estimates from limited calibration data. We propose DASH-Q, a robust PTQ framework using diagonal Hessian approximation and iterative weighted least squares. By discarding noise-prone dependencies, DASH-Q filters sampling noise while prioritizing the preservation of salient feature power. We outperform other PTQ baselines in ultra low-bit regime, improving zero-shot accuracy by 7.01% on average and up to 14.01% over the strongest baselines across five baseline LLM models, while showing robust and stable performance with very small calibration data. CCS Concepts: • Computing methodologies → Machine learning.

low-bit quantization is often very expensive due to the finetuning overhead to recover accuracy. Recently, GPTQ [18], a post-training quantization (PTQ) method without fine-tuning, proposes using an approximate Hessian from a calibration set to enable low-bit quantization. It mitigates quantization error by propagating (i.e., compensating) the error across feature channels. However, prior work has observed that this strategy can degrade generation quality, particularly at low bit-widths [28, 29]. The reason for this problem, in our analysis, is that off-diagonal Hessian entries are highly susceptible to sampling noise (batch-to-batch variance), making the resulting cross-channel compensation prone to overfitting. Motivated by this, we propose DASH-Q, a PTQ framework that discards noisy feature correlations and retains stable feature importance. Using a diagonal Hessian yields a reliable weighting and decouples quantization into independent weighted least square problems, each with a closed-form solution for the quantization parameters. As a result, DASH-Q enables robust ultra low-bit quantization with strong accuracy and marginal quantization overhead.

Keywords: Deep learning systems, Quantization ACM Reference Format: Jaemin Kim, Sungkyun Kim, Junyeol Lee, and Jiwon Seo. 2026. Robust Ultra Low-Bit Post-Training Quantization via Stable Diagonal Curvature Estimate. In The 6th Workshop on Machine Learning and Systems (EuroMLSys ’26), April 27–30, 2026, Edinburgh, Scotland Uk. ACM, New York, NY, USA, 12 pages. https://doi.org/10.1145/ 3805621.3807619

1

Introduction

Large Language Models (LLMs) are proven to be useful across many application domains, but their scale makes it challenging to deploy them, especially in resource-limited environments. Quantization is a standard approach for reducing the memory footprint of neural networks; however, for LLMs, ∗ Corresponding author

This work is licensed under a Creative Commons Attribution 4.0 International License. EuroMLSys ’26, Edinburgh, Scotland Uk © 2026 Copyright held by the owner/author(s). ACM ISBN 979-8-4007-2605-7/26/04 https://doi.org/10.1145/3805621.3807619

2

Related Works

Given the high computational cost of modern neural networks, prior work has explored a broad range of optimizations spanning training, inference, and efficient deployment [25, 34–36, 39]. Among these directions, quantization have emerged as practical tools for improving efficiency [12, 21, 24, 30]. Within LLM PTQ, early methods focus on activation outliers. SmoothQuant [44] reduces activation quantization difficulty by migrating activation variance into weights. AWQ [29] scales weights using activation statistics, motivated by the observation that values near range-boundaries tend to incur smaller errors. However, pushing outlier magnitudes into weights often makes ultra low-bit quantization complicated. Other approaches isolate outliers or reorganize channels to mitigate precision loss: LLM.int8() [11] and OWQ [27] keep a small set of sensitive channels in higher precision, while other methods such as [23, 46, 49] use channel permutation and grouping to better fit low-precision constraints. Another line of research formulates PTQ as reconstructionerror minimization under local curvature. GPTQ [18] performs layer-wise quantization using a second-order approximation and iteratively compensates the error introduced by previously quantized coordinates. Complementary to these

EuroMLSys ’26, April 27–30, 2026, Edinburgh, Scotland Uk

Jaemin Kim, Sungkyun Kim, Junyeol Lee, and Jiwon Seo

solvers, several methods [3, 6, 15, 48] apply orthogonal transformations to redistribute outlier energy across channels to obtain representations that are easier to quantize at low bitwidths. However, they typically still rely on Hessian-based error compensation after the transformation. Despite their effectiveness, these Hessian-based compensation schemes face limitations in ultra low-bit regimes. Recent studies [28, 29] report that full-Hessian optimization with sparse calibration data can overfit, resulting in a gap between perplexity and downstream task accuracy. Arai et al. [2] further note that such local approximations can amplify error accumulation across layers.

3

4

Backgrounds

Weight-only quantization maps high-precision weights 𝑊 ∈ Rdout ×din to a low-bit representation 𝑄 such that the reconstructed weights 𝑊ˆ incur minimal error. A widely used approach is uniform affine quantization, defined as:  𝑄 = clip 𝑠=

  𝑊 + 𝑧 , 0, 2𝑏 − 1 , 𝑠

max(𝑊 ) − min(𝑊 ) , 2𝑏 − 1

𝑊ˆ = 𝑠 · (𝑄 − 𝑧)

(1)

min(𝑊 ) 𝑠

(2)

𝑧 =−

where 𝑠 and 𝑧 are the scale and zero-point for 𝑏-bit quantization. In LLMs, quantization is typically applied at a per-group granularity to reduce quantization error. To reduce accuracy loss, PTQ often minimizes a layer-level reconstruction error. Given calibration inputs 𝑋 , a standard objective is min ∥𝑊 𝑋 − 𝑊ˆ 𝑋 ∥ 22 = min ∥𝑊𝑖,:𝑋 − 𝑊ˆ 𝑖,:𝑋 ∥ 22 . (3) 𝑊ˆ

𝑊ˆ 𝑖,:

Since the objective is separable across output channels, the minimization decomposes into independent row-wise subproblems, one for each row of 𝑊 . Recently GPTQ and its variants [17, 18] derive a solution for Eq. (3) under a second-order approximation. For a single row 𝑤 = 𝑊𝑖,: , the reconstruction loss can be written as: ˆ ≈ (𝑤 − 𝑤) ˆ H (𝑤 − 𝑤) ˆ ⊤, L (𝑤)

H ≈ Ĥ = 𝑋𝑋 ⊤,

(4)

where H is the Hessian matrix capturing second-order information, and Ĥ denotes its estimate computed from the calibration data. Based on Eq. (4), they incrementally quantize the weight parameters by iteratively selecting the coordinate expected to incur the smallest increase in loss. The most widely used selection method (and the corresponding error compensation for the remaining columns) relies on the inverse Hessian as follows, where 𝑤 𝑗 denotes the 𝑗’th column of 𝑤: 𝑤 𝑗 = arg min 𝑤𝑗

(𝑤 𝑗 − 𝑤ˆ 𝑗 ) 2 , H−1 𝑗𝑗

Ideally, this framework is theoretically optimal for the objective in Eq. (3) when the full Hessian H, computed over the true input distribution, can be obtained. In practice, however, H is replaced by Ĥ, an estimate from a small calibration set, thus making the method susceptible to sampling noise. In our experiments, this estimation error becomes especially problematic at low bit-widths and can lead to the well-known degradation in generation quality [28, 29]. This motivates our central question: under limited calibration data, are all entries of Ĥ equally reliable, or do certain parts remain statistically stable while others are dominated by noise?

𝛿𝑤 = −

𝑤 𝑗 − 𝑤ˆ 𝑗 −1 H 𝑗,: . H−1 𝑗𝑗

(5)

The diagonal term H−1 𝑗 𝑗 acts as a sensitivity score for coordinate 𝑗, while the off-diagonal entries in H−1 𝑗,: propagate the quantization error to other coordinates.

Motivation

In our preliminary analysis, we observe that the off-diagonal entries of the estimated Hessian are highly unstable and sensitive to the choice of calibration samples, whereas the diagonal entries remain consistent. This aligns with conventional observations that limited-sample estimation of high-dimensional covariance or curvature is often unreliable, and that off-diagonal fluctuations of the entries can aggregate into estimation error [5, 16, 26, 40, 41]. However, in Hessian-based PTQ for LLMs [7, 17, 18, 43], this issue has mostly been discussed through calibration sensitivity or numerical robustness, rather than as a stability problem of the error compensation across calibration batches. This is partly because they are designed to leverage an accurate local curvature estimate for better quantization, rather than to enforce consistency across different calibration samples. We argue that the essential focus should instead be batchstable optimization: find the curvature estimate that yield consistent accuracy and performance throughout batches. To quantify this effect, we estimate Hessians using calibration sets ranging from 8 to 2048 samples and compare them to a reference Hessian computed from 4096 samples. Fig. 1(a) reports the 𝐿1 difference between each estimate and the reference, computed separately over diagonal and off-diagonal entries. The diagonal entries stabilize quickly even with small calibration sets (yielding a small 𝐿1 ), while the off-diagonal entries remain unstable even with much larger sets. Thus, the overall difference in Hessian estimates is dominated by the off-diagonal terms, i.e., feature-correlation components. To systematically examine this issue, we adopt a linear shrinkage estimator for the Hessian. Let Ĥ = 𝑋𝑋 ⊤ denote the Hessian estimate from a calibration set 𝑋 , and decompose it into its diagonal and off-diagonal components: D = diag( Ĥ) and O = Ĥ − D. We then define a shrinkage family H̃(𝜌) = D + 𝜌O, which scales the off-diagonal terms by 𝜌 ∈ [0, 1]. To evaluate statistical stability, we compute H̃(𝜌) from two independent calibration sets 𝐴 and 𝐵, and measure their discrepancy via the normalized 𝐿1 difference: 𝑅(𝜌) =

∥ H̃𝐴 (𝜌) − H̃𝐵 (𝜌) ∥ 1 ∥ H̃𝐴 (𝜌) ∥ 1

=

∥ΔD + 𝜌ΔO∥ 1 , ∥D𝐴 + 𝜌O𝐴 ∥ 1

(6)

Robust Ultra Low-Bit PTQ via Stable Diagonal Curvature Estimate

EuroMLSys ’26, April 27–30, 2026, Edinburgh, Scotland Uk

Figure 2. Normalized histogram of diagonal and off-diagonal SNR values. Diagonal entries show a sharp high-SNR peak, while off-diagonal entries are dominated by low-SNR tail. Measured at the 10th transformer layer of Llama-2-7B. Figure 1. (a) Relative ℓ1 error of the Hessian estimate computed from 𝑛 calibration samples against a reference with 𝑛 = 4096 samples, (b) Relative ℓ1 error between two independent 128-sample sets over 100 trials. Results are measured at the 10th transformer layer of Llama-2-7B (𝑠𝑒𝑞 = 2048). where ΔD=D𝐴 −D𝐵 and ΔO=O𝐴 −O𝐵 . As 𝜌 increases beyond 𝜌 0 ≈ ∥D𝐴 ∥ 1 /∥O𝐴 ∥ 1 , the metric quickly saturates to 𝑅(𝜌) ≈ ∥ΔO∥ 1 /∥O𝐴 ∥ 1 . This is because O contains 𝑂 (𝑑 2 ) entries, so typically ∥O𝐴 ∥ 1 ≫ ∥D𝐴 ∥ 1 , making the off-diagonal terms to dominate the estimate and its variability. We observe this saturation consistently across 100 random pairs of calibration sets 𝐴 and 𝐵 (Fig. 1(b)): even for small 𝜌 > 0, the Hessian estimates remain highly sensitive to sampling noise. Moreover, we examine individual entries of the Hessian estimate and empirically measure their signal-to-noise ratio (SNR) across calibration samples, defined as: |E(𝐻ˆ𝑖 𝑗 )| . (7) SNR𝑖 𝑗 = Std(𝐻ˆ𝑖 𝑗 ) As shown in Figure 2, the SNR is substantially lower for off-diagonal entries than for diagonal ones. We can also interpret this from the perspective of sampleto-sample (i.e., batch-to-batch) variation. Consider two randomly sampled calibration sets 𝐴 and 𝐵 and the difference in (𝐵) the off-diagonal entries, Δ𝑂𝑖 𝑗 = 𝑂𝑖(𝐴) 𝑗 −𝑂 𝑖 𝑗 . Since Std(Δ𝑂 𝑖 𝑗 ) = √ 2 Std(𝑂𝑖 𝑗 ), the normalized variation of this difference is √ Std(Δ𝑂𝑖 𝑗 )/ E[𝑂𝑖 𝑗 ] ≈ 2/SNR𝑖 𝑗 . As shown in Fig. 2, SNR𝑖 𝑗 is very low for off-diagonal entries; hence, the batch-to-batch variation of off-diagonal terms is large and even comparable to their average magnitude. This in turn induce large ∥ΔO∥ 1 /∥O𝐴 ∥ 1 , consistent with Fig. 1(b) when 𝜌 > 0. Based on the above analysis, we emphasize the importance of batch-to-batch stability in Hessian-based quantization. While off-diagonal curvature terms can in principle reduce bias by accounting for empirical feature correlations, their reliable estimation requires much larger calibration sets that can be used for practical efficiency. In realistic PTQ settings, where only hundreds to a few thousand samples are used, these off-diagonal estimates remain statistically unstable

with high variance. Consequently, aggressively shrinking the off-diagonal component (i.e., 𝜌 ≈ 0) leads to more stable curvature estimates (with increased bias but substantially reduced variance) and, as we will show, better quantization performance – particularly in ultra low-bit configuration. Motivated by this, we propose DASH-Q, which exploits only diagonal components of Hessian for stable optimization. By doing so, we can decouple the optimization into independent weighted least-squares subproblems, improving both robustness and efficiency for low-bitwidth quantization.

5

Methodology

Based on the motivation, we approximate the Hessian H as a diagonal matrix D = diag(ℎ 11, ℎ 22, . . . , ℎ𝑑𝑖𝑛 𝑑𝑖𝑛 ), where Í𝑁 2 ℎ 𝑗 𝑗 = 𝑘=1 𝑥 𝑗𝑘 represents the feature importance of the 𝑗-th input channel. By substituting H with D in the reconstruction objective, the previously coupled multivariate optimization problem is decomposed into 𝑑𝑖𝑛 independent scalar sub-problems for each weight element 𝑤𝑖 𝑗 within a row: L (𝑊ˆ 𝑖,: ) ≈

𝑑𝑖𝑛 ∑︁

ℎ 𝑗 𝑗 (𝑤𝑖 𝑗 − 𝑤ˆ 𝑖 𝑗 ) 2 .

(8)

𝑗=1

This intra-row decoupling eliminates the dependency between input channels during optimization. Consequently, the quantization of each weight element can be treated as an independent 1D weighted least square problem, where the diagonal Hessian elements ℎ 𝑗 𝑗 serve as importance weights that prioritize the preservation of key features. For a given quantization group G, we formulate the task of finding quantization parameters 𝑠 and 𝑧 as follows: ∑︁ min ℎ 𝑗 𝑗 (𝑤 𝑗 − (𝑠 · 𝑞 𝑗 − 𝑧)) 2 + 𝜆𝑠 2 . (9) 𝑠,𝑧 𝑗∈G

The ridge regularization 𝜆𝑠 2 ensures numerical stability by preventing the scaling factor 𝑠 from diverging in sparse weight groups where the weighted variance is minimal.

EuroMLSys ’26, April 27–30, 2026, Edinburgh, Scotland Uk

Jaemin Kim, Sungkyun Kim, Junyeol Lee, and Jiwon Seo

By setting the partial derivatives with respect to 𝑠 and 𝑧 to zero, we derive the optimal closed-form solutions as follows: Í ¯ℎ) Covℎ (𝑊 , 𝑄) 𝑗 ∈ G ℎ 𝑗 𝑗 (𝑞 𝑗 − 𝑞¯ℎ ) (𝑤 𝑗 − 𝑤 ∗ 𝑠 = = Í (10) 2 Varℎ (𝑄) + 𝜆 𝑗 ∈ G ℎ 𝑗 𝑗 (𝑞 𝑗 − 𝑞¯ℎ ) + 𝜆 𝑧 ∗ = 𝑠 ∗ · 𝑞¯ℎ − 𝑤¯ ℎ ,

(11)

Í ℎ 𝑤 where 𝑤¯ ℎ and 𝑞¯ℎ denote the weighted means Í ℎ𝑗 𝑗𝑗 𝑗 𝑗 and Í ℎ 𝑞 Í 𝑗 𝑗 𝑗 , respectively. The solution ensure optimality within ℎ𝑗 𝑗

the objective for a fixed set of quantized integers 𝑄. Detailed mathematical derivations for the optimal parameters 𝑠 ∗ and 𝑧 ∗ are provided in Appendix A. Since the optimal quantized integers 𝑄 depend on 𝑠 and 𝑧, and vice versa, we employ an iterative optimization process following the coordinate descent algorithm. Starting from an initial estimate of 𝑠 and 𝑧 as following: max(𝑊 ) − min(𝑊 ) , 𝑧 (0) = − min(𝑊 ), (12) 2𝑏 − 1 we refine the parameters by alternating between two steps: 1. Integer Refinement: Fix 𝑠 (𝑡 ) and 𝑧 (𝑡 ) , and update the   quantized integers 𝑄 𝑗(𝑡 ) = clip( (𝑊 𝑗 + 𝑧 (𝑡 ) )/𝑠 (𝑡 ) ). 2. Parameter Regression: Fix 𝑄 (𝑡 ) , and compute 𝑠 (𝑡 +1) and 𝑧 (𝑡 +1) using Eq. (10) and (11). This repetitive approach rapidly converges to a stable solution, typically within a few iterations, effectively minimizing the reconstruction error while remaining robust to the sampling noise. The entire process is expressed in Algorithm 1. 𝑠 (0) =

6

Evaluation

6.1

Experimental Setup

We implement DASH-Q and all comparative baselines using Python 3.12.3 and PyTorch 2.9.0, executing all experiments on a compute node equipped with an AMD EPYC 9755 CPU and an NVIDIA RTX PRO 6000 GPU under CUDA 13.0. We evaluate six LLMs: Llama-3.1-8B-Instruct [14], Qwen3-14B [45], DeepSeek-Moe-16B [10], Phi-3.5-Moe [1], and Mixtral-8x7B-Instruct-v0.1 [22] for accuracy, and Llama-2-7b [42] for qualitative analysis. For each model and methods, we perform calibration using 128 samples with each sample with sequence length of 2048 randomly drawn from the WikiText-2 [31] training set and evaluate perplexity on the test set. General reasoning capabilities are assessed across eight zero-shot tasks—ARC (Easy/Challenge) [9], PIQA [4], Hellaswag [47], Winogrande [37], BoolQ [8], SIQA [38], and OpenBookQA [32] —via the LM Evaluation Harness [20]. All experiments perform weight-only quantization with group sizes of 128 for 4-bit, 64 for 3-bit, and 32 for 2-bit precision. Each baseline is implemented based on their official repository and recommended configuration to ensure a fair comparison. DASH-Q is optimized with 𝑇 = 9 iterations (𝛼 = 0.5, 𝜆 = 10−2 ), while AWQ [29] employs grid search size of 20. Second-order Hessian based methods, including GPTQ [18],

Algorithm 1: Layer-wise DASH-Q Procedure Input: Pre-trained model M with 𝐿 layers, calibration data 𝑋 , bit-width 𝑏, iterations 𝑇 , ridge 𝜆, smoothing 𝛼 Output: Quantized model M̂ 1: for each layer 𝑙 = 1 to 𝐿 do 2: 𝑋 (𝑙 ) ← Accumulated Input activations of layer 𝑙 Í 3: 𝐷 (𝑙 ) ← diag( (𝑋 (𝑙 ) ) 2 ) 4: for each weight group 𝑊𝑔(𝑙 ) ∈ 𝑊 (𝑙 ) do 5: Initialize 𝑠, 𝑧 of 𝑊𝑔(𝑙 ) (Eq. (12)) 6: for 𝑡 = 0 to 𝑇 − 1 do 7: Step A: Coordinate Descent m  j 𝑄 (𝑡 ) ← clip (𝑊𝑔(𝑙 ) + 𝑧)/𝑠 , 0, 2𝑏 − 1 9: Step B: Weighted least squares (Eq.(10)) 10: 𝑠 ← Covℎ (𝑊𝑔(𝑙 ) , 𝑄 (𝑡 ) )/(Varℎ (𝑄 (𝑡 ) ) + 𝜆) 11: 𝑧 ← 𝑠𝑞¯ℎ − 𝑤¯ ℎ 12: end for ˆ 13: M𝑔(𝑙 ) ← 𝑄 (𝑇 ) , 𝑠, 𝑧 14: end for 15: propagate 𝑋 (𝑙+1) ← Mˆ(𝑙 ) (𝑋 (𝑙 ) ) 16: end for 17: return M̂ 8:

QuIP [6], OWQ [27], and QuaRot [3], utilize a block size of 128. Rotation-based schemes are implemented by applying hadamard (for QuaRot) or butterfly (for QuIP) rotation prior to error compensation process and subsequently reverting the transformation to simulate quantized inference. For OWQ, the outlier count is set to 128. Both scaling factor and zero-points are kept in fp16 for all methods. 6.2

Overall Accuracy

Table 1 compares the performance of DASH-Q against six PTQ baselines across five evaluation models, covering both dense and MoE architectures. RTN denotes the naive baseline that quantizes all parameters using Eq. (1), (2). Under 4-bit precision on Llama-3.1-8B, DASH-Q achieves 66.90% average zero-shot accuracy, closely matching AWQ (67.07%) while requiring 64.7× less quantization time. DASH-Q also remains comparable to second-order and rotation-based approaches such as OWQ (66.98%) and QuaRot (66.70%). The advantage of DASH-Q is more pronounced in ultra low-bit regimes. At 2-bit on Llama-3.1-8B, DASH-Q preserves reasoning quality with 56.52% average accuracy, outperforming OWQ by 14.01% (42.51%) and improving over GPTQ by 1.59× (35.66%). While both ours and OWQ prioritize salient feature reconstruction, OWQ’s dependence on full-Hessian compensation is more susceptible to fitting spurious feature correlations under limited calibration data. In contrast, our diagonal approximation suppresses such noise and yields markedly stronger downstream reasoning. This trend extends consistently to larger models. A notable observation appears in the Qwen3-14B results at 2-bit precision, where DASH-Q and QuaRot achieve nearly identical

Robust Ultra Low-Bit PTQ via Stable Diagonal Curvature Estimate

EuroMLSys ’26, April 27–30, 2026, Edinburgh, Scotland Uk

Table 1. Performance evaluation of DASH-Q against six PTQ baselines on Llama-3.1-8B-Instruct, Qwen3-14B, DeepSeek-MoE-16B, Phi-3.5-MoE, and Mixtral-8x7B-Instruct-v0.1. We report WikiText-2 perplexity (PPL), zero-shot accuracies, and total quantization time. W-bit / GS represent weight bits and group size, respectively. Model

Zero-Shot Reasoning Tasks

Method W-bit / GS PPL ↓

Time (s)

ARC-C ARC-E BoolQ Hella OBQA PIQA SIQA Wino Avg ↑ Baseline

Llama-3.1-8B

Qwen3-14B

7.22

55.20

79.71

85.41

79.54

45.00

81.07 42.27 73.88 67.76

RTN 4 / 128 AWQ 4 / 128 GPTQ 4 / 128 QuIP 4 / 128 QuaRot 4 / 128 OWQ 4.30 / 128 DASH-Q 4 / 128

7.89 7.53 7.48 7.46 7.45 7.42 7.55

53.41 54.01 52.56 52.30 52.99 53.07 54.10

77.06 78.58 76.85 77.90 78.07 79.00 78.41

84.53 84.31 84.53 85.17 85.29 84.07 83.82

78.15 78.92 77.00 78.79 78.55 78.91 78.52

43.00 43.60 41.00 43.60 43.20 44.20 43.80

81.01 41.45 74.27 66.61 2.19 80.14 42.27 74.74 67.07 4466.09 79.65 39.71 73.72 65.63 405.38 80.41 41.20 74.35 66.72 705.70 79.76 41.76 73.95 66.70 768.65 80.52 41.56 74.51 66.98 397.54 81.01 42.37 73.16 66.90 69.06

RTN AWQ GPTQ QuIP QuaRot OWQ DASH-Q

3 / 64 3 / 64 3 / 64 3 / 64 3 / 64 3.32 / 64 3 / 64

11.02 8.90 8.33 8.24 8.19 8.10 8.52

44.54 46.16 24.15 49.06 51.37 48.63 50.00

65.74 71.59 44.28 76.18 75.93 74.75 75.13

75.63 80.92 39.82 84.01 80.92 82.57 82.78

71.38 75.27 37.55 76.23 76.36 76.99 76.27

37.40 41.00 28.60 41.00 41.40 42.00 42.60

77.04 35.98 69.69 59.67 1.87 79.43 39.15 71.67 63.15 4458.15 60.94 34.14 53.99 40.43 405.99 78.29 40.17 72.77 64.71 705.04 78.78 40.02 70.80 64.45 765.55 78.89 39.51 72.61 64.49 397.64 80.03 40.38 72.69 64.99 69.59

RTN AWQ GPTQ QuIP QuaRot OWQ DASH-Q

2 / 32 2 / 32 2 / 32 2 / 32 2 / 32 2.33 / 32 2 / 32

42446.44 24.57 128.34 26.11 28.33 25.94 21.20 26.02 20.78 27.39 15.80 25.94 16.98 38.57

27.19 33.21 26.47 36.07 36.62 38.72 66.79

37.83 50.12 38.17 50.06 49.33 51.96 76.12

26.71 36.38 25.96 37.01 41.29 46.79 62.39

28.20 24.00 30.60 28.00 29.20 30.80 34.80

52.23 34.08 51.07 35.23 1.91 57.51 34.08 50.04 38.93 4430.02 53.97 34.49 49.64 35.66 406.31 57.07 33.32 51.62 39.90 700.24 57.94 33.67 52.57 41.00 777.88 58.65 34.54 52.64 42.51 399.97 72.91 34.14 66.46 56.52 69.97

Baseline

FP16

8.65

60.6

82.7

89.4

78.7

46.2

80.1

RTN AWQ GPTQ QuIP QuaRot OWQ DASH-Q

2 / 32 2 / 32 2 / 32 2 / 32 2 / 32 2.26 / 32 2 / 32

299.83 15.12 13.65 11.81 11.79 11.85 11.38

23.46 37.54 33.70 36.01 37.29 36.01 53.75

32.87 61.07 52.10 54.34 58.50 56.90 79.88

48.10 68.29 72.45 71.47 67.71 75.02 86.12

31.75 61.67 59.26 64.86 64.76 63.15 68.35

26.00 35.40 36.20 35.60 35.40 36.00 42.80

56.69 33.47 48.54 37.61 5.34 71.60 37.10 57.22 53.74 7719.32 68.66 35.36 56.35 51.76 801.29 70.51 37.41 60.30 53.81 1553.46 71.65 36.13 60.06 53.94 2096.17 72.74 34.90 59.27 54.25 804.41 75.46 39.10 69.93 64.42 126.39

FP16

6.51

45.82

69.53

73.18

77.17

43.80

79.49 39.46 70.32 62.35

2 / 32 2 / 32 2 / 32 2 / 32 2 / 32 2.92 / 32 2 / 32

1914.07 477.26 11.22 10.40 10.82 9.25 9.82

22.44 24.49 25.43 31.57 30.29 34.64 36.35

29.55 29.80 35.65 50.67 55.51 60.61 63.38

43.91 38.44 59.79 63.58 58.17 58.65 63.27

26.96 26.77 40.57 59.99 55.73 65.15 66.13

23.40 28.00 23.80 32.40 30.80 34.20 39.40

53.26 35.62 52.01 35.89 7.03 52.94 33.88 47.51 35.23 10299.41 60.45 34.70 51.14 41.44 1049.60 71.93 33.88 58.41 50.30 1368.66 71.38 35.31 56.51 49.21 1334.33 73.07 35.01 61.64 52.87 1028.94 76.06 35.72 64.40 55.59 138.31

Baseline RTN AWQ GPTQ DeepSeek-MoE-16B QuIP QuaRot OWQ DASH-Q

Phi-3.5-MoE

Mixtral-8x7B

FP16

44.5

72.8

69.4

-

Baseline

FP16

3.98

53.33

66.12

88.53

79.80

50.40

78.02 42.53 76.40 66.89

-

RTN AWQ GPTQ QuIP QuaRot OWQ DASH-Q

2 / 32 2 / 32 2 / 32 2 / 32 2 / 32 2.36 / 32 2 / 32

52.40 9.54 6.80 67.49 6.38 6.18 5.88

32.59 42.66 39.51 29.10 43.77 46.93 52.13

41.71 55.26 52.23 38.26 54.50 55.09 64.10

58.47 72.29 81.13 50.83 83.85 84.50 86.18

43.76 66.88 68.17 41.48 73.10 74.18 76.79

29.20 39.00 40.40 30.20 44.20 44.60 45.80

57.78 34.39 55.01 44.11 72.47 35.77 66.77 56.39 68.06 36.69 65.19 56.42 58.81 32.50 51.46 41.58 v72.52 37.36 63.93 59.16 72.47 39.41 67.80 60.62 77.31 41.50 73.40 64.65

15.22 8843.61 1394.85 2116.30 2080.00 1395.27 213.04

Baseline

FP16

4.14

65.53

84.26

88.69

86.46

50.00

84.22 46.93 77.51 72.95

RTN AWQ GPTQ QuIP QuaRot OWQ DASH-Q

2 / 32 2 / 32 2 / 32 2 / 32 2 / 32 2.31 / 32 2 / 32

1103.32 7.53 12.74 5.94 5.94 5.80 5.51

25.68 43.60 25.43 53.24 51.02 51.28 58.36

25.42 66.25 30.93 74.75 75.08 74.62 80.81

40.09 69.91 44.19 80.09 79.63 77.68 87.06

26.99 65.21 30.35 78.47 77.96 72.65 80.35

24.00 36.00 25.40 43.80 44.80 40.00 46.80

51.20 35.01 48.22 34.58 76.71 36.03 57.54 56.41 58.32 35.52 50.12 37.53 79.43 41.86 73.24 65.61 77.58 40.74 70.24 64.63 79.00 36.90 61.40 61.69 81.99 45.39 76.95 69.72

33.17 9683.59 2313.41 4571.50 5002.99 2288.56 235.46

EuroMLSys ’26, April 27–30, 2026, Edinburgh, Scotland Uk

Jaemin Kim, Sungkyun Kim, Junyeol Lee, and Jiwon Seo

Figure 3. Each plot shows the mapping of original weights (𝑊 ) to quantized levels (𝑄) by the affine mapping (blue line). Points are colored by their normalized log importance (𝑙𝑜𝑔(𝑑𝑖𝑎𝑔( Ĥ))). Points closer to the blue line indicate lower quantization error. perplexity (11.38 vs. 11.79), yet our method maintains a substantial 10.48% point lead in reasoning accuracy. As analyzed in Section 4, complex second-order methods can fit noisy feature dependencies rather than preserving global logic. The same pattern holds across the remaining architectures. On DeepSeek-MoE-16B, Phi-3.5-MoE, and Mixtral-8x7B, DASH-Q consistently achieves the best average zero-shot accuracy, outperforming the strongest competing baseline by 2.72%, 4.03%, and 4.11% points, respectively. In particular, the DeepSeek-MoE result again shows that lower perplexity does not necessarily translate into better reasoning performance, whereas on Phi-3.5-MoE and Mixtral-8x7B, DASH-Q achieves both the best accuracy and the lowest perplexity. Overall, in the 2-bit regime, DASH-Q achieves the highest average zero-shot accuracy on all five evaluation models, improving over the strongest competing baseline by 1.14× on average (7.01%) and by up to 1.33× (14.01%). It also maintains strong perplexity score although perplexity alone does not fully represent downstream reasoning quality. Furthermore, our method is up to 74.5× faster in quantization time than other optimization-based PTQ baselines, demonstrating that a diagonal, noise-robust curvature approximation scales effectively across both dense and MoE architectures.

7

Ablation Study

7.1

Qualitative Analysis of Weight Reconstruction

To qualitatively assess the effectiveness of the proposed weighted regression, we visualize the weight reconstruction behavior across different quantization schemes in Fig. 3. We extract a randomly selected weight group from Llama2-7b model to analyze the mapping precision. Each plot in Fig. 3 illustrates the mapping from original weights (𝑊 ) to discrete quantized levels (𝑄), where the solid line represents the affine mapping 𝑦 = 𝑠 · 𝑞 − 𝑧. Points closer to this solid line indicate a lower quantization error, as the weights are more accurately preserved during the mapping process.

The visualization reveals distinct failure modes in existing baselines. RTN and GPTQ show significant rounding error, since weights are uniformly mapped based on a rigid min-max interval without considering feature importance. Although GPTQ attempts to mitigate this error by compensating for quantization errors through subsequent features, individual weight mapping remains suboptimal. Conversely, AWQ attempts to preserve salient weights by scaling them to the limits of the dynamic range (indicated by white points). This expansion of the quantization scale leads to significant increase in grid size, which leads to majority of the remaining weights compressed into a few quantization grids, increasing the overall distortion. In contrast, QuIP employs a randomized orthogonal transformation to achieve importance homogenization, reflected in its uniform color distribution. This process mitigates the risk of catastrophic errors by ensuring that no single critical feature suffers from disproportionately high quantization noise. However, by spreading importance uniformly, QuIP inherently forfeits the opportunity to achieve better representation for truly salient feature. Unlike these baselines, DASH-Q achieves the tightest alignment with the ideal mapping. By treating quantization as a weighted regression problem, our method explicitly prioritizes the reconstruction of salient features (indicated by darker red nodes). This approach ensures that the most critical weights are accurately restored without sacrificing the resolution of the overall distribution, effectively mitigating both rounding noise and grid collapse. 7.2

Sensitivity to Calibration Data Size

We evaluate the sensitivity of DASH-Q to the calibration data size (𝑛) compared to GPTQ using the Llama-2-7b model (Fig. 4). Each sample contains 𝑠𝑒𝑞 = 2048 tokens. DASH-Q demonstrates surprisingly stable performance across all evaluated sizes, maintaining perplexity between 8.22 and 8.36 even with a scarce calibration sample (𝑛 = 2). In contrast, GPTQ suffers from severe numerical instability, with perplexity diverging beyond 102 for 𝑛 ≤ 4. Although GPTQ

Robust Ultra Low-Bit PTQ via Stable Diagonal Curvature Estimate

Figure 4. Comparison of perplexity between DASH-Q and GPTQ on Llama-2-7b across varying calibration sample sizes. It shows robustness against calibration data scarcity. eventually stabilizes as 𝑛 increases, its performance consistently remains above the perplexity compared to our method. Notably, GPTQ’s perplexity begins to fluctuate or even slightly degrades as 𝑛 exceeds 28 , suggesting that even our large-scale empirical Hessian may yet have failed to capture a robust representation for each entry. This observation emphasizes the theoretical analysis in Section 4, illustrating that while full Hessian estimation in second-order methods is batch-sensitive due to the feature dependency noises, our diagonal approximation effectively filters such noise to ensure robust quantization. This high efficiency is particularly advantageous in scenarios where calibration data is restricted or rapid, low-overhead quantization is required. 7.3

Analysis on Optimization Stability

To validate the efficiency and convergence of DASH-Q’s coordinate descent solver, we track perplexity and the scaling factors 𝑠 across varying iteration steps 𝑇 . As shown in Fig. 5 (left), the perplexity drops sharply within a few iterations and reaches a stable floor with negligible fluctuations thereafter. In the right plot, we quantify numerical convergence by aggregating the normalized scale change 𝛿𝑠 of quantization groups that contain important feature channels from all layers. Despite the early convergence in perplexity, the right plot shows that it continues to decrease over multiple steps. This behavior suggests that while the closed-form update in Eq. (10) corrects the dominant reconstruction error from salient features in the initial steps, subsequent iterations refine quantization boundaries for the remaining lessimportant features to better align the overall distribution. Because accuracy improvements become marginal beyond 𝑇 = 10, we fix 𝑇 = 9 for all experiments to achieve a practical balance between quantization time and performance. 7.4

Inference Optimization and Deployment

DASH-Q is compatible with standard inference engines, since it preserves the original model structure and avoids

EuroMLSys ’26, April 27–30, 2026, Edinburgh, Scotland Uk

Figure 5. (Left) Perplexity and quantization time accross iteration steps. (Right) Convergence of scaling factors (|𝑠𝑡 − 𝑠𝑡 −1 |/|𝑠 0 |) for quantization groups containing key features across layers. Both are measured with Llama-2-7b model.

the auxiliary runtime operations or architectural modifications required by several prior PTQ schemes. In contrast, rotation-based methods such as QuaRot and QuIP, as well as outlier-aware approaches such as OWQ, typically introduce additional inference-time components, including Hadamard transforms or customized kernels. DASH-Q instead operates directly on discrete weights under the standard affine quantization form, allowing the memory and bandwidth benefits of low-bit weight-only quantization to be realized without changing the inference pipeline. As a result, DASH-Q can be readily deployed on existing LLM inference engines such as vLLM [25] and TensorRT-LLM [33]. Its standard quantized representation is also compatible with optimized weightonly quantization backends, including Marlin [19] and GemLite [13], without requiring custom kernel implementations.

8

Conclusion

This paper introduces DASH-Q, a statistically robust PTQ framework designed to overcome the overfitting limitations of second-order optimization in ultra low-bit regimes. By identifying off-diagonal Hessian elements as a primary source of sampling noise, we leverage a stable diagonal approximation to redefine weight reconstruction as an iterative weighted least square problem. Our results confirm that DASH-Q consistently outperforms SOTA PTQ baselines in 2-bit precision, achieving downstream zero-shot accuracy improvements of 1.14× on average and up to 1.33× over the strongest competing baseline, while showing competitive perplexity and robust performance with very small calibration data. Crucially, by maintaining a standard weight format without auxiliary transformations, DASH-Q allows deployment on production-ready inference engines with no additional overhead. Ultimately, this work provides a scalable solution for the practical ultra low-bit compressed LLMs.

EuroMLSys ’26, April 27–30, 2026, Edinburgh, Scotland Uk

Acknowledgement This work was supported by Institute of Information & communications Technology Planning & Evaluation (IITP) grant funded by the Korea government(MSIT) (NO.RS-2021II211343, Artificial Intelligence Graduate School Program (Seoul National University), No.RS-2025-02214497, Development of low-level optimization program API technology for AI semiconductors, No.RS-2025-02263167, Development of Integrated Resource Management Technology for AI Semiconductors, IITP-2026-RS-2021-II211817, ITRC(Information Technology Research Center)), This work was also supported by the Basic Science Research Program through the National Research Foundation of Korea(NRF) funded by the Ministry of Education(RS-2026-25476387), and Automation and System Research Institute at Seoul National University (No.041820250030). Jiwon Seo is the corresponding author.

References [1] Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, et al. Phi-3 technical report: A highly capable language model locally on your phone, 2024. URL https://arxiv. org/abs/2404.14219, 2(6):4, 2024. [2] Yamato Arai and Yuma Ichikawa. Quantization error propagation: Revisiting layer-wise post-training quantization. arXiv preprint arXiv:2504.09629, 2025. [3] Saleh Ashkboos, Amirkeivan Mohtashami, Maximilian L Croci, Bo Li, Pashmina Cameron, Martin Jaggi, Dan Alistarh, Torsten Hoefler, and James Hensman. Quarot: Outlier-free 4-bit inference in rotated llms. Advances in Neural Information Processing Systems, 37:100213–100240, 2024. [4] Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. Piqa: Reasoning about physical commonsense in natural language. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pages 7432–7439, 2020. [5] Richard H Byrd, Samantha L Hansen, Jorge Nocedal, and Yoram Singer. A stochastic quasi-newton method for large-scale optimization. SIAM Journal on Optimization, 26(2):1008–1031, 2016. [6] Jerry Chee, Yaohui Cai, Volodymyr Kuleshov, and Christopher M De Sa. Quip: 2-bit quantization of large language models with guarantees. Advances in Neural Information Processing Systems, 36:4396–4429, 2023. [7] Everlyn Asiko Chimoto, Mostafa Elhoushi, and Bruce A Bassett. Calibrating beyond english: Language diversity for better quantized multilingual llm. arXiv preprint arXiv:2601.18306, 2026. [8] Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. Boolq: Exploring the surprising difficulty of natural yes/no questions. arXiv preprint arXiv:1905.10044, 2019. [9] Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018. [10] Damai Dai, Chengqi Deng, Chenggang Zhao, RX Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Yu Wu, et al. Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1280–1297, 2024. [11] Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale.

Jaemin Kim, Sungkyun Kim, Junyeol Lee, and Jiwon Seo Advances in neural information processing systems, 35:30318–30332, 2022. [12] Zhen Dong, Zhewei Yao, Amir Gholami, Michael W Mahoney, and Kurt Keutzer. Hawq: Hessian aware quantization of neural networks with mixed-precision. In Proceedings of the IEEE/CVF international conference on computer vision, pages 293–302, 2019. [13] Dropbox AI. Gemlite: Fast low-bit matmul kernels in triton, 2024. [14] Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The llama 3 herd of models. arXiv e-prints, pages arXiv–2407, 2024. [15] Vage Egiazarian, Roberto L Castro, Denis Kuznedelev, Andrei Panferov, Eldar Kurtic, Shubhra Pandit, Alexandre Marques, Mark Kurtz, Saleh Ashkboos, Torsten Hoefler, et al. Bridging the gap between promise and performance for microscaling fp4 quantization. arXiv preprint arXiv:2509.23202, 2025. [16] Michael Fleermann and Johannes Heiny. High-dimensional sample covariance matrices with curie-weiss entries. arXiv preprint arXiv:1910.12332, 2019. [17] Elias Frantar and Dan Alistarh. Optimal brain compression: A framework for accurate post-training quantization and pruning. Advances in Neural Information Processing Systems, 35:4475–4488, 2022. [18] Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. Gptq: Accurate post-training quantization for generative pre-trained transformers. arXiv preprint arXiv:2210.17323, 2022. [19] Elias Frantar, Roberto L Castro, Jiale Chen, Torsten Hoefler, and Dan Alistarh. Marlin: Mixed-precision auto-regressive parallel inference on large language models. arXiv preprint arXiv:2408.11743, 2024. [20] Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. The language model evaluation harness, 07 2024. [21] Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, and Dmitry Kalenichenko. Quantization and training of neural networks for efficient integerarithmetic-only inference. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2704–2713, 2018. [22] Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. Mixtral of experts. arXiv preprint arXiv:2401.04088, 2024. [23] Jaemin Kim, Hongjun Um, Sungkyun Kim, Yongjun Park, and Jiwon Seo. Flexiq: Adaptive mixed-precision quantization for latency/accuracy trade-offs in deep neural networks. arXiv preprint arXiv:2510.02822, 2025. [24] Youngseok Kim, Junyeol Lee, Younghoon Kim, and Jiwon Seo. Robust quantization of deep neural networks. In Proceedings of the 29th International Conference on Compiler Construction, pages 74–84, 2020. [25] Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedattention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles, 2023. [26] Olivier Ledoit and Michael Wolf. A well-conditioned estimator for large-dimensional covariance matrices. Journal of multivariate analysis, 88(2):365–411, 2004. [27] Changhun Lee, Jungyu Jin, Taesu Kim, Hyungjun Kim, and Eunhyeok Park. Owq: Outlier-aware weight quantization for efficient fine-tuning and inference of large language models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 13355–13364, 2024.

Robust Ultra Low-Bit PTQ via Stable Diagonal Curvature Estimate [28] Jemin Lee, Sihyeong Park, Jinse Kwon, Jihun Oh, and Yongin Kwon. Exploring the trade-offs: Quantization methods, task difficulty, and model size in large language models from edge to giant. arXiv preprint arXiv:2409.11055, 2024. [29] Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, WeiChen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han. Awq: Activation-aware weight quantization for on-device llm compression and acceleration. Proceedings of machine learning and systems, 6:87–100, 2024. [30] Shuming Ma, Hongyu Wang, Lingxiao Ma, Lei Wang, Wenhui Wang, Shaohan Huang, Li Dong, Ruiping Wang, Jilong Xue, and Furu Wei. The era of 1-bit llms: All large language models are in 1.58 bits. arXiv preprint arXiv:2402.17764, 2024. [31] Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models. arXiv preprint arXiv:1609.07843, 2016. [32] Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. Can a suit of armor conduct electricity? a new dataset for open book question answering. arXiv preprint arXiv:1809.02789, 2018. [33] NVIDIA Corporation. Tensorrt-llm, 2023. [34] Hyungjun Oh, Kihong Kim, Jaemin Kim, Sungkyun Kim, Junyeol Lee, Du-seong Chang, and Jiwon Seo. Exegpt: Constraint-aware resource scheduling for llm inference. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2, pages 369–384, 2024. [35] Hyungjun Oh, Junyeol Lee, Hyeongju Kim, and Jiwon Seo. Out-oforder backprop: An effective scheduling technique for deep learning. In Proceedings of the Seventeenth European Conference on Computer Systems, pages 435–452, 2022. [36] Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. Zero: Memory optimizations toward training trillion parameter models. In SC20: international conference for high performance computing, networking, storage and analysis, pages 1–16. IEEE, 2020. [37] Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99–106, 2021. [38] Maarten Sap, Hannah Rashkin, Derek Chen, Ronan LeBras, and Yejin Choi. Socialiqa: Commonsense reasoning about social interactions. arXiv preprint arXiv:1904.09728, 2019. [39] Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. Megatron-lm: Training multibillion parameter language models using model parallelism. arXiv preprint arXiv:1909.08053, 2019. [40] Alexander Soen and Ke Sun. On the variance of the fisher information for deep learning. Advances in Neural Information Processing Systems, 34:5708–5719, 2021. [41] Alexander Soen and Ke Sun. Trade-offs of diagonal fisher information matrix estimators. Advances in Neural Information Processing Systems, 37:5870–5912, 2024. [42] Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and finetuned chat models. arXiv preprint arXiv:2307.09288, 2023. [43] Miles Williams and Nikolaos Aletras. On the impact of calibration data in post-training quantization and pruning. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 10100–10118, 2024. [44] Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han. Smoothquant: Accurate and efficient post-training quantization for large language models. In International conference on machine learning, pages 38087–38099. PMLR, 2023. [45] An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025.

EuroMLSys ’26, April 27–30, 2026, Edinburgh, Scotland Uk [46] Zhihang Yuan, Lin Niu, Jiawei Liu, Wenyu Liu, Xinggang Wang, Yuzhang Shang, Guangyu Sun, Qiang Wu, Jiaxiang Wu, and Bingzhe Wu. Rptq: Reorder-based post-training quantization for large language models. arXiv preprint arXiv:2304.01089, 2023. [47] Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830, 2019. [48] Jinhao Zhang Yunquan Zhang, Boyang Zhang, Jun Sun, Daning Cheng, et al. Hero-q: A general framework for stable low bit quantization via hessian conditioning. arXiv e-prints, pages arXiv–2601, 2026. [49] Yilong Zhao, Chien-Yu Lin, Kan Zhu, Zihao Ye, Lequn Chen, Size Zheng, Luis Ceze, Arvind Krishnamurthy, Tianqi Chen, and Baris Kasikci. Atom: Low-bit quantization for efficient and accurate llm serving. Proceedings of Machine Learning and Systems, 6:196–209, 2024.

A

Derivation of 𝑠 ∗ and 𝑧 ∗ (Eq. (10), (11))

For a group G, consider the weighted quadratic objective  2 ∑︁ min ℎ 𝑗 𝑗 𝑤 𝑗 − (𝑠 · 𝑞 𝑗 − 𝑧) + 𝜆𝑠 2, (13) 𝑠,𝑧 𝑗∈G

where ℎ 𝑗 𝑗 ≥ 0 denotes the diagonal Hessian importance, 𝜆 ≥ 0 is a scale regularizer, and {𝑞 𝑗 } 𝑗 ∈ G are the (fixed) integer codes for the current update. For brevity, we write ℎ 𝑗 := ℎ 𝑗 𝑗 . Note that we use quantization function 𝑞 = ⌊(𝑤 + 𝑧)/𝑠⌉ and a reconstruction 𝑤ˆ = 𝑠 · 𝑞 − 𝑧, where 𝑧 is a full-precision offset. This is algebraically equivalent to the conventional affine form ⌊𝑤/𝑠 + 𝑧𝑝 ⌉ by defining 𝑧𝑝 := 𝑧/𝑠, since (𝑤 + 𝑧)/𝑠 = 𝑤/𝑠 + 𝑧/𝑠. Thus, our formulation differs only by reparameterizing the zero-point offset before scaling. When 𝑞 𝑗 are fixed, Eq. (13) is a convex quadratic in (𝑠, 𝑧). Fix any scale 𝑠 and minimize over 𝑧: 2 ∑︁  𝑧 ∗ (𝑠) := arg min ℎ 𝑗 𝑤 𝑗 − (𝑠 · 𝑞 𝑗 − 𝑧) . (14) 𝑧 𝑗∈G

Here 𝑧 ∗ (𝑠) denotes the best 𝑧 for a given 𝑠 (i.e., the optimizer of the inner problem). Define the residual 𝑟 𝑗 (𝑠, 𝑧) := 𝑤 𝑗 − 𝑠 · 𝑞 𝑗 + 𝑧; then ∑︁ ∑︁ 𝜕 ∑︁ ℎ 𝑗 𝑟 𝑗 (𝑠, 𝑧) 2 = 2 ℎ 𝑗 𝑟 𝑗 (𝑠, 𝑧) = 2 ℎ 𝑗 (𝑤 𝑗 −𝑠 ·𝑞 𝑗 +𝑧). 𝜕𝑧 𝑗 𝑗 𝑗 (15) Setting the derivative to zero yields ∑︁ ∑︁ ∑︁ ℎ𝑗𝑤 𝑗 − 𝑠 ℎ𝑗𝑞𝑗 + 𝑧 ℎ 𝑗 = 0. (16) 𝑗

Let 𝐻 :=

Í

𝑗

𝑗

𝑗 ∈ G ℎ 𝑗 , and define weighted means

𝑤¯ ℎ :=

1 ∑︁ ℎ𝑗𝑤 𝑗, 𝐻 𝑗∈G

𝑞¯ℎ :=

1 ∑︁ ℎ𝑗𝑞𝑗 . 𝐻 𝑗∈G

(17)

Dividing Eq. (16) by 𝐻 gives the closed form 𝑧 ∗ (𝑠) = 𝑠 · 𝑞¯ℎ − 𝑤¯ ℎ .

(18)

Intuitively, Eq. (18) aligns the weighted mean of the reconstruction 𝑠 · 𝑞 𝑗 − 𝑧 with the weighted mean of 𝑤 𝑗 for the given scale 𝑠.

EuroMLSys ’26, April 27–30, 2026, Edinburgh, Scotland Uk

Jaemin Kim, Sungkyun Kim, Junyeol Lee, and Jiwon Seo

Now minimize Eq. (13) over 𝑠 and 𝑧 jointly. Taking the partial derivative of Eq. (13) with respect to 𝑠 gives #

"

∑︁ 𝜕 ∑︁ ℎ 𝑗 (𝑤 𝑗 − 𝑠𝑞 𝑗 + 𝑧) 2 + 𝜆𝑠 2 = 2 ℎ 𝑗 (𝑤 𝑗 −𝑠·𝑞 𝑗 +𝑧) (−𝑞 𝑗 )+2𝜆𝑠. 𝜕𝑠 𝑗 𝑗 (19)

Setting it to zero yields the stationarity condition ∑︁ ℎ 𝑗 𝑞 𝑗 (𝑤 𝑗 − 𝑠 · 𝑞 𝑗 + 𝑧) = 𝜆𝑠.

(20)

𝑗∈G

Substitute 𝑧 = 𝑧 ∗ (𝑠) from Eq. (18) and expand: ∑︁ ∑︁ ∑︁ 𝜆𝑠 = ℎ𝑗𝑞𝑗𝑤 𝑗 − 𝑠 ℎ 𝑗 𝑞 2𝑗 + 𝑧 ∗ (𝑠) ℎ𝑗𝑞𝑗 =

𝑗

𝑗

∑︁

∑︁

ℎ𝑗𝑞𝑗𝑤 𝑗 − 𝑠

𝑗

Using

Í

 ∑︁

𝑗

ℎ 𝑗 𝑞 2𝑗 + (𝑠𝑞¯ℎ − 𝑤¯ ℎ )

∑︁

𝑗

𝑗 ℎ 𝑗 𝑞 𝑗 = 𝐻 𝑞¯ℎ and

ℎ𝑗𝑞𝑗 .

Í

¯ ℎ , we obtain 𝑗 ℎ 𝑗𝑤 𝑗 = 𝐻𝑤

  ∑︁  ℎ 𝑗 𝑞 𝑗 𝑤 𝑗 − 𝐻 𝑞¯ℎ𝑤¯ ℎ = 𝑠 ℎ 𝑗 𝑞 2𝑗 − 𝐻 𝑞¯ℎ2 + 𝜆𝑠.

𝑗

(21)

𝑗

(22)

𝑗

Define the weighted covariance and variance ∑︁ Covℎ (𝑊 , 𝑄) := ℎ 𝑗 (𝑞 𝑗 − 𝑞¯ℎ ) (𝑤 𝑗 − 𝑤¯ ℎ )

(23)

𝑗∈G

=

∑︁

ℎ 𝑗 𝑞 𝑗 𝑤 𝑗 − 𝑞¯ℎ 𝐻 𝑤¯ ℎ − 𝑤¯ ℎ 𝐻 𝑞¯ℎ + 𝑞¯ℎ𝑤¯ ℎ 𝐻

𝑗

=

∑︁

ℎ 𝑗 𝑞 𝑗 𝑤 𝑗 − 𝐻 𝑞¯ℎ𝑤¯ ℎ ,

𝑗

Varℎ (𝑄) :=

∑︁

ℎ 𝑗 (𝑞 𝑗 − 𝑞¯ℎ ) 2

(24)

𝑗∈G

=

∑︁

=

∑︁

ℎ 𝑗 𝑞 2𝑗 − 2𝑞¯ℎ 𝐻 𝑞¯ℎ + 𝑞¯ℎ2 𝐻

𝑗

ℎ 𝑗 𝑞 2𝑗 − 𝐻 𝑞¯ℎ2 .

𝑗

Then Eq. (22) becomes  Covℎ (𝑊 , 𝑄) = 𝑠 Varℎ (𝑄) + 𝜆 , which yields the closed-form solution Í ¯ℎ) Covℎ (𝑊 , 𝑄) 𝑗 ∈ G ℎ 𝑗 (𝑞 𝑗 − 𝑞¯ℎ ) (𝑤 𝑗 − 𝑤 ∗ 𝑠 = = Í . 2 Varℎ (𝑄) + 𝜆 𝑗 ∈ G ℎ 𝑗 (𝑞 𝑗 − 𝑞¯ℎ ) + 𝜆

(25)

(26)

Finally, the global optimum for 𝑧 is obtained by plugging 𝑠 ∗ into Eq. (18): 𝑧 ∗ = 𝑧 ∗ (𝑠 ∗ ) = 𝑠 ∗ · 𝑞¯ℎ − 𝑤¯ ℎ .

(27)

Eq. (26) and Eq. (27) correspond to Eq. (10) and Eq. (11), respectively.

B

MT-bench results

We evaluate instruction-following and multi-turn reasoning quality via MT-Bench. To ensure reproducibility and avoid reliance on proprietary APIs, we employ a publicly available judge model, Llama-3.3-70B-Instruct, under a single-judge protocol. Following the FastChat pipeline, we generate model responses and score them using the standard judge prompts. Table 2 reports MT-Bench scores alongside perplexity (PPL) and zero-shot reasoning averages. The data reveals distinct performance patterns across precision regimes. Moderate Precision (3–4 bit). DASH-Q maintains competitive or superior averages across all models, matching (and in several cases slightly exceeding) FP16 scores throughput the models. Category-wise gains are spread across multiple dimensions of MT-Bench, with notable strengths in Coding (7.10) on Llama-3.1-8B and Reason (7.10) on Mixtral-8x7B. We note that the margins at 4-bit are modest, and MT-Bench variance may contribute to small absolute differences. Ultra Low-bit (2-bit). The transition from 3-bit to 2-bit induces sharp behavioral degradation for several baselines, and is particularly discriminative on the compact dense model Llama-3.1-8B. In this regime, most PTQ baselines concentrate near the minimum MT-Bench range (∼1.0), whereas DASH-Q preserves functional responses with a 2.91 average, retaining non-trivial quality in Writing (5.15) and Extraction (4.40). On Qwen3-14B, DASH-Q retains a 6.12 average at 2bit, while the strongest baseline (OWQ) reaches 2.93. For Mixtral-8x7B, the MoE architecture appears more tolerant to quantization noise, allowing DASH-Q to achieve 6.56 at 2-bit and outperform rotation-based schemes such as QuaRot (4.88). Overall, these patterns are consistent with the view that importance weighting based on diagonal Hessian signals may help stabilize generation behavior under ultra-low precision, although we do not claim a causal attribution from MT-Bench alone. The results also highlight multiple cases where tokenlevel perplexity fails to predict interactive utility. On 4-bit Llama-3.1-8B, OWQ yields the lowest PPL (7.42) but a lower MT-Bench average than DASH-Q (7.47 vs. 7.62), and similar mismatches appear on 4-bit Qwen3-14B (OWQ: 7.65 vs. DASH-Q: 7.78) and Mixtral-8x7B (OWQ: 7.09 vs. DASH-Q: 7.37). More critically, some quantization strategies can induce behavioral degeneration that is not reflected by PPL: at 3-bit on Llama-3.1-8B, GPTQ maintains a reasonable PPL (8.33) and non-trivial zero-shot accuracy (40.43%), yet drops to a near-minimum MT-Bench score (0.98), consistent with malformed or repetitive outputs. Furthermore, DASH-Q achieves higher utility than OWQ in several settings despite using a lower nominal bit-width (e.g., 3.0 vs. 3.32 bits), suggesting that improved importance weighting can be more effective than relying on increased effective precision via outlier handlin

Robust Ultra Low-Bit PTQ via Stable Diagonal Curvature Estimate

EuroMLSys ’26, April 27–30, 2026, Edinburgh, Scotland Uk

Table 2. MT-Bench evaluation with perplexity (PPL) and average zero-shot reasoning accuracy (from Table 1). We report category-wise scores and overall average score using a single-judge protocol with meta-llama/Llama-3.3-70B-Instruct. Model

Llama-3.1-8B

Qwen3-14B

Mixtral-8x7B

Method

W-bit / GS

PPL ↓

MT-Bench

Zero-shot Avg ↑

Coding

Extract

Human

Math

Reason Role STEM Writing

Avg ↑

Baseline

FP16

7.22

67.76

6.70

8.05

8.60

6.55

5.50

8.40

8.55

8.00

7.54

RTN AWQ GPTQ QuIP QuaRot OWQ DASH-Q

4 / 128 4 / 128 4 / 128 4 / 128 4 / 128 4.30 / 128 4 / 128

7.89 7.53 7.48 7.46 7.45 7.42 7.55

66.61 67.07 65.63 66.72 66.70 66.98 66.90

6.35 6.18 5.00 6.40 6.05 6.00 7.10

8.30 8.45 7.80 8.25 8.30 8.50 8.55

8.55 8.60 8.50 8.75 8.75 8.75 8.80

5.50 6.25 4.70 5.95 5.75 5.95 6.10

5.45 5.35 5.60 5.70 6.30 5.60 5.20

8.60 8.40 7.90 8.20 8.30 7.90 8.40

8.65 8.40 8.20 8.65 8.70 8.55 8.45

8.05 8.05 7.90 7.80 8.05 7.80 8.35

7.43 7.46 6.95 7.46 7.53 7.38 7.62

RTN AWQ GPTQ QuIP QuaRot OWQ DASH-Q

3 / 64 3 / 64 3 / 64 3 / 64 3 / 64 3.32 / 64 3 / 64

11.02 8.90 8.33 8.24 8.19 8.10 8.52

59.67 63.15 40.43 64.71 64.45 64.49 64.99

2.55 4.50 0.95 4.15 3.90 5.25 5.45

6.50 7.95 1.00 8.25 8.00 8.05 8.40

7.65 8.30 0.95 8.40 8.40 8.60 8.10

3.10 5.90 0.80 4.65 5.15 5.15 6.55

3.90 4.85 1.00 5.70 6.15 5.75 5.20

7.30 7.75 1.05 7.55 8.10 8.45 8.30

7.10 7.90 1.00 8.00 8.35 8.50 8.20

7.35 7.95 1.00 8.00 7.95 7.90 8.00

5.68 6.89 0.97 6.84 7.00 7.21 7.28

RTN AWQ GPTQ QuIP QuaRot OWQ DASH-Q

2 / 32 2 / 32 2 / 32 2 / 32 2 / 32 2.33 / 32 2 / 32

42446.44 128.34 28.33 21.20 20.78 15.80 16.98

35.23 38.93 35.66 39.90 41.00 42.51 56.52

1.00 1.00 1.00 1.00 0.95 1.00 1.20

1.00 1.00 1.00 1.00 1.00 1.40 4.40

1.00 1.00 1.00 1.00 1.00 1.15 1.75

1.00 1.00 0.80 1.05 1.00 1.00 1.70

1.00 1.00 0.95 1.00 1.00 1.05 2.60

1.00 1.00 1.00 1.05 1.05 1.05 3.50

1.00 1.00 0.90 1.00 1.00 1.05 2.95

1.00 1.00 1.00 1.00 0.90 1.25 5.15

1.00 1.00 0.96 1.01 0.99 1.12 2.91

Baseline

FP16

8.65

69.40

5.20

8.80

8.45

6.40

6.55

8.55

8.25

8.30

7.56

RTN AWQ GPTQ QuIP QuaRot OWQ DASH-Q

4 / 128 4 / 128 4 / 128 4 / 128 4 / 128 4.24 / 128 4 / 128

9.43 8.87 8.87 8.78 8.75 8.74 8.91

68.58 68.95 68.93 68.68 68.71 68.92 69.06

4.75 5.00 5.20 5.30 4.90 4.98 5.75

7.75 7.95 8.05 8.53 8.18 8.35 8.60

8.25 8.00 8.45 8.50 8.65 8.60 8.10

7.00 6.90 7.75 6.80 7.50 7.45 7.80

6.80 5.90 6.70 6.55 6.40 6.50 7.10

8.60 8.60 8.45 8.70 8.65 8.60 8.70

7.95 7.55 7.80 7.85 8.00 7.95 7.70

8.20 8.45 8.45 8.30 8.45 8.75 8.45

7.41 7.29 7.61 7.57 7.59 7.65 7.78

RTN AWQ GPTQ QuIP QuaRot OWQ DASH-Q

3 / 64 3 / 64 3 / 64 3 / 64 3 / 64 3.25 / 64 3 / 64

11.61 9.51 9.38 9.09 9.03 9.16 9.37

63.20 67.46 66.58 66.39 67.57 67.67 68.31

3.15 5.00 4.90 4.95 5.15 4.10 4.50

5.15 7.50 8.40 7.65 8.50 8.15 8.30

6.90 8.40 7.55 8.30 8.15 8.45 8.30

5.55 7.00 6.30 7.05 7.05 6.50 7.30

3.90 5.65 5.20 6.40 6.55 6.50 6.00

7.05 7.50 7.80 8.35 8.55 8.60 8.75

5.80 7.75 7.20 7.70 6.55 7.65 7.90

7.60 7.95 7.70 8.35 8.20 8.45 8.40

5.64 7.09 6.88 7.34 7.34 7.30 7.43

RTN AWQ GPTQ QuIP QuaRot OWQ DASH-Q

2 / 32 2 / 32 2 / 32 2 / 32 2 / 32 2.26 / 32 2 / 32

299.83 15.12 13.65 11.81 11.79 11.85 11.41

37.61 53.74 51.76 53.81 53.94 54.25 64.42

0.70 1.50 1.00 1.40 1.20 1.30 3.55

0.90 2.25 1.25 3.70 3.45 3.20 6.10

1.20 3.15 1.05 3.90 2.75 3.45 6.55

0.75 1.85 1.05 2.05 1.25 2.95 6.10

0.80 2.20 1.50 2.25 1.45 2.00 3.60

0.75 3.40 1.25 2.20 2.35 3.40 7.75

0.90 2.90 1.85 3.05 3.05 3.10 7.30

1.30 4.15 1.35 3.65 3.30 4.45 8.00

0.91 2.68 1.29 2.78 2.35 2.98 6.12

Baseline

FP16

4.14

72.95

5.80

7.75

8.35

5.40

6.35

7.95

8.00

8.10

7.21

RTN AWQ GPTQ QuIP QuaRot OWQ DASH-Q

4 / 128 4 / 128 4 / 128 4 / 128 4 / 128 4.28 / 128 4 / 128

4.45 4.29 4.34 4.23 4.23 4.23 4.27

70.61 71.50 71.13 71.94 72.39 72.46 72.40

5.85 5.50 6.10 6.25 6.55 5.85 6.55

6.15 7.15 6.95 7.85 6.95 7.95 7.80

8.50 8.35 8.60 8.55 8.55 7.95 8.35

4.60 4.25 5.85 4.35 4.50 5.55 4.75

5.10 6.40 5.10 5.45 7.15 5.15 7.10

7.85 8.45 7.90 8.30 8.30 8.25 8.20

8.15 8.00 7.95 8.10 8.35 7.85 7.95

7.95 8.20 8.20 8.35 8.25 8.15 8.25

6.77 7.04 7.08 7.15 7.33 7.09 7.37

RTN AWQ GPTQ QuIP QuaRot OWQ DASH-Q

3 / 64 3 / 64 3 / 64 3 / 64 3 / 64 3.30 / 64 3 / 64

5.43 4.68 4.88 4.48 4.48 4.47 4.53

68.99 70.70 67.24 71.03 70.06 71.26 71.35

5.00 4.85 4.95 5.30 5.80 4.90 6.05

5.95 7.70 7.30 7.45 8.15 7.75 7.45

7.95 8.35 8.35 8.35 8.15 8.60 8.30

3.35 5.90 5.00 6.00 4.95 5.80 5.40

5.85 5.35 6.40 5.70 5.90 6.30 6.60

7.95 8.00 7.90 8.30 8.05 8.15 8.00

7.75 8.30 7.95 8.00 8.00 8.00 7.85

7.90 7.85 7.95 8.00 8.25 8.30 8.30

6.46 7.04 6.98 7.14 7.16 7.23 7.24

RTN AWQ GPTQ QuIP QuaRot OWQ DASH-Q

2 / 32 2 / 32 2 / 32 2 / 32 2 / 32 2.31 / 32 2 / 32

1103.32 7.53 12.74 5.94 5.94 5.80 5.51

34.58 56.41 37.53 65.61 64.63 61.69 69.72

1.00 1.50 0.95 1.85 1.85 1.90 5.20

1.00 2.75 1.45 5.15 4.80 5.40 7.05

1.00 5.05 1.30 6.95 7.65 7.65 7.85

1.00 1.65 1.10 1.80 1.95 2.90 4.50

1.00 3.25 1.45 3.80 3.20 3.50 5.20

1.00 4.50 1.15 6.40 6.75 7.45 7.85

1.00 5.40 1.15 6.55 6.45 6.55 7.45

1.00 4.95 1.10 6.40 6.30 7.50 7.35

1.00 3.63 1.21 4.86 4.87 5.36 6.56

EuroMLSys ’26, April 27–30, 2026, Edinburgh, Scotland Uk

Overall, the MT-Bench ablation suggests that DASH-Q provides a comparatively stable behavior-preserving quantization strategy across model families, with the largest separations emerging under 2-bit compression where behavioral degeneration is most pronounced. These observations are

Jaemin Kim, Sungkyun Kim, Junyeol Lee, and Jiwon Seo

broadly consistent with the hypothesis that emphasizing diagonal Hessian importance may reduce sensitivity to limited calibration data and mitigate unstable feature correlation fitting, thereby improving instruction-following utility under aggressive quantization.

Record · ID 14047 · SHA-256 e513b4f2796942d9
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.