Conceptio › Archive › arXiv CS
arXiv CSopen access

Enhancing SignSGD: Small-Batch Convergence Analysis and a Hybrid Switching Strategy

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Enhancing SignSGD: Small-Batch Convergence Analysis and a Hybrid Switching Strategy Haoran Chen, Wentao Wang

arXiv:2604.25550v1 [cs.LG] 28 Apr 2026

École polytechnique Palaiseau, France [email protected], [email protected]

Abstract—SignSGD compresses each stochastic gradient coordinate to a single bit, offering substantial memory and communication savings, but its 1-bit quantization removes magnitude information and is known to leave a generalization gap relative to well-tuned SGD. We revisit SignSGD from a 1-bit quantization and dithering perspective and contribute three improvements. First, we derive a small-batch convergence rate for SignSGD under unimodal symmetric gradient noise using a signal-to-noise weighted stationarity measure, removing the large-batch assumption of prior analyses. Second, we inject annealed Gaussian noise before the sign operator, which acts as a classical dithering mechanism and probabilistically restores magnitude information lost to hard thresholding. Third, we adapt the SWATS strategy to sign-based updates with a projection-based learning-rate calibration that smoothly transitions from SignSGD to SGD. Single-worker experiments on ResNet-18 isolate optimizer effects from communication aspects: pre-sign dithering surpasses Adam on CIFAR-100, and the calibrated switch reaches 92.18% test accuracy on CIFAR-10, outperforming both pure SGD (91.38%) and pure SignSGD with momentum (90.82%). Index Terms—Stochastic optimization, 1-bit quantization, dithering, sign-SGD, communication-efficient training, deep learning.

I. I NTRODUCTION Training deep neural networks at scale stresses both compute and communication budgets, motivating compressed gradient updates for memory-constrained and distributed settings. SignSGD [1] replaces each stochastic gradient coordinate g̃i by its sign, yielding a 1-bit update with strong communicationreduction properties when combined with majority-vote aggregation across workers. From a signal-processing viewpoint, sign(g̃i ) is a coarse 1-bit quantizer applied to a noisy signal, and many of the algorithm’s behaviors—robustness to sparse gradient noise, biased fixed points, and sensitivity to small-magnitude coordinates—are natural consequences of this quantization. Despite its appeal, sign-based optimizers exhibit a persistent generalization gap relative to well-tuned SGD on vision benchmarks [1], mirroring the gap observed for adaptive methods more broadly [8], and now also visible in modern variants such as Lion [4], which is in essence a sign-of-momentum optimizer. The gap is commonly attributed to two factors: (i) hard thresholding discards magnitude information that would otherwise modulate the step direction near small or noisy coordinates, and (ii) the uniform ±δ step can prevent convergence to flat minima associated with good generalization [2].

Contributions. Through a 1-bit-quantization-and-dithering lens, we make three contributions: • Small-batch convergence analysis (Sec. III): we derive a signal-to-noise (SNR) weighted stationarity rate for SignSGD with constant batch size under unimodal symmetric noise, complementing the large-batch rate of [1]. • Pre-sign dithering (Sec. IV): we add annealed Gaussian noise before the sign operator, the optimizer-side analogue of classical subtractive dithering in 1-bit A/D conversion [10], [11], and show it consistently improves test accuracy on CIFAR-100, surpassing Adam. • Calibrated switch to SGD (Sec. V): we adapt SWATS [8] to sign-based updates with a projection-based learningrate calibration that matches the effective step size of SignSGD with momentum at the switching iterate, avoiding the magnitude mismatch of a naive switch. Scope. All experiments use a single worker and are intended to isolate the optimization-side effects of 1-bit quantization, dithering, and switching from communication-side considerations. The dithering and switching mechanisms are compatible with majority-vote aggregation, but a multi-worker evaluation is left to future work. II. BACKGROUND AND R ELATED W ORK SignSGD and variants. Bernstein et al. [1] established √ that SignSGD attains an O(1/ N ) rate on the average ℓ1 gradient norm under coordinate-wise smoothness and bounded variance, and that majority-vote aggregation across M work√ ers reduces the variance term by M . Subsequent work showed that vanilla SignSGD can fail to converge under non-symmetric noise and proposed error-feedback variants [2]; communication-compressed Adam variants such as 1-bit Adam [3] extend the same idea to adaptive optimizers. Lion [4], recently popular for large-scale pretraining, is in essence a sign-of-momentum optimizer and inherits the geometry studied here. An early predecessor is 1-bit SGD applied to speech DNN training [7]. Quantized and compressed SGD. Beyond 1-bit, QSGD [5], TernGrad [6], and Top-K sparsification trade compression rate against unbiasedness. SignSGD sits at the extreme compression end and trades unbiasedness for simplicity. Gradient noise and dithering. Adding Gaussian noise to gradients improves training of very deep networks [9]; in

signal processing, dithering before a coarse quantizer is a classical mechanism to decorrelate quantization error from the input and to encode sub-quantum-step information probabilistically [10], [11]. Our pre-sign noise injection makes this connection explicit on the optimizer side. Adaptive-to-SGD switching. SWATS [8] starts with Adam and switches to SGD with a learning rate determined by a projection condition. We transplant this idea to sign-based updates with a momentum-aware projection. III. S MALL -BATCH C ONVERGENCE U NDER U NIMODAL S YMMETRIC N OISE We adopt the standard assumptions of [1]: f is lower bounded by f ⋆ , has a coordinate-wise Lipschitz constant ⃗ with Li the Lipschitz constant of ∂i f , and the unvector L biased single-sample stochastic gradient oracle has coordinate variance bounded by σi2 . We additionally assume that the percoordinate noise is unimodal and symmetric, so that the signfailure-probability bound of [1]—a consequence of Gauss’s inequality [12]—applies coordinate-wise. The mini-batch estimator g̃k at iteration k is the average of n independent samples of the oracle, so that Var(g̃k,i | xk ) ≤ σi2 /n.p P P Let L̃1 := i Li , σ̃1 := i σi , sk,i := Var(g̃k,i | xk ), and Sk,i := |gk,i |/sk,i denote the per-coordinate signalto-noise ratio at iteration k. We define the SNR-weighted stationarity measure d X

d X

  |gk,i |2 min |gk,i |, . sk,i i=1 i=1 (1) Φk smoothly interpolates between ∥gk ∥1 on high-SNR coordinates and a quadratic-in-|gk,i | measure on low-SNR coordinates, reflecting the diminishing usefulness of the sign bit when noise dominates the signal. This is a more refined stationarity criterion than ∥gk ∥1 : it weights each coordinate by its informativeness rather than its raw magnitude. Φk :=

|gk,i | min(1, Sk,i ) =

Theorem 1 (Small-batch SNR-weighted rate of SignSGD). Run SignSGD for K iterations with constant mini-batch size p nk ≡ n and constant stepsize δk ≡ 1/ L̃1 K. Then " K−1 # p  1 X 3 L̃1 E Φk ≤ √ f0 − f ⋆ + 21 , (2) K K k=0 " K−1 # p  1 X 3 L̃1 σ̃1 E ∥gk ∥1 ≤ √ f0 − f ⋆ + 21 + √ . (3) K n K k=0

Proof. By coordinate-wise smoothness with y = xk+1 , x = xk , δ2 fk+1 − fk ≤ −δk gk⊤ sign(g̃k ) + k L̃1 . (4) 2 Let pk,i := P[sign(g̃k,i ) ̸= sign(gk,i ) | xk ], so that E[gk,i sign(g̃k,i ) | xk ] = |gk,i |(1 − 2pk,i ).

(5)

The Gauss-inequality bound for unimodal symmetric noise [1], [12] p gives, after a slight relaxation of the split point at S = 2/3,  p 2   Sk,i > 2/3,  2 , 9Sk,i pk,i ≤ (6) 1 Sk,i    − √ , otherwise, 2 2 3 which yields 1 − 2pk,i ≥ 13 min(1, Sk,i ) in both regimes. Summing over coordinates and combining with (5) gives 1 Φk . (7) 3 Substituting (7) into (4), taking total expectation and telescoping over k = 0, . . . , K − 1, E[gk⊤ sign(g̃k ) | xk ] ≥

f0 − f ⋆ ≥

1X L̃1 X 2 δk E[Φk ] − δk . 3 2 k

(8)

k

p With δk ≡ 1/ L̃1 K, rearranging yields (2). For (3), the elementary inequality |a| ≤ min(|a|, a2 /s)+s valid for a ∈√R, √ s > 0, applied with sk,i ≤ σi / n gives ∥gk ∥1 ≤ Φk +σ̃1 / n. Averaging over k and combining with (2) completes the proof. Discussion. Compared to Theorem 1 of [1], which requires nk ≡ K (a large-batch regime where increasing the optimization horizon implicitly raises the batch size), Theorem 1 holds √ for any constant n, at the price of a non-vanishing σ̃1 / n floor in (3). This floor is an honest reflection of the cost of 1-bit quantization at small batch sizes: when the percoordinate SNR is low, the sign bit carries little information about the gradient direction, and no amount of iteration √ can drive ∥gk ∥1 below the noise level encoded in σ̃1 / n. The SNR-weighted measure Φk in (2), by contrast, vanishes at √ the standard O(1/ K) rate, providing a clean stationarity guarantee in the small-batch regime. IV. P RE -Q UANTIZATION D ITHERING Mechanism. The sign operator is a coarse 1-bit quantizer; its output ignores all magnitude information in its argument. Subtractive and non-subtractive dithering—adding controlled noise to a signal before quantization—is a textbook technique to decorrelate quantization error from the input and to encode sub-quantum-step information probabilistically [10], [11]. We apply this idea to SignSGD with momentum (SignSGD-M), summarized in Algorithm 1. We anneal the dither variance following [9]: σk2 = α(1 + k)−γ ,

γ = 0.55.

(9)

For comparison we also evaluate post-sign injection, xk+1 = xk − δ(sign(mk+1 ) + ξk ) ,

(10)

which perturbs the already-normalized step and is closer in spirit to exploration noise on a constant-magnitude update. Effect on the sign-failure probability. Consider a single coordinate with momentum signal mi > 0 and pre-sign Gaussian noise ξi ∼ N (0, σk2 ). The probability of correct sign

Algorithm 1 Dithered SignSGD-M (pre-sign noise) Require: learning rate δ, momentum β ∈ (0, 1), dither schedule σk2 = α(1 + k)−γ 1: Initialize x0 , m0 ← 0 2: for k = 0, 1, . . . , K − 1 do 3: Sample mini-batch and compute g̃k 4: mk+1 ← βmk + (1 − β)g̃k 5: Sample dither ξk ∼ N (0, σk2 I) 6: xk+1 ← xk − δ sign(mk+1 + ξk ) 7: end for

Algorithm 2 Hybrid SignSGD-M → SGD with projection Require: SignSGD-M lr δ, momentum β ∈ (0, 1), EMA decay η, switch epoch Tswitch 1: Initialize x0 , m0 ← 0, λ̄ ← 0 2: for k = 0, 1, . . . , K − 1 do 3: Sample mini-batch and compute g̃k 4: if epoch < Tswitch then 5: mk+1 ← βmk + (1 − β)g̃k 6: λk ← δ |⟨sign(mk+1 ), g̃k ⟩|/(∥g̃k ∥22 + ϵ) 7: λ̄ ← η λ̄ + (1 − η)λk 8: xk+1 ← xk − δ sign(mk+1 ) 9: else 10: xk+1 ← xk − λ̄ g̃k ▷ SGD with calibrated lr 11: end if 12: end for

update. V. C ALIBRATED S WITCHING TO SGD

Fig. 1. Test accuracy on CIFAR-100 (ResNet-18) for SignSGD-M with preand post-sign Gaussian dithering, against SGD and Adam baselines. Pre-sign injection with α = 0.1 is the best sign-based variant, surpassing Adam and approaching SGD; post-sign injection is unstable. Stepwise jumps reflect a learning-rate schedule shared across all methods.

is Φ(mi /σk ), where Φ is the standard normal CDF. As σk → 0 this recovers the deterministic sign of mi ; for moderate σk , coordinates with |mi | ≲ σk produce dithered (probabilistic) p signs whose expectation is 2Φ(mi /σk ) − 1 ≈ mi 2/π/σk to first order. This is, in spirit, the optimizer-side analogue of classical dithering theory [10], [11]: the average update on small-magnitude coordinates becomes proportional to mi rather than its sign, partially restoring the magnitude information lost to quantization. The annealing schedule (9) ensures that this effect is strongest early in training—when the gradient signal is informative but coarsely quantized—and decays so as not to interfere with late-stage convergence. Empirical results. We evaluate on CIFAR-100 with ResNet-18 [14] (Fig. 1). Pre-sign dithering with α = 0.1 reaches roughly 76% test accuracy, the best among all signbased variants, surpassing Adam [13] (∼67%) and approaching the SGD baseline (∼77%). Larger α = 0.5 also helps but plateaus lower, while α = 0.01 provides marginal benefit, indicating that the dither variance must be matched to the typical magnitude of the momentum signal. Post-sign injection, which perturbs the already-normalized step, fails to provide a meaningful gain: it shows at most a marginal improvement over clean SignSGD-M (within 1%) and degrades for larger α, consistent with random-walk behavior on a fixed-magnitude

Motivation. A naive hand-off from SignSGD-M to SGD at a fixed epoch is unstable: the normalized ±δ updates of SignSGD-M and the gradient-proportional updates of SGD are on different magnitude scales, and reusing either learning rate produces stagnation or divergence. Following SWATS [8], we calibrate the SGD learning rate at the switching iterate so that the SGD update matches the projection of the sign-based update onto the gradient direction. Projection rule. Let δ be the SignSGD-M learning rate, mk+1 = βmk + (1 − β)g̃k the just-updated momentum at iteration k, and g̃k the current stochastic mini-batch gradient. The SignSGD-M step actually applied at iteration k is −δ sign(mk+1 ). We seek a scalar λk ≥ 0 such that the SGD update −λk g̃k equals the projection of −δ sign(mk+1 ) onto g̃k : λk ∥g̃k ∥22 = δ ⟨sign(mk+1 ), g̃k ⟩. (11) Enforcing a non-negative magnitude for stability, λk = δ

|⟨sign(mk+1 ), g̃k ⟩| , ∥g̃k ∥22 + ϵ

(12)

with small ϵ > 0 for numerical stability. During the SignSGDM phase we maintain an exponential moving average (EMA) of λk to reduce stochastic variance, and at a fixed switching epoch Tswitch we transition to SGD with the EMA stepsize. We use a fixed Tswitch to isolate the effect of the projection mechanism from the automatic-trigger heuristic of [8]. The full procedure is given in Algorithm 2. Empirical results. On CIFAR-10 (Fig. 2), the hybrid method enjoys the rapid early progress of SignSGD-M— crossing 80% test accuracy several epochs earlier than pure SGD—and after the switch fine-tunes smoothly to 92.18%, outperforming both pure SGD (91.38%) and pure SignSGDM (90.82%). On CIFAR-100 (Fig. 3), the hybrid method substantially reduces the stagnation gap left by SignSGD-M (∼69% → ∼72%), but does not close the gap to well-tuned SGD (∼77%); we attribute this to the larger generalization

TABLE I F INAL TEST ACCURACY (%) ON R ES N ET-18, SINGLE - WORKER TRAINING . D ITHERING ROWS (S EC . IV) AND HYBRID ROW (S EC . V) COME FROM INDEPENDENT RUNS WITH MATCHED ARCHITECTURES BUT SEPARATELY TUNED LEARNING - RATE SCHEDULES ; CIFAR-100 SIGN - BASED NUMBERS ARE READ FROM F IG . 1 AND ROUNDED TO THE NEAREST INTEGER . B EST PER COLUMN IN BOLD .

Method Fig. 2. ResNet-18 on CIFAR-10. The hybrid method (green) follows the SignSGD-M trajectory through epoch 25 (red dashed line), then transitions to SGD with the projection-calibrated learning rate from (12). Final test accuracy: hybrid 92.18%, SGD 91.38%, SignSGD-M 90.82%.

CIFAR-10

CIFAR-100

91.38 – 90.82 – – – 92.18

77 67 65 66 73 76 72

SGD (well-tuned) Adam SignSGD-M (clean) SignSGD-M + post-noise (α=0.1) SignSGD-M + pre-noise (α=0.5) SignSGD-M + pre-noise (α=0.1) Hybrid (SignSGD-M → SGD)

100 the switch removes SignSGD-M’s stagnation but does not match SGD, suggesting that the projection rule alone cannot recover the implicit regularization that SGD enjoys throughout training. Combining dithering during the SignSGD-M phase with the calibrated switch is a natural next step that we did not exhaustively explore here. Fig. 3. ResNet-18 on CIFAR-100. The hybrid method substantially mitigates the late-stage stagnation of SignSGD-M (∼69%), reaching ∼72%, but does not match well-tuned SGD (∼77%) on this benchmark. Switch at epoch 60.

sensitivity of CIFAR-100 to the precise late-training trajectory and to the heuristic choice of Tswitch . The CIFAR-100 result indicates that calibrated switching is best understood as a remedy for sign-induced stagnation rather than a uniform improvement over SGD. Practical considerations. The projection rule (12) requires only inner products and norms of quantities that SignSGDM already computes, so the EMA tracking adds negligible cost. The dithering scheme of Sec. IV can be combined with the switching scheme of Sec. V: dithering operates inside the SignSGD-M phase and does not modify the projection rule. We did not observe instability from this combination in preliminary experiments, although a systematic study is beyond the scope of this paper. VI. R ESULTS S UMMARY AND D ISCUSSION Table I consolidates the final test accuracies on ResNet-18 across all settings considered. Three observations stand out. First, plain SignSGD-M trails well-tuned SGD by roughly 0.6% on CIFAR-10 and by approximately 12% on CIFAR100, confirming that the generalization cost of 1-bit quantization grows with task difficulty. Second, pre-sign dithering is the only one of our two improvements that closes most of the CIFAR-100 gap to SGD without changing the optimizer family, while post-sign injection provides at most a marginal benefit over the clean baseline and destabilizes training at higher noise levels—a clear empirical separation between dithering before and after the quantizer. Third, the calibrated switch is most effective on CIFAR-10, where the late-training landscape is benign enough that a short SignSGD-M warmup followed by SGD beats either pure method; on CIFAR-

VII. C ONCLUSION We revisited SignSGD through a 1-bit-quantization-anddithering lens and contributed (i) a small-batch SNR-weighted convergence rate under unimodal symmetric noise, removing the large-batch assumption of prior analyses; (ii) annealed pre-sign Gaussian dithering, the optimizer-side analogue of classical dither, which restores magnitude information lost to hard thresholding; and (iii) a SWATS-style calibrated switch from sign-based updates to SGD via a momentum-aware projection rule. Single-worker experiments on ResNet-18 isolate optimizer effects from communication and show that presign dithering surpasses Adam on CIFAR-100, while the calibrated switch outperforms both pure SGD and SignSGDM on CIFAR-10 and mitigates SignSGD-M stagnation on CIFAR-100 without fully closing the gap to SGD. Multiworker experiments under majority-vote aggregation, errorfeedback combinations, and an automatic switching trigger calibrated by the EMA stability of λk are natural next steps. R EFERENCES [1] J. Bernstein, Y.-X. Wang, K. Azizzadenesheli, and A. Anandkumar, “signSGD: Compressed optimisation for non-convex problems,” in Proc. Int. Conf. Machine Learning (ICML), 2018. arXiv:1802.04434. [2] S. P. Karimireddy, Q. Rebjock, S. U. Stich, and M. Jaggi, “Error feedback fixes SignSGD and other gradient compression schemes,” in Proc. Int. Conf. Machine Learning (ICML), 2019. arXiv:1901.09847. [3] H. Tang, S. Gan, A. A. Awan, S. Rajbhandari, C. Li, X. Lian, J. Liu, C. Zhang, and Y. He, “1-bit Adam: Communication efficient large-scale training with Adam’s convergence speed,” in Proc. Int. Conf. Machine Learning (ICML), 2021. arXiv:2102.02888. [4] X. Chen, C. Liang, D. Huang, E. Real, K. Wang, Y. Liu, H. Pham, X. Dong, T. Luong, C.-J. Hsieh, Y. Lu, and Q. V. Le, “Symbolic discovery of optimization algorithms,” in Adv. Neural Inform. Process. Syst. (NeurIPS), 2023. arXiv:2302.06675. [5] D. Alistarh, D. Grubic, J. Li, R. Tomioka, and M. Vojnovic, “QSGD: Communication-efficient SGD via gradient quantization and encoding,” in Adv. Neural Inform. Process. Syst. (NeurIPS), 2017. arXiv:1610.02132.

[6] W. Wen, C. Xu, F. Yan, C. Wu, Y. Wang, Y. Chen, and H. Li, “TernGrad: Ternary gradients to reduce communication in distributed deep learning,” in Adv. Neural Inform. Process. Syst. (NeurIPS), 2017. arXiv:1705.07878. [7] F. Seide, H. Fu, J. Droppo, G. Li, and D. Yu, “1-bit stochastic gradient descent and application to data-parallel distributed training of speech DNNs,” in Proc. Interspeech, 2014, pp. 1058–1062. [8] N. S. Keskar and R. Socher, “Improving generalization performance by switching from Adam to SGD,” arXiv:1712.07628, 2017. [9] A. Neelakantan, L. Vilnis, Q. V. Le, I. Sutskever, L. Kaiser, K. Kurach, and J. Martens, “Adding gradient noise improves learning for very deep networks,” arXiv:1511.06807, 2015. [10] L. Schuchman, “Dither signals and their effect on quantization noise,” IEEE Trans. Commun. Technol., vol. 12, no. 4, pp. 162–165, Dec. 1964. [11] R. M. Gray and T. G. Stockham, “Dithered quantizers,” IEEE Trans. Inf. Theory, vol. 39, no. 3, pp. 805–812, May 1993. [12] F. Pukelsheim, “The three sigma rule,” The American Statistician, vol. 48, no. 2, pp. 88–91, May 1994. [13] D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Proc. Int. Conf. Learning Representations (ICLR), 2015. [14] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), 2016, pp. 770–778.

Record · ID 141492 · SHA-256 7d286b1cdd084a2f
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.