Evolution of Optimization Methods: Algorithms, Scenarios, and Evaluations Tong Zhang1 , Jiangning Zhang1† , Zhucun Xue1 , Juntao Jiang1 , Yicheng Xu1 , Chengming Xu2 , Teng Hu3 , Xingyu Xie4 , Xiaobin Hu5† , Yabiao Wang1 , Yong Liu1† , Shuicheng Yan4 Zhejiang University, APRIL Lab, 2 Fudan University, 3 Shanghai Jiaotong University, 4 National University of Singapore
arXiv:2604.12968v1 [cs.LG] 14 Apr 2026
1
Balancing convergence speed, generalization capability, and computational efficiency remains a core challenge in deep learning optimization. First-order gradient descent methods, epitomized by stochastic gradient descent (SGD) and Adam, serve as the cornerstone of modern training pipelines. However, large-scale model training, stringent differential privacy requirements, and distributed learning paradigms expose critical limitations in these conventional approaches regarding privacy protection and memory efficiency. To mitigate these bottlenecks, researchers explore second-order optimization techniques to surpass first-order performance ceilings, while zeroth-order methods reemerge to alleviate memory constraints inherent to large-scale training. Despite this proliferation of methodologies, the field lacks a cohesive framework that unifies underlying principles and delineates application scenarios for these disparate approaches. In this work, we retrospectively analyze the evolutionary trajectory of deep learning optimization algorithms and present a comprehensive empirical evaluation of mainstream optimizers across diverse model architectures and training scenarios. We distill key emerging trends and fundamental design trade-offs, pinpointing promising directions for future research. By synthesizing theoretical insights with extensive empirical evidence, we provide actionable guidance for designing next-generation highly efficient, robust, and trustworthy optimization methods. Date: April 15, 2026 Project: https://github.com/APRIL-AIGC/Awesome-Optimizer
Contents 1 Introduction
1
2 Background 2.1 Optimization Problem Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2.2 Deep Learning Optimization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2.3 Evaluation Metrics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2.4 Optimization for Large Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2.5 History and Roadmap . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
3 3 3 4 4 4
3 Optimization Algorithms 3.1 Unified Mathematical Perspective . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.2 First-Order Algorithms . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.2.1 Accelerating Convergence Rate . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.2.2 Adaptive Step-Size Control . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.2.3 Variance Adaptation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.2.4 Enhancing Training Stability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.2.5 Learning Rate Scheduling . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.2.6 Enhancing Generalization Ability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.2.7 Hybrid Methods . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
5 6 7 8 10 14 15 16 18 19
3.2.8 Towards LLMs Training . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . Second-Order Algorithms . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.3.1 Hessian Approximation & Estimation . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.3.2 Fisher Information Matrix Applications . . . . . . . . . . . . . . . . . . . . . . . . . . 3.3.3 Quasi-Newton Methods . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.3.4 Second-Order Moment Fusion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . Zeroth-Order Algorithms . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.4.1 Adaptive Methods . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.4.2 Perturbation Optimization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.4.3 Zeroth-First Order Hybrid . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.4.4 Memory-Efficient Methods . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.4.5 Variance Reduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.4.6 Distributed Zero-Order Optimization . . . . . . . . . . . . . . . . . . . . . . . . . . . . Distributed Optimization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.5.1 Gradient Compression & Quantization . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.5.2 Local Update Strategies . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.5.3 Decentralized Communication . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.5.4 Federated Learning Optimization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.5.5 Communication Scheduling&Threshold . . . . . . . . . . . . . . . . . . . . . . . . . . 3.5.6 Distributed Hybrid Optimization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . Privacy-Preserving Optimization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.6.1 Differential Privacy Optimization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.6.2 Gradient Noise Injection . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.6.3 Privacy-Utility Tradeoff . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.6.4 Federated Privacy Enhancement . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.6.5 Privacy-Aware Gradient Clipping . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . Memory-Efficient Optimization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.7.1 Low-Memory Optimizer Design . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.7.2 Low-Rank Algorithms . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.7.3 Optimizer State Compression . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.7.4 Low-Rank Gradient Storage . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.7.5 Stateless Optimization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.7.6 Memory-Efficient for Large Models . . . . . . . . . . . . . . . . . . . . . . . . . . . . . Tailored Optimization Approaches . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.8.1 Auto-Designed Optimizers . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3.8.2 Robust Optimization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
21 22 23 24 25 26 27 27 29 30 30 32 32 33 34 35 36 37 39 39 40 41 42 42 42 43 43 44 44 45 45 45 46 46 46 47
4 Experiments 4.1 Datasets . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4.2 Metrics and Settings . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4.3 Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4.3.1 Learning Rate Sensitivity on ViT-S . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4.3.2 Cross-Architecture Generalization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4.3.3 Long-term Training Scalability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4.3.4 Correlation of Optimizers . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4.3.5 Comprehensive Optimizer Evaluation. . . . . . . . . . . . . . . . . . . . . . . . . . . . 4.3.6 Computational cost disclosure. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
47 48 48 49 49 51 52 53 54 54
5 Future Prospect
55
6 Conclusion
57
3.3
3.4
3.5
3.6
3.7
3.8
1
Introduction
work to evaluate 23 optimizers across continuous vision tasks (ResNet He et al. (2016), ViT Dosovitskiy (2020)) and discrete causal language modeling (Llama Llama (2023)). Our extensive benchmarking yields several critical insights: First, foundational FO algorithms like SGD Robbins and Monro (1951) suffer catastrophic training collapse on the highly anisotropic loss landscapes of LLMs (Fig. 9), dictating the absolute necessity of adaptive scaling. Second, while standard adaptive methods (e.g., Adam Kingma (2015)) exhibit severe sensitivity to large learning rates, emerging methods utilizing matrix orthogonalization (e.g., Muon Jordan et al. (2024)) or gradient correction (e.g., MARS Yuan et al. (2025)) demonstrate superior hyperparameter robustness and crossarchitecture generalization (Sec. 4.3.2).
Deep learning has historically been driven by foundational optimization algorithms, with early methods like stochastic gradient descent (SGD) Robbins and Monro (1951) establishing the basic training framework, and the Adam Kingma (2015) family later becoming the standard for various tasks. However, the unprecedented scale of modern foundation models has fundamentally shifted this optimization paradigm. During the training of large language models (LLMs), these traditional workhorses have exposed severe limitations. The primary bottleneck is no longer merely the theoretical convergence rate, but the physical and systemic feasibility of training: the memory wall imposed by backpropagation, the communication wall in decentralized networks, and the privacy wall for sensitive data. This survey argues that the recent diversification of optimization algorithms represents a unified, albeit fragmented, response to this feasibility crisis. We demonstrate that traditional mathematical primitives, including first-order (FO) workhorses, second-order (SO) curvature-aware methods, and the recently revived memory-efficient zeroth-order (ZO) techniques, are being fundamentally re-architected. By evolving into scenario-oriented paradigms, including distributed systems and rigorously differentially private frameworks, modern optimizers are moving beyond pure algorithmic design to become highly constrained, systems-aware engineering solutions. Optimization methods can generally be divided into four mainstream research fields: 1) FO Robbins and Monro (1951); Kingma (2015) methods rely on first-order gradients (and their derived statistics) to achieve superior convergence with low computational overhead, strictly avoiding the explicit approximation of true second-order curvature. 2) SO methods Moré and Sorensen (1982); Byrd et al. (1995); Amari (1998) explicitly construct and incorporate true curvature information (e.g., the Hessian or true FIM) to precondition the update direction, aiming to overcome the fundamental performance limits of FO algorithms. 3) ZO methods Maryak and Chin (2001) aim to approximate gradient directions via function evaluations to avoid the massive memory burden of backpropagation. And one overarching application framework: 4) Scenario-oriented paradigms, which focus on re-engineering these base primitives for distributed communication, differential privacy, and other specialized use cases. The development of foundational techniques has laid the groundwork for these specialized paradigms, but the field now requires a more unified perspective to advance.
Existing surveys Fucheng et al. (2025); Mustapha et al. (2020) suffer from three critical limitations that hinder the development of the field. First, most reviews are limited in scope, concentrating nearly entirely on conventional FO and SO approaches while neglecting ZO optimization and scenario-oriented frameworks. These two fast-developing paradigms are now essential for memory-efficient LLM training, distributed learning, and privacy-preserving applications. Second, existing works lack a unified, mathematically rigorous taxonomy, resulting in fragmented, inconsistent categorization of algorithms that obscures the intrinsic connections and evolutionary logic between different methods. Third, none of the existing surveys provide a standardized, large-scale empirical benchmark for fair cross-architecture evaluation of modern optimizers, leading to conflicting, non-reproducible conclusions in the literature and leaving researchers and engineers without reliable guidance for optimizer selection in practical scenarios. To address these critical gaps, this survey provides a systematic, panoramic review of deep learning optimization methods, with a unified theoretical framework, standardized empirical benchmark, and actionable insights for both theoretical research and engineering practice. Contributions. In this survey, we make three key contributions to the field of deep learning optimization: 1) Unified Taxonomy: We unify disparate conceptual definitions and establish a rigorous mathematical taxonomy (Sec. 2), tracing the evolution of fundamental optimization primitives from FO to SO and ZO methods. This taxonomy clarifies the intrinsic connections and evolutionary logic between different optimization approaches, providing a coherent framework for understanding the field’s development. 2) Scenario-Oriented Analysis: We demonstrate how these foundational algorithms are being
Also, we introduce a standardized evaluation frame-
1
SGDO
GCSAM NAdam
LightSAM SAMPa AdamP
Adan
SAM
GAM
ESAM
AdaBelief
Lion
MADGRAD LOMO
LAMB
AdamW
Muon
RAdam
ScheduleFree
BADAM
MARS
Adabound SWAN
Adopt
Adam
K-FAC
APOLLO
OCAR Shampoo
SPAM
TeZO LORENZA LOZO MeZO
NIGT
4bitSPlus Shampoo ASGO R-AdaZO
Clipped-SGD
TK-FAC AdaFisher
Natural Gradient
MAC ZO-AdaMM
MeZO-SVRG
SGN FZOO QZO
Color Ranking
Sophia
First-Order Methods
ZO-AdaMU
LeZO
Citation Ranking
GD
BFGS
ADAHESSIAN Q-Newton
S-BFGS mL-BFGS
Base Methods NAG RMSProp
Adadelta AdaGrad
Momentum SGD
Figure 1 The evolutionary tree of optimization methods. Rooted in classic base methods, the development trajectory branches into first-order, second-order, and zeroth-order paradigms. Node sizes reflect citation impact, while distinct clusters illustrate the progression from foundational gradient updates towards advanced, scenario-tailored frameworks.
fundamentally re-architected into scenario-oriented paradigms to address severe physical bottlenecks, such as distributed communication barriers and strict differential privacy constraints (Sec. 3). This analysis highlights the shift from pure algorithmic design to systems-aware engineering solutions that balance theoretical guarantees with practical constraints. 3) Standardized Evaluation: We introduce a rigorously controlled evaluation framework that separates algorithmic performance from large-scale engineering optimizations. We develop a standardized testbed to evaluate 23 distinct optimizers across diverse architectural proxies, including CNN LeCun et al. (2002) and Transformer-based Vaswani et al. (2017) models. This extensive evaluation systematically isolates and examines learning rate sensitivity, long term training scalability, and cross-architecture generalization, providing quantitative and qualitative insights that guide the design of next generation’s optimizers.
tion frameworks. Given the vast volume of literature, including both published articles and preprints, we primarily include representative and impactful works. Survey pipeline. Fig. 1 illustrates the evolutionary trajectory of optimization methods, systematically classified by their respective gradient orders. We first cover the background knowledge, including the concept of fundamental optimization methods (Sec. 2). Next, we conduct a comprehensive review of various optimization methods categorized by their gradient information usage and application scenarios (Sec. 3). We then provide comprehensive experimental evaluations of representative optimizers (Sec. 4). In Sec. 5, we critically examine remaining challenges and highlights potential directions for future research. Finally, Sec. 6 provides a concise summary of the survey. We closely follow the latest developments in this project.
Survey scope. This survey primarily focuses on four core optimization paradigms: FO, SO, and ZO optimization, along with scenario-oriented optimiza2
2
Background
of optimization algorithms, we formally define convexity and strong convexity. Let f : Rd → R be a differentiable function representing the objective loss.
This section establishes the core conceptual foundations, critical challenges, and standardized evaluation metrics underpinning modern deep learning optimization. We first formalize the canonical optimization problem setup for empirical risk minimization, then analyze the unique non-convex loss landscape of deep neural networks, and finally survey key evaluation metrics and hardware-aware scaling strategies tailored to large-scale model training.
2.1
A differentiable function f is convex if, for all x, y ∈ Rd , its first-order Taylor approximation provides a global underestimator: f (y) ≥ f (x) + ⟨∇f (x), y − x⟩.
This property implies that every local minimum is necessarily a global minimum. If f is twice differentiable, convexity is equivalent to the condition that the Hessian matrix is positive semi-definite, denoted as ∇2 f (x) ⪰ 0, for all x ∈ Rd .
Optimization Problem Setup
Objective function. The fundamental goal of optimization is to find a set of parameters θ ∈ Rd such that the model f (·; θ) performs optimally on unseen data. Assuming data samples z = (x, y) are drawn from an unknown true distribution P, our objective is to minimize the expected risk, denoted as R(θ): R(θ) = Ez∼P [ℓ(f (x; θ), y)],
A function f is µ-strongly convex with parameter µ > 0 if, for all x, y ∈ Rd , the following inequality holds: µ f (y) ≥ f (x) + ⟨∇f (x), y − x⟩ + ∥y − x∥2 . (5) 2 This property guarantees the existence of a unique global optimal solution. Intuitively, strong convexity implies that the function grows at least quadratically. For a twice-differentiable function, this is equivalent to requiring the eigenvalues of the Hessian to be bounded below by µ:
(1)
where ℓ(·, ·) represents the loss function. Since the distribution P is unknown, R(θ) cannot be computed directly. Following the principle of Empirical Risk Minimization Vapnik (1999), we use a training set S = {z1 , . . . , zN } consisting of N independent and identically distributed (i.i.d.) samples to construct a surrogate objective function, the empirical risk J(θ):
∇2 f (x) ⪰ µI,
Convergence. Under classical convex optimization theory Bottou (2010); Nesterov et al. (2018), gradient descent (GD) Cauchy et al. (1847) typically exhibits sublinear convergence with a rate of O(1/T ) for convex functions. Due to gradient noise variance, SGD Robbins and √ Monro (1951) has a slower sublinear rate of O(1/ T ). For strongly convex functions, GD Cauchy et al. (1847) achieves linear convergence with a rate of O((1 − µη)T ), where η is the learning rate. SGD Robbins and Monro (1951) exhibits linear convergence to a fixed error radius, which depends on the gradient noise variance and batch size; however, to attain exact convergence, the rate must reduce to sublinear.
(2)
While J(θ) converges to R(θ) as N → ∞, in finitesample regimes, excessive optimization of J(θ) may lead to overfitting. Therefore, the design of an optimizer must consider not only the minimization of training loss but also the generalization gap, often addressing this through regularization or implicit bias. Stochastic gradients. Computing the full gradient PN ∇J(θ) = N1 i=1 ∇ℓ(zi ; θ) is computationally prohibitive for large datasets. Consequently, modern optimizers universally adopt stochastic gradients. At iteration t, a mini-batch Bt is sampled to compute the gradient estimate gt (θ): 1 X gt (θ) = ∇ℓ(z; θ). |Bt |
(6)
where I denotes the identity matrix.
N
1 X ℓ(f (xi ; θ), yi ). min J(θ) = N i=1 θ∈Rd
(4)
2.2
Deep Learning Optimization
While convex theory provides intuition, the loss landscape of deep neural networks (DNNs) LeCun et al. (2015) is highly non-convex, presenting unique challenges that violate standard assumptions. Non-convexity. The loss function of deep neural networks contains numerous local minima. Deep learning optimization leverages inherent gradient noise and high-dimensional geometry to escape saddle points,
(3)
z∈Bt
For theoretical analysis, it is standard to assume that gt is an unbiased estimator of the true gradient, i.e., E[gt ] = ∇J(θ), with bounded variance E[∥gt − ∇J∥2 ] ≤ σ 2 . Convexity. To analyze the convergence properties 3
Total Count
Count
100
Second-Order Methods
First-Order Methods
First-Order: Derivative-based, computationally efficient.
Derivative-based
80
First-Order: Relies on firstderivative information of the objective function.
48
Gradient-free
40
20
0
6
11 1
11
10
< 2019
2019
3
1 2
77
18
10
7
Zeroth-Order: Relies only on sampled function values, with no gradient information.
10
18
18 1
vations and optimizer states. Activations are intermediate results generated during forward propagation that must be retained for gradient calculation during backward propagation, though their memory footprint can be reduced via gradient checkpointing Chen et al. (2016). Optimizer states introduce substantial overhead because methods like Adam Kingma (2015) require storing first-order and second-order moment estimates. Furthermore, privacy and security considerations are necessary when processing sensitive data. Techniques such as differentially private SGD (DP-SGD) Abadi et al. (2016) provide privacy guarantees through gradient clipping and the addition of Gaussian noise, but this often reduces convergence speed and final model accuracy.
Zeroth-Order: Gradient-free, random perturbation.
Observation: Recent Sharp Increase(2024-2025): This rapid growth corresponds to the emergence of massive deep learning models.
Second-Order: Incorporates second-derivative (Hessian or its approximation) information.
60
Zeroth-Order Methods
Second-Order: Curvature-based, higher precision.
14 1
12
1 1
15
15
11
11
13
2020
2021
2022
2023
8 31
49
8
2024
2025
2026
Year
Figure 2 Quantitative evolution of optimization methods over time. The stacked bars categorize algorithms into FO, SO, and ZO algorithms, with the line graph tracking the total annual count. A significant surge is observed starting in 2024, driven by the rapid development and optimization demands of massive deep learning models. Please note that the statistics for 2026 are incomplete, covering data only up to April. Based on the prevailing upward trend, the total count for 2026 is expected to continue growing; thus, the current partial figure should not be misinterpreted as a decline in research activity.
2.4
As model parameter counts scale toward the billion or even trillion level, optimization strategies must account for hardware constraints. Mixed precision training. Using FP16 or BF16 instead of FP32 reduces memory footprint and enhances computational efficiency. Notably, FP16 requires loss scaling Micikevicius et al. (2018) to prevent numerical underflow of gradients, whereas BF16 typically does not need additional loss scaling because its dynamic range is consistent with that of FP32. Parallelism strategies. It is necessary to combine data parallelism, tensor/model parallelism Shoeybi et al. (2019), and pipeline parallelism Huang et al. (2019b) to overcome single-GPU memory limitations. Batch size and learning rate. Increasing batch size enhances parallel efficiency; however, to maintain convergence, the learning rate must typically be scaled according to specific rules (e.g., the linear scaling rule η ∝ B Goyal et al. (2017)). Gradient accumulation. When memory constraints prevent the use of large batch sizes, gradients are accumulated over multiple micro-steps before performing a single parameter update, thereby simulating the effect of large-batch training.
typically converging to flat minima that yield strong generalization performance. Saddle point. In high-dimensional spaces, local minima are relatively scarce, and saddle points, which are local minima in some directions and a local maximum in others, are far more prevalent. For first-order methods like GD Cauchy et al. (1847), saddle points cause stagnation and slow convergence; SGD Robbins and Monro (1951) can often escape saddle points due to stochastic gradient noise, making it more effective for nonconvex optimization. Over-parameterization regimes. Modern deep networks are typically over-parameterized, meaning the number of parameters far exceeds the number of training samples. Under this regime, the optimization process exhibits distinctive properties. The neural tangent kernel (NTK) Jacot et al. (2018) theory shows that, in the infinite-width limit with small initialization and gradient flow, the training dynamics of neural networks can be approximated by kernel regression. This provides a theoretical framework for understanding why over-parameterized networks converge to global optima and generalize well.
2.3
Optimization for Large Models
2.5
History and Roadmap
Before detailing modern optimizers, we summarize their growth trends (Fig. 2) and evolution (Fig. 3). SGD Robbins and Monro (1951) established the foundation of first-order training, while second-order methods like K-FAC Martens and Grosse (2015) sought curvature-aware efficiency despite high computational costs. The rise of large language models has shifted the paradigm toward resource efficiency. To bypass backpropagation’s memory bottlenecks,
Evaluation Metrics
Evaluation metrics are as critical as model accuracy in large-scale training. Computational complexity is typically measured in floating-point operations (FLOPs), a metric that directly correlates with training duration and hardware costs. Training memory consumption primarily stems from two sources: acti4
AdaGrad
SGD
[Ann. Math. Statist.]
Adam
[JMLR]
2011
1951
K-FAC
[ICLR]
[ICML]
2015
[NeurIPS]
mL-BFGS [TMLR]
2023
2023 Sophia
[ICLR]
2024
Lion
2015
AdamP
2018
LOMO
ADOPT
[NeurIPS]
2024
2024
Muon
2019 ZO-AdaMM LAMB
[NeurIPS]
[ICLR]
2019
2020
2021
2021
2023
[ACL]
[AAAI]
[ICLR]
[NeurIPS]
a supplementary classification based on application scenarios, such as distributed optimization and differential privacy. To facilitate a structured analysis of algorithmic evolution (Fig. 4, Fig. 5, Fig. 6, Fig. 7), we have unified the mathematical notation system (Tab. 1). This framework combines universal vanilla symbols with superscripted algorithm-specific notations to highlight the core differences between various methods. Furthermore, it explicitly delineates the specific improvements made by each algorithm relative to its predecessors. Specifically, Fig. 4 dissects the evolutionary logic of base methods from a formulation perspective: SGD Robbins and Monro (1951) improves upon GD Cauchy et al. (1847) by computing gradients based on minibatches, significantly enhancing computational efficiency; momentum Polyak (1964) introduces a variable vt to accumulate historical gradient information, ensuring parameter updates rely on more than just the current gradient; NAG Nesterov (1983) further refines this by computing the gradient at a "lookahead" position (θt + γvt−1 ), correcting potential overshooting caused by momentum Polyak (1964); AdaGrad Duchi et al. (2011) introduces a cumulative term Gt = Gt−1 + gt2 to achieve parameter-wise adaptive learning rates; RMSprop Tieleman and Hinton (2012) replaces AdaGrad’s direct accumulation with an exponential moving average E[g 2 ]t , effectively preventing premature learning rate decay caused by an infinitely growing denominator; Adadelta Zeiler (2012) eliminates the reliance on a manually set global learning rate η, instead utilizing the root mean square p of parameter updates E[∆θ2 ]t−1 for adaptive adjustment. We trace the evolutionary paths of firstorder (Fig. 5), second-order (Fig. 6), and zeroth-order (Fig. 7) algorithms through representative methods, providing detailed theoretical discussions for each category in Sec. 3.2, Sec. 3.3, and Sec. 3.4. Motivation of survey organization. Current research on emerging optimizers lacks a systematic classification framework and comparison of core logic, for several reasons: 1) Coarse classification granularity: Existing surveys often lump diverse algorithms under generic labels such as "adaptive methods" without dissecting underlying mechanisms, such as specific momentum improvement strategies (e.g., scheduled reset vs. lookahead) or different types of gradient clipping. 2) Disconnect from application scenarios: Current literature lacks practical guidance on which optimization logic applies best to specific engineering hurdles, such as high memory consumption in large models or noise sensitivity in differential privacy. 3) Expanded optimization scope: The scope of optimization has expanded beyond narrow gradient
AdamW [ICLR]
ADAHESSIAN
MeZO
Shampoo [ICML]
MARS
[GitHub]
[ICML]
2024
2025
LOZO [ICLR]
2025
FANoS [arXiv]
2026
HomeAdam
[arXiv]
2026
Figure 3 Timeline of prominent optimization algorithms. The evolution highlights key algorithmic milestones, associated research institutions, and publication venues over time.
zeroth-order methods and memory-efficient designs have emerged. Recent developments, including automated rules like Lion Chen et al. (2023), matrix-wise approaches like Muon Jordan et al. (2024), and lowrank adaptations like LOZO Chen et al. (2025b), demonstrate a clear trend toward hardware-aware optimization. Detailed taxonomies of these methods are provided in Sec. 3.
3
Optimization Algorithms
In this section, we conduct a comprehensive review of various optimization algorithms categorized by their gradient information usage and application scenarios. We systematically delineate the fundamental mathematical principles of neural network optimization across different gradient orders, FO, SO, and ZO, along with their evolutionary development and prevailing challenges (Secs. 3.2 to 3.4). However, deploying these theoretical paradigms in the engineering practice of modern deep learning scenarios encounters severe bottlenecks. We therefore shift our research perspective from pure algorithmic design to scenario-driven optimization paradigms, systematically investigating how the aforementioned FO, SO, and ZO optimization primitives can be re-engineered to breach the memory wall, mitigate communication bottlenecks, and enforce rigorous privacy guarantees (Secs. 3.5 to 3.8). As illustrated in Fig. 1, to provide a comprehensive overview of the development of optimization algorithms, we visualize the existing landscape as an "evolutionary tree". Rooted in classic base methods such as GD Cauchy et al. (1847), SGD Robbins and Monro (1951), and momentum Polyak (1964), this tree branches into three distinct categories based on technical trajectories: first-order, second-order, and zeroth-order algorithms. Furthermore, the latter part of this section provides 5
GD Efficiency
Minimizes an objective function by updating parameters along the negative direction of the gradient
History
Approximates the global gradient based on minibatches
Replaces accumulation with an exponential moving average, tracking only the gradient
NAG
Momentum
SGD
Lookahead
Introduces momentum, leveraging past gradients for parameter updates
Prevents momentum overshoot using lookahead gradient estimation
RMSprop AdaGrad EMA Adadelta
Replaces accumulation with an exponential moving average
Dynamic learning rate scaling via the sum of squared historical gradients
Figure 4 Evolution and formulations of base methods. We details the developmental trajectory from GD Cauchy et al. (1847), SGD Robbins and Monro (1951), Momentum Polyak (1964), NAG Nesterov (1983) to AdaGrad Duchi et al. (2011), RMSprop Zeiler (2012), Adadelta Tieleman and Hinton (2012), and highlights the core mathematical update rules and key transition mechanisms at each stage. Table 1 Nomenclature of key variables and hyperparameters in gradient-based optimization. This table summarizes the unified mathematical notations consistently used throughout the paper.
updates to include broad algorithmic strategies like zeroth-order estimation for memory efficiency and communication-efficient compression for distributed systems. This makes it difficult for researchers to quickly identify suitable application scenarios and for engineers to select optimal algorithms. By establishing a unified classification taxonomy and systematically analyzing algorithmic behaviors across diverse scenarios, this survey aims to inspire the structural design and theoretical advancement of nextgeneration optimization algorithms.
3.1
Symbol Description
Unified Mathematical Perspective
To provide a unified basis for comparing modern optimizers, we depart from conventional categorizations and formulate a generalized constrained optimization framework. Most methods discussed in this survey can be cast as specific instantiations of the following discrete-time dynamical system:
Model parameters (weights and biases) evaluated at iteration t.
η
Learning rate, determining the step size of parameter updates .
∇J(θt )
Exact gradient of the objective function with respect to parameters at step t
gt
Stochastic gradient (mini-batch gradient) of the objective function at step t
mt
Biased first moment estimate (exponential moving average of . past gradients)
vt
Biased second raw moment estimate (exponential moving average of past squared gradients)
β1 , β2
Decay rates for the first and second moment estimates
ϵ
. .
. .
Small smoothing constant added to prevent division by zero .
Gradient estimator E(·). This operator dictates how objective landscape information is acquired. 1) First-order: E yields the standard stochastic gradient gt . 2) Zeroth-order: Bypassing backpropagation, E constructs gradients via finite-difference forward queries. For instance, using random directions ui ∼ N (0, I) and a smoothing radius µ, it is typically Pq (θt −µui ) formulated as g̃tZO = 1q i=1 f (θt +µui )−f ui . 2µ
g̃t = E(f, θt , ξt ), ĝt = Tscenario (g̃t ), mt = ϕ(mt−1 , ĝt ), θt+1 = PΘ θt − ηt Mt−1 mt − ηt λθt ,
θt
(7)
where ξt denotes the stochastic data batch, ηt is the step size at iteration t, and λ represents weight decay. This unified formulation disentangles the optimization process into four distinct dimensions, allowing us to systematically categorize the evolutionary trajectories of existing methods:
Preconditioner Mt . This matrix warps the gradient vector to accelerate convergence in ill-conditioned spaces. 1) First-order adaptation. Mt is an operator or transformation matrix derived solely from first-order infor6
mation, such as historical gradient statistics or the inherent geometric structure of the gradient matrix itself. Crucially, approaches utilizing empirical Fisher Information Matrix (EFIM) ( Gupta et al. (2018)) are classified under this first-order category, as they inherently capture gradient variance rather than explicitly approximating true second-order curvature Kunstner et al. (2019). This spans from strict diagonal matrices √ (e.g., Mt = diag( vt + ϵ) in Adam Kingma (2015)) to matrix-level structural normalizations under specific operator norms (e.g., spectral orthogonalization in Muon Jordan et al. (2024)), all operating without computing the true Hessian or FIM. 2) Second-order algorithms. Mt explicitly incorporates high-order geb t + γI, ometry, typically taking the form Mt = H 2 f (θ ), F b t ∈ {∇\ bt } is an approximation of the where H t Hessian or the FIM, and γ is a damping factor.
sided preconditioning methods (such as ASGO An et al. (2025), where Mt uses low-rank gradient decomposition). In this framework, both approaches execute the same mathematical operation: they apply a dimensionality-reduced mapping on the Rd space to offset computational constraints on either the gradient estimator E or the scenario transformation Tscenario . Rather than treating memory-efficient optimizers as isolated heuristics, this perspective systematically shows they operate as joint compromises between Mt−1 and PΘ . Systematic design space. Furthermore, Eq. (7) outlines a comprehensive design space that guides the development of new optimizers. Abstracting existing algorithms into these four operators reveals fundamental theoretical gaps. For example, although Kronecker-factored methods like Shampoo Gupta et al. (2018) are computationally efficient, their reScenario-aware transformation Tscenario (·). This liance on the empirical Fisher information restricts non-linear operator introduces environmental contheir preconditioner Mt to capturing gradient varistraints directly into the gradient flow, serving as the ance rather than true geometric curvature. The procore mechanism for modern deployment challenges: posed formulation suggests that this limitation can be 1) Privacy-preserving. T (g̃t ) = Clip(g̃t , C)+N (0, σ 2 C 2 I),addressed by redesigning the interaction between the gradient estimator and the preconditioner. A pracensuring differential privacy via norm bounding and tical direction involves integrating memory-efficient Gaussian noise injection. 2) Distributed optimization. gradient estimators for E (such as those utilizing T (g̃t ) = C(g̃t ), where C(·) denotes a lossy compresmodel-distribution sampling to capture true Fisher sion operator (e.g., quantization, sparsification) to information) with structured or localized precondimitigate communication bottlenecks. tioners for Mt . Investigating such hybrid configuraStructural projection PΘ (·). In highly constrained tions provides a clear pathway to move beyond the scenarios, updates must be projected onto specific variance-adaptation limits of empirical methods, enstructural manifolds. For memory-efficient optimizers abling genuine curvature-aware optimization under (e.g., low-rank training), PΘ restricts updates to a strict memory constraints for large language models. low-dimensional subspace to circumvent the storage By viewing modern optimizers through the lens of Eq. (7), overhead of dense optimizer states. we shift the narrative from isolated algorithmic tricks As a theoretical framework, Eq. (7) is not intended to a holistic engineering co-design. The subsequent as a universal convergence proof framework for all sections will unfold how researchers systematically existing algorithms, but rather as a modular decouengineer E, Mt , T , and P to navigate the complex pling perspective. By abstracting complex optimizatrade-offs among theoretical convergence, hardware tion trajectories into four distinct structural operaefficiency, and systemic constraints. tors (E, T , ϕ, P), this theoretical framework reveals structural equivalences across distinct optimization 3.2 First-Order Algorithms paradigms. Identifying these equivalences simultaneously exposes theoretical gaps within the design We define FO methods as parameter update schemes space. Researchers can use these insights to intuthat construct an adaptive transformation or preitively recombine mathematical primitives, guiding conditioning operator Mt by exclusively utilizing the algorithmic design of next-generation optimizers. first-order gradient information, specifically includTheoretical insights. The unified formulation reing Empirical FIM, which capture historical gradient veals non-trivial structural equivalences across disvariance rather than explicitly approximating true tinct optimization paradigms. Examining the strucsecond-order curvature. Despite their widespread tural projection PΘ and the preconditioner Mt−1 adoption, vanilla stochastic gradient descent suffers highlights a connection between memory-efficient ZO from several fundamental limitations, including senmethods (such as LOZO Chen et al. (2025b), which sitivity to the choice of learning rate and instability restricts PΘ to a low-rank subspace) and FO singleunder noisy gradients. To address these challenges,
7
the research community has developed a rich ecosystem of fo variants that enhance the basic update rule across multiple complementary dimensions. This section systematically reviews these advancements, organizing them into eight core themes. Each theme represents a distinct evolutionary trajectory, yet together they form a coherent narrative of how fo algorithms have evolved to meet the escalating demands of contemporary deep learning architectures. We also trace the evolutionary trajectory of representative algorithms (Fig. 5) and systematically evaluate their strengths and limitations (Tab. 2). 3.2.1
which may interrupt the continuous accumulation of effective historical gradients and necessitate meticulous hyperparameter tuning. Accelerated momentum. Certain approaches surpass traditional momentum strategies by introducing complex momentum computations or look-ahead mechanisms to achieve faster convergence rates. From a unified derivation perspective, these methods estimate gradients at extrapolated future positions, denoted as gt (θt−1 − ηβmt−1 ), or apply explicit bias correction to the standard formulation. Specifically, adaNAPG Zhu et al. (2025a) and SGDO Kopal et al. (2025) focus on obtaining superior gradient estimates by looking ahead to future parameter configurations. The method adaNAPG Zhu et al. (2025a) combines the accelerated gradient of Nesterov Nesterov (1983) with adaptive sampling to handle stochastic composite optimization with optimal complexity. SGDO Kopal et al. (2025) directly improves the update direction by computing gradients at future weights, achieving zero additional memory overhead. Conversely, RSGDM Qin et al. (2024) and SQuARM-SGD Singh et al. (2021) introduce correction mechanisms to rectify inherent biases in vanilla momentum. The momentum implementation of MADGRAD Defazio and Jelassi (2022) eschews the exponential moving average of gradients, making it highly suitable for sparse models. RSGDM Qin et al. (2024) applies exponential moving averages on both the gradient and the differential of the gradient to proactively correct bias and lag within SGDM Rumelhart et al. (1986). SQuARM-SGD Singh et al. (2021) integrates the momentum of Nesterov Nesterov (1983) with local stochastic gradient descent in distributed learning, where event-triggered communication acts as a conditional correction of communicated gradients. The primary advantage of this category is the attainment of higher-order convergence performance through lookahead estimation. The corresponding limitation is the increased computational complexity per step and potential instability when future gradient approximations contain high variance. Double momentum mechanism. To address the limitations of single momentum configurations in complex optimization scenarios, recent techniques employ multiple momentum terms or recursive estimators to capture gradient information comprehensively. The unified formal derivation entails the construction of a secondary momentum buffer or recursive estimator, (2) (1) formulated conceptually as mt = f (mt , gt ), to reduce variance and bias systematically. The algorithm µ2 -SGD Dahan and Levy (2025) integrates anytime averaging with corrected momentum, achieving optimal convergence rates and enabling a fixed learning rate. AdEMAMix Pagliardini et al. (2025) utilizes
Accelerating Convergence Rate
While momentum techniques significantly accelerate optimizer convergence speed, traditional momentum frameworks exhibit inherent limitations in complex optimization landscapes. The general mathematical framework of momentum methods typically formulates the update rule as mt = βmt−1 + γgt and θt = θt−1 − ηmt . To transcend the boundaries of this standard formulation, recent studies explore several core improvement dimensions. These dimensions encompass the adjustment of momentum resetting and weighting, the enhancement of computational look-ahead and direction alignment, the introduction of dual momentum mechanisms, and the application of novel analytical perspectives from the frequency domain and dynamical systems. Scheduled momentum reset. Several methodologies address issues such as update bias or overfitting to embedding layers in vanilla momentum methods under non-stationary objectives. These issues typically emerge from inappropriate accumulation of historical information. The unified formal approach involves active resetting or selective application of momentum states, conceptually modifying the update to mt = γt mt−1 +gt , where γt acts as a binary or continuous decay trigger. For instance, SRSGD Wang et al. (2022) resets the momentum buffer via linear or exponential scheduling to prevent the accumulation of the error of Nesterov accelerated gradient under stochastic conditions, thereby achieving simplicity and efficiency. Adam-Rel Ellis et al. (2024) targets non-stationarity in reinforcement learning by resetting time steps to balance momentum estimation, which avoids large updates caused by abrupt gradient changes. Furthermore, the LazyOptimizer Su (2019) selectively updates active parameters through zerogradient checks. The advantage of these methods lies in their capacity to enhance the adaptability of optimizers in dynamic tasks through adaptive momentum resetting. However, the inherent limitation of this category is the reliance on heuristic reset triggers,
8
AdamW
Adam
Decoupled
First-Order Methods
Lion
Sign
Dynamically adapts individual learning rates
Weight decay term λ is decoupled from the gradient update
Completely removes the second moment, retaining only sign information
Nadam
MADGRAD
Adan
Accelerating
Revise
Dual
Introduces NAG, replacing standard momentum
Replaces Nesterov momentum with momentumized dual averaging
Reformulates the Nesterov look-ahead mechanism
4bit-Shampoo
ASGO
Shampoo Quantization
Captures matrix structural geometry via preconditioning
Applies 4-bit quantization to the dense matrices
Light
Utilizes a single-sided preconditioner to reduce memory
Figure 5 Evolution and formulations of typical first-order methods. We show the prominent algorithms from Sec. 3.2, categorizing these algorithms by adaptive strategies ( Kingma (2015); Loshchilov and Hutter (2019); Chen et al. (2023)) acceleration ( Dozat (2016); Defazio and Jelassi (2022); Xie et al. (2024b)), and curvature approximation ( Gupta et al. (2018); Wang et al. (2024f); An et al. (2025)), and analyze their key transitional mechanisms through the lens of mathematical formulations.
historical momentum mt−1 and the current gradient gt . AngularGrad Roy et al. (2021) utilizes consecutive gradient angles to control the step size, thereby reducing zig-zag effects and ensuring smooth convergence paths. The Cautious Optimizer Liang et al. (2024) modifies momentum methods by strictly aligning update directions with gradients, which preserves convergence guarantees and is implementable in a single line of code. The advantage of these alignment methods is the effective improvement of convergence smoothness and speed by preventing historical updates from deviating from current optimization trends. The limitation is that enforcing strict alignment may discard beneficial acceleration provided by historical momentum in narrow ravines, potentially slowing down the escape from shallow local minima. Dynamic momentum weight. To overcome the constraints of fixed momentum coefficients, several studies dynamically adapt weights based on gradient variations. The unified derivation replaces the static coefficient with a dynamic function, computing the state as mt = βt mt−1 + (1 − βt )gt , where βt adapts strictly to current gradient variations. For
a mixture of fast and slow exponential moving averages, adaptively balancing gradient relevance through schedulers to mitigate the forgetting of models during training. YOGI Zaheer et al. (2018) employs an additive adaptive update for the second-moment accumulator with signed and bounded increments, which prevents abrupt increases in the effective learning rate. MARS Yuan et al. (2025) unifies variance reduction with adaptive methods via scaled stochastic recursive momentum. The advantage of double momentum mechanisms is the substantial improvement in optimization robustness and gradient estimation precision. The inherent limitation is the introduction of additional memory overhead to store multiple momentum states and the challenge of balancing the interaction between fast and slow moving averages. Momentum-gradient alignment. Addressing the persistent zig-zag effect and suboptimal convergence paths in vanilla momentum methods, this category of research aligns momentum directions with current gradients. The formal mechanism operates by modulating the step size or update vector based on the cosine similarity or angular metric between the
9
example, DEAM Bai et al. (2020) computes adaptive weights using the discriminative angle between historical momentum and the current gradient, employs a backtrack term to restrict redundant updates, and reduces hyperparameters compared to the mechanism of Adam Kingma (2015). Furthermore, SGDF Yao et al. (2023b) leverages the principles of the Wiener filter to dynamically adjust gradient estimation, minimizing mean square error while balancing noise reduction and signal preservation. The advantage of dynamic weighting is the enhancement of convergence and generalization through tailored momentum adjustments. The primary limitation is the computational overhead required to evaluate the dynamic weight function at each iteration, along with the risk of vanishing momentum if the coefficient decays excessively. Frequency domain momentum analysis. This dimension provides a novel perspective to interpret and optimize momentum methods via frequency domain insights. By conceptualizing the momentum operation as a time-varying filter H(ω, t) in the frequency domain, the framework presented in Li et al. (2025b) demonstrates that high-frequency components become detrimental during the late stages of optimization, whereas low-frequency signals require amplification. FSGDM Li et al. (2025b) dynamically adjusts momentum filtering characteristics based on these specific insights. The distinct advantage of this analytical paradigm is the ability to systematically isolate and suppress harmful noise frequencies. However, the limitation resides in the theoretical complexity of mapping time-domain gradients to frequency-domain filters in highly non-linear deep neural networks, which complicates practical hardware implementations. Momentum damping mechanism. To resolve convergence instability and overshoot issues in large-scale machine learning tasks, researchers model the optimization process as a controlled physical or dynamical system. The unified formal approach involves introducing a damping factor or feedback controller, conceptually modifying the standard discrete update into a robust differential equation representation, such as θ̈ + γ θ̇ + ∇f (θ) = 0. The PIDAO Chen et al. (2024) framework and PID-based approaches An et al. (2018) treat traditional SGDM Rumelhart et al. (1986) as a controlled heavy ball system, incorporating a PID controller driven by gradient history to suppress overshoot. This idea of dynamically regulating system states via feedback mechanisms finds a deeper physical extension in FANoS Dhiman (2025), which formulates the momentum update as a discretized second-order dynamical system and introduces a thermostat as an integral controller. Furthermore,
VRAdam Vaidhyanathan et al. (2025) dampens excessive velocity via a dynamic learning rate, while AdamP Heo et al. (2021) geometrically removes the radial momentum component responsible for unwarranted weight norm growth. SNGM Zhao et al. (2024c) eliminates the dependence on batch size by normalizing gradients prior to momentum accumulation. The advantage of damping mechanisms is the critical stabilization of parameter evolution. The inherent limitation is the introduction of complex control dynamics that may require precise tuning of damping coefficients to prevent premature convergence. Regarding the theoretical convergence analysis, the unified mathematical framework across these diverse dimensions typically guarantees optimal convergence √ complexities, such as O(1/ T ) for non-convex objectives, by strictly bounding the variance of momentum estimators. However, the theoretical boundaries and limitations of these methods often assume bounded gradients or specific smoothness conditions. These mathematical assumptions are frequently violated in the highly non-convex loss landscapes of modern deep learning architectures, which restricts the theoretical guarantees from perfectly translating into empirical success. In terms of core design paradigms and evolution logic, the field has witnessed a milestone paradigm shift from static, heuristic momentum accumulation to dynamic, feedback-driven momentum modulation. This evolutionary trajectory transitions from basic scalar adjustments to complex geometric, physical, and frequency-domain interventions. Despite these comprehensive advancements, the core gap in existing research remains the lack of a universal, hyperparameter-free momentum framework. Current methodologies still struggle to automatically identify and adapt to the spectral properties of the loss landscape without incurring prohibitive memory overhead or computational costs. The future trajectory necessitates bridging the divide between rigorous continuous-time dynamical system models and the discrete, highly stochastic nature of empirical optimization. 3.2.2
Adaptive Step-Size Control
Adaptive learning rates are widely utilized to reduce the cost of manual hyperparameter tuning. The generic update framework of these methods can be formulated as θt+1 = θt −ηt ϕ(g1:t )/ψ(g1:t ), where ϕ(·) and ψ(·) denote the first-order momentum and the second-order preconditioner, respectively. Although the classic optimizer, such as Adam Kingma (2015), establishes the foundation for this paradigm, it ex10
Table 2 Taxonomy and Comparison of Representative Optimization Methods. This table presents a systematic evaluation of various optimizers categorized into FO, SO, and ZO paradigms, shows their limitations and highlights. Optimizer
Venue
Type
SGDM Rumelhart et al. (1986)
Nature’1986
Accelerating Convergence Rate
Adam Kingma (2015)
ICLR’15
AdaBelief NeurIPS’20 Zhuang et al. (2020)
Limitation First-Order Methods ( Sec. 3.2)
Strategies for Adaptive Step-Size Estimation Strategies for Adaptive Step-Size Estimation Approximating Curvature Information Approximating Curvature Information
Complex hyperparameter tuning.
First proposed to introduce a momentum mechanism to accelerate convergence.
High memory usage, worse generalization.
First proposed a method combining adaptive step-size estimation with momentum.
ϵ requires task-specific tuning.
Compared to adaptive methods like Adam, it achieves better generalization.
Computationally expensive.
First proposed to exploit tensor structures to maintain preconditioners
Shampoo Gupta et al. (2018)
ICML’18
4-bit Shampoo Wang et al. (2024f)
NeurIPS’24
AdaNorm Dubey et al. (2023)
WACV’23
Enhancing training stability
Add computational overhead.
AdaBound Luo et al. (2019)
ICLR’19
Learning Rate Scheduling
Convergence may slow in later stages.
SAM Foret et al. (2021)
ICLR’21
GAM CVPR’23 Zhang et al. (2023c)
Enhancing Generalization Ability Enhancing Generalization Ability
Highlight
Computationally expensive.
Additional gradient computation steps. Computationally expensive.
AdaSGD Wang and Wiens (2020)
arXiv’20
Hybrid Methods
Inferior convergence speed compared to Adam.
LOMO Lv et al. (2024)
ACL’24
Towards LLMs Training
Slow training throughput.
AAAI’21
Hessian Approximation&Estimation
Compared to standard Shampoo, it significantly reduces memory usage by quantizing optimizer states to 4-bit. Compared to Adam, it proposes to solve the slow convergence issue caused by gradient anomalies through adaptively scaling updates. Compared to Adam, it achieves SGD-level generalization capability through dynamic bound transitions. First proposed to enhance model generalization by simultaneously minimizing loss value and sharpness. Compared to SAM, it proposes first-order flatness as a stronger measure to further enhance generalization. First proposed to combine Adam’s convergence speed with SGD’s generalization via an adaptive global learning rate. First proposed to reduce gradient memory usage in large model training by fusing gradient computation and parameter updates.
Second-Order Methods( Sec. 3.3) AdaHessian Yao et al. (2021)
K-FAC Martens and Grosse ICML’15 (2015) MAC Seung et al. (2025)
arXiv’25
mL-BFGS Niu et al. (2023)
TMLR’23
Sophia Liu et al. (2024a)
ICLR’24
SGDHess Tran and Cutkosky (2022)
NeurIPS’22
ZO-AdaMM Chen et al. (2019)
NeurIPS’19
MeZO NeurIPS’23 Malladi et al. (2023) Addax Li et al. (2025c)
ICLR’25
LOZO Chen et al. (2025b)
ICLR’25
MeZO-SVRG Gautam et al. (2024)
ICLR’24
ZoPro Wang et al. (2024a)
CDC’24
First proposed to use Hutchinson’s method to estimate the Hessian diagonal distribution, achieving superior convergence accuracy and speed. First proposed factoring the FIM into Fisher Information Higher computational cost. Kronecker products, thereby achieving faster Matrix Application convergence. Compared to other second-order methods, it Fisher Information Inferior accuracy relative to greatly simplifies curvature estimation, Matrix Application other second-order methods. achieving memory usage and training speed comparable to SGD. Introduce additional Compared to standard L-BFGS, it effectively Quasi-Newton memory and computational overcomes its instability by introducing a Methods complexity. momentum mechanism. Compared to traditional methods, it achieves Curvature-Guided Lack the maturity. faster convergence through curvature-guided Preconditioning preconditioning. Higher per-step First proposed to utilize Hessian-vector Second-Order computational overhead and products to correct momentum bias in SGD, Moment Fusion complexity. achieving a faster convergence speed. Zeroth-Order Methods( Sec. 3.4) First proposed to generalize the adaptive Adaptive Methods Slow convergence speed. momentum method to zeroth-order optimization. Compared to other methods, it adapts Perturbation Require more iterations zeroth-order SGD to operate in-place, Optimization than first-order methods. drastically reducing memory footprint. First proposed to dynamically select between Lower convergence speed Zeroth-First Order first-order and zeroth-order gradient compared to first-order Hybrid computations, achieving faster convergence methods. speed compared to pure zeroth-order methods. Exhibits a performance gap First proposed to leverage the intrinsic Memory-efficient compared to first-order low-rank structure of gradients to achieve faster Methods methods. zeroth-order convergence. Exhibits a performance gap First proposed to incorporate SVRG into Variance compared to first-order zeroth-order optimization, improving Reduction methods. convergence stability and downstream accuracy. Theoretically only converges Compared to other zeroth-order methods, it Distributed to a neighborhood of the incorporates gradient approximations into a Zero-Order optimum rather than the distributed framework, achieving faster Optimization exact solution. convergence. Higher per-step computational cost and memory footprint.
11
hibits limitations regarding extreme step sizes and domains. Similarly, FAdam Hwang (2024) integrates non-convergence in complex landscapes. Subsequent adaptive adjustments through natural gradient coradvancements systematically evolve along structural rections based on Riemannian geometry. These apdimensions: refining temporal moment estimation, proaches effectively prevent division by zero anomaexpanding spatial adaptation granularity, decoupling lies, yet their theoretical bounds often rely on empirarchitectural components, and minimizing state overical approximations. head. These developments share a unified objective Prediction deviation adaptation. By substitutof bounding the gradient variance while ensuring theing the standard gradient variance with the discreporetical convergence under non-convex settings. To ancy between observed and predicted gradients, Adaddress the temporal dynamics of gradient sequences, aBelief Zhuang et al. (2020) scales the update direcoptimizations primarily target the internal formulations to achieve rapid convergence and strong gention of ϕ(·) and ψ(·). eralization. Aida Zhang et al. (2023a) extends this Bias correction rules adaptation. Initial training belief-based scaling by utilizing first-momentum proinstabilities are mitigated by modifying the earlyjections for second-momentum estimation to strictly stage estimators. For instance, AdamD John (2021) suppress step size ranges. This framework optimally removes the bias correction on the first-order mocouples the learning rate with gradient uncertainty, almentum estimate to generate smaller initial updates, though it remains susceptible to highly noisy stochaswhich stabilizes the warm-up period. While this tic batches. strategy enhances early alignment with the gradient Momentum-based adaptation. The integration geometry, the inherent limitation lies in the marginal of momentum mechanisms modifies ϕ(·) to accelimpact on late-stage convergence. erate convergence through phase-shifted accumulaSecond-order moment adaptation. To bound extions. Nadam Dozat (2016) and Adan Xie et al. treme step sizes and resolve the non-convergence (2024b) incorporate Nesterov Nesterov (1983) moof variance estimates, a multitude of strategies rementum, while RAdam Liu et al. (2020a) rectifies formulate the exponential moving average of ψ(·). the variance of early updates. Other structural inteNosAdam Huang et al. (2019a) increases the weight of grations include utilizing extrapolated first moments historical gradients, whereas ADOPT Taniguchi et al. in Adam+ Liu et al. (2020b), quasi-hyperbolic terms (2024) eliminates the dependency on the decay rate in QHAdam Ma and Yarats (2019), and adapted β2 . Furthermore, AdamNX Zhu et al. (2025c) adaprate models in MoMo-Adam Schaipp et al. (2024). tively adjusts the exponential decay rate, and SETConversely, Lion Chen et al. (2023) isolates the sign Adam Zhang (2024) calibrates moments to prevent intopology of the momentum to eliminate second-order finitesimally small step sizes. HomeAdam Huang et al. dependencies. Further augmentations involve second(2026a) removes the square-root in the second-order order injections in AdaInject Dubey et al. (2022), momentum to bound unstable step sizes and improve alternating curvature thresholds in EAGLE Fujimoto generalization. Additional variants incorporate differand Nishi (2025), recursive scaling in MARS Yuan ential privacy corrections in DP-AdamBC Tang et al. et al. (2025), and Newton-Schulz soft-thresholding (2024), higher-order moments in HAdam Jiang et al. in ROOT He et al. (2025b). INNAprop Shin et al. (2019), sign-power transformations in AdamPower Wang (2025b) additionally merges structural information et al. (2025d), dynamic weighting in MADAM Zhu with RMSprop Tieleman and Hinton (2012). The et al. (2021), direct variance adaptation in M-SVAG Balles collective advantage of these methods is the signifiand Hennig (2018), probabilistic generalization in cant acceleration of the optimization trajectory, but VSGD Kuzina et al. (2025), and memory-efficient facthe inherent limitation is the accumulation of historitorization in AdaLomo Lv et al. (2023). The primary cal momentum may become dynamically inconsistent advantage of these methods is the rigorous stabiwith the current gradient field, thereby introducing lization of convergence trajectories; however, they update bias in complex or non-stationary loss landinherently introduce heightened sensitivity to auxilscapes. iary hyperparameters. Beyond temporal accumulation, the spatial granularDynamic epsilon adjustment. To further refine ity of adaptation shifts from scalar parameter-level the preconditioner without auxiliary overhead, alupdates to broader structural geometries. gorithms dynamically adjust the stabilization term Layer-wise adaptation. Uniform learning rates freϵ. EAdam Yuan and Gao (2020) applies the addiquently fail to accommodate the structural heterotive term directly to the uncorrected accumulators geneity and anisotropic scaling of deep networks. To to implicitly reduce the step size for small moments, facilitate massive batch training, LARS You et al. outperforming the baseline methods Kingma (2015); (2017) and LAMB You et al. (2020) scale the updates Liu et al. (2020a); Chen et al. (2025b) across diverse 12
of AlexNet Krizhevsky et al. (2012), ResNet-50 He et al. (2016), and BERT Devlin et al. (2019) by applying layer-wise normalization with non-convex convergence guarantees. NovoGrad Ginsburg et al. (2019) halves the memory requirements through layerwise norms while decoupling the weight decay, and AdaL Zhang et al. (2021) transforms raw gradients prior to accumulation. Coupled Adam Stollenwerk and Stollenwerk (2025) resolves the anisotropic embeddings of Adam by enforcing a uniform second moment. Transitioning to matrix-wise orthogonalization, Muon Jordan et al. (2024) utilizes NewtonSchulz iterations to standardize the spectral norm of updates across layers. CaAdam Genet and Inzirillo (2024) further integrates structural depth and connectivity through multi-strategy scaling. These layer-wise normalizations effectively resolve gradient vanishing issues in heterogeneous architectures, yet they fundamentally assume that the gradient distributions within a single layer are uniformly bounded. Neuron-level adaptation. At a finer granularity, AdaAct Seung et al. (2024a) adjusts the gradient updates inversely with the square root of the activation variance. By utilizing the exponential moving average of activation variances rather than gradient variances, it shares the learning rates across parameters that process identical input features. This approach isolates the variance at the activation level to maximize output stability, though the inherent limitation is the substantial memory footprint required for networks with expansive width.
(1983) acceleration in S3 Peng et al. (2025), asynchronous second-moment centering in ACProp Zhuang et al. (2021), and gradient difference step sizing in DiffGrad Dubey et al. (2019). Parameter regularization is isolated in AdamW Loshchilov and Hutter (2019), and the adaptivity is interpolated with signed methods in LaProp Ziyin et al. (2020). Finally, AdaMuon Si et al. (2025) incorporates variance adaptivity into orthogonal updates, and AdaFamily Fassold (2022) establishes a continuous algorithmic spectrum via hyperparameter blending. The primary advantage of hybridization is the simultaneous mitigation of multiple failure modes, but the inherent limitation is the resultant computational complexity and the intractability of theoretical convergence analysis. As model dimensions scale exponentially, the memory overhead of maintaining historical states necessitates optimal approximation strategies. Stateless adaptation. To eliminate the redundant storage of second-order momentum, AlphaGrad Sane (2025) employs L2 normalization and hyperbolic transformations supported by non-convex proofs. AutoDrop Wang et al. (2024c) leverages angular velocity saturation to adjust rates without supplementary hyperparameters. AdamS Zhang et al. (2025a) replaces the temporal variance with a momentum-current gradient denominator, thereby matching the memory efficiency of SGD while inheriting the configuration of AdamW Loshchilov and Hutter (2019). AEGDM Liu and Tian (2022) merges momentum with transformed gradient sums to ensure energy stability over baseline methods Rumelhart et al. (1986); Kingma (2015). These designs approach the memory limits of firstorder methods while retaining adaptivity, but they strictly rely on the assumption of locally stationary gradient distributions. Kalman filtering based method. RLEKF Hu et al. (2023) reorganizes network layers to accommodate the covariance tracking of Kalman filtering, approximating the dense error matrix with a sparse diagonal block structure. This efficiently utilizes second-order covariance information; however, the truncation errors of the sparse approximation inevitably accumulate across deep architectures.
To synthesize the complementary strengths of isolated paradigms, contemporary methods decouple orthogonal components and hybridize discrete algorithms. Decoupled learning rate and adaptability. Standard optimizers frequently conflate the global step size with local adaptivity. AvaGrad Savarese et al. (2021) explicitly decouples these components to optimize performance across diverse modalities, offering superior task-specific tuning capabilities. While this isolation prevents the degradation of the global convergence rate, it inherently expands the hyperparameter search space. Hybrid adaptive strategy. By fusing disparate adaptation criteria, hybrid models establish generalized optimization frameworks. AdaSGD Wang and Wiens (2020) merges the uniform trajectory of SGD Robbins and Monro (1951) with adaptive scaling, while EXAdam Adly (2024) introduces debiasing terms for moment interactions. Memory reduction is achieved through structured projections in APOLLO Zhu et al. (2025b). Furthermore, theoretical bounds are fortified by p-th order momentums and Nesterov Nesterov
The evolution of adaptive step-size control demonstrates a definitive paradigm shift from temporal scalar correction to spatial structural optimization, and ultimately towards memory-efficient geometric approximations. Early methodologies predominantly conceptualized the optimization landscape as an independent sequence of historical gradients, focusing on bounding the variance through first-order and second-order moment refinements. As network archi13
tectures increased in depth and heterogeneity, the paradigm transitioned to layer-wise and neuron-level normalizations. Currently, the trajectory converges upon stateless and hybrid adaptations, prioritizing the mathematical decoupling of momentum, variance, and weight decay to maximize the efficiency of largescale pre-training. Despite these milestones, a critical gap remains in the existing literature: the theoretical convergence guarantees of hybrid and structurally adaptive methods largely depend on strict assumptions of convexity or bounded smoothness, which fail to accurately characterize the highly non-convex and singular optimization landscapes of modern deep neural networks. 3.2.3
per-step runtime. Building upon this framework, 4bit Shampoo Wang et al. (2024f) adopts eigenvector matrix quantization technology to reduce memory costs. PSGD Li (2017) introduces a noise-robust criterion for preconditioner estimation that equilibrates the perturbations of preconditioned gradients with the perturbations of parameters. Additionally, SPlus Frans et al. (2025) enhances the stability of the framework of Shampoo Wang et al. (2024f) through instant-sign normalization and iterate averaging. It incorporates shape-aware scaling for adaptation to network width and combines bounded updates with historical eigenbases to ensure stability even under lower-frequency matrix inversions. Theoretically, these dual-metric preconditioners efficiently approximate natural gradient descent, accelerating convergence in highly non-convex landscapes. Nevertheless, their boundary of applicability is constrained by the substantial memory overhead required to store and invert covariance matrices for large-scale models. Preconditioners based on a single metric. To address the severe resource demands, subsequent research evolved to leverage a unified single metric. Within this unified formulation, algorithms adjust the update step based on a single structural proxy to balance efficiency and performance effectively. For instance, ASGO An et al. (2025) employs a singleside preconditioner that preserves the matrix structure, utilizing low-rank gradients to reduce computational overhead while maintaining convergence properties compared to the method of Gupta et al. (2018). Similarly, NYSACT Seung et al. (2024b) applies an eigenvalue-shifted Nyström approximation for scalable covariance estimation, bridging the gap between FO and SO algorithms through improved estimation accuracy and reduced resource consumption. Furthermore, the method of Hessian-aware scaling Smee et al. (2025) adjusts gradients based on local curvature to guarantee a unit step size. Theoretical convergence analysis indicates that these methods effectively improve the condition number of the optimization landscape, facilitating faster convergence under convex assumptions with a significantly lower memory footprint. However, the inherent limitation of single-metric approaches resides in their reduced capacity to capture the complex inter-dimensional correlations fully represented in dual-metric designs.
Variance Adaptation
Classic FO algorithms flatten parameters into vectors, thereby discarding inherent structural information. In contrast, preconditioning methods formulate the optimization update through a unified mathematical framework, typically expressed as θt+1 = θt − ηPL ∇J(θt )PR , where gradients are multiplied by approximations often derived from the empirical FIM. This formulation can be uniformly incorporated into the aforementioned framework Eq. (7), where the preconditioner Mt−1 is instantiated as the leftright multiplication operator PL · PR . While traditionally motivated as achieving curvature adaptation by approximating the Hessian, recent perspectives suggest that the EFIM primarily captures the noncentral second moment of gradients Kunstner et al. (2019). Therefore, these preconditioning techniques are increasingly understood as performing variance adaptation, mitigating gradient noise in stochastic optimization, while preserving the structural information of parameters. The core improvement dimensions of these methods involve efficiently constructing these preconditioning matrices while minimizing computational overhead, ultimately balancing optimization efficiency with potential generalization. The evolutionary logic of this field transitions from complex dimension-wise preconditioning to highly efficient single-metric approximations to mitigate severe resource bottlenecks. Preconditioners based on two metrics. To capture intricate structural information, the initial design paradigm utilized two distinct metrics, typically approximating the row and column covariance matrices. Shampoo Gupta et al. (2018) pioneers an online structure-aware algorithm by maintaining separate dimension-wise preconditioners constructed from second-order gradient statistics. This formulation achieves faster convergence compared to traditional optimizers while preserving a comparable
The core design paradigm of curvature approximation has evolved from Kronecker-factored dimension-wise preconditioning back to streamlined single-matrix preconditioning, driven by the critical need for extreme scalability in modern architectures. The advantage of methods based on two metrics is their comprehensive structural awareness, but their inherent limitation
14
is the high memory footprint. Conversely, the advantage of methods based on a single metric is their low computational barrier and memory efficiency, yet they suffer from limited curvature representation. The core gap in existing research remains the lack of fully dynamic and hardware-friendly preconditioners that can achieve the representational capacity of dualmetric methods while maintaining the algorithmic efficiency of single-metric methods, without requiring complex hyperparameter tuning or suffering from precision degradation during aggressive quantization. 3.2.4
tion utility. The inherent limitation is the increased computational overhead required to persistently track and integrate error feedback states. Dynamic gradient clipping. Transitioning from static to adaptive thresholds, dynamic clipping adjusts constraints based on historical gradient statistics to accommodate varying parameter sensitivities and sudden gradient spikes. Stable-SPAM Huang et al. (2025a) augments Adam Kingma (2015) with AdaClip to apply entry-wise clipping derived from exponential moving averages. Similarly, AdaGC Wang et al. (2025b) employs per-parameter moving average thresholds, maintaining broad compatibility with existing optimizers Kingma (2015) Zhuang et al. (2020). Both methods share a unified derivation relying on historical momentum to dictate local clipping boundaries. The advantage of dynamic adaptation is the universal applicability across diverse architectures and modalities. Conversely, the inherent limitation is the potential lag in threshold adjustment during rapid topological shifts in the loss landscape. Spike-aware gradient clipping. For highly specific non-convex scenarios where standard dynamic adjustment fails to react to extreme anomalies, specialized clipping targets abrupt variations. SPAM Huang et al. (2025b) integrates spike-aware clipping mechanisms directly into the Adam Kingma (2015) optimizer to mitigate the severe impact of gradient spikes. The advantage of this approach is the targeted enhancement of large-scale training stability against sparse anomalies. The inherent limitation is its narrow applicability, as the mechanism is predominantly effective only in regimes exhibiting extreme gradient variance. Element-wise gradient scaling. Moving beyond magnitude truncation, fine-grained scaling resolves training instability by individually adjusting each gradient element, thereby addressing the limitations of uniform scaling. SPAM Huang et al. (2025b) integrates momentum reset, spike-aware clipping, and sparse momentum into Adam to mitigate spikes and optimize memory consumption. AdaNorm Dubey et al. (2023) utilizes exponential moving averages of past norms to correct individual gradients, ensuring optimization consistency. Additionally, pbSGD Zhou et al. (2020) applies powerball functions for flexible scaling, providing theoretical convergence guarantees for non-convex objectives. The advantage of elementwise scaling is the maximum flexibility afforded to individual parameter updates. The inherent limitation is the substantial memory footprint required to maintain distinct statistical states for every parameter. Layer-wise gradient normalization. To balance optimization stability with memory efficiency, layer-wise methods normalize gradients at macroscopic group
Enhancing Training Stability
To address gradient-related optimization challenges such as gradient explosion and variance imbalance, contemporary stabilization methods can be unified under a general framework that modulates the gradient update step via functional transformations. The core improvement dimensions of these transformations are primarily categorized by processing granularity and statistical adaptation, evolving from static constraints to dynamic, distribution-aware calibrations. Basic fixed gradient clipping. Bounding gradient magnitudes provides a fundamental mechanism to stabilize convergence under heavy-tailed noise and satisfy constraints such as privacy preservation. Within this paradigm, Clipped-SGD Gorbunov et al. (2020) is an accelerated stochastic method utilizing fixed gradient clipping to suppress heavy-tailed noise. Furthermore, DP-SGD Yang et al. (2022) integrates gradient clipping with Gaussian noise injection for differentially private non-convex optimization. A unified characteristic of these approaches is the application of a predefined scalar threshold to restrict the gradient norm. The primary advantage of this category is the provision of robust noise resilience and strict privacy bounds. However, the inherent limitation lies in the reliance on manually tuned static thresholds, which often induce gradient bias and hinder optimal convergence. DP-enhanced gradient clipping. To address the biased optimization trajectory caused by rigid truncation, recent methods augment traditional clipping with error correction frameworks. DiceSGD Zhang et al. (2024b) introduces an error feedback mechanism to eliminate clipping bias in differentially private stochastic gradient descent Robbins and Monro (1951). This formulation allows for flexible clipping thresholds independent of problem-specific parameters while maintaining rigorous utility and privacy guarantees through updated convergence analysis. The distinct advantage of this method is the theoretical decoupling of privacy thresholds from optimiza-
15
levels. Stable-SPAM Huang et al. (2025a) incorporates AdaGN to compute moving averages of per-layer gradient norms, rescaling current gradients by the ratio of the historical mean to the root second moment to suppress norm explosions. MultiAdam Yao et al. (2023a) categorizes loss terms into separate groups and maintains second-order momentum individually to balance diverse objectives, improving convergence in systems such as physics-informed neural networks. AuON Maity (2025) automatically suppresses unstable updates by normalizing the entire momentum matrix. The advantage is the effective improvement of convergence in specialized multi-objective scenarios with reduced memory overhead. The inherent limitation is the assumption of uniform gradient behavior within a layer, which may inadvertently suppress heterogeneous feature learning. Mean-removal normalization. Targeting gradient noise and non-smooth optimization trajectories, statistical centering stabilizes updates by shifting gradients or corresponding moments. GCSAM Hassan et al. (2025) integrates gradient centralization into the ascent step of sharpness-aware minimization Foret et al. (2021) to reduce noise and computational overhead while enhancing generalization. AdamMC Sadu et al. (2023) normalizes the first-order moments of adaptive optimizers to yield shorter and smoother optimization paths. These methods share a unified core of mean-removal but operate on different analytical components. The advantage of centralization is the improved smoothness of the training path and enhanced generalization. The inherent limitation is the assumption of symmetrically distributed gradient noise, which can degrade performance under heavily skewed stochastic distributions. Noise-robust normalization. To ensure convergence consistency across varied modalities under atypical gradient conditions, advanced normalization techniques integrate comprehensive historical statistics. AdaGC Wang et al. (2025b) utilizes exponential moving averages to adaptively bound parameters while maintaining standard non-convex convergence rates comparable to Adam Loshchilov and Hutter (2019) Chen et al. (2023). AdaNorm Dubey et al. (2023) corrects gradients through historical norm averages, creating variants that counteract atypical updates. Furthermore, NIGT Cutkosky and Mehta (2020) modifies the momentum formulation to accelerate convergence on second-order smooth objectives, effectively eliminating the reliance on large batch sizes. The advantage of noise-robust methods is the mitigation of batch size constraints and universal stabilization across tasks. The inherent limitation is the complexity of theoretical convergence analysis, which often requires strict assumptions regarding the
underlying landscape smoothness. The evolutionary logic of these stability enhancement techniques reveals a clear milestone paradigm shift: progressing from rigid constraints to granular rescaling, and ultimately culminating in adaptive statistical calibrations. While existing literature provides a robust empirical foundation for gradient manipulation, a fundamental gap persists in bridging local gradient corrections with global topological guarantees. Current methodologies operate predominantly as reactive mechanisms based on historical observations. The critical missing link is a proactive, unified mathematical framework that inherently couples stability constraints with the intrinsic geometry of the loss landscape, thereby eliminating the reliance on auxiliary tuning parameters and heuristic statistical tracking.
3.2.5
Learning Rate Scheduling
Learning rate scheduling constitutes a fundamental mechanism to dynamically adjust the step size of the optimizer, balancing the speed of convergence and the stability of model training across distinct optimization phases. The general mathematical framework of these methods can be formulated as θt+1 = θt − ηt Φ(g1:t , Ht ), where Φ denotes the transformation function conditioned on historical gradients and structural priors. This formulation aligns with the unified perspective in Eq. (7) by absorbing the momentum and preconditioning operations into Φ, and specializing to the case of identity projection and no explicit weight decay. The core improvement dimensions of recent studies focus on the scheduling basis, the granularity of adaptation, the reduction of manual dependency, and the enhancement of theoretical stability. Through targeted strategies, these dimensions collectively improve training efficiency and generalization. Lightweight scaling and batch-aware scheduling. Addressing computational efficiency and batch dynamics, this paradigm optimizes scaling properties and layer-wise learning rates. SGD-SaI Xu et al. (2024) leverages a scaling mechanism for the initial learning rate based on the signal-to-noise ratio of gradients. This approach eliminates the necessity for complex adaptive methods by maintaining constant adjustments to the signal-to-noise ratio, which reduces memory consumption by half compared to previous approaches while preserving competitive performance. Furthermore, addressing the critical challenge of scaling batch sizes for faster deep learning training, LAMB You et al. (2020) utilizes layer-wise adaptive learning rates. This batch-aware scheduling 16
overcomes a key barrier to efficient large-scale deployment by preventing the degradation of model performance in tasks requiring both rapid convergence and high accuracy. The advantage of this category lies in its exceptional memory efficiency and computational simplicity, while the inherent limitation is the reliance on specific empirical initial values that may require careful tuning. Element adaptation and gradient angle scheduling. Progressing towards finer granularity and directional awareness, the second paradigm dynamically calibrates step sizes using dimension-level constraints and angular information. AdaBound Luo et al. (2019) applies dynamic bounds to element-wise learning rates to mitigate the adverse effects of extreme step sizes. This strategy balances the rapid initial training of adaptive methods and the superior generalization of SGD Robbins and Monro (1951), supported by rigorous theoretical convergence proofs. Simultaneously, to address suboptimal direction alignment and oscillatory updates in high-dimensional scenarios, ACMo Huang et al. (2021) operates as a first-moment optimizer that eschews second-moment estimates. It employs angle calibration to ensure that descent directions form acute angles with both current and historical gradients, which delivers convergence rates comparable to the methods of the Adam family. Similarly, HGM Sarkar (2025) introduces a hindsight mechanism to evaluate the cosine similarity between current gradients and accumulated momentum. It increases learning rates in coherent gradient regions and decreases them in oscillatory areas, retaining an efficiency similar to Adam Kingma (2015). Furthermore, AdaBFE Cao (2022) extends binary forward exploration by adapting learning rates per parameter dimension utilizing the information of the forward loss function. The advantage of these methods lies in their ability to correct suboptimal update directions, whereas the inherent limitation is the increased computational overhead for continuous angular or metric evaluation. Loss-sensitive and stability-aware scheduling. To further enhance the robustness of optimization in complex non-convex topologies, subsequent paradigms integrate loss landscape information and rigorous stability controls. DecGD Shao et al. (2025) decomposes gradients into surrogate gradients and vectors based on the loss function, adjusting learning rates according to loss information rather than squared gradients to achieve rapid convergence and better generalization. Addressing the critical challenge of maintaining optimization stability in scenarios where instability from large step sizes or stochastic noise causes deviation, LyAm Mirzabeigi et al. (2025) integrates the adaptive moments of Adam with stability mecha-
nisms based on the Lyapunov function. This integration dynamically adjusts learning rates, enhancing robustness against noise while providing guarantees for non-convex optimization. Multistage SGDM Liu et al. (2020d) also utilizes a Lyapunov function to govern momentum deviation, enabling larger initial step sizes and smaller subsequent ones while matching the convergence rates of SGD. Additionally, the analysis of large step size SGD Andriushchenko et al. (2023) models the dynamics through stochastic differential equations, revealing loss stabilization and implicit sparsity emerging from multiplicative noise. These methods provide the advantage of theoretical rigor and enhanced stability, but they suffer from the inherent limitation of relying on stringent theoretical assumptions that may not perfectly align with highly stochastic real-world data distributions. Scheduler-free adaptation. The most recent milestone paradigm shift focuses on the complete elimination of manual learning rate scheduling through data-driven metrics and unified theoretical models. Schedule-Free Defazio et al. (2024) unifies scheduling with momentum interpolation. AutoDrop Wang et al. (2024c) utilizes angular velocity saturation without requiring extra tuning. Other parameter-free or automated variants include AdaS Hosseini and Plataniotis (2020), which utilizes knowledge gain per block, and SGD-G2 Ayadi and Turinici (2021), which combines the Runge-Kutta method with adaptive SGD Robbins and Monro (1951). Furthermore, AutoSGD Surjanovic et al. (2025) evaluates neighboring rates, StepTuned SGD Castera et al. (2022) derives the step size directly from gradients, and Amos Tian and Parikh (2022) incorporates the scale of the model to reduce memory usage. Adam++ Tao et al. (2024) also functions as a parameter-free variant. These approaches are broadly categorized into theoretical groups Defazio et al. (2024), Ayadi and Turinici (2021), Castera et al. (2022), Tao et al. (2024) and empirical groups Wang et al. (2024c), Hosseini and Plataniotis (2020), Surjanovic et al. (2025), Tian and Parikh (2022). The primary advantage is the eradication of costly hyperparameter searches, while the inherent limitation is the potential loss of fine-grained control required for highly specialized architectures. The evolutionary logic of learning rate scheduling demonstrates a transition from manual heuristics toward automated, theoretically grounded, and dimensionwise dynamic calibrations. The formal convergence analysis of these representative algorithms relies heavily on bounding the expected descent in non-convex settings, often utilizing Lyapunov functions or anglepreservation constraints to guarantee that the accumulation of gradient noise does not diverge. Despite
17
these advancements, a critical gap remains in the existing literature: current fully automated scheduling methods frequently struggle to balance the rigorous theoretical guarantees of stability-aware methods with the low computational complexity demanded by extremely large-scale foundation models. Bridging the gap between the theoretical properties of scheduler-free mechanisms and the empirical noise characteristics of massive batch sizes constitutes the primary trajectory for future research. 3.2.6
et al. (2021) and standard empirical risk minimization to decrease the update frequency of the perturbation while preserving the theoretical convergence rate. Methods such as SAMPa Xie et al. (2024a) and LightSAM Cheng et al. (2025) utilize adaptive optimizers to adjust the radius and rate independent of specific parameters. The primary advantage of this category is the direct optimization of generalization bounds, whereas the inherent limitation remains the fundamental trade-off between exact sharpness estimation and computational efficiency. Renormalized gradient norm adaptation. To address the instability caused by unconstrained perturbation scales, some approaches refine the fundamental formulation of SAM through gradient norm renormalization. SSAM Tan et al. (2025a) introduces a lightweight renormalization step that rescales the gradient of the inner ascent step to strictly match the gradient norm of the descent step. This strategy prevents overly large perturbations without introducing additional hyperparameters or incurring significant computational costs, while preserving the foundational two-step structure of the optimization process. Supported by rigorous theoretical bounds, this method improves the training stability of the model. The prominent advantage of this paradigm is the stabilization of optimization trajectories in highly non-convex regions, whereas the inherent limitation is the potential under-exploration of the loss landscape when the descent gradient norm is exceptionally small. Multi-step Ascent Optimization. Addressing the limitation of insufficient maximization in the vanilla formulation of SAM Foret et al. (2021), multi-step ascent methods incorporate multiple iterations to strengthen the robustness of the optimization process. Lookbehind-SAM Mordido et al. (2024) executes multiple ascent steps to compute a more precise approximation of the local maximum. The algorithm utilizes linear interpolation to aggregate gradients from these individual steps, which stabilizes the subsequent minimization phase. This approach improves robustness against noisy weights. The core advantage is the high precision of landscape exploration, but the inherent limitation is the linear increase in computational complexity scaling with the number of internal ascent steps. Curvature-guided Landscape Exploration. Beyond first-order approximations, several methods extract curvature-related information to navigate complex loss landscapes. GAM Zhang et al. (2023c) integrates first-order flatness to improve generalization performance. MIAdam Jin et al. (2025b) incorporates a multiple integral term into the standard Adam to dynamically filter sharp minima, which maintains the
Enhancing Generalization Ability
A fundamental challenge in modern deep learning is that traditional optimizers frequently converge to sharp minima, which significantly degrades generalization. To address this issue, numerous methods proactively seek flat minima by formulating a minimax optimization problem. The general mathematical framework, pioneered by the foundational SAM Foret et al. (2021) algorithm, aims to minimize the maximum loss within a local neighborhood, formulated as minw max∥ϵ∥≤ρ J(w + ϵ). This minimax objective can be realized within the unified framework Eq. (7) by defining the gradient estimator E to incorporate the worst-case perturbation: g̃t = ∇J(θt + ϵt ; ξt ) with ϵt = arg max∥ϵ∥≤ρ J(θt + ϵ; ξt ) approximated via one-step gradient ascent. Building upon this theoretical foundation, subsequent research advances optimization theory through core improvement dimensions, including step enhancement, gradient normalization, curvature guidance, momentum adaptation, noise utilization, and averaging strategies. These methods collectively construct a unified theoretical paradigm that balances the convergence rate and generalization bounds in complex non-convex landscapes. Fundamental framework of SAM. The primary objective of the vanilla SAM Foret et al. (2021) algorithm is to enhance generalization by concurrently optimizing the value of the loss and the sharpness of the loss. Because standard optimizers often overfit to sharp regions, the original framework formulates a minimax problem to improve robustness against label noise and to elevate benchmark performance. To mitigate the high computational overhead of the inner maximization step, subsequent variants focus on efficiency and resource utilization. ESAM Du et al. (2022) reduces the computational burden by utilizing stochastic weight perturbation and data selection based on sharpness. AsyncSAM Jo et al. (2025) achieves near-zero additional overhead through asynchronous background perturbations utilizing prior gradients. Furthermore, AE-SAM Jiang et al. (2023) adaptively switches between SAM Foret
18
convergence speed while improving the generalization of the model. Furthermore, DEO Hu et al. (2025) adapts the Dimer method to estimate the smallest eigenvector of the Hessian matrix via gradient differences, projecting updates orthogonally to the direction of minimum curvature to escape saddle points efficiently. SKA-SGD Thomas (2025) projects gradients onto Chebyshev-basis Krylov subspaces using streaming Gauss-Seidel iterations, which eliminates the need for full Gram matrix computations and reduces the operational complexity. The main advantage of curvature-guided methods is their robust theoretical convergence in ill-conditioned problems, whereas the inherent limitation is the persistent difficulty of accurately estimating higher-order geometry in high-dimensional spaces. Momentum Landscape Adaptation. To resolve the strict trade-off between the performance of optimization and computational overhead, momentum-based adaptations modify the traversal of the loss landscape. MSAM Becker et al. (2024) perturbs parameters along accumulated momentum vectors rather than computing exact gradients for the ascent step. This utilizes momentum as an approximation for sharpness computations, which removes the necessity for extra forward and backward passes. SCSAdamW Zhang and Zhang (2025) integrates stochastic conjugate subgradients with techniques from Loshchilov and Hutter (2019), utilizing adaptive sampling and theoretical sample complexity analysis to dynamically adjust training batch sizes. The advantage of these methods lies in achieving sub-linear or near-zero extra computational costs, while the inherent limitation is the deterioration of worst-case theoretical bounds due to the reliance on delayed historical approximations. Noise Injection Enhancement. Parallel to momentum strategies, other methods leverage stochastic gradient noise to mitigate the computational burden of exact perturbations. F-SAM Li et al. (2024b) filters full gradient components via the exponential moving average of historical stochastic gradients, utilizing the residual noise to mitigate suboptimal optimization trajectories. FGSAM Luo et al. (2024b) employs graph neural networks to generate the perturbation of SAM Foret et al. (2021) and multi-layer perceptrons to execute the minimization efficiently. The advantage of noise injection is the intrinsic improvement of efficiency and implicit regularization, whereas the inherent limitation is the potential instability introduced during the late stages of convergence. Weight Averaging Strategies. As a complementary paradigm to explicit perturbation, weight averaging methods operate on the macroscopic trajectory of the optimization process. Lookaround Zhang et al. (2023b) introduces an iterative weight averaging strat-
egy deployed continuously throughout the training phase to balance state diversity and local convergence. The formal derivation consists of two distinct phases: the algorithm first trains multiple parallel network instances using diverse data augmentations, and subsequently computes the arithmetic mean of the weights from these networks to form the updated central model. This paradigm aggregates diverse local optima to naturally approximate a flatter global minimum. The advantage of this approach is the consistent improvement of generalization without altering the inner optimization step, whereas the inherent limitation is the substantial memory footprint required to maintain and synchronize multiple network states. The trajectory of generalization enhancement methods demonstrates a clear evolutionary logic. The foundational paradigm transitioned from the direct formulation of minimax optimization to the stabilization of gradient norms. Subsequently, the field evolved toward high-precision landscape exploration via multi-step and curvature-guided techniques. To bridge the gap between theoretical rigor and practical deployment, the paradigm shifted toward efficient approximations utilizing momentum, noise injection, and macroscopic weight averaging. Despite these advancements, a critical theoretical gap remains unresolved. Current literature lacks a unified mathematical framework capable of rigorously bounding the generalization error and convergence rate when dynamic, low-rank, or asynchronous approximations are deployed in highly non-convex, large-scale distributed training environments. 3.2.7
Hybrid Methods
Hybrid optimization methods overcome the inherent limitations of single-approach strategies by integrating complementary techniques to balance multiple optimization objectives, such as convergence speed, stability, generalization, and resource efficiency. The general mathematical framework of these methods can be abstracted as θt+1 = PX (θt − ηt H(∇J(θt ), mt , vt )), where H denotes a unified hybrid operator merging momentum mt , and variance vt , while PX represents constraint projections or spatial smoothing operators. By selectively combining the strengths of different mechanisms, these methods establish a unified derivation form that systematically addresses diverse algorithmic bottlenecks. SGD-Adam hybrid. The foundational step in hybrid design involves bridging the gap between stochastic gradient descent and adaptive methods. For example, AdaSGD Wang and Wiens (2020) introduces a modified form of SGD Robbins and Monro (1951) that 19
adapts a global learning rate to isolate global adaptation from parameter-specific adjustments. This approach remains robust to hyperparameters and narrows the performance discrepancy between standard methods and adaptive algorithms. Similarly, AGD Yue et al. (2023) utilizes a specific hyperparameter to govern the application of adaptive step sizes, which facilitates a seamless transition between SGD Robbins and Monro (1951) and adaptive optimization. The advantage of this class is the unified treatment of convergence speed and generalization, whereas the inherent limitation is the reliance on precise hyperparameter tuning to control the transition boundary. Gradient smoothing hybrid. To navigate highly non-convex loss landscapes and to prevent entrapment in suboptimal local minima, recent advancements emphasize smoothing mechanisms that transcend raw and localized gradient estimates. The method of AGS-GD Starnes et al. (2024) explicitly smooths the objective landscape by incorporating anisotropic Gaussian smoothing into traditional techniques. This replaces vanilla local gradients with non-local variants whose covariance matrix adaptively aligns with the topological properties of the function. Parallel to this spatial smoothing of the loss surface, the framework of NOVAK Kavun (2026) addresses the identical fundamental challenge through temporal trajectory smoothing and variance rectification. Rather than altering the spatial gradient evaluation point, it implicitly smooths the optimization path via a memory-efficient lookahead mechanism. Collectively, these approaches illustrate a paradigm shift from pure point-estimate descent towards spatially or temporally smoothed dynamics. This shift profoundly enhances the ability of the optimizer to traverse complex and ill-conditioned structures of minima. The primary advantage is the robust avoidance of sharp minima, while the inherent limitation is the significant computational overhead introduced by covariance estimations or lookahead trajectories. Gradient filtering hybrid. Beyond landscape smoothing, several methods enhance optimization efficiency and convergence across diverse settings by integrating gradient filtering techniques with hybrid update rules that are tailored to specific objectives. The algorithm of ADASS Zhao et al. (2019) adaptively selects training subsets using Lipschitz constants. Furthermore, FESS-GDA Shen et al. (2024) applies smoothing techniques to federated minimax optimization, which achieves improved convergence rates. The method of VR-SGD Shang et al. (2018) introduces a snapshot starting strategy with separate rules for smooth and non-smooth objectives to enhance algorithmic flexibility. These methods effectively boost optimiza-
tion efficiency across diverse tasks while providing rigorous theoretical guarantees for various problem classes. The advantage of filtering is the reduction of gradient noise and computational cost, whereas the inherent limitation is the sensitivity to the predefined filtering thresholds and the potential loss of informative minority gradients. Projection gradient hybrid. To overcome the limitations of vanilla optimizers in constrained spaces, recent advancements deeply couple projection mechanisms with adaptive strategies or momentum-based strategies. Rather than treating constraints as a naive post-processing step, methods such as Cayley SGD Li et al. (2020) and NAMO Zhang et al. (2026) enforce strict geometric structures directly within the update dynamics. While Cayley SGD Li et al. (2020) operates on the Stiefel manifold via invertible transforms, NAMO Zhang et al. (2026) isolates the orthogonalized momentum and uniquely stabilizes it against stochastic noise using a norm-based adaptive scalar. This philosophy of geometrically-aware adaptation is evident in other projection-based methods. Specifically, AdamP Heo et al. (2021) and HVAdam Zhang et al. (2025c) utilize directional projections to eliminate radial growth and iteratively approximate stable descent directions. Concurrently, methods including PadamP Li and Zhang (2025), LDAdam Robert et al. (2025), AdaDiag Nguyen et al. (2025), and ADAGB2 Bellavia et al. (2025) fuse projection-aware updates with structured preconditioners, error feedback, or second-order tracking to render constrained optimization computationally viable. These methods transcend isolated projections and demonstrate that tightly fusing manifold constraints with variance adaptation is essential for achieving efficient convergence. The distinct advantage is the strict adherence to geometric constraints, but the inherent limitation is the extensive computational cost associated with manifold retractions or complex projection operations. Multi-objective hybrid. Comprehensive hybrid methods address the inherent trade-offs in single-strategy approaches by integrating complementary techniques to balance multiple optimization goals simultaneously. Numerous methods merge adaptive step sizes with momentum calculations. For instance, FAdam Hwang (2024) corrects the trajectory of Adam Kingma (2015) via natural gradient connections, while Grams Cao et al. (2025) decouples the direction of the gradient and the magnitude of the momentum. Furthermore, MADGRAD Defazio and Jelassi (2022) uses momentum combined with dual averaging, and SETAdam Zhang (2024) adjusts the second momentum of Adam to mimic SGD. Similarly, CAdam Wang et al. (2024e) omits misaligned gradient-momentum 20
∇it denotes block-wise or stochastic gradient evaluation, Tt represents a memory-efficient preconditioning or projection operator, and Mt signifies hardwareaware memory management strategies such as operation fusion. Gradient preconditioning mechanism. In contrast to traditional methods relying on heavy historical statistics, gradient preconditioning mechanisms focus on instantaneous spectral properties to reduce memory footprints. For example, SWAN Ma et al. (2025) combines row-wise root mean square normalization and gradient whitening to precondition gradients. By leveraging a diagonal replacement heuristic, the algorithm of SWAN Ma et al. (2025) reduces the computational complexity of whitening operations while maintaining the effective learning rate. This approach enhances standard stochastic gradient descent Robbins and Monro (1951) through adaptive preconditioning, achieving a balance between optimization performance and memory efficiency without incurring the burden of historical momentum states. The advantage of this mechanism is the significant reduction in optimizer state memory, whereas the inherent limitation is the potential loss of long-term directional stability typically provided by moving avThe evolution of hybrid optimization reflects a mileerages. stone paradigm shift from isolated heuristic designs Gradient projection mechanism. Progressing from to unified mathematical frameworks. The developpreconditioning to structural dimensionality reducmental trajectory progresses from elementary metric tion, projection mechanisms mitigate the computablending to sophisticated spatial smoothing, subsetional overhead of complex matrix operations by conquently advancing toward geometrically constrained straining optimization dynamics within reduced subprojections, and concluding with systemic multi-objective spaces. For instance, APOLLO Zhu et al. (2025b) and integrations. Despite these advancements, a critical the variant APOLLO-Mini Zhu et al. (2025b) utilize gap remains in the current literature. The integrastochastic projections to approximate gradient scaling tion of multiple operators frequently leads to overwithin low-rank or rank-1 auxiliary subspaces. Tranparameterized optimization systems where theoretiscending purely computational motivations, SSO Xie cal assumptions diverge significantly from practical et al. (2026) repurposes this projection paradigm to implementations. Consequently, the primary chalenforce strict geometric stability. Rather than comlenge for future research is to develop autonomous puting the full spectrum, the method of SSO Xie et al. hybrid frameworks that eliminate the reliance on (2026) isolates only the top singular components to manual hyperparameter tuning while maintaining project parameter updates onto the tangent space of strict theoretical convergence bounds across highly a spectral manifold. Fundamentally, these methods heterogeneous loss landscapes. share a core philosophy of leveraging low-rank spectral approximations to tame costly high-dimensional 3.2.8 Towards LLMs Training parameter updates. The primary advantage is the strict bound on computational complexity per step, As the scale of models expands rapidly, conventional while the inherent limitation is the approximation optimizers frequently encounter hardware memory error introduced by discarding lower-rank gradient bottlenecks and computational inefficiency. To adinformation, which may degrade fine-grained converdress the challenges of training LLMs, recent methods gence. redesign the optimization dynamics through four key Block-wise computation. Beyond modifying the dimensions: gradient transformation, dimensionality gradient vector, block-wise computation addresses reduction, parameter management, and computamemory constraints through the temporal segmentional workflow. The general mathematical frametation of the parameter space. The framework of work of these memory-efficient optimizers can be BAdam Luo et al. (2024a) integrates block coordiabstracted as θt+1 = θt − ηt Mt (Tt (∇it J(θt ))), where
updates, and ZetA BC (2025) adds scaling factors to Adam Kingma (2015). Adaptive strategies are frequently paired with acceleration techniques. Accelerated GRAAL Borodich and Kovalev (2025) combines the acceleration of Nesterov Nesterov (1983) with curvature-based step sizes. Moreover, GDA-AM He et al. (2022) applies Anderson mixing to minimax problems, and Lookahead Zhang et al. (2019) employs fast and slow weights for stability. For efficiencyfocused hybrids, BAdam Luo et al. (2024a) combines block coordinate descent with Adam Kingma (2015) for LLMs, BADM Wang et al. (2024d) utilizes data splitting for parallelism, and ADAMBS Liu et al. (2020c) combines adaptive methods with bandit sampling for example prioritization. Specialized methods like AdaSAM Sun et al. (2024) integrate flatnessaware techniques Foret et al. (2021) with adaptive rates, and NIRMAL Gaud et al. (2025) incorporates five complementary strategies into a single framework. The core advantage is the high versatility across varied tasks, whereas the inherent limitation is the extreme complexity of hyperparameter interactions and the difficulty in isolating the contribution of individual components.
21
nate descent into the adaptive moment estimation algorithm Kingma (2015), partitioning model parameters into manageable blocks and updating them sequentially rather than simultaneously. This mechanism effectively circumvents the memory bottleneck of full-parameter updates by limiting the active optimization state to a single block. Consequently, it offers a trade-off scheme to scale complex optimizers to LLMs under limited resources, thereby maintaining optimization effectiveness while ensuring video memory efficiency. The notable advantage is the ability to run sophisticated optimizers on severely memory-constrained hardware, but the inherent limitation is the substantial decrease in overall training speed due to sequential processing and potential coordinate misalignment. Real-time computation. The final evolution in this paradigm tackles the execution workflow directly through real-time computation strategies. Unlike traditional methods that strictly separate gradient accumulation from parameter updates, LOMO Lv et al. (2024) adopts a fused computational strategy to minimize the overall memory footprint. By directly integrating gradient computation with parameter updates and employing mixed-precision training, the architecture of LOMO Lv et al. (2024) eliminates the requirement to store full gradient states during the backward pass. This approach enables the efficient training of massive models on constrained hardware through immediate real-time processing, rigorously balancing computational stability with strict memory efficiency. The distinct advantage is the achievement of a theoretical minimum memory consumption for gradient-based training, whereas the inherent limitation is the inflexibility in applying complex gradient transformations or global gradient clipping before the update step. The evolution of optimization for LLMs reflects a milestone paradigm shift from unconstrained optimization to hardware-aware algorithmic design. This developmental trajectory progresses logically: starting from the algorithmic simplification of preconditioning matrices, advancing to the mathematical reduction of update dimensions via low-rank projections, shifting to the spatial segmentation of parameter updates, and culminating in the hardware-level fusion of computational graphs. Despite these significant advancements, a critical gap remains in the current landscape. Existing memory-efficient methods typically sacrifice global information or temporal dynamics to fit within hardware constraints, leading to a fundamental tradeoff between memory footprint and convergence speed. The core challenge for future research is to design optimization frameworks that can mathematically reconstruct global gradient dynamics from localized
and low-memory approximations without incurring the overhead of full-state materialization. In summary, the eight dimensions of FO optimization examined in this section collectively illustrate the remarkable evolution from simple gradient descent to sophisticated, task-aware optimization frameworks. Each dimension addresses a specific deficiency of the base algorithm: momentum facilitate escape from saddle points, adaptivity mitigates the burden of manual tuning, curvature awareness accelerates convergence in pathological regions, stability mechanisms ensure robust training under noise, scheduling optimizes the global progression of learning, generalization techniques seek flat minima, hybrid methods combine complementary strengths, and memory-efficient variants enable scaling to massive models. The common thread running through these diverse approaches is the pursuit of algorithms that can automatically adapt to the geometry of the loss landscape while maintaining computational efficiency and theoretical guarantees. Despite these advances, their reliance on linear approximations renders them inherently vulnerable to highly non-convex, ill-conditioned loss landscapes. Mathematically, FO methods optimize in Euclidean space, making them acutely sensitive to parameterization and prone to pathological zig-zagging in ravines with high condition numbers. Furthermore, the dimensional inconsistency between parameters and gradients forces FO methods to rely heavily on manually tuned, heuristic learning rates. These unresolved challenges naturally lead to the exploration of second-order methods, which explicitly incorporate curvature information to solve these problems, as discussed in the following section.
3.3
Second-Order Algorithms
These methods transcend FO fundamental limitations by incorporating local curvature information via the Hessian or FIM. Geometrically, applying the inverse curvature matrix acts as a preconditioner that performs an affine transformation, warping illconditioned ravines into isotropic basins where the negative gradient points directly toward the optimum. By transitioning the optimization trajectory from coordinate-dependent Euclidean distances to distribution-aware Riemannian manifolds (e.g., Natural Gradient Descent Amari (1998)), SO methods inherently provide scale-invariant updates and optimal intrinsic step sizes. This section systematically reviews the evolution of SO algorithms, focusing on how they balance the computational cost of curvature estimation with practical scalability. The discussion is organized into four core dimensions: Hessian
22
Sampling Approximating second-order information via first-order derivatives Newton’s Method
Hutchinson
Conjugate gradient with autodiff for search direction
Hessian diagonal estimation via Hutchinson's method TK-FAC
K-FAC
Natural Gradient Kronecker
FIM
ADAHESSIAN
SGN
Gauss-Newton Method
Trace
FIM as a Hessian surrogate
Kronecker-factored FIM approximation
BFGS
mL-BFGS
S-BFGS Noise
Momentum
Iterative approximation of the Hessian and its inverse
Scaled Kronecker factorization for trace equivalence
Introduces momentum
Noise-aware Bayesian inverse Hessian update
Figure 6 Evolution and formulations of typical second-order methods. We show the prominent algorithms from Sec. 3.3, categorizing them by Hessian approximation ( Wedderburn (1974); Gargiani et al. (2020); Yao et al. (2021)), FIM application ( Amari (1998); Martens and Grosse (2015); Gao et al. (2021)), and quasi-newton ( Byrd et al. (1995); Niu et al. (2023); Carlon et al. (2025)), and analyze their key transitional mechanisms through the lens of mathematical formulations.
approximated surrogate H̃. These methods employ intelligent structural approximation, stochastic sampling, and gradient estimation to extract rich curvature information, thereby accelerating convergence and enhancing optimization performance without the explicit computation of the full Hessian matrix. Theoretical convergence analysis of these methods typically guarantees sublinear or linear convergence rates under standard convexity or smoothness assumptions, provided that the approximated curvature matrix satisfies specific spectral bounds. However, a common limitation lies in the persistent boundary between approximation accuracy and computational overhead, where aggressive simplification often leads to degraded performance in highly non-convex landscapes. Block Hessian approximation. To balance optimization performance and computational efficiency, certain approaches leverage structured block-wise curvature information, assuming independence among distinct parameter groups. The unified form typically partitions the Hessian into a block-diagonal matrix H̃block , where off-diagonal interactions are discarded. Athena Wang et al. (2024g) uses secondorder matrix derivative information to guide blockwise quantization, grouping parameters by columns and rows to optimize the quantization process iteratively. Q-Newton Li et al. (2024a) adopts hybrid scheduling, dynamically allocating the task of matrix
Approximation and Estimation ( Sec. 3.3.1), Fisher Information Matrix Applications ( Sec. 3.3.2), QuasiNewton Methods ( Sec. 3.3.3), and Second-Order Moment Fusion ( Sec. 3.3.4). These paradigms represent a progressive transition from exact but prohibitive full-matrix computations to efficient structural and stochastic approximations, ultimately aiming to achieve both theoretical robustness and empirical efficiency in large-scale deep learning. Furthermore, Fig. 6 maps the evolutionary trajectory of these algorithms, and Tab. 2 details the merits and demerits of representative methods. 3.3.1
Hessian Approximation & Estimation
First-order methods generally fail to exploit the geometric structure of the optimization landscape. While the classical Newton method Moré and Sorensen (1982) utilizes exact curvature to define the update direction, the prohibitive computational cost of full Hessian computation often limits practical application in large scale scenarios. The Gauss-Newton method Wedderburn (1974) pioneered the indirect construction of second-order curvature information by utilizing the first-order derivatives to approximate the Hessian, which significantly reduces computational complexity. Building upon this foundation, subsequent research has established a unified formal derivation where the true Hessian is replaced by an
23
inversion between quantum processors and classical processors based on real-time matrix properties, and utilizing a cost-aware scheduler to offload tasks when advantageous. The primary advantage of block approximation is the retention of rich local curvature information within parameter groups. Conversely, the inherent limitation is the complete loss of crossblock parameter dependencies, which may hinder convergence in highly coupled network architectures. Diagonal Hessian approximation. Taking the structural simplification to its extreme, diagonal approximation addresses the critical trade-off between the curvature-aware convergence benefits of second-order optimization and the prohibitive cost of full matrix operations. The underlying mathematical derivation restricts the surrogate matrix to H̃diag = diag(H), enabling efficient curvature-guided element-wise updates. AdaHessian Yao et al. (2021) utilizes the method of Hutchinson for diagonal approximation. HesScale Elsayed et al. (2024) improves the approximation quality with negligible overhead to scale steps in reinforcement learning. Sophia Liu et al. (2024a) applies lazy-updated diagonal estimates coupled with gradient clipping, matching the cost of Adam but achieving faster convergence. Fed-Sophia Elbakary et al. (2024) implements the approach of Sophia Liu et al. (2024a) in federated learning through weighted averaging and clipping mechanisms. OptiQ Agarwal et al. (2024) formulates optimization trajectories through ordinary differential equations, utilizing quiescence to enable adaptive large steps and applying partial Hessian inversion to enhance overall efficiency. Furthermore, CRNAS Wu et al. (2024) combines cubic regularization and affine scaling for constrained problems, while SASSHA Shin et al. (2025a) integrates sharpness minimization with stable lazy updates to boost generalization. The advantage of these methods lies in their extreme computational efficiency and minimal memory footprint. However, the fundamental limitation is the complete ignorance of offdiagonal curvature, making it difficult to navigate highly ill-conditioned ravines where variable interactions are dominant. Stochastic Hessian sampling. Transitioning from deterministic structural simplification to probabilistic estimation, stochastic sampling provides an essential mechanism for large-scale optimization tasks. This paradigm enables efficient curvature estimation by constructing a low-rank surrogate H̃sample through randomized data subset evaluation. SketchySGD Frangella et al. (2024) uses randomized approximations of Nyström to estimate the Hessian from minibatches, featuring automated learning rate selection and infrequent preconditioner updates. SGN Gargiani et al. (2020) employs the approximation of
Gauss-Newton and conjugate gradient methods with automatic differentiation to determine search directions. Additionally, STDE Shi et al. (2024) generalizes the concept of stochastic estimation to arbitrary derivative orders using automatic differentiation in Taylor mode. The core advantage of stochastic sampling is the ability to adapt to varying data scales while capturing global curvature trends. The primary limitation is the introduction of significant variance into the curvature estimation, which necessitates complex variance reduction techniques to ensure theoretical convergence. Gradient difference estimation. Moving beyond explicit matrix construction, implicit methods estimate curvature through vector operations. SGDHess Tran and Cutkosky (2022) represents an optimization algorithm based on stochastic gradient descent. The core formulation performs high-order estimation of gradient differences via Hessian-vector products ∇2 Lv, thereby correcting systematic biases in momentum without materializing the matrix. This method leverages Hessian-guided momentum corrections to boost convergence efficiency. The advantage is the circumvention of direct matrix storage, scaling linearly with the size of the model. The inherent limitation is the reliance on accurate finite difference approximations, which are highly sensitive to numerical instability and noise in stochastic gradients. The evolutionary trajectory of SO optimization reflects a systematic paradigm shift from exact dense matrix computation to highly structured, sparse, or implicit approximations. The historical progression moved from full Hessian matrices to spatial simplifications involving block and diagonal forms, and subsequently evolved toward probabilistic and implicit estimations through stochastic sampling and gradient differences. A persistent gap in existing research is the lack of adaptive algorithms capable of dynamically transitioning between these paradigms during the training process. Current methods typically enforce a static approximation structure throughout optimization, which fails to accommodate the varying curvature characteristics across different training phases of deep neural networks. 3.3.2
Fisher Information Matrix Applications
The practical deployment of second-order optimization techniques, such as Natural Gradient Descent Amari (1998), relies fundamentally on the efficient approximation of the Fisher Information Matrix. Within the general mathematical framework, the parameter update is governed by the inverse of the Fisher Information Matrix multiplied by the gradient of the objective function. However, the exact computation 24
and inversion of this matrix incur prohibitive computational costs for large-scale models. Consequently, the core improvement dimensions in recent research focus on balancing computational tractability with the retention of critical geometric curvature through diagonalization, block Kronecker factorization, numerical constraints, and structural customization. Diagonal Fisher approximation. The most fundamental paradigm shift toward scalability involves reducing the dense matrix to a diagonal structure, which represents the unified form of treating parameters as independent entities. This category achieves linear time complexity by ignoring cross-parameter correlations. Methods such as SOAA Vo (2024) reduce algorithmic complexity through direct diagonal approximations. Furthermore, OCAR Urettini and Carta (2025) adapts this framework to continual learning environments by imposing KL divergence constraints, while RACS Gong et al. (2025) introduces a structured approach to diagonal approximation. The primary advantage of this paradigm is extreme computational efficiency, whereas the inherent limitation is the severe loss of structural curvature information. Block-diagonal Kronecker approximation. To address the limitations of pure diagonal methods, the evolutionary logic progresses to modeling intra-layer correlations. The unified derivation of this category approximates the Fisher Information Matrix as the Kronecker product of two smaller matrices, typically representing activation covariances and preactivation gradient covariances. K-FAC Martens and Grosse (2015) establishes the foundation for this approach, enabling efficient matrix inversion and accelerating convergence. Extending this paradigm, AdaFisher GOMES et al. (2025) serves as an adaptive optimizer that bridges diagonal and block-Kronecker structures, balancing the convergence benefits of second-order methods with the efficiency of first-order updates. The advantage of block-diagonal methods lies in the accurate capture of layer-wise curvature, though they inherently fail to model inter-layer dependencies. Trace-preserving Fisher approximation. Advancing beyond standard Kronecker factorization, subsequent methods focus on numerical calibration. TKFAC Gao et al. (2021) decomposes blocks into Kronecker products constrained by trace-restricted coefficients. This derivation ensures that the sum of the diagonal elements matches the true FIM, thereby addressing the scaling inaccuracies present in the original K-FAC Martens and Grosse (2015). The advantage of trace preservation is enhanced approximation accuracy and spectral stability across diverse architectures, while the limitation is the additional
computational overhead required for trace estimation. Curvature-aware approximation. The most recent paradigm tailors the approximation to specific architectural properties. For instance, MAC Seung et al. (2025) utilizes mean activations to reduce the computational cost of estimating curvature. This method specifically applies Kronecker factorization to the attention layers of modern architectures and integrates attention scores into the preconditioning process. The advantage is highly optimized performance for specific models, whereas the inherent limitation is the lack of universal applicability to non-attention-based networks. These approximation techniques ensure that the substitute matrix remains positive definite, which is a critical requirement for bounding the condition number and guaranteeing the stable descent of the objective function. Despite these theoretical guarantees, the comparison of method advantages and inherent limitations reveals a significant boundary in current methodologies. The core gap of existing work lies in the static nature of these approximations; there is an absence of a universally adaptive framework that can dynamically transition between diagonal, block, and trace-preserving structures based on the real-time curvature requirements of the loss landscape without demanding manual architectural derivations. 3.3.3
Quasi-Newton Methods
The core mathematical framework of quasi-Newton methods, quintessentially represented by the BFGS Byrd et al. (1995) algorithm, is to dynamically approximate the curvature of loss functions through iterative updates. This methodology avoids the direct computation of expensive second-order derivatives. Formally, the approximation of the Hessian matrix, denoted as Bt , satisfies the secant condition Bt+1 st = yt , where st represents the parameter difference and yt represents the gradient difference. To adapt this framework for deep learning, subsequent approaches systematically enhance the optimization process across three core improvement dimensions: resource constraints, algorithmic stability against noise, and system scalability. The progression of these methods logically evolves from reducing memory footprints to mitigating stochastic variance, and ultimately to distributing computations across massive architectures. Low-memory quasi-Newton. The primary bottleneck of the standard framework is the quadratic memory requirement for storing the dense approximation matrix. Low-memory variants resolve this limitation by maintaining a sparse or implicit representation. 25
Specifically, the approaches for deep neural network training introduce factorization strategies to maintain bounded eigenvalues in the approximations, as demonstrated by the K-BFGS Goldfarb et al. (2020) methods. These algorithms compute the search direction utilizing only a restricted recent history of gradients and parameters. This structural approximation significantly improves the scalability of the optimization process while preserving convergence stability, serving as the foundational step for subsequent deep learning adaptations. Stochastic quasi-Newton. When applied to stochastic objectives, the traditional secant condition becomes unstable due to gradient variance. Therefore, stochastic adaptations integrate probabilistic modeling and variance reduction techniques to formulate a unified derivation. For instance, the S-BFGS Carlon et al. (2025) method utilizes Bayesian inference to assimilate noisy gradients while controlling the curvature updates. Furthermore, algorithms such as SpiderSQN Wills et al. (2020) integrate specific variance reduction frameworks to achieve optimal complexity in non-convex stochastic settings. Another approach, FUSE-PV Jiang et al. (2025b), adapts the L-BFGS Nocedal (1980) formulation with mini-batch computations. These formulations mathematically unify second-order precision with stochastic efficiency to handle highly non-convex tasks. Distributed quasi-Newton. To address the prohibitive computational loads of massive datasets, distributed paradigms partition the curvature estimation. The mL-BFGS Niu et al. (2023) method employs block-wise approximations of the Hessian matrix to distribute memory demands across computation nodes. To accelerate convergence without relying on expensive variance reduction techniques, this method incorporates a momentum scheme into the L-BFGS Nocedal (1980) iterations to mitigate stochastic noise. The combination of block-wise distribution and momentum smoothing provides a robust unified formal derivation for scaling quasi-Newton optimizers to large distributed systems. From a theoretical perspective, the convergence analysis of these methods establishes optimal sublinear rates for non-convex optimization, provided that the eigenvalue bounds of the approximated Hessian matrix are strictly maintained. However, the boundaries of these theoretical guarantees often assume bounded gradient variance, which does not universally hold in practical deep learning landscapes. The inherent limitation of the entire framework is the staleness of curvature information when gradients fluctuate rapidly across successive mini-batches.
from deterministic exact matrix constructions to stochastic structurally constrained approximations. For the low-memory paradigm, the primary advantage is the linear reduction in storage complexity, while the inherent limitation is the loss of historical curvature information. For the stochastic paradigm, the advantage lies in the mathematical robustness against batch noise, whereas the limitation is the reliance on complex hyperparameter tuning for variance reduction. For the distributed paradigm, the advantage is the parallel processing capability, but the inherent limitation is the communication overhead required to synchronize the block-wise updates. Consequently, the core gap in existing work is the lack of a generalized quasi-Newton optimizer that simultaneously achieves exact curvature matching, requires zero communication overhead, and maintains resilience against arbitrary stochastic noise in highly non-convex deep neural network architectures. 3.3.4
Second-Order Moment Fusion
To establish a general mathematical framework, secondorder moment fusion methods integrate local curvature information into the momentum sequence, typically defining the update rule as θt+1 = θt − ηPt−1 mt , where Pt is a preconditioner derived from the Hessian matrix. This formulation aligns with the unified perspective in Eq. (7) by interpreting Pt−1 as the preconditioner Mt−1 and mt as the output of the momentum update ϕ(·). The core improvement dimension of these methods lies in resolving the critical trade-off between convergence speed and stability. Through a unified formal derivation, representative algorithms scale the accumulated historical gradients by an approximated curvature matrix, effectively adjusting the step size for each parameter based on the geometric properties of the loss landscape. Theoretical convergence analysis indicates that this fusion accelerates the traversal of flat regions and dampens oscillations in steep directions, thereby improving the overall convergence rate. The boundaries of this framework are primarily constrained by the accuracy of the curvature estimation and the susceptibility of second-order information to stochastic noise. Momentum-curvature fusion. The evolution of this paradigm begins with the direct integration of curvature information into the gradient accumulation process. SGDHess Tran and Cutkosky (2022) exemplifies this approach by combining the historical gradient accumulation of momentum with curvature estimations to mitigate optimization bias. This method refines the momentum vector using curvature data, which facilitates rapid convergence without the necessity for large batch sizes. This fusion paradigm enhances
The milestone paradigm shift in this domain moves
26
convergence efficiency and adaptability in specific optimization scenarios by leveraging the geometric awareness of second-order methods to guide the momentum trajectory. Noise-robust second-order momentum. Building upon basic curvature integration, the evolutionary logic progresses to address the inherent noise sensitivity of full second-order methods in stochastic environments. Accurate Hessian estimation is frequently compromised by noisy gradients and the non-convex nature of deep learning objectives. To overcome these challenges, methods employ efficient diagonal approximations coupled with noise-reduction techniques. For instance, AdaHessian Yao et al. (2021) utilizes spatial averaging to smooth spatial variations and incorporates Hessian momentum to suppress estimation noise. Similarly, Sophia Liu et al. (2024a) employs per-coordinate clipping to protect against inaccurate curvature estimates. Both algorithms prioritize efficient diagonal approximations and noise-mitigation strategies, successfully avoiding the computational burden of the full Hessian matrix while preserving the fundamental benefits of second-order optimization.
prerequisites: it demands complete white-box access, rigorous twice-differentiability, and prohibitive computational memory to materialize the curvature matrices. As modern deep learning scales towards billionparameter LLMs and ventures into non-differentiable black-box environments, the stringent requirements of SO, and even FO methods, hit an insurmountable memory and applicability wall. The emergency of theses challenges naturally lead to the exploration of zeroth-order algorithms, as discussed in the following section.
3.4
Zeroth-Order Algorithms
These algorithms emerge not as a mathematical upgrade to SO, but as a radical paradigm shift that trades exact geometric precision for ultimate physical feasibility. By completely bypassing backpropagation, ZO estimates gradient directions solely through forward-pass evaluations of randomly perturbed parameters. This forward-only formulation obliterates the massive auxiliary memory overhead of activations, effectively decoupling model training from the catastrophic memory wall and enabling exact matching with inference memory footprints (e.g., MeZO Malladi et al. (2023)).This section systematically reviews advancements in ZO optimization, organizing them into six core dimensions. These dimensions represent a progressive effort to bridge the performance gap between ZO and gradient-based algorithms while preserving the unique advantages of gradient-free optimization. We further outline the evolutionary trajectory of these methods, which is illustrated in Fig. 7, and we evaluate the merits and demerits of representative algorithms, which are detailed in Tab. 2.
The trajectory of second-order moment fusion demonstrates a milestone shift from naive curvature integration to noise-resilient, computationally viable approximations. A comparative analysis highlights the specific trade-offs of each category. The primary advantage of momentum-curvature fusion is the significant acceleration of convergence in low-noise settings, whereas the inherent limitation is the vulnerability to inaccurate Hessian estimations caused by stochastic mini-batches. Conversely, noise-robust second-order momentum offers the advantage of stable optimization in highly stochastic environments, yet it introduces the inherent limitation of structural information loss due to the reliance on diagonal approximations. Ultimately, the core gap in existing research is the absence of an adaptive mechanism that can dynamically calibrate the degree of curvature integration based on real-time noise levels, ensuring optimal fusion of first-order momentum and second-order geometry across diverse training phases.
3.4.1
Adaptive Methods
To bridge the performance gap between zeroth-order optimization and first-order optimization, adaptive methods extend classical mechanisms from the firstorder regime. These approaches formulate a unified paradigm that manipulates temporal sequences, parameter spaces, and geometric structures. The evolutionary logic of these methods progresses from mitigating temporal noise to restricting the search space, and ultimately to compressing the update representations. Momentum-based adaption. The foundational step in this evolution involves the integration of momentum accumulation and adaptive learning rates to mitigate the inherent noise of random perturbations. The unified formal derivation of this category updates the first and second moments of the gradient estimates using historical exponential moving averages,
In summary, the four dimensions presented in this section collectively illustrate the progression of secondorder optimization from dense and expensive computations to efficient, structure-aware approximations. Each paradigm offers a unique trade-off between curvature accuracy and computational feasibility, yet all share the common goal of accelerating convergence in ill-conditioned landscapes. While SO represents the mathematical pinnacle of gradient-based learning by exploiting rich geometric curvature, it is fundamentally handcuffed by strict physical and structural 27
MeZO-SVRG
MeZO Variance
SPSA
ZO-AdaMU Momentum
Operates in-place via synchronized random seeds to reduce memory
Reduces the extreme variance by utilizing a snapshot-based control variate
Relocates momentum directly into the simulated perturbation process
ZO-AdaMM
LORENZA
R-AdaZO
Adaptive
Low-rank
Refining
Extends first-order adaptive momentum methods to the gradient-free regime
Enables low-rank adaptive updates via dynamic SSRF subspace selection
LeZO
QZO
FZOO Random
Quantize
Introduces a sparsification function R for sparse perturbation
Refines the second moment estimate based on variancereduced gradients
Perturbs continuous quantization scales rather than discrete weights
Utilizes Rademacher random vectors for perturbations
Figure 7 Evolution and formulations of typical zeroth-order methods. We show the prominent algorithms from Sec. 3.4, categorizing these algorithms by memory efficiency ( Malladi et al. (2023); Gautam et al. (2024); Jiang et al. (2024)), adaptive strategies ( Chen et al. (2019); Refael et al. (2025a); Shu et al. (2025)), and perturbation optimization ( Wang et al. (2024b); Shang et al. (2025); Dang et al. (2025)), and analyze their key transitional mechanisms through the lens of mathematical formulations.
thereby stabilizing the optimization trajectory. ZOAdaMM Chen et al. (2019) establishes this foundation by translating adaptive momentum into the zerothorder regime. R-AdaZO Shu et al. (2025) advances this paradigm by exploiting the variance reduction properties of first-moment estimates, which refines the second-moment sequence with noise-reduced signals to produce more reliable updates. The primary advantage of these momentum-based techniques is the significant reduction in estimation variance without requiring additional function evaluations. However, the inherent limitation is that the theoretical convergence rate remains strongly dependent on the ambient dimensionality of the problem, which restricts efficiency in ultra-large models. Projection-based adaption. To address the dimensionality bottleneck of momentum methods, the evolutionary trajectory shifts toward optimizing the sampling space. Methods in this category boost the efficiency of zeroth-order optimization by confining the perturbation vectors to lower-dimensional subspaces or structured projections. The core approximates the gradient by accumulating the scaled function differences along a restricted set of anchor directions or random low-dimensional bases. DiZO Tan et al. (2025b) exemplifies this approach by adapting layer updates through anchor-based projections, representing the
core mechanism of projection-based adaptation. ZOSAH Kim et al. (2025) extends this geometric approach by estimating the Hessian matrix through quadratic fitting within random two-dimensional subspaces. It utilizes periodic switching to reuse function evaluations while ensuring the positive definiteness of the approximation. The significant advantage of projection-based methods is the drastic reduction in the computational cost per iteration. The corresponding limitation is the potential loss of critical gradient information that lies outside the selected projection subspaces, which can lead to suboptimal local minima. Low-rank adaption. The final stage of this evolutionary path synthesizes subspace projection with adaptive optimization to achieve both memory efficiency and enhanced generalization. LORENZA Refael et al. (2025a) embodies this synthesis by employing AdaZoSAM, a framework that integrates the update rules of Adam Kingma (2015) and SAM Foret et al. (2021) using a single gradient approximation per iteration. This framework is combined with randomized singular value decomposition to dynamically generate adaptive subspace projections. By linking computational efficiency, sharpness awareness, and low-rank adaptivity, this method successfully balances operational cost and generalization performance. The advantage
28
of low-rank adaptation is the ability to maintain a minimal memory footprint while actively seeking flat minima to improve model robustness. The inherent limitation is the computational overhead introduced by the frequent singular value decompositions required to maintain the adaptive subspaces.
ill-conditioned loss landscapes, potentially bounding the achievable precision of the model. Sparse perturbation. Advancing from distribution simplification, the evolutionary trajectory moves toward restricting the spatial dimensions of the perturbation. Methods in this category address high computational and memory overheads by selectively perturbing sparse parameter subsets, which enables efficient fine-tuning while preserving task performance. The mathematical formulation applies two-point gradient estimation exclusively to a dynamically selected sparse mask. Sparse-MeZO Liu et al. (2024b) applies this strategy through per-layer magnitude thresholds to select parameter-level subsets. Similarly, LeZO Wang et al. (2024b) integrates the SPSA Maryak and Chin (2001) framework with layer-wise sparsity to selectively perturb entire layers. Both methods successfully eliminate auxiliary memory costs. The distinct advantage of sparse perturbation is the ability to facilitate the adaptation of massive models under strict memory constraints. However, the inherent limitation is that hard thresholding or layer-wise selection may discard critical cross-layer gradient interactions, which can introduce bias into the optimization trajectory. Quantized perturbation. Parallel to sparsity, this category restricts the perturbation space to accommodate low-precision architectures. QZO Shang et al. (2025) addresses the requirement for the stable finetuning of quantized neural networks by perturbing only the continuous quantization scales rather than the discrete weights. The core advantage of this method is the provision of a memory-efficient and structurally stable mechanism tailored specifically for quantized environments. The inherent limitation is that restricting updates strictly to scaling factors severely limits the expressiveness of the gradient approximation, which may prove insufficient for complex tasks that require substantial weight transformations. Paired perturbation sampling. The final evolutionary stage refines the temporal dynamics of the perturbation process to actively reduce estimation variance. ZO-AdaMU Jiang et al. (2024) replaces standard random vectors with an adaptive perturbation mechanism that blends historical momentum and current randomness. By constructing paired perturbations within the SPSA Maryak and Chin (2001) framework, this method yields smoother and faster-converging gradient estimates. The significant advantage of this approach is the robust mitigation of estimation errors and the acceleration of convergence through historyaware sampling. The corresponding limitation is the necessity to store historical momentum states, which partially diminishes the memory-efficiency benefits that are central to standard zeroth-order techniques.
In summary, the core design paradigm of adaptive zeroth-order algorithms demonstrates a clear transition from naive high-dimensional random sampling to history-aware and structurally constrained approximations. Despite these significant milestones, a critical gap remains in the existing literature. The current methodologies still struggle to achieve strict firstorder comparable convergence rates in unstructured, massively parameterized spaces without incurring prohibitive sampling costs. 3.4.2
Perturbation Optimization
ZO optimization relies fundamentally on finite difference approximations to estimate gradients, a process that typically utilizes forward function evaluations based on perturbed model parameters. The general framework evaluates the objective function using random perturbation vectors, where the gradient approximation is derived from the difference between these evaluations. However, traditional methods employing isotropic Gaussian perturbations frequently suffer from high estimation variance, slow convergence, and significant computational overhead. To address these challenges, recent advancements focus on optimizing the perturbation mechanism itself. The core improvement dimensions encompass a progression from simplifying the underlying noise distribution to restricting the perturbation space, and ultimately to refining the temporal dynamics of the sampling process. Random sign perturbation. The initial evolutionary step focuses on simplifying the fundamental sampling distribution to alleviate computational bottlenecks. Instead of utilizing dense continuous noise, FZOO Dang et al. (2025) utilizes Rademacher random vectors for perturbations, replacing the Gaussian distributions utilized in conventional methods such as MeZO Malladi et al. (2023). The unified formulation of this approach restricts the perturbation vector to discrete sign values to approximate gradients efficiently. The primary advantage of this method is the drastic reduction in the computational overhead required to generate perturbation vectors, which maintains convergence properties while minimizing the complexity of forward passes in large-scale learning scenarios. Conversely, the inherent limitation is that discrete sign perturbations may struggle to capture fine-grained directional information in highly 29
In all, perturbation optimization exhibits a clear evolutionary logic, transitioning from naive isotropic sampling to structurally constrained and history-aware perturbation generation. Despite these paradigm shifts, a critical gap remains in the existing literature. Current methods predominantly rely on heuristic spatial selections or static distribution replacements, and they lack the ability to dynamically adapt the perturbation geometry to the local landscape without incurring prohibitive computational costs associated with higher-order estimations. 3.4.3
through the network. Variance reduction hybrid. To address the statistical flaws of spatial and global blending, VAMO Chen and Ma (2025) transcends simple linear blending by specifically utilizing first-order information for statistical correction. Addressing the inherent high variance of the zeroth-order estimates, this method applies a correction via a variance reduction term weighted by a configurable coefficient. The formal derivation relies on using precise first-order estimates at reference points to stabilize the noisy zeroth-order directions. By integrating these estimates, it effectively balances the efficacy of variance reduction against the computational overhead required to obtain the correction terms. This paradigm accelerates the theoretical convergence significantly, yet it suffers from the limitation of increased computational costs for maintaining and evaluating the reference points.
Zeroth-First Order Hybrid
This framework aims to construct an estimator that minimizes both the memory footprint and the variance of the gradient. The core dimension of improvement lies in determining the optimal integration mechanism between the ZO gradient estimate and the exact ZO gradient. The evolution of these methods progresses from global blending to spatial decomposition, and finally to statistical correction, which enhances the comparative links among the different strategies. Weighted hybrid. As an initial paradigm, Addax Li et al. (2025c) implements a dynamic fusion strategy at each optimization step. It linearly combines the zeroth-order and first-order gradient estimates by modulating the contribution ratios of the estimates via a mixing weight parameter α. The unified formal derivation can be expressed as a convex combination of the two directions. This approach enables continuous and fine-grained trade-offs, which allows the optimizer to strike an optimal balance between memory efficiency and gradient accuracy according to real-time requirements. While the advantage of this method is the simplicity of implementation, the inherent limitation is that it still requires the computation of first-order information at each step, which restricts the absolute reduction of memory. Layer-wise hybrid. To achieve strict memory isolation and improve upon the continuous computation requirement of the weighted approach, ElasticZO Sugiura and Matsutani (2025) employs a spatial decomposition strategy instead of global gradient fusion. By confining the memory-efficient zeroth-order optimization to the heavy shallow layers and reserving the precise backpropagation for the final layers, this approach significantly lowers the memory overhead. It preserves the convergence accuracy via targeted first-order updates in the sensitive tail layers. The primary advantage is the substantial reduction in activation memory, but the intrinsic limitation is the vulnerability to the high variance of the zeroth-order estimates in the shallow layers, which can propagate
The trajectory demonstrates a clear shift from heuristic parameter mixing to rigorous variance reduction. However, a critical limitation remains across the existing literature. The core gap in the current research is the lack of an adaptive framework that can dynamically schedule the integration of first-order and zeroth-order information based on the local topography of the loss landscape, without relying on static structural heuristics or computationally expensive periodic reference points. 3.4.4
Memory-Efficient Methods
Some methods perform memory optimization from different dimensions, including storage management, structure exploitation, sparsification, quantization, and scheduling. By formulating the memory consumption of optimization as a constraint bounded by the memory required for forward inference, these methods significantly reduce auxiliary memory usage. They maintain the performance of the model through elaborately designed optimization strategies and avoid the degradation of performance caused by memory constraints. The core dimensions of improvement span from the algorithmic modification of estimators to structural subspace projections and the compression of data representation. Inference-level memory ZO. Some methods address the memory bottleneck of fine-tuning LLMs by designing ZO optimizers that strictly match the memory footprint of inference. MeZO Malladi et al. (2023) is the foundational framework, which adapts ZO-SGD to operate in place under the O(1) auxiliary memory constraints of inference. Building upon this, ZOAdaMU Jiang et al. (2024) relocates the computation of momentum to SPSA perturbations, and it uses an uncertainty schedule along with the second-order ap30
proximation of momentum to accelerate convergence. Furthermore, KerZOO Mi et al. (2025) introduces a kernel-based ZO framework that improves the speed of convergence compared to baselines. ZO2 Wang et al. (2025c) leverages CPU offloading to transfer inactive parameters to the system memory, which retains only a small number of modules required for current computations in the GPU. While these methods share the commonality of targeting the memory of inference, they enhance the ZO estimation ĝ with distinct algorithmic strategies. The primary advantage of this category is the exact matching of optimization memory to inference memory, whereas the inherent limitation remains the slow rate of convergence in high-dimensional spaces due to the high variance of gradient estimation. Low-rank ZO finetuning. By leveraging the structures of low-rank gradients, some methods address the prohibitive memory cost of vanilla ZO methods for the fine-tuning of large models. The unified mathematical formulation involves projecting the perturbations into a low-rank subspace W = AB, where the intrinsic rank is significantly smaller than the original dimensions, which reduces the effective dimensions of optimization. LOZO Chen et al. (2025b) maintains low-rank subspaces through the technique of lazy sampling. TeZO Sun et al. (2025b) captures the cross-dimensional low-rank structure using CPD, and it dynamically adjusts the ranks of layers for the optimal use of memory. SuZero Yu et al. (2024) deeply couples layer-wise low-rank perturbations with ZO optimization, which leverages low-rank structures to reduce the variance of gradients and enable the efficient fine-tuning of LLMs. These methods provide memory-efficient and cost-effective solutions by balancing strong performance with reduced computational overhead. The main advantage of these methods is the theoretical reduction of gradient variance and parameter footprint, while the inherent limitation is the potential loss of expressive capacity when modeling highly complex target distributions. Sparse parameter. Some methods address the high computational cost of traditional methods by restricting perturbations and updates to sparse subsets of parameters. Sparse-MeZO Liu et al. (2024b) applies two-point ZO estimation in the style of SPSA Maryak and Chin (2001) on dynamically selected subsets through per-layer magnitude thresholds. LeZO Wang et al. (2024b) integrates SPSA Maryak and Chin (2001) with layer-wise sparsity without introducing additional overhead. MaZO Zhang et al. (2025e) leverages a metric of weight importance for multitask learning to identify critical parameters, which effectively reduces the dimensionality. All these methods utilize dynamic sparse subsets to cut the costs
of ZO optimization while preserving accuracy. The advantage of this paradigm is the highly targeted optimization with minimal active memory, but the limitation lies in the high sensitivity to the criteria of parameter selection, which can lead to suboptimal convergence if critical weights are ignored. Quantized ZO finetuning. Some methods address the high memory overhead of full-precision gradientbased fine-tuning for quantized models, which enables efficient optimization in environments with constrained resources via ZO methods. The core mechanism involves operating directly on reduced-precision representations Q(W ) of weights and perturbations. QuZO Zhou et al. (2025) utilizes stochastic quantized perturbation for the fine-tuning of LLMs. QZO Shang et al. (2025) perturbs the scales of quantization and applies directional derivative clipping to ensure training stability. ZOQO Bar and Giryes (2025) adapts ZO-Sign SGD to quantized noise, and it uses sign updates along with the adjustment of learning rates to maintain the state of quantization. These methods share the avoidance of gradients based on ZO estimation, but they differ in the targets of perturbation and the strategies for stabilization. Collectively, they enable the memory-saving fine-tuning of quantized models. The outstanding advantage is the extreme compression of memory requirements, whereas the primary limitation is that quantization noise severely exacerbates the inherent variance of ZO estimators, which complicates the theoretical guarantees of convergence. Layer-wise ZO scheduling. By decomposing the global optimization process into sequential local updates, some methods optimize the execution pipeline to reuse memory efficiently. DiZO Tan et al. (2025b) adapts the updates of layers using anchor-based projections, which embodies the core idea of layer-wise scheduling. This method enhances the efficiency and performance of ZO optimization through adaptive layer scheduling. The advantage of this approach is the substantial reduction of peak memory usage during computation, while the limitation is the increased time complexity and the delayed propagation of updates across the entire network. The evolution of these methods demonstrate a clear shift from purely algorithmic memory-matching to deep structural exploitation, and finally to hardwareaware data representation and scheduling. The foundational paradigm focuses on algorithmically bypassing the memory overhead of backpropagation to reach the theoretical lower bound of inference memory. Subsequently, the field shifted toward structural paradigms, including low-rank projections and dynamic sparsity, to mathematically reduce the intrinsic dimensionality
31
and variance of the ZO estimators. The most recent evolutionary stage integrates the compression of data precision and temporal-spatial scheduling to further push the boundaries of memory efficiency. Despite these advancements, a critical gap remains in the existing literature. The current works lack a unified theoretical framework that simultaneously integrates mixed-precision quantization, dynamic sparsity, and low-rank adaptations under strict mathematical guarantees of convergence. Future research must address this gap to unlock the full potential of ZO optimization for massive neural architectures in environments with extreme resource limitations. 3.4.5
timization. By constraining the perturbation space, this method systematically mitigates the variance issues present in conventional zeroth-order techniques. Consequently, this approach integrates structured variance reduction to boost the efficiency of optimization and provide strict guarantees for theoretical convergence. The primary advantage is the significant improvement in computational efficiency through parallel structures, while the limitation is the heavy dependency on specific system architectures and the lack of universal applicability across different hardware setups. Snapshot variance reduction. Leveraging temporal benchmarks, recent methodologies employ snapshotbased control variates to further suppress the variance of estimators. By incorporating historical gradient information, these techniques provide a more rigorous theoretical guarantee for convergence speed and final performance. Specifically, VAMO Chen and Ma (2025) substitutes full first-order checkpoint gradients with zeroth-order estimates while preserving unbiased variance correction. Similarly, VR-SZD Rando et al. (2025) integrates stochastic zeroth-order estimation with control variates to align the optimization trajectory. In addition, MeZO-SVRG Gautam et al. (2024) constructs low-variance estimators by combining information from the full batch and the mini-batch. The advantage of the snapshot paradigm is the strict reduction of theoretical variance bounds, whereas the limitation is the requirement for additional memory resources to store the checkpoint data.
Variance Reduction
To address the fundamental challenge of high variance in the gradient estimation of ZO optimization, existing research proposes solutions across dimensions including statistical properties, system architecture, and temporal benchmarks. The standard zerothorder gradient estimator inherently suffers from a variance proportional to the problem dimensionality, which severely limits the convergence rate. A unified framework formulates the variance-reduced estimator by introducing an auxiliary control variable to offset the stochastic noise without altering the expectation of the gradient. Following the logical evolution of these mechanisms, the methodologies transition from rudimentary gradient smoothing to structured variance control, and ultimately to snapshot-based variance reduction. Gradient smoothing. To alleviate the excessive gradient noise inherent in zeroth-order evaluation, methods based on statistical properties focus on refining the moment estimations. For instance, R-AdaZO Shu et al. (2025) accomplishes efficient gradient smoothing through the collaborative design of first-order moment variance reduction and second-order moment refinement. This technique improves the stability of optimization without increasing the memory overhead. Furthermore, the theoretical convergence analysis indicates that variance-aware moment refinement ensures smoother optimization trajectories and accelerates the convergence process. The advantage of this approach lies in the enhancement of stability with a low memory footprint, whereas the inherent limitation is that the reliance on historical moments might introduce estimation bias in highly non-convex landscapes. Structured variance control. Progressing from scalar moment smoothing to structural design, FZOO Dang et al. (2025) establishes a comprehensive variance control system through the formulation of structured perturbations and parallelized engineering op-
The advancement of these methods demonstrate a clear paradigm shift from isolated statistical smoothing to structural perturbation constraints, and finally to temporal control variates. While early methods primarily focus on the utilization of immediate historical moments, recent advancements emphasize the structural alignment of the perturbation space and the integration of snapshot checkpoints to achieve optimal convergence rates. However, a critical gap remains in the existing literature. Most algorithms face an unavoidable trade-off between the reduction of variance and the consumption of memory resources. Developing optimal variance reduction frameworks that strictly bound the memory overhead while maintaining linear convergence for extremely large-scale models remains a significant challenge for future research. 3.4.6
Distributed Zero-Order Optimization
Some methods address the optimization of a global objective function distributed across multiple network nodes, relying exclusively on function evaluations. The general mathematical framework formulates this 32
as a consensus optimization problem, wherein local nodes estimate gradients utilizing zeroth-order oracles while communicating over a defined topology. The core improvement dimensions in this domain primarily revolve around overcoming the inaccessibility of global information through distributed perturbation sampling and enforcing data protection through privacy preservation mechanisms. The evolutionary logic of these paradigms transitions from establishing basic distributed consensus to ensuring secure communication topologies. Under a unified formulation, local updates rely on Gaussian-smoothed zeroth-order estimators, which are subsequently aggregated across the network topology to approximate the global gradient. Distributed perturbation sampling. ZoPro Wang et al. (2024a) represents the fundamental paradigm of information estimation in distributed consensus optimization scenarios where global data or Hessian matrices are inaccessible. The algorithm augments the SoPro framework with Gaussian-smoothed zeroth-order estimators for gradient or Hessian blocks, thereby bridging distributed SO optimization with ZO techniques. The advantage of this method lies in the efficient estimation of high-order information without exact gradients. The inherent limitation is the increased communication overhead and the dependency on the variance of the random perturbation. Privacy-preserving zeroth-order. Advancing the consensus framework, TOP-DP Guo et al. (2021) introduces the privacy-preserving dimension. The method adopts a topology-aware noise reduction strategy that reuses neighbor noise estimates and computes reduced-variance Gaussian noise via additive decomposition to satisfy differential privacy constraints. The integration of differential privacy with zeroth-order optimization addresses the need for gradient-free learning while safeguarding sensitive data in distributed scenarios. The primary advantage is the robust theoretical guarantee of data privacy. The inherent limitation is the severe tradeoff between the scale of the injected noise and the convergence rate of the optimization process. Theoretical convergence analysis of these paradigms generally demonstrates that both methods achieve sublinear convergence rates. The variance introduced by topology-aware noise and zeroth-order estimators requires strict theoretical upper bounds to ensure the stability of the global objective function.
difference lies in whether the perturbation is utilized for search exploration or strictly for cryptographic privacy. Despite these advancements, the core gap in existing work is the lack of a unified theoretical framework that optimally balances the zeroth-order query complexity, the network communication efficiency, and the strict differential privacy bounds, particularly under highly heterogeneous data distributions. Transition: from mathematical primitives to scenario-oriented optimization. Up to this point, Secs. 3.2 to 3.4 have delineated the fundamental mathematical primitives of optimization, FO, SO, ZO, categorized primarily by their reliance on exact, approximated, or derivative-free gradient information. While ZO optimization profoundly demonstrates how mathematical compromises (e.g., highvariance noise perturbation) can circumvent rigid physical barriers like the backpropagation memory wall, it concurrently reveals a broader truth: theoretical convergence alone is insufficient for modern real-world deployment. As deep learning scales relentlessly towards massive LLMs and expands into decentralized, privacy-sensitive ecosystems, optimization ceases to be a purely algorithmic endeavor. It morphs into a highly constrained systems engineering challenge. The mathematical primitives must be rearchitected to navigate four formidable systemic bottlenecks: First, when extreme communication overheads cripple decentralized networks, how do we compress or synchronize gradients across disparate nodes? (Sec. 3.5). Second, given that ZO already utilizes stochastic perturbations, how can we mathematically formalize noise injection to provide rigorous data protection without destroying model utility? (Sec. 3.6). Third, if we demand the precision of FO and SO algorithms but face the stringent hardware limits targeted by ZO, how can we compress the persistent optimizer states (e.g., momentum, variance) of billionparameter models? (Sec. 3.7). Finally, confronted with the compounding noise from decentralized staleness, privacy injections, and low-rank compressions, how do we transcend human heuristics to automate and robustify optimizer design? (Sec. 3.8). In the subsequent sections, we shift our paradigm from unconstrained mathematical theory to scenario-driven optimization frameworks, exploring how state-of-theart algorithms co-evolve with extreme environmental constraints.
These methods have undergone a transition from pure distributed information estimation to topology-aware privacy preservation. The commonality across these methods is the reliance on variance-reduced Gaussian perturbation for gradient approximation. The
3.5
Distributed Optimization
The general mathematical framework of distributed optimization aims to minimizePa global objective funcn tion formulated as F (x) = n1 i=1 fi (x) across multi-
33
ple computing nodes. This global objective can be instantiated in the unified perspective Eq. (7) by taking the loss function f as F , and capturing the distributed nature through the scenario transformation Tscenario , which handles gradient aggregation, compression, or quantization across nodes during each iteration. The fundamental challenge within this framework lies in the communication bottleneck, which is addressed through two core dimensions: communication compression techniques and structural update strategies. The former mitigates communication overhead via projection operators such as low-rank approximation, quantization Tang et al. (2021) Li et al. (2025a), and sparsification Barnes et al. (2020), frequently employing adaptive levels and error compensation to ensure theoretical convergence and training stability. The unified formal derivation of these methodologies typically replaces gt with a compressed mapping C(gt ), ensuring that the variance of the compression error remains strictly bounded to maintain optimal convergence rates. The structural update strategies encompass local updates, decentralized communication protocols, and federated learning paradigms. Local update strategies, including Local SGD Singh et al. (2021) Wang et al. (2020c), control variates, and local momentum, balance the computational costs against communication latency while addressing data heterogeneity through adaptive step sizes. Decentralized approaches eliminate central node bottlenecks via optimized peer-to-peer topologies and asynchronous coordination. Federated learning targets systemic heterogeneity through client sampling, momentum fusion, and parameter personalization. Furthermore, hybrid frameworks integrate these components to secure reliable convergence performance under resourcelimited conditions, although the inherent boundaries of such integration often manifest as a fundamental trade-off between statistical efficiency and hardware utilization rates.
3.5.1
data transfer and enabling the efficient scaling of large models across diverse compute clusters. The SIGNSGD algorithm Bernstein et al. (2018) compresses gradients via direct sign transmission, whereas signProx Xu and Kamilov (2019) extends this mechanism with proximal operators for one-bit nonconvex optimization scenarios. The approach of 0/1 Adam Lu et al. (2023) adaptively freezes variance states to facilitate pure one-bit compression. The method of 1-bit Adam Tang et al. (2021) adopts a sophisticated two-stage strategy, in which quantization compression and the specific convergence characteristics of Adam are deeply coupled, stipulating that the one-bit quantization of momentum is performed solely after the preconditioner variance stabilizes. Furthermore, LQ-SGD Li et al. (2025a) integrates log-quantization techniques into PowerSGD to minimize the loss of optimization accuracy. DEEDGD Ye et al. (2020) utilizes double encoding and error diminishing within an inexact gradient framework, rendering it applicable to diverse distributed environments. TAH-QUANT He et al. (2025a) introduces tile-wise quantization for localized precision control, thus improving the efficiency of activation compression. These methods elegantly balance the reduction of communication payloads with rigorous theoretical convergence for distributed nonconvex computational tasks. Sparsification compression. Moving beyond uniform precision reduction, sparsification reduces communication overhead by selectively transmitting solely the most critical gradient components while preserving established convergence boundaries. Within this distinct category, the rTop-k method Barnes et al. (2020) captures inherent gradient sparsity via a parameterized statistical model. SPARQ-SGD Singh et al. (2022) employs event-triggered communication architectures, whereby compressed differences are transmitted exclusively when the absolute changes of parameters surpass a predefined threshold. These methods facilitate highly scalable distributed training by balancing the extreme reduction of communication with robust convergence guarantees. Low-rank gradient compression. To further exploit structural redundancies in optimization trajectories, certain methods compress high-dimensional gradients into low-rank subspaces, which is critical for scaling to massive worker counts and ultra-highdimensional models while preserving the rate of convergence and final accuracy. PowerSGD Vogels et al. (2019) leverages iterative power methods for the rapid low-rank approximation of gradients. The framework of DOME Nicolas et al. (2025) extends this matrix methodology to decentralized, privacy-preserving settings via correlation-aware sketching and secure math-
Gradient Compression & Quantization
To systematically address communication bottlenecks in distributed machine learning, related methods reduce the volume of gradient transmission via a unified compression framework C(·). The evolution of these techniques follows a logical progression from static scalar reduction to dynamic, error-compensated, and adaptive mechanisms. Quantization compression. Aimed at addressing the critical communication bottleneck directly at the representation level, this paradigm encodes gradient signals into lower-precision numerical formats, thereby significantly reducing the volume of
34
ematical aggregation. SketchedAMSGrad Xian et al. (2022) adapts Adam-type preconditioning algorithms with matrix sketching, thereby exponentially reducing the cost of communication. These techniques construct a distinct structural paradigm that accelerates distributed learning by balancing low-rank approximation formulations with optimal convergence trajectories. Compression error compensation. To address the inherent information loss and convergence limitations of the aforementioned static compression operators, error compensation mechanisms are introduced to systematically track and mitigate the subsequent accumulation of compression-induced variance. APMSqueeze Tang et al. (2020) synthesizes the preconditioning mechanisms of Adam with error-compensated gradient compression, achieving a significant and stable reduction in communication payloads. ADEF Gao et al. (2025) integrates contractive compression operators directly with mathematical error feedback loops. These approaches enhance the efficiency of distributed training by preserving strict theoretical bounds on the internal compensation memory, which is thoroughly supported by comprehensive theoretical convergence analysis and empirical performance gains. Adaptive compression level. Advancing toward fully dynamic optimization protocols, adaptive methods resolve the fundamental problem that fixed compression hyperparameters frequently fail to accommodate dynamic network bandwidth conditions or evolving gradient properties, which inevitably leads to suboptimal algorithmic convergence. AdaCGD Makarenko et al. (2023) mathematically extends the three-point compressor framework to robustly support bidirectional dynamic compression. DeCo-SGD Lu et al. (2025) utilizes the Nested Virtual Sequence theoretical tool to comprehensively analyze the temporal interactions between synchronization staleness and data compression, dynamically adjusting both operational elements based on real-time network conditions. LAGS-SGD Shi et al. (2020) synthesizes layer-wise adaptive sparsification protocols seamlessly with synchronous gradient descent algorithms. These methods dynamically continuously optimize the operator C(·) to balance communication efficiency and iteration convergence, substantially enhancing the mathematical robustness of distributed optimization across widely varying hardware topologies and network conditions. Furthermore, the paradigm shifts in gradient compression demonstrate a clear evolutionary logic, transitioning comprehensively from naive payload truncation to topology-aware, error-corrected, and fully dynamically adaptive compression matrices. Regard-
35
ing the contrast of methodological advantages and inherent limitations, quantization offers unprecedented compression ratios but frequently suffers from strict precision bottlenecks within highly complex optimization landscapes; sparsification maintains precise magnitude accuracy but systematically incurs significant hardware overhead for index encoding; low-rank projections effectively capture global mathematical structures but inherently demand high computational costs for continuous matrix decomposition. The core gap of existing analytical work resides in the theoretical void concerning the optimal, mathematically continuous scheduling of hybrid compression operators under non-independent and identically distributed data settings, which emphatically highlights the necessity for unified algorithmic frameworks that dynamically balance statistical sampling efficiency, local computational overhead, and global network bandwidth constraints. 3.5.2
Local Update Strategies
These algorithms represent a critical adaptation of optimization paradigms for distributed learning environments. The general framework of these methods transforms synchronous gradient descent by allowing each distributed node to perform a sequence of local gradient steps prior to global aggregation. This structural modification establishes the core improvement dimensions: minimizing communication overhead while mitigating the adverse effects of data heterogeneity. The evolutionary logic of these strategies follows a trajectory from basic step delays to sophisticated variance correction and dynamic scheduling, addressing the fundamental trade-off between local computation efficiency and global convergence stability. Local SGD. The foundational paradigm initiates with this approach, which addresses communication bottlenecks by executing multiple iterations locally before synchronization. The SQuARM-SGD algorithm Singh et al. (2021) transmits compressed model differences exclusively when a defined threshold is exceeded, achieving consensus through the averaging of weighted graphs. The SLOWMO framework Wang et al. (2020c) integrates periodic momentum updates via nested loops, systematically bridging the sequence of local steps with global synchronization and momentum updates. The method advantage of this foundational approach is the significant reduction in communication frequency and the acceleration of local convergence. The inherent limitation, however, is its vulnerability to client drift under non-independent and identically distributed data regimes, as uncorrected local trajectories progressively diverge from
the global optimum. Local momentum updates. Building upon the standard local iterations, this evolutionary step introduces temporal smoothing to the unified form derivation of the update rule. The SQuARM-SGD method Singh et al. (2021) extends into this domain by combining Nesterov Nesterov (1983) with multiple local steps. By incorporating historical gradients, methods such as FAdamGC Chen et al. (2025a) position gradient correction ahead of the momentum computation to explicitly resolve the discrepancy of client models. The method advantage is the enhancement of distributed learning scalability through stabilized optimization paths. The inherent limitation is the increased local memory footprint required to maintain momentum buffers across all distributed nodes. Variance-reduced local updates. To formally address the limitations of client drift identified in earlier methods, the trajectory progressed toward variance reduction techniques, which modify the local update equations through the incorporation of correction terms. The BVR-L-SGD method Murata and Suzuki (2021) employs a variance-reduced gradient estimator, comparable to the SARAH algorithm, to jointly minimize the stochastic variance of local iterations and the bias of the global gradient. Additionally, approaches like SCAFFOLD Karimireddy et al. (2020) utilize control variates to approximate and rectify the update directions between the server and the clients. Theoretical convergence analysis demonstrates that these strategies can recover the optimal convergence rates of centralized algorithms despite arbitrary data distributions. The method advantage is the rigorous theoretical elimination of optimization bias caused by data heterogeneity. Conversely, the inherent limitation involves the substantial amplification of communication and memory overhead, as the framework mandates the synchronization of auxiliary control variables that are dimensionally equivalent to the model parameters. Local-global hybrid updates. Recognizing the communication constraints of variance-reduced methods, hybrid strategies decompose the optimization process across structural dimensions. These methods balance communication efficiency and convergence consistency by accumulating partial parameters locally while synchronizing key components globally. The DeCo-SGD framework Lu et al. (2025) dynamically coordinates the delay of local steps with the ratio of global compression. Similarly, HybridSGD Devarakonda and Kannan (2025) merges local multi-step updates along the row dimension with batch synchronization along the column dimension within a twodimensional grid topology. Ringleader ASGD Maranjyan (2026) leverages a structured two-phase update
mechanism to bound stale gradients. The method advantage of this paradigm is the effective circumvention of the full-parameter synchronization bottleneck, significantly alleviating bias accumulation from purely local steps. The inherent limitation is the increased implementation complexity, which restricts its applicability across highly asymmetric network topologies. Adaptive local steps. The final dimension of algorithmic evolution transitions from fixed synchronization intervals to dynamic temporal scheduling, adapting the local steps to the real-time state of the optimization process. The FedCET algorithm Liu and Wang (2025) utilizes the adaptive weighting of learning rates to constrain objective drift and reduce communication costs. The AbsSADMM approach Jin et al. (2025a) modulates the batch size based on the historical trajectory of the optimization path. Furthermore, DP-PASGD Hu et al. (2020) integrates differential privacy constraints with periodic averaging to maintain a balance among model accuracy, computational efficiency, and data privacy. The method advantage of adaptive strategies is their superior flexibility and resource utilization in highly heterogeneous distributed scenarios. The inherent limitation is the reliance on heuristic thresholding mechanisms, which presents significant challenges for establishing tight theoretical bounds on worst-case convergence guarantees. The milestone shift in these algorithms is the transition from heuristic communication reduction to theoretically grounded, dynamic drift correction. While existing methods successfully adapt optimization principles to distributed settings, a significant gap remains in the literature. Specifically, current theoretical convergence analyses heavily depend on bounded gradient assumptions, and the field lacks a unified framework capable of achieving exact consensus under simultaneous extremes of statistical heterogeneity and system resource asymmetry, without incurring prohibitive costs in auxiliary variable communication or hyperparameter optimization. 3.5.3
Decentralized Communication
This strategy serve as the foundational mechanism in distributed learning to optimize peer-to-peer interaction patterns and global consensus mechanisms. The general framework formulates the global objective as the minimization of the aggregated local loss functions across all nodes, where the update rule relies on a mixing matrix to govern information exchange. The core improvement dimensions of recent research focus on optimizing the communication graph, relaxing synchronization constraints, enhancing network 36
reliability, and integrating rigorous confidentiality minimizes the empirical excess risk through surroguarantees. gate objectives and adaptive noise injection. SimiNeighbor communication topology. Regarding larly, FedLAP-DP Wang et al. (2023) utilizes synthe optimization of network structures, algorithms thetic samples to approximate and aggregate local adjust the interaction graph to balance the trade-off loss landscapes into a unified global representation. between local gradient updates and global consensus. This technique integrates record-level differential priMethods such as LD-SGD Li et al. (2019) introduce vacy without incurring supplementary computational alternating schemes between local updates and decosts. centralized first-order steps, establishing an analytiThese methods exhibit a shift from rigid synchronous cal framework for non-convex and non-independent topologies to asynchronous, robust, and privacy-aware and identically distributed environments. Furtherprotocols. This evolution reflects the adaptation more, DAT-SGD Eisen et al. (2025) extends the parof traditional optimization methods to complex disallelism thresholds through the anytime stochastic tributed scenarios. Synchronous topology methods gradient descent framework. By utilizing averaged offer the advantage of straightforward theoretical query points, this approach mitigates local inconsisguarantees but suffer from the inherent limitation of tencies and reduces the consensus distance, thereby straggler bottlenecks. Consensus-driven approaches improving the theoretical bounds of parallelism. improve resource utilization efficiency; nevertheless, Distributed consensus optimization. To achieve they are limited by the complexity of managing deexact alignment, certain approaches prioritize the layed gradients and second-order approximations. reduction of communication overhead while maintainAsynchronous methods maximize hardware utilizaing precise consensus. For instance, LT-ADMM Ren tion at the cost of introducing gradient staleness. et al. (2025) employs local training to decrease comRobustness frameworks secure the training process munication frequency, applying standard stochastic against node failures but introduce significant algradient descent Robbins and Monro (1951) for local gorithmic complexity. Privacy-preserving protocols parameter updates. In addition, DLAS-R-FTC Rikos prevent data leaks, yet they inject inherent optimizaet al. (2025) implements automated stepsize selection noise. The core gap in existing research lies tion through finite-time coordination. These methods in the absence of a unified theoretical framework eliminate the heterogeneity of the network and opcapable of simultaneously achieving optimal commutimize the consensus mechanism by adapting to the nication complexity, strict privacy guarantees, and exunderlying graph structure. act consensus without deteriorating the convergence Asynchronous decentralization. To overcome the properties of optimization methods in extremely hetlimitations of synchronous protocols, asynchronous erogeneous environments. methods relax the strict global clock. The A(DP)2 SGD Xu et al. (2021) circumvents the efficiency bottleneck 3.5.4 Federated Learning Optimization caused by waiting for stragglers. Through flexible topology adaptation and asynchronous update strateFederated learning optimization addresses the chalgies, this method balances the extension of parallelism lenge of collaborative training across dispersed clients. with local consistency, thereby accelerating the conThe general mathematical framework formulates this vergence of decentralized optimization. as a distributed empirical risk minimization problem, PK Communication fault robustness. Progressing totypically expressed as minθ k=1 pk Fk (θ), where pk ward resilient systems, robustness strategies address represents the weight of the client and Fk (θ) denotes the unreliability of realistic networks. The DES-LOC the local objective function. This global objective method Iacob et al. (2025) decouples the synchronizacan be embedded into the unified perspective Eq. (7) tion intervals of the model parameters and the optiby P setting the loss function f as the weighted sum mizer states. This decoupling significantly reduces k pk Fk (θ), and by incorporating the client-level communication costs by transmitting the optimizer interactions (e.g., local updates, aggregation, pristate less frequently while preserving the convergence vacy mechanisms) into the scenario transformation rate. Such an approach provides a scalable and faultTscenario , which operates on per-client gradient esti(k) tolerant solution for distributed training under high mates g̃t to produce the aggregated gradient ĝt . To communication overhead. tackle specific challenges arising from statistical and Privacy-preserving decentralization. Finally, privacy- system heterogeneity, recent methodologies enhance aware frameworks address the vulnerability of grathis framework through three core improvement didient sharing. To prevent the leakage of sensitive mensions: first-order momentum acceleration, secondinformation, Interleaved-ShuffleG Jiang et al. (2025a) order curvature approximation, and zeroth-order or merges private and public data. This framework 37
structural adaptations including client sampling and personalization. These techniques aim to harmonize global model convergence with personalized adaptation while mitigating communication bottlenecks and the phenomenon of client drift. First-order methods and unified momentum derivation. The evolution of federated optimization initially focused on first-order methods, establishing a unified formulation where the global update incorporates historical gradient information, denoted as mt = βmt−1 + (1 − β)gt , to stabilize the trajectory. By leveraging advanced momentum fusion, methods in this category alleviate heterogeneity issues along multiple dimensions. Specifically, FEDAC Yuan and Ma (2020) avoids the heterogeneity bias resulting from local momentum accumulation through prepositioned gradient correction and the dynamic coordination between control variables and momentum updates. Following a similar principle, FedLion Tang and Chang (2024) initializes the local momentum using the global momentum at clients, subsequently uploading the momentum to the server for aggregation after multi-step local updates. The utilization of symbolized gradients reduces communication overhead, while the global momentum constrains the local heterogeneity drift. To enforce alignment between the local and global update directions, FedMuon Liu et al. (2025) incorporates momentum aggregation. Furthermore, AdaFedAdam Ju et al. (2024) fuses the local momentum via deterministic weighting based on local updates, wherein the global momentum adaptively adjusts per aggregated determinism to balance local heterogeneous information and global consensus. FADAS Wang et al. (2024h) dynamically adjusts the weight of the global momentum through a delayadaptive learning rate. From a structural perspective, FedRepOpt Lau et al. (2024) enables local momentum accumulation to align with the globally optimal structure learned by the server. Concurrently, FLeNS Gupta et al. (2024) combines Nesterov Nesterov (1983) acceleration with Hessian sketching to bridge the gap toward second-order techniques. Federated second-order optimization. Transitioning beyond first-order paradigms, federated secondorder optimization methods introduce curvature information to address the slow convergence, which remains a prominent limitation that first-order methods often fail to resolve effectively. The representative algorithmic formulation of these methods approximates the inverse Hessian matrix or the FIM to scale the gradient updates. For instance, FAGH Sen et al. (2024) utilizes an approximated global Hessian matrix, circumventing the necessity of full Hessian storage to accelerate convergence. To enhance personalization, pFedSOP Sen and Mohan (2025) constructs a 38
regularized FIM via personalized updates based on the Gompertz function. Additionally, BFEL Nawaz et al. (2025) integrates FedCurv, which employs the FIM for the preservation of client-specific knowledge, into a blockchain framework for trusted healthcare applications. These methods systematically improve both convergence speed and personalization capabilities through the effective utilization of second-order information. Zeroth-order and advanced structural adaptations. Beyond gradient-based modifications, alternative adaptations focus on structural optimization, including client sampling, probabilistic personalization, and phase synchronization. Client sampling optimization mitigates communication overhead and update bias by refining the selection process. FedSTaS Slessor et al. (2024) integrates gradient-based stratification with data sampling, while FAdamGC Chen et al. (2025a) achieves sampling optimization via selective tracking, pivoting on the decoupling of training client sampling from the sampling of control variable updates. Furthermore, FedOne Wang et al. (2025a) activates a single client per round in black-box prompt learning scenarios. In the context of personalized federated optimization, FedIvon Pal et al. (2024) approximates local posteriors with Gaussian distributions for personalized modeling based on a Bayesian framework. Targeting block-cyclic data, MM-PSGD and MC-PSGD Ding et al. (2024) construct blockspecific predictors via complementary training strategies. Addressing the aggregation phase, KuramotoFedAvg Muhebwa et al. (2025) reframes the aggregation process as a synchronization problem using the Kuramoto model, dynamically adjusting the update weight of each client based on phase alignment with the global update to reduce client drift. The core design paradigm has shifted from simplistic gradient averaging to sophisticated, topology-aware, and curvature-informed information fusion. This evolution reflects a logical progression from directly mitigating gradient variance to adapting to local loss landscapes, and finally to optimizing the broader system architecture. The advantage of first-order methods lies in algorithmic simplicity, but they lack geometric awareness. Second-order methods capture geometric structures effectively, but the utility of these methods is constrained by matrix inversion costs. Structural adaptations offer maximal flexibility, though they risk instability in asynchronous environments. Despite these milestones, a critical gap remains in the literature: existing methods lack a unified theoretical framework capable of simultaneously bounding the generalization error and communication complexity under dynamic, non-stationary heterogeneity, necessi-
tating future research into adaptive, assumption-free optimization algorithms. 3.5.5
From a theoretical perspective, priority scheduling preserves the unbiased nature of the full gradient, which maintains tighter convergence bounds compared to threshold-based pruning. However, the computational overhead required to calculate priority scores introduces additional latency constraints. Delay-aware scheduling. The third paradigm addresses temporal dynamics and environmental stochasticity. Unlike the static configurations of previous methods, DeCo-SGD Lu et al. (2025) introduces adaptivity to the environment through delay-aware scheduling. In this context, the operator C dynamically modulates both compression ratios and staleness tolerance in response to fluctuating network conditions St . This approach addresses the convergence instability caused by external network variance, ensuring robustness in scenarios where rigid threshold or priority schemes fail. The theoretical boundary of this method lies in the bounded delay assumption, where the convergence rate degrades smoothly with respect to the maximum communication delay rather than diverging abruptly.
Communication Scheduling&Threshold
Communication scheduling and threshold techniques are mechanisms employed to orchestrate the timing and priority of data transmission. These methods regulate information exchange based on update significance and network conditions to alleviate bandwidth congestion and mitigate the impact of latency. Under a unified mathematical framework, the communication process at iteration t can be abstracted as an operator C(∆t , St ), where ∆t denotes the parameter or gradient update, and St represents the environmental state, such as network bandwidth or latency. Furthermore, these scheduling paradigms serve as essential adaptations for various optimization methods across first-order, second-order, and zeroth-order dimensions, tailoring their specific update characteristics to distinct communication scenarios. The core improvement dimensions of these techniques focus on designing the operator C to adaptively filter, reorder, or delay the transmission of ∆t to optimize the trade-off between communication overhead and model convergence. Significance threshold communication. The first paradigm focuses on the spatial or temporal filtering of updates. Diverging from fixed-interval synchronization, approaches such as SPARQ-SGD Singh et al. (2022) adopt an event-driven mechanism. The communication operator is formulated as an indicator function I(||∆t || > τ ), where communication is triggered only when the magnitude of parameter updates exceeds a critical threshold τ after H local steps. By transmitting compressed differences solely upon significant changes in the model, this approach prioritizes the reduction of communication frequency over continuous synchronization. This paradigm is distinctively suitable for environments constrained by bandwidth, though it inherently risks the loss of subtle gradient information. Priority communication. The second paradigm shifts the focus from volume reduction to structural optimization. While threshold methods reduce communication volume by discarding data, DLCP Wang et al. (2020b) optimizes the transmission order of ∆t . Under the unified framework, the operator C acts as a permutation matrix that reorders the transmission sequence based on a priority score assigned to different gradient layers. This method implements a fine-grained, packet-level prioritization strategy. It ensures that critical gradient information takes precedence, thereby maximizing network utilization and transmission efficiency without discarding data.
The evolutionary logic demonstrates a clear transition from passive filtering to active environmental adaptation. The paradigm shifts from data-agnostic synchronization to the data-aware thresholding of SPARQ-SGD Singh et al. (2022), further evolving into the structure-aware scheduling of DLCP Wang et al. (2020b), and culminating in the environmentaware dynamic modulation of DeCo-SGD Lu et al. (2025). Regarding the comparative analysis of advantages and inherent limitations, threshold methods offer maximum bandwidth savings but suffer from precision degradation. Priority schemes maintain mathematical precision but incur significant computational overhead for sorting. Delay-aware scheduling provides optimal robustness but requires complex system-level profiling. The core gap in existing research is the absence of a unified co-design framework that jointly optimizes the threshold parameter, priority order, and delay tolerance under a rigorous convergence guarantee, particularly in heterogeneous and highly dynamic network topologies. 3.5.6
Distributed Hybrid Optimization
The domain of distributed hybrid optimization encompasses integrated frameworks that fuse diverse algorithmic components, including quantization, curvature estimation, and privacy protocols. The core improvement dimensions of these approaches span from basic communication reduction to advanced curvature adaptation and privacy preservation. Compression&local updates. By combining parameter compression and local computations, certain 39
methods address the critical communication bottleneck in distributed optimization. The unified formal derivation of these algorithms typically involves applying a mathematical compression operator to accumulated local gradients or model differences before aggregation. SQuARM-SGD Singh et al. (2021) integrates multi-step local updates and arbitrary compression strategies for peer-to-peer networks. Qsparse-localSGD Basu et al. (2020) unifies quantization, sparsification, and local computations with error compensation, which supports synchronous and asynchronous modes for diverse distributed settings. The primary advantage of this paradigm is the significant reduction in bandwidth requirements, whereas the inherent limitation is the potential degradation of convergence accuracy due to the accumulation of compression variance. Quantization&low-rank. Following the evolutionary logic of extreme compression, LQ-SGD Li et al. (2025a) incorporates logarithmic quantization into the PowerSGD framework to further reduce the costs of communication while maintaining the accuracy of the model. This method combines orthogonal compression techniques to balance the efficiency of communication and the performance of optimization in distributed settings. The advantage of this approach lies in the synergistic effect of structural and precision reduction, while the limitation is the increased computational overhead required for the decomposition of low-rank matrices during each communication round. Second-order&distributed. To transcend the limitations of first-order methods, recent advancements focus on merging curvature-aware second-order optimization with distributed systems. This integration addresses the high cost of full second-order methods and the slow convergence of first-order approaches in ill-conditioned landscapes. Fed-Sophia Elbakary et al. (2024) utilizes a lightweight estimation of the Hessian diagonal, which allows the algorithm to incorporate curvature information without computing the full Hessian matrix, thus ensuring the efficiency of both communication and computation. The advantage is the accelerated convergence in complex topographies, while the limitation remains the sensitivity to the estimation variance of the diagonal preconditioner under heterogeneous data distributions. Privacy&distributed. Addressing the orthogonal requirement of data security, optimization frameworks must integrate rigorous protection mechanisms. TOPDP Guo et al. (2021) rigorously verifies theoretical privacy guarantees through the framework of differential privacy, which is suitable for visual tasks such as image classification. The advantage is the provable bound on information leakage, whereas the limitation
is the inherent trade-off between the scale of noise injection and the final utility of the model. The evolution of distributed hybrid optimization exhibits a paradigm shift from heuristic single dimensional compression to theoretically guaranteed, multi dimensional joint designs. Initial methods primarily focused on reducing communication via simple gradient quantization, which gradually evolved into sophisticated frameworks that simultaneously manage variance reduction, curvature approximation, and privacy noise. Despite these advancements, a significant gap remains in the existing work. Specifically, there is an absence of a unified theoretical framework that can simultaneously bound the convergence errors introduced by extreme compression, the inaccurate estimation of second-order curvature, and the strict constraints of differential privacy, particularly under conditions of severe data heterogeneity.
3.6
Privacy-Preserving Optimization
The framework of privacy-preserving optimization can generally be divided into fundamental algorithmic frameworks and trade-off management strategies. The primary distinction between these aspects relies on whether the objective focuses on the mechanism of privatization or the mitigation of adverse side effects on convergence. Within the fundamental frameworks, differential privacy optimization serves as the core paradigm. This approach typically revolves around modifying vanilla optimizers to incorporate rigorous privacy guarantees while addressing computational inefficiencies. By formulating the privatized update rule as θt+1= θt − P 2 2 ηt B1 i∈Bt ΠC (∇J(xi , θt ))+N (0, σ C I) , where ΠC represents the per-example gradient clipping operator and σ denotes the noise multiplier, this foundational paradigm can be seamlessly encapsulated into our unified formulation Eq. (7), from which various methods are extended across different optimization orders. Furthermore, specific attention is directed toward gradient noise injection, which strategically perturbs computations to enforce privacy protocols. These techniques vary from adaptive intensity adjustment and scalar injection to noise-robust designs and geometric perturbations that align noise with the direction of the gradient. To manage the side effects, trade-off management strategies aim to reconcile data protection with model fidelity, predominantly relying on adaptive mechanisms to mitigate the impact of noise. Several studies focus on dynamic noise scheduling or post-processing optimization to refine gradients after noise injection. A closely related technique is privacy-aware gradient clipping,
40
which regulates the magnitude of gradients to prevent instability. These methods encompass global clipping strategies, adaptive clipping based on gradient statistics, and curvature-aware clipping to guide updates in non-convex landscapes. Federated privacy enhancement extends these concepts to collaborative environments, where approaches focus on neutralizing noise amplification during aggregation through protection protocols and structural adaptations. The protection protocols involve federated gradient shuffling to mix private and public data. The structural adaptations target federated noise aggregation, employing techniques such as SVD reparameterization or SAM-based optimizers to ensure robustness against aggregated noise. Additionally, low-rank adaptations are utilized to eliminate truncation bias in heterogeneous settings.
3.6.1
tional diffusion modelsTransitioning to zeroth-order optimization, SPARTA Makni et al. (2025) implements structured perturbations to achieve efficiency under privacy constraints. Additionally, DP-SGDJL Bu et al. (2021) leverages Johnson-Lindenstrauss projections to reduce the computation overhead of gradient norms, DOPPLER Zhang et al. (2024a) applies frequency-domain filtering to boost the signalto-noise ratio, and RaCO-DP Yaghini et al. (2025) adapts gradient ascent for rate-constrained tasks with theoretical convergence guarantees. Joint differential privacy. To provide a theoretical convergence analysis, studies focus on deriving bounds for population loss. The method ANSGD Huang et al. (2023) derives theoretical bounds on population loss in stochastic convex optimization, and empirical studies on multi-class classification effectively validate the validity of these insights. It provides specialized solutions to balance differential privacy and optimization performance in joint learning scenarios. Dynamic noise scheduling. Moving beyond static parameters, dynamic scheduling algorithms adjust noise injection based on the training context or data properties to address the rigid limitations of static noise levels. TOP-DP Guo et al. (2021) employs topology-aware noise reuse and time-aware decay to reduce noise progressively over the training process. Similarly, DC-SGD Wei et al. (2025) dynamically sets clipping thresholds utilizing differential privacy histograms of gradient norms. These approaches balance privacy guarantees and model utility through context-aware adjustments. Privacy-utility balance. The critical trade-off between differential privacy guarantees and model utility necessitates advanced balancing strategies, as excessive noise degrades performance while insufficient noise compromises privacy. In distributed settings, TOP-DP Guo et al. (2021) uses topology-aware noise reuse to maintain utility. Alternatively, InterleavedShuffleG Jiang et al. (2025a) combines private and public data within a hybrid framework to balance differential privacy and utility through tailored data usage.
Differential Privacy Optimization
Differential privacy optimization encompasses specialized algorithmic frameworks implemented within privacy-preserving training to resolve the tension between rigorous data protection and model fidelity. These techniques function to mitigate the adverse effects of noise injection and gradient bias through adaptive mechanisms and error correction, thereby maximizing utility and convergence stability without compromising privacy guarantees. DP-SGD variants. Addressing the key limitations of the vanilla DP-SGD Yang et al. (2022), including high noise impact, biased optimizer estimates, and computational inefficiencies, requires targeted modifications across optimization orders. For firstorder methods, TOP-DP Guo et al. (2021) reduces noise variance through topology-aware reuse and time decay, whereas DPIS Wei et al. (2022) utilizes importance sampling and adaptive clipping to achieve lower variance and privacy cost. DP-λCGD Kalinin et al. (2026) further enhances first-order privacy training by introducing lightweight iterative noise correlation. To address second-order momentum adaptations, DPMicroAdam Hudişteanu et al. (2025) utilizes top-k sparsification to filter out updates with low signal-tonoise ratios for bias reduction, and it employs an error feedback accumulation formula to recover lost signals. Furthermore, Corrected DP-Adam Tang and Lécuyer (2023) corrects the bias in the second moment estimation under differential privacy, while DP-AdamW Sun et al. (2025a) applies decoupling to make the regularization strength independent of instantaneous gradient noise levels. DP-aware AdaLN-Zero Huang et al. (2026b) mitigates conditioning-induced heavy-tailed gradients that trigger excessive clipping in condi-
The evolutionary trajectory of privacy-preserving optimization demonstrates a fundamental shift from static, uniform noise injection to adaptive, structureaware perturbations. The milestone paradigm transition involves adapting the three major optimization orders to differential privacy constraints. First-order modifications primarily focus on adaptive clipping and variance reduction. SO adaptations attempt to correct the biased moment estimations caused by noise injection. ZO methods explore structured perturbations for derivative-free privacy. The primary
41
advantage of these methodologies is the provision of rigorous mathematical guarantees for data protection. However, the inherent limitation is the unavoidable degradation of convergence rates and model utility, particularly in highly non-convex landscapes. The core gap in existing research remains the lack of a unified theoretical framework capable of optimally bounding the privacy-utility trade-off across diverse optimization orders without relying on extensive hyperparameter tuning or prior knowledge of the data distribution.
tuning preventing underfitting. ZO methods minimize communication but assume strict convexity for convergence bounds. A unified framework guaranteeing tight theoretical convergence, minimal memory footprint, and geometry aware calibration across heterogeneous nonconvex landscapes remains absent. 3.6.3
Privacy-Utility Tradeoff
To mitigate the utility loss caused by DP, existing studies have optimized from two dimensions: gradient clipping and noise handling. DC-SGD Wei et al. (2025) focuses on dynamic adjustment during the gra3.6.2 Gradient Noise Injection dient clipping phase. This framework adaptively sets Gradient noise injection enforces differential privacy the clipping threshold using a differentially private in model training via the update θt+1 = θt −η(Clip(gt )+ histogram of gradient norms, reducing the burden of N (0, σ 2 I)), perturbing bounded gradients. Recent manual hyperparameter tuning while improving opmethods calibrate noise using adaptive and geomettimization efficiency. In contrast, DP-LSSGD Wang ric mechanisms and enhance optimization resilience et al. (2020a) devotes itself to signal recovery in the via second order and zeroth order approximations to post-processing stage. By introducing a real-time preserve theoretical convergence and empirical perdenoising mechanism based on Laplacian smoothing, formance. it smooths the gradients injected with Gaussian noise. Adaptive noise intensity. First order optimizaThese two methods effectively enhance model utility tion conventionally applies uniform vector perturwhile maintaining strict DP guarantees, by optimizbation Yang et al. (2022). This method dynamically ing the input clipping threshold and output noisy calibrates this; DC-SGD Wei et al. (2025) adjusts gradients, respectively. clipping thresholds via gradient norm statistical distributions to optimize computation privacy tradeoffs. 3.6.4 Federated Privacy Enhancement Geometric noise injection. GeoDP-SGD Duan Federated privacy enhancement uses specialized proet al. (2025) structurally decomposes perturbations, tocols in collaborative learning ecosystems to forwhich separates gradient direction and magnitude, tify data confidentiality during aggregation. The offering granular control over isotropic noise to immathematical framework defines the server update prove structural efficiency. PK (i) as θt+1 = θt − η i=1 wi M(∆θt ), with M(·) beNoise robust optimization. Addressing privacy ining a randomized privacy mechanism for local upduced variance, DP-Adam Tang and Lécuyer (2023) dates. This update rule can be instantiated in the corrects second moment estimation bias when stanunified perspective Eq. (7) by interpreting the perdard Adam Kingma (2015) processes noisy gradi(i) ents. For landscape smoothing, DP-FedSAM Shi client local update ∆θt as the gradient estimate (i) (i) et al. (2023) integrates SAM Foret et al. (2021) to g̃t = Ei (f, θt , ξt ) from client i, embedding the prilocate flatter loss regions, naturally resisting gradient vacy mechanism and aggregation into the scenario P (i) perturbations and stabilizing high noise convergence. transformation Tscenario such that ĝt = i wi M(g̃t ). Scalar noise injection. In communication conCore improvements neutralize noise amplification strained vertical federated learning. DPZV Zhang and prevent information leakage via architectural et al. (2025b) replaces high dimensional vector perreparameterization and algorithmic resilience. Theturbations with scalar noise. Discarding explicit FO oretical convergence analysis shows that controlling directional derivatives, it theoretically matches their distributed noise variance enables algorithms to apconvergence rates, balancing strict privacy with low proach non-private baseline convergence rates. communication overhead to achieve distributed memArchitecture collaboration. To mitigate utility loss ory efficiency. under strict differential privacy constraints, this dimension optimizes gradient interaction modes. The The design paradigm evolves from isotropic perturmethod in Jiang et al. (2025a) introduces Interleavedbation to geometry aware scaling, landscape regularShuffle, reformulating collaborative architecture and ization, and memory efficiency. Adaptive methods transmission mechanisms. Utilizing encrypted graoffer real time calibration but increase per step comdient transmission, it distributes noise calibration putational complexity. Robust optimizers stabilize across clients and the server. This structural adaptaconvergence but demand rigorous hyperparameter 42
tion bounds aggregation function sensitivity, reducing privacy leakage risks without altering fundamental first-order optimization dynamics. Optimizer-level robustness. Progressing from transmission modifications to intrinsic algorithmic stability, DP-FedSAM Shi et al. (2023) enhances algorithmic tolerance to aggregated noise. Addressing noise amplification during server aggregation, it integrates the client-side SAM optimizer. Rather than modifying transmission structures, this approach identifies flat minima within local loss landscapes. Exploring a broader solution space grants the objective function landscape-aware robustness, improving resistance to gradient perturbations while maintaining theoretical convergence guarantees.
histograms of gradient norms to dynamically approximate the underlying distribution. Conversely, StableSPAM Huang et al. (2025a) adopts an EMA of the maximum gradient for entry-wise clipping. These data-driven techniques adapt to varying gradient scales, outperforming static clipping in privacy efficiency and empirical stability. Curvature-aware clipping. Advancing to secondorder landscape utilization, some methods address naive scalar clipping limitations in highly non-convex environments. Traditional methods uniformly scale all parameter dimensions, ignoring local landscape geometry and yielding suboptimal updates. Resolving this, Sophia Liu et al. (2024a) leverages curvature estimates to precondition the clipping operation. Incorporating second-order insights makes the clipping threshold structurally aware of the parameter space, fundamentally improving convergence stability and update efficiency in complex tasks.
The design paradigm evolves from structural transmission modifications to landscape-aware algorithmic resilience. Architecture collaboration provides strict privacy guarantees via transmission encryption but exhibits high communication overhead. Optimizerlevel robustness mitigates local noise amplification and improves generalization, yet requires substantial local computational resources. A unified framework simultaneously minimizing communication overhead, optimizing local landscape flatness, and adapting to heterogeneous data distributions under strict privacy constraints remains absent. 3.6.5
The trajectory advances from static global constraints to dynamic statistical calibration and curvature informed structural clipping, reflecting the sequential adaptation of FO, dynamic FO, and approximated SO optimization. Global clipping offers straightforward implementation but suffers suboptimal performance across heterogeneous training phases. Adaptive clipping improves the utility tradeoff via real-time calibration but incurs auxiliary computational overhead. Curvature-aware clipping maximizes landscape utilization to stabilize optimization but requires computationally expensive curvature approximations. The core gap remains the absence of a unified clipping operator seamlessly integrating ZO memory efficiency with SO curvature awareness under strict DP constraints.
Privacy-Aware Gradient Clipping
Privacy-aware gradient clipping bounds algorithmic sensitivity in DP and stochastic optimization. The mathematical framework updates gradients via θt+1 = θt −ηĝt , with the clipped gradient ĝt = min(1, ∥g̃Ct ∥ )g̃t restricting update magnitude through a threshold C. Core improvements transition from static global constraints to dynamic statistical adaptations and curvature-informed reparameterization. Global gradient clipping. Some methods enforces a uniform optimization constraint. Addressing DP in stochastic optimization, AClipped-dpSGD Jin et al. (2024) uses a one-time global clipping strategy for convex problems. For similarity-based objectives, Logit-DP Kong et al. (2025) clips logit gradients before loss computation. Although targeting distinct scenarios, both rely mathematically on predefined global thresholds to restrict sensitivity, reduce noise, and establish fundamental theoretical convergence bounds. Adaptive clipping. Somw methods introduce dynamic first-order modulation. By continuously estimating gradient sequence statistics, these methods calibrate clipping thresholds in real-time. For example, DC-SGD Wei et al. (2025) constructs DP
3.7
Memory-Efficient Optimization
These methods bifurcates into structural design strategies (algorithmic reformulation) and numerical compression techniques (data representation scaling). Regarding structural design, Low-memory optimizer design minimizes storage overhead through mechanisms like statelessness Sane (2025), approximate second-moment storage Shazeer and Stern (2018), and state quantization or blocking. Stateless optimization methods further eliminate historical dependencies via real-time gradient statistics Zaheer et al. (2018), parameter-driven updates He et al. (2025a), and curvature estimation Shang et al. (2018). Additionally, memory-efficient fine-tuning supports massive architectures using stateless backpropagation and selective updates. Regarding numerical compression, methods focus on low-rank methods
43
and optimizer state compression. Low-rank methods compress statistics into lower-dimensional subspaces, encompassing general strategies like adaptive projections and specific low-rank gradient Storage techniques utilizing tensor decomposition. Optimizer state compression reduces the variable footprint through sparse compression Li et al. (2019), low-rank approximation, and state sharing Ma et al. (2025). 3.7.1
torization) to structural efficiency, Adam-mini Zhang et al. (2025d) partitions parameters into Hessianstructure-based blocks. By assigning a single learning rate per block, it matches AdamW’s performance with significantly fewer states. This structural approach addresses memory and communication bottlenecks in large-scale training by optimizing parameter grouping rather than just compressing individual state values. Target projection strategy. Diverging fundamentally from the previous methods that focus on managing internal optimizer states, tpSGD Lomnitz et al. (2023) redefines the training paradigm itself. By employing random label projections to generate local targets, it enables layer-by-layer training. This effectively eliminates the need for global backward passes, reducing memory overhead by altering the gradient propagation flow rather than the optimizer’s storage mechanism.
Low-Memory Optimizer Design
It refers to algorithmic techniques engineered to minimize the storage overhead inherent in training processes through mechanisms such as state quantization, approximation, and structural partitioning. These methods alleviate critical memory bottlenecks, enabling scalable and high-throughput learning by effectively reconciling limited hardware resources with the demands of maintaining model performance. Stateless optimizers. Representing the most aggressive approach to memory efficiency, AlphaGrad Sane (2025) addresses the overhead of stateful optimizers by completely eliminating historical gradient storage. By relying on a single steepness parameter for gradient transformation, it introduces a stateless optimization mechanism that serves as a baseline for reducing memory footprints without maintaining a history window. Approximate second-moment storage. Moving beyond the binary choice of stateful versus stateless, Adafactor Shazeer and Stern (2018) seeks to retain the benefits of adaptivity while mitigating costs. Unlike Sane (2025)’s removal of state, Adafactor Shazeer and Stern (2018) reduces the high memory overhead of second-moment storage through matrix decomposition. This approach bridges the gap between efficiency and performance by managing optimizer states through approximation rather than elimination. Quantized optimizer states. While Shazeer and Stern (2018) achieves compression through decomposition, 4-bit Shampoo Wang et al. (2024f) pursues state reduction via quantization. It preserves the structure of preconditioner eigenvector matrices but quantizes them to 4-bits. FlashOptim Ortiz et al. (2026) further advances memory-efficient optimization by combining ULP-based weight splitting and tailored companding functions for 8-bit quantization of momentum and variance states. This method demonstrates that high-performance state management can be maintained even with reduced numerical precision, offering an alternative compression strategy to factorization for large-scale tasks. Blocked optimizer states. Shifting focus from numerical compression (as seen in quantization or fac-
3.7.2
Low-Rank Algorithms
These methods denote algorithmic strategies that compress gradients or optimizer statistics into lowerdimensional subspaces using adaptive projections and randomized approximations. These techniques substantially mitigate memory constraints and communication overheads, effectively harmonizing the tradeoff between resource efficiency and approximation fidelity to support large-scale model training. Low-rank projection. Focusing on direct dimensionality reduction, AdaRankGrad Refael et al. (2025c) mitigates the memory footprint of Adam by applying low-rank projections to the gradients. Unlike fullrank updates, it employs a randomized SVD scheme to efficiently approximate the gradient subspace. This integrates adaptive rank reduction with Adam’s learning rates, establishing a baseline for reducing memory via efficient matrix factorization. Dynamic rank adjustment. While Refael et al. (2025c) applies projections to gradients, Adapprox Zhao et al. (2024b) specifically targets the memory-intensive second moment of the optimizer. Furthermore, it distinguishes itself from static low-rank methods by introducing an adaptive rank selection mechanism. This allows the algorithm to dynamically adjust the rank of the approximation during training, striking a more flexible balance between optimization accuracy and memory overhead for large-scale models. Low-Rank&quantization. Moving beyond matrix factorization, hybrid approaches combine low-rank structures with data precision reduction. QuZO Zhou et al. (2025) complements low-rank adaptation with stochastic quantized perturbation to specifically reduce gradient estimation bias. Similarly, LQ-SGD Li et al. (2025a) extends the low-rank framework of Vo-
44
gels et al. (2019) by incorporating logarithmic quantization. These methods demonstrate that combining low-rank constraints with quantization can simultaneously lower communication costs and memory usage without compromising model performance. 3.7.3
and stable model training. Gradient low-rank projection. While both methods aim to reduce storage by projecting high-dimensional gradients onto low-dimensional subspaces, they diverge in their mathematical foundations. SUMO Refael et al. (2025b) relies on matrix decomposition, utilizing truncated SVD of per-layer gradients to form adaptive projection subspaces. In contrast, GWT Wen et al. (2025) approaches compression through signal processing, applying wavelet transforms to compress Adam Kingma (2015) states. This distinction highlights two viable paths—algebraic decomposition versus spectral transformation—for achieving significant memory reduction in optimizer states. Dynamic gradient rank. Moving beyond static projection, these methods introduce adaptivity into the rank selection process. Adapprox Zhao et al. (2024b) directly targets memory efficiency by explicitly adjusting the storage rank of the second-moment matrix via an adaptive selection mechanism. HELENE Zhao et al. (2024a), conversely, achieves a similar effect implicitly; rather than manipulating rank parameters directly, it employs layer-wise adaptive clipping and annealing to prioritize gradient information. Both strategies ultimately enhance large model optimization, but distinguish themselves by whether the dynamic ranking is an explicit architectural choice or an emergent property of gradient regularization. Gradient reconstruction. Distinct from the projection or ranking of existing gradients, TeZO Sun et al. (2025b) addresses memory through structured estimation within a zeroth-order framework. It employs canonical polyadic decomposition (CPD) to reconstruct gradients by capturing low-rank temporal and cross-model structures. Unlike methods that compress a calculated gradient, TeZO Sun et al. (2025b) reduces training costs by dynamically adjusting the estimator’s rank based on layer characteristics, optimizing memory usage through tensor decomposition rather than direct state compression.
Optimizer State Compression
It refers to a suite of methodologies that drastically reduce the memory footprint of optimization variables by employing sparsity, low-rank approximations, and state-sharing mechanisms. These strategies mitigate storage constraints, securing a viable compromise between minimized resource consumption and the preservation of training stability and convergence performance. Sparse state compression. By sparsifying state, some methods reduce memory usage while preserving performance and stability for efficient training. SPAM Huang et al. (2025b) integrates spike-aware clipping into Adam, and uses sparse momentum to cut memory while maintaining performance. MICROADAM Modoranu et al. (2024) compresses gradients before optimizer state input, retaining convergence guarantees competitive. These methods balance memory efficiency with training performance and stability via sparse states. Low-rank state approximation. Adapprox Zhao et al. (2024b) employs randomized low-rank matrix approximation for Adam’s second moment, introduces adaptive rank selection to balance accuracy and memory. It addresses the high memory and computational overhead of full-rank second-moment estimation in Adam. State sharing. By sharing state information across parameters, some methods address excessive memory overhead in adaptive optimizers. Adafactor Shazeer and Stern (2018) proposes that for matrix parameters, instead of storing individual states for each element, it enables parameters in the same row or column to share states. AdaAct Seung et al. (2024a) shares neuron-wise rates for same input features. State sharing methods boost optimization efficiency, stability, and adaptability via cross-parameter state reuse. 3.7.4
3.7.5
Stateless Optimization
These methods eliminate the dependency on historical data accumulation by deriving update directives directly from instantaneous computations and intrinsic parameter attributes. These strategies leverage immediate feedback mechanisms, ranging from realtime gradient statistics to lightweight curvature estimates, to drastically minimize memory overhead and computational latency, thereby ensuring efficient convergence and adaptability without the burden of maintaining persistent optimizer states. Real-time gradient statistics. SWAN Ma et al.
Low-Rank Gradient Storage
It alleviates the substantial memory footprint associated with high-dimensional optimization by approximating full-matrix gradients through compressed, lower-dimensional representations. These approaches utilize techniques such as subspace projection, dynamic rank adaptation, and tensor decomposition to drastically curtail state storage requirements while preserving essential information fidelity for efficient
45
(2025) integrates core preprocessing, efficiency tweaks, and variants for flexible trade-offs, all relying on instantaneous stats. It addresses the overhead and bias of exponential moving averages in adaptive optimizers by using instantaneous gradient statistics for stateless updates. Parameter characteristic-driven updates. Some methods focus on adapting updates using parameterspecific properties. AlphaGrad Sane (2025) uses a single steepness parameter for memory-efficient scaling. SGD-SaI Xu et al. (2024) adjusts initial learning rates through gradient signal-to-noise ratio. BFE Cao (2022) tunes learning rates through forward loss information; AdaBFE Cao (2022) extends it to per-parameter adaptation for high-dimensional spaces. These methods enhance optimization by parameter-specific adaptations, balancing efficiency, performance, and adaptability across tasks. Real-time curvature estimation. Sophia Liu et al. (2024a) uses a lightweight diagonal Hessian estimate as a preconditioner, eliminating the dependency of traditional second-order optimizers on extensive storage via lightweight real-time curvature estimation. Such method efficiently use real-time curvature information to boost convergence without extra computational burden. 3.7.6
for large models. Sparse-MeZO Liu et al. (2024b) dynamically selects a sparse parameter subset through layer-wise magnitude thresholds. LeZO Wang et al. (2024b) integrates SPSA with layer-wise sparsity, selectively updating layers to achieve efficient finetuning without additional memory. These methods enable cost-effective large model fine-tuning by dynamic parameter selection, balancing efficiency and performance.
3.8
Tailored Optimization Approaches
These methods are categorized into automated design paradigms and robustness enhancement mechanisms, distinguishing between rule discovery and engineered resilience. Auto-designed optimizers transcend human heuristics by automating rule formulation. These include programmatic search within structured spaces Kingma (2015) Nocedal (1980), neural controllers modeling physical transport, and symbolic derivation. Additionally, techniques leverage evolutionary strategies to escape local optima or meta-learning for dynamic, task-specific adaptation. Robust optimization mitigates training vulnerabilities through gradient resilience and systemic adaptability. Gradient resilience ensures stability via noise-robust gradients with adaptive clipping Wang et al. (2024a) and hardware noise countermeasures Sane (2025). Systemic adaptability addresses broader constraints, encompassing distribution shift robustness for federated learning and structure-aware optimization for large scale feasibility, collectively ensuring reliable convergence under volatile conditions.
Memory-Efficient for Large Models
It circumvents the prohibitive resource demands associated with full-parameter training by innovating gradient computation and application protocols. These strategies deploy techniques ranging from stateless backpropagation and zeroth-order kernel estimation to sparse parameter selection, effectively minimizing memory footprints to enable the feasible optimization of massive architectures on constrained hardware while maintaining computational effectiveness. ZO fine-tuning. KerZOO Mi et al. (2025) is a kernelfunction-based ZO framework using kernel smoothing to mitigate lower-order bias, enhancing convergence speed. This method boosts ZO fine-tuning efficiency for LLMs through kernel-based gradient estimation improvement. Staless fine-tuning. LOMO Lv et al. (2024) reduces memory usage by computing the gradient for each parameter, immediately updating the parameter, clearing the gradient afterward, and retaining the gradient of at most one parameter in memory. It enables staleless fine-tuning by fusing backpropagation with parameter update. Selective parameter fine-tuning. By dynamically selecting subsets of parameters to perturb and update, some methods address the prohibitive computational and memory overhead of full-parameter fine-tuning
3.8.1
Auto-Designed Optimizers
These optimizers transcend the rigid boundaries of manually engineered heuristics by automating the discovery and formulation of update rules through algorithmic exploration. These systems harness diverse generative paradigms, ranging from symbolic derivation and evolutionary search to meta-learning and programmatic synthesis, to uncover novel optimization mechanisms that significantly enhance adaptability, generalization, and convergence capabilities beyond the reach of human intuition. Programmatic search optimizers. Some methods address the limitation of hand engineered optimizers by automating rule discovery through structured program spaces, exploring beyond human heuristics to find efficient update mechanisms. The method in Li and Malik (2017) replaces hand-engineered optimization rules through programmatic policy representation and automated search mechanisms, which carries significant heuristic value in the field of op-
46
3.8.2
timizer design. The method in Chen et al. (2023) frames algorithm discovery as a program search in infinite sparse spaces, uses evolutionary techniques to explore efficiently, and introduces Lion Chen et al. (2023), an optimizer with sign momentum updates that reduce memory usage. The method in Bello et al. (2017) employs a controller trained with RL to discover update rules such as PowerSign and AddSign. These methods share programmatic representation of rules but differ in search paradigms, uniting to automate optimizer discovery. Neural controller optimizers. KO Feng et al. (2025) models updates as stochastic collisions via Boltzmann transport equation to boost parameter diversity and generalization. It achieves precise regulation of parameter dispersion through hard and soft collision mechanisms, thus providing a physically inspired, theoretically provable, and engineering-efficient new paradigm for neural controller optimizers. Symbolic derivation optimizers. AGS-GD Starnes et al. (2024) reconstructs the gradient update rule through the symbolic mathematical derivation of Gaussian smoothing, converting local gradients into non-local gradients and addressing the pain point of traditional optimizers getting trapped in suboptimal local minima at the theoretical root. This method addresses the pain point of traditional optimizers getting trapped in suboptimal local minima at the theoretical root by means of symbolic reasoning. Evolutionary algorithm optimizers. Some methods address the local optima limitation of traditional gradient-based optimizers by integrating evolutionary strategies. GADAM Zhang and Gouza (2018) combines Adam Kingma (2015) with genetic algorithms using crossover and mutation to escape local optima. BGADAM Bai et al. (2021) extends this framework by adding a boosting strategy to generate diverse training sets. These optimizers merge evolutionary and gradient-based approaches to improve exploration and avoid local optima. Meta-learning optimizers. They address limitations of fixed-update optimizers by learning specific task’s update rules through meta-learning, enabling better adaptation to diverse objectives. WarpAdam Pan et al. (2024) inserts a learnable distortion matrix into Adam’s pipeline. HyperAdam Wang et al. (2019) uses an RNN Elman (1990) framework to adapt decay rates and weights, combining ensemble updates from varying Adam Kingma (2015) configurations. MADA Ozkara et al. (2024) unifies optimizers via hyper-gradient descent. These meta-learned optimizers dynamically tune update rules via meta-learning, enhancing task adaptability and outperforming fixed heuristics in diverse scenarios.
Robust Optimization
These methods mitigate the vulnerability of training procedures to systemic irregularities, ranging from stochastic gradient noise and hardware quantization errors to data distribution shifts. By incorporating resilience enhancing mechanisms such as adaptive clipping, drift correction, and structural decomposition, these strategies counteract destabilizing factors to ensure reliable convergence and consistent model performance amidst volatile or heterogeneous conditions. Noise-robust gradients. Some methods focus on improving the noise-robust of optimizer through different techniques. AdaGC Wang et al. (2025b) uses exponential moving averages for per-parameter local clipping thresholds. AdaNorm Dubey et al. (2023) corrects norms by EMA of past gradients, resolving low and atypical gradient issues. These methods enhance robustness of the gradient, improving the stability of optimizer across tasks. Distribution shift robustness. Some methods address client drift and local global objective misalignment under heterogeneous data distributions in federated learning, where vanilla methods suffer from performance degradation or high communication costs. FedCET Liu and Wang (2025) uses learning rates to weight client information. FAdamGC Chen et al. (2025a) integrates gradient correction into adaptive federated optimization. These methods enhance federated optimizers’ robustness to distribution shift and reduce communication overhead through tailored strategies. Hardware noise robustness. QuZO Zhou et al. (2025) reduces gradient bias via stochastic quantized perturbation. It enhances hardware noise robustness, enabling stable optimization in constrained setups. Structure-aware optimization. BC-ADMM Pan and Wu (2025) decomposes large scale problems into parallelizable subproblems allowing larger steps and faster convergence, maintains solution feasibility throughout the process. Such structure-aware methods enhance optimization efficiency and feasibility for large-scale tasks by tailored decomposition and relaxation strategies.
4
Experiments
Although Sec. 3 details the evolution of optimizers for constrained environments, such as distributed or differentially private training, fair empirical comparisons are often hindered by hardware and engineering variations. We therefore present a controlled evaluation protocol designed to isolate fundamental algorithmic
47
behavior from implementation details. Rather than benchmarking peak system throughput, we focus on evaluating the mathematical mechanisms underlying modern optimizers. We analyze hyperparameter sensitivity, convergence dynamics, and cross-architecture generalization to determine whether methods that perform well on vision tasks can transfer reliably to language modeling. The remainder of this section details our datasets, evaluation metrics, quantitative results, and qualitative analysis.
(2016) and Llama Llama (2023) models. Vision tasks. To assess the performance of different optimizers on image classification, we use Top1 accuracy as the primary metric. We utilize the DeiT Touvron et al. (2021) training setting for ViTS Dosovitskiy (2020). For ResNet-50 He et al. (2016), we evaluate models for 100 and 300 epochs, employing the A3 and A2 Wightman et al. (2021) training settings, respectively. Specifically, we consider three regular training settings for ImageNet-1K Hudişteanu et al. (2025) classification experiments: 1) DeiT setting Deng et al. (2009): Designed for 4.1 Datasets Transformer and modern CNN architectures like ViTTo comprehensively evaluate the optimizers, we seS Dosovitskiy (2020), this setting uses several adlected widely adopted deep learning benchmarks spanvanced augmentations, including RandAugment Cubuk ning image classification and language modeling. et al. (2020), Mixup Zhang et al. (2018), CutMix Yun Vision tasks. We used the standard ImageNetet al. (2019), and Random Erasing Zhong et al. 1K Deng et al. (2009) dataset, which comprises ap(2020). proximately 1.28M training and 50,000 validation 2) A2/A3 settings Wightman et al. (2021): Deimages across 1,000 classes. We selected ResNetsigned for CNNs LeCun et al. (2002) to match the 50 He et al. (2016) and ViT-Small (ViT-S) Dosovitperformance and convergence speeds of ViTs Dososkiy (2020) to represent CNN LeCun et al. (2002) vitskiy (2020), these settings reduce augmentation and vision transformer architectures, respectively. strengths and replace the cross-entropy (CE) loss Language tasks. We trained a 60M-parameter with binary cross-entropy (BCE) loss compared to Llama Llama (2023) model from scratch on the WikiText- the DeiT setting. 103 Merity et al. (2016) dataset. We emphasize that Ingredients and hyperparameters used for image clasthis 60M configuration is not designed to replicate sification training settings are detailed in Tab. 3. emergent behaviors at the scale of full-sized LLMs. InWe note that certain optimizers Robbins and Monro stead, it serves as a highly representative architectural (1951); Wang et al. (2025c); Zhuang et al. (2020); proxy for modern autoregressive Transformers, retainShazeer and Stern (2018); Taniguchi et al. (2024); Li ing the core architectural features of Llama Llama (2017); You et al. (2017); Zhang et al. (2019); Defazio (2023) while remaining computationally tractable. and Jelassi (2022) and SGDP Heo et al. (2021) exThis setup enables us to create a rigorous testbed hibit unstable or failed convergence on the ResNet He for assessing the cross-architecture transferability of et al. (2016) architecture when trained with BCE loss; optimizers, addressing the key question: Do optifor these cases, we used CE loss instead. mizers heavily tuned for CNN LeCun et al. (2002) Language tasks. We use perplexity (PPL) to evalor ViT Dosovitskiy (2020) perform catastrophically uate the performance of different optimizers on lanpoorly when faced with the highly anisotropic loss guage modeling tasks. Our training configuration landscapes characteristic of causal attention mechaaligns with that of Semenov et al. (2025), utilizing a nisms? batch size of 256 and a sequence length of 512. However, due to computational constraints, we employ a 60M-parameter Llama-like architecture, deviating 4.2 Metrics and Settings from the model size used in the reference. FurtherTo rigorously evaluate the performance of the diverse more, the total number of training tokens is deteroptimizers, we apply a rich set of evaluation metrics mined following the Chinchilla scaling laws Hoffmann tailored to the nature of each benchmark task. et al. (2022). Specifically, we consider a standard Hyperparameter setting. For a fair comparison, we training setting for the Llama Llama (2023) architecfocus exclusively on optimizing one common hyperture (60M parameters) on the WikiText-103 Merity parameter: the learning rate. Starting with default et al. (2016) dataset. This training scheme includes values suggested in the original literature for each specific data configurations, optimization setups, and optimizer, we employ a grid search strategy that exregularization tricks, using AdamW Loshchilov and plores the vicinity of these baselines by scaling them Hutter (2019) as an example: with a peak learnwith factors of 0.1, 0.2, 1.0, 5.0, and 10.0. The oping rate of 1 × 10−3 , a weight decay of 0.5, and an timal learning rates identified through this process optimizer ϵ of 10−8 . We use a cosine learning rate are subsequently transferred to the ResNet He et al. scheduler with a warmup phase of 2000 steps. To 48
Table 3 Ingredients and hyper-parameters used for image classification training settings. Taking ImageNet-1K as the template setups, the settings of A2/A3 Touvron et al. (2021) take ResNet-50 He et al. (2016) as the examples, the DeiT Touvron et al. (2021) setting takes ViT-S Dosovitskiy (2020) as the example. Gray regions should be modified for each optimizer. Procedure Dataset Epochs Batch size Optimizer Learning rate Optimizer Momentum Weight decay LR decay Warmup epochs Label smoothing ϵ Dropout Stochastic Depth Repeated Augmentation Gradient Clip. Horizontal flip RandomResizedCrop Rand Augment Auto Augment Mixup α Cutmix α Erasing probability ColorJitter EMA CE loss BCE loss
DeiT IN-1K 300 1024 AdamW 1 × 10−3 0.9, 0.999 0.05 Cosine 5 0.1 × 0.1 ✓ × ✓ ✓ 9/0.5 × 0.8 1.0 0.25 × × ✓ ×
A2 IN-1K 300 1024 AdamW 1 × 10−3 0.9, 0.999 0.05 Cosine 5 × × 0.05 ✓ × ✓ ✓ 7/0.5 × 0.1 1.0 × × × × ✓
Table 4 Comparison of various optimizers’ hyperparameter robustness and learning rate settings.
A3 IN-1K 100 1024 AdamW 1 × 10−3 0.9, 0.999 0.05 Cosine 5 × × × × × ✓ ✓ 6/0.5 × 0.1 1.0 × × × × ✓
Analysis
We systematically analyze the practical performance of various optimizers on diverse tasks and investigate their underlying mechanisms. These dynamics are visually corroborated by the training curves on the ViT-S Dosovitskiy (2020), ResNet-50 He et al. (2016), and Llama-60m Llama (2023) models (as detailed in Figs. 8, 9 and 11). 4.3.1
0.1×LR
0.2×LR
1.0×LR
5.0×LR
10.0×LR Best LR
SGD Robbins and Monro (1951) RMSprop Tieleman and Hinton (2012) Adam Kingma (2015) AdaBelief Zhuang et al. (2020) Adafactor Shazeer and Stern (2018) RAdam Liu et al. (2020a) AdamP Heo et al. (2021) SGDP Heo et al. (2021) AdamW Loshchilov and Hutter (2019) Adan Xie et al. (2024b) ADOPT Taniguchi et al. (2024) Kron Li (2017) LAMB You et al. (2020) LAPROP Ziyin et al. (2020) LARS You et al. (2017) Lion Chen et al. (2023) Lookahead Zhang et al. (2019) MADGRAD Defazio and Jelassi (2022) MARS Yuan et al. (2025) Muon Jordan et al. (2024) Nadam Dozat (2016) NadamW Wightman (2019) NovoGrad Ginsburg et al. (2019)
48.056 54.254 66.234 29.248 66.294 66.220 67.336 49.800 67.082 74.562 62.408 75.878 24.346 66.834 8.3020 63.192 38.492 72.218 62.414 73.844 66.594 67.332 10.290
55.918 43.416 70.772 40.210 70.834 70.826 71.856 57.662 72.076 74.956 66.552 76.432 37.750 71.834 12.426 71.092 48.896 74.150 70.300 75.944 70.742 71.872 16.686
64.360 20.512 72.240 58.562 74.630 71.190 75.444 66.536 75.196 75.190 64.684 77.304 69.100 75.050 32.892 76.334 62.898 74.634 74.474 78.252 72.180 75.278 48.892
60.590 6.1140 66.528 73.344 72.606 61.182 70.136 38.170 75.858 6.3640 53.660 72.254 64.320 68.946 74.608 75.206 74.982 66.514
59.262 68.642 60.416 0.2640 29.670 69.956 58.276 65.988 61.962 61.526 72.542 64.258
0.1 1e-4 1e-3 1e-2 1e-3 1e-3 1e-3 0.1 1e-3 1.25e-2 2e-4 1e-3 5e-3 1e-3 1 1e-4 0.5 1e-3 5e-3 2e-3 1e-3 1e-3 5e-3
High-performance robustness. Kron Li (2017) and Adan Xie et al. (2024b) show minimal variance at lower scales (0.1× to 1.0×) but are prone to divergence at high learning rates (LR). This occurs because Adan Xie et al. (2024b) reformulates Nesterov Nesterov (1983) momentum and integrates adaptive gradient mechanisms, accelerating convergence via a targeted look-ahead momentum strategy. However, the coupling of momentum terms with adaptive gradients creates high update inertia; when the base LR is scaled up, large gradient magnitudes combined with strong momentum inertia cause drastic oscillations of parameter updates on the loss landscape, negating the stabilizing effect of the adaptive denominator. Muon Jordan et al. (2024) stands out as the most hyperparameter-insensitive, maintaining > 75% accuracy from 0.2× to 5.0×, indicating superior ease of tuning. It constrains parameter updates to a limited manifold through iterative orthogonalization of the accumulated momentum matrix, regardless of how the external learning rate is scaled. This intrinsic regularization makes it highly insensitive to hyperparameter variations. High lower-bound stability. Despite notable degradation outside optimal settings, Lion Chen et al. (2023), MADGRAD Defazio and Jelassi (2022), and MARS Yuan et al. (2025) maintain a "safe" baseline (> 60% accuracy) across the full 0.1× to 10× range. The core of Lion Chen et al. (2023) is its parameter update mechanism, which uses only the sign of the momentum-weighted gradient (sign(ct )) and discards the exact magnitude. This means that regardless of whether gradients explode or vanish due to learning rate scaling, the sign update term remains bounded to {−1, 0, +1}. This ensures the core step size of parameter updates is determined solely by the LR and is identical across all dimensions for the sign com-
ensure training stability and generalization, gradient clipping is applied with a threshold of 0.5. The training is conducted on 8 GPUs with a sequence length of 512 and a per-device batch size of 32 (resulting in a global batch size of 256). Ingredients and hyperparameters used for this language modeling setting are strictly controlled to benchmark optimization performance.
4.3
Optimizer
Learning Rate Sensitivity on ViT-S
To systematically assess learning rate sensitivity, we report the Top-1 accuracy across a broad spectrum of varying learning rates, explicitly highlighting the identified optimal values that yield peak performance (see Tab. 4 and Fig. 8).
49
Learning Rate Setting — Base LR×10 — Base LR×5
— — —
60
Base LR/10
0
60 40 20 40
60
80
20
20
80
20
40
40
Base LR Base LR/5
80 RAdam
0
AdamP
0
20
40
60
80
100
20 0
20
40
60
80
100
0
20
60
40
40
20
20
20
20
0
20
40
60
80
100
LAMB
0
20
40
60
80
100
0
20
40
60
80
60
80
100
40
60
80
100
40
60
80
100
40
60
80
100
20 0
100
40
20
40
60
80
100
0
20
80 LAPROP
80 LARS
80 Lion
80 Lookahead
60
60
60
60
60
40
40
40
40
40
20
20
20
20
100
0
40
20
40
60
80
0
100
MARS
20
40
60
80
100
Muon
20
40
60
80
100
20
40
60
80
100
20 0
20
80 NadamW
60
60
60
40
40
40
20 0
0
80 Nadam
80
20 100
100
60
60
80
80
40
40
60
60
60
60
40
40
40
80
20
20
40
20
60
80 MADGRAD
0
0
40
20
40
100
20
100
40
Adafactor
60
20 80
80
60
80 ADOPT
40
60
60
80
60
80 Adan
60
40
40
80 AdaBelief
60
80 AdamW
80
20
20
80 Adam
80 SGDP
80 Kron
0
60 RMSprop
SGD
80
20 0
20
40
60
80
100
40
60
80
100
20
40
60
80
100
20
50 30
20 0
0
70 NovoGrad
10 0
20
40
60
80
100
0
20
Figure 8 Comparison of Top-1 accuracy and learning rate sensitivity analysis of various optimizers. Each subfigure demonstrates the robustness and sensitivity of individual optimizers across a range of scaled learning rates (from 0.1× to 10× the base learning rate). Figures demonstrate that convergence stability and final Top-1 accuracy heavily depend on precise hyperparameter tuning.
ponent. This mechanism inherently protects against numerical instability caused by anomalous gradient magnitudes, resulting in a remarkably high training floor. MADGRAD Defazio and Jelassi (2022) combines an averaging-based update mechanism with adaptive scaling properties, centered around a dual averaging framework. Instead of applying adaptive scaling to gradients step by step, it smoothly accumulates historical gradients with weights, followed by normalization via a cube-rooted adaptive denominator. This mechanism naturally buffers against extreme gradient shocks. Large learning rate sensitivity (Adam family). These methods are highly vulnerable to aggressive LRs. Scaling by 5× or 10× typically leads to non-convergence or sharp performance declines. Specific preferences. RMSprop Tieleman and Hinton (2012) exhibits limited robustness and prefers conservative LRs. In contrast, AdaBelief Zhuang et al. (2020), LAMB You et al. (2020), LARS You et al. (2017), and NovoGrad Ginsburg et al. (2019) face convergence difficulties at small LRs, necessitating larger learning rates. The core of LARS You et al. (2017) and LAMB You et al. (2020) is a trust ratio: they normalize the update step size for each layer by computing the ratio of the parameter norm to the update vector norm, ensuring stable update magnitudes across layers. This mechanism has a critical weakness: if the global base learning rate is set excessively small, multiplying it by the trust ratio will excessively compress the actual parameter update step size. This prevents the model from escaping initial flat regions in the early training stage. Consequently, these optimizers require larger learning rates 50
to drive normalized updates and produce sufficient parameter movement. SGD resilience. While SGD Robbins and Monro (1951) lags in peak accuracy (64.36%), it demonstrates exceptional robustness to large learning rates, remaining usable even at 10× scaling where other optimizers fail. The update formula for adaptive optimizers like Adam Kingma (2015) involves scaling the first moment of gradients by the square root of the second moment of gradients plus a small epsilon value, all multiplied by the negative learning rate. For parameters with extremely small gradients, the square root is also minuscule, and the division yields an enormous multiplier. When the base learning rate is further increased by 10×, the actual step size for these parameters becomes astronomically large, instantly corrupting the model weights. In contrast, the update rule for SGD Robbins and Monro (1951) is straightforward: it simply multiplies the negative learning rate by the gradient directly. Even if the LR is scaled up 10×, the update step size remains linearly proportional to the gradient. While a large LR causes SGD Robbins and Monro (1951) to oscillate violently around the optimal point, it never triggers numerical explosion due to division by a tiny value. Efficacy of decoupled weight decay modifiers The integration of decoupled weight decay consistently yields measurable performance improvements over foundational algorithms. At the 1.0× learning rate, AdamW Loshchilov and Hutter (2019) achieves 75.196, outperforming standard Adam Kingma (2015) by a notable margin. A parallel improvement is evident when comparing NadamW Wightman (2019), which reaches 75.278, to its unmodified counterpart
Table 5 Results on ViT-S Dosovitskiy (2020), ResNet50 He et al. (2016), and Llama-60m Llama (2023) models across various optimizers. We compare different learning rate multipliers on Llama-60m Llama (2023) and show that Muon Jordan et al. (2024) consistently achieves strong performance across all architectures. Optimizer SGD Robbins and Monro (1951) RMSprop Tieleman and Hinton (2012) Adam Kingma (2015) AdaBelief Zhuang et al. (2020) Adafactor Shazeer and Stern (2018) RAdam Liu et al. (2020a) AdamP Heo et al. (2021) SGDP Heo et al. (2021) AdamW Loshchilov and Hutter (2019) Adan Xie et al. (2024b) ADOPT Taniguchi et al. (2024) Kron Li (2017) LAMB You et al. (2020) LAPROP Ziyin et al. (2020) LARS You et al. (2017) Lion Chen et al. (2023) Lookahead Zhang et al. (2019) MADGRAD Defazio and Jelassi (2022) MARS Yuan et al. (2025) Muon Jordan et al. (2024) Nadam Dozat (2016) NadamW Wightman (2019) NovoGrad Ginsburg et al. (2019)
ViT-S
ResNet-50
Acc (%)
Acc (%)
Base LR
64.360 54.254 72.240 68.642 74.630 71.190 75.444 66.536 75.196 75.190 66.552 77.304 75.858 75.050 58.276 76.334 64.320 74.634 74.608 78.252 72.180 75.278 66.514
75.926 71.536 74.480 74.232 76.808 74.612 76.818 74.066 76.314 76.770 74.584 75.226 77.590 76.146 76.196 75.582 76.872 76.940 77.294 77.996 74.732 76.514 71.694
522.57 614.93 43.27 13.78 667.83 13.79 13.42 14.11 655.08 12.58 13.56 15.32 210.44 13.80 473.88 12.12 12.40 627.00 14.95 1176.20
onto restricted manifolds. This intrinsic geometric regularization effectively flattens exaggerated step sizes in the highly anisotropic loss landscape of large language models. MARS Yuan et al. (2025), on the other hand, introduces a scaling parameter to control the strength of variance reduction and subtracts the gradient of the previous parameters on the current batch from the current step’s gradient. This gradient correction mathematically cancels the inherent stochastic noise introduced by diverse network architectures and complex data distributions. High transferability with constraints. While showing strong general performance, Adafactor Shazeer and Stern (2018), AdamP Heo et al. (2021), AdamW Loshchilov and Hutter (2019), Adan Xie et al. (2024b), LAMB You et al. (2020), LAPROP Ziyin et al. (2020), Lion Chen et al. (2023), NadamW Wightman (2019), and Kron Li (2017) slightly trail Muon Jordan et al. (2024) and MARS Yuan et al. (2025). Kron Li (2017) serves as a prime example: despite achieving a toptier PPL of 12.57 on Llama at the base LR (demonstrating a high ceiling), its performance collapses at 5× LR (PPL 336.2). This reveals a significantly narrower hyperparameter "safety margin" compared to robust alternatives. One core mechanism enabling optimizers like AdamW Loshchilov and Hutter (2019) and NadamW Wightman (2019) to remain competitive across diverse architectures is their successful decoupling of adaptive gradient updates from weight decay. When transitioning from shallow vision models to large language models with drastically different parameter scales, traditional regularization is easily distorted by the uneven accumulation of historical second moments. By placing the penalty term outside the scaling denominator, these methods ensure consistent and faithful regularization regardless of the absolute magnitudes of the parameters in each layer. LAMB You et al. (2020) adopts layer-wise adaptation to rigorously constrain step sizes. This mathematically decouples the magnitude of stochastic gradients from the applied updates across networks of varying depths. Similarly, Lion uses sign-based updates to maintain uniform magnitudes for all parameter updates, rendering the optimizer insensitive to gradient spikes common in complex attention mechanisms. Methods such as Adan Xie et al. (2024b) and AdamP Heo et al. (2021) significantly improve cross-architecture generalization by introducing sophisticated momentum estimation and trajectory correction. Architecture-specific instability. A large group of optimizers including RMSprop Tieleman and Hinton (2012), Adam Kingma (2015), MADGRAD Defazio and Jelassi (2022), and Adabelief Zhuang et al. (2020) performs acceptably on vision backbones but strug-
Llama-60m 5.0×LR 269.27 357.95 41.89 16.55 399.59 20.01 14.70 17.14 44.22 336.21 12.28 17.37 289.03 19.72 473.79 14.46 12.21 988.49 19.96 1192.70
0.2×LR 1178.34 964.46 72.43 16.48 1063.62 15.97 592.58 15.91 12.41 1208.53 13.39 18.72 15.86 164.25 13.52 642.97 12.75 14.01 999.46 17.75 309.63
Nadam Dozat (2016), which achieves 72.18. Furthermore, AdamP Heo et al. (2021) also demonstrates superior performance over standard Adam Kingma (2015) by reaching 75.444, highlighting the general effectiveness of refined parameter update rules. 4.3.2
Cross-Architecture Generalization
To evaluate cross-architecture generalization, we introduce a parameter transfer test and report the resulting Top-1 accuracy and PPL in Tab. 5. Specifically, the optimal learning rates identified for ViTS Dosovitskiy (2020) are directly applied to ResNet50 He et al. (2016) and Llama-60M Llama (2023). Under this setting, the core learning rate remains strictly fixed while auxiliary training strategies and weight decay are adapted to architecture-specific best practices. Furthermore, we conduct learning rate scaling experiments on Llama-60M Llama (2023) to comprehensively assess the hyperparameter robustness of different optimizers on language model architectures. Superior robustness. Muon Jordan et al. (2024) and MARS Yuan et al. (2025) achieve excellent results across ResNet He et al. (2016) and Llama Llama (2023). They are particularly effective in maintaining low and stable PPL (PPL ≈ 12-14) on Llama Llama (2023), even under extreme learning rate scaling. Muon Jordan et al. (2024)’s cross-architecture generalization stems from matrix-level orthogonalization in its algorithmic design. Unlike methods relying on scalar variance histories, it projects update matrices
51
Learning Rate Setting
10
—
Base LR×5
—
Base LR
—
Base LR/5
RAdam
12 SGD 11
8
8
6
10 500 10
10 Adam
10 RMSprop
2k
4k
6k
6 8k 9.1k 500
AdamP
12
2k
4k
6k
10
10 Adafactor
6
6
2
4 8k 9.1k 500
SGDP
10 AdaBelief
2k
4k
6k
AdamW
10
6
6
8
6
6
2
2
4
2
2
500 20
2k
4k
6k
Kron
8k 9.1k 500 10
2k
4k
6k
8k 9.1k
LAMB
10
15 10 5 500
2k
4k
6k
MADGRAD
500
2k
4k
6k
8k 9.1k 500
LAPROP
10
2k
4k
6k
8k 9.1k
LARS
10
6
6
6
2
2
2
2
2k 2k
4k 4k
6k 6k
8k 8k 9.1k
500
MARS
2k 2k
4k 4k
6k 6k
8k 8k 9.1k 500
Muon
2k
4k
6k
8k 9.1k
Nadam
10
10
6
6
6
6
2
2
2
2
2
500
2k
4k
10
6k
8k 9.1k 500
2k
4k
6k
8k 9.1k
500
2k
4k
6k
8k 9.1k 500
2k
10
4k
6k
8k 9.1k
4k
6k
8k 9.1k 500 10
2k
4k
6k
8k 9.1k
4k
6k
8k 9.1k
6k
8k 9.1k
6k
8k 9.1k
ADOPT
6 2 2k
4k
6k
Lion
8k 9.1k
500 12
2k
Lookahead
8 4
500
6
10
2k
Adan
500
6
8k 9.1k 500
2
8k 9.1k 500
2k
4k
6k
NadamW
500
8k 9.1k 500 14
2k
4k
NovoGrad
10 6 2 2k
4k
6k
8k 9.1k 500
2k
4k
Figure 9 Optimizers’ performance on Llama-60m. Subfigures show validation loss curves of various optimizers across three learning rate settings. The robustness to learning rate variations differs significantly among methods, with some optimizers (e.g., Muon Jordan et al. (2024), MARS Yuan et al. (2025)) maintaining stable convergence across all scales while others (e.g., SGD Robbins and Monro (1951), SGDP Heo et al. (2021)) rapidly diverge. Missing data indicates gradient explosion (loss becomes NaN) at the initial step or the point of discontinuation.
gles to converge on Llama Llama (2023). The core mechanism behind this generalization failure lies in gradient outliers common in large language model training. Relying on highly biased scalar variance histories causes adaptive methods like Adam Kingma (2015) to suffer from second-moment contamination by extreme gradients, leading the optimizer to incorrectly scale per-dimension step sizes during parameter updates and deviate from valid optimization trajectories in high-dimensional nonconvex spaces. Training collapse. While SGD series optimizers are viable in vision architectures, they face catastrophic training collapse (as shown in Fig. 9) when applied to the Llama Llama (2023) architecture. The extreme depth and complex attention mechanisms of large language models result in massive disparities in gradient magnitudes and parameter scales across different layers. Lacking adaptive parameter-wise or layerwise scaling mechanisms, pure first-order momentum methods fail to simultaneously satisfy the convergence requirements of all layers. Consequently, a uniform global learning rate triggers exploding gradients in certain layers while inducing vanishing gradients in others, ultimately leading to the complete loss of the model’s representational capacity.
protocol to 300 epochs on both ViT-S Dosovitskiy (2020) and ResNet-50 He et al. (2016) models. We detail the Top-1 accuracy at both training milestones and quantify the resulting performance improvements (∆) in Tab. 6, with the overall convergence trajectories visualized (Fig. 11), and comprehensive optimizer performance profiles illustrated through radar charts (Fig. 10). Architecture-specific scaling. The SGD series exhibits strong scalability, with SGD Robbins and Monro (1951) gaining 9.41% on ViT-S Dosovitskiy (2020). Mechanistically, these optimizers rely on first-order momentum without aggressive elementwise variance accumulation. This mathematical design prevents rapid step-size decay; while it leads to slower initial convergence, it enables the model to extensively explore the loss landscape extensively. Consistent scalability. When training is extended to 300 epochs, advanced optimizers including Muon Jordan et al. (2024), Lion Chen et al. (2023), Kron Li (2017) and AdamW Loshchilov and Hutter (2019) exhibit consistent yet modest accuracy gains, ranging from +1.896% to +4.690%, across various architectures. For instance, Muon Jordan et al. (2024) achieves a high baseline of 78.252 on ViT-S Dosovitskiy (2020) at 100 epochs, with its 300-epoch gain capped at +2.622%. This phenomenon occurs because these methods leverage sophisticated preconditioning, orthogonalization, or decoupled weight decay to rapidly traverse the optimization trajectory. By
4.3.3 Long-term Training Scalability To thoroughly investigate the long-term scalability of these optimizers, we extend the 100-epoch training
52
ViT-S
Muon Kron Lion Adafactor
ResNet
AdamP NadamW AdamW LAPROP LAMB MADGRAD MARS Adam Adan RAdam
AdaBelief
Nadam AdaBelief SGDP NovoGrad ADOPT SGD Lookahead LARS RMSprop
AdaBelief
Figure 10 Performance Comparison of Various Optimizers. We illustrate the ranking shifts of various optimizers on ViT-S Dosovitskiy (2020) and ResNet-50 He et al. (2016) after 100 and 300 training epochs.
efficiently exploiting local curvature and structural priors, they converge toward high quality’s optima within the initial 100 epochs. Consequently, further training iterations yield diminishing returns, as these optimizers have already extracted the vast majority of the network’s representational capacity during the early stages. Performance regression. Conversely, RMSprop Tieleman and Hinton (2012) exhibits a detrimental response to prolonged training on ResNet-50 He et al. (2016), experiencing a performance regression between 100 and 300 epochs. This degradation underscores the structural vulnerability of relying on an uncorrected EMA of squared gradients over protracted training cycles. Lacking the stabilizing effect of firstorder momentum or decoupled regularization mechanisms, the continuous accumulation of late-stage gradient noise disproportionately distorts parameterwise step sizes. Over the course of 300 epochs, this imbalanced scaling disrupts the fragile equilibrium within local minima, forcing the optimizer to escape optimal basins and leading to a decline in representational capacity in the late stage.
an optimizer escapes suboptimal regions or enters rapid descent phases. By stripping away the cumulative scale of the loss, the differencing method reveals the intrinsic rhythm and pacing of the underlying update mechanisms. Architectural influence on convergence homogeneity The empirical correlation matrices reveal that model architecture exerts a profound influence on the optimization trajectory, occasionally superseding the optimizer’s algorithmic design. In the ViT Dosovitskiy (2020) experiments, the correlation matrix displays overwhelming homogeneity across all evaluated optimizers. This phenomenon indicates that the highly constrained training regime of ViT Dosovitskiy (2020), which heavily relies on prolonged linear learning rate warmup and stringent regularization, dictates the loss landscape traversal. Consequently, the scheduling strategy effectively overrides the idiosyncratic step-size adaptations of individual optimizers. Conversely, the ResNet-50 He et al. (2016) experiments exhibit significant algorithmic bifurcation. The inherent inductive bias and smoother loss landscape of CNNs grant optimizers the geometric freedom to explore distinct descent paths, resulting in highly variable inter-optimizer correlations. Algorithmic categorization and trajectory divergence It should be noted that in all presented correlation heatmaps, the optimizers are arranged from left to right along the horizontal axis and from top to bottom along the vertical axis, strictly following the sequence listed in Fig. 11. Within the less constrained ResNet-50 He et al. (2016) environment, these heatmaps reveal distinct optimizer taxonomy clusters based on their dynamic behavior. Adaptive family. Methods such as Adam Kingma (2015), RAdam Liu et al. (2020a), AdamP Heo et al. (2021), AdamW Loshchilov and Hutter (2019), Nadam Dozat
4.3.4 Correlation of Optimizers We present a correlation analysis (Fig. 11) of optimizer performance on ViT-S Dosovitskiy (2020) and ResNet-50 He et al. (2016). To accurately capture the dynamic behavior of various optimizers, we analyze the first-order differencing of the validation loss rather than the raw loss values. Raw loss trajectories often exhibit high collinearity due to the overarching trend of network convergence. In contrast, first-order differencing isolates the step-wise acceleration and deceleration of the loss function. This approach effectively highlights the temporal inflection points where
53
4.3.5
Comprehensive Optimizer Evaluation.
A rigorous empirical comparison of more than twenty optimizers is detailed in Tab. 7, which employs a multifaceted star evaluation paradigm across five essential criteria to clearly demonstrate that an increased star count reflects enhanced algorithmic proficiency. Specifically, the star ratings indicate relative performance percentiles among the evaluated
Table 6 Results on ViT-S Dosovitskiy (2020) and ResNet50 He et al. (2016) models across various optimizers, where we compare the 100 epoch accuracy and 300 epoch accuracy and show that ViT-S Dosovitskiy (2020) achieves larger performance improvements from extended training than ResNet-50 He et al. (2016). 100 Epoch 300 Epoch Improvement Acc (%) Acc (%) (∆)
ViT-S
Model Optimizer SGD Robbins and Monro (1951) RMSprop Tieleman and Hinton (2012) Adam Kingma (2015) AdaBelief Zhuang et al. (2020) Adafactor Shazeer and Stern (2018) RAdam Liu et al. (2020a) AdamP Heo et al. (2021) SGDP Heo et al. (2021) AdamW Loshchilov and Hutter (2019) Adan Xie et al. (2024b) ADOPT Taniguchi et al. (2024) Kron Li (2017) LAMB You et al. (2020) LAPROP Ziyin et al. (2020) LARS You et al. (2017) Lion Chen et al. (2023) Lookahead Zhang et al. (2019) MADGRAD Defazio and Jelassi (2022) MARS Yuan et al. (2025) Muon Jordan et al. (2024) Nadam Dozat (2016) NadamW Wightman (2019) NovoGrad Ginsburg et al. (2019)
64.360 54.254 72.240 68.642 74.630 71.190 75.444 66.536 75.196 75.190 66.552 77.304 75.858 75.050 58.276 76.334 64.320 74.634 74.608 78.252 72.180 75.278 66.514
73.774 57.090 77.274 74.640 80.108 77.032 79.916 74.620 79.886 77.238 74.058 80.510 79.672 79.850 70.530 80.258 72.220 79.080 78.800 80.874 76.978 79.912 74.616
+ 9.414 + 2.836 + 5.034 + 5.998 + 5.478 + 5.842 + 4.472 + 8.084 + 4.690 + 2.048 + 7.506 + 3.206 + 3.814 + 4.800 + 12.254 + 3.924 + 7.900 + 4.446 + 4.192 + 2.622 + 4.798 + 4.634 + 8.102
ResNet-50
(2016), and NadamW Wightman (2019) form a tightly correlated cluster. Their shared reliance on the exponential moving average of squared gradients yields synchronized adaptation to landscape steepness, causing their rapid descent phases to align temporally. Preconditioning and orthogonalization. Notably, Kron Li (2017) and Muon Jordan et al. (2024) exhibit an exceptionally high intra-group correlation while remaining distinctly divergent from the adaptive scalar family. This empirical divergence perfectly mirrors their theoretical foundations. By leveraging structural gradient statistics for preconditioning or employing matrix orthogonalization techniques, these methods navigate the parameter space via structural matrix updates rather than simple diagonal scaling, resulting in fundamentally different loss traversal trajectories. Mechanistic outliers. Optimizers with structural deviations, such as Lookahead Zhang et al. (2019), MADGRAD Defazio and Jelassi (2022), and SGDP Heo et al. (2021), present consistently low correlations with standard methods. For instance, the dual-weight interpolation mechanism of Lookahead Zhang et al. (2019) introduces a temporal lag in the first-order difference sequence, shifting its inflection points relative to standard step-based methods. Temporal saturation and cross-epoch homogenization. Comparing short-term and long-term training regimes reveals a clear temporal smoothing effect. During the initial training phase, the exploratory divergence among optimizers is maximized, highlighting their unique mathematical characteristics. However, as training extends from 100 to 300 epochs, the correlation matrices demonstrate an overall increase in homogeneity. This shift occurs because the majority of optimizers eventually converge into a localized basin of attraction. In this exploitation phase, learning rates decay to minimal values, and the first-order differencing sequences are dominated by near-zero fluctuations. This asymptotic saturation mathematically forces the correlation coefficients to converge, illustrating that diverse algorithmic paths ultimately lead to similar geometrical destinations over extended temporal horizons.
SGD Robbins and Monro (1951) RMSprop Tieleman and Hinton (2012) Adam Kingma (2015) AdaBelief Zhuang et al. (2020) Adafactor Shazeer and Stern (2018) RAdam Liu et al. (2020a) AdamP Heo et al. (2021) SGDP Heo et al. (2021) AdamW Loshchilov and Hutter (2019) Adan Xie et al. (2024b) ADOPT Taniguchi et al. (2024) Kron Li (2017) LAMB You et al. (2020) LAPROP Ziyin et al. (2020) LARS You et al. (2017) Lion Chen et al. (2023) Lookahead Zhang et al. (2019) MADGRAD Defazio and Jelassi (2022) MARS Yuan et al. (2025) Muon Jordan et al. (2024) Nadam Dozat (2016) NadamW Wightman (2019) NovoGrad Ginsburg et al. (2019)
75.926 71.536 74.480 74.232 76.808 74.612 76.818 74.066 76.314 76.770 74.584 75.226 77.590 76.146 76.196 75.582 76.872 76.940 77.294 77.996 74.732 76.514 71.694
78.726 71.152 77.274 76.914 78.940 77.252 79.252 78.534 78.946 77.818 76.818 77.216 79.404 79.066 79.124 78.616 78.610 79.028 78.478 79.892 77.226 79.076 77.472
+ 2.800 - 0.384 + 2.794 + 2.682 + 2.132 + 2.640 + 2.434 + 4.468 + 2.632 + 1.048 + 2.234 + 1.990 + 1.814 + 2.920 + 2.928 + 3.034 + 1.738 + 2.088 + 1.184 + 1.896 + 2.494 + 2.562 + 5.778
methods. The quantitative basis for these ratings is grounded in specific empirical evaluations: accuracy aggregates the comprehensive performance across three different architectures (see Tab. 5); generalization is demonstrated through cross-architecture generalization experiments (see Sec. 4.3.2); and hyperparameter robustness is evaluated via learning rate scaling experiments (see Tab. 4). Additionally, convergence speed and memory efficiency are quantified by the iteration steps to reach a target loss and the peak GPU memory footprint, respectively. Overall, Muon Jordan et al. (2024) achieves the best comprehensive performance. 4.3.6
Computational cost disclosure.
To guarantee the internal validity and reproducibility of our findings, we enforce a strict grid-search protocol for hyperparameters across all 23 evaluated optimizers. However, this rigorous methodology induces an exponential explosion in computational overhead. As
54
—SGD
—RMSprop
—Adan
—Adam
—ADOPT
—AdaBelief —Kron
—Adafactor
—MADGRAD
—AdamP
—LAMB
—SGDP
—LARS
ViT-S-100epoch
80
—AdamW
—MARS
—Lookahead
—Muon
—Nadam
—NadamW —NovoGrad —LAPROP
ViT-S-300epoch
80
60
—RAdam
—Lion
60 1.0
40
1.0
40
0.8
0.8
0.6
Correlation Heatmap of Optimizers on ViT (100 epochs)
20
0.6
0.4
Correlation Heatmap of Optimizers on ViT (300 epochs)
20
0.4
0.2
0.2
0.0
80
0
20
40
60
80
0.0
100
ResNet-50-100epoch
60
80
0
50
100
150
200
250
300
ResNet-50-300epoch
60
1.0
40
Correlation Heatmap of Optimizers on ResNet (100 epochs)
20
1.0
40
0.8
0.8
0.6
0.6
0.4
Correlation Heatmap of Optimizers on ResNet (300 epochs)
20
0.2
0.4 0.2
0.0
0
20
40
60
80
0.0
100
0
50
100
150
200
250
300
Figure 11 Top-1 accuracy and trajectory correlation of popular optimizers. The main line plots display the Top-1 accuracy performance of 23 optimizers on ViT-S Dosovitskiy (2020) and ResNet-50 He et al. (2016) models over 100 and 300 epochs. The inset heatmaps present the Pearson correlation matrices calculated via first-order differencing of the validation metrics, revealing the algorithmic categorization and trajectory divergence across different architectures and training horizons. Table 7 Comparison of various optimizers across five key dimensions. We evaluate convergence speed, accuracy, memory efficiency, generalization, and hyperparameter robustness, and show that Muon achieves the best overall performance. Convergence Speed SGD Robbins and Monro (1951) ★ RMSprop Tieleman and Hinton (2012) ★ Adam Kingma (2015) ★★★ AdaBelief Zhuang et al. (2020) ★★★★ Adafactor Shazeer and Stern (2018) ★★ RAdam Liu et al. (2020a) ★★★★ AdamP Heo et al. (2021) ★★★ SGDP Heo et al. (2021) ★★★ AdamW Loshchilov and Hutter (2019) ★★★★ Adan Xie et al. (2024b) ★★★★ ADOPT Taniguchi et al. (2024) ★★★ Kron Li (2017) ★★★★★ LAMB You et al. (2020) ★★★★ LAPROP Ziyin et al. (2020) ★★★ LARS You et al. (2017) ★ Lion Chen et al. (2023) ★★★★★ Lookahead Zhang et al. (2019) ★ MADGRAD Defazio and Jelassi (2022) ★★★ MARS Yuan et al. (2025) ★★★★ Muon Jordan et al. (2024) ★★★★★ Nadam Dozat (2016) ★★★ NadamW Wightman (2019) ★★★★ NovoGrad Ginsburg et al. (2019) ★★ Optimizer
Accuracy ★★ ★ ★★★ ★★ ★★★ ★★★ ★★★★ ★★ ★★★★ ★★★★ ★★ ★★★★★ ★★★★ ★★★★ ★ ★★★★★ ★★ ★★★ ★★★ ★★★★★ ★★★ ★★★★ ★★
Memory Efficiency ★★★★★ ★★★★ ★★★ ★★★ ★★★★ ★★★ ★★★ ★★★★ ★★★ ★★ ★★★ ★★ ★★★ ★★★ ★★★★ ★★★★ ★★ ★★★ ★★★★ ★★★★ ★★★ ★★★ ★★★★
Generalization ★★ ★★ ★★ ★★ ★★★ ★★ ★★★ ★★ ★★★★ ★★★★★ ★★ ★★★★ ★★★★ ★★★ ★★ ★★★★ ★★ ★★ ★★★★★ ★★★★★ ★★ ★★★★ ★★
development directions for optimization algorithms. Challenges in current optimization frameworks. Optimizing large models is fundamentally a highly constrained, multi objective trade-off. When current methods pursue specific metrics, such as convergence speed, memory efficiency, or distributed scaling, they often sacrifice generalization, introduce computational latency, or exacerbate system instability. Therefore, a core challenge for existing optimization algorithms is to break these interconnected bottlenecks under a unified design framework, effectively balancing stability, computational cost, and noise robustness.
Hyperparameter Robustness ★ ★★ ★★★ ★★ ★★★ ★★★ ★★★ ★ ★★★ ★★★★★ ★★★ ★★★★★ ★★ ★★★ ★★ ★★★★ ★★ ★★★★ ★★★★ ★★★★★ ★★★ ★★★ ★★
Generalization and stability margins. Current adaptive methods frequently trap models in sharp minima, resulting in worse generalization than standard SGD Robbins and Monro (1951). While preconditioned approaches use structural approximations to improve convergence, they exhibit a narrow stability margin where large learning rates cause training divergence Li (2017). Furthermore, stateless memory-efficient designs lose the historical accumulation of second moments, removing their ability to average extreme gradient shocks and risking training divergence in non-convex loss landscapes Shazeer and Stern (2018).
detailed in Tab. 8, the complete execution of our empirical framework, encompassing multi-scale learning rate sweeps and extended 300-epoch horizons, amasses an estimated footprint of over 1073 single NVIDIA A100 hours.
5
Future Prospect
We summarize the current challenges in the field of optimization. Based on challenges and the preceding experimental analysis, we explore the potential
Memory and computational overheads. As models scale to billions of parameters, dense optimizer 55
Table 8 Comprehensive estimation of computational cost (GPU Hours). The table details the estimated single NVIDIA A100 GPU hours required for the entire empirical benchmark. The workload accounts for 23 optimizers, including rigorous grid searches across 5 learning rate scales and extended 300-epoch training horizons. The massive total computational footprint (∼1073 GPU Hours) underscores the rigorousness of our evaluation and highlights the fundamental computational barrier in modern optimizer benchmarking. Benchmark Task
Hardware Type
ResNet-50 He et al. (2016) NVIDIA A100 (ImageNet-1K Hudişteanu et al. (2025)) ViT-S Dosovitskiy (2020) NVIDIA A100 (ImageNet-1K Hudişteanu et al. (2025)) LlaMA-60M Llama (2023) NVIDIA A100 (WikiText-103 Merity et al. (2016)) Total Aggregated Cost
Hyperparameter Search Final Evaluation Total GPU Hours (5 LR Scales × 23 Optimizers) (Best LR × 23 Optimizers) -
∼288 Hours(100/300 epochs) ∼288 Hours
∼544 Hours (100 epochs)
∼213 Hours (300 epochs)
∼757 Hours
∼28 Hours (Chinchilla scaling)
-
∼28 Hours
Equivalent to ∼45 Days on a single NVIDIA A100 GPU ∼1073 Hours
states create a primary memory bottleneck Rajbhandari et al. (2020). Structural approximations attempt to balance precision and memory overhead, but matrix inversions and structural overhead significantly decrease global efficiency Martens and Grosse (2015). Although certain methods theoretically require fewer iterations to converge, their per-step computational latency often negates these gains, limiting their application in large-scale settings.
automated generation of architecture-specific optimizers, enabling models to inherently navigate complex loss landscapes without manual intervention. 2) Preconditioning and orthogonalization: Given that methods like Kron Li (2017) and Muon Jordan et al. (2024) exhibit fundamentally different loss traversal trajectories compared to the adaptive scalar family, future research should focus on advancing structural matrix updates. By leveraging structural gradient statistics for preconditioning or employing matrix orthogonalization techniques, these approaches overcome the representational bottlenecks of simple diagonal scaling, opening up novel and highly efficient pathways for navigating the parameter space. 3) Dynamic precision scaling: To alleviate the severe memory bottlenecks imposed by dense optimizer states, future research should integrate adaptive lowprecision arithmetic that significantly reduces memory footprints without compromising convergence stability.
Noise amplification and estimation variance. In the anisotropic loss landscapes typical of deep attention networks, the stochastic noise from mini-batch training is heavily amplified during curvature estimation, reducing update stability Liu et al. (2024a). Similarly, random perturbations Maryak and Chin (2001) used to estimate directional gradients suffer from approximation variance that increases rapidly with parameter dimensionality. Privacy-preserving optimizers face a parallel utility tradeoff, as the required noise injection for differential privacy degrades gradient fidelity and destabilizes convergence Chen et al. (2020).
Second-order optimization. To address these prohibitive bottlenecks, future development of secondorder frameworks should emphasize structural innovation across three key areas: 1) Hardware-algorithm co-design: To overcome emerging memory bottlenecks, research should pivot toward optimizing global computational efficiency by integrating structureaware adaptation and low-precision arithmetic to balance computational intensity with convergence stability. 2) Robust curvature estimation: Future designs should develop mathematically sound mechanisms to isolate extreme gradient outliers, expanding the dangerously narrow safety margin without inflating memory footprints. 3) Sparse preconditioning: To mitigate prohibitive per-step computational latency, future architectures should explore highly efficient sparse matrix inversion techniques that preserve the intrinsic advantages of second-order information while achieving global efficiency comparable to first-
Scalability and consensus tensions. Optimization trajectories remain highly sensitive to hyperparameters, requiring extensive manual tuning for successful scaling. In distributed settings, combining gradient compression with local updates increases client drift Basu et al. (2020), creating a persistent tension between communication efficiency and global consensus. Additionally, optimizing within low-rank subspaces to manage high-dimensional search spaces can prematurely constrain the trajectory and prevent access to global optima Lialin et al. (2023). First-order optimization. To address these limitations, future development of first-order optimization could focus on three key directions: 1) Automated symbolic discovery: The methodological focus should shift from fragile heuristic tuning to the
56
order baselines.
sentational limits of simple diagonal scaling. Methodologically, the focus should shift from heuristic tuning and static noise injection to automated symbolic discovery and exact noise cancellation. Combined with the geometric corrections of preconditioning and orthogonalization, these advancements enable architecture-specific optimizers that handle highly anisotropic loss landscapes and resist extreme gradient outliers. In distributed and privacy-preserving settings, frameworks must also advance beyond empirical trade-offs. Establishing theoretical lower bounds for communication and utility, along with topologyaware consensus mechanisms, will ensure the scalable and reliable deployment of these specialized architectures.
Zeroth-order optimization. To bridge this substantial performance gap, the future development of gradient-free frameworks must focus on noise and space management. 1) Exact noise cancellation: Future frameworks must mathematically cancel the inherent stochastic noise introduced by diverse network architectures, potentially drawing inspiration from exact gradient correction mechanisms that subtract historical perturbation bias to stabilize the highly variable update steps. 2) Dimensionality robustness: Research should discover advanced subspace projection methods that safely constrain the search space without prematurely restricting access to high quality global solutions. 3) Adaptive perturbation scheduling: Rather than relying on static random perturbations, future strategies should dynamically adjust the variance and direction of blind estimates based on local landscape geometry, thereby preventing continuous disruption of the fragile equilibrium within local minima.
6
Conclusion
This survey presents a comprehensive, systematic review of the latest advancements in deep learning optimization methods. We first establish a rigorous theoretical foundation by reviewing core background concepts in optimization theory, providing a unified framework for understanding the evolution of optimization algorithms. We then systematically categorize and analyze technical approaches across four main domains: first-order, second-order, and zeroth-order optimization algorithms, as well as specialized scenario-oriented optimization frameworks. Our analysis examines both the structural design principles and functional characteristics of each class of optimization methods. To provide practical, actionable insights, we conduct a rigorous, controlled empirical evaluation that benchmarks representative optimizers across diverse task domains and model architectures. Our experimental analysis reveals critical trade-offs between hyperparameter sensitivity, cross-architecture generalization capability, and longterm training scalability, offering valuable guidance for algorithm selection and hyperparameter tuning in real-world applications. Finally, we identify key open challenges in the field and outline promising future research directions to guide the development of next-generation high-efficiency, robust, and trustworthy optimization technologies. By synthesizing theoretical insights with extensive empirical evidence, this survey aims to serve as a comprehensive, authoritative resource for researchers and practitioners seeking to advance the state-of-the-art in deep learning optimization.
Scenario-oriented frameworks. To ensure scalable and trustworthy deployment in critical infrastructure, future development of these specialized frameworks must shift from empirical compromises to rigorous guarantees across three key dimensions: 1) Rigorous theoretical bounds: Distributed and privacy preserving frameworks must establish strict mathematical lower bounds for communication efficiency and privacy utility. 2) State fidelity preservation: Future memory efficient architectures must be designed to retain just enough historical context to smoothly average out extreme gradient shocks without hitting the rigid fidelity ceiling or exceeding memory constraints. 3) Heterogeneous consensus mechanisms: To resolve the persistent tension between communication efficiency and client drift, distributed setups must introduce topology aware gradient compression techniques that naturally accommodate the massive disparities in gradient magnitudes across diverse network architectures. Discussion. Future optimization frameworks will likely move beyond isolated algorithmic improvements by integrating FO, SO, and ZO algorithms. To address current memory bottlenecks and structural compromises, research must shift from minimizing iteration complexity to improving global wall-clock efficiency. This transition requires hardware-algorithm co-design. Specifically, structural preconditioning and matrix orthogonalization are critical. Integrating these techniques with dynamic low-precision arithmetic balances computational cost and convergence stability, allowing optimizers to overcome the repre-
57
tion without computational overhead. arXiv preprint arXiv:2401.12033, 2024. 19
References Martin Abadi, Andy Chu, Ian Goodfellow, H Brendan McMahan, Ilya Mironov, Kunal Talwar, and Li Zhang. Deep learning with differential privacy. In CCS, 2016. 4
Stefania Bellavia, Serge Gratton, Benedetta Morini, and Ph L Toint. Fast stochastic second-order adagrad for nonconvex bound-constrained optimization. arXiv preprint arXiv:2505.06374, 2025. 20
Ahmed M Adly. Exadam: The power of adaptive crossmoments. arXiv preprint arXiv:2412.20302, 2024. 13
Irwan Bello, Barret Zoph, Vijay Vasudevan, and Quoc V Le. Neural optimizer search with reinforcement learning. In ICML, 2017. 47
Aayushya Agarwal, Larry Pileggi, and Ronald Rohrer. Second-order optimization via quiescence. arXiv preprint arXiv:2410.08033, 2024. 24 Shun-Ichi Amari. Natural gradient works efficiently in learning. Neural computation, 1998. 1, 22, 23, 24
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Animashree Anandkumar. signsgd: Compressed optimisation for non-convex problems. In ICML, 2018. 34
Kang An, Yuxing Liu, Rui Pan, Yi Ren, Shiqian Ma, Donald Goldfarb, and Tong Zhang. Asgo: Adaptive structured gradient optimization. arXiv preprint arXiv:2503.20762, 2025. 7, 9, 14
Ekaterina Borodich and Dmitry Kovalev. Nesterov finds graal: Optimal and adaptive gradient method for convex optimization. arXiv preprint arXiv:2507.09823, 2025. 21
Wangpeng An, Haoqian Wang, Qingyun Sun, Jun Xu, Qionghai Dai, and Lei Zhang. A pid controller approach for stochastic optimization of deep networks. In CVPR, 2018. 10
Léon Bottou. Large-scale machine learning with stochastic gradient descent. In COMPSTAT, 2010. 3 Zhiqi Bu, Sivakanth Gopi, Janardhan Kulkarni, Yin Tat Lee, Hanwen Shen, and Uthaipon Tantipongpipat. Fast and memory efficient differentially private-sgd via jl projections. In NeurIPS, 2021. 41
Maksym Andriushchenko, Aditya Vardhan Varre, Loucas Pillaud-Vivien, and Nicolas Flammarion. Sgd with large step sizes learns sparse features. In ICML, 2023. 17
Richard H Byrd, Peihuang Lu, Jorge Nocedal, and Ciyou Zhu. A limited memory algorithm for bound constrained optimization. SIAM Journal on scientific computing, 1995. 1, 23, 25
Imen Ayadi and Gabriel Turinici. Stochastic runge-kutta methods and adaptive sgd-g2 stochastic gradient descent. In ICPR, 2021. 17
Xin Cao. Bfe and adabfe: A new approach in learning rate automation for stochastic optimization. arXiv preprint arXiv:2207.02763, 2022. 17, 46
Jiyang Bai, Yuxiang Ren, and Jiawei Zhang. Deam: adaptive momentum with discriminative weight for stochastic optimization. In ASONAM, 2020. 10
Yang Cao, Xiaoyu Li, and Zhao Song. Grams: Gradient descent with adaptive momentum scaling. In ICLR Workshop, 2025. 20
Jiyang Bai, Yuxiang Ren, and Jiawei Zhang. Bgadam: Boosting based genetic-evolutionary adam for neural network optimization. In IJCNN, 2021. 47
André Carlon, Luis Espath, and Raul Tempone. Efficient stochastic bfgs methods inspired by bayesian principles. arXiv preprint arXiv:2507.07729, 2025. 23, 26
Lukas Balles and Philipp Hennig. Dissecting adam: The sign, magnitude and variance of stochastic gradients. In ICML, 2018. 12
Camille Castera, Jérôme Bolte, Cédric Févotte, and Edouard Pauwels. Second-order step-size tuning of sgd for non-convex optimization. Neural Process. Lett., 2022. 17
Noga Bar and Raja Giryes. Zoqo: Zero-order quantized optimization. In ICASSP, 2025. 31 Leighton Pate Barnes, Huseyin A Inan, Berivan Isik, and Ayfer Özgür. rtop-k: A statistical estimation approach to distributed sgd. JSAIT, 2020. 34
Augustin Cauchy et al. Méthode générale pour la résolution des systemes d’équations simultanées. Comp. Rend. Sci. Paris, 1847. 3, 4, 5, 6
Debraj Basu, Deepesh Data, Can Karakus, and Suhas N Diggavi. Qsparse-local-sgd: Distributed sgd with quantization, sparsification, and local computations. JSAIT, 2020. 40, 56
Evan Chen, Shiqiang Wang, Jianing Zhang, Dong-Jun Han, Chaoyue Liu, and Christopher Brinton. Gradient correction in federated learning with adaptive optimization. arXiv preprint arXiv:2502.02727, 2025a. 36, 38, 47
Samiksha BC. Zeta: A riemann zeta-scaled extension of adam for deep learning. arXiv preprint arXiv:2508.02719, 2025. 21
Jiahe Chen and Ziye Ma. Vamo: Efficient large-scale nonconvex optimization via adaptive zeroth order variance reduction. arXiv preprint arXiv:2505.13954, 2025. 30, 32
Marlon Becker, Frederick Altrock, and Benjamin Risse. Momentum-sam: Sharpness aware minimiza-
58
Song Chen, Jiaxu Liu, Pengkai Wang, Chao Xu, Shengze Cai, and Jian Chu. Accelerated optimization in deep learning with a proportional-integral-derivative controller. Nature Communications, 2024. 10
Aditya Devarakonda and Ramakrishnan Kannan. Communication-efficient, 2d parallel stochastic gradient descent for distributed-memory optimization. arXiv preprint arXiv:2501.07526, 2025. 36
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. Training deep nets with sublinear memory cost. arXiv, 2016. 4
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In NAACL, 2019. 13
Xiangning Chen, Chen Liang, Da Huang, Esteban Real, Kaiyuan Wang, Hieu Pham, Xuanyi Dong, Thang Luong, Cho-Jui Hsieh, Yifeng Lu, et al. Symbolic discovery of optimization algorithms. In NeruIPS, 2023. 5, 9, 12, 16, 47, 49, 51, 52, 54, 55
Nalin Dhiman. Fanos: Friction-adaptive nos\’e–hoover symplectic momentum for stiff objectives. arXiv preprint arXiv:2601.00889, 2025. 10 Yucheng Ding, Chaoyue Niu, Yikai Yan, Zhenzhe Zheng, Fan Wu, Guihai Chen, Shaojie Tang, and Rongfei Jia. Distributed optimization over block-cyclic data. In MMAsia, 2024. 38
Xiangyi Chen, Sijia Liu, Kaidi Xu, Xingguo Li, Xue Lin, Mingyi Hong, and David Cox. Zo-adamm: Zeroth-order adaptive momentum method for black-box optimization. In NeurIPS, 2019. 11, 28
Alexey Dosovitskiy. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. 1, 48, 49, 51, 52, 53, 54, 55, 56
Xiangyi Chen, Steven Z Wu, and Mingyi Hong. Understanding gradient clipping in private sgd: A geometric perspective. NeurIPS, 2020. 56 Yiming Chen, Yuan Zhang, Liyuan Cao, Kun Yuan, and Zaiwen Wen. Enhancing zeroth-order fine-tuning for language models with low-rank structures. In ICLR, 2025b. 5, 7, 11, 12, 31
Timothy Dozat. Incorporating nesterov momentum into adam. ICLR Workshop, 2016. 9, 12, 49, 51, 53, 54, 55 Jiawei Du, Hanshu Yan, Jiashi Feng, Joey Tianyi Zhou, Liangli Zhen, Rick Siow Mong Goh, and Vincent Tan. Efficient sharpness-aware minimization for improved training of neural networks. In ICLR, 2022. 18
Yifei Cheng, Li Shen, Hao Sun, Nan Yin, Xiaochun Cao, and Enhong Chen. Lightsam: Parameteragnostic sharpness-aware minimization. arXiv preprint arXiv:2505.24399, 2025. 18
Jiawei Duan, Haibo Hu, Qingqing Ye, and Xinyue Sun. Analyzing and optimizing perturbation of dp-sgd geometrically. In ICDE, 2025. 42
Ekin D Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V Le. Randaugment: Practical automated data augmentation with a reduced search space. In CVPRW, 2020. 48
Shiv Ram Dubey, Soumendu Chakraborty, Swalpa Kumar Roy, Snehasis Mukherjee, Satish Kumar Singh, and Bidyut Baran Chaudhuri. diffgrad: an optimization method for convolutional neural networks. IEEE TNNLS, 2019. 13
Ashok Cutkosky and Harsh Mehta. Momentum improves normalized sgd. In ICML, 2020. 16 Tehila Dahan and Kfir Yehuda Levy. Do stochastic, feel noiseless: Stable stochastic optimization via a double momentum mechanism. In ICLR, 2025. 8
Shiv Ram Dubey, SH Shabbeer Basha, Satish Kumar Singh, and Bidyut Baran Chaudhuri. Adainject: Injection-based adaptive gradient descent optimizers for convolutional neural networks. ITAI, 2022. 12
Sizhe Dang, Yangyang Guo, Yanjun Zhao, Haishan Ye, Xiaodong Zheng, Guang Dai, and Ivor Tsang. Fzoo: Fast zeroth-order optimizer for fine-tuning large language models towards adam-scale speed. arXiv preprint arXiv:2506.09034, 2025. 28, 29, 32
Shiv Ram Dubey, Satish Kumar Singh, and Bidyut Baran Chaudhuri. Adanorm: Adaptive gradient norm correction based optimizer for cnns. In WACV, 2023. 11, 15, 16, 47
A Defazio and S Jelassi. Adaptivity without compromise: A momentumized, adaptive, dual averaged gradient method for stochastic optimization. JMLR, 2022. 8, 9, 20, 48, 49, 50, 51, 54, 55
John Duchi, Elad Hazan, and Yoram Singer. Adaptive subgradient methods for online learning and stochastic optimization. JMLR, 2011. 5, 6 Ofri Eisen, Ron Dorfman, and Kfir Yehuda Levy. Enhancing parallelism in decentralized stochastic convex optimization. In ICML, 2025. 37
Aaron Defazio, Xingyu Yang, Ahmed Khaled, Konstantin Mishchenko, Harsh Mehta, and Ashok Cutkosky. The road less scheduled. In NeurIPS, 2024. 17
Ahmed Elbakary, Chaouki Ben Issaid, Mohammad Shehab, Karim Seddik, Tamer ElBatt, and Mehdi Bennis. Fed-sophia: A communication-efficient second-order federated learning algorithm. In ICC, 2024. 24, 40
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, 2009. 48
59
Benjamin Ellis, Matthew T Jackson, Andrei Lupu, Alexander D Goldie, Mattie Fellows, Shimon Whiteson, and Jakob Foerster. Adam on local time: Addressing nonstationarity in rl with relative adam timesteps. In NeurIPS, 2024. 8
Tanmay Gautam, Youngsuk Park, Hao Zhou, Parameswaran Raman, and Wooseok Ha. Variancereduced zeroth-order methods for fine-tuning language models. In ICLR, 2024. 11, 28, 32 Rémi Genet and Hugo Inzirillo. Caadam: Improving adam optimizer using connection aware methods. arXiv preprint arXiv:2410.24216, 2024. 13
Jeffrey L Elman. Finding structure in time. Cognitive science, 1990. 47 Mohamed Elsayed, Homayoon Farrahi, Felix Dangel, and A. Rupam Mahmood. Revisiting scalable hessian diagonal approximations for applications in reinforcement learning. In ICML, 2024. 24
Boris Ginsburg, Patrice Castonguay, Oleksii Hrinchuk, Oleksii Kuchaiev, Vitaly Lavrukhin, Ryan Leary, Jason Li, Huyen Nguyen, Yang Zhang, and Jonathan M Cohen. Stochastic gradient methods with layer-wise adaptive moments for training of deep networks. arXiv preprint arXiv:1905.11286, 2019. 13, 49, 50, 51, 54, 55
Hannes Fassold. Adafamily: A family of adamlike adaptive gradient methods. arXiv preprint arXiv:2203.01603, 2022. 13
Donald Goldfarb, Yi Ren, and Achraf Bahamou. Practical quasi-newton methods for training deep neural networks. In NeurIPS, 2020. 26
Mingquan Feng, Yixin Huang, Yifan Fu, Shaobo Wang, and Junchi Yan. Ko: Kinetics-inspired neural optimizer with pde simulation approaches. arXiv preprint arXiv:2505.14777, 2025. 47
Damien MARTINS GOMES, Yanlei Zhang, Eugene Belilovsky, Guy Wolf, and Mahdi S. Hosseini. Adafisher: Adaptive second order optimization via fisher information. In ICLR, 2025. 25
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur. Sharpness-aware minimization for efficiently improving generalization. In ICLR, 2021. 11, 16, 18, 19, 21, 28, 42
Wenbo Gong, Meyer Scetbon, Chao Ma, and Edward Meeds. Towards efficient optimizer design for llm via structured fisher approximation with a low-rank extension. arXiv preprint arXiv:2502.07752, 2025. 25
Zachary Frangella, Pratik Rathore, Shipu Zhao, and Madeleine Udell. Sketchysgd: reliable stochastic optimization via randomized curvature estimates. SIAM, 2024. 24
Eduard Gorbunov, Marina Danilova, and Alexander Gasnikov. Stochastic optimization with heavy-tailed noise via accelerated gradient clipping. In NeurIPS, 2020. 15
Kevin Frans, Sergey Levine, and Pieter Abbeel. A stable whitening optimizer for efficient neural network training. arXiv preprint arXiv:2506.07254, 2025. 14
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. Accurate, large minibatch sgd: Training imagenet in 1 hour. arXiv, 2017. 4
Deng Fucheng, Wang Wanjie, Gong Ao, Wang Xiaoqi, and Wang Fan. Gradient descent algorithm survey. arXiv preprint arXiv:2511.20725, 2025. 1
Shangwei Guo, Tianwei Zhang, Guowen Xu, Han Yu, Tao Xiang, and Yang Liu. Topology-aware differential privacy for decentralized image classification. IEEE Trans, 2021. 33, 40, 41
Takumi Fujimoto and Hiroaki Nishi. eagle: early approximated gradient based learning rate estimator. arXiv preprint arXiv:2502.01036, 2025. 12 Kaixin Gao, Xiaolei Liu, Zhenghai Huang, Min Wang, Zidong Wang, Dachuan Xu, and Fan Yu. A tracerestricted kronecker-factored approximation to natural gradient. In AAAI, 2021. 23, 25
Sunny Gupta, Mohit Jindal, Pankhi Kashyap, Pranav Jeevan, and Amit Sethi. Flens: Federated learning with enhanced nesterov-newton sketch. In BigData, 2024. 38
Yuan Gao, Anton Rodomanov, Jeremy Rack, and Sebastian U Stich. Accelerated distributed optimization with compression and error feedback. arXiv preprint arXiv:2503.08427, 2025. 35
Vineet Gupta, Tomer Koren, and Yoram Singer. Shampoo: Preconditioned stochastic tensor optimization. In ICML, 2018. 7, 9, 11, 14 Mohamed Hassan, Aleksandar Vakanski, Boyu Zhang, and Min Xian. Gcsam: Gradient centralized sharpness aware minimization. arXiv preprint arXiv:2501.11584, 2025. 16
Matilde Gargiani, Andrea Zanelli, Moritz Diehl, and Frank Hutter. On the promise of the stochastic generalized gauss-newton method for training dnns. arXiv preprint arXiv:2006.02409, 2020. 23, 24
Guangxin He, Yuan Cao, Yutong He, Tianyi Bai, Kun Yuan, and Binhang Yuan. Tah-quant: Effective activation quantization in pipeline parallelism over slow network. arXiv preprint arXiv:2506.01352, 2025a. 34, 43
Nirmal Gaud, Surej Mouli, Preeti Katiyar, and Vaduguru Venkata Ramya. Comparative analysis of novel nirmal optimizer against adam and sgd with momentum. arXiv preprint arXiv:2508.04293, 2025. 21
60
Huan He, Shifan Zhao, Yuanzhe Xi, Joyce Ho, and Yousef Saad. GDA-AM: ON THE EFFECTIVENESS OF SOLVING MIN-IMAX OPTIMIZATION VIA ANDERSON MIXING. In ICLR, 2022. 21
How to train in 4-bit more stably than 16-bit adam. In ICLR, 2025a. 15, 16, 43 Tianjin Huang, Ziquan Zhu, Gaojie Jin, Lu Liu, Zhangyang Wang, and Shiwei Liu. SPAM: Spike-aware adam with momentum reset for stable LLM training. In ICLR, 2025b. 15, 45
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016. 1, 13, 48, 49, 51, 52, 53, 54, 55, 56
Xunpeng Huang, Runxin Xu, Hao Zhou, Zhe Wang, Zhengyang Liu, and Lei Li. Acmo: angle-calibrated moment methods for stochastic optimization. In AAAI, 2021. 17
Wei He, Kai Han, Hang Zhou, Hanting Chen, Zhicheng Liu, Xinghao Chen, and Yunhe Wang. Root: Robust orthogonalized optimizer for neural network training. arXiv preprint arXiv:2511.20626, 2025b. 12
Yangsibo Huang, Haotian Jiang, Daogao Liu, Mohammad Mahdian, Jieming Mao, and Vahab Mirrokni. Learning across data owners with joint differential privacy. arXiv preprint arXiv:2305.15723, 2023. 41
Byeongho Heo, Sanghyuk Chun, Seong Joon Oh, Dongyoon Han, Sangdoo Yun, Gyuwan Kim, Youngjung Uh, and Jung-Woo Ha. Adamp: Slowing down the slowdown for momentum optimizers on scale-invariant weights. In ICLR, 2021. 10, 20, 48, 49, 51, 52, 53, 54, 55
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, et al. Gpipe: Efficient training of giant neural networks using pipeline parallelism. NeurIPS, 2019b. 4
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. Training compute-optimal large language models. In NeurIPS, 2022. 48
Mihaela Hudişteanu, Edwige Cyffers, and Nikita P Kalinin. Dp-microadam: Private and frugal algorithm for training and fine-tuning. arXiv preprint arXiv:2511.20509, 2025. 41, 48, 56
Mahdi S Hosseini and Konstantinos N Plataniotis. Adas: Adaptive scheduling of stochastic gradients. arXiv preprint arXiv:2006.06587, 2020. 17
Dongseong Hwang. Fadam: Adam is a natural gradient optimizer using diagonal empirical fisher information. arXiv preprint arXiv:2405.12807, 2024. 12, 20
Rui Hu, Yuanxiong Guo, E Paul Ratazzi, and Yanmin Gong. Differentially private federated learning for resource-constrained internet of things. arXiv preprint arXiv:2003.12705, 2020. 36
Alex Iacob, Lorenzo Sani, Mher Safaryan, Paris Giampouras, Samuel Horváth, Andrej Jovanovic, Meghdad Kurmanji, Preslav Aleksandrov, William F Shen, Xinchi Qiu, et al. Des-loc: Desynced low communication adaptive optimizers for training foundation models. arXiv preprint arXiv:2505.22549, 2025. 37
Siyu Hu, Wentao Zhang, Qiuchen Sha, Feng Pan, LinWang Wang, Weile Jia, Guangming Tan, and Tong Zhao. Rlekf: An optimizer for deep potential with ab initio accuracy. In AAAI, 2023. 13
Arthur Jacot, Franck Gabriel, and Clément Hongler. Neural tangent kernel: Convergence and generalization in neural networks. NeurIPS, 2018. 4
Yue Hu, Zanxia Cao, and Yingchao Liu. Dimer-enhanced optimization: A first-order approach to escaping saddle points in neural network training. arXiv preprint arXiv:2507.19968, 2025. 19
Shuli Jiang, Pranay Sharma, Zhiwei Steven Wu, and Gauri Joshi. The cost of shuffling in private gradient based optimization. arXiv preprint arXiv:2502.03652, 2025a. 37, 41, 42
Feihu Huang, Guanyi Zhang, and Songcan Chen. Homeadam: Adam and adamw algorithms sometimes go home to obtain better provable generalization. arXiv preprint arXiv:2603.02649, 2026a. 12
Shuoran Jiang, Qingcai Chen, Youcheng Pan, Yang Xiang, Yukang Lin, Xiangping Wu, Chuanyi Liu, and Xiaobao Song. Zo-adamu optimizer: Adapting perturbation by the momentum and uncertainty in zeroth-order optimization. In AAAI, 2024. 28, 29, 30
Haiwen Huang, Chang Wang, and Bin Dong. Nostalgic adam: Weighting more of the past gradients when designing the adaptive learning rate. In IJCAI, 2019a. 12
Weisen Jiang, Hansi Yang, Yu Zhang, and James Kwok. An adaptive policy to employ sharpness-aware minimization. In ICLR, 2023. 18
Tao Huang, Jiayang Meng, Xu Yang, Chen Hou, and Hong Chen. Dp-aware adaln-zero: Taming conditioninginduced heavy-tailed gradients in differentially private diffusion. arXiv preprint arXiv:2602.22610, 2026b. 41
Zhanhong Jiang, Aditya Balu, Sin Yong Tan, Young M Lee, Chinmay Hegde, and Soumik Sarkar. On higherorder moments in adam. In NeurIPS, 2019. 12
Tianjin Huang, Haotian Hu, Zhenyu Zhang, Gaojie Jin, Xiang Li, Li Shen, Tianlong Chen, Lu Liu, Qingsong Wen, Zhangyang Wang, and Shiwei Liu. Stable-SPAM:
Zhanhong Jiang, Md Zahid Hasan, Aditya Balu, Joshua R Waite, Genyi Huang, and Soumik Sarkar. Fuse: First-
61
order and second-order unified synthesis in stochastic optimization. In CAI, 2025b. 26
future gradients in momentum-based stochastic optimization. arXiv preprint arXiv:2501.09556, 2025. 8
Chenhan Jin, Kaiwen Zhou, Bo Han, James Cheng, and Tieyong Zeng. Efficient private sco for heavy-tailed data via averaged clipping. Mach. Learn., 2024. 43
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. NeurIPS, 2012. 13
Jiachen Jin, Kangkang Deng, Boyu Wang, and Hongxia Wang. Stochastic admm with batch size adaptation for nonconvex nonsmooth optimization. arXiv preprint arXiv:2505.06921, 2025a. 36
Frederik Kunstner, Philipp Hennig, and Lukas Balles. Limitations of the empirical fisher approximation for natural gradient descent. NeurIPS, 2019. 7, 14 Anna Kuzina, Haotian Chen, Babak Esmaeili, and Jakub M. Tomczak. Variational stochastic gradient descent for deep neural networks. TMLR, 2025. 12
Long Jin, Han Nong, Liangming Chen, and Zhenming Su. A method for enhancing generalization of adam by multiple integrations. In AAAI, 2025b. 18
Kin Wai Lau, Yasar Abbas Ur Rehman, Pedro Porto Buarque de Gusmão, Lai-Man Po, Lan Ma, and Yuyang Xie. Fedrepopt: Gradient re-parametrized optimizers in federated learning. In ACCV, 2024. 38
Junhyuk Jo, Jihyun Lim, and Sunwoo Lee. Asynchronous sharpness-aware minimization for fast and accurate deep learning. arXiv preprint arXiv:2503.11147, 2025. 18
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 2002. 2, 48
John St John. Adamd: Improved bias-correction in adam. arXiv preprint arXiv:2110.10828, 2021. 12 Keller Jordan, Yuchen Jin, Vlado Boza, You Jiacheng, Franz Cecista, Laker Newhouse, and Jeremy Bernstein. Muon: An optimizer for hidden layers in neural networks, 2024. URL https://kellerjordan.github.io/ posts/muon/. 1, 5, 7, 13, 49, 51, 52, 54, 55, 56
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. Deep learning. nature, 2015. 3 Hongyang Li, Lincen Bai, Caesar Wu, Mohammed Chadli, Said Mammar, and Pascal Bouvry. Trustworthy efficient communication for distributed learning using lqsgd algorithm. arXiv preprint arXiv:2506.17974, 2025a. 34, 40, 44
Li Ju, Tianru Zhang, Salman Toor, and Andreas Hellander. Accelerating fair federated learning: Adaptive federated adam. IEEE Transactions on Machine Learning in Communications and Networking, 2024. 38
Jun Li, Fuxin Li, and Sinisa Todorovic. Efficient riemannian optimization on the stiefel manifold via the cayley transform. In ICLR, 2020. 20
Nikita P Kalinin, Ryan McKenna, Rasmus Pagh, and Christoph H Lampert. DP-λCGD: Efficient noise correlation for differentially private model training. arXiv preprint arXiv:2601.22334, 2026. 41
Ke Li and Jitendra Malik. Learning to optimize. In ICLR, 2017. 46 Pingzhi Li, Junyu Liu, Hanrui Wang, and Tianlong Chen. Q-newton: Hybrid quantum-classical scheduling for accelerating neural network training with newton’s gradient descent. arXiv preprint arXiv:2405.00252, 2024a. 23
Sai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank Reddi, Sebastian Stich, and Ananda Theertha Suresh. Scaffold: Stochastic controlled averaging for federated learning. In ICML, 2020. 36 Sergii Kavun. Novak: Unified adaptive optimizer for deep neural networks. arXiv preprint arXiv:2601.07876, 2026. 20
Tao Li, Pan Zhou, Zhengbao He, Xinwen Cheng, and Xiaolin Huang. Friendly sharpness-aware minimization. In CVPR, 2024b. 19
Dongyoon Kim, Sungjae Lee, Wonjin Lee, and Kwang In Kim. Subspace-based approximate hessian method for zeroth-order optimization. arXiv preprint arXiv:2507.06125, 2025. 28
Xi-Lin Li. Preconditioned stochastic gradient descent. IEEE TNNLS, 2017. 14, 48, 49, 51, 52, 54, 55, 56 Xiang Li, Wenhao Yang, Shusen Wang, and Zhihua Zhang. Communication-efficient local decentralized sgd methods. arXiv preprint arXiv:1910.09126, 2019. 37, 44
Diederik P Kingma. Adam: A method for stochastic optimization. In ICLR, 2015. 1, 4, 7, 9, 10, 11, 12, 13, 15, 17, 20, 21, 22, 28, 42, 45, 46, 47, 49, 50, 51, 52, 53, 54, 55
Xianliang Li, Jun Luo, Zhiwei Zheng, Hanxiao Wang, Li Luo, Lingkun Wen, Linlong Wu, and Sheng Xu. On the performance analysis of momentum method: A frequency domain perspective. In ICLR, 2025b. 10
Weiwei Kong, Mónica Ribero, et al. Differentially private optimization for non-decomposable objective functions. In ICLR, 2025. 43
Yongqi Li and Xiaowei Zhang. Adaptive moment estimation optimization algorithm using projection gradient for deep learning. In CAMMIC, 2025. 20
Jakub Kopal, Michal Gregor, Santiago de Leon-Martinez, and Jakub Simko. Overshoot: Taking advantage of
62
Zeman Li, Xinwei Zhang, Peilin Zhong, Yuan Deng, Meisam Razaviyayn, and Vahab Mirrokni. Addax: Utilizing zeroth-order gradients to improve memory efficiency and performance of SGD for fine-tuning language models. In ICLR, 2025c. 11, 30
Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In ICLR, 2019. 9, 13, 16, 19, 48, 49, 50, 51, 52, 53, 54, 55 Rongwei Lu, Jingyan Jiang, Chunyang Li, Haotian Dong, Xingguang Wei, Delin Cai, and Zhi Wang. Decosgd: Joint optimization of delay staleness and gradient compression ratio for distributed sgd. arXiv preprint arXiv:2507.17346, 2025. 35, 36, 39
Vladislav Lialin, Namrata Shivagunde, Sherin Muckatira, and Anna Rumshisky. Relora: High-rank training through low-rank updates. arXiv preprint arXiv:2307.05695, 2023. 56 Kaizhao Liang, Lizhang Chen, Bo Liu, and Qiang Liu. Cautious optimizers: Improving training with one line of code. arXiv preprint arXiv:2411.16085, 2024. 9
Yucheng Lu, Conglong Li, Minjia Zhang, Christopher De Sa, and Yuxiong He. Maximizing communication efficiency for large-scale training via 0/1 adam. In ICLR, 2023. 34
Hailiang Liu and Xuping Tian. An adaptive gradient method with energy and momentum. Ann. Appl. Math., 2022. 13
Liangchen Luo, Yuanhao Xiong, and Yan Liu. Adaptive gradient methods with dynamic bound of learning rate. In ICLR, 2019. 11, 17
Hong Liu, Zhiyuan Li, David Leo Wright Hall, Percy Liang, and Tengyu Ma. Sophia: A scalable stochastic second-order optimizer for language model pre-training. In ICLR, 2024a. 11, 24, 27, 43, 46, 56
Qijun Luo, Hengxu Yu, and Xiao Li. Badam: A memory efficient full parameter optimization method for large language models. In NeurIPS, 2024a. 21
Jie Liu and Yongqiang Wang. Communication efficient federated learning with linear convergence on heterogeneous data. arXiv preprint arXiv:2503.15804, 2025. 36, 47
Yihong Luo, Yuhan Chen, Siya Qiu, Yiwei Wang, Chen Zhang, Yan Zhou, Xiaochun Cao, and Jing Tang. Fast graph sharpness-aware minimization for enhancing and accelerating few-shot node classification. In NeurIPS, 2024b. 19
Junkang Liu, Fanhua Shang, Junchao Zhou, Hongying Liu, Yuanyuan Liu, and Jin Liu. Fedmuon: Accelerating federated learning with matrix orthogonalization. arXiv preprint arXiv:2510.27403, 2025. 38
Kai Lv, Hang Yan, Qipeng Guo, Haijun Lv, and Xipeng Qiu. Adalomo: Low-memory optimization with adaptive learning rate. arXiv preprint arXiv:2310.10195, 2023. 12
Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han. On the variance of the adaptive learning rate and beyond. In ICLR, 2020a. 12, 49, 51, 53, 54, 55
Kai Lv, Yuqing Yang, Tengxiao Liu, Qipeng Guo, and Xipeng Qiu. Full parameter fine-tuning for large language models with limited resources. In ACL, 2024. 11, 22, 46
Mingrui Liu, Wei Zhang, Francesco Orabona, and Tianbao Yang. Adam+ : A stochastic method with adaptive variance reduction. arXiv preprint arXiv:2011.11985, 2020b. 12
Chao Ma, Wenbo Gong, Meyer Scetbon, and Edward Meeds. SWAN: SGD with normalization and whitening enables stateless LLM training. In ICML, 2025. 21, 44, 45
Rui Liu, Tianyi Wu, and Barzan Mozafari. Adam with bandit sampling for deep learning. NeurIPS, 2020c. 21
Jerry Ma and Denis Yarats. Quasi-hyperbolic momentum and adam for deep learning. In ICLR, 2019. 12
Yanli Liu, Yuan Gao, and Wotao Yin. An improved analysis of stochastic gradient descent with momentum. In NeurIPS, 2020d. 17
Dipan Maity. Auon: A linear-time alternative to orthogonal momentum updates. arXiv preprint arXiv:2509.24320, 2025. 16
Yong Liu, Zirui Zhu, Chaoyu Gong, Minhao Cheng, ChoJui Hsieh, and Yang You. Sparse mezo: Less parameters for better performance in zeroth-order llm finetuning. arXiv preprint arXiv:2402.15751, 2024b. 29, 31, 46
Maksim Makarenko, Elnur Gasanov, Abdurakhmon Sadiev, Rustem Islamov, and Peter Richtárik. Adaptive compression for communication-efficient distributed training. TMLR, 2023. 35 Mehdi Makni, Kayhan Behdin, Gabriel Afriat, Zheng Xu, Sergei Vassilvitskii, Natalia Ponomareva, Rahul Mazumder, and Hussein Hazimeh. Sparta: An optimization framework for differentially private sparse fine-tuning. In KDD, 2025. 41
Hugo Touvron Llama. Llama: Open and efficient foundation language models, 2023. 1, 48, 49, 51, 52, 56 Michael Lomnitz, Zachary Daniels, Saurabh Farkya, Michael Isnardi, David Zhang, and Michael Piacentino. Learning with local gradients at the edge. In SMC, 2023. 44
Sadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian, Jason D Lee, Danqi Chen, and Sanjeev Arora.
63
Fine-tuning language models with just forward passes. In NeurIPS, 2023. 11, 27, 28, 29, 30
field. In International conference on smart applications and data analysis, 2020. 1
Artavazd Maranjyan. First provably optimal asynchronous sgd for homogeneous and heterogeneous data. arXiv preprint arXiv:2601.02523, 2026. 36
Anum Nawaz, Muhammad Irfan, Xianjia Yu, Zhuo Zou, and Tomi Westerlund. Blockchain-enabled privacypreserving second-order federated edge learning in personalized healthcare. arXiv preprint arXiv:2506.00416, 2025. 38
James Martens and Roger Grosse. Optimizing neural networks with kronecker-factored approximate curvature. In ICML, 2015. 4, 11, 23, 25, 56
Y Nesterov. A method of solving a convex programming problem with convergence rate O( k12 ). Sov. Math. Dokl., 1983. 5, 6, 8, 12, 13, 21, 36, 38, 49
John L Maryak and Daniel C Chin. Global random optimization by simultaneous perturbation stochastic approximation. In Proceedings of the 2001 American control conference.(Cat. No. 01CH37148), 2001. 1, 29, 31, 56
Yurii Nesterov et al. Lectures on convex optimization. Springer, 2018. 3
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models. arXiv preprint arXiv:1609.07843, 2016. 48, 56
Son Nguyen, Bo Liu, Lizhang Chen, and Qiang Liu. Improving adaptive moment optimization via preconditioner diagonalization. arXiv preprint arXiv:2502.07488, 2025. 20
Zhendong Mi, Qitao Tan, Xiaodong Yu, Zining Zhu, Geng Yuan, and Shaoyi Huang. Kerzoo: Kernel function informed zeroth-order optimization for accurate and accelerated llm fine-tuning. arXiv preprint arXiv:2505.18886, 2025. 31, 46
Julien Nicolas, Mohamed Maouche, Sonia Ben Mokhtar, and Mark Coates. Communication efficient, differentially private distributed optimization using correlationaware sketching. arXiv preprint arXiv:2507.03545, 2025. 34
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, et al. Mixed precision training. In ICLR, 2018. 4
Yue Niu, Zalan Fabian, Sunwoo Lee, Mahdi Soltanolkotabi, and Salman Avestimehr. ml-bfgs: A momentum-based l-bfgs for distributed large-scale neural network optimization. TMLR, 2023. 11, 23, 26 Jorge Nocedal. Updating quasi-newton matrices with limited storage. Mathematics of computation, 1980. 26, 46
Elmira Mirzabeigi, Sepehr Rezaee, and Kourosh Parand. Lyam: Robust non-convex optimization for stable learning in noisy environments. arXiv preprint arXiv:2507.11262, 2025. 17
Jose Javier Gonzalez Ortiz, Abhay Gupta, Chris Renard, and Davis Blalock. Flashoptim: Optimizers for memory efficient training. arXiv preprint arXiv:2602.23349, 2026. 44
Ionut-Vlad Modoranu, Mher Safaryan, Grigory Malinovsky, Eldar Kurtić, Thomas Robert, Peter Richtárik, and Dan Alistarh. Microadam: Accurate adaptive optimization with low space overhead and provable convergence. In NeurIPS, 2024. 45
Kaan Ozkara, Can Karakus, Parameswaran Raman, Mingyi Hong, Shoham Sabach, Branislav Kveton, and Volkan Cevher. MADA: Meta-adaptive optimizers through hyper-gradient descent. In ICML, 2024. 47
Goncalo Mordido, Pranshu Malviya, Aristide Baratin, and Sarath Chandar. Lookbehind-SAM: k steps back, 1 step forward. In ICML, 2024. 18
Matteo Pagliardini, Pierre Ablin, and David Grangier. The adEMAMix optimizer: Better, faster, older. In ICLR, 2025. 8
Jorge J Moré and Danny C Sorensen. Newton’s method. Technical report, Argonne National Lab.(ANL), Argonne, IL (United States), 1982. 1, 23
Shivam Pal, Aishwarya Gupta, Saqib Sarwar, and Piyush Rai. Federated learning with uncertainty and personalization via efficient second-order optimization. arXiv preprint arXiv:2411.18385, 2024. 38
Aggrey Muhebwa, Khotso Selialia, Fatima Anwar, and Khalid K Osman. Kuramoto-fedavg: Using synchronization dynamics to improve federated learning optimization under statistical heterogeneity. arXiv preprint arXiv:2505.19605, 2025. 38
Chengxi Pan, Junshang Chen, and Jingrui Ye. Warpadam: A new adam optimizer based on meta-learning approach. arXiv preprint arXiv:2409.04244, 2024. 47
Tomoya Murata and Taiji Suzuki. Bias-variance reduced local sgd for less heterogeneous federated learning. In ICML, 2021. 36
Zherong Pan and Kui Wu. Bc-admm: An efficient nonconvex constrained optimizer with robotic applications. arXiv preprint arXiv:2504.05465, 2025. 47
Aatila Mustapha, Lachgar Mohamed, and Kartit Ali. An overview of gradient descent algorithm optimization in machine learning: Application in the ophthalmology
Hanyang Peng, Shuang Qin, Yue Yu, Fangqing Jiang, Hui Wang, and Wen Gao. Softsignsgd (s3): An enhanced optimizer for practical dnn training and loss
64
spikes minimization beyond adam. arXiv:2507.06464, 2025. 13
arXiv preprint
convergence of convolutional neural networks. arXiv preprint arXiv:2105.10190, 2021. 9
Boris T Polyak. Some methods of speeding up the convergence of iteration methods. USSR Comput. Math. Math. Phys., 1964. 5, 6
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams. Learning representations by backpropagating errors. nature, 1986. 8, 10, 11, 13
Honglin Qin, Hongye Zheng, Bingxing Wang, Zhizhong Wu, Bingyao Liu, and Yuanfang Yang. Reducing bias in deep learning optimization: The rsgdm approach. In CCSB, 2024. 8
Sumanth Sadu, Shiv Ram Dubey, and SR Sreeja. Moment centralization-based gradient descent optimizers for convolutional neural networks. In CVMI. Springer, 2023. 16
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. Zero: Memory optimizations toward training trillion parameter models. In SC20: international conference for high performance computing, networking, storage and analysis, 2020. 56
Soham Sane. Alphagrad: Non-linear gradient normalization optimizer. arXiv preprint arXiv:2504.16020, 2025. 13, 43, 44, 46 Krisanu Sarkar. Hindsight-guided momentum (hgm) optimizer: An approach to adaptive learning rate. arXiv preprint arXiv:2506.22479, 2025. 17
Marco Rando, Cheik Traoré, Cesare Molinari, Lorenzo Rosasco, and Silvia Villa. A structured proximal stochastic variance reduced zeroth-order algorithm. arXiv preprint arXiv:2506.23758, 2025. 32
Pedro Savarese, David McAllester, Sudarshan Babu, and Michael Maire. Domain-independent dominance of adaptive methods. In CVPR, 2021. 13
Yehonathan Refael, Iftach Arbel, Ofir Lindenbaum, and Tom Tirer. Lorenza: Enhancing generalization in lowrank gradient llm training via efficient zeroth-order adaptive sam. arXiv preprint arXiv:2502.19571, 2025a. 28
Fabian Schaipp, Ruben Ohana, Michael Eickenberg, Aaron Defazio, and Robert M Gower. Momo: Momentum models for adaptive learning rates. In ICML, 2024. 12 Andrei Semenov, Matteo Pagliardini, and Martin Jaggi. Benchmarking optimizers for large language model pretraining. arXiv preprint arXiv:2509.01440, 2025. 48
Yehonathan Refael, Guy Smorodinsky, Tom Tirer, and Ofir Lindenbaum. Sumo: Subspace-aware momentorthogonalization for accelerating memory-efficient llm training. arXiv preprint arXiv:2505.24749, 2025b. 45
Mrinmay Sen and Chalavadi Krishna Mohan. pfedsop: Accelerating training of personalized federated learning using second-order optimization. arXiv preprint arXiv:2506.07159, 2025. 38
Yehonathan Refael, Jonathan Svirsky, Boris Shustin, Wasim Huleihel, and Ofir Lindenbaum. Adarankgrad: Adaptive gradient rank and moments for memoryefficient LLMs training and fine-tuning. In ICLR, 2025c. 44
Mrinmay Sen, A Kai Qin, et al. Fagh: Accelerating federated learning with approximated global hessian. arXiv preprint arXiv:2403.11041, 2024. 38
Xiaoxing Ren, Nicola Bastianello, Karl H Johansson, and Thomas Parisini. Communication-efficient stochastic distributed learning. arXiv preprint arXiv:2501.13516, 2025. 37
Hyunseok Seung, Jaewoo Lee, and Hyunsuk Ko. An adaptive method stabilizing activations for enhanced generalization. In ICDMW, 2024a. 13, 45
Apostolos I Rikos, Nicola Bastianello, Themistoklis Charalambous, and Karl H Johansson. Distributed optimization and learning for automated stepsize selection with finite time coordination. arXiv preprint arXiv:2508.05887, 2025. 37
Hyunseok Seung, Jaewoo Lee, and Hyunsuk Ko. Nysact: A scalable preconditioned gradient descent using nyström approximation. In BigData, 2024b. 14 Hyunseok Seung, Jaewoo Lee, and Hyunsuk Ko. Mac: An efficient gradient preconditioning using mean activation approximated curvature. arXiv preprint arXiv:2506.08464, 2025. 11, 25
Herbert Robbins and Sutton Monro. A stochastic approximation method. Ann. Math. Stat., 1951. 1, 3, 4, 5, 6, 13, 15, 17, 19, 20, 21, 37, 48, 49, 50, 51, 52, 54, 55 Thomas Robert, Mher Safaryan, Ionut-Vlad Modoranu, and Dan Alistarh. LDAdam: Adaptive optimization from low-dimensional gradient statistics. In ICLR, 2025. 20
Fanhua Shang, Kaiwen Zhou, Hongying Liu, James Cheng, Ivor W Tsang, Lijun Zhang, Dacheng Tao, and Licheng Jiao. Vr-sgd: A simple stochastic variance reduction method for machine learning. TKDE, 2018. 20, 43
Swalpa Kumar Roy, Mercedes Eugenia Paoletti, Juan Mario Haut, Shiv Ram Dubey, Purbayan Kar, Antonio Plaza, and Bidyut B Chaudhuri. Angulargrad: A new optimization technique for angular
Sifeng Shang, Jiayi Zhou, Chenyu Lin, Minxian Li, and Kaiyang Zhou. Fine-tuning quantized neural networks with zeroth-order optimization. arXiv preprint arXiv:2505.13430, 2025. 28, 29, 31
65
Zhou Shao, Hang Zhou, and Tong Lin. A new adaptive gradient method with gradient decomposition. Machine Learning, 2025. 17
Oscar Smee, Fred Roosta, and Stephen J Wright. Firstish order methods: Hessian-aware scalings of gradient descent. arXiv preprint arXiv:2502.03701, 2025. 14
Noam Shazeer and Mitchell Stern. Adafactor: Adaptive learning rates with sublinear memory cost. In ICML, 2018. 43, 44, 45, 48, 49, 51, 54, 55
Andrew Starnes, Guannan Zhang, Viktor Reshniak, and Clayton Webster. Anisotropic gaussian smoothing for gradient-based optimization. arXiv preprint arXiv:2411.11747, 2024. 20, 47
Wei Shen, Minhui Huang, Jiawei Zhang, and Cong Shen. Stochastic smoothed gradient descent ascent for federated minimax optimization. In AISTATS, 2024. 20
Felix Stollenwerk and Tobias Stollenwerk. Better embeddings with coupled adam. ACL, 2025. 13
Shaohuai Shi, Zhenheng Tang, Qiang Wang, Kaiyong Zhao, and Xiaowen Chu. Layer-wise adaptive gradient sparsification for distributed deep learning with convergence guarantees. ECAI, 2020. 35
Jianlin Su. Implementation of two optimizers in keras: Lookahead and lazyoptimizer, Jul 2019. URL https: //spaces.ac.cn/archives/6869. Blog post. 8 Keisuke Sugiura and Hiroki Matsutani. Elasticzo: A memory-efficient on-device learning with combined zeroth-and first-order optimization. arXiv preprint arXiv:2501.04287, 2025. 30
Yifan Shi, Yingqi Liu, Kang Wei, Li Shen, Xueqian Wang, and Dacheng Tao. Make landscape flatter in differentially private federated learning. In CVPR, 2023. 42, 43
Hao Sun, Li Shen, Qihuang Zhong, Liang Ding, Shixiang Chen, Jingwei Sun, Jing Li, Guangzhong Sun, and Dacheng Tao. Adasam: Boosting sharpness-aware minimization with adaptive learning rate and momentum for training deep neural networks. Neural Networks, 2024. 21
Zekun Shi, Zheyuan Hu, Min Lin, and Kenji Kawaguchi. Stochastic taylor derivative estimator: Efficient amortization for arbitrary differential operators. In NeurIPS, 2024. 24 Dahun Shin, Dongyeop Lee, Jinseok Chung, and Namhoon Lee. Sassha: Sharpness-aware adaptive second-order optimization with stable hessian approximation. In ICML, 2025a. 24
Lillian Sun, Kevin Cong, Je Qin Chooi, and Russell Li. Dp-adamw: Investigating decoupled weight decay and bias correction in private deep learning. In ICML, 2025a. 41
Dahun Shin, Dongyeop Lee, Jinseok Chung, and Namhoon Lee. Sassha: Sharpness-aware adaptive second-order optimization with stable hessian approximation. arXiv preprint arXiv:2502.18153, 2025b. 12
Yan Sun, Tiansheng Huang, Liang Ding, Li Shen, and Dacheng Tao. Tezo: Empowering the lowrankness on the temporal dimension in the zerothorder optimization for fine-tuning llms. arXiv preprint arXiv:2501.19057, 2025b. 31, 45
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. Megatron-lm: Training multi-billion parameter language models using model parallelism. arXiv, 2019. 4
Nikola Surjanovic, Alexandre Bouchard-Côté, and Trevor Campbell. Autosgd: Automatic learning rate selection for stochastic gradient descent. arXiv preprint arXiv:2505.21651, 2025. 17
Yao Shu, Qixin Zhang, Kun He, and Zhongxiang Dai. Refining adaptive zeroth-order optimization at ease. In ICML, 2025. 28, 32
Chengli Tan, Jiangshe Zhang, Junmin Liu, Yicheng Wang, and Yunda Hao. Stabilizing sharpness-aware minimization through a simple renormalization strategy. JMLR, 2025a. 18
Chongjie Si, Debing Zhang, and Wei Shen. Adamuon: Adaptive muon optimizer. arXiv preprint arXiv:2507.11005, 2025. 13
Qitao Tan, Jun Liu, Zheng Zhan, Caiwei Ding, Yanzhi Wang, Xiaolong Ma, Jaewoo Lee, Jin Lu, and Geng Yuan. Harmony in divergence: Towards fast, accurate, and memory-efficient zeroth-order llm fine-tuning. arXiv preprint arXiv:2502.03304, 2025b. 28, 31
Navjot Singh, Deepesh Data, Jemin George, and Suhas Diggavi. Squarm-sgd: Communication-efficient momentum sgd for decentralized optimization. JSAIT, 2021. 8, 34, 35, 36, 40 Navjot Singh, Deepesh Data, Jemin George, and Suhas Diggavi. Sparq-sgd: Event-triggered and compressed communication in decentralized optimization. TAC, 2022. 34, 39
Hanlin Tang, Shaoduo Gan, Samyam Rajbhandari, Xiangru Lian, Ji Liu, Yuxiong He, and Ce Zhang. Apmsqueeze: A communication efficient adampreconditioned momentum sgd algorithm. arXiv preprint arXiv:2008.11343, 2020. 35
Jordan Slessor, Dezheng Kong, Xiaofen Tang, Zheng En Than, and Linglong Kong. Fedstas: Client stratification and client level sampling for efficient federated learning. arXiv preprint arXiv:2412.14226, 2024. 38
Hanlin Tang, Shaoduo Gan, Ammar Ahmad Awan, Samyam Rajbhandari, Conglong Li, Xiangru Lian, Ji Liu, Ce Zhang, and Yuxiong He. 1-bit adam: Com-
66
munication efficient large-scale training with adam’s convergence speed. In ICML, 2021. 34
James Vo. Efficient second-order neural network optimization via adaptive trust region methods. arXiv preprint arXiv:2410.02293, 2024. 25
Qiaoyue Tang and Mathias Lécuyer. DP-adam: Correcting DP bias in adam’s second moment estimation. In ICLR, 2023. 41, 42
Thijs Vogels, Sai Praneeth Karimireddy, and Martin Jaggi. Powersgd: Practical low-rank gradient compression for distributed optimization. In NeurIPS, 2019. 34, 44
Qiaoyue Tang, Frederick Shpilevskiy, and Mathias Lécuyer. Dp-adambc: Your dp-adam is actually dp-sgd (unless you apply bias correction). In AAAI, 2024. 12
Bao Wang, Quanquan Gu, March Boedihardjo, Lingxiao Wang, Farzin Barekat, and Stanley J Osher. Dp-lssgd: A stochastic optimization method to lift the utility in privacy-preserving erm. In PMLR, 2020a. 42
Zhiwei Tang and Tsung-Hui Chang. Fedlion: Faster adaptive federated optimization with fewer communication. In ICASSP, 2024. 38
Bao Wang, Tan Nguyen, Tao Sun, Andrea L Bertozzi, Richard G Baraniuk, and Stanley J Osher. Scheduled restart momentum for accelerated stochastic gradient descent. SIAM Journal on Imaging Sciences, 2022. 8
Shohei Taniguchi, Keno Harada, Gouki Minegishi, Yuta Oshima, Seong Cheol Jeong, Go Nagahara, Tomoshi Iiyama, Masahiro Suzuki, Yusuke Iwasawa, and Yutaka Matsuo. Adopt: Modified adam can converge with any β2 with the optimal rate. In NeurIPS, 2024. 12, 48, 49, 51, 54, 55
Chengan Wang, Zichong Ou, and Jie Lu. A zeroth-order proximal algorithm for consensus optimization. In CDC, 2024a. 11, 33, 46
Yuanzhe Tao, Huizhuo Yuan, Xun Zhou, Yuan Cao, and Quanquan Gu. Towards simple and provable parameter-free adaptive gradient methods. arXiv preprint arXiv:2412.19444, 2024. 17
Fei Wang, Li Shen, Liang Ding, Chao Xue, Ye Liu, and Changxing Ding. Simultaneous computation and memory efficient zeroth-order optimizer for fine-tuning large language models. arXiv preprint arXiv:2410.09823, 2024b. 28, 29, 31, 46
Stephen Thomas. Streaming krylov-accelerated stochastic gradient descent. arXiv preprint arXiv:2505.07046, 2025. 19
Ganyu Wang, Jinjie Fang, Maxwell J Ying, Bin Gu, Xi Chen, Boyu Wang, and Charles Ling. Fedone: Query-efficient federated learning for black-box discrete prompt learning. arXiv preprint arXiv:2506.14929, 2025a. 38
Ran Tian and Ankur P Parikh. Amos: An adam-style optimizer with adaptive weight decay towards modeloriented scale. arXiv preprint arXiv:2210.11693, 2022. 17
Guoxia Wang, Shuai Li, Congliang Chen, Jinle Zeng, Jiabin Yang, Tao Sun, Yanjun Ma, Dianhai Yu, and Li Shen. Adagc: Improving training stability for large language model pretraining. arXiv preprint arXiv:2502.11034, 2025b. 15, 16, 47
Tijmen Tieleman and Geoffrey Hinton. Lecture 6.5rmsprop: Divide the gradient by a running average of its recent magnitude. COURSERA: Neural networks for machine learning, 2012. 5, 6, 12, 49, 50, 51, 53, 54, 55
Hao Wang, Jingrong Chen, Xinchen Wan, Han Tian, Jiacheng Xia, Gaoxiong Zeng, Weiyan Wang, Kai Chen, Wei Bai, and Junchen Jiang. Domain-specific communication optimization for distributed dnn training. arXiv preprint arXiv:2008.08445, 2020b. 39
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou. Training data-efficient image transformers & distillation through attention. In ICML, 2021. 48, 49
Hui-Po Wang, Dingfan Chen, Raouf Kerkouche, and Mario Fritz. Fedlap-dp: Federated learning by sharing differentially private loss approximations. arXiv preprint arXiv:2302.01068, 2023. 37
Hoang Tran and Ashok Cutkosky. Better sgd using secondorder momentum. In NeurIPS, 2022. 11, 24, 26 Edoardo Urettini and Antonio Carta. Online curvatureaware replay: Leveraging 2nd order information for online continual learning. CoRR, 2025. 25
Jianyu Wang, Vinayak Tantia, Nicolas Ballas, and Michael Rabbat. Slowmo: Improving communicationefficient distributed sgd with slow momentum. In ICLR, 2020c. 34, 35
Pranav Vaidhyanathan, Lucas Schorling, Natalia Ares, and Michael A Osborne. A physics-inspired optimizer: Velocity regularized adam. arXiv preprint arXiv:2505.13196, 2025. 10
Jiaxuan Wang and Jenna Wiens. Adasgd: Bridging the gap between sgd and adam. arXiv preprint arXiv:2006.16541, 2020. 11, 13, 19
Vladimir N Vapnik. An overview of statistical learning theory. IEEE transactions on neural networks, 1999. 3
Jing Wang, Yunfei Teng, and Anna Ewa Choromanska. Autodrop: Training deep learning models with automatic learning rate drop. In UAI, 2024c. 13, 17
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. NeurIPS, 2017. 2
67
Liangyu Wang, Jie Ren, Hang Xu, Junxiao Wang, Huanyi Xie, David E Keyes, and Di Wang. Zo2: Scalable zeroth-order fine-tuning for extremely large language models with limited gpu memory. arXiv preprint arXiv:2503.12668, 2025c. 31, 48
Adrian Wills, Thomas B Schön, and Carl Jidling. A fast quasi-newton-type method for large-scale stochastic optimisation. IFAC-PapersOnLine, 2020. 26 Chenyu Wu, Nuozhou Wang, Casey Garner, Kevin Leder, and Shuzhong Zhang. Novel optimization techniques for parameter estimation. arXiv preprint arXiv:2407.04235, 2024. 24
Mingze Wang, Jinbo Wang, Jiaqi Zhang, Wei Wang, Peng Pei, Xunliang Cai, Lei Wu, et al. Gradpower: Powering gradients for faster language model pre-training. arXiv preprint arXiv:2505.24275, 2025d. 12
Wenhan Xian, Feihu Huang, and Heng Huang. Communication-efficient adam-type algorithms for distributed data mining. In ICDM, 2022. 35
Ouya Wang, Shenglong Zhou, and Geoffrey Ye Li. Badm: Batch admm for deep learning. arXiv preprint arXiv:2407.01640, 2024d. 21
Tian Xie, Haoming Luo, Haoyu Tang, Yiwen Hu, Jason Klein Liu, Qingnan Ren, Yang Wang, Wayne Xin Zhao, Rui Yan, Bing Su, et al. Controlled llm training on spectral sphere. arXiv preprint arXiv:2601.08393, 2026. 21
Shaowen Wang, Anan Liu, Jian Xiao, Huan Liu, Yuekui Yang, Cong Xu, Qianqian Pu, Suncong Zheng, Wei Zhang, Di Wang, et al. Cadam: Confidence-based optimization for online learning. arXiv preprint arXiv:2411.19647, 2024e. 20
Wanyun Xie, Thomas Pethick, and Volkan Cevher. Sampa: Sharpness-aware minimization parallelized. In NeurIPS, 2024a. 18
Shipeng Wang, Jian Sun, and Zongben Xu. Hyperadam: A learnable task-adaptive adam for network training. In AAAI, 2019. 47
Xingyu Xie, Pan Zhou, Huan Li, Zhouchen Lin, and Shuicheng Yan. Adan: Adaptive nesterov momentum algorithm for faster optimizing deep models. TPAMI, 2024b. 9, 12, 49, 51, 54, 55
Sike Wang, Pan Zhou, Jia Li, and Hua Huang. 4-bit shampoo for memory-efficient network training. In NeurIPS, 2024f. 9, 11, 14, 44
Jie Xu, Wei Zhang, and Fei Wang. A (DP)2 SGD: Asynchronous decentralized parallel stochastic gradient descent with differential privacy. TPAMI, 2021. 37
Yanshu Wang, Wenyang He, and Tong Yang. Athena: Efficient block-wise post-training quantization for large language models using second-order matrix derivative information. arXiv preprint arXiv:2405.17470, 2024g. 23
Minghao Xu, Lichuan Xiang, Xu Cai, and Hongkai Wen. No more adam: Learning rate scaling at initialization is all you need. arXiv preprint arXiv:2412.11768, 2024. 16, 46
Yujia Wang, Shiqiang Wang, Songtao Lu, and Jinghui Chen. Fadas: Towards federated adaptive asynchronous optimization. In ICML, 2024h. 38
Xiaojian Xu and Ulugbek S Kamilov. Signprox: One-bit proximal algorithm for nonconvex stochastic optimization. In ICASSP, 2019. 34
Robert WM Wedderburn. Quasi-likelihood functions, generalized linear models, and the gauss—newton method. Biometrika, 1974. 23
Mohammad Yaghini, Tudor Cebere, Michael Menart, Aurélien Bellet, and Nicolas Papernot. Private rateconstrained optimization with applications to fair learning. arXiv preprint arXiv:2505.22703, 2025. 41
Chengkun Wei, Weixian Li, Gong Chen, and Wenzhi Chen. Dc-sgd: Differentially private sgd with dynamic clipping through gradient norm distribution estimation. TIFS, 2025. 41, 42, 43
Xiaodong Yang, Huishuai Zhang, Wei Chen, and Tie-Yan Liu. Normalized/clipped sgd with perturbation for differentially private non-convex optimization. arXiv preprint arXiv:2206.13033, 2022. 15, 41, 42
Jianxin Wei, Ergute Bao, Xiaokui Xiao, and Yin Yang. Dpis: An enhanced mechanism for differentially private sgd with importance sampling. In CCS, 2022. 41
Jiachen Yao, Chang Su, Zhongkai Hao, Songming Liu, Hang Su, and Jun Zhu. Multiadam: Parameter-wise scale-invariant optimizer for multiscale training of physics-informed neural networks. In ICML, 2023a. 16
Ziqing Wen, Ping Luo, Jiahuan Wang, Xiaoge Deng, Jinping Zou, Kun Yuan, Tao Sun, and Dongsheng Li. Wavelet meets adam: Compressing gradients for memory-efficient training. arXiv preprint arXiv:2501.07237, 2025. 45
Zhewei Yao, Amir Gholami, Sheng Shen, Mustafa Mustafa, Kurt Keutzer, and Michael Mahoney. Adahessian: An adaptive second order optimizer for machine learning. In AAAI, 2021. 11, 23, 24, 27
Ross Wightman. Pytorch image models. https://github. com/rwightman/pytorch-image-models, 2019. 49, 50, 51, 54, 55 Ross Wightman, Hugo Touvron, and Hervé Jégou. Resnet strikes back: An improved training procedure in timm. arXiv preprint arXiv:2110.00476, 2021. 48
Zhipeng Yao, Guiyuan Fu, Ying Li, Yu Zhang, Dazhou Li, and Rui Yu. Signal processing meets sgd: From
68
momentum to filter. arXiv preprint arXiv:2311.02818, 2023b. 10
Adaptive gradient transformation contributes to convergences and generalizations. CoRR, 2021. 13
Tian Ye, Peijun Xiao, and Ruoyu Sun. Deed: A general quantization scheme for communication efficiency in bits. arXiv preprint arXiv:2006.11401, 2020. 34
Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimization. In ICLR, 2018. 48
Yang You, Igor Gitman, and Boris Ginsburg. Large batch training of convolutional networks. arXiv preprint arXiv:1708.03888, 2017. 12, 48, 49, 50, 51, 54, 55
Huishuai Zhang, Bohan Wang, and Luoxin Chen. Adams: Momentum itself can be a normalizer for llm pretraining and post-training. arXiv preprint arXiv:2505.16363, 2025a. 13
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh. Large batch optimization for deep learning: Training bert in 76 minutes. In ICLR, 2020. 12, 16, 49, 50, 51, 54, 55
Jiangtao Zhang, Shunyu Liu, Jie Song, Tongtian Zhu, Zhengqi Xu, and Mingli Song. Lookaround optimizer: k steps around, 1 step average. In NeurIPS, 2023b. 19 Jianing Zhang, Evan Chen, Chaoyue Liu, and Christopher G Brinton. Dpzv: Elevating the tradeoff between privacy and utility in zeroth-order vertical federated learning. arXiv preprint arXiv:2502.20565, 2025b. 42
Ziming Yu, Pan Zhou, Sike Wang, Jia Li, Mi Tian, and Hua Huang. Zeroth-order fine-tuning of llms in random subspaces. arXiv preprint arXiv:2410.08989, 2024. 31 Honglin Yuan and Tengyu Ma. Federated accelerated stochastic gradient descent. In NeurIPS, 2020. 38
Jiawei Zhang and Fisher B Gouza. Gadam: geneticevolutionary adam for deep neural network optimization. arXiv preprint arXiv:1805.07500, 2018. 47
Huizhuo Yuan, Yifeng Liu, Shuang Wu, zhou Xun, and Quanquan Gu. MARS: Unleashing the power of variance reduction for training large models. In ICML, 2025. 1, 9, 12, 49, 51, 52, 54, 55
Michael Zhang, James Lucas, Jimmy Ba, and Geoffrey E Hinton. Lookahead optimizer: k steps forward, 1 step back. In NeurIPS, 2019. 21, 48, 49, 51, 54, 55
Wei Yuan and Kai-Xin Gao. Eadam optimizer: How ϵ impact adam. arXiv preprint arXiv:2011.02150, 2020. 12
Minxin Zhang, Yuxuan Liu, and Hayden Scheaffer. Adam improves muon: Adaptive moment estimation with orthogonalized momentum. arXiv preprint arXiv:2602.17080, 2026. 20
Yun Yue, Zhiling Ye, Jiadi Jiang, Yongchao Liu, and Ke Zhang. Agd: An auto-switchable optimizer using stepwise gradient difference for preconditioning matrix. In NeurIPS, 2023. 20
Xingxuan Zhang, Renzhe Xu, Han Yu, Hao Zou, and Peng Cui. Gradient norm aware minimization seeks first-order flatness and improves generalization. In CVPR, 2023c. 11, 18
Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. Cutmix: Regularization strategy to train strong classifiers with localizable features. In ICCV, 2019. 48
Xinwei Zhang, Zhiqi Bu, Mingyi Hong, and Meisam Razaviyayn. Doppler: Differentially private optimizers with low-pass filter for privacy noise reduction. In NeurIPS, 2024a. 41
Manzil Zaheer, Sashank Reddi, Devendra Sachan, Satyen Kale, and Sanjiv Kumar. Adaptive methods for nonconvex optimization. In NeurIPS, 2018. 9, 43
Xinwei Zhang, Zhiqi Bu, Steven Wu, and Mingyi Hong. Differentially private SGD without clipping bias: An error-feedback approach. In ICLR, 2024b. 15
Matthew D Zeiler. Adadelta: an adaptive learning rate method. arXiv preprint arXiv:1212.5701, 2012. 5, 6
Yiheng Zhang, Shaowu Wu, Yuanzhuo Xu, Jiajun Wu, Shang Xu, Steve Drew, and Xiaoguang Niu. Hvadam: A full-dimension adaptive optimizer. In AAAI, 2025c. 20
Di Zhang and Yihang Zhang. Beyond first-order: Training llms with stochastic conjugate subgradients and adamw. arXiv preprint arXiv:2507.01241, 2025. 19
Yushun Zhang, Congliang Chen, Ziniu Li, Tian Ding, Chenwei Wu, Diederik P Kingma, Yinyu Ye, Zhi-Quan Luo, and Ruoyu Sun. Adam-mini: Use fewer learning rates to gain more. In ICLR, 2025d. 44
Guoqiang Zhang. On suppressing range of adaptive stepsizes of adam to improve generalisation performance. In ECML, 2024. 12, 20 Guoqiang Zhang, Kenta Niwa, and W. Bastiaan Kleijn. A DNN optimizer that improves over adabelief by suppression of the adaptive stepsize range. TMLR, 2023a. 12
Zhen Zhang, Yifan Yang, Kai Zhen, Nathan Susanj, Athanasios Mouchtaris, Siegfried Kunzmann, and Zheng Zhang. Mazo: Masked zeroth-order optimization for multi-task fine-tuning of large language models. arXiv preprint arXiv:2502.11513, 2025e. 31
Hongwei Zhang, Weidong Zou, Hongbo Zhao, Qi Ming, Tijin Yan, Yuanqing Xia, and Weipeng Cao. Adal:
Huaqin Zhao, Jiaxi Li, Yi Pan, Shizhe Liang, Xiaofeng Yang, Wei Liu, Xiang Li, Fei Dou, Tianming Liu,
69
and Jin Lu. Helene: Hessian layer-wise clipping and gradient annealing for accelerating fine-tuning llm with zeroth-order optimization. arXiv preprint arXiv:2411.10696, 2024a. 45
Momentum centering and asynchronous update for adaptive gradient methods. In NeurIPS, 2021. 13 Liu Ziyin, Zhikang T Wang, and Masahito Ueda. Laprop: Separating momentum and adaptivity in adam. arXiv preprint arXiv:2002.04839, 2020. 13, 49, 51, 54, 55
Pengxiang Zhao, Ping Li, Yingjie Gu, Yi Zheng, Stephan Ludger Kölker, Zhefeng Wang, and Xiaoming Yuan. Adapprox: Adaptive approximation in adam optimization via randomized low-rank matrices. arXiv preprint arXiv:2403.14958, 2024b. 44, 45 Shen-Yi Zhao, Hao Gao, and Wu-Jun Li. Adass: Adaptive sample selection for training acceleration. arXiv preprint arXiv:1906.04819, 2019. 20 Shen-Yi Zhao, Chang-Wei Shi, Yin-Peng Xie, and WuJun Li. Stochastic normalized gradient descent with momentum for large-batch training. Sci. China Inf. Sci., 2024c. 10 Zhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, and Yi Yang. Random erasing data augmentation. In AAAI, 2020. 48 Beitong Zhou, Jun Liu, Weigao Sun, Ruijuan Chen, Claire J Tomlin, and Ye Yuan. pbsgd: Powered stochastic gradient descent methods for accelerated non-convex optimization. In IJCAI, 2020. 15 Jiajun Zhou, Yifan Yang, Kai Zhen, Ziyue Liu, Yequan Zhao, Ershad Banijamali, Athanasios Mouchtaris, Ngai Wong, and Zheng Zhang. Quzo: Quantized zerothorder fine-tuning for large language models. arXiv preprint arXiv:2502.12346, 2025. 31, 44, 47 Chen Zhu, Yu Cheng, Zhe Gan, Furong Huang, Jingjing Liu, and Tom Goldstein. Maxva: Fast adaptation of step sizes by maximizing observed variance of gradients. In ECML, 2021. 12 Dongxuan Zhu, Weihuan Huang, and Caihua Chen. Boosting accelerated proximal gradient method with adaptive sampling for stochastic composite optimization. arXiv preprint arXiv:2507.18277, 2025a. 8 Hanqing Zhu, Zhenyu Zhang, Wenyan Cong, Xi Liu, Sem Park, Vikas Chandra, Bo Long, David Z. Pan, Zhangyang Wang, and Jinwon Lee. APOLLO: SGDlike memory, adamw-level performance. In MLSys, 2025b. 13, 21 Meng Zhu, Quan Xiao, and Weidong Min. Adamnx: An adam improvement algorithm based on a novel exponential decay mechanism for the second-order moment estimate. arXiv preprint arXiv:2511.13465, 2025c. 12 Juntang Zhuang, Tommy Tang, Yifan Ding, Sekhar C Tatikonda, Nicha Dvornek, Xenophon Papademetris, and James Duncan. Adabelief optimizer: Adapting stepsizes by the belief in observed gradients. In NeurIPS, 2020. 11, 12, 15, 48, 49, 50, 51, 54, 55 Juntang Zhuang, Yifan Ding, Tommy Tang, Nicha Dvornek, Sekhar C Tatikonda, and James Duncan.
70