Conceptio › Archive › arXiv CS
arXiv CSopen access

AdamX: Cosine similarity meets gradient descent

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

AdamX: Cosine similarity meets gradient descent Francisco Caldas1[0000−0001−5090−0216] , Ruben Belo1[0009−0006−8516−7732] , and Cláudia Soares1[0000−0003−3071−6627]

arXiv:2609.11867v1 [cs.LG] 10 Sep 2026

NOVA School of Science of Technology Universidade Nova de Lisboa, Caparica, Portugal [email protected]

Abstract. We introduce AdamX, a first-order optimizer that incorporates cosine similarity as an adaptive mechanism for controlling update magnitudes. The proposed method is scalable, model-agnostic, and straightforward to integrate into existing training pipelines. We further introduce a variance rectification scheme that promotes smoother optimization during the early stages of training. Overall, we provide empirical evidence that AdamX achieves competitive convergence rates across a range of benchmark datasets and architectures. Performance is evaluated in terms of the number of epochs required to reach predefined performance thresholds under a fixed hyperparameter budget. Code and Experiments available at: https://github.com/FranciscoCaldas/adamX Keywords: First Order Optimizer · Online Convex Optimization · Cosine Similarity.

1

Introduction

Gradient-based optimization is central to modern machine learning, where model training requires minimizing high-dimensional and generally non-convex objectives. While stochastic gradient descent (SGD) remains a foundational approach [4], adaptive first-order methods have become widely used because they adjust update magnitudes according to observed gradient statistics. Prominent examples include AdaGrad [8], RMSProp [1], and Adam [12]. Adam combines exponential moving averages of the gradient and its coordinatewise squared magnitude to produce momentum-based, adaptively normalized updates [12]. This combination has made Adam a practical default for many deep learning tasks, as it often provides stable training behavior with limited task-specific tuning. However, its adaptive normalization can also lead to problematic update dynamics and, in some settings, a failure to converge [15]. Existing Adam-type methods primarily improve the treatment of magnitude information in the optimization trajectory. AMSGrad enforces a monotone second-moment envelope to recover convergence guarantees in online convex optimization [15]; AdamW decouples weight decay from adaptive updates [14]; RAdam addresses instability in the early variance estimate [13]; and AdaBelief modifies the second-moment statistic to reflect deviation from the predicted

2

F. Caldas et al.

gradient direction [23]. More recent optimizers, such as Lion [5] and Muon [11], further reconsider the form of the update rule or its preconditioning structure. Despite these developments, the alignment between successive gradients remains comparatively underused as a direct mechanism for modulating update magnitudes. Directional alignment provides a computationally inexpensive signal for adapting update magnitudes. Consecutive gradients that point in similar directions indicate locally consistent optimization progress, whereas poorly aligned or opposing gradients may indicate oscillation or rapidly changing trajectory information. Cosine similarity captures this signal independently of gradient scale. Closely related to this motivation, GALA [10] adapts the learning rate using consecutive-gradient alignment together with a local curvature estimate, formulated through a one-dimensional online learning problem. AdamX instead incorporates alignment through a bounded multiplicative cosine controller within an Adam/AMSGrad-style coordinate-wise adaptive update. This design preserves the practical structure of adaptive moment methods while explicitly exploiting directional consistency. Second-order and preconditioned optimization methods also exploit geometric information to improve training dynamics. Methods such as Shampoo and SOAP construct richer approximations to curvature or preconditioning structure, and can improve optimization performance in large-scale learning problems [9,20]. In contrast, our objective is to investigate whether a lightweight scalar signal derived from consecutive gradient directions can provide useful geometric adaptivity while retaining the implementation simplicity and scalability of first-order Adam-type methods. Our contributions are threefold. First, we introduce AdamX, an adaptive first-order optimizer that integrates a bounded cosine-similarity controller into an Adam-style moment-normalized update with a monotone second-moment envelope. Second, we provide an OCO analysis of a simplified momentum-free AdamX variant, showing that a bounded cosine controller can be incorporated into an adaptive projected-gradient scheme without worsening the standard convex regret rate. Third, we evaluate AdamX across benchmark datasets and architectures, with default settings, measuring the number of epochs required to reach predefined test-performance thresholds.

2

Method

We consider stochastic optimization of an expected loss over parameters θ ∈ Rd . Let D denote the data distribution and let B ∼ D be a randomly sampled mini-batch. The training objective is min EB∼D [L(θ; B)] .

θ∈Rd

(1)

At iteration t, the optimizer observes the stochastic gradient gt = ∇θ L(θt ; Bt ).

(2)

AdamX: Cosine similarity meets gradient descent

3

AdamX builds on the moment-normalized update used by Adam. Specifically, it maintains exponential moving averages of the stochastic gradient and its coordinate-wise square mt = β1 mt−1 + (1 − β1 )gt ,

vt = β2 vt−1 + (1 − β2 )gt2 ,

(3)

where β1 , β2 ∈ [0, 1) and all squares are taken element-wise. The corresponding bias-corrected estimates are vt mt , v̂t = . (4) m̂t = 1 − β1t 1 − β2t AdamX addresses two complementary aspects of adaptive optimization. First, Adam may fail to converge even in simple convex settings because its adaptive denominator can produce unfavorable effective stepsizes. AMSGrad addresses this issue by retaining a coordinate-wise maximum of past second-moment estimates [15]. Second, adaptive learning rates may exhibit high variance during the early stages of training, motivating rectification mechanisms such as RAdam [13]. AdamX retains an AMSGrad-style monotone denominator and augments it with a bounded controller derived from directional agreement between consecutive gradients. The distinctive component of AdamX is a cosine-similarity controller that modulates the magnitude of updates. For t ≥ 2, we define ct =

⟨gt , gt−1 ⟩ , max {∥gt ∥2 ∥gt−1 ∥2 , δ}

γt = exp(λct ),

(5)

where λ ≥ 0 controls the strength of the adaptation and δ > 0 prevents division by zero. We set γ1 = 1, since no preceding gradient is available at the first iteration. Because ct ∈ [−1, 1], the controller is bounded as e−λ ≤ γt ≤ eλ .

(6)

Thus, aligned consecutive gradients increase the effective update magnitude, while opposing gradients reduce it, without making the controller dependent on (k) gradient scale. Also note that, for each parameter group k, γt is computed from the cosine similarity between the current and previous gradients of the group. Gradient alignment has previously been used to adapt learning rates in hypergradient-based methods [2,17,3]. AdamX uses this signal in a different optimizer structure: cosine similarity acts as a bounded multiplicative controller on top of an Adam-style moment-normalized update with a monotone secondmoment envelope. Consequently, setting λ = 0 removes the alignment controller and recovers the AMSGrad algorithm. The normalization component of AdamX uses an AMSGrad-style variance envelope. We define ṽt = max {ṽt−1 , v̂t } , (7) where the maximum is evaluated coordinate-wise. This construction ensures that the adaptive denominator is coordinate-wise non-decreasing, preventing increases in effective coordinate-wise stepsizes that arise solely from decreases in the second-moment estimate.

4

F. Caldas et al.

The resulting AdamX update combines moment normalization, the monotone variance envelope, and the cosine controller. Given a base learning rate η > 0, t the parameters are updated as θt = θt−1 −ηγt √ṽm̂+ϵ1 , where all vector operations t in the denominator are coordinate-wise.

Algorithm 1 AdamX Optimizer Require: Learning rate η, decay rates β1 , β2 ∈ [0, 1), epsilon ϵ Require: Scaling parameter λ, initial parameters θ0 1: m0 ← 0,v0 ← 0, γ1 ← 1 2: gprev ← 1 {Initialize previous gradient} 3: for t = 1 to T do 4: gt ← ∇θ ft (θt−1 ) {Get gradients w.r.t. stochastic objective at t} 5: mt ← β1 mt−1 + (1 − β1 )gt 6: vt ← β2 vt−1 + (1 − β2 )gt2 7: m̂t ← mt /(1 − β1t ) 8: v̂t ← vt /(1 − β2t ) 9: if t > 1 then 10: γt ← exp (λ · cosinesimilarity(gt , gprev )) 11: end if 12: ṽt ← max(ṽt , v̂t ) {variance envelope} t 13: θt ← θt−1 − η · √γṽttm̂ {Update parameters} +ϵ1 14: gprev ← gt {Store gradient for next iteration} 15: end for 16: return θt

Our regret analysis considers a modified AdamX update designed for online convex optimization. In particular, the analyzed variant removes momentum, uses the current subgradient as the update direction, projects onto a convex feasible set in an adaptive diagonal metric, and controls the alignment-scaled learning rate through a non-increasing envelope. These modifications isolate the effect of the bounded cosine controller while enabling a standard adaptive onlinelearning analysis.

3

Regret Guarantees for AdamX-OCO

Scope of the analysis. The practical AdamX optimizer in Algorithm 1 uses momentum and applies the raw alignment multiplier γt to the update magnitude. To obtain a transparent regret guarantee, we analyze an OCO variant that removes momentum, projects onto a convex feasible set using an adaptive diagonal metric, and replaces the raw alignment-scaled step size with a non-increasing envelope. This variant isolates the effect of the bounded cosine controller while retaining the AMSGrad-style monotone second-moment envelope.

AdamX: Cosine similarity meets gradient descent

5

Algorithm 2 AdamX for OCO Require: Convex compact set K ⊂ Rd , sequence αt > 0, parameters β2 ∈ [0, 1), ϵ > 0, λ ≥ 0, ρ ∈ [0, 1] Require: Initial point θ1 ∈ K 1: v0 ← 0, v̄0 ← 0, q0 ← +∞, g0 ← ⊥, γ1 ← 1 2: for t = 1 to T do 3: Play θt and observe gt ∈ ∂ft (θt ) 4: vt ← β2 vt−1 + (1 − β2 )gt2 5: v̂t ← vt /(1 − β2t ) 6: ṽt ← max{ṽt−1 , v̂t } 7: if t > 1 then 8: ct ← ⟨gt , gt−1 ⟩ /(∥gt ∥2 ∥gt−1 ∥2 ) 9: γt ← exp(λct ) {AdamX} 10: end if 11: qt ← min{qt−1 √ , αt γt } {OCO Adaptation} 12: Ht ← diag( ṽt + ϵ1) 2 13: θt+1 ← argminθ∈K θ − (θt − qt Ht−1 gt ) H t 14: end for 15: return θT +1

2

Adaptive projected-gradient interpretation. Define the weighted norm ∥x∥Ht = x⊤ Ht x and the effective metric At =

Ht . qt

(8)

Since multiplication of the projection metric by a positive scalar does not change the projection, Algorithm 2 can equivalently be written as   1 2 (9) θt+1 = argminθ∈K ⟨gt , θ⟩ + ∥θ − θt ∥At . 2 The monotone envelope qt is introduced solely for analysis: together with the monotone second-moment envelope, it ensures that At is coordinate-wise nondecreasing. Convex Regret Guarantee Online convex optimization setting. At round t, the learner chooses θt ∈ K, observes a convex loss ft : K → R, and receives a subgradient gt ∈ ∂ft (θt ). For any comparator u ∈ K, the regret is RT (u) =

T X

(ft (θt ) − ft (u)) .

(10)

t=1

By convexity, RT (u) ≤

T X t=1

⟨gt , θt − u⟩ .

(11)

6

F. Caldas et al.

Assumption 1 (Bounded domain and gradients) There exist constants D∞ > 0 and G∞ > 0 such that, for every θ, u ∈ K and every t, ∥θ − u∥∞ ≤ D∞ ,

∥gt ∥∞ ≤ G∞ .

(12)

Bounded alignment controller. The stabilized cosine similarity satisfies ct ∈ [−1, 1], and therefore the AdamX alignment multiplier obeys e−λ ≤ γt = exp(λct ) ≤ eλ

(13)

for every t (with γ1 = 1 by definition). Lemma 1 (One-step adaptive projected-gradient bound). For every u ∈ K,  1 1 2 2 2 ∥θt − u∥At − ∥θt+1 − u∥At + ∥gt ∥A−1 . (14) ⟨gt , θt − u⟩ ≤ t 2 2 Equivalently, ⟨gt , θt − u⟩ ≤

 q 1  t 2 2 2 ∥θt − u∥Ht − ∥θt+1 − u∥Ht + ∥gt ∥H −1 . t 2qt 2

(15)

Proof. The optimality condition of the projected update gives ⟨gt + At (θt+1 − θt ), u − θt+1 ⟩ ≥ 0. Combining this inequality with 2

2

2

2 ⟨θt − θt+1 , At (θt − u)⟩ = ∥θt − θt+1 ∥At + ∥θt − u∥At − ∥θt+1 − u∥At .

(16)

Applying Young’s inequality to ⟨gt , θt − θt+1 ⟩ yields the result. Convex Regret Bound for AdamX-OCO Theorem 1 (Convex√OCO regret). Suppose Assumption 1 holds. Run Algorithm 2 with αt = η/ t for some η > 0. Then for every u ∈ K, d T d 2 2 X 1 X X gt,i D∞ HT,i + qt . RT (u) ≤ 2qT i=1 2 t=1 i=1 Ht,i

(17)

Moreover, using the coarse bounds ϵ ≤ Ht,i ≤ G∞ + ϵ, RT (u) ≤ Consequently,

2 dD∞ (G∞ + ϵ) √ ηeλ dG2∞ √ T + T. 2ηe−λ ϵ

√ RT (u) = O( T ).

(18)

AdamX: Cosine similarity meets gradient descent

7

Proof. By convexity, RT (u) ≤

T X

⟨gt , θt − u⟩ .

t=1

Applying Lemma 1 and summing over t gives T T  1X 1 X 2 2 2 ∥θt − u∥At − ∥θt+1 − u∥At + ∥gt ∥A−1 . RT (u) ≤ t 2 t=1 2 t=1

Because ṽt is coordinatewise nondecreasing and qt is nonincreasing, the matrix sequence At = Ht /qt is positive semidefinite nondecreasing. Therefore the first sum telescopes with an additional nonnegative metric-growth term and can be bounded as T T  1 1X 1 X 2 2 2 2 ∥θt − u∥At −At−1 . ∥θt − u∥At − ∥θt+1 − u∥At ≤ ∥θ1 − u∥A1 + 2 t=1 2 2 t=2

Since At is diagonal and ∥θt − u∥∞ ≤ D∞ , this is at most d d 2 X D∞ D2 X AT,i = ∞ HT,i . 2 i=1 2qT i=1

The second term is T T d 2 1 X X gt,i 1X 2 ∥gt ∥A−1 = qt . t 2 t=1 2 t=1 i=1 Ht,i

This proves the first bound. √ Since γt ∈ [e−λ , eλ ] and αt = η/ t, the envelope satisfies ηe−λ qT ≥ √ , T

ηeλ qt ≤ √ . t

Also Ht,i ≥ ϵ and, under ∥gt ∥∞ ≤ G∞ , the second-moment and scalar rectification terms are bounded so that Ht,i ≤ G∞ + ϵ. Hence d 2 X 2 D∞ dD∞ (G∞ + ϵ) √ HT,i ≤ T. 2qT i=1 2ηe−λ

For the second term, T d T 2 1 X X gt,i 1 X ηeλ dG2∞ ηeλ dG2∞ √ √ qt ≤ ≤ T. 2 t=1 i=1 Ht,i 2 t=1 t ϵ ϵ

Combining the two inequalities gives the stated result. Remark 1 (Effect of the alignment parameter). The regret rate is unchanged by the alignment factor, but the constants scale with eλ . This is expected: the multiplier γt is bounded between e−λ and eλ .

8

F. Caldas et al.

Using Raw Alignment Instead of the Envelope The monotone envelope qt = min{qt−1 , αt γt } is theoretically convenient, but it removes part of the intended behavior of AdamX: when gradients become strongly aligned, the method cannot re-increase the effective step size if the envelope has already decreased. If one instead uses the raw effective step size rt = αt γt , and defines At =

Ht , rt

then At need not be monotone, even when Ht is monotone. The proof still yields a data-dependent variation bound. Proposition 1 (Variation-dependent regret with raw alignment). Consider the AdamX-OCO update with rt = αt γt instead of the monotone envelope qt . Then for every u ∈ K, RT (u) ≤

T T 1 1X 1X 2 2 2 ∥θ1 − u∥A1 + ∥θt − u∥(At −At−1 )+ + ∥gt ∥A−1 , t 2 2 t=2 2 t=1

(19)

where (At − At−1 )+ denotes the positive part of the symmetric matrix At − At−1 . This statement is more faithful to the practical optimizer. It says that AdamXOCO keeps sublinear regret when the metric variation induced by the denominator and the alignment factor is controlled. In adversarial sequences, however, the cosine signal can oscillate, and the variation term may be large. Momentum and the Full AdamX Algorithm The full AdamX update uses m̂t rather than gt . In OCO, convexity gives ft (θt ) − ft (u) ≤ ⟨gt , θt − u⟩ , whereas the projected update controls a term involving m̂t . Consequently, ⟨gt , θt − u⟩ = ⟨m̂t , θt − u⟩ + ⟨gt − m̂t , θt − u⟩ . The first term can be handled by adaptive mirror descent. The second is a momentum-bias term. A direct bound gives T X t=1

⟨gt − m̂t , θt − u⟩ ≤ D∞

T X

∥gt − m̂t ∥1 ,

t=1

which can be linear in T for adversarial gradient sequences [7].

(20)

AdamX: Cosine similarity meets gradient descent

4

9

Experiments

We evaluate the proposed algorithm against comparable optimizers using an experimental protocol inspired by DeepOBS [19] and AlgoPerf [6]. When designing empirical evaluations for deep learning optimizers, we focus on three key aspects. (1) Generalization. The goal of optimization in deep learning is to learn models that generalize well to unseen data. Although some prior studies focus primarily on training metrics, improvements in training loss do not necessarily translate into better test performance. Accordingly, our evaluation emphasizes test-set performance throughout. (2) Stochasticity. The observed performance can vary substantially due to random initialization. To mitigate the influence of these sources of randomness and ensure fair comparisons, all optimizers are evaluated using the same set of five random seeds, and results are reported as averages across runs. (3) Realistic evaluation setting. Optimizer performance is highly dependent on the model architecture and dataset. Consequently, we adopt established benchmark architectures from DeepOBS [19] together with widely used datasets, ensuring evaluation on representative and commonly studied tasks. For consistency, we evaluate all optimizers, including ours, with the default hyperparameters [18]. Following this principles, the main evaluation tool is the number of epochs necessary to achieve a predetermined test set accuracy. Unlike AlgoPerf[6], which is more focused on algorithmic speed, we do not evaluate on wall-clock runtime, which has well-known drawbacks, such as dependency on hardware or weak reproducibility. By evaluating on epochs, we evaluate performance against the number of gradient evaluations, which typically dominates the total computational costs. Baseline Algorithms To evaluate AdamX, we compare against widely used first-order optimizers spanning adaptive, momentum-based, and non-adaptive methods. Specifically, we consider SGD [16]; Adagrad, which accumulates squared historical gradients [8]; RMSProp, which replaces Adagrad’s cumulative statistic with an exponential moving average [1]; Adam [12]; AdamW, which decouples weight decay from adaptive updates [14]; AMSGrad, which enforces a non-decreasing second-moment estimate [15]; RAdam, which introduces variance rectification during early training [13]; Yogi, which controls excessive growth of the variance estimate [22]; Lion, which updates parameters using the sign of the momentum vector [5]; and Adan, which incorporates Nesterov-style momentum into adaptive moment estimation [21]. MNIST On MNIST, we use a three-layer CNN with default hyperparameters and measure the epochs required to reach 0.994 test accuracy, up to 100 epochs. Figure 1 shows that AdamX is competitive with the best-performing optimizers, AMSGrad and Yogi. SGD does not reach the target, RMSprop fails in all runs due to gradient collapse, and Adagrad shows the largest variance across

10

F. Caldas et al. 

































7UDLQORVV

(SRFK













[

P DGD





\RJ

L



DGDP[ \RJL VJG UDGDP OLRQ DPVJUDG DGDQ DGDPZ DGDP DGDJUDG



VJG

UP

RS VSU

U

P DGD

OLRQ

DP

DG VJU

2SWLPL]HU

Q DGD

D

Z GDP

P DGD

DG DJU

DG

Fig. 1. Number of epochs to reach the test accuracy target. Each optimizer is evaluated over five seeds. Yogi, AMSGrad, and AdamX achieve the best performance, while SGD and RMSProp fail to reach the target.







(SRFK







Fig. 2. Training loss over 100 epochs on MNIST. Methods with better generalization (AMSGrad, Yogi, AdamX) also exhibit lower training loss. RMSProp is omitted due to significantly higher loss values.

seeds. The training-loss curves in Figure 2 are consistent with these results, with AMSGrad, AdamX, and Yogi among the fastest methods to reach the target. CIFAR-10 CIFAR-10 is more challenging than MNIST; to focus on optimizer behavior, we use the fixed CifarNet architecture [11]. The target test accuracy is 0.84, with a maximum of 100 epochs. Figure 3 shows that five out of eleven optimizers fail to reach the target, indicating that the threshold captures a demanding training regime. AdamX, AMSGrad, and Adam are the strongest methods, with AdamX achieving the lowest mean number of epochs and outperforming the closely related AMSGrad baseline. Figure 4 further shows that AdamX, AMSGrad, and Adan exhibit smoother training-loss trajectories, whereas RAdam, Adam, and AdamW display larger oscillations. Table 1 summarizes the results. Overall, AdamX and AMSGrad require the fewest gradient evaluations to reach the target, with AdamX comparing favorably on CIFAR-10. The results also illustrate that lower training loss does not necessarily imply better generalization; for example, Adagrad obtains a low MNIST training loss but requires more epochs to reach the test-accuracy target.

5

Conclusions

We presented AdamX, an adaptive first-order optimizer that augments Adam/AMSGradstyle updates with a bounded cosine similarity controller. A simplified OCO analysis shows that, under a monotone envelope on the cosine-scaled step size, the alignment mechanism is compatible with standard adaptive regret guarantees. Empirically, AdamX is competitive with ten established optimizers and

AdamX: Cosine similarity meets gradient descent 











adamx yogi sgd rmsprop radam lion amsgrad adan adamw adam adagrad

100

10 1

Train loss

(SRFK





 

11

10 2









P[

DGD







\RJ

L

VJG

UPV

SUR

S

UDG

DP

OLRQ

VJU

DP

DG

Q

DGD

2SWLPL]HU

DGD

PZ

10 3

P DG DGD DGDJU

Fig. 3. Number of epochs to reach the test accuracy target. Each optimizer is evaluated over five seeds. AdamX obtains the lowest mean number of epochs to reach the target, with similar values obtained by AMSGrad and Adam.

0

20

40

Epoch

60

80

100

Fig. 4. Training loss over 100 epochs on CIFAR-10. Methods with better generalization (AMSGrad, Yogi, AdamX) also exhibit lower training loss. RMSProp is omitted due to significantly higher loss values.

Table 1. Number of epochs until target accuracy, and training loss at 100 Epochs. Lower is better. Maximum number of runs is 100. Best, second-best, and third-best results are highlighted. MNIST Optimizer

Epochs (↓) Train Loss (↓) (×10

AdamX (Ours) 13.2 ± 2.00 Yogi SGD RMSProp Radam Lion AMSGrad Adan AdamW Adam Adagrad

CIFAR-10 −5

9.2 ± 0.734 100 100 25.2 ±6.67 42.0 ±10.36 9.2 ±1.35 44.4 ± 10.19 24.2 ± 2.73 28.0 ± 4.09 64.4 ±21.39

) Epochs (↓) Train Loss(↓)

2.17 ± 0.41

6.6 ± 0.54

1.96e-03

0.79 ± 0.13 25.00 ± 0.52 1609.97 ± 295.46 90.39 ± 13.29 397.40 ± 4.07 0.44 ± 0.02 89.99 ± 12.00 112.03 ± 67.39 100.33 ±35.05 7.32 ± 1.645

100 100 24.2 ± 28.94 30.2 ±24.93 100.0 7.8 ± 1.30 100 16.6 ± 11.65 9.6 ± 5.86 100

9.59e-03 2.30 2.64e-03 2.72e-03 6.42e-03 2.34e-03 2.16e-03 1.60e-03 1.72e-03 5.08e-02

achieves the best result in the considered CIFAR-10 setting. These results indicate that directional alignment is a promising lightweight source of adaptivity, motivating future work on hyperparameter robustness, second-order extensions, and larger-scale training regimes.

12

F. Caldas et al.

Acknowledgments. This work was partially supported by NOVA LINCS (UID/04516) funded by FCT IP, and the Neuraspace AI Fights Space Debris project (C62644988900463050), co-funded by Recovery and Resilience Plan and NextGeneration EU Funds, www.recuperarportugal.gov.pt. The authors have no competing interests to declare that are relevant to the content of this article. NextGenerationEU

References 1. Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude. COURSERA: Neural networks machine learning 4(2), 26 (2012) 2. Almeida, L.B., Langlois, T., Amaral, J.D., Plakhov, A.: Parameter adaptation in stochastic optimization. In: On-Line Learning Neural Networks. CUP (1998) 3. Baydin, A., Cornish, R., Rubio, D., Schmidt, M., Wood, F.: Online learning rate adaptation with hypergradient descent. In: ICLR (2018) 4. Cauchy, A.L.: Méthode générale pour la résolution des systèmes d’équations simultanées. Comptes Rendus Hebd Seances Acad Sci 25, 536–538 (1847) 5. Chen, X., Liang, C., Huang, D., Real, E., Wang, K., Pham, H., Dong, X., Luong, T., Hsieh, C.J., Lu, Y., Le, Q.V.: Symbolic discovery of optimization algorithms. In: NeurIPS (2023) 6. Dahl, G.E., Schneider, F., Nado, Z., Agarwal, N., Sastry, C.S., Hennig, P., Medapati, S., Eschenhagen, R., Kasimbeg, P., Suo, D., Bae, J., Gilmer, J., Peirson, A.L., Khan, B., Anil, R., Rabbat, M., Krishnan, S., Snider, D., Amid, E., Chen, K., Maddison, C.J., Vasudev, R., Badura, M., Garg, A., Mattson, P.: Benchmarking Neural Network Training Algorithms. arXiv preprint arXiv:2306.07179 (2023) 7. Défossez, A., Bottou, L., Bach, F., Usunier, N.: A simple convergence proof of Adam and Adagrad. TMLR (2022) 8. Duchi, J., Hazan, E., Singer, Y.: Adaptive subgradient methods for online learning and stochastic optimization. JMLR 12, 2121–2159 (2011) 9. Gupta, V., Koren, T., Singer, Y.: Shampoo: Preconditioned stochastic tensor optimization. In: ICML. pp. 1842–1850. PMLR (2018) 10. Jiang, R., Kavis, A., Mokhtari, A.: Online learning-guided learning rate adaptation via gradient alignment. arXiv preprint arXiv:2506.08419 (2025) 11. Jordan, K., Jin, Y., Boza, V., Jiacheng, Y., Cesista, F., Newhouse, L., Bernstein, J.: Muon: An optimizer for hidden layers in neural networks (2024) 12. Kingma, D.P., Ba, J.: Adam: A method for stochastic optimization. In: ICLR (2015) 13. Liu, L., Jiang, H., He, P., Chen, W., Liu, X., Gao, J., Han, J.: On the variance of the adaptive learning rate and beyond. In: ICLR (2020) 14. Loshchilov, I., Hutter, F.: Decoupled weight decay regularization. In: ICLR (2019) 15. Reddi, S.J., Kale, S., Kumar, S.: On the convergence of adam and beyond. In: ICLR (2018) 16. Robbins, H., Monro, S.: A stochastic approximation method. Annals Mathematical Statistics 22(3), 400–407 (1951) 17. Rubio, D.M.: Convergence Analysis of an Adaptive Method of Gradient Descent. Msc thesis, U. Oxf. (2017) 18. Schmidt, R.M., Schneider, F., Hennig, P.: Descending through a crowded valley benchmarking deep learning optimizers. In: ICML (2021)

AdamX: Cosine similarity meets gradient descent

13

19. Schneider, F., Balles, L., Hennig, P.: DeepOBS: A deep learning optimizer benchmark suite. In: ICLR (2019) 20. Vyas, N., Morwani, D., Zhao, R., Shapira, I., Brandfonbrener, D., Janson, L., Kakade, S.: SOAP: Improving and stabilizing shampoo using adam for language modeling. In: ICLR (2025) 21. Xie, X., Zhou, P., Li, H., Lin, Z., Yan, S.: Adan: Adaptive nesterov momentum algorithm for faster optimizing deep models. IEEE TPAMI (2024) 22. Zaheer, M., Reddi, S., Sachan, D., Kale, S., Kumar, S.: Adaptive methods for nonconvex optimization. In: NeurIPS. vol. 31 (2018) 23. Zhuang, J., Tang, T., Ding, Y., Tatikonda, S.C., Dvornek, N., Papademetris, X., Duncan, J.: Adabelief optimizer: Adapting stepsizes by the belief in observed gradients. NeurIPS (2020)

Record · ID 673500 · SHA-256 05fbbf6b1f5588b6
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.