Conceptio › Archive › arXiv CS
arXiv CSopen access

Musec: MomentUm SpEctral Clipping for Stable Muon-type Training

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Preprint. Under review.

M USEC : M OMENT U M S P E CTRAL C LIPPING FOR S TABLE M UON - TYPE T RAINING Zhuanghua Liu National University of Singapore

Menglian Wang Wonders Information Co., Ltd.

arXiv:2609.11655v1 [cs.LG] 10 Sep 2026

Luo Luo Fudan University

A BSTRACT Muon has emerged as a highly effective optimizer for large language model training, often achieving superior convergence and performance compared with the widely adopted Adam and AdamW optimizers. Nevertheless, Muon is prone to training instability due to its spectral flattening, manifested by loss spikes and unbounded growth of model weights. Existing approaches primarily rely on weight or attentionlogit clipping, which require architecture-specific modifications and do not directly address instability across all model components. We propose MomentUm SpEctral Clipping (Musec), which replaces Muon’s spectral flattening with spectral clipping: rather than setting all singular values of the momentum matrix to approximately one, Musec clips singular values that exceed a threshold while preserving the underlying spectral structure of the momentum. Our strategy provides an optimizer-level, architecture-agnostic mechanism for stabilizing Muon training. We further develop Soft Musec, an efficient implementation that uses a smooth spectral saturation function approximated by coupled Newton–Schulz iterations. Theoretically, we establish convergence guarantees for Musec in nonconvex nonsmooth stochastic optimization. To the best of our knowledge, this is the first convergence guarantee for Muon-type methods in the nonconvex nonsmooth setting. We provide empirical studies to show that Soft Musec consistently improves training stability over existing Muon variants across a wide range of learning rates and model sizes. Notably, Soft Musec remains stable in settings where existing Muon variants diverge, while matching their performance under well-tuned configurations.

1

I NTRODUCTION

Matrix-aware optimizers have recently emerged as strong alternatives to Adam (Kingma & Ba, 2015) and AdamW (Loshchilov & Hutter, 2019) for training large language models. Methods such as Shampoo (Gupta et al., 2018), SOAP (Vyas et al., 2025), Muon (Jordan et al., 2024; Liu et al., 2025), and Scion (Pethick et al., 2025) exploit the two-dimensional structure of weight matrices to achieve faster convergence. Among these, Muon stands out as the simplest and has already been adopted in several large-scale commercial LLM training efforts, including Kimi K2 (Kimi et al., 2025), GLM-5 (Zeng et al., 2026), and DeepSeek-V4 (Xu et al., 2026). Despite the fast convergence, Muon training is known to suffer from training instability, frequently exhibiting exploding weight norms and divergent attention logits (Kimi et al., 2025). Muon replaces the singular values of its momentum matrix with approximately one, producing an update with a flat singular spectrum. This spectral flattening can inject substantial update magnitude into directions that have relatively small singular values in the original momentum matrix, contributing to unstable weight growth (Kimi et al., 2025, Appendix E). Several approaches have been proposed to address this instability. Logit soft-capping (Gemma et al., 2024) bounds attention scores after computation, but does not prevent the underlying query-key dot products from growing excessively. QK-Norm (Dehghani et al., 2023; Wortsman et al., 2024) normalizes query and key vectors directly, but is incompatible with architectures such as multi-head latent attention (MLA) (Liu et al., 2024), where key matrices are not explicitly materialized during inference. MuonClip (Kimi et al., 2025) takes a 1

Preprint. Under review.

different approach by applying weight clipping to the query-key matrices to constrain attention logits. However, it is inherently architecture-specific and only constrains the query-key weight matrices, leaving the value-output and MLP weights entirely unaddressed. These approaches treat the symptoms of Muon’s instability rather than its root cause, which lies in Muon’s spectral transformation itself. In this paper, we propose MomentUm SpEctral Clipping (Musec), which replaces Muon’s spectral flattening with a spectral clipping operator. While Muon sets all singular values to one, Musec only clips singular values that exceed a threshold. The resulting update preserves the spectral skewness of the momentum matrix, thereby avoiding the amplification of directions associated with relatively small momentum singular values. Unlike existing remedies that target specific architectural components, such as capping attention logits (Gemma et al., 2024) or clipping query-key weights (Kimi et al., 2025), Musec operates directly within the optimizer and is therefore architecture-agnostic, stabilizing all weight matrices uniformly. Moreover, we show that Musec’s update admits a principled interpretation as constrained steepest descent: it solves a Frobenius-norm-penalized quadratic subproblem subject to a spectral-norm constraint, complementing the unconstrained steepest descent interpretation of Muon established by Bernstein & Newhouse (2024). We demonstrate the effectiveness of Musec both theoretically and empirically. Despite the growing empirical adoption of Muon-type optimizers, their convergence properties beyond the smooth setting are less explored. Since training objectives in modern deep neural networks are generally nonconvex and may involve nonsmooth components, such as piecewise linear activations and discrete routing mechanisms, we provide convergence guarantees for Musec in the general nonconvex nonsmooth setting. We summarize our main contributions: • Theoretically, we consider the following stochastic optimization problem: min

W ∈Rm×n

f (W ) = Eξ∼P [F (W ; ξ)],

(1)

where P is some unknown distribution, the objective function f (W ) is ρ-weakly convex and the stochastic component function F (W ; ξ) is possibly nonconvex and nonsmooth. We show that Musec converges to a (δ, ϵ)-Goldstein stationary point of the objective with O(r3/2 δ −1 ϵ−3 + δ 2 ρ3 r−3/2 + r3/4 δ −1 ) stochastic gradient oracle calls, where r = min(m, n) is the smaller matrix dimension. As long as the weak convexity parameter satisfies ρ ≤ O(rδ −1 ϵ−1 ), the dominant term of our complexity is O(r3/2 δ −1 ϵ−3 ). The dependence on the accuracy parameters δ and ϵ matches that of the optimal stochastic firstorder method for nonconvex nonsmooth optimization (Cutkosky et al., 2023). To the best of our knowledge, this constitutes the first convergence analysis of Muon-type algorithms in the nonconvex nonsmooth setting. • We further develop an efficient implementation, Soft Musec, which replaces hard spectral clipping with a smooth saturation function and approximates the resulting matrix transformation using coupled Newton-Schulz iterations. Empirically, Soft Musec achieves more stable training than existing Muon variants across multiple datasets, including FineWeb, OpenWebText, and C4, over a substantially wider range of learning rates. Notably, Soft Musec maintains stable convergence in settings where existing Muon variants diverge, while matching their performance under well-tuned configurations. Concurrent work. Concurrently with our work, Jiang et al. (2026) proposed spectral clipping as an optimizer-agnostic wrapper, SPECTRA, that can be applied to various base optimizers, including AdamW (Loshchilov & Hutter, 2019) and Signum (Bernstein et al., 2018). Yi (2026) studied singularvalue clipping as a replacement for Muon’s polar step (MuCon), with a focus on the numerical challenges of SVD-free approximation. We independently identify spectral clipping as a natural replacement for Muon’s spectral flattening, motivated by its ability to preserve the spectral structure of momentum matrices while suppressing excessively large singular values. Beyond this Muonspecific formulation and motivation, our work provides theoretical and empirical contributions that complement these concurrent studies: to the best of our knowledge, we present the first convergence analysis of Muon-type algorithms in the nonconvex nonsmooth setting, and demonstrate consistent stability improvements across multiple datasets and model sizes. Please refer to Section 4.1 and Appendix D for a detailed discussion. Paper Organization. Section 2 reviews related work on matrix-aware optimizers and nonconvex nonsmooth optimization. Section 3 establishes notation, formalizes assumptions, and reviews the 2

Preprint. Under review.

Muon optimizer along with its training instability. Section 4 introduces Musec and presents its convergence analysis. Section 5 develops Soft Musec, an efficient SVD-free implementation via coupled Newton–Schulz iterations. Section 6 provides experimental evaluation on NanoGPT models across multiple datasets and model scales. Section 7 concludes the paper.

2

R ELATED W ORK

Matrix-Aware Optimizers Matrix-aware optimizers exploit the two-dimensional structure of weight matrices to improve upon coordinate-wise methods such as Adam (Kingma & Ba, 2015). Shampoo (Gupta et al., 2018) and SOAP (Vyas et al., 2025) leverage matrix structure to construct Kronecker-factored or eigenspace-based preconditioners that incorporate approximate second-order information. Muon (Jordan et al., 2024) takes a fundamentally different approach by orthogonalizing the momentum matrix via Newton–Schulz iterations. Bernstein & Newhouse (2024) showed that this orthogonalization corresponds to steepest descent under the matrix spectral norm. Scion (Pethick et al., 2025) further develops this viewpoint through the Frank–Wolfe framework, interpreting the orthogonalized update as a linear minimization oracle over the spectral-norm ball. Subsequent work has explored variants of Muon, such as NorMuon (Li et al., 2026), which introduces neuron-wise adaptive scaling after orthogonalization. Concurrently with our work, SPECTRA (Jiang et al., 2026) proposed spectral clipping as an optimizer-agnostic wrapper applicable to various base optimizers, with convergence analysis limited to the convex smooth setting. MuCon (Yi, 2026) studied singularvalue clipping as a replacement for Muon’s polar step, focusing on numerical challenges without providing convergence guarantees or empirical validation. Yukhimchuk et al. (2026) independently proposed spectral clipping of gradient matrices to handle heavy-tailed noise. Their method clips gradients before the optimizer, whereas Musec integrates clipping into the momentum recurrence as a replacement for Muon’s polar step. On the theoretical side, Shen et al. (2025) provided a comprehensive convergence analysis of Muon under smoothness assumptions, characterizing conditions under which Muon can outperform gradient descent, and Gluon (Riabinin et al., 2025) established convergence guarantees for LMO-based optimizers under generalized smoothness. Recently, Yang et al. (2026) analyzed Spectral Descent in the nonsmooth setting, but under convexity and sharpness conditions. These existing results do not cover the general nonconvex nonsmooth stochastic setting, for which we establish the first convergence guarantees. Nonconvex Nonsmooth Optimization The foundations of nonsmooth optimization trace back to the seminal work of Clarke (1975) and Goldstein (1977). Non-asymptotic complexity guarantees for general nonconvex nonsmooth settings remained elusive until Zhang et al. (2020) established the first non-asymptotic complexity guarantees for subgradient methods converging to Goldstein stationary points. This breakthrough led to a series of subsequent developments (Davis et al., 2022; Tian et al., 2022; Kornowski & Shamir, 2022; Cutkosky et al., 2023; Jordan et al., 2023). In particular, Cutkosky et al. (2023) introduced the online-to-nonconvex (O2NC) framework, achieving the first optimal convergence rates for stochastic nonconvex nonsmooth optimization. Our analysis considers the class of weakly convex functions (Duchi & Ruan, 2018; Davis & Drusvyatskiy, 2019; Davis & Grimmer, 2019), which encompasses a broad class of objectives arising in neural network training. For this class of problem, Ji & Yuan (2026) showed that the O2NC framework can be derandomized while preserving the optimal O(δ −1 ϵ−3 ) complexity. However, convergence of Muon-type algorithms for general nonconvex nonsmooth stochastic objectives remains open. We address this gap from the view of online learning. By applying spectral clipping to the matrix-valued setting, we establish the first convergence guarantees for Muon-type algorithms in the nonconvex nonsmooth regime.

3

P RELIMINARIES AND BACKGROUND

In this section, we establish notation, formalize the problem setting and assumptions, and review the Muon optimizer along with its training instability. 3.1

N OTATIONS

Throughout this paper, ∥·∥2 and ∥·∥F denote the spectral norm and the Frobenius norm of a matrix, respectively, and ⟨·, ·⟩ denotes the Frobenius inner product. For a set of matrices Ω ⊆ Rm×n , we 3

Preprint. Under review.

denote dist(0, Ω) := inf M ∈Ω ∥M ∥F and conv(Ω) for the convex hull of Ω. For any positive integer N , we abbreviate [N ] = {1, . . . , N }. We use Bδ (W ) = {V ∈ Rm×n : ∥V − W ∥F ≤ δ} to denote the closed ball of radius δ centered at W ∈ Rm×n . For a scalar w ≥ 0 and a threshold D > 0, we define the clipping operator as clip(w, D) = min{w, D}. For a diagonal matrix M ∈ Rr×r , we apply the clipping entrywise: Clip(M , D) = Diag(clip(M1,1 , D), . . . , clip(Mr,r , D)), where Mi,i is the (i, i)-th entry of the matrix M . 3.2

P ROBLEM F ORMULATION

We now formalize the problem setting and state the assumptions used throughout this paper. We begin with the definitions of Lipschitz continuity and weak convexity. Definition 3.1. We say a function h : Rm×n → R is L-Lipschitz continuous if |h(W1 ) − h(W2 )| ≤ L ∥W1 − W2 ∥F , for all W1 , W2 ∈ Rm×n . The Clarke subdifferential (Clarke, 1990) of a Lipschitz function h at W ∈ Rm×n is denoted by ∂h(W ). We next define weakly convex functions. Definition 3.2. We say a function h : Rm×n → R is ρ-weakly convex for some ρ > 0 if the 2 quadratically regularized function h(·) + (ρ/2) ∥·∥F is convex, or equivalently ρ 2 h(W1 ) ≥ h(W2 ) + ⟨g, W1 − W2 ⟩ − ∥W1 − W2 ∥F , 2 for all W1 , W2 ∈ Rm×n , g ∈ ∂h(W2 ). In the remainder of this paper, we suppose Problem (1) satisfies the following assumptions. Assumption 3.3. For any ξ ∼ P, the loss function F (W ; ξ) is L-Lipschitz continuous with respect to its first argument. Furthermore, the objective f (W ) is ρ-weakly convex. Assumption 3.4. For each W ∈ Rm×n , we can access to a stochastic oracle G(W ; ξ) that is an unbiased estimator of a Clarke subgradient such that E[G(W ; ξ)] ∈ ∂f (W ). Given a positive constant σ > 0, we further assume bounded variance: 2

E[∥G(W ; ξ) − E[G(W ; ξ)]∥F ] ≤ σ 2 . We assume that the objective is bounded below, i.e., f ∗ = inf W ∈Rm×n f (W ) > −∞ and we denote ∆f := f (W0 ) − f ∗ . For general nonsmooth nonconvex objectives, convergence is standardly measured via Goldstein stationarity (Goldstein, 1977). Definition 3.5. The Goldstein δ-subdifferential of a Lipschitz function f at a point W ∈ Rm×n is the convex hull of all Clarke subgradients at points in a δ-ball around W , i.e.,    [  ∂δ f (W ) := conv ∂f (V ) .   V ∈Bδ (W )

A point W ∈ Rm×n is called a (δ, ϵ)-stationary point if dist(0, ∂δ f (W )) ≤ ϵ. 3.3

T HE M UON O PTIMIZER

Muon (Jordan et al., 2024) is an optimizer designed for matrix-valued parameters. At each iteration, it orthogonalizes the momentum matrix, replacing its nonzero singular values with one and thereby flattening the singular spectrum. The complete procedure is presented in Algorithm 1. Since computing the full SVD is prohibitively expensive in practice, Muon is typically implemented using Newton–Schulz iterations (Higham, 2008) to approximate this orthogonalization (Jordan et al., 2024; Bernstein & Newhouse, 2024). 4

Preprint. Under review.

Algorithm 1: Muon (Jordan et al., 2024) 1 Input: Initial point W0 , momentum parameter β ∈ [0, 1), learning rate η. 2 for n = 0, 1, . . . , N − 1 do 3 Sample ξn ∼ P 4 Gn = G(Wn ; ξn ) 5 if n = 0 then 6 M0 = G0 7 else 8 Mn = (1 − β)Mn−1 + βGn 9 end 10 (Un , Sn , Vn ) = SVD(Mn ) 11 Wn+1 = Wn − ηUn Vn⊤ 12 end

Algorithm 2: Musec: MomentUm SpEctral Clipping 1 Input: Initial point W0 , momentum parameter β ∈ [0, 1), learning rate η, clipping threshold D > 0, positive integers K and T 2 N =K ×T 3 for n = 0, 1, . . . , N − 1 do 4 Sample ξn ∼ P 5 Gn = G(Wn ; ξn ) 6 if n = 0 then c0 = G0 7 M 8 else cn = (1 − β)Mn−1 + βGn 9 M 10 end bn , Vn ) = SVD(M cn ) 11 (Un , S bn , D) 12 Sn = clip(S 13 Mn = Un Sn Vn⊤ 14 Wn+1 = Wn − ηMn 15 end 16 Set W

(k)

= T1

PT −1 t=0

(k)

Wt

(k)

where Wt

17 Return: W T ∼ Uniform({W

(k)

= W(k−1)T +t for ∀k ∈ [K].

: k ∈ [K]}).

Training instability of Muon. Let Mn = Un Sn Vn⊤ denote the SVD of the momentum matrix at step n. The Muon update Un Vn⊤ replaces all nonzero singular values of Mn with one, discarding the spectral magnitude information entirely. Consequently, directions associated with small momentum singular values receive updates of the same magnitude as those associated with large singular values, amplifying weakly represented directions and contributing to weight growth and training instability (Kimi et al., 2025, Appendix E). These observations motivate replacing spectral flattening with spectral clipping, which preserves the spectral structure of the momentum while constraining only excessively large singular values.

4

M ETHODOLOGY

In this section, we introduce MomentUm SpEctral Clipping (Musec) and establish its convergence guarantees for nonconvex nonsmooth optimization. 5

Preprint. Under review.

4.1

T HE A LGORITHM

We present the complete procedure of Musec in Algorithm 2. The key difference from Muon lies in Step 13, where we replace the orthogonalized momentum with a spectrally clipped momentum: bn , D). Mn = Un Sn Vn⊤ , where Sn = Clip(S bn , Vn ) = SVD(M cn ) is the singular value decomposition of the momentum M cn , Here, (Un , S b with Sn the diagonal matrix of its singular values, and D > 0 is the clipping threshold. Unlike Muon’s polar step, which maps all nonzero singular values to one, the clipping operator only truncates singular values exceeding the threshold D while leaving smaller singular values unchanged. Thus, Musec preserves the spectral structure of the momentum while controlling large spectral components. In contrast to SPECTRA (Jiang et al., 2026), which applies spectral clipping as a post-processing step to the accumulated momentum while carrying the unclipped momentum state forward, Musec feeds the clipped momentum Mn−1 back into the exponential moving average (EMA) update at the next step, integrating spectral clipping directly into the momentum recurrence. This ensures that the momentum state itself remains spectrally bounded, allowing singular values to decay in directions where recent gradients are small, rather than being dominated by historical accumulation. The return procedure in Steps 16–17 partitions the iterates into K consecutive epochs of length T , averages within each epoch, and returns one epoch average uniformly at random. This two-level averaging scheme is standard in the nonconvex nonsmooth optimization literature (Cutkosky et al., 2023; Ji & Yuan, 2026) and is required to establish convergence to Goldstein stationary points. In practice, it is sufficient to return the final iterate Wn , as is conventional in deep learning. 4.2

M USEC AS C ONSTRAINED S TEEPEST D ESCENT

Bernstein & Newhouse (2024) showed that Muon’s polar update solves the unconstrained steepest descent problem under the spectral norm. We establish an analogous characterization for Musec: its spectrally clipped update arises as the solution to a Frobenius-norm-penalized quadratic subproblem subject to a spectral-norm constraint. Proposition 4.1. Let G1 , . . . , GL be gradient matrices, let λ > 0 be a sharpness parameter, and let D > 0. For each l = 1, . . . , L, consider the problem   λ 2 min ⟨Gl , ∆l ⟩ + ∥∆l ∥F , s.t. ∥∆l ∥2 ≤ D/λ. (2) ∆l 2 where ⟨·, ·⟩ denotes the Frobenius inner product and ∆l has the same shape as Gl . Let Gl ∈ Rml ×nl have reduced SVD of the form Gl = Ul Sl Vl⊤ with Sl = diag(σ1 , . . . , σdl ) and dl = min(ml , nl ). Then Problem (2) is solved by ∆l = −η · Ul Ŝl Vl⊤ , where η =

1 , Ŝl = clip(Sl , D). λ

The spectral-norm constraint ∥∆Wl ∥2 ≤ D/λ clips singular values that exceed D while preserving smaller ones. Thus, rather than being an ad hoc modification of Muon, Musec admits a principled interpretation as constrained steepest descent under the Frobenius norm with a spectral-norm budget. 4.3

C ONVERGENCE A NALYSIS

In this subsection, we present the convergence analysis of Musec for solving the nonconvex nonsmooth optimization problem. We define the filtration Fn = σ(G1 , G2 , . . . , Gn ). Under Assumption 3.4, we can infer that (k) (k) Hn := E[Gn | Fn−1 ] ∈ ∂f (Wn ). Analogous to Wt , we define variables Mt = M(k−1)T +t (k)

= H(k−1)T +t for any k ∈ [K] and t ∈ {0} ∪ [T − 1], and we denote the epoch average PT (k) = t=1 Ht /T . We first establish the following convergence result for a single epoch k.

and Ht H

(k)

6

Preprint. Under review.

Lemma 4.2. Let γ = β/η and η ≤ 1/ρ where β ≤ 1/8, then for any k ∈ [K], the sequence (k) {Wt }Tt=1 generated by Algorithm 2 satisfies h i (k) (k) E f (WT ) − f (W0 ) # " √   T −1 X 2 βDT βDσ T β βT 1 β 2 L2 T rD2 (k) (k) ≤−E Mt + + H + + D2 + + . γ 8γ γ γ γ γ γ F F t=0 By combining Lemma 4.2 with a bound relating H

(k) F

to the Goldstein δ-subdifferential

dist(0, ∂δ f (W k )), we obtain the following convergence guarantee for Musec (Algorithm 2) in finding a (δ, ϵ)-Goldstein stationary point of nonconvex nonsmooth problem. √ Theorem 4.3. Let γ = β/η and η ≤ 1/ρ where β ≤ 1/8, Then for any δ ≥ ηT D r, the sequence {W

(k) K }k=1 generated by Algorithm 2 satisfies

"

#   K 1 X σ 2r βL2 γ∆f (k) E dist(0, ∂δ f (W )) ≤ √ + 1 + D+ + . K βT D βDT K T k=1

The following corollary, which follows directly from Theorem 4.3, establishes the oracle complexity of Musec for finding a (δ, ϵ)-Goldstein stationary point. Corollary 4.4. By choosing the parameters as T = (δN )2/3 , β=

r1/2 , (δN )2/3

K = δ −2/3 N 1/3 , γ=

r , 4/3 δ N 1/3

D = (δN )−1/3 , η=√

δ 2/3 , rN 1/3

then it takes Algorithm 2 at most N =O

r3/2 (σ 3 + L6 + ∆3f ) r3/4 δ 2 ρ3 + 3/2 + δ δϵ3 r

!

to obtain a (δ, ϵ)-Goldstein stationary point. Remark 4.5. When the weak convexity parameter satisfies ρ ≤ O(rδ −1 ϵ−1 ), the dominant term of the complexity reduces to O(r3/2 δ −1 ϵ−3 ). The dependence on the accuracy parameters δ and ϵ matches that of the optimal stochastic first-order method for nonconvex nonsmooth optimization (Cutkosky et al., 2023). The factor of r3/2 arises from the matrix-valued structure of the updates. For comparison, the convergence analysis of Muon under smoothness assumptions by Shen et al. (2025) incurs a factor of r2 in the stochastic setting, though we note that the two results target different stationarity concepts and operate under different assumptions.

5

A P RACTICAL I MPLEMENTATION

Computing the exact spectral clipping operator requires a full SVD decomposition at each iteration (step 11 in Algorithm 2), which is prohibitively expensive for the large weight matrices encountered in LLM training. In this section, we develop an efficient SVD-free approximation of the clipping operation. We term the resulting algorithm Soft Musec. Specifically, we replace the hard clipping function clip(w, D) = min(w, D) with the smooth saturation function Dw h(w, D) = √ , w2 + D2 which satisfies that h(w, D) ≈ w when w ≪ D and h(w, D) ≈ D when w ≫ D. Applying h(·, D) to each singular value of a matrix M ∈ Rm×n while preserving its singular vectors yields the soft spectral clipping operator: H(M , D) = D(M M ⊤ + D2 I)−1/2 M . 7

(3)

Preprint. Under review.

The key computational challenge in evaluating the operator (3) is the matrix inverse square root (M M ⊤ + D2 I)−1/2 . We approximate this using coupled Newton–Schulz iterations (Higham 2008, Chapter 6; Jiang et al. 2026), which involve only matrix–matrix multiplications and are therefore well-suited to modern GPU hardware. The addition of D2 I shifts all eigenvalues away from zero, ensuring numerical stability and rapid convergence of the iteration. In practice, we find that five Newton–Schulz iterations yield strong training performance in bfloat16 arithmetic. The complete Soft Musec algorithm is provided in Algorithm 4 in Appendix C.

6

E XPERIMENTS

In this section, we evaluate the effectiveness and stability of Soft Musec on LLM training. 6.1

E XPERIMENTAL S ETUP

We compare Soft Musec against three baselines: Muon (Jordan et al., 2024), MuonClip (Kimi et al., 2025), and SPECTRA (Jiang et al., 2026). For SPECTRA, we use SGDM (stochastic gradient descent with momentum) as the base optimizer, which corresponds to Muon without the orthogonalization step. This isolates the effect of spectral clipping from Muon’s orthogonalization. We note that Jiang et al. (2026) evaluated SPECTRA on AdEMAMix (Pagliardini et al., 2025), AdamW (Loshchilov & Hutter, 2019), and Signum (Bernstein et al., 2018); to our knowledge, our experiments provide the first empirical evaluation of spectral clipping applied to Muon in this setting. For all methods, we tune weight decay over {0.096, 0.3, 0.6, 1.2}. All methods use five Newton–Schulz iterations. For Soft Musec and SPECTRA, we additionally tune the clipping threshold D over {0.05, 0.125, 0.25, 0.5, 0.75, 1.0}. We adopt the modded-nanogpt codebase1 , a decoder-only Transformer with modern architectural modifications, including RMSNorm (Zhang & Sennrich, 2019), Rotary Positional Embeddings (RoPE) (Su et al., 2024), and square ReLU activations (So et al., 2021). We evaluate three model configurations—NanoGPT-Small (491M parameters), NanoGPT-Medium (613M parameters), and NanoGPT-Wide (1.63B parameters)—across three datasets: FineWeb (Penedo et al., 2024), OpenWebText (Gokaslan & Cohen, 2019), and C4 (Raffel et al., 2020). For all experiments, the respective optimizer is applied to matrix parameters in non-embedding layers (attention and MLP weight matrices), while AdamW is used for embedding layers and vector parameters. Full architectural and hyperparameter details are provided in Appendix E. 6.2

S TABILITY ACROSS L EARNING R ATES

We evaluate each method across learning rates ranging from 0.01 to 0.8 for each model configuration on FineWeb. The exact grid for each configuration is provided in Appendix E.3. Results are shown in Figure 1; runs that diverge during training are omitted. For NanoGPT-Small, all methods achieve comparable validation loss at smaller learning rates. As the learning rate increases, Soft Musec maintains stable convergence, while Muon and MuonClip begin to degrade. At the largest learning rate, Muon fails to converge across all model sizes. MuonClip also fails to converge on NanoGPT-Small and NanoGPT-Medium, and converges to substantially worse validation loss on NanoGPT-Wide. In contrast, Soft Musec and SPECTRA remain stable across all configurations. Across all configurations, Soft Musec achieves performance comparable to SPECTRA, with slight improvements in several settings. Consistent results on OpenWebText and C4 are reported in Appendix F.1. 6.3

T RAINING DYNAMICS

To examine training behavior in greater detail, we train NanoGPT-Medium at a learning rate of 0.2 on all three datasets and plot the validation loss over training steps in Figure 2. All methods exhibit a transient loss spike near step 400, which coincides with a scheduled transition in the modded-nanogpt training pipeline: the learning rate increases by 52%, the batch size doubles, and the attention window widens simultaneously. Notably, Soft Musec and SPECTRA recover from 1

https://github.com/kellerjordan/modded-nanogpt

8

Preprint. Under review.

3.6 3.4 3.2

10 1 Learning rate (a) NanoGPT-small

3.4

Muon MuonClip SPECTRA Soft Musec

Muon MuonClip SPECTRA Soft Musec

Final validation loss

Muon MuonClip SPECTRA Soft Musec

Final validation loss

Final validation loss

3.7 3.6 3.5 3.4 3.3 10 2

3.2 3.0 2.8

3.0 10 2

10 1 Learning rate (b) NanoGPT-medium

10 1 Learning rate (c) NanoGPT-wide

10 2

Figure 1: Validation loss versus learning rate on FineWeb across three model configurations. this spike quickly and smoothly, while Muon recovers slowly and converges to a substantially higher final validation loss. MuonClip improves over Muon but displays noticeable oscillations throughout training, suggesting that QK-weight clipping alone does not fully stabilize the optimization trajectory. In contrast, both Soft Musec and SPECTRA converge smoothly and reach comparable final validation loss. Complete training dynamics across all datasets and model configurations are reported in Appendix F.2.

4 3

5 4 3

0

2000 4000 Iterations (a) FineWeb

0

2000 4000 Iterations (b) OpenWebText

Muon MuonClip SPECTRA Soft Musec

6

Validation loss

Validation loss

5

Muon MuonClip SPECTRA Soft Musec

6

Validation loss

Muon MuonClip SPECTRA Soft Musec

6

5 4 3

0

2000 Iterations (c) C4

4000

Figure 2: Validation loss versus steps of training NanoGPT-Medium on the three datasets.

6.4

W EIGHT N ORM A NALYSIS

To understand the source of the stability improvements, we track the spectral norm of weight matrices across three component types: query-key (QK) projections, value-output (VO) projections, and MLP weights, during NanoGPT-Medium training on FineWeb at learning rate η = 0.2. Results are shown in Figure 3. Muon exhibits severe spectral-norm inflation across all three components, with norms reaching 300–400 and displaying large oscillations. This is consistent with the spectral flattening mechanism discussed in Section 3: by assigning comparable update magnitudes to all singular directions, Muon amplifies weakly represented directions, leading to substantial weight growth. The unchecked growth of QK norms is particularly concerning, as it can drive attention logit explosion (Kimi et al., 2025). MuonClip partially mitigates QK norm growth but leaves VO and MLP norms elevated, confirming that it only regulates the query-key matrices. In contrast, both Soft Musec and SPECTRA maintain spectral norms below 10 across all components, demonstrating that optimizer-level spectral clipping provides uniform stabilization across the entire model.

7

C ONCLUSION

In this paper, we introduced Musec, which replaces Muon’s spectral flattening with a spectral clipping operator that preserves the spectral structure of the momentum matrix while constraining excessively large singular values. We established the first convergence analysis of Muon-type algorithms in the nonconvex nonsmooth setting, with a rate matching the optimal dependence on the accuracy parameters for stochastic nonconvex nonsmooth optimization. On the practical side, we developed Soft Musec, an efficient SVD-free implementation based on coupled Newton–Schulz iterations, and 9

Preprint. Under review.

0

2000 4000 Iterations (a) QK projections

Muon MuonClip SPECTRA Soft Musec

300

Muon MuonClip SPECTRA Soft Musec

200

MLP

400 300 200 100 0

VO

Muon MuonClip SPECTRA Soft Musec

QK

400 300 200 100 0

100

0

2000 4000 Iterations (b) VO projections

0

0

2000 4000 Iterations (c) MLP weights

Figure 3: Spectral norm of weight matrices over training steps for NanoGPT-Medium on FineWeb. demonstrated that it substantially improves training stability over existing Muon variants across multiple datasets and model sizes.

R EFERENCES Jeremy Bernstein and Laker Newhouse. Old optimizer, new norm: An anthology. arXiv preprint arXiv:2409.20325, 2024. Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Animashree Anandkumar. signsgd: Compressed optimisation for non-convex problems. In International conference on machine learning, pp. 560–569. PMLR, 2018. Frank H Clarke. Generalized gradients and applications. Transactions of the American Mathematical Society, 205:247–262, 1975. Frank H. Clarke. Optimization and Nonsmooth Analysis, volume 5 of Classics in Applied Mathematics. Society for Industrial and Applied Mathematics (SIAM), Philadelphia, 1990. Ashok Cutkosky, Harsh Mehta, and Francesco Orabona. Optimal stochastic non-smooth non-convex optimization through online-to-non-convex conversion. In International Conference on Machine Learning, pp. 6643–6670. PMLR, 2023. Damek Davis and Dmitriy Drusvyatskiy. Stochastic model-based minimization of weakly convex functions. SIAM Journal on Optimization, 29(1):207–239, 2019. Damek Davis and Benjamin Grimmer. Proximally guided stochastic subgradient method for nonsmooth, nonconvex problems. SIAM Journal on Optimization, 29(3):1908–1930, 2019. Damek Davis, Dmitriy Drusvyatskiy, Yin Tat Lee, Swati Padmanabhan, and Guanghao Ye. A gradient sampling method with complexity guarantees for lipschitz functions in high and low dimensions. Advances in neural information processing systems, 35:6692–6703, 2022. Mostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski, Jonathan Heek, Justin Gilmer, Andreas Peter Steiner, Mathilde Caron, Robert Geirhos, Ibrahim Alabdulmohsin, et al. Scaling vision transformers to 22 billion parameters. In International conference on machine learning, pp. 7480–7512. PMLR, 2023. John C Duchi and Feng Ruan. Stochastic methods for composite and weakly convex optimization problems. SIAM Journal on Optimization, 28(4):3229–3259, 2018. Quentin Garrido, Randall Balestriero, Laurent Najman, and Yann Lecun. Rankme: Assessing the downstream performance of pretrained self-supervised representations by their rank. In International conference on machine learning, pp. 10929–10974. PMLR, 2023. Team Gemma, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, et al. Gemma 2: Improving open language models at a practical size. arXiv preprint arXiv:2408.00118, 2024. 10

Preprint. Under review.

Aaron Gokaslan and Vanya Cohen. Openwebtext corpus. http://Skylion007.github.io/ OpenWebTextCorpus, 2019. Allen A Goldstein. Optimization of lipschitz continuous functions. Mathematical Programming, 13 (1):14–22, 1977. Vineet Gupta, Tomer Koren, and Yoram Singer. Shampoo: Preconditioned stochastic tensor optimization. In International Conference on Machine Learning, pp. 1842–1850. PMLR, 2018. Nicholas J. Higham. Functions of Matrices: Theory and Computation. SIAM, 2008. Fanfan Ji and Xiaotong Yuan. Derandomized online-to-non-convex conversion for stochastic weakly convex optimization. In International Conference on Learning Representations, volume 2026, pp. 125389–125414, 2026. Xiaowen Jiang, Andrei Semenov, and Sebastian U. Stich. Enhancing llm training via spectral clipping. In Proceedings of the 43rd International Conference on Machine Learning, 2026. Keller Jordan, Yuchen Jin, Vlado Boza, You Jiacheng, Franz Cesista, Laker Newhouse, and Jeremy Bernstein. Muon: An optimizer for hidden layers in neural networks, 2024. URL https: //kellerjordan.github.io/posts/muon/. Michael Jordan, Guy Kornowski, Tianyi Lin, Ohad Shamir, and Manolis Zampetakis. Deterministic nonsmooth nonconvex optimization. In The Thirty Sixth Annual Conference on Learning Theory, pp. 4570–4597. PMLR, 2023. Team Kimi, Yifan Bai, Yiping Bao, Y Charles, Cheng Chen, Guanduo Chen, Haiting Chen, Huarong Chen, Jiahao Chen, Ningxin Chen, et al. Kimi k2: Open agentic intelligence. arXiv preprint arXiv:2507.20534, 2025. Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In International Conference on Learning Representations, 2015. Guy Kornowski and Ohad Shamir. Oracle complexity in nonsmooth nonconvex optimization. Journal of Machine Learning Research, 23(314):1–44, 2022. Zichong Li, Liming Liu, Chen Liang, Weizhu Chen, and Tuo Zhao. NorMuon: Making Muon more efficient and scalable. In Proceedings of the 43rd International Conference on Machine Learning, 2026. Aixin Liu, Bei Feng, Bin Wang, Bingxuan Wang, Bo Liu, Chenggang Zhao, Chengqi Dengr, Chong Ruan, Damai Dai, Daya Guo, et al. Deepseek-v2: A strong, economical, and efficient mixture-ofexperts language model. arXiv preprint arXiv:2405.04434, 2024. Jingyuan Liu, Jianlin Su, Xingcheng Yao, Zhejun Jiang, Guokun Lai, Yulun Du, Yidao Qin, Weixin Xu, Enzhe Lu, Junjie Yan, et al. Muon is scalable for llm training. arXiv preprint arXiv:2502.16982, 2025. Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations, 2019. Matteo Pagliardini, Pierre Ablin, and David Grangier. The ademamix optimizer: Better, faster, older. In International Conference on Learning Representations, volume 2025, pp. 64715–64757, 2025. Guilherme Penedo, Hynek Kydlı́ček, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, Thomas Wolf, et al. The fineweb datasets: Decanting the web for the finest text data at scale. Advances in Neural Information Processing Systems, 37:30811–30849, 2024. Thomas Pethick, Wanyun Xie, Kimon Antonakopoulos, Zhenyu Zhu, Antonio Silveti-Falls, and Volkan Cevher. Training deep learning models with norm-constrained LMOs. In Proceedings of the 42nd International Conference on Machine Learning, volume 267, pp. 49069–49104, 2025. Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21(140):1–67, 2020. 11

Preprint. Under review.

Artem Riabinin, Egor Shulgin, Kaja Gruntkowska, and Peter Richtárik. Gluon: Making muon & scion great again!(bridging theory and practice of lmo-based optimizers for llms). arXiv preprint arXiv:2505.13416, 2025. Olivier Roy and Martin Vetterli. The effective rank: A measure of effective dimensionality. In 2007 15th European signal processing conference, pp. 606–610. IEEE, 2007. Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, and Tri Dao. Flashattention-3: Fast and accurate attention with asynchrony and low-precision. Advances in Neural Information Processing Systems, 37:68658–68685, 2024. Wei Shen, Ruichuan Huang, Minhui Huang, Cong Shen, and Jiawei Zhang. On the convergence analysis of muon. arXiv preprint arXiv:2505.23737, 2025. David So, Wojciech Mańke, Hanxiao Liu, Zihang Dai, Noam Shazeer, and Quoc V Le. Searching for efficient transformers for language modeling. Advances in neural information processing systems, 34:6010–6022, 2021. Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 568:127063, 2024. Lai Tian, Kaiwen Zhou, and Anthony Man-Cho So. On the finite-time complexity and practical computation of approximate stationarity concepts of lipschitz functions. In International Conference on Machine Learning, pp. 21360–21379. PMLR, 2022. Nikhil Vyas, Depen Morwani, Rosie Zhao, Itai Shapira, David Brandfonbrener, Lucas Janson, and Sham Kakade. Soap: Improving and stabilizing shampoo using adam for language modeling. In International Conference on Learning Representations, volume 2025, pp. 93423–93444, 2025. Mitchell Wortsman, Peter Liu, Lechao Xiao, Katie Everett, Alexander Alemi, Ben Adlam, John D Co-Reyes, Izzeddin Gur, Abhishek Kumar, Roman Novak, et al. Small-scale proxies for largescale transformer training instabilities. In International Conference on Learning Representations, volume 2024, pp. 49844–49869, 2024. Anyi Xu, Bangcai Lin, Bing Xue, Bingxuan Wang, Bingzheng Xu, Bochao Wu, Bowei Zhang, Chaofan Lin, Chen Dong, Chenchen Ling, et al. Deepseek-v4: Towards highly efficient milliontoken context intelligence. arXiv preprint arXiv:2606.19348, 2026. Yixuan Yang, Yuqing He, and Song Li. Convergence of spectral descent for non-smooth optimization. arXiv preprint arXiv:2605.26977, 2026. Albert Yi. Mucon: Clipped muon updates for llm training. arXiv preprint arXiv:2605.26459, 2026. Alexander Yukhimchuk, Mladen Kolar, Martin Takáč, and Sayantan Choudhury. Gradient clipping beyond vector norms: A spectral approach for matrix-valued parameters. arXiv preprint arXiv:2605.11838, 2026. Aohan Zeng, Xin Lv, Zhenyu Hou, Zhengxiao Du, Qinkai Zheng, Bin Chen, Da Yin, Chendi Ge, Chenghua Huang, Chengxing Xie, et al. Glm-5: from vibe coding to agentic engineering. arXiv preprint arXiv:2602.15763, 2026. Biao Zhang and Rico Sennrich. Root mean square layer normalization. Advances in neural information processing systems, 32, 2019. Jingzhao Zhang, Hongzhou Lin, Stefanie Jegelka, Suvrit Sra, and Ali Jadbabaie. Complexity of finding stationary points of nonconvex nonsmooth functions. In International conference on machine learning, pp. 11173–11182. PMLR, 2020.

12

Preprint. Under review.

A PPENDIX The appendix is organized as follows. Section A proves that Musec’s update solves a Frobeniusnorm-penalized quadratic subproblem subject to a spectral-norm constraint. Section B provides the detailed convergence analysis of Musec for nonconvex nonsmooth optimization. Section C presents the complete Soft Musec algorithm with coupled Newton–Schulz iterations. Section D compares the update rules of Musec and SPECTRA. Section E details the architecture configurations, training configurations, and hyperparameter choices for the NanoGPT experiments. Section F provides additional learning rate sweep results, training dynamics, effective rank analysis, computational overhead, and ablation studies.

A

P ROOF OF P ROPOSITION 4.1

In this section, we give a formal proof of Proposition 4.1. We first restate the proposition as follows. Proposition A.1. Let G1 , . . . , GL be gradient matrices, let λ > 0 be a sharpness parameter, and let D > 0. For each l = 1, . . . , L, consider the problem   D λ 2 (4) arg min ⟨Gl , ∆l ⟩ + ∥∆l ∥F , s.t. ∥∆l ∥2 ≤ . 2 λ ∆l where ⟨·, ·⟩ denotes the Frobenius inner product and ∆l has the same shape as Gl . Let Gl ∈ Rml ×nl have reduced SVD of the form Gl = Ul Sl Vl⊤ with Sl = Diag(σ1 , . . . , σdl ) and dl = min(ml , nl ). Then Problem (2) is solved by 1 ∆l = −η · Ul Ŝl Vl⊤ , where η = , Ŝl = Clip(Sl , D). λ Proof. The Frobenius norm ∥∆l ∥F and the spectral norm ∥∆l ∥2 depend on ∆l only through its singular values, while by Von Neumann’s trace inequality the inner product ⟨Gl , ∆l ⟩ is minimized, for any fixed spectrum of ∆l , by aligning ∆l ’s singular vectors with those of −Gl . It therefore suffices to restrict attention to ∆l = −Ul Al Vl⊤ ,

Al = diag(a1 , . . . , adl ),

a1 ≥ a2 ≥ · · · ≥ adl ≥ 0.

Under this parametrization, ⟨Gl , ∆l ⟩ = −

dl X

σi ai ,

2 ∥∆l ∥F =

i=1

dl X

a2i ,

∥∆l ∥2 = a1 .

i=1

Since a1 ≥ · · · ≥ adl ≥ 0, the constraint a1 ≤ D/λ is equivalent to 0 ≤ al ≤ D/λ for every i, and Problem (2) reduces to the separable box-constrained program  dl  X λ min −σi ai + a2i . (5) 2 ai ∈[0,D/λ] i=1 Each scalar subproblem is a quadratic in ai with unconstrained minimizer σi /λ ≥ 0; clipping to the interval [0, D/λ] gives   σi D 1 a∗i = min , = min(σi , D). λ λ λ The ordering σ1 ≥ · · · ≥ σdl ensures a∗1 ≥ · · · ≥ a∗dl , so the monotonicity of the parametrization is automatically satisfied. Since a∗i = min(σi , D)/λ, we have A∗l = λ1 Ŝl with Ŝl = Clip(Sl , D), and hence ∆l = −Ul A∗l Vl⊤ = −η · Ul Ŝl Vl⊤ , with η = 1/λ, as claimed.

B

C ONVERGENCE A NALYSIS

In this section, we provide the detailed convergence analysis of the Musec algorithm. 13

Preprint. Under review.

B.1

R EGRET A NALYSIS OF O NLINE G RADIENT D ESCENT

Before presenting the convergence analysis of the Musec algorithm, we first establish a convergence result for online gradient descent on a quadratic objective, which will subsequently serve as the basis for analyzing the clipping operation in Musec. Consider a quadratic function of the form: β β 2 ft (X) := − ⟨Gt+1 , X⟩ + ∥X∥F γ 2γ in the matrix-valued variable X. We analyze the following standard online gradient descent (OGD) method with stepsize γ > 0 over a convex compact set C: Xt+1 :=ΠC (Xt − γ∇ft (Xt )) =ΠC (Xt + βGt+1 − βXt ) =ΠC ((1 − β)Xt + βGt+1 ) ,

(6)

where ΠC (·) denotes the Euclidean Projection associated with the convex and compact set C. The OGD step (6) is consistent with the update of the momentum variable Mn in Algorithm 2 when we choose C = {X : ∥X∥ ≤ D}. Let RegretT (X) denote the regret of the OGD algorithm with respect to some X ∈ C after T iterations, defined as:

RegretT (X) :=

T −1 X

ft (Xt ) −

t=0

T −1 X

ft (X).

t=0

We present the regret analysis of the OGD algorithm as follows. Lemma B.1. Suppose that β ≤ 1/8, then the OGD update (6) applied to the sequence of quadratic −1 functions {ft (X)}Tt=0 over the convex constraint set C guarantees that for all T ≥ 1 and all X ∈ C: RegretT (X) 2  2 T −1  X ∥X0 ∥F + X F β β2 β 2 2 2 ∥Gt+1 ∥F + . ≤ − ∥Xt ∥F + X F+ 8γ 2γ γ γ t=0

Proof. Fix any X ∈ C. By the non-expansiveness of the projection operator, 2

Xt+1 − X F 2

= ΠC (Xt − γ∇ft (Xt )) − X F 2

≤ Xt − γ∇ft (Xt ) − X F 2

2

= Xt − X F + γ 2 ∥∇ft (Xt )∥F − 2γ⟨∇ft (Xt ), Xt − X⟩. Rearranging the terms yields 2

2

2 Xt − X F − Xt+1 − X F γ ∥∇ft (Xt )∥F ⟨∇ft (Xt ), Xt − X⟩ ≤ + . 2γ 2

14

(7)

Preprint. Under review.

Accordingly, we can show that RegretT (X) =

T −1 X

ft (Xt ) −

t=0

≤

ft (X)

t=0

T −1  X t=0

T −1 X

β 2 Xt − X F ⟨∇ft (Xt ), Xt − X⟩ − 2γ



2 2 T −1 −1 2 X0 − X F XT − X F TX γ ∥∇ft (Xt )∥F β X 2 ≤− Xt − X F + − + 2γ t=0 2γ 2γ 2 t=0 2

2 2 T −1 −1 γ − β Gt+1 + β Xt ∥X0 ∥F + X F TX γ γ β X 2 F Xt − X F + ≤− + 2γ t=0 γ 2 t=0 ! ! 2 2 T −1 −1 2 2 ∥X0 ∥F + X F TX β 2 ∥Gt+1 ∥F β2 β X ∥Xt ∥F 2 2 − X F + + + ∥Xt ∥F ≤− 2γ t=0 2 γ γ γ t=0 2  2 T −1  X ∥X0 ∥F + X F β 1 β β2 2 2 2 = (β − ) ∥Xt ∥F + X F+ ∥Gt+1 ∥F + . γ 4 2γ γ γ t=0

The first inequality is due to the β/γ-strong convexity of ft (·); the second inequality follows from (7); the third inequality follows from Young’s inequality. The last inequality is due to the reverse Young’s 2 2 2 inequality: ∥A − B∥F ≥ ∥A∥F /2 − ∥B∥F . Since β ≤ 1/8, then we have 2  2 T −1  X ∥X0 ∥F + X F β β2 β 2 2 2 RegretT (X) ≤ − ∥Xt ∥F + X F+ ∥Gt+1 ∥F + . 8γ 2γ γ γ t=0

B.2

P ROOF OF L EMMA 4.2

We first restate Lemma 4.2 as follows. Lemma B.2. Let γ = β/η and η ≤ 1/ρ where β ≤ 1/8, then for any k ∈ [K], the sequence (k) {Wt }Tt=1 generated by Algorithm 2 satisfies h i (k) (k) E f (WT ) − f (W0 ) # " √   T −1 X 2 βT β βDσ T 1 β 2 L2 T rD2 βDT (k) (k) + + + + . ≤−E H + Mt D2 + γ 8γ γ γ γ γ γ F F t=0 Proof. Let us consider the filtration Fn = σ(G1 , G2 , . . . , Gn ), where σ(·) denotes the generated σ-field. From Assumption 3.4, we can infer that Hn := E[Gn | Fn−1 ] ∈ ∂f (Wn ). The weak convexity of f (·) implies that for all n ≥ 1 and η −1 ≥ ρ, 1 2 f (Wn+1 ) − f (Wn ) ≤⟨Hn+1 , Wn+1 − Wn ⟩ + ∥Wn+1 − Wn ∥F 2η   1 2 =E ⟨Gn+1 , Wn+1 − Wn ⟩ + ∥Wn+1 − Wn ∥F | Fn . 2η The law of total expectation implies that   1 2 ∥Wn+1 − Wn ∥F . E[f (Wn+1 ) − f (Wn )] ≤ E ⟨Gn+1 , Wn+1 − Wn ⟩ + 2η 15

Preprint. Under review.

(k)

Fix an arbitrary k ∈ [K]. Recall that Wt (k)

(k)

define the notions of Gt , Ht

= W(k−1)T +t where t ∈ {0} ∪ [T − 1] and we similarly

(k)

and Mt . We can express the previous inequality as

(k) (k) E[f (Wt+1 ) − f (Wt )] ≤ E

  2 1 (k) (k) (k) (k) (k) ⟨Gt+1 , Wt+1 − Wt ⟩ + Wt+1 − Wt . 2η F

Summing the above inequality over t = 0, . . . , T − 1 yields that

(k) (k) E[f (WT ) − f (W0 )] ≤ E

"T −1  # X 2 1 (k) (k) (k) (k) (k) ⟨Gt+1 , Wt+1 − Wt ⟩ + Wt+1 − Wt . (8) 2η F t=0

(k)

From the update rule of Wt+1 (step 14 in Algorithm 2), we have (k) (k) Wt − Wt+1 (k) Mt = ,

η

then we can rewrite the inequality (8) as:

(k) (k) E[f (WT ) − f (W0 )] ≤ E

"T −1  X

η (k) (k) −η⟨Gt+1 , Mt ⟩ +

2

t=0

(k) Mt

2

# .

(9)

F

If we choose β = ηγ, and we denote the quadratic function β β (k) (k) 2 ∥X∥F . ft (X) = − ⟨Gt+1 , X⟩ + γ 2γ

(10)

Since β ≤ 1/8, we can apply Lemma B.1 for the regret analysis with respect to the quadratic function (k)

(k)

ft (X) by taking Xt = Mt

and X = M

(k)

, C = {X : ∥X∥2 ≤ D}:

RegretT (M ) :=

T −1 X

(k)

(k)

ft (Mt ) −

T −1 X

t=0

≤

(k)

) (11)

t=0

T −1  X t=0

(k)

ft (M

−

2 2 β β β2 (k) 2 (k) (k) Mt + M + Gt+1 8γ 2γ γ F F F

 +

(k) M0

2

+ M F

(k) 2 F

γ

.

Substitute the definition of the quadratic function (10) into (11): T −1  X t=0

≤

T −1  X t=0

2 β β (k) (k) (k) Mt − ⟨Gt+1 , Mt ⟩ + γ 2γ F



β (k) β (k) (k) 2 − ⟨Gt+1 , M ⟩ + M γ 2γ F



(k) M0

+

2

+ M F

t=0

(k) 2 F

γ

+

T −1  X

.

16

2 2 β β β2 (k) 2 (k) (k) − Mt + M + Gt+1 8γ 2γ γ F F F



Preprint. Under review.

Taking expectations on both sides yields that: "T −1  # X 2 β (k) β (k) (k) Mt E − ⟨Gt+1 , Mt ⟩ + γ 2γ F t=0 "T −1   X β (k) β (k) (k) 2 ≤E − ⟨Gt+1 , M ⟩ + M γ 2γ F t=0 +

T −1  X

−

t=0

" =E −

+

2

β β (k) + Mt M 8γ 2γ F

+

F

2

(k)

2

(k) 2

β γ

2

(k)

+

Gt+1

+ M

M0



F

γ

F

 (k) 2 F

 (12)

T −1 T −1 β X (k) β X (k) (k) (k) (k) ⟨Gt+1 − Ht+1 , M ⟩ − ⟨Ht+1 , M ⟩ γ t=0 γ t=0

T −1  X

−

t=0

2

β β (k) + Mt M 8γ γ F

(k) 2 F

(k)

2

+

β γ



2

(k)

+

Gt+1

M0

2

+ M F

γ

F

 (k) 2 F

.

We choose M

(k)

(k)

PT

H = D P t=1 t (k) T t=1 Ht

,

(13)

= D.

(14)

F

then we can show that M

(k)

≤ M 2

(k) F

(k)

(k)

In addition, using Jensen’s inequality and the fact that {Gt+1 − Ht+1 } is a martingale-difference sequence with variance bounded by σ 2 (Assumption 3.4): " # " # T −1 T −1  β X (k) β X  (k) (k) (k) (k) (k) E − ⟨G − Ht+1 , M ⟩ ≤E M Gt+1 − Ht+1 γ t=0 t+1 γ t=0 F F (15) √ βDσ T ≤ , γ where the last inequality follows from Assumption 3.4. In addition, the definition (13) implies that  * " # + PT T −1 T (k) X D H β X β (k) (k) (k)  ⟨Ht+1 , M ⟩ = − E  E − H , P t=1 t (k) T γ t=0 γ t=1 t t=1 Ht F (16) " # T X βD (k) H . =−E γ t=1 t F

(k) Gt+1

Furthermore, Assumption 3.3 implies that ≤ L. Step (13) in Algorithm 2 yields that F √ (k) M0 ≤ rD. Consequently, substituting (14), (15), (16) into (12), we have F

"T −1  # X 2 β (k) β (k) (k) E − ⟨Gt+1 , Mt ⟩ + Mt γ 2γ F t=0 " #  √  T −1 T X β 2 βD X (k) βT D2 β 2 L2 T rD2 D2 βDσ T (k) −E Ht + Mt + + + + ≤ (17) γ γ t=1 8γ γ γ γ γ F t=0 F " # √   T −1 X 2 βDT β βDσ T βT 1 β 2 L2 T rD2 (k) (k) =−E H + Mt + + + D2 + + . γ 8γ γ γ γ γ γ F F t=0

17

Preprint. Under review.

where in the last equality follows from the definition H then combining (9) and (17) yields that (k)

(k)

= T1

(k) t=1 Ht . Recall that β = ηγ,

PT

(k)

E[f (WT ) − f (W0 )] # " √   T −1 X 2 βDσ T βDT β βT 1 β 2 L2 T rD2 (k) (k) + Mt ≤−E + H + + D2 + + . γ 8γ γ γ γ γ γ F F t=0

Before proceeding, we establish an auxiliary lemma that will be used to control the deviation of the (k) iterates Wt from their average. Pn Lemma B.3. Let X1 , . . . , Xn be a set of matrices, and let X = n1 i=1 Xi . Then for all i ∈ [n], n n X 1X 2 2 2 Xi − X F ≤ ∥Xi − Xi′ ∥F ≤ n ∥∆j ∥F , n ′ j=1 i =1

where ∆j := Xj − Xj−1 and X0 can be chosen arbitrarily. Proof. Fix any i ∈ [n], we have n

2

Xi − X F = Xi −

1X Xi′ n ′ i =1

1X 2 ∥Xi − Xi′ ∥F n ′ i =1

F

2

i∨i X

i =1

j=i∧i′ +1

(Xj−1 − Xj ) F

2

′

i∨i n 1 X X ≤ n ′ ′ i =1

n

≤

′

n

1X = n ′

2

(18)

∥∆j ∥F 

j=i∧i +1

2  n n X X 2 ∥∆j ∥ , ∥∆j ∥  ≤ n ≤ F

F

j=1

j=2

where the first and the last inequalities are due to the Cauchy–Schwarz inequality; The second inequality follows from the triangle inequality. B.3

P ROOF OF T HEOREM 4.3

We now provide the proof of the main convergence theorem of the Musec. We first restate Theorem 4.3 as follows: √ Theorem B.4. Let γ = β/η and η ≤ 1/ρ where β ≤ 1/8, Then for any δ ≥ ηT D r, the sequence {W

(k) K }k=1 generated by Algorithm 2 satisfies " # K 1 X σ (k) E dist(0, ∂δ f (W )) ≤ √

K

  2r βL2 γ∆f + 1+ D+ + . βT D βDT K T

k=1

Proof. For any k ∈ [K], according to Lemma 4.2, one has (k)

(k)

E[f (WT ) − f (W0 )] " # √ T −1 X 2 βDT β βDσ T (k) (k) ≤−E H + Mt + γ 8γ γ F F t=0   βT 1 β 2 L2 T rD2 + + D2 + + γ γ γ γ √     βDT βDσ T βT 2r β 2 L2 T (k) + + D2 + . ≤−E H + γ γ γ γ γ F 18

Preprint. Under review.

(k)

By the definition that WT

(k+1)

= W0

, we obtain

(k+1) (k) E[f (W0 ) − f (W0 )]

√    βDσ T βDT βT 2r β 2 L2 T (k) + ≤−E H + + D2 + . γ γ γ γ γ F 

Rearranging the terms on both sides of the inequality, we have √     βDT βDσ T βT 2r β 2 L2 T (k) (k) (k+1) E H ≤ + + D2 + + E[f (W0 ) − f (W0 )]. γ γ γ γ γ F By telescoping both sides through k ∈ [K], we have "

K

βDT X (k) E H γ F k=1

#

√   βDσ T K βT 2r β 2 L2 T K (0) (K+1) ≤ + + D2 K + + E[f (W0 ) − f (W0 )] γ γ γ γ √   βT βDσ T K 2r β 2 L2 T K ≤ + + + ∆f , D2 K + γ γ γ γ

where the last inequality follows from the definition of ∆f . Dividing both sides by βDT K/γ yields that " #   K 1 X 2r γ∆f σ βL2 (k) √ E + 1+ + . H ≤ D+ K βT D βDT K F T k=1 To translate the bound on H (k)

iterate Wt that (k)

Wt

(k) F

into a Goldstein-stationarity bound at W

to lie within the δ-ball of W

−W

(k) F

(k)

(k)

, we need every

. Applying Lemma B.3, for any t ∈ [T ], we can show

v v u T u T −1 u X u X 2 2 √ (k) (k) (k) t ≤ T Wt − Wt−1 ≤ tT η 2 Mt ≤ ηT D r ≤ δ, F

t=1

F

t=0

where the second inequality follows from the step (14) in Algorithm 2 and the third inequality is due (k)

to Mt

2 F

≤ rD2 . It implies that dist(0, ∂δ f (W

(k)

T

1 X (k) H )) ≤ T t=1 t

= H

(k)

. F

F

Consequently, we obtain # "   K σ 2r γ∆f 1 X βL2 (k) dist(0, ∂δ f (W )) ≤ √ + 1 + + . E D+ K βT D βDT K T k=1

B.4

P ROOF OF C OROLLARY 4.4

We restate the complexity bound of Musec as follows. Corollary B.5. By choosing the parameters as T = (δN )2/3 , β=

r1/2 , (δN )2/3

K = δ −2/3 N 1/3 , γ=

r , 4/3 δ N 1/3

D = (δN )−1/3 , η=√

δ 2/3 , rN 1/3

then it takes at most N =O

r3/2 (σ 3 + L6 + ∆3f ) r3/4 δ 2 ρ3 + 3/2 + δ δϵ3 r

to obtain a (δ, ϵ)-Goldstein stationary point. 19

!

Preprint. Under review.

Proof. The theorem imposes four constraints on the parameters: √ ηT D r ≤ δ,

1 , η −1 ≥ ρ 8 We first specify the parameters, then verify that each constraint is satisfied under a mild lower bound on N , and finally translate the theorem’s bound into the stated oracle complexity. β≤

β = ηγ,

Step 1: Parameter choices. Set T = (δN )2/3 , β=

r1/2 , (δN )2/3

N = δ −2/3 N 1/3 , D = (δN )−1/3 , T r δ 2/3 γ = 4/3 1/3 , η = √ 1/3 δ N rN

K=

(19)

Step 2: Verifying the constraints. A direct computation gives ηγ = √

δ 2/3 r · 4/3 1/3 = β, 1/3 rN δ N

so β = ηγ holds by construction. For the step-size constraint, √ √ δ 2/3 ηT D r = √ 1/3 (δN )1/3 r ≤ δ, rN √ so ηT D r ≤ δ also holds. The remaining two constraints β ≤ 18 and η −1 ≥ ρ translate into lower bounds on N :  3/4  r ρ3 δ 2 N ≥O , N ≥ 3/2 . δ r Step 3: Evaluating the theorem’s bound. By the theorem 4.3 implies that " #   K 2r σ βL2 ∆f γ 1 X (k) dist(0, ∂δ f (W )) ≤ √ + 1 + E D+ + K βT D βDT K T k=1

≤

σ

1 +

(δN ) 3

1

1

1 +

(δN ) 3

2r 2

1

1 +

(δN ) 3

1

r 2 L2

∆f r 2

(δN ) 3

(δN ) 3

1 +

1

,

where the last inequality follows from the parameter choices in (19). To obtain (δ, ϵ)-Goldstein stationary point such that it satisfies " # K 1 X (k) dist(0, ∂δ f (W )) ≤ ϵ, E K k=1

it takes at most 3

r 2 (σ 3 + L6 + ∆3f ) r4 δ 2 ρ3 + 3 + δ δϵ3 r2 3

N =O

!

stochastic gradient oracle queries.

C

D ETAILED A LGORITHM OF S OFT M USEC

In Section 5, we introduced Soft Musec, which replaces exact SVD-based spectral clipping with a smooth saturation function approximated via coupled Newton–Schulz iterations. Here we provide the full algorithmic details. Recall from Section 5 that the soft spectral clipping operator is given by H(M , D) = D(M M ⊤ + D2 I)−1/2 M , 20

Preprint. Under review.

Algorithm 3: Soft Spectral Clipping (SSC) 1 Input: Matrix M ∈ Rm×n , clipping threshold D > 0, number of Newton–Schulz iterations K. 2 A = M M ⊤ + D2 I 3 α = ∥A∥F 4 Y0 = A/α 5 Z0 = I 6 for k = 0, 1, . . . , K − 1 do 7 Tk = 12 (3I − Zk Yk ) 8 Yk+1 = Yk Tk 9 Zk+1 = Tk Zk 10 end √ 11 Return: D · (ZK / α) · M .

Algorithm 4: Soft Musec 1 Input: Initial point W0 , momentum parameter β ∈ [0, 1), learning rate η, clipping threshold D > 0, positive integers K and T , number of Newton–Schultz steps Kns 2 N =K ×T 3 for n = 0, 1, . . . , N − 1 do 4 Sample ξn ∼ P 5 Gn = G(Wn ; ξn ) 6 if n = 0 then c0 = G0 7 M 8 else cn = (1 − β)Mn−1 + βGn 9 M 10 end cn , D, Kns ) // Algorithm 3 11 Mn = SSC(M 12 Wn+1 = Wn − ηMn 13 end PT −1 (k) (k) (k) = T1 t=0 Wt where Wt = W(k−1)T +t for ∀k ∈ [K] 14 Set W 15 Return: W T ∼ Uniform({W

(k)

: k ∈ [K]}).

where M ∈ Rm×n with m ≤ n w.l.o.g. We approximate the matrix inverse square root (M M ⊤ + D2 I)−1/2 using coupled Newton–Schulz iterations (Higham, 2008, Chapter 6), which involve only matrix– matrix multiplications and are therefore well-suited to GPU computation. The complete procedure is presented in Algorithm 3. The normalization by M M ⊤ + D2 I F ensures that the initial iterate Y0 has unit Frobenius norm, placing the eigenvalues in a range where the Newton–Schulz iteration converges. The addition of D2 I shifts all eigenvalues away from zero, ensuring numerical stability. We present the full Soft Musec procedure in Algorithm 4, which integrates the soft spectral clipping subroutine (Algorithm 3) into the Musec framework by replacing the exact SVD and hard clipping in Steps 11–14 of Algorithm 2.

D

C OMPARISON WITH SPECTRA

We provide a detailed comparison of the Musec and SPECTRA update rules to clarify their algorithmic differences. 21

Preprint. Under review.

Given a matrix M with reduced SVD M = U SV ⊤ , we define the spectral clipping operator Clipsc (M ) = U Clip(S, D)V ⊤ . Both Musec and SPECTRA employ this operator, but they differ in how the clipped output interacts with the momentum state. Musec maintains a clipped momentum buffer and computes:  cn = (1 − β)Mn−1 + βGn ,   M cn , D), Mn = Clipsc (M   W n+1 = Wn − ηClipsc Mn , Crucially, the clipped momentum Mn is propagated to the next iteration. This ensures that the momentum state is always spectrally bounded: ∥Mn ∥ ≤ D at every iteration. Ignoring weight decay, the SPECTRA update (Jiang et al., 2026, Eq. (4)) maintains an unclipped momentum buffer and applies spectral clipping only when computing the parameter update:  Mn = (1 − β)Mn−1 + βGn . Wn+1 = Wn − ηClipsc (Mn , D), Here, the unclipped momentum Mn is carried forward to the next iteration. Although the weight update is spectrally clipped, the momentum state itself is unconstrained. Key difference. The distinction lies in whether spectral clipping is integrated into the momentum recurrence. In Musec, the momentum state satisfies ∥Mn ∥ ≤ D at every iteration, so the input to the next EMA step is always spectrally bounded. In SPECTRA, the momentum state is unconstrained: ∥Mn ∥ can grow well beyond D over successive iterations, as clipping is only applied to the output update and does not feed back into the recurrence. Consequently, SPECTRA’s momentum can retain large spectral components accumulated over many past iterations, even when recent gradients are small in those directions. Musec’s feedback mechanism prevents this accumulation, providing tighter spectral control over the optimization trajectory.

E

E XPERIMENTAL D ETAILS

This section provides the detailed architecture, training, and hyperparameter configurations for our NanoGPT experiments. E.1

A RCHITECTURE C ONFIGURATIONS Table 1: Architecture specifications across model scales. Layers Attention heads Head dimension Model dimension MLP hidden dimension Max sequence length Core parameters

Small

Medium

Wide

11 6 128 768 3,072 2,048 ∼491M

16 8 128 1,024 4,096 4,096 ∼613M

16 16 128 2,048 8,192 4,096 ∼1.63B

In our experiments, we evaluate three model configurations, including NanoGPT-small, NanoGPTMedium, and NanoGPT-Wide. All three model configurations follow a decoder-only Transformer architecture based on the modded-nanogpt codebase.2 The architecture incorporates several modern modifications on top of the standard GPT-2 backbone, including RMSNorm (Zhang & Sennrich, 2019), squared ReLU activations (So et al., 2021), Rotary Position Embeddings (Su et al., 2024), sliding-window causal attention via Flash Attention 3 (Shah et al., 2024), and multi-token prediction as an auxiliary training objective. We refer to the codebase for full architectural details. The three configurations differ primarily in their width scaling. Table 1 summarizes their specifications. 2

https://github.com/KellerJordan/modded-nanogpt

22

Preprint. Under review.

E.2

T RAINING C ONFIGURATION Table 2: Training configuration across model scales. Total steps Final batch size (tokens) Precision Hardware

Small

Medium

Wide

1,390 393K bfloat16 H800

4,740 524K bfloat16 H800

8,040 1,049K bfloat16 H800

All models are trained using the GPT-2 BPE tokenizer in bfloat16 precision on NVIDIA H800 GPUs. We train all configurations on FineWeb (Penedo et al., 2024), OpenWebText (Gokaslan & Cohen, 2019), and C4 (Raffel et al., 2020). Each configuration employs a multi-phase training schedule that jointly ramps up batch size and learning rate. The learning rate follows a warmup-then-decay schedule. Table 2 summarizes the key training parameters. Full schedule details are available in the modded-nanogpt codebase. E.3

H YPERPARAMETERS

Table 3: Selected weight decay λ and clipping threshold D for each method and learning rate on NanoGPT-Small. Each entry corresponds to the best validation loss over the sweep. Soft Musec and SPECTRA share the same hyperparameter space. Muon

MuonClip

Soft Musec / SPECTRA

LR

λQK

λVO

λMLP

λQK

λVO

λMLP

λQK

λVO

λMLP

D

0.01 0.023 0.05 0.1 0.2 0.5

1.2 1.2 1.2 1.2 1.2 1.2

1.2 1.2 1.2 1.2 1.2 1.2

0.3 1.2 1.2 0.3 0.3 0.3

1.2 1.2 1.2 1.2 1.2 1.2

1.2 1.2 1.2 1.2 1.2 1.2

0.3 1.2 1.2 0.3 0.3 0.3

1.2 1.2 1.2 1.2 1.2 1.2

1.2 1.2 1.2 1.2 1.2 1.2

0.3 0.3 0.3 0.3 0.3 0.3

0.5 0.5 0.1 0.1 0.1 0.1

Table 4: Selected weight decay λ and clipping threshold D for each method and learning rate on NanoGPT-Medium. Each entry corresponds to the best validation loss over the sweep. Soft Musec and SPECTRA share the same hyperparameter space. Muon

MuonClip

Soft Musec / SPECTRA

LR

λQK

λVO

λMLP

λQK

λVO

λMLP

λQK

λVO

λMLP

D

0.01 0.015 0.05 0.1 0.2 0.3 0.5 0.8

1.2 1.2 1.2 1.2 1.2 1.2 0.3 0.3

1.2 1.2 1.2 1.2 1.2 1.2 0.3 0.3

1.2 1.2 1.2 0.3 0.3 0.3 0.3 0.1

1.2 1.2 1.2 1.2 1.2 1.2 0.3 0.3

1.2 1.2 1.2 1.2 1.2 1.2 0.3 0.3

1.2 1.2 1.2 1.2 1.2 0.3 0.3 0.1

1.2 1.2 1.2 1.2 1.2 1.2 0.3 0.3

1.2 1.2 1.2 1.2 1.2 1.2 0.3 0.3

0.3 0.3 0.3 0.3 0.3 0.3 0.3 0.1

0.25 0.25 0.05 0.05 0.05 0.05 0.05 0.05

For all methods, the respective optimizer is applied to attention and MLP projection matrices, while AdamW is used for embeddings, the language-model head, and scalar parameters. AdamW hyperparameters are held fixed across all methods. For each method, model configuration, and learning rate, we tune the weight decay over {0.096, 0.3, 0.6, 1.2}. For Soft Musec and SPECTRA, we additionally tune the clipping threshold D over {0.05, 0.125, 0.25, 0.5, 0.75, 1.0}. We report the best validation loss over these sweeps. Tables 3, 4, and 5 list the selected hyperparameters for NanoGPT-Small, NanoGPT-Medium, and NanoGPT-Wide, respectively. 23

Preprint. Under review.

Table 5: Selected weight decay λ and clipping threshold D for each method and learning rate on NanoGPT-Wide. Each entry corresponds to the best validation loss over the sweep. Soft Musec and SPECTRA share the same hyperparameter space. Muon

F

MuonClip

Soft Musec / SPECTRA

LR

λQK

λVO

λMLP

λQK

λVO

λMLP

λQK

λVO

λMLP

D

0.01 0.05 0.1 0.2 0.3 0.5

1.2 1.2 1.2 1.2 1.2 0.3

1.2 1.2 1.2 1.2 1.2 0.3

1.2 1.2 0.3 0.3 0.3 0.3

1.2 1.2 1.2 1.2 1.2 0.3

1.2 1.2 1.2 1.2 1.2 0.3

1.2 1.2 1.2 1.2 0.3 0.3

1.2 1.2 1.2 1.2 1.2 0.3

1.2 1.2 1.2 1.2 1.2 0.3

0.3 0.3 0.3 0.3 0.3 0.3

0.25 0.05 0.05 0.05 0.05 0.05

A DDITIONAL E XPERIMENTAL R ESULTS

This section presents additional experimental results comparing Soft Musec against baseline methods on NanoGPT training.

Final validation loss

3.7 3.6 3.5 3.4 3.3 10 2

Muon MuonClip SPECTRA Soft Musec

Muon MuonClip SPECTRA Soft Musec

3.4 3.2

10 1 Learning rate (a) FineWeb

Final validation loss

L EARNING R ATE S WEEP Final validation loss

F.1

3.6

Muon MuonClip SPECTRA Soft Musec

3.4

10 1 Learning rate (b) OpenWebText

10 2

10 2

10 1 Learning rate (c) C4

Figure 4: Validation loss versus learning rate for NanoGPT-Small across three datasets.

3.2 3.0 10 2

3.4

Muon MuonClip SPECTRA Soft Musec

3.2 3.0 2.8

10 1

Learning rate (a) FineWeb

3.6

Final validation loss

3.4

Muon MuonClip SPECTRA Soft Musec

Final validation loss

Final validation loss

3.6

3.4 3.2

Muon MuonClip SPECTRA Soft Musec

3.0

10 1

10 2

Learning rate (b) OpenWebText

10 2

10 1 Learning rate (c) C4

Figure 5: Validation loss versus learning rate for NanoGPT-Medium across three datasets. Figures 4–5 present the complete learning rate sweep results across all model configurations and datasets; runs that diverge during training are omitted. For NanoGPT-Small (Figure 4), all methods achieve comparable validation loss at small learning rates. As the learning rate increases, Muon degrades most rapidly, followed by MuonClip, while Soft Musec and SPECTRA maintain lower validation loss across the full range. For NanoGPT-Medium (Figure 5), the advantage of spectral clipping methods becomes more pronounced: at learning rates η ≥ 0.1, both Soft Musec and SPECTRA substantially outperform Muon and MuonClip, with Soft Musec achieving the best or near-best validation loss at the largest learning rates across all three datasets. 24

Preprint. Under review.

Muon MuonClip SPECTRA Soft Musec

Validation loss

η = 0.1

6 5 4 3 0

500 1000 Iterations Muon MuonClip SPECTRA Soft Musec

Validation loss

6

η = 0.2

500 1000 Iterations

5 4 0

Validation loss

10

Muon MuonClip SPECTRA Soft Musec

8 6 4

500 1000 Iterations

0

500 1000 Iterations

0

500 1000 Iterations Muon MuonClip SPECTRA Soft Musec

6 5 4

Validation loss

Validation loss

3

Adam AdamW Muon MuonClip SPECTRA Soft Musec

6 5 4 3

0

500 1000 Iterations Muon MuonClip SPECTRA Soft Musec

6 5 4

3

0

500 1000 Iterations Muon MuonClip SPECTRA Soft Musec

6 5 4

3

0

500 1000 Iterations Muon MuonClip SPECTRA Soft Musec

6

Validation loss

0

C4

5 4

3

0

500 1000 Iterations

10

Muon MuonClip SPECTRA Soft Musec

8 6 4 0

500 1000 Iterations

0

8 6 4

500 1000 Iterations Muon MuonClip SPECTRA Soft Musec

10

Validation loss

3

4

Validation loss

4

5

Validation loss

Validation loss

Best Tuned

5

Adam AdamW Muon MuonClip SPECTRA Soft Musec

6

Validation loss

Adam AdamW Muon MuonClip SPECTRA Soft Musec

6

η = 0.5

OpenWebText

Validation loss

FineWeb

0

500 1000 Iterations

Figure 6: Training loss over steps for NanoGPT-Small. Rows: learning rates; columns: datasets. Only learning rates where methods exhibit meaningfully different behavior are shown. NaN values within runs are omitted.

25

Preprint. Under review.

0

2000 Iterations

Validation loss

4

2000 Iterations

Muon MuonClip SPECTRA Soft Musec

6 5 4

2000 Iterations

4000 Muon MuonClip SPECTRA Soft Musec

Validation loss

5 4

0

2000 Iterations

4000 Muon MuonClip SPECTRA Soft Musec

6

Validation loss

0

5 4 3

3

0

2000 Iterations

4000

0

4000

12.5 10.0 7.5 5.0 2.5 0

Muon MuonClip SPECTRA Soft Musec

10 8 6 4 0

2000 Iterations

2000 Iterations

Validation loss

4 3

4000

3

6

Validation loss

0

4000 Muon MuonClip SPECTRA Soft Musec

Validation loss

η = 0.1

5 3

η = 0.2

4000 Muon MuonClip SPECTRA Soft Musec

6

3

5

2000 4000 Iterations

0

2000 Iterations

4000 Muon MuonClip SPECTRA Soft Musec

6

Validation loss

3

4

Adam AdamW Muon MuonClip SPECTRA Soft Musec

6

5 4 3

0

2000 Iterations

4000 Muon MuonClip SPECTRA Soft Musec

6

Validation loss

4

5

Validation loss

Validation loss

Best Tuned

5

C4

Adam AdamW Muon MuonClip SPECTRA Soft Musec

6

Validation loss

Adam AdamW Muon MuonClip SPECTRA Soft Musec

6

η = 0.5

OpenWebText

5 4 3

0

2000 Iterations

4000 Muon MuonClip SPECTRA Soft Musec

10

Validation loss

FineWeb

8 6 4 0

2000 Iterations

4000

Figure 7: Training loss over steps for NanoGPT-Medium. Rows: learning rates; columns: datasets. Only learning rates where methods exhibit meaningfully different behavior are shown. NaN values within runs are omitted.

26

Preprint. Under review.

F.2

A DDITIONAL T RAINING DYNAMICS

Figures 6–8 present training loss curves across all model configurations, datasets, and selected learning rates. At the best-tuned learning rate, all Muon-type optimizers converge to lower validation loss than Adam and AdamW. As the learning rate increases, the stability gap between methods widens, and this effect is amplified at larger model scales. For NanoGPT-Small, instability only manifests at the largest learning rate (η = 0.5), where Muon and MuonClip diverge on some datasets while Soft Musec and SPECTRA remain stable. For NanoGPT-Medium and NanoGPT-Wide, Muon and MuonClip exhibit increasingly severe instability starting from η = 0.1, with large loss spikes and sustained oscillations. Soft Musec and SPECTRA both maintain stable convergence across all configurations, with Soft Musec achieving comparable or slightly smoother training dynamics. F.3

E FFECTIVE R ANK A NALYSIS

To directly validate our theoretical motivation, we track the effective rank of the update matrices across QK projections, VO projections, and MLP weights during NanoGPT-Medium training on FineWeb at η = 0.5. Let us first recall the definition of the effective rank (Roy & Vetterli, 2007; Garrido et al., 2023) of a matrix A ∈ Rm×n with singular values σ1 ≥ σ2 ≥ · · · ≥ σr ≥ 0, where r = min(m, n). We define the singular value distribution as σi pi = Pr

j=1 σj

so that

i ∈ [r]

,

P

i pi = 1. We next compute the Shannon entropy of this distribution as

H(p1 , . . . , pr ) = −

r X

pi log(pi ).

i=1

The effective rank of the matrix A, denoted erank(A), is defined as erank(A) = exp(H(p1 , . . . , pr )). The complete results are shown in Figure 9. Muon exhibits near-full effective rank across all components, confirming that its polar step discards all spectral magnitude information. MuonClip does not meaningfully reduce the effective rank of the update, as it targets the weight matrices rather than the update itself. Both SPECTRA and Soft Musec achieve substantially lower effective rank than Muon, consistent with the effect of spectral clipping. Notably, as training progresses, Soft Musec maintains lower effective rank than SPECTRA across all three component types. This is consistent with the momentum feedback mechanism in Musec: by propagating the clipped momentum, spectral components are continuously regulated rather than allowed to accumulate in an unclipped buffer, as discussed in Section 4. F.4

C OMPUTATIONAL OVERHEAD

We compare the average wall-clock time per training step across all methods on FineWeb, with results shown in Figure 10. Across all three model configurations, Soft Musec introduces negligible computational overhead compared to Muon. For NanoGPT-Small, all methods are within 3% of each other (1021–1049 ms). For NanoGPT-Medium, all Muon-type methods achieve similar wallclock time (2550–2692 ms), with MuonClip being the slowest. For NanoGPT-Wide, all methods remain within 2% (15645–15977 ms). MuonClip is consistently the slowest method due to its additional weight clipping step. These results confirm that the spectral clipping operation in Soft Musec introduces minimal computational overhead relative to Muon, as both methods rely on Newton–Schulz iterations of similar complexity. F.5

A BLATION S TUDIES

In this subsection, we ablate two key hyperparameters of Soft Musec: the clipping threshold D and the number of Newton–Schulz iterations. 27

Preprint. Under review.

Validation loss

5 4 3 0

Muon MuonClip SPECTRA Soft Musec

6

Validation loss

Adam AdamW Muon MuonClip SPECTRA Soft Musec

6

5 4 3

2000 4000 6000 8000 Iterations

0

2000 4000 6000 8000 Iterations (b) η = 0.1

(a) Best Tuned

Validation loss

5 4 3 0

Muon MuonClip SPECTRA Soft Musec

10

Validation loss

Muon MuonClip SPECTRA Soft Musec

6

2000 4000 6000 8000 Iterations

8 6 4 0

2000 4000 6000 8000 Iterations

(c) η = 0.2

(d) η = 0.5

Figure 8: Training loss over steps for NanoGPT-Wide on FineWeb dataset. Only learning rates where methods exhibit meaningfully different behavior are shown. NaN values within runs are omitted.

Muon MuonClip SPECTRA Soft Musec

500 250 0

750

2000 4000 Iterations (a) QK projections

Muon MuonClip SPECTRA Soft Musec

500 250 0

0

1000 Effective rank

Effective rank

750

Effective rank

1000

1000

Muon MuonClip SPECTRA Soft Musec

750 500 250 00

0

2000 4000 Iterations (b) VO projections

2000 4000 Iterations (c) MLP

Figure 9: Effective rank of update matrices over training steps for NanoGPT-Medium on FineWeb.

3000 2500

800

2573.5 ms 2691.6 ms 2576.9 ms 2550.0 ms

2000

600

1000

200

15644.6 ms 15828.2 ms 15976.9 ms 15905.1 ms

10000

1500

400

15000

Average time (ms)

1020.9 ms 1048.5 ms 1045.8 ms 1035.0 ms

Average time (ms)

Average time (ms)

1000

500

5000

0 Muon MuonClip SPECTRA Soft Musec

0 Muon MuonClip SPECTRA Soft Musec

0 Muon MuonClip SPECTRASoft Musec

(a) NanoGPT-Small

(b) NanoGPT-Medium

(c) NanoGPT-Wide

Figure 10: Average wall clock time comparison of methods for training FineWeb.

28

Preprint. Under review.

5 4 3

0

2000 4000 Iterations

D=0.05 D=0.25 D=0.75 D=1.0

6 5 4 3

0

(a) η = 0.1

D=0.05 D=0.25 D=0.75 D=1.0

10.0

Validation loss

Validation loss

D=0.05 D=0.25 D=0.75 D=1.0

Validation loss

6

7.5 5.0

2000 4000 Iterations

0

(b) η = 0.2

2000 4000 Iterations

(c) η = 0.5

Figure 11: Effect of clipping threshold on Soft Musec training stability across learning rates. F.5.1

S ENSITIVITY TO C LIPPING T HRESHOLD

We study the sensitivity of Soft Musec to the clipping threshold D on NanoGPT-Medium trained on FineWeb across three learning rates (η ∈ {0.1, 0.2, 0.5}), with results shown in Figure 11. At moderate learning rates (η = 0.1 and η = 0.2), Soft Musec is robust to the choice of D: all values from D = 0.05 to D = 1.0 converge to comparable validation loss with similar training dynamics. At the largest learning rate (η = 0.5), the choice of D becomes more important. Smaller thresholds (D = 0.05 and D = 0.25) maintain smooth convergence, while larger thresholds (D = 0.75 and D = 1.0) exhibit loss spikes in the later stages of training. Notably, even with the largest clipping threshold (D = 1.0), Soft Musec still converges — in contrast to Muon and MuonClip, which diverge entirely at this learning rate (Figure 7). This demonstrates that spectral clipping provides a meaningful stability benefit regardless of the threshold, while a well-tuned D further improves training smoothness. N UMBER OF N EWTON -S CHULZ I TERATIONS

Validation loss

NS steps=3 NS steps=5 NS steps=8

5 4 3

NS steps=3 NS steps=5 NS steps=8

6 5 4

0

2000 4000 Iterations

(a) η = 0.1

3

NS steps=3 NS steps=5 NS steps=8

6

Validation loss

6

Validation loss

F.5.2

5 4

0

2000 4000 Iterations

(b) η = 0.2

3

0

2000 4000 Iterations

(c) η = 0.5

Figure 12: Effect of Newton–Schultz steps on Soft Musec training stability across learning rates. We study the sensitivity of Soft Musec to the number of Newton–Schulz iterations on NanoGPTMedium trained on FineWeb across three learning rates (η ∈ {0.1, 0.2, 0.5}), with the clipping threshold fixed at D = 0.05. Results are shown in Figure 12. Across all learning rates, Soft Musec is largely insensitive to the number of Newton–Schulz iterations. More iterations yield slightly better validation loss, consistent with a more accurate approximation of the spectral clipping operator. However, the differences are marginal, indicating that even a coarse approximation with three iterations provides most of the stability benefit. This robustness stems from cM c ⊤ + D2 I, whose eigenvalues are bounded below the well-conditioned nature of the matrix M 2 by D , enabling rapid convergence of the iteration. In practice, five iterations offer a good balance between approximation quality and computational cost.

29

Record · ID 673520 · SHA-256 c9cf5a03917f7157
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.