Tensorion: A Tensor-Aware Generalization of the Muon Optimizer
Vladimir Bogachev 1 Vladimir Aletov 1 2 Alexander Molozhavenko 1 Sergei Kudriashov 1 Maxim Rakhuba 1
arXiv:2606.25975v1 [cs.LG] 24 Jun 2026
Abstract
a particularly efficient method, achieving record-breaking performance on smaller models such as NanoGPT, and has recently been shown to scale to larger architectures (Liu et al., 2025; Team et al., 2026).
Common first-order optimizers, such as Adam, implicitly treat each parameter block as an unstructured vector, which disregards the multilinear weight structure present in many modern machine learning models. Recent work has shown that exploiting matrix structure can improve optimization dynamics. A notable example is Muon, which performs steepest descent under the spectral norm constraint. We take the next step and introduce Tensorion1 , a tensor-aware optimizer that extends Muon’s constrained optimization perspective from matrices to higher-order tensors. Tensorion is built around a linear minimization oracle (LMO) over a tensor norm ball. The norm is carefully chosen to balance two objectives: tightly bounding the tensor spectral norm, while still keeping the LMO tractable. This LMO becomes computable because it reduces to operations on adaptively selected unfolding matrices. Notably, when restricted to order-2 tensors (i.e., matrices), Tensorion recovers Muon exactly. Experiments on tensor-based computer vision problems suggest that Tensorion can offer improved convergence behavior and more stable gradient updates compared with Adam-based and existing tensor-aware baselines in the evaluated settings.
However, many modern machine learning models are inherently tensor-valued. Flattening tensors into matrices destroys the multilinear structure and mixes independent modes, which omits information crucial for optimization. Convolutional kernels, Multi-Head Attention (MHA) (Devlin et al., 2019), Mixture-of-Experts layers (MoE) (Oldfield et al., 2024), and other architectural components allow for a natural characterization as multilinear operators, i.e., higherorder tensors. In this work, we aim to extend the success of Muon to tensor-valued parameters. This setting naturally encompasses weights in neural networks with a convolution-based architecture and also arises in tensorization, where lowerorder objects (such as vectors) are reshaped into higherdimensional tensors (Novikov et al., 2015). To illustrate our approach and its connection to prior work, we consider the Linear Minimization Oracle (LMO) framework (Bernstein, 2025a). Given an ascent direction, for instance the gradient ∇L(W ) of the loss function L(·), one selects the search direction by solving max ⟨∇L (W ), X⟩.
X : ∥X∥≤1
(1)
The specific norm ∥ · ∥ chosen here essentially defines the geometry of the problem, and leads to completely different optimization trajectories (Pethick et al., 2025). If we take the Euclidean (Frobenius) norm, the solution recovers the standard gradient descent. This is unsurprising, as the Frobenius norm treats a matrix as a flat vector and ignores its inherent structure. In contrast, Muon employs the matrix spectral norm, which preserves the matrix structure and leads to an (approximate) orthogonalization of the gradient ∇L(W ) at each step. Other methods like signSGD (Bernstein et al., 2018) over ℓ∞ -norm balls also admit this type of interpretation.
1. Introduction The field of matrix optimization is emerging as a promising paradigm for the development of next-generation, efficient training algorithms for deep neural networks. In contrast to well-established methods like Adam (Kingma & Ba, 2014), which treat weights as a flat collection of numbers (parameter vectors), matrix-based approaches preserve and leverage the natural two-dimensional structure of the weight matrices. Among these, Muon (Jordan et al., 2024) has proven to be 1 HSE University 2 Basic Research of Artificial Intelligence Laboratory (BRAIn Lab). Correspondence to: Vladimir Bogachev <[email protected]>.
A natural step in extending this approach is to employ a norm specifically designed for tensors. However, unlike matrices, higher-order tensors exhibit fundamentally different properties, which significantly complicates such a general-
Preprint. June 25, 2026. 1 https://github.com/MTML-LAB/Tensorion
1
Tensorion: A Tensor-Aware Generalization of the Muon Optimizer
ization. In particular, one cannot straightforwardly use the well-known tensor generalization of the spectral norm.
Tensor methods have also been used to compress or structure neural network parameters, including tensor-structured filters (Rigamonti et al., 2013), Tensor Train layers (Novikov et al., 2015), stable tensor decompositions (Phan et al., 2020), and Tucker-style parameterizations (Molozhavenko & Rakhuba, 2025; Peshekhonov et al., 2024). Related tensor structures have been applied to multi-head attention (Gu et al., 2025; Li et al., 2026), mixture-of-experts models (XU et al., 2026), fixed-rank tensor optimization (Mo et al., 2025). These approaches typically modify the model parameterization, whereas Tensorion preserves the original parameters and uses tensor structure only in the optimizer. Finally, unlike (Zhang et al., 2026), which stacks layerwise gradients into auxiliary tensors, Tensorion operates directly on each layer’s tensor parameters.
Our contributions are as follows: 1. We show how to implicitly define a tensor norm that, on the one hand, serves as a tight upper bound to the spectral tensor norm (see Sec. 4.1), on the other hand, ensures a tractable solution to the LMO problem (see Sec. 4.2). In a nutshell, using our norm results in adaptively choosing unfolding matrices of a tensor depending on their nuclear matrix norms. We call the resulting optimizer “Tensorion”. 2. Based on theoretical results for randomly generated tensors (see Sec. 4.3) and an extensive evaluation of different unfolding strategies (see Sec. 5.1), we develop the best practical strategy for choosing the unfolding approach without requiring additional recomputation at each iteration.
3. Preliminaries 3.1. Norms, dual norms, and LMOs
3. For convolutional layers, we show (see Sec. 4.5) how the proposed method relates to the LMO problem for the actual Jacobian of the convolution operation (similarly to how W is a Jacobian of y = W x + b for fully-connected layers).
The LMO in Eq. (1) is naturally described through dual norms. For a norm ∥·∥ on a finite-dimensional inner-product space, its dual norm is
4. We evaluate the proposed optimizer on multiple CV classification datasets across CNN and transformerbased architectures (see Sec. 5) and observe consistent improvements over conventional SGD and Adambased optimizers.
Consequently, the maximizers of the linear oracle over the unit ball are exactly the subgradients of the dual norm:
∥M ∥† := max ⟨M, X⟩. ∥X∥≤1
(2)
Argmax⟨M, X⟩ = ∂∥M ∥† = ∥X∥≤1
Y
2. Related work
(3)
∥Y ∥ ≤ 1, ⟨M, Y ⟩ = ∥M ∥† ,
a standard consequence of convex duality (Rockafellar, 1997). The corresponding minimization oracle is obtained by changing the sign.
Tensorion is related to matrix-aware and tensor-aware optimization methods. Muon introduced matrix orthogonalization into optimization (Jordan et al., 2024), with later work improving its training recipe (Liu et al., 2025), connecting it to spectral constraints (Chen et al., 2025), and extending it through manifold constraints (Bernstein, 2025b), fixed-rank fine-tuning (Bogachev et al., 2026), distributed orthonormalized updates (Ahn et al., 2025), and general norm-based LMOs (Riabinin et al., 2025). A parallel line of structured optimizers uses matrix or tensor structure for preconditioning: K-FAC uses Kronecker-factored curvature (Martens & Grosse, 2015), with extensions to convolutional layers (Grosse & Martens, 2016) and transformers (Eschenhagen et al., 2023); Shampoo maintains modewise tensor preconditioners (Gupta et al., 2018), with further analysis in (Morwani et al., 2024); SOAP combines Shampoo eigenspaces with Adam-like updates (Vyas et al., 2024); and (Bernstein & Newhouse, 2024) give an operator-norm view. In contrast to these preconditioning methods, Tensorion does not maintain auxiliary matrices and instead acts directly through tensor unfoldings.
For Muon, the relevant primal norm is the matrix spectral norm, ∥M ∥2 = σ1 (M ), whose dual is the nuclear norm ∥M ∥†2 = ∥M ∥∗ =
r X
σi (M ),
(4)
i=1
where r = rank(M ) and σ1 (M ) ≥ · · · ≥ σr (M ) > 0 are the nonzero singular values. Let M = U ΣV ⊤ be a compact SVD. Then U V ⊤ is a maximizer since ⟨M, U V ⊤ ⟩ = Tr(Σ) = ∥M ∥∗ . Thus, U V ⊤ ∈ ∂∥M ∥∗ , while −U V ⊤ solves the corresponding minimization oracle. This dual-norm viewpoint is the basic mechanism behind the Muon update and will be the starting point for our tensor extension. For completeness, for a convex function f : V → R, the subdifferential at x is ∂f (x) := {g ∈ V | f (y) ≥ f (x) + ⟨g, y − x⟩, ∀y ∈ V } . 2
(5)
Tensorion: A Tensor-Aware Generalization of the Muon Optimizer
3.2. The tensor case
4.1. Relaxed tensor spectral norm
Let X, Y ∈ Rn1 ×···×nd be d-th order tensors. We use the Euclidean tensor inner product
The tensor spectral norm in Eq. (8) is the natural analogue of the matrix spectral norm and is therefore the most direct candidate for extending Muon beyond matrices. However, computing the tensor spectral norm, and hence solving the corresponding exact LMO, is NP-hard in general (Hillar & Lim, 2009). This makes the direct tensor-spectral formulation impractical for modern tensor-valued neural-network parameters.
⟨X, Y ⟩ :=
n1 X
···
i1 =1
nd X
Xi1 ,...,id Yi1 ,...,id ,
(6)
id =1
p and the associated Frobenius norm ∥X∥F := ⟨X, X⟩. The matrix spectral norm admits the variational form: ∥X∥2 =
max
∥u∥2 =∥v∥2 =1
|u⊤ Xv| =
max
∥u∥2 =∥v∥2 =1
A standard relaxation is obtained by upper-bounding the tensor spectral norm with the spectral norm of a matrix unfolding. Namely, for any τ ⊂ {1, . . . , d}, it is well known (Wang et al., 2017) that
|⟨X, uv ⊤ ⟩|.
(7) Then its multilinear analogue is a direct generalization of a matrix case: D E ∥X∥σ := sup X, u(1) ⊗ · · · ⊗ u(d) ,
∥X∥σ ≤ ∥X(τ ) ∥2 .
This observation motivates unfolding-based heuristics, including the empirical strategy used in (Jordan et al., 2024) for convolutional layers. However, selecting a single unfolding hard-codes a particular tensor geometry. Different layers, and even different tensor shapes within the same architecture, may favor different groupings of modes. Thus, the choice of unfolding becomes an additional design decision rather than a property of the optimization problem itself.
u(k) ∈Rnk , ∥u(k) ∥2 =1 k=1,...,d
(8) (1) (d) where u(1) ⊗ · · · ⊗ u(d) i ,...,i = ui1 · · · uid . For d = 2, 1 d this definition reduces to the usual matrix spectral norm. For d ≥ 3, however, the tensor spectral norm and its dual are substantially harder to compute and analyze (Friedland & Lim, 2018). This motivates the tractable relaxation developed in Section 4.
A natural attempt to remove this ambiguity is to aggregate several candidate unfoldings. Given a family of mode subsets T , for example T = {{1}, {1, 2}} or T = 2{1,2,...,d} , one may consider
We will frequently use tensor unfoldings. For τ ⊂ {1, . . . , d}, let τ c denote its complement and define Nτ :=
Y
nk ,
Nτ c :=
Y
nk ,
(9)
min ∥X(τ ) ∥2
k∈τ /
k∈τ
τ ∈T
with the convention that an empty product equals one. The τ -unfolding of X, denoted by X(τ ) , is the matrix obtained by grouping the modes in τ into rows and the remaining modes into columns: X(τ ) jτ ,j c := Xj1 ,...,jd , τ
(12)
or
max ∥X(τ ) ∥2 . τ ∈T
(13)
The first construction is generally not a norm: the minimum of norms need not satisfy the triangle inequality. The second construction is a norm, but it is often too conservative. In particular, it can collapse to the Frobenius norm in the presence of small or singleton modes, which are common in convolutional kernels.
X(τ ) ∈ RNτ ×Nτ c , (10)
To see this, consider the extreme case X ∈ R1×n2 ×···×nd and suppose that {1} ∈ T . Then
where jτ and jτ c are the corresponding lexicographically ordered multi-indices. The standard mode-k matricization is the special case
max ∥X(τ ) ∥2 = ∥X({1}) ∥2 = ∥ vec(X)⊤ ∥2 = ∥X∥F , τ ∈T
X(k) := X({k}) .
(14) where vec(·) denotes column-major vectorization. Hence, the relaxation becomes the Frobenius norm, which may substantially overestimate the tensor spectral norm and lose sensitivity to the multilinear structure of X. Moreover, this behavior prevents the relaxation from recovering the matrix Muon update on tensors of size 1 × m × n, which should effectively behave as matrices.
(11)
4. Tensorion This section introduces Tensorion, our tensor-valued extension of Muon. The central difficulty is that the natural LMO over the tensor spectral-norm ball is computationally intractable. Tensorion addresses this obstacle by replacing the tensor spectral norm with a duality-based relaxation that is both tighter than any unfolding spectral norm and computable.
Tensorion avoids this pathology by moving the relaxation to the dual side. Instead of aggregating primal unfolding spectral norms, we construct a lower bound on the dual tensor 3
Tensorion: A Tensor-Aware Generalization of the Muon Optimizer
spectral norm using unfolding nuclear norms. Specifically, we define: ∥X∥†Σ := max ∥X(τ ) ∥∗ . (15)
Proof is in Supplementary Materials Section A, Proposition A.4. Importantly, Proposition 4.2 gives us a practical way to solve the LMO in Eq. (19) by finding an unfolding with the largest nuclear norm. In particular, for any τ ∈ I, one admissible choice is obtained by folding U V ⊤ back along the unfolding τ , where U and V are the left and right singular vectors from the singular value decomposition M(τ ) = U ΣV ⊤ .
τ ∈T
The relaxed tensor spectral norm is then defined by duality: † ∥X∥Σ := ∥X∥†Σ .
(16)
This dual construction is the key relaxation underlying Tensorion. It preserves the desirable upper-bound property required for a spectral-norm relaxation, while avoiding the overly conservative behavior of the maximum over primal unfolding norms. Supplementary Material D provides an empirical analysis of relaxation tightness along optimization trajectories.
The best unfolding can be different for different layers and even change during the optimization process. We found that searching for optimal unfoldings during every iteration of the optimization process is excessive (see numerical experiments). Even though it is possible to do it efficiently, using, e.g., Newton-Schulz iteration run in parallel, there are simpler ways that lead to similar optimization trajectory in practice.
The following proposition formalizes the relationship between the proposed norm, the tensor spectral norm, and unfolding-based spectral relaxations. In particular, Tensorion always upper-bounds the tensor spectral norm, yet it is no looser than the best spectral norm among the selected unfoldings.
4.3. Motivation for the choice of unfolding shape Observe that for any unfolding M(τ ) ∈ Rmτ ×nτ
Proposition 4.1. Let X ∈ Rn1 ×···×nd be any tensor. Then ∥X∥†Σ ≤ ∥X∥†σ ,
∥M ∥2F = ∥M(τ ) ∥2F =
rτ 2 X (τ ) σj , rτ = rank (M(τ ) ), j=1
(17)
(22) = i∈τ ni , nτ = i∈τ c ni . Since unfolding only permutes tensor entries, the Frobenius norm of the unfolding is independent of the choice of τ . Hence, for a fixed Frobenius norm, maximizing the nuclear norm reduces to maximizing the sum of singular values subject to a fixed sum of their squares. By the Cauchy–Schwarz inequality, the maximum is attained when the norm budget is distributed uniformly across all available singular values. Consequently, the maximal nuclear norm grows with the number of potentially nonzero singular values, motivating the heuristic of choosing an unfolding with the largest possible number of nonzero singular values.
(τ ) where {σj } denote Q Q the singular values of M(τ ) , mτ
and consequently ∥X∥σ ≤ ∥X∥Σ ≤ min ∥X(τ ) ∥2 . τ ∈T
(18)
Proof is in Supplementary Materials, Section A, Proposition A.3. 4.2. Relaxed LMO solution Similar in spirit to the Muon optimizer, our method solves an LMO over the ∥ · ∥Σ at each iteration. Given a tensor M (e.g., the gradient or a momentum estimate at some iteration), we compute X opt ∈ Arg max ⟨M, X⟩.
The following formal statement further supports the heuristic of choosing an unfolding that maximizes the effective number of nonzero singular values.
(19)
X : ∥X∥Σ ≤1
Proposition 4.3 (High-Probability Nuclear Norm Bound). Let M ∈ Rn1 ×···×nd be a random tensor with entries i.i.d. Mi1 ,...,id ∼ P, where P is a zero-mean sub-exponential distribution. Consider any unfolding M(τ ) ∈ Rmτ ×nτ with Qd mτ ≤ nτ such that mτ nτ = i=1 ni . Then, there exists a constant C > 0 (depending on the sub-exponential norm of P) such that with probability at least 1 − δ: √ M(τ ) ∗ ≤ C · mτ nτ + log(1/δ). (23)
Proposition 4.2 characterizes the solution set of Eq. (19). Proposition 4.2. The solution set of the LMO in Eq. (19) is the subdifferential of the Tensorion dual norm at M : X opt ∈ ∂∥M ∥†Σ = conv ∂∥M(τ ) ∥∗ , (20) τ ∈I
where conv (·) denotes a convex hull and n o I = τ ∥M(τ ) ∥∗ = ∥M ∥†Σ ⊆ 2{1,...,d}
(21)
Moreover, for a fixed total dimension mτ nτ , the value of this upper bound is maximized when the unfolding is square, √ i.e., mτ = nτ = mτ nτ .
is a nonempty set of unfolding indices attaining the maximal nuclear norm. 4
Tensorion: A Tensor-Aware Generalization of the Muon Optimizer
Algorithm 1 Tensorion
the output tensor. The convolution operation can be viewed as a linear mapping
n1 ×···×nd
Input: Weight tensor W ∈ R , step sizes {ηk }, momentum coefficient β. Output: Weight tensor W ′ ∈ Rn1 ×···×nd . M0 = 0; Q Q τ := arg max Q τ minQ i∈τ ni , i̸∈τ ni ; mτ , nτ := i∈τ ni , i̸∈τ ni ; for k := 1, . . . do Compute gradient Gk ∈ Rn1 ×···×nd ; Mk := Gk + βMk−1 ; Xk := fold (NS p(unfold(Mk , τ )) , τ ); η̂k := ηk · 0.2 max(mτ , nτ ); Wk+1 := (1 − γηk )Wk − η̂k Xk ; end for return Wk+1 .
vec(Y ) = TK vec(X),
Therefore, unfoldings whose row and column dimensions are more balanced, in expectation potentially yield larger nuclear norms. Motivated by this observation, we consider the following as our primary strategy to select the unfolding index prior to the optimization loop: τ
:= arg max (min {mτ , nτ }) ,
2
(25)
where vec(·) denotes the vectorization operator and TK is a matrix of the layer determined by the kernel K. Note that TK is a large matrix with additional structure, i.e., it is a sparse block matrix where each block represents a two-level Toeplitz structure (Grishina et al., 2025). It is used solely for theoretical analysis purposes and is never formed explicitly. The spectral properties of this convolution operator have been shown to affect generalization and training stability (Singla & Feizi, 2019; Agarwal et al., 2019). According to (Grishina et al., 2025), the spectral norm of the linear map TK is bounded by the spectral tensor norm as stated in Theorem 4.4. Theorem 4.4 (Grishina et al. (2025)). Let TK ∈ 2 2 Rcout n ×cin n be the Jacobian matrix of a convolutional layer with stride s = 1 and zero padding with parameter p ≥ 0, or circular padding. Let K ∈ Rcout ×cin ×h×w be the convolution kernel. Then √ (26) ∥K∥σ ≤ ∥TK ∥2 ≤ hw ∥K∥σ .
Proof is in Supplementary Materials Section B, Proposition B.1.
opt
2
TK ∈ Rcout n ×cin n ,
As a corollary to this theorem, the same inequality holds for gradients: √ ∥∇K L(K)∥σ ≤ ∥∇TK L(TK )∥2 ≤ hw∥∇K L(K)∥σ , (27)
(24)
τ ∈T
then we restrict the orthogonalization step to the τ opt unfolding.
where L denotes the loss function, ∇K L(K) ∈ Rcout ×cin ×h×w is the gradient with respect to the convolu2 2 tion kernel, and ∇TK L(TK ) = T∇K L(K) ∈ Rcout n ×cin n is the corresponding structured gradient with respect to the convolution matrix.
4.4. Tensorion algorithm The proposed theoretical framework leads to our main contribution, the Tensorion Optimizer. Algorithm 1 presents the version of the method used in the experimental section.
Applying the Muon optimizer directly to the matrix T∇K L(K) is not practical, as it breaks the sparse multilevel structure of the matrix. However, if we try to solve the LMO problem with Toeplitz-like structure constraints, the inequality (27) becomes useful. Indeed, it allows us to relax the LMO problem to the one with ∥∇K L(K)∥σ instead of ∥∇TK L(TK )∥2 . Finally, using Eq. (18), we further relax the LMO and arrive at the Tensorion approach for the convolution kernel tensor. To conclude, the Tensorion approach is a relaxation of the LMO problem that is applied directly to the matrix of the actual linear convolution operator.
The first three lines contain the momentum initialization and the computation of the τ opt -unfolding from Eq. (24), which is used to obtain a faster LMO solution. Within the main optimization loop, Tensorion follows the Muon update scheme, with the key modification appearing in line 7. First, the function unfold(·, τ ) maps a d-dimensional tensor to its τ -unfolding. Then, Newton–Schulz orthogonalization, denoted NS(·), is applied. Finally, fold(·, τ ) maps the result back to the original tensor shape. Changes with respect to the Muon implementation of (Liu et al., 2025) are highlighted in blue.
5. Experiments 4.5. Tensorion for Convolution Layers
To validate the derived heuristic, we first ablate the choice of tensor unfolding index set τ and study its effect on the optimization dynamics (Section 5.1). We then evaluate whether the resulting improvements translate into higher final accuracy on standard CNN (Section 5.2) and Vision Transformer
In this subsection, let us discuss what it means to apply Tensorion to a convolutional kernel tensor. Let X ∈ Rcin ×n×n be the input to a convolutional layer and let K ∈ Rcout ×cin ×h×w be its kernel. Let Y ∈ Rcout ×n×n denote 5
Tensorion: A Tensor-Aware Generalization of the Muon Optimizer
Test accuracy
(ViT) classification benchmarks (Section 5.3). Additional details, including hyperparameters (Section F) and further experiments on the Laplacian eigenproblem (Section C.3) and CIFAR datasets (Section C.1), are provided in the Supplementary Material.
0.8 0.5
0.90
0.2
0.85
0
5.1. Ablation on an Unfolding Set Test loss
We isolate the impact of the unfolding index set τ and evaluate our heuristic choice τ opt from Eq. (24), referred to as the offline strategy. As a brute-force reference, we also compare against an online strategy that evaluates all admissible unfoldings at each iteration and selects the one maximizing Eq. (21). While potentially more adaptive, this online strategy is substantially more expensive in practice.
10
5
10 15 Steps
20
5 × 10
−1
4 × 10
−1
25
τ = {1} τ = {2} τ = {3} τ = {4} τ = {1, 2} τ = {1, 3} τ = {1, 4} τ opt τ per iter SGD-M AdamW
0
0
5
10
15
20
25
15 Steps
20
25
Steps Train loss
For this ablation, we evaluate all seven nontrivial unfoldings of a 4D convolutional kernel up to the complement symmetry τ ∼ τ c . Static baselines use one unfolding across all layers, while the offline strategy selects τ opt per layer. To isolate unfolding selection from approximation effects, we use exact SVDs instead of NS at line 7 in Algorithm 1 and train ResNet-18 on CIFAR-10 for 25 epochs with identical settings and independently tuned constant learning rates; see Supplementary Materials, Section E.1. Figure 1 and Table 1 show that the choice of unfolding has a substantial effect on performance. The tested unfoldings separate into two accuracy regimes: the lower-performing group consists of τ = {3}, τ = {4}, and τ = {1, 2}, whereas the remaining choices form a higher-performing group. The differences within each group are not statistically significant under this protocol, and therefore we do not attempt to rank individual unfoldings inside the high-accuracy group.
10
0
0
5
10
Figure 1. Ablation on the unfolding set τ for ResNet-18 on CIFAR10 (25 epochs; constant LR, no scheduler) using exact SVD for orthogonalization. Curves labeled “τ = {·}” use a fixed unfolding index set τ ⊂ {1, 2, 3, 4}.“τ opt ” selects τ once per layer via the offline heuristic in Eq. (24). “τ per iteration” performs online selection by explicitly maximizing ∥X(τ ) ∥∗ over τ at each iteration as in Eq. (21). SGD-M: SGD+Momentum.
constant-learning-rate protocol. To evaluate the applicability of Tensorion in a larger-scale image classification setting, we next provide a set of experiments on various computer vision problems.
The online strategy, which explicitly maximizes Eq. (21) at every iteration, achieves the highest mean accuracy, as expected from its additional adaptivity. However, this comes at a substantially higher computational cost, and the observed gap to the other high-accuracy methods is not statistically resolved in this ablation. In contrast, the offline heuristic τ opt is inexpensive and still falls into the high-accuracy regime. This suggests that the heuristic succeeds at avoiding unfavorable unfoldings and recovers a competitive fixed-per-layer choice without requiring brute-force online selection. Moreover, the elapsed time for Tensorion with offline τ -selection strategy is comparable to the Muon optimizer.
Table 1. Ablation on the unfolding set τ for ResNet-18 on CIFAR10 after 25 epochs. Reported values are mean ± standard deviation over 10 runs; SGD-M: SGD+Momentum. Method
Test acc. ↑
Test loss ↓
τ = {1} 90.67± 0.32 0.424± 0.026 τ = {2} 90.62± 0.44 0.404± 0.024 τ = {3} 87.40± 1.00 0.485± 0.051 τ = {4} 88.05± 0.69 0.423± 0.028 τ = {1, 2} 88.83± 0.39 0.430± 0.022 τ = {1, 3} 90.30± 0.51 0.426± 0.031 τ = {2, 4} 90.77± 0.45 0.432± 0.038 τ opt 90.44± 0.36 0.441± 0.024 online τ 91.10± 0.39 0.416± 0.026 SGD-M 86.10± 1.82 0.527± 0.095 AdamW 86.04± 1.05 0.507± 0.041
We include the fixed unfolding τ = {1} as the Muon comparison, since this choice recovers the original Muon update in our framework. The main conclusion of the ablation is not a fine-grained ordering among the competitive methods, but rather that Tensorion is sensitive to the unfolding choice and that the proposed heuristic provides a practical way to avoid poor unfoldings without per-iteration search. SGD-M and AdamW (Loshchilov & Hutter, 2017) are included as baselines and achieve lower test accuracy under the same 6
Tensorion: A Tensor-Aware Generalization of the Muon Optimizer
5.2. Tiny ImageNet with ResNet50
As shown in Figure 2, Tensorion consistently outperforms both baselines in test accuracy throughout training. The advantage is even more pronounced in terms of test loss: AdamW exhibits clear overfitting, with its test loss increasing after roughly 50 steps, whereas Tensorion achieves the lowest and monotonically decreasing loss. These findings are consistent with our ResNet-18/-34 experiments on CIFAR-10/-100 reported in Appendix C.1 and summarized alongside all main results in Table 2.
We train a ResNet-50 model on Tiny ImageNet (Yin et al., 2023) for 200 epochs, comparing Tensorion against SGD and AdamW baselines. Tensorion is applied only to tensorvalued parameters of order d ≥ 3 (i.e., convolutional kernels), while all remaining parameters (d ≤ 2) are optimized with SGD. We report mean ± std over multiple seeds.
0.6
0.8
0.2
Accuracy
Accuracy
0.9
0.4
Tensorion SGD AdamW
0
25
50
75
100
125
150
175
0.7 0.6 0.5 Tensorion-SGD Tensorion-AdamW
200 0.4
Epochs
100
200
300
500
600
700
Steps
(a) Test accuracy vs Epoch
(a) Swin Base
6 5
0.9
4
0.8
Accuracy
Test Loss
400
SGD AdamW
3 2 0
25
50
75
100
125
150
175
0.7 0.6 0.5
200
Epochs
Tensorion-SGD Tensorion-AdamW
0.4 100
RN-18
C100
SGD-M 0.89±0.13 AdamW 1.33±0.1 RN-34 T-SGD 0.98±0.09 T-AdamW 4.89±0.11
Tiny-IN RN-50
SGD-M AdamW T-SGD
600
700
5.3. Tensorion consistently improves performance of Vision Transformers. Although Vision Transformers mostly consist of common matrix layers, the convolutional patch-embedding module is crucial for optimization stability and model generalization (Xiao et al., 2021). We aim to show that applying Tensorion to this module alone is sufficient to improve performance across multiple scales and architectures. Specifically, we initialize ViT (Dosovitskiy et al., 2020) (Small, Base, Medium) and Swin Transformer (Liu et al., 2021) (Tiny, Small, Base) models from pretrained checkpoints in the timm library, fine-tune them on the Imagenette dataset (Howard, 2019) for 5 epochs with a fixed learningrate scheduler and weight decay across optimizers, and evaluate performance at the end of each epoch. Tensorion is
Acc. ↑
SGD-M 0.25±0.02 94.1±0.13 AdamW 0.44±0.02 93.7±0.3 T-SGD 0.21±0.01 95.3±0.1 T-AdamW 0.59±0.08 94.34±0.3
C10
500
Figure 3. Imagenette results for Swin and ViT models trained for 5 epochs. Tensorion is used only for tensor-valued blocks with order d ≥ 3 (e.g., convolutional kernels), while all remaining parameters (d ≤ 2) are optimized with either SGD or AdamW.
Table 2. Main results. Final test loss and test accuracy. C10: CIFAR-10; C100: CIFAR-100; Tiny-IN: Tiny ImageNet; RN: ResNet; SGD-M: SGD+Momentum; T-SGD: Tensorion+SGD; T-AdamW: Tensorion+AdamW. Loss ↓
400
(b) ViT Base
Figure 2. Tiny ImageNet results for ResNet-50 trained for 200 epochs. We report mean ± std over multiple seeds for test accuracy and test loss Tensorion is used only for tensor-valued blocks with order d ≥ 3 (e.g., convolutional kernels), while all remaining parameters (d ≤ 2) are optimized with a first-order method (SGD).
Opt.
300
Steps
(b) Test loss vs Epoch
Dataset Model
200
SGD AdamW
78.2±0.2 74.2±0.1 79.1±0.1 73.8±0.2
1.79±0.06 61.01±0.02 1.33±0.05 57.4±0.1 1.65±0.06 62.7±0.1
7
Tensorion: A Tensor-Aware Generalization of the Muon Optimizer
applied only to the convolutional patch-embedding layer. All remaining parameters are optimized with AdamW or SGD, respectively. Optimal learning rates are tuned via grid search for each setup. We report the best performance across algorithms in Fig. 3. Detailed results are reported in Section C.2.
Bernstein, J. Modular manifolds. Thinking Machines Lab: Connectionism, 2025b. doi: 10.64434/tml.20250926. https://thinkingmachines.ai/blog/modular-manifolds/.
6. Conclusion
Bernstein, J. and Newhouse, L. Old optimizer, new norm: An anthology. arXiv preprint arXiv:2409.20325, 2024.
https://jeremybernste.in/writing/ deriving-muon.
We propose Tensorion, a novel optimizer that generalizes Muon to tensors via a norm that lower bounds the spectral norm of any tensor unfolding, leading to a practical relaxation of the tensor spectral-norm LMO problem. A heuristic offline method for choosing the unfolding achieves performance comparable to the online optimal choice. Experiments across multiple CV tasks and architectures show consistent improvements over conventional optimizers.
Bernstein, J., Wang, Y.-X., Azizzadenesheli, K., and Anandkumar, A. signsgd: Compressed optimisation for nonconvex problems. In International conference on machine learning, pp. 560–569. PMLR, 2018. Bogachev, V., Aletov, V., Molozhavenko, A., Bobkov, D., Soboleva, V., Alanov, A., and Rakhuba, M. LoRA meets riemannion: Muon optimizer for parametrizationindependent low-rank adapters. In The Fourteenth International Conference on Learning Representations, 2026. URL https://openreview.net/forum? id=WtbXgc9GVA.
Limitations Our study is limited primarily to computer vision models. We do not evaluate video models or architectures operating on sequential visual domains. We also do not assess whether similar behavior appears for convolutional or any other tensor-valued layers in other domains beyond vision. Finally, our experiments are limited in scale: while the observed trends are encouraging, we cannot determine from the present evidence whether the method scales as reliably as Muon or other established optimizers in substantially larger training regimes.
Chen, L., Li, J., and Liu, Q. Muon optimizes under spectral norm constraints. arXiv preprint arXiv:2506.15054, 2025. Contributors, M. Openmmlab’s pre-training toolbox and benchmark. https://github.com/ open-mmlab/mmpretrain, 2023. Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers), pp. 4171–4186, 2019.
Acknowledgements The work was supported by the grant for research centers in the field of AI provided by the Ministry of Economic Development of the Russian Federation in accordance with the agreement 000000C313925P4E0002 and the agreement with HSE University № 139-15-2025-009. This research was supported in part through computational resources of HPC facilities at HSE University (Kostenetskiy et al., 2021).
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. Eschenhagen, R., Immer, A., Turner, R., Schneider, F., and Hennig, P. Kronecker-factored approximate curvature for modern neural network architectures. Advances in Neural Information Processing Systems, 36:33624–33655, 2023.
References Agarwal, C., Nguyen, A., and Schonfeld, D. Improving robustness to adversarial examples by encouraging discriminative features. In 2019 IEEE international conference on image processing (ICIP), pp. 3801–3805. IEEE, 2019. doi: 10.1109/ICIP.2019.8803601.
Friedland, S. and Lim, L.-H. Nuclear norm of higher-order tensors. Mathematics of Computation, 87(311):1255– 1281, 2018. doi: 10.1090/mcom/3239.
Ahn, K., Xu, B., Abreu, N., Fan, Y., Magakyan, G., Sharma, P., Zhan, Z., and Langford, J. Dion: Distributed orthonormalized updates. arXiv preprint arXiv:2504.05295, 2025. Bernstein,
J.
Deriving muon,
2025a.
Girsanov, I. V. Lectures on mathematical theory of extremum problems. Springer Science & Business Media, 2012. Grishina, E., Gorbunov, M., and Rakhuba, M. Tight and efficient upper bound on spectral norm of convolutional
URL 8
Tensorion: A Tensor-Aware Generalization of the Muon Optimizer
layers. In Leonardis, A., Ricci, E., Roth, S., Russakovsky, O., Sattler, T., and Varol, G. (eds.), Computer Vision – ECCV 2024, pp. 19–34, Cham, 2025. Springer Nature Switzerland. ISBN 978-3-031-73024-5. doi: 10.1007/ 978-3-031-73024-5 2.
Lim, L.-H. and Comon, P. Blind multilinear identification. IEEE Transactions on Information Theory, 60(2):1260– 1280, 2014. doi: 10.1109/TIT.2013.2291876. Liu, J., Su, J., Yao, X., Jiang, Z., Lai, G., Du, Y., Qin, Y., Xu, W., Lu, E., Yan, J., et al. Muon is scalable for llm training. arXiv preprint arXiv:2502.16982, 2025.
Grosse, R. and Martens, J. A kronecker-factored approximate fisher matrix for convolution layers. In International Conference on Machine Learning, pp. 573–582. PMLR, 2016.
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 10012–10022, 2021.
Gu, Y., Zhou, W., Iacovides, G., and Mandic, D. Tensorllm: Tensorising multi-head attention for enhanced reasoning and compression in llms. In 2025 International Joint Conference on Neural Networks (IJCNN), pp. 1–8. IEEE, 2025.
Loshchilov, I. and Hutter, F. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017. Martens, J. and Grosse, R. Optimizing neural networks with kronecker-factored approximate curvature. In International conference on machine learning, pp. 2408–2417. PMLR, 2015.
Gupta, V., Koren, T., and Singer, Y. Shampoo: Preconditioned stochastic tensor optimization. In International Conference on Machine Learning, pp. 1842–1850. PMLR, 2018.
Mo, Z., Huang, L.-K., and Pan, S. J. Parameter and memory efficient pretraining via low-rank riemannian optimization. In The Thirteenth International Conference on Learning Representations, 2025.
He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016.
Molozhavenko, A. and Rakhuba, M. V. Optimization on the extended tensor-train manifold with shared factors. Computational and Applied Mathematics, 45, 2025. URL https://api.semanticscholar. org/CorpusID:280950046.
Hillar, C. J. and Lim, L.-H. Most tensor problems are np-hard. ArXiv, abs/0911.1393, 2009. doi: 10.1145/ 2512329. Howard, J. Imagenette: A smaller subset of 10 easily classified classes from imagenet, March 2019. URL https://github.com/fastai/imagenette.
Morwani, D., Shapira, I., Vyas, N., Malach, E., Kakade, S., and Janson, L. A new perspective on shampoo’s preconditioner. arXiv preprint arXiv:2406.17748, 2024.
Jordan, K., Jin, Y., Boza, V., You, J., Cesista, F., Newhouse, L., and Bernstein, J. Muon: An optimizer for hidden layers in neural networks, 2024. URL https: //kellerjordan.github.io/posts/muon/.
Novikov, A., Podoprikhin, D., Osokin, A., and Vetrov, D. P. Tensorizing neural networks. Advances in neural information processing systems, 28, 2015. doi: 10.5555/2969239.2969289.
Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. CoRR, abs/1412.6980, 2014. URL https: //doi.org/10.48550/arXiv.1412.6980.
Oldfield, J., Georgopoulos, M., Chrysos, G. G., Tzelepis, C., Panagakis, Y., Nicolaou, M. A., Deng, J., and Patras, I. Multilinear mixture of experts: scalable expert specialization through factorization. In Proceedings of the 38th International Conference on Neural Information Processing Systems, NIPS ’24, Red Hook, NY, USA, 2024. Curran Associates Inc. ISBN 9798331314385.
Kostenetskiy, P., Chulkevich, R., and Kozyrev, V. Hpc resources of the higher school of economics. In Journal of Physics: Conference Series, volume 1740, pp. 012050. IOP Publishing, 2021.
Peshekhonov, I., Arzhantsev, A., and Rakhuba, M. Training a tucker model with shared factors: a riemannian optimization approach. In International Conference on Artificial Intelligence and Statistics, pp. 3304–3312. PMLR, 2024.
Krizhevsky, A., Hinton, G., et al. Learning multiple layers of features from tiny images. 2009. Li, Y., Guo, Z., Yin, M., and Li, B. Lestd: Llm compression via learning-based sparse tensor decomposition. In The Fourteenth International Conference on Learning Representations, 2026.
Pethick, T., Xie, W., Antonakopoulos, K., Zhu, Z., SilvetiFalls, A. J., and Cevher, V. Training deep learning models 9
Tensorion: A Tensor-Aware Generalization of the Muon Optimizer
with norm-constrained lmos. International Conference on Machine Learning, 2025. doi: 10.48550/arXiv.2502. 07529.
Y., Zhang, Y., Zhang, Y., Zhang, Y., Zhang, Z., Zhao, H., Zhao, Y., Zhao, Z., Zheng, H., Zheng, S., Zhong, L., Zhou, J., Zhou, X., Zhou, Z., Zhu, J., Zhu, Z., Zhuang, W., and Zu, X. Kimi k2: Open agentic intelligence, 2026. URL https://arxiv.org/abs/2507.20534.
Phan, A.-H., Sobolev, K., Sozykin, K., Ermilov, D., Gusak, J., Tichavskỳ, P., Glukhov, V., Oseledets, I., and Cichocki, A. Stable low-rank tensor decomposition for compression of convolutional neural network. In European Conference on Computer Vision, pp. 522–539. Springer, 2020.
Vyas, N., Morwani, D., Zhao, R., Kwun, M., Shapira, I., Brandfonbrener, D., Janson, L., and Kakade, S. Soap: Improving and stabilizing shampoo using adam. arXiv preprint arXiv:2409.11321, 2024.
Riabinin, A., Shulgin, E., Gruntkowska, K., and Richtárik, P. Gluon: Making muon & scion great again!(bridging theory and practice of lmo-based optimizers for llms). arXiv preprint arXiv:2505.13416, 2025.
Wang, M., Duc, K. D., Fischer, J., and Song, Y. S. Operator norm inequalities between tensor unfoldings on the partition lattice. Linear algebra and its applications, 520: 44–66, 2017. doi: 10.1016/j.laa.2017.01.017.
Rigamonti, R., Sironi, A., Lepetit, V., and Fua, P. Learning separable filters. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2754– 2761, 2013.
Xiao, T., Singh, M., Mintun, E., Darrell, T., Dollár, P., and Girshick, R. Early convolutions help transformers see better. In Proceedings of the 35th International Conference on Neural Information Processing Systems, NIPS ’21, Red Hook, NY, USA, 2021. Curran Associates Inc. ISBN 9781713845393.
Rockafellar, R. T. Convex analysis, volume 28. Princeton university press, 1997. Singla, S. and Feizi, S. Fantastic four: Differentiable bounds on singular values of convolution layers. arXiv preprint arXiv:1911.10258, 2019.
XU, Y., WANG, Y., Peng, X., Zang, H., Minghao, C., Xia, P., and Wen, Z. TD-moe: Tensor decomposition for moe models. In The Fourteenth International Conference on Learning Representations, 2026. URL https:// openreview.net/forum?id=D9cnZNZfxX.
Team, K., Bai, Y., Bao, Y., Charles, Y., Chen, C., Chen, G., Chen, H., Chen, H., Chen, J., Chen, N., Chen, R., Chen, Y., Chen, Y., Chen, Y., Chen, Z., Cui, J., Ding, H., Dong, M., Du, A., Du, C., Du, D., Du, Y., Fan, Y., Feng, Y., Fu, K., Gao, B., Gao, C., Gao, H., Gao, P., Gao, T., Ge, Y., Geng, S., Gu, Q., Gu, X., Guan, L., Guo, H., Guo, J., Hao, X., He, T., He, W., He, W., He, Y., Hong, C., Hu, H., Hu, Y., Hu, Z., Huang, W., Huang, Z., Huang, Z., Jiang, T., Jiang, Z., Jin, X., Kang, Y., Lai, G., Li, C., Li, F., Li, H., Li, M., Li, W., Li, Y., Li, Y., Li, Y., Li, Z., Li, Z., Lin, H., Lin, X., Lin, Z., Liu, C., Liu, C., Liu, H., Liu, J., Liu, J., Liu, L., Liu, S., Liu, T. Y., Liu, T., Liu, W., Liu, Y., Liu, Y., Liu, Y., Liu, Y., Liu, Z., Lu, E., Lu, H., Lu, L., Luo, Y., Ma, S., Ma, X., Ma, Y., Mao, S., Mei, J., Men, X., Miao, Y., Pan, S., Peng, Y., Qin, R., Qin, Z., Qu, B., Shang, Z., Shi, L., Shi, S., Song, F., Su, J., Su, Z., Sui, L., Sun, X., Sung, F., Tai, Y., Tang, H., Tao, J., Teng, Q., Tian, C., Wang, C., Wang, D., Wang, F., Wang, H., Wang, H., Wang, J., Wang, J., Wang, J., Wang, S., Wang, S., Wang, S., Wang, X., Wang, Y., Wang, Y., Wang, Y., Wang, Y., Wang, Y., Wang, Z., Wang, Z., Wang, Z., Wang, Z., Wei, C., Wei, Q., Wu, H., Wu, W., Wu, X., Wu, Y., Xiao, C., Xie, J., Xie, X., Xiong, W., Xu, B., Xu, J., Xu, L. H., Xu, L., Xu, S., Xu, W., Xu, X., Xu, Y., Xu, Z., Xu, J., Xu, J., Yan, J., Yan, Y., Yang, H., Yang, X., Yang, Y., Yang, Y., Yang, Z., Yang, Z., Yang, Z., Yao, H., Yao, X., Ye, W., Ye, Z., Yin, B., Yu, L., Yuan, E., Yuan, H., Yuan, M., Yuan, S., Zhan, H., Zhang, D., Zhang, H., Zhang, W., Zhang, X., Zhang, Y., Zhang, Y., Zhang, Y., Zhang, Y., Zhang,
Yin, Z., Xing, E., and Shen, Z. Squeeze, recover and relabel: Dataset condensation at imagenet scale from a new perspective. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2023. Zhang, R., Zhao, Y., Liu, Z., Wang, Z., Li, D., Su, Y., Liu, S., and Zhang, Z. Teon: Tensorized orthonormalization beyond layer-wise muon for large language model pretraining, 2026. URL https://arxiv.org/abs/ 2601.23261. Zhao, Y. Concentration inequalities for sub-weibull random tensors, 2026. URL https://arxiv.org/abs/ 2509.03439.
10
Tensorion: A Tensor-Aware Generalization of the Muon Optimizer
A. Primal norm Theorem A.1 (Primal relaxed spectral norm). Let X ∈ Rn1 ×...×nd then X Ξτ(τ ) ∥X∥Σ = min τ (Ξ )τ ∈T P Ξτ =X
τ ∈T
, 2
(28)
τ ∈T
where Ξτ ∈ Rn1 ×...×nd , τ ∈ T are auxiliary variables. Proof. Since T is finite, the constraint
∥M ∥†Σ = max ∥M(τ ) ∥∗ ≤ 1
(29)
∥M(τ ) ∥∗ ≤ 1
(30)
τ ∈T
is equivalent to the collection of constraints ∀τ ∈ T .
Hence, by introducing auxiliary tensors Aτ that duplicate M , we can rewrite the definition of the dual norm as n o τ τ ∥X∥Σ = max ⟨M, X⟩ : M = A , ∥A ∥ ≤ 1 ∀τ ∈ T . (τ ) ∗ τ
(31)
Equivalently, we consider the minimization problem n o τ τ min −⟨M, X⟩ : M = A , ∥A ∥ ≤ 1 ∀τ ∈ T . ∗ (τ ) τ
(32)
M,(A )τ ∈T
M,(A )τ ∈T
Its optimal value is −∥X∥Σ . This is a finite-dimensional convex problem, and Slater’s condition holds (for instance, M = 0 and Aτ = 0 strictly satisfy all inequality constraints). Therefore, strong duality applies. Introduce dual variables Ξτ for the constraints M = Aτ and nonnegative multipliers λτ for the constraints ∥Aτ(τ ) ∥∗ ≤ 1. The Lagrangian is i Xh L(M, A, λ, Ξ) = −⟨M, X⟩ + λτ ∥Aτ(τ ) ∥∗ − 1 + ⟨Ξτ , M − Aτ ⟩ . (33) τ ∈T
Rearranging terms gives * L(M, A, λ, Ξ) =
−X +
+ X
τ
Ξ , M
τ ∈T
+
i Xh λτ ∥Aτ(τ ) ∥∗ − ⟨Ξτ , Aτ ⟩ − λτ .
(34)
τ ∈T
We now minimize over the primal variables. First, if X
Ξτ ̸= X,
(35)
τ ∈T
then the term linear in M is nonzero, and minimizing over M yields −∞. Thus the dual function is finite only if X Ξτ = X.
(36)
τ ∈T
Assume from now on that this constraint holds. Then the minimization over Aτ separates across τ . For each τ , we use that unfolding preserves the Frobenius inner product: ⟨Ξτ , Aτ ⟩ = ⟨Ξτ(τ ) , Aτ(τ ) ⟩.
(37)
Next, recall the standard duality between the nuclear and spectral norms: ⟨U, V ⟩ ≤ ∥U ∥2 ∥V ∥∗ . 11
(38)
Tensorion: A Tensor-Aware Generalization of the Muon Optimizer
Applying this with U = Ξτ(τ ) and V = Aτ(τ ) gives ⟨Ξτ , Aτ ⟩ ≤ ∥Ξτ(τ ) ∥2 ∥Aτ(τ ) ∥∗ .
(39)
λτ ∥Aτ(τ ) ∥∗ − ⟨Ξτ , Aτ ⟩ − λτ ≥ λτ − ∥Ξτ(τ ) ∥2 ∥Aτ(τ ) ∥∗ − λτ .
(40)
Therefore,
If λτ ≥ ∥Ξτ(τ ) ∥2 , the right-hand side is bounded below by −λτ , and this bound is attained at Aτ = 0. Hence the minimum over Aτ equals −λτ . If instead λτ < ∥Ξτ(τ ) ∥2 , then the coefficient in front of ∥Aτ(τ ) ∥∗ is negative. Using the fact that the spectral norm is dual to the nuclear norm, one can choose a matrix B τ with ∥B τ ∥∗ = 1,
⟨Ξτ(τ ) , B τ ⟩ = ∥Ξτ(τ ) ∥2 .
(41)
eτ and sending t → ∞ shows that the eτ be the tensor whose τ -matricization is B τ . Evaluating the Lagrangian at tA Let A infimum is then −∞. Thus the dual function is G(λ, Ξ) =
P − λτ , τ ∈T
−∞,
P τ ∈T
Ξτ = X and λτ ≥ ∥Ξτ(τ ) ∥2 ∀τ ∈ T ,
(42)
otherwise.
The dual problem is therefore max G(λ, Ξ) = − λ,Ξ
P
X
minτ
λτ .
τ ∈T Ξ =X τ ∈T λτ ≥∥Ξτ(τ ) ∥2 , ∀τ ∈T
(43)
For fixed (Ξτ )τ ∈T , the minimum over λτ is attained at λτ = ∥Ξτ(τ ) ∥2 .
(44)
Hence the dual optimal value is −
X
min τ
P(Ξ )τ ∈T τ τ ∈T τ ∈T Ξ =X
∥Ξτ(τ ) ∥2 .
(45)
X
(46)
Finally, by strong duality, −∥X∥Σ = −
min τ
P(Ξ )τ ∈T τ τ ∈T τ ∈T Ξ =X
∥Ξτ(τ ) ∥2 ,
and multiplying by −1 yields ∥X∥Σ =
min τ
X
P(Ξ )τ ∈T τ τ ∈T τ ∈T Ξ =X
∥Ξτ(τ ) ∥2 .
(47)
This proves the claim. Lemma A.2. Let ∥ · ∥A , ∥ · ∥B be any norms defined on the same finite-dimensional linear space V and ∥ · ∥†A , ∥ · ∥†B are their duals. Let also ∥v∥A ≤ ∥v∥B , v ∈ V, (48) then ∥u∥†A ≥ ∥u∥†B , 12
u ∈ V ∗.
(49)
Tensorion: A Tensor-Aware Generalization of the Muon Optimizer
Proof. By definition of the dual norm: ∥u∥†A =
∥u∥†B =
max ⟨v, u⟩, v∈BA (0,1)
max ⟨v, u⟩, v∈BB (0,1)
(50)
where BA (0, 1) := {v ∈ V | ∥v∥A ≤ 1} ,
BB (0, 1) := {v ∈ V | ∥v∥B ≤ 1} .
(51)
From Eq. (48) it follows that BB (0, 1) ⊆ BA (0, 1), therefore maximum over BA (0, 1) in Eq. (50) is not smaller than maximum over BB (0, 1) yielding the proposed inequality. Proposition A.3. Let X ∈ Rn1 ×···×nd be any tensor then ∥X∥†Σ ≤ ∥X∥†σ and ∥X∥σ ≤ ∥X∥Σ ≤ min ∥X(τ ) ∥2 . τ ∈T
(52)
Proof. We begin by comparing the dual norms. By the definition of ∥ · ∥†Σ (see Eq. (15)), ∥M ∥†Σ = max ∥M(τ ) ∥∗ . τ ∈T
(53)
Applying Lemma A.2 to ∥X(τ ) ∥2 ≥ ∥X∥σ for each unfolding yields ∥M(τ ) ∥∗ ≤ ∥M ∥∗
∀τ ∈ T .
(54)
Therefore, ∥M ∥†Σ = max ∥M(τ ) ∥∗ ≤ ∥M ∥∗ . τ ∈T
(55)
Finally, the tensor nuclear norm is dual to the tensor spectral norm (Lim & Comon, 2014, Lemma 21), so ∥M ∥∗ = ∥M ∥†σ . Hence ∥M ∥†Σ ≤ ∥M ∥†σ . We now turn to the lower bound on ∥X∥Σ . Applying Lemma A.2 to ∥M ∥†Σ ≤ ∥M ∥†σ
(56)
∥X∥σ ≤ ∥X∥Σ
(57)
yields
It remains to establish the upper bound. By the variational representation of ∥ · ∥Σ (see Eq. (28)), ∥X∥Σ =
min τ
X
P(Ξ )τ ∈T τ τ ∈T τ ∈T Ξ =X
Fix any τ ′ ∈ T , and consider the feasible decomposition ( Ξτ =
∥Ξτ(τ ) ∥2 .
X, τ = τ ′ , 0, τ ̸= τ ′ .
Substituting this choice into the objective yields X ′ ∥X∥Σ ≤ ∥Ξτ(τ ) ∥2 = ∥Ξτ(τ ′ ) ∥2 = ∥X(τ ′ ) ∥2 .
(58)
(59)
(60)
τ ∈T
Since τ ′ was arbitrary, we conclude that ∥X∥Σ ≤ min ∥X(τ ) ∥2 . τ ∈T
This completes the proof. 13
(61)
Tensorion: A Tensor-Aware Generalization of the Muon Optimizer
Proposition A.4. The solution set of the LMO in Eq. (19) is the subdifferential of the relaxed spectral dual norm at point M: X opt ∈ ∂∥M ∥†Σ = conv ∂∥M(τ ) ∥∗ , (62) τ ∈I
where conv (·) denotes a convex hull and n o I = τ ∥M(τ ) ∥∗ = ∥M ∥†Σ ⊆ 2{1,...,d}
(63)
is a nonempty set of unfolding indices attaining the maximal nuclear norm. Proof. First, by definition of the relaxed spectral dual norm (see Eq. (15)), ∥M ∥†Σ = max ∥M(τ ) ∥∗ .
(64)
fτ (M ) := ∥M(τ ) ∥∗ .
(65)
τ ∈T
For every τ ∈ T , define Since the unfolding map M 7→ M(τ ) is linear and the nuclear norm is convex, each function fτ is convex in M . Because the index set T is finite, we may apply the Dubovitskii-Milyutin subdifferential formula (see, e.g.(Girsanov, 2012)) for the maximum of finitely many convex functions: ! [ ∂ max fτ (M ) = conv ∂fτ (M ) , (66) τ ∈T
where
τ ∈I
I=
τ ∈T
fτ (M ) = max fθ (M ) . θ∈T
(67)
Substituting the definition of fτ , we obtain
with
∂∥M ∥†Σ = convτ ∈I ∂∥M(τ ) ∥∗ ,
(68)
n I= τ ∈T
(69)
o ∥M(τ ) ∥∗ = ∥M ∥†Σ .
It remains to identify this subdifferential with the solution set of the LMO in Eq. (19). For any norm ∥ · ∥A and its dual norm ∥ · ∥†A (Rockafellar, 1997, Theorem 23.5) n o ∂∥x∥A = Argmax ⟨z, x⟩ ∥z∥†A ≤ 1 , x ̸= 0. (70) Applying this general result with we obtain
∥ · ∥A = ∥ · ∥†Σ ,
∥ · ∥†A = ∥ · ∥Σ ,
∂∥M ∥†Σ = Argmax {⟨X, M ⟩ | ∥X∥Σ ≤ 1} .
(71) (72)
Combining this identity with the representation established above completes the proof.
B. High-Probability Nuclear Norm Bound Proposition B.1 (High-Probability Nuclear Norm Bound). Let M ∈ Rn1 ×···×nd be a random tensor with entries i.i.d. Mi1 ,...,id ∼ P, where P is a zero-mean sub-exponential distribution. Consider any unfolding M(τ ) ∈ Rmτ ×nτ Qd with mτ ≤ nτ such that mτ nτ = i=1 ni . Then, there exists a constant C > 0 (depending on the sub-exponential norm of P) such that with probability at least 1 − δ: √ M(τ ) ∗ ≤ C · mτ nτ + log(1/δ) (73) Moreover, for a fixed total dimension mτ nτ , the value of this upper bound is maximized when the unfolding is square, i.e., √ mτ = nτ = mτ nτ . 14
Tensorion: A Tensor-Aware Generalization of the Muon Optimizer
Proof. Let x := vec(M ) ∈ RN ,
N=
d Y
ni .
(74)
i=1
Since unfolding only permutes the entries of the tensor, we have ∥M(τ ) ∥F = ∥M ∥F = ∥x∥2
(75)
for every admissible τ ∈ T . We first recall the standard comparison between the nuclear and Frobenius norms. Let A ∈ Rm×n
(76)
σ1 ≥ · · · ≥ σr > 0,
(77)
have singular values where r = rank(A) ≤ min(m, n). Then r X
∥A∥∗ =
σi
(78)
i=1
and ∥A∥F =
r X
!1/2 σi2
.
(79)
i=1
By the Cauchy–Schwarz inequality, ∥A∥∗ =
r X
σi · 1 ≤
i=1
r X
!1/2 σi2
i=1
r X
!1/2 12
=
√
r ∥A∥F ≤
p min(m, n) ∥A∥F .
(80)
i=1
In particular, since mτ ≤ nτ , ∥M(τ ) ∥∗ ≤
√
mτ ∥M(τ ) ∥F =
√
mτ ∥x∥2 .
(81)
It therefore remains to control ∥x∥2 . Since the entries of M are i.i.d., centered, unit-variance, and sub-exponential, (Zhao, 2026, Lemma 4.1) applied with α = 1 yields the concentration bound √ P ∥x∥2 − N ≥ t ≤ 2 exp −c min(t2 , t) , t ≥ 0, (82) for some constant c > 0 depending only on the sub-exponential norm of P. Now fix δ ∈ (0, 1). Choose tδ > 0 so that 2 exp −c min(t2δ , tδ ) ≤ δ. Then, with probability at least 1 − δ, ∥x∥2 ≤
(83)
√ N + tδ .
Since N ≥ 1, after enlarging the constant if necessary, we may write this bound in the simpler form √ ∥x∥2 ≤ Cδ N .
(84)
(85)
Consequently, for every fixed τ we also have the slightly weaker but algebraically convenient estimate √ 1 ∥x∥2 ≤ Cδ N + √ log(1/δ), mτ
(86)
again with probability at least 1 − δ. Combining this with the nuclear-versus-Frobenius estimate, we obtain √ √ √ 1 ∥M(τ ) ∥∗ ≤ mτ ∥x∥2 ≤ mτ Cδ N + √ log(1/δ) . mτ 15
(87)
Tensorion: A Tensor-Aware Generalization of the Muon Optimizer
Therefore, ∥M(τ ) ∥∗ ≤ Cδ
p
mτ N + log(1/δ).
(88)
Using
we conclude that
N = m τ nτ ,
(89)
√ ∥M(τ ) ∥∗ ≤ Cδ mτ nτ + log(1/δ),
(90)
as claimed. Finally, let us justify the statement about the “most square” unfolding. For fixed total dimension N and under the constraint mτ ≤ nτ , we have N (91) nτ = mτ and hence √
mτ nτ = mτ
r
p N = N mτ . mτ
(92)
Thus the leading term is a monotone increasing function in mτ . Since admissible unfoldings are determined by partitions of the tensor modes, not every factorization of N need to be realizable. Therefore, among the admissible unfoldings, the largest upper bound is attained by the unfolding with the largest possible value of mτ subject to mτ ≤ nτ , namely by the most balanced admissible unfolding.
C. Additional experiments C.1. CIFAR-10 and CIFAR-100 This section presents additional training results for ResNet-18 and ResNet-34 in the CIFAR10 and CIFAR100 datasets respectively. Unlike previous experiments, here we also evaluate the Tensorion Algorithm 1 using the AdamW optimizer for the non-tensor blocks; this setting is denoted as tensorion adamw in Fig. 4 and 5. It is also important to analyze how the algorithm’s behavior depends on the optimizer chosen for the non-tensor parameters of the model. Indeed, the reported experiments show that the behavior strongly depends on this choice. When SGD is used, we observe improvements in both convergence and accuracy compared to applying SGD to all model parameters. For AdamW, we observe a roughly similar picture, where both algorithms achieve approximately the same performance. C.2. ViT & Swin models Here we present some additional details on evaluations on transformer-based models. We initialize all the models from pretrained checkpoints available in the timm library and finetune the models for 5 epochs on Imagenette dataset. Tensorion is only applied to convolutional layers, while we revert to either AdamW or SGD for other layers to isolate the effects. Notably, applying Tensorion to weight matrices in Transformer blocks would recover Muon updates. We fix weight decay and LR scheduler parameters for this experiment, and only sweep over optimal learning rates for the respective choice of optimizers. We report best runs for each choice of the optimizer in 6 and 7. C.3. Beyond CNNs: Laplacian Spectrum Tensor trains (TT). Tensor Train (TT) decomposition is a low-rank factorization for high-order tensors that represents a d-way tensor x ∈ Rn1 ×···×nd as a product of TT-cores: x(i1 , . . . , id ) = G(1) (i1 ) G(2) (i2 ) · · · G(d) (id ),
(93)
where each core G(k) (ik ) ∈ Rrk−1 ×rk is a matrix indexed by ik ∈ {1, . . . , nk }, with boundary ranks r0 = rd = 1. The tuple (r1 , . . . , rd−1 ) are the TT-ranks; small TT-ranks yield a compressed representation that enables scalable optimization in very high dimensions. 16
Accuracy
Tensorion: A Tensor-Aware Generalization of the Muon Optimizer
0.8
0.6 Tensorion-SGD Tensorion-AdamW
0.4 0
25
50
75
100
125
SGD AdamW
150
175
200
Epochs (a) Test accuracy vs Epoch Tensorion-SGD Tensorion-AdamW
2.0
SGD AdamW
Train Loss
Test Loss
1.5
1.0
0.5
Tensorion-SGD Tensorion-AdamW
SGD AdamW
1.5 1.0 0.5 0.0
0
25
50
75
100
125
150
175
200
0
25
Epochs
50
75
100
125
150
175
200
Epochs
(b) Test loss vs Epoch
(c) Train loss vs Epoch
Figure 4. CIFAR-10 results for ResNet-18 trained for 200 epochs under the standard CIFAR pipeline of (Contributors, 2023). We report mean ± std over multiple seeds for test accuracy (Figure 4a), test loss (Figure 4b), and train loss (Figure 4c). Tensorion is used only for tensor-valued blocks with order d ≥ 3 (e.g., convolutional kernels), while all remaining parameters (d ≤ 2) are optimized with a first-order method (SGD or AdamW).
Discrete Laplacian on [0, 1]d . Consider a uniform tensor-product grid on [0, 1]d with n points per dimension and step 1 size h = n+1 (Dirichlet boundary conditions). Let L ∈ Rn×n be the standard 1D second-difference matrix L=
1 tridiag(1, −2, 1), h2
(94)
where tridiag(1, −2, 1) denotes tridiagonal matrix. d
d
The d-dimensional discrete Laplacian ∆d ∈ Rn ×n on the tensor-product grid is given by the Kronecker sum ∆d =
d X
I ⊗(k−1) ⊗ L ⊗ I ⊗(d−k) ,
(95)
k=1
where I ∈ Rn×n is the identity and ⊗ denotes the Kronecker product. We treat ∆d as a linear operator acting on an order-d tensor x ∈ Rn×···×n via vectorization. TT-operator form. Each term I ⊗(k−1) ⊗ L ⊗ I ⊗(d−k) in (95) is a rank-1 Kronecker product of d factors and thus admits an exact TT-operator representation with TT-rank 1. Consequently, the full Laplacian ∆d , being a sum of d such Kronecker terms, admits a compact TT-operator representation with small TT-ranks (bounded by a constant independent of nd ), which allows for efficient application ∆d x and inner products ⟨x, ∆d x⟩ in TT format. 17
Tensorion: A Tensor-Aware Generalization of the Muon Optimizer
Accuracy
0.8 0.6 0.4 0.2
Tensorion-SGD Tensorion-AdamW
0
50
100
150
SGD AdamW
200
250
Epochs (a) Test accuracy vs Epoch Tensorion-SGD Tensorion-AdamW
SGD AdamW
Tensorion-SGD Tensorion-AdamW
4
Train Loss
Test Loss
4
3
2
SGD AdamW
3 2 1
1
0 0
50
100
150
200
250
0
Epochs
50
100
150
200
250
Epochs
(b) Test loss vs Epoch
(c) Train loss vs Epoch
Figure 5. CIFAR-100 results for ResNet-34 trained for 250 epochs under the standard CIFAR pipeline of (Contributors, 2023). We report mean ± std over multiple seeds for test accuracy (Figure 5a), test loss (Figure 5b), and train loss (Figure 5c). Tensorion is used only for tensor-valued blocks with order d ≥ 3 (e.g., convolutional kernels), while all remaining parameters (d ≤ 2) are optimized with a first-order method (SGD or AdamW).
Low-rank structure of eigenvectors. Our goal is to compute the low end of the spectrum of ∆d in a compressed form. We restrict the search space to eigenvectors representable with small TT-ranks, i.e., x is parameterized by TT-cores with ranks bounded by a prescribed r. This reflects the practical regime in high-dimensional PDE discretizations where one seeks approximate eigenmodes within a low-rank manifold. Rayleigh quotient optimization. Let A := −∆d be the discrete positive semidefinite operator. The smallest eigenpair of A can be obtained by minimizing the Rayleigh quotient min R(x) = min x̸=0
x̸=0
⟨x, Ax⟩ , ⟨x, x⟩
(96)
In our experiments, we optimize (96) over the TT-parameterization of x with bounded TT-ranks, comparing Tensorion to gradient descent baselines on this non-deep, structured optimization task.
D. Relaxation Tightness The empirical analysis of relaxation tightness along optimization trajectories of ResNet-18, CIFAR-10 is demonstrated in Fig. 9. The results show that the actual tensor spectral norm of the LMO solutions with relaxed tensor spectral norm on optimization trajectories of ResNet-18 is very close to 1. 18
Tensorion: A Tensor-Aware Generalization of the Muon Optimizer
Tensorion-SGD Tensorion-AdamW
0.8
1.5
Accuracy
Train loss
2.0
0.9
SGD AdamW
1.0
0.7 0.6
0.5 0.5
Tensorion-SGD Tensorion-AdamW
0.0 0
100
200
300
400
500
600
700
100
200
300
Steps
600
700
(b) Swin Tiny Accuracy
Tensorion-SGD Tensorion-AdamW
SGD AdamW
0.90 0.85
1.5
Accuracy
Train loss
500
Steps
(a) Swin Tiny Train Loss 2.0
400
SGD AdamW
1.0
0.80 0.75
0.5 Tensorion-SGD Tensorion-AdamW
0.70
0.0 0
100
200
300
400
500
600
700
100
Steps
300
400
500
600
700
Steps
(c) Swin Small Train Loss
(d) Swin Small Accuracy
Tensorion-SGD Tensorion-AdamW
2.0
200
SGD AdamW
SGD AdamW
0.9
Accuracy
Train loss
0.8 1.5 1.0
0.7 0.6
0.5
0.5
0.0
0.4 0
100
200
300
400
500
600
700
Tensorion-SGD Tensorion-AdamW
100
200
300
Steps
400
500
Steps
(e) Swin Base Train Loss
(f) Swin Base Accuracy
Figure 6. Test accuracy and training loss across all Swin models.
19
SGD AdamW
600
700
0.80
2.5
0.75
2.0
Train loss
Accuracy
Tensorion: A Tensor-Aware Generalization of the Muon Optimizer
0.70 0.65 0.60 Tensorion-SGD Tensorion-AdamW
0.55 100
200
300
400
500
1.5 1.0
0.0 0
700
100
200
300
(a) ViT Small Accuracy
500
600
700
(b) ViT Small Train Loss Tensorion-SGD Tensorion-AdamW
2.5
0.9
SGD AdamW
2.0
Train loss
0.8
Accuracy
400
Steps
Steps
0.7 0.6 0.5
Tensorion-SGD Tensorion-AdamW
0.4 100
200
300
400
500
1.0 0.5
SGD AdamW
600
1.5
0.0
700
0
100
200
300
Steps
400
500
600
700
Steps
(c) ViT Base Accuracy
(d) ViT Base Train Loss
0.9
Tensorion-SGD Tensorion-AdamW
3.0
0.8
SGD AdamW
2.5
Train loss
0.7
Accuracy
SGD AdamW
0.5
SGD AdamW
600
Tensorion-SGD Tensorion-AdamW
0.6 0.5 0.4 Tensorion-SGD Tensorion-AdamW
0.3 100
200
300
400
500
1.5 1.0 0.5
SGD AdamW
600
2.0
0.0
700
0
100
200
Steps
300
400
500
Steps
(e) ViT Medium Accuracy
(f) ViT Medium Train Loss
Figure 7. Test accuracy and training loss across ViT models.
20
600
700
105
105
10
4
104
10
3
103
Residual
residual
Tensorion: A Tensor-Aware Generalization of the Muon Optimizer
102 101
101
100 10
102
100 SGD Tensorion
−1
10−7
10−8
GD Tensorion
10−1 10−5
10−6
10−4
10−3
0
1000
2000
lr
3000
4000
5000
Step
(a) Residual vs learning rate
(b) Residual vs step, optimal learning rate
Figure 8. Minimization of the Rayleigh quotient (96) for the smallest eigenpair of A = −∆d on a d-dimensional tensor-product grid, with the eigenvector x parameterized in TT format with bounded TT-ranks. Left: final functional-residual as a function of the learning rate. Right: functional-residual versus optimization step at the best-tuned learning rate for each method.
0.996 0.995
Mean layers
0.994 0.993 0.992 0.991 0.990 0.989
0
20
40
Step
60
80
100
Figure 9. Actual tensor spectral norm of the LMO solutions with relaxed tensor spectral norm on optimization trajectories of ResNet-18, CIFAR-10
E. Experimental Details E.1. Tau ablation We enumerate all proper non-empty subsets τ ⊊ {1, 2, 3, 4}. Since unfolding with τ and with its complement τ c produces matrices that are transposes of each other, and the nuclear norm is invariant to this operation, ∥X(τ ) ∥∗ = ∥X(τ c ) ∥∗ , it suffices to consider τ up to the equivalence τ ∼ τ c , yielding (24 − 2)/2 = 7 distinct cases for 4-dimensional convolution kernel. The convolutional kernel has shape (Cout , Cin , kh , kw ). 21
In ResNet architectures, it typically holds that
Tensorion: A Tensor-Aware Generalization of the Muon Optimizer
Cout ≥ Cin ≫ kh , kw . In case of static τ selection, the unfolding is fixed across all layers, while the offline strategy selects τ opt for each layer individually. To decouple the effect of τ selection from numerical approximation artifacts, we replace the Newton–Schulz routine (line 7 in Algorithm 1) with an exact SVD throughout this ablation. AdamW was used for layers with weight tensors of order d ≤ 2. We train ResNet-18 (He et al., 2016) on CIFAR-10 (Krizhevsky et al., 2009) for a fixed budget of 25 epochs (no early stopping). Across all optimizers, we keep the batch size, weight decay, and image preprocessing identical. We do not use an LR scheduler; instead, the learning rate is tuned independently for each method via a grid search (see Table 4).
F. Hyperparameters The optimal hyperparameter configurations reported throughout the experiments are summarized in Table 3 and Table 4. Table 3. Hyperparameter configurations used across datasets and optimization setups.
Setup / Method
SGD
AdamW
Tensorion
Tensorion AdamW
CIFAR10 Scheduler CIFAR10 LR CIFAR10 WD
Multistep 0.075 0.0005
Multistep 0.0001 0.0001
Onecycle 0.001 0.0001
Onecycle 0.075 0.0005
CIFAR100 Scheduler CIFAR100 LR CIFAR100 WD
Multistep 0.05 0.0005
Multistep 0.00025 0.0005
Onecycle 0.0075 0.0005
Onecycle 0.00005 0.0005
Tiny Scheduler Tiny LR Tiny WD
Onecycle 0.01 0.0005
Onecycle 0.0005 0.0001
Onecycle 0.075 0.0005
– – –
Imagenette Scheduler Imagenette LR Imagenette WD
Cosine 0.005 0.0005
Cosine 0.0001 0.0005
Cosine 0.0005 0.0005
Cosine 0.0001 0.0005
Table 4. Learning rates used in the ablation on the unfolding set τ for ResNet-18 on CIFAR-10 after 25 epochs. The learning rate is tuned per method via grid search.
{2}
Tensorion (static τ )
Offline
Online
Baselines
{3}
τ opt
online τ
SGD+M. AdamW
0.01
0.005
Hyperparameter
{1}
{4}
{1, 2} {1, 3} {1, 4}
Learning rate
0.01 0.001 0.001 0.005 0.005
0.001
0.005
0.05
0.0005
G. Compute resources All experiments were conducted on a shared high-performance computing infrastructure. The experiments reported in Section 5.1 required approximately 300 GPU-hours on NVIDIA V100-SXM2-32GB accelerators. All remaining experiments, including baseline training, ablation studies, and final model evaluations, consumed approximately 2,000 GPU-hours on NVIDIA A100-SXM-80GB accelerators. In total, the computational budget for the reported results amounts to roughly 2,300 GPU-hours. Hyperparameter search, preliminary runs, and failed attempts are included in these estimates. We confirm that the paper reports the full set of experiments conducted within this budget, and no selective reporting based on compute availability was applied.
22