FedSPDnet: Geometry-Aware Federated Deep Learning with SPDnet Thibault Pautrel1 , Florent Bouchard1 , Ammar Mian3 , Guillaume Ginolhac3 1
arXiv:2604.22494v1 [stat.ML] 24 Apr 2026
2
CentraleSupélec, L2S, France Université Savoie Mont-Blanc, LISTIC, France
architectures that assume that all clients use the SPDnet network. To develop a manifold-based federated learning model, the main challenge lies in aggregating parameters: while local training on clients leverages optimization methods adapted to the manifold, the server-side averaging step must also be designed to obtain a result that belongs to the manifold. The optimal solution is therefore to use a Riemannian mean [7], which allows to obtain the barycenter of a set of parameters belonging to the specific manifold. Unfortunately, this step is tricky for SPDnet. This is because SPDnet network parameters are mainly orthogonal matrices that belong to the Stiefel manifold. However, it is well known [16] that calculating the Riemannian mean in this manifold is computationally expensive and therefore impossible to use for large-scale federated learning. Inspired by the geometric tools from [16], we propose two I. I NTRODUCTION lightweight, geometry-preserving aggregation strategies for Federated learning (FL) enables collaborative model training SPDnet: (i) ProjAvg, which computes Euclidean averages of without centralizing raw data, typically by iteratively averaging local weights and maps them back onto the Stiefel manifold parameters on a central server as in FedAvg [1]–[3]. Subsequent via polar decomposition, and (ii) RLAvg, which approximates refinements such as FedProx [3], [4] and SCAFFOLD [5] tangent-space Riemannian averaging using retractions and improve robustness under data heterogeneity and client drift, liftings. Both methods are computationally efficient, simple to but remain limited to Euclidean parameter spaces and do not implement, and guarantee that Stiefel constraints are preserved extend to non-Euclidean geometries. throughout training. By combining strong geometric motivaIn parallel, the last decade has seen a surge of interest tions with practical algorithmic design, our work proposes a in Riemannian optimization, which generalizes classical op- scalable deployment of SPDnet in federated learning applicatimization to parameters living on manifolds [6], [7]. Such tions. In particular, the proposed architecture is more robust techniques are crucial for learning problems with geometric than the one proposed in [14] because it does not depend on constraints, including low-rank matrix factorization, dimen- the local optimizer used in the clients. Numerical experiments sionality reduction, and orthogonal parameterizations. In deep on real EEG data show that the performance of our federated learning, Riemannian models have proven particularly effective models matches that of a centralized model for a small number when operating on structured data such as symmetric positive of clients. definite (SPD) covariance matrices, leading to architectures II. BACKGROUND like SPDnet [8]. SPDnet integrates bilinear mappings on the Stiefel manifold with nonlinear eigenvalue rectification, This section presents the key elements that are needed to ensuring that intermediate representations remain SPD while develop our proposed federated learning method adapted to reducing dimensionality in a geometry-preserving way. This the SPDnet architecture. Since SPDnet weights are orthogonal, architecture has proven its effectiveness in several applications we first introduce Riemannian geometry and optimization on such as micro-Doppler radar [9], electroencephalography [10], the Stiefel manifold. Then, we provide details on the SPDnet and ground-penetrating radar [11]. architecture. Some works [12]–[15] have proposed adapting federated learning architectures to Riemannian manifolds, but these A. Riemannian geometry and optimization on Stiefel studies are fairly general and do not concern the SPDnet The Stiefel manifold [6], [7] is defined as Stp,k = {W ∈ network. In this paper, we therefore propose new federated Rp×k : W ⊤ W = Ik }. At W ∈ Stp,k , the tangent space Abstract—We introduce two federated learning frameworks for the classical SPDnet model operating on symmetric positive definite (SPD) matrices with Stiefel-constrained parameters. Unlike standard Euclidean averaging, which violates orthogonality, our approach preserves geometric structure through two efficient aggregation strategies: ProjAvg, projecting arithmetic means onto the Stiefel manifold, and RLAvg, approximating tangentspace averaging via retractions and liftings. Both methods are computationally efficient, independent of the optimizer, and enable scalable federated learning for signal processing applications whose features are SPD matrices. Simulations on EEG motor imagery benchmarks show that FedSPDnet outperforms federated EEGnet in F1 score and robustness to federation and partial participation, while using fewer parameters per communication round. Index Terms—Federated learning, Riemannian manifold, symmetric positive definite matrices
is TW Stp,k = {X ∈ Rp×k : W ⊤ X + X ⊤ W = 0}. We endow it with the Euclidean metric: for X, Y ∈ TW Stp,k , ⟨X, Y ⟩W = tr(X ⊤ Y ). From there, in the context of federated learning, we need geometrical tools on Stp,k to: (i) perform Riemannian optimization, and (ii) aggregate points on Stp,k . Concerning the optimization of a cost function, one needs the Riemannian gradient to obtain a descent direction on the tangent space and a retraction to get a new iterate on the manifold from it. Given a cost function f : Stp,k → R, the Riemannian gradient gradf (W ) at W is the unique tangent vector such that, ∀X ∈ TW Stp,k , ⟨gradf (W ), X⟩W = D f (W )[X], where D f (W )[X] denotes the directional derivative of f at W in direction X. For Stp,k , it can be obtained from the Euclidean gradient ∇f (W ) through gradf (W ) = PW (∇f (W )), where PW : Rp×k → TW Stp,k is the orthogonal projection [6], [7] 1 PW (X) = X − W (W ⊤ X + X ⊤ W ). (1) 2 On Stiefel, an advantageous retraction RW : TW Stp,k → Stp,k at W is [6], [7] RW (X) = P(W + X) = uf(W + X),
(2)
the aggregation step performed in the central server, which is delicate for the SPDnet model. Indeed, despite the works proposed to develop federated learning on manifolds [12]–[15], they are unfortunately not appropriate in the context of SPDnet due to the geometric structure of the parameters (the weights belong to the Stiefel manifold). As stated in the background, this is because the Riemannian mean of parameters living on the Stiefel manifold is either untractable or computationally expensive [7]. Therefore, this strategy is impossible to use for large-scale federated learning. In this paper, we propose two new strategies to aggregate weight matrices in Stiefel. The first is inspired by [12] but unlike computing the true Riemannian mean, the efficient mean on Stiefel from [16] is exploited. Our second strategy is inspired by [15], which avoids having to compute a mean on the manifold by averaging local (transported) Riemannian gradients instead. However, this approach is limited to using a specific Riemannian stochastic gradient descent for the local optimizer on every client. Our proposed method is more general and allows to exploit any optimizer (e.g. SGD, Adam, . . . ) for each client, which works well with SPDnet. FedSPDnet is summarized in Algorithm 1.
A. Proposed model where P : Rp×k → Stp,k is the projection on Stiefel obtained Let C = {1, . . . , N } denote the set of all clients. The client by taking the orthogonal factor of the polar decomposition. i has the training dataset To aggregate points on a Riemannian manifold, the natural n o|D(i) | solution is to exploit the Riemannian mean. To be able (i) (i) ++ (i) D = (Σ , y ) ∈ S × J1, KK , (3) to compute it, one needs three elements: the Riemannian j j d0 j=1 exponential and logarithm mappings, and the Riemannian (i) (i) distance [6], [7]. These three objects rely on geodesics, which where yj is the class label associated with the covariance Σj generalize the concept of straight lines on the manifold. and |D(i) | is the cardinal of the dataset D(i) . Given SPDnet Unfortunately, for Stp,k , the Riemannian logarithm and distance learnable parameters θ = (W1 , . . . , WL , ξ, β), with, ∀ℓ ∈ are unknown. Hence, the Riemannian mean is not available J1, LK, BiMap weights Wℓ ∈ Stdℓ−1 ,dℓ , and softmax parameters and another solution needs to be found to aggregate points. It ξ ∈ RK×dL (dL +1)/2 and β ∈ RK , the empirical risk based on will be the contributions of Section III-B. cross-entropy [17] on client i is B. SPDnet architecture
(i)
|D | K −1 X X (i) Fi (θ) = (i) δy(i) =k log [fθ (Σj )]k , |D | j=1 k=1 j
(4) SPDnet [8] is a lightweight deep learning architecture (see Figure 1) operating on SPD (covariance) matrices in a SPD-preserving way while reducing dimension thanks where δ (i) is the Kronecker delta and fθ (Σ(i) j ) denotes the yj =k to BiRe blocks. Such blocks combine a BiMap layer, (i) performing bilinear projection for dimensionality reduction: forward pass of Σj through SPDnet. On the server, the objective is to minimize the global Σ̄ℓ = Wℓ⊤ Σℓ−1 Wℓ , Wℓ ∈ Stdℓ−1 ,dℓ , and a ReEig layer empirical risk over all clients, i.e., solve inducing non-linearity by eigenvalues rectification: Σℓ = N U max(εIdℓ , Λ)U ⊤ , Σ̄ℓ = U ΛU ⊤ . After L BiRe blocks, 1 X argmin Fi (θ). (5) the representation ΣL−1 is mapped from the SPD manifold N i=1 θ to the Euclidean space through the LogEig layer ΣL = U log(Λ)U ⊤ , ΣL−1 = U ΛU ⊤ and finally fed to a softmax To find a solution collaboratively, T client-server communicaclassifier via half-vectorization, ŷ = softmax(ξvech(ΣL ) + tion rounds are performed. At each round t, a subset St ⊂ C β), ξ ∈ RK×dL (dL +1)/2 , β ∈ RK , where K is the number of of M clients is selected uniformly at random. classes. The learnable parameters are θ = (W1 , . . . , WL , ξ, β). The first step of the round t consists in optimizing every local SPDnet model on the selected clients. For each client III. F EDERATED L EARNING WITH SPD NET A RCHITECTURE i ∈ S , the global parameters θ = (W , . . . , W , ξ , β ) t t 1,t L,t t t In this section, we give the main contribution of the paper: are set to the local SPDnet architecture. Then, Ei local epochs a new federated learning framework adapted to the SPDnet of standard SPDnet training are performed. All parameters are architecture, called FedSPDnet. The main challenge lies in optimized with Adam [18]: Stiefel weights Wℓ,t , ℓ ∈ J1, LK,
BiRe block L
BiRe block 1 Input Σ0 ∈ Sd++ 0
BiMap W1
ReEig(ε)
···
BiMap WL
ReEig(ε)
LogEig
vech
softmax ξ, β
Output ŷ
Figure 1. General SPDnet architecture [8]: a chain of L stacked BiRe (BiMap + ReEig) blocks, followed by LogEig mapping to Euclidean space and a fully connected softmax classifier. Learnable parameters are Stiefel weights W1 , . . . , WL and Euclidean weights ξ, and bias β.
are handled via tangent-space local trivialization [19], while strategy with a different rewriting of the Euclidean average. If softmax parameters ξt , βt are updated in the standard Euclidean the weights are Euclidean, we get the same update formula as way. After Ei local epochs, client i returns updated parameters in (6), which can be rewritten as (i) θt to the server. 1 X 1 X (i) (i) W (Wℓ,t − Wℓ,t ). W = = W + ℓ,t+1 ℓ,t ℓ,t From there, the second step, which is the key issue in this |St | |St | (i) i∈St i∈St work, is to aggregate the set of parameters {θt }i∈St from (8) every selected client to obtain the new global parameters θt+1 . One can then notice that we have For the Euclidean softmax parameters, the usual averaging ! 1 X technique from [1] can be employed, i.e., (i) (9) logWℓ,t (Wℓ,t ) , Wℓ,t+1 = expWℓ,t |St | 1 X (i) 1 X (i) i∈St ξt+1 = ξt , βt+1 = βt . (6) |St | |St | where expW (X) = W +X and logW (W̄ ) = W̄ −W denote i∈St i∈St However, this standard procedure cannot be applied to the the Euclidean exponential and logarithm mappings, respectively. Stiefel weights. This issue is overcome in the next subsection. Unfortunately, we cannot apply this formula directly on the Stiefel manifold because the Riemannian logarithm is unknown and the Riemannian exponential is numerically expensive and B. Orthogonality-preserving aggregation In this subsection, we propose two strategies to aggregate might not be advantageous [7]. Instead, as in [16], we exploit (i) weights {Wℓ,t }i∈St in the Stiefel manifold Stdℓ−1 ,dℓ . The some approximation of these tools. The Riemannian exponential first one is a straightforward adaptation of the usual federated and logarithm can be replaced with a retraction R and a lifting learning averaging method [1]. It relies on computing a suitable L, respectively. For Stiefel, as discussed in [16], the best mean on the Stiefel manifold. As discussed in Section II-A, choices appear to be the retraction (2) and the lifting the Riemannian mean on Stiefel is computationally intractable. LW (W̄ ) = PW (W̄ − W ), (10) Following [16], we instead project the arithmetic mean of the local weights onto the manifold. Concretely, given weights where P is defined in (1). Finally, the aggregation formula is ! (i) {Wℓ,t }i∈St , the aggregated weight reads 1 X (i) Wℓ,t+1 = RWℓ,t LWℓ,t (Wℓ,t ) . (11) ! |St | i∈St 1 X (i) Wℓ,t+1 = P Wℓ,t , (7) |St | In the following, the corresponding FedSPDnet method is i∈St denoted RLAvg. The resulting algorithm is provided in Algowhere the projection P is defined in (2). As explained in [16], rithm 1. this projection is the one that best suits the geometry of IV. N UMERICAL EXPERIMENTS the Stiefel manifold. In the following, the corresponding FedSPDnet method is denoted ProjAvg. Our second strategy We evaluate FedSPDnet on EEG-based motor imagery is inspired by [15]. In this latter work, authors avoid having to classification, where spatial covariance matrices are effective average parameters on the manifold by noticing that the usual descriptors [20], [21] and multi-site data confidentiality natfederated learning average used to get the new global parameter urally motivates federated learning. Unlike traditional BCI can be rewritten as the sum of the previous global parameter pipelines [22], [23] that rely on within-subject calibration, we and an average of some gradients of the local empirical risks target population-level models that generalize across subjects Fi . They then adapt this Euclidean formula to the Riemannian without per-user data, reflecting realistic federated scenarios. case to obtain a new federated learning method. This allows As a Euclidean baseline, we consider EEGnet [24], a compact them to only have to compute an average in the tangent space convolutional network operating on raw signals, federated with of the previous global parameters, which is fairly simple, before standard FedAvg [1]. going back on the manifold through a retraction to get the new global parameters. While this strategy is very appealing, it still A. Experimental setup suffers from a fundamental limitation. Indeed, it requires the We use two motor imagery datasets from the MOABB user to employ a specific stochastic gradient descent method benchmark [25]: Weibo2014 (10 subjects, 60 channels, samfor the local optimization procedure on the clients, which might pling frequency fs = 200 Hz, 7 classes, ∼80 trials/class) and not be adapted. To overcome this issue, we adopt the same PhysionetMI (106 subjects, 64 channels, sampling frequency
(B, 1, nchan , T )
Input Trial
(B, F1 ×D, 1, T /p1 )
Block 1 Temporal & Spatial Feature Extraction
(B, F2 , 1, T /(p1 p2 ))
Block 2 Separable Convolution
(B, ncl )
Class. Head Flatten → Linear
Output ŷ
Hyperparams: F1 =8, D=2, F2 =16, p1 =4, p2 =8 Kernels: Kt = max(⌊0.5fs ⌉, 32) Ks = max(⌊fs /(2p1 )⌉, 8)
fs = 160 Hz, 4 classes, ∼23 trials/class). Three PhysionetMI subjects were excluded due to incomplete recordings. All signals are band-pass filtered in [8, 32] Hz. EEGnet [24] operates on z-score normalized raw trials Xj ∈ Rnchan ×T through temporal, depthwise spatial, and separable convolutions followed by a classification head (Figure 2). SPDnet [8] operates on the sample covariance matrix (SCM) of each centered trial Σj = T 1−1 X̄j X̄j⊤ ∈ Sn++ . A single chan BiRe block is used, with hidden dimension d and ReEig threshold ε chosen per dataset. Prior spectral analysis of the sample covariances sets ε to balance stability against discriminative information (ε = 10−1 for Weibo2014, 10−2 for PhysionetMI), and a sweep over d retains the smallest value maximising validation F1 (d = 22 for Weibo2014, d = 18 for PhysionetMI), yielding models more compact than EEGnet. Both architectures are trained with cross-entropy loss using Adam (through local tangent space parametrisation for SPDnet), initial learning rate 10−3 , batch size 64, and ReduceLROnPlateau scheduling (patience 20, factor 0.5). All results are reported over 10 independent runs. B. Centralized baselines As an upper bound on federated performance, we pool all subjects and perform a stratified split into training (75%), validation (10%), and test (15%) sets. Training runs for at most 300 epochs with early stopping (patience 75), retaining the best validation checkpoint [26]. Results (Table I) show comparable
FedEEGNet 100%
ProjAvg 100%
RLAvg 100%
FedEEGNet 50%
ProjAvg 50%
RLAvg 50%
40
30
20
10 0
50
100
150
Round Figure 3. Convergence on Weibo2014 under full (solid) and half (dashed) participation. Shaded bands: ±1 std over 10 seeds. ProjAvg/RLAvg curves slightly offset horizontally. FedEEGNet 100%
ProjAvg 100%
RLAvg 100%
FedEEGNet 50%
ProjAvg 50%
RLAvg 50%
40
Test F1 (%)
Algorithm 1 FedSPDnet Input: number of rounds T , number of clients per round M , number of local epochs E 1: Initialize randomly θ0 = (W1,0 , . . . , WL,0 , ξ0 , β0 ) 2: for t = 0, . . . , T − 1 do 3: Randomly sample St ⊂ C with |St | = M 4: for each client i ∈ St in parallel do 5: Initialize local SPDnet with θt (i) 6: Train SPDnet over E epochs to get θt (i) 7: Send θt to the server 8: end for 9: Get ξt+1 and βt+1 with (6) 10: for ℓ = 1, . . . , L do 11: Compute Wℓ,t+1 with either (7) or (11) 12: end for 13: θt+1 = W1,t+1 , . . . , WL,t+1 , ξt+1 , βt+1 14: end for Output: θT
Test F1 (%)
Figure 2. EEGnet overall pipeline.
30
20
10 0
50
100
150
Round Figure 4. Convergence on PhysionetMI under full (solid) and half (dashed) participation. Shaded bands: ±1 std over 10 seeds. ProjAvg/RLAvg curves slightly offset horizontally.
centralized performance: SPDnet slightly outperforms EEGnet on Weibo2014 while EEGnet leads on PhysionetMI. C. Federated learning experiments We simulate a multi-institutional setting by assigning two subjects per client – mimicking small clinical sites – yielding N = 5 clients for Weibo2014 and N = 53 for PhysionetMI. Each client applies the same 75/10/15 stratified split to its local data. Training follows Algorithm 1 with T = 150 rounds and E = 2 local epochs, matching the centralized budget of 300 total epochs. We consider both full participation (all clients every round) and half participation (M = ⌊N/2⌋ clients). For FedEEGnet, we apply FedAvg excluding batch normalization statistics from aggregation [27]. D. Results and analysis a) Aggregation scheme equivalence: ProjAvg and RLAvg share the same per-round complexity of O(M nchan d + nchan d2 ).
Table I F INAL TEST F1 (%) FOR CENTRALIZED AND FEDERATED APPROACHES . M EAN ± STD OVER 10 SEEDS ; BEST PER ROW IN BOLD .
Dataset
EEGnet (FedAvg)
SPDnet (FedSPDnet)
5 303 50.7 ± 1.5 39.9 ± 2.4 36.4 ± 4.1
4 715 51.7 ± 0.8 43.3 ± 1.0 41.2 ± 0.8
Weibo2014
# params Centr. Full Half
PhysionetMI
# params 3 284 2 452 Centr. 50.8 ± 0.7 43.1 ± 0.5 Full 38.3 ± 1.8 39.5 ± 0.6 Half 38.0 ± 2.7 39.5 ± 0.9
As shown in Figures 3 and 4, their validation F1 curves overlap throughout all communication rounds on both datasets, under both full and partial participation. Since ProjAvg has slightly lower constant factors and does not require storing the previous global iterate, it is adopted as the default in Table I. b) Federated performance: FedSPDnet consistently retains more of its centralized performance than FedEEGnet across both datasets and participation rates, despite using fewer parameters (Table I). This advantage holds even on PhysionetMI where SPDnet’s centralized score is well below EEGnet’s, suggesting that geometry-aware aggregation provides intrinsic regularisation absent in Euclidean averaging. FedSPDnet also converges faster, reaching its plateau in fewer communication rounds while FedEEGnet continues to oscillate or improve slowly (Figures 3 and 4), and exhibits lower variance across seeds. V. C ONCLUSION We proposed FedSPDnet, a federated learning framework for SPDnet with two lightweight, optimizer-agnostic, geometrypreserving aggregation strategies (ProjAvg and RLAvg) that rely only on standard linear algebra. Both are empirically equivalent; ProjAvg is recommended for its simpler, storagefree implementation. On two EEG motor imagery benchmarks, FedSPDnet retains most of its centralized accuracy, degrades gracefully under partial participation and data heterogeneity, and converges faster than federated EEGnet while communicating fewer parameters per round. R EFERENCES [1] B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y Arcas, “Communication-efficient learning of deep networks from decentralized data,” in Artificial intelligence and statistics. PMLR, 2017, pp. 1273– 1282. [2] H. Yu, S. Yang, and S. Zhu, “Parallel restarted sgd with faster convergence and less communication: Demystifying why model averaging works for deep learning,” in Proceedings of the AAAI conference on artificial intelligence, vol. 33, 2019, pp. 5693–5700. [3] T. Li, A. K. Sahu, M. Zaheer, M. Sanjabi, A. Talwalkar, and V. Smith, “Federated optimization in heterogeneous networks,” Proceedings of Machine learning and systems, vol. 2, pp. 429–450, 2020. [4] X. Yuan and P. Li, “On convergence of fedprox: Local dissimilarity invariant bounds, non-smoothness and beyond,” Advances in Neural Information Processing Systems, vol. 35, pp. 10 752–10 765, 2022. [5] S. P. Karimireddy, S. Kale, M. Mohri, S. Reddi, S. Stich, and A. T. Suresh, “Scaffold: Stochastic controlled averaging for federated learning,” in International conference on machine learning. PMLR, 2020, pp. 5132–5143.
[6] P.-A. Absil, R. Mahony, and R. Sepulchre, Optimization algorithms on matrix manifolds. Princeton University Press, 2008. [7] N. Boumal, An introduction to optimization on smooth manifolds. Cambridge University Press, 2023. [8] Z. Huang and L. Van Gool, “A riemannian network for spd matrix learning,” in Proceedings of the AAAI conference on artificial intelligence, vol. 31, 2017. [9] D. Brooks, O. Schwander, F. Barbaresco, J.-Y. Schneider, and M. Cord, “Riemannian batch normalization for spd neural networks,” Advances in Neural Information Processing Systems, vol. 32, 2019. [10] R. J. Kobler, J. ichiro Hirayama, Q. Zhao, and M. Kawanabe, “SPD domain-specific batch normalization to crack interpretable unsupervised domain adaptation in EEG,” in Neurips, 2022. [11] D. Jafuno, A. Mian, G. Ginolhac, and N. Stelzenmuller, “Classification of buried objects from ground penetrating radar images by using second order deep learning models,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2025. [12] J. Li and S. Ma, “Federated learning on riemannian manifolds,” arXiv preprint arXiv:2206.05668, 2022. [13] J. Zhang, J. Hu, A. M.-C. So, and M. Johansson, “Nonconvex federated learning on compact smooth submanifolds with heterogeneous data,” Advances in Neural Information Processing Systems, vol. 37, pp. 109 817– 109 844, 2024. [14] Z. Huang, W. Huang, P. Jawanpuria, and B. Mishra, “Federated learning on riemannian manifolds with differential privacy,” arXiv preprint arXiv:2404.10029, 2024. [15] ——, “Riemannian federated learning via averaging gradient stream,” arXiv preprint arXiv:2409.07223, 2024. [16] F. Bouchard, N. Laurent, S. Said, and N. Le Bihan, “Beyond Rbarycenters: An effective averaging method on Stiefel and Grassmann manifolds,” IEEE Signal Processing Letters, vol. 32, pp. 1950–1954, 2025. [17] I. Goodfellow, Y. Bengio, A. Courville, and Y. Bengio, Deep learning. MIT press Cambridge, 2016, vol. 1. [18] D. P. Kingma, “Adam: A method for stochastic optimization,” arXiv preprint arXiv:1412.6980, 2014. [19] P. Ablin, S. Vary, B. Gao, and P.-A. Absil, “Infeasible deterministic, stochastic, and variance-reduction algorithms for optimization under orthogonality constraints,” Journal of Machine Learning Research, vol. 25, no. 389, pp. 1–38, 2024. [20] A. Barachant, S. Bonnet, M. Congedo, and C. Jutten, “Multiclass brain–computer interface classification by riemannian geometry,” IEEE transactions on biomedical engineering, vol. 59, no. 4, pp. 920–928, 2012. [21] P. Li, J. Xie, Q. Wang, and W. Zuo, “Is second-order information helpful for large-scale visual recognition?” in Proceedings of the IEEE international conference on computer vision, 2017, pp. 2070–2078. [22] F. Lotte, L. Bougrain, A. Cichocki, M. Clerc, M. Congedo, A. Rakotomamonjy, and F. Yger, “A review of classification algorithms for EEG-based brain–computer interfaces: a 10 year update,” Journal of neural engineering, vol. 15, no. 3, p. 031005, 2018. [23] I. Carrara, B. Aristimunha, M.-C. Corsi, R. Y. de Camargo, S. Chevallier, and T. Papadopoulo, “Geometric neural network based on phase space for bci-eeg decoding,” Journal of Neural Engineering, vol. 22, no. 1, p. 016049, 2025. [24] V. J. Lawhern, A. J. Solon, N. R. Waytowich, S. M. Gordon, C. P. Hung, and B. J. Lance, “Eegnet: a compact convolutional neural network for eeg-based brain–computer interfaces,” Journal of neural engineering, vol. 15, no. 5, p. 056013, 2018. [25] S. Chevallier, I. Carrara, B. Aristimunha, P. Guetschel, S. Sedlar, B. Lopes, S. Velut, S. Khazem, and T. Moreau, “The largest eeg-based bci reproducibility study for open science: the moabb benchmark,” arXiv preprint arXiv:2404.15319, 2024. [26] L. Prechelt, “Early stopping—but when?” in Neural Networks: Tricks of the trade. Springer, 1998, pp. 55–69. [27] X. Li, M. Jiang, X. Zhang, M. Kamp, and Q. Dou, “Fedbn: Federated learning on non-iid features via local batch normalization,” arXiv preprint arXiv:2102.07623, 2021.